Colophon
How this site is built.
The site is part of the portfolio. Here’s what’s under it, and why. The code is on GitHub.
01 / Stack
Next.js with the App Router and TypeScript, styled with Tailwind CSS and deployed on Vercel. Every page except two small API routes is generated at build time, so most of the site is plain HTML on a CDN.
Case studies are MDX files. Each one exports typed metadata that’s validated with Zod at build time, so adding a case study means adding one file, and a missing field fails the build instead of shipping a broken page.
02 / Data pipelines
Each Lab dataset is a small data product. Python pulls raw rows from the source API. A dbt project on DuckDB builds bronze, silver and gold layers, with tests and enforced contracts on the gold tables.
One definition, two runtimes. Each publish exports the SQL dbt compiled for every model and test. The Lab runs that same SQL in DuckDB compiled to WebAssembly in your browser. The models, tests and results match.
- Orchestrated. An Airflow DAG runs extract, load, dbt build, model training and publish. GitHub Actions runs it weekly with
airflow dags test, so there is no server to host. Vercel redeploys on the commit. - Tested. Uniqueness, nulls, accepted values, ranges, referential integrity, freshness and row-count guards. Flames shot-level goals are reconciled against the official final score for every game.
- Contracted. An error-level failure stops the DAG before anything is written, so the last good snapshot keeps serving. Warnings are real issues in the source data, shown rather than hidden.
- Incremental. Live permit runs use the City’s
:updated_atfield as a watermark, so status changes on old permits are picked up along with new applications. Changed rows merge on the primary key.
Read the dbt project, the Airflow DAG, or browse the dbt docs and lineage.
03 / The home value model
The housing tab’s model is gradient-boosted trees in Python, using XGBoost. It trains on 390,000 homes in the weekly Airflow run. The same Python file trains on a sample in your browser through Pyodide.
The trees are exported to JSON, and a small TypeScript scorer explains each estimate. Tests check its predictions match XGBoost’s.
- Honest evaluation. One home in five is held out, chosen by hashing its roll number, so the test set is stable across runs. The model is compared against the obvious baseline (the community median for that property type), not against nothing.
- No leakage. Community, type and zoning are target-encoded out of fold, so a home’s own value never feeds its features.
- Explained. Every estimate is broken into per-feature effects by following the home’s path through each tree, and comes with a range from the holdout error distribution.
- Gated. The model retrains only after every error-level data test passes.
04 / Queries in your browser
The Lab runs DuckDB compiled to WebAssembly. The engine and Parquet files load only when you scroll to a demo, and each table downloads only when a query first needs it. After that, filters are SQL queries that run locally in milliseconds, with no API to wait on and no per-query cost.
Every chart panel has a “Show SQL” toggle. If a number looks wrong, you can check it.
05 / Observability
The Telemetry panel (on the Lab tabs, or ⌘K → Open telemetry) shows what this tab has done: Core Web Vitals measured with the web-vitals library, the DuckDB engine’s start-up time, every table downloaded, every query with its latency and p50 and p95, each analytics event, and every pipeline run.
It’s all in memory in your browser. Nothing in the panel is sent anywhere.
06 / Analytics and the tracking plan
Page views come from Vercel Web Analytics: no cookies, no cross-site tracking, no personal data. Product events follow a tracking plan written as code. Each event is validated against the plan before it’s sent, so a typo or an unplanned property is caught in development and dropped in production. The table below is generated from that plan.
No free text is ever sent. Questions you type into “Ask the data” are tracked only as a length.
| Event | Properties |
|---|---|
lab_tab_viewedA Lab dataset tab is opened. | dataset: permits | housing | flames |
lab_filter_changedA dashboard control changes value. | dataset: permits | housing | flamescontrol: stringvalue: string |
lab_sql_viewedSomeone opens the SQL behind a chart. | dataset: permits | housing | flamespanel: string |
ask_submittedA question is sent to Claude. The question text is not tracked. | dataset: permits | housing | flamessource: typed | examplelength: number |
ask_completedClaude's answer comes back and runs, or fails. | dataset: permits | housing | flamesoutcome: answered | declined | errorlatency_ms: number |
ask_sql_rerunSomeone edits the generated SQL and runs it themselves. | dataset: permits | housing | flames |
pipeline_run_startedA live pipeline run starts in the browser. | dataset: permits | housing | flames |
pipeline_run_completedA live pipeline run finishes. | dataset: permits | housing | flamesoutcome: success | failedrows_in: numbertests_failed: numberduration_ms: number |
model_estimatedThe value model produces an estimate. Inputs other than property type are not tracked. | property_group: stringmodel: production | yours |
model_trainedSomeone trains their own value model in the browser. | rows: numbertrees: numberdepth: numberlearning_rate: numbermedian_error_pct: numberduration_ms: number |
telemetry_openedThe Telemetry panel is opened. | location: string |
command_menu_openedThe ⌘K menu is opened. | none |
contact_submittedThe contact form is sent. The message itself is never tracked. | reason: hiring | project | otheroutcome: sent | error |
cta_clickedA contact or hiring call to action is clicked. | cta: contact | email | resume | linkedin | booking | github | case_studies | lab | role | consulting | frontierlocation: string |
07 / The AI feature and its guardrails
“Ask the data” sends your question to Claude through a small server route. Claude replies in a fixed JSON shape with one SQL query, a short explanation and a chart suggestion. The guardrails, in order:
- Narrow input. Questions are capped at 300 characters, and the dataset is picked from a fixed list.
- Narrow context. Claude sees the question and the table schemas. No keys, no personal data, no tools, no browsing.
- Prompt injection is expected. The question is wrapped and labelled as untrusted visitor input. Off-topic requests get a short “here’s what I can answer” instead of a query.
- Structured output. The response must match a JSON schema, so there’s no free text to parse or render as HTML.
- SQL is checked twice. Once on the server and once in the browser: a single SELECT, no file, network or settings functions, and a hard 200-row cap.
- The blast radius is tiny by design. The query runs in your own browser against public data. The worst case is a slow query in your own tab.
- Rate limits and a spend cap. Per-visitor limits on the route, and a monthly cap on the API account.
The NHL’s API doesn’t allow browser requests, so Flames live runs go through a proxy that allows exactly two URL shapes, caches at the edge and rate-limits each visitor.
08 / Design decisions
A written brand guide in the repo sets every visual and writing rule. Fraunces for headings, Geist for text and Geist Mono for numbers. One deep teal accent marks links, the main action and the key metric. Charts use a separate four-color data palette.
Every color, size, space and radius comes from one tokens file. The styling layer can only produce values from that file. A brand check in CI fails on hardcoded colors, off-scale sizes, em dashes, banned phrases and case studies that skip the format.
Charts are hand-written SVG and canvas, not a charting library. Motion is limited to hover and focus, and turns off when your system asks for reduced motion. Light and dark themes follow your system until you pick one.
Want to see it running? Open the Lab.