Hospital Quality Explorer
Look up any of 5,419 U.S. hospitals and see how it compares with its peers on the federal quality measures. Built on public CMS Care Compare data (799k rows across six datasets) in a Postgres star schema, with the benchmarks computed in SQL instead of loaded from benchmark tables. The interactive explorer lets you define a peer group and get a report card, linked comparisons, a map, weighted rankings, and an in-browser SQL console (DuckDB-WASM) over the same tables. Version 2 adds a live question-answering demo: plain-English questions are answered with model-written SQL (validated, run under a read-only login) or quoted CMS documentation, or declined, and the case study covers how it was evaluated and what broke.
PostgreSQL (Supabase) SQL Python ETL D3 DuckDB-WASM Tableau Public FastAPI Azure Container Apps pgvector
Identity Resolution Lab
769 synthetic messy customer records from four source systems, matched and merged into one record per customer, live in the browser. Every match shows which rule fired, and precision and recall are measured against known ground truth and update as you drag the match threshold. At the default threshold it finds 451 customers against a ground truth of 450. Under the hood: normalization, blocking, fuzzy matching (nickname-aware Jaro-Winkler, weighted field scores), union-find clustering, and survivorship.
Entity resolution Data quality Vanilla JS Python data generator
Pronunciation Coach (in-browser)
The core feedback loop of a pronunciation-learning app, running entirely in the browser: a wav2vec2 phoneme model (ONNX/WebAssembly), Viterbi forced alignment, and per-phoneme goodness-of-pronunciation scoring — no server, audio never leaves the device. Python reference implementation and ONNX export in the open repo; companion write-up benchmarks Whisper vs Qwen3-ASR on faint conversational speech.
wav2vec2 ONNX Runtime Web Forced alignment (Viterbi) GOP scoring Vanilla JS
Houston Eats
An interactive web map for discovering Houston restaurants, with location filtering, marker clustering, and a CSV→geocode→JSON data pipeline. A personal project.
React 19 Vite Leaflet Supabase Tailwind CSS
Graduation Name Pronouncer
A prototype for getting names right at graduation ceremonies, especially non-English names. It generates a pronunciation with multilingual grapheme-to-phoneme models and a curated lexicon, and lets students record their own. IPA is the source of truth, and voice data stays self-hosted.
Python Flask espeak-ng Piper / Kokoro TTS scikit-learn