I PRON yo turn V convierto messy data NP datos desordenados into working tools. PP en herramientas que funcionan
From 2024 to 2026 I was a research technician, and in practice the data person, in a neurosurgery lab at Baylor College of Medicine. We recorded single neurons in the brains of epilepsy patients while they listened and talked, and I built what turned those recordings into results: the sorting and quality-control pipeline the lab still runs, the speech transcription and alignment, and the statistical models. I also wrote the step-by-step guides, trained the staff who run those tools now, and was first-line technical support for recording sessions at two hospitals.
Now I build the same kind of thing outside the lab. A pronunciation coach whose speech model runs entirely in your browser. A record matcher that shows which rule fired on every match. AI agent workflows that carry a ten-step procedure from start to finish, with a person approving the result before anything ships.
B.A. Cognitive Sciences, Rice University, 2025
What I do
Data
Pipelines that clean, match and check messy records.
- Spike-sorting and quality-control pipeline: 98% agreement with expert curation, manual review cut from 28% of the data to 5%, adopted lab-wide.
- Record matching at 98.9% precision and 90.2% recall against known ground truth, with the rule that fired shown for every match.
- 799K rows of federal hospital data behind an explorer with an in-browser SQL console.
Language
Speech and language models, measured on real audio.
- A pronunciation coach whose speech model, cut from 380 MB to 123 MB, runs entirely in the browser.
- A speech-recognition benchmark on faint conversational audio: 13.1% word error rate for Whisper against 44.1% for Qwen3-ASR (25.2% after tuning).
- 26 language models compared on a GPU cluster, in a job that finishes in about two minutes.
AI automation
Agent workflows with checks built in and a person at the gate.
- An agent workflow that runs a ten-step brain-imaging procedure end to end: 203 electrode contacts on 19 leads localized within 0.16 mm of the hand-checked result.
- A research engine where every verdict has to survive two agents trying to refute it.
- Planner, worker and reviewer agents with caps, deny lists and an escalation path, used daily.
Try the work
All projects →Pronunciation Coach (in-browser)
The core feedback loop of a pronunciation-learning app, running entirely in the browser: a wav2vec2 phoneme model (ONNX/WebAssembly), Viterbi forced alignment, and per-phoneme goodness-of-pronunciation scoring — no server, audio never leaves the device. Python reference implementation and ONNX export in the open repo; companion write-up benchmarks Whisper vs Qwen3-ASR on faint conversational speech.
Live demo →
Hospital Quality Explorer
Look up any of 5,419 U.S. hospitals and see how it compares with its peers on the federal quality measures. Built on public CMS Care Compare data (799k rows across six datasets) in a Postgres star schema, with the benchmarks computed in SQL instead of loaded from benchmark tables. The interactive explorer lets you define a peer group and get a report card, linked comparisons, a map, weighted rankings, and an in-browser SQL console (DuckDB-WASM) over the same tables. Version 2 adds a live question-answering demo: plain-English questions are answered with model-written SQL (validated, run under a read-only login) or quoted CMS documentation, or declined, and the case study covers how it was evaluated and what broke.
Ask the data (live) →
Identity Resolution Lab
769 synthetic messy customer records from four source systems, matched and merged into one record per customer, live in the browser. Every match shows which rule fired, and precision and recall are measured against known ground truth and update as you drag the match threshold. At the default threshold it finds 451 customers against a ground truth of 450. Under the hood: normalization, blocking, fuzzy matching (nickname-aware Jaro-Winkler, weighted field scores), union-find clustering, and survivorship.
Live demo →
Autonomous Research Engine
An unattended multi-agent LLM system that harvests literature claims, investigates them, and subjects every verdict to adversarial refutation — two refuter agents with different attack lenses; majority refutation kills; unfalsifiable claims are parked, not answered. Headless Claude Code + Task Scheduler + a plain-markdown ledger; no servers, no database.
Case study →
Where it comes from
The lab work was research on how the human brain represents language: we recorded single neurons while people listened and talked, then asked what each one was tracking. It produced nine co-authored papers: three published in Nature, Nature Neuroscience and Cell, and a fourth in press at Nature Human Behaviour.
1,008
single neurons, recorded one at a time, in 14 patients.
81,370
words of live conversation, plus a 7,346-word podcast, each word timed against the spikes.
2 of 3
brain regions tracked grammar more strongly than meaning: the hippocampus and anterior cingulate, but not orbitofrontal cortex. The same answer held for a podcast and for live conversation.
26
language models compared with the brain, from 0.1 to 32.6 billion parameters. Bigger models matched the hippocampus no better; the best match was GPT-2 medium, one of the smallest.
From analyses for a manuscript that is still in preparation. Several earlier results did not survive stricter controls and were dropped; these are the ones that did.
- Plasticity and language in the anaesthetized human hippocampus Nature, 2026
- Attention is all you need (in the brain): Semantic contextualization in human hippocampus Nature Human Behaviour, 2026
- A population code for semantics in human hippocampus Nature Neuroscience, 2026
- Shared neural geometries for bilingual semantic representations in human hippocampal neurons Cell, 2026
Languages I speak
English, Spanish (professional) · French, Portuguese, Japanese (intermediate)
Plus a habit of learning writing systems, which is why my name keeps changing script up top. Press play on any of them.





































