Research
Research Technician · Hayden Lab, Department of Neurosurgery, Baylor College of Medicine · May 2024 to September 2026
The lab records single neurons in the brains of epilepsy patients while they listen and talk. This page is
what I did there: the recordings I ran, the pipelines I built, the papers that work went into, and one
analysis project of my own.
What I did
Linguistics
The language side of the lab’s analyses.
- Did all of the linguistics for the lab’s Nature Human Behaviour paper (356 hippocampal neurons, 10 patients): every word of the podcasts tagged for part of speech, dependency relations, syntactic depth, clause boundaries, position within the clause, and word frequency.
- Extracted the word embeddings those neurons were compared against, from five models: GPT-2, Llama-3 and DeBERTa for words in context, GloVe and Word2Vec for words on their own.
- Built a 57-feature linguistic annotation of every word for my own grammar-and-meaning analysis.
Recording
Getting the data, at the hospital bedside.
- Ran research recording sessions with epilepsy patients at two hospitals, Baylor St. Luke’s and Texas Children’s, working alongside neurosurgeons, epileptologists and nurses.
- Recruited and enrolled patients into studies, under IRB protocols and HIPAA.
- Traced noisy electrode bundles to an electrical fault rather than their location in the brain, and showed which part of the noise re-referencing could and could not remove.
- Wrote the lab’s standard operating procedures for data handling and electrode reconstruction, and trained lab staff on them.
Pipelines
Turning raw recordings into data people can analyse.
- Spike sorting and quality control, built solo: 98% agreement with expert curation, with manual review cut from 28% of the data to 5%. Adopted lab-wide and run daily by other staff.
- Electrode localization: 203 contacts on 19 leads localized within 0.16 mm of the hand-validated result with no manual steps. The lab’s ten-step reconstruction procedure now runs end to end, with a person approving the result before anything is uploaded.
- Speech transcription and word-level alignment for patient recordings, run locally so the audio never left the site.
- Automatic redaction of patient identifiers from clinical log files.
Modelling
Asking what each neuron was tracking.
- Poisson and logistic regression models of single-neuron firing, checked with cross-validation, permutation tests, bootstrap confidence intervals and confound-matched controls.
- Datasets of 435 neurons in the main analysis and 1,008 pooled across 14 patients.
- Embeddings from 26 language models, extracted on a GPU cluster as a job array that finishes in about two minutes.
Papers it went into
I am a co-author on nine papers from the lab. These four are in journals; the other 5 are preprints.
The first is the one I contributed most to.
- Attention is all you need (in the brain): Semantic contextualization in human hippocampus Nature Human Behaviour, 2026 · In press
Equal-contribution second author. I did the linguistics and the language-model embeddings, and helped design the analysis, run it and write the paper. It shows that hippocampal neurons track where a word sits in its clause, and that their response to a word carries a weighted mix of the words before it, weighted much as a language model’s attention would.
- Plasticity and language in the anaesthetized human hippocampus Nature, 2026 · Published · 654, 714–723
- A population code for semantics in human hippocampus Nature Neuroscience, 2026 · Published online 30 Sep 2026
- Shared neural geometries for bilingual semantic representations in human hippocampal neurons Cell, 2026 · Published · 189(16), 5065–5080
All nine, with citations →
Click a figure to open it full size.
Analysis I built for other studies
- Music
- Poisson regression with nested likelihood-ratio tests for a piano-listening study, including pitch encoding and correction for clock drift between recording systems.
- Bilingual listening
- Decoders for grammatical gender and verb conjugation in recordings from people listening to Spanish.
Methods and tools
Poisson and logistic GLMs Cross-validation Permutation tests Bootstrap confidence intervals Confound-matched controls PCA / CCA UMAP Python MATLAB R PyTorch scikit-learn Hugging Face SLURM GPU clusters