Demo · runs entirely in your browser
Pick a sentence, read it aloud, and get a score for every phoneme — the core feedback loop of a pronunciation coach, built from forced alignment and goodness-of-pronunciation scoring.
Three text-to-speech recordings, scored by the exact same pipeline (no download needed). The second one reads the wrong sentence on purpose — the score collapses, which is the point: the system measures what you said, not whether you said something.
1. Listen. A wav2vec2 acoustic model (charsiu/en_w2v2_fc_10ms) turns your audio into a probability over 40 English phonemes every 10 milliseconds.
2. Align. A Viterbi forced alignment matches those frames against the sentence's canonical phoneme sequence (with optional silences between words), so every phoneme owns a time span — the same idea behind the Montreal Forced Aligner.
3. Score. Each phoneme gets a goodness-of-pronunciation (GOP) score: how strongly the model believes you produced the expected sound, versus the best competing sound, averaged over its frames.
Honestly: this is a demo, not a product. Scores are model posteriors, not a validated assessment; the model is English-only and fairest to North American accents; very noisy rooms and cut-off recordings will score poorly for the wrong reasons. The Python reference implementation, the ONNX export, and this page's source are all in the repository.