Note · July 2026
Whisper vs Qwen3-ASR on faint conversational speech
13.1% vs 44.1% word error rate out of the box — but more than half of that gap was my own under-tuned configuration, not the model. Both numbers matter.
The setup
I needed to know which open ASR model to trust for hard real-world audio: 74 minutes of faint,
far-field, five-speaker conversational speech recorded by a room microphone under an IRB-approved
research protocol (no audio or transcript content is shared here). The reference transcript was
hand-corrected — 15,061 words after normalization. Scoring used Whisper's own
EnglishTextNormalizer and jiwer, and both models received the identical
16 kHz mono file. Contenders: OpenAI's Whisper large-v3 and Alibaba's
Qwen3-ASR-1.7B, the largest open Qwen3 ASR checkpoint.
Round one: out of the box
Whisper large-v3, with robust long-form settings (condition_on_previous_text=False,
no_speech_threshold=0.9, temperature fallback): 13.1% WER (10.6% CER),
transcribing 94% of the reference words. Qwen3-ASR-1.7B with defaults: 44.1% WER —
and the failure mode was specific. Of the errors, 5,408 were deletions: it produced words
for only 64% of the reference. On faint multi-speaker audio it doesn't hallucinate; it goes quiet.
Round two: tune both sides
A 3× gap between two modern models usually means someone's configuration is unfair, so I rematched
after fixing mine. Three changes recovered most of Qwen's deficit: windowed automatic gain control
to lift the faint segments; forcing language="English" (the full name — the ISO code
raises an error), which stops quiet chunks from being dropped on failed language detection; and
requesting timestamps with Qwen's forced-aligner companion model, which switches processing to
180-second chunks and removes a token cap that silently truncates long files.
Tuned Qwen: 25.2% WER, coverage up from 64% to 85%, deletions cut from 5,408 to 2,240. About 60% of what looked like model failure was recoverable configuration. Whisper still wins — by roughly 12 points instead of 31 — and it got there with no babysitting.
What I take from it
- Test on your own audio. Leaderboard WER on clean benchmarks predicted none of this; the failure mode (systematic deletion of faint speech) only shows up on audio like yours.
- Tune both sides before declaring a winner. My first result overstated the gap by ~19 points. A benchmark you haven't tried to break is marketing, not measurement.
- Turnkey robustness is a real feature. Whisper's margin isn't just accuracy; it's that the accuracy arrived without the failure-mode archaeology Qwen needed.
- The closed
Qwen3-ASR-FlashAPI reportedly beats the open weights, but API-only is a non-starter for protected audio — a constraint that quietly decides many real deployments.
Pipeline details: both models ran on one NVIDIA L40 via SLURM; the harness feeds one shared 16 kHz mono conversion, scores with normalized WER/CER, and diffs error types (substitutions/insertions/deletions) per model. Related: this pipeline's alignment and pronunciation-scoring layer is public at pronunciation-scoring, with a live in-browser demo.









