Skip to content

Note · July 2026

Whisper vs Qwen3-ASR on faint conversational speech

13.1% vs 44.1% word error rate out of the box — but more than half of that gap was my own under-tuned configuration, not the model. Both numbers matter.

The setup

I needed to know which open ASR model to trust for hard real-world audio: 74 minutes of faint, far-field, five-speaker conversational speech recorded by a room microphone under an IRB-approved research protocol (no audio or transcript content is shared here). The reference transcript was hand-corrected — 15,061 words after normalization. Scoring used Whisper's own EnglishTextNormalizer and jiwer, and both models received the identical 16 kHz mono file. Contenders: OpenAI's Whisper large-v3 and Alibaba's Qwen3-ASR-1.7B, the largest open Qwen3 ASR checkpoint.

Round one: out of the box

Whisper large-v3, with robust long-form settings (condition_on_previous_text=False, no_speech_threshold=0.9, temperature fallback): 13.1% WER (10.6% CER), transcribing 94% of the reference words. Qwen3-ASR-1.7B with defaults: 44.1% WER — and the failure mode was specific. Of the errors, 5,408 were deletions: it produced words for only 64% of the reference. On faint multi-speaker audio it doesn't hallucinate; it goes quiet.

Round two: tune both sides

A 3× gap between two modern models usually means someone's configuration is unfair, so I rematched after fixing mine. Three changes recovered most of Qwen's deficit: windowed automatic gain control to lift the faint segments; forcing language="English" (the full name — the ISO code raises an error), which stops quiet chunks from being dropped on failed language detection; and requesting timestamps with Qwen's forced-aligner companion model, which switches processing to 180-second chunks and removes a token cap that silently truncates long files.

Tuned Qwen: 25.2% WER, coverage up from 64% to 85%, deletions cut from 5,408 to 2,240. About 60% of what looked like model failure was recoverable configuration. Whisper still wins — by roughly 12 points instead of 31 — and it got there with no babysitting.

What I take from it

Pipeline details: both models ran on one NVIDIA L40 via SLURM; the harness feeds one shared 16 kHz mono conversion, scores with normalized WER/CER, and diffs error types (substitutions/insertions/deletions) per model. Related: this pipeline's alignment and pronunciation-scoring layer is public at pronunciation-scoring, with a live in-browser demo.