Clinical ASR safety benchmark: do speech-to-text systems transcribe the words that can hurt a patient? 400 clips, 86 African English accents, 5 commercial configurations.
# ScribeCheck
A clinical ASR safety and equity benchmark.
Word error rate is the number speech-to-text vendors publish and the number
buyers compare. This measures something else: whether the words that can hurt a
patient survive transcription. Drug names, dosage values, negations. And whether
that holds across African English accents.
The claim under test is that headline WER is the wrong acceptance metric for
clinical dictation. The benchmark either supports that with numbers or it does
not, and both outcomes are reportable.
**Status: all 2,000 transcriptions complete and scored. The labelling sheet has
been rebuilt as 150 individual errors after four defects were found by using it
on real rows, and human labelling is outstanding.** See RESULTS.md
for the numbers and docs/PRD_EVAL_V2.md for what the
rebuild fixed and why.
## What is measured
Five system configurations across four vendors: Whisper large-v3 hosted by Groq,
Deepgram nova-3, Deepgram nova-3-medical, AssemblyAI universal-3-5-pro, and
Gemini 3.5 Flash Lite. Deepgram appears twice because the general-against-medical
delta is a pricing and product finding in itself.
The result, in one line: Whisper large-v3 and Deepgram nova-3 differ by 0.0009
on WER and 10.6 points on drug-name recall, p = 0.0015 paired across the 109
clips that contain a drug. Full table in RESULTS.md.
Whisper runs through Groq rather than OpenAI because every configuration here
sits on a free tier, which keeps the benchmark reproducible by anyone without a
card. Two consequences, both stated again wherever the numbers appear: the model
is large-v3 where OpenAI's `whisper-1` serves large-v2, and the cost and latency
columns for that row describe Groq's serving rather than OpenAI's. The quality
columns belong to the model, the speed and price columns belong to the host.
See `docs/DECISIONS.md` D016.
400 clips, 71.5 minutes, 86 accents, drawn from the AfriSpeech-200 test split
under seed 42 and stratified into three accent tiers …