Logo Lanfrica

suwaidakhan/ScribeCheck

Domain:

healthcarenatural language processing

Record type:

dataset
Creator:
suw
Host:
Clinical ASR safety benchmark: do speech-to-text systems transcribe the words that can hurt a patient? 400 clips, 86 African English accents, 5 commercial configurations. # ScribeCheck A clinical ASR safety and equity benchmark. Word error rate is the number speech-to-text vendors publish and the number buyers compare. This measures something else: whether the words that can hurt a patient survive transcription. Drug names, dosage values, negations. And whether that holds across African English accents. The claim under test is that headline WER is the wrong acceptance metric for clinical dictation. The benchmark either supports that with numbers or it does not, and both outcomes are reportable. **Status: all 2,000 transcriptions complete and scored. The labelling sheet has been rebuilt as 150 individual errors after four defects were found by using it on real rows, and human labelling is outstanding.** See RESULTS.md for the numbers and docs/PRD_EVAL_V2.md for what the rebuild fixed and why. ## What is measured Five system configurations across four vendors: Whisper large-v3 hosted by Groq, Deepgram nova-3, Deepgram nova-3-medical, AssemblyAI universal-3-5-pro, and Gemini 3.5 Flash Lite. Deepgram appears twice because the general-against-medical delta is a pricing and product finding in itself. The result, in one line: Whisper large-v3 and Deepgram nova-3 differ by 0.0009 on WER and 10.6 points on drug-name recall, p = 0.0015 paired across the 109 clips that contain a drug. Full table in RESULTS.md. Whisper runs through Groq rather than OpenAI because every configuration here sits on a free tier, which keeps the benchmark reproducible by anyone without a card. Two consequences, both stated again wherever the numbers appear: the model is large-v3 where OpenAI's `whisper-1` serves large-v2, and the cost and latency columns for that row describe Groq's serving rather than OpenAI's. The quality columns belong to the model, the speed and price columns belong to the host. See `docs/DECISIONS.md` D016. 400 clips, 71.5 minutes, 86 accents, drawn from the AfriSpeech-200 test split under seed 42 and stratified into three accent tiers …