MedTerm-90: seven speech-to-text engines, judged the way a clinician would judge them
Appen's MedTerm-90 benchmark scores seven speech-to-text APIs on 90.5 hours of real physician dictation — judged on clinical entity recall, not word error rate.
Medical dictation is where speech recognition earns its keep or causes real harm. A transcript that is “95% accurate” by word error rate can still drop the one medication, lab value, or blood pressure reading that mattered. MedTerm-90 asks a harder question: does the medical terminology survive, in a form a clinician would accept? The benchmark scores seven commercial speech-to-text APIs, brings Google’s Gemini 3.5 Transcribe and Meta’s Muse Transcribe into the field, and introduces the strictest scoring standard we have published.
Real data, from real environments
We ran seven commercial APIs over 1,413 real physician dictations: ElevenLabs Scribe v2, Microsoft MAI-Transcribe 2, OpenAI gpt-transcribe, Meta Muse Transcribe, Amazon Transcribe Medical, Google Gemini 3.5 Transcribe, and xAI Grok STT 1.0. The corpus covers 90.5 hours across 14 specialties. Seventy percent of it is telephony, GSM-compressed recordings from telephone dictation lines, and the rest splits between MP3 and uncompressed WAV from desktop and handheld recorders, complete with background noise, disfluencies, and the occasional passage nobody could transcribe. All of it is dictation audio: one speaker, close to the microphone. That matters, because it means the conditions favor the models. There is no crosstalk or far-field capture to blame, so the errors in this benchmark belong to the engines themselves. This is the audio real transcription pipelines receive, not a curated test set.
Every recording is real data from real clinical environments: actual physicians dictating actual patient encounters, paired with the notes their human scribes produced. Nothing is synthetic, scripted, or studio-read. The corpus comes from Appen’s licensed dataset library, and we used only audio that has never been sold to or acquired by any of the companies tested. No model could have seen this data in training.
Scoring that thinks like a clinician
Results come from the Appen Clinical Entity Scorer, a pipeline judging 30,017 clinician-validated entities. First it normalizes formatting, so “126/62”, “126 over 62” and “one twenty-six over sixty-two” all count as the same blood pressure. Then it handles what normalization cannot: character similarity finds near-misses but cannot tell a formatting difference from a misspelling. So every one of the 12,326 distinct near-miss pairs was reviewed the way a clinician would read it, under one rule: credit any written form clinicians actually use, reject anything they would flag. “Para vertebral” for “paravertebral” passes, and a model that writes “b.i.d.” where the note says “twice daily” loses nothing. “Unglyza” for the diabetes drug “Onglyza” does not pass. Only 21.2% of near-misses survived the review.
Fairness is verified rather than assumed. Identical files, entities, thresholds and judges for every model. No vocabulary hints. A cohort of medical experts reviewed the dataset and spot-checked the final outputs of every model. Decoy testing shows statistically indistinguishable false-match rates across all engines, and every number carries a bootstrap 95% confidence interval.
What we found
ElevenLabs Scribe v2 leads and takes three of seven clinical categories. OpenAI gpt-transcribe and Microsoft MAI-Transcribe 2 are a statistical tie for second with different characters: gpt-transcribe spells medicine cleanly, while MAI-2 pairs near-top recall with 94× real-time speed, the benchmark’s speed-accuracy frontier. Google’s Gemini 3.5 debuts fourth at 80.65%, within the statistical margin of third; it takes diagnosis and lab values outright and needed the fewest corrections of any model tested. Meta’s Muse holds fifth with a balanced category profile, and Amazon keeps the best medication recall in the field (83.0%). Grok remains the fastest engine we have measured, at the lowest recall.
Two patterns hold across the field. Every model bottoms out on dictations of one to three minutes. And on 28.0% of files, no model reached 85% recall.
When a miss is not just a miss
Alongside the benchmark we ran a clinical risk review: every near-miss rendering, all 6,425 of them, classified by consequence on a five-level scale, from fatal (acting on the transcript could directly cause life-threatening harm) through serious, moderate and minor down to none (clinically equivalent). The result is the strongest argument for strict scoring we have. 191 renderings were rated potentially fatal and 1,027 serious. 516 of those are outright drug-name garbles, and 444 were made by a single model; no other engine made the same mistake. All seven engines reduced the inhaler Trelegy to a stray “g y”, erasing the drug name entirely, and all seven wrote hyponatremia as “hypernatremia”, opposite conditions treated in opposite ways. MAI-2 and Grok turned Toradol, an NSAID, into tramadol, an opioid. Three engines wrote “fludrocortisone” for fluticasone, swapping an inhaled steroid for a systemic one. Three others turned Lantus 10 units into 18.
The most telling errors are the ones only a single model made; no other engine made the same mistake on the same audio. MAI-2 alone turned a Cardene drip, a cardiac medication, into a “codeine drip”, an opioid. Scribe v2 alone turned Mucinex into amoxicillin and Zoloft into zolpidem. gpt-transcribe alone turned diclofenac into Xanax and Sinemet into senna. Gemini alone wrote “hema log” for Humalog and swapped calcium chloride for potassium chloride. Muse alone turned tramadol into Toradol. Grok alone reduced hydrochlorothiazide to “hydrochloride” and Florinef to “fluorine”. Amazon alone rendered QVAR as “quavar”. These are not audio problems; they are each model’s private confusions. Sound-alike drug pairs are a known hazard in human transcription, and the models reproduce the same failure at machine speed.
These errors read plausibly. They survive casual review. Character-similarity scoring would have credited many of them. That is why recall percentages, on their own, understate deployment risk, and why every number in this benchmark passed through a judge that reads like a clinician.
Appen’s vantage point
The frontier of AI is no longer starved for data. It is starved for expert data. The web has been consumed; what moves models now is the speech no crawler can reach: a physician dictating a real patient encounter over a telephone line, in the vocabulary of their specialty, with the noise and disfluency of an actual clinic. That is what Appen provides. The MedTerm-90 corpus is real dictations from real doctors, captured in live clinical transcription workflows and licensed for AI development through Appen’s data partnerships, validated by a global network of vetted specialists that includes clinicians.
That access is why this evaluation looks the way it does. Authentic audio instead of synthetic test sets. Expert-validated ground truth held to the same strict standard as the models. Clinically weighted metrics instead of raw word error rate, and provenance controls that guarantee none of the tested models had seen the data before. The corpus is one of many expert domain datasets in Appen’s data library, and the same collection, validation, and evaluation machinery is available to teams who want to benchmark or train models on their own domain, against their own risk profile.
Results reflect September 2026 model versions and the corpus and configurations described in the full report; they may not generalize to other audio or settings. Model names are trademarks of their respective owners. Full methodology, judge policy, per-specialty results, and limitations: see the Appen MedTerm-90 Benchmark report.