research

ORPHEUS achieves top-class results on two public medical benchmarks

Philipp Mühl

Product & Policy Strategy, IDM gGmbH

September 30, 2026
ORPHEUS achieves top-class results on two public medical benchmarks

Measurements as of 28 September 2026

out of 100 words correct
96
MedTerm, without vocabulary
out of 100 words correct
98+
MedTerm, with vocabulary
out of 100 words correct
93+
MedDictate, real voices
of spoken commands carried out
95%
both benchmarks

In everyday clinical practice, medical speech recognition has one job above all: to turn spoken language, complex medical terms, numbers and dosages reliably into usable text. On MedTerm, a public benchmark that recreates exactly these demands with 200 synthetic clinical dictations, ORPHEUS recognises 96 out of 100 dictated words correctly. If the most difficult medical terms are entered in the vocabulary beforehand, this share rises to more than 98 out of 100 words; ORPHEUS then also writes 98 percent of the entered terms themselves exactly right.

Recognition stays high with real voices too, and ORPHEUS reliably carries out spoken commands. On MedDictate, a second benchmark with real instead of synthetic voices, which are more demanding because of changing pace, individual pronunciation and pauses, ORPHEUS recognises more than 93 out of 100 words correctly. ORPHEUS carries out 95 percent of commands such as “Punkt” (full stop) or “neue Zeile” (new line), which doctors use to structure their text while dictating, on both benchmarks.

The results show that ORPHEUS combines a low word error rate with reliable recognition of medical terms and spoken formatting commands. Spoken language thus becomes not just a transcript but directly usable medical text, with few corrections, structured reports and the option of adapting the recognition specifically to the terminology of an institution.

Both benchmarks were published by the Danish medical AI company Corti, and we measured with its open evaluation tool [3]. Readers who want to go deeper will find all the details in the technical report by Dr Jan Brederecke, who is responsible for ORPHEUS model training at IDM, and Leo Harmsen, the lead of the ORPHEUS team: how the datasets are built and the measurement method, the key metrics with their statistical uncertainty, the full comparison values for other systems, and a detailed discussion of the counting conventions and the limits of the measurement.

Technical report

How well does ORPHEUS understand medical dictation? We tested it

Jan Brederecke, Leo Harmsen

How we measured ORPHEUS on the public datasets MedTerm and MedDictate: method, counting conventions, all metrics with confidence intervals, comparison values and limitations.

Read the technical report

Bar chart: word error rate on the MedTerm benchmark, German part. ORPHEUS with vocabulary 1.8 percent and without vocabulary 4.0 percent, alongside the values of other speech recognition systems.

Figure 1: Word error rate on the public MedTerm benchmark, German part; lower is better. The chart shows ORPHEUS alongside the systems that Corti evaluates in its paper on the benchmark [4], and alongside three open models we measured on the same recordings. The technical report explains the counting method and its limitations.

Even without stored vocabulary, ORPHEUS masters the demanding MedTerm test with synthetic voices

MedTerm deliberately puts the difficult parts of medical dictation to the test, and ORPHEUS recognises 96 out of 100 words correctly with its standard vocabulary alone. For everyday use, this means the report is almost finished in the document once the dictation is done. Doctors proofread and adjust individual words instead of rewriting passages. The 200 dictations are synthetically spoken, last 4.1 hours in total and contain rare medical terms such as “Malgaigne-Beckenfraktur”, dates, dosages and measurements as well as spoken commands [1]. The measure is the word error rate, the standard metric for speech recognition, which counts wrong, missing and extra words and is shown in Figure 1.

Using the ORPHEUS vocabulary feature improves the result further, especially for rare medical terms

With the vocabulary feature, recognition across all words rises from 96 to more than 98 out of 100. In daily practice, ORPHEUS users can store specific terms for recognition, such as rare diagnoses, medications, proper names or in-house designations. This way ORPHEUS adapts to a department’s terminology without retraining. To reflect this practical use, we entered the rare medical terms that MedTerm specifies for each dictation and tests separately, up to three per dictation. These few entries alone more than halved the number of misrecognised words.

The vocabulary feature also raises the hit rate for rare medical terms, one of the ultimate tests in the MedTerm benchmark, from 61 to 98 percent. For this, MedTerm contains, in addition to the measurement across all words, a separate subtest that checks only the specified rare medical terms, such as “Malgaigne-Beckenfraktur” or “serielle transverse Enteroplastie”. Counting is strict: a term only counts as a hit if it appears in the text completely and exactly right; if even one of the three words of “serielle transverse Enteroplastie” is wrong, the whole term counts as missed.

ORPHEUS maintains its high level with real voices too

On MedDictate, which uses real instead of synthetic voices, ORPHEUS recognises more than 93 out of 100 words correctly without vocabulary. Real voices bring more variation in pace, pronunciation and manner of speaking than synthetic ones. ORPHEUS is built for exactly these conditions: we continuously test and develop the system further in live operation. MedDictate consists of nine real dictations with a total of 33 minutes of audio [2].

For the German MedDictate dictations, this is among the best results published so far. The technical report explains how the statistical uncertainty of nine dictations affects this assessment. ORPHEUS also achieves strong results on the medical terms that MedDictate tests separately: even without vocabulary, 84 percent come out exactly right.

Besides speech, ORPHEUS also reliably carries out spoken formatting commands

ORPHEUS correctly carries out 95 percent of the commands for punctuation marks and line breaks. In addition to word recognition, both benchmarks separately measure whether such spoken commands are carried out. If a doctor says “Punkt” (full stop), ORPHEUS inserts a full stop. If she says “neue Zeile” (new line), the next text starts on a new line (Figure 2). The dictation thus turns directly into structured text.

Example of a medical dictation: on the left, the spoken text with dictation commands such as full stop, comma and new line; on the right, the formatted text written by ORPHEUS.

Figure 2: Our own example in the style of MedTerm. On the left, what is spoken, including the dictation commands; on the right, the text ORPHEUS writes from it (⏎ = line break).

This result pays off in everyday use because ORPHEUS writes the structured text directly where the report is created. ORPHEUS writes at the cursor position in practically any hospital or practice software; the dictation thus lands in the document already structured and without detours. ORPHEUS runs on the desktop and as an app, can be operated in a German cloud or in the institution’s own data centre, and has now been rolled out across 40+ hospitals and 500+ practices.

We are staying on it: for speech recognition that doctors can rely on every day

Our goal is speech recognition that gives doctors time back for their patients: measurable, sovereign and at home in the language of their specialty. Public benchmarks make this claim verifiable, and we will continue to put ourselves to the test. The next benchmark is everyday life on the ward, with background noise, changing microphones and time pressure, which today’s datasets only begin to capture. We are developing ORPHEUS further for this and will publish new results as soon as they are available.

We thank Corti for publishing the data and the evaluation method.

Questions, ideas for collaboration or criticism: hallo@idmedizin.de

References

Technical report: Dr J. Brederecke, L. Harmsen. “How well does ORPHEUS understand medical dictation? We tested it.” September 2026.

[1] Corti. MedTerm, German subset (corti/med-term on Hugging Face, revision 47886572): 200 recordings of synthetic speech, 4.1 hours of audio. Licence CDLA-Permissive-2.0 with Corti’s usage addendum restricting the data to evaluation.

[2] Corti. MedDictate, German subset (corti/med-dictate on Hugging Face, revision 2a840558): 9 recordings of real dictation, 33 minutes of audio. Licence as [1].

[3] Corti. bewer, Evaluation and analysis framework for automatic speech recognition in Python, version 0.1.0a17, MIT licence. https://github.com/corticph/bewer

[4] A. Nix, R. James, L. Borgholt, A. B. Ekner, L. Krumm, J. Severin, D. Engel, L. Maaløe, J. Havtorn. “Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces.” arXiv:2605.16545v2 [cs.LG], May 2026. https://doi.org/10.48550/arXiv.2605.16545

Download the figures

Both figures in English and German, as PNG (200 dpi) or SVG (vector). The browser asks before saving.

About the author

Philipp Mühl

Philipp Mühl

Product & Policy Strategy, IDM gGmbH

Verantwortet bei der IDM die Gemeinnützigkeit, die Außenkommunikation und den Austausch mit Politik und Gesellschaft.

LinkedIn →