How well does ORPHEUS understand medical dictation? We tested it
Jan Brederecke, Leo Harmsen
In short
We tested ORPHEUS on two public datasets and found that it makes fewer errors on medical dictation than real-time variants of large general-purpose systems from ElevenLabs and OpenAI. The same holds for common open models such as OpenAI Whisper, Qwen3-ASR or NVIDIA Parakeet. According to these benchmarks, ORPHEUS comes close to other speech recognition models specialised in medicine, such as Symphony by Corti and Scribe v2 Medical by ElevenLabs.
Background
ORPHEUS, our software for medical speech recognition, has now been rolled out across 40+ hospitals and 500+ practices. We develop ORPHEUS as a non-profit subsidiary of the University Medical Center Hamburg-Eppendorf. We also train and test its speech recognition models ourselves and locally.
Naturally, we cannot publish our test data, which is why a verifiable comparison with other speech recognition systems has been difficult so far.
This spring, the Danish company Corti released its medical speech recognition system Symphony [1] together with two datasets (MedTerm [2], MedDictate [3]), on which we have now tested ORPHEUS extensively (as of 28 September 2026). Both datasets cover several languages; since ORPHEUS is built for German dictation, we evaluate only the German part of each.
First dataset: MedTerm
The published version of MedTerm contains synthetic German recordings of clinical notes that cover much of what makes medical dictation hard: rare medical terms ("Malgaigne-Beckenfraktur", "serielle transverse Enteroplastie"), dates, dosages, measurements, plus spoken formatting commands such as "Punkt" (full stop), "Komma" (comma), "neue Zeile" (new line) or "neuer Absatz" (new paragraph). Each recording comes with a formatted reference text and lists of the medical terms and numeric expressions it contains. Figure 1 shows an example of what such a dictation sounds like and how ORPHEUS would transcribe it.
Figure 1: Our own example in the style of MedTerm. On the left, what is spoken, including the dictation commands; on the right, how ORPHEUS would write it (⏎ = line break).
The evaluation method is public too; it comprises several metrics:
- Word error rate (WER): how many errors occur per 100 words of the reference, ignoring case and punctuation.
- Recall and precision on medical terms: how many of the listed medical terms come out exactly right, and whether the system inserts medical terms where none were spoken.
- Formatting: whether dates and numbers appear exactly as written in the reference.
- Dictated punctuation: whether spoken punctuation marks and line breaks are rendered, and whether the system adds punctuation nobody dictated.
Corti reports these values in Table 4 of its paper [1] for Symphony and for four general-purpose systems (OpenAI, ElevenLabs, Whisper, Parakeet) on the 499 German recordings of MedTerm. Only 200 of them are publicly available, however.
Method
Every value we measured was computed with Corti's own openly available evaluation library, bewer [4]. To put the values on the public recordings into context, we also ran three open general-purpose models through the same evaluation on these 200 recordings: Voxtral Mini by Mistral [5], Whisper large-v3 by OpenAI [6] and Qwen3-ASR 1.7B by Alibaba [7]. All three ran in their default configuration, without fine-tuning, extra prompting or other interventions. They are the only comparisons on exactly the same recordings and also serve as a rough plausibility check. They all landed in the range where the paper also places systems of this class. Our Whisper value, 15.9 percent, is close to the 15.6 percent the paper reports for its Whisper variant. This suggests that the public recordings behave similarly to the paper's 499, but it does not prove it.
We measured ORPHEUS itself in two configurations that occur in the clinic, i.e. both with and without its vocabulary feature, in which users can enter difficult terms before dictating. For the measurement with vocabulary, we gave the system the medical terms the dataset itself provides for each recording (up to three terms per recording). This corresponds to the most favourable case, in which someone has adapted ORPHEUS to their own usage by entering, over time, the terms that were not reliably recognised. In the Corti paper's vocabulary experiment, on English data only, each recording received the whole vocabulary of the dataset [1, Tab. 7].
Conventions
The references write dates as "3. Juli 1995" and ordinal numbers up to nine as words ("der vierten Woche", "the fourth week"). ORPHEUS writes "03.07.1995" and "der 4. Woche". Both are correct German, but the word error rate counts such deviations as errors: for a written-out date, two or three at once. For our measurements we therefore report a style-neutral word error rate, which only counts genuine recognition errors: before counting, reference and output are brought to one convention, symmetrically. Its rules follow the categories of the published style guide [1, Sec. 3.2]. The style-neutral word error rate cannot be computed for the systems from the paper, because Corti has not published their output texts. Symphony writes 99.5 percent of the dates and numbers exactly as the references do [1, Tab. 4], so its value without alignment should be close to the style-neutral one. For the other systems, by contrast, the values without alignment also contain such spurious errors and therefore probably come out somewhat worse. For every system we measured ourselves on MedTerm, the alignment accounts for about one point.
The second convention is a gesture. The references render "neuer Absatz" (new paragraph) as a blank line and "neue Zeile" (new line) as a line break. ORPHEUS was trained to treat both commands as one gesture with one line break. The style-neutral values therefore count a break as a break; results for both ways of counting are in Table 1.
Results on MedTerm
Style-neutral, ORPHEUS reaches a word error rate of 4.0 percent without and 1.8 percent with vocabulary, Symphony 2.5 percent according to the paper. On the same recordings, the three open models make (style-neutral) just under three to just under four times as many word errors as ORPHEUS without vocabulary, and six to just over eight times as many as ORPHEUS with vocabulary. The real-time variants of the commercial systems from OpenAI and ElevenLabs reach 13.0 and 16.4 percent without alignment. This is not a strict ranking, because the recordings are not the same. Figure 2 puts the word error rates of all systems side by side.
Without vocabulary, ORPHEUS finds 61 percent of the medical terms, the general-purpose systems 48 to 71, and Symphony 76 percent. With vocabulary, ORPHEUS gets 98 percent of the medical terms exactly right, and precision only drops from 99 to 98 percent. Since the vocabulary here contained only terms that actually occur in the recording, this does not show how a large vocabulary with many unspoken terms behaves. How much Symphony gains from its own vocabulary feature on the German data has not been published. Figure 3 compares recall without and with vocabulary.
Figure 2: Word error rate on MedTerm, German part. For our own measurements the bars show the style-neutral value and the diamonds the value without alignment. For the systems from the paper only the value without alignment exists (hatched), since their transcripts are not available.
Figure 3: Recall of ORPHEUS on medical terms on MedTerm, without and with vocabulary.
For dictated punctuation, ORPHEUS without vocabulary renders 95 percent of the spoken punctuation marks and line breaks (style-neutral, i.e. with the paragraph difference factored out); Symphony reaches 94. ORPHEUS does, however, add punctuation nobody dictated more often (precision 85 versus 91 percent). The general-purpose systems hit at most about half, presumably because they place punctuation at their own discretion instead of listening to the spoken commands.
Table 1: MedTerm, German part. Values in percent, except Rec.
| System | Rec. | Word error rate | Medical terms | Formatting | Punctuation | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| WER | WER° | Recall | Precision | Recall | Recall° | Recall | Recall° | Precision | Precision° | ||
| ORPHEUS, with the annotated terms as vocabulary | 200 | 2.9 [2.5, 3.3] | 1.8 [1.5, 2.1] | 98.1 [96.7, 99.3] | 97.8 | 60.5 | 91.2 | 84.8 | 98.4 | 83.2 | 89.1 |
| ORPHEUS, without vocabulary | 200 | 5.1 [4.6, 5.5] | 4.0 [3.6, 4.3] | 61.3 [57.0, 65.7] | 99.4 | 60.0 | 90.5 | 81.8 | 95.3 | 79.1 | 84.8 |
| Voxtral Mini | 200 | 12.4 [11.5, 13.2] | 11.3 [10.6, 12.1] | 52.9 [48.9, 57.5] | 99.8 | 57.5 | 84.5 | 41.1 | 52.2 | 25.6 | 29.3 |
| Qwen3-ASR 1.7B | 200 | 14.9 [14.2, 15.7] | 13.8 [13.1, 14.5] | 48.3 [44.3, 52.7] | 99.2 | 50.0 | 74.2 | 38.9 | 51.3 | 23.9 | 27.7 |
| Whisper large-v3 | 200 | 15.9 [15.1, 16.8] | 14.7 [13.9, 15.5] | 52.4 [47.9, 56.9] | 99.8 | 44.4 | 74.2 | 36.9 | 49.0 | 22.6 | 26.7 |
| Symphony (Corti) | 499 | 2.5 [2.3, 2.7] | n/c | 75.6 [73.4, 77.9] | n/p | 99.5 | n/c | 94.1 | n/c | 90.8 | n/c |
| OpenAI Realtime (Corti) | 499 | 13.0 [12.6, 13.4] | n/c | 70.6 [68.0, 73.0] | n/p | 84.8 | n/c | 40.3 | n/c | 25.6 | n/c |
| ElevenLabs Realtime (Corti) | 499 | 16.4 [15.9, 17.0] | n/c | 62.9 [60.3, 65.5] | n/p | 84.5 | n/c | 38.8 | n/c | 19.2 | n/c |
| Whisper, real-time variant (Corti) | 499 | 15.6 [15.1, 16.0] | n/c | 57.1 [54.8, 59.7] | n/p | 81.9 | n/c | 33.9 | n/c | 20.1 | n/c |
| Parakeet (Corti) | 499 | 15.9 [15.5, 16.4] | n/c | 51.1 [48.6, 53.7] | n/p | 81.3 | n/c | 32.4 | n/c | 20.0 | n/c |
Second dataset: MedDictate, real voices
Shortly before this post was published, ElevenLabs released the new Scribe v2 Medical, a variant of its speech recognition adapted to clinical audio [8]. ElevenLabs reports results for it on the second public dataset, MedDictate [3]: on the German recordings a word error rate of 7.2 percent, 9.6 for OpenAI's GPT Transcribe, 11.7 for the general Scribe v2 and 13.3 percent for Muse Voice Transcribe 1.0. The MedTerm figures in the same post pool English, French and German; they therefore cannot be compared with Table 1.
The German part of MedDictate consists of nine real dictations totalling 33 minutes. The speakers include Corti staff. Each is between just under two and seven and a half minutes long and contains 35 to 139 medical terms. Structure and evaluation match MedTerm. This removes the main reservation about MedTerm, its synthetic voices. On the other hand, the dataset is small, and our confidence intervals are correspondingly wide. There are no published results for Symphony on the German MedDictate data.
For the best possible comparability we measured ORPHEUS here without vocabulary. Style-neutral, it reaches a word error rate of 6.6 percent, with a confidence interval of 4.3 to 9.1. The 7.2 percent ElevenLabs states for Scribe v2 Medical lies inside this interval, so both systems are in the same range. The figures ElevenLabs gives for GPT Transcribe, the general Scribe v2 and Muse Voice Transcribe 1.0 lie above our interval. An important caveat applies here as well: ElevenLabs has not published how the texts were normalised before counting, and we did not measure these systems ourselves. We therefore do not know whether both numbers rest on exactly the same procedure. Figure 4 puts the values side by side.
Of the annotated medical terms, 84 percent come out exactly right here, and precision is above 99 percent. Style-neutral, ORPHEUS renders 95 percent of the dictated punctuation.
Figure 4: Word error rate on MedDictate, German part. For ORPHEUS the bar shows the style-neutral value and the diamond the value without alignment. The other values come from ElevenLabs (hatched); how they were normalised has not been published.
Table 2: MedDictate, German part. Values in percent, except Rec.; our row computed with bewer [4], the others from ElevenLabs' publication [8].
| System | Rec. | WER | WER° | Terms, recall | Precision | Formatting, recall | Recall° | Punctuation, recall | Recall° | Precision |
|---|---|---|---|---|---|---|---|---|---|---|
| ORPHEUS, without vocabulary | 9 | 6.9 [4.4, 9.4] | 6.6 [4.3, 9.1] | 83.8 [78.0, 88.8] | 99.6 | 86.7 | 88.0 | 89.3 | 95.0 | 97.0 |
| Scribe v2 Medical (ElevenLabs, own figure) | 9 | 7.2 | n/c | n/p | n/p | n/p | n/c | n/p | n/c | n/p |
| GPT Transcribe (OpenAI, figure by ElevenLabs) | 9 | 9.6 | n/c | n/p | n/p | n/p | n/c | n/p | n/c | n/p |
| Scribe v2 (ElevenLabs, own figure) | 9 | 11.7 | n/c | n/p | n/p | n/p | n/c | n/p | n/c | n/p |
| Muse Voice Transcribe 1.0 (figure by ElevenLabs) | 9 | 13.3 | n/c | n/p | n/p | n/p | n/c | n/p | n/c | n/p |
How to read Tables 1 and 2
- Rec.: number of evaluated recordings. Rows marked "(Corti)" are taken from the paper [1, Tab. 4].
- WER: word error rate, i.e. substitutions, deletions and insertions per 100 reference words, ignoring case and punctuation. Lower is better.
- Medical terms, recall and precision: share of annotated medical terms that come out exactly right. A term only counts if every one of its words is correct. Precision shows whether a system inserts medical terms where none were spoken; this matters above all for the vocabulary row.
- Formatting: share of annotated dates and numeric expressions exactly as written. The column is only partly comparable: Whisper large-v3 reaches 74 percent style-neutral on the 200 public recordings, while the paper reports 82 for its Whisper variant.
- Punctuation, recall and precision: share of dictated punctuation marks and line breaks that are rendered. The references contain only what was spoken; punctuation a system adds on its own counts against precision.
- °: the same value after bringing both texts to one convention (ordinals, number words, numeric dates, times, measurements, ranges and paragraph to line break), symmetrically, so the alignment cannot repair a wrong value.
- n/c: not computable without transcripts; n/p: not published.
- [a, b]: 95% confidence interval from 1,000 bootstrap samples over the recordings. We computed it ourselves for our measurements; for the rows from the paper it is taken from [1, Tab. 4]. For space reasons it is only shown for word error rate and term recall; for the other columns the table gives the point value.
- Vocabulary: the medical terms (up to three) the benchmark annotates for each recording, given to the system.
Interpretation
First: on MedTerm, ORPHEUS makes markedly fewer word errors than the real-time variants of the commercial systems from OpenAI and ElevenLabs as well as freely available models. In word error rate it comes close to Corti's Symphony, and it reliably renders dictated punctuation and line breaks.
Second: the measurement on real voices suggests that the result holds, although the dataset, with nine recordings, is small. On MedDictate, ORPHEUS reaches a style-neutral word error rate of 6.6 percent, in the same range as the new Scribe v2 Medical.
Third: without vocabulary, ORPHEUS does recognise rare medical terms less reliably than Symphony, but using the vocabulary feature closes this gap: it lowers the error rate and markedly raises term recall, without a notable loss in precision. In everyday use the effect will realistically be smaller, because nobody knows all the difficult terms in advance. The feature does, however, allow targeted adaptation to one's own usage without retraining.
Fourth: apart from ORPHEUS, every system, in the form in which it is compared here, is a programming interface (API) or a freely available model. For a practice or hospital, an API or a model file alone is not yet a dictation system: it still lacks an application at the workstation, the connection to microphones and to the hospital or practice software, user management and support, and, for a model, its own GPU servers and their operation. ORPHEUS, by contrast, is complete software: it writes directly at the cursor position in practically any hospital or practice software, runs on the desktop and as an app, and is operated in a German cloud or in the hospital's own data centre.
Limitations
The comparison with Symphony and the other systems we did not measure ourselves is an approximation: the recordings, the measurement mode and in part the normalisation differ. Only our own measurements on the same recordings are directly comparable. Our MedTerm values come only from the 200 public recordings. Whether they are among the paper's 499 and how they were selected is not documented. We did not measure ORPHEUS in streaming mode; all rows marked "(Corti)" in Table 1, Symphony included, come from real-time (streaming) measurements [1, Tab. 4]. We could compute style-neutral values only for our own measurements; that Symphony's value without alignment is close to the style-neutral one is an assumption. The three open comparison models ran in their default configuration and would likely do better still with targeted fine-tuning. The comparison does not include established dictation systems from other vendors, as used in many hospitals and practices: we know of no results for them on these datasets, and apart from ORPHEUS we measured only freely available models ourselves. The MedTerm recordings are clean synthetic speech, a kind of data that has so far not been part of ORPHEUS's training; MedDictate consists of real voices but is very small, and the comparison figures there come from ElevenLabs itself. Neither dataset says much about noisy hospital wards, more distant microphones or hesitant speakers.
Outlook
We keep working on ORPHEUS and will continue to show new results here. The benchmarks show us clearly where to start next: recognising rare medical terms without vocabulary, and punctuation, above all not adding marks nobody dictated. Our aim remains high-quality, sovereign and transparent medical software and AI for hospitals and practices.
We thank Corti for publishing the data and the evaluation method.
Questions, ideas for collaboration, criticism: hallo@idmedizin.de
References
[1] A. Nix, R. James, L. Borgholt, A. B. Ekner, L. Krumm, J. Severin, D. Engel, L. Maaløe, J. Havtorn. "Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces." arXiv:2605.16545v2 [cs.LG], May 2026. https://doi.org/10.48550/arXiv.2605.16545
[2] Corti. MedTerm, German subset (corti/med-term on Hugging Face, revision 47886572): 200 recordings of synthetic speech, 4.1 hours of audio. Licence CDLA-Permissive-2.0 with Corti's usage addendum restricting the data to evaluation.
[3] Corti. MedDictate, German subset (corti/med-dictate on Hugging Face, revision 2a840558): 9 recordings of real dictation, 33 minutes of audio. Licence as [2].
[4] Corti. bewer, Evaluation and analysis framework for automatic speech recognition in Python, version 0.1.0a17, MIT licence. https://github.com/corticph/bewer
[5] A. H. Liu, A. Ehrenberg, A. Lo, C. Denoix, C. Barreau et al. (Mistral AI). "Voxtral." arXiv:2507.13264v1 [cs.SD], July 2025. Model measured: mistralai/Voxtral-Mini-3B-2507, licence Apache-2.0.
[6] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever. "Robust Speech Recognition via Large-Scale Weak Supervision." arXiv:2212.04356v1 [eess.AS], December 2022. Model measured: openai/whisper-large-v3, licence Apache-2.0.
[7] X. Shi, X. Wang, Z. Guo, Y. Wang, P. Zhang et al. (Qwen Team, Alibaba). "Qwen3-ASR Technical Report." arXiv:2601.21337v2 [cs.CL], January 2026. Model measured: Qwen/Qwen3-ASR-1.7B, licence Apache-2.0.
[8] ElevenLabs. "Scribe v2 Medical is now available to everyone." Blog post, 22 September 2026, last updated 27 September 2026. https://elevenlabs.io/blog/scribe-v2-medical-generally-available
Download the figures
All four figures in English and German, as PNG (200 dpi) or SVG (vector). The browser asks before saving.
| Figure 1: dictation example | English · PNG English · SVG German · PNG German · SVG |
|---|---|
| Figure 2: word error rate on MedTerm | English · PNG English · SVG German · PNG German · SVG |
| Figure 3: vocabulary feature | English · PNG English · SVG German · PNG German · SVG |
| Figure 4: word error rate on MedDictate | English · PNG English · SVG German · PNG German · SVG |
About the authors

Leo Harmsen
Team Lead Development, IDM gGmbH
Softwareentwickler und Team Lead Development im ORPHEUS-Team der IDM.
LinkedIn →