Overview
Today, we're releasing DAI-ASR-I18N, David AI's multilingual benchmark for conversational speech recognition and speaker diarization. DAI-ASR-I18N was built from a dataset of 147 hours of unscripted conversations collected across 21 languages.
We evaluated eight commercial systems and six open-weight models on tasks they natively support. We showcase a detailed benchmarking analysis on the public evaluation dataset and report model performance on the public, private, and combined datasets. We also released our inference and evaluation harness with an improved I18N normalization pipeline. DAI-ASR-I18N gives model developers and the research community a shared baseline for how systems perform on natural conversation in each language.
What’s in the Benchmark
DAI-ASR-I18N is built on audio data collected following a common collection protocol. Native speakers were paired on our web-based audio collection platform and given open-ended topics in roughly 15-minute sessions, producing natural speech with interruptions and backchannels. Each speaker was recorded on a separate synchronized channel, providing benchmark ground-truth speaker attribution and allowing the same recording to be used for both speech recognition and speaker diarization.
Transcripts reflect the speech verbatim, keeping fillers, repetitions, and false starts, and were produced under shared transcription guidelines and consistent quality standards. Each transcript was reviewed by two human annotators as well as a language-expert reviewer.
For each language, we stratified two hours of data for public evaluation (Hugging Face), and held out five hours as a speaker-disjoint private test set to avoid contaminating the leaderboards. We targeted similar amounts of audio and a similar share of short-form and long-form audio in every language. Languages still differ in overlapping speech, speaker variety, and recording conditions (see Limitations). The resulting distribution is below.
| AR | BN | DE | EN | ES | FR | HI | ID | IT | JA | KO | MR | PT | RU | TA | TE | TH | TL | TR | VI | ZH | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Clips | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Conversations | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Speakers | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Hours | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Overlap (%) | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| Dataset | Speakers | Gender (F / M / Other) | Hours | Overlap (%) |
|---|---|---|---|---|
| Combined evaluation | — | — / — / — | — | — |
We present a few examples to illustrate conversational phenomena such as backchanneling, interruptions, filler words, and pauses that naturally occur in our audio corpus.
Conversational phenomena
For full details, see the Evaluation Methodology section.
Results
We present our analysis on the public evaluation set and hold out the private set to avoid leakage. The leaderboards include results on both datasets.
Key Findings
- Automatic speech recognition (ASR) systems score within a few points of each other in English, but in most other languages the gap between the best and the typical system is several times larger.
- Microsoft AI MAI-Transcribe-2 has the lowest transcription error rate in most of the 21 languages and stays under 10% in more of them than any other system.
- When speaker audio is collapsed into a single mono channel, as in many real recordings, most systems' error rates rise sharply while MAI-Transcribe-2’s error rates remain relatively stable.
- In speaker diarization, every dedicated diarization model we tested is more accurate than every combined transcription-and-diarization system, with NVIDIA Nemotron 3 Diarization leading in nearly all languages.
- Across both speech recognition and diarization tasks, every system makes the most errors when speakers talk over each other.
Speech Recognition
We compare transcript accuracy using normalized word error rate (WER); for some languages (Mandarin, Japanese, Korean, Thai), we use character error rate (CER). Due to differences in how languages mark word boundaries and how much meaning a single word carries, the same error rate can reflect very different amounts of error from one language to the next. We therefore primarily compare ASR systems within each language.
English benchmarks understate how much system choice matters. In the five languages where systems perform most consistently (English, Korean, Mandarin, Spanish, and Russian), the standard deviation of error rates across systems averages 3.7 percentage points. In the five with the widest spread (Bengali, Hindi, Marathi, Tamil, and Telugu), it averages 25.4 points, about seven times larger. For teams serving non-English native speakers, the spread between systems is wider, and widest in Indic languages. This becomes clear only when each language is measured directly.
MAI-Transcribe-2 stays accurate across more languages than any other system. On the public set, MAI-Transcribe-2 has the lowest error rate in 18 of 21 languages and stays under 10% error in 16. Meanwhile, the next-best system, ElevenLabs Scribe v2, achieves the lowest error rate across the remaining 3 languages, including English and achieved sub-10% error in 10 languages. No other system came under 10% in more than 8 languages. Teams that need strong performance in specific languages should check the per-language results, since the best system varies by language.
| Model | Coverage (langs) | 1st-place | Avg rank | Median rank |
|---|---|---|---|---|
| 21 | 18 | 1.19 | 1 | |
| 21 | 3 | 3.24 | 3 | |
| 20 | 0 | 4.38 | 4 | |
| 16 | 0 | 4.38 | 2 | |
| 20 | 0 | 4.90 | 5 | |
| 21 | 0 | 5.29 | 6 | |
| 17 | 0 | 7.19 | 7 | |
| 21 | 0 | 7.90 | 8 | |
| 21 | 0 | 8.57 | 9 | |
| 12 | 0 | 9.57 | 9 |
Table 3. Per-language rankings and coverage per model on public data
Most systems lose accuracy when speaker audio is mixed into a single mono channel, but MAI-Transcribe-2 is comparatively robust. The results above score each speaker’s channel separately, yet many real recordings, such as phone calls, meetings, and field recordings, capture both speakers on a single channel. When we mix the channels and score each system's speaker-attributed transcript using concatenated minimum-permutation WER/CER (cpWER/cpCER), every system's error rate rises, and MAI-Transcribe-2 shows the smallest relative increase. The mixing penalty grows on recordings with more overlapping speech, linking mixed-channel robustness to how well a system handles overlap.
Combined transcription and speaker attribution accuracy on public data
Speaker Diarization
We separately evaluate the subset of systems that natively support speaker diarization and produce speaker-activity timestamps that we can score directly against the DAI-ASR-I18N reference. We compare these systems on diarization error rate (DER), using forced-aligned references with no collar, with additional analysis on conversations split into tiers by overlapping-speech ratio, the share of each recording in which both speakers talk at once (low <10%, medium 10–20%, high ≥20%).
Dedicated diarization models are more accurate than every combined system we tested. NVIDIA Nemotron 3 Diarization, BUT-FIT DiariZen, NVIDIA Streaming Sortformer v2.1, and pyannote Community-1 have pooled DERs of 18.8–25.8%, compared with 32.8–44.3% for systems that produce speaker labels as part of transcription. Their lead holds under harder conditions: the highest-error dedicated model beats the lowest-error combined system by 7 percentage points overall and still edges it out in high-overlap conversations, though this tier is small (15 recordings). For teams that need accurate speaker attribution, these results make a strong case for pairing their transcription system with a dedicated diarization model.
Every system has its highest diarization error on high-overlap recordings. Across model evaluations, about 66% of diarization error comes from false alarms, or predicted speech activity beyond what the forced-aligned reference marks. Missed speech accounts for 25%, while speaker confusion contributes 8%. Because our references mark speech word by word, even brief pauses between words count as silence, so some of these false alarms come from systems treating those pauses as continued speech. Even so, most error comes from detecting when someone is speaking, rather than telling the two speakers apart. From low-overlap to high-overlap recordings, DER rises by 8.9–12.1 percentage points for dedicated diarization models and 6.1–21.8 points for combined systems. Dedicated models retain the lowest error rates under high overlap.
DER: Overall, Low and High Overlapping Speech on public data
DER (%) · lower is better
Low overlap
High overlap
Limitations
DAI-ASR-I18N focuses on two-speaker conversations recorded on separate microphones, which may give diarization systems recording-condition cues that are less available in single-microphone or multi-speaker settings. Diarization results are also sensitive to evaluation choices such as pause-merging thresholds and collar size, so we report our primary configuration explicitly. And while 21 languages have broad coverage, many lower-resource languages remain outside the benchmark's scope.
Within the benchmark, languages also differ in ways beyond the language itself. Some languages contain naturally occurring acoustic quality issues; the share of overlapping speech ranges from 3% to 14.5%, and the number of distinct speakers ranges from 9 in Italian and French to 91 in English. Scoring choices can add to these differences: because we remove unintelligible stretches from the reference before scoring, any words a system outputs there count as errors, which may affect error rates in recordings with more unintelligible speech. Per-language results therefore reflect a model’s performance on each language's recordings and speakers along with the language itself.
Conclusion
DAI-ASR-I18N provides the speech community a shared, multilingual benchmark for conversational ASR and diarization under real-world conditions. As the industry-leading audio data research lab, we are uniquely positioned to build this benchmark and address current gaps in language coverage and representation in natural speech model performance benchmarks.
Our results show that English-only evaluation hides much of what matters: systems that score within a few points of each other in English can be far apart in other languages. MAI-Transcribe-2 stands out for staying accurate across most languages and for holding steady when both speakers' audio is mixed into one channel. In diarization, NVIDIA Nemotron 3 Diarization and other dedicated diarization models outperform every combined system, suggesting that pairing a transcription system with a dedicated diarization model is the strongest approach for speaker attribution. Across both tasks, overlapping speech remains the hardest problem, and the one most likely to affect systems in real deployments.
At David AI, we are committed to pursuing open evaluation and benchmarking to push the frontiers of speech technology.
For inquiries about model evaluation, please reach out to evals@withdavid.ai
For future research collaboration, reach out to research@withdavid.ai
License and Access
For open research and reproducibility, we are releasing the project code and public dataset of DAI-ASR-I18N:
Code. We are releasing the full evaluation harness, including inference adapters, scoring, and the per-language normalizers, under the MIT License at https://github.com/withdavid-ai/dai-asr-i18n.
Data. The public evaluation set (audio, verbatim transcripts, and word-level timestamps, across 21 languages) is released for non-commercial research under Download Upon Agreement (DUA) on Hugging Face. Our data collection is in full compliance with the appropriate regulatory standards.
Commercial use. Commercial use of the dataset is not covered by CC-BY-NC-4.0. If you would like to use the data commercially, contact evals@withdavid.ai to arrange a commercial license.
Private test set. To keep the ASR leaderboard uncontaminated and competitive, the private test set is not distributed. Contact us at evals@withdavid.ai if you would like your system evaluated against it.
Responsible use. The recordings are natural conversations collected from consenting participants for research and benchmarking. Please use the data accordingly and do not attempt to re-identify speakers.
Citation. If you use DAI-ASR-I18N, please add the following citation in your papers and reports:
@techreport{davidai2026dai-asr-i18n,
title = {DAI-ASR-I18N: A Multilingual, Dual-Channel Conversational Benchmark for Speech Recognition and Speaker Diarization},
author = {{David AI}},
institution = {David AI},
year = {2026},
url = {https://research.withdavid.ai/},
}References
- [1]H. Bredin, “pyannote.metrics: A toolkit for reproducible evaluation, diagnostic, and error analysis of speaker diarization systems,” in Proc. Interspeech, 2017, pp. 3587–3591, doi: 10.21437/Interspeech.2017-411.
- [2]J. Han, F. Landini, J. Rohdin, A. Silnova, M. Diez, and L. Burget, “Leveraging self-supervised learning for speaker diarization,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2025.
- [3]J. Han, P. Pálka, M. Delcroix, F. Landini, J. Rohdin, J. Černocký, and L. Burget, “Efficient and generalizable speaker diarization via structured pruning of self-supervised models,” arXiv preprint arXiv:2506.18623, 2025. https://arxiv.org/abs/2506.18623
- [4]S. Horiguchi, N. Tawara, T. Ashihara, A. Ando, and M. Delcroix, “Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?,” in Proc. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025, arXiv:2507.09226. https://arxiv.org/abs/2507.09226
- [5]M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,” in Proc. Interspeech, 2017, pp. 498–502, doi: 10.21437/Interspeech.2017-1386.
- [6]I. Medennikov et al., “Streaming Sortformer: Speaker cache-based online speaker diarization with arrival-time ordering,” in Proc. Interspeech, 2025, pp. 5238–5242, doi: 10.21437/Interspeech.2025-2244.
- [7]Mistral AI, “Voxtral,” arXiv preprint arXiv:2507.13264, 2025. https://arxiv.org/abs/2507.13264
- [8]T. Park et al., “Sortformer: A novel approach for permutation-resolved speaker supervision in speech-to-text systems,” in Proc. 42nd Int. Conf. Machine Learning (ICML), vol. 267, 2025, pp. 48153–48169. https://proceedings.mlr.press/v267/park25h.html
- [9]A. Plaquet and H. Bredin, “Powerset multi-class cross entropy loss for neural speaker diarization,” in Proc. Interspeech, 2023, pp. 3222–3226, doi: 10.21437/Interspeech.2023-205.
- [10]A. Polok et al., “DiCoW: Diarization-conditioned Whisper for target speaker automatic speech recognition,” arXiv preprint arXiv:2501.00114, 2024. https://arxiv.org/abs/2501.00114
- [11]V. Pratap et al., “Scaling speech technology to 1,000+ languages,” Journal of Machine Learning Research, vol. 25, no. 97, pp. 1–52, 2024. https://jmlr.org/papers/v25/23-1318.html
- [12]Qwen Team, “Qwen3-ASR technical report,” arXiv preprint arXiv:2601.21337, 2026. https://arxiv.org/abs/2601.21337
- [13]A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proc. 40th Int. Conf. Machine Learning (ICML), vol. 202, 2023, pp. 28492–28518. https://proceedings.mlr.press/v202/radford23a.html
- [14]T. von Neumann, C. Boeddeker, M. Delcroix, and R. Haeb-Umbach, “MeetEval: A toolkit for computation of word error rates for meeting transcription systems,” in Proc. 7th Int. Workshop Speech Processing in Everyday Environments (CHiME), 2023, pp. 27–32, doi: 10.21437/CHiME.2023-6.
- [15]S. Watanabe et al., “CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,” arXiv preprint arXiv:2004.09249, 2020. https://arxiv.org/abs/2004.09249
Appendix
Inference Config
| Provider | System / Model | Access | ASR | Diarization | Endpoint or host | API model/checkpoint | Key request params |
|---|---|---|---|---|---|---|---|
| BUT-FIT | DiariZen | Open | — | Yes | Self-hosted | diarizen-wavlm-large-s80-md-v2 | Published DiariZen pipeline configuration with VBx clustering; no language hint or supplied speaker count. Model weights licensed under CC BY-NC 4.0. |
| ElevenLabs | Scribe v2 | Proprietary | Yes | Yes | https://api.elevenlabs.io/v1/speech-to-text | model_id = scribe_v2 | language_code set explicitly; diarize=true, timestamps_granularity=word, no_verbatim=false, tag_audio_events=false, use_multi_channel=false, temperature=0, seed=0. Audio supplied through a presigned URL; client timeout 1,800 s. |
| Google DeepMind | Gemini 3.5 Transcribe | Proprietary | Yes | Yes | generativelanguage.googleapis.com/v1beta/interactions | gemini-3.5-transcribe | Verbatim transcription with speaker diarization, word-level timestamps, and an explicit BCP-47 language hint. Audio uploaded through the Files API. |
| Meta | Muse Voice Transcribe 1.0 | Proprietary | Yes | Yes | api.meta.ai/v1/asr/transcribe | muse-voice-transcribe-1.0 | File transcription in DIARIZATION mode, using 24 kHz PCM WAV and an explicit language-name hint. Recordings exceeding the 600 s request limit are split into shorter windows; speaker embeddings link identities across mixed-audio windows. |
| Microsoft AI | MAI-Transcribe-2 | Proprietary | Yes | Yes | Azure Speech | MAI-Transcribe-2 | Enhanced MAI-Transcribe-2 mode with verbatim transcription, word-level timestamps, diarization enabled, and explicit language locales; Azure Speech API version 2025-10-15. |
| Mistral | Voxtral Mini Transcribe 2 | Proprietary* | Yes | Yes | api.mistral.ai/v1/audio/transcriptions | voxtral-mini-2602 | diarize=true, timestamp_granularities=["segment"], temperature=0, stream=false. Automatic language detection; no explicit language hint is sent. |
| NVIDIA | Nemotron 3 Diarization | Open | — | Yes | Self-hosted | nvidia/Nemotron-3-Diarization | Released checkpoint with NeMo’s offline profile: speaker cache 264, FIFO 40, chunk length 340, right context 40, and cache update period 300 frames. No language hint or supplied speaker count. |
| NVIDIA | Streaming Sortformer 4spk v2.1 | Open | — | Yes | Self-hosted | nvidia/diar_streaming_sortformer_4spk-v2.1 (NeMo) | NeMo offline profile: speaker cache 188, FIFO 40, chunk length 340, right context 40, and cache update period 300 frames. No language hint or supplied speaker count. |
| OpenAI | GPT-Transcribe | Proprietary | Yes | — | api.openai.com/v1/audio/transcriptions | gpt-transcribe | response_format=json with explicit languages[]; audio uploaded as 16 kHz, 128 kbps MP3. |
| OpenAI | Whisper large v3 | Open | Yes | — | Self-hosted | large-v3 (Whisper CT2 + pyannote align) | Long-form inference using faster-whisper/CTranslate2 on CUDA in FP16, with beam_size=5, an explicit language hint, and word-level timestamps. |
| OpenAI | gpt-4o-transcribe-diarize | Proprietary | Yes | Yes | api.openai.com/v1/audio/transcriptions | gpt-4o-transcribe-diarize | response_format=diarized_json, chunking_strategy=auto, and an explicit language hint. Audio uploaded as 16 kHz, 128 kbps MP3; native segment-level timestamps retained. |
| pyannote | pyannote Community-1 | Open | — | Yes | Self-hosted | pyannote/speaker-diarization-community-1 | HF-gated pipeline with default settings and no supplied speaker count. Overlapping speaker turns are retained. |
| Qwen (Alibaba) | Qwen3-ASR-1.7B | Open | Yes | — | Self-hosted | qwen3-asr-1.7b (qwen-asr 0.0.4) | Pinned vLLM runtime with qwen-asr==0.0.4, automatic language detection, and greedy decoding. Audio processed in windows of at most 120 s; outputs reaching the 4,096-token limit are recorded as failures. |
| SpaceXAI | Grok Voice Transcribe 2.0 | Proprietary | Yes | Yes | api.x.ai/v1/stt | xai-stt, default model | Explicitly pinned model=grok-voice-transcribe-2.0, with diarize=true, format=false, filler_words=true, and vad_threshold=0.5. |
Normalization Details
For each transcript (both ground truths and candidates), we apply the following text normalization pipeline:
| Step | Action | Details |
|---|---|---|
| 1 | Decode formatting escapes | Literal \u00A0, \u200C, \u200D, \u3000 are turned into real characters (a closed 4-item allowlist, so arbitrary escapes aren't reinterpreted). |
| 2 | Lowercase | Convert text to lowercase. |
| 3 | Remove annotation tags and parenthesized content | [...], <...>, (...) spans (such as unintelligible annotations, emotion tags) are removed. |
| 4 | I18N normalization | Customized normalization components are implemented per language (described below). |
| 5 | Transform punctuation and symbols to spaces | All Unicode punctuation (P*) and symbol (S*) categories are replaced with a space, while combining marks are preserved (so diacritics and tone marks survive) with some exceptions for Thai and Turkish. |
| 6 | Collapse whitespace and strip | Remove redundant spaces and trim leading/trailing whitespace. |
For baseline normalization, we referenced Whisper’s English normalizer for English and its basic normalizer for other languages.
DAI-ASR-I18N adds language-specific normalizers for accurate and fair metrics. The table lists their behavior. Implementation details are in our GitHub repository.
| Languages | Codes | Normalizer | Main behavior | Metric |
|---|---|---|---|---|
| English | en | whisper_english | Use Whisper English normalization for numbers, contractions and spelling | WER |
| Spanish, French, German, Italian, Portuguese, Russian, Indonesian, Tagalog, Vietnamese | es fr de it pt ru id tl vi | MarkPreservingNormalizer | Apply NFKC, remove all punctuation and symbols; Unicode marks and diacritics are preserved. | WER |
| Turkish | tr | TurkicNormalizer | Apply Turkish-aware casing; preserve dotted versus dotless i. | WER |
| Arabic | ar | ArabicNormalizer | Remove harakat/tatweel; fold alef variants (آ/أ/إ/ٱ) to bare alef (ا); replace alef maqsura (ى, U+0649) with ya (ي, U+064A); preserve the distinction between ta marbuta (ة) and ha (ه). | WER |
| Hindi, Marathi, Bengali, Tamil, Telugu | hi mr bn ta te | ConservativeIndicNormalizer | Apply script-specific NFC and conservative Indic cleanup; preserve vowel marks, nukta and virama | WER |
| Japanese, Korean | ja ko | CJKNormalizer | Apply NFKC and punctuation cleanup. Japanese: remove spaces between kanji/kana. Korean: preserve Korean spaces for formatting consistency but exclude spaces in CER calculation | CER |
| Mandarin | zh | SimplifiedChineseNormalizer | Similar to the CJKNormalizer; for Mandarin we also convert traditional characters to simplified characters with OpenCC | CER |
| Thai | th | ThaiNormalizer | Remove punctuation, preserving Thai vowel and tone marks (combining marks), ๆ (mai yamok, repetition mark), Thai digits ๐–๙. | CER |
Forced-Alignment pipeline for Word-level timestamps
Our internal experiments found that no single forced aligner performs best across all languages. For each language, we used the aligner that produced the most accurate word-level timestamps in our tests.
We checked the following conditions across every run:
- The original annotated references reproduced our diarization error rates.
- The word-merge threshold could increase the amount of reference speech, but could not reduce it.
- Missed speech, false alarm, and speaker confusion summed to the reported error rate.
We treat adjacent same-speaker segments as continuous speech when the gap is below 0.20 seconds. This primary threshold follows the reference from multi-speaker ASR corpora[4].
| Aligner | Languages | Clips |
|---|---|---|
| Qwen3-ForcedAligner-0.6B | de, en, es, fr, it, ja, ko, pt, ru, zh | 569 |
| Meta MMS 300M multilingual forced aligner, fairseq implementation | ar, bn, hi, id, mr, ta, te, tl, tr, vi | 562 |
| Montreal Forced Aligner | th | 49 |

