EvaluationSeptember 14, 2026

DAI-S2S-ST: Evaluating S2S Model Quality with Human Preference Data

The DAI-S2S-ST Human Preference Leaderboard measures which voice models people prefer in open-ended conversations, not just whether a task was completed.

Comparative ratings
153k
Distinct raters
283
Recorded prompts
819
Rating questions
12
Models evaluated
7

Overview

Today we’re releasing DAI-S2S-ST, our single-turn S2S (speech-to-speech) human preference leaderboard, built on 153k comparative human preference ratings across 7 models, 819 distinct prompts, and 12 questions. For this evaluation, we grade model responses to single-turn, pre-recorded input prompts. Our initial results tell a different story than current S2S leaderboards.

We believe the future of voice AI is one where people spend hours a day interacting with models across apps, devices, and new physical interfaces. For that to happen, models need to be more than capable. They need to be compelling enough that people actually want to keep talking to them.

That is different from how voice assistants have traditionally been used. Most were built for short transactions: set a timer, check the weather, play a song. If the model understands the request and completes the task then the interaction is successful, even if the voice sounds robotic or awkward.

Interacting with models for hours a day raises the bar — and requires us to evaluate models differently. Naturalness, personality, empathy, and the overall quality of the interaction become critical. Most voice benchmarks measure whether the model understood the input, answered correctly, or completed the task. These questions matter, but they don’t tell us whether someone would actually want to keep talking to the model. To measure that, we need to ask humans.

DAI-S2S-ST is a first step toward measuring not just what a model can do, but what it feels like to interact with one:

  1. Human preference across 12 dimensions. Existing human preference leaderboards (e.g., Voice Showdown, now retired, and Speech Agent Arena) typically collapse preference into a single overall judgment. DAI-S2S-ST asks twelve distinct side-by-side questions to understand what actually drives preference.
  2. Focused on consumer use cases. We evaluate eleven categories of consumer-oriented scenarios, where the qualities that drive a good interaction differ from what most task-oriented voice-agent leaderboards cover (e.g., τ³-Voice and Eva Bench).

This initial version focuses on single-turn interactions. Future work will extend this evaluation to longer, multi-turn interactions.

What’s In the Evaluation

For full details on our methodology, see the methodology section.

1. Recording

David AI’s evaluation team created 819 single-turn evaluation scripts across eleven objective categories. These scripts span five different difficulty levels and include scene direction across robustness categories (such as background environment).

User-recorded prompts

Emotion Response

The speaker shares a personal, emotionally charged situation; the model must recognize the feeling and respond with fitting empathy and tone.

Example prompt
0:07
Transcript

“I keep wondering if there’s something actually wrong with me since I never seem to click with anyone. Is it just in my head?”

Our US-based voice contributor network recorded in a variety of environments that simulate real user conversations.

  • Indoor Background NoiseCoffee shop
    0:35
  • Outdoor Background NoiseOutdoor public park
    0:36

2. Evaluation

Each recording was run through inference with each of the relevant models, and then David AI’s paid rater panel completed side-by-side preference judgements against our rubric.

These questions are slightly modified from those in our LALM-as-judge vs HITL post based on rater feedback and analysis of those earlier results. Additional details on the evaluation methodology for this leaderboard can be found in the methodology section.

Comparison interface

Input prompt
0:23
Model A
0:23
Model B
0:57

“Overall, which response do you prefer?”

Comparison interface with sample prompt and model outputs. For one pair of responses, raters answer every question in the rubric below.
Rubric questions
AxisParameterRater question
OverallOverall“Overall, which response do you prefer?”
HumannessEngagement“Which response was more engaging?”
Emotional appropriateness*“Which response’s emotional tone responded more appropriately to the user’s emotional state and the situation?”
Naturalness“Which response’s delivery sounded more like a natural speaker?”
Pleasantness“Which response was more pleasant to listen to?”
Conversational register“Which response used more natural, spoken-conversation language, rather than sounding like written text being read aloud?”
Technical QualityResponse quality“Which response was more helpful in achieving what the user was trying to accomplish?”
Reasoning*“Which response’s reasoning — when explaining its thinking or answering a question requiring judgment — was more correct and logically sound?”
Instruction following*“Which response followed the specific instructions or constraints more closely?”
Perceived audio quality“Which response had better audio quality?”
Pronunciation accuracy“Which response had more accurate pronunciation?”
Voice consistency“Which response kept a more consistent voice throughout (less noticeable change in character, timbre, or quality over time)?”
Every question a rater answers: overall preference, then 5 on Humanness and 6 on Technical Quality.* Only when relevant to input prompt

Results

Current leaderboards do not capture naturalness or acoustic quality

Because of our focus on (1) human perception of conversation quality and (2) consumer-oriented scenarios, our overall rankings diverge from those produced by other public leaderboards with coverage over the same models.

Overall preference by model

Overall preference per model on the CMOS Elo scale, with 95% confidence intervals.

For example, Grok Think Fast 2.0 and Qwen3 both perform very well on leaderboards that measure task completion on transactional voice-agent scenarios with verifiable outcomes, but perform significantly worse on our leaderboard.

Our rubrics resolve into two distinct groups, which we have labeled “Humanness” and “Technical Quality”. We found them by running hierarchical clustering over per-rubric rater responses (after adjusting for rater- and prompt-level effects — see methodology section), which gave a stable clustering at k=2 that consistently reproduced across rater-prompt bootstrapping:

  • Humanness: conversational register, emotional appropriateness, engagement, naturalness, pleasantness
  • Technical Quality: instruction following, perceived audio quality, pronunciation accuracy, reasoning, response quality, voice consistency

In previous experiments we grouped questions into “content”, “acoustic” and “interaction” categories. However, during this analysis, we found that these groupings did not capture meaningfully different information — not only did they produce the same model stack ranks, they produced Elo distributions that were statistically indistinguishable from the “overall” human preference Elo.

CMOS Elo by rubric metric

CMOS Elo per model per rubric metric, with 95% confidence intervals
MetricGPT-Realtime 2.1GPT-Realtime 2.0Gemini 3.1 Flash LiveGrok voice-think-fast 2.0Nova 2 SonicCascadeQwen3-Omni
Conversational register1160 [1140, 1179]1131 [1109, 1152]1122 [1100, 1143]981 [961, 1002]921 [888, 956]844 [808, 875]841 [798, 880]
Engagement1138 [1118, 1155]1102 [1080, 1120]1105 [1080, 1133]971 [952, 990]988 [957, 1029]883 [853, 912]813 [782, 847]
Emotional appropriateness1179 [1157, 1208]1138 [1104, 1173]1073 [1044, 1105]977 [950, 1008]964 [921, 1003]866 [811, 916]802 [747, 847]
Naturalness1176 [1160, 1196]1149 [1128, 1170]1127 [1106, 1147]977 [954, 998]960 [935, 988]810 [771, 844]800 [759, 840]
Pleasantness1164 [1146, 1183]1134 [1115, 1155]1107 [1085, 1132]987 [966, 1005]978 [950, 1003]847 [809, 885]783 [748, 821]
Response quality1113 [1099, 1128]1105 [1087, 1125]1009 [991, 1027]1006 [990, 1021]1050 [1020, 1080]947 [918, 974]770 [742, 798]
Instruction following1075 [1038, 1106]1064 [1023, 1103]1015 [982, 1049]1015 [979, 1048]966 [902, 1026]1003 [940, 1083]861 [815, 914]
Reasoning1118 [1095, 1145]1103 [1073, 1137]979 [949, 1008]1007 [978, 1037]1109 [1055, 1162]927 [870, 981]757 [700, 809]
Perceived audio quality1071 [1048, 1093]1068 [1047, 1090]1103 [1082, 1126]1015 [993, 1039]915 [883, 949]990 [947, 1033]837 [796, 873]
Pronunciation accuracy1056 [1035, 1079]1064 [1038, 1087]1073 [1053, 1094]1018 [996, 1040]1026 [1000, 1048]1002 [964, 1039]762 [716, 802]
Voice consistency1040 [1015, 1062]1066 [1041, 1087]1039 [1013, 1068]1023 [999, 1044]1020 [992, 1048]985 [949, 1022]827 [781, 875]
Model

GPT-Realtime 2.1Overall CMOS Elo: 1149

CMOS Elo for the selected model across the 11 rubric metrics, with Humanness metrics on the left and Technical Quality on the right.Humanness: conversational register, emotional appropriateness, engagement, naturalness, pleasantness · mean pairwise ρ = 0.68 — Technical Quality: instruction following, perceived audio quality, pronunciation accuracy, reasoning, response quality, voice consistency · mean pairwise ρ = 0.40

The root cause of the correlation of these dimensions across models is likely traceable to the model training pipeline: perhaps encoded in the inductive bias of modern architectures, influenced by related latent patterns in the training data, or implicit in the internal metrics different research labs are hill climbing.

These two clusters are a statement about the variation in performance of current models on our prompt distribution. As models improve unevenly across capabilities, or as the use cases evaluated change, these relationships may separate. It is entirely possible that a future model will excel in perceived audio quality and voice consistency but struggle at instruction following and reasoning, just as it is entirely possible that transactional use cases show no strong relationship between human preference and the “humanness” of the model.

Raters prefer certain voices over others

Further evidence for the importance of aesthetics to human preference comes from examining model performance disaggregated by speaking voice. By default, each model is evaluated using two of its supported speaking voices (one male and one female).

Win rate between a model’s two voices

GPT-Realtime 2.1/2.0

Gemini 3.1 Flash Live

Grok voice-think-fast 2.0

Nova 2 Sonic

Win rate in direct voice comparison for different models. Tied comparisons are set aside before computing each percentage. Qwen3-Omni is absent because its two voices were never paired against each other in the comparative collection.Brackets show the 95% confidence interval for the leading voice’s share. n.s. — the interval includes an even split, so the lead is not significant at 95%.

Sometimes, as was the case with Grok Ara vs. Rex voices, this preference was very prominent.

Rater comments by Grok voice

Ara
0:34

Rater free text comments

  • “The models response to the user was accurate as well as consistent and very natural. The overall model response was excellent”
Rex
0:42

Rater free text comments

  • “model sounds a bit robotic.”
  • “response needs to sound more energetic.”
  • “conversation felt a bit unnatural because of models tone of voice.”
Each Grok voice with the free-text comments raters left on it, sampled separately for each — one on Ara, three on Rex, and no comment set against any other; comments appear as raters typed them, spelling and punctuation included.

Users display a strong preference for Ara on Humanness dimensions, and a weaker preference for Ara on Technical Quality dimensions.

Ara’s margin over Rex (Grok voice-think-fast 2.0)

Grok Ara against Rex by dimension: ratings, CMOS margin, 95% confidence interval and win rate.
DimensionRatingsAra’s marginAra win rate
CMOS points95% CIties excluded
Overall
Overall preference446+0.93[0.63, 1.21]74%
Humanness
Pleasantness446+1.03[0.77, 1.29]82%
Naturalness446+0.98[0.74, 1.22]83%
Conversational register446+0.81[0.58, 1.05]79%
Engagement446+0.76[0.52, 1.00]78%
Emotional appropriateness126+0.72[0.28, 1.14]74%
Technical Quality
Instruction following78+0.40[-0.05, 0.88]n.s.78%
Reasoning122+0.36[-0.08, 0.78]n.s.63%
Response quality446+0.32[0.08, 0.56]64%
Perceived audio quality446+0.19[-0.03, 0.40]n.s.62%
Pronunciation accuracy446+0.18[0.02, 0.35]62%
Voice consistency446+0.14[-0.08, 0.37]n.s.55%
CMOS margin (−3 to +3 scale) and win rate for the two Grok voices, Ara vs. Rex, on Grok voice-think-fast 2.0. Tied comparisons are set aside before computing each win rate. Emotional appropriateness, instruction following and reasoning carry fewer ratings than the rest because those questions are only asked when the prompt calls for them.Bars are the CMOS margin in Ara’s favor, with the 95% confidence interval marked at the bar’s end; 0 is no preference between the two voices. n.s. — the interval includes 0, so the lead is not significant at 95%.

Notably, three Technical Quality dimensions should not vary by voice: reasoning, instruction following, and response quality. In most modern S2S architectures, one model backs every voice, so the underlying quality should be the same. For two models, our data disagrees; Grok and GPT-Realtime both show a voice gap on these dimensions. Raters prefer Grok’s Ara voice on response quality, and its reasoning and instruction-following margins lean the same way. GPT-Realtime shows the same lean between its voices. The reasoning and instruction-following results should be interpreted cautiously; they are gated (applying to only some prompts) and have smaller samples with wide confidence intervals. It’s possible this dynamic is caused by a halo effect from the Humanness preference, or that some models have deeper architectural or training data distribution explanations for the differences in voice performance.

Regardless of the underlying reason, these results demonstrate that you cannot assume comparable performance across model voices as most S2S leaderboards today do. Moreover, they demonstrate that voice-specific factors can strongly influence overall human preference on model outputs.

Higher thinking level does not increase performance on our leaderboard

In contrast to speaking voices, our data showed no material difference in overall model performance based on model thinking level:

Overall Elo by model and thinking level

Highlight rows
Only the three models that ship more than one thinking level appear here — Nova 2 Sonic, Cascade and Qwen3-Omni offer a single setting, so there is nothing to compare. Ordered by Elo within this subset; each plotted range is a 95% confidence interval.

This pattern holds even in isolated dimensions where additional thinking would be most likely to differentiate performance. Higher-thinking variants did not meaningfully out-perform their lower-thinking counterparts in instruction following or reasoning, for example.

We attribute the lack of differentiation among thinking modes primarily to the nature of our task distribution: the single-turn mix does not contain the sort of difficult multi-step problems where you would expect extra thinking to pay off. However, it is notable that the overall differentiation of model thinking levels across S2S leaderboards is mixed — Speech Agent Arena shows relatively little difference between different thinking levels, while τ³-Voice shows a larger effect, but only for certain models. Overall, the public evidence suggests that higher thinking can improve specific reasoning capabilities without reliably improving broader S2S model performance.

Conclusion

Our findings demonstrate that human preference is shaped as much by communication style as by content. This analysis relies on preference data from 283 participants, underscoring the value of human evaluation for measuring outcomes that are difficult to quantify.

At David AI we continue to evaluate public models, pre-release checkpoints, and agentic speech-to-speech workflows using a mixture of human and automated evaluation to better understand real-world interaction quality.

For inquiries regarding model evaluation or future research collaborations, contact evals@withdavid.ai. We are also currently hiring for the Evaluations team.

Bibtex citation:
@misc{davidai2026s2seval,
  title        = {Evaluating Human Preference with David AI’s Single-Turn S2S (DAI-S2S-ST) Human Preference Leaderboard},
  author       = {{David AI Research}},
  year         = {2026},
  howpublished = {\url{https://research.withdavid.ai/}},
  abstract     = {A human-preference leaderboard for speech-to-speech models, built from
                  comparative mean-opinion-score (CMOS) votes by 283 participants. We find
                  that human preference is shaped as much by communication style as by
                  content, underscoring the value of human evaluation for outcomes that are
                  hard to quantify automatically.},
  note         = {Ordinal Bradley-Terry Elo from CMOS preference votes, with two-way
                  prompt-by-rater cluster-bootstrap intervals. Contact: evals@withdavid.ai}
}

Appendix

Extended Data and Robustness Analysis

Evaluation Data

At launch, our leaderboard collection consists of ~153k comparative preference ratings from a pool of 283 raters across model responses to 819 human-recorded audio prompts. A single comparative preference rating is defined as a standard CMOS (-3 to +3 with ties allowed) rating for one of 12 rubric questions for a given pair of audio samples. A rated audio sample pair consists of the output of two model configurations that have been fed the same single-turn human-recorded audio input prompts. A model configuration is defined as the combination of <model, voice, thinking level>, and 19 distinct model configurations were evaluated (generally 2 voices and 2 thinking levels for each model, with a few exceptions in the model selection section). Not all model configurations were compared against all other model configurations for a given input prompt, but every model configuration had at least one comparison against another model configuration for each input prompt.

In addition to CMOS, we also collected MOS ratings for internal validation and comparative analysis with CMOS data. This MOS data is not included in the leaderboard or methodology section, but is discussed throughout the appendix.

The rater panel that supplied these ratings was roughly balanced on gender, with a mean age of 40.9 and coverage across age bands from 18 to 65+. Per collection, 283 raters took part in the CMOS collection and 251 raters in the MOS collection; 166 raters did both MOS and CMOS evaluations.

Table A1: Rater pool gender and age distribution
GenderPool18-2526-3536-4546-5556-6565+Total
MMOS23332819144121
CMOS25423328143145
FMOS12253127292126
CMOS12273229294133
OtherMOS1300004
CMOS1301005
TotalMOS36615946436251
CMOS38726558437283

Inter-rater agreement

In our data, we found a Krippendorff’s alpha for the “Overall” preference question of 0.216, and a Gwet AC2 of 0.211:

Table A2: Inter-rater agreement for CMOS metrics (Krippendorff’s alpha, Gwet AC2). In our case (same number of raters per comparison), Krippendorff’s alpha is numerically equivalent to ICC-1 value and hence reported as such. Low α with high AC2 indicates heavily tie-concentrated distributions.
CMOS Rubricα / ICC(1)Gwet AC2n units
Overall0.2160.2115653
Humanness0.2340.49524206
Conversational register0.2420.5085653
Engagement0.2070.4595653
Pleasantness0.2340.5015653
Naturalness0.2500.5435653
Emotional appropriateness0.2370.4031594
Technical quality0.1320.68325175
Instruction following0.2250.5381010
Perceived audio quality0.0570.6695653
Pronunciation accuracy0.0540.8805653
Voice consistency0.0470.7115653
Response quality0.2070.5015653
Reasoning0.1840.5771553

However, on average each sample pair has only ~2 independent ratings, making traditional inter-rater agreement metrics difficult to interpret. Instead, we prefer split-half Spearman correlation of overall model ranking, which measures how similar the model rankings are when computed from two disjoint random halves of the rater pool. The table below shows the mean correlation of split-half rankings over 20,000 random splits:

Table A3: Split-half Spearman correlation for our rankings (CMOS) and standard deviation
FamilyRubricρsd
OverallOverall preference0.9640.013
HumannessEmotional appropriateness (gated)0.9610.018
Naturalness0.9550.023
Engagement0.9340.028
Conversational register0.9230.026
Pleasantness0.9170.029
Technical QualityResponse quality0.9380.028
Perceived audio quality0.9250.042
Reasoning (gated)0.8670.031
Pronunciation accuracy0.8130.073
Instruction following (gated)0.8020.069
Voice consistency0.6740.122

Together, these metrics tell us that per-rating agreement is highly variable across raters, but in aggregate the overall model ranking produced is very stable. More discussion of this phenomenon can be found in the appendix section of our previous blog post.

MOS vs. CMOS Agreement

Our published leaderboard relies on CMOS ratings data to drive Elo scores, but we also collected human MOS ratings on the same set of rubrics on a distinct set of input prompts.

Table A4: CMOS vs. MOS (Overall) with 95%-CI per model; MOS score is calculated by averaging MOS scores for each sample and averaging all samples.
ModelCMOS EloMOS (sample-averaged)
GPT-Realtime 2.11149 [1130, 1168]4.13 [4.03, 4.23]
GPT-Realtime 2.01128 [1109, 1149]4.08 [3.98, 4.18]
Gemini 3.1 Flash Live1071 [1050, 1091]4.26 [4.19, 4.32]
Grok voice-think-fast 2.0993 [977, 1009]3.64 [3.53, 3.75]
Nova 2 Sonic979 [948, 1014]3.59 [3.46, 3.73]
Cascade897 [864, 928]3.62 [3.47, 3.74]
Qwen3-Omni-30B-A3B-Instruct784 [751, 817]3.04 [2.91, 3.18]

While these two metrics largely agree, there are two notable differences:

  1. MOS flips the Gemini 3.1 Flash Live and GPT-Realtime-2.0/2.1 rankings at the top.
  2. MOS considered the cascade baseline to be competitive with Grok and Nova 2 Sonic, while CMOS considered the cascade baseline to lag far behind.

There are several plausible explanations for why you might see a difference between MOS and CMOS ratings that are worth examining.

Case 1: MOS penalizes catastrophic failures more harshly in saturated domains

The acoustic quality of modern TTS systems is high enough to saturate MOS ratings - almost everything is “pretty good” on an objective scale, and model response that’s in the lowest quartile of output quality might still get a 4 on the MOS scale. In contrast, CMOS asks about relative preference, and the extreme values of the scale simply indicate a “strong” subjective preference for one output over another. A model might have a perfectly good output, but a competing model might still be strongly preferred by the rater. This means that a catastrophic failure earns -3 on CMOS, but so does a merely-strongly-dispreferred output, so CMOS caps the penalty of catastrophic samples while MOS doesn't.

The implication is that model A could be preferred over model B in general, but also have more catastrophic failures, which would lead to model A having a higher CMOS rating but a lower MOS rating. In the case of Gemini 3.1 Flash Live and GPT-Realtime-2.1, we see the following distributions:

Table A5: Distribution of “overall” MOS ratings for each model
12345
Gemini 3.1 Flash Live0.9%2.1%11.8%41%44.2%
GPT-Realtime-2.12.2%3.2%14.5%40.3%39.8%

In our data, it appears that this effect partially explains the MOS / CMOS discrepancy, but does not completely resolve the question. It is true that GPT-Realtime-2.1 MOS shows more catastrophic failures (MOS scores of 1 and 2) than Gemini 3.1 Flash Live, but this is part of an overall left-shift across the whole distribution; there are also more 3’s and fewer 4’s and 5’s.

Case 2: Heterogeneity in model performance across the input distribution impacts MOS and CMOS asymmetrically

If one model performs characteristically differently on a subset of input prompts, it could have a similar distributional effect to the catastrophic failure case discussed above; model A might be significantly worse at a subset of the eval, but marginally better at the majority, which would lead to stronger CMOS rankings and weaker MOS rankings. To explore this in our data, we examine whether MOS score difference is specific to the difficulty of the prompt:

Table A6: MOS values for GPT-Realtime-2.1 and Gemini 3.1 Flash Live per Difficulty Level (L1 - L5 + degraded category).
DifficultyGPT-Realtime 2.1Gemini 3.1 Flash Livegapn (GPT-Realtime 2.1 / Gemini)
L14.164.30+0.13770 / 1140
L24.164.29+0.13809 / 1168
L34.154.25+0.10844 / 1162
L44.074.21+0.141223 / 1222
L54.134.23+0.111075 / 1076
Degraded4.084.24+0.16349 / 348

As expected, the overall MOS score is marginally higher for easier prompts and lower for harder prompts, but the gap between models remains relatively consistent across difficulty levels.

Case 3: Certain dimensions of quality impact overall quality score differently than they impact overall preference

For example, if a rater strongly prefers one speaking voice over another, this could result in a high overall MOS ratings for model A while still showing a strong overall preference for model B on overall CMOS preference. To investigate, we examine the relationship between the MOS / CMOS delta between models across rubric dimensions:

Table A7: MOS vs. CMOS per metrics for GPT-Realtime 2.1 - Gemini 3.1 Flash Live. The gated metrics are only applicable to certain prompts. CMOS value of the average head-to-head CMOS score on a scale of (-3..+3)
RubricMOS Δ (GPT 2.1 − Gemini)CMOS Δ (GPT 2.1 vs Gemini)
Overall preference−0.13+0.58
Humanness
Conversational register−0.09+0.23
Engagement−0.26+0.18
Emotional appropriateness (gated)+0.06+0.65
Naturalness−0.07+0.29
Pleasantness−0.00+0.35
Technical quality
Response quality+0.06+0.60
Reasoning (gated)+0.02+0.70
Instruction following (gated)+0.04+0.38
Perceived audio quality−0.19−0.15
Pronunciation accuracy−0.11−0.04
Voice consistency−0.17+0.04

Here again, we find the overall MOS ratings favoring Gemini 3.1 Flash Live across most dimensions, while the CMOS ratings favor GPT-Realtime-2.1 across most dimensions. Interestingly, we do see that perceived audio quality favors Gemini for both MOS and CMOS. One interpretation of this might be that in isolation raters tended to penalize GPT-Realtime for audio quality issues, but this factor did not strongly influence their overall preference across model outputs.

Our conclusion is that MOS and CMOS are simply measuring different things; it is entirely plausible that there are some model quality shortcomings that are relatively more obvious in a standalone pointwise setting, and others that are relatively more obvious in a comparative setting.

Rubric clustering and correlation

Each rubric category bundles several sub-metrics which are themselves positively correlated across prompts within a cluster. Humanness is highly coherent: its five delivery/affect sub-questions (naturalness, pleasantness, engagement, emotional appropriateness, conversational register) move together at a mean pairwise correlation of 0.68, with naturalness and pleasantness the tightest pair (0.82).

Technical Quality is more loosely coupled, with a mean correlation of 0.40. One explanation for this is that it spans two distinct sub-groups: the lexical content items — response quality, instruction following and reasoning — track each other closely at 0.82, while the acoustic content items — perceived audio quality, pronunciation accuracy and voice consistency — cohere more modestly at 0.47. These two sub-groups of Technical Quality correlate only weakly with each other (0.28).

Table A8: Metrics correlation within a category group and across category groups
Group / clusterRubricsMean pairwise ρ
Humannessnaturalness, pleasantness, engagement, emotional appropriateness, conversational register0.68
Technical Qualityresponse quality, instruction following, reasoning, perceived audio quality, pronunciation accuracy, voice consistency0.40
Technical Quality - lexicalresponse quality, instruction following, reasoning0.82
Technical Quality - audioperceived audio quality, pronunciation accuracy, voice consistency0.47
lexical × audio sub-groupcross-group0.28

Humanness can be interpreted as a single tight group, while Technical Quality is a broader axis holding a strong lexical content block and a looser audio content block.

These two sub-groups are exactly the groupings that emerge for a k=3 clustering of our data, so it is worth saying why we do not separate them out: we do not find this third cluster to reproduce reliably across rater-prompt bootstrapping (see methodology section). We hypothesize that with more evaluation data and more models being evaluated we would see this third grouping emerge reliably, but our data as yet does not justify treating these as separate clusters, so for the time being we continue using the k=2 grouping of Humanness and Technical Quality that robustly reproduces.

Insights from free text analysis

Our collection protocol includes the ability for raters to leave free text comments explaining their ratings, rationale, or any issues they want to flag that are not captured in existing rubrics. We perform clustering analysis around themes from free-text rater comments and display the results in the table below. Theme mention rate is calculated as percentage of the model’s rater comments from the MOS ratings, where (as discussed in previous post) we see more robust and informative free text comments. Note that this analysis shows the relative differences between the way raters tend to characterize the responses given by each model, rather than the absolute number of comments that addressed the particular theme. Our hope is that these qualitative insights are informative to model trainers in identifying areas of strength and weakness of their model, and that they may influence future model training priorities.

Table A9: Theme mention rate by model, as a percentage of each model’s MOS rater comments. Shading compares across a row, not down a column. Themes overlap, so rates do not sum to 100.
ThemeGPT-Realtime 2.1GPT-Realtime 2.0Gemini 3.1 FlashGrok voice-think-fastNova 2 SonicCascadeQwen3-Omni
Natural / human17.015.417.010.59.210.78.5
Robotic / monotone5.47.74.419.516.923.615.2
Warm / pleasant9.18.08.56.06.14.44.1
Helpful / informative33.933.734.530.429.237.221.3
Too long / verbose2.82.70.62.114.30.80.5
Pacing (fast / slow)5.55.32.83.65.58.214.6
Audio artifact11.610.85.813.324.18.015.2
Accent / mispronunciation4.95.43.14.42.95.310.3
Emotional / empathetic6.35.25.83.22.12.72.5
Breathy / breathing1.91.85.20.70.90.11.6

Rubric-Level MOS & CMOS Rankings

Below is the full list of CMOS Elo ranking and MOS scores per model per metric with 95% CI; the MOS score is calculated using sample-averaging. CMOS Elo rankings follow the same BT fitting approach discussed in the methodology section, simply using the specific rubric preference question rather than the overall preference question.

Table A10: CMOS Elo and sample-averaged MOS per model for every rubric, with 95% CIs in brackets.
FamilyCMOS EloMOS 1–5 (sample-averaged)
Overall preference
GPT-Realtime 2.11149 [1130, 1168]4.13 [4.03, 4.23]
GPT-Realtime 2.01128 [1109, 1149]4.08 [3.98, 4.18]
Gemini 3.1 Flash Live1071 [1050, 1091]4.26 [4.19, 4.32]
Grok voice-think-fast 2.0993 [977, 1009]3.64 [3.53, 3.75]
Nova 2 Sonic979 [948, 1014]3.59 [3.46, 3.73]
Cascade897 [864, 928]3.62 [3.47, 3.74]
Qwen3-Omni-30B-A3B-Instruct784 [751, 817]3.04 [2.91, 3.18]
Humanness
Conversational register
GPT-Realtime 2.11160 [1140, 1179]4.24 [4.16, 4.33]
GPT-Realtime 2.01131 [1109, 1152]4.17 [4.07, 4.26]
Gemini 3.1 Flash Live1122 [1100, 1143]4.33 [4.25, 4.42]
Grok voice-think-fast 2.0981 [961, 1002]3.71 [3.58, 3.83]
Nova 2 Sonic921 [888, 956]3.59 [3.44, 3.73]
Cascade844 [808, 875]3.44 [3.26, 3.61]
Qwen3-Omni-30B-A3B-Instruct841 [798, 880]3.14 [2.95, 3.32]
Engagement
GPT-Realtime 2.11138 [1118, 1155]4.00 [3.91, 4.10]
Gemini 3.1 Flash Live1105 [1080, 1133]4.25 [4.18, 4.33]
GPT-Realtime 2.01102 [1080, 1120]3.90 [3.78, 4.01]
Nova 2 Sonic988 [957, 1029]3.46 [3.31, 3.60]
Grok voice-think-fast 2.0971 [952, 990]3.54 [3.42, 3.65]
Cascade883 [853, 912]3.41 [3.25, 3.57]
Qwen3-Omni-30B-A3B-Instruct813 [782, 847]3.12 [2.96, 3.26]
Emotional appropriateness
GPT-Realtime 2.11179 [1157, 1208]4.43 [4.35, 4.52]
GPT-Realtime 2.01138 [1104, 1173]4.36 [4.24, 4.47]
Gemini 3.1 Flash Live1073 [1044, 1105]4.37 [4.28, 4.46]
Grok voice-think-fast 2.0977 [950, 1008]3.77 [3.64, 3.91]
Nova 2 Sonic964 [921, 1003]3.98 [3.81, 4.16]
Cascade866 [811, 916]3.68 [3.48, 3.86]
Qwen3-Omni-30B-A3B-Instruct802 [747, 847]3.14 [2.94, 3.35]
Naturalness
GPT-Realtime 2.11176 [1160, 1196]4.14 [4.06, 4.24]
GPT-Realtime 2.01149 [1128, 1170]4.10 [4.01, 4.20]
Gemini 3.1 Flash Live1127 [1106, 1147]4.22 [4.13, 4.30]
Grok voice-think-fast 2.0977 [954, 998]3.52 [3.41, 3.65]
Nova 2 Sonic960 [935, 988]3.60 [3.47, 3.73]
Cascade810 [771, 844]3.11 [2.90, 3.30]
Qwen3-Omni-30B-A3B-Instruct800 [759, 840]2.78 [2.62, 2.93]
Pleasantness
GPT-Realtime 2.11164 [1146, 1183]4.18 [4.09, 4.27]
GPT-Realtime 2.01134 [1115, 1155]4.14 [4.04, 4.24]
Gemini 3.1 Flash Live1107 [1085, 1132]4.18 [4.10, 4.26]
Grok voice-think-fast 2.0987 [966, 1005]3.57 [3.45, 3.69]
Nova 2 Sonic978 [950, 1003]3.73 [3.61, 3.86]
Cascade847 [809, 885]3.39 [3.23, 3.56]
Qwen3-Omni-30B-A3B-Instruct783 [748, 821]2.94 [2.79, 3.08]
Technical quality
Response quality
GPT-Realtime 2.11113 [1099, 1128]4.39 [4.33, 4.46]
GPT-Realtime 2.01105 [1087, 1125]4.41 [4.33, 4.48]
Nova 2 Sonic1050 [1020, 1080]4.14 [4.04, 4.24]
Gemini 3.1 Flash Live1009 [991, 1027]4.33 [4.26, 4.40]
Grok voice-think-fast 2.01006 [990, 1021]4.19 [4.12, 4.27]
Cascade947 [918, 974]4.16 [4.06, 4.26]
Qwen3-Omni-30B-A3B-Instruct770 [742, 798]3.71 [3.60, 3.82]
Instruction following
GPT-Realtime 2.11075 [1038, 1106]4.52 [4.43, 4.61]
GPT-Realtime 2.01064 [1023, 1103]4.53 [4.41, 4.64]
Gemini 3.1 Flash Live1015 [982, 1049]4.49 [4.39, 4.57]
Grok voice-think-fast 2.01015 [979, 1048]4.41 [4.30, 4.51]
Cascade1003 [940, 1083]4.39 [4.22, 4.55]
Nova 2 Sonic966 [902, 1026]4.33 [4.16, 4.48]
Qwen3-Omni-30B-A3B-Instruct861 [815, 914]4.00 [3.79, 4.19]
Reasoning
GPT-Realtime 2.11118 [1095, 1145]4.70 [4.63, 4.76]
Nova 2 Sonic1109 [1055, 1162]4.47 [4.35, 4.59]
GPT-Realtime 2.01103 [1073, 1137]4.70 [4.62, 4.78]
Grok voice-think-fast 2.01007 [978, 1037]4.55 [4.47, 4.64]
Gemini 3.1 Flash Live979 [949, 1008]4.68 [4.62, 4.74]
Cascade927 [870, 981]4.45 [4.33, 4.57]
Qwen3-Omni-30B-A3B-Instruct757 [700, 809]4.16 [4.00, 4.30]
Perceived audio quality
Gemini 3.1 Flash Live1103 [1082, 1126]4.51 [4.45, 4.58]
GPT-Realtime 2.11071 [1048, 1093]4.32 [4.23, 4.42]
GPT-Realtime 2.01068 [1047, 1090]4.31 [4.22, 4.39]
Grok voice-think-fast 2.01015 [993, 1039]4.16 [4.04, 4.27]
Cascade990 [947, 1033]4.17 [4.04, 4.29]
Nova 2 Sonic915 [883, 949]3.91 [3.77, 4.05]
Qwen3-Omni-30B-A3B-Instruct837 [796, 873]3.68 [3.48, 3.86]
Pronunciation accuracy
Gemini 3.1 Flash Live1073 [1053, 1094]4.76 [4.71, 4.81]
GPT-Realtime 2.01064 [1038, 1087]4.63 [4.57, 4.70]
GPT-Realtime 2.11056 [1035, 1079]4.65 [4.58, 4.71]
Nova 2 Sonic1026 [1000, 1048]4.64 [4.56, 4.72]
Grok voice-think-fast 2.01018 [996, 1040]4.64 [4.58, 4.70]
Cascade1002 [964, 1039]4.55 [4.45, 4.63]
Qwen3-Omni-30B-A3B-Instruct762 [716, 802]4.01 [3.87, 4.14]
Voice consistency
GPT-Realtime 2.01066 [1041, 1087]4.53 [4.45, 4.62]
GPT-Realtime 2.11040 [1015, 1062]4.51 [4.42, 4.60]
Gemini 3.1 Flash Live1039 [1013, 1068]4.68 [4.62, 4.73]
Grok voice-think-fast 2.01023 [999, 1044]4.40 [4.32, 4.49]
Nova 2 Sonic1020 [992, 1048]4.55 [4.47, 4.63]
Cascade985 [949, 1022]4.35 [4.23, 4.47]
Qwen3-Omni-30B-A3B-Instruct827 [781, 875]4.04 [3.90, 4.16]
Model Configuration Details
OpenAIGPT-Realtime-2.0-xhigh (Marin)
Config (simplified)
Model: GPT-Realtime 2.0 Voice: Marin Reasoning: xhigh Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gpt-realtime-2" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "marin" } } reasoning: { effort: "xhigh" } instructions: "You are a helpful voice assistant."
OpenAIGPT-Realtime-2.0-xhigh (Cedar)
Config (simplified)
Model: GPT-Realtime 2.0 Voice: Cedar Reasoning: xhigh Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gpt-realtime-2" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "cedar" } } reasoning: { effort: "xhigh" } instructions: "You are a helpful voice assistant."
OpenAIGPT-Realtime-2.1-xhigh (Marin)
Config (simplified)
Model: GPT-Realtime 2.1 Voice: Marin Reasoning: xhigh Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gpt-realtime-2.1" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "marin" } } reasoning: { effort: "xhigh" } instructions: "You are a helpful voice assistant."
OpenAIGPT-Realtime-2.1-xhigh (Cedar)
Config (simplified)
Model: GPT-Realtime 2.1 Voice: Cedar Reasoning: xhigh Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gpt-realtime-2.1" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "cedar" } } reasoning: { effort: "xhigh" } instructions: "You are a helpful voice assistant."
OpenAIGPT-Realtime-2.1-minimal (Marin)
Config (simplified)
Model: GPT-Realtime 2.1 Voice: Marin Reasoning: minimal Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gpt-realtime-2.1" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "marin" } } reasoning: { effort: "minimal" } instructions: "You are a helpful voice assistant."
OpenAIGPT-Realtime-2.1-minimal (Cedar)
Config (simplified)
Model: GPT-Realtime 2.1 Voice: Cedar Reasoning: minimal Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gpt-realtime-2.1" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "cedar" } } reasoning: { effort: "minimal" } instructions: "You are a helpful voice assistant."
xAIGrok Voice Think Fast 2.0 - High (Ara)
Config (simplified)
Model: Grok Voice Think Fast 2.0 Voice: Ara Reasoning: high Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "grok-voice-think-fast-2.0" session: { instructions: "You are a helpful voice assistant.", voice: "ara", reasoning: { effort: "high" }, turn_detection: { type: null }, audio: { input: { format: { type: "audio/pcm", rate: 24000 } }, output: { format: { type: "audio/pcm", rate: 24000 } } } } Supported voices: ara, eve, leo, rex, sal
xAIGrok Voice Think Fast 2.0 - High (Rex)
Config (simplified)
Model: Grok Voice Think Fast 2.0 Voice: Rex Reasoning: high Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "grok-voice-think-fast-2.0" session: { instructions: "You are a helpful voice assistant.", voice: "rex", reasoning: { effort: "high" }, turn_detection: { type: null }, audio: { input: { format: { type: "audio/pcm", rate: 24000 } }, output: { format: { type: "audio/pcm", rate: 24000 } } } } Supported voices: ara, eve, leo, rex, sal
xAIGrok Voice Think Fast 2.0 - None (Ara)
Config (simplified)
Model: Grok Voice Think Fast 2.0 Voice: Ara Reasoning: none (xAI's own default is high, so this is sent explicitly) Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "grok-voice-think-fast-2.0" session: { instructions: "You are a helpful voice assistant.", voice: "ara", reasoning: { effort: "none" }, turn_detection: { type: null }, audio: { input: { format: { type: "audio/pcm", rate: 24000 } }, output: { format: { type: "audio/pcm", rate: 24000 } } } } Supported voices: ara, eve, leo, rex, sal
xAIGrok Voice Think Fast 2.0 - None (Rex)
Config (simplified)
Model: Grok Voice Think Fast 2.0 Voice: Rex Reasoning: none (xAI's own default is high, so this is sent explicitly) Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "grok-voice-think-fast-2.0" session: { instructions: "You are a helpful voice assistant.", voice: "rex", reasoning: { effort: "none" }, turn_detection: { type: null }, audio: { input: { format: { type: "audio/pcm", rate: 24000 } }, output: { format: { type: "audio/pcm", rate: 24000 } } } } Supported voices: ara, eve, leo, rex, sal
GoogleGemini 3.1 Flash Live Preview - thinking-high (Zephyr)
Config (simplified)
Model: Gemini 3.1 Flash Live Preview Voice: Zephyr Thinking: high — thoughts not returned Audio: PCM16, 16 kHz in / 24 kHz out Turn detection: off — automatic activity detection disabled Extras: output audio transcription on Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gemini-3.1-flash-live-preview" responseModalities: ["AUDIO"] systemInstruction: "You are a helpful voice assistant." speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Zephyr" } } } realtimeInputConfig: { automaticActivityDetection: { disabled: true } } thinkingConfig: { thinkingLevel: "HIGH", includeThoughts: false } outputAudioTranscription: {} realtimeInput: { audio: { data: <b64>, mimeType: "audio/pcm;rate=16000" } } Output audio: 24000 Hz PCM16
GoogleGemini 3.1 Flash Live Preview - thinking-high (Iapetus)
Config (simplified)
Model: Gemini 3.1 Flash Live Preview Voice: Iapetus Thinking: high — thoughts not returned Audio: PCM16, 16 kHz in / 24 kHz out Turn detection: off — automatic activity detection disabled Extras: output audio transcription on Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gemini-3.1-flash-live-preview" responseModalities: ["AUDIO"] systemInstruction: "You are a helpful voice assistant." speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Iapetus" } } } realtimeInputConfig: { automaticActivityDetection: { disabled: true } } thinkingConfig: { thinkingLevel: "HIGH", includeThoughts: false } outputAudioTranscription: {} realtimeInput: { audio: { data: <b64>, mimeType: "audio/pcm;rate=16000" } } Output audio: 24000 Hz PCM16
GoogleGemini 3.1 Flash Live Preview - thinking-minimal (Zephyr)
Config (simplified)
Model: Gemini 3.1 Flash Live Preview Voice: Zephyr Thinking: minimal — thoughts not returned Audio: PCM16, 16 kHz in / 24 kHz out Turn detection: off — automatic activity detection disabled Extras: output audio transcription on Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gemini-3.1-flash-live-preview" responseModalities: ["AUDIO"] systemInstruction: "You are a helpful voice assistant." speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Zephyr" } } } realtimeInputConfig: { automaticActivityDetection: { disabled: true } } thinkingConfig: { thinkingLevel: "MINIMAL", includeThoughts: false } outputAudioTranscription: {} realtimeInput: { audio: { data: <b64>, mimeType: "audio/pcm;rate=16000" } } Output audio: 24000 Hz PCM16
GoogleGemini 3.1 Flash Live Preview - thinking-minimal (Iapetus)
Config (simplified)
Model: Gemini 3.1 Flash Live Preview Voice: Iapetus Thinking: minimal — thoughts not returned Audio: PCM16, 16 kHz in / 24 kHz out Turn detection: off — automatic activity detection disabled Extras: output audio transcription on Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "gemini-3.1-flash-live-preview" responseModalities: ["AUDIO"] systemInstruction: "You are a helpful voice assistant." speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Iapetus" } } } realtimeInputConfig: { automaticActivityDetection: { disabled: true } } thinkingConfig: { thinkingLevel: "MINIMAL", includeThoughts: false } outputAudioTranscription: {} realtimeInput: { audio: { data: <b64>, mimeType: "audio/pcm;rate=16000" } } Output audio: 24000 Hz PCM16
AmazonNova 2 Sonic (Tiffany)
Config (simplified)
Model: Nova 2 Sonic (Bedrock, us-east-1) Voice: Tiffany Reasoning: n/a — no effort control Audio: LPCM16 mono, 16 kHz in / 24 kHz out Turn detection: cannot be disabled — endpointing LOW (~2.0 s pause), the most patient setting Sampling: model defaults (no max tokens, top-p or temperature set) Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "amazon.nova-2-sonic-v1:0" (bedrock, us-east-1) sessionStart.inferenceConfiguration: {} # maxTokens / topP / temperature all unset sessionStart.turnDetectionConfiguration: { endpointingSensitivity: "LOW" } promptStart.textOutputConfiguration: { mediaType: "text/plain" } promptStart.audioOutputConfiguration: { mediaType: "audio/lpcm", sampleRateHertz: 24000, sampleSizeBits: 16, channelCount: 1, audioType: "SPEECH", encoding: "base64", voiceId: "tiffany" } contentStart(SYSTEM,TEXT) + textInput.content: "You are a helpful voice assistant." contentStart(USER,AUDIO).audioInputConfiguration: { mediaType: "audio/lpcm", sampleRateHertz: 16000, sampleSizeBits: 16, channelCount: 1, audioType: "SPEECH", encoding: "base64" }
AmazonNova 2 Sonic (Matthew)
Config (simplified)
Model: Nova 2 Sonic (Bedrock, us-east-1) Voice: Matthew Reasoning: n/a — no effort control Audio: LPCM16 mono, 16 kHz in / 24 kHz out Turn detection: cannot be disabled — endpointing LOW (~2.0 s pause), the most patient setting Sampling: model defaults (no max tokens, top-p or temperature set) Prompt: default ("You are a helpful voice assistant.")
System prompt
You are a helpful voice assistant.
Full parameters
model: "amazon.nova-2-sonic-v1:0" (bedrock, us-east-1) sessionStart.inferenceConfiguration: {} # maxTokens / topP / temperature all unset sessionStart.turnDetectionConfiguration: { endpointingSensitivity: "LOW" } promptStart.textOutputConfiguration: { mediaType: "text/plain" } promptStart.audioOutputConfiguration: { mediaType: "audio/lpcm", sampleRateHertz: 24000, sampleSizeBits: 16, channelCount: 1, audioType: "SPEECH", encoding: "base64", voiceId: "matthew" } contentStart(SYSTEM,TEXT) + textInput.content: "You are a helpful voice assistant." contentStart(USER,AUDIO).audioInputConfiguration: { mediaType: "audio/lpcm", sampleRateHertz: 16000, sampleSizeBits: 16, channelCount: 1, audioType: "SPEECH", encoding: "base64" }
Alibaba (self-hosted, Modal)Qwen3-Omni-30B-A3B-Instruct (Chelsie)
Config (simplified)
Model: Qwen3-Omni 30B-A3B Instruct — self-hosted on Modal (H200) Voice: Chelsie Reasoning: n/a Audio: WAV PCM16, 16 kHz in / 24 kHz out Turn detection: n/a — one clip per call, batched Prompt: Qwen's own official assistant prompt. Caps replies at 50 words, forbids formatting.
System prompt
You are a virtual voice assistant with no gender or age. You are communicating with the user. In user messages, "I/me/my/we/our" refer to the user and "you/your" refer to the assistant. In your replies, address the user as "you/your" and yourself as "I/me/my"; never mirror the user's pronouns—always shift perspective. Keep original pronouns only in direct quotes; if a reference is unclear, ask a brief clarifying question. Interact with users using short (no more than 50 words), brief, straightforward language, maintaining a natural tone. Never use formal phrasing, mechanical expressions, bullet points, overly structured language. Your output must consist only of the spoken content you want the user to hear. Do not include any descriptions of actions, emotions, sounds, or voice changes. Do not use asterisks, brackets, parentheses, or any other symbols to indicate tone or actions. You must answer users' audio or text questions, do not directly describe the video content. You should communicate in the same language strictly as the user unless they request otherwise. When you are uncertain (e.g., you can't see/hear clearly, don't understand, or the user makes a comment rather than asking a question), use appropriate questions to guide the user to continue the conversation. Keep replies concise and conversational, as if talking face-to-face
Full parameters
model: "Qwen/Qwen3-Omni-30B-A3B-Instruct" (self-hosted) endpoint: modal://main/qwen3-omni-instruct/Qwen3OmniS2SBatched.generate task: "s2s" speaker: "chelsie" messages: [ { role: "system", content: [{ type: "text", text: <official prompt> }] }, { role: "user", content: [{ type: "audio", audio_b64: <WAV b64> }] } ] Input audio: WAV PCM16 16000 Hz Output audio: 24000 Hz PCM16
Alibaba (self-hosted, Modal)Qwen3-Omni-30B-A3B-Instruct (Ethan)
Config (simplified)
Model: Qwen3-Omni 30B-A3B Instruct — self-hosted on Modal (H200) Voice: Ethan Reasoning: n/a Audio: WAV PCM16, 16 kHz in / 24 kHz out Turn detection: n/a — one clip per call, batched Prompt: Qwen's own official assistant prompt. Caps replies at 50 words, forbids formatting.
System prompt
You are a virtual voice assistant with no gender or age. You are communicating with the user. In user messages, "I/me/my/we/our" refer to the user and "you/your" refer to the assistant. In your replies, address the user as "you/your" and yourself as "I/me/my"; never mirror the user's pronouns—always shift perspective. Keep original pronouns only in direct quotes; if a reference is unclear, ask a brief clarifying question. Interact with users using short (no more than 50 words), brief, straightforward language, maintaining a natural tone. Never use formal phrasing, mechanical expressions, bullet points, overly structured language. Your output must consist only of the spoken content you want the user to hear. Do not include any descriptions of actions, emotions, sounds, or voice changes. Do not use asterisks, brackets, parentheses, or any other symbols to indicate tone or actions. You must answer users' audio or text questions, do not directly describe the video content. You should communicate in the same language strictly as the user unless they request otherwise. When you are uncertain (e.g., you can't see/hear clearly, don't understand, or the user makes a comment rather than asking a question), use appropriate questions to guide the user to continue the conversation. Keep replies concise and conversational, as if talking face-to-face
Full parameters
model: "Qwen/Qwen3-Omni-30B-A3B-Instruct" (self-hosted) endpoint: modal://main/qwen3-omni-instruct/Qwen3OmniS2SBatched.generate task: "s2s" speaker: "ethan" messages: [ { role: "system", content: [{ type: "text", text: <official prompt> }] }, { role: "user", content: [{ type: "audio", audio_b64: <WAV b64> }] } ] Input audio: WAV PCM16 16000 Hz Output audio: 24000 Hz PCM16
ElevenLabs, OpenAI, ElevenLabsCascade
Config (simplified)
Pipeline: ElevenLabs Scribe v2 (ASR) -> OpenAI GPT-5.6-terra (LLM) -> ElevenLabs Flash v2.5 (TTS) Voice: Rachel (ElevenLabs legacy default, retires 2026-12-31) Reasoning: low — on the LLM leg Audio: PCM, 16 kHz in / 24 kHz out Turn detection: n/a — one clip per call Language: auto (not pinned) Prompt: pinned by the deployed app, cannot be overridden. Caps replies at 100 words, forbids markdown.
System prompt
You are a helpful assistant handling a voice chat with a user. # Important Voice Considerations 1. Respond naturally and conversationally as you would in a real conversation 2. Try to be helpful and always follow the instructions below. 3. Keep the response short and concise. Do not exceed 100 words. You are writing a final script for text-to-speech (TTS). Your response will be synthesized directly into speech. Follow the duration instruction as strictly as possible. Output only the final spoken text, with natural punctuation. Do not output markdown, bullets, JSON, XML tags, stage directions, or extra commentary. Do not mention these instructions.
Full parameters
endpoint: modal://main/s2s-cascaded/SingleTurnCascade.run_turn asr_model: "scribe_v2" (ElevenLabs Scribe v2) llm_model: "gpt-5.6-terra" (OpenAI) tts_provider: "elevenlabs" tts_model: "eleven_flash_v2_5" voice: "21m00Tcm4TlvDq8ikWAM" (Rachel; legacy default, retires 2026-12-31) language: null reasoning_effort: "low" messages: [{ role: "user", content: [{ type: "audio", audio_b64: <WAV b64>, filename: "prompt.wav" }] }] Input audio: 16000 Hz Output audio: 24000 Hz PCM No system-prompt input: the deployed app pins its own and 400s on a system turn.
References

Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4), 324–345.

Chiang, W.-L., et al. (2024). Chatbot Arena: an open platform for evaluating LLMs by human preference.

Elo, A. E. (1978). The Rating of Chessplayers, Past and Present. Arco.

Eyben, F., Scherer, K. R., Schuller, B. W., et al. (2016). The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing. IEEE Transactions on Affective Computing, 7(2), 190–202.

Krippendorff, K. (2004). Content Analysis: An Introduction to Its Methodology (2nd ed.). Sage.

Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29–48.

McCullagh, P. (1980). Regression models for ordinal data. JRSS: Series B, 42(2), 109–142. (cumulative-link / ordinal model)

Massey, K. (1997). Statistical models applied to the rating of sports teams. Bluefield College.

Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.