Spokes.wiki Search About
Defined Term concept updated Mon Aug 03 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

TTS benchmarks

How text-to-speech quality is measured — and why “best” is multi-axis and volatile. The leading public ranking is the Speech Arena Elo board (tts-arena-leaderboard).

The metrics

  • Elo (Speech Arena) — blind A/B human preference votes → an Elo rating; captures perceived naturalness but nothing about latency or accuracy. The headline number.
  • MOS (Mean Opinion Score) — rated naturalness, typically on ~10-second clips (UTMOS training bounds it); kokoro scores ~4.5 MOS. The metric is older than TTS: it’s a five-point absolute-category scale (5 Excellent → 1 Bad), averaged over listeners, standardized for telephony in ITU-T P.800 (mean-opinion-score). The number to distrust isn’t the rating — it’s comparing ratings across tests: ITU itself warns MOS values “should not be compared across experiments” unless designed to be, because listener pool, clips, and instructions shift the absolute scale. So “4.5 MOS” only ranks a model against the others in its own test — the standards basis for this page’s snapshot discipline.
  • WER / CER — word/character error rate via round-trip ASR transcription; depends on the ASR model, so it’s noisy. fish-audio-s2-pro ~3.5% WER / 1.2% CER (English, vendor figure).
  • TTFA (time-to-first-audio) — the latency metric that matters for real-time UX; Sonic 3.5 ~82ms, Deepgram Aura-2 <90ms, misotts claimed 110ms.

Caveats (why these are snapshots)

  • Rankings shift weeklytts-models-2026-benchmark stresses leaderboard positions are dated snapshots, not fixed truth (this wiki dates every figure).
  • Sources “disagree” mostly because they are dated differently. This wiki recorded Gemini at 1216 vs 1217 and Fish S2 Pro at 1123 vs 1128 and filed it as methodological disagreement between vote pools. The 2026-08-03 refresh argues it was mostly drift: over two months Gemini moved 1217 → 1212, Fish S2 Pro 1128 → 1123, Inworld’s Realtime TTS-2 fell 1206 → 1189. Single-digit gaps between sources published days apart are the board moving, not the boards differing. Real methodological difference would show as ordering changes, and those happened too — so date every figure and compare orderings, not points.
  • Vendor-reported numbers (latency, WER) have incentives; treat as indicative.
  • Metrics are partial: Elo ≠ accuracy ≠ latency. A model wins or loses per axis.

tts-arena-leaderboard · text-to-speech · open-weight-tts · tts-models-2026-benchmark · mean-opinion-score · speech-audio-ai