Artificial Analysis — Text-to-Speech Leaderboard
The neutral, Elo-based TTS leaderboard (the speech analog of an LLM leaderboard), ranking text-to-speech models by blind human preference in a “Speech Arena”: listeners compare two samples and vote which sounds more natural. The wiki’s anchor for tts-benchmarks. Every number here is a dated snapshot — the board churns weekly.
Top of the board — 2026-08-03 snapshot (all closed/API-only)
| # | Model | Provider | Elo | $/1M chars |
|---|---|---|---|---|
| 1 | Simba 3.2 | Speechify | 1229 | $10.0 |
| 2 | Qwen-Audio-3.0-TTS-Plus | Alibaba | 1227 | $27.6 |
| 3 | gemini 3.1 Flash TTS | 1212 | $18.3 | |
| 4 | Luna TTS | VUI Labs | 1205 | $80.0 |
| 5 | Sonic 3.5 | Cartesia | 1202 | $49.0 |
| 6 | StepAudio 2.5 TTS | StepFun | 1199 | $85.0 |
| 7 | Realtime TTS 1.5 Max | Inworld | 1196 | $26.0 |
| 8 | Lightning V3.1 Pro | Smallest.ai | 1190 | $19.5 |
| 9 | Realtime TTS-2 (research preview) | Inworld | 1189 | $20.8 |
| 10 | Eleven v3 | ElevenLabs | 1173 | $100.0 |
What moved since the 2026-06-05 snapshot
The prior snapshot is kept rather than overwritten, because the movement is the finding.
| Model | Jun 05 | Aug 03 | Δ |
|---|---|---|---|
| Alibaba’s leader (Fun-Realtime-TTS → Qwen-Audio-3.0-TTS-Plus) | 1227 (#1) | 1227 (#2) | renamed/relined, same score |
| gemini 3.1 Flash TTS | 1217 (#2) | 1212 (#3) | −5 |
| Sonic 3.5 | 1206 (#4) | 1202 (#5) | −4 |
| Realtime TTS 1.5 Max | 1199 (#5) | 1196 (#7) | −3 |
| Realtime TTS-2 | 1206 (#3) | 1189 (#9) | −17, six places |
A new leader, and it is the one the corpus wrote off. Speechify’s Simba tops the board at 1229. Two months ago tts-models-2026-benchmark listed “Speechify SIMBA 3.0” in a trailing clause as the budget option. It is now #1 at $10/1M characters — the cheapest of the entire top ten, against ElevenLabs v3 at $100 for tenth place. Whatever the board is measuring, price and rank have come apart completely at the top.
Also new since June: Luna TTS (VUI Labs) and Smallest.ai’s Lightning enter the top ten; StepFun’s StepAudio appears as a closed entry at 1199, where June’s board had “Step Audio EditX” as the #2 open model at 1112.
Highest open-weight models — and their true rank
| Model | Elo | Board rank | $/1M |
|---|---|---|---|
| fish-audio-s2-pro | 1123 | 20 | $15.0 |
| kokoro 82M v1.0 | 1055 | 48 | $0.7 |
| OpenVoice v2 | 951 | 77 | $8.3 |
The board rank is the number this page failed to record in June, and it changes the picture. Saying “no open-weight model cracks the leaders” understated it: the best open model sits twentieth, and the efficiency favourite sits forty-eighth. Both drifted down over the two months (fish-audio-s2-pro 1128 → 1123, kokoro 1064 → 1055) while the top of the board rose, so the gap to the leader widened from 99 to 106 Elo. See open-weight-tts.
Kokoro’s price is listed as $0.7/1M characters here, against the $0.65 this wiki recorded in June — a rounding-level difference, noted rather than resolved.
Method & caveats
Elo from blind A/B votes in the Speech Arena. Reflects perceived naturalness, not WER/latency — pair with the tts-benchmarks page’s other axes (CER, MOS, TTFA). Vote populations and sample sets bias results; treat as indicative. Cross-source check: the MarkTechPost survey tts-models-2026-benchmark cites slightly different Elo values (e.g. Gemini 1216 vs 1217 in June) — expected snapshot drift, flagged in synthesis. That article is a fixed 2026-05-30 publication, so its figures are superseded by this snapshot rather than disagreeing with it.
Method & caveats
Elo from blind A/B votes in the Speech Arena. Reflects perceived naturalness, not WER/latency — pair with the tts-benchmarks page’s other axes (CER, MOS, TTFA). Vote populations and sample sets bias results; treat as indicative. Cross-source check: the MarkTechPost survey tts-models-2026-benchmark cites slightly different Elo values (e.g. Gemini 1216 vs 1217 here) — expected snapshot drift, flagged in synthesis.
Related
tts-benchmarks · text-to-speech · open-weight-tts · fish-audio-s2-pro · kokoro · tts-models-2026-benchmark