Spokes.wiki Search About
Dataset source ↗ source url updated Mon Aug 03 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Artificial Analysis — Text-to-Speech Leaderboard

The neutral, Elo-based TTS leaderboard (the speech analog of an LLM leaderboard), ranking text-to-speech models by blind human preference in a “Speech Arena”: listeners compare two samples and vote which sounds more natural. The wiki’s anchor for tts-benchmarks. Every number here is a dated snapshot — the board churns weekly.

Top of the board — 2026-08-03 snapshot (all closed/API-only)

#ModelProviderElo$/1M chars
1Simba 3.2Speechify1229$10.0
2Qwen-Audio-3.0-TTS-PlusAlibaba1227$27.6
3gemini 3.1 Flash TTSGoogle1212$18.3
4Luna TTSVUI Labs1205$80.0
5Sonic 3.5Cartesia1202$49.0
6StepAudio 2.5 TTSStepFun1199$85.0
7Realtime TTS 1.5 MaxInworld1196$26.0
8Lightning V3.1 ProSmallest.ai1190$19.5
9Realtime TTS-2 (research preview)Inworld1189$20.8
10Eleven v3ElevenLabs1173$100.0

What moved since the 2026-06-05 snapshot

The prior snapshot is kept rather than overwritten, because the movement is the finding.

ModelJun 05Aug 03Δ
Alibaba’s leader (Fun-Realtime-TTS → Qwen-Audio-3.0-TTS-Plus)1227 (#1)1227 (#2)renamed/relined, same score
gemini 3.1 Flash TTS1217 (#2)1212 (#3)−5
Sonic 3.51206 (#4)1202 (#5)−4
Realtime TTS 1.5 Max1199 (#5)1196 (#7)−3
Realtime TTS-21206 (#3)1189 (#9)−17, six places

A new leader, and it is the one the corpus wrote off. Speechify’s Simba tops the board at 1229. Two months ago tts-models-2026-benchmark listed “Speechify SIMBA 3.0” in a trailing clause as the budget option. It is now #1 at $10/1M characters — the cheapest of the entire top ten, against ElevenLabs v3 at $100 for tenth place. Whatever the board is measuring, price and rank have come apart completely at the top.

Also new since June: Luna TTS (VUI Labs) and Smallest.ai’s Lightning enter the top ten; StepFun’s StepAudio appears as a closed entry at 1199, where June’s board had “Step Audio EditX” as the #2 open model at 1112.

Highest open-weight models — and their true rank

ModelEloBoard rank$/1M
fish-audio-s2-pro112320$15.0
kokoro 82M v1.0105548$0.7
OpenVoice v295177$8.3

The board rank is the number this page failed to record in June, and it changes the picture. Saying “no open-weight model cracks the leaders” understated it: the best open model sits twentieth, and the efficiency favourite sits forty-eighth. Both drifted down over the two months (fish-audio-s2-pro 1128 → 1123, kokoro 1064 → 1055) while the top of the board rose, so the gap to the leader widened from 99 to 106 Elo. See open-weight-tts.

Kokoro’s price is listed as $0.7/1M characters here, against the $0.65 this wiki recorded in June — a rounding-level difference, noted rather than resolved.

Method & caveats

Elo from blind A/B votes in the Speech Arena. Reflects perceived naturalness, not WER/latency — pair with the tts-benchmarks page’s other axes (CER, MOS, TTFA). Vote populations and sample sets bias results; treat as indicative. Cross-source check: the MarkTechPost survey tts-models-2026-benchmark cites slightly different Elo values (e.g. Gemini 1216 vs 1217 in June) — expected snapshot drift, flagged in synthesis. That article is a fixed 2026-05-30 publication, so its figures are superseded by this snapshot rather than disagreeing with it.

Method & caveats

Elo from blind A/B votes in the Speech Arena. Reflects perceived naturalness, not WER/latency — pair with the tts-benchmarks page’s other axes (CER, MOS, TTFA). Vote populations and sample sets bias results; treat as indicative. Cross-source check: the MarkTechPost survey tts-models-2026-benchmark cites slightly different Elo values (e.g. Gemini 1216 vs 1217 here) — expected snapshot drift, flagged in synthesis.

tts-benchmarks · text-to-speech · open-weight-tts · fish-audio-s2-pro · kokoro · tts-models-2026-benchmark