Spokes.wiki Search About
Defined Term concept updated Mon Aug 03 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Open-weight TTS

The segment of text-to-speech models whose weights are downloadable and self-hostable — the speech analog of llm-providers-wiki’s open-weight-models. Defined by two facts in 2026:

1. It trails the closed frontier, and by more than this page used to say

No open-weight model cracks the tts-arena-leaderboard top tier (all closed/API-only). The open leader is fish-audio-s2-pro (Elo 1123, 5B), and the 2026-08-03 board puts a number on “below” that this page previously lacked: it ranks 20th overall, with kokoro 48th, against a closed leader at 1229.

This section used to read “but not by much.” That was wrong, and it got wronger. Over the summer the gap to the leader widened from 99 to 106 Elo — the top of the board rose while both open models drifted down (fish-audio-s2-pro 1128 → 1123, kokoro 1064 → 1055). Reporting only the Elo delta and never the rank made a twentieth-place finish sound like a near miss. The proprietary-premium-vs-open-wedge dynamic llm-providers-wiki tracks for text does hold for voice, but the wedge is thinner here than the earlier phrasing implied.

2. License is a first-class axis, not a footnote

The open field splits on terms:

  • Permissive (Apache-2.0 / MIT): kokoro, orpheus, sesame-csm, Chatterbox, Dia, Higgs Audio V2 — freely commercial-usable open-source-tts-models. Coqui’s XTTS belongs here too: a peer-reviewed, publicly released zero-shot cloner across 16 languages (xtts-paper) — the open field’s evidence that even voice-cloning ships without a research-license gate.
  • Research-only / paid-commercial: fish-audio-s2-pro (the highest-Elo open model is not freely commercial) and misotts‘s “modified MIT.” So the practical “best open model” depends on whether you can use it commercially, not just its score.

How lopsided the split actually is (2026-08-03). A 36-model census of the TTS field (free-voice-clone-list, April 2026 snapshot) counts 27 Apache-2.0, 4 MIT, 1 MIT/Apache-2.0, 2 under Liquid’s LFM licence, 1 Research License (fish-audio-s2-pro) and 1 OpenRAIL-M (supertonic-2). The restrictive corner is not a wing of the field, it is two models — but one of them is the quality leader, which is precisely why counting licences cannot settle the ceiling question in synthesis. Permissive availability is near-total; permissive at the top is still unproven on this branch.

The shape of the field

Efficiency-first (kittentts 15M and supertonic-2 66M, then kokoro 82M — the last two without voice-cloning) → all-rounder (Chatterbox 0.5B) → expressive/streaming (orpheus) → conversational (sesame-csm) → emotive (misotts) → top-quality but restricted (fish-audio-s2-pro). Many are Llama-based — a bridge to the text-LLM market (gemini cross-wiki).

text-to-speech · tts-benchmarks · tts-arena-leaderboard · kokoro · kittentts · supertonic-2 · fish-audio-s2-pro · open-source-tts-models · free-voice-clone-list · xtts-paper · voice-cloning · open-weight-models