Spokes.wiki Search About

speech-audio-wiki

log

Synthesis — Speech & Audio AI

The evolving thesis. Spun out as tts-wiki from the hub _inbox speech-audio-models cluster on 2026-06-05 — seeded by the parked misotts and grown with three TTS field sources the router curated at the user’s request — then broadened to all speech & audio AI and renamed speech-audio-wiki the same day, and immediately grown with four more router-curated sources covering STT/ASR and audio/music generation. The wiki now spans the three production branches of speech-audio-ai — synthesis, recognition, generation — plus an emerging fourth, audio comprehension (audio-flamingo-3); TTS is still the deepest corner.

Current thesis

The 2026 speech & audio AI market — across text-to-speech (synthesis), speech-to-text (recognition), and audio-music-generationrhymes with the text-LLM market (llm-providers-wiki): a closed frontier leading on the headline metric, an open-weight field competing on cost, control, and licensing. The same four dynamics recur in all three branches (see speech-audio-ai):

  1. A closed frontier on the headline metric — and on TTS the lead is not thin. Proprietary leaders top each branch’s ranking: TTS Elo (tts-arena-leaderboard, 2026-08-03: Speechify Simba 3.2 at 1229, Alibaba’s Qwen-Audio-3.0-TTS-Plus 1227, gemini 3.1 Flash TTS 1212), STT WER (ElevenLabs Scribe ~3.3% EN, Deepgram 5.26% — stt-apis-comparison), music Elo (music-arena-leaderboard, 2026-08-03: suno V5.5 at 1189 instrumental / 1171 vocals — the earlier “~1293” came from a T4 survey on a different scale). Corrected 2026-08-03: this bullet used to say the open field was “close behind and closing.” On STT that holdscanary-qwen at 5.63% nearly matches the commercial APIs, and four Apache-2.0 models now sit below it. On TTS it does not: the open leader fish-audio-s2-pro ranks 20th overall and kokoro 48th, and over the summer the gap to the leader widened from 99 to 106 Elo. The wiki had been reporting Elo deltas without ranks, which made twentieth place read like a near miss. The branches diverge, and the earlier sentence generalized the recognition branch’s good news across all three.
  2. License/rights — not just the score — decides, and the axis has three positions (revised 2026-08-03). The constraint sharpens as you move across branches: TTS has a research-license ceiling (fish-audio-s2-pro is the top open model but paid-commercial); STT is mostly permissive (whisper MIT, Canary CC-BY); music escalates to outright copyright war (ai-music-copyright: suno in Sony litigation vs the clean udio / stable-audio). The best-sounding option is often the least legally safe — but mureka, the quality runner-up, fits neither pole: it asserts full commercial rights and discloses no training-data provenance, no rights deals and no operating company. Contested and licensed were the two positions this wiki knew; undisclosed is a third, and a buyer comparing on legal safety cannot place it at all.
  3. “Best” is multi-axis and weekly-volatile. Each branch sorts by its own binding constraint — quality (Elo/MOS), accuracy (WER/CER), latency (TTFA / streaming RTFx), languages, capabilities, cost (tts-benchmarks). “No single model wins.” Every figure here is a dated snapshot.
  4. Convergence on the LLM stack. TTS rides Llama backbones + RVQ neural-audio-codec; STT’s accuracy leaders are SALM models bolting an LLM decoder onto a speech encoder (canary-qwen = FastConformer + Qwen3; Granite-Speech; Qwen3-ASR). “Speech as language modeling” in both directions — the structural bridge to llm-providers-wiki (gemini straddles). A fourth branch — audio comprehension — now completes the pattern: audio-flamingo-3 (NVIDIA, a fully-open Large Audio-Language Model) doesn’t transcribe or synthesize, it reasons over speech, sound, and music (unified AF-Whisper encoder, on-demand chain-of-thought, ~10-min audio), SOTA on 20+ understanding benchmarks. Where codec-token TTS and SALM STT were the production directions of “the LLM eats audio from both ends,” AF3 is the pure comprehension vertex — and it re-confirms dynamic 2 (license, not score, decides): open weights and data, but a non-commercial research license, the same shippability ceiling as fish-audio-s2-pro and musicgen.

Unifying tension: as in text, capability concentrates at a few closed providers while cost and access are democratized from below by open weights — but audio adds two twists text rarely faces: latency (TTFA / RTFx) as a make-or-break product axis, and rights (research-license ceilings, voice-cloning consent, music copyright) as a constraint that can outweigh quality outright.

Recurring reads

  • Efficiency vs. controllability split — the smallest models (kokoro 82M) drop voice-cloning to hit footprint; larger models add cloning/emotion/streaming. Capability tends to cost parameters. (STT mirror: distilled/streaming models like Parakeet/Distil-Whisper trade some accuracy for huge RTFx.)
  • The LLM is eating audio from both ends — codec-token TTS (neural-audio-codec: misotts‘s Mimi, fish-audio-s2-pro‘s dual-AR+RVQ) and SALM STT (canary-qwen‘s Qwen3 decoder). Both directions are now “next-token over audio/text,” importing llm-providers-wiki’s architectures. The STT baseline they build past, whisper, is now primary-grounded (whisper-paper, Radford et al. 2022): its dominance traces to weak-supervision-at-scale (680k hours, zero-shot), not to topping WER — which is exactly why the SALM successors can beat it on accuracy yet not displace it.
  • Latency is a product axis — TTFA (~82ms Sonic, ~90ms Aura-2, 110ms MisoTTS) and STT RTFx (Parakeet >2,000×) gate real-time agents — a constraint with no clean text-LLM analog.
  • The open wedge appears in every branch — but not equally, and in music barely at all (revised 2026-08-03). kokoro/fish-audio-s2-pro (TTS) and whisper/canary-qwen (STT) are real alternatives: usable, self-hostable, a known distance behind. Music is different in kind. music-arena-leaderboard puts musicgen last of eighteen (Elo 865, 300+ behind the leader) and stable-audio 2.0 fifteenth. The self-hostable option here isn’t trading polish for control; it is off the competitive board entirely. Ranking the branches by how real the open wedge is — STT strongest, TTS middling (open leader 20th), music negligible — tracks how hard the output is to evaluate automatically, which is worth watching as a hypothesis rather than a finding. musicgen keeps its role regardless of its rank: a single-stage transformer over EnCodec RVQ tokens (the music-branch instance of the LLM-stack convergence) under MIT code / CC-BY-NC weights — open-weight yet not shippable, the music echo of TTS’s research-license ceiling, and with a clean-rights training set that keeps it clear of the litigation hitting suno. It is the branch’s reference implementation and its rights-clean option; it is not a competitive one.
  • Compression is now its own STT claim, and it doesn’t degrade evenly. The small/local pole used to be about choosing a small model (vosk, moonshine) or a faster runtime (whisper-cpp). vibevoice-asr-bitnet instead compresses a large one and sells the recipe — 1.58 GB from 4.62, ternary-format decoder weights plus 8-bit tokenizer, served by llama.cpp, real-time on 3 CPU threads and claimed 1.86× whisper.cpp. What its own table shows is that quantization damage is not uniform: most benchmarks lose 0.2–0.6 WER points, but the far-field Chinese meeting set loses 7.6 and the lower-resource languages 1.3–2.2. Compression costs least where the task is easy and the data is thick — which is the opposite of where edge deployment usually needs it to hold up.
  • The closed-frontier archetype is elevenlabs — the commercial leader the spoke keeps citing (Scribe WER, TTS Elo) and the rare provider spanning all branches (TTS + voice cloning + dubbing
    • Scribe STT + Eleven Music), at an $11B valuation (Feb 2026). It sits on both sides of the rights axis: its cloning powers audio-deepfake fraud, yet it ships an AI Speech Classifier to detect synthetic audio.

What the word-error rates were measured on

The three corpora every comparison here quotes now have pages, and reading them together changes how the spoke should treat its own numbers.

All three are read speech. librispeech is audiobooks, common-voice is prompts read into a laptop microphone, fleurs is translated sentences read aloud in 102 languages. Nothing in the spoke’s evidence base measures spontaneous conversation, overlapping talkers, or a microphone across a room — which is what the products these numbers sell are used for. The WERs are real and they are all from one register.

“Clean” is a circular label. LibriSpeech’s clean/other split ranks speakers by the error rate of an acoustic model trained on WSJ si-84 and cuts at the median: “the lower-WER speakers designated as ‘clean’ and the higher-WER speakers designated as ‘other’.” So test-clean is the half of the speakers that a 2015 system already handled. Improvements reported there are improvements on material pre-selected for tractability, which is worth remembering next to the observation in open-asr-models-2026-comparison that the top ten models sit inside one WER point.

The licence that makes Common Voice useful also makes it suspect as a test set. CC0 means it is in nearly every open model’s reach, and the spoke has no source measuring how much of it ended up in training data. A Common Voice score is uninterpretable without that.

This is the same shape as the spoke’s standing complaint about vendor latency benchmarks — the number is precise, the conditions it was taken under are the part nobody states.

Open questions

  • Does the open ceiling go permissive? If a future Apache-2.0/MIT model matches fish-audio-s2-pro‘s Elo, the research-license ceiling collapses — the key thing to watch. Unchanged on TTS after the 2026-08-03 board refresh: fish-audio-s2-pro (research licence) is still the top open model at 1123, with the best permissive model well behind it. Two months passed and the ceiling did not move — which is itself a data point, since the licence census in free-voice-clone-list shows 27 of 36 open TTS models are Apache-2.0. Permissive weights are abundant; permissive quality is not. Answered on STT (2026-07-24), still open on TTS and music. The open STT board is now led by Apache-2.0 models — ARK-ASR-3B 5.04, Granite Speech 4.1 5.33, Cohere Transcribe 5.42, MOSS-Transcribe-preview under 5.33 — all ahead of the CC-BY-4.0 canary-qwen at 5.63 (open-asr-models-2026-comparison). Microsoft’s vibevoice-asr-bitnet ships MIT. So on the recognition branch, permissive licensing and the accuracy lead now sit on the same models, and the roundup treats licence as the first filter a buyer applies rather than a tiebreak. The restrictive-licence ceiling this question was written about is a TTS and music phenomenon now, not a branch-wide one.
  • How real are the latency claims? TTFA figures are largely vendor-reported under unstated conditions; a neutral latency benchmark would be high-value. Unchanged but now worse-behaved (2026-08-03): supertonic-2 publishes a table claiming 12,164 chars/sec at RTF 0.001 on an RTX 4090, 42× elevenlabs Flash v2.5 — comparing a local GPU against a hosted API, conditions unstated, and putting kokoro at RTF 1.3 (slower than realtime), which nothing else here reports. Another vendor table, not an answer.
  • Remaining coverage gaps: safety/consent for voice-cloning now sourced (2026-06-09): audio-deepfake adds the misuse/consent/detection axis (CEO-voice fraud, fake-Biden robocalls, non-consensual cloning; ASVspoof detection + SynthID watermarking + the FCC robocall ban) — the dark mirror of the cloning capability, and the lived form of “rights can outweigh quality.” Also added wavenet (DeepMind 2016), the origin point of neural audio the whole stack descends from — its sample-by-sample latency problem is the first instance of the spoke’s TTFA tension. Still thin: non-English depth, audio understanding beyond transcription (audio LLMs), and a neutral latency benchmark. Sharpened 2026-08-03: the safety gap is not only about harms and detection but about what the supply side omitsfree-voice-clone-list distributes 53 cloning-capable models with zero consent, watermarking or misuse language anywhere in it. What is still missing is any source describing a catalogue, hub or model card that does carry such a line, which would turn this from an observed silence into a comparison.
  • Does the open frontier flip? Watch whether a permissive open model takes a branch’s headline metric — closest on STT (canary-qwen 5.63% vs commercial ~3–5%), furthest on music (closed suno well ahead). The TTS research-license ceiling (fish-audio-s2-pro) is the other thing to watch.
  • Snapshot drift: rankings disagree across sources (TTS Elo: Gemini 1216 vs 1217; STT WER varies by test set) — treat all as dated. Largely explained 2026-08-03: refreshing tts-arena-leaderboard after two months showed the board itself moving several points per model (Gemini 1217 → 1212, Fish S2 Pro 1128 → 1123, Inworld Realtime TTS-2 1206 → 1189, a six-place fall). Single-digit gaps between sources published days apart are drift, not methodological disagreement. What is worth tracking is ordering change, and there the board turned over properly: a new #1, two new entrants in the top ten, and Alibaba’s leader renamed.
  • Where does distillation put the open TTS ceiling? (added 2026-07-29) luxtts is a 4-step distillation of ZipVoice with a swapped 48 kHz vocoder — Apache-2.0, 1 GB VRAM, 150× realtime, cloning from three seconds, 4.9k★. If a distilled community derivative can reach the quality of the models it was cut down from, the “does the open ceiling go permissive?” question stops being about whether a lab releases permissive weights and starts being about whether anyone needs them to: a permissive base plus a distillation pass produces a permissive fast model without the original licence-holder’s participation. LuxTTS asserts parity with “models 10x larger” and publishes no MOS, WER or Elo, so it is a live version of the question rather than an answer. Getting it onto tts-arena-leaderboard would settle more than another vendor claim.

Growth edges

Ranked; each names the kind of source that would close it (see ../QUALITY.md → Growth edges).

  1. A neutral latency benchmark. Every latency figure in this spoke — TTFA, RTFx, “ultra-low latency”, Cerebras-class claims — is a vendor number measured under unstated conditions. It is the axis the spoke calls make-or-break and the only major one with no independent board, where TTS Elo and STT WER both have one. — needs: a T1/T2 independent latency measurement across providers with stated hardware and concurrency · hunted 2026-08-08 — nothing ≥ bar. The whole result set is the vendor-comparison genre: Deepgram, Gladia and Inworld each publish a “best STT APIs 2026” page on which they rank well. The closest thing to a third party is Coval, an eval-tooling company whose own headline for the sibling piece is “why vendor benchmarks lie” — and which still states no hardware and no concurrency, the two conditions this edge exists to demand. The figures circulating (ElevenLabs Scribe v2 Realtime “under 150 ms”, Inworld “~92 ms time-to-first-token”) are clean-studio marketing numbers, which is the same defect as supertonic-2‘s table, not its cure. Not re-hunted before 2026-08-22.
  2. One neutral board per branch, held permanently. The music corner ran two months on a T4 marketing survey’s cast list. TTS and STT have boards; audio understanding now has none, and it is the branch growing fastest. — needs: a T2 leaderboard or independent evaluation covering LALMs and fixed-head understanding models on the same tasks.
  3. Anything measuring an assembled loop rather than a component. The corpus grades models; the artifacts that matter are S2ST pipelines, conversational agents and on-device stacks. No metric here applies to them. — needs: a T1/T2 end-to-end evaluation of a composite speech system.
  4. The distilled/derivative tier, measured. luxtts claims parity with “models 10x larger” and publishes nothing; the benchmark pages cover labs, not derivatives, exactly as derivatives become the interesting artifacts. — needs: any independent measurement of a distilled open TTS model.

Coverage edges (added 2026-08-08, at the curator’s request for a wider backlog). These widen what the spoke covers instead of answering an open question above; one ordinary solid source closes any.

  1. The pipeline around the model. Diarization, voice activity detection and forced alignment are what a real transcript needs beside whisper or vosk, and the spoke holds models only. — needs: the pyannote paper or an equivalent, T1.
  2. The corpora. CLOSED 2026-08-09 (research pass) — librispeech, common-voice and fleurs written from their own papers, all T1. Successor, and it is a real gap: all three are read speech, so every WER in this wiki was earned on someone reading aloud, while the products those numbers sell are used on meetings, calls and far-field capture. — needs: a conversational or far-field benchmark with published WERs — CHiME, AMI, Switchboard or an equivalent — plus, separately, any measurement of Common Voice contamination in open models’ training sets.
  3. Wake words and keyword spotting. The always-on tier below full transcription, running on a microcontroller budget, is absent. — needs: a paper or an on-device toolkit’s docs. Cross-spoke: the hardware side is embedded-iot-wiki.
  4. The real-time plumbing. Growth edge 1 wants a neutral latency benchmark for a loop that is mostly transport — WebRTC, echo cancellation, noise suppression, jitter buffers — and no page describes any of it. — needs: the WebRTC specs or a measured deployment.

The unit is the loop, not the model, and the loop is built at every scale

The three branches aren’t only sold separately — they compose, and tracking that composition has turned out to matter more than tracking any single model’s score. The pattern shows up three times over, at three different scales, and the sections below were written a month apart before it was clear they were one argument.

The composite task

The three branches aren’t only sold separately — they compose. speech-to-speech-translation (S2ST) is the first applied/composite task in the wiki: STT + machine translation + TTS in one pipeline. Its 2026 exemplar, gemini-live-3-5-translate (Google; 70+ languages, 2000+ pairs, near-real-time, voice-preserving, SynthID-watermarked), is the purest case of “the LLM eats audio from both ends” — both ends in one streamed, end-to-end model. Two threads it sharpens: (1) latency becomes a quality trade-off, not just a number — simultaneous translation must choose between waiting for context and translating immediately; (2) provenance/watermarking (SynthID on all output) enters the wiki as an audio-safety axis (alongside voice-cloning consent and music copyright). It also extends the closed-frontier read: gemini now leads (by its own framing) at TTS, plays at STT, and stakes the S2ST frontier — a single proprietary provider spanning all of it.

Second composite (2026-06-13): the real-time conversational agent. elevenlabs-expressive-mode (ElevenLabs’ Conversational AI) is the dialogue sibling of S2ST’s translation loop: Scribe v2 Realtime STT → reasoning → Eleven V3 Conversational TTS in one low-latency turn-by-turn loop. It adds a genuinely new axis to the spoke — emotion/affect: the agent infers emotion from prosody (pitch, pacing, exclamations) and applies “tone cue cards,” making paralinguistic understanding the headline differentiator rather than fidelity or WER. Two thesis hooks: (1) it confirms the branches compose into products — leaders increasingly sell the integrated loop, not components; (2) another “ultra-low latency” claim with unstated conditions (T3 vendor page) — more weight on the still-open call for a neutral latency benchmark. Caveat: a marketing landing page; CSAT/latency figures unverified.

Third entrant (2026-06-14): a single-model frontier voice agent. gpt-realtime-2 (OpenAI’s Realtime-API voice model, “GPT-5-class reasoning,” speech-to-speech over WebRTC) is the conversational loop collapsed into one end-to-end model — the gemini-live-3-5-translate shape aimed at general dialogue rather than translation, and the contrast to ElevenLabs’ assembled STT→LLM→TTS pipeline. It pushes the “the LLM eats audio from both ends” read to its limit: a speech-to-speech network carrying the reasoning of the text frontier, not just its tokens. Its document-context feature also adds a grounding axis to spoken dialogue (answer by voice about pasted text) — retrieval entering the voice loop. So the real-time-voice frontier now has three closed entrants (Google, ElevenLabs, OpenAI). Caveat: secondhand practitioner source (T4), no latency/quality numbers.

The same loop, disaggregated onto the user’s own hardware

Every source above assembles the loop in the cloud. The consumption layer builds it from the other direction — locally, out of parts anyone can download — and the binding constraints flip with it.

Every source so far looked at speech/audio AI as models and their metrics — WER, Elo, latency, license. vocalinux (open-source Linux voice dictation) opens a different layer: the end-user application that assembles those engines into something you use, where the binding constraints flip from accuracy to privacy, offline operation, and OS integration. It’s a wrapper, not a recognizer — it runs a chosen local backend (whisper-cpp, whisper, or vosk) and spends its real engineering on injecting text under both X11 and Wayland. Three threads fold in. (1) It gives the recurring open wedge a concrete consumer end-point: the whole value proposition is no cloud, no subscription — private dictation built entirely from off-the-shelf open engines, the thing the open-weight field makes possible that the closed frontier structurally can’t. (2) It surfaces a small/local pole the WER pages omit — vosk (Kaldi, ~50 MB, 20+ languages, runs on a Pi) is chosen for constraint, not accuracy, and whisper-cpp (quantization applied to ASR) is what makes Whisper responsive on a laptop; “best” here means runs on my hardware, offline, a fifth sorting axis under the speech-to-text leaderboard. VOSK also sits at the clean end of the license axisApache-2.0, so unlike the TTS research-license ceiling, the small-STT corner’s constrained option is also the freely-shippable one; here the recurring “best-sounding is least legally safe” tension doesn’t bite. (3) It’s the inverse of the frontier voice agents — where gpt-realtime-2 and gemini-live-3-5-translate collapse the loop into one cloud model, Vocalinux disaggregates it onto the local desktop for privacy. Caveat: single T3 product write-up; the app inherits whatever accuracy its backend has, so there’s no new benchmark here, only a new use.

The pole extends to the microcontroller (added 2026-07-23). moonshine-pico-voice-toolkit pushes the small/local pole one order of magnitude smaller than vocalinux: Moonshine AI‘s micro toolkit runs a whole offline voice loop — VAD → SpellingCNN STT → neural TTS — on a Raspberry Pi Pico 2 W in ~3.6 MiB flash / 468 KiB SRAM. It sharpens the “best = fits my hardware” axis into a design principle: at MCU scale you don’t ship open-vocabulary ASR, you ship a retrainable 50-token command recognizer that fits, and you accept a bounded vocabulary as the price of no cloud, no OS, no network. Two firsts for the spoke fold in — TTS on an MCU and on-device VAD — and moonshine (previously only the “edge, from 27M params” mention in open-source-stt-models) becomes a paged subject anchoring the sub-laptop end of the size axis, below vosk/whisper-cpp. It also flips the vocalinux “assemble off-the-shelf engines” pattern: micro is a purpose-built tiny stack, three models co-designed to fit, not a wrapper choosing among general backends. Same T3 caveat, harder here — the impressive claim is the fit (concrete), while recognition quality on a 50-token SpellingCNN at this size is entirely unmeasured in the source.

The loop collapsed into one pass

The frontier collapses the loop into one model; the edge shrinks it until it fits. There is a third move: collapse it into one pass, and fix at training time which questions it will answer.

sensevoice (Alibaba/FunAudioLLM) returns ASR, spoken-language ID, a seven-category emotion label and audio-event tags from a single non-autoregressive pass over the audio. It does the job this wiki opened with audio-flamingo-3 — treat sound as something to interpret, not only to transcribe — but with fixed task heads instead of a language model, and it is fast for exactly that reason: no token-by-token decode. The README claims more than 5x Whisper-Small and 15x Whisper-Large at comparable size.

That splits the audio-understanding branch in two, and the split is architectural rather than a matter of scale. A LALM can answer a question nobody anticipated; SenseVoice answers four that were fixed at training time, cheaply enough to run over everything. Neither replaces the other, and the spoke should stop treating “audio understanding” as one shelf.

It also sharpens how the STT corner is graded here. Every model on that shelf competes on word error rate and on how small it gets. SenseVoice competes on how many questions one pass answers, which no benchmark in this corpus measures. The emotion output is the weakest part of it: seven discrete categories, no per-category accuracy, no speaker demographics, no cross-corpus evaluation — the exact conditions under which speech-emotion results are known not to transfer. The audio-event labels are the defensible half.

Two openness caveats join the standing pattern. MIT code with the weights under the FunASR licence is another instance of the split-licence shape the spoke keeps recording, and only SenseVoiceSmall is released while the benchmark tables include a Large nobody outside can run.

Read together, the three scales say the same thing about how this spoke should grade work. A model card reports WER, Elo or latency for one component. None of the artifacts above is one component, and the corpus has no measurement that applies to an assembled loop — which is why the open questions below keep asking for a neutral latency benchmark and never get one.

Openness arrives three ways, and the safety machinery rides on none of them

The spoke started by treating openness as something a lab grants: a licence attached to weights at release, with the standing question being whether the permissive tier can reach the quality tier (open-weight-tts, fish-audio-s2-pro‘s research-licence ceiling). Two later sources show that this is only the first of three routes, and that what fails to travel along any of them is the same thing.

Route two — the census, which counts availability rather than quality

audio-deepfake has carried the safety axis since June, and every response on it sits downstream of the models: detection (ASVspoof), watermarking (SynthID), the FCC robocall ban. free-voice-clone-list shows what the upstream looks like. It is a 59 KB catalogue of 53 open audio models, 35 of its 36 TTS entries advertising cloning from a three-to-ten-second reference sample, and it contains no occurrence of consent, ethics, watermark, misuse, deepfake, responsible, abuse or disclaimer. Zero, not few.

This wiki had already flagged the same absence on luxtts five days earlier, on a project shipping three-second cloning. One project is an oversight; a catalogue of 53 is the genre’s default. So the spoke’s safety axis gains a second half: the harms are documented in journalism, regulation and detection research, and the layer that puts the capability in a user’s hands carries no corresponding sentence. The one licence in the catalogue with behavioural use restrictions (OpenRAIL-M) belongs to the one model that cannot clone a voice — a single data point, noted rather than built on.

Two other things the census settles, both about the shape of the open field rather than its ceiling:

  • Permissive licensing is near-total by count. 27 of 36 TTS entries are Apache-2.0, 5 more MIT or MIT/Apache. The restrictive corner open-weight-tts tracks is two models — and one of them is fish-audio-s2-pro, the quality leader. Availability is overwhelmingly permissive; the ceiling question stays open precisely because a census counts models rather than ranking them.
  • Cloning stopped being a dividing line. voice-cloning was written around the observation that cloning correlates with size and the smallest models drop it. At 35 of 36, it is table stakes, and kittentts claims it at 15M parameters — a fifth of kokoro‘s footprint. The small pole now runs 15M (kittentts) and 66M (supertonic-2) below the 82M this wiki has called the efficiency leader since June.

The source’s own weakness is the caveat on all of it: no MOS, no WER, no Elo, no methodology, one pseudonymous maintainer (0xsojalsec), and a last push of 2026-04-09 read four months later, in a field this wiki’s conventions call weekly-volatile.

Route three — distillation, which needs no release decision at all

The spoke has tracked openness as a property labs grant: a licence attached to weights at release, with the standing question being whether the permissive tier can reach the quality tier (open-weight-tts, fish-audio-s2-pro‘s research-licence ceiling). luxtts arrives by a different route. It is not a released model; it is someone else’s model, distilled to four steps and fitted with a different vocoder, published Apache-2.0 by a person who did not train the original.

Three consequences worth holding:

Capability propagates without a release decision. k2-fsa’s ZipVoice licence permitted the derivative; the derivative is faster, higher-fidelity at 48 kHz, and now has more community tooling than the base. Openness here is downstream of what the base licence allowed, not of what any lab chose to promote.

The safety infrastructure does not propagate with it. audio-deepfake records SynthID watermarking and ASVspoof detection as the countermeasures to non-consensual cloning. Those attach to specific model outputs and specific vendors. A distillation with a swapped vocoder inherits the capability and none of the provenance machinery, and luxtts ships no consent language at all. The cheapest, freest cloning in this corpus is also the least traceable, and that correlation is structural rather than incidental.

Evidence quality drops as the artifact count rises. The distilled model claims parity with “models 10x larger” and publishes nothing. The corpus’s benchmark pages (tts-benchmarks, tts-arena-leaderboard, mean-opinion-score) exist for exactly this and cover the labs, not the derivatives. The measurement layer is organised around the release model at the moment the interesting artifacts have stopped being releases.

The common finding is the one worth carrying: consent language, watermarking and provenance attach to a release — a specific lab’s specific model — and every mechanism above weakens that attachment. The census shows capability distributed without the sentence; distillation shows it distributed without the releaser. The safety axis this spoke tracks was built for a world where openness is granted, and it is now mostly not.

The corpus’s real weakness is what nobody measured, at both ends of the market

Tier discipline weighs claims. Two 2026 findings, arriving from opposite ends of the field, show the same blind spot: neither was a false claim that a tier check would have caught. One was an omission, the other a vacuum.

At the bottom — a marketing survey supplied the set of things that exist

A staleness pass refreshed the music corner and found the branch narrative built on one source: ai-music-generators-2026, a T4 marketing survey. Adding a single neutral board (music-arena-leaderboard, T2) changed three things at once.

udio is not suno‘s close rival. This wiki said it traded “a slight quality gap” for clean rights. It ranks 16th of 18 on instrumental, below stable-audio 2.0, and ~235 Elo behind Suno V5.5. The rights case for Udio is untouched — but “slight gap for lower legal risk” and “large gap for lower legal risk” are different decisions, and the wiki was recommending the first.

A top-two provider was missing. mureka holds ranks 2 and 3 on both boards, within 2 Elo of the leader on instrumental, and appeared nowhere in this hub. The survey never mentioned it. Its own docs, read the same day, make the omission worse rather than better: Mureka has the widest API surface in the branch — generation, vocal cloning, transcription, recognition, stem separation, and text-to-speech and podcasts — undercuts suno and udio on price, and asserts full commercial rights while disclosing no training-data provenance and not naming its operating company. The branch’s quality runner-up is also its least legible participant, which is exactly what a marketing survey has no incentive to surface.

The open wedge does not exist in music. musicgen is last of eighteen.

The generalizable failure is not that a T4 source was wrong about a number — the wiki already discounts T4 claims. It is that a marketing survey silently supplied the set of things that exist. Tier discipline weighs assertions; it does nothing about a source’s omissions, because an omission makes no claim to weigh. Two months of careful reasoning ran on a cast list nobody had checked.

The cheap defence, and the reason this pass was worth running: for any branch the wiki ranks, hold at least one neutral board — not to source the individual claims, but to enumerate the field. TTS had one from the start (tts-arena-leaderboard) and its corresponding error was smaller and numeric (reporting Elo without rank). Music had none, and the error was structural.

At the top — a frontier lab shipped with nothing attached

Music generation had been the branch the frontier labs stayed out of. They shipped speech (gemini-live-3-5-translate) and open instrumental research (musicgen, stable-audio); the songs-with-vocals market belonged to specialists (suno, udio, ElevenLabs Music) who were also the ones carrying the copyright exposure. lyria-3-5 ends that division: google now ships vocals, lyrics and prompt-controlled song structure as a consumer product (Flow Music), competing on exactly the capability the open wedge lacks.

What the announcement withholds counts as much as what it claims. Four adjectival bullets, one demo track, and no architecture, no baseline version, no listening test, no rights statement, and — from the lab that made SynthID a headline feature of its speech model six weeks earlier — no watermarking claim. The same evidence vacuum that luxtts showed at the open end (“on par with models 10x larger,” nothing published) now appears at the closed frontier, from a lab that has the evaluation apparatus and chose not to use it here.

So the branch’s two poles are converging on a shared problem for this wiki rather than on quality: the artifacts that matter most are arriving with the least measurement attached. On music specifically the spoke can still rank suno by Elo and place stable-audio by licence, but it cannot place the newest entrant at all until someone independent measures it.

So both ends fail the same way, and the failure is not dishonesty. A T4 survey and a frontier announcement are equally unfalsifiable when one omits the competitors and the other omits the numbers. The defence in both cases is the same and is cheap: hold one neutral board per branch, and treat an artifact with no independent measurement as unplaced rather than as unranked.

Contradictions / tensions

  • Does Kokoro clone voices? RESOLVED same day (2026-08-03). free-voice-clone-list‘s table marked kokoro as having zero-shot voice-cloning, against tts-models-2026-benchmark and open-source-tts-models. Checked against the Hugging Face model card, which neither side had consulted: Kokoro-82M does not support voice cloning, and its voices/ directory holds exactly 54 .pt voicepack tensors — a voice is loaded from a file, with no reference-audio path. The roundups were right; the catalogue’s checkbox is wrong. Two lessons the spoke should keep, both larger than the fact. First: the disputed cell was settleable in about a minute, and nobody had looked — not the catalogue, not the roundups, not this wiki, which had been citing the claim as a founding example since June. Second: the same check found an error herekokoro read “~15 languages” against the card’s 8, from no source cited anywhere. The catalogue’s failure was a missing method; ours was an uncited number that survived because it was never contradicted. When a primary source is one fetch away, “two secondary sources agree” is not the end of the inquiry.
  • Leaderboard disagreement (minor). tts-models-2026-benchmark and tts-arena-leaderboard report slightly different Elo values and orderings (e.g. whether Fun-Realtime-TTS or Gemini leads). Not a fact conflict — different snapshot dates and vote pools. Recorded, not resolved; both cited with dates.
  • The STT leaderboard lost its authority (2026-07-24). This spoke recorded canary-qwen in June as topping the Open ASR Leaderboard at 5.63% WER. Six weeks later four open models sit below it (open-asr-models-2026-comparison), the top ten are inside one WER point, the reported figures aren’t measured over the same benchmark sets (8, or 7 with TED-LIUM dropped, or LibriSpeech alone), and the board reorders when Appen’s private-track data is toggled on. Both the June and July positions are kept per the record-don’t-overwrite rule; what’s superseded is the framing, not the number. The practical replacement the source argues for: licence, language coverage, streaming-vs-batch, and cost per audio-hour, with WER measured on your own audio.
  • Averaged WER hides the audio type it was measured on. vibevoice-asr-bitnet‘s 15-benchmark table has it beating whisper by 5–11 points on meeting and far-field speech while losing to it on clean read speech, with Parakeet taking every clean set and losing every meeting set. No single ranking survives that. Two sources arriving the same day reach the same conclusion by different routes — one by disaggregating a benchmark, one by distrusting the aggregate.

Cross-spoke adjacency

  • llm-providers-wiki — the text/multimodal LLM market. This spoke is its speech sibling: shared architecture (Llama backbones, open-weight-models dynamics, quantization/codec ideas) and the same closed-vs-open structure. gemini straddles both — Gemini 3.1 Flash TTS is a TTS leader and a Gemini model; linked cross-wiki, not duplicated.
  • llm-inference-wiki — TTS latency/streaming and codec decoding are an inference-mechanics story; adjacent but not yet bridged.
  • Now covered (after the 2026-06-05 broaden + ingest): STT/ASR (speech-to-text: whisper, canary-qwen) and audio/music generation (audio-music-generation: suno, udio, stable-audio) are sourced and integrated under the speech-audio-ai umbrella. The cross-wiki bridge to llm-providers-wiki widened: SALM STT decoders are literally Qwen3 LLMs.

Index — Speech & Audio AI Wiki

Catalog of every page, grouped by schema.org @type. Spine: synthesis (thesis), log.md (history), this file (catalog). Domain = speech & audio AI models across three branches — TTS (synthesis · deepest), STT/ASR (recognition), audio/music generation. Some wiki-links resolve cross-wiki to llm-providers-wiki (gemini, open-weight-models, quantization) — intentional bridge links. Elo / WER / latency facts are dated snapshots.

DefinedTerm (concepts)

SoftwareApplication (models)

  • voicecraft — Peng et al. (UT Austin/Meta, arXiv:2403.16973): token-infilling neural codec LM, 330M/830M on GigaSpeech XL over a 4-codebook EnCodec. The wiki’s first speech-editing model — alter words inside a real recording, speaker and room intact — with zero-shot TTS as the degenerate case. CC BY-NC-SA + Coqui CPML (non-commercial); an explicit consent prohibition, the exception audio-deepfake expected not to find; quiet since March 2025 · source · T1 · github.com
  • TTS: kittentts (KittenML; 15M params under 25 MB, no GPU, Apache-2.0 — claims cloning at a fifth of Kokoro’s size; single T4 source, unmeasured) · supertonic-2 (Supertone; 66M ONNX on-device, OpenRAIL-M, no cloning; vendor table claims 12,164 chars/sec / RTF 0.001 and 42× ElevenLabs Flash v2.5) · luxtts (Apache-2.0 ZipVoice distillation, 4.9k★: flow matching in 4 steps, 48 kHz vocoder, 150× realtime in 1 GB VRAM, 3-second cloning; “SOTA… on par with models 10x larger” with no MOS/WER/Elo published, and no consent language · source · T1 · github.com) · misotts (Miso 8B emotive; RVQ/Mimi · source) · kokoro (82M efficiency leader) · fish-audio-s2-pro (5B; highest-Elo open) · orpheus (Llama family; cloning+streaming) · sesame-csm (1B conversational) · wavenet (DeepMind 2016; the foundational neural-audio ancestor · source)
  • STT: whisper (OpenAI; MIT; ecosystem default) · canary-qwen (NVIDIA; 5.63% SALM; displaced 2026-07 — four Apache-2.0 models now below it) · whisper-cpp (ggml C/C++ Whisper inference; quantized, on-device) · vosk (Alpha Cephei; Kaldi; ~50 MB offline; Apache-2.0; 20+ langs; the small/embedded pole · source · T1 · github.com) · moonshine (Moonshine AI; edge STT from 27M params; micro runs VAD+STT+neural-TTS on a Pico 2 W MCU — the sub-laptop pole) · vibevoice-asr-bitnet (Microsoft Research; MIT; 1.58 GB from 4.62 via I2_S+Q6_K decoder / I8_S tokenizer, llama.cpp-served; RTF 0.77 at 3 CPU threads, claimed 1.86× whisper.cpp; wins far-field/meeting audio, loses clean read speech · source · T1 · huggingface.co)
  • Music: mureka (ranks 2 and 3 on both boards, V9 within 2 Elo of Suno; docs still describe V7.5/O1. Widest API surface here — generation + cloning + transcription + stem separation + TTS/podcast — at $8/$24 a month. Commercial rights asserted, training data undisclosed, operator unnamed: a third position on the rights axis) · source · T3 · lyria-3-5 (Google, 2026-07-29; vocals+lyrics in Flow Music — the first frontier-lab song generator here, announced with no benchmark, no rights statement, no SynthID claim · source · T3) · suno (leader, Elo ~1293; Sony litigation) · udio (licensing-clean; UMG deal) · stable-audio (Stability AI; open-weight, instrumental) · musicgen (Meta AudioCraft; open research reference; MIT code / CC-BY-NC weights · source)
  • S2ST: gemini-live-3-5-translate (Google; end-to-end speech-to-speech translation, 70+ langs, voice-preserving, SynthID · source)
  • Conversational S2S: gpt-realtime-2 (OpenAI; Realtime-API speech-to-speech, GPT-5-class reasoning, WebRTC, document-context grounding · source · T4)
  • Speech understanding (multi-task): sensevoice (Alibaba/FunAudioLLM, repo under a QwenAudio org; MIT code but FunASR licence on the weights, ~9k★): ASR + language ID + 7-category emotion + audio-event detection in one non-autoregressive pass — claimed >5× Whisper-Small / 15× Whisper-Large; only Small released, the benchmarked Large is not downloadable; the cheap non-generative form of the LALM branch · source · T3 · github.com
  • Audio understanding (LALM): audio-flamingo-3 (NVIDIA; reasons over speech+sound+music; AF-Whisper encoder, CoT, 10-min audio; open weights/data, non-commercial; the 4th branch · source)

SoftwareApplication (apps over the engines)

  • abogen — denizsafak, MIT, 5.5k★: ebook/PDF/subtitle → audiobook on kokoro, with synced captions, M4B chaptering, a voice mixer answering Kokoro’s no-cloning tradeoff, and LLM-assisted text normalization; PyQt6 desktop + Flask web UI — the TTS consumption layer · source · T1 · github.com
  • vocalinux — open-source (GPL-3.0) system-wide voice dictation for Linux; wraps whisper.cpp/Whisper/VOSK locally + Silero VAD; X11+Wayland text injection, Vulkan accel without CUDA; offline, no cloud — the STT consumption layer · source · T1 · github.com

Organization (providers)

  • elevenlabs — leading commercial voice-AI company; TTS + cloning + dubbing + Scribe STT + Eleven Music; closed; $11B (2026) · source · wikipedia

Person (curators)

  • 0xsojalsec — pseudonymous GitHub curator of local/open-weight model censuses, one per modality: free-voice-clone-list here and llms-local-list in llm-inference-wiki. Paged on recurrence across two spokes; evidence is the two repos and nothing else · entity

Collection (sources)

  • free-voice-clone-list — 0xSojalSec’s free-voice-clone: 53 open audio models across TTS/music/any-to-audio/restoration/ASR. 27 of 36 TTS entries Apache-2.0; 35 of 36 claim zero-shot cloning; and zero occurrences of consent/watermark/misuse/deepfake in 59 KB. Census, not benchmark — no MOS/WER/Elo, last pushed 2026-04-09. The one row anyone verified (kokoro cloning) was wrong · source · T4 · github.com

ScholarlyArticle (sources)

  • whisper-paper — Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (OpenAI, 2022); the primary behind whisper — 680k hours, zero-shot, no fine-tuning · source · T1 · arxiv.org
  • soundstream-paper — Zeghidour et al., SoundStream (Google, 2021); the origin of the RVQ neural-audio-codec — one model for speech+music, 3–18 kbps, real-time on a phone · source · T1 · arxiv.org
  • encodec-paper — Défossez et al., High Fidelity Neural Audio Compression / EnCodec (Meta, 2022); SoundStream’s successor and the codec under musicgen — multiscale spectrogram adversary, ~40% transformer compression · source · T1 · arxiv.org
  • xtts-paper — Casanova et al., XTTS (Coqui, INTERSPEECH 2024); open zero-shot voice-cloning across 16 languages, publicly released · source · T1 · arxiv.org

WebPage (sources)

  • elevenlabs-expressive-mode — ElevenLabs Conversational AI / Expressive Mode: V3 Conversational TTS + Scribe v2 Realtime STT; emotion-from-prosody; the real-time conversational composite · source · T3 · join.elevenlabs.io

Dataset (benchmark corpora)

  • librispeech — Panayotov et al., ICASSP 2015 (T1): 1000 h of read English from LibriVox audiobooks, CC BY 4.0, aligned against Gutenberg texts. Subset table + the definition that matters — “clean” means speakers a 2015 WSJ-trained model already transcribed well, split at the median WER, not an acoustic measurement · source · T1 · danielpovey.com
  • common-voice — Ardila et al., 2019 (T1): Mozilla’s crowd-recorded, crowd-validated corpus — 2,500 h, 50,000+ contributors, 29 languages released / 38 collecting, CC0. Breadth and public-domain licensing at the cost of controlled recording; the licence is also why it may be in the training data of the models being tested on it · source · T1 · arxiv.org
  • fleurs — Conneau et al., 2022 (T1): 102 languages, n-way parallel, ~12 h each, recorded from FLoRes-101 sentences; an evaluation budget, not a training one. Same content in every language, so cross-language scores are comparable · source · T1 · arxiv.org

Article / BlogPosting / Dataset (sources)

  • tts-models-2026-benchmark — MarkTechPost: TTS benchmark comparison (proprietary + open) · source · marktechpost.com
  • tts-arena-leaderboard — Artificial Analysis: neutral TTS Elo Speech Arena board. Refreshed 2026-08-03: new #1 (Speechify Simba 3.2, 1229 — and the cheapest of the top ten at $10/1M); open leader fish-audio-s2-pro is only 20th, kokoro 48th, and the open→closed gap widened 99 → 106 Elo · source · T2 · artificialanalysis.ai
  • open-source-tts-models — Modal: open-weight TTS deep-dive (Higgs, Kokoro, Dia, Chatterbox, Orpheus, CSM) · source · modal.com
  • open-source-stt-models — Northflank: open-source STT/ASR + Open ASR Leaderboard WER · source · northflank.com
  • moonshine-pico-voice-toolkit — Hackaday: Moonshine AI’s micro toolkit — VAD + SpellingCNN STT + neural TTS on a Raspberry Pi Pico 2 W (~3.6 MiB flash/468 KiB SRAM), 50-token retrainable vocab; the microcontroller extreme of the on-device pole · source · T3 · hackaday.com
  • open-asr-models-2026-comparison — MarkTechPost: 15 open ASR models compared on WER/languages/latency/licence; argues rank is no longer the deciding variable (top ten inside one WER point, board reorders on private data) — licence, coverage, streaming-vs-batch and cost/audio-hour instead; surfaces Meta Omnilingual ASR at 1600+ languages, diarization and NAR variants · source · T3 · marktechpost.com
  • stt-apis-comparison — Future AGI: commercial STT APIs (Deepgram, AssemblyAI, Chirp, ElevenLabs) · source · futureagi.com
  • ai-music-generators-2026 — Chartlex: music generators + the copyright fault line. Superseded on rankings 2026-08-03 — and it had set the branch’s whole cast, omitting mureka · source · T4 · chartlex.com
  • music-arena-leaderboard — Artificial Analysis: neutral music Elo, split instrumental / vocals (2026-08-03). suno V5.5 leads both (1189/1171); mureka V9 is 2 Elo behind; udio is 16th of 18 and musicgen last — the open wedge that exists in TTS/STT does not exist here · source · T2 · artificialanalysis.ai
  • stable-audio-3 — MindStudio: Stable Audio 3.0 open-weight music generation · source · mindstudio.ai

Synthesis

  • synthesis — the thesis: across all three branches, a closed frontier vs. an open-weight wedge; rights/latency as audio-specific axes; convergence on the LLM stack

Bridge nodes (live in sibling wikis, linked cross-wiki)

gemini · open-weight-models · quantization (llm-providers-wiki) · llms-local-list (llm-inference-wiki — the same curator’s text-LLM census, see 0xsojalsec)