Log — Speech & Audio AI Wiki
Append-only history. Each entry starts with ## [YYYY-MM-DD] <op> | <title> where
<op> is ingest, query, lint, split, or broaden, so grep "^## \[" log.md | tail -5 works.
[2026-06-05] ingest | STT/ASR + audio/music generation sources (router-curated, user request)
User: “seek STT and audio generation sources.” Following the curate-at-request pattern, web-searched
both landscapes and ingested 4 sources + 13 new pages, generalizing the wiki from TTS-only to
all three branches of speech-audio-ai.
Sources (URL-only, source: true):
- open-source-stt-models (Northflank) — open STT/ASR + Open ASR Leaderboard WER.
- stt-apis-comparison (Future AGI) — commercial STT APIs.
- ai-music-generators-2026 (Chartlex) — music generators + copyright.
- stable-audio-3 (MindStudio) — open-weight music generation. New pages (13): concepts speech-audio-ai (new top umbrella), speech-to-text, audio-music-generation, ai-music-copyright; models whisper, canary-qwen (STT), suno, udio, stable-audio (music). Updated text-to-speech (now one branch of the umbrella) + index + synthesis (thesis generalized: the closed-vs-open / per-axis / LLM-convergence / rights dynamics now shown across all three branches; key new insight — SALM STT decoders are literally Qwen3 LLMs, widening the llm-providers-wiki bridge; and music’s binding axis is copyright, the sharpest form of the rights theme). Wiki 13 → 26 pages. Site rebuilt + verified. Remaining gaps noted: voice-cloning safety, non-English depth, audio-understanding, neutral latency.
[2026-06-05] broaden | tts-wiki → speech-audio-wiki (user directive “broaden the domain”)
Hours after spin-out, the user directed broadening the domain. Widened scope from TTS only to
all speech & audio AI models — TTS / speech synthesis (founding), STT / ASR (speech
recognition), and audio / music generation — and renamed the spoke tts-wiki →
speech-audio-wiki to match (the registry-vs-name drift rule; same play as the
ai-search-wiki → search-marketing-wiki broaden+rename). Updated: directory mv; this spoke’s
CLAUDE.md (title + domain header + boundary — STT/ASR & audio-gen now in scope, not parked),
index.md (title + scope note), synthesis.md (title + thesis caveat + the former “out of scope” lines
flipped to “in scope, not yet sourced”). Existing 13 TTS pages unchanged and still valid — TTS is now
the deepest corner of a broader map. Slug-based wikilinks unaffected by the dir rename; no other
wiki links into this one, so no external references to fix. Hub: wikis.md block renamed + broadened,
hub log rename/broaden entry; spoke count unchanged (11). Clean rebuild + verify. Priority next
sources: STT/ASR (Whisper/Deepgram) and audio/music gen (Suno/Udio/MusicGen) to generalize the
thesis beyond TTS.
[2026-06-05] split | tts-wiki created from _inbox speech-audio-models (4 sources, router-curated)
User asked the router to “seek for more tts models knowledge.” With the parked misotts as seed
(1), the router web-searched the TTS landscape and curated 3 quality sources → 4 total → spun out
tts-wiki (the speech-synthesis model market — the speech sibling of llm-providers-wiki). Scaffolded
from CLAUDE.template.md; registered in wikis.md; hub spoke count 10 → 11.
Sources ingested (all URL-only, source: true):
- misotts — Miso Labs 8B emotive (RVQ/Mimi; 110ms; modified-MIT) — migrated from _inbox; model+source.
- tts-models-2026-benchmark — MarkTechPost benchmark comparison (proprietary + open-weight).
- tts-arena-leaderboard — Artificial Analysis neutral Elo Speech Arena board.
- open-source-tts-models — Modal open-weight deep-dive (Higgs, Kokoro, Dia, Chatterbox, Orpheus, CSM).
Pages created (13 total): 5 concepts (text-to-speech, tts-benchmarks, open-weight-tts, voice-cloning, neural-audio-codec) + 5 models (misotts, kokoro, fish-audio-s2-pro, orpheus, sesame-csm) + 3 source summaries (misotts doubles as model+source).
Synthesis thesis: the TTS market rhymes with the text-LLM market (closed Elo frontier vs.
open-weight field) with three twists — license decides at the open ceiling (fish-audio-s2-pro
is research-only), latency (TTFA) is a make-or-break axis, and TTS is converging on the LLM
stack (Llama backbones + RVQ neural-audio-codec). Cross-wiki bridges to llm-providers-wiki
(gemini straddles both as Gemini 3.1 Flash TTS; open-weight-models/quantization).
Deleted the parked _inbox/misotts.md. STT/ASR + music-gen explicitly out of scope. Site rebuilt + verified.
Flagged tension: the two leaderboard sources disagree slightly on Elo (snapshot drift), recorded in synthesis.
[2026-06-09] ingest | Gemini 3.5 Live Translate (end-to-end speech-to-speech translation)
Hub-routed (Telegram, blog.google). New model/source page gemini-live-3-5-translate (SoftwareApplication, url) — Google’s end-to-end speech-to-speech translation audio model (70+ languages, 2000+ combinations, near-real-time/simultaneous, voice-preserving intonation/pacing/pitch, SynthID-watermarked; Gemini Live API + Google Translate + Meet, 2026-06). New concept speech-to-speech-translation (DefinedTerm) — the first composite/applied task in the wiki (STT+MT+TTS, streaming), added to the speech-audio-ai umbrella as a cross-branch task. Folded into synthesis (“composite tasks fuse the branches”; purest case of “LLM eats audio from both ends” — both ends in one model; latency-as-quality-tradeoff; SynthID introduces a provenance/watermark axis; gemini now spans TTS+STT+S2ST at the closed frontier). Cross-linked gemini (llm-providers, cross-wiki). Index updated. Caveat: vendor announcement; language/latency/voice claims not independently benchmarked. speech-audio-wiki 26 → 28 pages.
[2026-06-09] ingest | +2 safety axis + origin (audio deepfake, WaveNet) — all-spokes cron test
Filled the flagged “safety/consent” coverage gap and added the field’s origin point: audio-deepfake (DefinedTerm, src — voice-cloning misuse: CEO-voice fraud, fake-Biden robocalls, non-consensual cloning; ASVspoof detection + SynthID watermarking + FCC robocall ban — the lived form of “rights outweigh quality”) and wavenet (SoftwareApplication, src — DeepMind 2016, dilated causal convolutions, the neural-audio ancestor; its sample-by-sample latency → Parallel WaveNet is the first instance of the spoke’s TTFA tension). Synthesis coverage-gap question updated. Both Wikipedia url-only. 28 → 30 pages.
[2026-06-10] ingest | ElevenLabs + MusicGen — all-spokes pass (commercial archetype + open music reference)
Two pages at opposite ends of the closed-vs-open structure (spoke recently grown, kept to 2). elevenlabs (Organization, url, Wikipedia) — the leading commercial closed voice-AI company, referenced throughout but unpaged; spans all three branches (TTS/cloning/dubbing/Scribe STT/Eleven Music), $11B valuation Feb 2026; sits both sides of the rights axis (cloning → audio-deepfake + ships an AI Speech Classifier detector). musicgen (SoftwareApplication, url, Meta announcement — MusicGen/AudioCraft Wikipedia pages 404’d, substituted the official Meta source) — the open research reference for the music branch: single-stage transformer over EnCodec RVQ tokens, 300M/1.5B/3.3B, melody-conditioned; MIT code but CC-BY-NC (non-commercial) weights = open-but-not-shippable, the music echo of TTS’s research-license ceiling; clean-rights training data (clear of ai-music-copyright suits hitting suno). Folded into synthesis (new 2026-06-10 section) + index (new Organization group + music row). No contradictions. 30 → 32 pages.
[2026-06-12] ingest | Audio Flamingo 3 (NVIDIA) — audio-understanding LALM
All-spokes daily expansion. Added audio-flamingo-3 (@type SoftwareApplication) — the wiki’s first audio-understanding model, opening a fourth branch (comprehension) beside TTS/STT/music. A fully-open Large Audio-Language Model that reasons over speech+sound+music (AF-Whisper unified encoder, on-demand CoT, ~10-min audio, voice-to-voice; SOTA on 20+ benchmarks; trained on open data). Completes “the LLM eats audio from both ends” (the comprehension vertex) and re-confirms “license, not score, decides” (open weights+data but non-commercial — same ceiling as fish-audio-s2-pro/musicgen). Seeds the flagged “audio understanding beyond transcription” gap. synthesis “fourth branch” note; index gains an Audio- understanding row. 1 new page. Caveat: vendor-reported SOTA, young benchmarks, non-commercial license.
[2026-06-13] ingest | ElevenLabs Expressive Mode / Conversational AI (join.elevenlabs.io)
Routed from hub (Telegram drop). ElevenLabs’ real-time conversational-voice-agent product: Eleven V3
Conversational TTS + Scribe v2 Realtime STT, emotion inferred from prosody (“tone cue cards”), 70+
languages, “ultra-low latency.” Quality gate: tier T3 (vendor landing/signup ad — CSAT/NPS/latency claims
unverified; model names are first-party specs); elevenlabs already paged → new source summary
elevenlabs-expressive-mode (not a dedup of the Org page); gap-relevance — advances the open “how real
are latency claims?” question and adds an emotion/affect quality axis; integrated as the second composite
task beside S2ST in synthesis + a new product bullet on elevenlabs.
freshness: volatile. url provenance. Site rebuild + commit follow.
[2026-06-14] ingest | GPT-Realtime-2 / OpenAI WebRTC Audio Session (simonwillison.net)
Routed from hub (Telegram drop). OpenAI’s GPT-Realtime-2 voice model for the Realtime API — speech-to-speech
over WebRTC, “first voice model with GPT-5-class reasoning,” plus a document-context feature (paste text,
explore by voice). Quality gate: tier T4 (Simon Willison personal-blog demo write-up; primary is OpenAI’s own
release); not previously paged → new model page gpt-realtime-2. Gap-relevance: third entrant in the
real-time conversational-voice thread (after elevenlabs-expressive-mode + gemini-live-3-5-translate) —
the loop collapsed into one end-to-end model carrying frontier reasoning; adds a grounding axis (document
context). Integrated into synthesis (conversational composite thread). Cross-spoke: OpenAI text frontier →
llm-providers; RAG → research-wiki (noted, not paged). freshness: volatile. Site rebuild + commit follow.
[2026-06-15] ingest | Whisper primary paper (Radford et al. 2022) — T1 anchor for the STT baseline
Quality cycle, T1-floor raise. whisper was a concept page sourced only from comparison roundups
(open-source-stt-models, stt-apis-comparison, T3/T4); added the primary paper as a source:
whisper-paper (ScholarlyArticle, T1) — Radford, Kim, Xu, Brockman, McLeavey, Sutskever, Robust
Speech Recognition via Large-Scale Weak Supervision, arXiv:2212.04356 (ICML 2023). Grounds why
Whisper is the multilingual default: weak-supervision-at-scale (680k hours), zero-shot transfer with no
fine-tuning, human-approaching robustness, MIT release — not WER leadership. Threaded into synthesis
(“LLM eats audio from both ends” recurring read: SALM successors beat it on accuracy yet can’t displace
the weak-supervision baseline). Concept/paper split mirrors optimization (NFL) + llm-inference
(FlashAttention). Found via WebSearch; figures from the abstract, architecture from established
knowledge (noted on page). Linked from concept page, synthesis, index (new ScholarlyArticle section).
1 new page.
[2026-06-18] ingest | DefinedTerm enrichment pass (subagent)
Thinnest-first deepening of the spoke’s [DefinedTerm] pages with fetched primary/standard sources.
Enriched four 33-line concept pages and added four source pages:
- neural-audio-codec — grounded the RVQ line in its primaries: soundstream-paper (Zeghidour et al., Google 2021 — one model for speech+music, 3–18 kbps via quantizer dropout, real-time on a phone, 3 kbps beats Opus 12 kbps) and encodec-paper (Défossez et al., Meta 2022 — multiscale spectrogram adversary, loss balancer, ~40% transformer token compression; the codec under musicgen).
- tts-benchmarks — anchored MOS in mean-opinion-score (ITU-T P.800 5-point Excellent→Bad scale; the “don’t compare across experiments” rule that is the standards basis for snapshot discipline).
- voice-cloning + open-weight-tts — added xtts-paper (Casanova et al., Coqui, INTERSPEECH 2024) as the peer-reviewed primary for open zero-shot cloning across 16 languages, publicly released — the autoregressive (Tortoise-lineage) route to cloning, permissive side of the open license split. New pages: soundstream-paper (T1), encodec-paper (T1), xtts-paper (T1) — ScholarlyArticle; mean-opinion-score (T2, DefinedTerm+WebPage). All figures from fetched abstracts/reference; volatile claims dated. index.md updated (new ScholarlyArticle entries + MOS under DefinedTerm). Cross-wiki bridge reused: open-weight-models (llm-providers-wiki). No build, no commit.
[2026-07-19] ingest | Vocalinux — open-source Linux voice dictation (Linuxiac)
Routed here by the hub (dominant substance = STT/ASR; the Linux-desktop/FOSS angle is a facet, not a spoke). Opens a new layer in the spoke: the consumption/application layer over STT engines, distinct from the model+WER pages that preceded it. vocalinux is a system-wide dictation app (hotkey toggle/push-to-talk, injects text into any app) that wraps a local backend — whisper.cpp (default), OpenAI Whisper, or VOSK — runs offline, no cloud, handles both X11 and Wayland text injection, Vulkan GPU accel. New nodes: vocalinux (SoftwareApplication, source, T3 — single Linuxiac product write-up), vosk (Alpha Cephei / Kaldi; ~50 MB offline; the small/embedded pole the WER pages skip), whisper-cpp (ggml quantized Whisper inference — the “how it runs on-device” counterpart to whisper‘s model story). Reused cross-wiki openai (llm-providers-wiki) and quantization. Updated whisper (+whisper.cpp/Vocalinux) and speech-to-text (new “application layer — on-device dictation” section). Synthesis: new thread “The consumption layer — packaged apps over the engines” (open-wedge consumer end-point; the small/local sorting axis; the inverse of the cloud voice agents gpt-realtime-2/gemini-live-3-5-translate). avoid-ai-writing self-pass. +3 pages, 3 updated. Per-route verify deferred per hub policy.
[2026-07-19] ingest | VOSK API repo (github.com/alphacep/vosk-api) — primary upgrade of vosk
Re-seen subject: vosk was created hours earlier from a secondary mention in vocalinux; the primary
repo now upgrades it in place (mention-grounded → source: true, url, T1). New primary facts: Apache-2.0
(fully permissive — not a research/NC license), 20+ languages, ~50 MB models, streaming zero-latency +
speaker identification (a capability beyond transcription), runs “Pi/Android → big clusters,” bindings for
Python/Java/Node/C#/C++/Rust/Go; ~15k stars, latest v0.3.50 (Apr 2024) = mature/slow-moving. Synthesis: added
the license-axis point to the consumption-layer thread — VOSK sits at the clean end (Apache-2.0), so in the
small-STT corner the constrained option is also the freely-shippable one, and the spoke’s recurring
“best-sounding = least legally safe” tension (fish-audio-s2-pro ceiling) doesn’t bite. Idempotent refresh
(no new page). 1 page upgraded, index + synthesis updated. Per-route verify deferred per hub policy.
[2026-07-23] ingest | Vocalinux — primary repo (refresh of the 2026-07-19 page)
Re-seen subject, better source. The page was created 2026-07-19 from a Linuxiac writeup (T3); the user sent the primary GitHub repo (jatinkrmalik/vocalinux). Per HUB re-seen rule, refreshed the existing vocalinux page in place — no new page, no duplicate. Upgrades from the primary source: tier T3 → T1 (repo facts); resolves the earlier open license caveat → GPL-3.0; adds Silero VAD neural voice-activity detection (ONNX Runtime), editing voice commands (“new line”/“delete that”/“capitalize”), multi-language (FR/DE/RU), the Vulkan-without-CUDA detail, the “Voca” ecosystem (VocaMac beta / VocaWin planned) + maintainer @jatinkrmalik, and stack/DE coverage. Kept the Linuxiac URL as a secondary source in frontmatter. Unchanged verdict: feature claims are still vendor self-description (no independent benchmark), and Vocalinux stays untestable as a recognizer by design — accuracy is inherited from the chosen backend/model. Dedup: whisper-cpp, whisper, vosk, speech-to-text already exist and are linked; Silero VAD left as an inline mention, not a new page (single mention — thin-node caution; page it if VAD recurs). Entity: maintainer @jatinkrmalik not paged (solo-dev name only, evidence-only rule). Synthesis untouched — no thesis change, just a provenance/detail upgrade on an existing consumption-layer node. Index line updated (T1, GPL-3.0, github). Page count unchanged (refresh). Verify deferred per hub policy (content-only, no page moves). avoid-ai-writing run.
[2026-07-23] ingest | Voice control toolkit comes to a Pico near you (Hackaday)
Routed from the hub (route in ../log.md). URL-only, T3 — Hackaday secondary coverage of Moonshine
AI’s micro toolkit; no WER/latency/accuracy, license not stated (repo would resolve).
New pages: moonshine-pico-voice-toolkit (TechArticle source) and moonshine (SoftwareApplication —
edge STT family / Moonshine AI). The toolkit chains VAD → SpellingCNN STT → neural TTS on a Raspberry
Pi Pico 2 W in ~3.6 MiB flash / 468 KiB SRAM, 50-token retrainable vocab, C++ loop.
Dedup: Moonshine was a dangling inline mention in open-source-stt-models (“edge, from 27M params”)
with no page — now paged and both mentions there linked to moonshine. vocalinux (yesterday’s
consumption-layer page) is the nearest neighbour; linked, not duplicated.
Gap-relevance: extends the small/local pole vocalinux opened by an order of magnitude — desktop →
microcontroller. Two firsts for the spoke: TTS on an MCU and on-device VAD. Synthesis: added a
paragraph under the consumption-layer section — the “best = fits my hardware” axis hardened into a design
principle (retrainable 50-token recognizer instead of open-vocabulary ASR), and the flip vs Vocalinux
(purpose-built co-designed tiny stack vs wrapper over general engines).
Caveat recorded: the impressive claim is the fit (concrete); recognition quality at this size is
unmeasured. T3, secondary source.
Entities: Moonshine AI noted inline on moonshine but not separately paged (no distinct evidence
beyond the model line + toolkit); author Julian Scheffers (Hackaday) not paged (byline, evidence-only).
Index: added moonshine to the STT models row and moonshine-pico-voice-toolkit to sources.
Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-24] ingest | microsoft/VibeVoice-ASR-BitNet (Hugging Face model card)
Routed from the hub (route in ../log.md). URL-only, T1 (first-party model card); all speed and
WER figures are self-reported on one machine.
New page: vibevoice-asr-bitnet (SoftwareApplication + source: true, the vosk pattern).
Facts: MIT; 1.58 GB total from 4.62 GB FP16 (“2.9× compression”), VAE tokenizer I8_S 1.31→0.65 GB,
LM decoder I2_S+Q6_K 3.32→0.92 GB; served by llama.cpp; RTF 1.98/1.08/0.77/0.63/0.49/0.42 at
1–8 threads on an AMD EPYC 7V13, crossing real-time at 3 threads with a claimed 1.86× whisper.cpp;
AVX2/NEON, no GPU; 7 languages; 15-benchmark WER table against the uncompressed VibeVoice-ASR-7B,
Parakeet, Whisper, SenseVoice and FunASR.
Two readings recorded, both new here: (1) quantization damage is uneven — 0.2–0.6 WER points on most
sets but +7.6 on AISHELL4, +4.4 AliMeeting, +1.3 MLC-KO, +2.2 MLC-VI, so the loss lands on far-field
Chinese and lower-resource languages; (2) it beats Whisper on meeting/far-field and loses on clean read
speech, with Parakeet inverted, so the single-number ranking on speech-to-text hides the split.
Naming caveat recorded: the card body never says “BitNet”, “1-bit” or “ternary” — I2_S is the ternary
format but only on the decoder, alongside Q6_K and an 8-bit tokenizer, so it’s mixed precision; the
“1.58 GB” total is not the 1.58-bits-per-weight figure.
Contradiction recorded, unresolved: the card’s “0.3B params” field can’t be reconciled with its own
4.62 GB FP16 figure or with the VibeVoice-ASR-7B parent. Also unstated: training data, limitations,
streaming, diarization, and any link to the VibeVoice TTS line.
Dedup: no VibeVoice page existed. Updated whisper-cpp (it is now the baseline a vendor quotes) and
speech-to-text. Cross-wiki quantization linked, not duplicated.
Entities: none paged — Microsoft Research is the publisher, no byline, no distinct evidence beyond the card.
Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-24] ingest | Best open ASR models in 2026 — WER, languages, latency, licence (MarkTechPost)
Routed from the hub (route in ../log.md). URL-only, T3 — a secondary editorial roundup
aggregating vendor cards and leaderboard figures, same outlet as tts-models-2026-benchmark. Arrived
minutes after the VibeVoice card and reaches the same conclusion from the other end, so the two were
ingested as a pair.
New page: open-asr-models-2026-comparison (Article source) with the 15-model table.
The argument: “rank is no longer the deciding variable” — the top ten open models sit inside one WER
point, the board reorders when Appen’s private-track data is toggled on, and the quoted WERs aren’t
measured over the same sets (8 · 7 without TED-LIUM · LibriSpeech only). Its five-step replacement:
licence → language coverage → streaming vs batch → WER on your own audio → cost per audio-hour;
“a procurement question rather than a research one.”
Four things it changes here: canary-qwen displaced (ARK-ASR-3B 5.04, Granite Speech 4.1 5.33,
Cohere Transcribe 5.42, MOSS-Transcribe-preview under 5.33 all below its 5.63, and all Apache-2.0 against
its CC-BY-4.0); throughput and streaming latency separated from accuracy as tiers (Parakeet TDT RTFx
3333; Voxtral Mini 80–1200 ms configurable; Kyutai STT 0.5 s) — latency had been a commercial-API
property here, now an open-weight one; Meta Omnilingual ASR at 1600+ languages (5400+ zero-shot,
~10% CER on 78%) against Whisper’s 99; and diarization + non-autoregressive decoding appear as
product axes with no page yet (gap noted).
Dedup: distinct from open-source-stt-models (Northflank, June) — same subject, newer field, different
argument; both kept and cross-linked. Updated canary-qwen (displacement section, both figures kept per
record-don’t-overwrite) and speech-to-text (new “the leaderboard stopped deciding things” section).
Synthesis: three folds — the standing “does the open ceiling go permissive?” question is now
answered on STT (permissive licences hold the accuracy lead there; the research-licence ceiling is a
TTS/music phenomenon), a new contradiction entry for the leaderboard’s lost authority, and a recurring-read
bullet on compression as an STT claim with uneven damage.
Entities: none paged — MarkTechPost already the publisher behind tts-models-2026-benchmark, no byline given.
Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-07-29] ingest | LuxTTS (ysharma3501)
Routed from the hub. T1, repository and model card read directly: Apache-2.0, 4,867★/634 forks, created 23 January 2026, last pushed 5 June. New page luxtts.
What it is, stated by its author: a distillation of k2-fsa’s ZipVoice to 4 sampling steps with an improved sampling technique and a custom 48 kHz Vocos vocoder replacing the default 24 kHz. Not a new architecture. That makes it the spoke’s first flow-matching source — an architecture family the domain names and had no instance of. Claims 150× realtime on one GPU, faster than realtime on CPU, 1 GB VRAM, cloning from a 3-second reference. float32 today; fp16 unshipped.
The claim to record and not repeat: “SOTA voice cloning on par with models 10x larger” with no
MOS, WER, Elo, comparison table or named baseline. Sharpened by the fact that the README uses WER as a
tuning axis (t_shift trades quality against it) — the author works with the metric and publishes no
value. The 48 kHz output is the one differentiator that needs no benchmark, being a vocoder property.
Real evidence that does exist is adoption: five independent community projects (Gradio, ComfyUI, ONNX,
a UI wrapper) plus fal.ai hosting it.
No consent language anywhere — no ethics note, no watermarking, no acceptable-use statement. Filed against audio-deepfake, and it generalised into a new synthesis section: distillation is the corpus’s third openness mechanism, after permissive release and research-licence ceilings. Capability propagates without a release decision (the base licence permitted it), the safety infrastructure does not propagate with it (SynthID and ASVspoof attach to vendors and outputs, not to a fork with a swapped vocoder), and the benchmark layer this spoke maintains is organised around labs’ releases at the point where the interesting artifacts have stopped being releases. Added to open questions as where does distillation put the open TTS ceiling? — a live version of the existing “does the open ceiling go permissive?” question, since a distilled derivative reaching parity would make lab permissiveness optional.
No page for ZipVoice or for flow matching as an architecture: one source each, folded into luxtts per the ENTITIES recursion discipline. Both become warranted the moment a second source touches them — flagged here so the next router sees it. No person node either (a GitHub handle, an email and a HF account, with no role in the corpus beyond this repo). avoid-ai-writing run. Verify deferred per hub policy (content-only, no page moves).
[2026-07-29] ingest | Lyria 3.5 in Google Flow Music (blog.google)
Routed from the hub (Telegram). lyria-3-5 created as a combined model + source page, following the
gemini-live-3-5-translate precedent for a Google announcement of an audio model: one node, T3,
freshness: volatile.
The source is four adjectival bullets and a demo video — thin even by vendor-post standards. Recorded what it claims (musicality, lyrics, vocals, tempo/duration control) and, at more length, what it omits: no architecture, no baseline version named, no benchmark, no pricing, no rights or training-data statement, and no SynthID mention — notable because Google led with SynthID on gemini-live-3-5-translate six weeks earlier. Unstated, not denied; written that way.
Folded into audio-music-generation and a new synthesis section: the frontier labs had left vocal song generation to the specialists carrying the ai-music-copyright exposure, and that division is now over. The evidence gap pairs with the luxtts read from earlier today — open end and closed end both shipping the branch’s most significant artifacts with nothing measurable attached.
No page for Flow Music as a product (one source, folded in) and no Google Labs org node — google already exists in llm-providers-wiki and is linked cross-wiki per the bridge convention. avoid-ai-writing run. Verify deferred per hub policy (content-only, no page moves).
[2026-08-03] ingest | free-voice-clone (0xSojalSec) — 53 open speech/audio models
Routed from the hub (Telegram). T4 — one pseudonymous maintainer, README-only, no inclusion criteria, no MOS/WER/Elo/listening test, no licence on the repo, and last pushed 2026-04-09 (read 2026-08-03) in a field this spoke’s conventions call weekly-volatile. Ingested with the weakness recorded per the quality gate; its value is coverage and a licence census, both checkable against the linked repos. New pages: free-voice-clone-list (Collection, source), kittentts + supertonic-2 (SoftwareApplication), 0xsojalsec (Person). Updated: open-weight-tts (licence census), voice-cloning (the dividing line stopped dividing), audio-deepfake (the supply-side silence), kokoro (date only), synthesis, index. Dedup: kokoro, fish-audio-s2-pro, luxtts, audio-flamingo-3 and VibeVoice-ASR all already paged — linked, not duplicated. Only two of the 53 models were paged; the rest are recorded as a census, since paging 53 unmeasured rows would inflate the spoke with claims nothing verifies. Gap-relevance: (1) sharpens the safety/consent axis from “harms exist” to “the distribution layer says nothing” — the source’s most checkable fact; (2) hardens the licence half of the does-the-open-ceiling-go-permissive question with counts (27/36 Apache-2.0) while explicitly not answering it, since a census ranks nothing; (3) moves the small pole from 82M to 15M. Counted twice. First pass read the TTS table as 38 rows with two non-cloning entries; the table is 36 rows with one (supertonic-2). All four pages were corrected before commit. Contradiction flagged: the list marks kokoro as having zero-shot cloning, against tts-models-2026-benchmark and open-source-tts-models, which both record preset voices and no cloning as the price of its 82M footprint. Both kept; the pages keep the roundups’ reading and treat the conflict as a warning about the catalogue’s other 35 checkboxes. The finding held to its evidence: zero occurrences of consent/ethics/watermark/misuse/deepfake/ responsible/abuse/disclaimer in 58,997 characters (grep, reproducible). Recorded as a gap in what gets said, with no claim that the omission causes harm. Also noted: the one OpenRAIL-M entry (behavioural use restrictions) is the one model that cannot clone a voice — one data point, recorded, not built on. Entity: 0xsojalsec paged on recurrence across two spokes (deferred as a thin creator in llm-inference-wiki on 2026-06-21; the same curator’s audio census triggers the “page when one recurs” rule). Canonical node lives here; llms-local-list over in llm-inference-wiki was given a back-link. elevenlabs reused, not re-paged. Document integrity flaws noted on the source page: a comparison-table row (Voxtral-4B-TTS-2603) pointing at an anchor with no detail section, and an empty “ComfyUI Integrations” heading. Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-08-03] lint | record the Kokoro cloning dispute on the disputed page itself
Follow-up to today’s free-voice-clone-list ingest. The contradiction was recorded in synthesis.md,
free-voice-clone-list, voice-cloning and supertonic-2 — but not on kokoro, which had
received only an updated: 2026-08-03 bump. That page still asserted “No voice cloning” flatly under a
fresh date, so it read as reviewed-and-confirmed on the day a source disputed it, and a reader landing
there directly saw no sign of the conflict. Added a “Disputed since 2026-08-03” paragraph that keeps this
wiki’s reading, states why (two sources discussing the tradeoff vs one unmethodical checkbox), names the
claim as load-bearing for voice-cloning and open-weight-tts, and records that it is resolvable
against the Hugging Face model card — which nobody here has checked. Related links extended.
Also fixed a miscount introduced in the same ingest: free-voice-clone-list said “Three sources here
say the opposite” and then named two. Corrected to two.
[2026-08-03] query | resolve the Kokoro cloning dispute against the primary
Asked to check the model card. Fetched huggingface.co/hexgrad/Kokoro-82M — the primary neither the
catalogue nor the roundups had consulted.
Resolved: Kokoro-82M does not support voice cloning. The card says so, and the repository settles it
structurally — voices/ holds exactly 54 .pt voicepack tensors, a voice is loaded by selecting one
of those files, and there is no reference-audio path because a voicepack is the speaker representation.
Apache-2.0; base model yl4579/StyleTTS2-LJSpeech, which independently confirms the StyleTTS2 lineage
this wiki already recorded. The two roundups were right; free-voice-clone-list‘s checkbox is wrong.
Second finding, against this wiki. kokoro read “~54 preset voices and ~15 languages.” The 54 is
exactly right. The ~15 languages is wrong: the card says 8, and the voicepack filenames agree — 9
language prefixes (a/b English US+GB, e Spanish, f French, h Hindi, i Italian, j Japanese, p Portuguese,
z Mandarin), i.e. 8 languages counting English once. The figure was uncited and appears to have been
wrong since the page was created in June. Corrected.
Updated: kokoro (dispute resolved + language correction), free-voice-clone-list (the row is
wrong, and it is the only row anyone checked), voice-cloning (now on primary evidence),
synthesis (contradiction moved to RESOLVED), index.
Kept rather than deleted, per record-don’t-overwrite: the resolution is more useful than a silent fix,
because of what it says about the corpus. The disputed cell was settleable in about a minute and nobody
had looked — not the catalogue, not the roundups, not this wiki, which had cited the claim as a founding
example since June. And the same check caught our own uncited number, which survived only because
nothing contradicted it. Standing lesson recorded in synthesis: when a primary is one fetch away, “two
secondary sources agree” is not the end of the inquiry.
[2026-08-03] lint | refresh the TTS leaderboard corner (quality cycle, staleness)
The quality pass flagged 19 volatile pages past their 30-day window here, the worst in the batch and structurally so — this spoke’s own manual says Elo/WER/latency “churn weekly.” Refreshed the two load-bearing pages against the live board rather than bumping dates. Refreshed: tts-arena-leaderboard against artificialanalysis.ai (2026-08-03 snapshot). Propagated: kokoro (Elo 1064 → 1055, $0.65 → $0.7/1M), fish-audio-s2-pro (1128 → 1123, rank now known: 20th), open-weight-tts, tts-benchmarks, synthesis, index. Kept the June snapshot beside the August one per record-don’t-overwrite — the movement is the finding. Three findings, two of them corrections to this wiki’s own framing.
- New #1, and it is the model the corpus dismissed. Speechify Simba 3.2 tops the board at 1229. tts-models-2026-benchmark mentioned “Speechify SIMBA 3.0” once, in a trailing clause, as the budget option. It is now first and the cheapest of the top ten ($10/1M vs ElevenLabs v3 at $100 for tenth). Price and rank have come apart at the top.
- “Trails but not by much” was wrong. open-weight-tts said the open field trails the closed frontier “but not by much,” and synthesis said it was “close behind and closing.” The board ranks the open leader 20th and kokoro 48th, and the gap to the leader widened over the summer (99 → 106 Elo) as the top rose and both open models drifted down. The error was reporting Elo deltas without ever recording rank, which made twentieth place read like a near miss. Corrected on both pages, with the old claim quoted rather than deleted. Note the branch split: “close and closing” still holds on STT; the synthesis had generalized recognition’s good news to all three.
- The “sources disagree” tension was mostly drift. Gemini 1216-vs-1217 and Fish S2 Pro 1123-vs-1128
were filed as vote-pool disagreement. Two months of movement (Gemini 1217 → 1212, Inworld Realtime
TTS-2 1206 → 1189 and six places) shows single-digit gaps between sources dated days apart are the
board moving. Reframed on tts-benchmarks: compare orderings, not points.
Method note recorded on tts-models-2026-benchmark: it is a fixed 2026-05-30 publication, so it
can only be superseded, never refreshed. Re-marked
freshness: stable(the artifact is stable) with a note that its volatile claims are superseded — otherwise the quality cycle re-flags it forever with no action available. That distinction is worth generalizing to other dated articles. Not done: the other ~17 stale pages here are individual model pages whose facts move more slowly than the board. Left for a later pass rather than touched to reset dates.
[2026-08-03] lint | clear the staleness backlog (17 pages) — and find the music branch was mis-cast
Follow-on to the leaderboard refresh. 17 pages remained past their volatile window.
First, a bug of mine. Three pages edited earlier today (fish-audio-s2-pro, tts-benchmarks,
synthesis.md) had their bodies changed without updated: being bumped, so they claimed June dates
while carrying August content — the same staleness-laundering problem in reverse. Fixed.
New pages: music-arena-leaderboard (Dataset, T2), mureka (SoftwareApplication).
Updated: suno, udio, musicgen, stable-audio, ai-music-generators-2026,
elevenlabs, synthesis, index.
The finding: this wiki’s entire music-branch cast came from one T4 marketing survey. Adding one
neutral board changed three things.
- udio is not suno‘s close rival. The page said “a slight quality gap” for clean rights. The board has it 16th of 18 on instrumental (954, below stable-audio 2.0) and 14th of 15 on vocals, ~235 Elo behind Suno V5.5. Udio’s rights case is untouched and still the reason to care — but “slight gap for lower legal risk” and “large gap for lower legal risk” are different recommendations, and the wiki was making the first.
- A top-two provider was missing entirely. mureka holds ranks 2 and 3 on both boards, V9 within 2 Elo of the leader on instrumental, and appeared nowhere in this hub. Paged, with the page saying plainly that two Elo scores is all this wiki knows — no architecture, licence or rights posture, which matters most in this branch.
- The open wedge does not exist in music. musicgen ranks last of eighteen (865, 300+
behind). Revised the synthesis’s “open wedge in every branch” read: STT strongest, TTS middling
(open leader 20th), music negligible.
Also recorded, not resolved: the wiki’s “suno v5 ~1293” has no counterpart on the neutral board
(Suno V5 = 1158/1097). Too large to be drift — different scale. Both kept, board preferred.
The generalizable failure, written into synthesis: tier discipline weighs what a source claims
and does nothing about what it omits, because an omission makes no claim to weigh. A T4 survey
silently supplied the set of things that exist, and two months of careful reasoning ran on an
unchecked cast list. Cheap defence: for any branch this wiki ranks, hold at least one neutral board —
to enumerate the field, not to source the claims. TTS had one and its error was merely numeric.
Freshness reclassification (the rest of the backlog). Five dated blog posts
(stt-apis-comparison, stable-audio-3, open-source-tts-models, open-source-stt-models,
ai-music-generators-2026) →
stable: a fixed publication can only be superseded, never refreshed, so marking it volatile re-flags it forever with no action available. Six model pages (orpheus, sesame-csm, misotts, audio-flamingo-3, gemini-live-3-5-translate, gpt-realtime-2) →stablewith an explicit delegation note: identity facts (architecture, parameters, licence) are stable, competitive standing is volatile and now lives on the dated board pages. This is a classification change, not a data refresh, and is logged as such. elevenlabs re-fetched: valuation unchanged at $11B, ~400 staff, plus a March 2026 $1B voice-restoration pledge — noted because it sits opposite the audio-deepfake harms its cloning enables. Left undone, deliberately: elevenlabs-expressive-mode is a live marketing page, so it genuinely can change and genuinely needs a fetch. Keptvolatileand annotated on the page that its figures are still the 2026-06-13 reading, rather than date-bumped. Backlog: 17 → 1.
[2026-08-03] ingest | Mureka platform docs — the branch’s runner-up, documented
Follow-up to the music-board refresh: mureka was paged the same day as a leaderboard entry with a name attached; its own docs were pulled to make it a real page. T3 — first-party documentation and marketing, fine for API surface, model names and pricing, and no evidence at all behind the rights claim. The Elo scores it is paged for stay sourced to music-arena-leaderboard (T2). Updated: mureka (rewritten from stub), ai-music-copyright (a third position on the axis), synthesis, index. Findings:
- Widest API surface in the branch. Song and instrumental generation, lyrics generation and extension, vocal cloning, song extension, recognition, description, transcription, soundtrack, remixing, stem separation — plus text-to-speech and podcast creation. Recognition/transcription /description are analysis, adjacent to audio-flamingo-3‘s comprehension corner; the TTS endpoints make Mureka the second provider after elevenlabs to span speech and music, which is evidence for the synthesis’s convergence read.
- The docs are two versions behind the board. They describe V7.5 and O1; the board ranks V9 and V8. So this wiki holds a neutral benchmark for a model the vendor has not documented. Recorded, not reconciled. (“O1” is a curation layer selecting among generated takes rather than a larger generator — noted as a pattern to watch.)
- Pricing undercuts both incumbents: free tier, Basic $8/mo (400 songs), Pro $24/mo (1,600 songs, WAV/stems, voice cloning) against suno and udio at $10/$30.
- Rights are asserted, not evidenced. “Every piece of music comes with full commercial rights…”, output “royalty-free”, Spotify/TikTok/YouTube named. Undisclosed: training-data sources or their licensing, rights-holder agreements, ownership transfer, any limit on the grant — and the operating company is not named on its own site or docs. A search result attributes it to Skywork AI (Kunlun Tech); recorded as unverified because no first-party page says so. The load-bearing distinction, now on ai-music-copyright: an output-side promise (“you may use this commercially”) is not an input-side claim (“we were entitled to train on this”), and the lawsuits are about the second. suno and udio are fought and settled on the input side, stable-audio answers with a licensed dataset. Mureka answers only the output question. So the axis gains a third position — not contested, not licensed, undisclosed — and the branch’s quality runner-up is the one a buyer comparing on legal safety cannot place at all. Also noted on the page: vocal cloning ships on the Pro tier with no consent language anywhere in the material read — the same supply-side silence free-voice-clone-list found across the open catalogue and audio-deepfake tracks. Entity: operating company not paged — it is not named by any first-party source. Verify deferred per hub policy (content-only). avoid-ai-writing run.
[2026-08-04] ingest | abogen (denizsafak)
Routed here by the hub (runner-up: none). T1 — the project’s own repository. 1 new page: abogen. Updated kokoro (a downstream note) and index.
Dedup mattered: abogen runs on kokoro, already held here as the efficiency leader, so this is an application over a model the spoke has rather than a new engine.
Second source in the consumption layer, after vocalinux on the STT side, and the pair is starting to say something. This spoke measures models — naturalness, Elo, CER, price per million characters — and nothing measures the layer that turns a model into a usable artifact. abogen’s actual work is chapter detection, M4B chaptering, caption alignment, ePub/PDF extraction and text normalization; none of it is synthesis, and all of it is what stands between kokoro and an audiobook.
Kokoro’s tradeoff propagates into the product. The model drops voice-cloning to hit its footprint; abogen answers with a voice mixer blending the 54 presets. Recorded on kokoro too, because it shows the efficiency-vs-controllability split is not a benchmark artifact — it reaches the end user.
Flagged, unmeasured: the Web UI puts LLM-assisted normalization in front of the speech model, a quality lever sitting entirely outside the TTS benchmarks this spoke tracks. No scores of any kind in the source; it is a tool, not an evaluation.
[2026-08-05] ingest | SenseVoice (Alibaba / FunAudioLLM)
Routed from the hub (runner-up: llm-providers-wiki). New page sensevoice (source, T3 — first-party README, self-run numbers only). Multi-task speech understanding: ASR (zh/yue/en/ja/ko) + language ID + 7-category emotion + audio-event detection in one non-autoregressive pass. MIT code, FunASR licence on the weights, ~9k★.
Filed as a new row in the models index rather than under STT, because it is not competing on WER or on size like whisper, canary-qwen, vosk, moonshine and vibevoice-asr-bitnet. It answers what language, what emotion and what background sounds, with the transcript as one output of four — the audio-flamingo-3 job done with fixed heads instead of an LLM, and fast precisely because it drops the autoregressive decoder that reasoning needs. The two are not substitutes and the page says so.
Three things recorded against it: only SenseVoiceSmall is released while the benchmark tables include an unavailable Large; the repo sits under a QwenAudio org whose relationship to the FunAudioLLM team is not established from the page; and the emotion output has no per-category accuracy, no demographics and no cross-corpus evaluation, which is where SER results usually fall apart.
[2026-08-06] ingest | VoiceCraft (jasonppy/VoiceCraft)
Routed here by the hub. New page: voicecraft (T1, repo + arXiv:2403.16973).
It adds a task the spoke did not have. Everything paged here so far produces speech; VoiceCraft’s first job is editing an existing recording — token infilling over codec tokens, so words inside real audio can be replaced with the speaker, prosody and room left intact. Zero-shot text-to-speech is the degenerate case of the same model. voice-cloning got a section on the distinction: “can it clone a voice” now has a sharper sibling, “can it alter a recording that already exists”, and no leaderboard in this wiki ranks the second.
Architecture sits on the lineage already paged: 4 RVQ codebooks × 2048 codes in the EnCodec design (neural-audio-codec), against misotts’ 32 codebooks — much less codebook depth, which is the editing-versus-fidelity trade.
The finding worth the ingest went into audio-deepfake. That page had recorded a gap: capability ships without consent language (luxtts flagged for it, the 36-model census carrying almost none). VoiceCraft is the counter-example — an explicit prohibition on generating or editing anyone’s speech without consent, naming political figures and celebrities, plus non-commercial licences with actual teeth. And it is the project with the most consequential capability of the lot, which is the opposite of what the pattern predicted. Recorded as a norm-not-law point: a README line has no enforcement.
Held back: every benchmark claim is the authors’. Nothing independent, tts-benchmarks has no entry,
and the README’s last improvement is dated March 2025 — a well-starred repo that has been quiet for
over a year. freshness: volatile for that reason.
[2026-08-09] ingest | LibriSpeech, Common Voice, FLEURS — the corpora behind every WER here (via research pass)
Coverage edge 6 closed with the three dataset papers, all T1: Panayotov et al. (ICASSP 2015), Ardila et al. (2019, Mozilla), Conneau et al. (2022, Google). The LibriSpeech PDF would not read through WebFetch and was extracted locally with pypdf.
New pages: librispeech, common-voice, fleurs.
The gap was a provenance hole, not a coverage one. open-asr-models-2026-comparison, stt-apis-comparison and the leaderboard pages all rank systems on these three sets, and the wiki could not say what any of them contained.
Two findings worth more than the page count.
First, LibriSpeech’s clean label is defined by a model, not by acoustics: speakers were transcribed
by an acoustic model trained on WSJ si-84, ranked by that model’s WER, and split at the median. So
test-clean is the half of the speakers a 2015 system already found easy, and the “other” dev/test
sets were then chosen deliberately to be harder. A 2% test-clean figure is a figure on pre-selected
material.
Second, all three corpora are read speech — audiobooks, prompts, and translated sentences read aloud. Every WER this spoke quotes comes from someone reading, while the products are sold for meetings and calls. That is now edge 6’s successor: a conversational or far-field benchmark.
Recorded and not resolved: Common Voice is CC0, which is why it is everywhere, including in the training data of models that then report scores on it. No source here measures that overlap.
Entities: 0 created. Mozilla and Google both already exist as nodes in sibling spokes and are linked cross-wiki rather than duplicated.
[2026-08-12] lint | Quality cycle: freshness only, zero new pages
Both volatile pages past the 60-day window were regraded stable, and neither needed a refresh:
- musicgen — a Meta announcement from August 2023. A dated announcement records what was said on that date; re-reading the URL returns the same text, so the flag could never be cleared.
- wavenet — the encyclopedic page on a 2016 model. Nothing on it is a current-state quantity a reader would act on.
This is the case ../QUALITY.md grades stable, not a decision to stop checking them. stale: 2 → 0.
The spoke’s own top edge — a neutral latency benchmark with stated hardware and concurrency — is on cooldown until 2026-08-22 after the 2026-08-08 hunt found nothing above the bar. Its best coverage edge (a conversational or far-field corpus: CHiME, AMI, Switchboard, against the three read-speech corpora here) was not hunted this cycle: the batch had already spent its source budget in five other spokes, and padding a sixth is what the rubric forbids. Named here so the next cycle can take it first.