Spokes.wiki Search About
Collection source ↗ source url updated Mon Aug 03 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

free-voice-clone (0xSojalSec’s list)

A single 59 KB README cataloguing 53 open-weight audio models across five sections — TTS, music generation, anything-to-audio, restoration/enhancement, and ASR — each with a comparison table and a collapsible detail block (parameters, languages, streaming, licence, links to GitHub and Hugging Face). 450 stars, 83 forks, no licence of its own. Maintained by 0xsojalsec, the same curator behind the LLMs-local list this hub already holds in llm-inference-wiki (llms-local-list).

It is a census, not a benchmark. No Elo, no MOS, no WER, no listening test, and no stated inclusion criteria — a model is here because the maintainer added it.

What the census shows

Permissive licensing is the norm, overwhelmingly. Of the 36 TTS entries: 27 Apache-2.0, 4 MIT, 1 MIT/Apache-2.0, 2 Liquid’s LFM, 1 Research License (fish-audio-s2-pro), 1 OpenRAIL-M (supertonic-2). The restrictive corner this wiki has tracked in open-weight-tts is real but small — and it is where the quality leader sits, which is exactly why a census can’t settle the ceiling question.

The other branches are less uniform. Anything-to-audio is Apache-2.0 with one Research Only (HunyuanVideo-Foley). Restoration has NVIDIA A2SB under NVIDIA Non-Commercial beside Apache-2.0 NovaSR and MIT AudioSR. Music is MIT/Apache-2.0 plus Stability AI’s own terms for Foundation-1.

Voice cloning is the default, not a differentiator. 35 of 36 TTS entries claim zero-shot cloning, and the sole exception is supertonic-2. The list sorts on languages, streaming and licence instead, which is what a capability looks like once it is table stakes — a shift from the field voice-cloning describes, where cloning was the dividing line.

One row is simply wrong, and it was checked. The table marks kokoro as having voice cloning, against tts-models-2026-benchmark and open-source-tts-models. Resolved 2026-08-03 against the primary: the Hugging Face model card states Kokoro-82M does not support voice cloning, and its voices/ directory holds exactly 54 .pt voicepack tensors — a voice is selected by loading one, and there is no reference-audio path at all.

That matters more than one bad cell. It is the only row anyone has verified, it was verifiable in about a minute, and it is wrong — which is the strongest available reason to treat the other 35 checkboxes as unverified claims rather than as data. The same check also caught an error on this wiki’s own kokoro page (an unsourced “~15 languages” against the card’s 8), so the catalogue is not uniquely careless; it is that nobody, here included, had looked.

The small pole moved down again. kittentts claims 15M parameters under 25 MB running without a GPU; supertonic-2 claims 66M at up to 167× realtime. Both sit below kokoro‘s 82M, which this wiki has called the efficiency leader since June.

The thing that isn’t in it

Across 58,997 characters and 53 models there is no occurrence of “consent”, “ethics”, “watermark”, “misuse”, “deepfake”, “responsible”, “abuse”, or “disclaimer”. Not a weak treatment — zero. A document whose title is voice-clone and whose entries advertise cloning from a three-to-ten-second reference sample carries no line about whose voice you may clone.

This is the same absence luxtts was flagged for on 2026-07-29, and finding it again across an entire catalogue makes it a property of the genre rather than one project’s oversight. See audio-deepfake, where this wiki keeps the harms this silence sits next to.

One incidental counterpoint: the single entry under a use-restricting licence (OpenRAIL-M, supertonic-2) is also the single entry that cannot clone a voice. The list gives no reason for either fact and one case supports no conclusion — recorded because it is the only place in the document where use restrictions appear at all.

Freshness and quality

The repo was last pushed 2026-04-09 and read here on 2026-08-03, so it is a roughly four-month-old snapshot of a field this wiki’s own conventions call weekly-volatile. Release dates inside the entries run to February 2026. Anything after early April is missing by construction — the list predates lyria-3-5 and the July ASR reordering in open-asr-models-2026-comparison.

Two integrity flaws in the document itself: the TTS comparison table’s first row, Voxtral-4B-TTS-2603, links to an anchor that has no detail section, and the “ComfyUI Integrations” heading under Additional Resources is empty.

T4 — curated by one pseudonymous maintainer, no methodology, no measurements, no editorial process, and stale. Its value is coverage and the licence census, both of which are checkable against the linked repos. Every capability figure in it is the model author’s own claim passed through unverified, and this wiki records them as such.

open-weight-tts · voice-cloning · audio-deepfake · tts-models-2026-benchmark · open-source-tts-models · kittentts · supertonic-2 · kokoro · fish-audio-s2-pro · luxtts · 0xsojalsec · llms-local-list · synthesis