KittenTTS
A 15M-parameter open-weight-tts model from KittenML, Apache-2.0, described as “state-of-the-art TTS model under 25MB … running without GPU on any device” (free-voice-clone-list, v0.8.1, February 2026). The parameter range given is 15M–80M, so 15M is the smallest variant rather than the whole family.
It takes the bottom of this wiki’s size ladder from kokoro‘s 82M down by roughly a factor of five, and it does so without dropping voice-cloning — the list marks zero-shot cloning and emotion control present, streaming supported, English plus unspecified others.
Why that combination matters
kokoro is paged here as the clean case of the efficiency-versus-controllability trade: it hit 82M by shipping fixed preset voices and no cloning. KittenTTS claims a fifth of that footprint with cloning, which — if it holds up — breaks the correlation voice-cloning describes between cloning and larger models conditioned on reference audio.
Nothing here tests that claim. The only source is a curated list with no MOS, no WER, no Elo and no listening test, reporting the model author’s own description. The <25 MB and no-GPU figures are checkable against the repo; the quality implied by “state-of-the-art” is not checked anywhere in this wiki. Treat it as a claim with a size attached, not a ranking.
Related
kokoro · supertonic-2 · open-weight-tts · voice-cloning · text-to-speech · free-voice-clone-list · tts-benchmarks