Spokes.wiki Search About
Scholarly Article source ↗ source url updated Thu Jun 18 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

XTTS: A Massively Multilingual Zero-Shot TTS Model (Casanova et al., 2024)

Edresson Casanova et al. (Coqui), arXiv:2406.04904, INTERSPEECH 2024 — a peer-reviewed primary for the zero-shot voice-cloning capability the voice-cloning page lists only from roundups (open-source-tts-models, tts-models-2026-benchmark). XTTS is the widely-used open zero-shot cloner that predates and seeds much of the open field.

What it is

A massively multilingual zero-shot TTS model: it synthesizes speech in a target speaker’s voice from a short reference clip, with no per-voice fine-tuning — the defining property the voice-cloning page calls the clean dividing line. Per the abstract it “builds upon the Tortoise model and adds several novel modifications to enable multilingual training, improve voice cloning, and enable faster training and inference.”

The facts that matter

  • 16 languages, SOTA in most. “XTTS was trained in 16 languages and achieved state-of-the-art (SOTA) results in most of them” — establishing that one zero-shot model can clone across many languages, the multilingual-cloning point the voice-cloning page didn’t have a primary for.
  • Publicly released. The authors commit to “making publicly available the XTTS system” — an open zero-shot cloner, which puts it on the permissive side of the open-weight-tts license split and makes it a concrete instance of “cloning is available off-the-shelf,” sharpening the consent concern the page flags.
  • Lineage from Tortoise. It descends from Tortoise-TTS, locating it in the autoregressive-cloning line rather than the codec-token line — a different architectural route to the same capability.

(The abstract states the language count, SOTA claim, Tortoise lineage, and public release; it does not detail the conditioning encoder or per-language metrics, so those are left out.)

Tier

T1 — peer-reviewed primary (INTERSPEECH 2024) from the model’s authors. Figures (16 languages, SOTA, public release, Tortoise lineage) from the abstract; finer architecture left unclaimed.

voice-cloning · text-to-speech · open-weight-tts · open-source-tts-models · speech-audio-ai