Spokes.wiki Search About
Dataset source ↗ source url updated Sun Aug 09 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Common Voice

Ardila, Branson, Davis, Henretty, Kohler, Meyer, Morais, Saunders, Tyers and Weber (Mozilla), Common Voice: A Massively-Multilingual Speech Corpus, 2019 (LREC 2020). T1 — the corpus paper.

2,500 hours of transcribed speech from over 50,000 contributors, released CC0 — public domain, not merely permissive. At the time of writing: 29 languages in the release, 38 collecting data. The paper’s claim is “the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages.”

Collection and validation are both crowdsourced: contributors read prompts, other contributors vote on whether a clip matches its text. That design is the source of both its virtues and its weaknesses. Coverage extends to languages no commercial vendor would fund, and it grows continuously — but the recordings are consumer microphones in uncontrolled rooms, the speaker distribution follows whoever volunteered, and quality control is a majority vote rather than a trained annotator.

Against the other two

librispeech is one language, read from books, aligned mechanically against a known text. Common Voice is many languages, read from short prompts, validated by strangers. fleurs is parallel across 102 languages with about 12 hours each. So a model reporting on all three is reporting on read speech three times over, in three different recording regimes — none of them conversational, none far-field.

The CC0 licence is the reason this dataset shows up in nearly every open model’s training set, and therefore a live contamination question for anyone quoting Common Voice as a test set. The spoke holds no source measuring that overlap.

Cross-spoke

Mozilla also maintains gecko and firefox (web-browsers-wiki) — the same organization’s public-good funding argument, in a different field.