FLEURS
Conneau, Ma, Khanuja, Zhang, Axelrod, Dalmia, Riesa, Rivera and Bapna (Google), FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech, 2022. T1 — the benchmark paper.
An n-way parallel speech dataset in 102 languages, about 12 hours per language, built by recording the sentences of the FLoRes-101 machine-translation benchmark. Parallel is the whole point: the same sentences exist in every language, so a score in Yoruba and a score in Finnish are measured on the same content. It supports ASR, speech language identification, speech translation and retrieval, and the stated aim is to “enable speech technology in more languages and catalyze research in low-resource speech understanding.”
What its shape implies
Twelve hours per language is an evaluation budget, not a training one — FLEURS is a test set with a small train split, which is why the paper frames it as few-shot. And because the sentences come from a translation benchmark, the domain is written, translated prose read aloud: no spontaneous speech, no code-switching, no dialect variation within a language. A high FLEURS score means a model handles many languages’ phonetics on clean read sentences. It does not mean the model works for speakers of those languages in their own conversational register.
Against the other two
librispeech measures English audiobook reading in depth; common-voice measures crowd-recorded prompts in breadth with uneven quality; FLEURS measures breadth on identical content with even, small per-language coverage. The three are complementary by construction and identical in one respect that the spoke should keep saying out loud: all three are read speech. Every headline WER in this wiki was earned on someone reading.