Mean Opinion Score (MOS)
The definitional anchor under one of the tts-benchmarks metrics. MOS is the arithmetic mean of listeners’ naturalness ratings of synthesized clips — the tts-benchmarks page quotes kokoro at “~4.5 MOS” but didn’t pin down what the scale is; this source does.
The scale
A standardized five-point absolute-category scale, where each rating maps to a label:
| Rating | Label |
|---|---|
| 5 | Excellent |
| 4 | Good |
| 3 | Fair |
| 2 | Poor |
| 1 | Bad |
The score is MOS = (R₁ + R₂ + … + Rₙ) / N — the mean of N subjects’ individual ratings R.
Where it comes from, and the catch
MOS originates in subjective telephony measurement, standardized by ITU-T Recommendation P.800 — listeners in controlled acoustic conditions rating call quality. TTS borrowed it to score perceived naturalness. The catch the tts-benchmarks caveats already gesture at is made explicit here: MOS values should not be compared directly across experiments unless the studies were designed to be comparable, because context (listener pool, clips, instructions) shifts the absolute numbers. A “4.5 MOS” is only meaningful against the other systems in the same test — exactly why the spoke treats these as dated, in-context snapshots rather than a global scale.
Tier
T2 — reference encyclopedia entry tracing to the primary standard (ITU-T P.800). Used for a stable definition, not a volatile figure; the scale and non-comparability point are standards-grounded.