Spokes.wiki Search About
Scholarly Article source ↗ source url updated Sat Aug 08 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

“Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI

nithya-sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Kumar Paritosh and Lora Mois Aroyo, CHI 2021 (ACM), doi 10.1145/3411764.3445518. Interviews with 53 AI practitioners in India, East and West African countries, and the USA, all working in high-stakes domains: cancer detection, suicide prevention, landslide detection, wildlife poaching, loan allocation.

This is the first source in this wiki about the data layer itself rather than about a model that happens to need data.

The finding

A data cascade is a compounding failure: a decision taken upstream about data — a proxy label, an unrepresentative sample, a hand-off to people who were never told what the model would do with their annotations — produces downstream damage that surfaces much later, usually as a model that will not work in the field.

  • 92% of practitioners reported at least one cascade.
  • 45.3% reported two or more in a single project.
  • The paper’s characterization: cascades are “pervasive, invisible, delayed, but often avoidable.”

The framing sentence is the one worth keeping, because it is a claim about incentives rather than technique: “data is the most under-valued and de-glamorised aspect of AI.” The title says the same thing in a practitioner’s own words.

Why this closes a gap here rather than adding to a pile

The founding gap in this wiki was stated as no source on training a model from data end to end — no pretraining, no data collection or labelling, no evaluation methodology. Evaluation closed in a tutorial (timesfm-2-5-forecasting-tutorial), the training loop closed in a tutorial collection (transformers-tutorials), and both arrived as method carried incidentally by teaching material. The wiki’s reading was that the middle layer has no literature of its own.

That reading was wrong, and this is the correction: the data layer does have a literature, it is in HCI rather than ML, and it studies practitioners rather than models. That is why searching the ML side of the field never turned it up.

Against the corpus’s existing failure claim

ai-projects-fail-infrastructure-people holds that AI projects fail on infrastructure and people. It is T4, its body was never readable, and this wiki has been carrying its headline as a framing rather than an argument.

Data Cascades does not refute it so much as reach a different layer with better evidence. Both are saying the failure is not model quality. They disagree about where it is: the T4 headline says the serving stack and the org chart, the CHI paper says decisions taken over the data before any model existed. 53 interviews against an unreadable article is not a close contest, but the corpus keeps both — see demo-to-production-gap, where the disagreement is recorded.

The number to be careful with is 92%. It is prevalence among practitioners who were interviewed about their data practice, in high-stakes domains, recruited for that purpose. It is not a population failure rate for AI projects, and the corpus’s standing open question — how often projects fail, against what definition — is still open.

pervasive-label-errors · training-data-quality · demo-to-production-gap · ai-projects-fail-infrastructure-people · nithya-sambasivan · machine-learning · synthesis