Spokes.wiki Search About
Scholarly Article source ↗ source url updated Mon Aug 10 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou and Steffen Eger — NLLG Lab, University of Technology Nuremberg, with the Interdisciplinary Transformation University, Austria. arXiv 2606.02255, 1 June 2026. Preprint, not yet peer-reviewed; tiered T1 as a primary empirical study reporting its own method and data.

This wiki went looking for a documented labelling operation — guidelines, annotator training, sampling design, cost — and kept not finding one. This paper explains why, by measuring how often those things are written down at all.

What it did

A task-level audit of annotation reporting across ACL-venue papers, 2018–2025. Two datasets: ANNOTATED GOLD, a human-adjudicated standard of 41 papers and 72 annotation tasks used to validate the pipeline, and ANNOTATED LLM, 2,667 extracted annotation tasks from 1,603 papers.

The unit is the task, not the paper, and the authors argue that choice matters: one paper may run several annotation tasks with different annotators, instructions and controls, so counting papers hides exactly the variation the audit is about.

Extraction is LLM-assisted, and they validate it rather than assert it — the best model reaches Krippendorff’s α = 0.606 against adjudicated labels, versus 0.585 for human–human agreement on the same data. Reliability is uneven by field: demographics, compensation and IAA metrics score α > 0.8, while categories needing interpretation score lower.

The numbers

Reporting rates across the 2,667 tasks:

ReportedRate
Total items annotated90%
Recruitment86%
Annotator expertise level86%
Annotations per item65%
Total annotators63%
Compensation56%
Education40%
IAA metric40%
IAA value35%
Guidelines released34%
Crowd screening25%
Quality control24%
Language proficiency24%
Items per annotator23%
Adjudication20%
Annotator training19%
Nation of residence14%
Gender6%
Age5%
Nation of origin2%
Political orientation1%

The paper’s own framing of the split: papers “frequently report operational details such as recruitment strategies, annotator expertise, and annotation volume, but often omit details needed to assess annotation validity, including training, language proficiency, compensation, socio-demographics, adjudication, and agreement values, especially in model-evaluation studies.” Precise figures for the weak end: training 18.7%, language proficiency 24.0%, released guidelines 34.1% — against recruitment 90.4%, expertise 86.5%, total items 86.0%.

Reporting has improved over the period but remains uneven, and the authors note the absence of evidence that the ACL Responsible NLP Checklist produced measurable gains.

Why this closes the wiki’s gap the wrong way round

The training-data-quality thread has been assembled from sources that each identify a failure and none of which document a working practice. data-cascades shows the downstream damage, pervasive-label-errors shows the labels are wrong at measurable rates, and measuring-annotator-agreement supplies an instrument while finding it ill-suited to most real label shapes. The open question asked for the practice itself.

This paper says the practice is mostly not published. Fewer than one task in five reports whether annotators were trained; two in three do not release the guidelines; and an agreement value — the output of the instrument measuring-annotator-agreement analyses — appears for 35%. So the corpus was not failing to find the documented operation through bad searching. The base rate is low.

That reframes rather than settles the question, and the reframing is sharper than the original. Agreement statistics are the field’s standard evidence of label quality, and the two numbers sit badly together: the instrument is reported for 40% of tasks, its value for 35%, and the conditions that would let a reader interpret that value — who the annotators were, whether they were trained, what they were told — are reported less often still. A kappa without the guidelines that produced it is a number about an unspecified procedure.

Note what this source is and is not. It is a meta-study of reporting, not the documented operation the edge asked for. It measures the shape of the absence with authority; it does not supply the missing artifact. The edge is re-specified rather than closed — see synthesis.

training-data-quality · measuring-annotator-agreement · data-cascades · pervasive-label-errors · synthesis