人工标注是大量自然语言处理研究的实证基础,从数据集构建到模型评估皆不例外,但论文往往未明确说明标注由谁完成以及标注过程如何受控。我们首次对主要自然语言处理会议中的人工标注报告进行了大规模、任务层面的审计,探究哪些标注细节被记录、哪些被遗漏,以及报告方式随时间、主题、会议类型和人工判断预期用途的变化。
我们引入了一套统一的标注报告实践分类体系,并针对 Annotated-gold(一个由 41 篇论文和 72 个标注任务组成、经人工裁决的黄金标准数据集)验证了基于大语言模型的辅助提取流程,其中最佳模型与裁决标签达到与人类相当的一致性,Krippendorff 的 alpha 值为 0.606,而人类之间的一致性为 0.585。
利用该流程,我们构建了 Annotated-llm 数据集,覆盖 2018 至 2025 年 ACL 会议论文,从 1603 篇论文中提取出 2667 个标注任务,并发现论文经常报告招募策略、标注者专业知识和标注量等操作细节,但往往省略评估标注有效性所需的细节,包括培训、语言能力、报酬、社会人口统计信息、裁决和一致性数值,在模型评估研究中尤为突出。
我们的结果表明,自然语言处理领域的标注报告质量随时间有所改善,但仍不均衡,同时我们建立了一个可扩展的框架和最低限度报告建议,旨在使人工标注更加可靠、可复现和可解释。
Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled. We provide the first large-scale, task-level audit of human annotation reporting across major NLP venues, asking which annotation details are documented, which are missing, and how reporting varies across time, topic, venue, and intended use of human judgment. We introduce a unified taxonomy of annotation-reporting practices and validate an LLM-assisted extraction pipeline against Annotated-gold, a human-adjudicated gold standard of 41 papers and 72 annotation tasks, where the best model reaches human-comparable agreement with adjudicated labels, with Krippendorff's alpha of 0.606 versus 0.585 for human-human agreement.
Using this pipeline, we construct Annotated-llm, a dataset covering ACL-venue papers from 2018-2025, with 2,667 extracted annotation tasks from 1,603 papers, and find that papers frequently report operational details such as recruitment strategies, annotator expertise, and annotation volume, but often omit details needed to assess annotation validity, including training, language proficiency, compensation, socio-demographics, adjudication, and agreement values, especially in model-evaluation studies. Our results show that annotation reporting in NLP has improved over time but remains uneven, and they establish a scalable framework and bare-minimum reporting recommendations for making human annotation more reliable, reproducible, and interpretable.