paper-with-me

홈 › Papers

Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation

2026-05-25 · Yi Ren arxiv

Evaluating structured-information extraction from historical documents at scale requires high-precision ground-truth annotations, yet traditional manual labeling is expensive and fully automated pipelines built on large language models are prone to hallucination. We propose Double Triangle Annotation, a two-layer human-in-the-loop framework that leverages cross-model consensus to automate the majority of annotation work while ensuring high-precision outputs. In the first layer, two architecturally independent Multimodal Large Language Models annotate each document in parallel; when they agree, the label is auto-accepted, and disagreements are routed to a human jury. A second layer cross-checks two such systems against each other, escalating residual conflicts to a domain expert. The framework rests on a single assumption -- error independence between models -- requires no distributional priors or task-specific calibration, and becomes more autonomous as model capability improves. On the Guides Rosenwald, a corpus of French medical directories spanning 1887-1906, the framework achieves a final Word Error Rate of 0.003. Applied at scale, model consensus auto-accepts over 85% of 13,595 fields. We release the resulting benchmark -- the first structured-extraction ground truth for the Rosenwald Guides -- to support future work on historical document processing.

📄 PDF Abstract BibTeX arXiv:2605.25781

Code (0)

등록된 구현이 없습니다.

Tasks

Information Extraction

Similar Papers 제목 키워드 기반

Scalable Data Annotation Pipeline for High-Quality Large Speech Datasets Development

2021-09-01 · Mingkuan Liu, Chi Zhang, Hua Xing, Chao Feng 외

This paper introduces a human-in-the-loop (HITL) data annotation pipeline to generate high-quality, large-scale speech datasets. The pipeline combines human and machine advantages to more quickly, accurately, and cost-ef…

Vocal Bursts Intensity Prediction

The SemDaX Corpus ― Sense Annotations with Scalable Sense Inventories

2016-05-01 · LREC 2016 5 · Bolette Pedersen, Anna Braasch, Anders Johannsen, H{\'e}ctor Mart{\'\i}nez Alonso 외

We launch the SemDaX corpus which is a recently completed Danish human-annotated corpus available through a CLARIN academic license. The corpus includes approx. 90,000 words, comprises six textual domains, and is annotat…

Pushing the Limit of Range Resolution Beyond Bandwidth Constraint with Triangle FMCW

2024-04-06 · Yanbo Zhang

This paper proposes a novel signal processing technique that doubles the range resolution of FMCW~(Frequency Modulated Continuous Wave) sensing without increasing the required bandwidth. The proposed design overcomes the…

Deep Reinforcement Learning for Efficient Measurement of Quantum Devices

2020-09-30 · V. Nguyen, S. B. Orbell, D. T. Lennon, H. Moon 외

Deep reinforcement learning is an emerging machine learning approach which can teach a computer to learn from their actions and rewards similar to the way humans learn from experience. It offers many advantages in automa…

Decision MakingDeep Reinforcement LearningNavigatereinforcement-learning+2

Scalable Agent-Based Modeling for Complex Financial Market Simulations

2023-12-22 · Aaron Wheeler, Jeffrey D. Varner

In this study, we developed a computational framework for simulating large-scale agent-based financial markets. Our platform supports trading multiple simultaneous assets and leverages distributed computing to scale the …

Decision MakingDistributed Computing