paper-with-me

홈 › Papers

From Documents to Spans: Scalable Supervision for Evidence-Based ICD Coding with LLMs

2026-03-16 · Xu Zhang, Wenxin Ma, Chenxu Wu, Rongsheng Wang, Zhiyang He, Xiaodong Tao, Kun Zhang, S. Kevin Zhou arxiv

International Classification of Diseases (ICD) coding assigns diagnosis codes to clinical documents and is essential for healthcare billing and clinical analysis. Reliable coding requires that each predicted code be supported by explicit textual evidence. However, existing public datasets provide only code labels, without evidence annotations, limiting models' ability to learn evidence-grounded predictions. In this work, we argue that dense, document-level evidence annotation is not always necessary for learning evidence-based coding. Instead, models can learn code-specific evidence patterns from local spans and use these patterns to support document-level evidence-based coding. Based on this insight, we propose Span-Centric Learning (SCL), a training framework that strengthens LLMs' coding ability at the span level and transfers this capability to full clinical documents. Specifically, we use a small set of annotated documents to supervise evidence recognition, aggregation, and code assignment, while leveraging a large collection of lightweight evidence spans to reinforce span-level reasoning. Due to their compactness, span annotations are scalable and can be further augmented through synthesis. Under the same Llama3.1-8B backbone, our approach achieves an 8.2-point improvement in macro-F1 at only 20% of the training cost of standard SFT, and provides explicit supporting evidence for each predicted code, enabling human auditing and revision.

📄 PDF Abstract BibTeX arXiv:2603.15270

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MDACE: MIMIC Documents Annotated with Code Evidence

2023-07-07 · ACL 2023 7 · Hua Cheng, Rana Jafari, April Russell, Russell Klopfer 외

We introduce a dataset for evidence/rationale extraction on an extreme multi-label classification task over long medical documents. One such task is Computer-Assisted Coding (CAC) which has improved significantly in rece…

Document ClassificationExtreme Multi-Label ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

Long Context Question Answering via Supervised Contrastive Learning

2021-12-16 · NAACL 2022 7 · Avi Caciularu, Ido Dagan, Jacob Goldberger, Arman Cohan

Long-context question answering (QA) tasks require reasoning over a long document or multiple documents. Addressing these tasks often benefits from identifying a set of evidence spans (e.g., sentences), which provide sup…

Contrastive LearningQuestion AnsweringVideo Generation

Seq2seq is All You Need for Coreference Resolution

2023-10-20 · Wenzheng Zhang, Sam Wiseman, Karl Stratos

Existing works on coreference resolution suggest that task-specific models are necessary to achieve state-of-the-art performance. In this work, we present compelling evidence that such models are not necessary. We finetu…

Allcoreference-resolutionCoreference Resolution

ECPO: Evidence-Coupled Policy Optimization for Evidence-Certified Candidate Ranking

2026-05-21 · Miaobo Hu, Shuhao Hu, BoKun Wang, Yina Sa 외 arxiv

Ranking systems used in decision-support settings should not only order candidates but also expose evidence that can be independently checked. We study evidence-certified candidate ranking: given an intent_id, a predefin…

Explicit Evidence Grounding via Structured Inline Citation Generation

2026-06-05 · Anar Yeginbergen, Amelie Wührl, Anna Rogers, Rodrigo Agerri arxiv

As AI systems become more widely adopted, the demand for factual and faithful generation grows. Properly attributing information through citations becomes, therefore, crucial. This work introduces FullCite, a framework t…

Question Answering