paper-with-me

홈 › Papers

When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity

2025-10-20 · Nisrine Rair, Alban Goupil, Valeriu Vrabie, Emmanuel Chochoy arxiv

Language models are often evaluated with scalar metrics like accuracy, but such measures fail to capture how models internally represent ambiguity, especially when human annotators disagree. We propose a topological perspective to analyze how fine-tuned models encode ambiguity and more generally instances. Applied to RoBERTa-Large on the MD-Offense dataset, Mapper, a tool from topological data analysis, reveals that fine-tuning restructures embedding space into modular, non-convex regions aligned with model predictions, even for highly ambiguous cases. Over $98\%$ of connected components exhibit $\geq 90\%$ prediction purity, yet alignment with ground-truth labels drops in ambiguous data, surfacing a hidden tension between structural confidence and label uncertainty. Unlike traditional tools such as PCA or UMAP, Mapper captures this geometry directly uncovering decision regions, boundary collapses, and overconfident clusters. Our findings position Mapper as a powerful diagnostic tool for understanding how models resolve ambiguity. Beyond visualization, it also enables topological metrics that may inform proactive modeling strategies in subjective NLP tasks.

📄 PDF Abstract BibTeX arXiv:2510.17548

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Everyone's Voice Matters: Quantifying Annotation Disagreement Using Demographic Information

2023-01-12 · Ruyuan Wan, Jaehyung Kim, Dongyeop Kang

In NLP annotation, it is common to have multiple annotators label the text and then obtain the ground truth labels based on the agreement of major annotators. However, annotators are individuals with different background…

Cover Learning for Large-Scale Topology Representation

2025-03-12 · Luis Scoccola, Uzu Lim, Heather A. Harrington

Classical unsupervised learning methods like clustering and linear dimensionality reduction parametrize large-scale geometry when it is discrete or linear, while more modern methods from manifold learning find low dimens…

Dimensionality ReductionTopological Data Analysis

The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

2024-05-09 · Eve Fleisig, Su Lin Blodgett, Dan Klein, Zeerak Talat

Longstanding data labeling practices in machine learning involve collecting and aggregating labels from multiple annotators. But what should we do when annotators disagree? Though annotator disagreement has long been see…

Position

When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization

2026-09-06 · Eyal Hanania, Daniel Arkushin, Naveh Ayal, Jonathan Benvenisti 외 arxiv

Annotators routinely disagree on laughter boundaries and subtle chuckles, yet temporal laughter localization typically evaluates against a single reference annotation. We show that this disagreement is structured rather …

Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the Truth

2020-05-01 · ICLR 2020 1 · Igor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei 외

In most machine learning tasks unambiguous ground truth labels can easily be acquired. However, this luxury is often not afforded to many high-stakes, real-world scenarios such as medical image interpretation, where even…

BIG-bench Machine LearningBinary Classification