It Takes Two to Tango: Navigating Conceptualizations of NLP Tasks and Measurements of Performance
Progress in NLP is increasingly measured through benchmarks; hence, contextualizing progress requires understanding when and why practitioners may disagree about the validity of benchmarks. We develop a taxonomy of disagreement, drawing on tools from measurement modeling, and distinguish between two types of disagreement: 1) how tasks are conceptualized and 2) how measurements of model performance are operationalized. To provide evidence for our taxonomy, we conduct a meta-analysis of relevant literature to understand how NLP tasks are conceptualized, as well as a survey of practitioners about their impressions of different factors that affect benchmark validity. Our meta-analysis and survey across eight tasks, ranging from coreference resolution to question answering, uncover that tasks are generally not clearly and consistently conceptualized and benchmarks suffer from operationalization disagreements. These findings support our proposed taxonomy of disagreement. Finally, based on our taxonomy, we present a framework for constructing benchmarks and documenting their limitations.
Code (0)
등록된 구현이 없습니다.
Tasks
coreference-resolutionCoreference ResolutionQuestion AnsweringSurveySimilar Papers 제목 키워드 기반
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requir…
Vision-Language NavigationIt Takes Two to Tango: Directly Optimizing for Constrained Synthesizability in Generative Molecular Design
Constrained synthesizability is an unaddressed challenge in generative molecular design. In particular, designing molecules satisfying multi-parameter optimization objectives, while simultaneously being synthesizable and…
Drug Discoveryreinforcement-learningReinforcement LearningIt Takes Two to Tango: Combining Visual and Textual Information for Detecting Duplicate Video-Based Bug Reports
When a bug manifests in a user-facing application, it is likely to be exposed through the graphical user interface (GUI). Given the importance of visual information to the process of identifying and understanding such bu…
Optical Character RecognitionOptical Character Recognition (OCR)RetrievalText RetrievalTangoBERT: Reducing Inference Cost by using Cascaded Architecture
The remarkable success of large transformer-based models such as BERT, RoBERTa and XLNet in many NLP tasks comes with a large increase in monetary and environmental cost due to their high computational load and energy co…
CPUReading ComprehensionSST-2text-classification+1