paper-with-me

Papers

CoDA21: Evaluating Language Understanding Capabilities of NLP Models With Context-Definition Alignment

2021-09-17 · ACL ARR September 2021 9 · Anonymous

Pretrained language models (PLMs) have achieved superhuman performance on many benchmarks, creating a need for harder tasks. We introduce CoDA21 (Context Definition Alignment), a challenging benchmark that measures natural language understanding (NLU) capabilities of PLMs: Given a definition and a context each for k words, but not the words themselves, the task is to align the k definitions with the k contexts. CoDA21 requires a deep understanding of contexts and definitions, including complex inference and world knowledge. We find that there is a large gap between human and PLM performance, suggesting that CoDA21 measures an aspect of NLU that is not sufficiently covered in existing benchmarks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingWorld Knowledge

Similar Papers 제목 키워드 기반

CoDA21: Evaluating Language Understanding Capabilities of NLP Models With Context-Definition Alignment

2022-03-11 · ACL 2022 5 · Lütfi Kerem Senel, Timo Schick, Hinrich Schütze

Pretrained language models (PLMs) have achieved superhuman performance on many benchmarks, creating a need for harder tasks. We introduce CoDA21 (Context Definition Alignment), a challenging benchmark that measures natur…

Natural Language UnderstandingWorld Knowledge

CoDA: Coding LM via Diffusion Adaptation

2025-09-27 · Haolin Chen, Shiyu Wang, Can Qin, Bo Pang 외 arxiv

Diffusion language models promise bidirectional context and infilling capabilities that autoregressive coders lack, yet practical systems remain heavyweight. We introduce CoDA, a 1.7B-parameter diffusion coder trained on…

Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

2024-04-16 · Kai Chen, Yanze Li, Wenhua Zhang, Yanxin Liu 외

Large Vision-Language Models (LVLMs) have received widespread attention for advancing the interpretable self-driving. Existing evaluations of LVLMs primarily focus on multi-faceted capabilities in natural circumstances, …

Autonomous DrivingVisual Reasoning

SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning

2025-02-19 · Renxi Wang, Honglin Mu, Liqun Ma, Lizhi Lin 외

Evaluating large language models' (LLMs) long-context understanding capabilities remains challenging. We present SCALAR (Scientific Citation-based Live Assessment of Long-context Academic Reasoning), a novel benchmark th…

Long-Context Understanding

Retrieval or Global Context Understanding? On Many-Shot In-Context Learning for Long-Context Evaluation

2024-11-11 · Kaijian Zou, Muhammad Khalifa, Lu Wang

Language models (LMs) have demonstrated an improved capacity to handle long-context information, yet existing long-context benchmarks primarily measure LMs' retrieval abilities with extended inputs, e.g., pinpointing a s…

16kBenchmarkingIn-Context LearningRetrieval