paper-with-me

홈 › Papers

Are Emergent Abilities of Large Language Models a Mirage?

2023-04-28 · NeurIPS 2023 11 · Rylan Schaeffer, Brando Miranda, Sanmi Koyejo

Recent work claims that large language models display emergent abilities, abilities not present in smaller-scale models that are present in larger-scale models. What makes emergent abilities intriguing is two-fold: their sharpness, transitioning seemingly instantaneously from not present to present, and their unpredictability, appearing at seemingly unforeseeable model scales. Here, we present an alternative explanation for emergent abilities: that for a particular task and model family, when analyzing fixed model outputs, emergent abilities appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale. Specifically, nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous predictable changes in model performance. We present our alternative explanation in a simple mathematical model, then test it in three complementary ways: we (1) make, test and confirm three predictions on the effect of metric choice using the InstructGPT/GPT-3 family on tasks with claimed emergent abilities; (2) make, test and confirm two predictions about metric choices in a meta-analysis of emergent abilities on BIG-Bench; and (3) show to choose metrics to produce never-before-seen seemingly emergent abilities in multiple vision tasks across diverse deep networks. Via all three analyses, we provide evidence that alleged emergent abilities evaporate with different metrics or with better statistics, and may not be a fundamental property of scaling AI models.

📄 PDF Abstract BibTeX arXiv:2304.15004

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Step Reasoning in Korean and the Emergent Mirage

2025-01-10 · Guijin Son, Hyunwoo Ko, Dasol Choi

We introduce HRMCR (HAE-RAE Multi-Step Commonsense Reasoning), a benchmark designed to evaluate large language models' ability to perform multi-step reasoning in culturally specific contexts, focusing on Korean. The ques…

MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments

2025-01-03 · Yin Cai, Zhouhong Gu, Zhaohan Du, Zheyu Ye 외

Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly in interactive role-playing contexts. Thi…

Decision Making

Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables

2025-11-21 · Anshul Singh, Rohan Chaudhary, Gagneet Singh, Abhay Kumary arxiv

The impressive performance of VLMs is largely measured on benchmarks that fail to capture the complexities of real-world scenarios. Existing datasets for tabular QA, such as WikiTableQuestions and FinQA, are overwhelming…

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?

2026-07-10 · Abhinav Rao, Liancheng Gong, Bin Hu, Atharva Naik arxiv

Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, alongside evidence that this behavior can…

MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence

2025-05-15 · Chonghan Liu, Haoran Wang, Felix Henry, Pu Miao 외

Spatial perception and reasoning are core components of human cognition, encompassing object recognition, spatial relational understanding, and dynamic reasoning. Despite progress in computer vision, existing benchmarks …

AttributeObjectObject RecognitionRelation+1