paper-with-me

홈 › Papers

The ARC of Progress towards AGI: A Living Survey of Abstraction and Reasoning

2026-03-09 · Sahar Vahdati, Andrei Aioanei, Haridhra Suresh, Jens Lehmann arxiv

The Abstraction and Reasoning Corpus (ARC-AGI) has become a key benchmark for fluid intelligence in AI. This survey presents the first cross-generation analysis of 82 approaches across three benchmark versions and the ARC Prize 2024-2025 competitions. Our central finding is that performance degradation across versions is consistent across all paradigms: program synthesis, neuro-symbolic, and neural approaches all exhibit 2-3x drops from ARC-AGI-1 to ARC-AGI-2, indicating fundamental limitations in compositional generalization. While systems now reach 93.0% on ARC-AGI-1 (Opus 4.6), performance falls to 68.8% on ARC-AGI-2 and 13% on ARC-AGI-3, as humans maintain near-perfect accuracy across all versions. Cost fell 390x in one year (o3's $4,500/task to GPT-5.2's $12/task), although this largely reflects reduced test-time parallelism. Trillion-scale models vary widely in score and cost, while Kaggle-constrained entries (660M-8B) achieve competitive results, aligning with Chollet's thesis that intelligence is skill-acquisition efficiency. Test-time adaptation and refinement loops emerge as critical success factors, while compositional reasoning and interactive learning remain unsolved. ARC Prize 2025 winners needed hundreds of thousands of synthetic examples to reach 24% on ARC-AGI-2, confirming that reasoning remains knowledge-bound. This first release of the ARC-AGI Living Survey captures the field as of February 2026, with updates at https://nimi-ai.com/arc-survey/

📄 PDF Abstract BibTeX arXiv:2603.13372

Code (0)

등록된 구현이 없습니다.

Tasks

Test-time AdaptationProgram Synthesis

Similar Papers 제목 키워드 기반

Neural-guided, Bidirectional Program Search for Abstraction and Reasoning

2021-10-22 · Simon Alford, Anshula Gandhi, Akshay Rangamani, Andrzej Banburski 외

One of the challenges facing artificial intelligence research today is designing systems capable of utilizing systematic reasoning to generalize to new tasks. The Abstraction and Reasoning Corpus (ARC) measures such a ca…

ARCProgram SynthesisVisual Reasoning

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes

2026-06-09 · Avinash Anand, Mahisha Ramesh, Avni Mittal, Ashutosh Kumar 외 arxiv

Large Language Models (LLMs) have achieved strong performance across natural language processing tasks, yet reliable reasoning remains an open challenge. Although modern LLMs show progress in structured inference, multi-…

Mathematical ReasoningCommon Sense ReasoningReinforcement LearningDomain Generalization

Reasoning or Pattern Matching? Probing Large Vision-Language Models with Visual Puzzles

2026-01-20 · Maria Lymperaiou, Vasileios Karampinis, Giorgos Filandrianos, Angelos Vlachos 외 arxiv

Puzzles have long served as compact and revealing probes of human cognition, isolating abstraction, rule discovery, and systematic reasoning with minimal reliance on prior knowledge. Leveraging these properties, visual p…

When LLM Meets Tree Search: A Systematic View of Inference as Search in Large Language Models

2026-08-31 · Jiaqi Wei, Xiang Zhang, Yuejin Yang, Wenxuan Huang 외 arxiv

As pretraining scaling laws approach saturation, Test-Time Scaling (TTS) has emerged as an important direction for improving reasoning by allocating inference-time compute to a fixed model prior. Viewed at a high level, …

Do AI Models Perform Human-like Abstract Reasoning Across Modalities?

2025-10-02 · Claas Beger, Ryan Yi, Shuhao Fu, Kaleda Denton 외 arxiv

OpenAI's o3-preview reasoning model exceeded human accuracy on the ARC-AGI-1 benchmark, but does that mean state-of-the-art models recognize and reason with the abstractions the benchmark was designed to test? Here we in…