paper-with-me

Papers

Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning

2026-05-26 · Zhe Yu, Wenpeng Xing, Yunzhao Wei, Jie Chen, Hongzhi Wang, Xuyang Teng, Meng Han arxiv

Post-training is routinely evaluated through aggregate benchmark scores that treat multi-hop reasoning as a single capability -- as if a model that answers more questions correctly must be better at assembling facts. We show that this assumption can be misleading: recipes with statistically indistinguishable atomic knowledge produce composition behaviour separated by over 40 percentage points, a phenomenon we call composition collapse: the systematic failure to assemble stably-known facts into chains, invisible to aggregate metrics. We introduce a double-gate protocol that changes the estimand from an aggregate compositionality gap to residual composition failure conditioned on stable atomic access, decomposing post-training gains into three independent channels: atomic stability, residual composition, and critical depth. On a benchmark of temporal factual chains spanning depths 2--11 across four post-training recipes, this decomposition reveals that post-training objectives shift composition capability in directions that aggregate metrics mask, and suggests that claims about multi-hop reasoning improvement should be accompanied by atomic-gate-controlled composition metrics. Diagnostic probes further show that a substantial share of measured composition failure reflects generation-time computation constraints rather than permanent inability to compose.

📄 PDF Abstract BibTeX arXiv:2605.26789

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"

2026-02-04 · Dhruv Madhwal, Lyuxin David Zhang, Dan Roth, Tomer Wolfson 외 arxiv

Large language models often struggle to recognize their knowledge limits in closed-book question answering, leading to confident hallucinations. While decomposed prompting is typically used to improve accuracy, we invest…

Question Answering

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

2026-08-06 · Xi Zeng, Haojie Ren, Ziying Song arxiv

We propose PhyLatent, a dynamics-relevant training objective for JointEmbedding Predictive Architecture (JEPA) world models. Our key observation is that preventing global latent collapse does not ensure that a representa…

COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking

2026-05-14 · Shukun Jia, Shiyu Hu, Yipei Wang, Ximeng Cheng 외 arxiv

Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between the high-discriminability demand and the sparse semantic supervision. This mismatch is particularly acute in highly homogeneous …

Multi-Object Tracking

The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence

2026-06-19 · David Nordfors arxiv

The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs. No content is given in advance; the contestants create all of it -- a new kind …

Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training

2025-09-05 · Figarri Keisha, Zekun Wu, Ze Wang, Adriano Koshiyama 외 arxiv

Large language models increasingly rely on synthetic data due to human-written content scarcity, yet recursive training on model-generated outputs leads to model collapse, a degenerative process threatening factual relia…

Computational Efficiency