paper-with-me

Papers

XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition

2026-05-14 · Gong Zhiren, Tiantong Wu, Jiaming Zhang, Fuyao Zhang, Che Wang, Yurong Hao, Yikun Hou, Foo Ping, Yilei Zhao, Fei Huang, Chau Yuen, Wei Yang Bryan Lim arxiv

Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on single-turn restricted scenarios, failing to capture the capability boundaries exposed by real-world interactive scientific workflows. To address this, we introduce XDomainBench, a diagnostic benchmark for interactive interdisciplinary scientific reasoning. We formalize the composition order and mixture structure to enable systematic stress-testing from single-discipline to inter-disciplinary, comprising 8,598 interactive sessions across 20 domains and 4 task categories, with 8 realistic trajectory patterns covering difficulty and domain-mixture dynamics, simulating real AI4S scenarios. Large-scale evaluation of LLMs reveals a systematic reasoning collapse as composition order increases, stemming from two root causes: (i) direct difficulty increases induced by domain composition, and (ii) indirect interaction-amplified failures where trajectory patterns trigger error accumulation, reasoning breaks, and domain confusion, ultimately leading to session collapse.

📄 PDF Abstract BibTeX arXiv:2605.14754

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

2026-07-12 · Keqin Peng, Chen Li, Yuanxin Ouyang, Yancheng Yuan 외 arxiv

On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxically degrades downstream performance. In th…

Diagnosing CFG Interpretation in LLMs

2026-04-22 · Hanqi Li, Lu Chen, Kai Yu arxiv

As LLMs are increasingly integrated into agentic systems, they must adhere to dynamically defined, machine-interpretable interfaces. We evaluate LLMs as in-context interpreters: given a novel context-free grammar, can LL…

Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents

2026-05-29 · Yunpeng Zhou arxiv

Modular visual reasoning systems increasingly rely on shared working memory for multi-step collaboration, yet the failure dynamics of intermediate state evolution in low-capacity regimes remain underexplored. We study fa…

Visual Question AnsweringVisual Reasoning

Diagnosing Korean-Language LLM Political Bias via Census-Grounded Agent Simulation

2026-05-18 · Sungwoo Kang arxiv

Large language models (LLMs) exhibit systematic political biases in voter simulations, but their underlying mechanisms and cross-lingual generalizations remain poorly understood. We introduce Dynamo-K, a census-grounded …

Intention Collapse: Intention-Level Metrics for Reasoning in Language Models

2026-01-03 · Patricio Vera arxiv

Language generation maps a rich, high-dimensional internal state to a single token sequence. We study this many-to-one mapping through the lens of intention collapse: the projection from an internal intention space I to …