paper-with-me

홈 › Papers

Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning

2025-09-17 · Haodong Zhao, Chenyan Zhao, Yansi Li, Zhuosheng Zhang, Gongshen Liu arxiv

The capacity of Large Language Models (LLMs) to reason is fundamental to their application in complex, knowledge-intensive domains. In real-world scenarios, LLMs are often augmented with external information that can be helpful, irrelevant, or even misleading. This paper investigates the causal impact of such auxiliary information on the reasoning process of LLMs with explicit step-by-step thinking capabilities. We introduce SciAux, a new dataset derived from ScienceQA, to systematically test the robustness of the model against these types of information. Our findings reveal a critical vulnerability: the model's deliberative "thinking mode" is a double-edged sword. While helpful context improves accuracy, misleading information causes a catastrophic drop in performance, which is amplified by the thinking process. Instead of conferring robustness, thinking reinforces the degree of error when provided with misinformation. This highlights that the challenge is not merely to make models "think", but to endow them with the critical faculty to evaluate the information upon which their reasoning is based. The SciAux dataset is available at https://huggingface.co/datasets/billhdzhao/SciAux.

📄 PDF Abstract BibTeX arXiv:2509.18163

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reconstructing Groups of People with Hypergraph Relational Reasoning

2023-08-30 · ICCV 2023 1 · Buzhen Huang, Jingyi Ju, Zhihao LI, Yangang Wang

Due to the mutual occlusion, severe scale variation, and complex spatial distribution, the current multi-person mesh recovery methods cannot produce accurate absolute body poses and shapes in large-scale crowded scenes. …

3D Multi-Person Mesh RecoveryPose EstimationRelational Reasoning

Not Too Short, Not Too Long: How LLM Response Length Shapes People's Critical Thinking in Error Detection

2026-03-06 · Natalie Friedman, Adelaide Nyanyo, Kevin Weatherwax, Lifei Wang 외 arxiv

Large language models (LLMs) have become common decision-support tools across educational and professional contexts, raising questions about how their outputs shape human critical thinking. Prior work suggests that the a…

Rethinking On-Policy Self-Distillation for Thinking Models

2026-07-06 · Simran Kaur, Narutatsu Ri, Yinghui He, Liam Fowl 외 arxiv

Self-distillation is a promising recipe for self-improvement in language models. In this setting, a model can serve as its own teacher when given privileged information, such as a solution to a math problem. This seems e…

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs

2026-05-12 · Houcheng Jiang, Jiajun Fu, Junfeng Fang, Chen Gao 외 arxiv

Multimodal large language models are increasingly expected to perform thinking with images, yet existing visual latent reasoning methods still rely on explicit textual chain-of-thought interleaved with visual latent toke…

Visual Reasoning

SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression

2025-08-18 · Yuyang Xu, Yi Cheng, Haochao Ying, Zhuoyun Du 외 arxiv

Test-time scaling has proven effective in further enhancing the performance of pretrained Large Language Models (LLMs). However, mainstream post-training methods (i.e., reinforcement learning (RL) with chain-of-thought (…

Reinforcement Learning