paper-with-me

홈 › Papers

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

2026-07-22 · Anmol Kankariya, Sercan Ö. Arık arxiv

While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when models encounter novel abstractions or rigorous domain constraints. We introduce PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents: (1) Adversarial Refinement Agent, (2) Hierarchical strategic Planning Agent, (3) Spectrum Search Agent, and (4) Direct Chain Agent. A final Task-Adaptive Aggregation Layer dynamically reconciles these perspectives -- via final candidate selection, semantic synthesis, or neuro-symbolic verification -- to produce a robust global solution. We evaluate PoTRE on three frontier benchmarks: ARC-AGI-2, Humanity's Last Exam (HLE), and PRBench Finance. PoTRE achieves state-of-the-art accuracy of 49.92% on HLE, surpassing the previous best official score. We demonstrate that this architectural heterogeneity achieves improved reasoning performance using similar or fewer inference tokens compared to heavily scaled homogeneous baselines.

📄 PDF Abstract BibTeX arXiv:2607.20268

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning

2025-11-28 · Yang Li, Zhiyuan He, Yuxuan Huang, Zhuhanling Xiao 외 arxiv

Recent Vision-Language Models (VLMs) exhibit strong perceptual reasoning abilities, yet they often struggle to adapt efficiently when encountering novel tasks at test time. In contrast, humans leverage the metacognitive …

Reinforcement LearningTest-time AdaptationAtari Games

A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models

2024-02-28 · Xiujie Song, Mengyue Wu, Kenny Q. Zhu, Chunhao Zhang 외

Large Vision-Language Models (LVLMs), despite their recent success, are hardly comprehensively tested for their cognitive abilities. Inspired by the prevalent use of the "Cookie Theft" task in human cognition test, we pr…

Image DescriptionQuestion AnsweringVisual Question Answering

Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models

2025-12-17 · Jinwu Hu, Dongjin Yang, Langyu Bian, Zhiquan Wen 외 arxiv

Large language models (LLMs) have demonstrated impressive performance across various language tasks. However, existing LLM reasoning strategies mainly rely on the LLM itself with fast or slow mode (like o1 thinking) and …

Reinforcement Learning

Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching

2025-03-07 · Simon A. Aytes, Jinheon Baek, Sung Ju Hwang

Recent advances in large language models (LLMs) have enabled strong reasoning capabilities through Chain-of-Thought (CoT) prompting, which elicits step-by-step problem solving, but often at the cost of excessive verbosit…

TangramSR: Can Vision-Language Models Reason in Continuous Geometric Space?

2026-02-05 · Yikun Zong, Cheston Tan arxiv

Humans excel at spatial reasoning tasks like Tangram puzzle assembly through cognitive processes involving mental rotation, iterative refinement, and visual feedback. Inspired by how humans solve Tangram puzzles through …

Spatial Reasoning