paper-with-me

홈 › Papers

CodePercept: Code-Grounded Visual STEM Perception for MLLMs

2026-03-11 · Tongkun Guan, Zhibo Yang, Jianqiang Wan, Mingkun Yang, Zhengtao Guo, Zijian Hu, Ruilin Luo, Ruize Chen, Songtao Jiang, Peng Wang, Wei Shen, Junyang Lin, Xiaokang Yang arxiv

When MLLMs fail at Science, Technology, Engineering, and Mathematics (STEM) visual reasoning, a fundamental question arises: is it due to perceptual deficiencies or reasoning limitations? Through systematic scaling analysis that independently scales perception and reasoning components, we uncover a critical insight: scaling perception consistently outperforms scaling reasoning. This reveals perception as the true lever limiting current STEM visual reasoning. Motivated by this insight, our work focuses on systematically enhancing the perception capabilities of MLLMs by establishing code as a powerful perceptual medium--executable code provides precise semantics that naturally align with the structured nature of STEM visuals. Specifically, we construct ICC-1M, a large-scale dataset comprising 1M Image-Caption-Code triplets that materializes this code-as-perception paradigm through two complementary approaches: (1) Code-Grounded Caption Generation treats executable code as ground truth for image captions, eliminating the hallucinations inherent in existing knowledge distillation methods; (2) STEM Image-to-Code Translation prompts models to generate reconstruction code, mitigating the ambiguity of natural language for perception enhancement. To validate this paradigm, we further introduce STEM2Code-Eval, a novel benchmark that directly evaluates visual perception in STEM domains. Unlike existing work relying on problem-solving accuracy as a proxy that only measures problem-relevant understanding, our benchmark requires comprehensive visual comprehension through executable code generation for image reconstruction, providing deterministic and verifiable assessment. Code is available at https://github.com/TongkunGuan/Qwen-CodePercept.

📄 PDF Abstract BibTeX arXiv:2603.10757

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationImage ReconstructionCode TranslationVisual Reasoning

Similar Papers 제목 키워드 기반

VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding

2025-09-29 · Yizhuo Ding, Mingkang Chen, Zhibang Feng, Tong Xiao 외 arxiv

Multimodal large language models (MLLMs) often struggle to ground reasoning in perceptual evidence. We present a systematic study of perception strategies-explicit, implicit, visual, and textual-across four multimodal be…

Reinforcement LearningMultimodal Reasoning

Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models

2026-04-02 · Zekai Ye, Qiming Li, Xiaocheng Feng, Ruihan Chen 외 arxiv

While Reinforcement Learning from Verifiable Rewards (RLVR) has advanced reasoning in Large Vision-Language Models (LVLMs), prevailing frameworks suffer from a foundational methodological flaw: by distributing identical …

Reinforcement LearningMultimodal Reasoning

Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning

2025-12-19 · Siqi Yang, Zilve Gao, Haibo Qiu, Fanfan Liu 외 arxiv

Multimodal Large Language Models (MLLMs) demonstrate significant potential but remain brittle in complex, long-chain visual reasoning tasks. A critical failure mode is "visual forgetting", where models progressively lose…

Reinforcement LearningMultimodal ReasoningLogical ReasoningVisual Grounding

ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation

2026-08-20 · Linhan Cao, Siyuan Li, Jun Lan, Liangbo He 외 arxiv

Large multimodal models (LMMs) have demonstrated strong OCR recognition capabilities, yet remain vulnerable to adversarial visual text that is readable to humans but challenging for models to localize and recognize. Exis…

Visual Question Answering

Perceptual Reality Transformer: Neural Architectures for Simulating Neurological Perception Conditions

2025-08-13 · Baihan Lin arxiv

Neurological conditions affecting visual perception create profound experiential divides between affected individuals and their caregivers, families, and medical professionals. We present the Perceptual Reality Transform…