paper-with-me

홈 › Papers

Unlocking Cognitive Capabilities and Analyzing the Perception-Logic Trade-off

2026-02-27 · Longyin Zhang, Shuo Sun, Yingxu He, Won Cheng Yi Lewis, Muhammad Huzaifah Bin Md Shahrin, Hardik Bhupendra Sailor, Heng Meng Jeremy Wong, Tarun Kumar Vangani, Yi Ma, Qiongqiong Wang, Minh Duc Pham, Ridong Jiang, Jingtao Li, Jingyi Liao, Zhuohan Liu, Yanfeng Lu, Manas Gupta, Ai Ti Aw arxiv

Recent advancements in Multimodal Large Language Models (MLLMs) pursue omni-perception capabilities, yet integrating robust sensory grounding with complex reasoning remains a challenge, particularly for underrepresented regions. In this report, we introduce the research preview of MERaLiON2-Omni (Alpha), a 10B-parameter multilingual omni-perception tailored for Southeast Asia (SEA). We present a progressive training pipeline that explicitly decouples and then integrates "System 1" (Perception) and "System 2" (Reasoning) capabilities. First, we establish a robust Perception Backbone by aligning region-specific audio-visual cues (e.g., Singlish code-switching, local cultural landmarks) with a multilingual LLM through orthogonal modality adaptation. Second, to inject cognitive capabilities without large-scale supervision, we propose a cost-effective Generate-Judge-Refine pipeline. By utilizing a Super-LLM to filter hallucinations and resolve conflicts via a consensus mechanism, we synthesize high-quality silver data that transfers textual Chain-of-Thought reasoning to multimodal scenarios. Comprehensive evaluation on our newly introduced SEA-Omni Benchmark Suite reveals an Efficiency-Stability Paradox: while reasoning acts as a non-linear amplifier for abstract tasks (boosting mathematical and instruction-following performance significantly), it introduces instability in low-level sensory processing. Specifically, we identify Temporal Drift in long-context audio, where extended reasoning desynchronizes the model from acoustic timestamps, and Visual Over-interpretation, where logic overrides pixel-level reality. This report details the architecture, the data-efficient training recipe, and a diagnostic analysis of the trade-offs between robust perception and structured reasoning.

📄 PDF Abstract BibTeX arXiv:2602.23730

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation

2025-04-09 · Luo Ling, Bai Qianqian

Enhancing the spatial perception capabilities of mobile robots is crucial for achieving embodied Vision-and-Language Navigation (VLN). Although significant progress has been made in simulated environments, directly trans…

HallucinationSpatial ReasoningVision and Language Navigation

ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge Base

2023-05-10 · Siyu Yuan, Jiangjie Chen, Changzhi Sun, Jiaqing Liang 외

Analogical reasoning is a fundamental cognitive ability of humans. However, current language models (LMs) still struggle to achieve human-like performance in analogical reasoning tasks due to a lack of resources for mode…

Knowledge Graphs

Improving cognitive diagnostics in pathology: a deep learning approach for augmenting perceptional understanding of histopathology images

2025-03-10 · Xiaoqian Hu

In Recent Years, Digital Technologies Have Made Significant Strides In Augmenting-Human-Health, Cognition, And Perception, Particularly Within The Field Of Computational-Pathology. This Paper Presents A Novel Approach To…

DiagnosticImage CaptioningMedical Image Analysis

Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning

2025-11-27 · Zhaoyang Wei, Wenchao Ding, Yanchao Hao, Xi Chen arxiv

Models capable of "thinking with images" by dynamically grounding their reasoning in visual evidence represent a major leap in multimodal AI. However, replicating and advancing this ability is non-trivial, with current m…

Reinforcement LearningVisual Reasoning

Perception Graph for Cognitive Attack Reasoning in Augmented Reality

2025-08-30 · Rongqian Chen, Shu Hong, Rifatul Islam, Mahdi Imani 외 arxiv

Augmented reality (AR) systems are increasingly deployed in tactical environments, but their reliance on seamless human-computer interaction makes them vulnerable to cognitive attacks that manipulate a user's perception …