paper-with-me

홈 › Papers

ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

2026-02-26 · Yiran Guan, Sifan Tu, Dingkang Liang, Linghao Zhu, Jianzhong Ju, Zhenbo Luo, Jian Luan, Yuliang Liu, Xiang Bai arxiv

Omni-modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing omni-modal large language models (OLLM) excel at perceiving diverse modalities, they lack the complex reasoning abilities of recent large reasoning models (LRM). However, enhancing the reasoning ability of OLLMs through additional training presents significant challenges, including the need for high-quality data, task-specific adaptation, and substantial computational costs. To address these limitations, we propose ThinkOmni, a training-free and data-free framework that lifts textual reasoning to omni-modal scenarios. ThinkOmni introduces two key components: 1) LRM-as-a-Guide, which leverages off-the-shelf LRMs to guide the OLLM decoding process; 2) Stepwise Contrastive Scaling, which adaptively balances perception and reasoning signals without manual hyperparameter tuning. Experiments on six multi-modal reasoning benchmarks demonstrate that ThinkOmni consistently delivers performance improvements, with main results achieving 70.2 on MathVista and 75.5 on MMAU. Overall, ThinkOmni offers a flexible and generalizable solution for omni-modal reasoning and provides new insights into the generalization and application of reasoning capabilities. Code is publicly available at https://github.com/1ranGuan/thinkomni

📄 PDF Abstract BibTeX arXiv:2602.23306

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

2026-05-21 · Yifan Dai, Zhenhua Wu, Bohan Zeng, Daili Hua 외 arxiv

Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evidence from both modalities. A central lim…

Visual Reasoning

OmniBench: Towards The Future of Universal Omni-Language Models

2024-09-23 · Yizhi Li, Ge Zhang, Yinghao Ma, Ruibin Yuan 외

Recent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We in…

Instruction Following

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

2026-08-26 · Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of…

OmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D Scenes

2025-10-30 · Yukun Huang, Jiwen Yu, Yanning Zhou, Jianan Wang 외 arxiv

There are two prevalent ways to constructing 3D scenes: procedural generation and 2D lifting. Among them, panorama-based 2D lifting has emerged as a promising technique, leveraging powerful 2D generative priors to produc…

Scene Generation

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning

2025-05-28 · Shifang Zhao, Yiheng Lin, Lu Han, Yao Zhao 외

While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To address this gap, we introduce OmniAD, a novel framework that unifies anom…

Anomaly DetectionMultimodal ReasoningText GenerationVisual Reasoning