paper-with-me

홈 › Papers

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

2024-11-27 · Qing Jiang, Gen Luo, Yuqin Yang, Yuda Xiong, Yihao Chen, Zhaoyang Zeng, Tianhe Ren, Lei Zhang

Perception and understanding are two pillars of computer vision. While multimodal large language models (MLLM) have demonstrated remarkable visual understanding capabilities, they arguably lack accurate perception abilities, e.g. the stage-of-the-art model Qwen2-VL only achieves a 43.9 recall rate on the COCO dataset, limiting many tasks requiring the combination of perception and understanding. In this work, we aim to bridge this perception gap from both model designing and data development perspectives. We first introduce ChatRex, an MLLM with a decoupled perception design. Instead of having the LLM directly predict box coordinates, we feed the output boxes from a universal proposal network into the LLM, allowing it to output the corresponding box indices to represent its detection results, turning the regression task into a retrieval-based task that LLM handles more proficiently. From the data perspective, we build a fully automated data engine and construct the Rexverse-2M dataset which possesses multiple granularities to support the joint training of perception and understanding. After standard two-stage training, ChatRex demonstrates strong perception capabilities while preserving multimodal understanding performance. The combination of these two capabilities simultaneously unlocks many attractive applications, demonstrating the complementary roles of both perception and understanding in MLLM. Code is available at \url{https://github.com/IDEA-Research/ChatRex}.

📄 PDF Abstract BibTeX arXiv:2411.18363

Code (1)

idea-research/chatrex 공식 구현 pytorch

Similar Papers 제목 키워드 기반

TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models

2025-12-01 · Zhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou 외 arxiv

Unified multimodal models (UMMs) aim to jointly perform multimodal understanding and generation within a single framework. We present TUNA, a native UMM that builds a unified continuous visual representation by cascading…

Video GenerationImage Editing

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

2024-12-19 · CVPR 2025 1 · Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya 외

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (l…

Audio GenerationAudio SynthesisAudio-Visual SynchronizationVideo-to-Sound Generation

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning

2025-07-17 · Penglei Sun, Yaoxian Song, Xiangru Zhu, Xiang Liu 외

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have primarily focused on indoor household tasks, …

Question AnsweringScene Understanding

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

2026-06-17 · Yueyi Sun, Yuhao Wang, Jason Li, Ye Tian 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks…

Perceive What Matters: Relevance-Driven Scheduling for Multimodal Streaming Perception

2026-03-13 · Dingcheng Huang, Xiaotong Zhang, Kamal Youcef-Toumi arxiv

In modern human-robot collaboration (HRC) applications, multiple perception modules jointly extract visual, auditory, and contextual cues to achieve comprehensive scene understanding, enabling the robot to provide approp…

Scene Understanding