paper-with-me

Papers

CrystaL: Spontaneous Emergence of Visual Latents in MLLMs

2026-02-24 · Yang Zhang, Danyang Li, Yuxuan Li, Xin Zhang, Tianyu Xie, Mingming Cheng, Xiang Li arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable performance by integrating powerful language backbones with large-scale visual encoders. Among these, latent Chain-of-Thought (CoT) methods enable implicit reasoning in continuous hidden states, facilitating seamless vision-language integration and faster inference. However, existing heuristically predefined supervision signals in latent CoT provide limited guidance for preserving critical visual information in intermediate latent states. To address this limitation, we propose CrystaL (Crystallized Latent Reasoning), a single-stage framework with two paths to process intact and corrupted images, respectively. By explicitly aligning the attention patterns and prediction distributions across the two paths, CrystaL crystallizes latent representations into task-relevant visual semantics, without relying on auxiliary annotations or external modules. Extensive experiments on perception-intensive benchmarks demonstrate that CrystaL consistently outperforms state-of-the-art baselines, achieving substantial gains in fine-grained visual understanding while maintaining robust reasoning capabilities.

📄 PDF Abstract BibTeX arXiv:2602.20980

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs

2025-10-28 · Huanyu Zhang, Wenshan Wu, Chengzu Li, Ning Shang 외 arxiv

While Multimodal Large Language Models (MLLMs) excel at visual understanding, they often struggle in complex scenarios that require visual planning and imagination. Inspired by how humans use sketching as a form of visua…

Multimodal Reasoning

Bridging Compressed Image Latents and Multimodal Large Language Models

2024-07-29 · Chia-Hao Kao, Cheng Chien, Yu-Jen Tseng, Yi-Hsin Chen 외

This paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large lan…

Image Compression

Look-Back: Implicit Visual Re-focusing in MLLM Reasoning

2025-07-02 · Shuo Yang, Yuwei Niu, Yuyang Liu, Yang Ye 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in multimodal reasoning. However, they often excessively rely on textual information during the later stages of inference, neglecting the crucial…

Multimodal Reasoning

Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs

2025-12-18 · Jintao Tong, Jiaqi Gu, Yujing Lou, Lubin Fan 외 arxiv

While Multimodal Large Language Models (MLLMs) excel at visual understanding tasks through text reasoning, they often fall short in scenarios requiring visual imagination. Unlike current works that take predefined extern…

MEGC2026: Micro-Expression Grand Challenge on Visual Question Answering

2026-03-09 · Xinqi Fan, Jingting Li, John See, Moi Hoon Yap 외 arxiv

Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stak…

Visual Question AnsweringVideo Question AnsweringMultimodal Reasoning