paper-with-me

Papers

LanteRn: Latent Visual Structured Reasoning

2026-03-26 · André G. Viveiros, Nuno Gonçalves, Matthias Lindemann, André Martins arxiv

While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs). As a result, most LMMs default to verbalizing perceptual content into text, a strong limitation for tasks requiring fine-grained spatial and visual understanding. While recent approaches take steps toward thinking with images by invoking tools or generating intermediate images, they either rely on external modules, or incur unnecessary computation by reasoning directly in pixel space. In this paper, we introduce LanteRn, a framework that enables LMMs to interleave language with compact latent visual representations, allowing visual reasoning to occur directly in latent space. LanteRn augments a vision-language transformer with the ability to generate and attend to continuous visual thought embeddings during inference. We train the model in two stages: supervised fine-tuning to ground visual features in latent states, followed by reinforcement learning to align latent reasoning with task-level utility. We evaluate LanteRn on three perception-centric benchmarks (VisCoT, V*, and Blink), observing consistent improvements in visual grounding and fine-grained reasoning. These results suggest that internal latent representations provide a promising direction for more efficient multimodal reasoning.

📄 PDF Abstract BibTeX arXiv:2603.25629

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMultimodal ReasoningVisual GroundingVisual Reasoning

Similar Papers 제목 키워드 기반

LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding

2024-10-04 · Doohyuk Jang, Sihwan Park, June Yong Yang, Yeonsung Jung 외

Auto-Regressive (AR) models have recently gained prominence in image generation, often matching or even surpassing the performance of diffusion models. However, one major limitation of AR models is their sequential natur…

Image Generation

LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models

2025-02-10 · Sihwan Park, Doohyuk Jang, Sungyub Kim, Souvik Kundu 외

Speculative decoding has been widely used to accelerate auto-regressive (AR) text generation. However, its effectiveness for visual AR models remains limited due to token selection ambiguity, where multiple tokens share …

Text Generation

Temperature sensitivity of pest reproductive numbers in age-structured PDE models, with a focus on the invasive spotted lanternfly

2021-12-21 · Stephanie M. Lewkiewicz, Sebastiano De Bona, Matthew R. Helmus, Benjamin Seibold

Invasive pest establishment is a pervasive threat to global ecosystems, agriculture, and public health. The recent establishment of the invasive spotted lanternfly in the northeastern United States has proven devastating…

Sensitivity

LANTERN: Scalable Distillation of Large Language Models for Job-Person Fit and Explanation

2025-10-07 · Zhoutong Fu, Yihan Cao, Yi-Lin Chen, Aman Lunia 외 arxiv

Large language models (LLMs) have achieved strong performance across a wide range of natural language processing tasks. However, deploying LLMs at scale for domain specific applications, such as job-person fit and explan…

Knowledge DistillationPrompt Engineering

Lantern: A Minimalist Robotic Object Platform

2026-01-29 · Victor Nikhil Antony, Zhili Gong, Guanchen Li, Clara Jeon 외 arxiv

Robotic objects are simple actuated systems that subtly blend into human environments. We design and introduce Lantern, a minimalist robotic object platform to enable building simple robotic artifacts. We conducted in-de…