paper-with-me

홈 › Papers

Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight

2025-11-20 · Yi Yang, Xueqi Li, Yiyang Chen, Jin Song, Yihan Wang, Zipeng Xiao, Jiadi Su, You Qiaoben, Pengfei Liu, Zhijie Deng arxiv

Recent advances in Vision-Language-Action (VLA) models demonstrate that visual signals can effectively complement sparse action supervisions. However, letting VLA directly predict high-dimensional visual states can distribute model capacity and incur prohibitive training cost, while compressing visual states into more compact supervisory signals inevitably incurs information bottlenecks. Moreover, existing methods often suffer from poor comprehension and reasoning capabilities due to the neglect of language supervision. This paper introduces Mantis, a novel framework featuring a Disentangled Visual Foresight (DVF) to tackle these issues. Specifically, Mantis decouples visual foresight prediction from the backbone with the combination of meta queries and a diffusion Transformer (DiT) head. With the current visual state provided to the DiT via a residual connection, a simple next-state prediction objective enables the meta queries to automatically capture the latent actions that delineate the visual trajectory, and hence boost the learning of explicit actions. The disentanglement reduces the burden of the VLA backbone, enabling it to maintain comprehension and reasoning capabilities through language supervision. Empirically, pretrained on human manipulation videos, robot demonstrations, and image-text pairs, Mantis achieves a 96.7% success rate on LIBERO benchmark after fine-tuning, surpassing powerful baselines while exhibiting high convergence speed. Real-world evaluations show that Mantis outperforms $π_{0.5}$, a leading open-source VLA model, particularly in instruction-following capability, generalization to unseen instructions, and reasoning ability. Code and weights are released to support the open-source community.

📄 PDF Abstract BibTeX arXiv:2511.16175

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MANTIS: Interleaved Multi-Image Instruction Tuning

2024-05-02 · Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei 외

Large multimodal models (LMMs) have shown great results in single-image vision language tasks. However, their abilities to solve multi-image visual language tasks is yet to be improved. The existing LMMs like OpenFlaming…

MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation

2026-05-23 · Dahoon Park, Jahyun Koo, Sangwoo Hwang, Jaeha Kung arxiv

As the demand for deep learning grows, cost reduction through quantization has become essential for both training and inference. In 2022, the Open Compute Project (OCP) consortium standardized narrow precision formats fo…

Multimodal Conditionality for Natural Language Generation

2021-09-02 · Michael Sollami, Aashish Jain

Large scale pretrained language models have demonstrated state-of-the-art performance in language understanding tasks. Their application has recently expanded into multimodality learning, leading to improved representati…

DescriptiveLanguage ModelingLanguage ModellingText Generation

MantisV2: Closing the Zero-Shot Gap in Time Series Classification with Synthetic Data and Test-Time Strategies

2026-02-19 · Vasilii Feofanov, Songkang Wen, Jianfeng Zhang, Lujia Pan 외 arxiv

Developing foundation models for time series classification is of high practical relevance, as such models can serve as universal feature extractors for diverse downstream tasks. Although early models such as Mantis have…

Time Series ClassificationHuman Activity Recognition

Medical Large Vision Language Models with Multi-Image Visual Ability

2025-05-25 · Xikai Yang, Juzheng Miao, Yuchen Yuan, Jiaze Wang 외

Medical large vision-language models (LVLMs) have demonstrated promising performance across various single-image question answering (QA) benchmarks, yet their capability in processing multi-image clinical scenarios remai…

Question AnsweringVisual Question Answering (VQA)