paper-with-me

Papers

OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer

2025-11-13 · Haosong Peng, Hao Li, Yalun Dai, Yushi Lan, Yihang Luo, Tianyu Qi, Zhengshen Zhang, Yufeng Zhan, Junfei Zhang, Wenchao Xu, Ziwei Liu arxiv

General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., camera intrinsics, poses, and depth maps). To address this issue, we introduce OmniVGGT, a novel framework that can effectively benefit from an arbitrary number of auxiliary geometric modalities during both training and inference. In our framework, a GeoAdapter is proposed to encode depth and camera intrinsics/extrinsics into a spatial foundation model. It employs zero-initialized convolutions to progressively inject geometric information without disrupting the foundation model's representation space. This design ensures stable optimization with negligible overhead, maintaining inference speed comparable to VGGT even with multiple additional inputs. Additionally, a stochastic multimodal fusion regimen is proposed, which randomly samples modality subsets per instance during training. This enables an arbitrary number of modality inputs during testing and promotes learning robust spatial representations instead of overfitting to auxiliary cues. Comprehensive experiments on monocular/multi-view depth estimation, multi-view stereo, and camera pose estimation demonstrate that OmniVGGT outperforms prior methods with auxiliary inputs and achieves state-of-the-art results even with RGB-only input. To further highlight its practical utility, we integrated OmniVGGT into vision-language-action (VLA) models. The enhanced VLA model by OmniVGGT not only outperforms the vanilla point-cloud-based baseline on mainstream benchmarks, but also effectively leverages accessible auxiliary inputs to achieve consistent gains on robotic tasks.

📄 PDF Abstract BibTeX arXiv:2511.10560

Code (0)

등록된 구현이 없습니다.

Tasks

Camera Pose EstimationDepth Estimation

Similar Papers 제목 키워드 기반

OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models

2026-07-03 · Shijie Cao, Qingyu Zhang, Boxi Yu, Yuzhong Zhang 외 arxiv

Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they generate large token sequences under audio-visual inputs, leading to substantial in…

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering

2026-04-09 · Yiduo Jia, Muzhi Zhu, Hao Zhong, Mingyu Liu 외 arxiv

To extend the reinforcement learning post-training paradigm to omni-modal models for concurrently bolstering video-audio understanding and collaborative reasoning, we propose OmniJigsaw, a generic self-supervised framewo…

Reinforcement Learning

OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

2026-05-02 · Detao Bai, Shimin Yao, Weixuan Chen, Chengen Lai 외 arxiv

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, au…

Sign Language RecognitionComputational EfficiencySpeaker Identification

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

2026-05-09 · Tao Yu, yiming ding, Shenghua Chai, Minghui Zhang 외 arxiv

Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively search for cross-modal evidence remains …

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

2026-02-05 · Zhangquan Chen, Jiale Tao, Ruihuang Li, Yihao Hu 외 arxiv

While humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings, existing omnivideo models still face substantial challenges on audio-visu…

Self-Supervised LearningContrastive LearningVisual Reasoning