paper-with-me

Papers

AFT-VO: Asynchronous Fusion Transformers for Multi-View Visual Odometry Estimation

2022-06-26 · Nimet Kaygusuz, Oscar Mendez, Richard Bowden

Motion estimation approaches typically employ sensor fusion techniques, such as the Kalman Filter, to handle individual sensor failures. More recently, deep learning-based fusion approaches have been proposed, increasing the performance and requiring less model-specific implementations. However, current deep fusion approaches often assume that sensors are synchronised, which is not always practical, especially for low-cost hardware. To address this limitation, in this work, we propose AFT-VO, a novel transformer-based sensor fusion architecture to estimate VO from multiple sensors. Our framework combines predictions from asynchronous multi-view cameras and accounts for the time discrepancies of measurements coming from different sources. Our approach first employs a Mixture Density Network (MDN) to estimate the probability distributions of the 6-DoF poses for every camera in the system. Then a novel transformer-based fusion module, AFT-VO, is introduced, which combines these asynchronous pose estimations, along with their confidences. More specifically, we introduce Discretiser and Source Encoding techniques which enable the fusion of multi-source asynchronous signals. We evaluate our approach on the popular nuScenes and KITTI datasets. Our experiments demonstrate that multi-view fusion for VO estimation provides robust and accurate trajectories, outperforming the state of the art in both challenging weather and lighting conditions.

📄 PDF Abstract BibTeX arXiv:2206.12946

Code (0)

등록된 구현이 없습니다.

Tasks

Motion EstimationSensor FusionVisual Odometry

Similar Papers 제목 키워드 기반

Text-Guided Texturing by Synchronized Multi-View Diffusion

2023-11-21 · Yuxin Liu, Minshan Xie, Hanyuan Liu, Tien-Tsin Wong

This paper introduces a novel approach to synthesize texture to dress up a given 3D object, given a text prompt. Based on the pretrained text-to-image (T2I) diffusion model, existing methods usually employ a project-and-…

Texture Synthesis

Learning When to Denoise: Optimizing Asynchronous Schedules for Latent Diffusion

2026-06-18 · Bingshuo Qian, Xiang Cheng arxiv

Multi-representation diffusion models can improve visual synthesis by denoising complementary views of an image, but their performance depends critically on the asynchronous schedule that determines when each representat…

Long Context Tuning for Video Generation

2025-03-13 · Yuwei Guo, Ceyuan Yang, Ziyan Yang, Zhibei Ma 외

Recent advances in video generation can produce realistic, minute-long single-shot videos with scalable diffusion transformers. However, real-world narrative videos require multi-shot scenes with visual and dynamic consi…

Video Generation

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising

2026-04-29 · Jun Guo, Qiwei Li, Peiyan Li, Zilong Chen 외 arxiv

We propose X-WAM, a Unified 4D World Model that unifies real-time robotic action execution and high-fidelity 4D world synthesis (video + 3D reconstruction) in a single framework, addressing the critical limitations of pr…

3D Reconstruction

Dual Memory Neural Computer for Asynchronous Two-view Sequential Learning

2018-02-02 · Hung Le, Truyen Tran, Svetha Venkatesh

One of the core tasks in multi-view learning is to capture relations among views. For sequential data, the relations not only span across views, but also extend throughout the view length to form long-term intra-view and…

DecoderMULTI-VIEW LEARNINGVocal Bursts Valence Prediction