paper-with-me

Papers

Understanding Multi-View Transformers

2025-10-28 · Michal Stary, Julien Gaubil, Ayush Tewari, Vincent Sitzmann arxiv

Multi-view transformers such as DUSt3R are revolutionizing 3D vision by solving 3D tasks in a feed-forward manner. However, contrary to previous optimization-based pipelines, the inner mechanisms of multi-view transformers are unclear. Their black-box nature makes further improvements beyond data scaling challenging and complicates usage in safety- and reliability-critical applications. Here, we present an approach for probing and visualizing 3D representations from the residual connections of the multi-view transformers' layers. In this manner, we investigate a variant of the DUSt3R model, shedding light on the development of its latent state across blocks, the role of the individual layers, and suggest how it differs from methods with stronger inductive biases of explicit global pose. Finally, we show that the investigated variant of DUSt3R estimates correspondences that are refined with reconstructed geometry. The code used for the analysis is available at https://github.com/JulienGaubil/und3rstand .

📄 PDF Abstract BibTeX arXiv:2510.24907

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D Space

2022-06-23 · Jinghuan Shang, Srijan Das, Michael S. Ryoo

Humans are remarkably flexible in understanding viewpoint changes due to visual cortex supporting the perception of 3D structure. In contrast, most of the computer vision models that learn visual representation from a po…

Action Recognitionimage-classificationImage ClassificationVideo Alignment

Multiview Transformers for Video Recognition

2022-01-12 · CVPR 2022 1 · Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu 외

Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations. Although transformer architectures have recently advanced the…

Action ClassificationAction RecognitionVideo Understanding

Transformers and Language Models in Form Understanding: A Comprehensive Review of Scanned Document Analysis

2024-03-06 · Abdelrahman Abdallah, Daniel Eberharter, Zoe Pfister, Adam Jatowt

This paper presents a comprehensive survey of research works on the topic of form understanding in the context of scanned documents. We delve into recent advancements and breakthroughs in the field, highlighting the sign…

Form

Optimizing Multi-Class Text Classification: A Diverse Stacking Ensemble Framework Utilizing Transformers

2023-08-19 · Anusuya Krishnan

Customer reviews play a crucial role in assessing customer satisfaction, gathering feedback, and driving improvements for businesses. Analyzing these reviews provides valuable insights into customer sentiments, including…

ClassificationMulti Class Text Classificationtext-classificationText Classification

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

2022-05-29 · Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu 외

Large-scale pretrained transformers have created milestones in text (GPT-3) and text-to-image (DALL-E and CogView) generation. Its application to video generation is still facing many challenges: The potential huge compu…

Text-to-Video GenerationVideo Generation