paper-with-me

Papers

Seeing without Pixels: Perception from Camera Trajectories

2025-11-26 · Zihui Xue, Kristen Grauman, Dima Damen, Andrew Zisserman, Tengda Han arxiv

Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigate this seemingly implausible question. Towards this end, we propose a contrastive learning framework to train CamFormer, a dedicated encoder that projects camera pose trajectories into a joint embedding space, aligning them with natural language. We find that, contrary to its apparent simplicity, the camera trajectory is a remarkably informative signal to uncover video content. In other words, "how you move" can indeed provide valuable cues about "what you are doing" (egocentric) or "observing" (exocentric). We demonstrate the versatility of our learned CamFormer embeddings on a diverse suite of downstream tasks, ranging from cross-modal alignment to classification and temporal analysis. Importantly, our representations are robust across diverse camera pose estimation methods, including both high-fidelity multi-sensored and standard RGB-only estimators. Our findings establish camera trajectory as a lightweight, robust, and versatile modality for perceiving video content.

📄 PDF Abstract BibTeX arXiv:2511.21681

Code (0)

등록된 구현이 없습니다.

Tasks

Camera Pose EstimationContrastive Learning

Similar Papers 제목 키워드 기반

Seeing Objects in a Cluttered World: Computational Objectness from Motion in Video

2024-02-02 · Douglas Poland, Amar Saini

Perception of the visually disjoint surfaces of our cluttered world as whole objects, physically distinct from those overlapping them, is a cognitive phenomenon called objectness that forms the basis of our visual percep…

Object

InfiniteNature-Zero: Learning Perpetual View Generation of Natural Scenes from Single Images

2022-07-22 · Zhengqi Li, Qianqian Wang, Noah Snavely, Angjoo Kanazawa

We present a method for learning to generate unbounded flythrough videos of natural scenes starting from a single view, where this capability is learned from a collection of single photographs, without requiring camera p…

Perpetual View Generation

Pixelis: Reasoning in Pixels, from Seeing to Acting

2026-03-26 · Yunpeng Zhou arxiv

Most vision-language systems are static observers: they describe pixels, do not act, and cannot safely improve under shift. This passivity limits generalizable, physically grounded visual intelligence. Learning through a…

Visual Reasoning

BatMobility: Towards Flying Without Seeing for Autonomous Drones

2023-07-21 · Emerson Sie, Zikun Liu, Deepak Vasisht

Unmanned aerial vehicles (UAVs) rely on optical sensors such as cameras and lidar for autonomous operation. However, such optical sensors are error-prone in bad lighting, inclement weather conditions including fog and sm…

Collision AvoidanceOptical Flow Estimation

Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories

2026-04-10 · Wonbong Jang, Shikun Liu, Soubhik Sanyal, Juan Camilo Perez 외 arxiv

Recovering camera parameters from images and rendering scenes from novel viewpoints have been treated as separate tasks in computer vision and graphics. This separation breaks down when image coverage is sparse or poses …

Video GenerationPose Estimation