paper-with-me

Papers

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation

2025-10-09 · Kang Liao, Size Wu, Zhonghua Wu, Linyi Jin, Chao Wang, Yikai Wang, Fei Wang, Wei Li, Chen Change Loy arxiv

Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric multimodal model that extends spatial awareness along the camera dimension. Puffin integrates language regression and diffusion-based generation to interpret and create scenes from arbitrary viewpoints. To bridge the modality gap between cameras and vision-language, we introduce a novel paradigm that treats camera as language, enabling thinking with camera. This guides the model to align spatially grounded visual cues with photographic terminology while reasoning across geometric context. Puffin is trained on Puffin-4M, a large-scale dataset of 4 million vision-language-camera triplets. We incorporate both global camera parameters and pixel-wise camera maps, yielding flexible and reliable spatial generation. Experiments demonstrate Puffin superior performance over specialized models for camera-centric generation and understanding. With instruction tuning, Puffin generalizes to diverse cross-view tasks such as spatial imagination, world exploration, and photography guidance. We will release the code, models, dataset pipeline, and benchmark to advance multimodal spatial intelligence research.

📄 PDF Abstract BibTeX arXiv:2510.08673

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EgoM2P: Egocentric Multimodal Multitask Pretraining

2025-06-09 · Gen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao 외

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction. These capabilities en…

Depth EstimationGaze PredictionMonocular Depth Estimation

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

2025-11-06 · Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li 외 arxiv

The "Thinking with Text" and "Thinking with Images" paradigms significantly improve the reasoning abilities of large language models (LLMs) and Vision-Language Models (VLMs). However, these paradigms have inherent limita…

Multimodal ReasoningVideo Generation

Ego-Grounding for Personalized Question-Answering in Egocentric Videos

2026-04-02 · Junbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela Yao arxiv

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this …

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

2026-07-20 · Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang 외 hf

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal…

Depth Estimation

EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera

2026-05-12 · Christen Millerdurai, Shaoxiang Wang, Yaxu Xie, Vladislav Golyanik 외 arxiv

Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulatio…