paper-with-me

Papers

ProCap: Projection-Aware Captioning for Spatial Augmented Reality

2026-04-01 · Zimo Cao, Yuchen Deng, Haibin Ling, Bingyao Huang arxiv

Spatial augmented reality (SAR) directly projects digital content onto physical scenes using projectors, creating immersive experience without head-mounted displays. However, for SAR to support intelligent interaction, such as reasoning about the scene or answering user queries, it must semantically distinguish between the physical scene and the projected content. Standard Vision Language Models (VLMs) struggle with this virtual-physical ambiguity, often confusing the two contexts. To address this issue, we introduce ProCap, a novel framework that explicitly decouples projected content from physical scenes. ProCap employs a two-stage pipeline: first it visually isolates virtual and physical layers via automated segmentation; then it uses region-aware retrieval to avoid ambiguous semantic context due to projection distortion. To support this, we present RGBP (RGB + Projections), the first large-scale SAR semantic benchmark dataset, featuring 65 diverse physical scenes and over 180,000 projections with dense, decoupled annotations. Finally, we establish a dual-captioning evaluation protocol using task-specific tokens to assess physical scene and projection descriptions independently. Our experiments show that ProCap provides a robust semantic foundation for future SAR research. The source code, pre-trained models and the RGBP dataset are available on the project page: https://ZimoCao.github.io/ProCap/.

📄 PDF Abstract BibTeX arXiv:2604.00912

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Imagine How To Change: Explicit Procedure Modeling for Change Captioning

2026-03-06 · Jiayang Sun, Zixin Guo, Min Cao, Guibo Zhu 외 arxiv

Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on static image pairs, thus ignoring the rich temporal dynamics of the chang…

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

2026-07-23 · Debjyoti Das Adhikary, Aritra Hazra, Partha Pratim Chakrabarti arxiv

Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected object…

Video Captioning

RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning

2025-09-19 · Xiaosheng Long, Hanyu Wang, Zhentao Song, Kun Luo 외 arxiv

Recent retrieval-augmented image captioning methods incorporate external knowledge to compensate for the limitations in comprehending complex scenes. However, current approaches face challenges in relation modeling: (1) …

Image Captioning

OmniAI: A Surface-Adaptive Aerial Projection Interface for Human--Drone Interaction

2026-08-01 · Nikita Kuzmin, Yuhua Jin, Georgii Demianchuk, Mariya Lezina 외 arxiv

Drones in human environments often lack spatially grounded in- terfaces for situated communication. We present OmniAI, an em- bodied aerial agent that supports surface-adaptive interaction by switching projection between…

Towards Fine-Grained Human Motion Video Captioning

2025-10-24 · Guorui Song, Guocun Wang, Zhe Huang, Jing Lin 외 arxiv

Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motion details, resulting in vague or semanti…

Human Mesh RecoveryVideo Captioning