paper-with-me

홈 › Papers

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

2025-10-16 · Gabriel Fiastre, Antoine Yang, Cordelia Schmid arxiv

Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the ability to understand spatio-temporal details and describe them in natural language. Due to the complexity of the task and the high cost associated with manual annotation, previous approaches resort to training strategies with limited data, potentially leading to suboptimal performance. To circumvent this issue, we propose to generate captions about spatio-temporally localized entities leveraging a state-of-the-art VLM, and extend the LVIS and LV-VIS datasets with our synthetic captions (LVISCap and LV-VISCap). Moreover, we introduce an end-to-end model, CaptionFormer, capable of jointly detecting, segmenting, tracking and captioning object trajectories. CaptionFormer achieves state-of-the-art DVOC results on three existing benchmarks, VidSTG, VLN and BenSMOT. The datasets and code are available at https://www.gabriel.fiastre.fr/captionformer/.

📄 PDF Abstract BibTeX arXiv:2510.14904

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

2025-04-07 · Yunlong Tang, Jing Bi, Chao Huang, Susan Liang 외

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects through time. CAT-V integrates three ke…

Boundary DetectionObjectSemantic SegmentationVideo Captioning

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

2026-05-20 · Aditya Chetan, Eric Cai, Peeyush Kushwaha, Bharath Raj Nagoor Kani 외 arxiv

The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus predominantly on coarse-grained tasks such as action segmentation, cla…

Action Segmentation

Learning Spatio-Appearance Memory Network for High-Performance Visual Tracking

2020-09-21 · Fei Xie, Wankou Yang, Bo Liu, Kaihua Zhang 외

Existing visual object tracking usually learns a bounding-box based template to match the targets across frames, which cannot accurately learn a pixel-wise representation, thereby being limited in handling severe appeara…

Object TrackingSegmentationSemantic SegmentationVideo Object Segmentation+4

DVIS++: Improved Decoupled Framework for Universal Video Segmentation

2023-12-20 · Tao Zhang, Xingye Tian, Yikang Zhou, Shunping Ji 외

We present the \textbf{D}ecoupled \textbf{VI}deo \textbf{S}egmentation (DVIS) framework, a novel approach for the challenging task of universal video segmentation, including video instance segmentation (VIS), video seman…

Contrastive LearningDenoisingInstance SegmentationPanoptic Segmentation+6

Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

2025-06-09 · Xiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu 외

We present Genesis, a unified framework for joint generation of multi-view driving videos and LiDAR sequences with spatio-temporal and cross-modal consistency. Genesis employs a two-stage architecture that integrates a D…

NeRFScene Generation