paper-with-me

홈 › Papers

Track and Caption Any Motion: Query-Free Motion Discovery and Description in Videos

2025-12-11 · Bishoy Galoaa, Sarah Ostadabbas arxiv

We propose Track and Caption Any Motion (TCAM), a motion-centric framework for automatic video understanding that discovers and describes motion patterns without user queries. Understanding videos in challenging conditions like occlusion, camouflage, or rapid movement often depends more on motion dynamics than static appearance. TCAM autonomously observes a video, identifies multiple motion activities, and spatially grounds each natural language description to its corresponding trajectory through a motion-field attention mechanism. Our key insight is that motion patterns, when aligned with contrastive vision-language representations, provide powerful semantic signals for recognizing and describing actions. Through unified training that combines global video-text alignment with fine-grained spatial correspondence, TCAM enables query-free discovery of multiple motion expressions via multi-head cross-attention. On the MeViS benchmark, TCAM achieves 58.4% video-to-text retrieval, 64.9 JF for spatial grounding, and discovers 4.8 relevant expressions per video with 84.7% precision, demonstrating strong cross-task generalization.

📄 PDF Abstract BibTeX arXiv:2512.10607

Code (0)

등록된 구현이 없습니다.

Tasks

Text Retrieval

Similar Papers 제목 키워드 기반

LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning

2024-10-09 · Zhe Li, Weihao Yuan, Yisheng He, Lingteng Qiu 외

Language plays a vital role in the realm of human motion. Existing methods have largely depended on CLIP text embeddings for motion generation, yet they fall short in effectively aligning language and motion due to CLIP'…

Large Language ModelMotion CaptioningMotion GenerationRepresentation Learning+2

MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

2024-07-08 · Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan 외

Sora's high-motion intensity and long consistent videos have significantly impacted the field of video generation, attracting unprecedented attention. However, existing publicly available datasets are inadequate for gene…

Video AlignmentVideo Generation

Similar Scenes arouse Similar Emotions: Parallel Data Augmentation for Stylized Image Captioning

2021-08-26 · Guodun Li, Yuchen Zhai, Zehao Lin, Yin Zhang

Stylized image captioning systems aim to generate a caption not only semantically related to a given image but also consistent with a given style description. One of the biggest challenges with this task is the lack of s…

Data AugmentationImage CaptioningSentence

Socratis: Are large multimodal models emotionally aware?

2023-08-31 · Katherine Deng, Arijit Ray, Reuben Tan, Saadia Gabriel 외

Existing emotion prediction benchmarks contain coarse emotion labels which do not consider the diversity of emotions that an image and text can elicit in humans due to various reasons. Learning diverse reactions to multi…

Articles

Motion Prediction Performance Analysis for Autonomous Driving Systems and the Effects of Tracking Noise

2021-04-16 · Ameni Trabelsi, Ross J. Beveridge, Nathaniel Blanchard

Autonomous driving consists of a multitude of interacting modules, where each module must contend with errors from the others. Typically, the motion prediction module depends upon a robust tracking system to capture each…

Autonomous Drivingmotion predictionPrediction