paper-with-me

홈 › Papers

IAE-VTG: Interaction-Aligned Action-Entity Video Temporal Grounding

2026-09-09 · Shiwen Zhao, Qi Zhang, Sezer Karaoglu, Theo Gevers, Martin R. Oswald arxiv

Video Temporal Grounding (VTG) localizes the video segment that matches a natural-language query. Many queries describe an action performed by a particular entity. Existing methods often encode the query as a whole or use general video-text interactions, without explicitly checking whether the action and entity occur together. They may therefore select a segment that contains both concepts but not the event described by the query. We propose Interaction Aligned Action-Entity Video Temporal Grounding (IAE-VTG), which models this rela?tionship at both the representation and training assignment levels. First, the Fine-grained Disentangled Interaction Module (FDIM) separates action and entity related query information and aligns it with complementary motion and appearance features. It then combines token-level interactions to build representations that capture the relationship between the action and entity. Second, Interaction-Sensitive Assignment (ISA) adds this interaction evidence to bipartite matching, so training targets are selected using both temporal overlap and semantic compatibility. This reduces supervision from temporally plausible but semantically incorrect proposals. Experiments on QVHighlights, Charades?STA, and TACoS show that IAE-VTG consistently improves strong baselines and achieves competitive or state-of-the-art performance on standard grounding metrics. Additional analyses show that the method is especially effective when similar actions or entities appear at multiple times and produces more reliable assignments for complex events.

📄 PDF Abstract BibTeX arXiv:2609.09736

Code (1)

arxivsub/arXivSub_daily_arxiv ★ 4

Similar Papers 제목 키워드 기반

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

2025-06-11 · Zhenzhi Wang, Jiaqi Yang, Jianwen Jiang, Chao Liang 외

End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject…

Human AnimationHuman-Object Interaction Detection

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

2024-12-05 · Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan 외

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency…

Portrait AnimationVideo Generation

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

2026-05-02 · Alejandro Aparcedo, Akash Kumar, Aaryan Garg, Dalton Pham 외 arxiv

Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, m…

Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

2026-07-25 · Narthana Sivalingam, Santhirarajah Sivasthigan, Buddhi Wijenayake, Roshan Godaliyadda 외 arxiv

Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approaches perform well but often rely on deep stacks of layers to learn these heteroge…

Group Activity Recognition

Identity-Aware Human-Object Interaction Motion Captioning

2026-08-21 · Yiming Wang, Yonghao Dang, Huilai Li, Jiawei Tu 외 arxiv

Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subje…

Motion Captioning