paper-with-me

Papers

SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation

2025-01-01 · CVPR 2025 1 · Hao Du, Bo Wu, Yan Lu, Zhendong Mao

Vision-language temporal alignment is a crucial capability for human dynamic recognition and cognition in real-world scenarios. While existing research focuses on capturing vision-language relevance, it faces limitations due to biased temporal distributions, imprecise annotations, and insufficient compositionally. To achieve fair evaluation and comprehensive exploration, our objective is to investigate and evaluate the ability of models to achieve alignment from a temporal perspective, specifically focusing on their capacity to synchronize visual scenarios with linguistic context in a temporally coherent manner. As a preliminary step, we present the statistical analysis of existing benchmarks and reveal the existing challenges from a decomposed perspective. To this end, we introduce SVLTA, the synthetic vision-language temporal alignment derived via a well-designed and feasible control generation method within a simulation environment. The approach considers commonsense knowledge, manipulable action, and constrained filtering, which generates reasonable, diverse, and balanced data distributions for diagnostic evaluations. Our experiments reveal diagnostic insights through the evaluations in temporal question answering, distributional shift sensitiveness, and temporal alignment adaptation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDiagnosticQuestion Answering

Similar Papers 제목 키워드 기반

Benchmarking Cross-Lingual Semantic Alignment in Multilingual Embeddings

2025-12-29 · Wen G. Gong arxiv

With hundreds of multilingual embedding models available, practitioners lack clear guidance on which provide genuine cross-lingual semantic alignment versus task performance through language-specific patterns. Task-drive…

Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties

2025-02-24 · Zhenglin Wang, Jialong Wu, Pengfei Li, Yong Jiang 외

Temporal reasoning is fundamental to human cognition and is crucial for various real-world applications. While recent advances in Large Language Models have demonstrated promising capabilities in temporal reasoning, exis…

Benchmarking

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation

2026-06-26 · Anya Ji, Abhijith Varma Mudunuri, David M. Chan, Alane Suhr arxiv

While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code for webpages, charts, or SVGs, it remains unclear whether they can recover temporal…

Visual ReasoningCode Generation

Video-Language Alignment via Spatio-Temporal Graph Transformer

2024-07-16 · Shi-Xue Zhang, Hongfa Wang, Xiaobin Zhu, Weibo Gu 외

Video-language alignment is a crucial multi-modal task that benefits various downstream applications, e.g., video-text retrieval and video question answering. Existing methods either utilize multi-modal information in vi…

Contrastive LearningQuestion AnsweringRetrievalText Retrieval+2

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

2026-08-26 · Eunjee Choi, JungHoon Sung, Seongwhan Cho, Chu Xin 외 arxiv

Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, provid…

Sign Language RecognitionRepresentation Learning