paper-with-me

홈 › Papers

AutoArabic: A Three-Stage Framework for Localizing Video-Text Retrieval Benchmarks

2025-09-19 · Mohamed Eltahir, Osamah Sarraj, Abdulrahman Alfrihidi, Taha Alshatiri, Mohammed Khurd, Mohammed Bremoo, Tanveer Hussain arxiv

Video-to-text and text-to-video retrieval are dominated by English benchmarks (e.g. DiDeMo, MSR-VTT) and recent multilingual corpora (e.g. RUDDER), yet Arabic remains underserved, lacking localized evaluation metrics. We introduce a three-stage framework, AutoArabic, utilizing state-of-the-art large language models (LLMs) to translate non-Arabic benchmarks into Modern Standard Arabic, reducing the manual revision required by nearly fourfold. The framework incorporates an error detection module that automatically flags potential translation errors with 97% accuracy. Applying the framework to DiDeMo, a video retrieval benchmark produces DiDeMo-AR, an Arabic variant with 40,144 fluent Arabic descriptions. An analysis of the translation errors is provided and organized into an insightful taxonomy to guide future Arabic localization efforts. We train a CLIP-style baseline with identical hyperparameters on the Arabic and English variants of the benchmark, finding a moderate performance gap (about 3 percentage points at Recall@1), indicating that Arabic localization preserves benchmark difficulty. We evaluate three post-editing budgets (zero/ flagged-only/ full) and find that performance improves monotonically with more post-editing, while the raw LLM output (zero-budget) remains usable. To ensure reproducibility to other languages, we made the code available at https://github.com/Tahaalshatiri/AutoArabic.

📄 PDF Abstract BibTeX arXiv:2509.16438

Code (0)

등록된 구현이 없습니다.

Tasks

Video-Text RetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

Boundary Proposal Network for Two-Stage Natural Language Video Localization

2021-03-15 · Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji 외

We aim to address the problem of Natural Language Video Localization (NLVL)-localizing the video segment corresponding to a natural language description in a long and untrimmed video. State-of-the-art NLVL methods are al…

Vocal Bursts Valence Prediction

Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding

2024-03-18 · CVPR 2024 1 · Chaolei Tan, JianHuang Lai, Wei-Shi Zheng, Jian-Fang Hu

Video Paragraph Grounding (VPG) is an emerging task in video-language understanding, which aims at localizing multiple sentences with semantic relations and temporal order from an untrimmed video. However, existing VPG a…

Multiple Instance Learning

Span-based Localizing Network for Natural Language Video Localization

2020-04-29 · ACL 2020 6 · Hao Zhang, Aixin Sun, Wei Jing, Joey Tianyi Zhou

Given an untrimmed video and a text query, natural language video localization (NLVL) is to locate a matching span from the video that semantically corresponds to the query. Existing solutions formulate NLVL either as a …

Temporal Sentence Grounding

UAL-Bench: The First Comprehensive Unusual Activity Localization Benchmark

2024-10-02 · Hasnat Md Abdullah, Tian Liu, Kangda Wei, Shu Kong 외

Localizing unusual activities, such as human errors or surveillance incidents, in videos holds practical significance. However, current video understanding models struggle with localizing these unusual events likely beca…

Unusual Activity LocalizationVideo Understanding

Actor-centered Representations for Action Localization in Streaming Videos

2021-04-29 · Sathyanarayanan N. Aakur, Sudeep Sarkar

Event perception tasks such as recognizing and localizing actions in streaming videos are essential for scaling to real-world application contexts. We tackle the problem of learning actor-centered representations through…

Action Localization