paper-with-me

Papers

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model

2025-03-12 · Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper

Integrating audio and visual data for training multimodal foundational models remains challenging. We present Audio-Video Vector Alignment (AVVA), which aligns audiovisual (AV) scene content beyond mere temporal synchronization via a Large Language Model (LLM)-based data curation pipeline. Specifically, AVVA scores and selects high-quality training clips using Whisper (speech-based audio foundation model) for audio and DINOv2 for video within a dual-encoder contrastive learning framework. Evaluations on AudioCaps, VALOR, and VGGSound demonstrate that this approach can achieve significant accuracy gains with substantially less curated data. For instance, AVVA yields a 7.6% improvement in top-1 accuracy for audio-to-video retrieval on VGGSound compared to ImageBind, despite training on only 192 hours of carefully filtered data (vs. 5800+ hours). Moreover, an ablation study highlights that trading data quantity for data quality improves performance, yielding respective top-3 accuracy increases of 47.8, 48.4, and 58.0 percentage points on AudioCaps, VALOR, and VGGSound over uncurated baselines. While these results underscore AVVA's data efficiency, we also discuss the overhead of LLM-driven curation and how it may be scaled or approximated in larger domains. Overall, AVVA provides a viable path toward more robust, text-free audiovisual learning with improved retrieval accuracy.

📄 PDF Abstract BibTeX arXiv:2503.09205

Code (0)

등록된 구현이 없습니다.

Tasks

AudioCapsContrastive LearningLarge Language ModelRetrievalVideo Retrieval

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation

2025-02-12 · Jinda Xu, Yuhao Song, Daming Wang, Weiwei Zhao 외

In an era overwhelmed by vast amounts of data, the effective curation of web-crawl datasets is essential for optimizing model performance. This paper tackles the challenges associated with the unstructured and heterogene…

Quality over Quantity: Demonstration Curation via Influence Functions for Data-Centric Robot Learning

2026-03-10 · Haeone Lee, Taywon Min, Junsu Kim, Sinjae Kang 외 arxiv

Learning from demonstrations has emerged as a promising paradigm for end-to-end robot control, particularly when scaled to diverse and large datasets. However, the quality of demonstration data, often collected through h…

Scalable Vision Language Model Training via High Quality Data Curation

2025-01-10 · Hongyuan Dong, Zijian Kang, Weijie Yin, Xiao Liang 외

In this paper, we introduce SAIL-VL (ScAlable Vision Language Model TraIning via High QuaLity Data Curation), an open-source vision language model (VLM) of state-of-the-art (SOTA) performance with 2B parameters. We intro…

Instruction FollowingLanguage ModelingLanguage Modelling

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

2025-08-05 · Lin Zhang, Zefan Cai, Yufan Zhou, Shentong Mo 외 arxiv

Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specif…

Bridging High-Quality Audio and Video via Language for Sound Effects Retrieval from Visual Queries

2023-08-17 · Julia Wilkins, Justin Salamon, Magdalena Fuentes, Juan Pablo Bello 외

Finding the right sound effects (SFX) to match moments in a video is a difficult and time-consuming task, and relies heavily on the quality and completeness of text metadata. Retrieving high-quality (HQ) SFX using a vide…

Contrastive LearningRetrieval