paper-with-me

홈 › Papers

Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event Parser

2023-05-27 · NeurIPS 2023 11 · Yung-Hsuan Lai, Yen-Chun Chen, Yu-Chiang Frank Wang

Audio-visual learning has been a major pillar of multi-modal machine learning, where the community mostly focused on its modality-aligned setting, i.e., the audio and visual modality are both assumed to signal the prediction target. With the Look, Listen, and Parse dataset (LLP), we investigate the under-explored unaligned setting, where the goal is to recognize audio and visual events in a video with only weak labels observed. Such weak video-level labels only tell what events happen without knowing the modality they are perceived (audio, visual, or both). To enhance learning in this challenging setting, we incorporate large-scale contrastively pre-trained models as the modality teachers. A simple, effective, and generic method, termed Visual-Audio Label Elaboration (VALOR), is innovated to harvest modality labels for the training events. Empirical studies show that the harvested labels significantly improve an attentional baseline by 8.0 in average F-score (Type@AV). Surprisingly, we found that modality-independent teachers outperform their modality-fused counterparts since they are noise-proof from the other potentially unaligned modality. Moreover, our best model achieves the new state-of-the-art on all metrics of LLP by a substantial margin (+5.4 F-score for Type@AV). VALOR is further generalized to Audio-Visual Event Localization and achieves the new state-of-the-art as well. Code is available at: https://github.com/Franklin905/VALOR.

📄 PDF Abstract BibTeX arXiv:2305.17343

Code (1)

franklin905/valor 공식 구현 pytorch

Tasks

audio-visual event localizationaudio-visual learning

Similar Papers 제목 키워드 기반

Curriculum Learning Meets Weakly Supervised Modality Correlation Learning

2022-12-15 · Sijie Mai, Ya Sun, Haifeng Hu

In the field of multimodal sentiment analysis (MSA), a few studies have leveraged the inherent modality correlation information stored in samples for self-supervised learning. However, they feed the training pairs in a r…

Multimodal Sentiment AnalysisSelf-Supervised LearningSentiment Analysis

Weakly Supervised Visible-Infrared Person Re-Identification via Heterogeneous Expert Collaborative Consistency Learning

2025-07-17 · Yafei Zhang, Lingqi Kong, Huafeng Li, Jie Wen

To reduce the reliance of visible-infrared person re-identification (ReID) models on labeled cross-modal samples, this paper explores a weakly supervised cross-modal person ReID method that uses only single-modal sample …

Person Re-Identification

WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-only Supervised Text Spotting

2024-07-28 · Jingjing Wu, Zhengyao Fang, Pengyuan Lyu, Chengquan Zhang 외

Transcription-only Supervised Text Spotting aims to learn text spotters relying only on transcriptions but no text boundaries for supervision, thus eliminating expensive boundary annotation. The crux of this task lies in…

Contrastive LearningText Spotting

AI as a Teaching Partner: Early Lessons from Classroom Codesign with Secondary Teachers

2025-12-12 · Alex Liu, Lief Esbenshade, Shawon Sarkar, Zewei Tian 외 arxiv

This report presents a comprehensive account of the Colleague AI Classroom pilot, a collaborative design (co-design) study that brought generative AI technology directly into real classrooms. In this study, AI functioned…

Learning When to Trust Which Teacher for Weakly Supervised ASR

2023-06-21 · Aakriti Agrawal, Milind Rao, Anit Kumar Sahu, Gopinath Chennupati 외

Automatic speech recognition (ASR) training can utilize multiple experts as teacher models, each trained on a specific domain or accent. Teacher models may be opaque in nature since their architecture may be not be known…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition