paper-with-me

Papers

Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio-Visual Event Perception

2023-01-01 · CVPR 2023 1 · Junyu Gao, Mengyuan Chen, Changsheng Xu

With only video-level event labels, this paper targets at the task of weakly-supervised audio-visual event perception (WS-AVEP), which aims to temporally localize and categorize events belonging to each modality. Despite the recent progress, most existing approaches either ignore the unsynchronized property of audio-visual tracks or discount the complementary modality for explicit enhancement. We argue that, for an event residing in one modality, the modality itself should provide ample presence evidence of this event, while the other complementary modality is encouraged to afford the absence evidence as a reference signal. To this end, we propose to collect Cross-Modal Presence-Absence Evidence (CMPAE) in a unified framework. Specifically, by leveraging uni-modal and cross-modal representations, a presence-absence evidence collector (PAEC) is designed under Subjective Logic theory. To learn the evidence in a reliable range, we propose a joint-modal mutual learning (JML) process, which calibrates the evidence of diverse audible, visible, and audi-visible events adaptively and dynamically. Extensive experiments show that our method surpasses state-of-the-arts (e.g., absolute gains of 3.6% and 6.1% in terms of event-level visual and audio metrics). Code is available in github.com/MengyuanChen21/CVPR2023-CMPAE.

📄 PDF Abstract BibTeX

Code (1)

mengyuanchen21/cvpr2023-cmpae 공식 구현 pytorch

Similar Papers 제목 키워드 기반

LiDAR-Event Stereo Fusion with Hallucinations

2024-08-08 · Luca Bartolomei, Matteo Poggi, Andrea Conti, Stefano Mattoccia

Event stereo matching is an emerging technique to estimate depth from neuromorphic cameras; however, events are unlikely to trigger in the absence of motion or the presence of large, untextured regions, making the corres…

Stereo Matching

Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding

2026-03-31 · Jingqi Xu arxiv

Vision-Language Models (VLMs) have demonstrated strong capabilities across a wide range of multimodal tasks. However, recent studies have shown that VLMs, such as CLIP, perform poorly in understanding negation expression…

Text Retrieval

Treatment Effects in Extreme Regimes

2023-06-20 · Ahmed Aloui, Ali Hasan, Yuting Ng, Miroslav Pajic 외

Understanding treatment effects in extreme regimes is important for characterizing risks associated with different interventions. This is hindered by the unavailability of counterfactual outcomes and the rarity and diffi…

counterfactual

Traffic Incident Database with Multiple Labels Including Various Perspective Environmental Information

2023-12-17 · Shota Nishiyama, Takuma Saito, Ryo Nakamura, Go Ohtani 외

A large dataset of annotated traffic accidents is necessary to improve the accuracy of traffic accident recognition using deep learning models. Conventional traffic accident datasets provide annotations on traffic accide…

Presence-absence reasoning for evolutionary phenotypes

2014-10-14 · James P. Balhoff, T. Alexander Dececchi, Paula M. Mabee, Hilmar Lapp

Nearly invariably, phenotypes are reported in the scientific literature in meticulous detail, utilizing the full expressivity of natural language. Often it is particularly these detailed observations (facts) that are of …