paper-with-me

Papers

Creating HAVIC: Heterogeneous Audio Visual Internet Collection

2012-05-01 · LREC 2012 5 · Stephanie Strassel, Am Morris, a, Jonathan Fiscus, Christopher Caruso, Haejoong Lee, Paul Over, James Fiumara, Barbara Shaw, Brian Antonishek, Martial Michel

Linguistic Data Consortium and the National Institute of Standards and Technology are collaborating to create a large, heterogeneous annotated multimodal corpus to support research in multimodal event detection and related technologies. The HAVIC (Heterogeneous Audio Visual Internet Collection) Corpus will ultimately consist of several thousands of hours of unconstrained user-generated multimedia content. HAVIC has been designed with an eye toward providing increased challenges for both acoustic and video processing technologies, focusing on multi-dimensional variation inherent in user-generated multimedia content. To date the HAVIC corpus has been used to support the NIST 2010 and 2011 TRECVID Multimedia Event Detection (MED) Evaluations. Portions of the corpus are expected to be released in LDC's catalog in the coming year, with the remaining segments being published over time after their use in the ongoing MED evaluations.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Event Detection

Similar Papers 제목 키워드 기반

Leave No Stone Unturned: Uncovering Holistic Audio-Visual Intrinsic Coherence for Deepfake Detection

2026-03-25 · Jielun Peng, Yabin Wang, Yaqi Li, Long Kong 외 arxiv

The rapid progress of generative AI has enabled hyper-realistic audio-visual deepfakes, intensifying threats to personal security and social trust. Most existing deepfake detectors rely either on uni-modal artifacts or a…

DeepFake Detection

From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

2024-09-27 · Kun Su, Xiulong Liu, Eli Shlizerman

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interpl…

Audio ClassificationAudio GenerationRepresentation Learning

Adversarial-Metric Learning for Audio-Visual Cross-Modal Matching

2021-01-12 · IEEE Transactions on Multimedia 2021 1 · Aihua Zheng, Menglan Hu, Bo Jiang *, Yan Huang 외

Audio-visual matching aims to learn the intrinsic correspondence between image and audio clip. Existing works mainly concentrate on learning discriminative features, while ignore the cross-modal heterogeneous issue betwe…

audio-visual learningMetric LearningRetrieval

Temporarily-Aware Context Modelling using Generative Adversarial Networks for Speech Activity Detection

2020-04-02 · Tharindu Fernando, Sridha Sridharan, Mitchell McLaren, Darshana Priyasad 외

This paper presents a novel framework for Speech Activity Detection (SAD). Inspired by the recent success of multi-task learning approaches in the speech processing domain, we propose a novel joint learning framework for…

Action DetectionActivity DetectionMulti-Task Learning

Heterogeneous Graph Learning for Acoustic Event Classification

2023-03-05 · Amir Shirian, Mona Ahmadian, Krishna Somandepalli, Tanaya Guha

Heterogeneous graphs provide a compact, efficient, and scalable way to model data involving multiple disparate modalities. This makes modeling audiovisual data using heterogeneous graphs an attractive option. However, gr…

Classificationgraph constructionGraph Learning