paper-with-me

홈 › Papers

AudioPairBank: Towards A Large-Scale Tag-Pair-Based Audio Content Analysis

2016-07-13 · Sebastian Sager, Benjamin Elizalde, Damian Borth, Christian Schulze, Bhiksha Raj, Ian Lane

Recently, sound recognition has been used to identify sounds, such as car and river. However, sounds have nuances that may be better described by adjective-noun pairs such as slow car, and verb-noun pairs such as flying insects, which are under explored. Therefore, in this work we investigate the relation between audio content and both adjective-noun pairs and verb-noun pairs. Due to the lack of datasets with these kinds of annotations, we collected and processed the AudioPairBank corpus consisting of a combined total of 1,123 pairs and over 33,000 audio files. One contribution is the previously unavailable documentation of the challenges and implications of collecting audio recordings with these type of labels. A second contribution is to show the degree of correlation between the audio content and the labels through sound recognition experiments, which yielded results of 70% accuracy, hence also providing a performance benchmark. The results and study in this paper encourage further exploration of the nuances in audio and are meant to complement similar research performed on images and text in multimedia analysis.

📄 PDF Abstract BibTeX arXiv:1607.03766

Code (0)

등록된 구현이 없습니다.

Tasks

TAG

Similar Papers 제목 키워드 기반

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

2024-11-28 · Jisheng Bai, Haohe Liu, Mou Wang, Dongyuan Shi 외

With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to the time-intensive and labour-heavy demand…

Audio captioningAudio to Text RetrievalCaption GenerationRetrieval+1

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

2025-01-30 · Yuqin Cao, Xiongkuo Min, Yixuan Gao, Wei Sun 외

Many video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual qualit…

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

2026-05-18 · Haojie Zheng, Yixin Yang, Siqi Yang, Shuchen Weng 외 arxiv

Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, …

Instruction FollowingVideo Generation

Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning

2023-09-20 · Luoyi Sun, Xuenan Xu, Mengyue Wu, Weidi Xie

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limit…

Audio captioningCaption GenerationImage Captioningobject-detection+5

Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal

2025-12-14 · Weihan Xu, Kan Jen Cheng, Koichi Saito, Muhammad Jehanzeb Mirza 외 arxiv

Joint editing of audio and visual content is crucial for precise and controllable content creation. This new task poses challenges due to the limitations of paired audio-visual data before and after targeted edits, and t…

Semantic correspondence