paper-with-me

Papers

SemanticAC: Semantics-Assisted Framework for Audio Classification

2023-02-12 · Yicheng Xiao, Yue Ma, Shuyan Li, Hantao Zhou, Ran Liao, Xiu Li

In this paper, we propose SemanticAC, a semantics-assisted framework for Audio Classification to better leverage the semantic information. Unlike conventional audio classification methods that treat class labels as discrete vectors, we employ a language model to extract abundant semantics from labels and optimize the semantic consistency between audio signals and their labels. We verify that simple textual information from labels and advanced pretraining models enable more abundant semantic supervision for better performance. Specifically, we design a text encoder to capture the semantic information from the text extension of labels. Then we map the audio signals to align with the semantics of corresponding class labels via an audio encoder and a similarity calculation module so as to enforce the semantic consistency. Extensive experiments on two audio datasets, ESC-50 and US8K demonstrate that our proposed method consistently outperforms the compared audio classification methods.

📄 PDF Abstract BibTeX arXiv:2302.05940

Code (0)

등록된 구현이 없습니다.

Tasks

Audio ClassificationClassificationLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging

2026-03-15 · Shunlong Wu, Hai Lin, Shaoshen Chen, Tingwei Lu 외 arxiv

Existing KV cache compression methods generally operate on discrete tokens or non-semantic chunks. However, such approaches often lead to semantic fragmentation, where linguistically coherent units are disrupted, causing…

TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics

2025-09-30 · Yi-Cheng Lin, Yu-Hua Chen, Jia-Kai Dong, Yueh-Hsuan Huang 외 arxiv

Large audio-language models are advancing rapidly, yet most evaluations emphasize speech or globally sourced sounds, overlooking culturally distinctive cues. This gap raises a critical question: can current models genera…

Question Generation

Segment Relevance Estimation for Audio Analysis and Weakly-Labelled Classification

2019-11-12 · Juliano Henrique Foleiss, Tiago Fernandes Tavares

We propose a method that quantifies the importance, namely relevance, of audio segments for classification in weakly-labelled problems. It works by drawing information from a set of class-wise one-vs-all classifiers. By …

Audio ClassificationClassificationGeneral Classification

Semantics-Aware Human Motion Generation from Audio Instructions

2025-05-29 · Zi-An Wang, Shihao Zou, Shiyao Yu, Mingyuan Zhang 외

Recent advances in interactive technologies have highlighted the prominence of audio signals for semantic encoding. This paper explores a new task, where audio signals are used as conditioning inputs to generate motions …

Motion Generation

MPN: Multimodal Parallel Network for Audio-Visual Event Localization

2021-04-07 · Jiashuo Yu, Ying Cheng, Rui Feng

Audio-visual event localization aims to localize an event that is both audible and visible in the wild, which is a widespread audio-visual scene analysis task for unconstrained videos. To address this task, we propose a …

audio-visual event localizationGeneral Classification