paper-with-me

Papers

DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation

2025-11-23 · Weichuang Shao, Iman Yi Liao, Tomas Henrique Bode Maul, Tissa Chandesa arxiv

Existing Test-time Adaptation (TTA) studies rely heavily on static and homogeneous corruption protocols, such as ImageNet-C and CIFAR-10-C/100-C, leading to inconsistent evaluation settings and potentially inflated robustness estimates that are compared with real-world situations. TTA lacks a standardized evaluation infrastructure capable of modeling realistic heterogeneous acoustic degradation. We introduce DHAuDS, a standardized benchmark suite for evaluating audio classification TTA robustness under dynamic corruption severity and heterogeneous noise mixtures. Rather than proposing a new TTA algorithm, DHAuDS focuses on exposing robustness limitations that remain hidden under conventional fixed-noise evaluation protocols.

📄 PDF Abstract BibTeX arXiv:2511.18421

Code (0)

등록된 구현이 없습니다.

Tasks

Test-time AdaptationAudio Classification

Similar Papers 제목 키워드 기반

CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges

2025-09-26 · Hui Li, Changhao Jiang, Hongyu Wang, Ming Zhang 외 arxiv

The ability to reason from audio, including speech, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks mainly focus on static or single-scene …

Accommodating Audio Modality in CLIP for Multimodal Processing

2023-03-12 · Ludan Ruan, Anwen Hu, Yuqing Song, Liang Zhang 외

Multimodal processing has attracted much attention lately especially with the success of pre-training. However, the exploration has mainly focused on vision-language pre-training, as introducing more modalities can great…

AudioCapsContrastive LearningLanguage ModelingLanguage Modelling+3

Heterogeneous bimodal attention fusion for speech emotion recognition

2025-03-09 · Jiachen Luo, Huy Phan, Lin Wang, Joshua Reiss

Multi-modal emotion recognition in conversations is a challenging problem due to the complex and complementary interactions between different modalities. Audio and textual cues are particularly important for understandin…

Contrastive LearningEmotion RecognitionSpeech Emotion Recognition

Qwen3.5-Omni Technical Report

2026-04-17 · Qwen Team arxiv

In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and suppor…

Scene SegmentationVisual GroundingSpeech Synthesis

Heterogeneous Graph Learning for Acoustic Event Classification

2023-03-05 · Amir Shirian, Mona Ahmadian, Krishna Somandepalli, Tanaya Guha

Heterogeneous graphs provide a compact, efficient, and scalable way to model data involving multiple disparate modalities. This makes modeling audiovisual data using heterogeneous graphs an attractive option. However, gr…

Classificationgraph constructionGraph Learning