paper-with-me

홈 › Papers

Leveraging Prediction Entropy for Automatic Prompt Weighting in Zero-Shot Audio-Language Classification

2026-01-08 · Karim El Khoury, Maxime Zanella, Tiffanie Godelaine, Christophe De Vleeschouwer, Benoit Macq arxiv

Audio-language models have recently demonstrated strong zero-shot capabilities by leveraging natural-language supervision to classify audio events without labeled training data. Yet, their performance is highly sensitive to the wording of text prompts, with small variations leading to large fluctuations in accuracy. Prior work has mitigated this issue through prompt learning or prompt ensembling. However, these strategies either require annotated data or fail to account for the fact that some prompts may negatively impact performance. In this work, we present an entropy-guided prompt weighting approach that aims to find a robust combination of prompt contributions to maximize prediction confidence. To this end, we formulate a tailored objective function that minimizes prediction entropy to yield new prompt weights, utilizing low-entropy as a proxy for high confidence. Our approach can be applied to individual samples or a batch of audio samples, requiring no additional labels and incurring negligible computational overhead. Experiments on five audio classification datasets covering environmental, urban, and vocal sounds, demonstrate consistent gains compared to classical prompt ensembling methods in a zero-shot setting, with accuracy improvements 5-times larger across the whole benchmark.

📄 PDF Abstract BibTeX arXiv:2601.05011

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Classification

Similar Papers 제목 키워드 기반

Multi-head Reward Aggregation Guided by Entropy

2025-03-26 · Xiaomin Li, Xupeng Chen, Jingxuan Fan, Eric Hanchen Jiang 외

Aligning large language models (LLMs) with safety guidelines typically involves reinforcement learning from human feedback (RLHF), relying on human-generated preference annotations. However, assigning consistent overall …

Attribute

AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition

2024-03-18 · SooHwan Eom, Eunseop Yoon, Hee Suk Yoon, Chanwoo Kim 외

In Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a r…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding

2026-08-24 · Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu 외 arxiv

Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising object integrity in lightweight VLMs. Existin…

STaR-DRO: Stateful Tsallis Reweighting for Group-Robust Structured Prediction

2026-04-09 · Samah Fodeh, Ganesh Puthiaraju, Elyas Irankhah, Afshan Khan 외 arxiv

Structured prediction with large language models requires outputs that are label-accurate, ontology-constrained, structurally valid, and evidence-grounded under label imbalance and heterogeneous group difficulty. We pres…

Structured PredictionPrompt Engineering

Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

2026-06-12 · Ravi Ranjan, Utkarsh Grover, Xiaomin Lin, Agoritsa Polyzou arxiv

Transformer-based automatic speech recognition (ASR) models such as Whisper are highly accurate, but their predictions remain difficult to interpret. Existing explainable AI (XAI) methods often lack faithfulness and prec…

Speech Recognition