Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models
Audio-language models (ALMs) often follow text that conflicts with audio, even when the audio evidence is clear. This raises a basic question: is the audio-supported answer unavailable, or is it represented but overridden by the conflicting text? We examine this question using a same-audio counterfactual that keeps the audio fixed, removes only the conflicting text, and measures the resulting shift in model preference. Across five ALMs and four conflict tasks, 64.1% of conflict samples show a sign flip: the same-audio branch prefers the audio-supported answer, whereas the joint branch prefers the text-supported answer. This pattern suggests that the relevant audio evidence is encoded but loses in arbitration. Activation patching further localizes the reversal to answer-position computation, and patching effects closely track output candidate-score differences (Spearman rho=0.93). Using this diagnostic, we propose Gated Audio Counterfactual Logit Correction (GACL), a training-free decoding rule that interpolates between joint and same-audio scores. Under a strict 5 pp faithfulness-drop budget, GACL improves nAUC by 17.8 points over the best contrastive baseline and transfers without retuning to vision-text arbitration (up to +40.5 pp).
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration
Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliability of multimodal large language models (MLLMs) in real-world deployme…
DAFE: LLM-Based Evaluation Through Dynamic Arbitration for Free-Form Question-Answering
Evaluating Large Language Models (LLMs) free-form generated responses remains a challenge due to their diverse and open-ended nature. Traditional supervised signal-based automatic metrics fail to capture semantic equival…
FormInstruction FollowingQuestion AnsweringDominance reversals: The resolution of genetic conflict and maintenance of genetic variation
Beneficial reversals of dominance reduce the costs of genetic trade-offs and can enable selection to maintain genetic variation for fitness. Beneficial dominance reversals are characterized by the beneficial allele for a…
Don't Kill the Baby: The Case for AI in Arbitration
Since the introduction of Generative AI (GenAI) in 2022, its ability to simulate human intelligence and generate content has sparked both enthusiasm and concern. While much criticism focuses on AI's potential to perpetua…
FairnessPacos: Modeling Users' Interpretable and Context-Dependent Choices in Preference Reversals
Choice problems refer to selecting the best choices from several items, and learning users' preferences in choice problems is of great significance in understanding the decision making mechanisms and providing personaliz…
Decision Making