paper-with-me

Papers

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

2025-01-03 · Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta, Yaoting Wang, Mohamed Elhoseiny, Ruohan Gao, Dinesh Manocha

With the rapid advancement of Multi-modal Large Language Models (MLLMs), several diagnostic benchmarks have recently been developed to assess these models' multi-modal reasoning proficiency. However, these benchmarks are restricted to assessing primarily the visual aspect and do not examine the holistic audio-visual (AV) understanding. Moreover, currently, there are no benchmarks that investigate the capabilities of AVLLMs to calibrate their responses when presented with perturbed inputs. To this end, we introduce Audio-Visual Trustworthiness assessment Benchmark (AVTrustBench), comprising 600K samples spanning over 9 meticulously crafted tasks, evaluating the capabilities of AVLLMs across three distinct dimensions: Adversarial attack, Compositional reasoning, and Modality-specific dependency. Using our benchmark we extensively evaluate 13 state-of-the-art AVLLMs. The findings reveal that the majority of existing models fall significantly short of achieving human-like comprehension, offering valuable insights for future research directions. To alleviate the limitations in the existing approaches, we further propose a robust, model-agnostic calibrated audio-visual preference optimization based training strategy CAVPref, obtaining a gain up to 30.19% across all 9 tasks. We will publicly release our code and benchmark to facilitate future research in this direction.

📄 PDF Abstract BibTeX arXiv:2501.02135

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial AttackDiagnostic

Similar Papers 제목 키워드 기반

Multi-level Diagnosis and Evaluation for Robust Tabular Feature Engineering with Large Language Models

2025-09-20 · Yebin Lim, Susik Yoon arxiv

Recent advancements in large language models (LLMs) have shown promise in feature engineering for tabular data, but concerns about their reliability persist, especially due to variability in generated outputs. We introdu…

Feature Engineering

Robustness quantification: a new method for assessing the reliability of the predictions of a classifier

2025-03-28 · Adrián Detavernier, Jasper De Bock

Based on existing ideas in the field of imprecise probabilities, we present a new approach for assessing the reliability of the individual predictions of a generative probabilistic classifier. We call this approach robus…

Uncertainty Quantification

Robustness Quantification and Uncertainty Quantification: Comparing Two Methods for Assessing the Reliability of Classifier Predictions

2026-03-24 · Adrián Detavernier, Jasper De Bock arxiv

We consider two approaches for assessing the reliability of the individual predictions of a classifier: Robustness Quantification (RQ) and Uncertainty Quantification (UQ). We explain the conceptual differences between th…

MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues

2024-08-26 · Kuluhan Binici, Abhinav Ramesh Kashyap, Viktor Schlegel, Andy T. Liu 외

Automatic Speech Recognition (ASR) systems are pivotal in transcribing speech into text, yet the errors they introduce can significantly degrade the performance of downstream tasks like summarization. This issue is parti…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationIn-Context Learning+2

Assessing the Reliability of Deep Learning Classifiers Through Robustness Evaluation and Operational Profiles

2021-06-02 · Xingyu Zhao, Wei Huang, Alec Banks, Victoria Cox 외

The utilisation of Deep Learning (DL) is advancing into increasingly more sophisticated applications. While it shows great potential to provide transformational capabilities, DL also raises new challenges regarding its r…