paper-with-me

홈 › Papers

AQUALLM: Audio Question Answering Data Generation Using Large Language Models

2023-12-28 · Swarup Ranjan Behera, Krishna Mohan Injeti, Jaya Sai Kiran Patibandla, Praveen Kumar Pokala, Balakrishna Reddy Pailla

Audio Question Answering (AQA) constitutes a pivotal task in which machines analyze both audio signals and natural language questions to produce precise natural language answers. The significance of possessing high-quality, diverse, and extensive AQA datasets cannot be overstated when aiming for the precision of an AQA system. While there has been notable focus on developing accurate and efficient AQA models, the creation of high-quality, diverse, and extensive datasets for the specific task at hand has not garnered considerable attention. To address this challenge, this work makes several contributions. We introduce a scalable AQA data generation pipeline, denoted as the AQUALLM framework, which relies on Large Language Models (LLMs). This framework utilizes existing audio-caption annotations and incorporates state-of-the-art LLMs to generate expansive, high-quality AQA datasets. Additionally, we present three extensive and high-quality benchmark datasets for AQA, contributing significantly to the progression of AQA research. AQA models trained on the proposed datasets set superior benchmarks compared to the existing state-of-the-art. Moreover, models trained on our datasets demonstrate enhanced generalizability when compared to models trained using human-annotated AQA data. Code and datasets will be accessible on GitHub~\footnote{\url{https://github.com/swarupbehera/AQUALLM}}.

📄 PDF Abstract BibTeX arXiv:2312.17343

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Question AnsweringQuestion Answering

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

2023-08-22 · Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, Ying Shan

Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU…

Caption GenerationLarge Language ModelMultimodal Music GenerationMusic Captioning+2

Event-Grounded Question Answering over Long Audio via Structured Retrieval

2026-02-16 · Kartik Hegde, Arvind Krishna Sridhar, Naveen Vakada, Yinyi Guo 외 arxiv

Answering natural-language questions over multi-hour audio requires reliable event recognition, temporal grounding, and efficient retrieval. We present LA-RAG (Long Audio Retrieval-Augmented Generation), a structured fra…

Question AnsweringMoment Retrieval

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity

2026-02-04 · Gaia A. Bertolino, Yuwei Zhang, Tong Xia, Domenico Talia 외 arxiv

As conversational multimodal AI tools are increasingly adopted to process patient data for health assessment, robust benchmarks are needed to measure progress and expose failure modes under realistic conditions. Despite …

Question Answering

Audiopedia: Audio QA with Knowledge

2024-12-29 · Abhirama Subramanyam Penamakuri, Kiran Chhatre, Akshat Jain

In this paper, we introduce Audiopedia, a novel task called Audio Question Answering with Knowledge, which requires both audio comprehension and external knowledge reasoning. Unlike traditional Audio Question Answering (…

Audio Question AnsweringEntity LinkingQuestion Answering

Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

2025-02-24 · Tianpeng Li, Jun Liu, Tao Zhang, Yuanbo Fang 외

We introduce Baichuan-Audio, an end-to-end audio large language model that seamlessly integrates audio understanding and generation. It features a text-guided aligned speech generation mechanism, enabling real-time speec…

Language ModelingLanguage ModellingLarge Language ModelQuestion Answering