paper-with-me

홈 › Papers

AC/DC: LLM-based Audio Comprehension via Dialogue Continuation

2025-06-12 · Yusuke Fujita, Tomoya Mizumoto, Atsushi Kojima, Lianbo Liu, Yui Sudo

We propose an instruction-following audio comprehension model that leverages the dialogue continuation ability of large language models (LLMs). Instead of directly generating target captions in training data, the proposed method trains a model to produce responses as if the input caption triggered a dialogue. This dialogue continuation training mitigates the caption variation problem. Learning to continue a dialogue effectively captures the caption's meaning beyond its surface-level words. As a result, our model enables zero-shot instruction-following capability without multitask instruction tuning, even trained solely on audio captioning datasets. Experiments on AudioCaps, WavCaps, and Clotho datasets with AudioBench audio-scene question-answering tests demonstrate our model's ability to follow various unseen instructions.

📄 PDF Abstract BibTeX arXiv:2506.10312

Code (0)

등록된 구현이 없습니다.

Tasks

AudioCapsAudio captioningInstruction FollowingQuestion Answering

Similar Papers 제목 키워드 기반

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

2025-12-28 · Kai Liu, Jungang Li, Yuchong Sun, Shengqiong Wu 외 arxiv

This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoder-LLM-decoder architecture, which has a SyncFu…

Realtime-Venus: A full-duplex interaction system with asynchronous delegation

2026-09-12 · Ruixiang Zhao, Hualei Wang, Renhe Sun, Enzhi Zhou 외 hf

Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the co…

Question Answering

Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias

2025-09-26 · Shree Harsha Bokkahalli Satish, Harm Lameris, Olivier Perrotin, Gustav Eje Henter 외 arxiv

Speech Continuation (SC) is the task of generating a coherent extension of a spoken prompt while preserving both semantic context and speaker identity. Because SC is constrained to a single audio stream, it offers a more…

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

2026-09-18 · Haolin He, Yunfei Chu, Qi Chen, Wen Huang 외 hf

We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return tex…

Reinforcement LearningSpeech RecognitionVideo Generation

Covo-Audio Technical Report

2026-02-10 · Wenfu Wang, Chenxing Li, Liqiang Zhang, Yiyang Zhao 외 arxiv

In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture. Through large-scale curated pretrainin…

Instruction Following