paper-with-me

Papers

AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

2025-12-18 · Sanjoy Chowdhury, Karren D. Yang, Xudong Liu, Fartash Faghri, Pavan Kumar Anasosalu Vasu, Oncel Tuzel, Dinesh Manocha, Chun-Liang Li, Raviteja Vemulapalli arxiv

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning tracking who speaks, maintaining roles, and grounding events across time. These scenarios are central to multimodal audio-video understanding, where models must jointly reason over audio and visual streams in applications such as conversational video assistants and meeting analytics. We introduce AMUSE, a benchmark designed around tasks that are inherently agentic, requiring models to decompose complex audio-visual interactions into planning, grounding, and reflection steps. It evaluates MLLMs across three modes zero-shot, guided, and agentic and six task families, including spatio-temporal speaker grounding and multimodal dialogue summarization. Across all modes, current models exhibit weak multi-speaker reasoning and inconsistent behavior under both non-agentic and agentic evaluation. Motivated by the inherently agentic nature of these tasks and recent advances in LLM agents, we propose RAFT, a data-efficient agentic alignment framework that integrates reward optimization with intrinsic multimodal self-evaluation as reward and selective parameter adaptation for data and parameter efficient updates. Using RAFT, we achieve up to 39.52\% relative improvement in accuracy on our benchmark. Together, AMUSE and RAFT provide a practical platform for examining agentic reasoning in multimodal models and improving their capabilities.

📄 PDF Abstract BibTeX arXiv:2512.16250

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation

2024-03-18 · Jun Yu, Wangyuan Zhu, Jichao Zhu

In this paper, we present the solution to the Emotional Mimicry Intensity (EMI) Estimation challenge, which is part of 6th Affective Behavior Analysis in-the-wild (ABAW) Competition.The EMI Estimation challenge task aims…

AMUSE: Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

2024-06-01 · IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) 2024 6 · Chhatre K., Danecek R., Athanasiou N., Becherini G. 외

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output anim…

Gesture GenerationRhythm

Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

2023-12-07 · CVPR 2024 1 · Kiran Chhatre, Radek Daněček, Nikos Athanasiou, Giorgio Becherini 외

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output anim…

Gesture GenerationRhythm

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

2026-05-25 · Tengfei Liu, Yang Shi, Xuanyu Zhu, Jiafu Tang 외 arxiv

Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to short-form settings. Existing benchmarks primarily focus on 5--10 secon…

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation

2025-12-23 · Jingqi Tian, Yiheng Du, Haoji Zhang, Yuji Wang 외 arxiv

Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by integrating auditory and visual cues. However, existing methods often struggle with multi-source entanglement and audio-visua…

Contrastive Learning