paper-with-me

홈 › Papers

MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders

2024-09-10 · Wenyu Zhang, Shuo Sun, Bin Wang, Xunlong Zou, Zhuohan Liu, Yingxu He, Geyu Lin, Nancy F. Chen, Ai Ti Aw

The rapid advancements in large language models (LLMs) have significantly enhanced natural language processing capabilities, facilitating the development of AudioLLMs that process and understand speech and audio inputs alongside text. Existing AudioLLMs typically combine a pre-trained audio encoder with a pre-trained LLM, which are subsequently finetuned on specific audio tasks. However, the pre-trained audio encoder has constrained capacity to capture features for new tasks and datasets. To address this, we propose to incorporate mixtures of `weak' encoders (MoWE) into the AudioLLM framework. MoWE supplements a base encoder with a pool of relatively light weight encoders, selectively activated based on the audio input to enhance feature extraction without significantly increasing model size. Our empirical results demonstrate that MoWE effectively improves multi-task performance, broadening the applicability of AudioLLMs to more diverse audio tasks.

📄 PDF Abstract BibTeX arXiv:2409.06635

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs

2025-06-07 · Wenyu Zhang, Yingxu He, Geyu Lin, Zhuohan Liu 외

Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches ofte…

Emotion Recognitionspeech-recognitionSpeech Recognition

FinAudio: A Benchmark for Audio Large Language Models in Financial Applications

2025-03-26 · Yupeng Cao, Haohang Li, Yangyang Yu, Shashidhar Reddy Javaji 외

Audio Large Language Models (AudioLLMs) have received widespread attention and have significantly improved performance on audio tasks such as conversation, audio understanding, and automatic speech recognition (ASR). Des…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Financial Analysisspeech-recognition+1

WEE-Therapy: A Mixture of Weak Encoders Framework for Psychological Counseling Dialogue Analysis

2025-09-24 · Yongqi Kang, Yong Zhao arxiv

The advancement of computational psychology requires AI tools capable of deeply understanding counseling dialogues. Existing audio language models (AudioLLMs) often rely on single speech encoders pre-trained on general d…

Emotion Recognition

AudioBench: A Universal Benchmark for Audio Large Language Models

2024-06-23 · Bin Wang, Xunlong Zou, Geyu Lin, Shuo Sun 외

We introduce AudioBench, a universal benchmark designed to evaluate Audio Large Language Models (AudioLLMs). It encompasses 8 distinct tasks and 26 datasets, among which, 7 are newly proposed datasets. The evaluation tar…

Audio Scene UnderstandingInstruction FollowingScene Understanding

Memory Augmented Language Models through Mixture of Word Experts

2023-11-15 · Cicero Nogueira dos santos, James Lee-Thorp, Isaac Noble, Chung-Ching Chang 외

Scaling up the number of parameters of language models has proven to be an effective approach to improve performance. For dense models, increasing model size proportionally increases the model's computation footprint. In…

Mixture-of-Experts