paper-with-me

Papers

PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models

2026-01-07 · Yuwen Wang, Xinyuan Qian, Tian-Hao Zhang, Jiaran Gao, Yuchen Pan, Xin Wang, Zhou Pan, Chen Wei, Yiming Wang arxiv

Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is largely generic (e.g., summarizing spoken content) and fails to adequately support personalized question answering (e.g., summarizing what my best friend says). In contrast, human conditions their interpretation and decision-making on each individual's personal context. To bridge this gap, we formalize the task of Personalized LALMs (PALM) for recognizing personal concepts and reasoning within personal context. Moreover, we create the first benchmark (PALM-Bench) to foster the methodological advances in PALM and enable structured evaluation on several tasks across multi-speaker scenarios. Our extensive experiments on representative open-source LALMs, show that existing training-free prompting and supervised fine-tuning strategies, while yield improvements, remains limited in modeling personalized knowledge and transferring them across tasks robustly. Data and code will be released.

📄 PDF Abstract BibTeX arXiv:2601.03531

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

FedPalm: A General Federated Learning Framework for Closed- and Open-Set Palmprint Verification

2025-03-05 · Ziyuan Yang, Yingyu Chen, Chengrui Gao, Andrew Beng Jin Teoh 외

Current deep learning (DL)-based palmprint verification models rely on centralized training with large datasets, which raises significant privacy concerns due to biometric data's sensitive and immutable nature. Federated…

Federated LearningPrivacy Preserving

AudioPaLM: A Large Language Model That Can Speak and Listen

2023-06-22 · Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna 외

We introduce AudioPaLM, a large language model for speech understanding and generation. AudioPaLM fuses text-based and speech-based language models, PaLM-2 [Anil et al., 2023] and AudioLM [Borsos et al., 2022], into a un…

Language ModelingLanguage ModellingLarge Language Modelspeech-recognition+5

Towards Efficient Unconstrained Palmprint Recognition via Deep Distillation Hashing

2020-04-07 · Huikai Shao, DEXING ZHONG, Xuefeng Du

Deep palmprint recognition has become an emerging issue with great potential for personal authentication on handheld and wearable consumer devices. Previous studies of palmprint recognition are mainly based on constraine…

Knowledge Distillation

PaLM: Scaling Language Modeling with Pathways

2022-04-05 · Google Research 2022 4 · Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma 외

Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed t…

Auto DebuggingCode GenerationCommon Sense ReasoningCoreference Resolution+19

Talk3D: High-Fidelity Talking Portrait Synthesis via Personalized 3D Generative Prior

2024-03-29 · Jaehoon Ko, Kyusun Cho, Joungbin Lee, Heeji Yoon 외

Recent methods for audio-driven talking head synthesis often optimize neural radiance fields (NeRF) on a monocular talking portrait video, leveraging its capability to render high-fidelity and 3D-consistent novel-view fr…

NeRF