paper-with-me

Papers

Roadmap towards Superhuman Speech Understanding using Large Language Models

2024-10-17 · Fan Bu, Yuhao Zhang, Xidong Wang, Benyou Wang, Qun Liu, Haizhou Li

The success of large language models (LLMs) has prompted efforts to integrate speech and audio data, aiming to create general foundation models capable of processing both textual and non-textual inputs. Recent advances, such as GPT-4o, highlight the potential for end-to-end speech LLMs, which preserves non-semantic information and world knowledge for deeper speech understanding. To guide the development of speech LLMs, we propose a five-level roadmap, ranging from basic automatic speech recognition (ASR) to advanced superhuman models capable of integrating non-semantic information with abstract acoustic knowledge for complex tasks. Moreover, we design a benchmark, SAGI Bechmark, that standardizes critical aspects across various tasks in these five levels, uncovering challenges in using abstract acoustic knowledge and completeness of capability. Our findings reveal gaps in handling paralinguistic cues and abstract acoustic knowledge, and we offer future directions. This paper outlines a roadmap for advancing speech LLMs, introduces a benchmark for evaluation, and provides key insights into their current limitations and potential.

📄 PDF Abstract BibTeX arXiv:2410.13268

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionWorld Knowledge

Similar Papers 제목 키워드 기반

What's the Meaning of Superhuman Performance in Today's NLU?

2023-05-15 · Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajic 외

In the last five years, there has been a significant focus in Natural Language Processing (NLP) on developing larger Pretrained Language Models (PLMs) and introducing benchmarks such as SuperGLUE and SQuAD to measure the…

PositionReading Comprehension

Superalignment with Dynamic Human Values

2025-03-17 · Florian Mai, David Kaczér, Nicholas Kluge Corrêa, Lucie Flek

Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they do not simultaneously account for 2). W…

CoDA21: Evaluating Language Understanding Capabilities of NLP Models With Context-Definition Alignment

2021-09-17 · ACL ARR September 2021 9 · Anonymous

Pretrained language models (PLMs) have achieved superhuman performance on many benchmarks, creating a need for harder tasks. We introduce CoDA21 (Context Definition Alignment), a challenging benchmark that measures natur…

Natural Language UnderstandingWorld Knowledge

CoDA21: Evaluating Language Understanding Capabilities of NLP Models With Context-Definition Alignment

2022-03-11 · ACL 2022 5 · Lütfi Kerem Senel, Timo Schick, Hinrich Schütze

Pretrained language models (PLMs) have achieved superhuman performance on many benchmarks, creating a need for harder tasks. We introduce CoDA21 (Context Definition Alignment), a challenging benchmark that measures natur…

Natural Language UnderstandingWorld Knowledge

Humanly Certifying Superhuman Classifiers

2021-09-16 · Qiongkai Xu, Christian Walder, Chenchen Xu

Estimating the performance of a machine learning system is a longstanding challenge in artificial intelligence research. Today, this challenge is especially relevant given the emergence of systems which appear to increas…