paper-with-me

홈 › Papers

Salmon: A Suite for Acoustic Language Model Evaluation

2024-09-11 · Gallil Maimon, Amit Roth, Yossi Adi

Speech language models have recently demonstrated great potential as universal speech processing systems. Such models have the ability to model the rich acoustic information existing in audio signals, beyond spoken content, such as emotion, background noise, etc. Despite this, evaluation benchmarks which evaluate awareness to a wide range of acoustic aspects, are lacking. To help bridge this gap, we introduce SALMon, a novel evaluation suite encompassing background noise, emotion, speaker identity and room impulse response. The proposed benchmarks both evaluate the consistency of the inspected element and how much it matches the spoken text. We follow a modelling based approach, measuring whether a model gives correct samples higher scores than incorrect ones. This approach makes the benchmark fast to compute even for large models. We evaluated several speech language models on SALMon, thus highlighting the strengths and weaknesses of each evaluated method. We make the code and data publicly available at https://pages.cs.huji.ac.il/adiyoss-lab/salmon/ .

📄 PDF Abstract BibTeX arXiv:2409.07437

Code (1)

slp-rl/salmon 공식 구현 pytorch

Tasks

Language Model EvaluationLanguage ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

A Multi-purpose Tracking Framework for Salmon Welfare Monitoring in Challenging Environments

2025-09-30 · Espen Uri Høgstedt, Christian Schellewald, Annette Stahl, Rudolf Mester arxiv

Computer Vision (CV)-based continuous, automated and precise salmon welfare monitoring is a key step toward reduced salmon mortality and improved salmon welfare in industrial aquaculture net pens. Available CV methods fo…

Pose Estimation

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

2024-06-22 · Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen 외

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end a…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+1

SALMONN: Towards Generic Hearing Abilities for Large Language Models

2023-10-20 · Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen 외

Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types o…

Audio captioningAutomatic Speech RecognitionEmotion RecognitionLanguage Modelling+8

video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language Models

2025-06-18 · Changli Tang, Yixuan Li, Yudong Yang, Jimin Zhuang 외

Videos contain a wealth of information, and generating detailed and accurate descriptions in natural language is a key aspect of video understanding. In this paper, we present video-SALMONN 2, an advanced audio-visual la…

Audio captioningLarge Language ModelQuestion AnsweringVideo Captioning+2

Patch Ensembles for Robust Salmon Re-Identification with Weak Trajectory Labels

2026-05-18 · Espen Uri Høgstedt, Christian Schellewald, Annette Stahl, Rudolf Mester arxiv

Salmon re-identification in commercial net-pens is challenging due to large populations, which impose strict accuracy requirements and make large-scale labeled data acquisition infeasible. Trajectory IDs can be used as p…