paper-with-me

Papers

LSSED: a large-scale dataset and benchmark for speech emotion recognition

2021-01-30 · Weiquan Fan, Xiangmin Xu, Xiaofen Xing, Weidong Chen, DongYan Huang

Speech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED, a challenging large-scale english speech emotion dataset, which has data collected from 820 subjects to simulate real-world distribution. In addition, we release some pre-trained models based on LSSED, which can not only promote the development of speech emotion recognition, but can also be transferred to related downstream tasks such as mental health analysis where data is extremely difficult to collect. Finally, our experiments show the necessity of large-scale datasets and the effectiveness of pre-trained models. The dateset will be released on https://github.com/tobefans/LSSED.

📄 PDF Abstract BibTeX arXiv:2102.01754

Code (1)

tobefans/LSSED 공식 구현 pytorch

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

LSSED: A Robust Segmentation Network for Inflamed Appendix from CT Images

2023-05-05 · ICASSP 2023 5 · Wing W Y. Ng, Peixin Zheng, Ting Wang, Jianjun Zhang 외

Acute appendicitis (AA) is one of the most prevalent surgical acute abdominal condition diseases. The treatment management of A A is highly dependent on the CT image diagnosis. However, the in-flamed appendix exhibits bl…

DecoderManagementSegmentation

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

2026-06-15 · Sejal Bhalla, Larry Kieu, Aina Merchant, Eyal de Lara 외 arxiv

Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems. Current clinical speech AI methods have largely progressed through isolated conditio…

SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning

2025-04-12 · Prabhat Pandey, Rupak Vignesh Swaminathan, K V Vijay Girish, Arunasish Sen 외

We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50M is built from publicly available speec…

Instruction Following

A High-Quality and Large-Scale Dataset for English-Vietnamese Speech Translation

2022-08-08 · Linh The Nguyen, Nguyen Luong Tran, Long Doan, Manh Luong 외

In this paper, we introduce a high-quality and large-scale benchmark dataset for English-Vietnamese speech translation with 508 audio hours, consisting of 331K triplets of (sentence-lengthed audio, English source transcr…

SentenceTranslation

SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

2026-08-01 · Ruidong Zhang, Jiacheng Liu, François Guimbretière, Cheng Zhang arxiv

Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies. Approaches achieving larger vocabularies require obtrusive hardware such as facial electrodes. We present SoniSpeech, the first large-sc…

Speech Recognition