paper-with-me

홈 › Papers

SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors

2025-03-20 · Yang Chen, Hui Wang, Shiyao Wang, Junyang Chen, Jiabei He, Jiaming Zhou, Xi Yang, Yequan Wang, Yonghua Lin, Yong Qin

While voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal characteristics like presbyphonia and dialectal variations. The limited data available on super-aged individuals in existing elderly speech datasets, coupled with overly simple recording styles and annotation dimensions, exacerbates this issue. To address the critical scarcity of speech data from individuals aged 75 and above, we introduce SeniorTalk, a carefully annotated Chinese spoken dialogue dataset. This dataset contains 55.53 hours of speech from 101 natural conversations involving 202 participants, ensuring a strategic balance across gender, region, and age. Through detailed annotation across multiple dimensions, it can support a wide range of speech tasks. We perform extensive experiments on speaker verification, speaker diarization, speech recognition, and speech editing tasks, offering crucial insights for the development of speech technologies targeting this age group.

📄 PDF Abstract BibTeX arXiv:2503.16578

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker DiarizationSpeaker Verificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Zero-shot Cross-lingual Conversational Semantic Role Labeling

2022-04-11 · Findings (NAACL) 2022 7 · Han Wu, Haochen Tan, Kun Xu, Shuqi Liu 외

While conversational semantic role labeling (CSRL) has shown its usefulness on Chinese conversational tasks, it is still under-explored in non-Chinese languages due to the lack of multilingual CSRL annotations for the pa…

Response GenerationSemantic Role Labeling

Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset

2022-03-31 · Zehui Yang, Yifan Chen, Lei Luo, Runyan Yang 외

This paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC. The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native spe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeaker-diarization+5

Zero-shot Cross-lingual Conversational Semantic Role Labeling

2021-11-16 · ACL ARR November 2021 11 · Anonymous

While conversational semantic role labeling (CSRL) has shown its usefulness on Chinese conversational tasks, it is still under-explored in non-Chinese languages due to the lack of multilingual CSRL annotations for the pa…

Response GenerationSemantic Role LabelingTranslation

ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation

2021-12-12 · LREC 2022 6 · Holy Lovenia, Samuel Cahyawijaya, Genta Indra Winata, Peng Xu 외

Code-switching is a speech phenomenon occurring when a speaker switches language during a conversation. Despite the spontaneous nature of code-switching in conversational spoken language, most existing works collect code…

Towards Better Understanding of User Satisfaction in Open-Domain Conversational Search

2022-04-06 · Zhumin Chu, Qingyao Ai, Zhihong Wang, Yiqun Liu 외

With the increasing popularity of conversational search, how to evaluate the performance of conversational search systems has become an important question in the IR community. Existing works on conversational search eval…

Conversational SearchSemantic SimilaritySemantic Textual Similarity