Increasing the Accessibility of Time-Aligned Speech Corpora with Spokes Mix
Code (0)
등록된 구현이 없습니다.
Tasks
Speech RecognitionSimilar Papers 제목 키워드 기반
Speaking at the Right Level: Literacy-Controlled Counterspeech Generation with RAG-RL
Health misinformation spreading online poses a significant threat to public health. Researchers have explored methods for automatically generating counterspeech to health misinformation as a mitigation strategy. Existing…
Reinforcement LearningWho Gets Left Behind? Auditing Disability Inclusivity in Large Language Models
Large Language Models (LLMs) are increasingly used for accessibility guidance, yet many disability groups remain underserved by their advice. To address this gap, we present taxonomy aligned benchmark1 of human validated…
TVD: A Reproducible and Multiply Aligned TV Series Dataset
We introduce a new dataset built around two TV series from different genres, The Big Bang Theory, a situation comedy and Game of Thrones, a fantasy drama. The dataset has multiple tracks extracted from diverse sources, i…
Dynamic Time WarpingInformation RetrievalRetrievalSentiment Analysis+1Exploring Generative Error Correction for Dysarthric Speech Recognition
Despite the remarkable progress in end-to-end Automatic Speech Recognition (ASR) engines, accurately transcribing dysarthric speech remains a major challenge. In this work, we proposed a two-stage framework for the Speec…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionFacial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
We propose FEIM-TTS, an innovative zero-shot text-to-speech (TTS) model that synthesizes emotionally expressive speech, aligned with facial images and modulated by emotion intensity. Leveraging deep learning, FEIM-TTS tr…
Emotional Speech SynthesisSpeech Synthesistext-to-speechText to Speech