paper-with-me

홈 › Papers

TRESTLE: Toolkit for Reproducible Execution of Speech, Text and Language Experiments

2023-02-14 · Changye Li, Weizhe Xu, Trevor Cohen, Martin Michalowski, Serguei Pakhomov

The evidence is growing that machine and deep learning methods can learn the subtle differences between the language produced by people with various forms of cognitive impairment such as dementia and cognitively healthy individuals. Valuable public data repositories such as TalkBank have made it possible for researchers in the computational community to join forces and learn from each other to make significant advances in this area. However, due to variability in approaches and data selection strategies used by various researchers, results obtained by different groups have been difficult to compare directly. In this paper, we present TRESTLE (\textbf{T}oolkit for \textbf{R}eproducible \textbf{E}xecution of \textbf{S}peech \textbf{T}ext and \textbf{L}anguage \textbf{E}xperiments), an open source platform that focuses on two datasets from the TalkBank repository with dementia detection as an illustrative domain. Successfully deployed in the hackallenge (Hackathon/Challenge) of the International Workshop on Health Intelligence at AAAI 2022, TRESTLE provides a precise digital blueprint of the data pre-processing and selection strategies that can be reused via TRESTLE by other researchers seeking comparable results with their peers and current state-of-the-art (SOTA) approaches.

📄 PDF Abstract BibTeX arXiv:2302.07322

Code (1)

LinguisticAnomalies/harmonized-toolkit 공식 구현

Similar Papers 제목 키워드 기반

FunCodec: A Fundamental, Reproducible and Integrable Open-source Toolkit for Neural Speech Codec

2023-09-14 · Zhihao Du, Shiliang Zhang, Kai Hu, Siqi Zheng

This paper presents FunCodec, a fundamental neural speech codec toolkit, which is an extension of the open-source speech processing toolkit FunASR. FunCodec provides reproducible training recipes and inference scripts fo…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech Synthesis+3

ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit

2019-10-24 · Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura 외

This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit supports state-of-the-art E2E-TTS models, i…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

ESPnet-ST: All-in-One Speech Translation Toolkit

2020-04-21 · ACL 2020 6 · Hirofumi Inaguma, Shun Kiyono, Kevin Duh, Shigeki Karita 외

We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework. ESPnet-ST is a new project inside end-to-end speech processing toolkit, ESPnet, which integ…

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translation+6

ESPnet-SpeechLM: An Open Speech Language Model Toolkit

2025-02-21 · Jinchuan Tian, Jiatong Shi, William Chen, Siddhant Arora 외

We present ESPnet-SpeechLM, an open toolkit designed to democratize the development of speech language models (SpeechLMs) and voice-driven agentic applications. The toolkit standardizes speech processing tasks by framing…

Language ModelingLanguage Modelling

ESPnet-SLU: Advancing Spoken Language Understanding through ESPnet

2021-11-29 · Siddhant Arora, Siddharth Dalmia, Pavel Denisov, Xuankai Chang 외

As Automatic Speech Processing (ASR) systems are getting better, there is an increasing interest of using the ASR output to do downstream Natural Language Processing (NLP) tasks. However, there are few open source toolki…

Spoken Language Understandingtext-to-speechText to Speech