paper-with-me

홈 › Papers

On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models

2024-06-13 · Jinchuan Tian, Yifan Peng, William Chen, Kwanghee Choi, Karen Livescu, Shinji Watanabe

The Open Whisper-style Speech Model (OWSM) series was introduced to achieve full transparency in building advanced speech-to-text (S2T) foundation models. To this end, OWSM models are trained on 25 public speech datasets, which are heterogeneous in multiple ways. In this study, we advance the OWSM series by introducing OWSM v3.2, which improves on prior models by investigating and addressing the impacts of this data heterogeneity. Our study begins with a detailed analysis of each dataset, from which we derive two key strategies: data filtering with proxy task to enhance data quality, and the incorporation of punctuation and true-casing using an open large language model (LLM). With all other configurations staying the same, OWSM v3.2 improves performance over the OWSM v3.1 baseline while using 15% less training data.

📄 PDF Abstract BibTeX arXiv:2406.09282

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelSpeech-to-Text

Similar Papers 제목 키워드 기반

An Adaptive Kernel Approach to Federated Learning of Heterogeneous Causal Effects

2023-01-01 · Thanh Vinh Vo, Arnab Bhattacharyya, Young Lee, Tze-Yun Leong

We propose a new causal inference framework to learn causal effects from multiple, decentralized data sources in a federated setting. We introduce an adaptive transfer algorithm that learns the similarities among the dat…

Causal InferenceFederated Learning

Combining observational and experimental data to find heterogeneous treatment effects

2016-11-08 · Alexander Peysakhovich, Akos Lada

Every design choice will have different effects on different units. However traditional A/B tests are often underpowered to identify these heterogeneous effects. This is especially true when the set of unit-level attribu…

Time SeriesTime Series Analysis

A Tree-based Model Averaging Approach for Personalized Treatment Effect Estimation from Heterogeneous Data Sources

2021-03-10 · Xiaoqing Tan, Chung-Chou H. Chang, Ling Zhou, Lu Tang

Accurately estimating personalized treatment effects within a study site (e.g., a hospital) has been challenging due to limited sample size. Furthermore, privacy considerations and lack of resources prevent a site from l…

Federated Learning

Topic Stability over Noisy Sources

2015-08-05 · WS 2016 12 · Jing Su, Oisín Boydell, Derek Greene, Gerard Lynch

Topic modelling techniques such as LDA have recently been applied to speech transcripts and OCR output. These corpora may contain noisy or erroneous texts which may undermine topic stability. Therefore, it is important t…

Model SelectionOptical Character Recognition (OCR)Topic Models

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

2026-04-24 · Chunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang 외 arxiv

Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous control paradigms. Unifying these modal…