paper-with-me

Papers

ASTRA: Aligning Speech and Text Representations for Asr without Sampling

2024-06-10 · Neeraj Gaur, Rohan Agrawal, Gary Wang, Parisa Haghani, Andrew Rosenberg, Bhuvana Ramabhadran

This paper introduces ASTRA, a novel method for improving Automatic Speech Recognition (ASR) through text injection.Unlike prevailing techniques, ASTRA eliminates the need for sampling to match sequence lengths between speech and text modalities. Instead, it leverages the inherent alignments learned within CTC/RNNT models. This approach offers the following two advantages, namely, avoiding potential misalignment between speech and text features that could arise from upsampling and eliminating the need for models to accurately predict duration of sub-word tokens. This novel formulation of modality (length) matching as a weighted RNNT objective matches the performance of the state-of-the-art duration-based methods on the FLEURS benchmark, while opening up other avenues of research in speech processing.

📄 PDF Abstract BibTeX arXiv:2406.06664

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment

2026-08-03 · Junyu Zhu, Hao Zhu, Xinzhuo Zhang, Xu Zhang 외 arxiv

Dynamic 3D scene reconstruction has made significant progress with multi-camera systems, often relying on temporally aligned observations across views. However, in real-world scenarios, temporal asynchrony among capturin…

Unified Multi-Foundation-Model Slide Representation for Pan-Cancer Recognition and Text-Guided Tumor Localization

2026-04-21 · Tianyang Wang, Ziyu Su, Abdul Rehman Akbar, Usama Sajjad 외 arxiv

The expanding ecosystem of pathology foundation models has produced powerful but fragmented tile-level representations, limiting their use in clinical tasks that require unified slide-level reasoning and interpretable li…

Representation LearningCancer Classification

ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots

2026-06-16 · Ethan Chew, Enjia Wu, Iruss Eng, Ian Lim 외 arxiv

Air Traffic Control Operators (ATCOs) are vital in ensuring the safe, orderly, and efficient flow of air traffic, yet training capacity is constrained by reliance on specialized human trainers known as simpilots, who mus…

Speech Recognition

The C-ORAL-BRASIL I: Reference Corpus for Spoken Brazilian Portuguese

2012-05-01 · LREC 2012 5 · Tommaso Raso, Heliana Mello, Maryual{\^e} Malvessi Mittmann

C-ORAL-BRASIL I is a Brazilian Portuguese spontaneous speech corpus compiled following the same architecture adopted by the C-ORAL-ROM resource. The main goal is the documentation of the diaphasic and diastratic variatio…

text-to-speechText to Speech

Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation

2025-03-13 · Henglyu Liu, Andong Chen, Kehai Chen, Xuefeng Bai 외

Recent advancement of large language models (LLMs) has led to significant breakthroughs across various tasks, laying the foundation for the development of LLM-based speech translation systems. Existing methods primarily …

Cross-Modal RetrievalTranslation