paper-with-me

Papers Speech-to-Text

“Speech-to-Text” 태그가 달린 논문 403편 · 필터 해제

An Empirical Evaluation of AI-Powered Non-Player Characters' Perceived Realism and Performance in Virtual Reality Environments

2025-07-14 · Mikko Korkiakoski, Saeid Sheikhi, Jesper Nyman, Jussi Saariniemi 외

Advancements in artificial intelligence (AI) have significantly enhanced the realism and interactivity of non-player characters (NPCs) in virtual reality (VR), creating more engaging and believable user experiences. This…

Speech-to-Texttext-to-speechText to Speech

LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization

2025-06-20 · DaeJin Jo, Jeeyoung Yun, Byungseok Roh, Sungwoong Kim

With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approache…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+6

End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data

2025-06-19 · Aishwarya Pothula, Bhavana Akkiraju, Srihari Bandarupalli, Charan D 외

The scarcity of high-quality annotated data presents a significant challenge in developing effective end-to-end speech-to-text translation (ST) systems, particularly for low-resource languages. This paper explores the hy…

SentenceSpeech-to-TextSpeech-to-Text TranslationTranslation

I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs

2025-06-17 · Yu Qi, Lipeng Gu, Honghua Chen, Liangliang Nan 외

Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often …

3D visual groundingContrastive LearningSpeech-to-TextVisual Grounding

S2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamless Speech-Text Alignment and Streaming Speech Generation

2025-06-11 · Yu Pan, Yuguang Yang, Yanni Hu, Jianhao Ye 외

Multilingual speech-to-speech translation (S2ST) aims to directly convert spoken utterances from multiple source languages into fluent and intelligible speech in a target language. Despite recent progress, several critic…

Reading ComprehensionSpeech SynthesisSpeech-to-Speech TranslationSpeech-to-Text+5

Advancing STT for Low-Resource Real-World Speech

2025-06-10 · Flavio D'Intino, Hans-Peter Hutter

Swiss German is a low-resource language represented by diverse dialects that differ significantly from Standard German and from each other, lacking a standardized written form. As a result, transcribing Swiss German invo…

SentenceSpeech-to-Text

Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios

2025-05-30 · Gerard I. Gállego, Oriol Pareras, Martí Cortada Garcia, Lucas Takanori 외

We propose a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and zero-resource settings. By introducing ph…

Cross-Lingual TransferPhoneme RecognitionSpeech-to-TextSpeech-to-Text Translation+1

Improving Language and Modality Transfer in Translation by Character-level Modeling

2025-05-30 · Ioannis Tsiamas, David Dale, Marta R. Costa-jussà

Current translation systems, despite being highly multilingual, cover only 5% of the world's languages. Expanding language coverage to the long-tail of low-resource languages requires data-efficient methods that rely on …

Speech-to-TextSpeech-to-Text TranslationTransfer LearningTranslation

The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence

2025-05-29 · Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti 외

Training large-scale models presents challenges not only in terms of resource requirements but also in terms of their convergence. For this reason, the learning rate (LR) is often decreased when the size of a model is in…

Speech-to-Text

BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System

2025-05-29 · Matthew Raffel, Victor Agostinelli, Lizhong Chen

This paper discusses the construction, fine-tuning, and deployment of BeaverTalk, a cascaded system for speech-to-text translation as part of the IWSLT 2025 simultaneous translation task. The system architecture employs …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+4

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

2025-05-24 · Binhao Ma, Hanqing Guo, Zhengping Jay Luo, Rui Duan

Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced the naturalness and flexibility of human computer interaction by enabling seamless understanding across text, vision, and audio moda…

Adversarial AttackSpeech TokenizationSpeech-to-TextSpeech-to-Text Translation

Conversational Recommendation System using NLP and Sentiment Analysis

2025-05-17 · Piyush Talegaonkar, Siddhant Hole, Shrinesh Kamble, Prashil Gulechha 외

In today's digitally-driven world, the demand for personalized and context-aware recommendations has never been greater. Traditional recommender systems have made significant strides in this direction, but they often lac…

Conversational RecommendationDynamic Time WarpingMarketingRecommendation Systems+2

Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation

2025-04-27 · Pengchao Feng, Ziyang Ma, Wenxi Chen, Yao Li 외

In recent years, end-to-end speech-to-speech (S2S) dialogue systems have garnered increasing research attention due to their advantages over traditional cascaded systems, including achieving lower latency and more natura…

RAGRetrievalRetrieval-augmented GenerationSpeech-to-Text

MEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED BENGALI-ENGLISH TRANSLATOR FOR CLINICAL APPLICATIONS

2025-04-25 · medRxiv 2025 4 · Promila Ghosh, Sunipun Talukder

Code-switching in multilingual healthcare settings challenges automated transcription systems as facilities adopt AI documentation tools. To tackle this issue, we developed a cost-effective solution using the MediBeng Wh…

Clinical Language TranslationMachine TranslationSpeech-to-TextSpeech-to-Text Translation+1

Acquisition of high-quality images for camera calibration in robotics applications via speech prompts

2025-04-15 · Timm Linder, Kadir Yilmaz, David B. Adrian, Bastian Leibe

Accurate intrinsic and extrinsic camera calibration can be an important prerequisite for robotic applications that rely on vision as input. While there is ongoing research on enabling camera calibration using natural ima…

Camera CalibrationSpeech-to-TextTAG

LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect

2025-04-03 · Hedi Naouara, Jean-Pierre Lorré, Jérôme Louradour

Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+2

Transformer-Based Named Entity Recognition for Automated Server Provisioning

2025-04-01 · Conference 2025 4 · Hossein Damavandi, Hasan Jalali, Boshra Pishgoo

This paper introduces a novel method for automated server provisioning by integrating Transformerbased Named Entity Recognition models with Automated Speech Detection using OpenAI's Whisper. Leveraging advanced Transform…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+3

Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit

2025-03-26 · Aniket Abhishek Soni

Although speech recognition algorithms have developed quickly in recent years, achieving high transcription accuracy across diverse audio formats and acoustic environments remains a major challenge. This work explores ho…

speech-recognitionSpeech RecognitionSpeech-to-Text

AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation

2025-03-18 · Findings (ACL) 2021 8 · Wuwei Huang, Dexin Wang, Deyi Xiong

In end-to-end speech translation, acoustic representations learned by the encoder are usually fixed and static, from the perspective of the decoder, which is not desirable for dealing with the cross-modal and cross-lingu…

DecoderSpeech-to-TextSpeech-to-Text TranslationTranslation

Focusing Robot Open-Ended Reinforcement Learning Through Users' Purposes

2025-03-16 · Emilio Cartoni, Gianluca Cioccolini, Gianluca Baldassarre

Open-Ended Learning (OEL) autonomous robots can acquire new skills and knowledge through direct interaction with their environment, relying on mechanisms such as intrinsic motivations and self-generated goals to guide le…

Large Language Modelreinforcement-learningReinforcement LearningSpeech-to-Text
1–20 / 403 다음 →