Papers Speech-to-Text
“Speech-to-Text” 태그가 달린 논문 403편 · 필터 해제
An Empirical Evaluation of AI-Powered Non-Player Characters' Perceived Realism and Performance in Virtual Reality Environments
Advancements in artificial intelligence (AI) have significantly enhanced the realism and interactivity of non-player characters (NPCs) in virtual reality (VR), creating more engaging and believable user experiences. This…
Speech-to-Texttext-to-speechText to SpeechLM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approache…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+6End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
The scarcity of high-quality annotated data presents a significant challenge in developing effective end-to-end speech-to-text translation (ST) systems, particularly for low-resource languages. This paper explores the hy…
SentenceSpeech-to-TextSpeech-to-Text TranslationTranslationI Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs
Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often …
3D visual groundingContrastive LearningSpeech-to-TextVisual GroundingS2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamless Speech-Text Alignment and Streaming Speech Generation
Multilingual speech-to-speech translation (S2ST) aims to directly convert spoken utterances from multiple source languages into fluent and intelligible speech in a target language. Despite recent progress, several critic…
Reading ComprehensionSpeech SynthesisSpeech-to-Speech TranslationSpeech-to-Text+5Advancing STT for Low-Resource Real-World Speech
Swiss German is a low-resource language represented by diverse dialects that differ significantly from Standard German and from each other, lacking a standardized written form. As a result, transcribing Swiss German invo…
SentenceSpeech-to-TextSpeech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
We propose a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and zero-resource settings. By introducing ph…
Cross-Lingual TransferPhoneme RecognitionSpeech-to-TextSpeech-to-Text Translation+1Improving Language and Modality Transfer in Translation by Character-level Modeling
Current translation systems, despite being highly multilingual, cover only 5% of the world's languages. Expanding language coverage to the long-tail of low-resource languages requires data-efficient methods that rely on …
Speech-to-TextSpeech-to-Text TranslationTransfer LearningTranslationThe Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
Training large-scale models presents challenges not only in terms of resource requirements but also in terms of their convergence. For this reason, the learning rate (LR) is often decreased when the size of a model is in…
Speech-to-TextBeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System
This paper discusses the construction, fine-tuning, and deployment of BeaverTalk, a cascaded system for speech-to-text translation as part of the IWSLT 2025 simultaneous translation task. The system architecture employs …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+4Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced the naturalness and flexibility of human computer interaction by enabling seamless understanding across text, vision, and audio moda…
Adversarial AttackSpeech TokenizationSpeech-to-TextSpeech-to-Text TranslationConversational Recommendation System using NLP and Sentiment Analysis
In today's digitally-driven world, the demand for personalized and context-aware recommendations has never been greater. Traditional recommender systems have made significant strides in this direction, but they often lac…
Conversational RecommendationDynamic Time WarpingMarketingRecommendation Systems+2Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
In recent years, end-to-end speech-to-speech (S2S) dialogue systems have garnered increasing research attention due to their advantages over traditional cascaded systems, including achieving lower latency and more natura…
RAGRetrievalRetrieval-augmented GenerationSpeech-to-TextMEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED BENGALI-ENGLISH TRANSLATOR FOR CLINICAL APPLICATIONS
Code-switching in multilingual healthcare settings challenges automated transcription systems as facilities adopt AI documentation tools. To tackle this issue, we developed a cost-effective solution using the MediBeng Wh…
Clinical Language TranslationMachine TranslationSpeech-to-TextSpeech-to-Text Translation+1Acquisition of high-quality images for camera calibration in robotics applications via speech prompts
Accurate intrinsic and extrinsic camera calibration can be an important prerequisite for robotic applications that rely on vision as input. While there is ongoing research on enabling camera calibration using natural ima…
Camera CalibrationSpeech-to-TextTAGLinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we p…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+2Transformer-Based Named Entity Recognition for Automated Server Provisioning
This paper introduces a novel method for automated server provisioning by integrating Transformerbased Named Entity Recognition models with Automated Speech Detection using OpenAI's Whisper. Leveraging advanced Transform…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+3Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit
Although speech recognition algorithms have developed quickly in recent years, achieving high transcription accuracy across diverse audio formats and acoustic environments remains a major challenge. This work explores ho…
speech-recognitionSpeech RecognitionSpeech-to-TextAdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
In end-to-end speech translation, acoustic representations learned by the encoder are usually fixed and static, from the perspective of the decoder, which is not desirable for dealing with the cross-modal and cross-lingu…
DecoderSpeech-to-TextSpeech-to-Text TranslationTranslationFocusing Robot Open-Ended Reinforcement Learning Through Users' Purposes
Open-Ended Learning (OEL) autonomous robots can acquire new skills and knowledge through direct interaction with their environment, relying on mechanisms such as intrinsic motivations and self-generated goals to guide le…
Large Language Modelreinforcement-learningReinforcement LearningSpeech-to-Text