paper-with-me

Papers

A Benchmarking on Cloud based Speech-To-Text Services for French Speech and Background Noise Effect

2021-05-07 · Binbin Xu, Chongyang Tao, Zidu Feng, Youssef Raqui, Sylvie Ranwez

This study presents a large scale benchmarking on cloud based Speech-To-Text systems: {Google Cloud Speech-To-Text}, {Microsoft Azure Cognitive Services}, {Amazon Transcribe}, {IBM Watson Speech to Text}. For each systems, 40158 clean and noisy speech files about 101 hours are tested. Effect of background noise on STT quality is also evaluated with 5 different Signal-to-noise ratios from 40dB to 0dB. Results showed that {Microsoft Azure} provided lowest transcription error rate $9.09\%$ on clean speech, with high robustness to noisy environment. {Google Cloud} and {Amazon Transcribe} gave similar performance, but the latter is very limited for time-constraint usage. Though {IBM Watson} could work correctly in quiet conditions, it is highly sensible to noisy speech which could strongly limit its application in real life situations.

📄 PDF Abstract BibTeX arXiv:2105.03409

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingSpeech-to-Text

Similar Papers 제목 키워드 기반

Scaling up to the cloud: Cloud technology use and growth rates in small and large firms

2024-09-25 · Bernardo Caldarola, Luca Fontanelli

Recent empirical evidence shows that investments in ICT disproportionately improve the performance of larger firms versus smaller ones. However, ICT may not be all alike, as they differ in their impact on firms' organisa…

Towards a Language Service Infrastructure for Mobile Environments

2016-05-01 · LREC 2016 5 · Ngoc Nguyen, Donghui Lin, Takao Nakaguchi, Toru Ishida

Since mobile devices have feature-rich configurations and provide diverse functions, the use of mobile devices combined with the language resources of cloud environments is high promising for achieving a wide range commu…

text-to-speechText to SpeechTranslation

Emotion Filtering at the Edge

2019-09-18

Voice controlled devices and services have become very popular in the consumer IoT. Cloud-based speech analysis services extract information from voice inputs using speech recognition techniques. Services providers can t…

Privacy PreservingRaspberry Pi 4speech-recognitionSpeech Recognition

AI-Driven Modular Services for Accessible Multilingual Education in Immersive Extended Reality Settings: Integrating Speech Processing, Translation, and Sign Language Rendering

2026-04-07 · N. D. Tantaroudas, A. J. McCracken, I. Karachalios, E. Papatheou arxiv

This work introduces a modular platform that brings together six AI services, automatic speech recognition via OpenAI Whisper, multilingual translation through Meta NLLB, speech synthesis using AWS Polly, emotion classif…

Emotion ClassificationSpeech RecognitionSpeech Synthesis

Pantagruel: Unified Self-Supervised Encoders for French Text and Speech

2026-01-09 · Phuong-Hang Le, Valentin Pelloin, Arnault Chatelain, Maryem Bouziane 외 arxiv

We release Pantagruel models, a new family of self-supervised encoder models for French text and speech. Instead of predicting modality-tailored targets such as textual tokens or speech units, Pantagruel learns contextua…

Representation Learning