A Benchmarking on Cloud based Speech-To-Text Services for French Speech and Background Noise Effect
This study presents a large scale benchmarking on cloud based Speech-To-Text systems: {Google Cloud Speech-To-Text}, {Microsoft Azure Cognitive Services}, {Amazon Transcribe}, {IBM Watson Speech to Text}. For each systems, 40158 clean and noisy speech files about 101 hours are tested. Effect of background noise on STT quality is also evaluated with 5 different Signal-to-noise ratios from 40dB to 0dB. Results showed that {Microsoft Azure} provided lowest transcription error rate $9.09\%$ on clean speech, with high robustness to noisy environment. {Google Cloud} and {Amazon Transcribe} gave similar performance, but the latter is very limited for time-constraint usage. Though {IBM Watson} could work correctly in quiet conditions, it is highly sensible to noisy speech which could strongly limit its application in real life situations.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingSpeech-to-TextSimilar Papers 제목 키워드 기반
Scaling up to the cloud: Cloud technology use and growth rates in small and large firms
Recent empirical evidence shows that investments in ICT disproportionately improve the performance of larger firms versus smaller ones. However, ICT may not be all alike, as they differ in their impact on firms' organisa…
Towards a Language Service Infrastructure for Mobile Environments
Since mobile devices have feature-rich configurations and provide diverse functions, the use of mobile devices combined with the language resources of cloud environments is high promising for achieving a wide range commu…
text-to-speechText to SpeechTranslationEmotion Filtering at the Edge
Voice controlled devices and services have become very popular in the consumer IoT. Cloud-based speech analysis services extract information from voice inputs using speech recognition techniques. Services providers can t…
Privacy PreservingRaspberry Pi 4speech-recognitionSpeech RecognitionAI-Driven Modular Services for Accessible Multilingual Education in Immersive Extended Reality Settings: Integrating Speech Processing, Translation, and Sign Language Rendering
This work introduces a modular platform that brings together six AI services, automatic speech recognition via OpenAI Whisper, multilingual translation through Meta NLLB, speech synthesis using AWS Polly, emotion classif…
Emotion ClassificationSpeech RecognitionSpeech SynthesisPantagruel: Unified Self-Supervised Encoders for French Text and Speech
We release Pantagruel models, a new family of self-supervised encoder models for French text and speech. Instead of predicting modality-tailored targets such as textual tokens or speech units, Pantagruel learns contextua…
Representation Learning