CSVC-Net: Code-Switched Voice Command Classification using Deep CNN-LSTM Network
Colloquial Bengali has adopted many English words due to colonial influence. In conversational Bengali, it is quite common to speak in a mixture of English and Bengali, a phenomenon termed Code-switching (CS). To build a Voice Command Classifier in this era, when the usage of CS is ever-increasing, it is often necessary to map a single base command to its many different variants - spoken in multiple mixtures of languages. The works done with Bengali Speech have been primarily focused on single word classification and mostly incompetent in understanding the complex semantic relationships displayed in sentences. This paper proposes ‘CSVC-Net’, a CNN-LSTM based architecture for classifying spoken commands that exhibit code-switching between Bengali and English. To effectively reflect the scenario, it also presents a newly curated dataset named ‘Banglish’ containing 3,840 audio files of spoken computer commands belonging to 11 classes, considering 64 variations in total. The proposed pipeline passes the input audio signal through a series of appropriate transformation and augmentation steps enabling the model to achieve an accuracy of 92.08% on the curated dataset. Furthermore, the robustness of the proposed model has been justified by comparing with different architectures and tested under different noise levels with promising accuracy, which shows the applicability of the model in real-life scenarios.
Code (1)
Tasks
Voice Query RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Robust Sensor Fusion Algorithms Against Voice Command Attacks in Autonomous Vehicles
With recent advances in autonomous driving, Voice Control Systems have become increasingly adopted as human-vehicle interaction methods. This technology enables drivers to use voice commands to control the vehicle and wi…
Autonomous DrivingAutonomous VehiclesMultimodal Deep LearningSensor FusionImprove Cross-lingual Voice Cloning Using Low-quality Code-switched Data
Recently, sequence-to-sequence (seq-to-seq) models have been successfully applied in text-to-speech (TTS) to synthesize speech for single-language text. To synthesize speech for multiple languages usually requires multi-…
text-to-speechText to SpeechVoice CloningTowards Natural Bilingual and Code-Switched Speech Synthesis Based on Mix of Monolingual Recordings and Cross-Lingual Voice Conversion
Recent state-of-the-art neural text-to-speech (TTS) synthesis models have dramatically improved intelligibility and naturalness of generated speech from text. However, building a good bilingual or code-switched TTS for a…
Speech Synthesistext-to-speechText to SpeechVoice ConversionAnalyzing Multimodal Features of Spontaneous Voice Assistant Commands for Mild Cognitive Impairment Detection
Mild cognitive impairment (MCI) is a major public health concern due to its high risk of progressing to dementia. This study investigates the potential of detecting MCI with spontaneous voice assistant (VA) commands from…
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4 test sets which cover in total 113 uniqu…
Speech Recognition