paper-with-me

Papers

Survey on Deep Neural Networks in Speech and Vision Systems

2019-08-16 · Mahbubul Alam, Manar D. Samad, Lasitha Vidyaratne, Alexander Glandon, Khan M. Iftekharuddin

This survey presents a review of state-of-the-art deep neural network architectures, algorithms, and systems in vision and speech applications. Recent advances in deep artificial neural network algorithms and architectures have spurred rapid innovation and development of intelligent vision and speech systems. With availability of vast amounts of sensor data and cloud computing for processing and training of deep neural networks, and with increased sophistication in mobile and embedded technology, the next-generation intelligent systems are poised to revolutionize personal and commercial computing. This survey begins by providing background and evolution of some of the most successful deep learning models for intelligent vision and speech systems to date. An overview of large-scale industrial research and development efforts is provided to emphasize future trends and prospects of intelligent vision and speech systems. Robust and efficient intelligent systems demand low-latency and high fidelity in resource-constrained hardware platforms such as mobile devices, robots, and automobiles. Therefore, this survey also provides a summary of key challenges and recent successes in running deep neural networks on hardware-restricted platforms, i.e. within limited memory, battery life, and processing capabilities. Finally, emerging applications of vision and speech across disciplines such as affective computing, intelligent transportation, and precision medicine are discussed. To our knowledge, this paper provides one of the most comprehensive surveys on the latest developments in intelligent vision and speech applications from the perspectives of both software and hardware systems. Many of these emerging technologies using deep neural networks show tremendous promise to revolutionize research and development for future vision and speech systems.

📄 PDF Abstract BibTeX arXiv:1908.07656

Code (0)

등록된 구현이 없습니다.

Tasks

Cloud ComputingSurvey

Similar Papers 제목 키워드 기반

Accented Speech Recognition: A Survey

2021-04-21 · Arthur Hinsvark, Natalie Delworth, Miguel Del Rio, Quinten McNamara 외

Automatic Speech Recognition (ASR) systems generalize poorly on accented speech. The phonetic and linguistic variability of accents present hard challenges for ASR systems today in both data collection and modeling strat…

Accented Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Feature Engineering+3

WavChat: A Survey of Spoken Dialogue Models

2024-11-15 · Shengpeng Ji, Yifu Chen, Minghui Fang, Jialong Zuo 외

Recent advancements in spoken dialogue models, exemplified by systems like GPT-4o, have captured significant attention in the speech domain. Compared to traditional three-tier cascaded spoken dialogue models that compris…

speech-recognitionSpeech RecognitionSpoken Dialogue SystemsSurvey+2

Thank you for Attention: A survey on Attention-based Artificial Neural Networks for Automatic Speech Recognition

2021-02-14 · Priyabrata Karmakar, Shyh Wei Teng, Guojun Lu

Attention is a very popular and effective mechanism in artificial neural network-based sequence-to-sequence models. In this survey paper, a comprehensive review of the different attention models used in developing automa…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

A Survey of Code-switched Speech and Language Processing

2019-03-25 · Sunayana Sitaram, Khyathi Raghavi Chandu, Sai Krishna Rallabandi, Alan W. black

Code-switching, the alternation of languages within a conversation or utterance, is a common communicative phenomenon that occurs in multilingual communities across the world. This survey reviews computational approaches…

Survey

Deep Neural Networks for Automatic Speech Processing: A Survey from Large Corpora to Limited Data

2020-03-09 · Vincent Roger, Jérôme Farinas, Julien Pinquier

Most state-of-the-art speech systems are using Deep Neural Networks (DNNs). Those systems require a large amount of data to be learned. Hence, learning state-of-the-art frameworks on under-resourced speech languages/prob…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeaker Identification+2