Incremental processing of noisy user utterances in the spoken language understanding task
The state-of-the-art neural network architectures make it possible to create spoken language understanding systems with high quality and fast processing time. One major challenge for real-world applications is the high latency of these systems caused by triggered actions with high executions times. If an action can be separated into subactions, the reaction time of the systems can be improved through incremental processing of the user utterance and starting subactions while the utterance is still being uttered. In this work, we present a model-agnostic method to achieve high quality in processing incrementally produced partial utterances. Based on clean and noisy versions of the ATIS dataset, we show how to create datasets with our method to create low-latency natural language understanding components. We get improvements of up to 47.91 absolute percentage points in the metric F1-score.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language UnderstandingSpoken Language UnderstandingSimilar Papers 제목 키워드 기반
Natural Language Input for In-Car Spoken Dialog Systems: How Natural is Natural?
Recent spoken dialog systems are moving away from command and control towards a more intuitive and natural style of interaction. In order to choose an appropriate system design which allows the system to deal with natura…
Incremental Disfluency Detection for Spoken Learner English
Incremental disfluency detection provides a framework for computing communicative meaning from hesitations, repetitions and false starts commonly found in speech. One application of this area of research is in dialogue-b…
Refer-iTTS: A System for Referring in Spoken Installments to Objects in Real-World Images
Current referring expression generation systems mostly deliver their output as one-shot, written expressions. We present on-going work on incremental generation of spoken expressions referring to objects in real-world im…
Referring ExpressionReferring expression generationSpeech SynthesisText Generation+3Evaluating Models of Robust Word Recognition with Serial Reproduction
Spoken communication occurs in a "noisy channel" characterized by high levels of environmental noise, variability within and between speakers, and lexical and syntactic ambiguity. Given these properties of the received l…
Improving User Impression in Spoken Dialog System with Gradual Speech Form Control
This paper examines a method to improve the user impression of a spoken dialog system by introducing a mechanism that gradually changes form of utterances every time the user uses the system. In some languages, including…
Form