Data Augmentation for Improving Tail-traffic Robustness in Skill-routing for Dialogue Systems
Large-scale conversational systems typically rely on a skill-routing component to route a user request to an appropriate skill and interpretation to serve the request. In such system, the agent is responsible for serving thousands of skills and interpretations which create a long-tail distribution due to the natural frequency of requests. For example, the samples related to play music might be a thousand times more frequent than those asking for theatre show times. Moreover, inputs used for ML-based skill routing are often a heterogeneous mix of strings, embedding vectors, categorical and scalar features which makes employing augmentation-based long-tail learning approaches challenging. To improve the skill-routing robustness, we propose an augmentation of heterogeneous skill-routing data and training targeted for robust operation in long-tail data regimes. We explore a variety of conditional encoder-decoder generative frameworks to perturb original data fields and create synthetic training data. To demonstrate the effectiveness of the proposed method, we conduct extensive experiments using real-world data from a commercial conversational system. Based on the experiment results, the proposed approach improves more than 80% (51 out of 63) of intents with less than 10K of traffic instances in the skill-routing replication task.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationDecoderLong-tail LearningSimilar Papers 제목 키워드 기반
Enhancing Traffic Sign Recognition with Tailored Data Augmentation: Addressing Class Imbalance and Instance Scarcity
This paper tackles critical challenges in traffic sign recognition (TSR), which is essential for road safety -- specifically, class imbalance and instance scarcity in datasets. We introduce tailored data augmentation tec…
Data AugmentationImage GenerationTraffic Sign RecognitionNeural model robustness for skill routing in large-scale conversational AI systems: A design choice exploration
Current state-of-the-art large-scale conversational AI or intelligent digital assistant systems in industry comprises a set of components such as Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationNatural Language Understanding+2Hierarchical Classification of Transversal Skills in Job Ads Based on Sentence Embeddings
This paper proposes a classification framework aimed at identifying correlations between job ad requirements and transversal skill sets, with a focus on predicting the necessary skills for individual job descriptions usi…
ClassificationLanguage ModelingLanguage ModellingSentence+3Enhancing Encrypted Internet Traffic Classification Through Advanced Data Augmentation Techniques
The increasing popularity of online services has made Internet Traffic Classification a critical field of study. However, the rapid development of internet protocols and encryption limits usable data availability. This p…
Data AugmentationTraffic ClassificationSimple In-place Data Augmentation for Surveillance Object Detection
Motivated by the need to improve model performance in traffic monitoring tasks with limited labeled samples, we propose a straightforward augmentation technique tailored for object detection datasets, specifically design…
Data Augmentationobject-detectionObject Detection