paper-with-me

Papers

Benchmarking Language-agnostic Intent Classification for Virtual Assistant Platforms

2022-07-01 · NAACL (MIA) 2022 7 · Gengyu Wang, Cheng Qian, Lin Pan, Haode Qi, Ladislav Kunc, Saloni Potdar

Current virtual assistant (VA) platforms are beholden to the limited number of languages they support. Every component, such as the tokenizer and intent classifier, is engineered for specific languages in these intricate platforms. Thus, supporting a new language in such platforms is a resource-intensive operation requiring expensive re-training and re-designing. In this paper, we propose a benchmark for evaluating language-agnostic intent classification, the most critical component of VA platforms. To ensure the benchmarking is challenging and comprehensive, we include 29 public and internal datasets across 10 low-resource languages and evaluate various training and testing settings with consideration of both accuracy and training time. The benchmarking result shows that Watson Assistant, among 7 commercial VA platforms and pre-trained multilingual language models (LMs), demonstrates close-to-best accuracy with the best accuracy-training time trade-off.

📄 PDF Abstract BibTeX

Code (1)

posuer/benchmark-multilingual-intent-classification 공식 구현

Tasks

BenchmarkingClassificationintent-classificationIntent Classification

Similar Papers 제목 키워드 기반

MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

2022-04-18 · Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie 외

We present the MASSIVE dataset--Multilingual Amazon Slu resource package (SLURP) for Slot-filling, Intent classification, and Virtual assistant Evaluation. MASSIVE contains 1M realistic, parallel, labeled virtual assista…

intent-classificationIntent ClassificationNatural Language UnderstandingSlot Filling+3

Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models

2026-01-11 · Shaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng 외 arxiv

Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challen…

Meta learning to classify intent and slot labels with noisy few shot examples

2020-11-30 · Shang-Wen Li, Jason Krone, Shuyan Dong, Yi Zhang 외

Recently deep learning has dominated many machine learning areas, including spoken language understanding (SLU). However, deep learning models are notorious for being data-hungry, and the heavily optimized models are usu…

Benchmarkingintent-classificationIntent ClassificationMeta-Learning+1

MUPAX: Multidimensional Problem Agnostic eXplainable AI

2025-07-17 · Vincenzo Dentamaro, Felice Franchini, Giuseppe Pirlo, Irina Voiculescu

Robust XAI techniques should ideally be simultaneously deterministic, model agnostic, and guaranteed to converge. We propose MULTIDIMENSIONAL PROBLEM AGNOSTIC EXPLAINABLE AI (MUPAX), a deterministic, model agnostic expla…

Anatomical Landmark DetectionAudio ClassificationBenchmarkingFeature Importance+3

MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection

2025-10-27 · Anisha Saha, Varsha Suresh, Timothy Hospedales, Vera Demberg arxiv

Sarcasm is a specific type of irony which involves discerning what is said from what is meant. Detecting sarcasm depends not only on the literal content of an utterance but also on non-verbal cues such as speaker's tonal…

Sarcasm Detection