paper-with-me

Papers

ITALIC: An Italian Intent Classification Dataset

2023-06-14 · Alkis Koudounas, Moreno La Quatra, Lorenzo Vaiani, Luca Colomba, Giuseppe Attanasio, Eliana Pastor, Luca Cagliero, Elena Baralis

Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, the first large-scale speech dataset designed for intent classification in Italian. The dataset comprises 16,521 crowdsourced audio samples recorded by 70 speakers from various Italian regions and annotated with intent labels and additional metadata. We explore the versatility of ITALIC by evaluating current state-of-the-art speech and text models. Results on intent classification suggest that increasing scale and running language adaptation yield better speech models, monolingual text models outscore multilingual ones, and that speech recognition on ITALIC is more challenging than on existing Italian benchmarks. We release both the dataset and the annotation scheme to streamline the development of new Italian SLU models and language-specific datasets.

📄 PDF Abstract BibTeX arXiv:2306.08502

Code (1)

rita-nlp/italic 공식 구현 pytorch

Tasks

Classificationintent-classificationIntent Classificationspeech-recognitionSpeech RecognitionSpoken Language Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs

2026-05-08 · Andrea Sassella, Andrea Chizzola, Tommaso Bianchi, Luca Alessandrelli 외 arxiv

This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters. Performance is investigated across a…

PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin

2024-04-25 · Stephen Bothwell, Brian DuSell, David Chiang, Brian Krostenko

Computational historical linguistics seeks to systematically understand processes of sound change, including during periods at which little to no formal recording of language is attested. At the same time, few computatio…

Multi-lingual Intent Detection and Slot Filling in a Joint BERT-based Model

2019-07-05 · Giuseppe Castellucci, Valentina Bellomaria, Andrea Favalli, Raniero Romagnoli

Intent Detection and Slot Filling are two pillar tasks in Spoken Natural Language Understanding. Common approaches adopt joint Deep Learning architectures in attention-based recurrent frameworks. In this work, we aim at …

Intent DetectionNatural Language Understandingslot-fillingSlot Filling+2

Almawave-SLU: A new dataset for SLU in Italian

2019-07-17 · Valentina Bellomaria, Giuseppe Castellucci, Andrea Favalli, Raniero Romagnoli

The widespread use of conversational and question answering systems made it necessary to improve the performances of speaker intent detection and understanding of related semantic slots, i.e., Spoken Language Understandi…

Intent DetectionQuestion AnsweringSpoken Language Understanding

EmpLite: A Lightweight Sequence Labeling Model for Emphasis Selection of Short Texts

2020-12-15 · ICON 2020 12 · Vibhav Agarwal, Sourav Ghosh, Kranti Chalamalasetti, Bharath Challa 외

Word emphasis in textual content aims at conveying the desired intention by changing the size, color, typeface, style (bold, italic, etc.), and other typographical features. The emphasized words are extremely helpful in …