Almawave-SLU: A new dataset for SLU in Italian
The widespread use of conversational and question answering systems made it necessary to improve the performances of speaker intent detection and understanding of related semantic slots, i.e., Spoken Language Understanding (SLU). Often, these tasks are approached with supervised learning methods, which needs considerable labeled datasets. This paper presents the first Italian dataset for SLU. It is derived through a semi-automatic procedure and is used as a benchmark of various open source and commercial systems.
Code (0)
등록된 구현이 없습니다.
Tasks
Intent DetectionQuestion AnsweringSpoken Language UnderstandingSimilar Papers 제목 키워드 기반
Fauno: The Italian Large Language Model that will leave you senza parole!
This paper presents Fauno, the first and largest open-source Italian conversational Large Language Model (LLM). Our goal with Fauno is to democratize the study of LLMs in Italian, demonstrating that obtaining a fine-tune…
GPULanguage ModelingLanguage ModellingLarge Language Model+1Data on the annual aggregated income taxes of the Italian municipalities over the quinquennium 2007-2011
This dataset contains the annual aggregated income taxes of all the Italian municipalities over the years 2007-2011. Data are clustered over the Italian regions and provinces. The source of the data is the Italian Minist…
ITALIC: An Italian Intent Classification Dataset
Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, th…
Classificationintent-classificationIntent Classificationspeech-recognition+2Dialogue Act and Slot Recognition in Italian Complex Dialogues
Since the advent of Transformer-based, pretrained language models (LM) such as BERT, Natural Language Understanding (NLU) components in the form of Dialogue Act Recognition (DAR) and Slot Recognition (SR) for dialogue sy…
Natural Language UnderstandingKIND: an Italian Multi-Domain Dataset for Named Entity Recognition
In this paper we present KIND, an Italian dataset for Named-entity recognition. It contains more than one million tokens with annotation covering three classes: person, location, and organization. The dataset (around 600…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER