Leveraging study of robustness and portability of spoken language understanding systems across languages and domains: the PORTMEDIA corpora
The PORTMEDIA project is intended to develop new corpora for the evaluation of spoken language understanding systems. The newly collected data are in the field of human-machine dialogue systems for tourist information in French in line with the MEDIA corpus. Transcriptions and semantic annotations, obtained by low-cost procedures, are provided to allow a thorough evaluation of the systems' capabilities in terms of robustness and portability across languages and domains. A new test set with some adaptation data is prepared for each case: in Italian as an example of a new language, for ticket reservation as an example of a new domain. Finally the work is complemented by the proposition of a new high level semantic annotation scheme well-suited to dialogue data.
Code (0)
등록된 구현이 없습니다.
Tasks
Semantic CompositionSpeech RecognitionSpoken Language UnderstandingSimilar Papers 제목 키워드 기반
Robustesse et portabilit\'es multilingue et multi-domaines des syst\`emes de compr\'ehension de la parole : les corpus du projet PortMedia (Robustness and portability of spoken language understanding systems among languages and domains : the PORTMEDIA project) [in French]
Curriculum-based transfer learning for an effective end-to-end spoken language understanding and domain portability
We present an end-to-end approach to extract semantic concepts directly from the speech audio signal. To overcome the lack of data available for this spoken language understanding approach, we investigate the use of a tr…
POSPOS TaggingSpoken Language UnderstandingTransfer LearningA dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
Self-Supervised Learning is vastly used to efficiently represent speech for Spoken Language Understanding, gradually replacing conventional approaches. Meanwhile, textual SSL models are proposed to encode language-agnost…
Self-Supervised LearningSpoken Language UnderstandingSemantic enrichment towards efficient speech representations
Over the past few years, self-supervised learned speech representations have emerged as fruitful replacements for conventional surface representations when solving Spoken Language Understanding (SLU) tasks. Simultaneousl…
Spoken Language UnderstandingOn the Use of Semantically-Aligned Speech Representations for Spoken Language Understanding
In this paper we examine the use of semantically-aligned speech representations for end-to-end spoken language understanding (SLU). We employ the recently-introduced SAMU-XLSR model, which is designed to generate a singl…
Representation LearningSentenceSentence EmbeddingSentence-Embedding+2