paper-with-me

Papers

BaSCo: An Annotated Basque-Spanish Code-Switching Corpus for Natural Language Understanding

2022-06-01 · LREC 2022 6 · Maia Aguirre, Laura García-Sardiña, Manex Serras, Ariane Méndez, Jacobo López

The main objective of this work is the elaboration and public release of BaSCo, the first corpus with annotated linguistic resources encompassing Basque-Spanish code-switching. The mixture of Basque and Spanish languages within the same utterance is popularly referred to as Euskañol, a widespread phenomenon among bilingual speakers in the Basque Country. Thus, this corpus has been created to meet the demand of annotated linguistic resources in Euskañol in research areas such as multilingual dialogue systems. The presented resource is the result of translating to Euskañol a compilation of texts in Basque and Spanish that were used for training the Natural Language Understanding (NLU) models of several task-oriented bilingual chatbots. Those chatbots were meant to answer specific questions associated with the administration, fiscal, and transport domains. In addition, they had the transverse potential to answer to greetings, requests for help, and chit-chat questions asked to chatbots. BaSCo is a compendium of 1377 tagged utterances with every sample annotated at three levels: (i) NLU semantic labels, considering intents and entities, (ii) code-switching proportion, and (iii) domain of origin.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language Understanding

Similar Papers 제목 키워드 기반

BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions

2022-05-03 · LREC 2022 6 · Nayla Escribano, Jon Ander González, Julen Orbegozo-Terradillos, Ainara Larrondo-Ureta 외

Parliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcri…

Versatile Speech Databases for High Quality Synthesis for Basque

2012-05-01 · LREC 2012 5 · I{\~n}aki Sainz, Daniel Erro, Eva Navas, Inma Hern{\'a}ez 외

This paper presents three new speech databases for standard Basque. They are designed primarily for corpus-based synthesis but each database has its specific purpose: 1) AhoSyn: high quality speech synthesis (recorded al…

Emotional Speech SynthesisSpeech SynthesisVocal Bursts Intensity PredictionVoice Conversion

Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing

2022-06-10 · LREC 2022 6 · Elena Alvarez Mellado, Constantine Lignos

We present a new corpus of Twitter data annotated for codeswitching and borrowing between Spanish and English. The corpus contains 9,500 tweets annotated at the token level with codeswitches, borrowings, and named entiti…

Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives

2025-03-13 · Miguel Romero-Arjona, Pablo Valle, Juan C. Alonso, Ana B. Sánchez 외

The battle for AI leadership is on, with OpenAI in the United States and DeepSeek in China as key contenders. In response to these global trends, the Spanish government has proposed ALIA, a public and transparent AI infr…

Red Teaming

Overview of GUA-SPA at IberLEF 2023: Guarani-Spanish Code Switching Analysis

2023-09-12 · Luis Chiruzzo, Marvin Agüero-Torales, Gustavo Giménez-Lugo, Aldo Alvarez 외

We present the first shared task for detecting and analyzing code-switching in Guarani and Spanish, GUA-SPA at IberLEF 2023. The challenge consisted of three tasks: identifying the language of a token, NER, and a novel t…

ArticlesNERSingle Particle Analysis