paper-with-me

Papers

BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions

2022-05-03 · LREC 2022 6 · Nayla Escribano, Jon Ander González, Julen Orbegozo-Terradillos, Ainara Larrondo-Ureta, Simón Peña-Fernández, Olatz Perez-de-Viñaspre, Rodrigo Agerri

Parliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcripts facilitate research on political discourse from a computational social science perspective. In this paper we release the first version of a newly compiled corpus from Basque parliamentary transcripts. The corpus is characterized by heavy Basque-Spanish code-switching, and represents an interesting resource to study political discourse in contrasting languages such as Basque and Spanish. We enrich the corpus with metadata related to relevant attributes of the speakers and speeches (language, gender, party...) and process the text to obtain named entities and lemmas. The obtained metadata is then used to perform a detailed corpus analysis which provides interesting insights about the language use of the Basque political representatives across time, parties and gender.

📄 PDF Abstract BibTeX arXiv:2205.01506

Code (1)

ixa-ehu/basqueparl 공식 구현

Similar Papers 제목 키워드 기반

Adding the Basque Parliament Corpus to ParlaMint Project

2022-06-01 · ParlaCLARIN (LREC) 2022 6 · Jon Alkorta, Mikel Iruskieta Quintian

The aim of this work is to describe the colection created with transcript of the Basque parliamentary speeches. This corpus follows the constraints of the ParlaMint project. The Basque ParlaMint corpus consists of two ve…

A Parallel Corpus of Translationese

2015-09-11 · Ella Rabinovich, Shuly Wintner, Ofek Luis Lewinsohn

We describe a set of bilingual English--French and English--German parallel corpora in which the direction of translation is accurately and reliably annotated. The corpora are diverse, consisting of parliamentary proceed…

Machine TranslationTranslation

BaSCo: An Annotated Basque-Spanish Code-Switching Corpus for Natural Language Understanding

2022-06-01 · LREC 2022 6 · Maia Aguirre, Laura García-Sardiña, Manex Serras, Ariane Méndez 외

The main objective of this work is the elaboration and public release of BaSCo, the first corpus with annotated linguistic resources encompassing Basque-Spanish code-switching. The mixture of Basque and Spanish languages…

Natural Language Understanding

The Norwegian Parliamentary Speech Corpus

2022-01-26 · LREC 2022 6 · Per Erik Solberg, Pablo Ortiz

The Norwegian Parliamentary Speech Corpus (NPSC) is a speech dataset with recordings of meetings from Stortinget, the Norwegian parliament. It is the first, publicly available dataset containing unscripted, Norwegian spe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Open-Source High Quality Speech Datasets for Basque, Catalan and Galician

2020-05-01 · LREC 2020 5 · Oddur Kjartansson, Alex Gutkin, er, Alena Butryna 외

This paper introduces new open speech datasets for three of the languages of Spain: Basque, Catalan and Galician. Catalan is furthermore the official language of the Principality of Andorra. The datasets consist of high-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+3