paper-with-me

Papers

Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval

2023-11-03 · Jinrui Yang, Timothy Baldwin, Trevor Cohn

We present Multi-EuP, a new multilingual benchmark dataset, comprising 22K multi-lingual documents collected from the European Parliament, spanning 24 languages. This dataset is designed to investigate fairness in a multilingual information retrieval (IR) context to analyze both language and demographic bias in a ranking context. It boasts an authentic multilingual corpus, featuring topics translated into all 24 languages, as well as cross-lingual relevance judgments. Furthermore, it offers rich demographic information associated with its documents, facilitating the study of demographic bias. We report the effectiveness of Multi-EuP for benchmarking both monolingual and multilingual IR. We also conduct a preliminary experiment on language bias caused by the choice of tokenization strategy.

📄 PDF Abstract BibTeX arXiv:2311.01870

Code (1)

jrnlp/multi-eup 공식 구현

Tasks

BenchmarkingFairnessInformation RetrievalRetrievalText Reranking

Similar Papers 제목 키워드 기반

The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings

2023-09-18 · Michal Mochtak, Peter Rupnik, Nikola Ljubešić

The paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which are used in a series of experiments focused on training a robust sentiment identifier for parliamentary proce…

Decision MakingLanguage ModelingLanguage ModellingSentiment Analysis

Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification

2026-02-18 · Taja Kuzman Pungeršek, Peter Rupnik, Daniela Širinić, Nikola Ljubešić arxiv

This paper introduces ParlaCAP, a large-scale dataset for analyzing parliamentary agenda setting across Europe, and proposes a cost-effective method for building domain-specific policy topic classifiers. Applying the Com…

Europarl-ST: A Multilingual Corpus For Speech Translation Of Parliamentary Debates

2019-11-08 · Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerdà, Javier Jorge, Nahuel Roselló 외

Current research into spoken language translation (SLT),or speech-to-text translation, is often hampered by the lack of specific data resources for this task, as currently available SLT datasets are restricted to a limit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+4

DCEP -Digital Corpus of the European Parliament

2014-05-01 · LREC 2014 5 · Najeh Hajlaoui, David Kolovratnik, Jaakko V{\"a}yrynen, Ralf Steinberger 외

We are presenting a new highly multilingual document-aligned parallel corpus called DCEP - Digital Corpus of the European Parliament. It consists of various document types covering a wide range of subject domains. With a…

Machine TranslationSentenceTranslation

Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset

2025-10-23 · Paul Lerner, François Yvon arxiv

The political biases of Large Language Models (LLMs) are usually assessed by simulating their answers to English surveys. In this work, we propose an alternative framing of political biases, relying on principles of fair…