paper-with-me

홈 › Papers

P3B3: A Multi-Turn Conversational Benchmark for Measuring European and Brazilian Portuguese Variety Bias in LLMs

2026-06-15 · Rafael Ferreira, Inês Vieira, Inês Calvo, James Furtado, Iago Paulo, Diogo Tavares, Diogo Glória-Silva, David Semedo, João Magalhães arxiv

As Large Language Models (LLMs) become embedded in everyday communication, capturing regional linguistic variation is essential for reliable and equitable language use. In Portuguese, European (pt-PT) and Brazilian (pt-BR) varieties remain unevenly represented, with pt-BR dominating in data quantity, while LLM preference for Portuguese variants remains underexplored. To address this gap, we introduce P3B3, an expert-curated language variety agnostic benchmark of conversational prompts, along with an evaluation framework for measuring variety bias and controllability. Experiments on several models show that most LLMs exhibit a strong bias toward pt-BR, with variation in controllability across models. These results highlight the need for more balanced multilingual representation across language varieties.

📄 PDF Abstract BibTeX arXiv:2606.16753

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Finding Diamonds in Conversation Haystacks: A Benchmark for Conversational Data Retrieval

2025-10-03 · Yohan Lee, Yongwoo Song, Sangyeop Kim arxiv

We present the Conversational Data Retrieval (CDR) benchmark, the first comprehensive test set for evaluating systems that retrieve conversation data for product insights. With 1.6k queries across five analytical tasks a…

MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models

2025-10-18 · Young-Jun Lee, Byung-Kwan Lee, Jianshu Zhang, Yechan Hwang 외 arxiv

Vision-and-Language Models (VLMs) have shown impressive capabilities on single-turn benchmarks, yet real-world applications often demand more intricate multi-turn dialogues. Existing multi-turn datasets (e.g, MMDU, ConvB…

From Answers to Guidance: A Proactive Dialogue System for Legal Documents

2025-10-22 · Ashish Chouhan, Michael Gertz arxiv

The accessibility of legal information remains a constant challenge, particularly for laypersons seeking to understand and apply complex institutional texts. While the European Union provides open access to legislation, …

Goal Alignment in LLM-Based User Simulators for Conversational AI

2025-07-27 · Shuhaib Mehri, Xiaocheng Yang, Takyoung Kim, Gokhan Tur 외 arxiv

User simulators are essential to conversational AI, enabling scalable agent development and evaluation through simulated interactions. While current Large Language Models (LLMs) have advanced user simulation capabilities…

MultiTurnCleanup: A Benchmark for Multi-Turn Spoken Conversational Transcript Cleanup

2023-05-19 · Hua Shen, Vicky Zayats, Johann C. Rocholl, Daniel D. Walker 외

Current disfluency detection models focus on individual utterances each from a single speaker. However, numerous discontinuity phenomena in spoken conversational transcripts occur across multiple turns, hampering human r…