paper-with-me

홈 › Papers

GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

2026-05-28 · Fabian Mewes, Anne Lauscher, Vagrant Gautam arxiv

Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about reference. More recently, the interplay between reasoning and bias has been investigated with the task of pronoun fidelity, which assesses models' abilities to correctly reuse a previously-specified pronoun for a discourse entity, independent of other potentially distracting discourse entities mentioned in between. However, such research focuses on English, which is a language with limited grammatical gender and almost no gender agreement. In this paper we contribute a novel, large-scale dataset, GRUFF, to measure pronoun fidelity in German, covering four different gender agreement systems in nouns, and four sets of pronouns. With this dataset, we show that LLMs show strong grammatical agreement for masculine and feminine entities in the absence of explicit context, but not for neopronouns xier and en. Models are generally not robust to distractors, but encoder-only models are more robust in German than in English, reflecting the importance of grammatical gender. Finally, we show that occupational stereotypes in this context are poorly correlated across grammatical cases, and across most models, except ones with closely related architectures. We release all code and data to encourage further work on gender-inclusive language and referential reasoning in German.

📄 PDF Abstract BibTeX arXiv:2605.30214

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?

2024-04-04 · Vagrant Gautam, Eileen Bingert, Dawei Zhu, Anne Lauscher 외

Robust, faithful and harm-free pronoun use for individuals is an important goal for language model development as their use increases, but prior work tends to study only one or two of these characteristics at a time. To …

DecoderSentence

A Mechanistic Understanding of Pronoun Fidelity in LLMs

2026-06-15 · Katharina Trinley, Jesujoba O. Alabi, Dietrich Klakow, Vagrant Gautam arxiv

Faithful and robust pronoun use is important for fair and coherent generations, yet large language models largely fail when multiple referents use different pronouns. To study the interplay of reasoning, repetition, and …

Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs

2025-05-05 · Elisa Forcada Rodríguez, Olatz Perez-de-Viñaspre, Jon Ander Campos, Dietrich Klakow 외

One of the goals of fairness research in NLP is to measure and mitigate stereotypical biases that are propagated by NLP systems. However, such work tends to focus on single axes of bias (most often gender) and the Englis…

Fairness

A Pronoun Test Suite Evaluation of the English--German MT Systems at WMT 2018

2018-10-01 · WS 2018 10 · Liane Guillou, Christian Hardmeier, Ekaterina Lapshinova-Koltunski, Sharid Lo{\'a}iciga

We evaluate the output of 16 English-to-German MT systems with respect to the translation of pronouns in the context of the WMT 2018 competition. We work with a test suite specifically designed to assess system quality i…

Machine TranslationNMTTranslation

Findings of the 2016 WMT Shared Task on Cross-lingual Pronoun Prediction

2019-11-27 · WS 2016 8 · Liane Guillou, Christian Hardmeier, Preslav Nakov, Sara Stymne 외

We describe the design, the evaluation setup, and the results of the 2016 WMT shared task on cross-lingual pronoun prediction. This is a classification task in which participants are asked to provide predictions on what …

Language ModelingLanguage ModellingPOS