Detecting Independent Pronoun Bias with Partially-Synthetic Data Generation
We report that state-of-the-art parsers consistently failed to identify {`}hers{''} and {}theirs{''} as pronouns but identified the masculine equivalent {`}his{''}. We find that the same biases exist in recent language models like BERT. While some of the bias comes from known sources, like training data with gender imbalances, we find that the bias is {\_}amplified{\_} in the language models and that linguistic differences between English pronouns that are not inherently biased can become biases in some machine learning models. We introduce a new technique for measuring bias in models, using Bayesian approximations to generate partially-synthetic data from the model itself.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningSynthetic Data GenerationSimilar Papers 제목 키워드 기반
Automated detection of pronunciation errors in non-native English speech employing deep learning
Despite significant advances in recent years, the existing Computer-Assisted Pronunciation Training (CAPT) methods detect pronunciation errors with a relatively low accuracy (precision of 60% at 40%-80% recall). This Ph.…
Speech SynthesisGRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about reference. More recently, the interplay between reasoning and bias has been …
On Measuring Gender Bias in Translation of Gender-neutral Pronouns
Ethics regarding social bias has recently thrown striking issues in natural language processing. Especially for gender-related topics, the need for a system that reduces the model bias has grown in areas such as image ca…
EthicsImage CaptioningMachine TranslationSentence+1Type B Reflexivization as an Unambiguous Testbed for Multilingual Multi-Task Gender Bias
The one-sided focus on English in previous studies of gender bias in NLP misses out on opportunities in other languages: English challenge datasets such as GAP and WinoGender highlight model preferences that are "halluci…
TranslationComputer-assisted Pronunciation Training -- Speech synthesis is almost all you need
The research community has long studied computer-assisted pronunciation training (CAPT) methods in non-native speech. Researchers focused on studying various model architectures, such as Bayesian networks and deep learni…
AllSpeech Synthesistext-to-speechText to Speech