paper-with-me

홈 › Papers

How Contentious Terms About People and Cultures are Used in Linked Open Data

2023-11-13 · Andrei Nesterov, Laura Hollink, Jacco van Ossenbruggen

Web resources in linked open data (LOD) are comprehensible to humans through literal textual values attached to them, such as labels, notes, or comments. Word choices in literals may not always be neutral. When outdated and culturally stereotyping terminology is used in literals, they may appear as offensive to users in interfaces and propagate stereotypes to algorithms trained on them. We study how frequently and in which literals contentious terms about people and cultures occur in LOD and whether there are attempts to mark the usage of such terms. For our analysis, we reuse English and Dutch terms from a knowledge graph that provides opinions of experts from the cultural heritage domain about terms' contentiousness. We inspect occurrences of these terms in four widely used datasets: Wikidata, The Getty Art & Architecture Thesaurus, Princeton WordNet, and Open Dutch WordNet. Some terms are ambiguous and contentious only in particular senses. Applying word sense disambiguation, we generate a set of literals relevant to our analysis. We found that outdated, derogatory, stereotyping terms frequently appear in descriptive and labelling literals, such as preferred labels that are usually displayed in interfaces and used for indexing. In some cases, LOD contributors mark contentious terms with words and phrases in literals (implicit markers) or properties linked to resources (explicit markers). However, such marking is rare and non-consistent in all datasets. Our quantitative and qualitative insights could be helpful in developing more systematic approaches to address the propagation of stereotypes via LOD.

📄 PDF Abstract BibTeX arXiv:2311.10757

Code (1)

mnn-001-p/lodlit 공식 구현

Tasks

DescriptiveWord Sense Disambiguation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Linguistic Characterization of Divisive Topics Online: Case Studies on Contentiousness in Abortion, Climate Change, and Gun Control

2021-08-30 · Jacob Beel, Tong Xiang, Sandeep Soni, Diyi Yang

As public discourse continues to move and grow online, conversations about divisive topics on social media platforms have also increased. These divisive topics prompt both contentious and non-contentious conversations. A…

Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4

2025-05-23 · Zhuozhuo Joy Liu, Farhan Samir, Mehar Bhatia, Laura K. Nelson 외

LLMs have been demonstrated to align with the values of Western or North American cultures. Prior work predominantly showed this effect through leveraging surveys that directly ask (originally people and now also LLMs) a…

All

Probing Pre-Trained Language Models for Cross-Cultural Differences in Values

2022-03-25 · Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein

Language embeds information about social, cultural, and political values people hold. Prior work has explored social and potentially harmful biases encoded in Pre-Trained Language models (PTLMs). However, there has been …

A Multi-Cultural Repository of Automatically Discovered Linguistic and Conceptual Metaphors

2014-05-01 · LREC 2014 5 · Samira Shaikh, Tomek Strzalkowski, Ting Liu, George Aaron Broadwell 외

In this article, we present details about our ongoing work towards building a repository of Linguistic and Conceptual Metaphors. This resource is being developed as part of our research effort into the large-scale detect…

Risks of Cultural Erasure in Large Language Models

2025-01-02 · Rida Qadri, Aida M. Davani, Kevin Robinson, Vinodkumar Prabhakaran

Large language models are increasingly being integrated into applications that shape the production and discovery of societal knowledge such as search, online education, and travel planning. As a result, language models …

Language ModelingLanguage Modelling