An algorithm for controlled text analysis on Wikipedia
While numerous work has examined bias on Wikipedia, most approaches fail to control for possible confounding variables. In this work, given a target corpus for analysis (e.g. biography pages about women), we present a method for constructing a control corpus that matches the target corpus in as many attributes as possible, except the target attribute (e.g. the gender of the subject). This methodology can be used to analyze specific types of bias in Wikipedia articles, for example, gender or racial bias, while minimizing the influence of confounding variables.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesAttributeSimilar Papers 제목 키워드 기반
ToTTo: A Controlled Table-To-Text Generation Dataset
We present ToTTo, an open-domain English table-to-text dataset with over 120,000 training examples that proposes a controlled generation task: given a Wikipedia table and a set of highlighted table cells, produce a one-s…
Conditional Text GenerationData-to-Text GenerationSentenceTable-to-Text Generation+1The Role of Wikipedia in Text Analysis and Retrieval
This tutorial examines the characteristics, advantages and limitations of Wikipedia relative to other existing, human-curated resources of knowledge; derivative resources, created by converting semi-structured content in…
Coreference ResolutionInformation RetrievalRetrievalWikiIns: A High-Quality Dataset for Controlled Text Editing by Natural Language Instruction
Text editing, i.e., the process of modifying or manipulating text, is a crucial step in human writing process. In this paper, we study the problem of controlled text editing by natural language instruction. According to …
InformativenessEstimating Grammatical Gender Directions in Contextual Embeddings under Controlled and Natural Contexts
Contextual language models conflate grammatical gender and social semantic bias in gendered languages such as Spanish. Existing gender debiasing approaches only operate on static word embeddings leaving contextual repres…
Wikipedia Text Reuse: Within and Without
We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse b…
Retrieval