Deconfounded Lexicon Induction for Interpretable Social Science
NLP algorithms are increasingly used in computational social science to take linguistic observations and predict outcomes like human preferences or actions. Making these social models transparent and interpretable often requires identifying features in the input that predict outcomes while also controlling for potential confounds. We formalize this need as a new task: inducing a lexicon that is predictive of a set of target variables yet uncorrelated to a set of confounding variables. We introduce two deep learning algorithms for the task. The first uses a bifurcated architecture to separate the explanatory power of the text and confounds. The second uses an adversarial discriminator to force confound-invariant text encodings. Both elicit lexicons from learned weights and attentional scores. We use them to induce lexicons that are predictive of timely responses to consumer complaints (controlling for product), enrollment from course descriptions (controlling for subject), and sales from product descriptions (controlling for seller). In each domain our algorithms pick words that are associated with \textit{narrative persuasion}; more predictive and less confound-related than those of standard feature weighting and lexicon induction techniques like regression and log odds.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Inducing lexicons of in-group language with socio-temporal context
In-group language is an important signifier of group dynamics. This paper proposes a novel method for inducing lexicons of in-group language, which incorporates its socio-temporal context. Existing methods for lexicon in…
Cross-Lingual Emotion Lexicon Induction using Representation Alignment in Low-Resource Settings
Emotion lexicons provide information about associations between words and emotions. They have proven useful in analyses of reviews, literary texts, and posts on social media, among other things. We evaluate the feasibili…
SentenceTranslationImproving Bilingual Lexicon Induction for Low Frequency Words
This paper designs a Monolingual Lexicon Induction task and observes that two factors accompany the degraded accuracy of bilingual lexicon induction for rare words. First, a diminishing margin between similarities in low…
Bilingual Lexicon InductionExpanding Subjective Lexicons for Social Media Mining with Embedding Subspaces
Recent approaches for sentiment lexicon induction have capitalized on pre-trained word embeddings that capture latent semantic properties. However, embeddings obtained by optimizing performance of a given task (e.g. pred…
Word EmbeddingsInterpretable Language Modeling via Induction-head Ngram Models
Recent large language models (LLMs) have excelled across a wide range of tasks, but their use in high-stakes and compute-limited settings has intensified the demand for interpretability and efficiency. We address this ne…
Causal Language ModelingHuman fMRI response predictionLanguage ModelingLanguage Modelling