paper-with-me

홈 › Papers

Text Rationalization for Robust Causal Effect Estimation

2025-12-05 · Lijinghua Zhang, Hengrui Cai arxiv

Recent advances in natural language processing have enabled the increasing use of text data in causal inference, particularly for adjusting confounding factors in treatment effect estimation. Although high-dimensional text can encode rich contextual information, it also poses unique challenges for causal identification and estimation. In particular, the positivity assumption, which requires sufficient treatment overlap across confounder values, is often violated at the observational level, when massive text is represented in feature spaces. Redundant or spurious textual features inflate dimensionality, producing extreme propensity scores, unstable weights, and inflated variance in effect estimates. We address these challenges with Confounding-Aware Token Rationalization (CATR), a framework that selects a sparse necessary subset of tokens using a residual-independence diagnostic designed to preserve confounding information sufficient for unconfoundedness. By discarding irrelevant texts while retaining key signals, CATR mitigates observational-level positivity violations and stabilizes downstream causal effect estimators. Experiments on synthetic data and a real-world study using the MIMIC-III database demonstrate that CATR yields more accurate, stable, and interpretable causal effect estimates than existing baselines.

📄 PDF Abstract BibTeX arXiv:2512.05373

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Inference

Similar Papers 제목 키워드 기반

Towards Trustworthy Explanation: On Causal Rationalization

2023-06-25 · Wenbo Zhang, Tong Wu, Yunlong Wang, Yong Cai 외

With recent advances in natural language processing, rationalization becomes an essential self-explaining diagram to disentangle the black box by selecting a subset of input texts to account for the major variation in pr…

Causal Inference

D-Separation for Causal Self-Explanation

2023-09-23 · NeurIPS 2023 11 · Wei Liu, Jun Wang, Haozhao Wang, Ruixuan Li 외

Rationalization is a self-explaining framework for NLP models. Conventional work typically uses the maximum mutual information (MMI) criterion to find the rationale that is most indicative of the target label. However, t…

Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-Rationalization

2024-10-08 · Wei Liu, Zhiying Deng, Zhongyu Niu, Jun Wang 외

An important line of research in the field of explainability is to extract a small subset of crucial rationales from the full input. The most widely used criterion for rationale extraction is the maximum mutual informati…

Recon: Reconstruction-Guided Reasoning Synthesis for User Modeling

2026-05-26 · Alan Zhu, Mihran Miroyan, Carolyn Wang, Andrew Zhou 외 arxiv

User modeling aims to use language models (LMs) to mimic an individual's behavior from a corpus of past context-action pairs (e.g., conversation turns), enabling the simulation of users in settings like behavioral scienc…

Variable-lag Granger Causality for Time Series Analysis

2019-12-18 · Chainarong Amornbunchornvej, Elena Zheleva, Tanya Y. Berger-Wolf

Granger causality is a fundamental technique for causal inference in time series data, commonly used in the social and biological sciences. Typical operationalizations of Granger causality make a strong assumption that e…

Causal InferenceLeadership InferenceTime SeriesTime Series Analysis