Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond

Amir Feder; Katherine Keith; Emaad Manzoor; Reid Pryzant; Dhanya Sridhar; Zach Wood-Doughty; Jacob Eisenstein; Justin Grimmer; Roi Reichart; Margaret Roberts; Brandon Stewart; Victor Veitch; Diyi Yang

Vol. 10 (2022)

TACL approved

Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond

Published 2022-10-18

Amir Feder
Katherine Keith
Emaad Manzoor
Reid Pryzant
Dhanya Sridhar
Zach Wood-Doughty
Jacob Eisenstein
Justin Grimmer
Roi Reichart
Margaret Roberts
Brandon Stewart
Victor Veitch
Diyi Yang

Amir Feder
Technion

Katherine Keith
University of Massachusetts Amherst

Emaad Manzoor
University of Wisconsin - Madison

Reid Pryzant
Stanford University

Dhanya Sridhar
University of Montreal and Mila - Quebec Artificial Intelligence Institute

Zach Wood-Doughty
Johns Hopkins University

Jacob Eisenstein
Google Research

Justin Grimmer
Stanford University

Roi Reichart
Technion, Israel Institute of Technology

Margaret Roberts
University of California San Diego

Brandon Stewart
Princeton University

Victor Veitch
Google Research, University of Chicago

Diyi Yang
Georgia Tech

Abstract

A fundamental goal of scientific research is to learn about causal relationships. However, despite its critical role in the life and social sciences, causality has not had the same importance in Natural Language Processing (NLP), which has traditionally placed more emphasis on predictive tasks. This distinction is beginning to fade, with an emerging area of interdisciplinary research at the convergence of causal inference and language processing. Still, research on causality in NLP remains scattered across domains without unified definitions, benchmark datasets and clear articulations of the challenges and opportunities in the application of causal inference to the textual domain, with its unique properties. In this survey, we consolidate research across academic areas and situate it in the broader NLP landscape. We introduce the statistical challenge of estimating causal effects with text, encompassing settings where text is used as an outcome, treatment, or to address confounding. In addition, we explore potential uses of causal inference to improve the robustness, fairness, and interpretability of NLP models. We thus provide a unified overview of causal inference for the NLP community.

Article at MIT Press Presented at EMNLP 2022

Author Biography

Roi Reichart

Associate Professor, Faculty of Industrial Engineering and Management, The Technion, Israel Institute of Technology.