EconBase
← Back to paper

Mining Causality: AI-Assisted Search for Instrumental Variables

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

91,903 characters · 22 sections · 42 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Mining Causality: AI-Assisted Search for Instrumental Variables

abstractThe instrumental variables (IVs) method is a leading empirical strategy for causal inference. Finding IVs is a heuristic and creative process, and justifying its validity---especially exclusion restrictions---is largely rhetorical. We propose using large language models (LLMs) to search for new IVs through narratives and counterfactual reasoning, similar to how a human researcher would. The stark difference, however, is that LLMs can dramatically accelerate this process and explore an extremely large search space. We demonstrate how to construct prompts to search for potentially valid IVs. We contend that multi-step and role-playing prompting strategies are effective for simulating the endogenous decision-making processes of economic agents and for navigating language models through the realm of real-world scenarios, rather than anchoring them within the narrow realm of academic discourses on IVs. We apply our method to three well-known examples in economics: returns to schooling, supply and demand, and peer effects. We then extend our strategy to finding (i) control variables in regression and difference-in-differences and (ii) running variables in regression discontinuity designs. Keywords: Causal inference, instrumental variables, exclusion restrictions, artificial intelligence, large language models. JEL Codes: C26, C36, C5.

Introduction

Endogeneity is the major obstacle in conducting causal inference in observational settings. Since the credibility revolution angrist2010credibility and the causal revolution pearl2000causality, researchers in social science, statistics and other adjacent fields have developed various identification strategies to overcome endogeneity by restoring versions of quasi-experiments. A leading strategy is the instrumental variables (IVs) method. Over decades, researchers with their ingenuity have discovered IVs in various settings and justified their satisfaction of exclusion restrictions (e.g., IVs are conditionally exogenous of latent variables). With its various applicability, the IVs method has prevailed across all subfields of economics and beyond imbens1994identification,Blundell:2003wi,heckman2005structural,hernan2006instruments.\footnote{See mogstad2024instrumental for a more recent survey.}

Exclusion restrictions are fundamentally untestable assumptions.\footnote{An exception is a favorable situation where one enjoys over-identifying restrictions. We discuss this point in our context below. Unlike the exclusion restriction, the IV relevance is testable from data stock_yogo_2005,olea2013robust.} Often, in justifying them, researchers resort to rhetorical arguments specific to each setting. This non-statistical process follows the discovery of potential candidate IVs, which itself requires researchers' counterfactual reasoning and creativity---and sometimes luck. These elements all contribute to the heuristic processes employed by human researchers.

We demonstrate that large language models (LLMs) can facilitate the discovery of new IVs. Considering that narratives are the primary method of supporting IV exclusion, we believe that LLMs, with sophisticated language processing abilities, are well-suited to assist in the search for new valid IVs and justify them rhetorically, just as human researchers have done for decades. The stark difference, however, is that LLMs can accelerate this process at an exponentially faster rate and explore an extremely large search space, to an extent that human researchers cannot match. It is now recognized that artificial intelligence (AI) shows remarkable performances in conducting systematic searches for hypotheses and refining the search jumper2021highly,ludwig2024machine. Furthermore, LLMs are argued to be capable of conducting counterfactual reasoning---or, perhaps more precisely, exploring alternative scenarios---which makes them a promising tool for causal inference.

There are at least four benefits to pursuing this AI-assisted approach to discovering IVs. First, researchers can conduct a systematic search at a speedy rate, while adapting to the particularities of their settings. Second, interacting with AI tools can inspire ideas for possible domains for novel IVs. Third, the systematic search could increase the possibility of obtaining multiple IVs, which would then enable formal (i.e., statistical) testing of their validity via over-identifying restrictions. Fourth, having a list of candidate IVs would increase the chances of finding actual data that contain IVs or guide the construction of such data, including the design of experiments to generate IVs.

We show how to construct prompts in a way that guides LLMs to search for candidates for valid IVs. The text representation of exclusion restrictions (among others) is the main component of the prompts. We propose a multi-step approach in prompting that divides a discovery task into multiple subtasks, and thus separates counterfactual statements of different complexities. At the same time, we propose using role-playing prompts, arguing that they align with the very source of endogeneity, namely, agents' decisions.\footnote{Decisions of economic agents have been at the root of challenges for causal analyses in econometrics heckman1979sample,manski1993identification.} By doing so, we equip LLMs with the perspective of agents, enabling them to mimic agents' endogenous decision-making processes and gather contextual information in realistic scenarios. This approach also makes it convenient to impose statistical conditioning that qualifies the characteristics of the agent. Another benefit of multi-stage, role-playing prompts is that they help navigate language models through the realm of real-world scenarios, rather than anchor them within the narrow realm of academic discourses on IVs. Each stage's prompt focuses only on a portion of the IV assumptions, translated into an agent's real-world problem, thereby minimizing the likelihood that the LLM perceives the task as a search for IVs.\footnote{Even if LLMs exhibit memorization from academic texts, we still find value in the procedure as long as the list of discovered IVs includes those that are recognized as new by researchers.}

To prove the proposed concept and illustrate the actual performance of an LLM, we conduct discovery exercises using OpenAI's ChatGPT-4 (GPT4), one of the leading LLMs, to find IVs in three well-known examples in empirical economics: returns to schooling, supply and demand, and peer effects. In all three examples, GPT4 produced a list of candidate IVs, some of which appear to be new in the literature and provided rationale for their validity. The list also contains IVs that are popularly used in the literature. Our initial assessment of the results suggests that the proposed method can work in practice. In the peer effect example, we also demonstrate that the proposed method can be effective in exploring relatively new topics for empirical research, which may in turn increase the possibility of finding novel IVs.

From a broader perspective, the proposal is to systematically “search for exogeneity.” We extend the exercise to other causal inference methods: (i) searching for control variables in regression and difference-in-differences methods and (ii) searching for running variables in regression discontinuity designs. We construct relevant prompts and run them in well-known examples in the literature.

A list of candidate IVs or control variables produced as a result of the proposed method is not absolute. Rather, we hope that it serves as a valuable benchmark that inspires empirical researchers about which types of variables to consider and which domains to explore. The dialogue carried out with LLMs in the process can also help researchers solidify arguments or counterarguments for the validity of variables. After all, AI---like any machines---cannot be the ultimate authority (at least not yet). We believe a human researcher assisted by AI can choose research designs and conduct causal inference more effectively.

This essay contributes to a recent agenda in the social science literature on using AI to assist creative and heuristic parts of human research processes. This agenda views machine learning and AI as not only data-processing and prediction tools for economic research (mullainathan2017machine,athey2019machine), but also as tools that can improve conventional research practices themselves. In very interesting work, ludwig2024machine use generative models to systematically produce hypotheses that are comprehensible by humans in otherwise daunting settings. They make progress in research areas where the use of AI has been limited because, as they argue, establishing causal relationships in social science is an “open world” problem, unlike “closed world” problems in physical science.\footnote{The latter can be viewed as extremely difficult computation problems where machine learning makes significant progress; e.g., detecting new proteins using AlphaFold jumper2021highly or advances in particle physics and cosmology using machine learning carleo2019machine.} In related work, mullainathan2024predictive use predictive (neural network) algorithms to recover old anomalies and discover new ones in economic theory models. We do not attempt to generate hypotheses, although the new variables discovered implicitly maintain a range of hypotheses on their validity.

LLMs has only very recently been used in social science research. Notably, du2024labor use fine-tuned LLMs (Meta's LLaMA in particular) to predict job transitions and understand career trajectories in labor economics. They show that the prediction accuracy remarkably outperforms those from traditional job transition economic models. manning2024automated propose to use LLMs to automate the entire process of social scientific research, from data generation to testing causal hypotheses. We employ LLMs in statistical causal inference by incorporating specific structure from econometric assumptions and allowing for human intervention in discovery processes.

This paper also relates to the approach of using LLMs in causal discovery (ban2023causal,cohrs2024large,jiralerspong2024efficient,le2024multi,long2023can,takayama2024integrating); also see wan2024bridging for a recent survey and references therein. However, the fundamental difference of our approach to this line of work is that we use LLMs to systematically discover variables with particular causal structure rather than using LLMs to find causal links among a pre-determined set of variables.

The paper is organized as follows. Section (ref) states the IV assumptions and Section (ref) proposes the main idea of IV discovery along with the prompting strategies. Section (ref) provides the examples of discovered IVs. Sections (ref)--(ref) contain extensions: (i) the use of an adversarial LLM to review and refine the discovery process and (ii) the extension of the paper's approach to other causal inference settings. Section (ref) concludes.

Notation and IV Assumptions

We first formally state our discovery goal. Let $Y$ be the outcome of interest, $D$ be the potentially endogenous treatment, $\mathcal{Z}_{K}\equiv\{Z_{1},...,Z_{K}\}$ be the list of IVs $Z_{k}$'s with $K$ being the desired number of IVs to discover, and $X$ be the covariates. Let $Y(d,z_{k})$ be the counterfactual outcome given $(d,z_{k})$. Let “$\perp$” denote statistical independence. We say $Z_{k}$ is a valid IV if it satisfies the following two assumptions:

myas{REL}[Relevance]Conditional on $X$, the distribution of $D$ given $Z_{k}=z_{k}$ is a nontrivial function of $z_{k}$.
myas{EX}[Exclusion]For any $(d,z_{k})$, $Y(d,z_{k})=Y(d)$.
myas{IND}[Independence]For any $d$, $Y(d)\perp Z_{k}$ conditional on $X$.

The goal of our exercise is to search for IVs that satisfy Assumptions (ref), (ref) and (ref).\footnote{One can consider a weaker version of (ref) (i.e., mean independence and nonzero correlation). Although we do not believe our ultimate findings significantly differ from this relaxation, our prompts can reflect it.} Suppressing $X$, Figure (ref) depicts the causal direct acyclic graph (DAG) that implies (ref), (ref) and (ref) and with $Y(d)$ being a transformation of latent confounders $U$. This diagram is useful in describing our procedures.

figure[figure omitted — 672 chars of source]

Prompt Construction

We propose a two-step approach for IV discovery. In Step 1, we prompt an LLM to search for IVs that satisfy a verbal description of (ref) and (ref) (i.e., \raisebox{-0.3\height}{

tikzpicture[tikzpicture omitted — 285 chars of source]

}). In Step 2, we prompt the LLM to refine the search by selecting---among the IVs found in Step 1---those that satisfy a verbal description of (ref) (i.e., \raisebox{-0.3\height}{

tikzpicture[tikzpicture omitted — 284 chars of source]

}). In both steps, the prompts will involve counterfactual statements. In each step, we ask the LLM to provide rationale for its responses. This feature is useful for the user to understand the LLM's reasoning. The two steps can be conducted in the same session or in separate sessions. However, when submitting different queries, we recommend that each two-step query be conducted in a separate session to avoid interference across queries. In Appendix (ref), we also present a full three-step prompting that focuses on each of Assumptions (ref), (ref) and (ref) in each step.

We propose a multi-step approach for several reasons: First, LLMs are known to yield better performance when handling subtasks step-by-step, focusing on important details in interpreting the prompts and avoiding errors wu2022ai. Second, this approach creates more room for the user to inspect intermediate outputs, facilitating the evaluation of final outputs. In particular, Step 2 involves more complex counterfactual statements than Step 1, allowing the user (and the LLM) to apply varying degrees of attention when fine-tuning is needed. Third, intermediate outputs themselves can provide information and offer insights. Finally, this approach significantly reduces the likelihood that LLMs recognize the task as IV discovery and generate text from relevant academic sources.\footnote{This would be especially true when each step is conducted in a separate independent session.}

Alongside the multi-step method, we propose a role-playing approach. It has been reported that LLMs---including GPT4---gather better contextual information and generate more tailored and unique responses when prompts are structured as role-plays.\footnote{OpenAI Developer Forum: https://community.openai.com/t/make-chatgpt-better-for-roleplay-scenarios/344244} In fact, in most scenarios, the explanatory variable $D$ represents an economic agent's decision, which naturally facilitate role-playing. Additionally, role-playing prompts are more effective in guiding LLMs to respond as the relevant economic agent rather than as a researcher searching for IVs. In Appendix (ref), we compare our multi-step, role-playing prompting strategy with a more direct approach that explicitly states the goal of IV search, arguing that the former is more effective.

To simplify the exposition, in Sections (ref)--(ref), we first demonstrate the prompt construction without introducing covariates (in which case (ref) and (ref) should hold unconditionally). We then construct more realistic prompts with covariates in Section (ref). The prompts presented here can serve as a benchmark for more sophisticated prompts; we discuss them in Section (ref).

Step 1: Prompts to Search for IVs

For Step 1, Prompt (ref) is a role-playing prompt that queries the search for $K_{0}$ IVs (obtaining $\mathcal{Z}_{K_{0}}$) that satisfy verbal versions of (ref) and (ref) (with no $X$). In all prompts below, each bracketed term represents a user input: {[}treatment{]} is the treatment $D$, {[}agent{]} is the economic agent whose decision is $D$, {[}scenario{]} is the specific setting of interest, {[}outcome{]} is the outcome $Y$, and {[}K_0{]} is the desired number of variables $K_{0}$. When prompting, we ask the LLM to play the role of {[}agent{]} to make a \texttt{{[}treatment{]}} decision in a hypothetical \texttt{{[}scenario{]}}. Examples of these inputs are given in Section (ref).

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{1}[Search for IVs]\end{myprompt} you are {[}agent{]} who needs to make a {[}treatment{]} decision in {[}scenario{]}. what are factors that can determine your decision but do not directly affect your {[}outcome{]}, except through {[}treatment{]} (that is, factors that affect your {[}outcome{]} only through {[}treatment{]})? list {[}K_0{]} factors that are quantifiable. explain the answers.

}}

There are at least two variants of Prompt (ref) that may be useful in certain scenarios. First, instead of “list {[}K_0{]} factors that are quantifiable” one may simply write “list {[}K_0{]} factors” or even “list {[}K_0{]} factors that are hard to quantify.” This would return candidates of IVs that are harder to measure but can inspire creative data collection (e.g., text, images, or other unstructured data). Second, one can expand Prompt (ref) to be more specific about categorizing factors for relevant parties in a given setting. For example, in the schooling scenario (Section (ref)), we request separate lists for student factors and school factors. This approach can facilitate the user's evaluation of the results.

Step 2: Prompts to Refine the Search for IVs

Take the set of IVs, $\mathcal{Z}_{K_{0}}\equiv\{Z_{1},...,Z_{K_{0}}\}$, obtained by running Prompt (ref) in Step 1. Next, for Step 2 in the same session, Prompt (ref) is a role-playing prompt that queries the search for $K$ IVs (obtaining $\mathcal{Z}_{K}$, $K\le K_{0}$) within $\mathcal{Z}_{K_{0}}$ that satisfy a verbal version of (ref) (with no $X$). Below, {[}confounders{]} is the user input for unobserved confounders of concern and {[}K{]} is the user choice of $K$. In this prompt, we ask the LLM to continue playing the same role as in Prompt (ref).

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{2}[Refine IVs]\end{myprompt} you are {[}agent{]} in {[}scenario{]}, as previously described. among the {[}K_0{]} factors listed above, choose {[}K{]} factors that are most likely to be unassociated with {[}confounders{]}, which determine your {[}outcome{]}. the chosen factors can still influence your {[}treatment{]}. for each chosen factor, explain your reasoning.

}}

Unlike Prompt (ref), this prompt contains a statement about variables typically unobserved to researchers, which may pose challenges. We believe that incorporating the researcher's prior knowledge on latent confounders helps simplify the overall search process and yield more desirable results.\footnote{This relates to few-shot learning discussed in Section (ref).} For instance, in the schooling scenario, one can specify “innate ability and personality and school quality.” Alternatively, if the user prefers a more agnostic approach, they can list {[}confounders{]} as “other possible factors.” Another option is to systematically search for possible unobserved confounders; see Section (ref) for related prompting strategies. In Prompt (ref), we use the term “unassociated.” If the LLM ever captures the nuance of this word, it reflects the mean independence version of (ref), making the search easier. Interestingly, an alternative phasing such as “choose {[}K{]} factors that are purely random”, which may seem a straightforward way to impose (ref) without needing to specify unobserved confounders, often fails to produce intended outputs.

There are useful variants of Prompt (ref). First, one can omit {[}K{]} and instruct the LLM to “choose all factors” from $\mathcal{Z}_{K_{0}}$ that are likely to satisfy (ref), allowing the LLM to determine $K$ independently; we apply this strategy in all examples later. Second, as a sanity check, one can direct the LLM to select elements in $\mathcal{Z}_{K_{0}}$ that violate (ref) in addition to those that satisfy it. This can be achieved by adding “also choose factors that are, in contrast, associated with {[}confounders{]}.” To gain further insights, the user can request explanations for factors that she identifies as valid IVs in initial set $\mathcal{Z}_{K_{0}}$ from Step 1, but which are somehow not included in the final set $\mathcal{Z}_{K}$ by the LLM. We apply the last approach to the application in Section (ref).

Extension: Prompts to Search and Refine with Covariates

Typically, IVs are argued to be valid after conditioning on a list of covariates (as reflected in (ref)--(ref)). The IV discovery with covariates can be approached in at least two different ways. We can prompt the LLM to either (i) search for IVs conditional on predetermined covariates; or (ii) jointly search for IVs and covariates that satisfy (ref)--(ref). We focus on option (i); option (ii) is discussed in Appendix (ref). Whenever covariates are searched, option (i) can be viewed as initiating an IV search in a new independent session with the searched covariates.

We construct a prompt that introduces the notion of conditioning variables; role-playing prompts are suitable for this purpose. Here, we only modify Prompt (ref). Although (ref) also involves conditioning on $X$, we find that results are not sensitive to a relevant modification of Prompt (ref). Prompt (ref) qualifies both {[}agent{]} and {[}scenario{]} by {[}covariates{]}, the pre-determined user choice of covariates. It extends Prompt (ref) by modifying the first sentence. Prompt (ref) is intended to be run after completing Prompt (ref).

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{2$_{x}$}[Refine IVs with Covariates]\end{myprompt} suppose you are {[}agent{]} in {[}scenario{]} with {[}covariates{]}. among the {[}K_0{]} factors listed above, choose {[}K{]} factors that are most likely to be unassociated with {[}confounders{]}, which determine your {[}outcome{]}. the chosen factors can still influence your {[}treatment{]}. for each chosen factor, explain your reasoning.

}}

The recommended approach for incorporating {[}covariates{]} is to assign specific values for the covariates. For instance, in the schooling scenario, one can write “suppose you are an asian female high school student from california who considers attending a private college.”\footnote{One can run multiple queries across different values of covariates for robustness, although this does not appear to be necessary in most cases unless extreme values are assigned in the initial run.} Alternatively, one can simply use terms like “specific” or “particular” along with the name of chosen covariates (e.g., “suppose you are a high school student with specific gender, race, and regional origin who considers attending a college of specific type”).

Discovered IVs

Using Prompts (ref) and (ref) described in the previous section, we aim to identify candidates for IVs in four well-known examples in economics: returns to schooling, supply and demand, and peer effects. These examples are chosen for their significance in the empirical economics literature (representing labor economics, industrial organization, and development economics, respectively). They commonly employ the IVs method as an empirical strategy. The main purpose of this exercise is to evaluate the performance of LLMs in executing the proposed method and to demonstrate the practical applicability of the method.

To summarize the findings, in all the examples, LLMs appear to discover new candidates for IVs and candidates that are related to well-known IVs in the literature. When Also, many candidates demonstrate high levels of specificity to their context. With the results produced, we hope to spark debates and inspire the discovery of new and better IVs.

The prompts we construct in each example slightly deviate from the templates of Prompts (ref) and (ref) to better adapt to the scenario and enhance the flow of English language. For each example, we present results from the initial single run of the prompts without any curation or further refinement. When the results match the IVs described in the literature exactly, we include the corresponding references (to the best of our knowledge) and their citation counts. Results across sessions are largely consistent, although they can vary when different values of $K_{0}$ and $K$ are chosen. We use GPT4 as our LLM.

Returns to Education

Suppose we are interested in estimating the causal effects of educational attainment (e.g., college attendance, years of schooling) on earnings. The main latent confounders in this setting is unobserved individual and school characteristics (e.g., student ability and personality, school quality) that affect both the schooling decisions and future earnings. To address this endogeneity and recover meaningful causal effects (e.g., local average treatment effects imbens1994identification), IVs such as distance to schools, tuition fees, and compulsory schooling laws have been widely used in the literature card1999causal.

Returns of College Attendance

As the first example, we focus on the returns to college attendance. The following is the prompts we use. We choose $K_{0}=40$ and let GPT4 choose $K$. We explicitly request separate lists for individual factors and school factors.

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{1-1}[Example: Returns to College]\end{myprompt} you are a high school graduate. you need to make a college attendance decision. what would be factors (factors of schools and factors of yourself) that can determine your decision but that do not directly affect your future earnings, except through college attendance (that is, that affect your earnings only through college attendance)? list forty factors that are quantifiable, twenty for school factors and twenty for factors of yourself. explain the answers.

}}

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{2$_{x}$-1}[Example: Returns to College]\end{myprompt} suppose you are a student with family income \$10K per year, who is asian female from california, whose parents have college education, who is catholic. among the forty factors listed above, choose all factors that are not associated with your innate ability and personality and school quality, which determine earnings. create separate lists for school factors and factors of yourself. for each factor chosen, explain your reasoning.

}}

table[table omitted — 3,811 chars of source]

Table (ref) presents the results from a single session of running Prompts (ref) and (ref). It contains IVs suggested by GPT4 and GPT4's rationale for the suggestions.\footnote{GPT4 also provide the summary of overall rationale, which is not reported here for succinctness.} In the table, we find IVs that are already popular in the literature (e.g., \#1, 3, 5) as well as IVs that seem to be new (to our best knowledge) (e.g., \#6, 7, 9, 11, 12, 13, 14). The latter have potential to be valid, especially after being conditioned on additional covariates that are not considered in the prompt. Producing all these results took less than one minute in total. The rationale given by GPT4 can be elaborated further by requesting it in the same session, which we do not present here for brevity.\footnote{Appendix (ref) explains an effective way of adjusting the length of responses via a system message.}

Returns to Years of Education

As a second example, we consider the returns to years of schooling. We omit the prompts as they are similar as before, except that we impose the following role: “you are a student beginning high school in the united states. you need to make a decision on how many more years you will stay in school.” Table (ref) contains the IVs suggested by GPT4. Some candidates (e.g., \#3, 4) are similar to those found in the first example. Notably, GPT4 also finds variables (e.g., \#1, 2) that are related to compulsory schooling laws, the popular IVs in the literature. Local regional characteristics (e.g., \#5, 6, 7) are also interesting findings.

table[table omitted — 3,077 chars of source]

Supply and Demand

Production Function Estimation

Consider estimating a production function that captures the causal relationship between inputs and outputs. The key identification challenge is that input decisions can be correlated with unobserved productivity shocks, which directly influence outputs. To address this, IVs such as input prices have been proposed in the literature griliches1998production, which have been subsequently criticized olley1996dynamics,levinsohn2003estimating,ackerberg2015identification.

Here are the prompts we use. As before, we choose $K_{0}=40$ and let GPT4 choose $K$; we explicitly request separate lists for market factors and firm and manager factors. Note that in Prompt (ref), we use a loose description of covariates (unlike in Prompt (ref) where we assign specific values).

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{1-2-1}[Example: Production Functions]\end{myprompt} you are a manager at a manufacturing firm. you need to make a decision on how much labor and capital inputs to use to produce outputs. what would be factors (factors of markets and economy and factors of yourself) that can determine your decision but that do not directly affect your output productions, except through the input choices (that is, that affect your firm's outputs only through inputs)? list forty factors that are quantifiable, twenty for market factors and twenty for managerial factors. explain the answers.

}}

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{2$_{x}$-2-1}[Example: Production Functions]\end{myprompt} suppose you are a manager at a firm with specific level of capital intensity and specific scale of operations, which has a specific market share in a specific industry. among the forty factors listed above, choose all factors that are not influenced by productivity shocks of your firm, which determine outputs. create separate lists for market factors and managerial factors. for each factor chosen, explain your reasoning.

}}

Table (ref) presents the results from a single session of running Prompts (ref) and (ref). It contains IVs suggested by GPT4 and GPT4's rationale. Interestingly, IVs that are suggested in the literature (i.e., input prices) are not chosen by GPT4 although they appear in the answer to Prompt (ref) (not shown here for brevity).\footnote{This result was consistent over multiple runs.} This suggests that these IVs are not deemed by GPT4 to satisfy (ref), aligning with similar concerns in the literature olley1996dynamics,levinsohn2003estimating,ackerberg2015identification. However, GPT4 suggests IVs that may influence input prices (e.g., \#1, 2, 3, 5, 6, 14), some of which can be arguably exogenous. There are a handful of other IVs suggested as market-related and managerial factors. Among the latter, there are variables related to long-term decisions of the firm, which are argued by GPT4 to not influence short-term productivity shocks. However, long-term decisions affect long-term outputs, which may or may not be relevant to the short-term outputs of concern. Overall, the explanations given by GPT4 are more detailed than those in Table (ref), reflecting the random nature of the LLM's responses.

table[table omitted — 3,360 chars of source]
table[table omitted — 2,408 chars of source]

Demand Estimation

Consider estimating demand for consumers in a given market. In estimating the effect of price on demand, the main concern is that price is endogenous, as it is an equilibrium outcome. Researchers have used supply-side IVs that are excluded from the demand equation in a simultaneous system for supply and demand angrist2000interpretation or motivated by structural models berry1995automobile.

Below are the prompts we use, focusing on the setting of angrist2000interpretation. As before, we choose $K_{0}=40$ and let GPT4 choose $K$. Note that in Prompt (ref), covariate information is omitted.

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{1-2-2}[Example: Demand]\end{myprompt} you are a dealer at a fish market. you need to set the prices of fish. what would be factors that can determine your decision but that do not directly affect the customers' demand for fish, except through the price you set (that is, that affect the demand only through fish prices). list forty factors that are quantifiable. explain your answer.

}}

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{2-2-2}[Example: Demand]\end{myprompt} suppose you are a dealer at the fish market who is selling fish and setting its prices on a day of the week. among the factors listed above, choose all factors are not influenced by fish market conditions or customers' characteristics that determine demand for fish. for each factor chosen, explain your reasoning.

}}

table[table omitted — 4,055 chars of source]

Table (ref) presents the results from a single session of running Prompts (ref) and (ref). It contains IVs suggested by GPT4 and GPT4's rationale. Many supply-side factors (e.g., costs) are chosen by GPT4, which are reasonable candidates for IVs. The “weather conditions” variable used as IVs in angrist2000interpretation appears in the list (\#3). Interestingly, another supply-side factor “labor costs” produced by Prompt (ref) is not included in the final list produced by Prompt (ref). When asked “explain why you didn't include “labor costs” in the final list,” GPT4 responded that “labor costs are somewhat flexible and can be adjusted in response to changes in market conditions and customer demand, making them more dynamic than some of the other factors listed.”

Peer Effects

Suppose we are interested in the causal effects of peers on an individual's outcomes within a social network. We consider two well-known examples: (i) the effects of peer farmers on the adoption of new farming technologies foster1995learning,conley2010learning; (ii) the effects of peers on teenage smoking gaviria2001school. In both examples, the main source of endogeneity is latent factors that determine the formation of network (e.g., latent homophily). To address this, the literature on peer effects sometimes uses friends of friends as IVs bramoulle2009identification,angrist2014perils. In both examples, we construct prompts similar to the first two examples, except that we choose $K_{0}=20$.

Effects of Peer Farmers on New Technology Adoption

Here are the prompts. It is worth noting that, in this example, the role-playing is done from the peer's perspective, rather than from the perspective of the individual whose outcome is of concern.

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{1-3-1}[Example: Peer Effects on Technology Adoption]\end{myprompt} you are a farmer in a village in rural india. you want to influence your peer farmers in the same village to introduce a new farming technologies that you introduced. what would be factors (factors of farming and village, and factors of yourself) that can determine your influence on peers but that do not directly affect your peers' technology adoption decisions, except through your influence (that is, that affect your peers' decisions only through your influence)? list twenty factors that are quantifiable. explain your answer.

}}

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{2$_{x}$-3-1}[Example: Peer Effects on Technology Adoption]\end{myprompt} suppose you are a 40 year old male farmer of a specific crop in the village in rural india. among the factors listed above, which factors are not influenced by factors (e.g., similar background and preferences) that brought you and your peers in the same neighborhood and social network from the first place? for each factor chosen, explain your reasoning.

}}

Table (ref) presents the results in a single session of running Prompts (ref) and (ref). It contains IVs suggested by GPT4 and GPT4's rationale. Interestingly, some IVs that are suggested in the literature (e.g., friends of friends) are not chosen by GPT4. It can be because GPT4 either views them as invalid or is incapable of identifying them. On the other hand, \#10 seems to relate to the IV used in conley2010learning, which exploits variation in the presence of experienced farmers. Additionally, there are other IVs that seem to be new, notably \#9. Finally, GPT4 fails to identify the IV used in foster1995learning, namely, endowed land size. The validity of this variable is justified in their paper by the specific historical backgrounds of the Indian villages studied. This implies that, for some IVs, providing LLMs with institutional details can be necessary; see Appendix (ref) for implementing this via system messages.

table[table omitted — 3,831 chars of source]

Effects of Peer Teenagers on Smoking Behavior

Here are the prompts. Again, in this example, the role-playing is done from the peer's perspective. We choose a teenager in urban Indonesia for its relevance, given that the teenage smoking rate in Indonesia has been recently reported as one of the world's highest fithria2021indonesian. We consider a social media network in the scenario to illustrate the effectiveness of our approach in exploring relatively recent topics in the literature, thereby highlighting the potential to discover novel IVs.

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{1-3-2}[Example: Peer Effects on Smoking]\end{myprompt} you are a teenager in indonasia who smokes. you want to influence your peers in your social media network to smoke. what would be factors (factors of social media, your school and region, and factors of yourself) that can determine your influence on peers but that do not directly affect your peers' smoking decisions, except through your influence (that is, that affect your peers' decisions only through your influence)? list twenty factors that are quantifiable. explain your answer.

}}

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{2$_{x}$-3-2}[Example: Peer Effects on Smoking]\end{myprompt} suppose you are a teenage boy in urban indonesia who goes to high school and is from middle-income family. among the factors listed above, which factors are not influenced by factors (e.g., similar background and preferences) that brought you and your peers in the same social network from the first place? for each factor chosen, explain your reasoning.

}}

table[table omitted — 3,684 chars of source]

Table (ref) presents the results from a single session of running Prompts (ref) and (ref). It contains IVs suggested by GPT4 and GPT4's rationale. It is important to note that, given that the prompts are written from the perspective of peers, the variables in the table should be understood as factors influencing peers of the focal individual. Given that the setup incorporates modern elements such as social media, we identify many potentially new and interesting IVs, particularly from the social media category (i.e., \#1, 2, 3, 4, 7). Interestingly, \#7 can be viewed as a “friends of friends” IV.\footnote{There were other “friends of friends” IVs that are produced from Step 1 but did not survive Step 2.}

Adversarial Large Language Models

One way to refine the answers of an LLM is to request another LLM (or a different session of the same LLM) to play the role of an adversary and review the responses produced by the first LLM, namely the defender LLM. In the adversarial stage, we fully disclose the IV discovery task and ask the adversarial LLM to provide counter-arguments, but without using econometric jargon. These counter-arguments are then given to the defender LLM, which is asked to refine the previous answers. Overall, this process produces more sophisticated responses. We find this to be a useful exercise, which mimics the mental process of a human researcher.

The following is an example prompt given to the adversarial LLM. In the prompt, {[}the list of variables by Defender{]} is the list provided the defender LLM, generated from running, for example, Prompts (ref)--(ref).

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{AD1}\end{myprompt} you are a researcher who wants to find instrumental variables to estimate the effect of {[}treatment{]} on {[}outcome{]}. below is a list of candidate instrumental variables. for each variable in the list, provide arguments as to why it may not be a valid instrument: \\ \\ {[}the list of variables by Defender{]}

}}

Then, the counter-arguments generated from Prompt (ref) are presented to the defender LLM using the following prompt, which follows Prompts (ref)--(ref).

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{3}\end{myprompt} below are counter-arguments for each of your previous answers. based on these arguments, revise your selection and provide a list: \\ \\ {[}the arguments by Adversary{]}

}}

When running Prompt (ref), it is important to encourage the defender to create a list, otherwise, it might become pessimistic and reject all the previous selections.

We apply this adversarial process to our examples in Section (ref). Overall, the defender provides more sophisticated answers in the end. In the “Returns to College” example (Section (ref)), the variables related to geographical proximity or transportation have been retained. The other variables related to school facilities (e.g., library size, technology integration in classrooms) have been modified to factors that may better satisfy exogeneity (e.g., the percentage of renewable energy used on campus, the number of green spaces on campus, campus medical facilities). New variables also appeared (e.g. language spoken at home). Interestingly, many variables now appear to be weak IVs, which is sensible. In the “Returns to Years of Education” examples (Section (ref)), the “state laws” variable survived, but was renamed to the more descriptive “mandatory minimum years of schooling required by state.” Other variables were also further refined. “GED program availability” emerged. In the “Demand Estimation” example (Section (ref)), the “weather condition” variable was refined to “global climate patterns (e.g., El Ni\ {n}o),” because the counter-argument was that “bad weather can also discourage customers from coming to the market.” Many other variables were refined to reflect global and macroeconomic conditions. Again, these variables suffer from being weak IVs or having less individual variation.

Variables Search in Other Causal Inference Methods

In this section, we demonstrate how prompting strategies similar to those for the IV discovery can be used to find (i) control variables under which treatments are conditionally independent (i.e., exogenous); (ii) control variables under which parallel trends are likely to hold in difference-in-differences; and (iii) running variables in regression discontinuity designs.

Conditional Independence

Using the same notation as in Section (ref), consider a conditional independence (CI) assumption that assigns a more crucial role to the vector of control variables $X\equiv(X_{1},...,X_{L})$:

myas{CI} For any $d$, $D\perp Y(d)|X$.

Assumption (ref) is commonly introduced in causal inference settings, especially when combined with machine learning to estimate nuisance functions; e.g., debiased/double machine learning methods chernozhukov2024applied. More traditionally, this assumption is closely related to matching and propensity score matching techniques heckman1998matching. The mean independence version of (ref) (i.e., $E[Y(d)|D,X]=E[Y(d)|X]$) is relevant to regression methods.

We propose using LLMs to systematically search for $X$ that satisfies a verbal version of (ref). The prompt writing is slightly simpler than that for IVs. In particular, we construct prompts that solicit the relationship between $X$ and $D$ (Step 1) and $X$ and $Y(d)$ (Step 2). Therefore, only the second-step prompt involves a counterfactual statement. Let $L_{0}$ be the number of controls to be found in Step 1 ($L_{0}\ge L$). One may want to choose the value of $L_{0}$ to be larger than one would normally use for $K_{0}$ and leave $L$ unspecified.

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{C1}[Search for Control Variables]\end{myprompt} you are {[}agent{]} who needs to make a {[}treatment{]} decision in {[}scenario{]}. what factors determine your decision? list {[}L_0{]} factors that are quantifiable. explain the answers.

}}

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{C2}[Refine Control Variables]\end{myprompt} among the {[}L_0{]} factors listed above, choose all factors that directly determine your {[}outcome{]}, not only indirectly through {[}treatment{]}. the chosen factors can still influence your {[}treatment{]}. for each factor chosen, explain your reasoning.

}}

The prompts are constructed to search for confounders and need to be controlled for. Researchers sometimes mistakenly control for “colliders” and/or “mediators” pearl2000causality, which are intended to be excluded from the search. Note that Prompts (ref)--(ref) can also be adapted to jointly search for covariates and latent confounders in the IV search. In this case, one can distinguish $X$ from latent confounders by referring to the former as “quantifiable.” Also, one may want to use the phrase “demographic factors” to refer to $X$, as they are common control variables in many empirical applications.

Difference in Differences

The difference-in-differences (DiD) method is popular in empirical research, partly due to the simplicity and intuitiveness of its main assumption, namely, the parallel trend assumption (stated below). However, this assumption is not directly testable and typically hard to justify ghanem2022selection,rambachan2023more. It is believed that conditioning on the right control variables can make this assumption more justifiable, which can motivate the search for such controls.

myas{PT}$E[\Delta Y(0)|D,X]=E[\Delta Y(0)|X]$ where $\Delta Y(0)\equiv Y_{after}(0)-Y_{before}(0)$.

Assumption (ref) can be viewed as a mean independence version of (ref), where the counterfactual outcome is replaced with the temporal difference of counterfactual (untreated) outcomes before and after the event. Therefore, Prompts (ref)--(ref) can be directly used to search for $X$ that satisfy a verbal version of (ref). This can be done by inputting “average temporal changes in {[}outcome_t{]} during the time of no {[}treatment{]}” for {[}outcome{]} in Prompt (ref), where {[}outcome_t{]} refers to $Y_{t}$ for $t\in\{before,after\}$. The example of such prompts is constructed to revisit the classical empirical example, namely, the effects of minimum wage on the fast food industry's labor markets card1994minimum; see Appendix (ref) for the actual prompts. Table (ref) contains the control variables suggested by GPT4, conditional on which the parallel trend is likely to hold, and GPT4's rationale. On the list, \#3, 4, 7, 10, 11 are particularly interesting and \#11 seems particularly novel. In the table, the first four rows (\#1, 2, 3, 4) are chosen by GPT4 from an additional prompt that emphasizes the requirement with respect to $\Delta Y(0)$: “be sure to choose all factors that do not determine the average wage level but only determine the temporal changes in average wages.” Nonetheless, controls that satisfy the mean version of (ref) with the level, $Y_{t}(0)$ for $t\in\{before,after\}$, are also valid controls for (ref).

table[table omitted — 3,508 chars of source]

Regression Discontinuity

Regression discontinuity designs (RDDs) are another well-known method for causal inference that closely relates to the IVs method lee2010regression.\footnote{For example, the fuzzy RDD estimand can be viewed as the two-stage least squares estimand.} The key for this method to work is to find a running variable (i.e., assignment variable) that satisfies the following:

myas{RD} There exists a variable $R_{j}$ and a cutoff $r_{0}$ such that $D=1$ if $R_{j}\ge r_{0}$ and $D=0$ if $R_{j}<r_{0}$.

One can use LLMs to systematically search for running variables $\{R_{1},...,R_{J}\}$ for a given $D$ and $Y$ of interest. We provide the example of prompts here. It is worth noting that, unlike in all the previous cases, none of the prompts below involve counterfactual statements. Therefore, if LLMs outperform a traditional search for running variables, it would be due to their automated and comprehensive search behavior.\footnote{In further refining the candidates of running variables to ensure that RRD's continuity assumptions are satisfied, counterfactual prompting would be necessary; see Appendix (ref).} Similarly as above, we only specify initial $J_{0}$ and leave $J$ unspecified.

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{R1}[Search for Running Variables]\end{myprompt} you are {[}agent{]} who needs to make a {[}treatment{]} decision in {[}scenario{]}. what would be the possible criteria based on which your eligibility for {[}treatment{]} is determined? provide {[}J_0{]} of the most relevant criteria that are (1) quantifiable and (2) have specific cutoffs determining eligibility. explain the answers.

}}

{\fboxsep 10pt\fbox{

minipage[t]{1\columnwidth - 2\fboxsep - 2\fboxrule} \begin{myprompt}{R2}[Refine Running Variables]\end{myprompt} among the {[}J_0{]} criteria listed above, choose all criteria that involve continuous or ordered measures and have precise cutoffs determining eligibility. also report the cutoff value for each criterion from verifiable sources only (ensuring no fabricated or hypothetical numbers are used). explain the answers.

}}

Note that when Prompt (ref) is run on GPT4, it will engage in a series of automated web searches. The request for cutoff values may lead the LLM to provide hypothetical numbers as possibilities. When one wants to get the actual values from verifiable sources, it is important to explicitly state that, as we do above. We apply Prompts (ref)--(ref) to a range of famous examples in the literature where RDDs are used as empirical strategies. Table (ref) presents the results obtained by running the prompts, which are adapted to each specific context and country of the empirical example. In most cases, a handful of new possible running variables are suggested by GPT4 with specific cutoffs obtained from web sources. Except for one case (i.e., \#4), GPT4 also identifies the running variables used in the literature.

table[table omitted — 3,919 chars of source]

Conclusions

This essay proposes the agenda to use LLMs to systematically search for variables in designing causal inference. It merely serves as a starting point, and there are potential next steps that can follow. In constructing prompts for IVs, there are many possible ways for sophistication: First, one can consider using previously known IVs in the literature to guide LLMs to discover new ones. This can be done by adding textual demonstration of how Assumptions (ref)--(ref) are satisfied with known IVs before starting the proposed prompts. This approach would evoke few-shot learning in LLMs brown2020language, which can enhance their performances. This approach would also “orthogonalize” the search ludwig2024machine to focus on novel IVs. Second, the elaborated search can be directed toward finding IVs that are more policy-relevant imbens1994identification,heckman2005structural) by specifying targeted policies in the prompt. Third, none of the results reported in the current essay are findings aggregated across sessions. To account for and potentially leverage the stochastic nature of LLMs' responses, exploring the possibility of aggregation (e.g., taking the union or intersection of $\mathcal{Z}_{K}$'s across sessions) would be beneficial. Fourth, we can explore the use of multiple LLM agents each assuming distinct roles in the discovery process, analogous to the collaborative and critical interactions among human researchers. An example of this approach is discussed in Section (ref), where the responses of one LLM are reviewed and critiqued by another acting as a critic.

Additionally, we can consider having a horse race among multiple LLMs or using an open-source LLM to fine-tune it du2024labor for our purpose. A potential challenge is that the performance metric is hard to define in our context due to the lack of ground truth for valid IVs. In fact, this is the very reason we propose using LLMs from the first place: for any IVs found by human researchers or the machine, there are only more compelling narratives or less compelling ones. In later stages, when data eventually come into play, over-identification tests can potentially be a fruitful framework for the evaluation of LLMs. More broadly, it would be interesting to apply the proposed approach of variable search in other empirical examples and other causal inference methods.