EconBase
← Back to paper

Mining Causality: AI-Assisted Search for Instrumental Variables

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

91,847 characters

Mining Causality: AI-Assisted Search for Instrumental Variables


\title{Mining Causality:\\
AI-Assisted Search for Instrumental Variables\thanks{I appreciate Guido Imbens for his thoughtful feedback and encouragement.
I also thank Susan Athey, Orazio Attanasio, Kevin Chen, Phil Haile,
Hide Ichimura, Pat Kline, Michal Kosinski, Lihua Lei, Benjamin Manning,
Mark Rosenzweig, Jesse Shapiro and Jann Spiess and participants at
the Causal Data Science Meeting 2024 for helpful discussions.}}
\author{Sukjin Han\thanks{School of Economics, University of Bristol. \UrlFont\protect[email removed]}}
\date{March 8, 2025}

\maketitle
\vspace{-0.2cm}

\begin{abstract}
The instrumental variables (IVs) method is a leading empirical strategy
for causal inference. Finding IVs is a heuristic and creative process,
and justifying its validity---especially exclusion restrictions---is
largely rhetorical. We propose using large language models (LLMs)
to search for new IVs through narratives and counterfactual reasoning,
similar to how a human researcher would. The stark difference, however,
is that LLMs can dramatically accelerate this process and explore
an extremely large search space. We demonstrate how to construct prompts
to search for potentially valid IVs. We contend that multi-step and
role-playing prompting strategies are effective for simulating the
endogenous decision-making processes of economic agents and for navigating
language models through the realm of real-world scenarios, rather
than anchoring them within the narrow realm of academic discourses
on IVs. We apply our method to three well-known examples in economics:
returns to schooling, supply and demand, and peer effects. We then
extend our strategy to finding (i) control variables in regression
and difference-in-differences and (ii) running variables in regression
discontinuity designs.

\vspace{0.1in}

\noindent \textit{Keywords:} Causal inference, instrumental variables,
exclusion restrictions, artificial intelligence, large language models.

\noindent \textit{JEL Codes:} C26, C36, C5.
\end{abstract}

\section{Introduction\label{sec:Introduction}}

Endogeneity is the major obstacle in conducting causal inference
in observational settings. Since the credibility revolution \citep{angrist2010credibility}
and the causal revolution \citep{pearl2000causality}, researchers
in social science, statistics and other adjacent fields have developed
various identification strategies to overcome endogeneity by restoring
versions of quasi-experiments. A leading strategy is the instrumental
variables (IVs) method. Over decades, researchers with their ingenuity
have discovered IVs in various settings and justified their satisfaction
of \emph{exclusion restrictions} (e.g., IVs are conditionally exogenous
of latent variables). With its various applicability, the IVs method
has prevailed across all subfields of economics and beyond \citep[e.g.,][]{imbens1994identification,Blundell:2003wi,heckman2005structural,hernan2006instruments}.\footnote{See \citet{mogstad2024instrumental} for a more recent survey.}

Exclusion restrictions are fundamentally untestable assumptions.\footnote{An exception is a favorable situation where one enjoys over-identifying
restrictions. We discuss this point in our context below. Unlike
the exclusion restriction, the IV relevance is testable from data
\citep{stock_yogo_2005,olea2013robust}.} Often, in justifying them, researchers resort to \emph{rhetorical}
arguments specific to each setting. This non-statistical process follows
the discovery of potential candidate IVs, which itself requires researchers'
\emph{counterfactual reasoning} and creativity---and sometimes luck.
These elements all contribute to the heuristic processes employed
by human researchers.

We demonstrate that large language models (LLMs) can facilitate the
discovery of new IVs. Considering that narratives are the primary
method of supporting IV exclusion, we believe that LLMs, with sophisticated
language processing abilities, are well-suited to assist in the search
for new valid IVs and justify them rhetorically, just as human researchers
have done for decades. The stark difference, however, is that LLMs
can accelerate this process at an exponentially faster rate and explore
an extremely large search space, to an extent that human researchers
cannot match. It is now recognized that artificial intelligence (AI)
shows remarkable performances in conducting systematic searches for
hypotheses and refining the search \citep[e.g.,][]{jumper2021highly,ludwig2024machine}.
Furthermore, LLMs are argued to be capable of conducting counterfactual
reasoning---or, perhaps more precisely, exploring alternative scenarios---which
makes them a promising tool for causal inference.

There are at least four benefits to pursuing this AI-assisted approach
to discovering IVs. First, researchers can conduct a systematic search
at a speedy rate, while adapting to the particularities of their
settings. Second, interacting with AI tools can inspire ideas for
possible domains for novel IVs. Third, the systematic search could
increase the possibility of obtaining multiple IVs, which would then
enable formal (i.e., statistical) testing of their validity via over-identifying
restrictions. Fourth, having a list of candidate IVs would increase
the chances of finding actual data that contain IVs or guide the construction
of such data, including the design of experiments to generate IVs.

We show how to construct prompts in a way that guides LLMs to search
for candidates for valid IVs. The text representation of exclusion
restrictions (among others) is the main component of the prompts.
We propose a multi-step approach in prompting that divides a discovery
task into multiple subtasks, and thus separates counterfactual statements
of different complexities. At the same time, we propose using role-playing
prompts, arguing that they align with the very source of endogeneity,
namely, agents' decisions.\footnote{Decisions of economic agents have been at the root of challenges for
causal analyses in econometrics \citep[e.g., ][]{heckman1979sample,manski1993identification}.} By doing so, we equip LLMs with the perspective of agents, enabling
them to mimic agents' endogenous decision-making processes and gather
contextual information in realistic scenarios. This approach also
makes it convenient to impose statistical conditioning that qualifies
the characteristics of the agent. Another benefit of multi-stage,
role-playing prompts is that they help navigate language models through
the realm of real-world scenarios, rather than anchor them within
the narrow realm of academic discourses on IVs. Each stage's prompt
focuses only on a portion of the IV assumptions, translated into an
agent's real-world problem, thereby minimizing the likelihood that
the LLM perceives the task as a search for IVs.\footnote{Even if LLMs exhibit \emph{memorization} from academic texts, we still
find value in the procedure as long as the list of discovered IVs
includes those that are recognized as new by researchers.}

To prove the proposed concept and illustrate the actual performance
of an LLM, we conduct discovery exercises using OpenAI's ChatGPT-4
(GPT4), one of the leading LLMs, to find IVs in three well-known examples
in empirical economics: returns to schooling, supply and demand, and
peer effects. In all three examples, GPT4 produced a list of candidate
IVs, some of which appear to be new in the literature and provided
rationale for their validity. The list also contains IVs that are
popularly used in the literature. Our initial assessment of the results
suggests that the proposed method can work in practice. In the peer
effect example, we also demonstrate that the proposed method can be
effective in exploring relatively new topics for empirical research,
which may in turn increase the possibility of finding novel IVs.

From a broader perspective, the proposal is to systematically ``search
for exogeneity.'' We extend the exercise to other causal inference
methods: (i) searching for control variables in regression and difference-in-differences
methods and (ii) searching for running variables in regression discontinuity
designs. We construct relevant prompts and run them in well-known
examples in the literature.

A list of candidate IVs or control variables produced as a result
of the proposed method is not absolute. Rather, we hope that it serves
as a valuable benchmark that inspires empirical researchers about
which types of variables to consider and which domains to explore.
The dialogue carried out with LLMs in the process can also help researchers
solidify arguments or counterarguments for the validity of variables.
After all, AI---like any machines---cannot be the ultimate authority
(at least not yet). We believe a human researcher assisted by AI can
choose research designs and conduct causal inference more effectively.

This essay contributes to a recent agenda in the social science literature
on using AI to assist creative and heuristic parts of human research
processes. This agenda views machine learning and AI as not only data-processing
and prediction tools for economic research (\citealp{mullainathan2017machine,athey2019machine}),
but also as tools that can improve conventional research practices
themselves. In very interesting work, \citet{ludwig2024machine} use
generative models to systematically produce hypotheses that are comprehensible
by humans in otherwise daunting settings. They make progress in research
areas where the use of AI has been limited because, as they argue,
establishing causal relationships in social science is an ``open
world'' problem, unlike ``closed world'' problems in physical science.\footnote{The latter can be viewed as extremely difficult computation problems
where machine learning makes significant progress; e.g., detecting
new proteins using AlphaFold \citep{jumper2021highly} or advances
in particle physics and cosmology using machine learning \citep{carleo2019machine}.} In related work, \citet{mullainathan2024predictive} use predictive
(neural network) algorithms to recover old anomalies and discover
new ones in economic theory models. We do not attempt to generate
hypotheses, although the new variables discovered implicitly maintain
a range of hypotheses on their validity.

LLMs has only very recently been used in social science research.
Notably, \citet{du2024labor} use fine-tuned LLMs (Meta's LLaMA in
particular) to predict job transitions and understand career trajectories
in labor economics. They show that the prediction accuracy remarkably
outperforms those from traditional job transition economic models.
\citet{manning2024automated} propose to use LLMs to automate the
entire process of social scientific research, from data generation
to testing causal hypotheses. We employ LLMs in statistical causal
inference by incorporating specific structure from econometric assumptions
and allowing for human intervention in discovery processes.

This paper also relates to the approach of using LLMs in causal discovery
(\citet{ban2023causal,cohrs2024large,jiralerspong2024efficient,le2024multi,long2023can,takayama2024integrating});
also see \citet{wan2024bridging} for a recent survey and references
therein. However, the fundamental difference of our approach to this
line of work is that we use LLMs to systematically discover variables
with particular causal structure rather than using LLMs to find causal
links among a \emph{pre-determined} set of variables.

The paper is organized as follows. Section \ref{sec:Notation-and-IV}
states the IV assumptions and Section \ref{sec:Prompt-Construction}
proposes the main idea of IV discovery along with the prompting strategies.
Section \ref{sec:Discovered-IVs} provides the examples of discovered
IVs. Sections \ref{sec:Adversarial-Large-Language}--\ref{sec:Variables-Search-in}
contain extensions: (i) the use of an adversarial LLM to review and
refine the discovery process and (ii) the extension of the paper's
approach to other causal inference settings. Section \ref{sec:Conclusions}
concludes.

\section{Notation and IV Assumptions\label{sec:Notation-and-IV}}

We first formally state our discovery goal. Let $Y$ be the outcome
of interest, $D$ be the potentially endogenous treatment, $\mathcal{Z}_{K}\equiv\{Z_{1},...,Z_{K}\}$
be the list of IVs $Z_{k}$'s with $K$ being the desired number of
IVs to discover, and $X$ be the covariates. Let $Y(d,z_{k})$ be
the counterfactual outcome given $(d,z_{k})$. Let ``$\perp$''
denote statistical independence. We say $Z_{k}$ is a valid IV if
it satisfies the following two assumptions:

\begin{myas}{REL}[Relevance]\label{as:REL}Conditional on $X$, the
distribution of $D$ given $Z_{k}=z_{k}$ is a nontrivial function
of $z_{k}$.\end{myas}

\begin{myas}{EX}[Exclusion]\label{as:EX}For any $(d,z_{k})$, $Y(d,z_{k})=Y(d)$.\end{myas}

\begin{myas}{IND}[Independence]\label{as:IND}For any $d$, $Y(d)\perp Z_{k}$
conditional on $X$.\end{myas}

The goal of our exercise is to search for IVs that satisfy Assumptions
\ref{as:REL}, \ref{as:EX} and \ref{as:IND}.\footnote{One can consider a weaker version of \ref{as:IND} (i.e., mean independence
and nonzero correlation). Although we do not believe our ultimate
findings significantly differ from this relaxation, our prompts can
reflect it.} Suppressing $X$, Figure \ref{fig:DAG} depicts the causal direct
acyclic graph (DAG) that implies \ref{as:REL}, \ref{as:EX} and \ref{as:IND}
and with $Y(d)$ being a transformation of latent confounders $U$.
This diagram is useful in describing our procedures.
\begin{figure}
\begin{centering}
\begin{tikzpicture}
\tikzstyle{node} = [circle, draw, minimum size=20pt, inner sep=0pt]
\tikzstyle{dashed_node} = [circle, draw, dashed, minimum size=20pt, inner sep=0pt]
\tikzstyle{arrow} = [->, thin]
\tikzstyle{dashed_arrow} = [->, thin, dashed]
\node[node] (Z) {$Z_k$};
\node[node, right=of Z] (D) {$D$};
\node[node, right=of D] (Y) {$Y$};
\node[dashed_node, right=0.13cm of D, yshift=0.9cm] (U) {$U$};
\draw[arrow] (Z) -- (D);
\draw[arrow] (D) -- (Y);
\draw[dashed_arrow] (U) -- (D);
\draw[dashed_arrow] (U) -- (Y);
\end{tikzpicture}
\vspace{3pt}
\par\end{centering}
\caption{Causal DAG for a Valid IV ($X$ suppressed)}
\label{fig:DAG}

\end{figure}


\section{Prompt Construction\label{sec:Prompt-Construction}}

We propose a two-step approach for IV discovery. In Step 1, we prompt
an LLM to search for IVs that satisfy a verbal description of \ref{as:REL}
and \ref{as:EX} (i.e., \raisebox{-0.3\height}{
\begin{tikzpicture}[scale=0.5]
\tikzstyle{node} = [circle, draw, minimum size=20pt, inner sep=0pt]
\tikzstyle{arrow} = [->, thin]
\node[node] (Z) {$Z_k$};
\node[node, right=0.5cm of Z] (D) {$D$};
\node[node, right=0.5cm of D] (Y) {$Y$};
\draw[arrow] (Z) -- (D);
\draw[arrow] (D) -- (Y);
\end{tikzpicture}
}). In Step 2, we prompt the LLM to refine the search by selecting---among
the IVs found in Step 1---those that satisfy a verbal description
of \ref{as:IND} (i.e., \raisebox{-0.3\height}{
\begin{tikzpicture}[scale=0.5]
\tikzstyle{node} = [circle, draw, minimum size=20pt, inner sep=0pt]
\tikzstyle{dashed_node} = [circle, draw, dashed, minimum size=20pt, inner sep=0pt]
\tikzstyle{arrow} = [->, thin]
\node[node] (Z) {$Z_k$};
\node[dashed_node, right=0.5cm of Z] (U) {$U$};
\end{tikzpicture}
}). In both steps, the prompts will involve counterfactual statements.
In each step, we ask the LLM to provide rationale for its responses.
This feature is useful for the user to understand the LLM's reasoning.
The two steps can be conducted in the same session or in separate
sessions.  However, when submitting different queries, we recommend
that each two-step query be conducted in a separate session to avoid
interference across queries. In Appendix \ref{subsec:Three-Step-Prompts},
we also present a full three-step prompting that focuses on each of
Assumptions \ref{as:REL}, \ref{as:EX} and \ref{as:IND} in each
step.

We propose a multi-step approach for several reasons: First, LLMs
are known to yield better performance when handling subtasks step-by-step,
focusing on important details in interpreting the prompts and avoiding
errors \citep{wu2022ai}. Second, this approach creates more room
for the user to inspect intermediate outputs, facilitating the evaluation
of final outputs. In particular, Step 2 involves more complex counterfactual
statements than Step 1, allowing the user (and the LLM) to apply varying
degrees of attention when fine-tuning is needed. Third, intermediate
outputs themselves can provide information and offer insights. Finally,
this approach significantly reduces the likelihood that LLMs recognize
the task as IV discovery and generate text from relevant academic
sources.\footnote{This would be especially true when each step is conducted in a separate
independent session.}

Alongside the multi-step method, we propose a role-playing approach.
It has been reported that LLMs---including GPT4---gather better
contextual information and generate more tailored and unique responses
when prompts are structured as role-plays.\footnote{OpenAI Developer Forum: https://community.openai.com/t/make-chatgpt-better-for-roleplay-scenarios/344244}
In fact, in most scenarios, the explanatory variable $D$ represents
an economic agent's decision, which naturally facilitate role-playing.
Additionally, role-playing prompts are more effective in guiding LLMs
to respond as the relevant economic agent rather than as a researcher
searching for IVs. In Appendix \ref{subsec:Comparison-to-Direct},
we compare our multi-step, role-playing prompting strategy with a
more direct approach that explicitly states the goal of IV search,
arguing that the former is more effective.

To simplify the exposition, in Sections \ref{subsec:Step-1:-Prompts}--\ref{subsec:Step-2:-Prompts},
we first demonstrate the prompt construction without introducing covariates
(in which case \ref{as:REL} and \ref{as:IND} should hold unconditionally).
We then construct more realistic prompts with covariates in Section
\ref{subsec:Extension:-Prompts-to}. The prompts presented here can
serve as a benchmark for more sophisticated prompts; we discuss them
in Section \ref{sec:Conclusions}.


\subsection{Step 1: Prompts to Search for IVs\label{subsec:Step-1:-Prompts}}

For Step 1, Prompt \ref{P1} is a role-playing prompt that queries
the search for $K_{0}$ IVs (obtaining $\mathcal{Z}_{K_{0}}$) that
satisfy verbal versions of \ref{as:REL} and \ref{as:EX} (with no
$X$). In all prompts below, each bracketed term represents a user
input:\texttt{ {[}treatment{]}} is the treatment $D$, \texttt{{[}agent{]}}
is the economic agent whose decision is $D$, \texttt{{[}scenario{]}}
is the specific setting of interest, \texttt{{[}outcome{]}} is the
outcome $Y$, and \texttt{{[}K\_0{]}} is the desired number of variables
$K_{0}$. When prompting, we ask the LLM to play the role of \texttt{{[}agent{]}}
to make a \texttt{{[}treatment{]}} decision in a hypothetical \texttt{{[}scenario{]}}.
Examples of these inputs are given in Section \ref{sec:Discovered-IVs}.

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{1}[Search for IVs]\label{P1}\end{myprompt}\vspace{-0.3cm}

\texttt{you are {[}agent{]} who needs to make a {[}treatment{]} decision
in {[}scenario{]}. what are factors that can determine your decision
but do not directly affect your {[}outcome{]}, except through {[}treatment{]}
(that is, factors that affect your {[}outcome{]} only through {[}treatment{]})? list
{[}K\_0{]} factors that are quantifiable. explain the answers.}
\end{minipage}}}

\bigskip{}

There are at least two variants of Prompt \ref{P1} that may be useful
in certain scenarios. First, instead of ``\texttt{list {[}K\_0{]}
factors that are quantifiable}'' one may simply write ``\texttt{list
{[}K\_0{]} factors}'' or even ``\texttt{list {[}K\_0{]} factors
that are hard to quantify}.'' This would return candidates of IVs
that are harder to measure but can inspire creative data collection
(e.g., text, images, or other unstructured data). Second, one can
expand Prompt \ref{P1} to be more specific about categorizing factors
for relevant parties in a given setting. For example, in the schooling
scenario (Section \ref{subsec:Returns-to-Education}), we request
separate lists for student factors and school factors. This approach
can facilitate the user's evaluation of the results.

\subsection{Step 2: Prompts to Refine the Search for IVs\label{subsec:Step-2:-Prompts}}

Take the set of IVs, $\mathcal{Z}_{K_{0}}\equiv\{Z_{1},...,Z_{K_{0}}\}$,
obtained by running Prompt \ref{P1} in Step 1. Next, for Step 2 in
the same session, Prompt \ref{P2} is a role-playing prompt that queries
the search for $K$ IVs (obtaining $\mathcal{Z}_{K}$, $K\le K_{0}$)
within $\mathcal{Z}_{K_{0}}$ that satisfy a verbal version of \ref{as:IND}
(with no $X$). Below, \texttt{{[}confounders{]}} is the user input
for unobserved confounders of concern and \texttt{{[}K{]}} is the
user choice of $K$. In this prompt, we ask the LLM to continue playing
the same role as in Prompt \ref{P1}.\bigskip{}

\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{2}[Refine IVs]\label{P2}\end{myprompt}\vspace{-0.3cm}

\texttt{you are {[}agent{]} in {[}scenario{]}, as previously described. among
the {[}K\_0{]} factors listed above, choose {[}K{]} factors that are
most likely to be unassociated with {[}confounders{]}, which determine
your {[}outcome{]}. the chosen factors can still influence your {[}treatment{]}. for
each chosen factor, explain your reasoning.}
\end{minipage}}}

\bigskip{}
Unlike Prompt \ref{P1}, this prompt contains a statement about variables
typically unobserved to researchers, which may pose challenges. We
believe that incorporating the researcher's prior knowledge on latent
confounders helps simplify the overall search process and yield more
desirable results.\footnote{This relates to \emph{few-shot learning} discussed in Section \ref{sec:Conclusions}.}
For instance, in the schooling scenario, one can specify ``\texttt{innate
ability and personality and school quality}.'' Alternatively, if
the user prefers a more agnostic approach, they can list \texttt{{[}confounders{]}}
as ``\texttt{other possible factors}.'' Another option is to systematically
search for possible unobserved confounders; see Section \ref{subsec:Conditional-Independence}
for related prompting strategies. In Prompt \ref{P2}, we use the
term ``\texttt{unassociated}.'' If the LLM ever captures the nuance
of this word, it reflects the mean independence version of \ref{as:IND},
making the search easier. Interestingly, an alternative phasing such
as ``\texttt{choose {[}K{]} factors that are purely random}'', which
may seem a straightforward way to impose \ref{as:IND} without needing
to specify unobserved confounders, often fails to produce intended
outputs.

There are useful variants of Prompt \ref{P2}. First, one can omit
\texttt{{[}K{]}} and instruct the LLM to ``\texttt{choose all factors}''
from $\mathcal{Z}_{K_{0}}$ that are likely to satisfy \ref{as:IND},
allowing the LLM to determine $K$ independently; we apply this strategy
in all examples later. Second, as a sanity check, one can direct the
LLM to select elements in $\mathcal{Z}_{K_{0}}$ that \emph{violate}
\ref{as:IND} in addition to those that satisfy it. This can be achieved
by adding ``\texttt{also choose factors that are, in contrast, associated
with {[}confounders{]}}.'' To gain further insights, the user can
request explanations for factors that she identifies as valid IVs
in initial set $\mathcal{Z}_{K_{0}}$ from Step 1, but which are somehow
not included in the final set $\mathcal{Z}_{K}$ by the LLM. We apply
the last approach to the application in Section \ref{subsec:Demand-Estimation}.

\subsection{Extension: Prompts to Search and Refine with Covariates\label{subsec:Extension:-Prompts-to}}

Typically, IVs are argued to be valid after conditioning on a list
of covariates (as reflected in \ref{as:REL}--\ref{as:IND}). The
IV discovery with covariates can be approached in at least two different
ways. We can prompt the LLM to either (i) search for IVs conditional
on predetermined covariates; or (ii) jointly search for IVs and covariates
that satisfy \ref{as:REL}--\ref{as:IND}. We focus on option (i);
option (ii) is discussed in Appendix \ref{subsec:Alternative-Prompts-with}.
Whenever covariates are searched, option (i) can be viewed as initiating
an IV search in a new independent session with the searched covariates.

We construct a prompt that introduces the notion of conditioning variables;
role-playing prompts are suitable for this purpose. Here, we only
modify Prompt \ref{P2}. Although \ref{as:REL} also involves conditioning
on $X$, we find that results are not sensitive to a relevant modification
of Prompt \ref{P1}. Prompt \ref{P2x} qualifies \emph{both} \texttt{{[}agent{]}}
and \texttt{{[}scenario{]}} by \texttt{{[}covariates{]}}, the pre-determined
user choice of covariates. It extends Prompt \ref{P2} by modifying
the first sentence. Prompt \ref{P2x} is intended to be run after
completing Prompt \ref{P1}.\bigskip{}

\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{2$_{x}$}[Refine IVs with Covariates]\label{P2x}\end{myprompt}\vspace{-0.3cm}

\texttt{suppose you are {[}agent{]} in {[}scenario{]} with {[}covariates{]}. among
the {[}K\_0{]} factors listed above, choose {[}K{]} factors that are
most likely to be unassociated with {[}confounders{]}, which determine
your {[}outcome{]}. the chosen factors can still influence your {[}treatment{]}. for
each chosen factor, explain your reasoning.}
\end{minipage}}}\bigskip{}

The recommended approach for incorporating \texttt{{[}covariates{]}}
is to assign specific values for the covariates. For instance, in
the schooling scenario, one can write ``\texttt{suppose you are an
asian female high school student from california who considers attending
a private college}.''\footnote{One can run multiple queries across different values of covariates
for robustness, although this does not appear to be necessary in most
cases unless extreme values are assigned in the initial run.} Alternatively, one can simply use terms like ``\texttt{specific}''
or ``\texttt{particular}'' along with the name of chosen covariates
(e.g., ``\texttt{suppose you are a high school student with specific
gender, race, and regional origin who considers attending a college
of specific type}'').

\section{Discovered IVs\label{sec:Discovered-IVs}}

Using Prompts \ref{P1} and \ref{P2x} described in the previous section,
we aim to identify candidates for IVs in four well-known examples
in economics: returns to schooling, supply and demand, and peer effects.
These examples are chosen for their significance in the empirical
economics literature (representing labor economics, industrial organization,
and development economics, respectively). They commonly employ the
IVs method as an empirical strategy. The main purpose of this exercise
is to evaluate the performance of LLMs in executing the proposed method
and to demonstrate the practical applicability of the method.

To summarize the findings, in all the examples, LLMs appear to discover
new candidates for IVs and candidates that are related to well-known
IVs in the literature. When Also, many candidates demonstrate high
levels of specificity to their context. With the results produced,
we hope to spark debates and inspire the discovery of new and better
IVs.

The prompts we construct in each example slightly deviate from the
templates of Prompts \ref{P1} and \ref{P2x} to better adapt to the
scenario and enhance the flow of English language. For each example,
we present results from the \emph{initial} \emph{single} run of the
prompts without any curation or further refinement. When the results
match the IVs described in the literature exactly, we include the
corresponding references (to the best of our knowledge) and their
citation counts. Results across sessions are largely consistent, although
they can vary when different values of $K_{0}$ and $K$ are chosen.
We use GPT4 as our LLM.

\subsection{Returns to Education\label{subsec:Returns-to-Education}}

Suppose we are interested in estimating the causal effects of educational
attainment (e.g., college attendance, years of schooling) on earnings.
The main latent confounders in this setting is unobserved individual
and school characteristics (e.g., student ability and personality,
school quality) that affect both the schooling decisions and future
earnings. To address this endogeneity and recover meaningful causal
effects (e.g., local average treatment effects \citep{imbens1994identification}),
IVs such as distance to schools, tuition fees, and compulsory schooling
laws have been widely used in the literature \citep{card1999causal}.

\subsubsection{Returns of College Attendance\label{subsec:Returns-of-College}}

As the first example, we focus on the returns to college attendance.
The following is the prompts we use. We choose $K_{0}=40$ and let
GPT4 choose $K$. We explicitly request separate lists for individual
factors and school factors.

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{1-1}[Example: Returns to College]\label{P1-1}\end{myprompt}\vspace{-0.3cm}

\texttt{you are a high school graduate. you need to make a college
attendance decision. what would be factors (factors of schools and
factors of yourself) that can determine your decision but that do
not directly affect your future earnings, except through college attendance
(that is, that affect your earnings only through college attendance)? list
forty factors that are quantifiable, twenty for school factors and
twenty for factors of yourself. explain the answers.}
\end{minipage}}}

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{2$_{x}$-1}[Example: Returns to College]\label{P2x-1}\end{myprompt}\vspace{-0.3cm}

\texttt{suppose you are a student with family income \$10K per year,
who is asian female from california, whose parents have college education,
who is catholic. among the forty factors listed above, choose all
factors that are not associated with your innate ability and personality
and school quality, which determine earnings. create separate lists
for school factors and factors of yourself. for each factor chosen,
explain your reasoning.}
\end{minipage}}}\bigskip{}

\begin{table}
\begin{centering}
\begin{tabular}{|c|>{\centering}p{4cm}|>{\raggedright}p{8cm}|>{\raggedright}p{2cm}|}
\hline
\# & IVs Suggested & \centering{}Rationale Provided & \centering{}Citations\tabularnewline
\hline
\hline
1 & {\footnotesize{}Distance from Home to College } & {\footnotesize{}This is purely geographical and independent of personal
attributes. } & \centering{}{\footnotesize{}\nocite{card1995using}Card ('95), 1955;
\nocite{card2001estimating}Card ('01), 3134}\tabularnewline
\hline
2 & {\footnotesize{}Personal Savings for Education } & {\footnotesize{}While this might be influenced by a student's ability
to save, it's largely determined by external economic conditions and
family support. } & \centering{}\tabularnewline
\hline
3 & {\footnotesize{}Number of Siblings Attending College } & {\footnotesize{}This depends on family decisions and the educational
pursuits of siblings, unrelated to the student's personal capabilities
or characteristics. } & \centering{}{\footnotesize{}\nocite{griliches1977estimating}Griliches
('77), 2393}\tabularnewline
\hline
4 & {\footnotesize{}Parents' Educational Background } & {\footnotesize{}A demographic factor that precedes the student's own
choices and abilities. } & \centering{}\tabularnewline
\hline
5 & {\footnotesize{}Religious Affiliations } & {\footnotesize{}A personal and family cultural factor, not necessarily
a direct influence on personal academic or career potential. } & \centering{}\tabularnewline
\hline
\hline
6 & {\footnotesize{}Campus Crime Rates } & {\footnotesize{}Safety measures and local crime statistics are external
to the student's attributes. } & \centering{}\tabularnewline
\hline
7 & {\footnotesize{}Diversity of Student Body } & {\footnotesize{}Reflects the school's admissions policies and demographic
outreach, independent of any one student's traits. } & \centering{}\tabularnewline
\hline
8 & {\footnotesize{}Library Size (Volumes Held) } & {\footnotesize{}A resource factor that's determined by the institution's
funding and priorities. } & \centering{}\tabularnewline
\hline
9 & {\footnotesize{}Availability of Online Courses } & {\footnotesize{}Indicates the institution's commitment to flexible
learning options, independent of student characteristics. } & \centering{}\tabularnewline
\hline
10 & {\footnotesize{}Campus Facilities Quality (Gyms, Labs, etc.) } & {\footnotesize{}Dependent on the school's investment in infrastructure. } & \centering{}\tabularnewline
\hline
11 & {\footnotesize{}Campus Housing Capacity } & {\footnotesize{}Dictated by the university's accommodations and facilities,
not by student traits. } & \centering{}\tabularnewline
\hline
12 & {\footnotesize{}Proximity to Urban Centers } & {\footnotesize{}Geographic location of the school relative to cities. } & \centering{}\tabularnewline
\hline
13 & {\footnotesize{}Environmental Sustainability Rating } & {\footnotesize{}Reflects the institution's environmental policies
and infrastructure. } & \centering{}\tabularnewline
\hline
14 & {\footnotesize{}Technology Integration in Classrooms } & {\footnotesize{}Indicates the level of modern technological adoption
by the school. } & \centering{}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{Returns to College: Suggested IVs and Rationale for Validity}
\label{tab:schooling}

\noindent\begin{minipage}[t]{1\columnwidth}
\medskip{}

\emph{\scriptsize{}Notes}{\scriptsize{}: IVs for college attendance
are presented. All IVs are discovered and explained by GPT4 from a
single run of Prompts \ref{P1-1} and \ref{P2x-1} with $K_{0}=40$
and $K$ left unspecified. The first 5 rows are categorized by GPT4
to be student-related factors, and the next 9 rows to be school-related
factors. The total running time was less than 1 minute.}{\scriptsize\par}
\end{minipage}
\end{table}
Table \ref{tab:schooling} presents the results from a single session
of running Prompts \ref{P1-1} and \ref{P2x-1}. It contains IVs suggested
by GPT4 and GPT4's rationale for the suggestions.\footnote{GPT4 also provide the summary of overall rationale, which is not reported
here for succinctness.} In the table, we find IVs that are already popular in the literature
(e.g., \#1, 3, 5) as well as IVs that seem to be new (to our best
knowledge) (e.g., \#6, 7, 9, 11, 12, 13, 14). The latter have potential
to be valid, especially after being conditioned on additional covariates
that are not considered in the prompt. Producing all these results
took less than one minute in total. The rationale given by GPT4 can
be elaborated further by requesting it in the same session, which
we do not present here for brevity.\footnote{Appendix \ref{subsec:Three-Step-Prompts} explains an effective way
of adjusting the length of responses via a system message.}

\subsubsection{Returns to Years of Education\label{subsec:Returns-to-Years}}

As a second example, we consider the returns to years of schooling.
We omit the prompts as they are similar as before, except that we
impose the following role: ``\texttt{you are a student beginning
high school in the united states. you need to make a decision on how
many more years you will stay in school}.'' Table \ref{tab:schooling-1}
contains the IVs suggested by GPT4. Some candidates (e.g., \#3, 4)
are similar to those found in the first example. Notably, GPT4 also
finds variables (e.g., \#1, 2) that are related to compulsory schooling
laws, the popular IVs in the literature. Local regional characteristics
(e.g., \#5, 6, 7) are also interesting findings.

\begin{table}
\begin{centering}
\begin{tabular}{|c|>{\centering}p{4cm}|>{\raggedright}p{8cm}|>{\raggedright}p{2cm}|}
\hline
\# & IVs Suggested & \centering{}Rationale Provided & \centering{}Citations\tabularnewline
\hline
\hline
1 & {\footnotesize{}Credits Required for Graduation} & {\footnotesize{}This is a regulatory requirement and does not depend
on your personal qualities or the inherent quality of your school
but can determine the length of time you stay in school. } & \centering{}\tabularnewline
\hline
2 & {\footnotesize{}State Education Laws} & {\footnotesize{}Laws mandating education until a certain age or credits
affect all students uniformly, irrespective of individual abilities
or school quality. } & \centering{}{\footnotesize{}\nocite{angrist1991does}Angrist \& Krueger
('91), 3951; \nocite{staiger1997instrumental}Staiger \& Stock ('97),
10656}\tabularnewline
\hline
3 & {\footnotesize{}Distance to Nearest College/University} & {\footnotesize{}The geographical proximity to higher education institutions
is an external factor that doesn't depend on your ability but can
influence the decision to pursue further education based on accessibility. } & \centering{}{\footnotesize{}\nocite{card1995using}Card ('95), 1955;
\nocite{card2001estimating}Card ('01), 3134}\tabularnewline
\hline
4 & {\footnotesize{}Transportation Availability} & {\footnotesize{}The ease of getting to and from educational institutions
is a logistical concern that affects your ability to attend and continue
education, independent of personal academic qualities or school resources. } & \centering{}\tabularnewline
\hline
5 & {\footnotesize{}Local Education Funding per Student} & {\footnotesize{}This reflects the financial investment in education
at a community or state level, impacting the resources available to
you, which can influence educational continuation decisions, yet is
not a reflection of personal ability. } & \centering{}\tabularnewline
\hline
6 & {\footnotesize{}Local College Enrollment Rates} & {\footnotesize{}Community trends in college enrollment can create
a social environment that encourages further education, which is more
about local culture and less about individual capacity. } & \centering{}\tabularnewline
\hline
7 & {\footnotesize{}Local Unemployment Rates} & {\footnotesize{}Higher local unemployment might incentivize staying
in school longer as a way to defer entering a tough job market, a
decision influenced by economic conditions rather than personal academic
prowess. } & \centering{}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{Returns to Years of Schooling: Suggested IVs and Rationale for Validity}
\label{tab:schooling-1}

\noindent\begin{minipage}[t]{1\columnwidth}
\medskip{}

\emph{\scriptsize{}Notes}{\scriptsize{}: IVs for years of schooling
are presented. All IVs are discovered and explained by GPT4 from a
single run of Prompts \ref{P1-1} and \ref{P2x-1} (modified for this
example) with $K_{0}=40$ and $K$ left unspecified. The total running
time was less than 1 minute.}{\scriptsize\par}
\end{minipage}
\end{table}


\subsection{Supply and Demand}

\subsubsection{Production Function Estimation\label{subsec:Production-Functions}}

Consider estimating a production function that captures the causal
relationship between inputs and outputs. The key identification challenge
is that input decisions can be correlated with unobserved productivity
shocks, which directly influence outputs. To address this, IVs such
as input prices have been proposed in the literature \citep{griliches1998production},
which have been subsequently criticized \citep{olley1996dynamics,levinsohn2003estimating,ackerberg2015identification}.

Here are the prompts we use. As before, we choose $K_{0}=40$ and
let GPT4 choose $K$; we explicitly request separate lists for market
factors and firm and manager factors. Note that in Prompt \ref{P2x-2-1},
we use a loose description of covariates (unlike in Prompt \ref{P2x-1}
where we assign specific values).

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{1-2-1}[Example: Production Functions]\label{P1-2-1}\end{myprompt}\vspace{-0.3cm}

\texttt{you are a manager at a manufacturing firm. you need to make
a decision on how much labor and capital inputs to use to produce
outputs. what would be factors (factors of markets and economy and
factors of yourself) that can determine your decision but that do
not directly affect your output productions, except through the input
choices (that is, that affect your firm's outputs only through inputs)? list
forty factors that are quantifiable, twenty for market factors and
twenty for managerial factors. explain the answers.}
\end{minipage}}}

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{2$_{x}$-2-1}[Example: Production Functions]\label{P2x-2-1}\end{myprompt}\vspace{-0.3cm}

\texttt{suppose you are a manager at a firm with specific level of
capital intensity and specific scale of operations, which has a specific
market share in a specific industry. among the forty factors listed
above, choose all factors that are not influenced by productivity
shocks of your firm, which determine outputs. create separate lists
for market factors and managerial factors. for each factor chosen,
explain your reasoning.}
\end{minipage}}}\bigskip{}

Table \ref{tab:production} presents the results from a single session
of running Prompts \ref{P1-2-1} and \ref{P2x-2-1}. It contains IVs
suggested by GPT4 and GPT4's rationale. Interestingly, IVs that are
suggested in the literature (i.e., input prices) are not chosen by
GPT4 although they appear in the answer to Prompt \ref{P1-2-1} (not
shown here for brevity).\footnote{This result was consistent over multiple runs.}
This suggests that these IVs are not deemed by GPT4 to satisfy \ref{as:IND},
aligning with similar concerns in the literature \citep{olley1996dynamics,levinsohn2003estimating,ackerberg2015identification}.
However, GPT4 suggests IVs that may influence input prices (e.g.,
\#1, 2, 3, 5, 6, 14), some of which can be arguably exogenous. There
are a handful of other IVs suggested as market-related and managerial
factors. Among the latter, there are variables related to long-term
decisions of the firm, which are argued by GPT4 to not influence short-term
productivity shocks. However, long-term decisions affect long-term
outputs, which may or may not be relevant to the short-term outputs
of concern. Overall, the explanations given by GPT4 are more detailed
than those in Table \ref{tab:schooling}, reflecting the random nature
of the LLM's responses.

\begin{table}
\begin{centering}
\begin{tabular}{|c|>{\centering}p{4cm}|>{\raggedright}p{10cm}|}
\hline
\# & IVs Suggested & \centering{}Rationale Provided\tabularnewline
\hline
\hline
1 & {\footnotesize{}Interest Rates} & {\footnotesize{}Influence the cost of borrowing, affecting decisions
on how much capital to acquire or lease for production without altering
current productivity levels.}\tabularnewline
\hline
2 & {\footnotesize{}Exchange Rates} & {\footnotesize{}Affect the price of imported inputs, which can change
input costs and investment decisions in foreign capital, again without
directly impacting the productivity of existing inputs.}\tabularnewline
\hline
3 & {\footnotesize{}Tariffs and Trade Policies} & {\footnotesize{}Alter the cost structures for imported and exported
goods, impacting decisions on where to source inputs or sell outputs,
but do not affect the productivity of the inputs themselves.}\tabularnewline
\hline
4 & {\footnotesize{}Regulatory Changes} & {\footnotesize{}Can necessitate changes in production processes or
input materials, which may affect costs and sourcing decisions without
directly impacting the productivity of the inputs once deployed.}\tabularnewline
\hline
5 & {\footnotesize{}Transportation Costs} & {\footnotesize{}Affect the total cost of inputs and can influence
decisions regarding sourcing and logistics. These costs don't directly
alter how effectively inputs are converted into outputs.}\tabularnewline
\hline
6 & {\footnotesize{}Real Estate Prices} & {\footnotesize{}Influence decisions about expanding or relocating
production facilities, affecting capital expenditure and operational
scale without altering current production efficiency.}\tabularnewline
\hline
7 & {\footnotesize{}Environmental Regulations} & {\footnotesize{}Stricter environmental regulations may force a firm
to switch to greener, possibly more expensive inputs or production
technologies, affecting input choices and costs but not directly modifying
the productivity of any specific input set.}\tabularnewline
\hline
8 & {\footnotesize{}Patents and Intellectual Property Rights} & {\footnotesize{}Governed by legal systems and not directly influenced
by a firm\textquoteright s internal productivity shocks.}\tabularnewline
\hline
9 & {\footnotesize{}Political Stability} & {\footnotesize{}Affect market predictability and investment decisions,
impacting where and how much to invest in capital and labor, without
changing how these inputs produce once acquired.}\tabularnewline
\hline
10 & {\footnotesize{}Global Economic Conditions} & {\footnotesize{}Influence market opportunities and risks associated
with foreign investments, guiding input acquisition strategies without
changing input productivity.}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{(a) Production Functions: Suggested Candidates for IVs (Market Factors)}
\label{tab:production}

\noindent\begin{minipage}[t]{1\columnwidth}
\medskip{}

\emph{\scriptsize{}Notes}{\scriptsize{}: IVs for inputs are presented.
All IVs are discovered and explained by GPT4 from a single run of
Prompts \ref{P1-2-1} and \ref{P2x-2-1} with $K_{0}=40$ and $K$
left unspecified. Panel (a) is categorized by GPT4 to be market-related
factors, and Panel (b) to be managerial factors. The total running
time was less than 1 minute.}{\scriptsize\par}
\end{minipage}
\end{table}
\begin{table}
\begin{centering}
\begin{tabular}{|c|>{\centering}p{4cm}|>{\raggedright}p{10cm}|}
\hline
\# & IVs Suggested & \centering{}Rationale Provided\tabularnewline
\hline
\hline
10 & {\footnotesize{}Capital Availability} & {\footnotesize{}Determines its ability to invest in both labor and
capital inputs, shaping the scale of operations and types of technology
employed, without directly changing how productive these inputs are}\tabularnewline
\hline
11 & {\footnotesize{}Company\textquoteright s Risk Tolerance} & {\footnotesize{}Can decide the extent to which a firm is willing to
invest in new, potentially more efficient, but riskier technologies
or markets, affecting input decisions rather than the productivity
of current inputs}\tabularnewline
\hline
12 & {\footnotesize{}Strategic Objectives} & {\footnotesize{}Long-term strategic objectives may dictate prioritizing
certain types of inputs or production scales, influencing the firm\textquoteright s
approach to markets and technology investments without affecting current
input productivity.}\tabularnewline
\hline
13 & {\footnotesize{}Financial Health of the Company} & {\footnotesize{}The overall financial stability can limit or expand
the firm\textquoteright s ability to procure and utilize inputs optimally,
shaping how inputs are managed and financed rather than directly influencing
their productivity}\tabularnewline
\hline
14 & {\footnotesize{}Compliance and Legal Considerations:} & {\footnotesize{}Driven by external legal requirements and internal
ethics, not by short-term productivity}\tabularnewline
\hline
15 & {\footnotesize{}Corporate Social Responsibility (CSR) Initiatives} & {\footnotesize{}Strategic decisions about CSR are influenced by long-term
planning and brand image considerations.}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{(b) Production Functions: Suggested Candidates for IVs (Managerial
Factors)}
\label{tab:production-1}

\noindent\begin{minipage}[t]{1\columnwidth}
\medskip{}

\emph{\scriptsize{}Notes}{\scriptsize{}: IVs for production inputs
are presented. All IVs are discovered and explained by GPT4 from a
single run of Prompts \ref{P1-2-1} and \ref{P2x-2-1} with $K_{0}=40$
and $K$ left unspecified. Panel (a) is categorized by GPT4 to be
market-related factors, and Panel (b) to be managerial. The total
running time was less than 1 minute.}{\scriptsize\par}
\end{minipage}
\end{table}


\subsubsection{Demand Estimation\label{subsec:Demand-Estimation}}

Consider estimating demand for consumers in a given market. In estimating
the effect of price on demand, the main concern is that price is endogenous,
as it is an equilibrium outcome. Researchers have used supply-side
IVs that are excluded from the demand equation in a simultaneous system
for supply and demand \citep{angrist2000interpretation} or motivated
by structural models \citep{berry1995automobile}.

Below are the prompts we use, focusing on the setting of \citet{angrist2000interpretation}.
As before, we choose $K_{0}=40$ and let GPT4 choose $K$. Note that
in Prompt \ref{P2-2-2}, covariate information is omitted.

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{1-2-2}[Example: Demand]\label{P1-2-2}\end{myprompt}\vspace{-0.3cm}

\texttt{you are a dealer at a fish market. you need to set the prices
of fish. what would be factors that can determine your decision but
that do not directly affect the customers' demand for fish, except
through the price you set (that is, that affect the demand only through
fish prices). list forty factors that are quantifiable. explain your
answer.}
\end{minipage}}}

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{2-2-2}[Example: Demand]\label{P2-2-2}\end{myprompt}\vspace{-0.3cm}

\texttt{suppose you are a dealer at the fish market who is selling
fish and setting its prices on a day of the week. among the factors
listed above, choose all factors are not influenced by fish market
conditions or customers' characteristics that determine demand for
fish. for each factor chosen, explain your reasoning.}
\end{minipage}}}\bigskip{}

\begin{table}
\begin{centering}
\begin{tabular}{|c|>{\centering}p{2cm}|>{\raggedright}p{10cm}|>{\raggedright}p{2cm}|}
\hline
\# & IVs Suggested & \centering{}Rationale Provided & \centering{}Citations\tabularnewline
\hline
\hline
1 & {\footnotesize{}Fuel Costs} & {\footnotesize{}Costs for fuel typically depend on broader economic
factors such as oil prices and are not directly affected by market
demand for fish or specific customer characteristics. } & \centering{}\tabularnewline
\hline
2 & {\footnotesize{}Fishing Equipment Depreciation} & {\footnotesize{}The depreciation of fishing gear like nets and boats
is largely a function of use and time, rather than market conditions
or customer demand. } & \centering{}\tabularnewline
\hline
3 & {\footnotesize{}Weather Conditions} & {\footnotesize{}Weather affects fishing operations directly but is
not influenced by market conditions or the specific characteristics
of customers. } & \centering{}{\footnotesize{}\nocite{angrist2000interpretation}Angrist,
Graddy \& Imbens ('00), 454}\tabularnewline
\hline
4 & {\footnotesize{}Regulatory Costs} & {\footnotesize{}Costs associated with licensing and compliance with
fishing regulations are dictated by regulatory bodies and do not fluctuate
based on customer demand or market conditions.} & \centering{}\tabularnewline
\hline
5 & {\footnotesize{}Interest Rates} & {\footnotesize{}The rates charged on loans for purchasing boats or
equipment are determined by financial markets and policy, rather than
direct market conditions of the fish market or customer demands. } & \centering{}\tabularnewline
\hline
6 & {\footnotesize{}Insurance Costs} & {\footnotesize{}Premiums for insuring fishing equipment and operations
are typically based on risk assessments and sector-wide data, rather
than being directly influenced by day-to-day market conditions or
characteristics of customers. } & \centering{}\tabularnewline
\hline
7 & {\footnotesize{}Utility Costs at Sales Points} & {\footnotesize{}The cost of utilities like electricity and water at
sales points tends to be fixed or based on usage rates that are independent
of market demand specifics. } & \centering{}\tabularnewline
\hline
8 & {\footnotesize{}Economic Conditions} & {\footnotesize{}Broader economic factors that affect overall spending
and investment patterns influence operational costs but are not dictated
by the fish market conditions or customer preferences. } & \centering{}\tabularnewline
\hline
9 & {\footnotesize{}Technological Advances} & {\footnotesize{}Investments in technology to improve fishing or sales
operations are usually planned based on long-term business strategy
and efficiency gains rather than immediate market conditions or specific
customer demographics. } & \centering{}\tabularnewline
\hline
10 & {\footnotesize{}Government Subsidies} & {\footnotesize{}Subsidies to the fishing industry are determined by
government policies, which are independent of daily market conditions
or customer demands in the fish market. } & \centering{}\tabularnewline
\hline
11 & {\footnotesize{}Tariffs on Imports} & {\footnotesize{}Tariffs imposed on imported fish are a matter of international
trade policy and do not change based on daily fluctuations in market
demand or customers characteristics. } & \centering{}\tabularnewline
\hline
12 & {\footnotesize{}Employee Training Costs} & {\footnotesize{}Costs of training employees in handling and selling
fish relate to operational efficiency, which are not directly influenced
by the day's market conditions or customer demands. } & \centering{}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{Demand Estimation: Suggested Candidates for IVs}
\label{tab:demand}

\noindent\begin{minipage}[t]{1\columnwidth}
\medskip{}

\emph{\scriptsize{}Notes}{\scriptsize{}: IVs for price are presented.
All IVs are discovered and explained by GPT4 from a single run of
Prompts \ref{P1-2-2} and \ref{P2-2-2} with $K_{0}=40$ and $K$
left unspecified. The total running time was less than 1 minute.}{\scriptsize\par}
\end{minipage}
\end{table}
Table \ref{tab:demand} presents the results from a single session
of running Prompts \ref{P1-2-2} and \ref{P2-2-2}. It contains IVs
suggested by GPT4 and GPT4's rationale. Many supply-side factors (e.g.,
costs) are chosen by GPT4, which are reasonable candidates for IVs.
The ``weather conditions'' variable used as IVs in \citet{angrist2000interpretation}
appears in the list (\#3). Interestingly, another supply-side factor
``labor costs'' produced by Prompt \ref{P1-2-2} is not included
in the final list produced by Prompt \ref{P2-2-2}. When asked ``\texttt{explain
why you didn't include ``labor costs'' in the final list},'' GPT4
responded that ``labor costs are somewhat flexible and can be adjusted
in response to changes in market conditions and customer demand, making
them more dynamic than some of the other factors listed.''

\subsection{Peer Effects\label{subsec:Peer-Effects}}

Suppose we are interested in the causal effects of peers on an individual's
outcomes within a social network. We consider two well-known examples:
(i) the effects of peer farmers on the adoption of new farming technologies
\citep{foster1995learning,conley2010learning}; (ii) the effects of
peers on teenage smoking \citep{gaviria2001school}. In both examples,
the main source of endogeneity is latent factors that determine the
formation of network (e.g., latent homophily). To address this, the
literature on peer effects sometimes uses friends of friends as IVs
\citep{bramoulle2009identification,angrist2014perils}. In both examples,
we construct prompts similar to the first two examples, except that
we choose $K_{0}=20$.

\subsubsection{Effects of Peer Farmers on New Technology Adoption\label{subsec:Effects-of-Peer-1}}

Here are the prompts. It is worth noting that, in this example, the
role-playing is done from the peer's perspective, rather than from
the perspective of the individual whose outcome is of concern.

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{1-3-1}[Example: Peer Effects on Technology Adoption]\label{P1-3-1}\end{myprompt}\vspace{-0.3cm}

\texttt{you are a farmer in a village in rural india. you want to
influence your peer farmers in the same village to introduce a new
farming technologies that you introduced. what would be factors (factors
of farming and village, and factors of yourself) that can determine
your influence on peers but that do not directly affect your peers'
technology adoption decisions, except through your influence (that
is, that affect your peers' decisions only through your influence)? list
twenty factors that are quantifiable. explain your answer.}
\end{minipage}}}

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{2$_{x}$-3-1}[Example: Peer Effects on Technology Adoption]\label{P2x-3-1}\end{myprompt}\vspace{-0.3cm}

\texttt{suppose you are a 40 year old male farmer of a specific crop
in the village in rural india. among the factors listed above, which
factors are not influenced by factors (e.g., similar background and
preferences) that brought you and your peers in the same neighborhood
and social network from the first place? for each factor chosen, explain
your reasoning.}
\end{minipage}}}\bigskip{}

Table \ref{tab:peer1} presents the results in a single session of
running Prompts \ref{P1-3-1} and \ref{P2x-3-1}. It contains IVs
suggested by GPT4 and GPT4's rationale. Interestingly, some IVs that
are suggested in the literature (e.g., friends of friends) are not
chosen by GPT4. It can be because GPT4 either views them as invalid
or is incapable of identifying them. On the other hand, \#10 seems
to relate to the IV used in \citet{conley2010learning}, which exploits
variation in the presence of experienced farmers. Additionally, there
are other IVs that seem to be new, notably \#9. Finally, GPT4 fails
to identify the IV used in \citet{foster1995learning}, namely, endowed
land size. The validity of this variable is justified in their paper
by the specific historical backgrounds of the Indian villages studied.
This implies that, for some IVs, providing LLMs with institutional
details can be necessary; see Appendix \ref{subsec:Three-Step-Prompts}
for implementing this via system messages.

\begin{table}
\begin{centering}
\begin{tabular}{|c|>{\centering}p{4cm}|>{\raggedright}p{10cm}|}
\hline
\# & IVs Suggested & \centering{}Rationale Provided\tabularnewline
\hline
\hline
1 & {\footnotesize{}Number of Training Sessions Attended} & {\footnotesize{}This reflects an individual\textquoteright s initiative
to seek additional education and may vary widely among farmers with
similar backgrounds due to personal choices, availability of time,
or perception of the value of formal education}\tabularnewline
\hline
2 & {\footnotesize{}Number of Demonstrations Conducted} & {\footnotesize{}This is influenced by an individual\textquoteright s
willingness to share knowledge and the resources they are willing
or able to allocate for such activities, which might not be uniformly
distributed even among neighbors}\tabularnewline
\hline
3 & {\footnotesize{}Increase in Yield After Technology Adoption} & {\footnotesize{}This factor depends significantly on how effectively
an individual implements and adapts new technologies to their specific
farming practices, which can differ greatly even among farmers growing
the same crops}\tabularnewline
\hline
4 & {\footnotesize{}Access to Credit} & {\footnotesize{}Individual financial management skills, credit history,
and relationships with lending institutions can influence this factor,
which isn't necessarily shared among all peers in a community}\tabularnewline
\hline
5 & {\footnotesize{}Participation in Local Decision Making Bodies} & {\footnotesize{}Involvement in local governance or committees often
depends on personal ambition, political skills, and community perception
of an individual's leadership qualities, which can vary independently
of shared initial factors}\tabularnewline
\hline
6 & {\footnotesize{}Awards or Recognitions Received} & {\footnotesize{}Achieving recognition for farming practices or community
service is heavily dependent on individual effort and excellence and
is not uniformly influenced by community-shared characteristics}\tabularnewline
\hline
7 & {\footnotesize{}Communications Skills Assessment} & {\footnotesize{}Communication effectiveness can be developed through
personal experience, education, and innate ability, which vary among
individuals regardless of their initial common background}\tabularnewline
\hline
8 & {\footnotesize{}Number of Articles or Case Studies Published} & {\footnotesize{}This is influenced by an individual\textquoteright s
ability to write, research, and engage with broader audiences, skills
that are not common to all individuals in a network}\tabularnewline
\hline
9 & {\footnotesize{}Engagement in Social Media or Technology Platforms} & {\footnotesize{}The level of engagement in modern communication and
technology platforms depends on personal interest in technology, availability
of resources, and individual attitudes toward innovation}\tabularnewline
\hline
10 & {\footnotesize{}Success Stories Documented and Shared} & {\footnotesize{}The ability and willingness to document and share
success stories can be a personal choice influenced by an individual\textquoteright s
perspective on self-promotion and communication}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\addtocounter{table}{-1}\caption{Peer Effects on Technology Adoption: Suggested Candidates for IVs}
\label{tab:peer1}

\noindent\begin{minipage}[t]{1\columnwidth}
\medskip{}

\emph{\scriptsize{}Notes}{\scriptsize{}: IVs for peers' decisions
are presented. Given the perspective in the prompts, the variables
should be understood as factors of }\emph{\scriptsize{}peers}{\scriptsize{}.
IVs for All IVs are discovered and explained by GPT4 from a single
run of Prompts \ref{P1-3-1} and \ref{P2x-3-1} with $K_{0}=20$ and
$K$ left unspecified. The total running time was less than 1 minute.}{\scriptsize\par}
\end{minipage}
\end{table}


\subsubsection{Effects of Peer Teenagers on Smoking Behavior\label{subsec:Effects-of-Peer}}

Here are the prompts. Again, in this example, the role-playing is
done from the peer's perspective. We choose a teenager in urban Indonesia
for its relevance, given that the teenage smoking rate in Indonesia
has been recently reported as one of the world's highest \citep{fithria2021indonesian}.
We consider a social media network in the scenario to illustrate the
effectiveness of our approach in exploring relatively recent topics
in the literature, thereby highlighting the potential to discover
novel IVs.

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{1-3-2}[Example: Peer Effects on Smoking]\label{P1-3-2}\end{myprompt}\vspace{-0.3cm}

\texttt{you are a teenager in indonasia who smokes. you want to influence
your peers in your social media network to smoke. what would be factors
(factors of social media, your school and region, and factors of yourself)
that can determine your influence on peers but that do not directly
affect your peers' smoking decisions, except through your influence
(that is, that affect your peers' decisions only through your influence)? list
twenty factors that are quantifiable. explain your answer.}
\end{minipage}}}

\bigskip{}
\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{2$_{x}$-3-2}[Example: Peer Effects on Smoking]\label{P2x-3-2}\end{myprompt}\vspace{-0.3cm}

\texttt{suppose you are a teenage boy in urban indonesia who goes
to high school and is from middle-income family. among the factors
listed above, which factors are not influenced by factors (e.g., similar
background and preferences) that brought you and your peers in the
same social network from the first place? for each factor chosen,
explain your reasoning.}
\end{minipage}}}\bigskip{}

\begin{table}
\begin{centering}
\begin{tabular}{|c|>{\centering}p{4cm}|>{\raggedright}p{10cm}|}
\hline
\# & IVs Suggested & \centering{}Rationale Provided\tabularnewline
\hline
\hline
1 & {\footnotesize{}Engagement Rate} & {\footnotesize{}While the number of followers might be initially influenced
by shared backgrounds, the engagement rate depends more on the content
quality and how it resonates with the audience at any given time,
rather than the reasons why the audience initially formed}\tabularnewline
\hline
2 & {\footnotesize{}Frequency of Posts} & {\footnotesize{}This is largely a personal choice and strategic decision,
independent of peer background. It reflects personal commitment to
the platform rather than shared interests or backgrounds}\tabularnewline
\hline
3 & {\footnotesize{}Content Reach} & {\footnotesize{}This metric depends on the algorithm and how engaging
the content is, which can transcend the initial reasons for network
formation}\tabularnewline
\hline
4 & {\footnotesize{}Number of Platforms Used} & {\footnotesize{}Choosing to diversify presence across various platforms
is a strategic decision that doesn\textquoteright t necessarily correlate
with shared backgrounds or interests}\tabularnewline
\hline
\hline
5 & {\footnotesize{}School Policies on Smoking} & {\footnotesize{}These are institutional decisions that apply to all
students, regardless of the factors that brought them together}\tabularnewline
\hline
6 & {\footnotesize{}Local Regulations on Tobacco Advertising} & {\footnotesize{}Like school policies, these are determined by local
government and affect all residents uniformly, irrespective of social
network formations}\tabularnewline
\hline
7 & {\footnotesize{}Access to Smartphones Among Peers} & {\footnotesize{}This might vary widely even within similar socio-economic
backgrounds due to individual family decisions or priorities}\tabularnewline
\hline
\hline
8 & {\footnotesize{}Visibility at School Events} & {\footnotesize{}Active participation and visibility in school events
are personal choices and efforts that go beyond shared backgrounds,
reflecting individual initiative}\tabularnewline
\hline
9 & {\footnotesize{}History of Disciplinary Actions at School} & {\footnotesize{}This is generally a result of personal behavior and
choices rather than group influence}\tabularnewline
\hline
10 & {\footnotesize{}Academic Performance} & {\footnotesize{}Although there could be a correlation with socio-economic
status, individual effort and capability play significant roles, making
this somewhat independent of why peers might group together initially}\tabularnewline
\hline
11 & {\footnotesize{}Extracurricular Leadership Roles} & {\footnotesize{}Holding leadership positions is often based on personal
qualities, skills, and choices rather than the shared preferences
and backgrounds that might define a social network initially}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{Peer Effects on Smoking: Suggested Candidates for IVs}
\label{tab:peer1-1}

\noindent\begin{minipage}[t]{1\columnwidth}
\medskip{}

\emph{\scriptsize{}Notes}{\scriptsize{}: IVs for peers' decisions
are presented. Given the perspective in the prompts, the variables
should be understood as factors of }\emph{\scriptsize{}peers}{\scriptsize{}.
All IVs are discovered and explained by GPT4 from a single run of
Prompts \ref{P1-3-2} and \ref{P2x-3-2} with $K_{0}=20$ and $K$
left unspecified. The first four rows concern social media factors;
the next three rows concern school and regional factors; and the last
four rows concern personal factors. The total running time was less
than 1 minute.}{\scriptsize\par}
\end{minipage}
\end{table}
Table \ref{tab:peer1-1} presents the results from a single session
of running Prompts \ref{P1-3-2} and \ref{P2x-3-2}. It contains IVs
suggested by GPT4 and GPT4's rationale. It is important to note that,
given that the prompts are written from the perspective of peers,
the variables in the table should be understood as factors influencing
\emph{peers} of the focal individual. Given that the setup incorporates
modern elements such as social media, we identify many potentially
new and interesting IVs, particularly from the social media category
(i.e., \#1, 2, 3, 4, 7). Interestingly, \#7 can be viewed as a ``friends
of friends'' IV.\footnote{There were other ``friends of friends'' IVs that are produced from
Step 1 but did not survive Step 2.}

\section{Adversarial Large Language Models\label{sec:Adversarial-Large-Language}}

One way to refine the answers of an LLM is to request another LLM
(or a different session of the same LLM) to play the role of an adversary
and review the responses produced by the first LLM, namely the defender
LLM. In the adversarial stage, we fully disclose the IV discovery
task and ask the adversarial LLM to provide counter-arguments, but
without using econometric jargon. These counter-arguments are then
given to the defender LLM, which is asked to refine the previous answers.
Overall, this process produces more sophisticated responses. We find
this to be a useful exercise, which mimics the mental process of a
human researcher.

The following is an example prompt given to the adversarial LLM. In
the prompt, \texttt{{[}the list of variables by Defender{]}} is the
list provided the defender LLM, generated from running, for example,
Prompts \ref{P1}--\ref{P2}. \bigskip{}

\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{AD1}\label{P_AD1}\end{myprompt}\vspace{-0.3cm}

\texttt{you are a researcher who wants to find instrumental variables
to estimate the effect of {[}treatment{]} on {[}outcome{]}. below
is a list of candidate instrumental variables. for each variable in
the list, provide arguments as to why it may not be a valid instrument:}~\\
\texttt{}~\\
\texttt{{[}the list of variables by Defender{]}}
\end{minipage}}}

\bigskip{}

Then, the counter-arguments generated from Prompt \ref{P_AD1} are
presented to the defender LLM using the following prompt, which follows
Prompts \ref{P1}--\ref{P2}.\bigskip{}

\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{3}\label{P3}\end{myprompt}\vspace{-0.3cm}

\texttt{below are counter-arguments for each of your previous answers. based
on these arguments, revise your selection and provide a list:}~\\
\texttt{}~\\
\texttt{{[}the arguments by Adversary{]}}
\end{minipage}}}\bigskip{}

When running Prompt \ref{P3}, it is important to encourage the defender
to create a list, otherwise, it might become pessimistic and reject
all the previous selections.

We apply this adversarial process to our examples in Section \ref{sec:Discovered-IVs}.
Overall, the defender provides more sophisticated answers in the end.
In the ``Returns to College'' example (Section \ref{subsec:Returns-of-College}),
the variables related to geographical proximity or transportation
have been retained. The other variables related to school facilities
(e.g., library size, technology integration in classrooms) have been
modified to factors that may better satisfy exogeneity (e.g., the
percentage of renewable energy used on campus, the number of green
spaces on campus, campus medical facilities). New variables also appeared
(e.g. language spoken at home). Interestingly, many variables now
appear to be weak IVs, which is sensible. In the ``Returns to Years
of Education'' examples (Section \ref{subsec:Returns-to-Years}),
the ``state laws'' variable survived, but was renamed to the more
descriptive ``mandatory minimum years of schooling required by state.''
Other variables were also further refined. ``GED program availability''
emerged. In the ``Demand Estimation'' example (Section \ref{subsec:Demand-Estimation}),
the ``weather condition'' variable was refined to ``global climate
patterns (e.g., El Ni\~{n}o),'' because the counter-argument was
that ``bad weather can also discourage customers from coming to the
market.'' Many other variables were refined to reflect global and
macroeconomic conditions. Again, these variables suffer from being
weak IVs or having less individual variation.

\section{Variables Search in Other Causal Inference Methods\label{sec:Variables-Search-in}}

In this section, we demonstrate how prompting strategies similar to
those for the IV discovery can be used to find (i) control variables
under which treatments are conditionally independent (i.e., exogenous);
(ii) control variables under which parallel trends are likely to hold
in difference-in-differences; and (iii) running variables in regression
discontinuity designs.

\subsection{Conditional Independence\label{subsec:Conditional-Independence}}

Using the same notation as in Section \ref{sec:Notation-and-IV},
consider a conditional independence (CI) assumption that assigns a
more crucial role to the vector of control variables $X\equiv(X_{1},...,X_{L})$:

\begin{myas}{CI}\label{as:CI} For any $d$, $D\perp Y(d)|X$.\end{myas}

Assumption \ref{as:CI} is commonly introduced in causal inference
settings, especially when combined with machine learning to estimate
nuisance functions; e.g., debiased/double machine learning methods
\citep{chernozhukov2024applied}. More traditionally, this assumption
is closely related to matching and propensity score matching techniques
\citep{heckman1998matching}. The mean independence version of \ref{as:CI}
(i.e., $E[Y(d)|D,X]=E[Y(d)|X]$) is relevant to regression methods.


We propose using LLMs to systematically search for $X$ that satisfies
a verbal version of \ref{as:CI}. The prompt writing is slightly simpler
than that for IVs. In particular, we construct prompts that solicit
the relationship between $X$ and $D$ (Step 1) and $X$ and $Y(d)$
(Step 2). Therefore, \emph{only} the second-step prompt involves a
counterfactual statement. Let $L_{0}$ be the number of controls
to be found in Step 1 ($L_{0}\ge L$). One may want to choose the
value of $L_{0}$ to be larger than one would normally use for $K_{0}$
and leave $L$ unspecified.\bigskip{}

\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{C1}[Search for Control Variables]\label{PC1}\end{myprompt}\vspace{-0.3cm}

\texttt{you are {[}agent{]} who needs to make a {[}treatment{]} decision
in {[}scenario{]}. what factors determine your decision? list {[}L\_0{]}
factors that are quantifiable. explain the answers.}
\end{minipage}}}

\bigskip{}

\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{C2}[Refine Control Variables]\label{PC2}\end{myprompt}\vspace{-0.3cm}

\texttt{among the {[}L\_0{]} factors listed above, choose all factors
that directly determine your {[}outcome{]}, not only indirectly through
{[}treatment{]}. the chosen factors can still influence your {[}treatment{]}. for
each factor chosen, explain your reasoning.}
\end{minipage}}}

\bigskip{}

The prompts are constructed to search for confounders and need to
be controlled for. Researchers sometimes mistakenly control for ``colliders''
and/or ``mediators'' \citep{pearl2000causality}, which are intended
to be excluded from the search. Note that Prompts \ref{PC1}--\ref{PC2}
can also be adapted to jointly search for covariates and latent confounders
in the IV search. In this case, one can distinguish $X$ from latent
confounders by referring to the former as ``\texttt{quantifiable}.''
Also, one may want to use the phrase ``\texttt{demographic factors}''
to refer to $X$, as they are common control variables in many empirical
applications.

\subsection{Difference in Differences\label{subsec:Difference-in-Differences}}

The difference-in-differences (DiD) method is popular in empirical
research, partly due to the simplicity and intuitiveness of its main
assumption, namely, the parallel trend assumption (stated below).
However, this assumption is not directly testable and typically hard
to justify \citep{ghanem2022selection,rambachan2023more}. It is believed
that conditioning on the right control variables can make this assumption
more justifiable, which can motivate the search for such controls.

\begin{myas}{PT}\label{as:PT}$E[\Delta Y(0)|D,X]=E[\Delta Y(0)|X]$
where $\Delta Y(0)\equiv Y_{after}(0)-Y_{before}(0)$.\end{myas}

Assumption \ref{as:PT} can be viewed as a mean independence version
of \ref{as:CI}, where the counterfactual outcome is replaced with
the temporal difference of counterfactual (untreated) outcomes before
and after the event. Therefore, Prompts \ref{PC1}--\ref{PC2} can
be directly used to search for $X$ that satisfy a verbal version
of \ref{as:PT}. This can be done by inputting ``\texttt{average
temporal changes in {[}outcome\_t{]} during the time of no {[}treatment{]}}''
for \texttt{{[}outcome{]}} in Prompt \ref{PC2}, where \texttt{{[}outcome\_t{]}}
refers to $Y_{t}$ for $t\in\{before,after\}$. The example of such
prompts is constructed to revisit the classical empirical example,
namely, the effects of minimum wage on the fast food industry's labor
markets \citep{card1994minimum}; see Appendix \ref{subsec:Search-for-Control}
for the actual prompts. Table \ref{tab:DID} contains the control
variables suggested by GPT4, conditional on which the parallel trend
is likely to hold, and GPT4's rationale. On the list, \#3, 4, 7, 10,
11 are particularly interesting and \#11 seems particularly novel.
In the table, the first four rows (\#1, 2, 3, 4) are chosen by GPT4
from an additional prompt that emphasizes the requirement with respect
to $\Delta Y(0)$: ``\texttt{be sure to choose all factors that do
not determine the average wage level but only determine the temporal
changes in average wages}.'' Nonetheless, controls that satisfy the
mean version of \ref{as:CI} with the level, $Y_{t}(0)$ for $t\in\{before,after\}$,
are also valid controls for \ref{as:PT}.

\begin{table}
\begin{centering}
\begin{tabular}{|c|>{\centering}p{3.5cm}|>{\raggedright}p{10cm}|}
\hline
\# & Control Variables Suggested & \centering{}Rationale Provided\tabularnewline
\hline
\hline
1 & {\footnotesize{}Inflation Rates} & {\footnotesize{}Higher inflation can decrease the real income of workers
if wages don't adjust accordingly, making a case for introducing a
minimum wage to maintain purchasing power.}\tabularnewline
\hline
2 & {\footnotesize{}Consumer Price Index (CPI)} & {\footnotesize{}A rising CPI without wage increases can reduce workers'
real earnings, suggesting a minimum wage could help offset cost of
living increases.}\tabularnewline
\hline
3 & {\footnotesize{}Job Vacancy Rates} & {\footnotesize{}Low vacancies could keep wages down due to ample labor
supply; high rates may drive wages up, affecting the urgency and approach
to minimum wage legislation.}\tabularnewline
\hline
4 & {\footnotesize{}Labor Productivity Growth} & {\footnotesize{}This affects temporary wage increases linked to efficiency
improvements. Productivity gains can lead to bonus payments or higher
wages temporarily but don't define starting wage rates.}\tabularnewline
\hline
\hline
5 & {\footnotesize{}Employment Growth Rates} & {\footnotesize{}Robust job growth in the sector might show the industry
can support higher wages, while slow growth suggests caution in wage
hikes.}\tabularnewline
\hline
6 & {\footnotesize{}Labor Force Participation Rate} & {\footnotesize{}High participation may suppress wages due to increased
labor supply, supporting the need for a minimum wage to ensure fair
earnings.}\tabularnewline
\hline
7 & {\footnotesize{}Union Membership Rates} & {\footnotesize{}Low unionization in fast food often means weaker wage
bargaining, possibly justifying minimum wage laws to improve wage
standards.}\tabularnewline
\hline
8 & {\footnotesize{}Turnover Rates} & {\footnotesize{}High turnover might indicate wage dissatisfaction,
pointing to the need for a minimum wage to stabilize the workforce.}\tabularnewline
\hline
9 & {\footnotesize{}Corporate Profit Trends} & {\footnotesize{}If profits are high but wages are stagnant, it could
suggest that the industry can afford to pay more, bolstering the argument
for a minimum wage.}\tabularnewline
\hline
10 & {\footnotesize{}Economic Diversity Score} & {\footnotesize{}Dependency on low-wage sectors like fast food in less
economically diverse areas might necessitate a minimum wage to ensure
living standards.}\tabularnewline
\hline
11 & {\footnotesize{}Percentage of Workforce in Gig Economy} & {\footnotesize{}Increased gig work could pressure fast food employers
to offer competitive wages, influencing when and how to implement
minimum wage laws.}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{Difference-in-Differences for Minimum Wage: Suggested Control Variables}
\label{tab:DID}

\noindent\begin{minipage}[t]{1\columnwidth}
\medskip{}

\emph{\scriptsize{}Notes}{\scriptsize{}: All controls are discovered
and explained by GPT4 from a single run of Prompts \ref{PC1}--\ref{PC2},
adapted to Assumption \ref{as:PT} with $L_{0}=40$ and $L$ left
unspecified. Among them, the first four row are factors that are chosen
from the additional emphatic prompt: ``}\texttt{\scriptsize{}be sure
to choose all factors that do not determine the average wage level
but only determine the temporal changes in average wages}{\scriptsize{}.''
The total running time was less than 1 minute.}{\scriptsize\par}
\end{minipage}
\end{table}


\subsection{Regression Discontinuity\label{subsec:Regression-Discontinuity}}

Regression discontinuity designs (RDDs) are another well-known method
for causal inference that closely relates to the IVs method \citep{lee2010regression}.\footnote{For example, the fuzzy RDD estimand can be viewed as the two-stage
least squares estimand.} The key for this method to work is to find a running variable (i.e.,
assignment variable) that satisfies the following:

\begin{myas}{RD}\label{as:CI-1} There exists a variable $R_{j}$
and a cutoff $r_{0}$ such that $D=1$ if $R_{j}\ge r_{0}$ and $D=0$
if $R_{j}<r_{0}$.\end{myas}

One can use LLMs to systematically search for running variables $\{R_{1},...,R_{J}\}$
for a given $D$ and $Y$ of interest. We provide the example of prompts
here. It is worth noting that, unlike in all the previous cases, none
of the prompts below involve counterfactual statements. Therefore,
if LLMs outperform a traditional search for running variables, it
would be due to their automated and comprehensive search behavior.\footnote{In further refining the candidates of running variables to ensure
that RRD's continuity assumptions are satisfied, counterfactual prompting
would be necessary; see Appendix \ref{subsec:Further-Refining-Search}.} Similarly as above, we only specify initial $J_{0}$ and leave $J$
unspecified.\bigskip{}

\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{R1}[Search for Running Variables]\label{PR1}\end{myprompt}\vspace{-0.3cm}

\texttt{you are {[}agent{]} who needs to make a {[}treatment{]} decision
in {[}scenario{]}. what would be the possible criteria based on which
your eligibility for {[}treatment{]} is determined? provide {[}J\_0{]}
of the most relevant criteria that are (1) quantifiable and (2) have
specific cutoffs determining eligibility. explain the answers.}
\end{minipage}}}

\bigskip{}

\noindent{\fboxsep 10pt\fbox{\begin{minipage}[t]{1\columnwidth - 2\fboxsep - 2\fboxrule}
\begin{myprompt}{R2}[Refine Running Variables]\label{PR2}\end{myprompt}\vspace{-0.3cm}

\texttt{among the {[}J\_0{]} criteria listed above, choose all criteria
that involve continuous or ordered measures and have precise cutoffs
determining eligibility. also report the cutoff value for each criterion
from verifiable sources only (ensuring no fabricated or hypothetical
numbers are used). explain the answers.}
\end{minipage}}}

\bigskip{}

Note that when Prompt \ref{PR2} is run on GPT4, it will engage in
a series of automated web searches. The request for cutoff values
may lead the LLM to provide hypothetical numbers as possibilities.
When one wants to get the actual values from verifiable sources, it
is important to explicitly state that, as we do above. We apply Prompts
\ref{PR1}--\ref{PR2} to a range of famous examples in the literature
where RDDs are used as empirical strategies. Table \ref{tab:RD} presents
the results obtained by running the prompts, which are adapted to
each specific context and country of the empirical example. In most
cases, a handful of new possible running variables are suggested by
GPT4 with specific cutoffs obtained from web sources. Except for one
case (i.e., \#4), GPT4 also identifies the running variables used
in the literature.

\begin{table}
\begin{centering}
\begin{tabular}{|c|>{\centering}p{2.5cm}|>{\centering}p{2.5cm}|>{\centering}p{3cm}|>{\raggedright}p{6cm}|}
\hline
\# & {\footnotesize{}Outcome(s) (Country)} & {\footnotesize{}Treatment(s)} & {\footnotesize{}Suggested Running Variable, Same as the Literature} & \begin{centering}
{\footnotesize{}Other Suggested Running Variables}{\footnotesize\par}
\par\end{centering}
\centering{}{\footnotesize{}(Cutoffs for Eligibility)}\tabularnewline
\hline
\hline
1 & {\footnotesize{}Spending on schools, test scores (US)} & {\footnotesize{}State education aid} & {\footnotesize{}Relative average property values (\nocite{guryan2001does}Guryan,
2001)} & {\footnotesize{}- Percentage of low-income students (e.g., Equity
Multiplier 2023-2024, above 70\%)}{\footnotesize\par}

{\footnotesize{}- Mobility rate (e.g., Equity Multiplier, above 25\%)}{\footnotesize\par}

{\footnotesize{}- Age (e.g., Transitional Kindergarten (TK) expansion
2023-24, 15th b-day by April 2)}{\footnotesize\par}

{\footnotesize{}- Local Control Funding Formula (LCFF) (California)$^{*}$}\tabularnewline
\hline
2 & {\footnotesize{}College enrollment (US)} & {\footnotesize{}Financial aid offer} & {\footnotesize{}SAT scores, GPA (\nocite{vanderklaauw2002estimating}Van
der Klaauw, 2002)} & {\footnotesize{}- Expected family contribution (EFC) (e.g., the Pell
Grant 2023-2024: below \$6,656)}\tabularnewline
\hline
3 & {\footnotesize{}Overall insurance coverage (US)} & {\footnotesize{}Medicaid eligibility} & {\footnotesize{}Age (\nocite{card2004using}Card and Shore-Sheppard,
2004)} & {\footnotesize{}- Federal Poverty Level (FPL) (e.g., Washington D.C.:
below 215\% and below 221\% (family of 3); equiv. annual incomes below
\$31,347 and \$54,940, reps.)}{\footnotesize\par}

{\footnotesize{}- Household Size (e.g., Modified Adjusted Gross Income
(MAGI) rules: expressed as \% of FPL, adjusted by 5\% FPL disregard)}\tabularnewline
\hline
4 & {\footnotesize{}Employment rates (Italy)} & {\footnotesize{}Job training program} & {\footnotesize{}Attitudinal test score (\nocite{battistin2002testing}Battistin
and Rettore, 2002)$^{\dagger}$} & {\footnotesize{}- Age (e.g., below 35; source: National Policies Platform)}{\footnotesize\par}

{\footnotesize{}- Income: (e.g., below 60\%; source: National Policies
Platform)}{\footnotesize\par}

{\footnotesize{}- Salary (e.g., EU Blue Card: above 3/2 of average
Italian salary; source: ETIAS Italy)}\tabularnewline
\hline
5 & {\footnotesize{}Re-employment probability (UK)} & {\footnotesize{}Job search assistance, training, education} & {\footnotesize{}Age at end of unemployment spell (\nocite{degiorgi2005long}De
Giorgi, 2005)} & {\footnotesize{}- Age (e.g., Jobseeker\textquoteright s Allowance
(JSA): above 18, with exceptions for some 16 or 17; source: UK Rules) }{\footnotesize\par}

{\footnotesize{}- Minimum Salary (e.g., Skilled Worker visa: above
\pounds38,700 or going rate for job type, whichever is higher; source:
GOV.UK)}{\footnotesize\par}

{\footnotesize{}- Residency Duration (e.g., JSA: above 3 months prior
to claim, for new or returning UK nationals; source: UK Rules)}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{Regression Discontinuity: Suggested Candidates for Running Variables}
\label{tab:RD}

\noindent\begin{minipage}[t]{1\columnwidth}
\medskip{}

\emph{\scriptsize{}Notes}{\scriptsize{}: All running variables are
discovered and explained by GPT4 from a single run of Prompts \ref{PR1}--\ref{PR2},
adapted to each context with $J_{0}=20$ and $J$ left unspecified.
All running variables used in the literature (Column 4) are also found
by GPT4, except \#4. The total running time for each row was less
than 1 minute (even with an automated web search for Prompt \ref{PR2}).
The sources indicated are given by GPT4 with links. $*$: A formula,
not a running variable. $\dagger$: Not found by GPT4.}{\scriptsize\par}
\end{minipage}
\end{table}


\section{Conclusions\label{sec:Conclusions}}

This essay proposes the agenda to use LLMs to systematically search
for variables in designing causal inference. It merely serves as a
starting point, and there are potential next steps that can follow.
In constructing prompts for IVs, there are many possible ways for
sophistication: First, one can consider using previously known IVs
in the literature to guide LLMs to discover new ones. This can be
done by adding textual demonstration of how Assumptions \ref{as:REL}--\ref{as:IND}
are satisfied with known IVs \emph{before} starting the proposed prompts.
This approach would evoke \emph{few-shot learning} in LLMs \citep{brown2020language},
which can enhance their performances. This approach would also
``orthogonalize'' the search \citep{ludwig2024machine} to focus
on novel IVs. Second, the elaborated search can be directed toward
finding IVs that are more policy-relevant \citep{imbens1994identification,heckman2005structural})
by specifying targeted policies in the prompt. Third, none of the
results reported in the current essay are findings aggregated across
sessions. To account for and potentially leverage the stochastic nature
of LLMs' responses, exploring the possibility of aggregation (e.g.,
taking the union or intersection of $\mathcal{Z}_{K}$'s across sessions)
would be beneficial. Fourth, we can explore the use of multiple
LLM agents each assuming distinct roles in the discovery process,
analogous to the collaborative and critical interactions among human
researchers. An example of this approach is discussed in Section \ref{sec:Adversarial-Large-Language},
where the responses of one LLM are reviewed and critiqued by another
acting as a critic.

Additionally, we can consider having a horse race among multiple LLMs
or using an open-source LLM to fine-tune it \citep[e.g., ][]{du2024labor}
for our purpose. A potential challenge is that the performance metric
is hard to define in our context due to the lack of ground truth for
valid IVs. In fact, this is the very reason we propose using LLMs
from the first place: for any IVs found by human researchers or the
machine, there are only more compelling narratives or less compelling
ones. In later stages, when data eventually come into play, over-identification
tests can potentially be a fruitful framework for the evaluation of
LLMs. More broadly, it would be interesting to apply the proposed
approach of variable search in other empirical examples and other
causal inference methods.

\newpage{}