Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 128 chars of source]
center[center omitted — 142 chars of source]
\address{Department of Economics, University of Houston, 3623 Cullen Boulevard - Office 221-C
Houston, TX, 77204, USA}
\address{Vancouver School of Economics, University of British Columbia, 6000 Iona Drive, Vancouver, BC, V6T 1L4, Canada}
\fontsize{12}{14} \selectfont
abstractStructural models that admit multiple reduced forms, such as game-theoretic models with multiple equilibria, pose challenges in practice, especially when parameters are set-identified and the identified set is large. In such cases, researchers often choose to focus on a particular subset of equilibria for counterfactual analysis, but this choice can be hard to justify. This paper shows that some parameter values can be more “desirable” than others for counterfactual analysis, even if they are empirically equivalent given the data. In particular, within the identified set, some counterfactual predictions can exhibit more robustness than others, against local perturbations of the reduced forms (e.g. the equilibrium selection rule). We provide a representation of this subset which can be used to simplify the implementation. We illustrate our message using moment inequality models, and provide an empirical application based on a model with top-coded data.
{ Key words: Counterfactual Analysis, Multiple Equilibria, Equilibrium Selection Rule, Locally Robust Refinement}
{ JEL Classification: C30, C57}
bibunit[econometrica]
\@startsection{section}{1}
\z@{.5\linespacing\@plus.5\linespacing}{.4\linespacing}{Introduction}
Economists often use a structural model to analyze the effect of a policy that has never been implemented, such as potential mergers or changes to legislation. For such analysis, it is crucial to have a plausible specification of the causal relationship between the policy variable and the outcome variable of interest. However, a plausible specification can come short of determining the causal relationship uniquely. Prominent examples are found in game theoretic models, where the presence of multiple equilibria admits multiple reduced forms. (Tamer:03:ReStud called models with multiple reduced forms incomplete.)\footnote{In this paper, a reduced form refers to a system of equations where the variables on the right hand side are all external variables (i.e. inputs) and the variables on the left hand side are all internal variables (i.e. generated through the system of equations using the inputs) in the sense of Heckman:10:QJE (p.56). The term “reduced form” in this sense appears in econometrics at least as early as in Koopmans:49:Eca.} The multiplicity of reduced forms is often the source of set-identification of structural parameters.
When a parameter is set-identified, all the values in the identified set are empirically equivalent, in the sense that the observed data cannot be used to distinguish between the values. The identified set represents the empirical content of the model for the parameter (i.e., the amount of information on the value of the parameter that can be extracted from data through the model). However, when a parameter is set-identified with a large identified set, the identified set can be of little use for counterfactual analysis in practice. It is not unusual that predictions from counterfactual analysis based on different equilibria yield self-contradictory results (Aguirregabiria/Nevo:12:WP, p.111). In this situation, just as in the calibration approach in macroeconomics, the researcher typically uses additional constraints on the parameter values or chooses a specific equilibrium to obtain a meaningful analysis.\footnote{For example, in an entry model, Jia:08:Eca focuses on the equilibrium that is most profitable to one of the firms, and Ackerberg/Gowrisankaran:06:RAND use an equilibrium selection rule which randomizes between a Pareto-best equilibrium and a Pareto-worst equilibrium with an unknown probability. Another example is Roberts/Sweeting:13:AER in the context of auctions, who consider the particular equilibrium with the lowest signal for agents in the model so they enter the auction.} In other words, the researcher considers a refinement of the identified set, by introducing additional restrictions. However, these restrictions are often only weakly motivated.
In our view, a desirable approach in this situation is to distinguish between the identification step and the refinement step, and lay out transparently what criterion (or motivation) is used for the refinement step. There is a small body of literature that explores this refinement step. Song:14:ET analyzed the problem of choosing a point from the identified interval based on the local asymptotic minimax regret criterion. Aryal/Kim:13:JBES considered an incomplete English auction model where the seller is ambiguity averse, and employing the $\Gamma$-maximin expected utility framework of Gilboa/Schmeidler:89:JME, suggested a point decision on the optimal reserve price. Jun/Pinkse:17:WPa considered an incomplete English auction set-up similar to that in Haile/Tamer:03:JPE and proposed a point decision on the optimal reserve price using a maximum entropy method. Jun/Pinkse:19:JOE studied a discrete two-player complete information game and established sharp identification of counterfactual predictions of the game. Furthermore, they investigated point decisions on the probability distributions through a maximum entropy method and what they called a Dirichlet approach, and compared the two methods.
The main goal of this paper is to draw attention to the fact that some values in the identified set can give counterfactual predictions that are more robust to the change of certain aspects of model specification than others. Thus, it is reasonable to give priority to those values that produce most “reliable” counterfactual predictions. Generally speaking, predictions that are sensitive to a change of the aspects of the model that the researcher is not sure about cannot be reliable. Given that the researcher is often least sure about how a reduced form is determined among multiple reduced forms, we consider a subset of the identified set which produces counterfactual predictions that are more robust to local perturbations of the reduced forms (e.g. local perturbations of the equilibrium selection rule in a game setting) than others.
More specifically, consider a causal relationship between endogenous variable $y$ and exogenous variables $x,\varepsilon$:
\begin{align}
y = \rho_{\beta,\gamma}(x,\varepsilon;\eta),
\end{align}
where $\beta$ is the parameter of interest, $\gamma$ a nuisance parameter, and $\eta$ is a random variable with distribution $G$. Without observing $\eta$ or knowing $G$, the model admits multiple reduced forms. For example, in the case of games with multiple equilibria, $G$ plays the role of the equilibrium selection rule.\footnote{As shown in the paper, our approach applies to causally incomplete models other than game-theoretic models, such as models involving top-coded data.}
Suppose that the counterfactual setting of interest involves a counterfactual distribution $\tilde F_X$ for $X$. Consider the average structural function (ASF) (Blundell/Powell:04:ReStud):
\begin{align*}
ASF_{\beta,\gamma,G}(x) = \int \int \rho_{\beta,\gamma}(x,\varepsilon;\eta)dG(\eta|x,\varepsilon)dF_\varepsilon(\varepsilon),
\end{align*}
where $F_\varepsilon$ denotes the distribution of $\varepsilon$, and $G(\cdot|x,\varepsilon)$ the conditional distribution of $\eta$ given $x$ and $\varepsilon$. Then our proposal is that we choose $\gamma$ that minimizes
\begin{align}
\sup_{G' \in B(G;\delta)}\left| \int ASF_{\beta,\gamma,G}(x)d\tilde F_X(x) - \int ASF_{\beta,\gamma,G'}(x)d\tilde F_X(x)\right|,
\end{align}
where $B(G;\delta)$ denotes a $\delta$-neighborhood of $G$ for some small $\delta>0$. (As we show later, given $\beta$, the set of $\gamma$'s that minimize ((ref)) does not depend on the choice of $\delta$ or $G$.) We call the resulting set of values $(\beta,\gamma)$ in the identified set \textit{the Locally Robust Refinement (LRR)}. In other words, our refinement is motivated by the desirability of a stable behavior in the counterfactual predictions when the structural parameters are partially identified.
In many set-ups, the LRR is a substantially smaller set than the identified set. Furthermore, the LRR is often very intuitive. For example, as we show in Section 3, the LRR in an entry game picks nuisance parameter values that minimize the region of multiple equilibria. In the example of a switching regression model with interval data, it picks nuisance parameters that minimize a difference across models in switching regression.
However, unlike the identified set, the size of the LRR does not indicate the empirical content of the model. Rather the LRR gives parameter values which produce counterfactual predictions that are most robust to the change of the reduced forms among their multiple forms. Due to this robustness property, one may want to pay more attention to those parameter values than others when conducting policy exercises, even if all the values in the identified set are observationally equivalent.
The formulation in ((ref)) gives the impression that the LRR involves additional optimization over an infinite dimensional space in estimation, adding to computational cost. Using the linearity of $\textsf{ASF}_{\beta,\gamma,G}(x)$ in $G$ and Hilbert space geometry, this paper gives a simple characterization of the LRR so that focusing on the LRR does not cause extra computational cost. In fact, due to the additional constraints that come from the LRR, the computational cost is substantially reduced from a benchmark case without using the LRR. Focusing on moment inequality models, this paper presents a method of bootstrap inference on the LRR from the identified set and shows its uniform asymptotic validity. We provide results from Monte Carlo simulations, and illustrate the LRR through an empirical application on top-coded data. These exercises illustrate that the LRR's predictions are indeed more stable than those with the identified set and that they may affect policymaking in applied contexts (i.e. policies suggested by the identified set could be driven by parameter values that are not robust, and are not part of the LRR).
There has long been interest in studying model sensitivity and robustness to misspecifications in econometric models. Recent contributions in this line include Andrews/Gentzkow/Shapiro:18:NBER, Armstrong/Kolesar:20:QE and Bonhomme/Weidner:20:WP who, among many others, proposed a method of measuring local misspecification and performing inference that is robust to local misspecifications. In the context of instrumental variable models, Nevo/Rosen:12:ReStat permitted the instrumental variable to be correlated with the error term and presented bounds for the parameter of interest under an assumption on the correlation between the instrumental variable and the error term. Oster:19:JBES proposed a method of extracting information on the omitted variable bias from data in linear regression models. Masten/Poirier:18:Ecma provided an identification analysis of various treatment effect parameters from treatment effect models where the conditional independence restriction is weakened to a partial dependence condition.\footnote{Robustness to misspecification has also been considered in the macroeconomics literature (e.g. in optimal control theory as in Hansen/Sargent:01:AER, see Hansen/Marinacci:16:SS as well), where one can consider perturbations of the decision maker's model, but less emphasis is placed on the identification and inference problems in these models. Finally, Andrews/Gentzkow/Shapiro:17:QJE focus on the case where the model is point-identified, and present a statistic on how estimates vary when a model's moments changes.} Christensen/Connault:19:WP proposed a computationally tractable method of performing inference on counterfactual quantities that arise in structural models, while permitting a global misspecification of the model. Li:19:WP uses a semiparametric restriction (instead of parametric ones) to generate bounds for counterfactuals of interest in moment inequality models.
All these authors are concerned about robustness of inference to misspecifications. Their robust inference often leads to a larger confidence set, and a wider range of policy predictions. In contrast, we consider settings where the identified set is already too large to be useful for a meaningful policy analysis. This can happen \textit{even if} the model is correctly specified (see Aguirregabiria/Nevo:12:WP for a discussion, particularly in the context of multiple equilibria). Instead of enlarging an existing identified set by “robustifying” it as above, we explore a rationale for focusing on a subset of the identified set for policy analysis. Our criterion is based on the stability of counterfactual predictions as we perturb the reduced-form selection rules. One can relate our guiding philosophy to the literature of equilibrium refinements in game theory, where one focuses on a subset of the set of Nash equilibria based on its stability against various perturbations of the game. (See Chapter 5 of Myerson:91:GameTheory.) Then the idea of local robustness in the paper, i.e., the stability of reduced forms against local perturbations of reduced form selection rules is related to Fudenberg/Kreps/Levine:88:JET who considered robustification of the set of strict Nash equilibria against local perturbations of the games.
The paper is organized as follows. The next section gives a general set-up of a structural model and presents examples that fit our framework. Section 3 presents the approach of LRR and provides its characterization that is useful for implementation in practice. Focusing on generic moment inequality models, Section 4 proposes inference for the LRR and presents results from Monte Carlo simulations. An empirical application is presented in Section 5. Section 6 concludes. In the Appendix, we provide some mathematical proofs. The Supplemental Note to this paper includes additional information on the data for the empirical application, and presents results on the uniform validity of our inference procedure.
\@startsection{section}{1}
\z@{.5\linespacing\@plus.5\linespacing}{.4\linespacing}{A Counterfactual Analysis with Multiple Reduced Forms}
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Models with Multiple Reduced Forms}
Suppose that data $\{(Y_i',X_i')'\}_{i=1}^n$ are drawn from the joint distribution $P$ of $(Y',X')'$ which obeys the following \textit{reduced form}: for $Y \in \mathbf{R}$ and $W = (X',\varepsilon')'$,
\begin{align}
Y = \rho(W;\eta),
\end{align}
for some function $\rho$, where $X$ is observed but $\varepsilon$ and $\eta$ are not. For each value $\eta$, we regard the map $\rho(\cdot;\eta)$ as representing a functional-causal relationship between internal variable $Y$ and external variables $W$.\footnote{Here we follow the proposal by Heckman:10:QJE (p.56) and use the terms \textit{internal variables} and \textit{external variables}. External variables are specified from outside of the structural equations as inputs, and internal variables are generated through the system of equations from the inputs. External variables are not necessarily exogenous variables. When $X$ is correlated with $\varepsilon$, the variable $X$ is called endogenous, although $(X,\varepsilon)$ is jointly given externally as inputs, and thus is a vector of external variables. Here a reduced form refers to a system of equations where all the variables on the left hand side are internal variables and all the variables on the right hand side are external variables.} This relationship is fully determined once the value of $\eta$ is realized. Typically, $\varepsilon$ has a structural meaning such as unobserved costs, whereas $\eta$ does not and is used only as an index for a reduced form. For example, in a game-theoretic model, $W = (X',\varepsilon')'$ represents observed and unobserved payoff components and $\eta$ an index for an equilibrium. The response of $Y$ to a change in $W$ is ambiguous unless the value of $\eta$ is fixed. In this sense, there is a multiplicity of reduced forms in ((ref)). To describe the generation of observed outcome $Y$, we assume that Nature selects the value of $\eta$ from the conditional distribution $G(\cdot|w)$ given $W = w$. In a game setting, $G$ is often called the equilibrium selection rule. However, the multiplicity of reduced forms can arise in a setting that has nothing to do with a game, such as in the case of switching regressions with interval data which we describe below. Hence we call $G$ the \textit{reduced-form selection rule} in this paper.
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Counterfactual Analysis Under Partial Identification}
We focus on a counterfactual experiment where the distribution $F$ of the external vector $W$ is changed to a counterfactual one $\tilde F$. Such a counterfactual experiment includes cases where one “shuts off” the effect of a certain covariate, “fixing” of the covariate to a certain value, or changing the distribution of $\varepsilon$. Suppose that the counterfactual quantity of interest is generically denoted by $C(\rho,\tilde F,G)$. This quantity could be, for example, the prediction of game outcomes or expected payoffs under a counterfactual scenario:
\begin{align*}
C(\rho,\tilde F,G) = \int \rho(w;\eta)dG(\eta|w) d\tilde F(w).
\end{align*}
Or in other applications, $C(\rho,\tilde F,G)$ could be the average increase in welfare. Let $\mathcal{G}$ be the space of the reduced-form selection rules $G$ and let $\mathcal{R}$ be the identified set for the function $\rho$. Then the identified set for the counterfactual prediction $C(\rho,\tilde F,G)$ is given by
\begin{align*}
\{C(\rho,\tilde F,G):G \in \mathcal{G}, \rho \in \mathcal{R}\}.
\end{align*}
A standard approach to make inference based on this identified set faces multiple challenges in practice. The set of predictions can be of little practical use, when the estimated set is too large. Furthermore, estimation of the identified set itself and statistical inference can be computationally intensive.
Due to these difficulties, it is not unusual that empirical researchers focus on a particular subset of $\mathcal{G}$ (such as focusing on a particular equilibrium) or consider a subset of $\mathcal{R}$ (such as by restricting the admissible values of parameters as in the case of calibration). Such a subset selection approach, while inevitable in practice, is not easy to motivate.
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Examples}
For brevity, we consider two examples here to illustrate our set-up: a finite game of complete information and a switching regression with interval data.
\@startsection{subsubsection}{3}
\z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Finite Game of Complete Information}
Consider the setting of a complete information $N$ player game with finite actions, and let $A$ be the finite action space for each player.\footnote{A special case is the two player entry game used by Bresnahan/Reiss:91:JOE and Tamer:03:ReStud, among others.} Suppose that the payoff for player $i$ playing $\tilde a \in A$ when all other players are playing $a_{-i} = (a_j)_{j \ne i}$ is given by $u_i(\tilde a, a_{-i},W)$ for a map $u_i$, where $W = (X,\varepsilon)$ is a payoff-relevant characteristic vector of player $i$. For each $a \in A^N$, we define
\begin{align}
\mathbb{W}(a) = \left\{w: u_i(a_i,a_{-i},w) \ge \max_{a' \in A} u_i(a',a_{-i},w), \forall i=1,...,N \right\}.
\end{align}
Hence if $w$ belongs to $\mathbb{W}(a)$, $a$ is a Nash equilibrium for the game with $w$, so that the set
\begin{align}
\mathcal{A}(w) = \left\{a \in A^N: w \in \mathbb{W}(a) \right\}
\end{align}
represents the set of equilibria for the game with $w$. Let $\eta$ be a random vector taking values in $A^N$ where the conditional distribution $G(\cdot|w)$ of $\eta$ given $w$ has a support in $\mathcal{A}(w)$. Let $Y = (Y_i)_{i =1 }^N$ be the observed action profile that comes from a Nash equilibrium $\eta$ of the game after Nature draws $\eta$ from $G(\cdot|w)$. Then, we can write for $i \in N$,
\begin{align}
Y_i = \rho_i(W;\eta),
\end{align}
where $\rho_i$ is the $i$-th entry of $\rho$ which is defined as
\begin{align*}
\rho(w;\eta) = \sum_{a \in \mathcal{A}(w)} 1\{\eta = a\} a.
\end{align*}
The conditional distribution $G(\cdot|w)$ over $\eta$ represents the equilibrium selection rule.
\@startsection{subsubsection}{3}
\z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{A Switching Regression Model with Interval Outcomes}
Consider an interval data set-up similar to one studied by Manski/Tamer:02:ECMA. More specifically, we consider top-coded data with threshold variables $Z_1 < Z_2$, where $Z_1, Z_2$ are observed and common across individuals. We observe an outcome random variable which equals a latent $Y$ if $Y \leq Z_1$ and is only known to be within $[Z_1, Z_2]$ otherwise, and a vector of covariates $X_1$. For example, $X_1$ represents a vector of covariates for each worker such as age, education, experience, gender, and $Y$ is the log of the hourly wage.
It is assumed that the latent $Y$ follows:
\begin{eqnarray*}
Y = \left\{\begin{array}{ll}
g_1\left(\beta_0 D + X_1'\gamma_0 \right) + \varepsilon, &\text{ if } \eta = 1, \text{ and }\\
g_2\left(\beta_0 D + X_1'\gamma_0 \right) + \varepsilon, & \text{ if } \eta = 0,
\end{array}
\right.
\end{eqnarray*}
where $D$ denotes the policy variable, and $\eta$ the binary indicator of a regime, and $g_1$ and $g_2$ are some known functions. As for the error term $\varepsilon$, we assume that $\mathbf{E}[\varepsilon|X_1,D] = 0$. We also assume that $Y \le Z_2$ with probability one. Let $\tilde Y_1 = \tilde Y_2 = Y$ if $Y \le Z_1$, and $\tilde Y_1 = Z_1$ and $\tilde Y_2 = Z_2$ if $Y > Z_1$. The econometrician observes $\tilde Y_1$ and $\tilde Y_2$. Let $\pi_0(X_1,D) = P\{\eta = 1 |X_1,D\}$ and $\Pi$ be the collection of functions that $\pi_0$ is known to belongs to.
It is straightforward that the identified set for $(\beta_0,\gamma_0)$ is given by
\begin{equation}
\Theta_P = \left\{(\beta_0,\gamma_0): \mathbf{E}[\tilde Y_1 |X_1,D] \leq \alpha(X_1,D;\beta_0,\gamma_0,\pi) \leq \mathbf{E}[\tilde Y_2 |X_1, D], \text{a.e.} \text{ for some } \pi \in \Pi \right\},
\end{equation}
where
\begin{align}
\alpha(X_1,D;\beta_0,\gamma_0,\pi) = g_1\left(\beta_0 D + X_1'\gamma_0 \right) \pi(X_1,D) + g_2\left(\beta_0 D + X_1'\gamma_0 \right) (1 - \pi(X_1,D)).
\end{align}
Now we can write
\begin{align*}
Y = \rho(X,\varepsilon;\eta),
\end{align*}
where $X = (X_1,D)$, and
\begin{align}
\rho(X,\varepsilon;\eta) = g_1\left(\beta_0 D + X_1'\gamma_0 \right) \eta + g_2\left(\beta_0 D + X_1'\gamma_0 \right) (1 - \eta) + \varepsilon.
\end{align}
The distribution $G$ of $\eta$ plays a role analogous to the equilibrium selection rule in a game-theoretic model. Here it selects from a given set of reduced forms.
\@startsection{section}{1}
\z@{.5\linespacing\@plus.5\linespacing}{.4\linespacing}{Locally Robust Refinement}
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Overview}
Suppose that the reduced-form $\rho(\cdot;\eta)$ in ((ref)) is parametrized as follows:
\begin{align}
Y = \rho_{\beta_0,\gamma_0}(W;\eta),
\end{align}
where $\beta_0$ is a parameter of interest taking values in a set $B$ and $\gamma_0$ is a nuisance parameter taking values in a set $\Gamma$, and $\eta$ belongs to some known compact set denoted by $H$. Let $\Theta_P$ be the identified set for $(\beta_0,\gamma_0)$. The counterfactual experiment involves changing the distribution $F$ of $W$ to a counterfactual distribution $\tilde F$. Let us consider the counterfactual average structural outcome:
\begin{align}
\textsf{ASO}_{\beta_0,\gamma_0}(G) = \int \int \rho_{\beta_0,\gamma_0}(w;\eta)dG(\eta|w) d\tilde F(w).
\end{align}
This is the average value of outcome $Y$ when a counterfactual change is made to the distribution of $W$ while everything else remains the same. Our focus is on a subset of $\Theta_P$ at which $\textsf{ASO}_{\beta, \gamma}(G)$ behaves robustly as we perturb the reduced-form selection rule $G$.
More formally, for any $K >0$, we introduce a $K$-sensitivity of $\textsf{ASO}_{\beta,\gamma}(G)$ as we perturb $G$ around in its $K$-neighborhood:
\begin{align}
S_{\beta,\gamma}(G;K) = \sup_{G' \in \mathcal{G}: \delta_{\mathcal{G}}(G',G) \le K} \frac{|\textsf{ASO}_{\beta,\gamma}(G') - \textsf{ASO}_{\beta,\gamma}(G)|}{\delta_{\mathcal{G}}(G',G)},
\end{align}
where $\mathcal{G}$ denotes the collection of $G$'s permitted in the model. The quantity $S_{\beta,\gamma}(G;K)$ measures the maximal rate of change of counterfactual average outcome when $G$ is perturbed within its $K$-neighborhood, where this neighborhood is defined using a metric $\delta_{\mathcal{G}}(G',G)$, formally defined below. We search for values of nuisance parameters $\gamma$ which minimize this sensitivity. For each $\beta \in B$, if we let
\begin{align}
\Gamma^{\mathsf{LRR}}(\beta) = \text{argmin}_{\gamma: (\beta,\gamma) \in \Theta_P} S_{\beta,\gamma}(G;K),
\end{align}
we define
\begin{align}
\Theta_P^{\mathsf{LRR}} = \{(\beta,\gamma) \in \Theta_P: \gamma \in \Gamma^{\mathsf{LRR}}(\beta) \}.
\end{align}
We call the subset $\Theta_P^{\mathsf{LRR}}$ the \textit{Locally Robust Refinement (LRR)} of $\Theta_P$. Hence, the LRR can be viewed as a collection of parameters at which the ASO is stable against local misspecifications of $G$. As we will see later in the characterization of the LRR set, the set is independent of the choice of $K$, $G$ and $\mathcal{G}$. The main reason behind this independence is that the functional $\textsf{ASO}_{\beta,\gamma}(G)$ is linear in $G$. (This is analogous to that the slope of a linear function is constant on the domain of the function.)\footnote{As long as the true reduced-form selection rule lies in the interior of $\mathcal{G}$, our characterization of the LRR set does not depends on the choice of $\mathcal{G}$. While it is possible to recover some information about the rule $G$ with further specifications, our procedure remains agnostic about $\mathcal{G}$, which is in line with many existing works in the econometrics literature of game-theoretic models - see Ciliberto/Tamer:09:Eca for one such example.} Note that, as we do not want our concern of local robustification against the reduced-form selection rules to directly affect the parameter of interest, our approach subjects only the nuisance parameters to the refinement.
The inference on $\beta$ based on the Locally Robust Refinement proceeds under the further restriction as follows.
\begin{align}
(\beta_0,\gamma_0) \in \Theta_P^{\mathsf{LRR}}.
\end{align}
We call this restriction \textit{the LRR Hypothesis}. Once we focus on the LRR set of parameters, this can be used for various exercises of counterfactual analysis, just as one does with the identified set of parameters.
Despite the impression that finding the set $\Theta_P^{\mathsf{LRR}}$ might be computationally intensive as it involves layers of optimization, it is not, due to our characterization result in the next section.
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Characterization of the Locally Robust Refinement}
To derive a characterization of the LRR, let us introduce some geometric structure on the space of reduced-form selection rules $G$. Let $\mathcal{G}$ be the collection of all the conditional distributions $G$ (of $\eta$ given $W$) which are dominated by the uniform distribution $\mu$ over $H$, and $w \in \mathcal{W}$, with $\mathcal{W}$ denoting the set $W$ takes values from. Let us introduce the following metric on $\mathcal{G}$: for $G,G' \in \mathcal{G}$,
\begin{align}
\delta_\mathcal{G}(G',G) = \sqrt{\int \int \left( \frac{dG'}{d\mu} - \frac{dG}{d\mu}\right)^2(\eta|w) d\mu(\eta) d\tilde F_P(w)}.
\end{align}
Let us define
\begin{align}
\Gamma(\beta) = \{\gamma \in \Gamma: (\beta,\gamma) \in \Theta_P\},
\end{align}
i.e., the set of nuisance parameters that are admitted in the counterfactual analysis together with given $\beta$. For $\kappa\ge0$, we define the $\kappa$-LRR set for $\gamma$ corresponding to $\beta$ as
\begin{align}
\quad \quad
\Gamma_{\kappa}^{\mathsf{LRR}}(\beta) = \left\{\gamma \in \Gamma(\beta): S_{\beta,\gamma}^2(G;K) \le S_{\beta,\tilde \gamma}^2(G;K)+\kappa, \forall \tilde \gamma \in \Gamma(\beta) \right\},
\end{align}
where $S_{\beta,\gamma}(G;K)$ is defined in ((ref)). The set $\Gamma^{\mathsf{LRR}}(\beta)$ defined in ((ref)) is a special case of $\Gamma_{\kappa}^{\mathsf{LRR}}(\beta)$ with $\kappa = 0$. However, as we discuss in Section 4.1, it is convenient to use the LRR set with a small positive number $\kappa>0$ which allows us to develop uniform inference. As shown below, this set does not depend on the choice of $K>0$ or that of $G \in \mathcal{G}$.
We now provide a useful characterization of the set $\Gamma_{\kappa}^{\mathsf{LRR}}(\beta)$. For this, we define
\begin{align*}
\Delta_{\beta,\gamma}(w;\eta) = \rho_{\beta,\gamma}(w;\eta) - \int \rho_{\beta,\gamma}(w;\eta) d\mu(\eta).
\end{align*}
The function $\Delta_{\beta,\gamma}(w;\eta)$ is a mean-deviation form of the structural function $\rho_{\beta,\gamma}(w;\eta)$. This function plays a central role in the characterization result. Let us make the following assumption that guarantees that the average variability of this mean-deviation is well defined.
\begin{assumption}
For each $\beta \in B$ and $\gamma \in \Gamma(\beta)$, $Q^\mathsf{LRR}(\beta,\gamma) < \infty$, where
\begin{align}
Q^\mathsf{LRR}(\beta,\gamma) = \int \int \Delta_{\beta,\gamma}^2(w;\eta) d\mu(\eta) d\tilde F(w).
\end{align}
\end{assumption}
In applications where we can compute $\rho_{\beta,\gamma}(w;\eta)$ for each value of $\eta$, we can evaluate $Q^\mathsf{LRR}(\beta,\gamma)$ easily. (See Section 3.3 below for examples.) The following theorem shows that the set $\Gamma_{\kappa}^{\mathsf{LRR}}(\beta)$ is fully characterized through the function $Q^\mathsf{LRR}(\beta,\gamma)$.
\begin{theorem}
Suppose that Assumption (ref) holds. Then for each $\beta \in B$ and $\gamma \in \Gamma(\beta)$,
\begin{align}
S_{\beta,\gamma}(G;K) = \sqrt{Q^\mathsf{LRR}(\beta,\gamma)},
\end{align}
for all $G \in \mathcal{G}$ and all $K>0$.
\end{theorem}
Thus the $\kappa$-LRR of $\gamma_0$ can be rewritten as
\begin{align}
\Gamma_{\kappa}^{\mathsf{LRR}}(\beta) = \left\{\gamma \in \Gamma(\beta): Q^\mathsf{LRR}(\beta,\gamma) \le Q^\mathsf{LRR}(\beta,\bar \gamma) + \kappa, \text{ for all } \bar \gamma \in \Gamma(\beta) \right\}.
\end{align}
The $\kappa$-LRR of the nuisance parameter vector $\gamma_0$ is the set of $\gamma$'s which minimize the average variability of $\rho_{\beta,\gamma}(w;\eta)$ (around its mean) as we perturb $\eta$. Thus, the $\kappa$-LRR for $(\beta_0,\gamma_0)$ is given by
\begin{align}
\Theta_{\kappa,P}^{\mathsf{LRR}} = \left\{(\beta,\gamma) \in \Theta_P: \gamma \in \Gamma_{\kappa}^{\mathsf{LRR}}(\beta) \right\}.
\end{align}
Sometimes, the target parameter may not be a parameter $\beta$ that determines the reduced form $\rho_{\beta,\gamma}$ but some aggregate quantity such as Average Treatment Effects or aggregate welfare. Suppose that such a target parameter is identified up to $\beta_0$ and $\gamma_0$, so that we can write it as, say $\psi(\beta_0,\gamma_0)$. (We give an example in Section (ref) below.) Then, the LRR set for $\psi(\beta_0,\gamma_0)$ is given by
\begin{align}
\Psi_{\kappa}^{\mathsf{LRR}} = \left\{ \psi(\beta,\gamma): (\beta,\gamma) \in \Theta_{\kappa,P}^{\mathsf{LRR}}\right\}.
\end{align}
To build familiarity with our characterization results for the LRR and show how it can be used in practice, we now derive $Q^{LRR}$ for two applied examples.
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Examples}
\@startsection{subsubsection}{3}
\z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Entry Game with Nash Equilibria}
Consider the setting of a 2x2 complete information entry games, such as the empirical setting in Bresnahan/Reiss:91:JOE and Ciliberto/Tamer:09:Eca, where there are two firms, $i=1,2$, deciding whether to enter a market or not. The entry decisions of the firms can then be formulated as
\begin{align}
D_1 &= 1\{\beta_{1,0} D_2 + X'\gamma_{1,0} \geq \varepsilon_1\}, \text{ and } \\
D_2 &= 1\{\beta_{2,0} D_1 + X'\gamma_{2,0} \geq \varepsilon_2\},
\end{align}
where $\beta_{1,0} D_2 + X'\gamma_{1,0} - \varepsilon_1$ represents the payoff difference for firm $1$ between entering and not entering the market, $D_i = 1$ represents the entry decision by firm $i$ and $D_i = 0$ the decision not to enter by firm $i$. The random vectors $X$ and $\varepsilon$ represent respectively the observed and unobserved characteristics of firms $1$ and $2$ combined. The coefficients $\beta_{1,0}$ and $\beta_{2,0}$ capture strategic interactions between firms and they are our coefficients of interest, whereas the coefficients $\gamma_{1,0}$ and $\gamma_{2,0}$ measure the roles of firm specific and market specific characteristics and, as such, are nuisance parameters.
\begin{center}
[INSERT FIGURE (ref) HERE]
\end{center}
In order to find a reduced form as in ((ref)), we first put
\begin{align*}
Y = (D_1,D_2),\text{ and } \varepsilon = (\varepsilon_1,\varepsilon_2),
\end{align*}
where $Y \in \mathcal{Y} = \{(0,0),(1,1),(0,1),(1,0)\}$. Define (with $\beta_0 = (\beta_{1,0}',\beta_{2,0}')'$ and $\gamma_0 = (\gamma_{1,0}',\gamma_{2,0}')'$)
\begin{align*}
A_{1,\beta_0,\gamma_0}(X) &= \{(e_1,e_2) \in \mathbf{R}^2: X'\gamma_{1,0} < e_1, X'\gamma_{2,0} < e_2\}\\
A_{2,\beta_0,\gamma_0}(X) &= \{(e_1,e_2) \in \mathbf{R}^2: \beta_{1,0} + X'\gamma_{1,0} \ge e_1, \beta_{2,0} + X'\gamma_{2,0} \ge e_2\}\\
A_{3,\beta_0,\gamma_0}(X) &= \{(e_1,e_2) \in \mathbf{R}^2: \beta_{1,0} + X'\gamma_{1,0} < e_1, X'\gamma_{2,0} \ge e_2\}, \text{ and }\\
A_{4,\beta_0,\gamma_0}(X) &= \{(e_1,e_2) \in \mathbf{R}^2: X'\gamma_{1,0} \ge e_1, \beta_{2,0} + X'\gamma_{2,0} < e_2\}.
\end{align*}
These regions are represented on the left panel of Figure (ref).
Let $\rho_{\beta,\gamma}(x,\varepsilon;\eta) = (\rho_{\beta,\gamma,1}(x,\varepsilon;\eta),\rho_{\beta,\gamma,2}(x,\varepsilon;\eta))$, where for a random variable $\eta \in \{0,1\}$,
\begin{align}
\quad \rho_{\beta,\gamma,1}(x,\varepsilon;\eta) &= 1\{(\varepsilon_1,\varepsilon_2) \in A_{2,\beta,\gamma}(x)\} + 1\{(\varepsilon_1,\varepsilon_2) \in A_{4 \setminus 3,\beta,\gamma}(x)\} \\
& \quad + 1\{(\varepsilon_1,\varepsilon_2) \in A_{3 \cap 4,\beta,\gamma}(x)\} \eta \notag \\
\rho_{\beta,\gamma,2}(x,\varepsilon;\eta) &= 1\{(\varepsilon_1,\varepsilon_2) \in A_{2,\beta,\gamma}(x)\} + 1\{(\varepsilon_1,\varepsilon_2) \in A_{3 \setminus 4,\beta,\gamma}(x)\} \\
&\quad + 1\{(\varepsilon_1,\varepsilon_2) \in A_{3 \cap 4,\beta,\gamma}(x)\} \eta \notag,
\end{align}
where $A_{3 \cap 4,\beta,\gamma}(x) = A_{3,\beta,\gamma}(x) \cap A_{4,\beta,\gamma}(x)$, $A_{4 \setminus 3,\beta,\gamma}(x) = A_{4,\beta_0,\gamma_0}(x) \setminus A_{3,\beta_0,\gamma_0}(x)$, and $A_{3 \setminus 4,\beta,\gamma}(x) = A_{3,\beta_0,\gamma_0}(x) \setminus A_{4,\beta_0,\gamma_0}(x)$.
Then we can write reduced forms for $Y_i$, for firm $i = 1,2$ as
\begin{align*}
Y_i = \rho_{\beta_0,\gamma_0,i}(X,\varepsilon;\eta).
\end{align*}
Let the distribution $G$ of $\eta$ be a discrete distribution with support $\{0,1\}$, and take $\mu$ to be a Bernoulli random variable with a mean $1/2$.\footnote{This example focuses on pure-strategy Nash equilibria following the empirical literature on entry-exit games (e.g. Tamer:03:ReStud, Jia:08:Eca, Ciliberto/Tamer:09:Eca, among many others).} It follows that for $i = 1,2$,
\begin{align}
\rho_{\beta,\gamma,i}(x,\varepsilon;\eta)-\int \rho_{\beta,\gamma,i}(x,\varepsilon;\eta) d\mu(\eta)
= \left(\eta-\frac{1}{2}\right) 1\{(\varepsilon_1,\varepsilon_2) \in A_{3 \cap 4,\beta,\gamma}(x)\}.
\end{align}
Hence, by Theorem (ref), the LRR set based on the ASO of $Y_1$ is the same as that based on the ASO of $Y_2$. Writing the mean-deviated reduced form on the left hand side of ((ref)) as $\Delta_{\beta,\gamma}(x,\varepsilon;\eta)$, we find that
\begin{align*}
\Delta_{\beta,\gamma}^2(x,\varepsilon;\eta) = \left(\eta-\frac{1}{2}\right)^2 1\{(\varepsilon_1,\varepsilon_2) \in A_{3 \cap 4,\beta,\gamma}(x)\}.
\end{align*}
Plugging this into ((ref)), we find that
\begin{align}
Q^\mathsf{LRR}(\beta,\gamma) &= \int \int \left(\eta-\frac{1}{2}\right)^2 d\mu(\eta) 1\{(\varepsilon_1,\varepsilon_2) \in A_{3 \cap 4,\beta,\gamma}(x)\} d\tilde F_{X,\varepsilon}(x,\varepsilon) \notag \\
&= \frac{1}{4}\int 1\{(\varepsilon_1,\varepsilon_2) \in A_{3 \cap 4,\beta,\gamma}(x)\} d\tilde F_{X,\varepsilon}(x,\varepsilon).
\end{align}
Under the assumption of independence of $(X, \varepsilon)$, we can rewrite ((ref)) as:
\begin{align}
Q^\mathsf{LRR}(\beta,\gamma) &= \frac{1}{4}\int P\left\{(\varepsilon_1,\varepsilon_2) \in A_{3 \cap 4,\beta,\gamma}(x)\right\} d\tilde F_X(x).
\end{align}
We find the LRR by using this $Q^\mathsf{LRR}(\beta,\gamma)$ in ((ref)).
When $(\varepsilon_1,\varepsilon_2)$ fall in the region $A_{4,\beta,\gamma}(x)\cap A_{3,\beta,\gamma}(x)$, the game exhibits multiple equilibria. Our counterfactuals are most robust if this region is minimal. We cannot altogether disregard this region, because observations in the data may permit this region and it is possible they may arise in the counterfactual setting. Nevertheless, we can consider the value of the nuisance parameter(s) (in their identified set) that minimize the average region of multiple equilibria for the counterfactuals. (See Figure (ref) for an illustration of this observation.)
\@startsection{subsubsection}{3}
\z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{A Switching Regression Model with Interval Outcomes }
Let us look for $Q^{\mathsf{LRR}}(\beta,\gamma)$ for the example from Section (ref). We take $G$ to be a discrete distribution over $\{0,1\}$ and $\mu$ be the uniform distribution over $\{0,1\}$. Hence if we write
\begin{align}
\rho_{\beta_0,\gamma_0}(X,\gamma_0;\eta) = g_1\left(\beta_0 D + X_1'\gamma_0 \right) \eta + g_2\left(\beta_0 D + X_1'\gamma_0 \right) (1 - \eta) + \varepsilon,
\end{align}
we find that
\begin{align*}
\Delta_{\beta_0,\gamma_0}(X,\varepsilon;\eta) &= \rho_{\beta_0,\gamma_0}(X,\varepsilon;\eta) - \int \rho_{\beta_0,\gamma_0}(X,\varepsilon;\eta) d\mu(\eta) \\
&= \left(\eta - \frac{1}{2}\right) \left(g_1\left(\beta_0 D + X_1'\gamma_0 \right) - g_2\left(\beta_0 D + X_1'\gamma_0 \right)\right).
\end{align*}
Therefore, given a counterfactual distribution $\tilde F(x,\varepsilon)$ of $(X,\varepsilon)$,
\begin{align}
Q^{\mathsf{LRR}}(\beta,\gamma) &= \int \int \Delta_{\beta,\gamma}^2(x,\varepsilon;\eta) d\mu(\eta) d\tilde F(x,\varepsilon) \\
&= \frac{1}{4} \int \left(g_1\left(\beta d + x_1'\gamma \right) - g_2\left(\beta d + x_1'\gamma\right)\right)^2 d \tilde F_{X_1,D}(x_1,d),
\end{align}
where $\tilde F_{X_1,D}$ denotes the joint distribution of $(X_1,D)$ under $\tilde F$. Thus, the $\kappa$-LRR of $\gamma_0$ can be rewritten as
\begin{align}
\Gamma_{\kappa}^{\mathsf{LRR}}(\beta) = \left\{\gamma \in \Gamma_P: Q^\mathsf{LRR}(\beta,\gamma) \le Q^\mathsf{LRR}(\beta,\bar \gamma) + \kappa, \text{ for all } \bar \gamma \in \Gamma_P \right\},
\end{align}
where $\Gamma_P$ is as defined in ((ref)). Then we obtain the LRR set:
\begin{align}
\Theta_{\kappa,P}^{\mathsf{LRR}} = \left\{(\beta,\gamma): \gamma \in \Gamma_{\kappa}^{\mathsf{LRR}}(\beta), \beta \in B \right\}.
\end{align}
As mentioned previously, using this LRR set, one may construct the LRR set for other target parameters. Suppose that our ultimate interest is the Average Treatment Effect (ATE) denoted by $\psi(\beta_0,\gamma_0,\pi_0)$, where
\begin{align}
\psi(\beta,\gamma,\pi) & = \mathbf{E}\left[ \alpha(X_1,1;\beta,\gamma,\pi) - \alpha(X_1,0;\beta,\gamma,\pi) \right],
\end{align}
where $\alpha(X_1,D;\beta,\gamma,\pi)$ is as defined in ((ref)). In this case, the LRR set for $\psi(\beta_0,\gamma_0,\pi_0)$ is given by
\begin{align}
\Psi_{\kappa,P}^{\mathsf{LRR}} = \left\{\psi(\beta,\gamma,\pi): (\beta,\gamma) \in \Theta_{\kappa,P}^{\mathsf{LRR}}(\pi),\pi \in \Pi \right\},
\end{align}
where $\Theta_{\kappa,P}^{\mathsf{LRR}}(\pi)$ is the same as $\Theta_{\kappa,P}^{\mathsf{LRR}}$, except that $\Theta_P$ is replaced by
\begin{equation}
\Theta_P(\pi) = \left\{(\beta,\gamma): \mathbf{E}[\tilde Y_1 |X_1,D] \leq \alpha(X_1,D;\beta,\gamma,\pi) \leq \mathbf{E}[\tilde Y_2 |X_1, D], \text{a.e.} \right\}.
\end{equation}
\@startsection{section}{1}
\z@{.5\linespacing\@plus.5\linespacing}{.4\linespacing}{Inference with the Locally Robust Refinement}
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Inference from Moment Inequality Models}
For the sake of concreteness, we focus on the case where the parameters are identified through moment inequality models and estimated using their sample moments.\footnote{Our LRR approach does not depend on a particular way in which partial identification of a parameter arises (i.e. whether by moment inequality models or alternative models). Hence, one can combine the LRR approach with other forms of partial identification by refining the identified set using the LRR criterion proposed in this paper. For the sake of concreteness, we focus on the case where partial identification arises from moment inequality restrictions as this is the way partial identification often arises in practice.} Let us consider the following moment inequality models:
\begin{align}
\mathbf{E}_P[m_j(Z_i;\beta_0,\gamma_0)] \le 0, \text{ for } j = 1,...,p,
\end{align}
for an observed i.i.d. random vector $Z_i$, where $m_j(\cdot;\beta_0,\gamma_0)$'s are moment functions. We assume that the moment restrictions define the identified set $\Theta_P$ so that for all $(\beta,\gamma) \in \Theta_P$, the moment restrictions are satisfied.
Define the sample analogue of the moment in ((ref)):
\begin{align}
\overline m_j(\beta,\gamma) = \frac{1}{n}\sum_{i=1}^n m_j(Z_i;\beta,\gamma).
\end{align}
Let $\hat \sigma_j^2(\beta,\gamma)$ be the estimated variance of $m_j(Z_i;\beta,\gamma)$, i.e.,
\begin{align}
\hat \sigma_j^2(\beta,\gamma) = \frac{1}{n}\sum_{i = 1}^n \left(m_j(Z_i;\beta,\gamma) - \overline m_j(\beta,\gamma)\right)^2.
\end{align}
We let for $\kappa \geq 0$ (for instance, $\kappa = 0.01, 0.02$ or $0.03$)\footnote{The parameter $\kappa$ makes the inference slightly more conservative, which facilitates the development of the uniform inference (see the Supplemental Note). The choice of a positive constant $\kappa$ does not affect the uniform validity of the inference, as long as it is a fixed constant. In practice, values of $\kappa$ between 0.01 to 0.03 appear to work very well, as we show in the Monte Carlo simulations in Section 4 and the empirical application (Section 5).}
\begin{align}
\hat Q_{\kappa}(\beta,\gamma) = \sum_{j=1}^p\left[\frac{\overline m_j(\beta,\gamma)}{\hat \sigma_j(\beta,\gamma)} - \kappa \right]_+,
\end{align}
where $[a]_+ = \max\{a,0\}$, $a \in \mathbf{R}$. When $\kappa = 0$, we simply write $\hat Q(\beta,\gamma) = \hat Q_{0}(\beta,\gamma)$.
Define
\begin{align}
\hat \Gamma_{\kappa}(\beta) =\left\{\gamma \in \Gamma: \hat Q_\kappa(\beta,\gamma) = 0 \right\}.
\end{align}
Then, we construct $T(\beta,\gamma)$ as
\begin{align}
T(\beta,\gamma) = \sqrt{n}\hat Q(\beta,\gamma).
\end{align}
As for critical values, we may consider two different approaches. The first approach is based on least favorable configurations. The second approach is a more refined version with enhanced power properties but with a higher computational cost.
As for the first approach, we consider the following:\footnote{We take this form of a test statistic for the sake of concreteness. One can use different functionals to produce different test statistics. See Andrews/Shi:13:Eca for various functionals.}
\begin{align}
\hat Q^*(\beta,\gamma) = \sum_{j=1}^p\left[\frac{\overline m_j^*(\beta,\gamma) - \overline m_j(\beta,\gamma)}{\hat \sigma_j(\beta,\gamma)}\right]_+,
\end{align}
where
\begin{align}
\overline m_j^*(\beta,\gamma) = \frac{1}{n}\sum_{i=1}^n m_j(Z_i^*;\beta, \gamma),
\end{align}
and $Z_i^*$'s are resampled from the empirical distribution of $Z_i$'s with replacement. Let
\begin{align}
\hat T^*(\beta,\gamma) = \sqrt{n} \hat Q^*(\beta,\gamma).
\end{align}
We take critical values $\hat c_{1-\alpha}(\beta,\gamma)$ to be the $1-\alpha$ quantile of the bootstrap distribution of $\hat T^*(\beta,\gamma)$. Then, we construct the confidence region for $(\beta_0,\gamma_0)$ as follows:
\begin{align}
C_{1-\alpha}^\mathsf{LRR} = \left\{ (\beta,\gamma) \in \Theta: T(\beta,\gamma) \le \hat c_{1-\alpha}(\beta,\gamma), \text{ and } \gamma \in \hat \Gamma_{\kappa}^{\mathsf{LRR}}(\beta) \right\},
\end{align}
where
\begin{align}
\hat \Gamma_{\kappa}^{\mathsf{LRR}}(\beta) = \left\{\gamma \in \hat \Gamma_{\kappa}(\beta): \hat Q^\mathsf{LRR}(\beta,\gamma) \le \inf_{\bar \gamma \in \hat \Gamma_{-\kappa}(\beta)} \hat Q^\mathsf{LRR}(\beta,\bar \gamma) + 2 \kappa \right\}
\end{align}
and $ \hat \Gamma_{\kappa}(\beta) $ is as defined in ((ref)), and $\hat Q^\mathsf{LRR}(\beta, \gamma)$ denotes a consistent estimator of $Q^\mathsf{LRR}(\beta,\gamma)$ defined in ((ref)). (For example, in models of entry games or interval observations, this consistent estimator can be obtained by applying the sample analogue principle to the moments in ((ref)) and ((ref)), once one parametrically specifies the distributions of $\varepsilon_1,\varepsilon_2,$ and $\varepsilon$.) We take the infimum of an empty set to be infinity. With $\gamma$ restricted to be in $\hat \Gamma_{\kappa}^{\mathsf{LRR}}(\beta)$, our confidence interval reflects the focus on the parameter values obeying the LRR hypothesis.
As for the second approach, there are various alternative ways to construct inference that improves upon the one based on least favorable configurations. For example, one may use moment selection procedures as in Hansen:05:ET, Andrews/Soares:10:Eca and Andrews/Shi:13:Eca, the contact set method in Linton/Song/Whang:10:JOE, or a Bonferroni-based procedure as in Romano/Shaikh/Wolf:14:Eca. Here we adapt the Bonferroni-based procedure to our context.
First, we take $0< \alpha_1 < \alpha$ (say, $\alpha_1 = 0.005$) and find $\hat \kappa_{\alpha_1}(\beta,\gamma)$ using a bootstrap procedure. We define $\hat \kappa_{\alpha_1}(\beta,\gamma)$ as the $\alpha_1$ quantile of the bootstrap distribution of
\begin{align}
\min_{1 \le j \le k} \frac{\sqrt{n}(\overline m_j^*(\beta,\gamma) - \overline m_j(\beta,\gamma))}{\hat \sigma_j(\beta,\gamma)} ,
\end{align}
which uses $ \overline m_j^*(\beta,\gamma)$ defined in ((ref)). The uniform asymptotic validity of such a bootstrap procedure is well known in the literature. (See also Section (ref) in the Supplemental Note for details.) Define
\begin{align}
\hat \lambda_{j,\alpha_1}(\beta,\gamma) = \min \left\{\overline m_j(\beta,\gamma) - \frac{\hat \sigma_j(\beta,\gamma) \hat \kappa_{\alpha_1}(\beta,\gamma)}{\sqrt{n}},0\right\}.
\end{align}
Let
\begin{align}
\tilde Q^*(\beta,\gamma) = \sum_{j=1}^p\left[\frac{\overline m_j^*(\beta,\gamma) - \overline m_j(\beta,\gamma) + \hat \lambda_{j,\alpha_1}(\beta,\gamma)}{\hat \sigma_j(\beta,\gamma)}\right]_+.
\end{align}
We construct a bootstrap test statistic as follows:
\begin{align}
\tilde T^*(\beta,\gamma) = \sqrt{n} \tilde Q^*(\beta,\gamma).
\end{align}
We take critical values $\tilde c_{1-\alpha + \alpha_1}(\beta,\gamma)$ to be the $1-\alpha + \alpha_1$ quantile of the bootstrap distribution of $\tilde T^*(\beta,\gamma)$. Then, we construct the confidence region for $(\beta_0,\gamma_0)$ as follows:
\begin{align}
\tilde C^{\mathsf{LRR}}_{1-\alpha} = \left\{ (\beta, \gamma)\in \Theta: T(\beta,\gamma) \le \tilde c_{1-\alpha + \alpha_1}(\beta,\gamma) \text{ and } \gamma \in \hat \Gamma_{\kappa}^{\mathsf{LRR}}(\beta) \right\}.
\end{align}
In the Supplemental Note, we establish the uniform asymptotic validity of the confidence region. In the next section, we describe how to implement this procedure for Example (ref) with both approaches to critical values.
As emphasized in the introduction, the LRR-based confidence region $\tilde C^{\mathsf{LRR}}_{1-\alpha}$ should not be confused with one based on the identified set. This LRR-based confidence region does not reflect the empirical content of the model. Rather it shows the confidence region when the true parameter belongs to the LRR set, that is, the true parameter belongs to the set of parameter values at which the ASF is robust to the local perturbations of the equilibrium selection rules. One may use $\tilde C^{\mathsf{LRR}}_{1-\alpha}$ as a complement to the confidence region based on the identified set for $(\beta_0, \gamma_0)$ in empirical and policy analysis.
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Monte Carlo Simulations}
\@startsection{subsubsection}{3}
\z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Simulation Design}
In this section, we present results from a Monte Carlo simulation study on the finite sample properties of the inference procedure developed previously. We focus on the switching regression model example considered in Section (ref), which will be the main specification of our empirical application introduced in the next section.
Let $\tilde Y_1, \tilde Y_2$ be as in Section (ref) and $\alpha(X_1,D; \beta_0,\gamma_0, \pi)$ as in ((ref)) with the following specification, also motivated by our empirical application:
\begin{align}
g_1(\beta_0 D + X_1'\gamma_0) &= \beta_0 D + X_1'\gamma_0, \text{ and } \\ \notag
g_2(\beta_0 D + X_1'\gamma_0) &= \log\left(g_1(\beta_0 D + X_1'\gamma_0)\right).
\end{align}
We take $X_1 = 1$ for simplicity, we draw $\eta$ independently and with equal probability (i.e. $\pi_0 = 0.5$, implying an equal share of $\eta = 1$ and $\eta = 0$ in the population), and treatment is drawn independently for all $i$ with $D_i = 1$ with probability 0.3. The moment inequalities we have are as follows\footnote{
These inequalities can be generalized for a multidimensional $X_1$ by changing the indicator function on the left hand side to accommodate the different values of $X_1$.}: for $x_1 = 1$ and $d \in \{0,1\}$,
\begin{align}
\mathbf{E}\left[(\tilde Y_1 - \alpha(1,d; \beta_0,\gamma_0, \pi_0)) 1\{D = d\}\right] &\le 0, \text{ and }\\
\mathbf{E}\left[(\alpha(1,d;\beta_0,\gamma_0, \pi_0) - \tilde Y_2) 1\{D = d\}\right] &\le 0.
\end{align}
We consider two cases in our simulations: one where parameters $(\gamma_0, \beta_0, \pi_0)$ are such that the moments are close to equalities, and the other where they are not. For the first case, which we call Specification 1, we set $(\gamma_0,\beta_0, Z_1, Z_2) = (1,0.15,2.4,7)$. The choice of $\beta_0 = 0.15$ implies a return to treatment (for $\eta=1$) equal to 15%. This specification corresponds to approximately 5% of observations being top-coded.\footnote{As can be seen by the identified set ((ref)), setting one moment as an equality leaves the other moment as a strict inequality. Therefore, we do not consider a “least-favorable” set-up with all moments holding as equalities.} The second specification uses parameter values for $(\gamma_0, \beta_0, Z_1, Z_2)= (1,0.15,1.5,7)$ so that the moment inequalities are far from binding. This is done by keeping $(\gamma_0, \beta_0, Z_2)$ the same as the first specification, but decreasing $Z_1$ so that there are now approximately 10% of observations top-coded, since we have an increased range of outcomes that are unobserved. In general, the identified set from the second specification will be larger than the first.
For both cases, we show results for the identified set and for using LRR with and without the Bonferroni-based method of Romano/Shaikh/Wolf:14:Eca. For the LRR, we set $\kappa = 0.03$ throughout. The confidence interval follows the computation in ((ref)). We set the nominal size at 0.05, the simulation number to 500, and the bootstrap number to 999 under a sample size of $n=500$. In this set-up, our parameter space is such that $\beta_0 \geq 0$ and the ATE given in ((ref)) equals 0.145.
\@startsection{subsubsection}{3}
\z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Results}
The results across specifications and procedures are shown in Table (ref). They show the coverage of both the identified set and the LRR, as well as the length of the confidence sets. From the results, it is immediate that the LRR provides a smaller set of ATE $\psi(\beta_0,\gamma_0,\pi_0)$, implying less variation on predicted (counterfactual) outcomes $Y_i$. (The identified set, on the other hand, presents all possible values consistent with the data.) The coverage of the theoretical LRR is appropriate (at least 95%) for values in the interior of the set, calculated using equations ((ref)) and ((ref)).\footnote{In simulations available with the authors, the results improve as $n$ grows, as can be expected by our theoretical results. However, the computational cost also grows in these cases, particularly for the Bonferroni-based modification.}
\begin{center}
[INSERT Table (ref) HERE]
\end{center}
We also find that the Bonferroni-based method of Romano/Shaikh/Wolf:14:Eca is effective at reducing the length of the confidence intervals, particularly for the identified set. It is most effective when the DGP is away from the configuration that is closest to moment equalities. However, this modification comes at the computational cost of calculating the term $\hat \kappa_{\alpha_1}$ via a bootstrap procedure for every value of $(\beta, \gamma)$ in a first step. This step could be time consuming when the dimension of $\gamma$ is large.
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Predictions Using the LRR and the Identified Set}
We conduct two additional Monte Carlo exercises to provide additional insights on the benefits of using the Locally Robust Refinement.
Our first exercise compares the predictions for our parameter of interest ($\psi$) when using the LRR compared to using the identified set across Specifications 1 and 2. It illustrates when the LRR's predictions are closest to the true parameter values and when they are most informative (i.e. when the predictions vary the least).
To do so, we draw $s=1,...,1000$ values $\psi^s$ from the average (conservative) 95% confidence set for the LRR or the identified set described in Table (ref). These sets would be the predictions reported by an empirical researcher, who does not observe or point-identify the true $\psi(\beta_0, \gamma_0, \pi_0)$ due to top-coding. We draw these vectors in 2 different ways: (i) uniformly (i.e. any parameter within the appropriate set is drawn with equal probability) and (ii) furthest-case (i.e. we draw the parameters in each set that generate average predictions furthest away from the true values).\footnote{These sampling schemes represent different scenarios that illustrate the performance of our refinement. They may be thought of as different researcher priors on the true parameter. One can also perform the exercise under intermediate weighting schemes, obtaining similar conclusions.} To showcase the results, in Figure (ref), we report the distribution of the absolute deviation $\left|\psi^s - \psi(\beta_0, \gamma_0,\pi_0)\right|$: i.e. how do the predictions using the LRR and the identified set differ from the true value.
\begin{center}
[INSERT FIGURES (ref) AND (ref) HERE]
\end{center}
As we can see, the LRR's predictions vary less than those with the identified set (they are often at least 30% shorter), regardless of the specification or how the parameters are drawn. The use of the LRR seems particularly favorable when (i) the identified set is large (Specification 2), and (ii) when there is a significant chance of drawing extreme values of parameters from the identified set. This difference is most salient when the furthest-care scenario is particularly important to the researcher (for instance, when a bad policy could be very costly in terms of welfare). Afterall, the LRR's objective is not about mean predictions, but rather as minimizing the worst case.
The specifications above satisfy the LRR hypothesis, since the true $\psi(\beta_0, \gamma_0,\pi_0)$ is in the LRR set. Does the LRR still preserve these favorable properties even if the LRR hypothesis is not satisfied? We investigate this question in a last exercise. We set $(\beta_0, Z_1, Z_2)$ to the same values as Specification 1, but we increase $\gamma_0$ to 2.5. As a result, $\gamma_0$ is no longer in the LRR. We then compare the performance of the LRR and of the identified set in predicting the Average Structural Outcome (ASO) obtained using ((ref)).\footnote{In this setting, $\psi(\beta_0, \gamma_0,\pi_0)$ will generally satisfy the LRR hypothesis due to the parametrizations in this example. Hence as a target parameter that is not in the LRR set in this design, we consider the ASO in this exercise.} This outcome is unobserved, but it can be predicted based on the LRR and the identified set. The results are shown in Figure (ref). In contrast to Figure (ref), some draws of the identified set can outpredict the LRR. Afterall, the former contains the true parameters which are now absent in the LRR. Nevertheless, the LRR still generates more stable predictions, and its gains are still most noticeable and relevant when the identified set is large and/or worst-case outcomes are likely. Hence, while our results rely on the simulation designs motivated by our empirical application, the approach of LRR seems to generates predictions with stable performance regardless of whether the LRR hypothesis is satisfied or not.
\@startsection{section}{1}
\z@{.5\linespacing\@plus.5\linespacing}{.4\linespacing}{Empirical Illustration}
We provide an empirical application to illustrate how our procedure works with data. The application considers the case of top-coded data, as in the Monte Carlo simulations. We focus on the returns to college education, a topic subject to a large body of work in economics (see Oreopoulos/Petronijevic:13:ERIC for a survey). For the purpose of illustration, we choose a simplified approach and focus on the ATE defined in ((ref)).\footnote{One could extend our empirical analysis to incorporate potential endogeneity of the policy variable by using instrumental variables, and considering a Local Average Treatment Effect (LATE) which is the returns of college education on wages for the subgroup of individuals who are induced to attend college by a change in the instrument. Or alternatively, one could focus on the Marginal Treatment Effect or a Marginal Policy Relevant Treatment Effect as emphasized by Carneiro/Heckman/Vytlacil:11:AER.} Depending on the question of interest, such a parameter could allow the evaluation of policies meant to increase enrollment, such as changing taxes or subsidies for college (see Dynarski/Scott:17:Tax for an overview of different policies in place in the U.S. and their associated costs). These policies could have large fiscal and welfare consequences.
In this set-up, the outcome ($Y_i$) is the log of the hourly wages, the policy treatment is attending college (i.e. $D_i \in \{0,1\}$, where $D_i = 1$ represents that $i$ attended college and $0$ otherwise) and $X_{1,i}$ are individual covariates such as gender and/or race. Finally, the decision to attend college may depend on unobserved types: some types may have higher wages and/or benefit more from college. We model this by the random variable $\eta$, which is assumed to be binary: the returns from treatment $D_i = 1$ may depend on whether unobserved type is high ($\eta = 1$) or low ($\eta = 0$) and this difference can exist throughout the distribution of $X_{1,i}$.\footnote{For illustration purposes, we assume $\eta$ is binary and independent of $X_1, D$. Both can be generalized, but that is beyond the scope of our exercise.}
Since top-coding is prevalent in wage data (it is present in such widely used datasets as the NLSY and the Current Population Survey - CPS) and such top-coding is not random (for example, college educated individuals are more likely to be top-coded since they have higher wages on average), this set-up is nested within Example (ref). Hence, the identified set is given in ((ref)) and the LRR in ((ref)) can be used with the expressions defined in and above ((ref)).
We follow Romano/Shaikh:10:Eca and use data from the 2000 Annual Demographic Supplement of the Current Population Survey (CPS), keeping individuals aged 20 to 24, white, with a primary source of income from wages and salaries and that worked at least 2 hours per week on average. Akin to their work, our covariates $X_{1,i}$ are a constant and a binary variable for Male/Female. We differ from their sample by keeping both college graduates and non-college graduates, since that is our counterfactual of interest. This yields a sample of 5816 individuals, of which 826 (14%) are treated ($D_i = 1$). Details on the data construction are provided in the Supplemental Note.\footnote{Romano/Shaikh:10:Eca took the empirical distribution of the data as a “true” probability distribution to perform Monte Carlo simulations. Here, as in typical empirical research, we regard the data as a sample from an unknown population.} We then explore the effects of different amounts of top-coding on the results.
Following the large literature on Mincerian equations (see Heckman/Humphries/Veramendi:18:JPE for a discussion), we use a linear specification for $g_1$:
\begin{align*}
g_1(\beta_0 D + X_1'\gamma_0) =\beta_0 D + X_1'\gamma_0.
\end{align*}
For the reduced form for $\eta=0$, we parametrize $g_2(\beta_0 D + X_1'\gamma_0) = \log(g_1(\beta_0 D + X_1'\gamma_0))$.\footnote{While this is an illustration and other specifications could be used, this nonlinearity allows for increasing differentials across types, while its nonseparability allows for heterogeneous returns to college based on observable characteristics (see Barrow/Rouse:AERPP:05 for one such study).} We present confidence sets for both the identified set for the ATE, $\psi_0 = \psi(\beta_0,\gamma_0,\pi_0)$, and the LRR for $\psi_0$ using both conservative and a Bonferroni-correction based inference. We present results for this data under different top-coding amounts, imposing either 5% or 10% top-coding in the sample, under a value of top possible wages of $Z_2 = 10^8$.\footnote{We also set the upper bound on $\beta_0$ to 100% (or 25% per year), above all estimates typically found in the literature (see Oreopoulos/Petronijevic:13:ERIC). This is a natural restriction using the literature's results, and helps speed up computation as we do not search for a priori unreasonable values of $\beta_0$).} The results are shown below. As standard, we report the returns to a year of college (i.e. $\psi_0/4$).
\begin{center}
[INSERT TABLE (ref) HERE]
\end{center}
Our estimates are largely consistent with the literature, including the ATE found in works as Carneiro/Heckman/Vytlacil:11:AER, despite the different sample. While the identified set is large, including ATE close to 0 and others close to 25%, the LRR is significantly shorter, between 3 and 11% (for 5% top-coding) and between 0 and 10% (for 10% top-coding). These results are also robust to different choices of $\kappa$.\footnote{We used $\kappa = 0.03$ which was shown to work well in the Monte Carlo simulations. When setting $\kappa = 0.01$, the LRR confidence sets are very similar (and shorter, since a larger $\kappa$ is more conservative): in the conservative inference case, they are $[0.061, 0.091]$ for 5% top-coding and $[0, 0.091]$ for 10% top-coding.}
To see the usefulness of the refinement in this context, consider one of the policies motivating this exercise. This could be a change in subsidies to college or expanding college loans with the purpose of a large increase in college attendance. In the presence of top-coding, the identified set cannot reject large rates of return for college (e.g. above 12% a year) - suggesting a rationale for financially costly policies. However, the LRR rejects these larger values, finding rates of return that are lower than 11% and more compatible with those observed in the literature. The LRR suggests that more costly counterfactual policies should not be undertaken. This rejection is not due to the informational content of the model, which is still captured by the identified set, but rather from the robustness of the counterfactual predictions. The larger returns to college found in the identified set seem to be driven by features of the model the researcher is less confident about.
\@startsection{section}{1}
\z@{.5\linespacing\@plus.5\linespacing}{.4\linespacing}{Conclusion}
This paper explores methods to deal with models with multiple reduced forms for counterfactual analysis. As mentioned by Eizenberg:14:ReStud, there is an inherent tradeoff within the partial identification literature: while researchers may prefer not to impose strong identifying assumptions, they must confront the issue of counterfactuals with larger identified sets when doing policy analysis.
This paper points out that not all the values of the parameter are equally desirable for counterfactual analysis despite their empirical relevance. In particular, this paper focuses on the robustness property of the parameter values when the reduced form selection rule is locally perturbed. Thus the refinement looks at the subset of counterfactual scenarios that give most reliable predictions against a range of perturbations of the reduced forms (such as through changing equilibrium selection rules).
There could be other ways to discriminate different values of the identified set depending on the purpose of the counterfactual analysis in the application. For example, one may consider local misspecification of causal relationships between certain variables. This is plausible when the researcher is not sure about the direction of causality between certain variables. Similarly as we do in this paper, one may want to find a refinement that is robust to this type of misspecification. A fruitful formulation of this approach is left for future research.
\putbib[counterfactual]
\@startsection{section}{1}
\z@{.5\linespacing\@plus.5\linespacing}{.4\linespacing}*{Tables and Figures}
figure[figure omitted — 898 chars of source]
figure[figure omitted — 1,248 chars of source]
figure[figure omitted — 1,008 chars of source]
table[table omitted — 1,896 chars of source]
table[table omitted — 1,139 chars of source]
\FloatBarrier
\@startsection{section}{1}
\z@{.5\linespacing\@plus.5\linespacing}{.4\linespacing}{Appendix: Mathematical Proofs}
\@startsection{subsection}{2}
\z@{.3\linespacing\@plus.3\linespacing}{-.5em}{\normalfont}{Proof of Theorem (ref)}
Proof of Theorem (ref): Let $U$ and $\mathcal{W}$ be the sets from which $\eta$ and $W = (X,\varepsilon)$ take values respectively. Let $\mathcal{H}$ be the collection of measurable maps $h: \mathcal{W} \times U \rightarrow \mathbf{R}$ such that
align[align omitted — 146 chars of source]
We endow $\mathcal{H}$ with the inner-product $\langle \cdot,\cdot \rangle$ as follows: for $h_1,h_2 \in \mathcal{H}$,
align[align omitted — 94 chars of source]
Then $(\mathcal{H},\langle \cdot,\cdot \rangle)$ is a Hilbert space (up to an equivalence class). Recall that all distributions $G \in \mathcal{G}$ are dominated by the uniform distribution $\mu$ over $H$. Define
$\| h \|^2 = \langle h, h \rangle$ and write
align[align omitted — 93 chars of source]
and $\textsf{ASO}_{\beta,\gamma}(G) = \int \int \rho_{\beta,\gamma}(w;\eta)dG(\eta|w)d\tilde F(w)$. If we define
align[align omitted — 121 chars of source]
we can write
align*[align* omitted — 396 chars of source]
where the first equality comes because $dG'/d\mu - dG/d\mu \in \mathcal{H}$ (from the fact that each density integrates to one) and the second equality comes from the fact that for all $h \in \mathcal{H}$, $\int h(w,\eta)d\mu(\eta) = 0$. Since $\Delta_{\beta,\gamma} \in \mathcal{H}$ by Assumption (ref), using ((ref)) and Cauchy-Schwarz inequality, we can achieve the last supremum by taking $h'$ in the supremum to be $K \Delta_{\beta,\gamma}/\|\Delta_{\beta,\gamma}\|$. The result is nothing but $ \sqrt{Q^{\mathsf{LRR}}(\beta,\gamma)}$, completing the proof.\footnote{We note that the right hand side is not a function of $\delta$ or $K$. This is because the ASO is linear in $G$, and the dependence on those terms is removed by the difference operator defined on the left hand side.} $\blacksquare$
\FloatBarrier