EconBase
← Back to paper

Counterfactual Sensitivity and Robustness

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

126,961 characters · 0 sections · 89 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Counterfactual Sensitivity and Robustness

\defaultbibliography{sm} \defaultbibliographystyle{chicago}

bibunit\shortcites{BenTal2013} \thispagestyle{empty} \begin{abstract} \singlespacing We propose a framework for analyzing the sensitivity of counterfactuals to parametric assumptions about the distribution of latent variables in structural models. In particular, we derive bounds on counterfactuals as the distribution of latent variables spans nonparametric neighborhoods of a given parametric specification while other “structural” features of the model are maintained. Our approach recasts the infinite-dimensional problem of optimizing the counterfactual with respect to the distribution of latent variables (subject to model constraints) as a finite-dimensional convex program. We also develop an MPEC version of our method to further simplify computation in models with endogenous parameters (e.g., value functions) defined by equilibrium constraints. We propose plug-in estimators of the bounds and two methods for inference. We also show that our bounds converge to the sharp nonparametric bounds on counterfactuals as the neighborhood size becomes large. To illustrate the broad applicability of our procedure, we present empirical applications to matching models with transferable utility and dynamic discrete choice models. Keywords: Robustness, ambiguity, model uncertainty, misspecification, global sensitivity analysis. JEL codes: C14, C18, C54, D81 \end{abstract} \setcounter{page}{1} \pagenumbering{arabic} \section{Introduction} Researchers frequently make parametric assumptions about the distribution of latent variables in structural models. These assumptions are typically made for computational convenience\footnote{Examples include the conventional Gumbel (or type-I extreme value) assumption in discrete choice models following McFadden1974, dynamic discrete choice models following Rust, and matching models with transferable utility following Dagsvik and ChooSiow. Models of static or dynamic discrete games often impose parametric assumptions about the distribution of payoff shocks---see, e.g., Berry1992, AM2007, BBL, and CilibertoTamer.} or because simulation-based methods are used for estimation. In many models, such as those we consider in this paper, the distribution of latent variables is not nonparametrically identified. This raises the possibility that model parameters and the outcomes of policy experiments, or counterfactuals, may be only partially identified when parametric assumptions are relaxed. That is, different distributions may fit the data equally well in-sample, but may yield different values of the counterfactual. It is therefore natural to question whether counterfactuals are sensitive or robust to researchers' parametric assumptions, especially when evaluating the credibility of structural modeling exercises. This paper proposes a framework for analyzing the sensitivity of counterfactuals to parametric assumptions about the distribution of latent variables in a class of structural models. In particular, we derive bounds on counterfactuals as the distribution of latent variables spans nonparametric neighborhoods of a given parametric specification while other “structural” features of the model are maintained. This approach is in the spirit of global sensitivity analysis advocated by Leamer1985 (see also Tamer2015). Global sensitivity analyses are important in this context: many structural models are nonlinear so policy interventions can have different effects at different points in the parameter space. But a major difficulty with implementing global sensitivity analyses is tractability. A more tractable alternative are local sensitivity analyses, which are based on small perturbations around a chosen specification. Because local approaches rely on linearization, they may fail to correctly characterize the range of counterfactuals predicted by a nonlinear model when the distribution differs nontrivially from the researcher's chosen parametric specification. Our main insight is to borrow from the robustness literature in economics pioneered by HS2001 (HS2001, HS2008) to simplify computation using convex programming.\footnote{Our approach is also related to the field of distributionally robust optimization in operations research. See, e.g., Shapiro2017, DuchiNamkoong, and references therein.} Following this literature, we define neighborhoods around the researcher's parametric specification using statistical divergence (e.g., Kullback--Leibler divergence), with the option to add certain shape restrictions as appropriate. For tractability, we restrict our attention to models that may be written as a finite number of moment (in)equalities, where the expectation is with respect to the distribution of latent variables. While restrictive, this class accommodates many important models of static and dynamic discrete choice, discrete games, and matching. To describe our procedure, consider the problem of minimizing or maximizing the counterfactual at a fixed value of structural parameters by varying the distribution of latent variables over a neighborhood, subject to the model's (in)equality restrictions. We use duality to recast this infinite-dimensional optimization problem as a finite-dimensional convex program. The value of this inner program is treated as a criterion function, which is optimized in an outer optimization with respect to structural parameters. Importantly, the dimension of the inner problem is independent of the neighborhood size, making our procedure tractable over both small and large neighborhoods. To further simplify computation, we develop an MPEC version of our procedure for models featuring endogenous parameters (e.g., value functions) defined by equilibrium constraints. We show that this implementation can produce significant computational gains for dynamic discrete choice models in particular. Our approach is conceptually different from nonparametric partial identification analyses which derive bounds on counterfactuals under minimal distributional assumptions. But as we show, bounds computed using our procedure converge to the (sharp) nonparametric bounds in the limit as the neighborhood size becomes large. Aside from sensitivity analyses, our methods may therefore be used to approximate nonparametric bounds by taking the neighborhood size to be large but finite. For estimation and inference, we propose simple plug-in estimators of the bounds and establish their consistency. We also propose and theoretically justify two methods for inference: a computationally simple but conservative projection procedure and a relatively more efficient bootstrap procedure. We illustrate our procedures with two empirical applications. The first revisits the “marital college premium” estimates reported in CSW, which relied on an i.i.d. Gumbel (type-I extreme value) assumption for the distribution of individuals' idiosyncratic marital preferences (see also ChooSiow). The second empirical application performs a counterfactual welfare analysis in the canonical dynamic discrete choice model of Rust. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Related literature.} Our approach has connections with global prior sensitivity in Bayesian analysis ChamberlainLeamer,Leamer1982,Berger1984, most notably GKU and Ho who consider sets of priors constrained by Kullback--Leibler divergence relative to a default prior. Motivated by questions of sensitivity, CTT study inference in semiparametric likelihood models using sieve approximations for the infinite-dimensional nuisance parameter (the distribution of latent variables in our setting). For the class of moment-based models we consider, our approach instead eliminates the infinite-dimensional nuisance parameter via a convex program of fixed dimension. Several other works have used convex duality to characterize identified sets in models with latent variables. Most closely related are EGH and Schennach.\footnote{Works using other notions of “duality” to construct identified sets include BMM, GalichonHenry2011, ChesherRosen2017, and Li.} The problem we study is different, both because of its focus on counterfactuals, rather than structural parameters, and because the optimization is performed over a neighborhood, rather than over all distributions. As a consequence, our estimation and inference methods are also quite different. Torgovitsky2019QE uses linear programming to characterize sharp identified sets in latent variable models defined by quantile restrictions. Within this class, his approach is more computationally convenient than ours for characterizing identified sets. Several important moments or counterfactuals cannot be expressed as quantile restrictions, such as social surplus in discrete choice models and Bellman equations in dynamic discrete choice models. Our approach is compatible with these moments and counterfactuals, thereby allowing the user to characterize identified sets in broader classes of model as well as to perform sensitivity analyses. There is also a literature deriving nonparametric bounds in specific latent variable models. Examples include Manski2007,Manski2014, AllenRehbeck, TTY, Laffers, Torgovitsky, and GualdaniSinha. Most closely related is NoretsTang, who construct identified sets of counterfactual conditional choice probabilities (CCPs) in dynamic binary choice models. Their approach is specific to counterfactual CCPs and to dynamic binary choice models. Our approach allows for a wider range of counterfactual (e.g., welfare), shape restrictions, and multinomial choice, in addition to performing sensitivity analyses.\footnote{KSS and KKLS consider the converse problem, in which flow payoffs are nonparametric (as they can be in our setting) but the distribution of latent payoff shocks is known.} Finally, our work is complementary to the recent literature on local sensitivity---see, e.g., KOE, AGS2017,AGS2018, AK, BW, and Mukhin. Much of this literature is concerned with local misspecification of moment conditions, which is different from the setting we consider. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Outline.} Section (ref) introduces our procedure, estimators of the bounds, and shows our approach recovers nonparametric bounds as the neighborhood size becomes large. Section (ref) discusses practical aspects and implementation details. Section (ref) gives guidance for interpreting the neighborhood size. Empirical applications are presented in Section (ref). Section (ref) discusses estimation and inference. The online appendix presents extensions of our methodology, connections with local sensitivity analyses, additional empirical results, and proofs of our main results. A secondary online appendix presents background material on Orlicz classes and supplemental proofs. \section{Procedure} We begin in Section (ref) by describing the class of models to which our procedure may be applied. Section (ref) describes our approach, Section (ref) shows how duality is used to simplify the bounds, and Section (ref) introduces our estimators of the bounds. Section (ref) shows our bounds converge to the sharp nonparametric bounds as the neighborhood size becomes large. \subsection{Setup} We consider a class of models that link a structural parameter $\theta \in \Theta \subset \mb R^{d_\theta}$, a vector of targeted moments $P_0 \in \mathcal P \subseteq \mathbb R^{d_P}$, and possibly an auxiliary parameter $\gamma_0 \in \Gamma$ (a metric space) via the moment restrictions \begin{subequations} \begin{align} \mb E^F[ g_1(U,\theta,\gamma_0)] & \leq P_{10} , \\ \mb E^F[ g_2(U,\theta,\gamma_0)] & = P_{20} , \\ \mb E^F[ g_3(U,\theta,\gamma_0)] & \leq 0 , \\ \mb E^F[ g_4(U,\theta,\gamma_0)] & = 0 , \end{align} \end{subequations} where $g_1,\ldots,g_4$ are vectors of moment functions, $P_0 = (P_{10},P_{20})$ is partitioned conformably, and $\mb E^F$ denotes expectation with respect to a vector of latent variables $U \sim F$. We assume that the researcher has consistent estimators $(\hat P,\hat \gamma)$ of $(P_0,\gamma_0)$. We also assume that the researcher is interested in a (scalar) counterfactual of the form \begin{equation} \kappa = \mb E^F[k(U,\theta,\gamma_0)] \,. \end{equation} This setup accommodates counterfactuals that do not depend explicitly on $U$, in which case ((ref)) reduces to $\kappa = k(\theta,\gamma_0)$. Note that $\kappa$ will still depend on the distribution of $U$ through $\theta$, whose values are disciplined by the moment conditions ((ref)). Several models and counterfactuals of interest fall into this framework. We review three examples before proceeding. \begin{example}[Discrete choice and consumer welfare] \normalfont Suppose an individual derives utility $h_j(X,\theta) + U_{j}$ from choice $j \in \mc J_0 := \{0, 1, \ldots, J\}$, where $X \in \mc X$ are observed covariates and $U = (U_{j})_{j \in \mc J_0}$ is latent (to the econometrician). We assume, as typical, that $U$ is drawn independently across individuals from a continuous distribution $F$. The probability that an individual with characteristics $x$ chooses $j$ is \begin{equation} p(j|x) = \mb P_F \left( h_j(x,\theta) + U_{j} = \textstyle \max_{j' \in \mc J_0} \left( h_{j'}(x,\theta) + U_{j'} \right) \right)\,, \end{equation} where $\mb P_F$ denotes probabilities when $U \sim F$. In empirical work, $\theta$ is typically estimated using a criterion that fits the model-implied choice probabilities ((ref)) to probabilities observed in the data. Welfare analyses are often based on the social surplus McFadden1978 \[ W(x) = \mb E^F\left[\textstyle \max_{j \in \mc J_0} \left( h_{j}(x,\theta) + U_{j} \right) \right] , \] which is the average utility consumers with characteristics $x$ derive from the choice problem. A related welfare measure is the change in surplus $\Delta W(x_a,x_b) = W(x_a) - W(x_b)$ associated with a shift from $x_b$ to $x_a$. In practice, it is common to assume the $U_j$ are i.i.d. Gumbel (type-I extreme value), as this yields closed-form expressions for choice probabilities and the welfare measures $W(x)$ and $\Delta W(x_a,x_b)$. Our approach may be used to perform a sensitivity analysis of $W(x)$ and $\Delta W(x_a,x_b)$ to parametric assumptions about $F$ when $\mathcal X$ is finite. A leading example is matching models with finitely many agent types---see Section (ref) and references therein. Understanding the sensitivity of $W(x)$ and $\Delta W(x_a,x_b)$ to $F$ is important in this case because $W(x)$ and $\Delta W(x_a,x_b)$ are not nonparametrically identified.\footnote{See, e.g., BerryHaile2010,BerryHaile2014 and AllenRehbeck for nonparametric identification of utilities and welfare measures in discrete choice models when characteristics have continuous support.} In our notation, $g_2$ collects indicator functions representing the choice probabilities ((ref)) across covariates $x \in \mc X$ and choices $j \in \mc J := \{1,\ldots,J\}$ ($j = 0$ is redundant): \[ g_2(U,\theta) = \left( 1\!\mathrm{l} \left\{ \textstyle h_j(x,\theta) + U_{j} = \max_{j' \in \mc J_0} \left( h_{j'}(x,\theta) + U_{j'} \right) \right\} \right)_{(j,x) \in \mc J \times \mc X} \] and $P_{20} = (\Pr(j|x))_{(j,x) \in \mc J \times \mathcal X}$ is the vector of true choice probabilities. There are no $g_1$, $g_3$, $g_4$, or $\gamma$ in this model. Finally, $k(U,\theta) = \max_{j \in \mc J_0} \left( h_{j}(x,\theta) + U_{j} \right)$ for $W(x)$ and $k(U,\theta) = \max_{j \in \mc J_0} \left( h_{j}(x_a,\theta) + U_{j} \right) - \max_{j \in \mc J_0} \left( h_{j}(x_b,\theta) + U_{j} \right)$ for $\Delta W(x_a,x_b)$. $\hfill \square$ \end{example} \begin{example}[Discrete games] \normalfont Following BresnahanReiss,BresnahanReiss1991, Berry1992, and Tamer2003, consider the complete-information game in Table (ref). \begin{table}[h] \begin{center} \begin{tabular}{*{4}{c}} \multicolumn{2}{c} & \multicolumn{2}{c}{Firm $2$\phantom{00000000}}\\ \multicolumn{1}{c} & & $0$ & $1$ \\[4pt] \multirow{2}*{Firm $1$} & $0\phantom{000}$ & $(0,0)$ & $(0,\beta_2'x + U_2)$ \\ & $1\phantom{000}$ & $(\beta_1' x + U_1 ,0)$ & $(\beta_1' x - \Delta_1 + U_1,\beta_2 ' x - \Delta_2 + U_2)$ \\ \end{tabular} \end{center} \caption{\label{table:simplegame:regressor} Payoff matrix for (Firm 1, Firm 2) when $X = x$.} \end{table} Here $U = (U_1,U_2)$ is the latent (to the econometrician) component of firms' profits, which is independent of covariates $X$. Suppose that the solution concept is restricted to equilibria in pure strategies. The econometrician may estimate the probabilities of the potential market structures $(0,0)$, $(0,1)$, $(1,0)$, $(1,1)$ (conditional on $X$) from data on a large number of markets. As the model is incomplete---there are values of $U$ for which there are multiple equilibria---moment inequality methods are typically used in empirical work to avoid restricting the equilibrium selection mechanism. However, strong parametric assumptions are often made about the distribution of $U$ (typically bivariate Normal) to derive the model-implied probabilities for different market structures; see, e.g., \cite{Berry1992}, \cite{CilibertoTamer}, \cite{BMM}, and \cite{KlineTamer2016}. It therefore seems natural to also question the sensitivity of counterfactuals to parametric assumptions for $U$. This model falls into our setup when the regressors $X$ have finite support $\mc X$.\footnote{Continuous regressors are often discretized in empirical applications; see, e.g., \cite{CilibertoTamer}, \cite{Grieco}, \cite{KlineTamer2016}, and \cite{CCT}.} In our notation, $g_1$ collects the moment inequalities that bound the probabilities of $(0,1)$ and $(1,0)$ across $x \in \mc X$, with $P_{10}$ denoting the corresponding true probabilities. The inequalities are typically expressed as upper bounds on the probabilities of $(0,1)$ and $(1,0)$; we flip the sign to be compatible with (\ref{e:mod:1}): \[ \begin{aligned} g_1(U,\theta) & = \left[ \begin{array}{c} \left( - 1\!\mathrm{l}\{ U_1 \geq -\beta_1'x ; U_2 \leq \Delta_2-\beta_2'x\} \right)_{x \in \mc X} \\ \left( - 1\!\mathrm{l}\{ U_1 \leq \Delta_1-\beta_1'x ; U_2 \geq -\beta_2'x \} \right)_{x \in \mc X} \end{array} \right] , & P_{10} & = \left[ \begin{array}{c} \left( - \Pr((1,0)|X = x) \right)_{x \in \mc X} \\ \left( - \Pr((0,1)|X = x) \right)_{x \in \mc X} \end{array} \right] , \end{aligned} \] where $\theta = (\Delta_1,\Delta_2,\beta_1,\beta_2)$. Similarly, $g_2$ and $P_{20}$ collect the moment conditions and probabilities for outcomes $(0,0)$ and $(1,1)$, which are always realized as the result of unique equilibria: \begin{align*} g_2(U,\theta) & = \left[ \begin{array}{c} \left( 1\!\mathrm{l}\{ U_1 \leq -\beta_1'x; \, U_2 \leq -\beta_2'x\} \right)_{x \in \mc X} \\ \left( 1\!\mathrm{l}\{ U_1 \geq \Delta_1 - \beta_1'x; \, U_2 \geq \Delta_2 -\beta_2'x \} \right)_{x \in \mc X} \end{array} \right] , & P_{20} & = \left[ \begin{array}{c} \left( \Pr((0,0)|X = x) \right)_{x \in \mc X} \\ \left( \Pr((1,1)|X = x) \right)_{x \in \mc X} \end{array} \right] . \end{align*} There is no $g_3$, $g_4$, or $\gamma$ in this model. CilibertoTamer compute upper bounds on the probability of entrants under a counterfactual payoff shift, say $\tau(\theta)$. The function $k(U,\theta) = 1\!\mathrm{l}\{ U_1 \geq \tau(\theta) - \beta_1'x\}$ corresponds to the upper bound on the probability of firm 1 entering when $X = x$ under this counterfactual. $\hfill \square$ \end{example} \begin{example}[Dynamic discrete choice] \normalfont Consider a canonical dynamic discrete choice (DDC) model following Rust. The decision maker solves \begin{equation} V(s) = \mathbb E^{F}\left[ \max_{d \in \mc D_0} \left( \pi_{d,s}(\theta_\pi) + U_d + \beta E[V(s')| d, s] \right) \right] , \end{equation} where $s \in \mc S$ is a Markov state variable, $\mc D_0 = \{0,1,\ldots,D\}$ is the set of actions, $\pi_{d,s}$ is the flow payoff for action $d$ in state $s$ which is parameterized by $\theta_\pi$, $U_d$ is a latent payoff shock, $\beta \in (0,1)$ is a discount parameter, and $E[\,\cdot\,|d, s]$ denotes expectation with respect to the future state $s'$. The distribution $F$ of $U = (U_d)_{d \in \mc D_0}$ is typically assumed to be continuous and independent of $s$. The CCP of action $d$ in state $s$ is \begin{equation} p(d|s) = \mathbb P_F \left( \pi_{d,s}(\theta_\pi) + U_d + \beta E[V(s')| d, s] = \max_{d' \in \mc D_0} \left( \pi_{d',s}(\theta_\pi) + U_{d'} + \beta E[V(s')| d', s] \right) \right) , \end{equation} where $\mathbb P_F$ denotes probabilities when $U \sim F$. It is standard to assume the $U_d$ are i.i.d. Gumbel, as this yields closed-form expressions for the expectation in ((ref)) and multinomial-logit expressions for the CCPs ((ref)). Parameters $\theta_\pi$ or $(\theta_\pi,\beta)$ are typically estimated using a criterion function that fits the model-implied CCPs ((ref)) to probabilities observed in the data. Counterfactuals are then computed by solving ((ref)) under alternative laws of motion, flow payoffs, or other interventions. When $\mc S$ is finite, model parameters, counterfactual CCPs, and counterfactual welfare measures are typically not identified without parametric restrictions on $F$. Our procedure may be used perform a sensitivity analysis of counterfactuals to parametric assumptions on $F$ as follows. Let $\theta = (\theta_\pi,v,\tilde v)$ or $\theta = (\theta_\pi,\beta,v,\tilde v)$, where $v = (V(s))_{s \in \mathcal S}$ and $\tilde v = (\tilde V(s))_{s \in \mathcal S}$ collect the baseline and counterfactual value functions across $s \in \mc S$. Also let $\gamma = (M_d)_{d \in \mc D_0}$ collect the transition matrices for $s$, $g_2$ collect indicator functions for the CCPs ((ref)) across states $s \in \mc S$ and choices $d \in \mc D := \{1,\ldots,D\}$ ($d = 0$ is redundant): \[ g_2(U,\theta,\gamma) = \left( 1\!\mathrm{l} \left\{ \pi_{d,s}(\theta_\pi) + U_d + \beta M_{d,s} v = \max_{d' \in \mc D_0} \left( \pi_{d',s}(\theta_\pi) + U_{d'} + \beta M_{d',s} v \right) \right\} \right)_{(d,s) \in \mc D \times \mathcal S} \] with $M_{d,s}$ denoting the $s$th row of $M_d$, and $P_{20} = (\Pr(d|s))_{(d,s) \in \mc D \times \mathcal S}$ collect the corresponding true CCPs. Finally, $g_4$ collects moment functions representing ((ref)) in the baseline model and under the counterfactual: \begin{equation} g_4(U, \theta, \gamma) = \left[ \begin{array}{c} ( \max_{d \in \mc D_0} \{ \pi_{d,s}(\theta_\pi) + U_d + \beta M_{d,s} v \} - v_{s} )_{s \in \mathcal S}\\[8pt] ( \max_{d \in {\mc D}_0} \{ \tilde \pi_{d,s}(\theta_\pi) + U_d + \tilde \beta \tilde M_{d,s} \tilde v \} - \tilde v_{s} )_{s \in \mathcal S} \end{array} \right] , \label{eq:ex.ddc.fp} \end{equation} where $v_s = V(s)$, $\tilde v_s = \tilde V(s)$, and $\tilde \pi$, $\tilde \beta$, $\tilde M_d$ denote counterfactual flow payoffs, discount factor, and law of motion.\footnote{If $\mb E^F[\max_{d \in \mc D_0}U_d]$ is finite, then $v \mapsto (\mb E^F[ \max_{d \in \mc D_0} \{ \pi_{d,s}(\theta_\pi) + U_d + \beta M_{d,s} v \} ])_{s \in \mc S}$ is a $\ell^\infty$-contraction of modulus $\beta$ on $\mb R^{|\mc S|}$. Hence, there is a unique $(v,\tilde v)$ solving $\mathbb{E}^F[ g_4(U,\theta,\gamma)] = 0$ at any fixed $(\theta_\pi, \beta, \tilde \beta, F)$. The solution $(v,\tilde v)$ must collect the solutions to (\ref{e:rust-emax}) in the baseline model and counterfactual across states: $v = (V(s))_{s \in \mathcal S}$ and $\tilde v = (\tilde V(s))_{s \in \mathcal S}$. It follows that $F$ satisfies $\mathbb{E}^F[g_4(U,\theta,\gamma)] = 0$ at $\theta = (\theta_\pi,\beta,v,\tilde v)$ if and only if $(v,\tilde v)$ corresponds to the value functions $V$ and $\tilde V$ under $F$.} We recommend including the location normalizations $\mathbb{E}^F[U_d] = 0$ for $d \in \mc D_0$ in $g_4$ for interpretability. We also recommend including scale normalizations in $g_4$ so that $\mathbb{E}^F[\max_{d \in \mc D_0} U_d]$ is finite. For instance, in Section~\ref{s:rust} we normalize $\mathbb{E}^F[U_d^2]$ for all $d \in \mc D_0$. Counterfactual CCPs can be computed using \[ k(U, \theta, \gamma) = 1\!\mathrm{l} \left\{ \tilde \pi_{d,s}(\theta_\pi) + U_d + \tilde \beta \tilde M_{d,s} \tilde v = \max_{d' \in {\mc D}_0} \left( \tilde \pi_{d',s}(\theta_\pi) + U_{d'} + \tilde \beta \tilde M_{d',s} \tilde v \right) \right\} . \] Change in average welfare corresponds to $k(\theta,\gamma) = w' (\tilde v - v)$ for a weight vector $w$. $\square$ \end{example} \begin{remark} We allow for conditional moments models with $\mb E[g_1(U,X,\theta,\gamma)|X=x] \leq P_{10}(x)$ (and similarly for ((ref))-((ref))) if $U$ is independent of $X$ and $X$ takes values in a finite set $\mc X$. Moment functions are then stacked across $x \in \mc X$ to form $g_1$, $g_2$, $g_3$, and $g_4$ (see Examples (ref)-(ref)). Appendix (ref) discusses extensions to conditional moment models where the distribution of $U$ may vary with the value of (discrete) covariates, and to non-separable models with discrete covariates. Models with continuous covariates fall outside the scope of our procedure. \end{remark} \begin{remark} Our setup relies on the counterfactual being expressible as ((ref)). If $k$ is vector-valued, our procedure can be applied to compute the support function\footnote{A closed convex set is determined by its support function---see Rockafellar (Rockafellar, Section 13).} of the identified set of counterfactuals: set $k^\tau(U,\theta,\gamma) = \tau'k(U,\theta,\gamma)$ for a conformable unit vector $\tau$ and replace ((ref)) with $\kappa^\tau = \mathbb{E}^F[k^\tau(U,\theta,\gamma_0)]$. Our setup excludes counterfactuals that are infinite-dimensional, such as the distribution of the number of firms in a market. \end{remark} \begin{remark} The distribution $F$ is not nonparametrically identified in any of the above examples or, more generally, in the class of models ((ref)) when the support of $U$ contains many more points than there are moment conditions (e.g., when $U$ is continuously distributed). \end{remark} In common practice, a seemingly reasonable or computationally convenient distribution, say $F_*$, is assumed by the researcher and maintained throughout the analysis (e.g., bivariate Normal in Example (ref) and i.i.d. Gumbel in Examples (ref) and (ref)). Given $F_*$ and estimates $\hat P = (\hat P_1,\hat P_2)$ of $P_0$ and $\hat \gamma$ of $\gamma_0$, the researcher computes an estimate $\hat \theta$ of $\theta$ using a criterion function based on the moment conditions \begin{equation} \begin{aligned} \mb E^{F_*}[ g_1(U,\theta,\hat \gamma)] & \leq \hat P_{1} \,, & \mb E^{F_*}[ g_2(U,\theta,\hat \gamma)] & = \hat P_{2} \,, \\ \mb E^{F_*}[ g_3(U,\theta,\hat \gamma)] & \leq 0 \,, & \mb E^{F_*}[ g_4(U,\theta,\hat \gamma)] & = 0 \,. \end{aligned} \end{equation} Finally, the researcher estimates the counterfactual using $\hat \kappa = \mb E^{F_*}[k(U,\hat \theta,\hat \gamma)]$. If $k$ does not depend on $U$, then the estimated counterfactual is simply $\hat \kappa = k(\hat \theta,\hat \gamma)$. In this case $\hat \kappa$ will still depend {implicitly} on $F_*$ through $\hat \theta$.\footnote{While this discussion has assumed point identification of $\theta$ and $\kappa$ for sake of exposition, our methods allow structural parameters and counterfactuals to be partially identified.} The researcher's chosen specification $F_*$ is used both for estimation of $ \theta$ and again when computing the counterfactual. A natural question is: to what extent does the counterfactual depend on the choice of distribution? The main contribution of this paper is to provide a tractable econometric framework for answering this question. \subsection{Our Approach} As a sensitivity analysis, we shall relax the researcher's parametric assumption and allow $F$ to vary over nonparametric neighborhoods $\mc N_\delta$ of $F_*$, where $\delta$ is a measure of neighborhood “size”. When we do so, there may be multiple pairs $(\theta,F) \in \Theta \times \mc N_\delta$ that satisfy ((ref)) but which yield different values of the counterfactual. Our objects of interest are the smallest and largest values of the counterfactual over all such $(\theta,F)$ pairs: \begin{align} \ul \kappa_\delta & = \inf_{\theta \in \Theta,F \in \mc N_\delta} \mb E^F[k(U,\theta,\gamma_0)] \quad \text{ subject to ((ref))} , \\ \ol \kappa_\delta & = \sup_{\theta \in \Theta,F \in \mc N_\delta} \mb E^F[k(U,\theta,\gamma_0)] \quad \text{ subject to ((ref))} . \end{align} By focusing on $\ul \kappa_\delta$ and $\ol \kappa_\delta$, our approach naturally accommodates models with partially-identified structural parameters and counterfactuals. Our approach also sidesteps having to compute the identified set of structural parameters. The optimization problems ((ref)) and ((ref)) are made tractable by a convenient choice of $\mc N_\delta$. Following HS2001 and MMR, we consider neighborhoods constrained by $\phi$-divergence Csiszar1975: \begin{equation} \begin{aligned} \mc N_\delta & = \{ F \in \mc F : D_\phi(F\|F_*) \leq \delta \} \,, \\ D_\phi(F\|F_*) & = \left[ \begin{array}{ll} \int \phi \left( \frac{\mr d F}{\mr d F_*} \right)\, \mr d F_* & \mbox{if $F \ll F_*$},\\ +\infty & \mbox{otherwise}, \end{array} \right. \end{aligned} \end{equation} where $\mc F$ denotes all probability measures on the support\footnote{That is, $\mc U$ is the set of all values that $U$ could conceivably take according to the model, which is possibly larger that the support of the measure $F_*$.} $\mc U$ of $U$ and $F \ll F_*$ denotes absolute continuity of $F$ with respect to $F_*$. The convex function $\phi : [0,\infty) \to \mb R_+ \cup \{+\infty\}$ penalizes deviations of $F$ from $F_*$. For example, $\phi(x) = x \log x - x + 1$ corresponds to Kullback--Leibler (KL) divergence, $\phi(x) = \frac{1}{2}(x - 1)^2$ corresponds to Pearson $\chi^2$ divergence, and \[ \phi(x) = \frac{x^p - 1 - p(x - 1)}{p(p-1)}\,, \quad (p > 1)\,, \] corresponds to $L^p$ divergence. If $F_*$ has positive (Lebesgue) density, then the absolute continuity condition merely rules out $F$ with mass points. \begin{remark} Normalizations and other shape restrictions may be added by augmenting the moment functions $g_1,\ldots,g_4$. Examples include: (i) location normalizations, e.g. $\mb E^F[U] =0$ or $\mb E^F[1\!\mathrm{l} \{U_i \leq 0\} - 0.5] = 0$ for each element $U_i$ of $U$; (ii) scale normalizations, e.g. $\mb E^F[U_i^2] =1$; (iii) covariance normalizations, e.g. $\mb E^F[UU'] = I$; and (iv) smoothness restrictions, e.g. $\mb E^F[1\!\mathrm{l}\{U_i \leq a_{k+1}\} - 1\!\mathrm{l}\{U_i \leq a_{k}\}] \leq C$ for $a_1 < \ldots < a_K$ and a positive constant $C$. \end{remark} \begin{remark} Appendix (ref) shows that shape restrictions including symmetry, exchangeability, and, more generally, invariance under a finite group of transforms, are also easy to impose. \end{remark} \subsection{Dual Formulation} We use convex duality to simplify computation of $\ul \kappa_\delta$ and $\ol \kappa_\delta$. We start by noting $\ul \kappa_\delta$ and $\ol \kappa_\delta$ may be written as the solution to two profiled optimization problems: \begin{align*} \ul \kappa_\delta & = \inf_{\theta \in \Theta} \ul K_\delta(\theta;\gamma_0,P_0) \,, & \ol \kappa_\delta & = \sup_{\theta \in \Theta} \ol K_\delta(\theta;\gamma_0,P_0) \,, \end{align*} where the criterion functions $\ul K_\delta(\theta;\gamma_0,P_0) $ and $\ol K_\delta(\theta;\gamma_0,P_0) $ are, respectively, the infimum and supremum of $\mb E^F[k(U,\theta,\gamma_0)]$ with respect to $F \in \mc N_\delta$ subject to the moment conditions ((ref)). In what follows, it is helpful to define the criterion functions at a generic $(\gamma,P)$. To do so, we say that the moment conditions ((ref)) hold “at $(\theta,\gamma,P)$” if they hold when $\gamma_0$ is replaced by $\gamma$ and $P_0$ is replaced by $P$. Then \begin{align} \ul K_\delta(\theta;\gamma,P) & = \inf_{F \in \mc N_\delta} \mb E^F[k(U,\theta,\gamma)] \quad \mbox{subject to ((ref)) holding at $(\theta,\gamma,P)$} \,, \\ \ol K_\delta(\theta;\gamma,P) & = \sup_{F \in \mc N_\delta} \mb E^F[k(U,\theta,\gamma)] \quad \mbox{subject to ((ref)) holding at $(\theta,\gamma,P)$} \,, \end{align} with the understanding that $ \ul K_\delta(\theta;\gamma,P) = +\infty$ and $ \ol K_\delta(\theta;\gamma,P) = -\infty$ if there does not exist a distribution in $\mc N_\delta$ for which the moment conditions ((ref)) hold at $(\theta,\gamma,P)$. We first impose some mild regularity conditions on $F_*$, $\phi$, and the moment functions to justify the dual formulation. Similar conditions are used in generalized empirical likelihood estimation (see, e.g., KR). Let $\Phi_0$ denote the set of all $\phi : [0,\infty) \to \mb R \cup \{+\infty\}$ such that $\phi$ is continuously differentiable on $(0,+\infty)$ and strictly convex, with $\phi(1) = \phi'(1) = 0$, $\phi(0) < +\infty$, $\lim_{x \downarrow 0} \phi'(x) < 0$, $\lim_{x \to +\infty} \phi(x)/x = +\infty$, $\lim_{x \to +\infty} \phi'(x) > 0$, and $\lim_{x \to +\infty} x \phi'(x)/\phi(x) < + \infty$. The functions inducing KL, $\chi^2$, and $L^p$ divergence all belong to $\Phi_0$. Let $\phi^\star(x) = \sup_{t\geq 0 : \phi(t) < +\infty} (tx - \phi(t))$ denote the convex conjugate of $\phi \in \Phi_0$ and let $\psi(x) = \phi^\star(x) - x$. Define $\mathcal E = \{ f : \mathcal U \to \mathbb R \mbox{ for which } \mathbb{E}^{F_*}[\psi(c|f(U)|)] < \infty \mbox{ for all } c > 0\}$. The class $\mc E$ is an Orlicz class of functions (see Appendix (ref) for details). For example, \begin{align*} \mc E & = \{f : \mc U \to \mb R : \mb E^{F_*}[e^{c|f(U)|}] < \infty \mbox{ for all } c > 0 \} & & \mbox{for KL divergence,} \\ \mc E & = \{f : \mc U \to \mb R : \mb E^{F_*}[f(U)^2] < \infty\} & & \mbox{for $\chi^2$ divergence, and} \\ \mc E & = \{f : \mc U \to \mb R : \mb E^{F_*}[|f(U)|^q] < \infty\} & & \mbox{for $L^p$ divergence ($p^{-1} +q^{-1} = 1$).} \end{align*} Let $g = (g_1,g_2,g_3,g_4)$ denote the vector formed by stacking each of the moment functions from ((ref))--((ref)). Our key regularity condition is the following: \begin{assumptionp}{\textPhi} (i)\hskip\labelsep $\phi \in \Phi_0$. \begin{enumerate}[topsep=-18pt,itemsep=0pt,parsep=0pt,partopsep=0pt] • $k(\,\cdot\,,\theta,\gamma)$ and each entry of $g(\,\cdot\,,\theta,\gamma)$ belong to $\mc E$ for each $\theta \in \Theta$ and $\gamma \in \Gamma$. \end{enumerate} \end{assumptionp} For KL divergence, the class $\mc E$ contains of bounded functions (e.g., indicator functions) and functions that are additively separable in $U$ provided $F_*$ has tails that decay faster than exponentially (e.g., Gaussian but not Gumbel). Assumption (ref) therefore fails for KL divergence in Examples (ref) and (ref), but holds for $\chi^2$ or $L^p$ divergence as these only require finite second or $q$th moments, respectively. Let $d = \sum_{i=1}^4 d_i$ where $d_i$ is the dimension of $g_i$, let $\Lambda = \mb R^{d_1}_+ \times \mb R^{d_2} \times \mb R^{d_3}_+ \times \mb R^{d_4}$, and let $\lambda_{12}$ denote the first $d_1 + d_2$ elements of $\lambda$. A derivation of the following criterion functions is presented in Appendix (ref). \begin{proposition} Suppose that Assumption (ref) holds. Then the criterion functions ((ref)) and ((ref)) may be restated as \begin{align} \ul K_\delta(\theta;\gamma,P) & = \sup_{\eta > 0, \zeta \in \mb R, \lambda \in \Lambda} -\eta \mathbb{E}^{F_*}\left[ {\textstyle \phi^\star \left(\frac{k(U,\theta,\gamma) + \zeta + \lambda' g(U,\theta,\gamma) }{-\eta} \right) } \right] - \eta \delta - \zeta - \lambda_{12}'P \,,\\ \ol K_\delta(\theta;\gamma,P) & = \inf_{\eta > 0, \zeta \in \mb R, \lambda \in \Lambda} \phantom{-}\eta \mathbb{E}^{F_*}\left[ {\textstyle \phi^\star \left( \frac{k(U,\theta,\gamma) - \zeta- \lambda' g(U,\theta,\gamma)}{\eta} \right) } \right] + \eta \delta + \zeta + \lambda_{12}'P \,. \end{align} Moreover, the value of ((ref)) is $+\infty$ (equivalently, the value of ((ref)) is $-\infty$) if and only if there is no distribution in $\mc N_\delta$ under which ((ref)) holds at $(\theta,\gamma,P)$. \end{proposition} \begin{remark} Problems ((ref)) and ((ref)) are convex in $(\eta, \zeta , \lambda)$. The parameter $\eta$ is the Lagrange multiplier for the constraint $D_\phi(F\|F_*) \leq \delta$. Similarly, $\lambda$ collects the Lagrange multipliers for the moment (in)equalities ((ref))--((ref)). These multipliers are non-negative if they correspond to inequality restrictions and unconstrained otherwise. Finally, $\zeta$ is the Lagrange multiplier for the constraint $\int \mathrm d F = 1$, which ensures that the optimization is over probability measures. \end{remark} Problems ((ref)) and ((ref)) simplify in some special cases. For KL neighborhoods, $\phi^\star(x) = e^x - 1$ and the multiplier $\zeta$ has a closed-form solution, leading to \begin{align*} \ul K_\delta (\theta;\gamma,P) & = \sup_{\eta > 0, \lambda \in \Lambda} -\eta \log \mathbb{E}^{F_*}\left[ e^{- (k(U,\theta,\gamma)+ \lambda' g(U,\theta,\gamma) )/\eta} \right] - \eta \delta - \lambda_{12}'P \,, \\ \ol K_\delta (\theta;\gamma,P) & = \inf_{\eta > 0,\lambda \in \Lambda} \eta \log \mathbb{E}^{F_*}\left[ e^{(k(U,\theta,\gamma)- \lambda' g(U,\theta,\gamma) )/\eta} \right] + \eta \delta + \lambda_{12}'P \,. \end{align*} Another special case is when $k(u,\theta,\gamma)$ does not depend on $u$. To analyze this case, consider \begin{equation} \Delta(\theta; \gamma, P) := \inf_{F} D_\phi(F\|F_*) \quad \mbox{subject to ((ref)) holding at $(\theta,\gamma,P)$} . \end{equation} The value $\Delta(\theta; \gamma, P)$ is the minimum $\phi$-divergence between $F_*$ and a distribution $F$ for which the moment conditions hold at $(\theta,\gamma,P)$. We show in Proposition (ref) that $\Delta(\theta; \gamma, P)$ has an equivalent dual formulation: \begin{equation} \Delta(\theta;\gamma,P) = \sup_{\zeta \in \mb R,\lambda \in \Lambda} - \mathbb{E}^{F_*}\Big[ \phi^\star(-\zeta - \lambda' g(U,\theta,\gamma)) \Big] - \zeta - \lambda_{12}'P\,. \end{equation} For KL divergence, $\zeta$ may be solved for in closed-form and problem ((ref)) simplifies to \[ \Delta(\theta;\gamma,P) = \sup_{\lambda \in \Lambda} \; - \log \mathbb{E}^{F_*}\Big[ e^{ - \lambda' g(U,\theta,\gamma)} \Big] - \lambda_{12}'P\,. \] When $k$ does not depend on $u$, by a change of variables\footnote{Substitute $\eta \zeta - k(\theta,\gamma)$ in place of $\zeta$ in ((ref)) and $\eta \zeta + k(\theta,\gamma)$ in place of $\zeta$ in ((ref)), then substitute $\eta \lambda$ in place of $\lambda$ in both ((ref)) and ((ref)).} we may then restate problems ((ref)) and ((ref)) as \begin{equation} \ul K_\delta (\theta; \gamma, P) = \left[ \begin{array}{l} k(\theta,\gamma) \\[2pt] + \infty \end{array} \right. , \quad \ol K_\delta (\theta; \gamma, P) = \left[ \begin{array}{ll} k(\theta, \gamma) & \mbox{if $\Delta(\theta; \gamma, P) \leq \delta$,} \\[2pt] - \infty & \mbox{if $\Delta(\theta; \gamma, P) > \delta$.} \end{array} \right. \end{equation} An important feature of our approach is that the optimization problems (\ref{e:dual:1}), (\ref{e:dual:2}), and (\ref{e:qual:dual}) are convex and their dimension does not increase with $\delta$. This feature is not shared by other seemingly natural approaches to flexibly model $F$, such as mixtures or other finite-dimensional sieves. As we show in Section~\ref{s:sharp}, our procedure may be used to approximate sharp nonparametric bounds on counterfactuals by taking $\delta$ to be large but finite. \subsection{Estimation}\label{s:estimators} We now propose simple estimators of the bounds $\ul \kappa_\delta$ and $\ol \kappa_\delta$ based on ``plugging in'' consistent estimators $(\hat P,\hat \gamma)$ of $(P_0,\gamma_0)$. Estimators $\hat{\ul \kappa}_\delta$ and $\hat{\ol \kappa}_\delta$ are computed by optimizing criterion functions with respect to $\theta$: \begin{equation*} \begin{aligned} \hat{\ul \kappa}_\delta & = \inf_{\theta \in \Theta} \hat{\ul K}_\delta(\theta) \,, \quad & \quad \hat{\ol \kappa}_\delta & = \sup_{\theta \in \Theta} \hat{\ol K}_\delta(\theta)\,, \end{aligned} \end{equation*} where \[ \hat{\ul K}_\delta(\theta) = \left[ \begin{array}{l} \ul K_\delta (\theta; \hat \gamma, \hat P) \\[2pt] + \infty \end{array} \right. , \quad \hat{\ol K}_\delta(\theta) = \left[ \begin{array}{ll} \ol K_\delta (\theta;\hat \gamma,\hat P) & \mbox{if $\Delta(\theta;\hat \gamma,\hat P) < \delta$,} \\[2pt] - \infty & \mbox{if $\Delta(\theta;\hat \gamma,\hat P) \geq \delta$,} \end{array} \right. \] and $\ul K_\delta (\theta; \hat \gamma, \hat P)$, $\ol K_\delta (\theta; \hat \gamma, \hat P)$, and $\Delta(\theta;\hat \gamma,\hat P)$ are the criterion functions ((ref)), ((ref)), and ((ref)) evaluated at $(\hat \gamma, \hat P)$. If $k(u,\theta,\gamma) = k(\theta,\gamma)$, then we simply have \[ \hat{\ul K}_\delta(\theta) = \left[ \begin{array}{l} k(\theta,\hat \gamma)\\ + \infty \end{array} \right., \quad \hat{\ol K}_\delta(\theta) = \left[ \begin{array}{ll} k(\theta,\hat \gamma) & \mbox{if $\Delta(\theta;\hat \gamma,\hat P) < \delta$,} \\ - \infty & \mbox{if $\Delta(\theta;\hat \gamma,\hat P) \geq \delta$.} \end{array} \right. \] In Section (ref) we establish consistency of $\hat{\ul \kappa}_\delta$ and $\hat{\ol \kappa}_\delta$ and derive their asymptotic distribution. \subsection{Nonparametric Bounds on Counterfactuals} We define the (nonparametric) identified set of counterfactuals as \[ \mc K = \left\{ \mb E^{F}[k(U,\theta,\gamma_0)] : \mbox{(\ref{e:mod}) holds for some $\theta \in \Theta$ and $F \in \mc F_\theta$} \right\} , \] where $\mc F_\theta = \{ F \in \mc F : \mb E^F[ g(U,\theta,\gamma_0)] \mbox{ is finite and } F \ll \mu \}$ denotes all distributions on $\mc U$ that are absolutely continuous with respect to a $\sigma$-finite dominating measure $\mu$ and for which the moments in ((ref)) are finite at $\theta$. We impose existence of a density with respect to $\mu$ as it is often a structural assumption used, e.g., to avoid ties in CCPs or to establish existence of equilibria. The main result of this section shows that $\ul \kappa_\delta$ and $\ol \kappa_\delta$ approach the sharp nonparametric bounds $\inf\mc K$ and $\sup \mc K$ as $\delta$ becomes large. We first introduce some additional regularity conditions. Say $k$ is “$\mu$-essentially bounded” if $|k(\cdot,\theta,\gamma_0)|$ has finite $\mu$-essential supremum\footnote{The $\mu$-essential supremum of a function $f$ is denoted $\mu\text{-}\mathrm{ess}\sup f$ and is the smallest value $c$ for which $\mu(\{u : f(u) > c\}) = 0$. The $\mu$-essential infimum, denoted $\mu\text{-}\mathrm{ess}\inf$, is defined analogously.} for each $\theta \in \Theta$. This holds trivially if $k$ is bounded (e.g., counterfactual CCPs in Examples (ref) and (ref) and change in average welfare in Example (ref)). Models with unbounded $k$ may be reparameterized (as a proof device) by setting $\tilde \theta = (\theta,\kappa)$, appending $k(U,\theta,\gamma_0) - \kappa$ as an element of $g_4$, and setting $k(U,\tilde \theta,\gamma_0) = \kappa$. We also require a constraint qualification condition. This is a sufficient condition for establishing equivalence of “nonparametric” primal and dual problems in Appendix (ref), which is an intermediate step in the proof of the following result. Let $0_{d_i}$ denote a $d_i \times 1$ vector of zeros, $\mc C = \mb R^{d_1}_+ \times \{0_{d_2}\} \times \mb R^{d_3}_+ \times \{0_{d_4}\}$, $\mc G(\theta,\gamma) = \{ \mathbb{E}^{F} [ g(U,\theta,\gamma) ] : F \in \mc N_\infty \}$ where $\mc N_\infty = \{F : D_\phi(F \|F_*) < \infty\}$, and $\vec P = (P,0_{d_3+d_4})$. For $A,B \subseteq \mb R^d$, we let $\mathrm{ri}(A)$ denote the relative interior of $A$ and $A + B = \{a + b : a \in A, b \in B\}$. \begin{definition} \emph{Condition S} holds at $(\theta,\gamma,P)$ if $\vec P \in \mathrm{ri}(\mc G(\theta,\gamma) + \mc C) $. \end{definition} Using relative interior instead of interior allows for moment functions that are collinear at some $\theta$ (i.e., some moments are redundant). To give some intuition, consider moment equality models. Condition S requires that ((ref)) holds at $(\theta, \gamma, P)$ under some $F \in \mc N_\delta$ that is “interior” to $\mc N_\infty$, in the sense that one can perturb the (non-redundant) moments in any direction by perturbing $F$. For moment inequality models, Condition S also requires that there is $F \in \mc N_\infty$ under which all moment inequalities hold strictly at $(\theta, \gamma, P)$. Let $\Theta_I = \{ \theta \in \Theta:$ ((ref)) holds for some $F \in \mc F_\theta\}$ denote the (nonparametric) identified set for $\theta$. Define the “nonparametric” objective function \begin{equation} \ul K_{np}(\theta;\gamma,P) = \inf_{F \in \mc F_\theta} \mb E^F[k(U,\theta,\gamma)] \quad \mbox{subject to ((ref)) holding at $(\theta,\gamma,P)$} \,, \end{equation} with the understanding that $\ul K_{np}(\theta;\gamma,P) = +\infty$ if the infimum runs over an empty set. Let $\ol K_{np}(\theta;\gamma,P)$ denote the analogous supremum. Evidently, \[ \inf \mc K = \inf_{\theta \in \Theta} \ul K_{np}(\theta;\gamma_0,P_0) \quad \mbox{and} \quad \sup \mc K = \sup_{\theta \in \Theta} \ol K_{np}(\theta;\gamma_0,P_0)\,. \] \begin{definition} $\Theta_I$ is \emph{$S$-regular} if for all $\epsilon > 0$ there exist $\ul \theta, \ol \theta \in \Theta_I$ such that Condition S holds at $(\ul \theta, \gamma_0, P_0)$ and $(\ol \theta, \gamma_0, P_0)$, $\ul K_{np}(\ul \theta; \gamma_0, P_0) < \inf \mc K + \epsilon$, and $\ol K_{np}(\ol \theta; \gamma_0, P_0) > \sup \mc K - \epsilon$. \end{definition} Intuitively, S-regularity requires that the values the counterfactual takes at “boundary” points of $\Theta_I$ (i.e., at which Condition S fails) are not materially more extreme than values it can take at points “inside” $\Theta_I$ (i.e., at which Condition S holds). This condition can be verified under more primitive continuity conditions on $k$ and $g$. A sufficient (but not necessary) condition for S-regularity is that Condition S holds at $(\theta,\gamma_0,P_0)$ for all $\theta \in \Theta_I$. \begin{theorem} Suppose that Assumption (ref) holds, $k$ is $\mu$-essentially bounded, $\Theta_I$ is $S$-regular, and $\mu$ and $F_*$ are mutually absolutely continuous. Then \begin{align*} \lim_{\delta \to \infty} \ul \kappa_\delta & = \inf \mc K \,, & \lim_{\delta \to \infty} \ol \kappa_\delta & = \sup \mc K \,. \end{align*} \end{theorem} Theorem (ref) shows that our procedure can be used to approximate the sharp nonparametric bounds $\inf \mc K$ and $\sup \mc K$ by setting $\delta$ to be large but finite. If $\mu$ is Lebesgue measure---which it often is in applications---then the mutual absolute continuity condition in Theorem (ref) is satisfied whenever $F_*$ has strictly positive density over $\mc U$. \begin{remark} Appendix (ref) presents the dual forms of $\ul K_{np}$ and $\ol K_{np}$. Unlike $\ul K_\delta$ and $\ol K_\delta$, the duals of $\ul K_{np}$ and $\ol K_{np}$ are min-max and max-min problems which involve an inner optimization over $u$. These problems may be computationally challenging, especially when $u$ is multivariate. Comparing Proposition (ref) with the duals in Appendix (ref), we see that setting $\delta < \infty$ replaces a “hard-max” (an optimization over $u$) with a “soft-max” (a convex expectation). In this respect, adding the constraint $F \in \mc N_\delta$ may be viewed as a regularization of the nonparametric objective functions, similar to the use of entropic penalization to regularize objective functions in optimal transport problems---see, e.g., Cuturi2013. Smaller values of $\delta$ impose a stronger regularization. \end{remark} Theorem (ref) is silent on the issue of how large $\delta$ needs to be so that $\ul \kappa_\delta$ and $\ol \kappa_\delta$ are close to the nonparametric bounds. While this is model- and counterfactual-specific, the following toy example suggests that relatively small values of $\delta$ may suffice in some problems where the counterfactual is a choice probability. \begin{example} \normalfont Consider the problem \[ \ol \kappa_\delta = \sup_{\theta \in \mb R, F \in \mc N_\delta} \mathbb{E}^F[ 1\!\mathrm{l}\{U \leq \theta \}] \quad \mbox{subject to} \quad \mathbb{E}^F[U - \theta] = 0, \] where $\mc N_\delta$ is defined by KL divergence and $F_*$ is the $N(0,1)$ distribution. When $F = F_*$, the only solution to $\mathbb{E}^F[U - \theta] = 0$ is $\theta = 0$. Therefore, the value of the counterfactual under $F_*$ is $\mathbb{E}^{F_*}[ 1\!\mathrm{l}\{U \leq 0 \}] = \frac{1}{2}$ whereas $\sup \mc K = 1$. In Appendix (ref), we derive the large-$\delta$ approximation $\ol \kappa_\delta = 1 - 2 \pi e^{-2\delta - 1}(1 + o(1))$. By symmetry, $\ul \kappa_\delta = 2 \pi e^{-2\delta - 1}(1 + o(1))$ and $\inf \mc K = 0$. Therefore, in this example, $\ul \kappa_\delta$ and $\ol \kappa_\delta$ converge rapidly to $\inf \mc K$ and $\sup \mc K$ as $\delta$ increases. $\square$ \end{example} More generally, suppose the dual problems ((ref)) and ((ref)) have unique solutions $\ul \eta$ and $\ol \eta$ for $\eta$, where the optimization is performed over $\eta \geq 0$.\footnote{Optimizing over $\eta \geq 0$ rather than $\eta > 0$ does not affect the optimal value---see Proposition (ref).} Under appropriate regularity conditions (see, e.g., MilgromSegal), it follows that \[ \frac{\partial \ul K_\delta(\theta;\gamma,P)}{\partial \delta} = -\ul \eta, \quad \quad \frac{\partial \ol K_\delta(\theta;\gamma,P)}{\partial \delta} = \ol \eta. \] One can therefore infer from $\ul \eta$ and $\ol \eta$ the extent to which, if at all, the bounds at any fixed $\theta$ would widen further if $\delta$ was increased. \section{Practical Considerations} We now discuss practical details for implementing our procedure. Section (ref) discusses computational methods, Section (ref) presents our MPEC approach, and Section (ref) discusses methods for dealing with over-identified models. \subsection{Computation} There are three aspects to computation: (i) computing the expectations with respect to $F_*$ in the objective functions, (ii) solving the inner optimization problems over Lagrange multipliers, and (iii) solving the outer optimization problems over $\theta$. The expectations in the objective functions ((ref)), ((ref)), and ((ref)) are available in closed form for certain settings,\footnote{An earlier draft derived closed-form expressions for a discrete game of complete information with Gaussian payoff shocks and KL neighborhoods---see \url{https://arxiv.org/abs/1904.00989v2}.} in which case the dimension of $u$ does not play a role in the computational complexity of our procedure. Otherwise, the expectations will need to be computed numerically. If so, the dimension of $u$ will play a role in terms of determining how many quadrature points or Monte Carlo draws are needed to control numerical approximation error. In the empirical applications we used a randomized quasi-Monte Carlo approach based on scrambled Halton sequences as in Owen2017. The inner optimization with respect to Lagrange multipliers can be solved rapidly: it is convex and gradients and Hessians are available in closed-form. The envelope theorem can be used to derive gradients for the outer optimization when $k$ and $g$ are differentiable in $\theta$.\footnote{In practice, we smoothed any non-smooth moments and used automatic differentiation to compute derivatives with respect to $\theta$ if these were not easily available analytically.} Our procedures were all implemented in Julia with the inner and outer optimizations solved using Knitro. A general-purpose implementation of our methods in Julia is provided in the supplemental material. As with parameter estimation in nonlinear structural models, the outer optimization with respect to $\theta$ is typically non-convex. In applications, we used an iterative multi-start procedure in an attempt to converge to global optima. Computation times are reported in the applications below. \subsection{MPEC Approach} We now describe and formally justify an MPEC version of our procedure in the spirit of SuJudd. This approach simplifies computation in models with endogenous parameters defined by equilibrium conditions (e.g., value functions defined by Bellman equations), resulting in significant computational gains for DDC models in particular. Suppose $\theta = (\theta_s, \theta_e)$ and $g_4 = (g_{4s}, g_{4e})$ where $\theta_s$ are “deep” structural parameters and $\theta_e$ are “endogenous” parameters that are defined implicitly by $g_{4e}$. That is, for any $(\theta_s,\gamma,F)$, the parameter $\theta_e = \theta_e(\theta_s,\gamma,F)$ solves \[ \mb E^F[ g_{4e}( U,( \theta_s, \theta_e), \gamma ) ] = 0 \,. \] For instance, in Example (ref) we have $\theta_s = \theta_\pi$ or $(\theta_\pi,\beta)$, while $\theta_e = (v,\tilde v)$ collects the value functions in the baseline model and counterfactual, and $g_{4e}$ collects the functions representing the corresponding Bellman equations, as in display ((ref)). Although our procedure can be implemented as described in Section (ref), that implementation does not make use of the fact that $\theta_e$ is defined implicitly by $g_{4e}$. To leverage this structure, consider the subset of moments conditions excluding $g_{4e}$: \begin{equation} \begin{aligned} \mb E^F[ g_1(U,\theta,\gamma_0)] & \leq P_{10} , & \mb E^F[ g_2(U,\theta,\gamma_0)] & = P_{20} , \\ \mb E^F[ g_3(U,\theta,\gamma_0)] & \leq 0 , & \mb E^F[ g_{4s}(U,\theta,\gamma_0)] & = 0 , & \end{aligned} \end{equation} and define criterion functions using these only: \begin{align} \ul K^s_{\delta}(\theta;\gamma,P) & = \inf_{F \in \mc N_\delta} \mb E^F[k(U,\theta,\gamma)] \quad \mbox{ subject to ((ref)) holding at $(\theta,\gamma,P)$} \,, \\ \ol K^s_{\delta}(\theta;\gamma,P) & = \sup_{F \in \mc N_\delta} \mb E^F[k(U,\theta,\gamma)] \quad \mbox{ subject to ((ref)) holding at $(\theta,\gamma,P)$} \,. \end{align} Under the conditions of Proposition (ref), these criterion functions may be restated as \begin{align} \ul K_{\delta}^s(\theta;\gamma,P) & = \sup_{\eta > 0, \zeta \in \mb R, \lambda \in \Lambda_s} -\eta \mathbb{E}^{F_*}\left[ {\textstyle \phi^\star \left( \frac{k(U,\theta,\gamma) + \zeta + \lambda' g_s(U,\theta,\gamma) }{-\eta}\right) } \right] - \eta \delta - \zeta - \lambda_{12}'P \,, \\ \ol K_{\delta}^s(\theta;\gamma,P) & = \inf_{\eta > 0, \zeta \in \mb R, \lambda \in \Lambda_s} \phantom{-} \eta \mathbb{E}^{F_*}\left[ {\textstyle \phi^\star \left( \frac{k(U,\theta,\gamma) - \zeta- \lambda' g_s(U,\theta,\gamma) }{\eta}\right) } \right] + \eta \delta + \zeta + \lambda_{12}'P \,, \end{align} with $g_s = (g_1,g_2,g_3,g_{4s})$ and $\Lambda_s = \mb R^{d_1}_+ \times \mb R^{d_2} \times \mb R^{d_3}_+ \times \mb R^{d_{4s}}$ with $d_{4s} = \dim(g_{4s})$. Problems ((ref)) and ((ref)) simplify analogously to ((ref)) when $k$ does not depend on $u$, with the minimum divergence problem $\Delta$ defined using $g_s$ in place of $g$. In our MPEC approach, the criterion functions ((ref)) and ((ref)) are optimized with respect to $\theta$, with the remaining moment conditions involving $g_{4e}$ appended as constraints. Importantly, these constraints are evaluated under the “least favorable” distributions $\ul F_{\delta,\theta}$ and $\ol F_{\delta,\theta}$ that solve problems ((ref)) and ((ref)), respectively. The following proposition formally justifies this approach. \begin{proposition} Suppose that Assumption (ref) holds. Then the problems \[ \inf_{\theta \in \Theta} \ul K_\delta(\theta;\gamma,P) \] and \[ \inf_{\theta \in \Theta} \ul K_{\delta}^s(\theta;\gamma,P) \;\mbox{ subject to } \; \mathbb{E}^{\ul F_{\delta,\theta}}[g_{4e}(U,\theta,\gamma)] = 0 \] have the same value. An analogous result holds for the upper bound. \end{proposition} To implement our MPEC approach, note that the expectations in the constraints may be expressed in terms of changes of measure. Let $\ul m_{\delta,\theta} = {\mr d \ul F_{\delta,\theta}}/{\mr d F_*}$ and $\ol m_{\delta,\theta} = {\mr d \ol F_{\delta,\theta}}/{\mr d F_*}$ so that \[ \begin{aligned} \mathbb{E}^{\ul F_{\delta,\theta}}[\, \cdot \,] & = \mathbb{E}^{F_*}[ \ul m_{\delta,\theta}(U) \, \, \cdot \,] \,, & \mathbb{E}^{\ol F_{\delta,\theta}}[\, \cdot \,] & = \mathbb{E}^{F_*}[ \ol m_{\delta,\theta}(U) \, \, \cdot \,] \,. \end{aligned} \] If $k$ depends on $u$, then we construct $\ul m_{\delta,\theta}$ and $\ol m_{\delta,\theta}$ from solutions to ((ref)) and ((ref)), say $(\ul \eta, \ul \zeta, \ul \lambda)$ and $(\ol \eta, \ol \zeta, \ol \lambda)$ (these solutions exist under the regularity conditions below). If $\ul \eta > 0$, then the distribution solving ((ref)) is unique and is induced by the change of measure \begin{equation} \ul m_{\delta,\theta}(u) = \dot \phi^\star\left( \frac{k(u,\theta,\gamma) + \ul \zeta + \ul \lambda' g_s(u,\theta,\gamma)}{ -\ul \eta} \right) , \end{equation} where $\dot \phi^\star(x) = \frac{d \phi^\star(x)}{dx}$. The function $\ol m_{\delta,\theta}(u)$ is constructed similarly, replacing $(\ul \eta, \ul \zeta, \ul \lambda)$ in ((ref)) by $(-\ol \eta, - \ol \zeta, - \ol \lambda)$. For KL divergence the change of measure simplifies to \[ \ul m_{\delta,\theta}(u) = \frac{e^{(k(u,\theta,\gamma) + \ul \lambda' g_s(u,\theta,\gamma))/ -\ul \eta}}{\mb E^{F_*}\left[e^{(k(u,\theta,\gamma) + \ul \lambda' g_s(u,\theta,\gamma))/ -\ul \eta}\right]} \,, \] and similarly for $\ul m_{\delta,\theta}(u)$. If $\ul \eta = 0$, then there may be multiple minimizing distributions. As shown in the proof of Proposition (ref), each such distribution must be supported on \[ \ul A_{\delta,\theta} := \{ u : k(u,\theta,\gamma) + \ul \lambda' g_s(u,\theta,\gamma) = F_*\text{-}\mathrm{ess} \inf ( k(\cdot,\theta,\gamma) + \ul \lambda' g_s(\cdot,\theta,\gamma))\}\,. \] Note $F_*(\ul A_{\delta,\theta}) > 0$ is required for $\ul \eta = 0$ to be a solution. Otherwise, any distribution supported on $\ul A_{\delta,\theta}$ is not absolutely continuous with respect to $F_*$ and is therefore not in $\mc N_\delta$. If $\ul \eta = 0$ and $F_*(\ul A_{\delta,\theta}) > 0$, then we construct $\ul m_{\delta,\theta}$ by restricting $F_*$ to $\ul A_{\delta , \theta}$ and rescaling: \[ \ul m_{\delta,\theta}(u) = 1\!\mathrm{l}\{u \in \ul A_{\delta,\theta}\}/F_*(\ul A_{\delta, \theta}). \] The function $\ol m_{\delta, \theta}(u)$ is constructed analogously, replacing $\ul \lambda$ with $-\ol \lambda$ and the set $\ul A_{\delta,\theta}$ with $\ol A_{\delta, \theta} = \{ u : k(u,\theta,\gamma) - \ol \lambda' g_s(u,\theta,\gamma) = F_*\text{-}\mathrm{ess} \sup (k(\cdot,\theta,\gamma) - \ol \lambda' g_s(\cdot,\theta,\gamma))\}$. If $k$ does not depend on $u$, then $\ul m_{\delta,\theta}$ and $\ol m_{\delta,\theta}$ are constructed from solutions to a version of problem ((ref)) with $g_s$ in place of $g$. Under the regularity conditions below, this program has a solution, say $(\ul \zeta, \ul \lambda)$. In this case, we define \begin{equation} \ul m_{\delta,\theta}(u) = \ol m_{\delta,\theta}(u) = \dot \phi^\star\left( - \ul \zeta - \ul \lambda' g_s(u,\theta,\gamma) \right) \,. \end{equation} For KL divergence the change of measure simplifies to \[ \ul m_{\delta,\theta}(u) = \ol m_{\delta,\theta}(u) = \frac{e^{- \ul \lambda' g_s(u,\theta,\gamma)}}{\mb E^{F_*}\left[e^{- \ul \lambda' g_s(u,\theta,\gamma)}\right]} \,. \] \begin{proposition} Suppose that Assumption (ref) holds, Condition S holds at $(\theta,\gamma,P)$, and there exists a distribution $F$ with $D(F\|F_*) < \delta$ under which ((ref)) holds at $(\theta,\gamma,P)$. Then the distributions $\ul F_{\delta,\theta}$ and $\ol F_{\delta,\theta}$ induced by $\ul m_{\delta,\theta}$ and $\ol m_{\delta,\theta}$ solve ((ref)) and ((ref)), respectively. \end{proposition} \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Example.} We consider a numerical example for the DDC model of Rust based on the parameterization in Section 5.4 of NoretsTang. The counterfactual they consider is a hypothetical change in the law of motion of the state. We follow these papers and use state-space of dimension 90. As $|\mc S| = 90$ and $\mc D_0 = \{0,1\}$, there are 90 functions in $g_2$ representing the observed CCPs. There are another 180 functions in $g_{4e}$ representing the Bellman equations in the baseline model and counterfactual across states. We also impose the normalization $\mathbb{E}^F[U_d] = 0$ for $d = 0,1$. Hence, $g_{4s}(U,\theta,\gamma) = (U_0,U_1)$. Our MPEC approach has 92 moments in the inner optimization (90 for CCPs and two mean-zero normalizations on the shocks) with the remaining 180 moments representing the Bellman equations appended as constraints. The full approach uses all 272 moments in the inner optimization. Table (ref) reports computation times for the inner optimization problems ((ref)) and ((ref)) (denoted $\ol K_\delta$) for maximizing the counterfactual CCP in the highest mileage state.\footnote{The times in Table (ref) are based on initializing the solver at $\eta = 1$, $\zeta = 0$, and $\lambda = 0$. When embedded in the outer optimization over $\theta$, computation times for the inner problem are reduced significantly by using a warm start that initializes at the solution to the inner problem at the previous value of $\theta$.} We also report times for solving the minimum divergence problem ((ref)) (denoted $\Delta$) using the full set of moment functions $g$ and its MPEC analogue using $g_s$. Neighborhoods are constrained by a hybrid of KL and $\chi^2$ divergence as in the empirical applications---see Section (ref). As can be seen, the inner optimization problems are solved at least 20 times faster for the MPEC implementation, with the relative efficiency increasing in $\delta$. \begin{table} \begin{center} \caption{Computation times (in seconds) for the inner problems} \begin{tabular}{lcccc} \hline \hline \multicolumn{1}{c}{Implementation} & \multicolumn{4}{c}{Objective} \\\cline{2-5} \\[-10pt] & $\ol K_{0.01}$ & $\ol K_{0.10}$ & $\ol K_{1.00}$ & $\Delta$ \\[2pt] MPEC (92 moments) & 0.207 & \phantom{1}0.232 & \phantom{1}0.256 & 0.108 \\[2pt] Full (272 moments) & 4.317 & 12.978 & 43.699 & 3.365 \\[2pt] \hline \\[-10pt] \end{tabular} \parbox{\textwidth}{\small \emph{Note:} Expectations are computed using 50,000 scrambled Halton draws. Computations are performed in Julia v1.6.4 and Knitro v12.4.0 on a 2.7GHz MacBook Pro with 16GB memory.} \end{center} \end{table} \subsection{Over-identification}\label{s:overid} In over-identified models (i.e., where the number of moment conditions $d$ exceeds the dimension $d_\theta$ of $\theta$), there might not exist $\theta \in \Theta$ for which the sample moment conditions (\ref{e:sample}) hold under $F_*$. We propose two methods for handling over-identified models. First, one may compute the smallest value of $\delta$ for which there exists $F \in \mc N_\delta$ consistent with the sample moment conditions (\ref{e:sample}) by solving the optimization problem \[ \hat \delta = \inf_{\theta \in \Theta} \Delta(\theta;\hat \gamma,\hat P). \] The interval $[\hat{\ul \kappa}_\delta,\hat{\ol \kappa}_\delta]$ will be nonempty for $\delta > \hat \delta$. If the model is correctly specified under $F_*$,\footnote{Neither our theoretical results developed in Section (ref) nor the estimation and inference results in Section (ref) require correct specification of the model under $F_*$.} then $\hat \delta$ will converge in probability to zero under the conditions of Theorem (ref). In this case, the interval $[\hat{\ul \kappa}_\delta,\hat{\ol \kappa}_\delta]$ will be nonempty with probability approaching one for each fixed $\delta > 0$. It is also possible that $\hat \delta = +\infty$ in correctly specified but over-identified models when $\hat P$ is incompatible with certain model restrictions. For instance, CCPs are often estimated nonparametrically using empirical choice frequencies. If some choices aren't observed in the data, then the estimated CCPs will be zero even though model-implied CCPs are strictly positive. This issue can be circumvented in models defined by equality restrictions only (hence $P_0 \equiv P_{20}$) using the following two-step approach. First, compute a preliminary estimator $\tilde \theta$ of $\theta$ based on ((ref)). Then, set $\hat P = \mb E^{F_*}[g_2(U,\tilde \theta, \hat \gamma)]$. This second-step estimator $\hat P$ is compatible with the model by construction, thereby ensuring that the interval $[\hat{\ul \kappa}_\delta, \hat{\ol \kappa}_\delta]$ is nonempty for each $\delta > 0$. The estimator $\hat P$ will be consistent and asymptotically normal under mild regularity conditions provided the model is correctly specified under $F_*$, so the consistency and inference results developed in Section (ref) will also apply. \section{Interpreting the Neighborhood Size} This section presents some theoretical results and practical methods to help interpret the neighborhood size $\delta$. Sections (ref) and (ref) discuss properties of $\phi$-divergences and their implications for interpreting $\delta$. Section (ref) shows how to construct the “least favorable” distributions that minimize or maximize the counterfactual. Section (ref) gives a practical, model-based metric for interpreting $\delta$. \subsection{Invariance} A defining property of $\phi$-divergences are their invariance to invertible transformations. That is, if $T$ is an invertible transformation and $G$ and $G_*$ denote the distributions of $T(U)$ when $U \sim F$ and $U \sim F_*$, respectively, then $D_{\phi}(F\|F_*) = D_{\phi}(G\|G_*)$.\footnote{See, e.g., LieseVajda. A more direct statement is in QiaoMinematsu.} An important consequence of invariance is that $\delta$ has the same interpretation under a change in units. For instance, if one researcher writes a model in terms of dollars with $U \sim F_*$ and another researcher uses thousands of dollars with $U \sim G_*$ for $G_*(u) = F_*(10^{-3} u)$, then $F$ is in $\mc N_\delta$ if and only if its rescaled counterpart $G$ is in a $\delta$-neighborhood of $G_*$. A second consequence is that neighborhood size is invariant under invertible location and scale transformations of $F_*$ (e.g., $N(\mu,\Sigma)$ versus $N(0,I)$). \subsection{Least Favorable Distributions} A useful feature of our approach is that the “least favorable” distributions (LFDs) that attain the smallest or largest values of the counterfactual may easily be recovered. To help interpret $\delta$, one may plot the LFDs and compute other quantities of interest (e.g., correlations or welfare measures) under them. Section (ref) describes how to construct LFDs when our MPEC approach is used. LFDs for our full (i.e., non-MPEC) approach are a special case with $g_4 = g_{4s}$. To briefly summarize, consider the LFD $\ul F_{\delta,\theta}$ solving the minimization problem ((ref)). First suppose that $k$ depends on $u$. Let $(\ul \eta, \ul \zeta, \ul \lambda)$ solve problem ((ref)). If $\ul \eta > 0$, then $\ul F_{\delta,\theta}$ is unique and its change-of-measure $\ul m_{\delta,\theta} = \mr d\ul F_{\delta,\theta}/\mr d F_*$ is given by \begin{equation} \ul m_{\delta,\theta}(u) = \dot \phi^\star\left( \frac{k(u,\theta,\gamma) + \ul \zeta + \ul \lambda' g(u,\theta,\gamma)}{ -\ul \eta} \right) . \end{equation} The LFD $\ol F_{\delta, \theta}$ solving the maximization problem ((ref)) is constructed similarly, replacing $(\ul \eta, \ul \zeta, \ul \lambda)$ in ((ref)) with $(-\ol \eta, - \ol \zeta, - \ol \lambda)$, where $(\ol \eta, \ol \zeta, \ol \lambda)$ solves ((ref)). If $\ul \eta = 0$ or $\ol \eta = 0$, then there may exist multiple distributions solving ((ref)) and ((ref)) at $\theta$. LFDs in this case are constructed analogously to the method described in Section (ref). Note that $\ul \eta = 0$ or $\ol \eta = 0$ is unlikely if $k$ and/or elements of $g$ are unbounded in $u$---see the discussion in Section (ref). If $k$ does not depend on $u$, then we set \begin{equation} \ul m_{\delta,\theta}(u) = \ol m_{\delta,\theta}(u) = \dot \phi^\star \left( - \ul \zeta - \ul \lambda' g(u,\theta,\gamma) \right) \end{equation} where $(\ul \zeta, \ul \lambda)$ solves ((ref)). While there may exist multiple distributions solving ((ref)) and ((ref)) in this case, the distribution induced by ((ref)) has smallest $\phi$-divergence relative to $F_*$ among all such distributions. \subsection{Viewing Neighborhood Size through the Lens of the Model} Another method for interpreting $\delta$ is based on measuring the variation in the moments at the distributions solving ((ref)) and ((ref)) relative to their values under $F_*$. Consider the sets of minimizing and maximizing values of $\theta$ at which $\ul \kappa_\delta$ and $\ol \kappa_\delta$ are attained, say $\ul \Theta_\delta$ and $\ol \Theta_\delta$. These are nonempty under the regularity conditions in Section (ref). While the moment conditions ((ref)) hold at any $\theta \in \ul \Theta_\delta \cup \ol \Theta_\delta$ under the corresponding LFD, they will typically not hold at $\theta$ under $F_*$. We therefore define \begin{multline*} size(\delta) = \sup_{\theta \in \ul \Theta_\delta \cup \ol \Theta_\delta} \max \Big\{ \big\| \left(\mb E^{F_*}[g_1(U,\theta,\gamma_0)] - P_{10}\right)_+ \big\|_\infty , \big\| \mb E^{F_*}[g_1(U,\theta,\gamma_0)] - P_{20} \big\|_\infty , \\ \big\| \left(\mb E^{F_*}[g_3(U,\theta,\gamma_0)] \right)_+ \big\|_\infty , \big\| \mb E^{F_*}[g_4(U,\theta,\gamma_0)] \big\|_\infty \Big\} \,, \end{multline*} where $(v)_+ = ( \max\{v_i,0\})_{i=1}^d$ for a vector $v \in \mb R^d$. The quantity $size(\delta)$ is the maximum degree to which the moments at $\theta \in \ul \Theta_\delta \cup \ol \Theta_\delta$ violate ((ref)) under $F_*$. This measure is informative about the extent to which the distortions to $F_*$ required to attain the smallest and largest values of the counterfactual over $\mc N_\delta$ are reflected in ((ref)). Small values of $size(\delta)$ indicate that the LFDs supporting $\ul \kappa_\delta$ and $\ol \kappa_\delta$ distort $F_*$ in a way that moves the counterfactual but barely moves the moments. Conversely, large values of $size(\delta)$ indicate that distortions required to increase or decrease the counterfactual also have a material impact on the moments. In practice, this measure can be computed by replacing $(P_0,\gamma_0)$ by estimators $(\hat P,\hat \gamma)$ and $\ul \Theta_\delta$ and $\ol \Theta_\delta$ by the minimizers and maximizers of the sample criterions or by the estimators of $\ul \Theta_\delta$ and $\ol \Theta_\delta$ introduced in Section (ref). \subsection{Relating Different Divergences} It is well known that $\phi$-divergences are equivalent over local neighborhoods (see, e.g., Theorem 4.1 of CsiszarShields). However, $\ul \kappa_\delta$ and $\ol \kappa_\delta$ may depend on the choice of $\phi$ when $\delta$ is not arbitrarily small. Bounds induced by different $\phi$ functions may be related as follows. Let $\mc N_{\delta,1}$ and $\mc N_{\delta,2}$ denote $\delta$-neighborhoods induced by $\phi_1$ and $\phi_2$, respectively. The quantity \[ \bar a = \sup_{x \geq 0, x \neq 1} \frac{\phi_1(x)}{\phi_2(x)} \] is a measure of relative neighborhood size: if $\bar a < \infty$ then $\mc N_{\delta,2} \subseteq \mc N_{\bar a \delta,1}$ for each $\delta > 0$, as shown formally in the proof of Proposition (ref) below. For instance, when comparing KL divergence ($\phi_1(x) = x \log x - x + 1$) and $\chi^2$ divergence ($\phi_2(x) = \frac{1}{2}(x-1)^2$) we obtain $\bar a = 2$. Therefore, $\delta$-neighborhoods under $\chi^2$ divergence are contained in $2\delta$-neighborhoods under KL divergence. Interchanging $\phi_1$ and $\phi_2$ produces $\bar a = +\infty$, which reflects the fact that KL divergence is weaker than $\chi^2$ divergence. Let $\ul \kappa_{\delta,1}$ and $\ul \kappa_{\delta,2}$ denote the smallest counterfactual from display ((ref)) over $\mc N_{\delta,1}$ and $\mc N_{\delta,2}$, respectively. Define $\ol \kappa_{\delta,1}$ and $\ol \kappa_{\delta,2}$ analogously. \begin{proposition} Suppose that Assumption (ref) holds for both $\phi_1$ and $\phi_2$ and $\bar a$ is finite. Then $[\ul \kappa_{\delta,2}, \ol \kappa_{\delta,2}] \subseteq [\ul \kappa_{\bar a \delta, 1}, \ol \kappa_{\bar a \delta, 1}] $ for each $\delta > 0$. \end{proposition} It follows from Proposition (ref) that bounds that are wide under $\phi_2$ must necessarily be wide under $\phi_1$. Similarly, narrow bounds under $\phi_1$ must also be narrow under $\phi_2$. Note also that the inclusion in Proposition (ref) holds for any counterfactual. \section{Empirical Applications} \subsection{Marital College Premium} CSW, henceforth CSW, study the evolution of marital returns to education using a frictionless matching model with transferable utility ChooSiow. Within this framework, the “marital college premium” is the additional expected utility that an individual would derive from the marriage market if they had a (counterfactually) higher level of education. CSW find that marital college premiums for women in the United States increased significantly across cohorts from the mid to late 20th century, particularly for the more highly educated. As is conventional following Dagsvik and ChooSiow, CSW assume latent variables representing individuals' idiosyncratic marital preferences are i.i.d. Gumbel. The marital college premium is only partially identified when the distribution of these latent variables is not specified. We therefore perform a sensitivity analysis of CSW's estimates to departures from this conventional parametric assumption. Our analysis makes several findings. First, it seems impossible to draw conclusions about whether marital college premiums have increased or decreased over time under small nonparametric relaxations of the i.i.d. Gumbel assumption. Interestingly, premiums have narrow nonparametric bounds at {fixed} parameter values, but a slight relaxation of the i.i.d. Gumbel assumption allows for significant variation in parameters which, in turn, produces uninformatively wide bounds. As parameters are just-identified under any fixed distribution of shocks GalichonSalanie, further restrictions on parameters or shape restrictions on the distribution are required to tighten the bounds. We show that imposing exchangeability can tighten the bounds significantly. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Model and Benchmark Estimates.} Agents are male or female and one of $J$ types (education levels). A type-$a$ male receives utility $\varepsilon_{a0}$ if he chooses to be unmatched and $z_{ab} + \varepsilon_{ab}$ if he matches with a type-$b$ female. Similarly, a type-$b$ female receives utility $e_{0b}$ if she chooses to be unmatched and $t_{ab} + e_{ab}$ if she matches with a type-$a$ male. The parameters $(z_{ab},t_{ab})_{a,b=1}^J$ represent the common deterministic component of marital preferences. The latent shocks $(\varepsilon_{a0},\ldots,\varepsilon_{aJ})$ and $(e_{0b},\ldots,e_{Jb})$ represent individuals' idiosyncratic marital preferences. Shocks are i.i.d. across individuals and have mean zero. The type $b$ to $b'$ marital education premium for females is the difference in expected marital utility between types $b$ and $b'$: \begin{equation} \kappa = \mathbb E^F \left[ \max_{a = 0,\ldots,J} \Big( t_{ab'} + e_{ab'} \Big) \right] - \mathbb E^F \left[ \max_{a = 0,\ldots,J} \Big( t_{ab} + e_{ab} \Big) \right] , \end{equation} where $F$ denotes the distribution of $(e_{0b},\ldots,e_{Jb'})$ and $t_{0b} = t_{0b'} = 0$. CSW use data from the American Community Survey. They form 28 cohorts indexed by female birth year from 1941 (cohort 1) to 1968 (cohort 28), each of which is treated as an independent marriage market. We focus on CSW's estimates for whites. There are $J = 5$ types: “high-school dropouts”, “high-school graduates”, “some college”, “college graduate”, and “college-plus”. We center our analysis on the “some college” to “college graduate” premium, though we obtained qualitatively similar results (not reported) for the “college graduate” to “college-plus” premium. Figure (ref) presents estimates and 95% confidence sets (CSs) for the premium under the i.i.d. Gumbel assumption (cf. Figure 21 in CSW) based on CSW's replication files. \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Implementation.} The model reduces to a standard individual-level discrete choice problem for each type (see CSW's Propositions 1 and 2). We assume that the distribution of females' preference shocks does not depend on their type, so we drop the $b$ subscript and consider a single random vector $U = (e_0,\ldots,e_J)$. We allow the distribution $F$ of $U$ to vary across cohorts and implement our procedures cohort-by-cohort.\footnote{In view of the just-identification results of GalichonSalanie, we would obtain the same bounds if $F$ was homogeneous across cohorts. Allowing for heterogeneity in own-type would result in wider bounds.} Under any fixed $F$, a cohort's parameters $(t_{ab})_{a=1}^J$ are just-identified from the marriage probabilities for that cohort's type-$b$ women GalichonSalanie. We therefore impose only the moment conditions involving the parameters $\theta = (t_{ab},t_{ab'})_{a = 1}^J$ appearing in ((ref)), as the remaining parameters can be chosen to fit the remaining marriage probabilities under the resulting least-favorable distribution. We form $g_2$ to explain the type $b$ and $b'$ marriage probabilities for women in a given cohort: \[ g_2(U,\theta) = \left[ \begin{array}{c} (1\!\mathrm{l} \{ t_{ab} + e_{a} = \max_{a' = 0,\ldots,J} ( t_{a'b} + e_{a'} ) \})_{a=1}^J \\ (1\!\mathrm{l} \{ t_{ab'} + e_{a} = \max_{a' = 0,\ldots,J} ( t_{a'b'} + e_{a'} ) \})_{a=1}^J \end{array} \right] \] and form $\hat P_2$ using CSW's estimates of the corresponding type-$b$ and $b'$ marriage probabilities. We set $g_4(U,\theta) = (e_j,e_j^2 - \pi^2/6)_{j=0}^{J}$ so that shocks have mean zero and the same variance as the Gumbel distribution. The scale normalization also ensures that the nonparametric bounds on the premium are finite at any fixed $\theta$. As $J = 5$, there are 22 moments (10 for marriage probabilities and 12 location/scale normalizations), and $\theta$ has dimension $10$. We consider a second implementation which imposes invariance of $F$ under rotations and reflections of potential spouse types, so that the model-implied marriage probabilities depend on $\theta$ but not the labeling of potential spouse types (though they may depend on their ordering).\footnote{Allowing dependence on the ordering of types seems desirable here as types correspond to education levels, which are naturally ordered.} Formally, this shape restriction corresponds to dihedral exchangeability (see Appendix (ref)); we refer to it simply as “exchangeability”. Under this shape restriction, $F$ must satisfy the 22 moment conditions under all 12 rotations and reflections of the elements of $U$. This implementation therefore imposes a total of $264$ moment conditions. Rather than including all 264 moments separately, it suffices to form $g_2$ and $g_4$ by taking the averages of the 22 moments across the 12 permutations (see Appendix (ref)). Both implementations therefore have inner optimization problems of the same dimension. Computations are performed as described in Section (ref). The first implementation uses 50,000 scrambled Halton draws to compute the expectations. The second uses 10,000 draws which are concatenated over the 12 permutations (see Remark (ref)), for a total of 120,000 draws. Computation times are reported in Appendix (ref). CSs for $\ul \kappa_\delta$ and $\ol \kappa_\delta$ are computed using the bootstrap procedure in Section (ref). Appendix (ref) discusses bootstrap details and presents projection CSs using the method from Section (ref). We define neighborhoods using a hybrid of KL and $\chi^2$ divergence: \[ \phi(x) = \left[ \begin{array}{ll} x \log x - x + 1 & \mbox{if $x \leq e$,} \\ {\textstyle \frac{1}{2e}(x-e)^2 + (x-e) + 1} & \mbox{if $x > e$.} \end{array} \right. \] We use this divergence because Assumption (ref)(ii) fails for KL divergence, whereas hybrid divergence only requires finite second moments for Assumption (ref)(ii). The LFDs under hybrid divergence are also everywhere positive, which is not guaranteed under $\chi^2$ divergence. We repeated our analysis with neighborhoods constrained by $\chi^2$ and $L^4$ divergences as robustness checks. Overall, our findings are not sensitive to $\phi$ (see Appendix (ref) for a discussion). \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Findings.} \begin{figure}[t] \begin{center} \begin{subfigure}[t]{\textwidth} \begin{center} \caption*{Smaller neighborhoods} \vskip -2pt \begin{subfigure}[t]{0.49\textwidth} \begin{center} \caption{Without exchangeability} \vskip -6pt \end{center} \end{subfigure} \begin{subfigure}[t]{0.49\textwidth} \begin{center} \caption{With exchangeability\phantom{out}} \vskip -6pt \end{center} \end{subfigure} \end{center} \end{subfigure} \vskip 10pt \begin{subfigure}[t]{\textwidth} \begin{center} \caption*{Larger neighborhoods} \vskip -2pt \begin{subfigure}[t]{0.49\textwidth} \begin{center} \caption{Without exchangeability} \vskip -6pt \end{center} \end{subfigure} \begin{subfigure}[t]{0.49\textwidth} \begin{center} \caption{With exchangeability\phantom{out}} \vskip -6pt \end{center} \end{subfigure} \end{center} \end{subfigure} \caption{Sensitivity analysis of the “some college” to “college graduate” premium across cohorts. \emph{Note:} Solid lines are estimates, dotted lines are (cohort-wise) 95% CSs. CSW's estimates and CSs correspond to $\delta = 0$.} \vskip -15pt \end{center} \end{figure} Figure (ref) presents a sensitivity analysis of the “some college” to “college graduate” premium. Cohort-wise estimates and CSs for $\ul \kappa_{\delta}$ and $\ol \kappa_{\delta}$ are presented, beginning at $\delta = 0.01$ and increasing $\delta$ by factors of 10 up to $\delta = 100$. Even with $\delta = 0.01$, estimates of $\ul \kappa_\delta$ and $\ol \kappa_\delta$ lie uniformly below and above zero across cohorts without exchangeability (see Figure (ref)). Imposing exchangeability can tighten the bounds, with the bounds for $\delta = 0.01$ significantly negative in early cohorts and significantly positive in later cohorts (see Figure (ref)). But the $\delta = 0.1$ bounds with exchangeability again contain zero across all cohorts. Bounds for larger $\delta$ presented in Figures (ref) and (ref) are uninformatively wide. \begin{figure} \begin{center} \begin{subfigure}[t]{\textwidth} \begin{center} \caption{Without exchangeability} \vskip -2pt \begin{subfigure}[t]{0.43\textwidth} \begin{center} \caption*{Type 0 (Unmatched)} \vskip -6pt \end{center} \end{subfigure} \begin{subfigure}[t]{0.43\textwidth} \begin{center} \caption*{Type 1 (High-school dropout)} \vskip -6pt \end{center} \end{subfigure} \begin{subfigure}[t]{0.43\textwidth} \begin{center} \caption*{Type 2 (High-school graduate)} \vskip -6pt \end{center} \end{subfigure} \begin{subfigure}[t]{0.43\textwidth} \begin{center} \caption*{Type 3 (Some college)} \vskip -6pt \end{center} \end{subfigure} \begin{subfigure}[t]{0.43\textwidth} \begin{center} \caption*{Type 4 (College graduate)} \vskip -6pt \end{center} \end{subfigure} \begin{subfigure}[t]{0.43\textwidth} \begin{center} \caption*{Type 5 (College-plus)} \vskip -6pt \end{center} \end{subfigure} \end{center} \end{subfigure} \vskip 10pt \begin{subfigure}[t]{\textwidth} \begin{center} \caption{With exchangeability (all types)} \vskip -2pt \begin{subfigure}[t]{0.45\textwidth} \begin{center} \end{center} \end{subfigure} \end{center} \end{subfigure} \vskip -20pt \caption{Marginal CDFs for the LFDs maximizing the “some college” to “college graduate” premium in cohort 1 across potential spouse types. } \end{center} \vskip -14pt \end{figure} \begin{table}[t] \begin{center} \caption{Metrics for interpreting $\delta$} { \begin{tabular}{ccccccccc} \hline \hline & & \multicolumn{3}{c}{Without exchangeability} & & \multicolumn{3}{c}{With exchangeability} \\ \cline{3-5} \cline{7-9} \\[-10pt] $\delta$ & & $\rho_{\max}$, $\ul \kappa_\delta$ & $\rho_{\max}$, $\ol \kappa_\delta$ & $size$ & & $\rho_{\max}$, $\ul \kappa_\delta$ & $\rho_{\max}$, $\ol \kappa_\delta$ & $size$ \\[2pt] 0.01 & & -0.015 & -0.014 & 0.010 & & -0.022 & \phantom{-}0.013 & 0.006 \\ 0.10 & & -0.071 & -0.073 & 0.038 & & -0.061 & \phantom{-}0.054 & 0.023 \\ 1 & & -0.247 & -0.197 & 0.112 & & -0.139 & \phantom{-}0.115 & 0.099 \\ 10 & & -0.502 & -0.496 & 0.242 & & -0.204 & \phantom{-}0.236 & 0.176 \\ 100 & & -0.620 & -0.576 & 0.266 & & -0.266 & \phantom{-}0.284 & 0.178 \\ \hline \\[-10pt] \end{tabular} } \parbox{\textwidth}{\small \emph{Note:} Averages across cohorts of the largest element of the correlation matrix for $U$ under the LFDs at which the estimated lower bounds ($\rho_{\max}$, $\ul \kappa_\delta$) and upper bounds ($\rho_{\max}$, $\ol \kappa_\delta$) are attained, and our $size$ measure from Section~\ref{s:lens}. Each is computed at the parameter values at which the estimated upper and lower bounds are attained.} \end{center} \vskip -15pt \end{table} To understand better what is meant by ``small'' and ``large'' neighborhoods, Figure~\ref{fig:csw_3_qq} plots marginal CDFs for the LFDs under which the upper bounds for cohort 1 are attained. Similar LFDs (not reported) were obtained for other cohorts and the lower bounds. Without exchangeability, the LFDs with $\delta = 0.1$ are almost identical to Gumbel (plots with $\delta = 0.01$ are indistinguishable from Gumbel). LFDs appear close to Gumbel across most potential spouse types with $\delta = 1$, while for $\delta = 10$ and $\delta = 100$ the LFDs have kinks and indicate shifts in mass from the center of the distribution to the tails. Under exchangeability (Figure~\ref{fig:csw_3_lfd_exch}), the marginal distribution of shocks is independent of potential spouse type. In this case the LFDs for $\delta = 1$ or smaller are virtually indistinguishable from Gumbel. LFDs with $\delta = 10$ and $\delta = 100$ are also less kinked than Figure~\ref{fig:csw_3_lfd} because distortions are spread more evenly across potential spouse types. We also computed the largest correlation of shocks under the LFDs at which the bounds are attained and our $size$ measure from Section~\ref{s:lens}. As these quantities are stable across cohorts, we present their averages across cohorts in Table~\ref{tab:csw_3}. Shocks are independent when $\delta = 0$ and only very weakly correlated for small $\delta$, while for large $\delta$ some shocks are strongly negatively correlated. The maximal correlations under exchangeability are smaller, especially for large $\delta$. Turning to the $size$ measure, the LFDs for $\delta = 0.01$ without exchangeability shift the model-implied marriage probabilities by 0.01 (on average, across cohorts) from their values under the i.i.d. Gumbel assumption. LFDs for $\delta = 10$ and $\delta = 100$ shift marriage probabilities around 0.25 (on average, across cohorts). Imposing exchangeability reduces the $size$ measure by around 25\% because model parameters do not vary as much under this shape restriction. In view of the small-$\delta$ bounds in Figure~\ref{fig:csw_3}, the LFDs in Figure~\ref{fig:csw_3_qq}, and the metrics in Table~\ref{tab:csw_3}, it seems impossible to draw conclusions about how the sign of the premium has changed over time under slight nonparametric relaxations of the i.i.d. Gumbel assumption. To help understand why, Figure~\ref{fig:csw_3_inner_both} plots bounds where $F$ is allowed to vary but $\theta$ is held fixed at CSW's estimates. These ``fixed-$\theta$'' bounds for $\delta = 10$ and $\delta = 100$ are almost identical, and are roughly the same width as the $\delta = 0.01$ bounds in Figure~\ref{fig:csw_3}. The width of the bounds in Figure~\ref{fig:csw_3} therefore seems largely attributable to the additional variation in $\theta$ that is permitted when parametric assumptions for $F$ are relaxed. Overall, our findings are complementary to \cite{GualdaniSinha} who perform a nonparametric reanalysis of CSW using the PIES methodology of \cite{Torgovitsky2019QE}. Although they do not derive nonparametric bounds on the marital education premium itself, only terms that contribute to it, they also find no evidence of an increase in premiums across cohorts. \subsection{Welfare Analysis in a Rust Model}\label{s:rust} Our second empirical illustration is a sensitivity analysis for welfare counterfactuals in the DDC model of \cite{Rust}. \medskip \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont\normalsize\bfseries}{Model and Benchmark Estimates.} We focus on the specification in Table IX of \cite{Rust} where maintenance costs are linear in the state (i.e., mileage). In the notation of Example~\ref{ex:DDC}, $|\mc S| = 90$, $\beta = 0.9999$, and $\theta_\pi = (RC, MC)$ where $RC$ is the replacement cost and $MC$ is a maintenance cost parameter. Our counterfactual of interest is the change in average welfare arising from a 10\% reduction in maintenance costs. Hence, $\pi_{1,s}(\theta_\pi) = \tilde \pi_{1,s}(\theta_\pi) = -RC$ and $\pi_{0,s}(\theta_\pi) = -0.001 MC \times s$ (baseline) and $\tilde \pi_{0,s}(\theta_\pi) = 0.9 \pi_{0,s}(\theta_\pi)$ (counterfactual). The counterfactual function is $k(\theta, \gamma) = w' (\tilde v - v)$ where $w$ is the stationary distribution of the state in the baseline model. Under the i.i.d. Gumbel assumption, the estimated counterfactual at the maximum likelihood estimate (MLE) of $\theta_\pi$ is 73.07 and its 95\% CS is [48.25,101.31].\footnote{We construct this CS by simulation. We draw $\hat \theta^*_\pi \sim N(\hat \theta_\pi,\hat \Sigma)$ where $\hat \theta_\pi$ is the MLE and $\hat \Sigma$ is an estimate of the inverse information matrix. For each $\hat \theta^*_\pi$ draw, we compute the baseline and counterfactual value functions $v^*$ and $\tilde v^*$, and hence the counterfactual $\hat \kappa^* = w'( \tilde v^* - v^*)$.} Note the counterfactual is point-identified under the i.i.d. Gumbel assumption because $\theta_\pi$ is point-identified. \medskip \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont\normalsize\bfseries}{Implementation.} We estimate CCPs using Rust's Group 4 data. Nonparametric estimates of the 90 CCPs are zero in many states, so we proceed as in Section~\ref{s:overid} and take the model-implied CCPs at the MLE of $\theta_\pi$ (under the i.i.d. Gumbel assumption) as our estimate $\hat P_2$. We drop moment conditions for CCPs in states where the replacement probability is less than $0.001$ to avoid numerical instabilities induced by including near-degenerate moments. This reduces the dimension of $g_2$ to 71. We normalize $F$ so that shocks have mean zero and the same variance as the Gumbel distribution by appending $\mathbb{E}^F[U_d] = 0$ and $\mathbb{E}^F[U_d^2 - \pi^2/6] = 0$, for $d = 0,1$, to $g_4$. In total, there are 255 moments (71 for CCPs, 180 for Bellman equations, and 4 location/scale normalizations) and $\theta = (\theta_\pi, v, \tilde v)$ has dimension 182. We implement our methods as described in Section~\ref{s:mpec}. The inner optimization uses 75 moments (71 for CCPs and 4 for normalizations), with the remaining 180 moments appended as constraints in the outer optimization. We define neighborhoods using hybrid divergence from Section~\ref{s:csw} so that Assumption~\ref{a:phi}(ii) holds. Similar results are obtained with $\chi^2$ and $L^4$ neighborhoods (see Appendix~\ref{ax:rust}). Expectations are computed using 50,000 scrambled Halton draws---see Appendix~\ref{ax:rust} for computation times. We compute 95\% CSs for $\ul \kappa_\delta$ and $\ol \kappa_\delta$ using the bootstrap procedure from Section~\ref{s:bootstrap} and projection procedure from Section~\ref{s:projection}. See Appendix~\ref{ax:rust} for details. \medskip \@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont\normalsize\bfseries}{Findings.} \begin{figure} \begin{center} \includegraphics[width=0.6\textwidth]{rust_bounds_2.pdf} \caption{\label{fig:rust_bounds} Sensitivity analysis for change in average welfare under a 10\% maintenance cost subsidy. \emph{Note:} Solid lines are estimates, dotted lines are bootstrap CSs, dashed lines are projection CSs.} \end{center} \vskip -14pt \end{figure} Estimates and CSs for $\ul \kappa_\delta$ and $\ol \kappa_\delta$ are plotted in Figure~\ref{fig:rust_bounds} for values of $\delta$ from $0.01$ to $100$.\footnote{The width of the bootstrap CSs relative to the bounds reduces as $\delta$ gets large. We re-estimated our bounds using several different draws of bootstrapped CCPs in place of $\hat P_2$ and obtained bounds that spanned a range similar to the bootstrap CSs for small $\delta$, but which for many draws converged to values close to our estimates of the bounds for large $\delta$. This corroborates the behavior of our bootstrap CSs. We conjecture that other features of the model are potentially more important than the numerical values of the CCPs in determining nonparametric bounds on the welfare counterfactual. } As can be seen, the bounds expand rapidly under slight relaxations of the i.i.d. Gumbel assumption then stabilize around $\delta = 1$, where the lower bound is 6.45 and the upper bound of 160.5 represents approximately 220\% of the value under the i.i.d. Gumbel assumption. \begin{figure}[t] \begin{center} \begin{subfigure}{0.43\textwidth} \begin{center} \caption*{Lower bound} \label{fig:rust_ldf_lower} \includegraphics[width=\linewidth]{rust_lfd_2.pdf} \end{center} \end{subfigure} \begin{subfigure}{0.43\textwidth} \begin{center} \caption*{Upper bound} \label{fig:rust_lfd_upper} \includegraphics[width=\linewidth]{rust_lfd_1.pdf} \end{center} \end{subfigure} \caption{\label{fig:rust_lfd} CDFs of $U_1 - U_0$ under the LFDs at which the estimated lower and upper bounds on the welfare counterfactual are attained.} \end{center} \end{figure} To interpret $\delta$, in Figure~\ref{fig:rust_lfd} we plot the CDFs of $U_1 - U_0$ under the LFDs at which the estimated bounds $\hat{\ul \kappa}_\delta$ and $\hat{\ol \kappa}_\delta$ are attained. LFDs were computed as described in Section~\ref{s:lfd} using the construction (\ref{e:m:delta}). The distributions appear very close to logistic (their distribution when $\delta = 0$) for $\delta = 0.01$. Therefore, we see that large differences in welfare counterfactuals can arise under very slight departures from the i.i.d. Gumbel assumption. LFDs for the upper bound shift increasing amounts of mass to the center of the distribution of $U_1 - U_0$ as $\delta$ increases. LFDs corresponding to the lower bound are relatively less distorted, but have increasing amounts of mass shifted into the right tail. These are similar for $\delta = 0.1$ through $\delta = 100$ because the estimated lower bound stabilizes for smaller values of $\delta$ than the upper bound (cf. Figure~\ref{fig:rust_bounds}). \begin{table}[t] \begin{center} \caption{\label{tab:rust}Metrics for interpreting $\delta$} { \begin{tabular}{ccccccccccc} \hline \hline & & \multicolumn{4}{c}{Lower bound} & & \multicolumn{4}{c}{Upper bound} \\ \cline{3-6} \cline{8-11} \\[-10pt] $\delta$ & & $corr$ & $size$ & $RC$ & $MC$ & & $corr$ & $size$ & $RC$ & $MC$ \\[2pt] 0 & & \phantom{-}0.000 & 0.000 & 10.208 & 2.294 & & \phantom{-}0.000 & 0.000 & 10.208 & 2.294 \\ 0.01 & & \phantom{-}0.036 & 0.010 & \phantom{1}7.357 & 1.411 & & -0.027 & 0.016 & 13.390 & 3.307 \\ 0.1 & & -0.058 & 0.039 & \phantom{1}5.186 & 0.553 & & \phantom{-}0.149 & 0.109 & 16.134 & 4.374 \\ 1 & & -0.045 & 0.039 & \phantom{1}4.023 & 0.203 & & \phantom{-}0.616 & 0.346 & 17.166 & 5.038 \\ 10 & & -0.040 & 0.039 & \phantom{1}4.022 & 0.202 & & \phantom{-}0.765 & 0.461 & 17.595 & 5.331 \\ 100 & & -0.063 & 0.039 & \phantom{1}3.931 & 0.176 & & \phantom{-}0.764 & 0.469 & 17.626 & 5.365 \\ \hline \\[-10pt] \end{tabular} } \parbox{\textwidth}{\small \emph{Note:} Correlation of $U_0$ and $U_1$ under the LFD at which the estimated lower and upper bounds are attained ($corr$), our $size$ measure from Section~\ref{s:lens}, and replacement and maintenance cost parameters at which the estimated lower and upper bounds are attained.} \end{center} \end{table} Table~\ref{tab:rust} lists other metrics to help interpret the neighborhood size. The first is the correlation of $U_0$ and $U_1$ under the LFDs at which $\hat{\ul \kappa}_\delta$ and $\hat{\ol \kappa}_\delta$ are attained. These are very small for $\delta = 0.01$ and remain small under the LFDs for $\hat{\ul \kappa}_\delta$ as $\delta$ increases, while $U_0$ and $U_1$ are strongly positively correlated under the LFDs for $\hat{\ol \kappa}_\delta$, especially for larger $\delta$ values. Given the asymmetry in distortions between the lower and upper values, we compute our $size$ measure separately for both. We measure distortions to the moments corresponding to the CCPs as these are most directly interpretable within the context of the model. We see that the LFDs for $\delta = 0.01$ are distorting $F_*$ in a manner that shifts the model-implied CCPs by at most 0.016. By contrast, the LFDs for $\delta = 10$ and $\delta = 100$ shift the model-implied CCPs from their values under the i.i.d. Gumbel assumption by at most $0.04$ for $\hat{\ul \kappa}_\delta$ and $0.47$ for $\hat{\ol \kappa}_\delta$. The parameters at which $\hat{\ul \kappa}_\delta$ and $\hat{\ol \kappa}_\delta$ are attained are also revealing about neighborhood size. Table~\ref{tab:rust} presents MLEs of $MC$ and $RC$, which are similar to the values reported in Table IX of \cite{Rust}. We see from Table~\ref{tab:rust} that $\hat{\ul \kappa}_\delta$ and $\hat{\ol \kappa}_\delta$ are attained at very different parameter values, with much smaller cost parameters for the lower bound and larger parameters for the upper bound, even for $\delta = 0.01$. Intuitively, a smaller $MC$ means that the saving from the subsidy---which is proportional---must be small. Correspondingly, a low $RC$ is needed to help the model to fit the observed CCPs at the smaller $MC$. While it is known that payoff parameters are not identified without parametric assumptions on $F$, it is perhaps surprising that these parameters vary by so much under slight relaxations of the i.i.d. Gumbel assumption. For instance, with $\delta = 0.01$ the lower bound is attained with cost parameters $RC = 7.357$ and $MC = 1.411$ while the upper bound is attained with cost parameters that are roughly double these values. \section{Estimation and Inference}\label{s:asymptotics} We begin in Section~\ref{s:consistency} by establishing consistency and the asymptotic distribution of the estimators $\hat{\ul \kappa}_\delta$ and $\hat{\ol \kappa}_\delta$ from Section~\ref{s:estimators}. We then present a bootstrap-based inference method in Section~\ref{s:bootstrap} and a projection-based inference method in Section~\ref{s:projection}. \subsection{Large-sample Properties of Plug-in Estimators}\label{s:consistency} We first introduce some regularity conditions. Recall the space $\mc E$ from Assumption~\ref{a:phi}. We equip $\mc E$ with the Orlicz norm (see Appendix~\ref{ax:Orlicz}) \[ \|f\|_\psi = \inf_{c > 0} \frac{1}{c} \left( 1 + \mathbb{E}^{F_*}[\psi(c|f(U)|)] \right). \] This norm is equivalent to the $L^2(F_*)$ norm for $\chi^2$ and hybrid divergence and equivalent to the $L^q(F_*)$ norm for $L^p$ divergence ($p^{-1} + q^{-1} = 1$), while for KL divergence it is stronger than any $L^p(F_*)$ norm with $p < \infty$ but weaker than the sup-norm. Say that a class of functions $\{f_a : a \in \mc A\} \subset \mc E$ indexed by a metric space $\mc A$ is \emph{$\mc E$-continuous in $a$} if $a' \to a$ in $\mc A$ implies $\|f_a - f_{a'}\|_\psi \to 0$. We also require a slightly stronger notion of constraint qualification than Condition S from Section (ref). \begin{definition} \emph{Condition S'} holds at $(\theta,\gamma,P)$ if $\vec P \in \mathrm{int}(\mc G(\theta,\gamma) + \mc C) $. \end{definition} Condition S' replaces “relative interior” in Condition S with “interior”. Finally, recall $\Delta(\theta;\gamma,P)$ from ((ref)) and let $\Theta_\delta(\gamma,P) = \{\theta \in \Theta : \Delta(\theta;\gamma,P) < \delta\}$. \begin{assumptionp}{M} (i)\hskip\labelsep $k(\cdot;\theta,\gamma)$ and each entry of $g(\cdot;\theta,\gamma)$ are $\mc E$-continuous in $(\theta,\gamma)$; \begin{enumerate}[topsep=-18pt,itemsep=0pt,parsep=0pt,partopsep=0pt] • $(\theta,\gamma) \mapsto \mb E^{F_*}[\phi^\star(a_1 + a_2 k(U,\theta,\gamma) + a_3' g(U,\theta,\gamma))]$ is continuous for each $(a_1,a_2,a_3) \in \mb R \times \mb R \times \mb R^{d}$; • $\Theta_\delta(\gamma_0,P_0)$ is nonempty and Condition S' holds at $(\theta,\gamma_0,P_0)$ for each $\theta \in \Theta_\delta(\gamma_0,P_0)$; • $\mathrm{cl}(\Theta_\delta(\gamma_0,P_0)) \supseteq \{ \theta \in \Theta : \Delta(\theta;\gamma_0,P_0) \leq \delta\}$; • $\Theta$ is a compact subset of $\mb R^{d_\theta}$. \end{enumerate} \end{assumptionp} Parts (i) and (ii) of Assumption (ref) are continuity conditions. If $k$ and $g$ consist of indicator functions, then these conditions hold provided the probabilities of the events under $F_*$ are continuous in $(\theta,\gamma)$. In models without $\gamma$, these conditions simply require continuity in $\theta$. There are two parts to Assumption (ref)(iii). The nonemptyness condition holds when the model is correctly specified under $F_*$ or, more generally, when there is at least one $F \in \mc N_\delta$ that satisfies ((ref)) for some $\theta$. The second part is a constraint qualification. This condition requires that for each $\theta \in \Theta_\delta(\gamma_0,P_0)$, there is a distribution $F$ under which ((ref)) holds at $(\theta,\gamma_0,P_0)$ that is “interior” to $\mc N_\infty$, in the sense that one can perturb the moments at $(\theta,\gamma_0,P_0)$ in all directions by perturbing $F$. Condition S' also requires that there is $F \in \mc N_\infty$ under which any inequality restrictions at $(\theta, \gamma_0, P_0)$ hold strictly. Note, however, that we do not require that this $F$ belongs to $\mc N_\delta$, only to $\mc N_\infty$. We therefore do not view this condition as overly restrictive. We also conjecture it could be relaxed using a notion similar to $S$-regularity from Section (ref). Assumption (ref)(iv) is made for convenience and can be relaxed; this condition simply ensures that there do not exist values of $\theta$ at which $\Delta(\theta;\gamma_0,P_0) = \delta$ that are separated from $\Theta_\delta(\gamma_0,P_0)$. Assumption (ref)(v) is standard and can be relaxed. \begin{theorem} Suppose that Assumptions (ref) and (ref) hold and $(\hat \gamma,\hat P) \to_p (\gamma_0,P_0)$ or, if there is no auxiliary parameter, $\hat P \to_p P_0$. Then $\hat{\ul \kappa}_\delta \to_p \ul \kappa_\delta$ and $\hat{\ol \kappa}_\delta \to_p \ol \kappa_\delta$. \end{theorem} To derive the asymptotic distribution of the estimators, we assume $\gamma_0$ is known and suppress dependence of all quantities on $\gamma$ for the remainder of this section. This entails no loss of generality for models without $\gamma$, such as Examples (ref) and (ref) and the application in Section (ref). In DDC models this presumes the law of motion of the state is known. The asymptotic distribution therefore reflects only sampling uncertainty from the estimated CCPs, which is the case for confidence sets reported when laws of motion are first estimated “offline”. Extending our approach to accommodate sampling variation in $\hat \gamma$ in a tractable manner appears to require exploiting application-specific model structure, which we defer to future work. Define \begin{equation} \begin{aligned} \ul b_\delta(P) & = \inf_{\theta \in \Theta_\delta(P)} \ul K_\delta(\theta;P) \,, & & & \ol b_\delta(P) & = \sup_{\theta \in \Theta_\delta(P)} \ol K_\delta(\theta;P) \,. \end{aligned} \end{equation} In this notation, $\ul \kappa_\delta = \ul b_\delta(P_0)$ and $\ol \kappa_\delta = \ol b_\delta(P_0)$ (see Lemma (ref)) and $\hat{\ul \kappa}_\delta = \ul b_\delta(\hat P)$ and $\hat{\ol \kappa}_\delta = \ol b_\delta(\hat P)$. We derive the asymptotic distribution of $\hat{\ul \kappa}_\delta$ and $\hat{\ol \kappa}_\delta$ by showing $\ul b_\delta$ and $\ol b_\delta$ are directionally differentiable and applying a suitable delta method. Say $f : \mb R^{d_1+d_2} \to \mb R$ is (Hadamard) directionally differentiable at $P_0$ if there is a continuous map $df_{P_0}[\,\cdot\, ] : \mb R^{d_1+d_2} \to \mb R$ such that \[ \lim_{n \to \infty} t_n^{-1} \left( f(P_0+t_n h_n)-f(P_0) \right) = df_{P_0}[h] \] for all sequences $t_n \downarrow 0$ and $h_n \to h$ Shapiro1990. If $df_{P_0}[h]$ is linear in $h$ then $f$ is (fully) differentiable at $P_0$. We introduce some additional notation used to define the directional derivatives of $\ul b_\delta$ and $\ol b_\delta$. Let \begin{align*} \ul \Xi_\delta(\theta;P) & = \mathrm{argsup}_{\eta \geq 0, \zeta \in \mb R, \lambda \in \Lambda} -\mathbb{E}^{F_*}\Big[ (\eta \phi)^\star( -k(U,\theta) - \zeta - \lambda' g(U,\theta)) \Big] - \eta \delta - \zeta - \lambda_{12}' P \,, \end{align*} where $(\eta\phi)^\star$ denotes the convex conjugate of $x \mapsto \eta \cdot \phi(x)$, and let $\ol \Xi_\delta(\theta;P)$ denote the analogous arginf for the minimization problem corresponding to the upper bound. Recall that $\ul \lambda_{12} = (\ul \lambda_1,\ul \lambda_2)$ collects the first $d_1 + d_2$ elements of $\ul \lambda$. Let \[ \ul \Lambda_\delta(\theta;P) = \{ ( \lambda_1, \lambda_2) : ( \eta, \zeta, \lambda_1, \lambda_2, \lambda_3, \lambda_4) \in \ul \Xi_\delta(\theta;P)\} \] denote the projection of $\ul \Xi_\delta(\theta;P)$ for $\ul \lambda_{12}$. We let $\ol \Lambda_\delta(\theta;P)$ denoting the analogous projection of $\ol \Xi_\delta(\theta;P)$. Finally, let \[ \begin{aligned} \ul \Theta_\delta(P_0) & = \mathrm{arg}\min_{\theta \in \Theta} \ul K_\delta(\theta;P_0)\,, & \ol \Theta_\delta(P_0) & = \mathrm{arg}\max_{\theta \in \Theta} \ol K_\delta(\theta;P_0)\,. \end{aligned} \] The sets $\ul \Theta_\delta(P_0)$ and $\ol \Theta_\delta(P_0)$ are nonempty and compact under Assumptions (ref) and (ref). The following regularity conditions are presented for the general case where $k$ depends on $u$. It may be possible to weaken some of these regularity conditions in the special case in which $k$ does not depend on $u$. \begin{assumptionp}{(ref) (continued)} (vi)\hskip\labelsep $\ul \Theta_\delta(P_0) \subseteq \Theta_\delta(P_0)$ and $\ol \Theta_\delta(P_0) \subseteq \Theta_\delta(P_0)$; \begin{enumerate}[topsep=-18pt,itemsep=0pt,parsep=0pt,partopsep=0pt] • $\theta \mapsto \ul \Lambda_\delta(\theta;P_0)$ and $\theta \mapsto \ol \Lambda_\delta(\theta;P_0)$ are lower hemicontinuous at each $\theta \in \ul \Theta_\delta(P_0)$ and $\theta \in \ol \Theta_\delta(P_0)$, respectively. \end{enumerate} \end{assumptionp} \begin{theorem} Suppose that Assumptions (ref) and (ref) hold. Then $\ul b_\delta$ and $\ol b_\delta$ are directionally differentiable at $P_0$, with \begin{align*} d\ul b_{\delta,P_0}[h] & = \min_{\theta \in \ul \Theta_\delta(P_0)} \max_{\ul \lambda_{12} \in \ul \Lambda_\delta(\theta;P_0)} -\ul \lambda_{12}' h \,, & d\ol b_{\delta,P_0}[h] & = \max_{\theta \in \ol \Theta_\delta(P_0)} \min_{\ol \lambda_{12} \in \ol \Lambda_\delta(\theta;P_0)} \ol \lambda_{12}' h \,. \end{align*} Moreover, if $\sqrt n (\hat P - P_0) \to_d Z \sim N(0,\Sigma)$ with $\Sigma$ finite, then \[ \sqrt n \left( \left( \begin{array}{c} \hat{\ol \kappa}_\delta \\ \hat{\ul \kappa}_\delta \end{array} \right) - \left( \begin{array}{c} \ol \kappa_\delta \\ \ul \kappa_\delta \end{array} \right) \right) \to_d \left( \begin{array}{c} d\ul b_{\delta,P_0}[Z] \\ d\ol b_{\delta,P_0}[Z] \end{array} \right) . \] \end{theorem} The asymptotic distribution presented in Theorem (ref) is non-Gaussian. In the special case in which $\cup_{\theta \in \ul \Theta_\delta(P_0)} \ul \Lambda_\delta( \theta;P_0) = \{\ul \lambda_{12}\}$, the asymptotic distribution of $\hat{\ul \kappa}_\delta$ simplifies to $N(0,\ul \lambda_{12}'\Sigma \ul \lambda_{12})$. An analogous simplification holds for $\hat{\ol \kappa}_\delta$ when $\cup_{\theta \in \ol \Theta_\delta(P_0)} \ol \Lambda_\delta( \theta;P_0)$ is a singleton. \subsection{Inference Procedure 1: Bootstrap} Our first inference procedure specializes the general approach of FangSantos for inference on directionally differentiable functions to the present setting. Define \begin{equation*} \begin{aligned} \widehat{d\ul b}_{\delta,P_0}[h] & = \inf_{\theta \in \hat{\ul \Theta}_{\delta,n}} \sup_{\ul \lambda_{12} \in \ul \Lambda_\delta(\theta;\hat P)} -\ul \lambda_{12}' h \,, & & & \widehat{d\ol b}_{\delta,P_0}[h] & = \sup_{\theta \in \hat{\ol \Theta}_{\delta,n}} \inf_{\ol \lambda_{12} \in \ol \Lambda_\delta(\theta;\hat P)} \ol \lambda_{12}' h \,, & \end{aligned} \end{equation*} where \begin{equation*}\begin{aligned} \hat{\ul \Theta}_{\delta,n} & = \{\theta \in \Theta_\delta(\hat P) : \ul K_\delta(\theta; \hat P) \leq \hat{\ul \kappa}_\delta + \hat \nu \sqrt{\log n/n}\} \,, \mbox{ and}\\ \hat{\ol \Theta}_{\delta,n} & = \{\theta \in \Theta_\delta(\hat P) : \ol K_\delta(\theta; \hat P) \geq \hat{\ol \kappa}_\delta - \hat \nu \sqrt{\log n/n}\} \,, \end{aligned} \end{equation*} with $\hat \nu$ a (possibly random) positive scalar tuning parameter for which $\hat \nu \to_p \nu > 0$. Any such $\hat \nu$ results in a confidence set with asymptotically correct coverage. We give some practical guidance for choosing $\hat \nu$ below. Let $\hat P^*$ denote a bootstrapped version of $\hat P$. In practice any bootstrap can be used provided it satisfies mild consistency conditions. In the empirical application in Section (ref) we simply draw $\hat P^* \sim N(\hat P,\hat \Sigma/n)$ where $\hat \Sigma$ is a consistent estimator of $\Sigma$. Let \begin{align*} \hat {\ul c}_{\alpha} & = \mbox{ $\alpha$-quantile of } \widehat{d\ul b}_{\delta,P_0}[\sqrt n(\hat P^* - \hat P)] \,, & \hat {\ol c}_{\alpha} & = \mbox{ $\alpha$-quantile of } \widehat{d\ol b}_{\delta,P_0}[\sqrt n(\hat P^* - \hat P)] \,, \end{align*} where the quantiles are computed by resampling $\hat P^*$ (conditional on the data). Lower, upper, and two-sided $100(1-\alpha)$% CSs for $\ul \kappa_\delta$ and $\ol \kappa_\delta$ are \begin{align*} & CS_{\delta,L}^{1-\alpha} = \left[ \hat{\ul \kappa}_\delta - { \frac{\hat {\ul c}_{1-\alpha}}{\sqrt n} }, +\infty\right) , \\ & CS_{\delta,U}^{1-\alpha} = \left(-\infty, \hat{\ol \kappa}_\delta - { \frac{\hat {\ol c}_{\alpha}}{\sqrt n} } \right] , & & CS_\delta^{1-\alpha} = \left[ \hat{\ul \kappa}_\delta - { \frac{\hat {\ul c}_{1-\alpha/2}}{\sqrt n}}, \hat{\ol \kappa}_\delta - { \frac{\hat {\ol c}_{\alpha/2}}{\sqrt n} } \right] . \end{align*} We require a slight strengthening of Assumption (ref)(vii) to establish validity of the procedure. As before, regularity conditions are presented for the general case where $k$ depends on $u$. It may be possible to weaken these conditions when $k$ does not depend on $u$. \begin{assumptionp}{(ref) (continued)} (vii')\hskip\labelsep $(\theta,P) \mapsto \ul \Lambda_\delta(\theta;P)$ and $(\theta,P) \mapsto \ol \Lambda_\delta(\theta;P)$ are lower hemicontinuous at $(\theta,P_0)$ for each $\theta \in \ul \Theta_\delta(P_0)$ and $\theta \in \ol \Theta_\delta(P_0)$, respectively. \end{assumptionp} \begin{theorem} Suppose that Assumptions (ref) and (ref)(i)--(vi),(vii') hold, $\sqrt n (\hat P - P_0) \to_d Z \sim N(0,\Sigma)$ with $\Sigma$ finite, and $\hat P^*$ satisfies Assumption 3 of FangSantos. Then the distribution of $\widehat{d\ul b}_{\delta,P_0}[\sqrt n(\hat P^* - \hat P)]$ and $\widehat{d\ol b}_{\delta,P_0}[\sqrt n(\hat P^* - \hat P)]$ (conditional on the data) is consistent for the asymptotic distribution derived in Theorem (ref). Moreover, if the CDFs of $d\ul b_{\delta,P_0}[Z]$ and $d\ol b_{\delta,P_0}[Z]$ are continuous and increasing at their $\alpha/2$, $\alpha$, $1-\alpha$, and $1-\alpha/2$ quantiles, then \begin{align*} \lim_{n \to \infty} \Pr( \ul \kappa_\delta \in CS_{\delta,L}^{1-\alpha}) & = 1-\alpha \,, \\ \lim_{n \to \infty} \Pr( \ol \kappa_\delta \in CS_{\delta,U}^{1-\alpha}) & = 1-\alpha \,, & \liminf_{n \to \infty} \Pr( [\ul \kappa_\delta,\ol \kappa_\delta] \subseteq CS_\delta^{1-\alpha}) & \geq 1-\alpha \,. \end{align*} \end{theorem} Any $\hat \nu$ that satisfies $\hat\nu \to_p \nu > 0$ results in asymptotically valid CSs. In view of the functional forms of $\widehat{d\ul b}_{\delta,P_0}[\,\cdot\,]$ and $\widehat{d\ul b}_{\delta,P_0}[\,\cdot\,]$, smaller $\hat \nu$ produce (weakly) wider CSs. In the CSW application, we set $\hat \nu$ equal to the minimum diagonal element of the covariance matrix of the moments evaluated at $(\hat \theta, \hat \gamma,\hat P)$ under $F_*$, where $\hat \theta$ is computed under $F_*$. We chose this quantity as it is related to the convexity of the inner problem for small $\delta$. In practice, this resulted in $\hat \nu$ between 0.001 and 0.01. We recommend setting $\hat \nu$ to be of a similarly small magnitude, then performing a sensitivity analysis to check that critical values aren't too dependent on $\hat \nu$. Setting $\hat \nu = 0$ and replacing $\hat{\ul \Theta}_{\delta,n}$ and $\hat{\ol \Theta}_{\delta,n}$ by $\{\hat{\ul \theta}_\delta\}$ and $\{\hat{\ol \theta}_\delta\}$ where $\hat{\ul \theta}_\delta$ and $\hat{\ol \theta}_\delta$ minimize and maximize the sample criterions is also valid, but may be conservative. \subsection{Inference Procedure 2: Projection} This second approach is computationally simple but possibly conservative.\footnote{We are grateful to a referee for suggesting this approach.} Suppose we have random vectors $\hat P_{1,U}^{1-\alpha}$, $\hat P_{2,U}^{1-\alpha}$, and $\hat P_{2,L}^{1-\alpha}$ that form a $100(1-\alpha)$% rectangular CS for $P_0$: \begin{equation} \liminf_{n \to \infty} \Pr\left( P_{10} \leq \hat P_{1,U}^{1-\alpha}\,,\, \hat P_{2,L}^{1-\alpha} \leq P_{20} \leq \hat P_{2,U}^{1-\alpha} \right) \geq 1-\alpha \,, \end{equation} where the inequalities should be understood to hold element-wise (we discuss how to construct a rectangular CS for $P_0$ below). The idea behind this approach is to replace any moment conditions involving $P$ by inequalities constructed from the rectangular CS. Define the criterion functions \[ \hat{\ul K}_{\delta,1-\alpha}(\theta) = \left[ \begin{array}{l} \ul K_{\delta,cs} (\theta; \hat P_{1-\alpha}) \\[2pt] + \infty \end{array} \right. , \quad \hat{\ol K}_{\delta,1-\alpha}(\theta) = \left[ \begin{array}{ll} \ol K_{\delta,cs} (\theta; \hat P_{1-\alpha}) & \mbox{if $\Delta_{cs}(\theta;\hat P_{1-\alpha}) < \delta$,} \\[2pt] - \infty & \mbox{if $\Delta_{cs}(\theta;\hat P_{1-\alpha}) \geq \delta$,} \end{array} \right. \] where $\ul K_{\delta,cs}$, $\ol K_{\delta,cs}$, and $\Delta_{cs}$ are versions of ((ref)), ((ref)), and ((ref)) formed using \begin{equation} \begin{aligned} \mb E^{F}[ g_1(U,\theta)] & \leq \hat P_{1,U}^{1-\alpha} \,, & \mb E^{F}[ g_2(U,\theta)] & \leq \hat P_{2,U}^{1-\alpha} \,, & \mb E^{F}[ -g_2(U,\theta)] & \leq -\hat P_{2,L}^{1-\alpha} \,, \end{aligned} \end{equation} as well as ((ref)) and ((ref)). In these criterions, $\Lambda$ is replaced by $\Lambda_{cs} = \mb R_+^{d_1 + 2 d_2 + d_3} \times \mb R^{d_4}$, $g$ is replaced by $g_{cs} = (g_1, g_2, -g_2, g_3, g_4)$, $P$ is replaced by $\hat P_{1-\alpha} = (\hat P_{1,U}^{1-\alpha},\hat P_{2,U}^{1-\alpha},-\hat P_{2,L}^{1-\alpha})$, and $\lambda_{12}$ denotes the first $d_1 + 2d_2$ elements of $\lambda$. Critical values are computed by optimizing the criterions $\hat{\ul K}_{\delta,1-\alpha}$ and $\hat{\ol K}_{\delta,1-\alpha}$ with respect to $\theta$: \begin{equation*} \begin{aligned} \hat{\ul \kappa}_{\delta,1-\alpha} & = \inf_{\theta \in \Theta} \hat{\ul K}_{\delta,1-\alpha}(\theta) \,, & \hat{\ol \kappa}_{\delta,1-\alpha} & = \sup_{\theta \in \Theta} \hat{\ol K}_{\delta,1-\alpha}(\theta)\,. \end{aligned} \end{equation*} Lower, upper, and two-sided $100(1-\alpha)$% CSs for $\ul \kappa_\delta$ and $\ol \kappa_\delta$ are then given by \begin{align*} & CS_{\delta,L}^{1-\alpha} = \left[ \hat{\ul \kappa}_{\delta,1-\alpha}, +\infty \right) \,, & & CS_{\delta,U}^{1-\alpha} = \left(-\infty, \hat{\ol \kappa}_{\delta,1-\alpha} \right] \,, & & CS_\delta^{1-\alpha} = \left[ \hat{\ul \kappa}_{\delta,1-\alpha}, \hat{\ol \kappa}_{\delta,1-\alpha} \right] \,. \end{align*} \begin{theorem} Suppose that Assumptions (ref) and (ref)(i),(iii)--(v) hold and $\hat P_{1-\alpha}$ satisfies ((ref)). Then \begin{align*} \liminf_{n \to \infty} \Pr( \ul \kappa_\delta \in CS_{\delta,L}^{1-\alpha}) & \geq 1-\alpha \,, \\ \liminf_{n \to \infty} \Pr( \ol \kappa_\delta \in CS_{\delta,U}^{1-\alpha}) & \geq 1-\alpha \,, & \liminf_{n \to \infty} \Pr( [\ul \kappa_\delta,\ol \kappa_\delta] \subseteq CS_\delta^{1-\alpha}) & \geq 1-\alpha\,. \end{align*} \end{theorem} To construct a rectangular CS for $P_0$ satisfying ((ref)), suppose $\sqrt n(\hat P - P_0) \to_d N(0,\Sigma)$ and we have a consistent estimator $\hat \Sigma$ of $\Sigma$. Let $\hat \sigma$ denote the vector formed by taking the square root of each diagonal entry of $\hat \Sigma$. Partition $\hat \sigma$ conformably as $\hat \sigma = (\hat \sigma_{(1)},\hat \sigma_{(2)})$ and set \begin{align*} \hat P_{1,L}^{1-\alpha} & = \hat P_1 + n^{-1/2} \hat c_{1-\alpha,1} \hat \sigma_{(1)} \,, & \hat P_{2,L}^{1-\alpha} & = \hat P_2 - n^{-1/2} \hat c_{1-\alpha,2} \hat \sigma_{(2)} \,, & \hat P_{2,U}^{1-\alpha} & = \hat P_2 + n^{-1/2} \hat c_{1-\alpha,2} \hat \sigma_{(2)} \,, \end{align*} where the (scalar) critical values $\hat c_{1-\alpha,1}$ and $\hat c_{1-\alpha,2}$ solve \[ \Pr\left(\max_{1 \leq i \leq d_1} Z_i/\hat \sigma_i \leq \hat c_{1-\alpha,1},\max_{d_1+1 \leq i \leq d_2} |Z_i/\hat \sigma_i| \leq \hat c_{1-\alpha,2}\right) = 1-\alpha \,, \quad Z \sim N(0, \hat \Sigma)\,. \] If $d_2 = 0$, then $\hat c_{1-\alpha,1}$ is the $(1-\alpha)$-quantile of $\max_{1 \leq i \leq d_1} Z_i/\hat \sigma_i$; similarly, if $d_1 = 0$, then $\hat c_{2,1-\alpha}$ is the $(1-\alpha)$-quantile of $\max_{1 \leq i \leq d_2} |Z_i/\hat \sigma_i|$. \section{Conclusion} This paper introduced a framework for analyzing the sensitivity of counterfactuals to parametric assumptions about the distribution of latent variables in structural models. In particular, we derived bounds on the set of counterfactuals obtained as the distribution of latent variables spans nonparametric neighborhoods of a given parametric specification while other “structural” model features are maintained. We illustrated our procedure with empirical applications to matching models and dynamic discrete choice. \let\oldbibliography\thebibliography \putbib
bibunit