Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,110 characters · 14 sections · 100 citation commands
Uniform inference for value functions
\thispagestyle{empty}
Keywords: uniform inference, value function, bootstrap, delta method, directional differentiability
JEL codes: C12, C14, C21
\setcounter{page}{1}
The (optimal) value function, that is, a function defined by optimizing an objective function over one of multiple arguments, is used widely in economics, finance, statistics and other fields. The statistical behavior of estimated value functions appears, for instance, in the study of distribution and quantile functions of treatment effects under partial identification (see, e.g., FirpoRidder08, FirpoRidder19 and FanPark10,FanPark12). More recently, investigating counterfactual sensitivity, ChristensenConnault19 construct interval-identified counterfactual outcomes as the distribution of unobservables varies over a nonparametric neighborhood of an assumed model specification. Statistical inference for value functions can be analyzed at a point as in Roemisch06 or CarcamoCuevasRodriguez19 and inference can be carried out using the techniques of HongLi18 or FangSantos19.
This paper extends these pointwise inference methods to uniform inference for value functions that are estimated with nonparametric plug-in estimators. Uniform inference has an important role for the analysis of statistical models.\footnote{For example, AngristChernozhukovVal06 discuss the importance of uniform inference vis-\`a-vis pointwise inference in a quantile regression framework. Uniform testing methods allow for multiple comparisons across the entire distribution without compromising confidence levels, and uniform confidence bands differ from corresponding pointwise confidence intervals because they guarantee coverage over the family of confidence intervals simultaneously.} However, uniform inference for the value function can be quite challenging. In order to use standard statistical inference procedures, one must show that the map from objective function to value function is Hadamard (i.e., compactly) differentiable as a map between two functional spaces, which may require restrictive regularity conditions. As a result, the delta method, as the conventional tool used to analyze the distribution of nonlinear transformations of sample data, may not apply to marginal optimization as a map between function spaces due to a lack of differentiability. We propose a solution to this problem that allows for the use of the delta method to conduct uniform inference for value functions.
Because in general the map of objective function to value function is not differentiable, our solution is to directly analyze statistics that are applied to the value function. This is in place of a more conventional analysis that would first establish that the value function is well-behaved and as a second step use the continuous mapping theorem to determine the distribution of test statistics. In settings involving a chain of maps, this can be seen as choosing how to define the “links” in the chain to which the chain rule applies. The results are presented for $L_p$ functionals for $1 \leq p \leq \infty$, that are applied to a value function, which should cover many cases of interest. In particular, this family of functionals includes Kolmogorov-Smirnov and Cram\'er-von Mises tests.
By considering $L_p$ functionals of the value function, we may bypass the most serious impediment to uniform inference. However, these $L_p$ functionals are only Hadamard directionally differentiable. A directional generalization of Hadamard (or compact) differentiability was developed for statistical applications in Shapiro90 and Duembgen93. There is a more recent growing literature of applications using Hadamard directional differentiability. See, for example, SommerfeldMunk17, ChetverikovSantosShaikh18, ChoWhite18, HongLi18, FangSantos19, ChristensenConnault19, MastenPoirier20, and CattaneoJanssonNagasawa20. We separately analyze $L_p$ functionals for $1 \leq p < \infty$ and the supremum norm, as well as one-sided variants used for tests that have a sense of direction. As an intermediate step towards showing the differentiability of supremum-norm statistics, we show directional differentiability of a minimax map applied to functions that are not necessarily convex-concave.
Directional differentiability of $L_p$ functionals provides the minimal amount of regularity needed for feasible statistical inference. We use directional differentiability to find the asymptotic distributions of test statistics estimated from sample data. The distributions generally depend on features of the objective function, and we show how to combine estimation with resampling to estimate the limiting distributions, as was proposed by FangSantos19. This technique can be tailored to impose a null hypothesis and is simple to implement in practice. We provide three examples throughout the paper to illustrate the proposed methods --- dependency bounds, stochastic dominance, and cost and profit functions. Moreover, we establish local size control of the resampling procedure in general and uniform size control for some functionals.
As a practical illustration of the developed framework, we consider bounds for the distribution function of a treatment effect. We use this model in both the Monte Carlo simulations and empirical application. Treatment effects models have provided a valuable analytic framework for the evaluation of causal effects in program evaluation studies. When the treatment effect is heterogeneous, its distribution function and related features, for example its quantile function, are often of interest. Nevertheless, it is well known (e.g. HeckmanSmithClements97, AbbringHeckman07) that features of the distribution of treatment effects beyond the average are not point identified unless one imposes significant theoretical restrictions on the joint distribution of potential outcomes. Therefore it has became common to use partial identification of the distribution of such features, and to derive bounds on their distribution without making any unwarranted assumptions about the copula that connects the marginal outcome distributions.\footnote{Bounds for the CDF and quantile functions at one point in the support of the treatment effects distribution were developed in Makarov82, Rueschendorf82, FrankNelsenSchweizer87, WilliamsonDowns90. These bounds are minimal in the sense that they rely only on the marginal distribution functions of the observed outcomes, not the joint distribution of (potential) outcomes.}
Monte Carlo simulations (reported in a supplementary appendix) assess the finite-sample properties of the proposed methods. The simulations suggest that Kolmogorov-Smirnov statistics used to construct uniform confidence bands have accurate empirical size and power against local alternatives. A Cram\'er-von Mises type statistic is used to test a stochastic dominance hypothesis using bound functions, similar to the breakdown frontier tests of MastenPoirier20. This test has accurate size at the breakdown frontier and power against alternatives that violate the dominance relationship. All results are improved when the sample size increases and are powerful using a modest number of bootstrap repetitions.
We illustrate the methods empirically using job training program data set first analyzed by LaLonde86 and subsequently by many others, including HeckmanHotz89, DehejiaWahba99, SmithTodd01, SmithTodd05, Imbens03, and Firpo07, without making any assumptions on dependence between potential outcomes in the experiment. This experimental data set has information from the National Supported Work Program used by DehejiaWahba99. We document strong heterogeneity of the job training treatment across the distribution of earnings. The uniform confidence bands for the treatment effect distribution function are imprecisely estimated at some parts of the earnings distribution. These large uniform confidence bands may in part be attributed to the large number of zeros contained in the data set, but are also inherent to the fact that the distribution function is everywhere only partially identified.
The remainder of the paper is organized as follows. Section (ref) defines the statistical model of interest for the value function. Section (ref) establishes directional differentiability for the functionals of interest. Inference procedures are established in Section (ref). An empirical application to job training is discussed in Section (ref). Finally, Section (ref) concludes. Monte Carlo simulation descriptions and results and all proofs are relegated to the supplementary appendix.
For any set $T \subseteq \mathbb{R}^d$ let $\ell^\infty(T)$ denote the space of bounded functions $f: T \rightarrow \mathbb{R}$ and let $\mathcal{C}(T)$ denote the space of continuous functions $f: T \rightarrow \mathbb{R}$, both equipped with the uniform norm $\| f \|_\infty = \sup_{t \in T} |f(t)|$. Given a sequence $\{f_n\}_n \subset \ell^\infty(T)$ and limiting random element $f$ we write $f_n \leadsto f$ to denote weak convergence in $(\ell^\infty(T), \| \cdot \|_\infty)$ in the sense of Hoffmann-J\o rgensen vanderVaartWellner96. Let $[x]_+ = \max\{x, 0\}$ and $[x]_- = \min\{x, 0\}$. A set-valued map (or correspondence) $S$ that maps elements of $X$ to the collection of subsets of $Y$ is denoted $S: X \rightrightarrows Y$. For set-valued map $S: X \rightrightarrows Y$, let $\textnormal{gr} S$ denote the graph of $S$ as a set in $X \times Y$.
Consider two sets $U \subseteq \mathbb{R}^{d_U}$ and $X \subseteq \mathbb{R}^{d_X}$, a set-valued map $A: X \rightrightarrows U$ that serves as a choice set, and an objective function $f \in \ell^\infty(\textnormal{gr} A)$. The objective function and the value function are linked by a marginal optimization step that maps one functional space to another. Let $\psi: \ell^\infty(\textnormal{gr} A) \rightarrow \ell^\infty(X)$ map the objective function to the value function obtained by marginally optimizing the objective $f$ with respect to $u \in A(x)$ for each $x \in X$. Without loss of generality, we consider only marginal maximization with respect to $u$ for each $x$:
The value function $\psi(f)$ is the object of interest and we would like to conduct uniform inference using a plug-in estimator $\psi(f_n)$. First we introduce examples of the value function model represented in equation (ref).
In Example (ref), the bounds on the quantile function are an example where $A(x)$ varies with $x$. It would be notationally simpler to consider optimization problems over the rectangle $\mathcal{X} \times \mathcal{U}$. However, given an $x$-varying set of constraints we would need to use “indicator functions” as they are used in the optimization literature, e.g. we would change the optimization problem to $\tilde{f}(u, x) = f(u, x) + \delta_A(u, x)$, where $\delta_A(u, x) = 0$ for $u \in A(x)$ and (for maximization) $\delta_A(u, x) = -\infty$ for $u \notin A(x)$. This takes us out of the realm of bounded functions and towards the consideration of epi- or hypographs, which may be more complicated. For example, see BuecherSegersVolgushev14. The literature on weak convergence in the space of bounded functions is voluminous and more familiar to researchers in economics.
To analyze the asymptotic properties of the above examples with plug-in estimates, one would typically rely on the delta method because $\psi$ is a nonlinear map from objective function to value function. In the next section we discuss the difficulties encountered with the delta method for value functions and a solution for the purposes of uniform inference.
We use the delta method to establish uniform inference methods for value functions. This depends on the notion of Hadamard (or compact) differentiability to linearize the map near the population distribution (see, for example, vanderVaartWellner96), and has the advantage of dividing the analysis into a deterministic part and a statistical part. We first consider Hadamard derivatives without regard to sample data, before explicitly considering the behavior of sample statistics in the next section.
The delta method between two metric spaces usually relies on (full) Hadamard differentiability vanderVaartWellner96. Shapiro90, Duembgen93 and FangSantos19 discuss Hadamard directional differentiability and show that this weaker notion also allows for the application of the delta method.\footnote{The derivative $\phi_f'$ in the definition is also called a semiderivative RockafellarWets98.}
In case the derivative map $\phi'_f$ is linear and $t_n$ may approach $0$ from both negative and positive sides, the map is fully Hadamard differentiable. The delta method and chain rule can be applied to maps that are either Hadamard differentiable or Hadamard directionally differentiable.
To discuss the differentiability of $\psi$ in equation (ref), for any $\epsilon \geq 0$ let $U_f: X \rightrightarrows U$ define the set-valued map of $\epsilon$-maximizers of $f(\cdot, x)$ in $u$ for each $x \in X$:
Results on the directional derivatives of the optimization map $f \mapsto \sup_X f$ date to Danskin67. For example, CarcamoCuevasRodriguez19 show that, in our notation, for fixed $x \in X$, $\psi$ is directionally differentiable in $\ell^\infty(A(x) \times \{x\})$, and for directions $h(\cdot, x)$
Assuming the stronger conditions that $A$ is continuous and compact-valued and that $f$ is continuous on the graph of $A$ ($\textnormal{gr} A$), the maximum theorem implies that $U_f$ is non-empty, compact-valued and upper hemicontinuous AliprantisBorder06, and CarcamoCuevasRodriguez19 show that tangentially to $\mathcal{C}(U \times \{x\})$, the derivative of $\psi(f)(x)$ simplifies to $\psi'_f(h)(x) = \sup_{u \in U_f(x, 0)} h(u, x)$. Further, when (for fixed $x$) $U_f(x, 0)$ is a singleton set, $\psi(f)(x)$ is fully differentiable, not just directionally so.
The pointwise directional differentiability of $\psi(f)(x)$ at each $x$ might lead one to suspect that $\psi$ is differentiable more generally as a map from $\ell^\infty(\textnormal{gr} A)$ to $\ell^\infty(X)$. However, this is not true. The following example illustrates a case where $\psi$ is not Hadamard directionally differentiable as a map from $\ell^\infty(\textnormal{gr} A)$ to $\ell^\infty(X)$, although $f$ and $h$ are both continuous functions.
The lack of uniformity of the convergence to $\psi_f'(h)(\cdot)$ appears to preclude the delta method for uniform inference, because the typical path of analysis would use a well-behaved limit of $r_n(\psi(f_n) - \psi(f))$ and the continuous mapping theorem to find the distribution of statistics applied to the limit. However, real-valued statistics are used to find uniform inference results, so we examine maps that include not only the marginal optimization step but also a functional applied to the resulting value function.
Consider $L_p$ norms (for $1 \leq p \leq \infty$) applied to $\psi(f)$ or $[\psi(f)]_+$. Letting $m$ denote the dominating measure, define $\lambda_j: \ell^\infty(\textnormal{gr} A) \rightarrow \mathbb{R}$ for $j = 1, \ldots, 4$ by
The maps $\lambda_2$ and $\lambda_4$ are included because one-sided comparisons may also be of interest, and the map $f \mapsto [f]_+$ is pointwise directionally differentiable but not differentiable as a map from $\ell^\infty(X)$ to $\ell^\infty(X)$. We illustrate its use in a stochastic dominance example below.
We make the following definitions in order to construct well-defined derivatives. First, let $\mu(f)$ be the supremum of $f$ over $\textnormal{gr} A$ and let $\sigma(f)$ be the maximin value of $f$ over $\textnormal{gr} A$:
Define $A_f^\mu \subseteq \textnormal{gr} A$ for any $\epsilon > 0$ to be the set of $\epsilon$-maximizers of $f$ in $(u, x)$ over $\textnormal{gr} A$. Next, let $X_f^\sigma \subseteq X$ collect the $\epsilon$-maximinimizers of $f$ in $x$ over $X$. That is, define
Also, note that $U_{(-f)}(x, \epsilon)$ is the set of $\epsilon$-minimizers of $f(u, x)$ in $u$ for given $x$.
CarcamoCuevasRodriguez19 show that the map (ref) has directional derivative
and Lemma (ref) in the \hyperref[appn]{Appendix} shows that (ref) has directional derivative
To analyze derivatives of the one-sided $\lambda_2(f)$ and $\lambda_4(f)$, define the contact set
The following two theorems describe the form of the four directional derivatives generally and under the condition that the value function is zero everywhere or on some subset of $X$.
The first part of Theorem (ref) essentially takes advantage of the identity $|x| = \max\{x, -x\}$, and breaks the evaluation of the supremum of the absolute value of the value function into two parts that are well-defined, a maximization and a minimization problem. The second part of the theorem is simpler because the $x \mapsto [x]_+$ map simplifies the evaluation of the supremum.
The next result shows that the $L_p$ functionals for $1 \leq p < \infty$ defined in (ref), are Hadamard directionally differentiable. The form of the derivatives for these maps also change depending on whether $\lambda_j(f) = 0$ or $\lambda_j(f) \neq 0$.
Theorem (ref) is related to the results of ChenFang19, who dealt with a squared $L_2$ statistic. It is interesting to note here that using an $L_2$ statistic instead of its square results in first-order (directional) differentiability of the map, unlike the squared $L_2$ norm, which has a degenerate first-order derivative but nondegenerate second-order derivative.
We conclude this section by revisiting Examples (ref), (ref), and (ref) and describing the computation of the directional derivatives in the corresponding cases.
In this section we derive asymptotic distributions for uniform statistics applied to value functions, and propose bootstrap estimators of their distributions that can be used for practical inference.
The derivatives developed in the previous subsections can be used in a functional delta method. We make a high-level assumption on the way that observations are related to estimated functions, but conditions under which such convergence holds are well-known, and examples will be shown below.
In case an $L_p$ statistic is used (with $p < \infty$), we restrict the space of allowable functions. In order to ensure that $L_p$ statistics are well-defined we make the assumption that the measure of $X$ is finite. While stronger than the assumption used for the supremum statistic, it is sufficient to ensure that the $L_p$ statistics of the value function are finite. This assumption also has the advantage of providing an easy-to-verify condition based on the objective function $f$, rather than the more direct, but less obvious condition $\psi(f) \in L_p(m)$. Lifting this restriction would require some other restriction on the objective function $f$ that is sufficient to ensure the value function is $p$-integrable.\footnote{We attempted to show $p$-integrability by assuming that $f$ is bounded and integrable (note that this would imply $f$ is $p$-integrable in $\textnormal{gr} A$), but were unable to show that this implies that the value function is integrable (in $X$). Note that if one were able to show that $f$ is bounded and integrable, then the elegant results in Kaji19 would apply for the purposes of verifying Assumption (ref).}
Given assumption (ref) and the derivatives of the last section, the asymptotic distribution of test statistics applied to the value function is straightforward.
This theorem is abstract and hides the fact that the limiting distributions may depend on features of $\lambda_j$ and $f$. Therefore it is only indirectly useful for inference. A resampling scheme is the subject of the next section.
The delta method as used in Theorem (ref) is tailored to arguments that have weak limits in the space $\ell^\infty(\textnormal{gr} A)$, which is an intrinsic feature of the delta method. That implies some limitations on the type of functions that may be considered. For example, functions estimated using kernel methods do not converge to a weak limit as assumed in Assumption (ref). One can in general only show the stochastic order of the sequence $r_n(f_n - f)$, and methods for uniform inference must shown using something other than the delta method. We leave uniform inference for value functions based on less regular estimators to future research.
Our bootstrap technique ahead is designed for resampling under the null hypothesis that $\lambda_j(f) = 0$. We assume that when $j = 1$ or $j = 3$ (that is, conventional two-sided statistics are used), the statistics are used to test the null and alternative hypotheses
Meanwhile, when $j = 2$ or $j = 4$ (the one-sided cases) we assume the hypotheses are
The probability measure $P$ is assumed to belong to a collection $\mathcal{P}$ which describes the set of probability measures allowed by the model. A few subcollections of $\mathcal{P}$ serve to organize the asymptotic results below. When $P$ is such that $\psi(f) \equiv 0$, we label $P \in \mathcal{P}_{00}^E$. For functional inequalities, the behavior of test statistics under the null is more complicated. If $P$ is such that $\psi(f) \leq 0$ everywhere, then label $P \in \mathcal{P}_0^I$. When $P \in \mathcal{P}_0^I$ makes $\psi(f)(x) = 0$ for at least one $x \in X$, then we label $P \in \mathcal{P}_{00}^I$.
The following corollary combines the derivatives from Theorem (ref) and Theorem (ref) with the result of Theorem (ref) for these distribution classes. Recall that the set $X_0$, defined in (ref), represents the subset of $X$ where $\psi(f)$ is zero.
The distributions described in the above corollary under the assumption that $P \in \mathcal{P}_{00}^E$ or $\mathcal{P}_{00}^I$ are those that we emulate using resampling methods in the next section.
In the previous section we established asymptotic distributions for the test statistics of interest. However, for practical inference we turn to resampling to estimate their distributions. This section suggests the use of a resampling strategy that was proposed by FangSantos19, and it combines a standard bootstrap procedure with estimates of the directional derivatives $\lambda_{jf}'$. Hadamard directional differentiability also implies that the numerical approximation methods of HongLi18 may be used to estimate these directional derivatives, an approach that we address in simulations (in the supplementary appendix).
The resampling scheme described below is designed to reproduce the null distribution under the assumption that the statistic is equal to zero, or in other terms, that $P \in \mathcal{P}_{00}^E$ or $P \in \mathcal{P}_{00}^I$. This is achieved by restricting the form of the estimates $\hat{\lambda}_{jn}'$. The bootstrap routine is consistent under more general conditions, as in Theorem 3.1 of FangSantos19. However, we discuss behavior of bootstrap-based tests under the assumption that one of the null conditions described in the previous section holds.
All the derivative formulas in Corollary (ref) require some form of an estimate of the near-maximizers of $f$ in $u$, that is, of the set $U_f(x, \epsilon)$ defined in (ref) for various $x$. The set estimators we use are similar to those used in LintonSongWhang10, ChernozhukovLeeRosen13 and LeeSongWhang18 and depend on slowly decreasing sequences of constants to estimate the relevant sets in the derivative formulas. For a sequence $a_n \rightarrow 0^+$ we estimate $U_f(x, \epsilon)$ with the plug-in estimator $U_{f_n}(x, a_n)$. The sequence $a_n$ should decrease more slowly than the rate at which $f_n$ converges uniformly to $f$, an assumption which will be formalized below.
The one-sided estimates $\hat{\lambda}_{2n}'$ and $\hat{\lambda}_{4n}'$ use a second estimate. For another sequence $b_n \rightarrow 0^+$, estimate the contact set $X_0$ defined in (ref) with
This is a method used by LintonSongWhang10 of achieving uniformity in the convergence of estimators that depend on contact set estimation under the null hypothesis $P \in \mathcal{P}_{00}^I$.
Define estimators of $\lambda_{1f}'$ and $\lambda_{2f}'$ by
and
These formulas impose the condition that $P \in \mathcal{P}_{00}^E$ or $\mathcal{P}_{00}^I$ on the derivative estimates, following the forms shown in Corollary (ref). For more general estimation we would need to devise more complex estimators to deal with the variety of limits discussed in Theorems (ref) and (ref), and also the technique of HongLi18 could be used in that case. However, inference is usually of primary interest, and generating a bootstrap distribution that respects the null hypothesis can improve the performance of bootstrap inference procedures. When $P \in \mathcal{P}_{00}^E$ the imposition amounts to the assumption that $\psi(f)(x) = 0$ for each $x$, so that the set of population $\epsilon$-maximizers in $\textnormal{gr} A$ is $U_f(X, \epsilon)$, and the estimator $\hat{\lambda}_{1n}'$ emulates that. The condition $P \in \mathcal{P}_{00}^I$ implies that $X_0$ is not empty, and $\hat{X}_0$ is meant to estimate this contact set, while using all of $X$ when the distribution is not as is hypothesized.
The estimators $\hat{\lambda}_{3n}'$ and $\hat{\lambda}_{4n}'$ are defined similarly to the supremum-norm estimators: let
and
We use an exchangeable bootstrap, which depends on a set of weights $\{W_i\}_{i=1}^n$ that put probability mass $W_i$ at each observation $Z_i$. This type of bootstrap describes many well-known bootstrap techniques including sampling with replacement from the observations vanderVaartWellner96.
We make the following assumptions to ensure the bootstrap is well-behaved.
We remark that it is difficult to propose a general method for choosing how to set $a_n$ and $b_n$ in practice. We numerically investigated their effect in simulations. We apply similar rules to those used in LintonSongWhang10 and LeeSongWhang18. However, in general situations, where the variance of the objective function changes significantly over its domain, one would need to estimate the pointwise standard deviation of the function. For example, FanPark10 provide pointwise asymptotic distributions for estimated dependency bound functions for $F_{X-Y}$. Nevertheless, we leave this extension for future research.
Assumption (ref) is used to ensure consistency of the $\epsilon$-maximizer set estimators used in the bootstrap algorithm. Assumption (ref) is a condensed version of Assumption 3 of FangSantos19 to ensure that $r_n(f_n^* - f_n)$ behaves asymptotically like $r_n(f_n - f)$.
Resampling routine to estimate the distribution of $r_n(\lambda_j(f_n) - \lambda_j(f))$
Then repeat steps (ref)-(ref) for $r = 1, \ldots R$:
Finally,
The consistency of this resampling procedure under the null hypothesis is summarized in the following theorem. To discuss weak convergence it is easiest to use the space of bounded Lipschitz functions, which indicates weak convergence. That is, defining $BL_1(\mathbb{R}) = \{g \in C_b(\mathbb{R}) : \sup_x |g(x)| \leq 1, \sup_{x \neq y} |g(x) - g(y)| \leq |x - y|\}$, $X_n$ converges weakly to $X$ if and only if $\sup_{g \in BL_1(\mathbb{R})} |\textnormal{E} \left[ g(X_n) \right] - \textnormal{E} \left[ g(x) \right]| \rightarrow 0$ vanderVaartWellner96.
We conclude this section by revisiting Examples (ref), (ref), and (ref) and describing the resampling procedures in the corresponding cases.
It is also of interest to examine how these tests behave under sequences of distributions local to distributions that satisfy the null hypothesis. We consider sequences of local alternative distributions $\{P_n\}$ such that for each $n$, $\{Z_i\}_{i=1}^n$ are distributed according to $P_n$, and $P_n$ converges towards a limit $P$ that satisfies the null hypothesis. To describe this process, for $t \geq 0$ define a path $t \mapsto P_t$, where $P_t$ is an element of the space of distribution functions $\mathcal{P}$, such that
where the score function $h \in L_2(F)$ satisfies $\textnormal{E} \left[ h \right] = 0$ and $P_0$ is a distribution that satisfies the null hypothesis. The direction that the sequence approaches the null is described asymptotically by the score $h$. We assume that by letting $t = c / r_n$ for $c \in \mathbb{R}$, we can parameterize distributions that are local to $P_0$ and for $t \geq 0$ denote $f(P_t)$ as the function $f$ under distribution $f(P_t)$ so that the unmarked $f$ described above can be rewritten $f = f(P_0)$. See, e.g., vanderVaartWellner96 for more details.
The following assumption ensures that $f$ remains suitably regular under such local perturbations to the null distribution.
Both parts of Assumption (ref) ensure that $f_n$ behave regularly as distributions drift towards $P_0 \in \mathcal{P}_{00}^E$ or $P_0 \in \mathcal{P}_{00}^I$. These additional regularity conditions allow us to describe the size and local power properties of the test statistics.
Theorem (ref) shows that the size of tests can be controlled locally to the null region (and the nominal rejection probability matches the intended probability) in only some cases. In particular, the tests for the null that $P_0 \in \mathcal{P}_{00}^E$ cannot be shown to control local size without more information about the direction of the local alternatives. This is because we must make assumptions about the underlying objective functions without being able to make similar assumptions about the corresponding value functions --- note that Assumption (ref) is made with regard to $f(P)$, not $\psi(f(P))$. To see where the problem lies we may take $\lambda_1$ as an example. The result of Theorem (ref) shows that when $P_n \in \mathcal{P}_{00}^E$ for each $n$,
Although $P \in \mathcal{P}_{00}^E$ implies $\lambda_f'(f'(c)) = 0$, as shown in the proof of Theorem (ref), it does not imply that $f'(c) \equiv 0$. For example, if $f$ are CDFs corresponding to a location-shift family of distributions, i.e., for all $c \in \mathbb{R}$, $G_{c}(x) = G_0(x - c)$ with differentiable densities that have square-integrable derivatives. Then since the $h$ in (ref) is $cg'_0(x) / g_0(x)$, the $f'(c)$ of assumption (ref) would be $-cg_0(x)$. This non-zero derivative can cause problems when optimizing it interacts with the absolute value function. Generally $|\sup_u (f(u) + g(u))| \not\leq |\sup_u f(u)| + |\sup_u g(u)|$ (take $f$ and $g$ strictly less than zero), and we do not achieve a simplification that would imply the probability on the right-hand side is less than or equal to $\alpha$. We cannot assume regularity of $\psi(f)$ because this would require the existence of a derivative $\psi_f'(\cdot) \in \ell^\infty(X)$, but this derivative cannot generally exist. It may be that size control could be shown for special cases (for example when theory implies a shape for the value function), but not in general. On the other hand, for the same reason, the size of one-sided test statistics is as intended. That is, when the value function is found through maximization, the positive-part map in the definitions of $\lambda_2$ and $\lambda_4$ interacts with it in a way that maintains size control. In contrast, the negative-part map would not guarantee size control. A stronger result in this vein will be shown in the next section.
The inequality in the second part of this theorem results from the one-sidedness of the test statistics $\lambda_2$ and $\lambda_4$. This is related to a literature in econometrics on moment inequality testing. Tests may exhibit size that is lower than nominal for local alternatives that are from the interior of the null region. A few possible solutions to this problem have been proposed. For example, one might evaluate the region where the moments appear to hold with equality, which leads to contact set estimates like in the bootstrap routine described above LintonSongWhang10. Alternatively, we may alter the reference distribution by shifting it in the regions where equality does not seem to hold AndrewsShi17.
Uniformity of inference procedures in the data distribution was introduced in GineZinn91 and SheehyWellner92, and uniformity has become a topic of great interest in the econometrics literature (see, for example, AndrewsGuggenberger10, LintonSongWhang10, RomanoShaikh12 or HongLi18 for an application like ours). Under stronger assumptions that ensure the underlying $f_n$ converge weakly to a limiting process uniformly over the set of possible data distributions, some of the above results can be extended from size control under local alternative distributions to size control that holds uniformly over the class of distributions that satisfy the null. To ensure uniformity we assume the following regularity conditions are satisfied.
The above assumptions help define the collection $\mathcal{P}$ on which we may assert that inference is valid uniformly for $P \in \mathcal{P}$. Assumptions (ref) and (ref) require that the sample objective functions converge to well-behaved limits uniformly on $\mathcal{P}$. Assumption (ref) requires that also the limiting distributions of the sample statistics $r_n(\lambda_j(f_n) - \lambda_j(f))$ are regular enough over $\mathcal{P}_0^I$ that their CDFs do not have extremely large atoms (i.e., extending above the $(1-\alpha)$-th quantile of the distribution) and are strictly increasing at the relevant critical value. This assumption may be simplified when the $\mathcal{G}_P$ are Gaussian since the functionals $\lambda_2$ and $\lambda_4$ are convex DavydovLifshitsSmorodina98. Assumption (ref) requires that the bootstrap objective functions also converge uniformly, as the original sample functions do. This last assumption may be implied by Assumption (ref), see Lemma A.2 of LintonSongWhang10, for example.
Uniformity for $P \in \mathcal{P}$ is maintained by the one-sided statistics when optimization matches the “direction” of the test. That is, when the value function matches the positive-part map in $\lambda_2$ and $\lambda_4$, we have uniformity over $P \in \mathcal{P}$, and analogously uniform inference would be possible for minimization and the negative-part map. Uniformity can be lost when optimization and the direction of the test do not match, which may be the case with two-sided tests.
In the next section we illustrate the usefulness of our results by providing details on the construction of uniform confidence bands around bound functions for the CDF of a treatment effect distribution. See the first section of the supplemental appendix for numerical simulations evaluating the finite-sample performance of the tests developed in Examples (ref) and (ref).
This section illustrates the inference methods with an evaluation of a job training program. We construct upper and lower bounds for both the distribution and quantile function of the treatment effects, confidence bands for these bound function estimates, and describe a few inference results. This application uses an experimental job training program data set from the National Supported Work (NSW) Program, which was first analyzed by LaLonde86 and later by many others, including HeckmanHotz89, DehejiaWahba99, SmithTodd01, SmithTodd05, Imbens03, and Firpo07.
Recent studies in statistical inference for features of the treatment effects distribution in the presence of partial identification include, among others, FirpoRidder08, FirpoRidder19, FanPark09, FanPark10, FanWu10, FanPark12, FanShermanShum14, FanGuerreZhu17. These studies have concentrated on distributions of finite-dimensional functionals of the distribution and quantile functions, including these functions themselves evaluated at a point. Additional work includes, among others, GautierHoderlein11, ChernozhukovLeeRosen13, Kim14 and ChesherRosen15. Each of these papers provides pointwise inference methods for bounds on the distribution or quantile functions, and often for more complex objects. We contribute to this literature by applying the general results in this paper to provide uniform inference methods for the bounds developed by Makarov and others, and hope that it may indicate the direction that pointwise inference for bounds in more involved models may be extended to be uniformly valid.
The data set we use is described in detail in LaLonde86. We use the publicly available subset of the NSW study used by DehejiaWahba99. The program was designed as an experiment where applicants were randomly assigned into treatment. The treatment was work experience in a wide range of possible activities, such as learning to operating a restaurant or a child care center, for a period not exceeding 12 months. Eligible participants were targeted from recipients of Aid to Families With Dependent Children, former addicts, former offenders, and young school dropouts. The NSW data set consists of information on earnings and employment (outcome variables), whether treated or not.\footnote{The data set also contains background characteristics, such as education, ethnicity, age, and employment variables before treatment. Nevertheless, since we only use the experimental part of the data we refrain from using this portion of the data.} We consider male workers only and focus on earnings in 1978 as the outcome variable of interest. There are a total of 445 observations, where 260 are control observations and 185 are treatment observations. Summary statistics for the two parts of the data are presented in Table (ref).
To provide a more complete overview of the data, we also compute the empirical CDFs of the treatment and control groups in Figure (ref). From this figure we note that the empirical treatment CDF stochastically dominates the empirical control CDF, and that there are a large number of zeros in each sample. In particular, $\mathbb{F}_{tr,n}(0) \approx 0.24$ and $\mathbb{F}_{co,n}(0) \approx 0.35$.
Suppose a binary treatment is independent of two potential outcomes $(X_{co}, X_{tr})$, where $X_{co}$ denotes outcomes under a control regime and $X_{tr}$ denotes outcomes under a treatment, and $X_{co}$ and $X_{tr}$ have marginal distribution functions $F_{co}$ and $F_{tr}$ respectively. Suppose that interest is in the distribution of the treatment effect $\Delta = X_{tr} - X_{co}$ but we are unwilling to make any assumptions regarding the dependence between $X_{co}$ and $X_{tr}$. In this section we study the relationship between the identifiable functions $F_{co}$ and $F_{tr}$ and functions that bound the distribution function of the unobservable random variable $\Delta$.
$F_\Delta(\cdot)$ is not point-identified because the full bivariate distribution of $(X_{co}, X_{tr})$ is unidentified and the analyst has no knowledge of the dependence between potential outcomes. However, $F_\Delta$ can be bounded. Suppose we observe samples $\{X_{ki}\}_{i=1}^{n_k}$ for $k \in \{co, tr\}$. Define the empirical lower and upper bound functions
These are plug-in estimates of analogous population bounds $L_\Delta$ and $U_\Delta$, and are similar to Example (ref) because $F_\Delta = F_{(X_{tr}-X_{co})}$. We make the following assumptions on the observed samples of treatment and control observations.
Under Assumptions (ref) and (ref), it is a standard result vanderVaart98 that for $k \in \{co, tr\}$, $\sqrt{n_k} (\mathbb{F}_{kn} - F_k) \leadsto \mathcal{G}_k$, where $\mathcal{G}_{co}$ and $\mathcal{G}_{tr}$ are independent $F_{co}$- and $F_{tr}$-Brownian bridges, that is, mean-zero Gaussian processes with covariance functions $\rho_k(x, y) = F_k(x \wedge y) - F_k(x) F_k(y)$. This implies in turn that \( \sqrt{n} (\mathbb{F}_n - F) \leadsto \mathcal{G}_F = ( \mathcal{G}_{co} / \sqrt{\nu_{co}}, \mathcal{G}_{tr} / \sqrt{\nu_{tr}} )\), where $\mathcal{G}_F$ is a mean-zero Gaussian process with covariance process $\rho_F(x, y) = \text{diag}\{\rho_k(x, y) / \nu_k\}$.
Now we focus on the calculation of a uniform confidence band for $L$ only. Because the bootstrap algorithm was described previously we only verify that the regularity conditions for this plan hold. Assumptions (ref) and (ref), along with the above discussion of the weak convergence of $\sqrt{n}(\mathbb{F}_n - F)$ imply that Assumption (ref) is satisfied. The class of functions $\{ (I(X \leq x), I(Y \leq y), x,y \in \mathbb{R} \}$ is uniform Donsker and the sequences $a_n$ and $b_n$ can be chosen to satisfy Assumption (ref) by making them converge to zero more slowly than $n^{-1/2}$. Assuming the weights are independent of the observations, assumption (ref) is satisfied by Lemma A.2 of LintonSongWhang10, which implies that the bootstrap algorithm described above is consistent, as described in Theorem (ref). It is also straightforward to verify that, under the high-level assumption (ref), both parts of Assumption (ref) are satisfied --- in the language of the assumption, $f(P_{c/\sqrt{n}})$ are the pair $(F_{co}^n, F_{tr}^n)$ under the local probability distribution $P_{c/\sqrt{n}}$, and $f' = (\int_{-\infty}^\cdot h_{co}^c, \int_{-\infty}^\cdot h_{tr}^c)$) for direction $h^c$ indexed by $c$ and $\sqrt{n}(\mathbb{F}_n - F^n) \leadsto \mathcal{G}_F$ vanderVaartWellner96.
The main objective is to provide uniform confidence bands for the CDF of the treatment effects distribution. We calculate the lower and upper bounds for the distribution function using the control and treatment samples as in equations (ref)--(ref). The bounds are computed on a grid with increments of \$100 dollars along the range of the common support of the bounds, which is roughly from -\$40,000 to \$60,000. We focus on the region between -\$40,000 and \$40,000, which contains almost all the observations (there is a single \$60K observation in the treated sample). The results are presented in Figure (ref) and given by the black solid lines in the picture. The most prominent feature is that, as expected, the upper bound for the CDF of treatment effects stochastically dominates the corresponding lower bound.
Next, we compute the uniform confidence bands as described in the text. They are shown in Figure (ref) as the dashed lines around the corresponding solid lines.
Due to the large number of zero outcomes in both samples, these bounds have some interesting features that we discuss further. First, we note that for any $\epsilon$ greater than zero but smaller than the next smallest outcome (about \$45), $\mathbb{F}_{tr,n}(0) - \mathbb{F}_{co,n}(-\epsilon) = 0.24$, which explains the jump in the lower bound estimate near zero (it is really for a point in the grid just above zero). Likewise, for the same $\epsilon$, $\mathbb{F}_{tr,n}(-\epsilon) - \mathbb{F}_{co,n}(0) = -0.35$, which explains the jump in the upper bound just below zero. Without these point masses at zero, both bounds would more smoothly tend towards 0 or 1. Second, the point masses at zero imply another feature of the bounds that can be discerned in the picture. The upper bound to the left of 0 is the same as $1 - \mathbb{F}_{co,n}(-x)$ and the lower bound to the right of zero is the same as $\mathbb{F}_{tr,n}(x)$. Taking the lower bound as an example, for each $x > 0$ find the closest observation from the control sample $y_{co,i^*}$, and set $x^*(x) = X_{co,i^*} + \epsilon$, leading to the supremum $\mathbb{F}_{tr,n}(X_{co,i} + \epsilon)$ at every point where $X_{co,i} + \epsilon < X_{tr,j}$ for all $j$ in the treated sample. It is identical to the empirical treatment CDF for the entire positive part of the support because the treatment first-order stochastically dominates the control. The situation would be different if there were a jump in the empirical control CDF at least as large as the jump in the empirical treatment CDF at zero. Because the opposite is the case for the upper bound, it does change slightly above the zero mark, tending from $1 - \mathbb{F}_{tr,n}(0) + \mathbb{F}_{co,n}(0)$ to 1 as $x$ goes from 0 to the right.
We can also use the confidence bands for the bound functions to construct a confidence band for the true distribution function of treatment effects. This is shown in Figure (ref). A $1 - \alpha$ confidence band can be constructed by using the upper $1 - \alpha / 2$ limit of the upper bound confidence band and the lower $\alpha / 2$ limit of the lower confidence band. This band is a uniform asymptotic confidence band for the true CDF, and uniform over correlation between the potential outcomes between samples. In other words, if $\mathcal{P}$ is the collection of bivariate distributions that have marginal distributions $F_{tr}$ and $F_{co}$, then
This confidence band is likely conservative, since
See Kim14 and FirpoRidder19 for a more thorough discussion of the sense in which these bounds are not uniformly sharp for the treatment effect distribution function. We leave more sophisticated, potentially tighter confidence bands for future research. Note that the technique of ImbensManski04 cannot be used to tighten these bounds, because the parameter, a function, could violate the null hypothesis at both sides of the confidence band.
We plot some other features in this figure for context. First, the dotted vertical line is positioned at $y = 0$, and it can be seen that we (just) reject the null hypothesis $H_0: P \left\{ \Delta = 0 \right\} = 1$. This hypothesis is closest to non-rejection, and it is clearer that one should reject the null that the treatment effect distribution is degenerate at any other point besides zero. This supports the notion that treatment effect heterogeneity is an important feature of these observations, especially because this band is completely agnostic about the form of the joint distribution. On the other hand, by examining the bands at horizontal levels, it can be seen that for the median effect and a wide interval in the center of the distribution, the hypothesis of zero treatment effect cannot be rejected (although these are uniform bounds and not tests of individual quantile levels).
The final feature in the figure is the dashed curve that represents the estimate that one would make under the assumption of comonotonicity (or rank invariance) --- the assumption that, had an individual been moved from the treatment to the control group, their rank in the control would be the same as their observed rank. Under this strong assumption the quantile treatment effects are the quantiles of the treatment distribution and they can be inverted to make an estimate. Clearly, the estimate under this assumption is just one point-identified treatment effect distribution function of many.
Finally, to provide context for the uniform confidence band results within the literature on inference for bounds, we compare the proposed uniform bands to the pointwise confidence intervals suggested by FanPark10. We used BickelSakov08's automatic procedure to choose subsample size and constructed confidence intervals for each individual point in the grid of the treatment effect support. This collection of pointwise confidence intervals are plotted along with uniform confidence bands in Figure (ref). The results show that the uniform bands are farther from the bound estimates than the set of pointwise confidence intervals.
This paper develops uniform statistical inference methods for optimal value functions, that is, functions constructed by optimizing an objective function in one argument over all values of the remaining arguments. Value functions can be seen as a nonlinear transformation of the objective function. The map from objective function to value function is not Hadamard differentiable, but statistics used to conduct uniform inference are Hadamard directionally differentiable. We establish the asymptotic properties of nonparametric plug-in estimators of these uniform test statistics and develop a resampling technique to conduct practical inference. Examples involving dependency bounds for treatment effects distributions are used for illustration. Finally, we provide an application to the evaluation of a job training program, estimating a conservative uniform confidence band for the distribution function of the program's treatment effect without making any assumptions about the dependence between potential outcomes.
We would like to thank Timothy Christensen, Zheng Fang, Hidehiko Ichimura, Lynda Khalaf, Alexandre Poirier, Pedro H.C. Sant'Anna, Andres Santos, Rami Tabri, Brennan Thompson, Tiemen Woutersen, and seminar participants at the University of Kentucky and New York Camp Econometrics XIV for helpful comments and discussions. We would also like to thank the referees who helped improve the paper, especially the comments of one referee that improved subsection 4.4 considerably. All the remaining errors are ours. This research was enabled in part by support from the Shared Hierarchical Academic Research Computing Network (SHARCNET: \url{https://www.sharcnet.ca}) and Compute Canada (\url{http://www.computecanada.ca}). Code used to implement the methods described in this paper and to replicate the simulations can be found at \url{https://github.com/tmparker/uniform_inference_for_value_functions}.