EconBase
← Back to paper

Uniform Confidence Band for Marginal Treatment Effect Function

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,015 characters · 14 sections · 62 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Uniform Confidence Band for Marginal Treatment Effect Function

abstractThis paper presents a method for constructing uniform confidence bands for the marginal treatment effect (MTE) function. The shape of the MTE function offers insight into how the unobserved propensity to receive treatment is related to the treatment effect. Our approach visualizes the statistical uncertainty of an estimated function, facilitating inferences about the function's shape. The proposed method is computationally inexpensive and requires only minimal information: sample size, standard errors, kernel function, and bandwidth. This minimal data requirement enables applications to both new analyses and published results without access to original data. We derive a Gaussian approximation for a local quadratic estimator and consider the approximation of the distribution of its supremum in polynomial order. Monte Carlo simulations demonstrate that our bands provide the desired coverage and are less conservative than those based on the Gumbel approximation. An empirical illustration regarding the returns to education is included. Keywords: uniform confidence band; marginal treatment effect; instrumental variables, empirical process; Gaussian approximation. JEL Classification: C12, C14, C21, C26.
refsection

Introduction

The marginal treatment effect (MTE) provides information on how the likelihood of receiving treatment relates to the treatment effects heckman1999local, heckman2005structural. We consider settings where unobserved factors that influence selection into treatment are correlated with the potential outcomes, and instrumental variables (IVs) are available. IVs affect the selection into treatment without directly affecting the outcomes. MTE represents the benefits to individuals on the cusp of deciding whether or not to participate in treatment. The shape of the MTE function provides insight into how the unobservable propensity is related to the treatment effect. For instance, when the MTE function decreases, those with a high propensity to receive treatment experience a relatively large treatment effect compared to those with low propensities. Conversely, an increasing MTE function implies that those with a low propensity to receive treatment exhibit a relatively large treatment effect compared to those with high propensities.

MTE functions have been examined in various empirical applications. For instance, carneiro2011marginalretrurn estimate the MTE of college attendance and find that the MTE function is decreasing. This result suggests that individuals who benefit from college education are more likely to attend college. YamaguchiAsaiKambayashi18 estimate the MTE function for the effect of early childcare. Their estimated MTE function is increasing. It suggests that “children who would benefit most from childcare are less likely to attend, implying inefficient allocation” YamaguchiAsaiKambayashi18. Other empirical examples of MTE include doyle2007child, kasahara2016does, and, brinch2017beyond. shea2023ivmte provide a wide list of empirical articles that use MTE.

This paper proposes a method for constructing uniform confidence bands for the MTE function. A uniform confidence band covers the true function with a prespecified probability. These bands are used to make statistical inferences on the shape of the MTE function. Uniform confidence bands are particularly useful because they enable the visualization of the statistical uncertainty underlying the estimated MTE function. All the above empirical articles present pointwise confidence intervals using the bootstrap method. A pointwise confidence interval quantifies the statistical uncertainty at a point of the MTE function, not the function itself. Pointwise procedures are not valid when we would like to evaluate the statistical uncertainty regarding the shape of the MTE function. For this purpose, uniform confidence bands are needed.

Our proposed method is straightforward to implement. The uniform confidence band takes a familiar form: "estimated function $\pm$ critical value $\times$ standard error." This format makes our confidence bands easy to visualize. The critical values can be obtained analytically, and our approach does not rely on computationally intensive procedures, such as bootstrapping, making it computationally efficient.

The proposed method applies not only to new analyses but also to published results without access to the original data. The computation of our critical value requires only the sample size, the kernel function, and the bandwidth used to estimate the MTE function. This information is typically available in published articles, allowing researchers to expand the pointwise confidence bands reported in existing studies using our critical value to create uniform confidence bands. This applicability is not available with bootstrap methods, which require access to the original data.

Our construction of uniform confidence bands is based on an asymptotic approximation of the supremum of the estimated MTE function over its domain of interest. We follow heckman2005structural for the setting. The semiparametric estimation of the finite-dimensional parameters is based on carneiro2009estimating. The local quadratic estimator estimates the MTE function as in heckman2006understanding. We approximate the supremum of a normalized version of the estimated function by the supremum of the Gaussian process using the result of chernozhukov2012gaussian. Then, we further approximate the Gaussian field by a stationary Gaussian field by the result of ghosal2000testing. We derive critical values using an analytic approximation of the distribution of the supremum of a stationary Gaussian process based on the arguments of piterbarg1996asymptotic. The theoretical analysis follows steps similar to those of lee2017doubly. However, our derivation necessitates non-trivial technical analysis due to the estimation of the MTE function as the first derivative of the nonparametric regression function. Additionally, we must carefully address the estimation error in the estimated propensity score. Overcoming this generated regressor problem constitutes our key technical contribution.

This new approximation is more precise than the existing one based on the Gumbel approximation. It is well-known that the supremum of a Gaussian process follows the Gumbel distribution asymptotically cramer67; however, the Gumbel approximation is only accurate up to a logarithmic order. This limitation can lead to overly conservative statistical inferences. In contrast, our approximation is valid up to a polynomial order, which provides more accurate results. It also offers tighter confidence bands than those derived from the Gumbel approximation.

We examine finite sample properties of the proposed uniform confidence band through Monte Carlo simulations. The results demonstrate that our confidence bands cover the true function with probability similar to the prespecified confidence level. Moreover, the simulation results show the superiority over alternative methods. Pointwise confidence bands yield much lower coverage, indicating the inadequacy of using them to make inferences regarding the shape of the MTE function. We also find that while uniform confidence bands are wider than pointwise confidence bands, they are sufficiently informative. The critical value for a 95% confidence band is 1.96 for the pointwise procedure. It is around 3 for our procedure. The uniform confidence bands based on the Gumbel approximation are much broader and may not be informative. The 95% critical value from the Gumbel approximation is around 4.

We demonstrate the application of our proposed procedure through an empirical study. Specifically, we examine the impact of university enrollment on future income Card95. While the estimated MTE function shows some fluctuations, it generally trends downward. Our confidence intervals do not contain a constant or monotonically increasing function, leading us to reject the null hypothesis that the MTE function is weakly increasing. Conversely, we cannot reject the null hypothesis that the MTE function is decreasing. Notably, while our uniform confidence band is wider than the pointwise confidence band, both bands arrive at the same conclusion regarding the shape of the MTE function. The band based on the Gumbel approximation is significantly wider and includes weakly increasing functions. These results suggest that while our method yields a broader confidence band, it still provides informative results.

The remainder of the paper is organized as follows. In the following subsection, we discuss the relation to the literature. In Section (ref), we describe the econometric model and discuss the identification of the MTE function. In Section (ref), we present a semiparametric estimation method for the MTE function. In Section (ref), we construct the uniform confidence band for the MTE function. Section (ref) provides the asymptotic justification of the proposed procedure. Section (ref) demonstrates our uniform confidence band's finite sample performances compared with other methods via Monte Carlo simulations. Section (ref) illustrates our method with an empirical application. Section (ref) concludes. The technical definition of VC-type classes is in Appendix (ref). Mathematical proofs are collected in the supplemental materials.

Related literature

This study contributes to the literature on the IV approach for causal inference. When evaluating the impacts of programs, researchers take into account the fact that responses to the treatments vary across individuals. The early work in the literature on heterogeneous treatment effects in the IV framework by ImbensAngrist94 allows for heterogeneity in both the outcome response to the treatment and the treatment selection response to the instruments. They clarify the interpretation of IV estimates as a local average treatment effect, which captures the average gain from the treatment for the compliers. heckman2005structural and heckman2006understanding generalize the marginal treatment effect (MTE) introduced by bjorklund1987estimation. They define the MTE as the mean gain from the treatment for marginal individuals entering the treatment (i.e., those who shift into or out of treatment by a marginal change in the propensity score). The MTE function is useful for causal inference for the following two reasons cornelissen2016from. First, the MTE function provides information about the relationship between the treatment effect and the unobserved propensity to obtain the treatment. Second, many conventional treatment effect parameters, such as the average treatment effect, can be expressed as different weighted averages of MTE. Our proposed method is useful for the first usage of the MTE function.

As discussed in the introduction, the shape of the MTE function is important in empirical applications. Many empirical papers include pointwise confidence bands when plotting the estimated MTE function carneiro2009estimating. This paper presents a straightforward method for computing uniform confidence bands, enabling the statistical inference of the shape of the MTE function. We also note that the shape of the MTE function is important not only for policy evaluation but also for optimal policy assignment, as discussed in chen2022personalized and liu2022policy.

This paper contributes to the literature on uniform confidence bands for functions estimated nonparametrically. Recent works on uniform confidence bands for nonparametric models include horowitz2012uniform, horowitz2017nonparametric, chen2018optimal, and chen2025adaptive, among others. However, few researchers have studied the uniform confidence band for derivatives of a function. chen2018optimal and chen2025adaptive establish a bootstrap uniform confidence band for a function and its derivatives. We mainly consider the confidence band for a function parameter estimated by a partial derivative with a local polynomial regression.

In the causal inference literature, a uniform confidence band is considered mainly for conditional average treatment effects functions. See, for example, lee2017doubly, semenova2021debiased, fan2022estimation, and baybutt2023doubly. We largely follow the approach taken by lee2017doubly. There are two critical differences. First, our target function is not a conditional mean function, but rather its derivative, which complicates the theoretical analysis. Second, the argument of our target function is the propensity score, which needs to be estimated; we must therefore handle the estimation error in the propensity score. This second difference poses technical challenges and constitutes our technical contribution.

Notation

Let $: = $ denote “equals by definition,” and let a.s. denote “almost surely.” Let $\mathbbm{1}\{\cdot\}$ denote the indicator function. For a random variable $X$, $f_{X}(\cdot)$ denotes the probability density function of $X$. $\|f\|_{\infty}$ denotes $\sup_{x\in\mathcal{X}}|f(x)|$, where $\mathcal{X}$ is the support of $f$. For an arbitrary set $T$, let $\ell^{\infty}(T)$ denote the class of functions $f$ that maps from $T$ to $\mathbb{R}$ such that $\|f\|_{\infty}<\infty$. Unless otherwise stated, $c>0$ refers to universal constants whose values may change from place to place.

Framework

This section presents the econometric framework and defines and identifies the MTE function. Our discussion is based on the standard potential outcome framework with selection on unobservables. In particular, we follow heckman2005structural.

Let $Y_1$ be the potential outcome under treatment, and $Y_0$ be the outcome without treatment. The treatment $D$ is a binary variable. We set $D=1$ when the treatment is received, and in this case, $Y_1$ is observed. Conversely, when the treatment is not received, $ D=0$, and $Y_0$ is observed. Thus, the observed outcome is

align*[align* omitted — 30 chars of source]

The potential outcomes are determined by observable covariates $X$ and unobservable factors $U_1$ and $U_0$. We write:

align*[align* omitted — 56 chars of source]

where $\mu_1$ and $\mu_0$ are unknown functions.

We assume the individual selects the treatment status according to the value of a vector of instruments $Z$ and an unobserved characteristic $U_D$:

align*[align* omitted — 89 chars of source]

where $\mu_D$ is an unknown function. We assume that, without loss of generality, $Z$ includes all elements of $X$.

We assume that selection into the treatment may be correlated with $(U_1, U_0)$ even after conditioning on $X$. This means that $ (U_1, U_0) $ and $ U_D$ may be correlated. We also assume that the instrument vector $Z$ is independent of the unobservable variables $(U_1, U_0, U_D)$. At least one element of $Z$ is not an element of $X$. This excluded instrument affects whether a unit receives the treatment, but does not directly affect the potential outcomes.

The selection process can be summarized using the propensity score, defined as the probability of receiving treatment based on the exogenous variables. Let $F_{U_D}$ denote the distribution function of $U_D$. We assume that $F_{U_D}$ is strictly increasing. Consequently, it holds $V:= F_{U_D} (U_D) \sim U(0,1)$. Let $P(Z) = F_{U_D} (\mu_D (Z) )$. Note that $D= \mathbbm{1} \{ P(Z) \geq V \} $. It holds that $ E (D | Z) = P (D=1 | Z) = P (Z) $. Thus, $P(Z)$ is the propensity score.

The MTE is the mean effect of the treatment conditional on observed characteristics $X$ at a particular value of $V$:

align*[align* omitted — 47 chars of source]

The MTE function is a function of $v$ ($MTE(\cdot, x)$ for a given value of $x$). It captures how the unobserved propensity, $V$, is related to the treatment effect. A small value of $V$ corresponds to a high propensity to receive the treatment, and a large value corresponds to a low propensity. Thus, if the MTE function is monotonically decreasing, those with a higher propensity to receive the treatment exhibit larger treatment effects.

For the identification, following heckman2005structural, we assume

align*[align* omitted — 85 chars of source]

Then, MTE can be identified by

align*[align* omitted — 68 chars of source]

over the support of distribution of $P(Z)$ conditional on $X=x$.

We make the following assumptions to simplify the MTE function formula, facilitating the estimation and statistical inferences. First, the potential outcomes are additive in $X$ and $(U_1, U_0)$. That is, we write

align*[align* omitted — 92 chars of source]

where $\mu_j (X) = E[ Y_j | X]$ for $j=0,1$.

assumption\ \begin{enumerate} • $X$ is independent of $(U_1, U_0, V)$. • $\mu_{j}(X)=\beta_{j}^\prime X$ for $j=0,1$. \end{enumerate}

Assumption (ref).(ref) states that the covariates are independent of all unobservables. It is a strong assumption. For example, it excludes heteroskedasticity. Nonetheless, it is employed, for instance, by carneiro2009estimating and brinch2017beyond, and arguably a standard assumption in the literature. As shown below, this can greatly simplify the formulation of the MTE function and its statistical analysis. Assumption (ref).(ref) imposes the linearity in the regression of a potential outcome on the covariates. Nonlinear models can be used as long as the parameters in the regression models can be estimated at the parametric rate.

Under these assumptions, the MTE function is separable in $X$ and $Z$. A straightforward calculation yields

align[align omitted — 92 chars of source]

where we define

align*[align* omitted — 70 chars of source]

Under Assumption (ref), $\Lambda(p,x)$ does not depend on $x$, so we write it as $\Lambda(p)$. We thus have

align[align omitted — 117 chars of source]

The MTE function is

align*[align* omitted — 124 chars of source]

In this paper, we consider this characterization of the MTE function and propose a method for constructing uniform confidence bands for it.

Estimation

This section provides a semiparametric estimation method of the MTE. Our estimation strategy combines the semiparametric estimation approach of carneiro2009estimating with the local quadratic estimation method of heckman2006understanding. We assume an independent and identically distributed sample, $(Y_i,X_i,Z_i,D_i), i=1,\dots, n$, of $(Y, X, Z, D)$ is available, where $X_i\in\mathbb{R}^d$ and $Z_{i}\in\mathbb{R}^m$. We also assume that $Z_{i}$ has enough variation to generate a propensity score $P(Z_{i})$ continuously from $0$ to $1$ conditional on $X_{i}=x$ for all values of $x$ in the region we want our confidence band to cover.

We first estimate the propensity score $P(Z_i)$. Note that it is a binary choice problem where $D$ is the binary outcome and $Z_i$ is the vector of regressors. We often consider parametric models for $P(Z_i)$ and may use the logit or probit models. Let $\hat P(Z_i)$ be the estimated propensity score for $i$.

We consider a modified version of the partially polynomial estimator considered in carneiro2009estimating for the estimation of $\beta_{0}$ and $\beta_{1}$. Let $\hat \beta_0$ and $\hat \beta_1 $ denote the estimators. Our theory is agnostic about the choice of the coefficient estimator. Other estimators can be used if their convergence rates are sufficiently fast. We suggest applying the partially linear estimator to the regressions of $D_iY_i$ and $(1-D_i)Y_i$. A straightforward calculation gives

align*[align* omitted — 179 chars of source]

where

align*[align* omitted — 126 chars of source]

The estimation of each of these two partially linear regression models is carried out in two steps. In the first step, we use the local linear estimator for the regression of $(D_i Y_i, X_i \hat P(Z_i))$ on $\hat P(Z_i)$. In the second step, we regress the residual of $ D_iY_i$ on the residual of $X_i\hat P(Z_{i})$ to obtain $\hat \beta_1$. The estimator $\hat \beta_0$ is obtained by applying the same procedure to the regression model for $(1-D_i) Y_i$.\footnote{The original procedure in carneiro2009estimating applies the partially linear estimator to the regression of $Y_i$ on $X_i$ and $\hat P(Z_i)$ separately for the observations with $D_i=0$ and $D_i =1$. Our modification enables the use of all observations to estimate each equation. Alternatively, we may estimate $\beta_0$ and $\beta_1 - \beta_0$ simultaneously from the regression model of $E[ Y\mid X, P(Z)]$ as in heckman2006understanding. Extending our procedure to these alternative approaches is straightforward.}

Lastly, we estimate $\Lambda (\cdot)$. Let $\tilde{Y}_{i}=Y_{i}- \beta_{0}^\prime X_{i}-(\beta_{1}-\beta_{0})^\prime X_{i}P(Z_{i})$. From ((ref)), we obtain $\tilde{Y}_{i}=\Lambda (P(Z_{i}))+\varepsilon_{i}$. We thus have

align*[align* omitted — 63 chars of source]

This equation motivates us to consider a local polynomial estimator to estimate the derivative of $\Lambda(\cdot)$. Let $\hat{\tilde{Y}}_i$ denote $Y_{i}-\hat{\beta}_{0}^\prime X_{i}-(\hat \beta_{1}- \hat \beta_{0})^\prime X_{i}\hat{P}(Z_{i})$. For each $p$, the local polynomial estimator can be obtained by solving

align*[align* omitted — 249 chars of source]

where $K(\cdot)$ is a kernel function of $\mathbb{R}$ and $h_n$ is a sequence of bandwidths. $(\theta_0, \theta_1, \theta_2)$ correspond to the conditional mean, first derivative, and second derivative, respectively. Let $\hat \theta_1 (p)$ be the estimate of $\theta_1$ at $p$. We use the local quadratic estimator because it provides desirable properties for estimating the first derivative of the nonparametric function. carneiro2009estimating and carneiro2011marginalretrurn also use the local quadratic estimator.

The resulting estimate of $MTE(p,x)$ is

equation*[equation* omitted — 101 chars of source]

Uniform confidence band

This section explains the computation of critical values to construct a uniform confidence band. Our uniform confidence bands have a usual form of “estimated function $\pm$ critical value $\times $ standard error.” However, we use different critical values from conventional Gaussian critical values. The theoretical justification of our procedure is discussed in the next section.

The critical value for our uniform confidence solves the following equation:

align*[align* omitted — 38 chars of source]

where $F_{n,1}(t)=\exp\big(-e^{-t-t^2/2\ell_n^2}\big)$ and $\ell_n$ is the largest solution to the following equation:

align*[align* omitted — 88 chars of source]

with

align*[align* omitted — 113 chars of source]

$F_{n,1} (t)$ approximates the distribution of the supremum of the standardized version of the estimated MTE function. Solving the equation, we have

equation*[equation* omitted — 87 chars of source]

Our two-sided uniform confidence band has the form

align*[align* omitted — 159 chars of source]

where $\hat{s}(p)/\sqrt{nh_n^3}$ is a standard error of $\widehat{MTE(p,x)}$. Our confidence band has a usual form of confidence interval (i.e., estimate $\pm$ critical value $\times$ standard error). However, the critical value is $c(1- \alpha)$ and differs from the Gaussian critical value. For $\hat{s}(p)$, in the simulations and the empirical example, we use

align*[align* omitted — 161 chars of source]

where $\hat{f}_{P}(p)$ is the kernel density estimator of $f_P$, $\hat{\varepsilon}_i=\hat{\tilde{Y}}_i-\hat{\theta}_0(\hat{P}(Z_{i})) $, and $\nu_{2}(K) = \int u^2K^2(u) du$.

We would like to emphasize that calculating our confidence band is a straightforward process. To determine the critical value $ c(1 - \alpha) $, we need to know the kernel function $K$, the bandwidth $h_n$, and the range of the propensity score $ (a_0, b_0) $. Once we establish the critical value, we can compute the confidence band using the standard errors, the bandwidth, and the sample size.

Importantly, our method can be applied to existing studies since the necessary information is typically available in published research papers, without the need to access the original data. For example, suppose we have a figure that displays the MTE estimate along with its pointwise standard errors. In that case, we can overlay our confidence band on the figure without needing to review the original data. This capability makes our confidence band highly beneficial not only for original research but also for evaluating existing studies.

Asymptotic Theory

In this section, we provide an asymptotic justification for our proposed procedure. We begin by introducing the relevant notations and definitions. Next, we outline a set of assumptions that will be used in our asymptotic theory. The distribution of the estimated MTE function is approximated by a Gaussian process. Finally, we justify our critical value based on an approximation of the supremum of this Gaussian process.

We introduce some notation. Let $m(p)$ denote $E[\tilde{Y}_{i}|P( Z_{i})=p]$. Also let $\varepsilon_{i}=\tilde{Y}_{i}-m(P( Z_{i}))$. Let $\sigma^{2}(p)$ denote $E[\varepsilon^{2}_{i}|P(Z_{i})=p]$. Let $s_n^2(p)$ denote the first-order approximated version of the variance of $\widehat{MTE}(p,x)$:

align*[align* omitted — 143 chars of source]

where $\kappa_2(K)=\int u^2K(u)du$. We note that $\widehat{\beta_{1}-\beta_{0}}$ is estimated at the parametric rate, while $\hat \theta_1$ is nonparametrically estimated and its convergence rate is slower than $\widehat{\beta_{1}-\beta_{0}}$. Therefore, the asymptotic variance of $\widehat{MTE}(p,x)$ depends only on the asymptotic variance of $\hat \theta_1 $. The $s_n^2(p)$ formula corresponds to the variance of $\hat \theta_1$.

We impose the following conditions to establish asymptotic theory.

assumption\ \begin{enumerate} • $\mathcal{X}:= \Pi_{j=1}^d[a_j,b_j]$, where $a_j<b_j$, $j=1,\dots, d$, and $\mathcal{X}$ is a strict subset of the support of $X$. Let $\mathcal{P}:= [a_0,b_0]$, where $0 < a_0 < b_0 < 1$. • The distribution of $P(Z_{i})$ has a bounded Lebesgue density $f_P(\cdot)$ on $(0,1)$ and is bounded away from zero. Furthermore, $f_P(\cdot)$ is twice continuously differentiable in $(0,1)$ and the second derivative of $f_P(\cdot)$ is uniformly bounded on $(0,1)$. • $\sigma^{2}(p)$ is continuous on $p \in \mathcal{P}$, and $\sup_{p\in (0,1)}E[|\varepsilon_i|^4|P(Z_{i})=p]<\infty$ holds. Furthermore, $p\mapsto \sigma^{2}(p)f_P(p)$ is Lipschitz continuous. • $m(p)$ is three times continuously differentiable on $(0,1)$ with uniformly bounded derivatives. • Let $K(\cdot)$ be a kernel function on $\mathbb{R}$ compactly supported, bounded, and symmetric around zero and six times differentiable. • \begin{enumerate} • $\|\hat{P}(z)-P(z)\|_{\infty}=o_p(h_n^{3})$ where the supremum is taken in terms of the support of $Z$. • There exists a VC-type class $\mathcal{M}$ with the envelope 2 such that $\lim_{n\rightarrow\infty}\Pr(\hat{P}\in\mathcal{M})=1$ holds. \end{enumerate} • $h_n=Cn^{-\eta}$, where $C$ and $\eta$ are positive constant such that $\frac{1}{7}< \eta <\frac{1}{6}$ and $h_{n}\leq1$. • \begin{enumerate} • $\inf_{n\geq 1} \inf_{p\in \mathcal{P}}s_n(p)>0$ and $s_n(p)$ is continuous for each $n\geq 1$. • An estimator of $s_n^2(p)$, $\hat{s}^2(p)$, exists such that \begin{align*} \sup_{p\in \mathcal{P}}\big|\hat{s}^2(p)- s_n^2(p) \big|=O_p(n^{-c}). \end{align*} \end{enumerate} • \begin{enumerate} • There exists an estimator $\widehat{\beta}_{0}$ and $\widehat{\beta_{1}-\beta_{0}}$ such that \begin{equation*} \widehat{\beta}_{0}-\beta_{0}=O_{p}(n^{-1/2}),\quad \widehat{\beta_{1}-\beta_{0}}-(\beta_{1}-\beta_{0})=O_{p}(n^{-1/2}). \end{equation*} • For each $\ell\in\{1,\cdots,d\}$, $E[X_{i,\ell}|P(Z_{i})=p]$ is continuously differentiable in $(0,1)$. Especially, the derivative of $E[X_{i,\ell}|P(Z_{i})=p]$ is uniformly bounded on $(0,1)$. • For each $\ell\in\{1,\cdots,d\}$, $\sup_{p\in (0,1)}E[|X_{i,\ell}|^2|P(Z_{i})=p]<\infty$ holds. \end{enumerate} \end{enumerate}
remarkAssumption (ref).(ref) specifies the region over which we construct uniform confidence bands for $X$ and $p$. The requirement that the region for $X$ is a Cartesian product of closed intervals is typically not restrictive in practice. In empirical applications examining the MTE function, researchers commonly fix $X$ at a specific value, such as the sample mean. Additionally, the region of $p$ must be a closed interval excluding the neighborhoods around the boundary values 0 and 1. It is also important in practice, as the estimated MTE function tends to behave poorly near these extreme values.
remarkThe conditions on the functions in the data-generating process, as specified in Assumptions (ref).(ref)-(ref).(ref), are standard in the literature on nonparametric estimation. To establish uniform convergence results for local polynomial estimators, we follow the approach of chernozhukov2014gaussian and assume the existence of conditional fourth moments of error terms. The Lipschitz continuity condition is essential for approximating the supremum of a Gaussian process with the supremum of a stationary Gaussian process.
remarkWe assume that the kernel function is six times differentiable, as stated in Assumption (ref).(ref). This assumption is important for establishing the asymptotic theory for the maximum of a Gaussian homogeneous field, particularly because we follow Theorem 4.3 from piterbarg1996asymptotic. This assumption is also essential in analytically deriving the limit distribution of the supremum of a stationary Gaussian process. A practical implication is that it excludes specific kernels, such as the uniform kernel.
remarkAssumption (ref).(ref) outlines the conditions that the predicted propensity score, $\hat{P}(Z)$, must satisfy. In practice, $\hat{P}(Z)$ can be estimated using parametric models like probit or logit. In these cases, the estimation error is typically of the order $O_p(n^{-1/2})$. This convergence rate can meet the requirements outlined in Assumption (ref).(ref), provided that the sequence of $h_n$ satisfies the conditions in Assumption (ref).(ref). Semiparametric estimators, including those based on single-index models, are also options. However, appropriate higher-order kernels must be utilized to achieve a sufficiently fast convergence rate. To see this, let $\nu$ be the order of moment where $\int x^{\nu}K(x)dx\neq0$ holds. Under regular conditions, we need to have $3d<\nu$ to make Assumption (ref).(ref) feasible, where $d$ is the dimension of $Z$. Especially when considering a single-index model, we require at least a 4th-order kernel. The definition of the VC-type class related to Assumption (ref).(ref) is provided in Appendix (ref). It is important to note that this condition is relatively mild and is satisfied by many commonly used methods for estimating propensity scores.
remarkThe bandwidth, $h_n$, must meet the conditions given in Assumption (ref).(ref). This requirement is essential to ensure that the estimated MTE function is asymptotically unbiased. This rate is faster than the mean-squared-error optimal rate. In other words, we utilize undersmoothing to achieve asymptotically unbiased estimation. Furthermore, this condition, along with Assumption (ref).(ref) and the Lipschitz continuity of the kernel, allows us to asymptotically disregard the bias that arises from propensity score estimation in the first step.
remarkAssumption (ref).(ref) addresses the asymptotic variances of the MTE function. The conditions of strict positivity and continuity in Assumption (ref).(ref) are standard requirements. The rate condition in Assumption (ref).(ref) is necessary to prevent the inverse of the standard error from diverging to infinity. This rate can be achieved using the estimator provided at the end of the previous section. Alternatively, the formula for $s_n^2(p)$ suggests another estimator: \begin{align*} \tilde{s}^2(p) = \frac{1}{nh_n^3\hat{f}_P^2(p)\kappa_2^2(K)} \sum_{i=1}^{n} \hat{\varepsilon}_i^2 K^2\left(\frac{\hat{P}(Z_{i})-p}{h_n}\right)(\hat{P}(Z_{i})-p)^2. \end{align*} Simulations indicate that the estimator presented in the previous section behaves more stably than this alternative estimator.
remarkAssumption (ref).(ref) pertains to the impact of preliminary estimation before the final step in the MTE estimation process. The semiparametric estimator of the partially linear model can produce coefficient estimators that meet the rate conditions outlined in Assumption (ref).(ref), even when the propensity score $P(Z_i)$ is estimated. The coefficient estimator presented in carneiro2009estimating satisfies this condition. For more detailed assumptions required to achieve root-$n$ consistent estimators, refer to mammen2016semiparametric. Assumptions (ref).(ref) and (ref).(ref) indicate that the covariates must exhibit sufficient variation, even after conditioning on the propensity score.

The following lemma establishes a linear expansion of the semiparametric estimator.

lemmaLet Assumption (ref) and (ref) hold. Then, \begin{align*} \sup_{(p,x)\in \mathcal{P}\times \mathcal{X}}\sqrt{nh_n^{3}} &\left| \frac{\widehat{MTE}(p,x)-MTE(p,x)}{\hat{s}(p)}-\right.\\ & \left.\frac{1}{nh_n^{3}f_P(p)\kappa_2(K)s_n(p)}\sum_{i=1}^n\varepsilon_iK\left(\frac{P(Z_i)-p}{h_n}\right)(P(Z_i)-p) \right|\\ =O_P(n^{-c})&, \end{align*} for some positive constant $c>0$.

This lemma demonstrates that preliminary estimation of propensity scores and coefficients in the partial linear model does not asymptotically affect the estimation error of the MTE function. The proof addresses several potential sources of estimation error. First, we consider the difference in the dependent variables. While the local polynomial estimation of the MTE function would ideally use the true dependent variable $Y_{i}-\beta_{0}^\prime X_{i}+(\beta_{1}-\beta_{0})^\prime X_{i}P(Z_{i})$, we employ the feasible version $\hat{\tilde{Y}}_{i}$ defined as $Y_{i}-\hat{\beta}_{0}^\prime X_{i}+(\widehat{\beta_{1}-\beta_{0}})^\prime X_{i}\hat{P}(Z_{i})$. The difference between these is asymptotically negligible under Assumptions (ref).(ref) and (ref).(ref), noting that Assumptions (ref).(ref) and (ref).(ref) together ensure sufficiently fast convergence of the estimated propensity score. Second, we replace the asymptotic variance $s_n$ with its estimator $\hat{s}$, which is handled by Assumption (ref).(ref). Third, and most importantly, the nonparametric estimation uses the generated regressor $\hat{P}(Z_i)$ rather than the true propensity score. Assumptions (ref).(ref) and (ref).(ref) are essential for handling these generated regressor errors. The primary technical challenge arises because the estimated propensity score is a function, requiring us to manage an empirical process indexed by the propensity score.

Define

align*[align* omitted — 109 chars of source]

and

align*[align* omitted — 139 chars of source]

By Lemma (ref), we have

align*[align* omitted — 161 chars of source]

We approximate the supremum of the empirical process $c_n(p)\sqrt{nh_n^{3}}(T_n(p)-E[T_n(p)])$ by the supremum of a Gaussian process using the result of chernozhukov2012gaussian. Define

align*[align* omitted — 94 chars of source]
lemmaLet Assumption (ref) and (ref) hold. Then, for every $n\geq 1$, there is a tight Gaussian random variable $\tilde{B}_{n,1}$ in $\ell^{\infty} (\mathcal{P})$ with mean zero and covariance function \begin{align*} &E[ \tilde{B}_{n,1}(p) \tilde{B}_{n,1}(\check{p})] \notag \\ =& h_n^{-3}c_n(p)c_n(\check{p})E \left[\varepsilon_{i}^2K\left(\frac{P(Z_{i})-p}{h_n}\right) K\left(\frac{P(Z_{i})-\check{p}}{h_n}\right)(P(Z_{i})-p)^2(P(Z_{i})-\check{p})^2 \right], \end{align*} and there is a sequence $\tilde{W}_{n,1}$ of random variables such that $\tilde{W}_{n,1}$ is equal to $\sup_{p\in \mathcal{P}} \tilde{B}_{n,1}(p)$ in distribution and as $n \to \infty$: \begin{align*} |W_n-\tilde{W}_{n,1}|=O_P\{(nh_n^{3})^{-1/6}\log n +(nh_n^{3})^{-1/4}\log^{5/4}n+(nh_n^{6})^{-1/4}\log^{3/2}n \}. \end{align*}

Next, we show that $\tilde{W}_{n,1}$ can be further approximated by the supremum of a stationary Gaussian process.

lemmaLet Assumption (ref) and (ref) hold. For sufficiently large $n$, there is a tight Gaussian process $\tilde{B}_{n,2}$ in $\ell^{\infty} (\mathcal{P})$ with mean zero and covariance function \begin{align*} E[ \tilde{B}_{n,2}(p) \tilde{B}_{n,2}(\check{p})]=\mathbf{\rho}(p-\check{p}) \end{align*} for $p, \check{p} \in h_{n}^{-1}\mathcal{P}$, where we define \begin{align*} \rho(p)=\frac{\int uK(u)(u-p)K(u-p)du}{\int u^{2}K^2(u)du}. \end{align*} And there is a sequence $\tilde{W}_{n,2}$ of random variables such that $\tilde{W}_{n,2}$ is equal to $\sup_{p\in \mathcal{P}} \tilde{B}_{n,2}(h_{n}^{-1}p)$ in distribution and as $n \to \infty$: \begin{align*} |\tilde{W}_{n,1}-\tilde{W}_{n,2}|=O_P\left\{h_n^{1/2}\sqrt{\log h_n^{-(3/2 )}} \right\}. \end{align*}

From Lemma (ref) to Lemma (ref) with symmetry, there exist some positive constant $c$ such that

equation*[equation* omitted — 208 chars of source]

holds. The supremum of the studentized estimated MTE function can be approximated by the supremum of a stationary Gaussian process. Moreover, the approximation error is of a polynomial order.

Combining this with known results on the distribution of the supremum of a stationary Gaussian process, we obtain the following main theorem:

theoremSuppose Assumptions (ref) and (ref) hold. Then, uniformly in $t$ over any finite interval, the following result holds: \begin{align}\nonumber &\Pr\Big(\ell_n\Big[ \sup_{(p,x)\in \mathcal{P}\times \mathcal{X}}\sqrt{nh_n^3}\left|\frac{\widehat{MTE}(p,x)-MTE(p,x)}{\hat{s}(p)} \right|-\ell_n \Big]<t \Big) \\ =& \exp\big(-2e^{-t-t^2/2\ell_n^2} \big)+o(1), \end{align} as $n\to \infty$, where $\ell_n$ is the largest solution to the following equation: \begin{align*} (b_{0}-a_{0})\cdot h_n^{-1}\sqrt{\lambda}(2\pi)^{-1}\exp(-\ell_n^2/2)=1, \end{align*} where \begin{align*} \lambda:=-\frac{\int g(u)g^{\prime\prime}(u)du}{\int g^2(u)du}\quad and \quad g(u):=uK(u). \end{align*}

Theorem (ref) provides theoretical justification for our confidence band. For any $(p,x)\in\mathcal{P}\times \mathcal{X}$, we have

align*[align* omitted — 305 chars of source]
remarkThe result in Theorem (ref) is derived from a refined approximation of the distribution of the supremum of a stationary Gaussian process $\tilde{W}_{n,2}$, as presented by piterbarg1996asymptotic. While the commonly used Gumbel approximation $\Pr\Big(l_n \left(|\tilde{W}_{n,2}| - l_n \right) < t \Big) = \exp\big(-2e^{-t}\big) + o(1)$ has a logarithmic rate of convergence, piterbarg1996asymptotic demonstrates a more accurate approximation: $\Pr\Big(l_n \left(|\tilde{W}_{n,2}| - l_n \right) < t \Big) = \exp\big(-2e^{-t - t^2/2 \ell_n^2}\big) + o(n^{-c})$, where the approximation error is of polynomial order. The correction term $-t^2/2 \ell_n^2$ reduces the critical value relative to the Gumbel distribution, resulting in uniform confidence bands that achieve better coverage while being less conservative than Gumbel-based methods.

Monte Carlo Experiments

In this section, we present the results of Monte Carlo experiments. Our settings include cases with increasing, decreasing, and constant MTE functions. The results indicate that our confidence set has appropriate coverage in finite samples and is much narrower than the bands based on the Gumbel approximation.

Data generating process

We consider settings with two exogenous covariates and two excluded instruments. We generate $(X_{1},X_{2},Z_{1},Z_{2})$ as follows:

equation*[equation* omitted — 80 chars of source]

where we denote $\mathbf{0}_{4}$ and $I_{4}$ as a $4\times1$ zero vector and a $4\times4$ identity matrix, respectively. $X=(X_1, X_2)$ is the vector of covariates that affect potential outcomes, and $Z=(X_1, X_2, Z_1, Z_2)$ is the vector of instruments that affect the propensity score. The potential outcomes are generated by

align*[align* omitted — 75 chars of source]

The treatment $D$ is generated by

equation*[equation* omitted — 78 chars of source]

The vector of unobservable components $(U_{0},U_{1},U_D)^{\prime}$ is generated from $N(\mathbf{0}_{3},\Sigma)$, where $\Sigma$ depends on the design. We use the following three types of covariance matrices.

align*[align* omitted — 522 chars of source]

In the designs $\Sigma_1$ and $\Sigma_2$, the unobservable that affects selection into the treatment ($U_D$) is correlated with $U_1$ and $U_2$. Thus, there exists endogenous selection. In contrast, in design $\Sigma_{3}$, the treatment status is independent of the potential outcomes conditional on $Z$.

We now derive the MTE function in this setting. By the definition of normal distribution, the conditional distribution of $(U_1, U_0)$ given $V=\Phi (U_D)$, where $\Phi(\cdot)$ is the cumulative distribution function of the standard normal distribution, can be analytically calculated. Thus, the MTE function has a closed-form representation:

align*[align* omitted — 103 chars of source]

We focus on $MTE(p,E(X_1), E(X_2)) = MTE(p,0,0) = (Cov(U_{1},V)-Cov(U_{0},V))\times \Phi^{-1}(p) $. The shape of the MTE function varies across the types of $\Sigma$. $\Sigma_1$ yields an increasing MTE function, while the MTE function under $\Sigma_2$ is decreasing. In the case of $\Sigma_3$, the MTE function is constant.

Estimation Procedure

We estimate the MTE function using the following steps.

enumerate• We estimate the coefficients in $\mu_D$ using the maximum likelihood estimator (probit) and construct the estimated propensity score, $P(Z)$. The MTE function can only be estimated at the intersection of the support of the propensity score for individuals who received the treatment and the support of the propensity score for those who did not receive the treatment. Therefore, we exclude all observations that fall outside this intersection. • The partially linear estimator is used to estimate the regression coefficients. The estimation method is described in Section (ref). Let $\hat{\beta}_{dk}$ denote the coefficient estimator on $X_k$ in the regression for $Y_d$. • We estimate the $\Lambda (p) = E[U_1-U_0|V=p]$ and its first derivative through the local quadratic estimator with a Gaussian kernel. We use the rule of thumb bandwidth $h_{n}$ in fan1996local for the nonparametric estimation. We change the order of $h_{n}$ from $n^{-1/7}$ to $n^{-2/13}$ so that the bandwidth satisfies Condition (ref). • $MTE(p, E(X_1), E(X_2))$ is estimated as $(\hat \beta_{11} - \hat \beta_{21}) \bar X_1 + (\hat \beta_{12} - \hat \beta_{22})\bar X_2 + \hat \Lambda'(p)$, where $\hat \Lambda'(p)$ is the local quadratic estimator of the first derivative of $\Lambda (p)$.

Methods in comparison

We evaluate three confidence bands — pointwise, Gumbel, and our method — on the interval $[0.15, 0.85]$. While pointwise confidence intervals are widely used in applied research, they are not valid as a uniform confidence band. The confidence bands based on the Gumbel distribution are theoretically valid. However, the approximation by the Gumbel distribution is crude and yields a conservative confidence band.

Results

Table (ref) shows the results of the Monte Carlo experiments. The number of Monte Carlo replications is 1000. The sample size is $n=3000$.

table[table omitted — 1,269 chars of source]

The confidence band constructed by our method contains the true MTE function at a nominal coverage probability with a maximum error of 2.5% in all cases. The empirical coverage probabilities of the pointwise confidence band are much lower than the nominal levels. The pointwise confidence band is too narrow, failing to include the true MTE function at an intended probability. This result is not surprising because the pointwise confidence band is designed to include the true value of the MTE function at a particular point, not the range of the function. Therefore, pointwise confidence intervals are unsuitable for assessing the uncertainty of the estimated function. The confidence bands based on the Gumbel distribution attain the nominal coverage probabilities. However, this method is conservative and may not be informative.

Our procedure's critical value for the 95% confidence band is around 3. As expected for uniform coverage, we need to use a larger critical value than the conventional pointwise approach (1.96). However, this value is still smaller than the critical value obtained from the Gumbel distribution, which is approximately 4. Interestingly, the critical values do not vary significantly across different designs. This intermediate positioning makes our procedure offer more informative confidence bands than the Gumbel method.

Empirical Application

We illustrate the use of our proposed method in an empirical example. Using the MTE function, we analyze the impact of university enrollment on future income. We then construct our uniform confidence band for the MTE function and compare it with other methods.

We use the data from the National Longitudinal Survey of Young Men (NLSYM) for 1976. This data is originally used in Card95; we use the data set excerpted from hansen2022econometrics.\footnote{It is available at the website of hansen2022econometrics (\url{https://press.princeton.edu/books/hardcover/9780691235899/econometrics}).} The dependent variable $Y_{i}$ is the log of weekly wage. The treatment variable is a dummy variable $D_{i}$ representing university enrollment. In this analysis, if individuals have 13 years of education or more, we assume they entered universities, i.e., $D_{i}=\mathbbm{1}\{education_{i}\geq 13\}$. The included exogenous variable is $experience$ defined as $experience_{i}=age_{i}-(education_{i}+6)$. In particular, we consider quadratic regression models for $Y_{ji}$ for $j\in\{0,1\}$:

align*[align* omitted — 92 chars of source]

where $U_{ji}$ is an error term. Our instrument variables are $mothereducation$, $fathereducation$, $nearc2$ and $nearc4$. $mothereducation$ and $fathereducation$ denote the mother's and father's educational attainments, respectively. $nearc2$ and $nearc4$ are dummy variables indicating that an individual grew up in a county with a 2-year and 4-year college, respectively. The model for the propensity score is

align*[align* omitted — 232 chars of source]

where we assume $V_{i}\sim N(0,1)$. We thus consider the probit model for the propensity score. We exclude the observations with missing variables.

Figure (ref) displays histograms of the estimated propensity scores for both the treated and untreated groups. Although the instruments are discrete variables, the propensity score distributions for both groups range widely from 0 to 1, showing significant overlap between the treated and untreated populations. The extensive set of covariates included in the propensity score model results in considerable variation in the estimated propensity scores across different observations. These findings indicate that the assumption of a continuous propensity score for the MTE estimation can be effectively met by incorporating a comprehensive set of covariates. That is, the range of the propensity score can be sufficiently dense even when the underlying instruments are discrete.

We estimate the MTE function using the local quadratic estimator. The bandwidth $h_{n}$ is calculated using the method of fan1996local. We use the Gaussian kernel. We estimate the MTE function on the region $[0.01, 0.92]$. We note that the MTE function is estimable only on the intersection of the support of the propensity score given the treatment is taken and that of the propensity score given the treatment is not taken. The region $[0.01, 0.92]$ satisfies this condition. It contains 1925 observations. The propensity score takes 1226 points. We then conduct statistical inferences on the MTE function over $[0.15, 0.85]$.

figure[figure omitted — 409 chars of source]
figure[figure omitted — 314 chars of source]

Figure (ref) plots the three 95% uniform confidence bands and the estimated MTE function over $[0.15, 0.85]$. The bold black curve is the estimated MTE function. The estimated MTE function exhibits some variability but shows a downward trend. The dashed curves correspond to the confidence band derived from the Gumbel distribution, with a critical value of $4.16$. This dashed confidence band is so wide that we cannot reject any null hypotheses that the MTE is constant, monotone increasing, or monotone decreasing. The dotted curves correspond to the pointwise confidence interval, with a critical value of 1.96. This dotted confidence band is the narrowest of the three intervals, but it is not valid for a uniform evaluation of the MTE function. The band surrounded by solid curves is our confidence band, and the critical value is $2.99$. Our confidence band is valid for a uniform evaluation of the MTE function.

Our confidence band does not contain a constant or monotonically increasing function, which allows us to reject the null hypothesis of a weakly increasing MTE function. Conversely, we cannot reject the null hypothesis that the MTE function is decreasing. Interestingly, while our uniform confidence band is broader than the pointwise confidence band, both agree that the MTE function is not constant or monotonically increasing. If we focus on some particular range, they may disagree; for example, the pointwise confidence band does not support an invariant effect for $P(Z) < 0.28$, but our band does not exclude it. The band derived from the Gumbel approximation is substantially wider and encompasses weakly increasing functions. These findings suggest that while our approach yields a wider band, it provides meaningful and informative results.

Conclusion

This paper proposes a method to construct a uniform confidence band for the marginal treatment effect function. To estimate the MTE, we impose the linearity of the potential outcomes with respect to covariates and provide a semiparametric estimator. Our uniform confidence band relies on the approximation of the supremum of a Gaussian process combined with the Gaussian approximation of empirical processes. Empirical researchers are recommended to add our easy-to-implement uniform confidence bands when reporting MTE function estimation results.

Several avenues for future research emerge from this study. While our paper focuses on local polynomial estimation of the MTE function, alternative approaches such as sieve or series estimation, as employed by hoshino2022estimating, warrant exploration. Developing methods for uniform confidence bands for the MTE function estimated via sieve methods presents an intriguing challenge because it requires mathematical techniques different from those employed here. Relatedly, it is also a standard practice to use parametric polynomial models in MTE function estimation, as demonstrated by brinch2017beyond and sasaki2024welfare. Extending methodologies to accommodate such parametric approaches could benefit practitioners.

Another promising direction involves developing direct tests for the shape of the MTE function. While our method can be applied to shape testing, it is not specifically tailored to any particular hypotheses. Refining approaches to focus on specific hypotheses, such as the monotonicity of the MTE function, could potentially enhance statistical power.

\paragraph{Conflict of Interest statement:} The authors report there are no competing interests to declare.