Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
28,256 characters · 5 sections · 6 citation commands
Mode Treatment Effect
The effects of policies on the distribution of outcomes have long been of central interest in many areas of empirical economics. A policy maker might be interested in the difference of the distribution of outcome under treatment and the distribution of outcome in the absence of treatment. The empirical studies of distributional effects include but not are not limited to \citet*{freeman1980unionism}, \citet*{card1996effect}, \citet*{dinardo1995labor}, and \citet*{bitler2006mean}. Most researches use the difference of the averages or quantiles of the treated and untreated distribution, known as average treatment effect and quantile treatment effect, as a summary for the effect of treatment on distribution. The mode of a distribution, which is also an important summary statistics of data, has long been ignored in the literature. This paper fills up the gap by studying the mode treatment effect: the difference of the modes of the treated and untreated distribution. Compared to the average and the quantile treatment effect, the mode treatment effect has two advantages: (1) mode captures the most probable value of the distribution under treatment and in the absence of treatment. It provides a better summary of centrality than average and quantile when the distributions are highly skewed; (2) mode is robust to heavy-tailed distributions where outliers don't follow the same behavior as the majority of a sample. In economic studies, it is especially often to confront a skewed and heavy-tailed distribution when the outcome of interest is income or wage.
This paper discusses the estimation and inference of the mode treatment effect under the Strong Ignorability assumption \citep*{rosenbaum1983central}, which states that conditional on a vector of control variables the treatment is randomly assigned. The first estimator I propose is the kernel estimator. I estimate the density function of the outcome distribution using the kernel method and define the maximum of the estimated density function as the estimator of the mode. While the kernel estimator is a straightforward estimator, it requires to estimate the conditional density function in the process, and the estimation of the conditional density function may be difficult in practice when there exist more than two or three control variables, due to the curse of dimension. The kernel estimator is appropriate if there are less than three control variables. In practice, however, researchers may want to include as many control variables as possible in order to make their identification robust. In this circumstance, the curse of dimension may lead to inaccurate estimation and misleading inference.
To address this problem, I propose the ML estimator. The key feature of the proposed ML estimator is that it translates the estimation of the conditional density function into the estimation of conditional expectation, which we can apply a rich set of ML methods, such as Lasso, random forests, neural nets, and etc, to estimate. This feature provides researchers with the flexibility to apply ML methods to estimate the density function of the outcome distribution. By the virtue of ML methods, the proposed ML estimator can handle the situation when there exist many control variables, even the number of control variables is comparable to or more than the sample size. However, it is well-known that the regularization bias embedded in ML methods may lead to the bias of the final estimator and misleading inference \citep*{chernozhukov2018double}. To solve this problem, I further derive the Neyman-orthogonal scores \citep*{chernozhukov2018double} for each estimation which requires the first-step estimation of the conditional expectation. These Neyman-orthogonal scores, to my best knowledge, are new results. The proposed ML estimator is built on the newly derived Neyman-orthogonal score, and hence, it is robust to the regularization bias of the first-step ML estimation.
I derive the asymptotic properties for both the proposed kernel and ML estimators. I show that both estimators are consistent and asymptotically normal with the rate of convergence $\sqrt{Nh^{3}}$, where $N$ is the sample size and $h$ is the bandwidth of the chosen kernel, which is slower than the traditional rate of convergence $\sqrt{N}$ in the estimation of mean and quantile. In fact, this rate of convergence complies the intuition. While the estimators of mean and quantile are the weighted average of all the available observations, only a small portion of observations near the mode provides the information to the estimator of the mode. This explains the slower rate of convergence for the proposed estimators.
This paper contributes to the program evaluation literature which includes the studies of average treatment effect: \citet*{rosenbaum1983central}, \citet*{heckman1985alternative}, \citet*{heckman1997matching}, \citet*{hahn1998role}, and \citet*{hirano2003efficient}; the studies of quantile treatment effect: \citet*{abadie2002instrumental}, \citet*{chernozhukov2005iv}, and \citet*{firpo2007efficient}; the studies of mode estimation and mode regression: \citet*{parzen1962estimation}, \citet*{eddy1980optimum}, \citet*{lee1989mode}, \citet*{yao2014new}, and \citet*{chen2016nonparametric}; as well as the causal inference of ML methods: belloni2012sparse, Belloni14restud, chernozhukov2015valid, belloni2017program, chernozhukov2018double, and athey2019generalized. This paper is also closely related to the robustness of average treatment effect estimation discussed in \citep*{robins1995semiparametric} and the general discussion in \citep*{chernozhukov2016locally}. The asymptotic properties of the robust estimators discussed in these papers remain unaffected if only one of the first-step estimation with classical nonparametric method is inconsistent.
Plan of the paper. Section 2 sets up the notation and framework for the discussion of the mode treatment effect. Section 3 discusses the kernel method and derives the asymptotic properties. Section 4 presents the ML estimator for density estimation and the corresponding Neyman-orthogonal score. I combine the Neyman-orthogonal score with the cross-fitting algorithm to propose the ML estimator of the mode treatment effect, and derive its asymptotic properties. Section 5 concludes this paper.
Let Y be a continuous outcome variable of interest, D the binary treatment indicator, and X $d\times 1$ vector of control variables. Denote by $Y_{1}$ an individual's potential outcome when $D=1$ and $Y_{0}$ if $D=0$. Let $f_{Y_{1}}\left(y\right)$ and $f_{Y_{0}}\left(y\right)$ be the marginal probability density function (p.d.f.) of $Y_{1}$ and $Y_{0}$, respectively. The modes of $Y_{1}$ and $Y_{0}$ are the values that appear with the highest probability. That is, \[ \theta_{1}^{*}\equiv\operatorname*{arg\,max}_{y\in \mathcal{Y}_{1}}f_{Y_{1}}\left(y\right)\text{ and }\theta_{0}^{*}\equiv\operatorname*{arg\,max}_{y\in \mathcal{Y}_{0}}f_{Y_{0}}\left(y\right), \] where $\mathcal{Y}_{1}$ and $\mathcal{Y}_{0}$ are the supports of $Y_{1}$ and $Y_{0}$. Here I assume that $\theta_{1}^{*}$ and $\theta_{0}^{*}$ are unique, meaning that both $Y_{1}$ and $Y_{0}$ are unimodal. I also assume that the modes $\theta_{1}^{*}$ and $\theta_{0}^{*}$ are in the interior of the common supports of $Y_{1}$ and $Y_{0}$. These conditions are formally stated in the following assumption:
Assumption 1 has been widely adopted in many studies \citep*{parzen1962estimation, eddy1980optimum, lee1989mode, yao2014new}. Under Assumption 1, the mode treatment effect is uniquely defined as $\Delta^{*}\equiv\theta_{1}^{*}-\theta_{0}^{*}$. The following states the strong ignorability assumption \citep*{rosenbaum1983central}:
The first part of Assumption 2 assumes that potential outcomes are independent of treatment after conditioning on the observable covariates $X$. The second part states that for all values of $X$, both treatment status occur with a positive probability. Under the strong ignorability condition, both $f_{Y_1}$ and $f_{Y_0}$ can be identified from the observable variables $\left(Y,D,X\right)$ since
and thus \[ f_{Y_{1}}(y)=E\left[ f_{Y_{1}\mid X}\left(y\mid X\right)\right]=E\left[ f_{Y\mid D=1,X}\left(y\mid X\right)\right].\tag{2.1} \] Similarly, we have \[ f_{Y_{0}}\left(y\right)=E\left[ f_{Y\mid D=0,X}\left(y\mid X\right)\right].\tag{2.2} \] Equation (2.1) and (2.2) shows the identification result of the density function $f_{Y_{1}}$ and $f_{Y_{0}}$. Then it is straightforward to identify their modes $\theta_{1}^{*}$ and $\theta_{0}^{*}$: \[ \theta_{1}^{*}=\operatorname*{arg\,max}_{y\in\mathcal{Y}_{1}}E\left[ f_{Y\mid D=1,X}\left(y\mid X\right)\right]\text{ and }\theta_{0}^{*}=\operatorname*{arg\,max}_{y\in\mathcal{Y}_{0}}E\left[ f_{Y\mid D=0,X}\left(y\mid X\right)\right].\tag{2.3} \] If both $f_{Y\mid D=1,X}\left(y\mid X\right)$ and $f_{Y\mid D=0,X}\left(y\mid X\right)$ are differentiable with respect to $y$, we can further identify the modes using the first-order conditions under Assumption 1: \[ E\left[ f_{Y\mid D=1,X}^{\left(1\right)}\left(\theta_{1}^{*}\mid X\right)\right]=0\text{ and }E\left[ f_{Y\mid D=0,X}^{\left(1\right)}\left(\theta_{0}^{*}\mid X\right)\right]=0,\tag{2.4} \] where $m^{\left(s\right)}\left(y, x\right)\equiv \partial^{s} m\left(y, x\right)/ \partial y^{s}$ denotes the partial derivatives with respect to $y$.
Equation (2.1)-(2.4) provide us a direct way to estimate the modes $\theta_{1}^{*}$ and $\theta_{0}^{*}$. Intuitively, we estimate the density functions $f_{Y_{1}}(y)$ and $f_{Y_{0}}(y)$ in the first step and use the maximizers of the estimated density functions as the estimators of the modes. Section 3 and 4 presents the kernel and ML estimation method, respectively.
In this section, I propose kernel estimators for $\theta_{1}^{*}$, $\theta_{0}^{*}$, and the mode treatment effect $\Delta^{*}=\theta_{1}^{*}-\theta_{0}^{*}$. Let $K(\cdot)$ be a kernel function with bandwidth $h$. Define the estimators of the density functions $f_{Y_{1}}\left(y\right)$ and $f_{Y_{0}}\left(y\right)$ as, \[ \hat{f}_{Y_{1}}\left(y\right)=\frac{1}{n}\sum_{i=1}^{n}\hat{f}_{Y\mid D=1,X}\left(y\mid X_{i}\right), \] \[ \hat{f}_{Y_{0}}\left(y\right)=\frac{1}{n}\sum_{i=1}^{n}\hat{f}_{Y\mid D=0,X}\left(y\mid X_{i}\right) \] with the kernel estimators \[ \hat{f}_{Y\mid D=1,X}\left(y\mid x\right) = \frac{\sum_{j=1}^{n}D_{j}K_{h}\left(y-Y_{j}\right)K_{h}\left(x-X_{j}\right)}{\sum_{j=1}^{n}D_{j}K_{h}\left(x-X_{j}\right)}, \] \[ \hat{f}_{Y\mid D=0,X}\left(y\mid x\right) = \frac{\sum_{j=1}^{n}\left(1-D_{j}\right)K_{h}\left(y-Y_{j}\right)K_{h}\left(x-X_{j}\right)}{\sum_{j=1}^{n}\left(1-D_{j}\right)K_{h}\left(x-X_{j}\right)} \] where $K_{h}\left(y-Y_{j}\right)= h^{-1}K\left(\frac{y-Y_{j}}{h}\right)$ and \[ K_{h}\left(x-X_{j}\right)= h^{-d}K\left(\frac{x_{1}-X_{j1}}{h}\right)\times ... \times K\left(\frac{x_{d}-X_{jd}}{h}\right). \]
Then it is straightforward to define the estimators of the modes $\theta_{1}^{*}$ and $\theta_{0}^{*}$:
\[ \hat{\theta}_{1}\equiv \operatorname*{arg\,max}_{y} \hat{f}_{Y_{1}}\left(y\right), \] \[ \hat{\theta}_{0}\equiv\operatorname*{arg\,max}_{y} \hat{f}_{Y_{0}}\left(y\right). \] The estimator of the mode treatment effect $\Delta^{*}$ is $\hat{\Delta}\equiv \hat{\theta}_{1}-\hat{\theta}_{0}$. Through out the paper, I impose the following conditions on the kernel $K\left(\cdot\right)$:
The first part of Assumption 3 requires that $K\left(u\right)$ is bounded. Although the second part implies that $K\left(u\right)$ is a first-order kernel, the arguments in this paper can be easily extended to higher-order kernels. We assume the first-order kernel here just for simplicity. The third part imposes enough smoothness on $K\left(u\right)$.
Theorem 1 and 2 show that the asymptotic properties of the estimator of the mode treatment effect. We can see that the proposed estimators follows the asymptotic normality but with the rate of convergence slower than the regular rate $\sqrt{N}$. The intuition is that, unlike the estimation of the average and the quantile treatment effect, the estimation of modes only uses a small portion of total observations which are around the modes. The usage rate of observations determines that the rate of convergence is slower than the regular rate $\sqrt{N}$.
To estimate the asymptotic variances, we define $\pi_{0}\left(X\right)\equiv P\left(D=1\mid X\right)$ to be the propensity score. The consistent variance estimators are \[ \hat{M}_{1}=\frac{1}{n}\sum_{i=1}^{n}\hat{f}_{Y\mid X, D=1}^{\left(2\right)}\left(\hat{\theta}_{1}\mid X_{i}\right), \] \[ \hat{M}_{0}=\frac{1}{n}\sum_{i=1}^{n}\hat{f}_{Y\mid X, D=0}^{\left(2\right)}\left(\hat{\theta}_{0}\mid X_{i}\right), \] \[ \hat{V}_{1}=\kappa_{0}^{\left(1\right)}\frac{1}{n}\sum_{i=1}^{n}\frac{\hat{f}_{Y\mid X,D=1}\left(\hat{\theta}_{1}\mid X_{i}\right)}{\hat{\pi}\left(X_{i}\right)}, \] \[ \hat{V}_{0}=\kappa_{0}^{\left(1\right)}\frac{1}{n}\sum_{i=1}^{n}\frac{\hat{f}_{Y\mid X,D=0}\left(\hat{\theta}_{0}\mid X_{i}\right)}{\hat{\pi}\left(X_{i}\right)}. \]
In this section, I propose the ML estimator of the mode treatment effect. The ML estimator can accommodate a large number of control variables, potentially more than the sample size. This flexibility will enable researcher to include as many control variables they consider important to make their identification assumptions more plausible. The key to implement ML methods is to replace the estimation of the conditional density function with the estimation of the conditional expectation. To begin with, the estimation of the conditional density function in the traditional kernel estimation is
\[ \hat{f}_{Y\mid D=1,X}\left(y\mid x\right) = \frac{\sum_{j=1}^{n}D_{j}K_{h}\left(y-Y_{j}\right)K_{h}\left(x-X_{j}\right)}{\sum_{j=1}^{n}D_{j}K_{h}\left(x-X_{j}\right)}. \] Notice that we can divide both the numerator and the denominator by $\sum_{j=1}^{n}K_{h}\left(x-X_{j}\right)$ to obtain
The numerator is an kernel estimator of $E\left[DK_{h}\left(y-Y\right)\mid X\right]$ and the denominator is an kernel estimator of the propensity score $E\left[D\mid X\right]=\pi\left(X\right)$. Hence, $\hat{f}_{Y\mid D=1,X}\left(y\mid x\right)$ is an estimator of $E\left[DK_{h}\left(y-Y\right)\mid X\right]/\pi\left(X\right)$. Then the marginal density estimator \[ \hat{f}_{Y_{1}}\left(y\right)=\frac{1}{n}\sum_{i=1}^{n}\hat{f}_{Y\mid D=1,X}\left(y\mid X_{i}\right) \] defined in the previous section can be interpreted as an estimator of \[ E\left[\frac{E\left[DK_{h}\left(y-Y\right)\mid X\right]}{\pi\left(X\right)}\right]=E\left[\frac{DK_{h}\left(y-Y\right)}{\pi\left(X\right)}\right]. \] Therefore, we can use the machine learning estimator of $E\left[\frac{DK_{h}\left(y-Y\right)}{\pi\left(X\right)}\right]$ as an estimator for $f_{Y_{1}}\left(y\right)$. We have successfully translate the estimation of the conditional density function into the estimation of the conditional expectation, which is the propensity score $\pi(X)$.
Here we pursue a little bit further to construct the Neyman-orthogonal score \citep*{chernozhukov2018double} for the robustness of the first-step estimation: \[ m_{1}\left(Z, y, \eta_{10}\right)=\frac{DK_{h}\left(y-Y\right)}{\pi_{0}\left(X\right)}-\frac{D-\pi_{0}\left(X\right)}{\pi_{0}\left(X\right)}E\left[K_{h}\left(y-Y\right)\mid X, D=1\right],\tag{4.1} \] where $Z=\left(Y, D, X\right)$ and $\eta_{0}=\left(\pi_{0}, g_{10} \right)$ with $g_{10}\left(X\right)\equiv E\left[K_{h}\left(y-Y\right)\mid X, D=1\right]$. Similary, the Neyman-orthogonal score for $f_{Y_{0}}\left(y\right)$ is
\[ m_{2}\left(Z, y, \eta_{20}\right)=\frac{\left(1-D\right)K_{h}\left(y-Y\right)}{1-\pi_{0}\left(X\right)}-\frac{\pi_{0}-D\left(X\right)}{1-\pi_{0}\left(X\right)}E\left[K_{h}\left(y-Y\right)\mid X, D=0\right], \tag{4.2} \] where $\eta_{20}=\left(\pi_{0}, g_{20}\right)$ with $g_{20}\left(X\right)\equiv E\left[K_{h}\left(y-Y\right)\mid X, D=0\right]$. Equation (4.1) and (4.2), to my best knowledge, should be the new results for density estimation. The Neyman orthogonality will make the estimation of the density functions more robust to the first-step estimation. Now I combine (4.1) and (4.2) with the cross-fitting algorithm \citep*{chernozhukov2018double} to propose the new estimator:
\[ \sqrt{nh^{3}}\left(\hat{\theta}_{1}-\theta_{1}^{*}\right)\overset{d}\to N\left(0,M_{1}^{-1}V_{1}M_{1}^{-1}\right), \] \[ \sqrt{nh^{3}}\left(\hat{\theta}_{0}-\theta_{0}^{*}\right)\overset{d}\to N\left(0,M_{0}^{-1}V_{0}M_{0}^{-1}\right). \]
As for the variance estimation, recall that the kernel estimator of $M_{1}$ in the previous section is $\hat{M}_{1}=N^{-1}\sum_{i=1}^{N}\hat{f}_{Y\mid D=1,X}^{(2)}(\hat{\theta}_{1}\mid x)$, where
\[ \hat{f}_{Y\mid D=1,X}^{(2)}\left(y\mid x\right) = \frac{\sum_{j=1}^{n}D_{j}K_{h}\left(y-Y_{j}\right)^{(2)}K_{h}\left(x-X_{j}\right)}{\sum_{j=1}^{n}D_{j}K_{h}\left(x-X_{j}\right)}. \] Notice that we can divide both the numerator and the denominator by $\sum_{j=1}^{n}K_{h}\left(x-X_{j}\right)$ to obtain
Observe that the numerator is an kernel estimator of $E\left[DK_{h}\left(y-Y\right)\mid X\right]$ and the denominator is an kernel estimator of the propensity score $E\left[D\mid X\right]=\pi\left(X\right)$. Hence, $\hat{f}_{Y\mid D=1,X}\left(y\mid x\right)$ is an estimator of $E\left[DK_{h}^{(2)}\left(y-Y\right)\mid X\right]/\pi\left(X\right)$. Hence, we can use the machine learning estimator of $E\left[\frac{DK_{h}^{(2)}\left(y-Y\right)}{\pi\left(X\right)}\right]$ as an estimator for $M_{1}$. We can also construct a DML estimator using the Neyman-orthogonal functional form
\[ \frac{DK_{h}^{(2)}\left(y-Y\right)}{\pi\left(X\right)}-\frac{D-\pi_{0}\left(X\right)}{\pi_{0}\left(X\right)}E\left[K_{h}^{(2)}\left(y-Y\right)\mid X, D=1\right] \] In step 1, we use machine learning methods to estimate $\pi_{0}(X)$ and $E[K_{h}^{(2)}\left(\hat{\theta}_{1}-Y\right)\mid X, D=1]$ using auxiliary sample $I_{k}^{c}$. In Step 2, we construct the DML estimator of $M_{1}$:
\[ \hat{M}_{1}=\frac{1}{K}\sum_{k=1}^{K}\sum_{i\in I_{k}}\frac{D_{i}K_{h}^{(2)}\left(\hat{\theta}_{1}-Y_{i}\right)}{\hat{\pi}\left(X_{i}\right)}-\frac{D_{i}-\hat{\pi}_{0}\left(X_{i}\right)}{\hat{\pi}_{0}\left(X_{i}\right)}\hat{E}\left[K_{h}^{(2)}\left(\hat{\theta}_{1}-Y\right)\mid X_{i}, D=1\right]. \] By the general DML theory \citep*{chernozhukov2018double}, $\hat{M_{1}}$ is a consistent estimator of $M_{1}$. Similarly, we can construct the DML estimators for $V_{1}$, $M_{0}$, and $V_{0}$ using the following table:
This paper studies the estimation and inference of the mode treatment effect, which has been ignored in the treatment effect literature compared to the estimation of the average and the quantile treatment effect estimation. I propose both kernel and ML estimators to accommodate a variety of data sets faced by researchers. I also derive the asymptotic properties of the proposed estimators. I show that both estimators are consistent and asymptotically normal with the rate of convergence $\sqrt{Nh^{3}}$.