EconBase
← Back to paper

Robust Inference in High Dimensional Linear Model with Cluster Dependence

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

34,223 characters · 11 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust Inference in High Dimensional Linear Model with Cluster Dependence

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 35 chars of source]

} \fi

abstractCluster standard error (Liang and Zeger, 1986) is widely used by empirical researchers to account for cluster dependence in linear model. It is well known that this standard error is biased. We show that the bias does not vanish under high dimensional asymptotics by revisiting Chesher and Jewitt (1987)'s approach. An alternative leave-cluster-out crossfit (LCOC) estimator that is unbiased, consistent and robust to cluster dependence is provided under the high dimensional setting introduced by Cattaneo, Jansson and Newey (2018). Since LCOC estimator nests the leave-one-out crossfit estimator of Kline, Saggio and Sølvsten (2019), the two papers are unfied. Monte Carlo comparisons are provided to give insights on its finite sample properties. The LCOC estimator is then applied to Angist and Lavy's (2009) study of the effects of high school achievement award and Donohue III and Levitt’s (2001) study of the impact of abortion on crime

{\it Keywords:} Cluster standard error, many regressors, residual analysis, linear regression, jackknife estimator

\spacingset{1.8}

Introduction

In linear regression models, it's common to assume the observations can be sorted into clusters where observations are independent across clusters but correlated within the same cluster. One method to conduct accurate statistical inference under this setting is to estimate the regression model without controlling for within-cluster error correlation, and then proceed to compute the so called cluster-robust standard errors (White, 1984; Liang and Zeger, 1986; Arellano, 1987, Hasen, 2011).These cluster-robust standard errors do not require a model for within-cluster error structure for consistency, but do require additional assumptions such as the number of cluster tends to infinity or the size of the cluster tends to infinity. The cluster-robust standard errors had became popular among applied researchers after Rogers (1993) incorporated the method in Stata. For a comprehensive methodology review, see Cameron and Miller (2015) or Imbens and Kolesar (2016).

As demonstrated by the Monte Carlo results in Bell and McCaffrey (2002), BM thereafter, the cluster-robust variance is biased in finite sample in general. They devised a bias-correction method to reduce the finite sample bias. Since then, a large body of research has emerged to investigate and address this small sample problem. Young (2019) did an excellent meta analysis that involves 53 experimental paper from the journals of the American Economic Association and concluded that use of conventional clustered/robust standarderror in these paper yields more statistically significant results.

We shall see in later section that a sufficent condition for this finite sample bias to vanish as sample size, $n$, grows is that the maximum leverage of the sample, $\max_i h_{ii}$, tends to zero as sample size grows. This condition will be satisfied if the ratio of the number of parameters and samplel size, $\frac{p}{n}$, tends to zero as $n$ tends to infinity (see Huber, 1981). As an example, this sufficient condition is met in White's proof (1984) given his primitive conditions on the design matrix.

The requirement that the expected value of leverage is zero asymptotically or equivalently $\lim_{n\rightarrow\infty}\frac{p}{n}=0$ is unattractive in many modern day applications as many datasets involve a large set of control variables where the number of control variables could grow at the same rate as the sample size. At the time this article was written, little researches have been done to relax this requriment in data with cluster dependence. One such papaer is Verdier (2018), it considers cluster-robust inference in fixed effect models under high dimensional asymptoitcs using subsetting. His method accomodates instrumental variables estimation at the cost of efficiency and restricted invertibility. D'Adamo (2019) also provided a consistent estimator under one-way clustering but the estimator also suffers from invertibility issues. Most other researches seem to focus exclusively on data with conditional heteroskedasticity. In particular, Cattaeno et al (2018) give inference methods that allow for many covariates and heteroscedasticity. They derive consistency of the OLS estimate under high dimensional asymptotics and provide a consistent variance estimator under the condition that maximum leverage is less than or equal to 1/2 as sample size tends to infinity. Kline, Saggio, Solvsten (2020) propose an unbiased leave-one-out estimator that is consistent and only requires maximum leverage is bounded away from 1. Jochmans (2021) develops an approximate version of the leave-one-out estimator under the asymptotic setting of Cattaneo et al (2018).

In this paper, our main goal is to extend CJN (2018)'s framework to accomdate cluster dependence and follow up on KSS (2020)'s remark to develop a leave-cluster-out estension of their estimator. The asymptotic properties of our estimator are established under CJN (2018)'s asymptotic framework and the two papers are unified in the process. We found that, after some additional efforts, most results in CJN (2018) generalize to data with cluster depedence under some additional mild assumptions on the unobserved errors and no additional assumptions on the design matrix are required.

The rest of the paper is structured as follows, section 2 motivates the leave-cluster-out estimator by providing bounds on the bias of the traditional cluster-robust standard error and showing that the bound collapses if maximum leverage tends to zero as sample size grows. Section 3 introduces the setup and notations and establishes the unbiasedness and consistency of the leave-cluster-out crossfit (LCOC) estimator. Section 4 provides monte carlo experiment results to access the finite sample performance of our estimator. Section 5 uses the LCOC estimator to Angist and Lavy's (2009) study of the effects of high school achievement award and Donohue and Levit's (2002) study of the the casual impact of legalized abortion on crime reduction. The paper ends with a short conclusion. Appendix contains the proofs to all the theoretical results.

Bias of Cluster-Robust Standard Error

We start by bounding the bias of the popular cluster-robust standard errors in order to motivate our leave-cluster-out estimator. The approach here will follow Chesher and Jewitt (1987), CJ thereafter. In particular, we will analyze the eigenstructure of the data and provide bounds in terms of eigenvalues and elements of the hat matrix. As we will see, the bias is bounded by the maximum leverage of the sample and therefore will be asymptotic unbiased if the maximum leverage vanishes asymptotically.

OLS Estimator

Let $ \{1,\ldots,n\}$ be set of indexes for the sample where $n$ is equal to the sample size. Let $G$ be a partition of $ \{1,\ldots, n\}$ where $|G|$ is the number of clusters. We reserve $g$ to denote the element of $G$ only and each $g$ is an ordered subset of $\{1,\ldots,n\}$. Let $g(i)$ returns the index value of its $i$th element, $g[i]$ be the cluster that individual $i$ belongs and $|g|$ be the size of the cluster $g$. Notations regarding the data structure are given below \setcounter{MaxMatrixCols}{40}

align*[align* omitted — 407 chars of source]
align*[align* omitted — 331 chars of source]

Consider the linear model

align*[align* omitted — 66 chars of source]

The OLS estimator is given by

align*[align* omitted — 156 chars of source]

We are interested in the variance matrix conditional on $\tilde{X}$

align*[align* omitted — 503 chars of source]

where $\tilde{x}_{ig}$ and $\tilde{x}_{jg}$ denote the $i$th and $j$th row of $\tilde{X}_g$ respectively.

The cluster robust covariace matrix replaces $\Omega_g$ with the plug-in estimator

align*[align* omitted — 79 chars of source]

Formally, the cluster-robust estimate is consistent if

align*[align* omitted — 214 chars of source]

The consistency of the above estimator is established by White (1984), Liang and Zeger (1986) and Hansen (2007) with varying degree of restrictiveness.

For the purpose of this section, reader might assume we follow the asymptotic setting of White (1984) unless otherwise stated. For the proof of White (1984) to go through, we need the estimator $\hat{u}_g\hat{u}_g'$ to be an asymptotic unbiased estimator of $\Omega_g$. While this is true under White's regularity conditions, this is not necessary true when considering high dimensional asymptotics as it violates one of the regularity condition of White (1984) that the probability limit of $\frac{1}{n}\tilde{X}'\tilde{X}$ to be finite and positive definite.

Finite Sample Bias of Cluster-Robust Variance

It would be informative to directly analyze the finite bias of the cluster-robust variance estimator. We will generalize CJ (1987)'s analysis to a covariance matrix with non-diagonal elements. The cluster-specific bias is given by

align[align omitted — 398 chars of source]

where we define $H_{g}=\tilde{X}_g(\tilde{X}\tilde{X})^{-1}\tilde{X}'$ and $H_{g,g}=\tilde{X}_g(\tilde{X}'\tilde{X})^{-1}\tilde{X}_g$ (see appendix for detail). It is interesting to point out here that if $\Omega=\sigma^2 I$, then $B_g=-H_{g,g}\sigma^2$ which is negative definite. This is in line with the view that the bias is downward in general. It would be interesting to examine under what conditions would $B_g$ gurantee to be positive or negative definite and we leave this for future researches.

We now consider the proprtionate bias term introducded by CJ (1987)

align[align omitted — 385 chars of source]

where $w$ is any non-zero vector that has the same dimension as $\beta$.

If we define $z_g= \tilde{X}_g\Big(\sum_{g\in G} \tilde{X}_g' \tilde{X}_g\Big)^{-1}w$, then the above can be explicitly written as

align*[align* omitted — 257 chars of source]

which is a ratio of sum of quadratic forms. The above term is further bounded by two Rayleigh-quotient-like quantities $\max_{g\in G}\frac{z_g'B_g z_g}{z_g'\Omega_g z_g}$ and $\min_{g\in G}\frac{z_g'B_gz_g}{z_g'\Omega_g z_g}$. Applying result on ratio of quadratic forms from Rao p74 (1972) to these two quantities gives us the theorem below (see appendix for full proof).

thmThe proptionate bias is bounded by \begin{align*} \lambda_{\min}(B_{g_{\min}}\Omega_{g_{\min}}^{-1})\leq pb(V_{cluster})\leq \lambda_{\max}(B_{g_{\max}}\Omega_{g_{\max}}^{-1}), \end{align*} where $g_{\max}$ and $g_{\min}$ denote the cluster with greatest cluster specific bias and the cluster with smallest cluster specific bias respectively.

Our final result generalizes CJ (1987)'s insight to errors with cluster dependence.

thmIf $B_g$ is positive definite (or negative definite) for all $g$, $|g|=O(1)$, $\Omega_g=O(1)$ and $\lim_{n\rightarrow \infty} \max_i h_{ii}= 0$ , then \begin{align*} \lim_{n\rightarrow 0}pb(\widehat{Var}_{cluster})= 0 \end{align*}

In words, if the design is approximately balanced in that sense that the maximum leverage vanishes asymptotically and the size of the cluster is bounded, then the usual cluster-robust standard error (LZ, 1987) will be asymptotically unbiased. Theorem 2.2. motivates the use of an unbiased estaimtor in finite sample to avoid relying on the asympotic assumptions to kill the bias in large sample.

General Framework

Suppose $\{(y_i,x_i',w_i'):1\leq i \leq n\}$ is generated by

align[align omitted — 111 chars of source]

for $i=1,\ldots,n$ where $\alpha' = (\sum^n_{j=1}\text{E}[x_jw_j'])(\sum^n_{j=1}\text{E}[w_jw_j'])^{-1}$ and $v_i$ is the deviation of $x_i$ from the population linear projection.

We will refer equation (1) as the primary model and equation (2) as the auxiliary model. Our goal is to conduct valid inference on $\beta$. Note that the auxiliary model could be mis-specified in the sense that $ \alpha_n'w_i\neq \text{E}(x_i|\mathcal{W}_n)$. We assume $\text{E}(u_i|\mathcal{X}_n,\mathcal{W}_n)=0$ to take advantage of the unbiasedness of our variance estimator. This contrasts to CJN (2018)'s framework where they also allow for an asymptotic negligible amount of mis-specification errors in the primary model.

We will recyle the OLS notations used in section 2.1. Note that

align*[align* omitted — 219 chars of source]

Define $H_{\tilde{X}}=\tilde{X}(\tilde{X}'\tilde{X})^{-1}\tilde{X}'$ and $H_{W}=W(W'W)^{-1}W'$ as the hat matrices generated by $\tilde{X}$ and $W$ respectively. We have $M=I-H_{W}$and $\tilde{M}=I-H_{\tilde{X}}$. Then $M\bm{y}$ and $\bm{\hat{v}}=MX$ are the residuals after regressing $Y$ and $X$ on $W$ respectively. Define $H_{\bm{\hat{v}}}=\bm{\hat{v}}(\bm{\hat{v}}'\bm{\hat{v}})^{-1}\bm{\hat{v}}'$ as the hat matrix generated by $\bm{\hat{v}}$.

Under our asymptotic framework, it would be convient to re-state OLS estimator in the following format

align*[align* omitted — 92 chars of source]

which is just an applciation of the Frisch–Waugh–Lovell theorem.

{ Assumption 1}

quote$\max_{g\in G} |g|=O(1)$, where $|g|$ is the cardinality of $g$ and where $G=\{g_1,\ldots,g_{|G|}\}$ is a partition of $\{1,\ldots,n\}$ such that $\{(u_{i},V_{i}{'}\}: i \in g\}$ are independent across $g$ conditional on $(\mathcal{X}_n,\mathcal{W_n})$.

The assumption defines the sampling environment. It is a modified version of CJN (2018, assumption 1) that allows for within-cluster dependence that is common in panel data analysis. The asymptotics employed for the cluster structure is the same as the ones used in White (1984) where each cluster's size is bounded and $G$ is proportional to $n$.

{ Assumption 2}

quote$P[\lambda_{\min}(\sum^n_{i=1}w_{i}w_{i}')>0)>0]\rightarrow 1$, $\overline{\lim}_{n\rightarrow\infty}\frac{k}{n}<1$ and { \begin{align*} \max_{1\leq i\leq n}\Big\{&E[u_i{^4}|\mathcal{X}_n,\mathcal{W}_n]+E[\|V_i\|^4|\mathcal{W}_n]+\frac{1}{\lambda_{\min}(\Omega)}+\frac{1}{\lambda_{\min}(E[\frac{1}{n}\sum^n_{i=1}\tilde{V}_i\tilde{V}_i'|\mathcal{W}_n])}\Big\}=O_p(1) \end{align*}} where $\tilde{\Sigma}=\frac{1}{n}\sum_{g\in G}\sum_{i,j\in g}\tilde{V}_{i}\tilde{V}_{j}{'}E[u_{i}u_{j}|\mathcal{X}_n,\mathcal{W}_n]$ and $\tilde{V}=\sum^n_{j=1}M_{ij}V_j$.

First condition prevents the elements of the design matrix of the nuisance covariates from being too close to singularity. This is a generalization of the uniformly nonsingularity assumption (White, 1984, p22) where it allows the rank of $\sum^n_{i=1}w_{i}w_{i}'$ to grow as $n$ grows. This assumption is not restrictive as any linear dependent nuisance covariates can be dropped without impacting the OLS estimate. Second condition allows the number of parameters to be estimated to grow in line with sample size as long as we have slightly more than one observations per parameter. Third condition are moment conditions that restrict distributions of $u_{i}$ and $V_{i}$ from the main model and auxiliary model respectively. Note that the assumption differs from CJN (2018)'s assumption 2 in that we restrict the eigenstructure of $\Omega$ to control for the within-cluster correlations.

{ Assumption 3}

quote$E[u_{i}|\mathcal{X}_n,\mathcal{W}_n]=0\ \forall i$, $\chi=\frac{1}{n}\sum^n_{i=1}\text{E}(\|\underbrace{\text{E}[v_i|\mathcal{W}]}_{Q_i}\|^2)=O(1)$ and $\frac{\max_i\|\hat{v}_{i}\|}{\sqrt{n}}=o_p(1)$.

First condition is the usual exogeneity condition. This is necessary for the unbiasedness of our leave-cluster out estimator. One might relax this assumption to allow an asymptotic negligible amount of mis-specification bias in the primary model (see CJN 2018). Our estimator would then lose its unbiasedness but remain consistent. Second condition restricts the amount of inaccuracy permitted in the linear prediction of the conditional mean in the auxiliary model. The third condition is a necessary condition for the maximum leverage of design matrix of the second stage regression to vanish asymptotically.

{ Assumption 4}

quote\begin{enumerate}[label=\roman*.] • $\text{Pr}(\min_g\text{det}(\tilde{M}_{g,g})>0)\rightarrow 1$$\frac{1}{\min_g\lambda_{\min}(\tilde{M}_{g,g})}=O_p(1)$$\frac{\sum^n_{i=1}||\tilde{Q}_{i}||^4}{n}=O_p(1)$, where $\tilde{Q}_i=\sum^n_{j=1}M_{ij}Q_i$$\frac{\max_i||\mu_{i}||}{\sqrt{n}}=o_p(1)$ \end{enumerate}

Assumption 4 is a set of conditions needed for variance estimation. First and second conditions are there to control perfect and near perfect collinearity. The first one allows the estimator to exist in large sample with high probability while the second one prevents the variance of the estimator to blow up in large sample due to near perfect collinearity. Third assumption is needed to bound the foruth moment of the error in the auxiliary model. This in effect allows us to bound $\frac{1}{n}\sum^n_{i=1}\hat{v}_i^4$ which also contributes to the variance of the estimator. The last assumption restricts the amount of noise in the level variable $y_i$ that could come from the conditional mean $\mu_i$. This is needed because the crossfit estimator uses $y_i$ as a proxy for the unobserved error $u_i$.

Theoretical Results

Our first result extends Cattaneo's asymptotic normality result in high dimensional linear model to data with cluster dependence.

thmSuppose Assumptions 1-3 hold and $\lambda_{\min}(\sum^n_{i=1}w_{i}w_{i}')>0,\lambda_{\min}(\hat{\Gamma})>0$. Then, \begin{align*} \Omega^{-1/2}\sqrt{n}(\hat{\beta}-\beta)=\hat{\Gamma}^{-1}S\xrightarrow[]{d}\mathcal{N}(0,I),\ \ \ \Omega=\hat{\Gamma}^{-1}\Sigma\hat{\Gamma}^{-1}, \end{align*} where $$\hat{\beta}=1\{\lambda_{\min}(\hat{\Gamma})>0\}\hat{\Gamma}^{-1}\Big(\frac{1}{n}\sum_{1\leq i\leq n}\hat{v}_{i}y_{i}\Big),\ \ \ \hat{\Sigma}=\frac{1}{n}\sum_{g\in G}\sum_{i,j\in g}\hat{v}_{i}\hat{v}_{j}{'}E[u_{i}u_{j}|\mathcal{X}_n,\mathcal{W}_n],$$ \begin{align*} \hat{\Gamma}=\frac{1}{n}\sum_{1\leq i \leq n}\hat{v}_{i}\hat{v}_{i}' \ , \ S=\frac{1}{\sqrt{n}}\sum_{1\leq i\leq n}\hat{v}_{i}u_{i}\ \ and \ \ \ \hat{v}_{i}=\sum_{1\leq j\leq n}M_{ij}x_{j}. \end{align*}

We can now introduce our leave-cluster-out crossfit (LCOC) estimator.

defn[Leave-Cluster-Out estimator] \begin{align*} \hat{\Sigma}^LCOC=\frac{1}{n}\sum_{g\in G}\sum_{i,j\in g}\hat{v}_{i}\hat{v}_{i}{'}(y_{i}\hat{u}_{-g,j}+\hat{u}_{-g,i}y_{j}) \end{align*} where \begin{align*} \hat{u}_{-g,i}=\sum^n_{j=1}M_{ij}(y_{j}-x_{j}'\hat{\beta}_{-g}) \end{align*} and $\hat{\beta}_{-g}$ is the OLS estimator computed by excluding observations in the cluster $g$.

We now state the two results about the LCOC estimator.

thm[Unbiasedness] Suppose $\hat{\Sigma}^\text{LCOC}$ exists, then \begin{align*} E[\hat{\Sigma}^LCOC|\mathcal{X}_n,\mathcal{W}_n]=\Sigma. \end{align*}
thm[Consistency] Suppose assumption 1-4 holds, \begin{align*} \hat{\Sigma}^LCOC=\Sigma+o_p(1). \end{align*}

Numerical Results

We consider a setup that emulates our empirical example,

align*[align* omitted — 62 chars of source]

where $\beta =0.5$, $(x_{it},w_{it})\sim_{iid}N(0,1)$, $\gamma_t\sim U[-0.5,0.5]$, $\epsilon_{it}=\Big(0.8\epsilon_{it-1}+0.2u_{it}\Big)|x_{it}|$ and $u_{it}\sim N(0,1)$.

We generate a simple of $N$ individuals with $T$ periods. We perform a monte carlo simulation of 1,000 repetitions with the above setup. Note that model is both serially correlated and heteroskedastic. This specification gives a ration of parameter to observation of $\frac{181}{1000}=18.1\%$.

We find that the average bias of cluster-robust estimator is -0.1909 and the average bias of BM estimator is 0.0785. This supports the view that the LZ estimator is biased downward while the BM jackknife type estimator is biased upward in general. Our leave-cluster-out is unbiased so there is little surprise that the average bias is closed to zero. In terms of variance, the LZ estimator is the most precise (0.2438) while BM comes second (0.6038) and ours comes last (0.7552) . This illustrates the bias-variance trade-off and is expected because the BM estimator leaves out observations and LCOC estimtor uses the level variable which could be noisy. In terms of MSE performance, LZ comes out the ahead in this experiment. However, the differences among the three are small and all are of the same order of mangitude. Note that LCOC estimator is the only consistent estimator here so it will eventually outperform the other two by increasing the sample size of the monte carlor experiment. In terms of rejection rate, under the null that $\hat{\beta}=0.5$, our estimator has the best size control. Note that the t-statistic constructed using the LZ estimator (BM estimator) under-rejects (over-rejects).

figure[figure omitted — 141 chars of source]

Empirical Illustration

Angrist and Lavy (2009)

Angrist and Lavy (2009), AL thereafter, analyze the effect of high stakes high school achievement using cash incentives experiment. Their identication of treatment effect is given by the general model below

align*[align* omitted — 146 chars of source]

where $i$ indexes students, $j$ indexes schools, $\Lambda$ could be logistic function or identity function (OLS) depending the specification. Covariates include the treatment dummy (school level), $z_j$, a vector of school-level controls, $\bm{X}_j$, a vector of individual controls $\bm{W}_i$. and lagged test scores $\delta_q$.

AL (2009) assert that the bias of the LZ standard error is biased downward and use the Bell and Mcaffrey's (2009) Jackknife estimator in an attempt to address the finite sample bias problem of LZ (1987)'s estimator. Two potential concerns here are i) that there is no guarantee that bias of LZ is downward as seen in our bias analysis excersis and ii) that the BM standard error is also biased in finite sample. Thus, applying our cross-fit estimator here to their models would serve as excellent robustness checks to allievate the aforementioned concerns.

We will focus on the linear specifications where $\Lambda$ is the identity function and reproduce the AL's results in table 2's panel A. Two causual discoveries were made by AL (2009) from this table 1) they suggest that there is evidence that the Achievement Awards program increased Bagrut rates in 2001 and 2) the estimated treatment effect comes mainly from girls as suggested by the gender-specific regressions. However, they do note that most of their significant results are only “marginally significant”. A close examination of the implied p-value of the BM standard error seems to suggest that they mean significance level around or below 10% level.

Inferences based on our leave-cluster-out estimator support their conclusion with some additional insights. The two key specifications in this table are (1) SC + Q+ M and (3) SC + Q+ M + P. The p-value of the estimated coefficients for these two equations are significantly lower than their BM counter parts we have 0.0165 vs 0.0208 and 0.0274 vs 0.1055 respectively. One potential reason for the BM being more conservative is that there are manys discrete covariates. Eventhough the sample average of leverage points is low (low parameter to observation ratio), the regression design is still relatively imbalanced. This could result in the BM estimator being biased upward as shown in the simluation.

Donohue and Levitt (2001)

Donohue and Levitt (2001) concludes that legalized abortion has contributed significantly to crime reduction. They ran the following regression

align*[align* omitted — 109 chars of source]

where the left-hand-side variable is the logged crime rate per capita, ABORT$_{st}$ is the effective abortion rate for a given state, year and crime category, $X$ is a vector of state-level controls that includes prisoners and police per capita, a range of variables capturing state economic conditions, lagged state welfare generosity, the presence of concealed handgun laws, and per capita beer consumption. $\gamma_s$ and $\lambda_t$ represent state and year fixed effects. The cluster will be define at the state level. Clustering at the state level can be seen as a way to account for serial correlation in the sample. However, the cluster-robust standard error is biased due to the presence of serial correlation. If we assume textbook asymptotic setting, then the LCOC estimator can be seen as a finite sample bias correction to the cluster-robust standard error.

In this model, it is reasonable to allow for high dimensional asymptotics due to the inclusion of many state-level characteristics. The state-specific fixed effects are not identified in the leave-cluster-out sample, so we get around this by applying a within-transformation to get rid of the state-specific fixed effects. The baseline model is thus

align*[align* omitted — 149 chars of source]

We estimate the model above using OLS and compute the LZ, LCOC and BM estimator. The results are essentially replication of Table IV in DL (2009). In line with their results, the coefficients $\hat{\beta}$ is negative for all crime types. In terms of the estimated variances, we see that the LZ estimator is the smallest across the board while the BM estimator is the largest across the board. The LCOC estimator is then exactly in the middle across the board. In terms of significance level, nothing is changed when we switch from the LZ to LCOC estimator. However, the coefficient for muder crime is no longer significant at 1% level when we switch from LZ to BM estimator. This suggests that the result of DL (2009) might be less robust for more serious crimes. Overall, the original results remain relatively insensitive to the choice of estimator used. If we look at the left subfigure in figure 3, the histogram of sample leverage points for this specification is very balanced where the average value of leverage points is low (red line) and the spread is also small. This suggests that the finite sample biases of the two estimators would be likely small and thus leading to small differences across the three estimators.

Next, we consider a high dimensional specification that assumes the impact of the controls is time-varying. The model is given by

align*[align* omitted — 151 chars of source]

this gives a ratio of parameter to observations equals to $\frac{109}{624}\approx 17.5\%$ which is much larger than the ratio of the baseline model ($\frac{21}{624}\approx 3.4\%$).

The cofficient of interest $\hat{\beta}$ remains negative for all three crime types in this specification but the magnitudes of them are reduced by a non-trivial amount. This highlights the sensitivity to the choice of controls and supports the finding of Belloni, Chernozhukov and Hansen (2014). The estimated variances are larger than the baseline in across the board with again BM (LZ) being largest (smallest) across the board. Note that empirical distribution of leverage points in the levitt sample with high dimensional specification is quite similar to the empirical distribution of leverage points in this simulated sample. Thus, it seems reasonable that we observe similar pattern to that of the monte carlor experiment here. However, $\hat{\beta}$ is only highly signficant for property crime across the three estimators. Violent crime is no longer significant under 1% level for both LCO and BM. Murder crime is no longer significant under 10% level for BM but manages to stay significant under 10% level for both LZ and LCOC. This casts doubts into whether there is indeed a causal relationship between abortion rate and more serious crimes like murder and violent crimes. One potential explanation here is that the underlying unobserved dependence and heteroskedasticity are different across the crime types which leads to different inferential results.

Conclusion

To motivate the use of our bias corrected cluster-robust variance estimator, we derive the explicit bounds on the finite-sample bias of the cluster-robust standard error. The results show that the cluster robust standard error will be asymptotically unbiased if the maximum leverage point vanishes asympoticallty. Following KSS (2020)'s remark, we construct an unbiased variance estimator that is robust to cluster dependence under high dimensional setting introduced by CJN (2018). This estimator can be seen as a bias-corrected cluster-robust variance estimator (White, 1984; Liang and Zeger, 1986). Monte Carlo results show that the LCOC estimator is unbiased but could be less precise depending on the empirical hat matrix. As empirical illustrations, the leave-cluster-out estimator is applied to Angist and Lavy's (2009) study of the effects of high school achievement award and Donohue and Levit's (2002) study of the the casual impact of legalized abortion on crime reduction.

\nocite{https://doi.org/10.48550/arxiv.1806.07314} \nocite{10.2307/1911269} \nocite{doi:10.1080/01621459.2017.1328360} \nocite{jochmans2020} \nocite{10.1093/biomet/73.1.13} \nocite{https://doi.org/10.3982/ECTA16410} \nocite{10.1257/aer.99.4.1384} \nocite{ANATOLYEV201820} \nocite{10.1162/00335530151144050} \nocite{white1984} \nocite{CHESHER1991153} \nocite{https://doi.org/10.1111/j.1468-0084.1987.mp49004006.x} \nocite{RePEc:eee:econom:v:165:y:2011:i:2:p:137-151} \nocite{BM2002} \nocite{10.1093/qje/qjy029} \nocite{10.1162/rest_a_00807} \nocite{Huber2011} \nocite{CM2015} \nocite{10.1162/REST_a_00552} \nocite{RePEc:tsj:stbull:y:1994:v:3:i:13:sg17} \nocite{10.1257/jep.28.2.29}