Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,891 characters · 6 sections · 101 citation commands
Optimal Categorical Instrumental Variables
\ifpublic \fi
\pagenumbering{gobble} \pagenumbering{arabic}
Optimal instrumental variable estimators aim to improve statistical precision by maximizing the strength of the first stage. In the absence of functional form assumptions, optimal instruments need to be nonparametrically estimated which can introduce bias from over-fitting. In response to these challenges, a growing literature considers complexity-reducing assumptions on the first stage reduced form that -- when leveraged appropriately -- allow for second stage estimators with the same asymptotic variance as the infeasible oracle estimator that presumes knowledge of the optimal instrument. Increasingly popular in practice is the post-lasso IV estimator of belloni2012sparse that assumes approximate sparsity of the first stage reduced form gilchrist2016something, dhar2022reshaping.\footnote{Informally, approximate sparsity presumes that a slowly increasing unknown subset of instruments suffices to approximate the optimal instrument relative to the reduced form estimation error.} However, while approximate sparsity may be a well-suited assumption for economic settings with continuous instruments, it is often ill-suited when instruments are categorical. Simulations with categorical instruments in angrist2020machine and in this paper highlight that even in settings with many more observations than instruments, lasso-based IV estimators can have substantially worse finite sample behavior than even two-stage least squares (TSLS).\footnote{Recently, kolesar2023fragility formally characterize disadvantages of applying sparsity assumptions to settings with categorical variables. While their discussion focuses on settings with categorical controls, analogous arguments apply to sparsity assumptions for categorical instruments.}
This paper proposes a new optimal instrumental variable estimator for settings with a large number of categorical instruments. To approximate the practical settings in which the number of observations per category is small, I characterize the proposed categorical instrumental variable estimator (CIV) in asymptotic regimes that allow the expected number of observations per category to grow at arbitrarily small polynomial rate with the sample size. To obtain root-$n$ normality in these challenging “moderately many instruments” settings, I consider a regularization assumption designed specifically for categorical instruments: Fixed finite support of the optimal instrument. When the cardinality of the support of the optimal instrument is known, I show that CIV achieves the same asymptotic variance as the infeasible oracle two-stage least squares estimator that presumes knowledge of the optimal instrument, and is semiparametrically efficient under homoskedasticity. Further, under-specifying the number of support points maintains asymptotic normality but results in efficiency loss.\footnote{An implementation of CIV is provided in the R package \ifpublic \href{https://www.thomaswiemann.com/civ}{civ}\fi\ifanonymous civ\fi available on CRAN.}
The key idea of the categorical instrumental variable estimator is to leverage a latent categorical variable with fewer categories that achieves the same population-level fit in the first stage. Under the assumption that the support of the latent categorical variable is fixed with finite cardinality, it is possible to estimate a mapping from the observed categories to the latent categories. This estimated mapping can then be used to simplify the optimal instrumental variable estimator to a finite dimensional regression problem. Asymptotic properties of the CIV estimator then follow if the first-stage mapping can be estimated at a sufficiently fast rate. I provide sufficient conditions for estimation of the mapping at exponential rate using a $K$-Conditional-Means ($K$CMeans) estimator. The proposed $K$CMeans estimator is exact and computes very quickly with time polynomial in the number of observed categories, thus avoiding heuristic solution approaches otherwise associated with $K$Means-type problems.\footnote{$K$CMeans also has applications outside of instrumental variable settings. For example, $K$CMeans can easily be combined with the double/debiased machine learning framework of chernozhukov2018double. An implementation of $K$CMeans is provided in the R package \ifpublic \href{https://www.thomaswiemann.com/kcmeans}{kcmeans}\fi\ifanonymous kcmeans\fi \href{https://www.thomaswiemann.com/kcmeans} available on CRAN.}
The focus on categorical instrumental variables is motivated by the many examples in empirical economics, including leading examples in the many and weak instruments literature, featuring categorical instruments. In their analysis of returns to education, for example, angrist1991does consider interactions between quarter of birth indicators and year and place of birth indicators, resulting in an instrument with 180 categories. Extensions of their design consider the fully saturated first stage of all interactions between the three sets of indicators resulting in 1530 categories, each representing a unique combination of quarter, year, and place of birth mikusheva2020inference, angrist2020machine. More recently, a large empirical literature uses so-called examiner fixed effects as instruments kling2006incarceration,maestas2013does,aizer2015juvenile,dobbie2018effects,bhuller2020incarceration,agan2023misdemeanor. For example, dobbie2018effects use judge identities as instruments for pre-trial detention to analyze its effect on the probability of conviction.\footnote{Although judge identity is the most widely adopted examiner fixed effect instrument, other quasi-randomly assigned decision makers are also frequently considered. agan2023misdemeanor, for example, leverage random assignment of assistant district attorneys to non-violent misdemeanor cases. See chyn2024examiner for a comprehensive overview of empirical analyses using examiner fixed effect instruments.}
The empirical literature on examiner fixed effects also motivates the analysis of the proposed CIV estimator under “moderately many instruments” asymptotics. In these economic analyses using examiner fixed effects as instruments, it is common practice to estimate first-stage fitted values via a leave-one-out procedure -- i.e., to compute the fitted value of a particular decision without directly using the decision itself. These jackknifing procedures are often motivated as a solution to a many instruments problem arising due to only few observations per decision maker chyn2024examiner. However, unlike the econometric literature on many instruments that often focuses on asymptotic regimes in which the number of observations per instrument is fixed bekker1994alternative, practitioners appear to primarily consider settings that are approximated by regimes in which the number of observations per instruments grows slowly. This follows from the combined use of jackknifed fitted values and conventional (heteroskedasticity-robust or clustered) standard errors, the latter resulting in invalid inferential statements under the traditional “many instruments” asymptotic regime but leading to correct inference in regimes with a growing number of observations per instrument (even if these are growing very slowly) chao2012asymptotic.
In the moderately many instrument setting considered in this paper, traditional jackknife IV estimators (JIVE) phillips1977bias,angrist1999jackknife are first-order equivalent to CIV. Indeed, the results of chao2012asymptotic imply semiparametric efficiency of JIVE under weaker assumptions than I leverage for CIV as the former does not rely on a finite support restriction. Yet, as illustrated in the simulations and the application, JIVE does not uniformly outperform CIV: JIVE (and its IJIVE ackerberg2009improved and UJIVE kolesar2013estimation variants) can suffer from high instability and unnecessarily large standard errors -- a well-documented practical caveat of jackknife estimators that has been a topic of substantial discussion in the econometrics literature davidson2006case,ackerberg2006comment. In contrast, CIV is nearly as stable as TSLS while exhibiting substantially less bias and achieving better coverage.
The key difference between jackknife-based estimators with categorical instruments and CIV is how the bias from a large number of instruments is controlled. For a heuristic illustration, it is useful to consider the expansion
which plays a key role in the convergence of JIVE and CIV estimators. Here, $Z_i$ is the (categorical) instrument of the $i$th observation, $U_i$ is a structural residual, and $m_0(Z_i)$ and $\hat{m}_n(Z_i)$ are the true and estimated optimal instrument, respectively. The first term on the right hand side in (ref) is well-behaved under standard regularity conditions and converges in distribution to a Normal random variable. The second term can potentially diverge due to correlation between the residuals $U_i$ and the estimation error $\hat{m}_n(Z_i) - m_0(Z_i)$. For example, donald2001choosing highlight that when $\hat{m}_n$ is a first-stage linear regression estimator (as in TSLS), divergence can occur if the number of instruments grows faster than $\sqrt{n}$. JIVE resolves this problem of potential divergence by enforcing independence between the residuals and estimation errors for any particular observation via leave-one-out estimation. Under this independence, it then suffices for $\hat{m}_n$ to be consistent for $m_0$.\footnote{Note that leave-one-out estimation of the first stage does not immediately imply independence of the estimation error and the residuals if observations are correlated, a point previously highlighted in frandsen2023cluster. frandsen2023cluster propose a jackknife IV estimator for examiner fixed effect designs in which observations are clustered at a level finer than that of the examiner (e.g., at the shift level). However, if observations are arbitrarily correlated at the level of the examiner -- as suggested by current empirical literature that often clusters at the examiner level -- no applicable jackknife-based IV estimator exists. The regularization approach presented in this paper may be an appealing solution to the problem of correlated examiner decisions since the key concentration inequalities leveraged in the proof extend to weakly dependent data. Formally investigating the robustness of CIV to weakly dependent judge decisions may thus be a potentially fruitful topic of future research.} An alternative solution is to ensure that the $\hat{m}_n$ is not just consistent but to also control the estimation errors via regularization. Sufficiently tight control of estimation errors then allows the second term on the right hand side in (ref) to vanish without relying on leave-one-out estimation (or alternative sample splitting procedures). For example, belloni2012sparse show that under approximate sparsity, a lasso-based estimator of the optimal instrument is sufficiently well controlled to allow for root-$n$ normal inference. Similarly, in this paper, I show that CIV provides sufficiently tight control on the estimation errors for the second term in (ref) to vanish without the need to explicitly enforce independence via jackknifing.\footnote{Sample splitting and regularized estimation of the optimal instrument are not exclusive and can often be used jointly to weaken conditions for the control of estimation errors. For example, belloni2012sparse show that sample splitting allows for a denser representation of the optimal instrument. hk2014jive combine JIVE and ridge regularization to admit estimation with more instruments than observations in non-sparse regimes.}
The advantage of the proposed CIV estimator over alternative optimal IV estimators is that the regularization assumption employed for tight control of the estimation errors places restrictions on the data generating process that have straightforward economic interpretations when applied to categorical instruments such as examiner fixed effects. In particular, the regularization assumption I consider presumes existence of an unobserved combination of categories (or: examiners) that fully captures relevant variation from the instruments. In the application of dobbie2018effects, for example, this regularization assumption presumes that the leniency of judges is a discrete (unobserved) type. For the restriction of the optimal instrument to two support points, this would correspond to categorizing judges as either being “strict” or “lenient.” While semiparametric efficiency of CIV only follows if this regularization assumption holds exactly, I show that CIV remains root-$n$ normal even if judges are categorized into fewer “pseudo” types than their true types (e.g., if there are also “moderate” judges). CIV thus provides a solution to a “moderately many categorical instruments” problem under conditions that economists can readily understand and discuss.
Related Literature. The paper primarily draws from and contributes to three strands of literature. First, the literature on many instruments that develops estimators robust to asymptotic regimes in which the number of instruments is proportional to the sample size phillips1977bias, bekker1994alternative, angrist1995split, angrist1999jackknife,chao2005consistent, hansen2008estimation, chao2012asymptotic, hausman2012instrumental, kolesar2013estimation. Most closely related is bekker2005instrumental who provide limiting distributions of two-stage least squares (TSLS), limited information maximum likelihood (LIML), and heteroskedasticity-adjusted estimators under group asymptotics that consider replications of categorical instruments with a constant number of observations per category. In the less stringent asymptotic regimes I consider in this paper where the number of categories grows at a slower rate than the sample size, their results imply first-order equivalence of the LIML and the oracle IV estimator in the presence of heteroskedasticity when observations are equally distributed across categories and effects are constant.\footnote{Note also that Lemma 6.A of donald2001choosing implies that LIML using categorical instruments achieves first-order oracle equivalence when the number of categories grows below the sample rate and causal effects are constant.} Despite the favorable statistical properties of LIML in settings with categorical instrumental variables and homogeneous effects, its application to causal effects estimation in economics is limited by its strong reliance on constant effects in the linear IV model. kolesar2013estimation shows that under the nonparametric causal model of imbens1994identification, the LIML estimand cannot generally be interpreted as a positively (weighted) average of causal effects. In the terminology of blandhol2022tsls, LIML thus does not generally admit a weakly causal interpretation. In contrast, the proposed CIV estimator falls in the class of two-step estimators of kolesar2013estimation and therefore admits a weakly causal interpretation in the presences of unobserved heterogeneity.
Second, I draw from the literature on optimal instrumental variable estimators. Optimal instruments are conditional expectations that -- in the absence of functional form assumptions -- can be nonparametrically estimated amemiya1974multivariate,chamberlain1987asymptotic,newey1990efficient. newey1990efficient considers approximation of optimal instruments using polynomial sieve regression and characterizes the growth rate of series terms relative to the sample size that allows for root-$n$ consistency. In the setting of categorical variables, the restrictions imply that the number of categories should grow slower than root-$n$ to avoid the many instruments bias. CIV contributes to the literature on optimal IV estimation that leverages regularization assumptions on the first stage to allow for a larger number of considered instruments. In homoskedastic linear IV models, donald2001choosing propose instrument selection criteria, chamberlain2004random consider regularization via a random coefficient assumption, and okui2011instrumental suggests first stage estimation via Ridge regression ($\ell_2$ regularization). In linear IV models with heteroskedasticity, carrasco2012regularization consider $\ell_2$ regularization (including Tikhonov regularization) and provide conditions for asymptotic efficiency of the resulting IV estimator in settings when the number of instruments is allowed to grow at faster rate than the sample size. belloni2012sparse apply the lasso and post-lasso ($\ell_1$ regularization) to estimate optimal instruments in the setting with very many instruments. The authors provide sufficient conditions for the asymptotic efficiency of the resulting IV estimator, most notably, an approximate sparsity assumption, which presumes that a slowly increasing unknown subset of instruments suffices to approximate the optimal instrument relative to the reduced form estimation error. A common theme in the regularization approaches of these previous approaches is shrinkage of the first stage coefficients to zero. In the setting of categorical instruments, this corresponds to existence of one large latent base category (i.e., the constant) and only a few small deviating latent categories. Settings in which differing latent categories are approximately proportional are not admitted in these shrinkage-to-zero approaches as observed categories cannot be arbitrarily merged. CIV complements these existing optimal instrument estimators by leveraging an alternative regularization assumption that admits approximately proportional latent categories via arbitrary combination of observed categories.
Finally, I draw from the literature on estimation with finite support restrictions in longitudinal data settings including hahn2010panel, bonhomme2015grouped, bester2016grouped and su2016identifying. hahn2010panel show that finite support assumptions substantially decrease the incidental parameter problem associated with increasingly many fixed effects. bester2016grouped consider grouped fixed effects when the grouping is known. bonhomme2015grouped and su2016identifying, among others, consider settings with unknown groups and parameters. I adapt the $K$Means fixed effects estimator of bonhomme2015grouped for estimation of the optimal categorical instrumental variable. In doing so, I make two contributions to the theoretical analysis of $K$Means. First, I construct and characterize a $K$-Conditional-Means estimator suitable for cross-sectional regression. Leveraging arguments from fisher1958grouping and wang2011ckmeans, this estimator has the advantage of achieving global in-sample optimality in time polynomial in the number of observed categories, by-passing the NP-hard problem of $K$Means in multiple dimensions. I thus do not need to abstract away from optimization error as is usually necessary in applications of $K$Means.\footnote{Computational solutions to $K$Means applications in longitudinal data settings are an active literature. See, in particular, the ongoing work of chetverikov2022spectral and mugnier2022simple.} Second, I show that the $K$CMeans estimand can serve as an approximation to a conditional expectation function when the number of allowed-for support points is under-specified. In the instrumental variable setting considered in this paper, this property is leveraged to allow for root-$n$ consistent estimation given only a lower-bound on the support points of the optimal instrument.
Outline. The remainder of the paper is organized as follows: Section (ref) introduces the CIV estimator in the simple setting without additional covariates and discusses its motivating assumptions. Section (ref) states the CIV estimator with covariates and presents the main theoretical result of the paper. Section (ref) provides a simulation exercise to contrast the finite sample performance of CIV with competing estimators. Section (ref) revisits the application of dobbie2018effects and chyn2024examiner where bail judge identities are used as instruments for pre-trial release of defendants. Section (ref) concludes.
Notation. It is useful to clarify some notation. In the following sections, I characterize the law $P_n$ of the random vector $(Y, D, X^\top, Z^{(n)}, Z^{(0)}, U)$ associated with a single observation. Here, $Y$ denotes the outcome, $D$ is the scalar-valued endogenous variable of interest, $X$ is a $J$ dimensional vector of control variables, $Z^{(n)}$ is the observed instrumental variable, $Z^{(0)}$ is a latent instrumental variable, and $U$ are all other determinants of $Y$ other than $(D, X^\top)$. It is often convenient to use $W\equiv (D, X^\top)^\top$. Throughout, I maintain focus on a scalar-valued $D$ for ease of exposition but highlight that any fixed number of endogenous variables can easily be accommodated under the presented framework. The joint law $P_n$ is allowed to change with the sample size $n\in \mathbbm{N}$ to approximate settings with relatively few observations per observed category. To permit the study of semiparametric efficiency, however, I keep the marginal law of $(Y, D, X^\top, Z^{(0)}, U)$ fixed. Explicit references to $P_n$ are largely omitted for brevity. Further, for a random vector $S$, let $\mathcal{S}$ denote its support and $\vert \mathcal{S}\vert$ the cardinality of the support. For an i.i.d.\ sample $\{S_i\}_{i=1}^n$ from $S$, define the operators $\mathbbm{E}_n S \equiv \frac{1}{n}\sum_{i=1}^n S_i$ and $\mathbbm{G}_n S \equiv \frac{1}{\sqrt{n}}\sum_{i=1}^n (S_i - \mathrm{E} S)$. For a measurable set $\mathcal{A}$, the indicator function $\mathbbm{1}_{\mathcal{A}}(S)$ is equal to one if $S\in \mathcal{A}$ and zero otherwise. For a function $f:\mathcal{A}\to \mathbbm{R}$, let $f(\mathcal{A})$ denote its image. The $\ell_2$-norm is denoted by $\|\cdot\|$ where for a matrix $M$ the norm is $\|M\|=\operatorname{tr}(M^\top M)^{1/2}$.
This section introduces the categorical instrumental variable estimator and discusses its key motivating assumptions. I focus on the setting without control variables in this section, leaving the general specification for Section (ref) that provides the formal asymptotic analysis.
The instrumental variable model is
where $\tau_0$ is the parameter of interest. In the linear model with homogeneous effects considered here, the coefficient corresponds to the change in the outcome caused by a marginal change in endogenous variable of interest.
The mean-independence of $U$ and $Z$ implies the moment condition $ \mathrm{E}\big(Y - D\tau_0\big)\big(m(Z) - \mathrm{E} m(Z)\big) = 0$ for any measurable function $m:\mathcal{Z}\to \mathbbm{R}$. If in addition $\mathrm{Cov}(D, m(Z)) \neq 0$, a solution to the moment condition is given by
Replacing the covariances with their sample analogues, the instrumental variable model thus suggests potentially infinitely many estimators for $\tau_0$ as indexed by the function $m$, each consistent and root-$n$ asymptotically normal under regularity assumptions. To decide between these alternative estimators, econometricians have turned to study their efficiency.
Suppose the econometrician observes an i.i.d.\ sample $\{(Y_i, D_i, Z_i)\}_{i=1}^n$ from $(Y, D, Z)$ and $U$ is homoskedastic -- i.e., $\mathrm{E} [U^2\vert Z]\overset{a.s.}{=}\sigma^2$. If $f$ is chosen to be the conditional expectation $m_0(z)\equiv \mathrm{E}[D\vert Z=z]$, then the asymptotic variance of the corresponding oracle estimator
achieves the semiparametric efficiency bound $\sigma^2/\mathrm{Var}(m_0(Z))$ (see, e.g., chamberlain1987asymptotic). The transformed instruments $m_0(Z)$ are thus often termed “optimal instruments.”
Formulating estimators based on the moment solution (ref) has additional benefit of falling in the class of “two-step” estimators as defined by kolesar2013estimation. The author shows that even if the underlying structural model is not additively separable in the structural error $U$ as presumed here, two-step estimators admit interpretation as a convex combination of causal effects under the LATE assumptions of imbens1994identification.\footnote{In addition to stronger exogeneity assumptions, the LATE assumptions include a monotonicity assumption that prohibits simultaneous movements in-and-out of treatment for any increment of the optimal instrument.} This starkly contrast the LIML estimator and its variants, which do not generally permit a weakly causal interpretation in the LATE framework kolesar2013estimation. Estimation based on (ref) and the optimal instrument $m_0$ thus has both important economic and statistical benefits.
In economic applications, the conditional expectation $m_0$ is rarely known. The oracle estimator $\hat{\tau}^*$ is thus typically infeasible in practice. A growing literature focuses on estimating the optimal instruments such that the asymptotic distribution of the resulting estimator for $\tau_0$ achieves the same asymptotic variance as the infeasible estimator. For example, newey1990efficient considers nearest-neighbor and series regression to approximate $m_0$. In settings with growing numbers of instruments, belloni2012sparse and carrasco2012regularization consider regularized regression estimators.
This paper is concerned with estimation of optimal instruments in settings where the observed instrument is categorical. To provide a better asymptotic approximation to settings with relatively few observations per category, I allow the number of categories to grow with the sample size. Letting $Z^{(n)}$ denote the observed instrument to highlight this dependence on the sample size index $n \in \mathbbm{N}$, Assumption (ref) formally specifies the rate at which the number of categories is allowed to grow. In particular, the number of categories can increase such that the expected number of observations per category (i.e., $n\times \Pr(Z^{(n)}=z)$) grows at arbitrarily slow polynomial rate with the sample size.
If all $\lambda_z\in(0.5, 1]$ the number of categories grows sufficiently slowly such that the optimal instrument can be estimated at a sufficiently fast rate by simple least squares of $D$ on the (increasing) set of indicators $(\mathbbm{1}_z(Z^{(n)}))_{z\in\mathcal{Z}^{(n)}}$. The more interesting settings are thus if for some categories $\lambda_z\in(0, 0.5]$. Since $\lambda_z$ can be arbitrarily close to $0$, this regime can be viewed as approximating settings in which the number of observations per category is small or moderate. These settings seem of particular practical importance in economic applications. Indeed, as discussed in the introduction, the joint use of jackknife-based instrumental variable estimators and conventional (heteroskedasticity-robust or clustered) standard errors in the examiner fixed effects literature (see, e.g., chyn2024examiner) suggests these “moderately many instruments” regimes are the leading asymptotic approximation in current applied research that leverages judge identities as instruments.
To accommodate efficient estimation in these more challenging settings with growing number of categories, I make a complexity-reducing assumption on the structure of the optimal instrument. Assumption (ref) asserts that the optimal instrument, denoted $Z^{(0)}$, has finite support of cardinality $K_0\in \mathbbm{N}$. Unlike the observed instrument $Z^{(n)}$, the cardinality of the support of optimal instrument is thus fixed as the sample size increases.\footnote{The application of a finite support restriction as in Assumption (ref) along with a rate condition as in Assumption (ref) to cross-sectional regression on categorical variables appears novel, however, finite support assumptions have grown increasingly popular in longitudinal data settings where the categories are individual identifiers (see, in particular, bonhomme2015grouped). In these longitudinal data settings, Assumption (ref) corresponds to a group-fixed effects assumption and Assumption (ref) regulates the rates at which the cross-section and the time dimension grow.}
Assumption (ref) implies that all relevant information on the endogenous variable included in $Z^{(n)}$ is also captured by the latent instrument $Z^{(0)}$. That is, there exists a deterministic map from values of $Z^{(n)}$ to values of the optimal instrument $Z^{(0)}$ -- or -- a partition $(\mathcal{Z}^{(n,0)}_k)_{k=1}^{K_0}$ of $\mathcal{Z}^{(n)}$ such that the optimal instrument is constant within each $\mathcal{Z}^{(n,0)}_k$: $m_0^{(n)}(z) = m_0^{(n)}(z'), \forall z, z' \in \mathcal{Z}^{(n,0)}_k$. When this map is known, efficient estimation of $\tau_0$ simplifies to TSLS of $Y$ on $D$ using the (non-increasing) set of indicators $(\mathbbm{1}_{\mathcal{Z}^{(n,0)}_k}(Z^{(n)}))_{k \in \{1, \ldots, K_0\}}$ as instruments. In practice, the map is unknown and needs to be estimated.
To estimate the optimal instrument, I propose a $K$-Conditional-Means ($K$CMeans) estimator given by
where $\mathcal{M}\subset \mathbbm{R}$ is compact and $K\in\mathbbm{R}$ is the number of support points allowed-for by the researcher. In contrast to (unconditional) $K$Means which clusters observations into $K$ groups, $K$CMeans creates a partition of the support of the categorical variable $Z^{(n)}$ allowing its application to reduced form regression problems as required here. In practice, $K$CMeans is solved via a dynamic programming algorithm adapted from the algorithm for $K$Means discussed in wang2011ckmeans. An important feature of this approach is that $K$CMeans can be solved to a global minimum in time polynomial in $\vert \mathcal{Z}^{(n)}\vert$, avoiding the heuristic solution approaches to $K$Means problems in multiple dimensions. I thus do not need to abstract away from optimization error as is common in other applications of $K$Means estimators (see, e.g., bonhomme2015grouped).
The categorical instrumental variable (CIV) estimator is then simply given by
When $K_0$ is known and $\hat{m}_{K_0}^{(n)}$ is the $K$CMeans estimator (ref) with $K = K_0$, CIV is a feasible analogue to the infeasible oracle estimator $\hat{\tau}^*$ in (ref). As discussed in detail in Section (ref) and given the therein stated assumptions, $\hat{\tau}$ has the same asymptotic distribution as the infeasible estimator and is semiparametrically efficient under homoskedasticity.
The result of semiparametric efficiency of CIV depends on the correct choice of $K=K_0$. In applications, economic insight can occasionally provide concrete suggestions for a value of $K_0$. For example, Appendix (ref) provides a rationale for why $K_0=2$ in the application of angrist1991does. In other applications, knowledge of $K_0$ is less certain. In judge fixed effects applications, for example, a researcher may consider the classification of judges as “lenient” and “strict” as only an approximation to the true latent types of judges. In these settings where a concrete suggestions for $K_0$ is not available, the CIV estimator that uses $\hat{m}_{K}^{(n)}$ as an approximate optimal instrument maintains root-$n$ normality for any $K\in\{2,\ldots, K_0\}$ provided that the approximate optimal instrument remains relevant. In particular, for $K \in \{2, \ldots, K_0\}$, $\hat{m}_K^{(n)}$ estimates the approximate optimal instrument
As stated in Theorem (ref), when $m_K^{(n)}(Z^{(n)})$ satisfies the usual instrumental variable rank condition, the corresponding CIV estimator remains asymptotically normal even when $K<K_0$. Albeit at a loss of statistical efficiency, this implies that knowledge of a lower-bound on $K_0$ is often sufficient for inference (e.g., choosing $K=2$).\footnote{Importantly, under-specification of $K$ has no implication on the interpretation of the CIV estimand as a convex combination of causal effects. As illustrated in the proof of Lemma (ref), the $K$CMeans solution is contiguous in the sense of fisher1958grouping -- i.e., only adjacent support points of $Z^{(0)}$ are combined when constructing the approximate optimal instrument. As a consequence, if a monotonicity assumption applies to $Z^{(0)}=m^{(n)}_0(Z^{(n)})$ to warrant interpretation of the TSLS as a convex combination of causal effects as in angrist1995two, then a monotonicity assumption will also apply to approximate optimal instrument $m^{(n)}_K(Z^{(n)})$. Analogous arguments apply in a setting with covariates, albeit under the qualification of saturated controls which applies broadly to two-step estimators. See, for example, blandhol2022tsls for a discussion on the importance of saturated controls for the interpretation of TSLS estimands.}
This section provides a formal discussion of the CIV estimator. After defining the $K$CMeans and CIV estimators in the presence of additional exogenous variables of fixed dimension, I state the main result of the paper in Theorem (ref). Corollary (ref) provides the statement of semiparametric efficiency for known $K_0$.
Assumption (ref) defines the instrumental variable model with $Z^{(0)}$ being the unobserved optimal instrument previously characterized by Assumption (ref). The parameter of interest is the vector $\theta_0 \equiv (\tau_0, \beta_0^\top)^\top$.
For $K\in \mathbbm{N}$, define the CIV estimator as
where $\hat{F}_K \equiv (\hat{g}_K(Z,X), X^\top)^\top$ with $\hat{g}_K(Z,X) \equiv \hat{m}_K^{(n)}(Z) + X^\top \hat{\pi}$ and
with $\hat{\pi}$ being a root-$n$ consistent first-step estimator for $\pi_0$. Note that because $X$ is relatively low-dimensional, root-$n$ consistent estimators for $\pi_0$ can be obtained as within-category regression estimators of $D$ on $X$ with fixed effects defined by the categories of $Z^{(n)}$.
Assumption (ref) ensures that the tail probabilities of the conditional expectation function residual $V \equiv D - Z^{(0)} - X^\top\pi_0$ and the exogenous variables $X$ decay at exponential rate. Analogously to bonhomme2015grouped, Assumption (ref) along with a dependence restriction (see Assumption (ref) (d)) allows for application of exponential inequalities that are key to bound the probability of misclassifying values of $Z^{(n)}$ in estimation of the (approximate) optimal instruments.
Finally, Assumption (ref) (a) places moment restrictions on the second stage error used, in particular, for consistent estimation of standard errors. Assumption (ref) (b) requires compactness of the first stage coefficients. Assumption (ref) (c) requires $\hat{\pi}$ to be a root-$n$ consistent estimator for $\pi_0$. Assumption (ref) (d) asserts that the econometrician observes independent samples from $(Y, D, Z^{(n)}, X^\top)$.
\ifpublic \fi
Theorem (ref) states the main result of the paper. In particular, it shows that when the infeasible oracle estimator
where $F_K \equiv (g_K(Z^{(n)},X), X^\top)^\top$ with $g_K(Z^{(n)},X) \equiv m^{(n)}_K(Z^{(n)}) + X^\top \pi_0$, is root-$n$ normal, then assumptions (ref)-(ref) are sufficient for the CIV estimator $\hat{\theta}^K$ to be root-$n$ normal with the same asymptotic covariance matrix as long as $K\in \{2, \ldots, K_0\}$.
For the CIV estimator that uses $K=K_0$, Corollary (ref) further provides a semiparametric efficiency result under homoskedasticity.\footnote{Note that in the heteroskedastic setting, weighting observations proportional to their variance can improve the asymptotic variance. Since this approach follows standard general method of moments arguments, I omit further discussion here.}
\enlargethispage{\baselineskip}
This section discusses a Monte Carlo simulation exercise to illustrate finite sample behavior of the proposed CIV estimator and highlight key challenges of alternative optimal IV estimators for estimation with categorical instrumental variables.
For $i = 1, \ldots, n$, the data generating process is given by
where $(U_i, V_i)\sim \mathcal{N}(0, \left[
\right])$, $D_i$ is a scalar-valued endogenous variable, $X_i\simBernoulli(\frac{1}{2})$ is a binary covariate and $\beta_0 = \gamma_0 = 0$, and $Z_i$ is the categorical instrument taking values in $\mathcal{Z} = \{1, \ldots, 40\}$ with equal probability. To introduce correlation between $Z_i$ and $X_i$, I further set $\Pr(Z_i is odd\vert X_i = 0) = \Pr(Z_i is even\vert X_i = 1) = 0$. The optimal instrument $m_0$ is constructed by first partitioning $\mathcal{Z}$ into $K_0$ equal subsets and then assigning evenly-spaced values in the interval $[0, C]$.\footnote{For example, for $K_0=2$, $m_0(z)=0$ for $z \in \{1,\ldots, 20\}$ and $m_0(z)=C$ for $z \in \{21,\ldots, 40\}$.} I choose the scalars $\sigma_V^2$ and $C$ such that the variance of the first stage variable is fixed to 1 and the concentration parameter in the smallest considered sample is $\mu^2 = 180$.\footnote{In particular, $\sigma_V^2=0.9$, and $C\approx 0.85$ for $K_0=2$ and $C\approx 1.153$ for $K_0=4$ so that with $n=800$, the concentration parameter $nM_0^\top (\mathrm{Cov}(\mathbbm{1}_z(Z))_{z \in \mathcal{Z}}) M_0/\sigma_V^2 = 180$ where $M_0$ is the 40-dimensional vector of first stage coefficients associated with every category. Choosing $\sigma_V^2$ and $C$ in this manner is akin to the simulation setup in \citet{belloni2012sparse}.} As in the simulation considered in \citet{kolesar2013estimation}, the data generating process allows for individual treatment effects $\pi_0(X_i)$ to differ with covariates. Here, $\pi_0(X_i) = \tau_0 + 0.5(1 - 2X_i)$ so that the expected treatment effect is simply $\mathrm{E}\pi_0(X) = \tau_0.$ As a consequence, the second stage is heteroskedastic and -- unlike two-step IV estimators like CIV -- the LIML estimator is inconsistent for the average treatment effect.
I compare properties of thirteen estimators in the simulation: An infeasible oracle estimator $\tilde{\theta}^{K_0}$ with known optimal instrument, CIV with $K=2$ and $K=4$, TSLS, JIVE, IJIVE, UJIVE, and LIML that use the observed instruments, and five machine-learning based IV estimators that use lasso with cross-validated or plug-in penalty parameters, ridge regression, gradient tree boosting, or random forests to estimate the optimal instrument.
Table (ref) provides the bias, median absolute error (MAE), rejection probabilities of a 5% significance test (rp(0.05)),\footnote{Standard errors used for construction of rp(0.05) are heteroskedasitcity robust and do not include additional variance terms associated with traditional many instrument asymptotics. Inclusion of these additional terms has no qualitative consequence for the rejection probabilities of JIVE, IJIVE, or UJIVE.} and the inter-quantile range between 10th and 90th empirical quantiles (iqr(10, 90)) for a DGP in which $\tau_0 = 0$ and the optimal instrument has $K_0=2$ support points. All estimates are computed on sample sizes with 20, 25, 100, and 150 expected observations per observed instrument. As expected in the strong instruments setting considered here, the oracle estimator achieves small bias and nominal false rejection rates across all sample sizes. Its feasible analogues that attempt to estimate the optimal instrument in a first step, on the other hand, vary substantially across sample sizes.
The CIV estimator restricted to two support points achieves near-oracle performance at a moderate number of observations per category. This is in strong contrast to the alternative optimal instrument estimators in this setting. In particular, even at the much larger sample size with 150 observations per category, TSLS has a false rejection rate of 0.15, far above the 5% nominal level. Further, none of the considered machine-learning based optimal instrument estimators improve upon TSLS. Given that there are only 40 first stage instruments in a total sample size of up to 6000 ($=40 \times 150$), it may be surprising that application of the lasso does not result in better empirical performance. However, note that the shrinkage assumptions (implicitly) leveraged by any of the machine-learning based estimators are not suitable approximations of the categorical instrumental variable design considered here. Only the CIV estimator with over-specified number of support points has slightly lower bias and false rejection rates than TSLS, yet, remains inferior to the CIV estimator with correct number of groups. Note that the theory provided in this paper does not provide results for CIV estimators with $K>K_0$. Finally, all jackknife-based estimators exhibit small bias, but are substantially more dispersed than CIV. This dispersion prohibits JIVE, IJIVE, and UJIVE to accurately control size even at 150 expected observations per instrument. This dispersion seems in large part be due to the treatment effect heterogeneity present in the considered DGP: Replicating the results for a DGP without treatment effect heterogeneity shows that the jackknife-based estimators control size at all sample sizes (see Appendix (ref)).
Table (ref) replicates the simulation in a DGP where $K_0=4$. In contrast to the previous results on CIV with over-specified support points, the results show that estimated confidence intervals for CIV with under-specified number of support points can achieve correct coverage. While the first-stage estimation problem is substantially more challenging with $K_0=4$ as all support points are closer together, CIV with both $K=2$ and $K=4$ achieves near-oracle performance for the larger sample sizes. This is again in strong contrast to any of the competing optimal instrument estimators, whose biases are an order of magnitude larger and whose false rejection probabilities are two to three times those of the two CIV estimators. Similarly, the jackknife-based estimators are more dispersed than CIV and fail to control size in the data generating process with treatment effect heterogeneity.
Finally, as expected given the heterogeneous second-stage effects, LIML is heavily biased for the average treatment effect throughout in both Table (ref) and (ref).
For additional insights on the empirical performance of the considered estimators, Figure (ref) plots power curves for the hypothesis test $H_0: \tau_0 = 0$ at significance level $\alpha = 0.05$ in the design with $K_0=2$. Panels (a) and (b) keep the second stage coefficients constant, while panels (c) and (d) mirror the heterogeneous second stage effect design above. For brevity, the figure focuses on only a subset of estimators considered previously: CIV with $K=2$, the oracle estimator, TSLS, JIVE, LIML, and the lasso-based IV estimator with cross-validated penalty level. Appendix (ref) provides the figures corresponding to the remaining estimators.
When there is no treatment effect heterogeneity (panels (a) and (b)), the LIML and JIVE estimators achieves near oracle performance at even the smallest sample size considered. This mirrors the insights of donald2001choosing, bekker2005instrumental, and chao2012asymptotic, as well as the excellent empirical performance of the LIML estimator in the simulations of angrist2020machine. CIV with $K=2$ results in similar rejection rates for the moderate sample with 25 expected observations per category. In contrast to LIML and JIVE, CIV retains its near-oracle performance when second stage effects are heterogeneous (panels (c) and (d)). In all panels, TSLS and Lasso-IV (cv) are heavily biased.
This section applies the CIV estimator to a judge fixed effects instrumental variable analysis of the effect of pre-trial release on conviction. As in dobbie2018effects and chyn2024examiner, I consider a 2006-14 sample of misdemeanor and felony cases assigned to weekend bail hearings in the Miami-Dade County, Florida.\footnote{The data is publicly available from chyn2024examiner.} The data covers 94,355 cases for 186 bail judges. The goal of the example is to illustrate the application of CIV and alternative IV estimators in settings with potentially noisy estimates of judge leniency.
Following the previous literature analyzing the Miami-Dade data, the relationship of interest is the causal effect of pre-trial release on the conviction of the defendant. Pre-trial detention occurs if a defendant who has been assigned to a bail hearing does not post bail thereafter.\footnote{As described in chyn2024examiner, bail can also be posted after an initial bail value is set immediately following arrest. These observations are not in the sample.} Because the bail judge can change the bail amount, whether or not a judge is lenient may have a direct effect on the pre-trial release of a defendant. Using bail judge identity as an instrument is then motivated by the fact that bail judges are quasi-randomly assigned conditional on the court and date of the hearing. See, in particular, chyn2024examiner for a comprehensive discussion of the IV assumptions in the Miami-Dade setting.
I focus on the analysis of a simple instrumental variable specification:
with $\mathrm{E}[\varepsilon\vert Z, X] = \mathrm{E}[\nu\vert Z, X]=0$, where $Y$ is an indicator equal to one if the defendant is convicted, $D$ is an indicator equal to one if they have met bail, $X$ is a set of court-by-time fixed effects, $Z$ is the categorical instrument capturing bail judge identities, and $m_0(z)$ is the (unobserved) leniency of judge $z\in \mathcal{Z}$. The parameter of interest is the parameter $\tau_0$, which captures the causal effect of pre-trial release on conviction under correct model misspecification. Under stronger distributional independence and monotonicity assumptions, $\tau_0$ can also capture a convex combination of local average treatment effects in a data generating process with unobserved treatment effect heterogeneity frandsen2023judging,blandhol2022tsls.
A potential concern with estimating $\tau_0$ given an i.i.d.\ sample $\{(Y_i, D_i, X_i, Z_i)\}_{i=1}^n$ via TSLS is that the number of cases per judge can be moderate so that a many instrument bias arises. As shown in Figure (ref), the distribution of cases per judge indeed varies substantially, with most judges working on around 500 cases but a few completing less than 200 cases. The fact that leniency of judges may be estimated noisily may then motivate the use of jackknife-based estimation in dobbie2018effects and chyn2024examiner.\footnote{In contrast, the dimension of the court-by-time fixed effects in the Miami-Dade County setting is not a first order concern. Appendix (ref) shows that that almost all fixed effects are associated with more than 600 cases with no large left tail.}
Column (6) in Table (ref) presents estimates for $\tau_0$ using the sample of in chyn2024examiner which is restricted to judges with at least 200 cases. The considered estimators are the proposed CIV estimator with $K=2$ and $K=5$ support points, TSLS, JIVE, IJIVE, UJIVE, and the post-Lasso IV estimator of belloni2012sparse that selects among the judge fixed effects. I also provide OLS estimates for comparison. With exception of the post-Lasso IV estimator which is estimated with high variance,\footnote{The post-Lasso IV estimator using the plug-in penalty parameter of belloni2012sparse selects only three judge fixed effects and removes any others from the specification. The poor performance of the lasso-based estimator mirrors its poor performance in categorical instruments settings in Section (ref) and in angrist2020machine.} the point estimates of TSLS differ only marginally from those of the jackknife-based estimators or CIV with any level of regularization. There are substantial differences in the standard errors, yet, given the large share of observations associated with judges who have more than 500 cases, it might not be surprising that any potential bias from many instruments is small.
To gain a better understanding of estimator differences in settings with potentially noisy leniency estimates, I thus also consider artificially restricted sub-samples of the Miami-Dade data. These sub-samples are meant to represent examiner fixed effects settings with fewer decisions per decision maker so that bias from many instruments may be more likely to arise. For this purpose, columns (1)-(4) restrict the sample to judges with at most 400, 500, 600, and 700 cases, respectively. Throughout, I also restrict the sample to judges with at least 30 cases. Columns (1)-(4) may thus be well approximated by the “moderately many instruments” asymptotics outlined in this paper. For comparison, column (5) also provides results for judges with at least 30 cases but where the number of cases per judge has no upper bound. Throughout, I focus on comparison of estimators within a particular column. Within a column, TSLS and the jackknife-based estimators target the same estimand as CIV with a correct choice of $K$. However, because the distribution of judges changes across columns, the targeted estimand can also differ across columns when treatment effects are heterogeneous. These changes in estimands can further complicate comparisons across columns.
In the settings of columns (1)-(4) with moderately many cases per judge, regularization of the first stage via a support point restriction has a substantial impact on the estimates. The first considered CIV estimator restricts the estimated judge leniency to $K=2$ support points. This restriction can correspond to the assumption that all judges are either “lenient” or “strict.” Alternatively, $K=2$ can be viewed as optimally approximating a higher-dimensional vector of judge types with a binary “pseudo” type. For insights into the estimated binary (pseudo) type, Figure (ref) presents the empirical distribution of estimated judge leniency measures as well as how the leniency measures map to the binary (pseudo) type estimate. Here, the dotted vertical lines correspond to the leniency of the “lenient” and “strict” judges, respectively, and the dashed vertical line indicates the cut-off for assigning judges to these types.
The CIV estimator with $K=2$ effectively regularizes the first stage problem of estimating judge leniency. In columns (1)-(4), the CIV estimate differs from the (unregularized) TSLS estimate by about one standard error. Further, the CIV estimate is negative and statistically significant throughout, suggesting that pre-trial release has a negative causal effect on conviction of a defendant regardless of the particular subsample. This is in strong contrast to the estimates associated with JIVE, which does not result in statistically significant coefficients at conventional levels. The IJIVE and UJIVE estimates are numerically similar to the CIV (K=2) estimates but are associated with substantially larger standard errors throughout. For the smaller subsamples, IJIVE and UJIVE standard errors are between 30-60% larger than CIV standard errors.
The CIV estimator with $K=5$ (pseudo) types is included here to illustrate that lower regularization (i.e., higher $K$) results in the CIV estimator to tend towards TSLS. Indeed, for all considered subsamples, CIV with $K=5$ is similar to (unregularized) TSLS.
In summary, the presented results using the Miami-Dade County data suggest that the jackknife-estimators are either highly subsample-dependent (as is the case for JIVE) or are similar to highly regularized CIV with a binary (pseudo) type but with larger standard errors. The proposed CIV estimator (e.g., with a binary (pseudo) type) may thus provide a useful alternative to currently popular jackknife-based estimators in examiner fixed effects settings with moderately many decisions per decision maker.
This paper considers estimation with categorical instrumental variables when the number of observations per category is relatively small. The proposed categorical instrumental variable estimator is motivated by a first-stage regularization assumption that restricts the unknown optimal instrument to have fixed finite support. In a “moderately many instruments” regime that allows the number of observations to grow at arbitrarily slow polynomial rate with the sample, I show that when the number of support points of the optimal instrument is known, CIV achieves the same asymptotic variance as the infeasible oracle two-stage least squares estimator that presumes knowledge of the optimal instrument and is semiparametrically efficient under homoskedasticity. Further, under-specifying the number of support points maintains asymptotic normality but results in efficiency loss. A simulation exercise illustrates the finite sample performance of the proposed CIV estimator and highlights pitfalls associated with alternative optimal instrument estimators in the setting of categorical instruments. Similar to results in angrist2020machine, lasso-based IV estimators do not improve upon the bias of TSLS and fail to control size. In contrast, CIV successfully leverages the low-dimensional structure of the optimal instrument to obtain near-oracle estimates. Finally, the analysis of pre-trial release on conviction using bail judge identities as instruments as in dobbie2018effects and chyn2024examiner illustrates potential practical advantages of CIV over conventional jackknife-based instrumental variable estimators. In particular, highly regularized CIV with a binary (pseudo) type achieves lower standard errors than IJIVE and UJIVE while resulting in numerically similar point estimates.
\interlinepenalty=10000 \addcontentsline{toc}{section}{References} \interlinepenalty=10