EconBase
← Back to paper

A Locally Robust Semiparametric Approach to Examiner IV Designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

63,652 characters · 10 sections · 136 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Locally Robust Semiparametric Approach to Examiner IV Designs

abstractI propose a locally robust semiparametric framework for estimating causal effects using the popular examiner IV design, in the presence of many examiners and possibly many covariates relative to the sample size. The key ingredient of this approach is an orthogonal moment function that is robust to biases and local misspecification from the first step estimation of the examiner IV. I derive the orthogonal moment function and show that it delivers multiple robustness where the outcome model or at least one of the first step components is misspecified but the estimating equation remains valid. The proposed framework not only allows for estimation of the examiner IV in the presence of many examiners and many covariates relative to sample size, using a wide range of nonparametric and machine learning techniques including LASSO, Dantzig, neural networks and random forests, but also delivers root-n consistent estimation of the parameter of interest under mild assumptions. JEL Numbers: B23, C01, C14, C26, C31, C45 Keywords: examiner IV, semiparametric estimation, orthogonal moment function, machine learning.

Introduction

Examiner instrumental variable (IV) designs are increasingly popular in empirical economics. An early example of these examiner IV designs in applied economics is found in 10.1257/aer.96.3.863, who leverages random assignment of judges to estimate the effects of the duration of incarceration spells on two labor market outcomes: employment and earnings prospects. 10.1257/aer.96.3.863 uses the average incarceration length meted out by one’s randomly assigned judge across all cases adjudicated by that judge as an IV for the potentially endogenous duration of incarceration spell. Since 10.1257/aer.96.3.863’s seminal study, there has been a burst of new literature exploiting idiosyncratic allocation of examiners to cases in pursuit of answers to a wide variety of interesting economic questions in settings beyond the criminal justice system.

The key idea in these designs is to exploit exogenous variation in the treatment propensity of these randomly assigned examiners in an IV strategy. These designs differ from the textbook IV designs insofar as the examiner IV, the examiner’s propensity to assign treatment status, is a latent variable that is estimated from data. To estimate the examiner IV, the canonical examiner IV design exploits jackknife methods to deal with “own observation” bias, which arises when one's own treatment status is used in the estimation of the examiner's treatment propensity.

The standard approach to estimation and inference in examiner IV designs with a single binary endogenous treatment and covariates is the unbiased jackknife estimation method proposed by kolesarcowles. In most settings, examiners have control over which times (days or time of day) they can work, and where or in which office they can work. Random assignment of examiners is, therefore, conditional on this sorting. To the extent that the random assignment is conditional, the treatment propensity is valid as an IV conditional on these covariates (conditionally ignorable). In some settings, the treatment propensity also varies across case characteristics while in others random assignment happens within batches of cases rather than at the individual level. Different applications, therefore, admit different configurations of covariates to cleanse the examiner IV of the residual confounding variation. kolesarcowles's jackknife method allows for such covariates to avoid the well-known biases from canonical jackknife procedures in the presence of covariates.

However, kolesarcowles's approach excludes a range of applications. kolesarcowles's method hinges on the assumption that the approximation of the conditional expectation function for the treatment propensity is exactly linear in covariates. By construction, this linear approximation is exact in saturated specifications. In these specifications, empirical researchers necessarily include all the discrete covariates (e.g. strata indicators) and their interactions. With a few number of covariates, this approach works well but with even a moderate number of covariates, the dimension of the covariate vector scales up quite quickly relative to the sample size. On the other hand, in non-saturated designs, including when continuous covariates are included to predict the treatment propensity of examiners, the assumption that the linear approximation is exact is restrictive and generally unwarranted. Instead, researchers may want to estimate the treatment propensity nonparametrically.

In general, existing methods are not suitable for applications where the number of covariates is quite large relative to the number of observations, including in the saturated specifications with high dimensional fixed effects. In such settings, it seems eminently reasonable to use machine learning (ML) methods to predict the treatment propensity. Nevertheless, a growing recent literature suggests caution in use of ML in two-step estimation problems. This literature has established that plug-in estimators based on machine learning first steps introduce large regularization and model selection biases which contaminate estimation and inference in the second step (see, for example, chernozhukov2018double). In some settings, including applications involving naive LASSO plug-in estimators in two-estimation, the two-step estimators are not even root-n consistent (chernozhukov2022locally). I focus on these settings where empirical researchers estimate the conditionally ignorable treatment propensity nonparametrically with many examiners and possibly many covariates relative to the sample size.

I set out to answer the question: how can we perform reliable semiparametric estimation and inference in examiner IV designs where the examiner IV, the treatment propensity of each quasi-randomly examiner, is estimated nonparametrically with possibly many covariates relative to sample size? In an attempt to answer this question, I propose a locally robust semiparametric approach to estimation and inference in examiner IV designs. This method-of-moments approach entails constructing debiased sample moments based on an orthogonal/locally robust score which consists of an identifying moment function and a first step influence function. As I will argue, this approach is well-suited to settings where the examiner IV is estimated nonparametrically with possibly many covariates because it mitigates the machine learning (regularization and model selection) biases and the first step estimation error compared to simple plug-in approaches. This approach also takes care of the own observation bias that arises in this literature (chernozhukov2022locally). Furthermore, the framework that I propose delivers root-n consistent estimation using examiner IV in settings where the examiner IV is estimated in the first step using one of a wide range of nonparametric or ML methods.

This paper contributes to a growing theoretical literature on examiner IV designs. In important work, frandsen2023judging revisit these examiner IV designs and propose a new approach to surmount violations of some of the assumptions underpinning standard approaches. Unlike frandsen2023judging, I do not interrogate the identification issues in this literature. Instead, I focus exclusively on estimation and inference in the presence of covariates. I build on the excellent work of kolesarcowles who proposes an unbiased jackknife estimation method that accounts for covariates in the estimation of the treatment propensity. While kolesarcowles contends that the treatment propensity can also be estimated nonparametrically using his method, his method does not account for the biases from the first step if non-parametric methods or machine learning is used in the estimation of the treatment propensity. Another closely related piece of work is mueller2015criminal which, to the best of my knowledge, is the earliest application of machine learning to estimate the treatment propensity in examiner IV designs. Unlike mueller2015criminal's approach, which uses LASSO under strong sparsity assumptions without adjusting for regularization and model selections biases, the approach I propose allows for a wide range of ML first steps including LASSO, corrects for first estimation biases and delivers root-n consistent estimation that is not guaranteed for estimators based on naive plug-in LASSO and other ML first steps. My approach does not require strong sparsity assumptions and is robust to misspecification of some of the components estimated in the first step or the outcome model.

While I focus explicitly on examiner IV designs, this work also contributes to the broader IV literature. Particularly, it contributes to the optimal instruments literature with machine learning or nonlinear first steps, which spans the work of belloni2012sparse, HANSEN2014290, chernuzhukov2015, syrgkanis2019machine, and bai2010instrumental. belloni2012sparse and chernuzhukov2015 use LASSO under approximate sparsity assumptions, while HANSEN2014290 and bai2010instrumental use ridge regression and gradient boosting respectively. syrgkanis2019machine provides a more general framework for using ML to estimate conditional average treatment effects (CATE). In this paper, I leverage locally robust semiparametric theory to provide a unified framework for semiparametric estimation of the IV model with treatment effect heterogeneity in kolesarcowles for a wide range of ML first steps in the presence of many exogenous covariates. This framework admits a completely nonparametric first step estimation of the “optimal" instrument in presence of many exogenous covariates under mild conditions.

More recent work closely related to this work is wiemann2023optimal who propose using a version of K-means clustering to estimate the latent optimal instrument from a large set of candidate categorical instruments. While this approach can be applied to the examiner IV setting insofar as it accounts for many candidate categorical instruments, it does not deal with the aspect of many covariates that is relevant when the latent optimal instrument is conditionally ignorable and which motivates this work. jochmans2023many also proposes a group fixed effects characterization of the examiner fixed effects, inspired by the panel data literature. Like wiemann2023optimal, jochmans2023many does not deal with issues arising from having many covariates. However, an important insight that I exploit from jochmans2023many is that cross-fitting reduces the first stage estimation error in the construction of the latent instruments. Thus, besides dealing with “own observation bias" and lack of Donsker conditions which are not known to hold in high dimensional settings, the cross-fitting as applied in my framework also mitigates the first stage estimation error. mikusheva2021many also observes that cross-fitting reduces dependence between machine learning first steps and inference in the second step, which further bolsters the case for cross-fitting and alleviates concerns regarding possible distortions induced by machine learning in empirical practice that have been recently documented by angrist2022machine.

This paper draws mainly from a more recent but rapidly expanding literature on automatic debiased machine learning and locally robust semiparametric estimation (see chernozhukov2018double, chernozhukov2021automatic, chernozhukov2022automatic, chernozhukov2022debiased) chernozhukov2022locally and ichimura2022influence). chernozhukov2018double introduce various double debiased machine learning approaches to discipline use of machine learning in statistics and econometrics, with emphasis on reducing regularization and model selection biases from first step estimation. In this paper, I exploit the locally robust semiparametric approach that debiases the identifying moment function, which depends on a plug-in estimate of a first step nuisance function, by adding to it an influence function adjustment. In developing this approach, I primarily leverage theoretical insights from chernozhukov2022locally and ichimura2022influence. Using this approach, I provide an estimation method that allows for use of most of the off-the-shelf machine learning algorithms to estimate the examiner IV under mild regularity conditions in a setting where there are many examiners and possibly many covariates. The method also mitigates bias when other nonparametric methods are used, and allows for misspecification in the first step estimation, by virtue of a multiple robustness property of the orthogonal moment function.

The remainder of the paper is organized as follows: section 2 outlines and further motivates the estimation issues that arise in the examiner leniency design; section 3 presents the main results; and section 4 concludes and provides directions for future work.

The Examiner Leniency Design: Estimation Issues in Empirical Practice

To fix concepts, I provide a quick overview of the examiner leniency design using an example from a canonical empirical setting: the US criminal justice system.

Suppose we are interested in the causal effect of pretrial detention on conviction (see frandsen2023judging): $$Y_{i} = \delta T_{i}+\mathbf{X}_{i}^{\prime} \beta+\epsilon_{i}$$

where $Y_i$ is the outcome; $T_i$ is the bail decision (it is $1$ when a defendant is detained and $0$ otherwise); $X_i$ is a vector of defendant and case characteristics; and $\varepsilon$ includes idiosyncratic factors that also influence the outcome. The treatment status, $T$, is endogenous if there are unobservables that are correlated with both treatment and the outcome. For example, well-to-do defendants can afford better lawyers and are therefore less likely to be detained and less likely to be convicted. Those who get treated (detained) are, therefore, potentially systematically different from those who do not get treated. This selection biases the OLS estimates of $\delta$.

The idea behind the examiner IV design is to instrument the endogenous treatment status, $T_i$, with one's assigned judge's propensity to detain or release, $T_i$, with $E\left[T_i \mid Z_i\right]$ where $Z_i$ is a $J \times 1$ vector of examiner indicators. Since the examiners (judges, in this case) are quasi-randomly assigned to cases, this propensity to assign treatment status is independent of the unobservables. This treatment propensity, also known as the judge stringency or leniency measure (depending on how $T$ is defined), is a valid instrument under the plain vanilla IV assumptions, namely: the exclusion restriction, which requires that judges affect outcomes only through the bail decision; monotonicity, which requires that a defendant released by a more strict judge is also released by the more lenient judge; and relevance, which effectively requires that the examiners' treatment propensity varies non-trivially.

This principle of examiner (or judge) IV designs applies more generally in settings where there is an idiosyncratic assignment of individuals to a set of decision-makers, whose propensity to assign treatment status varies non-trivially across those decision-makers. It therefore comes as no surprise that examiner IV designs are being increasingly used for causal inference in economics and adjacent disciplines, given the wide range of settings where such idiosyncratic assignment of decision-makers takes place. These designs have been exploited to examine causal effects of foster care on various socioeconomic outcomes (e.g. gross2022temporary and bald2022economics); causal effects of bankruptcy on earnings, adverse financial events and foreclosure rates (e.g. dobbie2015debt); as well as causal effects of disability benefits on labor supply, household consumption and mortality (e.g. black2018effect), to pick but a few from the expanding list of applications.

Since 10.1257/aer.96.3.863's seminal study, several studies have used a leave-one-out sample analogue of $E\left[T_i \mid Z_i\right]$ as an examiner IV: $$\frac{1}{\left|i^{\prime} \neq i, Z_{i^{\prime}}=Z_i\right|} \sum_{i^{\prime} \neq i, Z_{i^{\prime}}=Z_i} \Tilde{T_i}$$

The leave-one-out approach is used to avoid “own observation bias" whereby one's own endogenous treatment status used in the estimation of the examiner IV sullies the IV with the same endogeneity we are grappling with. Moreover, in many applications, the assignment of examiners is not completely random. In the working example of the bail system in the US, judges have control over which courtrooms, days and shifts they work so there is clearly intentional sorting of judges into locations and times, which would otherwise invalidate the examiner IV. To cleanse the judge stringency measure of this confounding variation, the location-by-time fixed effects (and possibly other covariates) are partialled out from the treatment variable via an application of the Frisch-Waugh-Lovell theorem. Therefore, researchers use the residualized treatment, $\Tilde{T_i}$, instead of $T_i$. With this estimate, they then apply the two stage least squares (TSLS). This leave-one-out approach effectively comes down to the jackknife IV estimation proposed by angrist1999jackknife.

An alternative approach that is currently recommended is kolesarcowles's unbiased jackknife IV (UJIVE). While the jackknife IV (JIVE) based on angrist1999jackknife yields the sample average treatment for an examiner leaving out $i$'s own treatment status, the UJIVE residualizes the treatment status using a jackknife regression before it is used in an IV estimator. In the UJIVE approach of kolesarcowles, the covariates are partialled out linearly via jackknife regressions under the assumption that the first stage regression is saturated so that a linear approximation of the conditional expectation function is exact.

The UJIVE formulation of kolesarcowles is based on the following decomposition of the conditionally ignorable treatment propensity with the covariate vector partialled out:

align[align omitted — 166 chars of source]

where the second equality is the linear approximation of the conditionally ignorable treatment propensity. kolesarcowles's first step estimator of $\gamma_i$, $\hat{\gamma_i}$, is the sample analogue: \[\hat{\gamma}_i=Z_i^{\prime}\hat{\beta}_{1/i}+X_i^{\prime}\hat{\beta}_{2/i}-X_i \hat{\alpha} \] where $\hat{\beta_{1/i}}$, $\hat{\beta_{2/i}}$ and $\hat{\alpha}$ are jackknife estimators (least squares estimators with observation $i$ left out). A vector $\hat{\gamma}$ of the $\hat{\gamma_i}$'s is then plugged as an instrument into an IV estimator in the second step (see section 6.2 of kolesarcowles for more details).

First, note that since $\hat{\gamma}$, is an estimate of the true treatment propensity, the estimation error from the construction of $\gamma$, represented as $\hat{\gamma}-\gamma$, potentially contaminates inference in the second step (cf: hahnridder2013). Secondly, the construction above hinges crucially on the assumption that the linear approximation above is exactly linear in the examiner indicators, $Z_i$, and the covariate vector, $X_i$. With discrete controls, the first stage is saturated when all the controls and their interactions are included in the covariate vector, $X$. This often leads to high dimensional fixed effects, which effectively \enquote{chop up} the data into cells and the examiner IVs are estimated in those cells. This potentially raises concerns of precision and many-weak IV (see bhuller for an example of this in practice).

In many applications, researchers resort to throwing away “small' cells, but what are considered “small” cells often varies across settings. frandsen2023judging, for example, restrict their sample to judges who had seen at least 50 arraignments. Other researchers choose higher or lower thresholds. In the worst case, this ad hoc solution might induce a severe missing data problem. A less pernicious but still troubling concern is that if those discarded “small” cells belong to the tails of the distribution of examiners' treatment propensity, then the estimand the researchers effectively estimate might be different from the estimand they think they are estimating. To see this, consider a setting where senior judges are more stringent and senior judges end up with few cases. Figure 1 below shows the truncation of the examiner treatment propensity (defined here as “stringency”) distribution in this hypothetical setting:

center[center omitted — 275 chars of source]

The red segment denotes the treatment propensities thrown away along with the most stringent judges and cannot be estimated. Since the estimand from kolesarcowles yields a weighted average of pairwise LATEs between each pair of judges, for a fixed value $p$ we estimate the estimand as a weighted average of pairwise LATEs between $p$ and $p'$. In practice, therefore, all pairwise LATEs in the red segment are not part of the estimand that is estimated. The challenge is that we do not observe $p$ and $p'$, so we do not know empirically the estimand that we are recovering in this setting. This setting, while presented here as a hypothetical scenario, is not far-fetched. bhuller, for example, observes that in Norway, senior judges and junior judges are assigned different caseloads. If, in such a setting, senior judges tend to see few arraignments and also tend to be more stringent, this scenario (or variations thereof) would not be inconceivable. Moreover, since the threshold for what is considered a “small" cell is at the discretion of the researcher and there is no theory guiding researchers on the appropriate threshold in a given application, different researchers using the same data would come to different conclusions in their empirical analyses if the estimates are sensitive to the choice of the threshold.

In non-saturated specifications or when continuous covariates are included, the assumption of a perfectly linear conditional expectation function is dubious. As recently observed by blandhol2022tsls, when the first stage is not saturated and unless the first stage is estimated nonparametrically, the TSLS estimand does not yield a LATE within a convex hull of pairwise treatment effects. To use blandhol2022tsls's terminology, kolesarcowles's estimand has a weakly causal interpretation when the first stage is perfectly linear or estimated nonparametrically. In non-saturated specifications or when continuous controls are included, blandhol2022tsls's insight suggests estimating the treatment propensity using some nonparametric methods such as sieves or polynomial splines. These methods, however, tend to produce biases which get transmitted to the second step where the parameter of interest is estimated. This motivates an approach that takes into account these biases in the second step.

More generally, specifications where the number of examiners and number of covariates are very large relative to sample size tip researchers over into the realm of high dimensional estimation and inference, where machine learning methods are routinely used. This paper provides a more general framework with mild theoretical guarantees which permit researchers to use an array of machine learning techniques to estimate treatment propensity in settings with large numbers of examiners and covariates that are not vanishingly small relative to the sample size. The framework also allows for first step nonparametric methods such as sieves or polynomial splines in settings with low to moderate numbers of examiners and controls.

In the framework that I provide in the next section, I allow $X$ and $Z$ to be very large relative to the sample size, $N$. In such a high-dimensional setting, we can think of the dimensions of $X$ and $Z$ as increasing with the sample size (see chernozhukov2018double). In the standard examiner IV designs, $Z$ consists of examiner indicators. In my proposed approach, I allow for $Z$ to be more general and the examiner fixed effects characterization arises as a special case. For example, $Z$ can be examiner types, and therefore be identified with vectors of examiner's discrete and continuous characteristics such as age, race, experience and so on. Henceforth, in lieu of the linear approximation in equation ((ref)) and motivated by the issues raised in this section, I use a nonparametric specification where the difference in conditional expectations is a difference in a square-integrable function of both $X$ and $Z$ and a square-integrable function of $X$ only.

Results

In this section, I present the main theoretical results of this paper. I first construct the orthogonal moment function by adding an influence function to the identifying moment function. As the name suggests, this moment function is orthogonal or insensitive to the biases induced in the first step non-parametric estimation of the treatment propensity, and is used to construct debiased sample moments. I show that the orthogonal moment function is valid, namely; the function satisfies the key orthogonality properties that are requisite for the locally robust semiparametric framework. I argue that the orthogonal moment function does not only reduce the effect of the first step estimation error and machine learning biases, but it is also multiply robust in the sense that misspecification of some of the first step components or the outcome model leaves the estimating equation valid. I then provide the statistical guarantees for this approach, which mainly consist of mean square convergence rates and a few other primitive conditions on the components estimated in the first step. I show that the examiner IV estimator is asymptotically normal and provide an expression for the variance as well as its consistent estimator.

Construction of the Orthogonal Moment Function

The construction of the orthogonal moment function or locally robust score below follows chernozhukov2022locally and ichimura2022influence. The orthogonal score consists of two ingredients, namely: the identifying moment function (score) and the influence function adjustment (or correction term). The identifying moment function derives from the estimand of interest while the influence function adjustment follows from an orthogonality condition in a sense that I make precise later in this section.

I use the following IV estimand in kolesarcowles:

\[ \theta:=\frac{\operatorname{Cov}(Y, \gamma)}{\operatorname{Cov}(T, \gamma)}=\frac{\mathbb{E}[Y \gamma]}{\mathbb{E}[T \gamma]} \] where $\gamma:=\mathbb{E}[T \mid X, Z]-\mathbb{E}[T \mid X]$ is the conditionally ignorable treatment propensity given a vector of covariates, $X$; $T$ is a binary treatment status; $Z$ is a vector of examiner indicators or examiner types; and $Y$ is a scalar outcome variable. Informally, you may think of $\gamma$ as the treatment propensity after partialling out the covariates, $X$. More formally, treating each examiner as an instrument, we may interpret $\gamma$ as the strength of an instrument assigned to individual $i$ relative to other potential instruments conditional on covariates (see Section 3.2 in kolesarcowles)

kolesarcowles shows that, for heterogeneous treatment effects, this estimand yields a convex combination of local average treatment effects between pairs of judges under pairwise monotonicity, exclusion restriction and exogeneity. In this paper, I restrict my attention to this estimand on account of its growing popularity and acceptance in the examiner IV literature. frandsen2023judging provide alternative estimators when the stronger versions of monotonicity and exclusion restriction do not hold, while Frandsen2023 provide another estimator that takes into account the fact that in some settings such as bail hearings in the US, batches of cases rather than individual cases are randomly assigned. My framework can accommodate the latter estimator under suitable conditions by simply including batch indicators in the covariate vector, $X$. On the other hand, unfortunately, it does not immediately extend to the estimator in frandsen2023judging. I leave this extension to future work.

An oracle estimator based on the parameter above would simply be:

\[\hat{\theta}=\frac{\mathbb{E}_n[Y \gamma]}{\mathbb{E}_n[T \gamma]}\] where $\mathbb{E}_n$ is the empirical analogue of expectation and the oracle tells us the value of $\gamma$. This estimator is infeasible because, of course, there is no oracle to tell us the value of $\gamma$. For the feasible estimator, a researcher replaces the conditionally ignorable treatment propensity,$\gamma$, with an estimate. Then the estimate of the treatment propensity is plugged into the second step to obtain a two-step estimator for the parameter of interest. kolesarcowles estimates the treatment propensity using an unbiased jacknife estimation approach described in Section 2 to avoid own observation bias. As will be discussed in the section on estimation, besides eliminating the need for Donsker conditions, cross-fitting also eliminates this own observation bias. However, as pointed out earlier, kolesarcowles's approach, among other limitations, excludes a growing range of applications where the number of covariates and examiner fixed effects is large relative to the sample size, a setting that is suitable for use of ML to estimate the treatment propensity.

The identifying moment function is the function, $g (W, \gamma, \theta)$ such that $\mathbb{E}[g (W, \gamma, \theta)]=0$ whenever $\theta=\theta_0$. Here $W$ represents a concatenated data observation: $W=(Y,T,X)$. It follows from rearranging kolesarcowles's IV estimand that: $$g (W, \gamma, \theta)=[Y-\theta T]\gamma$$

To see this, notice that $\mathbb{E}[(Y-\theta T)\gamma]=\mathbb{E}[Y \gamma]-\theta \mathbb{E}[T \gamma]=0$ whenever $\theta=\theta_0$. The validity of this moment condition also follows from instrument exogeneity: if $\gamma$ is a valid instrument, then the idiosyncratic error term in the outcome model should be orthogonal to the instrument, implying $\mathbb{E}[(Y-\theta T) \gamma]=0$.

The next important ingredient for the orthogonal moment condition is the influence function adjustment. Use of influence functions to correct for biases in first step estimation has a long history in econometrics (see newey1994asymptotic, hahnridder2013, hahn1998role, and ai2003efficient. The construction of the influence function adjustment in this section is based on the work of chernozhukov2022locally and ichimura2022influence, who provide a general form of the influence function, provided a certain “exogenous" orthogonality condition is satisfied. chernozhukov2022locally and ichimura2022influence show that if the following “exogenous" orthogonality condition is satisfied:

equation[equation omitted — 117 chars of source]

then under some regularity conditions, the influence function adjustment takes the following general form:

equation[equation omitted — 99 chars of source]

where $\lambda(w, \gamma(x))$ is a nonparametric residual function and $\alpha(x, \theta)$ is a special function called a Riesz representer.

In our setting, observe that the conditionally ignorable treatment propensity, $\gamma$ , is just a difference of two conditional expectations, which are just regressions of T on X and Z, and T on X. Let $\gamma_1=\mathbb{E}[T \mid X,Z]$ and $\gamma_2=\mathbb{E}[T \mid X]$. By definition of conditional expectations, $\gamma_1$ is the projection of the endogenous treatment variable, $T$, on the space of square integrable functions of $X$ and $Z$, $\Gamma_1=L^2(X,Z)$; while, on the other hand, $\gamma_2$ is the projection of $T$ on the space of square integrable functions of $X$, $\Gamma_2=L^2(X)$.

The nice geometry of $L^2$ yields the orthogonality condition in equation ((ref)). It readily follows from $\gamma_1$ being a projection of $T$ on $L^2(X,Z)$ that $\mathbb{E}(\delta(X,Z)(T - \gamma_1(F)(X,Z)) ) = 0$ for all $\delta \in L^2(X,Z)$. Analogously for $\gamma_2$, $\mathbb{E}(\delta(X)(T - \gamma_2(F)(X)) ) = 0$ for all $\delta \in L^2(X)$ where, only for the time being, I index $\gamma_j$ for each $j=1,2$ by the underlying true distribution of the data, $F$, to emphasize this holds under the true underlying distribution. Henceforth, I suppress this index to avoid notational clutter. This characterization shows that $\gamma_1$ and $\gamma_2$ satisfy the \enquote{exogenous} orthogonality condition of ichimura2022influence, which is equation $2.10$ in chernozhukov2022locally. Therefore, consistent with the notation of equation 2.10 in chernozhukov2022locally, the orthogonality condition can be written as:

align*[align* omitted — 172 chars of source]

Analogously:

align*[align* omitted — 165 chars of source]

Hence the whole theory for a generalized regression, $\gamma$ in chernozhukov2022locally goes through with modifications depending on the identifying moment function. Since $\gamma$ consists of two first step functions, $\gamma_1$ and $\gamma_2$, the influence function adjustment consists of two correction terms, one for each first step function (see ichimura2022influence, chernozhukov2022locally and newey1994asymptotic). The terms of the influence function adjustment for the conditional expectations, $\gamma_1$ and $\gamma_2$, take the following more explicit form of equation ((ref)):

equation[equation omitted — 117 chars of source]

where $\alpha=(\alpha_1,\alpha_2)$ for Riesz Representers, $\alpha_1$ and $\alpha_2$; $\gamma=h(\gamma_1,\gamma_2$) is a function of $\gamma_1$ and $\gamma_2$; $\lambda(W,\gamma_1(X,Z))=T-\gamma_1(X,Z)$ and $\lambda(W,\gamma_2(X))=T-\gamma_2(X)$ are nonparametric residuals after projecting the endogeneous treatment indicator, $T$ on the respective spaces of square integrable functions of $X$ and $Z$, and square integrable functions of $X$. Since $T$ is binary, the predictions of $\gamma_1(X,Z)$ and $\gamma_2(X)$ need to be bounded between $0$ and $1$, but there is no guarantee these would fall between $0$ and $1$. It thus seems eminently reasonable to use an influence function adjustment of the following form:

\[\phi(W, \gamma, \alpha, \theta)=\alpha_1(X,Z)(T-G(\gamma_1(X,Z))+\alpha_2(X)(T-G(\gamma_2(X))\] where $G$ is a logit function, in which case the true $\gamma_1$ and $\gamma_2$ functions are logit projections of $T$ on $\Gamma_1$ and $\Gamma_2$ respectively. While this is plausible, I do not explore this suggestion further in this paper, for clarity of exposition. In practice, ML algorithms can be constrained to “spit out" values between 0 and 1 for $\gamma_1$ and $\gamma_2$, so I use the original influence function adjustment in ((ref)) without loss of generality.

While explicit expressions of Riesz representers may be available in closed form or can be calculated, they sometimes take complex analytical shapes and often involve components whose estimation can lead to complications. For conditional expectations, these Riesz Representers typically take the form of projections (see newey1994asymptotic). In other settings such as the average treatment effect (ATE), they involve inverse propensity scores which, if the scores are near zero or one for some individuals, can lead to numerical instability. Recent literature provides a valid way of learning these Riesz Representers directly from data using only the orthogonal moment function and without knowing the closed form of the Riesz representers. I use this “automatic" approach in this paper and therefore I defer further discussion of the estimation of Riesz Representers to the relevant section later.

Now that we have the identifying moment function and the influence function adjustment, we can write down the orthogonal score: \[\psi(W, \gamma, \theta, \alpha)=[Y-\theta T]\gamma+\alpha_1(X,Z)(T-\gamma_1(X,Z))+\alpha_2(X)(T-\gamma_2(X))\]

Notice that this orthogonal score is a valid moment condition since it is zero in expectation. To see why, note that from our orthogonality argument above, $\mathbb{E}(\delta(X,Z)(T - \gamma_1(F)(X,Z)) ) = 0$ for all $\delta \in L^2(X,Z)$, and $\alpha_1$ is an element in $L^2(X,Z)$ (more on this in the next section). It immediately follows that $\mathbb{E}(\alpha_1(X,Z)(T-\gamma_1(X,Z)))=0$. Analogously, $\mathbb{E}(\alpha_2(X)(T-\gamma_2(X,Z)))=0$ by the fact that $\alpha_2$ is an element in $L^2(X)$. The identifying moment function is zero in expectation, as shown earlier. Furthermore, as shown in ichimura2022influence and chernozhukov2022locally, the orthogonal moment function (or score) is also an influence function (see Appendix B or ichimura2022influence and chernozhukov2022locally). To distinguish between this influence function characterization of the orthogonal moment function and the influence function adjustment above, ichimura2022influence and chernozhukov2022locally refer to the latter as the \enquote{first step influence function} to emphasize its role as a bias correction term for the first step estimation.

Neyman Orthogonality and Multiple Robustness

The orthogonal moment function we have constructed ought to satisfy two orthogonality properties. The first property, known as Neyman orthogonality, requires that varying the first step function away from its true underlying value should have no effect locally on the average orthogonal moment function. More formally (see equation 2.4 in chernozhukov2022locally):

equation[equation omitted — 184 chars of source]

where $\gamma_0$ is the value of $\gamma$ under the true underlying distribution of the data; $\delta$ represents a deviation of $\gamma$ away from $\gamma_0$; the scalar, $t$, is the size of the deviation; and the derivative above is evaluated at $t=0$. In a setting with multiple steps like ours, this derivative is taken with respect to $t$ for each first step function, holding the other first step functions fixed.

Let us hold $\gamma_2$ fixed. Since the moment identifying function, $g(W,\gamma_1, \theta)$ is a continuous functional linear in the first step function $\gamma_1$ which lives on the Hilbert space, $\Gamma_1=L^2(X,Z)$, then by the Riesz Representation Theorem, there exist unique random variables, $\alpha_{01} \in \Gamma_1$ such that $g(W,\gamma_1, \theta)=\mathbb{E}(\alpha_{01} \eta(X,Z))$ for all $\eta(X,Z) \in \Gamma_1$. Since $\gamma_1 \in \Gamma_1$, it follows that $\alpha_{01}=\mathbb{E}[Y-\theta T|X,Z]=\mathbb{E}[Y-\theta T|\Gamma_1]$. Let $\gamma_{01}$ be the probability limit of $\hat{\gamma_1}$. It follows that:

equation[equation omitted — 513 chars of source]

If $\gamma_1$ is held fixed and we allow the identifying moment function to vary in $\gamma_2$, we find a similar result using the above argument. I omit the details for brevity.

Another property which must be satisfied is the following:

\[ \mathbb{E}\left[\phi\left(W, \gamma_0, \alpha, \theta\right)\right]=0 \text { for all } \theta \in \Theta \text { and } \alpha \in \mathcal{A} \]

This follows from the arguments towards the end of the previous section.

Theorem 4 in chernozhukov2022locally allows us to verify whether our orthogonal moment function is \enquote{multiply} robust. Multiple robustness, a generalization of double robustness from the case of one first step function to multiple first step functions, is a stronger requirement than Neyman orthogonality. Neyman orthogonality is a \enquote{local} property: varying the first step functions should not have an effect locally on the average moment function. Multiple (or double) robustness imposes a more stringent requirement: the outcome model or one or more of the first step functions may be misspecified, and the orthogonal score will still remain valid as an estimating equation. Theorem 4 of chernozhukov2022locally provides a condition under which Neyman orthogonality and multiple (or double) robustness coincide, namely: the orthogonal moment function should be affine in the first step functions. Typically, verifying this condition entails inspecting the identifying moment function and the first step influence function adjustment. If both are affine in the first step function, then the orthogonal moment function is affine in the first step function and the equivalence between Neyman orthogonality and multiple robustness follows immediately by Theorem 4.

We have already shown Neyman orthogonality and we can easily verify, by sheer inspection, that both the identifying moment function and the influence function adjustment in our setting are affine in one first step function (say, $\gamma_1$), holding the other (say, $\gamma_2$) fixed. The multiple robustness of our orthogonal moment function, therefore, follows immediately by invoking Theorem 4 of chernozhukov2022locally. The implication of this result is that the orthogonal moment function proposed here is not only robust to biases from the first step estimation, but it is also robust to misspecification of either some of the components in the first step estimation or the outcome model.

Estimation

The Method of Moments Estimator via Cross-Fitting

The method of moments estimator we present in this section combines the multiply robust orthogonal score constructed above with cross-fitting. The estimator is a root to the following debiased sample moments, which are the empirical analogue of population moment conditions based on the orthogonal score:

align[align omitted — 376 chars of source]

The double summation operator in the above display denotes a cross-fitting procedure, which is a generalization of sample-splitting. In this procedure, we first randomly partition the sample into $L$ folds: $I_l$ for $l=(1,2 \ldots L)$. For each fold, $I_l$ , we use the observations in the complement of this fold, ${I_l}^c$, to estimate the Riesz representers, $\alpha_{1 l}$ and $\alpha_{2 l}$, and the first step functions, $\gamma_{1 l}$ and $\gamma_{2 l}$. Then given this set of estimates of $\alpha_{1 l}$,$\alpha_{2 l}$,$\gamma_{1 l}$ and $\gamma_{2 l}$, the sample moment function, $\hat{\psi}$, averages over the observations that are in the folds as in equation ((ref)). The debiased examiner IV estimator, $\hat{\theta}$, is a solution to these debiased sample moments.

Cross-fitting is associated with several nice theoretical properties. The main reason for cross-fitting is that it obviates the need for Donsker conditions which are critical in classical semiparametric theory but are not known to hold in high-dimensional settings with machine learning first steps (see the discussion in Appendix B or the original references: chernozhukov2018double, chernozhukov2022debiased and chernozhukov2022locally). Like sample splitting (see angrist1999jackknife), cross-fitting also eliminates “own observation bias" chernozhukov2022locally, which is a prominent issue in the examiner IV designs (see discussion in the Introduction). In our setting, cross-fitting further reduces bias from estimation of treatment propensities of examiners (see jochmans2023many). In part addressing concerns regarding use of machine learning in IV estimation raised by angrist2022machine, mikusheva2021many observed that cross-fitting also reduces the dependence between flexible first stage estimation of optimal instrument and second step parametric estimation of an IV model. Arguably, this dependence does not really arise in our setting, since our estimation of the conditionally ignorable treatment propensity is based on exogenous covariates.

For the estimation of $\gamma_1$ and $\gamma_2$, any regression learner such as LASSO, boosting, neural nets, random forests and other ML tools may be used, provided the learner converges fast enough (this is formalized in Assumption 3.2(2) in Section 3). These convergence rates have been ascertained for a wide range of machine learning first steps (see Section 3.2.3 in chernozhukov2024automatic for references). On the other hand, the estimation of the Riesz representers is a bit more complicated and I devote the next subsection to its exposition.

Automatic Estimation of the Riesz Representers

In this section, I discuss estimation of the Riesz representers, which is arguably one of the most formidable challenges when working with locally robust semiparametric approaches. I describe an “automatic" approach which I use in this paper to eschew some of the difficulties associated with estimating these functions in practice.

In general, while estimators of Riesz representers can be obtained by estimating their components where the closed forms of these Riesz representers are available, directly estimating Riesz representers this way is not always a good idea. Directly estimating Riesz representers for conditional expectations, for example, entails working with some form of projection, while Riesz representers for an average treatment effect usually involve estimating propensity scores. Even average policy effects involve estimating a density ratio. As discussed in section 3.1, this direct estimation of Riesz representers is often hard or is marred by some complications including numerical instability of the estimates. In this work, I avoid these challenges by adopting a more recent and increasingly popular approach which estimates these Riesz representers in an \enquote{automatic} fashion. This approach is \enquote{automatic} in the sense that it only requires knowledge of the orthogonal score and data for construction.

In moderate dimensional settings where the number of parameters increases slowly with the sample size, the Riesz representers can be approximated by “parametric sieves" and researchers may use estimation methods from the well-established literature on sieve estimation (see, for example, CHEN20075549 for a handbook-chapter treatment of this broad literature on sieve estimation) but, for the purpose of the exposition in this paper, I work with regularized Riesz representers in a high dimensional setting. Particularly, I propose using LASSO to estimate the Riesz representers, focusing on estimation of $\alpha(X,Z)$ (analogous arguments apply to $\alpha(X)$). To this end, we start with the sample counterpart of the Gateaux derivative that characterizes Neyman orthogonality. Holding $\gamma_2$ fixed, we obtain:

\[

aligned& \frac{d}{d t} \frac{1}{n-n_l} \sum_{l^{\prime} \neq l} \sum_{i \in l}(Y_i- \tilde{\theta}_{ll^{\prime}} T_i) \left(\tilde{\gamma}_{1ll^{\prime}} + t \delta - \gamma_2\right) + \alpha_1\left(T_i - \tilde{\gamma}_{1ll^{\prime}} - t \delta\right) - \alpha_2\left(T_i -\gamma_2 \right) \\ &=\frac{1}{n-n_l} \sum_{l^{\prime} \neq l} \sum_{i \in l}\left(Y_i-\tilde {\theta}_{ll^{\prime}} T\right) \delta-\alpha_1 \delta=\frac{1}{n-n_l} \sum_{l^{\prime} \neq l} \sum_{i \in l}\left((Y_i-\tilde {\theta}_{ll^{\prime}} T_i)-\alpha_1\right) \delta

\] where $n_l$ is the total number of observations in $L_l$ (so the average is over observations not in $L_l$: $n-n_l$), whereas the index, $ll^{\prime}$, denotes observations not in $I_l$ and $I_{l^{\prime}}$.

Notice that the last expression corresponds to the population moment condition, $\mathbb{E}[(Y-\theta )-\alpha_{01}) \delta]=0$. This orthogonality condition corresponds to $\alpha_{01}= \mathbb{E}[(Y-\theta T) \mid X, Z]$, which is consistent with what we obtained earlier. In the high dimensional setting that we consider here, I follow chernozhukov2022locally and represent $\alpha_1(X,Z)$ as a linear combination of a dictionary of functions, $\rho^{\prime}b(X,Z)$, where $b(X,Z)=(b_1(X,Z) \ldots b_p(X,Z))^{\prime}$ \enquote{spans} the Hilbert space, $\Gamma$, and choose $\delta(X,Z)=b_j(X, Z)$ for some $j^{th}$ element of the dictionary.

Our sample moment then becomes: \[ \hat{\varphi}_\gamma(b_j ; \rho^{\prime}b(X,Z))=\frac{1}{n-n_l} \sum_{l^{\prime} \neq l} \sum_{i \in l}((Y_i-\tilde {\theta}_{ll^{\prime}} T_i)-\rho^{\prime} b(X_i,Z_i)) b_j(X_i, Z_i) \] Adding an $L1$ penalty or regularization, we obtain the LASSO estimator: \[\hat{\alpha}_l(x,z)=\hat{\rho}^{\prime} b(x,z) \] where, for a suitably chosen penalty level, $r$:

equation[equation omitted — 234 chars of source]

for some initial estimate of $\theta$: $\tilde{\theta}_{l l^{\prime}}$.

To obtain this initial estimate, $\tilde{\theta}_{l l^{\prime}}$, we start by designating a number of folds, $L$, as in the cross-fitting procedure described earlier. Let $I_l$ and $I_{l^{\prime}}$ be two hold-out folds from the $L$ folds. Then we use observations that are not in $I_l$ and $I_{l^{\prime}}$ to estimate $\tilde{\theta}_{l l^{\prime}}$ using only the original identifying moment condition (see section 2.2 in chernozhukov2022locally for an example of this construction). We then use this initial estimate in the LASSO estimator above, which averages over all observations not in $I_l$ as required for the debiased sample moments in equation ((ref)), to estimate the Riesz representers. As argued by chernozhukov2022locally, this initial estimate does not affect the asymptotic distribution of the debiased examiner IV estimator because of Neyman orthogonality.

For the LASSO estimator in equation ((ref)), researchers have to \enquote{choose} the penalty term, $r$, and the dictionary of functions, $b(X,Z)$. The penalty term, $r$, can be computed via a cross-validation procedure described in Appendix A.1.1 of chernozhukov2022automatic, but other choices are also possible (see some examples in chernozhukov2022locally). In practice, $b(X,Z)$ could be a fully-interacted specification with all covariates and examiner fixed effects or types, polynomials of continuous covariates, interactions of all covariates and examiner types, and interactions among covariates, as in chernozhukov2022debiased.

Instead of the LASSO estimator, we could also use a Generalized Dantzig Selector (DGS) suggested in chernozhukov2022debiased. Other approaches for learning the Riesz representers from data involve minimizing a stochastic loss over more general function spaces in which the true Riesz representer is hypothesized to live. Leveraging critical radius theory, chernozhukov2020adversarial provide an adversarial estimator of the Riesz representer over more general function spaces such as neural networks, random forests and reproducing kernel Hilbert spaces (in particular, the neural tangent kernel space). While this is more general than the LASSO approach adopted in this paper, the adversarial loss formulation is computationally harder. We stick to the LASSO approach and leave this and other more general approaches to future work.

Asymptotic Theory

The large sample theory for the method-of-methods estimator proposed in subsection (ref) requires mild theoretical guarantees. First, we need the debiased sample moment function to converge at the rate of $\sqrt{n}$ to a sample average of the population orthogonal moment function. To show this, I prove mean square consistency conditions for $\hat{\gamma}$ and $(\hat{\alpha}_l, \Tilde{\theta}_l)$ and a rate condition on an interaction term. Second, I show the asymptotic distribution of the estimator. This section does not provide the statistical guarantees for the \enquote{automatic} estimators of Riesz representers from the previous section, as the generic mean-square convergence rates of such estimators have been sufficiently documented in chernozhukov2022automatic.

assumption[Mean Square Consistency] $(i)$ For each $i \in \{1,2\}$, $\hat{\gamma}_{i l}$ is a consistent estimator of ${\gamma_{0 i}}$ in the mean square sense such that $ \|\hat{\gamma}_{il}-\gamma_{0 i} \| \overset{p}{\to} 0$; $(ii)$ For each $i \in \{1,2\}$, $\hat{\alpha}_{i l}$ is a consistent estimator of ${\alpha_{0 i}}$ in the mean square sense such that $\|\hat{\alpha}_{i l}-\alpha_{0 i} \| \overset{p}{\to} 0$, where, in both $(1)$ and $(2)$, $\|a\|=\sqrt{\mathbb{E}[a(W)]}$ is the mean-square norm.
assumption[Uniform convergence and convergence rates for first steps] $(i)$ $\theta_o \in \Theta^\mathrm{o}$, where $\Theta^\mathrm{o}$ is the interior of the parameter space, $\Theta$; and for each $i \in \{1,2\}$, $\hat{\alpha}_{il}$ converges uniformly over $\Theta^\mathrm{o}$; $(ii)$ For each $i \in \{1,2\}$, either $\hat{\alpha}_{il}$ or $\hat{\gamma}_{il}$ converge faster than $\sqrt{n}$ such that: \[\left\|\hat{\alpha}_l\left(\tilde{\theta}_l\right)-\alpha_0\left(\theta_0\right)\right\|\left\|\hat{\gamma}_l-\gamma_0\right\| =O_p\left(\frac{1}{\sqrt{n}}\right)\]
theorem[$\sqrt{n}$ convergence of the debiased sample moments] Under assumptions (ref) and (ref) and the mean square consistency conditions on $\hat{\gamma}$ and $(\hat{\alpha}_l, \Tilde{\theta}_l)$ in chernozhukov2022locally, $\sqrt{n} \hat{\psi}\left(\theta_0\right)=\frac{1}{\sqrt{n}} \sum_{i=1}^n \psi\left(W_i, \theta_0, \gamma_0, \alpha_0\right)+o_p(1)$

Theorem (ref) is an asymptotic orthogonality result. In essence, it asserts that the estimation of $\hat{\gamma}$ does not have an asymptotic effect on the orthogonal moment function. The theorem also reaffirms the validity of the orthogonal moment function as an influence function, and clarifies the conditions under which this holds. This follows from the fact that any asymptotically linear and locally regular estimator can be represented as an average of an influence function (see Appendix B or kosorok2008introduction). For simplicity, I do not make a distinction between the initial estimate of $\theta$ obtained via the nested cross-fitting procedure in equation ((ref)) and the initial estimate, $\Tilde{\theta}_l$, in this theorem because asymptotically they are equivalent.

$\Gamma_1$ and $\Gamma_2$ encode all restrictions on the distributions of the data, and since $\Gamma_1=L^2(X,Z)$ and $\Gamma_2=L^2(X)$, then the distributions in this paper are general and unrestricted. From classical semiparametric theory, for general unrestricted distributions, the tangent space is the entire Hilbert space (see Theorem 4.4 in tsiatis2006semiparametric) and therefore the influence function in Theorem (ref) coincides with the efficient influence function.

To prove Theorem (ref), I invoke Lemma 8 in chernozhukov2022locally which holds under some mean-square consistency conditions. My proof strategy, therefore, consists mainly of showing that these mean-square consistency conditions hold for the non-parametric first steps in this paper. The proof is rather tedious and is relegated to Appendix A.

Theorem (ref) yields an important corollary: the asymptotic normality of the estimator. This corollary forms the basis for inference. We, however, need two additional assumptions.

Let $\sigma_{01}^2(X,Z)=\mathbb{E}\left[\left(T-\gamma_{01}(X,Z)\right)^2 \mid X, Z \right]$ and $\sigma_{02}^2(X)=\mathbb{E}\left[\left(T-\gamma_{02}(X)\right)^2 \mid X \right]$.

assumption[Boundedness] $(i)$ For $i \in \{1,2\}$, $\alpha_{0 i}$ and $\alpha_{i l}$ are bounded; $(ii)$ $\sigma_{01}^2(X,Z)$ and $\sigma_{02}^2(X)$ are bounded, and $\mathbb{E}[(g(W, \alpha,\gamma_0, \theta_0)^2] < \infty$

We need another assumption that ensures that $\Bar{Q}_n=\partial \hat{\psi}(\bar{\theta}) / \partial \theta \rightarrow_p Q=\partial \hat{\psi}(\bar{\theta}) / \partial \theta_0$ for any $\bar{\theta} \rightarrow_p \theta_0$. Treating $\gamma$ as one function rather than a difference of two functions, we impose this assumption, which is an adaptation of Assumption 5 in chernozhukov2022locally:

assumption$Q$ exists and there is a neighborhood $\mathcal{N}$ of $\theta_0$ such that $\left.i\right)\left\|\hat{\gamma}-\gamma_0\right\| \rightarrow_p 0$; ii) for all $\left\|\gamma-\gamma_0\right\|$ and $\left\|\alpha-\alpha_0\right\|$ small enough $\psi\left(W, \gamma, \alpha, \theta\right)$ is differentiable in $\theta$ on $\mathcal{N}$ with probability approaching one; and there is $C>0$ and $d\left(W, \gamma, \alpha\right)$ such that $\mathbb{E}\left[d\left(W, \gamma, \alpha\right)\right] \leq C$ and such that for $\theta \in \mathcal{N}$ and $\left\|\gamma-\gamma_0\right\|$ and $\left\|\alpha-\alpha_0\right\|$ small enough $$ \left|\frac{\partial \psi\left(W, \gamma, \alpha, \theta\right)}{\partial \theta}-\frac{\partial \psi\left(W, \gamma, \alpha, \theta_0\right)}{\partial \theta}\right| \leq d\left(W, \gamma, \alpha\right)\left|\theta-\theta_0\right|^{1 / C}$$ (iii) For each $\ell=1, \ldots, L, \int\left|\partial \psi\left(w, \hat{\gamma}_{\ell}, \theta_0\right) / \partial \theta-\partial \psi\left(w, \gamma_0, \theta_0\right) / \partial \theta\right| F_0(d w) \xrightarrow{p}$ 0.
corollary[Asymptotic Normality of the Method of Moments Estimator] Under assumptions 3.1 to 3.4, the MoM estimator is asymptotically normal: $$ \sqrt{n}\left(\hat{\theta}-\theta_0\right) \xrightarrow{d} N(0, V) \quad \text { and } \quad \hat{V} \xrightarrow{p} V $$ where $V=Q^{-2}\operatorname{Var}[\psi^2]=Q^{-2}\mathbb{E}[\psi^2]$ and $\hat{V}$ is a consistent estimator of $V$.

The proof of this corollary is also relegated to Appendix A for the interested reader. Heuristically, the expression of the variance comes from the fact that the orthogonal moment function is the efficient influence function, combined with standard arguments from semiparametric theory. In particular, the influence function for the examiner IV estimator is proportional to the influence function for the average moment function (cf: section 21.1.4 in kosorok2008introduction). Therefore, the variance of the examiner IV estimator is also proportional to the variance of the average moment function which, in turn, is the second moment of its influence function. With this asymptotic normality result, construction of confidence intervals and test statistics follows in the standard manner.

Corollary (ref) presumes $\sqrt{n}$ consistency of the estimator, $\hat{\theta}$. Primitive conditions for this consistency are mild and are provided in Appendix F of the Supplemental Material in chernozhukov2022locally. In a nutshell, consistency of the estimator requires, inter alia, the mean square consistency conditions of Assumption 1 in chernozhukov2022locally, which I have proved for Theorem 3.1 in our setting; the rate condition in Assumption 2 in chernozhukov2022locally, which I have also proved for Theorem 3.1; and standard consistency assumptions, such as compactness of the parameter space, from NEWEY19942111.

Conclusion

This paper has provided a new framework for estimation and inference in examiner IV designs when there are many examiners and the examiner IV is valid conditional on possibly a large number of covariates relative to the sample size. The estimation method permits applied researchers to use a wide range of machine learning tools such as LASSO, random forests, neural networks and Dantzig in the first step estimation of the examiner IV. Based on the locally robust semiparametric theory of chernozhukov2022locally and ichimura2022influence, I provide mild theoretical guarantees for such machine learning first steps and the two-step examiner IV estimator, including root-n consistent estimation of the examiner IV estimator with mean-square consistency conditions on the first steps.

As an avenue of future research, it would also be interesting to explore locally robust semiparametric approaches for settings where examiners administer more than one treatment, such as in kamat2023identification. Furthermore, the method developed in this paper presumes that the canonical identification assumptions of these examiner IV designs obtain, but it might be instructive to extend this framework to settings where some of these assumptions do not hold, a la frandsen2023judging.