EconBase
← Back to paper

Leniency Designs: An Operator's Manual

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,110 characters · 6 sections · 80 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Leniency Designs: An Operator's Manual

High-stakes treatments are often determined by expert decision-makers: doctors decide whether to treat patients, bail judges decide whether to release defendants before trial, patent examiners decide whether to grant patents to firms, child welfare investigators decide whether to place children in foster homes, and loan officers decide whether to approve consumer loan applications. Usually, of course, decision-makers disagree on the appropriate course of action; some are more lenient than others, in that they tend to grant the treatment more often. Furthermore, which decision-maker is assigned to a given case is often either deliberately random (to ensure fairness), or, due to idiosyncrasies in the assignment process (such as rotations in shifts), as-good-as-random. A leniency design (also known as a judge or examiner instrument design) harnesses such exogenous variation to estimate the causal effects of the high-stakes treatment by using a measure of the decision-makers' leniencies as an instrument for the treatment in an instrumental variables regression. Such designs have exploded in popularity in recent years.\footnote{Prominent applications include dobbie2015debt,dobbie2018effects,dobbie2021measuring, farre2020patent, autor2010temporary, doyle2007child, mullainathan2022diagnosing. See Table 1 of frandsen2023judging for more.}

Leniency designs rest on a straightforward idea: many cases are close calls where a lenient decision-maker would say yes, but a stricter one would say no. If decision-makers are randomly assigned, assigned leniency is also as-good-as-random. And if assignment only affects outcomes through the treatment decisions, leniency is a valid instrument. But when implementing this logic, several practical questions arise. How exactly should we measure decision-maker leniency? What controls do we need to include in the analysis? Do we need to cluster standard errors, and if so, at what level? Does the fact that leniency is estimated matter for inference? And if we think treatment effects are heterogeneous, what are our estimates even capturing?

This paper develops a step-by-step guide to leniency designs, drawing on recent econometric literatures. The starting point of this operator's manual is a link between leniency designs and more conventional instrumental variable regressions, in which indicators for decision-maker assignment are used directly as instruments. This link follows from observing that, with no controls in this more conventional specification, the population first-stage regression of treatment on the assignment indicators gives the ideal leniency measure. Hence, the more conventional specification can be understood as also using a leniency measure as the instrument. Complementing a recent review by chynexaminer, we use this link to to develop answers to the above questions via a robust way of estimating the more conventional instrumental variable regression.

In particular, we show how the unbiased jackknife instrumental variables estimator (UJIVE) of kolesar13late is purpose-built for leveraging exogenous leniency variation while avoiding subtle biases from other leniency constructions with many decision-makers or controls. We further explain how the local average treatment effect framework of ImAn94 can be used to interpret UJIVE estimates under arbitrary treatment effect heterogeneity. As it turns out, UJIVE is not only useful for estimating treatment effects in this framework: building on abadie02 and kitagawa15, we show how it can also be used to test a key identifying assumption (average monotonicity). Finally, we show how knowledge of the decision-maker assignment process---or “design”---can guide standard error calculations, and that the heteroskedastic-robust (i.e., non-clustered) standard errors calculated by UJIVE are often appropriate.

We conclude with a practical checklist for applying examiner designs, assessing key assumptions, and probing external validity, all with a unified UJIVE estimation approach. We illustrate this checklist with a re-analysis of farre2020patent, who use quasi-random patent examiner assignment to estimate the value of patents to startups. Patent-granting significantly increases future patent applications, approvals, and citations on both the extensive and intensive margin, though the magnitude and precision of these estimates are sensitive to using the more robust UJIVE estimator. The key average monotonicity condition appears to hold, and the complier subpopulation which contributes treatment effects to the UJIVE estimates appears representative of the broader study population.

Motivating Example: Estimating the Value of Patents

We start by considering the simplest possible version of a leniency design, with two decision-makers who are completely randomly assigned. Concretely, consider a stylized version of the setting in farre2020patent, who again estimate the effect of patent-granting to startups on the firms' later inventiveness. We observe a set of $n$ applications, indexed by $i$, which are submitted by firms to the US Patent Office. The applications are then assigned by a random coin flip to one of two examiners, $s$ and $t$. The examiners decide whether to grant the patent, as indicated by the dummy variable $x_i\in\{0,1\}$. We refer to $x_i$ as treatment, following standard causal inference lingo. An outcome $y_i$ is then realized; for concreteness, let's suppose $y_i$ measures the number of future patent applications filed by the firm.

We want to leverage the randomness in examiner assignment to estimate the causal effect of $x_i$ on $y_i$. For now, we assume this effect is the same across all firms: being granted a patent causes each firm to submit $\beta$ additional patent applications, compared to being denied a patent. In practice, it is likely that the effects of patent-granting are heterogeneous (i.e., different for different firms); we allow for this possibility later. Ignoring treatment effect heterogeneity for now, we write an outcome equation for $y_i$ as

equation[equation omitted — 77 chars of source]

where the intercept $\gamma$ represents the average number of additional patent applications submitted when a firm is denied a patent. The actual number of additional patent applications varies across firms due to various unobservable factors captured by the error term $\varepsilon_{i}$.

The basic identification challenge is that estimating (ref) by simple ordinary least squares need not reveal $\beta$ because of selection bias: startups with high $\varepsilon_{i}$ who would produce many patents in the future even if their current application were denied (because, say, they have higher research and development budgets or employ better scientists) may submit stronger applications that are more likely to be granted. This results in a positive correlation between the error term and the treatment $x_i$, which, in turn, induces a positive selection bias: a least-squares regression of $y_i$ on $x_i$ will tend to overstate the true causal effect $\beta$.

A leniency design addresses selection bias by leveraging differences in the randomly-assigned examiners' tendency to grant patents, isolating variation in the treatment that is unrelated to $\varepsilon_{i}$. To formalize this, suppose that examiner $t$ is “tough”, while examiner $s$ is “soft” in that the overall patent-granting rate of examiner $s$ in the population of patent applications (which we refer to as their leniency) exceeds that of examiner $t$: ${p}_s>{p}_t$. Let $\hat{p}_s$ and $\hat{p}_t$ be in-sample estimates of these rates, corresponding to the fraction of patents we observe each examiner granting in the data. Further, let $z_i$ denote the dummy for being assigned to the soft examiner, so it equals $1$ if application $i$ is handled by examiner $s$ and $0$ if it is handled by examiner $t$. Then the estimated leniency of the examiner assigned to $i$ is given by $\hat{\ell}_i=\hat{p}_{t}+(\hat{p}_{s}-\hat{p}_{t})z_i$. A leniency design uses this estimated leniency as an instrument for patent-granting $x_i$. Specifically, we estimate $\beta$ with an instrumental variable (IV) regression, instrumenting $x_i$ with $\hat{\ell}_i$ in the outcome equation.

To understand why this IV regression works, note that here---with only two randomly assigned examiners---it is equivalent to a simpler IV regression which uses the examiner assignment dummy $z_i$ directly as an instrument. In other words, it does not matter whether we use the estimated leniency $\hat{\ell}_i$ as an instrument or the dummy $z_i$: the resulting estimates of $\beta$ are numerically identical. This is because the estimated leniency is just a version of $z_i$ that is scaled by $\hat{p}_{s}-\hat{p}_{t}$ and shifted by $\hat{p}_{t}$, and such instrument scaling and shifting leaves IV estimates unchanged.

Note further that $z_i$ is randomly assigned, like treatment assignments in a simple randomized controlled trial. Thus, so long as examiner assignment only affects the outcome $y_i$ through the patent-granting treatment $x_i$---the usual IV exclusion restriction---$z_i$ will be independent of $\varepsilon_{i}$ and therefore be a valid instrument for estimating $\beta$.\footnote{Consistency of these IV estimates also requires a relevance condition: that $z_i$ and $x_i$ have non-zero correlation. Here relevance holds because we assumed that the examiners vary in their leniency, $p_s>p_t$. This means that the population first-stage coefficient from regressing $x_i$ on $z_i$, given by $p_s-p_t$, is non-zero.} Exclusion seems reasonable to assume here, unless we thought examiners did other things besides ruling on the patent (such as giving advice to the firms). Furthermore, because $z_i$ is randomly assigned by a coin flip for each application, it is uncorrelated across applications. It follows that conventional heteroskedasticity-robust standard errors are appropriate for quantifying uncertainty in these IV estimates. There is no need to cluster the standard errors, analogous to how we don't need to cluster them in simple randomized controlled trials. By the numerical equivalence between the leniency version of the IV and the dummy version of it, the same is true when we use the estimated leniency as an instrument.

Real-world leniency designs are naturally more complicated than this stylized example, with more than two (and often many) decision-makers who are not assigned completely at random. Handling these features requires modifications to the above simple IV approach, as we next discuss. Still, the core equivalence between using estimated leniency versus examiner dummies directly as instruments will also prove a useful guide to thinking through estimation and inference in these more complex leniency designs.

Estimation with Many Decision-Makers and Controls

Decision-maker assignment is rarely completely random, but institutional knowledge can sometimes imply it is as-good-as-random once we condition on an appropriate set of control variables. We refer to these as necessary controls, as they must be included in every specification leveraging exogenous assignment. As we discuss more below, patent applications in the US are first classified depending on the technology being patented. This determines which group of specialist examiners---the so-called “art unit”---will review the application. Assignment within art units is then effectively random, conditional on the set of examiners working in the art unit at the time of the assignment. Assuming this set changes slowly over time, a researcher may take art unit-by-year fixed effects as the set of necessary controls.\footnote{One may optionally include precision controls in some specifications: pre-assignment characteristics of the firm that help explain variation in the outcome. Including these variables will soak up some variability in the error $\varepsilon_{i}$ and thus help reduce the standard errors. This is the IV analog of including baseline characteristics as controls in regression specifications estimating effects of treatments in randomized controlled trials, while including the necessary controls is analogous to controlling for strata indicators in randomized trials where the treatment assignment probability varies across strata.}

With many examiners, assignment is now captured by a vector $z_i$ of $K$ instruments: one for each examiner, with one examiner omitted in each art unit to prevent multicollinearity. This leads to a first-stage equation, a population regression specification linking the treatment, instrument, and controls:

equation[equation omitted — 68 chars of source]

where the vector $w_i$ comprises art unit-by-year dummies and $\nu_i$ is a regression residual. In the special case where we just have one year of data, so that $w_i$ reduces to art unit fixed effects, the first-stage regression coefficients have a simple interpretation: $\delta_j$ gives the overall leniency of the omitted reference examiner in art unit $j$, while $\pi_k$ measures the difference between the leniency of examiner $k$ and the omitted reference examiner in the art unit that $k$ works in.

If we knew the first-stage coefficients $\pi$ and $\delta$, we could use the population first-stage fitted values $\ell_{i} = z_i'\pi + w_i'\delta$ as our leniency measure. By the Frisch-Waugh-Lovell theorem, an IV regression using this instrument while controlling for $w_i$ is equivalent to running an IV regression without any controls, but instrumenting with the residual from projecting the fitted values onto the covariates, $\tilde{\ell}_i=\tilde{z}_i'{\pi}$. Here $\tilde{z}_i$ denotes the sample residual from regressing the instrument vector $z_i$ onto the controls $w_i$. This IV regression leads to the estimator:

equation[equation omitted — 111 chars of source]

the ratio of sample covariances between relative leniency $\tilde{\ell}_i$ and the outcome $y_i$ versus the treatment $x_i$. We use the term relative leniency to stress that $\tilde{\ell}_i$ measures the leniency of examiner handling application $i$ relative to other examiners who could have handled the application. In contrast, $\ell_i$ is a measure of absolute leniency. With just one year of data, these leniency measures take a simple form: $\ell_{i}$ is the overall patent-granting propensity of the examiner assigned to $i$, while $\tilde{\ell}_i$ is the examiner's patent-granting rate minus the overall patent-granting rate of the art unit handling application $i$.

Since examiner assignment is as-good-as-random given the controls $w_i$, and since we are netting out the differences across the controls by using the residual variation in examiner assignment $\tilde{z}_i$, relative leniency is itself as-good-as-randomly assigned. Provided the exclusion restriction holds, $\tilde\ell_i$ will therefore be uncorrelated with the outcome error $\varepsilon_i$ as in the simple motivating example. By definition of a regression residual, $\tilde\ell_i$ is also uncorrelated with the residual $\nu_i$ in the first-stage equation. These two facts imply that the expectation of the denominator in (ref) equals $\sum_i\tilde\ell_i^2$, while the expectation of the numerator equals $\beta\times \sum_i\tilde\ell_i^2$.\footnote{Here, and below, following the econometric literature, we implicitly condition on the observed instruments and covariates $z_i$ and $w_i$ such that randomness only comes from unobserved errors $\varepsilon_i$ and $\nu_i$. (ref) details this setup and gives formal derivations of all claims in this section.} Thus, approximating the expectation of a ratio by a ratio of expectations, we conclude that using the relative leniency instrument leads to an approximately unbiased estimate of $\beta$.

In practice, it is infeasible to use the true relative leniency as an instrument because we do not know the first-stage coefficients. A tempting alternative is to simply use least squares estimates of the first-stage coefficients, $\hat{\pi}$ and $\hat{\delta}$, in place of the unknown population values. This is exactly equivalent to using a two-stage least squares (2SLS) estimator which instruments with the full set of assignment dummies $z_i$. To see this, recall that the 2SLS estimator is equivalent to an instrumental variables estimator that uses a single instrument given by the sample first-stage fitted values, $z_i'\hat{\pi}+w_i'\hat\delta$. By the Frisch–Waugh-Lovell theorem, this is in turn equivalent to an IV regression without covariates that replaces the population relative leniency $\tilde\ell_i$ in (ref) with the sample analog $\hat{\ell}_i=\tilde{z}_i'\hat{\pi}$ (the residual from projecting the fitted values onto the controls).

In contrast to using the true relative leniency $\tilde{\ell}_i$, however, using the estimated relative leniency $\hat{\ell}_i$ will tend to produce biased treatment effect estimates when there are many examiners. This is the classic many-weak instrument bias problem for 2SLS bekker94,bound1995problems; it arises here because the relative leniency measure $\hat{\ell}_i$ is constructed in part from the treatment status of observation $i$ through the least squares estimate $\hat{\pi}$. With just one year of data, for instance, $\hat{\ell}_i$ is given by the fraction of patents granted by the examiner assigned to $i$ in the sample, minus the fraction of patents granted by the art unit handling the application---and both fractions are computed including the data from application $i$. But applicant $i$'s treatment likely correlates with the outcome error $\varepsilon_i$; this is, after all, the reason to use IV instead of simple least squares regression. Thus, the estimated leniency instrument is also likely correlated with the outcome error, and this generates bias in IV estimates in the direction of the naïve ordinary least squares estimates.

The bias from using the estimated leniency instrument can be severe, especially when examiner assignment only modestly affects the probability of treatment. A helpful rule of thumb for gauging the magnitude of the bias obtains when we assume the errors are homoskedastic; then, the 2SLS bias approximately equals the bias of ordinary least squares times $1 \big/ \left((1-R^2)\times E[F]\right)$, where $E[F]$ is expectation of the first-stage $F$ statistic for the hypothesis that the first-stage instruments are irrelevant (i.e., that $\pi=0$ in (ref)), and $R^2$ is the population partial R-squared from adding instruments to the first-stage regression. In practice, since $R^2$ tends to be near zero, the magnitude of $E[F]$ tells us how much smaller 2SLS bias is relative to ordinary least squares (it's never larger, since $E[F]$ always exceeds one). In fact, this relationship can be used to justify the popular $F>10$ rule of thumb proposed by staiger1997instrumental for identifying weak instruments. If the expectation of the first-stage $F$ statistic lies below 10, then 2SLS bias exceeds 10% of least squares bias, so the rule of thumb can be thought of as a diagnostic for whether we are in this large bias region.\footnote{\Textcite{StYo05} construct a formal statistical test based on $F$ for the null hypothesis that the 2SLS bias is large in the sense of potentially exceeding 10% of least squares bias, under homoskedastic errors. While the precise critical value depends on $K$, it is close to 10 for most values of $K$: see their Table 5.1.} With many examiners, the first-stage $F$ statistic can be modest even when examiner leniency explains economically meaningful variation in the treatment, because the formula for $F$ divides by $K$.

Knowing that the 2SLS bias comes from using own treatment status to estimate the relative leniency suggests a natural bias-free alternative: instrument with a leniency estimate $\hat{\ell}_{-i}$ that leaves out observation $i$'s own value of the treatment, thereby avoiding a mechanical correlation with $\varepsilon_i$. The unbiased jackknife instrumental variable estimator (UJIVE), studied in kolesar13late, implements this logic by setting the relative leniency estimate to $\hat{\ell}_{-i}=\tilde{z}_i^\prime\hat\pi_{-i}$ where $\tilde{z}_i$ are the same instrument residuals as before and $\hat\pi_{-i}$ is a least-squares estimate of $\pi$ which uses all observations except for $i$. Because this leave-one-out---or “jackknifed”---estimate is unbiased for $\pi$ and, in iid data, independent of the data for observation $i$, the same arguments used to show approximate unbiasedness of the infeasible estimator $\hat\beta^*$ can also be used to show approximate unbiasedness of the UJIVE estimator.

The unbiasedness property of UJIVE hinges on the fact that it uses leave-out estimation to directly estimate relative leniency $\tilde{\ell}_i$ accounting for controls through the residualized $\tilde{z}_i$. Leave-out estimation of absolute leniency $\ell_i$, in contrast, tends to work poorly when the number of covariates is large. In particular, PhHa77,aik99 propose an estimator known as JIVE (the jackknife IV estimator) that constructs a leave-out estimate of the absolute leniency and uses it as an instrument in a regression with covariates. By the Frisch-Waugh-Lovell theorem, this is equivalent to 2SLS without controls and instrumenting with least squares residuals from the regression of the leave-out absolute leniency estimate on the controls. Concretely, in a case with one year of data, the residuals of the leave-out absolute leniency are created by subtracting off the average leave-out-leniency estimates for observations in the art unit that handles application $i$---but this is just the (non-leave-out) average patent-granting rate of the art unit, which depends on the treatment status of application $i$. The subtraction thus reintroduces the own-observation bias of 2SLS, except that it typically runs in the opposite direction to 2SLS (since the own-observation component is now only the term being subtracted off). Under homoskedasticity, the JIVE bias approximately equals that of 2SLS multiplied by $-E[F] \times L/((E[F]-1)K-L)$ where $L$ is the number of controls. The JIVE bias is negligible when the number of controls is much smaller than the number of instruments (in fact, UJIVE and JIVE coincide in the absence of controls and a constant). But when many controls (e.g., art unit-by-year fixed effects) are needed to extract the randomness in many examiner assignment indicators, this formula shows that the magnitude of JIVE bias can be comparable to the bias of 2SLS\@.

While UJIVE is a natural solution to the 2SLS bias problem, there are other reasonable approaches. \Textcite{AcDe09} propose a clever modification of JIVE---the improved jackknife IV estimator (IJIVE)---which can greatly reduce its bias in the presence of controls by reversing the order of operations: first residualize the instruments and then compute a leniency measure based on leave-out fitted values. While reversing the order of operations doesn't entirely eliminate the own-observation bias, in practice IJIVE and UJIVE tend to produce similar estimates. More recently, csw23 propose an alternative to UJIVE, termed FEJIV\@. We show in the appendix that FEJIV can be interpreted as solving for a relative leniency measure with the smallest mean squared error under homoskedasticity, subject the constraints that the resulting estimator is free of own-observation bias and that the leniency measure is orthogonal to the covariates. The UJIVE leniency measure, in contrast, achieves the minimal mean squared error without the latter orthogonality constraint. The orthogonality property of FEJIV is attactive in that adding a linear function of covariates to the outcome does not affect the estimate; the price for this is that FEJIV leniency is slightly noisier. The resulting estimator also tends to be computationally demanding in large datasets, and imposes stronger data requirements (e.g., the estimator may not exist if there are high-leverage observations).

Another approach is to bias-correct the 2SLS formula, which replaces the infeasible instrument $\tilde{\ell}_i$ in (ref) with the estimated relative leniency $\hat{\ell}_i$. When 2SLS does this, it increases the expectation of the denominator from $\sum_i\tilde{\ell}_i^2$ to $\sum_i\tilde{\ell}_i^2+K\operatorname{var}(\nu_i)$ under homoskedasticity. The quantity $\sum_i\tilde{\ell}_i^2$ corresponds to the numerator of the population partial R-squared statistic; it measures the increase in the explained sum of squares from adding instruments to the first stage. Since 2SLS uses in-sample fitted values to estimate the predictive power of the instruments, it overstates this predictive power, exactly analogous to how the unadjusted R-squared overstates the predictive power of regression (the same issue happens in the numerator, and taking the ratio gives the rule-of-thumb bias formula for 2SLS). Paralleling how the adjusted R-squared fixes the issue by using a degrees of freedom correction, we can use degrees of freedom adjustments in both the denominator and the numerator of the 2SLS formula. Unfortunately, similar to the adjusted R-squared formula, the resulting bias-corrected 2SLS estimator (dating back to nagar1959bias) works only under homoskedasticity.

In practice, rather than using UJIVE or its cousins that compute the leniency measures internally, it is common for researchers to first construct an “external” leniency measure and then use it as an instrument in a just-identified IV regression. This is analogous to first computing first-stage fitted values and then using them as a single instrument, rather than directly using the standard 2SLS estimator which estimates everything in one step. Such “manual 2SLS” procedures are widely advised against angrist2009mostly and we would similarly advise against manual leniency IV estimation. Indeed, as shown above, subtle variations in how the leniency measure is exactly constructed and accounts for covariates can have potentially large consequences for the bias of the ultimate estimator. It is thus simpler and safer to use a one-step implementation like UJIVE, with established unbiasedness properties.

Two more comments on estimation are warranted before proceeding. First, it is common practice to report first-stage $F$ statistic with 2SLS estimates; as we have discussed, its expectation, $E[F]$, directly informs about 2SLS bias. But small first-stage $F$ statistics need not worry a researcher using UJIVE\@: our arguments for its approximate unbiasedness work even if $E[F]$ is small. In fact, UJIVE remains approximately unbiased and consistent even when the instruments are weak enough that $E[F]$ converges to zero in large samples, so long as $\sqrt{K}\times E[F]$ is large MiSu22. Second, the leave-out estimation approach leveraged by UJIVE assumes independence across observations. When observations are instead clustered together---a scenario we discuss more when discussing the calculation of standard errors---a leave-own-cluster-out modification of UJIVE may be needed to ensure unbiasedness frandsen2025cluster. Intuitively, correlations between treatment $x_i$ and errors $\varepsilon_j$ for observations $i$ and $j$ in the same cluster will generally reintroduce the mechanical bias between the leave-own-observation-out UJIVE instrument and the errors, again tending to bias estimates towards ordinary least squares. Hence, for examiner designs, clustering in the data is not just a consideration for inference but may also guide the choice of estimator.

Heterogeneous Treatment Effects

So far, we have assumed the parameter of interest is a constant treatment effect $\beta$. This is of course a strong assumption in practice. It means, for example, that any successful patent application raises the future innovativeness of all startups $i$ by the same constant amount. In reality, the value of patents is likely higher for some startups than for others. Luckily, the UJIVE estimator can retain a causal interpretation even in the presence of such heterogeneous effects---i.e., when patent effects $\beta_i$ vary arbitrarily across firms $i$.\footnote{In this section only, we assume the treatment is binary as in much of the literature on heterogeneous treatment effects. This is mostly for notational convenience, since otherwise effect heterogeneity could come both from differences across observations and from different margins of treatment response for a given observation. See kp25 for an extension of the local average treatment effect theorem to the case with multivalued or continuous treatment.}

The Nobel Prize--winning local average treatment effect theorem of ImAn94 offers a guide to the general heterogeneous-effect interpretation of leniency designs. The theorem maintains our earlier assumption of as-good-as-random decision-maker assignment and the IV exclusion restriction, while replacing the restriction of constant treatment effects with an assumption of first-stage monotonicity. In the patent setting, monotonicity means that if we can find a case that some examiner $a$ would grant a patent to but some other examiner $b$ would deny, it must be that $a$ is more lenient in all cases; i.e., there cannot be another case that $a$ would deny but $b$ would grant. In the simple example with just two examiners, this implies we can divide the universe of patents into two groups: a complier group of firms with patent applications that are granted when assigned to the soft examiner $s$ but not when assigned to the tough examiner $t$, and a non-responder group whose treatment status is unaffected by examiner assignment (they are always either granted or denied, regardless of which examiner handles the case). Monotonicity rules out the presence of a third group of defier firms that are granted their applications only if assigned to $t$. The local average treatment effect theorem states that, in the absence of defiers, the IV regression using an indicator variable for assignment to the soft examiner as an instrument estimates the average treatment effect for compliers. \Textcite{ImAn94} call this a local average treatment effect, to stress that we learn nothing about treatment effects for non-responders. In contrast, if we ran a randomized controlled trial that granted applications by a coin flip, we would learn the overall average treatment effect $E[\beta_i]$.

With many as-good-as-randomly assigned decision-makers, monotonicity implies there are many complier groups: one for each pair of decision-makers. In this case, an extension of the theorem shows that using the relative leniency $\tilde{\ell}_i$ as an instrument---or an approximately unbiased feasible version such as UJIVE---identifies a weighted average of complier-group specific treatment effects, weighted using the squared leniency differences of each decision-maker pair.\footnote{The appendix states the formal result. A general statement appears, for example, in kolesar13late, who builds on ImAn94 by allowing for linear controls $w_i$.} This is so because, in the population, leniency IV is equivalent to running a series of simple IV regressions with a single binary instrument that restricts the sample to a given decision-maker pair and then averaging these IV regressions together using squared pairwise leniency differences as weights. Since a leniency difference measures the share of compliers, leniency IV thus weights by the square of the group size. As a result, larger complier groups receive more weight. But importantly, the weighting is convex: i.e., there are no complier groups given negative weight.

While it is likely more palatable than assuming constant treatment effects, monotonicity is nevertheless a strong assumption. In particular, results in vytlacil02 imply it is equivalent to assuming each patent examiner has the same ranking of the appropriateness for granting patents; the examiners only differ in the cutoff they use for gauging whether the application is above the bar. The strength of this assumption for leniency designs was, in fact, first noted in the original ImAn94 analysis (see their Example 2, p. 472). There are two core issues. First, in many settings like our patent example, there are multiple dimensions to decision-makers' evaluation criteria which may make it unlikely that any two decision-makers would have the same ranking. For instance, patent applications are evaluated by their perceived usefulness, novelty, and non-obviousness, and different examiners likely place different weights on these criteria. Second, even if the weights were the same, variation in skill among decision-makers can induce ranking differences due to mistakes by less-skilled decision-makers. For instance, \Textcite{cgy22} show that in the case of radiologists variation in skill accounts for about 40% of variation in decision-maker leniency.

Luckily, first-stage monotonicity can be both weakened and tested. Without monotonicity, a simple IV regression restricted to a pair of decision-makers identifies a weighted difference between complier and defier treatment effects with weights given by the fraction of compliers and defiers (because defier treatments are shifted in the opposite direction by decision-maker assignment, relative to the compliers). Since leniency IV averages these pairwise IV regressions, it can be seen as identifying a weighted-average treatment effect for compliers and defiers, but with negative weight put on all defier groups. This is a problem for its causal interpretation, in general. For example, we risk “sign-reversals”: even if patent-granting, say, uniformly increases future firm productivity so that $\beta_i$ is positive for all $i$, the leniency IV could estimate a negative effect---namely, if the treatment effect for defiers is much larger than that for compliers.

This motivation for the monotonicity assumption suggests that it can be weakened in three ways. First, while the presence of defiers necessarily causes a negative weighting problem in the case with a single pair of examiners, there is a subtlety with multiple examiners. Since each firm appears in multiple pairwise examiner comparisons, the total weight we place on each firm corresponds to its average complier, or defier, status across all pairwise comparisons. Thus, as long as no firm is a defier “on average” the weights placed on individual firms remain positive. This is equivalent to assuming that the average leniency of the examiners who would grant the firm a patent exceeds the average leniency of those who would deny it, a condition frandsen2023judging term “average monotonicity.”\footnote{While the statement of the condition in frandsen2023judging is more complicated, we show in the appendix that our formulation is equivalent. With just two decision-makers, average monotonicity reduces to the usual monotonicity condition.}

Second, even if some weights are negative, it only presents a problem when the patent applications with negative weights systematically differ in their treatment effects; one may be willing to assume this is not the case.\footnote{\Textcite{heckman2006understanding} call this assumption no “essential” treatment effect heterogeneity. In terms of the roy1951some selection model, it imposes no “selection-on-gains.”} More generally, dechaisemartin17 points out that if one finds a subset of positively weighted applications that matches the treatment effects of the firms that are defiers on average, these compliers cancel out the negative treatment weights on the defiers, and we can interpret leniency IV as estimating a convex-weighted average of treatment effects for the remaining compliers.

Third, even if the negatively weighted observations differ systematically, this need not bias leniency IV estimates substantively unless there are many firms with non-negligibly negative weights and the systematic differences in treatment effects are very large. We never see two patent examiners evaluate the same case, so we cannot directly assess the common ranking assumption. But we do actually have some evidence on this in the context of judge assignment, because some court cases are assigned to a panel of judges: sigstad_monotonicity_2025 shows that while monotonicity is technically often violated in judicial panels, the ranking disagreements are not severe enough to create substantial bias in leniency IV estimates.

Typically, leniency IV designs do not feature multiple decision-makers being assigned to the same case, precluding such direct monotonicity tests in other contexts. But we can still indirectly test the assumption. One implication of monotonicity is that the average outcomes for cases assigned to two decision-makers with the same leniency must be the same in the population, because under monotonicity, the two decision-makers' decisions must exactly match. More generally, if the leniencies are similar, the average outcomes must also be similar since the number of observations on which the decision-makers disagree must be small. \Textcite{frandsen2023judging} formalize this idea in a statistical test.\footnote{Strictly speaking, both this test and the average monotonicity test described below are of the joint null hypothesis of as-good-as-random assignment, exclusion, and monotonicity, as a violation of any of these assumptions could drive a test rejection. Often they are interpreted as tests of monotonicity maintaining the first two assumptions, however.} While useful in some settings, the test has two limitations. First, to show its validity, frandsen2023judging rule out a large number of decision-makers (i.e., the original motivation for using UJIVE), and its implementation requires bounded outcomes (ruling out, e.g., looking at effects of patent-granting on future startup profits). Further, it tests the original ImAn94 monotonicity condition, not the weaker average monotonicity condition that is necessary and sufficient for positive individual weights.

To address both limitations, we propose here a test for average monotonicity---based on an earlier proposal by kitagawa15---which leverages a modification of the UJIVE treatment effect estimator.\footnote{Alternative informal tests have been used by researchers in the literature, such as in dobbie2015debt,dobbie2017consumer, where the researcher plots the average leniency of each decision-maker for one subgroup against their average leniency for the other subgroup. A positive relationship is taken as evidence that decision-makers are not switching their behavior differentially by subgroup. However, the formal properties of these procedures have not been established.} We develop our proposed monotonicity test by first noting, following abadie02, that the UJIVE estimator can not only estimate average treatment effects for compliers: it can also characterize compliers by their baseline characteristics and potential outcomes. Formally, consider a UJIVE estimator with the same treatment, instruments, and controls as before, but with a modified outcome. Rather than $y_i$, we put on the left-hand side $\tilde{y}_i=v_i\times x_i$, where $v_i$ equals some variable determined before the as-good-as-random assignment of $z_i$. The local average treatment effect theorem applies to this modified outcome as well, showing that under average monotonicity, the UJIVE estimator identifies a convex weighted average of treatment effects of $x_i$ on $\tilde{y}_i$. But by construction these “effects” are just $v_i$, since $\tilde{y}_i$ is moved by exactly this amount when $x_i$ is counterfactually increased by one unit. Moreover, the weights that UJIVE puts on these effects are exactly the same as the ones in the original treatment effect specification. Hence, by simply replacing $y_i$ with $\tilde{y}_i$, we can easily compute weighted averages of $v_i$ with the UJIVE weights.

To see why this result is useful for testing monotonicity, note that it is possible to obtain a logically invalid estimate when computing such weighted averages. If $v_i$ is a binary variable, which can only take on values zero or one, then its weighted average must not be negative or greater than one. Such a finding would suggest something has gone wrong in the local average treatment effect theorem, namely monotonicity (assuming we are confident in as-good-as-random assignment and exclusion). More generally, we can set $v_i$ to be an indicator that some baseline variable---or set of variables---takes on a particular set of values and check that the resulting UJIVE estimate with that $\tilde{y}_i$ as the outcome lies between zero and one. For binary $x_i$, we can even use the original outcome $y_i$ (by setting $\tilde{y}_i=y_i\times x_i$) since then UJIVE will estimate a weighted average of treated outcomes abadie02; weighted averages of untreated outcomes can be estimated by setting $\tilde{y}_{i}=y_i\times (x_1-1)$. Of course, monotonicity tests like these are as straightforward to implement via UJIVE, and they inherit its approximate unbiasedness property even with many decision-makers or controls.\footnote{An important caveat to this suggestion is that the power properties of this test have not yet been explored. With a binary instrument, setting $\tilde{y}_i$ to indicators for all possible outcome values interacted with $x_i$ and $x_i-1$ yields a version of the kitagawa15 test. With multiple instruments, our test is different from the multivalued-instrument extension of the kitagawa15 test chmw24, since we test average monotonicity not the stronger conventional monotonicity.}

We again close this section with two additional comments. First, one might wish to use the above method to estimate average baseline characteristics of compliers even when monotonicity is not a concern. This helps evaluate another more subtle concern with IV estimation given heterogeneous treatment effects: the external validity, or representativeness, of the estimated complier-average treatment effect. If compliers are representative of the broader study population through their observable characteristics, again simply by replacing the $y_i$ outcome with some $\tilde{y}_i=v_i\times x_i$ for baseline $v_i$, this concern can be lessened.\footnote{Another complementary way of gauging external validity is to report how complier-average treatment effects, computed by simple IV regressions restricted to a given examiner pair, vary with the pairs' leniencies. Ordering the examiners by their leniency and reporting the estimates for pairs of neighboring examiners then gives an estimate of the marginal treatment effect curve HeVy05. Intuitively, a flat curve suggests that effects for non-responders are likely similar to the UJIVE estimate. However, the properties of such procedures---or their more sophisticated versions mst18---have not to our knowledge been formalized in the case of many decision-makers or controls.}

Second, a further subtle issue can arise when the specification of controls in $w_i$ is insufficiently flexible. To ensure we estimate a convex weighted average of treatment effects free of omitted variable bias with a single instrument that is only conditionally as-good-as-randomly assigned, the covariates $w_{i}$ need to enter flexibly enough so that the conditional mean of the instrument vector given $w_i$, $E[z_i\mid w_i]$, is linear in the covariates (kolesar13late shows this condition is sufficient, while blandhol2022tsls show it is necessary). When the instrument is a vector of dummies, a sufficient condition for nonnegative weights is that the mean of each element of $z_i$ is linear in $w_i$ and the remaining elements of $z_i$ ghk24. In our application below, this would require interacting the examiner assignment with art unit-by-year fixed effects, similar to the specification in angrist1991does. However, as with the potential monotonicity violations discussed above, more parsimonious control specifications (such as only using examiner dummies and art unit-by-year fixed effects without interacting them) need not generate meaningful bias, and they may in fact yield more precise estimates. In practice, researchers can explore this potential bias-variance tradeoff by checking robustness to alternative control specifications.

Inference

Leniency designs with multiple decision-makers raise subtle inference problems. To see why, we start with a benchmark: suppose we knew the relative leniency $\tilde{\ell}_i$, and hence could compute the infeasible estimator $\hat{\beta}^*$ in (ref). With iid data, the correct standard errors are given by the square root of $\sum_i\hat{\varepsilon}_i^2 \tilde{\ell}_i^2/(\sum_i \tilde{\ell}_i^2)^2$, where $\hat{\varepsilon}_i$ is the residual from projecting $y_i-x_i\hat{\beta}^*$ on the covariates. Since the covariances in the numerator and denominator of (ref) are approximately normally distributed by the central limit theorem, the delta method tells us their ratio is also approximately normal with this standard error.

Standard software packages compute heteroskedasticity-robust standard errors for 2SLS using the feasible version of this formula, plugging in $\hat{\ell}_{i}$ for $\tilde{\ell}_i$ and the 2SLS estimate for $\hat{\beta}^*$. These standard errors are correct when treatment effects are homogeneous and many-weak instrument bias is negligible. Under these conditions, 2SLS is as efficient as the infeasible estimator $\hat\beta^*$.

Unfortunately, with many instruments, the same issue that causes bias in the 2SLS estimator also causes its standard errors to be small relative to the infeasible estimator. Recall that overfitting inflates the denominator in (ref) from $\sum_i\tilde{\ell}_i^2$ to $\sum_i\tilde{\ell}_i^2+K\operatorname{var}(\nu_i)$ under homoskedasticity. Since the same denominator appears in the standard error formula, overfitting shrinks standard errors too. Thus, 2SLS mimics not only the ordinary least squares estimate, but also its precision. As bound1995problems emphasize, this can mask the weak instrument problem: researchers see tight standard errors around an estimate close to ordinary least squares and may wrongly conclude the causal effect is precisely estimated with minimal selection bias.

UJIVE fixes the denominator problem by construction, but the numerator of the default hetero\-skedas\-ticity-robust standard error formula still has two issues. First, with heterogeneous treatment effects, we cannot match the infeasible estimator's precision even when instruments are strong: the correct numerator picks up a second term reflecting variation in complier treatment effects across complier groups. Conveniently, ImAn94 already derived the correct numerator for the standard error in their appendix lee18late. Second, with many examiners, a third term appears due to estimation noise in $\hat{\ell}_{i}$ (this is the bekker94 many instrument term; we give the details in the appendix). Luckily, if we use UJIVE along with the plug-in heterogeneity-robust formula, it turns out that this also takes care of the many instrument term---we use this formula in our empirical illustration below.\footnote{The plug-in formula for the standard error's numerator has an own-observation bias that makes it overshoot its target, similar to the upward bias in the denominator of the 2SLS estimator. By a lucky coincidence, this overshooting more than accounts for the many instrument term making the formula valid (albeit slightly conservative). Correcting the slight upward bias involves a much more complicated leave-three-out procedure: see AnSo23 and yap25.}

There are still two potential issues with the UJIVE standard error formula remaining. The first concerns the reliability of the delta method, which lets us argue that if the covariances in the numerator and denominator of (ref) are each approximately normal, then so is their ratio. Here the issue is a classic one, dating back to AnRu49, and arises when the first stage is very weak as measured by $\sqrt{K}\times E[F]$. Suppose, for example, that there is essentially no variation in leniency across patent examiners. Then, even for the infeasible estimator $\hat{\beta}^*$, the delta method does not apply as the numerator and denominator of (ref) are essentially zero in expectation; ratios of mean-zero normals follow a Cauchy distribution, which has much thicker tails than a normal distribution.

A vast literature has tackled this weak instrument problem over the years, beginning with AnRu49 in the case with a fixed number of instruments and most recently, several papers proposing jackknife extensions for many instruments MiSu22,MaOt24,yap25. The fix turns out to be quite simple. To test that the causal effect equals a particular value $\beta_0$, we compute the residual $\hat{\varepsilon}_{i0}$ by projecting $y_i-x_i\beta_0$ onto the covariates. If the null is true, $\hat{\varepsilon}_{i0}$ should be uncorrelated with relative leniency, so we just need to check whether this correlation is significantly different from zero. This turns out to be equivalent to computing the UJIVE standard error just like before and checking whether UJIVE is significantly different from $\beta_0$, with one tweak: we use $\hat{\varepsilon}_{i0}$ instead of the UJIVE residual---this is the test proposed by yap25.\footnote{The test proposed by MaOt24 uses the same idea, but doesn't use the heterogeneity-robust formula for the standard error. As a result, it is not robust to treatment effect heterogeneity, and neither is the many-instrument robust version of the AnRu49 test considered in MiSu22.}

Recently, for the single-instrument case, AngKo24 give a more optimistic take on the weak instrument problem: they observe that, for the case with known leniency, the usual standard error only leads to overly optimistic inference if the correlation between $\sum_i \varepsilon_i\tilde{\ell}_i$ (the numerator of $\hat{\beta}^*-\beta$) and $\sum_i x_i\tilde{\ell}_i$ (the denominator) is very high; if not, the weak instruments blow up the standard error with $\hat{\varepsilon}_i$ even more than if we used the robust yap25 approach that uses $\hat{\varepsilon}_{i0}$ (which in this case coincides with the Anderson-Rubin test). This correlation parameter can be thought of as a measure of endogeneity, since it corresponds to the correlation between $\varepsilon_i$ and $x_i$ under homoskedasticity. But severe selection biases leading to very high endogeneity are unlikely in many applications; hence, AngKo24 argue, the delta method approximation may be good enough for the usual standard errors not to be misleading. We show in the appendix that an analogous argument applies to the UJIVE estimator with many instruments and controls: the only difference is the correlation parameter uses the UJIVE leave-out leniency measure rather than $\tilde{\ell}_i$; while the correlation parameter no longer directly links to endogeneity due to the more complicated variance structure, it can be bounded by placing bounds on the treatment effects. In most applications these weak instrument issues are moot, as the variation in leniency (as measured by $\sqrt{K}\times E[F]$) will be substantial. But if variation is small and high correlation cannot be ruled out, using $\hat{\varepsilon}_{i0}$ to compute standard errors with the above weak-instrument test may be prudent.\footnote{\Textcite{MiSu22} develop a weak-instrument pretest, designed to check validity of delta-method based standard errors regardless of the value of the correlation parameter, from a variant of the JIVE estimator. As they discuss, using the pre-test comes with a power sacrifice: if one switches between conventional and weak-instrument robust inference procedures using the pre-test, one needs to use larger critical values than the conventional 1.96 to account for possible type I and type II errors in the pretest.}

A final complication arises when the data is not iid but clustered. This can happen for two reasons. Either the sampling is clustered (say the patent office, instead of sharing the full data, only shares application data submitted on a randomly selected subset of dates, and the researcher would like to generalize to the full set of data), or the assignment is clustered (there is a lottery for each date, and the examiner who wins the lottery is assigned all applications filed on that date). Then, even if the relative leniency $\tilde{\ell}_i$ was known, a clustered version of the standard error formula is needed. Often, a researcher observes the full population of interest (in our application, we observe all patents in a specific time frame), so the main consideration is the assignment process.

\Textcite{aaiw20,aaiw23} show that if the examiner assignment process is iid, one does not need to worry about the correlation patterns of the residual $\varepsilon_i$ across $i$. What matters for the standard error is the correlation pattern of the product, $\tilde{\ell}_i\varepsilon_i$, which will be uncorrelated due to relative leniency being consequentially iid. This contrasts with traditional considerations for whether and how to cluster, which treat the patent examiner assignment (and hence relative leniency) as fixed, necessitating worries about correlation patterns in the residuals. These worries go away under the random instrument view, and one can use the same rule of thumb as in randomized experiments: cluster at the level of variation in the assignment. Furthermore, whether clustering matters for the magnitude of the standard error does not help determine whether it is needed, so clustering the standard errors “just in case” may yield overly conservative inference. This reasoning also applies when leniency is unknown, but the number of examiners is fixed. While the aaiw20,aaiw23 arguments have not been formally extended to the many examiner case, it seems natural to let the sampling and assignment process guide the choice of whether and how to cluster.

A Checklist for Leniency Designs

We now present a step-by-step guide to leniency design implementation, illustrated with data from farre2020patent. Their setting examines how initial patent approval by a US Patent Office examiner affects subsequent innovation by US startups. Our reanalysis sample contains 32,514 first-time patent applications filed after 2001 with final decisions by 2013.\footnote{Our analysis uses the full sample from the original paper as found in its replication package. The initial dataset has 34,435 observations. We drop 1 observation that is missing citation data. We then drop 1,851 observations with singleton covariates or instruments, and 69 observations with leverage of 1. This causes us to drop 378 collinear controls and 1,676 collinear IVs. See the code documentation at \href{https://github.com/kolesarm/ManyIV}{https://github.com/kolesarm/ManyIV} for discussion on how this recursive algorithm is implemented to drop these observations and variables.} (ref) summarizes key variables. The treatment---approval of the first-time application---occurs for 65% of applicants. Following farre2020patent, we track several key outcomes: whether startups file additional patents (40% do), whether they have any subsequent patents that were approved (30% do), and counts of new applications, approvals, and citations. Besides standard covariates like patent class and claim counts, we observe venture capital funding rounds prior to application. For the subset of applications also submitted abroad, we observe a proxy of application quality: European or Japanese patent office decisions. We also construct a measure of local entrepreneurship using the number of startups in each state. As we show below, these additional observables can help gauge the plausibility of the leniency design's identifying assumptions and assess external validity.

Our checklist consists of five steps; to illustrate it, we use an original R package for UJIVE, available at \href{https://github.com/kolesarm/ManyIV}{https://github.com/kolesarm/ManyIV}.

\noindent1. Identify necessary controls for as-good-as-random assignment, and what estimator and standard errors are appropriate given the quasi-experimental design

A credible leniency design begins with institutional knowledge that justifies quasi-random decision-maker assignment. This justification naturally identifies any required controls. In the US patent office, for example, once patent applications are allocated to art units the assignment process to individual examiners is not standardized. However, LeSa12 argue, based on interviews with examiners, that the assignment process is effectively random conditional on the set of examiners working in the art unit at the patent office at the time of the assignment.\footnote{Many art units use the last digit of the application serial number for assignment. Others similarly use rules that imply an effectively random assignment, such as a “first-in-first-out” rule that assigns each incoming application to the first available examiner. See also the institutional discussions in SaWi19 and farre2020patent.} While the potential examiner set is not directly observable, we can assume the set of examiners working in a given art unit changes only slowly across time and specify the necessary controls as art unit-by-year fixed effects. Without these necessary controls, the analysis would conflate systematic differences across art units or changes over time with examiner-specific variation. From our discussion on heterogeneous effects, recall that it can be important to have sufficiently flexible controls to ensure a local average treatment effect interpretation of leniency design estimates.

Quasi-random assignment is only one piece of instrument validity in an examiner design; the second is the exclusion restriction, requiring decision-maker assignment to only affect relevant outcomes through the specified treatment. Here too, a clear argument should be made from institutional knowledge. In the patent setting example, exclusion is plausible since patent examiners make relatively narrow approval decisions and likely have no direct impact on the future innovativeness of the applying firm.

The assignment mechanism can also guide the appropriate level for clustering standard errors. Individual-level randomization in the patent setting, where patents are routed idiosyncratically to examiners within art unit-year, makes heteroskedasticity-robust standard errors appropriate. Based on our rule of thumb, clustering is not needed because each case is assigned independently. Contrast this with settings where groups of observations are assigned together: if all applications from a given month were assigned to a randomly chosen examiner, we would cluster by month. As discussed above, a justification for clustering standard errors may also justify modifying the UJIVE estimator (i.e.\ to a leave-cluster-out version).

\noindent2. Verify balance on other observables

The as-good-as-random assumption implies that the average predetermined characteristics should be equal across decision-makers, conditional on the necessary controls. We test this implication directly: for each observable characteristic, we run our UJIVE specification with that characteristic as the outcome. Significant coefficients indicate imbalance and potential threats to identification. This unified approach---using the same UJIVE specification for both balance tests and treatment effect estimation---offers important advantages over alternatives. Consider the common practice of regressing observables on constructed leniency measures from a first stage. With many examiners and fixed effects, this can generate mechanical correlations that masquerade as true imbalance.\footnote{Even if one uses a leave-out leniency measure, the coefficient will still be biased due to estimation error in the leniency measure, which generates an errors-in-variables bias. Another approach to testing balance is to regress the predetermined characteristics on the full set of examiner indicators and controls, and report the joint F-test for the examiner indicators. With many examiners, however, the default F-statistic is not valid AnSo23.} UJIVE avoids these finite-sample biases while maintaining the interpretability of our test: the magnitude of any imbalance directly translates to potential bias in our treatment effects. If venture capital funding has a positive coefficient in this UJIVE balance check, for example, this would imply that the randomly assigned examiners' tendency to approve patents is also positively correlated with patents' ex ante venture capital funding. Such a finding would raise questions about whether the examiners are truly randomly assigned, and the magnitude of the estimated coefficient can be used in quantifying sensitivity to omitted variables.

(ref) presents balance tests for the farre2020patent reanalysis. Each row reports a UJIVE coefficient from regressing a predetermined covariate on patent approval, instrumenting with examiner indicators and controlling for art unit-by-year fixed effects. The results support quasi-random assignment: seven of eight coefficients are statistically insignificant at the 5% level, and all are insignificant at the 1% level. As importantly, the magnitudes are economically negligible: coefficients are typically 10 times smaller than the estimated treatment effects discussed below. Consider the venture capital variable, which might raise particular concern if certain examiners systematically reviewed applications from better-funded startups. We estimate a close-to-zero effect, suggesting no systematic differences (coefficient = -0.024, standard error = 0.035). Similar null results emerge for technical complexity (independent claims), external quality validation (European patent approval), and local entrepreneurial environment (state startup density). The absence of systematic imbalance strengthens our confidence that examiner assignment generates plausibly exogenous variation in patent approval.

Balance tests can also examine post-assignment variables to detect exclusion restriction violations. These tests ask whether decision-makers affect outcomes through channels beyond the specified treatment. In the patent setting, for example, one concern is that examiner assignment affects not only whether an application is approved but also the speed of the application review---which plausibly has independent effects on follow-up innovation outcomes by reducing applicant uncertainty. A UJIVE regression which uses a measure of review speed as the outcome, keeping again the treatment, instrument, and controls, could be used to assess the significance of such potential exclusion restriction violations.\footnote{farre2020patent conduct a different test, including review speed as a second treatment while also including a second “leniency” instrument capturing the assigned examiner's average review speed for past cases (see also hlr22). Such multiple-treatment specifications can be difficult to interpret when the effects of patent-granting and review speed are heterogeneous BhSi24.}

\noindent3. Estimate treatment effects by UJIVE and alternative estimators

The UJIVE estimator is the preferred choice for leniency designs as it avoids the subtle potential biases of alternative approaches when there are many decision-makers and controls. Comparing UJIVE estimates and standard errors to those from these alternative approaches---such as 2SLS---is illustrative of these biases. As usual with instrumental variable designs, it can also be instructive to compare the UJIVE estimates to ordinary least squares estimates of the treatment effect as a way to probe the importance of selection bias.

(ref) presents treatment effect estimates across four specifications. Column 1 reports our preferred UJIVE estimates using examiner indicators as instruments with art unit-by-year controls. Columns 2 and 3 implement 2SLS with alternative instruments: the farre2020patent constructed leniency measure, based on examiner approval rates from earlier periods in Column 2, and the full set of examiner indicators in Column 3. Column 4 shows ordinary least squares results.

The UJIVE estimates reveal positive and significant effects of patent approval on subsequent innovation. Approved startups are more likely to file future patents, receive future approvals, and generate citations, with significant effects on both extensive and intensive margins. The pattern of bias across estimators is instructive. Ordinary least squares produces higher coefficients for subsequent applications but lower coefficients for approvals and citations, suggesting moderate selection that operates differently across outcomes. The 2SLS estimates using examiner indicators (Column 3) fall between ordinary least squares and UJIVE, pulled toward the biased ordinary least squares results. This intermediate position confirms our theoretical prediction: with many instruments, 2SLS suffers from finite-sample bias that partially reproduces the selection problem it aims to solve. The UJIVE estimates avoid this bias, providing our most credible evidence that initial patent approval causally stimulates follow-on innovation. These effects operate through multiple channels—approved firms not only attempt more patents but also succeed in getting them approved and generating scientific impact through citations.\footnote{In unreported results, we examine how the estimates change when we additionally control for pre-application covariates. Adding such precision controls can reduce standard errors, as (ref) discusses. This is not generally the case here, however, since adding the controls also reduces the effective estimation sample given the algorithm the UJIVE estimator uses to eliminate collinearities (see (ref)).} The 2SLS estimates using the farre2020patent measure of leniency tend to be larger than UJIVE\@. Since their leniency construction is similar to that of JIVE, this may reflect a many-covariate bias.

Standard errors also tell an important story in (ref). The many-instrument 2SLS (Column 3) produces standard errors roughly 3--4 times smaller than UJIVE (Column 1)—a difference that reflects statistical pathology rather than efficiency gains. As we have discussed, 2SLS standard errors and estimates are both pulled towards ordinary least squares with many examiners, creating a false sense of precision around the biased estimates. The shrinkage in standard errors occurs even when using the leave-out farre2020patent leniency measure in Column 2, since these do not reflect estimation error in the leniency measure. UJIVE avoids both failures: it delivers unbiased point estimates and standard errors that accurately capture sampling variation.

\noindent4. Test monotonicity to assess the plausibility of a LATE interpretation

To ensure that the UJIVE estimates have a clear causal interpretation under heterogeneous effects, we also need a version of first-stage monotonicity. Since the ImAn94 first-stage monotonicity condition is likely too strong in the examiner setting, we implement a UJIVE test of the weaker average monotonicity assumption: that the average leniency of the examiners who would grant the firm a patent exceeds the average leniency of those who would deny it.

(ref) tests average monotonicity using outcome-specific UJIVE regressions. For each possible outcome value, we create an indicator for that value (e.g., a dummy for zero subsequent applications) and interact it with treatment status as our dependent variable, maintaining examiner instruments and art unit-by-year controls. The specification is estimated as in (ref). The fact that all of these UJIVE estimates are positive, with tight 95% confidence intervals, builds support for average monotonicity holding.

\noindent5. Characterize compliers to assess external validity

Having established internal validity, we examine whether our complier population represents the broader sample. We estimate complier characteristics using UJIVE with modified outcomes: we regress the interaction of each covariate with the treatment indicator on the treatment itself, maintaining examiner instruments and controls. This recovers weighted averages of complier characteristics for treated compliers using the same weights as our main estimates. Using one minus the treatment as the treatment variable recovers the characteristics for untreated compliers. We efficiently pool the two specifications, following ANGRIST20231.\footnote{For a characteristic $v_i$, one can show this can be done by using $\tilde{x}_i=2x_i-1$ as the transformed treatment variable that takes on values $\{-1,1\}$ instead of $\{0,1\}$. We then run a UJIVE regression of $v_i\times \tilde{x}_i$ onto $\tilde{x}_i$, maintaining the controls and examiner instruments.}

(ref) reveals that compliers closely resemble the full sample. Column 2 reports complier means for eight characteristics, while Column 1 shows population means. None of the differences are statistically significant, while all the complier estimates fall within logical bounds---again supporting average monotonicity. Compliers have similar rates of venture capital funding, comparable technical complexity in their applications, and equivalent representation across patent classes. This representativeness matters for interpretation: our treatment effects likely approximate average effects for the full population of first-time patent applicants, not just an idiosyncratic subset whose outcomes depend on examiner assignment. The validity of our estimates thus appears to extend beyond the specific complier group to inform broader questions about how patent rights affect startup innovation.

Conclusion

Leniency designs allow for estimation of causal effects when expert decision-makers are as-good-as-randomly assigned and exert discretion on a high-stakes treatment. But, as with many things in both life and econometrics, the devil is in the details. Our review emphasizes how subtle choices in estimation and inference can have first-order consequences in practice. The UJIVE estimator emerges as a natural solution to the many-weak instrument and many control problems that plague two-stage least squares estimation, while correctly estimated heteroskedastic-robust standard errors are a natural choice when assignment operates at the individual level. Our practical checklist and reanalysis of farre2020patent shows how UJIVE can also be used for a number of important auxiliary analyses, such as checking balance, testing average monotonicity, and characterizing compliers. We hope this toolkit gives applied researchers more confidence when approaching leniency designs, and we encourage them to consult this manual whenever questions of operation arise.

\singlespacing \printbibliography

figure[figure omitted — 1,146 chars of source]
table[table omitted — 3,652 chars of source]
table[table omitted — 1,069 chars of source]
table[table omitted — 1,786 chars of source]
table[table omitted — 1,209 chars of source]