Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,250 characters · 21 sections · 102 citation commands
Negative Control Falsification Tests for Instrumental Variable Designs
The identification assumptions in instrumental variable (IV) designs cannot be directly tested. Instead, researchers often use indirect falsification, or “placebo,” tests. Reviewing the most-cited papers in five leading economics journals, we find that 51% of IV studies employ such falsification tests. The large majority of falsification tests fall into two categories: 75% of papers that implemented a falsification test examined that the IV is not associated with certain variables such as lagged outcomes, which we call negative control outcomes, (NCOs). Similarly, 24% examined that the outcome is not associated with other variables, which we call negative control instruments (NCIs). For example, they tested that the outcome is not correlated with variables that resemble the IV but do not affect the treatment. \phantomsectionExtensive literature has developed a theoretical framework for negative control falsification tests in classic causal settings lipsitch2010negative,shi2020selective, but not for the assumptions underlying IV designs.\footnote{In epidemiology and biomedical fields, researchers use the similar terminology of negative controls for falsification tests to detect potential confounding in an exposure-outcome relationship.} This paper aims to fill this gap.
We propose that practitioners first identify potential threats to their IV validity, which we formally define. Since variables posing these threats are typically unobserved, researchers should look for proxy variables, termed negative controls. We characterize the conditions that these proxies must satisfy and show how to use them to test the validity of the IV design using (conditional) independence tests. The theory highlights two common pitfalls, which can falsely flag problems in valid IV designs. First, NCI tests typically require conditioning on the IV---a step frequently overlooked in practice. Second, prevalent negative control tests may flag violations of unnecessary or replaceable functional form assumptions. We propose ways to separately test the validity of the IV design from these functional form assumptions. Our framework also suggests novel, underutilized negative control variables and testing methods.
\phantomsection We first introduce the concept of an alternative path variable, which represents a variable that poses a threat to the identification. In valid IV designs, which satisfy independence and the exclusion restriction, the only path between the IV and the outcome is through the treatment.\footnote{The independence assumption typically includes independence of the IV with the potential treatment and the potential outcome abadie2003semiparametric. The falsification tests we discuss in this paper focus only on outcome independence. See Section (ref).} Threats to identification can be characterized as alternative paths between the IV and the outcome through an alternative path variable, rather than through the treatment.
To construct a negative control test, researchers need to determine which type of alternative path threatens the IV design. We distinguish between two categories of variables that can create such paths. In the first category, alternative path outcome (APO) variables, the concern is that a variable that is associated with the outcome would also be associated with the IV. (ref) illustrates two examples of such cases.\footnote{Throughout the paper, we use directed acyclic graphs (DAGs) to visualize complex structures, as advocated by imbens2020potential. (ref) outlines the theory presented in this paper within the formal causal DAG framework pearl2009causality} Panel A shows a potential violation of the independence assumption. For concreteness, consider the context of martin2017bias, who examine the impact of Fox News viewership ($X$) on Republican vote shares ($Y$). As an IV, they use the local Fox News cable channel position ($Z$) since lower channel numbers induce higher viewership. One concern is that unobserved local conservativeness ($U$, the APO variable) affects not only the republication vote share (the outcome) but also the channel position (the IV), as marked with the dashed arrow. If that is the case, an alternative path between the channel position and Republican voting share exists via local conservativeness, the APO variable. This would violate the independence assumption and invalidate the design. Panel B offers an example of an APO variable ($U_2$) that is part of a potential violation of the exclusion restriction assumption.
In the second category, alternative path instrument (API) variables, the concern is that variables known to be associated with the IV would also be associated with the outcome. Panel C of (ref) provides an example of an API variable that potentially violates the independence assumption. For concreteness, consider the context of nunn2014us, who examine the effect of US food aid ($X$) on conflicts in recipient countries ($Y$). They use the US production of wheat ($Z$), a staple aid crop, as an IV for aid. Here, the API variable is unobserved weather conditions ($U_3$). It is known that weather conditions affect wheat production, the IV. The question is whether they also affect conflicts, the outcome, as marked with the dashed arrow. If so, an alternative path from wheat production to conflict exists via the API variable. Panel D offers an example of an API variable ($U_4$) that is part of a potential violation of the exclusion restriction assumption.
\phantomsection Since alternative path variables of both categories are often unobserved, researchers can utilize negative control variables as proxies for them. A negative control outcome (NCO) is a proxy for an APO variable. An association between an NCO and the IV implies the presence of an alternative path, indicating that the design is not valid. In martin2017bias, the lagged outcome, Republican vote share in 1996 ($NC_1$ in (ref)) is used as an NCO for unobserved conservativeness ($U_1$). If 1996 vote shares correlate with channel position, it would imply that cable companies consider the population conservativeness when placing the channels, such that the dashed arrow exists. Hence, there is an alternative path between the IV and the outcome, violating the independence assumption. Panel A of (ref) describes further examples of applications of NCO variables in economic research.
Similarly, a negative control instrument (NCI) can be used as a proxy for an API variable. For example, nunn2014us use orange production as an NCI ($NC_3$ in (ref)) for unobserved weather conditions ($U_3$)---orange production is affected by similar weather conditions as wheat but is not used for food aid. If orange production were associated with conflicts conditional on wheat production, this would indicate an alternative path between the IV and the outcome, violating the independence assumption. Panel B of (ref) lists additional examples of NCI variables in economic research.
The definition of negative controls clarifies which proxy variables researchers can use as NCOs or NCIs. This definition guarantees that these proxies can test for alternative paths using (conditional) independence tests. In particular, variables directly associated with the IV (not through the APO variable) cannot serve as NCOs. Similarly, variables associated directly with the outcome (not through the API variable or the IV) cannot serve as NCIs. Tests using such variables would falsely flag valid IV designs.
The theory highlights two common pitfalls in current practice. First, in our survey, the vast majority of papers using NCI variables implemented a test that will falsely flag problems in valid IV designs (with sufficient sample size). In almost all these papers, the NCI was a variable that is similar to the IV but does not affect the treatment (such as orange instead of wheat production in the previous example). Researchers typically test whether such an NCI is correlated with the outcome by plugging it instead of the original IV in the reduced form equation. This specification overlooks a key issue: NCIs are typically correlated with the original IV and therefore will be associated with the outcome due to this correlation, even in a valid IV design. For example, as shown in Panel C of (ref), both orange production ($NC_3$) and wheat production ($Z$) are influenced by weather conditions ($U_3$). This means that orange production and conflict ($Y$) would be correlated even if wheat production is a valid IV (i.e., the dashed arrow does not exist and there is no alternative path). We show that this problem can be avoided by controlling for the original IV in the NCI test. In the above example, this means controlling for wheat production when testing the association between orange production and conflict.
The second pitfall is that negative control tests may flag problems even in valid IV designs due to a misspecified functional form. In 2SLS specifications, researchers must choose a functional form for the IV and the control variables, typically assuming a simple linear-additive model. This structure often carries over into the execution of negative control tests. Consequently, even valid IV designs may fail negative control tests merely due to violations of functional form assumptions. However, unlike the independence and exclusion assumptions, which are necessary for IV validity, functional form assumptions can often be relaxed and, in some cases, are unnecessary for identifying causal effects. To address this problem, researchers can use alternative negative control tests that rely on weaker functional form assumptions.
The theory can also be used to identify new types of negative control variables, some of which have yet to be commonly employed in empirical research. For example, variables that causally influence the IV could serve as NCIs (e.g., observed weather conditions can serve as NCIs when the IV is wheat production).
\phantomsection This paper adds to prior econometrics work on tests for IV design validity and, more generally, the validity of causal designs. Recent work has suggested novel tests to examine the validity of IV designs kitagawa2015test, huber2015testing, mourifie2017testing, frandsen2023judging, chyn2024examiner. Previous work has also discussed robustness tests, not specific to IV, based on varying the set of controls altonji2005selection, oster2019unobservable, diegert2022assessing. However, pei2019poorly recommends using such control variables as NCOs instead. eggers2021placebo discuss the usage of placebo tests in the social sciences more broadly. We contribute to this literature by outlining a theoretical framework for the most common type of falsification tests for IV designs.
\phantomsection This paper also contributes to the growing literature on negative controls lipsitch2010negative,shi2020selective. In a standard causal design, negative controls are used to detect or even correct for bias due to unobserved treatment-outcome confounding. To this end, valid and, in some cases, even invalid IVs can serve as negative control exposures, and under further assumptions can be combined with an additional negative control to achieve point identification miao2018identifying, shi2020selective, tchetgen2024introduction, dukes2024using. We apply the theory of negative controls in a different setting, where researchers employ an IV design and seek to use negative controls specifically for testing the IV assumptions in order to assess the design's validity. Our work is related to davies2017compare, who use negative controls in IV contexts without developing a theoretical framework for such an approach. We find several important differences in the theoretical framework of negative controls for assessing IVs compared to their usage in the standard treatment-outcome setting.
The rest of this paper proceeds as follows. Section (ref) surveys the current practice of falsification tests for IV designs. Section (ref) presents the theory for negative control tests in IV designs. Section (ref) provides guidance for practitioners and demonstrates key findings using recent empirical studies. Section (ref) concludes.
To provide an overview of current practices in falsification testing for IV designs, we surveyed the most highly cited articles with an IV analysis published between 2013 and 2023 in top economics journals. We then classified the characteristics of the falsification tests used. (ref) provides additional details on the survey construction, and the results are summarized in (ref).
We highlight five key findings from this survey. First, falsification tests are widely used in IV analyses. Approximately half (51%) of all articles surveyed employ some form of falsification test (Column 2 of (ref)).
Second, most falsification tests fall within the negative control framework described in the introduction and formalized in this paper. As outlined earlier, these tests can be divided into two types: negative control outcome (NCO) tests, which check for associations between the IV and variables it should not be associated with, and negative control instrument (NCI) tests, which examine associations between the outcome and variables it should not be associated with. Among surveyed papers using falsification tests, 75% used NCO tests (Column 3) and 24% used NCI tests (Column 4). All other types of falsification tests combined were used in 21% of the papers (Column 5); (ref) lists these other, less common types.
Third, current applied work usually restricts itself to two simple types of negative control test specifications that rely on the 2SLS functional form assumptions. NCO tests typically involve estimating a revised reduced form equation using an alternative outcome (e.g., a lagged outcome) and testing if it is unrelated to the IV (Column 6). Such specifications account for 57% of all NCO tests. The remaining NCO tests often follow a similar logic.\footnote{For example, balance tables that regress various NCOs on the IV.} For NCI tests, researchers either replace the original IV in the reduced form equation with a similar variable that does not affect the treatment (Column 7), or add this variable to the reduced form equation (Column 8).
Fourth, most reported NCI tests are implemented incorrectly. As noted in the introduction and discussed in detail later, NCI tests should almost always control for the original IV. In practice, only 24% of the papers surveyed reported doing so. With sufficient sample size, this error leads to finding false problems in valid IV designs. It is likely that additional NCI tests were conducted incorrectly, finding false problems in valid designs, and were therefore not reported.
Finally, papers using falsification tests usually utilize only a few negative control variables. The median number of negative control variables used in the surveyed papers is 3.5 (Column 9), with 35% of of the papers using only one. This finding suggests that researchers use only a subset of the available relevant negative controls. As we demonstrate in Section (ref), the theory can guide a systematic search for negative control variables in existing data and suggest novel types of negative control variables researchers can use to evaluate their IV designs.
In this section, we present the theory of negative control tests for IV designs. The theory is constructed using the terminology of potential outcomes, and we use DAGs for examples and intuition. (ref) introduces the basic relevant concepts for DAG theory and replicates the key definitions and theorems using DAGs.
\phantomsection
Consider i.i.d. units indexed by $i=1,\ldots,n$. Denote the observed (endogenous) treatment status by $X_i$, and the candidate IV by $Z_i$. Let $Y_i(z,x)$ be the potential outcome for unit $i$ had $Z_i$ and $X_i$ been jointly set to the values $z$ and $x$, respectively.\footnote{This formulation implicitly assumes the stable unit treatment value assumption (SUTVA).} We make the standard assumption that the observed outcome $Y_i$ is given by $Y_i=Y_i(Z_i,X_i)$. Because units are assumed to be i.i.d., we omit the subscript $i$ when it improves clarity. All variables may be discrete or continuous.
The negative control tests we discuss in this paper examine whether there is an alternative path between the IV and the outcome, in addition to the standard path through the treatment. Such an additional path would violate one of the following two assumptions. The first assumption, outcome independence, maintains that IV assignment is independent of the potential outcomes.
This assumption is usually written as part of a more general independence assumption abadie2003semiparametric. Here, we distinguish between outcome independence and treatment independence, which requires $Z \protect\mathpalette{\protect\independenT}{\perp} X(z)$ for every value of $z$. Only outcome independence is tested in the negative control tests we discuss in this paper.
Outcome independence is violated if the IV is affected by a variable that also affects the outcome. As previously discussed, martin2017bias study the effect of Fox News on voting using channel positions as an IV. This example is illustrated in Panel A of (ref). The concern is that cable companies accounted for local conservativeness ($U_1$) when assigning Fox News channel position ($Z$), as illustrated by the dashed arrow. If so, an alternative path emerges between channel position and voting ($Y$) through local conservativeness. This path violates outcome independence.
The second assumption, exclusion restriction, maintains that the IV does not have a direct effect on the outcome. \phantomsectionLet $Y(x)$ be the potential outcome had the treatment $X$ been set to $x$, while $Z$ had not been set to any particular value and takes its natural value, i.e., $Y(x)=Y(Z,x)$.
The exclusion restriction is violated if the IV affects the outcome in alternative ways, in addition to its effect through the treatment.
\phantomsection Panel B of (ref) illustrates a potential violation of the exclusion restriction assumption. For example, angrist1996children study the effect of the number of children ($X$) on female labor supply ($Y$), using the sex composition of the first two children as an IV ($Z$) (because same-sex children induce further births for parents who have a preference for gender variety). One potential concern is that same-sex sibship could reduce housing expenditures ($U_2$) due to hand-me-downs, which could then affect female labor supply decisions. In this case, an alternative path between the sex composition of the first two children and female labor supply would exist through the effect on household expenditures. This violates the exclusion restriction.
Together, outcome independence and exclusion restriction imply\footnote{\phantomsectionWhen the exclusion restriction does not hold, $Y(x)$ is still properly defined but does not equal $Y(z,x)$ for every value of $z$. Therefore, recalling that $Y(x)=Y(Z,x)$, violation of the exclusion restriction implies $Z\cancel{\protect\mathpalette{\protect\independenT}{\perp}}Y(x)$.}
In loose terms, (ref) requires that there are no alternative paths between the IV and the outcome except through the treatment. Because potential outcomes are never observed, neither these two assumptions nor (ref) can be tested directly.
To identify a causal effect using an IV design, additional assumptions are also necessary. For example, the design of angrist1996identification also requires treatment independence, relevance, and monotonicity. However, the negative control tests presented in this paper do not test these other assumptions.
To formalize the notion of a threat to the identification, we introduce the concept of an alternative path variable---a variable that is part of a suspected alternative path between the IV and the outcome that, if such a path exists, would violate outcome independence or exclusion. For simplicity, we assume that only one potential threat to the IV validity exists. (ref) addresses a more general case with multiple threats. We distinguish between two types of alternative path variables that require different types of falsification tests.
\phantomsectionThe first type of identification threat involves alternative path outcome (APO) variables. These variables are presumably associated with the outcome. The threat is that they are also associated with the IV, which would generate an alternative path. In martin2017bias, conservativeness of the local population is an APO variable ($U_1$ in Panel A of (ref)). Since conservativeness certainly affects voting behavior, it poses a threat if it also affects cable companies' decisions for channel position (represented by the dashed arrow). Panel B of (ref) describes an APO variable for a potential violation of the exclusion restriction.
The formal definition of an APO variable is as follows.
Latent IV validity posits that had we observed and conditioned on the APO variable, the IV design would have been valid---both outcome independence and the exclusion restriction would hold conditional on the APO variable. This condition implies that imperfect proxies for variables posing identification threats cannot themselves be APO variables, as controlling for an imperfect proxy does not make the IV and the potential outcome conditionally independent. Using again the example of martin2017bias, the share of Republican votes in 1996 is only an imperfect proxy for the APO variable (latent conservativeness), and hence controlling for it does not eliminate the threat. Therefore, the Republican vote share in 1996 is not an APO variable as it does not satisfy the latent IV validity condition. \phantomsectionLatent IV validity is analogous to the latent exchangeability assumption appearing in recent literature on negative controls in epidemiology shi2020selective and statistics tchetgen2024introduction.
Path indication states that a valid IV is not associated with the APO variable. Its contrapositive ensures that an association between the IV and the APO variable implies an alternative path between the IV and the potential outcome. Path indication guarantees that if there is a path from the IV to the APO variable, the path continues from the APO variable to the outcome. Therefore, it excludes variables unrelated to the outcome, as they can be associated with the IV without implying anything about the design validity. While APO variables often causally affect the outcome, it is not mandatory (as demonstrated in (ref)). Path indication also rules out variables that could be related to both the IV and the outcome without generating a correlation between them. For example, this would occur if a variable is correlated with the outcome for some subpopulation but potentially correlated with the IV only for a separate subpopulation. Two examples are provided in (ref) and (ref).
The second type of alternative path variables is alternative path instrument (API) variables. API variables are known to be associated with the IV, and the concern is an association they might have with the outcome. This is in contrast to APO variables, which are known to be associated with the outcome, and the concern is their possible association with the IV. U sing the previous example of nunn2014us, illustrated in Panel C of (ref), weather conditions ($U_3$) is an API variable. Wheat production ($Z$) is known to be affected by weather. An alternative path that threatens identification may form if weather also affects conflicts ($Y$) directly (the dashed arrow exists). Panel D describes an API variable for a potential violation of the exclusion restriction.
Formally, an API variable satisfies the following definition.
This definition resembles the definition of APO variables (Definition (ref)). The first condition, latent IV validity, is exactly as before. The difference between API and APO variables is encapsulated in the second condition, path indication. For API variables, this condition requires that if $Z \protect\mathpalette{\protect\independenT}{\perp} Y(x)$, then the API variable must be independent of the observed outcome conditional on the IV (i.e., $U \protect\mathpalette{\protect\independenT}{\perp} Y|Z$). Through its contrapositive, this condition implies that an association between the API variable and the outcome, not via the IV, indicates that there exists an alternative path between the IV and the outcome. Therefore, the IV is invalid. Typically, path indication is satisfied when the API variable is associated with the IV.
For API variables, path indication rules out variables that are associated with the outcome through the treatment (conditional on the IV). Such variables are not informative about the validity of the IV design as they are associated with the outcome through the treatment, even if the IV design is valid. This is different from APO variables that could be associated with the treatment (even conditional on the IV). See (ref) for an example and a further discussion of this issue. This implies that APO variables can be associated with (or even identical to) the original confounder of the treatment-outcome relationship (which we mark by $W$ in our examples). However, an API variable cannot.
\phantomsection Negative control variables are observed proxies for the unobserved alternative path variables. Building on the definitions of alternative path variables, we are now ready to formalize the assumptions required for a random variable to serve as a negative control. The first type of negative control, negative control outcome, is a proxy for an APO variable. Observed variables can serve as NCOs if they satisfy the following definition.
\phantomsection
\phantomsection The NCO assumption guarantees that any path between the IV and the NCO must go through the APO variable $U$. It rules out variables that have other paths to the IV. \phantomsection Panel A of (ref) demonstrates the NCO assumption in a setting with a potential violation of outcome independence. For example, consider again the effect of Fox News on voting martin2017bias. The lagged outcome---Republican vote share in 1996---is an NCO ($NC_1$). The NCO assumption requires that any association between lagged voting and later channel position assignment ($Z$) arises only due to local conservativeness ($U_1$, the APO). Panel B of (ref) depicts an example of an NCO ($NC_2$) that satisfies this assumption where a violation of the exclusion restriction is the concern.
The NCO assumption is violated for variables that are directly related to the IV, not through an APO variable. This can occur if the IV affects the candidate for NCO, either directly or through the treatment or the outcome. For example, various IV studies on the impacts of exposure to air pollution on different outcomes use non-respiratory hospital admissions, a seemingly unrelated outcome, as NCOs. One might expect that these admissions would only correlate with flawed IVs for air pollution. However, guidetti2021placebo demonstrate otherwise. They find that air pollution increases non-respiratory admissions through hospital congestion caused by a surge in respiratory admissions. Therefore, non-respiratory admissions are not informative about the IV validity as they correlate with both flawed and valid IVs. Formally, non-respiratory admissions correlate with the IV, not through any APO variable but due to the unrelated mechanism of congestion. Hence, non-respiratory admissions violate the NCO assumption.
\phantomsection The $U$-comparability assumption guarantees that the NCO has a path to the APO variable. This assumption guarantees that the NCO is a relevant proxy for the APO variable. For example, voting in 1996 satisfied $U$-comparability as it is correlated with the APO, unobserved conservativeness in the region. This assumption rules out variables that are uninformative about the design validity because they are unrelated to the identification threat (and specifically to the APO variable).
\phantomsection The second type of negative control variable, negative control instrument, is a proxy for API variables. Observed variables can serve as NCIs if they satisfy the following definition.
\phantomsection The NCI assumption guarantees that any path between the outcome and the NCI goes through the API variable $U$ or the IV $Z$. It rules out variables that have other paths to the outcome.
While similar, the NCI assumption and the NCO assumption (Definition (ref)) differ in three key aspects. First, the alternative path variable $U$ is an API variable instead of an APO variable. Second, the conditional independence is between the NCI and the outcome instead of the IV. Due to these two differences, the NCI tests defined below test for a potential association with the outcome and not with the IV. The third difference is that the independence requirement is also conditional on the IV. This is because in valid IV designs, the NCI is often associated with the outcome through the IV, as we discuss in the next section.
Panel C of (ref) demonstrates the NCI assumption in a setting with a potential violation of outcome independence. For example, consider again the context of nunn2014us, which uses an alternative crop (e.g., oranges) production as an NCI ($NC_3$). Orange production is affected by similar weather conditions ($U_3$) as wheat production ($Z$) and, therefore, would be correlated with it. However, unlike wheat, oranges are not used as food aid ($X$). Therefore, orange production is unrelated to conflicts ($Y$), conditional on both weather and wheat production.
In many applications, the NCI assumption rules out a large class of observed variables because of their association with the outcome. For example, demographic variables often exhibit an association with the outcome, even conditional on the IV and the API variable, and therefore cannot serve as NCIs. Variables that are associated with the treatment are also not NCIs as they are also associated with the outcome conditional on the IV and the API variable. Moreover, in cases where the IV effect on the outcome is heterogeneous, any variable associated with the source of heterogeneity cannot serve as an NCI. In practice, the NCI assumption is more restrictive than the NCO assumption. The reason is that the NCO assumption requires conditional independence with the IV, which is typically more plausible than conditional independence with the outcome.
\phantomsection The $U$-comparability assumption for NCIs implies that the NCI is indeed a proxy for the API variable. In contrast to $U$-comparability for NCOs, for NCIs, their association with the API variable must exist conditionally on the IV. This assumption rules out variables that are not informative about the IV validity because they are unrelated to the identification threat (and specifically to the API variable), conditional on the IV.
\phantomsection The NCO and NCI definitions are analogous to the conditions that were formalized in previous literature on negative controls. In particular, U-comparability is common in the literature on negative controls lipsitch2010negative,shi2020selective. The NCO and NCI assumptions are similar to the conditional independence assumption of tchetgen2024introduction (see equations 12 and 13).
\phantomsectionDefinitions (ref) and (ref) imply that alternative path variables are themselves negative controls. This is because APO and API variables trivially satisfy both conditions in the definitions. For example, in the previously discussed design of angrist1996children, the concern is that the sex composition of the first two children (the IV) may influence household expenditures (the APO variable) due to hand-me-downs, forming an alternative path to female labor supply (the outcome). rosenzweig2000natural explore this by using a dataset in which clothing expenditures are observed, and use this APO variable as an NCO.
The NCO and NCI assumptions can be weakened to cover more variables that are informative about the validity of the IV design. In (ref), we offer a more general definition of negative controls that allows for direct associations between the NCO and the IV or the NCI and the outcome, not through the alternative path variable if the design is invalid.
A negative control outcome test (NCO test) is any statistical test of independence between the IV and an NCO. The null hypothesis is $H_0$: $Z \protect\mathpalette{\protect\independenT}{\perp} NC$. For example, martin2017bias regress their NCO, Republican vote share in 1996, on the IV, Fox channel positioning. Under the null, the coefficient on the IV in this regression should equal zero. Indeed, they found no evidence to reject this hypothesis, which supports their design validity.
The following theorem states that rejecting the null hypothesis implies a violation of outcome independence or the exclusion restriction.
All proofs are given in (ref). The appendix proof covers a more general version of Theorem (ref) for designs that include control variables (discussed in Section (ref)). For the case without controls, the sketch of the proof is as follows. By the NCO assumption, the dependence between the IV and the NCO implies an association between the IV and an APO variable ($Z \cancel{\protect\mathpalette{\protect\independenT}{\perp}} U$). By path indication, $Z \cancel{\protect\mathpalette{\protect\independenT}{\perp}} U$ indicates an alternative path between the IV and the outcome ($Z\cancel{\protect\mathpalette{\protect\independenT}{\perp}}Y(x)$); i.e., the IV design is invalid.
Similarly, a negative control instrument test (NCI test) examines whether the outcome and the NCI are independent, conditional on the IV. Formally, the statistical test is for the null hypothesis $H_0:$ $NC \protect\mathpalette{\protect\independenT}{\perp} Y|Z$. If the NCI is associated with the outcome conditional on the IV, this necessarily implies that the IV design is not valid, as stated in the following theorem.
For example, nunn2014us regress their outcome, conflicts, on various alternative crop production (e.g., oranges), which are the NCIs, controlling for the original IV, wheat production. They are unable to reject a zero coefficient on alternative crops. Hence, they do not find an indication of a problem with the IV.
NCI tests typically require conditioning on the IV, as the NCI may be associated with the outcome even in valid IV designs. This association arises because the NCI is often associated with the IV, which in turn influences the outcome through the treatment. For example, in Panels C and D of (ref), the NCI and the outcome are associated through the IV, even if no alternative path exists and the IV design is valid. In nunn2014us, orange production (the NCI) is associated with conflicts (the outcome), as both are associated with wheat production (the IV).
However, if the NCI and IV are independent, conditioning on the IV is not required. In such cases, researchers can use an unconditional independence test for the null $H_0:$ $NC \protect\mathpalette{\protect\independenT}{\perp} Y$, as formalized in the following theorem.
\phantomsection (ref) provides a version of this theorem with control variables, in which the IV and NCI need to be conditionally independent only given the set of controls.
\phantomsectionSituations where $NC\protect\mathpalette{\protect\independenT}{\perp} Z$ (so, per the theorem, unconditional NCI tests may be valid) can occur when considering violations of the exclusion restriction assumption. Panel A of (ref) provides an example. Consider the context of jacob2007crime, who study the effect of lagged crime ($X$) on current crime ($Y$). They use lagged weather as an IV ($Z$) for lagged crime. The API variable is temporal displacement of economic activity ($U$): Lagged weather can postpone economic activity to the current period, which could in turn affect current crime ($Y$), thus violating the exclusion restriction. In this context, a different variable that displaces economic activity can be used as an NCI. For example, payday cycles ($NC$) are known to impact the timing of economic activity Hastings2010. Payday timing is independent of weather. Therefore, an association between payday and crime would imply a violation of the exclusion restriction assumption. In this case, no conditioning on $Z$ is needed.\footnote{The NCI assumption is that payday timing only correlates with crime only through the timing of economic activity.} By contrast, in contexts where a violation of outcome independence is suspected, the IV and the NCI are typically associated as well (as in Panel C of (ref)). Therefore, the NCI test should condition on the IV.
As a result, unconditional independence tests between a negative control and the outcome are unique to IV settings. In non-IV settings, there is no exclusion restriction, and therefore, independence tests between a negative control and the outcome, carried out to detect unmeasured confounding, are always done conditionally.\footnote{The analog of NCI in non-IV settings is negative control exposure (NCE). NCE tests always condition on the exposure.}
\phantomsection Nevertheless, researchers can choose to always control for the IV. Since both $NC$ and $Z$ are observed, the condition $NC \protect\mathpalette{\protect\independenT}{\perp} Z$ can be empirically tested. However, researchers might opt to skip this test and condition on the IV anyway. In a linear model, adding an additional control that is uncorrelated with $NC$ will not affect the coefficient estimate for $NC$ asymptotically. Furthermore, if $Z$ has a causal effect on $Y$, including it in the regression can improve the precision of the estimation.
\phantomsection In many cases, the IV is believed to be valid only conditionally on certain control variables. For example, in papers that use judge assignment as an IV, the assignment of judges is quasi-random only within date and location kling2006incarceration. Therefore, the independence assumption is satisfied only conditionally, and the IV design is valid only once controlling for date and location.
Formally, let $C$ be the vector of controls. Similar to the case without controls, outcome independence and exclusion restriction together imply $Z \protect\mathpalette{\protect\independenT}{\perp} Y(x)|\ C$. (ref) presents the theory of negative controls when control variables are included.
When the IV is presumably valid only conditional on a vector of control variables $C$, an NCO test is a test for the null hypothesis
Similarly, for NCIs, the null hypothesis is
While accounting for controls in an IV analysis can be done in a variety of ways abadie2003semiparametric, the large majority of applications use a two-stage least squares (2SLS) specification. This specification makes additional functional form assumptions. Most negative control tests used in practice adopt the same functional form as the 2SLS.
In particular, NCO tests typically adopt the functional form for how the IV depends on the control variables. To avoid excessive notation, let $C$ also denote the set of controls in a 2SLS specification.\footnote{The vector $C$ may include, for example, a quadratic function of one of the original controls or interactions. For ease of notation, $C$ would always include the intercept.} \phantomsection blandhol2022tsls show that 2SLS requires the following linearity assumption to satisfy their definition of a weakly causal estimand.\footnote{ A weakly causal estimand is a positively weighted average of subgroup-specific treatment effects.}
Combining the null hypothesis of NCO tests (ref) and rich covariates, we expect that
This equation provides a more specific null hypothesis for conditional independence testing. This hypothesis can be tested by regressing the IV on the vector of controls and the NCO. The following corollary formalizes this argument.
In many cases, researchers run the reverse regression in which the NCO is the outcome variable. This practice is equivalent, as formalized in the following corollary.
NCI tests typically adopt the functional form of the relationship between the outcome and the IV and the control variables. Specifically, NCI tests often use the same structure as the reduced form equation. Therefore, they implicitly make the following assumption.
Combining the null hypothesis (ref) with the CSRF assumption, we expect that
This equation also provides a more specific null hypothesis, which can be tested with OLS. The following corollary shows that such an OLS jointly tests IV violations due to an alternative path and CSRF.
Corollaries (ref), (ref), and (ref) imply that, in the tests discussed, the null hypothesis can be rejected in IV designs that satisfy outcome independence and exclusion if functional form assumptions are violated. For NCO tests, the null can be rejected because the rich covariates assumption is not satisfied. In such cases, researchers can still estimate a causal effect by modifying the functional form or using methods other than 2SLS blandhol2022tsls. For NCI tests, the null can be rejected because the CSRF assumption is violated. However, unlike rich covariates, CSRF is not a necessary assumption for 2SLS analysis, implying that negative control tests sensitive to this assumption could reject perfectly valid IV designs.
For example, an NCI test can reject the null in designs where the IV is randomly assigned due to CSRF violation. Random assignment guarantees that outcome independence and rich covariates hold. Assuming the exclusion restriction holds, the design is valid. However, CSRF could still be violated if the IV has a nonlinear effect on the outcome or a heterogeneous effect across control vector values. In such cases, an NCI test may reject the null in (ref), despite the design being valid.
This section offers guidelines for the implementation of negative control tests. We recommend that researchers follow four steps, summarized in (ref). First, when possible, researchers should articulate specific threats to the validity of the IV design and characterize the alternative paths variables, as discussed in Section (ref). Second, researchers should survey available data to identify suitable negative controls---variables that can serve as proxies for the unobserved alternative path variables. These proxies should satisfy the NCO or NCI assumption (see Definitions (ref) and (ref)). Examples are discussed in Section (ref).
Third, researchers should choose a statistical test for independence between the NCO and the IV or between the NCI and the outcome, conditioning on the IV. For IV designs that require conditioning on a set of controls, negative control tests should also condition on these controls. Section (ref) discusses particular test specifications, their validity, and the assumptions they test, which in some cases also include functional form assumptions. Fourth and finally, researchers should interpret the result and conduct further diagnostics if the test rejects the null, as discussed in Section (ref).
To illustrate these recommendations, we apply them to IV designs used in prior work. We chose four widely cited papers published in the American Economic Review with publicly posted replication data. We use autor2013china and deming2014using to discuss NCO tests and ashraf2013out and nunn2014us to discuss NCI tests.\footnote{ashraf2013out and nunn2014us are the two most cited AER papers published since 2013 that use an NCI test. Similarly, autor2013china is the most cited AER paper published after 2013 that uses an NCO test. deming2014using was selected to demonstrate how our proposed follow-up analysis can be used to diagnose and correct problems with the IV design in Section (ref).} (ref) summarizes the IV designs in these papers and the negative controls they used in their falsification tests. (ref) provides additional details on our analyses.
Two guiding questions can help researchers characterize potential violations of IV validity. This characterization can assist in selecting appropriate negative control variables and determining which hypothesis should be tested. The first question is whether the primary concern is a violation of outcome independence (Assumption (ref)) or the exclusion restriction (Assumption (ref)). Both types of violations introduce an alternative path between the IV and the outcome.
Outcome independence is violated when this path is through some factor that affects both the IV and the outcome. As previously discussed, martin2017bias examine whether Fox News viewership influences Republican vote shares using cable channel positions as an IV. The concern is that cable companies may place Fox News in lower channel numbers in conservative locations, where voters lean republican regardless. Violations of outcome independence are illustrated in panels A and C of (ref).
The exclusion restriction is violated when the IV affects the outcome through channels other than its effect through the treatment. For example, as previously discussed, in angrist1996children, the sex composition of the first two children (the IV) may affect female labor force participation (the outcome) not only through its effect on family size (the treatment) but also through its effects on household expenditures due to hand-me-downs. Exclusion restriction violations can occur even with randomly assigned IVs, as in a randomized controlled trial. Violations of the exclusion restriction are illustrated in panels B and D of (ref).
The second question is whether the threat (the alternative path) operates through an APO or an API variable. That is, does the alternative path variable have a known association with the outcome, and the concern is that it may also be associated with the IV? Or does it have a known association with the IV, and the concern is that it may also be related to the outcome? In the first case, the alternative path operates through an APO variable; in the second, it operates through an API variable. In the context of martin2017bias, unobserved conservativeness is an APO variable---it certainly influences voting for Republican candidates (the outcome), yet it is unclear whether it is also associated with Fox News channel placement (the IV). By contrast, in nunn2014us, weather conditions are an API variable---they surely affect wheat production (the IV), and the concern is that they may also directly affect the conflicts in aid recipient countries (the outcome). Note that both outcome independence and exclusion restriction can be violated through either APO or API variables. In (ref), panels A and B demonstrate this for APO variables and panels C and D for API variables.
These two questions can assist researchers in finding relevant negative control variables and choosing the right negative control tests. For APO variables, researchers should search for NCOs and test their association with the IV. For API variables, researchers should search for NCIs and test their association with the outcome, conditionally on the IV. The type of violation (outcome independence or exclusion restriction) can be useful for thinking of relevant negative control types.
In this section, we discuss different types of negative controls, both commonly used in practice and novel ones suggested by the theoretical framework.
\phantomsection Predetermined Variables. Variables fixed before the IV is determined are frequently used in NCO tests. Common examples include lagged outcome variables and demographic characteristics such as gender, race, and age. Predetermined variables are useful for testing outcome independence. In some cases, researchers may choose to use predetermined variables as NCOs even without clear knowledge of which exact APO variable they proxy for. If the IV is associated with a predetermined variable, it could imply that it is affected by something that affects the outcome as well.
However, not every predetermined variable is a valid NCO. First, NCOs need to satisfy U-comparability (Definition (ref))---predetermined variables that are completely unrelated to the outcome are uninformative and should not be used. Second, not all predetermined variables satisfy the NCO assumption. In particular, certain predetermined variables may influence the IV, even if the underlying IV design is valid. For example, when the IV is the child's quarter of birth angrist1991does, the parents' quarter of marriage is not a valid NCO as it likely influences the child's quarter of birth, even if the design is valid.
\phantomsection When IVs are assumed to be quasi-randomly assigned (e.g., lotteries), they cannot be affected by predetermined variables. Researchers could therefore use predetermined variables as NCOs to evaluate the claim that the IV is quasi-random.\footnote{In this case, predetermined variables are NCOs based on the more general Definition (ref). This definition allows for the NCO to be directly associated with the IV if the IV is not quasi-random as claimed.} If the IV is associated with a predetermined variable, it is unlikely to be quasi-random, and hence, outcome independence might not hold. For example, we found multiple predetermined variables in the replication data from deming2014using, which uses an IV constructed based on school lotteries. We use these predetermined variables as NCOs to evaluate outcome independence. We have found an association of the IV with the predetermined variables. In particular, we found that the construction of the IV involved non-random components that require additional controls; see Section (ref).
\phantomsection Predetermined variables are also useful NCOs when the IV is not quasi-random. For example, autor2013china use a shift-share IV for commute-zone exposure to Chinese imports to evaluate their impact on employment. To evaluate this IV, they use predetermined local labor market manufacturing employment as NCOs. The concern is that since industry exposure to Chinese imports is non-random, it might be associated with other labor market conditions, which in turn could be associated with the outcome (e.g., Chinese imports are more pronounced in regions with industries that were declining in Western countries regardless). Such an association would violate outcome independence. We found many additional predetermined variables in the original paper's replication data that could proxy for latent local labor market conditions (e.g., past unemployment). These variables can also serve as NCOs. The NCO assumption requires that any association of the IV with the predetermined variables used as NCOs is driven by an APO variable, i.e., by something that also affects the outcome. This would be violated if, for example, Chinese import penetrated industries due to economic factors that were only relevant in the past and are no longer relevant in the studied period.
IV Leads and Lags. Certain IVs are predicated on serendipitous or chance occurrences (“strokes of luck”). Because such unexpected shocks should not be autocorrelated, leads and lags of the variable used as the IV can serve as NCOs. For example, jager2022substitutable use a worker's premature death as an IV for employee turnover, under the assumption that such deaths occur randomly across firms. A potential concern is that deaths are non-random and reflect riskier conditions in the firm (the APO variable) that directly impact wages (the outcome). This would violate outcome independence. To rule this out, jager2022substitutable use subsequent premature deaths in the same firm as an NCO that proxies for potentially unobservable riskier conditions. The NCO assumption here stipulates that given the risk conditions, premature deaths should not be autocorrelated. A recurring pattern of premature deaths would cast doubt on the assumption that such deaths occur randomly across firms.
Alternative Outcomes. Alternative or unrelated outcomes can also serve as NCOs for two different types of APO variables. First, APO variables can potentially affect the IV, forming an alternative path that violates outcome independence (as in Panel A of (ref)). Alternative outcomes that are affected by the same APO variable can then serve as NCOs. For example, chetty2014measuring leverage teachers' moves between schools to evaluate middle-school teacher value-added measures. The concern is that high-quality teachers may tend to move to schools that experience simultaneous improvements in student quality. Here, the APO variable is the unobserved changes in school quality. To evaluate this threat, chetty2014measuring use as NCOs test scores from subjects not taught by the teacher in question. If the NCO test finds that teacher quality is associated with better outcomes in subjects they do not teach, it would cast doubt on the design's validity. chetty2014measuring focus on middle-school teachers, as opposed to elementary school teachers who teach multiple topics. This is because the NCO assumption requires that the IV will not affect the alternative outcome directly. The IV should also not affect alternative outcomes indirectly via the treatment or outcome.
The second type of APO variables potentially violates the exclusion restriction. The concern is that the IV affects an additional factor (the APO variable), which in turn affects the outcome (as in Panel B of (ref)). In the previously discussed example of angrist1996children, the concern is that same-sex sibship IV may affect female labor supply due to hand-me-downs, thus forming an alternative path. To evaluate this concern, rosenzweig2000natural check the correlation of same-sex sibship and an alternative outcome---clothing expenditure.\footnote{\phantomsection In this example, the NCO is the APO variable itself, so the NCO assumption is trivially satisfied (see Section (ref))}
Variables Similar to the IV That Do Not Affect the Treatment. Researchers often choose NCIs that are similar to the IV but are presumed not to influence the treatment variable. These NCIs typically test outcome independence. They usually share many similarities with the IV and are thus likely to be correlated with the API variable. For example, as previously discussed (and illustrated in Panel C of (ref)), nunn2014us use US wheat production as their IV for US aid and consider the US production of other crops unrelated to US aid (e.g., oranges) as NCIs. These variables are similar, as they are both affected by the same API variables such as weather conditions. Similarly, ashraf2013out replace their original IV, distance from Addis-Ababa, with distance from London, Tokyo, and Mexico City.
\phantomsection In some cases, researchers generate variables similar to the IV on their own. They construct a variable in a similar way to how the IV was constructed but remove the impact on the treatment. For example, de2020consumption study the effect of peers' consumption on own consumption. As an IV, they use economic shocks to firms of distant peers. This IV will be correlated with shocks to large firms (as statistically, they are more likely to affect all workers, including distant peers), which could potentially affect the outcome in other ways. To test this, they use an NCI which they call a “placebo” IV---they calculate the same IV when replacing the real allocation of workers to employers with a random allocation, keeping firm sizes constant.
IV Leads. Future instances of the IV (IV leads) can often serve as effective NCIs for testing outcome independence. For example, moretti2021effect studies the effect of the size of high-tech clusters on productivity. As an IV, he uses predicted cluster size based on the expansion of local firms outside the cluster. moretti2021effect then ascertains that future predicted cluster size is also not correlated with productivity. This relies on the fact that IV leads, which are based on events that occur after the outcome, cannot influence it. IV leads share similarities with the IV and are therefore likely to be associated with the API variable. To satisfy the NCI assumption, the outcome must not influence future realizations of the IV. In the example of moretti2021effect, regional productivity cannot affect the expansion of local firms in other locations.
When practitioners observe IV leads, they need to consider whether they expect the IVs to be autocorrelated. When the IVs are expected to be autocorrelated moretti2021effect an NCI strategy can be used. When the IVs are presumably uncorrelated, and NCO strategy can be applied, as discussed in the previous section.
\phantomsection Causes of the IV. In some cases, researchers may suspect violations of outcome independence through API variables that affect the IV and potentially also affect the outcome. In these cases, researchers can use variables that causally affect the IV as NCIs. This approach does not require full knowledge of the API variable. In the example from angrist1991does, the parents' quarter of marriage influences the child's quarter of birth (the IV) and qualifies as a valid NCI. In this case, API variables are any factors that influence a child's quarter of birth and are suspected to affect wages (the outcome).
Panel B of (ref) illustrates how such NCIs work. When both the API variable and the NCI influence the IV, they are associated conditional on the IV pearl2009causality. If, conditional on the IV, the NCI is also associated with the outcome, it implies that a path exists between the NCI and the outcome through the IV and the API variable. This, in turn, implies that the API variable affects the outcome, violating outcome independence.
\phantomsection To satisfy the NCI assumption, these NCIs should have no association with the outcome other than through the IV.\footnote{If such an association does exist, these variables must be used as controls.} In particular, the NCI cannot directly affect either the treatment or the outcome. In the example of angrist1991does, using parents' quarter of marriage as an NCI requires assuming that marriage timing does not directly influence child schooling or wages.
IV Side Effect Proxies. An IV that not only affects the treatment but also produces a side effect may violate the exclusion restriction. This occurs if the side effect also affects the outcome. In such cases, proxies for the side effect (the API variable) may serve as NCIs.
\phantomsection Panel D of (ref) illustrates a scenario where the IV influences the NCI through the API variable (therefore, the NCI is itself a side effect). For example, in the previously discussed context of jacob2007crime, lagged weather ($Z$) serves as an IV for lagged crime ($X$) to study its impact on current crime ($Y$). Lagged weather also creates intertemporal displacement of economic activity (the API variable $U_4$). The concern is that intertemporal displacement of economic activity affects subsequent crime, thus violating the exclusion restriction. To evaluate whether this alternative path exists, Jacob et al. use traffic patterns, a proxy for economic activity, as an NCI ($NC_4$). Testing whether traffic patterns are correlated with crime, conditional on lagged weather, constitutes an NCI test for this alternative path. If a correlation exists, it suggests that the exclusion restriction is violated, as weather influences crime not only through past crime but also through displaced economic activity. Alternatively, Panel A of (ref) presents another type of side-effect proxy that influences the API variable rather than being influenced by it. For jacob2007crime, that could be other factors that displace economic activity (e.g., payday schedule; see the discussion of this example in Section (ref)).
Negative control variables must satisfy U-comparability. This implies that these variables are indeed associated with the alternative path variable. Some negative controls might satisfy this condition but have only a weak association with the alternative path variable. In this case, if the IV design is not valid, the association between the NCO and the IV or the NCI and the outcome would be difficult to detect without having a large dataset. Therefore, power considerations suggest excluding negative control variables that have only a weak association with the alternative path variable, as they can lower test power. This mirrors the effect of irrelevant control variables in OLS. This issue is especially acute in NCI variables that are intentionally similar to the original IV. Such variables are often strongly correlated with the IV but only weakly correlated with the API variable conditional on the IV.
As discussed in Section (ref), negative control tests assess whether a negative control variable is (conditionally) independent of the IV or outcome. The choice of statistical test for conditional independence should take into account three primary considerations: the estimation method (e.g., 2SLS); the anticipated functional form of the relationship between the negative control and either the IV or the outcome (and the controls); and statistical power. This section discusses both commonly used and underused conditional independence tests and presents examples using replication data from existing work.
The most commonly used NCO falsification tests are based on the original IV reduced-form equation, $Y = \alpha_Z Z + \alpha'_CC + \epsilon$, but replace the outcome with an NCO (e.g., past outcomes). That is, the test estimates the model
and evaluates the null hypothesis $H_0:\beta_Z=0$. This test evaluates outcome independence or exclusion restriction, as well as the rich covariates assumption (Corollary (ref)). Because in 2SLS the rich covariates assumption is necessary for causal interpretation, this test provides useful information regarding both the validity of the IV design and of the 2SLS specification. Since this test uses the same inferential framework as the original study, it can also expose errors in the inference method eggers2021placebo.
\phantomsection When multiple negative controls are available, Model (ref) can be estimated separately for each negative control. However, correcting for multiple hypotheses is necessary, which reduces statistical power. An alternative approach is to jointly incorporate multiple negative controls using the model
where $NC$ now represents a vector of NCOs.\footnote{In most applications, if a vector of negative controls is associated with the IV, at least one of its components will be as well. (ref) provides a theoretical counterexample, but such cases are unlikely in practice as small parameter changes would reverse the result.} \phantomsectionAn F-test can be used to evaluate the null hypothesis $H_0: \gamma'_{NC}=0'$ under standard assumptions.\footnote{With robust standard errors, the common implementation calculates the Wald statistic, divides it by the degrees of freedom, and calculates a p-value from an F-distribution.}
For NCI tests, a similar approach applies, but the outcome is regressed on the NCI, controlling for the IV. Specifically, the NCI (denoted again NC) is added to the reduced-form equation
and the null hypothesis is $H_0:\theta_{NC} =0$. A common mistake in practice is omitting the IV from this regression (see Section (ref)). Doing so can lead researchers to reject the null, even when the IV design is valid. The IV can be omitted in the (rare) event that the NCI is independent of the IV.
This NCI test can still reject the null for valid IV designs (even when conditioning on the IV). Corollary (ref) shows that even if the IV design is valid, the null could still be rejected due to a violation of CSRF (Assumption (ref)). As discussed in Section (ref), the CSRF assumption is not necessary for causal identification and can be violated even under random assignment.\footnote{An exception is binary IV without controls, in which case CSRF is always satisfied.} Therefore, tests based on Model (ref) may reject the null hypothesis, even though the IV design can identify a causal estimand with 2SLS. To relax this sensitivity to linearity assumptions, researchers may opt for semi-parametric or non-parametric tests, which are discussed next.
The tests discussed above inherit the basic functional form assumption as in the 2SLS specification. The NCO tests replace the outcome with the NCO in the reduced-form equation or posit the reverse linear model for the IV. The NCI tests include the NCI additively in the reduced form. Moreover, these tests only test the mean independence of the IV with the NCO or the outcome with the NCI. In some cases, researchers should use more general tests.
\phantomsection Researchers should consider replacing the functional form assumptions in two cases. First, tests with more flexible functional forms can help relax undesirable functional form assumptions. In particular, when using an NCI, researchers should test for non-linear associations whenever possible to avoid testing the CSRF assumption, which is unnecessary. By contrast, the previously discussed linear NCO tests examine the rich covariates assumption, which is necessary for 2SLS. Therefore, tests based on Models (ref) or (ref) are often preferable in such contexts.
Second, more flexible tests are useful when researchers suspect a non-linear association between the NCO and the IV or between the NCI and the outcome. Such non-linear associations also indicate a violation of the IV assumptions and should therefore be tested whenever possible. For example, in studies using crops as an IV nunn2014us, one might be concerned that extreme weather conditions, such as unusually high or low temperatures, could affect both crop yield and conflict incidence (the outcome). In this case, Model (ref) might fail to detect a non-monotonic relationship between an average temperature NCI and the outcome.
To estimate more complex functional forms, researchers can include higher-order polynomial terms or interactions in regression models. Alternatively, they can use semi- or non-parametric tests. The next section discusses a few examples, which we implement in our application examples.
When using estimation methods other than 2SLS, researchers should consider more general conditional independence tests that assess relationships beyond mean independence. For example, IV quantile regression chernozhukov2008instrumental relies on the broader notion of conditional independence. In such cases, researchers can use quantile regression of the IV on the NCO or the outcome on the NCI, with appropriate controls. A variety of other conditional independence tests can also be considered heinze2018invariant, li2020nonparametric. The choice of test depends on the specific context, as there is no uniformly optimal conditional independence test.\footnote{Shah2020hardness show that for continuous distributions, there is no conditional independence test that is uniformly valid and is simultaneously powerful against any conditional dependence types.} Naturally, more complex tests require larger datasets or low-dimensional covariates for reliable implementation. Therefore, these tests may be less informative in small samples or in settings with many covariates.
(ref) presents the NCO tests results using data from autor2013china and deming2014using. Column (2) shows that a test based on a single NCO fails to reject the null hypothesis in both cases.\footnote{For autor2013china, Column (1) replicates the original NCO test from their paper, without control variables. They find the IV is significantly associated with the lagged outcome, albeit with the opposite sign from the main analysis. This association becomes insignificant when all controls are included.} However, Column (3) demonstrates that using multiple NCOs and applying a Bonferroni correction leads to rejection of the null, as does a joint F-test (Column 4). These results underscore how theory-guided inclusion of additional NCOs enhances test power.
To examine non-linear associations, we implement the common semi-parametric approach of Generalized Additive Models hastie1990generalized,wood2006generalized. We express the IV as an additive combination of smooth functions of the controls ($C$) and NCOs ($NC$):
where $f_j$ and $g_k$ are smooth functions estimated via splines. We test whether $g_k = 0$ for all $k$ to assess whether the NCOs are conditionally independent of the IV.\footnote{A similar GAM test can be implemented for NCI tests, by adding smooth functions to Model (ref).}
To test rich covariates (Assumption (ref)) within this framework, researchers can restrict $f_j$ to be linear (i.e., $f_j(C_j) = \gamma_j C_j$) while allowing $g_k$ to remain nonlinear. Such a model is useful when using 2SLS, which requires rich covariates, while suspecting a strong nonlinear association between the NCO and the IV. Columns (5) and (6) in (ref) implement a GAM test with and without assuming linearity in the controls. As expected, the GAM test underperforms in smaller samples.
(ref) presents the NCI test results for nunn2014us and ashraf2013out. Columns (1) and (2) show the results of regressing the outcome on the NCI, both with and without conditioning on the IV. In both studies, conditioning on the IV is necessary. For nunn2014us, failure to condition leads to rejection of the null, suggesting a false rejection due to the path between the NCI and the outcome through the IV. For ashraf2013out, the null is not rejected in either case, likely due to limited statistical power. Columns (3) and (4) of (ref) implement multiple separate NCI tests with Bonferroni corrections and joint F-tests using all NCIs. In both approaches, the null is not rejected.\footnote{Due to the small sample sizes relative to the number of control variables, proper estimation of GAM models is infeasible for both studies. For nunn2014us, we estimate a GAM model assuming linear controls; see (ref).}
Rejection of the Null. Rejecting the null hypothesis in a parametric linear negative control test, such as an F-test, may indicate a violation of the IV assumptions or failure of the linearity assumption (rich covariates for NCO, CSRF for NCI). Researchers can further investigate by directly testing the linearity assumption ramsey1969tests. With sufficient sample size, semi- or non-parametric tests can also test outcome independence or the exclusion restriction without relying on strict functional form assumptions.
When using multiple negative controls, identifying which ones drive the rejection can provide important insights. A diagnostic scatter plot of the correlation of each negative control with the IV against its correlation with the outcome can highlight potential alternative pathways. (ref) displays such diagnostics for deming2014using. \phantomsectiondeming2014using uses predicted school value-added, based on school lotteries, to evaluate school value-added measures. Specifically, the IV is the value added of the student's preferred school if they won the lottery and of their neighborhood school if they did not (see (ref) for details). This IV satisfies outcome independence only when controlling for the relevant school value-added measures and the probability of winning the lotteries. The scatter plot reveals that the NCO with the strongest correlation with the IV is the value added of the neighborhood school. This occurs because the original 2SLS analysis does not control for the neighborhood school value added. This diagnostic thus identifies a fixable problem in the IV construction. This problem is resolved when using the original lottery results as an IV, as outcome independence is satisfied conditional only on the winning probabilities (i.e., the schools the student applied to).\footnote{While estimates using the original lottery as the IV are noisier, we cannot reject the main conclusions.}
One caveat of this diagnostic exercise is that correlations with negative controls may not directly reflect the strength of alternative paths. Negative controls are proxies, and so the strength of their correlations with the IV or outcome depends on the strength of their correlation with the alternative path variables. Thus, weak correlations between a negative control and the IV or outcome might still mask strong alternative paths.
Non-Rejection of the Null. As discussed in Section (ref), failure to reject the null does not imply that the IV is valid. Two concerns remain. First, the IV may still be invalid due to alternative path variables not captured by the NCO or NCI used in the test. For example, a quasi-random allocation to teachers that is found to be uncorrelated with students' neighborhoods could still be correlated with students' abilities within neighborhoods. Second, an invalid IV design may pass the test due to limited statistical power.
This paper provides a thorough examination of the assumptions underlying negative control tests for IV designs. Our analysis clarifies existing practices and emphasizes several issues of direct practical relevance. First, most current implementations of NCI tests fail to condition on the original IV, which could lead to the unwarranted rejection of valid IV designs. Second, common negative control tests assess not only the outcome independence and exclusion restriction assumptions but also assess specific functional form assumptions. Because these assumptions are replaceable and sometimes unnecessary, researchers should distinguish between the essential IV identification conditions and the ancillary functional form assumptions when interpreting and considering additional tests. Third, our analysis clarifies what variables can serve as negative controls. These include variables that are rarely used in practice, such as variables that causally affect the IV. Moreover, in some cases, negative control variables are readily available in researchers' datasets and should be used to construct more powerful negative control tests. We hope this paper will foster a more systematic and efficient use of negative control falsification tests in empirical IV designs.
\phantomsection While this paper focused on the role of negative controls in testing the IV assumptions, negative controls can also be used for other purposes. As discussed, when the IV design is valid, NCOs can be used to test functional form assumptions. NCOs can also improve the estimation precision, for example, when included as control variables. Since NCOs are correlated with the outcome but not with the IV, including them as controls can improve precision. By contrast, NCIs cannot be used similarly. They do not test a necessary functional form assumption, as we discussed in Section (ref). Moreover, including NCIs in the reduced form will only decrease precision because they are correlated with the IV and not with the outcome.\footnote{In case of heteroskedasticity, at face value, NCIs may be used to improve precision in valid designs by including additional moment conditions for their orthogonality with the error as in cragg1983more. However, this can also be achieved by including other functions of the IV.}
{ \setstretch{1.15}
}
\addtab{tab_applications_nco}{Illustrative Applications of Negative Control Outcome Tests}{ This table presents $p$-values from different NCO tests using data from autor2013china and deming2014using. Column (1) replicates one of the original falsification analyses, in which autor2013china replaced the outcome with the NCO in the same 2SLS specification as their main analysis (ibid., Table 2, Part II). deming2014using conducted no falsification tests. Columns 2--6 report $p$-values obtained from additional tests that include the same controls as in the most exhaustive specification of the original analyses. Column (2) reports a single test using one NCO, where the outcome is replaced with the NCO in the reduced form regression. For autor2013china the single NCO is the lagged outcome (in 1970), which is the same NCO reported in Column (1); for deming2014using this is lagged test scores (2002). Column (3) presents a Bonferroni-corrected $p$-value for multiple tests using all the NCOs, with the same specification as Column (2). Column (4) uses an F-test (Model (ref)) with all NCOs jointly. Columns (5) and (6) use GAM tests with linear and smoothed controls, respectively (Model (ref)). } \addtab{tab_applications_nci}{Illustrative Applications of Negative Control Instrument Tests}{ This table presents $p$-values from different NCI tests using data from nunn2014us and ashraf2013out, applying their original sets of NCIs (three and ten NCIs, respectively). Column (1) shows a single linear NCI test that, inappropriately, does not condition on the IV. The NCI with the lowest $p$-value is shown (grape production for nunn2014us and distance from Mexico City for ashraf2013out). Columns (2)--(4) condition on the IV: Column (2) implements a proper linear NCI test (Model (ref)) using the same NCI as Column (1); Column (3) applies Bonferroni correction for multiple linear NCI tests; Column (4) uses an F-test for all NCIs jointly.
}