Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
272,606 characters · 47 sections · 109 citation commands
Causal inference with misspecified exposure mappings: separating definitions and assumptions
\makeatletter
\makeatother
\doparttoc \faketableofcontents
Experimenters use exposure mappings to investigate complex causal effects involving interference between units. An exposure mapping is a terse representation of the nominal treatments assigned to the units under study. The representation facilitates both definition and estimation of causally relevant exposure effects, provided that the exposures are correctly specified. The exposures are correctly specified when they capture all causal information pertaining to the nominal treatments. Recognizing that it is difficult to construct correctly specified exposures, this paper considers estimation of exposure effects when the exposures are misspecified.
The central insight is that exposure mappings conventionally fill two roles. The first role is to capture aspects of the nominal treatments deemed relevant or interesting for the question at hand. To serve this purpose, exposure mappings do not need to be correctly specified; they can successfully capture relevant aspects of the causality operating in a certain context, including aspects of the interference, without necessarily capture all causal information. The second role is to encode assumptions about the causal structure in the experiment. It is convenient when an exposure mapping fills both roles simultaneously, as this allows experimenters to use standard causal inference techniques also in the presence of interference. However, this can typically only be achieved by making untenable assumptions. The current paper suggests that experimenters should focus primarily on the first role, making the exposures as relevant and interpretable as possible. However, doing so would mean that standard techniques and results no longer apply because the exposures would become misspecified.
The paper demonstrates that separating the two roles nevertheless is practically viable by showing that conventional estimators of exposure effects are consistent for a generalization of the exposure effects under misspecification given relatively mild conditions on the specification errors. Like the assumption of correctly specified exposure mappings, these conditions are generally not testable. Their advantage is instead that they are considerably more tenable than the prevailing assumptions. Assuming that the exposures are correctly specified is equivalent to assuming that the specification errors are zero uniformly. To achieve consistency, it is sufficient to assume that the errors are not strongly dependent. Weak dependence allows for potentially grave misspecification as long as the units' exposures are not misspecified in the same way. This condition does not require that the interference takes any particular structure, making the results widely applicable. This comes at the cost of being somewhat opaque, but the condition is sufficiently approachable to allow experimenters to reason about it in practical situations, as illustrated by several examples throughout the paper.
\citet*{Blattman2021Place} conduct an experiment in the city of Bogot\'a\ to investigate whether intensive policing reduces crime. Among 1,919 streets identified to be crime hot spots, the authors randomly selected 756 streets to be patrolled by police twice as much as the other hot spots. An important concern was displacement of crime. Intensive policing in some streets could cause criminals to move to other streets that are less patrolled, in which case crime is reduced in the targeted streets only by increasing crime elsewhere. The authors addressed this concern by estimating spillover effects. Among all non-hot spot streets that were within 250 meters of one of the 1,919 hot spots, the authors compared streets for which at least one of the neighboring hot spots was assigned intensive policing against streets for which none of the neighboring hot spots were assigned such policing. If there was a displacement effect, streets neighboring treated hot spots would be expected to experience more crime than other streets.
The authors effectively defined two exposures for the non-hot spot streets. The first exposure is to have at least one neighboring hot spot that is treated, and the other exposure is to have only untreated neighboring hot spots. Following the current literature, we would need to assume that these two exposures are correctly specified to estimate their effect. This means that as long as we hold the exposure of a non-hot spot street fixed, it cannot at all matter for its crime level how we otherwise assign policing resources in the city. For example, it cannot matter which hot spots within the 250 meter radius are treated. A treated hot spot 10 meters away must have the same effect as a hot spot 250 meters away. A treated hot spot where a local gang is known to gather must have the same effect as a treated hot spot that is a busy commercial street with many pickpockets. Furthermore, it cannot matter whether streets outside the 250 meter radius are treated. A treated hot spot 260 meters away must be the same as one 5,000 meters away, and the same as when no distant hot spots are treated. The number of treated hot spot streets within the 250 meter radius must also not matter. One treated neighboring hot spot must have the same effect as ten. The list goes on.
The assumption that the binary 250 meter radius exposure is correctly specified is not reasonable. A superficial solution is refine the exposures. For example, one could consider a set of concentric circles with radii $1, 2, 3, \dotsc$ meters centered at each non-hot spot street with a corresponding set of exposures capturing whether a treated hot spot falls in between two of the circles. In this case, a non-hot spot street would be assigned exposure 239 if there is a treated hot spot street 239 meters away. This will typically be an incomplete solution, for several reasons. First, many of the problems listed above still remain: the refined exposures still ignore potential interaction effects; they still ignore the identity of the streets; and it is doubtful whether bee-line distance is the relevant metric of closeness between streets. To make the exposures truly correctly specified, we would need thousands, if not millions, of them. Second, it is difficult to interpret and present the results from a study with more than a handful of exposures. Experimenters can hardly include tables with thousands or millions of estimated exposure effects. Third, an experiment with many exposures will often suffer from positivity violations. That is, most units will have no probability of being assigned most of the exposures. For example, a non-hot spot street that does not have a neighboring hot spot in the interval 100 to 200 meters cannot be assigned exposures 100 to 200. This means that the outcome for each exposure can be estimated only for a small subset of the experimental units. The subsets will differ between exposures, making it difficult, or impossible, to estimate exposure effects because comparisons between exposures will be confounded by differences in the composition of the subsets. While it sometimes is reasonable to make some refinements to the exposures, such refinements will rarely ensure that the exposures are correctly specified.
The suggestion of the current paper is to shift the focus from creating exposures that are correctly specified to creating exposures that are relevant for the question at hand. In the experiment in Bogot\'a, we should ask what exposures would be informative to the policymaker, rather than asking what exposures are correctly specified. If we suspect that potential displacement effects are strongest locally, the binary 250 meter radius exposure is a reasonable way to capture and measure those effects. In doing so, we do not necessarily mean to say that the displacement effects are all the same within that 250 meter radius, nor that there is no displacement effect beyond 250 meters. We simply want to measure the average difference in crime level between when at least one neighboring hot spot is assigned to be policed intensively and when all neighboring hot spots are assigned ordinary policing. While these exposures do not capture all of the interference dynamics related to policing and crime in Bogot\'a, they are relatively easy to interpret, also for laypeople, and they are informative for policy. The current literature would have us believe that it is not possible to precisely estimate this type of effect because the exposures are undoubtedly misspecified. The purpose of this paper is to show that it is possible.
The idea of exposure mappings has its origin in Halloran1995Causal, who consider causal inference under interference and provide foundational definitions. This initial work was extended by Sobel2006What and Hudgens2008, who describe effect estimators for exposures based on proportions of treated units in disjoint groups of units. Manski2013Identification and Aronow2017Estimating recognized that the key methodological tool in the prior literature was a terse description of the full treatment vector. They used this insight to generalize the approach beyond proportions of treated units in disjoint groups to arbitrary summaries of the nominal treatments, as formalized by exposure mappings. The terminology of exposures and exposure mappings used in this paper comes from Aronow2017Estimating, but the underlying idea is essentially identical to the concept of effective treatments in Manski2013Identification. This literature assumes that the exposures are correctly specified. The assumption transforms the interference problem into a causal inference problem without interference at the level of the exposures, which can be solved using standard techniques.
The assumption of correctly specified exposures has a direct parallel to the conventional no-interference assumption. The necessity of the no-interference assumption when estimating average treatment effects has recently been investigated by \citet*{Saevje2021Average}. These authors show that a generalization of the average treatment effect can be precisely estimated even in presence of unmodeled interference. The current paper connects these ideas with the ideas in Manski2013Identification and Aronow2017Estimating, extending the results to exposure mappings of arbitrary complexity. Additionally, the current paper derives its results using quantitative interference measures, whereas Saevje2021Average used a qualitative measure. That is, the interference measure used here takes into account not only whether interference exists between units, but also the strength of that interference. The differences between these two assumptions are discussed in detail in Section \suppref{sec:conditions-sah} in the online supplement.
To the best of my knowledge, the general case of misspecified exposures has not previously been investigated, but restricted forms of misspecification have. Partly building on insights of the current paper, Leung2022Causal develops methods to estimate exposure effects under misspecification when the interference structure is approximately known. Leung assumes that the strength of the interference between units decays in the geodesic distance between vertices representing the units in a known graph. This assumption is considerably weaker than the conventional assumption that the exposures are correctly specified, but it is more restrictive than the setting considered in this paper. However, Leung2022Causal can leverage the additional structure imposed by the interference graph to answer to some of the questions that remain open in the setting considered here. Section \suppref{sec:leung-connection} in the online supplement describes the connection to Leung2022Causal in more detail.
Egami2021Spillover studies estimation of spillover effects in partially unobserved interference networks, which is a way to formalize certain forms of misspecification. Li2022Random consider estimation of direct and spillover effects when the interference graph is generated by a graphon. They do not require knowledge of the graph or the graphon for some of their results, but such knowledge is required for the estimation of spillover effects. Wager2021Experimenting and Munro2021Treatment consider experimentation when units interfere through an equilibrium mechanism, such as a market price. Section \suppref{sec:example-market} in the online supplement explains how this setting can be understood in the framework of the current paper. \citet*{Wang2022Design} describe a new type of estimand that provides an alternative way to describe spillover effects, and they show how this effect can be precisely estimated under certain types of misspecification.
Consider a sample of $n$ units indexed by $\setb{1, \dotsc, n}$. Each unit is assigned one of two possible treatments $z_{i} \in \setb{0, 1}$. The assignments of all units are collected in $\boldsymbol{z} = \paren{z_{1}, \dotsc, z_{n}}$, and the set of all possible assignments is $\mathcal{Z} = \setb{0, 1}^n$. A function $y_{i} \colon \mathcal{Z} \to \mathbb{R}$ gives the realized outcome for unit $i$ under a specific, potentially counterfactual, assignment. That is, $\poi{\boldsymbol{z}}$ is the response of unit $i$ when the treatments are assigned as $\boldsymbol{z} \in \mathcal{Z}$. The elements of the image of the function are called potential outcomes. The potential outcomes are assumed to be well-defined throughout the paper, which requires that no hidden versions exist of the treatments in $\mathcal{Z}$. The function $y_{i}$ may depend on the full vector $\boldsymbol{z}$, so no restrictions are made at this point on the interference between units.
The treatments are assigned at random. Let $\boldsymbol{Z}$ be a random vector denoting the randomly selected treatments. The distribution of $\boldsymbol{Z}$ is called the assignment mechanism or experimental design. The support of $\boldsymbol{Z}$ may be smaller than $\mathcal{Z}$, so that some assignment vectors are never realized by the design. The design is the sole source of randomness under consideration in this paper, and the sample of units and their potential outcomes are considered non-random and fixed. The observed outcome $Y_{i}$ for unit $i$ is defined as the potential outcome corresponding to the randomly selected treatment vector: $Y_{i} = \poi{\boldsymbol{Z}}$.
The asymptotic regime used in the large sample investigation considers a sequence of fixed samples implicitly indexed by $n$. All quantities pertaining to the sample will thus have their own sequences. This type of regime has been used extensively in the literature on design-based sampling. It has more recently seen use in the design-based causal inference literature Lin2013Agnostic.
The potential outcomes contain all causal information pertaining to the treatments, and any causal quantity can be expressed solely using them. However, definitions of such causal quantities may be complex, and it is often difficult to formulate relevant and interesting effects when one is working with $2^n$ distinct potential outcomes for each unit. Exposures and exposure mappings are used to make the definitions more interpretable and intuitive. The idea is that different assignment vectors often share a similar interpretation. For example, in the illustration in Section (ref), having a treated hot-spot street at a distance of 50 meters might have the same causal interpretation as a treated street at 75 meters. The purpose of an exposure mapping is to encode this type of similarity. In particular, two assignment vectors are mapped to the same exposure if they are deemed similar with respect to the application at hand. The exposures can therefore be seen as nothing more than labels on subsets of $\mathcal{Z}$ that have the same or similar interpretation.
To state this formally, consider a set of exposure labels indexed by $\Delta \subset \mathbb{N}$. An exposure mapping is then a function $d_{i} \colon \mathcal{Z} \to \Delta$ for each unit that maps from all possible assignments to the exposures. The exposure of unit $i$ is $\exmi{\boldsymbol{z}}$ when the treatment assignments are $\boldsymbol{z}$. If $\exmi{\boldsymbol{z}} = \exmi{\boldsymbol{z}'}$, then $\boldsymbol{z}$ has a similar interpretation as $\boldsymbol{z}'$ with respect to unit $i$. The realized exposure is a random variable because the treatments are randomly assigned. Let $D_{i} = \exmi{\boldsymbol{Z}}$ denote the realized exposure for unit $i$, and let $\prei{d} = \Pr{D_{i} = d}$ describe its marginal distribution.
The current convention is to assume that exposure mappings are correctly specified. The assumption states that $\poi{\boldsymbol{z}} = \poi{\boldsymbol{z}'}$ whenever $\exmi{\boldsymbol{z}} = \exmi{\boldsymbol{z}'}$ for all units $i \in \setb{1, \dotsc, n}$ and assignments $\boldsymbol{z}, \boldsymbol{z}' \in \mathcal{Z}$. This implies that a function $\tilde{y}_{i} \colon \Delta \to \mathbb{R}$ exists for each unit such that $\pocsexi{\exmi{\boldsymbol{z}}} = \poi{\boldsymbol{z}}$ for all $\boldsymbol{z} \in \mathcal{Z}$, meaning that the exposures are assumed to capture the complete causal structure in the experiment. As noted in the introduction, this forces the exposures to serve two roles: to encode which assignment vectors have similar interpretation and to encode the (presumed) causal structure in the experiment.
If the exposures are indeed correctly specified, the full treatment vector provides no causal information in addition to what a unit's exposure already provides. This implies that we can use $\pocsexi{d}$ defined on $\Delta$ rather than the more cumbersome potential outcomes $\poi{\boldsymbol{z}}$ defined on the full $\mathcal{Z}$ without loss of information. The reduction in complexity can be considerable, because $\abs{\mathcal{Z}}$ grows exponentially in $n$ while $\abs{\Delta}$ typically is fixed. We can then define causal effects in the usual way, as contrasts between potential outcomes produced by the exposures. For example, the average causal effect of exposure $a \in \Delta$ relative to $b \in \Delta$ would be $n^{-1} \sum_{i=1}^n \braces[\big]{\pocsexi{a} - \pocsexi{b}}$. This is the definition in Aronow2017Estimating. The interpretation of these effects tends to be straightforward, because the exposures are typically chosen to be easy to interpret.
It is not possible to construct a function $\tilde{y}_{i} : \Delta \to \mathbb{R}$ such that $\pocsexi{\exmi{\boldsymbol{z}}} = \poi{\boldsymbol{z}}$ when the exposures are misspecified. One alternative is to extend $\Delta$ with more labels until the exposure mappings become correctly specified, which might require that $\abs{\Delta} \approx \abs{\mathcal{Z}} = 2^n$. But, as noted in Section (ref), this will generally not be a feasible or desirable solution.
A more productive alternative is to make the definition of the exposure effects robust to misspecification. We can achieve this by creating analogues of the exposure-based potential outcomes that remain unambiguous even when the exposures are misspecified. Let $\bar{y}_{i} : \Delta \to \mathbb{R}$ be a function such that $\poexi{d} = \Es{\poi{\boldsymbol{Z}} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = d}$, where the expectation is taken over the design. The interpretation of $\bar{y}_{i}$ is essentially the same as for $\tilde{y}_{i}$. The function captures the expected potential outcome under each exposure for each unit, so $\poexi{d}$ is the potential outcome we expect to be realized when unit $i$ is assigned to exposure $d$ under the current design. A definition of an exposure effect under misspecification is now immediate.
Effects building on this idea have previously been discussed in the literature. The earliest examples were introduced by Hudgens2008. The authors derive their main results assuming that the implicit exposures are correctly specified, but they define their effects allowing for some misspecification. They achieve this by marginalizing over all assignments that map to the same exposure, exactly as in Definition (ref). Aronow2017Estimating derive the expectation of their estimator after relaxing their assumption that the exposures are correctly specified. They show the expectation is a particular weighted average of the potential outcomes defined on the full treatment vector, and this average can be shown to coincide with Definition (ref). When the exposures are correctly specified, the expected exposure effect defined here coincides with the conventional exposure effect defined by Aronow2017Estimating.
Because the expected potential outcomes $\poexi{d}$ use the distribution of $\boldsymbol{Z}$ conditional on $D_{i} = d$ in the marginalization, the expected exposure effect $\ef{a, b}$ captures two aspects. The obvious captured aspect is the difference in the outcomes produced by changing the assigned exposure from $a$ to $b$. The second aspect is more subtle. When the exposures are misspecified, a unit's outcome could differ when other units' exposures or treatments are changed even holding its own exposure fixed. Because of the conditioning on $D_{i} = d$, the distribution of other units' treatments might be different in the contrast, and the expected exposure effect could capture this difference. Saevje2021Average discuss an alternative estimand that only captures the first aspect when the exposure mapping is $\exmi{\boldsymbol{z}} = z_{i}$. However, as discussed in Section \suppref{sec:other-estimands} in the online supplement, the nature of many exposure mappings prevents similar effects to be defined more generally, because the definition would involve potential outcomes that are inherently unrealizable. Similarly, when $\exmi{\boldsymbol{z}} = z_{i}$, one can eliminate the second aspect by assigning treatments independently between units, but that approach is not typically viable here because most exposure mappings introduce strong dependencies between exposures even if the nominal treatments are independent.
Consider Definition (ref) in the context of the experiment in Bogot\'a\ discussed in Section (ref). If there are spillover effects of crime between neighboring streets, we would expect the outcomes under the first exposure to be different than under the second exposure. However, because we allow for misspecification, there will be a distribution of outcomes for each of the exposures. The expected exposure effect compares the centers of these two outcome distributions. This comparison provides insights to the policymaker about whether displacement of crime is an important concern in Bogot\'a. If the effect is found to be small, the policymaker might have reason to target crime fighting measures exclusively on hot spot streets. A more targeted approach would likely be more effective in preventing crime in the hot spots themselves, and the small expected exposure effect indicates that we do not need to be concerned about crime displacement, at least not locally. But if the effect is found to be positive and large, the policymaker might have reason to consider a less targeted approach that would be better at preventing some of the displacement.
The Bogot\'a\ experiment also provides a context in which we can understand the second, more subtle aspect captured by the expected exposure effect, as discussed above. Consider a clustered design where Bogot\'a\ is divided into a grid of squares with one kilometer sides, and treatment is assigned to either all or none of the hot spot streets within each square. With this design, a street with no treated hot spots within a 250 meter radius tends to have fewer treated hot spots also within the annulus (i.e., ring) with inner and outer radii of 250 and 500 meters. Therefore, the exposures here act as proxies for treatment assignments beyond 250 meters. If we find a large effect using this design, we cannot conclude that it necessarily is the 250 meter radius that is causally relevant; the effect could have been zero using the same exposure mapping if we had used a different design, because the exposures may then not act as proxies in the same way. This illustrates that experimenters must carefully consider how their exposures interact with their design when they interpret exposure effects.
Misspecification introduces specification errors. The errors can be formalized as differences between the potential outcomes based on the full treatment vector, which is known to be correctly specified, and the potential outcomes based on the exposures, which may be misspecified.
The assumption that the exposures are correctly specified is the same as assuming that the specification errors are zero with probability one for all units. This insight suggests a way to weaken the assumption. Rather than assuming that the specification errors are zero, it may be sufficient to ensure that they are small, or perhaps only that they are sufficiently controlled in some other way.
Small specification errors are indeed sufficient to precisely estimate exposure effects, but such a condition is unnecessarily strong. Instead, the critical aspect is the dependence of errors between units. There are several ways to formalize this dependence. The route explored here is to measure the dependence by the expectation of the product of two units' specification errors conditional on the event that the units are assigned the same exposure: $\E{\varepsilon_{i}\varepsilon_{j} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = d}$. The expectation is defined to be zero if the event $D_{i} = D_{j} = d$ has probability zero. The overall dependence is captured by the following quantity.
We can understand this quantity to capture two sources of dependence. The first is the conditioning event itself, capturing the fact that knowledge about $j$'s exposure can provide information about $i$'s outcome in excess of the information provided by $i$'s exposure. For illustration, consider again the experiment in Bogot\'a. It is not reasonable to assume the binary exposures used here are correctly specified, so $\poi{\boldsymbol{Z}}$ will vary even when $D_{i}$ is fixed, and $\varepsilon_{i}$ will not be zero. This means that other units' exposures could provide information about a unit's specification error. If $i$ and $j$ are two non-hot spot streets with partially overlapping 250 meter radii, then knowing that $j$ has one or more neighboring treated hot spots gives us reason to suspect that $\varepsilon_{i} > 0$ even if we already knew that $i$ had one or more neighboring treated hot spots. This is because when both $i$ and $j$ have neighboring hot spots that are treated, we have reason to believe that more than one of unit $i$'s neighboring hot spots are treated, indicating a greater local police presence that potentially could displace crime.
The second source is dependence in excess of what can be explained by the conditioning event. This captures the fact that two units’ errors can be dependent if misspecified in the same way even if the exposures themselves provide no information about the outcomes. This might be information about the assignment vector $\boldsymbol{Z}$ that is altogether lost by the exposure mapping, or information that can only be captured by intricate combinations of many exposures. An example of this is general equilibrium effects. Returning once more to the experiment in Section (ref), we can imagine two or more equilibria for crime in Bogot\'a. The first could be an overall low-crime equilibrium achieved when sufficiently many hot spots are policed intensively. This could perhaps be because crime becomes too costly, so criminals find other ways to make a living or move to other cities. The second could be a high-crime equilibrium achieved when few hot spots are policed. The exposures might be correctly specified within each equilibrium, so that learning other units' exposures provide no more information about a unit's potential outcome as long as we remain in the same equilibrium. However, if the experimental design induces variation in which equilibrium is realized, the specification errors will typically be highly dependent between units. The equilibrium will act as a coordinating force for the specification errors, making $\bar{\varepsilon}_{d}$ large. This can be remedied either by ensuring that there are no such global equilibrium effects, or by picking an experimental design that does not induce such variation, so only one equilibrium is realized with high probability. The numerical example in Section (ref) considers a setting with this type of equilibrium phenomenon.
An extended discussion about the specification errors is provided in Section \suppref{sec:specification-errors} in the online supplement. This includes a decomposition of the specification error that formalizes the two sources of dependence discussed above, an investigation of how the specification errors relate to the interference conditions used by Saevje2021Average and Leung2022Causal, and simulation studies that elucidate what the specification errors are and how they relate to estimation of exposure effects in concrete settings.
The generality of the definitions in this section can sometimes make them difficult to reason about. One of the goals of the paper is to highlight the wide applicability of the idea that one can separate the two roles that exposure mappings traditionally have served, and the definitions were made to be general to emphasize this. If one were to find the definitions too opaque to be helpful in practice, one could see them as templates for constructing more concrete and situational definitions in particular contexts. Section \suppref{sec:error-examples} in the online supplement contains several illustrations of such concrete constructions.
Commonly used estimators for exposure effects build on ideas originally introduced in the survey sampling literature. The focus here is the Horvitz--Thompson estimator Horvitz1952:
where $D_{id} = \indicator{D_{i} = d}$ is an indicator denoting whether unit $i$ is assigned exposure $d \in \Delta$.
Proposition 8.1 in Aronow2017Estimating implies that the Horvitz--Thompson estimator is unbiased for the expected exposure effect $\ef{a, b}$. This is true no matter how severe the misspecification might be. See also Lemma \suppref{lem:unbiasedness} in the online supplement.
The results in this paper can be extended to several other common estimators, including the difference-in-means and H\'ajek estimators, and estimators making covariate adjustments. The investigation of these other estimators is relegated to Section \suppref{sec:other-estimators} in the online supplement in the interest of space.
Three regularity conditions, which are not directly related to misspecification, facilitate the following investigation. The first of these conditions is an assumption that the potential outcomes are bounded. This assumption can be weakened without materially changing the results, but it eases the exposition.
The second and third conditions concern the design and the distribution it induces on the exposures. The second condition is a positivity assumption. This assumption can be weakened as well, as explored in Section \suppref{sec:lack-positivity} in the online supplement.
The third condition concerns the dependence between exposures. The exposure mappings could potentially introduce strong dependencies between the realized exposures even if the design is well-behaved for the nominal treatments. For example, if the exposures were defined as the proportion of treated units in the whole sample, it would follow that $D_{1} = D_{2} = \dotsb = D_{n}$, which would make precise estimation impossible. Note that the concern here is not dependencies in the outcomes, which would relate to the actual interference structure in the experiment, but dependencies between the exposure labels selected by the experimenter. The following condition limits the dependence between exposures introduced by the design.
Condition (ref) might appear less familiar than Conditions (ref) and (ref), but it is a somewhat common assumption in the recent literature. Indeed, it corresponds to Condition 4 in Aronow2017Estimating. The current condition is slightly weaker, as it considers quantitative dependence between exposures rather than the binary dependence concept used by Aronow2017Estimating, but the interpretation carries over unchanged.
Note that the experimental design and the exposure mappings are known. It is therefore possible to verify the positivity condition and to measure the amount of design dependence in a particular sample. Indeed, design conditions like these are often seen as innocuous, because experimenters control their designs and can ensure that they hold. This may not be the case when estimating exposure effects. Exposure mappings tend to be complex, and it is often difficult to construct a design that would induce a desired distribution over the exposures. In particular, positivity as stipulated by Condition (ref) will sometimes be difficult to achieve. This is the motivation for the relaxation of the positivity assumption explored in Section \suppref{sec:lack-positivity} in the online supplement.
We now have the components needed to characterize the precision of the estimator. The proof of the following proposition is presented in the online supplement.
The proposition demonstrates that the variance of the estimator is governed by three aspects, as captured by the three terms in the bound. The first term captures variability induced by the fact that the exposures are randomly assigned. That is, even when the exposures are correctly specified and independent, the estimator would still vary because different potential outcomes would be observed for different assignments. The second term captures variability induced by dependence between exposures. That is, even when the exposures are correctly specified, the estimator tends to be less precise when exposures are highly dependent.
The final term of the bound captures variance stemming from misspecification. Recall that Definition (ref) captures the dependence between the specification errors. If the specification errors are strongly positively correlated, the estimator will exhibit more variability. The bound makes clear that the magnitude of the specification errors in itself is less of a concern. Large specification errors will affect the precision, but their effect is absorbed by the first term. The intuition for this is that large but uncorrelated specification errors tend to cancel when averaged.
In order for the estimator to become more precise as the sample increases in size, the overall dependence between the specification errors must be limited. This is captured by the following condition and corollary.
Because the estimator is unbiased, the variance bound directly describes the estimator's asymptotic behavior in a mean square sense. Control over the terms in the bound thus provides consistency through Chebyshev's inequality, resulting in Corollary (ref).
Condition (ref) states that the overall dependence between specification errors becomes smaller as the sample grows. One situation in which this is satisfied is when the interference is local, in the sense that $\E{\varepsilon_{i}\varepsilon_{j} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = d}$ is zero for most pairs of units. This is essentially the condition used by Saevje2021Average. However, Condition (ref) allows for global interference as long as it is limited. That is, $\E{\varepsilon_{i}\varepsilon_{j} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = d}$ can be non-zero for all pairs of units as long as it is not large and positive on average. Section (ref) below and Section \suppref{sec:error-examples} in the online supplement provide concrete examples of such limited global interference.
Consider Condition (ref) in the context of the experiment in Bogot\'a\ described in Section (ref). As discussed in Section (ref), one source of the error dependence is the conditioning event $D_{i} = D_{j} = d$. This is unlikely to be an issue in this setting, because only a small fraction of the streets will have overlapping 250 meter radii. The fraction of overlapping streets will also diminish as the sample grows, presuming that the 250 meter radius is held fixed and the streets do not become more dense. While exposures of more distant streets could occasionally be informative, that will not be the case for most pairs of streets.
The second source of error dependence could potentially be more problematic in Bogot\'a. If equilibrium effects of the type discussed in Section (ref) exist, and design induces variation in which equilibrium is realized, then Condition (ref) is unlikely to hold. For example, we could imagine that intensive policing in a certain combination of hot spots could lead to a prominent figure in one of the city's criminal gangs being apprehended, which in turn could cause an overall low-crime environment that otherwise would not have happened. However, global interference does not need to be a concern, and we can allow for equilibrium effects as long as the same equilibrium is realized with high probability. In fact, we can allow for unstable equilibria if the units' outcomes are not affected in the same way, meaning that we avoid strong positive dependencies. The concern with unstable equilibria is that the outcomes for most units will typically be affected in the same way when the equilibrium changes.
Note that positively correlated errors are the concern here. If the errors are negatively correlated, they act to stabilize the estimator, making it more precise than under correctly specified exposures. While this is theoretically possible, overall negatively correlated errors do not appear to be practically relevant.
As a concrete illustration of the ideas in this paper, consider an example study with global interference, meaning that all units can potentially interfere with all other units. Additional numerical examples and details about this example appear in Section \suppref{sec:error-examples} in the online supplement.
Suppose the experimenter is interested in evaluating the effect of a vaccine against a communicable disease, meaning that interference is expected to be present. The experimenter has access to an observed network through which units are hypothesized to interact, and this network is used to define the exposures. Denote an edge from unit $i$ to unit $j$ in the network as $g_{ij} = g_{ji} = 1$, and otherwise $g_{ij} = 0$. For simplicity, the graph is a cycle graph, meaning that $g_{ij} = 1$ if and only if $\abs{i - j} \in \braces{1, n - 1}$. A conventional Bernoulli design is used, meaning that treatments are assigned independently with equal probability.
There are two exposures of interest. The first exposure, denoted $D_{i} = 1$, is when the unit itself is untreated and both neighbors in the cycle graph are treated. The second exposure, denoted $D_{i} = 0$, is when the unit itself as well as both neighbors are untreated. Because of the Bernoulli design and cycle graph structure, the exposure probabilities are $12.5\%$ for both exposures and for all units. The treatment effect of interest is $\ef{1,0}$, which is the expected indirect effect of having two treated neighbors when being untreated compared to having no treated neighbors.
The potential outcomes are such that units interfere locally through the observed network, but they also interfere globally through behavior akin to herd immunity. The fact that herd immunity could occur is not known or hypothesized by the experimenter. This means that the exposures are misspecified. If the number of vaccinated units passes some threshold, there is no transmission of the disease in the community under study, and the outcome (e.g., viral load) is zero for all units. In more detail, the potential outcome functions are
where $\alpha_i$ is drawn uniformly at random from $[15, 25]$, $\beta_i$ is drawn from $[1, 10]$ and $\gamma_i$ is drawn from $[1, 2]$. The coefficients are drawn once and held fixed over the simulation rounds, mirroring the fact that all randomness under consideration comes from treatment assignment. The parameter $\phi \in (0, 1)$ sets the herd immunity threshold; if a share of $\phi$ units are treated, then herd immunity occurs, and the outcome is zero for all units.
Three versions of this data generating process will be considered, corresponding to three values of the herd immunity threshold $\phi$. The first version, labelled “Rarely,” uses $\phi = 0.52$, meaning that herd immunity occurs when at least $52\%$ of the units are treated. The second version, labelled “Infrequently,” uses $\phi = 0.51$, and the third version, labelled “Frequently,” uses $\phi = 0.5$. Note that $\sum_{j=1}^n Z_{j}$ follows a binomial distribution of $n$ trials with $0.5$ success probability. Herd immunity will therefore occur somewhat rarely under the first version of the data generating process ($\phi = 0.52$), and it grows increasingly rare as the sample size grows. However, when $\phi = 0.5$, herd immunity will be frequent also in large samples, because the distribution of $\sum_{j=1}^n Z_{j}$ is centered around $n / 2$. Indeed, herd immunity occurs with roughly $50\%$ probability under all three versions of the process when $n = 100$. But when $n = 10,000$, herd immunity occurs with a probability of only $0.003\%$ under the first version ($\phi = 0.52$), a probability of $2.33\%$ under the second version ($\phi = 0.51$), and a probability of $50.4\%$ under the third version ($\phi = 0.50$).
Interference is global under all three versions, because there are always situations in which changing a single unit's treatment assignment changes the outcome of essentially all other units. That is, the event $\sum_{i=1}^n Z_{i} < \phi n \leq 1 + \sum_{i=1}^n Z_{i}$ has a non-zero probability of occurring under all three versions, and changing any untreated unit to be treated in this case induces herd immunity. However, such global interference takes place only very rarely under the first version when the sample is large, because changing a single unit's treatment assignment under most assignments will not have global effects. In this sense, global interference does exist under all three versions, but it is not practically relevant for the first two versions of the data generating process, $\phi \in \braces{0.51, 0.52}$, provided that the sample is sufficiently large. This is not the case when $\phi = 0.5$, because there will be non-negligible variation in whether herd immunity occurs no matter how large the sample gets.
The results from the numerical example are presented in Figure (ref). Panel A presents the root of the sum of the magnitude of the overall error dependence for the two exposures, $\sqrt{\abs{\bar{\varepsilon}_{1}} + \abs{\bar{\varepsilon}_{0}}}$, as described in Definition (ref). As the sample grows in size, the error dependence decreases toward zero for both “Rarely” ($\phi = 0.52$) and “Infrequently” ($\phi = 0.51$). This mirrors the fact that variation in whether herd immunity occurs decreases as the sample size grows under both of these versions. However, the error dependence does not decrease for the version “Frequently,” instead flattening out slightly below the 1.25 mark. Therefore, Condition \mainref{cond:limited-error-dependence} does not hold under the third version, and the consistency result of the current paper does not apply.
Panel B presents the root mean square errors of the H{\'a}jek estimator, which in this case coincides with the difference-in-means estimator. The H{\'a}jek estimator is presented here because it is typically more precise than the Horvitz--Thompson estimator. Root mean square errors of the Horvitz--Thompson estimator are reported in the online supplement. The results are qualitatively the same, but differences between the data generating processes are less pronounced for the Horvitz--Thompson estimator because of the overall lower level of precision.
The figure shows that the precision of the H{\'a}jek estimator initially improves for all three versions of the data generating process. The precision continues to improve for the first two versions (“Rarely” and “Infrequently”), and the mean square error is essentially zero when $n=100,000$. However, under the third version (“Frequently”), the mean square error flattens out at around the $0.5$ mark, indicating that the estimator is not consistent. This mirrors what we learned about the behavior of the error dependence measure in Panel A. The reason the precision of the estimator initially improves also under the third version, despite no decrease in the error dependence, is that the first two terms of the variance bound in Proposition \mainref{prop:variance-bound} approach zero as the sample size grows even if the error dependence is large.
The motivating idea of this paper is that exposure mappings are primarily used to collect and describe sets of assignment vectors that share a similar interpretation with respect to the application at hand. It is rare that exposures that serve this purpose are also correctly specified in the sense that they provide a complete description of the causal structure. The results herein show that conventional point estimators perform well even if the exposures are misspecified, provided that the specification errors are only weakly dependent. This gives reassurance to experimenters studying complex causal effects under interference, including spillover effects, that their analyses remain informative even in the event their exposures omit some aspects of the causal structure. An expected exposure effect depends on the implemented design, which can make its interpretation challenging. However, an expected exposure effect will typically be more interpretable than the alternative: an exposure effect based on extremely refined exposures. These insights should prompt experimenters to focus on defining exposure mappings that are relevant for the question at hand. If they instead followed the conventional recommendation and focused on making the exposures correctly specified, the exposures would often be too granular to be relevant and useful.
The investigation highlights several open questions. The first question is whether estimators can be constructed to fully separate the two roles traditionally served by exposure mappings. That is, estimators that can incorporate knowledge about the interference structure in the estimation of an exposure effect without necessarily changing the effect that is being estimated. The results in this paper show that estimation is possible without such knowledge, but precision could perhaps be improved if we can take advantage of all available information about the interference, even if that information is imperfect or incomplete. A step in this direction is the effect definition based on linear functionals described by Harshaw2022Design.
The second open question concerns the limiting distribution of the point estimator under misspecification. Section \suppref{sec:limiting-distribution} in the online supplement describes two approaches to characterize the limiting distribution in this setting, but both approaches require considerably stronger assumptions than the limited error dependence condition used for the convergence results in Section (ref). It remains an open question whether the limiting distribution can be characterized under conditions resembling limited error dependence. A relevant result here is the central limit theorem by Kojevnikov2021Limit for network dependent observations under the assumption that the dependence decays in geodesic distance, which was used by Leung2022Causal to investigate the limiting distribution of estimators of expected exposure effects under approximate neighborhood interference. While this approach requires stronger assumptions than limited error dependence and requires that the interference structure is known approximately, it is undoubtedly an important step in the right direction.
A final set of open questions concerns variance estimation. Section \suppref{sec:variance-estimation} in the online supplement notes that variance estimation under both correctly specified and misspecified exposures is difficult, because complex exposure mappings can make many joint exposure probabilities small or zero. Experimenters can address this problem by making the estimator conservative, but the estimator will often be excessively conservative. Recent work by Harshaw2021Optimized describes methods that can mitigate the conservativeness when the exposures are correctly specified. However, variance estimation appears to be particularly challenging under misspecification, because the units' specification errors can be dependent in such a way to make variance estimators anti-conservative. Also in this case do Kojevnikov2021Limit and Leung2022Causal provide an encouraging result by describing a variance estimator to be used under misspecification when the interference structure is known approximately. Relatedly, the online supplement describes ways to leverage unstructured partial knowledge of the interference in an experiment to partially mitigate the problem. It nevertheless remains an open question whether practically useful variance estimators exist when little is known about the interference structure.
\startsupp{S}{Additional results and proofs}{supp:proofs}
This section provides additional results, explanations and illustrations to aid the understanding of the specification errors and the conditions imposed on them. Section (ref) provides a decomposition of the specification error into two parts to elucidate when error dependence arises and when it is controlled. Section (ref) relates the condition on the error dependence in the current paper to the interference conditions in Saevje2021Average. Section (ref) does the same for the interference conditions in Leung2022Causal. Section (ref) provides several concrete numerical examples that illustrate how the specification errors behave in specific settings.
The main paper defines the specification errors as $\varepsilon_{i} = \poi{\boldsymbol{Z}} - \poexi{D_{i}}$. It is possible to decompose this quantity into two parts, which may provide more intuition for the error. This section explores this decomposition.
Dependence between errors can be separated into two components. The first is the conditioning event itself, capturing the fact that knowledge about $j$'s exposure can provide information about $i$'s outcome in excess of the information provided by $i$'s exposure. An example is when unit $j$ interferes with unit $i$ in a way that is not captured in $i$'s exposure. The second source is dependence in excess of what can be explained by the conditioning event. This captures the fact two units' errors can be dependent if misspecified in the same way even if the exposures themselves provide no information about the outcomes.
We may gain a better understanding about the two components after realizing that the specification errors to some degree are in our control, because we decide how to define the exposures. Consider when $j$'s exposure provide information about $i$'s outcome in excess of the information provided by its own exposure, meaning that the conditioning event is informative. A simple way to eliminate this misspecification is to redefine $i$'s exposure to include also the exposure of $j$. If $i$'s redefined exposure is $\paren{D_{i}, D_{j}}$, no part of $i$'s specification error can be explained by $j$'s exposure because $i$'s exposure already contains this information.
Conceptually, it is straightforward to gradually remove misspecification by redefining the exposures, but such an approach will often prove impractical. The exposures would in that case increasingly depart from their primary purpose of producing an effect that is interpretable and relevant for theory or policy. In particular, if applied to all units in the sample, the redefined exposures would be the intersection of all units' nominal exposures, and much of the reduction in complexity the original exposures provided is lost. However, the idea of redefined exposures suggests a way to formalize the decomposition of the specification error that will prove useful.
Let $\bar{y}_{ij} : \Delta \times \Delta \to \mathbb{R}$ be a function such that $\porexij{d_{1}, d_{2}} = \E{\poi{\boldsymbol{Z}} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = d_{1}, D_{j} = d_{2}}$. That is, $\porexij{d_{1}, d_{2}}$ is the potential outcome of unit $i$ when defined over the exposures of both $i$ and $j$. Such a function may not be unambiguously defined if the event $D_{i} = d_{1}$ and $D_{j} = d_{2}$ has measure zero; that is, when $\preij{d_{1}, d_{2}} = \Pr{D_{i} = d_{1}, D_{j} = d_{2}} = 0$. To accommodate such cases, let the full definition be
This captures the intuition that learning $D_{j} = d_{2}$ provides no information about unit $i$'s outcome under $D_{i} = d_{1}$ if $D_{i} = d_{1}$ is not simultaneously realizable with $D_{j} = d_{2}$.
Because the combination of $D_{i}$ and $D_{j}$ provides more information about the treatment assignments than $D_{i}$ alone, the potential outcome $\porexij{d_{1}, d_{2}}$ based on the refined exposures is a more precise representation of unit $i$'s outcome than the potential outcome $\poexi{d_{1}}$ based on the original exposures. We may therefore interpret the difference between $\porexij{d_{1}, d_{2}}$ and $\poexi{d_{1}}$ as the part of the specification error for unit $i$ explainable by $j$'s exposure.
While $\porexij{d_{1}, d_{2}}$ provides more information than $\poexi{d_{1}}$, it will generally not be correctly specified. That is, $\porexij{d_{1}, d_{2}}$ will not provide complete causal information, in the sense that it does not provide the same information as $\poi{\boldsymbol{z}}$. The remaining error is that which cannot be explained by $j$'s exposure. This part is strictly speaking not unexplainable, because the full treatment vector will always perfectly explain the potential outcomes, but it is unexplainable with respect to pairwise refinements of the exposures. Similar to Definition \mainref{def:specification-error}, we may define the error not explainable by $j$'s exposure as the difference between the actual potential outcome and the outcome predicted by the redefined exposures.
The overall specification error can now be decomposed using the explainable and unexplainable specification errors. In particular, we have $\varepsilon_{i} = \serefij{D_{i}, D_{j}} + U_{ij}$ with probability one for any pair of units $i$ and $j$.
Definitions (ref) and (ref) capture the specification errors pertaining to any particular unit. The following definition aggregates these errors to an overall description of the specification errors in the experiment as a whole.
Definition (ref) captures pair-wise dependencies between errors of units. To understand the definitions, first consider the average explainable error dependence. If $\serefij{a, b} = 0$, then knowing that $D_{j} = b$ provides no insights about $Y_{i}$ in excess of knowing that $D_{i} = a$. Thus, $\serefij{d, d}\serefji{d, d}$ is non-zero only when the exposures of $i$ and $j$ both provide information about the other unit's outcome. This means that the magnitude of the explainable errors $\serefij{d, d}$ matters only insofar that the dependence make them large simultaneously. If the explainable errors are perfectly symmetric, so that $\serefij{d, d} = \serefji{d, d}$ for all pairs of units, then $\bar{e}_{d}$ collapses to a measure of magnitude. However, without perfect symmetry, $\bar{e}_{d}$ is a measure of both magnitude and between-unit coordination in the explainable errors. Indeed, $\bar{e}_{d}$ will be small, or even negative, if the pair-wise explainable errors tend to have opposite signs. These insights are perhaps made clear by the bound
which shows that the average explainable error dependence is upper bounded by the average magnitude of the explainable errors.
Consider a vaccination trial as an example. Unit $j$ in this trial is an asymptomatic potential carrier, meaning that $j$ would not get sick if infected but could potentially spread the pathogen to other units. Unit $i$ on the other hand will show symptoms if infected. Here, the exposure assigned to $j$ provides information about $i$'s outcome in excess of knowing $i$'s exposure, because $j$'s exposure provides information about whether unit $i$ is infected, and thus shows symptoms. Part of $i$'s error is thus explainable by $j$'s exposure, and $\serefij{d, d}$ is non-zero. However, $i$'s exposure contains no information about $j$'s outcome, because $j$ never shows symptoms, so $\serefji{d, d} = 0$. The lack of symmetry means that there is no dependence between the explainable errors of units $i$ and $j$ according to Definition (ref).
Next, consider the average unexplainable error dependence. The fact that $\bar{u}_{d}$ captures dependence is immediate by the use of a covariance in its definition. To build intuition, consider the vaccine trial again. Consider when the exposures capture whether units close to the unit in question are vaccinated (e.g., in their household, or in a neighborhood in a social network). For illustration, assume that the experiment is so large that the vaccinations in the experiment have the potential to induce herd immunity. The exposures of any pair of units will in this case provide little information about whether herd immunity is achieved, because whether or not a particular household is vaccinated matters little in that context. However, if the design of the experiment induces variation in whether head immunity is achieved, then the units' errors will exhibit great dependence even in cases where the explainable errors are small or zero, because pair-wise exposure cannot capture the global behavior. This example is explored numerically in Section (ref).
We can now extend the convergence result in the main paper to use the decomposition of the errors.
As noted in Section \mainref{sec:related-work} in the main paper, the key difference between the interference conditions in this paper and the corresponding conditions in Saevje2021Average is that the current paper uses a quantitative concept of interference, while Saevje2021Average used a qualitative (or counting) concept. It is possible to relate the two concepts. As the counting measure in Saevje2021Average is often easier to understand than a quantitative measure, the exercise provides understanding of the measure in the current paper. Furthermore, the exercise shows that the current measure is a relaxation of the type of measure used in Saevje2021Average. To abstract away from complications that do not provide useful insights for the current purpose, I will consider experimental designs that assign the treatments $z_{i}$ independently (i.e., Bernoulli designs).
Saevje2021Average defined an indicator for each pair of units, $i$ and $j$, that captured what the authors referred to as “interference dependence.” The authors used $d_{ij} \in \setb{0, 1}$ to indicate such dependence, but that notation collides with the exposure mapping notation used in this paper, so I will use $\delta_{ij} \in \setb{0, 1}$ to denote the interference dependence indicator in this discussion. Interference dependence exists between units $i$ and $j$, denoted $\delta_{ij} = 1$, in two situations. The first situation is when $i$ and $j$ are interfering directly with each other; that is, when $i$'s treatment affects $j$'s outcome, or vice versa. The second situation is when there exists a third unit $k$ such that $k$'s treatment assignment affects the outcome of both $i$ and $j$. Both these situations will induce dependence between the outcomes of $i$ and $j$ that is due to the interference. This could prevent the estimator from concentrating, and Saevje2021Average highlight that restricting this type of the interference ensures consistency. They impose the condition that $\sum_{i=1}^n \sum_{j=1}^n \delta_{ij}$ is dominated by $n^2$, meaning that a diminishing fraction of the $n^2$ possible pairs of units are interference dependent.
Saevje2021Average considered estimation of an ordinary average treatment effect, corresponding the exposure mapping $\exmi{\boldsymbol{z}} = z_{i}$. We must extend the idea of interference dependence to general exposure mappings to employ a counting concept of interference in the setting considered in this paper. Also in this setting, there are two situations to consider, corresponding to the two situations above. The first situation corresponds to the case of direct interference between $i$ and $j$. In particular, we say that $i$ and $j$ are interference dependent if $j$'s exposure affects $i$'s outcome while holding $i$'s exposure fixed. Using the notation from Section (ref), we have $\delta_{ij} = 1$ when $\porexij{d_{1}, d_{2}}$ depends on $d_{2}$ while we hold $d_{1}$ fixed, or vice versa. What this tells us is that $j$'s exposure contains information about the treatment of at least one unit that is interfering with $i$, and in this sense, $j$'s exposure interferes directly with $i$'s outcome. The second situation corresponds to the case when a third unit is interfering with both $i$ and $j$. That is, if changing some unit $k$'s treatment affects both the outcomes of $i$ and $j$, while holding their exposures fixed, then we set $\delta_{ij} = 1$. (Often, but not always, this will be equivalent to if changing some unit $k$'s exposure affects both the outcomes of $i$ and $j$.)
To connect the concept of interference dependence to specification errors, I will show that $\delta_{ij} = 0$ implies that the expected product of the units' specification errors conditional on their exposures is zero: $\E{\varepsilon_{i}\varepsilon_{j} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = d} = 0$. Recall that the expected product of the error is the basis of the interference measure used in the current paper.
First, recall from Section (ref) that we can decompose the error as $\varepsilon_{i} = \serefij{D_{i}, D_{j}} + U_{ij}$, where $\serefij{D_{i}, D_{j}}$ is the explainable error and $U_{ij}$ is the unexplainable error. If $\delta_{ij} = 0$, then $i$'s exposure is not affecting $j$'s outcome, and vice versa. This means that $\serefij{d_{1}, d_{2}} = \porexij{d_{1}, d_{2}} - \poexi{d_{1}} = 0$ for all exposures $d_{1}$ and $d_{2}$. Hence, when $\delta_{ij} = 0$, we can write $\varepsilon_{i} = U_{ij}$. As shown by Lemma (ref) in Section (ref), we have $\E{U_{ij} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = d} = 0$ by construction, so
Recall that the unexplainable error is defined as $U_{ij} = \poi{\boldsymbol{Z}} - \porexij{D_{i}, D_{j}}$. Hence, under a Bernoulli design, if there is not a third unit interfering with both $i$ and $j$, then $U_{ij}$ and $U_{ji}$ are uncorrelated conditional on the exposures of $i$ and $j$:
meaning that $\E{\varepsilon_{i}\varepsilon_{j} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = d}$ is zero in that case.
Next, consider when $\delta_{ij} = 1$. The expected product of the units' errors, $\E{\varepsilon_{i}\varepsilon_{j} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = d}$, will generally not be zero in this setting. The reason why the quantitative concept of interference used in the current paper is useful is that the expected product can be small (perhaps very small) even if it isn't zero, and that is sufficient for the estimator to be precise in large samples. What Saevje2021Average effectively are doing is to impose a worse-case bound on expected product whenever it is non-zero. While the approach is somewhat more sophisticated in Saevje2021Average, a simple such worse-case bound is to use Condition \mainref{cond:bounded-pos} in the current paper (i.e., bounded potential outcomes). We then have $\abs{\E{\varepsilon_{i}\varepsilon_{j} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = d}} \leq 4 k_1^2$, as shown in Lemma (ref).
This worst-case bound allows us to upper bound the interference measure in the current paper with the interference measure in Saevje2021Average:
The upper bound on the right-hand side is proportional to the quantity $d_{\textsc{avg}} / n$ in the notation of Saevje2021Average. The key condition in Saevje2021Average is that $d_{\textsc{avg}} = \littleO{n}$. Hence, if this condition holds, then $\bar{\varepsilon}_{d} = \littleO{1}$, which implies that Condition \mainref{cond:limited-error-dependence} (limited specification error dependence) in the current paper holds. In other words, the condition in Saevje2021Average implies the condition in the current paper, and in that sense, the current condition is weaker. The intuition for this is exactly that the condition in Saevje2021Average only considers whether there is interference between units, but not how strong that interference is. That is, Saevje2021Average require the interference to be strictly local, in the sense that sufficiently few pairs of units can interfere with each other. The current paper shows that the estimators under study concentrate also when the interference is global, as long as it is either sufficiently rare or sufficiently weak. This is illustrated in the examples in Section (ref).
Leung2022Causal uses a restricted interference assumption based on a graph, referred to as “approximate neighborhood interference.” The idea is that units in the experiment are represented as vertices in a graph, and as the shortest path between two vertices grows longer, the interference between the corresponding two units becomes weaker. This idea is formalized in Assumption 4 in Leung2022Causal.
Some additional notation is required to formally state this assumption. To the greatest degree possible, I use the same notation as in Leung2022Causal, with the exception that I suppress the indexing on the sample size ($n$) and the graphs ($A$). Let $\mathcal{N}(i, s)$ be the subset of unit indices for which the shortest path between unit $i$ and units $j \in \mathcal{N}(i, s)$ is $s$ or less. In the vernacular of the interference literature, $\mathcal{N}(i, s)$ is the $s$-hop neighborhood around $i$ in the graph. Let $\boldsymbol{Z}'$ be a vector of assignments drawn from the experimental design, independent of the actual assignments $\boldsymbol{Z}$. Define $\boldsymbol{Z}^{(i,s)}$ to be equal to $\boldsymbol{Z}$ on all coordinates $j \in \mathcal{N}(i, s)$ and equal to $\boldsymbol{Z}'$ on all coordinates $j \not\in \mathcal{N}(i, s)$. That is, the treatment assignments in $\boldsymbol{Z}^{(i,s)}$ are the same as the actual treatment assignment for all units in the $s$-hop neighborhood around $i$, and new assignments have been drawn for all units outside of the neighborhood. Critical to the investigation in Leung2022Causal is that the design is Bernoulli, so that all coordinates in $\boldsymbol{Z}$ are independent. This ensures that we can construct $\boldsymbol{Z}^{(i,s)}$ without being concerned about dependencies between assignments inside and outside the $s$-hop neighborhood around $i$.
Now, consider the expected difference in outcomes for unit $i$ under the actual assignment and the artificial assignment just constructed:
The quantity $\theta_{i,s}$ is a measure how much units outside of $i$'s $s$-hop neighborhood matters for its outcome. If the interference is strictly local, in the sense that only units in some $k$-hop neighborhood around $i$ affect $i$'s outcome, then $\theta_{i,s} = 0$ for all $s > k$. The approximate neighborhood interference assumption stipulates that
where $\sup_n$ denotes the supremum over all samples in the asymptotic sequence. The assumption ensures that the interference is local in the graph as long as the neighborhood is made sufficiently large. By itself, the approximate neighborhood interference assumption is somewhat vacuous, because we have $\theta_{i,s} = 0$ for all $i$, $n$ and $s \geq 2$, if we use the complete graph as the interference graph. Most of the action in Leung2022Causal is instead in Assumption 5, which restricts both the topology of the graph and the amount of interference within $s$-hop neighborhoods. We will return to this assumption below.
To connect the interference conditions in the current paper with those in Leung2022Causal, I will derive an upper bound for the overall error dependence (Definition \mainref{def:error-dependence} in the current paper) in terms of quantities used by Leung2022Causal. I will then show that Assumption 5 in Leung2022Causal implies that Condition \mainref{cond:limited-error-dependence} in the current paper holds; that is, that the overall error dependence diminishes. This connects the two papers, and specifically shows that the conditions in the current paper are weaker and more general. This is not meant as a critique of the results in Leung2022Causal, as his results are sharper than those in the current paper, and his assumptions might be easier to interpret in some contexts.
Recall the definition of the overall error dependence in the current paper (Definition \mainref{def:error-dependence}):
Let $\mathcal{N}^\partial(i, s)$ be the subset of unit indices for which the shortest path between unit $i$ and units $j \in \mathcal{N}^\partial(i, s)$ is exactly $s$. Leung2022Causal calls this the $s$-neighborhood boundary of $i$. Note that $\mathcal{N}^\partial(i, s)$ for $s \in \braces{1, 2, \dotsc, n}$ partitions the unit indices with the exception of $i$. (The index $i$ is excluded because $\mathcal{N}^\partial(i, 0) = \braces{i}$.) Using this partition, we can write the overall error dependence as
Similar to the current paper, Leung2022Causal considers generic exposure mappings. However, unlike the current paper, it is critical in Leung2022Causal that the exposure mappings are local to the observed graph. This is formalized in Assumption 1 in Leung2022Causal, which states that unit $i$'s exposure can only depend on the treatment assignments in a $K$-hop neighborhood around $i$ in the graph, where $K$ is a constant in the asymptotic sequence. We will break up the summation in the overall error dependence into two parts, the first over $1, \dotsc, 2K$ and the second over $2K + 1, \dotsc, n$:
Consider the first term in expression (ref). Because the potential outcomes are bounded, which is an assumption made both in the current paper and in Leung2022Causal, we can apply Lemma (ref) of the current paper, which states that $\E{\varepsilon_{i}\varepsilon_{j} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = d} \leq 4k_1^2$. This allows us to write
where $M^\partial(s) = n^{-1} \sum_{i=1}^n \abs{\mathcal{N}^\partial(i, s)}$ is the average size of the $s$-neighborhood boundary. The quantity $M^\partial(s)$ is used by Leung2022Causal, in Assumption 5 and elsewhere.
Next, consider the second term in expression (ref). Recall the definition of the specification errors (Definition \mainref{def:specification-error}): $\varepsilon_{i} = \poi{\boldsymbol{Z}} - \poexi{D_{i}}$. Let
so that the specification error can be written $\varepsilon_{i} = A_i(s) + B_i(s)$ for any $s$. For any $i$ and any $j \in \mathcal{N}^\partial(i, s)$ such that $s \geq 2K + 1$, we can write
where
Note that $B_i(\floor{s/2})$ depends on $\boldsymbol{Z}^{(i,\floor{s/2})}$ and $D_{i}$. When $s \geq 2K + 1$, both of these random variables only depend on assignments in $\boldsymbol{Z}$ within a $\floor{s/2}$-hop neighborhood around $i$. Furthermore, if $j \in \mathcal{N}^\partial(i, s)$, then the $\floor{s/2}$-hop neighborhood around $i$ is disjoint from the $\floor{s/2}$-hop neighborhood around $j$. The implication is that $B_i(\floor{s/2})$ and $B_j(\floor{s/2})$ are independent, because the design is Bernoulli. This means that we can write
These two factors are both zero, because
We therefore have
Because Leung2022Causal assumes that the exposures are defined within a $K$-hop neighborhood and because he assumes that the design is Bernoulli, two units that are at a geodesic distance of at least $2K + 1$ will have independent exposures. This tells us that, for any $i$ and any $j \in \mathcal{N}^\partial(i, s)$ such that $s \geq 2K + 1$,
where the probabilities on the right-hand side are bounded away from zero by the positivity assumption made in both the current paper and in Leung2022Causal. We can therefore write
where the bound $1 / \Pr{D_{i} = d} \leq k_2$ was used.
Using Hölder's inequality, we can write
where $\operatorname*{ess\,sup} X$ denotes the essential supremum of the random variable $X$. Note that $\Es[\big]{\abs{A_i(\floor{s/2})}}$ is equal to $\theta_{i,s}$, as defined in the beginning of this section. Also note that bounded potential outcomes implies that both $\operatorname*{ess\,sup} \abs{\varepsilon_{j}}$ and $\operatorname*{ess\,sup} \abs{B_i(\floor{s/2})}$ are bounded by $2k_1$. Let $\theta_{s} = \max_{i} \theta_{i,s}$, which corresponds to the definition in the beginning of Section 3 in Leung2022Causal. This allows us to write
Taken together, we have
for any $i$ and any $j \in \mathcal{N}^\partial(i, s)$ as long as $s \geq 2K + 1$. In turn, this gives us that the second term in expression (ref) is bounded as
where, as above, $M^\partial(s) = n^{-1} \sum_{i=1}^n \abs{\mathcal{N}^\partial(i, s)}$.
Combining the results for the two terms in expression (ref), we have
To connect the bound to the assumptions in Leung2022Causal, it will prove convenient to make it less tight by harmonizing the constants for the two terms and to extend the summation to include the case $s = 0$:
Next, define $\tilde{\theta}_s$ as
Note that this corresponds exactly to the definition of the same symbol in Theorem 1 in Leung2022Causal. This allows us to write
Now, Assumption 5 in Leung2022Causal is that
Given the bound we have just derived, this implies that $\bar{\varepsilon}_{d} = \littleO{1}$, which in turn implies that Condition \mainref{cond:limited-error-dependence} in the current paper holds. In other words, in the context studied by Leung2022Causal, the assumptions he makes imply the interference condition used in the current paper. In this sense, the conditions in the current paper are weaker and more general than those in Leung2022Causal. As noted above, this is not meant as a critique; instead, it highlights that the scopes of the results are different. Note that Leung2022Causal uses Assumption 6, which is a stronger version of Assumption 5, to prove asymptotic normality. Hence, he uses even stronger assumptions than the conditions in the current paper to derive this result, similar to how stronger assumptions are used in Section (ref) of the current paper to show asymptotic normality.
This section provides concrete, numerical examples to illustrate the specification errors and their connection to the precision of the estimators.
The basic setup is the same in all examples. There is an observed network through which units are hypothesized to interact. This network is used to define the exposures. Denote an edge from unit $i$ to unit $j$ in the network as $g_{ij} = g_{ji} = 1$, and the absence of an edge is denoted as $g_{ij} = 0$. For simplicity, the graph is a cycle graph, meaning that $g_{ij} = 1$ if and only if $\abs{i - j} \in \braces{1, n - 1}$. That is, unit $i$ is connected to units $i - 1$ and $i + 1$, with the exception of units $1$ and $n$, who are connected to each other. The qualitative results of this simulation do not depend on the use of a cycle graph; it is used purely for convenience.
Each unit is assigned a binary treatment independently at random with equal probability: $\Pr{Z_{i} = 1} = 1/2$ for all $i$. That is, the experimental design is Bernoulli. There are two exposures of interest. The first exposure, denoted $1$, is when the ego is untreated and both neighbors in the cycle graph are treated. The second exposure, denoted $0$, is when the ego as well as both neighbors are untreated. All other treatment vectors are mapped to a residual exposure, denoted $2$, which will not be used in the analysis. That is, the exposure mapping for unit $i$ is
Recall that $D_{i} = \exmi{\boldsymbol{Z}}$ denotes the realized exposure for unit $i$. Because of the Bernoulli design and cycle graph structure, the exposure probabilities are $12.5\%$ for both exposures and for all units:
The treatment effect of interest is
which is the expected indirect effect of having two treated neighbors when being untreated compared to having no treated neighbors. Because of its improved small sample performance, the main focus of the simulation study will be the H{\'a}jek estimator, which in this case coincide with the difference-in-means estimator. The results for the Horvitz--Thompson estimator are reported in Table (ref) at the end of the section on page (ref). The simulation consists of 50,000 rounds (for each setting and sample size), ensuring that the Monte Carlo errors are negligible, and no uncertain measures for Monte Carlo imprecision are reported.
The interference is local in the first example. This largely mirrors the setting studied by Saevje2021Average, and the discussion in Section (ref) tells us that the estimators should be consistent here. Two version of the data generating process will be considered.
In the first version, labelled “Correct,” the exposures are correctly specified with respect to the two exposures of interest. In particular, the potential outcome function is
where $\alpha_i$ is drawn uniformly at random from $[15, 25]$, $\beta_i$ is drawn from $[1, 10]$ and $\gamma_i$ is drawn from $[1, 2]$. The coefficients are drawn once and held fixed over the simulation rounds, mirroring the fact that all randomness under consideration comes from treatment assignment. Here, we have $\poi{\boldsymbol{z}} = \poi{\boldsymbol{z}'}$ whenever $\exmi{\boldsymbol{z}} = \exmi{\boldsymbol{z}'} \in \braces{0, 1}$, meaning that the exposures of interest, $d \in \braces{0, 1}$, are correctly specified. That is, the overall error dependence in Definition \mainref{def:error-dependence} is zero.
We can interpret this data generating process to describe a vaccination study. The treatment is whether a unit has received the vaccine, and the outcome is some measure of viral load. The coefficient $\alpha_i$ captures $i$'s baseline exposure and susceptibility to the virus, $\beta_i$ captures the direct effect of taking the vaccine oneself, and $\gamma_i$ captures the indirect effect of having vaccinated neighbors. Of course, this is a highly stylized data generating process that is unlikely to capture the dynamics of virus transmission in reality, but the process fills its role as an illustration.
The second version, labelled “Local,” has unmodeled interference, meaning that the exposures are not correctly specified. An additional network is added here, which is directed and unobserved by the investigator. The unobserved network is generated by randomly and independently drawing an arc from $i$ to $j$ with probability $n^{-4/5}$. Let $u_{ij} \in \setb{0, 1}$ denote whether there is an arc from $i$ to $j$ in this network. If an unobserved arc overlaps with an edge in the observed network, the unobserved arc is deleted, ensuring that $g_{ij}u_{ij}=0$. The unobserved network is generated once and held fixed over the simulation rounds. Because the observed network is completely unrelated to the unobserved network, approximate neighborhood interference, as defined in Leung2022Causal and discussed in Section (ref) above, does not hold here.
The potential outcome functions under the second version is
That is, vaccinated neighbors have the same effect in the observed and unobserved networks. The coefficients used in the first version of the data generating process are used here as well. We can interpret the second data generating process as also describing a vaccination study, but with unobserved transmission paths.
Note that the exposures are not correctly specified under the second version of the process. We have that the expected potential outcomes for unit $i$ under the two exposures of interest, $d \in \setb{0, 1}$, are
meaning that the specification error for unit $i$ conditional on $D_{i} \in \setb{0, 1}$ is
However, because the unobserved network is sparse, the overall error dependence in Definition \mainref{def:error-dependence} will diminish as the sample size grows. Indeed, the overall error dependence will be of order $n^{-16/25}$ in this setting, ensuring that Condition \mainref{cond:limited-error-dependence} is satisfied. The rate of the overall error dependence can be derived either by direct calculation or by using the approach described in Section (ref) above.
The results are presented in Figure (ref). Panel A presents the root of the sum of the magnitude of the overall error dependence measures defined in the paper for the two exposures, $\sqrt{\abs{\bar{\varepsilon}_{1}} + \abs{\bar{\varepsilon}_{0}}}$. For the first data generating process (“Correct”), the error dependence is constant at zero, indicating that the exposures indeed are correctly specified under this process. The measure is not zero under the second data generating process (“Local”). However, the error dependence decreases at a reasonably fast rate, indicating that Condition \mainref{cond:limited-error-dependence} indeed is satisfied, and that we can expect the estimators to be precise in large samples.
Panel B presents the root mean square errors of under the two versions of the data generating process. The estimator is somewhat less precise under the second version for all sample sizes, but the mean square error approaches zero at a reasonable rate under both versions. This is what we would expect given Proposition \mainref{prop:variance-bound} in the main paper and the diminishing error dependence shown in Panel A.
The second example introduces global interference. That is, all units can potentially interfere with all other units. The example builds on the first data generating process in the previous section by adding behavior akin to herd immunity. If the number of vaccinated units passes some threshold, there is no transmission of the virus in the community under study, and the outcome (e.g., viral load) is zero for all units. This is again a highly stylized example, but it suffices to illustrate the key insights.
The potential outcome function is here
where the same coefficients as in the previous data generating processes are used, and $\phi \in (0, 1)$. This process captures behavior reminiscent of herd immunity because if a share of $\phi$ units are treated, then the outcome is zero for all units.
Three versions of this data generating process will be considered, corresponding to three values of $\phi$. The first version, labelled “Rarely,” uses $\phi = 0.52$, meaning that herd immunity occurs when at least $52\%$ of the units are treated. The second version, labelled “Infrequently,” uses $\phi = 0.51$, and the third version, labelled “Frequently,” uses $\phi = 0.5$.
Note that $\sum_{j=1}^n z_{j}$ follows a binomial distribution of $n$ trials with $0.5$ success probability. Hence, herd immunity will occur somewhat rarely under the first version of the data generating process ($\phi = 0.52$), and it grows increasingly rare as the sample size grows. Herd immunity is more common, but still fairly infrequent, when $\phi = 0.51$. However, when $\phi = 0.5$, herd immunity will be frequent also in large samples, because the distribution of $\sum_{j=1}^n z_{j}$ is centered around $n / 2$. Indeed, when $n = 100$, herd immunity occurs with roughly the same probability under all three versions of the process. But when $n = 10,000$, herd immunity occurs with a probability of only $0.003\%$ under the first version ($\phi = 0.52$), a probability of $2.33\%$ under the second version ($\phi = 0.51$), and a probability of $50.4\%$ under the third version ($\phi = 0.50$).
Interference is global under all three versions, because there are situations in which changing a single unit's treatment assignment changes the outcome of all other units under all three versions. That is, the event $\sum_{i=1}^n Z_{i} < \phi n \leq 1 + \sum_{i=1}^n Z_{i}$ has a non-zero probability of occurring under all three versions, and changing any untreated unit to be treated in this case induces herd immunity, changing the outcomes of all other units. However, such global interference takes place only very rarely under the first version when the sample is large, because changing a single unit's treatment assignment under most assignments will not have global effects. In this sense, global interference does exist under all three versions, but it is not practically relevant for the first two versions of the data generating process, $\phi \in \braces{0.51, 0.52}$, if the sample is sufficiently large, because it will happen with a probability so low that it can be ignored. This is not the case when $\phi = 0.5$, because there will be non-negligible variation in whether herd immunity occurs no matter how large the sample gets.
To see this formally, let $H = \indicator[\big]{\sum_{i=1}^n Z_{i} \geq \phi n}$ be an indicator whether herd immunity occurs, and let $\pi_h = \E{H \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} Z_{1} = Z_{2} = Z_{3} = 0}$ be the probability of that event conditional on that three units are untreated. The expected potential outcomes for unit $i$ under exposure $d = 0$ is
and the specification error for unit $i$ conditional on $D_{i} = 0$ is
This implies that the overall error dependence will be of the same order as $\Var{H}$. That is, if there is variability in whether herd immunity occurs also in large samples, then the overall error dependence will not approach zero. The variance will be larger when $\phi = 0.51$ compared to $\phi = 0.52$, but we have $\Var{H} \to 0$ in both cases. However, when $\phi = 0.5$, $\Var{H}$ is asymptotically bounded away from zero, meaning that the overall error dependence does not approach zero in this setting. Hence, given Proposition \mainref{prop:variance-bound}, we expect convergence of the estimator when $\phi = 0.52$ and $\phi = 0.51$, but possibly not when $\phi = 0.5$.
The results are presented in Figure (ref). As in the previous figure, Panel A presents the sum of the magnitude of the overall error dependence measures for the two exposures. The error dependence increases initially, but as the sample grows in size, the error dependence decreases for both “Rarely” ($\phi = 0.52$) and “Infrequently” ($\phi = 0.51$). This mirrors the fact that variation in whether herd immunity occurs decreases as the sample size grows under both of these versions. However, the error dependence does not decrease for the version “Frequently,” instead flattening out slightly below the 1.25 mark. Therefore, Condition \mainref{cond:limited-error-dependence} does not hold, and the consistency result of the current paper does not apply.
Panel B presents the root mean square errors of under the three version of the data generating process. The precision of the estimator initially improves for all three versions. The precision continues to improve for the first two versions (“Rarely” and “Infrequently”), and the mean square error is very close to zero when $n=100,000$. However, under the third version (“Frequently”), the mean square error flattens out at around the $0.5$ mark, indicating that the estimator is not consistent. This mirrors what we learned about the behavior of the error dependence measure in Panel A.
The reason the precision of the estimator initially improves also under the third version, despite not seeing a corresponding decrease in the error dependence, is that two first terms of the variance bound in Proposition \mainref{prop:variance-bound} approach zero as the sample size grows even if the error dependence is large.
The third example considers interference transmitted through a market price. There are $n$ companies producing some good. The treatment under study is an improved production technology. There are spillovers in the production technology, in the sense that a company's production is improved if other neighboring companies have access to the improved technology. This could, for example, arise because technology-specific inputs (e.g., skilled labor) become more accessible in the local market if many companies use the same technology, because of technology transmission, or because of some other cluster effect in production.
Given market price $p$ and treatments $\boldsymbol{z}$, the production function for company $i$ is
That is, in absence of treatment, $\boldsymbol{z} = \boldsymbol{0}$, the production function is linear. If the company itself or any of its neighbors have access to the improved technology, the production function adds a quadratic component, making the production more efficient. The parameter $\gamma_i$ captures the size of the company, scaling the production function so that larger companies produce more goods, all else equal. The size of the supply side of the market is normalized to $100$, in the sense that $\sum_{i=1}^n \gamma_i = 100$.
The supply curve for the market is given by
where
The demand curve is
Given treatments $\boldsymbol{z}$, the equilibrium price $p^*(\boldsymbol{z})$ is the one that clears the market:
Note that when no company is treated, $\boldsymbol{z} = \boldsymbol{0}$, the equilibrium price is one, $p^*(\boldsymbol{z}) = 1$, because the market then clears when $100 p = 100 / p$. This was the price at baseline, before the experiment was run. If one or more companies are treated, the price will be lower than one. The exact price depends on both the number of treated companies and which companies are treated. As more large companies are treated (or large companies have treated neighbors), the lower the equilibrium price will be. Changing any company's treatment will always change the equilibrium price at least slightly.
Given treatments $\boldsymbol{z}$, the production in company $i$ is
Note that treatment has two effects. Holding the market price fixed, companies with access to the improved technology will have an increased production. But as more companies have access to the technology, the market price will decrease, leading to decreased production. The overall production will increase as more companies have access to the technology, but some of the companies might still decrease their production.
Company $i$'s baseline production (when $\boldsymbol{z} = \boldsymbol{0}$) was
The outcome of interest is the relative increase in production compared to baseline:
Unlike the example in Section (ref), the interference is global here. That is, changing a single unit's treatment could change the outcome of all other units. Unlike the example with herd immunity in Section (ref), which also considered global interference, it is common that interference occurs globally. That is, changing a single unit's treatment will always change the outcome of all other units. However, it is still possible to achieve consistency here, because the interference could be limited. In this case, all global interference is mediated through the market price. If the market price stabilizes in large samples, in the sense that it is close to some value with high probability, then the interference between most units will be negligible. This is because the production functions are smooth in the market price, so minor price disturbances will be inconsequential for production.
Three versions of this data generating process will be considered, differing in the distribution of the companies' sizes. In the first version, labelled “All small,” all companies are reasonably small: $\gamma_i$ for unit $i$ is proportional to the $i / (n + 1)$-th percentile of the standard log-normal distribution. When $n=1000$, the largest company has $\gamma_i = 1.35$, corresponding to a $1.35\%$ share of the total market at baseline. The five largest companies have a $5.07\%$ share of the market at baseline when $n=1000$. The market share of the top five companies will approach zero as the sample size grows.
In the second version, labelled “Some outliers,” most companies are small but there are some large outliers. The coefficients $\gamma_i$ are here proportional to the $i / (n + 1)$-th percentile of the log-normal distribution with mean parameter zero and standard deviation parameter three. When $n=1000$, the largest company has $\gamma_i = 19.6$, corresponding to a $19.6\%$ share of the total market at baseline. The five largest companies have a $46.3\%$ share of the market at baseline. The market share of the top five companies will approach zero as the sample size grows also in this case, but at a slower rate than in the first version.
In the third version, labelled “One large,” most companies are small but there is one company that has half of the market: $\gamma_i$ is proportional to the percentiles of the standard log-normal distribution, as in the first version, except for the first company, which instead has $\gamma_1 = 50$. For all sample sizes, the largest company has $\gamma_i = 50$, meaning a $50\%$ share of the total market at baseline. When $n=1000$, the five largest companies have a $52.1\%$ baseline share of the market. The market share of the top five companies will not approach zero as the sample size grows in this case.
The market price will stabilize in large markets with many small companies. Treatment assignment will introduce variability for individual companies, but the aggregated supply curve will be similar (but never identical) under most assignments. However, when there is one or a few very large companies in the market, the treatment assigned to those companies, or their neighbors, will have a large effect of the aggregated supply curve, and thus the market price. This is true even if the market as a whole is large. Hence, we expect the market price to stabilize under the first two versions of the data generating process (“All small” and “Some outliers”), but the market price might stabilize slowly when there are outliers as in the second version. The price will not stabilize under the third version (“One large”). The treatments assigned to the one large company and its neighbors will have a large effect on the market price, even if the market is very large.
To see this formally, consider the expected potential outcomes for unit $i$ under exposure $d = 0$:
which is the expected market price conditional on that unit $i$ has exposure $d = 0$. For ease of exposition, assume that $i$ is a small company with neighbors that also are small. This means that the conditioning event $D_{i} = 0$ is inconsequential for the distribution of the market price in large samples, and we can approximate the conditional expectation with the unconditional expected market price $\Et[\big]{p^*(\boldsymbol{Z})}$. That is, we have $\poexi{0} \approx \Et[\big]{p^*(\boldsymbol{Z})}$ with a negligible approximation error in large samples. Note that most companies will be small with small neighbors, even if there is one or a few large companies, so this case covers the almost all units in all three versions of this data generating process.
The specification error for unit $i$ conditional on $D_{i} = 0$ then becomes
This means that the error dependence between two small companies $i$ and $j$ approximately is
Like above, the approximation here uses the fact that the conditioning event $D_{i} = D_{j} = 0$ is inconsequential when companies $i$ and $j$, and their neighbors, are small. The number of companies that are large or have large neighbors will be a diminishing share of all companies under all three versions of the data generating process. The outcome is bounded for all companies, including large ones, meaning that $\E{\varepsilon_{i}\varepsilon_{j} \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = D_{j} = 0}$ is bounded for all units (see Lemma (ref)). This means the error dependence involving large companies can be ignored in large samples for all three versions, and we have
We need that $\Vars[\big]{ p^*(\boldsymbol{Z}) } \to 0$ for limited specification error dependence (Condition \mainref{cond:limited-error-dependence}) to hold. The variance of the market price will approach zero under the first two versions of the data generating process (“All small” and “Some outliers”), by the law of large numbers. However, the experimental design will induce variability in the price under the third version (“One large”), no matter the sample size. That is, $\Vars[\big]{ p^*(\boldsymbol{Z}) }$ is asymptotically bounded away from zero.
The results are presented in Figure (ref). Panel A shows that the error dependence diminishes quickly under the first version (“All small”). It also diminishes under the second version (“Some outliers”), but at a slower rate. The error dependence appears to be unrelated to the sample size under the third version (“One large”), indicating approximately the same level of market price variability for all sample sizes in this case. Condition \mainref{cond:limited-error-dependence} does not hold under the third version, and the consistency result in the current paper do not apply.
Panel B largely mirrors the results in the first panel. The mean square error diminishes under the first two versions of the data generating process (“All small” and “Some outliers”), and the rate is faster under the first version. While the precision initially improves slightly under the third version, it quickly flattens out, and there appears to be no improvements after about $n = 1,000$. The estimator does not appear to be consistent under the third version.
Saevje2021Average investigate estimation of average treatment effects under interference. They distinguish between two types of effects. The effect of primary interest in Saevje2021Average is the expected average treatment effect, or EATE, defined as
where $\poi{z; \boldsymbol{z}_{-i}}$ denotes the potential outcome for unit $i$ when $i$'s treatment assignment is $z$ and the assignment of all other units is $\boldsymbol{z}_{-i}$. This estimand captures the effect of changing a unit's own treatment while holding the treatments of all other units fixed, averaged over all units and marginalized over the experimental design. In this sense, it captures the “direct” treatment effect.
Saevje2021Average compare the EATE estimand with an estimand described by Hudgens2008. This alternative estimand is often referred to as the direct treatment effect, but Saevje2021Average refer to it as the average distributional shift effect (ADSE) to differentiate it with EATE. The ADSE estimand is defined as
The difference is that EATE marginalizes $\poi{z; \boldsymbol{Z}_{-i}}$ over the unconditional distribution of $\boldsymbol{Z}_{-i}$, while ADSE marginalizes $\poi{z; \boldsymbol{Z}_{-i}}$ over the conditional distribution of $\boldsymbol{Z}_{-i}$ given $Z_{i} = z$. Because the two conditional distributions, corresponding to $Z_{i} = 1$ and $Z_{i} = 0$, potentially are different, the ADSE estimand can capture both the direct of effect of $Z_{i}$ on the outcome of unit $i$ and effects due to the distributional shift of the marginalization.
For example, consider a sample with two units, for which $\poi{z_i; z_j} = z_j$ for both units. That is, a unit's outcome does not depend on its own treatment but it does depend on the other unit's treatment. The unit-level direct treatment effects are all zero here:
This means that the EATE estimand is zero, no matter the experimental design. However, the ADSE will not always be zero, because it could capture a distributional shift. For example, if the experimental design is such that $Z_{1} = 1 - Z_{2}$, so the two units have opposite treatments, the ADSE estimands is
The expected exposure effect, as defined in the current paper, is
As this estimand is marginalizing over conditional distributions, it is closer to ADSE than to EATE. Indeed, with the exposure mapping $\exmi{\boldsymbol{z}} = z_{i}$, expected exposure effect is the same as ADSE.
A relevant question in this setting is whether it is possible to define an estimand corresponding to EATE for exposure effects. The answer is that, in most cases, it will not be possible. The issue with an EATE-type exposure effect is that units' exposures typically cannot be independently manipulated. This is an inherent limitation due to the definition of the exposures, rather than a consequence of the experimental design.
To see this, note that in Saevje2021Average, the EATE estimand asks “What is the (expected, average) effect of changing a unit's treatment, holding all other units' treatments fixed?” This question makes sense because it is generally possible to change a unit's treatment without changing other units' treatment, even if that configuration of treatments is not in the support of the experimental design. An EATE-type exposure effect would ask “What is the (expected, average) effect of changing a unit's exposure, holding all other units' exposures (or treatments) fixed?” The issue is that we can essentially never change a unit's exposure while holding all other units' exposures/treatments fixed.
To illustrate this, consider a network setting with two exposures of interest: when no neighbors are treated ($D_i = 0$) and when at least one neighbor is treated ($D_i = 1$). The effect of interest is the contrast of the potential outcomes produced by these two exposures. Now, consider a group of three units ($A$, $B$, $C$), where there are two edges, $(A, B)$ and $(B, C)$, and there are no other units. That is, the network is $A \leftrightarrow B \leftrightarrow C$. To define an EATE-type exposure effect in this setting, we would need to compare the situation when unit $A$ is assigned exposure $D_A = 0$ to the situation when it is assigned $D_A = 1$, holding the exposures for $B$ and $C$ fixed. When $D_A = 0$, we know that unit $B$ is assigned to the control treatment (because $A$ does not have any treated neighbors), which means that unit $C$ also must be assigned to $D_C = 0$ (because it doesn't have any treated neighbors either). We would now want to consider a setting where we change unit $A$'s exposure while holding the exposures of all other units the same. But this is not possible. The only way to change unit $A$'s exposure to $D_A = 1$ is to change unit $B$'s treatment assignment from control to active treatment, ensuring that unit $A$ has a treated neighbor. As an inescapable consequence, unit $C$ will then also have a treated neighbor, so it will also switch to exposure $D_C = 1$. In this setting, it is inherently impossible to manipulate the units exposures independently.
The problem in this example is that we cannot separate the effect of having a neighbor that is treated from the effect of having a neighbor that in turn has a neighbor that is treated. This is similar to the problem that arises for the standard ADSE estimand, but the problem is more fundamental here. In the setting considered by Saevje2021Average, where the implicit exposure mapping is $\exmi{\boldsymbol{z}} = z_{i}$, we can almost always independently manipulate the treatments, at least in principle. That is, we can imagine changing the treatment assigned to one unit without changing any other units' treatment (even if that type of change is not realized by the experimental design). The problem is therefore entirely driven by dependence in treatment assignments introduced by the experimental design. Under a Bernoulli design, the ADSE and EATE estimands coincide when the exposure mapping is $\exmi{\boldsymbol{z}} = z_{i}$. The problem with exposure effects more generally is due to dependencies introduced by the nature of the exposures; the exposures are by construction not independently manipulable. In the example above, there is no way to assign the treatments (i.e., there exists no $\boldsymbol{z}$) so that unit $A$ has a treated neighbor while unit $C$ does not; units $A$ and $C$ have partially linked exposures by construction.
It is possible to define EATE-type exposure effects for specific types of exposures. For example, an EATE-type estimand can be defined if the exposures at least in principle are independently manipulable. One such case is when $\exmi{\boldsymbol{z}} = z_{i}$. More generally, if no two exposures depend on the same treatment assignment $z_{i}$, they are independently manipulable. An example when this is the case is when $\exmi{\boldsymbol{z}} = z_{\rho(i)}$ for some permutation $\rho(i)$ of the unit indices. Another example is if the sample is partitioned into clusters and one “focal” unit is selected within each cluster so that its exposure only depends on the treatments within its own cluster. If the analysis is restricted to only the focal units, their exposures will be independently manipulable. Another possibility is to extend treatment to include interventions to the interference structure itself, which could be used to isolate units from spillover effects. For example, the question above could be reformulated to ask what the effect is of having at least one treated neighbor versus being completely isolated (i.e., having no neighbors at all).
Slightly more generally, an EATE-type effect can often be defined if the exposure mapping is such that the exposure is given by the treatment assignments of a subset of the units. Specifically, let $\mathbf{m}_i \in \setb{0, 1}^n$ be a binary vector of length $n$ acting as a masking vector, meaning that for some $\boldsymbol{z} \in \setb{0, 1}^n$, the $j$th coordinate of the Hadamard product $\boldsymbol{z} \odot \mathbf{m}_i$ is equal to $z_{j}$ if and only if $m_i = 1$. If the exposure mappings are such that $\exmi{\boldsymbol{z}} = a$ if and only if $\boldsymbol{z} \odot \mathbf{m}_i = \mathbf{r}_{i,a}$ and $\exmi{\boldsymbol{z}} = b$ if and only if $\boldsymbol{z} \odot \mathbf{m}_i = \mathbf{r}_{i,b}$, where $\mathbf{m}_i$ is common for both exposures, then an EATE-type exposure effect for exposures $a$ and $b$ can be defined. This is done by fixing the coordinates of $\boldsymbol{Z}$ corresponding to coordinates with value one in $\mathbf{m}_i$ so that $\boldsymbol{Z} \odot \mathbf{m}_i$ is equal to either $\mathbf{r}_{i,a}$ or $\mathbf{r}_{i,b}$. The coordinates of $\boldsymbol{Z}$ corresponding to coordinates with value zero in $\mathbf{m}_i$ are marginalized over using their marginal distributions given the experimental design, without regard for other coordinates of $\boldsymbol{Z}$. This estimand would avoid the distributional shift artifact that affects the expected exposure effect. It is possible to slightly extend this effect to setting where $\exmi{\boldsymbol{z}} = a$ if and only if $\boldsymbol{z} \odot \mathbf{m}_i \in \mathcal{R}_{i,a}$ for some set $\mathcal{R}_{i,a} \subset \setb{0, 1}^n$. This would be done by repeating the exercise above for each element in $\mathcal{R}_{i,a}$, and then marginalize of these elements using an approriate distribution. Precise estimation of these EATE-type exposure effects would require additional conditions in line with those in Saevje2021Average; it is beyond the scope of this paper to describe and investigation those conditions.
The Horvitz--Thompson estimator is rarely a good choice in practice because of its instability in small samples. The analysis of the Horvitz--Thompson estimator serves as a foundation on which we can build understanding about the behavior of other estimators. This section investigates the behavior of common refinements of the Horvitz--Thompson estimator, which experimenters often prefer over the original estimator.
The first refinement accounts for the realized number of units assigned to the exposures of interest. The H{\'a}jek estimator Hajek1971 does this by dividing each term in the estimator with the sum of the reciprocals of the assignment probabilities for the units assigned to the exposure, rather than dividing by $n$. The change can absorb some of the variability in the estimator introduced by randomness in the number of units assigned to each exposure. The ratio structure introduces bias, but the bias is generally small enough to still grant improvements in mean square error sense. The denominator can generally be shown to be well-behaved, so the estimator's limited behavior can be linked to the Horvitz--Thompson estimator through linearization.
Experimenters often use estimators that implicitly adjust for the assignment probabilities. One such example is the difference-in-means estimator. This estimator can be shown to coincide with the H{\'a}jek estimator whenever the assignment probabilities are the same for all units: $\prei{d} = \prej{d}$ for all $i,j\in\mathcal{U}$. Proposition (ref) thus implies that the difference-in-means estimator can be used in similar situations also under misspecification. Experimenters should, however, not blindly use the difference-in-means estimator for exposure effects because the exposure mappings may not induce equal assignment probabilities on the exposures even if they are equal for the nominal treatments. Another estimator coinciding with the H{\'a}jek estimator is the ordinary least squares (OLS) estimator. The unweighted version requires equal assignment probabilities just like the difference-in-means estimator, but a weighted OLS estimator is equivalent to the H{\'a}jek estimator also with unequal assignment probabilities.
A disadvantage of both the Horvitz--Thompson and H{\'a}jek estimators is their inability to take advantage of auxiliary information. A modification of the Horvitz--Thompson estimator allows us to incorporate such information. The idea is that information beside the observed potential outcomes themselves might allow us to predict the potential outcomes we do not observe. If this prediction is sufficiently good, the predicted outcomes can be used to offset chance imbalances introduced by the randomization. Saerndal1992Model call it the difference estimator in a sampling setting, and the name will be used here as well.
The definition of the estimator reveals the idea that motivate its use. The first term is simply the average difference in predicted potential outcomes. If the predictions are of high quality, this term will be an accurate estimator of the exposure effect. The issue is that the predictions may have systematic errors. The second term is included to ensure unbiasedness. If the predictions are of low quality, this term will compensate for the errors in the first term, and it ensures that the estimator performs well in expectation. The estimator bears a resemblance in this regard to the class of doubly robust estimators used in observational studies when the assignment mechanism is unknown Robins2001Comment.
The properties of the difference estimator depend on the way the predictions are constructed. In particular, the estimator can be shown to retain the advantageous properties of the Horvitz--Thompson estimator if the predictions are external to the study. External here means that they do not depend on the treatment assignment. As the only randomness under consideration stems from the assignment mechanism, independence between $\poexesti{d}$ and $\boldsymbol{Z}$ implies that the predictions are non-random. The probability space can extended to accommodate random predictions if one wants to account for the consequences of external variability. Such variability could affect the rate of convergence if predictions are sufficiently dependent between units, but it is otherwise inconsequential to the results.
An alternative to focusing on the dependence between predictions is to consider their convergence. In particular, the average prediction dependence can be bounded by
which shows that Condition (ref) is satisfied if the predictions on average converges in mean square.
The difference estimator seemingly provides advantages at no cost. Good predictions of the potential outcomes confer improvements in finite samples, but the estimator has the same behavior as the Horvitz--Thompson estimator in large samples. The no-cost advantages are superficial. The mean square error may increase when the predictions are poor, so investigators should use the difference estimator only when the predictions are expected to be of reasonably high quality.
However, the quality of the predictions is less of a concern than their construction. Covariate information can be used to make the predictions, but the assigned exposures and the observed outcomes can generally not be used because it would induce dependence between the predictions and $\boldsymbol{Z}$. More precisely, if $\boldsymbol{x}_{i}$ denotes a vector of covariates describing characteristics of unit $i$, we can form the predictions as $\poexesti{d} = f\paren{d, \boldsymbol{x}_{i}}$ for some function $f$. The function $f$ can, however, not be constructed using $\paren{Y_{1}, Y_{2}, \dotsc, Y_{n}}$ or $\paren{D_{1}, D_{2}, \dotsc, D_{n}}$. This illustrates that the construction of $f$ truly needs to be external to treatment assignment when used for the predictions in the difference estimator. This severely limits its applicability. Split-sample or leave-one-out approaches Williams1961Generating that often are used to solve the issue cannot be used here because the misspecification may induce dependence between subsamples that otherwise appear isolated.
An estimator facilitating dependence between the predictions of the potential outcomes and the treatment assignments is inspired by the generalized regression estimator commonly used in the sampling literature. The estimator has received recent attention in the causal inference literature as well Lin2013Agnostic,Middleton2018Unified.
The estimator uses a linear working model for the relationship between the potential outcomes and the covariates. The working model is used to construct the predictions. Generally, $\poexesti{d} = \boldsymbol{x}_{i}^{\mathpalette\@tran{}} \coef{d}$ for some vector of coefficients $\coef{d}$ indexed by $d \in \Delta$, so different coefficients are used for different exposures. No assumptions are made about the validity of the model, but the quality of the predictions are related to how well the model can approximate the potential outcomes. It remains to pick the coefficients $\coef{d}$. The generalized regression estimator allows for dependence between the coefficients and the treatment assignments, so the coefficients can be estimated in the sample. For example, we may pick them as the minimizing solution to $\sum_{i=1}^n D_{id}\bracket{Y_{i} - \boldsymbol{x}_{i}^{\mathpalette\@tran{}} \coef{d}}^2$ as is often done in applications. But other choices exist, and the estimator is largely agnostic about how the coefficients were constructed.
The conventional approach to investigate the properties of the generalized regression estimator is to assume that the vector of coefficients constructed in the sample convergences to some fixed vector asymptotically. This ensures that the dependence between units' predictions is small in large samples, which provides consistency. The assumption can be weaken to only require that the magnitude of the vector of coefficients is asymptotically bounded, thereby bypassing the need of assuming a well-defined limit.
Positivity conditions are often seen as innocuous in experiments because the experimenter controls the design and can ensure their validity. But this is rarely the case when estimating exposure effects. Exposure mappings tend to be complex, and it may not be feasible to construct a design that would induce a desired distribution over the exposures. Experimenters will instead settle for heuristic choices for the design at the treatment level, and this could induce violations of Condition \mainref{cond:positivity}.
The positivity condition can fail in two ways. The first is when it is fundamentally impossible for a unit to be assigned a particular exposure. For example, in the experiment in Bogot\'a described in Section \mainref{sec:illustration} in the main paper, a non-hot spot street without any neighboring hot spot streets cannot be assigned to an exposure requiring that at least one neighboring hot spot street receives intensive policing. This may be formalized by saying that there is an exposure $d \in \Delta$ for which no $\boldsymbol{z} \in \mathcal{Z}$ exists with $\exmi{\boldsymbol{z}} = d$.
The consequences of such a failure are more than just statistical. If it is nonsensical to talk about some collection of units being assigned to a certain exposure, it is nonsensical to consider exposure effects that include those units in its average. Unless the experimenter is comfortable stipulating a metaphysical model allowing extrapolation to unrealizable potential outcomes, the only solution is to exclude such units from the average. The result may be that the number of units included in the analysis is fewer than the length of $\boldsymbol{z}$, but this is not an issue other than for efficiency. In the following discussion, it will be assumed that such exclusions have been made if necessary. That is, if the aim is to estimate the effect of exposures $a$ and $b$, then $\setb{a, b} \subseteq \setb{\exmi{\boldsymbol{z}} : \boldsymbol{z} \in \mathcal{Z}}$ for all units $i \in \mathcal{U}$.
The second way the positivity condition can fail is through the design; assignments $\boldsymbol{z} \in \mathcal{Z}$ exist so that $\exmi{\boldsymbol{z}} = d$, but the design is such that $\prei{d} = 0$. Statistical issues are the only sequelae in this case, which all have cures. Two situations must be considered. The first is when the assignment probability for some exposure is exactly zero, $\prei{d} = 0$. The second is when the probability approaches zero asymptotically. Both are problematic, but they have different solutions.
Superficially, the first situation appears most acute. There are two issues to consider. The first is that the definition of $\poexi{d}$ conditions on the measure-zero event $D_{i} = d$, rendering the definition ambiguous. To address this, extend the definition as follows:
That is, if $\prei{d} = 0$, then $\poexi{d}$ is the arithmetic mean of all potential outcomes for which the corresponding $\boldsymbol{z}$ maps to $\exmi{\boldsymbol{z}} = d$. The second concern is that the Horvitz--Thompson estimator now involves division by zero, rendering it ill-defined. However, for each term of the estimator with a zero denominator, the numerator is also zero with probability one. Hence, a straightforward solution is to define $0/0$ as zero, and that is the solution that will be used here. But to ensure that the estimator behaves well, the proportion of units with zero assignment probability must be small. To capture this, let $\zprei{d} = \indicator{\prei{d} = 0}$ denote whether unit $i$ has a zero assignment probability, and let
be the proportion of such units in the sample.
As for assignment probabilities that approaches zero, we must consider the rate at which they do so. The following norm-like quantity captures the average rate of convergence towards zero:
The quantities $\bar{s}_{d}$ and $\zprmom{d, p}$ allow us to weaken the positivity assumption in a controllable way. In particular, Condition \mainref{cond:positivity} is the same as $\bar{s}_{d} = 0$ and $\lim_{p \to \infty} \zprmom{d, p} \leq k_2 < \infty$. The following proposition shows that neither part is necessary for consistency. But the weakening comes at the cost of potentially slower convergence rates. This is captured by a strengthening of the definition of the design dependence, namely
This extended definition collapses to the definition in Condition \mainref{cond:limited-design-dependence} when $q = 1$, but generally $\bar{c}_{d} = \littleO[\big]{\ddepext{d}{q}}$ when $q > 1$. Thus, $\ddepext{d}{q} = \littleO{1}$ when $q > 1$ is stronger than the original condition. Extending the results in Delevoye2020Consistency to a setting with interference, this additional machinery admits a proof of consistency without positivity.
The proposition states that we can achieve consistency under misspecification even if positivity does not hold as long as the dependence between exposures, as captured by $\ddepext{d}{q}$, is sufficiently weak. While experimenters still should try to ensure that their designs and exposure mappings satisfy positivity, Proposition (ref) provides some reassurance that the results are not automatically invalidated in the case they are not perfectly successful and small violations to positivity occur.
The limiting distribution of the estimator is less tractable than its limit. Two approaches are explored in this section. The first is to tie the limiting behavior of the estimator under misspecification to its behavior when the exposures are correctly specified. Experimenters might find this result useful because familiar results for correctly specified exposures are directly extended to a setting with misspecification. The second approach is a direct proof of asymptotic normality using Stein's method. Both approaches require considerably stronger assumptions than those needed for consistency.
A situation where progress can be made is when the specification errors are very small relative to the sample size. The following aggregated measure of the misspecification is a strengthening of the error dependence measures used for the consistency results.
The definition is a strengthening of Definition \mainref{def:error-dependence} in two ways. First, the average total error dependence considers only the positive parts of the specification errors, while the dependence measures in Definition \mainref{def:error-dependence} includes the negative terms as well. It is possible that the dependence between some units' errors is negative, and this will have a compensatory effect in Definition \mainref{def:error-dependence}, making the average smaller. Definition (ref) ignores any such compensatory effects. Second, unlike the previous dependence measures, the average total error dependence includes the diagonal elements of the double sum, $\E{\varepsilon_{i}^2 \nonscript\:\delimsize\vert\allowbreak\nonscript\:\mathopen{} D_{i} = d}$, which generally will be larger than the off-diagonal elements.
To connect the behavior of the estimator under misspecification to its behavior under correctly specified exposures, we will impose a general condition on the experimental design and exposure mappings to ensure that the estimator is well-behaved when the exposures are correctly specified.
Under Condition \mainref{cond:bounded-pos} (bounded potential outcomes), the Horvitz--Thompson estimator of the exposure effect between exposures $a$ and $b$ has the limiting distribution $Q$ when the exposures are correctly specified if and only if Condition (ref) holds. Indeed, Condition (ref) is simply that statement written in a somewhat more general form. That is, if an experimenter is in a setting where they believe the estimator has some limiting distribution $Q$ if the exposures are correctly specified, then Condition (ref) holds.
The proposition states that the limiting distribution of the estimator is unchanged if the specification errors are small. However, observe that the condition $\bar{t}_{a} + \bar{t}_{b} = \littleO{R_n^{-2}}$ used in Proposition (ref) is considerably stronger than Condition \mainref{cond:limited-error-dependence} used for the consistency results. Indeed, while the previous conditions allowed for considerable misspecification also in large samples, this stronger condition is saying that any misspecification is negligible asymptotically relative to the variability induced by the randomization of treatments. Hence, the applicability of Proposition (ref) is limited, and experimenters should show caution before using the proposition to motivate any inferential statements. But, in the few situations in which Proposition (ref) is applicable, experimenters may find the following special case of the proposition particularly useful.
Stein's method has been extensively used in the recent literature on interference. An early example is Aronow2017Estimating. The result used here is due to Ross2011Fundamentals. This result has previously been used by Chin2019Central. Ogburn2022Causal provide a stronger result, but the result in Ross2011Fundamentals is used here due to its simplicity. There is nothing in the current application of Stein's method that is new compared to the previous results in the interference literature, but it might be of interest to note that these well-known results extend also to situations with misspecification.
Define the dependency neighborhood for a unit $i \in \mathcal{U}$ to be the smallest subset $\mathcal{N}_i \subset \mathcal{U}$ such that unit $i$'s exposure $D_{i}$ and specification error $\varepsilon_{i}$ are independent of the exposures and specification errors of all units not in $\mathcal{N}_i$:
Let $d_{\max} = \max_{i\in\mathcal{U}} \abs{\mathcal{N}_i}$ be the size of the largest dependency neighborhood in the sample.
Condition (ref) is considerably stronger than the conditions used in the main paper. It restricts the design to only exhibit local dependence with respect to the exposures, which rules out many common designs and some exposure mappings. It also restrict the specification errors to only be locally dependent. That is, the type of global error dependence that was found to be unproblematic in the main paper is not allowed here. While the condition is strong, it might nevertheless be found to be reasonable in some experiments. In that case, the sampling distribution of the estimator will be approximately normal in large samples, as the following proposition demonstrates.
Variance estimation for exposure effect estimators is challenging because the variance consists of pair-wise products of potential outcomes, and some of those outcomes are not simultaneously observable. The issue is not unique to exposure effects, but exposure mappings tend to induce complex distributions on the exposures, which exacerbates the problem.
The solution suggested by Aronow2017Estimating is to use Young's inequality for products to bound the unobservable parts of the variance expression. To better understand this idea, let $\preij{d_{1}, d_{2}} = \Pr{D_{i} = d_{1}, D_{j} = d_{2}}$ be the joint probability of unit $i$ and $j$'s exposures. If $\preij{d_{1}, d_{2}} = 0$, then the potential outcomes $\poexi{d_{1}}$ and $\poexj{d_{2}}$ are never observed simultaneously, which will complicate variance estimation because the variance depends on the product $\poexi{d_{1}} \poexj{d_{2}}$. Aronow2017Estimating tackle these products by using the bound
After having applied the bound to all problematic terms in the variance of the point estimator, they arrive at the estimator
where,
Aronow2017Estimating show that this variance estimator is conservative in expectation when exposures are correctly specified. However, what does not appear to be fully appreciated in the literature is that the bound on the problematic products may make the estimator excessively conservative. In fact, unless the assumption of correctly specified exposures is complemented with
the normalized variance estimator $n \EstVarAS[\big]{\est{a, b}}$ generally diverges to infinity. I will not offer a solution to this problem. The remark instead serves as an illustration of the difficulty of variance estimation for complex exposure effects. It also provides insights about the mechanics of the estimator, which will aid our understanding of its behavior under misspecification.
The analysis of the variance estimator by Aronow2017Estimating does not hold when the exposures are misspecified. The expectation of the variance estimator could both increase and decrease under misspecification relative to when the exposures are correctly specified. In some situations, the decrease is sizable, and the estimator may become anti-conservative, providing an unjustly optimistic estimate of the precision of the point estimator. The following proposition exactly characterizes the bias of the variance estimator under misspecification.
The terms on the right-hand side capture aspects that introduce bias of the variance estimator. These bias terms help us understand when variance estimation is possible under misspecification. The term $\varbiasf{1}{a, b}$ stems from what Holland1986Statistics describes as the fundamental problem of causal inference, namely that a unit cannot simultaneously be assigned to two different treatments. The joint distribution of the potential outcomes affects the variance, but the distribution can only be estimated if both potential outcomes are observed simultaneously. Such simultaneous observations would require simultaneous assignment of two different treatments to the same unit, but this is not possible. The first term captures the bias arising from our inability to estimate this aspect of the potential outcomes. The issue is not unique to the current setting, and similar bias terms arise for most causal inference problems in finite populations, including when the exposures are correctly specified.
As already noted, the distribution of the exposures is often complex, and the joint exposure probabilities may be zero for a considerable number of pairs of units. This issue is similar to the first source of bias, but it is now introduced by the design rather than being fundamental. The use of Young's inequality to tackle this problem introduces bias, and this bias is captured by the terms $\varbiasf{2}{a, b}$ and $\varbiasf{2}{b, a}$ in Proposition (ref). These biases arise also when the exposures are correctly specified.
At this point, we have replicated the result in Aronow2017Estimating. In particular, the remaining terms are zero by construction when the exposures are correctly specified. In other words, the bias of the variance estimator when the exposures are correctly specified is
which the same result as in Aronow2017Estimating, but presented in a different form. The fact that these three terms are non-negative by construction confirms that the estimator is conservative when the exposures are correctly specified. What Proposition (ref) shows is that we need to consider additional sources of bias when the exposures are misspecified.
The remaining biases stem from two sources. The first also arises from the use of Young's inequality in the construction of the estimator. If the probability that two units are simultaneously assigned to a certain combination of exposures is zero, then the corresponding specification errors cannot interact, and the errors do not affect the variance of the point estimator. This means that we do not need to adjust for any dependence between such errors, because we know it is zero. However, these pairs of units are exactly the ones for which we apply the bound on the unobserved product of potential outcomes to ensure conservativeness under correctly specified exposures. This inadvertently leads to that the variance estimator is affected by the magnitude of the corresponding errors. The terms $\varbiasf{3}{a}$ and $\varbiasf{3}{b}$ capture this part of the bias. Like the previous terms, these terms are non-negative by construction.
The terms of real concern are the last three: $\varbiasf{4}{a, b}$, $\varbiasf{4}{a, a}$, and $\varbiasf{4}{b, b}$. These capture the bias introduced by our inability to estimate the dependence in the specification errors. Unlike the previous terms, the signs of the terms are unknown, so they may introduce negative bias. The consequence is that we could systematically underestimate the variance when the specification errors are negatively dependent, and our inferences would then be anti-conservative.
The problem has no immediate solution, but some progress can be made. Similar to the approach taken in Section (ref), if the specification errors can be assumed to be negligible relative to the sample size, the terms $\varbiasf{4}{d_{1}, d_{2}}$ are negligible relative to the other terms, and the variance estimator is ensured to be asymptotically conservative.
An alternative approach is to incorporate more information about the structure of the interference in the variance estimator. In particular, the anti-conservative behavior of the estimator stems from negative interactions of errors in $\varbiasf{4}{d_{1}, d_{2}}$. One may remove such interactions by setting $\zpreij{d_{1}, d_{2}} = 1$ for the corresponding pairs of units, even if $\preij{d_{1}, d_{2}} > 0$ holds. This will move the corresponding terms from $\varbiasf{4}{d_{1}, d_{2}}$, where negative interactions are possible, to $\varbiasf{3}{d}$, where no interactions exist. It may be hard to discern whether the interaction between two specific units' errors is negative or positive. A conservative approach is to set $\zpreij{d_{1}, d_{2}} = 1$ for all pairs of units where an interaction of any type is suspected.
As an example, consider when units only interfere with each other within known disjoint groups, which is an interference structure often referred to as “partial interference.” One may here set $\zpreij{d_{1}, d_{2}} = 1$ if either $\preij{d_{1}, d_{2}}$ is zero or if unit $i$ and $j$ belong to the same group. A redefinition of $\zpreij{d_{1}, d_{2}}$ along these lines would ensure that $\varbiasf{4}{d_{1}, d_{2}}=0$, so the variance estimator remains conservative. Of course, such knowledge about the interference would allow for the definition of exposure mappings that are correctly specified, which would remove any concerns about misspecification. However, experimenters may want to keep the main exposure mapping simple to facilitate interpretation. They can then proceed with misspecified exposures for point estimation, and use the more intricate information about the interference structure only when estimating variance.
A third approach is a combination of the previous two. One may set $\zpreij{d_{1}, d_{2}} = 1$ for pairs of units where negative interaction are suspected to be particularly large. Even if one does not capture all pairs with negative interactions, the variance estimator will still be conservative as long as the missed interactions are small. In other words, setting $\zpreij{d_{1}, d_{2}} = 1$ for the terms deemed most problematic makes the assumption that the remaining errors are small more reasonable. It should also be noted that the other bias terms tend to be large and positive, and they therefore provide considerable leeway in the case $\varbiasf{4}{d_{1}, d_{2}}$ is negative.
The expectation of the variance estimator should be seen as a rough representation of its behavior more generally. The precision of the variance estimator will be poor if the joint exposure probabilities are small even if they are never exactly zero. Experimenters should be aware that the variance estimator could be imprecise even when the bias is positive, particularly when the number of exposures is large. This concern is not specific to misspecification.
\DeclarePairedDelimiterXPP\estorc[1]{\widehat{\bar{\tau}}}{\lparen}{\rparen}{#1}
\DeclarePairedDelimiterXPP\esterr[1]{\hat{\varepsilon}}{\lparen}{\rparen}{#1}
The linearization used to prove consistency for the H{\'a}jek estimator requires an alternative representation of the estimator.