EconBase
← Back to paper

Potential weights and implicit causal designs in linear regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

111,767 characters · 16 sections · 104 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Potential weights and implicit causal designs in linear regression

abstractWhen we interpret linear regression as estimating causal effects justified by quasi-experimental treatment variation, what do we mean? This paper formalizes a minimal criterion for quasi-experimental interpretation and characterizes its necessary implications. A minimal requirement is that the regression always estimates some contrast of potential outcomes under the true treatment assignment process. This requirement implies linear restrictions on the true distribution of treatment. If the regression were to be interpreted quasi-experimentally, these restrictions imply candidates for the true distribution of treatment, which we call implicit designs. Regression estimators are numerically equivalent to augmented inverse propensity weighting (AIPW) estimators using an implicit design. Implicit designs serve as a framework that unifies and extends existing theoretical results on causal interpretation of regression across starkly distinct settings (including multiple treatment, panel, and instrumental variables). They lead to new theoretical insights for widely used but less understood specifications.

Introduction

Linear regression is overwhelmingly popular in applied microeconomics for estimating causal effects. Users frequently justify it by arguing that a treatment variable is quasi-experimentally assigned pgp,angrist2010credibility,currie2020technology, rather than that it correctly specifies a structural model for potential outcomes. Under this view, a regression (e.g., $Y_i = W_i\tau + x_i'\gamma + \epsilon_i$ for treatment $W$ and covariates $x$) defines an estimand whose causal meaning---if any---comes from assumptions on treatment assignment, not functional-form assumptions on potential outcomes. Moreover, under these assumptions, $\tau$'s causal interpretation holds for arbitrary potential outcomes and heterogeneous treatment effects.\footnote {This quasi-experimental view reflects in angrist2010credibility, “With the growing focus on research design, it's no longer enough to adopt the language of an orthodox simultaneous-equations framework [\ldots] The new emphasis on a credibly exogenous source of variation has also filtered down to garden-variety regression estimates, in which researchers are increasingly likely to focus on sources of omitted-variables bias, rather than a quixotic effort to uncover the `true model' generating the data.”}

Practitioners appear optimistic that this quasi-experimental interpretation is typically available without taking the regression seriously as an outcome model. As angrist2008mostly put it in the preface to Mostly Harmless Econometrics,

quoteMost econometrics texts appear to take econometric models very seriously [\ldots Instead,] a principle [here] is that estimators in common use almost always have a simple interpretation that is not heavily model-dependent.

However, while regressions naturally represent an outcome model, they usually do not spell out an assignment model or the implied causal estimand under heterogeneous effects. As a result, applied work often proceeds by informally asserting that treatment is “as good as randomized,” choosing a specification, and interpreting its coefficients as causal---leaving implicit (i) what one must believe about treatment assignment to justify that interpretation and (ii) what weighting of heterogeneous effects the regression is estimating.

A large applied econometrics literature studies these questions in specific settings:\footnote{Among others, imbensangrist,angrist1998,lin2013agnostic,sloczynski2022interpreting,sloczynski2020should,blandhol2022tsls,aronow2016does,goldsmith2022contamination,borusyak2024negative,athey2018design,kline2011oaxaca,bugni2023decomposition,mogstad2024instrumental,arkhangelsky2023fixed,arkhangelsky2021double,chetverikov2023logit,kolesar2024dynamic,zhao2025interacted,arganaraz2024randomly.} Certain specifications have quasi-experimental interpretations under some treatment assignments, but others may not (e.g., produce negatively weighted causal effects). Seemingly small differences can be critical: With binary $W_i$, $Y_i = W_i\tau + x_i'\gamma + \epsilon_i$ estimates a weighted average treatment effect when the propensity score is linear in $x_i$ angrist1998, but the analogous specification with multi-valued $W$ produces uninterpretable estimands goldsmith2022contamination.\footnote{That is, for $W$ that takes values $\br{0,\ldots, J}$, the regression $Y_i = \sum_{j=1}^J \tau_j \one(W=j) + x_i'\gamma + \epsilon_i$ produces $\tau_j$ that suffers from contamination bias.}

Without a general principle, practitioners may struggle to navigate the many requirements for quasi-experimental interpretation. For instance, what treatment assignment assumptions are needed to interpret the interacted regression $Y_i = W_i\tau_0 + W_i x_i'\tau_1 + x_i'\gamma + \epsilon_i$ quasi-experimentally? What about a panel regression like $Y_ {it} = \alpha_i + \beta_t + W_{it}\tau + x_{it}'\gamma + \epsilon_{it}$? If they can be interpreted quasi-experimentally, what causal effects do they then target when treatment effects are heterogeneous?

This paper provides a general framework for quasi-experimental interpretation of arbitrary linear regressions with finite-valued treatments. For any specification, it computes candidate treatment-assignment processes (henceforth designs) and, for each candidate, the implied estimand. If the regression admits a quasi-experimental interpretation, the true assignment process must be one of these candidates; the corresponding estimand then makes explicit the regression's weighting of heterogeneous causal effects. We do this by formalizing a minimal criterion for quasi-experimental interpretation, then characterizing the designs and estimands it implies.

Adding to the applied econometrics literature, the framework unifies and extends several specification-specific results. These results can be obtained by mechanically computing the candidate designs and their implied estimands for the specification at hand. This computation recovers existing results and proves converses for them: Designs studied in the literature are the only admissible candidates. The same exercise also produces new results for common specifications: In particular, some specifications that otherwise posit reasonable outcome models do not admit any quasi-experimental interpretation at all. Finally, when a plausible design does exist, we show that the regression estimator numerically equals an augmented inverse-propensity weighting (AIPW) estimator computed under that design, yielding a doubly robust interpretation. More broadly, this framework contributes to a literature on the interpretation of estimators under misspecification, as reviewed by andrews2025purpose.

For practitioners, the framework provides a transparent way to make regressions' identifying content and target estimands explicit. Reporting the implied design clarifies what “as-good-as-random” must mean for the chosen specification, while reporting the implied estimand clarifies what causal effect is being aggregated under heterogeneity. The framework provides diagnostics for when a quasi-experimental interpretation is impossible and guidance for how to refine specifications to improve interpretability and robustness.

To build this framework, we first formalize quasi-experimental interpretation by asking what a regression coefficient would estimate in the idealized experiment that redraws treatment according to the true assignment mechanism. Let $\pi_i^*$ denote the true assignment probabilities of unit $i$'s treatment $W_i$. Given any potential outcomes $y_i(\cdot)$, imagine repeatedly drawing $W_i \sim \pi_i^*$, observing $Y_i = y_i(W_i)$, and estimating the regression coefficient on each draw. This process's large sample limit defines the regression estimand $\tau$ as a functional of $y_i(\cdot)$.

We contend that a minimal requirement for calling $\tau$ “quasi-experimental” is that it be a contrast under this experiment, regardless of potential outcomes:

enumerate[label=[MQE]] • Under $\bm{\pi}^*$, the estimand $\tau$ is a contrast\footnote{If there are $J+1$ treatments $\br{0,\ldots,J}$ and $n$ units $\br{1,\ldots, n}$, a contrast of individual potential outcomes $y_{i}(j)$ is defined to be a parameter $\frac{1}{n} \sum_{i=1}^n \sum_{j=0}^{J} \omega_ {i}(j) y_i(j)$, for weights that sum to zero across $j$, $\sum_{j=0}^{J} \omega_ {i} (j) = 0$. With binary treatments, these are parameters of the form $\frac{1} {n}\sum_{i=1}^n \omega_i (y_i(1) - y_i(0))$, where the $\omega_i$'s are permitted to be negative.} of individual potential outcomes for any potential-outcome distribution---even a worst-case one.

(ref)---for minimally quasi-experimental---codifies two requirements. First, it enforces the “treatment-based” logic practitioners appeal to: If a coefficient is causal because of treatment variation, then its causal meaning should not depend on special features of the potential outcomes.\footnote{While weak, this requirement excludes specifications whose causal interpretation hinges on correctly modeling outcomes. For instance, difference-in-differences, whose validity hinges on outcome-dependent parallel trends assumptions roth2023parallel, does not qualify as quasi-experimental per (ref). Nevertheless, studying (ref) is informative for difference-in-differences, because researchers often appeal to treatment variation as justifying parallel trends and (ref) evaluates these arguments. For instance, martinez2022rise write (emphasis ours), “we study the introduction of [Chinese local] elections in the 1980s and 1990s [\ldots] We document that the timing of the first election is uncorrelated with a large set of village characteristics. This suggests that timing was quasi-random [\ldots] Thus, we exploit the staggered timing of the introduction of elections across villages {to estimate a difference-in-difference effect} of the introduction of elections.”} Second, it imposes level independence blandhol2022tsls: If all individual treatment effects are zero, then the estimand should always be zero.

If a regression satisfies (ref), practitioners can safely interpret its estimates as some---though not necessarily {useful}---causal effects, identified in the idealized experiment $\bm{\pi}^*$. With binary treatment, for instance, (ref) requires that $\tau$ be a weighted average of $y_i(1) - y_i(0)$ and does not require these weights be convex. This permissiveness is deliberate---e.g., it allows for calling treatment effect differences quasi-experimental. This permissiveness is also not a limitation: If one wants stronger properties---e.g., convex weights---those can be checked after computing the implied estimand. Practitioners can retarget by reweighting if desired.

Because (ref) depends on the unknown true design $\bm{\pi}^*$, we cannot test it directly. We therefore decompose it into an existence and a correctness question:

enumerate[label=[MQE-\arabic*]] • Does any design $(\pi_1,\ldots,\pi_n)$, consistent with the data, satisfy (ref)? • Is any design in (ref) equal to $\pi_1^*,\ldots, \pi_n^*$?

(ref) is an existence question: Does an idealized experiment even exist that supports quasi-experimental interpretation? This question can be objectively answered because it reduces to verifiable restrictions computable from data. (ref) is a separate, context-specific correctness question---whether the design in (ref) is correctly specified. Like any model specification question, evaluating it requires subjective judgement.

Our framework enables practitioners to objectively and systematically evaluate (ref) through (ref). (ref) alone is not sufficient for quasi-experimental interpretation, but it is useful: It either rejects quasi-experimental interpretation outright, or produces a narrow set of concrete candidate designs that makes important debates about (ref) explicit rather than implicit. (ref) alone also adds to interpretation of the regression estimator: When a candidate design satisfies (ref), the regression estimator itself is numerically equivalent to an AIPW estimator targeting the corresponding estimand under that design.

Having defined quasi-experimental interpretation, we constructively characterize designs that satisfy (ref). In particular, for any linear regression coefficient $\tau$, there exists potential weights $\rho_i(w)$---known functions of the regression specification---such that $\tau$ is a linear combination of potential outcomes: \[ \underbrace{\tau}_{\text{Regression estimand, } \R^k} = \frac{1}{n}\sum_{i=1}^n \sum_{{w} \in \mathcal{W}} \underbrace{ { {\pi_i^* ({w})}}}_{\text{true design, } \R} \cdot \underbrace{{\rho_i ({w})}}_ {\text{potential weights, } \R^{k \times T}} \cdot \underbrace{{y}_i ({w})}_{\text{potential outcomes, } \R^T}. \numberthis \label{eq:aggregation} \] If $\tau$ satisfies (ref), it must be invariant to adding a constant to all of unit $i$'s potential outcomes. This requirement translates into linear restrictions on the assignment mechanism $\bm{\pi}^*$: the true design must satisfy $\sum_ {w\in \mathcal W} \pi_i^*(w) \rho_i(w) = 0$, for all $i$, if $\tau$ satisfies (ref). We call any design $ (\pi_1,\ldots, \pi_n)$ that solves these linear equations an implicit design of the regression. In many leading specifications the linear restrictions are sharp---often yielding either no solution or a unique one---so (ref) can be highly informative.

Given an implicit design, the regression also targets a corresponding causal contrast, determined together with the implicit design. We refer to this contrast as the implicit estimand. For an implicit design $\bm{\pi}$, its corresponding implicit estimand puts weight $\omega_i(w) \equiv \pi_i(w) \rho_i(w)$ on the unit-$i$ potential outcome $y_i (w)$. These weights $\omega_i(w)$ can be explicitly computed.

Together, implicit designs and implicit estimands characterize (i) which idealized experiments are consistent with the regression and (ii) which causal contrast a regression targets. They also endow the OLS estimator with a doubly robust interpretation as an AIPW estimator for the implicit estimand using the implicit design, building on bruns2025augmented and robins2007comment. Our tools make these implicit choices explicit, enabling researchers to transparently assess their validity.

We then use the framework to deliver payoffs for both theory and practice: For applied econometrics, it provides a common language for when regressions do (and do not) admit design-based causal meaning. For practitioners, it turns otherwise implicit assumptions about “the ideal experiment” and otherwise implicit choices about heterogeneous treatment effects into objects that can be computed, inspected, and stress-tested.

On the theory side, first, (ref) captures the shared logic across starkly different settings angrist1998,blandhol2022tsls,goldsmith2022contamination,kline2011oaxaca,athey2018design. Computing implicit designs and implicit estimands for a regression recovers the designs and estimands posited in these papers and delivers their converses---namely, that quasi-experimental interpretation is only possible under exactly those designs.

Second, we uncover new results for specifications that interact treatment with covariates lin2013agnostic,miratrix2013adjusting,imbens2009recent,kline2011oaxaca,zhao2025interacted and for two-way fixed effects (TWFE). In both cases, quasi-experimental interpretation can be fragile in the sense that implicit designs need not exist outside special cases. Taken together, these results suggest that quasi-experimental interpretation of regression is perhaps less generic than predicted by angrist2008mostly.

Third, we extend the framework to two-stage least-squares (TSLS). There, our framework characterizes requirements on the instrument assignment process for interpreting TSLS coefficients as instrument-on-outcome contrasts (i.e., intent-to-treat effects). The implicit estimand here additionally pins down restrictions for treatment compliance patterns for interpreting TSLS estimands as reasonable treatment-on-outcome effects. Our framework similarly unifies and extends the TSLS literature\footnote{We share a focus on unified analysis with related papers by navjeevan2023identification and goff2024does. Compared to these papers, our starting point is the interpretation of a particular TSLS estimator. } by recovering converses to results in blandhol2022tsls,imbensangrist,behaghel2013robustness,sloczynski2020should,bhuller20242sls---they even help clarify a small gap in recent work on TSLS with multiple treatments.

For applied work, our primary recommendation is to compute, evaluate, and report implicit designs and implicit estimands whenever regressions are interpreted quasi-experimentally. If causal interpretation hinges on an idealized experiment and an induced aggregation of heterogeneous effects, then those choices should be made explicit. To facilitate that, we discuss practical diagnostics that (i) check whether the implicit design is proper and calibrated and (ii) evaluate its functional form statistically and economically. In addition, once an implicit design is deemed plausible, we use the implicit estimand to diagnose sensitivity to heterogeneity (including whether some units receive negative weight), and, when the implicit estimand is not substantively meaningful, we show how to retarget alternative estimands by reweighting. We illustrate these recommendations with re-analyses of blakeslee2020way and cervellati2024random, so that a reader can see what the regression is implicitly “assuming” and “averaging” rather than taking either on faith.

This paper proceeds as follows. (ref) contains our main results. To build intuition, (ref) starts with a simple setting with cross-sectional data and binary treatments. (ref) then formalizes (ref), (ref), and (ref) and their relation to implicit designs. (ref) applies our framework to a litany of regression specifications, yielding new theoretical results. (ref) extends the framework to TSLS. (ref) illustrates our diagnostics with two empirical applications. (ref) concludes.

Potential weights and implicit designs

Consider a finite population of units $i \in [n] \equiv \br{1,\ldots, n}$. Each unit receives one treatment ${w}$ from a finite set $\mathcal{W}$ of size $J+1$. Each unit has covariates ${x}_i$ and vector-valued potential outcomes of length $T$, $\br{{y}_i ({w}) \in \R^T : {w} \in \mathcal{W}}$.\footnote{For expositional clarity, we assume that the dimension of the outcome vector is the same across individuals (i.e., balanced panels). (ref) discusses imbalanced panels.} We denote by ${W}_i$ the realized treatment. After assignment, we observe a corresponding realized outcome ${Y}_i = {y}_i ({W}_i)$.

To emphasize that identification comes from variation in treatment assignment, we isolate this variation by thinking of $({x}_i, {y}_i(\cdot))$ as fixed numbers and only considering the randomness in ${W}_i$. Cosmetically, this design-based perspective aligns with how quasi-experimentalists argue identification and with how we compute implicit designs. Substantively, it allows for treatment assignment to be correlated. Importantly, adopting a design-based setup does not drive different conclusions from a sampling one---since it conditions on the sampled $(y_i(\cdot), x_i)$; see (ref).

Let $\bm{\pi}^*$ denote the marginal treatment assignment probabilities (i.e., propensity scores):\[\bm{\pi}^* = (\pi^*_1, \ldots, \pi^*_n) \text{ where } \pi_i^*({w}) = \P({W}_i = {w}). \] We call $\bm{\pi} =(\pi_1 (\cdot),\ldots,\pi_n(\cdot))$ a design. In principle, these probabilities may be arbitrarily different across units. Write a regression generically as ${Y}_{it} = z_t({x}_i, {W}_i)'\beta + \epsilon_{it}$ with known $z_t(\cdot, \cdot)$. For a known matrix $\Lambda \in \R^{k \times K}$, we would like to interpret certain coefficient contrasts $\tau \equiv \Lambda\beta \in \R^k$ as causal effects. To emphasize, this regression does not specify a structural model; it simply specifies an estimand $\tau$ given $(\bm{\pi}^*, ({y}_i(\cdot), {x}_i)_ {i=1}^n)$. Since it is common in practice to specify a regression first and interpret its estimated coefficients as causal effects, our analysis starts with a regression and investigates which configurations of $\bm{\pi}^*$ are compatible with interpreting the regression under (ref).

This setup is general: It encompasses cross-sectional ($T=1$), panel ($T > 1$), $ (J+1)$-valued treatment, scalar contrast ($k = 1$), and multiple contrasts ($k > 1$) settings.\footnote{Our results do extend to continuous treatments, but they become much less powerful, essentially because there are only finitely many restrictions for infinitely many objects. } (ref) extends these results to TSLS. To illustrate (ref), we start with the binary-treatment, scalar-outcome, and scalar-contrast case $(T=J=k=1)$. Our main results then push this intuition to the general case.

Core intuition

To motivate the framework, blakeslee2020way study the impact of water loss in rural India on employment and income. Water loss is measured by a binary $W_i$, indicating whether the first borewell household $i$ drilled has failed. The authors motivate quasi-experimental identification by emphasizing that well failure “depends on highly irregular, quasi-random subsurface properties” (p. 206). The true design $\bm{\pi}^*$---the natural process of borewell failure---is unknown, but the authors argue that failure is difficult to predict, making treated and untreated households plausibly comparable, and they marshal detailed hydrogeological evidence in support of this claim.

blakeslee2020way then estimate a simple regression across multiple outcomes: For $i$ a household and covariates $x_i$, \[Y_i = \tau W_i + x_i'\gamma + \epsilon_i, \text{ for which } \numberthis z (x_i, W_i) = [W_i, x_i']', \beta = [\tau, \gamma']', \Lambda = [1,0_{\dim(x)}']. \label{eq:angrist98intro} \] From the perspective of quasi-experimental interpretation, the key tension is that the regression itself does not encode the substantive discussion of $\bm{\pi}^*$: Instead, $\bm{\pi}^*$ is left implicit as whatever assignment process that would justify interpreting (ref) quasi-experimentally. What, then, must a reader believe about $\bm{\pi}^*$ for (ref) to have a quasi-experimental interpretation, and which causal contrast does $\tau$ represent when effects are heterogeneous?

To answer these questions, let us return to a regression of a scalar $Y_i$ on some known transform $z (x_i, W_i)$. The population regression coefficient is defined as: \[ \beta \equiv \pr{{\frac{1}{n}\sum_{i=1}^n \E_{W_i \sim \pi_i^*} \bk{z (x_i, W_i) z(x_i, W_i)'}}}^{-1} \pr{\frac{1}{n}\sum_{i=1}^n \E_{W_i \sim \pi_i^*} [z (x_i, W_i) y_i (W_i)]}. \] This definition is simply the design-based analogue of the usual “$\E[x_i x_i']^{-1}\E [x_i y_i]$” formula. Let $G_n \equiv G_n(\bm{\pi}^*) \equiv {\frac{1}{n}\sum_ {i=1}^n \E_{W_i \sim \pi_i^*} \bk{z(x_i, W_i) z(x_i, W_i)'}}$ denote the population Gram matrix of this regression. Since $G_n$ is consistently estimable, we treat it as known.\footnote{In particular, since the regression estimator replaces $G_n$ with $\hat G_n = \frac{1}{n} \sum_{i=1}^n z(x_i, W_i) z(x_i, W_i)'$, it is implausible that the regression estimator is consistent but $\hat G_n$ is far from $G_n$. (ref) provide formal guarantees for $\hat G_n$.

Because $G_n$ depends on the unknown $\bm{\pi}^*$, treating $G_n$ as known implicitly restricts $\bm{\pi}^*$ to those designs that are consistent with the realized treatment assignment. We discuss its interpretation further in (ref).}

Under these definitions, $\tau$ admits the representation (ref): For $\pi_i^* = \pi_i^*(1)$, \[ \tau = \Lambda\beta = \frac{1}{n}\sum_{i=1}^n \pi_i^* \underbrace{\Lambda G_n^{-1} z(x_i, 1)}_{\rho_i(1)} y_i (1) + (1-\pi_i^*) \underbrace{\Lambda G_n^{-1} z(x_i, 0)}_{\rho_i(0)} y_i(0). \numberthis \label{eq:simple_aggregation} \] Here, the potential weights $\rho_i(w) = \Lambda G_n^{-1} z(x_i, w)$ are known up to $G_n$. In the case of (ref) where $x_i$ includes a constant, we can compute $\rho_i(w)$ in closed form:

align*[align* omitted — 212 chars of source]

$\rho_i(w)$ is proportional to $w-x_i'\delta$, where $\delta$ is the projection coefficient of $\pi^*$ on $x_i$.

If the regression is quasi-experimental in the sense of (ref), the true design $\bm{\pi}^*$ is such that the estimand (ref) satisfies level independence blandhol2022tsls:

defnWe say that $\tau$ is minimally quasi-experimental under $\bm{\pi}^*$ if $\tau$ is always unchanged when we replace all potential outcomes $y_i(w)$ with $y_i (w) + c_i$ for arbitrary $c_i \in \R$, holding fixed $(\bm{\pi}^*, x_1,\ldots, x_n)$. Since $\tau$ is a linear aggregation, equivalently, $\tau$ is minimally quasi-experimental if there is some $\omega_1,\ldots, \omega_n \in \R$, not dependent on $y_i(\cdot)$, such that $ \tau = \frac{1}{n} \sum_{i=1}^n \omega_i (y_i(1) - y_i(0))$ for all choices of $y_i(1), y_i (0) \in \R$.

(ref) is a natural minimal requirement for quasi-experimental estimands. It imposes that a quasi-experimental estimand should be invariant to any changes to the potential outcomes that do not change individual treatment effects---holding fixed the treatment assignment process. For linear estimands, this condition is equivalent to $\tau$ being a weighted average treatment effect (these weights $\omega_i$ may be negative).\footnote {Negative weights are intended, for example, when the estimand is meant as a contrast of subgroup average effects. Thus to preserve generality, we allow for negative weights. Alternatively, blandhol2022tsls term an estimand “weakly causal” if it additionally satisfies $\omega_i \ge 0$.}

Allowing for negative weights is admittedly lenient, but we do not view that as a limitation. Since we could recover the estimand itself, we could additionally inspect whether the weighting is convex or whether it satisfies further restrictions. Practitioners can opt to reweight the estimand if dissatisfied with the regression-chosen weighting, and they can examine empirically whether treatment effect heterogeneity correlates sufficiently with these weights for the reweighting to drive conclusions.

Importantly, estimands that rely on modeling $y_i(0)$---e.g., difference-in-differences estimands that rely on parallel trends---do not qualify as quasi-experimental per our definition. These estimands do not mimic a randomized experiment (though quasi-experimental assignment is often invoked to informally justify, e.g., parallel trends). While looser definitions of quasi-experiments are reasonable card2022design, we argue this stricter one is both principled and useful. It is principled by taking very seriously that quasi-experiments should emulate randomized experiments angrist2010credibility,leamer1983let. It is also not overly stringent---the theoretical applications in (ref) show that much of the applied econometrics literature is consistent with this definition.

Returning to (ref), observe that $\tau$ satisfies (ref) under $\bm{\pi}^*$ if and only if \[ \pi_i^* \rho_i(1) + (1-\pi_i^*)\rho_i(0) = 0 \text{ for all $i=1,\ldots,n$.} \numberthis \label{eq:equation_simple} \] We separate two questions: (i) which assignment vectors $\bm{\pi} = (\pi_1,\ldots, \pi_n) $ solve (ref), and (ii) whether the true assignment vector $\bm{\pi}^*$ is plausibly among those solutions. The first question is (ref). Viewing (ref) as an equation in $\bm{\pi}^*$, we can solve to obtain \[ \pi_i = \frac{-\rho_i(0)}{\rho_i(1) - \rho_i(0)}. \] and we call such a $\bm{\pi}$ an implicit design. The second question is [MQE2]: it requires that the implicit and true designs coincide, i.e. $\pi_i^* = \pi_i$.

Identifying $\bm{\pi}$ immediately pinpoints the estimand. If $\pi_i$ were $\pi_i^*$, then, for $\omega_i(\bm{\pi}, w) \equiv \pi_i(w)\rho_i(w)$ and $\omega_i \equiv \omega_i (\bm{\pi}, 1) = -\omega_i(\bm{\pi}, 0)$, $\tau$ is a weighted average treatment effect \[ \tau = \frac{1}{n} \sum_{i=1}^n \omega_i(\bm{\pi}, 1) y_i(1) + \omega_i(\bm{\pi}, 0) y_i(0) = \frac{1}{n} \sum_{i=1}^n \omega_i (y_i(1) - y_i(0)). \] Here, $\omega_i(\bm{\pi}, 1) = -\omega_i(\bm{\pi}, 0)$ because $\pi_i$ satisfies (ref). Thus, simply solving (ref) yields both candidate designs $\bm{\pi}$ and their corresponding implicit estimands.

Computing the implicit designs and estimands is helpful for assessing a regression's quasi-experimental interpretation and for making empirical work more transparent. A necessary requirement for (ref) is (ref), which the implicit designs objectively assess. The mere existence of implicit designs, of course, is not sufficient, since it is possible that none is how the treatment was actually randomized. Nevertheless, computing them makes validating (ref)---inherently a subjective judgement---less abstract.

Finally, how regressions aggregate heterogeneous treatment effects is inherently tied to how they implicitly model treatment assignment. Implicit estimands further clarify whether this aggregation is substantively informative and allow practitioners to enforce stricter standards. For instance, one could decide that (ref) is too lax and require that the implicit estimand be, say, the ATE---reweighting any regression that fails this test towards estimating the ATE instead.

There are at least two ways in which $\pi_i$ cannot possibly equal $\pi^*_i$, leading to a rejection of (ref). The more obvious one is if $\pi_i \not\in [0,1]$ for any $i$ or if $\rho_i(1) = \rho_i(0) \neq 0$, occurring when $\rho_i (1)$ and $\rho_i (0)$ are on the same side of zero. When this happens, the implicit design is not even a probability distribution. More subtly, $\pi_i$ is also indefensible if it generates a Gram matrix that is different from $G_n(\bm{\pi}^*)$: \[G_n(\bm{\pi}) = \frac{1}{n}\sum_{i=1}^n \pi_i z(x_i, 1)z(x_i, 1)' + (1-\pi_i) z(x_i, 0)z(x_i, 0)' \neq G_n(\bm{\pi}^*). \numberthis \label{eq:gram_criterion} \] This restriction is useful when we analyze specifications theoretically under this framework. It is harder to implement when we do not know and have to estimate $G_n$, though, with a confidence set for $G_n$, one could use it as a basis for inference on $\bm{\pi}^*$ (see (ref)).

We summarize these results in the following corollary of (ref), to be introduced.

restatable{cor}{corbinary} When $k=T=J=1$, $\tau$ is minimally quasi-experimental if and only if \begin{enumerate} • $\rho_i(1) \rho_i(0) \le 0$ for all $i$. Some implicit design $\bm{\pi}$ satisfies (ref) and has $\pi_i = \frac{-\rho_i (0)} {\rho_i (1) - \rho_i(0)}$ for all $i$ with one of $\rho_i(1)$ and $\rho_i(0)$ nonzero. • For all units $i$ with one of $\rho_i(1)$ and $\rho_i(0)$ nonzero, $\pi^*_i = \frac{-\rho_i(0)}{\rho_i(1) - \rho_i(0)}$. \end{enumerate} When this happens, the implicit estimand is \[ \tau = \frac{1}{n} \sum_{i=1}^n \omega_i^* (y_i(1) - y_i(0)) \text { for } \omega_i^* \equiv \omega_i(\bm{\pi}^*, 1) = \pi_i^* \rho_i(1). \] The weight $\omega_i^* < 0$ if and only if $\rho_i(1) < 0 < \rho_i(0)$.

The two conditions in (ref) separate (ref) into (ref) and (ref). (ref)(1) formalizes (ref). If an implicit design exists, it is uniquely and explicitly defined (up to units with $\rho_i(1) = \rho_i(0) = 0$). (ref)(2) formalizes (ref), which requires that $\pi_i^*$ is equal to the unique implicit design $\frac{-\rho_i(0)}{\rho_i(1) - \rho_i(0)}$. In this case, the implicit estimand is a weighted average treatment effect, where weights $\omega_i$ are all nonnegative provided no unit has $\rho_i(1) < 0 < \rho_i(0)$. Applied to (ref), (ref) shows that the implicit design is precisely $\pi_i = x_i'\delta$ and the corresponding estimand is a weighted ATE\[ \tau = \frac{1}{n} \sum_{i=1}^n \omega_i(y_i(1) - y_i(0)) \quad \omega_i = \frac{\pi_i (1-\pi_i)} {\frac{1}{n}\sum_{j=1}^n\pi_j (1-\pi_j)} \numberthis \label{eq:angrist_estimand_simple}. \] Simply computing them thus recovers results in angrist1998 and blandhol2022tsls.\footnote{Both angrist1998 and blandhol2022tsls consider a superpopulation sampling setup. angrist1998 considers a binary $x_i$ in his equation (9), but the argument can be easily generalized, e.g., in borusyak2024negative,goldsmith2022contamination. Corollary 1 in blandhol2022tsls---which specializes their TSLS result to OLS---shows that assuming unconfoundedness, $\tau$ is a positively weighted average treatment effect if and only if the propensity score is linear. This is effectively what we find, and thus we view our result (formally in (ref)((ref))) as a reinterpretation of theirs. Additionally, (ref) clarifies how our results relate to Theorem 1 in blandhol2022tsls. }

To summarize, our analysis proceeds in four steps:

mdframed\begin{enumerate}[label=(\roman*)] • We treat the triplet $(\br{z(x_i, \cdot)}_ {i=1}^n, \Lambda, G_n)$ as known (at least in the population). • We write the population regression estimand $\tau$ in the form (ref) and (ref). Because we treat $G_n$ as known, the potential weights $\rho_i (w)$ are known for all units. • We observe that (ref) imposes linear restrictions on $\pi_i^*$, where the coefficients are the potential weights. • Separating (ref) into (ref), we call the solutions to these linear equations implicit designs. Computing implicit designs also yields the corresponding estimands by (ref). If (ref) holds, then $\tau$ has a quasi-experimental interpretation, and one may then assess the extent that $\tau$ is substantively relevant. \end{enumerate}

We generalize these steps in the next subsection and in (ref) and show that the OLS estimator has a doubly robust interpretation. We conclude this subsection by stating the superpopulation analogue of these results.

rmksq[Superpopulation] Suppose instead $ (Y_i (0), Y_i(1), W_i, X_i) \iid P$. We can convert the sampling setup to a design-based setup by setting $\pi_i^* = P(W_i = 1 \mid X_i, Y_i (1), Y_i(0))$ and conditioning on $(X_i, Y_i(1), Y_i(0))$. Now, consider a hypothetical set of potential outcomes $Y_i' (w) = Y_i (w) + C_i$ where $C_i$ is some random variable satisfying $C_i \indep W_i \mid X_i, Y_i(\cdot)$. This independence restriction makes sure that $Y_i'(\cdot)$ does not introduce new selection concerns: $P(W=1 \mid Y_i'(1), Y_i'(0), Y_i(1), Y_i(0), X_i) = \pi_i^*.$ The sampling analogue of (ref) is that $\tau$ is unchanged for all such $Y_i'$: \begin{align*} \tau &= \E[\underbrace{\Lambda \E[z(X_i, W_i) z(X_i, W_i)']^{-1} z(X_i, W_i)}_{\rho_i (W_i)} Y_i(W_i)] = \E[ \rho_i(W_i) Y_i' (W_i)]. \numberthis \end{align*} By the law of iterated expectations, conditioning on $(Y_i(1), Y_i(0), C_i, X_i)$, (ref) is equivalent to $0 = \E\bk{C_i \pr{\pi_i^* \rho_i (1) + (1-\pi_i^*) \rho_i(0)}} $. Since we can choose $C_i$ as an arbitrary function of $X_i, Y_i(1), Y_i(0)$ and in particular as $C_i = \pi_i^* \rho_i(1) + (1-\pi_i^*) \rho_i(0)$, we can force the following condition, which is the analogue of (ref): \[ \pi_i^* \rho_i(1) + (1-\pi_i^*) \rho_i(0) = 0 \quad \text{$P$-almost surely.} \] See (ref) for a formalized analogue with general $J,k,T$.

General setup

We now generalize to panel data and multivalued treatments. Consider a regression of ${Y}_ {it}$ on some transform $z_t ({x}_i, {W}_i) \in \R^K$ of covariates and treatment. The population Gram matrix is \[ G_n (\bm{\pi}^*) \equiv \frac{1}{n} \sum_{i=1}^n \sum_{t=1}^T \E_{{W}_i \sim \pi_i^*} \bk{z_t ({x}_i, {W}_i) z_t({x}_i, {W}_i)'}. \] Following (ref), let ${z}({x}_i,\cdot) \in \R^{T\times K}$ stack $z_t ({x}_i, \cdot)$; we treat $ (\Lambda, G_n, {z} ({x}_1, \cdot), \ldots, {z}({x}_n, \cdot))$ as known and refer to this tuple as a population regression specification.

rmksqThere are two subtleties for panel settings. First, since treating $G_n$ as known is motivated by its consistent estimation, we require representing fixed effects through the within-transformation for ${z}({x}, \cdot)$, rather than through unit-level dummy variables.\footnote{That is, individual fixed effects should be incorporated by setting $\sum_t z_t({x}_i, \cdot) = 0$, rather than by considering unit dummies as covariates. To see this, for unit $i$, let $\tilde z_t ({x}_i, {W}_i)$ denote the covariate transforms that exclude the unit dummy. Assume $z_t({x}_i, {W}_i)$ includes a unit dummy. Then $\sum_{t=1}^T \E[\tilde z_t({x}_i, {W}_i)]$ is in the Gram matrix (it is the interaction between $\tilde z_t$ and the unit-$i$ dummy variable). However, this quantity is not consistently estimable as unit $i$ is only observed once.} Second, assuming ${z}({x}, \cdot)$ is known precludes mediators (e.g. lagged outcomes) in the right-hand side of the regression, since we do not know counterfactual values of the mediator.

As in (ref), the regression estimand is:

align*[align* omitted — 254 chars of source]

which verifies the representation (ref). Relative to the simple case (ref), we sum over $J+1$ values, and potential weights $ \rho_i({w}) \equiv \Lambda G_n^{-1} {z}({x}_i, {w})' $ are matrices of dimension $k \times T$.\footnote{(ref) shows that the potential weights for a given contrast do not depend on how the regression is parametrized. For instance, it does not matter if we write $Y_i = \alpha + \tau W_i + \epsilon_i$ instead as $Y_i = \mu_1 \one(W_i=1) + \mu_0\one (W_i=0) + \epsilon_i$ and consider $\tau = \mu_1 - \mu_0$. (ref) also shows that the potential weights are suitably invariant under the Frisch--Waugh--Lovell transform.}

For (ref), a natural generalization of (ref) imposes that the estimand is invariant to shifts in potential outcome paths that do not alter treatment effects:

defn[Minimally quasi-experimental] $\tau$ is minimally quasi-experimental if it is always unchanged when we replace all potential outcomes ${y}_{it}({w})$ with ${y}_{it} ({w}) + c_ {it}$ for arbitrary $c_{it} \in \R$, fixing $\bm{\pi}^*, {x}_1,\ldots, {x}_n$. For linear estimands, this is equivalent to \begin{align*} \tau &= \frac{1}{n} \sum_{i=1}^n \sum_{{w}\in \mathcal{W}} {{\omega}_{i} ({w})} {y}_ {i}({w}) for some ${\omega}_{i} ({w}) \in \R^{k\times T} $ where 0 = \sum_ {{w} \in \mathcal{W}} {\omega}_{i} ({w}) \end{align*}

(ref) is equivalent to the following linear system \[\text{For $i = 1,\ldots, n$}, \sum_{{w} \in \mathcal{W}}\pi_i^*({w}) \rho_i({w}) = 0 ,\quad \sum_ {{w} \in \mathcal{W}} \pi_i^* ({w}) = 1. \numberthis \label{eq:pop_level_irrelevance_condition} \] Since $\rho_i(w)$ is a $k\times T$ matrix and $|\mathcal W| - 1 = J$, there are $kT$ restrictions in $J$ unknowns. We call any solution an implicit design. Implicit designs are typically unique when they exist, because often the number of equations $kT$ is greater than the number of unknowns $J$.\footnote{For instance, $J + 1$ treatments generate $k = J$ contrasts; panels under staggered adoption admit fewer unique treatment times $(J + 1)$ than time horizon $T$. (ref) proves uniqueness when $T=1$.}

For a given implicit design, the corresponding implicit estimand is the following, for ${\omega}_i(\bm{\pi},{w}) \equiv \pi_i({w}) \rho_i({w})$:

align*[align* omitted — 284 chars of source]

These observations result in the following theorem formalizing how implicit designs answer (ref) and (ref), as in (ref). For a given implicit design $\bm{\pi}$, we call it proper if all $\pi_i(\cdot)$ are probability distributions. We say it generates $G_n$ if it satisfies (ref): $G_n(\bm{\pi}) = G_n$.

restatable{theorem}{thmmain} $\tau$ is minimally quasi-experimental if and only if \begin{enumerate} • Some implicit design $\bm{\pi}$ exists, is proper, and generates $G_n$, and • The true design $\bm{\pi}^*$ is equal to $\bm{\pi}$. \end{enumerate} When this happens, the estimand $\tau$ is equal to the implicit estimand under $\bm{\pi}$.

(ref) separates (ref) into an objectively computable question (ref) and a substantive question (ref). Proper implicit designs that generate $G_n$ answer (ref). If none exist, then $\tau$ cannot be minimally quasi-experimental. Judging whether the true design is plausibly $\bm{\pi}$ (ref) is context-specific. Computing implicit designs makes this judgment concrete and transparent.

Implicit designs also enable a doubly robust interpretation for the OLS estimator $\Lambda \hat\beta$, which is useful even when the regression is primarily viewed as an outcome model. Fix a hypothesized design $\bm{\pi}$ and target estimand weights $\omega_i(w)$. Given an estimated outcome regression $\hat m(w,x_i)$ meant to approximate $\E[Y(w)\mid X=x_i]$, consider the corresponding augmented inverse propensity weighting (AIPW) estimator \[ \hat\tau_{\mathrm{AIPW}} \equiv \frac{1}{n}\sum_{i=1}^n \sum_{w\in\mathcal W} \omega_i(w)\left[ \frac{\one(W_i=w)}{\pi_i(w)}\bigl(Y_i-\hat m(w,x_i)\bigr)+\hat m(w,x_i) \right]. \numberthis \label{eq:aipw} \] It is well-known that $\hat\tau_{\mathrm{AIPW}}$ is doubly robust bang2005doubly: It recovers the target estimand if either $\hat m(w,x)$ is correctly specified or the hypothesized design $\bm{\pi}$ equals the true design $\bm{\pi}^*$.

The next theorem shows that, for implicit designs, OLS is exactly such an AIPW estimator. In particular, when $\bm{\pi}$ is a proper implicit design that generates $G_n$ and $\omega_i(\cdot)$ describes its implicit estimand, choosing $\hat m$ to be the fitted values from the regression makes the AIPW formula coincide with $\Lambda\hat\beta$ in finite samples.

restatable[Double robustness of $\Lambda \hat\beta$ under (ref)] {theorem} {thmaipw} Let $\bm{\pi}$ be some proper implicit design that generates $G_n$ and let $\omega_i(w)$ be its corresponding implicit estimand. Assume that $\pi_i(w) = 0$ only if $\rho_i(w) = 0$. Then the OLS estimator $\Lambda \hat\beta$ is numerically equivalent to an AIPW estimator\[ \hat\tau_{\mathrm{AIPW}} \equiv \frac{1}{n}\sum_{i=1}^n \sum_{w\in \mathcal{W}} \omega_i(w) \bk{ \frac{\one(W_i=w)}{\pi_i(w)} (Y_i - \hat m(w, x_i)) + \hat m(w, x_i) } = \Lambda \hat\beta \equiv \hat\tau_{\mathrm{OLS}}, \] where $\hat m(w, x_i) = z(w, x_i)\hat\beta$ is the predicted value of the regression.

(ref) strengthens the dual interpretation in angrist1998 from a population identity to a numerical equivalence, and it applies to arbitrary regression specifications and general $(k,T,J)$.\footnote{(ref) is closely related to Proposition 3.2 in bruns2025augmented and to section 3 of robins2007comment. Applying Proposition 3.2 in bruns2025augmented would show that $\hat\tau_{ \mathrm{AIPW}}$ is numerically equivalent to the imputation estimator targeted to implicit estimand $ \frac{1} {n} \sum_ {i} \sum_w \omega_i (w) \hat m (w, x_i)$, and further algebra shows that this imputation estimator is numerically equivalent to the OLS coefficients $\hat\tau_{\mathrm{OLS}} = \Lambda \hat\beta$. Discussions in bruns2025augmented and robins2007comment mainly focus on cases where regressions are fit within treatment groups; (ref) allows the regression specification $\hat m(w,x)$ to be arbitrary over the entire sample. sloczynski2025covariate show related numerical equivalence results for estimators of average treatment effects. } Crucially, such a doubly robust interpretation is enabled by designs $\bm{\pi}$ that satisfy (ref).\footnote{It also exists simultaneously for all such designs, since all such designs, combined with their corresponding implicit estimand, describe the same parameter.} For a regression meant as an outcome model, a design satisfying (ref) thus endows it with an additional failsafe: The outcome model may be misspecified if (ref) holds.

Takeaways for practitioners

An extremely common workflow in practice, like blakeslee2020way, is to informally argue that treatment is “as good as randomized,” specify a regression, and interpret coefficients as causal effects---leaning on angrist2008mostly-style optimism that regressions “almost always have” outcome-model-free interpretation. This workflow leaves two gaps. First, the justification for causal interpretation typically hinges on a model of treatment assignment, yet the regression does not force researchers to articulate what that model is. Second, the regression itself chooses how heterogeneous treatment effects are aggregated, so the reported estimand need not be substantively interesting mogstad2024instrumental.

Our theoretical applications in (ref) show that these gaps matter: Some regressions that otherwise specify reasonable outcome models do not admit any treatment-based interpretation at all. On the other hand, (ref) shows that if these gaps are closed, then regression estimators are attractive as AIPW estimators---robust to its misspecification as an outcome model or to the misspecification of its implicit design.

Our results help close these gaps in this popular workflow. Practitioners under this workflow simply need to justify (ref). To this end, computing implicit designs---most straightforwardly by replacing $G_n$ with $\hat G_n$\footnote{Certain joint distribution of treatment implies that $\hat G_n = G_n$ almost surely, in which case there is no estimation error in $G_n$ to account for and (ref) is applicable as-is. See (ref). Otherwise, we prove estimation consistency in (ref). }---checks the objective implications (ref) and facilitates subjective evaluation of (ref). If (ref) passes these tests, then practitioners can safely and transparently interpret regression as quasi-experimentally estimating some causal effect.

After computing implicit designs, practitioners can evaluate whether an implicit design is plausible and consistent with economic intuition. Beyond whether the implicit design exists and is proper, a simple exercise is to verify whether the implicit design is a calibrated prediction of treatment assignment.\footnote{That is, among units with approximately $x\%$ probability of being assigned to treatment $w$, do approximately $x\%$ of those units have $W_i=w$?} The implicit design should also be consistent with substantive knowledge of the assignment mechanism. If no implicit design is plausible, the regression does not have quasi-experimental interpretation and should be interpreted as an outcome model; it can be combined with explicit treatment modeling through doubly robust estimators wager2024causal.

After determining that the implicit design is plausibly the true design, the implicit estimand informs how robust the regression is to heterogeneous treatment effects. A popular consideration is whether any unit's treatment effect contributes negatively to the estimand (poirier2024quantifying provide further diagnostics). If the implicit estimand is a weighted average treatment effect that is not substantively relevant, practitioners can target alternative estimands by reweighting the regression.\footnote{A simple recipe is to use the AIPW estimator (ref) for a user-chosen target estimand $\omega_i$ (e.g. the ATE) and a user-supplied design $\pi_i$. If some estimated implicit design satisfies (ref) and is plausible, it could serve as a candidate for propensity scores $\pi_i$ when they are unknown. The numerical equivalence of (ref) would no longer apply for a user-chosen estimand.} These diagnostics and refinements on implicit designs and estimands are illustrated in (ref) for blakeslee2020way and cervellati2024random.

Theoretical applications and examples

This section applies our framework to assess (ref) across a wide swath of regression specifications and discusses them in self-contained vignettes. To emphasize, our results essentially reduce the problem to computing the potential weights and the set of implicit designs. This unifies results across starkly distinct settings.

Several specifications have known causal interpretations under specific designs angrist1998,goldsmith2022contamination,imbens2009recent,lin2013agnostic,kline2011oaxaca,athey2018design. Applied to these specifications, the implicit designs recover these results and supply a converse ((ref)). Specifically, we show that quasi-experimental interpretation analyzed in these settings is tenable only under those designs assumed in the literature. These specifications target exactly those estimands found in the literature. Thus (ref) underpins much of the existing work. Imposing it establishes that sufficient conditions in the literature are necessary---that is, there is no weaker or alternative set of conditions on the design to prove the results in the literature.

Applying our framework also reveals new theoretical results: Two classes of specifications admit quasi-experimental interpretations only under stringent conditions. First, cross-sectional regressions with $W \times x$ interactions qualify essentially only when $x$ is saturated (discrete) or $W$ is randomly assigned independently of $x$ ((ref)). Second, TWFE regressions with time-varying covariates or imbalanced panels lack implicit designs whenever treatment timing covaries with covariates or observation patterns ((ref)). These results show that regressions that otherwise specify reasonable outcome models can have no quasi-experimental interpretation.

A unified analysis of quasi-experimental interpretation in regression

Assume throughout that the population Gram matrix is invertible.

restatable{theorem}{thmzoo} We compute the implicit designs and estimands of the regression specifications (1)--(5) described in (ref). In every specification, the implicit design exists uniquely. The implicit design generates $G_n$ regardless of whether $\bm{\pi}=\bm{\pi}^*$, for all specifications except ((ref)). \begin{enumerate} • \begin{enumerate} • $\pi_i = x_i'\delta$ for $\delta = \pr{\frac{1}{n}\sum_{i=1}^n x_ix_i'}^{-1} \frac{1}{n} \sum_{i=1}^n \pi_i^* x_i$$\pi_i^* = \pi_i$ if and only if $\pi_i^* = x_i'\delta$$\omega_i \equiv \omega_i(\bm{\pi}, 1) = -\omega_i(\bm{\pi},0) = \frac{\pi_i (1-\pi_i)} {\frac{1}{n} \sum_{i=1}^n \pi_i (1-\pi_i)}$. When $\pi_i^* = \pi_i$, $\omega_i \ge 0$. \end{enumerate} • \begin{enumerate} • $\pi_i(j) = x_i'\delta_j$ for $\delta_j = \pr{\frac{1}{n}\sum_{i=1}^n x_ix_i'}^{-1} \frac{1}{n} \sum_{i=1}^n \pi_i^*(j) x_i$$\pi_i^* = \pi_i$ if and only if $\pi_i^*(j) = x_i'\delta_j$ for all $j \in [J]$ • The implicit estimand is shown in (ref). This estimand is generally contaminated (that is, $\omega_{ij}(\bm{\pi}, \ell) \neq 0$ for some $j \in [J]$ and $\ell \not\in \br{0,j}$). \end{enumerate} • \begin{enumerate} • The implicit design equals the mean of $\pi_i^*$ among the units with the same $x_i$-value • $\pi_i^* = \pi_i$ if and only if $\pi_i^*$ is the same for all units with the same $x_i$-value • The implicit estimand is the ATE. That is, $\omega_i = \omega_i(\bm{\pi}, 1) = -\omega_i(\bm{\pi},0) = 1$. \end{enumerate} • \begin{enumerate} • $\pi_i = \frac{\delta_0 + (x_i-\bar x)'\delta_1}{1+\delta_0 + (x_i-\bar x)'\delta_1 }$, where $\delta_0, \delta_1$ are equal to the population weighted least-squares coefficients of $\pi_i^*/(1-\pi_i^*)$ on $x_i-\bar x$ and a constant, weighted by $1-\pi_i^*$$\pi_i^* = \pi_i$ if and only if $\pi_i^*/(1-\pi_i^*) = \delta_0 + \delta_1'(x_i-\bar x)$ • When $\bm{\pi}^* = \bm{\pi}$, the implicit estimand is the ATT: $\omega_i = \omega_i (\bm{\pi}, 1) = \frac{\pi_i}{\frac{1}{n} \sum_{i=1}^n \pi_i}$. \end{enumerate} • \begin{enumerate} • The implicit design is constant in $i$ and is unique, $\pi_i({w}) = \frac{1} {n} \sum_ {i=1}^n \pi_i^* ({w})$$\pi_i^* = \pi_i$ if and only if $\pi_i^*$ is the same for all $i$ • The implicit estimand is shown in (ref), which matches Theorem 1(ii) in athey2018design under staggered adoption.\footnote{One might wish to further impose that the post-treatment weights are nonnegative (i.e., ${\omega}_{it}(\bm{\pi}^*, {w}) \ge 0$ if ${w}_t = 1$). Failure of this condition implies that post-treatment units are severely used as comparisons for newly treated units, echoing the “forbidden comparison” issue in the recent difference-in-differences literature roth2023s,borusyak2024revisiting,de2020two,goodman2021difference. (ref) shows that when $\mathcal{W}$ only has two elements and includes a never treated unit, all weights post treatment are non-negative, but such forbidden comparisons are possible in all other cases.} \end{enumerate} \end{enumerate}

(ref) computes implicit designs and estimands for several specifications individually analyzed in the literature. Simply examining (ref) shows the implicit design exists, is unique, and matches the form studied; the implicit estimand matches as well. (ref) is thus a set of converses to the existing results---the regression estimand satisfies (ref) only if $\bm{\pi}^* = \bm{\pi}$ and the target causal effect is the implicit estimand. These necessity results are new, to our knowledge, except for (ref)((ref)). These calculations, combined with (ref), also immediately imply regression estimators are equivalent to AIPWfor the implicit estimand---regardless of (ref)---for all but ((ref)).

{

landscape\begin{table}[h] \begin{tabularx}{1\linewidth}{@ l l l l X @} \toprule \# & Setting & Specification & Contrast & Additional conditions\\ \midrule (1) & $k=T=J=1$ & $Y_i=\tau W_i + x_i'\gamma +\epsilon_i$ & $\tau$ & $x_i$ includes a constant \\ (2) & $k=J, T=1$ & $Y_i=\sum_{j=1}^J \tau_j W_{ij} + x_i'\gamma + \epsilon_i$ & $(\tau_1,\ldots, \tau_J)$ & $x_i$ includes a constant. $\mathcal W = \br{0,\ldots, J}$, $W_ {ij}=\one (W_i = j)$ \\ (3) & $k=T=J=1$ & $Y_i = \alpha_0 + \gamma_1'x_i + \tau W_i + W_i (x_i-\bar x)'\gamma_2 + \epsilon_i$ & $\tau$ & $x_i$ saturated for some discrete covariate $x_i^*$ taking values in $ \br{0,\ldots, L}$: $x_i = [x_ {i1},\ldots, x_ {iL}]'$ for $x_ {i\ell} = \one (x^*_i = \ell)$, $\bar x = \frac{1}{n} \sum_i x_i$ \\ (4) & $k=T=J=1$ & $Y_i = \alpha_0 + \gamma_1'x_i + \tau W_i + W_i(x_i - \bar x_1)'\gamma_2 + \epsilon_i$ & $\tau$ & $\bar x_1 = \frac{\sum_i \pi_i^* x_i} {\sum_i \pi_i^*}$ \\ (5) & $T > 1$ & $Y_{it} = \alpha_i + \mu_t + \tau W_{it} + \epsilon_{it}$ & $\tau$ & $\mathcal W \subset \br{0,1}^T$ is the set of treatment paths. The nonzero elements of $\mathcal W$ are linearly independent vectors whose span excludes $1_T = (1,\ldots,1)'$. This condition is satisfied by staggered adoption that excludes always-treated units. \\\bottomrule \end{tabularx} \caption{Regression specifications analyzed in (ref)} \begin{proof}[Notes] (1) is discussed in angrist1998 and section 2.1 of blandhol2022tsls; (2) is discussed in goldsmith2022contamination; (3) is discussed in miratrix2013adjusting,imbens2009recent,lin2013agnostic, among others; (4) is discussed in kline2011oaxaca; (5) is discussed in athey2018design. \end{proof} \end{table}

}

Rather than detail every vignette, we highlight two notable findings. First, (ref)((ref)) and ((ref)) leave a few puzzles, which (ref) resolves. Both regressions pick a contrast from the interacted specification $Y_i=\gamma_0+\gamma_1'x+\tau_0 W_i+\tau_1'x_iW_i+\epsilon_i$.\footnote{ (ref)((ref)) takes $\tau_0 + \tau_1' \bar x$ while (ref)((ref)) takes $\tau_0 + \tau_1' \bar x_1$.} Curiously, ((ref)) requires saturated covariates; ((ref)) does not. Moreover, (ref)((ref)) is asymmetric. If we flip treatment and control, (ref)((ref)) would show that the average treatment effect on the untreated (ATU) estimand is minimally quasi-experimental only if the reciprocal propensity odds $ (1-\pi_i^*)/\pi_i^*$ is linear in $x_i$. Thus, worryingly, the same specification yields ATT and ATU interpretations under different designs.

Second, (ref)((ref)) shows the TWFE estimand fails to be minimally quasi-experimental unless treatment timing is fully randomized---which athey2018design study. (ref) extends this by showing TWFE's quasi-experimental interpretation is additionally fragile. (ref) extends the analysis to one-way FE and event-study designs.

Interactions and impossibility of regression estimation of ATE

Assume $T=J=1$ and split $x_i$ into subvectors $x_{1i},x_{2i}$ (possibly overlapping). Consider the specification \[Y_i = \gamma_0 + \tau_0 W_i + \tau_1' W_i x_{1i} + \gamma_1'x_{2i} + \epsilon_i. \numberthis \label{eq:interaction} \] Viewed as an outcome model, $\tau_0$ is the treatment effect for a baseline covariate value, and $\tau_1$ captures how treatment effect varies with $x_1$. One might hope that even without the outcome model, $\tau = (\tau_0, \tau_1')$ retains causal interpretation in a more flexible manner than the specification (ref) without interactions. This hope generally fails: Quasi-experimental interpretation of $\tau$ necessitates that both $\pi_i^*$ and $\pi_i^* x_ {1i}$ be linear in $x_ {2i}$. When this fails, some contrast $\tau_0 + \tau_1'x_{1}$ does not satisfy level independence.\footnote{This result was novel at the time of a working paper draft of this paper (arXiv:2407.21119v2, January 13, 2025); concurrent and independent work by zhao2025interacted (arXiv:2502.00251, February 1, 2025) provides a similar result.}

restatable{prop}{propforbidden} Consider the specification (ref) and let $\tau = (\tau_0, \tau_1')'$ be the coefficients of interest. Then the corresponding implicit design exists if and only if, for some conformable matrices $(\Gamma_0, \Gamma_1)$ and all $i$, $(\delta_0 + \delta_1'x_ {2i}) x_ {1i} = \Gamma_0 + \Gamma_1 x_{2i}$, where $\delta_0, \delta_1$ are population projection coefficients of $\pi_i^*$ on $x_ {2i}$. When this happens, the unique implicit design is $\pi_i = \delta_0 + \delta_1' x_ {i2}$. Therefore, if $\tau$ satisfies (ref), then $\pi_i^* = \delta_0 + \delta_1'x_{2i}$ and $\pi_i^*x_{1i} = \Gamma_0 + \Gamma_1 x_{2i}$.

The necessary condition for interpreting $\tau$ as minimally quasi-experimental is that both the propensity score $\pi_i^*$ and its interaction with the covariates $\pi_i^* x_ {1i}$ are linear functions of $x_{2i}$. When $x_{1i}$ is included in $x_{2i}$, this condition is unlikely to hold in general, as $\pi_i^* x_{1i}$ would involve nonlinear transformations of $x_ {1i}$ and thus cannot be linear. This condition does hold if $\pi_i^*$ is constant or if $x_{1i}$ represents a saturated categorical variable and $x_ {2i}$ contains all other covariates interacted with $x_ {1i}$.\footnote{That is, $x_ {1i}$ contains mutually exclusive binary random variables, and $x_{2i}$ contains $x_ {1i}$, some set of other covariates $x_{3i}$, and all interactions $x_{3ik}x_{1i\ell}$.}

Why can we not interpret $\tau_0 + \tau_1'x_ {1i}$ as a linear approximation of the conditional average treatment effect? One could think of (ref) as two regressions, one on the treated $W=1$ and one on the untreated $W=0$. Both regressions are indeed best linear approximations to $\E [Y(1) \mid x, W=1]$ and $\E[Y(0) \mid x, W=0]$, which are equal to the mean potential outcomes $\E [Y (1) \mid x], \E[Y(0) \mid x]$ under unconfoundedness. The contrast $\tau_0 + \tau_1'x_1$ is then the difference of the fitted values of these two regressions. However, the two regressions are best linear approximations with respect to different distributions of the covariates ($x_i \mid W=1$ vs. $x_i \mid W=0$). Thus, their difference is not a best linear approximation to the conditional average treatment effect under any particular distribution of $x$. Shifting $Y(1)$ and $Y(0)$ by the same arbitrary amount therefore causes asymmetric behavior in the two regressions, leading to a failure of level irrelevance.

When $x_{1i} = x_{2i} = x_i$, this result supplements (ref)((ref))--((ref)) by showing different contrasts necessitate incompatible designs.\footnote{(ref) shows formally that requiring (ref) for the contrast $\tau_\lambda = \lambda_0 \tau_0 + \lambda_1'\tau_1$ in this regression implies implicit designs $\pi_\lambda$, generally fractional-linear in $x_i$, that depends on the contrast $\lambda_0, \lambda_1$.

In particular, these results are relevant for the ATE contrast $\tau_0+\tau_1'\bar x$, which is separately studied in Theorem 1 in chattopadhyay2023implied. chattopadhyay2023implied show that if we insist that $\tau_0+\tau_1'\bar x$ equal the ATE, then we need both propensity odds and reciprocal odds to be linear. In contrast, we show that if $\tau_0+\tau_1'\bar x$ is only required to be some treatment effect contrast (not necessarily the ATE), the implicit design exists but is fractional-linear $\pi_i = \frac{\theta_{0} + \theta_1'(x-\bar x)}{1-\Gamma_2'(x-\bar x)}$. However, the requirement (ref) that $\pi_i = \pi_i^*$ then imposes additional (unpleasant) restrictions on $ (\theta_0, \theta_1, \Gamma_2)$, formalized in (ref).} Insisting on all such contrasts being minimally quasi-experimental imposes a knife-edge condition for the design. Without saturated covariates, (ref)((ref)) shows that particular contrasts (e.g., the ATT) maintains causal interpretation, at the expense of others.

Taken together, interacted regressions are less robust in terms of (ref) than the simple regression (ref), contrasting with the qualitative takeaway in lin2013agnostic and negi2021revisiting. The uninteracted regression introduces variance weighting for the estimand, but maintains validity under a simple design. The interacted regression removes this weighting with saturated covariates but loses quasi-experimental interpretation in general.

Is there a simple regression that targets the ATE under linear propensity scores? Unfortunately, the next proposition shows that the answer is no, at least not with specifications that are linear in $[1, x_i, W_i, W_ix_i]$.\footnote{One could estimate the uninteracted regression and weigh by $1/ (\pi_i(1-\pi_i))$ to remove the variance weighting, but this approach requires estimating $\pi_i$ separately. } As a result, targeting the ATE under the same implicit design as (ref) necessitates moving beyond regression estimators.

restatable[No simple regression estimates the ATE under linear design]{prop} {propnoatereg} Let $n \ge 3$. Let $W_i \in \br{0,1}$, covariates $x_i \in \R^d$, and $y_i(\cdot) \in \R$. Suppose the true design is linear $\pi_i^* = \delta_0 + \delta_1'x_i$ for some $\delta_0 \in \R, \delta_1 \in \R^d$. There is no regression $ (\Lambda, z(x, w))$---where $\Lambda$ may\footnote{This is to accommodate for estimands like the model-based ATT, where we may consider contrasts that depend on $\bar x_1 = \sum_i \pi_i^* x_i/\sum_i \pi_i^*$} depend on $x_{1:n}, \pi^*_{1:n}$---such that: \begin{enumerate} • (Regression is linear in covariates) For all $m$ and all $w$, the $m$\th entry of $z (x_i, w)$ is of the form $a_m(w) + b_m(w)'x_i$ for some fixed conformable $a_m (\cdot), b_m(\cdot)$. • ($\Lambda\beta$ is the ATE) The corresponding estimand $\Lambda\beta \in \R$ is equal to the ATE, regardless of the configuration of $d, x_{1:n}, \delta_0, \delta_1$ (such that $\pi_i^* \in [0,1]$ for all $i$). \end{enumerate}

The fragility of quasi-experimental TWFE

(ref)((ref)) shows TWFE is minimally quasi-experimental only under totally randomized treatment paths. We now show that adding time-varying covariates often destroys even that.

restatable{prop}{proptimevarying} Assume $\mathcal{W} \subset \br{0,1}^T$. Consider the regression ${Y}_{it} = \alpha_i + \gamma_t + \tau {W}_{it} + \delta'{x}_{it}$ where $\tau$ is the coefficient of interest. Let $\beta_{w\to x}$ be the population projection coefficient of $W_{it}$ on $x_{it}$ under $\bm{\pi}^*$, with individual and time fixed effects. If an implicit design exists, then, for $x_i \in \R^{T \times \dim(x_{it})}$ that stacks the covariates $x_ {it}$, \[\pr{{x}_i - \frac{1}{n}\sum_{j=1}^n {x}_j}\beta_{w\to x} \in \Span(\mathcal{W} \cup \br{1_T})\, \text{ for all $i=1,\ldots,n$}.\] When $\beta_{w\to x} = 0$, if $\mathcal{W}$ contains linearly independent vectors whose span excludes $1_T$, the implicit design is uniquely equal to $\pi_i ({w}) = \frac{1}{n} \sum_{i=1}^n \pi_i^*({w})$ for all $i$ as in (ref)((ref)).

An implicit design exists only if a linear combination of demeaned covariates lies in the span of $\mathcal{W}$ and $1_T$ for every unit. This condition arises because we essentially need that the mean treatment $\E[{W}_i] = \sum_{{w} \in \mathcal{W}} \pi_i^* ({w}) {w}$ is exactly described by two-way fixed effects with time-varying covariates, analogous to the intuition for (ref). This then restricts the space of covariates, since they need to generate vectors that lie in the linear span of $\mathcal{W}$.

With staggered adoption, $\Span\pr{\mathcal{W} \cup \br{1_T}}$ is the subspace of vectors that are piecewise constant between adjacent adoption dates. This subspace is highly restrictive if there are relatively few adoption dates. If $\beta_ {w\to x} \neq 0$, it is thus knife-edge that $ \pr{{x}_i - \frac{1} {n}\sum_ {j=1}^n {x}_j} \beta_{w\to x}$ happens to be located in that subspace, unless columns of ${x}_i$ happens to be piecewise constant over $t$ as well.\footnote{This is plausible if the time-varying covariates are interactions of fixed covariates with the time fixed effects (${x}_ {it}'\delta = x_i'\delta_t$). (ref) shows that for this specification, causal interpretation is possible necessarily under linear generalized propensity scores $\pi_i({w}) = \delta_0({w}) + \delta_1({w})'x_i$.} On the other hand, if $\beta_ {w \to x}$ is zero under $\bm{\pi}^*$, including the covariates makes no difference to the coefficient on $W_{it}$. Thus TWFE with time-varying covariates rarely retains a quasi-experimental interpretation. Researchers using such a specification either believe that the covariates do not affect treatment assignment and are irrelevant for identification, or they are embedding outcome modeling assumptions.

Finally, a similar fragility inflicts regressions with imbalanced panels. Such a regression only has a quasi-experimental interpretation when the missingness pattern is uncorrelated with the treatment assignment pattern, in which case the design must again be total randomization of treatment paths. We detail this result in (ref).

Extension: Two-stage least-squares

Similar ideas to (ref) extend to two-stage least-squares (TSLS): We can use level irrelevance to recover some design---now a distribution of the instrument $W_i$---under which TSLS estimands have a causal interpretation in the instrument $W$ (cf. intent-to-treat effects). Interestingly, the implicit estimand also provides necessary conditions on compliance behavior for TSLS to estimate properly weighted causal effects in terms of the endogenous treatment.

For instance, examining the implicit estimand for a binary treatment, binary instrument TSLS regression recovers (strong) monotonicity as a necessary condition imbensangrist,sloczynski2020should. Doing so for TSLS with multiple treatments yields a compliance restriction in bhuller20242sls. These results are recovered simply by enumerating which compliance types for each unit are consistent with the implicit estimand assigning proper weights to said unit's potential outcomes in the treatment.

Consider the following TSLS specification of a scalar outcome on a covariate transform \[ Y_i = t(D_i, x_i)'\beta + \epsilon_i, \] instrumenting $t(D_i, x_i)$ with $z(W_i, x_i)$. Here, $D_i = d^*_i(W_i) \in \mathcal{D}$ is the endogenous treatment, $d^*_i(\cdot)$ is the compliance type for unit $i$, and $t (\cdot, \cdot), z (\cdot, \cdot)$ are again known transforms. Assume the exclusion restriction holds so that $y_i(d_i^* (w), w) = y_i(d_i^*(w))$. In this notation, a binary treatment, binary instrument TSLS regression can be represented by $t(D_i, x_i) = [1, D_i]'$ and $z (W_i, x_i) = [1, W_i]'$.

We extend steps (ref)--(ref) in (ref). For (ref), define the TSLS estimand $\tau = \Lambda\beta$ as \[ \tau = \Lambda \pr{G_{tz} G_{zz}^{-1} G_{zt}}^{-1} \pr{G_{tz} G_{zz}^{-1} \frac{1} {n} \sum_{i=1}^n \E_{W_i\sim \pi_i^*}[z(W_i, x_i) y_i(W_i)]}, \] where $ G_{tz} \equiv \frac{1}{n}\sum_{i=1}^n \E_{W_i \sim \pi_i^*}\bk{ t(d_i^*(W_i), x_i) z(W_i, x_i)' }$, $G_{zt} \equiv G_{tz}'$, $G_{zz} \allowbreak \equiv \allowbreak \frac{1} {n}\sum_ {i=1}^n \allowbreak\E_{W_i \sim \pi_i^*}\bk{ z(W_i, x_i) z(W_i, x_i)'} $. This representation simply replaces all averages in the TSLS estimator with expectations over the instrument $W_i$. Let $H_n \equiv \pr{G_{tz} G_{zz}^{-1} G_{zt}}^{-1} G_{tz} G_{zz}^{-1}$. $H_n$ is the analogue of the inverse Gram matrix $G_n^{-1}$.\footnote{Indeed, if $t(d_i(W_i), x_i) = z (W_i, x_i)$ so that the TSLS specification is equivalent to OLS, then $H_n$ is exactly the inverse Gram matrix.} Like $G_n^{-1}$, $H_n$ is known in population and consistently estimable in sample. Thus, by the same reasoning, we treat $H_n$ as known.

Next, for (ref), write $\tau$ in the form of (ref): \[ \tau = \frac{1}{n} \sum_{i=1}^n \sum_{w \in \mathcal{W}} \pi_i^*(w)\Lambda H_n z(w, x_i) y_i (w) \equiv \frac{1}{n} \sum_{i=1}^n \sum_{w \in \mathcal{W}} \pi_i^*(w) \underbrace{\rho_i (w)}_{k \times 1} y_i(d_i(w)). \] We define potential weights analogously: $\rho_i(w) \equiv \Lambda H_n z(w, x_i)$. For (ref)--(ref), the requirement that $\tau$ is minimally quasi-experimental ((ref)) continues to be reasonable. If $\tau$ were a comparison of different potential outcomes $y_i (d)$, then it should be invariant to shifting all $\br{y_i(w), w\in\mathcal{W}}$ by arbitrary $c_i$. Maintaining this restriction again yields (ref) for $\pi_i^*(w)$, whose solutions we continue to call implicit designs. They continue to be plausible candidates for the true design $\pi_i^*(\cdot)$ in the sense of (ref).

Just-identified TSLS specifications have enough equations\footnote{For an TSLS specification to be non-collinear, an instrument that takes $J+1$ values can support $k\le J$ endogenous coefficients of interest. Since $\pi_i(\cdot)$ is a $J$-dimensional unknown vector, we need $k \ge J$ restrictions to have a unique implicit design. } to pin down an implicit design $\pi_i (\cdot)$. If there are more distinct instrument values than coefficients of interest, then we may have too few restrictions on $\pi_i^*(w)$ from $\tau$ alone. However, it may be reasonable to also impose level irrelevance for certain first-stage coefficients, which would add more restrictions to recover a unique implicit design.

The estimand for TSLS---in terms of $y_i(k)$ rather than $y_i(d_i^*(w))$---depends on units' unknown compliance types $d_i^* (\cdot)$. Therefore, interpreting the estimand as a causal effect of the treatment $d$ implicitly restricts compliance patterns. This can be operationalized as follows. Given an implicit design $\pi_i (\cdot)$, the corresponding implicit estimand $\tau$ can be written as a weighted sum of individual potential outcomes, which can be grouped into treatment conditions:

align*[align* omitted — 541 chars of source]

(ref) represents the estimand as an aggregation of $w$-on-$y$ causal effects. (ref) then groups together $w$ values that lead to the same $d_i^*(w) = k$, thereby translating (ref) to $d$-on-$y$ effects. In (ref), the weight on the $k$\th treatment is ${\omega}^*_i(k; \bm{\pi}, d_i^*) \equiv \sum_{w: d_i^*(w)=k} {\omega}_i(w; \bm{\pi})$, which is known given $d_i^* (\cdot)$. If $\tau$ were to have a causal interpretation, we can then enumerate all compliance types $d_i^*$ for each unit and check which ones lead to weights ${\omega}^*_i (k; \bm{\pi}, d_i^*)$ that are consistent with the causal interpretation.

To illustrate, consider a particular class of TSLS specifications: For $x_i$ that includes a constant, consider a just-identified specification with $J+1$ values of an unordered treatment $\mathcal D = \br{0,\ldots, J}$

align*[align* omitted — 155 chars of source]

In this TSLS specification, the coefficients of interest are $\tau = (\tau_1,\ldots, \tau_{J})'$, where $\tau_k$ is the coefficient on $\one(d=k)$, meant to capture the causal effect of $d=k$ relative to $d=0$.

Examining entries in (ref), we have \[ \tau_k = \frac{1}{n} \sum_{i=1}^n \sum_{k' = 1}^{J} \omega_{i}^{(k,k')}(d_i^*) (y_i(k') - y_i(0)) \quad \text{ for } \quad \omega_i^{(k, k')} \equiv ({\omega}^*_{i}(k'; \bm{\pi}, d_i^*))_k. \numberthis \label{eq:estimand_tsls} \] If $\tau_k$ is to be interpreted as a causal effect of $d=k$ relative to $d=0$, then we should at least restrict $\omega_i^{(k,k)} \ge 0$ and $\omega_i^{(k,k')} = 0$ for $k \neq k'$. If so, $\tau_k$ equals a convex aggregation of $y_i(k) - y_i (0)$ that is not contaminated by treatment effects of some other arm $y_i(\ell) - y_i (0)$. If this is true and if $\bm{\pi}^* = \bm{\pi}$, following bhuller20242sls, we say that TSLS assigns proper weights.\footnote{When $J=2$, $\tau$ is minimally quasi-experimental and assigns proper weights if and only if it is weakly causal in the sense of blandhol2022tsls.}

Given $\omega_i^{(k,k')}(\cdot)$, for each unit, we can then enumerate all compliance types $d_i(\cdot)$ and retain those consistent with proper weights. Analogous to implicit designs, we refer to each element of the following set as an implicit compliance profile: For $\mathcal D = \br{0,\ldots, J}$,\[ \br{ (d_1(\cdot),\ldots, d_n(\cdot)): \text{ for all $i$, $k\neq k' \in \mathcal D$}, \omega_i^ {(k, k)} (d_i) \ge 0 \text{ and } \omega_i^{(k, k')}(d_i) = 0 } \numberthis \label{eq:implicit_compliance_profile}. \] The following proposition summarizes these results:

restatable{prop}{propivmain} In TSLS, $\tau$ is minimally quasi-experimental if and only if \begin{enumerate} • An implicit design $\bm{\pi}$ exists • $\bm{\pi}^* = \bm{\pi}$. \end{enumerate} Additionally, $\tau$ from the specification (ref) assigns proper weights if and only if the following holds for the implicit estimand under $\bm{\pi}$: \begin{enumerate}[resume] • An implicit compliance profile $d_1(\cdot), \ldots, d_n(\cdot)$ in (ref) exists • Some implicit compliance profile $d_1(\cdot), \ldots, d_n(\cdot)$ is equal to $d_i^*(\cdot),\ldots, d_i^*(\cdot)$. \end{enumerate}

Like (ref), (ref) separates requirements for causal interpretation into objectively and subjective components. We can directly compute items (1) and (3) since the potential weights, implicit design, and implicit estimand are known in the population. Results from this computation are plausible candidates for items (2) and (4)---if no such candidate is found, then causal interpretation must be rejected.

These computations are informative. To illustrate, simply computing the implicit design and compliance profiles recovers necessary conditions to several recent results in the instrumental variables literature. To introduce, we first give terminology to compliance patterns.

defn[Compliance restrictions] \begin{itemize} • With $J+1=2$, we say that a profile $d_1(\cdot), \ldots, d_n (\cdot)$ satisfies strong monotonicity if either $d_i(1) \ge d_i(0)$ for all $i$ or $d_i(1) \le d_i (0)$ for all $i$. • With $J+1>2$, for $k = 1,\ldots, J$, we say that $d(\cdot)$ is a $k$-always taker if $d (\cdot) = k$; it is a $k$-never taker if $d(\cdot) \neq k$; otherwise we say $d(\cdot)$ is a $k$-complier. We say $d(\cdot)$ is a full complier if it is a $k$-complier for all $k$. • We say that a compliance profile $ (d_1 (\cdot),\ldots, d_n(\cdot))$ satisfies common compliance if for any $k = 1,\ldots, J$ and any two $k$-compliers $d_i (\cdot), d_j(\cdot)$, we have $d_i(w) = k$ if and only if $d_j(w) = k$. • We say that a compliance profile satisfies extended monotonicity if there exists some permutation $f (\cdot)$ of the instrument values $\br{0,\ldots, J}$ such that, for all $i$, either (i) for all $w$, $d_i (f (w)) \in \br{0,w}$ or (ii) $d_i(\cdot)$ is constant.\footnote{For three instrument values, up to permutation of the instruments, extended monotonicity limits $d (\cdot)$ to one of six types $(d(0), d(1), d(2)) \in \br{(000), (111),(222), (010),(002), (012)}$---for, respectively, never-taker, 1-always-taker, 2-always-taker, 1-complier, 2-complier, or full complier bhuller20242sls. This condition is a generalization of Assumption 3 in behaghel2013robustness, who call this assumption extended monotonicity. Indeed, the condition is equivalent to that, for all $i$, $w\neq 0$ and $w', w'' \neq w$, $ \one\br{d_i(f(w)) = w} \ge \one\br{d_i(f(w')) = w} = \one\br{d_i(f(w'')) = w}. $.} \end{itemize}
restatable{prop}{proptsls} Consider the TSLS specification in (ref), \begin{enumerate} • The unique implicit design satisfies $\pi_i(j) = x_i'\delta_j$ for $\delta_j = \pr{ \sum_{i=1}^n x_ix_i'}^{-1} \sum_{i=1}^n \pi_i^*(j) x_i$. • When $\bm{\pi}^* = \bm{\pi}$, the implicit compliance profiles relative to the implicit design satisfy: \begin{enumerate} • When $J+1=2$, all implicit compliance profiles satisfy strong monotonicity. • When $J+1>2$ and $x_i$ is a constant, all implicit compliance profiles satisfy {common compliance}; all implicit compliance profiles containing a full complier satisfy extended monotonicity. \end{enumerate} \end{enumerate}

(ref) recovers several results for TSLS. (ref)(1) and (2)(a) recover the necessary direction for Corollary 3.4 in sloczynski2020should and Theorem 1 in blandhol2022tsls: With binary treatment, monotonicity is required for interpreting the TSLS coefficient causally, in the sense that it assigns proper weights.\footnote {Theorem 1 in blandhol2022tsls imposes exogeneity and monotonicity and finds that $\tau$ is minimally quasi-experimental and has proper weights if and only if $\pi_i^*$ is linear. (ref)(1) and (2)(a) in turn show that if $\tau$ is minimally quasi-experimental and has proper weights, then linear propensity scores and monotonicity are satisfied (see (ref) for details). Likewise, Corollary 3.4 in sloczynski2020should shows that strong monotonicity implies proper weights, but not the converse. } Without covariates, this is a converse to imbensangrist.

(ref)(2)(b) recovers---and corrects---Propositions 5 and B.1 in bhuller20242sls. Proposition 5 in bhuller20242sls claims that if TSLS assigns proper weights, then compliance satisfies extended monotonicity---that is, up to permutation of the instrument values, we can think of instrument $w$ as an encouragement to take up treatment $w$ from $0$, with no effect on other treatment takeup nor substitution from other $w'\neq w$ to $w$. Unfortunately, just assuming TSLS assigns proper weights does not suffice for this conclusion.\footnote{See (ref) for a counterexample. We are grateful to Henrik Sigstad for discussion. } Instead, the essence of their argument implies that compliance profiles satisfy common compliance; their conclusion in turn stands if it is known that some full complier exists. Both implications are captured by (ref)(2)(b).

Empirical illustration of diagnostics

This section illustrates how the framework can be used in applied work to make quasi-experimental interpretations of regressions more transparent. Our practical recommendation is a simple workflow:

enumerate• Compute the implicit design and check whether it is proper (ref). • Towards (ref), evaluate whether the resulting assignment model is substantively and statistically plausible---focusing on calibration, functional-form plausibility, and economic plausibility. • Conditional on a plausible design, inspect the implicit estimand to understand what the regression is weighting (including the prevalence and concentration of negative weights) and, when the implied contrast is not substantively aligned with the question of interest, retarget alternative estimands by reweighting.

We organize the empirical illustrations around different parts of this workflow. cervellati2024random provide a setting in which the true design is known, so the implicit design can be benchmarked directly; in their setting, evaluating implicit designs complements balance tests and diagnoses concerns on sample selection. blakeslee2020way provide a setting in which the true design is unknown, and we use it to walk through the full workflow above.

Diagnostics for the implicit design

Known design

cervellati2024random study Italian elections. Parties in these elections are organized into coalitions at the ballot box. Due to a quirk of ballot design, the party in the middle of a coalition on the ballot paper (in the focal position) receives more votes, all else equal. Since ballot order is random, the authors use this feature to study the effect on outcomes, including fiscal spending towards various policies. The causal identification is explicitly framed as coming from this random assignment.

The true design here is known. The authors define the focal position as the middle position if the coalition has an odd number of parties and the middle two positions if the coalition has an even number. Thus, if a coalition has $x$ parties, the probability of being treated is $1/x$ for odd $x$ and $2/x$ for even $x$. In this setting, therefore, both (ref) and (ref) can be directly tested.

Table IV in cervellati2024random studies the impact of this focal treatment on fiscal policy for major political parties. Since only winning coalitions control fiscal policy, the authors restrict to “ruling coalitions that include each of the major parties” (p.1570--1571, cervellati2024random) and consider a specification like \[ Y_{i} = \tau W_i + x_{i}'\gamma + \tilde x_i'\tilde \gamma + \epsilon_{i}, \numberthis \label{eq:cervellati_spec} \] where $i$ indexes a party in a given municipal election. $Y_{i}$ denotes budgetary item on the salient policy area of each party for the legislature session after the election of $i$, $W_i$ denotes the focal position treatment, $x_i$ denotes saturated dummies on the number of parties in the same coalition as party-election $i$, and $\tilde x_i$ denote other covariates.\footnote{In cervellati2024random, Table IV, column (3) and equation (E3) consider a panel version of (ref): \[ Y_{it} = \tau W_i + x_{i}'\gamma + \tilde x_{it}'\tilde \gamma + \epsilon_{it} \] where $t$ indexes calendar year. The only time-varying covariates in $\tilde x_{it}$ are year-of-legislature fixed effects and year fixed effects interacted with party-type fixed effects. Since panel specifications with time-varying covariates are unlikely to have implicit designs per (ref), we aggregate to a cross-sectional setup, by replacing time-varying covariates with party-type fixed effects interacted with election-year fixed effects, which do not vary within $i$. Doing this aggregation changes the Table IV (3) coefficient and standard errors from 0.058 (0.025) to 0.62 (0.028). In the replication files, the authors redefine a small proportion of treatment--—when any main party in a winning coalition is treated, any other main party in that coalition is considered treated as well. We are unable to find documentation of this change in the paper. If we further use instead the treatment variable before this redefinition, then the same coefficient obtains an estimate of 0.047 (0.029). }

figure[figure omitted — 1,131 chars of source]

We use this setting to illustrate steps (1)--(2): whether an implicit design is plausibly $\bm{\pi}^*$. (ref) plots the estimated implicit designs from this specification. In the full sample (left panel), the implicit design computed from the specification tracks the benchmark assignment probabilities almost perfectly. However, this changes once we restrict to major parties in winning coalitions. After this restriction, the implicit design no longer resembles $\bm{\pi}^*$. This raises concerns about sample selection, especially since whether a coalition wins is plausibly affected by treatment.

figure[figure omitted — 739 chars of source]

To investigate, we test the hypothesis that the implicit design in the selected sample is equal to the true design. Within each coalition, we can redraw placebo treatment statuses by permuting the ballot order. The distribution of any test statistic across these draws is then equal to its distribution under the true design. We choose the test statistic to be the prediction error for the true design $T = \pr{ \frac{1}{n} \sum_i (\hat\pi_i - \pi_i^*)^2}^{0.5}$. Reassuringly, (ref) implements this test and finds at worst suggestive evidence against the null. It is thus plausible that the divergence in (ref) is an artifact of noise.

This exercise complements and is consistent with the covariate balance tests in cervellati2024random. Covariate balance tests directly inform internal validity when viewing regression as an outcome model.\footnote{Under a linear outcome model, imbalance in $y(0)$ across treatment and control can only arise due to imbalance in $\tilde x$.} Meanwhile, our exercise directly informs whether the implicit design is the true design. If it were, then regression is even an AIPW estimator using the true design ((ref))---in many ways a natural estimator in causal inference settings. Notably, this AIPW estimator is for an estimand that weights treatment effects by $\pi_i^* (1-\pi_i^*)$. A party in a 5-party coalition receives 72% of the weight that a party in a 3-party coalition receives. Practitioners can opt to reweight such estimands, which we illustrate with an application to blakeslee2020way.

Unknown design

We return to blakeslee2020way who use borewell failure ($W_i$) as a quasi-experimentally assigned treatment. They consider a range of income and employment outcomes and conclude that(i) well failure causes a decline in agricultural income and employment, but reallocation to off-farm offsets the lost income, and (ii) those living in high economic development areas adapt more easily. For evaluating (ii), blakeslee2020way consider the regression that interacts treatment with an indicator $h_i$ of whether the village $v(i)$ of household $i$ has high economic development \[ Y_i = \tau_0 W_i + \tau_1 W_i h_i + \tilde x_i'\mu + \epsilon_{i}, \numberthis \label{eq:blakeslee_interact} \] where $\tau_0$ is interpreted as a treatment effect for those with $h_i = 0$ and $\tau_1$ is interpreted as a difference of treatment effects among $h_i = 1$ versus $h_i = 0$. See their Table 9 for the choice of covariates $\tilde x_i$.

(ref) shows that $\tau_0$ and $\tau_1$ in (ref) are both minimally quasi-experimental only if $\pi_i^*$ is linear in $\tilde x_i$ and $\pi_i^* h_i$ is also linear in $\tilde x_i$. Because $\tilde x_i$ includes village fixed effects and $h_i$ is their span, it is easy to check that $\pi_i^* h_i$ is linear if $\pi_i^*$ is linear---and thus the implicit design exists. We compute it in (ref).

Here the true design is unknown, so following (1) we first check whether the implicit design even looks like a coherent model of treatment. This basic check already raises concerns: The estimated implicit design places 55 out of 786 observations outside of $[0,1]$, which immediately raises concerns about (ref). Next, we consider some stress tests for (2). Calibration performance of the implicit design is reasonable ((ref)); however, Ramsey's RESET test blandhol2022tsls against the linearity in $\tilde x_i$ does decisively reject ($p$-value: 0.00). We also examine whether the implicit design is economically plausible and concurs with descriptions in the paper.\footnote{In terms of predictiveness, the implicit design accounts for about 25% of variation in $W_i$, indicating that observable characteristics of households do predict treatment. Consistent with blakeslee2020way's explanation, most of the predictive power comes from the village and drill-time fixed effects (the within-$R^2$ is only 0.8%). blakeslee2020way (p.220) worry about selection on unobserved confounders, most plausibly “wealthier and more skilled farmers being less likely to experience borewell failure.” The estimated implicit design from their specification does not appear to show this. We assess this by regressing $\hat \pi_i$ on indicators for whether a household owns a $\br{\text{tractor}, \text{seed drill}, \text{thresher}, \text{motorcycle}}$ before they drilled their first borewell. None of these covariates, jointly or separately, is statistically significant at the conventional level. The largest $|t|$-statistic among these is 1.16. These covariates are not included in the specification (ref).}

{

figure[figure omitted — 338 chars of source]

}

{

figure[figure omitted — 350 chars of source]

}

Taken together, these diagnostics undermine a quasi-experimental interpretation of the interacted specification, i.e., they push toward answering (ref) in the negative and treating the regression primarily as an outcome model. This latter interpretation is also not straightforward: The same regression specification is used across multiple outcomes (some positive, some bounded by $[0,1]$), and it is not obvious why they are reasonably modeled by the same specification.

Refining the implicit design and implicit estimand

A researcher is then left with a practical question: Even if the implicit design is rejected, is the misspecification consequential for the reported conclusions, and can we assess sensitivity? One simple patch is to treat the implicit design as an estimated propensity score and recalibrate it by binning predicted probabilities and replacing them with within-bin treated frequencies, as suggested by imbens2015causal,zhao2022regression,van2024stabilized,\footnote{This is known by subclassification in imbens2015causal and histogram binning in the calibration literature zadrozny2001obtaining.} which enforces basic calibration properties by construction. A complementary robustness check is trimming: Set the weights for observations with out-of-bounds propensities to zero and see whether they were materially driving the regression. A third option is to model the design directly and implement, say, an AIPW estimator.

Beyond concerns about how the regression models treatment, we may also be concerned with various choices in the implicit estimand. These concerns are economically relevant: In blakeslee2020way, the estimand $\tau_1$ in (ref) is a difference of two variance-weighted estimands. A priori, we cannot rule out that the difference is driven by the weighting scheme compared to the difference in conditional average treatment effects. To assess sensitivity, we can check whether outcomes correlates with implicit designs. If not, different weighting schemes are unlikely to make a difference. If they do, we can retarget alternative estimands.

{

figure[figure omitted — 1,570 chars of source]

}

(ref) includes a battery of alternative estimates that address these strands of concerns. First, we make the implicit design less obviously misspecified. The simplest assessment is whether the units with out-of-bounds implicit design contribute substantially to the regression estimate. To that end, the variance-weighted estimates {\color{ALICE}$\bm{\times}$} uses the same implicit design, targets the same estimand, but removes the out-of-bounds units. These {\color{ALICE}$\bm{\times}$} estimates are almost identical to the regression estimates, indicating that the regression estimates put little weight on out-of-bounds units. The estimates {\color{ALICE}$\star$} patch the implicit design by recalibrating it.\footnote{We subclassify on the propensity scores following Chapter 17 in imbens2015causal. The binning in the subclassification uses the data-driven procedure in imbens2015causal, which recursively partitions the estimated propensity scores until either bins are too small or the mean propensity score is similar among treated and untreated units within a bin.} This also does not meaningfully alter the estimate.

Second, we may assess whether the variance-weighting in the implicit estimand matters by considering weighting schemes that treat units more equally.\footnote{Since the propensity score estimates are often close to or equal to zero and one, overlap violations make estimating the average treatment effect infeasible. Thus, we trim the propensity scores to $[0.02, 0.98]$ and construct corresponding estimators for the trimmed average treatment effect crump2009dealing.} These estimates---especially the patched estimates {\color{CORAL}$\star$}---are more different from the regression estimates, though not substantially so compared to sampling noise.

Finally, moving entirely away from the implicit designs in the regression, we also compute estimates by augmented inverse propensity weighting ({\color{RUBY}$\bm{A}$}IPW) by using a simple logit model for the propensity scores and a linear model for the outcome means. These alternative estimates are again similar to the regression estimates, indicating that the outcomes {in this application} are not so adversarially configured: The implicit design, while clearly rejected, nevertheless produces estimates that are similar to alternative estimates.

figure[figure omitted — 197 chars of source]

Why do alternative weighting schemes not make a difference? (ref) partitions the implicit designs into 7 bins by quantile, so that each bin within $[0,1]$ contains the same number of units.\footnote{In (ref), if everyone is treated in a bin, we treat the mean control outcome as zero, and vice versa. Thus the “treatment effects” for $\hat\pi_i \le 0$ represent negative mean control outcomes, and the “treatment effects” for $\hat\pi_i > 1$ represent mean treated outcomes.} On each bin, it displays the treatment effect difference as well as the weight placed on each bin by the implicit estimand. As the bin size becomes small, computing the difference between the {\color{ALICE} teal} curve---weighted by the {\color{ALICE} teal} weights---and the{\color{RUBY} crimson} curve---weighted by the {\color{RUBY} weights}---approximates the regression estimate. The weighting schemes for $h_i = 1$ versus $h_i=0$ are indeed different in $\tau_1$: High employment area ($h_i=1$) puts larger {\color{ALICE} weight} for households more likely to lose water access---peaking at the bin {\color{ALICE} $(0.6, 0.72]$} as opposed to at {\color{RUBY} $(0.44, 0.6]$}. Thus the apparent treatment effect difference reflects in part the difference in weighting. But since the differences in conditional average treatment effects are effectively constant and zero, the weighting again makes little difference to the bottom-line estimate.

Conclusion

Linear regressions are ubiquitous. Interpreting their results as causal, thanks to quasi-random assignment, is similarly commonplace. This paper studies the necessary conditions that this interpretation imposes on treatment assignment. We do so by studying the comparisons that regression estimands make under random assignment. Requiring that a regression be minimally quasi-experimental imposes linear restrictions in the design. The set of designs that satisfy these restrictions can be thought of as models of treatment assignment that the regression implicitly specifies. Each design also pinpoints a corresponding estimand that the regression implicitly chooses. Indeed, the regression is numerically equivalent to an AIPW estimator with such a treatment model for such an estimand.

Understanding quasi-experimental interpretation of regressions in this way essentially reduces to mechanical computations that can be scaled and automated. These computations can aid in examining new theoretical properties of particular specifications, itself the subject of a highly influential recent literature. In several theoretical vignettes, these computations unify and strengthen disparate strands of the literature. Additionally, we find that regressions with interactions and with two-way fixed effects have fragile design-based interpretations. This calls for caution and nuance when using them and presenting their results.

Directly computing implicit designs and estimands in practice provides a set of simple diagnostics for practitioners who wish to understand the quasi-experimental properties of a given regression. Doing so makes transparent the statistical and economic choices masked by a regression specification. Having opened up the black box, we can examine each of its components: e.g., evaluating whether the implicit design is plausible, assessing whether the regression targets an economically interesting estimand, and constructing estimates for alternative estimands. Additionally, making these implicit choices transparent may nudge practitioners to choose methods that model treatment assignment more directly.

\FloatBarrier