EconBase
← Back to paper

Misspecified regressions with mixed regressors: robust inference and causal interpretation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,564 characters · 28 sections · 65 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Misspecified regressions with mixed regressors: robust inference and causal interpretation

\onehalfspacing

abstractFor analytic convenience, existing statistical frameworks either assume random or fixed regressors. However, it is a little awkward that they do not cover the practical case of estimating the average treatment effect in experiments with randomized treatments and non-randomized, fixed pretreatment covariates. We unify the literature by providing the theory for regressions with mixed regressors that contain both random and fixed components. Importantly, our theory allows for misspecification of the regression functions. We first establish general results for estimating equations with both random and fixed components and then use it to analyze misspecified linear regression, with applications to completely randomized experiments. We focus on the causal interpretation of the regression coefficients and standard errors even when the models are wrong. We start with the theory for independent data and then extend the discussion to clustered data.

KEYWORDS: Causal inference; instrumental variable; misspecification; regression adjustment; robust standard error

Introduction to statistical inference with misspecified models

Empirical researchers use models to extract information from data. However, models are only approximations to real world phenomena and are often misspeficied. The theory of misspecified models has been of continuing interest in statistics and econometrics Cox1961, Huber1967, White1980a, White1982, Imbens1997, AbadieImbensZheng2014, BujaBrownBerk2019, BujaBrownKuchibhotla2019. A well-known result from Huber1967 and White1982 is that we must use the Huber--White (HW) (also known as robust or sandwich) standard errors when the models are misspecified.

AbadieImbensZheng2014 pointed out an interesting distinction between random regressors and fixed regressors in misspecified regressions: while the robust standard errors are consistent with random regressors, they are only conservative with fixed regressors (see also White1983 and Chow1984). In the theoretical literature, the choice between random regressors and fixed regressors is often driven by analytic convenience. They do not cover the practical setting of randomized experiments with randomized treatments and fixed pretreatment covariates. Motivated by this setting due to its relevance for regression-based analysis for causal inference, we unify the literature by developing the theory for misspecified regressions with mixed regressors. We show that similar to the setting of fixed regressors, the robust standard errors are conservative with mixed regressors. This constitutes our first contribution.

As Freedman2006 critically pointed out, although robust standard errors can be useful for approximating the large-sample uncertainty of the estimators based on misspecified models, they do not solve the first-order problem that the targeted parameters themselves may not be meaningful in general. Geer2019 made a similar point. A more optimistic quote from BoxDraper1987 is that “Essentially, all models are wrong, but some are useful.” To complement BoxDraper1987, we argue that when we use models for inference, we must first verify that the parameters from misspecified models have meaningful interpretations. In linear regression, a celebrated result is that least squares gives the best linear approximation to the conditional mean function of the outcome given the regressors White1980a, AngristPischke2009, BujaBrownBerk2019. While this is a correct mathematical statement, it does answer the question whether linear approximation is a good approximation in the first place (Sims2010; Ding2023). We report some positive results about linear regression for causal inference in randomized experiments. In particular, we analyze various regressions for causal inference in randomized experiments with covariate adjustment, and extend them to the local average treatment effect framework ImbensAngrist1994, AngristImbensRubin1996. This constitutes our second contribution.

We further extend the theory to deal with clustered data. Moreover, we detail the theory for regression analysis of cluster randomized experiments BarriosDiamondImbens2012, SuDing2021, AbadieAtheyImbens2023, BugniCanayShaikh2025. In particular, we propose a novel correction to the usual Liang--Zeger (LZ) cluster robust standard error LiangZeger1986, Arellano1987, AbadieAtheyImbens2023 when the covariates are treated as random in the regression with treatment-covariates interaction. This constitutes our third contribution.

Table (ref) summarizes the main results of our paper. The remainder of the paper is organized as follows. Section (ref) displays the general results of $Z$-estimation with independent and identically distributed (i.i.d.) data. Section (ref) displays the results from linear regressions. Section (ref) displays the results under complete randomization. Section (ref) displays the general results of $Z$-estimation under clustered data, and apply it to cluster randomization. Section (ref) studies the finite-sample performance of our point and variance estimators based on simulation. Section (ref) provides the concluding remarks. The Supplementary Material contains the results for the local average treatment effect framework, all proofs, and several intermediate results.

table[table omitted — 3,161 chars of source]

\paragraph{Notation} We use $\mathbb{N}$ to denote the set of all non-negative integers. Let $1(\cdot)$ denote the indicator function. Let ${I}_{m}$ be a $m \times m$ identity matrix. We suppress the dimension $m$ when it is clear from the context. Let $\|\cdot\|$ denote the Euclidean norm, i.e., $\|w\| = \sqrt{w^{\mathrm{T}} w}$ for $w\in\mathbb{R}^v$. Unless stated otherwise, all vectors are assumed to be column vectors. Let $X$, $Y$, and $\varepsilon$ be the $n \times K$ matrix with $i$th row equal to $x_i^{\mathrm{T}}$, the $n$-vector with $i$th element equal to $y_i$, and the $n$-vector with $i$th element equal to $\varepsilon_i$, respectively. Define $\dot{x}_i = x_i - \mathbb{E}(x_i)$ and $\ddot{x}_i = x_i - \bar{x}$ where $\bar{x} = n^{-1}\sum_{i=1}^n x_i$, following the notation in NegiWooldridge2021. Let $[\cdot]_{(a,a)}$ denote the $(a,a)$th element of the matrix inside $[\cdot]$. We use $\text{lm}(y_i \sim x_i)$ to denote the least-squares regression of $y_i$ on $x_i$ and focus on the associated Eicker--Huber--White (EHW) variance estimator and $\text{lm}(y_{ij} \sim x_{ij})$ to denote the least-squares regression of $y_{ij}$ on $x_{ij}$ and focus on the associated LZ variance estimator clustered at the level of cluster $i$. The terms “regression” and “EHW variance” refer to the numerical outputs of the least-squares fit without any modeling assumptions. Let \(\mathbb{E}_\circ[\cdot]\) denote the relevant expectation operator (unconditional or conditional depending on the design): \(\mathbb{E}_\circ = \mathbb{E}_{x}\) refers to the expectation under the joint distribution of \(x\), while \(\mathbb{E}_\circ = \mathbb{E}_{x_1|x_2}\) refers to the expectation under the conditional distribution of \(x_1\) given \(x_2\). The same convention applies to the probability measure $\mathbb{P}_\circ$. We use $o(1;\mathbb{P}_{\circ})$ as a less cluttered notation for $o_{\mathbb{P}_{\circ}}(1)$, denoting a sequence of random variables that converges to zero in $\mathbb{P}_{\circ}$-probability. We use \(\mathbb{V}\) to denote variance. To present our asymptotic results, we introduce the notion of conditional convergence in distributions. Let $\mathcal{L}(t;W \mid X)=\mathbb{P}(W \le t \mid X)$ denote the cumulative distribution function of the random variable $W$. We say that $W_n \mid X_n \overset{\textup{d}}{\to} W $ a.s. if $\limsup_{n\to\infty}\;\sup_{t\in\mathbb{R}} \bigl|\mathcal{L}\bigl(t;W_n \mid X_n\bigr) -\mathcal{L}\bigl(t;W\bigr)\bigr| =0$, $\mathbb{P}$-a.s. We omit the measure-theoretic language “a.s." below. Let $v^{\otimes 2} = vv^{\mathrm{T}}$ denote the outer product of a vector $v$.

Robust inference based on $Z$-estimation

We begin with the familiar random-design setting with i.i.d. observations and establishes a benchmark in Section (ref). Because this is the standard framework for \(Z\)-estimation, we use it to introduce the basic notation and variance estimator before turning to fixed and mixed designs. Sections (ref) and (ref) then develop the corresponding fixed- and mixed-design results. The key distinction across these designs is the source of randomness: random-design inference is unconditional, fixed-design inference conditions on all covariates, and mixed-design inference conditions on only a subset of the covariates.

Random design

Under random design, we observe independent observations \( w_i \), and the parameter of interest \( \beta^\textup{r} \) is identified through a population moment condition \[ \mathbb{E}[\psi(w_i; \beta^\textup{r})] = 0, \] where we use $\beta^\textup{r}$ to denote the estimand under random design. We focus on the class of estimators \( \hat{\beta} \) that can be written as the unique solution to the corresponding sample estimating equations:

equation[equation omitted — 88 chars of source]

where $W = (w_1,\cdots, w_n)$ collects all observations. A first-order Taylor expansion of $\bar{\psi}(W;\hat{\beta})$ around $\beta^\textup{r}$ gives \[ 0=\bar{\psi}(W;\hat{\beta}) \approx \bar{\psi}(W;{\beta}^\textup{r}) +\Gamma(\beta^\textup{r})\,(\hat\beta-\beta^\textup{r}), \] where ${\Gamma}(\beta^\textup{r}) = \left.\frac{\partial}{\partial b^{\mathrm{T}}} \bar{\psi}(W;{b}) \right|_{b=\beta^\textup{r}}$. This yields the linearization of $Z$-estimator in (ref) as $\hat{\beta} \approx \beta^\textup{r} - \Gamma(\beta^\textup{r})^{-1}\bar{\psi}(W;{\beta}^\textup{r})$, and the HW variance estimator for the asymptotic variance of $\hat{\beta}$:

equation[equation omitted — 124 chars of source]

where

align*[align* omitted — 237 chars of source]

We use the label HW for the general $Z$-estimation setting. Maximum likelihood estimation with possibly misspecified models is a special case of this framework with $\psi$ as the score function Cox1961,Huber1967,White1980a, White1982; we therefore focus on the general $Z$-estimation.

We apply the general $Z$-estimation results to linear regression in Section (ref) and to randomized experiments in Section (ref). The framework accommodates not only Lin-type fully interacted regression adjustment, but also the IV regression, with the local average treatment effect as a leading application of the latter ImbensAngrist1994. Throughout this section, we write $w_i=(y_i,x_i)$, where $y_i$ is the response and $x_i$ denotes the covariates. Specifically, we define the estimand $\beta^\textup{r}$ as the unique solution to the estimating equation: \[ \mathbb{E}\left[\psi(y_i, x_i;\beta^\textup{r})\right] = 0. \] Define the asymptotic variance of $\hat{\beta}$ under random design as $V^\textup{r} = (\Gamma^\textup{r})^{-1}\Delta^\textup{r}(\Gamma^{\textup{r} })^{-\mathrm{T}}$, with

align*[align* omitted — 266 chars of source]

More generally, to present the three cases in a unified notation, we use \(\beta^{\diamond}\) to denote the estimand under different designs, and \(\mathbb{E}_{\circ}\) and \(\mathbb{P}_{\circ}\) to denote the corresponding expectation operator and probability measure, respectively. Specifically, let \(\diamond = \textup{r}\) under a random design with \(\circ = (y, x)\); \(\diamond = \textup{f}\) under a fixed design with \(\circ = y| x \); and \(\diamond = \textup{m}\) under a mixed design with \(\circ = (y, x_1) | x_2\). This notation allows us to present the following assumption uniformly across the three design regimes.

assumptionLet \(\beta^{\diamond} \in \Theta \subset \mathbb{R}^p\) be the target parameter. Suppose that $w_i$, $i=1,\ldots,n$, are i.i.d.. \begin{enumerate}[(a)] • The parameter space $\Theta$ is compact, \(\beta^{\diamond}\) lies in the interior of the \(\Theta\), and $\beta^{\diamond}$ is the unique solution to \(\mathbb{E}_\circ[\psi(w, b)] = 0\). • The function $\psi(w,b)$ is continuous in $b\in\Theta$ almost surely, and $\mathbb{E}\left[\sup_{b\in\Theta}\|\psi(w,b)\|\right]<\infty$. • There exists a neighborhood $\mathscr{N}$ of $\beta^\diamond$ such that $\psi(w,b)$ is continuously differentiable in $b\in \mathscr{N}$ almost surely. • $\mathbb{E}_\circ\left[ \sup_{b\in\mathscr{N}}\|\psi(w,b)\|^{2+\delta} \right]<\infty$ for some $\delta>0$ and $\mathbb{E}_\circ\left[ \sup_{b\in\mathscr{N}}\|\nabla_b\psi(w,b)\| \right]<\infty$. • \(\Gamma^\diamond := \mathbb{E}_\circ[\nabla_b \psi(w, \beta^{\diamond})]\) is nonsingular. \end{enumerate}

Theorem (ref) below reviews the asymptotic normality of the $Z$-estimator $\hat{\beta}$ and the consistency of the HW variance estimator $\hat{V}_{\textsc{hw}}$ in (ref) under random design.

theoremAssume that $W = \{y_i, x_i\}_{i=1}^n$ are i.i.d.. Under random design and Assumption (ref) with \(\beta^{\diamond} = \beta^{\textup{r}}\), and \(\mathbb{E}_\circ = \mathbb{E}_{(y,x)}\) and \(\mathbb{P}_\circ = \mathbb{P}_{(y,x)}\), we have \[ (V^\textup{r})^{-1/2} \sqrt{n}(\hat{\beta} - \beta^\textup{r}) \overset{\textup{d}}{\to} \mathcal{N}\left(0, I \right) \text{ and } n \hat{V}_{\textsc{hw}} = V^\textup{r} + o(1;\mathbb{P}_{(y,x)}). \]

Theorem (ref) is the standard random-design result for $Z$-estimation equations; see, for example, NeweyMcFadden1994. We include it to establish notation and to serve as the benchmark for the fixed- and mixed-design results below.

Fixed design

Under fixed design, the $x_i$'s are fixed, or, equivalently, we condition on them. Let $X = (x_1, \ldots, x_n)^{\mathrm{T}}.$ Define the estimand $\beta^\textup{f}$ as the solution to the estimating equations: \[ \mathbb{E}\left[ \frac{1}{n} \sum_{i=1}^n \psi(y_i,x_i;\beta^\textup{f} ) \mid {X} \right] = 0. \] Define the conditional variance $n\cdot \mathbb{V}(\hat{\beta} \mid X)$ as $V^\textup{f} = (\Gamma^\textup{f})^{-1}\Delta^\textup{f}(\Gamma^{\textup{f} })^{-\mathrm{T}}$ with

align*[align* omitted — 292 chars of source]

and the asymptotic bias of $\hat{V}_{\textsc{ehw}}$ in (ref) for $V^\textup{f}$ as

equation[equation omitted — 253 chars of source]
theoremAssume that $\{y_i, x_i\}_{i=1}^n$ are i.i.d.. Under fixed design and Assumption (ref) with \(\beta^{\diamond} = \beta^{\textup{f}}\), \(\mathbb{E}_\circ = \mathbb{E}_{y|x}\) and \(\mathbb{P}_\circ = \mathbb{P}_{y|x}\), we have \[ (V^\textup{f})^{-1/2} \sqrt{n}(\hat{\beta} - \beta^\textup{f}) \mid X \overset{\textup{d}}{\to} \mathcal{N}\left(0, I \right) \text{ and } n \hat{V}_{\textsc{hw}} = V^\textup{f} + B^{\textup{f}} + o(1;\mathbb{P}_{y|x}). \]

Theorem (ref) is the fixed-design counterpart of the standard $Z$-estimation result in Theorem (ref). Theorem (ref) indicates that under fixed design, $\hat{V}_{\textsc{hw}}$ in (ref) is conservative with asymptotic bias $B^{\textup{f}}$ in (ref). AbadieImbensZheng2014 also study general $Z$-estimators under misspecification and distinguish between population and covariate-conditional estimands. Their asymptotic results, however, are formulated under i.i.d.\ sampling and unconditional inference, whereas our fixed-design analysis conditions on the realized regressors.

Mixed design

We partition all the regressors $X= (X_1, X_2)$. Under mixed design, our analysis conditions on part of the regressors, i.e., ${X}_2$, and take the remaining regressors ${X}_1$ as random. Define the estimand $\beta^\textup{m}$ as the solution to the estimating equations: \[ \mathbb{E}\left[\frac{1}{n} \sum_{i=1}^n \psi(y_i,x_i;\beta^\textup{m})\mid {X}_2 \right] = 0. \] Define the conditional variance $n\cdot \mathbb{V}(\hat{\beta} \mid X_2)$ as $V^\textup{m} = (\Gamma^\textup{m})^{-1}\Delta^\textup{m}(\Gamma^{\textup{m} })^{-\mathrm{T}}$ with

align*[align* omitted — 297 chars of source]

and the asymptotic bias of $\hat{V}_{\textsc{ehw}}$ in (ref) for $V^\textup{m}$ as

equation[equation omitted — 256 chars of source]
theoremAssume that $\{y_i, x_i\}_{i=1}^n$ are i.i.d.. Under mixed design and Assumption (ref) with \(\beta^{\diamond} = \beta^{\textup{m}}\), \(\mathbb{E}_\circ = \mathbb{E}_{(y,x_1)|x_2}\) and \(\mathbb{P}_\circ = \mathbb{P}_{(y,x_1)|x_2}\), we have \begin{equation*} (V^m)^{-1/2} \sqrt{n}(\hat{\beta} - \beta^m) \mid X_2 \overset{d}{\to} \mathcal{N}\left(0, I \right) and n \hat{V}_{hw} = V^m + B^\textup{m} + o(1;\mathbb{P}_{(y,x_1)\mid x_2}). \end{equation*}

Theorem (ref) generalizes the random- and fixed-design benchmarks to mixed designs, where only part of the regressors are conditioned on. This formulation is useful for settings such as randomized experiments with fixed or pre-determined covariates. Under mixed design, $\hat{V}_{\textsc{hw}}$ remains conservative for estimating $V^\textup{m}$ under misspecification, with the additional term $B^{\textup{m}}$ in (ref) capturing the approximation-error component induced by conditioning on \(X_2\). In the special case in which either $X_2$ is empty or $\mathbb{E}\!\left[ \psi(y_i,x_i;\beta^{\textup{m}}) \mid x_{i2} \right] =0$, the HW variance estimator $\hat{V}_{\textsc{hw}}$ is consistent rather than strictly conservative.

OLS estimation as a special case

We now specialize the framework in Section (ref) to OLS. The structure of this section parallels that of Section (ref). Assume:

equation[equation omitted — 74 chars of source]

with $y_i$ being the outcome of interest, $x_i$ a $K$-vector of observed covariates, possibly including an intercept, and $\varepsilon_i$ an unobserved error. If $\sum_{i=1}^n x_ix_i^{\mathrm{T}}>0$, the OLS estimator $\hat{\beta}$ solves

equation[equation omitted — 137 chars of source]

Define the residual from the OLS fit as $\hat{\varepsilon}_i = y_i-x_i^{\mathrm{T}}\hat{\beta}$. Specializing the general sandwich estimator $\hat{V}_{\textsc{hw}}$ in (ref) to the OLS estimating equations in (ref) gives the Eicker--Huber--White (EHW) variance estimator Eicker1967,Huber1967,White1980a:

equation[equation omitted — 275 chars of source]

We use the label EHW to emphasize that, for OLS, the general HW sandwich estimator coincides with Eicker1967's heteroskedasticity-robust covariance estimator.

In this section, we impose the following assumption.

assumption\begin{enumerate}[{(a)}] • The variables $(y_i, x_i)$, $i=1, \ldots, n$, are i.i.d.; • $\mathbb{E}_\circ(y_i^4)<\infty$; • $\mathbb{E}_\circ\|x_i\|^4<\infty$; • $\mathbb{E}_\circ(x_ix_i^{\mathrm{T}})$ is positive definite. \end{enumerate}

Below, we study the performance of the OLS estimator $\hat{\beta}$ in (ref) and $\hat{V}_{\textsc{ehw}}$ in (ref) under different sources of randomness.

Random Design

Under random design, the estimand of interest $\beta^{\textup{r}}$ is the solution to the estimating equations:

align[align omitted — 98 chars of source]

Equivalently, $\beta^{\textup{r}}$ is the population OLS projection coefficient because our theory does not assume the linear model (ref) to be correctly specified. Define the projection error $\varepsilon_i^{\textup{r}} = y_i-x_i^{\mathrm{T}}\beta^{\textup{r}}$. The OLS estimator $\hat{\beta}$ in (ref) is biased for $\beta^{\textup{r}}$ since $\mathbb{E}(\hat{\beta}-\beta^{\textup{r}})\ne 0$, but is consistent for $\beta^{\textup{r}}$ by the law of large numbers under Assumption (ref). Define

align[align omitted — 212 chars of source]
theoremUnder random design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,x)}\), we have \begin{equation} (V^{r})^{-1/2} \sqrt{n}(\hat{\beta} - \beta^{r}) \overset{d}{\to} \mathcal{N}(0, I) and n \hat{V}_{ehw} = V^{r} + o(1;\mathbb{P}_{(y,x)}). \end{equation}

Theorem (ref) states that the EHW variance estimator $\hat{V}_{\textsc{ehw}}$ in (ref) is consistent for estimating $V^{\textup{r}}$ in (ref). Theorem (ref) is the standard random-design OLS result under possible misspecification; see White1980, White1982 and NeweyMcFadden1994.

Fixed Design

Under fixed design, we condition on all regressors $X$. Define the estimand $\beta^{\textup{f}}$ as the solution to the estimating equations:

align[align omitted — 125 chars of source]

When conditional on ${X}$, $\hat{\beta}$ is an unbiased estimator for $\beta^{\textup{f}}$ since $\mathbb{E}(\hat{\beta}-\beta^{\textup{f}} \mid {X})=0$, and also consistent for $\beta^{\textup{f}}$ by the law of large numbers under Assumption (ref). Define the conditional variance $n\cdot \mathbb{V}(\hat{\beta} \mid X)$ as

equation[equation omitted — 247 chars of source]

and the asymptotic bias of $\hat{V}_{\textsc{ehw}}$ in (ref) for $V^{\textup{f}}$ as

equation[equation omitted — 290 chars of source]

where $\varepsilon_i^{\textup{f}} = y_i-x_i^{\mathrm{T}}\beta^{\textup{f}}$.

theoremUnder fixed design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{y|x}\), we have \begin{align*} (V^{f})^{-1/2} \sqrt{n}(\hat{\beta} - \beta^{f}) \mid {X} \overset{d}{\to} \mathcal{N} (0, I) and n \hat{V}_{ehw} = V^{f} + B^{\mathrm{f}} + o(1;\mathbb{P}_{y|x}). \end{align*}

Theorem (ref) states that under fixed design, the EHW variance estimator $\hat{V}_{\textsc{ehw}}$ in (ref) is a conservative estimator for $V^{\textup{f}}$ in (ref) with asymptotic bias $B^{\mathrm{f}}$ in (ref). Theorem (ref) is the OLS specialization of the mixed-design $Z$-estimation result in Theorem (ref).

AbadieImbensZheng2014 also study the covariate-conditional OLS estimand and derive the corresponding asymptotic distribution. Their analysis is unconditional, so the realized design $X$, and hence the conditional estimand $\beta^{\mathrm f}(X)$, varies across repeated samples. By contrast, Theorem (ref) establishes asymptotic normality conditional on the realized design $X$. AbadieImbensZheng2014 show that the conditional variance of the OLS estimator $\hat{\beta}$ converges in probability to \[ V_{\text {cond }} = \operatorname{plim}(n \cdot \mathbb{V}(\hat{\beta} \mid {X})) = \left(\mathbb{E}\left[x_i x_i^{\mathrm{T}}\right]\right)^{-1} \left(\mathbb{E}\left[\sigma^2\left(x_i\right) x_i x_i^{\mathrm{T}} \right]\right)\left(\mathbb{E}\left[x_i x_i^{\mathrm{T}}\right]\right)^{-1}. \] The middle part of $V^{\textup{r}}$ can be decomposed into conditional variance and approximation error: \[ \mathbb{E}\left[(y_i-x_i^{\mathrm{T}}\beta^{\textup{r}})^2 \mid x_i \right] = \mathbb{E}\left[(y_i-\mu(x_i)+\mu(x_i)-x_i^{\mathrm{T}}\beta^{\textup{r}})^2 \mid x_i \right] = \sigma^2(x_i) + (\mu(x_i)-x_i^{\mathrm{T}}\beta^{\textup{r}})^2, \] where $\sigma^2(x_i)= \mathbb{V}(y_i \mid x_i)$ and $\mu(x_i)=\mathbb{E}(y_i \mid x_i)$. Comparing $V_{\text {cond }}$ with $V^{\textup{r}}$ shows that the latter is generally larger: $V^{\textup{r}} = V_{\text {cond }} + \mathbb{V}(\beta^{\textup{f}})$ where $\mathbb{V}(\beta^{\textup{f}}) = \operatorname{plim}n \cdot \mathbb{E}[(\beta^{\textup{f}} - \beta^{\textup{r}})(\beta^{\textup{f}} - \beta^{\textup{r}})^{\mathrm{T}}]$. The variance estimator $\hat{V}_{\textsc{ehw}}$ is conservative for $V^{\textup{f}}$ since $\beta^{\textup{f}}$ cannot be consistently estimated.

Mixed Design

We partition the regressors as \(X = (X_1, X_2)\). Under the mixed-design setting, our analysis conditions on part of the regressors, namely \(X_2\), while treating the remaining regressors \(X_1\) as random. A leading application of this framework arises in randomized controlled trials, where \(X_1\) represents the randomized treatment and \(X_2\) consists of fixed or pretreatment covariates. The estimand $\beta^{\textup{m}}$ is therefore defined as the solution to the estimating equations:

align[align omitted — 128 chars of source]

With $\varepsilon_i^{\textup{m}} = y_i-x_i^{\mathrm{T}}\beta^{\textup{m}}$, define the conditional variance $n\cdot \mathbb{V}(\hat{\beta} \mid X_2)$ as

equation[equation omitted — 305 chars of source]

and the asymptotic bias of $\hat{V}_{\textsc{ehw}}$ in (ref) for $V^{\mathrm{m}} $ as

equation[equation omitted — 346 chars of source]
theoremUnder mixed design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,x_1)|x_2}\), we have \[ (V^{\textup{m}})^{-1/2} \sqrt{n}(\hat{\beta} - \beta^{\textup{m}}) \mid {X}_2 \overset{\textup{d}}{\to} \mathcal{N}\left(0, I \right) \text{ and } n \hat{V}_{\textsc{ehw}} = V^{\mathrm{m}} + B^{\mathrm{m}} + o(1;\mathbb{P}_{(y,x_1)\mid x_2}). \]

Theorem (ref) indicates that under mixed design, the EHW variance estimator $\hat{V}_{\textsc{ehw}}$ in (ref) is conservative for $V^{\mathrm{m}}$ in (ref) with asymptotic bias $B^{\mathrm{m}}$ in (ref). Theorem (ref) is the OLS specialization of the mixed-design $Z$-estimation result in Theorem (ref), and the result is new in the literature. Related work by ChetverikovHahnLiao2023 studies OLS when a regressor of interest is randomly assigned, but does not consider our conditional mixed-design framework or the resulting conservativeness of the EHW variance estimator under misspecification.

Interpretation of misspecified regressions for causal inference

Freedman2008 and Geer2019 question the relevance of inference based on $\hat{V}_{\textsc{hw}}$ when the target parameters lack a meaningful interpretation under model misspecification. To address this concern, we provide a detailed discussion of the causal interpretations of coefficients from regressions commonly used in causal inference when the working linear models may be misspecified. Our main objective is not merely to derive asymptotic variance formulas, which follow from the general results in Sections (ref) and (ref), but to clarify which causal parameters the regression coefficients identify. We study completely randomized experiments and focus on the comparison of random and mixed designs with covariates in Section (ref), Section (ref) because treatment $Z$ is randomized. We do not consider inference conditional on $Z$ because the coefficient does not have causal interpretation if the linear model is incorrect. For this reason, we omit the fixed-design analysis, since it does not quantify the advantages of randomization with a misspecified linear model. We focus here on OLS and relegate the IV results for the local average treatment effect framework to Section (ref) in the Supplementary Material .

Consider an experiment with a binary treatment $Z_i\in\{0,1\}$. Let $y_i(z)$ denote the potential outcome of unit $i$ under treatment status $z\in\{0,1\}$. This notation allows us to characterize the causal interpretation of the OLS coefficient even when the working linear model is misspecified. The observed outcome $y_i$ is connected with potential outcomes in the usual manner: $y_i = Z_i Y_i(1) + (1-Z_i)Y_i(0)$. Researchers also observe the covariate vector $x_i = (x_{i1}, \ldots, x_{iJ})$ for $i = 1, \ldots, n$.

In this section, we impose the following assumption.

assumption\begin{enumerate}[(a)] • The variables $(Y_i(1), Y_i(0), Z_i, x_i)$, $i=1, \ldots, n$, are i.i.d.; • $(Y_i(0), Y_i(1), x_i) \perp \!\!\! \perp Z_i$; • $\mathbb{P}(Z_i=1)= e \in(0,1)$; • $\mathbb{E}_\circ[ |Y_{i}(z)|^{4} ] < \infty $ for all $z \in \{0,1\}$ and $\mathbb{E}_\circ[ \left\|x_{i}\right\|^{4} ] < \infty $ for some $\delta > 0$; • $\mathbb{E}_\circ(x_ix_i^{\mathrm{T}})$ is positive definite. \end{enumerate}

Simple difference in means and its regression implementation

To estimate the average treatment effect (ATE), we consider the linear regression:

equation[equation omitted — 60 chars of source]

Let $\hat{\beta}$ denote the coefficient of $Z_i$ from the above OLS fit and $\hat{\varepsilon}_{i}$ denote the residual from the same OLS fit. Define $\hat{V}_{\textsc{ehw}}$ as the $(2,2)$ entry of the EHW variance estimator in the form of (ref) with regressors $x_{i} = (1, Z_i)^{\mathrm{T}}$; this is the variance of the coefficient on $Z_i$.

The estimands are $(\alpha^{\textup{r}}, (\beta^{\textup{r}})^{\mathrm{T}} )^{\mathrm{T}} = \mathbb{E}(x_{i} x_{i}^{\mathrm{T}} )^{-1}\mathbb{E}(x_{i} y_i )$. By the general theory in Section (ref), we derive the asymptotic variance of $\hat{\beta}$ as

equation[equation omitted — 161 chars of source]

with $\varepsilon^{\textup{r}}_i(z) = y_i(z) - \mathbb{E}(y_i(z))$ for $z\in \{0,1\}$.

theoremUnder random design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z)}\), we have $\beta^{\textup{r}} = \mathbb{E}(Y_i(1) - Y_i(0) )$ and \begin{equation*} (V^{r})^{-1/2}\sqrt{n}(\hat{\beta} - \beta^{r} ) \overset{d}{\to} \mathcal{N}(0, 1) and n \hat{V}_{ehw} = V^{r} + o(1;\mathbb{P}_{(y,Z)}). \end{equation*}

The coefficient of $Z_i$ from the projection of $y_i$ on $x_{i}$ identifies the ATE. Theorem (ref) also indicates that under random design, $\hat{V}_{\textsc{ehw}}$ is consistent for estimating $V^{\textup{r}}$ in (ref). Theorem (ref) is the standard difference-in-means result for randomized experiments; see Wooldridge2020.

Fisher's analysis of covariance

We consider the additive regression:

equation[equation omitted — 67 chars of source]

Let $\hat{\beta}_{\textsc{f}}$ denote the coefficient of $Z_i$ from the OLS in (ref). We use the subscript “F” to signify Fisher1935. Denote $\hat{\varepsilon}_{i,\textsc{f}}$ as the residual from the same OLS fit. Define $\hat{V}_{\textsc{ehw},\textsc{f}}$ as the $(2,2)$ entry of the EHW variance estimator in the form of (ref) with regressors $x_{i,\textsc{f}} = (1, Z_i, x_i^{\mathrm{T}})^{\mathrm{T}}$.

\paragraph*{Random Design} By the general theory in Section (ref), we derive the asymptotic variance of $\hat{\beta}_{\textsc{f}}$ as

equation*[equation* omitted — 181 chars of source]

where $\varepsilon_{i,\textsc{f}}^{\textup{r}}(z) = y_i(z) - \mathbb{E}(y_i(z)) - \dot{x}_i^{\mathrm{T}}\gamma_\textsc{f}^{\textup{r}}$ for $z\in\{0, 1\}$ with $\gamma^{\textup{r}}_{\textsc{f}}$ being the probability limit of the coefficient of $x_i$ from the OLS fit in (ref).

theoremUnder random design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z,x)}\), we have $\beta^{\textup{r}} = \mathbb{E}(Y_i(1) - Y_i(0) )$ and \begin{equation*} (V_f^{r} )^{-1/2}\sqrt{n}(\hat{\beta}_{f} -\beta^{r} ) \overset{d}{\to} \mathcal{N}(0, 1) and n \hat{V}_{\textsc{ehw,f}} = V_\textsc{f}^{\textup{r}} + o(1;\mathbb{P}_{(y,Z,x)}). \end{equation*}

The coefficient of $Z_i$ from the projection of $y_i$ on $x_{i,\textsc{f}}$ identifies the ATE. Theorem (ref) indicates that under random design, $\hat{V}_{\textsc{ehw},\textsc{f}}$ is consistent. This specification corresponds to Fisher's additive regression adjustment, whose design-based properties are studied by Freedman2008 and Lin2013. In their finite-population framework, the usual EHW variance estimator is generally conservative, whereas our random-design formulation, closer to the perspective of NegiWooldridge2021, yields consistency. We include the result to connect the random-design OLS theory in Section (ref) to Fisher's analysis of covariance adjustment and to provide a benchmark for the mixed-design result below.

\paragraph*{Mixed Design} Define $\varepsilon_{i,\textsc{f}}^{\textup{m}}(z) = (y_i(z) - n^{-1} \sum_{i=1}^n \mathbb{E}(y_i(z) \mid x_i)) - \ddot{x}_i^{\mathrm{T}}\gamma_{\textsc{f}}^{\textup{m}}$ for $z=0, 1$ with $\gamma^{\textup{m}}_{\textsc{f}}$ being the probability limit of the coefficient of $x_i$ from the OLS fit in (ref) under mixed design. By the general theory in Section (ref), we derive the conditional variance of $n\cdot \mathbb{V}(\hat{\beta}_{\textsc{f}} \mid X)$ as

equation*[equation* omitted — 389 chars of source]

and the asymptotic bias of $\hat{V}_{\textsc{ehw},\textsc{f}}$ for $V_{\textsc{f}}^{\textup{m}}$ as

equation[equation omitted — 214 chars of source]
theoremUnder mixed design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z)|x}\), we have $\beta^{\textup{m}} = n^{-1} \sum_{i=1}^n \mathbb{E}(Y_i(1) - Y_i(0) \mid x_i )$ and \begin{equation*} (V_{f}^{m})^{-1/2}\sqrt{n}(\hat{\beta}_{f}-\beta^{m} ) \mid X \overset{d}{\to} \mathcal{N}(0, 1) and n \hat{V}_{\textsc{ehw},{\textsc{f}}} = V_{\textsc{f}}^{\textup{m}} + B_{\textsc{f}}^{\textup{m}} + o(1;\mathbb{P}_{(y,Z) \mid x}). \end{equation*}

The coefficient of $Z_i$ from the projection of $y_i$ on $x_{i,\textsc{f}}$ identifies the average of the empirical conditonal average treatment effect, which is also known as the mixed average treatment effect by LiDingMealli2023. Theorem (ref) also indicates that under mixed design, $\hat{V}_{\textsc{ehw},\textsc{f}}$ is conservative in general with asymptotic bias $B^{\textup{m}}_{\textsc{f}}$ in (ref).

Lin's fully interacted adjustment

Now consider the regression with fully-interacted covariates, i.e.,

equation[equation omitted — 95 chars of source]

where $\ddot{x}_i = x_i - \bar{x}$. Let \(\hat{\beta}_{\textsc{l}}\) denote the coefficient on \(Z_i\) from the OLS regression in (ref). We use the subscript “L” to reference Lin2013. The coefficient on \(Z_i\) from the OLS fit in (ref) provides an estimate of the ATE when covariates are centered around their sample mean. Without centering, we need to estimate the ATE using \(\hat{\beta}_{\textsc{l}} + \hat{\xi}_{\textsc{l}} \bar{x}\) where $\hat{\xi}_{\textsc{l}}$ denotes the coefficient on the interaction term \(Z_i x_i\). Denote $\hat{\varepsilon}_{i,\textsc{l}}$ as the residual from the OLS fit in (ref). Define $\hat{V}_{\textsc{ehw},\textsc{l}}$ as the $(2,2)$ entry of the EHW variance estimator in the form of (ref) with regressors $x_{i,\textsc{l}} = (1, Z_i,\ddot{x}_i, Z_i\ddot{x}_i)^{\mathrm{T}}$.

\paragraph*{Random Design} We cannot naively apply the OLS theory without modifications because it ignores the uncertainty in $\bar x$. We can either modify the OLS theory or apply the general $Z$-estimation theory. We adopt the second strategy here. We apply the $Z$-estimation framework in Section (ref) to derive the asymptotic variance of the estimator $\hat{\beta}_{\textsc{l}}$. The variance estimator $\hat{V}_{\textsc{ehw},\textsc{l}}$, obtained by directly applying the OLS results in Section (ref), is incorrect because it ignores the additional uncertainty introduced by centering the covariates. In particular, obtaining a consistent EHW covariance estimator requires using augmented estimating equations Newey1984. Specifically, $\hat{\beta}_{\textsc{l}}$ can be expressed as $\hat{Y}_{\textsc{l}}(1) - \hat{Y}_{\textsc{l}}(0)$ where $\left(\mu_1, \gamma_1, \mu_0, \gamma_0, \mu_x\right)= (\hat{Y}_{\textsc{l}}(1), \hat{\gamma}_1, \hat{Y}_{\textsc{l}}(0), \hat{\gamma}_0, \bar{x})$ jointly solves the estimating equations:

equation[equation omitted — 422 chars of source]

where the last line accounts for the estimation of \(\mu_x\). With these equations, we can construct the variance estimator using (ref) and apply Theorem (ref) to demonstrate the consistency of the EHW variance estimator.

Let $\gamma_z^{\textup{r}}$ be the coefficient of $x_{i}$ in the OLS fit of $y_{i}$ on $1$ and $\ddot{x}_{i}$ over $\{i: Z_i = z\}$ under random design. Define $\varepsilon_{i,\textsc{l}}^{\textup{r}}(z) = y_i(z) - \mathbb{E}(y_i(z)) - \dot{x}_i^{\mathrm{T}}\gamma_z^{\textup{r}}$ for $z\in \{0, 1\}$. By the general theory in Section (ref), we derive the asymptotic variance of $\hat{\beta}_{\textsc{l}}$ under random design as

equation[equation omitted — 329 chars of source]

Let $\Sigma_x = n^{-1} \sum_{i=1}^n \ddot{x}_i\ddot{x}_i^{\mathrm{T}}$. We estimate $V^{\textup{r}}_{\textsc{l}}$ using the following estimator:

align[align omitted — 224 chars of source]

Theorem (ref) below establishes the asymptotic normality of $\hat{\beta}_{\textsc{l}}$ and the consistency of $\hat{V}_{\textsc{ehw},\textsc{l},\text{adj}}$ for estimating $V_{\textsc{l}}^{\textup{r}}$ in (ref) under random design.

theoremUnder random design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z,x)}\), we have $\beta^{\textup{r}} = \mathbb{E}(Y_i(1) - Y_i(0) )$ and \[ (V^{\textup{r}}_{\textsc{l}})^{-1/2}\sqrt{n}(\hat{\beta}_{\textsc{l}} - \beta^{\textup{r}}) \overset{\textup{d}}{\to} \mathcal{N}(0, 1) \text{ and } n \hat{V}_{\textsc{ehw},\textsc{l},\text{adj}} = V_{\textsc{l}}^{\textup{r}} + o(1;\mathbb{P}_{(y,Z,x)}). \] Moreover, \[ n \hat{V}_{\textsc{ehw},\textsc{l}} = V_{\textsc{l}}^{\textup{r}} - B_{\textsc{l}}^{\textup{r}} + o(1;\mathbb{P}_{(y,Z,x)}) \] with \begin{equation} B_{l}^{r} = (\gamma^{r}_1-\gamma^{r}_0)^{\mathrm{T}} \mathbb{V}(x_i) (\gamma^{r}_1-\gamma^{r}_0). \end{equation}

Under random design, the coefficient of $Z_i$ from the projection of $y_i$ on $x_{i,\textsc{l}}$ identifies the ATE. The result of asymptotic normality in Theorem (ref) is also proved in NegiWooldridge2021, whereas our proof relies on the general $Z$-estimation framework developed in Section (ref). NegiWooldridge2021 and ZhaoDing2021 similarly propose correcting the EHW variance estimator by adding the second term on the right-hand side of (ref); see also Ding2023.

As noted above, under a random design, centering around $\bar{x}$ introduces additional uncertainty, and the theory of OLS with i.i.d. data does not apply here. Consequently, using the OLS-based variance estimator $\hat{V}_{\textsc{ehw},\textsc{l}}$ from (ref) results in an anti-conservative variance estimator. To address this issue, a correction term in (ref) is necessary. This contrasts with the design-based framework of Lin2013, where the covariates are treated as fixed. In that setting, centering by the sample mean $\bar x$ introduces no additional uncertainty, and the usual EHW variance estimator for Lin's fully interacted regression is asymptotically conservative for the randomization variance.

remarkWe can construct the variance estimator in two ways: by applying the HW variance formula to the augmented estimating equations in (ref) and then using a plug-in approach as in (ref), or by adding the correction term to the EHW variance estimator $\hat{V}_{\textsc{ehw},\textsc{l}}$ as in (ref). Although these two methods are asymptotically equivalent, they are not numerically identical due to small finite-sample differences.

\paragraph*{Mixed Design} We can apply the results of OLS in Section (ref) because the covariates $\{x_i\}_{i=1}^n$ are fixed. Let $\gamma_z^{\textup{m}}$ be the coefficient of $x_{i}$ in the OLS fit of $y_{i}$ on $1$ and $\ddot{x}_{i}$ over $\{i: Z_i = z\}$ under mixed design. Define $\varepsilon_{i,\textsc{l}}^{\textup{m}}(z) = (y_i(z) - n^{-1}\sum_{i=1}^n \mathbb{E}(y_i(z) \mid x_i)) - \ddot{x}_i^{\mathrm{T}}\gamma_z^{\textup{m}}$ for $z\in\{0, 1\}$. By the general theory in Section (ref), we derive the asymptotic variance of $\hat{\beta}_{\textsc{l}}$ under mixed design as

equation*[equation* omitted — 357 chars of source]

Define the asymptotic bias of $\hat{V}_{\textsc{ehw},{\textsc{l}}}$ for $V^{\textup{m}}_{\textsc{l}}$ as

equation[equation omitted — 205 chars of source]
theoremUnder mixed design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z)|x}\), we have $\beta^{\textup{m}} = n^{-1} \sum_{i=1}^n \mathbb{E}(Y_i(1) - Y_i(0) \mid x_i )$ and \begin{equation*} (V^{m}_{l})^{-1/2} \sqrt{n}(\hat{\beta}_{l}-\beta^{m} ) \mid X \overset{d}{\to } \mathcal{N}(0, 1 ) and n \hat{V}_{\textsc{ehw},{\textsc{l}}} = V_{\textsc{l}}^{\textup{m}} + B_{\textsc{l}}^{\textup{m}} + o(1;\mathbb{P}_{(y, Z) \mid x}). \end{equation*}

The coefficient of $Z_i$ from the projection of $y_i$ on $x_{i,\textsc{l}}$ identifies empirical average of the conditional average treatment effects under mixed design. Theorem (ref) indicates that under mixed design, $\hat{V}_{\textsc{ehw},{\textsc{l}}}$ is conservative in general with asymptotic bias $B^{\textup{m}}_{\textsc{l}}$ in (ref). The mixed design formulation here has the advantage of simplifying the estimation of the asymptotic variance of $\hat{\beta}_{\textsc{l}}$.

Both the Fisher specification and the Lin specification target the same causal estimand under the random and mixed designs considered here. Their distinction therefore concerns efficiency rather than interpretation. Proposition (ref) shows that, under mixed design, the Lin specification is asymptotically no less efficient than the Fisher specification.

propositionUnder mixed design, $V_{\textsc{l}}^{\mathrm{m}} \leq V_{\textsc{f}}^{\mathrm{m}}$. If $\Sigma_x$ is positive definite, equality holds if and only if either $e=1/2$ or $\gamma_1^{\mathrm{m}}=\gamma_0^{\mathrm{m}}$.

This result complements existing efficiency comparisons under alternative sources of randomness. In the finite-population, design-based framework, Freedman2008 shows that additive covariate adjustment need not improve precision relative to the difference-in-means estimator, whereas Lin2013 shows that fully interacted adjustment is asymptotically no less efficient, even under misspecification. Under random sampling, NegiWooldridge2021 establishes an analogous ranking: full regression adjustment is asymptotically no less efficient than either the difference-in-means estimator or pooled regression adjustment, without requiring linear conditional mean functions. Proposition (ref) extends the comparison between additive and fully interacted adjustment to the mixed-design setting.

Clustered Data

Clustered data are common in empirical work, especially when observations are grouped by schools, classrooms, villages, firms, or geographic units. In such settings, researchers routinely use cluster-robust standard errors to account for within-cluster dependence. AbadieAtheyImbens2023 emphasize that the interpretation of LZ cluster robust standard error depends on the source of randomness, and SuDing2021 provide design-based theory for regression estimators in cluster-randomized experiments. We extend the $Z$-estimation framework in Section (ref) to clustered data and apply it to cluster-randomized experiments, focusing on parameter interpretation and robust inference under misspecified models and different sources of randomness.

M-estimation

Parallel to Section (ref), we develop a general \(Z\)-estimation framework for clustered data. We begin with the familiar random-design setting as a benchmark in Section (ref) and then extend the analysis to fixed- and mixed-design settings in Sections (ref) and (ref), respectively.

Random design

Under random design, we observe a random sample of clusters. Let $w_{ij}$ denote the observation for unit $j$ in cluster $i$, for $j=1,\ldots, n_i$, $i=1, \ldots M$, and let the total sample size be $N=\sum_{i=1}^M n_i$. Let $\sum_{i j}=\sum_{i=1}^M \sum_{j=1}^{n_i}$ denote the summation over all units. We assume that $n_i$ is fixed for each $i$. The parameter of interest \( \beta^\textup{r} \) is identified through a population moment condition \[ \mathbb{E}\left[ \sum_{ij} \psi(w_{ij};\beta^\textup{r}) \right] = 0. \] where we use $\beta^\textup{r}$ to denote the estimand under random design. Let $\hat{\beta}$ be the solution to the following equation:

equation[equation omitted — 137 chars of source]

where $W=\{w_{ij}: j=1,\ldots,n_i; i=1,\ldots,M\}$ collects all observations. Define $\psi_{i}(\hat{\beta})$ as an $n_{i} \times p$ matrix with row $j$ equaling $\psi(w_{ij}; \hat{\beta})$, for $j=1, \ldots, n_{i}$. A first-order Taylor expansion of $\bar{\psi}(W;\hat{\beta})$ around $\beta$ gives \[ 0=\bar{\psi}(W;\hat{\beta}) \approx \bar{\psi}(W;{\beta^\textup{r}}) +\Gamma(\beta^\textup{r})\,(\hat\beta-\beta^\textup{r}), \] where ${\Gamma}(\beta^\textup{r}) = \left.\frac{\partial}{\partial b^{\mathrm{T}}} \bar{\psi}(W;{b}) \right|_{b=\beta^\textup{r}}$. This yields the linearization of $Z$-estimator as $\hat{\beta} \approx \beta^\textup{r} - \Gamma(\beta^\textup{r})^{-1}\bar{\psi}(W;{\beta}^\textup{r})$, and the cluster-robust variance estimator:

equation[equation omitted — 406 chars of source]

Throughout this section, we write $w_{ij}=(y_{ij},x_{ij})$ to emphasize the response $y_{ij}$ and covariate $x_{ij}$. We impose the following assumptions on the clustered data.

assumption\begin{enumerate}[(a)] • The variables $(y_{ij}, x_{ij})$ have the same marginal distribution across $(i = 1, \ldots, M; j = 1, \ldots, n_i)$; • The variables $\{(y_{ij}, x_{ij})\}_{j=1}^{n_i}$ are independent across cluster $i$, but can be arbitrarily correlated within the same cluster $i$. \end{enumerate}

Specifically, we define the estimand $\beta^\textup{r}$ as the unique solution to the estimating equation under random design: \[ \mathbb{E}\left[ \sum_{ij} \psi(y_{ij}, x_{ij};\beta^\textup{r}) \right] = 0. \] Define the asymptotic variance of $\hat{\beta}$ under random design as $V^\textup{r} = (\Gamma^\textup{r})^{-1}\Delta^{\textup{r}}(\Gamma^{\textup{r}})^{-\mathrm{T}}$ with

align*[align* omitted — 343 chars of source]

Theorem (ref) below indicates that under random design, $\hat{V}_{\textsc{lz}}$ in (ref) is consistent for $V^\textup{r}$.

theoremUnder random design and Assumptions (ref) and (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,x)}\) and \(\mathbb{P}_\circ = \mathbb{P}_{(y,x)}\), we have \[ (V^\textup{r})^{-1/2} \sqrt{M}(\hat{\beta} - \beta^\textup{r}) \overset{\textup{d}}{\to} \mathcal{N}\left(0, I \right) \text{ and } M \hat{V}_{\textsc{lz}} = V^\textup{r} + o(1;\mathbb{P}_{(y,x)}). \]

Theorem (ref) extends the usual $Z$-estimation asymptotic normality result with i.i.d. data to the cluster-level setting; see, for example, LiangZeger1986 for estimating equations $Z$-estimation and NeweyMcFadden1994 for general large-sample theory for $Z$-estimation. We include it here to establish notation and to serve as the random-design benchmark for the fixed- and mixed-design results below.

Fixed design

Under fixed design, the $x_{ij}$'s are fixed, or, equivalently, we condition on them. Let $X = (x_{ij})_{1\le i\le M, 1\le j \le n_i}$ denote the stacked covariate vector under fixed design. Define the estimand $\beta^\textup{f}$ as the unique solution to \[ \mathbb{E}\left[ \frac{1}{N} \sum_{ij} \psi(y_{ij},x_{ij};\beta^\textup{f} ) \mid {X} \right] = 0. \] Define $V^\textup{f} = (\Gamma^\textup{f})^{-1}\Delta^\textup{f}(\Gamma^\textup{f})^{-\mathrm{T}}$ with

align*[align* omitted — 338 chars of source]

and the asymptotic bias of $\hat{V}_{\textsc{lz}}$ in (ref) for $V^\textup{f}$ as

equation[equation omitted — 284 chars of source]
theoremUnder fixed design and Assumptions (ref) and (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{y|x}\) and \(\mathbb{P}_\circ = \mathbb{P}_{y|x}\), we have \[ (V^\textup{f})^{-1/2} \sqrt{M}(\hat{\beta} - \beta^\textup{f}) \mid X \overset{\textup{d}}{\to} \mathcal{N}\left(0, 1 \right) \text{ and } M \hat{V}_{\textsc{lz}} = V^\textup{f} + B^{\textup{f}} + o(1;\mathbb{P}_{y|x}). \]

Theorem (ref) is the clustered analogue of the fixed-regressor conservativeness result in AbadieImbensZheng2014 and Theorem (ref). Theorem (ref) indicates that under fixed design, $\hat{V}_{\textsc{lz}}$ in (ref) is conservative with asymptotic bias $B^{\textup{f}}$ in (ref). Conditional on the fixed regressors, the cluster-level estimating equations can have nonzero conditional means under misspecification, so the LZ middle matrix estimates the sum of the conditional variance and a positive semidefinite approximation-error component.

Mixed design

We partition the regressors as $X = (X_1, X_2)$. Under mixed design, our analysis is based on conditioning on part of the regressors, i.e., ${X}_2$, and take the remaining regressors ${X}_1$ as random. Define the estimand $\beta^\textup{m}$ as the solution to \[ \mathbb{E}\left[ \frac{1}{N} \sum_{ij} \psi(y_{ij},x_{ij};\beta^\textup{m})\mid {X}_2 \right] = 0. \] Define $V^\textup{m} = (\Gamma^\textup{m})^{-1}\Delta^\textup{m}(\Gamma^\textup{m})^{-\mathrm{T}}$, with

align*[align* omitted — 343 chars of source]

and

equation[equation omitted — 283 chars of source]
theoremUnder mixed design and Assumptions (ref) and (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,x_1)|x_2}\) and \(\mathbb{P}_\circ = \mathbb{P}_{(y,x_1)\mid x_2}\), we have \begin{equation*} (V^m)^{-1/2} \sqrt{M}(\hat{\beta} - \beta^m) \mid X_2 \overset{d}{\to} \mathcal{N}\left(0, I \right) and M \hat{V}_{lz} = V^m + B^\textup{m} + o(1;\mathbb{P}_{(y,x_1)\mid x_2}). \end{equation*}

Theorem (ref) is the clustered analogue of the mixed-design result in Theorem (ref). This setting is particularly relevant for cluster-randomized experiments with fixed or pre-determined covariates. The theorem shows that under mixed design, the usual LZ variance estimator $\hat{V}_{\textsc{lz}}$ in (ref) remains conservative under misspecification, with the additional term $B^\textup{m}$ capturing the approximation-error component induced by conditioning on the fixed part of the regressors.

remarkAnalogous to Section (ref), the $Z$-estimation framework in Section (ref) applies directly to the OLS regression as a special case: $y_{ij} = x_{ij}^\mathrm{T}\beta + \varepsilon_{ij}$, where $\beta$ is the coefficient in the population linear projection of \(y_{ij}\) on \(x_{ij}\). From that theory one can show that the cluster-robust variance estimator is consistent under a random design and remains conservative under fixed or mixed designs. We omit the details to avoid repetitiveness and turn directly to its role in cluster randomized trials in Section (ref).

Cluster randomization and regression analysis

Consider a study with $N$ units, clustered, for example, by classrooms or villages. Cluster $i$ has $n_i$ units $(i=1, \ldots, M)$, and the total number of units is $N=\sum_{i=1}^M n_i$. Let $(i, j)$ index the $j$th unit within cluster $i$ $(i=1, \ldots, M ; j=1, \ldots, n_i)$. Unit $(i, j)$ has covariates $x_{i j}$. Let $Z_i$ be the treatment indicator for cluster $i$ and $Z_{i j}$ be the treatment indicator for unit $(i, j)$. In a cluster-randomized experiment, units within a cluster receive identical treatment levels. So if cluster $i$ receives treatment, then $Z_{i j}=Z_i=1$; if cluster $i$ receives control, then $Z_{i j}=Z_i=0$. Let $\mathcal{T}= \{(i, j): Z_{i j}=1 \}$ be the indices of units under treatment and $\mathcal{C}=\{(i, j): Z_{i j}=0\}$ be the indices of units under control. Their cardinalities $n_{\mathcal{T}}=\sum_{i j} Z_{i j}$ and $n_{\mathcal{C}}=\sum_{i j}(1-Z_{i j})$ represent the total numbers of units under treatment and control, respectively. Similar to Section (ref), we omit the fixed-design analysis, since it does not quantify the advantages of randomization with a misspecified linear model.

For unit $(i, j)$, let $y_{ij}(1)$ and $y_{ij}(0)$ be the potential outcomes under treatment and control, respectively. The observed outcome is related to the potential outcomes through $y_{ij}=Z_{ij}y_{ij}(1)+(1-Z_{ij})y_{ij}(0)$, for $i=1,\ldots,M$ and $j=1,\ldots,n_i$. A central goal in analysing a cluster-randomized experiment is to make inference using the observed data $\{(Z_{ij}, y_{ij}, x_{ij}): i = 1, \ldots, M; j = 1, \ldots, n_i\}$. SuDing2021 introduced the following notation to measure the heterogeneity of the cluster sizes: \[ \omega_i= \frac{n_i}{N}, \quad \Omega=\max _{1 \leq i \leq M} \omega_i, \quad \tilde{\omega}_i=\omega_i M= \frac{n_i}{N / M}. \] When all clusters have equal sizes, $\omega_i = \Omega = 1/M$.

In this subsection, we impose the following assumption.

assumption\begin{enumerate}[(a)] • The variables $(y_{ij}(0),y_{ij}(1),x_{ij})$ have identical marginal distribution across $(i = 1, \ldots, M; j = 1, \ldots, n_i)$; • The variables $((y_{ij}(0),y_{ij}(1),x_{ij}):j = 1, \ldots, n_i)$ are independent across $i$ but allow for arbitrary dependence within cluster; • $(y_{ij}(0), y_{ij}(1), x_{ij}:j = 1, \ldots, n_i) \perp \!\!\! \perp Z_i$ for $i = 1, \ldots, M$; • $\mathbb{P}_\circ(Z_i=1)= e$; • $\mathbb{E}_\circ(y_{ij}^4)<\infty$; • $\mathbb{E}_\circ\left( \|x_{ij}\|^4 \right) <\infty$; • $\Omega = o(M^{-2/3})$. \end{enumerate}

Without covariates

We consider the OLS fit with individual-level data:

equation[equation omitted — 73 chars of source]

Denote by $\hat{\beta}_{\textsc{i}}$ the coefficient of $Z_{ij}$ from the above OLS fit and $\hat{\varepsilon}_{i j}$ the residual from the same OLS fit. Define $X_{i}$ as an $n_{i} \times 2$ matrix with row $j$ equaling $(1, Z_{i j}), j=1, \ldots, n_{i}$. Stack $X_{i}$ together to obtain an $n \times 2$ matrix $X$. Define an $n_{i} \times n_{i}$ matrix $\hat{U}_{i}=\left(\hat{\varepsilon}_{i j} \hat{\varepsilon}_{i k}\right)_{1 \leq j, k \leq n_{i}}$. The cluster-robust variance estimator of $\hat{\beta}_{\textsc{i}}$ can be derived as the $(2,2)$ entry of $\hat{V}_{\textsc{lz}}$ in (ref):

align[align omitted — 255 chars of source]

Define $\varepsilon_{ij}(z) = y_{ij}(z) - \mathbb{E}(y_{ij}(z))$ and ${\varepsilon}_{i\cdot,\textsc{i}}(z) = \sum_{j=1}^{n_i}\varepsilon_{ij}(z)M/N$ for $z=0,1$. By the theory in Section (ref), we derive the asymptotic variance of $\hat{\beta}_{\textsc{i}}$ as

equation*[equation* omitted — 214 chars of source]
theoremUnder random design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z)}\) and \(\mathbb{P}_\circ = \mathbb{P}_{(y,Z)}\), we have $\beta^{\textup{r}} = \mathbb{E}(y_{ij}(1) - y_{ij}(0))$ and \begin{equation*} \left( V_{i}^{r} \right)^{-1/2} M^{1/2}(\hat{\beta}_{i} - \beta^{r}) \overset{d}{\to} \mathcal{N} \left( 0, 1 \right) and M \hat{V}_{\textsc{lz,i}} = V_{\textsc{i}}^{\textup{r}} + o(1;\mathbb{P}_{(y,Z)}). \end{equation*}

The coefficient of $Z_i$ from individual-level OLS fit in (ref) identifies the ATE. Theorem (ref) provides a random-design counterpart to existing design-based results for cluster-randomized experiments. The identification of $\beta^{\textup{r}}$ follows from cluster-level random assignment, while its asymptotic normality and the consistency of the LZ variance estimator follow from clustered-regression asymptotics; see LiangZeger1986 and HansenLee2019a. Related design-based results for individual-level regressions in cluster-randomized experiments appear in Schochet2013 and SuDing2021.

With additive covariates

We consider the OLS fit with additive regressors:

equation[equation omitted — 84 chars of source]

Let $\hat{\beta}_{\textsc{f}}$ denote the coefficient of $Z_{ij}$ from the above OLS fit and $\hat{\varepsilon}_{ij,\textsc{f}}$ denote the residual from the same OLS fit. Define $X_{i,\textsc{f}}$ as an $n_{i} \times(2+ p_{x})$ matrix with row $j$ equaling $(1, Z_{i j}, {x}_{i j}^{\mathrm{T}}), j=1, \ldots, n_{i}$. Stack $X_{i,\textsc{f}}$ together to obtain an $n \times(2+p_{x})$ matrix $X_{\textsc{f}}$. Define $\hat{\varepsilon}_{ij,\textsc{f}}$ as the residual from the OLS fit in (ref). Define an $n_{i} \times n_{i}$ matrix $\hat{U}_{i,\textsc{f}}=(\hat{\varepsilon}_{ij,\textsc{f}} \hat{\varepsilon}_{ik,\textsc{f}})_{1 \leq j, k \leq n_{i}}$. The cluster-robust variance estimator of $\hat{\beta}_{\textsc{f}}$ equals

equation[equation omitted — 350 chars of source]

\paragraph*{Random Design} Define ${\varepsilon}^{\textup{r}}_{i\cdot,\textsc{f}}(z) = \sum_{j=1}^{n_i} (\varepsilon^{\textup{r}}_{i j}(z) -\dot{x}_{ij}^{\mathrm{T}} \gamma^{\textup{r}}_{\textsc{f}}) M/N $ for $z\in\{0,1\}$ with $\gamma^{\textup{r}}_{\textsc{f}}$ being the probability limit of the coefficient of ${x}_{ij}$ from the OLS fit in (ref). By the theory in Section (ref), we derive the asymptotic variance of $\hat{\beta}_{\textsc{f}}$ under random design as

equation*[equation* omitted — 237 chars of source]
theoremUnder random design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z,x)}\) and \(\mathbb{P}_\circ = \mathbb{P}_{(y,Z,x)}\), we have $\beta^{\textup{r}} = \mathbb{E}(y_{ij}(1) - y_{ij}(0))$ and \begin{equation*} \left( V_{{f}}^{r} \right)^{-1/2} M^{1/2}(\hat{\beta}_{f} - \beta^{r}) \overset{d}{\to} \mathcal{N} \left( 0, 1 \right) and M \hat{V}_{\textsc{lz,f}} = V_{{\textsc{f}}}^{\textup{r}} + o(1;\mathbb{P}_{(y,Z,x)}), \end{equation*}

The coefficient of $Z_{ij}$ from individual-level OLS fit in (ref) identifies the ATE. Under random design, Theorem (ref) further establishes the asymptotic normality of $\hat\beta_{\textsc{f}}$ and the consistency of $\hat{V}_{\textsc{lz,f}}$ for $V_{\textsc{f}}^{\textup{r}}$. As in Theorem (ref), these conclusions follow from the general clustered-sample asymptotic theory of LiangZeger1986 and HansenLee2019a.

\paragraph*{Mixed Design} Stack the covariates in cluster $i$ to obtain $x_i = (x_{ij}: j = 1, \ldots, n_i)$. Define ${\varepsilon}^{\textup{m}}_{i\cdot,\textsc{f}}(z) = \sum_{j=1}^{n_i} (\varepsilon^{\textup{m}}_{i j}(z) - \ddot{x}_{ij}^{\mathrm{T}} \gamma^{\textup{m}}_{\textsc{f}}) M/N$ for $z\in\{0,1\}$ with $\gamma^{\textup{m}}_{\textsc{f}}$ being the probability limit of the coefficient of $\ddot{x}_{ij}$ from the OLS fit in (ref). By the theory in Section (ref), we derive the asymptotic variance of $\hat{\beta}_{\textsc{f}}$ under mixed design as

align*[align* omitted — 421 chars of source]

and the asymptotic bias of $\hat{V}_{\textsc{lz,f}} $ for $V^{\textup{m}}_{\textsc{f}}$ as

equation[equation omitted — 239 chars of source]
theoremUnder mixed design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z)|x}\) and \(\mathbb{P}_\circ = \mathbb{P}_{(y,Z)|x}\), we have $\beta^{\textup{m}} = N^{-1}\sum_{ij}\mathbb{E}\left(y_{ij}(1) - y_{ij}(0) \mid x_{ij} \right)$ and \begin{equation*} \left( V_{{f}}^{m} \right)^{-1/2} M^{1/2}(\hat{\beta}_{{f}} - \beta^{m}) \mid X \overset{d}{\to} \mathcal{N} \left( 0, 1 \right) and M \hat{V}_{\textsc{lz,f}} = V_{{\textsc{f}}}^{\textup{m}} + B_{{\textsc{f}}}^{\textup{m}} + o(1;\mathbb{P}_{(y,Z)|x}). \end{equation*}

The coefficient of $Z_{ij}$ from individual-level OLS fit in (ref) identifies the ATE. In parallel to Theorem (ref), Theorem (ref) indicates that under mixed design, $\hat{V}_{\textsc{lz,f}}$ in (ref) is conservative with asymptotic bias $B^{\textup{m}}_{\textsc{f}}$ in (ref).

With fully-interacted covariates

Define $\bar{x} = N^{-1} \sum_{ij} x_{ij}$ and $\ddot{x}_{ij} = x_{ij} - \bar{x}$. Consider the OLS fit with fully-interacted covariates:

equation[equation omitted — 112 chars of source]

Let $\hat{\beta}_{\textsc{l}}$ denote the coefficient of $Z_{ij}$ from the above OLS fit and $\hat{\varepsilon}_{ij,\textsc{l}}$ denote the residual from the same fit. Define $X_{i,\textsc{l}}$ as an $n_{i} \times (2+2 p_{x})$ matrix with row $j$ equaling $(1, Z_{i j}, \ddot{x}_{i j}^{\mathrm{T}}, Z_{ij} \ddot{x}_{i j}^{\mathrm{T}}), j=1, \ldots, n_{i}$. Stack $X_{i,\textsc{l}}$ together to obtain an $n \times(2+2 p_{x})$ matrix $X_{\textsc{l}}$. Define $\hat{\varepsilon}_{ij,\textsc{l}}$ as the residual from the above OLS fit. Define an $n_{i} \times n_{i}$ matrix $\hat{U}_{i,\textsc{l}}=(\hat{\varepsilon}_{ij,\textsc{l}} \hat{\varepsilon}_{i k,\textsc{l}})_{1 \leq j, k \leq n_{i}}$.

\paragraph*{Random Design} We apply the $Z$-estimation theory under random design in Section (ref) to obtain the asymptotic variance of the estimator $\hat{\beta}_{\textsc{l}}$. Define $\varepsilon_{ij}^{\textup{r}}(z) = y_{ij}(z) - \mathbb{E}(y_{ij}(z) )$ and let $\gamma^{\textup{r}}_z$ be the coefficient of $\ddot{x}_{ij}$ in the OLS fit of $\varepsilon^{\textup{r}}_{ij}(z)$ on $\ddot{x}_{ij}$ over $\{i: Z_i = z\}$. Define ${\varepsilon}^{\textup{r}}_{i\cdot,\textsc{l}}(z) = \sum_{j=1}^{n_i} (\varepsilon^{\textup{r}}_{ij}(z) - \ddot{x}_{i j}^{\mathrm{T}} \gamma^{\textup{r}}_z )M/N$ and ${x}_{i\cdot} = \sum_{j=1}^{n_i} x_{i j}(z) M/N$. By the theory in Section (ref), we derive the asymptotic variance of $\hat{\beta}_{\textsc{l}}$ under random design as

equation[equation omitted — 387 chars of source]

which is the the asymptotic variance of $\hat{\beta}_{\textsc{l}}$ under random design, as established in Theorem (ref). We consider the following variance estimator for $V^{\textup{r}}_{\textsc{l}} $:

align[align omitted — 294 chars of source]

where

equation[equation omitted — 351 chars of source]

is the LZ variance estimator from regression in (ref). Similar to Theorem (ref), we require a correction term in the variance estimator to account for the additional uncertainty introduced by centering the covariates. This issue cannot be resolved by directly applying Theorem (ref) to the OLS fit in (ref). Accordingly, we employ augmented estimating equations to derive the consistent variance estimator in (ref), in parallel with the development in Section (ref).

theoremUnder random design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z,x)}\) and \(\mathbb{P}_\circ = \mathbb{P}_{(y,Z,x)}\), we have $\beta^{\textup{r}} = \mathbb{E}(y_{ij}(1) - y_{ij}(0))$ and \[ \left( V_{\textsc{l}}^{\textup{r}} \right)^{-1/2} M^{1 / 2}\left(\hat{\beta}_{\textsc{l}}-\beta^{\textup{r}} \right) \overset{\textup{d}}{\to} \mathcal{N}(0,1) \text{ and } M \hat{V}_{\textsc{lz},\textsc{l},\text{adj}} = V^{\textup{r}}_{\textsc{l}} + o(1;\mathbb{P}_{(y,Z,x)}). \] Moreover, \[ M \hat{V}_{\textsc{lz,l}} = V_{\textsc{l}}^{\textup{r}} - B_{\textsc{l}}^{\textup{r}} + o(1;\mathbb{P}_{(y,Z,x)}) \] with \begin{equation} B_{l}^{r} = (\gamma^{r}_1 - \gamma^{r}_0)^{\mathrm{T}} \frac{1}{M} \sum_{i=1}^M \mathbb{V}\left( {x}_{i\cdot} \right) (\gamma^{r}_1 - \gamma^{r}_0). \end{equation}

The coefficient of $Z_i$ from individual-level OLS fit in (ref) identifies the ATE. Theorem (ref) indicates that under random design, $\hat{V}_{\textsc{lz,l},\text{adj}}$ in (ref) is consistent. Importantly, the adjusted estimator $\hat V_{\textsc{lz},\textsc{l},\text{adj}}$ in (ref) is a novel contribution. Moreover, under random design, the cluster-robust variance estimator $\hat{V}_{\textsc{lz},\textsc{l}}$ in (ref) is anti-conservative with asymptotic bias $B_{\textsc{l}}^{\textup{r}}$ in (ref). SuDing2021 establish the validity of $\hat V_{\textsc{lz,l}}$ under design-based framework. where $x$ is fixed.

remarkSimilar to Remark (ref), we can also obtain consistent variance estimator based on the cluster-robust covariance estimator with the augmented estimating equations, though with finite sample difference to $\hat{V}_{\textsc{lz},\textsc{l},\text{adj}}$.

\paragraph*{Mixed Design} Define $\varepsilon_{ij}^{\textup{m}}(z) = y_{ij}(z) - N^{-1} \sum_{ij} \mathbb{E}(y_{ij}(z) \mid x_{i} )$ and let $\gamma^{\textup{m}}_z$ be the coefficient of $\ddot{x}_{ij}$ in the OLS fit of $\varepsilon^{\textup{m}}_{ij}(z)$ on $\ddot{x}_{ij}$ over $\{i: Z_i = z\}$. Define ${\varepsilon}^{\textup{m}}_{i\cdot,\textsc{l}}(z) = \sum_{j=1}^{n_i} (\varepsilon^{\textup{m}}_{ij}(z) - \ddot{x}_{i j}^{\mathrm{T}} \gamma^{\textup{m}}_z )M/N$. By the theory in Section (ref), we derive the asymptotic variance of $\hat{\beta}_{\textsc{l}}$ under mixed design as

align*[align* omitted — 421 chars of source]

and the asymptotic bias of $\hat{V}_{\textsc{lz,l}}$ in (ref) for $V^{\textup{m}}_{\textsc{l}}$ as

equation[equation omitted — 237 chars of source]
theoremUnder mixed design and Assumption (ref) with \(\mathbb{E}_\circ = \mathbb{E}_{(y,Z)|x}\) and \(\mathbb{P}_\circ = \mathbb{P}_{(y,Z)|x}\), we have $\beta^{\textup{m}} = N^{-1}\sum_{ij}\mathbb{E}\left(y_{ij}(1) - y_{ij}(0) \mid x_{i} \right)$ and \begin{equation*} \left( V_{l}^{m} \right)^{-1/2} M^{1 / 2}\left(\hat{\beta}_{l}-\beta^{m} \right) \mid X \overset{d}{\to} \mathcal{N}(0,1) and M \hat{V}_{\textsc{lz,l}} = V_{\textsc{l}}^{\textup{m}} + B_{\textsc{l}}^{\textup{m}} + o(1;\mathbb{P}_{(y,Z)|x}). \end{equation*}

The coefficient of $Z_{ij}$ from individual-level OLS fit in (ref) identifies the conditional average treatment effect. Theorem (ref) indicates that under mixed design, $\hat{V}_{\textsc{lz,l}}$ in (ref) is conservative with asymptotic bias $B^{\textup{m}}_{\textsc{l}}$ in (ref). Unlike in the random-design case, centering around covariates introduces additional uncertainty, whereas under a mixed design, no such issue arises, and the variance estimator remains conservative.

Simulation

We report simulation results for complete randomization and cluster randomization. For each setting, we consider: random and mixed designs, because the fixed design formulation assumes away the benefits of randomization. In the random design, all random variables are independently redrawn in each simulation iteration. In the mixed design, we hold $\{x_i\}_{i=1}^n$ fixed while redrawing all other random variables. To examine robustness to model misspecification, the outcome equations include higher order of covariate $x_i$. For each case, we report the estimand (“estimand”), the simulation estimate (“estimate”), the asymptotic standard error (“asym SE”), the estimated standard error (EHW for i.i.d. data under “EHW SE” and LZ for clustered data under “LZ SE”), and the empirical coverage of the 95% confidence interval based on the corresponding variance estimator (“coverage”). Rows labeled “Fisher” correspond to regressions with covariates, and rows labeled “Lin” correspond to regressions with fully interacted covariates. For the fully interacted specification, we also report the corrected EHW or LZ variance estimator and the coverage of the associated 95% CI, marked with an asterisk (*).

Complete randomization

We conduct a Monte Carlo simulation with $n = 1000$ observations per sample and $B = 5000$ simulation replications. We generate the covariate $x_i$ independently from a standard normal distribution: $x_i \overset{\text{i.i.d.}}{\sim} \mathcal{N}(0,1)$. Treatment assignment is randomized such that: $\mathbb{P}(Z_i = 1) = 0.5$. Potential outcomes are generated as follows:

align*[align* omitted — 214 chars of source]

where the parameters are set as $(\mu_1, \mu_0, \gamma_1, \gamma_0) = (3,2,2,1)$. The observed outcome is defined as: \[ y_i = Z_i Y_i(1) + (1 - Z_i) Y_i(0). \] We present the results for three OLS regression specifications: without covariates, with covariates, and with fully interacted covariates. As established in Theorem (ref), the difference-in-means estimator is consistent under the random design, and the EHW variance estimator is also consistent in this setting. Similarly, Theorems (ref) and (ref) verifies that the estimator from the regression with covariates is consistent under both the random and mixed designs. The EHW variance estimator remains consistent under the random design but is conservative under the mixed design. For the regression with fully interacted covariates, Theorems (ref) and (ref) confirms that the estimator is consistent under both random and mixed designs. However, the EHW variance estimator is anti-conservative under the random design and conservative under the mixed design. Nonetheless, the corrected EHW variance estimator is consistent under the random design, as verified by Theorem (ref).

Cluster randomization

We conduct a Monte Carlo simulation with $M = 160$ clusters per sample and $B = 5000$ simulation replications. The cluster sizes are drawn from a uniform distribution. Each cluster $i$ contains $n_i$ units where $n_i=\operatorname{round}\!\left(\frac{1000}{M}U_i\right)$ with $U_i \sim \operatorname{Unif}(0.6, 1.4)$. where: The cluster sizes $n_i$ are fixed across simulation replications. Each unit $(i, j)$ has a covariate drawn independently from a standard normal distribution: $x_{i j} \overset{\text{i.i.d.}}{\sim} \mathcal{N}(0,1)$. The covariance matrix $\Sigma_i$ for each cluster $i$ is generated as: \[ \Sigma_{i} = \frac{(\tilde{\Sigma}_{ij}: j=1,\ldots,n_i)(\tilde{\Sigma}_{ij}: j=1,\ldots,n_i)^{\mathrm{T}}}{\sqrt{(\tilde{\Sigma}_{ij}: j=1,\ldots,n_i)(\tilde{\Sigma}_{ij}: j=1,\ldots,n_i)^{\mathrm{T}}}_{ii}}, \] where the elements $\tilde{\Sigma}_{ij}$ are drawn independently from a uniform distribution: $\tilde{\Sigma}_{ij} \overset{\text{i.i.d.}}{\sim} \text{Unif}[0,1]$. The potential outcomes for each unit in cluster $i$ are generated as:

align*[align* omitted — 229 chars of source]

The error terms are drawn from multivariate normal distributions with covariance matrices that vary by treatment status, with the treated potential outcome having a larger variance structure than the control.

We present the results for three OLS regression specifications: without covariates, with covariates, and with fully interacted covariates. As established in Theorem (ref), the difference-in-means estimator is consistent under the random design, and the LZ variance estimator is also consistent in this setting. Similarly, Theorems (ref) and (ref) verify that the estimator from the regression with covariates is consistent under both the random and mixed designs. The LZ variance estimator remains consistent under the random design but is conservative under the mixed design. For the regression with fully interacted covariates, Theorems (ref) and (ref) confirm that the estimator is consistent under both random and mixed designs. However, the LZ variance estimator is anti-conservative under the random design and conservative under the mixed design. Nonetheless, the corrected LZ variance estimator is consistent under the random design, as verified by Theorem (ref).

\newcolumntype{P}[1]{>{\arraybackslash}p{#1}} \newcolumntype{Y}{>{\arraybackslash}X}

table[table omitted — 1,885 chars of source]

Discussion

We unify the literature by developing a general theory for $Z$-estimation with random, fixed, and mixed regressors, covering both independent and clustered data. We clarify how regression coefficients should be interpreted under misspecification and different sources of randomness, and derive the corresponding inference theory. The usual robust variance estimator is consistent under random design but generally conservative under fixed and mixed designs. We apply these results to OLS and regression adjustment in completely randomized and cluster-randomized experiments, characterize the associated causal estimands, and provide corrections for conventional variance estimators in fully interacted specifications with random covariates.

Our framework complements the design-based inference literature AbadieAtheyImbens2020 by clarifying how different sources of sampling randomness affects both parameter interpretation and robust inference. Although robust variance estimators may be conservative under both mixed and design-based analyses, the sources of conservativeness differ. In our mixed-design framework, the additional term arises from nonzero conditional means of the estimating equations under misspecification, whereas in design-based inference conservativeness typically reflects unidentified treatment-effect heterogeneity.

Although our applications focus on linear and instrumental-variable regressions, the $Z$-estimation inference result applies more broadly to nonlinear models. In such settings, however, the interpretation of the resulting parameters is more challenging and requires further study. Under misspecification, such models generally target pseudo-true parameters defined by their population objectives or estimating equations. Their interpretation depends on the estimation criterion: likelihood-based estimators minimize the expected Kullback--Leibler divergence from the true conditional distribution, whereas nonlinear least squares provides an $L^2$ approximation to the conditional mean. Under mixed design, these objectives are evaluated conditional on the fixed regressors. The resulting pseudo-true parameter may therefore depend on which regressors are treated as random and which are conditioned on. This distinction also matters for inference. Under likelihood misspecification White1982, the sandwich variance accounts separately for the curvature of the objective and the variance of the score. More generally, for $Z$-estimation, it accounts separately for the sensitivity and variability of the estimating equations. Under mixed design, the additional term $B^{\mathrm m}$ arises when the estimating equations have nonzero conditional means given the fixed regressors. Valid inference for a pseudo-true parameter, however, does not by itself provide that parameter with a causal or otherwise substantive interpretation.

AngristChernozhukovFernandez-Val2006 study the inference and interpretation of quantile regression under random design, whereas AbadieImbensZheng2014 consider a covariate-conditional estimand but derive unconditional asymptotic results. Extending these ideas to fixed and mixed designs is a natural direction. Under misspecification, quantile-regression coefficients can be viewed as weighted linear approximations to conditional quantile functions. In a mixed-design framework, however, the approximation target is evaluated conditional on the realized fixed regressors. Hence both the interpretation and the asymptotic variance may depend on the source of randomness. Because quantile regression is nonsmooth, this extension requires separate asymptotic arguments beyond our differentiable \(Z\)-estimation framework.

Acknowledgement

Peng Ding is partially supported by the U.S. National Science Foundation \# 2514234.

Supplementary material

The supplementary material includes the results for the local average treatment effect framework and proofs of all theorems.