EconBase
← Back to paper

Bootstrapping with AI/ML-generated labels

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

74,167 characters · 20 sections · 27 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bootstrapping with AI/ML-generated labels

\onehalfspacing

\thispagestyle{empty}

abstract\singlespacing AI/ML methods are increasingly used in economics to generate binary variables (or labels) via classification algorithms. When these generated variables are included as covariates in regressions, even small misclassification errors can induce large biases in OLS estimators and invalidate standard inference. We study whether the bootstrap can correct this bias and deliver valid inference. We first show that a seemingly natural fixed-label bootstrap, which generates data using estimated labels but relies on a corrupted version in estimation, is generally invalid unless a strong independence condition between the latent true labels and other covariates holds. We then propose a coupled-label bootstrap that jointly resamples the true and imputed labels, and show it is valid without this condition. Two finite-sample adjustments further improve coverage: a variance correction for uncertainty in estimated misclassification rates and a Hessian rotation for near-singular designs. We illustrate the methods in simulations and apply them to investigate the relationship between wages and remote work status. Keywords: bootstrap, bias correction, binary labels, misclassification, AI/ML-generated data, inference.

\setcounter{page}{1}

\pagenumbering{arabic}

Introduction

It is now standard practice in economics and across the social sciences to generate new variables and measurements using artificial intelligence (AI) and machine learning (ML) methods. A leading use case is the generation of binary variables, or labels, which are imputed using classification algorithms. To give just a few examples within economics, goldsmith-pinkhamGenderGapHousing2023 impute the gender of borrowers from borrower names, bursztynImmigrantNextDoor2024 impute the race of charitable donors from donor names, adams-prasslFirmConcentrationJob2023 and hansenRemoteWorkJobs2026 classify the flexibility of work arrangements from job postings metadata, and \citetalias{dupasGenderDifferencesEconomics2026}\ (dupasGenderDifferencesEconomics2026) impute speaker gender and tone-of-voice from audio recordings of economics seminars. These imputed variables are rarely of interest themselves. Rather, they are used as inputs to econometric models for downstream inference tasks.

While the easy generation of labels has opened the door for new research, it is not without significant methodological challenges. battagliaInferenceRegressionVariables2025 (battagliaInferenceRegressionVariables2025, \citetalias{battagliaInferenceRegressionVariables2025} hereafter) show theoretically and empirically that even high performance classifiers with error rates below 1% can generate substantial biases in parameter estimates due to misclassification error in the generated labels. Moreover, failure to correct for these biases leads to invalid inference on regression parameters. There is therefore a need for methods that correct bias and restore valid inference.

A central challenge in developing such methods is data constraints. A widely discussed approach for correcting bias requires large validation samples, in which both the imputed variables and ground-truth labels are observed alongside other variables in the model.\footnote{In economics, this approach goes back to boundExtentMeasurementError1991, carrollSemiparametricEstimationLogistic1991, sepanskiSemiparametricQuasilikelihoodVariance1993, boundEvidenceValidityCrossSectional1994, and leeEstimationLinearNonlinear1995, among others, and has connections with the literature on auxiliary data chenMeasurementErrorModels2005,chenSemiparametricEfficiencyGMM2008. Recent papers proposing these methods within the context of AI/ML-generated data in economics include ludwigLargeLanguageModels2025 and carlsonUnifyingFrameworkRobust2026. Similar approaches are taken in the recent statistics literature on “prediction-powered inference” angelopoulosPredictionpoweredInference2023,angelopoulosPPIEfficientPredictionPowered2023. See also fongMachineLearningPredictions2021 and egamiUsingLargeLanguage2024, among others, for related work in political science.} In many economic applications, however, such data are unavailable. Instead, researchers may only observe an external sample from which the accuracy of the classifier can be measured. For instance, bursztynImmigrantNextDoor2024 estimate the accuracy of their classifier using an external sample of North Carolina voter registration data which contains self-reported ethnicity, but is missing data on charitable donations. Moreover, because collecting ground-truth labels is typically much more costly than generating labels with a classifier, the size $m$ of this external sample is usually much smaller than the sample size $n$ used for model estimation ($m \ll n$). \citetalias{battagliaInferenceRegressionVariables2025} propose analytic bias corrections that account for both these data constraints.

In this paper, we study whether the bootstrap can correct bias and deliver valid inference. There is a long tradition of using the bootstrap to reduce bias and improve the coverage of confidence intervals efronNonparametricStandardErrors1981,hallBootstrapEdgeworthExpansion1992. In our setting, however, the problem is deceptively difficult. With binary labels, measurement error is necessarily “nonclassical” aignerRegressionBinaryIndependent1973: a misclassified 1 must be a 0, and vice versa. Consequently, common intuitions and methods developed in settings with “classical” measurement error do not easily carry over.

More formally, consider the regression model \[

alignedY_i & = X_i' \beta + u_i , \\ X_i & = g(\theta_i, Z_i),

\] where $\theta_i$ is a latent binary variable, $Y_i$ and $Z_i$ are observed, and the functional form of $g$ is known. For example, $g(\theta_i,Z_i) = (\theta_i, Z_i')'$ when $\theta_i$ enters the regression additively, or $g(\theta_i,Z_i) = (\theta_iZ_i', Z_i')'$ when it enters through interactions. In practice, researchers replace $\theta_i$ with an estimate $\hat \theta_i$, regress $Y_i$ on $\hat X_i \equiv g(\hat \theta_i, Z_i)$, and report confidence intervals and standard errors for $\beta$ treating $\hat X_i$ as though it is the truth. Misclassification ($\hat \theta_i \neq \theta_i$ for some observations) induces measurement error in the $\hat X_i$, so regression estimators can be biased and standard inference can be invalid.

As in \citetalias{battagliaInferenceRegressionVariables2025}, we assume the researcher observes a large sample $(Y_i, Z_i, \hat \theta_i)_{i=1}^n$ and a much smaller external sample $(\theta_i, \hat \theta_i)_{i=1}^m$ used to estimate classifier accuracy.\footnote{In fact, neither we nor \citetalias{battagliaInferenceRegressionVariables2025} need the full external sample, just estimates $\hat F_+$ and $\hat F_-$ of the false positive and negative rates of the classifier, and the sample size $m$ used to estimate those rates.} The question we address is: how should one bootstrap in this scenario?

Naively resampling $(Y_i,\hat X_i)$, or using equivalent residual-based approaches such as the wild bootstrap, does not work. These methods do not introduce any measurement error in the covariates in the bootstrap world, because the same covariates are used in data generation and parameter estimation. As a result, they fail to reproduce the bias of the OLS estimator and yield invalid inference.

A seemingly natural approach is a fixed-label bootstrap, a variant of the wild bootstrap that builds in measurement error. Outcomes are generated as $Y_i^* = \hat X_i' \hat \beta + u_i^*$, where $\hat \beta$ is the OLS estimator from regression of $Y_i$ on $\hat X_i$, and $u_i^*$ is a bootstrap version of the OLS residual. Then, $Y_i^*$ is regressed on $\hat X_i^* \equiv g(\hat \theta_i^*, Z_i)$, where $\hat \theta_i^*$ is a corrupted version of $\hat \theta_i$ constructed so that the false positive and negative rates for predicting $\hat \theta_i$ with $\hat \theta_i^*$ match those observed in the external sample. Analogues of this bootstrap have been used to correct bias in other settings with generated covariates, including factor-augmented regressions goncalvesBootstrappingFactoraugmentedRegression2014,goncalvesBootstrapInferenceGroup2025 and panel data models goncalvesBootstrapInferenceLinear2015,higginsBootstrapInferenceFixedEffect2024,liConfidenceIntervalsTreatment2024. However, we show that this bootstrap is generally inconsistent: it has the correct asymptotic variance but fails to replicate the asymptotic bias of the OLS estimator unless certain independence conditions hold between $\theta_i$ and $Z_i$.

To address this problem, we introduce a novel coupled-label bootstrap. Outcomes are generated as $Y_i^* = X_i^{*\prime} \hat \beta + u_i^*$, where $X_i^* = g(\theta_i^*,Z_i)$, $\theta_i^*$ is a bootstrap version of the true latent $\theta_i$, and $\hat \beta$ and $u_i^*$ are defined above. Estimation then proceeds by regressing $Y_i^*$ on $\hat X_i^*$ as before. The key innovation is to generate the pairs $(\theta_i^*, \hat \theta_i^*)$ jointly so that the misclassification rates for predicting $\theta_i^*$ with $\hat \theta_i^*$ match those in the external sample. We show that this bootstrap reproduces the correct asymptotic bias and variance of the OLS estimator, which justifies the use of confidence sets and hypothesis tests constructed via this bootstrap.

While our results establish asymptotic validity of the coupled-label bootstrap, two further modifications improve its finite-sample performance. The first is a variance correction that accounts for additional sampling error from estimating the misclassification rates. While \citetalias{battagliaInferenceRegressionVariables2025} develop analytic corrections, we show that a similar adjustment can be implemented via a simple modification of the bootstrap. The second is a rotation of the inverse Hessian, which is useful in near-singular designs where only a small fraction of $\theta_i$ are 0 (or 1). In these settings, misclassification can introduce discrepancies between the bootstrap and estimated inverse Hessian that are asymptotically negligible but material in finite samples. Simulations calibrated to our empirical application show that these modifications substantially improve finite-sample performance. We therefore recommend the coupled-label bootstrap with both modifications.

The remainder of the paper is organized as follows. Section (ref) provides high-level conditions for bootstrap validity with generated covariates. Section (ref) develops theory for the binary labels case, analyzing a general wild bootstrap and providing results for the fixed-label and coupled-label bootstraps. Section (ref) introduces two modifications that improve finite-sample performance. Section (ref) presents simulation evidence in a design calibrated to the application, and Section (ref) revisits hansenRemoteWorkJobs2026, who study the relationship between wages and remote work status. All proofs are presented in the appendix.

\paragraph{Notation} Let $\mathbb{P}^*$ denote the bootstrap probability measure conditional on the data, and $\mathbb{E}^*$ the corresponding expectation. For bootstrap random vectors $Y_n^*$, we write $Y_n^* \to_{p^*} 0$ or $Y_n^* = o_{p^*}(1)$ if $\mathbb{P}^*(\|Y_n^*\| > \epsilon) \to_p 0$ for all $\epsilon > 0$. Similarly, $Y_n^* = O_{p^*}(1)$ if $\limsup_{n \to \infty} \mathbb{P}^*(\|Y_n^*\| > M) \to_p 0$ as $M \to \infty$. Finally, we write $Z_n^* \to_{d^*} Z$ to denote that $\mathbb{P}^*(Z_n^* \leq z) - F(z) \to_p 0$ for all continuity points $z$ of the CDF $F$ of $Z$.

General Case

In this section, we consider the general case in which $Y$ is regressed on a vector of generated covariates $\hat X$. We first recap results from \citetalias{battagliaInferenceRegressionVariables2025}, who show that regressing $Y$ on $\hat X$ induces a bias in the asymptotic distribution of the OLS estimator that invalidates standard inference. We then provide high-level conditions under which the bootstrap can match the asymptotic bias and variance of the OLS estimator, justifying the use of bootstrap percentile intervals to perform valid inference. The central question is then whether one can design a bootstrap satisfying these conditions in settings of interest. Section (ref) addresses this question for the binary labels case.

Asymptotic Distribution of the OLS Estimator

We wish to perform inference on the parameter vector $\beta$ in the regression model \[ Y_i = \beta' X_i + u_i, \] where the true covariates $X_i$ are latent and generated covariates $\hat X_i$ are used in their place. We are interested in the large-sample properties of the OLS estimator \[ \hat{\beta}=\Big(n^{-1}\sum_{i=1}^{n}\hat{X}_i\hat{X}^{\prime}_i\Big)^{-1}n^{-1}\sum_{i=1}^{n}\hat{X}_i Y_i \] from a regression of $Y_i$ on a vector of generated covariates $\hat X_i$.

The following regularity conditions are similar to those in \citetalias{battagliaInferenceRegressionVariables2025}:

assumption\begin{enumerate}[label={(\roman*)}] • $\frac{1}{n} \sum_{i=1}^n \|\hat X_i - X_i\|^2 \to_p 0$; • $\frac{1}{\sqrt n} \sum_{i=1}^n \hat X_i (\hat X_i - X_i)' \to_p B$; • $\frac{1}{\sqrt n} \sum_{i=1}^n (\hat X_i - X_i) u_i \to_p 0$. \end{enumerate}
assumption\begin{enumerate}[label={(\roman*)}] • $\frac{1}{n} \sum_{i=1}^n X_i X_i' \to_p Q \equiv \mathbb{E}[X_i X_i'] > 0$, for $\mathbb{E}[\|X_i\|^2] < \infty$; • $\frac{1}{\sqrt n} \sum_{i=1}^n X_i u_i \to_d N(0,\Sigma)$, for $\Sigma = \mathbb{E}[ X_i X_i' u_i^2]$. \end{enumerate}

Assumption (ref) imposes conditions on the generated variable $\hat X_i$ while Assumption (ref) is standard and only concerns the latent variable $X_i$. Assumption (ref) allows for both classical and nonclassical error in $\hat X_i$. The key condition is Assumption (ref)(ref). In the classical case, this condition requires the measurement error variance to be $O(n^{-1/2})$, similar to the small measurement error asymptotic frameworks of chesherEffectMeasurementError1991 and evdokimovSimpleEstimationSemiparametric2023. This scaling ensures that measurement error persists but does not dominate sampling error in the asymptotic distribution of $\hat \beta$. As a result, asymptotics deliver tractable and useful approximations to the finite-sample distribution of $\hat \beta$, where both sources of error play a role. Section (ref) provides sufficient conditions for Assumption (ref) for the binary labels case.

The next result follows by similar arguments to \citetalias{battagliaInferenceRegressionVariables2025} and so a proof is omitted:

theoremSuppose that Assumptions (ref) and (ref) hold. Then as $n \to \infty$, \[ \sqrt n \big(\hat \beta - \beta\big) \to_d N(b, V), \] where $b = - Q^{-1} B \beta$, and $V = Q^{-1} \Sigma Q^{-1}$.

As noted in \citetalias{battagliaInferenceRegressionVariables2025}, Theorem (ref) shows that if $b\ne 0$, then $\hat \beta$ has a first-order asymptotic bias and the usual approach to inference is invalid.

Consistency of the Bootstrap

Suppose that we generate the bootstrap data as \[ Y^*_i =\hat{\beta}^\prime X^*_i+u^*_i \] where $u^*_i$ is a bootstrap version of $\hat{u}_i=Y_i-\hat{\beta}^\prime \hat{X}_i$. For now, we do not specify a particular bootstrap method used in generating $u^*_i$. Later we will consider a wild bootstrap for i.i.d. data. Similarly, $X^*_i$ is a bootstrap analogue of the true latent $X_i$. A natural choice for $X^*_i$ is $\hat{X}_i$, as in a standard fixed-regressor bootstrap. As we will see below, allowing for $X^*_i$ to differ from $\hat{X}_i$ turns out to be important for the binary labels application, so we allow for this possibility here.

Let $\hat{X}^*_i$ denote a bootstrap analogue of $\hat{X}_i$ which will be context-specific. The bootstrap analogue of $\hat{\beta}$ is

flalign*\hat{\beta}^*=\Big(n^{-1}\sum_{i=1}^{n}\hat{X}^*_i\hat{X}^{*^\prime}_i\Big)^{-1}n^{-1}\sum_{i=1}^{n}\hat{X}^*_i Y^*_i.

The following set of high-level conditions on $(u^*_i,X^*_i,\hat{X}^*_i)$ ensure the bootstrap distribution of $\sqrt{n}(\hat{\beta}^*-\hat{\beta})$ is consistent for the distribution of $\sqrt{n}(\hat{\beta}-\beta)$ given in Theorem (ref) as $n \to \infty$. \setcounter{assumption}{0}

assumption[] \begin{enumerate}[label={(\roman*)}] • $\frac{1}{n}\sum_{i=1}^{n}\|\hat{X}^*_i-X^*_i\|^2\rightarrow_{p^*}0$ and $\frac{1}{n}\sum_{i=1}^{n}\|X^*_i-\hat{X}_i\|^2\rightarrow_{p^*}0$. • $\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\hat{X}^*_i(\hat{X}^*_i-X^*_i)^\prime \rightarrow_{p^*} B$. • $\frac{1}{\sqrt{n}}\sum_{i=1}^{n}(\hat{X}^*_i-X^*_i)u^*_i\rightarrow_{p^*} 0$ and $\frac{1}{\sqrt{n}}\sum_{i=1}^{n}(X^*_i-\hat{X}_i)u^*_i\rightarrow_{p^*} 0$. \end{enumerate}
assumption$\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\hat{X}_iu^*_i \rightarrow_{d^*} N(0,\Sigma)$.

Assumption (ref) is the bootstrap analogue of Assumption (ref). Part (ref) is used to show that $n^{-1}\sum_{i=1}^{n}\hat{X}^*_i\hat{X}^{*\prime}_i \to_{p^*} Q$, given $n^{-1}\sum_{i=1}^{n}\hat{X}_i\hat{X}_i'\to_p Q$ under Assumptions (ref) and (ref). Part (ref) and Assumption (ref) imply that $n^{-1/2}\sum_{i=1}^{n}\hat{X}^*_i u^*_i\to_{d^*} N(0,\Sigma)$. In particular, Assumption (ref) is the bootstrap analogue of Assumption (ref)(ref) and follows by the application of a bootstrap CLT. Assumption (ref) (ref) is the crucial assumption that ensures that the bootstrap mimics the bias term $B$ in Assumption (ref)(ref).

theoremSuppose that Assumptions (ref) and (ref) hold. If $(u^*_i,X^*_i,\hat{X}^*_i)$ satisfy Assumptions (ref) and (ref), then as $n\rightarrow\infty$, \begin{flalign*} \sqrt{n}\big(\hat{\beta}^*-\hat{\beta}\big)\rightarrow_{d^*} N(b,V), \end{flalign*} where $b = - Q^{-1} B \beta$, and $V = Q^{-1} \Sigma Q^{-1}$.

Theorem (ref) shows that the bootstrap correctly replicates the limiting distribution of $\hat{\beta}$, including its bias and variance, thereby justifying the construction of bootstrap percentile confidence intervals for the elements of $\beta$. Specifically, a $100(1-\alpha)\%$ percentile interval for the $j$-th element $\beta_j$ of $\beta$ is given by \[ [\hat{\beta}_j-\hat{q}_{1-\alpha/2,j},\hat{\beta}_j-\hat{q}_{\alpha/2,j}], \] where $\hat{q}_{\alpha,j}$ denotes the $\alpha$-quantile of the bootstrap distribution of $\hat{\beta}^*_j-\hat{\beta}_j$. Unlike asymptotic theory-based intervals, this interval requires neither an explicit bias correction nor a variance estimator.

Binary Labels Case

We now turn to the binary-label case. We begin by describing the setup and specializing Theorem (ref) to this setting. We then consider two bootstrap methods, both of which are instances of the wild bootstrap. The first is a seemingly natural fixed-label bootstrap, analogous to the standard fixed-regressor bootstrap, which is generally inconsistent unless strong independence conditions hold between $\theta_i$ and $Z_i$. We next present a coupled-label bootstrap and show that it is consistent in settings where the fixed-label bootstrap fails. We therefore recommend the coupled-label bootstrap, together with the finite-sample modifications described in the next section.

Setup

In this setting, \[ X_i = g(\theta_i, Z_i), \] where $g : \{0,1\} \times \mathbb R^{d_z} \to \mathbb R^k$ is known, $\theta_i$ is a latent binary random variable, and $Z_i \in \mathbb R^{d_z}$ is an observed random vector. Since $\theta_i$ is latent, it is common practice to generate an estimate $\hat \theta_i$ of $\theta_i$ using a classification algorithm, then regress $Y_i$ on $\hat X_i = g(\hat \theta_i, Z_i)$. For ease of exposition we will focus on the case of scalar $\theta_i$, though it is straightforward to extend our analysis to the vector case.

exampleThis framework includes models where $\theta_i$ is interacted with a subset of the $Z_i$ variables, such as \begin{equation} Y_i = \beta_1' \theta_i Z_{1i} + \beta_2' Z_{2i} + u_i, \end{equation} to capture, e.g., effects that vary by membership of a race or gender group. In this case, $g(\theta_i, Z_i) = (\theta_i Z_{1i}', Z_{2i}')' \in \mathbb R^{k}$.
exampleExample (ref) includes as a special case the model studied in \citetalias{battagliaInferenceRegressionVariables2025}, in which \begin{equation} Y_i = \beta_1 \theta_i + \beta_2' Z_{2i} + u_i, \end{equation} where $Z_{1i} = 1$ and $Z_{2i} = Z_i$. For this model, $g(\theta_i,Z_i)=(\theta_i,Z_i^\prime)^\prime \in \mathbb R^{k}$.

To maintain flexibility, we will not take a stand on the algorithm with which $\hat \theta_i$ is generated. Instead, we shall assume that $(Y_i,Z_i,\hat \theta_i,\theta_i)_{i=1}^n$ are drawn i.i.d. from a distribution $P_n$. Here we use the notation $P_n$ because our asymptotic framework introduced below allows the distribution of the data to vary with $n$, so that both misclassification and sampling error feature in the asymptotic distribution of the OLS estimator. The econometrician observes $(Y_i,Z_i, \hat \theta_i)_{i=1}^n$ and, for the purposes of bias correction, has access to an external sample $(\theta_i,\hat \theta_i)_{i=1}^m$, which can be used to assess the accuracy of the classifier. As in \citetalias{battagliaInferenceRegressionVariables2025}, we allow the external sample size $m$ to be of much smaller magnitude than $n$. This reflects common empirical settings with large amounts of generated data and small amounts of ground-truth observations.

Asymptotic Bias of the OLS Estimator

We now provide primitive conditions for Assumption (ref) for the binary labels case. To this end we impose some structure on $\hat \theta_i$ and $\theta_i$. We follow \citetalias{battagliaInferenceRegressionVariables2025} and adopt an asymptotic framework where the (unconditional) false-positive rate $\mathbb{E}[\hat \theta_i(1-\theta_i)]$ and (unconditional) false-negative rate $\mathbb{E}[\theta_i(1-\hat \theta_i)]$ both tend to zero as $n \to \infty$, so that both sampling error in the regression and measurement error in the $\hat \theta_i$ remain relevant in large samples. This ensures the asymptotic distribution mimics the finite-sample setting where both play roles.

To introduce the assumptions, note that we can write \[ \hat{X}_i - X_i =\hat{\theta}_i(1-\theta_i)(g(1,Z_i) - g(0,Z_i)) +\theta_i(1-\hat{\theta}_i)(g(0,Z_i) - g(1,Z_i)), \] from which it follows that \[ \hat{X_i}(\hat{X}_i - X_i )'=\hat{\theta}_i(1-\theta_i)D_{+i} +\theta_i(1-\hat{\theta}_i)D_{-i}, \] where \[

alignedD_{+i} & = g(1,Z_i)(g(1,Z_i) - g(0,Z_i))' , \\ D_{-i} & = g(0,Z_i)(g(0,Z_i) - g(1,Z_i))' .

\] Let $G_i = \max\{\|g(0,Z_i)\|,\|g(1,Z_i)\|\}$ and let $\delta > 0$.

assumption\begin{enumerate}[label={(\roman*)}] • $\mathbb{E}[X_i u_i] = 0$ and $\mathbb{E}[|u_i|^{4+\delta}] < \infty$; • $\mathbb{E}[X_i X_i'] > 0$ and $\mathbb{E}[G_i^{4+\delta}] < \infty$; • $\mathbb{E}[(\hat X_i - X_i)u_i] = o(n^{-1/2})$ as $n \to \infty$; • $\sqrt n \mathbb{E}[\hat \theta_i(1-\theta_i)] \to \kappa_+$ and $\sqrt n \mathbb{E}[\theta_i(1-\hat \theta_i)] \to \kappa_-$ as $n \to \infty$, for $0 \leq \kappa_+, \kappa_- < \infty$; • $\sqrt n ( \mathbb{E} [ \hat \theta_i(1-\theta_i) D_{+i} + \theta_i(1-\hat \theta_i) D_{-i} ] - \mathbb{E}[\hat \theta_i(1-\theta_i)] \mathbb{E}[D_{+i}] - \mathbb{E}[\theta_i(1-\hat \theta_i)] \mathbb{E}[D_{-i}] ) \to 0$ as $n \to \infty$. \end{enumerate}

Assumption (ref)(ref)--(ref) are standard regularity conditions. Assumption (ref)(ref) allows for some correlation between the measurement error in $\hat X_i$ and $u_i$, provided it vanishes at a sufficiently fast rate. For instance, suppose $\theta_i$ and $\hat \theta_i$ are functions of $Z_i$ and some auxiliary information $\mathcal U_i$. Then by iterated expectations,

flalign*\mathbb{E}[(\hat{X}_i - X_i)u_i] &=\mathbb{E}\big[\hat{\theta}_i(1-\theta_i)(g(1,Z_i) - g(0,Z_i))\mathbb{E}[u_i|Z_i,\mathcal{U}_i]\big]\\ &\quad+\mathbb{E}\big[\theta_i(1-\hat{\theta}_i)(g(0,Z_i) - g(1,Z_i))\mathbb{E}[u_i|Z_i,\mathcal{U}_i]\big],

so it suffices that $\mathbb{E}[u_i|Z_i,\mathcal{U}_i]=0$ for Assumption (ref) to be satisfied.

Assumption (ref)(ref) formalizes the asymptotic framework described above. This condition requires the false positive and false negative rates to vanish at the same order as sampling uncertainty in the regression (i.e., $O(n^{-1/2})$). It provides a more robust starting point for empirical work than the standard approach of simply regressing $Y_i$ on $\hat X_i$ and performing OLS inference, which implicitly assumes $\kappa_+$ and $\kappa_-$ are zero.

Finally, Assumption (ref)(ref) limits the dependence between $g(0,Z_i)$ and $g(1,Z_i)$ and the classification errors in $\hat \theta_i$ (but not the dependence between $Z_i$ and $\theta_i$). This condition relates both to the functional form of $g$ and the statistical nature of the classification errors. This condition allows the conditional probabilities of misclassification $\mathbb{E}[\hat \theta_i(1-\theta_i)|Z_i]$ and $\mathbb{E}[\theta_i(1-\hat \theta_i)|Z_i]$ to depend on $Z_i$, provided the misclassification errors are asymptotically uncorrelated with $D_{+i}$ and $D_{-i}$. As this assumption is important, before proceeding we discuss some sufficient conditions for it.

continuance{ex:bchs} Consider ((ref)) as in \citetalias{battagliaInferenceRegressionVariables2025}. Then \[ \begin{aligned} D_{+i} = \begin{pmatrix} 1 & 0 \\ Z_i & 0 \end{pmatrix}, & & D_{-i} & = \begin{pmatrix} 0 & 0 \\ -Z_i & 0 \end{pmatrix} , \end{aligned} \] so a sufficient condition for Assumption (ref)(ref) is \begin{equation} \begin{aligned} \sqrt n \mathbb{E}[\theta_i - \hat \theta_i] & \to 0, \\ \sqrt n \mathbb{E}[(\theta_i - \hat \theta_i)Z_i] & \to 0. \end{aligned} \end{equation} Suppose that $\hat \theta_i$ is generated by a logistic classifier with (potentially many) variables $W_{i,n}$, including $Z_i$ and a constant. That is, \begin{equation} \mathbb{P}(\hat \theta_i = 1|W_{i,n}) = \frac{e^{W_{i,n}'\gamma_n}}{1+e^{W_{i,n}'\gamma_n}}. \end{equation} It is important to note that we are not assuming the logistic model is the DGP for $\theta_i$, just that it generates $\hat \theta_i$. Suppose the classifier has been pre-trained on a large external data set of $(\theta_i, W_{i,n})$. The parameters $\gamma_n$ of the classifier maximize the population likelihood \[ L(\gamma) = \mathbb{E}[ \theta_i W_{i,n}' \gamma - \log (1 + e^{W_{i,n}'\gamma})], \] and therefore satisfy the first-order condition \[ 0 = \mathbb{E}\left[ \left( \theta_i - \frac{e^{W_{i,n}'\gamma_n}}{1+e^{W_{i,n}'\gamma_n}} \right) W_{i,n} \right]. \] Hence, by ((ref)) and iterated expectations, we have \[ \mathbb{E}[(\theta_i - \hat \theta_i)W_{i,n}] = \mathbb{E}[(\theta_i - \mathbb{E}[\hat \theta_i|W_{i,n}])W_{i,n}] = 0, \] from which it follows that ((ref)) holds. Note that the first condition in ((ref)) together with $\sqrt n \mathbb{E}[\hat \theta_i(1-\theta_i)] \to \kappa_+$ implies that $\sqrt n \mathbb{E}[\theta_i(1-\hat \theta_i)] \to \kappa_+$ also holds, and thus $\kappa_+ = \kappa_-$.
lemmaSuppose that Assumption (ref) holds. Then Assumptions (ref) and (ref) hold with \[ B = \kappa_+ \mathbb{E}[D_{+i}] + \kappa_- \mathbb{E}[D_{-i}] . \]

Before proceeding, we characterize the bias term $B$ in the context of the two running examples.

continuance{ex1} Consider the interactions model ((ref)). Then \[ B = \begin{pmatrix} \kappa_+ \mathbb{E}[Z_{1i}Z_{1i}'] & 0 \\ (\kappa_+ - \kappa_-) \mathbb{E}[Z_{2i}Z_{1i}'] & 0 \end{pmatrix}. \]
continuance{ex:bchs} Consider the additive model ((ref)). Then \[ B = \begin{pmatrix} \kappa_+ & 0 \\ (\kappa_+ - \kappa_-) \mathbb{E}[Z_i] & 0 \end{pmatrix}. \] If $\kappa_+ = \kappa_-$, then \[ B = \begin{pmatrix} \kappa_+ & 0 \\ 0 & 0 \end{pmatrix}, \] as in \citetalias{battagliaInferenceRegressionVariables2025}.

Wild Bootstrap with Binary Labels

Both the fixed-label and coupled-label bootstraps are versions of the following wild bootstrap with binary labels, but differ in how $(\theta^*_i,\hat{\theta}^*_i)$ are drawn:

equation[equation omitted — 218 chars of source]

where $\eta_i$ are i.i.d. $(0,1)$ and generated independently of the bootstrap labels $(\theta^*_i,\hat{\theta}^*_i)$, conditional on the original sample. To understand when these methods are valid, we use the following set of conditions on $(\theta_i^*, \hat \theta_i^*)$.

\setcounter{assumption}{0}

assumption[] \begin{enumerate}[label={(\roman*)}] • $\frac{1}{n}\sum_{i=1}^{n}\mathbb{E}^*[\mathbb{I}[\theta^*_i\ne \hat{\theta}_i]]\to_p 0$. • $\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\mathbb{E}^*[\hat{\theta}^*_i(1-\theta^*_i)]\to_p \kappa_+$, $\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\mathbb{E}^*[\theta^*_i(1-\hat{\theta}^*_i)]\to_p \kappa_{-}$. • $\mathrm{Var}^*\big(\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\hat{\theta}^*_i(1-\theta^*_i)\big)\to_p 0$ and $\mathrm{Var}^*\big(\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\theta^*_i(1-\hat{\theta}^*_i)\big)\to_p 0$. • $\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\hat{\theta}^*_i(1-\theta^*_i)\big(D_{+i}-\mathbb{E}[D_{+i}]\big)+\theta^*_i(1-\hat{\theta}^*_i)\big(D_{-i}-\mathbb{E}[D_{-i}]\big)\to_{p^*} 0$. \end{enumerate}

Assumption (ref)(ref) requires the bootstrap “true” labels $\theta^*_i$ to match the estimated labels $\hat{\theta}_i$ on average with probability converging to one. This condition is automatically satisfied if $\theta^*_i=\hat{\theta}_i$. Assumption (ref)(ref) is the bootstrap analogue of Assumption (ref)(ref). It requires the average bootstrap false positive and negative rates, when scaled by $\sqrt n$, to converge in probability to $\kappa_{+}$ and $\kappa_{-}$. Together with Assumptions (ref)(ref) and (ref)(ref), this condition ensures that the bootstrap reproduces the asymptotic bias in the OLS estimator. Assumption (ref)(ref) can be viewed as the bootstrap analogue of Assumption (ref)(ref).

The following lemma shows that the high-level Assumptions (ref) and (ref) hold for the wild bootstrap with binary labels provided the method for generating $(\theta_i^*, \hat \theta_i^*)$ satisfies Assumption (ref).

lemmaSuppose Assumption (ref) holds. Let $u^*_i=\hat{u}_i\eta_i$ where $\eta_i$ are i.i.d. $(0,1)$ independently of $(\theta^*_i,\hat{\theta}^*_i)$ such that $\mathbb{E}^*[|\eta_i|^{2+\delta}]<\infty$ for $\delta>0$. If, in addition, $(\theta^*_i,\hat{\theta}^*_i)$ satisfy Assumption (ref), then Assumptions (ref) and (ref) hold.

Specific Bootstrap Methods

Here we study two methods for generating $(\theta^*_i,\hat{\theta}^*_i)$. The first is a standard fixed-regressor bootstrap which sets $\theta^*_i=\hat{\theta}_i$. Its validity requires $\mathbb{E}[D_{+i}|{\theta}_i=0]=\mathbb{E}[D_{+i}]$ and $\mathbb{E}[D_{-i}|{\theta}_i=1]=\mathbb{E}[D_{-i}]$, which holds if $D_{+i}$ and $D_{-i}$ are mean independent of ${\theta}_i$, but not necessarily otherwise. We therefore introduce a coupled-label bootstrap that jointly resamples $(\theta^*_i,\hat{\theta}^*_i)$ independently of the original sample to mimic the bias $B$ without this strong independence assumption.

Fixed-Label Bootstrap

This method sets $\theta^*_i=\hat{\theta}_i$ in the bootstrap DGP for $Y^*_i$, treating the bootstrap latent labels as fixed. The bootstrap estimated labels $\hat{\theta}^*_i$ are generated as i.i.d. Bernoulli random variables, conditional on $\hat{\theta}_i$:

flalign*\mathbb{P}^*(\hat{\theta}^*_i=\theta|\hat{\theta}_i=1)= \begin{cases} \hat{F}_{-}/\hat{\pi}& if \theta=0,\\ 1-\hat{F}_{-}/\hat{\pi}& if \theta=1, \end{cases}

and

flalign*\mathbb{P}^*(\hat{\theta}^*_i=\theta|\hat{\theta}_i=0)= \begin{cases} \hat{F}_{+}/(1-\hat{\pi})& if \theta=1,\\ 1-\hat{F}_{+}/(1-\hat{\pi})& if \theta=0. \end{cases}

Here, $\hat F_+$ and $\hat F_-$ are such that $\sqrt{n}\hat{F}_+\to_p \kappa_+$ and $\sqrt{n}\hat{F}_-\to_p \kappa_-$, respectively, and $\hat{\pi}=n^{-1}\sum_{i=1}^{n}\hat{\theta}_i$. To this end, we follow \citetalias{battagliaInferenceRegressionVariables2025} and rely on an external sample $(\theta_i, \hat \theta_i)_{i=1}^m$ used to assess the accuracy of the classifier. \citetalias{battagliaInferenceRegressionVariables2025} show that \[ \hat \kappa_+ = \sqrt n \hat F_+, \quad \text{with} \quad \hat F_+ = \frac 1m \sum_{i=1}^m \hat \theta_i(1-\theta_i) \] is consistent as $n, m \to \infty$ with $n/m^2 \to 0$ under Assumption (ref)(ref). An analogous result holds for \[ \hat \kappa_- = \sqrt n \hat F_-, \quad \text{with} \quad \hat F_- = \frac 1m \sum_{i=1}^m \theta_i(1-\hat \theta_i). \] Importantly, these results allow for the external sample size $m$ to be much smaller than the original sample size $n$ used in the regression.

For this method, Assumption (ref)(ref) is automatically satisfied since $\theta^*_i=\hat{\theta}_i$. One can also show that Assumptions (ref)(ref)-(ref) are satisfied provided $\sqrt{n}\hat{F}_+\to_p \kappa_+$ and $\sqrt{n}\hat{F}_-\to_p \kappa_-$. However, Assumption (ref)(ref) fails whenever $(D_{+i},D_{-i})$ are not mean independent of $\hat{\theta}_i$. This implies that the fixed-label bootstrap only replicates $B$ if we assume that this strong independence condition holds. To see this, note that the bootstrap bias is determined by (the probability limit of) the bootstrap expectation of \[ \hat{B}^*\equiv \frac{1}{\sqrt{n}}\sum_{i=1}^{n}(\hat{\theta}^*_i(1-\theta^*_i)D_{+i}+ \theta^*_i(1-\hat{\theta}^*_i)D_{-i}), ~\text{where~}\theta^*_i=\hat{\theta}_i. \] Fixing the “latent” bootstrap labels $\theta_i^*$ at $\hat{\theta}_i$ implies that the bootstrap expectation of $\hat{B}^*$ is equal to

flalign*\hat{B}\equiv \mathbb{E}^*[\hat{B}^*]&=\hat{\kappa}_+ \frac{1}{1-\hat{\pi}}\frac{1}{n}\sum_{i=1}^{n}(1-\hat{\theta}_i)D_{+i} + \hat{\kappa}_- \frac{1}{\hat{\pi}}\frac{1}{n}\sum_{i=1}^{n}\hat{\theta}_iD_{-i}\\ &\to_p \kappa_+ \mathbb{E}[D_{+i}|{\theta}_i=0]+\kappa_{-}\mathbb{E}[D_{-i}|{\theta}_i=1],

whereas \[ B = \kappa_+ \mathbb{E}[D_{+i}] + \kappa_- \mathbb{E}[D_{-i}] . \] Hence, $\hat B$ does not consistently estimate the bias $B$ in general, unless $(D_{+i},D_{-i})$ are mean independent of ${\theta}_i$. We summarize this result next.

corollaryUnder Assumption (ref), the fixed-label bootstrap satisfies the conditions of Theorem (ref) if $\sqrt{n}\hat{F}_ +\to_p \kappa_+$ and $\sqrt{n}\hat{F}_ -\to_p \kappa_-$, and $(D_{+i},D_{-i})$ are mean independent of ${\theta}_i$.
continuance{ex:bchs} One instance in which this bootstrap is valid is the special case of model ((ref)) where $Z_i$ contains just a constant for the intercept. Then, $D_{+i}$ and $D_{-i}$ are non-stochastic, and so $\hat B \to_p B$.

Coupled-Label Bootstrap

The coupled-label bootstrap generates the pair $(\theta_i^*, \hat{\theta}_i^*)$ jointly for each observation $i = 1, \ldots, n$, independently across $i$ and independently of $u_i^* = \hat{u}_i\eta_i$, according to the following categorical distributions:

equation*[equation* omitted — 411 chars of source]

and

equation*[equation* omitted — 413 chars of source]

The key feature of this method is that the probability of discordant pairs $(\theta_i^* , \hat{\theta}_i^*) = (1, 0)$ and $(0, 1)$ equals $\hat{F}_-$ and $\hat{F}_+$, respectively, for every observation $i$, regardless of $\hat{\theta}_i$. As a result, the bootstrap false positive and negative rates $\mathbb{E}^*[\hat{\theta}_i^*(1-\theta_i^*)] = \hat{F}_+$ and $\mathbb{E}^*[\theta_i^*(1-\hat{\theta}_i^*)] = \hat{F}_-$ are constant across $i$. This ensures that the bootstrap correctly replicates the asymptotic bias $B$ without requiring independence between $\theta_i$ and $Z_i$, since the bootstrap misclassification indicators are decoupled from $(D_{+i},D_{-i})$.

The conditional probabilities assigned to the concordant pairs $(\theta_i^*, \hat{\theta}_i^*) = (1,1)$ and $(0,0)$ ensure that the bootstrap Hessian $\hat{Q}^*$ converges to the empirical Hessian $\hat{Q}$ conditional on the data. Since $\hat{Q}^*$ depends only on $\hat{\theta}_i^*$ interacted with $g(1,Z_i)g(1,Z_i)'$ and $g(0,Z_i)g(0,Z_i)'$, its bootstrap expectation is determined by $\mathbb{E}^*[\hat{\theta}_i^*]$. As the discordant pair probabilities are set equal to $\hat{F}_+$ and $\hat{F}_-$, the key requirement is that $\mathbb{P}^*(\theta_i^*=1, \hat{\theta}_i^*=1|\hat{\theta}_i=1)\to_p 1$ and $\mathbb{P}^*(\theta_i^*=1, \hat{\theta}_i^*=1|\hat{\theta}_i=0)\to_p 0$, both of which are satisfied by the proposed construction. Under this distribution, $\mathbb{E}^*[\hat{\theta}_i^*]=\hat{\theta}_i(1-\hat{F}_-/\hat{\pi})+(1-\hat{\theta}_i)\hat{F}_+/(1-\hat{\pi})$, which ensures that the coupled-label bootstrap Hessian matches the fixed-label bootstrap Hessian in expectation. Other choices are nevertheless possible.

The following result formally establishes the validity of this bootstrap:

theoremSuppose that Assumption (ref) holds. Let $u^*_i=\hat{u}_i\eta_i$ where $\eta_i$ are i.i.d. $(0,1)$ independently of $(\theta^*_i,\hat{\theta}^*_i)$ such that $\mathbb{E}^*[|\eta_i|^{2+\delta}]<\infty$ for $\delta>0$. If, in addition, $\sqrt{n}\hat{F}_ +\to_p \kappa_+$ and $\sqrt{n}\hat{F}_ -\to_p \kappa_-$, then the coupled-label bootstrap satisfies the conditions of Theorem (ref) and, consequently, \begin{flalign*} \sqrt{n}\big(\hat{\beta}^*-\hat{\beta}\big)\rightarrow_{d^*} N(b,V), \end{flalign*} as $n \to \infty$, with $b = - Q^{-1} B \beta$ and $V = Q^{-1} \Sigma Q^{-1}$.

Theorem (ref) shows that the coupled-label bootstrap correctly replicates the bias and variance of the asymptotic distribution of the OLS estimator in Theorem (ref), thus justifying the use of percentile intervals for valid inference. However, as our simulations show, finite-sample coverage can deviate from the nominal level. We therefore recommend using the coupled-label bootstrap with the two finite-sample modifications described in the next section.

Finite-Sample Modifications

The asymptotic validity of the bootstrap percentile interval requires the bootstrap to correctly mimic the asymptotic distribution of $\sqrt{n}(\hat{\beta}-\beta)$, including its mean and variance. From the proofs of Lemma (ref) and Theorem (ref),

equation[equation omitted — 101 chars of source]

where $Z^* \to_{d^*} N(0,V)$ and

flalign*\hat{b}^* &= -\hat{Q}^{*-1}\bigg(\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\hat{\theta}^*_i(1-\theta^*_i)D_{+i} + \theta^*_i(1-\hat{\theta}^*_i)D_{-i}\bigg)\hat{\beta}\equiv -\hat{Q}^{*-1}\hat{B}^*\hat{\beta},

where $\hat{Q}^*=n{-1}\sum_{i=1}^{n}\hat{X}^*_i\hat{X}^{*\prime}_i$ is the bootstrap Hessian matrix. To show that $\hat{b}^*\to_{p^*}b$, we add and subtract appropriately to obtain

flalign\hat{b}^* &= -\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\hat{\theta}^*_i(1-\theta^*_i)\Gamma_{+}\beta + \theta^*_i(1-\hat{\theta}^*_i)\Gamma_{-}\beta + o_{p^*}(1)\equiv b^*+o_{p^*}(1),

where $\Gamma_{+}=Q^{-1}\mathbb{E}[D_{+i}]$ and $\Gamma_{-}=Q^{-1}\mathbb{E}[D_{-i}]$. While the coupled-label bootstrap bias $b^*$ is consistent for $b$, finite-sample distortions may arise. In particular, as noted in \citetalias{battagliaInferenceRegressionVariables2025}, estimating $F_+$ and $F_-$ introduces additional noise that can inflate the variance of $b^*$, requiring a variance correction to improve coverage of bias-corrected confidence intervals when $m/n$ is small. The same issue arises here, but a similar correction can easily be incorporated into the coupled-label bootstrap. A second source of distortion arises from the $o_{p^*}(1)$ term representing the difference between $\hat{b}^*$ and $b^*$. This term depends on the difference between $\hat{Q}^{*-1}$ and $\hat{Q}^{-1}$, which can be non-negligible in finite samples, especially when the proportion of $\hat{\theta}_i$ equal to 1 or 0 is small. In such cases, the bootstrap mean of $\hat{b}^*$ is different from that of $b^*$, introducing a location bias in the bootstrap. This motivates rotating $\hat{\beta}^*-\hat{\beta}$ by $\hat{R}^*=\hat{Q}^{-1}\hat{Q}^*$.

We discuss these two modifications next. We first describe the variance correction for the coupled-label bootstrap, establish its validity, and show it is able to replicate the two leading higher-order variance terms. We then describe the rotation that resolves the mismatch between the bootstrap Hessian and its sample analogue.

Coupled-Label Bootstrap with Variance Correction

To generate each bootstrap sample, first draw $V_+^* \sim \text{Binomial}(m, \hat F_+)$ and $V_-^* \sim \text{Binomial}(m, \hat F_-)$ and set $\hat F_+^* = V_+^*/m$ and $\hat F_-^* = V_-^*/m$. Conditional on $\hat F_+^*$ and $\hat F^*_-$, the pairs $(\theta_i^*, \hat{\theta}_i^*)$ are generated jointly for each observation $i = 1, \ldots, n$, independently across $i$ and independently of $u_i^* = \hat{u}_i\eta_i$, according to the following distributions:

equation[equation omitted — 468 chars of source]

and

equation[equation omitted — 471 chars of source]

Note that the $\hat F^*_+$ and $\hat F_-^*$ are generated only once per bootstrap sample and are held fixed when generating the $n$ pairs $(\theta_i^*, \hat{\theta}_i^*)$ for $i = 1,\ldots,n$.

The following result formally establishes validity of this bootstrap:

theoremSuppose that Assumption (ref) holds. Let $u^*_i=\hat{u}_i\eta_i$ where $\eta_i$ are i.i.d. $(0,1)$ independently of $(\theta^*_i,\hat{\theta}^*_i)$ such that $\mathbb{E}^*[|\eta_i|^{2+\delta}]<\infty$ for $\delta>0$. If, in addition, $\sqrt{n}\hat{F}_ +\to_p \kappa_+$ and $\sqrt{n}\hat{F}_ -\to_p \kappa_-$ and $n/m^2 \to 0$, then the coupled-label bootstrap with variance correction satisfies the conditions of Theorem (ref) and, consequently, \begin{flalign*} \sqrt{n}\big(\hat{\beta}^*-\hat{\beta}\big)\rightarrow_{d^*} N(b,V), \end{flalign*} as $n \to \infty$, with $b = - Q^{-1} B \beta$ and $V = Q^{-1} \Sigma Q^{-1}$.

Theorem (ref) shows that this bootstrap is able to reproduce the asymptotic bias and variance of the OLS estimator. We now show it is able to match key higher-order variance terms as well.

Variance Analysis

Finite-Sample Correction from \citetalias{battagliaInferenceRegressionVariables2025}

We first consider the bias-corrected estimator proposed in \citetalias{battagliaInferenceRegressionVariables2025}. That estimator can be written as \[ \hat \beta^{bc} = (I + \hat \Gamma_+ \hat F_+ + \hat \Gamma_- \hat F_-) \hat \beta, \] with $\hat Q = \frac 1n \sum_{i=1}^n \hat X_i \hat X_i'$, $\hat \Gamma_+ = \hat Q^{-1} \left( \frac 1n \sum_{i=1}^n D_{+i} \right)$, and $\hat \Gamma_-$ defined analogously. We may analyze the effect of $\hat F_+$ and $\hat F_-$ on the variance of $\hat \beta^{bc}$ as follows. First, as $\mathrm{Var}(\sqrt n(\hat \beta - \beta)) \approx V$, we have \[

aligned\mathbb{E}[\mathrm{Var}(\hat \beta^{bc} | \hat F_+, \hat F_-)] & \approx \frac 1n \mathbb{E}[ (I + \Gamma_+ \hat F_+ + \Gamma_- \hat F_-) V (I + \Gamma_+ \hat F_+ + \Gamma_- \hat F_-)'] \\[4pt] & = \frac 1n \left( V + O(n^{-1/2}) \right) ,

\] because $F_+ = \mathbb{E}[\hat F_+]$, $F_- = \mathbb{E}[\hat F_-]$ and $\sqrt n F_+, \sqrt n F_- = O(1)$. Moreover, \[

aligned\mathrm{Var}(\mathbb{E}[\hat \beta^{bc} |\hat F_+,\hat F_-]) & \approx \mathrm{Var}((I + \Gamma_+ \hat F_+ + \Gamma_- \hat F_-) \beta) \\[4pt] & = \frac{ F_+(1-F_+)}{m} \Gamma_+ \beta \beta' \Gamma_+' + \frac{ F_-(1-F_-)}{m} \Gamma_- \beta \beta' \Gamma_-' \\ & \quad - \frac{ F_+ F_- }{m} ( \Gamma_+ \beta \beta' \Gamma_-' + \Gamma_- \beta \beta' \Gamma_+'),

\] with $F_+ = \mathbb{E}[\hat F_+]$ and $F_- = \mathbb{E}[\hat F_-]$. The law of total variance implies that the variance of $\hat \beta^{bc}$ is approximately equal to the sum of these two terms. As $\sqrt n F_+, \sqrt n F_- = O(1)$ and $\sqrt{n}/m = o(1)$, we have

multline[multline omitted — 237 chars of source]

Under the framework adopted in \citetalias{battagliaInferenceRegressionVariables2025} where $m/n \to 0$ with $n/m^2 \to 0$, the second and third terms are $O(\sqrt n/m)$ and therefore dominate the remaining $O(n^{-1/2})$ term. Thus, it is these two terms that are most critical to match for the purposes of finite-sample variance correction.

The Bootstrap Matches the Finite-Sample Variance Correction

We now show that the coupled-label bootstrap with variance correction is able to match the above two $O(\sqrt n/m)$ terms. Recalling the bootstrap expansion ((ref)), the bootstrap expectation of the bias term $b^*$ is \[ \mathbb{E}^*[b^*] = - \left( \sqrt n \hat F_+ \Gamma_+ \beta + \sqrt n \hat F_- \Gamma_- \beta \right) \to_p b. \] Therefore, drawing $\hat F^*_+$ and $\hat F^*_-$ from a distribution centered at $\hat F_+$ and $\hat F_-$ does not affect the bootstrap's ability to reproduce the leading bias term.

To analyze the bootstrap variance of the bias term, first note that conditional on the data and $\hat F_+^*, \hat F_-^*$, the $\hat{\theta}^*_i(1-\theta^*_i)$ and $\theta^*_i(1-\hat{\theta}^*_i)$ are i.i.d. Bernoulli and satisfy \[

aligned\mathbb{E}^*[\hat{\theta}^*_i(1-\theta^*_i)|\hat F_+^*, \hat F_-^*] & = \hat F_+^* , & & & \mathbb{E}^*[\theta^*_i(1-\hat{\theta}^*_i)|\hat F_+^*, \hat F_-^*] & = \hat F_-^* , \\ \mathrm{Var}^*(\hat{\theta}^*_i(1-\theta^*_i)|\hat F_+^*, \hat F_-^*) & = \hat F_+^*(1-\hat F_+^*), & & & \mathrm{Var}^*(\theta^*_i(1-\hat{\theta}^*_i)|\hat F_+^*, \hat F_-^*) & = \hat F_-^*(1-\hat F_-^*),

\] and \[ \mathrm{Cov}^*(\hat{\theta}^*_i(1-\theta^*_i),\theta^*_i(1-\hat{\theta}^*_i)|\hat F_+^*, \hat F_-^*) = - \hat F_+^* \hat F_-^*. \] Using these expressions and the fact that $m\hat F^*_+$ and $m\hat F^*_-$ are independent Binomial$(m,\hat F_+)$ and Binomial$(m,\hat F_-)$ random variables conditional on the data, we have \[

aligned\mathrm{Var}^*\left( \mathbb{E}^*[b^*|\hat F_+^*,\hat F^*_-] \right) & = \mathrm{Var}^*\left( \sqrt n \hat F^*_+ \Gamma_+ \beta + \sqrt n \hat F^*_- \Gamma_- \beta \right) \\ & = \frac{n \hat F_+(1-\hat F_+)}{m}\Gamma_+ \beta \beta' \Gamma_+' + \frac{n \hat F_-(1-\hat F_-)}{m}\Gamma_- \beta \beta' \Gamma_-'.

\] Moreover,

multline*[multline* omitted — 276 chars of source]

from which we see that \[

aligned\mathbb{E}^* \left[ \mathrm{Var}^*(b^*|\hat F_+^*,\hat F^*_-) \right] & = \left( 1 - \frac{1}{m} \right) \left( \hat F_+(1-\hat F_+) \Gamma_+ \beta \beta' \Gamma_+' + \hat F_-(1-\hat F_-) \Gamma_- \beta \beta' \Gamma_-' \right) \\ & \quad - \hat F_+ \hat F_- ( \Gamma_+ \beta \beta' \Gamma_-' + \Gamma_- \beta \beta' \Gamma_+') \\ & = O_p(n^{-1/2}).

\] Hence, by the law of total variance, we have, \[ \mathrm{Var}^*(b^*) = \frac{n \hat F_+(1-\hat F_+)}{m}\Gamma_+ \beta \beta' \Gamma_+' + \frac{n \hat F_-(1-\hat F_-)}{m}\Gamma_- \beta \beta' \Gamma_-' + O_p(n^{-1/2}). \] And since $\sqrt{n}\hat{F}_ +\to_p \kappa_+$ and $\sqrt{n}\hat{F}_ -\to_p \kappa_-$, we have \[ \mathrm{Var}^*(b^*) = \frac{n F_+(1-F_+)}{m}\Gamma_+ \beta \beta' \Gamma_+' + \frac{n F_-(1-F_-)}{m}\Gamma_- \beta \beta' \Gamma_-' + o_p(\sqrt n/m) . \] These leading terms match $O(\sqrt n/m)$ terms in (ref), as required.

Rotation

As discussed previously, an additional source of finite-sample distortions associated with the coupled-label bootstrap (and its variance-corrected variant) is the discrepancy between the inverse bootstrap Hessian matrix $\hat{Q}^{*-1}$ and its sample analogue $\hat{Q}^{-1}$. If this difference is large, the remainder term in the bootstrap expansion ((ref)) is not negligible, introducing a bias in the bootstrap distribution.

One way to solve this problem is to bootstrap the distribution of $\sqrt{n}\hat{R}^*(\hat{\beta}^*-\hat{\beta})$, where $\hat{R}^*=\hat{Q}^{-1}\hat{Q}^*$. This is equivalent to bootstrapping the distribution of \[ \sqrt{n}(\tilde{\beta}^*-\hat{R}^*\hat{\beta}),\quad \tilde{\beta}^*=(\hat{X}^{\prime}\hat{X})^{-1}\hat{X}^{*\prime}Y^*, \]where $\hat{X}^*$ and $Y^*$ are obtained by the variance-corrected coupled-label bootstrap. With this rotation, the bias term of the bootstrap expansion in ((ref)) becomes

flalign*\hat{R}^*\hat{b}^* &= -\hat{Q}^{-1}\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\hat{\theta}^*_i(1-\theta^*_i)D_{+i}\hat{\beta} + \theta^*_i(1-\hat{\theta}^*_i)D_{-i}\hat{\beta},

thus no longer depending on $\hat{Q}^{*-1}$. Since $\hat{R}^*\to_{p^*} I_k$, the rotated coupled-label bootstrap with variance correction is asymptotically valid under the same conditions as those stated in Theorem (ref).

Simulations

We analyze the finite-sample properties of the bootstrap methods by extending the simulation design of \citetalias{battagliaInferenceRegressionVariables2025} in two directions. First, we consider an interactions model in which a standard normal covariate $Z_{i}$ is interacted with the binary label $\theta _{i}$ so that the effect of $Z_{i}$ on $Y_{i}$ differs between $\theta _{i}=1$ and $\theta _{i}=0$. Second, we allow $p_{i}=\mathbb{P}(\theta _{i}=1)$ to vary across observations to introduce correlation between $\theta _{i}$ and the matrices $D_{+i}$ and $D_{-i}$, which, as shown above, invalidates the fixed-label bootstrap. The model is:

align*[align* omitted — 165 chars of source]

where $F\left( \cdot \right) $ is the cumulative distribution function of a $ \chi ^{2}\left( 1\right) $ random variable, which implies that $ p_{i}\thicksim U[0,1].$ The joint distribution of the label pair $(\theta _{i},\hat{\theta}_{i})^{\prime }$ is given by

align*[align* omitted — 274 chars of source]

We set $\kappa _{+}=\kappa _{-}=\kappa $ and $F_{+}=F_{-}=\kappa /\sqrt{n}$. To ensure all probabilities remain non-negative, we map $p_{i}$ into

equation*[equation* omitted — 65 chars of source]

so that $\tilde{p}_{i}$ is uniformly distributed between $F_{+},$ and $2\bar{ p}-F_{+}$, where $\bar{p}$ is the average probability that $\theta _{i}=1$.

We consider two values of $\bar{p}\in \{0.05,\,0.5\}$, matching the constant probabilities used by \citetalias{battagliaInferenceRegressionVariables2025} . The correlation between $\theta _{i}$ and $Z_{i}^{2}$ depends strongly on $ \bar{p}$: it is approximately $0.02$ when $\bar{p}=0.05$ but rises to approximately $0.31$ when $\bar{p}=0.5$. As theory predicts, this has important consequences for the performance of the fixed-label bootstrap, which requires mean independence between $\theta _{i}$ and $(D_{+i}$, $ D_{-i})$.

The bootstrap DGP takes the form

flalign*Y_i^* &= \hat\beta'X^*_i + u_i^*, \qquad u_i^* = \hat{u}_i \eta_i, \quad \eta_i \sim i.i.d.\ N(0,1),\\ X^*_i&=(1,\theta_i^* Z_i, Z_i)^\prime,

where $\hat{\beta}$ is the sample OLS estimator and $\hat{u} _{i}$ are the OLS residuals. In all experiments, we use the wild bootstrap with standard normal weights to generate $u_{i}^{\ast }$. Four methods for generating $(\theta _{i}^{\ast },\hat{\theta}_{i}^{\ast })$ are considered. One is a naive bootstrap that sets $\theta _{i}^{\ast }=\hat{\theta} _{i}^{\ast }=\hat{\theta}_{i}$. This bootstrap method is denoted as \textquotedblleft no-label resampling\textquotedblright\ in the tables. By not introducing any measurement error in the bootstrap world, it implicitly sets the bootstrap bias to zero. This method is only included as a benchmark to demonstrate the need to resample labels in the bootstrap world. We then consider the fixed-label and coupled-label methods from Section (ref). Finally, we also include the variant of the coupled-label bootstrap with the finite-sample modifications discussed in Section (ref). We use $B=499$ bootstrap replications and $10{,}000$ Monte Carlo replications throughout. The sample sizes we consider are $n\in \{8{,}000;\,16{,}000;\,32{,}000\}$, and the misclassification rates are $ \kappa \in \{0.5,\,1,\,1.5\}.$ We vary the size of the external sample $m$ to keep the ratio $\sqrt{n}/m$ constant at 0.1265, corresponding to $m=707$ for $n=8{,}000,$ $m=1{,}000$ for $n=16{,}000$, and $m=1{,}414$ for $n=32{,}000.$

We focus on inference for the slope coefficient on the interaction term $ \theta_i Z_i$. For each configuration we report three statistics: the median bias, the empirical coverage rate of a nominal 95% confidence interval, and the median interval length. For bootstrap methods, we report equal-tailed percentile intervals, which dominate symmetric intervals in this setting.

Table (ref) presents results for $\bar{p}=0.5$. Each panel contains six estimators: OLS using the imputed regressor $\hat\theta_i Z_i$, the bias-corrected estimator of \citetalias{battagliaInferenceRegressionVariables2025} using their finite-sample variance adjustment, and the four bootstrap methods.

The use of an imputed label biases OLS downward, with the bias increasing in $\kappa $. The resulting undercoverage is severe: even in the mildest configuration ($n=8{,}000$, $\kappa =0.5$), OLS coverage is only $15.0\%$. The \citetalias{battagliaInferenceRegressionVariables2025} method removes virtually all the bias and delivers reliable inference across configurations, though coverage deteriorates modestly as $\kappa $ rises, and the intervals are substantially wider than those based on OLS.

The no-label-resampling bootstrap simply mimics the OLS estimator, inheriting its bias and producing similarly poor coverage (at best 15.5%).

The fixed-label bootstrap overcorrects: it introduces bias in the opposite direction. Coverage remains well below nominal and far below the \citetalias{battagliaInferenceRegressionVariables2025} benchmark.

The coupled-label bootstrap, by contrast, correctly replicates the bias of the bias-corrected estimator. Coverage remains below nominal level, however, because the basic coupled-label bootstrap does not account for the additional sampling uncertainty from estimating $F_{+}$ and $F_{-}$: interval lengths are close to those of OLS, and the improvement in coverage relative to no-label resampling comes entirely from correctly centering the bootstrap distribution. The coupled-label bootstrap with rotation and variance correction achieves coverage closest to the nominal $95\%$ among all methods, slightly edging the \citetalias{battagliaInferenceRegressionVariables2025} analytic correction, while maintaining similar interval lengths.

table[table omitted — 3,040 chars of source]
table[table omitted — 3,121 chars of source]

Table (ref) presents results for $\bar{p}=0.05$. This near-singular design is very challenging: the sample contains very few observations with $ \theta _{i}=1$, and the majority of these are misclassified when $\kappa \geq 1$. As a consequence, the OLS bias is roughly twice as large as in the $ \bar{p}=0.5$ case. Because the correlation between $\theta _{i}$ and $ (D_{+i},D_{-i})$ is low, the fixed-label bootstrap performs well and even slightly better than the coupled-label bootstrap. However, both maintain coverage below the \citetalias{battagliaInferenceRegressionVariables2025} benchmark. The variance correction and Hessian rotation improve coverage, and the gap with the analytic correction from \citetalias{battagliaInferenceRegressionVariables2025} decreases with sample size.

We conclude that the coupled-label bootstrap with finite-sample modifications is the preferred bootstrap method in all situations as it delivers coverage close to nominal coverage while producing confidence intervals of competitive length.

Application

As an illustration, we revisit \citetalias{battagliaInferenceRegressionVariables2025}' regressions investigating the wage premium for remote work. The data are from hansenRemoteWorkJobs2026, who construct a dataset that imputes remote work status for a corpus of online job postings provided by Lightcast.\footnote{See \url{https://wfhmap.com/}.} For each job posting $i$, they impute a binary variable $\hat \theta_i$ representing whether or not the job offers remote work. The variable is generated using DistilBERT, a LLM, applied to metadata on occupation, firm location, job title, posted wage, and a textual job description. As in \citetalias{battagliaInferenceRegressionVariables2025}, we regress log posted wage on $\hat \theta_i$, with and without occupation (SOC2) and full-time/part-time fixed effects (FEs). We focus on a sample of $n = 16{,}315$ job postings from 2022-2023 for the accommodation and food services sector in San Diego, CA. This is a near-singular design: very few jobs in this industry offer remote work ($\hat \pi = 0.024$), making bias correction more challenging. We take $\hat F_+ = 0.009$, as estimated by \citetalias{battagliaInferenceRegressionVariables2025} using $m = 1{,}000$ postings. Since there is no reason to expect the classifier to systematically over- or under-predict, we set $\hat F_- = 0.009$. As a robustness check, we present a second set of results with $\hat F_- = 0.018$.

table[table omitted — 1,192 chars of source]
table[table omitted — 1,072 chars of source]

The first set of results with $\hat F_+ = \hat F_- = 0.009$ is reported in Table (ref). We present OLS estimates and 95% confidence intervals for the coefficient of $\theta_i$, alongside results based on the analytical correction in \citetalias{battagliaInferenceRegressionVariables2025}, and bias-corrected estimates and confidence intervals using the naive no-label resampling bootstrap, the fixed-label bootstrap, and the coupled-label bootstrap with and without modifications. OLS estimates are around 30% smaller than those using the analytical correction. As expected, the naive no-label resampling method produces estimates and CIs that are near identical to those produced by OLS. In the baseline specification without fixed effects, both the fixed-label and coupled-label bootstrap are valid. The estimates and CIs they produce are very similar, as predicted. Our preferred coupled-label bootstrap with variance correction and rotation produces estimates and CIs that are close to those produced by the analytical correction. Similar findings are obtained when FEs are included.

As a robustness check, Table (ref) presents results with $\hat F_-$ doubled. The estimates and CIs are largely similar to those presented in Table (ref), except for the fixed-label and coupled-label bootstrap (without rotation). For those methods, estimates and CIs are shifted to the right, well above those obtained using analytical methods. This illustrates the importance of using the rotation in near-singular designs.

Conclusion

In this paper, we study whether the bootstrap can correct bias and deliver valid inference in regressions with binary labels that have been generated by AI or ML methods. We show that a seemingly natural fixed-label bootstrap, which generates data using the estimated labels while relying on a corrupted version in estimation, is generally invalid unless a strong independence condition between the latent true labels and other covariates holds. We propose a coupled-label bootstrap that jointly resamples the pair of true and imputed labels and show it is valid without such independence restrictions. We also introduce two further finite-sample modifications: a variance correction and a Hessian rotation for near-singular designs. Our recommendation for empirical practice is to use the coupled-label bootstrap with both of these modifications.