EconBase
← Back to paper

Inference after data-driven control-unit selection in difference-in-differences with estimated covariance

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

76,584 characters

Inference after data-driven control-unit selection in difference-in-differences with estimated covariance


\maketitle
\begin{abstract}
In difference-in-differences (DiD), researchers may use pre-treatment trends to select a control group for which the parallel-trends assumption appears plausible, with the aim of estimating the average treatment effect on the treated (ATT). Our earlier paper, \citet{NH2026} in \emph{Economics Letters}, and the present paper jointly provide the first selective-inference approach to the ATT that explicitly accounts for this control selection. We generalize our exact Gaussian procedure with known covariance to allow the covariance matrix to be estimated from the same individual-level data used for control selection and DiD estimation. We use this estimate to compute the variance, conditioning direction, residual, and truncation set. With fixed numbers of regions and periods, we establish uniform conditional coverage for selection events with probabilities bounded away from zero, and marginal coverage of the selected target without that restriction. We allow unequal regional sample sizes, heterogeneous covariances, ties in population fit, and regional sample shares that converge to zero. We establish asymptotic equivalence between the plug-in and known-covariance interval endpoints and derive rates for interval length. For staggered adoption, the control pools may differ across cohorts and periods, controls may be not yet treated, observations may be reused, and treatment effects may be heterogeneous. We also construct inference conditional on unions of selection paths that leave the reported parameter unchanged, together with simultaneous confidence bands for finitely many event-time effects. Under parallel trends and the other identifying conditions, the coverage results apply to the ATT. We give sufficient sampling conditions for individual panels and independent repeated cross-sections.
\end{abstract}
\noindent Keywords: difference-in-differences, control selection, post-selection inference, covariance estimation, within-region sample means.\\
\noindent JEL classification: C12, C21, C23.

\section{Introduction}
Difference-in-differences (DiD) estimates an average treatment effect on the treated (ATT) by comparing outcome changes in treated and untreated populations. Its central identifying restriction is parallel trends: in the absence of treatment, the relevant mean outcome change in the treated population would equal that in the control population. When several untreated groups are available, researchers may select from a candidate pool a control group for which parallel trends appears plausible, using similarities in observed pre-treatment trends. Using the same pre-treatment observations for control selection and DiD estimation can make control selection and the ATT estimator statistically dependent and alter the coverage of conventional confidence intervals. Inference must therefore account for how the control was selected rather than treat it as a comparison fixed before observing the data.

Our earlier paper, \citet{NH2026} in \emph{Economics Letters}, and the present paper jointly provide the first selective-inference approach to the ATT in DiD that explicitly accounts for selecting a control group from a candidate pool on the basis of pre-treatment trends to make parallel trends plausible. \citet{NH2026} construct exact inference conditional on control selection under a Gaussian model with known covariance. The present paper generalizes that method to feasible inference with covariance estimated from the same individual-level data and to control selection under staggered treatment adoption.

Under parallel trends, the selected population DiD contrast equals the ATT. Choosing a control by comparing candidates can nevertheless change the sampling distribution of the DiD estimator. We account for that selection event when constructing confidence intervals, while retaining parallel trends as an identifying assumption: observed pre-treatment fit alone does not establish post-treatment parallel trends.

\citet{Roth2022} studies the power of pre-trend tests and the bias and conventional confidence-interval coverage of estimators conditional on nonrejection for a given treated--control comparison. The paper does not construct ATT confidence intervals after choosing a control group from a candidate pool by pre-treatment fit. We instead choose the comparison group from a candidate pool and construct ATT inference conditional on that choice.

\citet{AKM2024} develop conditional, projection, and hybrid inference for targets chosen on the basis of estimated scores; a leading example is inference on the true effect of a program selected for its high estimated effect. The paper does not formulate and construct ATT intervals for DiD after selecting a control group by pre-treatment fit. We select a control group for estimating a given treated population's counterfactual outcome, not a treatment with a high estimated effect, and derive DiD inference conditional on that control-selection event.

\citet{RR2023} develop identification, robust confidence sets, and sensitivity analysis under restrictions on departures from parallel trends. Their analysis addresses violations of the identifying restriction; ours addresses selection from a control pool by pre-treatment fit. The two approaches are complementary. Using selected pre-trend estimates to calibrate a sensitivity analysis requires accounting for their selection as well.

The statistical foundations include inference under quadratic selection constraints \citep{Loftus2015}, selective asymptotics \citep{Tian2017}, and uniform inference with unknown variance \citep{TRTW2018}. \citet{MXT2018} use joint asymptotic normality of selection summaries and target statistics and study covariance estimation after selection. We retain comparisons of pre-treatment fit as quadratic functions of the regional means, without requiring the squared fit scores themselves to be asymptotically Gaussian. By normalizing the coefficients of the comparison inequalities, we use the same procedure for strict population rankings, ties at positive values of the fit criterion, ties at perfect fit, and sequences moving between these cases. Under the regional sampling model, we obtain a consistent estimator of the joint covariance matrix with unequal sample sizes and region-specific within-person serial dependence. We use the estimated covariance matrix to calculate the variance, conditioning direction, residual, and truncation set, allowing dependence between the sample mean and the covariance estimator.

We establish uniform asymptotic validity of the selective intervals and asymptotic equivalence of their endpoints to those computed with the reference covariance. With a fixed control pool, the single-comparison interval has stochastic order $(n_0^{-1}+n_s^{-1})^{1/2}$. The interval may therefore shrink even when the selection adjustment persists relative to the standard error. This stochastic rate does not imply a finite expected length, and a tie alone does not imply that adjustment is necessary.

For staggered adoption, we define cohort--period effects using never-treated or not-yet-treated controls and allow effects to differ across cohorts and event times, consistent with the distinction between identification and aggregation in \citet{CS2021}. We use explicit cohort--period ATT contrasts rather than interpret unrestricted two-way-fixed-effects lead and lag coefficients as those effects \citep{SA2021}. We express the relevant control choices as a joint event defined by quadratic inequalities and derive feasible selective inference, including simultaneous confidence bands for event-time effects. Shared controls and overlapping calendar periods are represented in a single joint outcome vector. When several selection paths lead to the same reported parameter, we may group them in advance and condition on their union. The resulting coverage guarantee conditions on the grouped event, rather than on every other cohort's control choice.

Section~\ref{sec:setup} presents the observation design, causal target, and control selection. Sections~\ref{sec:pivot} and \ref{sec:asymptotics} give the interval and its asymptotic guarantees. Section~\ref{sd:section} develops staggered adoption and joint inference, and Section~\ref{sec:simulation} reports the single-comparison simulations. Proofs of the additional results are in Appendix~\ref{sd:appendix}.

\section{Regional individual-level observations, control selection, and the target of inference}\label{sec:setup}
\subsection{Observation units and covariance estimation}
Let $j=0$ denote the treated region and $j=1,\ldots,J$ the candidate controls, and fix the number of regions $K=J+1$ and the number of periods $T$. From individual $i$ in region $j$ we observe the panel vector
\[
 X_{nji}=(X_{nji1},\ldots,X_{njiT})',\qquad
 \bar X_{nj}=n_j^{-1}\sum_{i=1}^{n_j}X_{nji}.
\]
Here $n$ indexes the experiment, and $n_j=n_j(n)$ is a deterministic sample size. We assume that the individual vectors are independent and identically distributed within each region and that the samples are independent across regions. We place no restriction on dependence over time within the same individual. We write
\begin{align}
 Y_n&=(\bar X_{n0}',\ldots,\bar X_{nJ}')',\quad
 \mu_n=(\mu_{n0}',\ldots,\mu_{nJ}')',
 \quad \mu_{nj}=\mathbb E X_{nji},\label{eq:regional-means}\\
 L_n&=\operatorname{blockdiag}_{j=0}^J(n_j^{-1/2}I_T),\quad
 V_n=\operatorname{blockdiag}_{j=0}^J(\Gamma_{nj}),\quad
 \Gamma_{nj}=\operatorname{Var}(X_{nji}).\label{eq:regional-scale}
\end{align}
The dimension $d=KT$ is fixed, and $\operatorname{Var}(Y_n)=\Sigma_n=L_nV_nL_n'$. The vector $Y_n$ includes the post-treatment means used for DiD as well as the pre-treatment means.

From the same individual-level observations we compute the unbiased sample covariance
\[
 \widehat\Gamma_{nj}=\frac1{n_j-1}\sum_i
 (X_{nji}-\bar X_{nj})(X_{nji}-\bar X_{nj})'
\]
and use
\begin{equation}\label{eq:regional-cov-est}
 \widehat V_n=\operatorname{blockdiag}_{j=0}^J(\widehat\Gamma_{nj}),
 \qquad \widehat\Sigma_n=L_n\widehat V_nL_n'
 =\operatorname{blockdiag}_{j=0}^J(\widehat\Gamma_{nj}/n_j).
\end{equation}
If positive definiteness has to be enforced in finite samples, we raise the eigenvalues of the standardized $\widehat V_n$ to at least $\varepsilon_n\downarrow0$ and then multiply by $L_n$. Adjusting the standardized covariance before rescaling preserves the different sampling-error scales across regions. The convention used when no such correction is applied is stated in Section~\ref{sec:asymptotics}.

\subsection{Selection by pre-treatment fit}
We fix, before the analysis, the set of pre-treatment periods $P$, the set of post-treatment periods $A$, and the candidate set $\{1,\ldots,J\}$, with $|P|\ge2$, $|A|\ge1$, and $P\cap A=\varnothing$. The candidates are assumed to be untreated through the last period in $A$. With $D_{njt}=\bar X_{n0,t}-\bar X_{nj,t}$ and $\bar D_{nj,P}=|P|^{-1}\sum_{t\in P}D_{njt}$, we define
\begin{equation}\label{eq:selection}
 Q_j(Y_n)=\sum_{t\in P}(D_{njt}-\bar D_{nj,P})^2
 =\|B_jY_n\|^2=Y_n'M_jY_n,
 \qquad S_n=\arg\min_{1\le j\le J}Q_j(Y_n).
\end{equation}
Here $B_j$ is a known centered comparison matrix and $M_j=B_j'B_j$. We fix a tie-breaking rule in advance and assume $M_j-M_s\ne0$ for distinct candidates. The selection criterion is unweighted by the estimated covariance matrix.

Apart from ties, the event that candidate $s$ is selected is
\begin{equation}\label{eq:joint}
 \widetilde E_{n,s}=\bigcap_{j\ne s}\{Y_n'A_{j,s}Y_n\ge0\},\qquad
 A_{j,s}=M_j-M_s.
\end{equation}
The actual selection event $E_{n,s}=\{S_n=s\}$ includes the fixed tie-breaking rule and coincides with the closed constraint set $\widetilde E_{n,s}$ outside the boundary. The candidate-specific DiD and its mean contrast are
\begin{equation}\label{eq:target}
 T_{n,s}=|A|^{-1}\sum_{t\in A}D_{nst}-|P|^{-1}\sum_{t\in P}D_{nst}
 =c_s'Y_n,\qquad \theta_{n,s}=c_s'\mu_n,
\end{equation}
where $c_s$ is a known nonzero vector. The parameter reported after selection is $\theta_{n,S_n}$ and may differ across selected controls.

\paragraph{Causal target and parallel trends.}
Let $\mu^0_{nj,t}$ denote the mean outcome in region $j$ at time $t$ under no treatment. For region 0, let $\mu^1_{n0,t}$ denote the mean under its actual treatment regime, which begins after the periods in $P$ and before the periods in $A$. Assume consistency, no spillovers across regions, no anticipation in $P$, and that every candidate is untreated through $\max A$. The average treatment effect on the treated (ATT) is
\begin{equation}\label{sd:single-att}
 \tau_n=\frac1{|A|}\sum_{t\in A}
          (\mu^1_{n0,t}-\mu^0_{n0,t}).
\end{equation}
The required parallel-trends restriction for a candidate $s$ is
\begin{equation}\label{sd:single-pt}
 \frac1{|A|}\sum_{t\in A}(\mu^0_{n0,t}-\mu^0_{ns,t})
 =\frac1{|P|}\sum_{u\in P}(\mu^0_{n0,u}-\mu^0_{ns,u}).
\end{equation}
This restriction equates the average untreated changes in the treated and control groups; the stronger restriction that the untreated mean gap is constant at every observed time is sufficient but not necessary. Define $\Delta_{n,s}$ as the left side minus the right side of \eqref{sd:single-pt}. Consistency and no anticipation give the identity
\begin{equation}\label{sd:single-decomposition}
 \theta_{n,s}=c_s'\mu_n=\tau_n+\Delta_{n,s}.
\end{equation}
Thus DiD identifies $\tau_n$ for this comparison when \eqref{sd:single-pt} holds.

We specify the control pool on substantive and design grounds before observing the fit scores, and then select a control using the pre-treatment criterion in equation~\eqref{eq:selection}. Subtracting the pre-treatment mean of each treated--control gap removes the level difference, so the criterion measures differences in their pre-treatment changes. This selection rule assesses pre-treatment fit, not the unobserved restriction \eqref{sd:single-pt} itself. If the restriction holds for every selectable candidate, all selected population contrasts identify the same $\tau_n$. More generally, the causal coverage bound in Corollary~\ref{sd:causal-coverage} allows a pool containing invalid candidates when the probability of selecting an invalid comparison tends to zero. That property requires a substantive restriction linking valid controls to the selection rule; it is not implied by observing good pre-treatment fit.

The selective interval corrects inference for the random choice of $s$. It does not subtract $\Delta_{n,s}$ or assert unbiasedness of the selected point estimate. Without an identifying restriction, its guarantee is for $\theta_{n,S_n}$.

\section{The truncated-normal pivot and computation of the interval}\label{sec:pivot}
In this section we suppress the subscript $n$. We first consider the case in which $Y\sim N_d(\mu,\Sigma)$ and $\Sigma\succ0$ is known. For candidate $s$, let
\begin{equation}\label{eq:decomp}
 T_s=c_s'Y,\quad v_s=c_s'\Sigma c_s,\quad
 b_s=\Sigma c_s/v_s,\quad R_s=Y-b_sT_s.
\end{equation}
Then $T_s$ and $R_s$ are independent. With $R_s=r$ fixed, $Y$ lies on the line $r+b_st$, which we call the conditional line, and the selection event is equivalent to $T_s$ belonging to
\begin{equation}\label{eq:fiber}
 \mathcal T_s(r;\Sigma)=\bigcap_{j\ne s}
 \{t:(r+b_st)'A_{j,s}(r+b_st)\ge0\}.
\end{equation}
This set is a finite union of intervals. For a variance $v>0$ and a truncation set $\mathcal T$, we define
\begin{equation}\label{eq:pivot}
 F_\theta(t;v,\mathcal T)=
 \frac{\int_{\mathcal T\cap(-\infty,t]}\exp\{-(x-\theta)^2/(2v)\}\,dx}
 {\int_{\mathcal T}\exp\{-(x-\theta)^2/(2v)\}\,dx}.
\end{equation}

\begin{lemma}[Gaussian reference model with known covariance]\label{lem:gaussian}
For each $s$ with $P(E_s)>0$, the distribution of $T_s\mid(E_s,R_s)$ is $N(\theta_s,v_s)$ restricted to $\mathcal T_s(R_s;\Sigma)$. Hence
\begin{equation}\label{eq:selci}
 C^{\mathrm{sel}}_\alpha(Y;\Sigma)=
 \{\theta:\alpha/2\le F_\theta(T_s;v_s,\mathcal T_s(R_s;\Sigma))\le1-\alpha/2\}
\end{equation}
has coverage exactly $1-\alpha$ conditional on $E_s,R_s$ and satisfies
\begin{equation}\label{eq:coverage}
 P\{\theta_s\inC^{\mathrm{sel}}_\alpha(Y;\Sigma)\mid E_s\}=1-\alpha.
\end{equation}
On the selection event, the observed point lies in the interior of the truncation set almost surely, and the two endpoints of the interval are finite and unique.
\end{lemma}

With unknown covariance, we use the same interval construction with the estimated covariance matrix. We replace $\Sigma$ by $\widehat\Sigma_n$ and recompute all of $v_s,b_s,R_s,\mathcal T_s$. We call the resulting interval the plug-in selective interval, or the plug-in interval for short. The lower and upper endpoints of the interval solve $F_\theta(T_s)=1-\alpha/2$ and $F_\theta(T_s)=\alpha/2$, respectively. The conventional interval is the normal interval $T_s\pm z_{1-\alpha/2}\sqrt{c_s'\widehat\Sigma_nc_s}$, where $z_q$ is the $q$-quantile of the standard normal distribution; \citet{NH2026} call its known-covariance counterpart the naive interval.

To obtain the truncation set, we partition the real line at the real roots of each constraint of degree at most two and check all constraints on each interval. We integrate over every admissible interval, including components that do not contain the observed point. A procedure that computes $b_s$, $R_s$, $\mathcal T_s$ from a fixed working covariance $W_0$ other than $\widehat\Sigma_n$ and replaces only the variance $v_s$ by its estimate is not covered by the guarantee of the next section. If $W_0c_s/(c_s'W_0c_s)$ differs from $\Sigma c_s/(c_s'\Sigma c_s)$, then $T_s$ and the residual $Y-\{W_0c_s/(c_s'W_0c_s)\}T_s$ are correlated, and the conditional distribution of $T_s$ given the residual is in general not a truncated $N(\theta_s,v_s)$ distribution.

\section{Asymptotic inference with estimated covariance}\label{sec:asymptotics}
We state the asymptotic results under a joint normal approximation and consistent covariance estimation, allowing their application to estimators other than within-region sample means. We fix $d,J,B_j,c_s$ and write the finite set of candidates as $\mathcal S=\{1,\ldots,J\}$. Let $\mathcal P_n$ be a class of data-generating distributions; $\mu_n=\mu_n(P)$ is unrestricted. In this section $\mu_n$ denotes the deterministic center of the normal approximation; for a general asymptotically linear estimator it can be taken to be the estimand and need not equal the finite-sample expectation. For the sample means in \eqref{eq:regional-means}, $\mu_n=\mathbb E Y_n$. The distance $d_{\mathrm{BL}}$ is the supremum of the difference in expectations over real functions that are bounded by 1 in absolute value and have Lipschitz constant at most 1.

To accommodate known linear transformations and jointly estimated statistics, we state the theorem for the following class of selection events. For each candidate in a fixed finite set $\mathcal S$, suppose that the selection event, apart from ties, is
\[
 \widetilde E_{n,s}=\bigcap_{k=1}^{m_s}\{p_{n,s,k}(Y_n)\ge0\}.
\]
Here $m_s$ is fixed, and $p_{n,s,k}$ is a known deterministic polynomial of degree at most two that is not identically zero; its coefficients may vary with $n$. In this general form, $\mathcal S$ denotes a fixed finite set of selection labels, whose elements need not be indices of single regions. The relation between the actual selection event and the closed constraint set is specified below. The contrast coefficient $c_{n,s}\ne0$ is also deterministic. The original selection rule is the case in which comparison $k$ corresponds to a candidate $j\ne s$, $p_{n,s,k}(y)=y'A_{j,s}y$, and $c_{n,s}=c_s$.

\begin{assumption}[Standardized normal approximation and covariance estimation]\label{ass:array}
Let $L_n$ be a known nonsingular deterministic matrix and $\widehat V_n$ a symmetric-matrix-valued estimator. Let $Z_n=L_n^{-1}(Y_n-\mu_n)$, and suppose that for some $0<\underline v<\overline v<\infty$,
\[
 \underline v I_d\preceq V_n(P)\preceq\overline v I_d
 \qquad(P\in\mathcal P_n).
\]
We also assume
\begin{align}
 &\sup_{P\in\mathcal P_n}
 d_{\mathrm{BL}}\{\mathcal L_P(Z_n),N_d(0,V_n(P))\}
 \longrightarrow0,\label{eq:uniform-clt}\\
 &\sup_{P\in\mathcal P_n}
 P\{\|\widehat V_n-V_n(P)\|>\epsilon\}
 \longrightarrow0\quad(\epsilon>0).\label{eq:uniform-cov}
\end{align}
The norm is the operator norm. Independence of $Y_n$ and $\widehat V_n$ is not assumed.
\end{assumption}

Let $\widehat\Sigma_n=L_n\widehat V_nL_n'$; we compute the variance, the conditioning direction, the residual, and the truncation set with this matrix.
We define the closed constraint set of each candidate and the boundary set by
\[
 \widetilde E_{n,s}
  =\bigcap_{k=1}^{m_s}\{p_{n,s,k}(Y_n)\ge0\},\qquad
 \mathcal B_n
  =\bigcup_{s\in\mathcal S}\bigcup_{k=1}^{m_s}
       \{p_{n,s,k}(Y_n)=0\}.
\]
The actual selection events $E_{n,s}=\{S_n=s\}$, $s\in\mathcal S$, are assumed to form a partition of the sample space through a fixed
tie-breaking rule and to satisfy
$E_{n,s}\mathbin{\triangle}\widetilde E_{n,s}\subseteq\mathcal B_n$ for every $s$.
An empty intersection is the whole space, and an empty union is the empty set.

From here on, $\widehat V_n$ denotes the matrix whose eigenvalues have been adjusted on the standardized scale where necessary.
If $\widehat V_n^{\mathrm{raw}}$ is the original symmetric estimator,
its eigenvalues may be raised to at least $\varepsilon_n\downarrow0$.
By the eigenvalue lower bound and the consistency in Assumption~\ref{ass:array}, the probability that this operation
changes the estimator tends to zero uniformly, and
\eqref{eq:uniform-cov} continues to hold after the adjustment.

On $\mathcal B_n$, or when the covariance estimator in use is not positive definite,
we define the pivot to be $1/2$ and the confidence interval to be $\mathbb R$ for every hypothesized parameter value.
The interval endpoints used below are defined on the event that excludes these exceptions.
Under Assumption~\ref{ass:array}, the probability of these exceptions tends to zero uniformly:
for $\mathcal B_n$ by Lemma~\ref{rev:boundary} in the appendix, and for positive definiteness
by the eigenvalue lower bound $\lambda_{\min}\{V_n(P)\}\ge\underline v$ and \eqref{eq:uniform-cov}.
Even when positive definiteness is enforced, the selection score itself is not changed.


\begin{theorem}[Plug-in selective inference for general means and unbalanced scales]\label{thm:plugin}
Under Assumption~\ref{ass:array}, fix $0<\alpha<1$. Let $S_n$ be the realized selection label, $E_{n,s}=\{S_n=s\}$, and $\theta_{n,s}=c_{n,s}'\mu_n$. For the original selection rule, $c_{n,s}=c_s$. Let $U_{n,s}$ denote the pivot obtained by substituting $\widehat\Sigma_n$ into \eqref{eq:pivot}, with the corresponding polynomial constraints and $c_{n,s}$ in the general form, and evaluated at the true target $\theta_{n,s}$. Outside $E_{n,s}$, $U_{n,s}$ may be defined arbitrarily. For the joint probability with the selection event,
\begin{equation}\label{eq:pluginjoint}
 D_n:=\sup_{P\in\mathcal P_n}\max_{s\in\mathcal S}
 \sup_{u\in[0,1]}
 \left|P(U_{n,s}\le u,E_{n,s})-uP(E_{n,s})\right|
 \longrightarrow0.
\end{equation}
For any fixed $p_0>0$,
\begin{equation}\label{eq:pluginpivot}
 \sup_{\substack{P\in\mathcal P_n:\ P(E_{n,s})\ge p_0}}
 \sup_{u\in[0,1]}
 \left|P(U_{n,s}\le u\mid E_{n,s})-u\right|
 \longrightarrow0.
\end{equation}
Consequently,
\begin{equation}\label{eq:plugincoverage}
 \sup_{\substack{P\in\mathcal P_n:\ P(E_{n,s})\ge p_0}}
 \left|P\{\theta_{n,s}\inC^{\mathrm{sel}}_\alpha(Y_n;\widehat\Sigma_n)
          \mid E_{n,s}\}-(1-\alpha)\right|
 \longrightarrow0.
\end{equation}
Moreover, without assuming a lower bound on the selection probability of each candidate, the unconditional coverage for the target $\theta_{n,S_n}$ of the realized candidate satisfies
\begin{equation}\label{eq:pluginmarginal}
 \sup_{P\in\mathcal P_n}
 \left|P\{\theta_{n,S_n}\in
       C^{\mathrm{sel}}_\alpha(Y_n;\widehat\Sigma_n)\}-(1-\alpha)\right|
 \longrightarrow0.
\end{equation}
The conditions $\mu_n=m+h/\sqrt n$, $B_jm=0$, a lower bound on the population score differences, and positive limits for the regional sample shares are not required. The conclusion is unchanged if the supremum over finitely many candidates is taken in \eqref{eq:pluginpivot}--\eqref{eq:plugincoverage}.
\end{theorem}

Theorem~\ref{thm:plugin} uses the same interval construction for a strict population ranking, a tie at a positive value of the fit criterion, a tie at perfect fit, and mean sequences moving among these cases. The procedure does not require a preliminary test to select among limiting approximations. Uniformity is over the distribution class satisfying \eqref{eq:uniform-clt}--\eqref{eq:uniform-cov}; a pointwise central limit theorem alone is insufficient for this conclusion. Because the error in the candidate-specific conditional distribution function of the pivot is at most $D_n/P(E_{n,s})$, \eqref{eq:pluginpivot}--\eqref{eq:plugincoverage} require a uniform lower bound $p_0$ on the selection probability. The guarantee \eqref{eq:pluginmarginal} sums the joint probabilities over finitely many candidates and does not extend to the conditional coverage of each candidate whose probability approaches zero. Nor is it a finite-sample guarantee conditional on each realization of the estimated residual.

If $P(E_{n,s})>0$, the error in the conditional distribution function of the pivot and
the error in the conditional coverage of the two-sided interval satisfy
\begin{align*}
 \sup_{u\in[0,1]}
 \left|P(U_{n,s}\le u\mid E_{n,s})-u\right|
 &\le \frac{D_n}{P(E_{n,s})},\\
 \left|P\{\theta_{n,s}\inC^{\mathrm{sel}}_\alpha(Y_n;\widehat\Sigma_n)
           \mid E_{n,s}\}-(1-\alpha)\right|
 &\le \frac{2D_n}{P(E_{n,s})}.
\end{align*}
The error in the unconditional coverage of the selected target, in contrast, is at most $2|\mathcal S|D_n$.
Thus the candidate-specific conditional guarantee uses a lower bound on the selection probability,
whereas the unconditional guarantee over a fixed finite number of candidates does not require this bound.
Because the rate at which $D_n\to0$ is not specified here, these expressions
do not give a numerical error bound for an available sample size.

In the proof we abbreviate the comparison matrices of candidate $s$ as $A_k$ and divide the coefficients of
\begin{align}
 q_{kn}(z)&=(\mu_n+L_nz)'A_k(\mu_n+L_nz)\notag\\
 &=z'H_{kn}z+\ell_{kn}'z+a_{kn},\label{eq:normalized-comparison}\\
 H_{kn}&=L_n'A_kL_n,\quad
 \ell_{kn}=2L_n'A_k\mu_n,\quad a_{kn}=\mu_n'A_k\mu_n\notag
\end{align}
by
\[
 \rho_{kn}=\{\|H_{kn}\|_F^2+\|\ell_{kn}\|^2+a_{kn}^2\}^{1/2}>0.
\]
In the general form we set $q_{kn}(z)=p_{n,s,k}(\mu_n+L_nz)$ and normalize the coefficients of its symmetric quadratic term, linear term, and constant term in the same way. A nonsingular affine transformation maps a nonzero polynomial to a nonzero polynomial, so $\rho_{kn}>0$ in this case as well. Division by a positive constant does not change the selection event. The normalized coefficients lie in a compact set, and any subsequence has a further subsequence whose limit is a nonzero polynomial. This operation also handles cases in which the relative magnitudes of the constant, linear, and quadratic terms change. The quantities $\mu_n$ and $\rho_{kn}$ are used only in the proof and are not estimated in the computation. Appendix~\ref{app:proofs} gives the complete proof.

\begin{theorem}[Plug-in equivalence of interval endpoints]\label{rev:endpoints}
Assume the conditions of Theorem~\ref{thm:plugin} and fix $0<\alpha<1$.
Let $\Sigma_n=L_nV_nL_n'$ be the Gaussian reference covariance, and let
\[
 r_{n,s}=\|L_n'c_{n,s}\|,\qquad
 \sigma_{n,s}=(c_{n,s}'\Sigma_nc_{n,s})^{1/2},\qquad
 \widehat\sigma_{n,s}=(c_{n,s}'\widehat\Sigma_nc_{n,s})^{1/2}.
\]
For the $s$ that is actually selected, write the two intervals based on the same $Y_n$ as
\[
 C^\circ_{n,s}=[\ell^\circ_{n,s},u^\circ_{n,s}],\qquad
 \widehat C_{n,s}=[\widehat\ell_{n,s},\widehat u_{n,s}].
\]
The former uses $\Sigma_n$ and the latter $\widehat\Sigma_n$;
in each case the variance, the conditional line, the residual, and the truncation set are computed with the respective covariance.
On the event without exceptions, let
\[
 A_{n,s}=\frac{
  \max\{|\widehat\ell_{n,s}-\ell^\circ_{n,s}|,
        |\widehat u_{n,s}-u^\circ_{n,s}|\}}{\sigma_{n,s}}
\]
and define $A_{n,s}=\infty$ on the exceptions.
For any $\epsilon>0$,
\begin{equation}\label{rev:eq:endpoint-equivalence}
 \sup_{P\in\mathcal P_n}\max_s
 P\{E_{n,s},\ A_{n,s}>\epsilon\}\longrightarrow0.
\end{equation}
Further, let
\[
 R_{n,s}=
 \frac{\max\{|\widehat\ell_{n,s}-\theta_{n,s}|,
             |\widehat u_{n,s}-\theta_{n,s}|\}}{\sigma_{n,s}}
\]
with $R_{n,s}=\infty$ on the exceptions. Then
\begin{align}
 &\lim_{M\to\infty}\limsup_{n\to\infty}
   \sup_{P\in\mathcal P_n}\max_s
       P\{E_{n,s},R_{n,s}>M\}=0,\label{rev:eq:endpoint-tight}\\
 &\lim_{\eta\downarrow0}\limsup_{n\to\infty}
   \sup_{P\in\mathcal P_n}\max_s
       P\{E_{n,s},|\widehat C_{n,s}|/\sigma_{n,s}<\eta\}=0.
                                                        \label{rev:eq:width-lower}
\end{align}
These also hold for $C^\circ_{n,s}$. In addition, for the selected label,
\begin{equation}\label{rev:eq:length-ratio}
 \frac{|\widehat C_{n,S_n}|}{|C^\circ_{n,S_n}|}\to_p1
\end{equation}
uniformly over the class of distributions. On exceptions where the ratio is undefined, any value
may be substituted.

Each statement also holds conditionally on a candidate with $P(E_{n,s})\ge p_0>0$.
Moreover, by summing over the fixed finite number of selection labels,
the unconditional conclusions evaluated at $S_n$ do not require a lower bound on the selection probability.
The conclusions are unchanged if the denominator $\sigma_{n,s}$ is replaced by $\widehat\sigma_{n,s}$.
\end{theorem}

The proof uses subsequential limits: even if the mean sequence or comparison coefficients do not converge,
every subsequence has a further subsequence along which the standardized interval converges to an interval obtained from a truncated normal distribution
with finite, positive length. The limiting distribution need not be the same for every mean sequence.
The reference covariance $\Sigma_n=L_nV_nL_n'$ is determined by the $V_n$ used in the normal approximation of Assumption~\ref{ass:array};
for a general asymptotically linear estimator it need not equal $\operatorname{Var}(Y_n)$.
Also, \eqref{rev:eq:endpoint-equivalence} does not give a rate of convergence.
In the setting of Corollary~\ref{cor:regional}, if the individual-level observations do not have a finite fourth moment,
the convergence of $\widehat V_n$ can be slower than $n_j^{-1/2}$, so the convergence of the endpoint difference can also be slow.



\begin{corollary}[Heterogeneous regional covariances and sample sizes]\label{cor:regional}
In the setting of \eqref{eq:regional-means}--\eqref{eq:regional-cov-est}, suppose that for some $\delta>0$ and finite $M$,
\[
 \sup_{n,j}\mathbb E\|X_{nji}-\mu_{nj}\|^{2+\delta}\le M,
 \qquad
 \underline v I_T\preceq\Gamma_{nj}\preceq\overline v I_T,
 \qquad \min_j n_j\longrightarrow\infty.
\]
Then Assumption~\ref{ass:array} holds for the class of all distributions that satisfy these constraints with the same constants. Hence Theorem~\ref{thm:plugin} applies. There may be regions for which $n_j/\sum_l n_l$ converges to zero, and the covariance may be estimated from the same individual-level observations as those used for the selection and DiD.
\end{corollary}

For example, if $n_0=r$, $n_1=r^2$, $n_2=2r$ and $r\to\infty$, the shares of regions 0 and 2 in the total sample tend to zero, but the conditions of Corollary~\ref{cor:regional} are satisfied. The corollary does not cover cases in which the minimum sample size does not increase and the normal approximation for the regional means needed for inference is unavailable. In the more restrictive case with $n_{\mathrm{tot}}=\sum_jn_j$, $n_j/n_{\mathrm{tot}}\to\lambda_j>0$, and $\Gamma_{nj}\to\Gamma_j$, we obtain the usual representation on a common scale,
\[
 \sqrt{n_{\mathrm{tot}}}(Y_n-\mu_n)\mathrel{\Rightarrow} N(0,\Omega),\qquad
 n_{\mathrm{tot}}\widehat\Sigma_n\mathrel{\to_{\!p}}\Omega,
 \quad
 \Omega=\operatorname{blockdiag}_j(\Gamma_j/\lambda_j).
\]

\begin{corollary}[Regional effective sample size and shrinkage of the interval]\label{rev:regional-rate}
Assume the conditions of Corollary~\ref{cor:regional}. Let $w\in\mathbb R^T$ be the vector equal to
$|A|^{-1}$ in the post-treatment periods, $-|P|^{-1}$ in the pre-treatment periods, and 0 in all other periods, and define
\[
 n^{\mathrm{eff}}_{n,s}
   =\left(n_0^{-1}+n_s^{-1}\right)^{-1}
   =\frac{n_0n_s}{n_0+n_s}.
\]
Then
\[
 \sigma_{n,s}^2=
 \frac{w'\Gamma_{n0}w}{n_0}+
 \frac{w'\Gamma_{ns}w}{n_s},\qquad
 \|w\|^2=|A|^{-1}+|P|^{-1}.
\]
For the interval of the selected target, uniformly over the class of distributions,
\begin{align}
 |\widehat C_{n,S_n}|
   &=\Theta_p\left((n^{\mathrm{eff}}_{n,S_n})^{-1/2}\right),
                                      \label{rev:eq:regional-width}\\
 \max\{ |\widehat\ell_{n,S_n}-\theta_{n,S_n}|,
         |\widehat u_{n,S_n}-\theta_{n,S_n}|\}
   &=O_p\left((n^{\mathrm{eff}}_{n,S_n})^{-1/2}\right).
                                      \label{rev:eq:regional-ends}
\end{align}
Expression \eqref{rev:eq:regional-width} means that the interval length divided by the scale on the right-hand side
is uniformly tight and that its reciprocal is also uniformly tight.
The probability that the interval equals $\mathbb R$ on the exceptions tends to zero uniformly.

Hence, under $\min_jn_j\to\infty$, for any $\epsilon>0$,
\begin{equation}\label{rev:eq:shrinking}
 \sup_{P\in\mathcal P_n}
 P\left\{\max\{
   |\widehat\ell_{n,S_n}-\theta_{n,S_n}|,
   |\widehat u_{n,S_n}-\theta_{n,S_n}|
                   \}>\epsilon\right\}\longrightarrow0.
\end{equation}
If $P(E_{n,s})\ge p_0>0$ for candidate $s$, the corresponding candidate-specific conditional
conclusions also hold. If $\Delta_{n,s}=0$ for all candidates,
$\theta_{n,S_n}$ may be replaced by $\tau_n$ in these expressions.
\end{corollary}

With a fixed number of regions, the interval can shrink even when the pre-treatment fit is tied.
Any persistent effect of selection adjustment appears in the distribution of the interval after normalization by the standard error.
For example, if $n_0=r$ and $n_s=r^2$, then $n^{\mathrm{eff}}_{n,s}\sim r$, and
increasing the sample of the control region alone leaves the stochastic order of the interval at $r^{-1/2}$.
Expression \eqref{rev:eq:shrinking} states that the estimation error of the endpoints tends to zero;
it does not state that the coverage probability tends to one.



\begin{corollary}[When the selection becomes asymptotically deterministic]\label{cor:separated}
Consider a single sequence of distributions that satisfies the conditions of Theorem~\ref{thm:plugin}.
If $P(E_{n,s})\to1$ for some fixed label $s$, then
\begin{align*}
 \frac{\widehat\ell_{n,s}
        -(c_{n,s}'Y_n-z_{1-\alpha/2}\widehat\sigma_{n,s})}
      {\widehat\sigma_{n,s}}&\to_p0,\\
 \frac{\widehat u_{n,s}
        -(c_{n,s}'Y_n+z_{1-\alpha/2}\widehat\sigma_{n,s})}
      {\widehat\sigma_{n,s}}&\to_p0.
\end{align*}
The endpoints may be defined arbitrarily when $s$ is not selected.
For the original quadratic-score selection, this condition holds if $\mu_n\to\mu$, $\|L_n\|\to0$, and
$\mu'A_{j,s}\mu>0$ for all comparisons.
\end{corollary}

Corollary~\ref{cor:separated} establishes convergence without specifying its rate. When population fit differences
diverge slowly relative to the standard error, convergence to the conventional interval can remain slow even if the selection probability is close to one.


\begin{proposition}[When the selection does not change along the conditional line]\label{rev:no-adjustment}
In the Gaussian model with known covariance, write each constraint for candidate $s$ as
$p_k(y)=y'A_ky+d_k'y+e_k\ge0$, $A_k=A_k'$.
If, for $b_s=\Sigma c_s/(c_s'\Sigma c_s)$,
\begin{equation}
 A_kb_s=0,\qquad d_k'b_s=0\qquad\text{(for all comparisons $k$)},
                                                        \label{rev:eq:invariance}
\end{equation}
then the truncation set on $E_s$ is $\mathbb R$, and the equal-tailed selective interval coincides exactly with the conventional
interval based on $N(c_s'\mu,c_s'\Sigma c_s)$.
For the original homogeneous quadratic scores, a sufficient condition is $A_{j,s}\Sigma c_s=0$ ($j\ne s$).
When there are two or more candidates, this is equivalent to
$B_j\Sigma c_s=0$ for all candidates $j$, and the latter is easier to check.

In the asymptotic setting of Theorem~\ref{rev:endpoints}, if
\eqref{rev:eq:invariance} holds for the reference covariance $\Sigma_n$ for each $n$,
then the difference between the endpoints of the plug-in selective interval and those of the conventional interval for the same candidate,
divided by $\widehat\sigma_{n,s}$, converges to zero in probability jointly with the selection event, in the sense of \eqref{rev:eq:endpoint-equivalence}.
For a candidate with $P(E_{n,s})\ge p_0>0$, the same conclusion holds conditionally.
\end{proposition}



\begin{remark}[An example in which no adjustment is needed even under an exact tie]\label{rev:tie-example}
Place the same number of independent individual-level observations in each region $j=0,1,2$, and let
$V_{ji}$ and $H_{ji}$ be independent standard normal variables.
With $S=\arg\min_{j=1,2}|\bar V_0-\bar V_j|$ and
$T_s=\bar H_0-\bar H_s$, the population fit of both candidates is zero and tied, and
by exchangeability $P(S=1)=P(S=2)=1/2$.
However, the selection is a function of $V$ only and the estimator is a function of $H$ only, so for each candidate $s$,
$T_s$ is independent of $E_s$. With known covariance the conventional interval is already exact, and
also when the covariance estimate from the same individual-level observations is used, the plug-in selective interval approaches the conventional interval
asymptotically. This shows that the need for an adjustment cannot be concluded from a tie alone.
With an independent $B$ added, this example can be written as a positive-definite three-period panel
$(B-V/2,B+V/2,B+H)$.
In this panel $P=\{1,2\}$, $A=\{3\}$, and $w=(-1/2,-1/2,1)'$;
if $\operatorname{Var}(B)>0$, the within-individual covariance $\Gamma$ of each region is positive definite and satisfies $\Gamma w=(0,0,1)'$.
Because the pre-treatment components of $\Gamma w$ are zero, $B_j\Sigma_nc_s=0$ for all $j$, and
the condition of Proposition~\ref{rev:no-adjustment} holds.
\end{remark}

The same condition also holds in the following case. In the setting of \eqref{eq:regional-means}--\eqref{eq:regional-scale},
suppose that $\Gamma_{n0}$ of the treated region and $\Gamma_{ns}$ of candidate $s$ are both compound symmetric,
that is, the sum of a constant multiple of the identity matrix and a matrix with all entries equal.
Because the components of $w$ in Corollary~\ref{rev:regional-rate} sum to zero, $\Gamma_{n0}w$ and $\Gamma_{ns}w$ are
constant multiples of $w$, and their pre-treatment components are constant.
By centering, $B_j\Sigma_nc_s=0$ for all $j$.
Therefore, in this setting, the selection adjustment for candidate $s$ can remain asymptotically on the standard-error scale only if
$\Gamma_{n0}$ or $\Gamma_{ns}$ has heteroskedasticity over time
or serial correlation that is not compound symmetric.
Even in the compound-symmetric case, the plug-in selective interval in finite samples need not coincide exactly with the conventional interval.


When the population fit is exactly tied or nearly tied, uncertainty about the selection can remain
even as the number of individual-level observations increases. This alone, however, does not imply that a selection adjustment
is needed. If the selection event is constant along the conditional line, as in Proposition~\ref{rev:no-adjustment},
the selective interval equals the conventional interval.
If, in contrast, the selection event varies in the direction of inference, the effect of the adjustment can remain
on the scale normalized by the standard error.
For example, when $\Sigma_n=\Omega/n$, the correlation between $B_jY_n$ and $c_s'Y_n$
does not depend on $n$. The unstandardized covariance
$B_j\Sigma_nc_s\to0$ alone must not be taken as evidence of asymptotic independence.
In all these cases, the procedure computes the conditional line and truncation set from the estimated covariance without a preliminary test.

\section{Control selection under staggered adoption}\label{sd:section}
\subsection{Cohort--period comparisons and causal effects}
We keep the numbers of regions $K$ and periods $T$ fixed and include each region--period mean once in $Y_n$. In this section, regions are indexed by $j=1,\ldots,K$. Region $j$ first receives an irreversible treatment at $G_j\in\{2,\ldots,T,\infty\}$, with $G_j=\infty$ for never-treated regions. Treatment dates, reporting periods, and the weights below are fixed design quantities. If the design is random and inference is conditional on it, the following assumptions must hold under that conditional distribution; conditioning alone does not establish them.

Let $\mathcal G$ be the finite set of treated cohorts, $\mathcal J_g=\{j:G_j=g\}$, and let $q_{gj}\ge0$ be fixed weights on $\mathcal J_g$ summing to one. These weights specify the population average of interest, not a requirement that regional sample sizes be proportional to population sizes. Write
\[
 \bar Y_{ng,u}=\sum_{j\in\mathcal J_g}q_{gj}\bar X_{nj,u}.
\]
Fix a finite reporting set $\mathcal A\subset\{(g,t):g\in\mathcal G,\ t\ge g\}$. We allow a known anticipation bound $\kappa\ge0$: the observed mean in region $j$ equals its untreated mean at every $u<G_j-\kappa$. The case $\kappa=0$ is no anticipation. Assume consistency and no spillovers.

For a cell $a=(g,t)\in\mathcal A$, choose $P_a\subset\{u:u<g-\kappa\}$ with $|P_a|\ge2$ and fixed weights $\pi_{a,u}\ge0$, $\sum_{u\in P_a}\pi_{a,u}=1$. Let $\mathcal C_a$ be a fixed finite nonempty candidate pool. Each candidate $k\in\mathcal C_a$ is a prespecified group with weights $v_{akj}\ge0$ satisfying
\begin{equation}\label{sd:eligible}
 \sum_jv_{akj}=1,\qquad
 v_{akj}>0\ \Longrightarrow\ G_j>t+\kappa.
\end{equation}
Thus the candidate group is untreated and unaffected by anticipation through the comparison time. A single candidate region is the special case of a unit weight. Different candidate groups may overlap, and the same region may be reused for different cohorts or periods. Define
\begin{align}
 D_{n,ak,u}&=\bar Y_{ng,u}-\sum_jv_{akj}\bar X_{nj,u},\\
 T_{n,ak}&=D_{n,ak,t}-\sum_{u\in P_a}\pi_{a,u}D_{n,ak,u}
             =d_{n,ak}'Y_n,\label{sd:cell-contrast}\\
 Q_{n,ak}(Y_n)&=\sum_{u\in P_a}
  \left(D_{n,ak,u}-|P_a|^{-1}\sum_{v\in P_a}D_{n,ak,v}\right)^2
       =\|B_{n,ak}Y_n\|^2=Y_n'M_{n,ak}Y_n.\label{sd:cell-score}
\end{align}
Here $M_{n,ak}=B_{n,ak}'B_{n,ak}$, and all coefficient matrices are deterministic. The control-group weights are fixed rather than estimated by continuous optimization using the same outcomes; a finite collection of prespecified weighting schemes may instead define the candidate pool.

Let $\mu^0_{nj,u}$ be the untreated mean and $\mu^{g}_{nj,t}$ the mean under adoption at $g$. Put
\begin{align}
 \tau_{n,g,t}&=\sum_{j\in\mathcal J_g}q_{gj}
                     (\mu^g_{nj,t}-\mu^0_{nj,t}),\label{sd:att}\\
 m^0_{n,ak,u}&=\sum_{j\in\mathcal J_g}q_{gj}\mu^0_{nj,u}
                           -\sum_jv_{akj}\mu^0_{nj,u},\\
 \Delta_{n,ak}&=m^0_{n,ak,t}
                     -\sum_{u\in P_a}\pi_{a,u}m^0_{n,ak,u}.\label{sd:bias}
\end{align}
The candidate-specific parallel-trends restriction is $\Delta_{n,ak}=0$.

\begin{proposition}[Identification and aggregation]\label{sd:identification}
Under consistency, no spillovers, the anticipation restriction, and \eqref{sd:eligible},
\begin{equation}\label{sd:decomposition}
 d_{n,ak}'\mu_n=\tau_{n,g,t}+\Delta_{n,ak}.
\end{equation}
For a path $s=(s_a)_{a\in\mathcal A}$ and any fixed reporting weights $w_{\ell,a}$, define
\begin{align}
 c_{n,\ell,s}&=\sum_{a\in\mathcal A}w_{\ell,a}d_{n,a,s_a},&
 \vartheta_{n,\ell,s}&=c_{n,\ell,s}'\mu_n,\label{sd:aggregate}\\
 \tau_{n,\ell}&=\sum_{a=(g,t)\in\mathcal A}w_{\ell,a}\tau_{n,g,t}.\nonumber
\end{align}
Then
\begin{equation}\label{sd:aggregate-id}
 \vartheta_{n,\ell,s}=\tau_{n,\ell}
                       +\sum_aw_{\ell,a}\Delta_{n,a,s_a}.
\end{equation}
In particular, parallel trends for the contributing comparisons identifies the aggregate effect, without homogeneity of treatment effects across cohorts or times.
\end{proposition}

Nonnegative reporting weights summing to one give an average effect; signed weights give differences of effects. An event-time effect at $e\ge0$ uses nonzero weights only for cells $(g,g+e)$, with fixed cohort weights summing to one. The simultaneous inference results below apply to a fixed finite list of these event-time effects. A reference coefficient that is identically zero is reported as zero, rather than treated as a nonzero random contrast. This construction separates cohort--period identification and aggregation as in \citet{CS2021}; it does not rely on a homogeneous-effect interpretation of two-way-fixed-effects event-study coefficients \citep{SA2021}.

\subsection{Selection paths and the conditioning event}
Choose
\[
 S_{n,a}=\arg\min_{k\in\mathcal C_a}Q_{n,ak}(Y_n),\qquad
 S_n=(S_{n,a})_{a\in\mathcal A},\qquad
 \mathcal S=\prod_{a\in\mathcal A}\mathcal C_a,
\]
with a fixed tie-breaking rule. If several candidates have scores that are identical as functions of $Y_n$, we retain in advance the candidate preferred by that rule. For the remaining comparisons assume $M_{n,ak}-M_{n,ak'}\ne0$ whenever $k\ne k'$. Some paths may be impossible; the statements below allow zero-probability labels. Outside algebraic boundaries, the event $S_n=s$ equals
\begin{equation}\label{sd:path}
 \widetilde{\mathcal E}_{n,s}
 =\bigcap_{a\in\mathcal A}\bigcap_{k\ne s_a}
  \{y:y'(M_{n,ak}-M_{n,a,s_a})y\ge0\}.
\end{equation}
To use one control for all reported periods of a cohort, use a common score and tie-breaking rule and restrict the pool to groups eligible through the last reported period.

Let $\rho:\mathcal S\to\mathcal H$ be a prespecified map to a finite set and put $H_n=\rho(S_n)$. The identity map conditions on the entire path. A coarser map is permitted for reports $\ell=1,\ldots,q$ if
\begin{equation}\label{sd:compatibility}
 \rho(s)=\rho(s')=h\quad\Longrightarrow\quad
 c_{n,\ell,s}=c_{n,\ell,s'}=:c_{n,\ell,h}\ne0
 \quad(\ell=1,\ldots,q).
\end{equation}
The estimator and target are then determined by $h$, and the conditioning set is
\begin{equation}\label{sd:union-event}
 \widetilde{\mathcal E}_{n,h}
      =\bigcup_{s:\rho(s)=h}\widetilde{\mathcal E}_{n,s},
 \qquad E_{n,h}=\{H_n=h\}.
\end{equation}
For example, reporting only one cohort's effect may allow other cohorts' control choices to be left unconditioned. The guarantee is then conditional on $H_n$, not on the finer path $S_n$. The map must not itself be chosen after comparing confidence intervals. Coarsening does not imply a pointwise ordering of interval lengths.

For a positive definite matrix $W$, suppress $n$ and put
\begin{align}
 v_{\ell,h}(W)&=c_{\ell,h}'Wc_{\ell,h},&
 b_{\ell,h}(W)&=Wc_{\ell,h}/v_{\ell,h}(W),\\
 R_{\ell,h}(W)&=Y-b_{\ell,h}(W)c_{\ell,h}'Y,&
 \mathcal T_{\ell,h}(r;W)&=
    \{z:r+b_{\ell,h}(W)z\in\widetilde{\mathcal E}_{h}\}.
                                                        \label{sd:fiber}
\end{align}
The truncation set is a finite union of intervals: first intersect the quadratic restrictions within each path, then take the union over paths. Merge overlapping intervals so that each is integrated only once. Define
\begin{equation}\label{sd:pivot}
 U_{\ell,h}(\vartheta;W)
 =\frac{\displaystyle\int_{\mathcal T_{\ell,h}(R_{\ell,h}(W);W)\cap(-\infty,c_{\ell,h}'Y]}
                     e^{-(z-\vartheta)^2/(2v_{\ell,h}(W))}\,dz}
        {\displaystyle\int_{\mathcal T_{\ell,h}(R_{\ell,h}(W);W)}
                     e^{-(z-\vartheta)^2/(2v_{\ell,h}(W))}\,dz},
\end{equation}
and let
\begin{equation}\label{sd:ci}
 C_{\ell,h}(\alpha;W)
   =\{\vartheta:\alpha/2\le U_{\ell,h}(\vartheta;W)\le1-\alpha/2\}.
\end{equation}
During inversion, we vary only the hypothesized mean; $W$, the residual, and the truncation set are computed from the observed data and held fixed. In the plug-in procedure, however, each of these quantities is recomputed using $\widehat\Sigma_n$ rather than the reference covariance.

\begin{theorem}[Joint selection, coarsened paths, and estimated covariance]\label{sd:main}
Assume the fixed-dimensional design above, \eqref{sd:compatibility}, and fixed finite $\mathcal A,\mathcal C_a,\mathcal H$, and $q$. Fix $0<\alpha<1$.

\textnormal{(i)} If $Y\sim N_d(\mu,\Sigma)$ with known $\Sigma\succ0$, then, for every $h$ with $P(E_h)>0$, the pivot in \eqref{sd:pivot}, evaluated at $\vartheta_{\ell,h}=c_{\ell,h}'\mu$, is uniform conditional on $E_h$ and its own residual $R_{\ell,h}(\Sigma)$, for almost every residual under that event. Consequently $C_{\ell,h}(\alpha;\Sigma)$ has exact conditional coverage $1-\alpha$ given $E_h$.

\textnormal{(ii)} Suppose Assumption~\ref{ass:array} holds. Use $W=\widehat\Sigma_n=L_n\widehat V_nL_n'$ in \eqref{sd:pivot}--\eqref{sd:ci}. On any comparison boundary, or if the used covariance is not positive definite, use the convention $U=1/2$ and $C=\mathbb R$. Write $\widehat U_{n,\ell,h}$ for the pivot at $\vartheta_{n,\ell,h}=c_{n,\ell,h}'\mu_n$. Then
\begin{equation}\label{sd:joint-limit}
 D_n^{\mathrm{sd}}:=\sup_{P\in\mathcal P_n}\max_{\ell,h}\sup_{u\in[0,1]}
 \left|P(\widehat U_{n,\ell,h}\le u,E_{n,h})-uP(E_{n,h})\right|
 \longrightarrow0.
\end{equation}
For $P(E_{n,h})>0$ the conditional coverage error of $C_{n,\ell,h}(\alpha;\widehat\Sigma_n)$ is at most $2D_n^{\mathrm{sd}}/P(E_{n,h})$. It therefore tends to zero uniformly when $P(E_{n,h})\ge p_0>0$. For the target actually reported,
\begin{equation}\label{sd:marginal}
 \sup_{P\in\mathcal P_n}
 \left|P\{\vartheta_{n,\ell,H_n}\in
 C_{n,\ell,H_n}(\alpha;\widehat\Sigma_n)\}-(1-\alpha)\right|
 \le2|\mathcal H|D_n^{\mathrm{sd}}\longrightarrow0.
\end{equation}
No positive lower bound on individual selection probabilities is needed for \eqref{sd:marginal}. No independence among the different cohorts' selections, no common regional covariance, no positive limiting regional sample share, and no separation of population fit scores is assumed.
\end{theorem}

For the identity map $\rho$, the asymptotic assertion follows from the polynomial-event theorem after encoding \eqref{sd:path}. The union in \eqref{sd:union-event} need not itself be an intersection of quadratic inequalities; the proof in Appendix~\ref{sd:appendix} supplies this additional step directly. The same proof covers any fixed finite union of intersections of nonzero deterministic polynomials of degree at most two. Part (ii) guarantees coverage conditional on $E_{n,h}$, not uniform coverage conditional on every value of an estimated continuous residual.

\begin{corollary}[Simultaneous event-time and aggregate inference]\label{sd:bands}
Under Theorem~\ref{sd:main}, choose fixed $\alpha_\ell>0$ with $\sum_{\ell=1}^q\alpha_\ell\le\alpha<1$ and form $C_{n,\ell,h}(\alpha_\ell;\widehat\Sigma_n)$ for each report. Then
\begin{align}
 P\{\exists\ell:\vartheta_{n,\ell,h}\notin C_{n,\ell,h},\ E_{n,h}\}
   &\le \alpha P(E_{n,h})+2qD_n^{\mathrm{sd}},\label{sd:band-joint}\\
 \inf_{P\in\mathcal P_n}P\{\vartheta_{n,\ell,H_n}\in C_{n,\ell,H_n}
                      (\ell=1,\ldots,q)\}
   &\ge1-\alpha-2q|\mathcal H|D_n^{\mathrm{sd}}.\label{sd:band-marginal}
\end{align}
Conditional simultaneous coverage is at least $1-\alpha-2qD_n^{\mathrm{sd}}/P(E_{n,h})$. Under the exact Gaussian model with known covariance the corresponding bounds hold with $D_n^{\mathrm{sd}}=0$.
\end{corollary}

Each component interval is computed with its own residual and truncation set. The corollary applies a union bound to probabilities conditional only on the common selection event; it does not condition simultaneously on all component residuals. It allows arbitrary dependence among event-time estimators. The resulting rectangular confidence band has the stated simultaneous coverage; the critical values are obtained by Bonferroni correction, without an optimality result.

\begin{corollary}[Causal coverage and selection of valid comparisons]\label{sd:causal-coverage}
In addition to the conditions of Corollary~\ref{sd:bands}, assume the causal model of Proposition~\ref{sd:identification}. Define
\begin{equation}\label{sd:bad-selection}
 \eta_n=\sup_{P\in\mathcal P_n}P\left\{
  \exists\ell:\sum_aw_{\ell,a}\Delta_{n,a,S_{n,a}}\ne0\right\}.
\end{equation}
The simultaneous coverage of the causal effects satisfies
\begin{equation}\label{sd:causal-bound}
 \inf_{P\in\mathcal P_n}P\{\tau_{n,\ell}\in C_{n,\ell,H_n}
                         (\ell=1,\ldots,q)\}
 \ge1-\alpha-2q|\mathcal H|D_n^{\mathrm{sd}}-\eta_n.
\end{equation}
If all contributing candidate comparisons satisfy parallel trends, $\eta_n=0$. More generally, $\eta_n\to0$ suffices. For a fixed label with $P(E_{n,h})\ge p_0$, conditional causal coverage has lower bound $1-\alpha-(2qD_n^{\mathrm{sd}}+\eta_n)/p_0$. A single selected control and a single report give the corresponding result for \eqref{sd:single-att}.
\end{corollary}

Equation~\eqref{sd:bad-selection} imposes a condition on the probability of selecting a comparison that violates the identifying restriction; it does not follow from the pre-treatment score comparison. The pivot does not estimate or remove departures from parallel trends. Near-zero but nonzero violations require a further approximation or a sensitivity model; they are not covered by setting $\eta_n=0$.

\begin{corollary}[Individual panels and independent repeated cross-sections]\label{sd:sampling}
For individual panels, suppose the independent regional samples satisfy the common moment bounds, eigenvalue bounds, and $\min_j n_j\to\infty$ of Corollary~\ref{cor:regional}. Use the joint $Y_n$ and $\widehat\Sigma_n$ in \eqref{eq:regional-means}--\eqref{eq:regional-cov-est}, with each region included once. Then Theorem~\ref{sd:main} and Corollary~\ref{sd:bands} hold for the staggered design.

For independent repeated cross-sections, let the samples in the region--period cells be mutually independent, with cell sizes $n_{jt}$, centered $(2+\delta)$ moments uniformly bounded, and individual variances $v_{n,jt}$ bounded above and away from zero. If $\min_{j,t}n_{jt}\to\infty$, use
\begin{equation}\label{sd:rcs}
 L_n=\operatorname{diag}_{j,t}(n_{jt}^{-1/2}),\qquad
 V_n=\operatorname{diag}_{j,t}(v_{n,jt}),\qquad
 \widehat\Sigma_n=\operatorname{diag}_{j,t}(\widehat v_{n,jt}/n_{jt}),
\end{equation}
where $\widehat v_{n,jt}$ is the within-cell sample variance. The same conclusions hold. Neither case requires positive limits of the regional or cell sample shares.
\end{corollary}

Even if regional blocks or region--period cells are independent, overlapping comparisons generally are not. For two reports the relevant covariance is
\[
 \operatorname{Cov}(c_{n,\ell,s}'Y_n,c_{n,m,s}'Y_n)
       =c_{n,\ell,s}'\Sigma_nc_{n,m,s}.
\]
Duplicating a region in the outcome vector and assigning independent covariance blocks to its repeated appearances would give a different, incorrect covariance. With estimated aggregation weights or nonlinear covariate-adjusted DiD estimators, one must account for all terms in their joint first-order expansion and establish the representation of the selection event. Treating estimated coefficients as fixed does not justify applying the theorem.

\subsection{Why a cohort-by-cohort check can be insufficient}
\begin{proposition}[Orthogonality to the joint selection statistics]\label{sd:orthogonality}
In the exact Gaussian model, fix a report and a realized label with coefficient $c$. Let $B$ stack all the pre-treatment comparison matrices used in \eqref{sd:path}. Then $c'Y$ is independent of $BY$ if and only if
\begin{equation}\label{sd:orthogonality-eq}
 B\Sigma c=0.
\end{equation}
Under \eqref{sd:orthogonality-eq}, for every feasible residual the selection is constant along the conditioning line, and \eqref{sd:ci} equals the conventional Gaussian interval, for the full path or any compatible coarsening. It is enough to check the matrices determining the chosen conditioning event. The condition is not claimed to be necessary for independence from the coarser event itself.
\end{proposition}

\begin{proposition}[Another cohort's selection can matter under spherical covariance]\label{sd:counterexample}
Let all 16 observations $Y_{j,u}$, $j\in\{E,L,A,B\}$, $u=1,\ldots,4$, be independent $N(0,1)$. The early cohort starts treatment at 3, the late cohort at 4, and $A,B$ are never treated. Fix $A$ as the early control and set
\[
 T=Y_{E,3}-\tfrac12(Y_{E,1}+Y_{E,2})
      -Y_{A,3}+\tfrac12(Y_{A,1}+Y_{A,2}),\qquad W=T/\sqrt3.
\]
The late cohort uses periods 2 and 3 to select between $A$ and $B$ by centered squared fit. If its selected label is $S_L$, then
\begin{equation}\label{sd:counterexample-moments}
 \mathbb E(W^2\mid S_L=A)=1-\frac{\sqrt3}{4\pi},\qquad
 \mathbb E(W^2\mid S_L=B)=1+\frac{\sqrt3}{4\pi}.
\end{equation}
Consequently, the usual standardized early-cohort estimator is not standard normal conditional on the full selection path, although the outcome covariance is the identity and its own control was fixed.
\end{proposition}

If only the early effect is reported under a prespecified rule that does not condition on $S_L$, its ordinary Gaussian interval remains exact. The proposition concerns the stronger, full-path conditional guarantee. The second-moment calculation alone does not establish the direction of error for a particular 95\% interval. It also shows why a single-adoption orthogonality check cannot simply be repeated separately for each cohort when all selections are included in the conditioning event.

\section{Simulations}\label{sec:simulation}
\subsection{Designs and procedures compared}
We fix the treated region 0 and the candidate regions 1 and 2, and generate two variables $(V_{ji},H_{ji})$ for each of the independent individual-level observations within each region. The variable $V$ represents the pre-treatment change, and $H$ the difference between the post-treatment value and the pre-treatment mean. The selection is
\[
 S=\arg\min_{j=1,2}|\bar V_0-\bar V_j|,
 \qquad T_S=\bar H_0-\bar H_S.
\]
This corresponds to the centered criterion and DiD with two pre-treatment periods and one post-treatment period. Specifically, with an auxiliary variable $B$, we can write $X_1=B-V/2$, $X_2=B+V/2$, $X_3=B+H$. Then $Q_j=(\bar V_0-\bar V_j)^2/2$, and the DiD is $\bar H_0-\bar H_j$. Because the computation requires only $(V,H)$, we generate these two individual-level variables directly. With an independent $B\sim N(0,1)$ added, the design can also be represented as a three-period panel with a positive-definite covariance matrix.

The regional covariances of the individual-level observations are
\begin{equation}\label{eq:sim-cov}
 G_0=\begin{pmatrix}.9&.65\\.65&1\end{pmatrix},\qquad
 G_1=\begin{pmatrix}1.2&.30\\.30&1.4\end{pmatrix},\qquad
 G_2=\begin{pmatrix}.7&.50\\.50&.8\end{pmatrix}.
\end{equation}
We obtain the centered individual-level observations by transforming independent standard normal variables, or independent $(\mathrm{Gamma}(2,1)-2)/\sqrt2$ variables, with the Cholesky factor. We set the mean of every $H$ to 0, and the true target is 0 for either candidate. We specify the means of $V$ and the regional sample sizes in four ways.
\begin{enumerate}
\item Exact tie (T0): $\mathbb E V=(0,0,0)$, $(n_0,n_1,n_2)=(r,2r,3r)$.
\item Tie at a positive fit value (T+): $\mathbb E V=(0,.5,-.5)$; sample sizes as in T0.
\item Separated population ranking (S): $\mathbb E V=(0,.12,.65)$; sample sizes as in T0.
\item Unbalanced sample sizes with drifting means (U): $\mathbb E V=(0,.4r^{-1/4},-.4r^{-1/4}-.5r^{-1/2})$, $(n_0,n_1,n_2)=(r,\operatorname{round}(r^{1.3}),2r)$.
\end{enumerate}
We use $r=25,100,400$, combining two distributions, four mean and sample-size settings, and three values of $r$ to obtain 24 designs with 5,000 replications each. In U, the sample shares of regions 0 and 2 tend to zero as $r\to\infty$. We compute the regional sample covariances from the same individual-level observations used for the selection and DiD, and divide them by the sample sizes to estimate the covariance of the means. We use no external calibration sample and no sample splitting.

We compare the conventional interval, the plug-in selective interval, which uses the covariance estimated from the same individual-level observations, and the truncated-normal interval based on the true covariance. The last is an exact finite-sample benchmark for normal observations, but for Gamma observations it is a benchmark that relies on the normal approximation. We determine coverage by whether the pivot at the true target lies in $[.025,.975]$. This is equivalent to obtaining the interval by inversion and checking whether it contains the target, and it avoids endpoint-search error in each replication.

\subsection{Results}
Table~\ref{tab:main} reports the unconditional coverage of the selected target. The coverage of the plug-in selective interval is 94.06--95.60\% across the 24 designs and 94.62--95.60\% in the 8 designs with $r=400$. The Monte Carlo standard error (MCSE) in each design is about 0.3 percentage points. The finite-sample coverage rates need not equal the nominal level or improve monotonically with sample size.
\begin{table}[tbp]
\centering\small
\caption{Unconditional coverage (\%) of 95\% intervals: 5,000 replications per design}\label{tab:main}
\begin{tabular}{llrrrrrr}
\toprule
& &\multicolumn{3}{c}{Normal observations}&\multicolumn{3}{c}{Gamma observations}\\
\cmidrule(lr){3-5}\cmidrule(lr){6-8}
Design&$r$&Conventional&Plug-in&True $\Sigma$&Conventional&Plug-in&True $\Sigma$\\
\midrule
Exact tie & 25 & 95.62 & 94.34 & 95.10 & 96.26 & 94.88 & 95.34\\
 & 100 & 95.68 & 94.66 & 94.88 & 96.14 & 94.86 & 95.12\\
 & 400 & 96.52 & 94.94 & 95.08 & 96.20 & 95.00 & 95.06\\
\addlinespace
Positive-fit tie & 25 & 94.68 & 94.50 & 94.94 & 94.32 & 94.66 & 95.46\\
 & 100 & 95.34 & 95.16 & 95.24 & 94.86 & 95.16 & 95.26\\
 & 400 & 94.32 & 94.62 & 94.88 & 95.10 & 94.76 & 94.82\\
\addlinespace
Separated & 25 & 94.60 & 94.82 & 95.24 & 94.54 & 94.26 & 94.76\\
 & 100 & 94.44 & 94.44 & 94.74 & 95.00 & 95.04 & 95.04\\
 & 400 & 95.60 & 95.60 & 95.68 & 94.80 & 94.80 & 94.76\\
\addlinespace
Unbalanced, drifting & 25 & 94.16 & 94.06 & 94.78 & 94.64 & 94.82 & 95.12\\
 & 100 & 94.88 & 94.82 & 94.86 & 94.96 & 95.20 & 95.04\\
 & 400 & 95.32 & 95.38 & 95.36 & 94.88 & 94.66 & 94.34\\
\bottomrule
\end{tabular}
\par\vspace{2pt}{\footnotesize\noindent Note: Conventional: normal interval that treats the selection as fixed. Plug-in: selective interval based on the covariance estimated from the same individual-level observations. True $\Sigma$: truncated-normal interval based on the population covariance; for Gamma observations it does not imply finite-sample exactness. The Monte Carlo standard error of the plug-in selective interval is 0.29--0.34 percentage points in each row.\par}
\end{table}


To measure coverage at the smallest sample size more precisely, we ran the 8 designs with $r=25$ again with a different seed and 40,000 replications each (Table~\ref{tab:r25} in Appendix~\ref{app:simulation}). The unconditional coverage of the plug-in selective interval is 94.12--94.59\% (Monte Carlo standard error about 0.12 percentage points), that is, 0.4--0.9 percentage points below the nominal level. The candidate-specific conditional coverage is 92.97--95.32\% for candidates selected in at least 40\% of the replications (Table~\ref{tab:r25-conditional}). The minimum occurs for candidate 2 in the design with Gamma observations and a tie at a positive fit value, where the Monte Carlo standard error is 0.18 percentage points. In these designs a small undercoverage remains at $r=25$, and larger differences can arise for individual candidates.

In the exact-tie designs, even at $r=400$ the coverage of the conventional interval, which treats the selection as fixed, is 96.52\% for normal observations and 96.20\% for Gamma observations, higher than the 94.94\% and 95.00\% of the plug-in selective interval. Figure~\ref{fig:tie} shows this comparison and the Monte Carlo error. The conventional interval is therefore conservative in these designs, illustrating that control selection need not produce undercoverage. In the separated designs, candidate 1 was selected in all replications at $r=400$, and the coverage of the conventional and the plug-in selective intervals coincided. This behavior is consistent with Corollary~\ref{cor:separated}.

\begin{figure}[tbp]
\centering\includegraphics[width=\linewidth]{figures/exact_tie_coverage.pdf}
\caption{Coverage in the exact-tie designs. The dashed line is 95\%; the error bars are the estimated coverage $\pm1.96$ MCSE and are not simultaneous confidence bands. The horizontal axis is the sample size $r$ of region 0; regions 1 and 2 have $2r,3r$, respectively.}\label{fig:tie}
\end{figure}

For Gamma observations, the correlation across replications between the pre-treatment sample mean of region 0 and the off-diagonal element of the sample covariance was about 0.55--0.58. The Gamma designs therefore allow dependence between the sample mean and covariance estimator, as permitted by the asymptotic results, which require consistency of the standardized covariance estimator rather than independence.

The coverage conditional on each candidate, its Monte Carlo standard error, and the number of replications in which the candidate was selected are given in Appendix~\ref{app:simulation}. For a candidate that is rarely selected in the separated design, the number of replications is small or zero. These rare selections provide too few replications to assess candidate-specific conditional coverage reliably; the theorem requires a positive lower bound on the selection probability for that guarantee.

\subsection{Interval endpoints, shrinkage rates, and a design in which the adjustment vanishes}\label{sec:endpoint-sim}
To examine the behavior described by Theorem~\ref{rev:endpoints} and Corollary~\ref{rev:regional-rate} in addition to coverage,
we ran 12 designs, for the exact-tie and the unbalanced, drifting settings with normal and Gamma observations and
$r=25,100,400$, with 1,000 replications each and a separate seed.
The regional covariances, means, and sample sizes are stated in Appendix~\ref{app:endpoint-details}.
We also added 3 exact-tie designs in which $V,H$ are independent standard normal variables in each region and the sample sizes are $(r,r,r)$.
These correspond to the example with independent $V$ and $H$ that accompanies Proposition~\ref{rev:no-adjustment}.
The preceding experiment assessed coverage from the pivot. In this experiment, we also inverted the pivot in every replication
to obtain both endpoints of the plug-in interval and the interval based on the reference covariance.
\begin{table}[tbp]
\centering\small
\caption{Endpoint experiment: interval endpoints and length, 1,000 replications per design}\label{tab:endpoint-main}
\begin{tabular}{llrrrrr}
\toprule
Design&$r$&$E_{50}$&$E_{90}$&$W_{50}$&$W_{95}$&$R_{50}$\\
\midrule
N--T0 & 25 & 0.799 & 5.484 & 6.137 & 35.925 & 0.975\\
N--T0 & 100 & 0.415 & 2.795 & 6.313 & 50.347 & 0.997\\
N--T0 & 400 & 0.204 & 1.405 & 6.146 & 43.500 & 0.999\\
N--U & 25 & 0.614 & 4.613 & 5.651 & 41.194 & 0.992\\
N--U & 100 & 0.296 & 2.195 & 5.453 & 37.607 & 0.997\\
N--U & 400 & 0.161 & 1.283 & 5.478 & 45.093 & 0.997\\
G--T0 & 25 & 0.884 & 5.794 & 6.085 & 41.827 & 0.976\\
G--T0 & 100 & 0.492 & 3.110 & 6.121 & 39.912 & 0.989\\
G--T0 & 400 & 0.225 & 1.645 & 6.188 & 42.888 & 0.992\\
G--U & 25 & 0.745 & 4.464 & 5.444 & 23.698 & 0.950\\
G--U & 100 & 0.415 & 2.917 & 5.516 & 39.826 & 0.980\\
G--U & 400 & 0.219 & 1.412 & 5.432 & 34.677 & 0.997\\
N--I & 25 & 0.229 & 3.001 & 4.114 & 10.650 & 1.050\\
N--I & 100 & 0.087 & 0.675 & 3.962 & 7.076 & 1.011\\
N--I & 400 & 0.039 & 0.143 & 3.934 & 4.988 & 1.004\\
\bottomrule
\end{tabular}
\par\vspace{2pt}{\footnotesize\noindent Note: N: normal observations; G: Gamma observations. T0: exact tie; U: unbalanced, drifting; I: exact tie with independent $V,H$.
$E$ is the maximum endpoint difference from the interval based on the reference covariance, divided by the true standard error; $W$ is the length of the plug-in interval divided by $(n_0^{-1}+n_s^{-1})^{1/2}$; $R$ is the ratio of the length to that of the interval based on the reference covariance.
Subscripts denote empirical quantiles (\%). Both intervals use the same observations and the same selected candidate.\par}
\end{table}


For normal observations with an exact tie, the median of the maximum endpoint difference divided by the standard error is
0.799, 0.415, 0.204 at $r=25,100,400$, and 0.884, 0.492, 0.225 for Gamma observations.
By contrast, the median of the normalized interval length in the normal design is 6.137, 6.313, 6.146, which is
consistent with the interval on the original scale shrinking as the sample size increases.
In the same normal design, the median ratio of the plug-in interval length to that of the conventional interval is 1.536, 1.563, 1.556.
Thus the error due to covariance estimation decreases while the selection adjustment persists relative to the standard error.

In the design with independent $V,H$, the interval based on the reference covariance coincides exactly with the conventional known-variance interval.
Between the plug-in interval and the interval based on the reference covariance, the median of the maximum endpoint difference divided by the standard error decreases to 0.229, 0.087, 0.039.
The median ratio of the plug-in interval length to that of the interval based on the reference covariance (in this design, the conventional known-variance interval) is
1.050, 1.011, 1.004, and the median ratio to the conventional interval based on the same estimated covariance is
1.001, 1.000, 1.000.
These results distinguish the independent design, in which adjustment vanishes despite the tie, from designs in which selection and the DiD estimator remain dependent.
Large upper quantiles remain, so these results provide neither a common error bound for all replications nor evidence that the expected interval length converges. All quantiles and the candidate-specific coverage are given in Tables~\ref{tab:endpoint-extra}--\ref{tab:endpoint-coverage}.

The coverage in these 15 designs is based on 1,000 replications each, and the Monte Carlo standard error is about 0.6--0.8 percentage points.
The unconditional coverage of the plug-in interval is 92.4--96.1\%, and the minimum occurs in the design with Gamma observations, an exact tie, and $r=25$.
In the same design, the coverage is 94.88\% in the experiment with 5,000 replications (Table~\ref{tab:main}) and
94.39\% in the additional run with 40,000 replications (Table~\ref{tab:r25}).

\section{Scope and interpretation}\label{sec:discussion}
We generalized the selective inference for the ATT introduced by \citet{NH2026} in \emph{Economics Letters} to estimated covariance and staggered adoption. The causal question is estimation of the ATT using a control group for which parallel trends is credible. Conditioning on the control selected from a pool restricts the sampling distribution used to construct our intervals. Parallel trends then equates the selected population DiD contrast with the ATT. Proposition~\ref{sd:identification} and Corollary~\ref{sd:causal-coverage} make the distinction explicit, including a pool in which invalid comparisons are selected with probability tending to zero. Good observed pre-treatment fit alone does not imply that the probability of selecting an invalid comparison tends to zero. Sensitivity analysis for departures from parallel trends \citep{RR2023} is complementary, but using selected pre-trend estimates to calibrate it requires accounting for their joint selection. The intervals constructed here adjust for control selection; combining them with sensitivity analysis would require a separate procedure.

For staggered adoption, we condition jointly on the control choices included in the stated coverage guarantee. Theorem~\ref{sd:main} covers cohort--period-specific pools, predetermined groups of eligible controls, heterogeneous treatment effects, and reused observations. Corollary~\ref{sd:bands} gives simultaneous inference for a fixed list of dynamic or aggregate effects. When the reported parameter depends on only part of the selection path, a coarsening that satisfies the compatibility condition leaves the other control choices unconditioned. The coverage guarantee then concerns the coarsened event, while the identifying restriction is unchanged. Coarsening, simultaneous Bonferroni bands, and the hybrid intervals in \citet{AKM2024} are different constructions; we make no optimality or pointwise length comparison among them.

The asymptotic analysis keeps the numbers of regions, periods, candidate groups, and reported effects fixed while within-region sample sizes increase. Sample sizes may grow at different rates across regions or region--period cells. The full-path bound involves $|\mathcal S|=\prod_{a\in\mathcal A}|\mathcal C_a|$, while a compatible coarsening uses $|\mathcal H|$. These are fixed but potentially large constants, and the theorem supplies no numerical finite-sample accuracy bound or growing-pool guarantee. Conditional validity for a specific rare path is not implied by marginal validity of the selected target. Nor does the plug-in result assert finite-sample validity conditional on every estimated residual. The existing simulations assess the single-comparison procedures; the staggered and simultaneous extensions here are theoretical results, not additional simulation evidence.

Independent individual panels retain within-person temporal dependence; independent repeated cross-sections have a different covariance structure, as stated in Corollary~\ref{sd:sampling}. Increasing the number of independently sampled individuals within a region does not remove uncertainty from a time-varying shock common to that region. The few-group inference issues studied by \citet{DL2007}, \citet{CT2011}, and \citet{FP2019} concern this distinct source of uncertainty. Applying the general selection theorem in such an experiment requires a joint normal approximation and covariance estimator appropriate to its independent sampling or assignment units. The present within-region corollary does not by itself provide them.

For the single-control design, we also established asymptotic equivalence of interval endpoints, the interval-length rate determined by effective sample size, and conditions under which selection adjustment is unnecessary. A tie alone does not require adjustment, and an adjustment that persists relative to the standard error is compatible with a shrinking interval. The expected interval length can nevertheless be infinite. We thus obtain feasible ATT inference after control selection with unknown covariance and extend the procedure to staggered adoption.

\section*{Declaration of competing interest}
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

\section*{Funding}
This work was supported by JST SPRING, Japan [grant number JPMJSP2123] and JSPS KAKENHI [grant number 25K21651].

\section*{Data availability}
No empirical data are used. Simulation code and numerical results are provided as ancillary files. The LaTeX source includes the figure and tables used to compile this manuscript.

\section*{Declaration of generative AI and AI-assisted technologies in the manuscript preparation process}
During the preparation of this work the authors used Claude (Anthropic) and ChatGPT (OpenAI) to assist in drafting the manuscript, implementing the simulation code, and checking the content. After using these tools, the authors reviewed, verified, and edited the content as needed and take full responsibility for the content of the publication.