EconBase
← Back to paper

Simple Inference on a Simplex-Valued Weight

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,778 characters · 0 sections · 65 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 107 chars of source]
center[center omitted — 155 chars of source]

\address{Department of Economics, University of Warwick, Coventry, CV4 7AL, United Kingdom} \email{[email removed]} \address{Vancouver School of Economics, University of British Columbia, 6000 Iona Drive, Vancouver, BC, V6T 1L4, Canada} \email{[email removed]}

\fontsize{12}{14} \selectfont

abstractIn many applications, the parameter of interest involves a simplex-valued weight which is identified as a solution to an optimization problem. Examples include synthetic control methods with group-level weights and various methods of model averaging and forecast combinations. The simplex constraint on the weight poses a challenge in statistical inference due to the constraint potentially binding. In this paper, we propose a simple method of constructing a confidence set for the weight using an adaptive test based on the projection on a polyhedral cone and prove that the method is asymptotically uniformly valid. The procedure does not require tuning parameters or simulations to compute critical values. The confidence set accommodates both the cases of point-identification or set-identification of the weight. We illustrate the method with an empirical example. { Key words: Simplex-Valued Weight; Synthetic Control; Model Averaging; Forecast Combination; Uniform Asymptotic Validity} { JEL Classification: C30, C54}

\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{Introduction}

We consider the following problem. Suppose that the population weight $w_0 \in \Delta_{K-1}$ in the $(K-1)$-simplex $\Delta_{K-1}$ is defined as a solution to the following optimization problem:

align[align omitted — 110 chars of source]

where $Q(w)$ is the population objective function that is convex and differentiable in $w \in \mathbf{R}^{K}$. The dimension $K$ is fixed and not allowed to depend on the sample size $n$. Our focus is on constructing a $(1-\alpha)$-level confidence set $C_{1-\alpha}$ for $w_0$ that is uniformly asymptotically valid.

This inference problem arises in many contexts of applications. For example, in the causal inference literature, various methods of synthetic control design choose a weight as a solution to ((ref)) (see Abadie:21:JEL for a survey of this literature). Meanwhile, in the literature of forecast combinations, the final forecast is constructed as a weighted average of forecasts from different methods or experts. (See Timmermann:06:Handbook.) While not in consensus, the use of a simplex-constrained weight has been part of the main approaches considered in this literature. The least squares model averaging proposed by Hansen:07:Eca also falls into this framework. In particular, Hansen:07:Eca proposed minimizing the Mallows metric to obtain the optimal weight. The population version of the minimization problem can be written as the optimization problem in ((ref)).

The main challenge in the theory of statistical inference on the weight is that the parameter is potentially on the boundary, as $w_0$ can fall on an edge or a vertex of the simplex $\Delta_{K-1}$. Hence, even if $w_0$ is point-identified and $\sqrt{n}$-consistently estimable, it is far from obvious how to construct a pivotal test statistic whose limiting distribution is invariant uniformly over the probabilities in the model.

Despite the challenge, there are methods which can be applied to develop an asymptotic inference procedure. For example, when $w_0$ satisfies the first order condition for the optimization problem ((ref)) - an assumption this paper does not assume, we may build a quadratic approximation of the objective function to derive the limiting distribution that depends on the localization parameter (Geyer:94:AS, Andrews:99:Ecma and Moon/Schorfheide:09:JOE), or apply the conditional likelihood ratio approach by Ketz:18:JOE. Alternatively, we could adapt the inference based on the random draws from a quasi-posterior in Chen/Christensen/Tamer:18:Eca to this setting or employ the bootstrap-based projection method of Fang/Seo:21:Eca. These methods target a substantially more general setting than simplex-valued weights. However, they require bootstrap or simulations to obtain critical values and often involve tuning parameters that require a delicate choice to ensure a good finite sample performance.

In this paper, we propose a simple method of constructing a confidence set for the simplex-valued weight which is free of tuning parameters and does not require simulations to compute critical values. Furthermore, our method does not require point-identification of the weight $w_0$ and, hence, does not rely on the quadratic approximation of the objective function. The method is shown to produce confidence sets that are asymptotically valid uniformly over the behavior of the population weight. We provide simulation results and an empirical application to demonstrate the merits of our method.

Our method relies on a likelihood ratio-type test statistic constrained to a polyhedral cone. When the test statistic is formed from a multivariate normal random vector, its distribution is known to follow a mixture of a $\chi^2$ distribution with generally unknown weights (see Silvapulle/Sen:05:ConstrainedStatInference and references therein). In this setting, AlMohamad/vanZwet/Cator/Goeman:20:Biometrika developed an adaptive inference method that selects the relevant $\chi^2$ distribution automatically. In econometrics, Breunig/Chen:20:WP, Breunig/Chen:24:Eca proposed a related adaptive approach for testing equality or inequality restrictions on nonparametric functions identified in a nonparametric instrumental variable model. A related idea is also found in Cox/Shi:23:ReStud who considered a general moment inequality model with an additive nuisance parameter and proposed asymptotic inference methods. These approaches simplify the inference procedure for testing problems with constraints by using $\chi^2$ distributions with data-dependent degrees of freedom. Our paper builds on this line of research by developing a simple adaptive procedure for asymptotic inference on simplex-valued weights. However, both the specific design of our proposal and the proof of its uniform validity are new, to the best of our knowledge.\footnote{Related to this literature, recently, Li:25:WP developed an interesting method of bootstrap-based inference on parameters identified as a constrained optimizer. Li:25:WP does not require quadratic approximation of the objective function, yet involves a sequence of tuning parameters that converge to zero. In contrast to Li:25:WP, our paper focuses on the inference on a simplex-valued weight, and as such, our proposal is tailored to this case, and does not require point-identification of the weight or tuning parameters that go to zero at a certain rate.}

The rest of the paper proceeds as follows. In Section 2, we introduce a basic set-up with examples, and present our main proposal to construct a confidence set for a simplex-valued weight. In Section 3, we provide our main result of uniform asymptotic validity of the confidence set, and Monte Carlo simulations that study the finite sample properties of our statistical procedure. An empirical application is presented in Section 4 as an illustration. In Section 5, we conclude. Mathematical proofs of the results in this paper are found in the appendix.

\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{A Simple Confidence Set for a Simplex-Valued Weight}

\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{A Simplex-Valued Weight}

Let us formally introduce the problem. Suppose that $H_P$ is a $K \times K$ matrix and $h_P$ is a $K$-dimensional vector which potentially depends on the population distribution $P$. Let $\Delta_{K-1}$ denote the simplex in $\mathbf{R}^K$, i.e., $\Delta_{K-1}\coloneqq \{w \in \mathbf{R}^K: w \ge 0, w'\mathbf{1} = 1\}$, where $\mathbf{1}$ is the $K \times 1$ vector of ones. (Here, the inequality between two vectors should be understood as the set of pointwise inequalities between the corresponding entries.)

Given a map $Q_P: \mathbf{R}^K \rightarrow \mathbf{R}$, we define the argmin set of weights, $\mathbb{W}_P$, as follows:

align[align omitted — 143 chars of source]

We assume that $Q_P$ is convex and differentiable on $\mathbf{R}^K$ and there exists a true weight $w_0$ in the set $\mathbb{W}_P$.

example[Synthetic Control with Group-Level Weights] Suppose that we have $K+1$ large groups of individual units which are observed over $T+1$ periods. The units in group $0$ are treated at $T+1$, but all the other units stay untreated. For each time $t=1,...,T+1$, $[Y_{i,t},G_{i,t}]$'s are i.i.d.\ across $i$'s, $Y_{i,t}$ denotes the observed outcome variable and $G_{i,t}$ the group membership taking values in $\{0,1,2,...,K\}$. Let $\mu_{j,t} = \mathbf{E}[Y_{i,t} \mid G_{i,t} = j]$. Let $Y_{i,T+1}(1)$ be the potential outcome of an individual $i$ at time $T+1$ when her group $G_{i,T+1}$ is treated first time at time $T+1$ and $Y_{i,T+1}(0)$ that when her group is never treated. We are interested in the average treatment effect on the treated for the target group $0$: \begin{align*} \theta_0 \coloneqq \mathbf{E}[Y_{i,T+1}(1) - Y_{i,T+1}(0) \mid G_{i,T+1} = 0]. \end{align*} For the identification of $\theta_0$, synthetic control with group-level weights replaces a counterfactual untreated average outcome for the target group 0 by a weighted average of donor pool group outcomes. More specifically, synthetic control imposes the following assumption: \begin{align*} \mu_{0,T+1}(0) = \sum_{j=1}^K \mu_{j,T+1}(0) w_{0,j}, \end{align*} where $\mu_{j,T+1}(0) = \mathbf{E}[Y_{i,T+1}(0) \mid G_{i,T+1} = j]$ and $w_0 = [w_{0,1},...,w_{0,K}]'$ is chosen to be such that \begin{align*} w_0 \in \operatorname*{\arg\!\min}_{w \in \Delta_{K-1}} \frac{1}{2} \sum_{t=1}^T\left( \mu_{0,t} - \sum_{j=1}^K w_{j} \mu_{j,t} \right)^2. \end{align*} Then, the target parameter $\theta_0$ is identified as $\theta_0 = \theta_P(w_0) \in \mathbf{R}$, with \begin{align*} \theta_P(w) = \mu_{0,T+1} - \sum_{j=1}^K \mu_{j,T+1} w_j. \end{align*} This setting of a synthetic control method is different from more frequently studied settings where the weights are assigned at the individual level (Abadie:21:JEL for a review of this literature.) In contrast, causal inference with groupwise matching assumes large groups of cross-sectional units observed over a short period of time. See Rincon/Song:25:WP for a study on this causal inference setting and references therein. To see how this example maps to our setting, let $\mu_t = [\mu_{1,t},...,\mu_{K,t}]'$, and define \begin{align*} H_P = \frac{1}{T}\sum_{t=1}^T \mu_t \mu_t' and h_P = \frac{1}{T}\sum_{t=1}^T \mu_t \mu_{0,t}. \end{align*} We take \begin{align*} Q_P(w) = \frac{1}{2} w' H_P w - w'h_P, \end{align*} and formulate the population optimization problem as follows: \begin{align} w_0 \in \operatorname*{\arg\!\min}_{w \in \Delta_{K-1}} \enspace Q_P(w). \end{align} As for the weight $w_0$, synthetic control adopts the population-level perfect pretreatment fit condition as follows (Ferman/Pinto:21:QE): \begin{align} \sum_{j=1}^K \mu_{j,t} w_{0,j} = \mu_{0,t}, for all t = 1,...,T. \end{align} In this case, we have $\partial Q_P(w_0)/\partial w = 0$. If $H_P$ is invertible, we have \begin{align} w_0 = H_P^{-1} h_P. \end{align} Hence, the perfect pre-treatment fit ((ref)) implies that the regression-based weight $w_0 = H_P^{-1} h_P$ falls automatically into the simplex $\Delta_{K-1}$. $\blacksquare$
example[Distributional Synthetic Control] Chen:JAE:2020 and Gunsilius:23:ECMA propose methods for constructing counterfactual quantiles of potential outcome distributions using synthetic controls. Chen:JAE:2020 allows the synthetic control weights to vary across quantiles, whereas Gunsilius:23:ECMA employs quantile-invariant weights. In this example, we consider Gunsilius:23:ECMA's approach. Suppose that we have $j = 0, 1, ..., K$ populations, where each population $j$ at time $t$ consists of individuals $i$ with observed outcomes $Y_{i,t}$. Let $Y_{i,t}(0)$ denote the potential outcome of individual $i$ in time $t$ for the untreated state. As in Example (ref), we assume that for each time $t$, the random vector of outcome variable and the group membership, $[Y_{i,t},G_{i,t}]$, are i.i.d.\ across $i$'s. For each $j=0,1,...,K$ and $\tau \in (0,1)$, let $\rho_{j,t}(\tau)$ denote the $\tau$-th quantile of the distribution of $Y_{i,t}$ in population $j$ in time $t$ and $\rho_{j,t}(\tau;0)$ that of $Y_{i,t}(0)$. For simplicity, we assume that we have two periods $t=1,2$, where $t=1$ denotes the pre-treatment period and $t=2$ the post-treatment period. The treatment occurs only in the target population $j=0$. Our focus of interest is the following quantile treatment effect at a given quantile $\tau^*$: \begin{align*} \theta_0 \coloneqq \rho_{0,2}(\tau^*) - \rho_{0,2}(\tau^*;0), \end{align*} where $\rho_{0,2}(\tau;0)$ denotes the $\tau$-quantile of $Y_{i,t}(0)$ for individual $i$ in population $0$. Gunsilius:23:ECMA proposes the following matching condition: \begin{align*} \rho_{0,2}(\tau;0) = \sum_{j=1}^K \rho_{j,2}(\tau) w_{j,0}, \end{align*} where $w_0 = [w_{0,1},...,w_{0,K}]' \in \Delta_{K-1}$ is identified by solving the following optimization problem: \begin{align*} w_0 \in \operatorname*{\arg\!\min}_{w \in \Delta_{K-1}}\int_0^1 \left( \sum_{j=1}^K \rho_{j,1}(\tau) w_j - \rho_{0,1}(\tau) \right)^2 d\tau. \end{align*} Then, the target parameter $\theta_0$ is identified as $\theta_P(w)$ with \begin{align*} \theta_P(w) = \rho_{0,2}(\tau^*) - \sum_{j=1}^K \rho_{j,2}(\tau^*) w_{j}. \end{align*} Let $U_P(\tau)$ be the $K \times K$ matrix whose $(j,k)$-th entry is given by $\rho_{j,1}(\tau) \rho_{k,1}(\tau)$ and $u_P(\tau)$ the $K$-dimensional vector whose $k$-th entry is given by $\rho_{k,1}(\tau) \rho_{0,1}(\tau)$. Define \begin{align*} H_P = \int_0^1 U_P(\tau) d\tau and h_P = \int_0^1 u_P(\tau) d\tau. \end{align*} Now, we can reformulate the optimization problem as in ((ref)) with $Q_P(w) = \frac{1}{2} w' H_P w - w' h_P$. $\blacksquare$
example[Forecast Combination] In the literature of forecast combination, the final forecast is constructed as a weighted average of multiple forecasts: \begin{align*} \hat w \in \operatorname*{\arg\!\min}_{w \in \Delta_{K-1}} \enspace \frac{1}{n}\sum_{t=1}^n (y_t - \mathbf{\hat y}_t' w)'(y_t - \mathbf{\hat y}_t' w), \end{align*} where $y_t$ is the target outcome and $\mathbf{\hat y}_t$ is the $K$ dimensional vector of outcomes used for constructing the forecast at time $t$. (See Timmermann:06:Handbook for a review of this literature.) The population version of this optimization problem is given by \begin{align*} w_0 \in \operatorname*{\arg\!\min}_{w \in \Delta_{K-1}} \enspace \frac{1}{n}\sum_{t=1}^n \mathbf{E}_P\left[ (y_t - \mathbf{y}_t' w)'(y_t - \mathbf{y}_t' w) \right], \end{align*} where $\mathbf{y}_t$ is the $K$ dimensional vector of population outcomes corresponding to $\mathbf{\hat y}_t$. Then, we can rewrite the optimization problem as the minimization of $Q_P(w) = \frac{1}{2} w' H_P w - w'h_P$ over $w \in \Delta_{K-1}$, where \begin{align*} H_P = \frac{1}{n}\sum_{t=1}^n \mathbf{E}_P\left[\mathbf{y}_t \mathbf{y}_t'\right] and h_P = \frac{1}{n}\sum_{t=1}^n \mathbf{E}_P\left[y_t \mathbf{y}_t\right]. \end{align*} The value of restricting the weight to the simplex has been debated in the literature and defended on the ground that it reduces the variability of the resulting forecast. (See Timmermann:06:Handbook.) Despite the popular use of forecast combinations, the theory of statistical inference on the weighted forecasts appears under-developed (see Wang/Hyndman/Li/Kang:23:IJF, p.1539.) $\blacksquare$
example[Least Squares Model Averaging] Hansen:07:Eca proposed a model averaging estimator using the Mallows metric. Let $(y_i,x_i)$, $i=1,...,n$, be a random sample, where $y_i \in \mathbf{R}$ but $x_i = (x_{i1},x_{i2},...)$ is countably infinite. The model for the outcome $y_i$ is given by \begin{align*} y_i = \mu_i + e_i, \end{align*} where $\mu_i = \sum_{\ell = 1}^\infty a_\ell x_{i \ell}$, $\mathbf{E}[e_i \mid x_i] = 0$, and $a_\ell$'s are constants. Let $0 \le \ell_1 < \ell_2 < ... < \ell_K$, where $\ell_k$'s are integers. For each $k=1,...,K$, we let $\mathbf{a}_k = [a_{1},...,a_{\ell_k}]'$ be the $\ell_k$ dimensional vector of the first $\ell_k$ coefficients. The focus here is on the model average estimator of $\mathbf{a}_K$. Let $\mathbf{\hat a}_k$ be the least squares estimator of $\mathbf{a}_k$ from projecting $Y=[y_1,...,y_n]'$ onto the column space of $X_k$, where $X_k$ denotes the $n \times \ell_k$ matrix whose $(i,j)$-th entry is given by $x_{ij}$. The model averaging estimator $\mathbf{\overline a}_K(w)$ of $\mathbf{a}_K$ takes the following form: \begin{align*} \mathbf{\overline a}_K(w) \coloneqq \hat A w, \end{align*} where $\hat A$ is the $\ell_K \times K$ matrix whose $k$-th column vector has the first $\ell_k$ dimensional subvector equal to $\mathbf{\hat a}_k$ and the rest of its entries as zeros. In determining the weight $w = [w_1,...,w_K]'$, Hansen:07:Eca proposed to minimize the Mallows metric: \begin{align*} \hat w \in \operatorname*{\arg\!\min}_{w \in \Delta_{K-1}} (Y - X_K \mathbf{\overline a}_K(w))'(Y - X_K \mathbf{\overline a}_K(w)) + 2 \sigma^2 w' \boldsymbol{\ell}, \end{align*} where $\boldsymbol{\ell} = [\ell_1,...,\ell_K]'$, and showed its asymptotic optimality property. Note that we can write \begin{align*} \hat w \in \operatorname*{\arg\!\min}_{w \in \Delta_{K-1}} \enspace \frac{1}{2} w' \hat H w - w' \hat h, \end{align*} where $\hat H = \frac{1}{n} \hat A' \left( X_K' X_K \right) \hat A$ and $\hat h = \frac{1}{n} \left(\hat A' X_K' Y - \sigma^2 \boldsymbol{\ell} \right)$. Thus, for inference, the population (pseudo-true) weight $w_0$ can be viewed as the solution to the optimization problem ((ref)), where \begin{align*} H_P = \frac{1}{n} A_P' \mathbf{E}_P\left[ X_K' X_K \right]A_P and h_P = \frac{1}{n} \left(A_P' \mathbf{E}_P\left[ X_K' Y \right] - \sigma^2 \boldsymbol{\ell} \right), \end{align*} and $A_P$ is the $\ell_K \times K$ matrix whose $k$-th column vector consists of the first $\ell_k$ dimensional subvector which is equal to $(\mathbf{E}_P[X_k'X_k])^{-1} \mathbf{E}_P[X_k' Y]$ and the rest of its entries as zeros. $\blacksquare$

\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{A Simple Confidence Set}

Suppose that the true weight $w_0$ belongs to $\mathbb{W}_P$. Our goal is to construct a confidence set for $w_0$ that is asymptotically uniformly valid over the class of population distributions $P$ under consideration. We will give a set of conditions later. For this section, we present the procedure of constructing the confidence set.

First, we consider the $K \times K$ orthogonal matrix $[\mathbf{1}/\sqrt{K},B_2]$, where $B_2$ is the $K \times (K-1)$ matrix consisting of orthonormal column vectors that are orthogonal to $\mathbf{1}$.\footnote{The computation of $B_2$ is straightforward. First note that $B_2'B_2 = I$. Hence, $B_2 B_2' = I - \mathbf{1} \mathbf{1}'/K$, i.e., the projection matrix projecting onto the orthogonal complement of the span of $\mathbf{1}$. To obtain $B_2$, we obtain a spectral decomposition : $ I - \mathbf{1} \mathbf{1}'/K = U D U'$. From this, we set $B_2$ to be the $K \times (K-1)$ matrix after removing the eigenvector from $U$ that corresponds to the zero diagonal element of $D$.} Define

align*[align* omitted — 77 chars of source]

and take $\hat \varphi(w)$ as an estimator of $\varphi_P(w)$ such that

align*[align* omitted — 97 chars of source]

as $n \rightarrow \infty$, for some positive definite matrix $V_P(w)$. Suppose that we have a consistent estimator of $V_P(w)$, denoted by $\hat V(w)$. Using $\hat \varphi$ and $\hat V$ only, we can construct a confidence set for $w_0$ as explained in Algorithm (ref).\footnote{The algorithm applies to cases where $w_0$ is partially identified. When $w_0$ is point-identified and has a consistent estimator $\hat w$, we can substitute $\hat w$ into $\hat V(w)$, replacing $\hat V(w)$ with $\hat V(\hat w)$. This reduces the computation of $\hat \lambda(w)$ to a quadratic programming problem, significantly improving computational speed.} The construction of the confidence set $C_{1-\alpha}$ is simple, involving no simulations or a sequence of tuning parameters that require a judicious choice in finite samples.

algorithm[algorithm omitted — 1,380 chars of source]

The computation of $V_P(w)$ and its consistent estimation can be done in a standard manner. For example, consider the case where

align*[align* omitted — 118 chars of source]

Then, we have

align*[align* omitted — 94 chars of source]

As for $\hat H$ and $\hat h$, suppose that these estimators admit an asymptotic linear representation as follows:

align*[align* omitted — 186 chars of source]

where $\Psi_{i,H}$ and $\psi_{i,h}$ are influence functions that are i.i.d.\ across $i$'s and have mean zero. Then, we can find that

align[align omitted — 84 chars of source]

where $\xi_i(w) = B_2'(\Psi_{i,H} w - \psi_{i,h})$. A consistent estimator $\hat V(w)$ of $V_P(w)$ is obtained as follows:

align[align omitted — 95 chars of source]

where $\hat \xi_i(w) = B_2'(\hat \Psi_{i,H} w - \hat \psi_{i,h})$, and $\hat \Psi_{i,H}$ and $\hat \psi_{i,h}$ are appropriate estimators of $\Psi_{i,H}$ and $\psi_{i,h}$. Using $\hat V(w)$, we construct the confidence set $C_{1-\alpha}$ is constructed as explained in Algorithm (ref).

example[Synthetic Control with Group-Level Weights, Cont'd] We revisit Example (ref). Recall that for each $t=1,...,T + 1$, $(Y_{i,t},G_{i,t})$, $i \in N_t$, are i.i.d.\ random vectors. (We allow their distributions to vary over time.) We let $N = \bigcup_{t=1}^{T+1} N_t$ and $n = |N|$. The conditional mean $\mu_{j,t}$ is estimated as a within-group average of the outcomes: for $j=0,1,...,K$ and $t=1,...,T + 1$, \begin{align*} \hat \mu_{j,t} = \frac{1}{n_{j,t}}\sum_{i \in N_{j,t}} Y_{i,t}, \end{align*} where $N_{j,t} = \{i \in N: G_{i,t} = j\}$. For each group $j$, we have \begin{align*} \sqrt{n}(\hat \mu_{j,t} - \mu_{j,t}) = \frac{1}{\sqrt{n}}\sum_{i \in N} \psi_{ij,t} + o_P(1), \end{align*} where, with $p_{j,t} = P\{G_{i,t} = j\}$, \begin{align*} \psi_{ij,t} = \frac{n}{n_t} \frac{1\{G_{i,t} = j\}}{p_{j,t}}(Y_{i,t} - \mu_{j,t}), \end{align*} and $n_t = |N_t|$. If we let $\psi_{i,t} = [\psi_{i1,t},...,\psi_{iK,t}]'$, we have \begin{align*} \Psi_{i,H} = \frac{1}{T}\sum_{t=1}^T \left(\psi_{i,t} \mu_t' + \mu_t \psi_{i,t}'\right) and \psi_{i,h} = \frac{1}{T}\sum_{t=1}^T \left( \mu_t \psi_{i0,t} + \psi_{i,t} \mu_{0,t} \right). \end{align*} We obtain $V_P(w) = \mathbf{E}_P\left[ \xi_i(w) \xi_i(w)'\right]$, where $\xi_i(w) = B_2'(\Psi_{i,H} w - \psi_{i,h})$. Suppose that the perfect pre-treatment fit ((ref)) is satisfied. In this case, we have \begin{align} \Psi_{i,H} w - \psi_{i,h} = \frac{1}{T}\sum_{t=1}^T \mu_t (\psi_{i,t}' w - \psi_{i0,t}) = \frac{1}{T}\sum_{t=1}^T \mu_t \left( \sum_{j=1}^K \psi_{ij,t} w_j - \psi_{i0,t} \right). \end{align} Thus, we obtain the variance formula: \begin{align} V_P(w) = B_2' Var_P \left( \frac{1}{T}\sum_{t=1}^T \mu_t \left( \sum_{j=1}^K \psi_{ij,t} w_j - \psi_{i0,t} \right)\right) B_2, \end{align} and estimate this by $\hat V(w)$: \begin{align*} \hat V(w) = B_2' \frac{1}{n}\sum_{i \in N} \left( \frac{1}{T}\sum_{t=1}^T \hat \mu_t \left( \sum_{j=1}^K \hat \psi_{ij,t} w_j - \hat \psi_{i0,t} \right)\right) \left( \frac{1}{T}\sum_{t=1}^T \hat \mu_t \left( \sum_{j=1}^K \hat \psi_{ij,t} w_j - \hat \psi_{i0,t} \right)\right)' B_2, \end{align*} where, with $\hat p_{j,t} = n_{j,t}/n_t$, \begin{align*} \hat \psi_{ij,t} = \frac{n}{n_t} \frac{1\{G_{i,t} = j\}}{\hat p_{j,t}}(Y_{i,t} - \hat \mu_{j,t}). \end{align*} When the sample is repeated cross-sections, the observations are independent across time. In this case, we can obtain sharper inference by modifying $\hat V(w)$ as follows: \begin{align} \hat V(w) = \frac{1}{n}\sum_{i \in N} \frac{1}{T^2}\sum_{t=1}^T \left( \sum_{j=1}^K \hat \psi_{ij,t} w_j - \hat \psi_{i0,t} \right)^2 B_2' \hat \mu_t \hat \mu_t' B_2. \end{align} $\blacksquare$
example[Distributional Synthetic Control, Cont'd] We revisit Example (ref). Let $\hat\rho_{j,t}(\tau)$ be the empirical $\tau$-quantile of $Y_{i,t}$, $i \in N_{j,t}$ for $j = 0,1,...,K$, where $N_{j,t} = \{i \in N: G_{i,t} = j\}$. For brevity, we assume that there exist constants $c,C>0$ such that for all $\tau \in (0,1)$, $f_{j,t}(\rho_{j,t}(\tau)) \in (c,C)$, where $f_{j,t}$ is the density of $Y_{i,t}$, $i \in N_{j,t}$, in time $t$. Then, we can show that \begin{align*} \sqrt{n}(\hat \rho_{j,t}(\tau) - \rho_{j,t}(\tau)) = \frac{1}{\sqrt{n}}\sum_{i \in N} \psi_{ij,t}(\tau) + o_P(1), \end{align*} where the $o_P(1)$ term is uniform over $\tau \in (0,1)$, and \begin{align} \psi_{ij,t}(\tau) = \frac{n 1\{i \in N_{j,t}\}}{n_{j,t}} \cdot \frac{\tau - 1\{Y_{i,t} \le \rho_{j,t}(\tau)\}}{f_{j,t}(\rho_{j,t}(\tau))}. \end{align} (See e.g. Section 2.5 of Serfling:80:Approx.\footnote{The uniformity of $o_P(1)$ in $\tau \in (0,1)$ can be shown using the empirical process theory in the standard way.}) Then, the $(j,k)$-th entry of $\Psi_{i,H}$ is given as follows: \begin{align*} \int_0^1 \left( \psi_{ij,1}(\tau) \rho_{k,1}(\tau) + \psi_{ik,1}(\tau) \rho_{j,1}(\tau) \right)d\tau, \end{align*} and the $k$-th entry of $\psi_{i,h}$ is given by \begin{align*} \int_0^1 \left( \psi_{i0,1}(\tau) \rho_{k,1}(\tau) + \psi_{ik,1}(\tau) \rho_{0,1}(\tau) \right)d\tau. \end{align*} We obtain $V_P(w) = \mathbf{E}_P\left[ \xi_i(w) \xi_i(w)'\right]$, where $\xi_i(w) = B_2'(\Psi_{i,H} w - \psi_{i,h})$. $\blacksquare$
algorithm[algorithm omitted — 1,479 chars of source]

In some applications, the computation and estimation of $V_P(w)$ can be cumbersome for practitioners, because the practitioner has to find an analytical form of $V_P(w)$ in their application to find its consistent estimator. In such cases, an alternative is to use a bootstrap estimator of $V_P(w)$. As shown by Hahn/Liao:21:Ecma, the asymptotic validity of the inference procedure is preserved under the bootstrap, while the inference procedure is conservative in general.\footnote{In some settings, we can modify the bootstrap variance estimator using a truncation method proposed by Shao:92:Stat to render the inference procedure asymptotically non-conservative. See also Goncalves/White:05:JASA.} Suppose that $\hat \varphi^*(w)$ is constructed from the bootstrap sample, such that the bootstrap distribution of $\sqrt{n}B_2'(\hat \varphi^*(w) - \hat \varphi(w))$ converges almost surely to $N(0,V_P(w))$. Suppose that $w_0$ is point-identified and the minimizer $\hat w$ of $\hat Q(w)$ over $w \in \Delta_{K-1}$ is consistent. Then, we obtain the bootstrap variance estimator:

align*[align* omitted — 168 chars of source]

where $\mathbf{E}^*$ denotes the expectation with respect to the bootstrap distribution. This suggests the modified inference procedure in Algorithm (ref). Note that $\hat V^*(\hat w)$ is outside the iteration over $w \in \Delta_{K-1}$. Hence, once the bootstrap estimator $\hat V^*(\hat w)$ is computed, the cost of the computation remains the same as before.

\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{Heuristics}

\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Construction of the Test Statistic}

Let us give a heuristic motivation behind the construction of the confidence set in Algorithm (ref). To construct a test statistic, we form the Lagrangian of the constrained optimization in ((ref)):

align*[align* omitted — 127 chars of source]

where $\tilde \lambda$ and $\lambda$ are Lagrange multipliers. The necessary and sufficient condition for $w_0$ to be a solution to the constrained optimization is:

align*[align* omitted — 75 chars of source]

for some $\tilde \lambda \in \mathbf{R}$ and $\lambda \in U(w_0)$, where

align*[align* omitted — 114 chars of source]

We concentrate out $\tilde \lambda$ to obtain the following equality:

align*[align* omitted — 91 chars of source]

for some $\lambda \in U(w_0)$. The equality has $K$ equations but these equations are not linearly independent. To extract maximally linearly independent equations out of these, we premultiply $B_2'$ to obtain the following equality restriction:

align[align omitted — 89 chars of source]

for some $\lambda \in U(w_0)$. The equality restrictions motivate the test statistic $T(w)$ in Algorithm (ref).

\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Construction of a Critical Value}

To provide an intuition behind the construction of the critical value, it helps to begin with the result of AlMohamad/vanZwet/Cator/Goeman:20:Biometrika. Suppose that $Y \sim N(\mu,V)$, for some $\mu \in \mathbf{R}^K$ and a symmetric positive definite matrix $V$. We introduce a polyhedral cone $\Lambda(w)$ as follows:

align[align omitted — 153 chars of source]

Define the norm $\| \cdot \|_{V}$ as $\|x\|_{V} = \sqrt{x' V^{-1} x}$, $x \in \mathbf{R}^K$. Then we focus on the test statistic of the form:

align*[align* omitted — 79 chars of source]

where $\Pi_{V}(Y \mid C)$ for a closed convex set $C$ denotes the projection of $Y$ on $C$ along $\| \cdot \|_{V}$. Let $\Lambda^\circ(w)$ be the polar cone of $\Lambda(w)$ and let $F_{\ell}$, $\ell=1,...,L$, be the faces of $\Lambda^\circ(w)$ and $\text{ri}(F_\ell)$ their relative interiors. Furthermore, let $k_\ell$ be the rank of the projection matrix of $Y$ onto the linear span of $F_{\ell}$. Then, Theorem 1 of AlMohamad/vanZwet/Cator/Goeman:20:Biometrika gives the following result:

align*[align* omitted — 94 chars of source]

where\footnote{AlMohamad/vanZwet/Cator/Goeman:20:Biometrika did not consider the case where the projection of $Y$ onto the polar cone becomes zero with positive probability. Thus, here we take $\max\{k_\ell,1\}$ in place of $k_\ell$ as in their paper.}

align[align omitted — 177 chars of source]

and $G(\cdot;k)$ denotes the CDF of the $\chi^2$ distribution with $k$ degrees of freedom. In order to adapt their result to our setting, we take two steps. First, we find a characterization of the critical value in our setting. Second, we extend the validity result to the desired asymptotic validity result.

We illustrate the main idea here, assuming that $V = I$ and is known. We write simply $q_{1-\alpha}(Y;w)$ instead of $q_{1-\alpha,V}(Y;w)$ and $\Pi(Y \mid C)$ instead of $\Pi_{V}(Y \mid C)$. The general case where $V$ is unknown and consistently estimated is dealt with in the proof of the main result in the appendix.

Step 1 (Critical Value Characterization): We first establish the following characterization of the critical value $q_{1-\alpha}(Y;w)$:

align*[align* omitted — 60 chars of source]

where

align*[align* omitted — 123 chars of source]

and $J_0[x] = \{j: x_j = 0\}$, the set of the indices of zeros in a vector $x$. To show this, we find the polar cone $\Lambda^\circ(w)$ of $\Lambda(w)$ as

align*[align* omitted — 102 chars of source]

where for any vector $a$, $[a]_J$ denotes the subvector of $a$ indexed by $J$. From this, the faces of $\Lambda^\circ(w)$ and their relative interiors are seen to be of the following form: for $J \subset J_0[w]$,

align*[align* omitted — 307 chars of source]

The linear span of $\text{ri}(\Lambda_{J}^\circ(w))$ is $L_J^\circ := \{x \in \mathbf{R}^{K-1}: [B_2 x]_J = 0\}$. Hence, the rank of the projection matrix onto this space is $K-1 - |J|$. Furthermore, since the relative interiors, $\text{ri}(\Lambda_{J}^\circ(w))$, partition the polyhedral cone $\Lambda^\circ(w)$, we find that

align[align omitted — 176 chars of source]

The projection of $Y$ onto $\Lambda^\circ(w)$ falls on $\text{ri}(\Lambda_{J}^\circ(w))$ if and only if the projection falls on the latter's linear span $L_J^\circ$. Therefore, only one of the terms in the sum on the right-hand side of ((ref)) realizes. Noting the decomposition

align*[align* omitted — 73 chars of source]

we write this term as $G^{-1}(1-\alpha;k(w))$.

To illustrate the degrees of freedom $k(w)$ for the $\chi^2$ distribution in the critical value, consider the case $K=3$ and $w = [1,0,0]'$. In this case, the cone $\Lambda(w)$ and its polar cone $\Lambda^\circ(w)$ are given as follows:

align*[align* omitted — 219 chars of source]

There are four faces of $\Lambda^\circ(w)$: $\Lambda_{\{2,3\}}^\circ(w)$, $\Lambda_{\{2\}}^\circ(w)$, $\Lambda_{\{3\}}^\circ(w)$, and $\Lambda_{\varnothing}^\circ(w)$. The relative interiors of these faces and their corresponding degrees of the $\chi^2$ distribution in the critical value are depicted in Figure (ref).

figure[figure omitted — 1,137 chars of source]

Step 2 (Extension to Asymptotic Validity): So far the result is for a normally distributed random vector $Y$. In order to accommodate random vectors that are asymptotically normal, we extend the previous result. This extension consists of two components: the asymptotic approximation of the test statistic and that of the critical values. For simplicity, we maintain the assumption that $\hat V_n$ is known to be $I$.

Consider a setting where

align*[align* omitted — 80 chars of source]

as $n \rightarrow \infty$. (Note that we allow $\|\mu_n\| \rightarrow \infty$ as $n \rightarrow \infty$.) From Skorohod representation, there exists a probability space and a sequence of random vectors $\tilde Z_n$ that have the same distribution as $Z_n$ and $\tilde Z_n \rightarrow_{a.s.} Z$ for a random vector $Z \sim N(0,I)$. We let $\tilde Y_n = \tilde Z_n + \mu_n$ and $Y_n^* = Z + \mu_n$. We focus on the limit behavior of the following test statistic:

align*[align* omitted — 124 chars of source]

Since the projection map on a nonempty, closed convex set is a contraction map, we have

align*[align* omitted — 154 chars of source]

as $n \rightarrow \infty$. Hence, it is not hard to see that, with $T_n^*(w) := \left\|Y_n^* - \Pi(Y_n^* \mid \Lambda(w)) \right\|^2$,

align*[align* omitted — 95 chars of source]

The main challenge is to show the asymptotic validity of the critical values. For this, we prove that for any subsequence of $\{n\}$, there exists a further subsequence $\{n'\} \subset \{n\}$, with large probability,

align[align omitted — 156 chars of source]

This yields the result that the critical value $c_n^*$ based on $Y_n^*$ is less than the critical value $\tilde c_n$ based on $\tilde Y_n$ eventually. Hence, the rejection probability $P\{T_n^*(w) > c_n^*\}$ is bounded below by $P\{\tilde T_n(w) > \tilde c_n\}$. By the construction of the critical value $c_n^*$ using the $\chi^2$ distribution, the former rejection probability is bounded by $\alpha$, delivering the asymptotic validity of the test.

It remains to show ((ref)). Since $|J_0[x]|$ is discontinuous in $x$, this result does not come directly from the convergence, $\tilde Y_n - Y_n^* \rightarrow_{a.s.} 0$. We begin by choosing any subsequence and find a further subsequence, along with some $J,J' \subset J_0[w]$ such that

align*[align* omitted — 172 chars of source]

We show that the set of values of $\tilde Y_n$ such that $J$ is not contained in $J'$ eventually has measure zero. In light of ((ref)), this yields ((ref)). While this summarizes the basic insights briefly, delivering the proof considering random variance matrices $\hat \Omega$ and the sequences of weights $w_n \in \Delta_{K-1}$ requires a considerable care for subtleties in the proofs. We refer the reader to the appendix for details.

\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{Discussions}

In this section, we discuss related methods developed in the literature and compare them with ours. For brevity, we discuss two methods that are most closely related to ours. First, we discuss the method of quadratic approximation of the sample objective function as proposed by Geyer:94:AS and Andrews:99:Ecma and a related proposal by Ketz:18:JOE. Second, we discuss Cox/Shi:23:ReStud who proposed inverting a test whose critical values are based on a $\chi^2$ distribution with a data dependent degree of freedom. Like our approach, their procedure does not involve tuning parameters that require a judicious choice to ensure stable finite sample properties.

\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Quadratic Approximation}

Our sample objective function $\hat Q(w)$ is a quadratic function of $w$. Hence, we may consider applying the method of quadratic approximation to address the issue of $w_0$ on the boundary of the simplex $\Delta_{K-1}$. To see how this method applies to our setting, consider the example of synthetic control with groupwise matching which takes the sample objective function $\hat Q(w)$ of the form:

align*[align* omitted — 66 chars of source]

where $\hat H$ and $\hat h$ are $\sqrt{n}$-consistent and asymptotically normal estimators of a positive definite matrix $H$ and a vector $h$. Now, we can write

align[align omitted — 137 chars of source]

for $w_0 \in \operatorname*{\arg\!\min}_{w \in \Delta_{K-1}} Q(w)$. The quadratic approximation method requires

align*[align* omitted — 86 chars of source]

for some random variable $\zeta$. Under regularity conditions, this requirement is met if

align[align omitted — 46 chars of source]

More generally, the quadratic approximation requires the following: at $w = w_0$,

align*[align* omitted — 59 chars of source]

This means that the global minimizer of $Q_P$ lies in the simplex $\Delta_{K-1}$. However, when the simplex constraint is binding, we do not have this equality and cannot the methods based on the quadratic approximation. The setting of synthetic control naturally uses additional equality restrictions ((ref)) which represent perfect pre-treatment fit at the population level. In this case, the condition ((ref)) is satisfied, and we can apply the method of quadratic approximation. However, as it is well known, due to the constraint imposed on the estimator $\hat w$, the limiting distribution of $\sqrt{n}(\hat w - w)$ is characterized as a minimizer of a random function. Thus, in order to obtain the critical value, we need to first simulate the random function and obtain a minimizer. In contrast to these approaches, our method is simple, as it does not require simulating the limiting distribution of the test statistic. More importantly, our method does not require that the objective function admit a quadratic approximation or the parameter $w_0$ be point-identified.

Ketz:18:JOE proposed a simple, interesting way to deal with a parameter on the boundary when the sample objective function admits a quadratic approximation. The estimator is obtained by applying a single Newton-Raphson adjustment to the constrained estimator. Let us study his approach in the context of the synthetic control setting. Let $\hat w$ be the constrained estimator of $w_0$. Then, from the quadratic approximation in ((ref)), his approach results in the following estimator:

align*[align* omitted — 93 chars of source]

This is the regression estimator by regressing $\hat \mu_{0,t}$ onto $\hat \mu_{1,t}$,...,$\hat \mu_{K,t}$. Hence, the estimator $\tilde w$ is an unconstrained estimator, allowed to take values outside $\Delta_{K-1}$. Due to this, the estimator is $\sqrt{n}$ consistent and asymptotically normal, under a regularity conditions where $w_0$ is $\sqrt{n}$-consistently estimable.

To map the construction of this test statistic to our setting, note that using the pre-treatment fit condition in synthetic control: $H w_0 = h$, we can write

align*[align* omitted — 154 chars of source]

This motivates the following quasi-likelihood ratio test statistic:

align*[align* omitted — 82 chars of source]

Under regularity conditions, we expect that $T_{\mathsf{NR}}(w_0) \to_d \chi_{K-1}^2$, as $n \to \infty$. Thus, the resulting method is non-adaptive to the possibility of $w_0$ being on the boundary.\footnote{Hsieh/Shi/Shum:22:JoE proposed a method of statistical inference on equality or inequality restrictions using a linear constraint formulation via Karush-Kuhn-Tucker conditions. We can adapt their proposal to this setting of simplex-valued weights. Like ours, their main proposal does not require a tuning parameter or simulating critical values. However, when adapted to our setting, the critical values are taken from the $\chi^2$ distribution with a degree of freedom $K+1$ (see (39) on page 258 of their paper.) Thus, the test is more conservative than the non-adaptive version above.} In contrast, our method is based on the idea from AlMohamad/vanZwet/Cator/Goeman:20:Biometrika and adaptive to $w_0$ being on the boundary. For comparison, our test statistic takes the form:

align*[align* omitted — 110 chars of source]

and the critical values are taken from $\chi_{\hat k(w)}^2$. While there is no uniform dominance of one test over the other, AlMohamad/vanZwet/Cator/Goeman:20:Biometrika present the power comparison between the two approaches and report simulation results in favor of the adaptive test.

\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Adaptive Moment Inequality Tests}

Closely related to our proposal is that of Cox/Shi:23:ReStud who proposed an adaptive size-exact subvector inference from moment inequality restrictions. They considered the moment inequality model:\footnote{Importantly, Cox/Shi:23:ReStud's framework accommodates a model with a nuisance parameter that enters the moment inequality restrictions in an additive form. To facilitate the comparison, we only discuss a simplified version of their model without a nuisance parameter here.}

align*[align* omitted — 59 chars of source]

where $A$ is a $d_A \times d_m$ matrix, $b$ is a $d_A$ dimensional vector, and $\overline m_n(\theta)$ is the sample average of the moment function $m(W_i,\theta)$:

align*[align* omitted — 77 chars of source]

To construct a confidence set, they considered the following test statistic:

align[align omitted — 154 chars of source]

where

align*[align* omitted — 141 chars of source]

Like our proposal, they constructed the critical value from a $\chi^2$ distribution with data-dependent degrees of freedom.

While our problem appears similar to theirs, their methods and asymptotic validity results do not subsume ours. First, our testing framework is not necessarily motivated from moment inequality restrictions. Hence, we cannot directly use their uniform asymptotic validity result without modification in our setting. Second, our method is simpler than theirs in our case. To see this, we rewrite $T(w)$ as follows:

align*[align* omitted — 168 chars of source]

where $\Lambda(w)$ is the polyhedron defined in ((ref)) and $\hat f(w) = B_2'\hat \varphi(w)$. To the best of our knowledge, there does not seem to be an explicit solution for a matrix $A(w)$ and a vector $b(w)$ such that

align*[align* omitted — 96 chars of source]

While one may develop an algorithm to compute $A(w)$ and $b(w)$ and apply the approach of Cox/Shi:23:ReStud, our proposal is already simpler than this, as we do not need to compute $A(w)$ and $b(w)$ for each $w \in \Delta_{K-1}$.

Alternatively, we may consider rewriting $T(w)$ as follows:

align*[align* omitted — 183 chars of source]

to map this to the setting of Cox/Shi:23:ReStud. In this case, we can explicitly find matrices $A(w)$ and a vector $b(w)$ such that

align*[align* omitted — 153 chars of source]

However, unlike $\hat \Sigma(\theta)$ in ((ref)), the matrix $B_2 \hat V(w)^{-1} B_2'$ is a singular matrix both in finite samples and in the limit. Hence, the results and proofs of Cox/Shi:23:ReStud are not directly applicable here. The singularity problem in our setting stems from the linear constraint $w'\mathbf{1} = 1$ rather than from accommodating distributions with singular covariance matrices in uniformly asymptotically valid inference. Therefore, simply removing such distributions from consideration cannot resolve this issue.

\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{Extension: Inference on a Function of $w_0$} In many applications, the main parameter of interest is an identified function of $w_0$:

align*[align* omitted — 40 chars of source]

for some map $\theta_P: \Delta_{K-1} \to \Theta$ and $\Theta \subset \mathbf{R}^L$ is the parameter space for $\theta_0$.\footnote{Recently, Dufour/Tuvaandorj:25:arXiv developed a quasi-likelihood ratio test when the parameter of interest is a known function of a nuisance parameter that is on the boundary, and showed its pointwise asymptotic validity. Their test is simple to use, involving no tuning parameters or bootstrap for critical values. VanDijcke/Gunsilius/Wright:24:arXiv developed bootstrap inference for treatment effect parameters identified via the distributional synthetic control approach of Gunsilius:23:ECMA and proved its pointwise asymptotic validity. However, their proof requires Hadamard differentiability of the parameter functional at distributions outside a certain exceptional set, thereby precluding uniform asymptotic validity over any region containing this set.} In this setting, we consider two approaches for constructing a confidence set for $\theta_0$.

The first approach simply uses the Bonferroni method. More specifically, suppose that the estimator $\hat \theta$ satisfies the following asymptotic linear representation:

align[align omitted — 238 chars of source]

where $\Sigma_P(w_0)$ is a symmetric positive definite matrix, and $n$ denotes the size of the sample from the population that identifies $\theta_P(\cdot)$. Then, for any $\alpha \in (0,1)$, we can use the confidence set $C_{1-\kappa}$ for $w_0$ with level $100(1-\kappa)\%$, $\kappa \in (0,\alpha)$ (in Algorithm (ref)) and construct the $100(1-\alpha)\%$-level confidence interval for $\theta$ as

align[align omitted — 246 chars of source]

where $\hat \Sigma(w)$ is a consistent estimator of $\Sigma_P(w)$ and $q_{1-\alpha - \kappa}$ denotes the $(1-\alpha-\kappa)$ quantile of the $\chi^2$ distribution with degree of freedom $L$.

The second approach is to construct a joint confidence set for $[w_0',\theta_P(w_0)']'$ and project the confidence set on the space for $\theta_P(w_0)$.\footnote{We are grateful to a referee from our previous submission for suggesting this idea.} First, let us assume that for each $w \in \Delta_{K-1}$,

align[align omitted — 198 chars of source]

as $n \to \infty$, for a symmetric positive definite covariance matrix $\Omega_P(w)$, where $\hat \varphi(\cdot)$ is a $\sqrt{n}$-consistent estimator of $\varphi_P(\cdot)$. (Note that we allow for the sample size for $\hat \theta(\cdot)$ to be different from that for $\hat \varphi(\cdot)$.) We let

align*[align* omitted — 254 chars of source]

Then, using these and a consistent estimator $\hat \Omega(w)$ of $\Omega_P(w)$, we construct a confidence set $C_{1-\alpha}^{\mathsf{joint}}$ as in Algorithm (ref). The confidence set for $\theta$ can be obtained from projecting the joint confidence set onto $\Theta$:

align[align omitted — 189 chars of source]

While the projection method can be among various methods of subvector inference, we have opted for this method due to its simplicity, being free of any tuning parameters. We relegate the investigation of improvement on this method to future research.

algorithm[algorithm omitted — 1,697 chars of source]
example[Synthetic Control with Group-Level Weights, Cont'd] Let us revisit the synthetic control setting in Example (ref). For an estimator of $\theta_P(w)$, we consider $\hat \theta(w) = \hat \mu_{0,T+1} - \hat \mu_{T+1}' w.$ Recall that \begin{align*} \sqrt{n} B_2' (\hat \varphi(w) - \varphi_P(w)) = \frac{1}{\sqrt{n}} \sum_{i \in N} \frac{1}{T}\sum_{t=1}^T B_2'\mu_t \left( \sum_{j=1}^K \psi_{ij,t} w_j - \psi_{i0,t} \right) + o_P(1), \end{align*} where $\psi_{ij,t}$, $j=0,1,...,T$, are defined in Example (ref). Also, note that \begin{align*} \sqrt{n}(\hat \theta(w) - \theta_P(w)) = \frac{1}{\sqrt{n}} \sum_{i \in N} \left( \psi_{i0,T+1} - \sum_{j=1}^K \psi_{ij,T+1} w_j \right) + o_P(1), \end{align*} where $\hat \theta(w) = \hat \mu_{0,T+1} - \sum_{j=1}^K \hat \mu_{j,T+1} w_j.$ When we apply the Bonferroni method, we can construct \begin{align*} \hat \sigma^2(w) = \frac{1}{n} \sum_{i \in N} \left( \hat \psi_{i0,T+1} - \sum_{j=1}^K \hat \psi_{ij,T+1} w_j \right)^2, \end{align*} with \begin{align*} \hat \psi_{ij,T+1} = \frac{n}{n_{T+1}} \frac{1\{G_{i,T+1} = j\}(Y_{i,T+1} - \hat \mu_{j,T+1})}{\hat p_{j,T+1}}, \end{align*} where $\hat p_{j,T+1} = n_{j,T+1}/n_{T+1}$. Thus, we obtain \begin{align} \tilde C_{1-\alpha}^{\mathsf{Bonf}} \coloneqq \left\{\theta \in \Theta : \inf_{w \in C_{1-\kappa}} \frac{n (\hat \theta(w) - \theta)^2}{\hat \sigma^2(w)} \le q_{1 - \alpha + \kappa} \right\}, \end{align} with $q_{1 - \alpha + \kappa}$ defined as before, except with $L=1$. As for the projection method, we construct \begin{align} \hat \Omega(w) &= \begin{bmatrix} \hat \Omega_{11}(w) & \hat \Omega_{12}(w) \\ \hat \Omega_{21}(w) & \hat \Omega_{22}(w) \end{bmatrix} , \end{align} where \begin{align*} \hat \Omega_{11}(w) &= \frac{1}{n} \sum_{i \in N} B_2' \left( \frac{1}{T}\sum_{t=1}^T \hat \mu_t \left(\sum_{j=1}^K \hat \psi_{ij,t} w_j - \hat \psi_{i0,t}\right) \right) \left( \frac{1}{T}\sum_{t=1}^T \hat \mu_t \left(\sum_{j=1}^K \hat \psi_{ij,t} w_j - \hat \psi_{i0,t}\right) \right)' B_2\\ \hat \Omega_{12}(w) &= \frac{1}{n} \sum_{i \in N} B_2' \left( \frac{1}{T}\sum_{t=1}^T \hat \mu_t \left(\sum_{j=1}^K \hat \psi_{ij,t} w_j - \hat \psi_{i0,t}\right) \right) \left( \hat \psi_{i0,T+1} - \sum_{j=1}^K \hat \psi_{ij,T+1} w_j \right), and \\ \hat \Omega_{22}(w) &= \frac{1}{n} \sum_{i \in N} \left(\hat \psi_{i0,T+1} - \sum_{j=1}^K \hat \psi_{ij,T+1} w_j \right)^2. \end{align*} When the sample is repeated cross-sections, we can utilize independence across time and obtain sharper joint inference for $(w_0,\theta_0)$, by setting \begin{align*} \hat \Omega(w) = \begin{bmatrix} \hat \Omega_{11}(w) & 0\\ 0 & \hat \Omega_{22}(w) \end{bmatrix}, \end{align*} where \begin{align*} \hat \Omega_{11}(w) = \frac{1}{n} \sum_{i \in N} \frac{1}{T^2}\sum_{t=1}^T \left(\sum_{j=1}^K w_j \hat \psi_{ij,t} - \hat \psi_{i0,t} \right)^2 B_2' \hat \mu_t \hat \mu_t' B_2, \end{align*} and $\hat \Omega_{22}(w)$ is the same as defined above. Note that $\Omega_{12}(w)$ is zero, because it consists of covariances between quantities involving observations at $t=1,...,T$ and those involving observations at $t = T+1$. Now, using $\hat \Omega(w)$, we obtain the joint confidence set $C_{1-\alpha}^{\mathsf{joint}}$ as in Algorithm (ref). From this, we obtain the confidence set for $\theta_0$ as in ((ref)). $\blacksquare$
example[Distributional Synthetic Control, Cont'd] Let us revisit Example (ref). In this setting, we have \begin{align*} \sqrt{n} B_2' (\hat \varphi(w) - \varphi_P(w)) = \frac{1}{\sqrt{n}} \sum_{i \in N} B_2'(\Psi_{i,H} w - \psi_{i,h}) + o_P(1), \end{align*} where $\Psi_{i,H}$ and $\psi_{i,h}$ are defined in Example (ref). Define \begin{align*} \hat \theta(w) = \hat \rho_{0,2}(\tau^*) - \sum_{j=1}^K \hat \rho_{j,2}(\tau^*) w_j. \end{align*} As for the estimated quantile treatment effect, \begin{align*} \sqrt{n}(\hat \theta(w) - \theta_P(w)) = \frac{1}{\sqrt{n}} \sum_{i \in N} \left( \psi_{i0,2}(\tau^*) - \sum_{j=1}^K \psi_{ij,2}(\tau^*) w_j \right) + o_P(1), \end{align*} where $\psi_{ij,t}(\tau)$ is defined in ((ref)). When we apply the Bonferroni method, we can construct \begin{align*} \hat \sigma^2(w) = \frac{1}{n} \sum_{i \in N} \left( \hat \psi_{i0,2}(\tau^*) - \sum_{j=1}^K \hat \psi_{ij,2}(\tau^*) w_j \right)^2, \end{align*} with \begin{align*} \hat \psi_{ij,2}(\tau^*) = \frac{n}{n_{j,2}} \frac{1\{i \in N_{j,2}\}(\tau^* - 1\{Y_{i,2} \le \hat \rho_{j,2}(\tau^*)\})}{\hat f_{j,2}(\hat \rho_{j,2}(\tau^*))}, \end{align*} where $\hat f_{j,2}(t)$ is a consistent and asymptotically normal estimator of $f_{j,2}(t)$. Using $\hat \theta(w)$ and $\hat \sigma^2(w)$, we can construct the confidence set using the Bonferroni method as before. As for the projection method, we construct \begin{align} \hat \Omega(w) &= \begin{bmatrix} \hat \Omega_{11}(w) & \hat \Omega_{12}(w) \\ \Omega_{21}(w) & \Omega_{22}(w) \end{bmatrix} , \end{align} where \begin{align*} \hat \Omega_{11}(w) &= \frac{1}{n} \sum_{i \in N} B_2' (\hat \Psi_{i,H} w - \hat \psi_{i,h})(\hat \Psi_{i,H} w - \hat \psi_{i,h})' B_2\\ \hat \Omega_{12}(w) &= \frac{1}{n} \sum_{i \in N} B_2' (\hat \Psi_{i,H} w - \hat \psi_{i,h})\left( \hat \psi_{i0,2}(\tau^*) - \sum_{j=1}^K \hat \psi_{ij,2}(\tau^*) w_j \right), and \\ \hat \Omega_{22}(w) &= \frac{1}{n} \sum_{i \in N} \left( \hat \psi_{i0,2}(\tau^*) - \sum_{j=1}^K \hat \psi_{ij,2}(\tau^*) w_j \right)^2, \end{align*} and $\hat \Psi_{i,H}$ and $\hat \psi_{i,h}$ are estimated versions of $\Psi_{i,H}$ and $\psi_{i,h}$ with $\psi_{ij,t}$ replaced by $\hat \psi_{ij,t}$. Using $\hat \Omega(w)$, we obtain the joint confidence set $C_{1-\alpha}^{\mathsf{joint}}$ as in Algorithm (ref).

\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{Uniform Asymptotic Validity}

\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{A Preliminary Result}

We first provide a preliminary result. Suppose that $Y_n \in \mathbf{R}^{K+L-1}$, $n \ge 1$, is a sequence of random vectors such that

align[align omitted — 80 chars of source]

where $Z_n \in \mathbf{R}^{K+L-1}$ is a sequence of asymptotically standard normal random vectors, $\hat \Omega$ is a sequence of symmetric positive definite $(K+L-1) \times (K+L-1)$ random matrices, and $\mu_n \in \mathbf{R}^{K+L-1}$ is a sequence of nonstochastic vectors. We do not assume that the sequence $\mu_n$ is bounded.

Our main focus is on the test statistic that is based on the squared error from the projection of $Y_n$ onto a polyhedral cone along the norm $\|\cdot \|_{\hat \Omega}$, where $\| x \|_{\hat \Omega}^2 = x' \hat \Omega^{-1} x$, $x \in \mathbf{R}^{K+L-1}$, and $\Lambda(w)$ is defined in ((ref)). Recall that $G(\cdot;k)$ denotes the CDF of the $\chi^2$ distribution with the degrees of freedom equal to $k$ and $J_0[z] = \{1\le j \le K: z_j = 0\}$ for any vector $z = [z_j] \in \mathbf{R}^K$. For each $w \in \Delta_{K-1}$, we define

align[align omitted — 312 chars of source]

The test statistic takes the form of the squared residuals from projecting $Y_n$ onto $\tilde \Lambda(w)$:

align*[align* omitted — 99 chars of source]

where $\Pi_{\hat \Omega}(Y_n\mid \tilde \Lambda(w))$ denotes the projection of $Y_n$ onto $\tilde \Lambda(w)$ along $\| \cdot \|_{\hat \Omega}$.

As for the critical value, let $\hat \delta(w) \coloneqq B \hat \Omega^{-1} (Y_n - \Pi_{\hat \Omega}(Y_n \mid \tilde \Lambda(w)))$ and

align*[align* omitted — 99 chars of source]

with $B$ defined in ((ref)) and $0_{K \times L}$ denoting the $K \times L$ dimensional matrix of zeros. Then, the construction of the confidence set for $w$ is based on the critical value of the following form:

align*[align* omitted — 83 chars of source]

This critical value is random, depending on the data through its dependence on the number of zeros in $\hat \delta(w)$.

We first introduce an assumption that requires the consistency of $\hat \Omega$ and the asymptotic normality of $Z_n$ along a sequence of probabilities $P_n$.

assumptionFor each $t \in \mathbf{R}$ and $\epsilon>0$, \begin{align*} P_n\left\{Z_n \le t \right\} \to \Phi(t) and P_n\left\{ \left\|\hat \Omega - \Omega_n \right\| > \epsilon \right\} \rightarrow 0, \end{align*} for some sequence of nonstochastic $(K+L-1) \times (K+L-1)$ symmetric positive definite matrices $\Omega_n$ such that for some $\overline B, \epsilon >0$, $\| \Omega_n\| \le \overline B$ and $\lambda_{\min}(\Omega_n) \ge \epsilon$, for all $n \ge 1$, where $\lambda_{\min}(\Omega_n)$ denotes the minimum eigenvalue of $\Omega_n$ and $\Phi$ the CDF of $N(0,I)$.

The asymptotic normality of $Z_n$ and the consistency of $\hat \Omega$ are typically satisfied in the applications under consideration in this paper. The following lemma is crucial for the uniform asymptotic validity result.

lemmaSuppose that Assumption (ref) holds. Then, for any $\alpha \in (0,1)$, and for any sequence $w_n \in \Delta_{K-1}$ and $\mu_n \in \tilde \Lambda(w_n)$, we have \begin{align*} \limsup_{n \rightarrow \infty} P_n\left\{\left\| Y_n - \Pi_{\hat \Omega}(Y_n\mid \tilde \Lambda(w_n)) \right\|_{\hat \Omega}^2 > \hat c_{1-\alpha}(Y_n;w_n)\right\} \le \alpha. \end{align*}

This lemma shows that reading the critical values from the $\chi^2$ distribution with the degrees of freedom equal to $\hat k(w)$ generates a uniformly asymptotically valid procedure. The proof of the lemma can be potentially useful for many similar settings where the test statistic involves a projection onto a polyhedral cone.

\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{The Main Result}

Let us revisit Section (ref) and present the main result of the asymptotic validity of the confidence set for $(w_0,\theta_0)$. We focus on the uniform asymptotic validity of the joint confidence set $C_{1-\alpha}^{\mathsf{joint}}$ in Algorithm (ref), as it is more general than $C_{1-\alpha}$ in Algorithm (ref). First, we introduce a set of assumptions. From here on, $\mathcal{P}_n$ denotes the set of probability distributions of data that are under consideration. Our goal is to establish the uniform asymptotic validity of the confidence set $C_{1-\alpha}^{\mathsf{joint}}$ for $w_0$ over the class of probability distributions $\mathcal{P}_n$.

assumptionFor each sequence of probabilities $P_n \in \mathcal{P}_n$ and each $w \in \Delta_{K-1}$, there exists a sequence of symmetric positive definite matrices, $\Omega_{P_n}(w)$, such that the following conditions are satisfied for any sequence of weights $w_n \in \mathbb{W}_{P_n}$. (i) For each $t \in \mathbf{R}^{K + L -1}$, we have \begin{align*} P_n\left\{ \Omega_{P_n}^{-1/2}(w_n) \sqrt{n} \begin{bmatrix} B_2'(\hat \varphi(w_n) - \varphi_{P_n}(w_n)) \\ \hat \theta(w_n) - \theta_{P_n}(w_n) \end{bmatrix} \le t \right\} \rightarrow \Phi(t), \end{align*} as $n \rightarrow \infty$, where $\Phi$ is the CDF of the standard normal random vector. (ii) $\limsup_{n \rightarrow \infty} \left\|\Omega_{P_n}(w_n) \right\| < \infty$ and $\liminf_{n \rightarrow \infty}\lambda_{\min}(\Omega_{P_n}(w_n)) > 0$. (iii) For any $\epsilon>0$, and along any sequence $w_n \in \Delta_{K-1}$, \begin{align*} P_{n}\left\{ \left\|\hat \Omega(w_n) - \Omega_{P_n}(w_n) \right\| > \epsilon \right\} \rightarrow 0. \end{align*}

Condition (i) requires the asymptotic normality of the estimators that is uniform over the probabilities in $\mathcal{P}_n$. This condition is satisfied in many settings where maps, $B_2' \varphi_P(w)$, $\theta_P(w)$, are “regular”, i.e., behave “smoothly” in the perturbation of $P$. Condition (ii) requires that the limit distribution is non-degenerate uniformly over the probabilities in $\mathcal{P}_n$. Condition (iii) requires the variance estimator $\hat \Omega(w_n)$ to be consistent for $\Omega_P(w_n)$ uniformly over the probabilities in $\mathcal{P}_n$. Condition (iv) requires that the minimizer of $Q_{P_n}(w)$ over $w \in \Delta_{K-1}$ is close to the global minimizer. This condition is easily satisfied with the synthetic control examples with the pretreatment fit. In these cases, we have $\varphi_{P_n}(w) = 0$ for all $w \in \mathbb{W}_{P_n}$. The constant $C$ is permitted to be unknown, and hence the condition (iv) is too weak to apply the approach of quadratic approximation.

Thus, the conditions in Assumption (ref) essentially define the scope of the method and are satisfied in many settings. By applying Lemma (ref), we obtain the following result.

theoremSuppose that Assumption (ref) holds. Then, for any $\alpha \in (0,1)$, \begin{align*} \liminf_{n \rightarrow \infty} \inf_{P \in \mathcal{P}_n} \inf_{w \in \mathbb{W}_P} P\left\{ (w,\theta_{P}(w)) \in C_{1-\alpha}^{\mathsf{joint}}\right\} \ge 1 - \alpha. \end{align*}

The proof is found in the appendix of this paper. To see how Lemma (ref) gives this validity result, observe that

align*[align* omitted — 139 chars of source]

where

align[align omitted — 152 chars of source]

Thus, the test statistic $T(w,\theta)$ is nothing but the squared length of the projection error in projecting $Y_n(w)$ onto the polyhedral cone $\Lambda(w)$ along the norm $\| \cdot \|_{\hat \Omega(w)}$. Since

align*[align* omitted — 79 chars of source]

with

align[align omitted — 311 chars of source]

the random vector $Y_n(w,\theta_{P_n}(w))$ is asymptotically normal, yet with a potentially diverging shift $\mu_n(w)$. This setting maps to the one in Lemma (ref).\footnote{Note that we allow for the possibility that $\liminf_{n \rightarrow \infty}\varphi_{P_n}(w_n) \ne 0$ for $w_n \in \mathbb{W}_{P_n}$ such as when the minimizer of $Q_{P_n}(w)$ over $w \in \mathbf{R}^K$ lies outside $\Delta_{K-1}$. In this case, the sequence $\|\mu_n(w_n)\|$ may diverge to $\infty$, as $n \rightarrow \infty$.} The uniform asymptotic validity of the confidence set $C_{1-\alpha}$ for $w_0$ follows from the lemma.

example[Synthetic Control with Group-Level Weights, Cont'd] Let us consider low level conditions for Assumption (ref) for the synthetic control setting in Example (ref). It is not hard to see that Assumption (ref) is satisfied if the following conditions are satisfied. (i) For each $t=1,...,T$, the random vectors, $[Y_{i,t},G_{i,t}]$, $i \in N_t$, are i.i.d., and \begin{align*} \limsup_{n \ge 1} \sup_{P \in \mathcal{P}_n} \max_{0 \le j \le K} \mathbf{E}_P\left[ Y_{i,t}^4 \mid G_{i,t} = j \right] < \infty and \liminf_{n \ge 1} \inf_{P \in \mathcal{P}_n} \min_{0 \le j \le K} P\left\{G_{i,t} = j \right\} > 0. \end{align*} (ii) $\liminf_{n \ge 1} \inf_{P \in \mathcal{P}_n} \inf_{w \in \Delta_{K-1}} \lambda_{\min}\left(\Omega_P(w)\right)>0$, where $\Omega_P(w)$ denotes the population version of $\hat \Omega(w)$ defined in ((ref)). Assumption (ref)(i) is satisfied by the Central Limit Theorem and Law of Large Numbers. Note that Assumption (ref)(ii) is violated if $Y_{i,t}$ is identically distributed across time $t$. For example, in the case of a linear factor model for $Y_{i,t}$ with a short panel, the probabilities here need to be replaced by the conditional probabilities given the factors. Hence, $Y_{i,t}$ is identically distributed across time $t$ only if the factors are time-invariant, which is unlikely in practice. $\blacksquare$

\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{Monte Carlo Simulations}

We now turn to Monte Carlo simulations to evaluate the finite sample performance of our procedure. Our set-up follows Example (ref) and the literature on synthetic control with individual data, which are of significant applied interest (see Abadie/LHour:21:JASA). This will also be the focus of our empirical application below.

We consider a treated group (0) and $K$ untreated groups, where $K \in \{3,5,7\}$. Each group $j$ has a sample size $n_j \in \{100, 200, 1000\}$. The outcomes for individual $i$ in group $j$, $Y_{ij,t}$, are realized over $T$ periods, but only the first $T_0 = 10$ periods are used to match the treated and untreated groups.

We generate the pre-treatment outcomes for the untreated groups $j=1,...,K$ using the linear model:

eqnarray*[eqnarray* omitted — 59 chars of source]

where $\mu_{j,t} = 0.5 + 0.5(-1)^{j-1}\frac{t}{T} + 0.5 \eta_{j,t}$, $\eta_{j,t} \sim_{i.i.d.} N(0,1)$ and $\varepsilon_{ij,t} \sim_{i.i.d.} N(0,1)$. Thus, the mean outcome for $j$, $\mu_{j,t}$ is composed by a deterministic component and a random component. The former grows over time for odd-numbered $j$, while it decreases over time for even-numbered groups. The random component is fixed throughout simulations. Individual outcomes in a group are noisy observations of the associated means.

We draw data for the treated group as:

eqnarray*[eqnarray* omitted — 59 chars of source]

where $\mu_{0,t} = w_{0}' \mu_{t}$, $\mu_t = (\mu_{1,t},...,\mu_{K,t})'$, $w_{0} \in \Delta_{K-1}$, and $\varepsilon_{i0,t} \sim_{iid} N(0,1)$ across $i$ and $t$. Thus, the expected outcomes for the treated group are a weighted (synthetic) combination of expected outcomes from the untreated groups with weights, $w_0$.\footnote{While this may resemble the assumption in the synthetic control literature that the outcomes of the treated group have a perfect match in the untreated ones, there is a major distinction: we impose the matching on the population quantities rather than on their sample counterparts. It is important for us to distinguish between population and sample. Otherwise, there is no inference to be made on $w_0$. For the same reason, the observed outcomes are noisy observations of the population means in our setting.} We consider the setting where $T > K$ and there is no linear dependency between the population means $\mu_{j,t}$, $j=0,...,K$. Hence, the weights $w_0$ are point-identified.\footnote{As for our grid for $w$, we use two complementary approaches that allow us to extensively explore both the interior and the boundary of the simplex. The first draws a fine grid of $w$ uniformly over its simplex using a procedure based on Rubin:81:AoS. This is done by first drawing a vector of dimension $K-1$, with each element i.i.d. from the uniform distribution with support $[0,1]$. Then, we include 0 and 1 into that drawn vector, sort it, and produce the vector of differences across adjacent elements of $w$. These are all nonnegative and sum up to 1 by construction. As for the boundary points, we first define a step size, generate a vector of $[0,1]$ with points separated by that step-size, and then generate a $K$-dimensional meshed grid of all possible combinations of those values. We only keep the points that (i) add up to 1, (ii) have a zero element. The final grid is the union of the gridpoints generated by both approaches.}

The weight $w_0$ is the key object in inference. We consider two main specifications, each one presenting a different case of whether $w_0$ is interior or on the boundary of the simplex. In the former (Specification 1), we set $w_0 = (0.2, 0.8 \mathbf{1}_{K-1}'/(K-1))'$, so that $w_0$ lies in the interior of $\Delta_{K-1}$, while, in the latter (Specification 2), $w_0 = (0.5, 0.5, \mathbf{0}_{K-2}')'$, so that $w_0$ lies on the boundary of $\Delta_{K-1}$, where $\mathbf{1}_{d}$ and $\mathbf{0}_{d}$ are $d$-dimensional vectors of 1's and 0's, respectively. In each of $1,000$ simulations, we independently draw $(\varepsilon_{i0,t}, \varepsilon_{i1,t},...,\varepsilon_{K,t})'$ for $t=1,...,T_0$ and $i=1,...,n_j$, while keeping $\mu_{j,t}$ fixed. Since the outcomes $Y_{ij,t}$'s are independent across time, the data generating process represents a repeated cross-section design. All results are shown in Table (ref).

table[table omitted — 1,031 chars of source]

Table (ref) shows that our procedure works very well in finite samples across scenarios. Let us start with the interior weights case (Specification 1). In line with the theory, coverage probabilities are very close to nominal levels across specifications. This holds even for the smallest sample sizes ($n_j=100, K = 3$ and $n_j = 100, K = 5$), with empirical coverage probabilities between 93.6-96.2%. Coverage probabilities become very close to 95% even when $n_j = 200$, across specifications (94.7, 94.6 and 94.1% for $K \in \{3,5,7\}$, respectively). Similarly, the results continue to hold as $n_j$ or $K$ grows because our asymptotic results are based on the total sample size (numbers of individuals times number of groups). This is seen both with $K=5$ and $K=7$, with coverage probabilities at or close to 95% for $n_j = 1,000$.

The results from Specification 2 show that our procedure also works well even when $w_0$ is on the boundary of the simplex. Indeed, coverage probabilities are close to the nominal 95% across specifications and sample sizes. This holds both for small sample sizes (96.8% and 94.1% for the smallest values of $K$ and $n_j$), but also as both grow. Indeed, they become very close to 95% for large sample sizes: with $n_j = 200$, coverage probabilities are 94.9% for $K=3$, 94.6% when $K=5$ and 93.9% for $K=7$. Similar results hold when $n_j = 1,000$.

Sometimes researchers are interested in inference on an individual component of $w_0$ rather than inference on the full vector $w_0$. For example, one may be interested in the weight that the synthetic method gives to a specific untreated group (e.g., state, country, firm, or forecaster in the case of predictions). To that end, one can combine our approach with projection-based subvector inference. Table (ref) presents the results of this approach for the first dimension of $w_0$ in the specifications of Table (ref). We report both the empirical coverage probability for $w_{0,1}$ and the mean length of the confidence interval. This also provides information on the power of our inference approach.

table[table omitted — 1,654 chars of source]

The projection-based inference works well in finite samples. As expected, it is more conservative than inference on the full $w_0$ (c.f. Table (ref)). For $K=3, 5$ and interior $w_0$, the former is typically 2-4 percentage points more conservative than the latter, although such conservativeness decreases with sample size (e.g., it is below 1 percentage point when $n_j = 1,000$ and $K=5$ or $K=7$, regardless of specification). In spite of its conservative nature, projection-based inference is still informative in our setting. Even when sample size is small ($n_j = 100, K=3$), the average length of the confidence interval can be as short as 0.14-0.16 and it is far from spanning $[0,1]$. Furthermore, for a fixed $K$ and weight $w_0$, as sample size $n$ increases, the length of the confidence interval quickly decreases and gets even close to 0: when $n_j = 1,000$, the average length is below 0.04 across specifications. The latter would be very sharp for inference with a sample size often found in empirical applications with individual-level data. Thus, our simulations suggest that, when $w_0$ is point-identified and the sample size is large enough, projection-based inference for the weights can be informative and not overly conservative.

\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{Empirical Application}

As an empirical illustration of our method, we study how a large increase in minimum wages in Alaska in 2003, from US\$5.65 to US\$7.15, affected average family income. To do so, we use individual-level data and state-level weights, as in Example (ref). This follows a recent strand of the literature that uses synthetic control methods to obtain credible estimates of the effects of minimum wage increases on economic outcomes (e.g., Allegretto/Dube/Reich/Zipperer:17:ILR, Neumark/Wascher:17:ILR, Powell:22:JBES) and the analysis of this policy in Gunsilius:23:ECMA.

Following Dube:19:AEJ and Gunsilius:23:ECMA, our outcome of interest is a measure of household income, called average family income. This is a measure which adjusts for family size and composition (i.e., “equivalized”) and it is measured in multiples of the federal poverty threshold. We use the main sample from Dube:19:AEJ (i.e., individuals aged under 65), originally drawn from the Current Population Survey (CPS), and focus on the subsample from 1998-2002 (i.e., the five years before the policy).

To form a synthetic control, we form a pool of states that resemble Alaska in terms of population and economic activity. We start with states that, like Alaska, are among the top-10 oil producing states in 2002 (i.e., prior to the policy), e.g., see Figure 7 in Ismayilova:07:Thesis. We then drop states that have a minimum wage increase during the period of interest (California) and states that have a population over 3 million (Texas, Louisiana and Oklahoma). By comparison, Alaska had a population of approximately 650,000 in 2002, Census2002. Our final set of control states is, thus: Kansas, Mississippi, New Mexico, North Dakota and Wyoming ($K=5$). From this pool, we construct a synthetic comparison state by deriving state-level weights from estimating $\mu_{j,t}$: the mean family income in each state $j$ and year $t$.

We use our statistical procedure outlined in Section (ref) and report the projection-based inference for weights and Bonferroni-based confidence interval for a parameter of interest, $\theta_0$, defined as the difference in mean outcomes in Alaska in 2003 (the year post-policy) relative to the mean in the synthetic control. This can be written as:

eqnarray[eqnarray omitted — 95 chars of source]

where $\mu_{\text{ctrl},2003} = [\mu_{1,2003}, \mu_{2,2003},...,\mu_{K,2003}]'$ is a vector of the mean average family income across families in the control states and $\theta_0 = \theta(w_0)$. We can compute the asymptotic variance $\Omega(w)$ as explained in Example (ref).

table[table omitted — 1,941 chars of source]

The first row of Table (ref) presents the confidence interval for the effect of the increase in minimum wages on average family income in 2003 across specifications. The results are robust across specifications and inference methods, with confidence intervals centered around -0.2. From our results, we can reject the presence of any meaningful positive effects, as the upper bound of our confidence intervals are all very close or at 0 (and below a 1% effect). This suggests that this minimum wage policy, at best, did not affect average family income. At worst, it decreased average family income by up to 13.4% relative to a pre-policy (1998-2002) average income of 3.43 times the federal poverty threshold. While it is likely that the poorest households benefit from this policy, the average family may have lost a small amount of income, possibly due to increased unemployment. Our confidence intervals account for inference on the weights, which contrasts with the standard approach in synthetic control, where the synthetic weights are assumed to hold exactly for any realization of the sample. While the results are robust across specifications (i.e., with or without Mississippi) and methods (Bonferroni or projection-based), we do find that the projection method presented in Algorithm (ref) seems to outperform the Bonferroni counterpart in this setting. Indeed, the average length of the projection-based confidence interval for $\theta_0$ is 16% (0.4/0.48) shorter than the Bonferroni counterpart, and the former is a subset of the latter. This empirical gain comes at an increased computational cost.

Table (ref) further reports projection-based confidence intervals for weights across two choices of control units: with and without Mississippi. Kansas often received the largest weights (i.e., with confidence intervals centered around 0.8-0.85), followed by Wyoming, while the others have relatively similar confidence intervals centered around 0-0.05. Meanwhile, the average length of the confidence intervals is around 0.15. This provides further support to the findings in the simulations that projection-based inference for $w$ can be informative in this context.

\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{Conclusion}

In this paper, we propose a new method for inference on the simplex-constrained weights. Our method is based on the projection of an asymptotically normal statistic onto a polyhedral cone. We show that the critical value for the test statistic can be read from the $\chi^2$ distribution with data-dependent degrees of freedom. This leads to an asymptotically uniformly valid confidence set for the weights. We provide a proof of this result and show that it is applicable to a wide range of settings. We also show that the method works well in finite samples through Monte Carlo simulations. We apply our method to a synthetic control example and show that it can be useful in practice.