Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,778 characters · 0 sections · 65 citation commands
\address{Department of Economics, University of Warwick, Coventry, CV4 7AL, United Kingdom} \email{[email removed]} \address{Vancouver School of Economics, University of British Columbia, 6000 Iona Drive, Vancouver, BC, V6T 1L4, Canada} \email{[email removed]}
\fontsize{12}{14} \selectfont
\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{Introduction}
We consider the following problem. Suppose that the population weight $w_0 \in \Delta_{K-1}$ in the $(K-1)$-simplex $\Delta_{K-1}$ is defined as a solution to the following optimization problem:
where $Q(w)$ is the population objective function that is convex and differentiable in $w \in \mathbf{R}^{K}$. The dimension $K$ is fixed and not allowed to depend on the sample size $n$. Our focus is on constructing a $(1-\alpha)$-level confidence set $C_{1-\alpha}$ for $w_0$ that is uniformly asymptotically valid.
This inference problem arises in many contexts of applications. For example, in the causal inference literature, various methods of synthetic control design choose a weight as a solution to ((ref)) (see Abadie:21:JEL for a survey of this literature). Meanwhile, in the literature of forecast combinations, the final forecast is constructed as a weighted average of forecasts from different methods or experts. (See Timmermann:06:Handbook.) While not in consensus, the use of a simplex-constrained weight has been part of the main approaches considered in this literature. The least squares model averaging proposed by Hansen:07:Eca also falls into this framework. In particular, Hansen:07:Eca proposed minimizing the Mallows metric to obtain the optimal weight. The population version of the minimization problem can be written as the optimization problem in ((ref)).
The main challenge in the theory of statistical inference on the weight is that the parameter is potentially on the boundary, as $w_0$ can fall on an edge or a vertex of the simplex $\Delta_{K-1}$. Hence, even if $w_0$ is point-identified and $\sqrt{n}$-consistently estimable, it is far from obvious how to construct a pivotal test statistic whose limiting distribution is invariant uniformly over the probabilities in the model.
Despite the challenge, there are methods which can be applied to develop an asymptotic inference procedure. For example, when $w_0$ satisfies the first order condition for the optimization problem ((ref)) - an assumption this paper does not assume, we may build a quadratic approximation of the objective function to derive the limiting distribution that depends on the localization parameter (Geyer:94:AS, Andrews:99:Ecma and Moon/Schorfheide:09:JOE), or apply the conditional likelihood ratio approach by Ketz:18:JOE. Alternatively, we could adapt the inference based on the random draws from a quasi-posterior in Chen/Christensen/Tamer:18:Eca to this setting or employ the bootstrap-based projection method of Fang/Seo:21:Eca. These methods target a substantially more general setting than simplex-valued weights. However, they require bootstrap or simulations to obtain critical values and often involve tuning parameters that require a delicate choice to ensure a good finite sample performance.
In this paper, we propose a simple method of constructing a confidence set for the simplex-valued weight which is free of tuning parameters and does not require simulations to compute critical values. Furthermore, our method does not require point-identification of the weight $w_0$ and, hence, does not rely on the quadratic approximation of the objective function. The method is shown to produce confidence sets that are asymptotically valid uniformly over the behavior of the population weight. We provide simulation results and an empirical application to demonstrate the merits of our method.
Our method relies on a likelihood ratio-type test statistic constrained to a polyhedral cone. When the test statistic is formed from a multivariate normal random vector, its distribution is known to follow a mixture of a $\chi^2$ distribution with generally unknown weights (see Silvapulle/Sen:05:ConstrainedStatInference and references therein). In this setting, AlMohamad/vanZwet/Cator/Goeman:20:Biometrika developed an adaptive inference method that selects the relevant $\chi^2$ distribution automatically. In econometrics, Breunig/Chen:20:WP, Breunig/Chen:24:Eca proposed a related adaptive approach for testing equality or inequality restrictions on nonparametric functions identified in a nonparametric instrumental variable model. A related idea is also found in Cox/Shi:23:ReStud who considered a general moment inequality model with an additive nuisance parameter and proposed asymptotic inference methods. These approaches simplify the inference procedure for testing problems with constraints by using $\chi^2$ distributions with data-dependent degrees of freedom. Our paper builds on this line of research by developing a simple adaptive procedure for asymptotic inference on simplex-valued weights. However, both the specific design of our proposal and the proof of its uniform validity are new, to the best of our knowledge.\footnote{Related to this literature, recently, Li:25:WP developed an interesting method of bootstrap-based inference on parameters identified as a constrained optimizer. Li:25:WP does not require quadratic approximation of the objective function, yet involves a sequence of tuning parameters that converge to zero. In contrast to Li:25:WP, our paper focuses on the inference on a simplex-valued weight, and as such, our proposal is tailored to this case, and does not require point-identification of the weight or tuning parameters that go to zero at a certain rate.}
The rest of the paper proceeds as follows. In Section 2, we introduce a basic set-up with examples, and present our main proposal to construct a confidence set for a simplex-valued weight. In Section 3, we provide our main result of uniform asymptotic validity of the confidence set, and Monte Carlo simulations that study the finite sample properties of our statistical procedure. An empirical application is presented in Section 4 as an illustration. In Section 5, we conclude. Mathematical proofs of the results in this paper are found in the appendix.
\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{A Simple Confidence Set for a Simplex-Valued Weight}
\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{A Simplex-Valued Weight}
Let us formally introduce the problem. Suppose that $H_P$ is a $K \times K$ matrix and $h_P$ is a $K$-dimensional vector which potentially depends on the population distribution $P$. Let $\Delta_{K-1}$ denote the simplex in $\mathbf{R}^K$, i.e., $\Delta_{K-1}\coloneqq \{w \in \mathbf{R}^K: w \ge 0, w'\mathbf{1} = 1\}$, where $\mathbf{1}$ is the $K \times 1$ vector of ones. (Here, the inequality between two vectors should be understood as the set of pointwise inequalities between the corresponding entries.)
Given a map $Q_P: \mathbf{R}^K \rightarrow \mathbf{R}$, we define the argmin set of weights, $\mathbb{W}_P$, as follows:
We assume that $Q_P$ is convex and differentiable on $\mathbf{R}^K$ and there exists a true weight $w_0$ in the set $\mathbb{W}_P$.
\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{A Simple Confidence Set}
Suppose that the true weight $w_0$ belongs to $\mathbb{W}_P$. Our goal is to construct a confidence set for $w_0$ that is asymptotically uniformly valid over the class of population distributions $P$ under consideration. We will give a set of conditions later. For this section, we present the procedure of constructing the confidence set.
First, we consider the $K \times K$ orthogonal matrix $[\mathbf{1}/\sqrt{K},B_2]$, where $B_2$ is the $K \times (K-1)$ matrix consisting of orthonormal column vectors that are orthogonal to $\mathbf{1}$.\footnote{The computation of $B_2$ is straightforward. First note that $B_2'B_2 = I$. Hence, $B_2 B_2' = I - \mathbf{1} \mathbf{1}'/K$, i.e., the projection matrix projecting onto the orthogonal complement of the span of $\mathbf{1}$. To obtain $B_2$, we obtain a spectral decomposition : $ I - \mathbf{1} \mathbf{1}'/K = U D U'$. From this, we set $B_2$ to be the $K \times (K-1)$ matrix after removing the eigenvector from $U$ that corresponds to the zero diagonal element of $D$.} Define
and take $\hat \varphi(w)$ as an estimator of $\varphi_P(w)$ such that
as $n \rightarrow \infty$, for some positive definite matrix $V_P(w)$. Suppose that we have a consistent estimator of $V_P(w)$, denoted by $\hat V(w)$. Using $\hat \varphi$ and $\hat V$ only, we can construct a confidence set for $w_0$ as explained in Algorithm (ref).\footnote{The algorithm applies to cases where $w_0$ is partially identified. When $w_0$ is point-identified and has a consistent estimator $\hat w$, we can substitute $\hat w$ into $\hat V(w)$, replacing $\hat V(w)$ with $\hat V(\hat w)$. This reduces the computation of $\hat \lambda(w)$ to a quadratic programming problem, significantly improving computational speed.} The construction of the confidence set $C_{1-\alpha}$ is simple, involving no simulations or a sequence of tuning parameters that require a judicious choice in finite samples.
The computation of $V_P(w)$ and its consistent estimation can be done in a standard manner. For example, consider the case where
Then, we have
As for $\hat H$ and $\hat h$, suppose that these estimators admit an asymptotic linear representation as follows:
where $\Psi_{i,H}$ and $\psi_{i,h}$ are influence functions that are i.i.d.\ across $i$'s and have mean zero. Then, we can find that
where $\xi_i(w) = B_2'(\Psi_{i,H} w - \psi_{i,h})$. A consistent estimator $\hat V(w)$ of $V_P(w)$ is obtained as follows:
where $\hat \xi_i(w) = B_2'(\hat \Psi_{i,H} w - \hat \psi_{i,h})$, and $\hat \Psi_{i,H}$ and $\hat \psi_{i,h}$ are appropriate estimators of $\Psi_{i,H}$ and $\psi_{i,h}$. Using $\hat V(w)$, we construct the confidence set $C_{1-\alpha}$ is constructed as explained in Algorithm (ref).
In some applications, the computation and estimation of $V_P(w)$ can be cumbersome for practitioners, because the practitioner has to find an analytical form of $V_P(w)$ in their application to find its consistent estimator. In such cases, an alternative is to use a bootstrap estimator of $V_P(w)$. As shown by Hahn/Liao:21:Ecma, the asymptotic validity of the inference procedure is preserved under the bootstrap, while the inference procedure is conservative in general.\footnote{In some settings, we can modify the bootstrap variance estimator using a truncation method proposed by Shao:92:Stat to render the inference procedure asymptotically non-conservative. See also Goncalves/White:05:JASA.} Suppose that $\hat \varphi^*(w)$ is constructed from the bootstrap sample, such that the bootstrap distribution of $\sqrt{n}B_2'(\hat \varphi^*(w) - \hat \varphi(w))$ converges almost surely to $N(0,V_P(w))$. Suppose that $w_0$ is point-identified and the minimizer $\hat w$ of $\hat Q(w)$ over $w \in \Delta_{K-1}$ is consistent. Then, we obtain the bootstrap variance estimator:
where $\mathbf{E}^*$ denotes the expectation with respect to the bootstrap distribution. This suggests the modified inference procedure in Algorithm (ref). Note that $\hat V^*(\hat w)$ is outside the iteration over $w \in \Delta_{K-1}$. Hence, once the bootstrap estimator $\hat V^*(\hat w)$ is computed, the cost of the computation remains the same as before.
\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{Heuristics}
\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Construction of the Test Statistic}
Let us give a heuristic motivation behind the construction of the confidence set in Algorithm (ref). To construct a test statistic, we form the Lagrangian of the constrained optimization in ((ref)):
where $\tilde \lambda$ and $\lambda$ are Lagrange multipliers. The necessary and sufficient condition for $w_0$ to be a solution to the constrained optimization is:
for some $\tilde \lambda \in \mathbf{R}$ and $\lambda \in U(w_0)$, where
We concentrate out $\tilde \lambda$ to obtain the following equality:
for some $\lambda \in U(w_0)$. The equality has $K$ equations but these equations are not linearly independent. To extract maximally linearly independent equations out of these, we premultiply $B_2'$ to obtain the following equality restriction:
for some $\lambda \in U(w_0)$. The equality restrictions motivate the test statistic $T(w)$ in Algorithm (ref).
\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Construction of a Critical Value}
To provide an intuition behind the construction of the critical value, it helps to begin with the result of AlMohamad/vanZwet/Cator/Goeman:20:Biometrika. Suppose that $Y \sim N(\mu,V)$, for some $\mu \in \mathbf{R}^K$ and a symmetric positive definite matrix $V$. We introduce a polyhedral cone $\Lambda(w)$ as follows:
Define the norm $\| \cdot \|_{V}$ as $\|x\|_{V} = \sqrt{x' V^{-1} x}$, $x \in \mathbf{R}^K$. Then we focus on the test statistic of the form:
where $\Pi_{V}(Y \mid C)$ for a closed convex set $C$ denotes the projection of $Y$ on $C$ along $\| \cdot \|_{V}$. Let $\Lambda^\circ(w)$ be the polar cone of $\Lambda(w)$ and let $F_{\ell}$, $\ell=1,...,L$, be the faces of $\Lambda^\circ(w)$ and $\text{ri}(F_\ell)$ their relative interiors. Furthermore, let $k_\ell$ be the rank of the projection matrix of $Y$ onto the linear span of $F_{\ell}$. Then, Theorem 1 of AlMohamad/vanZwet/Cator/Goeman:20:Biometrika gives the following result:
where\footnote{AlMohamad/vanZwet/Cator/Goeman:20:Biometrika did not consider the case where the projection of $Y$ onto the polar cone becomes zero with positive probability. Thus, here we take $\max\{k_\ell,1\}$ in place of $k_\ell$ as in their paper.}
and $G(\cdot;k)$ denotes the CDF of the $\chi^2$ distribution with $k$ degrees of freedom. In order to adapt their result to our setting, we take two steps. First, we find a characterization of the critical value in our setting. Second, we extend the validity result to the desired asymptotic validity result.
We illustrate the main idea here, assuming that $V = I$ and is known. We write simply $q_{1-\alpha}(Y;w)$ instead of $q_{1-\alpha,V}(Y;w)$ and $\Pi(Y \mid C)$ instead of $\Pi_{V}(Y \mid C)$. The general case where $V$ is unknown and consistently estimated is dealt with in the proof of the main result in the appendix.
Step 1 (Critical Value Characterization): We first establish the following characterization of the critical value $q_{1-\alpha}(Y;w)$:
where
and $J_0[x] = \{j: x_j = 0\}$, the set of the indices of zeros in a vector $x$. To show this, we find the polar cone $\Lambda^\circ(w)$ of $\Lambda(w)$ as
where for any vector $a$, $[a]_J$ denotes the subvector of $a$ indexed by $J$. From this, the faces of $\Lambda^\circ(w)$ and their relative interiors are seen to be of the following form: for $J \subset J_0[w]$,
The linear span of $\text{ri}(\Lambda_{J}^\circ(w))$ is $L_J^\circ := \{x \in \mathbf{R}^{K-1}: [B_2 x]_J = 0\}$. Hence, the rank of the projection matrix onto this space is $K-1 - |J|$. Furthermore, since the relative interiors, $\text{ri}(\Lambda_{J}^\circ(w))$, partition the polyhedral cone $\Lambda^\circ(w)$, we find that
The projection of $Y$ onto $\Lambda^\circ(w)$ falls on $\text{ri}(\Lambda_{J}^\circ(w))$ if and only if the projection falls on the latter's linear span $L_J^\circ$. Therefore, only one of the terms in the sum on the right-hand side of ((ref)) realizes. Noting the decomposition
we write this term as $G^{-1}(1-\alpha;k(w))$.
To illustrate the degrees of freedom $k(w)$ for the $\chi^2$ distribution in the critical value, consider the case $K=3$ and $w = [1,0,0]'$. In this case, the cone $\Lambda(w)$ and its polar cone $\Lambda^\circ(w)$ are given as follows:
There are four faces of $\Lambda^\circ(w)$: $\Lambda_{\{2,3\}}^\circ(w)$, $\Lambda_{\{2\}}^\circ(w)$, $\Lambda_{\{3\}}^\circ(w)$, and $\Lambda_{\varnothing}^\circ(w)$. The relative interiors of these faces and their corresponding degrees of the $\chi^2$ distribution in the critical value are depicted in Figure (ref).
Step 2 (Extension to Asymptotic Validity): So far the result is for a normally distributed random vector $Y$. In order to accommodate random vectors that are asymptotically normal, we extend the previous result. This extension consists of two components: the asymptotic approximation of the test statistic and that of the critical values. For simplicity, we maintain the assumption that $\hat V_n$ is known to be $I$.
Consider a setting where
as $n \rightarrow \infty$. (Note that we allow $\|\mu_n\| \rightarrow \infty$ as $n \rightarrow \infty$.) From Skorohod representation, there exists a probability space and a sequence of random vectors $\tilde Z_n$ that have the same distribution as $Z_n$ and $\tilde Z_n \rightarrow_{a.s.} Z$ for a random vector $Z \sim N(0,I)$. We let $\tilde Y_n = \tilde Z_n + \mu_n$ and $Y_n^* = Z + \mu_n$. We focus on the limit behavior of the following test statistic:
Since the projection map on a nonempty, closed convex set is a contraction map, we have
as $n \rightarrow \infty$. Hence, it is not hard to see that, with $T_n^*(w) := \left\|Y_n^* - \Pi(Y_n^* \mid \Lambda(w)) \right\|^2$,
The main challenge is to show the asymptotic validity of the critical values. For this, we prove that for any subsequence of $\{n\}$, there exists a further subsequence $\{n'\} \subset \{n\}$, with large probability,
This yields the result that the critical value $c_n^*$ based on $Y_n^*$ is less than the critical value $\tilde c_n$ based on $\tilde Y_n$ eventually. Hence, the rejection probability $P\{T_n^*(w) > c_n^*\}$ is bounded below by $P\{\tilde T_n(w) > \tilde c_n\}$. By the construction of the critical value $c_n^*$ using the $\chi^2$ distribution, the former rejection probability is bounded by $\alpha$, delivering the asymptotic validity of the test.
It remains to show ((ref)). Since $|J_0[x]|$ is discontinuous in $x$, this result does not come directly from the convergence, $\tilde Y_n - Y_n^* \rightarrow_{a.s.} 0$. We begin by choosing any subsequence and find a further subsequence, along with some $J,J' \subset J_0[w]$ such that
We show that the set of values of $\tilde Y_n$ such that $J$ is not contained in $J'$ eventually has measure zero. In light of ((ref)), this yields ((ref)). While this summarizes the basic insights briefly, delivering the proof considering random variance matrices $\hat \Omega$ and the sequences of weights $w_n \in \Delta_{K-1}$ requires a considerable care for subtleties in the proofs. We refer the reader to the appendix for details.
\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{Discussions}
In this section, we discuss related methods developed in the literature and compare them with ours. For brevity, we discuss two methods that are most closely related to ours. First, we discuss the method of quadratic approximation of the sample objective function as proposed by Geyer:94:AS and Andrews:99:Ecma and a related proposal by Ketz:18:JOE. Second, we discuss Cox/Shi:23:ReStud who proposed inverting a test whose critical values are based on a $\chi^2$ distribution with a data dependent degree of freedom. Like our approach, their procedure does not involve tuning parameters that require a judicious choice to ensure stable finite sample properties.
\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Quadratic Approximation}
Our sample objective function $\hat Q(w)$ is a quadratic function of $w$. Hence, we may consider applying the method of quadratic approximation to address the issue of $w_0$ on the boundary of the simplex $\Delta_{K-1}$. To see how this method applies to our setting, consider the example of synthetic control with groupwise matching which takes the sample objective function $\hat Q(w)$ of the form:
where $\hat H$ and $\hat h$ are $\sqrt{n}$-consistent and asymptotically normal estimators of a positive definite matrix $H$ and a vector $h$. Now, we can write
for $w_0 \in \operatorname*{\arg\!\min}_{w \in \Delta_{K-1}} Q(w)$. The quadratic approximation method requires
for some random variable $\zeta$. Under regularity conditions, this requirement is met if
More generally, the quadratic approximation requires the following: at $w = w_0$,
This means that the global minimizer of $Q_P$ lies in the simplex $\Delta_{K-1}$. However, when the simplex constraint is binding, we do not have this equality and cannot the methods based on the quadratic approximation. The setting of synthetic control naturally uses additional equality restrictions ((ref)) which represent perfect pre-treatment fit at the population level. In this case, the condition ((ref)) is satisfied, and we can apply the method of quadratic approximation. However, as it is well known, due to the constraint imposed on the estimator $\hat w$, the limiting distribution of $\sqrt{n}(\hat w - w)$ is characterized as a minimizer of a random function. Thus, in order to obtain the critical value, we need to first simulate the random function and obtain a minimizer. In contrast to these approaches, our method is simple, as it does not require simulating the limiting distribution of the test statistic. More importantly, our method does not require that the objective function admit a quadratic approximation or the parameter $w_0$ be point-identified.
Ketz:18:JOE proposed a simple, interesting way to deal with a parameter on the boundary when the sample objective function admits a quadratic approximation. The estimator is obtained by applying a single Newton-Raphson adjustment to the constrained estimator. Let us study his approach in the context of the synthetic control setting. Let $\hat w$ be the constrained estimator of $w_0$. Then, from the quadratic approximation in ((ref)), his approach results in the following estimator:
This is the regression estimator by regressing $\hat \mu_{0,t}$ onto $\hat \mu_{1,t}$,...,$\hat \mu_{K,t}$. Hence, the estimator $\tilde w$ is an unconstrained estimator, allowed to take values outside $\Delta_{K-1}$. Due to this, the estimator is $\sqrt{n}$ consistent and asymptotically normal, under a regularity conditions where $w_0$ is $\sqrt{n}$-consistently estimable.
To map the construction of this test statistic to our setting, note that using the pre-treatment fit condition in synthetic control: $H w_0 = h$, we can write
This motivates the following quasi-likelihood ratio test statistic:
Under regularity conditions, we expect that $T_{\mathsf{NR}}(w_0) \to_d \chi_{K-1}^2$, as $n \to \infty$. Thus, the resulting method is non-adaptive to the possibility of $w_0$ being on the boundary.\footnote{Hsieh/Shi/Shum:22:JoE proposed a method of statistical inference on equality or inequality restrictions using a linear constraint formulation via Karush-Kuhn-Tucker conditions. We can adapt their proposal to this setting of simplex-valued weights. Like ours, their main proposal does not require a tuning parameter or simulating critical values. However, when adapted to our setting, the critical values are taken from the $\chi^2$ distribution with a degree of freedom $K+1$ (see (39) on page 258 of their paper.) Thus, the test is more conservative than the non-adaptive version above.} In contrast, our method is based on the idea from AlMohamad/vanZwet/Cator/Goeman:20:Biometrika and adaptive to $w_0$ being on the boundary. For comparison, our test statistic takes the form:
and the critical values are taken from $\chi_{\hat k(w)}^2$. While there is no uniform dominance of one test over the other, AlMohamad/vanZwet/Cator/Goeman:20:Biometrika present the power comparison between the two approaches and report simulation results in favor of the adaptive test.
\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Adaptive Moment Inequality Tests}
Closely related to our proposal is that of Cox/Shi:23:ReStud who proposed an adaptive size-exact subvector inference from moment inequality restrictions. They considered the moment inequality model:\footnote{Importantly, Cox/Shi:23:ReStud's framework accommodates a model with a nuisance parameter that enters the moment inequality restrictions in an additive form. To facilitate the comparison, we only discuss a simplified version of their model without a nuisance parameter here.}
where $A$ is a $d_A \times d_m$ matrix, $b$ is a $d_A$ dimensional vector, and $\overline m_n(\theta)$ is the sample average of the moment function $m(W_i,\theta)$:
To construct a confidence set, they considered the following test statistic:
where
Like our proposal, they constructed the critical value from a $\chi^2$ distribution with data-dependent degrees of freedom.
While our problem appears similar to theirs, their methods and asymptotic validity results do not subsume ours. First, our testing framework is not necessarily motivated from moment inequality restrictions. Hence, we cannot directly use their uniform asymptotic validity result without modification in our setting. Second, our method is simpler than theirs in our case. To see this, we rewrite $T(w)$ as follows:
where $\Lambda(w)$ is the polyhedron defined in ((ref)) and $\hat f(w) = B_2'\hat \varphi(w)$. To the best of our knowledge, there does not seem to be an explicit solution for a matrix $A(w)$ and a vector $b(w)$ such that
While one may develop an algorithm to compute $A(w)$ and $b(w)$ and apply the approach of Cox/Shi:23:ReStud, our proposal is already simpler than this, as we do not need to compute $A(w)$ and $b(w)$ for each $w \in \Delta_{K-1}$.
Alternatively, we may consider rewriting $T(w)$ as follows:
to map this to the setting of Cox/Shi:23:ReStud. In this case, we can explicitly find matrices $A(w)$ and a vector $b(w)$ such that
However, unlike $\hat \Sigma(\theta)$ in ((ref)), the matrix $B_2 \hat V(w)^{-1} B_2'$ is a singular matrix both in finite samples and in the limit. Hence, the results and proofs of Cox/Shi:23:ReStud are not directly applicable here. The singularity problem in our setting stems from the linear constraint $w'\mathbf{1} = 1$ rather than from accommodating distributions with singular covariance matrices in uniformly asymptotically valid inference. Therefore, simply removing such distributions from consideration cannot resolve this issue.
\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{Extension: Inference on a Function of $w_0$} In many applications, the main parameter of interest is an identified function of $w_0$:
for some map $\theta_P: \Delta_{K-1} \to \Theta$ and $\Theta \subset \mathbf{R}^L$ is the parameter space for $\theta_0$.\footnote{Recently, Dufour/Tuvaandorj:25:arXiv developed a quasi-likelihood ratio test when the parameter of interest is a known function of a nuisance parameter that is on the boundary, and showed its pointwise asymptotic validity. Their test is simple to use, involving no tuning parameters or bootstrap for critical values. VanDijcke/Gunsilius/Wright:24:arXiv developed bootstrap inference for treatment effect parameters identified via the distributional synthetic control approach of Gunsilius:23:ECMA and proved its pointwise asymptotic validity. However, their proof requires Hadamard differentiability of the parameter functional at distributions outside a certain exceptional set, thereby precluding uniform asymptotic validity over any region containing this set.} In this setting, we consider two approaches for constructing a confidence set for $\theta_0$.
The first approach simply uses the Bonferroni method. More specifically, suppose that the estimator $\hat \theta$ satisfies the following asymptotic linear representation:
where $\Sigma_P(w_0)$ is a symmetric positive definite matrix, and $n$ denotes the size of the sample from the population that identifies $\theta_P(\cdot)$. Then, for any $\alpha \in (0,1)$, we can use the confidence set $C_{1-\kappa}$ for $w_0$ with level $100(1-\kappa)\%$, $\kappa \in (0,\alpha)$ (in Algorithm (ref)) and construct the $100(1-\alpha)\%$-level confidence interval for $\theta$ as
where $\hat \Sigma(w)$ is a consistent estimator of $\Sigma_P(w)$ and $q_{1-\alpha - \kappa}$ denotes the $(1-\alpha-\kappa)$ quantile of the $\chi^2$ distribution with degree of freedom $L$.
The second approach is to construct a joint confidence set for $[w_0',\theta_P(w_0)']'$ and project the confidence set on the space for $\theta_P(w_0)$.\footnote{We are grateful to a referee from our previous submission for suggesting this idea.} First, let us assume that for each $w \in \Delta_{K-1}$,
as $n \to \infty$, for a symmetric positive definite covariance matrix $\Omega_P(w)$, where $\hat \varphi(\cdot)$ is a $\sqrt{n}$-consistent estimator of $\varphi_P(\cdot)$. (Note that we allow for the sample size for $\hat \theta(\cdot)$ to be different from that for $\hat \varphi(\cdot)$.) We let
Then, using these and a consistent estimator $\hat \Omega(w)$ of $\Omega_P(w)$, we construct a confidence set $C_{1-\alpha}^{\mathsf{joint}}$ as in Algorithm (ref). The confidence set for $\theta$ can be obtained from projecting the joint confidence set onto $\Theta$:
While the projection method can be among various methods of subvector inference, we have opted for this method due to its simplicity, being free of any tuning parameters. We relegate the investigation of improvement on this method to future research.
\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{Uniform Asymptotic Validity}
\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{A Preliminary Result}
We first provide a preliminary result. Suppose that $Y_n \in \mathbf{R}^{K+L-1}$, $n \ge 1$, is a sequence of random vectors such that
where $Z_n \in \mathbf{R}^{K+L-1}$ is a sequence of asymptotically standard normal random vectors, $\hat \Omega$ is a sequence of symmetric positive definite $(K+L-1) \times (K+L-1)$ random matrices, and $\mu_n \in \mathbf{R}^{K+L-1}$ is a sequence of nonstochastic vectors. We do not assume that the sequence $\mu_n$ is bounded.
Our main focus is on the test statistic that is based on the squared error from the projection of $Y_n$ onto a polyhedral cone along the norm $\|\cdot \|_{\hat \Omega}$, where $\| x \|_{\hat \Omega}^2 = x' \hat \Omega^{-1} x$, $x \in \mathbf{R}^{K+L-1}$, and $\Lambda(w)$ is defined in ((ref)). Recall that $G(\cdot;k)$ denotes the CDF of the $\chi^2$ distribution with the degrees of freedom equal to $k$ and $J_0[z] = \{1\le j \le K: z_j = 0\}$ for any vector $z = [z_j] \in \mathbf{R}^K$. For each $w \in \Delta_{K-1}$, we define
The test statistic takes the form of the squared residuals from projecting $Y_n$ onto $\tilde \Lambda(w)$:
where $\Pi_{\hat \Omega}(Y_n\mid \tilde \Lambda(w))$ denotes the projection of $Y_n$ onto $\tilde \Lambda(w)$ along $\| \cdot \|_{\hat \Omega}$.
As for the critical value, let $\hat \delta(w) \coloneqq B \hat \Omega^{-1} (Y_n - \Pi_{\hat \Omega}(Y_n \mid \tilde \Lambda(w)))$ and
with $B$ defined in ((ref)) and $0_{K \times L}$ denoting the $K \times L$ dimensional matrix of zeros. Then, the construction of the confidence set for $w$ is based on the critical value of the following form:
This critical value is random, depending on the data through its dependence on the number of zeros in $\hat \delta(w)$.
We first introduce an assumption that requires the consistency of $\hat \Omega$ and the asymptotic normality of $Z_n$ along a sequence of probabilities $P_n$.
The asymptotic normality of $Z_n$ and the consistency of $\hat \Omega$ are typically satisfied in the applications under consideration in this paper. The following lemma is crucial for the uniform asymptotic validity result.
This lemma shows that reading the critical values from the $\chi^2$ distribution with the degrees of freedom equal to $\hat k(w)$ generates a uniformly asymptotically valid procedure. The proof of the lemma can be potentially useful for many similar settings where the test statistic involves a projection onto a polyhedral cone.
\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{The Main Result}
Let us revisit Section (ref) and present the main result of the asymptotic validity of the confidence set for $(w_0,\theta_0)$. We focus on the uniform asymptotic validity of the joint confidence set $C_{1-\alpha}^{\mathsf{joint}}$ in Algorithm (ref), as it is more general than $C_{1-\alpha}$ in Algorithm (ref). First, we introduce a set of assumptions. From here on, $\mathcal{P}_n$ denotes the set of probability distributions of data that are under consideration. Our goal is to establish the uniform asymptotic validity of the confidence set $C_{1-\alpha}^{\mathsf{joint}}$ for $w_0$ over the class of probability distributions $\mathcal{P}_n$.
Condition (i) requires the asymptotic normality of the estimators that is uniform over the probabilities in $\mathcal{P}_n$. This condition is satisfied in many settings where maps, $B_2' \varphi_P(w)$, $\theta_P(w)$, are “regular”, i.e., behave “smoothly” in the perturbation of $P$. Condition (ii) requires that the limit distribution is non-degenerate uniformly over the probabilities in $\mathcal{P}_n$. Condition (iii) requires the variance estimator $\hat \Omega(w_n)$ to be consistent for $\Omega_P(w_n)$ uniformly over the probabilities in $\mathcal{P}_n$. Condition (iv) requires that the minimizer of $Q_{P_n}(w)$ over $w \in \Delta_{K-1}$ is close to the global minimizer. This condition is easily satisfied with the synthetic control examples with the pretreatment fit. In these cases, we have $\varphi_{P_n}(w) = 0$ for all $w \in \mathbb{W}_{P_n}$. The constant $C$ is permitted to be unknown, and hence the condition (iv) is too weak to apply the approach of quadratic approximation.
Thus, the conditions in Assumption (ref) essentially define the scope of the method and are satisfied in many settings. By applying Lemma (ref), we obtain the following result.
The proof is found in the appendix of this paper. To see how Lemma (ref) gives this validity result, observe that
where
Thus, the test statistic $T(w,\theta)$ is nothing but the squared length of the projection error in projecting $Y_n(w)$ onto the polyhedral cone $\Lambda(w)$ along the norm $\| \cdot \|_{\hat \Omega(w)}$. Since
with
the random vector $Y_n(w,\theta_{P_n}(w))$ is asymptotically normal, yet with a potentially diverging shift $\mu_n(w)$. This setting maps to the one in Lemma (ref).\footnote{Note that we allow for the possibility that $\liminf_{n \rightarrow \infty}\varphi_{P_n}(w_n) \ne 0$ for $w_n \in \mathbb{W}_{P_n}$ such as when the minimizer of $Q_{P_n}(w)$ over $w \in \mathbf{R}^K$ lies outside $\Delta_{K-1}$. In this case, the sequence $\|\mu_n(w_n)\|$ may diverge to $\infty$, as $n \rightarrow \infty$.} The uniform asymptotic validity of the confidence set $C_{1-\alpha}$ for $w_0$ follows from the lemma.
\@startsection{subsection}{2} \z@{.5\linespacing\@plus.7\linespacing}{.7\linespacing}{Monte Carlo Simulations}
We now turn to Monte Carlo simulations to evaluate the finite sample performance of our procedure. Our set-up follows Example (ref) and the literature on synthetic control with individual data, which are of significant applied interest (see Abadie/LHour:21:JASA). This will also be the focus of our empirical application below.
We consider a treated group (0) and $K$ untreated groups, where $K \in \{3,5,7\}$. Each group $j$ has a sample size $n_j \in \{100, 200, 1000\}$. The outcomes for individual $i$ in group $j$, $Y_{ij,t}$, are realized over $T$ periods, but only the first $T_0 = 10$ periods are used to match the treated and untreated groups.
We generate the pre-treatment outcomes for the untreated groups $j=1,...,K$ using the linear model:
where $\mu_{j,t} = 0.5 + 0.5(-1)^{j-1}\frac{t}{T} + 0.5 \eta_{j,t}$, $\eta_{j,t} \sim_{i.i.d.} N(0,1)$ and $\varepsilon_{ij,t} \sim_{i.i.d.} N(0,1)$. Thus, the mean outcome for $j$, $\mu_{j,t}$ is composed by a deterministic component and a random component. The former grows over time for odd-numbered $j$, while it decreases over time for even-numbered groups. The random component is fixed throughout simulations. Individual outcomes in a group are noisy observations of the associated means.
We draw data for the treated group as:
where $\mu_{0,t} = w_{0}' \mu_{t}$, $\mu_t = (\mu_{1,t},...,\mu_{K,t})'$, $w_{0} \in \Delta_{K-1}$, and $\varepsilon_{i0,t} \sim_{iid} N(0,1)$ across $i$ and $t$. Thus, the expected outcomes for the treated group are a weighted (synthetic) combination of expected outcomes from the untreated groups with weights, $w_0$.\footnote{While this may resemble the assumption in the synthetic control literature that the outcomes of the treated group have a perfect match in the untreated ones, there is a major distinction: we impose the matching on the population quantities rather than on their sample counterparts. It is important for us to distinguish between population and sample. Otherwise, there is no inference to be made on $w_0$. For the same reason, the observed outcomes are noisy observations of the population means in our setting.} We consider the setting where $T > K$ and there is no linear dependency between the population means $\mu_{j,t}$, $j=0,...,K$. Hence, the weights $w_0$ are point-identified.\footnote{As for our grid for $w$, we use two complementary approaches that allow us to extensively explore both the interior and the boundary of the simplex. The first draws a fine grid of $w$ uniformly over its simplex using a procedure based on Rubin:81:AoS. This is done by first drawing a vector of dimension $K-1$, with each element i.i.d. from the uniform distribution with support $[0,1]$. Then, we include 0 and 1 into that drawn vector, sort it, and produce the vector of differences across adjacent elements of $w$. These are all nonnegative and sum up to 1 by construction. As for the boundary points, we first define a step size, generate a vector of $[0,1]$ with points separated by that step-size, and then generate a $K$-dimensional meshed grid of all possible combinations of those values. We only keep the points that (i) add up to 1, (ii) have a zero element. The final grid is the union of the gridpoints generated by both approaches.}
The weight $w_0$ is the key object in inference. We consider two main specifications, each one presenting a different case of whether $w_0$ is interior or on the boundary of the simplex. In the former (Specification 1), we set $w_0 = (0.2, 0.8 \mathbf{1}_{K-1}'/(K-1))'$, so that $w_0$ lies in the interior of $\Delta_{K-1}$, while, in the latter (Specification 2), $w_0 = (0.5, 0.5, \mathbf{0}_{K-2}')'$, so that $w_0$ lies on the boundary of $\Delta_{K-1}$, where $\mathbf{1}_{d}$ and $\mathbf{0}_{d}$ are $d$-dimensional vectors of 1's and 0's, respectively. In each of $1,000$ simulations, we independently draw $(\varepsilon_{i0,t}, \varepsilon_{i1,t},...,\varepsilon_{K,t})'$ for $t=1,...,T_0$ and $i=1,...,n_j$, while keeping $\mu_{j,t}$ fixed. Since the outcomes $Y_{ij,t}$'s are independent across time, the data generating process represents a repeated cross-section design. All results are shown in Table (ref).
Table (ref) shows that our procedure works very well in finite samples across scenarios. Let us start with the interior weights case (Specification 1). In line with the theory, coverage probabilities are very close to nominal levels across specifications. This holds even for the smallest sample sizes ($n_j=100, K = 3$ and $n_j = 100, K = 5$), with empirical coverage probabilities between 93.6-96.2%. Coverage probabilities become very close to 95% even when $n_j = 200$, across specifications (94.7, 94.6 and 94.1% for $K \in \{3,5,7\}$, respectively). Similarly, the results continue to hold as $n_j$ or $K$ grows because our asymptotic results are based on the total sample size (numbers of individuals times number of groups). This is seen both with $K=5$ and $K=7$, with coverage probabilities at or close to 95% for $n_j = 1,000$.
The results from Specification 2 show that our procedure also works well even when $w_0$ is on the boundary of the simplex. Indeed, coverage probabilities are close to the nominal 95% across specifications and sample sizes. This holds both for small sample sizes (96.8% and 94.1% for the smallest values of $K$ and $n_j$), but also as both grow. Indeed, they become very close to 95% for large sample sizes: with $n_j = 200$, coverage probabilities are 94.9% for $K=3$, 94.6% when $K=5$ and 93.9% for $K=7$. Similar results hold when $n_j = 1,000$.
Sometimes researchers are interested in inference on an individual component of $w_0$ rather than inference on the full vector $w_0$. For example, one may be interested in the weight that the synthetic method gives to a specific untreated group (e.g., state, country, firm, or forecaster in the case of predictions). To that end, one can combine our approach with projection-based subvector inference. Table (ref) presents the results of this approach for the first dimension of $w_0$ in the specifications of Table (ref). We report both the empirical coverage probability for $w_{0,1}$ and the mean length of the confidence interval. This also provides information on the power of our inference approach.
The projection-based inference works well in finite samples. As expected, it is more conservative than inference on the full $w_0$ (c.f. Table (ref)). For $K=3, 5$ and interior $w_0$, the former is typically 2-4 percentage points more conservative than the latter, although such conservativeness decreases with sample size (e.g., it is below 1 percentage point when $n_j = 1,000$ and $K=5$ or $K=7$, regardless of specification). In spite of its conservative nature, projection-based inference is still informative in our setting. Even when sample size is small ($n_j = 100, K=3$), the average length of the confidence interval can be as short as 0.14-0.16 and it is far from spanning $[0,1]$. Furthermore, for a fixed $K$ and weight $w_0$, as sample size $n$ increases, the length of the confidence interval quickly decreases and gets even close to 0: when $n_j = 1,000$, the average length is below 0.04 across specifications. The latter would be very sharp for inference with a sample size often found in empirical applications with individual-level data. Thus, our simulations suggest that, when $w_0$ is point-identified and the sample size is large enough, projection-based inference for the weights can be informative and not overly conservative.
\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{Empirical Application}
As an empirical illustration of our method, we study how a large increase in minimum wages in Alaska in 2003, from US\$5.65 to US\$7.15, affected average family income. To do so, we use individual-level data and state-level weights, as in Example (ref). This follows a recent strand of the literature that uses synthetic control methods to obtain credible estimates of the effects of minimum wage increases on economic outcomes (e.g., Allegretto/Dube/Reich/Zipperer:17:ILR, Neumark/Wascher:17:ILR, Powell:22:JBES) and the analysis of this policy in Gunsilius:23:ECMA.
Following Dube:19:AEJ and Gunsilius:23:ECMA, our outcome of interest is a measure of household income, called average family income. This is a measure which adjusts for family size and composition (i.e., “equivalized”) and it is measured in multiples of the federal poverty threshold. We use the main sample from Dube:19:AEJ (i.e., individuals aged under 65), originally drawn from the Current Population Survey (CPS), and focus on the subsample from 1998-2002 (i.e., the five years before the policy).
To form a synthetic control, we form a pool of states that resemble Alaska in terms of population and economic activity. We start with states that, like Alaska, are among the top-10 oil producing states in 2002 (i.e., prior to the policy), e.g., see Figure 7 in Ismayilova:07:Thesis. We then drop states that have a minimum wage increase during the period of interest (California) and states that have a population over 3 million (Texas, Louisiana and Oklahoma). By comparison, Alaska had a population of approximately 650,000 in 2002, Census2002. Our final set of control states is, thus: Kansas, Mississippi, New Mexico, North Dakota and Wyoming ($K=5$). From this pool, we construct a synthetic comparison state by deriving state-level weights from estimating $\mu_{j,t}$: the mean family income in each state $j$ and year $t$.
We use our statistical procedure outlined in Section (ref) and report the projection-based inference for weights and Bonferroni-based confidence interval for a parameter of interest, $\theta_0$, defined as the difference in mean outcomes in Alaska in 2003 (the year post-policy) relative to the mean in the synthetic control. This can be written as:
where $\mu_{\text{ctrl},2003} = [\mu_{1,2003}, \mu_{2,2003},...,\mu_{K,2003}]'$ is a vector of the mean average family income across families in the control states and $\theta_0 = \theta(w_0)$. We can compute the asymptotic variance $\Omega(w)$ as explained in Example (ref).
The first row of Table (ref) presents the confidence interval for the effect of the increase in minimum wages on average family income in 2003 across specifications. The results are robust across specifications and inference methods, with confidence intervals centered around -0.2. From our results, we can reject the presence of any meaningful positive effects, as the upper bound of our confidence intervals are all very close or at 0 (and below a 1% effect). This suggests that this minimum wage policy, at best, did not affect average family income. At worst, it decreased average family income by up to 13.4% relative to a pre-policy (1998-2002) average income of 3.43 times the federal poverty threshold. While it is likely that the poorest households benefit from this policy, the average family may have lost a small amount of income, possibly due to increased unemployment. Our confidence intervals account for inference on the weights, which contrasts with the standard approach in synthetic control, where the synthetic weights are assumed to hold exactly for any realization of the sample. While the results are robust across specifications (i.e., with or without Mississippi) and methods (Bonferroni or projection-based), we do find that the projection method presented in Algorithm (ref) seems to outperform the Bonferroni counterpart in this setting. Indeed, the average length of the projection-based confidence interval for $\theta_0$ is 16% (0.4/0.48) shorter than the Bonferroni counterpart, and the former is a subset of the latter. This empirical gain comes at an increased computational cost.
Table (ref) further reports projection-based confidence intervals for weights across two choices of control units: with and without Mississippi. Kansas often received the largest weights (i.e., with confidence intervals centered around 0.8-0.85), followed by Wyoming, while the others have relatively similar confidence intervals centered around 0-0.05. Meanwhile, the average length of the confidence intervals is around 0.15. This provides further support to the findings in the simulations that projection-based inference for $w$ can be informative in this context.
\@startsection{section}{1} \z@{0.8\linespacing\@plus\linespacing}{.7\linespacing}{Conclusion}
In this paper, we propose a new method for inference on the simplex-constrained weights. Our method is based on the projection of an asymptotically normal statistic onto a polyhedral cone. We show that the critical value for the test statistic can be read from the $\chi^2$ distribution with data-dependent degrees of freedom. This leads to an asymptotically uniformly valid confidence set for the weights. We provide a proof of this result and show that it is applicable to a wide range of settings. We also show that the method works well in finite samples through Monte Carlo simulations. We apply our method to a synthetic control example and show that it can be useful in practice.