Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
105,972 characters · 23 sections · 19 citation commands
Supplement to “Uncertainty Quantification in Synthetic Controls with Staggered Treatment Adoption”
We summarize the notation used throughout the paper in the following table.
In Section 4.2 we discuss the approach for quantifying the out-of-sample uncertainty based on the non-asymptotic bounds. We briefly describe two strategies below.
Section 4.4 constructs prediction intervals with simultaneous coverage. We briefly describe two other common approaches below.
When the causal predictand of interest depends on treatment effects on multiple treated units, such as the time-specific unit-averaged predictand $\tau_{\mathcal{Q}k}$, there is an alternative method for in-sample uncertainty quantification that may yield tighter prediction intervals. It relies on the fact that if the SC weights $\widehat{\boldsymbol{\beta}}^{[i]}$ for each ever-treated unit $i$ are obtained through a separate optimization process using data $\mathbf{A}^{[i]}$, $\mathbf{B}^{[i]}$ and $\mathbf{C}^{[i]}$, then, as outlined in Remark 1, a sequence of restrictions must be obeyed: for each $i\in\mathcal{E}$,
Then, analogously to (6.5) in the main text, for any predictand $\tau$ defined before, we can set
conditional on the data, with $\tilde{\mathcal{M}}_{\mathbf{G}}^\star=\{\boldsymbol{\delta}=(\boldsymbol{\delta}^{[1]\prime},\cdots, \boldsymbol{\delta}^{[J_1]\prime})': \; \boldsymbol{\delta}^{[i]}\in\Delta^{[i]^\star},\; \boldsymbol{\delta}^{[i]\prime}\widehat\mathbf{Q}^{[i]}\boldsymbol{\delta}^{[i]}- 2(\mathbf{G}^{[i]\star})'\boldsymbol{\delta}^{[i]}\leq 0,\; i\in\mathcal{E}\}$ and $\mathbf{G}^\star=(\mathbf{G}^{[1]\prime},\cdots, \mathbf{G}^{[J_1]\prime})'|\mathsf{Data}\sim \mathsf{N}(\mathbf{0},\widehat{\boldsymbol{\Sigma}})$. Again, $\Delta^{[i]}$ is replaced with its feasible (enlarged) version $\Delta^{[i]\star}$, and $\widehat{\boldsymbol{\gamma}}^{[i]}-\boldsymbol{\gamma}^{[i]}$ is replaced with the approximating Gaussian vector $\mathbf{G}^{[i]\star}$. As a reminder, this strategy is not equivalent to doing in-sample uncertainty quantification for each time-specific unit-specific prediction $\widehat{\tau}_{ik}$ and then combining their bounds together, since the whole vector $\mathbf{G}^\star$ has the same covariance as $\widehat{\boldsymbol{\gamma}}=(\widehat{\boldsymbol{\gamma}}^{[1]\prime}, \cdots, \widehat{\boldsymbol{\gamma}}^{[J_1]\prime})'$, thus maintaining the correlation structure among different treated units.
Technically, the above optimization procedure for constructing SC weights is equivalent to a special case of (6.1), where the weighting matrix $\mathbf{V}$ is block diagonal, taking the form $\mathbf{V}=\mathsf{diag}(\mathbf{V}^{[1]}, \cdots, \mathbf{V}^{[J_1]})$, and the feasible set $\Delta=\mathcal{W}\times \mathcal{R}$ is the Cartesian product of $\Delta^{[i]}=\mathcal{W}^{[i]}\times \mathcal{R}^{[i]}$ for each subvector $\boldsymbol{\beta}^{[i]}$ (hence, there is no cross-treated-unit constraint). The general in-sample uncertainty quantification strategy presented earlier in (6.5) of the main text relies on the fact that $(\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)'\widehat{\mathbf{Q}}(\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)- 2(\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma})'(\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)= \sum_{i\in\mathcal{E}}[(\widehat{\boldsymbol{\beta}}^{[i]}-\boldsymbol{\beta}_0^{[i]})'\widehat{\mathbf{Q}}^{[i]}(\widehat{\boldsymbol{\beta}}^{[i]}-\boldsymbol{\beta}_0^{[i]})- 2(\widehat{\boldsymbol{\gamma}}^{[i]}-\boldsymbol{\gamma}^{[i]})'(\widehat{\boldsymbol{\beta}}^{[i]}-\boldsymbol{\beta}_0^{[i]})]\leq 0$ and $\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0\in\Delta=\times_{i\in\mathcal{E}}\Delta^{[i]}$, which are immediately implied by (ref). In this sense, the restriction (ref) is stricter, making the alternative bounds in (ref) tighter (or at least no looser) than those in (6.5) in the main text. Table (ref) quantifies the gains from this restriction in terms of the length of the prediction intervals in our leading empirical application.
Due to the complexity of the feasibility set $\tilde{\mathcal{M}}_{\mathbf{G}}^\star$, the alternative bounds (ref) may not be theoretically justified using the same argument for (6.5) in the main text. However, other techniques, such as strong approximations, can be employed to establish the validity of (ref), albeit at the expense of additional technical complexity. We do not pursue further justification of (ref) and leave it for future research.
There are different ways to justify the synthetic control method, including the cointegrated system motivated by our empirical application. In this section, we briefly discuss an alternative justification that assumes the data are generated through a linear factor model: \[ Y_{it}(\infty)=\lambda_t\mu_i + \nu_{it}, \qquad 1\leq i\leq N,\; 1\leq t\leq T. \] For simplicity, assume that only the first unit is treated, the common factor $\lambda_t$ is i.i.d. over $1\leq t\leq T$, $\nu_{it}$ is i.i.d. across $1\leq i\leq N$ and over $1\leq t\leq T$, and $\{\lambda_t:1\leq t\leq T\}$ and $\{\nu_{it}:1\leq i\leq N, 1\leq t\leq T\}$ are independent of each other. In addition, we assume that $\mathbb{E}[\nu_{it}]=0$, the factor loadings $\{\mu_i:1\leq i\leq N\}$ are fixed, and only the pre-intervention outcomes are used in the SC construction. Then, $\mathscr{H}=\{(Y_{2t}(\infty), \cdots, Y_{Nt}(\infty)): 1\leq t\leq T\}$. Accordingly, $\mathbf{w}_0$ is given by the following expression: \[ \mathbf{w}_0=\underset{\mathbf{w}\in\mathcal{W}}{\operatorname*{arg\,min}}\; \frac{1}{T_0}\sum_{t=1}^{T_0}\mathbb{E}\Big[\Big((\mu_1-\boldsymbol{\mu}_c'\mathbf{w})\lambda_t+(\nu_{1t}-\bm{\nu}_{t,c}'\mathbf{w})\Big)^2\Big|\mathscr{H}\Big], \] where $\boldsymbol{\mu}_c=(\mu_2, \cdots, \mu_N)'$ and $\bm{\nu}_{t,c}=(\nu_{2t}, \cdots, \nu_{Nt})'$, assuming the expectation (exists and) is finite. Define $\mathscr{H}_t=\{Y_{2t}(\infty), \cdots, Y_{Nt}(\infty)\}$. Then, $\mathbf{w}_0$ can be further written as
assuming these expectations are finite. Given our assumptions on $\lambda_t$ and $\nu_{it}$, we expect that $$\frac{1}{T_0}\sum_{t=1}^{T_0}\mathbb{E}[\lambda_t^2|\mathscr{H}_t] \approx \mathbb{E}[\lambda_t^2], \quad \frac{1}{T_0}\sum_{t=1}^{T_0}\mathbb{E}[\bm{\nu}_{t,c}\bm{\nu}_{t,c}'|\mathscr{H}_t] \approx \mathbb{E}[\nu_{1t}^2]\mathbf{I}_{N-1}, \quad\text{and}\quad \frac{1}{T_0}\sum_{t=1}^{T_0}\mathbb{E}[\lambda_t\bm{\nu}_{t,c}'|\mathscr{H}_t]\approx \bm{0} $$ with high probability when $T_0$ is large, which can be shown under additional moment conditions. As a consequence, our expression for $\mathbf{w}_0$ is similar to that in Ferman-Pinto_2021_wp.
Then, the in-sample error for the prediction of the TSUS effect $\tau_{1k}$ is given by \[ \mathsf{InErr}(\tau_{1k})=-\mathbf{Y}_{\mathcal{N}t}'(\widehat{\mathbf{w}}-\mathbf{w}_0), \] which represents the error in estimating the weights. The out-of-sample error is given by \[ \mathsf{OutErr}(\tau_{1k})=Y_{1(T_1+k)}-\mathbf{Y}_{\mathcal{N}(T_1+k)}'\mathbf{w}_0= \nu_{1(T_1+k)}-\bm{\nu}_{T_1+k,c}'\mathbf{w}_0+ (\mu_1-\boldsymbol{\mu}_c'\mathbf{w}_0)\lambda_{T_1+k}. \] The first term $\nu_{1(T_1+k)}$ represents the treated unit's innovation in the post-treatment period $T_1+k$; the second term $\bm{\nu}_{T_1+k,c}'\mathbf{w}_0$ represents the weighted average of the innovations for the donor units in the post-treatment period $T_1+k$; and the third term $(\mu_1-\boldsymbol{\mu}_c'\mathbf{w}_0)\lambda_{T_1+k}$ can be thought of as the impact of the “bias” of the SC weights in $T_1+k$, if one thinks the target weight in this context is the one recovering the true factor loading $\mu_1$ of the treated unit. Given the conditions imposed, our method is still applicable to this scenario, and importantly, the “bias” due to $(\mu_1-\boldsymbol{\mu}_c'\mathbf{w}_0)\lambda_{T_1+k}$ is taken into account.
As explained in the main paper, by convexity of the constraint set $\mathcal{W}\times\mathcal{R}$ and the optimality of $\widehat{\boldsymbol{\beta}}$, $$ \inf_{\boldsymbol{\delta}\in\mathcal{M}_{\widehat\boldsymbol{\gamma}-\boldsymbol{\gamma}}}-\mathbf{p}_\tau'\boldsymbol{\delta} \leq -\mathbf{p}_\tau'(\widehat\boldsymbol{\beta}-\boldsymbol{\beta}_0)\leq \sup_{\boldsymbol{\delta}\in\mathcal{M}_{\widehat\boldsymbol{\gamma}-\boldsymbol{\gamma}}}-\mathbf{p}_\tau'\boldsymbol{\delta}, $$ where $\mathcal{M}_{\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma}}= \{\boldsymbol{\delta}\in\Delta:\boldsymbol{\delta}'\widehat{\mathbf{Q}}\boldsymbol{\delta}- 2(\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma})'\boldsymbol{\delta}\}$. Thus, condition (i) in Theorem 1 indeed requires that $\widehat\boldsymbol{\gamma}-\boldsymbol{\gamma}$ can be approximated by a Gaussian vector $\mathbf{G}$. Corollary 1 in the paper provided a verification of this condition in the special case of cointegrated data. In this section, we provide a more general way to verify condition (i) by imposing a conditional independence assumption on the pseudo-true residuals. The extension that allows for weakly dependent errors can be established using the idea of Theorem A in Cattaneo-Feng-Titiunik_2021_JASA. For simplicity, we assume that only $T_0$ pre-treatment periods are used to obtain the weights. Also, we write $\mathbf{U}^{[i]}=(u_{it,1}, \cdots, u_{it,M})'$, which is the vector of pseudo-true residuals corresponding to the treated unit $i$.
Our proposed strategy in Section 6.2 for constructing the constraint set in the simulation relies on the tuning parameters $\varrho_\ell$'s, which are used to differentiate the binding constraints from the non-binding ones. The recommended choice described in (6.8) and (6.9) can be rationalized as follows.
For ease of notation, suppose that there is only one treated unit, and thus we can omit the superscript $[i]$ for quantities like $\boldsymbol{\beta}_0^{[i]}$ or $\varrho^{[i]}$ with no confusion. Assume the $\ell$-th constraint $m_{\leq, \ell}(\boldsymbol{\beta})\leq 0$ is binding ($m_{\leq, \ell}(\boldsymbol{\beta}_0)=0$), and $m_{\leq, \ell}(\cdot)$ is sufficiently smooth. By the mean value theorem, $|m_{\leq, \ell}(\widehat{\boldsymbol{\beta}})|=|\frac{\partial}{\partial\boldsymbol{\beta}'}m_{\leq, \ell}(\check\boldsymbol{\beta})(\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0)|\leq \|\frac{\partial}{\partial\boldsymbol{\beta}}m_{\leq, \ell}(\check\boldsymbol{\beta})\|_2\|\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0\|_2$, where $\check{\boldsymbol{\beta}}$ is some point between $\boldsymbol{\beta}_0$ and $\widehat{\boldsymbol{\beta}}$ . The parameters $\varrho_\ell^{[i]}$ are used to characterize this upper bound. For $\frac{\partial}{\partial\boldsymbol{\beta}}m_{\leq, \ell}(\check\boldsymbol{\beta})$, a natural “estimator” would be $\frac{\partial}{\partial\boldsymbol{\beta}}m_{\leq, \ell}(\widehat{\boldsymbol{\beta}})$. The difference between the two is bounded by $c_1\|\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0\|_2$ for some constant $c_1>0$ (with high probability). Then, it remains to characterize $\|\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0\|_2$.
Note that the basic inequality below always holds (deterministically) by optimality of $\widehat{\boldsymbol{\beta}}$:
We first discuss the denominator. Recall that $\widehat{\mathbf{Q}}=\mathbf{Z}'\mathbf{V}\mathbf{Z}$. Since an (unrestricted) intercept is often included in the SC construction, we assume the predictor variables included in $\mathbf{Z}$ are de-meaned (and the constant term is excluded from $\mathbf{Z}$). If $\mathbf{V}$ is an identity matrix and the predictor variables are (approximately) orthogonal to each other, then $s_{\min}(\widehat{\mathbf{Q}}/T_0)$ is the minimum diagonal entry of $\widehat{\mathbf{Q}}/T_0$, i.e., $\min_{1\leq j\leq J_0}\widehat{\sigma}_{b_j}^2$, which motivates the choice of the denominator in (6.9). On the other hand, recall that the numerator $\widehat{\boldsymbol{\gamma}}=\mathbf{Z}'\mathbf{u}$. As discussed in the paper, we can (conditionally) approximate $\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma}$ by the Gaussian vector $\mathbf{G}$. Then, by Gaussian tail bound, it can be shown that, for any large $\lambda>0$, $T_0^{-1/2}\|\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma}\|_\infty\leq (\sqrt{2\log d_\beta}+\lambda)\times\sqrt{\max_{1\leq j\leq J_0} \mathbb{V}[\widehat{\gamma}_j|\mathscr{H}]/T_0}$ with high probability. When $\mathbf{u}$ is independent of $\mathscr{H}$ and its components are independent over $t$, $\max_{1\leq j\leq J_0} \mathbb{V}[\widehat{\gamma}_j|\mathscr{H}]/T_0=\sigma_u^2\times \max_{1\leq j\leq J_0}\widehat{\sigma}_{b_j}^2$. Therefore, we can set \[ \mathcal{C}=\mathcal{C}_1:= \frac{2\sqrt{d_\beta}(\sqrt{2\log d_\beta}+\lambda)\max_{1\leq j \leq J_0}\widehat{\sigma}_{b_j}\widehat{\sigma}_{u}}{\min_{1\leq j \leq J_0}\widehat{\sigma}^2_{b_j}} \] for any large $\lambda$. In most applications, however, the SC weights are sparse due to the simplex- or lasso-type constraints. It is known from the sparse linear regression literature that the bound $\mathcal{C}_1$ can be further improved by replacing the factor $\sqrt{d_\beta}$ by $c_3\sqrt{\|\boldsymbol{\beta}_0\|_0}$, where $\|\cdot\|_0$ denotes the number of nonzeros in a vector and $c_3>0$ is some absolute constant (see, e.g., Wainwright_2019_Book). Assuming $\sqrt{\log T_0}\geq 4\sqrt{2}c_3$, we take $\lambda=\sqrt{\log d_\beta\log T_0}/(4c_3)$ and use $\|\widehat{\boldsymbol{\beta}}\|_0$ as a proxy for $\|\boldsymbol{\beta}_0\|_0$, which yields our recommended choice of $\mathcal{C}$ given in (6.9): \[ \mathcal{C}=\mathcal{C}_2:= \frac{\sqrt{\|\widehat{\boldsymbol{\beta}}\|_0\log d_\beta\log T_0}\max_{1\leq j \leq J_0}\widehat{\sigma}_{b_j}\widehat{\sigma}_{u}}{\min_{1\leq j \leq J_0}\widehat{\sigma}^2_{b_j}}. \] Furthermore, when $\widehat{\sigma}_{b_j}^2$ is roughly the same across $j$, then $\mathcal{C}_2$ can be further simplified to \[ \mathcal{C}=\mathcal{C}_3 = \frac{\sqrt{\|\widehat{\boldsymbol{\beta}}\|_0\log d_\beta\log T_0}\;\widehat{\sigma}_{u}}{\min_{1\leq j \leq J_0}\widehat{\sigma}_{b_j}}. \]
While these choices are only rules of thumb justified under specific assumptions on the data generating process, they at least have the correct order of magnitude and are valid at least when $T_0$ is large. For instance, for the cointegrated data considered in our basic setup, it can be shown that, for sufficiently small $c_4>0$, $\widehat{\sigma}^2_{b_j}\geq c_4T_0$ with high probability, and $\widehat{\sigma}_u^2$ is bounded with high probability. Therefore, the suggested $\varrho$ based on $\mathcal{C}_2$ has the order of $T_0^{-1}$, which is the well-known rate of convergence for cointegrated regression.
When SC weights are obtained by matching on both stationary and non-stationary features, the contribution of the stationary components to the concentration of $\widehat{\boldsymbol{\beta}}$ is negligible. In other words, the order of magnitude for the bound on $\|\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}\|_2$ is primarily determined by the non-stationary components. To see this, assume that we match on two features, a non-stationary $\mathbf{B}_{1}=(\mathbf{b}_{1,1}, \cdots, \mathbf{b}_{T_0,1})'\in\mathbb{R}^{T_0\times J_0}$ and a stationary $\mathbf{B}_{2}=(\mathbf{b}_{1,2}, \cdots, \mathbf{b}_{T_0,2})'\in\mathbb{R}^{T_0\times J_0}$. Consider the basic inequality (ref) again. $\widehat{\mathbf{Q}}/T_0$ can be written as $\frac{1}{T_0}\sum_{t=1}^{T_0}\mathbf{b}_{t,1}\mathbf{b}_{t,1}'+\frac{1}{T_0}\sum_{t=1}^{T_0}\mathbf{b}_{t,2}\mathbf{b}_{t,2}'=:\mathrm{I}+\mathrm{II}$. As explained previously, with high probability, $s_{\min}(\mathrm{I}/T_0)$ is bounded from below by $T_0$, up to a sufficiently small constant, whereas $\mathrm{II}/T_0$ is much smaller---under mild conditions, it is bounded with high probability. For the numerator $(\widehat{\boldsymbol{\gamma}}-\boldsymbol{\gamma})/T_0$, we can similarly decompose it into a non-stationary part and a stationary part. By the argument described before, it follows that the stationary part is much smaller than the non-stationary part, leading to a deviation bound for $\widehat{\boldsymbol{\beta}}$ on the order of $T_0^{-1}$. Therefore, we recommend applying the above formulas using the non-stationary features only in practice.
Finally, as a reminder, other choices of $\varrho$ can be proposed based on a slightly different basic inequality: $\|\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0\|_2\leq 2\|\widehat{\boldsymbol{\gamma}}\|_2/s_{\min}(\widehat{\mathbf{Q}})$. It also holds deterministically by optimality of $\widehat{\boldsymbol{\beta}}$. The denominator is the same as before, but the numerator $\widehat{\boldsymbol{\gamma}}$ is not demeaned. When the model is correctly specified ($\boldsymbol{\gamma}=\bm{0}$), the numerator may be small; however, when the model is misspecified, $\widehat{\boldsymbol{\gamma}}$ is not necessarily small, and it may be characterized by its sample analogue. Thus, we can set \[ \varrho=\frac{2\sqrt{d_\beta}\max_{1\leq j\leq J_o}\widehat{\sigma}^2_{b_ju}}{\min_{1\leq j\leq J_0}\widehat{\sigma}^2_{b_j}} \] where $\widehat{\sigma}^2_{b_ju}$ is the estimated covariance between the pseudo-true residual $\mathbf{u}$ and the $j$-th column of $\mathbf{B}$ (the features of the $j$-th control unit). However, compared to the alternatives described in the paper, this bound is generally loose and thus is not recommended.
As described in the main paper, we have an additional adjustment to nonlinear constraints in the simulation (the second term on the right-hand side of (6.10)). In this section we briefly discuss the necessity of this adjustment and the justification of the proposed strategy.
The constraints allowed in this paper is more general than those considered in Cattaneo-Feng-Titiunik_2021_JASA. Specifically, condition (T2.iii) in Cattaneo-Feng-Titiunik_2021_JASA requires the constraint set used in the simulation for in-sample uncertainty quantification be locally equal to the original constraint set in the synthetic control problem. In the notation of this paper, that condition means that, with some high probability, $\Delta^{\star}\cap\mathcal{B}(\bm{0},\varpi_\delta^\star)=\Delta\cap\mathcal{B}(\bm{0},\varpi_\delta^\star)$ for some (small) $\varpi_\delta^\star$-neighborhood of zero. This is generally true when constraints are formed by linear functions. To see this, suppose that there are two donor units, an L1 constraint $\mathcal{W}=\{(w_1, w_2): |w_1|+|w_2|\leq 1\}$ is imposed, and the pseudo-true value is $\mathbf{w}_0=(1,0)$. The constraint set $\Delta^\star$ used in the simulation relies on the estimated weights $\widehat{\mathbf{w}}$, which are generally close (but not exactly equal) to $\mathbf{w}_0$ with high probability. For instance, let $\widehat{\mathbf{w}}=(0.99, 0.01)$, and applying the suggested strategy described in Section 6.2, one correctly finds that two linear constraints among those derived from decomposing the original L1 constraint, $w_1+w_2\leq 1$ and $w_1-w_2\leq 1$, are binding. Thus, one defines $\Delta^\star=\{(\delta_1, \delta_2): (\delta_1+0.99)+(\delta_2+0.01)\leq 0.99+0.01, (\delta_1+0.99)-(\delta_2+0.01)\leq 0.99-0.01\}$, which is exactly the same as $\Delta=\mathcal{W}-\mathbf{w}_0$ locally around $(0,0)$.
However, such an exact equality is generally not possible when constraints are formed by nonlinear functions. Still consider the same example, but an L2 constraint $\mathcal{W}=\{(w_1, w_2): w_1^2+w_2^2\leq 1\}$ is imposed instead. Applying the same strategy, one correctly finds that the constraint should be binding and thus lets $\Delta^\star=\{(\delta_1,\delta_2):(\delta_1+0.99)^2+(\delta_2+0.01)^2\leq 0.99^2+0.01^2\}$, whereas the (centered) original constraint set is $\Delta=\{(\delta_1, \delta_2): (\delta_1+1)^2+\delta_2^2=1\}$. In other words, $\Delta$ and $\Delta^\star$ are two “circles” passing through $(0,0)$. Although they are locally close near $(0,0)$, their boundaries have different tangents at that point, and hence the two sets are not locally identical, no matter how small the neighborhood $\mathcal{B}(\bm{0},\varpi^\star)$ is. In this sense, the results in Cattaneo-Feng-Titiunik_2021_JASA are more applicable to synthetic controls with linear constraints (such as simplex and Lasso). By contrast, the local approximate equality in condition (iii) of Theorem 1 is much weaker and makes the results in this paper applicable to more general cases with possibly nonlinear constraints. Given the fact described above, when there are (binding) nonlinear constraints, we propose to enlarge the feasibility set defined by nonlinear constraints to satisfy the sufficient condition (iii) of Theorem 1.
In this section we briefly discuss the justification for the adjustment in (6.10). Our strategy is motivated by the proof of Lemma 1, which characterizes the distance between $\Delta\cap\mathcal{B}(\bm{0},\varpi_\delta^\star)$ and $\widehat{\Delta}$. For simplicity, assume that there is only one treated unit (so the superscript $[i]$ can be omitted with no confusion).
Suppose that we only have binding inequality constraints $\bm{m}(\boldsymbol{\beta})\leq \bm{0}$ (so $\bm{m}(\boldsymbol{\beta}_0)=\bm{0}$), where $\bm{m}(\cdot)$ is a $d_\leq$-vector of sufficiently smooth functions. For any feasible $\boldsymbol{\beta}\in\mathbb{R}^{d_\beta}$ close to $\boldsymbol{\beta}_0$ such that $\boldsymbol{\beta}-\boldsymbol{\beta}_0\in\Delta\cap\mathcal{B}(\bm{0}, \varpi_\delta^\star)$, \[ \bm{m}(\boldsymbol{\beta})=\bm{m}(\boldsymbol{\beta})-\bm{m}(\boldsymbol{\beta}_0)=\frac{\partial}{\partial\boldsymbol{\beta}'}\bm{m}(\boldsymbol{\beta}_0)(\boldsymbol{\beta}-\boldsymbol{\beta}_0)+\frac{1}{2}(\boldsymbol{\beta}-\boldsymbol{\beta}_0)'\Big[\frac{\partial^2}{\partial\boldsymbol{\beta}\partial\boldsymbol{\beta}'}\bm{m}(\check{\boldsymbol{\beta}})\Big](\boldsymbol{\beta}-\boldsymbol{\beta}_0) \] for some point $\check{\boldsymbol{\beta}}$ between $\boldsymbol{\beta}$ and $\boldsymbol{\beta}_0$, where $\partial^2\bm{m}(\check{\boldsymbol{\beta}})/\partial\boldsymbol{\beta}\partial\boldsymbol{\beta}'$ is a $d_\beta\times d_\beta\times d_\leq$ array, with each sheet the second-order derivative matrix for one constraint function. Thus, $(\boldsymbol{\beta}-\boldsymbol{\beta}_0)'[\partial^2\bm{m}(\check{\boldsymbol{\beta}})/\partial\boldsymbol{\beta}\partial\boldsymbol{\beta}'](\boldsymbol{\beta}-\boldsymbol{\beta}_0)$ is a $d_\leq$-vector, with each element corresponding to $(\boldsymbol{\beta}-\boldsymbol{\beta}_0)'[\partial^2m_\ell(\check{\boldsymbol{\beta}})/\partial\boldsymbol{\beta}\partial\boldsymbol{\beta}'](\boldsymbol{\beta}-\boldsymbol{\beta}_0)$ for $1\leq \ell\leq d_\leq$.
Let $\Gamma_0=\frac{\partial}{\partial\boldsymbol{\beta}}\bm{m}(\boldsymbol{\beta}_0)$ and $\widehat{\Gamma}=\frac{\partial}{\partial\boldsymbol{\beta}}\bm{m}(\widehat{\boldsymbol{\beta}})$. Assume $\Gamma_0$ is invertible in the neighborhood of $\boldsymbol{\beta}_0$. (When $d_\leq< d_\beta$, we can complement $\Gamma_0$ with additional rows and construct an invertible one, which is formalized in the proof of Lemma 1). Then, \[ \Gamma_0^{-1}\bm{m}(\boldsymbol{\beta})-(\boldsymbol{\beta}-\boldsymbol{\beta}_0) =\frac{1}{2}\Gamma_0^{-1}\Big[(\boldsymbol{\beta}-\boldsymbol{\beta}_0)'\frac{\partial^2}{\partial\boldsymbol{\beta}\partial\boldsymbol{\beta}'}\bm{m}(\check{\boldsymbol{\beta}})(\boldsymbol{\beta}-\boldsymbol{\beta}_0)\Big]=:\mathscr{L} \] It can be shown that $\boldsymbol{\lambda}:=\Gamma_0^{-1}\bm{m}(\boldsymbol{\beta})$ is (approximately) in $\widehat{\Delta}$. To see this, $$ \bm{m}(\widehat{\boldsymbol{\beta}}+\boldsymbol{\lambda})\approx\bm{m}(\widehat{\boldsymbol{\beta}})+\frac{\partial}{\partial\boldsymbol{\beta}'}\bm{m}(\boldsymbol{\beta}_0)\boldsymbol{\lambda}=\bm{m}(\widehat{\boldsymbol{\beta}})+(\bm{m}(\boldsymbol{\beta})-\bm{m}(\boldsymbol{\beta}_0))\leq \bm{m}(\widehat{\boldsymbol{\beta}}). $$ This is formalized in the proof of Lemma 1. For the moment, assume $\boldsymbol{\lambda}\in\widehat{\Delta}$. Then,
The errors introduced by “$\approx$” are of smaller order under mild conditions. The above calculation implies that any $\boldsymbol{\beta}-\boldsymbol{\beta}_0\in\Delta\cap\mathcal{B}(\bm{0},\varpi_\delta^\star)$ (approximately) belongs to the following adjusted version of $\widehat{\Delta}$ in (6.7): \[ \Big\{\boldsymbol{\delta}: \bm{m}(\widehat{\boldsymbol{\beta}}+\boldsymbol{\delta})\leq \bm{m}(\widehat{\boldsymbol{\beta}})+\frac{1}{2}(\boldsymbol{\beta}-\boldsymbol{\beta}_0)'\frac{\partial^2} {\partial\boldsymbol{\beta}\partial\boldsymbol{\beta}'}\bm{m}(\boldsymbol{\beta}_0)(\boldsymbol{\beta}-\boldsymbol{\beta}_0)\Big\}. \] Then, for each inequality constraint $m_{\leq,\ell}(\widehat{\boldsymbol{\beta}}+\boldsymbol{\delta})\leq m_{\leq, \ell}(\widehat{\boldsymbol{\beta}})$, the “adjustment” to the upper bound is at most \[ \frac{1}{2}s_{\max}\Big(\frac{\partial^2} {\partial\boldsymbol{\beta}\partial\boldsymbol{\beta}'}m_{\leq,\ell}(\boldsymbol{\beta}_0)\Big)\|\boldsymbol{\beta}-\boldsymbol{\beta}_0\|_2^2. \] Given the high-probability bound $\varrho$ for $\|\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}_0\|_2$, we propose the adjustment in (6.10).
\noindentNotes: Our enlargement strategy used to define $\Delta^\star$ in Section 6.2 can be conceptually understood as follows. A constraint set $\Delta$ can be generally written as $\Delta=\cap_{j=1}^L A_j$ for a sequence of convex sets $A_j$. Each $A_j$ is defined by one inequality. Accordingly, $\widehat{\Delta}$ in (6.7) (with no enlargement) can be written as $\widehat{\Delta}=\cap_{j=1}^L\widehat{A}_j$.
Let $\widehat{A}_{j,\varepsilon_j}$ be an $\varepsilon_j$-enlargement of $\widehat{A}_j$. We can apply Lemma 1 to each pair of $A_j$ and $\widehat{A}_j$, and then $A_j\subseteq\widehat{A}_{j,\varepsilon_j}$ for an appropriate (small) $\varepsilon_j$. For any point $g\in\Delta$, $g\in A_j$ for every $j$. So $g\in\widehat{A}_{j,\varepsilon_j}$. Then, $g\in\cap_{j=1}^L \widehat{A}_{j,\varepsilon_j}$. Since it holds for any $g$, we have $\Delta\subseteq\Delta^\star:=\cap_{j=1}^L \widehat{A}_{j,\varepsilon_j}$. The sufficient condition (iii) in Theorem 1 holds.
In this section, we first define three types of convex optimization problems that we will be relying on (for background knowledge and technical details, see Boyd_2004_BOOK). Second, we illustrate the link between these three families of convex problems. Third, we show that the optimization problems underlying the prediction/estimation and uncertainty quantification problems for SC presented in Section 6 of the main text are quadratically constrained quadratic program (QCQP) and quadratically constrained linear problem (QCLP), respectively, and show how to represent them as second-order cone program (SOCP). Finally, we provide two examples by showing how to write the L1-L2-type and Lasso-type constraints in conic form. These approaches are implemented in our companion general-purpose software \citep*{Cattaneo-Feng-Palomba-Titiunik_2025_JSS}, where we show that they lead to remarkable speed and scalability improvements.
QCQPs and QCLPs. A quadratically constrained quadratic program is an optimization problem of with the following form
where $\mathbf{P}_0, \mathbf{P}_1, \ldots, \mathbf{P}_m \in \mathcal{M}_{n\times n}(\operatorname{\mathbb{R}})$, $ \mathbf{q}_0,\mathbf{q}_1,\ldots,\mathbf{q}_m\in \operatorname{\mathbb{R}}^n$, $\mathbf{x}\in \operatorname{\mathbb{R}}^n$, $\mathbf{F} \in \mathcal{M}_{m \times n}(\operatorname{\mathbb{R}})$, $\mathbf{g} \in \operatorname{\mathbb{R}}^m$, and $r_0,r_1, \ldots, r_m, w \in \operatorname{\mathbb{R}}$. If all the matrices $\mathbf{P}_0,\mathbf{P}_1,\ldots, \mathbf{P}_m$ are positive semi-definite the QCQP is convex. Moreover, if $\mathbf{P}_0=\mathbf{0}$ the QCQP becomes a QCLP. For this reason, in what follows we will restrict our attention to QCQPs as they naturally embed QCLPs.
SOCPs. To define a SOCP, it is necessary to first give the definition of a second-order cone and then introduce the notion of associated generalized inequality.
\uline{Second-order cone definition.} A set $\mathcal{C}$ is called a cone if for every $\mathbf{x} \in \mathcal{C}$ and $\alpha\geq 0$ we have $\alpha \mathbf{x}\in \mathcal{C}$. A set $\mathcal{C}$ is a convex cone if it is convex and a cone, i.e. if $\forall\, \mathbf{x}_1,\mathbf{x}_2 \in \mathcal{C}, $and $\forall\, \alpha_1,\alpha_2 \geq 0$, we have \[\alpha_1 \mathbf{x}_1 + \alpha \mathbf{x}_2 \in \mathcal{C}.\] Now, consider any norm $||\cdot||$ defined on the Euclidean space $\operatorname{\mathbb{R}}^n$. The norm cone associated with the norm $||\cdot ||$ is defined to be the set \[\mathcal{C}=\{(\mathbf{x},t)\in \operatorname{\mathbb{R}}^{n+1}: ||\mathbf{x}|| \leq t\}\] and it is a convex cone by the standard properties of the norms. A second-order cone is the associated norm cone for the Euclidean norm and it is typically defined as
\uline{Generalized inequality.} A cone $\mathcal{C}$ is solid if it has non-empty interior and it is pointed if $\mathbf{x}\in \mathcal{C}, -\mathbf{x}\in \mathcal{C}$ implies that $\mathbf{x}=0$. We say that a cone $\mathcal{C}$ is proper if is it convex, closed, solid, and pointed. Proper cones in the Euclidean space $\operatorname{\mathbb{R}}^n$ are useful because they induce a partial ordering that enjoys almost all the properties of the basic one in $\operatorname{\mathbb{R}}$. Therefore, given a cone $\mathcal{C}\subseteq \operatorname{\mathbb{R}}^n$, we can define the generalized inequality $\preceq_\mathcal{C}$ for any two vectors $\mathbf{x},\mathbf{y}\in \operatorname{\mathbb{R}}^n$ \[\mathbf{x}\preceq_{\mathcal{C}} \mathbf{y} \iff \mathbf{y} -\mathbf{x} \in \mathcal{C}.\] From this definition we can see that quadratic constraints such as $||\mathbf{x}||_2 \leq h$ can be re-written as second-order cone constraints of the form $\mathbf{x}\preceq_\mathcal{C} h$ for some second-order cone $\mathcal{C}$. Note that if $\mathcal{C}=\operatorname{\mathbb{R}}_+^m$ then if $m=1$, $\preceq_\mathcal{C}$ is the standard inequality $\leq$ in $\operatorname{\mathbb{R}}$, whereas if $m>1$, $\preceq_\mathcal{C}$ is the component-wise inequality in $\operatorname{\mathbb{R}}^m$.
\uline{Second-order cone program.} Let $\mathcal{K}$ be a cone such that $\mathcal{K} = \mathbb{R}_+^m \times \mathcal{K}_1 \times \mathcal{K}_2 \times\cdots\times \mathcal{K}_L$ where $\mathcal{K}_l := \{(k_0,\mathbf{k}_1)\in \operatorname{\mathbb{R}}\times \operatorname{\mathbb{R}}^{l}: ||\mathbf{k}_1||_2\leq k_0\}, l= 1,\ldots, L$. Let $\preceq_\mathcal{K}$ be the generalized inequality associated with the cone $\mathcal{K}$. An optimization problem is called second-order cone program if it has the following form
Any QCQP can be converted to a SOCP Boyd_2004_BOOK. In other words, we can always rewrite an optimization problem such as (ref) in the form of (ref). First, we present the general result and then we explain all the necessary steps to reformulate QCQPs as SOCPs. Without loss of generality, assume that $w=0$ in (ref) and, to ease notation, let $m=1$ so that there is only a single quadratic inequality constraint. Moreover, given any positive semi-definite matrix $\mathbf{P}$, let $\mathbf{P}^{1/2}$ be the square root of $\mathbf{P}$, that is the unique symmetric positive semi-definite matrix $\mathbf{R}$ such that $\mathbf{R}\mathbf{R}=\mathbf{R}'\mathbf{R}=\mathbf{P}$. Then for any QCQP the following two formulations are equivalent
We can see that the logic beneath the conversion of a QCQP into a SOCP is to “linearize” all the non-linear terms appearing either in the objective function or in the inequality constraints. The “linearization” step does come at a cost, as it requires the introduction of a slack variable every time we rely on it. Indeed, above we linearized the objective function and the quadratic inequality constraint by introducing two auxiliary slack variables.
More formally, let $\mathbf{x}'\mathbf{P}\mathbf{x}$ be a symmetric positive semi-definite quadratic form and consider the constraint $\mathbf{x}'\mathbf{P}\mathbf{x}\leq y$. Then
We illustrate the approach described above for the L1-L2 constraint and the Lasso constraint. To ease notation, we do not consider the regularization to the local geometry of $\Delta$. Note that simplex, ridge, or least squares are particular cases of L1-L2. Throughout this section, we assume that $T_i\equiv T_0$ for all $i\in\mathcal{E}$.
\paragraph{L1-L2-type $\mathcal{W}$.} Consider first the prediction/estimation SC optimization problem, which relies on the following program:
where, as always, $\geq$ is understood as a component-wise inequality for vectors ($\mathbf{w}\in\operatorname{\mathbb{R}}^{J_0\cdot J_1}$). First, notice that the non-convex constraints $||\mathbf{w}^{[i]}||_1=1,i=1,\ldots,J_1$ can be replaced with the convex constraints $\mathbf{1}' \mathbf{w}^{[i]}=1,i=1,\ldots,J_1$ because of the non-negativity constraint on the elements of $\mathbf{w}$. Then, we can cast (ref) as a SOCP as follows
where $\mathcal{K}=\mathcal{C}_1\times\mathcal{C}_2\times\mathcal{C}_3^{J_1}\times\mathcal{C}_4^{J_1}=\mathbb{R}_+^{J_0\cdot J_1} \times \mathcal{K}_{\tilde{T}\cdot M+1}\times \operatorname{\mathbb{R}}_+^{J_1} \times \mathcal{K}_{J_0 +1}^{J_1}$ is the conic constraint for this program.
For uncertainty quantification, we need to solve the optimization problem underlying (6.5) in the main text. We discuss the lower bound only. Recalling that $\boldsymbol{\beta}=(\mathbf{w}',\mathbf{r}')'$, we have
We can cast the SC optimization problem in (ref) in conic form as follows:
where $\mathcal{K}=\mathcal{C}_1\times\mathcal{C}_2\times\mathcal{C}_3\times\mathcal{C}_4\times\mathcal{C}_5=\operatorname{\mathbb{R}}_+\times \operatorname{\mathbb{R}}_+^{J_0\cdot J_1}\times\operatorname{\mathbb{R}}_+^{J_1}\times \mathcal{K}_{1+J_0}^{J_1} \times \mathcal{K}_{1+(J_0+KM)\cdot J_1}$ is the conic constraint for this program, $\mathbf{a}=-2('\mathbf{Q}\widehat{\boldsymbol{\beta}} + \mathbf{G}^\star)'$, and $f=\widehat{\boldsymbol{\beta}}'\mathbf{Q}\widehat{\boldsymbol{\beta}} + 2\mathbf{G}^\star\widehat{\boldsymbol{\beta}}$.
\paragraph{Lasso-type $\mathcal{W}$.} We show how to write the QCQP as a SOCP when $\mathcal{W}$ has a lasso-type constraint. In this case, the SC weight construction (3.1) has the form:
We can write the optimization problem in (ref) as a SOCP of the following form
where $\mathcal{K}=\mathcal{C}_1\times\mathcal{C}_2^{J_1}\times\mathcal{C}_3\times\mathcal{C}_4=\mathcal{K}_{1+T_0\cdot M \cdot J_1}\times \operatorname{\mathbb{R}}_+^{J_1}\times \mathbb{R}_+^{J_0\cdot J_1}\times \mathbb{R}_+^{J_0\cdot J_1}$ is the conic constraint for this program and $\mathbf{z}:=(\mathbf{z}_1',\ldots,\mathbf{z}_{J_1}')'$.
For uncertainty quantification, we need to solve the optimization problem underlying (6.5) in the main text. Here we discuss the lower bound only for brevity. Recalling that $\boldsymbol{\beta}=(\mathbf{w}',\mathbf{r}')'$, we have
We can cast the SC optimization problem in (ref) in conic form as follows:
where $\mathcal{K}=\mathcal{C}_1\times\mathcal{C}_2^{J_1}\times\mathcal{C}_3\times\mathcal{C}_4\times\mathcal{C}_5=\operatorname{\mathbb{R}}_+\times\operatorname{\mathbb{R}}_+^{J_1}\times \mathbb{R}_+^{J_0\cdot J_1}\times \mathbb{R}_+^{J_0\cdot J_1} \times \mathcal{K}_{1+(J_0+KM)\cdot J_1}$ is the conic constraint for this program, $\mathbf{a}=-2('\mathbf{Q}\widehat{\boldsymbol{\beta}} + \mathbf{G}^\star)'$, and $f=\widehat{\boldsymbol{\beta}}'\mathbf{Q}\widehat{\boldsymbol{\beta}} + 2\mathbf{G}^\star\widehat{\boldsymbol{\beta}}$.
In this section, we first describe the variables in the Billmeier-Nannicini_2013_RESTAT (BN, henceforth) dataset and then go through the details of our empirical specification.
The original BN dataset contains data on some economic and political variables for 180 countries, over a period of time spanning from 1960 to 2005.\footnote{We downloaded the dataset from the Harvard Dataverse at \url{https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/28699}.} In detail, the variables available in the dataset are:
Despite having six candidate variables to match on, we end up matching on at most two variables because of missing data. Specifically, we always match on real GDP per-capita, whereas in the robustness exercise in Supplemental Appendix (ref) we also match on investment-to-GDP ratio.
A second difference of our final dataset from the original one used in BN is the final pool of countries - treated and donors - on which we conduct the analysis. In particular, we adopt the following criteria to select the countries to be included in our final dataset as either donors or treated units:
Table (ref) shows the final set of countries we select, together with their treatment date.
In this section, we describe all the details of the empirical application presented in the main text.
\paragraph{Constraint type.} Our preferred specification uses the L1-L2 constraint, i.e., \[\mathcal{W}_{\mathtt{L1-L2}} = \bigtimes_{i=1}^{J_1} \left\{\mathbf{w}^{[i]} \in \mathbb{R}_+^{J_0}:||\mathbf{w}^{[i]} ||_1 = 1, ||\mathbf{w}^{[i]} ||_2 \leq Q^{[i]}\right\},\] whereas the results in the Supplemental Appendix Section (ref) and (ref) use simplex- and Ridge-type constraints of the form \[\mathcal{W}_{\mathtt{S}} = \bigtimes_{i=1}^{J_1} \left\{\mathbf{w}^{[i]} \in \mathbb{R}_+^{J_0}:||\mathbf{w}^{[i]} ||_1 = 1\right\},\qquad \mathcal{W}_{\mathtt{R}} = \bigtimes_{i=1}^{J_1} \left\{\mathbf{w}^{[i]} \in \mathbb{R}^{J_0}: ||\mathbf{w}^{[i]} ||_2 \leq Q^{[i]}\right\}.\] Table (ref) shows the effective values for $Q^{[i]}, i=1,\ldots, J_1$ that we compute in our empirical application. Further below we explain in greater detail how these regularization parameters are computed in practice.
\paragraph{Selected features.} Our main specification uses only one feature ($M=1$)--the logarithm of real GDP per-capita--, uses the identify weighting matrix, and includes a constant term, that is \[\mathbf{B}^{[i]} =
, \quad \mathbf{C}^{[i]} = \mathbf{1}_{T_i}, \quad \mathcal{R}=\operatorname{\mathbb{R}}^{J_1}, \quad \mathbf{V}^{[i]} = \mathbf{I}_{T_i}, \quad i = 1,\ldots, J_1,\] where $\mathbf{Y}_j = (Y_{j1},\ldots, Y_{jT_i})', j=1,\ldots,J_0$ is the pre-treatment log-GDP per-capita of the $j$-th donor, $\mathbf{1}_{T_i}$ is a $T_i\times 1$ vector of ones, and $\mathbf{I}_{T_i}$ is the $T_i\times T_i$ identify matrix.
In Supplemental Appendix Section (ref), we present results using two features ($M=2$), where we also match on the investment-to-GDP ratio and control for a feature-specific constant term and linear trend, i.e. \[\mathbf{B}^{[i]} =
, \quad \mathbf{C}^{[i]} =
, \quad \mathcal{R}=\operatorname{\mathbb{R}}^{2\cdot J_1}, \quad \mathbf{V}^{[i]} =
,\quad i = 1,\ldots, J_1,\] where $\mathbf{IR}_j = (IR_{j1},\ldots, IR_{jT_i})', j=1,\ldots,J_0$ is the pre-treatment investment-to-GDP ratio of the $j$-th donor, $\mathbf{S}^{[i]}_Y = \mathrm{diag}(\widehat{\sigma}^{-1}_{Y,1},\ldots, \widehat{\sigma}^{-1}_{Y,T_i})$, $\mathbf{S}^{[i]}_{IR} = \mathrm{diag}(\widehat{\sigma}_{IR,1}^{-1},\ldots, \widehat{\sigma}^{-1}_{IR,T_i}),$ with \[\widehat{\sigma}_{W,t} = \left(\frac{1}{J_0-1} \sum_{j=1}^{J_0}(W_{jt} - \overline{W}_t)^2\right)^{1/2}, \quad \overline{W}_t = \frac{1}{J_0}\sum_{i=1}^{J_0}W_{jt}, \quad t=1,\ldots, T_i, \quad W\in\{Y,IR\},\] and $\mathrm{diag}(\mathbf{x})$ yields a square diagonal matrix with the elements of $\mathbf{x}$ on its main diagonal.
\paragraph{Tuning parameters.} Regarding the choice of $Q^{[i]}$, it is well-established that the Ridge regression problem can be equivalently expressed both as an unconstrained penalized optimization problem and as a constrained optimization problem. For simplicity, assume \( \mathbf{C} \) is not included and \( M = 1 \). The two Ridge-type optimization formulations are as follows: \[ \widehat{\mathbf{w}}^{[i]} = \arg \min_{\mathbf{w}^{[i]} \in \mathcal{W}} (\mathbf{A}^{[i]} - \mathbf{B}^{[i]}\mathbf{w}^{[i]})' \mathbf{V}^{[i]} (\mathbf{A}^{[i]} - \mathbf{B}^{[i]}\mathbf{w}^{[i]}) + \lambda^{[i]} ||\mathbf{w}^{[i]}||_2^2, \] where \( \lambda^{[i]} \geq 0 \) is a regularization parameter, and \[ \widehat{\mathbf{w}}^{[i]} = \arg \min_{\mathbf{w}^{[i]} \in \mathcal{W}, ||\mathbf{w}^{[i]}||_2^2 \leq \left(Q^{[i]}\right)^2} (\mathbf{A}^{[i]} - \mathbf{B}^{[i]}\mathbf{w}^{[i]})' \mathbf{V}^{[i]} (\mathbf{A}^{[i]} - \mathbf{B}^{[i]}\mathbf{w}^{[i]}), \] where \( Q^{[i]} \geq 0 \) is an explicit upper bound on the norm of \( \mathbf{w}^{[i]} \). Under the assumption of Gaussian errors, an optimal choice of the regularization parameter \( \lambda^{[i]} \) for risk minimization, as suggested by Hoerl-Kannard-Kent_1975_ridge, is: \[ \lambda^{[i]} = \frac{J_0 (\widehat{\sigma}^{[i]}_{\text{OLS}})^2}{\|\widehat{\mathbf{w}}^{[i]}_{\text{OLS}}\|_2^2}, \] where \( (\widehat{\sigma}^{[i]}_{\text{OLS}})^2 \) and \( \widehat{\mathbf{w}}^{[i]}_{\text{OLS}} \) are the estimates of the residual variance and the coefficients from the ordinary least squares (OLS) regression of $\mathbf{A}^{[i]}$ onto $\mathbf{B}^{[i]}$, respectively. Given the two optimization problems above, there exists a one-to-one correspondence between \( \lambda^{[i]} \) and \( Q^{[i]} \). For example, assuming the columns of \( \mathbf{B}^{[i]} \) are orthonormal, the closed-form solution for the Ridge estimator is: \[ \widehat{\mathbf{w}}^{[i]} = (\mathbf{I} + \lambda^{[i]} \mathbf{I})^{-1} \widehat{\mathbf{w}}^{[i]}_{\text{OLS}}, \] and if the constraint on the \( \ell_2 \)-norm is active, we have \( Q^{[i]} = \|\widehat{\mathbf{w}}^{[i]}\|_2 = \|\widehat{\mathbf{w}}^{[i]}_{\text{OLS}}\|_2 / (1 + \lambda^{[i]}) \). When more than one feature is considered (i.e., \( M > 1 \)), we compute the constraint size \( Q^{[i]}_\ell \) for each feature \( \ell = 1, \ldots, M \), and then choose \( Q^{[i]} \) as the most restrictive constraint to promote shrinkage of \( \mathbf{w}^{[i]} \): \[ Q^{[i]} := \min_{\ell=1, \ldots, M} Q^{[i]}_\ell. \]
\paragraph{In-sample Uncertainty.} In order to quantify the in-sample uncertainty from estimating the SC weights, we need to construct the bounds $\underline{M\mkern-4mu}\mkern4mu _{\mathtt{in}}$ and $\overline{\mkern-3muM\mkern-1.5mu}_{\mathtt{in}}$ on $\mathbf{p}_\tau'(\widehat\boldsymbol{\beta}-\boldsymbol{\beta}_0)$. The following strategy is adopted. First, we treat the synthetic control weights as possibly misspecified, thus estimating both the first and second conditional moments of the pseudo-true residuals $\mathbf{u}$. The conditional first moment $\mathbb{E}[\mathbf{u}\,|\,\mathscr{H}]$ is estimated feature-by-feature using a linear-in-parameters regression of the residual $\widehat{\mathbf{u}} = \mathbf{A} - \mathbf{B}\widehat{\mathbf{w}} -\mathbf{C}\widehat{\mathbf{r}}$ on $\mathbf{B}$ and the first lag of $\mathbf{B}$, whereas the conditional second moment $\mathbb{V}[\mathbf{u}\,|\, \mathscr{H}]$ is estimated with an HC1-type estimator. We then draw $S=200$ i.i.d. random vectors from the Gaussian distribution $\mathsf{N}(0,\widehat{\boldsymbol{\Sigma}})$, conditional on the data, to simulate the criterion function ${\ell^\star_{(s)}(\boldsymbol{\beta}-\boldsymbol{\beta}_0):=(\boldsymbol{\beta}-\boldsymbol{\beta}_0)'\widehat{\mathbf{Q}}(\boldsymbol{\beta}-\boldsymbol{\beta}_0)-2\mathbf{G}_{(s)}'(\boldsymbol{\beta}-\boldsymbol{\beta}_0)}$, $s=1,\ldots, 200$, and solve the following optimization problems
where $\Delta^\star$ is constructed as explained in Section 6.1. Finally, $\underline{M\mkern-4mu}\mkern4mu _{\mathtt{in}}$ is the $(\alpha_1/2)-$quantile of $\{l_{(s)}\}_{s=1}^S$ and $\overline{\mkern-3muM\mkern-1.5mu}_{\mathtt{in}}$ is the $(1-\alpha_1/2)-$quantile of $\{u_{(s)}\}_{s=1}^S$, where $\alpha_1$ is set to 0.05.
\paragraph{Out-of-sample Uncertainty.} In order to quantify the out-of-sample uncertainty from the stochastic error in the post-treatment period, we need to construct the bounds $\underline{M\mkern-4mu}\mkern4mu _{\mathtt{out}}$ and $\overline{\mkern-3muM\mkern-1.5mu}_{\mathtt{out}}$ on the out-of-sample error $e_\tau$ (associated with the $\tau$ prediction). We employ the non-asymptotic bounds described in (6.6), assuming that $e_\tau-\mathbb{E}[e_\tau|\mathscr{H}]$ is sub-Gaussian conditional on $\mathscr{H}$. Then, we take \[\underline{M\mkern-4mu}\mkern4mu _{\mathtt{out}}:=\mathbb{E}[e_\tau|\mathscr{H}]-\sqrt{2\sigma_{\mathscr{H}}^2\log(2/\alpha_2)} \qquad\text{and}\qquad \overline{\mkern-3muM\mkern-1.5mu}_{\mathtt{out}}:=\mathbb{E}[e_\tau|\mathscr{H}]+\sqrt{2\sigma_{\mathscr{H}}^2\log(2/\alpha_2)},\] We set $\alpha_2=0.05$, and the conditional mean $\mathbb{E}[e_\tau|\mathscr{H}]$ and the sub-Gaussian parameter $\sigma_{\mathscr{H}}$ are parametrized and estimated by a linear-in-parameters regression of the pre-treatment residuals on $\mathbf{B}$.
Finally, the prediction intervals for the counterfactual outcome and the treatment effect of interest are given by \[\left[\mathbf{p}_\tau'\widehat{\boldsymbol{\beta}}-\overline{\mkern-3muM\mkern-1.5mu}_{\mathtt{in}}+\underline{M\mkern-4mu}\mkern4mu _{\mathtt{out}}; \: \mathbf{p}_\tau'\widehat{\boldsymbol{\beta}}-\underline{M\mkern-4mu}\mkern4mu _{\mathtt{in}}+\overline{\mkern-3muM\mkern-1.5mu}_{\mathtt{out}}\right] \quad\text{and} \quad \left[\widehat{\tau}+\underline{M\mkern-4mu}\mkern4mu _{\mathtt{in}}-\overline{\mkern-3muM\mkern-1.5mu}_{\mathtt{out}};\: \widehat{\tau}+\overline{\mkern-3muM\mkern-1.5mu}_{\mathtt{in}}-\underline{M\mkern-4mu}\mkern4mu _{\mathtt{out}}\right],\] respectively.
\paragraph{Other assumptions.} Throughout all our specifications, we maintain the assumptions that (i) there is no anticipation of the treatment and (ii) $\mathbf{A}$ and $\mathbf{B}$ form a cointegrated system. When (i) is relaxed to allow for anticipation results remain qualitatively the same.