Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
126,241 characters · 13 sections · 132 citation commands
Robust Inference when Nuisance Parameters may be Partially Identified with Applications to Synthetic Controls
\onehalfspacing \affil[$\dagger$]{ Rutgers University, 75 Hamilton St, 08901, New Brunswick, New Jersey, United States\nopagebreak E-mail address: [email removed]}
In the context of method of moments estimation, if the moment conditions jointly identify a subvector of the parameters, then a standard estimation method such as General Method of Moments (GMM) is generally a consistent estimator for that subvector. However, if the remaining subvector is unidentified, then this remaining parameter cannot be consistently estimated and as a result the estimates for the identified subvector are not asymptotically normal (see, for example, AndrewsCheng2012), which significantly complicates inference. Additionally, if a nuisance parameter is at the boundary of the parameter space or close to the boundary relative to the sample size, then this can also result in our estimates for the parameter of interest not being asymptotically normal (see, for example, Andrews1999 and ConstrainedMEstimation). Lastly, if the full vector is high-dimensional, this can complicate standard asymptotic normality results even when the subvector we would like to perform inference on is low-dimensional. I propose an estimation method that aims to simultaneously overcome these complications to obtain an estimate for an identified parameter of interest that is asymptotically normal even when a nuisance parameter is partially identified, on or near the boundary of the parameter space, and, in some cases, high-dimensional. The procedure I propose can be decomposed into three steps. In the first step, regularized estimates of all parameters are found by minimizing a penalty function subject to the constraint that the sample moment conditions are close to zero. The primary purpose of this penalty function is to make the estimated nuisance parameter converge to a unique element of the identified set. We must therefore choose which element of the identified set we would like our estimate to converge to. Since this choice affects the asymptotic variance of our subsequent estimate of the parameter of interest, I base the penalty function on an estimate of the asymptotic variance as a function of the nuisance parameter, provided that this asymptotic variance function can be estimated sufficiently accurately. In cases where the asymptotic variance cannot be accurately estimated, such as when using time-series or panel data and the degree of temporal dependence is high relative to the number of time periods, I discuss alternative ways of choosing the penalty function in Section (ref). The second step is to use Neyman orthogonalization to construct a set of moment conditions that are orthogonal with respect to the nuisance parameter. This involves introducing a new parameter that is chosen to make the derivative of the moment conditions with respect to the original nuisance parameter equal to zero. After the estimated nuisance parameters have been plugged into the orthogonal moment conditions, the third step is to use these moment conditions to re-estimate the parameter of interest. In Section (ref), I give two suggestions of how this can be done and provide asymptotic normality results for both methods. My primary application of this approach is as a Synthetic Control Estimator (SCE) which gives an asymptotically normal estimate of the average treatment effect on the treated (ATT). Several other inference methods have been proposed for SCEs including the placebo method of Abadie2010, the subsampling method of Li2020, the cross-fitting t-test method of t-test2022, the conformal inference method of ConformalInference, and the End-of-Sample Instability Test, originally introduced by Andrews2003 and applied to SCE by CaoDowd2019. Additionally, CARVALHO2018 and SyntheticIV both provide asymptotic normality for estimates of the ATT using their versions of the SCE and several papers proposing methods closely related to the SCE have also provided asymptotic normality results for the estimated ATT. Arkhangelskyetal2021 introduce the Synthetic Difference-in-Differences estimator, which involves a weighted Difference-in-Differences regression. They establish consistency and asymptotic normality of the estimated ATT. However, the results of Arkhangelskyetal2021 and SyntheticIV rely on the Euclidean norm of the control weights converging to zero at a sufficiently fast rate and CARVALHO2018 and t-test2022 impose conditions for their estimators to be asymptotically unbiased that I do not. In section (ref), I discuss these alternatives further and compare my method with many of these approaches in simulations and show that it usually controls size and has high power relative to the other methods that control size. A large body of work considers inference in cases where a set of moment equations only partially identifies a vector of parameters (e.g., ChernozhukovHongTamer2007 and RomanoShaikh2010). Hansen1996 offers a method for inference when a nuisance parameter is unidentified under the null hypothesis, but it requires simulating the sampling distribution of the estimated nuisance parameter. CHAUDHURI2011 provide an inference method for GMM estimators when a nuisance parameter may be weakly identified, and AndrewsCheng2012 propose an inference method for extremum estimators when a subvector may be weakly identified. HanMcCloskey2019 generalize AndrewsCheng2012's results to the case where the entire vector is allowed to be weakly identified by introducing a method of reparameterization. cox2022weak also builds on this work for the case of, possibly constrained, minimum distance estimators. However, these methods generally do not allow the vector of parameters to be high-dimensional. Additionally, by only considering cases where the parameter of interest is identified, we can focus on conducting the estimation in a way that makes inference simpler for that parameter (by making its estimated values asymptotically normal) and possibly in a way that makes the estimates of that parameter more precise (by minimizing its asymptotic variance). On the other hand, methods that conduct inference on the whole identified set have the advantage that they can be used when both the parameter of interest and the nuisance parameter are partially identified. This work also extends the literature on the Neyman Orthogonalized Score. While the technique dates back to Neyman1959, several more recent papers have used it as a way to achieve asymptotic normality after obtaining an estimate of a high-dimensional nuisance parameter using a regularization penalty or machine learning technique (e.g., belloni2018highdimensional, NingLiu2017, ChernozhukovHansenSpindler2015, BelloniChernozhukovHansen2014, and ChernozhukovDoubleDebiased). Many of these methods estimate the nuisance parameter with LASSO. While this can be a powerful technique when the nuisance parameter is high-dimensional but has a sparse, point-identified value, if the nuisance parameter is partially identified, the LASSO penalty may often be insufficient for the estimates to converge to a specific vector. I help extend this literature by showing how the Neyman Orthogonalized Score can be applied in cases where the nuisance parameter is partially identified, as well as giving results that combine this literature with the fixed-smoothing asymptotics literature in section (ref). ChernozhukovDoubleDebiased's Neyman Near-Orthogonal Score is most closely related to my approach, as they use a regularized estimate of the parameter that is introduced to orthogonalize the moment conditions to choose among the values that make the moment conditions nearly orthogonal. The literature on optimal instruments is also related, and in particular, SinghHosanagarGandhi2020 take a similar approach as I do here, where they choose the function of the instruments that minimizes the estimated asymptotic variance for the parameter of interest. In section (ref), I give an example of how my method can be applied to the optimal choice of instruments. In Section (ref), I discuss the construction and use of the orthogonalized moment conditions. I discuss this first to show what properties we want the regularized estimates of the nuisance parameters to satisfy, and how the limiting values of the estimated nuisance parameters influence the asymptotic variance for the estimated parameter of interest. In Section (ref), I then show when the regularized estimates of parameters satisfy these conditions and discuss how to adjust the procedure to handle cases where the asymptotic variance cannot be consistently estimated. In Section (ref), I show when the high-level conditions in the previous sections are satisfied for a SCE under a linear factor model structure. I apply this SCE by replicating the work of Andersson2019 estimating the effect of Sweden's carbon pricing policies on CO2 emissions. My method finds that their results are statistically significant at commonly used levels, whereas several other methods do not. I also conduct simulations to compare the performance of my inference method with existing approaches for SCEs. Lastly, in Section (ref) I discuss other applications, extensions, and limitations of the method.
I assume that the researcher has a set of moment conditions they wish to use that are a function of a vector of parameters $\theta = (\beta, \delta)$ where the subvector $\beta \in B \subseteq \mathbb{R}^p$ is the parameter of interest and $\delta \in D_n \subseteq \mathbb{R}^J$ is a nuisance parameter. I let $\Theta_n = B \times D_n$ equal the parameter space for the entire vector. I use the subscript $n$ to highlight the fact that the parameter space for the nuisance parameter may change as the sample size grows. This is particularly relevant when $\delta$ is high-dimensional, so $J$ is allowed to grow with $n$. Then I let $g(\theta) = E[\sum_{i=1}^n g_i(\theta)/n]$ and $\hat{g}(\theta) = \sum_{i=1}^n g_i(\theta)/n$ denote $Q$-dimensional vectors for the population and sample moment conditions with a sample size of $n$. Note that since I wish to allow for cases where observations are not identically distributed, I allow $g$ to change with $n$. I am interested in the case where these moment conditions jointly identify the true value of the parameter of interest $\beta$, but may only partially identify $\delta$. The true value of the parameter of interest and the identified set of the nuisance parameter may also change with the sample size, so I denote the identified set as
where $\beta_{0,n}$ is the true value of $\beta$ and $D_{0,n}$ is the identified set for $\delta$. Allowing for drifting sequences of $\beta_{0,n}$ is useful for multiple reasons. First, it is useful for analyzing whether the method is robust to $\beta_{0,n}$ being close to the boundary of the parameter space and being close to values where $\delta$ is unidentified. Additionally, in the SCE application, $\beta_{0,n}$ corresponds to the average of a sequence of dynamic treatment effects, so we want to allow that average to change as our sample size changes. Using this set of the moment conditions, the orthogonalized moment conditions are given by
where $\eta \in H \subseteq \mathbb{R}^{m \times Q}$ is an additional nuisance parameter and $m$ is the number of orthogonalized moment conditions. Generally, I recommend having $\eta$ be unconstrained so that $H = \mathbb{R}^{m \times Q}$, however in specific applications such as the SCE case, there can be reason to constrain $\eta$ for achieving identification as I discuss in Section (ref).\footnote{I focus on cases where $Q$ is fixed. When the number of moment conditions is also large so that $vec(\eta)$ is high-dimensional, there may be a benefit to constraining $\eta$, similarly to the benefits of constraining $\delta$ when it is high-dimensional as discussed in Remark \hyperref[R3.1]{3.1} below.} Because the orthogonalized moments are linear combinations of the original moment conditions, there is no reason to choose $m > Q$. Furthermore, if the $q$-th element of the original moment conditions $g_q(\theta)$ does not vary with $\beta$ at all, then $g_q(\beta,\delta_{0,n}) = 0$ for any $\delta_{0,n} \in D_{0,n}$. Therefore, if $m$ is greater than the number of elements of $g$ which are nontrivial functions of $\beta$, the elements of $M(\beta,\delta_{0,n},\eta) = \eta g(\beta,\delta_{0,n})$ are linearly dependent functions of the elements of $g(\beta,\delta_{0,n})$, for any fixed values of $\eta$ and $\delta_{0,n} \in D_{0,n}$. This is relevant because the orthogonalized moment conditions are used to estimate $\beta$ after plugging in values of the nuisance parameters. Therefore, I set $m$ equal to the number of moment conditions in $g$ that are non-trivial functions of $\beta$. We want to choose $\eta$ so that $M(\theta,\eta)$ is insensitive to $\delta$. The first step in the estimation procedure is to obtain initial estimates of the parameters by picking values that minimize a penalty function among all the values that make the sample moment conditions close to zero, which I discuss how to do in Section (ref). Let $\hat{\theta} = (\hat{\beta},\hat{\delta})$ and $\hat{\eta}$ be equal to an initial regularized estimates of the parameters, and suppose that the distance from our estimates to the values $\theta_{0,n} = (\beta_{0,n},\delta_{0,n})$ and $\eta_{0,n}$ is converging to zero as $n \rightarrow \infty$.\footnote{Since the dimension of $\delta$ may be growing, which metrics this holds under is key for the results below. Assumption \hyperref[Ass2.1]{2.1} contains the details on under exactly which metrics this convergence must hold.} Then we want the sequence of matrices $\eta_{0,n}$ to satisfy
so that these moment conditions are not sensitive to $\delta$ at $(\theta_{0,n},\eta_{0,n})$. Generally, there may exist many choices of $\eta$ that satisfy this condition, particularly if $rank(\partial_\delta g(\theta_{0,n})) < Q$ which may happen when $\delta$ is partially identified. Therefore, $\eta$ itself can be thought of as a partially identified nuisance parameter. By construction, $\partial_{vec(\eta)} M(\theta_{0,n},\eta) = 0$ for any $\theta_{0,n} \in \Theta_{0,n}$ because $g(\theta_{0,n}) = 0$. Hence, $M(\theta_{0,n},\eta)$ is also orthogonal with respect to $\eta$. This allows $\eta$ to be estimated in the same manner as $\delta$, where a regularization penalty is used to make $\hat{\eta}$ converge to a specific element of the set $H_{0,n} \coloneqq \{\eta \in H : \partial_\delta M(\theta_{0,n}, \eta) = 0\}$. Note that this set depends on which $\delta_{0,n} \in D_{0,n}$ is selected, so it is necessary to either estimate $\delta$ first or estimate $\delta$ and $\eta$ jointly. The penalty function is also used to select $\hat{\eta}$ from among all the values of $\eta \in H$ which make the sample moment conditions approximately orthogonal with respect to $\delta$. This is similar to the Neyman Near-Orthogonal Score introduced by ChernozhukovDoubleDebiased, although they handle the estimation of $\delta$ differently since they impose that the original nuisance parameter is point-identified. Another property we want $\eta_{0,n}$ to satisfy is for $M(\beta,\delta_{0,n},\eta_{0,n}) = \eta_{0,n} g(\beta,\delta_{0,n})$ to identify $\beta$. However, this can be handled by choosing our penalty function so that it diverges to infinity at values of $\eta$ that make $\beta$ unidentified. This naturally arises when the penalty function is based on the asymptotic variance for our estimate of $\beta$, since the expression for the asymptotic variance generally diverges as $\eta$ approaches a point where $\beta$ is unidentified. After obtaining estimates of the nuisance parameters $\hat{\delta}$ and $\hat{\eta}$, they are plugged into the sample orthogonalized moments where
Because of equation (ref), under suitable conditions, the sample moment conditions are not sensitive to the values of the nuisance parameters when $\beta = \beta_{0,n}$ and $(\hat{\delta},\hat{\eta})$ is close to $(\delta_{0,n},\eta_{0,n})$. In my asymptotic results, I focus on the cases where the dimensions of the parameter of interest and the moment conditions are fixed but the dimension of $\delta$ may either be fixed or growing with the sample size. This is relevant to the SCE case where $\delta$ is the vector of control weights, which may be of a similar size to the number of time periods. I first give high-level conditions under which $\hat{M}(\beta_{0,n}, \hat{\delta},\hat{\eta})$ is asymptotically equivalent to $\hat{M}(\beta_{0,n},\delta_{0,n},\eta_{0,n})$.
Assumption 2.1\phantomsection As $ n\rightarrow \infty$ while $p$ and $Q$ are fixed and either $J$ fixed or $J \rightarrow \infty$, we have that
For Assumption \hyperref[Ass2.1]{2.1.2}, I show in Appendix \hyperref[ApB]{B}, that if a triangular array $\{X_{i}\}_{n \in \mathbb{N}}$ where $X_i = \{ X_{i1},...,X_{iJ}\}$ is $\alpha$-mixing with exponential speed, is mean-invariant, has uniformly bounded fourth moments, and has an exponential-type bound on the tails of their distributions, then $$\max_{1 \leq j \leq J} | E[\sum_{i=1}^n X_{ij}/n] - \sum_{i=1}^n X_{ij}/n| = O_p(\log (J)/\sqrt{n})$$ when $n,J \rightarrow \infty$ with $J/n^\gamma \rightarrow 0$ for some $\gamma > 0$. This allows for cases where there is a significant degree of dependence across the observations and $J$ is growing faster than $n$. This is important in some applications, such as the SCE case, where there may be a significant degree of temporal dependence and the length of the control weights (which corresponds to $J$) may be greater than $n$. Assumption \hyperref[Ass2.1]{2.1.3} imposes that there is a bound on how locally convex or concave $g$ is at $\theta_{0,n}$, which holds trivially when $g$ is linear in $\theta$. Assumption \hyperref[Ass2.1]{2.1.1} imposes that distances between $\hat{\delta}$ and $\delta_{0,n}$ as well as between $\hat{\eta}$ and $\eta_{0,n}$ using the L1 norm are converging to zero at the given rates. For $\hat{\delta}$, the faster that its dimension is growing, the faster its rate of convergence must be. Furthermore, an additional condition on its rate of convergence must be imposed when the moment conditions are a non-linear function of $\delta$, as shown in Assumption \hyperref[Ass2.1]{2.1.4}.\footnote{The reason why weaker conditions are needed when $\hat{g}$ is linear in $\delta$ is because in this case, making $\partial_\delta \hat{M}(\beta_{0,n},\delta_{0,n},\eta_{0,n})$ close to zero makes $\partial_\delta \hat{M}(\beta_{0,n}, \delta,\eta_{0,n})$ close to zero for any $\delta$. In cases where it is hard to achieve a rate of convergence for the nuisance parameters that is faster than $n^{-1/4}$, mackey18 shows that making the moment conditions $h$-th order orthogonal can allow this condition to be weakened to $o_p(n^{-1/(2h+2)})$.} I show how to obtain $\hat{\delta}$ and $\hat{\eta}$ so that Assumptions \hyperref[Ass2.1]{2.1.1} and \hyperref[Ass2.1]{2.1.4} are satisfied under plausible conditions in Section (ref). Together with the orthogonality condition, this gives the following adaptivity condition. Lemma 2.1 (Adaptivity Condition)\phantomsection Suppose $(\beta_{0,n},\delta_{0,n}) \in \Theta_{0,n}$, $\eta_{0,n}$ satisfies equation (ref), and Assumption \hyperref[Ass2.1]{2.1} holds. Then as $n \rightarrow \infty$ with $p$ and $Q$ fixed with either $J$ fixed or $J \rightarrow \infty$, $$\sqrt{n}(\hat{M}(\beta_{0,n}, \hat{\delta},\hat{\eta}) - \hat{M}(\beta_{0,n},\delta_{0,n},\eta_{0,n})) = o_p(1).$$ Because of this adaptivity condition, an estimator of $\beta$ using $\hat{M}(\beta, \hat{\delta},\hat{\eta})$ is asymptotically equivalent to an estimator using $\hat{M}(\beta,\delta_{0,n},\eta_{0,n})$. One way to use these orthogonalized moment conditions to estimate $\beta$ is via a GMM estimator:
where $W_n$ is a $m \times m$ weighting matrix. As is common for GMM estimators, when a consistent estimator of the asymptotic variance of $\sqrt{n} \hat{M}(\theta_{0,n},\eta_{0,n})$ is available, it is most efficient to let $W_n$ be equal to the inverse of the estimated asymptotic variance. I discuss the choice of $W_n$ further in Section (ref). In many cases, it is possible to show that $\sqrt{n}\hat{M}(\theta_{0,n},\eta_{0,n})$ is asymptotically normal since it is a linear function of a vector of averages. This allows for modified versions of standard arguments for the asymptotic normality of GMM estimators to be applied.
Assumption 2.2\phantomsection As $n \rightarrow \infty$ while $Q$ and $p$ are fixed and either $J$ fixed or $J \rightarrow \infty$, we have that
For the uniform convergence condition in Assumption \hyperref[Ass2.1]{2.1.1}, I provide an example of when this holds in Appendix \hyperref[ApB]{B}, provided $\hat{g}$ satisfies a stochastic Lipschitz continuity condition. Note that Assumption \hyperref[Ass2.2]{2.2.6} imposes that $\beta_{0,n}$ are interior points of $B$ and bounded away from the boundary of $B$ since extremum estimators like the one defined by equation (ref) are generally not be asymptotically normal when the parameter is on the boundary of the parameter space or close to the boundary. Even when $\hat{g}$ is linear in $\beta$, $\partial_\beta \hat{g}(\theta)$ may still vary with $\delta$ so Assumption \hyperref[Ass2.2]{2.2.2} is imposed to ensure that $\partial_\beta \hat{g}(\beta_{0,n},\hat{\delta})$ converges to $\partial_\beta \hat{g}(\beta_{0,n},\delta_{0,n})$ as $\hat{\delta}$ converges to $\delta_{0,n}$. This condition trivially holds when the cross partial derivatives $\partial_{\beta \delta} \hat{g}(\theta)$ are equal to zero. Assumption \hyperref[Ass2.2]{2.2.3} can be directly combined with the adaptivity condition to show the asymptotic normality of $\sqrt{n}\hat{M}(\beta_{0,n},\hat{\delta},\hat{\eta})$ and Assumption \hyperref[Ass2.2]{2.2.4} guarantees the identification of $\beta$. The penalty function can be chosen to help ensure that Assumption \hyperref[Ass2.2]{2.2.4} holds and $\eta_{0,n}$ is bounded. Because $\hat{M}(\theta_{0,n},\eta_{0,n})$ involves taking sample averages, in the case where $J$ is fixed it can be satisfied under standard conditions that allow for a Central Limit Theorem to be applied. I provide examples of this in Appendix \hyperref[ApB]{B}. In cases where $J \rightarrow \infty$, arguments justifying Assumption \hyperref[Ass2.2]{2.2.3} can be complicated. However, as I show with the SCE case, when $\delta_{0,n}$ is high-dimensional but sparse, Assumption \hyperref[Ass2.2]{2.2.3} holds under very similar conditions to the low-dimensional case. For the case when the parameter of interest may be close to or at the boundary, we can use a “One-Step” estimator $ \Tilde{\beta}_{OS}$, equal to
This estimator can be thought of as minimizing the quadratic approximation of the GMM objective function at the initial estimate $\Tilde{\beta}_{GMM}$. This means that when $\beta$ is unconstrained, the estimators $\Tilde{\beta}_{GMM}$ and $\Tilde{\beta}_{OS}$ are identical. The conditions for the asymptotic normality of the One-Step estimator are the same as for the GMM estimator, except $\beta_{0,n}$ is allowed to be on the boundary or close to the boundary of the parameter space. This follows from the same reasoning as Theorem 1 in KETZ2018 for his “quasi-unconstrained" estimator.
Proposition 2.1 (Asymptotic Normality)\phantomsection Suppose $(\beta_{0,n},\delta_{0,n}) \in \Theta_{0,n}$, $\eta_{0,n}$ satisfies equation (ref), and Assumptions \hyperref[Ass2.1]{2.1} and \hyperref[Ass2.2]{2.2.1}-\hyperref[Ass2.2]{2.2.5} hold. Then as $n \rightarrow \infty$ with $p$ and $Q$ fixed with either $J$ fixed or $J \rightarrow \infty$, $\sqrt{n}(\Tilde{\beta}_{GMM} - \beta_{0,n}) = O_p(1)$ and $$\sqrt{n}(\Tilde{\beta}_{OS} - \beta_{0,n}) \overset{d}{\rightarrow} N(0,V),$$ where $V = (M_\beta' W M_\beta)^{-1} M_\beta' W V_M W M_\beta (M_\beta' W M_\beta)^{-1}$. If Assumption \hyperref[Ass2.2]{2.6} additionally holds, then $$\sqrt{n}(\Tilde{\beta}_{GMM} - \beta_{0,n}) \overset{d}{\rightarrow} N(0, V).$$ Therefore, when the same weighting matrix $W_n$ is used, the GMM and One-Step estimators achieve the same asymptotic variance. When $W = V_M^{-1}$, $V$ simplifies to $(M_\beta'V_M^{-1}M_\beta)^{-1}$. In general, the choice of the nuisance parameters may influence the precision of $\Tilde{\beta}_{GMM}$ and $\Tilde{\beta}_{OS}$ both through changing how much variability there is in the sample moment conditions (i.e., $V_M$) and through how sensitive the population moment conditions are to $\beta$ (i.e., $M_\beta$).
I now discuss estimation of the nuisance parameters. The results in the previous section hold for many possible values of $(\delta_{0,n},\eta_{0,n})$ where $$(\beta_{0,n},\delta_{0,n},\eta_{0,n}) \in S_{0,n} := \{ (\theta,\eta) \in \Theta_{0,n} \times H : \partial_{\delta} M(\theta,\eta) = 0 \}.$$ For some penalty function $f(\theta,\eta)$, I define the optimal nuisance parameters as
The primary purpose of the penalty function is to select a unique pair of elements from the identified sets. Therefore, we want there to be unique elements of the identified sets that minimize $f(\theta,\eta)$. However, the penalty not only influences whether the estimated nuisance parameters each converge to a particular element of the identified sets, but also which elements of the identified sets they converge to. In some cases a form of relative asymptotic efficiency can be achieved by making the penalty function depend on the asymptotic variance of $\Tilde{\beta}_{GMM}$. However, since the asymptotic variance is not observed, if the penalty function depends on this, then $f(\theta,\eta)$ is also unknown. Since $S_{0,n}$ is unknown as well, the nuisance parameters are chosen to minimize an estimated penalty function $\hat{f}$ among all parameters that come close to setting the sample versions of the moment conditions equal to zero. We can define the feasible set for the regularized parameters as $$\hat{S}_0 = \{ (\theta,\eta) \in \Theta_n \times H : ||\hat{g}(\theta)||_\infty \leq \lambda_\delta,\; ||\partial_\delta \hat{M}(\theta, \eta)||_\infty \leq \lambda_\eta \},$$ where $\lambda_\delta$ and $\lambda_\eta$ are tuning parameters whose choice is discussed below. Then we can estimate the parameters using
\phantomsection
For analyzing the asymptotic properties of this estimator, I consider two cases as before: one where the dimension of $\delta$ is fixed and one where the dimension is allowed to grow with the sample size. In the case for which $J$ is fixed and $\hat{g}$ is linear in $\theta$, Assumption \hyperref[Ass2.1]{2.1} only requires that $||\hat{\delta} - \delta_{0,n}||_1 + ||\hat{\eta} - \eta_{0,n}||_1 = o_p(1/\log(n))$ and $||\hat{\eta}\partial_\delta \hat{g}(\beta_{0,n},\delta)||_\infty = O_p(n^{-1/2}/\log(n)) $. This allows for very slow rates of convergence for $\hat{\delta}$ and $\hat{\eta}$. Since $||\hat{\eta}\partial_\delta \hat{g}(\hat{\theta})||_\infty \leq \lambda_\eta$, in the linear case, the second condition can be directly achieved by choosing $\lambda_\eta$ to be $O_p(n^{-1/2}/\log(n))$. However, if $\hat{g}$ is non-linear in $\delta$, then we also want $||\hat{\delta} - \delta_{0,n}||_2 = o_p(n^{-1/4})$ and $||\hat{\eta} - \eta_{0,n}||_2 = o_p(n^{-1/4})$ in order for Assumption \hyperref[Ass2.1]{2.1.4} to hold. While this still allows for rates of convergence slower than the standard parametric $\sqrt{n}$ rate, stronger assumptions are imposed to satisfy these conditions for the non-linear case.
Assumption 3.1\phantomsection Let $S_n^\zeta \coloneqq \{(\theta,\eta) \in \Theta_{n} \times H :f(\theta,\eta) \leq f(\theta_{0,n},\eta_{0,n}) + \zeta\}$ and $S_{0,n}^\zeta \coloneqq \{(\theta,\eta) \in S_{0,n} :f(\theta,\eta) \leq f(\theta_{0,n},\eta_{0,n}) + \zeta\}$ for each $\zeta > 0$. As $n\rightarrow \infty$ while $p$ and $Q$ are fixed, and either $J$ fixed or $J \rightarrow \infty$, we have that
For the estimator defined by equation (ref), both the objective function and feasible set may be stochastic, so there can be uncertainty coming from both the estimated penalty function and the sample moment conditions. As a result, the rate of convergence for $\hat{\delta}$ and $\hat{\eta}$ depends both on the rate of convergence of the sample moments to the population moments and the rate of convergence of the estimated penalty function. Assumption \hyperref[Ass3.1]{3.1.1} allows for the parameter spaces to not be compact, as long as the feasible values of the parameters that make the penalty sufficiently small are compact. Assumption \hyperref[Ass3.1]{3.1.4} can be viewed as an extension of Assumption \hyperref[Ass2.2]{2.2.4}. It provides a strong identification condition for the identified set $S_{0,n}$. Strong partial identification conditions of this form are common in the literature on estimating identified sets and are imposed by others such as ChernozhukovHongTamer2007. It holds in a variety of applications including the case of a SCE under a linear factor model data-generating process in Section (ref) and the many instruments application in section (ref). Assumption \hyperref[Ass3.1]{3.1} weakens the assumptions imposed by ChernozhukovHongTamer2007 by not requiring that the parameter space or the identified sets to be compact, and it allows for the dimension of $\theta$ to be growing. Part of the reason that these conditions can be weakened is that I do not need to show that $\hat{S}_0$ are converging to $S_{0,n}$. Rather, I only need to show convergence of the feasible set on the subset of the parameter space where $f$ is small. Assumption \hyperref[Ass3.1]{3.1.2} strengths Assumption \hyperref[Ass2.2]{2.2.1} by imposing a specific rate for the uniform convergence of the sample moment conditions. I impose a rate of convergence for the estimated penalty function in Assumption \hyperref[Ass3.1]{3.1.3} and a condition relating $|f(\beta_{0,n},\delta,\eta) - f(\theta_{0,n},\eta_{0,n})|$ to $||\delta - \delta_{0,n}||_1$ and $||\eta - \eta_{0,n}||_1$ in Assumption \hyperref[Ass3.1]{3.1.5}. Note that for Assumption \hyperref[Ass3.1]{3.1.5}, it is only necessary that among elements of $S_{0,n}$ that are close to $\theta_{0,n}$ and $\eta_{0,n}$, $(||\delta - \delta_{0,n}||_1 + ||\eta - \eta_{0,n}||_1)^{\gamma_1}$ can be bounded by $f(\beta_{0,n},\delta,\eta) - f(\beta_{0,n}, \delta_{0,n},\eta_{0,n})$. This allows us to guarantee that if elements of the identified sets achieve close to the minimum value of $f$, then they must be close to $\theta_{0,n}$ and $\eta_{0,n}$. It may often be the case that $(\delta_{0,n},\eta_{0,n})$ is on the boundary of the identified set $S_{0,n}$ and is not a local minimum of $f(\beta_{0,n},\delta,\eta)$ on $D_n \times H$. This is why the second part of Assumption \hyperref[Ass3.2]{3.2.5} is needed, because it allows us to place a bound on the rate of change of $f(\theta,\eta)$ for values of the nuisance parameter that give a penalty value which is not significantly greater than the value at $(\theta_{0,n},\eta_{0,n})$. As a result, values just outside $S_{0,n}$ should not make $f$ much lower than $f(\theta_{0,n},\eta_{0,n})$. This still allows for the possibility that $f(\theta,\eta)$ may diverge to infinity at some values of $(\theta,\eta)$, such as values of $\eta$ where $\beta$ becomes unidentified when $f$ depends on $V$. When $f$ depends on the $V$, Assumption \hyperref[Ass3.1]{3.1.3} often holds when $\hat{f}$ is the same function of an estimate of the asymptotic variance and $\hat{V}$ is converging to $V$ at a rate of $c_n$. In subsection (ref) below, I discuss how the estimated variance is often converging at least as fast as a rate as the sample moment conditions. Alternatively, $f$ can be chosen to be a known function, in which case Assumption \hyperlink{Ass3.1}{3.1.3} holds trivially with $\hat{f}(\theta,\eta) = f(\theta,\eta)$, so $c_n = 0$. Under similar regularity conditions on the sample moment conditions as before, when $\Theta_n \times H$ is compact and $J$ is fixed, Assumption \hyperref[Ass3.1]{3.1.2} often holds with $a_n = 1/\sqrt{n}$. In the case where $J$ is growing, $a_n$ is generally growing with $J$ but adding constraints on $\delta$ can help ensure that it is growing slowly in $J$ (see Remark \hyperref[R3.1]{3.1}). For the SCE case, I use $f(\theta,\eta) = ||\delta||_2^2 + ||\eta||_2^2$ which, along with the properties of the identified set, allows for Assumption \hyperref[Ass3.1]{3.1.5} to hold with $\gamma_1 =2$ and $\gamma_2 = 1$.
Lemma 3.1 (Rate of Convergence of Regularized Estimates)\phantomsection Suppose that Assumption \hyperref[Ass3.1]{3.1} holds, and $\lambda_\delta, \lambda_\eta \overset{p}{\rightarrow} 0$ such that $b_n/\min\{\lambda_\delta,\lambda_\eta\} \overset{p}{\rightarrow} 0$ as $n \rightarrow \infty$ with $p$ and $Q$ fixed and either $J$ fixed or $J \rightarrow \infty$. Then $$||\hat{\delta} - \delta_{0,n}||_1 + ||\hat{\eta}-\eta_{0,n}||_1 = O_p(\max \{\lambda_\delta,\lambda_\eta,a_n,c_n\}^{\gamma_2/\gamma_1}).$$
This result has a similar form as Theorem 3.1 of ChernozhukovHongTamer2007 for the rate of convergence of the estimated identified set, except here we have the terms $c_n$, $\gamma_1$, and $\gamma_2$, which capture how the behavior of the penalty function influences the convergence of the penalized estimator. Lemma \hyperref[L3.1]{3.1} requires that $\lambda_\delta$ and $\lambda_\eta$ are shrinking more slowly than $b_n$. This guarantees that $S_{0,n} \subseteq \hat{S}_0$ with probability approaching one. Similarly as to with $a_n$ and $c_n$, when the identified sets are compact and $J$ is fixed, we often have that $b = 1/\sqrt{n}$. We can therefore satisfy this condition by having $\lambda_\eta$ and $\lambda_\delta$ shrinking at a slightly slower rate than $a_n$ and $b_n$ (e.g., $ \lambda_\delta,\lambda_\eta = O_p(n^{-1/2}\log (n))$). In such cases, we have $||\hat{\delta} - \delta_{0,n}||_1 + ||\hat{\eta} - \eta_{0,n}||_1 = O_p((n^{-1/2}\log(n))^{\gamma_2/\gamma_1})$. For the linear case, we only need $||\hat{\delta} - \delta_{0,n}||_1 = o_p(1/\log(n))$ when $J$ is fixed to satisfy Assumption \hyperref[Ass2.1]{2.1.1} and \hyperref[Ass2.1]{2.1.4}, so Assumption \hyperref[Ass3.1]{3.1.5} holding with any $\gamma_1,\gamma_2 > 0$ is sufficient. On the other hand, for the non-linear case, we want $||\hat{\delta} - \delta_{0,n}||_1 = o_p(n^{-1/4})$ so we need $\gamma_2/\gamma_1 > 1$. In cases where $\gamma_2/\gamma_1 \leq 1$, one potential solution is the approach of mackey18, where further nuisance parameters are introduced to make the moment conditions $h$-th order orthogonal, in which case we only need that $||\hat{\delta} - \delta_{0,n}||_2 = o_p(n^{-1/(2h+2)})$. When $J \rightarrow \infty$, $a_n$ as well as $c_n$ if $c_n \ne 0$ are generally growing with $J$. However, as long as it is growing slowly in $J$ and $J$ is not growing too much faster than $n$, we can satisfy Assumption \hyperref[Ass2.1]{2.1.1} and \hyperref[Ass2.1]{2.1.4} under similar conditions as when $J$ is fixed.\footnote{For example, if $a_n = b_n = c_n = n^{-1/2}\log (J)$ then we can choose $\lambda_\delta$ and $\lambda_\eta$ to be $O_p(n^{-1/2}\log(J) \log(n))$ so that $||\hat{\delta} - \delta_{0,n}||_1 = O_p((n^{-1/2}\log(J) \log(n))^{\gamma_2/\gamma_1})$. Then as long as $\gamma_2,\gamma_1 > 0$ and $J/n^\gamma \rightarrow 0$ for some $\gamma > 0$, we have that $\hat{\delta}$ satisfies Assumptions \hyperref[Ass2.1]{2.1.1} and \hyperref[Ass2.1]{2.1.4} for the linear case.} Having constraints on $\delta$ is beneficial for guaranteeing $a_n$ and $c_n$ grow slowly in $J$.
Remark 3.1 (Adding L1 norm Constraints)\phantomsection Suppose that we additionally add a constraint on the L1 norm of $\delta$ to the estimator defined by equation (ref) uses $\Tilde{\Theta}_n = \{ \theta \in \Theta_n: ||\delta||_1 \leq \lambda_1 \}$ as its parameter space, for some additional penalty term $\lambda_1$. Furthermore, suppose that $g$ and $\hat{g}$ are linear in $\theta$ so $g(\theta) = E[\sum_{i=1}^n( X_0 - X_1 \beta - X_2 \delta)/n]$ and $\hat{g}(\theta) = \sum_{i=1}^n (X_{0,i} - X_{1,i}\beta - X_{2,i}\delta)/n$. Then for the uniform convergence conditions in Assumption \hyperref[Ass3.2]{3.2.2}, $\sup_{\theta \in \Tilde{\Theta}_n}||\hat{g}(\theta) - g(\theta)||_\infty \leq ||\sum_{i=1}^nE[X_{0,i}]/n - \sum_{i=1}^nX_{0,i}/n||_\infty + \sup_{\beta \in B} ||(\sum_{i=1}^nE[X_{1,i}]/n - \sum_{i=1}^n X_{1,i}/n)\beta||_\infty + \lambda_1 \max_{1 \leq j \leq J}|\sum_{i=1}^nE[X_{2,j,i}]/n - \sum_{i=1}^n X_{2,j,i}/n|$. Note that the last term is also a bound on $\sup_{\theta \in \Tilde{\Theta}_n}||\partial_\delta \hat{g}(\theta) - \partial_\delta g(\theta)||_\infty$. In many cases, the maximum in the last term can be growing slowly in $J$. Additionally, if it results in $\delta_{0,n}$ being sparse, this can make it easier to show that $\sqrt{n}\hat{M}(\beta_{0,n},\delta_{0,n},\eta_{0,n})$ is asymptotically normal. On the other hand, if $\lambda_1$ is kept too small so that $\{\delta : ||\delta||_1 \leq \lambda_1\} \cap D_{0,n} = \emptyset$, then it causes Lemma \hyperref[L3.1]{3.1} to fail to hold. Also, further constraining $\delta$ may increase what the minimum asymptotic variance obtainable is. The conditions placed on $\lambda_\delta$ and $\lambda_\eta$ in Lemma \hyperref[L3.1]{3.1} allow for them to be stochastic, and therefore potentially chosen in a data-driven way, but do not tell us how to choose them in practice. Therefore, in order to reduce room for specification searching by researchers, it is important to have an algorithm for determining these tuning parameters. One plausible approach is based on estimating confidence sets for the identified set. When a parameter is partially identified by some population criterion function, one way to construct confidence sets for the identified set, used by ChernozhukovHongTamer2007 and others, is to use the set where the sample criterion function is below a particular value. This is similar to how $\hat{S}_0$ is constructed. However, if $\lambda_\delta$ and $\lambda_\eta$ are chosen so that $\hat{\Theta}_0$ and $\hat{H}_0$ are confidence set for $S_{0,n}$, this still leaves the choice of the confidence level. In order for $b_n/\min \{\lambda_\delta,\lambda_\eta\} \overset{p}{\rightarrow} 0$, we want this confidence level to be converging to $100\%$ as $n \rightarrow \infty$ and there is still the question of how to choose the confidence level in practice. Furthermore, these methods often require resampling and therefore may be computationally intensive. For the SCE in section (ref), I use a computationally simple method which guarantees that $\hat{S}_0$ is non-empty but also that $\lambda_\delta$ and $\lambda_\eta$ are shrinking at the appropriate rate.\footnote{This method involves finding the minimum values of $\lambda_\delta$ and $\lambda_\eta$ that makes $\hat{S}_0$ non-empty and multiplying these values by a function of the number of the number of time periods and number of controls.}
\phantomsection
As mentioned earlier, $$V = (M_\beta' W M_\beta)^{-1} M_\beta' W V_M W M_\beta (M_\beta' W M_\beta)^{-1}.$$ Therefore, we can estimate $V$ with $$\hat{V}(\hat{\theta},\hat{\eta}) = (\hat{M}_\beta' W_n \hat{M}_\beta)^{-1} \hat{M}_\beta' W_n \hat{V}_M(\hat{\theta},\hat{\eta}) W_n \hat{M}_\beta (\hat{M}_\beta' W_n \hat{M}_\beta)^{-1},$$ where $\hat{M}_\beta = \partial_\beta \hat{M}(\hat{\theta},\hat{\eta})$ and $\hat{V}_M(\hat{\theta},\hat{\eta})$ is an estimate of the asymptotic variance of the orthogonal moment conditions. In the case where the data are I.I.D., $V_M$ is simply the variance matrix for the moment conditions, and if $\hat{V}_M(\theta,\eta)$ is the sample variance matrix then under simple regularity conditions $\hat{V}_M(\hat{\theta},\hat{\eta})$ is $\sqrt{n}$-consistent when $J$ is fixed. However, in contexts with time series or panel data, $V_M$ often corresponds to the long-run variance of the moment conditions. When $V_M$ corresponds to the long-run variance and the structure of the dependence across observations is unknown, it is common to use estimators in the class of quadratic Heteroscedastic Autocorrelation Consistent (HAC) variance estimators that take the form:
where $Q_K(i,s)$ is a weighting function that depends on a smoothing parameter $K$. This includes kernel variance estimators such as those of Andrews1991 and NeweyWest1987 as well as the orthonormal series variance estimators such as that of Phillips_2005. For conventional Kernel HAC estimators, $Q_K(i, s) = \mathcal{K} ((i -s)/K)$ where $\mathcal{K}$ is a kernel and $K$ is the bandwidth. For Series HAC estimators, $Q_K(i,s) = \sum_{k=1}^K \phi_k (i) \phi_k(s)/K$ where $ \{ \phi_k(s) \}_{k=1}^K$ are orthonormal basis functions taking values in $[0,1]$. Asymptotic results for non-parametric long-run variance estimators that rely on $K \rightarrow \infty$ as $n \rightarrow \infty$ can often provide poor approximations in practice, particularly when the degree of temporal dependence is high relative to the sample size. Intuitively, this is because the uncertainty in our estimation of $V_M$ significantly contributes to our uncertainty in our test statistic in these cases, but this is not captured by increasing-smoothing asymptotic results. SCEs are often used is settings with small to moderate sample sizes and with data that display a high degree of dependence over time, making the use of increasing-smoothing asymptotic results particularly questionable. This is illustrated by t-test2022 who show that their inference procedure for their SCE performs very poorly when they calculate their standard errors using a HAC estimator and rely on increasing-smoothing asymptotic results to obtain critical values. The notion of fixed-smoothing asymptotics was first introduced by VogelsangKiefer2002a. As the name suggests, it involves keeping the degree of smoothing $K$ fixed as the sample size grows. For both Kernel and Series HAC estimators, this results in $\hat{V}_M(\hat{\theta},\hat{\eta})$ converging to a stochastic matrix. The fact that $\hat{V}_M(\hat{\theta},\hat{\eta})$, and therefore $\hat{V}(\hat{\theta},\hat{\eta})$, are converging to something stochastic has several implications for this procedure, since $\hat{V}_M$ is potentially being used three different times. It may be used to estimate $\hat{\delta}$ and $\hat{\eta}$, it can be used to estimate $\Tilde{\beta}_{GMM}$ and $\Tilde{\beta}_{OS}$ by setting $W_n = \hat{V}_M(\hat{\theta},\hat{\eta})^{-1}$, and lastly it may be used along with $\Tilde{\beta}_{GMM}$ or $\Tilde{\beta}_{OS}$ to construct a test statistic or confidence interval. If $\hat{V}$ is not converging to $V$, then $\hat{f}$ should not be a function of $\hat{V}$ in order to satisfy Assumption \hyperref[Ass3.1]{3.1.3}. Therefore, we want to define an alternative penalty function. One alternative is to have $f$ depend on a known function which upper bounds $V(\theta,\eta)$. In this case, Assumption \hyperref[Ass3.1]{3.1.3} can be trivially satisfied with $\hat{f} = f$. In Section (ref), I give an example of how to do this for the SCE case. For using $\hat{V}_M(\hat{\theta},\hat{\eta})$ to weight the moment conditions, HwangSun2018oneVsTwoStep compare the performance of the One-Step and Two-Step GMM estimator under a fixed-smoothing asymptotic framework. They show that whether the Two-Step GMM procedure outperforms the One-Step GMM procedure depends on the values of long-run correlation coefficients. Because these long-run correlations can also not be consistently estimated under the fixed-smoothing asymptotic framework, it is generally not clear whether the One-Step or Two-Step GMM estimators performs better. As shown by Sun2014TwoStep, under fixed-smoothing asymptotics, while the One-Step GMM estimator is still asymptotically normal, the Two-Step GMM estimator is asymptotically mixed normal. Because of this, it is easier to choose $\hat{\delta}$ and $\hat{\eta}$ to minimize an upper bound on the asymptotic variance and simpler to conduct inference when $W_n = I_m$. For these reasons, I focus on using $W_n = I_m$ and have the penalty function be based on the upper bound of the asymptotic variance when fixed-smoothing asymptotic results are relevant. However, even when $\Tilde{\beta}_{GMM}$ and $\Tilde{\beta}_{OS}$ are asymptotically normal, common test statistics can still have nonstandard limiting distributions. For example, the Wald test statistic, rather than converging to a chi-squared distribution, converges to a distribution which depends on the kernel or basis function and the smoothing parameter. For Kernel HAC estimators, Sun2014Fixedb provides conditions under which an adjusted Wald statistic and an adjusted $t$-statistic have asymptotic distributions that can be approximated by $F$ and $t$ distributions, but do not converge exactly to $F$ and $t$ distributions. On the other hand, Sun2013 gives versions of the Wald statistic and $t$-statistic that converge exactly to $F$ and $t$ distributions when $W_n = I_m$ and a Series HAC estimator is used. Furthermore, LazarusLewisStock2021 characterize the size-power frontier for Kernel and Series HAC estimators under a fixed-smoothing framework and find that there is little cost to restricting attention to tests which converge exactly to $t$ and $F$ distributions. I therefore focus on verifying that the conditions of Sun2013 for the GMM estimator with $W_n = I_m$ and $\hat{V}_M$ being a Series HAC estimator, with the test statistic being the standard Wald statistic defined as
and the $t$-statistic is defined as
when $p =1$. I impose the following conditions in order for this method to be able to conduct valid and standard inference in a fixed-smoothing asymptotic framework.
Assumption 3.2\phantomsection Suppose that with $p$, $K$, and $Q$ fixed, and either $J$ fixed or $J \rightarrow \infty$, as $n \rightarrow \infty$, we have that
Assumptions \hyperref[Ass3.2]{3.2.2}-\hyperref[Ass3.2]{3.2.4} contain the conditions imposed by Sun2013 adjusted to this setting.\footnote{Here, I imposed the conditions on the original moment conditions evaluated at $\theta_{0,n}$ rather than on the orthogonalized moment conditions used to estimate $\Tilde{\beta}_{GMM}$ and $\Tilde{\beta}_{OS}$. In the proof of Proposition \hyperref[P3.1]{3.1}, I verify that this implies that they hold for $\hat{M}(\beta,\hat{\delta},\hat{\eta})$ under Assumption \hyperref[Ass2.1]{2.1}.} Assumption \hyperref[Ass3.1]{3.1.2} is satisfied for commonly used based functions. For example, $\phi_k(x) = \sqrt{2}\sin (2\pi kx)$ and $\phi_k(x) = \sqrt{2}\cos (2\pi kx)$ satisfy the condition. However, there are basis functions such as $\phi_k(x) = \sqrt{2} \sin (\pi (0.5 - k)x)$ that do not satisfy it because it does not satisfy the mean-zero condition. Assumption \hyperref[Ass3.1]{3.1.3} is standard in the literature on fixed-smoothing asymptotics (see, for example, VogelsangKiefer2005), and easily holds in cases where $\hat{g}$ is linear in $\theta$. Assumption \hyperref[Ass3.2]{3.2.4} holds when $\hat{B}(r) = \sum_{i=1}^{\lfloor r n \rfloor} \hat{g}_i(\theta_{0,n})/\sqrt{n}$ is converging weakly to a Gaussian process with almost surely continuous sample paths and independent increments. In Appendix \hyperref[ApB]{B}, I show that this can hold in cases where $g_i(\theta_{0,n})$ is not stationary. One additional complication that is not present in Sun2013 or other previous fixed-smoothing results is the plugged-in values of the nuisance parameters $\hat{\delta}$ and $\hat{\eta}$. This is why Assumption \hyperref[Ass3.2]{3.2.1} is imposed. Assumption \hyperref[Ass3.2]{3.2.1} can be viewed as a slightly stronger version of Assumption \hyperref[Ass2.2]{2.2.2} and it is serving the same role of bounding how sensitive $\partial_\beta g_i (\theta)$ is to $\delta$. This has sufficient conditions similar to Assumption \hyperref[Ass2.2]{2.2.2} and also holds trivially when the cross partial derivatives of the moment conditions are equal to zero. Additionally, similarly to with the adaptivity condition in Lemma \hyperref[L2.1]{2.1}, we want $ \hat{V}(\beta_{0,n},\hat{\delta},\hat{\eta})$ to be asymptotically equivalent to $\hat{V}(\beta_{0,n},\delta_{0,n},\eta_{0,n})$. In order for this to be the case, I impose Assumption \hyperref[Ass2.1*]{2.1*}, which slightly strengthens some of the conditions of Assumption \hyperref[Ass2.1]{2.1} to hold with partial sums.
Assumption 2.1*\phantomsection As $ n\rightarrow \infty$ while $p$ and $Q$ are fixed and either $J$ fixed or $J \rightarrow \infty$, we have that
Under Assumptions \hyperref[Ass2.1]{2.1} and \hyperref[Ass2.1*]{2.1*}, when $\hat{V}_M$ is a Series HAC estimator, $$\hat{V}_M(\beta_{0,n},\hat{\delta},\hat{\eta}) - \hat{V}_M(\beta_{0,n},\delta_{0,n},\eta_{0,n}) \overset{p}{\rightarrow} 0.$$ Since the partial sums $\sum_{i=1}^t \eta g_i (\theta)$ include fewer terms than the full sum $\sum_{i=1}^n \eta g_i(\theta)$, they are generally smaller asymptotically, so the bounds on partial sums in Assumptions \hyperref[Ass2.1*]{2.1*} have similar sufficient conditions to before. Together with Assumption \hyperref[Ass3.2]{3.2}, we can show that the Wald and $t$ statistics converge to $F$ and $t$ distributions.
Proposition 3.1(Fixed-Smoothing Results) \phantomsection Suppose $\Tilde{\beta}_{GMM}$ is estimated using equation (ref) with $W_n = I_m$ and Assumptions \hyperref[Ass2.1]{2.1*}, \hyperref[Ass2.1]{2.1}, \hyperref[Ass2.2]{2.2}, and \hyperref[Ass3.2]{3.2} hold with $K \geq p$. When $n \rightarrow \infty$ with $K$, $p$, and $Q$ fixed and either $J$ fixed or $J \rightarrow \infty$,
where $F_{p,K-p+1}$ is an $F$ distribution with $p,K-p+1$ degrees of freedom and $t_K$ is a $t$-distribution with $K$ degrees of freedom. Here, I've focused on test statistics that use $\Tilde{\beta}_{GMM}$, but because $\Tilde{\beta}_{OS}$ has the same asymptotic distribution, Proposition \hyperref[P3.1]{3.1} holds for the Wald and $t$ statistics that use $\Tilde{\beta}_{OS}$ estimated with $W_n = I_m$ if we simply use $\Tilde{\beta}_{OS}$ in place of $\Tilde{\beta}_{GMM}$ in Assumption \hyperref[Ass3.2]{3.2.3}. If we were instead to focus on asymptotics with $K \rightarrow \infty$ where $V(\theta_{0,n},\eta_{0,n})$ is consistently estimated, then we have that $\mathbb{W}_n \overset{d}{\rightarrow} \chi^2_p/p$, where $\chi^2_p$ is chi-squared distribution with $p$ degrees of freedom, and $t_n \overset{d}{\rightarrow} N(0,1)$. Note that $F_{p,K-p+1} \overset{d}{\rightarrow} \chi^2_p/p$ and $t_K \overset{d}{\rightarrow} N(0,1)$ as $K \rightarrow \infty$, so critical values obtained from the fixed-smoothing asymptotic results are approximately the same as the critical values from increasing smoothing asymptotic results when $K$ is large. For applications in which increasing smoothing asymptotic results are likely to provide an accurate approximation, consistency for a variety of HAC estimators has been shown under fairly general conditions. For example, suppose that the sequence $\{ g_i(\theta_{0,n})\}_{i \in \mathbb{N}}$ is mean-zero, $\alpha$-mixing, and there exists $\nu > 1$ such that $\sup_{i} E[||g_i(\theta_{0,n})||_2^\nu] < \infty$ and $\sum_{s=1}^\infty s^2 \alpha(s)^{\frac{\nu - 1}{\nu }} < \infty$ where $\alpha(s)$ are the mixing coefficients. Andrews1991 shows that a class of Kernel HAC estimators are consistent under these conditions if $n,K \rightarrow \infty$ such that $K^2/n \rightarrow 0$.\footnote{This follows from Lemma 1 and Theorem 1(a) of Andrews1991. The class of Kernel estimators includes many commonly used ones such as the truncated, Bartlett, Parzen, Tukey-Hanning, and Quadratic Spectral kernels.} As noted in the literature, non-parametric long-run variance estimators can often be converging to the true long-run variance at a rate faster than $\sqrt{n}$. For Series HAC estimators, Phillips_2005 provides a Mean Squared Error (MSE)-optimal choice of $K$ under stationarity conditions such that $K = O(n^{4/5})$ and a convergence rate of $n^{4/5}$ is achieved. Whereas $K = O(n^{2/3})$ when using the Coverage Probability Error (CPE)-optimal choice of Sun2013. Andrews1991 provides rates of convergence for a class of Kernel HAC estimators when the data are weakly stationary. He shows that when using a Quadratic Spectral Kernel with the optimal choice of bandwidth, it is also possible to achieve a $n^{4/5}$ rate of convergence, while other choices still achieve a rate faster than $\sqrt{n}$. However, $\hat{V}(\theta,\eta)$ depends on $\hat{V}_M(\theta,\eta)$ and $\partial_\beta \hat{M}(\theta,\eta)$, and $\partial_\beta \hat{M}(\theta,\eta)$ is usually converging at the standard parametric rate $\sqrt{n}$ when $J$ is fixed. This means that when the penalty function $f$ depends on $V$ and the fixed-smoothing asymptotic results are not relevant, Assumption \hyperref[Ass3.1]{3.1} usually holds with $c_n = n^{-1/2}$ when $J$ is fixed, so the estimated penalty function is often converging at the same rate as the sample moment conditions. Intuitively, choosing the penalty function so that the nuisance parameters are minimizing some estimate of the asymptotic variance should provide a relatively more efficient estimator, at least when the variance estimator is sufficiently accurate. However, if $\beta$ is not a scalar, then there may be a trade-off between more precisely estimating different subvectors of $\beta$. However, in most applications, either $\beta$ is a scalar or, if $p > 1$, then the object of interest is $h(\beta)$ for some known function $h: \mathbb{R}^p \rightarrow \mathbb{R}$. When $\beta$ is a scalar, we can choose $\hat{f}(\theta,\eta) = \hat{V}(\theta,\eta)$. If the object of interest is $h(\beta)$ and $h$ is continuously differentiable, then $\sqrt{n}(h(\Tilde{\beta}_{GMM}) - h(\beta_{0,n}))$ is asymptotically normal by the Delta method and we can choose $\hat{f}(\theta,\eta) = \partial_\beta h(\beta) \hat{V}(\theta,\eta) \partial_\beta h(\beta)'$. In summary, in applications where fixed-smoothing asymptotic results provide a better approximation due to the degree of dependence being high relative to the sample size, I recommend: not having $\hat{f}$ depend on $\hat{V}$, having $W_n = I_m$, and using a Series HAC estimator with $K$ chosen according to Sun2013. Otherwise, greater efficiency can be achieved by having $\hat{f}$ depend on $\hat{V}$, having $W_n = \hat{V}_M(\hat{\theta},\hat{\eta})^{-1}$, and having $\hat{V}_M$ be either a Series or Kernel HAC estimator with $K$ chosen to provide the fastest rate of convergence.
First introduced by AandG and Abadie2010, SCEs have become popular in contexts with panel data where a single unit becomes and stays treated and there is a large pool of never-treated units which can be used as control units. The method involves constructing a weighted average of the control units or a Synthetic Control (SC) unit by minimizing the difference between this SC and the treated unit on a set of pre-treatment predictor variables, so that this SC can be used as an estimate of the treated unit's counterfactual outcomes in the post-treatment time periods. I focus on analyzing my method in this original context, where there is a single unit that becomes and stays treated. Researchers are most commonly interested in conducting inference on the average treatment effect on the treated unit and the weights on the control units are a nuisance parameter, so I use $\beta$ to denote the ATT and $\delta$ to denote the control weights. When applying my method as an SCE, I refer to my estimator as the Orthogonalized SCE.
\phantomsection I first discuss how to implement the estimator as an SCE and verify the conditions imposed in the formal results above hold when the data follow a linear factor model. Linear factor models, also known as interactive fixed effect models, have been a common setting to explore the properties of SCEs, beginning with Abadie2010. I index the units $\{0, 1,...,J\}$ where $i = 0$ is the treated unit and $\mathcal{J} = \{1,...,J\}$ denotes the set of control units. The control units never receive treatment, whereas the treated unit becomes and stays treated after a known point in time. I denote the sets of indices for time periods prior to its treatment and after its treatment as $\mathcal{T}_0$ with $T_0 = |\mathcal{T}_0|$ and $\mathcal{T}_1$ with $T_1 = |\mathcal{T}_1|$ respectively. Here, the asymptotics are slightly different from before because there are two variables $T_0$ and $T_1$ that capture our sample size rather than just $n$. I focus on verifying the assumptions in the previous sections for $n = \min \{T_0, T_1\}$.
Assumption 4.1 (Linear Factor Model)\phantomsection For all units $i \in \{0,...,J\}$ and time periods $t \in \mathcal{T}_0\bigcup \mathcal{T}_1$, outcomes follow a linear factor model with $R$ factors so that
where $d_{it}$ is an indicator function equal to 1 if and only if $i = 0$ and $t \in \mathcal{T}_1$ and equal to $0$ otherwise. The factor loadings $\mu_i$, dynamic treatment effects $\beta_t$, and treatment assignment $d_{it}$ are fixed, but the latent factors $f_t$ and idiosyncratic shocks $\epsilon_{it}$ are stochastic.\footnote{I have the treatment effects be fixed because the parameter of interest $\beta_{0,n}$ in the previous sections was assumed to be fixed. However, in the SCE case, generalizing the results to allow $\beta_{0,n}$ to be stochastic is straightforward, in which case confidence intervals can then be interpreted as prediction intervals.} I let $\mu$ denote the $R \times (J+1)$ matrix of factor loadings with $\mu_i$ being its $i$-th column, $f$ denotes the $(T_0+T_1) \times R$ matrix of realizations of the factors with $f_t$ being is $t$-th row, and $\epsilon$ denote the $(J+1) \times (T_0 + T_1)$ matrix of idiosyncratic shocks. Additionally, I use the subscript $\mathcal{J}$ to denote the sub-matrix for only the units $j \in \mathcal{J}$ and the superscripts $pre$ and $post$ to denote the sub-matrices for only pre-treatment and post-treatment values respectively. I define the average treatment effect on the treated to be $\beta_{0,n} = \sum_{t \in \mathcal{T}_1} \beta_t /T_1$. This means that $\beta_{0,n}$ may be changing as $T_1$ grows, just as $\beta_{0,n}$ is allowed to change with the sample size as in the previous sections. If the idiosyncratic shocks are mean-zero and the SC has the same factor loadings as the treated unit (i.e., $\mu_0 = \mu_{\mathcal{J}}\delta$), then we can identify $\beta_{0,n}$ using the moment condition $\sum_{t \in \mathcal{T}_1} E[Y_{0t} - \beta - Y_{\mathcal{J},t}'\delta]/T_1 = 0$. Therefore, we want to use this moment condition plus a set of moment conditions that identify the set $D_{0,n} = \{\delta \in \Delta^J: \mu_0 = \mu_{\mathcal{J}}\delta \}$, where $\Delta^J \coloneqq \{\delta \geq 0 : ||\delta||_1 = 1 \}$ is the $J-1$-dimensional unit simplex. Several of the difficulties mentioned earlier can arise in characterizing the asymptotic distribution of the estimated average treatment effect due to its dependence on $\hat{\delta}$. Not only may there be many $\delta \in D_{0,n}$ so $\delta$ is partially identified, but analysis of the asymptotic distribution of the control weights is also complicated by the fact that it is often high-dimensional (relative to $T_0$ and $T_1$). Lastly, even when $J$ is fixed and $\delta$ is point-identified, $\sqrt{T_0}(\hat{\delta} - \delta_{0,n})$ generally has a non-standard asymptotic distribution due to the constraints $\hat{\delta} \in \Delta^J$.\footnote{This is illustrated by Li2020 and fry2024 who use the method of Andrews1999 to characterize the asymptotic distribution of the estimated average treatment effect.} Other inference methods, like CaoDowd2019 and Li2020, also have the limitations of assuming that the control weights are point-identified and treating the number of units as fixed in their asymptotic results. One exception is Zhang_2023 who obtain a $\sqrt{n}$-consistent estimator while allowing $\delta$ to be partially identified, and another is t-test2022 whose method is discussed further below in subsection (ref). Several recent papers have discussed how to estimate the SC using a set of moment conditions, including fry2024, Powell2021, and shi2023. A linear instrumental variables approach along the lines of fry2024 or shi2023 is the most straightforward to verify the conditions of sections (ref) and (ref). In this case, the moment conditions are
where $Z_t$ contains the values of $Q - 1$ instruments in pre-treatment time periods. In fry2024, the vector of instruments is a constant and outcomes of other units which are not included in the set of controls but are also untreated in pre-treatment time periods. The exclusion restriction $E[Z_{qt}(Y_{0t} - Y_{\mathcal{J},t}\delta)] = 0 $ holds for these other units in pre-treatment time periods when the idiosyncratic shocks are uncorrelated across units and the factors are uncorrelated with the idiosyncratic shocks. Intuitively, when the factors are responsible for the covariance across units, we can guarantee that the SC has the same exposure to latent factors as the treated unit by estimating the SC to have the same covariance with other units. In the empirical application below, I provide a practical example for how the set of instruments can be chosen.\footnote{fry2024 also provides several model selection methods for splitting untreated units into a set of controls and set of instruments. However, because of the potential problems for inference that using a data-driven model selection procedure may introduce, I focus on cases where the units used as instruments are only the units which are known to not be valid controls but plausibly valid instruments. I discuss how this can be done when discussing the empirical application below.} Other potentially valid choices of instruments exist, such as using lagged values of the outcome variable or using shift-share instruments. Additionally, it may be possible to reframe other estimators, such as the Debiased OLS estimator of t-test2022 discussed below, as method of moments estimators, in which case the same orthogonalization technique could be employed. Following my recommendation from section (ref), I set $m = 1$ so $\eta$ is $1 \times Q$ and there is only a single orthogonal moment condition. This means that $\partial_\beta M(\theta,\eta) = -\eta_Q$ where $\eta_Q$ is the $Q$-th element of $\eta$. Also, since there is only a single orthogonalized moment condition that is linear in $\beta$, both the One-Step and GMM estimators in Section (ref) are equivalent to picking the value of $\beta$ that sets $\hat{M}(\beta,\hat{\delta},\hat{\eta})$ equal to zero for any positive definite weighting matrix $W_n$. Therefore, the estimator is simply given by:
Under Assumptions \hyperref[Ass4.2]{4.2} and \hyperref[Ass4.3]{4.3} below, the $l$-th row and $q$-th column of the asymptotic variance of the original moment conditions $V_g$ is given by
where $I_{q=Q}$ is an indicator function that is equal to $1$ if and only if $q = Q$. I define $\hat{V}_M$ to be a Series HAC estimator and $\hat{V}(\hat{\theta},\hat{\eta}) = \hat{V}_M(\hat{\theta},\hat{\eta})/\hat{\eta}_Q^2$. As mentioned before, fixed-smoothing asymptotics provide a better approximation for SCE applications because of the small sample sizes and high degree of temporal dependence in the data. Furthermore, the values of the smoothing parameter $K$ are often small in applications. For example, in the empirical application below, $K$ is equal to $4$ when the method of Sun2013 is used. Therefore, I use the method suggested in Section (ref) for such cases by not using $\hat{V}$ when constructing the penalty function $\hat{f}$. Also, for testing a null hypothesis $H_0 : \beta_{0,n} = \Bar{\beta}$, I use the test statistic $\sqrt{\min \{T_0,T_1\}} (\Tilde{\beta}_{GMM} - \Bar{\beta})/\sqrt{\hat{V}(\hat{\theta},\hat{\eta})}$ and critical values from a $t$ distribution with $K$ degrees of freedom. Even if the penalty function does not depend on the long-run variance estimator, we can still choose the penalty function to be equal to an upper bound of the asymptotic variance as a function of the nuisance parameters, where the upper bound is tight in special cases. Because $\hat{g}$ is linear in $\delta$, each of the elements of the asymptotic variance of the original moment conditions $V_g$ involves a quadratic form $\delta' \Omega \delta$ for some positive definite matrix $\Omega$. Also, since $V(\theta,\eta) = \eta V_g(\theta) \eta'/\eta_Q^2$ where $V_g(\theta)$ is positive definite for any fixed $\theta \in \Theta_n$, choosing $\hat{f}(\theta,\eta) = f(\theta,\eta) = ||\delta||_2^2 + ||\eta||_2^2/\eta_Q^2$ minimizes an upper bound on $V(\theta,\eta)$ and does not involve $\hat{V}(\theta,\eta)$. However, $||\eta||_2^2/\eta_Q^2 = \sum_{q=1}^{Q-1} (\eta_q/\eta_Q)^2 + 1$ may not have a unique minimum on $$ \{ \eta \in \mathbb{R}^Q : \eta
= 0 \}.$$ Therefore, I normalize $\eta_Q = 1$. I show in the proof of Proposition \hyperlink{P4.1}{4.1} that this allows $\eta_{0,n}$ to be identified. Then the penalty can simply be set to $\hat{f}(\theta,\eta) = f(\theta,\eta) = ||\delta||_2^2 + ||\eta||_2^2$ and have the parameter space for $\eta$ be equal to $H = \{ \eta \in \mathbb{R}^Q : \eta_Q = 1 \}$. I mention when the upper bound being minimized is tight below when discussing Assumption \hyperref[Ass4.3]{4.3}. In order for Assumption \hyperref[Ass3.1]{3.1} to be satisfied, I impose the following conditions:
Assumption 4.2\phantomsection As $T_0,T_1,J \rightarrow \infty$ while $Q$ is fixed,
Assumption \hyperref[Ass4.2]{4.2.1} ensures that the instruments satisfy the exclusion restriction, and Assumption \hyperref[Ass4.2]{4.2.3} guarantees that the instruments are relevant and that there are enough of them to identify $D_{0,n}$. Furthermore, following the reasoning discussed in Remark \hyperref[R3.1]{3.1}, by setting $D_n = \Delta^J$, the rate of convergence conditions in Assumption \hyperref[Ass4.2]{4.2} guarantee that the sample moment conditions converge to the population moment conditions at a rate of $\log (J)/\sqrt{\min \{T_0,T_1\}}$ uniformly in $\delta \in D_n$. The second part of Assumption \hyperref[Ass4.2]{4.2.2} guarantees that the set $D_{0,n}$ is strongly partially identified. Assumption \hyperref[Ass4.2]{4.2.5}, imposes that once enough control units are added, the factor loadings of the treated unit $\mu_0$ fall in the convex hull of the factor loadings of the control units. Note that the factors may not have the same average values before and after treatment, which allows for cases where there is a dependence between the values of the latent factors and the timing of treatment. In order for $\sqrt{n}\hat{M}(\beta_{0,n},\delta_{0,n},\eta_{0,n}) = \sqrt{\min \{T_0,T_1\}}\eta_{0,n} \hat{g}(\beta_{0,n},\delta_{0,n})$ to be asymptotically normal and for the other conditions of Propositions \hyperref[P2.1]{2.1} and \hyperref[P3.1]{3.1} to hold, I impose the following additional assumption:
Assumption 4.3\phantomsection As $T_0,T_1,J \rightarrow \infty$ while $Q$ and $K$ are fixed,
Assumption \hyperref[Ass4.3]{4.3.3} imposes that the optimal control weights are sparse. This sparsity condition is useful for verifying the identification condition in Assumption \hyperref[Ass3.1]{3.1.5}. It is also useful for obtaining the asymptotic normality of $ \hat{g}(\beta_{0,n},\delta_{0,n})$, since then it is sufficient to have asymptotic normality of sample averages involving the sparse set of control units that receive positive weight. In this context, sparsity of the optimal control weights is plausible for two reasons. First, because there are only $R$ latent factors, it is plausible that there exists $\delta \in \Delta^J$ such that $\mu_0 = \mu_{\mathcal{J}}\delta$ and $||\delta||_0$ is equal to or only slightly larger than $R$. Second, the simplex constraints have a tendency to select such sparse solutions, which is why most empirical applications with many controls find sparse weights. Assumption \hyperref[Ass4.3]{4.3.1} helps us specify what the asymptotic variance of the sample moment conditions is. Along with the mixing condition and requiring a bound on the moments of $\hat{g}(\beta_{0,n},\delta_{0,n})$, this is sufficient to apply a Functional Central Limit Theorem to the partial sums of the sample moment conditions, which allows for the estimator to be asymptotically normal and our $t$-statistic to have a $t$-distribution asymptotically. Also note that the upper bound of the asymptotic variance being minimized is tight in the special case where $\Sigma^{q,l} = I_{J+1}$ for all $q,s \in \{1,...,Q\}$ and $\Sigma_i$ is constant across units. This means that we should generally expect the upper bound to be closer to being tight when idiosyncratic shocks are homogeneous across units and the instruments are not redundant and have a similar level of relevance. Note that for $\delta_{0,n} \in D_{0,n}$, $$\hat{g}(\beta_{0,n},\delta_{0,n}) =
=
.$$ Therefore, using the adaptivity condition of Lemma \hyperref[L2.1]{2.1},
As a result, the asymptotic distribution of the estimator depends only on a finite number of averages of the idiosyncratic shocks and averages of the product of the idiosyncratic shocks with the instruments. When $T_0$ is large relative to $T_1$, the uncertainty in $\Tilde{\beta}_{GMM}$ comes from the idiosyncratic shocks in the post-treatment time periods, whereas when $T_1$ is large relative to $T_0$ the uncertainty in $\Tilde{\beta}_{GMM}$ comes from the instruments and idiosyncratic shocks in the pre-treatment time periods.
Proposition 4.1\phantomsection Suppose Assumptions \hyperref[Ass4.1]{4.1}-\hyperref[Ass4.3]{4.3} hold, $\lambda_\delta$ and $\lambda_\eta$ satisfy $$\max \{\lambda_\delta, \lambda_\eta\}^{1/2}\log(J) \overset{p}{\rightarrow} 0 \text{ and } \log(J) / (\sqrt{\min\{T_0,T_1\}}\min \{\lambda_\delta ,\lambda_\eta \}) \overset{p}{\rightarrow} 0,$$ and $T_1/T_0 \rightarrow a$ for some $a > 0$ as $T_0,T_1 \rightarrow \infty$, then we have that $$\sqrt{\min \{T_0,T_1\}} (\Tilde{\beta}_{GMM} -\beta_{0,n})/\sqrt{V} \overset{d}{\rightarrow} N(0, 1) \text{ and }$$ $$\sqrt{\min \{T_0,T_1 \}}(\Tilde{\beta}_{GMM} - \beta_{0,n})/\sqrt{\hat{V}(\hat{\theta},\hat{\eta})} \overset{d}{\rightarrow} t_{K}.$$ For empirical applications, I use an algorithm for selecting $\lambda_\delta$ and $\lambda_\eta$ so that $$\lambda_\delta,\lambda_\eta = O_p(\log(J)\log(\min\{T_0,T_1\})/\sqrt{\min\{T_0,T_1\}}).$$ Therefore, the condition on $\lambda_\delta$ and $\lambda_\eta$ in Proposition \hyperref[P4.1]{4.1} is satisfied as long as $$\log(J)^2\log(\min\{T_0,T_1\})/\sqrt{\min\{T_0,T_1\}} \rightarrow 0.$$ This allows for $J$ to be growing significantly faster than $T_0$ and $T_1$.
\phantomsection
I now explore how my method works in practice by replicating the work of Andersson2019 and then examining its performance in simulations fitted to their data. Andersson2019 evaluate the impact of Sweden's carbon tax on CO2 emissions from transport per capita in the country. The carbon tax was introduced at US \$30 per ton of CO2 in 1990 and increased slightly during the 1990s to US \$44 in 2000. Then, from 2001 to 2004, the rate was increased to US \$109, and as of 2023 it is around \$125. When the carbon tax was implemented, it complemented an existing energy tax and there was also an addition of a Value-Added-Tax (VAT) of 25 percent in 1990. The primary treatment effect they are interested in is the combined effect of the carbon tax and VAT starting in 1990 on CO2 per capita. While we could view this as a case with two continuous treatment variables (the carbon tax rate and the VAT tax rate), if we are only interested in estimating the average difference between actual CO2 per capita and CO2 per capita without the carbon tax and VAT, then we can view the carbon tax and VAT together as a single binary treatment. Andersson2019 uses pre-treatment time periods of 1960 through 1989 and the post-treatment periods of 1990 through 2005. The set of units they use as controls are the 14 OECD countries: Australia, Belgium, Canada, Denmark, France, Greece, Iceland, Japan, New Zealand, Poland, Portugal, Spain, Switzerland, and United States. They arrived at this set of 14 control countries by starting with 24 other OECD countries for which data was available. They excluded 10 of these countries: Ireland, Finland, Norway, Netherlands, Germany, Italy, United Kingdom, Austria, Turkey, and Luxembourg. In the case of Finland, Norway, the Netherlands, Germany, Italy, and the United Kingdom, their justification was based on these countries implementing a carbon tax or making significant changes to their fuel taxes. Because these units were experiencing similar interventions, using them as controls could lead us to underestimate the effect of the policies in Sweden. However, due in part to Sweden being one of the first countries to adopt a carbon tax, all these other interventions happened in the post-treatment years of 1990 to 2005. Since the post-treatment data for the instruments are not used, these policy changes do not necessarily present a problem for using these countries' CO2 per capita from transport data as instruments. In the case of Austria and Luxembourg, their justification for excluding them was based on concerns of “fuel tourism". They exclude Turkey because its CO2 emissions data was significantly different from the other OECD countries throughout the entire sample, and they exclude Ireland based on economic shocks that happened in the post-treatment time periods which did not also occur in Sweden.\footnote{More specifically, they cite the Celtic Tiger expansion period in Ireland.} I estimate the Orthogonalized SCE using the same set of controls as Andersson2019 and the set of instruments being Ireland, Finland, Norway, the Netherlands, Germany, Italy, the United Kingdom, and a constant, although I find similar results when also excluding Ireland from the set of instruments.
In the main specification of Andersson2019, their predictor variables are CO2 from transport per capita in 1970, 1980, and 1989, as well as GDP per capita, motor vehicles (per 1,000 people), gasoline consumption per capita, and urban population averaged for the period 1980–1989. They weight these predictors using the approach of Abadie2010. As is common when the SC's weights are constrained to be a convex combination of the control units, the control weights they find end up being rather sparse with only 6 of the 14 control countries included receiving weight greater than 1%: Belgium (0.195), New Zealand (0.177), Denmark (0.384), Greece (0.090), Switzerland (0.061), and the United States (0.088). Using this SC as their counterfactual, they estimate an effect of -0.29 metric tons of CO2 emissions per capita in an average year, which is a 10.9 percent reduction, for the 1990–2005 period. Aggregating over the total population and the 1990-2005 period, the total cumulative reduction in emissions for the post-treatment period is 40.5 million tons of CO2. They perform several placebo tests and robustness checks, including the popular placebo test of AandG and Abadie2010. This method involves additionally estimating a synthetic unit for every control unit using the other control units. A test statistic is then constructed for the treated unit and every control unit by calculating the MSPE (mean squared prediction error) of each synthetic unit in the post-treatment time periods, and then either dividing by the synthetic unit's pre-treatment MSPE or excluding certain synthetic units with especially large pre-treatment MSPE. A p-value can then be calculated by looking at what quantile the treated unit's test statistic falls in. In Andersson2019, when they exclude the synthetic units with a pre-treatment MSPE at least 20 times larger than Synthetic Sweden’s pre-treatment MSPE, it leaves 9 control countries. The gap in emissions for Sweden in the post-treatment period is the largest of all remaining countries, giving a p-value of 1/10 = .1. When using the ratio of post-treatment MSPE to pre-treatment MSPE, the p-value is 1/15 = .067. In both cases, the test statistic is the most extreme for the actually treated unit, but the p-value fails to fall below the most common thresholds for statistical significance because of the small number of control units.
When re-estimating the average treatment effect, I use the same set of pre-treatment time periods and post-treatment time periods. The weights of Synthetic Sweden are somewhat less sparse than before with the countries receiving positive weight being: Australia (0.087), Belgium (0.113), Denmark (0.322), Greece (0.089), Japan (0.0190), New Zealand (0.105), Switzerland (0.201), and the United States (0.064). It is unsurprising that the weights are still sparse but with slightly more countries receiving non-zero weight since the penalty function encourages the weights to be more spread out but the simplex constraints are still imposed. That said, the weights are quite similar, with all 6 countries that received positive weight before still receiving positive weight and Denmark still receiving the most weight. The weights on the moment conditions $\hat{\eta}$ are fairly spread out across the pre-treatment moment conditions with the largest weight being placed on the one using the United Kingdom as an instrument and the smallest weight being placed on the one using the Netherlands as an instrument.\footnote{More specifically, the weights on the pre-treatment moment conditions are: Finland (-6.166), Germany (-10.590), Ireland (7.883), Italy (7.912), Netherlands (-0.662), Norway (11.617), United Kingdom (-22.402), and the constant (17.882).} The average effect of the carbon tax and VAT from 1990 to 2005 is also approximately a decrease of 0.29 metric tons of CO2 per capita each year. Where the method introduced here allows for a notable difference is in terms of inference. Using the t-test described above, the p-value for the null hypothesis that these taxes had no average effect on CO2 emissions from transport per capita in the post-treatment time periods is 0.00018.\footnote{When estimating the long-run variance of the moment conditions, the smoothing parameter $K$ is estimated to be 4, illustrating the empirical relevance of the fixed-smoothing asymptotics. For the series $\phi_k(x)$, I choose $\phi_k(x) = \sqrt{2}\sin(2 \pi x k)$ for even $k$ and $\phi_k(x) = \sqrt{2}\cos (2\pi x k)$ for odd $k$. Therefore, Assumption \hyperref[Ass3.2]{3.2.2} is satisfied.} This supports the results of the original paper, by showing that under plausible assumptions, the results would be very surprising if these policies had no effect on average. However, in addition to the method of Abadie2010 not necessarily providing statistically significant results, if the inference methods of ConformalInference or CaoDowd2019 discussed below are used, we calculate p-values greater than .1 causing us to fail to reject the null of no effect at common levels of statistical significance.\footnote{Using the subsampling method with 300 iterations and a subsample size of 10 gives a p-value of .14. Using the conformal inference method with moving block permutations gives a p-value of 0.39. Using the End-of-Sample Instability test gives a p-value of 0.43. Using the t-test cross-fitting method of t-test2022 with $K = 3$ gives a p-value of .009.} This may be due to these methods having lower power in this case.
\phantomsection
The placebo method used by Andersson2019 is commonly used in practice and, as previously noted (see Abadie2010 and Abadie2015), corresponds to a traditional Fisher Randomization Test when treatment is randomly assigned. While this would mean that this test have exact size from a design-based perspective, this condition is unrealistic in most current SCE applications. Several other methods of inference for SCEs have recently been proposed in addition to the method of Abadie2010. I discuss the differences between these methods and the $t$-test using the Orthogonalized SCE and then compare their performance in simulations. The method that is most comparable to mine is the t-test cross-fitting procedure of t-test2022, where the pre-treatment time periods are split into $K$ blocks and control weights $\hat{\delta}_{k}$ are estimated using OLS withholding the $k$-th block, $H_k$. This handles the bias in estimating the control weights with OLS by subtracting $\sum_{t \in H_k} (Y_{0t} - Y_{\mathcal{J},t}' \hat{\delta}_{k})/|H_k|$ from $\sum_{t \in \mathcal{T}_1} (Y_{0t} - Y_{\mathcal{J},t}' \hat{\delta}_{k})/T_1$, giving $K$ different estimates $\hat{\beta}_k$ which can be averaged: $\hat{\beta}= \frac{1}{K} \sum_{k=1}^K \hat{\beta}_k$. Their test statistic then relies on standardizing $\hat{\beta}$ using the variation across the estimates $\hat{\beta}_k$. Following the suggestion of t-test2022, I use $K = 3$ in the simulations. While ImperfectFit show that using OLS to estimate $\delta$ can lead the SCE to be asymptotically biased, t-test2022's debiasing approach can fix this while avoiding the need for instruments. One similarity between our approaches is that their test statistic asymptotically follows a t-distribution and does not require consistently estimating the asymptotic variance of the estimated ATT. t-test2022 prove that their estimator is asymptotically normal, despite not orthogonalizing with respect to the weights. However, t-test2022's result relies on each of the sets of control weights $\hat{\delta}_k$ being approximately independent of the shocks that occur in $\mathcal{T}_1$ and $H_k$ and the bias of the OLS estimator being the constant over time. In cases where the timing of treatment is influenced by the values of the latent factors, we would usually expect the bias in the post-treatment time periods to be different from in the pre-treatment time periods. Using Neyman orthogonalization allows us to achieve asymptotic normality without having to impose these conditions. Li2020 proposes a subsampling method that they show has asymptotically correct size when both $T_0$ and $T_1$ are large. However, they use an I.I.D. subsampling method rather than a block subsampling method which requires stronger independence conditions and ProblemWithSubampling show that subsampling and m out of n bootstrap methods can have incorrect asymptotic size when the parameter is close to the boundary of the parameter space. Also, they use the OLS-SCE estimator, which, as mentioned, can be asymptotically biased in the factor model setting. ConformalInference provide a method for conformal inference that can be used with different SCEs provided that the estimator satisfies certain consistency conditions. For the simulations, I use the moving block version of the conformal inference method in order to better deal with the temporal dependence of the observations. The End-of-Sample Instability test was originally introduced by Andrews2003, suggested for SCE by SCandInference, and formally analyzed and extended by CaoDowd2019. It involves testing for a structural break in the sequence of differences between the treated unit and SC. CARVALHO2018's Artificial Counterfactual (ArCo) method uses a LASSO-based estimator of the ATT that is asymptotically normal, but their inference method relies on consistently estimating the long-run variance, unlike my method and the $t$-test of t-test2022. For the placebo method of Abadie2010, I include the version that excludes the synthetic units with a pre-treatment MSPE at least 20 times larger than the treated unit's pre-treatment MSPE. Using the version that uses the ratio of post-treatment MSPE to pre-treatment MSPE provides similar results, except has higher power when $\alpha = .1$.\footnote{For this inference method in the simulations, I also estimate the weights using OLS.} ConformalInference's and CaoDowd2019's methods require stronger conditions on the idiosyncratic shocks, but they both have the potential advantage that the sizes of their tests are asymptotically correct when $T_1$ is fixed and only $T_0 \rightarrow \infty$. This suggests that they may be preferable when $T_1$ is very small. On the other hand, these methods and the placebo method are designed to test the sharp null of no effect in every post-treatment time period, rather than testing a null hypothesis about the ATT. While it depends on context, usually testing the sharp null hypothesis is of less interest. When conducting the simulations for the sizes of the tests, $\beta_t = 0$ for all $t \in \mathcal{T}_1$, so both the sharp null and the null hypothesis of $\beta_{0,n} = 0$ are true. Also, the asymptotic results of Li2020 and CaoDowd2019 rely on keeping the number of control units fixed as the number of time periods grows, which can provide a poor approximation in applications where the number of controls is not small relative to the number of time periods. I conduct the simulations by fitting a linear factor model to the pre-treatment CO2 emissions from transport per capita data from Andersson2019. I first estimate the number of factors using the Singular Value Thresholding method of OptimalSVThresholding.\footnote{Here, the number of factors is chosen to be equal to the number of singular values greater than the median singular value times 2.858.} Using this method, I estimate that there are five factors. I then estimate the factor loadings and factor realizations using Principal Components Analysis. In order to allow the number of time periods to vary, I fit the estimated values of the factors $\hat{f}$ and the residuals $\hat{\epsilon}_{it} = Y_{it} - \sum_{r =1}^5 \hat{f}_{tr}\hat{\mu}_{ri}$ to models and use these models for a parametric bootstrap. For the factors, I fit each factor to an ARIMA model.\footnote{This is done using the auto.arima() function in the forecast package in R, where AIC is used for model selection and QMLE is used to estimate the parameters.} For the idiosyncratic shocks, I sample them from mean-zero normal distributions that are independent across both unit and time. I also set each of the variances of the idiosyncratic shocks to be the same over time but allow for heteroscedasticity across units by using the sample variance of $\{\hat{\epsilon}_{it}\}_{t=1960}^{1989}$ for each $i$, which is consistent with my formal results. The factor loadings are fixed across the simulation draws.
Figure (ref) shows the histogram and Quantile-Quantile plot for the estimates of $\beta_{0,n}$ for the Orthogonalized SCE from 10000 replications when the sample size is the same as in Andersson2019. We can see that even with this relatively small sample size, the normal approximation holds relatively well with only slightly greater concentration near zero and slightly more outliers than would be expected. Using the Anderson-Darling test for normality gives a p-value 0.05092.
Table (ref) contains the size results for the inference methods discussed above when the nominal size is $.1$, $.05$, and $.01$. Looking at the results, we can see that the rejection rates for the Orthogonalized SCE are generally below the nominal levels, except when the number of post-treatment time periods is very small. Since the asymptotic results involve both $T_0$ and $T_1$ growing, it unsurprisingly that the method can fail to control size when either $T_0$ or $T_1$ is very small. The conformal inference method, whose formal results are based on $T_1$ being fixed while $T_0$ grows, does succeed in controlling size even when $T_1$ is small. The placebo method never calculates a p-value below .05 by construction because $J$ is too small, and is also very conservative when $\alpha = .1$. On the other hand, the cross-fitting t-test, the End-of-Sample Instability test, ArCo method, and Subsampling method all experience over-rejection, with it generally being the largest for the Instability test.
Table (ref) contains the power for the same set of inference methods. The simulations are done under the same conditions as before, but now under the alternative hypothesis of $\beta_{t} = -.25$ in each post-treatment time period so $\beta_{0,n}$ is approximately equal to its estimated value in the empirical application. In Appendix \hyperref[ApC]{C}, I also include the size-adjusted power. The power is adjusted for size by finding the threshold for the p-value that makes the method's actual size equal to the nominal size under the null, and then seeing how often the p-value falls below this threshold under the alternative. While this adjustment is infeasible to do in application, it is useful for comparing the power of inference methods with very different sizes. The conformal inference method, subsampling method, and the placebo method generally have the lowest power, although the power of the conformal inference method starts to increase when $T_1$ is very small and the power of the placebo method would likely increase if $J$ was larger. The cross-fitting t-test and the ArCo method have similar, or in some cases larger, power than the Orthogonalized t-test. After adjusting for their over-rejection, the power is lower for the cross-fitting t-test but still larger in some cases for the ArCo method. Overall, however, the Orthogonalized SCE t-test generally has the highest power of the tests that control for size in these simulations. One potential point of caution is that in small samples, the p-values may be fairly sensitive to the choice of the smoothing parameter $K$ for the Series HAC estimator. In the simulations, I use the method of Sun2013 to choose $K$ and it appears to perform quite well. In applications, it may be worth checking the robustness of statistical significance of results to small changes in $K$. Also, one obvious drawback of this method is that it requires a set of units to be used as instruments while the others do not. This motivates further investigating what are alternative sets of moment conditions that can identify the ATT in such cases, such as using shift-share instruments as suggested by SyntheticIV.
In section (ref), I focused on the traditional SCE case, where there is a treated single unit that receives some binary intervention. However, the method can be extended to other contexts, provided that there is a set of moment conditions that identify whatever function of treatment effects is of interest. For example, if there is a set of treated units with indices in $\mathcal{N}_1$ who become treated at the same time, then the same moment equations could be used with $Y_{0t}$ replaced with $\sum_{i \in \mathcal{N}_1} Y_{it}/|\mathcal{N}_1|$. In a staggered adoption setting, this could then be extended by estimating separate control weights for each treatment block while using units in other treatment blocks as instruments.\footnote{Since units need to be untreated during the time periods for which they are being used as instruments, it would be important for each instrument's moment equation to only use time periods in which that instrument is untreated.} I now discuss other applications of the method, as well as potential extensions and limitations.
Another potential application of this method is estimation of a scalar regression coefficient $\beta$ using many instruments. Consider the linear model
where $Y$ is our outcome variable, $X$ is an endogenous explanatory variable, and $\epsilon$ is unobserved. We could also extend this to include covariates $W$ if we use that Frisch–Waugh–Lovell Theorem to project $Y$ and $X$ onto the orthogonal complement of $W$. Suppose we have a set of instruments $Z_i = (Z_{1i},...,Z_{Ji})$ which each satisfy the exclusion restriction so $E[Z_{ji} \epsilon_i] = 0$ for each $j \in \{ 1,...,J\}$, but many of them may be weak or irrelevant. We can use the moment condition
Because the instruments satisfy the exclusion restriction, $g(\theta) = E[(Z_i \delta)X_i (\beta_{0,n} - \beta)]$. Thus, $g(\theta) = 0$ if and only if $E[(Z_i \delta) X_i] =0 $ or $\beta = \beta_{0,n}$. The first option can be ruled out by choosing $\hat{\delta}$ to minimize our estimate for the asymptotic variance of $\Tilde{\beta}_{GMM}$ (i.e., $\hat{f}(\theta,\eta) = \hat{V}(\theta,\eta)$), since this expression diverges if the linear combination of instruments is chosen to be irrelevant. As a result, it is still true that $g(\beta,\delta_{0,n})$ identifies $\beta$. For this application, the moment condition is already orthogonal with respect to $\delta$ when $\beta = \beta_{0,n}$, so there is no need to perform the orthogonalization step. In other words, we can let $\hat{\eta} = \eta_{0,n} = 1$, so that the convergence condition for $\hat{\eta}$ is trivially satisfied. Since this is again a case where solving for both $\Tilde{\beta}_{GMM}$ and $\Tilde{\beta}_{OS}$ is equivalent to setting a single moment condition equal to zero, our estimate has the standard two-stage least squares form $$\Tilde{\beta} = \frac{\sum_{i=1}^n (Z_i \hat{\delta})Y_i/n}{\sum_{i=1}^n (Z_i \hat{\delta}) X_i/n}.$$ If the data are I.I.D., its asymptotic variance is the limit of $$V(\theta_{0,n}) = V_g(\theta_{0,n}) /E[(Z_i \delta_{0,n})X_i]^2 = E[((Z_i \delta_{0,n})\epsilon_i)^2]/E[(Z_i \delta_{0,n})X_i]^2.$$ However, this function $V(\theta)$ is not a convex function of $\delta$, making it more challenging to verify the conditions of Assumption \hyperref[Ass3.1]{3.1} for this choice of the penalty function. However, since linear instrumental variables models are well studied, we can use the existing result in the literature\footnote{See, for example, belloni2018highdimensional.} that the efficient linear combination of the instruments will be equal to $\delta_{0,n} = E[Z_i'X_i] E[Z_i'Z_i \epsilon_i^2]^{-1}$, or equivalently characterized by $E[Z_i' Z_i \epsilon_i^2]\delta - E[Z_i'X_i] = 0$. Therefore we could instead choose $$f(\theta) = ||E[Z_i' Z_i (Y_i - X_i \beta)^2]\delta - E[Z_i' X_i]||_2^2,$$ so that $f(\beta_{0,n},\delta)$ is a quadratic function of $\delta$ and uniquely minimized at $\delta_{0,n} = E[Z_i'X_i] E[Z_i'Z_i \epsilon_i^2]^{-1}$. So we can define
For the choice of $D_n$, if $Z_i$ is high-dimensional and $\delta_{0,n}$ is sparse, imposing the restriction that $D_n = \{\delta : ||\delta||_1 \leq \lambda_1 \}$, where $\lambda_1$ is an additional tuning parameter, can help us ensure that the necessary rate of convergence for $\hat{\delta}$ is achieved similarly as in the SCE case and as discussed in Remark \hyperref[R3.1]{3.1}. On the other hand, if $J$ is smaller than $n$ and there are many weak instruments so that $\delta_{0,n}$ is likely not sparse, alternative constraints, such as constraining the L2 norm of $\delta$, may be preferable. Other estimators have been proposed, such as BelloniChernozhukovHansen2012 who use a LASSO penalty in the first stage for the case of a sparse number of strong instruments and HANSEN2014 who use a Ridge penalty in the first stage for the case of many weak instruments. Further work is needed to compare the performance of this estimator to existing options and explore variations (e.g., using a jackknife estimator for $\Tilde{\beta}$).
This method could also be applied in cases where $\theta$ is point-identified but constrained. Generally, constraints can present a complication for inference if $\theta_{0,n}$ is at or close to the boundary of the parameter space, but the Neyman orthogonalization allows us to handle constraints on the nuisance parameter and the one-step estimator $\Tilde{\beta}_{OS}$ allows for the parameter of interest to be at or near the boundary of the parameter space. One example is models with control variables that have sign-restricted coefficients, and the number of controls can be large if regularization can be used to achieve the rate of convergence requirements in section (ref). This can also be applied to random coefficient models, such as that of BLP, since the variance parameters are constrained to be non-negative. So far, I have assumed that the parameter of interest is point-identified. However, ideas from this procedure may still be of use when $\beta$ is also partially identified. In some cases, it may be possible to reparametrize the model in order to obtain a point-identified subvector we want to conduct inference on. Alternatively, if we wish to test a null hypothesis that some vector $\Bar{\beta}$ is in the identified set for $\beta$, then we could use $n\hat{M}(\Bar{\beta},\hat{\delta},\hat{\eta})'\hat{V}_M(\Bar{\beta},\hat{\delta},\hat{\eta})^{-1}\hat{M}(\Bar{\beta},\hat{\delta},\hat{\eta})$ to form a test statistic, where $\delta_{0,n}$ and $\eta_{0,n}$ can be chosen to minimize $\hat{V}_M(\Bar{\beta},\delta,\eta)$ since this may increase the power of this test. While full vector partial identification can easily arise in models where the number of endogenous variables exceeds the number of instruments, partial identification of $\theta$ can also arise when the number of moment conditions exceeds the number of parameters to be estimated. For example, Chalak2024HigherOM study measurement error models where the latent factors and the observable proxies for the latent factors can both directly affect the outcome variable. They find that under independence conditions on the errors, higher order moment conditions can be used to partially identify the entire vector. Here, the identified set is a finite set of points. If we are interested in testing a particular null hypothesis for a subvector we could take the same approach discussed above. However, in this case, identification for the whole vector or a subvector can be achieved by placing sign restrictions on either the entire vector of parameters or a subvector respectively. This can place us in the case where $\beta_{0,n}$ is point-identified and $\theta$ is constrained, in which case this method can be employed as previously discussed. This method could also be extended to cover cases where $\delta$ is infinite dimensional. The idea of using a regularized estimate of a nuisance parameter when it is partially identified has been explored in the literature on non-parametric instrumental variables regression. For example, SANTOS2011 develops an asymptotically normal estimator for the slope coefficients in a partially linear instrumental variables model using a penalized non-parametric estimator of the infinite-dimensional parameter. In the case of a partially linear instrumental variables model where both right-hand-side variables are endogenous, FlorensEtAl2012 show that the completeness condition is necessary to identify the infinite-dimensional parameter, but it is not needed for the identification of the linear component. Relatedly, partial identification can also arise from ill-posed inverse problems. BabiiFlorens2017 show how spectral regularization can be used in this case and ChenPouzo2009 use penalization to deal with the ill-posed problem arising from discontinuities.
\phantomsection
Assumption \hyperref[Ass3.1]{3.1.4} can be considered a strong partial identification condition, as it bounds how far values of $\delta$ can be from its identified set in terms of how small the value of $\delta$ makes the population moment conditions at $\beta_{0,n}$. This means the results in the previous sections allow for $\delta$ to be strongly point-identified, strongly partially identified, or completely unidentified, but cases of weak identification have been ruled out. Weak identification of $D_{0,n}$ can present a problem for this method because it relies on being able to consistently estimate an element of the identified set $\delta_{0,n}$ which, roughly speaking, relies on being able to consistently estimate the identified set $D_{0,n}$. This cannot be done when $D_{0,n}$ is weakly identified, but whether $\hat{\delta}$ is converging to an element of $D_{0,n}$ be may still influence the asymptotic distribution of $\Tilde{\beta}_{GMM}$ and $\Tilde{\beta}_{OS}$. To illustrate the problem, consider the example of the non-linear regression model, $$Y_i = \beta_{0,n} h(X_i ,\delta_{0,n}) + U_i,$$ where $Y$ and $X$ are observed and the function $h$ is known. Suppose the moment conditions $$g(\beta,\delta) = E[X_i U_i] = E[X_i(Y_i - \beta h(X_i,\delta))] \text{ and } \hat{g}(\beta,\delta) = \sum_{i=1}^n (X_i(Y_i - \beta h(X_i,\delta)))/n$$ are used. Here, $D_{0,n} = \mathbb{R}^J$ when $\beta_{0,n} = 0$ and $D_{0,n}$ contains a single element $\delta_{0,n}$ otherwise under restrictions on $h$. When $\beta_{0,n}/\lambda_\delta \rightarrow b \in [0, \infty)$, the feasible values for $\hat{\delta}$ generally expand to include all of $\mathbb{R}^J$, whereas when $\beta_{0,n}/\lambda_\delta \rightarrow \infty$, $\hat{\Theta}_0$ shrinks to just include a single point. While standard estimators of $\beta$ are asymptotically normal when $\sqrt{n}\beta_{0,n} \rightarrow \infty$, they are $\sqrt{n}$-consistent but not asymptotically normal in the weak identification case of $\sqrt{n}\beta_{0,n} \rightarrow b \in [0,\infty)$ (see, AndrewsCheng2012 and HanMcCloskey2019). When $\delta$ is weakly identified, the orthogonalized estimators $\Tilde{\beta}_{GMM}$ and $\Tilde{\beta}_{OS}$ are $\sqrt{n}$-consistent by a similar reasoning as with standard estimators. However, they are generally not asymptotically normal. For example, in the case with $\sqrt{n}\beta_{0,n} \rightarrow b$, $\beta_{0,n}/\lambda_\delta \rightarrow 0$ so $\hat{\delta}$ converges to the global minimizer of $f(\beta_{0,n}, \delta,\eta)$ on $\mathbb{R}^J$, say $\delta^*$. This may in turn result in $\hat{\eta}$ converging to a different value $\eta^*$. Using the adaptivity condition, we would then expect that $\sqrt{n}\hat{M}(\beta_{0,n},\hat{\delta},\hat{\eta}) = \sqrt{n}\hat{M}(\beta_{0,n},\delta^*,\eta^*) + o_p(1)$. While $\sqrt{n}\hat{M}(\beta_{0,n},\delta^*,\eta^*)$ may still be asymptotically normal, it is not centered around zero unless $\sqrt{n} \beta_{0,n} \rightarrow 0$. As a result, $\Tilde{\beta}_{GMM}$ and $\Tilde{\beta}_{OS}$ generally have non-standard limiting distributions. On the other hand, when $\sqrt{n} \beta_{0,n} \rightarrow \infty$, it is possible to consistently estimate $\delta_{0,n}$, so the method works similarly to before. Hence, the problem arises when $\sqrt{n}\beta_{0,n} \rightarrow b \in (0,\infty)$. Intuitively, this is because we either want to be close enough to the unidentified case where the true value of $\delta$ is irrelevant (i.e., $\sqrt{n}\beta_{0,n} \rightarrow 0$ so $\sqrt{n}M(\beta_{0,n},\delta^*,\eta^*) \rightarrow 0$) or close enough to the strongly identified case where we can accurately estimate $\delta_{0,n}$ (i.e., $\sqrt{n}\beta_{0,n} \rightarrow \infty$ so $||\hat{\delta} - \delta_{0,n}||_1 \overset{p}{\rightarrow} 0$).