Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,384 characters · 10 sections · 46 citation commands
{3pt} {3pt}
The prevalence of big data and high-dimensional covariate spaces has grown significantly in applied economics (see jin2015significance for an overview). Consequently, the econometric literature strives to develop suitable tools to cater to the needs of applied research (e.g., belloni2014inference, farrell2015robust, belloni2017program, ning2020robust). One widely employed framework in this context is the $l_1$ regularization, known as lasso, which facilitates model selection and feature extraction by regulating covariate parameters. fan2006statistical provide an extensive discussion on the pivotal role of model selection for knowledge discovery and addressing high-dimensionality challenges.
Rigorous sparse model selection methods offer key advantages in terms of prediction accuracy and interpretability tibshirani1996regression. Beyond prediction tasks, lasso techniques have found substantial utility in causal analysis. Even though for causal identification strategies that rely on the unconfoundedness assumption rosenbaum1983central, the role of covariates narrows down to including all confounding variables in the data, model selection nevertheless offers significant features:
1) Common estimators often rely on estimating an outcome/propensity score model to evaluate unobserved potential outcomes or as a nuisance function. Precise estimation of those underlying models can lead to significant gains in prediction accuracy for causal estimands in finite samples.
2) In the absence of theoretical guidance on the true model, researchers may face a high-dimensional problem, where the number of parameters exceeds available observations (often referred to as '$p \gg n$' or 'The Curse of Dimensionality'). In such scenarios, estimators like the linear regression family are inapplicable due to insufficient degrees of freedom, necessitating dimensionality reduction.
3) The econometric literature has proposed various variants of post-regularization estimation methods chernozhukov2015valid. If the initial model selection step achieves sufficient sparsity, a second non-regularized estimation of the resulting parsimonious model may even become the oracle estimator when the true model is identified belloni2013least.
4) Recent years have witnessed a significant surge in interest in estimating and analyzing heterogeneous treatment effects kunzel2019metalearners. A model selection feature can offer valuable support and guidance for selecting potential heterogeneity dimensions.
All the aforementioned points are strongly supported by the prevalence of low-rank matrices in high-dimensional data, as demonstrated in udell2019big. It has been shown that lasso can consistently identify such sparse models while achieving convergence at the optimal rate concerning the parameters of interest (known as the oracle property) under specific conditions meinshausen2006variable,zou2006adaptive,meinshausen2007relaxed.
Our main contribution lies in extending the use of $l_1$ regularization to matrix completion estimators for causal panel data models. In their seminal work, athey2021matrix lay the foundation for employing matrix completion methods in causal panel data models. This estimator utilizes observed elements of control outcomes to impute the unobserved (missing) potential outcomes of treated observations in the untreated state. By interpreting the panel data structure as a matrix, this procedure approximates the complete matrix of control state potential outcomes by regularizing the complexity of the underlying factor model using the nuclear norm. Matrix completion estimation has been shown to outperform synthetic control estimators or regression methods based on the unconfoundedness assumption athey2019ensemble, athey2021matrix, and its initial applications have shown its practical value wood2020procedural, rafaty2020carbon, levy2022effects. Appendix (ref) provides background of the matrix completion estimator and an embedding in the context of causal panel data models.
In this work, we leverage the convex nature of nuclear norm minimization to regularize the rank of the general factor matrix. In a model with covariates, this allows for concurrent regularization of a potentially high-dimensional set of variables. Unlike other estimation methods, where $l_1$ regularization often increases the numerical complexity of an estimator by lacking a closed-form solution, preserves $l_1$ regularization of covariate parameters the convex optimization of rank regularization in matrix completion estimators, thus adding no additional complexity. The addition of $l_1$ regularization to the covariate space enhances the matrix completion framework with a model selection property, thus enhancing its performance in the context of causal panel data analysis without increasing the computational complexity. The methodology allows for a two-step approach similar to common post-regularisation estimators: Involving the selection of the optimal model as the first step and conducting a second estimation without covariate regularization ensuring unbiased estimates of model parameters.
We adopt the permutation-based inference procedure presented in chernozhukov2021exact to test the sharp null hypothesis of a zero treatment effect, accommodating a stationary and weakly dependent shock process. The validity of this inference procedure is established not only for the proposed estimator with integrated $l_1$ covariate regularization but also, as a second contribution, for settings with any treatment assignment scheme. This contribution extends the set of potential applications of matrix completion methods with integrated model selection properties by enabling their use in diverse settings with various treatment assignment mechanisms. This advancement enhances the robustness and versatility of our proposed approach, making it well-suited for a wide range of real-world scenarios.
A numerical simulation using the proposed estimator demonstrates that the introduced $l_1$ regularization on the covariate space within the matrix completion estimator effectively reduces the determined model size, as intended. Employing a cross-validation step to determine the optimal penalty parameter value justifies adopting a 1se optimality criterion to obtain the ideal model size. The resulting model sizes using the 1se condition are shown to be precise and appear to converge to the true values of the underlying model for large sample sizes..
The proposed estimator outperforms the matrix completion estimator without covariate regularization, in particular for small sample sizes or strong signal-to-noise ratios, and it remains equally precise in other settings. The permutation-based approach necessitates enforcing the null hypothesis for exact and valid inference. However, this requirement may lead to a biased treatment effect estimation if the null hypothesis does not hold. In Section (ref), we shed light on how this impacts the estimation procedure and identify the conditions under which a simple rule-of-thumb correction of the treatment effect estimate can be applied. The simulation results reveal that this correction partially mitigates the downward bias and achieves equal estimation accuracy as the unbiased estimates when discarding the null hypothesis required for inference. Executing a two-stage approach with a second non-regularized estimation on the initially determined model does not yield an improved estimation of the treatment effect on the treated.
An illustrative application examining the impact of travel restriction policies implemented during the Covid-19 pandemic on the incidence of infections requiring treatment in intensive care units underscores the advantageous properties of the model selection mechanism. Within a high-dimensional context characterized by a presumed minuscule signal-to-noise ratio, the proposed estimator adeptly identifies a notably sparse model. The empirical application using panel data aligns precisely with the anticipated behavior of the estimator, as derived from the simulation study. The findings reveal that the estimated effect of the mandatory testing requirement upon entry from foreign countries with heightened incidence rates is small coming along with a notably large p-value. This compelling evidence suggests that the influence of this particular travel restriction policy on public health outcomes is negligible.
To facilitate the adoption and replication of our methodology, we have developed an R-package that implements our proposed estimator. The package is publicly available for download on our GitHub repository.\footnote{\url{github.com/heinigersandro/MC_MS}} We encourage researchers to utilize our methodology, validate our findings, and contribute to advancements in the field.
The remainder of the article reads as follows: In Section (ref), we explain how the model selection property is incorporated into the estimator. Section (ref) discusses the theoretical properties of the proposed estimator. The behavior of the proposed estimator in simulations is presented in Section (ref). Section (ref) discusses the results of the illustrative application to Covid-19 infections in Germany. Finally, we conclude in Section (ref). In the Appendix, we provide an introduction to the matrix completion estimator for causal panel data models, additional estimation results and proofs for the theoretical findings.
We consider a panel of $N$ units observed over $T$ periods, where each unit $i$ potentially undergoes a binary treatment in time-period $t$, $W_{it} \in \{0,1\}$, with corresponding potential outcomes $Y_{it}(0) \coloneqq Y_{it}(W_{it}=0)$ and $Y_{it}(1) \coloneqq Y_{it}(W_{it}=1)$. Our object of interest is the average treatment effect on the treated\footnote{As reasoned in Appendix (ref), we assume most observations to be in the untreated state and therefore restrict the focus on the ATET.}, defined as follows:
We leverage the model with covariates proposed in athey2021matrix:
where $\mathbf{Y} \!\in\!\mathbb{R}^{N\!\times\!T}$ represents the realized outcomes, $\bm{\Theta} \!\in\!\mathbb{R}^{N\!\times\!T}$ is the matrix of potentially heterogeneous treatment effects, $\mathbf{W} \!\in\!\mathbb{R}^{N\!\times\!T}$ denotes the treatment allocation, $\mathbf{L}^* \!\in\!\mathbb{R}^{N\!\times\!T}$ represents the unobserved factor matrix of low-rank, $\mathbf{X} \!\in\!\mathbb{R}^{N\!\times\!P}$ contains unit-specific covariates, $\mathbf{Z}\!\in\!\mathbb{R}^{Q\!\times\!T}$ contains time-specific covariates, $\mathbf{V}_{it}\!\in\!\mathbb{R}^{J}$ contains the unit-time varying covariates for unit $i$ at time $t$, $\mathbf{H}^* \!\in\!\mathbb{R}^{P\!\times\!Q}$ and $\bm{\beta}^* \!\in\!\mathbb{R}^{J}$ are the unknown covariate parameters, $\bm{\Gamma}\!\in\!\mathbb{R}^{N}$ represents unit-level fixed effects, $\bm{\Delta}\!\in\!\mathbb{R}^{T}$ represents time-level fixed effects, and $\mathbf{U} \!\in\!\mathbb{R}^{N\!\times\!T}$ is a random shock term.
Following the literature on factor models, we use an additive and linear model in (ref). Despite appearing restrictive, the set of covariates is not limited in size, allowing for the inclusion of any functional forms if desired. More complex interactive terms are absorbed by the latent matrix $\mathbf{L}^*$ and potentially by the fixed effects as well, granting the model (ref) sufficient flexibility to cover a broad range of applications.
The structure of $\mathbf{H}$ in the outcome model (ref) allows only for linked unit- and time-covariate effects. To include linear terms in both $\mathbf{X}$ and $\mathbf{Z}$, a more comprehensive model can be defined as follows:
where $\tilde{\mathbf{X}}=[\mathbf{X}|\mathbf{I}_{N\!\times\!N}]$ and $\tilde{\mathbf{Z}}=[\mathbf{Z}^\top|\mathbf{I}_{T\!\times\!T}]^\top$. For brevity, we focus on the covariate setting as in (ref) since all adaptations to the richer model (ref) are straightforward.
We define the set of treated observations as $\mathcal{M}\coloneqq \left\{(i,t) \text{ with } W_{it}=1\right\}$ and the set of control observations as $\mathcal{O}\coloneqq \left\{(i,t) \text{ with } W_{it}=0\right\}$. Correspondingly, we define the matrices projecting on the respective treatment allocation space as follows:
Using the defined notation and supposing Assumption (ref) holds, we can rewrite the sample representation of the ATET, which is our object of interest, as follows:
To evaluate the estimator (ref), we only need the matrix of potential outcomes in the absence of treatment, denoted by $\mathbf{Y}(0)$. As for each observation, only one potential outcome is realized, some elements in $\mathbf{Y}(0)$ are missing as they are not observed in practice. The outcome model (ref) directly defines the potential outcome model as follows:
We estimate the unknown parameters of the potential outcome model (ref) using a matrix completion method similar to athey2021matrix and chernozhukov2021exact, but with regularization on the full covariate space\footnote{In athey2021matrix, the model with covariates drafts a penalty term on the link matrix $\mathbf{H}$ between unit- and time-varying covariates, but not on the unit-time-varying covariate parameters $\bm{\beta}$. Though, the authors do not discuss the implications of covariate space regularization on the model selection properties of the proposed estimator.}\footnote{In principle, the unit- and time-level fixed effects could be regularized as well. We follow the arguments of hastie2009elements that regularizing the intercept in $l_1$ regularization increases the bias (see also the discussion in athey2021matrix).}. We adopt the permutation-based procedure for valid finite sample inference described in chernozhukov2021exact. The method requires the enforcement of the null hypothesis of a zero treatment effect to obtain a valid and exact inference procedure.\footnote{The intuition for this is that the permutation-based inference requires the estimation to be provably accurate across the whole unit-time space to obtain equally distributed residuals. Enforcing the null allows us to use all information in the data, regardless of the treatment state.} In fact, chernozhukov2021exact show that the rejection rates are substantially erroneous if the null is not enforced. The presumed absence of a treatment effect implies that there is no missing data, which is conterminous to estimating the model parameters using all observations in the panel. This fundamentally differentiates the chernozhukov2021exact estimation approach from athey2021matrix which guarantees finite sample bounds for the effect estimates but no valid inference.
The estimation procedure is as follows:
where $\left\lVert \mathbf{A} \right\rVert_F^2=\sum_{(i,t)} \mathbf{A}_{it}^2$ is the squared Fröbenius norm, $\left\lVert \mathbf{A} \right\rVert_*=\sum_{i=1}^N \sigma_i(\mathbf{A})$ is the nuclear norm, and $\left\lVert \mathbf{A} \right\rVert_{1,e}=\sum_{(i,t)} \left\lvert \mathbf{A}_{it} \right\rvert$ is the element-wise $l_1$ norm.
The choice of the nuclear norm on $\mathbf{L}$ is crucial as it regularizes the rank of the matrix through its singular values $\sigma_i(\mathbf{L})$. Using other norms, such as the Frobenius or element-wise $l_1$ norm, would not be suitable, as the objective function would be minimized by setting $\mathbf{L}_{it}=0$ for all $(i,t) \in \mathcal{M}$ due to the projection on the control space $\mathbf{P}_\mathcal{O}$. The rank norm $\left\lVert \mathbf{A} \right\rVert_0=\sum_{i=1}^N \mathbf{1}\{\sigma_i(\mathbf{L})>0\}$,, which might be a preferred choice, is computationally infeasible as the rank norm optimization is NP-hard recht2010guaranteed. On the other hand, the choice of the other norms is straightforward. The Frobenius norm on the prediction error represents an MSE-like measure, while the $l_1$ regularization on the covariate parameters facilitates the model selection process.
The proposed estimator (ref) can be efficiently computed by iteratively applying a soft-impute step mazumder2010spectral to handle the nuclear norm rank regularization of the matrix of unobserved factors, and a gradient descent step friedman2010regularization with respect to the $l_1$ regularization of the covariate space. It is important to note that both steps remain within the realm of convex optimization, making the computation computationally efficient.
The values of the regularization parameters $\lambda_{L,H,\beta}$ are selected through cross-validation. To ensure that the estimation of the ATET produces precise potential outcomes under no treatment, we aim to find the $\lambda_{L,H,\beta}$ values that minimize the prediction error for potential outcomes. To maintain independence from whether the null hypothesis holds or not, we restrict the cross-validation sample to the observations in the control state. For a specified number of folds $K$, we randomly draw $K$ sets $\mathcal{O}_k\subset \mathcal{O}$ with a size $\frac{\left\lvert \mathcal{O}_k \right\rvert}{\left\lvert \mathcal{O} \right\rvert}=\frac{\left\lvert \mathcal{O} \right\rvert}{NT}$, corresponding to the overall share of untreated observations. The evaluation set for each fold is given by $\mathcal{O}_{k}^{\boldsymbol{-}} \coloneqq (i,t):(i,t)\!\in \!\mathcal{O} , (i,t)\! \notin\! \mathcal{O}_k$. For a chosen set $\Lambda$, we search for the optimal $\lambda$-configuration by minimizing the mean squared prediction error over all partitions as follows:
Here, $(\hat{\mathbf{L}'}_k,\hat{\mathbf{H}'}_k, \hat{\bm{\beta'}}_k, \hat{\bm{\Gamma'}}_k,\hat{\bm{\Delta'}}_k)$ are estimated by (ref) on each training set $\mathcal{O}_k$ using the penalty parameters $\left(\lambda_{L'},\lambda_{H'}, \lambda_{\beta'}\right)$.
To restrict the set of to-be-evaluate elements $\Lambda$, we identify the theoretical minimal values of $\lambda_{L,H,\beta}$ such that all parameter estimates in the respective objects are reduced to 0. Along with $\lambda_{L,H,\beta}=0$, this establishes two natural bounds for each penalization parameter. The optimization search can be conducted on a standard three-dimensional grid. The MCMS R-package utilizes a more advanced hyper-cube search that leverages the convex optimization to expedite convergence to the minimizing configuration. This approach enhances the efficiency of the parameter selection process.
A natural extension of estimators with integrated model selection is to implement a two-stage procedure. In the first stage, the estimation procedure is run as usual to identify the sparse model, and in the second stage, the estimation is performed without regularization on the identified subset of covariates chernozhukov2015valid. The objective of the second stage is to obtain an unbiased estimate of model parameters by dropping the regularization. For the matrix completion method, the analogous two-stage estimator first identifies the informative covariates and determines the rank of $\mathbf{L}$ in the first step. Then, in the second step, it applies an unregularized estimation under the restrictions on the covariate space and the rank of the factor model.
The estimation of the model (ref) is obviously biased if the null hypothesis is violated. In such cases, the treatment effect might be either evenly dispersed over all observations and periods or allocated to confounding variables, leading to a reduction in the ATET estimate. If the assessment of p-values is not a need, an estimation of the model parameters using only the observations in the absence of treatment yields more precise estimates of the treatment effect:
It is important to note that if the treatment effect is independent of the covariates or the unconfoundedness assumption holds, the treatment effect is distributed evenly across all observations. Under the assumption of an independent treatment effect, each observation receives a share of the average effect among the treated sample, given by $\frac{\left\lvert \mathcal{M} \right\rvert}{\left\lvert \mathcal{O} \right\rvert}$. If we assume homogeneity, this share is equal to the ATET. However, in situations with strongly heterogeneous treatment effects, the average effect in the sample may diverge from the ATET. Additionally, the treatment allocation mechanism also affects the dispersion of treatment effects among the observations under the unconfoundedness assumption. The split of the sample average effect assigns more weight to observations with a higher propensity score, as the treatment effect is partly allocated to confounding variables. To correct the ATET for marginal treatment effect and propensity score heterogeneity, a simple rule-of-thumb correction is given by:
While other test statistics based on element-wise p-norms are theoretically possible, the test statistic in Definition (ref) has been found to exhibit good properties when estimating average treatment effects chernozhukov2021exact.
Let the actual and estimated shock-discarded potential outcomes in the absence of treatment be denoted as:
Assumption (ref) is easy to verify for various types of estimators, including the basic non-regularized version of the matrix completion estimator, as shown in chernozhukov2021exact. The following lemma elaborates on the primitive conditions to ensure that the proposed estimator with covariate regularization satisfies Assumption (ref).
We make use of permutations to compute the p-values.
For the finite sample validity of the matrix completion estimator, the choice of $\Pi$ is irrelevant. However it is worth to note that set of one-to-one mappings on $ \{1,\dots,N\}\times\{1,\dots,T\}$ is larger than the set of horizontal moving block permutations allowing for the computation of more precise p-values under Assumption (ref).1.
Note that the results of Theorem (ref) are non-asymptotic, i.e. they hold in finite sample. The finite sample bounds of the size properties imply that the inference procedure is exact in $NT\rightarrow\infty$. Naturally, the power of the hypothesis testing depends on characteristics of the underlying model (ref). For instance, if the model is estimated more accurately or if the shocks have a smaller variance, the hypothesis tests tend to have higher power.
In the following section, the results are only presented for matrix $\mathbf{H}$ that links the unit-specific and the time-specific covariates. All findings equivalently apply for $\bm{\beta}$, the parameter of the unit-time-varying covariates. The corresponding figures are presented in Appendix (ref).
For the subsequent simulations, we use the following data-generating process: The outcome matrix follows the linear model in (ref) with a homogeneous treatment effect\footnote{The matrix completion estimator can naturally deal with heterogeneous treatment effects. As our object of interest is the average treatment effect on the treated, we stay with homogeneous effects for the simulation.}, interactive unit- and time-specific covariates, unit-time varying covariates, unit fixed-effects and time fixed-effect terms. The treatment allocation $\mathbf{W}$ follows a Bernoulli model. The latent factor model matrix $\mathbf{L}$ is random matrix restricted to a specified $\text{rank}_L$ with exponentially distributed singular values. The unit-specific covariates of size $p$ are multi-variate normal with an idiosyncratic unit factor, time-specific covariates of size $q$ are generated accordingly. The coefficients of $\mathbf{H}$, the link between those covariates, are normally distributed, but only a random fraction is active while the other elements are set to 0. The data-generating process omits linear terms in the unit- or time-specific covariates as discussed in Section (ref). The unit-time varying covariates $\mathbf{V}_{it}$ are independent standard normal and of the true coefficients $\bm{\beta}$, again, only a fraction is active. The exact formulations of all components and the default parameter values are described in Appendix (ref).
Figure (ref) shows the paths of the estimated coefficients in $\hat{\mathbf{H}}$ for different values of $\lambda_{H}$. We observe the known pattern of $l_1$ regularisation that coefficients with small attributed parameter values are wiped out for already low penalty parameters. With increasing regularization, one variable by another is eliminated until even the variables with largest absolute coefficient estimates are regularized out. The descriptive illustrations on the size and the number of coefficient parameters in Appendix (ref) confirm that the elimination process closely follows the order of parameter estimates in the absence of very strong correlation patterns among the covariates.
As discussed in Section (ref), the optimal values of the penalty parameters $\left(\lambda_{L}, \lambda_{H}, \lambda_{\beta}\right)$ are chosen by cross-validation. Figure (ref) shows the mean prediction error on the test sample over the cross-validation folds for different values of the regularization parameter $\lambda_H$. The convex optimization along the penalization parameter clearly shows itself.
As in each fold, the objective function is only evaluated on a subsample of the available data, the cross-validation optimization is not necessarily optimal for the out-of-sample observations. A known issue is the estimation error of risk curves using cross-validation samples hastie2009elements. Because of this estimation error, model selection using cross-validation tends to be too conservative while in fact, the smallest model is preferred among a set of indistinguishable models Chen2021. To address this, we implement the common 1se solution, which selects for the largest value of the penalization parameter for which the objective function remains below the minimum value plus one standard error over the cross-validation folds calculated at the optimal $\lambda_H$ position. The results on the number and size of the non-zero coefficients in Appendix (ref) strongly support the application of the 1se solution.
In this subsection, we evaluate the estimation error over multiple samples to remove the dependency of the measured estimation error on a specific randomly generated sample. The plots in the subsequent subsections use the abbreviation as described in Table (ref) to denote different estimation methods.
Figure (ref) shows the boxplot of ATET estimates by different sample sizes for the six versions of the matrix completion estimator as outlined in Table (ref). All estimators using covariate regularization perform a great deal better, in particular for small sample sizes. However, the versions imposing the null hypothesis exhibit a substantial downward bias, which is the price paid for enabling the inference procedure. With a decent sample size, the rule-of-thumb correction is able to eliminate the bias. The matrix completion estimator without imposed null, as in athey2021matrix, but using covariate regularization is unbiased and performs already excellently for very small sample sizes. The fit of the post-regularization model is consistently very poor. It can be shown that the estimates of $\hat{\mathbf{L}}$ are closely tied to the regularization by the penalty term. When dropping the penalization parameter in the post-regularization step, the coefficient in $\hat{\mathbf{L}}$ strongly diverge from the true values.
To undertake the estimation accuracy in a more detailed evaluation, Figure (ref) shows the median indexed squared error of the treatment effect where the squared errors in each sample are divided by the squared estimation error of the baseline imp0 estimator. Hence, a value below 1 denotes that an estimator performs on average better than the basic regularized estimator with imposed null.\footnote{Note that the median indexed squared error is an aggregate of a ratio and has to be interpreted accordingly.} This illustration refines and accentuates the insights from Figure (ref) that the enabled inference procedure by imposing the null hypothesis comes at a substantial price in treatment effect estimation accuracy. With increasing sample size, the rule-of-thumb correction achieves similar precision as the matrix completion estimator without the null hypothesis imposed. Thus, if inference on the treatment effect estimator is desired, applying the rule-of-thumb correction can partially compensate for the accuracy loss by the implied null hypothesis for the inference procedure.
The simulation results further evince that the reduction in model size confers benefits to estimation accuracy, particularly in the context of small sample sizes. The lower degrees of freedom of the fitted model result in better predictions for the potential outcomes under no-treatment of the treated observations with the available data. For larger sample sizes, there is enough information in the data such that fitting the full model without covariate regularization picks up less spurious correlations of non-explanatory covariates such that the estimation accuracy levels with the regularized version of the estimator. Nevertheless, the gains in accuracy achieved through by applying the estimate correction or performing an estimation without inference are still massive compared to the unregularized estimation.
In Appendix (ref), we demonstrate that the bias of the treatment effect estimates is not negatively affected by a low signal-to-noise ratio, in fact, the bias is slightly decreasing for weaker signals. However, for a low signal-to-noise ratio, the variance of the treatment effect estimate is substantial. The relative performance of the different estimator versions reflects the pattern observed in Figures (ref) and (ref): The estimator without imposing the null hypothesis is superior for all settings and the rule-of-thumb correction imp0_rot is consistently the best-performing estimator among the null hypothesis versions, in particular for low signal-to-noise ratios.
We assess the dimensions of the ascertained model, defined by the count of non-zero coefficients in the matrix $\mathbf{H}$, in comparison to the true size of the generated sample across varying sample sizes. It is pertinent to note that the post-regularization and rule-of-thumb correction estimates have no impact on the model size. Consequently, we display only the unregularized estimation, the regularization approach employing mse and 1se optimality criterion, and the estimator not imposing the null hypothesis.
Figure (ref) illustrates a substantial reduction in the number of non-zero coefficients within the estimated matrix $\mathbf{H}$ when utilizing a regularized estimator.\footnote{It is worth acknowledging that the actual count of non-zero coefficients in $\mathbf{H}$ may deviate from the DGP parameter $h_{\text{size}}$ under conditions where $p<\bar{p}$ or $q<\bar{q}$.} Notably, the MSE-optimal cross-validation results in an overestimation, maintaining approximately five times the number of covariates within the determined model across all sample sizes. Conversely, the 1se variant consistently selects a model size that closely aligns with the true model size. As the sample size increases, there appears to be a convergence of the estimated model size towards the true value.
Additionally, in Appendix (ref), we provide evidence that the MSE of the estimated coefficients in $\mathbf{H}$ exhibits a strong decrease as the sample size diminishes. The concurrent enhancements in both model size and coefficient accuracy contribute to the augmented precision in treatment effect estimates for larger sample sizes which has been observed in Section (ref).
In Appendix (ref), we illustrate that the determined model size persists is stable as the signal exhibits substantial strength. Nevertheless, as the signal-to-noise ratios approach very small magnitudes, rendering the task of segregating relevant information from noise more difficult, all estimators tend to identify a reduced number of informative parameters within the model. This discernment, in turn, can lead to estimated models that are more sparse than the true model.
To showcase the usefulness of the proposed estimator in an empirical context, we apply the developed methodology to scrutinize the impact of governmental regulations aimed at mitigating the Covid-19 pandemic. The period encompassing the peak of the global pandemic witnessed an intense public discourse regarding the appropriateness and effectiveness of public health measures designed to curb the spread of SARS-CoV-2. This discourse extended into the academic domain, putting forth extensive evaluations of, e.g., governmental responses to the pandemic Christensen2023Nordic, basseal2023key or public compliance with health regulations scandurra2023people.
In this illustrative analysis, we delve into the effect of international travel restrictions, a measure identified as highly effective by basseal2023key. Our focus is on providing a pure analytical assessment of travel restrictions as a public health intervention. The geopolitical implications of travel bans and the suspension of visa exemptions are thoughtfully explored in seyfi2023covid.
We analyze the impact of Covid-19 testing obligations upon entry from foreign countries with elevated incidence rates on the frequency of infections necessitating treatment in intensive care units (ICU). The testing obligation at entry from high-risk regions was partially implemented in Germany from the summer of 2020 to the spring of 2022.\footnote{The authority for public health regulations in Germany has been dispersed to various administrative levels with a general trend to transition from individual districts to state-harmonized policies over time. Some regions temporarily intensified regulations by imposing testing obligations on all incoming travellers, which is not being distinguished in this analysis.} To circumvent distortions in the outcome measure arising from limited testing capacities during the early stages of the pandemic, we confine the sample period to July 2020 to June 2022, utilizing weekly frequency. Additionally, to ensure robustness in our outcome measure, we exclude districts with fewer than 10 intensive care beds. It is noteworthy that, in the remaining districts, no district operated at maximum ICU occupancy for a significant duration, ensuring consistent reporting and treatment within the same district. With these restrictions, our analysis spans a panel of 342 units over 105 time periods.
The adoption of the ratio between patients treated in ICU due to Covid-19 infections and the number of reported infections as the outcome is appealing for two reasons. Firstly, travel restrictions primarily aim to avert the overwhelming strain on medical facilities by delaying the introduction of new and potentially threatening virus mutations. This delay allows more time for the refinement of vaccines to address emerging variants. Secondly, utilizing the frequency of severe infections as the outcome mitigates common concerns related to endogeneity and reverse causality when using plain incidence rates or vaccination rates as outcomes.\footnote{The median duration of ICU treatment for Covid-19 patients is between 10 to 14 days and exhibits a right-skewed distribution shryane2020length, kaccmaz2023covid. Consequently, during periods of declining infection rates, the ratio of patients admitted to the ICU to the reported incidence may exceed 1. A direct mapping of ICU patients to an infection is not possible. Consequently, the metric employed in this study should be interpreted as a proxy for the severity of infections. For a more comprehensive examination of the epidemiological landscape extending this illustrative example, the adoption of more sophisticated metrics becomes requisite.}
Given the extensive research on the medical, economic, and social impact of the pandemic, comprehensive macroeconomic data is available for the selected time frame. The German Corona-Datenplattform\footnote{\url{https://www.healthcare-datenplattform.de/}} serves as a primary data source, offering a rich collection of macroeconomic, infection, and policy data. The dataset encompasses 91 unit-time-specific covariates (encompassing various public health regulations, vaccination prevalence, and short-time work data), 55 unit-specific covariates (encompassing socio-demographics, medical care, infrastructure, and economic composition variables), and 11 time-specific covariates (covering general economic indicators, mobility, tourism, and the prevalence of threatening mutations among all reported infections). The complete datasets, along with a detailed description of all variables, can be accessed on the Harvard Dataverse DVN/JGGBQG_2024 related to this project.
The provided context and dataset serve to exemplify the estimation and model selection characteristics of the proposed estimator within an empirical application. The application of the panel data model (ref)requires the estimation of 1143 parameters.\footnote{An extension of the model to (ref), which incorporates linear terms in $\mathbf{X}$ and $\mathbf{Z}$, expands the parameter count to 10,678. While technically feasible, estimating this augmented model is deemed impractical due to the exceedingly small signal-to-noise ratio within the specified application setting.} Table (ref) summarizes the estimation outcomes and resultant model sizes across the various versions of the estimator, denoted by previously introduced abbreviations detailed in Table (ref). Notably, the estimated treatment effects are very close to 0. The p-values, approaching unity for the regularized estimators with imposed null hypothesis, suggest an increased precision in estimating potential outcomes for treated observations. This poses a challenge to the regularity assumption of the stochastic shock, as outlined in Assumption (ref). Given the non-rejection of the hypothesis asserting the absence of a treatment effect, there is no compelling argument to accord greater credibility to the results obtained through the not0 estimation procedure.
The lower segment of Table (ref) elucidates the resultant model sizes derived from the estimation process. Notably, the no_reg estimation, which solely regularizes the rank of the unobserved factor matrix $\mathbf{L}$ while leaving covariates unpenalized, manifests the full model in terms of $\mathbf{H}$ and $\bm{\beta}$ parameters.\footnote{Recall that the fixed effects remain unregularized throughout.} Estimators integrating covariate regularization substantially reduce the model size, yielding an exceedingly sparse model with merely one covariate when employing the 1se-criterion. This observation suggests the presence of very weak signal influencing the severity of Covid-19 infection trajectories within the variables of the dataset. The single parameter persisting the model selection process links the proportion of students and the proportion of infections attributed to the Alpha variant (B.1.1.7) of the Covid-19 virus. These findings hint at a notable association between this mutation and an increased incidence of patients necessitating intensive care unit (ICU) treatment, with young individuals in educational settings possibly serving as significant channel for viral transmission. However, it is essential to underscore that this interpretation hinges solely on the interpretation of model parameters and does not denote an identified causal effect.
We present novel findings concerning $l_1$ covariate regularization within matrix completion methods. A comprehensive simulation study illustrates that this form of regularization significantly diminishes the model size. The determination of penalization parameters through cross-validation, employing the 1se optimality criterion, yields a model size that converges to the true value as the panel size increases.
Moreover, we establish the applicability of the permutation-based inference procedure proposed by chernozhukov2021exact to the extended estimator incorporating covariate regularization. Additionally, we demonstrate its validity under any treatment assignment mechanism. While enforcing the null hypothesis of a treatment effect absence, a prerequisite for applying the inference procedure, introduces a downward bias in effect estimates, simulation results indicate that this bias can be substantially reduced, and even completely mitigated for larger sample sizes, through a simple rule-of-thumb correction.
The proposed estimator, utilizing both 1se-regularization and the estimate correction when inference is necessary, displays superior prediction accuracy compared to the baseline matrix completion estimator introduced by athey2021matrix. The latter has been demonstrated to outperform other common estimators in panel data regression methods. Apart from improved treatment effect estimates, our proposed estimator additionally features a robust model selection property and enables valid finite sample inference.
Future research avenues may include enhancing the two-stage procedure by integrating a more suitable post-model-selection estimation. The application of matrix completion methods in the second step, as evidenced by our simulation study, results in inferior prediction accuracy. Another potential direction for future work involves incorporating an alternative inference procedure that does not impose bias on treatment effect estimates. For instance, a Bayesian inference approach simulating posterior draws via a Markov Chain Monte Carlo sampler, as proposed by tanaka2021bayesian, could be explored.
Additionally, our proposed estimator can be linked to the broader literature that facilitates the application of matrix completion methods to patterns of missing data that are not completely random, as explored by bhattacharya2022matrix, agarwal2021causal, and bai2021matrix.