Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
132,109 characters · 17 sections · 105 citation commands
Double Machine Learning for Static Panel Data with Instrumental Variables: New Method and Applications
Panel data are the backbone of empirical research in modern economics. The structure of panel data enables researchers to account for time-invariant unobserved heterogeneity and exploit within-unit variation for causal inference. However, causal estimation remains challenging when treatment variables are endogenous and identification hinges on the validity of instrumental variables (IVs). In many applications, instruments are only valid conditional on rich sets of covariates angrist1995,abadie2003. Traditional econometric methods, such as two-stage least squares (2SLS), cannot handle the complexities introduced by high-dimensional confounders or nonlinear relationships. For instance, least squares estimation may be infeasible in high-dimensional settings due to rank deficiency induced by many irrelevant regressors, and it cannot capture nonlinear relationships unless these are explicitly parameterized. This creates a tension between the need to control for a rich set of confounders to justify instrument validity and the limitations of conventional methods under sparsity and model-selection constraints.
This paper presents a new method to address these challenges. We propose a Double Machine Learning (DML) procedure for static panels with endogenous treatments and instrumental variables, which we call `Panel IV DML'. Our approach extends the framework of chernozhukov2018 to accommodate serial dependence and unobserved individual heterogeneity, features typical of panel datasets, and that of clarke2023 to allow for treatment endogeneity. First, we derive a novel DML estimator for static panel data models with endogenous treatments based on Neyman orthogonal score functions that account for unobserved individual heterogeneity. The proposed method allows for flexible predictions of the nuisance functions with various machine learning algorithms. It further delivers first-order normal statistical inference for treatment effects, adjusted for regularization and over-fitting bias through the use of Neyman orthogonal score functions and block-k-fold cross-fitting, where each unit's entire time series is assigned to the same fold. Second, to ensure credible inference in the possible presence of weak instruments, we propose first-stage {\em F}-statistics and Anderson-Rubin (AR) test statistics and confidence sets allowing us to detect weak identification issues and, thus, ensure valid inference even when the instrument strength is limited. To our knowledge, this is among the first work to develop weak IV tests for DML procedures, providing empirical researchers with new tools for robust instrumental variables analysis.
This framework is particularly relevant in applied economics, as panel data models estimated with an instrumental variable method (e.g., 2SLS, generalized method of moments) account for 40% of the {\it American Economic Review} (AER) articles using panel data, published in the between 2011-2018.\footnote{This figure is based on a review of 477 empirical articles published in the American Economic Review between 2011 and 2018. In particular, $68\%(=326/477)$ of these inspected articles employ panel data methods. Among the set of these panel data articles, $40\%(=131/326)$ implement an instrumental variables estimation strategy. Detailed information on data collection is provided in Appendix (ref).} We hence illustrate the empirical relevance of our method by revisiting three highly influential studies in migration and political economy that rely on shift-share instruments. Such instruments typically require conditioning on rich sets of covariates to be plausibly exogenous, or as good as randomly assigned borusyak2025practical, hence these applications provide a natural setting for panel IV DML, which flexibly controls for high-dimensional and nonlinear confounding.
The first empirical study we revisit, tabellini2020, studies the political and economic effects of immigration in early 20th-century U.S. cities using a shift-share IV. Our panel IV DML increases the predictive power of the instrument (relevance) since we allow for flexible adjustment in the instrument equation. The panel IV DML estimates broadly align with conventional 2SLS in the baseline specifications for the political outcome, but find no effect for the economic outcome, unlike 2SLS. When additional controls are included in the specification as a robustness check, the main effects found in the original study for both economic and political outcomes disappear under both 2SLS and panel IV DML, suggesting that the results are sensitive to the inclusion of confounders rather than to the estimation method per se.
In contrast, the re-analyses of moriconi2019,moriconi2022, which examine the impact of immigration on political and economic attitudes of individual voters and parties, reveal weak-instrument concerns and, therefore, second-stage results should be interpreted with caution. Weak-identification diagnostics indicate that, while some reduced-form relationships may be present, the available variation is insufficient to support identification of the treatment effect. In these cases, panel IV DML helps assess instrument validity through flexible adjustment, while potentially reducing effective instrument relevance. In these exercises, panel IV DML estimator leads to more cautious inference, emphasising the value of robust weak-identification diagnostics, especially when instruments are only moderately strong with conventional 2SLS.
Our Monte Carlo simulation results support the findings of the empirical applications. Through simulation exercises calibrated to studied empirical designs, we show that panel IV DML outperforms conventional 2SLS in both bias and rate of convergence when confounding is nonlinear and the instrument is strong. When instruments are weak, our Panel IV DML method provides more reliable inference under weak identification than conventional 2SLS, and AR-based confidence sets more reliably reflect true uncertainty than standard inference.
Our work intersects with several strands of the econometric literature. First, the paper is situated within the broad field of causal inference methods that employ doubly/debiased (orthogonal) estimation to adjust for high-dimensional confounders belloni2014, belloni2014inference, belloni2016,chernozhukov2018,chernozhukov2022,knaus2022double,huber2023,scaillet2025,chernozhukov2025fisher,scaillet2025. The canonical article is chernozhukov2018, who introduced the DML method for cross-sectional models by combining the predictive power of machine learning (ML) algorithms to flexibly predict the functional form of the covariates with statistical methods to obtain valid statistical inference of the structural parameter. From this seminal paper, we extend the DML procedure for partially linear regression (PLR) models with instrumental variables to panel data. With repeated observations for each subject in the sample, the presence of the unobserved individual heterogeneity, usually assumed to be correlated with the confounding variables and treatment (i.e.\ fixed effects), poses serious challenges for machine learning and DML, which were not developed for serially correlated data.\footnote{For instance, naively applying to panel data the DML for partially linear regression models with instrumental variables set out by chernozhukov2018 (i.e.\ using `pooled' estimation) would ignore this heterogeneity and, thus, be inconsistent.} We thus use conventional panel data transformations (such as first-differences) for static panels to remove the linear effect of the unobserved heterogeneity. Then, provided that we can learn new nuisance functions (that is, transformations of the nuisance functions of the original partially linear panel model), we show how DML can be used to obtain a consistent estimator of the structural parameter with desirable asymptotic properties.
Second, within the literature on causal machine learning, our contribution relates to a growing strand that extends the DML framework to both static and dynamic panel data models targeting different causal estimands. In static panel settings, klosin2022 propose an estimator for the average partial effect in models with continuous exogenous treatments, using a first-difference transformation to eliminate unobserved time-invariant heterogeneity. Extensions of DML to dynamic panel data models address additional sources of endogeneity arising from predetermined regressors and lagged outcomes. semenova2023 develop a procedure for dynamic panel models with predetermined variables, binary treatments and fully heterogeneous treatment effects, relying on weak sparsity and linearity assumptions suitable for Lasso estimation. chernozhukov2024 further extend the framework by combining the Arellano-Bond estimation strategy with Lasso to estimate structural parameters in dynamic linear panel models with lagged dependent variables, predetermined covariates, and unobserved individual and time fixed effects. In contrast to the rest of the DML literature, arganaraz2025 allow for multivariate and general functionals of unobserved heterogeneity in high-dimensional panel settings, and by providing a full characterization of Neyman-orthogonal moments in models with nonparametric unobserved heterogeneity, even when nuisance components or the distribution of the heterogeneity are not fully identified.
The work most closely related to ours in this strand of the literature is clarke2023. The authors develop DML methods for partially linear panel regression models while accounting for low-dimensional unobserved individual heterogeneity, without imposing sparsity on the nuisance functions, thereby allowing the use of general machine learning algorithms within the DML framework. We extend their framework by relaxing the exogeneity assumption imposed on the treatment variable in their setting. Allowing for potentially non-random instruments requires explicitly modelling the structural equation of the instrument (uncommon in conventional IV procedures) and, consequently, learning an additional nuisance function for that equation. This IV setting also necessitates the construction of new Neyman-orthogonal score functions tailored to our panel data structure to remove regularization bias and deliver valid asymptotic inference under flexible machine learning estimation of the nuisance functions.
A complementary article to ours is marquez2025, who proposes a DML procedure for panel data to capture nonlinearity in the effects of otherwise valid instrumental variables which would be speciously weak using 2SLS. She uses a control function approach in which ML is used to learn the functional form of the instrument(s) and exogenous regressors in the first-stage. In contrast, we adopt a first-stage specification for the treatment which is linear in the instrument and nonlinear in the exogenous regressors. This setting gives us a direct connection to conventional 2SLS for scenarios where established weak IV procedures can be used. Our approach hence prioritizes improving instrument validity by flexibly controlling for a rich set of confounding variables using ML algorithms, making the instrument as good as randomly assigned, and by developing inference procedures that remain valid under weak identification.
We also interact with the literature on valid inference in IV settings under weak identification. Some studies focus primarily on developing inference tools (including $F$ and $t$-tests, adjusted critical values) for detecting weak instruments in low-dimensional structural models cragg1993,stock2005,kleibergen2006,olea2013,lee2022,windmeijer2025. Others built on Anderson-Rubin anderson1949 test statistics to provide robust inference regardless of instrumental relevance in settings with few instruments moreira2019 or in settings where the number of instruments grows with the sample size anatolyev2011,carrasco2016,crudu2021,mikusheva2022,dovi2024. Our paper contributes to this conversation by proposing AR-based inference for the panel IV DML estimator with few instruments and under heteroskedasticity, aligning with best practices emerged in the econometrics literature davidson2014,andrews2019,keane2023,keane2024.\footnote{The AR test statistic and confidence set are not as commonly used in applied work as the first-stage F-statistic, and is rarely reported in standard 2SLS regression tables. Some examples of empirical studies that report AR test statistics and/or AR confidence sets alongside the conventional weak IV tests are: clark2024 in labour economics; ronconi2012 and abayasekara2025 in health economics; cruz2005 in the economics of education, and khammo2024 in public economics; and fumagalli2021 in household behaviour and family economics. } Building on this established discussion, we leverage the desirable asymptotic properties of the DML estimator and the linearity of instrumental variable in the first-stage equation to derive statistical tests for the panel IV DML estimator under weak identification (first-stage {\em F}-statistic, Anderson-Rubin test and confidence sets), not originally proposed for the cross-sectional case. The proposed statistics incorporate flexible information on confounding variables, cross-fitting strategy, and orthogonal residuals that mitigate the bias arising from misspecified confounding structures. The linearity of the instrumental variable in the first-stage relationship between the instrument and the treatment is fundamental for invoking conventional weak IV asymptotic theory and associated inference procedures.\footnote{A methodological challenge for future research is extending weak IV asymptotics to more general scenarios where the first-stage relationship is nonlinear, as in chernozhukov2018 and marquez2025.}
Finally, we contribute to the broader discussion on the value added of machine learning methods to causal analysis and policy evaluation in economics. A growing body of empirical work deryugina2019,langen2023,strittmatter2023,baiardi2024b,baiardi2024a documents promising results from the use of causal machine learning techniques (e.g., causal forest, generic machine learning, and double machine learning) to address empirical research questions by uncovering nonlinearities and heterogeneous effects that would otherwise remain undetected with conventional estimation methods. These applications use cross-sectional data or panel data reshaped into a cross-sectional form. Our work complements this literature by extending these insights to panel data settings, where unobserved heterogeneity, nonlinearities, and high-dimensional confounding pose additional challenges for causal inference. Although panel data with instrumental variables is a common estimation strategy in empirical economics, the application of causal machine learning methods to this setting remains largely underdeveloped — a gap this paper directly addresses.
The remainder of the paper is structured as follows. Section (ref) introduces the notation and presents the econometric framework for panel IV DML. Section (ref) describes the estimation and inference with panel IV DML by defining the Neyman orthogonal score functions, the estimators, and the (robust) tests under weak identification. Section (ref) applies the method to three empirical applications with shift-share instruments in the context of the economics of migration. Section (ref) illustrates the finite sample properties of the panel IV DML estimator with the Monte Carlo simulation exercises. Section (ref) concludes.
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{lemma}{0} \setcounter{proposition}{0}
\setcounter{equation}{0}
Suppose that individual $i$ is randomly drawn from a population and that measures are collected for everyone in the sample over multiple time periods $t$ (or waves in survey studies). Let $\{(Y_{it},D_{it},\vector Z_{it},\vector X_{it}) : \ t=1,\ldots,T\}_{i=1}^{N}$ be independent and identically distributed (iid) random vectors for each of the $N$ individuals across all $T$ time points, where $Y_{it} \in \mathcal{Y}$ is the outcome variable, $D_{it}\in{\cal D}$ a continuous or binary endogenous treatment variable (or intervention), $\vector Z_{it}= (Z_{it,1},\dots,Z_{it,r})'\in{\cal Z}$ is a $r\times1$ vector of valid instruments (continuous or binary), and $\vector X_{it} = (X_{it,1},\dots,X_{it,p})'\in \mathcal{X}$ a $p\times 1$ vector of control (pre-determined) variables, usually including a constant term, able to capture time-varying confounding induced by non-random treatment selection.\footnote{Throughout, we use ${\bf v}^\prime$ to indicate the matrix transpose of arbitrary vector {\bf v} and take vectors to be column vectors unless the opposite is explicitly stated.} We denote the realizations of these random variables by $\{(y_{it}, d_{it}, \vector z_{it},\vector x_{it})\}$, respectively. For continuous $D_{it} \in \mathcal{D}\subset\mathbb{R}$, if $d_{it}\geq 0$ then it is assumed that a dose-response relationship is maintained with $D_{it}=0$ indicating null treatment; otherwise, $D_{it}$ is taken to be centered around its mean $\mu_D$ such that $D_{it} \equiv D_{it}-\mu_D$. For binary $D_{it}\in\{0,1\}$, $D_{it}=0$ is taken to indicate the absence and $D_{it}=1$ the presence of treatment. In the interval between times $t-1$ and $t$, it is assumed that the realizations of time-varying predictor $\vector X_{it}$ and instrument $\vector Z_{it}$ precede that of $(Y_{it},D_{it})$.
Consider the following partially linear panel regression (PLPR) model with instrumental variables, based on extending the cross-sectional model by robinson1988 (see also chernozhukov2018):
where $Y_{it}$ is the outcome of interest, $D_{it}$ an endogenous treatment variable (or policy intervention), $ \vector Z_{it}$ is a $r\times1$ vector of exogenous excluded instruments which is not fixed, $\vector X_{it}$ is a $p\times1$ vector of exogenous included regressors; $\bm{\eta}_0 = (l_0, r_0, \vector m_0)$ is a vector of possibly nonlinear nuisance functions of the covariates to be estimated using ML algorithms; $\theta_0$ is the structural parameter of interest to estimate consistently to conduct causal inference, and $\bm{\pi}_0$ a $r\times1$ vector of parameters for the excluded exogenous variables (or instruments). The $\alpha_i$, $\zeta_i$ and $\bm{\gamma_i}$ are unobserved individual-level, or fixed, effects; all three are functions of the omitted (time-invariant) individual-level omitted variables $\xi_i$.
Endogenous treatment selection induces a correlation between errors $U_{it}$ and $R_{it}$ such that $\mathbb{E}(U_{it}\mid D_{it},\vector X_i,\xi_i)\neq 0$ but $\mathbb{E}(U_{it}\mid \vector Z_{it},\vector X_{it},\xi_i)=0$. Unlike conventional 2SLS settings, the instruments $\vector Z_{it}$ are not treated as fixed but as wave $t$ realizations conditional on $\mathbf{X}_{it}$ but preceding $(Y_{it},D_{it})$. Based on ((ref)), the role of instrumental variable residuals ${\vector V}_{it}$ is crucial to the construction of our DML procedure. These are required to ensure a) the score function is Neyman orthogonal and, hence, that DML controls the size of the errors introduced by using ML-based estimation of the nuisance functions, and b) the nuisance functions for the first-difference estimator can be consistently learnt (see Remark (ref) below). Moreover, using $\vector V_{it}$ as the instrumental variable rather than ${\vector Z}_{it}$ enables (ref) and (ref) to be parametrized in terms of the same nuisance parameter $r_0(\mathbf{X}_{it})$, which simplifies the resulting DML algorithm.
The assumptions which must be satisfied by the data generating process in order that (ref)-(ref) holds, and the special role played by (ref), are now set out:
We employ the first-difference (FD) transformation to remove the unobserved individual heterogeneity (or fixed effects) in the PLPR Model (ref)-(ref). The FD transformation is preferable over other traditional panel data techniques, such as the within-group (or fixed effects) transformation and correlated random effects mundlak1978,chamberlain1984, because it imposes the fewest constraints on the data generating process and permits feasible estimation of the nuisance functions.
Let $\widetilde{Y}_{it} = Y_{it}-Y_{it-1}$ be the first-difference of the random variable $Y_{it}$, and $\widetilde{l}_0(\vector X_{it}) = l_0(\vector X_{it})-l_0(\vector X_{it-1})$ be the first-difference of the nuisance function $l_0$ for all $i=1,\dots,N$ and $t=2,\dots,T$. The first differences of the other random variables and nuisance functions are similarly defined using \ $\widetilde{}$ \ notation. Then, the first-differenced PLPR model with IV model based on model equations (ref)-(ref) is
where $\alpha_i,\zeta_i$ and $\bm{\gamma_i}$ are now absent from the system, and
where ${\bm\delta}_0={\bm\pi}_0\theta_0$ and $\widetilde{U}^*_{it}=\widetilde{U}_{it}+\widetilde{R}_{it}\theta_0$. Following andrews2019, we refer to (ref) as the structural model, (ref) as the first-stage model, and (ref) as the reduced-form model. The estimator of reduced-form model parameter ${\bm\delta}_0$ plays a crucial role in deriving asymptotic results for weak instruments.
Under the model above, naively using two-stage least squares (2SLS) to estimate the effect of endogenous treatment $\widetilde{D}_{it}$ on $\widetilde{Y}_{it}$ using $\widetilde{{\bf Z}}_{it}$ as an instrumental variable leads to bias. For example, in the simple $T=2$ case and letting $\widetilde{{\bf d}}=(\widetilde{D}_{12},\ldots,\widetilde{D}_{n2})^\prime$, $\widetilde{Z}=(\widetilde{{\bf Z}}_{12},\ldots,\widetilde{{\bf Z}}_{n2})^\prime$ and $\widetilde{X}=(\widetilde{{\bf X}}_{12},\ldots,\widetilde{{\bf X}}_{n2})^\prime$, it can be shown that
where $\widetilde{{\bf e}}$ is the residual of the population-level linear projection of $\widetilde{{\bf y}}=(\widetilde{Y}_{12},\ldots,\widetilde{Y}_{n2})^\prime$ on $\widetilde{X}$ and $\widetilde{Z}$, and $\widetilde{{\bf g}_2}=(\widetilde{g}_0({\bf X}_{12})\ldots,\widetilde{g}_0({\bf X}_{n2}))^\prime$; $M_X\widetilde{\bf d}$ is the residual of the linear projection of $\widetilde{\bf d}$ onto the space spanned by the columns of $\widetilde{X}$, or ${\rm span}(\widetilde{X})$, $M_X(\widetilde{{\bf e}}+\widetilde{{\bf g}_2})$ is that for $\widetilde{{\bf e}}+\widetilde{{\bf g}_2}$ onto ${\rm span}(\widetilde{X})$, and $M_{\hat{Z}}{\bf a}$ is the residual of the linear projection of any ${\bf a}$ onto ${\rm span}(M_X\widetilde{Z})$. The usual bias term for 2SLS is ${\bf b}\widetilde{{\bf e}}$, induced by the sample correlation between $\widetilde{{\bf Z}}_{it}$ and the structural error being non-zero, albeit negligible for strong $\widetilde{{\bf Z}}_{i2}$; but ${\bf b}\widetilde{{\bf g}}_2$, the (scaled) linear projection of $\widetilde{{\bf g}_2}$ onto ${\rm span}(\widetilde{X})$, is potentially non-zero whether $\widetilde{{\bf Z}}_{i2}$ is strong or weak.
The main challenge for implementing DML is to learn the transformed ex-ante unknown nuisance functions $\widetilde{l}_0(\vector X_{it})$, $\widetilde{r}_0(\vector X_{it})$, and $\widetilde{\vector m}_0(\vector X_{it})$. For this, we adopt the FD (exact) approach proposed by clarke2023 to learn these functions from the observed transformed data on $\widetilde{Y}_{it}$, $\widetilde{D}_{it}$ and $\widetilde{{\bf Z}}_{it}$, respectively, given ${\bf X}_{it-1}$ and $\mathbf{X}_{it}$. In the next sections, we handle the first source of bias and derive a consistent estimator of the target (or causal) parameter of interest when the treatment variable is endogenous.
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{lemma}{0} \setcounter{proposition}{0}
This section presents the estimation and inference of the panel IV DML estimator and the weak identification diagnostics. A detailed discussion of the algorithm is in the Online Appendix (ref).
In this section, we set out DML estimation of the panel IV model parameters first by deriving a Neyman-orthogonal score function for estimating structural model parameters ${\vector b}_0=(\theta_0,{\vector{\pi}}_0^\prime)^\prime$ needed for DML and second by establishing its asymptotic properties.\footnote{Equivalent expressions for reduced-form model parameters ${\vector b}_0^{\rm rf}=(\mathbf{\delta}_0^\prime,{\vector{\pi}}_0^\prime)^\prime$ to facilitate the use of weak instrument techniques are derived, but omitted from the main text.}
Denote the $N(T-1)\times 1$ random vectors of first differences $\widetilde{\vector Y} = (\widetilde{\vector Y}_1,\dots,\widetilde{\vector Y}_{N})'$, with $\widetilde{\vector Y}_i = (\widetilde{ Y}_{i2},\dots,\widetilde{Y}_{iT})'$ for individual $i$, and the first differences of the nuisance function $\widetilde{\vector l}_0 = (\widetilde{\vector l}_{01},\dots,\widetilde{\vector l}_{0N})'$, with $ \widetilde{\vector l}_{0i} = \big(\widetilde{l}_0({\bf X}_{i2}),\ldots, \widetilde{l}_0({\bf X}_{iT})\big)'$ for $i=1,\dots,N$. The other $N(T-1)$-vectors of first differences $\widetilde{\vector D},\widetilde{\vector U},\widetilde{\vector R},\widetilde{\vector U}^*$ and $\widetilde{\vector r}_{0} $ are similarly defined. The $N(T-1)\times r$ matrices relating to instrumental variable model (ref), where $r$ is the number of valid instrumental variables, are $\widetilde{\matrix Z}= (\widetilde{\matrix Z}_1,\dots,\widetilde{\matrix Z}_{N})'$ with $\widetilde{\matrix Z}_i = (\widetilde{\vector Z}_{i2},\dots,\widetilde{\vector Z}_{iT})'$, and $\widetilde{\matrix V}= \widetilde{\matrix Z}-\widetilde{\matrix M}_{0}$, where $\widetilde{\matrix M}_{0} = (\widetilde{\matrix M}_{01}, \dots,\widetilde{\matrix M}_{0N})'$ with $\widetilde{\matrix M}_{0i} = \big(\widetilde{{\vector m}_0}(\vector X_{i2}),\ldots,\widetilde{{\vector m}_0}(\vector X_{iT})\big)'$.\footnote{$\widetilde{\matrix M}_{0}$ should not be confused with projection matrix residual $M_X$ defined in the previous section.} Finally, let $\matrix X$ be a $NT\times p$ matrix of untransformed confounding variables. The complete set of random vectors and matrices with observed realizations $\matrix W=\{\widetilde{\vector Y},\widetilde{\vector D},\widetilde{\matrix Z},{\matrix X} \}$.
A proof is given in printed Appendix (ref). A formal definition of $T_N$ is later provided in the proof of Proposition (ref).
To present our main result for the estimation and inference of $\mathbf{b}_0$ using the panel IV DML estimator, let $\{\delta_N\}_{N=1}^{\infty}$ and $\{\Delta_N\}_{N=1}^{\infty}$ be two sequences of positive constants tending to zero as $N$ increases such that $\delta_N\ge N^{-1/2}$. Let $c,C>0$ be arbitrary fixed constants denoting lower and upper bounds, respectively. Further fixed constants:\ recall that $r\ge1$ is the number of instrumental variables, and let $K\ge2$ be the number of cross-validation folds chosen by the analyst.
As discussed above, the partially linear first-difference model (ref)-(ref) (and (ref)-(ref)) holds under Assumptions (ref)-(ref), where Assumptions (ref)(b) and (ref) permit $\widehat{\vector{l}}_{0i}$, $\widehat{\vector{r}}_{0i}$ and $\widehat{\mathbf{M}}_{0i}$ to be consistently learnt. Furthermore, {Propositions (ref) and (ref) present the Neyman orthogonal score and Regularity Conditions under which DML yields consistent and asymptotically normal estimators of $\theta_0$ and $\bm{\pi}_0$ (or $\bm{\delta}_0$ and $\bm{\pi}_0$). If we have $\sqrt{N}$-consistent estimators of $\Omega_{\theta\theta}$ and $\Omega_{\pi\pi}$ then, without loss of generality, we can set both equal to identity matrices in the following explanation. Then, following chernozhukov2018, a second-order Taylor series expansion of score (ref) around $\bm{\psi}=\bm{\psi}_0$, where $\bm{\psi}_0 = (\theta_0, \bm{\pi}_0, \widetilde{l}_0, \widetilde{r}_0, \widetilde{m}_0)$, gives
and
where DML allows us to lump into \(o_p(1)\) all cross-products involving at least one residual from \(\widetilde{\vector D} - \widetilde{\vector r}_{0}\), \(\widetilde{\vector R}_{0}\), \(\widetilde{\matrix V}_{0}\), and \(\widetilde{\vector U}_{0}\), and at least one of \(\widehat{\bm{\pi}}_{\mathrm{DML}} - \bm{\pi}_0\), \(\widehat{\vector l} - \widetilde{\vector l}_{0}\), \(\widehat{\vector r} - \widetilde{\vector r}_{0}\), and \(\widehat{\matrix M} - \widetilde{\matrix M}_{0}\) in the Taylor series (\(\widehat{\vector l}\), \(\widehat{\vector r}\), and \(\widehat{\matrix M}\) are, respectively, predictions of \(\widetilde{\vector l}\), \(\widetilde{\vector r}\), and \(\widetilde{\matrix M}\) from DML stage one); and the requirement that the base learners satisfy \(o_p(N^{-1/4})\) ensures the same for cross-products involving pairs from \(\widehat{\bm{\pi}}_{\mathrm{DML}} - \bm{\pi}_0\), \(\widehat{\vector l} - \widetilde{\vector l}_{0}\), \(\widehat{\vector r} - \widetilde{\vector r}_{0}\), and \(\widehat{\matrix M} - \widetilde{\matrix M}_{0}\). This allows weak-instrument asymptotics to be applied in the usual manner for both the first-stage \(F\) and Anderson-Rubin statistic.
Assessing instrument relevance is routine in applied work, as weak instruments bias 2SLS estimator and invalidate its normal asymptotic properties bound1995. In practice, applied researchers typically rely on the first-stage {\em F}-statistic, which is routinely reported in empirical tables. We therefore begin by showing how conventional first-stage {\em F}-statistic can be used within our DML framework, providing a diagnostic that is familiar and immediately interpretable for applied researchers. We then turn to the Anderson-Rubin test statistic and confidence set which, although more robust to weak identification, remains underutilized in empirical practice.
The F-statistic from the linear regression of the endogenous variable on the instrument(s) (first-stage regression) is routinely reported in empirical IV applications and often used as a diagnostic for instrument relevance. In practice, instruments are typically considered sufficiently strong when this statistic exceeds conventional critical values.\footnote{Conventional stock2005's critical values can be used for inferences in just-identified cases. In practice, a widely used rule-of-thumb is to consider an instrument strong if the first-stage {\em F}-statistic exceeds the value of 10 stock2019. A more rigorous approach is based on thresholds for first-stage F derived by stock2005, who consider the maximal tolerable t-test size distortion, or alternatively how often a 5% t-test rejects a true hypothesis. Recently, lee2022 claim these thresholds should be reconsidered. They show that only when the {\em F}-statistic exceeds 104.7 the t-tests are ensured not to reject the null hypothesis at a rate higher than the desired one (e.g., 5% rate).} Given the importance of assessing weak identification for valid inference, we derive expressions for a first-stage F-statistic adapted to our panel IV DML framework.\footnote{Our tests for weak IV can be easily adapted to the partially linear regression model with linear instrumental variables within cross-sectional DML framework by chernozhukov2018. These tests have not been yet developed in that setting at the time of writing.}
It is well established that statistical inference based on conventional 2SLS estimator is invalid when instruments are weak, especially when the OLS bias is large keane2024. Therefore, conventional F- and t-based inference becomes unreliable, even when instruments appear strong, whereas tests that are robust to weak identification and do not rely on assumptions about instrument relevance, such as the Anderson-Rubin (AR) test anderson1949, remain valid andrews2019,moreira2019,keane2023.
The AR test evaluates the null hypothesis on the structural parameter by testing the corresponding reduced-form restriction. Specifically, it assesses whether the instrument(s) have a statistically significant reduced-form effect on the outcome under the assumption of instrument exogeneity (i.e., that instruments are as good as randomly assigned). This approach provides valid inference on the structural parameter without requiring strong identification (i.e., relevance).
As in the conventional 2SLS estimation method, the solutions of the quadratic form associated with the AR test statistic are: (a) a bounded interval $[x,y]$, where $\{x,y\}$ with two real roots $x\le y$; (b) a disjoint set $(-\infty,x] \cup [y,+\infty)$ with two real roots $x< y$; (c) the real line $(-\infty,+\infty)$; and (d) the empty set, which is only possible in overidentified settings. AR CS (c) and (d) have no real roots.
The type of AR CS informs about the relevance of the instrument. When instruments are strong, the AR CS is bounded, as in case (a), and the model is well-identified. When instruments are sufficiently weak, the AR CS are usually unbounded (cases b-c) and uninformative, since a bounded set generally does not exist locally when the parameter is not identified davidson2014. Empty sets (case d) occur when the equality $\bm{\delta} = \bm{\pi}\theta_0$ fails due to either (a) treatment effect heterogeneity when the equality holds for $\theta\ne\theta_0$, or (b) invalidity of the instruments -- i.e., the failure of the overidentifying restrictions -- when there is no value of $\theta$ that satisfies the inequality andrews2019.
\setcounter{equation}{0}
In this section, we showcase the applicability of our panel IV DML method.\footnote{The panel IV DML estimation is conducted in R using the latest version of the xtivdml package accessible in its latest version at \url{https://github.com/POLSEAN/xtivdml} at the time of writing, which is built on R packages DoubleML DoubleML and xtdml xtdml. The conventional 2SLS estimation is implemented in Stata 19.5 using the community-contributed commands \texttt{xtivreg2} schaffer2005 and \texttt{twostepweakiv} sun2018} For a comprehensive illustration, we revisit three empirical studies that examine whether and how immigration affects the political views of native voters towards immigration in the United States (tabellini2020) and Europe (moriconi2019, moriconi2022).
As typical in the migration literature, an instrumental variable strategy (2SLS) is employed to account for the potentially endogenous decision of immigrants to settle across areas: on the one hand, they may be more likely to settle in areas where the economy or attitudes towards immigration are more favourable; on the other hand, they may decide to settle in places where housing is more affordable and which are economically declining. These three studies adopt modified versions of the widely used shift-share instrument card2001, which exploits the tendency of immigrants to settle in areas with larger pre-existing communities from the same country of origin or ethnic group.\footnote{The focus on articles that exclusively use a shift-share IV is motivated by the fact that the instrument is not randomly assigned, by construction, and a rich set of covariates is typically required to make the instrument approximately exogenous, namely as good as randomly assigned. This setting is well suited to our panel IV DML approach, which flexibly adjusts for high-dimensional confounding.} The version of the shift-share instrument used in these articles is constructed as the weighted average of new national inflows from each immigrant’s country of origin (the `shifts'), using fixed pre-existing local settlement patterns (the `shares') as weights. In formulae,
where $Pop$ is the total population in city/region $r$ in the pre-settlement period $T_0$, $Sh_{crT_0}= M_{crT_0} / \sum_r M_{crT_0}$ is a fixed share of pre-existing immigrants, $M_{crt}$ is the national stock (inflows) of immigrants, $\mathcal{C}$ is the set of countries of origin, and $\mathcal{R}$ is the set of receiving cities/regions.
Regarding the validity of the assumptions underlying the shift-share instrument, historical settlement patterns (the `shifts') are generally strong predictors of current immigrant inflows and therefore are often found to satisfy the relevance condition in conventional 2SLS settings. The exogeneity of these instruments and the validity of the exclusion restriction have recently been at the center of a number of papers, which highlight that one of the two components (either the `shifts' or the `shares') must be exogenous adao2019shift,borusyak2022quasi,goldsmith2020bartik. In our empirical applications, we abstract from this debate and take as given the arguments provided in the original papers in support of these assumptions. Instead, we focus on how our panel IV DML method can address the potential presence of a large number of possibly nonlinear confounders. Depending on the setting, the exogeneity of the instrument may only be valid once conditioning on certain controls borusyak2025practical. These controls may enter the true nuisance functions nonlinearly, and ignoring such nonlinearities can affect the estimation of treatment effects due to model misspecification.
In our re-analysis, we compare the results from our panel IV DML estimator to those obtained from conventional 2SLS with FD, which is the closest panel data estimation method to our estimation approach.\footnote{As previously discussed in Section (ref), the FD (exact) approach allows us to approximate the transformed unknown nuisance functions in nonlinear settings without imposing many constraints on the fixed effects, unlike the within-group transformation (or fixed effects) and the correlated random effects device. The other two approaches are possible but may lead to inconsistent estimates when the true functional form is highly nonlinear clarke2023.} As the original estimation approaches are either 2SLS with FE tabellini2020 or pooled ordinary least squares with region and year fixed effects moriconi2019,moriconi2022 due to the structure of the data, we additionally contrast the estimates obtained from a 2SLS regression with FE with 2SLS with FD in each empirical application. This preliminary step allows us to verify that any differences with panel IV DML method arise only from the use of a different estimation methodology and not from different specification choices (see Online Appendices (ref) and (ref) for an exhaustive discussion of conventional panel data estimation results).
In our panel IV DML regressions, the functional form of the confounding variables is learned flexibly using machine learning algorithms from each major class (i.e., Lasso for L-norm regularized linear models, gradient boosting with 1000 trees for tree-based methods, and a single-hidden-layer neural network for deep learning) allowing the model to capture a broad range of nonlinearities in the data. The hyperparameters of the base learners are tuned with grid search bergstra2012 (see Table (ref) in the Online Appendix (ref)). The set of covariates employed as inputs in panel IV DML estimation with neural network and gradient boosting includes raw covariates only, no interaction terms or polynomials are included because these base learners are designed to automatically capture nonlinearities in the data. Conversely, panel IV DML with Lasso uses an extended dictionary of nonlinear terms of the raw covariates (i.e., polynomials up to order three and interaction terms of all covariates) to satisfy weak sparsity assumption. The inputs used in panel IV DML estimation also include one-period lags of all control variables, following the FD (exact) approach discussed in Section (ref). In addition, the sample in panel IV DML regressions is divided in two folds, where the number of folds is chosen to account for the small cross-sectional dimension ($N$) in all empirical applications.\footnote{Dividing the sample into a higher number of folds reduces the size of the estimation sample considerably, which may differ in terms of observable characteristics from the prediction sample. In such cases, the learner would need to extrapolate the information, but is not desirable for flexible tree-based learners and neural networks. We allow for cross-fitting as desirable to restore efficiency.}
The first empirical application revisits the study by tabellini2020, which examines the political and economic effects of restrictive immigration policies induced by World War I and the Immigration Acts of the 1920s in the United States (U.S.) that reduced the quota of European immigrants. The analysis exploits the exogenous variation in immigration from European countries to 180 U.S. cities over three census years 1910, 1920 and 1930. Data are collected from various sources: U.S. Census of Population, Voteview, Census of Manufactures.
The author employs an instrumental variable strategy to address the potentially endogenous settlement patterns of European immigrants across U.S. cities. The instrument used is a leave-out version of the shift-share measure defined in equation (ref), which interacts historical settlement patterns of different ethnic groups in each city in 1900 with contemporary immigration flows of the same group, excluding immigrants who ultimately settle in the same city. The identifying assumption for the validity of the instrument is verified by the author with a pre-trend test, which confirms that, prior to 1900 (before any networks formation), European immigrants did not disproportionately settle in cities experiencing economic growth or political change. The original analysis is conducted using conventional 2SLS estimation with fixed effects (FE). The main findings of the article suggest that immigration induced hostile political reactions, such as the election of more conservative legislators, stronger support for anti-immigration legislation, and lower redistribution.
We re-analyze the baseline specifications from Table 3 (Column 4, Panel B) and Table 5 (Column 1, Panel B) in tabellini2020 using conventional 2SLS with FD and our panel IV DML method with the FD approach. We focus on two outcomes: for the political effects, we look at the Poole-Rosenthal DW Nominate Score, which ranks congressmen on an ideological scale from liberal to conservative using voting behaviour on previous roll-calls; for the economic effects, we consider the (log) occupational score, which is a proxy for natives' income and does not capture within occupation changes in earnings.\footnote{As explained by the author, wage data are not available since until 1940. Occupational scores are commonly used in the literature to proxy lifetime earnings, which are calculated by assigning the median income of an individual job category in 1950 to them.} The baseline specifications include only interaction terms of region and year fixed effects.
We then depart from the original analysis by augmenting the baseline specifications with additional control variables: predicted industrialization, immigrant and city population, value added manufacturing, skill ratios, fraction of blacks, value of products, employment share in manufacturing. In the original analysis, each of these control variables is added one at a time to the baseline specification to evaluate the robustness of the results.\footnote{ For instance see Tables D2-D3 in the Online Appendix of tabellini2020. We thank Marco Tabellini and Francesco Maria Toti Ognibene for providing us with the material to replicate Tables D2-D3 in the Online Appendix of tabellini2020.} In practice, the identification of the effect with inclusion of many irrelevant controls in the estimating equation may not be possible with conventional estimation methods, such as least squares, due to matrix singularity and multicollinearity issues. In contrast, DML methods (for both cross-sectional and panel data) can handle high-dimensional covariate spaces.
Before discussing panel IV DML results, we briefly compare conventional 2SLS estimates obtained using fixed effects (FE) and first-differences (FD) (see the Online Appendix (ref) for a detailed discussion). This preliminary step ensures that any differences observed against our panel IV DML estimator reflect methodological rather than specification choices. In the baseline specifications, the shift-share instruments appear strong with both FE and FD estimators for the political outcome (DW Nominate Score) and the economic outcome (Log Occupational Score). The second-stage coefficients are positive and significant at least at 5% level. When additional controls are included, an extension not implemented in the original analysis, the shift-share instrument becomes weak in the specification of the political outcome while remaining strong in the specification of the economic outcome with both panel data estimators. Regardless of the strength of the instrument, both estimators produce statistically insignificant effects, contrasting the results found in the baseline regression. In general, the two estimators produce very similar results in terms of sign, magnitude and statistical significance, ensuring that subsequent differences between 2SLS and panel IV DML can be explained by differences in the methodology.
We now proceed to evaluate how our panel IV DML estimates compare to the corresponding 2SLS with FD. Table (ref) reports the results of the baseline specifications in Panel A, and the augmented specifications with all control variables in Panel B for the two outcome variables of interest. Columns 1 and 5 in Table (ref) report the 2SLS estimates with FD, and Columns 2-4 and 6-8 the panel IV DML estimates with FD using Lasso, Neural Network (NNet), and Gradient Boosting (Boosting).
For each specification, we need to verify whether the shift-share instruments in panel IV DML regressions still predict immigrant shares across cities (i.e., instrument relevance) within the more flexible DML framework because our estimator controls for covariates in a more flexible way and, therefore, the effective strength of the instrument may differ. We start by discussing the baseline regression estimates (Panel A of Table (ref)) focusing on one outcome at a time. For the political outcome (Columns 2-4), the first-stage F-statistics from panel IV DML regressions are much larger than those from 2SLS, which slightly exceeds stock2005's cut-off of 16.30, but never exceed lee2022's threshold of 104.70. In this case, Panel IV DML strengthens the relevance of the shift-share instrument. The AR test statistics for the relevance of the reduced-form stage, when there is no second-stage effect regardless of the strength of the instrument, strongly reject the null hypothesis with panel IV DML, suggesting some effect of immigration on the political outcome. This is confirmed by the AR confidence sets (CS) at 95% level which are bounded, include the estimated second-stage coefficient and do not include zero, with the exception of neural network due to small sample size (further reduced by sample splitting). Given the strength of the instrument (based on the stock2005's threshold), we can comment on the statistical significance of the second-stage coefficient and make inference about the effect of immigration. All second-stage coefficients are positive and significant at 1% level; specifically, panel IV DML coefficients are larger in magnitudes than 2SLS, suggesting a stronger effect of the fraction of immigrants on the political outcome than originally found. Overall, panel IV DML findings indicate that a greater inflow of immigrants leads to the election of more conservative congressmen; the direction of the effect is aligned with the original 2SLS findings, which seem to be underestimated.
Moving to the economic outcome (Columns 5-8 of Panel A), the first-stage F-statistics from panel IV DML are smaller than those from 2SLS, but always well above the conventional critical value of 16.30 by stock2005 and below 104.70 by lee2022's. Therefore, the instrument can still be considered strong with panel IV DML. Unlike 2SLS results, the AR test statistic and AR CS from panel IV DML regressions highlight the absence of any effect for the economic outcome. That is, in Columns 6-8, the AR test never rejects the null hypothesis of no reduced-form effect when the treatment effect is assumed to be absent, and the AR CS always include zero as a possible value of the treatment effect. Given the strength of the IV and the results of the AR diagnostics, we can interpret the second-stage coefficients as not statistically different from zero with panel IV DML. By contrast, 2SLS finds a positive and statistically significant (at 5% level) second-stage coefficient. Overall, panel IV DML findings contradict conventional 2SLS results, suggesting that there is no evidence that employment gains for natives, induced by immigration, were accompanied by occupational or skill upgrading, as found in the original article.
We now focus on Panel B of Table (ref), when all controls are included simultaneously (not originally implemented). In both specifications (for political and economic outcomes), the first-stage F-statistics generally become smaller than those from Panel A. In Columns 2-4 of Panel B, the shift-share instrument in the panel IV DML regressions produces stronger first-stages ($F>16.30$) relative to the corresponding 2SLS regression, where the instrument is clearly weak ($F=9.26$). In Columns 5-8 of Panel B, the first-stage F-statistics obtained from both 2SLS and panel IV DML estimators are largely above stock2005's threshold, and in the case of the panel IV DML regression with Lasso even above lee2022's threshold of 104.70. Overall, results for Columns 2-8 provide supporting evidence of a strong instrument. This pattern suggests that allowing for flexible and potentially nonlinear control adjustment in the instrument equation strengthens the predictive power of the instrument with panel IV DML. The AR test statistics are insignificant with both estimators in both political and economic specifications, suggesting we cannot reject no second-stage effects. The absence of a causal effect for either outcome variable is supported by the AR CS, which include zero as a plausible value of the treatment effect, as well as by the statistical insignificance of the second-stage coefficients with 2SLS and panel IV DML estimators.
This first empirical application highlights that our method can enhance the plausibility of instrument's validity assumption, while controlling for more and potentially nonlinear confounding variables. We show that our panel IV DML estimator generally confirms the conventional 2SLS results, with the only exception of the economic outcome in the baseline specifications (Panel A of Table (ref)). Our robustness analysis with additional control variables (Panel B of Table (ref)) illustrates how the panel IV DML method in shift-share designs, where conventional estimators may fail. We find that the main effect of the original article (Columns 1-5 in Panel A of Table (ref)) for both outcomes disappears in both 2SLS and panel IV DML regressions.
In this section, we revisit two related studies by moriconi2019, moriconi2022 on the impact of high-skilled (HS) and low-skilled (LS) immigration on support for redistribution policies and perceived attitudes in European countries. Similarly to tabellini2020, endogeneity concerns arise as immigrants may decide to settle in areas with more favourable policies and attitudes towards immigration. To address these issues, an IV approach is employed with a modified version of the shift-share instrumental variable (ref) by skill-specific group (i.e., HS and LS). The instruments are constructed interacting the aggregate immigrant flows by skill group and country of origin (the `shift') with the initial distribution of immigrants by nationality across regions (the `share'), as in mayda2022. The authors show that the skill-specific shift-share instruments are uncorrelated with economic and demographic regional trends in the period before the analysis, so that the exclusion restriction is satisfied.
The data for the analyses in both articles are obtained from the European Social Survey (ESS), the European Labor Force Survey (EULFS), the Manifesto Project Database and Eurostat. The main analyses are implemented on a dataset of individual voters from twelve European countries observed over the election years between 2007 and 2016. The final dataset is not a conventional panel dataset, where the same subject is observed over multiple time periods, as contains information of randomly sampled native voters across regions of the same twelve 12 EU countries in each election year. In contrast, our panel IV DML method requires (balanced or unbalanced) panel data where the same subject is followed over at least two consecutive periods. To satisfy this requirement, we construct an aggregated balanced panel dataset from the individual data by averaging variables at the NUTS2 region-year level.\footnote{In the original analyses both the endogenous variables and the instruments are at the NUTS2 region-year level. Therefore, the level of aggregation of these variables does not change in our aggregated data, but it affects the outcome variables and the individual-level controls only.}
Both studies estimate the main specifications through pooled ordinary least-squares (POLS) regressions with regional and election-year fixed effects, using individual voters' data. Therefore, we first re-estimate the key specifications using both 2SLS with fixed effects (FE) and 2SLS with first-differences (FD) estimation approaches using our aggregated (unbalanced) panel dataset to establish a consistent comparison with our panel IV DML method.
The article investigates the effect of skilled and unskilled immigration on individual preferences for the expansion of the welfare state and public education in twelve EU countries. While the overall effect of migration varies, their main findings highlight a pro-redistribution effect with HS immigration from natives’ voters.
In this second empirical application, we re-examine the main results of moriconi2019 relative to Columns 2, 4 and 6 (Panels A and B) of their Table 4 (on individual voters' data).\footnote{In the Online Appendix (ref), we revisit the specifications in Table 5 of moriconi2019, estimated using parties' data. We thank the authors for providing us with the entire replication package to reproduce the main analysis.} The dependent variables are: Net Welfare State and Net Public Education. The control variables included in the regressions are the same as in the original article and include: the share of women, average age, share of tertiary/post-tertiary education, average GDP per capita (in log), share of tertiary sector (in log), average unemployment rate, and election year dummies. The regression analysis is separately estimated by skill-specific group of immigrants.
Before discussing the panel IV DML results, we briefly compare the results from conventional 2SLS estimators with FE and FD (see Online Appendix (ref) for more details). In general, both FE and FD estimators confirm that the instrument is strong in the HS immigration sample ($F\ll16.30$), but borderline weak in the LS immigration sample (as $10<F<16.30$) for both specifications. AR diagnostics, not computed in the original study, find the presence of a second-stage effect of HS immigrants and, unlike the original article, of LS immigrants on both outcomes. The AR CS from both panel estimators agree on the direction of the effect of HS immigrants on welfare expansion, and of LS immigrants on education expansion. Therefore, any difference observed between 2SLS with FD and panel IV DML later can be explained by differences in the adopted methodology, i.e. 2SLS versus panel IV DML.
Table (ref) reports our estimates obtained from 2SLS regressions with FD (Columns 1 and 5) and panel IV DML regressions using the FD approach (Columns 2-4, and 6-8), based on the aggregated panel dataset. The first-stage regression results in Panels A and B are identical for the same skill group, as they use the same subsample. In both panels, the first-stage F-statistics for HS and LS immigrants are substantially smaller in the panel IV DML specifications than in the corresponding 2SLS regressions.
Focusing on HS immigrants (Panels A and B of Columns 1-4 in Table (ref)), the first-stage F-statistics from panel IV DML regressions fall well below the conventional rule-of-thumb value of 10, whereas the corresponding F-statistics of 2SLS is above the stock2005 threshold ($F = 39.5$). Therefore, under panel IV DML, the shift-share instrument for HS immigrants appears extremely weak, unlike with conventional 2SLS. The AR test statistics, robust to weak instruments by construction, strongly reject the null hypothesis of no effect at 1% level with panel IV DML in both panels (with the exception of gradient boosting in Panel B), indicating the presence of an effect of HS immigration on both outcomes. By contrast, the AR test statistic from 2SLS regressions rejects the null hypothesis at 1% in Panel A only, and at 10% level in Panel B. The corresponding AR CS from the panel IV DML regressions are either disjoint sets that include zero or unbounded, reflecting limited information to identify the effect due to weak instruments, even when a causal effect may be present. More generally, all second-stage estimates of the treatment parameter in these specifications should be interpreted cautiously with both estimators, as the weakness of the shift-share instrument makes standard inference unreliable.
For LS immigrants (Columns 5-8 of Panels A and B in Table (ref)), the first-stage F-statistics from panel IV DML estimators are well below the rule-of-thumb threshold of 10 while F-statistic from 2SLS barely exceeds 10. This raises serious concerns about the relevance of the shift-share instrument for LS immigrations. Consistent with this, the AR test statistics generally fail to reject the null hypothesis with panel IV DML in both panels, with a few exceptions. In Panel A, the AR test statistic rejects the null hypothesis at the 1% level for boosting specification; however, given the known instability of the boosting in small samples, the latter result should be interpreted with caution. In Panel B, the AR test statistic is borderline significant (at 10% level) for the Lasso specification, but broadly consistent with the absence of a treatment effect with the other learners. The AR CS obtained from panel IV DML are unbounded (entire real line) in all cases, indicating insufficient information to identifying the range of causal effects, if any. By contrast, the AR confidence sets from the 2SLS specifications are always bounded and contain the corresponding second-stage point estimates, a pattern consistent with the tendency of 2SLS to inflate identification strength under weak instruments.
Overall, our panel IV DML does not fully support 2SLS findings mainly due concerns about the strength of the shift-share instruments. While the instruments appear moderately strong under standard 2SLS specifications, they are found to be weak once high-dimensional and potentially nonlinear confounding is flexibly accounted for using the panel IV DML estimator. The AR diagnostics further indicate that, although some effects may be present, weak IV prevents the identification of the causal effect. While panel IV DML makes the identifying assumptions more plausible under the included controls, in this setting this comes at the cost of reduced effective relevance, limiting the reliability of second-stage point estimates.
In line with the article discussed in the previous section, moriconi2022 examine the impact of high-skilled (HS) and low-skilled (LS) immigration on individual voting patterns towards parties with a nationalist agenda, and perceived attitudes towards politics and immigration. The authors find that a higher share of HS immigrants decreases the intensity of nationalist preferences of native voters, and an opposite effect for LS immigrants.
In this third empirical application, we revisit the baseline specifications of moriconi2022 on the effect of HS and LS immigrants on nationalism (Columns 2-3 from their Table 6), and on the change in attitudes towards politics and immigrants (Columns 1, 3, 4 and 6 from their Table 10). The dependent variables of interest for the reanalysis are: an index of nationalism intensity of parties (Nationalism), constructed by matching individual-level party votes in national elections with the content of each party’s political manifesto; an index of trust in country parliament as political attitude; and a measure for a better place to live as attitude towards migrants. The set of control variables employed in analysis are the same as those in Table (ref).
We first discuss the key insights from the comparison between conventional 2SLS estimates with FE and with FD transformation (see the Online Appendix (ref) for more details). First-stage strength is consistently higher under FE than FD, mainly because FD drops the first time-period and subjects without consecutive observations. The shift-share instruments in both FE and FD regressions are not always strong, in contrast with the original study, where $16.30 < F < 104.70$ in all first-stage specifications. While the instruments can be considered moderately strong in all specifications using the sample of HS immigrants, this is not always the case for LS immigrants samples. Specifically, for the nationalism outcome, the instrument is essentially irrelevant ($F \approx 0$) for LS immigration under both estimators, undermining the credibility of the relative estimates. The second-stage results seem to be estimator-sensitive. The effect of HS immigration on political attitudes (originally only marginally significant at 10% level) survives only in FE, and the LS immigration effect on immigration attitudes disappears entirely.
Moving to the panel IV DML results, Table (ref) displays conventional 2SLS with FD estimates (Columns 1 and 5) and our panel IV DML estimation results (Columns 2-4, and 6-8) by dependent variable. For HS immigrants (specifications in Columns 1-4 Panels A-C), the first-stage F-statistics obtained with panel IV DML never exceed even the rule-of-thumb threshold of 10 and are always substantially smaller than those from 2SLS (always $F>16.30$). This provides initial evidence of the presence of weak instruments in these specifications for HS immigrants with panel IV DML. In all panels, the AR test with 2SLS and panel IV DML fails to reject the null hypothesis of no reduced-form effect when assuming that the treatment has no effect. The associated AR CS are either bounded with zero included, suggesting no second-stage effect, or unbounded (real line), indicating that the causal parameter cannot be identified with the available variation in the data due to weak IV. Both estimators therefore point to the absence of a causal effect, although only panel IV DML explicitly reveals the weak-identification problem.
For LS immigrants we comment on the results of the three outcomes (panels) separately. In Columns 5-8 of Panel A, the first-stage F-statistics are well below 10 for both 2SLS and panel IV DML, clearly indicating the presence of weak instruments. Although AR test statistics from panel IV DML suggest the possible presence of a second-stage effect across all three learners, not detected by 2SLS, the associated AR confidence sets span the entire real line. This reflects the lack of information in the instrument that prevents identification of the causal parameter. In this case, without AR diagnostics, the researcher might draw the same conclusion from 2SLS and panel IV DML, i.e. that the instrument is weak and no inference on the second stage can be made. AR diagnostics can inform us if there is a second stage effect regardless of the strength of the instrument. In this case, only the AR diagnostics from panel IV DML allow us to conclude that there might be an effect, which cannot be identified due to the weakness of the instrument.
In Columns 5-8 of Panels B-C, where the first-stage regressions are identical across outcomes, panel IV DML again produces substantially smaller F-statistics than 2SLS, whose values now only marginally exceed the stock2005's threshold. In Panel B, AR tests fail to reject the null hypothesis with both estimators (and only at 10% level with neural network), indicating no evidence of a causal effect on political attitudes. In Panel C on migration attitudes, results are mixed: AR tests from panel IV DML with Lasso fails to reject the null of no effect while neural networks and boosting reject at 10% level only. In contrast, the result of the 2SLS AR test conveys the message that there is a strong second-stage effect. Consistent with weak identification, AR confidence sets from panel IV DML either cover the entire real line (NNet) or disjoint segments (boosting), whereas those from 2SLS are bounded and exclude zero. In contrast, AR CS from Lasso is bounded but includes zero, which is aligned with the AR test result. In general, given the weakness of the shift-share instrument, statistical inference based on conventional 2SLS estimator is unreliable, and the second-stage coefficient cannot be interpreted as a credible causal effect.
Overall, panel IV DML reveals that once instrument exogeneity is strengthened through flexible control for confounding, in this case the shift-share instruments lose substantial relevance. As a result, the available variation is insufficient to support reliable inference on second-stage effects, remarking the importance of weak-identification diagnostics in IV applications.
For the Monte Carlo simulations, we consider a data generating process (DGP) inspired by the estimating equations of the empirical applications discussed in Section (ref). The data are generated from a static panel data model with high-dimensional covariates and individual fixed effects, which are strongly correlated with the included variables. The outcome depends on an endogenous treatment; treatment endogeneity arises from two sources: (a) the correlation between the structural and first-stage error term, and (b) the presence of the fixed effect in both structural and treatment equations.\footnote{The first source of endogeneity can be addressed with an IV approach, while the second with a panel data estimation approach (i.e., FE, FD and correlated random effects).} The instrument is not randomly assigned, being affected by covariates, but is exogenous to the structural error. The nuisance functions linking controls to the outcome, treatment, and instrument are sparse (only few variables out of thirty are relevant) and nonlinear (interactions are included). We consider both a strong-instrument design and a weak-instrument design. More details on the DGP are provided in the Online Appendix (ref).
As in the empirical applications, we employ our Panel IV DML estimator with different learners -- i.e., Lasso, a single-layer neural network (NNet), and gradient boosting with 100 trees (Boosting) -- to predict the nuissance functions of the covariates in a flexible way, and use the conventional 2SLS estimator as benchmark to assess the finite sample properties of the proposed estimator. The set of confounding variables used in conventional 2SLS and panel IV DML with NNet and Boosting regressions includes all raw variables and no nonlinear terms because, in practice, the analyst is agnostic on the true functional form of the covariates.\footnote{Neural network and gradient boosting should be able to capture those nonlinearities, by construction, even if unspecified, unlike 2SLS.} By contrast, panel IV DML with Lasso requires the analyst to specify a rich set of variables, including polynomials and interactions of the raw covariates (extended dictionary) to satisfy the sparsity assumption.\footnote{The constructed extended dictionary does not include the interaction terms present in Equations (ref)-(ref) of the DGP, allowing us to assess how effectively Lasso can flexibly recover similar covariate complexity.} The hyperparameters of the base learners are tuned via grid search bergstra2012 as explained in the Online Appendix (ref).
Tables (ref) and (ref) report the average results across 100 bootstrap samples for the strong- and weak-instrument cases, respectively. Each table summarizes the averages (for numerical quantities) or proportions (for binary indicators) over all Monte Carlo replications. The metrics reported include: (a) the bias and root mean squared error (RMSE) of the estimated target parameter $\widehat{\theta}_n$, together with the ratio of its estimated standard error to its Monte Carlo standard deviation (SE/SD); (b) the RMSE of the three nuisance functions; (c) the first-stage F-statistic and threshold rules for assessing instrument strength; and (d) the p-value of the Anderson-Rubin (AR) test statistic together with the corresponding type of AR confidence set. Each panel displays the results by cross-sectional sample size, $N=\{100,500, 1000, 5000\}$, with $T=10$ fixed, where each row represents a different estimator.
We begin the discussion of the Monte Carlo results with Table (ref), which corresponds to the strong-instrument design. In this setting, inference is reliable and, hence, the focus is primarily on the performance of the panel IV DML estimator for the target parameter. Panel IV DML consistently outperforms conventional 2SLS across all sample sizes, reflecting its ability to capture nonlinearities in the covariates that 2SLS cannot accommodate, by construction. In particular, the panel IV DML estimator recovers the structural parameter with high accuracy and precision when using Lasso and neural networks: the bias remains small (below 0.10) even in small samples ($N=100$). Gradient boosting exhibits larger bias in small samples, likely due to overfitting when the training sample is small, but still performs slightly better than 2SLS even in small samples. As the sample size increases, its bias converges toward that of Lasso and neural networks. The RMSE of the panel IV DML estimator declines remarkably with sample size, being consistent with root-N convergence of the panel IV DML estimator, whereas the 2SLS estimator does not display the same improvement. Consistent with the design, first-stage F-statistics are large and frequently exceed the lee2022 threshold of 104.70, indicating strong identification. The main exceptions are for Lasso and neural network in very small samples ($N=100$), where the average F-statistics exceed the stock2005 critical value of 16.30 but remain below 104.70. This does not indicate genuine weak identification, but rather reflects the limited information available in small samples. The Anderson-Rubin (AR) diagnostics support this interpretation: both the AR test statistic and the AR confidence set (CS) provide evidence of the existence of a treatment effect under this DGP across estimators and sample sizes. This pattern mirrors our empirical reanalysis of tabellini2020, where F-statistics fell similarly below 104.70, but the AR test statistic rejected the null hypothesis of no effect and the AR CS remained bounded. Exceptions arise again for Lasso and neural networks in very small samples ($N=100$), where limited information implies that the absence of a second-stage effect cannot be fully ruled out.
When the instrument is designed to be weak (Table (ref)), the second-stage estimates and statistical inference are unreliable for all estimators, so the Monte Carlo metrics for $\widehat{\theta}_n$ (bias, RMSE, SE/SD) are largely uninformative.\footnote{As expected, both conventional 2SLS and panel IV DML produce biased estimates of the target parameter across all sample sizes. Surprisingly, neural network exhibits comparatively smaller average bias in large samples ($N=5,000$) though at the cost of reduced precision ($SE/SD\gg1$).} The key consideration, therefore, is whether the weak-identification diagnostics correctly detect the lack of instrument relevance when using conventional 2SLS and panel IV DML estimators.
Conventional 2SLS systematically fails to detect weak instruments. That is, the first-stage F-statistics always exceed 104.70, the AR test always rejects the null hypothesis of no effect at 5% level in all samples, and AR 95% CS are always bounded and exclude zero. These diagnostics incorrectly suggest strong identification and identification of the treatment effect, despite the DGP being designed with weak instruments. By contrast, panel IV DML with Lasso and NNet correctly detects weak identification in most cases, especially in small samples. For these learners, the average first-stage F-statistic does not exceed the rule-of-thumb threshold of 10 for $N\le1000$, is always below the stock2005's critical value of 16.30 in small samples, and never approaches the lee2022 threshold of 104.70 with exceptions in large samples ($N=5,000$). Boosting never reaches the lee2022's threshold, although its F-statistics are systematically larger (around 25 on average) than the other learners. The AR test further highlights the contrast between panel IV DML and conventional 2SLS. For Lasso and NNet, the AR rejection rates are low in small samples (on average around 30% at $N\le500$ and 46% at $N=1,000$), appropriately reflecting weak identification, and increase to about 95% at $N=5,000$ since larger samples provide more information. The corresponding AR CS are often unbounded (either disjoint or real line) in small samples, signalling weak instruments. As $N$ increases, the AR CS become more often bounded, but still include zero, correctly indicating that the causal effect cannot be precisely identified under weak instruments. Boosting behaves similarly to 2SLS in small samples, but converges toward the neural-network patterns as the sample size grows. In this regard, panel IV DML provides inference that is more reliable in finite samples than 2SLS, following the patterns observed in the empirical applications of moriconi2019,moriconi2022, where the instrument was found weak by panel IV DML but not by 2SLS and the AR CS were mostly bounded with zero excluded.
In conclusion, the Monte Carlo simulations indicate that conventional 2SLS produces biased and inconsistent estimates under both weak and strong instruments, with the magnitude of the bias remaining similar across sample sizes and failing to exhibit root-N convergence. Conversely, panel IV DML not only improves estimation and statistical inference of the treatment effect under strong identification, regardless of the sample size, but it is also able to provide more robust and informative inference when identification is weak and conventional 2SLS can be severely misleading.
This paper introduces novel DML estimation procedures for partially linear panel regression models with fixed effects under endogenous treatment and potentially nonlinear covariate effects. We show that panel IV DML is a powerful alternative to 2SLS, capturing nonlinearities in the data while delivering more precise and reliable finite-sample inference. The panel IV DML toolkit allows researchers to complement traditional estimation techniques, providing a flexible, theoretically grounded, and empirically practical approach to causal estimation in panel settings where conventional methods are falling short.
Another key contribution of this paper is the integration of weak-instrument diagnostics (i.e., the first-stage F-statistic and the Anderson-Rubin test and confidence sets) into the panel IV DML framework. Our results show that AR inference is essential for reliable conclusions, as it remains valid under weak identification and can uncover treatment effects even when the F-statistic suggests limited instrument strength. Therefore, we advocate reporting AR diagnostics alongside conventional measures in panel IV DML applications.