EconBase
← Back to paper

Estimating interaction effects with panel data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,117 characters · 24 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating interaction effects with panel data

\onehalfspacing

abstractThis paper analyzes how interaction effects can be consistently estimated under economically plausible assumptions in linear panel models with a fixed $T$-dimension. We advocate for a correlated interaction term estimator (CITE) and show that it is consistent under conditions that are not sufficient for consistency of the interaction term estimator that is most common in applied econometric work. Our paper discusses the empirical content of these conditions, shows that standard inference procedures can be applied to CITE, and analyzes consistency, relative efficiency, inference, and their finite sample properties in a simulation study. In an empirical application, we test whether labor displacement effects of robots are stronger in countries at higher income levels. The results are in line with our theoretical and simulation results and indicate that standard interaction term estimation underestimates the importance of a country's income level in the relationship between robots and employment and may prematurely reject a null hypothesis about interaction effects in the presence of misspecification. Keywords: panel data, interaction effects, correlated random coefficients, robots.

Introduction

In many econometric applications, the relationship between an outcome variable $Y$ and a regressor $X$ depends on a so-called interaction variable $H$. For example, an earlier macroeconomic literature postulated that the effect of foreign capital inflows $X$ on economic growth $Y$ depends on some host-country fundamentals $H$, such as institutional quality, financial development, or a sufficiently educated labor force.\footnote{See, for example, borenszteinHowDoesForeign1998, alfaroFDIEconomicGrowth2004, durham04, alfaroDoesForeignDirect2010, burnsideAidPoliciesGrowth2000. Economically, this literature assumes that foreign capital $X$ and interaction variables $H$ complement each other in the production of output $Y$.} In micro data, the effect of new technologies on firms' productivity may depend on firm size, reflecting scale effects. And how a worker responds to a labor market reform may depend on age, reflecting that older individuals value leisure higher and income less in their labor supply choice than younger workers.

Panel data have at least two dimensions, $i$ and $t$, which provides us with opportunities to estimate interaction effects that uni-dimensional data do not have. In particular, panel data give us more flexibility to model heterogeneities in the relationship between $Y$ and $X$, on top of variation in $H$. Yet, those opportunities are rarely exploited in applied econometric analysis, which possibly reflects that theoretical work to date has largely ignored the benefits of modeling heterogeneities across panel units for the estimation of interaction effects in panel data.

Consider the relationship between employment $Y$ and robotization $X$. We may hypothesize that this relationship is more negative in countries $i$ with a higher wage level $H_i$ because firms experience higher pressure to save labor costs. The standard approach to estimating such an interaction effect in a linear model with panel data is to regress employment $Y_{it}$ on robotization $X_{it}$, and its interaction with countries' wage levels $H_i$, together unit-specific fixed effects $\alpha_i$ and possibly additional variables in $Z_{it}$, where $t$ may index time, industries, firms etc.:

equation[equation omitted — 126 chars of source]

The conventional approach to estimate interaction effects in a linear model with panel data seems to be least squares based on equation (ref).\footnote{See, for example, burnsideAidPoliciesGrowth2000, Shambaugh2004, ListStrum2006, AmitiKonings2007, spilimbergoDemocracyForeignEducation2009, EpifaniGancia2009, dufloPeerEffectsTeacher2011, BloomSadunReenen2012, BermanMartinMayer2012, BloomDracaReenen2016, Storeygard2016, AlsanGoldin2019, HerreraTrebesch2020, manacordaLiberationTechnologyMobile2020.} We refer to this approach as the interaction term estimator (ITE). The partial effect of $X$ (robots) on $Y$ (employment) linearly depends on $H$ (wage level) in this model, as is easily verified by taking the partial derivative $\partial{Y}/\partial{X} = \beta + H\kappa$.

Our paper focuses on the interaction effect $\kappa$ and how it can be consistently estimated under economically plausible assumptions with panel data that have a short $T$ dimension. Standard linear regression assumptions in application to equation (ref) require the error term $\nu$ to be independent of the regressors $X, XH, Z$. This assumption is easily violated in many interaction models: the economic reasons for effect heterogeneity in the relationship between $X$ and $Y$ are extensive and rarely independent of the regressors. Recall the example of robots and employment. In a technology-minded country, firms $t$ may find various ways to substitute workers' tasks in the production process by robots. The partial relationship between robots $X$ and employment $Y$ will hence be more negative in those countries than in less tech-minded countries because a robot can replace more workers in the former. Since this aspect is unmodeled in equation (ref), and tech-mindedness is not directly observable, this heterogeneity across countries will end up in the error term $\nu$. And since tech-mindedness likely gives rise to higher robotization $X$ in the first place, we observe a standard endogeneity bias: $\nu$ is correlated with $X$ and $XH$ in equation (ref).

Panel data models for interaction effects can be deceptive because the inclusion of the additive fixed effects $\alpha_i$ may lull the researcher into a false sense of having controlled for unobserved heterogeneity. This is not the case because in the context of interaction effects we are not concerned about (additive) heterogeneity in levels of $Y$ and $X$, which are indeed absorbed through $\alpha_i$, but about effect heterogeneity in the relationship between $X$ and $Y$. Our research shows that a common source of such unobserved effect heterogeneity can be easily addressed with panel data. Our exposition focuses on the simple case where the interaction variable $H$ only varies in the $i$ dimension since this is plausibly the key dimension giving rise to heterogeneity in the partial effect of $X$ on $Y$. An extension to interaction variables that vary in the $it$ dimension and an application to higher-dimensionality panel data is provided in an earlier working paper version of this paper MurisWacker2022 and follows the same rationale: additive fixed effects do not control for important aspects of heterogeneity.

The key contribution of our paper is to advocate for a different approach to the estimation of interaction effects with panel data and to analyze the exogeneity conditions that are are sufficient for consistent estimation of $\kappa$. We therefore develop a framework of correlated unobserved effect heterogeneity that explicitly includes unit-specific effects $\beta_i$ as parameters in the relationship between $X$ and $Y$ Chamberlain1992. Building on this framework, we propose a correlated interaction term estimator (CITE) that obtains an estimate for the interaction term coefficient $\kappa$ by first running the regression $Y_{it} = \alpha_i + X_{it}\beta_i + Z_{it}^\prime\gamma + U_{it}$ and subsequently projecting the unit-specific effects $\widehat{\beta}_i$ onto $H_i$ in a second-step regression.\footnote{Such an approach has occasionally been used in applied work. See particularly couttenierViolentLegacyConflict2019a and macurdyEmpiricalModelLabor1981. giesselmannInteractionsFixedEffects2020 discuss CITE, but make a case for a double-demeaned estimator. Alternatively, running unit-by-unit regressions to recover common parameters is a strategy that has been used in different panel data settings, see e.g. fernandezvalNonadditiveUnobservedHeterogeneity2013a. In any case, we are not aware of existing work that demonstrates that ITE and CITE are distinct and that analyzes their asymptotic properties.} Since this second-step regression of $\widehat{\beta}_i$ onto $H_i$ is a simple cross-section OLS regression, it is easily to implement and can be assessed with the extensive, well-known statistical toolkit that has been developed for OLS regressions.\footnote{If one is interested in time-varying interaction variables, those can be considered part of $Z$ and the second step can be omitted. Unobserved effect heterogeneity across panel units is directly absorbed in $\widehat{\beta}_i$ in this case. See an earlier working paper version of our paper for a more explicit treatment of such time-varying interactions MurisWacker2022.}

A key result of the rigorous comparison of ITE and CITE that we provide in this paper is that the exogeneity restrictions that guarantee the consistency of CITE for $\kappa$ are not sufficient for the consistency of ITE. On top of conditions that apply to both estimators, CITE requires unobserved effect heterogeneity to be uncorrelated with the interaction variable $H$, while ITE requires this effect heterogeneity to be uncorrelated with all right-hand side regressors $H, X, Z$ for consistent estimation of $\kappa$. This suggests that researchers need strong exogeneity arguments when estimating interaction terms with ITE in panel data.

One plausible concern about our newly proposed CITE is that its superior bias protection comes at the cost of efficiency. By treating the unit-specific effects $\beta_i$ as parameters that need to be estimated, one may expect a higher variability of CITE for $\kappa$, compared to ITE. We present simulation results demonstrating that this is not necessarily the case when considering different Monte Carlo designs. There is hence no obvious penalty for using CITE in terms of efficiency, while its advantages in terms of consistency are clear.

Another advantage of CITE is that we can demonstrate that standard inference procedures, available in common software, can be applied under economically plausible assumptions about the structure of error terms. The finite sample performance of these inference procedures is also demonstrated in our simulation section. Finally, CITE does not require a large number of time periods. Our paper develops theoretical consistency results under fixed-$T$ and our simulations suggest that relative efficiency is often similar to ITE around $T=4$, depending on the simulation design. The fact that we can prove consistency of CITE for fixed-$T$ also sets our paper apart from the literature on heterogeneity in large-$T$ panels.\footnote{balliInteractionEffectsEconometrics2013 observe that unobserved effect heterogeneity can lead to inconsistency in the ITE, and suggest that “if the time-series dimension of the data is large, one may directly allow for country-varying slopes”. We show that such slopes should generally be preferred regardless of the time-series dimension.}

Related literature

Our framework builds on the literature on correlated random coefficient models Chamberlain1992,arellanoIdentifyingDistributionalCharacteristics2012,grahamIdentificationEstimationAverage2012,laage2024crcpublished,sasakiSlowMoversPanel2021a. Specifically, we build on Corollary 1 in arellanoIdentifyingDistributionalCharacteristics2012 to study the estimation of interaction effects in panel models. We adapt and extend the theoretical results in this literature to advocate for CITE over ITE.

Models with correlated coefficient heterogeneity have previously been used to study the identification and estimation of average treatment effects in various settings, see for example wooldridgeFixedEffectsRelatedEstimators2005, murtazashviliFixedEffectsInstrumental2008, verdierAverageTreatmentEffects2020,dechaisemartinTwoWayFixedEffects2022a, sloczynskiInterpretingOLSEstimands2022, chaisemartinMoreRobustEstimators2023, winkelmannNeglectedHeterogeneitySimpsons2024, dechaisemartinCredibleAnswersHard2024. This literature is mostly concerned with the identification and estimation of (convex combinations of) average coefficients, partial effects, and treatment effects, from cross-sectional and panel data. In contrast, we focus on the estimation of interaction effects from panel data in the presence of unobserved effect heterogeneity.

Organization and preview from an applied perspective

Section (ref) introduces the unobservable effect heterogeneity model. Section (ref) formally defines the two estimators. Section (ref) presents our main results about the consistency of CITE. Section (ref) discusses inference. Section (ref) contains a Monte Carlo simulation study. Section (ref) contains an empirical application where we investigate whether worker displacement effects of robots are stronger in countries at higher income levels. Section (ref) concludes and discusses some avenues for future research.

In sections (ref) - (ref), which focus on theory, we have mostly opted for a notation of within-transformed stacked variables to facilitate the theoretical exposition. Applied readers who are at less comfort with this notation, or less interested in the details of our theoretical results, are suggested to focus on the following aspects: the part of section (ref) prior to notational definitions to understand the unobservable effect heterogeneity model we have in mind and the first two paragraphs of section (ref) for an informal definition of ITE and CITE.

The crucial exogeneity conditions required for consistent estimation of interaction terms in panel data can be found in Assumptions (ref)-(ref) in Section (ref). It is particularly instructive to compare Assumption (ref), which is sufficient for consistency of CITE, to the stronger Assumption (ref), which is required for consistency of ITE. Remark (ref) compares both assumptions with an example from our application.

The essence of section (ref) for applied researchers is that conventional heteroskedasticity-robust standard errors in the second-stage projection of $\widehat{\beta}_i$ onto $H_i$ are sufficient for valid inference under CITE. Section (ref) is accessible without delving into technical or notational details, but can also be skipped if one is already convinced about the advantages of CITE in terms of consistency, relative efficiency, inference, and their finite sample properties. The application section (ref) will be most instructive for applied readers to understand the implementation of our approach in practice.

Model

We are interested in estimating interaction effects in static linear panel models with a short $T$ dimension. In our framework, the effect of $X$ on $Y$ linearly depends on observables $H$ and on additional unobservable sources of effect heterogeneity. We model this unobserved heterogeneity in the effect of $X$ on $Y$ as follows:

align[align omitted — 343 chars of source]

where $i$ indexes cross-sectional units, and $t$ indexes time periods (or other panel dimensions). The beginning of Section (ref) clarifies how the standard panel interaction term model relates to this framework.

The outcome equation (ref) of our framework describes how a dependent variable $Y_{it} \in \mathbb{R}$ responds to a change in a regressor $X_{it} \in \mathbb{R}$, \footnote{Our analysis is easily generalized to vector-valued $X_{it}$.} allowing for additive fixed effects $\alpha_i \in \mathbb R$, control variables $Z_{it} \in \mathbb{R}^{K_z}$, and an error term $U_{it} \in \mathbb R$. It is important to note that the control variables $Z_{it}$ may include interaction terms with $it$ variation. As an earlier working paper version of our paper shows more thoroughly, CITE poses less restrictive exogeneity conditions for their consistency as well MurisWacker2022.

The heterogeneity equation (ref) of our framework relates effect heterogeneity $\beta_i$ to observed, time-invariant interaction variables $H_i \in \mathbb{R}^{K_h}$. The associated interaction term coefficient $\kappa \in \mathbb{R}^{K_h}$ is the main object of interest in our paper.\footnote{Note that we obtain the constant $\kappa_0$ in the heterogeneity equation (ref) by setting the first element of all $H_i$ equal to 1 and that this constant can be interpreted as the average $\beta_i$ across panel units, conditional on all other elements of $H_i$.} The effect of $X$ on $Y$ may additionally vary, even when holding $H_i$ constant, because of unobserved effect heterogeneity $\epsilon_i \in \mathbb R$. The relationship between unobserved effect heterogeneity and $(X,H)$ plays an important role in our analysis (in a similar way as the relationship between unobserved heterogeneity $\alpha_i$ and $X$ is important in the conventional additive fixed effect model).

remarkThe index $t$ need not refer to time, but can refer to students $t$ for a given classroom $i$, counties $t$ within a given state $i$, employees $t$ within a given firm $i$, etc. The empirical application in Section (ref) provides an example and section 6.2 in an earlier working paper version illustrates how our framework applies to higher-dimensional panel data MurisWacker2022.

Some additional notation simplifies the exposition in the remainder of the paper. First, estimation will be based on the within-transformed outcome equation:

equation[equation omitted — 172 chars of source]

where \[ \widetilde{Y}_{it} = Y_{it} - \frac{1}{T} \sum_{s=1}^{T} Y_{is}, \quad \widetilde{X}_{it} = X_{it} - \frac{1}{T} \sum_{s=1}^{T} X_{is}, \] and $\widetilde Z_{it}$ and $\widetilde U_{it}$ are defined analogously.

Second, the reduced form of the within-transformed equations (ref)--(ref) is

align[align omitted — 407 chars of source]

with an error term $V_{it}$ that depends on the unobserved effect heterogeneity $\epsilon_i$.

Third, we collect the within-transformed outcome equation (ref) across $t=1,\cdots,T$ for a given $i$ to obtain

equation[equation omitted — 160 chars of source]

where \[ \widetilde{X}_i = (\widetilde{X}_{i1},\cdots,\widetilde X_{iT}) \in \mathbb R^T, \] and $\widetilde Y_{i} \in \mathbb R^T$, $\widetilde X_{i} $, $\widetilde Z_{i} \in \mathbb R^{T \times K_z}$, and $\widetilde U_{i} \in \mathbb R^T$. Accordingly, we obtain

align[align omitted — 126 chars of source]

where $\Psi_i = \widetilde X_i H_i^\prime \in \mathbb R^{T \times K_h}$ and $V_i = \widetilde U_i + \widetilde X_i \epsilon_i \in \mathbb R^T$ are stacked versions of $\Psi_{it}$ and $V_{it}$, respectively.

Estimators

We consider two estimators for the interaction term coefficient $\kappa$: the interaction term estimator (ITE) and the correlated interaction term estimator (CITE). The ITE is a fixed effects regression \footnote{By “fixed effects regression” we mean a regression that includes a dummy variable for each $i$, i.e. the within estimator.} of $Y$ on the interaction term $\Psi = XH$ and $Z$, as suggested by the reduced form (ref). It is the default approach in applied research when the effect of $X$ is expected to depend on $H$: “just add an interaction term”. This standard ITE panel interaction model is usually written along the lines $Y_{it} = \alpha_i + X_{it} \beta + X_{it} H_{1i} \kappa_1 + Z'_{it}\gamma + \nu_{it}$ in many applied papers, which is a special case of our framework that one obtains by substituting equation (ref) into equation (ref), imposing $\epsilon_i = 0 \ \forall \ i$, and setting $\beta = \kappa_0$. A formal definition of ITE is provided in Section (ref).

The CITE, defined formally in Section (ref), is a two-step estimator. The first step is a fixed effects regression of $Y$ on $Z$ and interactions of $X$ with dummy variables for each $i$: $Y_{it} = \alpha_i + X_{it}\beta_i + Z'_{it}\gamma + U_{it}$. The coefficients on the dummy variable interactions, $\widehat \beta_i$, are individual-specific effects of $X$ on $Y$ that capture unobserved effect heterogeneity. The second step is a regression of $\widehat \beta_i$ on $H$. CITE hence directly follows the structure in equations (ref)--(ref): the first-step regression is based on the outcome equation (ref), and the second-step regression is based on heterogeneity equation (ref). Practical implementation of CITE in standard software is simple: it just requires interacting $X$ with the unit-specific dummy variables, saving the estimated coefficients, and regressing them on $H$ in a second step.

The remainder of this manuscript assumes a random sample is available.

assumption[Random sampling] For each $i=1,\cdots,n$, the observed data is $$W_i = \left(Y_{i1},\cdots,Y_{iT},X_{i1},\cdots,X_{iT},Z_{i1},\cdots,Z_{iT}, H_i\right),$$ generated by equations (ref)--(ref). The random sequence $\{W_i\}_{i=1}^n$ is independent and identically distributed.

Correlated interaction term estimator

Sufficient variation in $\widetilde X_{i}$ for all $i$ is required for the CITE to be well-defined.

assumptionThere exists an $h>0$ such that \[ \inf_{i} \left( \sum_{t=1}^T \widetilde X_{it}^2 \right) \geq h. \]

This guarantees that the least squares estimators for the $(\beta_i)$ are well-defined, which avoids the identification issues addressed in grahamIdentificationEstimationAverage2012 and arellanoIdentifyingDistributionalCharacteristics2012. If Assumption (ref) does not hold for all $i$, our analysis applies to the subpopulation for which it does hold.

To obtain the within-transformed data, we define the residual maker matrix \[ M_{i} = I_T - \widetilde X_{i} \left( \widetilde X_{i}^\prime \widetilde X_{i} \right)^{-1} \widetilde X_{i}^\prime, \] where $I_T$ is the $T \times T$ identity matrix. Premultiplication by $M_i$ obtains

align[align omitted — 288 chars of source]

and estimation of $\gamma$ can be based on equation (ref).

assumptionThe matrix $E\left(\widetilde Z_{i}^\prime M_{i} \widetilde Z_{i}\right)$ is invertible.

Assumptions (ref)--(ref) guarantee that the following is well-defined for large $n$.

definitionCITE for $\gamma$ is given by \[ \widehat{\gamma}_{n} = \left( \sum_{i=1}^n \widetilde Z_{i}^\prime M_{i} \widetilde Z_{i} \right)^{-1} \sum_{i=1}^n \widetilde Z_{i}^\prime M_{i} \widetilde Y_{i}. \]

To define the CITE for $\kappa$, we also need sufficient variation in $H_i$.

assumptionThe matrix $E\left(H_{i}^\prime H_{i}\right)$ is invertible.

For each $i$, the estimator for the individual-specific effect of $X$ on $Y$ is

align*[align* omitted — 165 chars of source]

The CITE for $\kappa$ is obtained from a regression of $\widehat \beta_i$ on $H_i$.

definitionThe CITE for $\kappa$ is \[ \widehat{\kappa}_n = \left(\sum_{i=1}^n H_{i}^\prime H_{i}\right)^{-1} \sum_{i=1}^n H_{i}^\prime \widehat{\beta}_i. \]

Interaction term estimator

The ITE is based on equation (ref). To formally define it, we project out the fixed effects $\alpha_i$ using the within transformation. For the ITE to be well-defined, we require the following `no multicollinearity' condition:

assumptionThe matrix $E\left( [\Psi_{i} \; \widetilde Z_i]^\prime [\Psi_{i} \; \widetilde Z_i] \right)$ is positive definite.
definitionGiven Assumption (ref), the ITE for $\theta = (\kappa,\gamma)$ is well-defined: \begin{align*} \widecheck{\theta}_{n} &= \left(\sum_{i=1}^n [{\Psi}_{i} \; \widetilde Z_i]^\prime [{\Psi}_{i} \; \widetilde Z_i] \right)^{-1} \sum_{i=1}^n [{\Psi}_{i} \; \widetilde Z_i]^\prime \widetilde{Y}_{i} \\ &= (\widecheck \kappa_n, \widecheck \gamma_n). \end{align*}

Consistency

We now show that a set of exogeneity conditions that is sufficient for consistency of CITE is not sufficient for consistency of ITE. Throughout, we maintain the following strict exogeneity assumption.

assumption[Strict exogeneity] $E\left(\left.U_{i}\right|X_{i},Z_{i},H_{i}\right) = 0$.

This assumption is similar to the strict exogeneity assumption that is standard in the literature on correlated random coefficient panel models,\footnote{See Chamberlain1992, arellanoIdentifyingDistributionalCharacteristics2012, grahamIdentificationEstimationAverage2012. See laageCorrelatedRandomCoefficient2020b for an approach without exogeneous regressors.} and in textbook treatments of fixed-$T$ linear panel models with additive fixed effects. If one is only interested in $\gamma$, and not in $\kappa$, Assumption (ref) can be relaxed for CITE (but not for ITE) to $E\left(\left.U_{i}\right|X_{i},Z_{i}\right) = 0.$

CITE

It is sufficient for the consistency of CITE that the unobserved effect heterogeneity $\epsilon_i$ is orthogonal to the time-invariant interaction variable $H_i$.

assumption[Exogeneity, $\epsilon$]$E\left(\left.\epsilon_{i}\right|H_{i}\right)=0.$

This assumption is easily understood when recalling that the second step of CITE is a cross-sectional OLS regression: it is equivalent to the textbook OLS exogeneity condition that requires the regressor $H_i$ to be orthogonal to the error term $\epsilon_i$.

To understand the advantages of CITE in applied econometric work, it is crucial to emphasize that Assumption (ref) does not involve the regressors $\left(X,Z\right)$. We discuss this advantage in contrast with ITE, including an example, in Remark (ref) at the end of Section (ref).

thm(i) If Assumptions (ref)--(ref) and (ref) hold, and if $E\left\|Z_i\right\|^2$ and $E\left\|Y_i\right\|^2$ are bounded, then $\widehat{\gamma}_{n}\stackrel{p}{\to}\gamma \text{ as }n\to\infty.$ (ii) If additionally, Assumptions (ref) and (ref) hold, and if $E\|H_i\|^2$, $E\|\beta_i\|^2$, $E\|\widetilde X_i\|^4$, and $E\|\widetilde Z_i\|^4$ are bounded, then $\widehat{\kappa}\stackrel{p}{\to}\kappa \text{ as }n\to\infty.$

For consistency of CITE, we do not need to restrict the joint distribution of $(Z,X,\epsilon)$ beyond the existence of certain moments. The proof of consistency is standard and can be found in Appendix (ref). Because distribution theory for ITE and CITE is standard, we omit the remaining proofs.

Inconsistency of ITE

Theorem (ref) establishes that the exogeneity restrictions in Assumptions (ref) and (ref) guarantee the consistency of CITE for $\kappa$ and $\gamma$. We now show that the exogeneity restrictions are not sufficient for the consistency of ITE for $\kappa$.

Consider the special case that $H$ is scalar, and that there is no $Z$, so the reduced form simplifies to: \[ \widetilde Y_{it} = \widetilde X_{it} H_i \kappa + \left(\widetilde U_{it} + \widetilde X_{it} \epsilon_i\right). \] The ITE is simply a linear regression of $\widetilde Y$ on $\widetilde X H$. Consistency of OLS requires that the error term is orthogonal to the regressors. Given Assumption (ref), \[ E\left(\widetilde X_{it} H_i \widetilde U_{it}\right) = 0. \] However, under the exogeneity Assumptions (ref) and (ref) we may have \[ E\left(\widetilde X_{it}^2 H_i \epsilon_i\right) \neq 0. \] It is straightforward to construct data generating processes where this happens, see our simulation study in Section (ref).

To ensure the consistency of ITE, one could further strengthen the conditions. For example, strengthening Assumption (ref) to the following stronger correlated random effects assumption restores consistency:

assumption[Exogeneity, $\epsilon$, strengthened]$E\left(\left.\epsilon_{i}\right|H_{i},X_i,H_i\right)=0.$

This restriction requires the unobserved effect heterogeneity $\epsilon_i$ to be orthogonal to $X$ and $Z$ in addition to $H$. Consistency of ITE hence requires that unobserved effect heterogeneity is independent of all relevant regressors in the model.

remarkComparison of Assumptions (ref) and (ref) facilitates a key insight of our paper: the latter (which ensures consistency of ITE) is more demanding than Assumption (ref) (which is sufficient for consistency of CITE), since it requires additional orthogonality of $\epsilon_i$ to $X$ and $Z$ (in addition to $H$). In the context of our empirical application, where we estimate how the effect of robots $X$ on employment $Y$ depends on interaction variables $H$, ITE requires that heterogeneity in the relationship between robotization and employment across countries is uncorrelated with countries' robot adoption. As discussed in our introduction, this is implausible because certain countries may be more open to new technologies, in which case they are more likely to adopt robots $X$ and find ways to replace workers with robots, which gives rise to effect heterogeneity $\epsilon_i$. CITE does not require such effect heterogeneity to be independent of robotization $X$ for consistent estimation of the interaction effect.

Inference

In this section, we discuss how to conduct inference on the interaction effect $\kappa$ for both estimators. For both estimators, inference is standard and can be performed using widely available software. We sketch our arguments below for the case where $Z$ is absent. The general case with $Z$ is similar, and it is omitted because the additional notation obscures the argument.

In this paper, we are concerned with inference for the partial effects. Empirical researchers may also be interested in the distribution of partial effects. Inference for this object is significantly more challenging, see for example arellanoIdentifyingDistributionalCharacteristics2012 and fernandez2022dynamic.

ITE

Inference for the ITE follows the standard theory for linear panel data models with clustered errors. In panel data, cluster-robust standard errors account for arbitrary correlation patterns in the error terms within each panel unit while maintaining the assumption of independence across units. Clustered standard errors address the fact that observations from the same panel unit (like multiple years of data from the same country) are typically not independent.

Two features of our model make clustering necessary. First, the within transformation induces correlation in the transformed errors $\widetilde{U}_i$ within each panel unit $i$. Second, the presence of unit-specific unobserved effect heterogeneity $\epsilon_i$ in the composite error term $V_i = \widetilde{U}_i + \widetilde{X}_i \epsilon_i$ creates an additional source of within-unit correlation. However, our random sampling assumption (Assumption (ref)) ensures independence across panel units, making cluster-robust standard errors appropriate.

The asymptotic variance of the ITE takes the standard sandwich form for clustered data: \[ \text{var}(\widecheck{\kappa}_n) = \left(\sum_i \Psi_i^\prime \Psi_i\right)^{-1} \left(\sum_i \Psi_i^\prime \text{var}(V_i|\Psi_i) \Psi_i\right) \left(\sum_i \Psi_i^\prime \Psi_i\right)^{-1}. \] This variance can be consistently estimated in the usual way.

Implementation of cluster-robust standard errors is straightforward in standard statistical software. For instance, in current STATA versions, one would use the option vce(robust) for a standard panel command with interactions, like xtreg y x c.x\#c.h, fe vce(robust). Similar commands are available in R, Python, and other commonly used statistical packages. Our simulation study in Section (ref) documents the finite-sample performance of this inference approach.

CITE

Because of the two-step nature of CITE, it may be surprising that heteroskedasticity-robust standard errors in the second-step regression are sufficient for valid inference. To show that this is the case, recall that CITE for $\kappa$ is a regression of the estimated individual-specific coefficients $\widehat \beta_i$ on $H_i$, \[ \widehat{\kappa}_n = \left(\sum_{i=1}^n H_{i}^\prime H_{i}\right)^{-1} \sum_{i=1}^n H_{i}^\prime \widehat{\beta}_i. \] The corresponding regression equation is \[ \widehat \beta_i = \beta_i + (\widehat \beta_i - \beta_i) = H_i^\prime \kappa + \epsilon_i + (\widehat \beta_i - \beta_i), \] where the second equality follows from the heterogeneity equation (ref).

Therefore, $\widehat{\kappa}_n$ is based on a linear regression with a composite error term $\epsilon_i + (\widehat \beta_i - \beta_i)$. This composite error term is independent across $i$, and heteroskedastic by construction. Independence follows from two observations: first, Assumption (ref) implies that $\epsilon_i$ is independent across $i$, and second, the first-step estimation errors $\widehat \beta_i = \left(\widetilde X_{i}^\prime \widetilde X_{i}\right)^{-1} X_{i}^\prime Y_{i}$ are independent across $i$ because they only depend on $(\widetilde{X}_i, \widetilde{Y}_i)$, which are also independent across $i$ by Assumption (ref). The composite error term is heteroskedastic because the variance of the first-step estimation error $(\widehat \beta_i - \beta_i)$ depends on $\widetilde{X}_i$, even if $\epsilon_i$ is homoskedastic. Consequently, using heteroskedasticity-robust standard errors (whiteHeteroskedasticityConsistentCovarianceMatrix1980) in the second-step regression leads to valid inference.

Implementation is straightforward in standard statistical packages. For example, in STATA, after obtaining the unit-specific coefficients `bhat' via a first-stage regression, one would use reg bhat H, robust to obtain valid standard errors. Our simulation study in Section (ref) confirms good finite-sample performance of this approach. When control variables $Z$ are present, the argument remains valid asymptotically, because $\gamma$ is estimated at rate $\sqrt{nT}$. The contribution to sampling variation is negligible compared to that from $\widehat \beta_i$. The latter is estimated using $T$ observations.

Monte Carlo simulations

We use a Monte Carlo simulation to document the bias in ITE, the relative efficiency of ITE and CITE, and the finite sample performance of the inference procedures.

We use a data generating process in which all building blocks follow a normal distribution, are mutually independent, and are independent across $i$. For the heterogeneity equation, we let $H_i \sim \mathcal{N}\left(1,1\right)$, $\epsilon_i \sim \mathcal{N}(0,\sigma_\epsilon)$, and $\beta_i = \kappa H_i + \epsilon_i.$ For the outcome equation, we omit $Z$, and construct $X_{it}$ as follows:

align*[align* omitted — 181 chars of source]

The parameter $\Delta$ controls the amount of correlation between $X_{it}$ and $\epsilon_{it}$. Values of $\Delta$ away from zero favor CITE because they create a correlation between $X$ and unobserved effect heterogeneity $\epsilon$ (see Section (ref)). $\lambda_t$ creates common time trends for all panel units. Finally, we generate $U_{it} \sim \mathcal N(0,\sigma_u)$, $\alpha_i \sim \mathcal N(0,\sigma_a)$, and compute $Y_{it} = \alpha_i + X_{it} \beta_i + U_{it}.$

All $\sigma$ parameters default to 1. We will vary the number of time periods $T$ and the number of cross-section units $n$, as well as the value of the interaction term coefficient $\kappa$ and other design parameters.

The bias in ITE

We start with the design outlined above, with $n = 100$, $T=5$, and $\kappa = 0.5$. Figure (ref) plots the mean of ITE and CITE across 10000 simulations, as a function of $\Delta \in [-0.3,0.3]$. At $\Delta = 0$, there is no endogeneity bias in the ITE, and both estimators for $\kappa$ have a mean of 0.5. For any $\Delta \neq 0$, the ITE is biased, whereas the mean CITE remains at 0.5. The bias in the ITE can be severe.

Figure (ref) plots the mean of ITE and CITE as a function of $\kappa \in [-0.5,0.5]$ at $\Delta = 0.4$. An unbiased estimator should follow the 45-degree diagonal, which is achieved by CITE. Conversely, the bias in ITE is substantial ($\approx 0.3$) and does not change much when varying $\kappa$. The shaded area corresponds to values of $\kappa$ where the mean of ITE has the opposite sign of $\kappa$.

figure[figure omitted — 573 chars of source]

Relative efficiency

Next, we compare the variability of ITE and CITE. We continue using $n = 100$, set $\Delta = 0$ (otherwise ITE is biased), and set $\kappa = 0$ (results for other values are comparable). We will vary $T \in \{3,\cdots,8\}$. In our baseliness results, the $\sigma$ parameters are at $1$. We will consider the effect of changing them on the relative efficiency of ITE and CITE as a function of $T$.

Figure (ref) presents the baseline results. At $T=3$, the ITE is more efficient than CITE. At $T>3$, the CITE has a lower standard deviation. Figures (ref)--(ref) document that the profiles shift as we vary the values of $\sigma$. For example, an increase in effect heterogeneity $\sigma_\epsilon$ favors the CITE. For example, at $\sigma_\epsilon = 0.1$ (Figure (ref)), ITE is more efficient than CITE over the entire plotted range. Conversely, for $\sigma_\epsilon = 2$, CITE dominates ITE over $T$ (Figure (ref)).

Across all designs, the variability of both estimators decreases with $T$. The relative efficiency of the two estimators depends on the design, and there is no clear ranking that holds across all designs. Our main takeaway is that there is no clear penalty for using CITE to protect against bias.

figure[figure omitted — 1,517 chars of source]

Inference

In Section (ref), we established simple inference procedures for ITE and CITE: use cluster-robust standard errors for ITE, and use heteroskedasticity-robust standard errors in the second step of CITE.

In both cases, the asymptotic distribution theory is standard, and standardized test statistics are asymptotically standard normal. Figure (ref) plots the density estimates of the standardized test statistics across 100000 simulations in a design with $n=1000$, $T=20$, $\kappa=0.5$, $\Delta=0$, and all $\sigma$ parameters equal to 1. The dashed line is the pdf of a standard normal distribution. The dashed line is barely visible in the plot, because the distribution of $t-$statistics of both ITE and CITE are right on top. For the same design, Figure (ref) plots the density estimates of ITE and CITE. They appear bell-shaped, with CITE less variable than ITE. In conclusion, the results in Figure (ref) are in line with the asymptotic theory.

figure[figure omitted — 549 chars of source]

Figure (ref) provides further evidence on the asymptotic validity of our inferential procedure. We plot the rejection frequencies for $H_0:\kappa_0 = 0.5$, varying $\kappa$ in the data generating process along the horizontal axis. Both panels use 100000 simulations, and a nominal level of $\alpha = 0.05$ (indicated as a dashed horizontal line). Figure (ref) presents results for a design with $n=100$, $T=8$, $\Delta = 0$, and all $\sigma$ parameters equal to 1. In this design, ITE is consistent. At $\kappa = 0.5$, rejection frequencies are close to the nominal level of $0.05$ for both estimators. Rejection frequences away from $\kappa = 0.5$ are better for CITE: an incorrect null hypothesis is more frequently rejected. Figure (ref) presents results for a design with $n=500$, $T=8$, $\Delta = 0.3$, and all $\sigma$ parameters equal to 1. In this design, the ITE curve is shifted left due to the bias in ITE under this design. Compared to the left panel, due to the increased sample size, the CITE curve is steeper around 0.5, and the rejection frequency at $\kappa = 0.5$ is very close to $\alpha$.

figure[figure omitted — 610 chars of source]

What does the relation between robots and employment depend on?

The relationship between robots and employment has attracted considerable attention over recent years. Mondolo2022, Juratetal2023, and Restrepo2023 survey this literature, which is highly relevant beyond academia: a solid knowledge about the labor market consequences of technical change helps us understand whose jobs may be automated, and to design targeted policies.

Yet, the empirical evidence on the robot-employment nexus is mixed to date. A meta-analysis by Guarascioetal2024 summarizes 33 studies with 644 estimates and finds considerable variation across studies and estimates, with a very small average negative relationship between robotization and employment. Within this literature, our application is related to cross-country industry level studies in the spirit of GraetzMichaels2018 and deVriesetal2020.

One possible reason for the mixed findings about the robot-employment relationship is cross-country heterogeneity, as suggested by previous empirical results by Reljicetal2023 and ChenFrey2024, for example. Such heterogeneity is highly plausible on theoretical grounds. Consider the seminal partial equilibrium model of AcemogluRestrepo2020jpe, where employment $L$ in industry $t$ and region $i$ is affected by robotization $R$ through three channels: a negative labor displacement effect of robotization technology, a positive demand effect reflecting the productivity gains from robotization, and a composition effect, which captures that industries that robotize require more (non-automated) labor from non-expanding sectors. Formally, those three channels are captured in their equation:

equation[equation omitted — 346 chars of source]

where $\theta_t$ is a technology parameter indicating the range of tasks that can be automated (i.e., performed by robots), $Y_c$ is output, $\sigma>0$ is the elasticity of substitution (between goods of different industries) and $P_{it}$ is the output price of industry $t$ in $i$, and $1-\alpha$ is the share of non-robot capital in the production process.\footnote{AcemogluRestrepo2020jpe estimate a reduced-form version of equation (ref) for US commuting zones $i$ but their theoretical model is more general and can be applied across countries $i$. Note that our subscript $t$ here indexes industries, not time periods. An earlier working paper version of our paper MurisWacker2022 contains a textbook application from stockIntroductionEconometrics2015 that illustrates how ITE and CITE can differ in a standard panel setting with $t$ indexing time periods.} Those effects may operate into opposing directions and cause marginal effect heterogeneity. Heterogeneity across industries is thoroughly documented by Bekhtiaretal2024, while our application focuses on heterogeneity across countries.

We focus on two key candidates for effect heterogeneity across countries in the robot-employment relationship, which arise from the partial equilibrium model by AcemogluRestrepo2020jpe: income p.c. levels and demand effects. The demand effect is straightforward and reflected in the second right-hand side term of equation (ref). The importance of the income p.c. level for the robot-employment relationship arises from the fact that AcemogluRestrepo2020jpe assume the technology parameter $\theta$ to be homogeneous across $i$, which is plausible for US commuting zones but not across countries. Moreover, $\theta$ captures whether a task can technically be automated. Whether this is economically profitable depends on the cost savings from using robots ($\pi_i$ in AcemogluRestrepo2020jpe, who for simplicity assume $\pi_i>0 \ \forall i$). The cost-savings aspect unambiguously calls for a more negative effect of robotization on employment in higher-income countries: cost-saving potential $\pi_i$ positively depends on wage levels, which are higher in high-income countries.

The remainder of this section hence explores how income p.c. and its changes, reflecting demand effects, alter the robot-employment relationship. We will abstract from other factors that may give rise the cross-country effect heterogeneity in the robot-employment nexus for the sake of providing a traceable and instructive illustration of the performance of CITE and ITE in a relevant applications with plausible interaction term effects. Code documentation is available on the authors' GitHub repository (https://github.com/KMWacker/CITErobots).\\

Regression setup: ITE vs. CITE

In cross-country industry-level studies, the relationship between labor market outcomes $L$ and robots is typically estimated as a long-run first-difference equation (e.g., GraetzMichaels2018, deVriesetal2020, Bekhtiaretal2024):

equation[equation omitted — 125 chars of source]

The inclusion of country-specific fixed effects $a_i$ is convenient for isolating country-specific employment trends that are associated with robots. To capture the effect heterogeneity suggested above, we interact $\Delta Robots$ with the relevant interaction terms $H_1$ (income levels) and $H_2$ (demand changes) in an ITE settting:

equation[equation omitted — 180 chars of source]

where our (up to) $K=2$ interaction terms are country-specific income p.c. levels and demand changes. Recall that $\beta$ in this framework corresponds to $\kappa_0$ in the context of CITE.

The standard ITE approach consists in estimating equation (ref) through least squares (augmented with dummy variables $\alpha_i$). As thoroughly discussed in Section (ref), this requires strong exogeneity assumptions for consistent estimation of the interaction term coefficients $\kappa_1, \kappa_2$. In particular, we may be concerned about two cases that give rise to a correlation between the error term $u$ and the regressors.

One reason for concern is the behaviour of estimators in the presence of omitted interaction variables. Changes in demand $H_{2i}$ are likely correlated with income levels $H_{1i}$ and both may give rise to effect heterogeneity, leading to a bias in $\hat{\kappa}$ if one of them is omitted. We will hence assess to what extent it matters for ITE and CITE if $\kappa_1 H_1$ and $\kappa_2 H_2$ are jointly or individually included to gauge how both estimators behave in the presence of potentially omitted interaction variables.

A second reason for concern is more fundamental and concerns unobservable effect heterogeneity. For example, robots are frequently used for manual routine work in many industries, like health care. Patients, or clients more generally, in more technology-minded countries may be more open to being supported by robots. This gives rise to a higher worker substitutability, implying that the relationship between employment and robots is more negative in tech-minded countries. Tech-mindedness, which is not directly observable, will hence enter the error term $u$ in our ITE regression equation (ref) (through $X_{it}\epsilon_i$ in our general notation, see eq. (ref)-(ref)). Since tech-mindedness plausibly gives rise to higher robotization in the first place, the error term $u$ and $\Delta Robots$ will be positively correlated in equation (ref), which is a clear violation of the ITE exogeneity assumption (see Section (ref)).

Conversely, CITE explicitly models such country-specific effects. In the context of our application, CITE consists in first estimating the unit-specific parameters $\beta_i$ for each individual country $i$ through the outcome equation:

equation[equation omitted — 124 chars of source]

and subsequently exploring their relationship with country-specific income p.c. levels $H_1$ and demand changes $H_2$ by running the heterogeneity equation:

equation[equation omitted — 94 chars of source]

Both approaches, ITE and CITE, allow us to investigate if previous empirical studies have omitted an essential factor how robots influence labor market outcomes and facilitate a comparison of the performances of ITE and CITE.

It is important to recall that CITE absorbs all country-specific effect heterogeneity in the robot-employment relation in $\beta_{i}$, while ITE only allows for heterogeneity in the modeled $interation$ terms. Furthermore, note that equations (ref) - (ref) are long-run first-difference equations, which means that industry fixed effects are “differenced away” and that our subscript $t$ now indexes industries, not time periods, to illustrate the applicability of CITE in various panel setups. A standard textbook panel application (as well as a higher-dimensional panel application) can be found in an earlier working paper version of our paper MurisWacker2022.

Data

We use data from 15 manufacturing industries across 35 countries that is explained in full detail in DijkstraWacker2025. Robot stocks are calculated from IFR data (International Federabtion of Robotics) with a perpetual inventory method and divided by thousands of employees in the respective industry (from OECD TiM, see below). As suggested by GraetzMichaels2018, using raw or log changes in this robot density is not recommendable because of very low (or zero) starting values in the mid-2000s. We hence follow their recommendation to construct percentiles of (employment-weighted) changes in the robot stock.

For employment numbers, we rely on the 2023 edition of OECD's Trade in Employment (TiM) database (variable EMPN). Additionally, we merge the country-specific log of gross domestic product per capita, ln GDP p.c. (rgdpo/emp), from the Penn World Tables 10.0 feenstrapwt9, which captures income levels across countries and, in first differences, demand changes. Those are our interaction variables $H_1, H_2$, as suggested by the first two right-hand side terms of equation (ref).

For $\Delta \operatorname{ln} L$ and $\Delta Robots$, we construct 10-year changes comparing the 2014-2018 averages to 2004-2008 averages.\footnote{Comparing 10-year changes is standard in this literature. We construct 5-year averages at the start- and end-point to avoid year-specific variation in key variables (e.g., due to the global financial crisis in 2008).} For ln GDP p.c. we take country-specific averages over the 2004-2018 period. For changes in demand, we construct changes in ln GDP p.c. between the 2014-2018 and 2004-2008 averages (divided by 10 to obtain annual approximations). Further details on the data construction can be inferred from the code in the associated GitHub repository `CITErobots'.

table[table omitted — 734 chars of source]
figure[figure omitted — 160 chars of source]

Table (ref) reports the summary statistics of our sample, while Figure (ref) plots our key variables $\Delta \operatorname{ln} L$ and $\Delta Robots$ in a scatter that differentiates between lower- and higher-income countries. A number of features is worth highlighting. First, employment declines, on average (unweighted), with considerable heterogeneity: the standard deviation is more than double the mean. Second, since $\Delta Robots$ is constructed as percentile change, it ranges from 0 to 1 by construction, with a mean and IQR of 0.5. GDP p.c. is rather high, reflecting that most of our sample countries are high-income. Its rather low standard deviation reflects that this variable does not vary in the $t$ dimension of our panel (i.e., across industries of the same country). Figure (ref) separately illustrates countries at or above the ln GDP p.c. average (blue color) and those below average ln GDP p.c. (black). Third, there is some negative correlation between $\Delta Robots$ and ln GDP p.c., although this is not obviously visible from Figure (ref). Their correlation coefficient is -0.10 (not reported) and averages of $\Delta Robots$ are slightly higher for lower-income countries than for higher-income countries (0.54 vs. 0.47). Fourth, Figure (ref) suggests a positive descriptive correlation between $\Delta \operatorname{ln} L$ and $\Delta Robots$ in lower-income countries (dashed black line) and no correlation in higher-income countries (solid blue line), which our estimation will explore more thoroughly. Fifth, the sample is slightly unbalanced (with 509 observations), reflecting missing values in 1-4 industries across 9 sample countries. Finally, it is worth noticing that changes in demand are close to the historic average GDP growth rate of 2% per annum documented in PritchettSummers2014 and negatively correlated with ln GDP p.c. (-0.64, not reported), consistent with the idea of unconditional convergence pateletal2021.

Note that our third and fourth observation, taken together and at face value, may constitute a problem for ITE if an interaction of $\Delta Robots$ with ln GDP p.c. is omitted: $X$ ($\Delta Robots$) is plausibly correlated with the error term $\epsilon$ in the heterogeneity equation (ref). This violates the required Assumption (ref) for consistency of ITE and is likely on economic grounds: richer countries may exhibit certain features that simultaneously affect robotization and the robot-employment nexus.

Results

Table (ref) summarizes our regression results, with ITE in the upper panel A and CITE in the lower panel B. We start with a simple linear regression of $\Delta \operatorname{ln}L$ on $\Delta Robots$ as a baseline in column (1), which facilitates comparison to the literature. In our (manufacturing) sample, there seems to be an overall positive relationship between industries' robotization and employment changes.

In column (2) of Table (ref), we allow this relationship to vary with the interaction term for income levels (ln GDP p.c., $H_1$). In line with the above theoretical considerations (and already suggested by Figure (ref)), we find a negative interaction term coefficient $\hat{\kappa}_1$: the higher a country's income p.c. (and hence wage) level, the less favorable employment outcomes are associated with robotization. We will revert to the estimated marginal effects (which are graphically depicted in Figure (ref)) below but highlight a higher (absoulte) interaction term coefficient estimate for CITE, compared to ITE. Both estimators are relatively precise (interaction term standard errors $<1.7 \times |\hat{\kappa}_1|$), especially ITE.

table[table omitted — 2,216 chars of source]

In column (3) of Table (ref), we explore the hypthesis that heterogeneities in the robot-employment nexus are driven by demand changes $H_2$. Consistent with theory both point estimates of the interaction term coefficient suggest a more positive association between robotization and employment growth the more aggregate demand increases in a country. However, ITE would leave the researcher highly confident in this specification that the partial effect between robots and employment linearly depends on demand effects. The associated ITE p-value $< 0.01$ in column (3) suggests that the data are quite compatible with the specified interaction model, while CITE inference raises considerably more caution. Also notice the low R squared of the second stage CITE projection in column (3).

Column (4), which includes both interaction terms, illustrates that the demand channel identified in column (3) is most likely due to an omitted variable bias. Recall that demand changes and income p.c. levels are inversely correlated across countries (-0.64): richer countries experience lower increases in demand. The interaction term $\Delta Robots \times \Delta demand$ in column (3) hence captures all cross-country heterogeniety that is correlated with demand changes, including substantial heterogeneity due to income p.c. levels. It is, of course, highly likely that $\Delta demand$ (and ln GDP p.c.) are correlated with other sources of effect heterogeneity in the robot-employment nexus.\footnote{Likewise, it is plausible that our measure for demand changes is erroneous -- which substantiates our point that ITE inference in column (3) of panel A in Table (ref) is problematic.} As we know from our Monte Carlo inference simulations, ITE's rejection probabilities for t-tests of a null hypothesis about the interaction term coefficient can be highly distorted in this case. The stark ITE-related differences in p-values for such rejections when moving from column (3) to column (4) in panel A of Table (ref) should be a warning signal for applied researchers that inference of standard ITE may be highly susceptible and misleading.

Marginal effects

The results in Table (ref) suggest that cross-country heterogeneity in the robot-employment relationship is to a significant degree driven by differences in income per capita (and associated wage) levels. Column (2) of panel B suggests that 21% of cross-country heterogeneity can be explained by variation of ln GDP p.c. (see also Figure (ref)).

How much of a difference does the income level make for employment effects of robotization?\footnote{In this section, we treat the estimated relationship as an `effect', well-aware that other sources of endogeneity may pose challenges to a causal interpretation.} And how would that estimated marginal effect vary between ITE and CITE? Figure (ref) is instructive for answering this question. Let us consider income levels of ln GDP p.c. = (10.5, 11.4, 11.7). In our sample, the former corresponds to Bulgaria, which is at the lower end of new European Union member states and a plausible level for various middle-upper income countries (e.g., between Colombia and Chile). The second corresponds to Germany, the latter to the United States. Recall that an overall regression with no interaction effects would suggest a positive relationship equal to 0.13 in either case (red line in Figure (ref) and column (1) in panel A of Table (ref)).

figure[figure omitted — 151 chars of source]

Considering a difference in robotization rates of $\Delta Robots = 0.5$ and estimates from column (2) of Table (ref), ITE suggests an associated employment increase of 25% over a decade, for a country at a ln GDP level of 10.5.\footnote{Recall that $\Delta Robots$ is constructed as percentile changes, ranging from 0 to 1. The IQR of $\Delta Robots = 0.5$ hence compares the 25th percentile of robot adoption to the 75th percentile of robot adoption. Estimated marginal effect: $0.5 \cdot (\hat{\kappa}_0 + \hat{\kappa}_1 \operatorname{ln} GDP p.c.)$.} The magnitude is much larger with CITE, 60%, and Figure (ref) suggests that it more plausibly captures the estimated country-specific coefficients for various upper-middle income countries, while ITE seems heavily impacted by the low country-specific estimate of China, the lowest-income country in our sample.

For a ln GDP p.c. level of 11.4 (approximately Germany), the estimated marginal effects of CITE (9.1%) and ITE (4.6%) are both positive and relatively close to the magnitude without interaction effect (6% $ = 0.5 \cdot 0.13$, cf. column (1) panel A of Table (ref)). For a ln GDP p.c. level of 11.7 (approximately US), the estimated marginal effects are negative, and stronger for CITE (-7.9%) than for ITE (-2.2%). Considering that CITE is consistent under milder assumptions than ITE, our results suggest that ITE and homogeneous estimation without interaction terms both underestimate the negative effects of robotization for manufacturing employment in countries at the highest income levels in the sample (ln GDP p.c. $>$ 11.57; Ireland, Norway, Switzerland, United States).

If income (and associated wage) levels are indeed a main source of robot-employment effect heterogeneity, this may reconcile conflicting country level evidence in the literature. For example, AcemogluRestrepo2020jpe, find negative consequences of robot exposure on regional labor markets in the US, while Dauthetal2021 find no such aggregate employment effects for Germany, despite using a similar approach. This is quite consistent with our quantitative estimates about effect heterogeneity due to income differences across countries (especially when considering differences in sample coverage and methodologies).

CITE, of course, provides country-specific estimates of the robot-employment relationship, as they are illustrated in Figure (ref). The reason why we do not consider those country-specific estimates for our marginal effect calculations is that they suffer from incidental parameter bias. The purpose of our paper is to estimate interaction terms in exactly such a short-$T$ panel, where country-specific estimation is no option, and we have shown that CITE offers a consistent projection of those (individually biased) panel unit estimates on the interaction variable of interest. We hence consider marginal effects of a random country at an income level like, e.g., Bulgaria, rather than a Bulgaria-specific coefficient. Applied analysts working on specific countries may be willing to accept the incidental parameter bias in country-specific estimates, but this is not the purpose of our paper.

Application takeaways

An important insight of our analysis for the literature on employment effects of robotization is that economic displacement effects of robots are plausibly stronger in countries at higher income levels, where wages are usually higher. Econometrically, our application suggests that this income-interaction effect is underestimated by standard ITE and that ITE may lead to premature rejection of a null hypothesis about an interaction term in the presence of misspecification (e.g., when another relevant interaction term is omitted).

Our application results, and particularly Figure (ref), also highlight why CITE is attractive for applied econometricians. Because unit-specific coefficients are estimated in a first step and then further analyzed in a second step (through projection on the relevant interaction variables $H$), this offers a lot of flexibility and diagnostics tools in the second step.\footnote{For example, one may speculate to what extent Bulgaria, Colombia, China, and Chile are outliers. Indeed, those countries surpass a critical value for `Cook's distance' and when they are excluded from the analysis, we obtain ITE and CITE interaction coefficients for ln GDP p.c. of -0.58 and -0.59 (significant at the 1 and 10% level), respectively, which are very close to each other and fall between the point estimates reported in column (2) of Table (ref).}

Conclusion

Our paper has highlighted that the most common approach to estimate interaction effects with panel data neglects residual heterogeneity in the effect of $X$ on $Y$ across panel units. If this unobserved effect heterogeneity is systematically correlated with explanatory variables in the model, which is plausible in most econometric applications, this standard approach will not provide consistent estimates of the interaction effects.

We advocate for explicit modeling of effect heterogeneity across panel units with a correlated interaction term estimator (CITE) that requires less demanding exogeneity assumptions for consistency. Those findings are irrespective of whether interaction variables are constant within panels or not.

Our argument in favor of CITE resembles the case for additive fixed effects in linear panel data models. While additive fixed effects explicitly capture unobserved heterogeneity in unit-specific levels of $Y$, CITE explicitly models the effect heterogeneity $\beta_i$ that most common panel interaction estimates do not account for.

Our paper has focused on the comparison of ITE and CITE for the estimation of interaction effects in panel data. Two alternative approaches that are sometimes used in empirical practice are (i) subgroup analysis -- in which one conducts the analysis conditional on every value of the interaction variable (e.g. chettyEffectsExposureBetter2016a) -- and (ii) saturating the heterogeneity model and then running ITE (e.g., norrisEffectsParentalSibling2021). Either approach may bridge the gap between ITE and CITE. Although such approaches may work well when the interaction variables are discrete, they are not feasible or practical for continuous interaction variables. It would be interesting to extend our analysis to include a comparison with these alternative approaches.

An obvious extension of our research is to derive a test statistic for consistency of ITE when compared to CITE. Until such a test becomes available, we strongly recommend the use of CITE or at least reporting CITE results next to ITE results. Given the simple implementation of CITE, this should become standard in applied econometrics when estimating interaction terms in panel data.\footnote{If one is interested in interactions with varying variables $G$, it is sufficient to additionally include an interaction of the $X$ variable(s) with panel dummy variables. If one is interested in interactions with $H$ variables that are constant within panels, those panel specific estimates from the first step need to be saved and regressed on the $H$ variables in a second step. An adjustment of standard errors is necessary to account for first-step estimation.}

An additional advantage of CITE, beyond its less restrictive exogeneity assumptions, is that the unit-specific effect heterogeneity captured in $\beta_i$ can provide further insights about sources of heterogeneity in panel data. Recent advances in machine learning appear particularly promising to further explore those first-step CITE estimates of unit-specific effects.