EconBase
← Back to paper

Influence Analysis with Panel Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

43,252 characters · 11 sections · 38 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Influence Analysis with Panel Data

abstractThe presence of units with extreme values in the dependent and/or independent variables (i.e., vertical outliers, leveraged data) has the potential to severely bias regression coefficients and/or standard errors. This is common with short panel data because the researcher cannot advocate asymptotic theory. Example include cross-country studies, skill-cell analyses, and experimental studies. Available diagnostic tools may fail to properly detect these anomalies, because they are not designed for panel data. In this paper, we formalise statistical measures for panel data models with fixed effects to quantify the degree of leverage and outlyingness of units, and the joint and conditional influences of pairs of units. We first develop a method to visually detect anomalous units in a panel data set, and identify their type. Second, we investigate the effect of these units on LS estimates, and on other units’ influence on the estimated parameters. To illustrate and validate the proposed method, we use a synthetic data set contaminated with different types of anomalous units. We also provide an empirical example. \\ JEL codes: C13, C15, C23.\\ Keywords: Leverage, outliers, diagnostic measures, Cook's distance.

Introduction

Observational data often contain data points that exhibit extreme values in the response-space (vertical outliers, VO), in the covariate-space (good leverage points, GL), or in both directions (bad leverage points, BL) rousseeuw1990,silva2001. The presence of these anomalies creates bias in the least square (LS) estimates (i.e., coefficients and/or standard errors), and can potentially invalidate the results because econometric techniques based on the mean are more sensitive to extreme realisations donald1993,bramati2007,verardi2009. Their impact is more severe with panel data because these anomalies may be carried over the time series of a unit, exacerbating the bias bramati2007. This is particularly problematic in short panel data sets (few cross-sectional units $N$ are observed over multiple time periods $T$) because the small cross-sectional size does not allow to mobilise the benefits of the asymptotic theory. These data are common in some fields of economics due to the structure of the data or the nature of the study. Some examples are: worldwide country-level macro-data (e.g., 50 United States, 27 European countries); skill-cell data sets (e.g., with education-experience-state cells as unit of observation); or experimental data, where it is costly for the researchers to recruit many participants to follow in successive waves.

Popular tools for the detection of such data points include diagnostic plots (e.g., leverage-vs-residual plots), and measures of influence (e.g., \citeapos{cook1979} distance). These respectively serve to identify the type of anomaly, and assess the individual influence on the LS estimates by removing one unit at a time. Despite the ease of their interpretability and computation, these measures may fail to detect anomalous units when surrounded by other cases as they do not assess mutual influence of units. This phenomenon -- known as the masking effect -- is documented for Cook-like measures atkinson1985, chatterjee1988,rousseeuw1990,rousseeuw1991, but can be overcome by measures that delete two cases at a time from the sample, such as, joint and conditional measures for the influence lawrance1995. A visual inspection of these anomalies may be complicated with many variables rousseeuw1991,bramati2007, thus the need of a method that identifies these anomalies based on statistical measures.

In this paper, we develop a method to: (i) systematically detect anomalous units in a panel data set, and identify their type with leverage-vs-residual plots; and (ii) investigate the effect of these units on LS estimates, and on other units’ influence on the estimated parameters with influence plots. For this purpose, we first formalise statistical measures -- i.e., individual leverage and normalised residual squared -- that quantify the degree of leverage and outlyingness of units all over its time series. These are then used to produce `unit-wise' leverage-vs-residual plots. Second, we build on \citeapos{lawrance1995} pair-wise approach by proposing measures for joint and conditional influence of units $i$ and $j$ that are designed for panel data models with fixed effects. We illustrate the method with synthetic data that have been contaminated with anomalous units. We then apply our method to real data.

We show that our method detects problematic units that will not otherwise be detected by the conventional methods. Specifically, leverage-vs-residual plots provide an initial picture of the anomalies in the data, being more informative about the existence and type of anomaly. Then, influence plots complete the analysis by displaying the strength and direction of the connection between units $i$ and $j$, showing how the influential units affect the contribution of other units on the LS estimates. The strength of this method is that a unit, which is not individually influential according to Cook’s distance, will be detected if it is influential either jointly with another unit, or in the absence of another highly influential unit. The method can help the researcher understand how some units may be driving most of the effect on LS estimates after the main regression analysis, or explore the data set and identify potential anomalies before any regression analysis.

Once a unit is detected as potentially influential, the researcher should not proceed with its deletion because it might be a valid observation that should be properly addressed. For example, robust estimation accounts for anomalies that affect the estimated coefficients (i.e., VO and BL) bramati2007,verardi2009,aquaro2013,aquaro2014,Jiao2022, and jackknife-type standard errors deal with anomalies that affect the statistical inference (i.e., GL) mackinnon1985,chesher1987,davidson1993,mackinnon2013,belotti2020,polselli2022,mackinnon2023cluster,mackinnon2023.\footnote{These methods will not be discussed further in this study.}

This paper contributes to the literature on diagnostic and influential measures as follows. While Cook-like distances are still on of the most popular diagnostic measures for the detection of individually influential observations,\footnote{Several variants of \citeapos{cook1979} measure have been proposed for both cross-sectional and panel data models. In the cross-sectional framework, martin2009 extend the measure to logistic models, pinho2015 to generalized linear mixed models, and martin2015 to log-linear models. banerjee1997 are the first to formalise a leave-one-out diagnostic measure to detect outliers and influential points for panel data, specifically in linear longitudinal models with random effects. belotti2020 have extended \citeapos{banerjee1997} distance to linear panel data models with fixed effects.} authors have not yet addressed how to overcome the masking effect. We build on \citeapos{lawrance1995} pair-wise measures to investigate the reciprocal influence of units. We present an adaptation of the joint and conditional influence measures for linear panel data models with fixed effects. With this measures, we can detect units that would have not been detected otherwise by Cook's distance because not individually influential. Because the joint and conditional measures resemble a directed and weighted adjacency list from network analysis,\footnote{In a directed graph, the effect of unit $i$ on unit $j$ differs from $j$ to $i$; in a weighted graph, the intensity of the effect of each unit differs.} we propose a graphical representation of these measures that plots the influence of unit $j$ on $i$'s influence such as to show the existence and the strength of the connections between pairs of units. Using scatter plots or heat plots facilitates the visual detection and classification of anomalous data points, and their influence on other units. This representation is innovative in the literature on diagnostic measures, where \citeapos{atkinson1993} stalactite plot pioneers the visual inspection of multivariate outliers in the data.

This paper is closely related to mackinnon2023cluster,mackinnon2023. They derive a measure for the overall leverage of clusters for cross-sectional data. We follow their derivation to calculate the overall leverage of a unit all over its history. Because we have repeated measures of the variables for each unit, our formulation addresses the presence of fixed effects by transforming the variables with the within-group (WG) transformation.

The rest of the paper is structured as follows. In Section (ref), we start by defining the three types of anomalies. In Section (ref), we first introduce the panel data models and estimators, and then present the measures of influence for panel data. Section (ref) discusses the proposed method using a synthetic data set. Section (ref) applies the method to an empirical exercise. Section (ref) concludes.

Anomalous Units

Longitudinal data are extremely likely to contain units with values of the dependent and/or independent variables that follow a different pattern from the main cloud of the data in the response- and/or covariate-spaces. These are classifiable as vertical outliers (VO) if the extreme values are in the outcome variable, good leverage (GL) points if the extreme values are in the covariates, or bad leverage (BL) points if they are in both directions.

Vertical outliers (unit 30 in Figure (ref)) are anomalous in the response-factor space, follow the opposite trend of the main cloud of data points, and lay far from the regression line as if they were generated from a different process chatterjee1986, greene2012. They display large squared normalised residual, and their presence alters the estimated LS intercept (here, the fitted line shifts upwards), undermining the accuracy of the estimator but not its precision verardi2009.

Leveraged data points appear isolated from the rest of the data but, unlike VO, follow the same trend of the rest of the data chatterjee1986, greene2012. Leverage points can be distinguished in `good' or `bad' as follows. The former (unit 20 in Figure (ref)) exhibit unusually extreme values in the covariates and lie on the predicted regression line enhancing the precision of the regression fit; whereas, the latter (unit 10 in Figure (ref)) possess extreme values in both input and response direction, lie far from the plane where the bulk of data points are and are not fitted by the regression model rousseeuw1991,bramati2007. While BL realisations adversely affect the estimated LS coefficients (both the intercept and the slope of the regression), GL points add variability in the sample allowing for a better fit of the data at the cost of a deteriorated statistical inference mackinnon1985,chesher1987,silva2001,verardi2009.

In a panel data sets, anomalous realisations in the time series of a unit may appear either as isolated cases in the time-series of different units (cell-concentrated points), or concentrated in the time-series of the same units (block-concentrated points) bramati2007.

Measures of Influence in Panel Data Models

Let $\{y_{it},\mathbf{x}_{it}\}_{t=1}^T$ be repeated measurements of unit $i \in \mathcal{I}=\{1,\dots,N\}$ over $t$ time periods, and that $T_i=T$ for all $i$ (i.e., balanced panel).\footnote{The data is assumed to be balanced to simplify the notation, but it can be equally applied to unbalanced panel data with appropriate change in notation.} Consider a linear panel regression model with unobserved individual heterogeneity as follows

equation[equation omitted — 91 chars of source]

where $y_{it}$ is the response variable for the cross-sectional unit $i$ at time period $t$; $\mathbf{x}_{it}$ is a $k\times1$ vector of time-varying inputs, $\bm{\beta}$ is a $k\times1$ vector of parameters of interest; $\alpha_i$ is the individual-specific unobserved heterogeneity (or fixed effects); and $u_{it}$ is a stochastic error component.

Stacking observations over time $t$, model (ref) becomes

equation[equation omitted — 147 chars of source]

where $\mathbf{y}_i$ is a $T\times1$ vector of the dependent variable; $\mathbf{X}_i$ is a $T\times k$ matrix of time-varying independent variables; $\bm{\alpha}_i=\alpha_i\bm{\iota}$ is a $T\times1$ vector of unobserved individual fixed effects, and $\bm{\iota}$ is a vector of ones of order $T$; and $\mathbf{u}_i$ is a $T\times1$ vector of one-way error component.

Model (ref) can be consistently estimated using OLS on the time-demeaned model

equation[equation omitted — 167 chars of source]

where $\widetilde{\mathbf{y}}_i=(\mathbf{I}_T-\, T^{-1}\bm{\iota}\bm{\iota}')\mathbf{y}_i$ is $T\times1$; $\widetilde{\mathbf{X}}_i=(\mathbf{I}_T-\, T^{-1}\bm{\iota}\bm{\iota}')\mathbf{X}_i$ is $T\times k$; and $\widetilde{\mathbf{u}}_i=(\mathbf{I}_T-\, T^{-1}\bm{\iota}\bm{\iota}')\mathbf{u}_i$ is $T\times1$. Note that $(\mathbf{I}_T-\, T^{-1}\bm{\iota}\bm{\iota}')\bm{\alpha}_i=\bm{0}$ as $T^{-1} \bm{\iota}\bm{\iota}' \bm{\alpha}_i=\bm{\alpha}_i$. The OLS estimator of (ref) is the well-known within-group estimator of the true population parameter with formula

equation[equation omitted — 203 chars of source]

The estimator is unbiased and asymptotically normally distributed under conventional model assumptions.

The within-group estimator (ref) without the whole history of unit $i$ is the Leave-One-Out (L1O) estimator\footnote{The L1O estimator for panel data models is derived in banerjee1997 for mixed effects models, and belotti2020 for fixed effects models. We show the asymptotic distribution of $\widehat{\bm{\beta}}_{(i)}$ in Appendix (ref).} as follows

equation[equation omitted — 215 chars of source]

where $\mathbf{M}_i^{-1}=(\mathbf{I}_T-\mathbf{H}_{i})^{-1}$ with $\mathbf{H}_{i}=\widetilde{\mathbf{X}}_i (\widetilde{\mathbf{X}}'\widetilde{\mathbf{X}})^{-1} \widetilde{\mathbf{X}}_i'$, and $\widehat{\mathbf{u}}_i =\widetilde{\mathbf{y}}_i - \widetilde{\mathbf{X}}_i\widehat{\bm{\beta}}$ is the residual term. The L1O estimator estimates the effect of the covariates on the outcome variable by removing one unit at a time from the sample. This is a general measure of the influence exerted by that unit on the estimated coefficients or on the model fit.

The within-group estimator (ref) without the full history of pairs of units $\{i,j\}$ is the Leave-Two-Out (L2O) estimator\footnote{The derivation of Formula (ref) and its distribution in Appendices (ref) and (ref).} below

equation[equation omitted — 425 chars of source]

where $\mathbf{M}_j = \mathbf{I}_j-\mathbf{H}_j$ with $\mathbf{H}_{ij} = \widetilde{\mathbf{X}}_i (\widetilde{\mathbf{X}}'\widetilde{\mathbf{X}})^{-1} \widetilde{\mathbf{X}}_j' $, and $ \mathbf{H}_j = \widetilde{\mathbf{X}}_j (\widetilde{\mathbf{X}}'\widetilde{\mathbf{X}})^{-1} \widetilde{\mathbf{X}}_j' $ noting that $\mathbf{H}_{ji} = \mathbf{H}_{ij}'$. The deletion of pairs of units one at a time from the sample captures the influence exerted by that pair on the estimated parameters or on the model fit.

Leverage and Normalised Residual Squared

The change in the magnitude of the estimated coefficients or the model fit after the deletion of a unit from the sample is informative of the influence of that unit. This can be explained by high residuals, high individual leverage, or both. In this section, we construct measures for panel data that assess the individual influence of a unit along these two dimensions. We define the individual leverage and the normalised residual squared for panel data, that considers the full time series of a unit (`unit-wise' approach).

The leverage of a unit measures the distance of the values of the covariates of a unit from those of other units in the sample. This information is provided by the leverage matrix of unit $i$, $\mathbf{H}_{i} =\widetilde{\mathbf{X}}_i \bigl(\widetilde{\mathbf{X}}'\widetilde{\mathbf{X}}\bigr)^{-1}\widetilde{\mathbf{X}}_i'$, which is a $T\times T$ positive-definite and symmetric with diagonal elements $h_{ii,tt} = \widetilde{\mathbf{x}}_{it}'(\widetilde{\mathbf{X}}'\widetilde{\mathbf{X}})^{-1}\widetilde{\mathbf{x}}_{it}$, and off-diagonal elements $h_{ii,ts} = \widetilde{\mathbf{x}}_{it}'(\widetilde{\mathbf{X}}'\widetilde{\mathbf{X}})^{-1}\widetilde{\mathbf{x}}_{is}$.\footnote{With only one time period (cross-sectional data), the leverage matrix of unit $i$ is the scalar $h_i = \mathbf{x}_i'(\widetilde{\mathbf{X}}'\widetilde{\mathbf{X}})^{-1}\mathbf{x}_i$ with $\mathbf{x}_i$ a $k\times1$ vector of covariates and the constant term.}

We calculate the nuclear (or trace) norm of $\mathbf{H}_{i}$ to summarise the leverage exerted by each unit over their time series, as in mackinnon2023cluster,mackinnon2023 for clustered data.\footnote{Similarly, mackinnon2023 calculate the leverage matrix of cluster $G$ populated with $N_g$ units. They use the nuclear norm operator to provide an overall assessment of the influence of each group in a way that is computationally feasible. We face a similar challenge with panel data, where the unit can be seen as the $N$-th cluster populated with a total of $T$ time realisations. Therefore, their notion of leverage for cluster-level data is comparable to some extent to ours for panel data with some adjustments (i.e., the WG transformation to remove the fixed effects).} The overall leverage of unit $i$ can be calculated as $L_i = \mathrm{tr}\big(\widetilde{\mathbf{X}}_i'\widetilde{\mathbf{X}}_i \bigl(\widetilde{\mathbf{X}}'\widetilde{\mathbf{X}}\bigr)^{-1}\big)$. The average measure of the leverage exerted by all units in the sample is $ \frac{1}{N}\sum_{i=1}^N \mathrm{tr}(\mathbf{H}_{i})= k/N$ by rearranging terms in the trace operator, noting that $\sum_{i=1}^N\widetilde{\mathbf{X}}_i'\widetilde{\mathbf{X}}_i=\widetilde{\mathbf{X}}'\widetilde{\mathbf{X}}$, and $k=\mathrm{rank}(\widetilde{\mathbf{X}})$. This result (similar to the cross-sectional case) can be used to establish the cut-off value for a unit to be considered as highly influential. For instance, a unit with an individual leverage that exceeds twice the value of the individual leverage in the sample, $L_i>2k/N$, is a high leverage unit. High leverage units possess unusually extreme values (too small or too large) in the regressors, and can be classified as GL unit.

We now define a measure to locate anomalous units in the $y$-space. We define the normalised residual squared for panel data as $\widehat{u}^*_{it} = \big(\widehat{u}_{it}/\sqrt{\sum_i\widehat{u}^2_{it}}\big)^2$, which informs on the outlyingness of unit $i$ at time $t$. We define $\widehat{\mathbf{u}}^*_i = [\widehat{u}^*_{i1}\, \dots\, \widehat{u}^*_{iT} ]'$ a $T\times1$ vector that contains the information of the influence exerted by unit $i$ in each time period $t\in\{1,\dots,T\}$ on the least squares estimates.

We use the Euclidean norm of $\widehat{\mathbf{u}}^*_i$ to quantify the influence of unit $i$ in the output space such that $O_i= \sqrt{\sum_t \widehat{u}^{*2}_{it}}$. The average normalised residual squared over the full sample of units is $N^{-1}\sum_{i=1}^N\widehat{u}^*_{it}= 1/N$ (like in the cross-sectional case). This value can be used as a reference to compare the levels of influence of each units in terms of the dependent variables. That is, a unit is said to be influential in the $y$-space if its value of normalised residual squared exceeds twice the average value in the sample, i.e., $O_i> 2/N$.

Measures for Joint and Conditional Influence

Joint influence is the influence exerted by a pair $(i,j)$ on the LS estimates jointly. It is obtained by comparing the estimated coefficients with and without full history of the pair in the sample. In formulae,

equation[equation omitted — 305 chars of source]

where $N$ is the total number of cross-sectional units in the sample; $\nu_1=K$ is the total number of regressors including the constant, and $\nu_2 = (N-1)$ is the degrees of freedom at the denominator because of the clustering; $s^2=\nu_2^{-1}\sum_{i=1}^N\widehat{\mathbf{u}}_i'\widehat{\mathbf{u}}_i$ is the variance of the fitted model -- i.e., residual mean squared error (MSE) -- and is consistent for $\sigma^2$. $\mathrm{C}_{ij}(\widehat{\bm{\beta}})$ is symmetric for $i\ne j$. When $i=j$, the measure informs on the individual influence of unit $i$ on the estimated coefficients for $\beta$. This can be interpreted as the Cook's distance for panel data banerjee1997,belotti2020. Formula (ref) becomes

equation[equation omitted — 294 chars of source]

$\mathrm{C}_{ii}(\widehat{\bm{\beta}})$ is hence a special case of the more general measure for the deletion of pairs of units, $\mathrm{C}_{ij}(\widehat{\bm{\beta}})$.

The joint effect is the ratio

equation[equation omitted — 122 chars of source]

and informs on unit $i$'s influence within the $(i,j)$ pair. For large values of this measure, unit $j$ alters the individual effect of unit $i$ on the estimated coefficients. In other words, the most influential unit in the pair $j$ contributes to drive most of the effect on LS estimates by enhancing or reducing the effect of the least influential $i$. The assessment can be done in conjunction with the conditional effect below.

Because joint measures do not compare individual influences arising before and after the deletion of another unit, the notion of conditional influence is needed. Conditional influence is the influence exerted by unit $i$ on the LS coefficients conditional on removing unit $j$ from the sample. It shows how the absence of unit $j$ alters the influence of unit $i$ on the LS estimates. In formulae,

equation[equation omitted — 338 chars of source]

where $\widetilde{\mathbf{X}}_{i(j)}$ is a $(N-1)\times K$ matrix of time-demeaned regressors without unit $j$. The value of $\mathrm{C}_{i(j)}(\widehat{\bm{\beta}})$ is zero for $i=j$, and is not symmetric for $i\ne j$. The diagnostic measure $\mathrm{C}_{i(j)}(\widehat{\bm{\beta}})$ is not a test statistics but the knowledge of its empirical distribution can be used to extrapolate cut-off values to assess the conditional influence of units.

The conditional effect is the ratio

equation[equation omitted — 125 chars of source]

quantifying unit $i$'s influence before and after the deletion of unit $j$ from the sample. For $\mathrm{M}_{i(j)}\ge1$, unit $j$ is said to mask the influence of unit $i$ because the influence of $i$ increases without $j$ in the sample. When $\mathrm{M}_{i(j)}<1$, unit $j$ is said to boost the influence of unit $i$ because the influence of $i$ decreases without $j$. With this information, we can interpret the results of the joint effect as follows. When the influence of $i$ is masked (boosted) by $j$, $j$ is reducing (enhancing) the effect of $i$.

These measures follow a F-distribution with $\nu_1$ degrees of freedom at the numerator and $\nu_2$ degrees of freedom at the denominator.\footnote{The derivation is shown in Appendix (ref).} These two measures are not test statistics but the information on their empirical distributions can be used to select distributional cut-off values to identify influential units, as recommended practice in the cross-sectional framework cook1979,martin2009,martin2015. Cut-off values can be extrapolated from the median of the inverse cumulative F distribution with $\nu_1$ degrees of freedom at the numerator and $\nu_2$ degrees of freedom at the denominator. Alternatively, non-distributional cut-off values of 1 or $4/N$ can be chosen bollen1985.

The joint and conditional measures are constructed in a way that resembles a weighted and directed adjacency list from network analysis, which displays the existence and the strength of the links between pairs of units $(i,j)$. In a directed graph, the effect of unit $i$ on unit $j$ differs from $j$ to $i$; in a weighted graph, the intensity of the effect of each unit is different. This allows us to mobilise graphical tools from network analysis that plot unit $j$'s influence on $i$'s influence by means of two-way scatter plots, or heat plots.

The Method

In this section, we present the method to systematically detect anomalous units with panel data. Motivated by the fact that visual inspection of the data becomes complicated with more than two covariates rousseeuw1991,bramati2007, we propose a method based on statistical measures for panel data that facilitates the detection and classification of anomalous units with many covariates. The method consists of two steps: first, the detection of anomalous units by type; and, second, the influence analysis to determine their effect on the LS coefficients, and on other unit's the influence on the LS parameters.

This method is mainly designed for short panels because the bias arising from the presence of anomalous units in the sample is higher as the benefits of the asymptotic theory cannot be advocated. It is also thought to identify anomalies within clusters. The method can be employed before the regression analysis to identify possibly problematic units, or after to understand what units drive most of the average effect in the LS estimates.

It goes beyond the scope of this paper to examine how anomalous units should be treated once they are detected with the proposed method. The deletion of any anomalous unit from the sample is not always the best option because the econometric literature offers techniques to handle these anomalies properly. A strand of the literature focuses on developing robust estimators to consistently estimate the parameter of interest by handling VO and BL points bramati2007,verardi2009,aquaro2013,aquaro2014,Jiao2022. Another focuses on the downward bias of the statistical inference in the presence of GL points, recommending the use of jackknife-type standard errors mackinnon1985,chesher1987,davidson1993,mackinnon2013,belotti2020,polselli2022,mackinnon2023.

Detection of Anomalous Units

The first step consists in detecting the presence of anomalous units, and identifying their type from their location in the leverage-residual space. A graph that plots the individual leverage over the normalised residual squared (i.e., leverage-versus-residual plot) informs on the presence and the type of anomalous unit in the sample based on their location in the plane.

Figure (ref) plots the individual leverage over the normalised squared residual. The solid red lines mark the average values for individual leverage and normalised residuals squared. Cut-off values for high leverage and high residuals can be arbitrarily set as twice the average values. With this respect, plausible GL units possess high leverage but low squared residual and are located in the top-left quadrant; possible VO have high residual squared but low leverage and are located in the bottom-right quadrant; possible BL units display both high leverage and residual squared are in the top-right quadrant. Non-influential units are grouped in the cloud of points in the bottom-left quadrant or around the intersection of the two lines as they display low leverage and residual.

We illustrate the method to detect anomalous units in a leverage-residual plot with synthetic data that has been previously contaminated with anomalous units: unit 10 is a BL, unit 20 a GL, and unit 30 a VO. From the graph, the type of three units is correctly detected based on their leverage and residuals, and easily identifiable from their location on the plane. The rest of the units are not influential, as per data generating process, being confined to the bottom-left quadrant.

Influence Analysis

Once anomalous units are detected and classified with the leverage-vs-residuals plot, the second part of the influence analysis shows how anomalies affect the LS estimates, and other units' influence on the estimated coefficients in terms of joint and conditional influence. Because the output resembles a weighted and directed adjacency matrix, we can mobilise graphical tools from network analysis (e.g., heat plots and scatter plots) that plot the measure of pair-wise influence of unit $j$ on unit $i$. In this example, we visualise the links and strength between pairs of units using heat plots.

Figure (ref) shows four plots with values of the joint influence, joint effects, conditional influence, and conditional effects (respectively from top to bottom, left to right). The colour scale from dark blue to red shows the degree of influence/effect from the smallest to the largest value. The influence analysis consists of four steps as follows.

1. The joint influence plot displays the joint influence of pairs of units when $i\ne j$, and the individual influence when $i\ne j$. The diagonal elements represent Cook's distance for panel data. Units 10 and 20 -- generated as BL and GL units, respectively -- have high individual influence, largely exceeding the distributional cut-off 0.694 and the conventional Cook's Distance cutoff of 1. These units reciprocally have high joint influence in correspondence of another anomaly of the same type (e.g., with unit 30 as being a VO) and also with the rest of or almost all the units.

2. The joint effect plot shows joint effects $\mathrm{K}_{j|i}$ of pairs of units. The values of this measure are extremely high for pairs of units ${i,j}$ equal to $\{17, 10\}$, $\{17, 20\}$. The $j$-th units that exert most of the influence in the pair are the BL and GL units that was already detected with the leverage-vs-residual plot and with the joint influence plot whereas the $i$-th unit was not. Detected units $j=\{10, 20\}$ alter the individual effect of units $i=17$ by either reducing or enhancing its effect in the LS estimation. The direction of the effect (upwards or downwards) can be inferred from the outcome of the conditional effect in step 4.

3. The conditional influence plot highlights the influence of unit $i$ conditional on removing $j$ from the sample. The conditional influence is quite high for already individually influential units (i.e., 10, and 20), by construction; these are the BL and GL units. This measures are informative for the construction of the conditional effects.

4. The conditional effect plot plots the measure $\mathrm{M}_{i(j)}$ for the pair $\{i,j\}$. Units $j=\{10, 20\}$ are masking the influence of unit $i=17$ as $\mathrm{M}_{i(j)}\gg1$; no unit is boosting the effect of any other. This information is useful to conclude that the $j$-th units detected in step 2 are reducing the effect of unit 17.

Empirical example

In this section, we illustrate an application of the method for the detection of anomalous units in panel data sets. We use the data and the regression equation corresponding to Column (2a) of Table 4 in berka2018.\footnote{With this exercise, we do not intent to support or invalidate the original results.} The sample consists of 9 countries observed over the time period 1995--2007. The illustrative example is relevant to our case because of the small cross-sectional sample size, and the ease of interpretability of the regression output.

berka2018 study the relationship between real exchange rate and sectoral productivity in nine Eurozone countries finding a strong correlation between productivity and real exchange rates among high-income countries with floating nominal exchange rates. The estimating regression equation is as follows

equation[equation omitted — 110 chars of source]

where $RER_{it} $ is the log real exchange rate (expenditure-weighted) expressed as EU15 average relative to country $i$ (an increase is a depreciation) in period $t$; $TFP_{it}$ log of TFP level of traded relative to nontraded sector in EU12 relative to country i at time t; $\mathbf{x}_{it}$ are other covariates; $\alpha_i$ is the country fixed effects; and $u_{it}$ is the error term. The coefficients in Equation (ref) are estimated using OLS after the within-group transformation.

We are interested in identifying anomalous countries that may drive most of the effect on the LS estimate of $TFP$ based on Equation (ref). The leverage-vs-residual plot (Figure (ref)) highlights a group of six countries with individual leverage above the mean value -- where Italy may be a potential leverage units -- and only one country (Ireland) with the potential to be a BL or VO.

Figure (ref) shows the joint and conditional influence of pairs of countries. The labels refer to 1-Austria, 2-Belgium, 3-Finland, 4-France, 5-Germany, 6-Ireland, 7-Italy, 8-Netherlands, 9-Spain. On the main diagonal, individually influential countries correspond to 6 (Ireland) and 7 (Italy). These exert high joint influence (red and pink squares) with other non-influential countries -- i.e., Ireland on Netherlands (unit 8), and Italy on Austria (unit 1) and Belgium (unit 2). Looking at the joint effects, units 1, 3, 6, 7, and 8 alter the effect of unit 9 (Spain) on the LS estimates. The direction of the effect (upwards or downwards) can be inferred from the output of the conditional effect. Countries 6, 7, and 8 seems to be masking the effect of country 9 whereas countries 1 and 3 are boosting its contribution to the $TFP$ coefficient. Therefore, we can conclude by looking again at the joint effects plot that the former group of countries is reducing the effect of Spain while the latter is enhancing its influence.

Conclusion

The presence of outliers and leveraged data points has the potential to bias regression coefficients and standard errors, which necessitates appropriate methods to detect and identify them. Available measures -- such as, leverage-vs-residual plots and Cook-like measures -- may fail to correctly detect and classify them in panel data set. The method we propose overcomes this limitation by taking into account the panel structure of the data. We formalised the average leverage and average normalised residuals in a panel data setting to produce unit-wise leverage-residual plots. Then, we developed two diagnostic measures for panel data models -- based on \citeapos{lawrance1995} cross-sectional measures -- and showed their statistical distributions.

Overall, the method detected those units that exert more influence in the data set by taking into account the full history of a unit. The analysis of the joint and conditional effects showed how their presence alters the influence of other units in the sample and, hence, how their presence is driving the average effect on the LS estimates. The strength of this method is that a unit that is not individually influential according to Cook’s distance will always be detected if turns up to be influential jointly with or conditional on another highly influential unit. This will help the researcher understand how anomalous units drive their LS estimates. As a novelty, the individual and bilateral (joint and conditional) influence exerted by units in the sample are displayed with a network representation (e.g., heat plots and scatter plots).

The method can be used before the actual regression analysis to investigate the presence of anomalies in the sample, or after the regression analysis to understand what units are driving the estimated coefficients or standard errors.

Once anomalous units are properly detected and identified, the researcher should deal with their presence following the recommendations in the econometric literature. For example, the literature recommends the use of robust estimators of the median (e.g., M-estimators, S-estimators, etc.) in the presence of outliers bramati2007,verardi2009,aquaro2013,aquaro2014,Jiao2022, whereas heteroskedasticity-consistent standard errors based on jackknife methods in the presence of leveraged data hinkley1977,mackinnon1985,chesher1987,davidson1993,mackinnon2013,belotti2020,polselli2022,mackinnon2023cluster,mackinnon2023.

spacing{1}

Figures

figure[figure omitted — 834 chars of source]
figure[figure omitted — 841 chars of source]
figure[figure omitted — 698 chars of source]
figure[figure omitted — 475 chars of source]
figure[figure omitted — 654 chars of source]