EconBase
← Back to paper

Spatial Differencing for Sample Selection Models with Unobserved Heterogeneity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

67,179 characters · 8 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Spatial Differencing for Sample Selection Models with Unobserved Heterogeneity

\abstract{This paper derives identification, estimation and inference results using spatial differencing in sample selection models with unobserved heterogeneity. We show that under the assumption of smooth changes across space of the unobserved sub-location specific heterogeneities and inverse Mills ratio, key parameters of a sample selection model are identified. The smoothness of the sub-location specific heterogeneities implies a correlation in the outcomes. We assume that the correlation is restricted within a location or cluster and derive asymptotic results showing that as the number of independent clusters increases, the estimators are consistent and asymptotically normal. We also propose a formula for standard error estimation. A Monte-Carlo experiment illustrates the small sample properties of our estimator. The application of our procedure to estimate the determinants of the municipality tax rate in Finland shows the importance of accounting for unobserved heterogeneity. \\

Keywords: Sample selection, Spatial difference, Unobserved heterogeneity. }

Introduction

In linear models, spatial differencing has been used to deal with unobserved omitted variables. The availability of geographical locations allowed empirical papers to take advantage of the spatial dimension of the data and control for various unobserved heterogeneity (e.g. duranton2011assessing, black1999better or holmes1996effects). In general, spatial differencing offers an identification strategy in the situations when researchers face cross-sectional data with unobserved heterogeneity and lack suitable instrumental variables. This paper extends spatial differencing to a model with sample selection.

For economists, the question of omitted variables is a serious concern in the context of nonexperimental data. The solution is straightforward when omitted variables are simply the result of not including all relevant variables for which data exist - we add such variables to the model to avoid the bias induced by their omission. When omitted variables are unobserved, researchers have basically three options: they can use (i) proxies, (ii) instrumental variables, or (iii) differencing the data across time or space.

The proxies reduce the bias if they manage to capture the effect of the omitted variables such that what remains is uncorrelated with the error term. However, it is often the case that the proxies are imperfect, hence they may still be related to the unobserved heterogeneity, or the error term if it turns out to be endogenous, or it could also be irrelevant after controlling for observed covariates. In such cases, the inclusion of the proxy will not solve the bias problem, it may even exacerbate it.\footnote{ See todd2003specification for details on the use of proxy and oster2019unobservable for a rigorous treatment of the evaluation of the robustness to omitted variables.} The second solution - using a valid set of the instruments - may help alleviate the bias. However, as discussed in todd2003specification, the “quasi-experimental" local average treatment effect (LATE) obtained in the instrumental variable model may not correspond to the ceteris paribus and thus, may not correspond to the deep structural parameter of interest. Lastly, panel data sets allow researchers to control for the unobserved heterogeneity. They indeed help them to identify the causal effect when, for example, time constant unobserved heterogeneity might cause endogeneity problem and strong instruments satisfying exclusion restrictions cannot be found. However, there might be situations when such data sets are not available.

Our paper is a contribution to the literature identifying and estimating model parameters in the presence of unobserved omitted variables. We propose an identification strategy based on spatial differencing. As was discussed above, this approach has been used in the context of linear regressions. However, little is known about its performance in non-linear models. We extend spatial differencing into this direction, specifically to the case of cross-section data with sample-selection. We show that under justifiable assumptions on the smoothness of the unobserved heterogeneity (i.e. spatially close individuals have similar unobserved heterogeneity and the derivative of their Inverse Mill's ratio are similar), spatial differencing eliminates the unobserved effects even in the presence of a nonlinear element - in our case Mill's ratio. The parameters of interest of our sample selection model are estimated using a standard two-step approach of heckman1974shadow heckman1979sample. We derive asymptotic properties and propose a correction of standard errors accounting for the two-step nature of our estimation and spatial differencing. The asymptotic behavior of the estimator reveals important properties of spatial differencing that researchers would need to be cautious about. The new estimator and the standard errors correction are easy to implement.

The intuitions for the model of sample selection with unobserved spatial heterogeneity that we consider in this paper can be described as follows. Suppose we have a cross-sectional data on municipalities which are organized into larger geographical units called regions, and which have the authority to set the levels of local taxation. Municipality tax rates must be at least as high as the threshold set by a central government. As a result, municipalities self-select into those with the tax rates at the threshold and those above it. We are interested in what determines the municipalities' tax rates. The tax rate will depend on various socio-economic characteristics, such as age composition of population and income, but also on amenities. These can depend on the region where the municipalities are located: for example, regions with natural landscapes might have different level and composition of amenities than regions without them. We can control for them with region-specific dummies. However there can be a considerable unobserved heterogeneity at municipality level. Controlling for that with municipality-specific dummies might not be an option, since we may quickly run out of degrees of freedom. Therefore, we face a problem of a self-selected cross-sectional sample with unobserved heterogeneity which we cannot fully control with dummies, and which has two spatial dimensions: high-level which we call locations (in our example regions), and low-level which we call sub-locations (in our example municipalities).

Spatial differencing will eliminate the sub-location specific unobserved heterogeneity. It will, at the same time, also induce a correlation in the error terms. We take that correlation into account, and derive the asymptotic properties of our estimator using similar arguments to those present in the derivation of the asymptotic behavior of the clustered standard errors: the number of locations goes to infinity and the size of location is assumed random and bounded almost surely. We find that this result also extends to a linear model without sample selection. This has important implications that researchers need to be cautious about. Indeed, the consistency of estimator applied to the spatially differenced data requires (i) a large number of locations, (ii) a limited number of individuals in each location. Monte Carlo simulations also suggest that it would be better if the number of individuals in sub-locations were small as well. Before we continue, let's notice that locations in our model are equivalent to clusters and we use 'location' and 'cluster' interchangeably.

Since our estimator is derived for a clustered sample with unobserved heterogeneity, this paper contributes to the literature on the selection correction in panel data. In this literature, the main challenge is the presence of individual-specific unobserved heterogeneity in both the outcome and the selection equations. The existing solutions are based on either a full model specification or on a differencing procedure. wooldridge1995selection uses a Mundlak approach to specify the individual-specific unobserved heterogeneity in both equations. He also imposes a special functional form to the selection mechanism. kyriazidou1997estimation, on the other hand, does not impose strong restriction of the selection equation functional form and uses a nonparametric approach to difference out the unobserved fixed-effect. rochina1999new similarly relies on differencing to identify the parameters of the model, but she also imposes additional distributional assumptions to the selection equation. Even if our problem has similarities with the selection correction in panel data model literature, the main difference is that we are observing a clustered cross-section. In each cluster, there is a finer common sub-location specific unobserved heterogeneity shared by some individuals in that cluster. This heterogeneity, however, is different from the cluster and individual-specific ones studied in panel data models and implies a different cluster asymptotic. Since the outcome of the individuals are not independent in our model, while it is in the panel data case, our asymptotic results are thus derived using a large number of cluster asymptotic with heterogenous, random and bounded cluster size.

The clustered dependence created by the finer sub-location specific unobserved heterogeneity relates our asymptotic discussion to the papers dealing with clustering at variance level (see wooldridge2010econometric for a textbook treatment). The asymptotic in that literature is derived using either a large or a fixed number of clusters. The fixed number of cluster leads to non-normal asymptotic and discussion about recent contributions can be found in hansen2019asymptotic. A large number of cluster asymptotic was first derived by white1984asymptotic and has been investigated by several authors allowing either fixed cluster size or heterogeneous cluster. Recent developments include hansen2019asymptotic who propose conditions on the relation between the cluster sample sizes and the full sample in a regular asymptotic, or djogbenou2019asymptotic who derive asymptotic with varying cluster sizes and carry out a cluster wild bootstrap. Our results complement this literature by extending the cluster asymptotic to a sample selection model.

We present an empirical application of our new estimator. We examine the determinants of tax rates across four hundred and eleven Finish municipalities spread across nineteen Finish regions. In 1999, Finish central government has decided to raise the lower bound of the tax rate municipalities could set from 0.2% to 0.5%. This has created a sample selection mechanism which resulted with more than half of the municipalities opting for 0.5% while the rest charging higher tax rate. We use our spatial differencing estimator to control for unobserved municipalities effect which can be correlated with the error term, creating thus an endogeneity problem and rendering the standard sample selection estimator biased and inconsistent. Our results clearly show the presence of unobserved heterogeneity across municipalities, limitations of using only region dummies (nineteen in our case) to fully control for municipalities' unobserved heterogeneity, and the importance of spatial differencing to control for it.

The structure of the paper is as follows. First, we expand the spatial differing method in the case of linear regression model to the case of sample selection. Then we discuss identification assumptions, propose an estimation procedure, and derive the estimator of the corrected standard errors. Lastly, we conduct Monte Carlo simulations and present an empirical application of our estimator.

Sample Selection Models with Spatial Correlation

In many economic applications, we are interested in estimating the following regression equation:

equation[equation omitted — 112 chars of source]

where $x_{ij}^{\prime }$ is a vector of exogenous controls variables, $\gamma _{j}$ is location fixed effect, $\gamma _{j\alpha}$ is a sub-location specific effect for sub-location $\alpha$ which is at a finer spatial scale than location $j$, and $\varepsilon _{ij}$ is the error term.\footnote{The sub-location specific component, $\gamma_{j\alpha}$, is a simplification for $ \gamma_{j\alpha_i}$. We are implicitly assuming that the sub-location specific effects are the same for all its individuals.}

Examples of application of this model can be found in the estimation of the fertilizer effect on wheat crops yield in farms growing multiple crops, or the effect of local taxation on the growth of firms. The crop-yields depend on the soil quality of location $\gamma _{j}$ (e.g. a village), but also on the sub-location specific soil composition (e.g. a farm in the village) collins2006soil. Similarly, the impact of local taxation on the growth of firms may vary by county but also sub-locations such as neighborhoods as in duranton2011assessing. We can control for $\gamma _{j}$ with location dummy variables. However, they might not be enough to capture all unobserved heterogeneity related to location $j$ as there can be considerable heterogeneity at finer spatial scale of sub-locations: using the example above, the firms are located in various neighborhoods $\alpha$ which are sub-locations of location $j$. Furthermore, standard location fixed effect $\gamma _{j}$ relies upon an arbitrary specification of the comparison neighborhood group, as pointed out by gibbons2003valuing, making it an imperfect control for sub-location specific effect $\gamma _{j\alpha}$. If $\gamma _{j\alpha}$ is correlated with $x_{ij}$, OLS estimate of $\delta $ will be biased. In the absence of suitable instrumental variables for $x_{ij}$, the spatial differencing offers a solution by differencing out the unobserved sub-location specific effects $ \gamma _{j\alpha}$.

duranton2011assessing, black1999better or holmes1996effects use spatial differencing in the case of linear models to solve endogeneity problems arising from unobserved sub-location effect $\gamma _{j\alpha}$. They take advantage of the fact that for sufficiently small distances between sub-locations, their specific effect $\gamma _{j\alpha}$ changes smoothly across space, allowing thus to difference them out. This corresponds to the following assumption.\\

{Assumption I1: The sub-location specific unobservable effect is homogenous in the a neighborhood of the individual ie $\Delta _{d}\gamma _{j\alpha}=0$ for $ d $ small enough.\newline }

In several economic models, in addition to the sub-location specific fixed effects $\gamma _{j\alpha}$ , the outcome of interest is not observed for the selected sub-sample. The selection can be the result of the decision of individuals or the researcher. The presence of sample selection introduces nonlinearity to the model ((ref)).

We specify the model with sample selection as follows. Consider two latent dependent variables $y^*_{1ij}$, and $y^*_{2ij}$ in a cross-section which follow a regular linear model for individual $i$ in a location $j$:

$y_{1ij}^{\ast }=z_{ij}^{\prime }\beta +\theta _{j\alpha}+\theta _{j}+\varepsilon _{1ij}$ - selection equation,

$y_{2ij}^{\ast }=x_{ij}^{\prime }\delta +\gamma _{j\alpha}+\gamma _{j}+\varepsilon _{2ij}$ - \ outcome equation.

Individual error terms are $\varepsilon _{1ij}$ and $\varepsilon _{2ij}$ ; $\theta _{j\alpha}$ and $\gamma _{j\alpha}$ are sub-location specific effects for a sub-location $\alpha$ in location $j$, affecting the selection and the outcome equation respectively. The exogenous characteristics $x_{ij}$ affect the outcome. They could be correlated with $\gamma _{j\alpha}+\gamma _{j}$ but not with $\varepsilon _{1ij}$ and $\varepsilon _{2ij}$. The variables $z_{ij}$ are exogenous variables determining selection, they can be a subset of $x_{ij}$. However, for identification purposes, some elements of $z_{ij}$ are assumed to be absent from $x_{ij}$.

{Assumption I2: $\varepsilon _{1ij}$ and $\varepsilon _{2ij}$ are independent identically distributed normal random variables for all $i,j$.\\} The outcome is modelled in the form of a truncated sample selection model and is represented by equation ((ref)).

equation[equation omitted — 202 chars of source]

Let us consider the following conditions.

enumerate• Condition 1: $Cov[z_{ij}, \theta_{j\alpha}+\theta_j+\varepsilon_{1ij}]=0$; $z_{ij}$ is exogenous • Condition 2: $Cov[x_{ij}, \gamma_{j\alpha}+\gamma_j+\varepsilon_{2ij}]=0$; $x_{ij}$ is exogenous • Condition 3: and errors $(\varepsilon_{1ij}, \varepsilon_{2ij})$ satisfy $ \varepsilon_{2ij}=\rho \times \varepsilon_{1ij} +v_{ij} $ with $ \varepsilon_{1ij} \sim \mathcal{N}(0, 1)$ and independent of $v_{ij}$.

It is possible to consistently estimate $\delta$ by Tobit regression under these three conditions.\footnote{ Identification required an exclusion restriction ie a variable that affects $ y^*_{1ij}$ but not $y^*_{2ij}$. Otherwise, identification relies on the nonlinearity of the inverse Mills ratio.} In most applications, the Condition 1 and 2 are unlikely to hold because there is a possibility that, within a location, there could be a sub-location specific omitted variable affecting both the outcome and some observed characteristics of interest. Thus, it is possible that $Cov[z_{ij}, \theta_{j\alpha}+\theta_j]\neq0$ and $Cov[x_{ij}, \gamma_{j\alpha}+\gamma_j]\neq 0$. The standard way to deal with the correlation between $x_{ij}$ and $\gamma _{j\alpha}$ would be to find a suitable instrument for the $x_{ij}$ and run a IV Tobit or IV two-stage Heckit.

The very local nature of the sub-location specific effect means that it is not always evident to find a variable correlated with $x_{ij}$ and uncorrelated with $\gamma_{j\alpha}$. The exclusion restriction is likely to be violated and IV two-stage Heckit will yield inconsistent estimates for $\delta $. Another option is to use the finer location fixed effect and estimate the model using classic Heckman two-stage procedure, but this will in practice lead to a proliferation of variable and lose of degrees of freedom.

Identification via Spatial Differencing

This section investigates the application of this spatial differencing technique to the case of cross-section sample selection models. We denote $\Delta _{d}$ to be a spatial difference operator. One example is a pair-wise difference operator which takes the difference between each observation and another observation located at distance less than $d$ from that observation. In a location $j$, with individual $i$ and $k$ who are neighbours. The pair-wise differencing of the variable $A$ is: $$\Delta_d A=A_{ij}-A_{kj}.$$

Another example is the difference between the individual outcome and the average outcome of his/her neighbourhood $\mathcal{N}_{id}$. This operator is similar to the neighbourhood fixed effect operator, the difference being that the neighbourhoods can overlap. We call this operator the fixed-effect difference operator. Let $\mathcal{N}_{id}=\{ k,\hspace{0.2cm} in \hspace{0.2cm} neighbourhood \hspace{0.2cm}d \}$, the sample size of $\mathcal{N}_{id}$ is $N_d$, the differencing is given by: $$\Delta_{df} A=A_{ij}-\frac{1}{N_d}\sum_{k\in \mathcal{N}_{id}}A_{kj}.$$

A further possibility is to use a kernel as in kyriazidou1997estimation to weight neighbour in $\mathcal{N}_{id}$ according to how far they are, in term of observable characteristics. This operator is the kernel difference operator. $$\Delta_{dK} A=A_{ij}-\sum_{k\in \mathcal{N}_{id}}\psi(i,k) A_{kj}.$$ Where $\psi(i,k)=\frac{1}{h_{N_d}}K\left(\frac{(z_{ij}^{\prime }-z_{kj}^{\prime})\beta + (x_{ij}^{\prime }-x_{kj}^{\prime})\delta}{h_{N_d}}\right)$, $K$ is a kernel density function while $h_{N_d}$, is a sequence of bandwidths. To illustrate our identification strategy and for the asymptotic derivation, we use the pairwise spatial difference operator, while for the empirical application and for the Monte Carlo simulations, the fixed effect difference is used.

For the spatial difference operator $\Delta _{d}$, $\Delta _{d}y_{2ij}=y_{2ij}-y_{2kj}$ with $k$ an observation in the neighborhood $d$ of $i$. Let $\xi_{ij}\equiv\{ x_{ij}, z_{ij},y_{1ij}^{\ast }>0,\gamma _{id},\theta _{id} \}$ with $\gamma_{id}=\{ \gamma_{kj} \hspace{0.2cm} with \hspace{0.2cm} k \in \mathcal{N}_{id}\cup \{i\}\}$ and $\theta_{id}=\{ \theta_{kj} \hspace{0.2cm} with \hspace{0.2cm} k \in \mathcal{N}_{id}\cup \{i\}\}.$

eqnarray[eqnarray omitted — 540 chars of source]

where $\lambda (c)=\phi (c)/\Phi (c)$ is the inverse Mill's ratio while $\phi (c)$ and $\Phi (c)$ are respectively the density and distribution function of a normal random variable with mean zero and variance 1.

To go from Equation (3) to Equation (4) we use the linearity of expectation and the mean independence of $y_{2ij}$ and $ y_{1kj}^{\ast }$ conditional on $\xi_{ij}$, as well as the mean independence of $y_{2kj}$ and $y_{1ij}^{\ast }$ conditional on $\xi_{kj}$, since we have assumed in Assumption I2 that $\varepsilon _{1ij}$ and $\varepsilon _{2ij}$ are $iid$. The separation of the conditional set, $\xi_{ij}$ and $\xi_{kj}$, is possible because we are working with cross-sectional data. Such separation of the conditional set is not possible for panel data. Indeed, in the context of panel data with individual effects and sample selection, when the differencing is used to remove the fixed-effects, the conditional set cannot be separated as we have done to move from Equation (3) to (4). For example, kyriazidou1997estimation has to impose a “conditional exchangeability" assumption that is conditioned on the variable related to the two periods used in differencing. In case of models with censoring, lee2001first discusses conditions under which first-difference can be applied, and applies the linear implication of the "conditional exchangeability" assumption. In a similar context using first difference, rochina1999new imposes a joint normality between the difference in the error of the outcome equation and the error in the selections equation in the two time periods.\footnote{See dustmann2007selection for a review on selection correction in panel data models.}

Estimating equation (6) presents two challenges for the identification of the parameter of interest $\delta$ and the sample selection parameter $\rho$: the sub-location specific difference $\Delta _{d}\gamma _{j\alpha}$, and the sample selection term $\rho \Delta _{d}\lambda (z_{ij}^{\prime }\beta +\theta _{j\alpha}+\theta _{j})$. As for the sub-location specific difference $\Delta _{d}\gamma _{j\alpha}$, under Assumption I1 and I2, equation (6) becomes

equation[equation omitted — 187 chars of source]

These assumptions allow us to difference-out the sub-location specific unobserved effect $\gamma _{j\alpha}$, a strategy that was applied by duranton2011assessing.

As for the sample selection term $\rho \Delta _{d}\lambda (z_{ij}^{\prime }\beta +\theta _{j\alpha}+\theta _{j})$, we see that it depends on the unobservable sub-location specific and location effects $ \theta _{j\alpha}+\theta _{j}$. Because that sample selection term is a nonlinear function, a simple spatial differencing will not always work unlike the case of $\gamma _{ja}$. Therefore, the following assumption helps us to deal with this challenge:

Assumption I3:\\ (i) {The sub-location specific unobservable selection effect is homogeneous in a neighborhood of the individual i.e. $\Delta _{d}\theta _{ja}=0$ for $d$ small enough.}\newline

(ii) The changes in the inverse Mill's-Ratio in a neighborhood of the individual i.e.

equation[equation omitted — 359 chars of source]

for $i$ and $k$ in a neighborhood $d$ small enough, $\theta _{j\alpha_i}+\theta _{j}$ and $\theta _{j\alpha_k}+\theta _{j}$ both different from 0, $\lambda ^{\prime }(.)$ is the first derivative the inverse Mill's ratio, $c_i,$ and $c_k$ are, respectively, in the intervals formed by $[z_{ij}^{\prime }\beta,\hspace{0.2cm} z_{ij}^{\prime }\beta +\theta _{j\alpha_i}+\theta _{j}]$ and $[z_{kj}^{\prime }\beta,\hspace{0.2cm} z_{kj}^{\prime }\beta +\theta _{j\alpha_k}+\theta _{j}]$ such that Equation ((ref)) holds.

Assumption I3 (i) is similar to assumption I1. It seems plausible that if that assumption holds for the outcome equation, it will hold true for the selection equation as well.

Assumption I3 (ii) is novel and one of the contributions of this paper. It assumes that if the exact Taylor approximation is applied on the individual inverse Mill's ratio for individuals $i$ and $k$ in the location $j$, the intermediate points $c_i$ and $c_k$ should be similar. If the level of nonlinearity of $\lambda (.)$ is low, then the assumption will also hold. In the extreme case of local linearity of the inverse Mill's ratio, the Assumption 3 (ii) perfectly holds.

The combination of assumptions I3 (i) and I3 (ii) implies that $$\lambda (z_{ij}^{\prime }\beta +\theta _{j\alpha_i}+\theta _{j})-\lambda (z_{ij}^{\prime }\beta )=\lambda (z_{kj}^{\prime }\beta +\theta _{j\alpha_k}+\theta _{j})-\lambda (z_{kj}^{\prime }\beta ).$$ Thus, $\Delta_d\lambda (z_{ij}^{\prime }\beta)=\Delta _{d}\lambda (z_{ij}^{\prime }\beta +\theta _{j\alpha}+\theta _{j})$

{

thmLet us consider the sample selection model presented in Equation (ref). Under assumptions I1 to I3 the parameters $\delta$ and $\rho$ are identified.

}

Proof of Theorem 1\\ We have already shown that under the assumptions I1 and I2, we can obtain Equation ((ref)). Applying the assumption I3, to Equation ((ref)) leads to the following equation

equation[equation omitted — 156 chars of source]

Thus, assumptions I1 to I3 are sufficient for the identification of $\delta $ and $\rho$.\\

We have derived the results using the pairwise spatial difference operator. However, the identification result holds for other spatial difference operators as well. In the case of the average or kernel difference operator, the conditioning in equation (ref) is on $\xi_{kj}$ with $k \in \mathcal{N}_{id}$ for the average difference operator and $k$ is in the full sample for the kernel operator. Note that under assumptions I1 and I3, any difference of the weighted average in a neighborhood of the individual will enable us to remove the sub-location specific effect. The conditional expectation presented in Equation ((ref)) depends on exogenous observable variables and parameters of interest.

Estimation and Asymptotic Properties

In this section, we present an estimation procedure and derive asymptotic properties of the proposed estimator. The estimation procedure involves two-steps. In the first step, probit model is estimated and the inverse Mill's ratio predicted. In the second step, a spatial difference operator differences out both location and the sub-location specific unobserved heterogeneity. The model is then estimated using an ordinary least square estimator. When we have a sample of $N$ individuals, the estimation procedure is thus as follows:

itemize• Step 1: Estimate $\beta$ by probit with location effect $\gamma _{j}$; and calculate $\hat{\lambda}_i =\lambda (z_{ij}^{\prime }\hat{\beta})$. • Step 2: Estimate $\delta$ and $\rho$ in the OLS regression \begin{equation} \Delta_d y_{2ij}=\Delta_dx^{\prime }_{ij}\delta + \rho \Delta _d\lambda (z_{ij}^{\prime }\hat{\beta}) +w_{ikj}. \end{equation}

Since we used spatial differencing and $\lambda (z_{ij}^{\prime }\hat{ \beta})$ is estimated in the first step, a particular structure of the variance-covariance matrix emerges. Therefore, we also need to derive the correct estimator of standard errors $w_{ikj}$ which we will do in section 2.3.

We will now show that the estimator obtained by the above procedure is consistent and asymptotically normal. To derive the asymptotic properties we use similar arguments as those used to derive the asymptotic properties of the clustered standard errors. Specifically, the population size of each location is assumed random and bounded almost surely, and the law of large numbers is applied by letting the number of locations (clusters in case of clustered standard errors) go to infinity.

We consider a generic matrix of spatial difference $\Delta $. The matrix form notation of equation ((ref)) can be expressed in as \footnote{The variables without subscript represent vector or matrices of all observation in the sample.}

equation[equation omitted — 128 chars of source]

where $\eta _{ij}$ are the same error as in standard sample selection models.\footnote{We assume the notation that $\lambda(z^{\prime }\hat{\beta})$ is a vector with typical element $\lambda(z_{ij}^{\prime }\hat{\beta})$. } Let us denote $\theta =(\delta ,\rho )^{\prime }$ and $ W=[x^{\prime },\lambda (z^{\prime }\hat{\beta})]$. The simplified estimation Equation ((ref)) is $$ \Delta y_{2}=\Delta W\theta +\Delta \eta $$ and OLS estimator of $\theta $ is

equation[equation omitted — 114 chars of source]

The spatial nature of data implies that an observation $k$ with $n$ neighbours may appear in several pairs. This induces correlation in the error term $ \Delta \eta $ for all $n$ of these pairs because of the spatial differencing in the second step of the estimation procedure. As a result, a particular structure of the covariance matrix emerges, and we need to take that into account when calculating the standard errors.

To proceed further, we need to introduce assumptions under which the asymptotic properties of our estimator are derived. \\ Assumption E1: The sample is formed if $N$ individuals from the population.\\ $(i)$ We observed $\{ x_{ij}, z_{ij}\}$ independent and identically distributed random variable with $i=1,...,N$ and $j=1,...,J$. \\ $(ii)$ The number of individuals in a location $j$, $N_j$, is exogenous, random, identically distributed with $N_j<n_0$ almost surely and $E(N_j)<\infty$; where $n_0$ is a scalar.\\ $(iii)$ The outcomes and the latent variables are independent across location i.e. $j_1 \neq j_2$ the variables $y_{2ij_1}\bot y_{2ij_2}$ and $y^*_{1ij_1}\bot y^*_{1ij_2}$.

An implication of assumption E1 $(i)$ in conjunction with assumption I2 is that $\theta_j$ and $\gamma_j$ are $iid$. However, $within$ a location $j$, there is a certain level of correlation among individuals which operates through $\theta_{j\alpha_i}$ or $\gamma_{j\alpha_i}$. This means that our assumptions restrict how that within-location individual correlations occurs.

Assumption E1 $(ii)$ restricts the location size to be bounded and implies that the number of locations has to grow to achieve a large sample size in our asymptotic calculation. This assumption is similar to those held in the literature of cluster samples asymptotic and it leads to a “large number of cluster" asymptotic theory similar to the one discussed in wooldridge2010econometric, who assumes fixed cluster size. This assumption corresponds to a specific case of the Assumption 1 in hansen2019asymptotic, who allow for different cluster size ranging from fixed to infinite. We have, however, derived the asymptotic of our estimator under the more restrictive condition of Assumption E1 ($ii$). The reason is that it can be proven that under a joint asymptotic ($N,J \rightarrow \infty$), Assumption 1 is equivalent to assuming that the size of the sample in each location is bounded. If we instead allow for a sequential asymptotic where the number of locations is fixed and the sample size goes to infinity, then there exists at least one location with an infinite number of individuals and the inequality used in the proof of Hansen and Lee (2009)'s Theorem 1 becomes invalid.

To better illustrate our argument, let us consider the location sample size proposed by hansen2019asymptotic: $N_j=N^{\alpha}$ with $0\leq \alpha <1$; we can prove that $1-\alpha = \frac{ln(J)}{ln(N)}$. If we allow for a joint asymptotic, $\alpha$ is not define. If on the contrary we assume that the number of locations $J$ is fixed, then, $\alpha$ goes to $1$. In both cases, relaying on Hansen and Lee (2009)'s Assumption 1 seems not enough to warrant the desire asymptotic regularities.

Assumption E2: $z'$ and $W$ are full rank column, with each element having up to its $4^{th}$ moment.\\

{

thmWe consider the sample selection model presented in Equation (ref). Under assumptions I1 to I3, E1 and, E2.\\ $(i)$ $\hat{\theta} \rightarrow^p \theta$ as $N\rightarrow \infty$\\ $(ii)$ $\sqrt{N} (\hat{\theta}- \theta)\rightarrow^d \mathcal{N}(0, \Theta)$ with $\Theta= C \Gamma C'$\\ where $C^{-1}=E((\Delta W_{ij})'\Delta W_{ij}),$ $\Gamma=\rho^2 E[(\Delta W_{ij})'\Omega_{ij} \Delta W_{ij}]+E[(\Delta W_{ij})'\Delta e_{ij}\Delta e_{ij}(\Delta W_{ij})],$\\ and $\Omega_{ij}=[\lambda'(z_{ij}^{\prime }\beta)]^2 z_{ij}^{\prime}V_{\beta} z_{ij}$ taking $V_{\beta}$ as the first step probit variance-covariance matrix.

}

Proof of Theorem 2: In appendix.\\

It is important to notice that the same type of asymptotics should be used in a linear model. In this respect, we complement duranton2011assessing who propose a correction for the standard errors, but do not discuss the asymptotic properties of their estimators. Similarly, black1999better and holmes1996effects use spatial differencing, but do not account for the fact that differencing will lead to a correlation between pairs where an individual is present. Our asymptotic derivations do account for the presence of correlation between pairs, and are valid not only for a model with but also without sample selection (in our model, the absence of selection implies $\rho=0$). They also have important practical implications: the consistency of the estimator requires a large number of locations $\gamma _{j}$, and a small number of individuals in each sub-location $\gamma _{j\alpha}$

Estimator of Variance

This section derives a procedure estimate the variance-covariance of the estimator in Equation ((ref)) which has a particular structure arising from $(i)$ spatial differencing and $(ii)$ Heckman's two-step estimation procedure.

We consider $B= \left[ (\Delta W)^{\prime }\Delta W\right]^{-1}$ and $\Sigma=Var[(\Delta W)^{\prime }\Delta\eta]$ such that the conditional variance-covariance matrix of $\hat{\theta}$ is

equation*[equation* omitted — 54 chars of source]

Note that

equation*[equation* omitted — 71 chars of source]

This means that we need a consistent estimator of $Var(\Delta \eta )$ to compute correct standard error for $\hat{\theta}$.

Let us consider that $ Var(\Delta \eta )=V_{1}+V_{2}$ with

eqnarray*[eqnarray* omitted — 105 chars of source]

where $R$ a diagonal matrix of dimension $N$ (total number of observations), with $d_{ij}=1-\lambda (z_{ij}^{\prime }\beta )[z_{ij}^{\prime }\beta +\lambda (z_{ij}^{\prime }\beta )]$ as the diagonal elements.

eqnarray*[eqnarray* omitted — 204 chars of source]

where $D$ is the square, diagonal matrix of dimension $N$ with $1-d_{ij}$ as the diagonal elements; $z$ is the data matrix of selection equation; and $ V_{p}$ is the variance-covariance estimate from the probit estimation of the selection equation.

{

thmWe consider the sample selection model presented in Equation (ref). Under assumptions I1 to I3, E1 and, E2. The variance-covariance estimator of the $\hat{\theta}$ is given by \begin{equation} V_{twostep}=B(\Delta W)^{\prime }[\hat{V}_{1}+\hat{V}_{2}](\Delta W)B^{\prime } \end{equation} where $\hat{V}_1=\hat{\rho ^{2}}\Delta \hat{R} \Delta^{\prime }$ and $\hat{ V}_2=\hat{\rho}^2 \Delta \hat{D} z\hat{V}_{\beta}z^{\prime }\hat{D}\Delta^{\prime } $ with all unknown parameters replaced by their estimates. Moreover, this is a consistent estimator $Var(\hat{\theta})$.

}

Proof of Theorem 3:\\ The result holds by construction.\\

Monte Carlo Simulation

In this section we present the results of Monte Carlo simulations to (i) to describe the behavior of the estimator proposed in this paper and (ii) offer empirical guidance for applied research. Regarding the latter, we will pay a close attention to the implication of assumption E1 $(ii)$ according to which it is important to have a large number of locations relative to the number of individuals in the sub-locations. Monte Carlo experiments will offer empirical guidance as to when the number of locations is large enough.

The estimator developed in this paper is referred to as the “Sub-location Differencing" and it accounts for sub-location specific effect $\gamma _{j\alpha}$. To highlight its features, we compare it to two other estimators. One ignores the presence of both $\gamma _{j}$ and $\gamma _{j\alpha}$ and applies a simple two step estimator with no spatial differencing - we call it “No-Differencing" estimator. The other accounts only for the location fixed effects $\gamma _{j}$ and we call it “Location Differencing" estimator. For each estimator, the mean bias and the coverage rate for the 95% confidence level test are reported in tables (ref) to (ref).

The data is obtained using the following data generating process. We assume that there are $J=20, 30, 100$ non-overlapping locations, each location is divided into $s= 2, 4, 8$ sub-locations. There are $n_j=3, 5, 8, 10$ individuals sharing the same sub-location. The latent variables are $y^*_{1ij} = z_{ij}\beta + \theta_{ijs}+\theta_j+\varepsilon_{1ij}$ and $y^*_{2ij} = x_{ij}\delta + \gamma_{ijs}+\gamma_j+\varepsilon_{2ij}$, where $\theta_{ija}=10^{-5}j\times s $ and $\gamma_{ijs}=5j\times s$ is the sub-location specific effect, while $\theta_j=10^{-5} j$ and $\gamma_j=10j$ are the location effects; for all $i$ and $j$, $x_{ij} \sim \mathcal{N}(0,1)$ , $z_{ij} \sim U(0,1)$ each drawn independently; $\delta=1$, $\beta=0.2$. The error terms in both equations for all $i$ and $j$ are generated as follows: $\varepsilon_{1ij} \sim \mathcal{N}(0,1)$, $\varepsilon_{2ij}= \rho \varepsilon_{1ij}+v_{ij}$ where $v_{ij} \sim \mathcal{N}(0,1)$ is independent of $\varepsilon_{1ij}$ and $\rho=0.7$.

table[table omitted — 3,211 chars of source]
table[table omitted — 3,218 chars of source]
table[table omitted — 3,714 chars of source]

We summarize the main results of the simulations in the four points below but in general, the “Sub-location Differencing" estimator has the smallest mean bias and delivers a conservative coverage rate.\footnote{There is room for improvement concerning our inference strategy. Cluster robust inference is part of a large and growing literature and our work gives some insight as to how diffrencing can be used in cross-sectional data. Future work will investigate the importance of heteroscedasticity, and small sample procedures such as bootstrap will be used to improve inference.}

enumerate• As expected, the “No-Differencing" estimator has a larger mean bias in the presence of spatial heterogeneity. This result holds for both small and large numbers of locations as well as for few or many individuals having the same sub-location specific unobserved heterogeneity $\mathcal{N}_{id}$. • The mean bias of the “Sub-location Differencing" estimator is smaller than other estimators. It increases with the number of individuals in the sub-locations, and decreases with the number of locations. For example, in a sample of 600 individuals which are spread across 100 locations with 2 sub-locations and 3 individuals in each sub-location, the mean bias is of 0.007. However, for the same sample size but spread across 30 locations with 2 sub-locations and 10 individuals in each sub-locations, the mean bias is $-0.028$. This result is in line with our asymptotic derivations. • For a fixed number of locations, the bias increases with the number of individuals in sub-locations. The empirical consequence of this results is that our estimator should be applied when the size of sub-locations is small. • The coverage rate of the “Sub-location Differencing" estimator using the variance-covariance estimator in Equation ((ref)) and the normal asymptotic distribution suggests a conservative coverage.

We now turn to an empirical application of our estimation strategy.

Empirical Application

This section shows the empirical importance of spatial differencing methodology proposed in the previous sections. To illustrate the importance of our estimator, we ask what determines tax rates set by regional governing bodies.\footnote{There is a large literature which examines a range of factors influencing local tax rates e.g. charney1983intraurban, ashworth1997politicians, ross1999sorting, charlot2007market,charlot2015does, crowley2011does, baskaran2014identifying, buettner2016yardstick.} This question opens an important issue of identification since circular causation or omitted variable bias leads to biased and inconsistent estimators. We will use our spatial differencing method to examine the case of changes in the Finish local property tax rate at the turn of the millennium.

Finland consists of 411 municipalities (in 1999) spread across 19 regions which choose property tax rate within the limits set by the central government. In 1999, the central government decided to raise the lower limit for the year 2000 from 0.2% to 0.5%. This change created a probability mass of municipalities at the lower bound: more than half of municipalities have a taxation rate of 0.5%, making the data sample censored. We investigate what affected municipalities tax rate in the year 2000.

We estimate the parameter of the outcome equation in the model represented as in Equation (2). Specifically, the outcome variable is the level of general property tax in a municipality $i$ in a region $j$, and explanatory variables include municipality’s $i$ age structure of the population, level of municipality’s income, received subsidies, local income tax rate and a dummy for region $j$ in which the municipality $i$ is located. The selection equation determines whether the municipality sets its general tax rate at the mandatory minimum of 0.5% or above and contains all the variables which are in the outcome equation except for local income tax rate.

As illustrated in Equation ((ref)), there can be an unobserved sub-location specific effect operating at a finer spatial scale than region $j$, in our case at the level of municipalities which region $j$ consists of. Indeed, municipalities’ tax level can depend not only on its population size, income and subsidies received from the central government, but it can also depend on the level of amenities in the municipality. It is usually difficult to measure them. More importantly, even if we have a few measures of amenities or their proxies, they might not be able to capture all of them, leaving some amenities unobserved. In our case, unobserved amenities can be correlated with municipalities’ population, income level, or the level of subsidies which implies that not controlling for them will render the estimates biased and inconsistent. Therefore, using the fact that two municipalities from the same region sharing borders are neighbor, we use our spatial differencing method to tackle this problem.

We estimate equation ((ref)) with spatial differencing conducted as the difference between municipality $i$ and the average of its neighbours. Columns 1 and 2 present the results without spatial differencing, columns 3-6 with spatial differencing. Estimation is conducted with and without regional dummies respectively, and with two different estimators of standard errors: wild cluster bootstrap, and spatially-adjusted standard errors derived in Section 2.3. Clustering of the standard errors is done at the level of region $j$. Since there are only 19 regions, we use wild clustered bootstrap procedure developed by cameron2008bootstrap, which properties were studied by e.g. davidson2010wild, mackinnon2013thirty, and mackinnon2017wild. Specifically, we use a recently developed wild bootstrap package $boottest$ by roodman2019fast implemented in Stata.

We begin with discussing the results with spatial differencing: columns 3-6, Columns 3 and 4 present the results when we use spatial differencing with the standard errors calculated using the formula derived in Section 2.3. There is only one significant variable: the share of population older than 75 in column 4. This is not surprising, since our estimator of variance-covariance matrix is an asymptotic estimator, while the estimation is done on a sample with small number of clusters (nineteen regions). Therefore we use wild-bootstrap procedure which is known to be suitable for a small number of clusters. The results using wild-bootstrapping are shown in columns 5 and 6 and we see a considerable increase in the number of statistically significant results.

The comparison of columns 5 and 6 with columns 1 and 2 reveals the importance of spatial differencing. Controlling for the sub-location specific unobserved effect $\gamma _{j\alpha}$ by spatial differencing renders four variables statistically significant: share of population younger than 15, share of population older than 75, government grants, and income tax rate. This is in contrast to columns 1 and 2 in which income tax rate is the only significant variable. Not controlling for $\gamma _{j\alpha}$ leads to the omitted variable bias which, apart from rendering the estimates inconsistent, inflates the standard errors and makes the estimates mostly insignificant. Spatial differencing controls for this omitted variable bias, which means that they do not 'end up' in the error term and do not inflate the standard errors. In addition to comparing estimates with and without spatial differencing, it is also instructive to compare columns 5 and 6: spatial differencing with and without regional dummies. We see that controlling for the regional unobserved effect $\gamma_{j}$ does not help to control for sub-location specific effects $\gamma _{j\alpha}$. Indeed, the magnitude and the statistical significance of the estimates change very little, and even the one variable that looses its statistical significance after including regional dummies (municipality's income) is only marginally significant without these dummies.

Overall, our empirical analysis shows that controlling for the unobserved municipality effects matters. Estimations which control for spatial unobserved effects only at the regional suggest that the income tax rate is the only determinant of general tax rate set by the municipalities. However, after controlling for the sub-location specific unobserved effects, the tax rate depends not only on the income tax rate but also on the age composition of their population - the share of young as well as the share of elderly population. These results thus indicate that spatial differencing is an important tool to deal with omitted variable bias which often plagues empirical studies on local taxation.

table[table omitted — 4,805 chars of source]

Conclusion

This paper has investigated a sample selection model with unobserved heterogeneity at a very fine location level. It proposes spatial differencing as an alternative identification strategy when instrumental variable and/or a panel data are not available. We discuss the assumptions under which the parameters of the model are identified. The estimation of the parameters is done using the classic Heckman's two-step estimation procedure. The differecing and the two-step procedure lead to a novel estimator with properties that are also relevant for spatial differencing in linear models. To understand the behavior of the new estimator, we derive a cluster asymptotic of the estimator. The derivation reveals two important implications for its empirical implementation: $(i)$ the number of clusters needs to be large for inference to be based on normal distribution. ($ii$) each cluster should have a bounded number of individuals.

Monte Carlo experiments show that accounting for sub-location specific heterogeneity is crucial for identification. It also confirms the estimator's properties derived in our asymptotic. In particular, the estimator performs better with the increasing number of locations, and fewer individuals in sub-locations. In addition, ignoring sub-locations and applying spatial differencing only to more aggregate geographical units subsuming sub-locations, the mean bias is larger. The coverage rate of the test based on the corrected standard error has an empirical coverage lower that the theoretical one.

In the empirical application which looked at the determinants of municipal tax rate, we show that using spatial differencing in combination with cluster wild-bootstrap inference tools can be extremely useful. Indeed, the new estimator reveals several determinants of the municipal tax rate that would have been missed otherwise. The development of a bootstrap appropriate sample selection models is left for future research.