Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,110 characters · 9 sections · 75 citation commands
Threshold Regression with Nonparametric Sample Splitting
\thispagestyle{empty}\setcounter{page}{0}
\setcounter{page}{1}
Sample splitting and threshold regression models have spawned a vast literature in econometrics and statistics. Existing studies typically specify the sample splitting criteria in a parametric way as whether a single random variable or a linear combination of variables crosses some unknown threshold. See, for example, Hansen00a, Caner04, SeoLinton07, LeeSeoShin11, LiLing12, Yu12, LLSS18, Hidalgo19, and Yu19. In this paper, we study a novel extension to consider a nonparametric sample splitting model. Such an extension leads to new theoretical results and substantially generalizes the empirical applicability of threshold models.
Specifically, we consider a model given by
for $i=1,\ldots ,n$, where $\mathbf{1}\left[ \cdot \right] $ is the binary indicator. In this model, the marginal effect of $x_{i}$ to $y_{i}$ can be different across $i$ as $({\Greekmath 010C} _{0}+{\Greekmath 010E} _{0})$ or ${\Greekmath 010C} _{0}$ depending on whether $q_{i}\leq {\Greekmath 010D} _{0}\left( s_{i}\right) $ or not. The threshold function ${\Greekmath 010D} _{0}(\cdot )$ is unknown, and the main parameters of interest are ${\Greekmath 010C} _{0}$, ${\Greekmath 010E} _{0}$, and ${\Greekmath 010D} _{0}(\cdot )$. The novel feature of this model is that the sample splitting is determined by an unknown relationship between two variables $q_{i}$ and $s_{i}$, and their relationship is characterized by the nonparametric threshold function $ {\Greekmath 010D} _{0}(\cdot )$. In contrast, the classical threshold regression models assume ${\Greekmath 010D} _{0}\left( \cdot \right) $ to be a constant or a linear index. Our new specification can cover interesting cases that have not been studied. For example, we can consider the threshold to be heterogeneous and specific to each observation $i$ if we see ${\Greekmath 010D} _{0}\left( s_{i}\right) ={\Greekmath 010D} _{0i}$; or the threshold to be determined by the direction of some moment condition ${\Greekmath 010D} _{0}(s_{i})=\mathbb{E}[q_{i}|s_{i}]$. Apparently, when ${\Greekmath 010D} _{0}(s)={\Greekmath 010D} _{0}$ or ${\Greekmath 010D} _{0}(s)={\Greekmath 010D} _{0}s$ for some parameter ${\Greekmath 010D} _{0}$ and $s\neq 0$, it reduces to the standard threshold regression model.
The new model is motivated by the following two applications: estimating potentially heterogeneous thresholds in public economics and determining spatial sample splitting in urban economics. The first one is about the tipping point model proposed by Schelling71, who analyzes the phenomenon that a neighborhood's white population substantially decreases once the minority share exceeds a certain threshold, called the tipping point. Card08 empirically estimate the tipping point model by considering the constant threshold regression, $y_{i}={\Greekmath 010C} _{10}+{\Greekmath 010E} _{10}\mathbf{1}\left[ q_{i}\leq {\Greekmath 010D} _{0}\right] +x_{2i}^{\top }{\Greekmath 010C} _{20}+u_{i}$, where $y_{i}$ is the white population change in a decade and $ q_{i}$ is the initial minority share in the $i$th tract. The parameters $ {\Greekmath 010E} _{10}$ and ${\Greekmath 010D} _{0}$ denote the change size and the threshold, respectively. In Section VII of Card08, however, they find that the tipping point ${\Greekmath 010D} _{0}$ varies depending on the attitudes of white residents toward the minority. This finding raises the concern on the constant threshold model and motivates us to study the more general model ( (ref)) by specifying the tipping point ${\Greekmath 010D} _{0}$ as a nonparametric function of local demographic characteristics. We estimate such a tipping function in Section (ref).
For the second application, we use the model ((ref)) to define metropolitan area boundaries, which is a fundamental problem in urban economics. Recently, many studies propose to use nighttime light intensity collected from satellite imagery to define the metropolitan area. They set an ad hoc level of light intensity as a threshold and categorize a pixel in the satellite imagery as a part of the metropolitan area if the light intensity of that pixel is higher than the threshold. See, for example, Rozenfeld11, Henderson12, Dingel19, and Vogel19. In contrast, the model ((ref)) can provide a data-driven guidance of choosing the intensity threshold from the econometric perspective, if we let $y_{i}$ as the light intensity in the $i$th pixel and $(q_{i},s_{i})$ as the location information of that pixel (more precisely, the coordinate of a point on a rotated map as described in Section (ref)). In Section (ref), we estimate the metropolitan area of Dallas, Texas, especially its development from 1995 to 2010, and find substantially different results from the conventional approaches. To the best of our knowledge, this is the first study to nonparametrically determine the metropolitan area using a threshold model.
We develop a two-step estimation procedure of ((ref)), where we estimate ${\Greekmath 010D} _{0}\left( \cdot \right) $ by the local constant least squares. Under the shrinking threshold asymptotics as in Bai97b, Bai98, and Hansen00a, we show that the nonparametric estimator $ \widehat{{\Greekmath 010D} }(\cdot )$ is uniformly consistent and has a highly nonstandard limiting distribution. Based on such distribution, we develop a pointwise specification test of ${\Greekmath 010D} _{0}(s)$ for any given $s$, which enables us to construct a confidence interval by inverting the test. Besides, the parametric part $(\widehat{{\Greekmath 010C} }^{\top },\widehat{{\Greekmath 010E} } ^{\top })^{\top }$ is shown to satisfy the root-$n$ asymptotic normality.
We highlight some novel technical features of the new estimator as follows. First, since the nonparametric function ${\Greekmath 010D} _{0}\left( \cdot \right) $ is inside the indicator function, technical proofs of the asymptotic results are non-standard.\ In particular, we establish the uniform rate of convergence of $\widehat{{\Greekmath 010D} }\left( \cdot \right) $, which involves substantially more complicated derivations than the standard (constant) threshold regression model. Second, we find that, unlike the standard kernel estimator, $\widehat{{\Greekmath 010D} }(\cdot )$ is asymptotically unbiased even if the optimal bandwidth is used. Also, when the change size ${\Greekmath 010E} _{0}$ shrinks very slowly, the optimal rate of convergence of $\widehat{{\Greekmath 010D} } (\cdot )$ becomes close to the root-$n$ rate. In the standard kernel regression, such a fast rate of convergence can be obtained when the unknown function is infinitely differentiable, while we only require the second-order differentiability of ${\Greekmath 010D} _{0}\left( \cdot \right) $. Third, to limit the effect of estimating ${\Greekmath 010D} _{0}\left( \cdot \right) $ to $( \widehat{{\Greekmath 010C} }^{\top },\widehat{{\Greekmath 010E} }^{\top })^{\top }$, we propose to use the observations that are sufficiently away from the estimated threshold in the second-step parametric estimation. The choice of this distance is obtained by the uniform convergence rate of $\widehat{{\Greekmath 010D} }\left( \cdot \right) $. Fourth, we let the variables be cross-sectionally dependent by considering the strong-mixing random field as in Conley99 and Conley07. This generalization allows us to study nonparametric sample splitting of spatial observations. For instance, if we let $(q_{i},s_{i})$ correspond to the geographical location (i.e., latitude and longitude on the map), then the threshold $\mathbf{1}\left[ q_{i}\leq {\Greekmath 010D} _{0}\left( s_{i}\right) \right] $ identifies the unknown border yielding a two-dimensional sample splitting. In more general contexts, the model can be applied to identify social or economic segregation over interacting agents. Finally, noting that $\mathbf{1}\left[ q_{i}\leq {\Greekmath 010D} _{0}\left( s_{i}\right) \right] $ can be considered as the special case of $\mathbf{1} \left[ g_{0}\left( q_{i},s_{i}\right) \leq 0\right] $ when $g_{0}$ is monotonically increasing in $q_{i}$, we discuss how to extend the proposed method to such a more general case that leads to a threshold contour model.
The rest of the paper is organized as follows. Section (ref) sets up the model, establishes the identification, and defines the estimator. Section (ref) derives the asymptotic properties of the estimators and develops a likelihood ratio test of the threshold function. Section (ref) describes how to extend the main model to estimate a threshold contour. Section (ref) studies small sample properties of the proposed statistics by Monte Carlo simulations. Section (ref) applies the new method to estimate the tipping point function and to determine metropolitan areas. Section (ref) concludes this paper with some remarks. The main proofs are in the Appendix, and all the omitted proofs are collected in the supplementary material.
We use the following notations. Let $\rightarrow _{p}$ denote convergence in probability, $\rightarrow _{d}$ convergence in distribution, and $ \Rightarrow $ weak convergence of the underlying probability measure as $ n\rightarrow \infty $. Let $\left\lfloor r\right\rfloor $ denote the biggest integer smaller than or equal to $r$, $\mathbf{1}[E]$ the indicator function of a generic event $E$, and $\left\Vert A\right\Vert $ the Euclidean norm of a vector or matrix $A$. For any set $B$, let $|B|$ as the cardinality of $B$.
We assume spatial processes located on an evenly spaced lattice $\Lambda \subset \mathbb{R} ^{2}$, following Conley99, Conley07, and CarbonFrancqTran2007.\footnote{ It can be extended to an unevenly spaced lattice as in Bolthausen82 and Jenish09 with substantially more complicated notations.\ (cf.\ footnote 9 in Conley99).} We consider the threshold regression model given by ((ref)), which is
where the observations $\{(y_{i},x_{i}^{\top },q_{i},s_{i})^{\top }\in \mathbb{R}^{1+\dim (x)+1+1};i\in \Lambda _{n}\}$ are a triangular array of real random variables defined on some probability space with $\Lambda _{n}$ being a fixed sequence of finite subsets of $\Lambda $. In this setup, the cardinality of $\Lambda _{n}$, $n=|\Lambda _{n}|$, is the sample size and $ \sum_{i\in \Lambda _{n}}$ denotes the summation of all observations. For readability, we postpone the regularity conditions on $\Lambda _{n}$ in Assumption A later. The threshold function ${\Greekmath 010D} _{0}:\mathbb{R\rightarrow R}$ as well as the regression coefficients ${\Greekmath 0112} _{0}=({\Greekmath 010C} _{0}^{\top },{\Greekmath 010E} _{0}^{\top })^{\top }\in \mathbb{R}^{2\dim (x)}$ are unknown, and they are the parameters of interest.\footnote{ The main results of this paper can be extended to consider multi-dimensional $s_{i}$ using multivariate kernels. However, we only consider the scalar case for the expositional simplicity. Furthermore, the results are readily generalized to the case where only a subset of parameters differ between regimes.} Since we consider a shrinking threshold effect, the parameter $ {\Greekmath 010E} _{0}$ is to depend on the sample size $n$ as in Assumption A-(ii) below; hence ${\Greekmath 010E} _{0}$ and ${\Greekmath 0112} _{0}$ should be written as ${\Greekmath 010E} _{n0}$ and ${\Greekmath 0112} _{n0}$, respectively. However, we write ${\Greekmath 010E} _{0}$ and ${\Greekmath 0112} _{0}$ for simplicity. We let $\mathcal{Q}\subset \mathbb{R} $ and $\mathcal{S}\subset \mathbb{R} $ denote the supports of $q_{i}$ and $s_{i}$, respectively. Suppose the space of ${\Greekmath 010D} _{0}\left( s\right) $ for any $s$ is a compact set $\Gamma \subset \mathbb{R} $.
First, we establish the identification, which requires the following conditions.
\paragraph{Assumption ID}
Assumption ID is mild. The condition (i) excludes endogeneity, and (ii) is the full rank condition to identify the global parameters ${\Greekmath 010C} _{0}$ and $ {\Greekmath 010E} _{0}$. The conditions (ii) and (iii) require that the location of the threshold is not on the boundary of the support of $q_{i}$ for any $s\in \mathcal{S}$, which is inevitable for identification and has been commonly assumed in the existing threshold literature (e.g., Hansen00a). If $ {\Greekmath 010D} _{0}(s)$ reaches the boundary of $q_{i}$ for some $s\in \mathcal{S}$, then no observation can be generated from one side of the threshold function at this $s$, and identification is failed. The second condition in (iii) assumes the coefficient change exists (i.e., ${\Greekmath 010E} _{0}\neq 0$). Note that it does not require $\mathbb{E}\left[ x_{i}x_{i}^{\top }|q_{i}=q,s_{i}=s \right] $ to be of full rank, and hence $q_{i}$ or $s_{i}$ can be one of the elements of $x_{i}$ (e.g., the threshold autoregressive model by Tong83) or a linear combination of $x_{i}$. The condition (iv) requires the conditional density of $q_{i}$ given any $s_{i}$ is positive and bounded in $\Gamma $.
Under Assumption ID, the following theorem establishes the identification of the semiparametric threshold regression model ((ref)).
Given identification, we proceed to estimate this semiparametric model in two steps. First, for given $s\in \mathcal{S}$, we fix ${\Greekmath 010D} _{0}(s)={\Greekmath 010D} $ and obtain $\widehat{{\Greekmath 010C} }\left( {\Greekmath 010D} ;s\right) $ and $ \widehat{{\Greekmath 010E} }\left( {\Greekmath 010D} ;s\right) $ by the local constant least squares conditional on ${\Greekmath 010D} $:
where
for some kernel function $K\left( \cdot \right) $ and a bandwidth parameter $ b_{n}$. Then ${\Greekmath 010D} _{0}(s)$ is estimated by
for given $s$, where $\Gamma _{n}=\Gamma \cap \{q_{1},\ldots ,q_{n}\}$ and $ Q_{n}\left( {\Greekmath 010D} ;s\right) $ is the concentrated sum of squares defined as
To avoid any additional technical complexity, we focus on estimation of $ {\Greekmath 010D} _{0}(s)$ at $s\in \mathcal{S}_{0}\subset \mathcal{S}$ for some compact interior subset $\mathcal{S}_{0}$ of the support, say the middle 70% quantiles. Note that, given $s$, the nonparametric estimator $\widehat{ {\Greekmath 010D} }\left( s\right) $ can be seen as a local version of the standard (constant) threshold regression estimator. Therefore, the computation of ( (ref)) requires one-dimensional grid search of the threshold for only $n$ times as in the standard threshold regression estimation.
In the second step, we estimate the parametric components ${\Greekmath 010C} _{0}$ and $ {\Greekmath 010E} _{0}$. To minimize any potential effects from the first step estimation, we estimate ${\Greekmath 010C} _{0}$ and ${\Greekmath 010E} _{0}^{\ast }={\Greekmath 010C} _{0}+{\Greekmath 010E} _{0}$ using the observations that are sufficiently away from the estimated threshold. This is implemented by considering
for some constant ${\Greekmath 0119} _{n}>0$ satisfying ${\Greekmath 0119} _{n}\rightarrow 0$ as $ n\rightarrow \infty $, which is defined later. The change size ${\Greekmath 010E} $ can be estimated as $\widehat{{\Greekmath 010E} }=\widehat{{\Greekmath 010E} }^{\ast }-\widehat{{\Greekmath 010C} } $.
For the asymptotic behavior of the threshold estimator, the existing literature typically assumes martingale difference arrays (e.g., Hansen00a and LLSS18) or random samples (e.g., Yu12 and Yu19). In this paper, we allow for cross-sectional dependence by considering spatial ${\Greekmath 010B} $-mixing processes as in Bolthausen82 and Conley99. More precisely, for any indices (or locations) $i,j\in \Lambda $, we define the metric ${\Greekmath 0115} \left( i,j\right) =\max_{1\leq \ell \leq \dim \left( \Lambda \right) }\left\vert i_{\ell }-j_{\ell }\right\vert $ and the corresponding norm $\max_{1\leq \ell \leq \dim \left( \Lambda \right) }\left\vert i_{\ell }\right\vert $, where $i_{\ell }$ denotes the $ \ell $th component of $i$. The distance of any two subsets $\Lambda _{1},\Lambda _{2}\subset \Lambda $ is defined as ${\Greekmath 0115} (\Lambda _{1},\Lambda _{2})=\inf \{{\Greekmath 0115} (i,j):i\in \Lambda _{1},j\in \Lambda _{2}\} $. We let $\mathcal{F}_{\Lambda }$ be the ${\Greekmath 011B} $-algebra generated by a random sequence $\left( x_{i}^{\top },q_{i},s_{i},u_{i}\right) ^{\top }$ for $i\in \Lambda $ and define the spatial ${\Greekmath 010B} $-mixing coefficient as
where $\left\vert \Lambda _{1}\right\vert \leq k$ and $\left\vert \Lambda _{2}\right\vert \leq l$. Without loss of generality, we assume ${\Greekmath 010B} _{k,l}(0)=1$ and ${\Greekmath 010B} _{k,l}(m)$ is monotonically decreasing in $m$ for all $k$ and $l$.
The following conditions are imposed for deriving the asymptotic properties of our two-step estimator. Let $f\left( q,s\right) $\ be the joint density function of $(q_{i},s_{i})$ and
\paragraph{Assumption A}
We provide some discussions about Assumption A. First, we assume that $q_{i}$ and $s_{i}$ are continuous random variables as in the example in Section (ref). This setup also covers the two-dimensional \textquotedblleft spatial structural break\textquotedblright\ model as a special case as in the example in Section (ref). For the latter case, we denote $n_{1}$ and $n_{2}$ as the numbers of rows (latitudes) and columns (longitudes) in the grid of pixels, and normalize $q$ and $s$ in the way that $q\in \{1/n_{1},2/n_{1},\ldots ,1\}$ and $s\in \{1/n_{2},2/n_{2},\ldots ,1\}$. Under similar (and even weaker) regularity conditions as Assumption A, we can show that the asymptotic results in the following sections extend to this case once we treat $(q_{i},s_{i})^{\top }$ as independently and uniformly distributed random variables over $\left[ 0,1 \right] ^{2}$. Such similarity is also found in the standard structural break and the threshold regression models (e.g.,\ Proposition 5 in Bai98 and Theorem 1 in Hansen00a).
Second, Assumption A is mild and common in the existing literature. In particular, the condition (i) is the same as in Bolthausen82 to define the latent random field. Note that ${\Greekmath 0115} _{0}$ can be any strictly positive value, and hence we can impose ${\Greekmath 0115} _{0}>1$ without loss of generality. The condition (ii) adopts the widely used shrinking change size setup as in Bai97b, Bai98, and Hansen00a to obtain a simple limiting distribution. In contrast, a constant change size (when $ {\Greekmath 010F} =0$) leads to a complicated asymptotic distribution of the threshold estimator, which depends on nuisance parameters (e.g., Chan93). The condition (iii) is required to establish the maximal inequality and uniform convergence in a spatially dependent random field. We impose a stronger condition than Jenish09 to obtain the maximal inequality uniformly over ${\Greekmath 010D} $ and $s$. We could weaken this condition such that ${\Greekmath 010B} _{k,l}(m)$ decays at a polynomial rate (e.g., ${\Greekmath 010B} _{k,l}(m)\leq Cm^{-r}$ for some $r>8$ and a constant $C$ as in CarbonFrancqTran2007) if we impose higher moment restrictions in the condition (v). However, this exponential decay rate simplifies the technical proofs. The conditions (iv) to (viii) are similar to Assumption 1 of Hansen00a. The condition (ix) imposes restrictions on the bandwidth $b_{n}$ , which now depends on ${\Greekmath 010F} $ and ${\Greekmath 0127} $. The condition (x) holds for many commonly used kernel functions including the Gaussian and the uniform kernels.
Third, we assume ${\Greekmath 010D} _{0}\left( \cdot \right) $ to be a function from $ \mathcal{S}$ to $\Gamma $ in Assumption A-(vi), which is not necessarily one-to-one. For this reason, sample splitting based on $\mathbf{1}\left[ q_{i}\leq {\Greekmath 010D} _{0}\left( s_{i}\right) \right] $ can be different from that based on $\mathbf{1}\left[ s_{i}\geq \breve{{\Greekmath 010D}}_{0}\left( q_{i}\right) \right] $ for some function $\breve{{\Greekmath 010D}}_{0}\left( \cdot \right) $. Instead of restricting ${\Greekmath 010D} _{0}\left( \cdot \right) $ to be one-to-one in this paper, we presume that one knows which variables should be respectively assigned as $q_{i}$ and $s_{i}$ from the context. Alternatively, we can consider a function $g_{0}\left( q,s\right) $ such that $g_{0}$ is monotonically increasing in $q$ for any $s$. Then, $\mathbf{1 }\left[ q_{i}\leq {\Greekmath 010D} _{0}\left( s_{i}\right) \right] $ is viewed as a special case of $\mathbf{1}\left[ g_{0}\left( q_{i},s_{i}\right) \leq 0 \right] $ by inverting $g_{0}\left( \cdot ,s\right) =q^{\ast }$, where $ g_{0}\left( q^{\ast },s\right) =0$. We discuss such extension to identify a threshold contour in Section (ref).
We first obtain the asymptotic properties of $\widehat{{\Greekmath 010D} }\left( s\right) $. The following theorem derives the pointwise consistency and the pointwise rate of convergence of $\widehat{{\Greekmath 010D} }\left( s\right) $ at the interior points of $\mathcal{S}$.
The pointwise rate of convergence of $\widehat{{\Greekmath 010D} }\left( s\right) $ depends on two parameters, ${\Greekmath 010F} $ and $b_{n}$. It is decreasing in $ {\Greekmath 010F} $ like the parametric (constant) threshold case: a larger ${\Greekmath 010F} $ reduces the threshold effect ${\Greekmath 010E} _{0}=c_{0}n^{-{\Greekmath 010F} }$ and hence decreases the effective sampling information on the threshold. Since we estimate ${\Greekmath 010D} _{0}(\cdot )$ using the kernel estimation method, the rate of convergence depends on the bandwidth $b_{n}$ as well. As in the standard kernel estimator case, a smaller bandwidth decreases the effective local sample size, which reduces the precision of the estimator $\widehat{{\Greekmath 010D} } \left( s\right) $. Therefore, in order to have a sufficiently fast rate of convergence, we need to choose $b_{n}$ large enough when the threshold effect ${\Greekmath 010E} _{0}$ is expected to be small (i.e., when ${\Greekmath 010F} $ is close to $1/2$).
Unlike the standard kernel estimator, there seems no bias-variance trade-off in $\widehat{{\Greekmath 010D} }\left( s\right) $ in ((ref)), implying that we could improve the rate of convergence by choosing a larger bandwidth $b_{n}$ . However, as we can find in Theorem (ref) below, $b_{n}$ cannot be chosen too large to result in $n^{1-2{\Greekmath 010F} }b_{n}^{2}\rightarrow \infty $ , under which $n^{1-2{\Greekmath 010F} }b_{n}(\widehat{{\Greekmath 010D} }\left( s\right) -{\Greekmath 010D} _{0}\left( s\right) )$ is no longer $O_{p}(1)$. Therefore, we can obtain the optimal bandwidth using the restriction that $n^{1-2{\Greekmath 010F} }b_{n}^{2}$ does not diverge.
Under this restriction, we find the optimal bandwidth as $b_{n}^{\ast }=n^{-(1-2{\Greekmath 010F} )/2}c^{\ast }$ for some constant $0<c^{\ast }<\infty $, which yields the optimal pointwise rate of convergence of $\widehat{{\Greekmath 010D} } \left( s\right) $ as $n^{-(1-2{\Greekmath 010F} )/2}$. However, such a bandwidth choice is not feasible because of the unknown constant $c^{\ast }$ and the nuisance parameter ${\Greekmath 010F} $ that are not estimable. In practice, we suggest cross validation as we implement in Section (ref), although its statistical properties need to be studied further. Note that, when the change size ${\Greekmath 010E} _{0}$ shrinks very slowly with $n$ (i.e., $ {\Greekmath 010F} $ is close to $0$), the optimal rate of convergence of $\widehat{ {\Greekmath 010D} }(\cdot )$ is close to $n^{-1/2}$. This $\sqrt{n}$-rate is obtained in the standard kernel regression if the unknown function is infinitely differentiable, while we only require the second-order differentiability of $ {\Greekmath 010D} _{0}\left( \cdot \right) $.
The next theorem derives the limiting distribution of $\widehat{{\Greekmath 010D} } \left( s\right) $. We let $W(\cdot )$ be a two-sided Brownian motion defined as in Hansen00a:
where $W_{1}(\cdot )$ and $W_{2}(\cdot )$ are independent standard Brownian motions on $[0,\infty )$.
The drift term ${\Greekmath 0116} \left( r,{\Greekmath 0125} ;s\right) $ in ((ref)) depends on the constant $0<{\Greekmath 0125} <\infty $, which is the limit of $n^{1-2{\Greekmath 010F} }b_{n}^{2}=(n^{1-2{\Greekmath 010F} }b_{n})b_{n}$, and $|\dot{{\Greekmath 010D}}_{0}(s)|$, the steepness of ${\Greekmath 010D} _{0}(\cdot )$ at $s$. Interestingly, it resembles the typical $O(b_{n})$ boundary bias of the standard local constant estimator. However, this non-zero drift term is not because of the typical boundary effect but because of the inequality restriction inside the indicator function, $\mathbf{1}\left[ q_{i}\leq {\Greekmath 010D} _{0}\left( s_{i}\right) \right] $, which characterizes the sample splitting.
It is important to note that having this non-zero drift term in the limiting expression does not mean that the limiting distribution of $\widehat{{\Greekmath 010D} } \left( s\right) $ has a non-zero mean, even when we use the optimal bandwidth $b_{n}^{\ast }=O(n^{-(1-2{\Greekmath 010F} )/2})$ satisfying $ n^{1-2{\Greekmath 010F} }b_{n}^{\ast 2}\rightarrow {\Greekmath 0125} \in (0,\infty )$. This is mainly because the drift function ${\Greekmath 0116} \left( r,{\Greekmath 0125} ;s\right) $ is symmetric about zero and hence the limiting random variable $\arg \max_{r\in \mathbb{R} }\left( W\left( r\right) +{\Greekmath 0116} \left( r,{\Greekmath 0125} ;s\right) \right) $ is mean zero. In general, we can show that the random variable $\arg \max_{r\in \mathbb{R} }\left( W\left( r\right) +{\Greekmath 0116} \left( r,{\Greekmath 0125} ;s\right) \right) $ always has mean zero if ${\Greekmath 0116} \left( r,{\Greekmath 0125} ;s\right) $ is a non-random function that is symmetric about zero and monotonically decreasing fast enough. This result might be of independent research interest and is summarized in Lemma (ref) in the Appendix. Figure (ref) depicts the drift function ${\Greekmath 0116} \left( r,{\Greekmath 0125} ;s\right) $ for various kernels when ${\Greekmath 0118} (s)/\left( {\Greekmath 0125} |\dot{{\Greekmath 010D}}_{0}(s)|\right) =1$.
Since the limiting distribution in ((ref)) depends on unknown components, like ${\Greekmath 0125} $ and $\dot{{\Greekmath 010D}}_{0}(s)$, it is hard to use this result for further inference. We instead suggest undersmoothing for practical use. More precisely, if we suppose $n^{1-2{\Greekmath 010F} }b_{n}^{2}\rightarrow 0$ as $n\rightarrow \infty $, then the limiting distribution in ((ref)) simplifies to\footnote{ We let ${\Greekmath 0120} _{0}\left( r,0;s\right) =\int_{0}^{\infty }K\left( t\right) dt=1/2$ and ${\Greekmath 0120} _{1}\left( r,0;s\right) =\int_{0}^{\infty }tK\left( t\right) dt<\infty $.}
as $n\rightarrow \infty $, which appears the same as in the parametric case in Hansen00a except for the scaling factor $n^{1-2{\Greekmath 010F} }b_{n}$. The distribution of $\arg \max_{r\in \mathbb{R} }\left( W\left( r\right) -\left\vert r\right\vert /2\right) $ is known (e.g., Bhattacharya76 and Bai97b), which is also described in Hansen (2000, p.581). The ${\Greekmath 0118} \left( s\right) $ term determines the scale of the distribution at given $s$ in the way that it increases in the conditional variance $\mathbb{E}[u_{i}^{2}|x_{i},q_{i},s_{i}]$ and decreases in the size of the threshold constant $c_{0}$ and the density of $ (q_{i},s_{i})$ near the threshold.
Even when $n^{1-2{\Greekmath 010F} }b_{n}^{2}\rightarrow 0$ as $n\rightarrow \infty $ , the asymptotic distribution in ((ref)) still depends on the unknown parameter ${\Greekmath 010F} $ (or equivalently $c_{0}$) in ${\Greekmath 0118} \left( s\right) $ that is not estimable. Thus, this result cannot be directly used for inference of ${\Greekmath 010D} _{0}\left( s\right) $. Alternatively, given any $s\in \mathcal{S}_{0}$, we can consider a pointwise likelihood ratio test statistic for
\ which is given as
The following corollary obtains the limiting null distribution of this test statistic that is free of nuisance parameters. By inverting the likelihood ratio statistic, we can form a pointwise confidence interval for ${\Greekmath 010D} _{0}\left( s\right) $.
When $\mathbb{E}[u_{i}^{2}|x_{i},q_{i},s_{i}]=\mathbb{E}[u_{i}^{2}|s_{i}]$, which is the case of local conditional homoskedasticity, the scale parameter ${\Greekmath 0118} _{LR}\left( s\right) $ is simplified as ${\Greekmath 0114} _{2}$, and hence the limiting null distribution of $LR_{n}(s)$ becomes free of nuisance parameters and the same for all $s\in \mathcal{S}_{0}$. Though this limiting distribution is still nonstandard, the critical values in this case can be simulated using the same method as Hansen (2000, p.582) with the scale adjusted by ${\Greekmath 0114} _{2}$. More precisely, since the distribution function of ${\Greekmath 0110} =\max_{r\in \mathbb{R} }\left( 2W\left( r\right) -\left\vert r\right\vert \right) $ is given as $ \mathbb{P}({\Greekmath 0110} \leq z)=(1-\exp (-z/2))^{2}\mathbf{1}\left[ z\geq 0\right] $ , the distribution function of ${\Greekmath 0110} ^{\ast }={\Greekmath 0114} _{2}{\Greekmath 0110} $ is $ \mathbb{P}({\Greekmath 0110} ^{\ast }\leq z)=(1-\exp (-z/2{\Greekmath 0114} _{2}))^{2}\mathbf{1} \left[ z\geq 0\right] $, where ${\Greekmath 0110} ^{\ast }$ is the limiting random variable of $LR_{n}(s)$ given in ((ref)) under the local conditional homoskedasticity. By inverting it, we can obtain the critical values given a choice of $K(\cdot )$. For instance, the critical values for the Gaussian kernel is reported in Table (ref), where ${\Greekmath 0114} _{2}=(2\sqrt{{\Greekmath 0119} } )^{-1}\simeq 0.2821$ in this case.
In general, we can estimate ${\Greekmath 0118} _{LR}\left( s\right) $ by
where $\widehat{{\Greekmath 010E} }$ is from ((ref)) and ((ref)), and $ \widehat{{\Greekmath 011B} }^{2}(s)$, $\widehat{D}\left( \widehat{{\Greekmath 010D} }\left( s\right) ,s\right) $, and $\widehat{V}\left( \widehat{{\Greekmath 010D} }\left( s\right) ,s\right) $ are the standard Nadaraya-Watson estimators. In particular, we let $\widehat{{\Greekmath 011B} }^{2}(s)=\sum_{i\in \Lambda _{n}}{\Greekmath 0121} _{1i}(s)\widehat{u}_{i}^{2}$ with $\widehat{u}_{i}=y_{i}-x_{i}^{\top } \widehat{{\Greekmath 010C} }-x_{i}^{\top }\widehat{{\Greekmath 010E} }\mathbf{1}\left[ q_{i}\leq \widehat{{\Greekmath 010D} }\left( s_{i}\right) \right] $,
where
for some bivariate kernel function $\mathbb{K}(\cdot ,\cdot )$ and bandwidth parameters $(b_{n}^{\prime },b_{n}^{\prime \prime })$.
Finally, we show the $\sqrt{n}$-consistency of the semiparametric estimators $\widehat{{\Greekmath 010C} }$ and $\widehat{{\Greekmath 010E} }^{\ast }$ in ((ref)) and ( (ref)). For this purpose, we first obtain the uniform rate of convergence of $\widehat{{\Greekmath 010D} }\left( s\right) $.
Apparently, the uniform consistency of $\widehat{{\Greekmath 010D} }\left( s\right) $ follows when $\log n/(n^{1-2{\Greekmath 010F} }b_{n})\rightarrow 0$ as $ n\rightarrow \infty $. Based on this uniform convergence, the following theorem derives the joint limiting distribution of $\widehat{{\Greekmath 010C} }$ and $ \widehat{{\Greekmath 010E} }^{\ast }$. We let $\widehat{{\Greekmath 0112} }^{\ast }=(\widehat{ {\Greekmath 010C} }^{\top },\widehat{{\Greekmath 010E} }^{\ast \top })^{\top }$ and ${\Greekmath 0112} _{0}^{\ast }=({\Greekmath 010C} _{0}^{\top },{\Greekmath 010E} _{0}^{\ast \top })^{\top }$.
For the second-step estimator $\widehat{{\Greekmath 0112} }^{\ast }$, we use ((ref)) and ((ref)), instead of the conventional plug-in estimation, say $\arg \min_{{\Greekmath 010C} ,{\Greekmath 010E} }\sum_{i\in \Lambda _{n}}(y_{i}-x_{i}^{\top }{\Greekmath 010C} -x_{i}^{\top }{\Greekmath 010E} \mathbf{1}[q_{i}\leq \widehat{{\Greekmath 010D} }\left( s_{i}\right) ])^{2}\mathbf{1}[s_{i}\in \mathcal{S} _{0}]$. The reason is that the first-step nonparametric estimator $\widehat{ {\Greekmath 010D} }(\cdot )$ may not be asymptotically orthogonal to the second step. Unlike the standard semiparametric literature (e.g.,\ Assumption N(c) in Andrews94a), the asymptotic effect of $\widehat{{\Greekmath 010D} }\left( s\right) $ to the second-step estimation is not easily derived due to the discontinuity. The new estimation idea above, however, only uses the observations that are little affected by the estimation error in the first step to achieve asymptotic orthogonality. As we verify in Lemma (ref) in the Appendix, this is done by choosing a large enough ${\Greekmath 0119} _{n}$ in ((ref)) and ((ref)) such that the observations that are included in the second step are outside the uniform convergence bound of $\left\vert \widehat{{\Greekmath 010D} }\left( s\right) -{\Greekmath 010D} _{0}\left( s\right) \right\vert $. Thanks to the threshold regression structure, we can estimate the parameters on each side of the threshold even using these subsamples. Meanwhile, we also want ${\Greekmath 0119} _{n}\rightarrow 0$ fast enough to include more observations. By doing so, though we lose some efficiency in finite samples, we can derive the asymptotic normality of $\widehat{{\Greekmath 0112} }=(\widehat{{\Greekmath 010C} }^{\top }, \widehat{{\Greekmath 010E} }^{\top })^{\top }$ that has zero mean and achieves the same asymptotic variance as if ${\Greekmath 010D} _{0}(\cdot )$ was known.
By the delta method, Theorem (ref) readily yields the limiting distribution of $\widehat{{\Greekmath 0112} }=(\widehat{{\Greekmath 010C} }^{\top },\widehat{{\Greekmath 010E} }^{\top })^{\top }$ as
where
with $z_{i}=[x_{i}^{\top },x_{i}^{\top }\mathbf{1}\left[ q_{i}\leq {\Greekmath 010D} _{0}\left( s_{i}\right) \right] ]^{\top }$. The asymptotic variance expressions in ((ref)) and ((ref)) allow for cross-sectional dependence as they have the long-run variance (LRV) forms $\Omega ^{\ast }$ and $\Omega $. They can be consistently estimated by the robust estimator developed by Conley07 using $\widehat{u}_{i}=(y_{i}-x_{i}^{\top } \widehat{{\Greekmath 010C} }-x_{i}^{\top }\widehat{{\Greekmath 010E} }\mathbf{1}[q_{i}\leq \widehat{ {\Greekmath 010D} }\left( s_{i}\right) ])\mathbf{1}[s_{i}\in \mathcal{S}_{0}]$. The terms $\Sigma _{X}^{\ast }$ and $\Sigma _{X}$ can be estimated by their sample analogues.
When we consider sample splitting over a two-dimensional space (i.e., $q_{i}$ and $s_{i}$ respectively correspond to the latitude and longitude on the map), the threshold model ((ref)) can be generalized to estimate a nonparametric contour threshold model:
where the unknown function\ $g_{0}:\mathcal{Q}\times \mathcal{S}\mapsto \mathbb{R}$ determines the threshold contour on a random field that yields sample splitting. An interesting example includes identifying an unknown closed boundary over the map, such as a city boundary, and an area of a disease outbreak or airborne pollution. In social science, it can identify a group boundary or a region in which the agents share common demographic, political, or economic characteristics.
To relate this generalized form to the original threshold model ((ref) ), we suppose there exists a known center at $\left( q_{i}^{\ast },s_{i}^{\ast }\right) $ such that $g_{0}\left( q_{i}^{\ast },s_{i}^{\ast }\right) <0$. Without loss of generality, we can normalize $\left( q_{i}^{\ast },s_{i}^{\ast }\right) $ to be $\left( 0,0\right) $ and re-center the original location variables $(q_{i},s_{i})$ accordingly. In addition, we define the radius distance $l_{i}$ and angle $a_{i}^{\circ }$ of the $i$th observation relative to the origin as
where $\bar{a}_{i}^{\circ }=\arctan \left( \left\vert q_{i}/s_{i}\right\vert \right) $, and each of $(\mathbf{I}_{i},\mathbf{II}_{i},\mathbf{III}_{i}, \mathbf{IV}_{i})$ respectively denotes the indicator that the $i$th observation locates in the first, second, third, and forth quadrant.
We suppose that there is only one threshold at any angle and the threshold contour is star-shaped. For each chosen angle $a^{\circ }\in \lbrack 0^{\circ },360^{\circ })$, we rotate the original coordinate counterclockwise and implement the least squares estimation ((ref)) only using the observations in the first two quadrants after rotation. It will ensure that the threshold mapping after rotation is a well-defined function.
In particular, the angle relative to the origin is $a_{i}^{\circ }-a^{\circ } $ after rotating the coordinate by $a^{\circ }$ degrees counterclockwise, and the new location (after the rotation) is given as $(q_{i}\left( a^{\circ }\right) ,s_{i}\left( a^{\circ }\right) )$, where
After this rotation, we estimate the following nonparametric threshold model:
using only the observations $i$ satisfying $q_{i}\left( a^{\circ }\right) \geq 0$ and in the neighborhood of $s_{i}\left( a^{\circ }\right) =0$, where ${\Greekmath 010D} _{a^{\circ }}\left( \cdot \right) $ is the unknown threshold curve as in the original model ((ref)) on the $a^{\circ }$-degree-rotated coordinate plane. Such reparametrization guarantees that ${\Greekmath 010D} _{a^{\circ }}\left( \cdot \right) $ is always positive and it is estimated at the origin. Figure (ref) illustrates the idea of such rotation and pointwise estimation over a bounded support so that only the red cross points are included for estimation at different angles. Thus, the estimation and inference procedures developed in the previous sections are directly applicable, though we expect some efficiency loss as we only use the subsample with $q_{i}\left( a^{\circ }\right) \geq 0$ at each $a^{\circ }$.
We examine the small sample performance of the semiparametric threshold regression estimator by Monte Carlo simulations. We generate $n$ draws from
where $x_{i}=(1,x_{2i})^{\top }$ and $x_{2i}\in \mathbb{R}$. We let ${\Greekmath 010C} _{0}=({\Greekmath 010C} _{10},{\Greekmath 010C} _{20})^{\top }=0{\Greekmath 0113} _{2}$ and consider three different values of ${\Greekmath 010E} _{0}=({\Greekmath 010E} _{10},{\Greekmath 010E} _{20})^{\top }={\Greekmath 010E} {\Greekmath 0113} _{2}$ with ${\Greekmath 010E} =1,2,3,4$, where ${\Greekmath 0113} _{2}=(1,1)^{\top }$. For the threshold function, we let ${\Greekmath 010D} _{0}\left( s\right) =\sin (s)/2$. We consider the cross-sectional dependence structure in $\left( x_{2i},q_{i},s_{i},u_{i}\right) ^{\top }$ as follows:
where $\underline{\mathbf{u}}=(u_{1},\ldots ,u_{n})^{\top }$. The $(i,j)$th element of $\Sigma $ is $\Sigma _{ij}={\Greekmath 011A} ^{\lfloor \ell _{ij}n\rfloor } \mathbf{1}[\ell _{ij}<m/n]$, where $\ell _{ij}=\{\left( s_{i}-s_{j}\right) ^{2}+\left( q_{i}-q_{j}\right) ^{2}\}^{1/2}$ is the $L^{2}$-distance between the $i$th and $j$th observations. The diagonal elements of $\Sigma $ are normalized as $\Sigma _{ii}=1$. This $m$-dependent setup follows from the Monte Carlo experiment in Conley07 in the sense that each unit can be cross-sectionally correlated with at most $2m^{2}$ observations. Within the $ m$ distance, the dependence decays at a polynomial rate as indicated by $ {\Greekmath 011A} ^{\lfloor \ell _{ij}n\rfloor }$. The parameter ${\Greekmath 011A} $ describes the strength of cross-sectional dependence in the way that a larger ${\Greekmath 011A} $ leads to stronger dependence relative to the unit standard deviation. In particular, we consider the cases with ${\Greekmath 011A} =0$ (i.e., i.i.d. observations), $0.5$, and $1$. We consider the sample size $n=100$, $200$, and $500$, and set $\mathcal{S}_{0}$ to include the middle 70% observations of $s_{i}$.
First, Tables (ref) and (ref) report the small sample rejection probabilities of the LR test in ((ref)) for $H_{0}:{\Greekmath 010D} _{0}(s)=\sin (s)/2$ against $H_{1}:{\Greekmath 010D} _{0}(s)\neq \sin (s)/2$ at the 5% nominal level at three different locations $s=0$, $0.5$, and $1$. In particular, Table (ref) examines the case with no cross-sectional dependence (${\Greekmath 011A} =0$), while Table (ref) examines the case with cross-sectional dependence whose dependence decays slowly with ${\Greekmath 011A} =1$ and $m=10$. For the bandwidth parameter, we normalize $s_{i}$ and $q_{i}$ to have zero mean and unit standard deviation, and choose $b_{n}=0.5n^{-1/2}$ in the main regression. This choice is for undersmoothing so that $ n^{1-2{\Greekmath 010F} }b_{n}^{2}=n^{-2{\Greekmath 010F} }\rightarrow 0$. To estimate $ D\left( {\Greekmath 010D} _{0}\left( s\right) ,s\right) $ and $V\left( {\Greekmath 010D} _{0}\left( s\right) ,s\right) $, we use the rule-of-thumb bandwidths from the standard kernel regression satisfying $b_{n}^{\prime }=O(n^{-1/5})$ and $ b_{n}^{\prime \prime }=O(n^{-1/6})$. All the results are based on $1000$ simulations. In general, the test for ${\Greekmath 010D} _{0}$ performs better as (i) the sample size gets larger; (ii) the coefficient change gets more significant; (iii) the cross-sectional dependence gets weaker; and (iv) the target gets closer to the mid-support of $s$. When ${\Greekmath 010E} _{0}$ and $n$ are large, the LR test is conservative, which is also found in the classical threshold regression (e.g., Hansen00a).
Second, Table (ref) shows the finite sample coverage properties of the 95% confidence intervals for the parametric components $ {\Greekmath 010C} _{20}$, ${\Greekmath 010E} _{20}^{\ast }={\Greekmath 010C} _{20}+{\Greekmath 010E} _{20}$, and ${\Greekmath 010E} _{20}$. The results are based on the same simulation design as above with $ {\Greekmath 011A} =0.5$ and $m=3$. Regarding the tuning parameters, we use the same bandwidth choice $b_{n}=0.5n^{-1/2}$ as before and set the truncation parameter ${\Greekmath 0119} _{n}=\left( nb_{n}\right) ^{-1/2}$. Unreported results suggest that choice of the constant in the bandwidth matters particularly with small samples like $n=100$, but such effect quickly decays as the sample size gets larger. For the estimator of the LRV, we use the spatial lag order of $5$ following Conley07. Results with other lag choices are similar and hence omitted. The result suggests that the asymptotic normality is better approximated with larger samples and larger change sizes. Table (ref) shows the same results with\ a small sample adjustment of the LRV estimator for $\Omega ^{\ast }$ by dividing it by the sample truncation fraction $\sum_{i\in \Lambda _{n}}(\mathbf{1}[q_{i}> \widehat{{\Greekmath 010D} }(s_{i})+{\Greekmath 0119} _{n}]+\mathbf{1}[q_{i}<\widehat{{\Greekmath 010D} } (s_{i})-{\Greekmath 0119} _{n}])\mathbf{1}[s_{i}\in \mathcal{S}_{0}]/\sum_{i\in \Lambda _{n}}\mathbf{1}[s_{i}\in \mathcal{S}_{0}]$. This ratio enlarges the LRV estimator and hence the coverage probabilities, especially when the change size is small. It only affects the finite sample performance as it approaches one in probability as $n\rightarrow \infty $.
The first application is about the tipping point problem in social segregation, which stimulates a vast literature in labor, public, and political economics. Schelling71 initially proposes the tipping point model to study the fact that the white population decreases substantially once the minority share exceeds a certain tipping point. Card08 empirically estimate this model and find strong evidence for such a tipping point phenomenon. In particular, they specify the threshold regression model as
where for tract $i$ in a certain city, $q_{i}$ is the minority share in percentage at the beginning of a certain decade, $y_{i}$ is the normalized white population change in percentage within this decade, and $x_{2i}$ is a vector of control variables. They apply the least squares method to estimate the tipping point ${\Greekmath 010D} _{0}$. For most cities and for the periods 1970-80, 1980-90, and 1990-2000, they find that white population flows exhibit the tipping-like behavior, with the estimated tipping points ranging approximately from 5% to 20% across cities.
In Section VII of Card08, they also find that the location of the tipping point substantially depends on white residents' attitudes toward the minority. Specifically, they first construct a city-level index that measures white attitudes and regress the estimated tipping point from each city on this index. The regression coefficient is significantly different from zero, suggesting that the tipping point is heterogeneous across cities.
We go one step further by considering a more flexible model in the tract level given as
where ${\Greekmath 010D} _{0}(\cdot )$ denotes an unknown tipping point function and $ s_{i}$ denotes the attitude index. The nonparametric function ${\Greekmath 010D} _{0}(\cdot )$ here allows for heterogeneous tipping points across tracts depending on the level of the attitude index $s_{i}$ in tract $i$. Unfortunately, the attitude index by Card08 is only available at the aggregated city-level, and hence we cannot use it to analyze the census tract-level observations. For this reason, we instead use the tract-level unemployment rate as $s_{i}$ to illustrate the nonparametric threshold function, which is readily available in the original dataset. Such a compromise is far from being perfect but can be partially justified since race discrimination has been widely documented to be correlated with employment (e.g., DarityMason1998).
We use the data provided by Card08 and estimate the tipping point function ${\Greekmath 010D} _{0}(\cdot )$ over census tracts by the method introduced in Section (ref). As in their work, we drop the tracts where the minority shares are above 60 percentage points and use five control variables as $x_{2i}$, including the logarithm of mean family income, the fractions of single-unit, vacant, and renter-occupied housing units, and the fraction of workers who use public transport to travel to work. The bandwidth is set as $b_{n}=cn^{-1/2}$ for some $c>0$, so that it satisfies the technical conditions in the previous sections, where the constant $c$ is chosen by the leave-one-out cross validation. In particular, we first construct the leave-one-out estimate, $\widehat{{\Greekmath 010D} }_{-i}\left( s_{i}\right) $, of ${\Greekmath 010D} _{0}\left( s_{i}\right) $ as in ((ref)) without using the $i$th observation. Then, leaving the $i$th observation out, we construct $\widehat{{\Greekmath 010C} }_{-i}$ and $\widehat{{\Greekmath 010E} }_{-i}$ as in ((ref)) and ((ref)) with ${\Greekmath 0119} _{n}=\left( nb_{n}\right) ^{-1/2} $ using the bandwidth $b_{n}$ chosen in the previous step. We choose the bandwidth that minimizes $\sum_{i\in \Lambda _{n}}(y_{i}-\widehat{{\Greekmath 010C} } _{1,-i}-\widehat{{\Greekmath 010E} }_{1,-i}\mathbf{1}\left[ q_{i}\leq \widehat{{\Greekmath 010D} } _{-i}(s_{i})\right] -x_{2i}^{\top }\widehat{{\Greekmath 010C} }_{2,-i})^{2}\mathbf{1} \left[ s_{i}\in \mathcal{S}_{0}\right] $, where $\mathcal{S}_{0}$ again includes the middle 70% quantiles of $s_{i}$.
Figure (ref) depicts the estimated tipping points and the 95% pointwise confidence intervals by inverting the likelihood ratio test statistic ((ref)) in the years 1980-90 in Chicago, Los Angeles, and New York City, whose sample sizes are relatively large. For each city, the constant $c$ of the bandwidth $b_{n}=cn^{-1/2}$ chosen by the aforementioned cross validation is $3.20$, $4.87$, and $3.42$, respectively. We make the following comments. First, the estimates of the tipping points vary substantially in the unemployment rate within all three cities. Therefore, the standard constant tipping point model is insufficient to characterize the segregation fully. Second, the tipping points as functions of the unemployment rate do not exhibit the same pattern across cities, reinforcing the heterogeneous tipping points in the city-level as found in Card08 . Finally, the estimated tipping point $\widehat{{\Greekmath 010D} }\left( s\right) $ as a function of $s$ can be discontinuous, which does not contrast with Assumption A-(vi), that is, the true function ${\Greekmath 010D} _{0}\left( \cdot \right) $ is smooth. The discontinuity comes from the fact that $\widehat{ {\Greekmath 010D} }\left( s\right) $ is obtained by grid search and can only take values among the discrete points $\{q_{1},...,q_{n}\}$ in finite samples.
The second application is about determining the boundary of a metropolitan area, which is a fundamental question in urban economics. Recently, researchers propose to use nighttime light intensity obtained by satellite imagery to define metropolitan areas. The intuition is straightforward: metropolitan areas are bright at night while rural areas are dark.
Specifically, the National Oceanic and Atmospheric Administration (NOAA) collects satellite imagery of nighttime lights at approximately 1-kilometer resolution since 1992. NOAA further constructs several indices measuring the annual light intensity. Following the literature (e.g., Dingel19), we choose the \textquotedblleft average visible, stable lights\textquotedblright\ index that ranges from 0 (dark) to 63 (bright). For illustration, we focus on Dallas, Texas and use the data from the years 1995, 2000, 2005, and 2010. In each year, the data are recorded as a 240$ \times $360 grid that covers the latitudes from 32$^{\circ }$N to 34$^{\circ }$N and the longitudes from 98.5$^{\circ }$W to 95.5$^{\circ }$W. The total sample size is 240$\times $360=86400 each year. These data are available at NOAA's website and also provided on the authors' website. Figure (ref) depicts the intensity of the stable nighttime light of the Dallas area in 2010 as an example.
Let $y_{i}$ be the level of nighttime light intensity and $\left( q_{i},s_{i}\right) $ be the latitude and longitude of the $i$th pixel, which is normalized into the equally-spaced grids on $[0,1]^{2}$. To define the metropolitan area, existing literature in urban economics first chooses an ad hoc intensity threshold, say 95% quantile of $y_{i}$, and categorizes the $i$th pixel as a part of the metropolitan area if $y_{i}$ is larger than the threshold. See Dingel19, Vogel19, and references therein. In particular, on p.3 in Dingel19, they note that \textquotedblleft \lbrack ...] the choice of the light-intensity threshold, which governs the definitions of the resulting metropolitan areas, is not pinned down by economic theory or prior empirical research.\textquotedblright\ Our new approach can provide a data-driven guidance of choosing the intensity threshold from the econometric perspective.
To this end, we first examine whether the light intensity data exhibits a clear threshold pattern. We plot the kernel density estimates of $y_{i}$ in the year 2010 in Figure (ref). The bandwidth is the standard rule-of-thumb one. The estimated density exhibits three peaks at around the intensity levels 0, 8, and 63. They respectively correspond to the rural area, small towns, and the central metropolitan area. It shows that the threshold model is appropriate in characterizing such a mean-shift pattern.
Now we implement the rotation and estimation method described in Section (ref). In particular, we pick the center point in the bright middle area as the Dallas metropolitan center, which corresponds to the pixel point in the 181st column from the left and the 100th row from the bottom. Then for each $a^{\circ }$ over the 500 equally-spaced grid on $ [0^{\circ },360^{\circ }]$, we rotate the data by $a^{\circ }$ degrees counterclockwise and estimate the model ((ref)) with $x_{i}=1$. The bandwidth is chosen as $cn^{-1/2}$ with $c=1$. Other choices of $c$ lead to almost identical results, given the large sample size. Figure (ref) presents the estimated metropolitan areas using our nonparametric approach (red) and the area determined by the ad hoc threshold of the 95% quantile of $ y_{i}$ (black) in the years 1995, 2000, 2005, and 2010. It clearly shows the expansion of the Dallas metropolitan area over the 15 years of the sample period.
Several interesting findings are summarized as follows. First, the estimated boundary is highly nonlinear as a function of the angle. Therefore, any parametric threshold model could lead to a substantially misleading result. Second, our estimated area is larger than that determined by the ad hoc threshold, by 80.31%, 81.56%, 106.46%, and 102.09% in the years 1995, 2000, 2005, and 2010, respectively. In particular, our nonparametric estimates tend to include some suburban areas that exhibit strong light intensity and that are geographically close to the city center. For example, the very left stretch-out area in the estimated boundary corresponds to Fort Worth, which is 30 miles from downtown Dallas. Residents can easily commute by train or driving on the interstate highway 30. It is then reasonable to include Fort Worth as a part of the metropolitan Dallas area for economic analysis. Third, given the large sample size, the 95% confidence intervals of the boundary are too narrow to be distinguished from the estimates and therefore omitted from the figure. Such narrow intervals apparently exclude the boundary determined by the ad hoc method. Finally, the estimated value of ${\Greekmath 010C} _{0}+{\Greekmath 010E} _{0}$ is approximately $53$ in these sample periods, which corresponds to the 89% quantile of $y_{i}$ in the sample. This suggests that a more proper choice of the level of light intensity threshold is the 89% quantile of $y_{i}$, instead of the 95% quantile, if one needs to choose the light-intensity threshold to determine the Dallas metropolitan area.
This paper proposes a novel approach to conduct sample splitting. In particular, we develop a nonparametric threshold regression model where two variables can jointly determine the unknown threshold boundary. Our approach can be easily generalized so that the sample splitting depends on more numbers of variables, though such an extension is subject to the curse of dimensionality, as usually observed in the kernel regression literature. The main interest is in identifying the threshold function that determines how to split the sample. Thus our model should be distinguished from the smoothed threshold regression model or the random coefficient regression model. It instead could be seen as an unsupervised learning tool for clustering.
This new approach is empirically relevant in broad areas studying sample splitting (e.g., segregation and group-formation) and heterogeneous effects over different subsamples. We illustrate some of them with the tipping point problem in social segregation and metropolitan area determination using satellite imagery datasets. Though we omit in this paper, we also estimate the economic border between Brooklyn and Queens boroughs in New York City using housing prices.\footnote{ The result is available upon request.} The estimated border is substantially different from the existing administrative border, which was determined in 1931 and cannot reflect the dramatic city development. Interestingly, the estimated border coincides with the Jackson Robinson Parkway and the Long Island Railroad. This finding provides new evidence that local transportation corridors could increase community segregation (cf.\ Ananat11 and Heilmann18).
We list some related works, which could motivate potential theoretical extensions. First, while we focus on the local constant estimation in this paper, one could consider the local linear estimation using the threshold indicator $\mathbf{1}\left[ q_{i}\leq {\Greekmath 010D} _{1}+{\Greekmath 010D} _{2}(s_{i}-s)\right] $ in ((ref)). Although grid search is very difficult in determining the two threshold parameters (${\Greekmath 010D} _{1}$ and ${\Greekmath 010D} _{2}$), we could use the MCMC algorithm developed by Yu19 and the mixed integer optimization (MIO) algorithms developed by LLSS18. Besides the computational challenge, however, the asymptotic derivation is more involved since we need to consider higher-order expansions of the objective function. Second, while our nonparametric setup is on the threshold function ${\Greekmath 010D} _{0}(\cdot )$, some recent literature studies the nonparametric regression model with a parametric threshold, such as $y_{i}=m_{1}(x_{i})+m_{2}(x_{i})\mathbf{1} [q_{i}\leq {\Greekmath 010D} _{0}]+u_{i}$, where $m_{1}\left( \cdot \right) $ and $ m_{2}\left( \cdot \right) $ are different nonparametric functions. See, for example, Henderson17, Chiou18, Yu18, YuLiaoPhillips19, and DelgadoHidalgo2000.
\setcounter{equation}{0}\scalefont{0.96}\baselineskip=14pt