Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,711 characters · 14 sections · 54 citation commands
Transfer Learning for Spatial Autoregressive Models with Application to U.S. Presidential Election Prediction
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 {
} \fi
{\it Keywords:} Spatial autoregressive model, transfer learning, U.S. presidential election, spatial data analysis, high dimensional analysis.
\spacingset{1.9}
In recent years, amounts of Geographic Information Science (GIS) data have been used in election campaigns peters2004gis. Geography is the common denominator of where voters live, how residents engage in the voting process, how officials manage elections, and where elections campaign runs. Based on population data and geographic information, election district boundaries are resettled to draft multiple proposed plans and demonstrate the effects of several approaches. Political campaigns make full use of geographic information and related tools to obtain underlying support from voters, especially the ones from swing states, which are pivotal in determining the outcome of elections. Swing states, also known as battleground states, are crucial because they do not consistently support one party, making their election outcomes unpredictable and highly influential in election results. Shaping voter preference in swing states is an essential problem, especially for prediction. Including spatial information in data analysis brings a novel insight into the prediction. There are at least two challenges in using spatial information for predicting U.S. presidential elections: one is leveraging spatial information effectively to predict election outcomes with considering the impact of spatial factors, and the other is the limited data from swing states. Finding ways to use information from other states to analyze spatial factors and predict the final result is crucial to overcoming these challenges.
Spatial dependence among spatial units exists widely in various fields including economics anselin1988spatial, finance grennan2019dividend, political science peters2004gis, ecology verhoef2018spatial and climate sciences okunlola2021spatial and so on. The spatial autoregressive (SAR) model cliff1973spatial,ord1975estimation,cliff1975model and its variants are extensively applied in numerous empirical studies concerning spatial competition and spatial spillover effects. For instance, the SAR model is employed to investigate the crime rates in 49 areas of Columbus, Ohio anselin1988spatial, budget spillover effects and policy interdependence case1993budget, cigarette demands in various states in the U.S. baltagi2004prediction, agricultural land prices in Northern Ireland kostov2009spatial, the mutual influence on the stock returns of companies with the same headquarters location pirinsky2006does and the group effect of the dividend policy of listed companies grennan2019dividend.
When the sample size of the target data is limited, it is possible to obtain less precise estimates for most statistical models. This issue is particularly relevant in spatial econometrics, where sample sizes can often be quite limited. For instance, in the study investigating crime rates in Columbus, Ohio, the sample size is restricted to just 49 anselin1988spatial. This limitation underscores the necessity for additional research aimed at leveraging information from similar datasets to elevate estimation performance. Consequently, there is a pressing need for the exploration of methods to improve the estimation performance of the SAR models through the incorporation of information derived from similar source datasets. Transfer learning is a useful approach to incorporate information from source datasets pan2009survey, torrey2010transfer, zhuang2021comprehensive, hu2023optimal. The main idea of transfer learning is that the knowledge acquired in one context can be transferred to improve the performance of estimation or prediction when applied to another, even if the two tasks or domains are not identical. In transfer learning, there are typically two main components: (1) The first one is the target domain, to which the knowledge is transferred. The target domain is of our interest but particularly its sample size is limited. (2) The second one is the source domain, from which the knowledge is transferred. In general, the model is pre-trained using the source datasets.
In the recent literature, bastani2021predicting introduced a two-step estimator that employs high-dimensional techniques to capture biases between the target and source data and has provided an upper bound for the proposed estimators. li2022transfer considered transfer learning for estimation and prediction of a high dimensional linear regression and proposed a novel data-driven Trans-Lasso method for detecting the transferable sources. tian2023transfer introduced a novel entropy-based method for identifying transferable sources and further constructed confidence intervals for each coefficient under the generalized linear model. For more reference on related works, see cai2019transfer, cai2024transfera, cai2024transfer, and li2022transfera, li2022estimation.
In this paper, in order to deal with two aforementioned challenges, spatial dependence and sample limitation, in U.S. presidential election prediction problems, we present a transfer learning framework within the SAR model, which we refer to as tranSAR, designed to improve both estimation and prediction. Our study makes three contributions from methodology, statistical theory and empirical application in the literature of the SAR models. First, in methodology, we introduce a two-stage algorithm ($\mathcal{A}$-TranSAR), consisting of a transferring stage and a debiasing stage, to estimate the unknown parameters when the informative source data indexed by $\mathcal{A}$ are known. If we do not know which sources to transfer, a transferable source detection algorithm is proposed to detect informative sources data based on spatial residual bootstrap to retain the spatial dependence. To the best of our knowledge, this is the first work to introduce the transfer learning to the SAR models. Second, in statistical theory, we establish the theoretical convergence rates for the $\mathcal{A}$-TranSAR estimators under the SAR model and theoretically show that the transfer learning can improve the estimation accuracy with the help of the information from the transferable source data. We also derive the detection consistency of the transferable source detection algorithm to avoid the negative transfer when which sources to transfer is unknown. Third, in the empirical application, we apply our method to predict the election outcomes in swing states in the U.S. presidential election via utilizing polling data from the last U.S. presidential election along with other demographic and geographical data. The empirical results show that our method outperforms traditional spatial econometrics methods.
Here, we emphasize the differences between our work and tian2023transfer. Although our work shares a similar idea of transfer leaning with the previous literature, there are at least two significant differences from both theoretical and algorithmic perspectives. (1) tian2023transfer is based on generalized linear models for independent data. In contrast, our research focuses on spatially dependent data, which brings significant challenges for theoretical guarantees. This is because previous theories heavily rely on concentration inequalities applicable to independent data, and few concentration inequalities have been established for spatially dependent data. To address this, we develope a new theoretical foundation tailored to spatially correlated data. (2) For the algorithmic aspect, the transferable detection method in tian2023transfer relies on sample splitting to identify transferable sources, a technique that is inapplicable for spatial data as spatial data are not independent. To this end, we propose a new method based on spatial bootstrap and establish its theoretical guarantees.
In the rest of our paper, we firstly introduce the real application problem (predicting the U.S. presidential election result) and the associated transfer learning framework for the SAR model in Section (ref). Then, we introduce its estimation and source detection algorithm in Section (ref). In Section (ref), we evaluate the finite sample performance of our proposed method via simulation studies. Further, the application to predict U.S. presidential elections with transfer learning is conducted in Section (ref). Section (ref) establishes the estimation convergence rates of the transfer learning estimators for the SAR model and derive the detection consistency of the transferable source detection algorithm. We conclude the paper in Section (ref). Additional simulation results, theoretical analysis, and all proofs are provided in the supplementary materials.
\paragraph{Notations.} Let \([n] \equiv \{1, \ldots, n\}\). For an index set \(\mathcal{A} \subseteq [K]\) and a sequence of positive integers \(\{n_k, k = 1, \dots, K\}\), \(n_{\mathcal{A}} \equiv \sum_{k\in \mathcal{A}} n_{k}\). For a vector \(\mathbf{x} \in \mathbbm{R}^d\) and \(S\) being a subset of set \([d]\), \(\mathbf{x}_{S}\) is the sub-vector of \(\mathbf{x}\) with index \(S\). The notation \(a \lesssim b\) means there exists a positive constant \(C \text{ s.t. } a \leq C b\). The notation \(a \vee b\) means \(\max(a, b)\), and \(a \wedge b\) means \(\min(a,b)\). The notation \( a_n \asymp b_n\) means that \(a_n/b_n\) converges to a positive constant. For any random sequences \(a_n\) and \(b_n\), \(b_n = O_{\mathbb{P}}(a_n)\) means \(\forall \epsilon >0, \exists M>0 \text{ s.t. } \limsup_n \mathbb{P}(|b_n| \geq M |a_n|) \leq \epsilon \), and \(a_n \ll b_n\) means \(|a_n| = o_{\mathbb{P}}(|b_n|)\), i.e. \(\lim_{n \rightarrow \infty}\mathbb{P}(|\frac{a_n}{b_n}| \geq \varepsilon) = 0\) for any \(\varepsilon >0\). For a vector \(\mathbf{x} = {(x_1, \dots, x_d)}^\intercal \in \mathbbm{R}^d\), we define \(\|\mathbf{x}\|_1 = \sum_{j=1}^{d} |x_j|\), \(\|\mathbf{x}\|_2 = {(\sum_{j = 1}^{d} x_j^2)}^{1/2} \), and \(\|\mathbf{x}\|_\infty = \max_{1 \leq j \leq d} |x_j|\). For two column vectors \(\mathbf{x}\) and \(\mathbf{y} \in \mathbbm{R}^d\), their inner product is \(\langle \mathbf{x}, \mathbf{y} \rangle = \mathbf{x}^\intercal \mathbf{y}\). For an \(n \times n\) matrix \(A\), we define \(\|A\|_F \equiv {(\sum_{i=1}^m \sum_{j=1}^n\left|a_{i j}\right|^2)}^{1 / 2}\), \(\|{A}\|_{\infty} \equiv \max _i \sum_{j=1}^n\left|a_{i j}\right|\), its spectral norm \(\|A\|_2 \equiv\max _{x \in \mathbb{R}^n \backslash\{0\}} \frac{\|A x\|_2}{\|x\|_2}=\sqrt{\lambda_{\max}(A'A)}\), where \(\lambda_{\max}(A)\) is the maximum eigenvalue of \(A\) and \(\lambda_{\min}(A)\) is the minimum eigenvalue of \(A\). For an \(n \times m\) matrix \(A\), \(A_{(i,j)}\) represents the element in its \(i\)-th row and \(j\)-th column. \(A_{(\cdot,j)}\) refers to the \(j\)-th column of \(A\), and \(A_{(i,\cdot)}\) represents the \(i\)-th row of matrix \(A\).
In this work, we focus on predicting the U.S. presidential election results of each county in swing states by incorporating the spatial relationship among counties. We try to predict the election results in 2020 and 2024 elections using the polls from the 2016 and 2020 election U.S. presidential election and some other demographic data and geographic data, respectively. It is worth noting that swing states have undergone historical shifts over time. The process of identifying swing states in previous elections typically involves an assessment of the closeness of the vote margins in each state. This analysis incorporates historical election outcomes, opinion polls, political trends, recent developments since the previous election, and the specific attributes, strengths, or vulnerabilities of the candidates in contention. In Section (ref), we describe the motivation problem and data. Here, we reiterate the challenges we face, namely the utilization of spatial information and data scarcity. To this end, we propose a novel transfer learning framework based on the SAR model aimed at addressing above mentioned two challenges.
\paragraph{Problem:} We aim to handle the challenges of utilizing spatial information and analyzing spatial factors in predicting the U.S. presidential election. From Figure (ref), it is evident that the county-level support rates in the U.S. presidential election display a spatial relationship. Spatial relationships among the data are also prevalent, making it essential to incorporate spatial information into the analysis. In U.S. presidential elections, the distribution of votes exhibits significant spatial dependency, a phenomenon particularly obvious over counties. Specifically, the voting patterns of a county tend to align closely with those of its neighboring counties. This similarity can often be attributed to the shared socio-economic characteristics that arise from geographical proximity, such as industrial structure, income levels, and educational attainment fotheringham2021scale, kim2003spatial. Additionally, whether a county is urban or rural plays a substantial role in this spatial dependency. Urban counties generally show stronger support for the Democratic Party, whereas suburban and rural counties tend to favor the Republican Party. Thus, in analyzing voting distributions, the urban or rural status of a county becomes a critical factor, reflecting not only the geographical distribution of political preferences but also highlighting the political and social divides between urban and rural areas. Recognizing this spatial dependency helps in more accurately predicting election outcomes and in forming strategies for political parties. Thus the first challenge lies in how to incorporate spatial information into transfer learning techniques. \spacingset{1.5}
\spacingset{1.9} Though the method proposed by tian2023transfer provides a valuable framework to transfer learning and has been applied to the U.S. presidential election, their method is for independent data and does not account for spatial information during transfer learning. Furthermore, the second challenge is the issue of data scarcity. It is evident that the average sample size per state (the number of counties) is only around 60. Relying only on data from swing states for spatial analysis could harm the accuracy of the estimates.
\paragraph{Data sources and processing:} The digital boundary definitions of the United States congressional district are available at \url{https://cdmaps.polisci.ucla.edu/}, and the cartographic boundary shape files of counties in the USA are available at the United States Census Bureau {(\url{https://www2.census.gov/geo/tiger/})}. The spatial weight matrix is constructed using above two geographic datasets. Here we apply the “queen” contiguity-based spatial weights and row-normalized. The “queen” criterion defines neighbors as spatial units sharing a common edge or a common vertex. The response is the difference of support at the county level for presidential candidates from both parties in 2016 and 2020, available at \url{https://github.com/tonmcg/US_County_Level_Election_Results_08-20}. The county-level demographic information forms the predictors, collected at \url{https://data.census.gov/}. Among all states and the federal district, we exclude Alaska, Hawaii, and Washington, D.C\@.. In these 48 states, there are 3107 counties and we have 142 county-level predictors.
In order to deal with the two aforementioned challenges in U.S. presidential election prediction problems, we present a transfer learning framework within the spatial autoregressive (SAR) model. For a pre-specified swing state as the target data, we consider the following SAR model,
where the difference of support $Y^{(0)}$ is an $n_0\times 1$ vector of spatially correlated dependent variables, $W_{l}^{(0)}$ is an $n_0 \times n_0$ known non-stochastic spatial weight matrix for each $l \in [p]$, the county-level demographic information $X_{j}^{(0)}$ is an \(n_0 \times 1\) vector of exogenous regressors independent of the error term $V^{(0)}$ for each $j \in [q]$, \(V^{(0)}\) is an $n_0 \times 1$ vector of independently and identically distributed (i.i.d.) random variables with mean \(0\) and variance $\sigma^{2}>0$. Let $\lambda_0 = {(\lambda_{1 0},\dots,\lambda_{p 0})}^{\intercal}$, $\beta_0 = ( \beta_{10}, \dots, \beta_{q 0})^{\intercal}$ and \(\theta_0 \equiv \) \({({\lambda_0}^{\intercal}, {\beta_0}^{\intercal})}^{\intercal}\).
Similar to general transfer learning settings as in bastani2021predicting, li2021targeting, tian2023transfer, in addition to the target study (ref), additional samples are obtained from $K$ auxiliary studies (also termed as source datasets) from other states, where \(K\) is assumed to be fixed. For the $k$-th source dataset, we also consider the SAR model,
where \(Z^{(k)} \equiv (Y^{(k)}, X_{1}^{(k)},\dots, X_{q}^{(k)})\) is the $k$-th source dataset of sample size \(n_k\), $W_{l}^{(k)}$ is an $n_k \times n_k$ non-stochastic spatial weight matrix for each $l \in [p]$, and $V^{(k)}$ is an $n_k \times 1$ vector of i.i.d. random variables with mean zero and variance $\sigma^{2}_k$, independent of the covariates \(X_{j}^{(k)}, j \in [q]\). Let \(\lambda_{0}^{(k)} = {(\lambda_{1 0}^{(k)}, \dots, \lambda_{p 0 }^{(k)})}^{\intercal}\), $\beta_0^{(k)} ={(\beta_{1 0}^{(k)}, \dots, \beta_{q 0}^{(k)})}^{\intercal}$ and \(\omega^{(k)}_0 \equiv {({\lambda_{ 0 }^{(k)}}^{\intercal}, {\beta_0^{(k)}}^{\intercal})}^{\intercal}\) for the $k$-th source dataset. Let \(\mathbf{X}^{(k)} \equiv (X_{1}^{(k)}, \dots, X_{q}^{(k)})\), \(\mathbf{W}^{(k)} \equiv ({W_{1}}^{(k)}, \ldots,{W_{p}}^{(k)})\), and \(\mathfrak{D}^{(k)} \equiv (Y^{(k)}, \mathbf{X}^{(k)}, \mathbf{W}^{(k)})\), \(k = 0, \dots, K\). In transfer learning, a similar auxiliary model to the target model is called an informative model, and the “similarity” is characterized by the difference between $\omega_0^{(k)}$ and $\theta_0$, i.e.\ \(\delta^{(k)} \equiv \) \(\theta_0 - \omega^{(k)}_0\). Given a tolerance level $h>0$, the index set of the informative auxiliary samples is defined as
where \( \|\delta^{(k)}\|_1 \) is called the transferring level of the \(k\)-th source dataset. In this context, the value of $h$ is used to balance the bias and the variance introduced by transfer learning. When $h$ is large, more labeled datasets are used, which reduces the prediction variance but increases the bias due to source datasets, and vice versa. We will discuss the value of $h$ in the conditions for theoretical results in Section (ref).
In this section, we propose a two-step transfer learning procedure under the SAR model when the index set \(\mathcal{A}_h\) of the informative auxiliary samples is known, denoted as \( \mathcal{A} \)-TranSAR. In the first transferring step, we consolidate information from various sources by pooling all available auxiliary data, yielding a preliminary estimator which is typically biased, since \(\omega^{(k)}_0 \neq \theta_0\) in general, although they are close. In the second debiasing step, we correct the estimation bias by incorporating the target data via regularization. The details of the two-stage estimation procedure are present in the following.
\paragraph{The transferring stage} The first-step estimator, which is also called pre-training model in the transfer learning literature, with tuning parameter \(\lambda_{\omega}\) is defined as
where $\operatorname{\mathit{\ell}}_1 ( \cdot \mid \omega)$ is a generic loss function with parameter \( \omega \), and \(n_{\mathcal{A}} \equiv \sum_{k\in \mathcal{A}} n_{k}\) is the total sample size of all informative source data. For the SAR model, the estimator could be the two-stage least squares (2SLS) estimator kelejian1998generalized, lee2003best, generalized method of moments estimator (GMM) kelejian1999generalized,lee2007gmm,liu2010gmm, quasi-maximum likelihood estimator lee2004asymptotic and so on. The main requirement for the first step estimator is a suitable convergence rate discussed in detail in Section (ref). Here the tuning parameter \(\lambda_{\omega}\) could be zero in low dimensional cases. When the dimension of parameters is greater than the sample size in high dimensional cases, the penalized SAR estimation with nonzero \(\lambda_{\omega}\) could produce a sparse estimatorhiggins2023shrinkage. See also zhang2018spatial for spatial weight matrices selection with model averaging, and tibshirani1996regression, fan2001variable, zhang2010nearly, wang2011random, su2017false for more detail in regularization method.
\paragraph{The debiasing stage}
In the second step, the bias term estimator \(\widehat{\delta}\) with tuning parameter \(\lambda_{\delta}\) using the target data is defined as
Then, the bias-corrected \(\mathcal{A}\)-TranSAR estimator is defined as
The two-stage transfer learning procedure under the SAR model is summarized in Algorithm (ref). \spacingset{1}
\spacingset{1.9}
The \(\mathcal{A}\)-TranSAR estimator requires the true informative index set \(\mathcal{A}\) to be known correctly. However, in most empirical applications, the true informative set \(\mathcal{A}\) is unknown. Transferring {adversarial auxiliary} samples may not improve the performance of the target model and could even cause worse results. The effect of adversarial auxiliary samples is also called negative transfer, meaning that certain inferior results are caused by uninformative source data torrey2010transfer,tian2023transfer.
An intuitive approach might involve seeking a metric of similarity between each source and the target data. We compute the initial loss based on the target dataset. Subsequently, we combine the target dataset with each source dataset individually to estimate and compute the corresponding losses. If a source dataset is similar to the target dataset, we anticipate two loss values to be comparably close. Hence, we select a critical value to identify which source datasets are informative. For spatial datasets characterized by a network structure, data splitting used in tian2023transfer can not be directly applied since it could disrupt this network integrity. As a result, we introduce the spatial residual bootstrap approach as an alternative. This idea is illustrated in details as follows.
First, we compute an initial estimator \(\widehat{\theta}_{ini} = {({\widehat{\lambda}_{ini}}^{\intercal}, {\widehat{\beta}_{ini}}^{\intercal})}^{\intercal}\) using the target samples by certain estimation methods, such as 2SLS, QMLE or GMM, etc. Then, we obtain the residuals $ \hat{e} = Y^{(0)} - \mathbb{X}^{(0)} \hat{\theta}_{ini}, $ where \({\mathbb{X}}^{(0)} = ({W}_1^{(0)}{Y}^{(0)}, \ldots, {W}_p^{(0)}{Y}^{(0)} , \mathbf{X}^{(0)})\). Similar to residual bootstrap, we generate three independent copies, \(\widehat{e}^{(1)}\), \( \widehat{e}^{(2)}\) and \(\widehat{e}^{(3)}\) of sample size \(n_0\) from the empirical distribution of \(\widehat{e}\). Then, we generate three-fold bootstrap samples \(Y^{(0,r)}\) and \( \mathbf{X}^{(0,r)} \) by
for \(r = 1, 2, 3\). Let \(Z^{(0,r)} = (Y^{(0,r)}, \mathbf{X}^{(0,r)})\) for \(r = 1,2,3\). Let \(Z^{(0,-r)}\) be the concatenated bootstrap samples without the $r$-th sample, for example, $Z^{(0,-1)} = {({Z^{(0,2)}}^\intercal, {Z^{(0,3)}}^\intercal)}^\intercal$. Define \(\mathbf{X}^{(0, -r)}, \mathbb{X}^{(0, r)}\) and \(\mathbb{X}^{(0, -r)}\) similarly.
Second, for each \(r = 1, 2, 3\), we fit the model in (ref) on the connected bootstrap samples \(Z^{(0,-r)}\) and use the resulting estimator \(\widehat{\theta}^{(0, r)}\) to compute the loss value on the remaining bootstrap sample \(Z^{(0,r)}\), that is
Then, we compute the average of three loss values, \(\mathcal{L}^{(0)}=\sum_{r=1}^{3}\mathcal{L}^{[r]}(\widehat{\theta}^{(0, r)})/3\), which serves as the baseline loss value for the target dataset. We can also compute the standard deviation of three loss values, denoted by
which measures the dispersion of bootstrapped loss values $\mathcal{L}^{[r]}(\widehat{\theta}^{(0, r)})$ relative to their average $\mathcal{L}^{(0)}$. Next, for each \(r = 1, 2, 3\), we combine \(Z^{(0,-r)}\) with each $k$-th source dataset, fit the model in (ref) on the newly combined data and compute the loss denoted by \(\mathcal{L}^{(k,r)}\) on \(Z^{(0,r)}\). Similarly, the average loss for the $k$-th source dataset over three-fold bootstrap samples is \(\mathcal{L}^{(k)}\). Consequently, the difference between \(\mathcal{L}^{(k)}\) and \(\mathcal{L}^{(0)}\) provides a metric of the similarity between the $k$-th source dataset and the target data. Source datasets whose corresponding loss differences below certain threshold are recognized as transferable informative data. Following the idea of tian2023transfer, we choose $\hat{\sigma} \vee 0.01$ as the detection threshold. Then, the estimated transferable source index set \(\widehat{\mathcal{A}}\) is defined as
Once the transferable set $\mathcal{A}$ is estimated, we can apply the $\mathcal{A}$-TranSAR in Algorithm (ref) to obtain the transfer learning estimator for the SAR model, denoted by \(\widehat{\theta}_{\widehat{\mathcal{A}} \textit{-TranSAR}}\), which is termed as the TranSAR estimator. The procedure described above is summarized in Algorithm (ref).
\spacingset{1}
\spacingset{1.9}
In this section, we design several simulation experiments to investigate the performance of the proposed method (denoted as TranSAR). We compare the finite sample performance of our estimator with the classical 2SLS estimation (which is fitted only using the target dataset, denoted as SAR), the ${\mathcal{A}}$-TranSAR estimator (which is fitted based on the true informative index set $\mathcal{A}$, denoted as Oracle TranSAR), pooled-estimator (which is fitted using all the source datasets as the transferable sets, denoted as $[K]$-TranSAR), and the generalized linear model transfer learning algorithm (denoted as glmtrans) proposed by tian2023transfer. The \textit{glmtrans} ignores the spatial information and can be considered as a baseline estimator to show that it is necessarily important to consider the spatial correlations if they exist.
We generate the data according to Models (ref) and (ref). In these models, $n_{0}=256,n_{1}=n_{2}=\cdots=n_{K}=100,p=1,q=200,K=20$. The covariates $\mathbf{x}_{i}^{(k)}$ are generated by multivariate normal distribution with covariance matrix $\Sigma\in\mathbb{R}^{q\times q}$, whose the $(j,j')$-th entry is denoted as $\Sigma_{j,j'}$ for all $0\leq k\leq K$. We consider three designs for $\Sigma$: (1) $\Sigma_{j,j'}=1$ for $j=j'$ and $\Sigma_{j,j'}=0$ otherwise, (2) $\Sigma_{j,j'}=0.5^{|j-j'|}$, (3) $\Sigma_{j,j'}=0.9^{|j-j'|}$. We consider two different distributions for the i.i.d.\ error terms: (1) the standard normal distribution $N(0,1)$, and (2) the $t_{2}$ distribution. The spatial units are assumed to be located on a square grid, and we set up several candidate spatial weight matrices as follows: (1) the first \(W_{1}\) is the spatial weight matrix where spatial units interact with their neighbors on the left and right sides; (2) the second \(W_2\) is the spatial weight matrix where spatial units interact with their neighbor above and below; (3) the rest of the spatial weight matrices are assumed to be matrices where spatial units interact with their second-nearest, third-nearest, and so on, denoted as $W_{3},\ldots, W_{N}$, where $N<\sqrt{n_{0}}$. In each dataset, we randomly draw spatial weight matrices from these candidate matrices. The weight matrices ${\{W_{k,l}\}}_{k\in\{0\}\cup[K],l\in[p]}$ are drawn randomly from $\{W_{1},\ldots,W_{N}\}$ without replacement. In Models (ref) and (ref), we consider coefficients settings as follows. In the target Model (ref), $\lambda_{0}$ $=0.4$ and $\beta_{0}$ $=(\mathbf{1}_{3}^{\intercal},\mathbf{0}_{q-3}^{\intercal})^{\intercal}\in\mathbb{R}^{q}$. In Model (ref),
{where $\mathcal{H}_{k}$ is a random subset of $[q]$ with $|\mathcal{H}_{k}| = H$ if $k\in\mathcal{A}$ (i.e., \(H\) is the number of different coefficients in the informative set), or $|\mathcal{H}_{k}|=\frac{q}{2}$ if $k\notin\mathcal{A}$. So, on the informative datasets, the differences of different \(\beta_{j0}^{(k)}\)'s are 0.05; on uninformative datasets, the ones of different \(\beta_{j0}^{(k)}\)'s are 2.}
We compute the root mean squared errors (RMSE),
where \( {\widehat{\theta}}^{[r]}\) is the estimator of the \(r\)-th replication, based on the 2SLS loss with $q=200$ and display them under different settings in Figure (ref). In order to demonstrate the advantages of transfer learning, we let the size of informative auxiliary datasets $|\mathcal{A}|$ increase from $0$ to $20$ to show the decreasing trend in RMSE\@. As the number of informative sets increases, compared to the classical 2SLS estimation and the glmtrans method, we can see that $\widehat{\mathcal{A}}$-TranSAR and ${\mathcal{A}}$-TranSAR can significantly reduce the RMSE, resulting in more precise estimation. Additionally, we can see that the $\widehat{\mathcal{A}}$-TranSAR performs almost the same as the oracle transfer learning method $\mathcal{A}$-TranSAR. That means, our proposed detection method is able to select informative sets with a high probability. Moreover, from Figure (ref), when the number of informative sets is small, the negative transfer effect of the pooled method, $[K]$-TranSAR becomes apparent. Hence, it is necessary to select informative sets accurately. The same conclusion applies to all three types of design matrices and the three different settings for $H$. Detailed and comprehensive simulation results are compiled in Table (ref) for covariates designs type 1.
\spacingset{1}
\spacingset{1.9}
In this section, we apply the proposed method in Algorithm (ref) to empirical data from the U.S. presidential election. As shown in Section (ref), the response variable \(Y^{(k)}\) is the difference of support rates in the county level, and covariates \(\mathbf{X}^{(k)}\) include demographic and geographica data, such as including total population, total housing units, poverty people population, etc. With the observed spatial weight matrix \( \mathbf{W}^{(k)}\), each swing state is regarded as a target with \(k=0\), and all the other states are considered as sources \(k \in [47]\). The target model (ref) and sources models (ref) could formulated as
We run the proposed estimation method, TranSAR, using the polls from the 2016 U.S. presidential election and some other demographic data with geographic information. We also estimate the SAR model with Lasso penalty (due to high dimensional covariates) using only the target datasets (denoted as SAR) to compare the empirical performance. Due to the inherent stochastic nature of the resampling procedure in the algorithm, we conduct multiple replications (20 times) for both methods to obtain a stable result. The average prediction RMSE of the county-level vote results of the county-level 2020 U.S. presidential election is compared at Table (ref). For all swing states, the proposed TranSAR method performs better than the traditional SAR in terms of RMSE, which demonstrates the necessarity of transfer learning in the SAR framework.
Moreover, we also consider the practical rules of the U.S. presidential election, which operate on a “winner-takes-all” basis. In this context, the winner party secures all the state's electoral votes. Given the “winner-take-all” framework of the U.S. presidential election, we focus on predicting the winner party in each state. Leveraging the county-wise population, we compute the final difference of support in the state-level endorsement for presidential contenders belonging to the respective parties. We calculate the rate of support for state-level elections using the following formula, {
} Then, we compare the prediction rate of state-level support and the true one in $2020$ election, showing in Table (ref). It demonstrates that the transfer learning under the SAR framework is able to effectively improve the state-level prediction.
The calculated ratios of correct predictions out of the 20 election result forecasts are collected in Table (ref). As shown in Table (ref), the method we proposed exhibits better performance, especially in predicting election outcomes in Michigan and Minnesota. Simultaneously, it is noteworthy that for the election predictions in Georgia, neither of the methods works well. It can be explained by the historical context: from $1972$ to $2016$, Georgia consistently supported the Republican Party, except for Democratic candidates from the southern region. However, the state has become an increasingly competitive battleground. In 2020, the Democratic candidate, Joe Biden, secured a victory over Donald Trump by a narrow margin of $0.2\%$. This marks a significant departure from the history, as the Democrats secured the presidential election with a slim advantage. Evidently, Georgia has become an outlier in this regard. It is evident that predicting the ultimate election outcome for a single state is significantly more challenging than forecasting the rate of state-level support. Nevertheless, we have made commendable progress in this regard.
Next, we attempt to visualize transferable states, conducting a statistical analysis of the results after multiple repetitions. For a swing state, we record the frequency of occurrences for all transferable states and depict states with more than 50% occurrences on the map in green color. As observed in Figure (ref), some transferable states exhibit geographical information signals. However, for some targets, the spatial or geographical correlations may not be readily apparent. This is a commonplace occurrence. With the development of the economy, culture, and scientific technology, regional interconnections are no longer confined solely to spatial proximity. Latent economic and political affiliations become pivotal factors.
In addition, to provide a comprehensive demonstration of the superiority of the proposed TranSAR method, we treat each state as the target once a time. Subsequently, we conduct both the tranSAR and SAR procedures, and the outcomes have been summarized in Supplementary Materials. In conclusion, our tranSAR method consistently outperforms the traditional SAR method in the majority of cases, yielding superior prediction results.
Motivated by public interest, we also predict the outcome of the 2024 U.S. presidential election for all states using demographic and geographical data in 2022, based on the trained tranSAR model from the 2020 U.S. polling data along with other demographic and geographical data. We use the 2022 demographic data as covariates to predict county-level support rates. And we use the 2020 total vote counts of county-level as weights to calculate state-level support rates (because in the U.S., the “winner-takes-all" system weights each county by its vote count). We focus on the predictions of election outcomes for eight “swing states” following the rule based on whether there was a change in the supported party over the past three elections. Therefore, we select {Arizona, Georgia, Florida, Pennsylvania, Michigan, Wisconsin, Ohio, and Iowa} as swing states. For non-swing states, we use the results from their last three presidential elections as the outcome for this election, while for swing states, we use our tranSAR method for prediction. Considering the randomness inherent in the algorithm, we conduct multiple rounds of analysis using our {TranSAR} procedure. In these analyses, if the Democratic Party receive support in more than half of the iterations, the swing state was classified as supporting the Democratic Party; otherwise, it was classified as supporting the Republican Party. We obtain the supporting party for each state (see Figure (ref)) and calculate the final support vote counts combined with the electoral votes. The final prediction result shows that the Democratic Party receive {309} electoral votes, exceeding the threshold of 269 votes. Therefore, we predict that the outcome of the 2024 U.S. presidential election will favor {the Democratic Party}.
In this section, we will establish an estimation convergence rate for the proposed $\mathcal{A}$-TranSAR estimator by assuming $\underline{n}\equiv\min_{k\in[K]\cup\{0\}}n_{k}\rightarrow\infty$. Subsection (ref) is for general results and Subsection (ref) is based on the 2SLS estimation method. Finally, we prove the consistency of the transferable source detection algorithm in Subsection (ref).
For theoretical analysis, we define the population form for the above estimator as
where $\alpha_{k}=\frac{n_{k}}{n_{\mathcal{A}}}$ and $\bar{\mathit{\ell}}_{1}^{(k)}(\omega)$ is the corresponding population objective function for dataset $\mathfrak{D}^{(k)}$. Under some regularity conditions, the first-stage estimator in Eq. (ref) is consistent, i.e., $\hat{\omega}-\omega^{\mathcal{A}}\xrightarrow{p}0$. Analogously, define the “true parameter” based on the single source $k$ as
In this section, we verify the requirements in the above theorem for the 2SLS estimation and simplify the convergence rate of the \(\mathcal{A}\)-TranSAR estimator. As in Eq. (ref), let $\operatorname{\mathit{\ell}}_{1}(\cdot\mid\omega)$ be the loss function of the 2SLS. To elaborate, let $Y^{\mathcal{A}}$ denote the stack consisting of $Y^{(k)}$ for all $k\in\mathcal{A}$, i.e., $Y^{\mathcal{A}}={({Y^{(k)}}^{\intercal},k\in\mathcal{A})}^{\intercal}$. $\mathbf{X}^{\mathcal{A}}$ and ${\mathbb{X}}^{\mathcal{A}}$ are defined similarly. Then the objective function (ref) could be rewritten as
where $\widehat{\mathbb{X}}^{(k)}$ is defined similarly as $\widehat{\mathbb{X}}^{(0)}$ in Subsection (ref), and $\widehat{\mathbb{X}}^{\mathcal{A}}$ is their stack. Here the IV matrices are $Q^{(k)}$ and $Q^{\mathcal{A}}$. We impose several similar conditions on $\mathfrak{D}^{(k)}$ ($k\in\mathcal{A}$) in the following.
Under Assumptions (ref) and (ref), the convergence rate of the \(\mathcal{A}\)-TranSAR estimator is summarized in the following Theorem (ref).
Theorem (ref) indicates that if $h\ll s\sqrt{\log q/n_0}$ and $n_{\mathcal{A}}\gg n_0$, the \(\mathcal{A}\)-TranSAR estimator is better than the classical 2SLS estimator based on only the target data, which demonstrates the usefulness of transfer learning in the SAR models.
Finally, we will show that our proposed transferable source detection algorithm is consistent even for a high dimensional model, i.e., $q\gg n_{0}$.
Next, we establish the detection consistency of $\widehat \mathcal{A}$ in the following theorem.
To more accurately explore the spatial factors influencing swing states, we propose a method to enhance the estimation of spatial model in swing states by transferring knowledge from other states. We introduce a transfer learning method within the framework of SAR models, aimed at addressing spatially dependent data of small sample size. The empirical application of this methodology, specifically in predicting U.S. election results in swing states using demographic and geographical data, demonstrates its potential to enhance traditional spatial methods. Based on our comprehensive analysis and the application of the tranSAR model, our prediction for the 2024 U.S. presidential election is a victory for the Democratic party. Future research could focus on developing tests to detect informative sets and improving the accuracy of coefficient estimation within the transfer learning framework.