EconBase
← All papers

Data Synchronization at High Frequencies

Xinbing Kong, Cheng Liu, Bin Wu

arXiv 16 Jul 2025 · Econometrics

arXiv:2507.12220 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Asynchronous trading in high-frequency financial markets introduces significant biases into econometric analysis, distorting risk estimates and leading to suboptimal portfolio decisions. Existing synchronization methods, such as the previous-tick approach, suffer from information loss and create artificial price staleness. We introduce a novel framework that recasts the data synchronization challenge as a constrained matrix completion problem. Our approach recovers the potential matrix of high-frequency price increments by minimizing its nuclear norm -- capturing the underlying low-rank factor structure -- subject to a large-scale linear system derived from observed, asynchronous price changes. Theoretically, we prove the existence and uniqueness of our estimator and establish its convergence rate. A key theoretical insight is that our method accurately and robustly leverages information from both frequently and infrequently traded assets, overcoming a critical difficulty of efficiency loss in traditional methods. Empirically, using extensive simulations and a large panel of S&P 500 stocks, we demonstrate that our method substantially outperforms established benchmarks. It not only achieves significantly lower synchronization errors, but also corrects the bias in systematic risk estimates (i.e., eigenvalues) and the estimate of betas caused by stale prices. Crucially, portfolios constructed using our synchronized data yield consistently and economically significant higher out-of-sample Sharpe ratios. Our framework provides a powerful tool for uncovering the true dynamics of asset prices, with direct implications for high-frequency risk management, algorithmic trading, and econometric inference.

Citation extraction

45
references
78
in-text mentions
45
distinct cited
0
self-citations
22,397
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bollerslev T, Li J, Ren Y (2024) Optimal inference for spot regressions0.87482100%
2Ait-Sahalia Y, Xiu D (2017) Using principal component analysis to estimate a high dimensional factor model with high-frequency data0.87462100%
3Cui L, Hong Y, Li Y, Wang J (2024) A regularized high-dimensional positive definite covariance estimator with high-frequency data0.73732100%
4Hollstein F, Prokopczuk M, Wese Simen C (2020) The conditional capital asset pricing model revisited: Evidence from high-frequency betas0.73732100%
5Aẗ-Sahalia Y, Fan J, Xiu D (2010) High-frequency covariance estimates with noisy and asynchronous financial data0.64422100%
6Andersen TG, Thyrsgaard M, Todorov V (2021) Recalcitrant betas: Intraday variation in the cross-sectional dispersion of systematic risk0.64422100%
7Fan J, Li Y, Yu K (2012) Vast volatility matrix estimation using high-frequency data for portfolio selection0.64422100%
8Kong X, Lin JG, Liu C, Liu GY (2023) Discrepancy between global and local principal component analysis on large-panel high-frequency data0.64422100%
9Onatski A, Wang C (2024) Spurious factors in data with local-to-unit roots0.64422100%
10Recht B, Fazel M, Parrilo PA (2010) Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization0.64422100%

Showing the top 10 of 45 scored citations.