EconBase
← All papers

Optimal Control Variates for Survey Sampling and Causal Inference

Jinglong Zhao

arXiv 15 Aug 2026 · Econometrics

arXiv:2608.15333 · PDF · Extracted main text

Abstract

We propose a family of control variate estimators for variance reduction in design-based survey sampling and causal inference, with and without interference. In these settings, inverse probability weighting (IPW) estimators are widely used, but may have large variance when sampling, treatment, or exposure probabilities are small. Building on the observation that several common estimators, including the Hajek, normalized, and augmented inverse probability weighting (AIPW) estimators, correct the Horvitz-Thompson estimator by canceling part of its randomness, we provide a unified interpretation of these estimators as special cases of a general control variate estimator. We then construct optimal control variates that can further reduce the finite sample variance compared to these common estimators. We parameterize the proposed control variates by their bases and characterize the optimal bases through a stochastic optimization formulation. In survey sampling and causal inference without interference, the optimal bases are characterized by leading eigenvectors of matrices that depend on both the design-based sampling structure and the model-based outcome uncertainty. In causal inference under network interference, the optimal bases solve a nonconvex quadratic optimization problem; we provide a $\frac{1}{2}$-approximate solution and an alternating local search heuristic. We apply the control variate estimators to the Swiss Environmental Panel survey data and the Chinese social network data, and conduct extensive simulations to show that the proposed control variate estimators can achieve substantial variance reduction.

Citation extraction

67
references
126
in-text mentions
67
distinct cited
0
self-citations
57,964
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Cai J, Janvry AD, Sadoulet E (2015) Social networks and the decision to insure1.000103100%
2Quo F, Rudolph L, Gomm S, Wehrli S, Bernauer T (2021) Swiss environmental panel study 2018-2021, wave 1-6, cumulative data1.00053100%
3Ross N (2011) Fundamentals of stein’s method0.92844100%
4Fan K (1949) On a theorem of weyl concerning eigenvalues of linear transformations i0.87462100%
5Zhao J (2024) Experimental design for causal inference through an optimization lens0.84333100%
6Aronow PM, Samii C (2017) Estimating average causal effects under general interference, with application to a social network experiment0.81142100%
7Khan S, Ugander J (2023) Adaptive normalization for ipw estimation0.81142100%
8Horn RA, Johnson CR (2012) Matrix analysis0.73732100%
9Robins JM, Rotnitzky A, Zhao LP (1994) Estimation of regression coefficients when some regressors are not always observed0.73732100%
10Särndal CE, Thomsen I, Hoem JM, Lindley D, Barndorff-Nielsen O, Dale… (1978) Design-based and model-based inference in survey sampling [with discussion and reply]0.73732100%

Showing the top 10 of 67 scored citations.