Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
52,621 characters · 12 sections · 54 citation commands
Correcting Nonresponse Bias Using Panel Data on Data Requests and Responses
Researchers, governments, and private companies routinely request voluntary data submission, often more than once, from samples of populations of interest. These data gathering efforts usually only succeed with a portion of subjects. For example, this pattern occurs in experiments with delayed measurement of outcomes, such as those investigating effects of masks on COVID 19 transmission abaluck2022impact, effects of education on STIs and fertility duflo2015education, and effects of neighborhoods on health and safety katz2001moving. The same pattern arises in survey data gathered by universities requesting student course reviews, by political pollsters seeking to predict election results, by private businesses seeking feedback on product satisfaction, and by governments seeking information from individuals, businesses, and nonprofit institutions. \footnote{For instance, the Census Bureau spiers2022strategic, the Bureau of Labor Statistics nlsy97retention, and the Internal Revenue Service irscp59 all document procedures for making multiple attempts to elicit responses from surveyed individuals for data they collect.} If the individuals who voluntarily provide data are not representative of the population as a whole, inferences drawn from the data about the real world can be misleading.
An insight from heckman1976common,heckman1979sample is that variation in data observation rates between otherwise similar subpopulations can be used to correct for selection bias. To perform such a correction, it is necessary to retain data on nonrespondents along with respondents. Our core insight is that response behavior is observed for each subject each time they are sent a data request, rather than once. Retaining all such observations, at odds with standard practice, naturally provides variation in data observation rates between otherwise similar subpopulations. Briefly, we propose that researchers retain panel data on subjects' responses at the time of each data request so that cumulative request counts can be used as instrumental variables to correct for nonresponse bias. \footnote{ Methods that leverage instruments include parametric methods such as those of heckman1979sample and the extension to binary outcomes of van1981demand, as well as nonparametric bounding methods such as those of manski1990nonparametric, manski1994selection, and manski2000monotone. }
We join behaghel2015please and dutz2021selection in ordering subjects' responsiveness by their response timings to correct for nonresponse bias. Our primary distinction from their methods is to use response and request timing to construct a panel dataset containing subjects' responses as of each data request, rather than using response timing to order responsiveness between subjects in cross-sectional data. Constructing a panel dataset has two advantages. First, it makes explicit the assumption that potential responses are unaffected by time, which is made implicitly when potential responses are modeled cross-sectionally. Second, it converts response timing data into a format that accommodates existing nonresponse corrections that leverage instruments, as well as accommodating additional instruments that may or may not vary over time, such as randomized response incentives, without requiring revisions to estimation procedures.
Our main result is a proof that local average responses (LARs) for individuals who comply with each data request are identified from the joint distribution of responses and response rates for all requests. Identification of multiple local average responses is sufficient to estimate or bound population averages of requested variables using existing estimation methods that correct for selection bias, which are united by the marginal treatment response framework of heckman2005structural,heckman2007econometric. We establish identification under assumptions that resemble those of imbens1994identification and angrist1996identification. Identification follows regardless of whether requests are randomized or sent uniformly to all subjects if potential responses are time-invariant, and identification fails regardless of randomization if potential responses vary systematically with time.
In Section (ref), we demonstrate the use of nonrandom request instruments by estimating the gender gap in entrepreneurial career intentions among undergraduates using survey data from the University of Wisconsin-Madison. Estimating a parametric van1981demand model for binary outcomes, we find an 18-percentage-point nonresponse-corrected male-female gap in entrepreneurial intentions. We find statistically insignificant evidence of positive selection for men and negative selection for women driving the 20 percentage point uncorrected gap among respondents. This model is identified with two requests, so we also implement and pass a “predictable trends” overidentification test of the null that later requests have no direct effect on the outcome, supporting the validity of our parametric and nonparametric assumptions in this application. In Supplemental Appendix (ref), we estimate labor market outcomes using a synthetic version of the Norway in Corona Times survey constructed to match estimates reported by dutz2021selection. The 95% confidence intervals of our estimated population means cover ground truth averages from administrative data reported by these authors for five out of six survey variables, with respondent averages differing significantly from the ground truth for all variables.
We consider a panel of $N$ subjects indexed by $i$ who are observed in time periods $t=0,1,...,T$, with data collection commencing when $t=1$ and concluding when $t=T$. Subject $i$ receives data requests for an unobserved time-invariant requested variable, denoted by $Y_{i}^*$, with their observed requests accumulated by the end of time $t$ given by $R_{it}$. In each time period, we also observe their response choice $S_{it}$, with $S_{it}=1$ if they respond during time $t$ and $S_{it}=0$ if they do not, with their response, $Y_{it}$, observed if and only if $S_{it}=1$. We assume $S_{it}=1$ for subject $i$ at most once. Finally, we denote subject $i$'s retained response choice for time $t$ as $\hat S_{it} = \sum_{k=0}^t S_{ik}$, with their retained response for time $t$ given by $\hat Y_{it} = \sum_{k=0}^t Y_{ik}S_{ik}$.
In settings with nonresponse, the sample mean of retained responses in period $T$ among respondents is a consistent estimate of $\mathbb{E}[\hat Y_{iT}|\hat S_{iT}=1]$, with expected bias relative to $\mathbb{E}[Y_{i}^*]$ given by
With random nonresponse and no measurement error, $\mathbb{E}[\hat Y_{iT}|\hat S_{iT}=1] = \mathbb{E}[Y_{i}^*]$, and valid inferences on populations are possible using sample means of observed responses. If these conditions are not met, identification of population averages requires additional information.
We assume that observed variables are realizations of potential outcomes, following rubin1974estimating. First, let $S_{it} = S_{it}(R_{it}) \prod_{j=0}^{t-1} (1- S_{ij}(R_{ij}))$ where $S_{it}(r) \in \{0,1\}$ indicates subject $i$'s willingness to respond to $r$ requests at time $t$. \footnote{ We use this definition in lieu of $S_{it} = S_{it}(R_{it})$ because we assume that subjects respond at most once. } Similarly, let $Y_{it} = Y_{it}(S_{it},R_{it}) = Y_{it}(1,R_{it})S_{it} + Y_{it}(0,R_{it})(1-S_{it})$ with $Y_{it}(1,r)$ denoting the response subject $i$ would give at time $t$ if were to respond given $r$ accumulated requests, with $Y_{it}(0,r)$ missing for all $i$, $t$, and $r$. $S_{it}(r)$ and $Y_{it}(1,r)$ are defined for all $r \in \{0,...,\max(R_{it})\}$, respectively, while they are observed only for the realized $R_{it}$.
Following the marginal treatment response framework of heckman2005structural,heckman2007econometric, we assume that selection bias is driven by dependence between responses and response aversion. Specifically, we define the marginal survey response function $m(u) \equiv \mathbb{E}[Y_i^* | U_{it}=u]$ where $U_{it} \in [0,1]$ denotes the response aversion percentile of subject $i$ at time $t$. We assume that response aversion determines potential response choices such that
for all $r$, where $P(r) = \mathbb{E}[S_{it}(r)]$ gives the willing response propensity for request $r$. We then have
for all $r$ and $r'$.
It follows that we can learn about $m(u)$ if we can estimate $\mathbb{E}[Y_i^*|S_{it}(r)-S_{it}(r')=1]$, which we refer to as the local average response for $r-r'$ compliers. If $m(u)$ is parameterized, as in a heckman1979sample model, point estimation is possible if there are weakly more identified local average responses than parameters. Otherwise, $m(u)$ can be bounded using methods such as those of horowitz2000nonparametric or lee2009training. Local average responses are identified as
if requests are valid instruments for requests, so we proceed by establishing this result with assumptions similar to those used to establish instrument validity by imbens1994identification and angrist1996identification, which were shown by vytlacil2002independence to imply the relationship in ((ref)).
Our goal is to attribute differences in $\hat Y_{it}$ across different values of $R_{it}$ to the average of $Y_i^*$ for compliers who require more requests to induce response. If more requests are systematically sent to subjects or time periods with unrepresentative potential responses, we cannot reliably attribute differences in responses between requests to marginal compliers. We begin by assuming that requests are sent in such a way that subjects' potential responses and potential response choices across time are unrelated to their request histories.
This assumption strengthens the independence assumption of imbens1994identification by disallowing requests from being selectively sent to subjects with nonrepresentative potential responses to that request or any prior request. Randomizing requests and setting $T=1$, as proposed by dinardo2021practical, supports this assumption, with it also holding for $T>1$ with uniform requests if potential responses and response choices are time-invariant on average. Less extreme coarsening of the time index, for instance defining $t$ in weeks instead of minutes, can support this assumption by essentially averaging over temporal idiosyncrasies that occur in continuous time.
We next assume that requests do not directly affect potential responses, allowing us to write $Y_{it}(1) = Y_{it}(1,r)$.
The importance of avoiding experimenter demand effects or priming in surveys is well established stantcheva2023run. Our assumption expresses this as a special case of similar assumptions made in the context of sample selection by heckman1979sample and in the context of treatment effect identification by angrist1996identification. In contrast with intertemporal independence, exclusion is more likely to be violated due to continuous time variation in potential responses and requests if time periods are defined coarsely, as subjects receiving different numbers of requests in the same discrete time period may receive them at different points in continuous time.
With independence and exclusion, differences in responses between requests can only be explained by differences in responses between distinct complier groups, defined by common values of $S_{it}(r)$ for all $r$. Our next assumption restricts the possible number of these groups by imposing a natural ordering on them.
Our assumption strengthens that of imbens1994identification by imposing a sign restriction on the effects of requests on response choices, time-invariance on subjects' potential response choices, and intuitive restrictions on the accrual of requests and responses over time. Our monotonicity assumption implies that $S_{i}(R_{it}) = \hat S_{it}$, which enables identification of $P(r)$ and $\mathbb E[Y_{it}(1,R_{it})|S_i(R_{it})=1]$ from average retained response choices and retained responses. Relatedly, we assume that requests strictly increase response rates, slightly strengthening the first part of our monotonicity assumption.
As in treatment effects applications, relevance can be tested by estimating $P(r)$ for all $r$, with $P(r)=\mathbb{E}[\hat S_{it}|R_{it}=r]$ if intertemporal independence and intertemporal monotonicity hold.
Independence, exclusion, and monotonicity allow us to determine local average responses for compliers to each request. We also assume that responses are free of systematic measurement error.
This assumption is weaker than $Y_{it}(1,r) = Y_i^*$, which is common in the related literature, because it allows for idiosyncratic shocks that cancel out within complier groups, such as if students' potential responses regarding career intentions change over time in response to new information, such as grades on assignments. Intertemporal independence and exclusion rule out systematic time-varying measurement error and measurement error caused by requests, respectively, but they do not rule out systematic inaccuracies within or across complier groups. For example, if less conscientious subjects give systematically inaccurate responses to late requests, nonresponse corrections using request instruments will tend to replicate similar inaccuracy in their predictions for non-respondents.
Our main result proves that local average responses are identified under conditions 1–5.
Our result establishes conditions under which data requests are valid instruments for retained responses in data where subjects' requests and responses are observed multiple times. It follows that established methods that use instrumental variables to address selection bias, such as that of heckman1979sample, can be used in settings without explicitly randomized response incentives. The insight underlying these methods is that differences in responses between otherwise-similar subjects with different response rates can be explained by dependence between potential responses and response aversion. A more thorough review of these methods is provided by dutz2021selection, along with a generalized model of survey aversion with heterogeneous effects of instruments on responsiveness.
We use our method to estimate the gender gap in entrepreneurial aspirations among undergraduate students. There is a large gender gap in entrepreneurship among working-age adults aldrich2005entrepreneurship, perhaps driven at least in part by a gender gap in early-stage funding for new ventures canning2012women,greene2003women. Our paper contributes to this literature by investigating pre-market gender gaps in entrepreneurial career interest, which may be more likely to develop due to intrinsic interest gaps or pre-market discrimination rather than labor market discrimination. This application showcases the importance of researcher judgments in defining the time index, adjusting for nonrandom variation in requests, and testing modeling assumptions in cases where there are more requests than are needed for identification. In Supplemental Appendix (ref), we demonstrate our method's accuracy by showing that nonresponse-corrected estimates are closer to ground truth values from administrative data than respondent averages for the Norway in Corona Times Survey, using survey and ground truth averages reported by dutz2021selection.
We use a survey of the entrepreneurial intentions of the undergraduate population at University of Wisconsin–Madison that was implemented every fall from 2015 to 2022, as well as in spring 2020 and 2021. Students could choose between “Yes,” “No,” or “I Don't Know,” when responding to a question regarding whether they intend to pursue a career in entrepreneurship. We code an “Intention” variable equal to 1 if the student answered “Yes” and 0 if they answered “No” or “I Don't Know.” We observe term-specific survey request stratum indicators for groups of students who are sent requests at the same time through survey software. We observe timestamps for each request made to each stratum and for each student's response if they respond. Students who opted out of the survey in a prior year have no stratum indicators and no request or response timestamps in the raw data, whereas students who respond to early requests have stratum indicators and request timestamps for all intended requests in raw data, though they did not receive additional actual requests after responding.
Overall, 103,536 unique students and 333,201 total student-terms were surveyed. We merge the survey data to administrative records containing information on students regardless of whether they responded to the survey. Business Major and STEM Major are set to one if the student has any major in the term on the business school's list of majors or on the list of STEM majors provided by ice2020, respectively. GPA refers to the student's cumulative grade point average that varies across terms, ACT Math is the average of students' ACT math test scores on file, and ACT Verbal is the average of the sum of English and Reading portions of the ACT test. \footnote{We convert SAT scores into ACT scores for students without ACT scores using the 2018 ACT/SAT concordance tables provided by act.org.} Female, Racial Minority, and International are binary indicators for female gender from binary gender information in raw data, listing a race other than white, and having a recorded country of origin other than the U.S., respectively. We drop 6,075 student-terms with missing GPAs (students who dropped out in their first term) and 33,296 observations with missing ACT scores, leaving us with 290,588 student-terms and 82,000 unique students. There were 42,902 total survey responses, with a response rate of approximately 14% for men and 16% for women. Sample means of key variables broken down by response timing are provided in Table (ref).
In the raw data, men's entrepreneurial intention rate drops from 37.8% among early respondents to 36.6% among late respondents, while women's entrepreneurial intention rate rises from 16.9% among early respondents to 18.0% among late respondents. Most of the other variables we review do not exhibit any strong pattern, with the exception of racial minority status. We find that racial minorities are underrepresented among survey respondents, representing 23.5% of early respondents, 27.4% of late respondents, and 29.3% of nonrespondents.
Our results from Section (ref) involve panel data wherein subjects are observed multiple times as data requests are made, so we construct such a panel for each term. We index terms/semesters by $s$, with retained responses, retained response choices, and accumulated requests denoted by $\hat Y_{ist}$, $\hat S_{ist}$, and $R_{ist}$, respectively. In our setting, students who opted out of the survey in a prior term and students who respond early in a given term do not receive subsequent requests. It follows that actual requests almost surely violate the independence assumption in Section (ref), as students receiving each number of requests differ systematically. We address this by randomly assigning opt-outs to strata ex post and assuming $S_{ist}(r)=0$ for all $t$ and all $r$ for such students, then imputing intended request timestamps for each student using the request timestamps of nonrespondents in their stratum.
For each student-term in a stratum, we produce observations for $t$ in $0,1,...,R_i$ such that time periods are defined as intervals in which a subject has received a given number of intended requests, with $R_{ist}=t$ for all $i$, $s$, and $t$. Independence and monotonicity require that the times at which requests were made had representative potential responses and that subjects had sufficient time to respond between requests, respectively. We report request timestamps in Appendix (ref), revealing that the shortest interval between requests is three days, which we assume was adequate time for students to respond. We set $S_{ist}=1$ if subject $i$'s response timestamp is between those defining the beginning and end of period $t$, with $Y_{ist}$ given by their response to the survey if $S_{ist}=1$. We forward-impute retained response choices $\hat S_{ist}$ and retained responses $\hat Y_{ist}$ as described in Section (ref) for each value of $s$.
We begin by estimating the raw gender gap in entrepreneurial intentions without adjusting for the covariates listed in Table (ref). Because requests vary in number from two to seven between terms, we drop all observations with $R_{it}>2$ to support our independence assumption for the current specification. We estimate a parametric van1981demand selection model via maximum likelihood, separately for men and women. Specifically, we assume
and
where $X_{is} = 1$, and $Z_{ist} = [X_{is}, I(R_{ist}=2)]$. We assume $(\epsilon_{ist},u_{ist})$ are independent and identically distributed between $i$, with potential dependence allowed between $s$ and $t$ for each $i$. $\rho$ gives the correlation between outcome residuals $\epsilon_{ist}$ and response preference $u_{ist} = \Phi^{-1}(1-U_{ist})$, where $U_{ist}$ is $i$'s response aversion percentile in the notation of Section (ref) and $\Phi^{-1}(\cdot)$ denotes the inverse of the normal CDF. \footnote{ Our 15% response rate is insufficient for informative bounds from conservative approaches such as that of horowitz2000nonparametric, and we doubt that gender satisfies the monotonicity assumption of lee2009training or behaghel2015please, rendering the “local gender gap” among survey respondents unidentified, to say nothing of its policy-relevance. }
This gender-specific constant-only model implies that expected intention as a function of response aversion is $m(u) =\Phi((\beta+\rho\Phi^{-1}(1-u))/\sqrt{1-\rho^2})$, with the unconditional expectation expressed as $ \int_0^1m(u)du = \Phi(\beta)$. We graphically present our estimates of $m(u)$ for men and women in Figure (ref) for $u \in [0,1]$. Local average responses and sample averages for compliers to the first two requests are shown for reference. The levels of entrepreneurial intention for early and late respondents show a decreasing pattern for men and an increasing pattern for women, as in Table (ref). Differences in selection bias between men and women, reflected by their $\rho$ estimates, explain the difference between the 20 percentage point uncorrected gap estimate and the 12 percentage point corrected gap estimate.
In addition to estimating average entrepreneurial intention for men and women, we also decompose the gender intention gap using a nonlinear extension of methods developed by kitagawa1955components, oaxaca1973male, and blinder1973wage. Our decomposition estimates, for each covariate in $X_i$, the extent to which the gender gap would be reduced if women's average value of that single variable were equal to that of men, and how much it would be reduced if that variable's coefficient for women were equal to the coefficient for men. These results are shown in Table (ref).
To implement the decomposition, we estimate the model in ((ref))-((ref)) jointly for men and women via maximum likelihood using gender-specific parameters $(\beta_m,\alpha_m,\rho_m)$ for men and $(\beta_w,\alpha_w,\rho_w)$ for women. To avoid overweighting terms in which more requests were sent, we weight each observation by the inverse of the total number of requests sent to that student in that term. We define $X_{is}$ as all covariates listed in Table (ref) along with term fixed effects, and we define $Z_{ist}$ as binary indicators for each value in $R_{ist}=1,2...,7$ interacted with $X_{is}$. \footnote{ blandhol2022tsls recommend rich covariate interactions for instrumental variables when estimating local average treatment effects via two stage least squares. The estimator and target parameters are different here, but we find their argument convincing that if model assumptions hold conditional on $X_i$, then $X_i$ should be flexibly interacted with instruments to support identification. } We order the sample of $N_w$ women and $N_m$ men such that subjects with $i=1,...,N_w$ are women and those with $i = N_w+1,...,N$ are men, with $N= N_w+N_m$.
We estimate the contribution of $x_{is} \subset X_{is}$ to the entrepreneurial intention gap as
wherein we denote the sample average of $x_i$ for men and women as $\bar x_{m}$ and $\bar x_{w}$, respectively, and we denote male and female coefficients on $x_i$ as $\beta_{m,x}$ and $\beta_{w,x}$, respectively. \footnote{Fairlie's fairlie2005extension decomposition for nonlinear models considers the case in which the marginal distribution of $x_{is}$ for women is set equal to men's, while ours considers the case in which the mean of $x_{is}$ for women is set equal to men's via a level shift of women's distribution. We view the counterfactual we consider as likely simpler to understand and simpler to implement for policymakers, though describing $\beta_{w,x}$ as the “effect” of changing $x_{is}$ for women requires assumptions beyond those that either we or fairlie2005extension make. } We similarly estimate the contribution to the average gender gap of $x_{is}$'s differential effect on $Y^*_{is}$ between men and women as
In addition to estimating the contribution of each variable and its coefficient on the gender-intention gap, we also estimate the contribution of the entire vector $X_{is}$ on the gap as
defining the vector of sample means of all covariates in $X_{is}$ for men as $\bar X_{m}$ and for all women as $\bar X_{w}$. Similarly, we estimate the contribution of the entire vector of gender differences in coefficients on the gender-intention gap as
We estimate the remaining gap, which is driven by the interaction of male-female differences in $X$ and male-female differences in $\beta$, distributional differences in $X$ between men and women, and nonlinearity of the normal distribution as
Table (ref) contains estimated means of the variable or coefficient listed in each row for men, women, and the contribution of that item to the intention gap, using predicted values from a naive probit in columns 1–3, and using our proposed selection correction model in columns 4–6. $X_{is}$ is observed for all individuals regardless of survey response, so the means of $X_{is}$ by gender are the same for the uncorrected model and the corrected model. Columns 3 and 6 report estimated contributions of each row quantity to the gender gap as described in equations ((ref)), ((ref)), ((ref)), ((ref)), and ((ref)). We perform our decomposition for all variables in Table (ref), with term fixed effects included in $X_{is}$ contributing only to the unexplained gap.
We estimate the overall gender gap by calculating
where our estimates of $(\hat \beta_w, \hat \beta_m)$ differ by estimation method. The estimated gender-intention gap from our preferred specification is 18.3 percentage points, with a naive probit producing an estimate of 20.2 percentage points. Differences between the corrected and uncorrected model are due to the use of only observations with $t=T$ for the naive probit and the inclusion of the $\rho$ parameter in the corrected model (with $\rho=0$ fixed in the uncorrected model) which is positive if early respondents have higher intention and negative if they have lower intention. Our estimates of $\rho$ in the current specification are insignificant, suggesting limited nonresponse bias, though we see the same pattern with the signs of our point estimates here as we did in the no-controls specification shown in Figure (ref). Along these same lines, our richer specification repeats the finding in Figure (ref) that the uncorrected gap point estimate is larger than the corrected estimate, but here we fail to reject that they are the same in the population (p=0.449).
The difference in corrected estimates in Section (ref) and the corrected estimates in Table (ref) is due to the inclusion of controls and observations for requests three through seven, which were observed in some years but not others. The richer nonresponse correction specification in this section approximates a nonparametric specification that would fit a curve through local responses as in Figure (ref) for each value of $X_{is}$, where the slope of the curve is determined by a gender-specific $\rho$ that is held fixed across $X_i$. For both men and women, $\rho$ is attenuated in the richer model, suggesting that observed covariates account for both late responses and differences in intention between late and early responses. Nonetheless, the same general pattern from Figure (ref) persists, with positive selection for men ($\hat\rho_m>0)$ and negative selection for women $(\hat \rho_w<0)$, though these estimates are not statistically significant in this specification.
Our estimates are only reliable if the assumptions in Section (ref) and equation ((ref)) hold for this application. Under these assumptions, all local average responses should be explained by the marginal survey response function. We assess this by performing a “predictable trends” overidentification test wherein we include indicators for requests $R_{it}=3,...,7$ in $X_i$ with coefficient vector $\beta_R$, relying on requests one and two for identification. We fail to reject the null that later requests have no effect on responses at conventional levels for men (p=0.135), for women (p=0.411), and for both jointly (p=0.200). However, these point estimates are narrowly insignificant and mostly positive for men as shown in Figure (ref), suggesting that later requests drive corrected point estimates upward toward the uncorrected estimates for men when they are included, increasing the estimated gender gap.
Most of the entrepreneurial intention gender gap is unexplained by our observed variables. While the contribution of differences in observed covariates to the gender gap of 1.3 percentage points is statistically significant at conventional levels, it explains a small fraction of the overall gap. Meanwhile, the total corrected contribution of gender differences in $\beta$ to the gender gap is larger at 3 percentage points, but it is not significant at conventional levels. The largest statistically significant individual contributors to the gender intention gap are gender differences in the rates of choosing a STEM major and gender differences in the relative rates of intention between STEM and non-STEM students. Interestingly, these factors approximately cancel each other out. On one hand, there are more male STEM majors, who have somewhat lower entrepreneurial intention than male non-STEM majors. Meanwhile, there are fewer female STEM majors, but female STEM majors have much lower entrepreneurial intention than female non-STEM majors.
Results from our preferred specification suggest a large gender gap in undergraduate entrepreneurial intention, with minimal nonresponse bias in this survey. Despite the apparent lack of significant nonresponse bias, our corrected estimates support substantially different inferences than the uncorrected estimates because of their larger standard errors. Perhaps the clearest evidence of this is that the uncorrected gap estimate strongly rejects a true population average equal to the corrected point estimate. Given that our parametric correction gives substantially more precise estimates than conservative approaches such as that of horowitz1998censoring, we recommend against drawing inferences while assuming no nonresponse bias in applications such as ours.
We have established that data requests made to subjects are valid nonresponse correction instruments in panel data containing their accumulated requests and responses over time, under assumptions similar to those routinely made when analyzing survey data. Specifically, we assume that each data request must be sent to a representative sample of the population, requests should not directly affect responses, and average responses to any given request should be free of systematic measurement error. While these assumptions are strong, their importance for securing identification is not specific to our method. In addition to these assumptions, we make a monotonicity assumption similar to that of imbens1994identification which allows us to order subjects by their response timing and extrapolate from respondents to nonrespondents using existing nonresponse correction methods.
Our procedure for constructing request instruments imposes minimal challenges for data collection and analysis. It requires no knowledge of econometrics by data collection administrators and no financial incentives for subjects, which can bias responses if subjects maximize their effective wages by minimizing time spent on the survey, such as by using artificial intelligence to concoct responses westwood2025. Meanwhile, it naturally reveals rich variation in response rates when multiple reminders are sent, enabling model validation tests that compare late responses to their predicted values implied by early responses. These features enable routine use of our method when analyzing data obtained by contacting subjects, either via surveys or via similar procedures such as gathering lab data after medical experiments.
We illustrate our method by estimating the gender gap in entrepreneurial intentions among undergraduate students at the University of Wisconsin-Madison. To do so, we construct a panel dataset that contains a row for each student for each data request they receive, with their accumulated requests recorded in all rows and their responses recorded only in rows representing time periods after they respond. We find a large gender gap uncorrected estimates strongly rejecting the point estimate of the average gender gap obtained when correcting for nonresponse. In an additional application in Supplemental Appendix (ref), we accurately and precisely estimate ground truth averages from administrative data for five out of six labor market indicators using only aggregated statistics from early and late responses to the 45% response-rate Norway in Corona Times survey. Our empirical results suggest that correcting for nonresponse using nonrandom request instruments in subject contact panel data can improve inferences, while posing relatively small challenges for data collection and analysis.
\appendixpageoff