EconBase
← All papers

A Panel Quantile Approach to Attrition Bias in Big Data: Evidence from a Randomized Experiment

Matthew Harding, Carlos Lamarche

arXiv 9 Aug 2018 · Econometrics · publishedJournal of Econometrics (2018) · 13 citations (OpenAlex)

arXiv:1808.03364 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This paper introduces a quantile regression estimator for panel data models with individual heterogeneity and attrition. The method is motivated by the fact that attrition bias is often encountered in Big Data applications. For example, many users sign-up for the latest program but few remain active users several months later, making the evaluation of such interventions inherently very challenging. Building on earlier work by Hausman and Wise (1979), we provide a simple identification strategy that leads to a two-step estimation procedure. In the first step, the coefficients of interest in the selection equation are consistently estimated using parametric or nonparametric methods. In the second step, standard panel quantile methods are employed on a subset of weighted observations. The estimator is computationally easy to implement in Big Data applications with a large number of subjects. We investigate the conditions under which the parameter estimator is asymptotically Gaussian and we carry out a series of Monte Carlo simulations to investigate the finite sample properties of the estimator. Lastly, using a simulation exercise, we apply the method to the evaluation of a recent Time-of-Day electricity pricing experiment inspired by the work of Aigner and Hausman (1980).

Citation extraction

59
references
70
in-text mentions
59
distinct cited
1
self-citations
12,940
main-text words

appendix boundary found by appendix_command · 81% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Robins, Rotnitzky, and Zhao (1995) Analysis of Semiparametric Regression Models for Repeated Outcomes in the Presence of Missing Data0.58531100%
2Abrevaya and Dahl (2008) The Effects of Smoking and Prenatal Care on Birth Outcomes: Evidence from Quantile Regression Estimation on Panel Data0.51121100%
3Bhattacharya (2008) Inference in panel data models under attrition caused by unobservables0.51121100%
4Ridder (1992) An empirical evaluation of some models for non-random attrition in panel data0.51121100%
5Harding and Lamarche (2016) Empowering Consumers Through Data and Smart Technology: Experimental Evidence on the Consequences of Time-of-Use Electricity Pri…0.51121100%
6Canay (2011) A simple approach to quantile regression for panel data0.51121100%
7Hirano, Imbens, Ridder, and Rubin (2001) Combining Panel Data Sets with Attrition and Refreshment Samples0.51121100%
8Koenker (2004) Quantile Regression for Longitudinal Data0.51121100%
9Lipsitz, Fitzmaurice, Molenberghs, and Zhao (1997) Quantile Regression Methods for Longitudinal Data with Drop-outs: Application to CD4 Cell Counts of Patients Infected with the H…0.51121100%
10Chernozhukov, Fernández-Val, Hahn, and Newey (2013) Average and Quantile Effects in Nonseparable Panel Models0.51121100%

Showing the top 10 of 59 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
12004.051270.64432
2Post-Selection Inference in Three-Dimensional Panel Data0.40511
3Machine Learning Panel Data Regressions with Heavy-tailed Dependent Data: Theory and Application0.40511