EconBase
← Back to paper

Difference-in-Differences with a Misclassified Treatment

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

79,310 characters · 15 sections · 97 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Difference-in-Differences with a Misclassified Treatment

\begingroup \footnote{$^\dag$Department of Econometrics and Business Statistics, Monash University. Email: [email removed]. $^\ddag$Indira Gandhi Institute of Development Research, Mumbai. Email: [email removed]. \\ $^\ast$We would like to thank seminar participants at EWMES 2021 and ESAM 2022 for comments and suggestions on earlier versions of this paper.} \addtocounter{footnote}{-1} \endgroup

spacing{1.0} \begin{abstract} This paper studies identification and estimation of the average treatment effect on the treated (ATT) in difference-in-difference (DID) designs when the variable that classifies individuals into treatment and control groups (treatment status, $D$) is endogenously misclassified. We show that misclassification in $D$ hampers consistent estimation of ATT because 1) it restricts us from identifying the truly treated from those misclassified as being treated and 2) differential misclassification in counterfactual trends may result in parallel trends being violated with $D$ even when they hold with the true but unobserved $D^\ast$. We propose a solution to correct for endogenous one-sided misclassification in the context of a parametric DID regression which allows for considerable heterogeneity in treatment effects and establish its asymptotic properties in panel and repeated cross section settings. Furthermore, we illustrate the method by using it to estimate the insurance impact of a large-scale in-kind food transfer program in India which is known to suffer from large targeting errors. \end{abstract} JEL Classification Codes: C21, C23, C51 Keywords: Difference-in-differences, Panel data, Repeated cross-sections, Misclassification, Heterogeneous treatment effects

Introduction

The degree of measurement error in economic data and its influence on parameter estimates have been of much interest to econometricians. One case in particular is measurement error in a binary variable, also known as misclassification. Unlike classical measurement error, misclassification is necessarily non-classical since the error is negatively correlated with the truth aigner1973regression,bound2001measurement. This is an important case in the program evaluation literature, of which Difference-in-Differences is a workhorse empirical strategy, because the key regressor of interest, the treatment, is often binary in nature.\footnote{Evidence from the literature suggest that ignoring even a small amount of misclassification can have major repercussions for the estimated treatment effects millimet2011elephant,kreider2010regression.}

In this paper, we discuss identification and estimation of the average treatment effect on the treated (ATT) in a DID setting when the binary variable that classifies individuals into treatment or control groups (treatment status, $D$) is misclassified. The recent surge in the econometric literature on DID has seen a renewed interest in the identification, estimation, and interpretation of effects using the most common regression specifications. Despite such attention, the effects of a misclassified treatment on the DID estimand and its causal interpretation remain unknown.

Participation in programs or interventions is particularly prone to mismeasurment due to self-reporting or mistargeting bruckmeier2021misreporting,martinelli2009deception,coady2004targeting. For instance, meyer2015household document how misreporting of program receipt and conditional transfers from the government appear to be the biggest threat to household survey data quality for policy evaluation. Given that DID is often applied to data from nationally representative surveys groen2008effect,buchmueller2011effect, botosaru2018difference, it becomes all the more important to study the effects of misclassification on the DID estimand and its causal interpretation as the ATT.

We begin by showing that with $D$, the DID estimand can be decomposed into a sum of ATT for those misclassified as being treated (i.e. $D=1$) and the difference in counterfactual trends between those misclassified as being treated vis-$\grave{a}$-vis the misclassified controls. Therefore, the bias in DID is two-fold and confounds identification of the true ATT because: 1) it restricts us from identifying the truly treated from those misclassified as being treated (unless misclassification probabilities are known) and 2) the counterfactual trends in $D$ may not be parallel between those observed to be treated or non-treated even if parallel trends hold with the true unobserved $D^\ast$.

The first component of this bias is not new and has appeared in the misclassification literature including aigner1973regression, lewbel2007estimation, battistin2011misclassified among others. The second component, however, is a unique aspect of the DID design where parallel trends play an instrumental role in identifying the ATT. With $D$, parallel trends may not hold under the general case of differential misclassification in the counterfactual trends of the treated and control groups. But if we assume non-differential misclassification in the trends, then proposition (ref) in our paper presents a key relationship between counterfactual trends in $D$ vis-$\grave{a}$-vis $D^\ast$. It show that when parallel trends hold with $D^\ast$, then they will also hold with $D$ and vice versa except in the odd case where the probability of misclassification is the same as that of correct classification.

We outline our main approach in the context of a flexible linear specification of the potential outcome means and use it to arrive at parametric DID regressions for 2-group 2-period (2$\times$2) panel and repeated cross section settings. Our approach follows from nguimkeu2019estimation (NDT thereafter) who study endogenous participation and endogenous misreporting of social programs. We extend NDT to a 2$\times$2 setting and focus on the one-sided misclassification case which only explicitly accounts for errors of exclusion. Our regressions allow for heterogeneity by having interactions between covariates, treatment, and time. We then propose a two-step estimator which corrects for endogenous one-sided misclassification using a partial observability probit model (POP). Identification of the POP parameters and in turn the true ATT is achieved using a single exclusion restriction that only affects the misclassification probability but does not influence the true treatment or the outcome model. While NDT require two exclusion restrictions (one for endogenous treatment and the other for endogenous misclassification), we require only one since the DID design accounts for endogenous participation as long as it is based on time-invariant unobservables.

As an illustration of this method, we apply it to study the welfare impact of the Public Distribution System (PDS) which is a large-scale program for targeted distribution of highly subsidized food grains in India. One key argument in favor of in-kind transfers is that they provide implicit insurance to eligible households from commodity price risk gadenne2021kind,negi2022global. Using India as the setting, we quantify the welfare benefits of PDS for the poor who are most vulnerable to high food prices. The key challenge in estimating the welfare impact of PDS is that it is known to have large targeting errors where errors of wrong exclusion are disproportionately larger than errors of wrong inclusion dutta2001targeting,swaminathan2001errors, hirway2003identification,khera2008access,jha2013safety,pingali2019reimagining. We use the two-step estimator to correct for one-sided mistargeting of the treatment due to errors of wrong exclusion. Although our theory is in terms of misclassification, we show that it can be reformulated to fit the context of mistargeting, demonstrating the versatility of the method in different empirical settings.

Our paper makes three important contributions. The first contribution is to study and characterize the effects of misclassification on the DID estimand and its implications for identifying and estimating the ATT. To the best of our knowledge, this is the first paper to discuss the problem of misclassification within a DID framework and highlight the two important sources of bias in the DID estimand on account of using $D$ rather than $D^\ast$. We also argue how parallel trends in $D^\ast$ or even $D$ is no longer sufficient for identifying the true ATT and present how they relate to each other under the assumption of non-differential misclassification in counterfactual trends. Our second contribution is to propose a solution for the case of endogenous (or differential) one-sided misclassification in panel and repeated cross-section regressions and to characterize the inconsistency (asymptotic bias) in the first-differenced (FD) and pooled OLS (POLS) estimators of ATT for the panel and repeated cross section regressions, respectively. In doing so, we extend NDT's approach to the DID setting and establish consistency and asymptotic normality of the two-step DID estimators. A salient difference between NDT and us is that we require only one exclusion restriction which affects the misclassification probability but can still allow for an endogenous treatment as long as it is due to time-invariant factors. Our third contribution is empirical in nature where we apply the proposed method to estimate the insurance impact of PDS in the presence of one-sided mistargeting which is a reality in social welfare programs. We find that if the PDS subsidies were correctly targeted, the welfare impact of the program would have been much larger than what was actually observed.

Our paper contributes to two different strands of the causal inference literature. The most obvious contribution is to the misclassified binary regressor literature with antecedents such as aigner1973regression, mahajan2006identification, lewbel2007estimation, frazis2003estimating, meyer2017misclassification, haider2020correcting, calvietal2021, bollinger1996bounding, battistin2011misclassified, and ditraglia2019identifying. More recent contributions include ura2018heterogeneous, yanagi2019inference, tommasi2020bounding, and acerenza2021marginal. The other strand concerns the broader DID literature which includes abadie2005semiparametric, botosaru2018difference, goodman2018difference, de2020two, sant2020doubly, sun2020estimating, callaway2020difference, rambachan2022more, and wooldridgetwfe to name a few.

In particular, our paper is closely related to botosaru2018difference in the sense that they also depart from assuming perfect knowledge of the treatment status in DID studies and are rather interested in identifying the ATT when the treatment status in missing in one of the two periods. This is different than the main premise of our paper which assumes the treatment status to be observed with error or misclassified. Our paper is more general than botosaru2018difference since the problem of missing treatment status can be accommodated within our framework.\footnote{Suppose the missing treatment status is predicted in one of the two periods (as in botosaru2018difference), then the predicted status can be considered to be the observed but misclassified treatment status.}

The rest of this paper is organized as follows. Section (ref) presents a brief review of the binary misclassified regressor and the DID literatures. Section (ref) uses the standard DID framework to characterize the bias in the DID estimand when one uses the misclassified proxy $D$ in place of the true but unobserved treatment, $D^\ast$. This section also shows that under the simplifying assumption of non-differential misclassification, a relationship between parallel trends with $D$ and $D^\ast$ is easily established. Despite that, the true ATT remains unidentified. In this particular case, the problem reduces to the one explored in earlier misclassification papers. Section (ref) characterizes the bias in the coefficient on the treatment using the simple DID regression without covariates and explains why a standard instrumental variables strategy will not work. Section (ref) presents our main linear conditional specification derived using the potential outcome primitives to arrive at flexible 2$\times$2 panel and repeated cross section regressions. We maintain the assumption of conditional parallel trends and derive a more general bias expression for the case with covariates. Section (ref) presents the partial observability probit for one-sided misclassification and then establishes consistency and asymptotic normality of the two step FD and POLS DID estimators of the true ATT. Section (ref) applies the proposed method to correct for one-sided mistargeting in estimating the welfare impact of the Public Distribution System in India which suffers from targeting errors. Finally, section (ref) concludes.

Related Literature

One of the earliest treatments of the effects of misclassified independent variables in least squares regression framework is found in aigner1973regression. bollinger1996bounding studies the problem of identification with a binary mismeasured regressor and provides a method to bound the model parameters. While it is commonly understood that instrumental variables can be used in dealing with errors in variables, they produce biased parameter estimates when the mismeasured variable is binary loewenstein1997delayed,black2000bounding,kane1999estimating. However, black2003measurement show that, in the case of non-classical measurement errors, OLS and IV estimates can bound the true parameter. kane1999estimating and black2003measurement show that when two mismeasured reports on the variable of interest are available, point estimates can be recovered using the GMM framework. frazis2003estimating study the problem of misclassified and possibly endogenous regressors in linear regressions and propose a GMM based method to bound parameter estimates. Likewise, mahajan2006identification studies identification with misclassification but relaxes the assumption of independence between measurement error and other explanatory variables. mahajan2006identification proposes a solution akin to an IV strategy where marginal effects are identified using an additional variable correlated with the true value but uncorrelated with the measurement error. hu2008identification generalizes Mahajan's results to multivalued discrete variables. molinari2008partial derives partial identification results for discrete probability distributions under misclassification. kreider2010regression studies identification of regression coefficients in a linear probability framework when a binary regressor is fully arbitrarily misclassified.

While the literature above relates to misclassification in binary explanatory variables, its real relevance is reflected in the program evaluation literature where the focus is on the parameter estimate of the treatment indicator. lewbel2007estimation studies the identification and estimation of treatment effect with a misclassified treatment. He shows that under misclassification, the treatment effect has an attenuation bias and proposes an IV to overcome this bias. millimet2011elephant, using Monte Carlo methods, compares the performance of different treatment effect estimators with measurement errors in treatment assignment and finds most estimators performing poorly even with minor and infrequent errors.

In contrast to point identification approaches relying on repeated measurements or IV-type variables, kreider2012identifying propose partial identification methods to bound the average treatment effect when the participation is misreported and possibly endogenous using semiparametric and nonparametric approaches. nguimkeu2019estimation study point identification of treatment effects in parametric settings when participation is endogenously misreported. They propose a two-step estimation procedure that relies on poirier1980partial partial observability model for consistent estimation of the conditional average treatment effect.

In the context of the Local Average Treatment Effect (LATE) with missing or mismeasured treatment, studies have proposed methods for both point and partial identification. When two different measurements or an instrument for the misclassified treatment is available, studies have proposed methods to point identify the LATE battistin2014misreported,ditraglia2019identifying. yanagi2019inference shows that LATE can be point identified under misclassified treatment if exogenous covariates are available. calvietal2021 show that point identification of LATE is possible when the treatment status of some observation is missing if two mismeasured treatment indicators are available. Moreover, under misclassification, their proposed estimator has less bias than the standard LATE estimator. ura2018heterogeneous proposes a binary instrument for set identification in a nonparametric heterogeneous treatment effect framework with mismeasured and endogenous treatment assignment. Likewise, an IV strategy is also proposed by tommasi2020bounding to bound the LATE with misreported survey data. acerenza2021marginal bound the MTE under misclassification.

Although misclassified treatment in DID settings is unexplored, papers have looked at the impact of missing or misclassified treatment in longitudinal and repeated cross sectional settings. card1996effect studies the impact of labor unions on wages in longitudinal data while explicitly accounting for misclassification errors in reported union status. botosaru2018difference propose a method that identifies the Average Treatment Effect on the Treated (ATT) in a DID framework when individual group membership is missing in one of the time periods. Identification in their setup is achieved via the existence of proxy variables correlated with the treatment status. byker2016treatment evaluate the impact of a female sterilization campaign implemented in Peru between 1996 and 1997 using an inverse probability weighting (IPW) estimator that accounts for contamination in the self reported sterilization status. de2018fuzzy study identification in fuzzy DID when the share of treated units increases more in some groups than in others between the two time periods. As observed by botosaru2018difference, de2018fuzzy framework can also be applied to the case of missing treatment status in the baseline period.

DID, Parallel Trends, and Misclassification

To motivate the problem of misclassification, let $D^\ast$ denote the true but unobserved treatment status. However, suppose only we observe its proxy $D$ which falsely records receipt or non-receipt of the program thereby resulting in classification errors. We can write

equation[equation omitted — 51 chars of source]

where by definition $\varepsilon = D-D^\ast$. As illustrated in acerenza2021marginal, this is a general formulation and includes both; errors of inclusion and exclusion. \footnote{The triplet $(D, D^\ast, \varepsilon)$ can take any of the following values $\{(0,0,0), (1,0,1), (0,1, -1), (1,1,0)\}$. Then, $D = D^\ast \text{ if } \varepsilon \in \{0\} \text{ and } D= 1-D^\ast \text{ if } \varepsilon \in \{1, -1\} $} Let $S$ be a binary indicator representing the event $\varepsilon \in \{0\}$ (no misclassification). Then, we can also express ((ref)) as

equation[equation omitted — 70 chars of source]

Let $\{Y_t(0), Y_t(1)\}$ denote two potential outcomes (PO) where $Y_t(0)$ is the outcome in the control state and $Y_t(1)$ is the outcome in the treated state. Let $t=0$ represent the baseline period and $t=1$ represent the post-treatment period. Since $D^\ast=0$ for everyone in the baseline, we can write the observed outcome as

equation[equation omitted — 104 chars of source]

We assume that we either have a two-period panel or two repeated cross sections. Following wooldridgetwfe, we can decompose each PO into

equation[equation omitted — 55 chars of source]

which is a sum of the outcome at baseline and the gain over time, defined as $G_t(d) = Y_t(d)-Y_0(d)$ for each $d=0,1$. Therefore if $D^\ast$ were accurately observed, the DID estimand could be causally interpreted as the true $\text{ATT}$, defined as,

equation[equation omitted — 83 chars of source]

under the identifying assumption of unconditional parallel trends which says that \[\mathrm{DT}(D^\ast) \equiv \mathbb{E}[G_1(0)|D^\ast=1] -\mathbb{E}[G_1(0)|D^\ast=0] = 0.\] A crucial aspect of the above identification argument is that the treatment status, $D^\ast$, is measured accurately. In other words, we know individuals' group membership into treatment and control before conducting the DID analysis. However, with a misclassified $D$,

equation[equation omitted — 412 chars of source]

where $\tikz[baseline=(char.base)]{ \node[shape=circle,draw,inner sep=2pt] (char) {1};}$ is the ATT for those misclassified as being treated and $\tikz[baseline=(char.base)]{ \node[shape=circle,draw,inner sep=2pt] (char) {2};}$ is the difference in counterfactual trends when we use the misclassified $D$. Therefore, even if somehow $\mathrm{DT}(D) = 0 $ which would imply parallel trends in $D$, $\tikz[baseline=(char.base)]{ \node[shape=circle,draw,inner sep=2pt] (char) {1};}\neq \text{ATT}$. The following proposition presents a relationship between the quantities that are estimable using a misclassified $D$ vs. $D^\ast$ under the assumption of non-differential misclassification.

propositionIn a $2\times2$ design, if we assume non-differential misclassification in the mean i.e. $ \mathbb{E}[Y_1(1)-Y_1(0)|D, D^\ast] = \mathbb{E}[Y_1(1)-Y_1(0)|D^\ast]$ then \begin{equation*} \tikz[baseline=(char.base)]{ \node[shape=circle,draw,inner sep=2pt] (char) {1};}= \mathrm{ATT}\cdot \mathbb{P}(D^\ast=1|D=1)+\mathrm{ATU}\cdot \mathbb{P}(D^\ast=0|D=1) \end{equation*} where $\mathrm{ATU}$ refers to average treatment effect for those with $D^\ast=0$ (truly untreated). If in addition we assume non-differential misclassification in counterfactual trends, i.e. $\mathbb{E}[G_1(0)|D, D^\ast] = \mathbb{E}[G_1(0)|D^\ast]$ then, \begin{equation*} \tikz[baseline=(char.base)]{ \node[shape=circle,draw,inner sep=2pt] (char) {2};} = \mathrm{DT}(D^\ast)\cdot\left[\mathbb{P}(D^\ast=1|D=1)+\mathbb{P}(D^\ast=0|D=0)-1\right] \end{equation*}

Therefore, if one is willing to assume non-differential misclassification (which may be stronger than needed), one can demonstrate that the first component of $\text{DID}(D)$ in ((ref)) is actually a weighted average of the true ATT, true ATU, and the misclassification probabilities.

Under non-differential misclassification in counterfactual trends, the above relationship also implies that if parallel trends hold with the true but unobserved $D^\ast$ i.e. $\text{DT}(D^\ast)=0$, then they will also hold with the misclassified $D$ and vice versa unless $\mathbb{P}(D^\ast=1|D=1)=\mathbb{P}(D^\ast=1|D=0)$ which means that the probability of being misclassified is the same as that of being correctly classified (which is hard to justify). Note that even if $\text{DT}(D)=0$, with $D$ one can only identify the ATT for those misclassified as being treated and not the true ATT as defined in ((ref)).

Bias under Misclassification with Linear CEFs

Using ((ref)) and ((ref)), we can express the observed outcome in the post-treatment period as {

equation[equation omitted — 77 chars of source]

} The conditional mean of $Y_t$ given $D^\ast$ is {

align[align omitted — 194 chars of source]

} where the second equality in ((ref)) follows from two facts: 1) the decomposition given in ((ref)) and 2) the simple characterization that $\mathbb{E}[Y_1(1)-Y_1(0)|D^\ast] = D^\ast \cdot \mathbb{E}[Y_1(1)-Y_1(0)|D^\ast=1]+(1-D^\ast)\cdot \mathbb{E}[Y_1(1)-Y_1(0)|D^\ast =0]$ which implies $D^\ast\cdot \mathbb{E}[Y_1(1)-Y_1(0)|D^\ast] = D^\ast \cdot \mathbb{E}[Y_1(1)-Y_1(0)|D^\ast=1] = D^\ast \tau$. Since $D^\ast$ is binary, we can always express

equation[equation omitted — 87 chars of source]
assumption[Unconditional parallel trends in $D^\ast$] The trends in counterfactual outcomes in the absence of treatment between the truly treated and non-treated would not depend on $D^\ast$ i.e. $\mathrm{DT}(D^\ast) = 0$.

Parallel trends also imply that $\mathbb{E}[G_1(0)|D^\ast] = \mathbb{E}[G_1(0)] \equiv \theta$. Then combining equation ((ref)) and ((ref)) along with parallel trends, we obtain the simple DID equation\footnote{A similar equation can be derived for the repeated cross section setting by pooling observations for the two time periods and using the fact that the observed outcome is given by $Y = Y_0\cdot(1-T)+Y_1\cdot T$ where $T$ is a binary indicator for the post-treatment period. }

equation[equation omitted — 120 chars of source]

One can consistently estimate $\tau$ as the coefficient on $D^\ast$ from the first-differenced regression of $\Delta Y_i \text{ on } 1, D^\ast_i$ using a random sample of size $N$ from the population.

However, in the presence of misclassification, the coefficient on $D$ (say $\tau_{mis}$) is given by {

equation*[equation* omitted — 350 chars of source]

} In the context of the simple DID regression, we see that the inconsistency arises due to two reasons. The first is an attenuation bias as seen most notably in lewbel2007estimation, aigner1973regression, and battistin2011misclassified among others. The other part of the bias is due to counterfactual trends in $D$ not necessarily being parallel even though assumption ((ref)) is assumed to hold. This is because misclassification may lead to counterfactual trends diverging between the $D=0$ and $D=1$ groups even when parallel trends hold with the true $D^\ast$. Our solution allows for such endogenous or differential misclassification.\footnote{If, however, we assume non-differential misclassification between $S$ and counterfactual trends then a simple relationship between $\text{DT}(D)$ and $\text{DT}(D^\ast)$ emerges, as seen in section (ref).}

Would Instrumental Variables Work?

Can the instrumental variables approach help us in obtaining a consistent estimator of $\tau$? To see this, let's suppose we have a scalar instrument, $Z$, which is relevant for $D$. In addition to that, the IV also needs to be exogeneous to the regression error. Consider the first-dfferenced equation

equation*[equation* omitted — 67 chars of source]

where the IV must satisfy $ \mathrm{Cov}(Z, \Delta \epsilon) = 0$. However,

equation*[equation* omitted — 162 chars of source]

Being truly relevant means that $Z$ is also correlated with the misclassification error. Therefore the first covariance will be non-zero. This is because measurement error in a binary variable necessarily implies that the truth is negatively correlated with the error. Second, since misclassification may be correlated with $\Delta \xi$, $\mathrm{Cov}(Z, \Delta\xi)\neq 0$. Hence, a simple IV strategy will produce a biased estimator for $\tau$.

DID with Covariates

Often parallel trends are only plausible once we allow selection into treatment to be based on observables. To allow for this, we let $X = (X_1, X_2, \ldots, X_k)$ denote a vector of pre-treatment covariates.\footnote{For ease of derivations later on, we will use $R = (1, X)$ to be the set which includes an intercept.} Suppose that conditional parallel trends hold. Formally,

assumption[Conditional parallel trends] $\mathbb{E}[G_1(0)| X, D^\ast] = \mathbb{E}[G_1(0))|X]$

Assumption ((ref)) allows for covariate specific time trends in the potential outcome mean for the treated and control groups. DID studies that work with this assumption include abadie2005semiparametric, callaway2020difference, sant2020doubly, wooldridgetwfe etc.

assumption[Overlap] $\mathbb{P}(D^\ast=1) >0$ and $\mathbb{P}(D^\ast=1|X)<1$.

Assumption (ref) is the overlap condition required for the identification of the ATT.

We also assume that the researcher has access to either a two-period panel or two repeated cross sections. This is formalized in the assumption below.

assumption[Random sampling scheme] Assume that the data are either independent and identically distributed from i) Panel: $\{(Y_{it},X_i,D^\ast_i); \ i=1,2,\ldots,N\}$, for $t=0,1$ constitues an i.i.d draw from the population; ii) Repeated cross section: Conditional on $T=0$, the data are i.i.d from the distribution $\{(Y_{0}, D^\ast, X);\}$; conditional on $T = 1$, the data are i.i.d. from the distribution $\{(Y_{1}, D^\ast, X)\}$. \begin{equation*} \begin{split} \mathbb{P}(Y\leq y,D^\ast=d,X\leq x,T=t) &= t\cdot p\cdot P(Y_1\leq y,D^\ast=d,X\leq x|T=1) \\ &+(1-t)\cdot (1-p)\cdot P(Y_0 \leq y,D^\ast = d,X\leq x|T = 0), \end{split} \end{equation*} where $\lambda \equiv \mathbb{P}(T=1)\in (0,1)$ and $(y,d,x,t) \in \mathbb{R}\times\{0,1\}\times \mathbb{R}^k\times \{0,1\}$.

Assumption (ref) i) discusses the random sampling scheme in the case of a two-period panel where we have repeated observations on individuals for both the pre and post treatment periods. Part ii) discusses the sampling scheme in the case of two repeated cross sections where we observe each individual only once, either in the pre-treatment or the post-treatment period. This assumption rules out compositional changes in the sample (see hong2013measuring).

In the presence of covariates the conditional expectations, $\mathbb{E}[Y_t|X, D^\ast]$, are given as

equation[equation omitted — 180 chars of source]

where $\tau(X) = \mathbb{E}[Y_1(1)-Y_1(0)|X, D^\ast=1]$. Assume that

align[align omitted — 264 chars of source]

where $\mathring{X} = X-\mathbb{E}(X|D^\ast=1)$. Then substituting ((ref)), ((ref)), ((ref)) in ((ref)), we can generally express

equation[equation omitted — 234 chars of source]

where $\mathring{R} \equiv (1, \mathring{X})$, $W^\ast \equiv D^\ast \mathring{R}$, $\delta = (\delta_{01}, \delta_{11}^\prime)^\prime$, $\eta_1 = (\eta_{00}, \eta_{02}^\prime)^\prime$, $\eta_2 = (\eta_{01}, \eta_{03}^\prime)^\prime$, and $\theta = (\tau, \kappa)^\prime$. Note that our framework allows the covariates to affect the dynamics of the potential outcomes and also considers the treatment effect to be heterogeneous based on covariates. With a two-period panel, one can estimate the first-differenced equation

comment\begin{equation} \Delta Y = \delta_{01}+\mathring{X}\delta_{11}+D^\ast\tau+D^\ast\mathring{X}\kappa+\Delta \xi \end{equation}
equation[equation omitted — 197 chars of source]

\paragraph{Repeated cross sections} Often DID analysis is also carried out using two repeated cross sections since obtaining data on same observations before and after the treatment is not feasible. In that case, one can estimate the pooled regression

equation[equation omitted — 256 chars of source]

Note that the true treatment indicator, $D^\ast$, and covariates, $X$, are exogenous in equations ((ref)) and ((ref)) above. Therefore, if we perfectly observe $D^\ast$, we can consistently estimate $\theta$, and consequently $\tau$, as the coefficient on $D^\ast$ from the equations above depending on the kind of sample we have.

Partial Observability Probit

In the current paper, we propose a solution for only one-sided misclassification by assuming that

equation[equation omitted — 35 chars of source]

which means that $D=1$ only if $D^\ast=1$ and $S=1$. Apart from this, the one-sided formulation is only able to capture errors of exclusion where $D^\ast=1$ but $D=0$ which corresponds to $S=0$. Errors of inclusion (i.e. $D=1$ when $D^\ast=0$) are not explicitly accounted for in equation ((ref)). Suppose,

equation*[equation* omitted — 103 chars of source]

Then,

equation[equation omitted — 80 chars of source]

Let the conditional distribution of $(-U, -V)$ be bivariate normal with the CDF denoted by $F_{U, V}(\cdot, \cdot; \rho)$ where $\rho$ is the correlation coefficient (see assumption (ref)). Then,

equation*[equation* omitted — 116 chars of source]

Identification of parameters $\alpha$ and $\gamma$ in the partial observability probit model requires one exogenous variable in $Z$ that is excluded from $R$ in order to satisfy local identification assumptions in poirier1980partial.

The DID design allows for the treatment, $D^\ast$, to be endogenous only due to time invariant unobservables. Therefore, comparisons between treated and non-treated groups in the baseline and post-treatment periods provide a valid estimate of the ATT. The design, therefore, dictates a clear relationship between the outcome error and $U$. With respect to misclassification, since $S$ may be correlated with time varying unobservables affecting the outcome dynamics, we allow $V$ and the outcome error to be correlated. We already mention that this will lead to parallel trends not holding with $D$ even if we assume them to hold with $D^\ast$.\footnote{except when misclassification depends only on observables in which case we will obtain a similar result as section (ref) conditional on $X$.} We modify the assumptions in nguimkeu2019estimation to reflect these subtleties.

assumptionAssume that \begin{itemize} • The error term $\Delta\xi$ is independent of $R$, $Z$, with variance $\sigma^2$; and the error terms $(U,V)$ are independent of all covariates $R$, $Z$ and have unit variances. The correlations for the pairs $(\Delta \xi,V)$ and $(U, V)$ are denoted $\psi_v$ and $\rho$, respectively. • The error terms $(\Delta\xi, U, V)$, follow a trivariate normal distribution, conditional on all covariates $(R,Z)$. \begin{equation*} (\Delta\xi, U, V)^\prime|R, Z \sim N(0,\Sigma) where \Sigma = \begin{pmatrix} \sigma^2 & 0 & \psi_v\sigma \\ 0 & 1 & \rho \\ \psi_v\sigma & \rho & 1 \end{pmatrix} \end{equation*} \end{itemize} In the case of repeated cross sections, simply replace $\Delta \xi$ with the pooled error $\xi$.

Given that we observe $D$, the first differenced equation becomes,

equation[equation omitted — 247 chars of source]

Note that $\ddot{R}_i \equiv (1, \ddot{X}_i)$ where $\ddot{X}_i = X_i-\bar{X}_1$ with $\bar{X}_1$ being the sample mean of covariates for those misclassified as being treated $(D_i=1)$. Similarly, $\ddot{W}_i \equiv D_i\ddot{R}_i$. Define $\hat{\theta}_{FD}$ to be the estimator of the coefficient on $\ddot{W}_i$ from estimating the FD equation given in ((ref)).

\paragraph{Repeated cross-sections} In the case of repeated cross sections, we obtain the following pooled regression equation

equation[equation omitted — 159 chars of source]

where $\epsilon_i = \left[ (\mathring{R}_i-\ddot{R}_i)\eta_1+(W_i^\ast-\ddot{W}_i)\eta_2+T_i\cdot (\mathring{R}_i-\ddot{R}_i)\delta+T_i\cdot (W_i^\ast-\ddot{W}_i)\theta+\xi_i\right]$. Define $\hat{\theta}_{POLS}$ to be the estimator for $\theta$ from estimating equation ((ref)) using a pooled sample.

Since $D$ is endogenous to both equations, $\hat{\theta}_{FD}$ and $\hat{\theta}_{POLS}$ will be inconsistent for $\theta$. We characterize this asymptotic bias from using a misclassified $D$ in place of $D^\ast$.

theorem[Bias under misclassification with covariates] Under assumptions (ref), (ref), (ref), and (ref) \begin{equation*} \mathrm{plim}(\hat{\tau}_{FD})-\tau = \mathbb{E}[\dot{R}_i Q^{-1}(A\delta+B\theta+C)|D_i=1] \end{equation*} where { $Q = \mathbb{E}(\dot{W}_i^\prime \dot{W}_i)-\mathbb{E}(\dot{W}_i^\prime\dot{R}_i)\mathbb{E}(\dot{R}_i^\prime\dot{R}_i)^{-1}\mathbb{E}(\dot{R}_i^\prime\dot{W}_i)$, $A = \mathbb{E}(\dot{W}_i^\prime\mathring{R}_i)- \mathbb{E}(\dot{W}_i^\prime\dot{R}_i)[\mathbb{E}(\dot{R}_i^\prime\dot{R}_i)]^{-1}\mathbb{E}(\dot{R}_i^\prime\mathring{R}_i) $, $ B = \mathbb{E}[\dot{W}_i^\prime(W_i^\ast-\dot{W}_i)] - \mathbb{E}(\dot{W}_i^\prime\dot{R}_i)[\mathbb{E}(\dot{R}_i^\prime \dot{R}_i)]^{-1}\mathbb{E}[\dot{R}_i^\prime(W_i^\ast-\dot{W}_i)$ and $C = \sigma\psi_v\mathbb{E}\left[\dot{R}_i^\prime \phi(-Z_i\alpha)\Phi\left(\frac{R_i\gamma-\rho Z_i\alpha}{\sqrt{1-\rho^2}}\right)\right] $} and \begin{equation*} \mathrm{plim}(\hat{\tau}_{POLS})-\tau = \mathbb{E}[\dot{R}_i \left\{Q_1^{-1}(A_1\pi_1+B\pi_2+C_1)-Q_0^{-1}(A_1\eta_1+B\eta_2+C_0)\right\}|D_i=1] \end{equation*} where $Q, A, B$ and $C$ are defined for the $T=1$ and $T=0$ populations.

As we can see from the above theorem, the asymptotic bias is not a simple attenuation bias and allows for endogeneous misclassification in $D$.

Two-Step Solution

The first step of the two-step solution we propose estimates the partial observability probit parameters $(\gamma, \alpha, \rho)$ by maximizing the log-likelihood function given by

equation*[equation* omitted — 190 chars of source]

We then use the first step maximum likelihood estimate of $\alpha$ to obtain a predicted $D_i^\ast$ as $\hat{D}_i^\ast = \Phi(R_i\hat{\gamma})$ and plug it in true regression equations. The first-differenced equation becomes

equation[equation omitted — 261 chars of source]

and $\hat{R}^\ast \equiv (1, \hat{X}^\ast)$ where $\hat{X}^\ast = X-\hat{\bar{X}}_1^\ast$ and $ \hat{\bar{X}}_1^\ast = \frac{1}{\hat{N}^\ast}\sum_{i=1}^{N}\hat{D}^\ast_i\cdot X_i$ with $\hat{N}^\ast = \sum_{i=1}^{N}\hat{D}^\ast_i$ being the sum of predicted probabilities. Also, $\hat{W}^\ast \equiv \hat{D}^\ast \hat{R}^\ast$. Then, using Frisch-Waugh again, the two-step estimator will be obtained from the following regression

equation[equation omitted — 115 chars of source]

where $\hat{M}^\ast = I-\hat{R}^\ast(\hat{R}^{\ast\prime}\hat{R}^\ast)^{-1}\hat{R}^{\ast\prime}$ and all variables involved are expressed in vector and matrix notations. Then,

equation[equation omitted — 165 chars of source]

where $\hat{\theta}_{FD}^{2S}$ is the proposed two-step FD estimator.

For the case of a repeated cross section, the DID estimator is obtained from equation ((ref)) by plugging in $\hat{D}^\ast$ in place of $D$ to obtain the following equation

equation[equation omitted — 184 chars of source]

where $\varepsilon_i = \left[ (\mathring{R}_i-\hat{R}^\ast_i)\eta_1+(W_i^\ast-\hat{W}^\ast_i)\eta_2+T\cdot (\mathring{R}_i-\hat{R}^\ast_i)\delta+T\cdot (W_i^\ast-\hat{W}^\ast_i)\theta+\xi_i\right]$.

Next, we establish consistency and asymptotic normality (along with the variance-covariance expressions) of the two-step FD and POLS DID estimators.

theorem[Asymptotic distribution of the two-step estimator] Under assumptions (ref), (ref), (ref), and (ref) given above, \begin{equation*} \begin{split} \sqrt{N}(\hat{\tau}_{FD}^{2S}-\tau) &\overset{d}{\rightarrow} N(0, \Omega_{FD}) \\ \Omega_{FD} = \mathbb{E}(\mathring{R}_i|D_i^\ast=1)&\cdot Avar[\sqrt{N}(\hat{\theta}_{FD}^{2S}-\theta)]\cdot \mathbb{E}(\mathring{R}_i|D_i^\ast=1)^\prime \end{split} \end{equation*} where \begin{equation*} Avar[\sqrt{N}(\hat{\theta}_{FD}^{2S}-\theta)]=\Omega_{\Gamma}^{-1}\left(\Omega_{1\theta}+\Omega_{2\theta}+\Omega_{3\theta}\right)\Omega_{\Gamma}^{-1} \end{equation*} and \begin{equation*} \begin{split} \sqrt{N}(\hat{\tau}_{POLS}^{2S}-\tau) &\overset{d}{\rightarrow} N(0, \Omega_{POLS}) \\ \Omega_{POLS}& = \mathbb{E}(\mathring{R}_i|D_i^\ast=1)\cdot Avar[\sqrt{N}(\hat{\theta}_{POLS}^{2S}-\theta)]\cdot \mathbb{E}(\mathring{R}_i|D_i^\ast=1)^\prime \\ Avar[\sqrt{N}(\hat{\theta}_{POLS}^{2S}-\theta)] &=\Omega_{\Gamma}^{-1}\bigg\{\left(\frac{\Omega_{1\pi}}{\lambda}+\frac{\Omega_{1\eta}}{(1-\lambda)}\right)+\left(\frac{\Omega_{2\pi}}{\lambda}+\frac{\Omega_{2\eta}}{(1-\lambda)}\right)\\ &+\left(\frac{\Omega_{3\pi}}{\lambda}+\frac{\Omega_{3\eta}}{(1-\lambda)}\right)\bigg\}\Omega_{\Gamma}^{-1} \end{split} \end{equation*}

where $\pi$ are the regression parameters for the $T=1$ population and $\eta$ index the $T=0$ population regression.

Empirical Application

We apply the proposed method to study the insurance impact of a large-scale in-kind transfer program in India. The program in question is the Public Distribution System (PDS) of India. The PDS is the world's largest targeted food safety net program that distributes highly subsidized food grains, mainly rice and wheat, to close to a billion people in 180 million poor eligible households balani2013functioning,gadenne2021kind. Initially, the program had universal coverage, but in 1997 it was transformed into the Targeted Public Distribution System (TPDS), which emphasized targeted food subsidies for only the poor eligible households.

The TPDS differentiates between households below (BPL households) and above (APL households) the official poverty line. The state governments have the responsibility of identifying the BPL households. The state governments rely on the elected village governments who in turn identify the poor using multiple criteria, including qualitative parameters such as possession of land operated/owned; ownership of TVs, motorcycles, and other durables; and ownership of agricultural machinery and implements kochar2005can,ram2009understanding,kaushal2015consumer. The BPL and APL households are issued different ration cards that identify them as entitled to either the BPL or the APL subsidy. The BPL households are the priority households and are allocated food grains at much lower prices than the APL households. Although any household above the poverty line is entitled to the APL ration card, the main beneficiaries of the TPDS are the households with the BPL ration card.

We are interested in estimating the insurance impact of access to cheap grains through PDS on eligible or treated households in India. negi2022global shows that during the global food price surge in 2007-2008, Indian households with access to PDS subsidy were able to maintain their staple food consumption and total calorie intakes by substituting expensive market purchased food grains with subsidized PDS grains. gadenne2021kind in the context of India, show that in-kind transfers provide implicit insurance to eligible households from commodity price risk. Their key argument is that in-kind transfers have insurance benefits as the value of the transfer is naturally indexed to the market value of the commodity and hence rises as the market price of the commodity rises. This implies that if food prices go up, access to in-kind transfers will allow households to maintain their real value of income. This is the key argument we empirically test in this paper. The challenge is that the PDS is known to have large targeting errors dutta2001targeting, swaminathan2001errors, hirway2003identification, khera2008access. Moreover, there is some evidence that the errors of wrong exclusion are much larger than errors of wrong inclusion jha2013safety, pingali2019reimagining. Estimates suggest that targeting errors continue to be significant as only 28% in the bottom 40% of the households access PDS pingali2019reimagining. jha2013safety report that the proportion of poor who used the PDS in 2004-2005 was only 30% and the exclusion error was as high as 70%. They go on to report that a part of this exclusion error was due to targeting errors where some of the poor were not identified as poor and hence were ineligible for PDS subsidies.

In principle, targeting should increase the efficiency of the program by transferring scarce government resources directly to groups that need them the most. However, in reality, the performance of such programs varies based on the targeting methods and the quality of information on the population coady2004targeting. A detrimental factor influencing the efficacy of such programs is the precision with which the targeted group is identified in the population coady2004targeting. Errors in targeting are common and arise due to imperfect information about eligibility, corruption, political connections, and elite capture cornia1993two, coady2004targeting,pande2007understanding, besley2012just, alatas2012targeting, panda2015political. For these reasons, these programs generally suffer from errors of inclusion or identifying non-eligible as eligible and errors of exclusion or identifying eligible as non-eligible.

In general, a mistargeted program poses a challenge in terms of its impact evaluation. A routinely used method to assess the performance of targeted programs is difference-in-differences (DID), but in the presence of targeting errors, it's not clear what the estimated treatment effect would capture. In such a scenario, the treatment effect from a DID regression will not reflect the program's true impact as the observed treatment group would be different from the intended treatment group. Although mistargeting is a reality in targeted social welfare programs, few studies have taken account of it while estimating the welfare impacts of such programs emran2014assessing,tohari2019targeting. cameron2014can report an extreme case from Indonesia where mistargeting of a cash transfer program actually led to destruction of trust and social capital and increase in criminal behavior.

Data and Setting

We use data from the Indian Human Development Surveys (IHDS) conducted in 2004-05 and 2011-12 to assess the insurance impact of PDS subsidies on Indian households desai2010national,desai2018national. The IHDS are large scale, nationally represented panel household surveys conducted by the National Council of Applied Economic Research (NCAER) India, the University of Maryland, Indiana University, and the University of Michigan. The IHDS collect household and individual level data on a wide variety of indicators including household income, expenditure, assets, education, caste, gender relations, local infrastructure, and availability of facilities. More importantly, the IHDS include self reported ownership of ration cards and the type (BPL or APL) of ration cards owned.

Table (ref) presents the summary statistics on key variables of interest from the baseline and endline IHDS. We categorize households with no ration card or APL card as the control group as they are not the main beneficiaries of the PDS subsidies. The treatment group is defined as households owning a BPL or an AAY (Antyodaya Ann Yojna) ration card, as they are the intended beneficiaries of the program.\footnote{AAY ration card is issued to the poorest households within the BPL category with more generous subsidies.} Since the treatment group has to be the same in both the survey rounds, we retain only those households whose ration card status remained constant in both the rounds. We also remove households in the state of Tamil Nadu as it didn't go for TPDS and followed universal PDS during the period of analysis. Overall, we are left with 3035 households in the treated group and 6345 in the control group.

In terms of income and assets, the treated households are poorer than households in the control group. This is expected as BPL ration cards are primarily given to poor households. One important control variable in Table (ref) is the household's participation in wage work under Mahatma Gandhi National Rural Employment Guarantee Act (MGNREGA). MGNREGA is India's large-scale anti-poverty rural workfare program which was introduced in 2005 to provide around 100 days per year of minimum wage employment to working age individuals. The MGNREGA is primarily operational in rural areas and offers voluntary unskilled employment on local public work projects. Another critical control is monetary benefits received from other government programs, calculated as the sum of transfers received from scholarships, old age pension, maternity schemes, disability schemes, income generation programs other than MGNREGA, drought/flood compensation assistance, and insurance payouts. This variable is coded as zero if no transfer was recieved; otherwise, the monetary amount of transfer. Note that participation and benefits from both of these programs show an increasing trend during this period and hence are important controls in our empirical specification.

Interestingly, treated households report having higher membership in the local caste associations and links to local politicians than the control group. This is consistent with the observations made in the literature that elite capture and connections with local politicians is helpful in selecting beneficiaries for transfer programs in developing countries pande2007understanding,besley2012just,panda2015political. In the context of the PDS program in India, panda2015political shows that local political connections are conducive to being selected into getting a BPL ration card. These observations are important for us as we will use these variables to correct for the influence of mistargeting in our estimates.

To set the context, we present global and domestic price trends in Figure (ref). As can be seen from the figure, global and domestic rice and wheat prices were on the rise around 2007-08. Evidence from the literature suggests that global food price increase around this period was triggered by productions shocks in major producing countries and ensuing countercyclical trade policies negi2022global. Given pressure form rising global food prices, domestic food prices in India also started trending upwards so much so that the price of rice and wheat almost doubled from their levels in 2004-05. Figure (ref) shows the average market and PDS price of rice and wheat estimated from the IHDS for the baseline and the endline period. The market price of rice and wheat show an increase in the IHDS as well but the PDS price registers a marginal decline. The decline is probably on account of increased subsidies on PDS rice and wheat during this period gadenne2021kind.

Table (ref) shows suggestive evidence on the role of PDS in insulating eligible households' food consumption from high food prices. The total consumption of rice and wheat remains comparable for both the treated and control groups across the two rounds, but the shift from market purchased to subsidized rice and wheat is clearly visible for the treated group. In comparison, for the control group, the decline in market purchased rice and wheat is minor.

Our outcome of interest is the daily per capita calorie intakes which we calculate from item-wise food consumption data reported in the IHDS. We select total calories intakes as the outcome as it reflects food security and is a better measure in this context than other monetary measures of welfare. Figure (ref) shows the trends in the average per capita calorie intakes. Two observations are worth noting from Figure (ref). First, average calorie intakes of BPL ration card owning households are lower than the that of the APL ration card owning households. This basically reflects the fact that BPL ration card households are generally poorer households. The second and more interesting observation is that the calorie intakes of the BPL ration card owning households is stable across the baseline and the endline whereas for APL ration card owning households, it shows a marked decline. Our hypothesis is that, during the period of high food prices, BPL households could maintain their calorie intakes essentially because they could access cheap grains from the PDS. APL households had either limited or no access to government subsidized grains hence could not maintain their calorie intakes. In the next section, we illustrate how we use observations from Figures (ref) and (ref) to propose a DID strategy to estimate the insurance benefits of the PDS subsidy which is mainly targeted to the BPL ration card owning households.

Differences-in-Differences with Targeting Errors

We propose our key hypothesis and the identification strategy using Figure (ref) which is illustrative of the trends in the outcome of APL and BPL ration card households and the trends in PDS and market price of rice and wheat. The bold grey lines trace the daily per capita caloric intakes of the two sets of households. The dashed lines plot the movements in the market and PDS price of rice and wheat. The reference price of food for BPL households is the PDS price but for APL households it’s the market price. In reality, both sets of households are buying rice and wheat from the market but APL households primarily depend on market purchased rice and wheat (see Table (ref)). The price of food for the APL households is higher in 2010 in comparison to 2005 but for the BPL households it’s the same across all periods.

We are interested in showing that the APL cardholder households suffered a welfare loss due to the rise in food prices and that the PDS subsidy insures its main beneficiaries from food price increase. In terms of Figure (ref), this effect can be captured by the following double difference.

equation[equation omitted — 44 chars of source]

where $(E-B)$ is the difference in the outcomes of APL and BPL households in 2005 and $(D-A)$ is the difference in their outcomes in 2010. Since the main beneficiaries of the PDS subsidy are the BPL households, their nutritional outcomes are unaffected by the price increase but we expect the outcome of the APL household to worsen in 2010. Therefore, we expect $\tau>0$. Note that, the dashed grey line segments $EH$ and $BG$ reflect the counterfactual scenarios for the APL and BPL ration card households respectively. In terms of the figure, the insurance benefit of PDS subsidy, $\tau$, is represented by $(H-D)$ which by construction in equal to $(A-G)$. The parallel trends assumption can be tested by estimating the following double difference.

equation[equation omitted — 39 chars of source]

For parallel trends to be satisfied, $\tau^p = 0$.

Consider the following specification which builds on the ideas presented in Figure (ref).

equation[equation omitted — 161 chars of source]

where the dependent variable is the log of daily per capita total calories intakes. $\text{BPL}_i$ is the treatment dummy which is $1$ is the household has a BPL ration card and $0$ otherwise. $\text{POST}_t$ is $0$ for the baseline survey and $1$ for the endline survey. We are interested in the coefficient $\tau$, which quantifies the insurance role of in-kind transfers. $X$ is a vector of control variables.

Since we know that there are targeting errors in the in-kind food transfers due to misallocation of BPL ration cards, the observed BPL ration card owning status is not equal to the intended or targeted BPL ration card status. The targeting errors in ration cards can be expressed by the following relationship.

equation[equation omitted — 58 chars of source]

where $\text{BPL}_i$ is the observed treatment dummy and $\text{BPL}_{i}^*$ is the unobserved treatment status. $S_i$ is a dummy variable which captures one-sided targeting errors. Observed BPL ration card status is defined as a product of the correctly targeted BPL ration card allocation and the targeting error captured by $S_i$.

equation[equation omitted — 138 chars of source]

$R$ is a vector of variables that were supposed to be used for the classification of PDS beneficiaries like incomes, consumption expenditures, land ownership, and ownership of different assets. In that sense, variables in $R$ should predict the intended beneficiaries of the PDS subsidy and hence should predict the true BPL ration card status. Vector $Z$ contains additional variables that we believe will influence mistargeting. Due to cumbersome bureaucracy and high corruption at the local level, poor households with links to local politicians or membership in local caste associations have a greater likelihood of getting BPL ration cards panda2015political. Therefore, connections with local politicians or membership in local caste associations will independently influence the misallocation of the BPL ration cards. To implement the two step estimator proposed in this paper we first express the outcome equation in first difference form as:

equation[equation omitted — 108 chars of source]

The estimate of $\tau$ is biased as the observed treatment dummy is misclassified. To correct for this, we first estimate the POP model expressed in equation (ref). We use the predicted BPL status from the POP in the outcome equation to get the true estimate of PDS impact. Since we assume heterogenous treatment effects, to get the true treatment effect estimates, we demean all covariates by treatment group means.

Results

Table (ref) presents the estimates from equation (ref) where the dependent variables are the per capita consumption of rice and wheat from different sources. We observe that high food prices induced treated households to substitute expensive market foodgrains mainly with subsidized PDS grains. Table (ref) shows that targeted PDS subsidies were effective in insulating eligible households' cereal consumption and food expenditures from high food prices.

Table (ref) presents the main DID results where the outcome variable is the log of per capita per day calories intake. Specification (1) presents the DID estimates without any control variables. In specifications (2), (3), and (4), we introduce household level control variables and household participation in other government programs which may be correlated with food price changes.\footnote{See Table (ref) for the control variables included in the DID regressions.} Considering the first three specifications, we find the estimate of the insurance impact of in-kind transfers to be positive and statistically significant across all three specifications. This implies that PDS subsidies successfully insulated the BPL ration card owning households from high food prices. These estimates are robust to addition of other household level covariates and participation and benefits received from other government welfare programs.

Although positive, the estimated impact of PDS subsidies is biased due to mistargeting. Following the two-step estimator proposed in this paper, we first estimate a partial observability probit (POP) model where we use links with local politicians and memberships in caste associations as variables independently predicting the classification errors. The estimates from the POP are presented in Appendix Table (ref). In the second step, we estimate the impact of the program using the predicted BPL rations card status. The estimates from the two-step estimator is presented in specification (4) of Table (ref). The estimate in specification (4) is much larger than what we get in a simple DID with a mistargeted treatment. This indicates that mistargeting lead to a downward bias in our estimates. This result is intuitive as some of the poor households who should have had access to the PDS subsidies didn't get the BPL ration card and hence were not insulated from the high food prices. A more policy relevant way of interpreting this result is that if PDS subsidies were correctly targeted, the insurance impact of the PDS program would have been much larger than what was actually observed.

Finally, to show that these effects were not present when food prices were not rising, we use the fact that some households were surveyed in 2004 and some in 2005 in the baseline survey. A conventional parallel trend check is not feasible in our case as a panel of households for a period before 2005 is not available. However, we can use staggered household interviews during the baseline survey to estimate a DID regression. Strictly speaking, this is not a parallel trend check as we are using a cross sectional sample and comparing different sets of households surveyed over different years, but still acts as a useful placebo test. Table (ref) in the Appendix reports the results from the placebo check. The DID estimate is close to zero and is statistically insignificant in all specifications indicating no insurance benefits of the program during a period of stable food prices.

Conclusion

Much of the literature on difference-in-differences, both econometric and applied, assumes that the treatment receipt is measured accurately. There is sufficient evidence in the applied literature to suggest that participation in programs is prone to misclassification whether that is due to issues of misreporting by individuals or mistargeting due to errors in identifying the program beneficiaries. In this paper, we focus on identifying and estimating the ATT within a standard DID framework when the observed treatment status, $D$, wrongly classifies individuals into treatment and control groups. We show that the bias in the DID estimand is two-fold and hampers consistent estimation of the true ATT because 1) it restricts us from identifying those with $D^\ast=1$ from among those with $D=1$ and 2) differential misclassification in counterfactual trends may result in parallel trends being violated with $D$ even when we assume them to hold with the true and unobserved treatment status, $D^\ast$.

Our main approach considers the case of one-sided misclassification and extends NDT to two-period panel and repeated cross section settings using flexible parametric regression specifications which allow for interactions between covariates, treatment, and time. We then characterize the asymptotic bias in the first-differenced and pooled OLS estimators of the ATT. We finally propose a two-step procedure which corrects for such one-sided misclassification and delivers a consistent and asymptotically normal estimator of the ATT in DID studies. We apply the proposed method to estimate the welfare impact of the Public Distribution System (PDS) of India in insulating its main beneficiaries (households below the poverty line) from commodity price risk. PDS is know to suffer from targeting errors which have largely remained unaccounted for while estimating the program effects. We find that if the PDS subsidies were correctly targeted, the insurance impact of the program would have been much larger than what was actually observed.

\singlespacing

Figures

figure[figure omitted — 615 chars of source]
figure[figure omitted — 380 chars of source]
figure[figure omitted — 450 chars of source]
figure[figure omitted — 336 chars of source]

Tables

table[table omitted — 3,952 chars of source]
landscape\begin{table}[H] \begin{threeparttable} \caption{Monthly Rice and Wheat Consumption for Treatment and Conrol Households (in kilograms per person)} \begin{tabular}{lrrrrrrrr} \toprule\toprule Ration card type & Total & Homegrown & Market & PDS & Total & Homegrown & Market & PDS \\ \midrule & \multicolumn{ 4}{c}{Baseline (2005)} & \multicolumn{ 4}{c}{Endline (2012)} \\ Control: No ration card & 12.8 & 0.6 & 12.0 & 0.0 & 12.6 & 0.8 & 11.6 & 0.0 \\ Control: APL (above poverty line) ration card & 11.0 & 0.6 & 9.8 & 0.7 & 11.6 & 0.8 & 9.5 & 1.3 \\ Treatment: BPL (below poverty line) ration card & 11.2 & 0.7 & 7.5 & 3.0 & 12.0 & 1.0 & 5.6 & 5.4 \\ Treatment: AAY (antyodaya ann yojna) ration card & 14.2 & 0.8 & 7.6 & 5.7 & 14.6 & 1.0 & 5.4 & 8.1 \\ \bottomrule\bottomrule \end{tabular} \begin{tablenotes}[flushleft] • Notes: No ration card and APL ration card households are in the control group. AAY is given to poorest families from within the below poverty line households. BPL and AAY ration card household are the main benificiaries of the PDS subsidies hence form the treatment group. \end{tablenotes} \end{threeparttable} \end{table}
table[table omitted — 1,265 chars of source]
table[table omitted — 2,167 chars of source]