Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
35,711 characters · 8 sections · 19 citation commands
MTE with Misspecification
\normalem
\thispagestyle{empty}
Keywords: Marginal Treatment Effects, Misspecification, Weak Instruments.
\pagenumbering{arabic}
Marginal treatment effects (MTEs) have unified the identification theory of several policy parameters. While the MTE framework is essentially non-parametric,\footnote{Linearity is sometimes assumed to facilitate estimation. See, e.g., Appendix B in urzua2006} it is required that the recipient's participation into treatment follows a (generalized) Roy model. This is often referred to as additive separability: an “additive” comparison of costs and benefits determines selection. On the other hand, identification of the MTE is achieved via the local instrumental variable (LIV) approach (Heckman2001,Heckman2005). An excellent survey is provided by mogstad2018. An early effort to analyze MTE under misspecification can be found in the appendix of the seminal paper by Heckman2001. They consider a case where the additive separability in the selection equation does not hold. The most serious consequence is that the LIV approach does not identify the MTE curve.
In this paper we analyze a different type of misspecification. We model a situation in which, under additive separability, a proportion of the population does not take into account the instrumental variable when deciding whether to take up treatment or not. We refer to them as non-responders. To analyze the resulting bias, we define a pseudo-MTE curve which results from the LIV approach. Under no misspecification, the pseudo-MTE curve would coincide with the MTE curve. The resulting bias can be interpreted as a location-scale change of the MTE curve, parameterized by the proportion of non-responders and their propensity score.
We have two main results. The first one shows that the ability to recover the conditional average treatment effect (CATE) for the subpopulation of responders depends on the proportion of non-responders only through the support of the responders' propensity score. Indeed, when the support of the propensity score is the unit interval, it is possible to identify the CATE without having to recover the true MTE curve in the first place. In a nutshell, ignoring misspecification and integrating under the pseudo-MTE curve over the support of observed propensity score yields the correct CATE for the subpopulation of responders.
While the previous identification result for the CATE is independent of the proportion of non-responders, this is not true of the MTE curve and other parameters derived from it such as LATE and MPRTE. However, in our second result, we show how to recover the MTE curve for responders by undoing the location-scale change induced by the presence of non-responders. The correction is based on an estimate of the support of the propensity score and requires only observable data. It gives an estimator of the policy parameter of interest that is simple to implement. Cases where the propensity score is fully supported are relevant in practice. For a recent example, see the survey approach of Briggs2020 the probability of having a child is supported on the full unit interval.
Recently, kedadni2021 and vitor2021 focus on the effect of measurement error in treatment status on the MTE curve. We complement such results by noting that a simple change to our setup can cover the case of misclassification. In a setting where treatment status is misclassified, the observed outcome is generated with the true treatment status. In our setting of misclassification, the observed outcome can be regarded as a mixture of responders and non-responders. The proportion of non-responders is analogous to the proportion of misreporters. Indeed, our results also hold if instead of having a fraction of non-responders, we have a fraction of misreporters.
Another consequence of the presence of non-responders in the sample is that the effect of the instrumental variable on the propensity score is attenuated. Motivated by this, we model a situation where the proportion of non-responders approaches 1, analogous to the setting of weak instruments of stockstaiger1997. Thus, we can derive weak-instrument-like asymptotic distributions for the parameters derived from the MTE curve.
The rest of the paper is organized as follows: section (ref) introduces the model; section (ref) contains the main identification results; section (ref) provides bounds for the case where the propensity score is not fully supported in the unit interval; section (ref) traces the connection to the weak IV literature; and section (ref) concludes. While this paper only deals with identification, we expect to extend our results to cover estimation and inference.
In this section we introduce our model for misspecification in the MTE framework (Bjorklund1987, Heckman2001,Heckman2005). We analyze the consequences of misspecification from the identification point of view.
We start with a general non-separable potential outcome model
where $D^*$ is the observed treatment status, $X$ are observable covariates with support denoted by $\mathcal X$, and $\left\{Y(0),Y(1)\right\}, Y$ are potential and observed outcomes, respectively. The functions $h_0$ and $h_1$ are unknown.
We model misspecification as a situation where there are two types of individuals: responders and non-responders. Responders select into treatment taking into account the incentives in Z. Their selection equation is given by $ D=\mathds 1\left\{ \mu(X,Z)\geq V\right\}$. On the other hand, non-responders do not react to incentives in Z at all. Their selection equation is given by $\tilde D=\mathds 1\left\{ \tilde \mu(X)\geq \tilde V\right\}$. Notice how $Z$ is not featured in $\tilde{\mu}(\cdot)$. For the non-responders, Z fails the relevance condition of the standard MTE model.
Let $S$ be the latent status of an individual: $S=1$ for a responder and $S=0$ for a non-responder. The observed treatment status $D^*$ is given by:
We allow for the proportion of non-responders may vary with $X$. To this end, we define $\delta_X = \Pr(S=0|X)=\Pr(D^*= \tilde{D}|X)$. Thus, for every subpopulation with characteristics $X=x$ there is a proportion $\delta_x = \Pr(S=0|X=x) \in [0,1)$ of non-responders. We consider values where $\sup_{x \in \mathcal{X}} \delta_x < 1$ to avoid a situation where no-one responds to the instrumental variable.
The econometrician observes a cross section of $(Y_i, D^*_i, X_i, Z_i)$. When $\delta_X=0$ almost surely, then $D^*=D$ and we are in the familiar MTE framework of Heckman2001,Heckman2005. Otherwise, if $\delta_X \neq 0$ almost surely, for an observation of $D^*_i$, we do not know whether we are observing the treatment status of a non-responder or of a responder. That is, it is unknown if we are observing $D_i$ or $\tilde D_i$.
Assumption (ref) states that once we control for $X$, the latent status of a individuals does not vary with the instrumental variable Z.
Note that, for the subpopulation of non-responders, the instrument is valid but totally irrelevant. The larger the value of $\delta_x$, the “weaker” the instrument $Z$, since most participants with $X=x$ are non-responders. With the exception of the requirement that $\tilde V\perp Z\| X$, these are the same conditions of Heckman2001,Heckman2005. Our additional requirement covers the subpopulation of non-responders: neither the “cost” of treatment $\tilde V$ nor the “benefit” $\tilde \mu(X)$ depend on $Z$ when conditioned on $X$.
The misclassification structure of Equation (ref) allows to define three different propensity scores. An observed/identified one which is based on the observables $(D^*,X,Z)$, and two latent/unobserved propensity scores: one for the reponders and one for the non-responders. Formally, they are given by
The next result takes (mainly) advantage of Assumption (ref) to derive a useful linear relation between them.
For a fixed $X=x$, the result in Lemma (ref) shows that the observed propensity (still random through $Z$) is a linear transformation of the propensity score for the responders. If, additionally, we take two different values of $Z$, for example $z$ and $z'$, we can remove the contribution of $\tilde P(X)$, which is invariant with respect to $z$ and obtain\footnote{We write $P^*(x,z)$ for $\Pr(D^*=1|X=x,Z=z)$, and $P(x,z)$ for $\Pr(D=1|S=1,X=x,Z=z)$.}
Equation (ref) says that the changes on the observed propensity score induced by varying $Z$ are proportional to the changes on the true propensity score induced by varying $Z$. Thus, if we knew $\delta_x$, we could recover the change in the propensity score for the responders. When $Z$ is continuous, we can take a limiting version of this argument, e.g., as $z'\to z$, to obtain
Both the discrete (equation (ref)), and the continuous (equation(ref)) change in the propensity score play a role in the relationship between the MTE curve (defined below) and certain parameters of interest.
For the subpopulation of responders, the standard MTE framework holds. This motivates us to define an MTE curve for this subpopulation. In doing so, we are implicitly assuming that this is our object of interest. The reason for this is that many times we can also control the instrumental variable $Z$. Thus, to asses the effects of manipulations of $Z$ we look at the MTE curve for responders.
Let $\mathcal P_x$ and $\mathcal P^*_x$ denote the support of $P(x,Z):=\Pr(D=1|X=x,Z)$ and $P^*(x,Z):=\Pr(D^*=1|X=x,Z)$ respectively. For the subpopulation of responders, we rewrite the selection equation as $D=\mathds 1\left\{ P(X,Z)\geq U_D \right\}$ where $U_D\sim U_{(0,1)}$.\footnote{This follows from $D=\mathds 1\{ F_{V|S,X,Z}(\mu(X,Z)|1,X,Z)\geq F_{V|S,X,Z}(V|1,X,Z)\}$. Noting that by assumptions (ref).((ref)) and (ref), we have $D=\mathds 1\{ P(X,Z)\geq F_{V|S,X}(V|1,X)\}$. Finally, we take $U_D:=F_{V|S,X}(V|1,X)$.} Thus, we define the MTE curve for responders as $$\text{MTE}(u,x):=\mathbb E\left[Y(1)-Y(0)|S=1,U_D=u,X=x\right].$$ By the LIV approach we have the following equivalence result:\footnote{See Heckman2001 for sufficient conditions.}
Since we do not observe $P(X,Z)$, this is not an identification result in our setting. In a similar fashion, we define the following pseudo-MTE curve:
We emphasize that the pseudo-MTE curve is indexed by $\delta_x$ because it depends implicitly on the proportion of the nonresponders. From the data only, we can only compute $\text{MTE}^*(u,x;\delta_x)$, not $\text{MTE}(u,x)$. The pseudo-MTE curve is the curve that would be mistakenly taken to be the MTE curve. Indeed, in the absence of non-responders, $\text{MTE}^*(u,x;0)=\text{MTE}(u,x)$. If non-responders are present in the $X=x$ subpopulation, that is if $\delta_x> 0$, the observed $\text{MTE}^*(u,x;\delta_x)$ does not identify $\text{MTE}(u,x)$. In another words, the LIV approach is biased. We can now fully characterize the bias induced by $\delta_x$ on the MTE curve.
Lemma (ref) shows that the bias is in the form of both location and scale. Equation (ref), which is equivalent to Equation (ref),\footnote{Note the changes in the domain of integration between (ref) and (ref).} shows that $\text{MTE}^*$ is obtained by changing the location from $u$ to $u-\delta_x\tilde P(x)$, and rescaling by $(1-\delta_x)^{-1}$. Thus, as in a location-scale family of densities, we can regard $\text{MTE}^*$ as a family of curves, defined over $\mathcal P^*_x$, which is indexed by $\delta_x$ and $\tilde P(x)$.
We now introduce our two main results. We show that, for any subpopulation $X=x$ where the instrument is strong enough to induce a propensity score supported on the full unit interval $[0,1]$, the associated $CATE(x)$ can be identified for responders. This is true even if the $MTE^*(u,x,\delta_x)$ curve is biased for $MTE(u,x)$. We note that the identified $CATE(x)$ parameters corresponds to the subpopulation of responders.
Assumption (ref) says that the incentive in the instrument $Z$ is strong enough to induce any individual in the $X=x$ subpopulation into or out of treatment. Perhaps surprisingly, the $\text{CATE}(x)$, can be recovered only by resorting to the full support assumption. That is, to correctly compute the $\text{CATE}(x)$ we do not need to recover the true MTE curve for responders.
Unfortunately, the automatic “de-biasing" in Theorem (ref) does not hold for the other policy parameters that can be obtained via the MTE curve. On the other hand, we show that the full support assumption can be used to identify $\delta_x$ which allows an explicit “de-biasing" procedure. Given that $\mathcal P^*_x:= [\underline{p_x^*} , \overline{p_x^*} ] =[\delta_x\tilde P(x), (1-\delta_x)+ \delta_x\tilde P(x)]$ we can actually identify both $\delta_x$ and $\tilde P(x)$. It follows then from Lemma (ref) that we can recover the $\text{MTE}(u,x)$ curve.
The intuition for this result is simple. Because the original propensity score $P(Z,x)$, for any fixed $x$, is supported on the unit interval, the observed support $\mathcal P^*_x=[\underline{p_x^*} , \overline{p_x^*} ]$ will contain enough information to identify $\delta_x$. This is summarized Figure (ref).
Having identified $\delta_x$, then we use Equation (ref) to identify the MTE curve.
This corollary provides the correct “de-biasing” to be performed on the observed MTE curve to match the true MTE curve. However, it is possible to recover parameters that are based on the MTE curve without having to recover the MTE curve in the first place. We provide two examples.
In the previous examples, proceeding as if there were no misspecification, yields biased parameters. Thus, the automatic “de-biasing” in CATE is the exception rather than the rule.
Instead of assuming full support, now we allow for limited support of the propensity score $P(x,Z)$, but we still require that it is an interval.
Under Assumption (ref), and using (ref), we have that the observed support of $P^*(X,Z)$ is
Taking the difference we obtain that $\overline{p_x^*} - \underline{p_x^*}=(1-\delta_x)(\overline{p_x} -\underline{p_x} )$. Since $\overline{p_x} -\underline{p_x}\leq 1$, then $\overline{p_x^*} - \underline{p_x^*}\leq (1-\delta_x)$, so that a lower bound for $\delta_x$ is $\delta_x\geq 1-(\overline{p_x^*} - \underline{p_x^*})$.
In general, it is not possible to provide an upper bound for $\delta_x$. This is similar to the case of misclassification. Following that literature (see Assumption 4 in kedadni2021, and references therein), we assume it is known that for some $\overline \delta_x$: $\delta_x\leq \overline \delta_x<1$. Thus, we can write $1-(\overline{p_x^*} - \underline{p_x^*})\leq \delta_x\leq \overline\delta_x.$ The correction factor in Examples (ref) and (ref) is $(1-\delta_x)$. Now, it bounded by $ 1-\overline \delta_x \leq 1-\delta_x\leq \overline{p_x^*} - \underline{p_x^*}$. Thus, we can bound both LATE and MPRTE using this:
and
Naturally, if $\overline \delta_x$ is not known, we can only provide upper bounds.
Again, we stress that is not necessary to bound the MTE curve in the first place. Such a bound can be complicated to obtain since, by Lemma (ref), $\delta_x$ enters in three different ways in the observed MTE curve.
We can frame our model as the triangular scheme of stockstaiger1997 and consider a sequence $\left\{\delta_{x,n}\right\}_{n=1}^{\infty}$ such that $\lim_{n\to\infty}\delta_{x,n}=1$ at a certain rate as $n\to\infty$. Thus, as $n\to\infty$, the instrument becomes irrelevant in the model. A possible indicator of the presence of a large value of $\delta_{x,n}$ can be the average derivative of the observed propensity score. This equals an attenuated version of the average derivative of the true propensity score. For a given value of $\delta_{x,n}$, by equation (ref), we have
Thus a “small” value can be an indication that $\delta_{x,n}$ is close to 1. This is similar to a first stage regression in the linear model. We take the derivative with respect to $z$ to get rid of the propensity score that does not respond to $Z$. We average, because this likely to be a non-linear expression. Thus, $(1-\delta_{x,n})$ can be thought of as the counterpart of $C/\sqrt T$ in the notation of stockstaiger1997. Indeed, define
We have
and
Thus,
which is the covariance between the instrument and treatment status for the responders with $X=x$. To see the role of the rate at which $\delta_{x,n}$ converges to 1, suppose for a second that we know the functional form of $P^*(x,Z)$, and we estimate the average derivative using a sample mean:
Then
In order to investigate possible discontinuities in the limiting distributions, we follow kuersteiner2002, and we let $(1-\delta_{x,n})=n^{\nu_{x}}$, for $\nu_{x}<0$. We obtain
Then, we obtain a degenerate limit:
Now consider the MPRTE. Recall that, by Example (ref), under the full support guaranteed by Assumption (ref),
Assume that, if $\delta_x=0$, there exists $\hat{\text{MPRTE}}(x)$, a $\sqrt n$-consistent estimator of $\text{MPRTE}(x)$ such that
Thus, if $\nu_{x}=-1/2$, then $\hat{\text{MPRTE}}^*(x)$ does not converge in probability. In future work, we will use these results to construct confidence intervals for the parameters of interest.
In this paper we use the MTE framework to model a proportion of individuals who do not respond to the incentives of the instrumental variable. We show that in the special case where the observed propensity score is fully supported on the unit interval, i) the CATE is automatically identified regardless of the non-responders, and ii) we can identify the proportion of non-responders and use it to recover the MTE curve, and we can recover any parameter associated with it. We show that for some parameters, such as LATE and MPRTE, it is even possible to bypass the recovery of the MTE curve, and directly recover these parameters. Moreover, if the propensity has limited support, we find bounds for the LATE, the MPRTE, and the MTE curve. When we let the proportion of non-responders approach 1 at a certain rate, the framework resembles that of weak instruments. In future research we hope to leverage the results in this literature to construct valid confidence intervals for the MTE curve and related parameters.