EconBase
← Back to paper

Doubly weighted M-estimation for nonrandom assignment and missing outcomes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,945 characters · 16 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Doubly weighted M-estimation for nonrandom assignment and missing outcomes

\begingroup \footnote{$^\ast$I am grateful to Jeffrey M. Wooldridge, Steven Haider, Ben Zou, and Kenneth Frank. Special thanks to Tim Vogelsang, Wendun Wang, Alyssa Carlson, Christian Cox, Tymon S{\l}oczy{\'n}ski, and seminar & conference participants, for insightful comments and suggestions on earlier drafts of this paper. } \addtocounter{footnote}{-1} \endgroup

\begingroup \footnote{$^\dagger$Department of Econometrics and Business Statistics, Monash University. Email: [email removed], Webpage: \href{https://www.anegi.net/home}{www.anegi.net}} \addtocounter{footnote}{-1} \endgroup

spacing{1.0} \begin{abstract} This paper proposes a new class of M-estimators that double weight for the twin problems of nonrandom treatment assignment and missing outcomes, both of which are common issues in the treatment effects literature. The proposed class is characterized by a `robustness' property, which makes it resilient to parametric misspecification in either a conditional model of interest (for example, mean or quantile function) or the two weighting functions. As leading applications, the paper discusses estimation of two specific causal parameters; average and quantile treatment effects (ATE, QTEs), which can be expressed as functions of the doubly weighted estimator, under misspecification of the framework's parametric components. With respect to the ATE, this paper shows that the proposed estimator is doubly robust even in the presence of missing outcomes. Finally, to demonstrate the estimator's viability in empirical settings, it is applied to calonico2017women's reconstructed sample from the National Supported Work training program. \end{abstract} Keywords: { Unconfoundedness, Missing at random, Double weighting, M-estimation, Treatment effects} \\ JEL Classification: C13, C18, C31

Introduction

When interest lies in causal inference, the prevalence of missing data poses a major identification challenge. A common issue is that the outcome of interest is missing for some proportion of the sample. In this case, the complete data method that drops observations with missing outcomes is widely used. While dropping is practically convenient, it not only leads to substantial loss of information but more importantly creates a nonrandom sample for estimation. In turn, dropping can generally lead to inconsistent treatment effect estimates. This paper proposes an estimator that double weights for the twin problems of nonrandom treatment assignment and missing outcomes by using information on covariates.

Weighting has been used extensively in both the missing data [horvitz1952generalization, robins1994estimation, robins1995semiparametric, wooldridge2007inverse] and treatment effect [rosenbaum1983central, hahn1998role, hirano2001estimation, firpo2007efficient, sloczynski2018general] literatures. However, a weighting approach that corrects for general missingness in the outcome to estimate treatment effects using observational data is yet to be proposed. Previous studies have considered weighting to deal with specific missing data issues such as attrition and non-response in the presence of endogenous treatment selection [frolich2014treatment, huber2014treatment, fricke2020endogeneity]. Typically, the identification argument in these papers is based on one or more instruments with discussion centered around estimation of average treatment effects.

This paper introduces inverse probability weighting alongside propensity score (PS) weighting in a general M-estimation framework to address two prevalent problems in the causal inference literature. Moreover, the objective function being solved is permitted to be non-smooth in the underlying parameters thereby covering both average and quantile treatment effects. A key feature of the proposed estimator is its robustness to parametric misspecification in either a conditional model of interest (such as mean or quantile) or the two weighting functions. In addition, the ATE estimator which uses the proposed strategy is shown to be `doubly robust' [sloczynski2018general] even in the presence of missing outcomes.

The key identifying assumptions for consistency of the doubly weighted estimator of a population level parameter are unconfoundedness\footnote{This is a widely used assumption in the treatment effects literature and is known by a variety of names such as exogeneity, ignorability, selection on observables, and conditional independence assumption (CIA).} and missing at random. Put differently, the two restrictions imply that the treatment assignment and missing outcomes mechanisms are as good as randomly assigned after conditioning on covariates. With respect to missingness, the mechanism also allows sample observability to depend on the treatment status. As such it allows for differential non-response, attrition, and even non-compliance to the extent that conditioning variables predict it.

For many observational studies, unconfoundedness may be a reasonable assumption. Previous literature has found several situations where such an assumption is tenable, especially when pre-treatment values of the outcome variable are available. For example, lalonde1986evaluating and hotz2006evaluating have shown that controlling for pre-training earnings alone reduces significant bias between non-experimental and experimental estimates. The literature assessing teacher impact on student achievement has reported similar findings with pre-test scores [chetty2014measuring, kane2008estimating, and shadish2008can], indicating the plausibility of unconfoundedness in such settings.

Estimation then follows in two steps. The first step estimates the treatment and missing outcome probabilities using binary response maximum likelihood\footnote{As a practical matter, researchers typically follow the convention of estimating these probabilities as flexible logit functions.} and second step plugs in the estimated probabilities as weights to solve a general objective function. Given the parametric nature of the first and second steps, this paper highlights a robustness property which allows the estimator to remain consistent for a parameter of interest under misspecification of either a conditional model or the two probability weights. Consequently, the asymptotic theory in this paper distinguishes between these two halves. The first half focuses on misspecification of either a conditional expectation function (CEF) or a conditional quantile function (CQF), whereas the second half considers misspecification in the weighting functions.

As illustrative examples, the paper discusses robust estimation of two specific causal parameters, namely, the ATE and QTEs, expressed as functions of the doubly weighted estimator. Consistent estimation of the ATE is achievable under both misspecification scenarios. Of particular interest is the case when the conditional mean function is misspecified. For estimation of quantile treatment effects, the paper considers three different parameters, namely, conditional quantile treatment effect (CQTE), a linear approximation to CQTE, and unconditional quantile treatment effect (UQTE), each of which may be of interest to the researcher depending on whether features of the conditional or unconditional outcomes distribution are of interest. Simulations show that the doubly weighted ATE and QTE estimates have the lowest finite sample bias compared to alternatives which ignore one or both problems.\footnote{Such as the unweighted estimator which drops missing outcomes and does not weight or the ps-weighted estimator which drops the missing data and weights by the propensity score to correct for nonrandom assignment.}

Finally, the proposed method is applied to estimate average and distributional impacts of the National Supported Work (NSW) training program on earnings for the Aid to Families with Dependent Children (AFDC) target group. The sample is obtained from calonico2017women who recreate Lalonde's within-study analysis for the AFDC women. The idea behind choosing this empirical application is to utilize the presence of experimental and non-experimental comparison groups for evaluating whether the strategy of double weighting brings us close the experimental benchmark relative to other alternatives. The paper finds that the empirical bias for the doubly weighted estimate is much smaller than that for the unweighted estimate.

The rest of this paper is structured as follows. Section (ref) describes the basic potential outcomes framework and provides a short description of the population models with an introduction to the naive unweighted estimator. Section (ref) discusses the treatment assignment and missing outcome mechanisms which leads us directly to the identification lemma. Section (ref) develops the first half of the asymptotic theory for the doubly weighted estimator with a focus on misspecification of a conditional feature of interest. This half also requires the weights to be correct for delivering parameter identification. In contrast, section (ref) considers the other half where a conditional model of interest is correctly specified but the weights may be misspecified. Identification here relies on the parameter solving a conditional problem. Section (ref) studies the specifics of robustness for estimating ATE and QTEs in rigorous detail. Section (ref) provides supporting Monte Carlo evidence under three interesting cases of misspecification; correct conditional model with misspecified weights, misspecified conditional model with correct weights, and misspecified model and weights. Section (ref) applies the proposed method to job training data from calonico2017women and section (ref) concludes with directions for future research.

Potential outcomes and the population models

Consider the standard Neyman-Rubin causal model. Let $Y(1)$ and $Y(0)$ denote potential outcomes corresponding to the treatment and control states and let $W$ be an indicator for whether an individual received the treatment. Then observed outcome is

align[align omitted — 63 chars of source]

Also, let $\mathbf{X}$ be a vector of pre-treatment characteristics which includes an intercept.\footnote{As mentioned in negiwooldridge2020, $\mathbf{X}$ may include functions of covariates such as levels, squares, and interactions which will be chosen by the researcher. The dimension of the covariate vector is assumed fixed and does not grow with the sample size.} Some feature of the distribution of $(Y(g),\mathbf{X})\subset \mathfrak{R}^{\text{M}}$ is assumed to depend on a finite $P_g\times 1$ vector $\bm{\theta_g}$, contained in a parameter space $\bm{\Theta_g}\subset \mathfrak{R}^{P_g}$.\footnote{For generality, the dimension of $\bm{\theta_g}$ is allowed to be different for the treatment and control group problems and is also different than the dimension of $\mathbf{X}$, where $\mathbf{X}\in\mathfrak{X} \subset \mathfrak{R}^{\text{dim}(\mathfrak{X})}$} Let $q(Y(g), \mathbf{X}, \bm{\theta_g})$ be an objective function that depends on outcomes, covariates, and the parameter vector, $\bm{\theta_g}$. Then, the parameter of interest is defined to be a solution to the following M-estimation problem.

assumption{(Identification of $\bm{\theta_g^0}$)} The parameter vector $\bm{\theta_g^0}\in \bm{\Theta_g}$ is a unique solution to the population minimization problem \begin{equation} \underset{\bm{\theta_g}\in\bm{\Theta_g}}{\mathrm{min}}\mathbb{E}\left[q(Y(g),\mathbf{X},\bm{\theta_g})\right] \end{equation} for each $g=0,1$. \qed

Examples include the smooth ordinary least squares (OLS) function, $q(Y(g),\mathbf{X},\bm{\theta_g})= (Y(g)-\mathbf{X}\bm{\theta_g})^2$ or the non-smooth conditional quantile regression (CQR) of KoenkerBassett1978, $q(Y(g),\mathbf{X},\bm{\theta_g})= c_{\tau}(Y(g)-\mathbf{X}\bm{\theta_g})$\footnote{For a random variable $u$, $c_{\tau}(u) = (\tau-\mathbf{1}\{u<0\})u$ is the asymmetric loss function for estimating quantiles and $\mathbf{1}\{\cdot\}$ is an indicator function.}. Other examples of $q(\cdot)$ can be log-likelihood and quasi-log-likelihood (QLL) functions.

An implicit point in assumption (ref) is that $\bm{\theta_g^0}$ is not assumed to be correctly specified for a conditional feature like a conditional mean, variance, or even the full conditional distribution. It simply requires $\bm{\theta_g^0}$ to uniquely minimize the population problem in ((ref)). If $\bm{\theta_g^0}$ is correctly specified for any of the above mentioned quantities, then the parameter is of direct interest to researchers. However, if $\bm{\theta_g^0}$ is misspecified for any of these distributional features, assumption (ref) guarantees a unique pseudo true solution, $\bm{\theta_g^\ast}$ [white1982maximum]. In the case of misspecification, determining whether $\bm{\theta_g^\ast}$ is meaningful will depend on the conditional feature being studied and the estimation method used. For example, in the case of OLS, $\bm{\theta_g^0}$ will index a linear projection if one is agnostic about linearity of the CEF. Angristetal2006 establish analogous approximation properties for quantiles, where a misspecified CQF can still provide the best weighted mean square approximation to the true CQF.

Let $`S$' be a binary indicator such that $S=1$ if the outcome is observed and $S=0$ otherwise. The objective of this paper is to consistently estimate $\bm{\theta_g^0}$. In the presence of missing outcomes, a common empirical strategy is to solve the following M-estimation problems for the treatment and control groups, respectively.

equation[equation omitted — 317 chars of source]

Let us refer to the estimator that solves ((ref)) as the unweighted M-estimator and denote it as $\bm{\hat{\theta}_{g}^u}$. This estimator uses the available sample after dropping the missing data to estimate $\bm{\theta_g^0}$. Using the reverse analogy principle, $\bm{\hat{\theta}_{g}^u}$ will be consistent for $\bm{\theta_g^0}$ if it solves the population analogue of ((ref)), which may not be true. As an example, consider

equation*[equation* omitted — 153 chars of source]

In this case, even if the treatment is randomly assigned, missingness may still be correlated with the treatment, observable factors, or both. Hence, the population first order condition for the selected sample, $\mathbb{E}[S\cdot W \cdot \mathbf{X}^\prime U(g)]$, is not zero even though $\mathbb{E}[\mathbf{X}^\prime U(g)] = \mathbf{0}$. So identification of $\bm{\theta_g^0}$ is now confounded on two grounds; nonrandom assignment which renders the treatment and control groups incomparable and missing outcomes which violates the ‘random sampling’ assumption. The next section discusses the identification approach taken in this paper.

Identification of parameter of interest

Without imposing any structure on the assignment and missingness mechanisms in the population, estimating $\bm{\theta_g^0}$ remains difficult. To proceed with identification, I assume that the treatment is unconfounded on covariates.\footnote{Like most other assumptions, unconfoundedness is non-refutable. For methods that indirectly test for its validity, see huber2015test, de2014testing, and heckman1989choosing.} Formally,

assumption{(Strong ignorability)} Assume, \begin{equation} \{Y(0), Y(1)\mathpalette{\independenT}{\perp} W\} \bf{|} \ \mathbf{X} \end{equation} \begin{enumerate} • The vector of pre-treatment covariates, $\mathbf{X}$, is always observed for the entire sample. • For all $\mathbf{x}\in \mathfrak{X} \subset \mathfrak{R}^{\mathrm{dim}(\mathfrak{X})}$, define $p(\mathbf{x}) = \mathbb{P}(W=1|\mathbf{X}=\mathbf{x})$ such that $p(\mathbf{x})>\kappa$ for a constant $\kappa>0$. \qed \end{enumerate}

Equation ((ref)) indicates that conditioning on covariates is enough to parse out any systematic differences that may exist between the treatment and control groups. One advantage of unconfoundedness is that, intuitively, it has a better chance of holding once we control for a rich set of variables in $\mathbf{X}$.\footnote{For example, hirano2001estimation control for a rich set of prognostic factors to justify unconfoundedness while estimating the effects of right heart catheterization (RHC) on survival rates of patients.} Note that unconfoundedness not only includes cases where the treatment is a deterministic function of the covariates, for example stratified (or block) experiments, but also cases where the treatment is a stochastic function of covariates. Part i) requires that we observe these covariates for all individuals. Part ii) is an overlap condition which ensures that for all values of $\mathbf{x}$ in $\mathfrak{X}$, we observe units in both the treatment and control groups.\footnote{Methods for checking overlap involve calculating normalized sample average differences for each covariate and checking the empirical distribution of propensity scores.}

With respect to the missing outcomes mechanism, I assume selection on observables

assumption{(Missing at Random (MAR))} Assume, \begin{equation} \{Y(0), Y(1)\mathpalette{\independenT}{\perp} S\}{\bf{|}}\ \mathbf{X}, W \end{equation} \begin{enumerate} • In addition to $\mathbf{X}$, $W$ is always observed for the entire sample. • For each $(\mathbf{x},w)\in (\mathbf{X},W) \subset \mathfrak{R}^{\mathrm{dim}(\mathfrak{X})+1}$, define, $r(\mathbf{x},w) \equiv \mathbb{P}(S=1|\mathbf{X}=\mathbf{x},W=w)$ such that $\eta<r(\mathbf{x},w)<1$ for a constant $\eta>0$ and $w=0,1$. \qed \end{enumerate}

Equation ((ref)) states that conditional on covariates and the treatment status, the individuals whose outcomes are missing do not differ systematically from those who are observed. This implies that adjusting for $\mathbf{X}$ and $W$ renders the outcomes as good as randomly missing. In the statistics literature, this assumption is known as MAR and represents a mechanism wherein missingness only depends on observables and not on the missing values of the variable itself [little2019statistical]. Special cases covered under this mechanism are patterns such as missing completely at random (MCAR) and exogenous missingness considered in wooldridge2007inverse. Allowing the missingness probability to be a function of the treatment indicator is particularly useful in cases of differential nonresponse. For instance, in NSW, people assigned to the treatment group were less likely to drop out of the program compared to the control group. In such cases, covariates alone may not be sufficient for predicting missingness. To the extent that being observed in the sample is predicted by $\mathbf{X}$ and $W$, assumption (ref) can accommodate non-observability due to sampling design, item non-response, and attrition in a two period panel.\footnote{For the case of attrition, one must assume that second period missingness is ignorable conditional on initial period covariates and the treatment status.}

Part i) of the above assumption ensures that $\mathbf{X}$ and $W$ are fully observed and part ii) again imposes an overlap condition. It states that there is a positive probability of observing people in the sample for a given $\mathbf{X}$ and $W$.

Then solving the doubly weighted population problem given below is the same as solving the original M-estimation problem in ((ref)). The following lemma establishes this equality

lemma{(Identification)} Given assumptions (ref), (ref), (ref), assume i) $q(Y(g),\mathbf{X},\bm{\theta_g})$ is a real valued function for all $(Y(g), \mathbf{X})\subset \mathfrak{R}^M$ ii) $\mathbb{E}\left[\left. |q(Y(g),\mathbf{X},\bm{\theta_g})| \right. \right]< \infty$ for all $\bm{\theta_g} \in \bm{\Theta_g}$, $g=0,1$, then \begin{equation} \mathbb{E}\big[\omega_g\cdot q\left(Y(g), \mathbf{X}, \bm{\theta_g}\right)\big]= \mathbb{E}\big[q\left(Y(g), \mathbf{X}, \bm{\theta_g}\right)\big] \end{equation} where $\displaystyle \omega_1 = \frac{S\cdot W}{r(\mathbf{X},W)\cdot p(\mathbf{X})}$, $\displaystyle \omega_0 = \frac{S\cdot (1-W)}{r(\mathbf{X},W)\cdot (1-p(\mathbf{X}))}$. \qed

The proof uses two applications of the law of iterated expectations (LIEs) with unconfoundedness and MAR to arrive at the above result. It implies that one can now address the identification issue due to nonrandom assignment and missing outcomes by solving the doubly weighted population problem.\footnote{Define { $q\left(Y, \mathbf{X}, \bm{\theta}\right)=q(Y(1), \mathbf{X}, \bm{\theta_1}) \text{ for } W=1$ and $q(Y(0), \mathbf{X}, \bm{\theta_0}) \text{ for } W=0 $}, then $\mathbb{E}[\omega_g\cdot q\left(Y, \mathbf{X}, \bm{\theta}\right)]\equiv \mathbb{E}[\omega_g\cdot q\left(Y(g), \mathbf{X}, \bm{\theta_g}\right)]$ which makes it function of the observed random vector $\{(Y_i,\mathbf{X}_i,W_i,S_i): i=1,2,\ldots,N\}$.}

Asymptotic theory under weak identification

Lemma (ref) is important for us as it helps to illustrate the role of double weighting in dealing with the two issues at hand. However, to operationalize this argument, we first need to estimate $r(\mathbf{X},W)$ and $p(\mathbf{X})$ before introducing the estimator and studying its asymptotic properties.

The following assumptions posit that we have a correctly specified model for the two probabilities and that we estimate them using binary response maximum likelihood. Since both $W$ and $S$ are binary responses, estimation of $\bm{\gamma_0}$ and $\bm{\delta_0}$ using MLE will be asymptotically efficient under correct specification of these functions. Consistency and asymptotic normality for $\bm{\gamma_0}$ and $\bm{\delta_0}$ follow from theorems 2.5 and 3.3 of newey1994large.

assumption{(Correct parametric specification of propensity score)} Assume that i) There exists a known parametric function $G(\mathbf{X},\bm{\gamma})$ for $p(\mathbf{X})$ where $\bm{\gamma}\in \bm{\Gamma}\subset \mathfrak{R}^I$ and $0<G(\mathbf{X},\bm{\gamma})<1$ for all $\mathbf{X} \in \mathcal{X}$, $\bm{\gamma}\in \bm{\Gamma}$; ii) There exists $\bm{\gamma_0}\in \bm{\Gamma}$ s.t. $p(\mathbf{X})=G(\mathbf{X},\bm{\gamma_0})$; iii) $\bm{\hat{\gamma}}$ is the binary response maximum likelihood estimator that solves \begin{equation} \underset{\bm{\gamma}\in \bm{\Gamma}}{max}\sum_{i=1}^{N}\{W_ilogG(\mathbf{X}_i,\bm{\gamma})+(1-W_i)log(1-G(\mathbf{X}_i,\bm{\gamma}))\} \end{equation}

\qed

assumption{(Correct parametric specification of missing outcomes probability)} Assume that i) There exists a known parametric function $R(\mathbf{X},W,\bm{\delta})$ for $r(\mathbf{X},W)$ where $\bm{\delta}\in \bm{\Delta}\subset \mathfrak{R}^{K}$ and $R(\mathbf{X},W,\bm{\delta})>0$ for all $\mathbf{X} \in \mathcal{X}$, $\bm{\delta}\in \bm{\Delta}$; ii) There exists $\bm{\delta_0}\in \bm{\Delta}$ s.t. $r(\mathbf{X},W) = R(\mathbf{X},W,\bm{\delta_0})$; iii) $\bm{\hat{\delta}}$ is the binary response maximum likelihood estimator that solves \begin{equation} \underset{\bm{\delta}\in \bm{\Delta}}{max}\sum_{i=1}^{N}\{S_ilogR(\mathbf{X}_i,W_i,\bm{\delta})+(1-S_i)log(1-R(\mathbf{X}_i,W_i,\bm{\delta}))\} \end{equation}

\qed

The influence function representations for $\bm{\hat{\gamma}}$ and $\bm{\hat{\delta}}$ can then be written as

equation[equation omitted — 370 chars of source]

where $\mathbf{d}_i$ and $\mathbf{b}_i$ are scores of the binary response log-likelihood problems in ((ref)) and ((ref)) evaluated at the probability limits $\bm{\gamma_0}$ and $\bm{\delta_0}$, respectively. The doubly weighted estimator is then defined as:

equation[equation omitted — 189 chars of source]

where $ \widehat{\omega}_{i1}=\frac{S_i\cdot W_i}{R(\mathbf{X}_i,W_i,\bm{\hat{\delta}})\cdot G(\mathbf{X}_i,\bm{\hat{\gamma}})}$ and $ \widehat{\omega}_{i0} = \frac{S_i\cdot (1-W_i)}{R(\mathbf{X}_i,W_i,\bm{\hat{\delta}})\cdot (1-G(\mathbf{X}_i,\bm{\hat{\gamma}}))}$ are the estimated weights for solving the treatment and control group problems, respectively.\footnote{When necessary, the estimated weights will also be denoted as $\omega_g(\bm{\hat{\delta}}, \bm{\hat{\gamma}})\equiv\widehat{\omega}_g$.}

Given the two-step nature of the estimation problem; first step uses binary response MLE for estimating the probability weights and second step solves an objective function using the first-step weights, the asymptotic theory utilizes results for two-step estimators with a non-smooth objective function to establish the large sample properties of $\bm{\hat{\theta}_g}$. The following theorem fills in the primitive regularity conditions for applying the uniform law of large numbers.

theorem{(Consistency)} Suppose assumption (ref) holds and that i) $\{(Y_i, \mathbf{X}_i, W_{i}, S_i);\ i=1,2,\ldots, N\}$ are $i.i.d$ draws satisfying assumptions (ref) and (ref); ii) $\bm{\Theta_g}$ is compact for $g=0,1$; iii) $G(\mathbf{X},\bm{\gamma})$ satisfies assumption (ref) and is continuous for each $\bm{\gamma}$ on the support of $\mathbf{X}$. Similarly, $R(\mathbf{X},W,\bm{\delta})$ satisfies assumption (ref) and is continuous for each $\bm{\delta}$ on the support of $(\mathbf{X},W)$; iv) $q(Y(g),\mathbf{X},\bm{\theta_g})$ is continuous at each $\bm{\theta_g}\in\bm{\Theta_g}$ with probability one; v) $ \mathbb{E}\bigg[\underset{\bm{\theta_g}\in\bm{\Theta_g}}{\textup{sup}}|q(Y(g), \mathbf{X},\bm{\theta_g})|\bigg] <\infty $. Then, $\bm{\hat{\theta}_g}\overset{p}{\rightarrow} \bm{\theta_g^0} $. \qed

The proof follows from verifying the conditions in Lemma 2.4 of newey1994large. Under the dominance condition given in v), uniform convergence of sample averages holds quite generally.

For establishing asymptotic normality, I provide primitive conditions for the general case of non-smooth objective functions. Let the score of $q(Y(g),\mathbf{X},\bm{\theta_g})$ at the true parameter, $\bm{\theta_g^0}$, be denoted as $\mathbf{h}(Y(g),\mathbf{X},\bm{\theta_g^0}) \equiv \mathbf{h}_g$ and suppose it exists with probability one. Let the population problem be denoted as

equation*[equation* omitted — 113 chars of source]

and the sample analogue be given as

equation*[equation* omitted — 142 chars of source]

where $\hat{\rho}_g= N_g/N$ and $N\hat{\rho}_g \rightarrow \infty$ as $\hat{\rho}_g \rightarrow \rho_g$.\footnote{The sampling fractions $N_0 = \sum_{i=1}^{N}S_i\cdot W_i$ and $N_1 = \sum_{i=1}^{N}S_i\cdot (1-W_i)$ are random which implies that $N=N_0+N_1$ is also random as opposed to being fixed ahead of time.} For the sake of asymptotics, we may ignore the division by $\hat{\rho}_g$. The main condition needed for establishing asymptotic normality is stochastic equicontinuity of the empirical process

equation[equation omitted — 227 chars of source]

which will be sufficient to guarantee uniform convergence of the objective function to its population counterpart.

theorem{(Asymptotic Normality)} In addition to the conditions mentioned in Theorem (ref), assume i) $\bm{\theta_g^0}\in$ $int(\bm{\Theta_g})$; ii) $q(Y(g),\mathbf{X},\bm{\theta_g})$ is continuously differentiable on $int(\bm{\Theta_g})$ with probability one; iii) $\frac{1}{N}\sum_{i=1}^{N}\widehat{\omega}_{ig}\cdot \mathbf{h}(Y_i(g),\mathbf{X}_i,\bm{\hat{\theta}_g})= o_p(N^{-1/2})$; iv) { $\mathbb{E}\bigg[\underset{\bm{\theta_g}\in\bm{\Theta_g}}{\textup{sup}}\lVert \mathbf{h}(Y(g), \mathbf{X},\bm{\theta_g})\rVert^2\bigg] <\infty $}; v) $G(\cdot, \bm{\gamma})$ and $R(\cdot,\bm{\delta})$ are both twice continuously differentiable on $int(\Gamma)$ and $int(\Delta)$, respectively; vi) { $\mathbb{E}\bigg[\underset{\bm{\delta}\in\bm{\Delta}}{\textup{sup}}\lVert \mathbf{b}(\mathbf{X}, W, S,\bm{\delta})\rVert^2\bigg]<\infty$}, { $\mathbb{E}\bigg[\underset{\bm{\gamma}\in\bm{\Gamma}}{\textup{sup}}\lVert \mathbf{d}(\mathbf{X}, W,\bm{\gamma})\rVert^2\bigg]<\infty$}; vii) $\mathbb{E}\big[\omega_{g}\cdot\mathbf{h}(Y(g),\mathbf{X},\bm{\theta_g})\big]$ is continuously differentiable on $int(\bm{\Theta_g})$; viii) { $\mathbf{H_g} \equiv \nabla_{\bm{\theta_g}} \mathbb{E}\left[\omega_{g}\cdot\mathbf{h}(Y(g),\mathbf{X},\bm{\theta_g^0})\right] $} is nonsingular; ix) $\{\bm{v}_N(\bm{\theta_g}): N\geq 1\}$ is stochastically equicontinuous. Then, \begin{equation*} \sqrt{N}(\bm{\hat{\theta}_g}-\bm{\theta_g^0})\overset{d}{\rightarrow} N\left(\bm{0}, \mathbf{H}^{-1}_{\mathbf{g}}\mathbf{\Omega_g}\mathbf{H}^{-1}_{\mathbf{g}}\right)\end{equation*} where $\mathbf{\Omega_g} = \mathbb{E}\big(\mathbf{l}_{ig}\mathbf{l}_{ig}^{\prime}\big) - \mathbb{E}\left(\mathbf{l}_{ig}\mathbf{b}_i^{\prime}\right)\mathbb{E}\left(\mathbf{b}_i\mathbf{b}_i^{\prime}\right)^{-1}\mathbb{E}\big(\mathbf{b}_i\mathbf{l}_{ig}^{\prime}\big)-\mathbb{E}\left(\mathbf{l}_{ig}\mathbf{d}_i^{\prime}\right)\mathbb{E}\left(\mathbf{d}_i\mathbf{d}_i^{\prime}\right)^{-1}\mathbb{E}\big(\mathbf{d}_i\mathbf{l}_{ig}^{\prime}\big)$ for each $g=0,1$ and $\mathbf{l}_{ig}\equiv \omega_{ig}\mathbf{h}_{ig}$ is score of the weighted objective function evaluated at $\bm{\theta_g^0}$.\qed

Sufficient primitive conditions for stochastic equicontinuity may be found in andrews1994empirical. The asymptotic variance expression derived above offers some interesting insights. First, the middle term, $\bm{\Omega_g}$, represents the variance of the residual from the population regression of the weighted score, $\mathbf{l}_{ig}$, on the two binary response scores, $\mathbf{b}_i$ and $\mathbf{d}_i$. Note that even though $\bm{\Omega_g}$ would involve covariance between the two MLE scores, that term is zero on account of the two scores being conditionally independent.

Second, the expression for $\bm{\Omega_g}$ has an efficiency implication for the second step estimate, $\bm{\hat{\theta}_g}$. When a researcher is only willing to assume identification of $\bm{\theta_g^0}$ in the unconditional sense, it is potentially more efficient to estimate the two weights even when they are known. To show this formally, let us assume that $p(\mathbf{X})$ and $r(\mathbf{X}, W)$ are known and $\bm{\tilde{\theta}_g}$ is the doubly weighted estimator that uses known weights, $\omega_g$. Then,

corollary{(Efficiency gain with estimated weights)} Under the assumptions of theorem (ref), \begin{equation*} \begin{split} \mathrm{Avar}\big[\sqrt{N}\big(\bm{\tilde{\theta}_g}-\bm{\theta_g^0}\big)\big]- \mathrm{Avar}\big[\sqrt{N}\big(\bm{\hat{\theta}_g}-\bm{\theta_g^0}\big)\big]&= \mathbf{H}^{-1} _{\mathbf{g}}\mathbf{\Sigma_g}\mathbf{H}_{\mathbf{g}}^{-1}- \mathbf{H}_{\mathbf{g}}^{-1}\mathbf{\Omega_g}\mathbf{H}^{-1} _{\mathbf{g}}\\ &=\mathbf{H}^{-1} _{\mathbf{g}}\mathbf{\left(\Sigma_g-\Omega_g\right)}\mathbf{H}^{-1} _{\mathbf{g}} \end{split} \end{equation*} is positive semi-definite and where $\mathbf{\Sigma_g} = \mathbb{E}(\mathbf{l}_{ig}\mathbf{l}_{ig}^\prime)$. \qed

In other words, we do no worse, asymptotically, by estimating the weights even when we actually know them. This result can be seen an extension of wooldridge2007inverse to the case when one has two sets of probability weights being estimated in the first stage.\footnote{In the missing data literature, this result has also been called the “efficiency puzzle”. prokhorov2009gmm study this puzzle in a GMM framework using an augmented set of moment conditions, where the first set of moments correspond to the weighted objective function and the second set belongs to the missing outcomes (or selection in their case) problem.}

A conditional feature of interest is correctly specified

The asymptotic results in the previous section were derived under the assumption that some feature of the conditional distribution of outcomes may be misspecified. This was implicit in defining $\bm{\theta_g^0}$ as a solution to the unconditional M-estimation problem. Examples include estimating a misspecified linear conditional mean or quantile function. In contrast, this section highlights the other half of the asymptotic theory which is formalized using a strong version of the identification assumption and allowing the weights to be misspecified.

assumption{(Strong identification of $\bm{\theta_g^0}$)} The parameter vector $\bm{\theta_g^0}\in \bm{\Theta_g}$ is the unique solution to the population minimization problem \begin{equation} \underset{\bm{\theta_g}\in\bm{\Theta_g}}{\mathrm{min}}\mathbb{E}\left[q(Y(g),\mathbf{X},\bm{\theta_g})|\mathbf{X}\right]; g=0,1 \end{equation} under unconfoundedness (defined in (ref)) and MAR (defined in (ref)) for each $\mathbf{X} \in \mathfrak{X} \subset \mathfrak{R}^{\text{dim}(\mathfrak{X})}$. \qed

The above can be seen as a strengthening of the identification assumption in section (ref) since LIE implies that $\bm{\theta_g^0}$ is also a solution to the unconditional M-estimation problem. By requiring $\bm{\theta_g^0}$ to solve ((ref)), assumption (ref) is intended for situations where a conditional feature of interest is correctly specified. An implication of this strengthened identification is that $\bm{\theta_g^0}$ now solves the conditional score of the objective function i.e. $\mathbb{E}\big[\mathbf{h}(Y(g),\mathbf{X},\bm{\theta_g^0})|\mathbf{X}\big] = \mathbf{0}$.

For instance, the conditional score will be zero in the case of estimating a correctly specified CEF with either OLS or quasi maximum likelihood estimation (QMLE) in the linear exponential family (LEF). This would also hold for a correctly specified CQF estimated either using quantile regression or QMLE in the tick exponential family [Komunjer2005].

Delineating these two identification scenarios is important for determining which causal parameter can be estimated consistently under each setting. As we will see in the next section, it is possible to estimate the ATE under both cases of misspecification. However the same cannot be said for QTE parameters. In addition to assumption (ref), the asymptotic results in this half do not rely on correct specification of weights. In other words, assuming $R(\cdot, \cdot, \bm{\delta})$ and $G(\cdot, \bm{\gamma})$ to be correctly specified is rather restrictive and not required for the doubly weighted estimator to be consistent for $\bm{\theta_g^0}$.

assumption{(Parametric specification of propensity score)} Assume that conditions i) and iii) of assumption (ref) hold where condition ii) is defined for some $\bm{\gamma^\ast} \in \bm{\Gamma}$ such that plim$(\bm{\hat{\gamma}})=\bm{\gamma^\ast}$. \qed
assumption{(Parametric specification of missingness probability)} Assume that conditions i) and iii) of assumption (ref) hold where condition ii) is defined for some $\bm{\delta^\ast} \in \bm{\Delta}$ such that plim$(\bm{\hat{\delta}})=\bm{\delta^\ast}$. \qed

Note that assumptions (ref) and (ref) do not require the parametric models for the two probabilities to be correctly specified. Nevertheless, we continue to assume that $\bm{\hat{\gamma}}$ and $\bm{\hat{\delta}}$ solve the same binary response problem as in Assumptions (ref) and (ref) with probability limits given by pseudo true values $\bm{\gamma^\ast}$ and $\bm{\delta^\ast}$, respectively [white1982maximum]. To show that $\bm{\theta_g^0}$ is still a solution to the doubly weighted population problem with misspecified weights, a sketch of the argument is given below. Consider,

equation[equation omitted — 116 chars of source]

where $\omega_g^\ast$ are asymptotic weights which use $G(\mathbf{X}, \bm{\gamma^\ast})$ and $R(\mathbf{X},W,\bm{\delta^\ast})$. Using LIE along with unconfoundedness and MAR, I can rewrite the above expectation as

equation*[equation* omitted — 120 chars of source]

where $\xi_g(\mathbf{X})$ is a function of weights for $g=0,1$. The strong identification assumption implies

equation*[equation* omitted — 233 chars of source]

Further, since $\xi_g(\mathbf{X}) > 0$,

equation*[equation* omitted — 216 chars of source]

where the inequality is strict when $\bm{\theta_g}\neq \bm{\theta_g^0}$. Therefore, solving the doubly weighted problem identifies the parameter even if the weights are wrong. In general, the parameter that solves ((ref)) will be different from the one that solves the same problem with correct weights.\footnote{When $R(\mathbf{X}, W, \bm{\delta^\ast}) = r(\mathbf{X},W)$ and $G(\mathbf{X}, \bm{\gamma^\ast})=p(\mathbf{X})$, then solving ((ref)) will be the same as solving the problems in section (ref).} But as long as $\bm{\theta_g^0}$ is a unique solution, solving ((ref)) will identify it.

The following two theorems establish consistency and asymptotic normality of the doubly weighted estimator.

theorem{(Consistency under strong identification)} Under assumptions (ref), (ref), (ref), (ref), and (ref) with regularity conditions (1), (2) and (3) of Theorem (ref), $ \bm{\hat{\theta}_g}\overset{p}{\rightarrow} \bm{\theta_g^0} \text{ as } N\rightarrow \infty$ where $\bm{\hat{\theta}_g}$ is the doubly-weighted estimator that solves ((ref)). \qed
theorem{(Asymptotic Normality under strong identification)} Under the assumptions of theorem (ref) and the regularity conditions of theorem (ref) where MLE estimators $\bm{\hat{\gamma}}$ and $\bm{\hat{\delta}}$ have probability limits given by $\bm{\gamma^\ast}$ and $\bm{\delta^\ast}$, then $\sqrt{N}(\bm{\hat{\theta}_g}-\bm{\theta_g^0})\overset{d}{\rightarrow} N\big(\bm{0}, \mathbf{H}_{\bm{g}}^{-1}\bm{\Omega_g}\mathbf{H}_{\bm{g}}^{-1}\big) $ where $\bm{\Omega_g}=\mathbb{E}\big(\mathbf{l}_{ig}\mathbf{l}_{ig}^{\prime}\big)$ with $\mathbf{H}_{\bm{g}}$ and $\mathbf{l}_{ig}$ defined in Theorem (ref) except with asymptotic weights given by $\omega^\ast_{ig}$. \qed

Substantively, there is no real difference in the proof of the above theorem compared to those derived in section (ref) except that now $\bm{\hat{\gamma}}$ and $\bm{\hat{\delta}}$ are converging to probability limits that could be potentially different from those indexing the true treatment and missing outcome probabilities. A consequence of the objective function solving the conditional problem is reflected in the asymptotic variance expression. Compared to the previous section, $\bm{\Omega_g}$ now is simply the variance of weighted score of the objective function without first stage adjustment of the estimated probabilities. This is because under assumption (ref), $\mathbb{E}\big(\mathbf{l}_{ig}\mathbf{b}_i^\prime\big) =\mathbb{E}\big(\mathbf{l}_{ig}\mathbf{d}_i^\prime\big)=\mathbf{0}$. A sketch of the proof for $\mathbb{E}\big(\mathbf{l}_{ig}\mathbf{b}_i^\prime\big) =\mathbf{0}$ is provided below. The argument for $\mathbb{E}\big(\mathbf{l}_{ig}\mathbf{d}_i^\prime\big)$ follows analogously.

equation*[equation* omitted — 295 chars of source]

where $\zeta_g(\mathbf{X})$ is a function of weights. The first equality uses the definition of $\mathbf{l}_{ig}$ with misspecified weights and second equality applies LIEs with unconfoundedness and MAR. In other words, the reason for obtaining a simpler expression for $\bm{\Omega_g}$ is because the correlation between weighted score of the objective function and the two binary response scores is zero when $\bm{\theta_g^0}$ is correctly specified for a conditional feature of interest and we use an appropriate method to estimate it.

A simpler expression for $\bm{\Omega_g}$ also means that we can no longer exploit these correlations between scores to obtain asymptotic efficiency for estimating $\bm{\theta_g^0}$. Again, let $\bm{\tilde{\theta}_g}$ be the doubly weighted estimator that uses true weights, $\omega_g$, then

corollary{(No gain with estimated weights under strong identification)} Under the assumptions of theorem (ref) \begin{equation*} \mathrm{Avar}\big[\sqrt{N}\big(\bm{\tilde{\theta}_g}-\bm{\theta_g^0}\big)\big] = \mathrm{Avar}\big[\sqrt{N}\big(\bm{\hat{\theta}_g}-\bm{\theta_g^0}\big)\big] = \mathbf{H}_{\bm{g}}^{-1}\bm{\Omega_g}\mathbf{H}_{\bm{g}}^{-1} \end{equation*}

\qed

Hence knowledge of the weights does little when for instance we have a correctly specified CEF or CQF and we use either OLS or QR to estimate the parameters indexing these conditional models of interest.

A special case of weights misspecification is when $\omega_g^\ast$ is a constant. This is plausible since $R(\mathbf{X}, W, \bm{\delta^\ast})$ and $G(\mathbf{X}, \bm{\gamma^\ast})$ are allowed to be any bounded positive functions of $\mathbf{X}$ and $W$. In other words, the unweighted estimator, $\bm{\hat{\theta}_{g}^u}$, which does not weight to correct for either problem is also consistent for $\bm{\theta_g^0}$ under the results of theorem (ref). In fact, assumptions (ref) and (ref) suggest that any weighted estimator will suffice for estimating $\bm{\theta_g^0}$. In this case, one may turn to asymptotic efficiency to guide our choice between weighting or not weighting at all. The following result says that if the objective function satisfies the generalized conditional information matrix equality (GCIME), the unweighted estimator is asymptotically more efficient than any of its weighted counterpart (correctly specified weights or not).

corollary{(Efficiency gain with unweighted estimator under GCIME)} Under assumptions of theorem (ref) if we additionally suppose that the objective function satisfies GCIME in the population which is defined as: { \begin{equation} \mathbb{E}\big[\mathbf{h}(Y(g), \mathbf{X}, \bm{\theta_g^0})\mathbf{h}(Y(g), \mathbf{X}, \bm{\theta_g^0})^\prime|\mathbf{X}\big] = \sigma_{0g}^2\cdot \bm{\nabla}_{\bm{\theta_g}}\mathbb{E}\big[\mathbf{h}(Y(g), \mathbf{X}, \bm{\theta_g^0})|\mathbf{X}\big] = \sigma_{0g}^2 \cdot \mathbf{A}(\mathbf{X}, \bm{\theta_g^0}) \end{equation}} Then, { $\mathrm{Avar}\big[\sqrt{N}\big(\bm{\hat{\theta}_g}-\bm{\theta_g^0}\big) \big]= \mathbf{H}_{\bm{g}}^{-1}\bm{\Omega_g}\mathbf{H}_{\bm{g}}^{-1} \ \ \text{ and } \ \ \mathrm{Avar}\big[\sqrt{N}\big(\bm{\hat{\theta}_{g}^u}-\bm{\theta_{g}^0}\big) \big]= (\mathbf{H}_{\bm{g}}^{\mathbf{u}})^{-1}\bm{\Omega}_{\bm{g}}^{\mathbf{u}} (\mathbf{H}_{\bm{g}}^{\mathbf{u}})^{-1} $} and, { \[\mathrm{Avar}\big[\sqrt{N}\big(\bm{\hat{\theta}_g}-\bm{\theta_g^0}\big)\big] -\mathrm{Avar}\big[\sqrt{N}\big(\bm{\hat{\theta}_{g}^u}-\bm{\theta_{g}^0}\big) \big]\]} is positive semi-definite. \qed

The proof of this theorem follows from noting that we can express the difference in the two asymptotic variances as the expected outer product of population residuals from the regression of $\mathbf{B}_i $ on $\mathbf{D}_i$, which are weighted versions of square root of matrix $\mathbf{A}_i$ (See appendix (ref) for details). Hence the difference is positive semi-definite.

We know GCIME is known in a variety of estimation contexts. In the case of full maximum likelihood, GCIME holds for $q(Y(g),\mathbf{X},\bm{\theta_g}) = -\ln{f_g(Y|\mathbf{X},\bm{\theta_g})}$ where $f_g(\cdot|\cdot)$ is the true conditional density with $\sigma_{0g}^2=1$. For estimating conditional mean parameters using QMLE in the linear exponential family (LEF), GCIME holds if $\mathrm{Var}(Y(g)|\mathbf{X})=\sigma_{0g}^2\cdot v[m(\mathbf{X},\bm{\theta_g^0})]$. In other words, GCIME will be satisfied if $\mathrm{Var}(Y(g)|\mathbf{X})$ satisfies the generalized linear model assumption, irrespective of whether the higher order moments of the conditional distribution correspond to the chosen QLL or not. For estimation using nonlinear least squares, GCIME will hold for $q(Y(g),\mathbf{X},\bm{\theta_g}) = [Y(g)-m(\mathbf{X}, \bm{\theta_g})]^2$ with the homoskedasticity assumption. Hence in all these cases the unweighted estimator will be more efficient than its weighted counterpart. But when GCIME is not satisfied, the two may not be easy to rank.

Estimation of treatment effects

The asymptotic theory can now be used to discuss estimation of specific causal estimands like ATE and QTEs which can be expressed as functions of the doubly weighted estimator, $\bm{\theta_g^0}$.

Average treatment effect

As discussed in sloczynski2018general, DR estimators remain consistent for the population ATE despite misspecification in either the conditional mean function or the propensity score, but not both. The current doubly weighted framework along with results developed in sections (ref) and (ref) allow us to extend this result to the case with missing outcomes.

Let $m(\mathbf{X},\bm{\theta_g})$ be a parametric model for the conditional mean which is said to be correctly specified for the CEF if for some $\bm{\theta_g^0}\in \bm{\Theta_g}$

equation*[equation* omitted — 76 chars of source]

or equivalently, $Y(g) = m(\mathbf{X},\bm{\theta_g^0})+U(g)$ such that $\mathbb{E}[U(g)|\mathbf{X}] = 0$. Then, let us consider the following two scenarios in turn.

Double robustness

\paragraph{First half: Correct conditional mean} When the conditional mean function is correct, there is more than one estimation method that can be used to consistently estimate $\bm{\theta_g^0}$, namely, nonlinear least squares (NLS) and QMLE with LEF. For both these examples, results from section (ref) dictate that weighting is not needed for consistency. The fact that one could weight by the misspecified weights and still consistently estimate $\bm{\theta_g^0}$ is what forms the `first part' of the DR result with double weighting.

Once $\bm{\theta_g^0}$ has been estimated by solving the sample version of the NLS or QMLE problem, ATE can be estimated as follows, \[\hat{\Delta}_{\text{ate}} = \frac{1}{N}\sum_{i=1}^{N}m(\mathbf{X}_i, \bm{\hat{\theta}_1})- \frac{1}{N}\sum_{i=1}^{N}m(\mathbf{X}_i, \bm{\hat{\theta}_0})\] If in addition to having a correct conditional mean, I also assume the error variance of the outcomes to be homoskedastic $\big(\mathbb{E}[U^2(g)|\mathbf{X}]=\text{Var}[U(g)|\mathbf{X}] = \sigma^2_{0g}\big)$, then the NLS estimator that does not weight at all is the preferred alternative from an efficiency perspective. This is due to GCIME being satisfied with NLS under homoskedasticity.

\paragraph{Second half: Correct weights} If one acknowledges misspecification in the conditional mean model, there is no general way of consistently estimating the ATE. However, a useful mean fitting property of QMLEs in LEF along with double weighting can be used here to obtain consistent estimates of the unconditional means, $\mathbb{E}[Y(g)]$, despite misspecification in the conditional means, $\mathbb{E}[Y(g)|\mathbf{X}]$.\footnote{The property of QMLEs that we are most familiar with is the one where parameters in a correctly specified conditional mean can be consistently estimated if we choose $m(\mathbf{X},\bm{\theta_g})$ so that it's range corresponds to the chosen LEF density (or QLL function), irrespective of the range and nature of the outcomes. This property is used in the first half of DR.}

In the generalized linear model (GLM) literature, the link function, $h^{-1}(\cdot)$, relates the mean of the distribution to a linear index as follows

equation[equation omitted — 91 chars of source]

The estimation strategy then is to choose $m(\mathbf{X},\bm{\theta_g})$ to be the function, $h(\cdot)$, with the QLL corresponding to a choice of LEF density. Then the population first order conditions from solving this QMLE problem give us

equation[equation omitted — 214 chars of source]

where $v[h(\cdot)]$ is variance of the mean function and $\bm{\theta_g}^{\ast}$ denotes the pseudo true parameter indexing the misspecified conditional mean model [white1982maximum]. In particular, by choosing $h^{-1}(\cdot)$ to be the canonical link for the QLL associated with the density, the gradient in numerator of ((ref)) cancels with the variance term in the denominator. Note that this occurs only when one uses the canonical link function.

Such cancellation of terms ensures that if one includes an intercept in $\mathbf{X}$, the misspecified mean model fits the overall mean of the distribution (see wooldridge2010econometric chapter 13 for more detail) so that,

equation*[equation* omitted — 78 chars of source]

With nonrandom assignment and missing outcomes, solving the sample GLM FOC in ((ref)) would still not be sufficient for consistently estimating $\bm{\theta_g^\ast}$. Therefore, one would instead solve the doubly weighted FOC given below.

equation[equation omitted — 327 chars of source]

The role played by weighting is crucial here for $\bm{\hat{\theta}_g}$ to be consistent for the pseudo true parameter $\bm{\theta_g^\ast}$. This forms the `second half' of the DR result with double weighting.\footnote{Section (ref) in the online appendix provides a detailed proof of how population GLM FOCs identify the unconditional means (and hence the ATE).}

If $h(\cdot)$ is the identity function, the first order conditions above can be recognized as those belonging to OLS with the line of best fit passing through the mean of $Y$. This is because OLS is a QMLE with normal QLL and identity link function, typically used for outcomes with unrestricted support. Other combinations of QLLs and canonical link functions can be found in Table 2 of negiwooldridge2020 and have to be chosen depending on the range and nature of $Y$.

summaryDR estimation of ATE with double weighting \\ Case 1: Correct mean, misspecified weights \begin{enumerate} • Consistent estimates for the conditional mean parameters, $\bm{\theta_g^0}$, can be obtained by either using NLS or QMLE in LEF. • A consistent estimator of ATE is obtained as \begin{equation*} \hat{\Delta}_{ate} = \frac{1}{N}\sum_{i=1}^{N}m(\mathbf{X}_i, \bm{\hat{\theta}_1})-\frac{1}{N}\sum_{i=1}^{N}m(\mathbf{X}_i, \bm{\hat{\theta}_0}) \end{equation*} \end{enumerate} Case 2: Misspecified mean, correct weights \begin{enumerate} • Depending upon the range and nature of the outcome, $Y$, choose an appropriate QLL associated with an LEF density. Choose the mean function, $m(\mathbf{X},\bm{\theta_g}) = h(\mathbf{X}\bm{\theta_g})$, where $h(\cdot)$ is the inverse canonical link function associated with the chosen density. Using this combination of mean function and QLL, use the moment conditions in ((ref)) to obtain consistent estimates, $\bm{\hat{\theta}_g}$. • Consistent estimates of ATE can then be obtained as follows \begin{equation*} \hat{\Delta}_{ate} = \frac{1}{N}\sum_{i=1}^{N}h(\mathbf{X}_i\bm{\hat{\theta}_1})-\frac{1}{N}\sum_{i=1}^{N}h(\mathbf{X}_i\bm{\hat{\theta}_0}) \end{equation*} \end{enumerate}

where $\mathbf{X}$ includes an intercept and $\bm{\hat{\theta}_g}$ solves the GLM first order conditions.

Quantile treatment effects

Unlike the case of ATE, it is generally not possible to obtain UQTE by averaging CQTE over the distribution of $\mathbf{X}$. In this section, I use double weighting to illustrate estimation of three different quantile estimands, namely, UQTE, CQTE, and a weighted linear approximation (LP) to the true CQTE, each of which may be of interest to the researcher depending on whether features of the conditional or unconditional outcomes distribution are of interest. Whether $\bm{\theta_g^0}$ indexes the true CQF or an approximation depends on what is being assumed about the conditional quantile model and the estimation method used.

Let's assume that the two potential outcomes are continuous in $\mathfrak{R}$. It is typical to define the $\tau^{th}$ quantile of $Y(g)$ as

equation*[equation* omitted — 92 chars of source]

Then the UQTE for the $\tau^{th}$ quantile is defined as the difference in the marginal quantiles of the outcomes distributions,

equation*[equation* omitted — 81 chars of source]

Similarly, one may define the $\tau^{th}$ conditional quantile of $Y(g)$ for $\mathbf{X}=\mathbf{x}$ as,

equation*[equation* omitted — 115 chars of source]

where $F_g(\cdot|\mathbf{x})$ denotes the conditional distribution function of $Y(g)$ given $\mathbf{X}=\mathbf{x}$. Then, CQTE for the $\tau^{th}$ quantile for some subgroup defined by $\mathbf{X}$ is

equation*[equation* omitted — 117 chars of source]

Let $\mathfrak{q}_{\tau}(\mathbf{X},\bm{\theta_g}(\tau))$ be a parametric model for the $\tau^{th}$ conditional quantile of $Y(g)$ which is said to be correctly specified if for some $\bm{\theta_g^0}(\tau)\in\bm{\Theta_g}$

equation[equation omitted — 125 chars of source]

\paragraph{Estimation of CQTE$_\tau$:} Incidentally, much like conditional mean, if CQF$_\tau$ is correctly specified, there are two methods that will ensure consistent estimation of the CQF parameters, $\bm{\theta_g^0}(\tau)$. The first is CQR of KoenkerBassett1978 and the second is a class of QML estimators that use a special `tick-exponential' family of distributions to suggest consistent estimators of conditional quantile parameters. This QMLE class has been proposed by Komunjer2005. The method is analogous to estimating a correctly specified conditional mean function using QMLE in the linear exponential family.

For estimation that uses CQR, $\bm{\theta_g}(\tau)$ will actually solve the stronger conditional problem,

equation[equation omitted — 218 chars of source]

For estimation via QMLE, as long as the CQF is correct and we choose an appropriate QLL then,

equation[equation omitted — 233 chars of source]

where $\phi^{\tau}(\cdot, \cdot)$ is the density that belongs to the tick-exponential family.\footnote{$\phi^{\tau}(y,\eta)=\phi^{\tau}(y,\eta) = exp\left[-(1-\tau)[a(\eta)-b(y)]\mathbf{1}\{y\leq \eta\}+\tau[a(\eta)-c(y)]\mathbf{1}\{y>\eta\}\right]$ is a probability density and $\eta$ is the $\tau$-quantile of $\phi^{\tau}$ such that $\int_{-\infty}^{\eta} \phi^{\tau}(y,\eta) dy = \tau$. Komunjer2005 shows that CQR of KoenkerBassett1978 is a special case of this QMLE class.} As dictated by results in section (ref), weighting the QR or QML objective functions, irrespective of whether the weights are correctly specified or not will also deliver a consistent estimator of $\bm{\theta_g}(\tau)$.

Once we have obtained $\bm{\hat{\theta}_g}$ either by solving the QR or QML problem, the $\tau^{th}$ conditional quantile treatment effect for subgroup $\mathbf{X}$ can be estimated as $\widehat{\text{CQTE}}_\tau(\mathbf{X}) = \mathfrak{q}_{\tau}(\mathbf{X},\bm{\hat{\theta}_1}(\tau)) - \mathfrak{q}_{\tau}(\mathbf{X}, \bm{\hat{\theta}_0}(\tau))$.

\paragraph{Estimation of LP to CQTE$_\tau$:} The traditional literature on conditional quantile estimation has focused on correct specification. However, Angristetal2006 establish an approximation property of CQR that is analogous to the approximation property of linear regression. The main implication of such a result is that solving CQR with $\mathfrak{q}_{\tau}(\mathbf{X}, \bm{\theta_g}(\tau)) = \mathbf{X}\bm{\theta_g}(\tau)$ would still identify a weighted linear approximation to CQF$_\tau$. Therefore, the difference in LPs of $\tau$-quantile CQFs is interpretable as identifying an LP to the CQTE$_\tau$.

As before, weighting becomes crucial in the presence of nonrandom assignment and missing outcomes for identifying the LP parameters.

equation[equation omitted — 230 chars of source]

In other words, one would need to weight the CQR problem with correct weighting functions for $\bm{\hat{\theta}_g}(\tau)\overset{p}{\rightarrow}\bm{\theta_g^\ast}(\tau)$, which indexes the true LP to CQF$_\tau$ for group $g$. Then,

equation[equation omitted — 133 chars of source]

\paragraph{Direct estimation of UQTE$_\tau$:} As mentioned in the beginning of this section, estimating UQTE$_\tau$ from CQTE$_\tau$ is generally not possible even if we assume a correct model for the conditional quantiles of $Y(g)$. In other words, one cannot obtain unconditional quantiles from averaging conditional quantiles over the distribution of $\mathbf{X}$. In this case, we can directly estimate $\mathcal{Q}_{\tau,g}$ by running a quantile regression of the outcome on an intercept (similar to firpo2007efficient).\footnote{firpo2007efficient uses propensity score weighting to directly estimate unconditional quantiles in the presence of nonrandom assignment.} In the present case, the solution to the doubly weighted objective function gives us,

equation*[equation* omitted — 187 chars of source]

such that $\hat{\theta}_g(\tau) \overset{p}{\rightarrow}\mathcal{Q}_{\tau,g}$. Weighting by $G(\cdot)$ and $R(\cdot)$ is crucial here since these functions serve to remove biases arising due to nonrandom assignment and missing outcomes. One can then obtain the unconditional quantile treatment effect as, \[\widehat{\text{UQTE}}_{\tau} = \hat{\theta}_1(\tau)-\hat{\theta}_0(\tau)\] An alternative method of estimating UQTE$_\tau$ is to use recentered influence functions suggested by firpo2009unconditional (see appendix (ref)).

The next section discusses results from a Monte Carlo study which evaluates the finite sample behavior of doubly weighted ATE and QTE estimators under three different misspecification scenarios.

Simulations

This section compares the empirical distributions of ATE and QTEs using unweighted, ps-weighted, and d-weighted estimators.\footnote{Details of the simulation design are given in section (ref) of the online appendix.} The discussion is centered around three common misspecification scenarios that are interesting from an empirical standpoint. These cases are enumerated in tables (ref) and (ref) for estimating ATE and QTEs, respectively. Two of them describe situations implicit in the first and second half of the asymptotic theory, whereas the third case considers all three parametric components of the framework to be misspecified. Even though the theory developed in this paper is silent for the third case, simulation results appear to be promising.

Average treatment effect: Results

Case (1) in Table (ref) considers a misspecified mean function but correct probability weights. This is the principal case covered in section (ref) wherein weighting is crucial. As one can see, the empirical distribution of the doubly weighted estimator is centered on the true ATE whereas that for the unweighted estimator is shifted to the right (see figure (ref), Case 1).

Case (2) looks at what happens when everything, conditional mean and the two weights, is misspecified. The theory in this paper does not address this particular case. However, this characterizes an interesting possibility given that misspecification of all components is a valid concern. The simulation results do offer some insight here. The doubly weighted estimator seems to be the only choice that delivers the true ATE on average whereas the others distributions are shifted away from the truth (see figure (ref), Case 2).

Finally, case (3) depicts the possibility of a correctly specified conditional mean function but misspecified weights. Here weighting does not have any bite in resolving the identification issue, beyond what is already achieved from having a correct mean function. In figure (ref), case 3, the empirical distributions of the estimated ATE for the unweighted, ps-weighted, and d-weighted estimators all coincide and are centered on the true ATE.

[Figure (ref) here]

Quantile treatment effects: Results

As discussed earlier, there are really three parameters worth discussing when one talks about QTEs; CQTE, LP to CQTE, and UQTE. Misspecification in the CQF shifts attention to consistently estimating a linear projection to the true CQTE. First case in Table (ref) considers exactly such a scenario. Using the results in Angristetal2006, I interpret the solution to the doubly weighted problem given in ((ref)) as providing a consistent weighted linear projection to the true CQF which is then used to estimate an LP to the true CQTE. Case 1 of Figure (ref) plots the bias in estimated LP relative to the true LP as a function of $X_1$ for the three estimators. Note that weighting here is crucial for consistently estimating the LP. The relative bias of the doubly weighted estimator is the lowest amongst all and coincides with the line of no bias. Case 2 considers the situation when along with a misspecified CQF, the weights are also wrong. We still find the proposed estimator performing the best in terms of bias.

Finally figure (ref) considers a correctly specified CQF in which case we can estimate the CQTE.\footnote{See section (ref) of the online appendix for details regrading plotting the CQTE curve.} One can observe in the figure that the estimated function using double weighting coincides with the true CQTE irrespective of how we weight. All three estimators; unweighted, ps-weighted, and doubly weighted will be consistent for the true CQTE. Misspecification in the weights will not affect this result.

I also consider direct estimation of UQTE which does not require parametric specification of the CQF since it is simply a difference of the marginal quantiles. So the two weights are the only relevant components of the framework which will affect consistency of UQTE. In figure (ref), case 1, when both weights are correct, not weighting and double weighting both bring us close to the true parameter. For the second case where both probability models are misspecified, double weighting does a little worse than not weighting at all. However, the results at other quantiles reflect more favorably upon double weighting (see section (ref) of the online appendix for results at 50th and 75th quantiles). Propensity score weighting performs the worst in both cases suggesting instances where weighting for nonrandom assignment after dropping data that is missing may not be the preferred alternative.

[Figure (ref) here] [Figure (ref) here] [Figure (ref) here]

Returns to job training

In this section, I apply the proposed estimator to the Aid to Families with Dependent Children (AFDC) sample of women from the National Supported Work program compiled by calonico2017women (CS, thereafter). NSW was a transitional and subsidized work experience program which was implemented as a randomized experiment in the United States between 1975-1979. CS replicate lalonde1986evaluating's within-study analysis for the AFDC women in the program, where the purpose of such an analysis is to evaluate how training estimates obtained from using non-experimental identification strategies (for example, CIA) compare to experimental estimates. To compute the non-experimental estimates, CS combine the NSW experimental sample with two non-experimental comparison groups drawn from PSID, called PSID-1 and PSID-2.\footnote{The PSID-1 sample constructed by CS involves keeping all female household heads continuously from 1975-1979 who were between 20 and 55 years of age in 1975 and were not retired in 1975. The sample labeled PSID-2 further restricts PSID-1 to include only those women who received AFDC welfare in 1975.} In this paper, I utilize the within-study feature of this empirical application to estimate how close the doubly weighted estimates get to the experimental estimate compared with ps-weighting and unweighted estimates.

To construct these empirical bias measures, I first augment the CS sample to allow for women who had missing earnings information in 1979. This renders 26% of the experimental and 11% of the PSID samples missing. I then combine the experimental treatment group of NSW with three distinct comparison groups present in the CS dataset, namely, the experimental control group, and the two PSID samples, to compute the unweighted, ps-weighted, and d-weighted training estimates.\footnote{For details regarding sample construction and estimation of weights, see section (ref) of the online appendix.} The difference in the non-experimental estimate, obtained from using the doubly weighted estimator, and the experimental estimate provides the first measure of estimated bias associated with the proposed strategy. Combining the experimental control group with the non-experimental comparison group gives a second measure of estimated bias [hist]. Much like CS, I report both these estimates across a range of regression specifications for the average returns to training estimates.

Given the growing importance of estimating distributional impacts of job training programs, I also estimate returns to training at every 10th quantile of the 1979 earnings distribution. The role of double weighting is highlighted for the case of estimating marginal quantiles since covariates, which primarily serve to remove biases arising from nonrandom assignment and missing outcomes, enter the estimating equation only through the two weights.

Results

First, to evaluate whether women with missing earnings in 1979 were significantly different than those who were observed, Table (ref) reports the mean and standard deviation of the woman's age, years of schooling, pre-training earnings and other characteristics across the observed and missing samples. In terms of age, the women who were observed in the experimentally treated group of NSW and the PSID-1 sample were, on average, older than those who were missing. The observed women in PSID-1 were also more likely to be married. For the PSID-2 sample, women who were observed had, on average, more kids with higher pre-training earnings. Apart from these minor differences, the observed women did not appear to be systematically different that those who were missing, as measured through observable characteristics.

The presence of non-experimental control groups implies that assignment was nonrandom and therefore an issue in the sample. This is because the comparison groups were drawn from PSID after imposing only a partial version of the full NSW eligibility criteria. Table (ref) provides descriptive statistics for the covariates by the treatment status. As can be expected, the treatment and control groups of NSW are not observably different, indicating the strong role that randomization plays in producing comparable groups. In contrast, the women in PSID-1 and PSID-2 groups are statistically different than the treatment group members implying substantial scope for nonrandom assignment.

Table (ref) reports the d-weighted, ps-weighted and unweighted average returns to training estimates which using three different comparison groups; NSW control, PSID-1 and PSID-2. The unweighted (unadjusted and adjusted) experimental estimates given in row 1, are same as the estimates reported by CS in Table 3 of their paper. Overall, one can see that the doubly weighted experimental estimates are more stable than the single weighted or unweighted estimates across the different regression specifications, with a range between \$824-\$828.

For computing the ps-weighted and d-weighted non-experimental estimates, I first trim the sample to ensure common support between the treatment and comparison groups.\footnote{Appendix (ref) describes estimation of the two probability weights along with the sample trimming criteria.} This reduces the sample size from 1,248 to 1,016 observations for the PSID-1 estimates and from 782 to 720 observations for the PSID-2 estimates. A pattern that is consistent across the two sets of non-experimental estimates is that weighting gets us much closer to the benchmark relative to not weighting at all. For instance, the unweighted simple difference in means estimate of training, which uses the PSID-1 comparison group, is -\$799 whereas the weighted estimates are \$827 and \$803. For the PSID-2 comparison group, the unweighted estimate which controls for all covariates is \$335 whereas the weighted estimates are \$905 and \$904.

The second panel of Table (ref) reports the bias in training estimates from combining the experimental control group with the PSID comparison groups. A similar pattern is seen here with weighted bias estimates being much closer to zero than the unweighted estimates. For instance, the doubly weighted estimate that adjusts for all covariates using the PSID-1 comparison group is -\$21 whereas the unweighted estimates is -\$568. These results suggest that the argument for weighting is strong when using a non-experimental comparison group where nonrandom assignment and missing outcomes are significant problems.\footnote{Note that the large standard errors for the non-experimental estimates can be attributed to the small sample sizes and to the large residual variance of earnings in the PSID-1 and PSID-2 populations.}

Figure (ref) plots the relative bias in UQTE estimates at every 10th quantile of the 1979 earnings distribution. Much like the average training estimates, we see that the weighted estimates consistently lie below the unweighted estimates for most quantiles, irrespective of whether we use the PSID-1 or PSID-2 non-experimental group. Note that I do not plot UQTE estimates for quantiles less than 0.46, since these are all zero.\footnote{There are a lot of women in the experimental and PSID samples with zero real earnings in 1979.}

This empirical application illustrates the role of proposed estimator in both experimental and observational data contexts. The comparison involving the treatment and control group of NSW demonstrates its use in an experiment with missing outcomes, whereas the non-experimental sample demonstrates its use in the more realistic observational data setting.

[Table (ref) here] [Table (ref) here] [Table (ref) here] [Figure (ref) here]

Conclusion

In empirical research, the problems of nonrandom assignment and missing outcomes threaten identification of causal parameters. This paper proposes a new class of consistent and asymptotically-normal M-estimators that address these two issues using a double weighting procedure. The method combines propensity score weighting with weighting for missing outcomes in a general M-estimation framework, which can be applied to a range of estimation methods, such as ordinary least squares, quasi maximum likelihood, and quantile regression. In addition, the proposed class has a robustness property which allows us to estimate meaningful causal quantities of interest despite misspecification in either a conditional model of interest or the two weighting functions.

As leading applications, the paper discusses estimation of ATE and QTEs. A Monte Carlo study indicates that the doubly weighted estimates of average and quantile treatment effects have the lowest bias compared to naive alternatives (unweighted or propensity score weighted estimators) under three realistic cases of misspecification. Finally, the estimator is applied to the data on AFDC women from the NSW program compiled by Cal\'{o}nico and Smith (2017). The presence of experimental and non-experimental comparison groups in this application help to quantify the estimated bias in the doubly weighted returns to training estimates as well as the other two estimators.

Since the severity and magnitude of bias introduced from ignoring either problem cannot be assessed ex-ante, a safe bet from the practitioner's perspective is to report both doubly weighted and unweighted causal effect estimates. Practically, the doubly weighted estimator for the ATE is easy to implement. Appendix (ref) provides an example code that uses Stata's gmm command for implementing it. Computation of analytically correct standard errors, however, requires additional coding and is still a work in progress. Alternatively, one can use bootstrapped standard errors which will provide asymptotically correct inference.

Even though missing outcomes are a common concern in empirical analysis, it is equally common to encounter missing data on the covariates. A particularly important future extension can be to allow for missing data on both. In this case, using a generalized method of moments framework which incorporates information on complete and incomplete cases could provide efficiency gains over just using the observed data. A different possibility would be to relax the identifying restrictions to allow for selection on unobservables and possibly explore estimation of local average treatment effect (LATE).

References

spacing{1}