EconBase
← Back to paper

Generalized method of moments with partially missing data

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

33,451 characters

Generalized Method of Moments with Partially Missing Data



\maketitle

\begin{abstract}
\linespread{1.2}

We consider a generalized method of moments framework in which a part of the data vector is missing for some units in a completely unrestricted, potentially endogenous way.
In this setup, the parameters of interest are usually only partially identified.
We characterize the identified set for such parameters using the support function of the convex set of moment predictions consistent with the data.
This identified set is sharp, valid for both continuous and discrete data, and straightforward to estimate.
We also propose a statistic for testing hypotheses and constructing confidence regions for the true parameter, show that standard nonparametric bootstrap may not be valid, and suggest a fix using the bootstrap for directionally differentiable functionals of \citet{fang2019inference}.
A set of Monte Carlo simulations demonstrates that both our estimator and the confidence region perform well when samples are moderately large and the data have bounded supports.

\medskip

\noindent \textbf{Keywords:} endogenous attrition, partial identification, panel data, support function

\medskip

\noindent \textbf{JEL codes:} C14, C23
\end{abstract}

\newpage



\section{Introduction}

Missing data are ubiquitous in empirical research, particularly in repeated-measurement settings. \citet{rubin1976inference} classified mechanisms of missingness into three categories: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). In this paper, we consider the last case, where missingness may be correlated with the data in an arbitrary manner.
More specifically, we consider a generalized method of moments framework in which a portion of the data vector is missing for some units in a completely unrestricted, potentially endogenous way.
A representative example is a two-period panel where, in the second period, some units drop out of the sample in a way that is correlated with both outcomes and covariates in both periods.

When no restriction on missingness is imposed, the parameters of interest are usually only partially identified.
We characterize the identified set for the parameter of interest via the support function of the set of moments consistent with the observed data. We show that the identified set is sharp, is valid for both continuous and discrete data, and is straightforward to estimate.

We also show how to perform hypothesis testing and construct confidence regions using the estimate of the minimum of the support function as a test statistic. We derive the limit distribution of this test statistic and provide a procedure for computing critical values, following \citet{fang2019inference}. We also show that the test controls size locally uniformly using an abstract result in \citet{fang2019inference}.

Partial identification with missing data has been widely studied since Manski's seminal work, \citet{Manski1989anatomy} (see, e.g., \citet{Manski2005partial} and \citet{Molinari2020microeconometrics} for a survey). Our paper contributes to this extensive literature, with its key feature being that the identified set of moment predictions is convex, and the support function of this set can be estimated using a simple sample analog.

Our model is neither a moment inequality model with a convex identified set (e.g., \citet{KaidoSantos2014}), nor the intersection bounds model of \citet{ChernozhukovLeeRosen2013intersection}. It is similar to the model with convex moment predictions as in \citet{BeresteanuMolchanovMolinari2011}, the main difference being that our characterization of the sharp identified set does not rely on a representation via random closed sets.

As mentioned above, the primary application of our model is a panel regression with attrition that may be endogenous. The analysis of panel data with attrition has a long history, and numerous techniques for handling attrition under various assumptions have been proposed. These include modelling the attrition process parametrically (e.g., \citet{HausmanWise1979}), using auxiliary information in the form of refreshment samples (e.g., \citet{hirano2001combining}), imposing semiparametric assumptions on the attrition process (e.g., \citet{bhattacharya2008inference}), using inverse probability weighting (e.g., \citet{Wooldridge2002}), selection models with unobserved heterogeneity (e.g., \citet{SemykinaWooldridge2010}), and multiple imputation and Bayesian methods (e.g., \citet{DengEtAlSurvey}).

The rest of the paper is organized as follows.
Section \ref{sec:framework} introduces the general framework and characterizes the sharp identified set.
Section \ref{sec:estimation} suggests an estimator of the identified set.
Section \ref{sec:inference} develops a bootstrap procedure for hypothesis testing.
Section \ref{sec:simulation} explores the performance of our methodology in a set of Monte Carlo simulations.
Section \ref{sec:conclusion} concludes.
All proofs are given in the Appendix.


\section{Setup and the identified set}\label{sec:framework}

We consider the generalized method of moments framework when some of the observed variables may be missing for some observations.
Specifically, let $Z_{1i}$ be a data vector that we observe for each unit $i=1,\dots,n$ of the sample and let $Z_{2i}$ be a data vector that may be missing for some units.
Let $S_i\in \{0,1\}$ be the sample selection indicator, which is equal to 1 when $Z_{2i}$ is observed and 0 otherwise.
Dropping the subscript $i$, the observed vector is then $W=(S,Z_1',S Z_2')'$.
A key aspect of our framework is that we allow arbitrary dependence of $S$ on both observed and unobserved features.
In particular, $S$ can be arbitrarily correlated with $Z_2$, representing endogenous sample selection.
We do not allow for overlap between elements of $Z_1$ and $Z_{2}$.
More generally, no features of the distribution of $Z_2$ are assumed to be identified from the random sample of $Z_1$.
We are interested in a finite-dimensional parameter $\theta\in\Theta\subset \mathbb{R}^{d_\theta}$ that is defined by general moment conditions
\begin{align}
    \operatorname{\mathbb{E}}_{\pi}[\phi(Z_1,Z_2,\theta)]=0, \label{eq:GMM}
\end{align}
where $\phi$ is a $d_\phi$-dimensional moment function and $\pi$ is the joint distribution of $Z_1,Z_2$.
Of course, $\pi$ is not point-identified due to the presence of missing data (see below), and hence neither is $\theta$.

One special case of this framework is two-period panel data with unrestricted attrition in the second period.
To see this, let $Z_t=(X_t',Y_t) \in \mathbb{R}^d$ be the stacked vector of covariates and outcomes in period $t=1,2.$
In period 1, a random sample from $Z_1$ is observed.
In period 2, there is attrition, and so $Z_2$ is only observed when $S=1$, where $S\in\{0,1\}$ is the indicator of staying in the sample.
When the attrition is fully unrestricted and the parameter of interest is defined by a set of moment conditions \eqref{eq:GMM} (such as, for instance, the slope coefficient in the fixed effects linear regression), this model becomes a special case of the GMM with missing data described above.
For expositional clarity, we use the terminology related to this panel model throughout the rest of the paper.

We impose the following assumptions.

\begin{assumption}\label{a:moment-and-support}
    \begin{subassumption}
        \item \label{a:Theta-compact} The parameter space $\Theta$ is compact.
        \item \label{a:mu} The distribution $\pi$ of $(Z_1,Z_2)$ belongs to the set $\mathcal{P}(\mathcal{Z}_1\times \mathcal{Z}_2)$ of probability distributions on $\mathcal{Z}_1\times \mathcal{Z}_2$, where $\mathcal{Z}_1, \mathcal{Z}_2 \subset \mathbb{R}^d$ are known compact sets.
        \item \label{a:phi-continuous} For each $\theta\in\Theta$, the map $(z_1,z_2) \mapsto \phi(z_1,z_2,\theta)$ is continuous.
        \item \label{a:E-phi-continuous} For each $\pi \in \mathcal{P}(\mathcal{Z}_1\times \mathcal{Z}_2)$, the map $\theta \mapsto \operatorname{\mathbb{E}}_{\pi} [\phi(Z_1,Z_2,\theta)]$ is continuous.

    \end{subassumption}
\end{assumption}

\Cref{a:Theta-compact,a:E-phi-continuous,a:phi-continuous} are standard and impose compactness of the parameter space and the continuity of moments.
\Cref{a:mu} posits that the researcher knows the compact set $\mathcal{Z}_1 \times \mathcal{Z}_2$ that contains (but does not have to be equal to) the true support of the data.
We impose no restrictions on the data-generating process beyond these basic requirements.

To characterize the identified set for $\theta$, let $p=P(S=1) \in (0,1)$ be the unconditional selection probability and notice that $\pi$ can be decomposed into the point-identified distribution of $(Z_1,Z_2)|S=1$, denoted by $\pi^1$, and the partially identified distribution of $(Z_1,Z_2)|S=0$, denoted by $\pi^0$, viz.,
\begin{align*}
    \pi = p \pi^1 + (1-p)\pi^0.
\end{align*}
Denote by $\pi_1^0$ the point-identified probability distribution of $Z_1|S=0$.
The sharp identified set for $\pi^0$ is the convex set of the distributions satisfying \Cref{a:mu} with first marginal $\pi_1^0$, i.e.,
\begin{align*}
    \Pi^0 = \left\{\pi^0 \in \mathcal{P}(\mathcal{Z}_1\times \mathcal{Z}_2): \,\, \pi^0(A \times \mathcal Z_2) = \pi_1^0(A) \text{ for all Borel sets }A \subset \mathcal{Z}_1 \right\}.
\end{align*}
Therefore, the sharp identified set for $\pi$ is
\begin{align*}
    \Pi = \left\{p \pi^1 + (1-p) \pi^0, \,\, \pi^0 \in \Pi^0 \right\}.
\end{align*}
Write
\begin{align}
     \operatorname{\mathbb{E}}_{\pi}[\phi(Z_1,Z_2,\theta)] = p \operatorname{\mathbb{E}}_{\pi^1}[\phi(Z_1,Z_2,\theta)] + (1-p) \operatorname{\mathbb{E}}_{\pi^0} [\phi(Z_1,Z_2,\theta)] =: \nu_{\pi^0}(\theta).
\end{align}
Then the sharp identified set for $\theta$ is
\begin{align*}
    \Theta_I = \left\{\theta\in\Theta: \,\, \exists \pi^0 \in \Pi^0 \text{ such that } \nu_{\pi^0}(\theta)=0 \right\}.
\end{align*}
The next proposition establishes a convenient characterization of $\Theta_I$ in terms of the support function of the compact, convex set\footnote{See Step 1 in \Cref{app:id-set} for the proof of convexity and compactness of $N(\theta)$.} of moment predictions
\begin{align*}
    N(\theta) := \left\{\nu_{\pi^0}(\theta):\,\,\pi^0\in\Pi^0 \right\}.
\end{align*}

\begin{proposition}[identified set via support function]\label{prop:id-set}
    Suppose \Cref{a:moment-and-support} holds and let
    \begin{align*}
        \psi_{N(\theta)}(u) = \max_{\nu\in N(\theta)} u'\nu
    \end{align*}
    be the support function of $N(\theta)$. Define the criterion function
    \begin{align*}
        Q(\theta) = \min_{u\in \mathbb{B}} \psi_{N(\theta)}(u),
    \end{align*}
    where $\mathbb{B}$ is the unit ball in $\mathbb{R}^{d_\phi}$.
    Then $\Theta_I$ is compact and given by
    \begin{align}
        \Theta_I = \left\{ \theta\in\Theta: \,\, Q(\theta) = 0 \right\}.
    \end{align}
\end{proposition}

\begin{proof}
    See \Cref{app:id-set}.
\end{proof}

\begin{remark}
    In the above characterization, the Euclidean unit ball ~$\mathbb{B}$ can be replaced by any compact set containing the origin in its interior, such as the unit ball in the $\ell^1$ norm or the uniform norm. This may turn out to be convenient computationally.
\end{remark}

\begin{remark}
    A similar strategy was recently employed by \citet{franguridi2025inference} for characterizing the identified set in a partially identified GMM framework based on optimal transport.
\end{remark}

The support function $\psi_{N(\theta)}$ can be written as
\begin{align*}
    \psi_{N(\theta)}(u) &= \max_{\pi\in \Pi} u'\operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)]\\
    & = p \operatorname{\mathbb{E}}[u'\phi(Z_1,Z_2,\theta)|S=1] + (1-p) \max_{\pi^0\in\Pi^0} \operatorname{\mathbb{E}}_{\pi^0}[u'\phi(Z_1,Z_2,\theta]].
\end{align*}
The last term involves optimization over an infinite-dimensional set of distributions. Fortunately, we can convert it into a finite-dimensional program for every value of $Z_1$, as the following proposition shows.


\begin{proposition}[reduction formula] \label{prop:reduction} Suppose \Cref{a:moment-and-support} holds. Then
    \begin{align*}
        \max_{\pi^0\in\Pi^0} \operatorname{\mathbb{E}}_{\pi^0}[u'\phi(Z_1,Z_2,\theta)] = \operatorname{\mathbb{E}}\left[\max_{z_2\in\mathcal{Z}_2} u'\phi(Z_1,z_2,\theta) |S=0 \right].
    \end{align*}
\end{proposition}

\begin{proof}
    See \Cref{app:reduction}.
\end{proof}

\begin{figure}
    \centering
    \includegraphics[width=0.8\textwidth]{figures/univariate_figure.pdf}
    \caption{Set identification via support functions}
    \label{fig:univariate}
\end{figure}

\begin{example}\label{example:univariate}
Suppose there is only one moment condition, i.e., $\phi$ is scalar-valued.
In this case, the set of moment predictions is the closed interval
\begin{align*}
    N(\theta) = \left[ \min_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)],\,\, \max_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)] \right],
\end{align*}
The support function of $N(\theta)$ is a piecewise linear function given by
\begin{align*}
    \psi_{N(\theta)}(u) = \max_{\pi\in\Pi} u \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)] =
    \begin{cases}
        u \cdot \min_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)], &\text{if } u < 0,\\
        u \cdot \max_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)], &\text{if } u \ge 0,
    \end{cases}
\end{align*}
and hence the criterion function is
\begin{align*}
    Q(\theta) = \min_{u\in \mathbb{B}} \psi_{N(\theta)}(u) = \min\left\{0,\,\, \max_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)], \,\,- \min_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)] \right\}.
\end{align*}
The identification condition $Q(\theta) = 0$ can be easily seen to be equivalent to
\begin{align*}
    \min_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)] \le 0 \le \max_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)],
\end{align*}
which, of course, is just the condition that zero belongs to the set of moment predictions.
See \Cref{fig:univariate} for an illustration of how the argmax set of $\psi_{N(\theta)}$ determines the position of $\theta$ relative to the identified set $\Theta_I$.
Notice also that using the reduction formula in \Cref{prop:reduction}, we can write
\begin{align*}
    \min_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)] = \operatorname{\mathbb{E}}[f_{\min}(W,\theta)], \\
    \max_{\pi\in\Pi} \operatorname{\mathbb{E}}_\pi[\phi(Z_1,Z_2,\theta)] = \operatorname{\mathbb{E}}[f_{\max}(W,\theta)],
\end{align*}
where $W=(S,Z_1',S Z_2')'$ is the data vector and
\begin{align*}
    f_{\min}(W,\theta) = S \phi(Z_1,S Z_2,\theta) + (1-S) \min_{z_2\in\mathcal{Z}_2 } \phi(Z_1,z_2,\theta), \\
    f_{\max}(W,\theta) = S \phi(Z_1,S Z_2,\theta) + (1-S) \max_{z_2\in\mathcal{Z}_2} \phi(Z_1,z_2,\theta).
\end{align*}


\end{example}
\section{Estimation}\label{sec:estimation}

Combining \Cref{prop:id-set,prop:reduction} suggests the following estimator of the identified set:
\begin{align*}
    \hat\Theta_I = \left\{ \theta\in\Theta: \,\, \hat Q(\theta) \ge - \eta_n \right\},
\end{align*}
where $\eta_n>0$ is a tuning parameter and the sample criterion function
\begin{align*}
    \hat Q(\theta) = \min_{u\in\mathbb{B}} \hat \psi_{N(\theta)}(u) = \frac{1}{n} \sum_{i=1}^n \left( s_i u'\phi(z_{i1}, z_{i2},\theta) + (1-s_i) \max_{z_2\in\mathcal{Z}_2} u'\phi(z_{i1}, z_{2},\theta) \right).
\end{align*}
Define the distance from a point $x \in \mathbb{R}^{d_\phi}$ to a set $A \subset \mathbb{R}^{d_\phi}$ by
\begin{align*}
    d(x,A) = \inf_{a\in A} \|a-x\|.
\end{align*}
For any sets $A,B \subset \mathbb{R}^{d_{\theta}}$, define the Hausdorff distance
\begin{align*}
    d_H(A,B) = \max \left\{ \sup_{a\in A} d(a,B), \,\, \sup_{b\in B} d(b,A)  \right\}.
\end{align*}
We impose the following assumptions.

\begin{assumption}\label{as:separation}
    There exists a weakly increasing function $m:\mathbb{R}_+ \to \mathbb{R}_+$ such that $m(0)=0$, $m(\delta)>0$ for $\delta>0$ and $\min_{u\in \mathbb{B}} \psi_{N(\theta)}(u) \le - m(d(\theta,\Theta_I))$.
\end{assumption}

\begin{assumption}\label{as:eta}
    $\eta_n \downarrow 0$ and $n^{-1/2} = o(\eta_n)$.
\end{assumption}

\Cref{as:separation} imposes separation of the identified set $\Theta_I$ in terms of the criterion function $Q(\theta) = \min_{u\in \mathbb{B}} \psi_{N(\theta)}(u)$.
A sufficient condition can be formulated in terms of the gradient of $Q(\theta)$ away from the identified set.
\Cref{as:eta} states that $\eta_n$ has to converge to zero slower than $O(n^{-1/2})$.

The following proposition establishes the consistency of our estimator in the Hausdorff distance.

\begin{theorem}[Hausdorff consistency]\label{prop:consistency}
    Under \Cref{a:moment-and-support,as:separation,as:eta}, $d_H(\Theta_I,\hat\Theta_I)=o_p(1)$.
\end{theorem}

\begin{proof}
    See \Cref{app:consistency}.
\end{proof}
\section{Inference}\label{sec:inference}

We now wish to develop a test of the hypothesis $H_0:\,\theta=\theta_0$.
We suggest using the statistic
\begin{align*}
    \hat T(\theta_0) = \sqrt{n} \min_{u\in\mathbb{B}} \hat \psi_{N(\theta_0)}(u),
\end{align*}
which is the estimator of the scaled negative distance to the identified set of moment predictions $N(\theta_0)$.
Note that more negative values of $\hat T(\theta_0)$ indicate stronger evidence against the null.
To establish the asymptotic distribution of this statistic, we first prove that the function $\hat \psi_{N(\theta)}(u)$ converges to a Gaussian process uniformly over $u\in\mathbb{B}$ and $\theta\in\Theta$.

\begin{assumption}\label{a:phi-finite-variance}
    $\operatorname{\mathbb{E}}\left[ \sup_{\theta\in\Theta}\sup_{z_2\in\mathcal{Z}_2} \|\phi(Z_1,z_2,\theta)\|^2 \right]<\infty$.
\end{assumption}

\begin{assumption}\label{a:phi-Lipschitz}
    There exists a random variable $L(Z_1)$ such that $\operatorname{\mathbb{E}} L(Z_1)^2 < \infty$ and
    \begin{align*}
        \sup_{z_2\in\mathcal{Z}_2} \|\phi(Z_1,z_2,\theta_1)-\phi(Z_1,z_2,\theta_2)\| \le L(Z_1)\|\theta_1-\theta_2\|
    \end{align*}
    for all $\theta_1,\theta_2\in\Theta$, almost surely in $Z_1$.
\end{assumption}

\begin{theorem}[uniform CLT for the support function]\label{prop:uclt}
    Suppose \Cref{a:moment-and-support,a:phi-finite-variance,a:phi-Lipschitz} hold. Then there exists a Gaussian process $\mathbb{G}$ with trajectories in $C(\mathbb{B}\times\Theta)$ such that
    \begin{align*}
        \sqrt n (\hat \psi_{N(\theta)}(u) - \psi_{N(\theta)}(u)) \rightsquigarrow \mathbb{G}(u,\theta).
    \end{align*}
\end{theorem}

\begin{proof}
    See \Cref{app:uclt}.
\end{proof}

\begin{remark}
    The results of \citet{KaidoSantos2014} suggest that the estimator $\hat\psi_{N(\theta)}(u)$ is asymptotically semiparametrically efficient for $\psi_{N(\theta)}(u)$.
\end{remark}

An application of the functional delta method to \Cref{prop:uclt} yields the following result for the asymptotic distribution of the test statistic $\hat T(\theta)$ over $\theta\in\Theta_I$.

\begin{corollary}[asymptotic distribution of test statistic]\label{prop:lim-distribution}
    Under the assumptions of \Cref{prop:uclt},
    \begin{align*}
        \hat T(\theta) \rightsquigarrow \min_{u\in U(\theta)} \mathbb{G}(u,\theta)
    \end{align*}
    uniformly over $\theta\in\Theta_I$, where
    \begin{align*}
        U(\theta) = \operatorname*{argmin}_{u\in \mathbb{B}} \psi_{N(\theta)}(u).
    \end{align*}
\end{corollary}

\begin{proof}
    See \Cref{app:lim-distribution}.
\end{proof}

\begin{remark}
    Using this uniform convergence result, it should be possible to develop a specification test based on the statistic $\max_{\theta\in\Theta} \hat T(\theta)$, the population analog of which is zero if and only if the identified set $\Theta_I$ is nonempty.
    The critical values can be obtained using the bootstrap of \citet{fang2019inference}, in analogy to our treatment of the test statistic $\hat T(\theta_0)$ for a fixed $\theta_0$ below.
    Similar specification tests have been developed for moment inequality models, see, e.g., \citet{bugni2015specification} and references therein.
\end{remark}

Since the asymptotic distribution of $\hat T(\theta)$ is neither available in closed form nor is easy to simulate from, we need an alternative procedure for obtaining the critical values.
The results of \citet{fang2019inference} imply that, unless the Hadamard derivative $\min_{u \in U(\theta)} \mathbb{G}(u,\theta)$ is a linear functional of $\mathbb{G}$, the standard nonparametric bootstrap is invalid.

\setcounter{example}{0}
\begin{example}[continued]
    In the univariate case, the sample criterion function is
\begin{align*}
    \hat Q(\theta) = \min_{u\in \mathbb{B}} \hat \psi_{N(\theta)}(u) = \min\left\{0,\,\, \hat\operatorname{\mathbb{E}}[f_{\max}(W,\theta)], \,\,- \hat \operatorname{\mathbb{E}}[f_{\min}(W,\theta)] \right\}.
\end{align*}
Consider a data generating process and a parameter value $\theta$ satisfying $\operatorname{\mathbb{E}}[f_{\min}(W,\theta)] < 0$, making the sample moment $\hat \operatorname{\mathbb{E}}[f_{\min}(W,\theta)]$ irrelevant asymptotically.
We have
\begin{align*}
    Q(\theta) &= \min\left\{0,\,\, \operatorname{\mathbb{E}}[f_{\max}(W,\theta)] \right\}, \\
    \hat Q(\theta) &\overset{asy}{\sim} \min\left\{0,\,\, \hat\operatorname{\mathbb{E}}[f_{\max}(W,\theta)] \right\}.
\end{align*}
This is known to be an irregular problem when $\operatorname{\mathbb{E}}[f_{\max}(W,\theta)]=0$, which corresponds to $\theta$ being on the boundary of the identified set.
In particular, the standard nonparametric bootstrap is invalid for the asymptotic distribution $\min(0,N(0,1))$ of $\sqrt{n} \hat Q(\theta)$, see, e.g., \citet{andrews2000inconsistency}.
\end{example}

We show how to leverage the bootstrap for directionally differentiable functionals of \citet{fang2019inference}.
To this end, we propose an estimator of the derivative that satisfies their key Assumption 4.

For $\varepsilon_n>0$, define
\begin{align}
    \hat\chi'(h) = \min_{u\in \hat U(\varepsilon_n)} h(u), \label{eq:hdd-estimator}
\end{align}
where
\begin{align*}
    \hat U(\varepsilon_n) = \left\{u\in \mathbb{B}:\,\,\hat\psi_{N(\theta_0)}(u) \le \min_{v\in\mathbb{B}} \hat\psi_{N(\theta_0)}(v)+\varepsilon_n \right\}
\end{align*}
is the ``$\varepsilon_n$-enlarged'' argmin set of $\hat\psi_{N(\theta_0)}$.
We impose the following assumptions.

\begin{assumption}[sharp minima of $\psi_{N(\theta_0)}$]\label{as:sharp-minima}
    There exists $\kappa>0$ such that
    \begin{align*}
        \psi_{N(\theta_0)}(u) \ge \min_{v\in \mathbb{B}}\psi_{N(\theta_0)}(v) + \kappa \cdot d(u,U(\theta_0)) \text{ for all } u \in \mathbb{B}.
    \end{align*}
\end{assumption}

\begin{assumption}[bandwidth]\label{as:bandwidth} The sequence $\varepsilon_n$ satisfies $\varepsilon_n \downarrow 0$ and $\|\hat\psi_{N(\theta_0)}-\psi_{N(\theta_0)}\|_\mathbb{B} = o_p(\varepsilon_n)$.
\end{assumption}

\Cref{as:sharp-minima} posits that $\psi_{N(\theta_0)}$ grows sufficiently fast away from its argmin set $U(\theta_0)$.
The latter does not have to be a singleton, as illustrated in \Cref{example:univariate} for which \Cref{as:sharp-minima} holds.
This assumption is equivalent to every subgradient $\nabla\psi_{N(\theta_0)}(u)$ being bounded away from zero for $u\notin U(\theta_0)$.
It also suffices that it holds in a small neighborhood around $U(\theta_0)$ rather than on the entire set $\mathbb{B}$.
\Cref{as:bandwidth} restricts how fast $\varepsilon_n$ should converge to zero.
Since $\|\hat\psi_{N(\theta_0)}-\psi_{N(\theta_0)}\|_\mathbb{B}=O_p(1/\sqrt{n})$ by \Cref{prop:uclt}, we can take $\varepsilon_n \sim \log n/\sqrt{n}$.

Given our estimator \eqref{eq:hdd-estimator}, the \citet{fang2019inference} bootstrap algorithm for testing $H_0:\theta=\theta_0$ at the nominal size $\alpha\in (0,1)$ is as follows.
\begin{enumerate}
    \item Compute $\hat T(\theta_0)=\sqrt{n} \min_{u\in\mathbb{B}} \hat \psi_{N(\theta_0)}(u)$.
    \item Define $\hat U_n = \left\{u\in \mathbb{B}:\,\, \hat \psi_{N(\theta_0)}(u) \le \min_{v\in\mathbb{B}} \hat \psi_{N(\theta_0)}(v) + \varepsilon_n \right\}$.
    \item For each $b=1,\dots, B$:
    \begin{enumerate}
      \item Draw a random sample $W_i^{b}$, $i=1,\dots, n$, with replacement from the data $W_i = (S_i,Z_{1i},S_i Z_{2i})$, $i=1,\dots, n$.
      \item Compute $\hat \psi_{N(\theta_0)}^{*b}(u)$ on the sample $W_i^b$, $i=1,\dots,n$.
      \item Compute
      \begin{align*}
        \hat T^{*b} = \min_{u\in\hat U_n}\sqrt{n}\left( \hat \psi_{N(\theta_0)}^{*b}(u) - \hat \psi_{N(\theta_0)}(u) \right).
      \end{align*}
    \end{enumerate}
  \item Denote by $\hat c_{\alpha}^*$ the $\alpha$-quantile of $\hat T^{*1},\dots, \hat T^{*B}$.
  \item Reject $H_0$ if $\hat T(\theta_0) < \hat c_{\alpha}^*$.
  \end{enumerate}

The following theorem establishes that this bootstrap-based testing procedure controls size for a fixed data generating process.

\begin{theorem}[test validity, pointwise]\label{thm:bootstrap}
    Suppose \Cref{as:sharp-minima,as:bandwidth} hold and consider the testing procedure above with $B=\infty$. Then for every fixed data generating process, the asymptotic probability of rejecting the true null hypothesis is less than or equal to $\alpha$.
\end{theorem}

\begin{proof}
    See \Cref{app:bootstrap}.
\end{proof}

Although the previous theorem establishes that our test asymptotically controls size for a fixed data generating process, it is silent about uniform size control.
In other words, it could be possible that even for a large sample size, one can find a particularly unfavorable DGP that would lead our test to exhibit significant size distortion.
Such lack of uniform size control has been extensively studied in the context of moment inequality models, e.g., \citet{imbens2004confidence,romano2008inference,andrews2009validity,moon2009estimation,woutersen2006simple}.

Fortunately, we can rely on the abstract results of \citet{fang2019inference} to establish local size control.
Denote by $P$ the joint distribution of the data vector $W=(S,Z_1',S Z_2')'$ and let $\psi(P)$ be the corresponding population support function of $N(\theta_0)$, where $\theta_0$ is any value in the identified set $\Theta_I = \Theta_I(P)$.
We consider a sequence of DGPs $P_t$, $t \ge 0$, that is local to a given DGP $P_0$ in the following sense, cf. Assumption 5 in \citet{fang2019inference}.

\begin{assumption}[local sequence of DGPs]\label{as:local-analysis}
    The path $t \mapsto P_t$ is quadratic mean differentiable and satisfies for any $\lambda >0$ the following.
    \begin{subassumption}
    \item \label{suba:psi-diff} There exists $\psi'(\lambda)$ such that
    \[
        \|\sqrt{n}(\psi(P_{\lambda/\sqrt{n}})-\psi(P_0)) - \psi'(\lambda)\|_\mathbb{B} \to 0.
    \]
    \item \label{suba:psi-regular} $\sqrt{n} (\hat\psi_n - \psi(P_{\lambda/\sqrt{n}})) \overset{\lambda}{\to} \mathbb{G}_0$, where $\overset{\lambda}{\to}$ denotes convergence in distribution under $\{W_i\}_{i=1}^n$ i.i.d., with each $W_i$ distributed as $P_{\lambda/\sqrt{n}}$, and $\hat\psi_n$ is our estimator of $\psi(P_{\lambda/\sqrt{n}})$ based on $\{W_i\}_{i=1}^n$.
    \item \label{suba:local-distribution} $\mathbb{G}_0$ is tight and supported on $C(\mathbb{B})$.
    \end{subassumption}
\end{assumption}

\Cref{suba:psi-diff} imposes differentiability of the parameter $\psi(P)$ along the path $P_{\lambda/\sqrt{n}}$.
It is expected to hold for any quadratic mean differentiable path in our model under mild regularity conditions.
\Cref{suba:psi-regular} requires that the distribution of the estimator $\hat\psi_n$ is unaffected by local perturbations of the DGP, i.e. that $\hat\psi_n$ is \emph{regular}.
Finally, \Cref{suba:local-distribution} is expected to hold due to \Cref{prop:uclt}.


\begin{theorem}[test validity, locally uniform]\label{thm:uniform-test-validity}
    Suppose \Cref{as:sharp-minima,as:bandwidth,as:local-analysis} hold and consider the testing procedure above with $B=\infty$. Then this procedure controls size locally uniformly in the sense that
    \begin{align*}
        \limsup_{n \to \infty} P_{\lambda/\sqrt{n}}\left(\hat T_n(\theta_0) < \hat c_{\alpha}^* \right) \le \alpha.
    \end{align*}
\end{theorem}

\begin{proof}
    See \Cref{app:uniform-test-validity}.
\end{proof}

\section{Extension to more general missingness patterns}\label{sec:multi-period}

Our methodology can be extended to more general missingness patterns, as long as missingness is ``monotone'' (see below) and data contains no information on the distribution of the missing variables (i.e., missingness is ``fully unrestricted'').
We will illustrate this extension using the setup of panel data with more than two periods, where attrition happens monotonically, i.e., once the units drop out of the sample, they never return.
For expositional simplicity, we consider the case of $T=3$ periods.

Denote the data in the three periods by $Z_1,Z_2,Z_3$ and denote the indicators of staying in the sample in periods 2 and 3 by $S_2$ and $S_3$, respectively.
Since there is no return to the sample, $S_2=0$ implies $S_3=0$.
Also, let
\[
p_2=P(S_2=1), \quad p_3=P(S_3=1), \quad p_{2|0} = p(S_2=1|S_3=0).
\]
The parameter of interest $\theta$ is defined by the moment conditions
\[
\operatorname{\mathbb{E}}_\pi [\phi(Z_1,Z_2,Z_3,\theta)]=0,
\]
where the expectation is taken with respect to the latent, partially identified distribution $\pi$ of $(Z_1,Z_2,Z_3)$.

\begin{proposition}\label{prop:multi-period}
The identified set is $\Theta_I = \{\theta\in\Theta: \,\, \min_{u\in\mathbb{B}} \psi_\theta(u)=0 \}$, where
\begin{align*}
    \psi_\theta(u) &= p_3\operatorname{\mathbb{E}}[\phi(Z_1,Z_2,Z_3,\theta)|S_3=1] + (1-p_3)\left[p_{2|0} \operatorname{\mathbb{E}} [\max_{z_3} u'\phi(Z_1,Z_2,z_3,\theta) |S_2=1,S_3=0] \right. \\
    &+ \left. (1-p_{2|0}) \operatorname{\mathbb{E}} [\max_{z_2,z_3} u'\phi(Z_1,z_2,z_3,\theta) |S_2=S_3=0]
    \right].
\end{align*}
\end{proposition}

\begin{proof}
    See \Cref{app:multi-period}.
\end{proof}





\section{Monte Carlo simulation}\label{sec:simulation}

We consider a simple cross-sectional linear regression
\begin{align*}
    Y_i = \theta_0 + \theta_1 X_i + \varepsilon_i, \quad i=1,\dots,n,
\end{align*}
where $Y_i$ is observed for the entire sample, whereas $X_i$ is missing for observations $i$ with $S_i=0$, where $S_i \in \{0,1\}$ is the sample selection indicator.
The parameter of interest is $\theta=(\theta_0,\theta_1)$.
We assume that the support of $X_i$ is known, in agreement with \Cref{a:mu}.
Dropping the subscript $i$, the moment conditions are
\begin{align*}
    \operatorname{\mathbb{E}}(Y-\theta_0-\theta_1 X)&=0, \\
    \operatorname{\mathbb{E}}[X(Y-\theta_0-\theta_1 X)]&=0,
\end{align*}
and hence the support function is
\begin{align*}
    \psi_\theta(u_1,u_2) &= p \operatorname{\mathbb{E}}[u_1 (Y-\theta_0-\theta_1 X) + u_2 X(Y-\theta_0-\theta_1 X) \,|\,S=1  ] \\
    &+ (1-p)\operatorname{\mathbb{E}}\left[\max_{x\in[0,1]}\left\{ u_1 (Y-\theta_0-\theta_1 x) + u_2 x(Y-\theta_0-\theta_1 x) \right\}\right].
\end{align*}
Notice that the inner maximization is a constrained quadratic problem.
We assume that $X_i$ has a uniform[0,1] distribution and $\varepsilon_i$ has a uniform[-1,1] distribution independently of $X_i$. Selection is completely random with probability $p=P(S_i=1)=0.9$.
We set the sample size $n=1000$ and the tuning parameters $\eta_n=\varepsilon_n = 0.1\log n/\sqrt{n} \approx 0.022$.

\begin{figure}
    \centering
    \includegraphics[width=1\textwidth]{figures/results_idset_and_rejection.pdf}
    \caption{Estimates of the identified set and confidence regions with nominal sizes $1-\alpha \in \{0.10,0.90,0.95\}$ averaged over $1000$ simulations.}
    \label{fig:mc}
\end{figure}

The simulation results are illustrated in \Cref{fig:mc}.
The true identified set (red contour line) is approximated using a sample of $10,000$ observations.
Our estimates of the identified set track the true identified set closely.
The confidence regions (three black contour lines) collect the hypothesized parameter values that are not rejected by the bootstrap test described in Section \ref{sec:inference} with the size $\alpha \in \{0.05,0.10,0.90\}$ and the number of bootstrap samples $B=1000$.
They are found to be quite narrow even at high nominal confidence levels $1-\alpha$.
\section{Conclusion}\label{sec:conclusion}

In this paper, we study the partial identification of parameters defined by GMM moment conditions when components of the data vector are missing in a completely unrestricted manner. A leading motivation is panel data subject to endogenous attrition.

We characterize the sharp identified set through the support function of the convex set of moment predictions consistent with the observed data.
For inference, we employ the minimum of the sample analog of the support function as a test statistic and develop a valid inference procedure, drawing on the bootstrap for directionally differentiable functionals of \citet{fang2019inference}.
We demonstrate that our estimator and confidence regions perform well in a set of Monte Carlo simulations.




\bibliographystyle{ecta}
\singlespacing
\bibliography{ref}

\onehalfspacing
\frenchspacing
\newpage