The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
84,986 characters
Asymptotic Variance Theory for Trimmed Least Squares and Trimmed Least Absolute Deviations in Censored Panel Models with Fixed Effects
\title{Asymptotic Variance Theory for Trimmed Least Squares and Trimmed Least
Absolute Deviations in Censored Panel Models with Fixed Effects\footnote{We
thank Matias Cattaneo, Jin Hahn, Michael Jansson, and Whitney Newey for helpful
comments and suggestions.}}
\author[1]{Denis Chetverikov}
\author[2,4]{Jesper R.-V.~S{\o}rensen}
\author[3,4]{Bo Honor{\'e}}
\affil[1]{\small{University of California, Los Angeles}}
\affil[2]{\small{University of Copenhagen}}
\affil[3]{\small{Princeton University}}
\affil[4]{\small{Aarhus Center for Econometrics (ACE)}}
\date{\today}
\maketitle
\begin{abstract}
We study inference using trimmed least squares (TLS) and trimmed least
absolute deviations (TLAD) estimators of \citet{honore_trimmed_1992} in censored
two-period panel-data models with fixed effects. We show that the published
asymptotic variance formulas rely on additional regularity conditions that are
not fully stated in the original analysis. For TLS, the published Hessian formula requires that the regressor-difference index vanish only when the regressor difference itself is zero, a restriction not explicitly stated in the original paper and violated, for instance, with a zero parameter vector.
We derive the correct Hessian, establish asymptotic normality without imposing this restriction, and obtain a consistent plug-in variance estimator. We also show that the Hessian estimator proposed in \citet{honore_trimmed_1992} {\em is} actually consistent for the {\em correct} TLS asymptotic
variance. For TLAD, we show that the published variance formula omits a
conditional-probability term and that asymptotic normality requires additional
continuity conditions. Under these conditions, we derive the corrected
asymptotic variance and provide a tuning-parameter-free bootstrap variance
estimator.
\end{abstract}
\section{Introduction}
Inference for censored panel-data models with fixed effects often relies on
trimmed estimators whose objective functions are convex but nonsmooth.
\citet{honore_trimmed_1992} introduced two influential examples, the trimmed
least squares (TLS) and trimmed least absolute deviations (TLAD) estimators, and
derived their asymptotic distributions for a two-period censored regression
model with fixed effects. These estimators remain natural benchmark procedures
in this setting, both because they avoid parametric restrictions on the fixed
effects and because they provide tractable ways to exploit the structure of
two-period panel data under censoring. Valid inference for such estimators,
however, depends delicately on the form of the Hessian of the expected loss, and
the nonsmoothness created by trimming and censoring makes that Hessian less
straightforward than the published formulas suggest.
This paper revisits the asymptotic variance theory for the TLS and TLAD
estimators of \citet{honore_trimmed_1992}. We show that the published asymptotic
variance formulas rely on additional regularity conditions that are not fully
reflected in the original theorem statements. For TLS, the problem is that
censoring creates boundary points at which the score of the trimmed square loss
is nondifferentiable with positive probability. At such points, differentiation
under expectation need not be valid, so the Hessian of the expected loss can
differ from the expression obtained by differentiating inside the expectation.
For TLAD, the problem is that the expected loss can fail to be twice
differentiable when the conditional distribution of the model errors is not
sufficiently smooth. Our results identify how these issues enter the asymptotic
variance analysis for TLS and TLAD, derive the corrected variance formulas, and
clarify the scope of valid legacy inference.
Our starting point is the two-period censored regression model in which the
outcome variable $Y_\tau$ in period $\tau\in\{1,2\}$ is generated as
\begin{equation}\label{eq:Outcomes}
Y_\tau=\max\{0,Y_\tau^\ast\},\qquad
Y_\tau^\ast=\alpha+\boldsymbol{X}_\tau^\top\boldsymbol{\theta}_0+\varepsilon_\tau,
\end{equation}
where $\boldsymbol{X}_\tau\in\mathbb{R}^K$ is a vector of time-varying regressors, $\boldsymbol{\theta}_0\in
\mathbb{R}^K$ is the parameter of interest, $\alpha$ is an unobserved fixed effect, and
$\varepsilon_\tau$ is an unobserved error term. Let
$\{(Y_{i1},\boldsymbol{X}_{i1},Y_{i2},\boldsymbol{X}_{i2})\}_{i=1}^n$ be a random sample from the
distribution of $(Y_1,\boldsymbol{X}_1,Y_2,\boldsymbol{X}_2)$. \citet{honore_trimmed_1992} proposed
M-estimators of $\boldsymbol{\theta}_0$ based on a trimmed loss function
$m_\Xi:\mathbb{R}\times[0,\infty)\times[0,\infty)\to\mathbb{R}$ defined by
\begin{equation}\label{eq:TrimmedGenericLoss}
m_{\Xi}(t,\boldsymbol{y}):=\begin{cases}
\Xi(y_{1})-\left(y_{2}+t\right)\xi(y_{1}), &t\leqslant-y_2,\\
\Xi(y_{1}-y_{2}-t), &t\in\left(-y_{2},y_{1}\right),\\
\Xi(-y_{2})-\left(t-y_{1}\right)\xi(-y_{2}), &t\geqslant y_{1},
\end{cases}
\end{equation}
where $\Xi(\cdot)$ represents either the absolute loss $\left|\cdot\right|$ or
the (one-half) square loss
$\textstyle\frac{1}{2}(\cdot)^2$,\footnote{\citet[Assumption
L1]{honore_pairwise_1994} provide a list of conditions allowing other choices of
$\Xi$.} and $\xi(\cdot)$ is its derivative.\footnote{In the case of the absolute
loss $\Xi(\cdot)=\left|\cdot\right|$, we set $\xi(0) = 0$.} These choices yield
the TLAD and TLS estimators, respectively, defined for the corresponding
$\Xi(\cdot)$ as any solution
\begin{align}\label{eq:TrimmedEstimator}
\widehat{\boldsymbol{\theta}}_{\Xi} &\in \operatornamewithlimits{argmin}\limits_{\boldsymbol{\theta}\in\mathbb{R}^K}\cbr[3]{\frac{1}{n}\sum_{i=1}^n m_{\Xi}\big(\Delta\boldsymbol{X}_i^\top\boldsymbol{\theta},\boldsymbol{Y}_i\big)},
\end{align}
where we introduced the shorthands $\Delta\boldsymbol{X}_i := \boldsymbol{X}_{i1} - \boldsymbol{X}_{i2}$ and
$\boldsymbol{Y}_i:=(Y_{i1},Y_{i2})$. Under conditional exchangeability of
$(\varepsilon_1,\varepsilon_2)$ given $(\boldsymbol{X}_1,\boldsymbol{X}_2,\alpha)$ and other
regularity conditions, \citet{honore_trimmed_1992} established consistency and
asymptotic normality of both estimators.
Our first set of results concerns the TLS estimator. We show that the Hessian
formula appearing in \citet{honore_trimmed_1992} is valid only if
\begin{equation}\label{eq: Bo condition}
\mathbb{P}\del[1]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0 = 0\text{ and }\Delta\boldsymbol{X}\neq\boldsymbol{0}} = 0,
\end{equation}
a condition not explicitly stated in the original paper. This condition is
satisfied, for example, when the index $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ has a
continuous distribution, but it fails when all components of
$\boldsymbol{\theta}_0$ are equal to zero unless $\Delta\boldsymbol{X} = \boldsymbol{0}$ almost surely. When condition \eqref{eq: Bo condition} is not satisfied, the TLS score is nondifferentiable with positive probability, and the published Hessian formula need not equal the derivative of the expected score. We derive
the correct Hessian, provide an explicit representation of it, establish
asymptotic normality without excluding mass at zero, and obtain a consistent
plug-in estimator of the corresponding asymptotic variance.
The TLS analysis also yields a second result. Although the population Hessian
formula in \citet{honore_trimmed_1992} is not generally correct, the Hessian
estimator proposed there remains consistent for the correct TLS asymptotic
variance under the maintained assumptions. Thus, the paper does not only
identify where the published population formula fails; it also clarifies which
parts of the original inference procedure remain usable.
Our second set of results concerns the TLAD estimator. We show, first, that the
published asymptotic variance formula omits a conditional-probability term. More
importantly, we show that asymptotic normality requires additional continuity
conditions to ensure existence of the Hessian of the expected loss. We provide a
counterexample satisfying the conditions stated in \citet{honore_trimmed_1992}
for which that Hessian does not exist. Under additional continuity conditions,
we derive the corrected TLAD Hessian and corresponding asymptotic variance
formula. We then provide a tuning-parameter-free bootstrap variance estimator
for TLAD.
Taken together, these results clarify the role of boundary nondifferentiability
in the asymptotic variance theory for censored trimmed estimators. More
specifically, they show when legacy inference remains valid, when corrected
variance formulas are required, and how to implement inference under the
corrected theory.
As a by-product, our TLS analysis also yields a plug-in perspective on the
asymptotic variance of the cross-sectional trimmed least squares estimator
studied by \citet{honore_pairwise_1994}, thus avoiding the choice of tuning
parameters required by the numerical-differentiation-based variance estimator
proposed there.
The remainder of the paper is organized as follows. Section~\ref{sec:TLS}
revisits the TLS estimator, beginning with the asymptotic normality result in
\citet{honore_trimmed_1992}, then identifying the source of the problem,
deriving the correct Hessian, establishing asymptotic normality under the
corrected theory, and discussing variance estimation, including the status of
the legacy Hessian estimator. Section~\ref{sec:TLAD} carries out the analogous
analysis for TLAD, derives the corrected Hessian and asymptotic variance
formula, and develops bootstrap-based variance estimation. Proofs for the TLS results are in the appendix, while proofs for the TLAD results are in the supplemental appendix, which also contains technical lemmas establishing a generalized dominated convergence theorem, the existence of conditional PDFs,
and related measurability issues.
\paragraph{Notation.}
For a positive integer $k$, we write $\left[k\right]:=\{1,\ldots,k\}$. We use
$\left\|\cdot\right\|_2$ to denote the Euclidean norm. For any (suitably
differentiable) real-valued function $f$, whose first argument is a scalar, we
let $\dot{f}_1$ and $\ddot{f}_{11}$ denote the first- and second-order partial
derivatives of $f$ with respect to its first argument, respectively. For a
possibly vector-valued differentiable function $\boldsymbol{g}$, we let $\nabla \boldsymbol{g}$
denote the Jacobian of $\boldsymbol{g}$. If $\boldsymbol{g}$ is itself the gradient mapping
corresponding to some function $f$, then we write $\nabla^2 f$ for the Hessian
of $f$. The indicator $\mathbf{1}\{A\}$ equals one (zero) if the logical
statement $A$ is true (false). We use comma-separated events to denote the
intersection of events, e.g.~$\mathbf 1\{A,B\} = \mathbf 1\{A\cap B\}$. Unless
otherwise stated, all limits are understood as the sample size $n$ grows
without bound, holding everything else fixed. The symbols $\rightsquigarrow$,
$\to_{\text{a.s.}}$, and $\to_{\mathbb{P}}$ denote convergences in distribution,
almost surely, and in probability, respectively. We write $o_{\mathbb{P}}(1)$,
$o_{\mathrm{a.s.}}(1)$, and $o_{L^1}(1)$ to denote sequences of random
variables that converge to zero in probability, almost surely, and in
$L^1(\mathbb{P})$, respectively.
\section{Trimmed Least Squares}\label{sec:TLS} In this section, we focus on the
TLS estimator. We abbreviate the trimmed square loss, $m_{\Xi}$ in
\eqref{eq:TrimmedGenericLoss} with $\Xi=\textstyle{\frac{1}{2}}(\cdot)^2$, by
$m^{\texttt{tls}}$, which takes the explicit form
\begin{equation}\label{eq:TrimmedSqLoss}
m^{\texttt{tls}}(t,\boldsymbol{y})=\frac{1}{2}\cdot\begin{cases}
y_{1}^2-2\left(y_{2}+t\right)y_{1}, &t\leqslant-y_2,\\
\left(y_{1}-y_{2}-t\right)^2, &t\in\left(-y_{2},y_{1}\right),\\
y_{2}^2+2\left(t-y_{1}\right)y_{2}, &t\geqslant y_{1}.
\end{cases}
\end{equation}
This loss is continuously differentiable in its first argument with partial derivative given by
\begin{equation}\label{eq:TrimmedSqLossDerivative}
\dot{m}^{\texttt{tls}}_1(t,\boldsymbol{y})=\begin{cases}
-y_{1}, &t\leqslant-y_{2},\\
y_{2}-y_{1}+t, &t\in(-y_{2},y_{1}),\\
y_{2}, &t\geqslant y_{1}.
\end{cases}
\end{equation}
To facilitate our presentation below, note that it follows from
\eqref{eq:TrimmedSqLossDerivative} that $\dot{m}^{\texttt{tls}}_1(\cdot,\boldsymbol{y})$ is
Lipschitz continuous on $\mathbb{R}$ with Lipschitz constant one, uniformly in
$\boldsymbol{y}\in[0,\infty)\times[0,\infty)$. However, it fails to be differentiable at
$-y_2$ and $y_1$ (unless both $y_1$ and $y_2$ are zero), implying that the
trimmed square loss is not necessarily twice differentiable. Specifically, the
points of second-order \emph{non}-differentiability of
$m^{\texttt{tls}}(\cdot,\boldsymbol{y})$ are captured by set-valued function
$N:[0,\infty)\times[0,\infty)\rightrightarrows\mathbb{R}$, defined by
\begin{equation}\label{eq:TrimmedSqLossNonTwiceDiff}
N(\boldsymbol{y}):=\left\{t\in\mathbb{R}\middle| \ddot{m}_{11}^{\texttt{tls}}(t,\boldsymbol{y})\text{ does not exist}\right\}=
\begin{cases}
\emptyset, &y_{1}+y_{2}=0,\\
\{-y_{2},y_{1}\}, &y_{1}+y_{2}>0.
\end{cases}
\end{equation}
\subsection{Assumptions for Trimmed Least Squares}\label{sec:Assumptions-TLS}
For notational convenience, abbreviate the TLS estimator $\widehat\boldsymbol{\theta}_\Xi$
in \eqref{eq:TrimmedEstimator} with
$\Xi=\textstyle{\frac{1}{2}}\left(\cdot\right)^2$ by
$\widehat{\boldsymbol{\theta}}^{\texttt{tls}}$. Also, denote $\boldsymbol{W}:=(\boldsymbol{X}_1,\boldsymbol{X}_2,\alpha)$ and
let $\mathcal{W}\subseteq\mathbb{R}^{K}\times \mathbb R^K\times\mathbb R$ be the support of
$\boldsymbol{W}$, with $\boldsymbol{w} := (\boldsymbol{x}_1,\boldsymbol{x}_2,a)$ standing for a generic element of $\mathcal{W}$. In
addition, denote $\boldsymbol{\varepsilon}:=(\varepsilon_1,\varepsilon_2)$ and let $\mathcal E
:= \mathbb R\times\mathbb R$, with $\boldsymbol{e} = (e_1,e_2)$ standing for a generic
element of $\mathcal E$. \citet{honore_trimmed_1992} derived the asymptotic
distribution of the TLS estimator under the following
assumptions.\footnote{Assumptions
\ref{assu:Non-Degeneracy-TLS}--\ref{assu:RankRegressors-TLS}
are from Assumptions S.2, M.3, E.1, E.3 and R.1
, respectively, in \citet{honore_trimmed_1992}.}
\begin{assumption}[\textbf{Non-Degeneracy}]\label{assu:Non-Degeneracy-TLS}
The probability $\mathbb{P}(Y_1>0, Y_2>0)$ is strictly positive.
\end{assumption}
\begin{assumption}[\textbf{Integrability}]\label{assu:MomentConditions-TLS}
All of the following expectations are finite:
\[
\mathrm{E}[\|\boldsymbol{X}_1\|_2^4],\;\mathrm{E}[\|\boldsymbol{X}_2\|_2^4],\;\mathrm{E}[\alpha^2\|\Delta\boldsymbol{X}\|_2^4],\;
\mathrm{E}[\varepsilon_1^2\|\Delta\boldsymbol{X}\|_2^4],\;\text{and}\;
\mathrm{E}[\varepsilon_2^2\|\Delta\boldsymbol{X}\|_2^4].
\]
\end{assumption}
\begin{assumption}[\textbf{Continuity}]\label{assu:continuity tls}
The conditional distribution of $(\varepsilon_1,\varepsilon_2)$ given
$\boldsymbol{W}$ is absolutely continuous with respect to the Lebesgue measure.
\end{assumption}
\begin{assumption}[\textbf{Exchangeability}]\label{assu:Exchangeability-TLS}
Conditional on $\boldsymbol{W}$,
$\varepsilon_1$ and $\varepsilon_2$ are exchangeable.
\end{assumption}
\begin{assumption}[\textbf{Rank of Regressors}]\label{assu:RankRegressors-TLS}
There is no proper linear subspace of $\mathbb{R}^K$ containing the random variable
$\mathbf{1}\{\mathbb{P}\left(Y_1>0, Y_2>0\middle|\boldsymbol{X}_1,\boldsymbol{X}_2\right)>0\}\Delta\boldsymbol{X}$ with probability one.
\end{assumption}
Let $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$, mapping $\mathcal
W\times\mathcal E$ to $[0,\infty)$, be a version of the joint probability
density function (PDF) of the pair $\boldsymbol{\varepsilon} = (\varepsilon_1,\varepsilon_2)$
conditional on $\boldsymbol{W} = \boldsymbol{w}$ that is measurable in $(\boldsymbol{w},\boldsymbol{e})$ and is such that
$f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_1,e_2) = f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_2,e_1)$ for all
$\boldsymbol{w}\in\mathcal W$ and $\boldsymbol{e} = (e_1,e_2)\in\mathcal E$. As we
show in Appendix \ref{sec: measurability appendix}, the function
$(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$ does exist under Assumptions
\ref{assu:continuity tls} and \ref{assu:Exchangeability-TLS}. Also, let
$(\boldsymbol{w},\boldsymbol{e})\mapsto F_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e}) =
\int_{-\infty}^{e_1}\int_{-\infty}^{e_2}f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(u_1,u_2)\dif u_2\dif u_1$
be the corresponding version of the joint cumulative distribution function (CDF)
of $(\varepsilon_1,\varepsilon_2)$ conditional on $\boldsymbol{W} = \boldsymbol{w}$, and let
$(\boldsymbol{w},e)\mapsto F_{\varepsilon\mid\boldsymbol{w}}(e) = \lim_{u\to\infty}
F_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e,u)$ be the corresponding version of the common marginal
CDF of $\varepsilon_1$ and $\varepsilon_2$ conditional on $\boldsymbol{W} = \boldsymbol{w}$.
\subsection{Asymptotic Normality Result in Honor{\'e} (1992)} To state the TLS normality
result in \citet{honore_trimmed_1992}, introduce the $K\times K$ matrices
\[
\mathbf{V}_0^{\tt{tls}}:=\mathrm{E}\left[\dot{m}_1^{\tt{tls}}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0,\boldsymbol{Y})^2\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\right]
\]
and
\begin{equation}\label{eq:VarianceSandwichBread}
\boldsymbol{\Gamma}_0^{\tt{tls}}:=\mathrm{E}\left[\mathbf{1}\cbr[0]{-Y_2<\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<Y_1}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\right].\footnote{In
\citet{honore_trimmed_1992}, these matrices are denoted $V_4$ and $\Gamma_4$,
respectively.}
\end{equation}
\citet[Theorem 2(iv)]{honore_trimmed_1992} states that if Assumptions
\ref{assu:Non-Degeneracy-TLS}--\ref{assu:RankRegressors-TLS} hold, the
expectations involved in defining the matrices $\mathbf{V}_0^{\tt{tls}}$ and
$\boldsymbol{\Gamma}_0^{\tt{tls}}$ exist (in $\mathbb{R}^{K\times K}$), and both matrices are of full rank, then
\begin{equation}\label{eq:HonoreAsymptoticNormality-TLS}
\sqrt{n}\del[1]{\widehat\boldsymbol{\theta}^{\texttt{tls}}-\boldsymbol{\theta}_0}\rightsquigarrow\mathcal{N}\left(\mathbf{0},(\boldsymbol{\Gamma}_0^{\tt{tls}})^{-1}\mathbf{V}_0^{\tt{tls}}(\boldsymbol{\Gamma}_0^{\tt{tls}})^{-1}\right)\;\text{in}\;\mathbb{R}^K,
\end{equation}
with $\mathcal{N}(\boldsymbol{\mu},\boldsymbol{\Sigma})$ denoting the normal
distribution with mean $\boldsymbol{\mu}$ and covariance matrix
$\boldsymbol{\Sigma}$.
Among several steps, the proof of this result in
\citet{honore_trimmed_1992} involves:
\begin{enumerate}
\item Arguing that the gradient of the expected loss
$\boldsymbol{\theta}\mapsto\mathrm{E}[m^{\tt{tls}}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta},\boldsymbol{Y})]$ is equal to the
expected value of the model score [see \eqref{eq:ExpectedScore-TLS} below].
\item Establishing differentiability of the expected value of the model score
at the true parameter $\boldsymbol{\theta}_0$.
\end{enumerate}
In our notation, the second task amounts to arguing that the function
$\boldsymbol{G}:\mathbb{R}^K\to\mathbb{R}^K$ defined by
\begin{equation}\label{eq:ExpectedScore-TLS}
\boldsymbol{G}(\boldsymbol{\theta}):=\mathrm{E}\left[\frac{\partial}{\partial\boldsymbol{\theta}}m^{\tt{tls}}\del[0]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta},\boldsymbol{Y}}\right]
=\mathrm{E}\left[\dot{m}_1^\texttt{tls}\del[0]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta},\boldsymbol{Y}}\Delta\boldsymbol{X}\right],\quad\boldsymbol{\theta}\in\mathbb{R}^K,
\end{equation}
is differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$. Its Jacobian $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)\in\mathbb{R}^{K\times K}$ is then the Hessian of the expected loss.\footnote{In \citet{honore_trimmed_1992}, the function $\boldsymbol{G}$
is denoted $G_4$.} To carry out this task, \citet{honore_trimmed_1992} invoked the LDCT to interchange the
order of differentiation and expectation, and the matrix $\boldsymbol{\Gamma}_0^{\tt{tls}}$
is the result of that calculation. In the next subsection, however, we will show by example that
$\boldsymbol{\Gamma}_0^{\tt{tls}}$ is not always equal to the Jacobian of $\boldsymbol{G}$ at
$\boldsymbol{\theta}_0$.
\subsection{Counterexample}\label{sec:Counterexample} We here provide a DGP
which satisfies Assumptions
\ref{assu:Non-Degeneracy-TLS}--\ref{assu:RankRegressors-TLS}, for which the
function $\boldsymbol{G}$ \emph{is} differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$, but the
resulting Jacobian $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$ differs from $\boldsymbol{\Gamma}_0^{\tt{tls}}$.
To this end, let $K=1$, $\alpha\equiv0$, $\boldsymbol{\theta}_0=0$, $\boldsymbol{X}_1\equiv2$ and
$\boldsymbol{X}_2\equiv1$, and let $\varepsilon_1$ and $\varepsilon_2$ be independent of
each other and standard normally distributed. Then the outcome variables
$Y_{\tau}=\max\{0,\varepsilon_{\tau}\},\tau\in\{1,2\}$, are independent and
identically distributed as standard normals censored from below at zero. Case by
case inspection shows that Assumptions
\ref{assu:Non-Degeneracy-TLS}--\ref{assu:RankRegressors-TLS} are satisfied and
that (the here scalar) $\mathbf{V}_0^{\tt{tls}}$ and $\boldsymbol{\Gamma}_0^{\tt{tls}}$ are both
well-defined and of full rank (non-zero).
In this example DGP, the regressor difference $\Delta \boldsymbol{X} = \boldsymbol{X}_1 - \boldsymbol{X}_2$ is
identically one, so the (here scalar-valued) function $\boldsymbol{G}$ simplifies to
$G(\theta)=\mathrm{E}[\dot{m}^{\texttt{tls}}_1(\theta,\boldsymbol{Y})],\theta\in\mathbb{R}$. It is then straightforward to check that the function $G$ is given by
\[
G(\theta) =
\begin{cases}
\theta+\int_{0}^{-\theta}\Phi(u)\dif u, &\theta<0,\\
0, &\theta=0,\\
\theta-\int_{0}^{\theta}\Phi(u)\dif u, &\theta>0,
\end{cases}
\]
where $\Phi$ is the standard normal CDF. For example, if $\theta < 0$, then
letting $\varphi$ denote the standard normal PDF,
\begin{align*}
\mathrm{E}[\dot{m}^{\texttt{tls}}_1(\theta,\boldsymbol{Y})]
& = \mathrm{E}[-Y_1 + (Y_2 + \theta)\mathbf{1}\{Y_2 + \theta > 0\}]
= \theta - \mathrm{E}[(Y_1 + \theta)\mathbf{1}\{Y_1 + \theta \leqslant 0\}] \\
& = \theta - \int_{0}^{-\theta}u\varphi(u)\dif u - \theta\Phi(-\theta)
= \theta + \int_0^{-\theta}\Phi(u)\dif u,
\end{align*}
where the first equality follows from \eqref{eq:TrimmedSqLossDerivative}, the
second from (the here) identical distributions of $Y_1$ and $Y_2$, the third
from noting that $Y_1$ is a standard normal random variable truncated from
below at zero, and the fourth from integration by parts.
The graph of $G$ is shown in Figure
\ref{fig:ExpectedLoss}.
\begin{figure}
\caption{Graph of $\theta\mapsto G(\theta)$ in the
counterexample\label{fig:ExpectedLoss}}
\centering
\includegraphics{expected_loss_derivative_plot.pdf}
\end{figure}
As indicated by the figure, and immediately follows analytically, $G$
\emph{is} differentiable at $\theta_0(=0)$, with derivative $\dot{G}(0)=1 -
\Phi(0)=\textstyle{\frac{1}{2}}$. However, since $Y_1$ and $Y_2$ are
independently distributed as standard normals censored from below at zero, (the
here scalar) $\boldsymbol{\Gamma}_0^{\tt{tls}}$ equals
\[
\Gamma_0^{\tt{tls}}=\mathrm{E}[\mathbf{1}\{Y_1>0>-Y_2\}]=\mathbb{P}(Y_1>0)\cdot\mathbb{P}(Y_2>0)=\frac{1}{4},
\]
showing that $\dot{G}(0)\neq\Gamma_0^{\tt{tls}}$.
How does this discrepancy arise? When arguing differentiability of the function
$\boldsymbol{G}(\boldsymbol{\theta}) =
\mathrm{E}\sbr[1]{\dot{m}_1^\texttt{tls}\del[0]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta},\boldsymbol{Y}}\Delta\boldsymbol{X}}$ at
$\boldsymbol{\theta} = \boldsymbol{\theta}_0$, \citet{honore_trimmed_1992} called upon the LDCT to
interchange the order of differentiation and expectation. However, as captured
by $N(\boldsymbol{y})$ in \eqref{eq:TrimmedSqLossNonTwiceDiff}, the function
$\dot{m}^{\texttt{tls}}_1(\cdot,\boldsymbol{y})$ fails to be differentiable at $-y_2$ and
$y_1$, unless both $y_1$ and $y_2$ are zero. In turn, in a model with censoring,
one may see exactly one outcome equal to zero with (strictly) positive
probability. Thus, whenever the inner product $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ places
mass at zero, the (random) function $t\mapsto \dot{m}_1^\texttt{tls}(t,\boldsymbol{Y})$
fails to be differentiable at $t=\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta}_0$ with (strictly)
positive probability, invalidating the use of the LDCT.
\subsection{Extended Asymptotic Normality Result}\label{sec:DiffOfG} In this
subsection, we modify the argument in \citet{honore_trimmed_1992} to allow the
distribution of $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ to place mass at zero. We show that
Honor{\'e}'s assumptions, corresponding to our Assumptions
\ref{assu:Non-Degeneracy-TLS}--\ref{assu:RankRegressors-TLS}, imply
differentiability of $\boldsymbol{G}$ at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$. We also derive an explicit
expression for the Jacobian $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$, which coincides with the
Hessian of the expected loss. Our calculations further show that
$\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$ generally differs from $\boldsymbol{\Gamma}_0^{\tt{tls}}$, unless
mass at zero in the distribution of $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ is ruled out a
priori. By deriving equivalent representations of this Hessian, we obtain easily
interpretable necessary and sufficient conditions for its invertibility. We then
state an asymptotic normality result that allows $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ to
have mass at zero.
\begin{thm}[\textbf{Trimmed Least Squares Hessian
Existence}]\label{thm:JacobianExistence-TLS} Under Assumptions
\ref{assu:Non-Degeneracy-TLS}--\ref{assu:RankRegressors-TLS}, the function
$\boldsymbol{G}:\mathbb{R}^K\to\mathbb{R}^K$ defined by \eqref{eq:ExpectedScore-TLS} is differentiable at
$\boldsymbol{\theta}=\boldsymbol{\theta}_0$ with the resulting matrix
$\mathbf{J}^{\tt{tls}}_0:=\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$ given by
\begin{equation}\label{eq:Jacobian}
\mathbf{J}^{\tt{tls}}_0 =\mathrm{E}\sbr[2]{\del[2]{1-F_{\varepsilon\mid\boldsymbol{W}}\del[1]{-\alpha-\min\cbr[0]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top}.
\end{equation}
Equivalently,
\begin{align}
\mathbf{J}^{\tt{tls}}_0 =\mathrm{E}\Big[\Big(
&\mathbf{1}\{Y_1>0\}\del[1]{\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<0\}+\tfrac{1}{2}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0\}}\notag\\
+&\mathbf{1}\{Y_2>0\}\del[1]{\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0>0\}+\tfrac{1}{2}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0\}} \Big)\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\Big].\label{eq:JacobianMidpoint}
\end{align}
\end{thm}
The expressions in \eqref{eq:Jacobian} and \eqref{eq:JacobianMidpoint}
make it clear that $(Y_1,\boldsymbol{X}_1)$ and $(Y_2,\boldsymbol{X}_2)$ enter the Hessian
$\mathbf{J}^{\tt{tls}}_0$ in a symmetric manner, so that the (time) labeling is
irrelevant. The version in \eqref{eq:JacobianMidpoint} facilitates plug-in
estimation of $\mathbf{J}^{\tt{tls}}_0$, which we cover in Section
\ref{sec:TLSAsymptoticNormalityAndVarianceEstimation}. In turn, the version in
\eqref{eq:Jacobian} facilitates comparison with Honor{\'e}'s
$\boldsymbol{\Gamma}_0^{\tt{tls}}$. To this end, iterate expectations to write the latter as
\[
\boldsymbol{\Gamma}_0^{\tt{tls}}=\mathrm{E}\sbr[2]{\mathrm{E}\sbr[1]{\mathbf{1}\cbr[0]{-Y_2<\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<Y_1}\mid\boldsymbol{W}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top}.
\]
We next expand on the inner expectation. Rewrite the inner
expectation indicator as
\[
1-\mathbf{1}\{ Y_{1}\leqslant\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} -\mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} +\mathbf{1}\{ Y_{1}\leqslant \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} \mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\}.
\]
Because $Y_{1}$ and $Y_{2}$ are non-negative, the only way both indicators
appearing in the previous display can be turned on is if
$\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0$. Again using non-negativity, we deduce
\[
\mathbf{1}\{ Y_{1}\leqslant \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} \mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\}
=\mathbf{1}\{ \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0\} \mathbf{1}\{ Y_{1}=0\} \mathbf{1}\{ Y_{2}=0\},
\]
from which it follows that $\mathbf{1}\{-Y_2<\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<Y_1\}$ can be written as
\begin{align*}
&1-\mathbf{1}\{ Y_{1}\leqslant\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} -\mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} + \mathbf{1}\{ \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0\} \mathbf{1}\{ Y_{1}=0\} \mathbf{1}\{ Y_{2}=0\}.
\end{align*}
Taking conditional expectations, the inner expectation in $\boldsymbol{\Gamma}_0^{\tt{tls}}$
works out as
\begin{align*}
&\mathrm{E}\sbr[1]{\mathbf{1}\{-Y_2<\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<Y_1\}\mid\boldsymbol{W}}\\
&\qquad =\begin{cases}
1-F_{\varepsilon\mid\boldsymbol{W}}\del[0]{V_1}, & \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<0,\\
1-F_{\varepsilon\mid\boldsymbol{W}}\del[0]{V_1}-F_{\varepsilon\mid\boldsymbol{W}}\del[0]{V_2}+F_{\boldsymbol{\varepsilon}\mid\boldsymbol{W}}\del[0]{V_1,V_2}, & \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0,\\
1-F_{\varepsilon\mid\boldsymbol{W}}\del[0]{V_2}, & \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0>0,
\end{cases}
\end{align*}
where (due to space constraints) we introduced the shorthand notations
$V_1:=-(\alpha+\boldsymbol{X}_1^\top\boldsymbol{\theta}_0)$ and
$V_2:=-(\alpha+\boldsymbol{X}_2^\top\boldsymbol{\theta}_0)$.
Since $V_2-V_1=\boldsymbol{X}_1^\top\boldsymbol{\theta}_0-\boldsymbol{X}_2^\top\boldsymbol{\theta}_0=\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$,
we have $V_1\gtreqqless V_2$ if
$\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\lesseqqgtr0$. With this newly introduced shorthand
notation, the inner expectation in our Hessian \eqref{eq:Jacobian} can be
written as
$
1-F_{\varepsilon\mid\boldsymbol{W}}\del[0]{\max\cbr{V_1,V_2}},
$
which agrees with the previous display whenever $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\neq0$.
Hence, if this case can be ruled out
[i.e.,~$\mathbb{P}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0)=0$], then Honor{\'e}'s
$\boldsymbol{\Gamma}_0^{\tt{tls}}$ \emph{is} equal to $\mathbf{J}^{\tt{tls}}$.
However, when $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0$, so that $V_1$ and $V_2$ take a common
value (here:~$V$), the inner expectations differ by
\begin{align*}
&\sbr[1]{1-F_{\varepsilon\mid\boldsymbol{W}}\del[0]{V}}-\mathrm{E}\sbr[1]{\mathbf{1}\{ -Y_{2}<\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta}_{0}<Y_{1}\} \mid\boldsymbol{W}}\\
&\qquad =F_{\varepsilon\mid\boldsymbol{W}}\del[0]{V}-F_{\boldsymbol{\varepsilon}\mid\boldsymbol{W}}\del[0]{V,V}=\lim_{u\to\infty}F_{\boldsymbol{\varepsilon}\mid\boldsymbol{W}}\del[0]{V,u}-F_{\boldsymbol{\varepsilon}\mid\boldsymbol{W}}\del[0]{V,V},
\end{align*}
which is non-negative by monotone increasingness of CDFs. In the Section
\ref{sec:Counterexample} counterexample, we have both $V_{1}\equiv0$ and
$V_{2}\equiv0$, and the (conditional) joint CDF factors into the product of the
(standard normal) marginals,
$F_{\boldsymbol{\varepsilon}\mid\boldsymbol{W}}(u_{1},u_{2})=\Phi(u_{1})\Phi(u_{2})$. Since there
$\mathbb{P}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0)=1$, the difference in the previous display
captures the only relevant case. The difference is then
$\Phi(0)-\Phi(0)\Phi(0)=\frac{1}{4}$, which is precisely the previously
demonstrated discrepancy.
\begin{rem}
[\textbf{Alternative Hessian Expressions}]\label{rem:JacobianAlt1} As
established in the proof of Theorem \ref{thm:JacobianExistence-TLS}, the Hessian
$\mathbf{J}^{\tt{tls}}_0$ can also be expressed using either of the following
expressions:
\begin{align}
\mathbf{J}^{\tt{tls}}_0&=\mathrm{E}\sbr[2]{\del[2]{\mathbf{1}\{Y_1>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\leqslant0\}+\mathbf{1}\{Y_2>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0>0\}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top},\label{eq:JacobianAlt1} \\
\mathbf{J}^{\tt{tls}}_0&=\mathrm{E}\sbr[2]{\del[2]{\mathbf{1}\{Y_1>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<0\}+\mathbf{1}\{Y_2>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\geqslant0\}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top}\label{eq:JacobianAlt2}.
\end{align}
The version in \eqref{eq:JacobianMidpoint} is the average of the
\eqref{eq:JacobianAlt1} and \eqref{eq:JacobianAlt2} right-hand
sides.\hfill$\diamondsuit$
\end{rem}
\begin{rem}[\textbf{Relaxing
Exchangeability}]\label{rem:RelaxingExchangeability-TLS} Inspecting the proof of
Theorem \ref{thm:JacobianExistence-TLS} (specifically, the proof of Lemma
\ref{lem:ExpectedTrimmedSqLossDerivative}), we see that it does not actually
require full conditional exchangeability of $\varepsilon_1$ and $\varepsilon_2$
given $\boldsymbol{W}$. For the conclusions of Theorem \ref{thm:JacobianExistence-TLS} to
hold, Assumption \ref{assu:Exchangeability-TLS} can be replaced with the weaker
condition of conditional \emph{stationarity}: conditional on $\boldsymbol{W}$,
$\varepsilon_1$ and $\varepsilon_2$ are identically distributed. Some of the
generalizations made in \citet[Section 7.1]{arellano2001panel} for Tobit-type
models with fixed effects similarly only require the weaker assumption of
conditional stationarity. When their $\psi(\cdot)$ and $\xi(\cdot)$ are both the
identity mapping on the real line, the resulting estimator is precisely the
\citet{honore_trimmed_1992} TLS estimator studied here, but motivated
differently. \citet[p.~3274]{arellano2001panel} observe that ``[i]t follows from
standard results about extremum estimators that the resulting estimator will be
consistent and $\sqrt n$ asymptotically normal'', but do not comment on the form
of the limit variance components. The Hessian expressions in Theorem
\ref{thm:JacobianExistence-TLS} and Remark \ref{rem:JacobianAlt1} are therefore
also relevant for this \citet{arellano2001panel} generalization of the TLS
estimator.\hfill$\diamondsuit$
\end{rem}
To state the asymptotic normality result, we need to make sure that the Hessian
$\mathbf{J}^{\tt{tls}}_0$ is invertible. To this end, we will impose the following
condition:
\begin{assumption}[\textbf{Trimmed Least Squares Score Jacobian
Invertibility}]\label{assu:JacobianInvertibility-TLS} There is no proper
linear subspace of $\mathbb{R}^K$ containing the random variable
$(\mathbf{1}\{Y_1>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\leqslant0\}+\mathbf{1}\{Y_2>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0>0\})\Delta\boldsymbol{X}$
with probability one.
\end{assumption}
This assumption is easy to interpret and implies via
\eqref{eq:JacobianAlt1} that the matrix $\mathbf{J}^{\tt{tls}}_0$ is invertible.
Alternatively, we could state an assumption based on the expression for
$\mathbf{J}^{\tt{tls}}_0$ in \eqref{eq:JacobianAlt2}. Strictly speaking, Assumption
\ref{assu:JacobianInvertibility-TLS} does not appear in
\citet{honore_trimmed_1992} but the only purpose of this assumption is to ensure
the invertibility of the Hessian $\mathbf{J}^{\tt{tls}}_0$, and \citet[Theorem
2(iv)]{honore_trimmed_1992} did impose invertibility of the corresponding
Hessian $\boldsymbol{\Gamma}_0^{\tt{tls}}$.
Using our new understanding of the Hessian $\mathbf{J}^{\tt{tls}}_0$, we next state an
extended asymptotic normality result for the TLS estimator that does not require
the inner product $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ to have no mass at zero.
\begin{thm}[\textbf{Asymptotic Normality of Trimmed Least
Squares}]\label{thm:AsymptoticNormality-TLS} Let Assumptions
\ref{assu:Non-Degeneracy-TLS}--\ref{assu:JacobianInvertibility-TLS} hold, and
suppose that the expectations involved in defining the matrix
$\mathbf{V}_0^{\tt{tls}}$ exist (in $\mathbb{R}^{K\times K}$), and that this matrix is of full
rank. Then the TLS estimator satisfies
\[
\sqrt{n}\del[1]{\widehat\boldsymbol{\theta}^{\tt{tls}}-\boldsymbol{\theta}_0}\rightsquigarrow\mathcal{N}\left(\mathbf{0},(\mathbf{J}_0^{\tt{tls}})^{-1}\mathbf{V}_0^{\tt{tls}}(\mathbf{J}_0^{\tt{tls}})^{-1}\right)\;\text{in}\;\mathbb{R}^K.
\]
\end{thm}
The asymptotic normality result in Theorem \ref{thm:AsymptoticNormality-TLS}
differs from \citet[Theorem 2(iv)]{honore_trimmed_1992} in the sense that it
replaces the matrix $\boldsymbol{\Gamma}_0^{\tt{tls}}$ in
\eqref{eq:HonoreAsymptoticNormality-TLS} by $\mathbf{J}_0^{\tt{tls}}$. Whenever
$\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\neq 0$ with probability one, the two matrices coincide.
However, the matrices are in general different if $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0 = 0$
with (strictly) positive probability.
\subsection{Asymptotic Variance
Estimation}\label{sec:TLSAsymptoticNormalityAndVarianceEstimation} For the
asymptotic normality to yield a practical approximation, we need to consistently
estimate the asymptotic variance components. For $\mathbf{V}_0^{\tt{tls}}$, we use the
plug-in estimator from \citet{honore_trimmed_1992}:
\[
\widehat{\mathbf{V}}^{\tt{tls}}:=\frac{1}{n}\sum_{i=1}^n \dot{m}_1^{\tt{tls}}(\Delta\boldsymbol{X}_i^\top\widehat\boldsymbol{\theta}^{\tt{tls}},\boldsymbol{Y}_i)^2\Delta\boldsymbol{X}_i\Delta\boldsymbol{X}_i^\top.\footnote{In \citet{honore_trimmed_1992}, this estimator is denoted $\hat{V}_4$.}
\]
For the Hessian $\mathbf{J}_0^{\tt{tls}}$, we apply the analogy principle to
\eqref{eq:JacobianMidpoint} to arrive at
\begin{align}
\widehat\mathbf{J}^{\tt{tls}}=\frac{1}{n}\sum_{i=1}^n\bigg(
&\mathbf{1}\{Y_{i1}>0\}\del[2]{\mathbf{1}\{\Delta\boldsymbol{X}_i^\top\widehat\boldsymbol{\theta}^{\tt{tls}}<0\}+\frac{1}{2}\mathbf{1}\{\Delta\boldsymbol{X}_i^\top\widehat\boldsymbol{\theta}^{\tt{tls}}=0\}}\notag\\
+&\mathbf{1}\{Y_{i2}>0\}\del[2]{\mathbf{1}\{\Delta\boldsymbol{X}_i^\top\widehat\boldsymbol{\theta}^{\tt{tls}}>0\}+\frac{1}{2}\mathbf{1}\{\Delta\boldsymbol{X}_i^\top\widehat\boldsymbol{\theta}^{\tt{tls}}=0\}} \bigg)\Delta\boldsymbol{X}_i\Delta\boldsymbol{X}_i^\top.\label{eq:JacobianMidpointEstimator}
\end{align}
Alternatively, in order to estimate the Hessian $\mathbf{J}_0^{\tt{tls}}$, we could
use the equivalent expressions in \eqref{eq:JacobianAlt1} and
\eqref{eq:JacobianAlt2}.
These estimators are (strongly) consistent under the assumptions of Theorem
\ref{thm:AsymptoticNormality-TLS}:
\begin{thm}[\textbf{Plug-in Variance Estimator Consistency for
TLS}]\label{thm:VarianceConsistency-TLS} Let the assumptions of Theorem
\ref{thm:AsymptoticNormality-TLS} hold. Then
$\widehat\mathbf{V}^{\tt{tls}}\to_{\mathrm{a.s.}}\mathbf{V}_0^{\tt{tls}}$ and
$\widehat\mathbf{J}^{\tt{tls}}\to_{\mathrm{a.s.}}\mathbf{J}_0^{\tt{tls}}$.
\end{thm}
Theorem \ref{thm:VarianceConsistency-TLS} implies that the limit
variance $(\mathbf{J}_0^{\tt{tls}})^{-1}\mathbf{V}_0^{\tt{tls}}(\mathbf{J}_0^{\tt{tls}})^{-1}$ is
(strongly) consistently estimated by
$(\widehat\mathbf{J}^{\tt{tls}})^{-1}\widehat\mathbf{V}^{\tt{tls}}(\widehat\mathbf{J}^{\tt{tls}})^{-1}$,
which facilitates hypothesis testing and the construction of confidence
intervals.
\begin{rem}[\textbf{Comparison with \citet{honore2000estimation}}]
\citet[Section 2.1]{honore2000estimation} discuss estimation of the censored
regression model with fixed effects, allowing the number of time periods $T_i$
to exceed two and/or be individual specific (unbalanced panel data). While
\emph{ibid.}~(Section 2.1) covers a whole class of estimators, in the special
case of their $\xi(\cdot)$ being the identity mapping on the real line, the
estimator in their (6) becomes a TLS estimator based on all pairs of
time periods. In the balanced case $(T_i=T\geqslant2)$, their estimator is (in
our notation) any minimizer of
\[
\boldsymbol{\theta}\mapsto \frac{1}{n}\sum_{i=1}^n \sum_{1\leqslant \tau < \tau' \leqslant T}
{m}^{\tt{tls}}\del[1]{(\boldsymbol{X}_{i\tau}-\boldsymbol{X}_{i\tau'})^\top\boldsymbol{\theta},(Y_{i\tau},Y_{i\tau'})},
\]
which naturally generalizes the two-period TLS estimator studied in this paper.
Appropriately extending our assumptions to their multi-period case, one can
establish ($\sqrt n$-)consistency and asymptotic normality of their multi-period
TLS estimator and reuse the argument leading to Theorem
\ref{thm:JacobianExistence-TLS} to show that the Hessian of the expected loss in
the multi-period setting is given by an expression that naturally generalizes
\eqref{eq:JacobianMidpoint} to accommodate more than one pair of time periods.
One can then construct a (strongly) consistent estimator of this Hessian by
applying the analogy principle, as we did in
\eqref{eq:JacobianMidpointEstimator} for the $T=2$ case. \hfill$\diamondsuit$
\end{rem}
\begin{rem}[\textbf{A Plug-In Estimator for
\citet{honore_pairwise_1994}}]\label{rem:ComparisonWithHonorePowellJacobian}
Consider the {\em cross-sectional} censored regression model $Y =
\max\{0,Y^*\}$, where $Y^* = \boldsymbol{X}^{\top}\boldsymbol{\theta}_0 + \varepsilon$ and
$\varepsilon$ is independent of $\boldsymbol{X}$. This model was studied in
\citet{honore_pairwise_1994}, who developed, among other things, the TLS
estimator of the vector of parameters $\boldsymbol{\theta}_0$ in this model, proved its
($\sqrt{n}$-)consistency, and derived the corresponding asymptotic normality result. On
\emph{ibid.}~(p.~260), they gave an expression for the Hessian of the expected
loss, which enters the asymptotic variance formula. In our notation, their
Hessian takes the following form:
\begin{align}
\mathrm{E}\sbr[2]{\del[2]{1-F_{\varepsilon}\del[1]{-\min\cbr[0]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top},\label{eq: hessian cross section}
\end{align}
where the pairs $(Y_1,\boldsymbol{X}_1)$ and $(Y_2,\boldsymbol{X}_2)$ represent independent units of
observation and $F_{\varepsilon}$ is the CDF of $\varepsilon$. On \emph{ibid.}
(p.~261), they also proposed a generic numerical derivative estimator of this
Hessian. Due to numerical differentiation, however, their estimator introduces a
tuning parameter through the choice of the stepsize.
Our point in this remark is to show that one can actually estimate the Hessian
in \eqref{eq: hessian cross section} without introducing extra tuning
parameters. To see that, let $\boldsymbol{Z}_i := (Y_i,\boldsymbol{X}_i)$, $i\in\{1,2,\dots,n\}$, be
a random sample from the distribution of $\boldsymbol{Z}:=(Y,\boldsymbol{X})$, and observe that our
panel-data Hessian $\mathbf{J}^{\tt{tls}}_0$ in \eqref{eq:Jacobian} reduces to the
cross-sectional Hessian in \eqref{eq: hessian cross section} upon setting
$\alpha\equiv0$ and imposing independence of $\varepsilon_1$ and $\varepsilon_2$
from $\boldsymbol{W}$ in the former expression. Theorem \ref{thm:JacobianExistence-TLS}
thus reveals, via \eqref{eq:JacobianMidpoint}, that the cross-sectional Hessian
in \eqref{eq: hessian cross section} is equal to $\mathrm{E}[h(\boldsymbol{Z}_1,\boldsymbol{Z}_2;\boldsymbol{\theta}_0)]$,
where the function $h:\mathbb{R}^{1+K}\times\mathbb{R}^{1+K}\times\mathbb{R}^K\to\mathbb{R}^{K\times K}$ is
given by
\begin{align*}
h(\boldsymbol{z}_1,\boldsymbol{z}_2;\boldsymbol{\theta})
:=\Big(&\mathbf{1}\{y_1>0\}\del[1]{\mathbf{1}\{\Delta\boldsymbol{x}^\top\boldsymbol{\theta}<0\}+\tfrac{1}{2}\mathbf{1}\{\Delta\boldsymbol{x}^\top\boldsymbol{\theta}=0\}}\\
+&\mathbf{1}\{y_2>0\}\del[1]{\mathbf{1}\{\Delta\boldsymbol{x}^\top\boldsymbol{\theta}>0\}+\tfrac{1}{2}\mathbf{1}\{\Delta\boldsymbol{x}^\top\boldsymbol{\theta}=0\}}\Big)\Delta\boldsymbol{x}\Delta\boldsymbol{x}^\top,
\end{align*}
and we denote $\boldsymbol{z}_1:=(y_1,\boldsymbol{x}_1)$, $\boldsymbol{z}_2:=(y_2,\boldsymbol{x}_2)$, and $\Delta\boldsymbol{x} :=
\boldsymbol{x}_1 - \boldsymbol{x}_2$. Let $\widehat{\boldsymbol{\theta}}$ be the \citet{honore_pairwise_1994} TLS
estimator, i.e.~their estimator with
$\Xi\left(\cdot\right)=\textstyle{\frac{1}{2}}\left(\cdot\right)^2$. Then a
natural estimator of the cross-sectional Hessian in \eqref{eq: hessian cross
section} is of the plug-in form
\begin{align}
\frac{1}{n(n-1)}\sum_{1\leqslant i\neq j\leqslant n}h(\boldsymbol{Z}_i,\boldsymbol{Z}_j;\widehat\boldsymbol{\theta}),\label{eq:HonorePowellPlugInEstimator}
\end{align}
which is an approximate second-order U-statistic. Note that this estimator
involves no tuning parameter, and is therefore free of the stepsize problem
arising due to numerical differentiation. The consistency of this plug-in
estimator can be established using an argument paralleling the one used in the
proof of Theorem \ref{thm:VarianceConsistency-TLS}. The main difference is that, instead of establishing a uniform law of large
numbers involving \emph{simple} averages (or, equivalently, an empirical
process), to accommodate the expression in
\eqref{eq:HonorePowellPlugInEstimator}, one must now work with
\emph{generalized} averages (leading to a second-order U-process). To this end,
we refer the reader to \citet{nolan1987_u_processes_rates}.\hfill$\diamondsuit$
\end{rem}
\begin{rem}[\textbf{Comparison with
\citet{powell1984least,powell1986symmetrically}}]\label{rem:ComparisonWithPowell}
As noted in \citet{honore_trimmed_1992}, the rationale behind the TLS and TLAD
estimators considered in this paper can be viewed as a bivariate generalization
of the idea behind the \citet{powell1986symmetrically} symmetrically trimmed LS
estimators for censored Tobit models without fixed effects. The approach adopted
in \citet{powell1986symmetrically} is, in turn, closely related to the censored
LAD (CLAD) estimator proposed in \citet{powell1984least}---a key paper in the
censored regression literature.
In our notation, \citet{powell1984least} models the conditional median of a
non-negative response variable $Y$ given covariates $\boldsymbol{X}$ as
$\mathrm{Med}(Y|\boldsymbol{X})=\max\{0,\boldsymbol{X}^\top\boldsymbol{\theta}_0\}$. To ensure asymptotic
normality of the CLAD estimator, \citet[Assumption R.2]{powell1984least} rules
out regressors that are orthogonal to $\boldsymbol{\theta}_0$ with positive probability. The
possibility of such regressors is precisely the cause of difficulty in our
analysis, cf.~the discussion following Theorem
\ref{thm:JacobianExistence-TLS}.\hfill$\diamondsuit$
\end{rem}
\subsection{Consistency of the Honor{\'e} (1992) Hessian
Estimator}\label{sec:ConsistencyHonoreHessianEstimator-TLS} In
\citet{honore_trimmed_1992} the matrix $\boldsymbol{\Gamma}_0^{\tt{tls}}$ in
\eqref{eq:VarianceSandwichBread} is estimated by its empirical analogue
\begin{equation}\label{eq:Honore1992HessianEstimator}
\widehat\boldsymbol{\Gamma}^{\tt{H92}}:=\frac{1}{n}
\sum_{i=1}^{n}\boldsymbol{1}\{-Y_{i2}<\Delta\boldsymbol{X}_{i}^{\top}\widehat{\boldsymbol{\theta}}^{\tt{tls}}<Y_{i1}\}
\Delta\boldsymbol{X}_{i}\Delta\boldsymbol{X}_{i}^{\top}.\footnote{In
\cite{honore_trimmed_1992}, this estimator is denoted by $\widehat{\Gamma}_4$.}
\end{equation}
The TLS sandwich variance consistency results in \citet[Theorem
3]{honore_trimmed_1992} involve showing that $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ is a
consistent estimator of $\boldsymbol{\Gamma}_0^{\tt{tls}}$. As we have shown by example in
Section \ref{sec:Counterexample}, $\boldsymbol{\Gamma}_0^{\tt{tls}}$ is not equal to the
Hessian $\mathbf{J}_0^{\tt{tls}}$ of the expected loss in general. As
$\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ is targeting $\boldsymbol{\Gamma}_0^{\tt{tls}}$, one would
expect that $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ is not a consistent estimator of the
TLS Hessian $\mathbf{J}_0^{\tt{tls}}$ in general. Surprisingly, as we establish below,
$\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ nevertheless converges to the TLS Hessian
$\mathbf{J}_0^{\tt{tls}}$.
\begin{thm}[\textbf{Consistency of the \citet{honore_trimmed_1992} Hessian
Estimator for TLS}]\label{thm:ConsistencyHonoreHessianEstimator-TLS} Let the
assumptions of Theorem \ref{thm:AsymptoticNormality-TLS} hold. Then
$\widehat\boldsymbol{\Gamma}^{\tt{H92}}=\mathbf{J}_0^{\tt{tls}}+o_{L^1}(1)+o_{\mathrm{a.s.}}(1)=\mathbf{J}_0^{\tt{tls}}+o_{\mathbb{P}}(1)$.
\end{thm}
This counterintuitive result can be explained as follows. Decompose
$\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ into two parts:
\begin{align*}
\widehat{\mathbf{\Gamma}}^{\tt{H92}} & =\frac{1}{n}\sum_{i=1}^{n}\left(\boldsymbol{1}\{-Y_{i2}<\Delta\boldsymbol{X}_{i}^{\top}\widehat{\boldsymbol{\theta}}^{\tt{tls}}<0\}+\boldsymbol{1}\{0<\Delta\boldsymbol{X}_{i}^{\top}\widehat{\boldsymbol{\theta}}^{\tt{tls}}<Y_{i1}\}\right)\Delta\boldsymbol{X}_{i}\Delta\boldsymbol{X}_{i}^{\top}\tag{=:\ensuremath{\widehat{\mathbf{L}}^{\tt{H92}}}}\\
& \quad+\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{1}\{Y_{i1}>0,Y_{i2}>0,\Delta\boldsymbol{X}_{i}^{\top}\widehat{\boldsymbol{\theta}}^{\tt{tls}}=0\}\Delta\boldsymbol{X}_{i}\Delta\boldsymbol{X}_{i}^{\top}\tag{=:\ensuremath{\widehat{\mathbf{R}}^{\tt{H92}}}}.
\end{align*}
Although $\widehat\mathbf{L}^{\tt{H92}}$ and $\widehat\mathbf{J}^{\tt{tls}}$ are not
identical term by term, the leading component $\widehat\mathbf{L}^{\tt{H92}}$ is
closely related to our Hessian estimator in
\eqref{eq:JacobianMidpointEstimator}, which is the empirical analogue of
$\mathbf{J}_0^{\tt{tls}}$. The proof of Theorem
\ref{thm:ConsistencyHonoreHessianEstimator-TLS} shows that
$\widehat\mathbf{L}^{\tt{H92}}$ converges to $\mathbf{J}_0^{\tt{tls}}$, while the remainder
term $\widehat\mathbf{R}^{\tt{H92}}$ converges to the zero matrix. Key to these
findings is our extended asymptotic normality result in Theorem
\ref{thm:AsymptoticNormality-TLS}. Specifically, even if
$\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ has mass at zero, the absolutely continuous limiting
distribution of $\sqrt n(\widehat\boldsymbol{\theta}^{\tt{tls}}-\boldsymbol{\theta}_0)$ implies that,
for any fixed $\boldsymbol{x}\in\mathbb{R}^K\backslash\{\boldsymbol{0}\}$,
$\mathbb{P}(\boldsymbol{x}^\top\widehat\boldsymbol{\theta}^{\tt{tls}}=0)\to0$. We stress that Theorem
\ref{thm:ConsistencyHonoreHessianEstimator-TLS} is \emph{not} a simple
consequence of consistency of $\widehat\boldsymbol{\theta}^{\tt{tls}}$, established in
\citet[Theorem 1(iv)]{honore_trimmed_1992}. Indeed, even with a uniform law of
large numbers allowing us to replace sample averages in
$\widehat{\boldsymbol{\Gamma}}^{\tt{H92}}$ by their population counterparts (see the proof
of Theorem \ref{thm:ConsistencyHonoreHessianEstimator-TLS}), the resulting map
$\boldsymbol{\theta}\mapsto\boldsymbol{\Gamma}(\boldsymbol{\theta})$ need not be continuous at $\boldsymbol{\theta}_0$.
Consequently, a continuous mapping argument need not apply. In the
counterexample of Section \ref{sec:Counterexample}, one finds
$\Gamma(\theta)=[1-\Phi(|\theta|)]\boldsymbol{1}\{\theta\neq0\}+\tfrac{1}{4}\boldsymbol{1}\{\theta=0\}$,
which has a jump discontinuity at $\theta=0=\theta_0$. The jump size
$(\tfrac{1}{4})$ is exactly the difference between $\mathbf{J}_0^{\tt{tls}}$ and
$\boldsymbol{\Gamma}_0^{\tt{tls}}$ in that example.
A by-product of Theorem \ref{thm:ConsistencyHonoreHessianEstimator-TLS}
is that, under our maintained assumptions, the TLS Hessian consistency
statement in \citet[Theorem 3]{honore_trimmed_1992} does not hold
without additional regularity conditions ensuring continuity at $\boldsymbol{\theta}_0$.
\section{Trimmed Least Absolute Deviations}\label{sec:TLAD} In this section, we
focus on the TLAD estimator. We abbreviate the trimmed absolute loss, $m_{\Xi}$
in \eqref{eq:TrimmedGenericLoss} with $\Xi=\left|\cdot\right|$, by
$m^{\texttt{tlad}}$, which takes the form
\begin{equation}\label{eq:TrimmedAbsLoss}
m^{\texttt{tlad}}(t,\boldsymbol{y})=
\begin{cases}
|y_{1}|-\left(t+y_{2}\right)\mathrm{sgn}(y_{1}), & t\leqslant-y_{2},\\
|y_{1}-y_{2}-t|, & t\in\left(-y_{2},y_{1}\right),\\
|-y_{2}|-\left(t-y_{1}\right)\mathrm{sgn}(-y_{2}), & t\geqslant y_{1}.
\end{cases}
\end{equation}
\subsection{Assumptions for Trimmed Least Absolute
Deviations}\label{sec:Assumptions-TLAD} For notational convenience, abbreviate
the TLAD estimator $\widehat\boldsymbol{\theta}_\Xi$ in \eqref{eq:TrimmedEstimator} with
$\Xi=\left|\cdot\right|$ by $\widehat{\boldsymbol{\theta}}^{\texttt{tlad}}$. Also, let
$\boldsymbol{W}$, $\mathcal W$, $\boldsymbol{w}$, $\boldsymbol{\varepsilon}$, $\mathcal E$, and $\boldsymbol{e}$ be the same as
in Section \ref{sec:TLS}. Consider the following
assumptions.\footnote{Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:RankRegressors-TLAD} are from
Assumptions S.2, M.2, E.1, E.3, E.4 and R.1, respectively, in
\citet{honore_trimmed_1992}.}
\begin{assumption}[\textbf{Non-Degeneracy}]\label{assu:Non-Degeneracy-TLAD} The
probability $\mathbb{P}(Y_1>0, Y_2>0)$ is strictly positive. \end{assumption}
\begin{assumption}[\textbf{Integrability}]\label{assu:MomentConditions-TLAD} All
of the following expectations are finite:
\[
\mathrm{E}[\|\boldsymbol{X}_1\|_2^2],\;\mathrm{E}[\|\boldsymbol{X}_2\|_2^2],\;\mathrm{E}[\|\alpha\Delta\boldsymbol{X}\|_2],\;
\mathrm{E}[\|\varepsilon_1\Delta\boldsymbol{X}\|_2]\quad\text{and}\quad
\mathrm{E}[\|\varepsilon_2\Delta\boldsymbol{X}\|_2].
\]
\end{assumption}
\begin{assumption}[\textbf{Continuity}]\label{assu:AbsoluteContinuity-TLAD} The
conditional distribution of $(\varepsilon_1,\varepsilon_2)$ given $\boldsymbol{W}$ is
absolutely continuous with respect to the Lebesgue measure. \end{assumption}
\begin{assumption}[\textbf{Exchangeability}]\label{assu:Exchangeability-TLAD}
Conditional on $\boldsymbol{W}$, $\varepsilon_{1}$ and $\varepsilon_{2}$ are
exchangeable. \end{assumption}
Since Assumptions \ref{assu:AbsoluteContinuity-TLAD} and
\ref{assu:Exchangeability-TLAD} are the same as Assumptions \ref{assu:continuity
tls} and \ref{assu:Exchangeability-TLS}, following the reasoning in Section
\ref{sec:Assumptions-TLS}, there exists a function $(\boldsymbol{w},\boldsymbol{e})\mapsto
f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$, mapping $\mathcal W\times\mathcal E$ to
$[0,\infty)$, that is a version of the PDF of the pair $\boldsymbol{\varepsilon} =
(\varepsilon_1,\varepsilon_2)$ conditional on $\boldsymbol{W} = \boldsymbol{w}$, which is measurable
in $(\boldsymbol{w},\boldsymbol{e})$, and is such that $f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_1,e_2) =
f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_2,e_1)$ for all $\boldsymbol{w}\in\mathcal W$ and $\boldsymbol{e} =
(e_1,e_2)\in\mathcal E$. Also, let $(\boldsymbol{w},e)\mapsto f_{\varepsilon\mid\boldsymbol{w}}(e) =
\int_{\mathbb R}f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e,u)\dif u$ be the corresponding version of
the common marginal PDF of $\varepsilon_1$ and $\varepsilon_2$ conditional on
$\boldsymbol{W} = \boldsymbol{w}$ and let $(\boldsymbol{w},e)\mapsto f_{\varepsilon_1 - \varepsilon_2\mid\boldsymbol{w}}(e)
= \int_{\mathbb R} f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(u+e,u)\dif u$ be the corresponding
version of the PDF of the difference $\varepsilon_1 - \varepsilon_2$ conditional
on $\boldsymbol{W} = \boldsymbol{w}$.
\begin{assumption}[\textbf{Regularity}]\label{assu:Regularity-TLAD} There is a
constant $C\in(0,\infty)$ such that
$\sup_{e\in\mathbb{R}}f_{\varepsilon_{1}-\varepsilon_{2}\mid\boldsymbol{W}}(e)\leqslant C$ and
$\sup_{e\in\mathbb{R}}f_{\varepsilon\mid\boldsymbol{W}}(e)\leqslant C$ with probability one.
\end{assumption}
\begin{assumption}[\textbf{Rank of Regressors}]\label{assu:RankRegressors-TLAD}
There is no proper linear subspace of $\mathbb{R}^K$ containing the random variable
$\mathbf{1}\{\mathbb{P}\left(Y_1>0, Y_2>0\middle|\boldsymbol{X}_1,\boldsymbol{X}_2\right)>0\}\Delta\boldsymbol{X}$ with
probability one. \end{assumption}
\begin{assumption}[\textbf{Continuity, II}]\label{assu:extra continuity} The
functions $\boldsymbol{e}\mapsto f_{\boldsymbol{\varepsilon}\mid \boldsymbol{W}}(\boldsymbol{e})$, $e\mapsto
f_{\varepsilon\mid\boldsymbol{W}}(e)$, and $e\mapsto f_{\varepsilon_1 -
\varepsilon_2\mid\boldsymbol{W}}(e)$ are continuous with probability one.
\end{assumption}
Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:RankRegressors-TLAD} are the same as
the corresponding assumptions in \citet{honore_trimmed_1992}. Assumption
\ref{assu:extra continuity}, however, is not present in
\citet{honore_trimmed_1992}. We consider this assumption because the asymptotic
normality result in \citet{honore_trimmed_1992} may not hold without it. In
particular, in Section \ref{sec:Counterexample 2} below we provide a DGP that
satisfies Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:RankRegressors-TLAD} and is such that
the Hessian of the expected loss [see \eqref{eq:PopulationLossFunction-TLAD}
below], which appears in the asymptotic variance formula in
\citet{honore_trimmed_1992}, does not exist. We note also that Assumption
\ref{assu:extra continuity} may be stronger than necessary. For example, it
seems possible to obtain the asymptotic normality result assuming continuity of
the functions in Assumption \ref{assu:extra continuity} only on their respective
supports but we opt for a stronger than necessary conditions for clarity of the
argument.
Note also that Assumptions \ref{assu:Non-Degeneracy-TLAD},
\ref{assu:AbsoluteContinuity-TLAD}, \ref{assu:Exchangeability-TLAD}, and
\ref{assu:RankRegressors-TLAD} are the same as the corresponding assumptions for
the TLS estimator (Assumptions \ref{assu:Non-Degeneracy-TLS},
\ref{assu:continuity tls}, \ref{assu:Exchangeability-TLS}, and
\ref{assu:RankRegressors-TLS}, respectively). The TLAD moment conditions in
Assumption \ref{assu:MomentConditions-TLAD} are weaker than those for the TLS
estimator (Assumption \ref{assu:MomentConditions-TLS}), which reflects the fact
that the trimmed absolute loss is less sensitive to outliers than the trimmed
square loss. Assumptions \ref{assu:Regularity-TLAD} and \ref{assu:extra
continuity}, requiring certain PDFs to be bounded and continuous, do not have
analogs in the case of the TLS estimator.
\subsection{Asymptotic Normality in Honor{\'e} (1992)} To state the TLAD normality
result in \citet{honore_trimmed_1992}, introduce the $K\times K$ matrices\footnote{In
\citet{honore_trimmed_1992}, these matrices are denoted $V_3$ and $\Gamma_3$,
respectively. Honor{\'e}'s $\Gamma_3$ is stated in terms of probabilities and
densities of the censored $Y_1$ and $Y_2$, but it is clear from the underlying
proof that he means the latent variables $Y_1^\ast$ and $Y_2^\ast$,
respectively.}
\begin{equation}\label{eq:VarianceSandwichMeat-TLAD}
\left.\begin{aligned}
\mathbf{V}_0^{\tt{tlad}}:=
\mathrm{E}\Big[\Big(&\boldsymbol{1}\cbr[1]{Y_1>0}\boldsymbol{1}\cbr[1]{Y_1-Y_2>\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0}\\
+&\boldsymbol{1}\cbr[1]{Y_2>0}\boldsymbol{1}\cbr[1]{Y_1-Y_2<\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0}\Big)\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\Big]
\end{aligned}\right\}
\end{equation}
and
\begin{equation}\label{eq:VarianceSandwichBread-TLAD}
\left.\begin{aligned}
\boldsymbol{\Gamma}_0^{\tt{tlad}}
&:=\mathrm{E}\Big[\Big(2f_{Y_1^\ast-Y_2^\ast\mid\boldsymbol{W},Y_1^\ast>0,Y_2^\ast>0}\del[1]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0} \\
&\qquad+\mathbf{1}\cbr[1]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\geqslant0}\mathbb{P}\left(Y_2^\ast\leqslant0\middle|\boldsymbol{W}\right)f_{Y^\ast_1\mid\boldsymbol{W},Y_2^\ast\leqslant0}\del[1]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0} \\
&\qquad+\mathbf{1}\cbr[1]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<0}\mathbb{P}\left(Y_1^\ast\leqslant0\middle|\boldsymbol{W}\right)f_{Y^\ast_2\mid\boldsymbol{W},Y_1^\ast\leqslant0}\del[1]{-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0}\Big)\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\Big],
\end{aligned}\right\}
\end{equation}
where $(\boldsymbol{w},e) \mapsto
f_{Y_1^\ast-Y_2^\ast\mid\boldsymbol{w},Y_1^\ast>0,Y_2^\ast>0}\left(e\right)$ is the
conditional PDF of $Y_1^\ast-Y_2^\ast$ given $\boldsymbol{W}=\boldsymbol{w}$ and $\{Y_1^\ast>0\}\cap
\{Y_2^\ast>0\}$, $(\boldsymbol{w},e) \mapsto
f_{Y^\ast_1\mid\boldsymbol{w},Y_2^\ast\leqslant0}\left(e\right)$ is the conditional PDF of
$Y_1^\ast$ given $\boldsymbol{W}=\boldsymbol{w}$ and $Y_2^\ast\leqslant0$, and $(\boldsymbol{w},e)\mapsto
f_{Y^\ast_2\mid\boldsymbol{w},Y_1^\ast\leqslant0}\left(e\right)$ is the conditional PDF of
$Y_2^\ast$ given $\boldsymbol{W}=\boldsymbol{w}$ and $Y_1^\ast\leqslant0$.\footnote{Here and below, we
assume that the versions of all conditional PDFs and CDFs including latent
outcomes $Y_1^*$ and $Y_2^*$ are obtained (fixed) by combining the conditional
PDF $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$ with \eqref{eq:Outcomes}.}
\citet[Theorem 2(iii)]{honore_trimmed_1992} states that if Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:RankRegressors-TLAD} hold, the
expectations involved in defining the matrices $\mathbf{V}_0^{\tt{tlad}}$ and
$\boldsymbol{\Gamma}_0^{\tt{tlad}}$ exist (in $\mathbb{R}^{K\times K}$), and both matrices are of
full rank, then
\begin{equation}\label{eq:AsymptoticNormality-TLAD}
\sqrt{n}\del[1]{\widehat\boldsymbol{\theta}^{\texttt{tlad}}-\boldsymbol{\theta}_0}\rightsquigarrow\mathcal{N}\left(\mathbf{0},(\boldsymbol{\Gamma}_0^{\tt{tlad}})^{-1}\mathbf{V}_0^{\tt{tlad}}(\boldsymbol{\Gamma}_0^{\tt{tlad}})^{-1}\right)\;\text{in}\;\mathbb{R}^K.
\end{equation}
Among several steps, the proof of this result in \citet{honore_trimmed_1992}
includes establishing the existence of the Hessian of the expected loss [see
\eqref{eq:PopulationLossFunction-TLAD} below] at the true parameter value
$\boldsymbol{\theta}_0$. In our notation, this task corresponds to arguing that the function
$L:\mathbb{R}^K\to\mathbb{R}$ defined by
\begin{equation}\label{eq:PopulationLossFunction-TLAD}
L(\boldsymbol{\theta}):=\mathrm{E}[m^{\texttt{tlad}}(\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta},\boldsymbol{Y}) - m^{\texttt{tlad}}(0,\boldsymbol{Y})],\quad\boldsymbol{\theta}\in\mathbb{R}^{K},
\end{equation}
is twice differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$.\footnote{Note that because the
function $m^{\texttt{tlad}}(t,\boldsymbol{y})$ is Lipschitz continuous in its first
argument, it follows from Assumption \ref{assu:MomentConditions-TLAD} that the
expectation in \eqref{eq:PopulationLossFunction-TLAD} is well-defined. Also note
that, as in \citet{honore_trimmed_1992}, we work with the expected loss
\emph{difference} $\mathrm{E}[m^{\texttt{tlad}}(\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta},\boldsymbol{Y}) -
m^{\texttt{tlad}}(0,\boldsymbol{Y})]$ instead of
$\mathrm{E}[m^{\texttt{tlad}}(\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta},\boldsymbol{Y})]$ because the former loss
gives results under weaker regularity conditions.} To this end, Honor{\'e} used
the LDCT to differentiate once under the expectation, and then argued
differentiability at $\boldsymbol{\theta}_0$ of the resulting function to arrive at
$\boldsymbol{\Gamma}_0^{\tt{tlad}}$ in \eqref{eq:VarianceSandwichBread-TLAD}. In the next
subsection, however, we will show by example that Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:RankRegressors-TLAD} are not
sufficient to ensure the existence of the Hessian.
\subsection{Counterexample}\label{sec:Counterexample 2}
As in the Section \ref{sec:Counterexample} counterexample, let $K=1$,
$\alpha\equiv0$, $\boldsymbol{\theta}_0 = 0$, $\boldsymbol{X}_1\equiv2$ and $\boldsymbol{X}_2\equiv1$. Also, to
define the distribution of the pair $(\varepsilon_1,\varepsilon_2)$, let
$r:\mathbb R\to \mathbb R$ be a continuous function such that (i) $r(t)\geqslant 0$
for all $t\in\mathbb R$, (ii) $\int_{\mathbb R} r(t)\dif t = 1$, and (iii) $r(t)=0$
if $t\leqslant 1$ or $t\geqslant 3$. In addition, let $\mathscr E := \{0,2,4,\dots\}$,
$\mathscr O = \{1,3,5,\dots\}$, and let $\tilde h:[0,1]\to\{0,1\}$ be the function
defined by
$$
\tilde{h}(t)=\begin{cases}
1, & t\in(2^{-(k+1)},2^{-k}]\text{ for }k\in \mathscr E,\\
0, & t\in(2^{-(k+1)},2^{-k}]\text{ for }k\in \mathscr O,\\
0, & t=0.
\end{cases}
$$
Moreover, let $h:\mathbb R \to\mathbb R$ be the function defined by $h(t) =
3\tilde h(|t|) / 4$ for $t\in[-1,1]$ and $0$ otherwise. Note that $h(t)\geqslant0$
for all $t\in\mathbb R$ and
$$
\int_{\mathbb R} h(t)\dif t = \frac{3}{4}\int_{-1}^1 \tilde h(\left|t\right|)\dif t = \frac{3}{2}\int_{0}^1 \tilde h(t)\dif t = \frac{3}{2}\sum_{k\in\mathscr E}\frac{1}{2^{k+1}} = \frac{3}{2}\cdot \frac{1}{2}\cdot\frac{1}{1-\frac{1}{4}} = 1.
$$
Thus, the function $f\colon\mathbb R^2 \to \mathbb R$ defined by
$$
f(e_1,e_2) = 2h(e_1 - e_2)r(e_1 + e_2),\quad e_1,e_2\in\mathbb R,
$$
satisfies $f(e_1,e_2)\geqslant 0$ for all $e_1,e_2\in\mathbb R$ and
\begin{align*}
\int_{\mathbb{R}^2}f(\boldsymbol{e})\dif \boldsymbol{e}
& =2\int_{\mathbb{R}}\int_{\mathbb{R}}h(e_1-e_2)r(e_1+e_2)\dif e_1 \dif e_2\\
&=2\int_{\mathbb{R}}\int_{\mathbb{R}}h(s)r(s+2e_2)\dif s \dif e_2
=\int_{\mathbb{R}}\int_{\mathbb{R}}h(s)r(t)\dif s \dif t=1.
\end{align*}
Hence, $f$ is the PDF of a certain distribution on $\mathbb R^2$. Let
$(\varepsilon_1,\varepsilon_2)$ be a pair of random variables sampled from this
distribution. Because of the symmetry of the function $f$, the random variables
$\varepsilon_1$ and $\varepsilon_2$ are then exchangeable, and their common PDF
is
$$
f_{\varepsilon}(t) = 2\int_{\mathbb R}h(t - e_2)r(t+e_2)\dif e_2 = 2\int_{\mathbb R} h(u)r(u+2t)\dif u,\quad t\in\mathbb R,
$$
which implies that $0<\varepsilon_1<2$ and $0<\varepsilon_2<2$ with probability one. Moreover, the PDF of the difference $\varepsilon_1 - \varepsilon_2$ is
\begin{align*}
f_{\varepsilon_1 - \varepsilon_2}(t)
& = \int_{\mathbb R}f(e_2 + t,e_2)\dif e_2
= 2\int_{\mathbb R}h(t)r(2e_2 + t)\dif e_2 = \int_{\mathbb R}h(t)r(s)\dif s = h(t),\quad t\in\mathbb R.
\end{align*}
It is then straightforward to check that Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:RankRegressors-TLAD} are all
satisfied. Note, however, that Assumption \ref{assu:extra continuity} is {\em
not} satisfied because the function $f$ is not continuous.
Now, observe that since $\mathbb P(Y_1>0, Y_2>0) = 1$ by construction, a
calculation shows that the expected loss function $L$ in
\eqref{eq:PopulationLossFunction-TLAD} simplifies to
$$
L(\theta) = \mathrm{E}\left[m^{\texttt{tlad}}(\theta,\boldsymbol{Y}) - m^{\texttt{tlad}}(0,\boldsymbol{Y})\right] = \mathrm{E}\left[|Y_1 - Y_2 - \theta| - |Y_1 - Y_2|\right],\quad \theta\in\mathbb R.
$$
Since the integrand here is convex in $\theta$ and non-differentiable in
$\theta$ only at $Y_1 - Y_2$, which, for a given $\theta$, happens with
probability zero, it follows from \citet[Proposition
2.3]{bertsekas1973stochastic} that $L$ is (everywhere) differentiable with
derivative
$$
\dot{L}(\theta) = \mathrm{E}\left[2\cdot\mathbf{1}\{Y_1 - Y_2 \leqslant \theta\} - 1\right] = 2F_{\varepsilon_1 - \varepsilon_2}(\theta) - 1,
$$
where $F_{\varepsilon_1 - \varepsilon_2}$ denotes the CDF of the difference
$\varepsilon_1 - \varepsilon_2$. We claim that $F_{\varepsilon_1 -
\varepsilon_2}$ is non-differentiable at zero $(=\theta_0)$. Indeed, if $t_k =
2^{-k}$ for $k\in\mathscr E$, then
$$
\frac{F_{\varepsilon_1 - \varepsilon_2}(t_k) - F_{\varepsilon_1 - \varepsilon_2}(0)}{t_k} = 2^k \int_0^{2^{-k}} h(t)\dif t = \frac{3}{4}\cdot 2^k \cdot \sum_{l\in\mathscr E : l\geqslant k}\frac{1}{2^{l+1}} = \frac{1}{2}
$$
and if $t_k = 2^{-k}$ for $k\in\mathscr O$, then
$$
\frac{F_{\varepsilon_1 - \varepsilon_2}(t_k) - F_{\varepsilon_1 - \varepsilon_2}(0)}{t_k} = 2^k \int_0^{2^{-k}} h(t)\dif t = \frac{3}{4}\cdot 2^k \cdot \sum_{l\in\mathscr E : l\geqslant k+1}\frac{1}{2^{l+1}} = \frac{1}{4},
$$
which implies that the limit of $(F_{\varepsilon_1 - \varepsilon_2}(t) -
F_{\varepsilon_1 - \varepsilon_2}(0))/t$ as $t\to 0$ does not exist. Thus, the
function $\dot{L}$ is non-differentiable at zero ($=\theta_0$), implying that
the Hessian of $L$ at $\theta_0$ does not exist.
One can show that the TLAD estimator is still $\sqrt{n}$-consistent but not
(asymptotically) normal. Specifically, in Figure
\ref{fig:ECDF_Scaled_Estimates}, we show the (kernel) PDF of the scaled TLAD
estimates $\sqrt{n}\widehat{\theta}^{\tt{tlad}}$ based on $10,000$ Monte Carlo
samples of size $n=50,000$ using the DGP defined earlier in this
section.\footnote{The kernel density was created using the \texttt{R} package
\texttt{ggplot2} with \texttt{geom\_density}. We use a Gaussian kernel and the
\citet[Equation (3.31)]{silverman_density_1986} rule-of-thumb bandwidth (both
\texttt{geom\_density }defaults).} To facilitate comparison, we include the PDF
resulting from fitting a Gaussian distribution to the Monte Carlo dataset of
scaled estimates using maximum likelihood. The PDF of the scaled estimates looks
far from Gaussian. We therefore conclude that establishing the asymptotic
normality result for the TLAD estimator requires more than Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:RankRegressors-TLAD}. We do so by
imposing the continuity conditions in Assumption \ref{assu:extra continuity}.
\begin{figure}\caption{\normalsize{PDFs of Scaled TLAD Estimates (Solid) and Best Normal
Approximation (Dashed)}}
\centering
\includegraphics[width=\textwidth, trim={0 0 0 0.75cm}, clip]{kernel_PDF_Scaled_Estimates_n50000_50000.png}
\label{fig:ECDF_Scaled_Estimates}
\end{figure}
\subsection{Extended Asymptotic Normality
Result}\label{sec:HessianExistence-TLAD} In this subsection, we provide a
Hessian existence argument for the TLAD estimator using Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:extra continuity}. We also show that,
unless the pair of latent outcomes $(Y_1^\ast,Y_2^\ast)$ belongs to the first
quadrant with conditional probability given $\boldsymbol{W}$ equal to one almost surely,
$\boldsymbol{\Gamma}_0^{\tt{tlad}}$ will \emph{not} be the Hessian of the expected loss
(because of an apparent typographical error).
We then state the asymptotic normality result under Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:extra continuity} with the corrected Hessian.
To state the existence theorem, let $(\boldsymbol{w},\boldsymbol{y})\mapsto
f_{\boldsymbol{Y}^\ast\mid\boldsymbol{w}}\left(\boldsymbol{y}\right)$ denote the joint PDF of the latent outcomes
$Y_1^\ast$ and $Y_2^\ast$ conditional on $\boldsymbol{W}=\boldsymbol{w}$ for all $\boldsymbol{w}\in\mathcal{W}$.
\begin{thm}[\textbf{Trimmed Least Absolute Deviations Hessian
Existence}]\label{thm:LossHessianExistence-TLAD} Under Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:extra continuity}, the function
$L:\mathbb{R}^K\to\mathbb{R}$ defined by \eqref{eq:PopulationLossFunction-TLAD} is twice
differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$. Its Hessian matrix $\mathbf{H}_0^{\tt{tlad}} :=
\nabla^2L(\boldsymbol{\theta}_0)$ is given by
\begin{equation}\label{eq:LossHessian-TLAD}
\left.\begin{aligned}
\mathbf{H}_0^{\tt{tlad}}
&=\mathrm{E}\Bigg[\bigg(2\int_{0}^{+\infty}f_{\boldsymbol{Y}^{\ast}\mid\boldsymbol{W}}\del[1]{z+\max\cbr[0]{0,\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0},z-\min\cbr[0]{0,\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0}}\dif z\\
& \quad\qquad+\boldsymbol{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\geqslant0\}\int_{-\infty}^{0}f_{\boldsymbol{Y}^{\ast}\mid\boldsymbol{W}}\del[1]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0,z}\dif z\\
& \quad\qquad+\boldsymbol{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<0\}\int_{-\infty}^{0}f_{\boldsymbol{Y}^{\ast}\mid\boldsymbol{W}}\del[1]{z,-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0}\dif z\bigg)\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\Bigg].
\end{aligned}\right\}
\end{equation}
\end{thm}
\begin{rem}[\textbf{Comparison with \citet{honore_trimmed_1992}}] Comparing
$\mathbf{H}_0^{\tt{tlad}}$ in \eqref{eq:LossHessian-TLAD} with the matrix
$\boldsymbol{\Gamma}_0^{\tt{tlad}}$ in \eqref{eq:VarianceSandwichBread-TLAD}, and writing
out the definitions of the conditional PDFs involved, we see that the latter two
contributions to each expression are equal. (To align the two expressions, we
here interpret ``undefined'' times ``zero'' as ``zero.'') However, comparing the
\emph{first} terms of each matrix, we see that the conditional PDF
$f_{Y_1^\ast-Y_2^\ast\mid\boldsymbol{W},Y_1^\ast>0,Y_2^\ast>0}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0)$ in
Honor{\'e}'s $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ is missing multiplication by the
conditional probability $\mathbb{P}\left(Y_1^\ast>0, Y_2^\ast>0\middle|\boldsymbol{W}\right)$.
Hence, unless the pair of latent outcomes $(Y_1^\ast,Y_2^\ast)$ belongs to the
first quadrant with conditional probability given $\boldsymbol{W}$ equal to one almost
surely, $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ will \emph{not} be the Hessian of the expected
loss. Of course, if $(Y_1^\ast,Y_2^\ast)$ resides in the first quadrant with
probability one, then the model involves no censoring.\hfill$\diamondsuit$
\end{rem}
We now revise the asymptotic normality statement in
\eqref{eq:AsymptoticNormality-TLAD} for the TLAD estimator based on our new
understanding of the Hessian $\mathbf{H}_0^{\tt{tlad}}$.
\begin{thm}[\textbf{Asymptotic Normality of Trimmed Least Absolute
Deviations}]\label{thm:AsymptoticNormality-TLAD} Let Assumptions
\ref{assu:Non-Degeneracy-TLAD}--\ref{assu:extra continuity} hold, and suppose
that the expectations involved in defining the matrix $\mathbf{V}_0^{\tt{tlad}}$ exist
(in $\mathbb{R}^{K\times K}$), and that both matrices $\mathbf{V}_0^{\tt{tlad}}$ and
$\mathbf{H}_0^{\tt{tlad}}$ are of full rank. Then
\begin{equation}\label{eq:AsymptoticNormality-TLAD-2}
\sqrt{n}\del[1]{\widehat\boldsymbol{\theta}^{\tt{tlad}}-\boldsymbol{\theta}_0}\rightsquigarrow\mathcal{N}\left(\mathbf{0},(\mathbf{H}_0^{\tt{tlad}})^{-1}\mathbf{V}_0^{\tt{tlad}}(\mathbf{H}_0^{\tt{tlad}})^{-1}\right)\;\text{in}\;\mathbb{R}^K.
\end{equation}
\end{thm}
\begin{rem}[\textbf{Comparison with \citet{honore_pairwise_1994}}] We return to
the comparison with \citet{honore_pairwise_1994}, who considered censored
regression with cross-sectional data as discussed in Remark
\ref{rem:ComparisonWithHonorePowellJacobian}. On \emph{ibid.}~(p.~260), the
authors gave an expression for the Hessian of the expected loss at the true
parameter value $\boldsymbol{\theta}_0$ when using the trimmed absolute loss. Their
expression is (in our notation)
\begin{align*}
&\mathrm{E}\Bigg[\bigg(2\int_{0}^{+\infty}f_{\varepsilon}\del[1]{z-\min\cbr[1]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}^2\dif z\\
&\qquad+f_{\varepsilon}\del[1]{-\min\cbr[1]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}F_{\varepsilon}\del[1]{-\min\cbr[1]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}\bigg)
\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\Bigg],
\end{align*}
with the pairs $(Y_1,\boldsymbol{X}_1)$ and $(Y_2,\boldsymbol{X}_2)$ representing independent units,
where $f_{\varepsilon}$ and $F_{\varepsilon}$ denote the PDF of $\varepsilon$
and the CDF of $\varepsilon$, respectively. Upon setting $\alpha\equiv0$ in our
panel-data setting, and taking the model errors $(\varepsilon_1,\varepsilon_2)$
to be independent and identically distributed, the conditional PDF of
$(Y_1^\ast,Y_2^\ast)$ given $\boldsymbol{W}=\boldsymbol{w}$ factors as
$
f_{Y_1^\ast,Y_2^\ast\mid\boldsymbol{w}}\left(y_1^\ast,y_2^\ast\right)=f_{\varepsilon}\del[1]{y_1^\ast-\boldsymbol{x}_1^\top\boldsymbol{\theta}_0}f_{\varepsilon}\del[1]{y_2^\ast-\boldsymbol{x}_2^\top\boldsymbol{\theta}_0}.
$
A straightforward calculation then shows that our panel-data Hessian in
\eqref{eq:LossHessian-TLAD} reduces to the cross-sectional Hessian from
\citet{honore_pairwise_1994} in the previous display.\hfill$\diamondsuit$
\end{rem}
\subsection{Asymptotic Variance Estimation}\label{sec:LADAsymptoticNormalityAndVarianceEstimation}
We now discuss estimation of the asymptotic variance
$\boldsymbol{\Sigma}_0^{\tt{tlad}}:=(\mathbf{H}_0^{\tt{tlad}})^{-1}\mathbf{V}_0^{\tt{tlad}}(\mathbf{H}_0^{\tt{tlad}})^{-1}$
in \eqref{eq:AsymptoticNormality-TLAD-2}. Since the formula for
$\mathbf{H}_0^{\tt{tlad}}$ given in \eqref{eq:LossHessian-TLAD} includes a
nonparametric conditional PDF, we propose a bootstrap estimator instead of a
plug-in estimator. Also, since a \emph{na{\"i}ve} bootstrap variance estimator
may fail to be consistent for LAD-type estimators without additional
integrability conditions \citep{BF81,GPSB84}, we instead construct a
quantile-based robust bootstrap variance estimator that avoids reliance on
bootstrap second moments.
The \emph{robust} bootstrap variance estimator exploits the fact that quantiles
of linear combinations of a normal distribution scale with their standard
deviations. We therefore estimate bootstrap quantiles of suitably centered
parameter combinations and invert this scaling to recover the corresponding
covariance entries. This construction replaces unstable bootstrap variance
estimation based on second moments by identification of covariance entries
through the asymptotic normal geometry of the estimator.
To describe the robust bootstrap variance estimator in more detail, consider a
bootstrap sample
$\{(\widetilde{Y}_{i1},\widetilde{\boldsymbol{X}}_{i1},\widetilde{Y}_{i2},\widetilde{\boldsymbol{X}}_{i2})\}_{i=1}^n$.
obtained by sampling $n$ observation indices with replacement from
$\{1,2,\dots,n\}$ and extracting the corresponding observations from the
original sample. Also, let
\begin{equation}\label{eq: bootstrap}
\widetilde{\boldsymbol{\theta}}^{\tt{tlad}} := (\widetilde{\theta}^{\tt{tlad}}_1,\dots,\widetilde{\theta}^{\tt{tlad}}_K)^\top \in \operatornamewithlimits{argmin}\limits_{\boldsymbol{\theta}\in\mathbb{R}^K}\cbr[3]{\frac{1}{n}\sum_{i=1}^n m^{\tt{tlad}}\big((\Delta\widetilde{\boldsymbol{X}}_i)^\top\boldsymbol{\theta},\widetilde{\boldsymbol{Y}}_i\big)}
\end{equation}
be the bootstrap analog of $\widehat\boldsymbol{\theta}^{\tt{tlad}} =
(\widehat\theta^{\tt{tlad}}_1,\dots,\widehat\theta^{\tt{tlad}}_K)^\top$, where
we introduced shorthands $\Delta\widetilde{\boldsymbol{X}}_i := \widetilde{\boldsymbol{X}}_{i1} - \widetilde{\boldsymbol{X}}_{i2}$ and
$\widetilde{\boldsymbol{Y}}_i:=(\widetilde{Y}_{i1},\widetilde{Y}_{i2})$. Identification of the
covariance entries relies on the identities $\mathrm{var}(Z_k)=\Sigma^{\tt{tlad}}_{0kk}$
and $\mathrm{var}(Z_j+Z_k)=\Sigma^{\tt{tlad}}_{0jj} +\Sigma^{\tt{tlad}}_{0kk}
+2\Sigma^{\tt{tlad}}_{0jk}$ for a normal vector $\boldsymbol{Z}
\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{\Sigma}_0^{\tt{tlad}})$. Accordingly, we recover variances and
covariances from bootstrap quantiles of the corresponding linear combinations.
Motivated by these identities, we estimate the diagonal and off-diagonal entries
of $\boldsymbol{\Sigma}_0^{\tt{tlad}}=[\Sigma^{\tt{tlad}}_{0jk}]_{j,k=1}^K$ separately. For
the \emph{diagonal} elements, for each $k\in[K]$, we set
$\widehat\Sigma_{kk}^{\tt{tlad}}:=[\widehat{q}_{0.9,k}/\Phi^{-1}(0.95)]^2$,
where $\widehat{q}_{0.9,k}$ is the $0.9$ quantile of the conditional
distribution of $\sqrt n|\widetilde\theta^{\tt{tlad}}_k -
\widehat\theta^{\tt{tlad}}_k|$ given the original data and $\Phi^{-1}(0.95)$ is
the $0.95$ quantile of the standard normal distribution. For the
\emph{off-diagonal} elements, for each $(j,k)\in[K]\times[K]$ such that $j\neq
k$, we set $\widehat \Sigma_{jk}^{\tt{tlad}}:=
([\widehat{q}_{0.9,j,k}/\Phi^{-1}(0.95)]^2 - \widehat\Sigma_{jj}^{\tt{tlad}} -
\widehat\Sigma_{kk}^{\tt{tlad}})/2$, where $\widehat{q}_{0.9,j,k}$ is the $0.9$
quantile of the conditional distribution of $\sqrt
n|\widetilde\theta^{\tt{tlad}}_j + \widetilde\theta^{\tt{tlad}}_k -
\widehat\theta^{\tt{tlad}}_j - \widehat\theta^{\tt{tlad}}_k|$ given the original
data.\footnote{As the robust bootstrap covariance estimator
$\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ is recovered entrywise from quantiles of linear
combinations, it need not be positive semidefinite in finite samples. One can
construct a positive semidefinite variance estimator by projecting
$\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ onto the set of positive semidefinite matrices.}
In the following theorem, we prove consistency of
$\widehat\boldsymbol{\Sigma}^{\tt{tlad}}:=[\widehat\Sigma_{jk}^{\tt{tlad}}]_{j,k=1}^K$ under
the assumptions of Theorem \ref{thm:AsymptoticNormality-TLAD}.
\begin{thm}[\textbf{Consistency of the Robust Bootstrap Variance Estimator for
TLAD}]\label{thm: bootstrap consistency lad} Let the assumptions of Theorem
\ref{thm:AsymptoticNormality-TLAD} hold. Then
$\widehat\boldsymbol{\Sigma}^{\tt{tlad}}\to_{\mathbb{P}}\boldsymbol{\Sigma}_0^{\tt{tlad}}$.
\end{thm}
\begin{rem}[\textbf{Reducing Computational Burden via \citet{HH17}}]
Note that calculating the quantiles $\widehat{q}_{0.9,k}$ and
$\widehat{q}_{0.9,j,k}$ requires solving the $K$-dimensional optimization
problem \eqref{eq: bootstrap} for many draws of the bootstrap sample. To reduce
computational burden, one may instead employ the alternative bootstrap procedure
of \citet{HH17}, which replaces repeated $K$-dimensional optimization by a
sequence of one-dimensional problems. This approach is particularly attractive
in our setting because the ``meat'' matrix $\mathbf{V}_0^{\tt{tlad}}$ in
\eqref{eq:VarianceSandwichMeat-TLAD} admits a consistent plug-in estimator. We
refer the reader to the original paper for further details.\hfill$\diamondsuit$
\end{rem}
\begin{rem}[\textbf{Bootstrap Inference for Convex Pairwise Differencing Estimators}]
A related recent paper is \citet{cattaneo2025robust}, which develops
distribution theory and bootstrap-based inference for a broad class of convex
pairwise differencing estimators. Their framework includes trimmed absolute loss
as a special case. Our analysis in Theorem \ref{thm: bootstrap consistency lad}
is complementary in that it considers a panel setting, where the trimmed loss is
used to eliminate an individual-specific fixed effect and no bandwidth choice is
required. In addition, while \citet{cattaneo2025robust} establish bootstrap
validity for the distributional approximation, Theorem \ref{thm: bootstrap
consistency lad} establishes consistency of a robust bootstrap variance
estimator for the asymptotic variance of the TLAD estimator under the
assumptions of Theorem \ref{thm:AsymptoticNormality-TLAD}.\hfill$\diamondsuit$
\end{rem}
\begin{rem}[\textbf{Conservative Inference}]
The robust bootstrap variance estimator $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ is
consistent for $\boldsymbol{\Sigma}_0^{\tt{tlad}}$ under the assumptions of Theorem
\ref{thm:AsymptoticNormality-TLAD}. However, the construction of
$\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ is based on a particular choice of quantile level
(i.e., $0.9$ for the absolute differences, corresponding to $0.95$ for the
standard normal quantiles). This choice is not essential: the consistency
argument in the proof of Theorem \ref{thm: bootstrap consistency lad} continues
to apply for any fixed interior quantile. We use the particular quantile level
above as a convenient way to avoid reliance on bootstrap second moments, which
may be unstable for LAD-type estimators without additional integrability
conditions.
For comparison, \citet{hahn2021bootstrap} show that, under their conditions,
bootstrap variance estimators based on second moments, although not necessarily
consistent in LAD-type settings, can still be asymptotically conservative for
inference on fixed linear combinations. Thus, if one is willing to forgo
consistency of the variance estimator itself and instead target conservative
inference for linear combinations of the parameter vector, the na{\"i}ve
bootstrap provides an alternative route. This conservative-inference perspective
is distinct from the goal pursued here, namely consistent estimation of the full
asymptotic covariance matrix.\hfill$\diamondsuit$
\end{rem}
\section*{Funding}
Honor{\'e} and S{\o}rensen are grateful for research support from the Aarhus
Center for Econometrics (ACE) funded by the Danish National Research Foundation
grant number DNRF186. Honor{\'e} also acknowledges support from the Gregory
C.~Chow Econometric Research Program at Princeton University.
\begin{singlespace}
\bibliographystyle{ecta}
\bibliography{bo}
\end{singlespace}