EconBase
← Back to paper

Asymptotic Variance Theory for Trimmed Least Squares and Trimmed Least Absolute Deviations in Censored Panel Models with Fixed Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

84,986 characters · 15 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Asymptotic Variance Theory for Trimmed Least Squares and Trimmed Least Absolute Deviations in Censored Panel Models with Fixed Effects

\affil[1]{{University of California, Los Angeles}} \affil[2]{{University of Copenhagen}} \affil[3]{{Princeton University}} \affil[4]{{Aarhus Center for Econometrics (ACE)}}

abstractWe study inference using trimmed least squares (TLS) and trimmed least absolute deviations (TLAD) estimators of honore_trimmed_1992 in censored two-period panel-data models with fixed effects. We show that the published asymptotic variance formulas rely on additional regularity conditions that are not fully stated in the original analysis. For TLS, the published Hessian formula requires that the regressor-difference index vanish only when the regressor difference itself is zero, a restriction not explicitly stated in the original paper and violated, for instance, with a zero parameter vector. We derive the correct Hessian, establish asymptotic normality without imposing this restriction, and obtain a consistent plug-in variance estimator. We also show that the Hessian estimator proposed in honore_trimmed_1992 {\em is} actually consistent for the {\em correct} TLS asymptotic variance. For TLAD, we show that the published variance formula omits a conditional-probability term and that asymptotic normality requires additional continuity conditions. Under these conditions, we derive the corrected asymptotic variance and provide a tuning-parameter-free bootstrap variance estimator.

Introduction

Inference for censored panel-data models with fixed effects often relies on trimmed estimators whose objective functions are convex but nonsmooth. honore_trimmed_1992 introduced two influential examples, the trimmed least squares (TLS) and trimmed least absolute deviations (TLAD) estimators, and derived their asymptotic distributions for a two-period censored regression model with fixed effects. These estimators remain natural benchmark procedures in this setting, both because they avoid parametric restrictions on the fixed effects and because they provide tractable ways to exploit the structure of two-period panel data under censoring. Valid inference for such estimators, however, depends delicately on the form of the Hessian of the expected loss, and the nonsmoothness created by trimming and censoring makes that Hessian less straightforward than the published formulas suggest.

This paper revisits the asymptotic variance theory for the TLS and TLAD estimators of honore_trimmed_1992. We show that the published asymptotic variance formulas rely on additional regularity conditions that are not fully reflected in the original theorem statements. For TLS, the problem is that censoring creates boundary points at which the score of the trimmed square loss is nondifferentiable with positive probability. At such points, differentiation under expectation need not be valid, so the Hessian of the expected loss can differ from the expression obtained by differentiating inside the expectation. For TLAD, the problem is that the expected loss can fail to be twice differentiable when the conditional distribution of the model errors is not sufficiently smooth. Our results identify how these issues enter the asymptotic variance analysis for TLS and TLAD, derive the corrected variance formulas, and clarify the scope of valid legacy inference.

Our starting point is the two-period censored regression model in which the outcome variable $Y_\tau$ in period $\tau\in\{1,2\}$ is generated as

equation[equation omitted — 153 chars of source]

where $\boldsymbol{X}_\tau\in\mathbb{R}^K$ is a vector of time-varying regressors, $\boldsymbol{\theta}_0\in \mathbb{R}^K$ is the parameter of interest, $\alpha$ is an unobserved fixed effect, and $\varepsilon_\tau$ is an unobserved error term. Let $\{(Y_{i1},\boldsymbol{X}_{i1},Y_{i2},\boldsymbol{X}_{i2})\}_{i=1}^n$ be a random sample from the distribution of $(Y_1,\boldsymbol{X}_1,Y_2,\boldsymbol{X}_2)$. honore_trimmed_1992 proposed M-estimators of $\boldsymbol{\theta}_0$ based on a trimmed loss function $m_\Xi:\mathbb{R}\times[0,\infty)\times[0,\infty)\to\mathbb{R}$ defined by

equation[equation omitted — 287 chars of source]

where $\Xi(\cdot)$ represents either the absolute loss $\left|\cdot\right|$ or the (one-half) square loss $\textstyle\frac{1}{2}(\cdot)^2$,\footnote{honore_pairwise_1994 provide a list of conditions allowing other choices of $\Xi$.} and $\xi(\cdot)$ is its derivative.\footnote{In the case of the absolute loss $\Xi(\cdot)=\left|\cdot\right|$, we set $\xi(0) = 0$.} These choices yield the TLAD and TLS estimators, respectively, defined for the corresponding $\Xi(\cdot)$ as any solution

align[align omitted — 272 chars of source]

where we introduced the shorthands $\Delta\boldsymbol{X}_i := \boldsymbol{X}_{i1} - \boldsymbol{X}_{i2}$ and $\boldsymbol{Y}_i:=(Y_{i1},Y_{i2})$. Under conditional exchangeability of $(\varepsilon_1,\varepsilon_2)$ given $(\boldsymbol{X}_1,\boldsymbol{X}_2,\alpha)$ and other regularity conditions, honore_trimmed_1992 established consistency and asymptotic normality of both estimators.

Our first set of results concerns the TLS estimator. We show that the Hessian formula appearing in honore_trimmed_1992 is valid only if

equation[equation omitted — 164 chars of source]

a condition not explicitly stated in the original paper. This condition is satisfied, for example, when the index $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ has a continuous distribution, but it fails when all components of $\boldsymbol{\theta}_0$ are equal to zero unless $\Delta\boldsymbol{X} = \boldsymbol{0}$ almost surely. When condition (ref) is not satisfied, the TLS score is nondifferentiable with positive probability, and the published Hessian formula need not equal the derivative of the expected score. We derive the correct Hessian, provide an explicit representation of it, establish asymptotic normality without excluding mass at zero, and obtain a consistent plug-in estimator of the corresponding asymptotic variance.

The TLS analysis also yields a second result. Although the population Hessian formula in honore_trimmed_1992 is not generally correct, the Hessian estimator proposed there remains consistent for the correct TLS asymptotic variance under the maintained assumptions. Thus, the paper does not only identify where the published population formula fails; it also clarifies which parts of the original inference procedure remain usable.

Our second set of results concerns the TLAD estimator. We show, first, that the published asymptotic variance formula omits a conditional-probability term. More importantly, we show that asymptotic normality requires additional continuity conditions to ensure existence of the Hessian of the expected loss. We provide a counterexample satisfying the conditions stated in honore_trimmed_1992 for which that Hessian does not exist. Under additional continuity conditions, we derive the corrected TLAD Hessian and corresponding asymptotic variance formula. We then provide a tuning-parameter-free bootstrap variance estimator for TLAD.

Taken together, these results clarify the role of boundary nondifferentiability in the asymptotic variance theory for censored trimmed estimators. More specifically, they show when legacy inference remains valid, when corrected variance formulas are required, and how to implement inference under the corrected theory.

As a by-product, our TLS analysis also yields a plug-in perspective on the asymptotic variance of the cross-sectional trimmed least squares estimator studied by honore_pairwise_1994, thus avoiding the choice of tuning parameters required by the numerical-differentiation-based variance estimator proposed there.

The remainder of the paper is organized as follows. Section (ref) revisits the TLS estimator, beginning with the asymptotic normality result in honore_trimmed_1992, then identifying the source of the problem, deriving the correct Hessian, establishing asymptotic normality under the corrected theory, and discussing variance estimation, including the status of the legacy Hessian estimator. Section (ref) carries out the analogous analysis for TLAD, derives the corrected Hessian and asymptotic variance formula, and develops bootstrap-based variance estimation. Proofs for the TLS results are in the appendix, while proofs for the TLAD results are in the supplemental appendix, which also contains technical lemmas establishing a generalized dominated convergence theorem, the existence of conditional PDFs, and related measurability issues.

\paragraph{Notation.} For a positive integer $k$, we write $\left[k\right]:=\{1,\ldots,k\}$. We use $\left\|\cdot\right\|_2$ to denote the Euclidean norm. For any (suitably differentiable) real-valued function $f$, whose first argument is a scalar, we let $\dot{f}_1$ and $\ddot{f}_{11}$ denote the first- and second-order partial derivatives of $f$ with respect to its first argument, respectively. For a possibly vector-valued differentiable function $\boldsymbol{g}$, we let $\nabla \boldsymbol{g}$ denote the Jacobian of $\boldsymbol{g}$. If $\boldsymbol{g}$ is itself the gradient mapping corresponding to some function $f$, then we write $\nabla^2 f$ for the Hessian of $f$. The indicator $\mathbf{1}\{A\}$ equals one (zero) if the logical statement $A$ is true (false). We use comma-separated events to denote the intersection of events, e.g. $\mathbf 1\{A,B\} = \mathbf 1\{A\cap B\}$. Unless otherwise stated, all limits are understood as the sample size $n$ grows without bound, holding everything else fixed. The symbols $\rightsquigarrow$, $\to_{\text{a.s.}}$, and $\to_{\mathbb{P}}$ denote convergences in distribution, almost surely, and in probability, respectively. We write $o_{\mathbb{P}}(1)$, $o_{\mathrm{a.s.}}(1)$, and $o_{L^1}(1)$ to denote sequences of random variables that converge to zero in probability, almost surely, and in $L^1(\mathbb{P})$, respectively.

Trimmed Least Squares

In this section, we focus on the TLS estimator. We abbreviate the trimmed square loss, $m_{\Xi}$ in (ref) with $\Xi=\textstyle{\frac{1}{2}}(\cdot)^2$, by $m^{\texttt{tls}}$, which takes the explicit form

equation[equation omitted — 290 chars of source]

This loss is continuously differentiable in its first argument with partial derivative given by

equation[equation omitted — 219 chars of source]

To facilitate our presentation below, note that it follows from (ref) that $\dot{m}^{\texttt{tls}}_1(\cdot,\boldsymbol{y})$ is Lipschitz continuous on $\mathbb{R}$ with Lipschitz constant one, uniformly in $\boldsymbol{y}\in[0,\infty)\times[0,\infty)$. However, it fails to be differentiable at $-y_2$ and $y_1$ (unless both $y_1$ and $y_2$ are zero), implying that the trimmed square loss is not necessarily twice differentiable. Specifically, the points of second-order non-differentiability of $m^{\texttt{tls}}(\cdot,\boldsymbol{y})$ are captured by set-valued function $N:[0,\infty)\times[0,\infty)\rightrightarrows\mathbb{R}$, defined by

equation[equation omitted — 267 chars of source]

Assumptions for Trimmed Least Squares

For notational convenience, abbreviate the TLS estimator $\widehat\boldsymbol{\theta}_\Xi$ in (ref) with $\Xi=\textstyle{\frac{1}{2}}\left(\cdot\right)^2$ by $\widehat{\boldsymbol{\theta}}^{\texttt{tls}}$. Also, denote $\boldsymbol{W}:=(\boldsymbol{X}_1,\boldsymbol{X}_2,\alpha)$ and let $\mathcal{W}\subseteq\mathbb{R}^{K}\times \mathbb R^K\times\mathbb R$ be the support of $\boldsymbol{W}$, with $\boldsymbol{w} := (\boldsymbol{x}_1,\boldsymbol{x}_2,a)$ standing for a generic element of $\mathcal{W}$. In addition, denote $\boldsymbol{\varepsilon}:=(\varepsilon_1,\varepsilon_2)$ and let $\mathcal E := \mathbb R\times\mathbb R$, with $\boldsymbol{e} = (e_1,e_2)$ standing for a generic element of $\mathcal E$. honore_trimmed_1992 derived the asymptotic distribution of the TLS estimator under the following assumptions.\footnote{Assumptions (ref)--(ref) are from Assumptions S.2, M.3, E.1, E.3 and R.1 , respectively, in honore_trimmed_1992.}

assumption[Non-Degeneracy] The probability $\mathbb{P}(Y_1>0, Y_2>0)$ is strictly positive.
assumption[Integrability] All of the following expectations are finite: \[ \mathrm{E}[\|\boldsymbol{X}_1\|_2^4],\;\mathrm{E}[\|\boldsymbol{X}_2\|_2^4],\;\mathrm{E}[\alpha^2\|\Delta\boldsymbol{X}\|_2^4],\; \mathrm{E}[\varepsilon_1^2\|\Delta\boldsymbol{X}\|_2^4],\;\text{and}\; \mathrm{E}[\varepsilon_2^2\|\Delta\boldsymbol{X}\|_2^4]. \]
assumption[Continuity] The conditional distribution of $(\varepsilon_1,\varepsilon_2)$ given $\boldsymbol{W}$ is absolutely continuous with respect to the Lebesgue measure.
assumption[Exchangeability] Conditional on $\boldsymbol{W}$, $\varepsilon_1$ and $\varepsilon_2$ are exchangeable.
assumption[Rank of Regressors] There is no proper linear subspace of $\mathbb{R}^K$ containing the random variable $\mathbf{1}\{\mathbb{P}\left(Y_1>0, Y_2>0\middle|\boldsymbol{X}_1,\boldsymbol{X}_2\right)>0\}\Delta\boldsymbol{X}$ with probability one.

Let $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$, mapping $\mathcal W\times\mathcal E$ to $[0,\infty)$, be a version of the joint probability density function (PDF) of the pair $\boldsymbol{\varepsilon} = (\varepsilon_1,\varepsilon_2)$ conditional on $\boldsymbol{W} = \boldsymbol{w}$ that is measurable in $(\boldsymbol{w},\boldsymbol{e})$ and is such that $f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_1,e_2) = f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_2,e_1)$ for all $\boldsymbol{w}\in\mathcal W$ and $\boldsymbol{e} = (e_1,e_2)\in\mathcal E$. As we show in Appendix (ref), the function $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$ does exist under Assumptions (ref) and (ref). Also, let $(\boldsymbol{w},\boldsymbol{e})\mapsto F_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e}) = \int_{-\infty}^{e_1}\int_{-\infty}^{e_2}f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(u_1,u_2)\dif u_2\dif u_1$ be the corresponding version of the joint cumulative distribution function (CDF) of $(\varepsilon_1,\varepsilon_2)$ conditional on $\boldsymbol{W} = \boldsymbol{w}$, and let $(\boldsymbol{w},e)\mapsto F_{\varepsilon\mid\boldsymbol{w}}(e) = \lim_{u\to\infty} F_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e,u)$ be the corresponding version of the common marginal CDF of $\varepsilon_1$ and $\varepsilon_2$ conditional on $\boldsymbol{W} = \boldsymbol{w}$.

Asymptotic Normality Result in Honor{\'e} (1992)

To state the TLS normality result in honore_trimmed_1992, introduce the $K\times K$ matrices \[ \mathbf{V}_0^{\tt{tls}}:=\mathrm{E}\left[\dot{m}_1^{\tt{tls}}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0,\boldsymbol{Y})^2\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\right] \] and

equation[equation omitted — 330 chars of source]

honore_trimmed_1992 states that if Assumptions (ref)--(ref) hold, the expectations involved in defining the matrices $\mathbf{V}_0^{\tt{tls}}$ and $\boldsymbol{\Gamma}_0^{\tt{tls}}$ exist (in $\mathbb{R}^{K\times K}$), and both matrices are of full rank, then

equation[equation omitted — 315 chars of source]

with $\mathcal{N}(\boldsymbol{\mu},\boldsymbol{\Sigma})$ denoting the normal distribution with mean $\boldsymbol{\mu}$ and covariance matrix $\boldsymbol{\Sigma}$.

Among several steps, the proof of this result in honore_trimmed_1992 involves:

enumerate• Arguing that the gradient of the expected loss $\boldsymbol{\theta}\mapsto\mathrm{E}[m^{\tt{tls}}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta},\boldsymbol{Y})]$ is equal to the expected value of the model score [see (ref) below]. • Establishing differentiability of the expected value of the model score at the true parameter $\boldsymbol{\theta}_0$.

In our notation, the second task amounts to arguing that the function $\boldsymbol{G}:\mathbb{R}^K\to\mathbb{R}^K$ defined by

equation[equation omitted — 408 chars of source]

is differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$. Its Jacobian $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)\in\mathbb{R}^{K\times K}$ is then the Hessian of the expected loss.\footnote{In honore_trimmed_1992, the function $\boldsymbol{G}$ is denoted $G_4$.} To carry out this task, honore_trimmed_1992 invoked the LDCT to interchange the order of differentiation and expectation, and the matrix $\boldsymbol{\Gamma}_0^{\tt{tls}}$ is the result of that calculation. In the next subsection, however, we will show by example that $\boldsymbol{\Gamma}_0^{\tt{tls}}$ is not always equal to the Jacobian of $\boldsymbol{G}$ at $\boldsymbol{\theta}_0$.

Counterexample

We here provide a DGP which satisfies Assumptions (ref)--(ref), for which the function $\boldsymbol{G}$ is differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$, but the resulting Jacobian $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$ differs from $\boldsymbol{\Gamma}_0^{\tt{tls}}$. To this end, let $K=1$, $\alpha\equiv0$, $\boldsymbol{\theta}_0=0$, $\boldsymbol{X}_1\equiv2$ and $\boldsymbol{X}_2\equiv1$, and let $\varepsilon_1$ and $\varepsilon_2$ be independent of each other and standard normally distributed. Then the outcome variables $Y_{\tau}=\max\{0,\varepsilon_{\tau}\},\tau\in\{1,2\}$, are independent and identically distributed as standard normals censored from below at zero. Case by case inspection shows that Assumptions (ref)--(ref) are satisfied and that (the here scalar) $\mathbf{V}_0^{\tt{tls}}$ and $\boldsymbol{\Gamma}_0^{\tt{tls}}$ are both well-defined and of full rank (non-zero).

In this example DGP, the regressor difference $\Delta \boldsymbol{X} = \boldsymbol{X}_1 - \boldsymbol{X}_2$ is identically one, so the (here scalar-valued) function $\boldsymbol{G}$ simplifies to $G(\theta)=\mathrm{E}[\dot{m}^{\texttt{tls}}_1(\theta,\boldsymbol{Y})],\theta\in\mathbb{R}$. It is then straightforward to check that the function $G$ is given by \[ G(\theta) =

cases\theta+\int_{0}^{-\theta}\Phi(u)\dif u, &\theta<0,\\ 0, &\theta=0,\\ \theta-\int_{0}^{\theta}\Phi(u)\dif u, &\theta>0,

\] where $\Phi$ is the standard normal CDF. For example, if $\theta < 0$, then letting $\varphi$ denote the standard normal PDF,

align*[align* omitted — 333 chars of source]

where the first equality follows from (ref), the second from (the here) identical distributions of $Y_1$ and $Y_2$, the third from noting that $Y_1$ is a standard normal random variable truncated from below at zero, and the fourth from integration by parts.

The graph of $G$ is shown in Figure (ref).

figure[figure omitted — 167 chars of source]

As indicated by the figure, and immediately follows analytically, $G$ is differentiable at $\theta_0(=0)$, with derivative $\dot{G}(0)=1 - \Phi(0)=\textstyle{\frac{1}{2}}$. However, since $Y_1$ and $Y_2$ are independently distributed as standard normals censored from below at zero, (the here scalar) $\boldsymbol{\Gamma}_0^{\tt{tls}}$ equals \[ \Gamma_0^{\tt{tls}}=\mathrm{E}[\mathbf{1}\{Y_1>0>-Y_2\}]=\mathbb{P}(Y_1>0)\cdot\mathbb{P}(Y_2>0)=\frac{1}{4}, \] showing that $\dot{G}(0)\neq\Gamma_0^{\tt{tls}}$.

How does this discrepancy arise? When arguing differentiability of the function $\boldsymbol{G}(\boldsymbol{\theta}) = \mathrm{E}\sbr[1]{\dot{m}_1^\texttt{tls}\del[0]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta},\boldsymbol{Y}}\Delta\boldsymbol{X}}$ at $\boldsymbol{\theta} = \boldsymbol{\theta}_0$, honore_trimmed_1992 called upon the LDCT to interchange the order of differentiation and expectation. However, as captured by $N(\boldsymbol{y})$ in (ref), the function $\dot{m}^{\texttt{tls}}_1(\cdot,\boldsymbol{y})$ fails to be differentiable at $-y_2$ and $y_1$, unless both $y_1$ and $y_2$ are zero. In turn, in a model with censoring, one may see exactly one outcome equal to zero with (strictly) positive probability. Thus, whenever the inner product $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ places mass at zero, the (random) function $t\mapsto \dot{m}_1^\texttt{tls}(t,\boldsymbol{Y})$ fails to be differentiable at $t=\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta}_0$ with (strictly) positive probability, invalidating the use of the LDCT.

Extended Asymptotic Normality Result

In this subsection, we modify the argument in honore_trimmed_1992 to allow the distribution of $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ to place mass at zero. We show that Honor{\'e}'s assumptions, corresponding to our Assumptions (ref)--(ref), imply differentiability of $\boldsymbol{G}$ at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$. We also derive an explicit expression for the Jacobian $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$, which coincides with the Hessian of the expected loss. Our calculations further show that $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$ generally differs from $\boldsymbol{\Gamma}_0^{\tt{tls}}$, unless mass at zero in the distribution of $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ is ruled out a priori. By deriving equivalent representations of this Hessian, we obtain easily interpretable necessary and sufficient conditions for its invertibility. We then state an asymptotic normality result that allows $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ to have mass at zero.

thm[Trimmed Least Squares Hessian Existence] Under Assumptions (ref)--(ref), the function $\boldsymbol{G}:\mathbb{R}^K\to\mathbb{R}^K$ defined by (ref) is differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$ with the resulting matrix $\mathbf{J}^{\tt{tls}}_0:=\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$ given by \begin{equation} \mathbf{J}^{\tt{tls}}_0 =\mathrm{E}\sbr[2]{\del[2]{1-F_{\varepsilon\mid\boldsymbol{W}}\del[1]{-\alpha-\min\cbr[0]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top}. \end{equation} Equivalently, \begin{align} \mathbf{J}^{\tt{tls}}_0 =\mathrm{E}\Big[\Big( &\mathbf{1}\{Y_1>0\}\del[1]{\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<0\}+\tfrac{1}{2}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0\}}\notag\\ +&\mathbf{1}\{Y_2>0\}\del[1]{\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0>0\}+\tfrac{1}{2}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0\}} \Big)\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\Big]. \end{align}

The expressions in (ref) and (ref) make it clear that $(Y_1,\boldsymbol{X}_1)$ and $(Y_2,\boldsymbol{X}_2)$ enter the Hessian $\mathbf{J}^{\tt{tls}}_0$ in a symmetric manner, so that the (time) labeling is irrelevant. The version in (ref) facilitates plug-in estimation of $\mathbf{J}^{\tt{tls}}_0$, which we cover in Section (ref). In turn, the version in (ref) facilitates comparison with Honor{\'e}'s $\boldsymbol{\Gamma}_0^{\tt{tls}}$. To this end, iterate expectations to write the latter as \[ \boldsymbol{\Gamma}_0^{\tt{tls}}=\mathrm{E}\sbr[2]{\mathrm{E}\sbr[1]{\mathbf{1}\cbr[0]{-Y_2<\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<Y_1}\mid\boldsymbol{W}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top}. \] We next expand on the inner expectation. Rewrite the inner expectation indicator as \[ 1-\mathbf{1}\{ Y_{1}\leqslant\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} -\mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} +\mathbf{1}\{ Y_{1}\leqslant \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} \mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\}. \] Because $Y_{1}$ and $Y_{2}$ are non-negative, the only way both indicators appearing in the previous display can be turned on is if $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0$. Again using non-negativity, we deduce \[ \mathbf{1}\{ Y_{1}\leqslant \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} \mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} =\mathbf{1}\{ \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0\} \mathbf{1}\{ Y_{1}=0\} \mathbf{1}\{ Y_{2}=0\}, \] from which it follows that $\mathbf{1}\{-Y_2<\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<Y_1\}$ can be written as

align*[align* omitted — 285 chars of source]

Taking conditional expectations, the inner expectation in $\boldsymbol{\Gamma}_0^{\tt{tls}}$ works out as

align*[align* omitted — 586 chars of source]

where (due to space constraints) we introduced the shorthand notations $V_1:=-(\alpha+\boldsymbol{X}_1^\top\boldsymbol{\theta}_0)$ and $V_2:=-(\alpha+\boldsymbol{X}_2^\top\boldsymbol{\theta}_0)$. Since $V_2-V_1=\boldsymbol{X}_1^\top\boldsymbol{\theta}_0-\boldsymbol{X}_2^\top\boldsymbol{\theta}_0=\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$, we have $V_1\gtreqqless V_2$ if $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\lesseqqgtr0$. With this newly introduced shorthand notation, the inner expectation in our Hessian (ref) can be written as $ 1-F_{\varepsilon\mid\boldsymbol{W}}\del[0]{\max\cbr{V_1,V_2}}, $ which agrees with the previous display whenever $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\neq0$. Hence, if this case can be ruled out [i.e., $\mathbb{P}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0)=0$], then Honor{\'e}'s $\boldsymbol{\Gamma}_0^{\tt{tls}}$ is equal to $\mathbf{J}^{\tt{tls}}$.

However, when $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0$, so that $V_1$ and $V_2$ take a common value (here: $V$), the inner expectations differ by

align*[align* omitted — 438 chars of source]

which is non-negative by monotone increasingness of CDFs. In the Section (ref) counterexample, we have both $V_{1}\equiv0$ and $V_{2}\equiv0$, and the (conditional) joint CDF factors into the product of the (standard normal) marginals, $F_{\boldsymbol{\varepsilon}\mid\boldsymbol{W}}(u_{1},u_{2})=\Phi(u_{1})\Phi(u_{2})$. Since there $\mathbb{P}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0)=1$, the difference in the previous display captures the only relevant case. The difference is then $\Phi(0)-\Phi(0)\Phi(0)=\frac{1}{4}$, which is precisely the previously demonstrated discrepancy.

rem[Alternative Hessian Expressions] As established in the proof of Theorem (ref), the Hessian $\mathbf{J}^{\tt{tls}}_0$ can also be expressed using either of the following expressions: \begin{align} \mathbf{J}^{\tt{tls}}_0&=\mathrm{E}\sbr[2]{\del[2]{\mathbf{1}\{Y_1>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\leqslant0\}+\mathbf{1}\{Y_2>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0>0\}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top}, \\ \mathbf{J}^{\tt{tls}}_0&=\mathrm{E}\sbr[2]{\del[2]{\mathbf{1}\{Y_1>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<0\}+\mathbf{1}\{Y_2>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\geqslant0\}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top}. \end{align} The version in (ref) is the average of the (ref) and (ref) right-hand sides.$\diamondsuit$
rem[Relaxing Exchangeability] Inspecting the proof of Theorem (ref) (specifically, the proof of Lemma (ref)), we see that it does not actually require full conditional exchangeability of $\varepsilon_1$ and $\varepsilon_2$ given $\boldsymbol{W}$. For the conclusions of Theorem (ref) to hold, Assumption (ref) can be replaced with the weaker condition of conditional stationarity: conditional on $\boldsymbol{W}$, $\varepsilon_1$ and $\varepsilon_2$ are identically distributed. Some of the generalizations made in arellano2001panel for Tobit-type models with fixed effects similarly only require the weaker assumption of conditional stationarity. When their $\psi(\cdot)$ and $\xi(\cdot)$ are both the identity mapping on the real line, the resulting estimator is precisely the honore_trimmed_1992 TLS estimator studied here, but motivated differently. arellano2001panel observe that “[i]t follows from standard results about extremum estimators that the resulting estimator will be consistent and $\sqrt n$ asymptotically normal”, but do not comment on the form of the limit variance components. The Hessian expressions in Theorem (ref) and Remark (ref) are therefore also relevant for this arellano2001panel generalization of the TLS estimator.$\diamondsuit$

To state the asymptotic normality result, we need to make sure that the Hessian $\mathbf{J}^{\tt{tls}}_0$ is invertible. To this end, we will impose the following condition:

assumption[Trimmed Least Squares Score Jacobian Invertibility] There is no proper linear subspace of $\mathbb{R}^K$ containing the random variable $(\mathbf{1}\{Y_1>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\leqslant0\}+\mathbf{1}\{Y_2>0\}\mathbf{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0>0\})\Delta\boldsymbol{X}$ with probability one.

This assumption is easy to interpret and implies via (ref) that the matrix $\mathbf{J}^{\tt{tls}}_0$ is invertible. Alternatively, we could state an assumption based on the expression for $\mathbf{J}^{\tt{tls}}_0$ in (ref). Strictly speaking, Assumption (ref) does not appear in honore_trimmed_1992 but the only purpose of this assumption is to ensure the invertibility of the Hessian $\mathbf{J}^{\tt{tls}}_0$, and honore_trimmed_1992 did impose invertibility of the corresponding Hessian $\boldsymbol{\Gamma}_0^{\tt{tls}}$.

Using our new understanding of the Hessian $\mathbf{J}^{\tt{tls}}_0$, we next state an extended asymptotic normality result for the TLS estimator that does not require the inner product $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ to have no mass at zero.

thm[Asymptotic Normality of Trimmed Least Squares] Let Assumptions (ref)--(ref) hold, and suppose that the expectations involved in defining the matrix $\mathbf{V}_0^{\tt{tls}}$ exist (in $\mathbb{R}^{K\times K}$), and that this matrix is of full rank. Then the TLS estimator satisfies \[ \sqrt{n}\del[1]{\widehat\boldsymbol{\theta}^{\tt{tls}}-\boldsymbol{\theta}_0}\rightsquigarrow\mathcal{N}\left(\mathbf{0},(\mathbf{J}_0^{\tt{tls}})^{-1}\mathbf{V}_0^{\tt{tls}}(\mathbf{J}_0^{\tt{tls}})^{-1}\right)\;\text{in}\;\mathbb{R}^K. \]

The asymptotic normality result in Theorem (ref) differs from honore_trimmed_1992 in the sense that it replaces the matrix $\boldsymbol{\Gamma}_0^{\tt{tls}}$ in (ref) by $\mathbf{J}_0^{\tt{tls}}$. Whenever $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\neq 0$ with probability one, the two matrices coincide. However, the matrices are in general different if $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0 = 0$ with (strictly) positive probability.

Asymptotic Variance Estimation

For the asymptotic normality to yield a practical approximation, we need to consistently estimate the asymptotic variance components. For $\mathbf{V}_0^{\tt{tls}}$, we use the plug-in estimator from honore_trimmed_1992: \[ \widehat{\mathbf{V}}^{\tt{tls}}:=\frac{1}{n}\sum_{i=1}^n \dot{m}_1^{\tt{tls}}(\Delta\boldsymbol{X}_i^\top\widehat\boldsymbol{\theta}^{\tt{tls}},\boldsymbol{Y}_i)^2\Delta\boldsymbol{X}_i\Delta\boldsymbol{X}_i^\top.\footnote{In \citet{honore_trimmed_1992}, this estimator is denoted $\hat{V}_4$.} \] For the Hessian $\mathbf{J}_0^{\tt{tls}}$, we apply the analogy principle to (ref) to arrive at

align[align omitted — 598 chars of source]

Alternatively, in order to estimate the Hessian $\mathbf{J}_0^{\tt{tls}}$, we could use the equivalent expressions in (ref) and (ref).

These estimators are (strongly) consistent under the assumptions of Theorem (ref):

thm[Plug-in Variance Estimator Consistency for TLS] Let the assumptions of Theorem (ref) hold. Then $\widehat\mathbf{V}^{\tt{tls}}\to_{\mathrm{a.s.}}\mathbf{V}_0^{\tt{tls}}$ and $\widehat\mathbf{J}^{\tt{tls}}\to_{\mathrm{a.s.}}\mathbf{J}_0^{\tt{tls}}$.

Theorem (ref) implies that the limit variance $(\mathbf{J}_0^{\tt{tls}})^{-1}\mathbf{V}_0^{\tt{tls}}(\mathbf{J}_0^{\tt{tls}})^{-1}$ is (strongly) consistently estimated by $(\widehat\mathbf{J}^{\tt{tls}})^{-1}\widehat\mathbf{V}^{\tt{tls}}(\widehat\mathbf{J}^{\tt{tls}})^{-1}$, which facilitates hypothesis testing and the construction of confidence intervals.

rem[Comparison with honore2000estimation] honore2000estimation discuss estimation of the censored regression model with fixed effects, allowing the number of time periods $T_i$ to exceed two and/or be individual specific (unbalanced panel data). While ibid. (Section 2.1) covers a whole class of estimators, in the special case of their $\xi(\cdot)$ being the identity mapping on the real line, the estimator in their (6) becomes a TLS estimator based on all pairs of time periods. In the balanced case $(T_i=T\geqslant2)$, their estimator is (in our notation) any minimizer of \[ \boldsymbol{\theta}\mapsto \frac{1}{n}\sum_{i=1}^n \sum_{1\leqslant \tau < \tau' \leqslant T} {m}^{\tt{tls}}\del[1]{(\boldsymbol{X}_{i\tau}-\boldsymbol{X}_{i\tau'})^\top\boldsymbol{\theta},(Y_{i\tau},Y_{i\tau'})}, \] which naturally generalizes the two-period TLS estimator studied in this paper. Appropriately extending our assumptions to their multi-period case, one can establish ($\sqrt n$-)consistency and asymptotic normality of their multi-period TLS estimator and reuse the argument leading to Theorem (ref) to show that the Hessian of the expected loss in the multi-period setting is given by an expression that naturally generalizes (ref) to accommodate more than one pair of time periods. One can then construct a (strongly) consistent estimator of this Hessian by applying the analogy principle, as we did in (ref) for the $T=2$ case. $\diamondsuit$
rem[A Plug-In Estimator for honore_pairwise_1994] Consider the {\em cross-sectional} censored regression model $Y = \max\{0,Y^*\}$, where $Y^* = \boldsymbol{X}^{\top}\boldsymbol{\theta}_0 + \varepsilon$ and $\varepsilon$ is independent of $\boldsymbol{X}$. This model was studied in honore_pairwise_1994, who developed, among other things, the TLS estimator of the vector of parameters $\boldsymbol{\theta}_0$ in this model, proved its ($\sqrt{n}$-)consistency, and derived the corresponding asymptotic normality result. On ibid. (p. 260), they gave an expression for the Hessian of the expected loss, which enters the asymptotic variance formula. In our notation, their Hessian takes the following form: \begin{align} \mathrm{E}\sbr[2]{\del[2]{1-F_{\varepsilon}\del[1]{-\min\cbr[0]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top}, \end{align} where the pairs $(Y_1,\boldsymbol{X}_1)$ and $(Y_2,\boldsymbol{X}_2)$ represent independent units of observation and $F_{\varepsilon}$ is the CDF of $\varepsilon$. On ibid. (p. 261), they also proposed a generic numerical derivative estimator of this Hessian. Due to numerical differentiation, however, their estimator introduces a tuning parameter through the choice of the stepsize. Our point in this remark is to show that one can actually estimate the Hessian in (ref) without introducing extra tuning parameters. To see that, let $\boldsymbol{Z}_i := (Y_i,\boldsymbol{X}_i)$, $i\in\{1,2,\dots,n\}$, be a random sample from the distribution of $\boldsymbol{Z}:=(Y,\boldsymbol{X})$, and observe that our panel-data Hessian $\mathbf{J}^{\tt{tls}}_0$ in (ref) reduces to the cross-sectional Hessian in (ref) upon setting $\alpha\equiv0$ and imposing independence of $\varepsilon_1$ and $\varepsilon_2$ from $\boldsymbol{W}$ in the former expression. Theorem (ref) thus reveals, via (ref), that the cross-sectional Hessian in (ref) is equal to $\mathrm{E}[h(\boldsymbol{Z}_1,\boldsymbol{Z}_2;\boldsymbol{\theta}_0)]$, where the function $h:\mathbb{R}^{1+K}\times\mathbb{R}^{1+K}\times\mathbb{R}^K\to\mathbb{R}^{K\times K}$ is given by \begin{align*} h(\boldsymbol{z}_1,\boldsymbol{z}_2;\boldsymbol{\theta}) :=\Big(&\mathbf{1}\{y_1>0\}\del[1]{\mathbf{1}\{\Delta\boldsymbol{x}^\top\boldsymbol{\theta}<0\}+\tfrac{1}{2}\mathbf{1}\{\Delta\boldsymbol{x}^\top\boldsymbol{\theta}=0\}}\\ +&\mathbf{1}\{y_2>0\}\del[1]{\mathbf{1}\{\Delta\boldsymbol{x}^\top\boldsymbol{\theta}>0\}+\tfrac{1}{2}\mathbf{1}\{\Delta\boldsymbol{x}^\top\boldsymbol{\theta}=0\}}\Big)\Delta\boldsymbol{x}\Delta\boldsymbol{x}^\top, \end{align*} and we denote $\boldsymbol{z}_1:=(y_1,\boldsymbol{x}_1)$, $\boldsymbol{z}_2:=(y_2,\boldsymbol{x}_2)$, and $\Delta\boldsymbol{x} := \boldsymbol{x}_1 - \boldsymbol{x}_2$. Let $\widehat{\boldsymbol{\theta}}$ be the honore_pairwise_1994 TLS estimator, i.e. their estimator with $\Xi\left(\cdot\right)=\textstyle{\frac{1}{2}}\left(\cdot\right)^2$. Then a natural estimator of the cross-sectional Hessian in (ref) is of the plug-in form \begin{align} \frac{1}{n(n-1)}\sum_{1\leqslant i\neq j\leqslant n}h(\boldsymbol{Z}_i,\boldsymbol{Z}_j;\widehat\boldsymbol{\theta}), \end{align} which is an approximate second-order U-statistic. Note that this estimator involves no tuning parameter, and is therefore free of the stepsize problem arising due to numerical differentiation. The consistency of this plug-in estimator can be established using an argument paralleling the one used in the proof of Theorem (ref). The main difference is that, instead of establishing a uniform law of large numbers involving simple averages (or, equivalently, an empirical process), to accommodate the expression in (ref), one must now work with generalized averages (leading to a second-order U-process). To this end, we refer the reader to nolan1987_u_processes_rates.$\diamondsuit$
rem[Comparison with powell1984least,powell1986symmetrically] As noted in honore_trimmed_1992, the rationale behind the TLS and TLAD estimators considered in this paper can be viewed as a bivariate generalization of the idea behind the powell1986symmetrically symmetrically trimmed LS estimators for censored Tobit models without fixed effects. The approach adopted in powell1986symmetrically is, in turn, closely related to the censored LAD (CLAD) estimator proposed in powell1984least---a key paper in the censored regression literature. In our notation, powell1984least models the conditional median of a non-negative response variable $Y$ given covariates $\boldsymbol{X}$ as $\mathrm{Med}(Y|\boldsymbol{X})=\max\{0,\boldsymbol{X}^\top\boldsymbol{\theta}_0\}$. To ensure asymptotic normality of the CLAD estimator, powell1984least rules out regressors that are orthogonal to $\boldsymbol{\theta}_0$ with positive probability. The possibility of such regressors is precisely the cause of difficulty in our analysis, cf. the discussion following Theorem (ref).$\diamondsuit$

Consistency of the Honor{\'e} (1992) Hessian Estimator

In honore_trimmed_1992 the matrix $\boldsymbol{\Gamma}_0^{\tt{tls}}$ in (ref) is estimated by its empirical analogue

equation[equation omitted — 378 chars of source]

The TLS sandwich variance consistency results in honore_trimmed_1992 involve showing that $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ is a consistent estimator of $\boldsymbol{\Gamma}_0^{\tt{tls}}$. As we have shown by example in Section (ref), $\boldsymbol{\Gamma}_0^{\tt{tls}}$ is not equal to the Hessian $\mathbf{J}_0^{\tt{tls}}$ of the expected loss in general. As $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ is targeting $\boldsymbol{\Gamma}_0^{\tt{tls}}$, one would expect that $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ is not a consistent estimator of the TLS Hessian $\mathbf{J}_0^{\tt{tls}}$ in general. Surprisingly, as we establish below, $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ nevertheless converges to the TLS Hessian $\mathbf{J}_0^{\tt{tls}}$.

thm[Consistency of the honore_trimmed_1992 Hessian Estimator for TLS] Let the assumptions of Theorem (ref) hold. Then $\widehat\boldsymbol{\Gamma}^{\tt{H92}}=\mathbf{J}_0^{\tt{tls}}+o_{L^1}(1)+o_{\mathrm{a.s.}}(1)=\mathbf{J}_0^{\tt{tls}}+o_{\mathbb{P}}(1)$.

This counterintuitive result can be explained as follows. Decompose $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ into two parts:

align*[align* omitted — 651 chars of source]

Although $\widehat\mathbf{L}^{\tt{H92}}$ and $\widehat\mathbf{J}^{\tt{tls}}$ are not identical term by term, the leading component $\widehat\mathbf{L}^{\tt{H92}}$ is closely related to our Hessian estimator in (ref), which is the empirical analogue of $\mathbf{J}_0^{\tt{tls}}$. The proof of Theorem (ref) shows that $\widehat\mathbf{L}^{\tt{H92}}$ converges to $\mathbf{J}_0^{\tt{tls}}$, while the remainder term $\widehat\mathbf{R}^{\tt{H92}}$ converges to the zero matrix. Key to these findings is our extended asymptotic normality result in Theorem (ref). Specifically, even if $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ has mass at zero, the absolutely continuous limiting distribution of $\sqrt n(\widehat\boldsymbol{\theta}^{\tt{tls}}-\boldsymbol{\theta}_0)$ implies that, for any fixed $\boldsymbol{x}\in\mathbb{R}^K\backslash\{\boldsymbol{0}\}$, $\mathbb{P}(\boldsymbol{x}^\top\widehat\boldsymbol{\theta}^{\tt{tls}}=0)\to0$. We stress that Theorem (ref) is not a simple consequence of consistency of $\widehat\boldsymbol{\theta}^{\tt{tls}}$, established in honore_trimmed_1992. Indeed, even with a uniform law of large numbers allowing us to replace sample averages in $\widehat{\boldsymbol{\Gamma}}^{\tt{H92}}$ by their population counterparts (see the proof of Theorem (ref)), the resulting map $\boldsymbol{\theta}\mapsto\boldsymbol{\Gamma}(\boldsymbol{\theta})$ need not be continuous at $\boldsymbol{\theta}_0$. Consequently, a continuous mapping argument need not apply. In the counterexample of Section (ref), one finds $\Gamma(\theta)=[1-\Phi(|\theta|)]\boldsymbol{1}\{\theta\neq0\}+\tfrac{1}{4}\boldsymbol{1}\{\theta=0\}$, which has a jump discontinuity at $\theta=0=\theta_0$. The jump size $(\tfrac{1}{4})$ is exactly the difference between $\mathbf{J}_0^{\tt{tls}}$ and $\boldsymbol{\Gamma}_0^{\tt{tls}}$ in that example.

A by-product of Theorem (ref) is that, under our maintained assumptions, the TLS Hessian consistency statement in honore_trimmed_1992 does not hold without additional regularity conditions ensuring continuity at $\boldsymbol{\theta}_0$.

Trimmed Least Absolute Deviations

In this section, we focus on the TLAD estimator. We abbreviate the trimmed absolute loss, $m_{\Xi}$ in (ref) with $\Xi=\left|\cdot\right|$, by $m^{\texttt{tlad}}$, which takes the form

equation[equation omitted — 303 chars of source]

Assumptions for Trimmed Least Absolute Deviations

For notational convenience, abbreviate the TLAD estimator $\widehat\boldsymbol{\theta}_\Xi$ in (ref) with $\Xi=\left|\cdot\right|$ by $\widehat{\boldsymbol{\theta}}^{\texttt{tlad}}$. Also, let $\boldsymbol{W}$, $\mathcal W$, $\boldsymbol{w}$, $\boldsymbol{\varepsilon}$, $\mathcal E$, and $\boldsymbol{e}$ be the same as in Section (ref). Consider the following assumptions.\footnote{Assumptions (ref)--(ref) are from Assumptions S.2, M.2, E.1, E.3, E.4 and R.1, respectively, in honore_trimmed_1992.}

assumption[Non-Degeneracy] The probability $\mathbb{P}(Y_1>0, Y_2>0)$ is strictly positive.
assumption[Integrability] All of the following expectations are finite: \[ \mathrm{E}[\|\boldsymbol{X}_1\|_2^2],\;\mathrm{E}[\|\boldsymbol{X}_2\|_2^2],\;\mathrm{E}[\|\alpha\Delta\boldsymbol{X}\|_2],\; \mathrm{E}[\|\varepsilon_1\Delta\boldsymbol{X}\|_2]\quad\text{and}\quad \mathrm{E}[\|\varepsilon_2\Delta\boldsymbol{X}\|_2]. \]
assumption[Continuity] The conditional distribution of $(\varepsilon_1,\varepsilon_2)$ given $\boldsymbol{W}$ is absolutely continuous with respect to the Lebesgue measure.
assumption[Exchangeability] Conditional on $\boldsymbol{W}$, $\varepsilon_{1}$ and $\varepsilon_{2}$ are exchangeable.

Since Assumptions (ref) and (ref) are the same as Assumptions (ref) and (ref), following the reasoning in Section (ref), there exists a function $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$, mapping $\mathcal W\times\mathcal E$ to $[0,\infty)$, that is a version of the PDF of the pair $\boldsymbol{\varepsilon} = (\varepsilon_1,\varepsilon_2)$ conditional on $\boldsymbol{W} = \boldsymbol{w}$, which is measurable in $(\boldsymbol{w},\boldsymbol{e})$, and is such that $f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_1,e_2) = f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_2,e_1)$ for all $\boldsymbol{w}\in\mathcal W$ and $\boldsymbol{e} = (e_1,e_2)\in\mathcal E$. Also, let $(\boldsymbol{w},e)\mapsto f_{\varepsilon\mid\boldsymbol{w}}(e) = \int_{\mathbb R}f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e,u)\dif u$ be the corresponding version of the common marginal PDF of $\varepsilon_1$ and $\varepsilon_2$ conditional on $\boldsymbol{W} = \boldsymbol{w}$ and let $(\boldsymbol{w},e)\mapsto f_{\varepsilon_1 - \varepsilon_2\mid\boldsymbol{w}}(e) = \int_{\mathbb R} f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(u+e,u)\dif u$ be the corresponding version of the PDF of the difference $\varepsilon_1 - \varepsilon_2$ conditional on $\boldsymbol{W} = \boldsymbol{w}$.

assumption[Regularity] There is a constant $C\in(0,\infty)$ such that $\sup_{e\in\mathbb{R}}f_{\varepsilon_{1}-\varepsilon_{2}\mid\boldsymbol{W}}(e)\leqslant C$ and $\sup_{e\in\mathbb{R}}f_{\varepsilon\mid\boldsymbol{W}}(e)\leqslant C$ with probability one.
assumption[Rank of Regressors] There is no proper linear subspace of $\mathbb{R}^K$ containing the random variable $\mathbf{1}\{\mathbb{P}\left(Y_1>0, Y_2>0\middle|\boldsymbol{X}_1,\boldsymbol{X}_2\right)>0\}\Delta\boldsymbol{X}$ with probability one.
assumption[Continuity, II] The functions $\boldsymbol{e}\mapsto f_{\boldsymbol{\varepsilon}\mid \boldsymbol{W}}(\boldsymbol{e})$, $e\mapsto f_{\varepsilon\mid\boldsymbol{W}}(e)$, and $e\mapsto f_{\varepsilon_1 - \varepsilon_2\mid\boldsymbol{W}}(e)$ are continuous with probability one.

Assumptions (ref)--(ref) are the same as the corresponding assumptions in honore_trimmed_1992. Assumption (ref), however, is not present in honore_trimmed_1992. We consider this assumption because the asymptotic normality result in honore_trimmed_1992 may not hold without it. In particular, in Section (ref) below we provide a DGP that satisfies Assumptions (ref)--(ref) and is such that the Hessian of the expected loss [see (ref) below], which appears in the asymptotic variance formula in honore_trimmed_1992, does not exist. We note also that Assumption (ref) may be stronger than necessary. For example, it seems possible to obtain the asymptotic normality result assuming continuity of the functions in Assumption (ref) only on their respective supports but we opt for a stronger than necessary conditions for clarity of the argument.

Note also that Assumptions (ref), (ref), (ref), and (ref) are the same as the corresponding assumptions for the TLS estimator (Assumptions (ref), (ref), (ref), and (ref), respectively). The TLAD moment conditions in Assumption (ref) are weaker than those for the TLS estimator (Assumption (ref)), which reflects the fact that the trimmed absolute loss is less sensitive to outliers than the trimmed square loss. Assumptions (ref) and (ref), requiring certain PDFs to be bounded and continuous, do not have analogs in the case of the TLS estimator.

Asymptotic Normality in Honor{\'e} (1992)

To state the TLAD normality result in honore_trimmed_1992, introduce the $K\times K$ matrices\footnote{In honore_trimmed_1992, these matrices are denoted $V_3$ and $\Gamma_3$, respectively. Honor{\'e}'s $\Gamma_3$ is stated in terms of probabilities and densities of the censored $Y_1$ and $Y_2$, but it is clear from the underlying proof that he means the latent variables $Y_1^\ast$ and $Y_2^\ast$, respectively.}

equation[equation omitted — 414 chars of source]

and

equation[equation omitted — 846 chars of source]

where $(\boldsymbol{w},e) \mapsto f_{Y_1^\ast-Y_2^\ast\mid\boldsymbol{w},Y_1^\ast>0,Y_2^\ast>0}\left(e\right)$ is the conditional PDF of $Y_1^\ast-Y_2^\ast$ given $\boldsymbol{W}=\boldsymbol{w}$ and $\{Y_1^\ast>0\}\cap \{Y_2^\ast>0\}$, $(\boldsymbol{w},e) \mapsto f_{Y^\ast_1\mid\boldsymbol{w},Y_2^\ast\leqslant0}\left(e\right)$ is the conditional PDF of $Y_1^\ast$ given $\boldsymbol{W}=\boldsymbol{w}$ and $Y_2^\ast\leqslant0$, and $(\boldsymbol{w},e)\mapsto f_{Y^\ast_2\mid\boldsymbol{w},Y_1^\ast\leqslant0}\left(e\right)$ is the conditional PDF of $Y_2^\ast$ given $\boldsymbol{W}=\boldsymbol{w}$ and $Y_1^\ast\leqslant0$.\footnote{Here and below, we assume that the versions of all conditional PDFs and CDFs including latent outcomes $Y_1^*$ and $Y_2^*$ are obtained (fixed) by combining the conditional PDF $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$ with (ref).}

honore_trimmed_1992 states that if Assumptions (ref)--(ref) hold, the expectations involved in defining the matrices $\mathbf{V}_0^{\tt{tlad}}$ and $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ exist (in $\mathbb{R}^{K\times K}$), and both matrices are of full rank, then

equation[equation omitted — 316 chars of source]

Among several steps, the proof of this result in honore_trimmed_1992 includes establishing the existence of the Hessian of the expected loss [see (ref) below] at the true parameter value $\boldsymbol{\theta}_0$. In our notation, this task corresponds to arguing that the function $L:\mathbb{R}^K\to\mathbb{R}$ defined by

equation[equation omitted — 253 chars of source]

is twice differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$.\footnote{Note that because the function $m^{\texttt{tlad}}(t,\boldsymbol{y})$ is Lipschitz continuous in its first argument, it follows from Assumption (ref) that the expectation in (ref) is well-defined. Also note that, as in honore_trimmed_1992, we work with the expected loss difference $\mathrm{E}[m^{\texttt{tlad}}(\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta},\boldsymbol{Y}) - m^{\texttt{tlad}}(0,\boldsymbol{Y})]$ instead of $\mathrm{E}[m^{\texttt{tlad}}(\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta},\boldsymbol{Y})]$ because the former loss gives results under weaker regularity conditions.} To this end, Honor{\'e} used the LDCT to differentiate once under the expectation, and then argued differentiability at $\boldsymbol{\theta}_0$ of the resulting function to arrive at $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ in (ref). In the next subsection, however, we will show by example that Assumptions (ref)--(ref) are not sufficient to ensure the existence of the Hessian.

Counterexample

As in the Section (ref) counterexample, let $K=1$, $\alpha\equiv0$, $\boldsymbol{\theta}_0 = 0$, $\boldsymbol{X}_1\equiv2$ and $\boldsymbol{X}_2\equiv1$. Also, to define the distribution of the pair $(\varepsilon_1,\varepsilon_2)$, let $r:\mathbb R\to \mathbb R$ be a continuous function such that (i) $r(t)\geqslant 0$ for all $t\in\mathbb R$, (ii) $\int_{\mathbb R} r(t)\dif t = 1$, and (iii) $r(t)=0$ if $t\leqslant 1$ or $t\geqslant 3$. In addition, let $\mathscr E := \{0,2,4,\dots\}$, $\mathscr O = \{1,3,5,\dots\}$, and let $\tilde h:[0,1]\to\{0,1\}$ be the function defined by $$ \tilde{h}(t)=

cases1, & t\in(2^{-(k+1)},2^{-k}] for k\in \mathscr E,\\ 0, & t\in(2^{-(k+1)},2^{-k}] for k\in \mathscr O,\\ 0, & t=0.

$$ Moreover, let $h:\mathbb R \to\mathbb R$ be the function defined by $h(t) = 3\tilde h(|t|) / 4$ for $t\in[-1,1]$ and $0$ otherwise. Note that $h(t)\geqslant0$ for all $t\in\mathbb R$ and $$ \int_{\mathbb R} h(t)\dif t = \frac{3}{4}\int_{-1}^1 \tilde h(\left|t\right|)\dif t = \frac{3}{2}\int_{0}^1 \tilde h(t)\dif t = \frac{3}{2}\sum_{k\in\mathscr E}\frac{1}{2^{k+1}} = \frac{3}{2}\cdot \frac{1}{2}\cdot\frac{1}{1-\frac{1}{4}} = 1. $$ Thus, the function $f\colon\mathbb R^2 \to \mathbb R$ defined by $$ f(e_1,e_2) = 2h(e_1 - e_2)r(e_1 + e_2),\quad e_1,e_2\in\mathbb R, $$ satisfies $f(e_1,e_2)\geqslant 0$ for all $e_1,e_2\in\mathbb R$ and

align*[align* omitted — 274 chars of source]

Hence, $f$ is the PDF of a certain distribution on $\mathbb R^2$. Let $(\varepsilon_1,\varepsilon_2)$ be a pair of random variables sampled from this distribution. Because of the symmetry of the function $f$, the random variables $\varepsilon_1$ and $\varepsilon_2$ are then exchangeable, and their common PDF is $$ f_{\varepsilon}(t) = 2\int_{\mathbb R}h(t - e_2)r(t+e_2)\dif e_2 = 2\int_{\mathbb R} h(u)r(u+2t)\dif u,\quad t\in\mathbb R, $$ which implies that $0<\varepsilon_1<2$ and $0<\varepsilon_2<2$ with probability one. Moreover, the PDF of the difference $\varepsilon_1 - \varepsilon_2$ is

align*[align* omitted — 198 chars of source]

It is then straightforward to check that Assumptions (ref)--(ref) are all satisfied. Note, however, that Assumption (ref) is {\em not} satisfied because the function $f$ is not continuous.

Now, observe that since $\mathbb P(Y_1>0, Y_2>0) = 1$ by construction, a calculation shows that the expected loss function $L$ in (ref) simplifies to $$ L(\theta) = \mathrm{E}\left[m^{\texttt{tlad}}(\theta,\boldsymbol{Y}) - m^{\texttt{tlad}}(0,\boldsymbol{Y})\right] = \mathrm{E}\left[|Y_1 - Y_2 - \theta| - |Y_1 - Y_2|\right],\quad \theta\in\mathbb R. $$ Since the integrand here is convex in $\theta$ and non-differentiable in $\theta$ only at $Y_1 - Y_2$, which, for a given $\theta$, happens with probability zero, it follows from bertsekas1973stochastic that $L$ is (everywhere) differentiable with derivative $$ \dot{L}(\theta) = \mathrm{E}\left[2\cdot\mathbf{1}\{Y_1 - Y_2 \leqslant \theta\} - 1\right] = 2F_{\varepsilon_1 - \varepsilon_2}(\theta) - 1, $$ where $F_{\varepsilon_1 - \varepsilon_2}$ denotes the CDF of the difference $\varepsilon_1 - \varepsilon_2$. We claim that $F_{\varepsilon_1 - \varepsilon_2}$ is non-differentiable at zero $(=\theta_0)$. Indeed, if $t_k = 2^{-k}$ for $k\in\mathscr E$, then $$ \frac{F_{\varepsilon_1 - \varepsilon_2}(t_k) - F_{\varepsilon_1 - \varepsilon_2}(0)}{t_k} = 2^k \int_0^{2^{-k}} h(t)\dif t = \frac{3}{4}\cdot 2^k \cdot \sum_{l\in\mathscr E : l\geqslant k}\frac{1}{2^{l+1}} = \frac{1}{2} $$ and if $t_k = 2^{-k}$ for $k\in\mathscr O$, then $$ \frac{F_{\varepsilon_1 - \varepsilon_2}(t_k) - F_{\varepsilon_1 - \varepsilon_2}(0)}{t_k} = 2^k \int_0^{2^{-k}} h(t)\dif t = \frac{3}{4}\cdot 2^k \cdot \sum_{l\in\mathscr E : l\geqslant k+1}\frac{1}{2^{l+1}} = \frac{1}{4}, $$ which implies that the limit of $(F_{\varepsilon_1 - \varepsilon_2}(t) - F_{\varepsilon_1 - \varepsilon_2}(0))/t$ as $t\to 0$ does not exist. Thus, the function $\dot{L}$ is non-differentiable at zero ($=\theta_0$), implying that the Hessian of $L$ at $\theta_0$ does not exist.

One can show that the TLAD estimator is still $\sqrt{n}$-consistent but not (asymptotically) normal. Specifically, in Figure (ref), we show the (kernel) PDF of the scaled TLAD estimates $\sqrt{n}\widehat{\theta}^{\tt{tlad}}$ based on $10,000$ Monte Carlo samples of size $n=50,000$ using the DGP defined earlier in this section.\footnote{The kernel density was created using the R package ggplot2 with geom_density. We use a Gaussian kernel and the silverman_density_1986 rule-of-thumb bandwidth (both geom_density defaults).} To facilitate comparison, we include the PDF resulting from fitting a Gaussian distribution to the Monte Carlo dataset of scaled estimates using maximum likelihood. The PDF of the scaled estimates looks far from Gaussian. We therefore conclude that establishing the asymptotic normality result for the TLAD estimator requires more than Assumptions (ref)--(ref). We do so by imposing the continuity conditions in Assumption (ref).

figure[figure omitted — 271 chars of source]

Extended Asymptotic Normality Result

In this subsection, we provide a Hessian existence argument for the TLAD estimator using Assumptions (ref)--(ref). We also show that, unless the pair of latent outcomes $(Y_1^\ast,Y_2^\ast)$ belongs to the first quadrant with conditional probability given $\boldsymbol{W}$ equal to one almost surely, $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ will not be the Hessian of the expected loss (because of an apparent typographical error). We then state the asymptotic normality result under Assumptions (ref)--(ref) with the corrected Hessian.

To state the existence theorem, let $(\boldsymbol{w},\boldsymbol{y})\mapsto f_{\boldsymbol{Y}^\ast\mid\boldsymbol{w}}\left(\boldsymbol{y}\right)$ denote the joint PDF of the latent outcomes $Y_1^\ast$ and $Y_2^\ast$ conditional on $\boldsymbol{W}=\boldsymbol{w}$ for all $\boldsymbol{w}\in\mathcal{W}$.

thm[Trimmed Least Absolute Deviations Hessian Existence] Under Assumptions (ref)--(ref), the function $L:\mathbb{R}^K\to\mathbb{R}$ defined by (ref) is twice differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$. Its Hessian matrix $\mathbf{H}_0^{\tt{tlad}} := \nabla^2L(\boldsymbol{\theta}_0)$ is given by \begin{equation} \left.\begin{aligned} \mathbf{H}_0^{\tt{tlad}} &=\mathrm{E}\Bigg[\bigg(2\int_{0}^{+\infty}f_{\boldsymbol{Y}^{\ast}\mid\boldsymbol{W}}\del[1]{z+\max\cbr[0]{0,\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0},z-\min\cbr[0]{0,\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0}}\dif z\\ & \quad\qquad+\boldsymbol{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\geqslant0\}\int_{-\infty}^{0}f_{\boldsymbol{Y}^{\ast}\mid\boldsymbol{W}}\del[1]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0,z}\dif z\\ & \quad\qquad+\boldsymbol{1}\{\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<0\}\int_{-\infty}^{0}f_{\boldsymbol{Y}^{\ast}\mid\boldsymbol{W}}\del[1]{z,-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0}\dif z\bigg)\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\Bigg]. \end{aligned}\right\} \end{equation}
rem[Comparison with honore_trimmed_1992] Comparing $\mathbf{H}_0^{\tt{tlad}}$ in (ref) with the matrix $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ in (ref), and writing out the definitions of the conditional PDFs involved, we see that the latter two contributions to each expression are equal. (To align the two expressions, we here interpret “undefined” times “zero” as “zero.”) However, comparing the first terms of each matrix, we see that the conditional PDF $f_{Y_1^\ast-Y_2^\ast\mid\boldsymbol{W},Y_1^\ast>0,Y_2^\ast>0}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0)$ in Honor{\'e}'s $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ is missing multiplication by the conditional probability $\mathbb{P}\left(Y_1^\ast>0, Y_2^\ast>0\middle|\boldsymbol{W}\right)$. Hence, unless the pair of latent outcomes $(Y_1^\ast,Y_2^\ast)$ belongs to the first quadrant with conditional probability given $\boldsymbol{W}$ equal to one almost surely, $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ will not be the Hessian of the expected loss. Of course, if $(Y_1^\ast,Y_2^\ast)$ resides in the first quadrant with probability one, then the model involves no censoring.$\diamondsuit$

We now revise the asymptotic normality statement in (ref) for the TLAD estimator based on our new understanding of the Hessian $\mathbf{H}_0^{\tt{tlad}}$.

thm[Asymptotic Normality of Trimmed Least Absolute Deviations] Let Assumptions (ref)--(ref) hold, and suppose that the expectations involved in defining the matrix $\mathbf{V}_0^{\tt{tlad}}$ exist (in $\mathbb{R}^{K\times K}$), and that both matrices $\mathbf{V}_0^{\tt{tlad}}$ and $\mathbf{H}_0^{\tt{tlad}}$ are of full rank. Then \begin{equation} \sqrt{n}\del[1]{\widehat\boldsymbol{\theta}^{\tt{tlad}}-\boldsymbol{\theta}_0}\rightsquigarrow\mathcal{N}\left(\mathbf{0},(\mathbf{H}_0^{\tt{tlad}})^{-1}\mathbf{V}_0^{\tt{tlad}}(\mathbf{H}_0^{\tt{tlad}})^{-1}\right)\;in\;\mathbb{R}^K. \end{equation}
rem[Comparison with honore_pairwise_1994] We return to the comparison with honore_pairwise_1994, who considered censored regression with cross-sectional data as discussed in Remark (ref). On ibid. (p. 260), the authors gave an expression for the Hessian of the expected loss at the true parameter value $\boldsymbol{\theta}_0$ when using the trimmed absolute loss. Their expression is (in our notation) \begin{align*} &\mathrm{E}\Bigg[\bigg(2\int_{0}^{+\infty}f_{\varepsilon}\del[1]{z-\min\cbr[1]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}^2\dif z\\ &\qquad+f_{\varepsilon}\del[1]{-\min\cbr[1]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}F_{\varepsilon}\del[1]{-\min\cbr[1]{\boldsymbol{X}_1^\top\boldsymbol{\theta}_0,\boldsymbol{X}_2^\top\boldsymbol{\theta}_0}}\bigg) \Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\Bigg], \end{align*} with the pairs $(Y_1,\boldsymbol{X}_1)$ and $(Y_2,\boldsymbol{X}_2)$ representing independent units, where $f_{\varepsilon}$ and $F_{\varepsilon}$ denote the PDF of $\varepsilon$ and the CDF of $\varepsilon$, respectively. Upon setting $\alpha\equiv0$ in our panel-data setting, and taking the model errors $(\varepsilon_1,\varepsilon_2)$ to be independent and identically distributed, the conditional PDF of $(Y_1^\ast,Y_2^\ast)$ given $\boldsymbol{W}=\boldsymbol{w}$ factors as $ f_{Y_1^\ast,Y_2^\ast\mid\boldsymbol{w}}\left(y_1^\ast,y_2^\ast\right)=f_{\varepsilon}\del[1]{y_1^\ast-\boldsymbol{x}_1^\top\boldsymbol{\theta}_0}f_{\varepsilon}\del[1]{y_2^\ast-\boldsymbol{x}_2^\top\boldsymbol{\theta}_0}. $ A straightforward calculation then shows that our panel-data Hessian in (ref) reduces to the cross-sectional Hessian from honore_pairwise_1994 in the previous display.$\diamondsuit$

Asymptotic Variance Estimation

We now discuss estimation of the asymptotic variance $\boldsymbol{\Sigma}_0^{\tt{tlad}}:=(\mathbf{H}_0^{\tt{tlad}})^{-1}\mathbf{V}_0^{\tt{tlad}}(\mathbf{H}_0^{\tt{tlad}})^{-1}$ in (ref). Since the formula for $\mathbf{H}_0^{\tt{tlad}}$ given in (ref) includes a nonparametric conditional PDF, we propose a bootstrap estimator instead of a plug-in estimator. Also, since a na{\"i}ve bootstrap variance estimator may fail to be consistent for LAD-type estimators without additional integrability conditions BF81,GPSB84, we instead construct a quantile-based robust bootstrap variance estimator that avoids reliance on bootstrap second moments.

The robust bootstrap variance estimator exploits the fact that quantiles of linear combinations of a normal distribution scale with their standard deviations. We therefore estimate bootstrap quantiles of suitably centered parameter combinations and invert this scaling to recover the corresponding covariance entries. This construction replaces unstable bootstrap variance estimation based on second moments by identification of covariance entries through the asymptotic normal geometry of the estimator.

To describe the robust bootstrap variance estimator in more detail, consider a bootstrap sample $\{(\widetilde{Y}_{i1},\widetilde{\boldsymbol{X}}_{i1},\widetilde{Y}_{i2},\widetilde{\boldsymbol{X}}_{i2})\}_{i=1}^n$. obtained by sampling $n$ observation indices with replacement from $\{1,2,\dots,n\}$ and extracting the corresponding observations from the original sample. Also, let

equation[equation omitted — 389 chars of source]

be the bootstrap analog of $\widehat\boldsymbol{\theta}^{\tt{tlad}} = (\widehat\theta^{\tt{tlad}}_1,\dots,\widehat\theta^{\tt{tlad}}_K)^\top$, where we introduced shorthands $\Delta\widetilde{\boldsymbol{X}}_i := \widetilde{\boldsymbol{X}}_{i1} - \widetilde{\boldsymbol{X}}_{i2}$ and $\widetilde{\boldsymbol{Y}}_i:=(\widetilde{Y}_{i1},\widetilde{Y}_{i2})$. Identification of the covariance entries relies on the identities $\mathrm{var}(Z_k)=\Sigma^{\tt{tlad}}_{0kk}$ and $\mathrm{var}(Z_j+Z_k)=\Sigma^{\tt{tlad}}_{0jj} +\Sigma^{\tt{tlad}}_{0kk} +2\Sigma^{\tt{tlad}}_{0jk}$ for a normal vector $\boldsymbol{Z} \sim\mathcal{N}(\boldsymbol{0},\boldsymbol{\Sigma}_0^{\tt{tlad}})$. Accordingly, we recover variances and covariances from bootstrap quantiles of the corresponding linear combinations. Motivated by these identities, we estimate the diagonal and off-diagonal entries of $\boldsymbol{\Sigma}_0^{\tt{tlad}}=[\Sigma^{\tt{tlad}}_{0jk}]_{j,k=1}^K$ separately. For the diagonal elements, for each $k\in[K]$, we set $\widehat\Sigma_{kk}^{\tt{tlad}}:=[\widehat{q}_{0.9,k}/\Phi^{-1}(0.95)]^2$, where $\widehat{q}_{0.9,k}$ is the $0.9$ quantile of the conditional distribution of $\sqrt n|\widetilde\theta^{\tt{tlad}}_k - \widehat\theta^{\tt{tlad}}_k|$ given the original data and $\Phi^{-1}(0.95)$ is the $0.95$ quantile of the standard normal distribution. For the off-diagonal elements, for each $(j,k)\in[K]\times[K]$ such that $j\neq k$, we set $\widehat \Sigma_{jk}^{\tt{tlad}}:= ([\widehat{q}_{0.9,j,k}/\Phi^{-1}(0.95)]^2 - \widehat\Sigma_{jj}^{\tt{tlad}} - \widehat\Sigma_{kk}^{\tt{tlad}})/2$, where $\widehat{q}_{0.9,j,k}$ is the $0.9$ quantile of the conditional distribution of $\sqrt n|\widetilde\theta^{\tt{tlad}}_j + \widetilde\theta^{\tt{tlad}}_k - \widehat\theta^{\tt{tlad}}_j - \widehat\theta^{\tt{tlad}}_k|$ given the original data.\footnote{As the robust bootstrap covariance estimator $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ is recovered entrywise from quantiles of linear combinations, it need not be positive semidefinite in finite samples. One can construct a positive semidefinite variance estimator by projecting $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ onto the set of positive semidefinite matrices.} In the following theorem, we prove consistency of $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}:=[\widehat\Sigma_{jk}^{\tt{tlad}}]_{j,k=1}^K$ under the assumptions of Theorem (ref).

thm[Consistency of the Robust Bootstrap Variance Estimator for TLAD] Let the assumptions of Theorem (ref) hold. Then $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}\to_{\mathbb{P}}\boldsymbol{\Sigma}_0^{\tt{tlad}}$.
rem[Reducing Computational Burden via HH17] Note that calculating the quantiles $\widehat{q}_{0.9,k}$ and $\widehat{q}_{0.9,j,k}$ requires solving the $K$-dimensional optimization problem (ref) for many draws of the bootstrap sample. To reduce computational burden, one may instead employ the alternative bootstrap procedure of HH17, which replaces repeated $K$-dimensional optimization by a sequence of one-dimensional problems. This approach is particularly attractive in our setting because the “meat” matrix $\mathbf{V}_0^{\tt{tlad}}$ in (ref) admits a consistent plug-in estimator. We refer the reader to the original paper for further details.$\diamondsuit$
rem[Bootstrap Inference for Convex Pairwise Differencing Estimators] A related recent paper is cattaneo2025robust, which develops distribution theory and bootstrap-based inference for a broad class of convex pairwise differencing estimators. Their framework includes trimmed absolute loss as a special case. Our analysis in Theorem (ref) is complementary in that it considers a panel setting, where the trimmed loss is used to eliminate an individual-specific fixed effect and no bandwidth choice is required. In addition, while cattaneo2025robust establish bootstrap validity for the distributional approximation, Theorem (ref) establishes consistency of a robust bootstrap variance estimator for the asymptotic variance of the TLAD estimator under the assumptions of Theorem (ref).$\diamondsuit$
rem[Conservative Inference] The robust bootstrap variance estimator $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ is consistent for $\boldsymbol{\Sigma}_0^{\tt{tlad}}$ under the assumptions of Theorem (ref). However, the construction of $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ is based on a particular choice of quantile level (i.e., $0.9$ for the absolute differences, corresponding to $0.95$ for the standard normal quantiles). This choice is not essential: the consistency argument in the proof of Theorem (ref) continues to apply for any fixed interior quantile. We use the particular quantile level above as a convenient way to avoid reliance on bootstrap second moments, which may be unstable for LAD-type estimators without additional integrability conditions. For comparison, hahn2021bootstrap show that, under their conditions, bootstrap variance estimators based on second moments, although not necessarily consistent in LAD-type settings, can still be asymptotically conservative for inference on fixed linear combinations. Thus, if one is willing to forgo consistency of the variance estimator itself and instead target conservative inference for linear combinations of the parameter vector, the na{\"i}ve bootstrap provides an alternative route. This conservative-inference perspective is distinct from the goal pursued here, namely consistent estimation of the full asymptotic covariance matrix.$\diamondsuit$

Funding

Honor{\'e} and S{\o}rensen are grateful for research support from the Aarhus Center for Econometrics (ACE) funded by the Danish National Research Foundation grant number DNRF186. Honor{\'e} also acknowledges support from the Gregory C. Chow Econometric Research Program at Princeton University.

singlespace