Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,986 characters · 15 sections · 72 citation commands
Asymptotic Variance Theory for Trimmed Least Squares and Trimmed Least Absolute Deviations in Censored Panel Models with Fixed Effects
\affil[1]{{University of California, Los Angeles}} \affil[2]{{University of Copenhagen}} \affil[3]{{Princeton University}} \affil[4]{{Aarhus Center for Econometrics (ACE)}}
Inference for censored panel-data models with fixed effects often relies on trimmed estimators whose objective functions are convex but nonsmooth. honore_trimmed_1992 introduced two influential examples, the trimmed least squares (TLS) and trimmed least absolute deviations (TLAD) estimators, and derived their asymptotic distributions for a two-period censored regression model with fixed effects. These estimators remain natural benchmark procedures in this setting, both because they avoid parametric restrictions on the fixed effects and because they provide tractable ways to exploit the structure of two-period panel data under censoring. Valid inference for such estimators, however, depends delicately on the form of the Hessian of the expected loss, and the nonsmoothness created by trimming and censoring makes that Hessian less straightforward than the published formulas suggest.
This paper revisits the asymptotic variance theory for the TLS and TLAD estimators of honore_trimmed_1992. We show that the published asymptotic variance formulas rely on additional regularity conditions that are not fully reflected in the original theorem statements. For TLS, the problem is that censoring creates boundary points at which the score of the trimmed square loss is nondifferentiable with positive probability. At such points, differentiation under expectation need not be valid, so the Hessian of the expected loss can differ from the expression obtained by differentiating inside the expectation. For TLAD, the problem is that the expected loss can fail to be twice differentiable when the conditional distribution of the model errors is not sufficiently smooth. Our results identify how these issues enter the asymptotic variance analysis for TLS and TLAD, derive the corrected variance formulas, and clarify the scope of valid legacy inference.
Our starting point is the two-period censored regression model in which the outcome variable $Y_\tau$ in period $\tau\in\{1,2\}$ is generated as
where $\boldsymbol{X}_\tau\in\mathbb{R}^K$ is a vector of time-varying regressors, $\boldsymbol{\theta}_0\in \mathbb{R}^K$ is the parameter of interest, $\alpha$ is an unobserved fixed effect, and $\varepsilon_\tau$ is an unobserved error term. Let $\{(Y_{i1},\boldsymbol{X}_{i1},Y_{i2},\boldsymbol{X}_{i2})\}_{i=1}^n$ be a random sample from the distribution of $(Y_1,\boldsymbol{X}_1,Y_2,\boldsymbol{X}_2)$. honore_trimmed_1992 proposed M-estimators of $\boldsymbol{\theta}_0$ based on a trimmed loss function $m_\Xi:\mathbb{R}\times[0,\infty)\times[0,\infty)\to\mathbb{R}$ defined by
where $\Xi(\cdot)$ represents either the absolute loss $\left|\cdot\right|$ or the (one-half) square loss $\textstyle\frac{1}{2}(\cdot)^2$,\footnote{honore_pairwise_1994 provide a list of conditions allowing other choices of $\Xi$.} and $\xi(\cdot)$ is its derivative.\footnote{In the case of the absolute loss $\Xi(\cdot)=\left|\cdot\right|$, we set $\xi(0) = 0$.} These choices yield the TLAD and TLS estimators, respectively, defined for the corresponding $\Xi(\cdot)$ as any solution
where we introduced the shorthands $\Delta\boldsymbol{X}_i := \boldsymbol{X}_{i1} - \boldsymbol{X}_{i2}$ and $\boldsymbol{Y}_i:=(Y_{i1},Y_{i2})$. Under conditional exchangeability of $(\varepsilon_1,\varepsilon_2)$ given $(\boldsymbol{X}_1,\boldsymbol{X}_2,\alpha)$ and other regularity conditions, honore_trimmed_1992 established consistency and asymptotic normality of both estimators.
Our first set of results concerns the TLS estimator. We show that the Hessian formula appearing in honore_trimmed_1992 is valid only if
a condition not explicitly stated in the original paper. This condition is satisfied, for example, when the index $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ has a continuous distribution, but it fails when all components of $\boldsymbol{\theta}_0$ are equal to zero unless $\Delta\boldsymbol{X} = \boldsymbol{0}$ almost surely. When condition (ref) is not satisfied, the TLS score is nondifferentiable with positive probability, and the published Hessian formula need not equal the derivative of the expected score. We derive the correct Hessian, provide an explicit representation of it, establish asymptotic normality without excluding mass at zero, and obtain a consistent plug-in estimator of the corresponding asymptotic variance.
The TLS analysis also yields a second result. Although the population Hessian formula in honore_trimmed_1992 is not generally correct, the Hessian estimator proposed there remains consistent for the correct TLS asymptotic variance under the maintained assumptions. Thus, the paper does not only identify where the published population formula fails; it also clarifies which parts of the original inference procedure remain usable.
Our second set of results concerns the TLAD estimator. We show, first, that the published asymptotic variance formula omits a conditional-probability term. More importantly, we show that asymptotic normality requires additional continuity conditions to ensure existence of the Hessian of the expected loss. We provide a counterexample satisfying the conditions stated in honore_trimmed_1992 for which that Hessian does not exist. Under additional continuity conditions, we derive the corrected TLAD Hessian and corresponding asymptotic variance formula. We then provide a tuning-parameter-free bootstrap variance estimator for TLAD.
Taken together, these results clarify the role of boundary nondifferentiability in the asymptotic variance theory for censored trimmed estimators. More specifically, they show when legacy inference remains valid, when corrected variance formulas are required, and how to implement inference under the corrected theory.
As a by-product, our TLS analysis also yields a plug-in perspective on the asymptotic variance of the cross-sectional trimmed least squares estimator studied by honore_pairwise_1994, thus avoiding the choice of tuning parameters required by the numerical-differentiation-based variance estimator proposed there.
The remainder of the paper is organized as follows. Section (ref) revisits the TLS estimator, beginning with the asymptotic normality result in honore_trimmed_1992, then identifying the source of the problem, deriving the correct Hessian, establishing asymptotic normality under the corrected theory, and discussing variance estimation, including the status of the legacy Hessian estimator. Section (ref) carries out the analogous analysis for TLAD, derives the corrected Hessian and asymptotic variance formula, and develops bootstrap-based variance estimation. Proofs for the TLS results are in the appendix, while proofs for the TLAD results are in the supplemental appendix, which also contains technical lemmas establishing a generalized dominated convergence theorem, the existence of conditional PDFs, and related measurability issues.
\paragraph{Notation.} For a positive integer $k$, we write $\left[k\right]:=\{1,\ldots,k\}$. We use $\left\|\cdot\right\|_2$ to denote the Euclidean norm. For any (suitably differentiable) real-valued function $f$, whose first argument is a scalar, we let $\dot{f}_1$ and $\ddot{f}_{11}$ denote the first- and second-order partial derivatives of $f$ with respect to its first argument, respectively. For a possibly vector-valued differentiable function $\boldsymbol{g}$, we let $\nabla \boldsymbol{g}$ denote the Jacobian of $\boldsymbol{g}$. If $\boldsymbol{g}$ is itself the gradient mapping corresponding to some function $f$, then we write $\nabla^2 f$ for the Hessian of $f$. The indicator $\mathbf{1}\{A\}$ equals one (zero) if the logical statement $A$ is true (false). We use comma-separated events to denote the intersection of events, e.g. $\mathbf 1\{A,B\} = \mathbf 1\{A\cap B\}$. Unless otherwise stated, all limits are understood as the sample size $n$ grows without bound, holding everything else fixed. The symbols $\rightsquigarrow$, $\to_{\text{a.s.}}$, and $\to_{\mathbb{P}}$ denote convergences in distribution, almost surely, and in probability, respectively. We write $o_{\mathbb{P}}(1)$, $o_{\mathrm{a.s.}}(1)$, and $o_{L^1}(1)$ to denote sequences of random variables that converge to zero in probability, almost surely, and in $L^1(\mathbb{P})$, respectively.
In this section, we focus on the TLS estimator. We abbreviate the trimmed square loss, $m_{\Xi}$ in (ref) with $\Xi=\textstyle{\frac{1}{2}}(\cdot)^2$, by $m^{\texttt{tls}}$, which takes the explicit form
This loss is continuously differentiable in its first argument with partial derivative given by
To facilitate our presentation below, note that it follows from (ref) that $\dot{m}^{\texttt{tls}}_1(\cdot,\boldsymbol{y})$ is Lipschitz continuous on $\mathbb{R}$ with Lipschitz constant one, uniformly in $\boldsymbol{y}\in[0,\infty)\times[0,\infty)$. However, it fails to be differentiable at $-y_2$ and $y_1$ (unless both $y_1$ and $y_2$ are zero), implying that the trimmed square loss is not necessarily twice differentiable. Specifically, the points of second-order non-differentiability of $m^{\texttt{tls}}(\cdot,\boldsymbol{y})$ are captured by set-valued function $N:[0,\infty)\times[0,\infty)\rightrightarrows\mathbb{R}$, defined by
For notational convenience, abbreviate the TLS estimator $\widehat\boldsymbol{\theta}_\Xi$ in (ref) with $\Xi=\textstyle{\frac{1}{2}}\left(\cdot\right)^2$ by $\widehat{\boldsymbol{\theta}}^{\texttt{tls}}$. Also, denote $\boldsymbol{W}:=(\boldsymbol{X}_1,\boldsymbol{X}_2,\alpha)$ and let $\mathcal{W}\subseteq\mathbb{R}^{K}\times \mathbb R^K\times\mathbb R$ be the support of $\boldsymbol{W}$, with $\boldsymbol{w} := (\boldsymbol{x}_1,\boldsymbol{x}_2,a)$ standing for a generic element of $\mathcal{W}$. In addition, denote $\boldsymbol{\varepsilon}:=(\varepsilon_1,\varepsilon_2)$ and let $\mathcal E := \mathbb R\times\mathbb R$, with $\boldsymbol{e} = (e_1,e_2)$ standing for a generic element of $\mathcal E$. honore_trimmed_1992 derived the asymptotic distribution of the TLS estimator under the following assumptions.\footnote{Assumptions (ref)--(ref) are from Assumptions S.2, M.3, E.1, E.3 and R.1 , respectively, in honore_trimmed_1992.}
Let $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$, mapping $\mathcal W\times\mathcal E$ to $[0,\infty)$, be a version of the joint probability density function (PDF) of the pair $\boldsymbol{\varepsilon} = (\varepsilon_1,\varepsilon_2)$ conditional on $\boldsymbol{W} = \boldsymbol{w}$ that is measurable in $(\boldsymbol{w},\boldsymbol{e})$ and is such that $f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_1,e_2) = f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_2,e_1)$ for all $\boldsymbol{w}\in\mathcal W$ and $\boldsymbol{e} = (e_1,e_2)\in\mathcal E$. As we show in Appendix (ref), the function $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$ does exist under Assumptions (ref) and (ref). Also, let $(\boldsymbol{w},\boldsymbol{e})\mapsto F_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e}) = \int_{-\infty}^{e_1}\int_{-\infty}^{e_2}f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(u_1,u_2)\dif u_2\dif u_1$ be the corresponding version of the joint cumulative distribution function (CDF) of $(\varepsilon_1,\varepsilon_2)$ conditional on $\boldsymbol{W} = \boldsymbol{w}$, and let $(\boldsymbol{w},e)\mapsto F_{\varepsilon\mid\boldsymbol{w}}(e) = \lim_{u\to\infty} F_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e,u)$ be the corresponding version of the common marginal CDF of $\varepsilon_1$ and $\varepsilon_2$ conditional on $\boldsymbol{W} = \boldsymbol{w}$.
To state the TLS normality result in honore_trimmed_1992, introduce the $K\times K$ matrices \[ \mathbf{V}_0^{\tt{tls}}:=\mathrm{E}\left[\dot{m}_1^{\tt{tls}}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0,\boldsymbol{Y})^2\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top\right] \] and
honore_trimmed_1992 states that if Assumptions (ref)--(ref) hold, the expectations involved in defining the matrices $\mathbf{V}_0^{\tt{tls}}$ and $\boldsymbol{\Gamma}_0^{\tt{tls}}$ exist (in $\mathbb{R}^{K\times K}$), and both matrices are of full rank, then
with $\mathcal{N}(\boldsymbol{\mu},\boldsymbol{\Sigma})$ denoting the normal distribution with mean $\boldsymbol{\mu}$ and covariance matrix $\boldsymbol{\Sigma}$.
Among several steps, the proof of this result in honore_trimmed_1992 involves:
In our notation, the second task amounts to arguing that the function $\boldsymbol{G}:\mathbb{R}^K\to\mathbb{R}^K$ defined by
is differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$. Its Jacobian $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)\in\mathbb{R}^{K\times K}$ is then the Hessian of the expected loss.\footnote{In honore_trimmed_1992, the function $\boldsymbol{G}$ is denoted $G_4$.} To carry out this task, honore_trimmed_1992 invoked the LDCT to interchange the order of differentiation and expectation, and the matrix $\boldsymbol{\Gamma}_0^{\tt{tls}}$ is the result of that calculation. In the next subsection, however, we will show by example that $\boldsymbol{\Gamma}_0^{\tt{tls}}$ is not always equal to the Jacobian of $\boldsymbol{G}$ at $\boldsymbol{\theta}_0$.
We here provide a DGP which satisfies Assumptions (ref)--(ref), for which the function $\boldsymbol{G}$ is differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$, but the resulting Jacobian $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$ differs from $\boldsymbol{\Gamma}_0^{\tt{tls}}$. To this end, let $K=1$, $\alpha\equiv0$, $\boldsymbol{\theta}_0=0$, $\boldsymbol{X}_1\equiv2$ and $\boldsymbol{X}_2\equiv1$, and let $\varepsilon_1$ and $\varepsilon_2$ be independent of each other and standard normally distributed. Then the outcome variables $Y_{\tau}=\max\{0,\varepsilon_{\tau}\},\tau\in\{1,2\}$, are independent and identically distributed as standard normals censored from below at zero. Case by case inspection shows that Assumptions (ref)--(ref) are satisfied and that (the here scalar) $\mathbf{V}_0^{\tt{tls}}$ and $\boldsymbol{\Gamma}_0^{\tt{tls}}$ are both well-defined and of full rank (non-zero).
In this example DGP, the regressor difference $\Delta \boldsymbol{X} = \boldsymbol{X}_1 - \boldsymbol{X}_2$ is identically one, so the (here scalar-valued) function $\boldsymbol{G}$ simplifies to $G(\theta)=\mathrm{E}[\dot{m}^{\texttt{tls}}_1(\theta,\boldsymbol{Y})],\theta\in\mathbb{R}$. It is then straightforward to check that the function $G$ is given by \[ G(\theta) =
\] where $\Phi$ is the standard normal CDF. For example, if $\theta < 0$, then letting $\varphi$ denote the standard normal PDF,
where the first equality follows from (ref), the second from (the here) identical distributions of $Y_1$ and $Y_2$, the third from noting that $Y_1$ is a standard normal random variable truncated from below at zero, and the fourth from integration by parts.
The graph of $G$ is shown in Figure (ref).
As indicated by the figure, and immediately follows analytically, $G$ is differentiable at $\theta_0(=0)$, with derivative $\dot{G}(0)=1 - \Phi(0)=\textstyle{\frac{1}{2}}$. However, since $Y_1$ and $Y_2$ are independently distributed as standard normals censored from below at zero, (the here scalar) $\boldsymbol{\Gamma}_0^{\tt{tls}}$ equals \[ \Gamma_0^{\tt{tls}}=\mathrm{E}[\mathbf{1}\{Y_1>0>-Y_2\}]=\mathbb{P}(Y_1>0)\cdot\mathbb{P}(Y_2>0)=\frac{1}{4}, \] showing that $\dot{G}(0)\neq\Gamma_0^{\tt{tls}}$.
How does this discrepancy arise? When arguing differentiability of the function $\boldsymbol{G}(\boldsymbol{\theta}) = \mathrm{E}\sbr[1]{\dot{m}_1^\texttt{tls}\del[0]{\Delta\boldsymbol{X}^\top\boldsymbol{\theta},\boldsymbol{Y}}\Delta\boldsymbol{X}}$ at $\boldsymbol{\theta} = \boldsymbol{\theta}_0$, honore_trimmed_1992 called upon the LDCT to interchange the order of differentiation and expectation. However, as captured by $N(\boldsymbol{y})$ in (ref), the function $\dot{m}^{\texttt{tls}}_1(\cdot,\boldsymbol{y})$ fails to be differentiable at $-y_2$ and $y_1$, unless both $y_1$ and $y_2$ are zero. In turn, in a model with censoring, one may see exactly one outcome equal to zero with (strictly) positive probability. Thus, whenever the inner product $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ places mass at zero, the (random) function $t\mapsto \dot{m}_1^\texttt{tls}(t,\boldsymbol{Y})$ fails to be differentiable at $t=\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta}_0$ with (strictly) positive probability, invalidating the use of the LDCT.
In this subsection, we modify the argument in honore_trimmed_1992 to allow the distribution of $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ to place mass at zero. We show that Honor{\'e}'s assumptions, corresponding to our Assumptions (ref)--(ref), imply differentiability of $\boldsymbol{G}$ at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$. We also derive an explicit expression for the Jacobian $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$, which coincides with the Hessian of the expected loss. Our calculations further show that $\nabla\boldsymbol{G}(\boldsymbol{\theta}_0)$ generally differs from $\boldsymbol{\Gamma}_0^{\tt{tls}}$, unless mass at zero in the distribution of $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ is ruled out a priori. By deriving equivalent representations of this Hessian, we obtain easily interpretable necessary and sufficient conditions for its invertibility. We then state an asymptotic normality result that allows $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ to have mass at zero.
The expressions in (ref) and (ref) make it clear that $(Y_1,\boldsymbol{X}_1)$ and $(Y_2,\boldsymbol{X}_2)$ enter the Hessian $\mathbf{J}^{\tt{tls}}_0$ in a symmetric manner, so that the (time) labeling is irrelevant. The version in (ref) facilitates plug-in estimation of $\mathbf{J}^{\tt{tls}}_0$, which we cover in Section (ref). In turn, the version in (ref) facilitates comparison with Honor{\'e}'s $\boldsymbol{\Gamma}_0^{\tt{tls}}$. To this end, iterate expectations to write the latter as \[ \boldsymbol{\Gamma}_0^{\tt{tls}}=\mathrm{E}\sbr[2]{\mathrm{E}\sbr[1]{\mathbf{1}\cbr[0]{-Y_2<\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<Y_1}\mid\boldsymbol{W}}\Delta\boldsymbol{X}\Delta\boldsymbol{X}^\top}. \] We next expand on the inner expectation. Rewrite the inner expectation indicator as \[ 1-\mathbf{1}\{ Y_{1}\leqslant\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} -\mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} +\mathbf{1}\{ Y_{1}\leqslant \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} \mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\}. \] Because $Y_{1}$ and $Y_{2}$ are non-negative, the only way both indicators appearing in the previous display can be turned on is if $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0$. Again using non-negativity, we deduce \[ \mathbf{1}\{ Y_{1}\leqslant \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} \mathbf{1}\{ Y_{2}\leqslant-\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\} =\mathbf{1}\{ \Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0\} \mathbf{1}\{ Y_{1}=0\} \mathbf{1}\{ Y_{2}=0\}, \] from which it follows that $\mathbf{1}\{-Y_2<\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0<Y_1\}$ can be written as
Taking conditional expectations, the inner expectation in $\boldsymbol{\Gamma}_0^{\tt{tls}}$ works out as
where (due to space constraints) we introduced the shorthand notations $V_1:=-(\alpha+\boldsymbol{X}_1^\top\boldsymbol{\theta}_0)$ and $V_2:=-(\alpha+\boldsymbol{X}_2^\top\boldsymbol{\theta}_0)$. Since $V_2-V_1=\boldsymbol{X}_1^\top\boldsymbol{\theta}_0-\boldsymbol{X}_2^\top\boldsymbol{\theta}_0=\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$, we have $V_1\gtreqqless V_2$ if $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\lesseqqgtr0$. With this newly introduced shorthand notation, the inner expectation in our Hessian (ref) can be written as $ 1-F_{\varepsilon\mid\boldsymbol{W}}\del[0]{\max\cbr{V_1,V_2}}, $ which agrees with the previous display whenever $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\neq0$. Hence, if this case can be ruled out [i.e., $\mathbb{P}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0)=0$], then Honor{\'e}'s $\boldsymbol{\Gamma}_0^{\tt{tls}}$ is equal to $\mathbf{J}^{\tt{tls}}$.
However, when $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0$, so that $V_1$ and $V_2$ take a common value (here: $V$), the inner expectations differ by
which is non-negative by monotone increasingness of CDFs. In the Section (ref) counterexample, we have both $V_{1}\equiv0$ and $V_{2}\equiv0$, and the (conditional) joint CDF factors into the product of the (standard normal) marginals, $F_{\boldsymbol{\varepsilon}\mid\boldsymbol{W}}(u_{1},u_{2})=\Phi(u_{1})\Phi(u_{2})$. Since there $\mathbb{P}(\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0=0)=1$, the difference in the previous display captures the only relevant case. The difference is then $\Phi(0)-\Phi(0)\Phi(0)=\frac{1}{4}$, which is precisely the previously demonstrated discrepancy.
To state the asymptotic normality result, we need to make sure that the Hessian $\mathbf{J}^{\tt{tls}}_0$ is invertible. To this end, we will impose the following condition:
This assumption is easy to interpret and implies via (ref) that the matrix $\mathbf{J}^{\tt{tls}}_0$ is invertible. Alternatively, we could state an assumption based on the expression for $\mathbf{J}^{\tt{tls}}_0$ in (ref). Strictly speaking, Assumption (ref) does not appear in honore_trimmed_1992 but the only purpose of this assumption is to ensure the invertibility of the Hessian $\mathbf{J}^{\tt{tls}}_0$, and honore_trimmed_1992 did impose invertibility of the corresponding Hessian $\boldsymbol{\Gamma}_0^{\tt{tls}}$.
Using our new understanding of the Hessian $\mathbf{J}^{\tt{tls}}_0$, we next state an extended asymptotic normality result for the TLS estimator that does not require the inner product $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ to have no mass at zero.
The asymptotic normality result in Theorem (ref) differs from honore_trimmed_1992 in the sense that it replaces the matrix $\boldsymbol{\Gamma}_0^{\tt{tls}}$ in (ref) by $\mathbf{J}_0^{\tt{tls}}$. Whenever $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0\neq 0$ with probability one, the two matrices coincide. However, the matrices are in general different if $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0 = 0$ with (strictly) positive probability.
For the asymptotic normality to yield a practical approximation, we need to consistently estimate the asymptotic variance components. For $\mathbf{V}_0^{\tt{tls}}$, we use the plug-in estimator from honore_trimmed_1992: \[ \widehat{\mathbf{V}}^{\tt{tls}}:=\frac{1}{n}\sum_{i=1}^n \dot{m}_1^{\tt{tls}}(\Delta\boldsymbol{X}_i^\top\widehat\boldsymbol{\theta}^{\tt{tls}},\boldsymbol{Y}_i)^2\Delta\boldsymbol{X}_i\Delta\boldsymbol{X}_i^\top.\footnote{In \citet{honore_trimmed_1992}, this estimator is denoted $\hat{V}_4$.} \] For the Hessian $\mathbf{J}_0^{\tt{tls}}$, we apply the analogy principle to (ref) to arrive at
Alternatively, in order to estimate the Hessian $\mathbf{J}_0^{\tt{tls}}$, we could use the equivalent expressions in (ref) and (ref).
These estimators are (strongly) consistent under the assumptions of Theorem (ref):
Theorem (ref) implies that the limit variance $(\mathbf{J}_0^{\tt{tls}})^{-1}\mathbf{V}_0^{\tt{tls}}(\mathbf{J}_0^{\tt{tls}})^{-1}$ is (strongly) consistently estimated by $(\widehat\mathbf{J}^{\tt{tls}})^{-1}\widehat\mathbf{V}^{\tt{tls}}(\widehat\mathbf{J}^{\tt{tls}})^{-1}$, which facilitates hypothesis testing and the construction of confidence intervals.
In honore_trimmed_1992 the matrix $\boldsymbol{\Gamma}_0^{\tt{tls}}$ in (ref) is estimated by its empirical analogue
The TLS sandwich variance consistency results in honore_trimmed_1992 involve showing that $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ is a consistent estimator of $\boldsymbol{\Gamma}_0^{\tt{tls}}$. As we have shown by example in Section (ref), $\boldsymbol{\Gamma}_0^{\tt{tls}}$ is not equal to the Hessian $\mathbf{J}_0^{\tt{tls}}$ of the expected loss in general. As $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ is targeting $\boldsymbol{\Gamma}_0^{\tt{tls}}$, one would expect that $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ is not a consistent estimator of the TLS Hessian $\mathbf{J}_0^{\tt{tls}}$ in general. Surprisingly, as we establish below, $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ nevertheless converges to the TLS Hessian $\mathbf{J}_0^{\tt{tls}}$.
This counterintuitive result can be explained as follows. Decompose $\widehat\boldsymbol{\Gamma}^{\tt{H92}}$ into two parts:
Although $\widehat\mathbf{L}^{\tt{H92}}$ and $\widehat\mathbf{J}^{\tt{tls}}$ are not identical term by term, the leading component $\widehat\mathbf{L}^{\tt{H92}}$ is closely related to our Hessian estimator in (ref), which is the empirical analogue of $\mathbf{J}_0^{\tt{tls}}$. The proof of Theorem (ref) shows that $\widehat\mathbf{L}^{\tt{H92}}$ converges to $\mathbf{J}_0^{\tt{tls}}$, while the remainder term $\widehat\mathbf{R}^{\tt{H92}}$ converges to the zero matrix. Key to these findings is our extended asymptotic normality result in Theorem (ref). Specifically, even if $\Delta\boldsymbol{X}^\top\boldsymbol{\theta}_0$ has mass at zero, the absolutely continuous limiting distribution of $\sqrt n(\widehat\boldsymbol{\theta}^{\tt{tls}}-\boldsymbol{\theta}_0)$ implies that, for any fixed $\boldsymbol{x}\in\mathbb{R}^K\backslash\{\boldsymbol{0}\}$, $\mathbb{P}(\boldsymbol{x}^\top\widehat\boldsymbol{\theta}^{\tt{tls}}=0)\to0$. We stress that Theorem (ref) is not a simple consequence of consistency of $\widehat\boldsymbol{\theta}^{\tt{tls}}$, established in honore_trimmed_1992. Indeed, even with a uniform law of large numbers allowing us to replace sample averages in $\widehat{\boldsymbol{\Gamma}}^{\tt{H92}}$ by their population counterparts (see the proof of Theorem (ref)), the resulting map $\boldsymbol{\theta}\mapsto\boldsymbol{\Gamma}(\boldsymbol{\theta})$ need not be continuous at $\boldsymbol{\theta}_0$. Consequently, a continuous mapping argument need not apply. In the counterexample of Section (ref), one finds $\Gamma(\theta)=[1-\Phi(|\theta|)]\boldsymbol{1}\{\theta\neq0\}+\tfrac{1}{4}\boldsymbol{1}\{\theta=0\}$, which has a jump discontinuity at $\theta=0=\theta_0$. The jump size $(\tfrac{1}{4})$ is exactly the difference between $\mathbf{J}_0^{\tt{tls}}$ and $\boldsymbol{\Gamma}_0^{\tt{tls}}$ in that example.
A by-product of Theorem (ref) is that, under our maintained assumptions, the TLS Hessian consistency statement in honore_trimmed_1992 does not hold without additional regularity conditions ensuring continuity at $\boldsymbol{\theta}_0$.
In this section, we focus on the TLAD estimator. We abbreviate the trimmed absolute loss, $m_{\Xi}$ in (ref) with $\Xi=\left|\cdot\right|$, by $m^{\texttt{tlad}}$, which takes the form
For notational convenience, abbreviate the TLAD estimator $\widehat\boldsymbol{\theta}_\Xi$ in (ref) with $\Xi=\left|\cdot\right|$ by $\widehat{\boldsymbol{\theta}}^{\texttt{tlad}}$. Also, let $\boldsymbol{W}$, $\mathcal W$, $\boldsymbol{w}$, $\boldsymbol{\varepsilon}$, $\mathcal E$, and $\boldsymbol{e}$ be the same as in Section (ref). Consider the following assumptions.\footnote{Assumptions (ref)--(ref) are from Assumptions S.2, M.2, E.1, E.3, E.4 and R.1, respectively, in honore_trimmed_1992.}
Since Assumptions (ref) and (ref) are the same as Assumptions (ref) and (ref), following the reasoning in Section (ref), there exists a function $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$, mapping $\mathcal W\times\mathcal E$ to $[0,\infty)$, that is a version of the PDF of the pair $\boldsymbol{\varepsilon} = (\varepsilon_1,\varepsilon_2)$ conditional on $\boldsymbol{W} = \boldsymbol{w}$, which is measurable in $(\boldsymbol{w},\boldsymbol{e})$, and is such that $f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_1,e_2) = f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e_2,e_1)$ for all $\boldsymbol{w}\in\mathcal W$ and $\boldsymbol{e} = (e_1,e_2)\in\mathcal E$. Also, let $(\boldsymbol{w},e)\mapsto f_{\varepsilon\mid\boldsymbol{w}}(e) = \int_{\mathbb R}f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(e,u)\dif u$ be the corresponding version of the common marginal PDF of $\varepsilon_1$ and $\varepsilon_2$ conditional on $\boldsymbol{W} = \boldsymbol{w}$ and let $(\boldsymbol{w},e)\mapsto f_{\varepsilon_1 - \varepsilon_2\mid\boldsymbol{w}}(e) = \int_{\mathbb R} f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(u+e,u)\dif u$ be the corresponding version of the PDF of the difference $\varepsilon_1 - \varepsilon_2$ conditional on $\boldsymbol{W} = \boldsymbol{w}$.
Assumptions (ref)--(ref) are the same as the corresponding assumptions in honore_trimmed_1992. Assumption (ref), however, is not present in honore_trimmed_1992. We consider this assumption because the asymptotic normality result in honore_trimmed_1992 may not hold without it. In particular, in Section (ref) below we provide a DGP that satisfies Assumptions (ref)--(ref) and is such that the Hessian of the expected loss [see (ref) below], which appears in the asymptotic variance formula in honore_trimmed_1992, does not exist. We note also that Assumption (ref) may be stronger than necessary. For example, it seems possible to obtain the asymptotic normality result assuming continuity of the functions in Assumption (ref) only on their respective supports but we opt for a stronger than necessary conditions for clarity of the argument.
Note also that Assumptions (ref), (ref), (ref), and (ref) are the same as the corresponding assumptions for the TLS estimator (Assumptions (ref), (ref), (ref), and (ref), respectively). The TLAD moment conditions in Assumption (ref) are weaker than those for the TLS estimator (Assumption (ref)), which reflects the fact that the trimmed absolute loss is less sensitive to outliers than the trimmed square loss. Assumptions (ref) and (ref), requiring certain PDFs to be bounded and continuous, do not have analogs in the case of the TLS estimator.
To state the TLAD normality result in honore_trimmed_1992, introduce the $K\times K$ matrices\footnote{In honore_trimmed_1992, these matrices are denoted $V_3$ and $\Gamma_3$, respectively. Honor{\'e}'s $\Gamma_3$ is stated in terms of probabilities and densities of the censored $Y_1$ and $Y_2$, but it is clear from the underlying proof that he means the latent variables $Y_1^\ast$ and $Y_2^\ast$, respectively.}
and
where $(\boldsymbol{w},e) \mapsto f_{Y_1^\ast-Y_2^\ast\mid\boldsymbol{w},Y_1^\ast>0,Y_2^\ast>0}\left(e\right)$ is the conditional PDF of $Y_1^\ast-Y_2^\ast$ given $\boldsymbol{W}=\boldsymbol{w}$ and $\{Y_1^\ast>0\}\cap \{Y_2^\ast>0\}$, $(\boldsymbol{w},e) \mapsto f_{Y^\ast_1\mid\boldsymbol{w},Y_2^\ast\leqslant0}\left(e\right)$ is the conditional PDF of $Y_1^\ast$ given $\boldsymbol{W}=\boldsymbol{w}$ and $Y_2^\ast\leqslant0$, and $(\boldsymbol{w},e)\mapsto f_{Y^\ast_2\mid\boldsymbol{w},Y_1^\ast\leqslant0}\left(e\right)$ is the conditional PDF of $Y_2^\ast$ given $\boldsymbol{W}=\boldsymbol{w}$ and $Y_1^\ast\leqslant0$.\footnote{Here and below, we assume that the versions of all conditional PDFs and CDFs including latent outcomes $Y_1^*$ and $Y_2^*$ are obtained (fixed) by combining the conditional PDF $(\boldsymbol{w},\boldsymbol{e})\mapsto f_{\boldsymbol{\varepsilon}\mid\boldsymbol{w}}(\boldsymbol{e})$ with (ref).}
honore_trimmed_1992 states that if Assumptions (ref)--(ref) hold, the expectations involved in defining the matrices $\mathbf{V}_0^{\tt{tlad}}$ and $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ exist (in $\mathbb{R}^{K\times K}$), and both matrices are of full rank, then
Among several steps, the proof of this result in honore_trimmed_1992 includes establishing the existence of the Hessian of the expected loss [see (ref) below] at the true parameter value $\boldsymbol{\theta}_0$. In our notation, this task corresponds to arguing that the function $L:\mathbb{R}^K\to\mathbb{R}$ defined by
is twice differentiable at $\boldsymbol{\theta}=\boldsymbol{\theta}_0$.\footnote{Note that because the function $m^{\texttt{tlad}}(t,\boldsymbol{y})$ is Lipschitz continuous in its first argument, it follows from Assumption (ref) that the expectation in (ref) is well-defined. Also note that, as in honore_trimmed_1992, we work with the expected loss difference $\mathrm{E}[m^{\texttt{tlad}}(\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta},\boldsymbol{Y}) - m^{\texttt{tlad}}(0,\boldsymbol{Y})]$ instead of $\mathrm{E}[m^{\texttt{tlad}}(\Delta\boldsymbol{X}^{\top}\boldsymbol{\theta},\boldsymbol{Y})]$ because the former loss gives results under weaker regularity conditions.} To this end, Honor{\'e} used the LDCT to differentiate once under the expectation, and then argued differentiability at $\boldsymbol{\theta}_0$ of the resulting function to arrive at $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ in (ref). In the next subsection, however, we will show by example that Assumptions (ref)--(ref) are not sufficient to ensure the existence of the Hessian.
As in the Section (ref) counterexample, let $K=1$, $\alpha\equiv0$, $\boldsymbol{\theta}_0 = 0$, $\boldsymbol{X}_1\equiv2$ and $\boldsymbol{X}_2\equiv1$. Also, to define the distribution of the pair $(\varepsilon_1,\varepsilon_2)$, let $r:\mathbb R\to \mathbb R$ be a continuous function such that (i) $r(t)\geqslant 0$ for all $t\in\mathbb R$, (ii) $\int_{\mathbb R} r(t)\dif t = 1$, and (iii) $r(t)=0$ if $t\leqslant 1$ or $t\geqslant 3$. In addition, let $\mathscr E := \{0,2,4,\dots\}$, $\mathscr O = \{1,3,5,\dots\}$, and let $\tilde h:[0,1]\to\{0,1\}$ be the function defined by $$ \tilde{h}(t)=
$$ Moreover, let $h:\mathbb R \to\mathbb R$ be the function defined by $h(t) = 3\tilde h(|t|) / 4$ for $t\in[-1,1]$ and $0$ otherwise. Note that $h(t)\geqslant0$ for all $t\in\mathbb R$ and $$ \int_{\mathbb R} h(t)\dif t = \frac{3}{4}\int_{-1}^1 \tilde h(\left|t\right|)\dif t = \frac{3}{2}\int_{0}^1 \tilde h(t)\dif t = \frac{3}{2}\sum_{k\in\mathscr E}\frac{1}{2^{k+1}} = \frac{3}{2}\cdot \frac{1}{2}\cdot\frac{1}{1-\frac{1}{4}} = 1. $$ Thus, the function $f\colon\mathbb R^2 \to \mathbb R$ defined by $$ f(e_1,e_2) = 2h(e_1 - e_2)r(e_1 + e_2),\quad e_1,e_2\in\mathbb R, $$ satisfies $f(e_1,e_2)\geqslant 0$ for all $e_1,e_2\in\mathbb R$ and
Hence, $f$ is the PDF of a certain distribution on $\mathbb R^2$. Let $(\varepsilon_1,\varepsilon_2)$ be a pair of random variables sampled from this distribution. Because of the symmetry of the function $f$, the random variables $\varepsilon_1$ and $\varepsilon_2$ are then exchangeable, and their common PDF is $$ f_{\varepsilon}(t) = 2\int_{\mathbb R}h(t - e_2)r(t+e_2)\dif e_2 = 2\int_{\mathbb R} h(u)r(u+2t)\dif u,\quad t\in\mathbb R, $$ which implies that $0<\varepsilon_1<2$ and $0<\varepsilon_2<2$ with probability one. Moreover, the PDF of the difference $\varepsilon_1 - \varepsilon_2$ is
It is then straightforward to check that Assumptions (ref)--(ref) are all satisfied. Note, however, that Assumption (ref) is {\em not} satisfied because the function $f$ is not continuous.
Now, observe that since $\mathbb P(Y_1>0, Y_2>0) = 1$ by construction, a calculation shows that the expected loss function $L$ in (ref) simplifies to $$ L(\theta) = \mathrm{E}\left[m^{\texttt{tlad}}(\theta,\boldsymbol{Y}) - m^{\texttt{tlad}}(0,\boldsymbol{Y})\right] = \mathrm{E}\left[|Y_1 - Y_2 - \theta| - |Y_1 - Y_2|\right],\quad \theta\in\mathbb R. $$ Since the integrand here is convex in $\theta$ and non-differentiable in $\theta$ only at $Y_1 - Y_2$, which, for a given $\theta$, happens with probability zero, it follows from bertsekas1973stochastic that $L$ is (everywhere) differentiable with derivative $$ \dot{L}(\theta) = \mathrm{E}\left[2\cdot\mathbf{1}\{Y_1 - Y_2 \leqslant \theta\} - 1\right] = 2F_{\varepsilon_1 - \varepsilon_2}(\theta) - 1, $$ where $F_{\varepsilon_1 - \varepsilon_2}$ denotes the CDF of the difference $\varepsilon_1 - \varepsilon_2$. We claim that $F_{\varepsilon_1 - \varepsilon_2}$ is non-differentiable at zero $(=\theta_0)$. Indeed, if $t_k = 2^{-k}$ for $k\in\mathscr E$, then $$ \frac{F_{\varepsilon_1 - \varepsilon_2}(t_k) - F_{\varepsilon_1 - \varepsilon_2}(0)}{t_k} = 2^k \int_0^{2^{-k}} h(t)\dif t = \frac{3}{4}\cdot 2^k \cdot \sum_{l\in\mathscr E : l\geqslant k}\frac{1}{2^{l+1}} = \frac{1}{2} $$ and if $t_k = 2^{-k}$ for $k\in\mathscr O$, then $$ \frac{F_{\varepsilon_1 - \varepsilon_2}(t_k) - F_{\varepsilon_1 - \varepsilon_2}(0)}{t_k} = 2^k \int_0^{2^{-k}} h(t)\dif t = \frac{3}{4}\cdot 2^k \cdot \sum_{l\in\mathscr E : l\geqslant k+1}\frac{1}{2^{l+1}} = \frac{1}{4}, $$ which implies that the limit of $(F_{\varepsilon_1 - \varepsilon_2}(t) - F_{\varepsilon_1 - \varepsilon_2}(0))/t$ as $t\to 0$ does not exist. Thus, the function $\dot{L}$ is non-differentiable at zero ($=\theta_0$), implying that the Hessian of $L$ at $\theta_0$ does not exist.
One can show that the TLAD estimator is still $\sqrt{n}$-consistent but not (asymptotically) normal. Specifically, in Figure (ref), we show the (kernel) PDF of the scaled TLAD estimates $\sqrt{n}\widehat{\theta}^{\tt{tlad}}$ based on $10,000$ Monte Carlo samples of size $n=50,000$ using the DGP defined earlier in this section.\footnote{The kernel density was created using the R package ggplot2 with geom_density. We use a Gaussian kernel and the silverman_density_1986 rule-of-thumb bandwidth (both geom_density defaults).} To facilitate comparison, we include the PDF resulting from fitting a Gaussian distribution to the Monte Carlo dataset of scaled estimates using maximum likelihood. The PDF of the scaled estimates looks far from Gaussian. We therefore conclude that establishing the asymptotic normality result for the TLAD estimator requires more than Assumptions (ref)--(ref). We do so by imposing the continuity conditions in Assumption (ref).
In this subsection, we provide a Hessian existence argument for the TLAD estimator using Assumptions (ref)--(ref). We also show that, unless the pair of latent outcomes $(Y_1^\ast,Y_2^\ast)$ belongs to the first quadrant with conditional probability given $\boldsymbol{W}$ equal to one almost surely, $\boldsymbol{\Gamma}_0^{\tt{tlad}}$ will not be the Hessian of the expected loss (because of an apparent typographical error). We then state the asymptotic normality result under Assumptions (ref)--(ref) with the corrected Hessian.
To state the existence theorem, let $(\boldsymbol{w},\boldsymbol{y})\mapsto f_{\boldsymbol{Y}^\ast\mid\boldsymbol{w}}\left(\boldsymbol{y}\right)$ denote the joint PDF of the latent outcomes $Y_1^\ast$ and $Y_2^\ast$ conditional on $\boldsymbol{W}=\boldsymbol{w}$ for all $\boldsymbol{w}\in\mathcal{W}$.
We now revise the asymptotic normality statement in (ref) for the TLAD estimator based on our new understanding of the Hessian $\mathbf{H}_0^{\tt{tlad}}$.
We now discuss estimation of the asymptotic variance $\boldsymbol{\Sigma}_0^{\tt{tlad}}:=(\mathbf{H}_0^{\tt{tlad}})^{-1}\mathbf{V}_0^{\tt{tlad}}(\mathbf{H}_0^{\tt{tlad}})^{-1}$ in (ref). Since the formula for $\mathbf{H}_0^{\tt{tlad}}$ given in (ref) includes a nonparametric conditional PDF, we propose a bootstrap estimator instead of a plug-in estimator. Also, since a na{\"i}ve bootstrap variance estimator may fail to be consistent for LAD-type estimators without additional integrability conditions BF81,GPSB84, we instead construct a quantile-based robust bootstrap variance estimator that avoids reliance on bootstrap second moments.
The robust bootstrap variance estimator exploits the fact that quantiles of linear combinations of a normal distribution scale with their standard deviations. We therefore estimate bootstrap quantiles of suitably centered parameter combinations and invert this scaling to recover the corresponding covariance entries. This construction replaces unstable bootstrap variance estimation based on second moments by identification of covariance entries through the asymptotic normal geometry of the estimator.
To describe the robust bootstrap variance estimator in more detail, consider a bootstrap sample $\{(\widetilde{Y}_{i1},\widetilde{\boldsymbol{X}}_{i1},\widetilde{Y}_{i2},\widetilde{\boldsymbol{X}}_{i2})\}_{i=1}^n$. obtained by sampling $n$ observation indices with replacement from $\{1,2,\dots,n\}$ and extracting the corresponding observations from the original sample. Also, let
be the bootstrap analog of $\widehat\boldsymbol{\theta}^{\tt{tlad}} = (\widehat\theta^{\tt{tlad}}_1,\dots,\widehat\theta^{\tt{tlad}}_K)^\top$, where we introduced shorthands $\Delta\widetilde{\boldsymbol{X}}_i := \widetilde{\boldsymbol{X}}_{i1} - \widetilde{\boldsymbol{X}}_{i2}$ and $\widetilde{\boldsymbol{Y}}_i:=(\widetilde{Y}_{i1},\widetilde{Y}_{i2})$. Identification of the covariance entries relies on the identities $\mathrm{var}(Z_k)=\Sigma^{\tt{tlad}}_{0kk}$ and $\mathrm{var}(Z_j+Z_k)=\Sigma^{\tt{tlad}}_{0jj} +\Sigma^{\tt{tlad}}_{0kk} +2\Sigma^{\tt{tlad}}_{0jk}$ for a normal vector $\boldsymbol{Z} \sim\mathcal{N}(\boldsymbol{0},\boldsymbol{\Sigma}_0^{\tt{tlad}})$. Accordingly, we recover variances and covariances from bootstrap quantiles of the corresponding linear combinations. Motivated by these identities, we estimate the diagonal and off-diagonal entries of $\boldsymbol{\Sigma}_0^{\tt{tlad}}=[\Sigma^{\tt{tlad}}_{0jk}]_{j,k=1}^K$ separately. For the diagonal elements, for each $k\in[K]$, we set $\widehat\Sigma_{kk}^{\tt{tlad}}:=[\widehat{q}_{0.9,k}/\Phi^{-1}(0.95)]^2$, where $\widehat{q}_{0.9,k}$ is the $0.9$ quantile of the conditional distribution of $\sqrt n|\widetilde\theta^{\tt{tlad}}_k - \widehat\theta^{\tt{tlad}}_k|$ given the original data and $\Phi^{-1}(0.95)$ is the $0.95$ quantile of the standard normal distribution. For the off-diagonal elements, for each $(j,k)\in[K]\times[K]$ such that $j\neq k$, we set $\widehat \Sigma_{jk}^{\tt{tlad}}:= ([\widehat{q}_{0.9,j,k}/\Phi^{-1}(0.95)]^2 - \widehat\Sigma_{jj}^{\tt{tlad}} - \widehat\Sigma_{kk}^{\tt{tlad}})/2$, where $\widehat{q}_{0.9,j,k}$ is the $0.9$ quantile of the conditional distribution of $\sqrt n|\widetilde\theta^{\tt{tlad}}_j + \widetilde\theta^{\tt{tlad}}_k - \widehat\theta^{\tt{tlad}}_j - \widehat\theta^{\tt{tlad}}_k|$ given the original data.\footnote{As the robust bootstrap covariance estimator $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ is recovered entrywise from quantiles of linear combinations, it need not be positive semidefinite in finite samples. One can construct a positive semidefinite variance estimator by projecting $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}$ onto the set of positive semidefinite matrices.} In the following theorem, we prove consistency of $\widehat\boldsymbol{\Sigma}^{\tt{tlad}}:=[\widehat\Sigma_{jk}^{\tt{tlad}}]_{j,k=1}^K$ under the assumptions of Theorem (ref).
Honor{\'e} and S{\o}rensen are grateful for research support from the Aarhus Center for Econometrics (ACE) funded by the Danish National Research Foundation grant number DNRF186. Honor{\'e} also acknowledges support from the Gregory C. Chow Econometric Research Program at Princeton University.