EconBase
← Back to paper

Minimum Distance Estimation of Quantile Panel Data Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

449,989 characters · 34 sections · 116 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
This text was truncated for display. The citation measures were computed over the complete text.

Minimum Distance Estimation of Quantile Panel Data Models

frontmatter\runtitle{Minimum Distance Quantile Regression} \begin{aug} \address[add1]{ \orgdiv{Department of Economics}, \orgname{University of Bern}} \address[add2]{ \orgdiv{Department of Economics}, \orgname{University of Bern}} \end{aug} \begin{funding} We are grateful to Manuel Arellano, Bo Honoré, Aleksey Tetenov, and seminar participants at Bern University, Geneva University, UC Irvine, Fribourg University, the 27th International Panel Data Conference, the COMPIE 2022 Conference, the 2nd International Econometrics PhD Conference at Erasmus University Rotterdam, the 2023 Ski and Labor Seminar, the 2023 Young Swiss Economists Meeting, the 2023 IAAE Conference, the 2023 Summer Meeting of the Econometric Society, and the third Causal Inference Optimization-Conscious Econometrics Conference at the University of Chicago for useful comments. \end{funding} \begin{abstract} We propose a minimum distance estimation approach for quantile panel data models where unit effects may be correlated with covariates. This computationally efficient method involves two stages: first, computing quantile regression within each unit, then applying GMM to the first-stage fitted values. Our estimators apply to (i) classical panel data, tracking units over time, and (ii) grouped data, where individual-level data are available, but treatment varies at the group level. Depending on the exogeneity assumptions, this approach provides quantile analogs of classic panel data estimators, including fixed effects, random effects, between, and Hausman-Taylor estimators. In addition, our method offers improved precision for grouped (instrumental) quantile regression compared to existing estimators. We establish asymptotic properties as the number of units and observations per unit jointly diverge to infinity. Additionally, we introduce an inference procedure that automatically adapts to the potentially unknown convergence rate of the estimator. Monte Carlo simulations demonstrate that our estimator and inference procedure perform well in finite samples, even when the number of observations per unit is moderate. In an empirical application, we examine the impact of the food stamp program on birth weights. We find that the program's introduction increased birth weights predominantly at the lower end of the distribution, highlighting the ability of our method to capture heterogeneous effects across the outcome distribution. \end{abstract} \begin{keyword} \kwd{Quantile regression} \kwd{panel data} \kwd{grouped data} \end{keyword}

Introduction

Quantile regression, introduced by Koenker1978, is a powerful tool for analyzing the effect of policies on the distribution of an outcome variable. The quantile treatment effect function provides more information than the average treatment effect, allowing, for instance, evaluation of the treatment's impact on inequality. When panel data are available, new identification and estimation strategies become feasible. Researchers can alleviate endogeneity concerns, for instance, by allowing for correlated group effects. They can obtain more precise estimates using a random-effects estimator or exploit individual-level variables to identify the impact of group-level variables, e.g., with the Hausman1981 estimator.

We define panel data as a dataset structure where observations are organized along two dimensions. While classical panel data typically consist of individuals repeatedly measured over time, our results extend to group data, where individual-level observations are available, and treatments often vary at the group level. For example, Autor2013a use commuting zones in a given decade as groups, while Angrist2004 employ schools. In both cases, treatments vary only between groups, but individual data are essential for estimating the conditional distribution of outcomes within each group. This paper adopts a general notation ($i$ and $j$ subscripts) applicable to both classical panel and group data. We primarily use terminology (individuals and groups) that is more common to group data, as our application falls into this category.

As a first contribution of the paper, we propose a new class of minimum distance estimators for quantile panel data models. This class of estimators provides quantile analogs of established panel data methods, including fixed effects, random effects, between, and Hausman-Taylor estimators. Our estimation approach involves two stages. The first stage consists of group-level quantile regressions using individual-level covariates. In the second stage, we regress the fitted values from the first stage on individual-level and group-level variables. If these variables are potentially endogenous, an instrumental variable regression or, more generally, the generalized method of moments (GMM) estimator can be applied. This approach allows for the straightforward inclusion of external or internal instruments in the second stage. The proposed estimator is simple to implement, flexible, computationally fast, and applicable across various fields.\footnote{We provide general-purpose packages for both R and Stata.} While this two-step procedure may seem unconventional, we demonstrate in Section (ref) that it is numerically equivalent to standard estimators when the least squares estimator is used in the first stage along with appropriate instruments.

As a nonlinear estimator, first-stage quantile regression is subject to a finite-sample bias that diminishes as the number of observations per group increases. Thus, our inference procedures are justified within an asymptotic framework where the number of observations per group $n$ and the number of groups $m$ diverge to infinity.\footnote{Large $n$ (often called large $T$) asymptotics have been widely applied in the nonlinear and dynamic panel data literature. For seminal contributions, see Phillips1999, Hahn2002, and Alvarez2003.} The asymptotic variance of the sample moments has two components: one arising from the first-stage quantile regression and another from the second-stage GMM regression. Since the regressors vary within groups in the first stage, the relevant number of observations is $mn$, making the variance proportional to $1/(mn)$. The second-stage estimation error, on the other hand, stems from the randomness of the group effects, with the corresponding variance proportional to $1/m$. However, this second component is zero if no group effects exist or the instrument exploits only within variation. Thus, the asymptotic distribution of the estimator is dominated by the component with the slower rate of convergence, which is the second stage error unless there is no group heterogeneity or the instruments are uncorrelated with group membership. As a result, the asymptotic distribution of the estimator is non-standard, as the convergence rate of a coefficient depends on the presence of group heterogeneity and the variation used to identify that coefficient.

In Section (ref), we consider three specific scenarios before suggesting adaptive estimation and inference procedures in Section (ref). In the first case, we assume the presence of group heterogeneity and demonstrate that only the coefficients identified through within-group variation can be estimated at the $\sqrt{mn}$ rate. For example, this applies to the fixed effects estimator. In contrast, coefficients associated with variables that are constant within groups rely on between-group variation for identification and are only estimable at the slower $\sqrt m$ rate. Only the second-stage error shows up in the first-order asymptotic distribution for these coefficients. Consequently, in this scenario, some coefficients converge faster than others. Case 2 assumes no group heterogeneity, eliminating second-stage variance and yielding a uniform $\sqrt{mn}$ convergence rate for all coefficients. Finally, we consider the intermediate case when group-level heterogeneity is present but vanishes precisely at the correct rate such that both components of the variance matter asymptotically.

These asymptotic results provide valuable insights into the mechanics of our estimator, but applying them requires knowing whether group-level heterogeneity is present or not, leading to non-adaptive inference. As the second main contribution of the paper, Section (ref) introduces adaptive estimation and inference procedures that address this issue. The proposed methodology is robust to different degrees of heterogeneity, allowing the variance of group effects to be zero, bounded, or diminishing at arbitrary rates. We first show that the leading variance term can be adaptively estimated using a traditional cluster-robust variance estimator. This result offers a practical procedure that is simple to implement and circumvents the need to estimate challenging components like the variance of first-stage coefficients. By inverting this estimated variance, we obtain a GMM estimator that is uniformly efficient in the unknown relative convergence rates of the moments. This result is non-standard because of the potentially different convergence rates of the moments; the efficient weighting matrix may be asymptotically singular. Finally, we suggest an overidentification test, which provides, for instance, the quantile equivalent of the Hausman test for the exogeneity of the between variation.

In the context of group data, the most closely related work is the IV quantile regression estimator proposed by Chetverikov2016. They focus on the effect of variables that vary only between groups and implicitly assume the presence of group-level heterogeneity.\footnote{Their asymptotic distribution is degenerate in the absence of group effects, suggesting that the rate of convergence is faster in such cases.} While both their approach and ours share the same first-stage estimation, the second stage differs: we regress the fitted values on all variables, whereas they regress the estimated intercept on the group-level regressors. As a result, their estimator is not invariant to reparametrizations of the individual-level regressors. In Table (ref), simulations using the same data-generating process as Chetverikov2016 show that our minimum distance (MD) estimator exhibits substantially lower variance and mean squared error (MSE) across all sample sizes considered—reducing the MSE by a factor of up to 20. In Section (ref), we demonstrate and explain why our estimator is more precise than theirs. Furthermore, we contribute to this literature by providing efficient adaptive estimation and inference procedures that remain valid regardless of the degree of group-level heterogeneity, deriving the limiting distribution of the estimator for the coefficients on the individual-level variables, and relaxing the growth condition of $n$ relative to $m$.

Our class of estimators includes the MD estimators of Chamberlain1994 as a special case. We extend his framework by incorporating individual-level regressors and accommodating endogenous regressors and group effects but requiring the number of groups to approach infinity.\footnote{Chamberlain1994 uses different terminology because he focuses on cross-sectional regressions. He analyzes a quantile regression model with a finite number of combinations of regressor values, where the number of cells (groups in our terminology) is finite, and the regressors are constant within each cell.} Interestingly, in Chamberlain1994, all the variance arises from the first-stage estimation, consistent with classical MD estimation, while in Chetverikov2016, the variance originates entirely from the second stage. In our framework, the variance can stem from either the first or second stage, depending on the presence or absence of group effects, with the estimated standard errors capturing the relevant leading component.

Our paper also contributes to the literature on quantile panel data models.\footnote {The main text focuses on the large $n$ (often called large $T$) literature. In short panels, Chernozhukov2013a derive bounds for quantile effects. Arellano2016 introduce a class of correlated random effects quantile regression estimators that are consistent in finite $n$. They apply this approach to study earnings and consumption dynamics in Arellano2017.} Koenker2004 introduced a penalized quantile fixed effects estimator that treats individual heterogeneity as a pure location shift. Kato2012 extend this approach by allowing group effects to depend on the quantile of interest and by contributing to the asymptotic theory of the estimator. Galvao2015 propose a two-step minimum distance (MD) estimator as a computationally efficient method for estimating fixed effects quantile models. Our framework nests this estimator. However, their focus is solely on the effects of individual-level covariates without exploiting variation between individuals.\footnote{Galvao2015 consider a traditional panel data setting; thus, in their terminology, they focus on estimating the effects of time-varying regressors.} In contrast, we aim to estimate the effects of individual- and group-level regressors while allowing for internal and external instruments. Galvao2019 suggest using the usual pooled quantile regression estimator in the presence of random effects. Our random effects estimator differs in that it targets the conditional quantile function given the group effects (see Remark (ref) for a discussion on conditional effects). In other words, we estimate a different parameter for which quantile regression is inconsistent, even when the random effects are uncorrelated with the covariates.

Chernozhukov2013 introduced distribution regression as an alternative to quantile regression for estimating the entire conditional distribution of outcomes given covariates. Fernandez-Val2022a extend this approach by developing a dynamic distribution regression panel data model with heterogeneous coefficients across groups. In their framework, the first stage involves regressing the outcome on individual-level covariates using distribution regression, followed by a second stage where the coefficients are projected onto group-level instruments. We build on their proof strategy to demonstrate that our inference procedure remains uniformly valid with respect to the degree of heterogeneity. However, our approach differs in several key aspects: we use quantile regression instead of distribution regression, project the fitted values rather than the coefficients, consider both individual-level and group-level instruments, and optimally combine these instruments using GMM. Although our method is limited to continuous outcomes, we find that quantile regression coefficients are easier and more intuitive to interpret.

As a third main contribution, we show the practical relevance of our approach in an empirical application. Specifically, we extend the work of Almond2011 by estimating the distributional effect of the Food Stamp Program on birth weight. Following the enactment of the Food Stamp Act, the number of counties implementing the program increased substantially in the late 1960s and early 1970s. To apply our minimum distance estimator, we define groups as county-trimester cells. The subscript $j$ indexes a county-trimester cell, while the subscript $i$ defines an individual within this cell. We estimate the model separately for black and white mothers and find that the Food Stamp Program has a positive impact on the lower tail of the birth weight distribution, particularly among black mothers.

The remainder of the paper is structured as follows. Section (ref) introduces the model and the estimator and shows that traditional least-squares estimators can be implemented as MD estimators. Section (ref) develops the asymptotic theory. Section (ref) extends the discussion to grouped data and compares our estimator with the grouped IV quantile regression approach of Chetverikov2016. Section (ref) applies our framework to traditional panel data, proposing quantile analogs of the within, between, and random effects estimators and the Hausman test. Both Sections (ref) and (ref) include Monte Carlo simulations to evaluate finite sample performance. In Section (ref), we present the empirical application, and Section (ref) concludes.

Model and Minimum Distance Estimator

Quantile Model

We want to learn the effects of the individual-level variables $x_{1ij}$ and the group-level variables $x_{2j}$ on the distribution of an outcome $y_{ij}$. We observe these variables for the groups $j=1,\dots,m$ and individuals $i=1,\dots, n$.\footnote{We assume a balanced panel for notational simplicity. However, the results generalize to unbalanced datasets.} For some quantile index $0<\tau<1$, we assume that

equation[equation omitted — 137 chars of source]

where $Q(\tau, y_{ij}|x_{1ij}, x_{2j},v_j)$ is the $\tau$th conditional quantile function of the response variable $y_{ij}$ for individual $i$ belonging to group $j$ given the $K_1$-vector of individual-level regressors ${x}_{1ij}$, the $K_2$-vector of group-level variables $x_{2j}$, and an unobserved random vector $v_j$ of unrestricted and unknown dimension. In total, there are $K_1 + K_2 = K$ parameters to estimate. The parameters $\beta(\tau)$, $\gamma(\tau)$ and the unobserved group heterogeneity $\alpha(\tau,v_j)$ can depend on the quantile index $\tau$. Depending on the setting, $\beta(\tau)$ or $\gamma(\tau)$ (or both) might be the parameters of interest. We normalize $\mathbb{E}[\alpha(\tau,v_{j})]=0$, which is not restrictive because $x_{2j}$ includes a constant.

remark[ Conditional versus unconditional effects] In contrast to the average effect, the definition of a quantile treatment effect depends on the conditioning variables. In this paper, we model the distribution of $y_{ij}$ conditionally on the covariates and the group effect $\alpha(\tau,v_j)$. Thus, even if the group effects are independent of the regressors, we identify different parameters than those identified by quantile regression as introduced by Koenker1978 or by instrumental variable quantile regression as introduced by Chernozhukov2005. The following example illustrates the difference between these parameters. Consider an application where each group $j$ corresponds to a region and each unit $i$ to an individual within this region. We do not have any $x_{1ij}$ variable. We are interested in the effect of a binary treatment $x_{2j}$, which has been randomized and is, therefore, independent from $\alpha(\tau,v_{j})$. $\gamma(\tau)$ is the effect of this treatment for individuals that rank at the $\tau$ quantile of $y_{ij}$ in their region. On the other hand, the quantile regression of $y_{ij}$ on $x_{2j}$ identifies the effect for individuals that rank at the $\tau$ quantile in the whole country (given the treatment status). These are different parameters except if $\alpha(\tau,v_j) =0$ for all $j$ or if the treatment effect is homogeneous such that $\gamma(\tau)=\gamma$ for all $\tau$. Whether conditional or unconditional quantile treatment effects are of interest depends on the question. Conditional quantile treatment effects are particularly useful for studying within-group inequalities when groups might be regions or industries. For example, Autor2016a and Engelhardt2021 study the effect of the minimum wage on within-state inequality, while Autor2021 study the effect of trade shock on wage inequality within local labor markets. If the unconditional effect is of interest, one can naturally obtain the unconditional distribution functions by integrating out the group effects (and possibly the other variables) and then inverting the resulting distribution functions to obtain the unconditional quantile functions, see Chernozhukov2013.

When model ((ref)) holds, the $\tau$ quantile regression of ${y_{ij}}$ on $x_{1ij}$ and a constant using only observations for group $j$ identifies the slope $\beta(\tau)$ and the intercept $x_{2j}'\gamma(\tau)+\alpha(\tau,v_j)$. We need to consider variation across groups to identify the coefficient on the group-level variables. Note that model ((ref)) implies

equation*[equation* omitted — 192 chars of source]

If $\alpha(\tau,v_j)$ is exogenous with respect to $x_{1ij}$ and $x_{2j}$ and the linear model is correctly specified, $\mathbb{E}[\alpha(\tau,v_j)|x_{1ij},x_{2j}]=0$ and a linear regression identifies the parameters of interest.\footnote{Uncorrelation between $\alpha(\tau,v_j)$ and $x_{1ij}$ and $x_{2j}$ is sufficient to identify the linear projection.} The last representation suggests a two-step estimation strategy: (i) group-level quantile regression of $y_{ij}$ on $x_{1ij}$, (ii) OLS regression of the fitted values from the first stage on $x_{1ij}$ and $x_{2j}$.

When the group effects $\alpha(\tau,v_{j})$ are endogenous (possibly correlated with $x_{1ij}$ and $ x_{2j}$), we assume that there is a $L$-dimensional vector ($L \geq K$) of valid instruments $z_{ij}$ satisfying

equation[equation omitted — 196 chars of source]

Note that $\beta(\tau)$ is identified in model ((ref)) as long as there is some variation in $x_{1ij}$ within some groups. For instance, we can include the demeaned regressors, $\dot x_{1ij}=x_{1ij}-\bar x_{1j}$ with $\bar x_{1j}=n^{-1}\sum_{i = 1}^nx_{1ij}$, in the vector of instruments $z_{ij}$ because this variable will satisfy condition ((ref)) under strict exogeneity.\footnote{In the special case of traditional panel data, the demeaned regressors correspond to the within transformation.} On the other hand, we need additional instruments to identify $\gamma(\tau)$. Equation ((ref)) suggests a similar estimation strategy as in the exogenous case but with the instrumental variable estimator (or more generally the GMM estimator) in the second stage: (i) group-level quantile regression of $y_{ij}$ on $x_{1ij}$, (ii) GMM regression of the fitted values from the first stage on $x_{1ij}$ and $x_{2j}$ using $z_{ij}$ as instrument.

remark[ Skorohod representation] The following Skorohod representation implies the model defined in equation ((ref)): \begin{align*}y_{ij}&=x_{1ij}\beta(u_{ij})+x_{2j}\gamma(u_{ij})+\alpha(u_{ij},v_{j})\\ &=q(x_{1ij},x_{2j},u_{ij},v_j),\end{align*} where $q(x_{1ij},x_{2j},u_{ij},v_j)$ is strictly increasing in the third argument (while fixing the other arguments).\footnote{This is the same model as in Chetverikov2016, where a similar Skorohod representation is derived in their footnote 7.} We normalize $u_{ij}|x_{1ij},x_{2j},v_j\sim U(0,1)$ such that $q(x_{1ij},x_{2j},\tau,v_{j})$ is the $\tau$ conditional quantile function. $u_{ij}$ ranks the individuals within a group and $v_j$ captures the group heterogeneity. In this model, a sufficient condition for equation ((ref)) is $(u_{ij},v_j)\perp \!\!\! \perp z_{ij}$. If the instrument does not vary within groups, only $v_j\perp \!\!\! \perp z_j$ is sufficient.
remark[ Heterogeneous coefficients] Our model allows only the intercept to differ between groups. Now consider a more general model where we also allow the slopes to differ between groups: \begin{equation}y_{ij}=x_{1ij}'\beta(u_{ij},v_j)+x_{2j}'\gamma(u_{ij},v_j)+\alpha(u_{ij},v_j).\end{equation} If we maintain the conditional strict monotonicity assumption with respect to $u_{ij}$, this model implies that \begin{equation}Q(\tau,y_{ij}|x_{1ij}, x_{2j},v_j) = x'_{1ij} \beta(\tau,v_j) + x'_{2j} \gamma (\tau,v_j) + \alpha(\tau,v_j).\end{equation} In the exogenous case where $(x_{1ij},x_{2j})\perp \!\!\! \perp v_j$, it follows that \begin{align*}\mathbb{E}\left[Q\left(\tau,y_{ij}|x_{1ij}, x_{2j},v_{j}\right)|x_{1ij}, x_{2j}\right] &= x'_{1ij} \int\beta(\tau,v)dF_V(v) + x'_{2j} \int\gamma (\tau,v)dF_V(v) + \int\alpha(\tau,v)dF_V(v)\\ &=x_{1ij}'\bar\beta(\tau)+x_{2j}'\bar\gamma(\tau)\end{align*}because we have normalized $\mathbb{E}[\alpha(\tau,v_j)]=0$. This implies that the linear projection of $Q(\tau,y_{ij}|x_{1ij}, x_{2j},v_{j})$ on $x_{1ij}$ and $x_{2j}$ identifies the coefficients $\beta(\tau)$ and $\gamma(\tau)$ when the homogenous model ((ref)) holds and the average effect over all groups at the $\tau$ quantile of their conditional distribution when the heterogenous model ((ref)) holds.\footnote{In the endogenous case, we obtain the instrumental variable projection instead of the standard linear projection. For instance, if $x_{2j}$ is an endogenous binary variable and $z_{ij}$ is a binary instrument, we identify the average treatment effects for the compliers at the $\tau$ quantile of their conditional distribution.} Naturally, it is also possible to model the heterogeneity between groups by estimating more flexible linear projections of $Q(\tau,y_{ij}|x_{1ij}, x_{2j},v_{j})$. For instance, we can interact $x_{1ij}$ with observable characteristics $x_{2j}$.\footnote{Starting with model ((ref)), one can simultaneously analyze both within-group and inter-group heterogeneity by constructing a quantile function with two quantile indices: one for the heterogeneity across groups and one for the heterogeneity within groups. These heterogeneous coefficients are identified through a two-step quantile regression: (i) a group-by-group quantile regression of \( y_{ij} \) on \( x_{1ij} \), followed by (ii) a quantile regression of the fitted values from the first stage on \( x_{1ij} \) and \( x_{2j} \). Pons2024 explores a version of this model that defines different parameters and falls outside the scope of this paper.}

Quantile Minimum Distance Estimators

Motivated by the representation in equation ((ref)), we suggest the following the two-step procedure. In the first step, for each group $j$ and quantile $\tau$, we regress $y_{ij}$ on individual-level variables $x_{1ij}$ and a constant using quantile regression. The intercept of the first stage regression captures both the group effect $\alpha(\tau,v_j)$ and the term $x_{2j}' \gamma(\tau)$ as these vary only between groups. In a second step, we regress the fitted values of the first stage on $x_{1ij}$ and $ x_{2j}$, using GMM with instruments $z_{ij}$.

Formally, the first stage quantile regression solves the following minimization problem for each group and quantile separately:

equation[equation omitted — 252 chars of source]

where $\rho_\tau (x) = (\tau - 1 \{x < 0\})x$ for $x \in \mathbb{R}$ is the check function. The true vector of coefficients for group $j$ is given by $\beta_j(\tau) = (\alpha(\tau,v_j)+x_{2j}'\gamma(\tau), \beta(\tau)')'$. When the model does not contain any $x_{1ij}$ variables, quantile regression computes the sample quantiles in each group.

Notation. Throughout the paper, we will use the following notation. Let $\tilde x_{ij} = (1, x_{1ij}') '$ and $x_{ij} = (x_{1ij}', x_{2j}')'$. For each group $j$ we define the following matrices. The $n \times K_1$ matrix of individual-level regressors $ X_{1j} = ( x_{11j} , x_{12j}, \dots , x_{1nj} )'$, the $n \times K$ matrix containing all regressors $X_j=(x_{1j}, x_{2j},\dots, x_{nj})'$ and the $n \times L$ matrix of instruments $Z_j=(z_{1j},z_{2j},\dots,z_{nj})'$. Further, we define two matrices for all observations. The $mn\times K$ matrix of regressors for all groups $X =$ $(X_1', \dots , X_m')'$ and the $mn\times L$ matrix of instruments for all groups as $Z = (Z_1', \dots, Z'_m )'$. We let $Y$ be the response variable's $mn \times 1$ vector. The fitted value for individual $i$ in group $j$ at quantile $\tau$ is $\hat y_{ij}(\tau)=\hat\beta_{0,j}(\tau)+x_{1ij}'\hat\beta_{1,j}(\tau)$. We denote the $n \times 1$ column vector of fitted values for group $j$ by $\hat Y_j(\tau)=(\hat y_{1j}(\tau), \dots, y_{nj}(\tau))'$, and the $mn\times 1$ vector of fitted values by $\hat Y (\tau)= (\hat Y_1'(\tau), \dots, \hat Y_m'(\tau))'$.

remark[ Alternative first-stage estimators] The quantile regression estimator proposed by Koenker1978 is not necessarily efficient. Newey1990a suggest a semiparametrically efficient weighted estimator of $\beta_j(\tau)$. However, we opt for the unweighted quantile regression estimator due to the challenges associated with estimating the weights and the complicated interpretation of the estimates in cases of misspecification. In our model ((ref)), the variation within groups is assumed to be exogenous. If this assumption were violated, one could apply an instrumental variable (IV) quantile regression (see, e.g., Chernozhukov2006) in the first stage, followed by the second-stage GMM regression described below.\footnote{An IV extension of the MD estimator by Galvao2015 is suggested in Dai2021.} We do not pursue this (computationally intensive) extension in this paper.

The second stage consists of the linear GMM regression of $\hat Y(\tau)$ on $X$ using $Z$ as an instrument. The estimator has the following closed-form expression:

equation[equation omitted — 159 chars of source]

where $\hat W(\tau)$ is a $L \times L$ symmetric weighting matrix. When $L = K$, the second step estimator in equation ((ref)) simplifies to the IV estimator using $Z$ as an instrument, and we can drop the dependence on $\hat W(\tau)$.

Our two-step estimator is extremely simple to implement; it requires only routines performing quantile regression and GMM estimation, which are already available in many software applications. Quantile regression, which is computationally more demanding due to the absence of a closed-form solution, is used only in the first stage, where there are fewer observations and a limited number of parameters to estimate. The first stage is also embarrassingly parallelizable, increasing the computational speed. For this reason, our estimator remains computationally attractive in large datasets with numerous groups. The second stage is a straightforward GMM estimator, which includes OLS and two-stage least squares as special cases. Traditional panel data methods can also be used in the second stage. For instance, in our application, we observe individuals born in a given trimester in a given county. The subscript $j$ defines a county-trimester cell, while the subscript $i$ defines an individual within this cell. In the second stage, we include trimester, county, and state $\times$ year fixed effects to estimate the effect of food stamps on the birth weight distribution.

remark[ Interpretation as a minimum distance estimator] Our estimator can be written as an MD estimator, where the second stage imposes restrictions on the first-stage coefficients. For simplicity, in this remark, we consider the case where all the regressors are exogenous and $Z = X$. Define \begin{equation} \underset{\scriptscriptstyle (K_1+1) \times K}{R_j} = \begin{pmatrix} 0 & x_{2j}' \\ I_{K_1} & 0 \end{pmatrix} \end{equation} such that $\tilde X_j R_j = X_j$. It follows that our MD estimator minimizes \begin{align} \hat \delta(\tau) &= \operatorname*{arg\,min}_{\delta} \sum_{j = 1}^m (\tilde X_j\hat\beta_j(\tau)-X_j\delta)'(\tilde X_j\hat\beta_j(\tau)-X_j\delta)\nonumber \\ &= \operatorname*{arg\,min}_{\delta} \sum_{j = 1}^m (\tilde X_j\hat\beta_j(\tau)-\tilde X_j R_j \delta)'(\tilde X_j\hat\beta_j(\tau)-\tilde X_j R_j \delta)\nonumber \\ &= \operatorname*{arg\,min}_{\delta} \sum_{j = 1}^m (\hat\beta_j(\tau)-R_j \delta)'\tilde X_j'\tilde X_j(\hat\beta_j(\tau)-R_j \delta), \end{align} which corresponds to the definition of a weighted minimum distance estimator that imposes the linear restrictions $\beta_j(\tau)=R_j \delta(\tau)$ with weights $\tilde X_j'\tilde X_j$. Thus, our estimator is an MD estimator. However, it does not correspond to the textbook definition of a “classical minimum distance” estimator.\footnote{See section 14.6 in Wooldridge2010.} In the classical MD setup, all the sampling variance arises in the first stage: if we know the first stage coefficients, we know the final coefficients. It follows that the efficient weighting matrix $\tilde W(\tau)$ is the inverse of the first-stage variance. In our case, the second stage also contributes to the variance due to the presence of the group effects $\alpha(\tau,v_j)$. Even if we know $\beta_j(\tau)$ (for a finite number of groups), we cannot exactly pinpoint $\gamma(\tau)$. The group effects play a role similar to misspecification in the classical MD, but with our estimator, the resulting bias disappears asymptotically as the number of groups increases. This is the second important difference: the dimension of our first stage estimates increases with the sample size while it is fixed for classical MD estimators.

Least Squares Minimum Distance Estimators

This paper proposes a two-step estimator in which the first stage involves performing quantile regressions within each group. Although this approach might seem unusual and specific to quantile models, we demonstrate in this subsection that this method yields numerically identical results to traditional least squares panel estimators, provided that OLS is applied in the first stage. A more detailed discussion, including formal statements, is provided in Appendix (ref), with proofs available in Appendix (ref).

Consider first a model with group effects and individual-level regressors

equation[equation omitted — 77 chars of source]

Typically, when estimating this model with fixed effects, most researchers apply the within transformation, which is widely recognized as being equivalent to a dummy variable regression.\footnote{In the context of traditional panel data, we would refer to the $j$ units as “individuals” and the $i$ units as “time periods”. Thus, this transformation corresponds to the time-demeaning applied in the traditional panel data literature, eliminating time-invariant individual effects.} Yet, a third equivalent method exists for computing the least squares fixed effects estimator, which involves exploiting the exogenous within-group variation using instrumental variables. Setting $\dot x_{1ij}$ as an instrument for $x_{1ij}$ in an instrumental variable regression is numerically identical to the least squares fixed effects.

Corollary (ref) in Appendix (ref) presents a fourth way to compute the least squares fixed effects estimator, which aligns with our paper's approach. This minimum distance method involves two steps: first, regressing with OLS the dependent variable on $x_{1ij}$ within each group, then using IV to regress the first-stage fitted values on $x_{1ij}$, with $\dot x_{1ij}$ as the instrument. This approach offers the most computationally efficient alternative for quantile estimation, as it divides the problem into two convex optimization steps rather than a single nonconvex IV quantile regression, and it avoids the challenges of high-dimensional quantile regression.

The two-step procedure is not specific to fixed effects but applies to a wide range of estimators. Proposition (ref) in Appendix (ref) shows that the MD least squares estimator is algebraically identical to the one-step GMM regression of $y_{ij}$ on $x_{ij}$ under the condition that for each group $j$, the matrix of instruments lies in the column space of the matrix of first stage regressors.\footnote{The intuition is as follows. The fitted values of the first-stage least squares regression can be written as $P_{X_j} Y_j$ where $P_{X_j}$ is the first-stage least squares projection matrix of group $j$. If the instrument matrix, $Z_j$, is in the column space of $\tilde X_j$, it follows that $P_{X_j} Z_j = Z_j$. Therefore, $Z' \hat Y = Z'Y$ and the two GMM regressions are numerically identical.} For example, $\dot x_{1ij}$, $\bar x_{1j}$, and $x_{2j}$ satisfy the condition.

We extend the model by incorporating group-level regressors, $x_{2j}$:

equation[equation omitted — 92 chars of source]

By selecting different instrumental variables for the second-step GMM regression, we can numerically obtain the most common least squares panel data estimators. For example, using the group-averaged variables, $\bar{x}_{1j}$ and $x_{2j}$, as instruments yields the between estimator. Instrumental variable approaches are also available for random effects estimation. Although FGLS is the most common estimator for the random effects model, Im1999 show that the overidentified 3SLS estimator, with instruments $\dot x_{1ij}$, $\bar x_{1j}$, and $x_{2j}$, is numerically identical to the random effects estimator. Since 3SLS is a special case of a GMM estimator, using the first-stage fitted values as dependent variables does not change the estimates. Alternatively, the random effects estimator can be implemented using the theory of optimal instrument with a just identified 2SLS regression (see Im1999, Hansen2021). Additionally, the Hausman1981 estimator can be implemented by selecting the following instruments: $\dot x_{1ij}$ and the group average of the exogenous regressors. External instruments might also be included.

Asymptotic Theory

Preliminaries: Assumptions, Consistency, and Sample Moments

In this section, we state the assumptions and present the asymptotic results. All the proofs are included in Appendix (ref). For simplicity of notation, in the following, we write $\alpha_j(\tau)$ instead of $\alpha(\tau, v_j)$. We prove weak uniform consistency and weak convergence of the whole quantile regression process for $\tau\in\mathcal{T}$, where $\mathcal{T}$ is a compact set included in $(0,1)$. The symbol $\ell^\infty ( \mathcal{T} )$ denotes the set of component-wise bounded vector valued function of $\mathcal{T}$, $\rightsquigarrow$ denotes weak convergence, and for a random variable $h_{ij}$, $\mathbb{E}_{i|j}[h_{ij}]$ indicates the expectation over $i$ in group $j$.

We start by writing the sampling error of $\hat\delta(\hat W, \tau)$ as a sum of a component arising from the first stage estimation error of $\beta_j(\tau)$ and a component arising from the second stage noise $\alpha_j(\tau)$:

lemma[Sampling error] Assume that the model in equation ((ref)) holds, then $$\hat\delta(\hat W,\tau)-\delta(\tau)=\hat G(\tau) \frac{1}{mn} \sum_{j = 1}^m\sum_{i = 1}^n z_{ij} \left(\tilde x_{ij}'(\hat \beta_j(\tau)-\beta_j(\tau))+\alpha_j(\tau)\right),$$ where $ \hat G(\tau) = \left(S_{ZX}'\hat W(\tau)S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau)$ and $S_{ZX}=\frac{1}{mn}\sum_{j=1}^m\sum_{i = 1}^n z_{ij}x_{ij}'$.

To keep the notation light, we suppress the dependency of $\hat G(\tau)$ on $\hat W(\tau)$. We now state assumptions that ensure that both components are well-behaved. For the analysis of the first stage estimator, we rely on results derived in Galvao2020 and make the assumptions required in their Theorem 2:

assumption[Sampling] (i) The processes $\{(y_{ij}, x_{ij}, z_{ij}):i = 1, \dots, n\}$ are independent across $j$. (ii) For each $j$, the observations $(y_{ij}, x_{ij},z_{ij})_{i=1,\dots, n}$ are i.i.d. across $i$.
assumption[Covariates] (i) For all $j=1,\dots,m$ and all $i=1,\dots, n$, $\lVert x_{ij}\rVert \leq C$ almost surely. (ii) The eigenvalues of $\mathbb{E}_{i|j}[ \tilde x_{ij} \tilde x_{ij}']$ are bounded away from zero and infinity uniformly across $j$.
assumption[Conditional distribution] The conditional distribution $F_{y_{ij} | x_{1ij}, v_j} (y|x,v)$ is twice differentiable w.r.t. y, with the corresponding derivatives $f_{y_{ij} | x_{1ij}, v_j} (y|x,v)$ and $f'_{y_{ij} | x_{1ij}, v_j} (y|x,v)$. Further, assume that \begin{align*} f_{max} = \sup_j \sup_{y \in \mathbb{R}, x \in \mathcal{X}} |f_{y_{ij} | x_{1ij}, v_j} (y|x, v)| < \infty \end{align*} and \begin{align*} \bar f' = \sup_j \sup_{y \in \mathbb{R}, x \in \mathcal{X}} |f'_{y_{ij} | x_{1ij}, v_j} (y|x,v)| < \infty. \end{align*} where $\mathcal{X}$ is the support of $x_{1ij}$
assumption[Bounded density] There exists a constant $f_{min} < f_{max}$ such that \begin{align*} 0 < f_{min} \leq \inf_j \inf_{\tau \in \mathcal{T}} \inf_{x \in \mathcal{X}} f_{y_{ij} | x_{1ij}, v_j} (Q(\tau,y_{ij}|x, v)|x,v). \end{align*}

These are quite standard assumptions in the quantile regression literature. In Assumption (ref), we assume that the processes are independent across $j$; this assumption can also be relaxed by allowing for clustering between groups. We also assume that the observations are i.i.d. within groups, but this can be relaxed at the cost of a more complex notation by applying Theorem 4 in Galvao2020, which requires only stationarity and $\beta$-mixing. The estimator of the asymptotic variance that we suggest below is consistent in both cases. Assumption (ref) requires that the regressors are bounded and that $\mathbb{E}_{i|j}[\tilde x_{ij}\tilde x_{ij}']$ is invertible. Assumptions (ref) and (ref) impose smoothness and boundedness of the conditional distribution, the density, and its derivatives.

For the second stage GMM regression we impose the following assumptions:

assumption[Instruments] (i) For all $j=1,\dots,m$ and all $i=1,\dots, n$, $|| z_{ij} || \leq C$ a.s. (ii) For all $j=1,\dots,m$ and all $i=1,\dots, n$, $\mathbb{E}[z_{ij}\alpha_j(\tau)]=0$. (iii) For all $j=1,\dots,m$ and all $i=1,\dots, n$, $y_{ij}$ is independent of $z_{ij}$ conditional on $(x_{ij},v_j)$. (iv) As $m\rightarrow \infty$, $m^{-1}\sum_{j = 1}^m \mathbb{E}_{i|j}[z_{ij}x_{ij}']\rightarrow\Sigma_{ZX}$ where the singular values of $\Sigma_{ZX}$ are bounded from below and from above.
assumption[Group effects] \newline (i) For all $j=1,\dots,m$, $\mathbb{E}\left[\sup_{\tau \in \mathcal{T}}|\alpha_j(\tau)|^{4+\varepsilon_C}\right]$ $\leq C$ for $\varepsilon_C>0$. (ii) For some (matrix-valued) function $\Omega_2 : \mathcal{T} \times \mathcal{T} \rightarrow \mathbb{R}^{L \times L}$, $m^{-1}\sum_{j = 1}^m \mathbb{E}_{i|j}[\alpha_j(\tau_1) \alpha_j(\tau_2) z_{ij}z_{ij}']\underset{p}{\rightarrow} \Omega_{2}(\tau_1, \tau_2)$ uniformly over $\tau_1, \tau_2 \in \mathcal{T}$. (iii) For all $\tau_1, \tau_2 \in \mathcal{T}$, $|\alpha_j(\tau_2) - \alpha_j(\tau_1)| \leq C | \tau_2 - \tau_1 |$.

These assumptions are the same as in Chetverikov2016. For the instrumental variables, we assume that (i) they are bounded, (ii) they are not correlated with the group effect (exclusion restriction), (iii) they do not affect the first stage estimation (this is often satisfied by construction, e.g. when the instruments do not vary within individuals or are a linear transformation of the first stage regressors), and (iv) they satisfy the relevance conditions. For the group effects, we assume that they have a finite fourth moment, and the average variance of $z_{ij}\alpha_j(\tau)$ converges to a well-defined matrix.

Since the unobserved heterogeneity $\alpha_j(\tau)$ is group-specific, we require that the number of groups $m$ diverges to infinity. The first stage quantile regression estimator is a nonlinear estimator that is potentially biased in finite samples. Hence, the number of observations per group, $n$, must also diverge to infinity for consistency. Galvao2020 show that the bias is approximately of order $1/n$. For unbiased asymptotic normality, we need the bias to shrink faster than the standard deviation of the estimator. We will see that some elements of $\hat \delta(\hat W, \tau)$ converge at the $\sqrt m$ rate such that we need that $n$ goes to infinity more quickly than $\sqrt{m}$. On the other hand, other elements converge at the $\sqrt{mn}$ rate so that $n$ must go to infinity more quickly than $m$. We state these three different relative growth rates in the following assumption:

assumption[Growth rates] As $m\rightarrow \infty$, we have \begin{enumerate}[label=(\alph*)] • $\frac{\log m}{n}\rightarrow 0$, • $ \frac{ \sqrt{m} \log n }{n}\rightarrow 0$, • $ \frac{ m \left(\log n\right)^2 }{n}\rightarrow 0$. \end{enumerate}

Finally, we assume that the estimated weighting matrix uniformly converges to a strictly positive definite matrix that is continuous in the quantile index.

assumption[Full-rank weighting matrix] Uniformly in $\tau\in\mathcal T$, $\hat W(\tau)\underset{p}{\rightarrow} W(\tau)$ where $W(\tau)$ is strictly positive definite and, for all $\tau_1, \tau_2 \in \mathcal{T}$, $|| W (\tau_2) - W(\tau_1) || \leq C | \tau_2- \tau_1 |$.

Our first result establishes the uniform consistency of our estimator under the weakest growth rate condition:

theorem[Uniform consistency] Let the model in equation ((ref)), Assumptions (ref)-(ref), (ref)(a), and (ref) hold. Then, $$\underset{\tau\in\mathcal T}{\sup}\lVert \hat \delta(\tau)-\delta(\tau)\rVert = o_p(1).$$

We now study the asymptotic distribution of our estimator. In Lemma (ref), we see that the sample moment condition is the sum of two terms. It is useful to consider them separately:

align[align omitted — 262 chars of source]

such that total moment condition is the sum of both components: $\bar g_{mn}(\hat\delta,\tau)=\bar g^{(1)}_{mn}(\hat\delta,\tau)+\bar g^{(2)}_{mn}(\hat \delta,\tau)=\frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^ng_{ij}(\hat \delta,\tau)$. Lemma (ref) establishes joint asymptotic normality for the entire moment condition processes.

lemma[Asymptotic distribution of the sample moments] Let the model in equation ((ref)), and Assumptions (ref)-(ref) hold. \begin{enumerate}[label=(\roman*), wide, labelindent=0pt] • Under Assumption (ref)(c), as $m\rightarrow \infty$, \begin{equation} \sqrt{mn}\bar g^{(1)}_{mn}(\hat\delta, \cdot) \rightsquigarrow \mathbb Z_1(\cdot), in $\ell^\infty(\mathcal T)$, \end{equation} where $\mathbb Z_1(\cdot)$ is a mean-zero Gaussian process with uniformly continuous sample paths and covariance function $\Omega_{1}(\tau,\tau')=\mathbb{E}\left[\Sigma_{ZXj}V_j(\tau,\tau')\Sigma_{ZXj}'\right]$ with $\Sigma_{ZXj}=\mathbb{E}_{i|j}[z_{ij} \tilde x_{ij}']$ and $V_j(\tau,\tau')$ is the asymptotic variance-covariance matrix of $\hat\beta_j(\tau)$ and $\hat\beta_j(\tau')$: \begin{align*}V_j(\tau,\tau') = \mathbb{E}_{i|j}[f_{y|x}(Q_{y|x, \nu_j}(\tau|\tilde x_{ij} )|\tilde x_{ij})\tilde x_{ij} \tilde x_{ij}']^{-1} (\min(\tau,\tau')-\tau\tau') \mathbb{E}_{i|j}[\tilde x_{ij}\tilde x_{ij}']\\ \times \mathbb{E}_{i|j}[f_{y|x}(Q_{y|x, \nu_j}(\tau'|\tilde x_{ij} )| \tilde x_{ij})\tilde x_{ij} \tilde x_{ij}']^{-1}\end{align*} • Under Assumption (ref)(b), As $m\rightarrow \infty$, \begin{equation} \sqrt{m}\bar g^{(2)}_{mn}(\hat\delta, \cdot) \rightsquigarrow \mathbb Z_2(\cdot), in $\ell^\infty(\mathcal T)$, \end{equation} where $\mathbb Z_2(\cdot)$ is a mean-zero Gaussian process with uniformly continuous sample paths and covariance function $\Omega_2(\tau,\tau')$, which is defined in Assumption (ref)(ii). • Under Assumption (ref)(c), as $m\rightarrow \infty$, $\underset{\tau,\tau'\in\mathcal T}{\sup} \left\lVert \operatorname{Cov} \left (\bar g^{(1)}_{mn}(\hat\delta, \tau), \bar g^{(2)}_{mn}(\hat\delta, \tau')\right )\right\rVert=o_p\left(\frac{1}{\sqrt{mn}}\right)$. \end{enumerate}

$\bar g^{(1)}_{mn}(\hat\delta, \cdot)$ reflects the estimation error that arises in the first-stage quantile regression estimation. Since the first-stage regressors vary within groups, the relevant number of observations is $mn$, and correspondingly, the variance is proportional to $1/(mn)$. On the other hand, since the expected bias of the first-stage quantile regression is of order $1/n$, for asymptotic unbiasedness, we must require that $n$ goes to infinity slightly faster than $m$. In the proof, we build on results derived in Volgushev2019 and in Galvao2020. $\bar g^{(2)}_{mn}(\hat\delta, \cdot)$ reflects the estimation error due to the randomness in $\alpha_j(\tau)$. This moment can also be interpreted as the moment that would be relevant if we knew $\beta_j(\tau)$. Since $\alpha_j(\tau)$ varies only between groups, the relevant number of observations here is $m$ and, accordingly, the variance of this moment converges at the slower rate of $1/m$. For asymptotic unbiasedness, we need only the weaker condition (ref)(b), which requires that $n$ goes to infinity slightly faster than $\sqrt m$.

Asymptotic Distribution when the Degree of Heterogeneity is Known

The sample moment condition $\bar g_{mn}(\hat\delta,\tau)$ is thus the sum of two components that converge to zero at different rates. The asymptotic distribution of the estimator is dominated by the component with the slower rate of convergence, $\bar g_{mn}^{(2)}(\hat\delta,\tau)$, except its variance is zero, which is the case if $\operatorname{Var}(\alpha_j(\tau))=0$ or if $\bar z_j = \frac{1}{n}\sum_{i = 1}^n z_{ij} = 0$ for $j=1,\dots,n$. Since the degree of group-level heterogeneity affects this variance, it is useful to consider three cases: strong, no, and weak heterogeneity. In this subsection, we derive the asymptotic distribution of our estimator when the degree of heterogeneity is known. We suggest adaptive estimation and inference procedures in the following subsection.

Case 1: Strong group-level heterogeneity. We start with the case of strong heterogeneity that we define to be $\operatorname{Var}(\alpha_j(\tau))>\varepsilon>0$ uniformly in $\tau$. The variance of an element of the vector $\bar g^{(2)}_{mn}(\hat\delta, \tau)=\frac{1}{mn}\sum_{j=1}^m \bar z_j\alpha_j(\tau)$ equals to zero when the corresponding instrument satisfies $\bar z_{j}=0$ for all $j$. For this reason, we distinguish between two sorts of instruments: $L_{1}$ instruments in $z_{1ij}$ satisfy $\bar z_{1j}=n^{-1}\sum_{i=1}^nz_{1ij}=0$ for all $j$, while $L_{2}$ instruments in $z_{2ij}$ satisfy $\bar z_{2j}\neq 0$ at least for some groups $j$.\footnote{Note that all the instruments that vary only within groups can be normalized to have mean zero. For instance, we can identify the effect of the individual-level variable $x_{1ij}$ by using the instrument $\dot x_{1ij}$, which has a zero mean in all groups.} We order the instruments such that $z_{ij}=(z_{1ij}',z_{2ij}')'$. It follows that

equation*[equation* omitted — 335 chars of source]

where $g_{mn,1}(\hat\delta,\tau)$ is a $L_1\times 1$ vector and $g_{mn,2}(\hat\delta,\tau)$ is a $L_2\times 1$ vector. Thus, in the case of strong heterogeneity, some moments converge at the fast rate $\sqrt{mn}$ while others converge at the slow rate $\sqrt{m}$.

We order the coefficients in $\delta(\tau)$ (and the corresponding regressors in $x_{ij}$) such that the first $M_1$ elements are identified using only the $L_1$ fast moments while the remaining $M_2$ elements require the $L_2$ slow moments for identification. We denote by $\delta_1(\tau)$ the former and by $\delta_2(\tau)$ the latter coefficients. Formally, we partition $\Sigma_{ZX}$ such that it is block lower triangular:

align[align omitted — 206 chars of source]

where $\Sigma_{11}$ is a full column rank $L_1\times M_1$ matrix, $\Sigma_{12}$ is $L_1\times M_2$, $\Sigma_{21}$ is $L_2\times M_1$ and $\Sigma_{22}$ is $L_2\times M_2$. Assumption (ref)(iv) implies that $\Sigma_{22}$ also has full column rank. Note that $L_1$ and $M_1$ can be equal to zero such that equation ((ref)) is a definition and not an assumption. Accordingly, the adaptive estimation and testing procedures suggested in Section (ref) do not require the user to classify the instruments or the regressors.

In the exactly identified case, our estimator simplifies to the instrumental variable estimator such that

equation*[equation* omitted — 233 chars of source]

and the first-order asymptotic distributions can be written as

align*[align* omitted — 295 chars of source]

$\hat\delta_1(\tau)$ is only a function of $\bar g_{mn}^{(1)}$ and converges thus at the $\sqrt{mn}$ rate. $\hat\delta_2(\tau)$ depends on both $\bar g_{mn}^{(1)}$ and $\bar g_{mn}^{(2)}$ but its first-order asymptotic distribution is dominated by the slower $\bar g_{mn}^{(2)}$ term.

exampleWe can illustrate this notation with a simple example. Consider the case of one individual-level variable $x_{1ij}$, one group-level variable $\tilde x_{2j}$, and a constant. As instrumental variables, we use $(\dot x_{1ij}, \tilde x_{2j},1)$, which corresponds to applying the fixed effects estimator for the coefficient on $x_{1ij}$ and the between estimator for the other two coefficients. In this case, only the first instrument has mean zero in all groups such that $L_1=1$ and $L_2=2$. By construction, $\dot x_{1ij}$ is uncorrelated with $\tilde x_{2j}$ and with the constant such that $\Sigma_{ZX}$ is block lower diagonal as defined in equation ((ref)) with $M_1=1$ and $M_2=2$. The coefficient on $x_{1ij}$ converges at the fast rate because it is not affected by the group-level effects $\alpha_j(\tau)$. On the other hand, the intercept and the coefficient on the group-level variable $\tilde x_{2j}$ converge only at the slow rate of convergence $\sqrt m$ because they are affected by the group effects. In a many application, we see that $M_1=K_1$, but this is not always true. For instance, in the previous example, all coefficients would converge at the slow rate if we used $(x_{1ij},\tilde x_{2j},1)$ as instrumental variables such that $L_1$ and $M_1$ would be equal to zero.

In an overidentified model, using a full-rank weighting matrix \( W(\tau) \) as in Assumption (ref) can lead to contamination of the entire parameter vector $\hat\delta(\tau)$ by the slower-converging moments. This occurs because $\hat G(\tau)$ will not retain a block-lower-triangular structure, even if $S_{ZX}$ does. In Example (ref) below, this issue arises with the 2SLS estimator of $\delta_1(\tau)$, which converges at the slow $\sqrt{m}$ rate. In contrast, both the exactly identified IV estimator and the efficient GMM estimator achieve the faster $\sqrt{mn}$ rate.

To avoid contamination, we must give asymptotically infinitely higher weights to fast moments than slow ones. The following assumption imposes this critical condition on the weighting matrix.\footnote{The intuition behind the sequence $a_n(\tau)$ will be elucidated in the following subsection, where we delve into a comprehensive discussion of our efficient GMM estimator.}

assumptionp{\ref*{a:weighting}$'$}[Heterogeneous weighting matrix] Uniformly in $\tau\in\mathcal T$, \begin{equation*} \hat W(\tau) = \underbrace{\begin{pmatrix} W_{11}(\tau) & a_n(\tau) W_{12}(\tau) \\ a_n(\tau) W_{21}(\tau) & a_n(\tau) W_{22}(\tau) \end{pmatrix}}_{W_{mn}(\tau)} + \begin{pmatrix} o_p(1) & o_p \left (\sqrt{a_n(\tau)}\right ) \\ o_p \left (\sqrt{a_n(\tau)} \right ) & o_p({a_n}(\tau)) \end{pmatrix} \end{equation*} where $a_n(\tau)=\frac{1}{1+\operatorname{Var}(\alpha_j(\tau))n}$. $W_{11}(\tau)$ and $W_{22}(\tau)$ are respectively a $L_1\times L_1$ and a $L_2\times L_2$ full rank matrix. For all $l_1,l_2\in\{1,2\}$ and $\tau_1, \tau_2 \in \mathcal{T}$, $|| W_{l_1l_2} (\tau_2) - W_{l_1l_2}(\tau_1) || \leq C | \tau_2- \tau_1 |$.
exampleConsider an extension of Example (ref) where the vector of instrumental variable is now $(\dot x_{1ij}, \bar x_{1j}, x_{2j},1)$. This model is overidentified with $L_1=1$, $L_2=3$, and $K=3$. When we impose Assumption (ref), $W(\tau)$ is a full rank weighting matrix, and the estimator converges only at the $\sqrt m$ rate. We can obtain a faster-converging estimator by giving asymptotically infinitely more weight to $\dot x_{ij}$ than to the other instruments. When we impose Assumption (ref), the coefficient on $x_{1ij}$ is estimated at the $\sqrt{mn}$ rate. We formalize these results in Theorem (ref) below.

Case 2: no group-level heterogeneity. We now consider the case where there is no heterogeneity across groups, i.e., when $\alpha_j(\tau)=0$ uniformly in $j$ and $\tau$. When this occurs, $\bar g^{(2)}_{mn}(\hat\delta,\cdot)=0$ such that all the variance arises from the first stage moment, which can be estimated at the fast rate $\sqrt{mn}$. This case corresponds to the textbook definition of a classical minimum distance estimator. All the coefficients converge at the fast $\sqrt{mn}$ rate, and the asymptotic distribution of $\hat\delta(\tau)$ can be derived straightforwardly. By comparing Case 1 and 2, we notice that the limiting distribution and even the rate of convergence of the $M_2$ coefficients $\hat\delta_2(\tau)$ differ. Their asymptotic distribution is discontinuous at $0$ in the variance of $\alpha_j(\tau)$, and in many applications, we do not know whether there is group-level heterogeneity. For this reason, we discuss adaptive estimation and inference in the next subsection.

remark[Related literature] Chamberlain1994 considers a setting with a finite number of design points (groups in our terminology), exogenous group-level variables, and no individual-level variables. His correctly specified case corresponds to our Case 2 (no heterogeneity). Accordingly, our asymptotic distribution corresponds to his in this special case.\footnote{In the absence of group-level heterogeneity, the asymptotic distribution of $\hat\delta(\tau)$ stated in part (ii) of Theorem (ref) is also valid when the number of groups is fixed.} Chamberlain1994 also considers a misspecified case. We can interpret his misspecification errors as our group effects. However, in this case, he considers pseudo-true parameters that absorb a non-vanishing bias while we allow the number of groups to go to infinity to avoid bias. For this reason, our results fundamentally differ from his results in Case 1. Chetverikov2016 do not consider Case 2 (no heterogeneity) in their theoretical results even if, interestingly, it corresponds to one of their data-generating processes in their simulations. In Case 2, their matrix-valued function $J(u_1,u_2)$ defined in their Assumption 6(ii) is uniformly equal to zero. This implies that their asymptotic covariance function $\mathcal{C}(u_1,u_2)$ defined in their Theorem 1 is also uniformly equal to 0. This degenerate asymptotic distribution indicates that the convergence rate is faster than $\sqrt m$ in this case.

Case 3: weak group-level heterogeneity. In Case 1, the second-stage variance dominates the asymptotic distribution of $\hat\delta_2(\tau)$, while in Case 2, the first-stage variance dominates. It is interesting to consider the intermediate case when group-level heterogeneity is present but vanishes exactly at the right rate such that both components of the variance matter asymptotically. This should provide a good approximation for the applications where the first-stage and the second-stage variances are similar. Formally, we assume that $\Omega_2(\tau_1,\tau_2)$, the covariance function of $\bar g^{(2)}_{mn}(\hat\delta,\tau)$ defined in Assumption (ref)(ii), converges to zero at the $n$ rate: $$\Omega_2(\tau_1,\tau_2)=n^{-1}\bar\Omega_2(\tau_1,\tau_2).$$ Under this assumption, $\bar g^{(1)}_{mn}(\hat\delta,\cdot)$ and $\bar g^{(2)}_{mn}(\hat\delta,\cdot)$ converge at the same rate. Lemma (ref) implies then

equation[equation omitted — 159 chars of source]

where $\mathbb Z(\cdot)$ is a mean-zero Gaussian process with uniformly continuous sample paths and covariance function $\Omega_{1}(\tau,\tau')+\bar\Omega_2(\tau,\tau')$.

Theorem (ref) formally states the results for these three cases. Parts (i) provides the asymptotic distribution for the fast and slow coefficients when there is strong heterogeneity, part (ii) in the absence of heterogeneity, and part (iii) when there is weak heterogeneity.

theorem[Asymptotic distribution when the degree of heterogeneity is known] Let Assumptions (ref)-(ref) hold. \begin{enumerate}[label=(\roman*), wide, labelindent=0pt] • Case 1 (strong heterogeneity): $\operatorname{Var}(\alpha_j(\tau))>\varepsilon>0$ uniformly in $\tau$ and Assumption (ref) holds. \begin{enumerate} • In addition, let Assumption (ref)(c) hold. Then, \begin{equation} \sqrt{mn}(\hat\delta_1(\hat W(\cdot), \cdot)-\delta_1(\cdot))\rightsquigarrow G_{11}(\cdot) \mathbb Z_{11}(\cdot), in $\ell^\infty(\mathcal T)$,\end{equation} where $G_{11}(\tau) = \left( \Sigma_{11}' W_{11}(\tau) \Sigma_{11}\right)^{-1} \Sigma_{11}' W_{11}(\tau)$ and $\mathbb Z_{11}(\cdot)$ is the Gaussian process consisting of the first $L_1$ elements of $\mathbb Z_{1}(\cdot)$ defined in Lemma (ref)(i). • In addition, let Assumption (ref)(b) hold. Then, \begin{equation} \sqrt{m}(\hat\delta_2(\hat W(\cdot), \cdot)-\delta_2(\cdot))\rightsquigarrow G_{22}(\cdot) \mathbb Z_{22}(\cdot), in $\ell^\infty(\mathcal T)$,\end{equation} where $G_{22}(\tau) = \left( \Sigma_{22}' W_{22}(\tau) \Sigma_{22}\right)^{-1} \Sigma_{22}' W_{22}(\tau)$ and $\mathbb Z_{22}(\cdot)$ is the Gaussian process consisting of the last $L_2$ elements of $\mathbb Z_{2}(\cdot)$ defined in Lemma (ref)(ii). \end{enumerate} • Case 2 (no heterogeneity): $\alpha_j(\tau)=0$ uniformly in $j$ and $\tau$. In addition, Assumption (ref)(c) and (ref) hold. Then, \begin{equation} \sqrt{mn}(\hat\delta(\hat W(\cdot), \cdot)-\delta(\cdot))\rightsquigarrow G(\cdot) \mathbb Z_1(\cdot), in $\ell^\infty(\mathcal T)$,\end{equation} where $G(\tau) = \left( \Sigma_{ZX}' W(\tau) \Sigma_{ZX}\right)^{-1} \Sigma_{ZX}' W(\tau)$ and $\mathbb Z_1(\cdot)$ is defined in Lemma (ref)(i). • Case 3 (weak heterogeneity): $\Omega_2(\tau,\tau') = n^{-1} \bar\Omega_2(\tau,\tau')$. Assumption (ref) and (ref)(c) hold. Then, \begin{equation} \sqrt{mn}(\hat\delta(\hat W(\cdot), \cdot)-\delta(\cdot))\rightsquigarrow G(\cdot) \mathbb Z(\cdot) , in $\ell^\infty(\mathcal T)$,\end{equation} where $G(\tau) = \left( \Sigma_{ZX}' W(\tau) \Sigma_{ZX}\right)^{-1} \Sigma_{ZX}' W(\tau)$ and $\mathbb Z(\cdot)$ is a mean-zero Gaussian process with uniformly continuous sample paths and covariance function $\Omega_{1}(\tau,\tau')+\bar\Omega_2(\tau,\tau')$. \end{enumerate}

These asymptotic results are useful for understanding the mechanics behind our estimator, but they have several weaknesses. First, the asymptotic distribution of $\hat\delta_1(W, \tau)$ in part (i)-(a) is only a function of the fast instruments. The same asymptotic distribution can be obtained by ignoring the slow instruments. Consider Example (ref), including $\bar x_{1j}$ as an instrument does not reduce the asymptotic variance of $\hat\delta_1(W, \tau)$ even if this instrument is valid and the between-group variation in $x_{1ij}$ is non-negligible. In other words, the random effects estimator is asymptotically equivalent to the fixed effects estimator.\footnote{This is not specific to quantile models and also affects least squares models with large $n$ (see Ahn2014).} In some applications, the random effects estimator is appreciably more precise such that we would like to exploit the between-group variation efficiently. Second, the asymptotic distribution and the convergence rate of the slow coefficients $\hat \delta_2(W, \tau)$ are different, depending on whether there is group-level heterogeneity or not. Thus, performing inference based on Theorem (ref) requires knowing which case is relevant for the specific application. In other words, inference based directly on this result is not adaptive. Third, in Case 1, the variance coming from the first stage estimation does not appear in the asymptotic distribution of $\hat\delta_2(W, \tau)$ because it converges to zero at a quicker rate. Consequently, inference may have poor properties. We solve these issues in the following subsection by suggesting an efficient estimator and adaptive inference that are both valid in all three cases and more generally uniformly valid in the variance of $\alpha_j(\tau)$.

Adaptive Estimation and Inference

To establish asymptotic results that are uniformly valid in the variance of \( \alpha_j(\tau) \), we allow \( \operatorname{Var}(\alpha_j(\tau)) \) (and consequently \( \Omega_2(\tau,\tau) \)) to be exactly zero, bounded away from zero, or a sequence converging to zero at an arbitrary rate. This approach nests all three pointwise cases described in the previous section. The sequence $a_n(\tau)=\frac{1}{1 + \operatorname{Var}(\alpha_j(\tau)) n}$ defined in Assumption (ref) plays a central role because it is proportional to the rate of convergence of the slow moments relative to the fast moments $\parallel \operatorname{Var}(\bar g_{mn,2}(\delta,\tau)) \parallel^{-1} \parallel \operatorname{Var}(\bar g_{mn,1}(\delta,\tau) ) \parallel = O_p(a_n(\tau))$. It is equal to $1$ without heterogeneity, converges to $0$ at the $n^{-1}$ rate with strong heterogeneity, and is always bounded between $0$ and $1$. All matrices that are functions of $a_n(\tau)$ may also depend on the sample size. When necessary, we make the notation explicit with the subscript $mn$. For example, let

equation[equation omitted — 145 chars of source]

where $W_{mn}(\tau)$ is defined in Assumption (ref).

Similarly to Fernandez-Val2022a, we define the covariance kernel of the limiting process of $\hat\delta(\cdot)-\delta(\cdot)$. For a given integer $T>0$, let $\mathcal{T}_T=(\tau_1,\dots,\tau_T)$ be an arbitrary $T$-dimensional vector on $\otimes_{t=1}^T\mathcal{T}$. Let

equation*[equation* omitted — 146 chars of source]

and $\Sigma_{mn}(\tau)=\Sigma_{mn}(\tau,\tau)$. The covariance kernel is now given by the limit of the elements of the following $(KT)\times (KT)$ matrix

equation*[equation* omitted — 72 chars of source]

where

equation*[equation* omitted — 155 chars of source]
assumption[Covariance kernel] For any integer $T>0$ and any $T$-dimensional vector $\mathcal{T}_T=(\tau_1,\dots,\tau_T)$ on $\otimes_{t=1}^T\mathcal{T}$, there is a $(KT)\times(KT)$ matrix $H$, such that, almost surely, \begin{equation*} \underset{m,n\rightarrow\infty}{\lim}H_{mn}=H. \end{equation*} In addition, there is $c_{\mathcal{T}_T}>0$ such that for the smallest eigenvalue we have \begin{equation*} \lambda_{\min}(H)>c_{\mathcal{T}_T}. \end{equation*}

We can now state the asymptotic distribution of our estimator that is uniformly valid in $\operatorname{Var}(\alpha_j(\tau))$.

comment\begin{theorem}[Adaptive asymptotic distribution] Assumptions (ref)-(ref)(c) and (ref) hold. In addition, $\left\lVert\hat G(\tau)-G_{mn}(\tau)\right\rVert=o_p(\sqrt{G_{mn}(\tau)})$ uniformly in $\tau\in\mathcal T$. It follows that \begin{equation*} \operatorname{diag}(\Sigma_{mn}(\cdot))^{-1/2}(\hat\delta(\cdot)-\delta(\cdot))\rightsquigarrow \mathbb Z(\cdot) \end{equation*} where $\mathbb Z$ is a centered Gaussian process with a covariance function $H(\tau_k,\tau_l)$ as the $(k,l)$ submatrix of $H$. \end{theorem}
theorem[Adaptive asymptotic distribution] Assumptions (ref)-(ref)(c) and (ref) hold. In addition, either Assumption (ref) or Assumption (ref) holds. Then, \begin{equation*} \operatorname{diag}(\Sigma_{mn}(\cdot))^{-1/2}(\hat\delta(\cdot)-\delta(\cdot))\rightsquigarrow \mathbb G(\cdot) \end{equation*} where $\mathbb G$ is a centered Gaussian process with a covariance function $H(\tau_k,\tau_l)$.

Theorem (ref) allows for both types of weighting matrices: the asymptotically full-ranked weighting matrix of Assumption (ref) or the heterogeneous weighting matrix of Assumption (ref). The first case covers the exactly identified case and 2SLS. As we show in Proposition (ref) below, our estimated efficient weighting matrix satisfies the second condition.

commentEven if we do not know its convergence rate, Theorem (ref) shows that the correctly normalized estimator converges to a tight Gaussian process with mean zero. More precisely, we show in the proof that the convergence rate of the $k^{\text{th}}$ element of $\delta(\tau)$ is given by \begin{equation} \zeta(k,\tau)=\frac{1}{\sqrt{mn}}+\frac{\sqrt{\operatorname{Var}(\alpha(\tau))}}{\sqrt m} \sum_{l = L_1 + 1}^L \sqrt{G_{mn,kl}(\tau)}. \end{equation} where $G_{mn,k}$ is the $k^{\text{th}}$ row of $G_{mn}$. We also show that \begin{equation*} \hat\delta_k(\tau)-\delta_k(\tau)=\sum_{j=1}^m d_j(k,\tau)+o_p\left(\zeta(k,\tau)\right) \end{equation*} where $\underset{\tau\in\mathcal T, k\in\{1,\dots,K\}}{\sup}o_p\left(\zeta(k,\tau)\Sigma_{mn}(k,\tau)^{-1/2}\right)=o(1)$, $\Sigma_{mn}(k,\tau)$ is the $(k,k)$ element of $\Sigma_{mn}(\tau)$, and $d_j(k,\tau)$ is the score function for group $j$. This implies that the error we make when we approximate the distribution of $\delta_k(\tau)$ by the Gaussian process is asymptotically negligible uniformly in the variance of $\alpha_j(\tau)$.

In order to use these results for inference, we must provide a consistent estimator of the asymptotic variance $\Sigma_{mn}(\tau,\tau')$. The crucial ingredient is the asymptotic variance of the sample moments that we denote by

equation*[equation* omitted — 84 chars of source]

Remember that $\Omega_1(\tau,\tau')$ arises from the first stage estimation and $\Omega_2(\tau,\tau')$ from the presence of the group-level heterogeneity $\alpha_j(\cdot)$. The difficulty resides in that the leading term may arise from the first-stage or second-stage estimation depending on the variance of $\alpha_j(\cdot)$ and the type of instrument ($L_1$ or $L_2$). This expression may suggest that we must estimate these two components separately. However, and perhaps surprisingly, we find that we can estimate the leading term of $\Omega_{mn}(\tau,\tau')$ adaptively with a traditional cluster-robust estimator of the variance:

equation[equation omitted — 235 chars of source]

where $\hat u_{ij}(\tau)$ are the second-stage residuals.

proposition[Properties of $\hat \Omega(\tau, \tau')$] Let assumptions (ref)-(ref) and (ref)(c) hold. Further, assume that as $m\rightarrow\infty$, for each $l,l'\in \{1,\dots,L\}$ and uniformly in $\tau,\tau'\in\mathcal{T}^2$, \begin{equation} m^{-1}\sum_{j=1}^m \mathbb{E} \left[\left(\bar z_{jl} \bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau')-\Omega_{mn,2ll'}(\tau,\tau')\right)^2\right]\rightarrow C_{l,l'}(\tau,\tau')<\infty \end{equation} where $C_{l,l'}(\tau,\tau')$ is continuous in $\tau$ and $\tau'$. The estimator used to compute $\hat u_{ij}(\tau)$ satisfies (i) $\hat\beta(\tau)-\beta(\tau)=O_p\left (1/\sqrt{mn}\right)$, (ii) $\hat\gamma(\tau)-\gamma(\tau) = O_p(1/\sqrt{mn}) + O_p(\sqrt{\operatorname{Var}(\alpha_j(\tau))}/\sqrt m)$ uniformly in $\tau$. Then, for any $ll'$ entry of the $\hat \Omega(\tau, \tau')$ matrix with $l,l' \in \{1, \dots , L \}$ we have uniformly in $\tau, \tau' \in \mathcal{T}^2$, \begin{align*} \hat \Omega_{ll'}(\tau,\tau')= \Omega_{mn, ll'}(\tau,\tau') + o_p \left ( \sqrt{ \Omega_{mn,ll} (\tau) \Omega_{mn,l'l'}(\tau') } \right ). \end{align*}

Beyond the numbered assumptions, two additional conditions are necessary. First, equation ((ref)) requires the average variance of $\bar z_j\alpha_j(\tau)$ to converge to a constant. This holds automatically when we sample identically distributed groups but is required because we allow for heterogeneity across groups. Second, we require that the coefficients on the individual-level variables are estimated at the $1/\sqrt{mn}$ rate and the coefficients on the group-level variable at the $1/\sqrt{mn} + \sqrt{\operatorname{Var}(\alpha_j(\tau))}/\sqrt{m}$ rate. This corresponds to assuming that $M_1 = K_1$ so that $\beta = \delta_1$. This assumption ensures that estimation errors in slow coefficients do not contaminate the variance estimate of fast ones. This condition can always be achieved, for instance, by using $\dot x_{1ij}$ as an instrument for $x_{1ij}$ and a weighting matrix that satisfies Assumption (ref). We show below that our efficient weighting matrix satisfies this restriction.\footnote{The only exception in this paper is the between estimator from Section (ref). In that case, all the coefficients are estimated at the slow rate $1/\sqrt{mn} + \sqrt{\operatorname{Var}(\alpha_j(\tau))}/\sqrt{m}$. Proposition (ref) in the Appendix demonstrates the consistency of $\hat\Omega(\tau,\tau')$ in this case as well.}

The suggested estimator of the variance is straightforward to implement because it does not require estimating directly the variance of the first-stage quantile regression coefficients. This is a noteworthy advantage because the variance of these coefficients depends on the conditional density of the outcome given the covariates, which can be difficult to estimate. To understand this surprising result, consider the case where $\alpha_j(\tau)=0$ for all groups. This implies that the entire variance comes from the first-stage estimation. The clustered estimator of the variance can be interpreted as a subsampling estimator where the groups represent the subsamples. The variance of the coefficients across groups provides information about the precision of the first-stage estimates. Similar results for the cross-sectional bootstrap are provided in Liao2018, Lu2022, and Fernandez-Val2022a.

Proposition (ref) shows we can consistently estimate the diagonal elements of $\Omega_{mn}(\tau)$ in the sense that the error that we make vanishes more quickly than the true variance. On the other hand, if two sample moments vanish at different rates, we might not be able to estimate their covariance consistently. However, this does not preclude testing hypotheses about combinations of coefficients because the error that we make when we estimate the covariance vanishes faster than the variance of the slowest coefficient. We formalize this result in Proposition (ref) below for the case of a single linear null hypothesis.\footnote{We can similarly extend this proposition to multiple non-linear hypotheses.}

Define the $TK\times 1$ vector $\delta(\tau_1,\dots, \tau_T) = \left( \delta(\tau_1)', \dots, \delta(\tau_T)' \right) '$ whose covariance matrix is given by the elements of the following $(KT)\times (KT)$ matrix:

equation*[equation* omitted — 82 chars of source]

For each $\tau, \tau' \in \mathcal T \times \mathcal T$, $\Sigma_{mn}(\tau, \tau')$ is estimated by

align*[align* omitted — 106 chars of source]

where $\hat \Omega(\tau, \tau')$ is defined in equation ((ref)). The following proposition shows that a straightforward z-test provides valid inference that is adaptive in the degree of heterogeneity.

proposition[Asymptotic normality of the test statistic] Let the conditions for Theorem (ref) and Proposition (ref) hold. In addition, let $\eta \in \mathbb{R}^{T\times K}$ with $||\eta|| > \varepsilon > 0$. Then, \begin{equation*} \frac{\eta' \left (\hat \delta(\tau_1,\dots, \tau_T) -\delta(\tau_1,\dots, \tau_T) \right) }{ \sqrt{\eta' \hat \Sigma \eta }} \xrightarrow{d} N(0,1). \end{equation*}

The asymptotic variance derived in Theorem (ref) depends on the weighting matrix. Following standard GMM arguments, the efficient weighting matrix is given by

equation[equation omitted — 103 chars of source]

This weighting matrix automatically takes into account the different rates of convergence of the different moments. If some moments converge faster than others, then this matrix, asymptotically, gives infinitely more weight to the fast moments than the slow moments, so the parameters identified by the fast moments will converge at the faster rate.

Usually, we would simply plug in a consistent estimator of $W^*(\tau)$ and obtain a feasible efficient GMM estimator, but Proposition (ref) shows that the off-diagonal elements of $\Omega_{mn}(\tau)$ might not be consistently estimated when the rates of convergence of the corresponding coefficients are different. Proposition (ref) shows that the estimated efficient weighting matrix nevertheless satisfies Assumption (ref) such that Theorem (ref) and Proposition (ref) apply to the feasible efficient GM estimator. In addition, Theorem (ref) reveals that the (first-order) asymptotic distribution of the estimator that uses $W^*(\tau)$ as the weighting matrix is the same as that of the estimator that uses $\hat W^*(\tau)$ as the weighting matrix.

proposition[Adaptive Efficiency of the GMM Estimator] Assume that the conditions for Proposition (ref) hold. Then, the estimated efficient weighting matrix $\hat W^* (\tau) =\hat\Omega(\tau)^{-1}$ satisfies Assumption (ref).

If there are more moment conditions than parameters to estimate ($L > K$), it is possible to implement overidentification tests in the spirit of Sargan1958 and Hansen1982. More precisely, we can test the validity of the instrumental variables while maintaining the other assumptions. The test statistic is the GMM objective function evaluated at the efficient GMM estimator:

equation[equation omitted — 245 chars of source]

Proposition (ref) demonstrates the adaptive validity of the $J$-test implemented with the clustered covariance matrix.

proposition[Overidentification Test] Under the $\mathbb{H}_0: \ \mathbb{E}[z_{ij}\alpha_i(\tau)] = 0$ for all $\tau\in \mathcal T$ and the assumptions required for Proposition (ref), as $m \rightarrow \infty$, $J\left(\hat\delta\left(\hat W^*(\tau),\tau\right),\tau\right) \xrightarrow{d} \chi_{L-K}^2$.

We prove the result for a single quantile, but the proof easily extends to multiple quantiles. In Section (ref), we use this overidentification test to suggest a quantile analog of the Hausman test for the exogeneity of the between variation.

Grouped (IV) Quantile Regression Model

Chetverikov et al. (2016)

Chetverikov2016 introduce an IV quantile regression estimator for group-level treatments and propose an alternative two-step estimation procedure for model ((ref)). In the first stage, they perform quantile regression separately for each group $j$ and quantile $\tau$, regressing $y_{ij}$ on $x_{1ij}$ and a constant. In the second stage, they regress the intercept from the first stage on $x_{2j}$. The key distinction between their approach and ours is that CLP use the estimated intercept from the first stage as the dependent variable in the second stage, whereas we use the fitted values.

The intercept represents the fitted value for an observation with $x_{1ij} = 0$. Consequently, their estimator is not invariant to linear reparametrizations of the individual-level regressors. In finite samples, results can differ depending on how we code the variables. For instance, the estimated coefficients may change if age is recorded as years since birth versus years since 16 or if we measure temperature in Celsius versus Fahrenheit. Our estimator, by contrast, does not suffer from this issue. Beyond the undesirable dependence of the results on an arbitrary linear transformation of the individual-level variables, this property also affects the precision of the estimates and introduces potential misspecification bias.

Figure (ref) illustrates these concerns using artificial data. Panel (a) shows how the variance of the fitted values increases as we move away from the mean of the regressor. The solid line represents the median regression estimate, while the shaded area depicts the 95% confidence interval for the median fitted values. The intercept corresponds to the fitted value for an $x$ value that lies far outside the observed support of the variable. As a result, its variance is relatively large and would increase even more if we shifted the location of the individual-level variable.

Panel (b) highlights the misspecification bias using a similar artificial dataset. The dashed orange line represents the true median regression function, while the solid line shows the estimated linear model, which is slightly misspecified. Over the support of the covariates, the misspecification bias remains small because quantile regression minimizes the weighted mean squared error (see Angrist2006a). However, the intercept is more strongly biased, as it lies outside the covariate’s support, and this bias monotonically worsens as we increase the mean of the individual-level covariate.

figure[figure omitted — 30,475 chars of source]

Simulations

In this subsection, we compare the MD and CLP estimators of $\gamma(\tau)$ using Monte Carlo simulations. In the following subsection, we formally compare their asymptotic distributions to explain the simulation results. We use the same three data-generating processes (DGP) as in Chetverikov2016 and generate the outcome as follows:

equation[equation omitted — 128 chars of source]

where $x_{1ij}$ and $x_{2j}$ are $\textnormal{exp}(0.25 \cdot N[0,1])$ and individual heterogeneity is introduced via the rank variable $u_{ij}\sim U[0,1]$. The quantile coefficient functions are $\gamma(\tau)=\beta (\tau)=\sqrt{\tau}$ and $\beta_1(\tau) = \frac{\tau}{2}$ for $\tau\in [0,1]$. In the first DGP, we set $\alpha_j(u_{ij})=0$, so there is neither group heterogeneity nor endogeneity. In the second DGP, we introduce group-level heterogeneity as

equation[equation omitted — 84 chars of source]

where $\eta_{j}\sim U(0,1)$. There is no endogeneity because the group effects are uncorrelated with the regressors. As group-level heterogeneity is multiplied with the rank variable $u_{ij}$, there is only weak group heterogeneity in the lower tail of the distribution and strong heterogeneity in the upper tail. Finally, in the third DGP, we introduce endogeneity as

equation[equation omitted — 65 chars of source]

where $z_j$ and $\nu_j$ are each distributed $\textnormal{exp}(0.25 \cdot N[0,1])$. $z_j$ is a valid instrumental variable for the endogenous $x_{2j}$. We use the same sample sizes as CLP, that is, $(m,n) = \{(25,25),\allowbreak (200,25),\allowbreak (25,200),\allowbreak (200,200) \}$. We perform 10,000 Monte Carlo replications for the set of quantiles $ \tau \in \{0.1, 0.5, 0.9\}$. Since the CLP estimator does not directly estimate $\beta(\tau)$, we present only results for $\gamma(\tau)$.

table[table omitted — 3,350 chars of source]
table[table omitted — 2,827 chars of source]

The first stage, group-by-group quantile regressions, are the same for both estimators. In the second stage, CLP regress the estimated intercepts on $x_{2j}$ with OLS (DGP 1 and 2) or using $z_j$ as an instrument (DGP 3). We implement the MD estimator with the IV estimator instrumenting $(x_{1ij},x_{2j})$ with $(\dot x_{1ij},x_{2j})$ (DGP 1 and 2) or $(\dot x_{1ij},z_j)$ (DGP 3). With this choice of instruments, we do not exploit the between variation of $x_{1ij}$.\footnote{Using $x_{1ij}$ instead of $\dot x_{1ij}$ as an instrument has virtually no effect on the results since there is no variation in $x_{1ij}$ across groups.} Since the data generating process of Chetverikov2016 has a weak instrument when $m$ is small, one should pay attention when looking at the simulation results for the endogenous case.\footnote{With $m = 25$ in over 40% of the draws, the F-statistics of the first stage of the IV estimations is below 10. The issue disappears when $m = 200.$} It would be straightforward to compute weak instrument robust inference, for instance, using Anderson/Rubin confidence intervals in the second stage. However, we consider this to be outside the scope of this paper.

Table (ref) presents the bias, standard deviation, and the relative MSE of the MD estimator, defined as the MSE of the MD estimator divided by that of the CLP estimator. No clear pattern emerges regarding the bias of the estimators. Consistent with asymptotic theory, the bias diminishes as the number of observations increases. We observe more pronounced differences in the variance of the estimators. In the homogeneous case, the standard deviation of the MD estimator is four times smaller than that of the CLP estimator. This difference remains stable as the number of groups or individuals increases. The advantage of the MD estimator remains similar at the bottom of the distribution in the exogenous case, but the difference in variance narrows at the upper end. We explain this with the presence of weak heterogeneity at the bottom of the distribution and strong heterogeneity at the top. At the upper end of the distribution, the variance of the estimators converges as $n$ becomes large. Finally, in the endogenous case, the results resemble those of the exogenous case once we exclude findings affected by the weak instrument issue, precisely when $m=25$. The differences in precision explain the large discrepancies in MSE. The MSE of the CLP estimator is up to twenty times larger than that of the MD estimator when $\alpha_j(\tau) = 0$ and remains substantially larger in all scenarios considered except one with a weak instrument.\footnote{Although not shown here, we also computed the traditional quantile regression estimator. When there is heterogeneity, quantile regression is inconsistent for the parameter of interest; however, when $\alpha_j(\tau) = 0$, our estimator and quantile regression are practically indistinguishable in terms of bias and variance while the CLP estimator has a much larger variance.}

Table (ref) shows the performance of the $95\%$ confidence intervals suggested with our inference procedure. The table reports the coverage rate and the median length of the intervals of our estimator relative to that of the CLP estimator.\footnote{For the CLP estimator, we use heteroskedasticity robust standard errors as suggested in Chetverikov2016. Recall that the CLP estimator uses only one observation per group in the second stage. Hence, we would attain the same standard errors if we kept all observations and clustered the standard errors at the group level.} Our suggested inference procedure has coverage close to $95\%$ in all cases. Compared to the CLP estimator, our confidence bands are substantially shorter. In most cases, our estimator yields confidence bands that are less than half the length of those for the CLP estimator. This difference becomes even more pronounced in our empirical application in Section (ref), where the CLP estimator produces confidence bands that are, on average, 14 times wider than those of the MD estimator.

Comparison of the asymptotic distributions of CLP and MD estimators

CLP focus exclusively on $\gamma(\tau)$, the coefficients on the group-level variables. They assume strong group-level heterogeneity, imposing that $\operatorname{Var}(\alpha_j(\tau)) > 0$ uniformly in $\tau$.\footnote{Although they do not state this assumption explicitly, their asymptotic distribution would become degenerate without it, as they discuss in footnote 9.} Thus, their asymptotic distribution corresponds to case (i)-b of our Theorem (ref), where the variance from the first stage diminishes more rapidly than that from the second stage. In this subsection, we compare the variance of the CLP and MD estimators, allowing for weak or no heterogeneity.

To simplify notation, we consider the exogenous case where $x_{2j}$ can serve as its own instrumental variable, though the results also extend to the endogenous case. As discussed in Section (ref), the asymptotic variance consists of two components: one accounting for first-stage error and the other for second-stage noise. We will examine each component separately. First, we consider the scenario where the true first-stage coefficients are known to isolate the variance arising in the second stage. In this scenario, the CLP point estimates can be obtained numerically within our MD framework by regressing the true first-stage fitted values on $x_{1ij}$ and $x_{2j}$, using $\dot x_{1ij}$ and $x_{2j}$ as instruments. Thus, the second-stage variance resulting from the randomness of $\alpha_j(\tau)$ is identical for both estimators.

We can isolate the first-stage error by setting $\operatorname{Var}(\alpha_j(\tau))=0$. Under this condition, both estimators are classical MD estimators. We can express the CLP estimator as

equation[equation omitted — 242 chars of source]

where

equation*[equation* omitted — 178 chars of source]

and $l_j$ is a $m$-dimensional vector of zeros with a $1$ in the $j$ position. The restriction matrix $\tilde R_j$ differs from the restriction matrix of our estimator defined in equation ((ref)), as it does not impose equality of the first stage coefficients implied by the model. Thus, our estimator imposes $K_1\cdot (m-1)$ additional correct restrictions. A second difference is that CLP use an identity weighting matrix while we use $\tilde X_j'\tilde X_j$. When there is no group heterogeneity, we are in the classical MD framework, where the efficient weighting matrix is the inverse of the first-stage variance. It follows that weighting by $\tilde X_j'\tilde X_j$ is efficient when the first-stage error is separable and the density of $y$ given $x$ at the $\tau$ quantile is the same across groups. When, in addition, $\tilde X_j'\tilde X_j$ is constant across groups (balanced panel, identical distribution of $x_{1ij}$), then equal weighting is efficient.

Adding valid constraints within an efficient MD framework reduces variance. Therefore, if $\tilde{X}_j'\tilde{X}_j$ is the efficient weighting matrix, our estimator will necessarily have lower first-stage variance than the CLP estimator. However, with an inefficient weighting matrix, adding valid constraints might increase the variance in some relatively pathological cases.\footnote{See the discussion in Section 8 of Hansen2021.} While it would be possible to estimate the efficient weighing matrix, we prefer to avoid estimating the first-stage variance and instead opt for a more interpretable estimator in cases of misspecification.

To summarize the comparison, the MD estimator using $\dot{x}_{1ij}$ and $x_{2j}$ as instrumental variables exhibits the same variance due to the randomness of $\alpha_j(\tau)$ but a lower variance due to estimation of $\hat\beta_j(\tau)$ compared to the CLP estimator—this holds formally when the efficient weighting matrix is used. In light of these results, we can understand the simulation results. In the first DGP, all the variance arises from the first stage such that our estimator is more precise, even asymptotically. The relative MSE does not change as $n$ increases. In the second and third DGP, the variances of both estimators will converge as $n\rightarrow \infty$ because the second-stage variance will asymptotically dominate them. This convergence appears, however, to be relatively slow, especially at the bottom of the distribution, where heterogeneity is weak.

Note that CLP also consider a generalization of model ((ref)) in which they assume

align[align omitted — 204 chars of source]

By default, $\beta_{j,1}(\tau)$ is the first element of the vector $\beta_{j}(\tau)$, but it could represent any element of this vector. In contrast to model ((ref)), this approach allows the coefficient on $x_{1ij}$ to vary across groups. It also enables researchers to estimate interaction effects between group-level treatments and individual-level covariates. However, if the effect of a group-level variable varies with individual-level variables, the effect of $x_{2j}$ on the intercept no longer represents an average effect. Instead, it reflects the effect evaluated at $x_{1ij}=0$, which may not be meaningful or of interest. To address this issue, researchers would need to estimate all relevant interaction effects and combine them appropriately. This approach has not been discussed in CLP, nor has it been implemented in their simulations or in any applications of their estimator.

comment\section{Grouped (IV) Quantile Regression Model} \begin{figure} \begin{subfigure}[b]{0.49\textwidth} \resizebox{1\linewidth}{!} \caption{ Extrapolation } \end{subfigure} \begin{subfigure}[b]{0.49\textwidth} \resizebox{1\linewidth}{!} \caption{ Misspecification } \end{subfigure} \caption{First Stage Regressions} \floatfoot{The figure illustrates the venerability of the intercept to extrapolation and misspecification. Both panels show generated data for one group and a first-stage fit. Panel (a) uses the same DGP as in the simulation of CLP. The solid line shows the regression line estimated by median regression, and the shaded area shows the 95% confidence interval. Panel (b) uses a different DGP where $y = 15 - 0.5 x - 0.2x^2 + u$, where $u \sim N(0, 1), x \sim N(3, 1)$. The solid line is estimated by median regression without the quadratic term (misspecified). The dashed orange line is the true regression line.} \end{figure} \subsection{Chetverikov et al. (2016)} Chetverikov2016 propose an IV quantile regression estimator for group-level treatments. Their first model aligns with our model ((ref)). Since their simulations and all applications of their estimator adhere to this initial model, we begin by comparing their estimator with ours within this framework. Following that, we examine whether their second, more general model addresses some of the limitations of their estimator. For model ((ref)), CLP suggest an alternative two-step estimator. In the first stage, they perform quantile regression separately for each group $j$ and quantile $\tau$, regressing $y_{ij}$ on $x_{1ij}$ and a constant. In the second stage, they regress the intercept from the first stage on $x_{2j}$. The key difference between their approach and ours is that CLP use the estimated intercept from the first stage as the dependent variable in the second stage, whereas we use the fitted values. Let’s first compare the theoretical results. CLP focus exclusively on $\gamma(\tau)$, the coefficients on the group-level variables, and do not directly estimate $\beta(\tau)$. They assume strong group-level heterogeneity,\footnote{Although this assumption is not explicitly stated, their asymptotic distribution would become degenerate without it, as they discuss in footnote 9.} assuming $\operatorname{Var}(\alpha_j(\tau)) > 0$ uniformly in $\tau$. As a result, their theoretical findings are not adaptive with respect to $\operatorname{Var}(\alpha_j(\tau))$. In other words, their asymptotic distribution corresponds to case 1-a) in our Theorem (ref), where the variance from the first stage diminishes more rapidly than that from the second stage. Consequently, their asymptotic distribution is the same as if the true first-stage coefficients $\beta_j(\tau)$ were known. On a more technical note, when we focus on the coefficients of the group-level variables and assume group-level heterogeneity, we are able to relax the growth rate condition imposed by CLP, building on recent results from Volgushev2019. Next, we compare the variance of our estimator to that of the CLP estimator. To simplify notation, we consider the exogenous case where $x_{2j}$ can serve as its own instrumental variable, though the results extend to the endogenous case as well. As discussed in Section (ref), the asymptotic variance consists of two components: one accounting for first-stage error and the other for second-stage noise. We will examine each component separately. To isolate the variance arising in the second stage, we consider the scenario where the true first-stage coefficients are known. In this scenario, the CLP point estimates can be obtained numerically within our MD framework by regressing the true first-stage fitted values on $x_{1ij}$ and $x_{2j}$, using $\dot x_{1ij}$ and $x_{2j}$ as instruments. Thus, the second-stage variance resulting from the randomness of $\alpha_j(\tau)$ is identical for both estimators. Our framework allows exploiting, in addition, the between-group variation in $\bar x_{1j}$ using the efficient GMM estimator, which will (weakly) reduce the variance of the estimator. Other than this advantage, our approach is neither better nor worse than the CLP estimator concerning second-stage variance. We can isolate the first-stage error by setting $\operatorname{Var}(\alpha_j(\tau))=0$. Under this condition, our estimator becomes a classical MD estimator. To facilitate a comparison between both estimators, we can express the CLP estimator as a minimum distance estimator based on the same first-stage estimates: \begin{equation} \hat\delta_{CLP}(\tau)=\underset{\delta\in \mathbb R^{K_1\cdot m +K_2}}{\arg\min} \frac{1}{m} \sum_{j = 1}^m \left (\hat \beta_j (\tau) - \tilde R_j \delta \right )' \left (\hat \beta_j (\tau) - \tilde R_j \delta \right ), \end{equation} where \begin{equation*} \underset{\scriptscriptstyle (K_1+1) \times ( K_1 \cdot m + K_2)}{ \tilde R_j} = \begin{pmatrix} 0 & x_{2j}' \\ l_j' \otimes I_{K_1} & 0 \end{pmatrix}, \end{equation*} and $l_j$ is a $m$-dimensional vector of zeros with a $1$ in the $j$ position. The restriction matrix $\tilde R_j$ is different from the restriction matrix of our estimator defined in equation ((ref)), as it does not impose equality of the first stage coefficients implied by the model. Thus, our estimator imposes $K_1\cdot (m-1)$ additional restrictions that are correct. A second difference is that CLP use an identity weighting matrix while we use $\tilde X_j'\tilde X_j$. When there is no group heterogeneity, we are in the classical MD framework such that the efficient weighting matrix is the inverse of the first-stage variance. It follows that weighting by $\tilde X_j'\tilde X_j$ is efficient when the first-stage error is separable and the density of $y$ given $x$ at the $\tau$ quantile is the same across groups. When, in addition, $\tilde X_j'\tilde X_j$ is constant across groups (balanced panel, identical distribution of $x_{1ij}$), then equal weighting is efficient. Adding valid constraints within an efficient MD framework reduces variance. Therefore, if $\tilde{X}_j'\tilde{X}_j$ is the efficient weighting matrix, our estimator will necessarily have lower first-stage variance than the CLP estimator. However, with an inefficient weighting matrix, adding valid constraints might, in some relatively pathological cases, increase the variance.\footnote{See the discussion in Section 8 of Hansen2021.} While it is possible to develop a more efficient estimator by following the ideas discussed in Section (ref), we prefer to avoid estimating the first-stage variance and instead opt for a more interpretable estimator in cases of misspecification. To summarize the comparison, the MD estimator using $\dot{x}_{1ij}$ and $x_{2j}$ as instrumental variables exhibits the same variance due to the randomness of $\alpha_j(\tau)$ but a lower variance due to estimation of $\hat\beta_j(\tau)$ compared to the CLP estimator—this holds formally when the efficient weighting matrix is used. Note that CLP assume strong group heterogeneity so that only the variance arising in the second stage matters asymptotically. Consequently, it’s not surprising that they did not optimize their estimator to minimize the variance arising in the first stage. Intuitively, the first-stage intercept used by CLP represents the fitted value at an $x_{1ij}$ value that may lie outside of the variable’s support. For example, in a wage regression where one of the individual-level variables is $age$, the variance of the fitted value at $age = 0$ is higher than that at a value near the center of the $age$ distribution. Additionally, even a slight misspecification of the conditional quantile function is more likely to have a significant impact when we extrapolate beyond the support of the variable. Figure (ref) illustrates these issues with artificial data. CLP also consider a generalization of model ((ref)) in which they assume \begin{align} Q(\tau, y_{ij}| x_{1ij} , x_{2j}, v_j) =& \tilde {x}_{1ij}' \beta_j (\tau) \\ \beta_{j,1}(\tau) =& x_{2j}' \gamma(\tau) + \alpha(\tau, v_j). \end{align} In contrast to model ((ref)), this approach allows the coefficient on $x_{1ij}$ to vary across groups. By default, $\beta_{j,1}(\tau)$ is the first element of the vector $\beta_{j}(\tau)$, but it could represent any element of this vector. This allows the researcher to also estimate the interaction effects of the group-level treatment and a micro-level covariate. However, if the effect of a group-level variable varies with individual-level variables, the effect of $x_{2j}$ on the intercept will not represent an average effect, but the effect evaluated at $x_{1ij}=0$, which may not be of interest. In addition, the estimator of the effect on the intercept suffers from all the issues discussed above concerning the variance of the CLP estimator. To solve the problem, the researcher should estimate all interaction effects and combine them. This approach has not been discussed in CLP and has not been implemented in their simulations or in any application of their estimator. They suggest a two-step estimator. The first stage consists of regressing $y_{ij}$ on $x_{1ij}$ and a constant using quantile regression separately for each group $j$ and quantile $\tau$. In the second stage, they regress the intercept from the first stage on $x_{2j}$. Their estimator focuses on estimating $\gamma(\tau)$ and does not directly estimate $\beta(\tau)$. The main difference compared to our estimator is that CLP use the estimated intercept of the first stage (i.e., fitted values evaluated at $x_{1ij} = 0$) as a dependent variable in the second stage. Thus, as we discuss more in detail below, the CLP estimator is consistent for the quantile treatment effect at $x_{1ij} = 0$. The CLP estimator might have lower precision compared to the MD estimator. First, it only includes the individual-level covariates $x_{1ij}$ in the first stage, thereby not exploiting (potentially exogenous) variation in the individual-level covariates between groups. Second, it does not impose equality of the $\beta_j(\tau)$ in the first stage. Third, the estimator might extrapolate the intercept in the first stage, making it more vulnerable to misspecification and first-stage estimation error in finite samples. Figure (ref) illustrates these issues. The figure shows two groups from two different samples. Each panel shows an estimated first-stage regression line (solid dark line) and the true regression line (dashed orange line). Both panels show that if the support of the covariates does not cover zero in all groups, the intercept is extrapolated. As shown in Panel (a), a small estimation error in the slope parameter can lead to large estimation errors in the intercept. Further, the value of the (true) intercept and its estimation error change with the reparametrization of the individual-level covariates so that reparametrization leads to different results. In short, if the first stage slopes are allowed to change over groups, reparametrization of the covariates changes the estimand. By imposing equality of the slopes across groups, our estimator becomes invariant to reparametrizations of the covariates. Further, when $x_{1it}$ contains discrete variables, if some groups do not contain any observations in the base category or if some variables exhibit variation only within some group, a regression of the intercepts does not provide a meaningful comparison.\footnote{This is the case in our empirical application. To solve this issue, we exclude some groups from the analysis when using the CLP estimator.} Panel (b) shows how model misspecification in the first stage can lead to a large estimation error with the CLP estimator. The first stage regression is misspecified, as it fits a linear regression model instead of a quadratic one. The consequences of misspecification are substantially larger outside the support of the covariates, e.g., at $x_{1ij} = 0$. In comparison, the misspecification error in the fitted values is negligible. Now, we consider both estimators in the context of models ((ref)) and model ((ref))-((ref)). If model ((ref)) is correct, our estimator will, in general, outperform the CLP estimator for the reasons listed above. Further, imposing equality also makes the estimator invariant to reparametrizations of the individual-level variables. For this model, the CLP estimator is consistent for the treatment effect at $x_{1ij}=0$, which equals the quantile treatment effects. We obtain, however, a more precise estimator by estimating all the parameters simultaneously and imposing all the assumptions. On the other hand, if the model ((ref))-((ref)) is correct, we should not exploit the between-variation in the individual-level variables and use the demeaned individual-level regressors as instruments. Whereas, if the slopes are systematically correlated with the treatment variable, the treatment effect is heterogeneous, and CLP estimates the quantile treatment effects at $x_{1ij}=0$, which may not be particularly interesting. In such a case, one could parametrize the treatment effect on the random slope and estimate the effect on the intercept and the effect on the slope separately. In a second step, combining both estimates gives, for instance, an average (in $ x_{1ij}$) QTE. Using our approach, we estimate the best linear approximation of the treatment effect, and we can also allow for heterogeneous effects in a more natural way by including interaction terms between $ x_{1ij} $ and $x_{2j}$. We want to show that asymptotically our estimator has lower variance than the CLP estimator uniformly over different values of $\operatorname{Var}(\bar z_j \alpha_j(\tau))$.\footnote{The asymptotic results in Chetverikov2016 implicitly assume that $\operatorname{Var}(\bar z_j \alpha_j(\tau)) > \varepsilon > 0$ so that their estimator converges at the $\sqrt{m}$ rate.} From section (ref) we know that the variance comprises two terms, one that accounts for first-stage error and the other accounts for second-stage noise. The sum of these two terms determines the asymptotic behavior of the estimator. Thus, to show that our MD estimator is more precise, it suffices to show that both components of the variance are smaller. If $\operatorname{Var}(\alpha_j(\tau)) = 0$, there is no second stage noise, both estimators converge at the $\sqrt{mn}$ rate, and only the variance coming from the first stage matters. On the other hand, if $\operatorname{Var}(\alpha_j(\tau)) > \varepsilon > 0$, and $\operatorname{Var}(\bar z_j ) >\varepsilon > 0$, both estimators converge at the $\sqrt{m}$ rate, and the variance coming from the first stage does not enter the first-order asymptotic distribution. Thus, in the latter case, we will consider the estimators as if we knew the true first stage. In the following, we assume that the more widely used model ((ref)) is correct and focus on a case with exogenous regressors. We consider our MD estimator implemented using optimal instruments, as this estimator simultaneously minimizes both components of the variance. To apply optimal instruments, we need to impose the stronger assumption that $\mathbb{E}[\alpha_j (\tau) | X_j ] = 0$. Below, we provide some results without this assumption. First, we consider the variance arising from the first stage. To study this part of the variance, we can assume, without loss of generality, that $\bar g_{mn}^{(2)}(\hat \delta, \tau) = 0$. The optimal instrument is $Z_j^* = \left ( \tilde X_j V_j(\tau) \tilde X_j'\right )^+ X_j$, which yields an estimator that is algebraically identical to the efficient minimum distance estimator presented in Remark (ref) (see Proposition (ref)). The CLP estimator can also be written as a minimum distance estimator that minimizes \begin{equation} \frac{1}{m} \sum_{j = 1}^m \left (\hat \beta_j (\tau) - \tilde R_j \delta(\tau) \right )' \left (\hat \beta_j (\tau) - \tilde R_j \delta(\tau) \right ), \end{equation} where \begin{equation*} \underset{\scriptscriptstyle (K_1+1) \times ( K_1 \cdot m + K_2)}{ \tilde R_j} = \begin{pmatrix} x_{2j}' & 0 \\ 0 & l_j' \otimes I_{K_1} \end{pmatrix}, \end{equation*} and $l_j$ is a $m$-dimensional vector of zeros with a $1$ in the $j$ position. The restriction matrix $\tilde R_j$ is different from the restriction matrix of our estimator, as it does not impose equality of the first stage coefficients implied by the model. Further, CLP use an identity weighting matrix so that their estimator is inefficient relative to an efficient MD estimator with restriction matrix $\tilde R_j$. Since our estimator imposes the additional (correct) restriction, our efficient MD estimator has a smaller variance than any alternative (efficient) MD estimator with restriction matrix $\tilde R_j$, including the CLP estimator. In the special case of quantile independence, the weighting matrix of the efficient MD estimator reduces to $\hat W_j = \tilde X_j'\tilde X_j$, which corresponds to using OLS in the second stage. In this case, our estimator with a least squares second stage is efficient and will have a lower variance than the CLP estimator. Hence, if the estimators converge at a fast rate, our estimator provides a first-order improvement. Whereas, if the estimators converge at the slow rate, the improvement is of second order. Next, we focus on the component of the variance coming from the second-stage error. For this term, we can assume that we know the true first stage. We start by noting that we numerically obtain the CLP estimator by regressing the first stage fitted values on $x_{2j}$, $x_{1ij}\cdot d_1, \dots, x_{1ij}\cdot d_m$ with instruments $x_{2j}$, $\dot x_{1ij} \cdot d_1, \dots, \dot x_{1ij}\cdot d_m$ where $d_j$ is a group indicator. In the special case where we know the true first stage, we can recover the CLP point estimates if we regress the fitted values on $x_{1ij} $ and $x_{2j}$ with instruments $\dot x_{1ij} $ and $x_{2ij}$ without the interactions. From this representation, it is clear that the CLP estimator only exploits the within variation of $x_{1ij}$. Differently, our estimator uses the entire variation of $x_{1ij}$ efficiently. If we know the true first stage, the optimal instrument implied by the conditional moment restriction is $ Z_j^* = \mathbb{E}[ \alpha_j(\tau)^2 |X_j]^{-1} X_j$, which implies that our second stage is a GLS regression which is efficient. One backdrop of this analysis is that it relies on the stronger conditional moment restriction $\mathbb{E}[\alpha_j (\tau) | X_j ] = 0$. Nonetheless, we can show that regardless of the value of $\operatorname{Var}( \alpha_j(\tau))$ and $\operatorname{Var}(\bar z_j)$, we can implement an estimator that is more precise than the CLP estimator. More precisely, if $\operatorname{Var}(\alpha_j(\tau)) = 0$ for all $j$ or $\operatorname{Var}(\bar z_j) = 0$, the efficient minimum distance is optimal. Differently, if $\operatorname{Var}(\alpha_j(\tau)) > \varepsilon > 0$ and $\operatorname{Var}(\bar z_j) > \varepsilon > 0$ using an efficient GMM estimator with instruments $\dot x_{1ij}, \bar x_{1j}, x_{2j}$ yields more precise point estimates as it exploits all moment conditions efficiently. This GMM estimator exploits the between variation of $x_{1ij}$ by including $\bar x_{ij}$ in the instrument set. By adding an instrument, asymptotically, our estimator will have a weakly lower variance (see Proposition 4.51 in White2001). \subsection{Simulations} This subsection presents Monte Carlo simulations comparing the MD and CLP estimators. The simulations are based on the same data generating process and sample sizes as in Chetverikov2016. That is, $(m,n) = \{(25,25), (200,25), (25,200), (200,200) \}$. For both estimator we use a OLS (or 2SLS) second stage. The generated data include one individual-level regressor, one group-level regressor, and one instrument. Heterogeneity is introduced via a rank variable $u_{ij}$, and the data is generated as follows: \begin{equation} y_{ij} = \beta_{0}(u_{ij}) + x_{1ij} \beta(u_{ij}) + x_{2j} \gamma(u_{ij}) + \alpha_j(u_{ij}) , \end{equation} \begin{equation} z_j = x_{2j} + \eta_j + \nu_j , \end{equation} \begin{equation} \alpha_j(u_{ij}) = u_{ij} \eta_j - \frac{u_{ij}}{2}, \end{equation} where $x_{1ij}, x_{2j}$ and $\nu_j$ are distributed $\textnormal{exp}(0.25 \cdot N[0,1])$ and $\eta_{j}$ as well as the rank variable $u_{ij}$ are $U[0,1]$ distributed. The data generating process implies that $\mathbb{E}[\alpha (u_{ij})| x_{2j}] = \mathbb{E}[u_{ij} \eta_j - \frac{u_{ij}}{2}| x_{2j}] = \mathbb{E}[ \frac{u_{ij}}{2} - \frac{u_{ij}}{2}| x_{2j}] = 0$. At quantiles $\tau \in (0,1)$, the true parameters $\gamma(\tau) $ and $\beta (\tau)$ equal $\sqrt{\tau}$ and, $\alpha_1(\tau) = \frac{\tau}{2}$. Consequently, $\gamma(u_{ij}) = \beta (u_{ij})=\sqrt{u_{ij}}$ and $\beta_0(u_{ij})= \frac{u_{ij}}{2}$. \\ The simulations consider three cases. In the first one, $\alpha_j(\tau) = 0$ for all $j$ and all $\tau$. In this case, as there are no group effects, conditioning on the group does not affect the quantile function, and quantile regression is consistent for the same parameter. Further, both estimators are $\sqrt{mn}$-consistent. In the second case, there are group-specific effects ($\alpha_j(\tau) \neq 0$), which are uncorrelated with the regressors. As individual heterogeneity is multiplied with the rank variable $u_{ij}$, there is only weak group heterogeneity in the lower tail of the distribution, and we see faster convergence there. In the third case, $\alpha_{j}(\tau)$ is correlated with the regressor of interest, such that $x_{2j}$ is endogenous. For the implementation of the MD estimator, we use $\dot x_{1ij}$ as an instrument for $x_{1ij}$ so that, as with the CLP estimator, we do not exploit the between variation of the regressor.\footnote{Using the entire variation of $x_{1ij}$ does not affect the results since there is no between variation in $x_{1ij}$.} In the third case, as $x_{2j}$ is endogenous, we instrument $x_{2j}$ with $z_j$. Since the data generating process of Chetverikov2016 has a weak instrument when $m$ is small, one should pay attention when looking at the simulation results for the endogenous case.\footnote{With $m = 25$ in over 40% of the draws, the F-statistics of the first stage of the IV estimations is below 10. The issue disappears when $m = 200.$} In empirical research, it is straightforward to construct confidence intervals that are valid even if identification is weak. We perform 10,000 Monte Carlo replications for the set of quantiles $ \tau \in \{0.1, 0.5, 0.9\}$. Since the CLP estimator does not directly provide an estimate for $\beta(\tau)$, we present only results for $\gamma(\tau)$. Table (ref) reports the bias, standard deviation, and relative MSE of the estimators. The relative MSE reports the MSE of the MD estimator relative to that of the CLP estimator. Thus, a number smaller than 1 indicates that the MD estimator has a lower MSE. The CLP estimator seems to have a smaller bias than the MD estimator when $n = 25$. When $n$ increases to 200, the difference disappears. There are more remarkable differences in the variance of the estimators. The standard deviation of the MD estimator is four times smaller compared to that of the CLP estimator in the homogeneous case. The difference is somewhat smaller in the exogenous and endogenous cases but remains substantial.\footnote{The standard deviations in the endogenous case with $m = 25$ should be interpreted with caution due to the weak instrument.} The differences in the standard deviations of the estimator in the different cases also reflect the precision improvement of the MD estimator compared to the CLP estimator discussed above. The disparity is most remarkable in the case without group-level heterogeneity, where both estimators converge at the $\sqrt{mn}$ rate as we provide a first-order improvement compared to CLP. Differently, in the exogenous case, where $\operatorname{Var}(\alpha(\tau))$ increases over $\tau$ and therefore the convergence rate approaches $\sqrt{m}$, we see that the improvement becomes smaller at higher quantiles and with larger $n$. This difference in precision explains the large discrepancies in MSE. The MSE of the CLP estimator is over ten times larger than that of the MD estimator when $\alpha_j(\tau) = 0$ and remains substantially larger in all scenarios considered. If $\alpha_j(\tau) = 0$, quantile regression is a consistent estimator for $\beta(\tau)$. Although not shown here, simulation results comparing our estimator with traditional quantile regression show that the two estimators are indistinguishable in terms of bias and variance in large samples. Table (ref) show the performance of the $95\%$ confidence intervals suggested with our inference procedure. The table reports the coverage rate and the median length of the intervals of our estimator relative to that of the CLP estimator.\footnote{For the CLP estimator, we use heteroskedasticity robust standard errors as suggested in Chetverikov2016. Recall that the CLP estimator uses only one observation per group in the second stage. Hence, we would attain the same standard errors if we kept all observations and clustered the standard errors at the group level.} Our suggested inference procedure has coverage close to $95\%$ in all cases. Compared to the CLP estimator, our confidence bands are substantially shorter. In most cases, our estimator yields confidence bands less than half the length of those for the CLP estimator. This difference becomes even more pronounced in our empirical application in Section (ref), where the CLP estimator produces confidence bands that are, on average, 14 times wider than those of the MD estimator.

Traditional Quantile Panel Data Estimators

Fixed Effects, Random Effects and Between Estimators

In this subsection, we apply our results to derive quantile analogs of the fixed effects, between, random effects, and Hausman-Taylor estimators. As discussed in Section (ref), we can obtain these estimators by selecting appropriate instrumental variables in the second stage. For fixed effects estimation, model ((ref)) implies that $\dot{x}_{1ij}$ is a valid instrument since it varies only within groups and is uncorrelated with the group effects. This instrument automatically satisfies Assumption (ref). In this special case, the approach corresponds to the traditional MD estimator, where all variance originates in the first stage.

In the first stage, $\beta(\tau)$ is estimated separately for each group. The second stage then averages these group-level coefficients using weights proportional to $\tilde X_j'\tilde X_j$ (see equation (ref)). Galvao2015 propose an alternative approach, suggesting efficient weights proportional to the inverse of the variance of the first-stage estimators. Their weights are equivalent to ours when

equation[equation omitted — 148 chars of source]

for any $v$ and $v'$, i.e., when the conditional distribution of the group effects is the same across groups. Outside this specific case, our estimator may be less efficient but avoids the need to estimate the first-stage variance, which depends on the conditional densities in equation ((ref)). A third approach is to take the unweighted average of $\hat\beta_j(\tau)$, which, while not efficient when $\beta_j(\tau)$ are homogeneous, remains straightforward to interpret even if the model is misspecified.

We can implement a quantile between estimator using $\bar{x}_{1j}$ as an instrument to exploit only the variation across groups. On the other hand, combining within and between variations is more complex for quantile models than for least squares models. Applying quantile regression to the whole population without controlling for groups identifies parameters that differ from our intended parameters (see Remark (ref)). Using our MD estimator with $x_{1ij}$ as an instrument, which corresponds to using OLS in the second stage, consistently estimates $\beta(\tau)$ but only at the slow $\sqrt{m}$ rate because $x_{1ij}$ also varies between groups. The same occurs if both $\dot x_{1ij}$ and $\bar{x}_{1j}$ are combined with 2SLS, as the weights attributed to $\bar{x}_{1j}$ do not vanish asymptotically. Instead, we propose two efficient random effects estimators: an efficient GMM estimator and one using optimal instruments. It is worth highlighting that these estimators not only reduce the asymptotic variance of the estimator but also increase the rate of convergence from $\sqrt m$ to $\sqrt{mn}$ compared to 2SLS.

Given the first-stage estimation, we have the following moment condition:

equation[equation omitted — 156 chars of source]

When the instrument includes both the within-group variation $\dot x_{1ij}$ and the between-group average $\bar x_{1j}$, the efficient GMM estimator will optimally combine these two sources of variation. The weighting matrix is computed as shown in equation ((ref)). According to Propositions (ref) and Theorem (ref), this estimator is guaranteed to be both $\sqrt{mn}$ consistent and efficient. This random effects estimator has the same first-order asymptotic distribution as the fixed effects estimator that uses only the within-group variation $\dot x_{1ij}$. However, the random effects estimator is expected to have a lower variance in finite samples because it also incorporates the between-group variation. As the number of observations $n$ increases, the influence of the between-group variation diminishes, causing the random effects estimator to converge to the fixed effects estimator. This behavior is similar to what is observed in least squares models, as discussed by Baltagi2021 and Ahn2014.

If we impose the stronger assumption that the moment restriction in equation ((ref)) holds conditionally on $Z_j$, we can use the theory of optimal instruments to derive a more efficient random effects estimator. Optimal instruments are relevant when a researcher has a conditional moment restriction of the form $\mathbb{E}[g_{j}(\delta, \tau) | Z_j] = 0$. When a moment condition holds conditional on $Z_j$, an infinite set of valid moments exist, and one could use additional moments to increase efficiency. The goal is to select the instrument that minimizes the asymptotic variance, which takes the form $Z^*_j = \mathbb{E} [g_j(\delta, \tau)g_j(\delta, \tau)' | Z_j] ^{-1} R_j(\delta, \tau)$, with $R_j(\delta, \tau) = \mathbb{E} [ \frac{\partial}{\partial \delta} g_j(\delta, \tau) | Z_j] $ (see, e.g., Chamberlain1987 and Newey1993). To implement the random effect estimator with optimal instruments, we set $Z_j = X_j$. Under the additional assumption that $\mathbb{E}[\alpha_j^2(\tau) | X_j] = \sigma_\alpha^2(\tau)$,\footnote{We assume homoskedasticity of $\alpha_j(\tau)$ to obtain a simple estimator, similar to the classical least squares random effects estimator. If we were to drop this assumption, we would need to estimate $\mathbb{E}[\alpha_j(\tau)|X_j]$. Note that we do not assume homoscedasticity in the group-level model such that this assumption does not constrain the heterogeneity of $\beta(\tau)$ across different values of $\tau$.} the optimal instrument simplifies to

equation[equation omitted — 183 chars of source]

where $ V_j(\tau) $ is the asymptotic variance from the first stage for a group $j$, $\mathbf l_n$ is a $n$-dimenstional vector of ones, and $^+$ denotes the Moore-Penrose inverse.\footnote{Since the matrix $( \tilde X_j \frac{V_j(\tau)}{n}\tilde X_j' + \mathbf l_n'\mathbf l_n \sigma_\alpha^2(\tau) )$ is singular, we use the Moore-Penrose inverse.}

A few remarks about the optimal instruments follow. First, under standard random effects assumptions, the optimal instrument applied to least squares models is numerically identical to the FGLS estimator. Second, the optimal instrument depends on $n$ analogously to the efficient weighting matrix of the GMM estimator. As $n$ increases, the first stage variance converges to zero, and the generalized inverse will give infinitely more weights to the within variation and asymptotically converge to the fixed effects estimator. Third, if $\sigma_{\alpha}(\tau) = 0$, then all the variance arises in the first stage, and this estimator is identical to the efficient MD estimator (see Proposition (ref) in Appendix (ref)).

Hausman and Taylor Model

The Hausman-Taylor model provides a method for finding instrumental variables within the model itself. It represents a middle ground between the random effects model, which assumes orthogonality between $\alpha_j(\tau)$ and $x_{ij}$, and the fixed effects model, which only identifies the effect of individual-level variables. To estimate the effect of group-level variables, Hausman1981 assume that some elements of $x_{1ij}$ are uncorrelated with $\alpha_j(\tau)$. We consider model ((ref)) but partition the regressors into four types of variables, $x_{ij} = [x_{1ij}^{ex} \ x_{1ij}^{en} \ x_{2j}^{ex} \ x_{2j}^{en} ]$, where the superscript $ex$ indicates exogenous variables and the superscript $en$ indicates potentially endogenous variables. Thus,

align*[align* omitted — 116 chars of source]

The assumptions imply that we can estimate $\delta(\tau)$ using the instrument $z_{ij} = (\dot x_{1ij}^{ex}, \dot x_{1ij}^{en},$ $ \bar x_{1ij}^{ex}, x_{2j}^{ex})$. While $x_{2j}^{en}$ is potentially endogenous, the within variation is uncorrelated with $\alpha_j(\tau)$ as it varies only within $j$. Identification requires at least as many instruments as parameters to estimate, hence $dim(x_{1ij}^{ex}) \geq dim(x_{2j}^{en})$. In overidentified models, efficient GMM can be implemented, and if conditional moment restrictions are available, optimal instruments can be used. However, implementing optimal instruments is not straightforward, as it requires estimating $\mathbb{E} [x_{ij} |z_{ij}]$, typically done nonparametrically (see Newey1993). We do not contribute to this aspect in this paper.

Hausman Test

The random effects estimator's consistency relies on stronger orthogonality conditions compared to the fixed effects estimator. Under these stronger assumptions, both estimators are consistent, but the fixed effects estimator is inefficient. Hausman1978 proposed a test for the null hypothesis of random effects against the alternative of fixed effects. Various generalizations of the Hausman test have been suggested in the literature (e.g., Chamberlain1982, Mundlak1978, Wooldridge2019). Arellano1993 considers a heteroskedasticity and autocorrelation robust generalization based on a Wald test. Ahn1996 propose a GMM test based on a 3SLS regression as an equivalent method for the Hausman test.

This subsection explains how we can use the overidentification test presented in Section (ref) as a quantile version of the Hausman test for our two-step estimator. The assumption of correct specification of the first stage is maintained under both the null and alternative hypotheses. Compared to the fixed effects estimator, consistency of the random effects estimator additionally requires that $x_{1ij}$ is uncorrelated with $\alpha_j(\tau)$, so that $\mathbb{E}[\dot{x}_{1ij}' \alpha_j(\tau)] = 0$ and $\mathbb{E}[\bar{x}_{1j}' \alpha_j(\tau)] = 0$ are valid moment conditions. By contrast, the fixed effects estimator relies only on the moment condition $\mathbb{E}[\dot{x}_{1ij}' \alpha_j(\tau)] = 0$. Consequently, the overidentification test suggested in Proposition (ref) can be used as a test of the random effects orthogonality conditions. Compared to the traditional Hausman test, our test does not rely on the assumption of conditional homoskedasticity of the errors and is robust to clustering.

Galvao2019 also note that a quantile regression with all observations and fixed effects quantile regression identify different parameters. They propose a Hausman test based on an auxiliary quantile regression model that incorporates both $x_{1ij}$ and $\bar{x}_{1j}$ as regressors, and tests for the significance of $\bar{x}_{1j}$. However, this test differs from our approach in several ways: it starts from a different model (not conditional on the groups), relies on the correct specification of the auxiliary regression, and does not directly compare random effects and fixed effects estimators.

Simulations

table[table omitted — 3,188 chars of source]
table[table omitted — 1,969 chars of source]
table[table omitted — 2,061 chars of source]

This section presents simulation results for the different panel data estimators and the Hausman-type test presented in the previous subsections. These simulations focus on the estimation of $\beta(\tau)$. We consider the following data-generating process

align[align omitted — 73 chars of source]

where all variables are scalars, $\nu_{ij} \sim \mathcal{N}(0,1)$, and $x_{1ij} = h_j + 0.5u_{ij}$, with $u_{ij} \sim \mathcal{N}(0,1)$ and

equation*[equation* omitted — 238 chars of source]

If $\lambda \neq 0$, $x_{1ij}$ is correlated with $\alpha_j$. For the simulation of the panel data estimators, we let $\lambda = 0$ so that all estimators are consistent. In contrast, in the Monte Carlo study of the Hausman test, we set $\lambda = \{0,0.1, 0.2,0.3,0.4 \}$. The true coefficient takes the values $\beta(\tau) = 1 + 0.1 F^{-1}(\tau)$ where $F$ is the standard normal CDF. We consider the samples with $n = \{10,25, 200 \}$ and $m = \{ 25, 200\}$ and focus on the set of quantiles $\mathcal{T} = \{0.1, 0.5, 0.9\}$. All simulation results are based on 10,000 replications.

We compare the performance of four MD estimators, all based on the same first-stage regression but differing in their choice of instruments in the second stage. The `Pooled' estimator uses $x_{1ij}$ as the instrument,\footnote{We refer to this estimator as `Pooled' because the second stage involves a pooled OLS regression. It should not be confused with the one-step pooled quantile regression, which does not account for group structure.} the BE estimator uses $\bar x_{1j}$ as the instrument, the RE-GMM estimator combines $\dot x_{1ij}$ and $\bar x_{1j}$ using efficient GMM, and the RE-OI estimator combines the same exogenous variation using the single optimal instrument defined in equation ((ref)). We use the estimator of Powell1991 for $V_j(\tau)$ and the estimator of Nerlove1971 for $\sigma^2_\alpha(\tau)$.

Table (ref) shows that the estimators perform well also when both $m$ and $n$ are small. The RE-GMM estimator performs similarly to the RE-OI estimator except with small $n$ when the RE-GMM estimator outperforms the RE-OI both in terms of bias and variance. As expected, the RE-GMM, the RE-OI, and the fixed effects (FE) estimators become indistinguishable as $n$ increases. Whereas with small $n$, there is an apparent gain in using a random effects estimator. From the standard deviations, it is possible to see the different rates of convergence of the estimators. The precision of the fixed effects and random effects estimators increases in similar magnitude when $m$ or $n$ increases. In contrast, the standard deviation of the pooled and between estimators decreases only when $m$ increases. The pooled and the between estimators have the smallest bias but, in most cases, also the largest variance.

The coverage probabilities of the 95% confidence intervals are provided in Table (ref). The confidence intervals of the pooled and the fixed effects estimator perform well in all sample sizes considered. On the other hand, the confidence bands of the random effects estimators slightly undercover the true parameter mostly when $n$ is small. In larger samples, all the coverage probabilities are close to the theoretical level.

Table (ref) shows the rejection probabilities of the overidentification test for different values of $\lambda$. When $\lambda = 0$, the $\mathbb{H}_0$ is satisfied, so we should reject the null at a rate close to $5\%$. If $\lambda \neq 0$, $ x_{ij}$ is correlated with $\alpha_j$, and some moment conditions used by the RE-GMM estimator are not valid. In this case, higher rejection probabilities suggest a more powerful test. The first column shows that the empirical sizes of the test are close to the theoretical levels in most sample sizes. Columns 2-5 show that, as expected, the power of the test is higher in large samples and increases with the correlation between $\bar x_{1j}$ and the unobserved heterogeneity $\alpha_j$. An increase in $m$ substantially improves the power of the test, while a larger number of time periods $n$ improves the results to a lesser extent. In general, the test performs better both in terms of size and power when $m$ is large, which is most often the case in empirical applications. Even if the random effects estimator converges to the fixed effects estimator as $n$ increases and the random effects estimator of $\beta$ will be consistent even if $\lambda \neq 0$, the size and power of the test do not deteriorate. This result is consistent with the findings in Ahn2014.

Empirical Application: The Effect of the Food Stamps Program on Birth Weight

In this section, we apply our minimum distance approach to estimate the impact of the food stamp program on the birth weight distribution using grouped data. We complement the analysis of Almond2011 by providing distributional effects. Food stamps constitute an important means-tested program that gives entitled households coupons they can redeem at approved retail food stores. The Food Stamp Act (FSA) was introduced in 1964 and enabled counties to start their own federally funded food stamp program (FSP). In the subsequent years, counties increasingly adopted such programs, and in 1973, an amendment to the FSA required all counties to establish a FSP by 1975. Thus, the share of counties with an FSP increased steadily from 1964 to 1974, and identification exploits the variation in the timing of the adoption across counties. Almond2011 use data from 1968 (when about 40% of the counties had introduced the program) to 1977 (two years after the FSP was implemented everywhere) to analyze the effect of the program.

Given the negative consequences of low birth weight, besides estimating the effect of the policy on average weight, Almond2011 estimate the effect on the probability that birth weight falls below a certain threshold. As discussed in Melly2015a, this procedure leads to biased results unless there is no time effect or group effect or the outcome is uniformly distributed.

In this section, we use the subscripts $i$, $c$, and $t$ to denote the birth, the county, and the trimester of birth, respectively.\footnote{Using the same notation as in the paper, the $j$ units are county-trimester combinations, and the $i$ units index individual births within a county in a given trimester. In this section, we use three subscripts for clarity.} The variable of interest is a binary variable that is coded 1 if there was a food stamp program in place three months before birth. Hence, the treatment is assigned to county-month cells, and in around 1% of cases, it also varies within groups.

figure[figure omitted — 18,127 chars of source]

We consider the following model separately for black and white mothers:

equation[equation omitted — 172 chars of source]

where $Q(\tau, bw_{ict} | fsp_{ct}, x_{1ict}, x_{2ct}, v_{ct}) $ is the $\tau$th conditional quantile function of the outcome given all the variables. The variable $fsp_{ct}$ indicates whether there is a food stamp program in place, $x_{1ict}$ are variables related to the individual births, such as gender, mother age, and its square as well as the legitimacy status of the birth. Group-level control variables $x_{2ct}$ include annual county-level controls (real per capita income, government transfers to individuals, medical spending, and retirement and disability payments) and 1960 county-level characteristics (county population and the shares of urban population, black population, and of farmland) interacted with a linear time trend. Further, $x_{2ct}$ also includes county, state-year, and time fixed effects.

For the estimation, we drop groups that have less than 25 degrees of freedom.\footnote{If there are $K_1$ individual level variables we drop all groups with less than $K_1 + 1 + 25$ observations. Since some variables might vary at the individual level in some groups only, this threshold is group-specific.} The estimations are performed using a sample of 2,822,091 individual observations divided into 19,482 groups for blacks and 16,038,235 individual births divided into 80,289 groups for whites.\footnote{We have a different number of groups compared to Almond2011 due to multiple reasons. First, they give higher weights to births in groups where only 50% of the births are coded in the natality data; thus, when they drop groups with less than 25 births, the number of births in these groups is inflated. Second, since they take the group average, they keep births with missing values for birth weight. We drop those births as we work with individual-level data.} Figure (ref) illustrates the results. The results for black mothers are in the left panel, while the results for white mothers are in the right panel. The effect is substantially larger among black mothers. The results suggest a positive effect of the food stamp program on the lower tail of the conditional distribution. For instance, the estimates suggest that the food stamp program is associated with an increase in birth weight by almost 30 grams for blacks at the 5th percentile of the conditional distribution. For whites, there seems to be an effect only in the left tail of the distribution, and the effects are small. For blacks, the coefficients are large in the left tail and remain positive, albeit of small magnitude, until the 75% percentile. However, for higher quantiles, the effects are not statistically different from zero.

To compare our MD estimator with the CLP estimator on observational data, we also perform the estimation using the CLP procedure. For this comparison, we focus on the sample of black mothers. Before implementing the CLP estimator on this data, we need to solve a few issues that affect the CLP estimator but not the MD estimator. First, the treatment variable exhibits within-group variation in around 1% of the groups. If there are variables that vary only within some groups, the intercept is not comparable over groups, thus creating a problem for the CLP estimator. Hence, for the comparison, we drop groups for which the variable $fsp$ is varying within the group. A second but similar issue is due to the covariate controlling for the legitimacy status of the birth, as there are groups that do not contain all levels of the categorical variable. To solve this issue, we drop the variable from the conditioning set. Third, for one observation (out of over 2 million), one of the group-level control variables is minimally different from the value assigned to other members of the group. This is most likely due to a coding error. In this case, we replace the values with the group average. While these features of the data do not pose any problem with our estimation approach, they can lead to substantially different results using the CLP estimation method.

figure[figure omitted — 15,275 chars of source]

Figure (ref) shows the estimation results performed with our approach compared to the CLP approach using the sample of black mothers. Given that we estimate a slightly different model on a marginally different sample, we re-estimate the effects using the MD estimator to ensure a meaningful comparison. The results obtained using the MD estimator remain almost identical to those in Figure (ref). On the other hand, using the CLP estimator, we obtain substantially different results from those attained using the MD estimator, and the confidence intervals are, on average, 14 times larger, precluding any possibility of drawing informative conclusions. Further, the figure shows that the MD point estimates and confidence intervals are always inside the confidence intervals of the CLP estimator.

Conclusion

To summarize, our paper makes the following key contributions: First, we propose a new class of minimum distance estimators for quantile panel data models, applicable to both classical panel data settings, where units are observed over time, and grouped data settings, where individuals are divided into groups and treatment varies at the group level. This class of estimators provides quantile analogs of established panel data methods, including fixed effects, random effects, between, and Hausman-Taylor estimators. Additionally, it substantially improves upon the existing grouped instrumental variable estimator of Chetverikov2016 and remains computationally fast. Second, we establish uniform asymptotic properties under a general framework accommodating arbitrary degrees of group-level heterogeneity. We also introduce adaptive inference procedures that remain valid regardless of whether group effects are zero, bounded, or vanish at arbitrary rates. The adaptive estimator of the variance of the moments leads to a GMM estimator that is uniformly efficient across unknown relative convergence rates. Third, we demonstrate the practical relevance of our approach in an empirical application studying the impact of the Food Stamp Program on birth weight. Our results indicate that the policy’s positive effects are concentrated almost entirely in the lower tail of the birth weight distribution, particularly among black mothers.

While these advancements mark significant progress, our approach has limitations that invite further exploration. First, although our Monte Carlo simulations show that our estimator and suggested standard errors perform well in finite samples, our asymptotic framework relies on both the number of groups and the number of observations per group diverging to infinity. This assumption is reasonable in settings with large groups, such as our empirical application, but it is often violated in classical panel data contexts. Future work should explore finite $n$ inference or consider stronger assumptions for identification in such cases. Second, our estimator accommodates parametric time trends but not unrestricted time effects. Estimating time effects requires working with all units simultaneously, making a group-by-group computation strategy infeasible. Finally, while we focus on within-group heterogeneity, many applications also require understanding between-group heterogeneity. Pons2024 develops a framework that explicitly captures both, providing a promising avenue for further research.

appendix

Proofs

Proof of Lemma (ref): Sampling error

proof[Proof of lemma (ref)] Starting from the definition of the estimator we obtain \begin{align*} \hat \delta (\tau) &= \left( X' Z \hat W(\tau) Z' X \right) ^{-1} X' Z \hat W(\tau) Z' \hat Y(\tau)\\ &=\left( S_{ZX}' \hat W(\tau) S_{ZX} \right) ^{-1} S_{ZX}' \hat W(\tau) \frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\tilde x_{ij}' \hat \beta_j(\tau)\\ &=\left( S_{ZX}' \hat W(\tau) S_{ZX} \right) ^{-1} S_{ZX}' \hat W(\tau) \frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\left(\tilde x_{ij}'\left(\hat \beta_j(\tau)-\beta_j(\tau)\right)+\tilde x_{ij}'\beta_j(\tau)\right)\\ &=\left( S_{ZX}' \hat W(\tau) S_{ZX} \right) ^{-1} S_{ZX}' \hat W(\tau) \frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\left( \tilde x_{ij}'\left(\hat \beta_j(\tau)-\beta_j(\tau)\right)+ x_{ij}'\delta(\tau)+\alpha_j(\tau)\right)\\ &=\delta(\tau)+\left( S_{ZX}' \hat W(\tau) S_{ZX} \right) ^{-1} S_{ZX}' \hat W(\tau) \frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\left( \tilde x_{ij}'\left(\hat \beta_j(\tau)-\beta_j(\tau)\right)+\alpha_j(\tau)\right). \end{align*}

Proof of Theorem (ref): Uniform consistency

Lemma (ref)

As a preliminary step to prove the uniform consistency of our estimator, we show uniform (in $\tau$ and $j$) consistency of the group-level quantile regressions.

lemma[Uniform consistency of $\hat\beta_j(\tau)$] Under Assumptions (ref)-(ref) and (ref)(a), we have $$\underset{\tau\in\mathcal T} {\sup}\,\underset{1\leq j\leq m}{\max} \lVert \hat \beta_j(\tau)-\beta_j (\tau) \rVert = o_p(1).$$
proof[Proof of Lemma (ref)] Angrist2006a show uniform consistency of the quantile regression estimator in $\tau$ but not in $j$ (see their Theorem 3) while Galvao2015 show uniform consistency in $j$ but not in $\tau$ (see their Lemma 1). We show uniformity in both dimensions by following the steps of the proof in Galvao2015 and extending it. We define $\mathbb Q_{nj}(\beta, \tau)=\frac{1}{n}\sum_{i = 1}^n \rho_\tau(y_{ij}-\tilde x_{ij}'\beta) - \rho_\tau(y_{ij}-\tilde x_{ij}'\beta_j(\tau))$ and $Q_{j}(\beta,\tau)=\mathbb{E}_{i|j}[ \rho_\tau(y_{ij}-\tilde x_{ij}'\beta) - \rho_\tau(y_{ij}-\tilde x_{ij}'\beta_j(\tau))]$. As shown in Angrist2006a, the empirical processes $(\beta, \tau)\mapsto \mathbb Q_{nj}(\beta,\tau)$ for all groups $j$ are stochastically equicontinuous because $\left|\mathbb Q_{nj}(\beta',\tau')-\mathbb Q_{nj}(\beta'',\tau'') \right|\leq C_{1}\cdot |\tau'-\tau''|+C_{2}\cdot \lVert\beta'-\beta''\rVert$ where $C_{1}=2\cdot C\cdot \underset{\beta \in \mathcal{B}}{\sup}\lVert \beta \rVert$ for any compact set $\mathcal{B}$ and $C_{2}=2\cdot C$. The constant $C$ is defined in Assumption (ref). Note that $C_{1}$ and $C_{2}$ are neither functions of $j$ nor $\tau$. Fix any $\zeta>0$. Let $B_j(\zeta,\tau)=\{\beta:\lVert\beta-\beta_j(\tau)\rVert \leq \zeta\}$, the ball with center $\beta_j(\tau)$ and radius $\zeta$. For each $\beta\notin B_j(\zeta,\tau)$, define $\tilde \beta=r_j\beta+ (1-r_j)\beta_j(\tau)$ where $r_j=\frac{\zeta}{\lVert \beta-\beta_j(\tau)\rVert}$. So $\tilde \beta\in \partial B_j(\zeta,\tau)=\{\beta:\lVert\beta-\beta_j(\tau)\rVert=\zeta\}$, the boundary of $B_j(\zeta,\tau)$. Since $\mathbb Q_{nj}(\beta,\tau)$ is convex in $\beta$ for all $\tau$, and $\mathbb Q_{nj}(\beta_j(\tau),\tau)=0$, we have \begin{align} r_j \mathbb Q _{nj}(\beta,\tau) \geq\mathbb Q_{nj}(\tilde \beta,\tau)=Q_{j}(\tilde\beta,\tau)+\mathbb Q_{nj}(\tilde\beta,\tau)-Q_{j}(\tilde\beta,\tau)>\epsilon_\zeta+\mathbb Q_{nj}(\tilde\beta,\tau)-Q_j(\tilde\beta,\tau) \end{align} uniformly in $j$ and $\tau$, where \begin{equation*} \epsilon_\zeta=\underset{\tau\in\mathcal T}{\inf}\;\underset{1\leq j\leq m}{\inf}\;\underset{\lVert\beta-\beta_j(\tau)\rVert=\zeta}{\inf} \mathbb{E}_{i|j}\left[\int_0^{\tilde x_{ij}'(\beta-\beta_j(\tau))}\left(1(y_{ij}-\tilde x_{ij}'\beta_j(\tau)\leq s)-1(y_{ij}-\tilde x_{ij}'\beta_j(\tau)\leq 0)\right)ds\right] \end{equation*} by the identity of Knight1998 and $\epsilon_\zeta>0$ by Assumptions (ref) and (ref). Thus, we have that \begin{align*} \left\{\underset{\tau \in \mathcal{T}}{\sup}\,\underset{1\leq j\leq m}{\max} \lVert\hat\beta_j(\tau)-\beta_j(\tau)\rVert>\zeta\right\} &\overset{(a)}{\subseteq}\{\exists \tau_j\in\mathcal T,\exists\beta_j\notin B_j(\zeta,\tau_j) :\mathbb Q_{nj}(\beta_j,\tau_j)\leq 0\}\\ &\overset{(b)}{\subseteq} \cup_{j=1}^m \left\{\underset{\tau\in\mathcal T}{\sup}\,\underset{\beta_j\in B_j(\zeta,\tau_j)}{\sup} |\mathbb Q_{nj}(\beta_j,\tau_j)-Q_j(\beta_j,\tau_j)| \geq\epsilon_\zeta\right\}. \end{align*} Relation (a) holds because, by definition, $\hat\beta_j(\tau)$ minimizes $\mathbb Q_{nj}(\beta,\tau)$, and $\mathbb Q_{nj}(\beta_j(\tau),\tau)=0$. Relation (b) holds by the rightmost inequality of line ((ref)). Then, it follows that \begin{align*} {\mathrm{P}}\left\{\underset{\tau\in\mathcal T}{\sup}\, \underset{1\leq j\leq m}{\max} \lVert \hat \beta _j(\tau) -\beta_j(\tau)\rVert > \zeta\right\} &\leq {\mathrm{P}} \left\{\cup_{j=1}^m\left\{\underset{\tau\in\mathcal T}{\sup}\, \underset{\beta_j\in B_j(\zeta,\tau)}{\sup} |\mathbb Q_{nj}(\beta_j,\tau)-Q_j(\beta_j,\tau)| \geq \epsilon_\zeta \right\} \right\}\\ &\leq \sum_{j = 1}^m {\mathrm{P}}\left\{ \underset{\tau \in \mathcal T}{\sup} \,\underset{\beta_j\in B_j(\zeta,\tau)}{\sup} |\mathbb Q_{nj}(\beta_j,\tau)-Q_j(\beta_j,\tau)| \geq\epsilon_\zeta\right\}\\ &\leq m \underset{1\leq j\leq m}{\max} {\mathrm{P}}\left\{ \underset{\tau\in\mathcal T}{\sup}\,\underset{\beta_j\in B_j(\zeta,\tau)}{\sup} |\mathbb Q_{nj}(\beta_j,\tau)-Q_j(\beta_j,\tau) |\geq \epsilon_\zeta \right\} \end{align*} Therefore, if we can show that \begin{equation*} \underset{1\leq j\leq m}{\max} {\mathrm{P}}\left\{ \underset{\tau\in\mathcal T}{\sup}\,\underset{\beta_j\in B_j(\zeta,\tau)}{\sup} |\mathbb Q_{nj}(\beta_j,\tau)-Q_j(\beta_j,\tau) |\geq \epsilon_\zeta \right\}= o\left(\frac{1}{m}\right)\end{equation*} the proof of the lemma will be completed. Without loss of generality, we assume $\beta_j(\tau)=0$ for all $j$ and $\tau\in\mathcal T$. Then the balls $B_j(\zeta,\tau)$ for all $j$ and $\tau\in\mathcal T$ are identical and we denote them by $B(\delta)$. Because the closed ball $B(\zeta)$ is compact, there exist $K$ balls with center $\beta^k$, $k=1,...,K$, and radius $\frac{\epsilon_\zeta}{3C_2}$ such that the collection of them covers $B(\zeta)$. Since $\epsilon_\zeta>0$, we can find a finite $K$ that satisfies this condition and is independent of $j$ and $\tau$. Therefore, for any $\beta\in B(\delta)$, there is some $k\in \{1,...,K\}$ such that \begin{align*} |\mathbb Q_{nj}(\beta,\tau)-Q_j(\beta,\tau)|-|\mathbb Q_{nj}(\beta^k,\tau)-Q_j(\beta^k,\tau)|&\leq |\mathbb Q_{nj}(\beta,\tau)-Q_j(\beta,\tau)-\mathbb Q_{nj}(\beta^k,\tau)+Q_j(\beta^k,\tau)|\\ &\leq|\mathbb Q_{nj}(\beta,\tau)-\mathbb Q_{nj}(\beta^k,\tau)|+|Q_j(\beta,\tau)-Q_j(\beta^k,\tau)|\\ &\leq C_2 \frac{\epsilon_\zeta}{3C_2}+C_2 \frac{\epsilon_\zeta}{3C_2}=\frac{2\epsilon_\zeta}{3} \end{align*} uniformly in $j$ and $\tau\in\mathcal T$. Note that the third line is justified by the stochastic equicontinuity of $\mathbb Q_{nj}(\beta,\tau)$. It then follows that, $$\underset{\tau\in\mathcal T}{\sup}\, \underset{\beta\in B(\zeta)}{\sup} |\mathbb Q_{nj}(\beta,\tau)-Q_j(\beta,\tau)|\leq \underset{\tau\in\mathcal T}{\sup}\,\underset{1\leq k\leq K}{\max}|\mathbb Q_{nj}(\beta^k,\tau)-Q_j(\beta^k,\tau)| +\frac{2\epsilon_\zeta}{3},$$ and \begin{align*} {\mathrm{P}}\left\{\underset{\tau\in\mathcal T}{\sup}\,\underset{\beta\in B(\delta)}{\sup} |\mathbb Q_{nj}(\beta,\tau)-Q_j(\beta,\tau)>\epsilon_\zeta\right\}&\leq{\mathrm{P}}\left\{\underset{\tau\in\mathcal T}{\sup} \,\underset{1 \leq k \leq K} {\max} |\mathbb Q_{nj}(\beta^k,\tau)-Q_j(\beta^k,\tau)|+\frac{2\epsilon_\zeta}{3} > \epsilon_\zeta \right\}\\ &={\mathrm{P}}\left\{\underset{\tau\in\mathcal T}{\sup} \,\underset{1 \leq k \leq K} {\max} |\mathbb Q_{nj}(\beta^k,\tau)-Q_j(\beta^k,\tau)|>\frac{\epsilon_\zeta}{3}\right\}\\ &\leq \underset{\tau\in\mathcal T}{\sup} \sum_{k=1}^K {\mathrm{P}} \left\{|\mathbb Q_{nj}(\beta^k,\tau)-Q_j(\beta^k,\tau)|>\frac{\epsilon_\zeta}{3} \right\}. \end{align*} For each $\tau\in\mathcal T$, $\mathbb Q_{nj}(\beta^k,\tau)$ is the sample mean of $n$ i.i.d. terms bounded in absolute values by $2\cdot C\cdot \zeta$. By Hoeffding’s inequality, it follows that \begin{align*}\sum_{k=1}^K {\mathrm{P}} \left\{|\mathbb Q_{nj}(\beta^k,\tau)-Q_j(\beta^k,\tau)|>\frac{\epsilon_\zeta}{3}\right\} &\leq 2K\exp\left\{-\frac{2n\epsilon_\zeta^2}{3^22^2C^2\zeta^2} \right\}\\ & =2K\exp\left\{-\frac{n\epsilon_\zeta^2}{18C^2\zeta^2} \right\}\\ &=O(\exp(-n)).\end{align*} Since $\frac{\log m}{n}\rightarrow0$ by Assumption (ref)(a), it follows that $O_p(\exp(-n))=o_p(1/m)$. Note that the bound $\epsilon_\zeta/3>0$ is uniform in $\tau$, and $K$ is finite. By stochastic equicontinuity of $\mathbb Q_{nj}(\beta,\tau)$, as $n\rightarrow\infty$, uniformly in $\tau\in\mathcal T$, \begin{equation*} \mathbb Q_{nj}(\beta^k,\tau)=Q_j(\beta^k,\tau)+o_p(1). \end{equation*} It follows that \begin{equation*} \underset{\tau\in\mathcal T}{\sup} \sum_{k=1}^K {\mathrm{P}} \left\{|\mathbb Q_{nj}(\beta^k,\tau)-Q_j(\beta^k,\tau)|>\frac{\epsilon_\zeta}{3} \right\}=o_p\left(\frac{1}{m}\right). \end{equation*}

Lemma (ref)

lemma[Uniform consistency of $\hat G(\tau)$ with a full-rank weighting matrix] Assumptions (ref), (ref), (ref), and (ref) hold. Then, \begin{equation} \underset{\tau\in\mathcal{T}}{\sup}\left\lVert \left(S_{ZX}'\hat W(\tau) S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau)-\left(\Sigma_{ZX}'W(\tau) \Sigma_{ZX}'\right)^{-1}\Sigma_{ZX}'W(\tau)\right\rVert=o_p(1) \end{equation} and $\left(\Sigma_{ZX}'W(\tau) \Sigma_{ZX}'\right)^{-1}\Sigma_{ZX}'W(\tau)$ is uniformly bounded and continuous.
proof[Proof of Lemma (ref)] First, it follows from Assumptions (ref)(ii), (ref)(i) and (ref)(i) that $\operatorname{Var}\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}x_{ij}'\right)=o_p\left(\frac{1}{n}\right)$ and $\mathbb{E}\left[\frac{1}{n}\sum_{i = 1}^n z_{ij}x_{ij}'\right]=\mathbb{E}[z_{ij}x_{ij}']$. Hence, by Assumption (ref)(i), $\operatorname{Var}\left(\frac{1}{m}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n x_{ij}z_{ij}'\right)=o_p\left(\frac{1}{mn}\right)$. By Chebyshev’s inequality, $$\frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}x_{ij}'-\mathbb{E}[z_{ij}x_{ij}']\right)\underset{p}{\rightarrow}0.$$ In addition, by Assumption (ref)(iii), $m^{-1}\sum_{j = 1}^m\mathbb{E}[z_{ij}x_{ij}']\rightarrow\Sigma_{ZX}$. It follows that \begin{equation*} S_{ZX}\underset{p}{\rightarrow}\Sigma_{ZX}. \end{equation*} By assumption (ref), uniformly in $\tau \in \mathcal{T}$, $\hat W(\tau)\underset{p}{\rightarrow}W(\tau)$ where $W(\tau)$ is uniformly continuous and strictly positive definite. Since $\Sigma_{ZX}$ is bounded and has full column rank, it implies that $\Sigma_{ZX}'W(\tau)\Sigma_{ZX}$ is also uniformly continuous and invertible. The result of the lemma follows.

Theorem (ref)

proof[Proof of Theorem (ref)] By Lemma (ref), \begin{equation*} \hat \delta (\tau) - \delta(\tau) = \left( S_{ZX}' \hat W(\tau) S_{ZX} \right) ^{-1} S_{ZX}' \hat W(\tau) \frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\left( \tilde x_{ij}'\left(\hat \beta_j(\tau)-\beta_j(\tau)\right)+\alpha_j(\tau)\right). \end{equation*} By Lemma (ref), the first factor converges uniformly to $G(\tau)$: \begin{equation} \left(S_{ZX}'\hat W(\tau) S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau) \underset{p}{\rightarrow}\left(\Sigma_{ZX}'W(\tau) \Sigma_{ZX}'\right)^{-1}\Sigma_{ZX}'W(\tau)=G(\tau). \end{equation} By lemma (ref), $\hat\beta_j(\tau)$ is consistent for $\beta_j(\tau)$ uniformly in $j$ and $\tau$. Together with the boundedness of $x_{ij}$ in Assumption (ref)(i) and of $z_{ij}$ in Assumption (ref)(i), it follows that \begin{equation} \underset{\tau\in\mathcal T}{\sup} \ \frac{1}{mn} \sum_{j = 1}^m\sum_{i = 1}^n z_{ij} \tilde x_{ij}'(\hat \beta_j(\tau)-\beta_j(\tau))\underset{p}{\rightarrow}0 . \end{equation} By Assumption (ref)(ii), $\mathbb{E}[z_{ij}\alpha_j(\tau)]=0$ uniformly in $\tau$. By Assumption (ref), $\operatorname{Var}(z_{ij}\alpha_j(\tau))$ is uniformly bounded. In addition, $z_{ij}$ is bounded and $\alpha_j(\tau)$ is uniformly continuous in $\tau$. Hence, \begin{equation} \underset{\tau\in\mathcal T}{\sup} \ \frac{1}{mn} \sum_{j = 1}^m\sum_{i = 1}^n z_{ij} \alpha_j(\tau)\underset{p}{\rightarrow}0 . \end{equation} The result of the Theorem follows from equations ((ref)), ((ref)), and ((ref)).
commentWhen the rates of convergence of the moments are heterogeneous, i.e., when $\operatorname{Var}(\alpha(\tau))> \epsilon > 0$, $L_1>0$ instruments satisfy $\operatorname{Var}(\bar z_{1j})=0$ and $L_2>0$ instruments satisfy $\operatorname{Var}(\bar z_{2j})>0$, the convergence rates of the different elements of the efficient weighting matrix are also heterogeneous. In such a case, depending on the scaling of $W(\tau)$ (the estimator is invariant to the scaling of $W(\tau)$), either some elements converge to $0$ or other elements diverge to infinity such that Theorem (ref) does not apply. Nonetheless, Theorem (ref) shows that this does not preclude the consistency of the estimator. \begin{lemmap}{(ref)$'$}[Uniform consistency of $\hat G(\tau)$ when the weighting matrix is asymptotically singular] Assumptions (ref), (ref), (ref), and (ref) with $a_n(\tau) = o_p(1)$ hold. Then, \begin{equation*} \underset{\tau\in\mathcal{T}}{\sup}\left\lVert \left(S_{ZX}'\hat W(\tau) S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau)-\begin{pmatrix}G_{11}(\tau) & 0 \\ G_{21}(\tau) & G_{22}(\tau)\end{pmatrix} \right\rVert=o_p(1) \end{equation*} where $G_{11}(\tau)=\left(\Sigma_{11}'W_{11}(\tau)\Sigma_{11}\right)^{-1}\Sigma_{11}'W_{11}(\tau)$ is $M_1 \times L_1$, $G_{22}(\tau)=\left(\Sigma_{22}'W_{22}(\tau)\Sigma_{22}\right)^{-1}\Sigma_{22}'W_{22}(\tau)$ is $M_2 \times L_2$. $G_{11}(\tau)$, $G_{22}(\tau)$, and $G_{12}(\tau)$ are uniformly bounded and continuous. \end{lemmap} \begin{proof}[Proof of Lemma (ref)] The proof of this lemma is similar to the proof of Lemma (ref). In addition, we can exploit the structure of the weighting matrix to show that the $M_1\times L_2$ upper right submatrix of $G(\tau)$ converges to $0$ uniformly in $\tau$. To simplify the notation, in the following, we suppress the dependency of $W$ on $\tau$. Note that \begin{align*} \Sigma_{ZX}'W&=\begin{pmatrix}\Sigma_{11}' & \Sigma_{21}' \\ 0 & \Sigma_{22}' \end{pmatrix}\begin{pmatrix}W_{11} & a_n W_{12} \\ a_n W_{21} & a_n W_{22} \end{pmatrix}\\ &=\begin{pmatrix}\Sigma_{11}'W_{11}+a_n \Sigma_{21}'W_{21} & a_n\left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) \\ a_n \Sigma_{22}' W_{21} & a_n \Sigma_{22}' W_{22} \end{pmatrix}. \end{align*} Furthermore, \begin{align*} \Sigma_{ZX}'W\Sigma_{ZX}&= \begin{pmatrix}\Sigma_{11}' & \Sigma_{21}' \\ 0 & \Sigma_{22}' \end{pmatrix}\begin{pmatrix}W_{11} & a_n W_{12} \\ a_n W_{21} & a_n W_{22} \end{pmatrix}\begin{pmatrix}\Sigma_{11} & 0 \\ \Sigma_{21} & \Sigma_{22} \end{pmatrix}\\ &=\begin{pmatrix}A_{11}+a_n B_{11} & a_n B_{12} \\ a_n B_{21} & a_n B_{22} \end{pmatrix}, \end{align*} where $A_{11}=\Sigma_{11}'W_{11}\Sigma_{11}$, $B_{11}=\Sigma_{11}'W_{12}\Sigma_{21} + \Sigma_{21}' W_{21} \Sigma_{11} + \Sigma_{21}'W_{22}\Sigma_{21}$, $B_{12} = \Sigma_{11}'W_{12} \Sigma_{22} + \Sigma_{21}' W_{22} \Sigma_{22}$, $B_{21}= \Sigma_{22}W_{21} \Sigma_{11} +\Sigma_{22}W_{22}\Sigma_{21}$, and $B_{22}=\Sigma_{22}'W_{22}\Sigma_{22}$. By Assumption (ref), $\Sigma_{ZX}$ is full column rank. Since this matrix is block lower diagonal (see equation ((ref))), it implies that $\Sigma_{11}$ and $\Sigma_{22}$ are also full column rank. By Assumption (ref), $W_{11}$ and $W_{22}$ have full rank. It follows that $A_{11}$ and $B_{22}$ are invertible. By the inverse of a partitioned matrix, we obtain \begin{equation*} \left(\Sigma_{ZX}'W\Sigma_{ZX}\right)^{-1}=\\ \begin{pmatrix}C+a_n C B_{12}D B_{21}C & -C B_{12}D \\ -D B_{21}C & a_n^{-1}D \end{pmatrix} \end{equation*} where $C=\left(A_{11}+a_n B_{11}\right)^{-1}$ and $D=\left(B_{22}-a_n B_{21}CB_{12}\right)^{-1}$. Combining these results and taking into account that $a_n=o_p(1)$, we obtain \begin{align*} \left(\Sigma_{ZX}'W\Sigma_{ZX}\right)^{-1}\Sigma_{ZX}'W =\begin{pmatrix}A_{11}^{-1}\Sigma_{11}'W_{11} & 0 \\ -B_{22}^{-1}B_{21}A_{11}^{-1}\Sigma_{11}'W_{11} + B_{22}^{-1}\Sigma_{22}'W_{21} & B_{22}^{-1}\Sigma_{22}'W_{22} \end{pmatrix}+o_p(1) \end{align*} \begin{commentP} Note that by the Moore-Penrose Inverse the lower left element simplifies to \begin{align*} - B_{22}^{-1} \left(\Sigma_{22}W_{21} \Sigma_{11} +\Sigma_{22}W_{22}\Sigma_{21} \right)\left( \Sigma_{11}'W_{11}\Sigma_{11} \right)^{-1} \Sigma_{11}'W_{11} + B_{22}^{-1}\Sigma_{22} W_{21} \\ = - B_{22}^{-1} \Sigma_{22}W_{21} \Sigma_{11} \left( \Sigma_{11}'W_{11}\Sigma_{11} \right)^{-1} \Sigma_{11}'W_{11}^{1/2} W_{11}^{1/2}+B_{22}^{-1} \Sigma_{22}W_{22}\Sigma_{21} \left( \Sigma_{11}'W_{11}\Sigma_{11} \right)^{-1} \Sigma_{11}'W_{11}^{1/2} W_{11}^{1/2} \\ + B_{22}^{-1}\Sigma_{22} W_{21} \\ = - B_{22}^{-1} \Sigma_{22}W_{21} \Sigma_{11} \left( W_{11}^{1/2}\Sigma_{11} \right )^{+}W_{11}^{1/2}+B_{22}^{-1} \Sigma_{22}W_{22}\Sigma_{21} \left ( W_{11}^{1/2}\Sigma_{11} \right )^{+} W_{11}^{1/2} \\ + B_{22}^{-1}\Sigma_{22} W_{21} \\ = - B_{22}^{-1} \Sigma_{22}W_{21} -B_{22}^{-1} \Sigma_{22}W_{22}\Sigma_{21} \Sigma_{11}'^{+} + B_{22}^{-1}\Sigma_{22} W_{21} \\ = -B_{22}^{-1} \Sigma_{22}W_{22}\Sigma_{21} \Sigma_{11}'^{+} \end{align*} \end{commentP} All the terms in this matrix are finite. Uniform convergence follows from the boundedness of $x_{ij}$ and $z_{ij}$, the invertibility of $A_{11}$ and $B_{22}$ and the uniform continuity of $W(\tau)$ as a function of $\tau$. \end{proof} \color{gray} NEW VERSION (Don't think it is a good idea) \begin{lemmap}{(ref)$'$}[Uniform consistency of $\hat G(\tau)$ when the weighting matrix is asymptotically singular] Assumptions (ref), (ref), (ref), and (ref) hold. Then, \begin{equation*} \underset{\tau\in\mathcal{T}}{\sup}\left\lVert \left(S_{ZX}'\hat W(\tau) S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau)-\begin{pmatrix}G_{11}(\tau) & 0 \\ G_{21}(\tau) & G_{22}(\tau)\end{pmatrix} \right\rVert=O_p \left (\frac{a_n}{\sqrt{m}} \right) = o_p(1) \end{equation*} where $G_{11}(\tau)=\left(\Sigma_{11}'W_{11}(\tau)\Sigma_{11}\right)^{-1}\Sigma_{11}'W_{11}(\tau)$ and $G_{22}(\tau)=\left(\Sigma_{22}'W_{22}(\tau)\Sigma_{22}\right)^{-1}\Sigma_{22}'W_{22}(\tau)$. $G_{11}(\tau)$, $G_{22}(\tau)$, and $G_{12}(\tau)$ are uniformly bounded and continuous. \end{lemmap} \begin{proof}[Proof of Lemma (ref)] The proof of this lemma is similar to the proof of Lemma (ref). In addition, we can exploit the structure of the weighting matrix to show that the $M_1\times L_2$ upper right submatrix of $G(\tau)$ converges to $0$ uniformly in $\tau$. To simplify the notation, in the following, we suppress the dependency of $W$ on $\tau$. Note that \begin{align*} S_{ZX}'\hat W&=\begin{pmatrix}S_{11}' & S_{21}' \\ 0 & S_{22}' \end{pmatrix}\begin{pmatrix} \hat W_{11} & a_n \hat W_{12} \\ a_n \hat W_{21} & a_n \hat W_{22} \end{pmatrix}\\ &=\begin{pmatrix}S_{11}'W_{11}+a_n S_{21}'\hat W_{21} & a_n\left(S_{11}'\hat W_{12}+S_{21}'\hat W_{22}\right) \\ a_n S_{22}' \hat W_{21} & a_n S_{22}' \hat W_{22} \end{pmatrix} \end{align*} Furthermore, \begin{align*} S_{ZX}'\hat WS_{ZX}&= \begin{pmatrix} S_{11}' & S_{21}' \\ 0 & S_{22}' \end{pmatrix}\begin{pmatrix} \hat W_{11} & a_n \hat W_{12} \\ a_n \hat W_{21} & a_n \hat W_{22} \end{pmatrix}\begin{pmatrix}S_{11} & 0 \\ S_{21} & S_{22} \end{pmatrix}\\ &=\begin{pmatrix}A_{11}+a_n B_{11} & a_n B_{12} \\ a_n B_{21} & a_n B_{22} \end{pmatrix} \end{align*} where $A_{11}=S_{11}' \hat W_{11}S_{11}$, $B_{11}=S_{11}'\hat W_{12}S_{21} + S_{21}'\hat W_{21} S_{11} + S_{21}'\hat W_{22}S_{21}$, $B_{12} = S_{11}'\hat W_{12} S_{22} + S_{21}' \hat W_{22} S_{22}$, $B_{21}= S_{22}\hat W_{21} S_{11} +S_{22} \hat W_{22}S_{21}$, and $B_{22}=S_{22}'\hat W_{22}S_{22}$. By Assumption (ref), $S_{ZX}$ is full column rank. Since this matrix is block lower diagonal, it implies that $S_{11}$ and $S_{22}$ are also full column rank. By Assumption (ref), $W_{11}$ and $W_{22}$ have full rank. It follows that $A_{11}$ and $B_{22}$ are invertible. By the inverse of a partitioned matrix, we obtain \begin{equation*} \left(S_{ZX}'WS_{ZX}\right)^{-1}=\\ \begin{pmatrix}C+a_n C B_{12}D B_{21}C & -C B_{12}D \\ -D B_{21}C & a_n^{-1}D \end{pmatrix} \end{equation*} where $C=\left(A_{11}+a_n B_{11}\right)^{-1}$ and $D=\left(B_{22}-a_n B_{21}CB_{12}\right)^{-1}$. Combining these results and taking into account that $a_n=o_p(1)$, we obtain \begin{align*} \left(S_{ZX}' \hat WS_{ZX}\right)^{-1}S_{ZX}'W =\begin{pmatrix}A_{11}^{-1}S_{11}'\hat W_{11} & 0 \\ -B_{22}^{-1}B_{21}A_{11}^{-1}S_{11}'\hat W_{11} + B_{22}^{-1}S_{22}'\hat W_{21} & B_{22}^{-1}S_{22}'\hat W_{22} \end{pmatrix}+O_p(a_n) \end{align*} [IT IS A MESS TO SHOW IT THIS WAY.] All the terms in this matrix are finite. Uniform convergence follows from the boundedness of $x_{ij}$ and $z_{ij}$, the invertibility of $A_{11}$ and $B_{22}$ and the uniform continuity of $W(\tau)$ as a function of $\tau$. \end{proof} \color{black} \begin{theoremp}{(ref)$'$}[Uniform consistency when the weighting matrix is asymptotically singular] Let the model in equation ((ref)), Assumptions (ref)-(ref), (ref)(a), and (ref) hold. Then, $$\underset{\tau\in\mathcal T}{\sup}\; \lVert \hat \delta(\tau) - \delta(\tau) \rVert =o_p (1)$$ \end{theoremp} \begin{proof}[Proof of Theorem (ref)] We obtain the result by replacing Lemma (ref) with (ref) in the proof of Theorem (ref). \end{proof}

Proof of Lemma (ref) - Asymptotic distribution of sample moments

proof[Proof of Lemma (ref)] Part (i) Lemma 3 in Galvao2020 provides the uniform Bahadur representation for the group-level quantile regression coefficient under our assumptions: \begin{equation} \hat\beta_j(\tau)-\beta_j(\tau)=\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+R_{nj}^{(1)}(\tau)+R_{nj}^{(2)}(\tau), \end{equation} where \begin{equation} \phi_{j,\tau}(\tilde x_{ij},y_{ij})=-B_{j,\tau}^{-1} \tilde x_{ij}(1(y_{ij}\leq \tilde x_{ij}\beta_j(\tau))-\tau), \end{equation} with $B_{j,\tau}=\mathbb{E}_{i|j}[f_{y|x}(Q_{y|x, \nu_j}(\tau|\tilde x_{ij} )|\tilde x_{ij})\tilde x_{ij} \tilde x_{ij}']$ and \begin{equation}\underset{j}{\sup}\;\underset{\tau\in \mathcal{T}}{\sup}\left\lVert R_{nj}^{(2)}(\tau)\right\lVert=O_p\left(\frac{\log n}{n}\right),\end{equation} \begin{equation}\underset{j}{\sup}\;\underset{\tau\in \mathcal{T}}{\sup}\left\lVert \mathbb{E}_{i|j}\left[R_{nj}^{(1)}(\tau)\right]\right\lVert=O\left(\frac{\log n}{n}\right),\end{equation} \begin{equation} \underset{j}{\sup}\;\underset{\tau\in \mathcal{T}}{\sup}\left\lVert \mathbb{E}_{i|j} \left[\left(R_{nj}^{(1)}(\tau)-\mathbb{E}_{i|j}[R_{nj}^{(1)}(\tau)]\right)\left(R_{nj}^{(1)}(\tau)-\mathbb{E}_{i|j}[R_{nj}^{(1)}(\tau)]\right)' \right]\right\lVert =O\left(\left(\frac{\log n}{n}\right)^{3/2}\right).\end{equation} It follows that \begin{align} \frac{1}{mn} \sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\tilde x_{ij}' \left(\hat \beta_j(\tau)-\beta_j(\tau)\right)=& \frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'\right)\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right) \\ &+ \frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'\right)R_{nj}^{(1)}(\tau) \\ &+\frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'\right)R_{nj}^{(2)}(\tau). \end{align} Consider first the third term ((ref)). By assumptions (ref)(i) and (ref)(i), $x_{ij}$ and $z_{ij}$ are bounded by $C$ such that the sample mean of their product is also bounded. Therefore, ((ref)) implies that \begin{equation} \underset{\tau\in\mathcal T}{\sup} \ \frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'\right)R_{nj}^{(2)}(\tau)=O_p\left(\frac{\log n}{n} \right). \end{equation} Consider now the second term ((ref)). Since $\operatorname{Var}\left(R_{nj}^{(1)}(\tau)\right)=o\left(\frac{1}{n}\right)$ by ((ref)), $x_{ij}$ and $z_{ij}$ are bounded by assumptions (ref)(i) and (ref)(i), and observations are independent across groups, it follows that $\operatorname{Var}\left(\frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'\right)R_{nj}^{(1)}(\tau)\right)=o_p\left(\frac{1}{mn}\right)$. In addition, by ((ref)), $\underset{j}{\sup}\ \underset{\tau\in \mathcal T} {\sup} \ \mathbb{E}_{i|j} \left [ R_{nj}^{(1)}(\tau) \right ]=O\left(\frac{\log n}{n}\right)$ such that $\underset{\tau\in\mathcal T}{\sup}\ \mathbb{E}_{i|j} \left[\frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i=1}^n z_{ij}\tilde x_{ij}'\right)R_{nj}^{(1)}(\tau)\right]=O\left(\frac{\log n}{n}\right)$. Putting this together, by the Chebyshev inequality and under Assumption (ref)(c), \begin{equation} \underset{\tau\in\mathcal T}{\sup} \frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'\right)R_{nj}^{(1)}(\tau) = o_p\left(\frac{1}{\sqrt{mn}}\right). \end{equation} It follows that both remainder terms are $o_p\left(\frac{1}{\sqrt{mn}}\right)$ uniformly over $\tau$. Let $\Sigma_{ZXj}=\mathbb{E}_{i|j}[z_{ij}\tilde x_{ij}']$ and consider now the term ((ref)): \begin{align} \frac{1}{m}\sum_{j = 1}^m \left(\frac{1}{n}\sum_{i=1}^n z_{ij}\tilde x_{ij}'\right) &\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right) \nonumber \\ =& \frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'-\Sigma_{ZXj}\right) \left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right) \nonumber \\ & +\frac{1}{m}\sum_{j = 1}^m\Sigma_{ZXj} \left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right). \end{align} By the boundedness of $z_{ij}$ and $x_{ij}$ and the independence of the observations over time, it follows that $\left\lVert \frac{1}{n}\sum_{i=1}^n z_{ij}\tilde x_{ij}'-\Sigma_{ZXj}\right\rVert = o(1)$ uniformly in $j$. In addition, $\operatorname{Var} \left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)=O\left(\frac{1}{n}\right)$. Hence, \begin{equation*} \operatorname{Var}\left(\frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'-\Sigma_{ZXj}\right) \left(\frac{1}{n}\sum_{i = 1}^n\phi_{i,\tau}(\tilde x_{ij},y_{ij})\right)\right) = o\left(\frac{1}{mn}\right). \end{equation*} The model in equation ((ref)) and Assumption (ref)(iii) imply that $\mathbb{E}_{i|j} \left[1(y_{ij}\leq\tilde x_{ij}\beta_j(\tau))|\tilde x_{ij},z_{ij}\right]=\tau$, which implies that $$\mathbb{E} \left[\frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'-\Sigma_{ZXj}\right) \left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)\right]=0$$ uniformly in $\tau$. Therefore, by Chebyshev's inequality, \begin{equation} \frac{1}{m}\sum_{j = 1}^m\left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'-\Sigma_{ZXj}\right) \left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right) =o_p\left(\frac{1}{\sqrt{mn}}\right) \end{equation} uniformly in $\tau$. Since all other terms are $o_p\left(\frac{1}{\sqrt{mn}}\right)$ uniformly over $\tau$, the limiting distribution of the process $\frac{1}{mn} \sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\tilde x_{ij}' \left(\hat \beta_j(\tau)-\beta_j(\tau)\right)$ is the same as the limiting distribution of \begin{multline} \frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)\\=\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{-B_{j,\tau}^{-1}}{n}\sum_{i = 1}^n \tilde x_{ij}(1(y_{ij}\leq \tilde x_{ij}\beta_j(\tau))-\tau)\right)=\frac{1}{mn}\sum_{j = 1}^m \sum_{i = 1}^n s_{ij}(\tau).\end{multline} This is a sample mean over $mn$ independent (but not necessarily identically distributed) observations denoted by $s_{ij}(\tau)$. The model in equation ((ref)) and Assumption (ref)(iii) imply that $\mathbb{E} \left[1(y_{ij}\leq\tilde x_{ij}\beta_j(\tau))|\tilde x_{ij},z_{ij},v_j\right]=\tau$, which implies that $\mathbb{E}[s_{ij}(\tau)]=0$. In addition, \begin{align} \operatorname{Var}(s_{ij}(\tau))= \mathbb{E}[\Sigma_{ZXj}\operatorname{Var}(\phi_{j,\tau})\Sigma_{ZXj}']= \mathbb{E} [\Sigma_{ZXj}B_{j,\tau}^{-1}\tau(1-\tau)\mathbb{E}_{i|j}[\tilde x_{ij} \tilde x_{ij}']B_{j,\tau}^{-1}\Sigma_{ZXj}']. \end{align} Pointwise asymptotic normality follows from the Lindeberg CLT. Next we note that $\left\{\Sigma_{ZXj}\left(\frac{-B_{j,\tau}^{-1}}{n}\sum_{i = 1}^n \tilde x_{ij}(1(y_{ij}\leq \tilde x_{ij}\beta)-\tau)\right), \tau\in\mathcal{T},\beta\in\mathcal B\right\}$ is a Donsker class for any compact set $\mathcal B$. This follows by noting that $\left\{1(y_{ij}\leq\tilde x_{ij}\beta_j(\tau)), \tau\in\mathcal T,\beta\in\mathcal B\right\}$ is a VC subgraph class and hence a bounded Donsker class. Hence, $$\left\{\frac{1}{n}\sum_{i = 1}^n\tilde x_{ij}(1(y_{ij}\leq\tilde x_{ij}\beta)-\tau), \tau\in\mathcal T,\beta\in\mathcal B\right\}$$is also bounded Donsker with a square-integrable envelope $2\cdot\max_{i\in 1,...,n}\lVert \tilde x_{ij}\rVert \leq 2\cdot C$. The whole function is then Donsker by the boundedness of $\Sigma_{ZXj}$ and $B_{j,\tau}^{-1}$. The weak convergence result follows by application of the functional central limit theorem for independent but not identically distributed random variables, see, for instance, Theorem 3 in Brown1971. Part (ii) Follows directly by Lemma 3 in Chetverikov2016. Part (iii) The first moment is asymptotically equivalent to (up to a term, which is uniformly $o_p\left(\frac{1}{\sqrt{mn}}\right)$): $$\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{-B_{j,\tau}^{-1}}{n}\sum_{i = 1}^n \tilde x_{ij}(1(y_{ij}\leq \tilde x_{ij}\beta_j(\tau))-\tau)\right).$$ We have already shown that both moments have a mean of zero. By Assumption (ref), the observations are independent across $i$ and $j$ such that we only need to consider the correlation between both moments for the same individual and group: \begin{align*} \operatorname{Cov}&(\tilde x_{ij}(1(y_{ij}\leq \tilde x_{ij}\beta_j(\tau))-\tau),z_{ij}\alpha_j(\tau'))\\ &=\mathbb{E}[\tilde x_{ij}(1(y_{ij}\leq \tilde x_{ij}\beta_j(\tau))-\tau)z_{ij}'\alpha_j(\tau')]\\ &=\mathbb{E}[\tilde x_{ij}\mathbb{E}_{i|j}[(1(y_{ij}\leq \tilde x_{ij}\beta_j(\tau))-\tau)|x_{ij},z_{ij}]z_{ij}'\alpha_j(\tau)] =0. \end{align*} It then directly follows that \begin{align*} \underset{\tau,\tau'\in\mathcal T}{\sup} \left\lVert \operatorname{Cov} \left (\bar g^{(1)}_{mn}(\hat\delta, \tau), \bar g^{(2)}_{mn}(\hat\delta, \tau')\right )\right\rVert=o_p\left(\frac{1}{\sqrt{mn}}\right). \end{align*}

Lemma (ref): Uniform consistency of the G matrix with a heterogeneous weighting matrix

lemmaAssumptions (ref), (ref), (ref), and (ref) hold. Then, uniformly in $\tau$, we can partition $\hat G(\tau)$ into $\hat G_{11}(\tau)$ (with dimensions $M_1 \times L_1$), $\hat G_{12}(\tau)$ ($M_1 \times L_2$), $\hat G_{21}(\tau)$ ($M_2 \times L_1$), and $\hat G_{22}(\tau)$ ($M_2 \times L_2$) such that \begin{equation*} \hat G(\tau)=\begin{pmatrix} \hat G_{11}(\tau) & \hat G_{12}(\tau) \\ \hat G_{21}(\tau) & \hat G_{22}(\tau) \end{pmatrix}=\begin{pmatrix} G_{mn,11}(\tau) & G_{mn,12}(\tau) \\ G_{mn,21}(\tau) & G_{mn,22}(\tau) \end{pmatrix}+ \begin{pmatrix} o_p\left(1\right) & o_p \left (\sqrt{a_n(\tau)}\right ) \\ o_p \left (1/\sqrt{a_n(\tau)} \right) & o_p\left(1\right) \end{pmatrix}, \end{equation*} where $G_{mn}(\tau)$ is defined in equation ((ref)), $\underset{\tau\in \mathcal{T}}{\sup} \ G_{mn,12}(\tau)=O_p\left(a_n(\tau)\right)$ and the other elements of $G_{mn}(\tau)$ are $O_p\left(1\right)$ uniformly in $\tau$. \begin{comment} Then for entries $k\in\{K_1+1,\dots,K\}$ and $l\in\{1,\dots,L_1\}$ we have uniformly in $\tau \in \mathcal{T}$ \begin{equation*} \hat G_{kl}(\tau)-G_{mn,kl}(\tau)=o_p \left (1/\sqrt{a_n(\tau)} \right ). \end{equation*} and entries for $k\notin\{K_1+1,\dots,K\}$ and $l\notin\{1,\dots,L_1\}$ we have uniformly in $\tau \in \mathcal{T}$ \begin{equation*} \hat G_{kl}(\tau) = G_{mn, kl}(\tau) + o_p\left ( \sqrt{G_{mn, kl}(\tau)} \right), \end{equation*} Further, for $k \in \{1, \dots, K_1\}$ and $l \in \{ L_1 +1 , \dots , L \}$ it holds uniformly over $\tau$ that $G_{mn,kl} = O_p(a_n)$ and for all other elements $G_{mn, kl} = O_p(1)$. \end{comment}
proofTo simplify the notation, we omit the dependency on $\tau$. We partition $S_{ZX}'$ into four submatrices that have the same dimensions as the submatrices of $\hat G$. As in proposition (ref), we also partition $\hat W$ into four submatrices such that the diagonal submatrices are square matrices of dimensions $L_1$ and $L_2$. Note that \begin{align*} S'_{ZX} \hat W =\begin{pmatrix}S_{11}' & S_{21}' \\ 0 & S_{22}' \end{pmatrix}\begin{pmatrix} \hat W_{11} & \hat W_{12} \\ \hat W_{21} & \hat W_{22} \end{pmatrix}= \begin{pmatrix}S_{11}' \hat W_{11}+ S_{21}'\hat W_{21} & S_{11}'\hat W_{12}+S_{21}'\hat W_{22} \\ S_{22}'\hat W_{21} & S_{22}'\hat W_{22} \end{pmatrix}. \end{align*} Furthermore, \begin{align*} S_{ZX}'\hat WS_{ZX}&= \begin{pmatrix}S_{11}' & S_{21}' \\ 0 & S_{22}' \end{pmatrix}\begin{pmatrix} \hat W_{11} & \hat W_{12} \\ \hat W_{21} & \hat W_{22} \end{pmatrix}\begin{pmatrix}S_{11} & 0 \\ S_{21} & S_{22} \end{pmatrix}\\ &=\begin{pmatrix} \hat A_{11}+\hat B_{11} & \hat B_{12} \\ \hat B_{21} & \hat B_{22} \end{pmatrix}, \end{align*} where $\hat A_{11}=S_{11}'\hat W_{11}S_{11}$, $\hat B_{11}=S_{11}'\hat W_{12}S_{21} + S_{21}' \hat W_{21} S_{11} + S_{21}'\hat W_{22}S_{21}$, $\hat B_{12} = S_{11}' \hat W_{12} S_{22} + S_{21}' \hat W_{22} S_{22}$, $ \hat B_{21}= S_{22}'\hat W_{21} S_{11} +S_{22}'\hat W_{22}S_{21}$, and $\hat B_{22}=S_{22}'\hat W_{22}S_{22}$. For the non-zero elements of the matrix $S_{ZX}$, we have \begin{align*} S_{11} - \Sigma_{11} = o_p(1), \\ S_{22} - \Sigma_{22} = o_p(1) ,\\ S_{21} - \Sigma_{21} = o_p(1). \end{align*} Together with proposition (ref), this implies that uniformly over $\tau$, \begin{align} \hat A_{11} =& \Sigma_{11}'W_{11}\Sigma_{11} + o_p(1) , \\[1em] \hat B_{11}= & a_n\Sigma_{11}'W_{12}\Sigma_{21} + a_n \Sigma_{21}' W_{21} \Sigma_{11} + a_n \Sigma_{21}'W_{22}\Sigma_{21} + o_p(\sqrt{a_n}) = a_n B_{11} + o_p(\sqrt{a_n}) , \\[1em] \hat B_{12} = & a_n \Sigma_{11}'W_{12} \Sigma_{22} + a_n \Sigma_{21}' W_{22} \Sigma_{22} + o_p(\sqrt{a_n}) = a_n B_{12} + o_p(\sqrt{a_n}), \\[1em] \hat B_{21} = & a_n \Sigma_{22}' W_{21} \Sigma_{11} +a_n \Sigma_{22}' W_{22}\Sigma_{21} + o_p(\sqrt{a_n})= a_n B_{21} + o_p(\sqrt{a_n}),\\[1em] \hat B_{22}=& a_n\Sigma_{22}'W_{22}\Sigma_{22} + o_p(a_n) = a_n B_{22} + o_p(a_n). \end{align} \begin{comment} \begin{align*} \hat B_{12} = & S_{11}' \hat W_{12} S_{22} + S_{21}' W_{22} S_{22} = \left ( \Sigma_{11}' + o_p(1) \right) \left ( a_n W_{12} + o_p(\sqrt{a_n}) \right) \left ( \Sigma_{22} + o_p(1) \right) + S_{21}' W_{22} S_{22} \& = a_n \Sigma_{11}'W_{12} \Sigma_{22} + a_n \Sigma_{21}' W_{22} \Sigma_{22} + o_p(\sqrt{a_n}) \\ = & a_n B_{12} + o_p(\sqrt{a_n}) \end{align*} \end{comment} By the inverse of a partitioned matrix \begin{equation*} \left(S_{ZX}' \hat WS_{ZX}\right)^{-1}=\\ \begin{pmatrix}\hat C+\hat C \hat B_{12}\hat D \hat B_{21}\hat C & -\hat C \hat B_{12}\hat D \\ -\hat D \hat B_{21}\hat C & \hat D \end{pmatrix}, \end{equation*} where $\hat C=\left(\hat A_{11}+ \hat B_{11}\right)^{-1}$ and $\hat D=\left(\hat B_{22}-\hat B_{21}\hat C\hat B_{12}\right)^{-1}$. Hence uniformly over $\tau$, \begin{align} \hat C = & C + o_p(1), \\ \hat D = & a_n^{-1} D + o_p \left ( a_n^{-1} \right), \\ \hat D \hat B_{21} \hat C =& D B_{21} C + o_p(a_n^{-1/2}), \end{align} where $C$ and $D$ are strictly positive and bounded. Equation ((ref)) holds because $$a_n \hat D = \left ( a_n^{-1} \hat B_{22}-a_n^{-1}\hat B_{21}\hat C\hat B_{12} \right)^{-1} = \left( B_{22}- a_n B_{21} C B_{12} \right)^{-1} + o_p(1) = D + o_p(1).$$ Combining equation ((ref)) with equations ((ref))-((ref)), we obtain similarly \begin{equation} \hat C \hat B_{12} \hat D = C B_{12} D + o_p(a_n^{-1/2}). \end{equation} Now, we can consider each submatrix of $\hat G$ separately and derive their convergence rate to the corresponding element of the $G_{mn}$ matrix defined in equation ((ref)): $$G_{mn}(\tau) = \left ( \Sigma_{ZX}' W_{mn} (\tau) \Sigma_{ZX} \right) ^{-1} \Sigma_{ZX}' W_{mn}(\tau).$$ For the upper left term, we have \begin{align*} \hat G_{ 11}& = ( \hat C+ \hat C \hat B_{12}\hat D \hat B_{21}\hat C )(S_{11}'\hat W_{11}+ S_{21}'\hat W_{21} ) - \hat C \hat B_{12} \hat D S_{22}' \hat W_{21}\\ &= \left( C+ a_n C B_{12} D B_{21} C + o_p(1) \right) \left (\Sigma_{11}'W_{11}+a_n \Sigma_{21}'W_{21} + o_p(1) \right) \\ & \quad + \left ( C B_{12} D + o_p(a_n^{-1/2}) \right) \left (a_n \Sigma_{22}' W_{21} + o_p(\sqrt{a_n}) \right) \&= ( C+ a_n C B_{12} D B_{21} C )(\Sigma_{11}'W_{11}+a_n \Sigma_{21}'W_{21} ) - a_n C B_{12} D\Sigma_{22}' W_{21} + o_p(1)\\ & = G_{mn, 11} + o_p(1), \end{align*} uniformly over $\tau$. For the upper right term, we have uniformly over $\tau$, \begin{align*} \hat G_{12} =& ( \hat C+ \hat C \hat B_{12} \hat D \hat B_{21}\hat C ) \left(S_{11}'\hat W_{12}+S_{21}' \hat W_{22}\right) - \hat C \hat B_{12}\hat D S_{22}' \hat W_{22} \\ =& \left( C+ a_n C B_{12} D B_{21} C + o_p(1) \right) \left (a_n\left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) + o_p(\sqrt{a_n}) \right) \\ & - \left ( C B_{12} D + o_p(a_n^{-1/2}) \right) \left (a_n \Sigma_{22}' W_{22} + o_p(a_n) \right)\\ = &a_n( C+ a_n C B_{12} D B_{21} C ) \left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) - a_n C B_{12} D \Sigma_{22}' W_{22} + o_p(\sqrt{a_n}) \\ =& G_{mn,12} + o_p(\sqrt{a_n}). \end{align*} For the lower right term, \begin{align*} \hat G_{22} = & -\hat D \hat B_{21} \hat C \left(S_{11}'\hat W_{12}+S_{21}'\hat W_{22}\right) + \hat D S_{22}' \hat W_{22} \\ =&-\left( D B_{21} C + o_p({a_n}^{-1/2}) \right) \left(a_n\Sigma_{11}'W_{12}+a_n\Sigma_{21}'W_{22} + o_p(\sqrt{a_n} )\right) \\ = &- a_n D B_{21} C \left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) + D \Sigma_{22}' W_{22} + o_p(1) \\ = & G_{mn,22} + o_p(1) , \end{align*} uniformly over $\tau$. \begin{comment} \begin{align*} (A + B)^{-1} - A^{-1} = & A A^{-1} (A + B)^{-1} - (A+ B) (A + B)^{-1}A^{-1} \\ & = A A^{-1} (A + B)^{-1} - (A+ B) (A + B)^{-1}A^{-1} + (A+ B) A^{-1}(A + B)^{-1} - (A+ B) A^{-1}(A + B)^{-1}\\ &= [A - (A+ B) ]A^{-1}(A + B)^{-1} - (A+ B) [(A + B)^{-1}A^{-1} - A^{-1}(A + B)^{-1}] \\ & - B A^{-1}(A + B)^{-1} - (A+ B) [(A + B)^{-1}A^{-1} - A^{-1}(A + B)^{-1}] \end{align*} where the second term has non-zero entries only in the off-diagonal terms \end{comment} Finally, for the lower left term, uniformly of $\tau$, \begin{align*} \hat G_{21} = & -\hat D \hat B_{21} \hat C (S_{11}'\hat W_{11}+ S_ {21}'\hat W_{21} ) +\hat D S_{22}' \hat W_{21} \nonumber \\ =& -\left (D B_{21} C + o_p(a_n^{-1/2}) \right) \left (\Sigma_{11}' W_{11} + a_n \Sigma_{21}' W_{21} + o_p(\sqrt{a_n}) \right) \nonumber \\ &+\left (a_n^{-1} D + o_p(a_n^{-1} ) \right) \left( \Sigma_{22}' a_n W_{21} + o_p(\sqrt{a_n}) \right) \nonumber \\ =& - D B_{21} C \left ( \Sigma_{11}' W_{11} - \Sigma_{21}'a_n W_{21} \right) + D\Sigma_{22}' W_{21} + o_p({a_n}^{-1/2}) \\ = & G_{mn,21} + o_p \left ( a_n ^{-1/2}\right ). \end{align*} \begin{comment} OLD VERSION.\\ For the lower left term, we have \begin{align*} \hat G_{mn,21} = & -\hat D \hat B_{21} \hat C (S_{11}'\hat W_{11}+ S_ {21}'\hat W_{21} ) +\hat D S_{22}' \hat W_{21} \\ =& \hat D \bigg( {- \hat B_{21} \hat C S_{11}'\hat W_{11} }- \hat B_{21} \hat C S_ {21}'\hat W_{21} + { S_{22}' \hat W_{21}} \bigg) \\ =& \hat D \bigg( - \left({\color{red}S_{22}'\hat W_{21} S_{11}\hat C S_{11}'\hat W_{11} } +S_{22}' \hat W_{22}S_{21} \hat C S_{11}'\hat W_{11} \right) - \hat B_{21} \hat C S_ {21}'\hat W_{21} + {\color{red} S_{22}' \hat W_{21}} \bigg) \end{align*} The {goal} is to show that the terms inside the parenthesis converge to some expression $+ o_p(a_n).$ [we need $a_n$ because the entire expression is multiplied by $\hat D =a_n^{-1}D + o_p(a_n^{-1})$ ] For the terms in blue we have $$\hat B_{21} \hat C S_ {21}'\hat W_{21} = a_n B_{21} C S_ {21}'a_n W_{21} + o_p(a_n)$$ $$S_{22}' \hat W_{22}S_{21} \hat C S_{11}'\hat W_{11} = \Sigma_{22}' a_n W_{22}\Sigma_{21} C \Sigma_{11}'\hat W_{11} + o_p(a_n) $$ For the terms in red, we have \begin{align} S_{22}'\hat W_{21} & S_{11}\hat C S_{11}'\hat W_{11} - S_{22}' \hat W_{21} = S_{22}'\hat W_{21} \left( S_{11}\hat C S_{11}'\hat W_{11} - I \right) \\ = & S_{22}'\hat W_{21} \left( S_{11} \left( \hat A_{11}^{-1} - \hat A_{11}^{-1} \hat B_{11} ( \hat A_{11}+ \hat B_{11})^{-1}\right) S_{11}'\hat W_{11} - I \right) \\ = & S_{22}'\hat W_{21} \left (S_{11} \hat A_{11}^{-1} S_{11}'\hat W_{11} - I \right) - S_{22}'\hat W_{21} \hat A_{11}^{-1} \hat B_{11} ( \hat A_{11}+ \hat B_{11})^{-1} S_{11}'\hat W_{11} \end{align} For the second term $$ S_{22}'\hat W_{21} \hat A_{11}^{-1} \hat B_{11} ( \hat A_{11}+ \hat B_{11})^{-1} S_{11}'\hat W_{11} = \Sigma_{22}'a_n W_{21} A_{11}^{-1} a_n B_{11} ( A_{11}+ a_n B_{11})^{-1} \Sigma_{11}' W_{11} + o_p(a_n) $$ The first term contains an annihilator matrix. Let $ M = \left( S_{11}\hat A_{11}^{-1} S_{11}'\hat W_{11} - I \right)$. $-M$ has $L_1-K_1$ positive (= 1) eigenvalues and $K_1$ zero eigenvalues. $Rank(M) = 1.$ Hence, there is a plane full of vectors that is squeezed into the origin. If $K_1 = 1$ is trivial to show that $M = 0$ because $A_{11}$ is a scalar. \\ If $L_1 = K_1$, $A_{11} = (S_{11}' W_{11} S_{11})^{-1} = S_{11}^{-1} W_{11}^{-1} S_{11}^{-1} $ so that $M = 0.$ It looks like an algebraic argument works only if $K_1$ or $K_1 = L_1$. But in simulations, it works even with $L_1 = 3$, $K_1 = 2$, $K_2 = 2$, $L_2 = 4$ and everything is correlated. Then, \begin{align*} \hat G_{21} = & -\hat D \hat B_{21} \hat C (S_{11}'\hat W_{11}+ S_ {21}'\hat W_{21} ) +\hat D S_{22}' \hat W_{21} \\ =& \hat D \bigg( - \hat B_{21} \hat A_{11}^{-1} S_{11}'\hat W_{11}- \hat B_{21} \hat A_{11}^{-1} S_ {21}'\hat W_{21} \\ & - \hat B_{21}[(\hat A_{11} + \hat B_{11} ) \hat A_{11}]^{-1} \hat B_{11} (S_{11}'\hat W_{11}+ S_ {21}'\hat W_{21} ) + S_{22}' \hat W_{21} \bigg)\\ =&\hat D \left( {- S_{22}'\hat W_{21} }-S_{22}'\hat W_{22}S_{21} S_{11}'^{+}- \hat B_{21} \hat A_{11}^{-1} S_ {21}'\hat W_{21} { +S_{22}' \hat W_{21} } + \hat E \right)\\ = &a_n^{-1} D \left ( - a_n \Sigma_{22}' W_{22} \Sigma_{21} \Sigma_{11}'^{+} -a_n B_{21} A_{11}^{-1} \Sigma_{21}' a_n W_{21} + a_n^2 E + o_p(a_n) \right) \\ = & - D \Sigma_{22}' W_{22} \Sigma_{21} \Sigma_{11}'^{+} +a_n D B_{21} A_{11}^{-1} \Sigma_{21} W_{21} + a_n D E + o_p(1) \end{align*} Where $^+$ is the Moore-Penrose Inverse. The second line uses equation ((ref)). The third and the fourth lines use $$\hat E = \hat B_{21}[(\hat A_{11} + \hat B_{11} ) \hat A_{11}]^{-1} \hat B_{11} (S_{11}'\hat W_{11}+ S_ {21}'\hat W_{21} ) = a_n^2 E + o_p(a_n)$$ with $$E = B_{21}[( A_{11} + a_n B_{11} ) A_{11}]^{-1} B_{11} (\Sigma_{11}' W_{11}+ \Sigma_{21}'a_n W_{21} ). $$ In the third line, we insert $\hat B_{21} = S_{22}'\hat W_{21} S_{11} +S_{22}'\hat W_{22}S_{21}$ and $\hat A_{11} = S_{11}'\hat W_{11}S_{11}$. The second-last line follows because the sufficient conditions for continuity of the Moore-Penrose inverse are satisfied (see Rakocevic1997). \begin{commentP} More detailed derivation: By the properties of the Moore-Penrose Inverse, we have \begin{align*} -\hat B_{21} \hat A_{11}^{-1} S_{11}'\hat W_{11} = & - \left ( S_{22}\hat W_{21} S_{11} +S_{22}\hat W_{22}S_{21}\right) \left (S_{11}'\hat W_{11}S_{11}\right)^{-1} S_{11}'\hat W_{11} \\ = & - \left ( S_{22}\hat W_{21} S_{11} +S_{22}\hat W_{22}S_{21}\right) \left (\hat W_{11}^{1/2}S_{11}\right)^{+} W_{11}^{1/2} \\ & = - S_{22}\hat W_{21} -S_{22}\hat W_{22}S_{21} S_{11}'^{+} \end{align*} so that \begin{align*} \hat D \left( - \hat B_{21} \hat A_{11}^{-1} S_{11}'\hat W_{11}- \hat B_{21} \hat A_{11}^{-1} S_ {21}'\hat W_{21} + S_{22}' \hat W_{21} + o_p( {a_n}) \right)\\ \hat D \left( {\color{blue} - S_{22}\hat W_{21} }-S_{22}\hat W_{22}S_{21} S_{11}'^{+}- \hat B_{21} \hat A_{11}^{-1} S_ {21}'\hat W_{21} { +\color{blue}S_{22}' \hat W_{21} } + o_p( {a_n}) \right)\\ = a_n^{-1} D \left ( - a_n \Sigma_{22} W_{22} \Sigma_{21} \Sigma_{11}'^{+} -a_n B_{21} A_{11}^{-1} \Sigma_{21} a_n W_{21} + o_p(a_n) \right) \\ = - D \Sigma_{22} W_{22} \Sigma_{21} \Sigma_{11}'^{+} +a_n D B_{21} A_{11}^{-1} \Sigma_{21} W_{21} + o_p(1) \end{align*} Where the second line follows because the conditions for continuity of the Moore-Pensore inverse are satisfied (see Rakocevic1997). \end{commentP} \\ Finally, for the lower right term, \begin{align*} \hat G_{mn,22} = & -\hat D \hat B_{21} \hat C \left(S_{11}'\hat W_{12}+S_{21}'\hat W_{22}\right) + \hat D S_{22}' \hat W_{22} + o_p(1) \\ = &- a_n D B_{21} C \left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) + D \Sigma_{22}' W_{22} + o_p(1) \end{align*} \begin{align*} \hat G_{22} = & -\hat D \hat B_{21} \hat C \left(S_{11}'\hat W_{12}+S_{21}'\hat W_{22}\right) + \hat D S_{22}' \hat W_{22} + o_p(1) \\ = &-(\left( D B_{21} C + o_p(1) \right) \left(a_n\Sigma_{11}'W_{12}+a_n\Sigma_{21}'W_{22} + o_p(\sqrt{a_n} )\right) + \left( a_n^{-1}D + o_p(1) \right) \left( a_n \Sigma_{22}' W_{22} + o_p(a_n) \right) \end{align*} \end{comment} Hence, \begin{equation*} \sup_\tau \left \rVert \hat G(\tau)- G_{mn}(\tau) \right \rVert = \begin{pmatrix} o_p(1) & o_p \left (\sqrt{a_n(\tau)}\right ) \\ o_p \left (1/\sqrt{a_n(\tau)} \right ) & o_p(1) \end{pmatrix} \end{equation*} where $G_{mn,12}(\tau)=O_p(a_n(\tau))$ and the other elements of $G_{mn}(\tau)$ are $O_p(1)$ uniformly over $\tau$.
commentThus $$\hat C = C + o_p(1)$$ $$ \hat D= \left (\hat B_{22}-\hat B_{21}\hat C\hat B_{12} \right)^{-1} = \left ( a_n( B_{22}- a_n B_{21} C B_{12} ) \right)^{-1} + o_p(1) = a_n^{-1} D + o_p(1 / \sqrt{a_n})$$ where $D = O_p(1)$, $C = O_p(1)$ \begin{equation*} \left(S_{ZX}' \hat WS_{ZX}\right)^{-1}=\\ \begin{pmatrix} C+ a_n C B_{12} D B_{21} C & - C B_{12} D \\ - D B_{21} C & a_n^{-1}D \end{pmatrix} + \begin{pmatrix} o_p(1) & o_p(1/ \sqrt{a_n}) \\ o_p(1/ \sqrt{a_n}) & o_p(a_n^{-1}) \end{pmatrix} \end{equation*} \begin{align*} S'_{ZX} \hat W & = \begin{pmatrix}\Sigma_{11}' + o_p(1)& \Sigma_{21}' + o_p(1) \\ 0 & \Sigma_{22}' + o_p(1) \end{pmatrix}\begin{pmatrix} W_{11} + o_p(1)& a_n W_{12} + o_p(\sqrt{a_n}) \\ a_n W_{21} + o_p(\sqrt{a_n}) & a_n W_{22} + o_p({a_n}) \end{pmatrix} \\ & = \begin{pmatrix}\Sigma_{11}'W_{11}+a_n \Sigma_{21}'W_{21} + o_p(1) & a_n\left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) + o_p(\sqrt{a_n})\\ a_n \Sigma_{22}' W_{21} + o_p(\sqrt{a_n}) & a_n \Sigma_{22}' W_{22} + o_p({a_n}) \end{pmatrix} \end{align*} \begin{align*} \hat G = \left(S_{ZX}' \hat WS_{ZX}\right)^{-1}S'_{ZX} \hat W = & \begin{pmatrix} C+ a_n C B_{12} D B_{21} C + o_p(1) & - C B_{12} D + o_p(a_n) \\ - D B_{21} C + o_p(a_n) & a_n^{-1}D + o_p(a_n) \end{pmatrix} \\ & \begin{pmatrix}\Sigma_{11}'W_{11}+a_n \Sigma_{21}'W_{21} + o_p(1) & a_n\left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) + o_p(\sqrt{a_n}) \\ a_n \Sigma_{22}' W_{21}+ o_p(\sqrt{a_n}) & a_n \Sigma_{22}' W_{22} + o_p(a_n) \end{pmatrix} \end{align*} \begin{equation} \hat G_{11}= ( C+ a_n C B_{12} D B_{21} C )(\Sigma_{11}'W_{11}+a_n \Sigma_{21}'W_{21} ) - a_n C B_{12} D\Sigma_{22}' W_{21} + o_p(1) \end{equation} \begin{equation} \hat G_{[12]mn} = a_n( C+ a_n C B_{12} D B_{21} C ) \left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) - a_n C B_{12} D \Sigma_{22}' W_{22} + o_p(\sqrt{a_n}) \end{equation} \begin{equation} \hat G_{[21]mn} = -D B_{21} C (\Sigma_{11}'W_{11}+a_n \Sigma_{21}'W_{21} ) + D \Sigma_{22}' W_{21} + o_p(1) \end{equation} \begin{equation} \hat G_{22} = -a_n D B_{21} C \left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) + D \Sigma_{22}' W_{22} + o_p(1) \end{equation}
comment\begin{lemma} Define the matrix \begin{align*} G_{mn} (\tau)= \left ( \Sigma_{ZX}'W(\tau)\Sigma_{ZX} \right )^{-1}\Sigma_{ZX}'W(\tau) \end{align*} Under assumptions (ref), (ref), (ref), and (ref) hold. for an element in the $k$-th row and $l$-th column of $\hat G(\tau)$ and $G_{mn}(\tau)$ we have $\hat G_{kl} - G_{mn, kl} = o_p(G_{mn,kl})$. \end{lemma} \begin{proof} From the proof of Lemma (ref), we have that all elements of $G_{mn}$ are $O_p(1)$. Similarly, $\hat G - G_{mn} = o_p(1)$. Consider now the $M_1 \times L_2$ top-right submatrix. Using the same derivation as in the proof of Lemma (ref), we have \begin{equation*} \hat G_{12}= a_n \left [ ( \hat C+a_n \hat C \hat B_{12}\hat D \hat B_{21}\hat C ) \left(S_{11}'\hat W_{12}+S_{21}'\hat W_{22}\right) -\hat C \hat B_{12}\hat D S_{22}' \hat W_{22} \right ] \end{equation*} \begin{equation*} G_{mn, 12} =a_n \left [( C+a_n C B_{12}D B_{21}C ) \left(\Sigma_{11}'W_{12}+\Sigma_{21}'W_{22}\right) -C B_{12}D \Sigma_{22}' W_{22} \right ] \end{equation*} Note that all matrices in $\hat G_{12}$ converge in probability to the liming matrices. Then by the continuous mapping theorem $\hat G_{12} - G_{mn,12} = o_p(a_n)$. Further, since all matrices in $G_{mn,12}$ are finite, $G_{mn,12} = O_p(a_n)$. Combining, we get that for all elements $k$, $s$, of the matrix we have: \begin{equation*} \hat G_{ks} - G_{mn, ks} = o_p(G_{mn, ks}) \end{equation*} \end{proof}

Proof of Theorem (ref): Degree of heterogeneity is known

proof[Proof of Theorem (ref)] From the definition of the estimator, \begin{align} \hat\delta(\tau)-\delta(\tau)&=\left(S_{ZX}'\hat W(\tau)S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau) \bar g_{mn}(\hat \delta,\tau).\end{align} In part (i) Lemma (ref) applies uniformly in $\tau$ with $a_n(\tau)=O_p(1/n)$ such that \begin{align*} \left(S_{ZX}'\hat W(\tau) S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau)&=\begin{pmatrix} \hat G_{11}(\tau) & \hat G_{12}(\tau)\\ \hat G_{21}(\tau) & \hat G_{22}(\tau) \end{pmatrix}\\ &=\begin{pmatrix} G_{11}(\tau)+o_p(1) & o_p(1/\sqrt n) \\ o_p(\sqrt n) & G_{22}(\tau)+o_p(1) \end{pmatrix}. \end{align*} By the definition of the fast instruments, the first $L_1$ elements of $\bar g^{(2)}_{mn}(\hat\delta,\tau)$ are equal to zero. Together with Lemma (ref)(i), this implies that \begin{equation} \sqrt{mn}\bar g_{mn,1}(\hat \delta,\cdot)\rightsquigarrow \mathbb Z_{11}(\cdot), \end{equation} where $\bar g_{mn,1}(\hat \delta,\cdot)$ contains the first $L_1$ elements of $\bar g_{mn}(\hat \delta,\cdot)$. For the remaining $L_2$ elements, we have $\sqrt{m}\bar g_{mn,2}(\hat \delta,\cdot) = O_p(1)$. It follows that \begin{align*} \sqrt{mn}\left(\hat\delta_1(\cdot)-\delta_1(\cdot)\right)&=\hat G_{11}\sqrt{mn}\bar g_{mn,1}(\hat\delta,\cdot)+\hat G_{12}\bar g_{mn,2}(\hat\delta,\cdot)\\ &=G_{11}\sqrt{mn}\bar g_{mn,1}(\hat\delta,\cdot)+o_p(1)\sqrt{mn}\bar g_{mn,1}(\hat\delta,\cdot)+o_p(1/\sqrt{n})\sqrt{mn}\bar g_{mn,2}(\hat\delta,\cdot)\\ &=G_{11}\sqrt{mn}\bar g_{mn,1}(\hat\delta,\cdot)+o_p(1)\sqrt{mn}\bar g_{mn,1}(\hat\delta,\cdot)+o_p(1)\sqrt{m}\bar g_{mn,2}(\hat\delta,\cdot)\\ &\rightsquigarrow G_{11} \mathbb Z_{11}(\cdot)+o_p(1), \end{align*} which proves part (i)-(a) of the theorem. For the remaining $M_2$ coefficients, applying the same results, we obtain \begin{align*} \sqrt{m}\left(\hat\delta_2(\cdot)-\delta_2(\cdot)\right)&=\hat G_{21}\sqrt{m}\bar g_{mn,1}(\hat\delta,\cdot)+\hat G_{22}\sqrt{m}\bar g_{mn,2}(\hat\delta,\cdot)\\ &=o_p(\sqrt{n})\sqrt{m}\bar g_{mn,1}(\hat\delta,\cdot)+G_{22}\sqrt{m}\bar g_{mn,2}(\hat\delta,\cdot)+o_p(1)\sqrt{m}\bar g_{mn,2}(\hat\delta,\cdot)\\ &\rightsquigarrow G_{22}(\cdot) \mathbb Z_{22}(\cdot)+o_p(1). \end{align*} The result of part (i)-(b) of the theorem follows. In part (ii) and (iii), Lemma (ref) applies such that \begin{equation} \left(S_{ZX}'\hat W(\tau) S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau)\underset{p}{\rightarrow}G(\tau) \end{equation} uniformly in $\tau\in\mathcal{T}$. In part (ii), $\bar g_{mn}(\hat \delta,\cdot)=\bar g_{mn}^{(1)}(\hat \delta,\cdot)$. It follows from Lemma (ref)(i) that \begin{equation} \sqrt{mn}\bar g_{mn}(\hat \delta,\cdot)\rightsquigarrow \mathbb Z_{1}(\cdot). \end{equation} The result of part (ii) of the theorem follows from equations ((ref)), ((ref)), and ((ref)). In part (iii), $\bar g_{mn}^{(1)}(\hat\delta,\tau)$ and $\bar g_{mn}^{(2)}(\hat\delta,\tau)$ converge at the same rate such that Lemma (ref) implies \begin{equation} \sqrt{mn}\bar g_{mn}(\hat \delta,\cdot)\rightsquigarrow \mathbb Z(\cdot), \end{equation} where $\mathbb Z(\cdot)$ is a mean-zero Gaussian process with uniformly continuous sample paths and covariance function $\Omega_1(\tau,\tau')+\bar\Omega_1(\tau,\tau')$. The result of part (iii) of the theorem follows from equations ((ref)), ((ref)), and ((ref)).
comment\subsubsection{Weak convergence of the estimated quantile process} \begin{theorem}[Asymptotic distribution of $\hat \gamma$] Assume that conditions (ref)-(ref), (ref)(b), (ref), and (ref) hold. Then \begin{equation*} \sqrt{m} \left ( \hat \gamma(\cdot ) - \gamma(\cdot) \right) \Rightarrow \mathbb{G}( \cdot ) \quad in \quad \ell ^\infty(\mathcal{T}) \end{equation*} where $\mathbb{G} $ is a zero-mean Gaussian process with uniformly continuous sample paths and covariance function $QJ(\tau_1, \tau_2) Q'$, where $Q = \left( \Sigma_{22}' W_{22} \Sigma_{22}\right)^{-1} \Sigma_{22}' W_{22} $. \end{theorem} \begin{proof}[Proof of Theorem (ref)] From equation ((ref)) we have \begin{align*} \sqrt N \bar g_{mn}(\delta,\tau)= & \frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij} \left(\tilde x_{ij}'\left (\frac{1}{n}\sum_{i=1}^n\phi(\tilde x_{ij},y_{ij})+R_{ij}^{(1)}(\tau)+R_{ij}^{(2)}(\tau) \right )+\alpha_j(\tau)\right)\\ = & \frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij} \tilde x_{ij}'\frac{1}{n}\sum_{i=1}^n\phi(\tilde x_{ij},y_{ij})+ \frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'R_{ij}^{(1)}(\tau)\\ & + \frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}' R_{ij}^{(2)}(\tau)+\frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij}\alpha_j(\tau) \end{align*} For the last term, by Lemma 3 in Chetverikov2016 follows that \begin{align*} \frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij}\alpha_j(\cdot ) = \frac{1}{\sqrt N}\sum_{j = 1}^m \bar z_{j}\alpha_j(\cdot ) \rightarrow \mathbb G^0(\cdot), \quad in \quad l^\infty(\mathcal{T}) \end{align*} where $\mathbb{G}^0$ is a zero-mean Gaussian process with uniformly continuous sample paths and covariance function $J (\tau_1, \tau_2) $. For the third term, Lemma 3 in Galvao2020 implies $\sup_j \sup_{\tau \in \mathcal{T}} || R^{(2)}_{ij} || = O_p\left ( \frac{\log n}{n} \right )$, and since $\tilde x_{ij}$ and $z_j$ are bounded, \begin{align*} \frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij} \tilde x_{ij}' R_{ij}^{(2)}(\tau) = O_p \left ( \frac{\log n \sqrt{m}}{n} \right ). \end{align*} uniformly over $\tau$. For the second term, note that Lemma 3 in Galvao2020 implies that $\operatorname{Var}(R_{ij}^{(1)}(\tau) ) = o \left (\frac{1}{n} \right) $ and $\mathbb{E}[ \sqrt N \sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'R_{ij}^{(1)}(\tau) ] = O \left (\frac{\log n \sqrt{m}}{n} \right) $ uniformly over $\tau$. Thus, \begin{align*} \frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'R_{ij}^{(1)}(\tau) = O \left ( \frac{\log n \sqrt{m}}{n} \right) + o_p \left ( \frac{1}{\sqrt{n}}\right) \end{align*} uniformly over $\tau$. Then, \begin{align*} \sqrt N \bar g_{mn}(\delta,\tau)=& \frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij} \tilde x_{ij}'\frac{1}{n}\sum_{i=1}^n\phi(\tilde x_{ij},y_{ij})+\frac{1}{\sqrt N}\sum_{j = 1}^m\frac{1}{n}\sum_{i = 1}^n z_{ij}\alpha_j(\tau) + o_p(1) \end{align*} uniformly over $\tau$. Using the results of the proof of Lemma (ref) it follows that \begin{equation*} \sqrt{m} \left ( \hat \gamma( \cdot ) - \gamma( \cdot ) \right ) \Rightarrow \mathbb{G}( \cdot ) \quad in \quad \ell ^\infty(\mathcal{T}) \end{equation*} where $\mathbb{G}( \cdot ) $ is a zero-mean Gaussian process with uniformly continuous sample path and covariance function $QJ(\tau_1, \tau_2) Q'$ where $Q = \left( \Sigma_{22}' W_{22} \Sigma_{22}\right)^{-1} \Sigma_{22}' W_{22} $. Whereas, the bias will dominate the asymptotic behavior of $\hat \beta$. \end{proof}
commentWe verify the conditions of proposition E.1 in Fernandez-Val2022a. I try to use the same strategy as in their proof of Theorem 5.1. I am not sure that it is easy for GMM. Maybe we have again to take care of the different rates of convergence of the weighting matrix. I start with the exactly identified IV. In this case, \begin{equation*} \hat\delta(\tau)-\delta(\tau)=S_{ZX}^{-1}\bar g_{mn}(\hat\delta,\tau) \end{equation*} From the proof of Lemma (ref), we have $\left\lVert S_{ZX}-\Sigma_{ZX}\right\rVert=O_p\left(\frac{1}{\sqrt{m}}\right)$. From the proof of Lemma (ref)(i) we have \begin{equation*} \bar g_{mn}^{(1)}(\hat\delta,\tau)=\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+o_p\left(\frac{1}{\sqrt{mn}}\right) \end{equation*} uniformly in $\tau\in\mathcal T$. Together with the boundedness of $S_{ZX}$, it implies that \begin{align*} S_{ZX}\bar g_{mn}^{(1)}(\hat\delta,\tau)&=S_{ZX}\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+o_p\left(\frac{1}{\sqrt{mn}}\right) \end{align*} uniformly in $\tau\in\mathcal T$. In addition, Lemma 2(i) implies that $\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})=O_p\left(\frac{1}{\sqrt{mn}}\right)$ uniformly in $\tau\in\mathcal T$. This, in turn, implies that \begin{equation*} \left(S_{ZX}-\Sigma_{ZX}\right)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})=o_p\left(\frac{1}{\sqrt{mn}}\right) \end{equation*} uniformly in $\tau\in\mathcal T$. Combining the last two displayed equation, we obtain \begin{align} S_{ZX}\bar g_{mn}^{(1)}(\hat\delta,\tau)=\Sigma_{ZX}\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+o_p\left(\frac{1}{\sqrt{mn}}\right) \end{align} uniformly in $\tau\in\mathcal T$. Lemma (ref)(ii) implies that $\left\lVert \bar g^{(2)}_{mn}(\hat\delta,\tau)\right\rVert=\left\lVert\frac{1}{mn}\sum_{j=1}^m\sum_{i=1}^n z_{ij}\alpha_j(\tau)\right\rVert=O_p\left(\frac{1}{\sqrt{m}}\right) \left\lVert \Omega_2(\tau)\right\rVert^{1/2}$ uniformly in $\tau\in\mathcal T$. It implies that \begin{equation} S_{ZX}\bar g_{mn}^{(2)}(\hat\delta,\tau)=\Sigma_{ZX}\bar g_{mn}^{(2)}(\hat\delta,\tau)+o_p\left(\frac{1}{\sqrt m}\right)\left\lVert \Omega_2(\tau)\right\rVert^{1/2} \end{equation} Write \begin{equation} \zeta(\tau)=\frac{1}{\sqrt{mn}}+\frac{1}{\sqrt m}\left\lVert\Omega_2(\tau)\right\rVert ^{1/2}. \end{equation} Thus, uniformly in $\tau\in\mathcal T$, \begin{align*} \hat\delta(\tau)-\delta(\tau)=&S_{ZX}\bar g^{(1)}_{mn}(\hat\delta,\tau)+S_{ZX}\bar g^{(2)}_{mn}(\hat\delta,\tau)\\ =&\Sigma_{ZX}\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+o_p\left(\frac{1}{\sqrt{mn}}\right)\\ &+\Sigma_{ZX}\frac{1}{m}\sum_{j=1}^m \bar z_{j}\alpha_j(\tau)+o_p\left(\frac{1}{\sqrt{m}}\right) \left\lVert \Omega_2(\tau)\right\rVert^{1/2}\\ =&\Sigma_{ZX}\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+\Sigma_{ZX}\frac{1}{m}\sum_{j=1}^m \bar z_{j}\alpha_j(\tau)+o_p\left(\zeta(\tau)\right) \end{align*} where the first equality follows equations (ref) and (ref) and the second from the definition $\zeta$ in (ref). Hence $\hat\delta(\tau)-\delta(\tau)$ can be written as in equation (E.1) in Fernandez-Val2022a \begin{equation*} \hat\delta(\tau)-\delta(\tau)=\frac{1}{m}\sum_{j=1}^m\left[\frac{1}{\sqrt n}d_{\psi,j}+d_{\gamma_i}(\tau)\right]+o_p\left(\zeta(\tau)\right) \end{equation*} with $\tau$ replacing $y$ as the argument of the process, our index $j$ replacing their index $i$, our index $i$ replacing their index $t$, and notation \begin{align*} d_{\psi,j}(\tau)&=\Sigma_{ZX}\Sigma_{ZXj}\frac{1}{\sqrt n}\sum_{i=1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\\ d_{\gamma,j}(\tau)&=\Sigma_{ZX}\bar z_{j}\alpha_j(\tau) \end{align*} We now check that Assumption E.1 in Fernandez-Val2022a is satisfied. For part (i), the model in equation (ref) and Assumption (ref)(iv) imply that $E[d_{\psi,j}(\tau)]=0$. Assumption (ref)(ii) implies that $E[d_{\gamma,j}(\tau)]=0$. The arguments in the proof of part (iii) of Lemma (ref) imply that $E[d_{\psi,j}(\tau)d_{\gamma,j}(\tau')]=0$ for $\tau,\tau'\in\mathcal T$. For part (ii), assumptions (ref)-(ref) imply that the eigenvalues of $\operatorname{Var}(d_{\psi,j}(\tau))=\Omega_1(\tau)$ are bounded away from zero and from above. In particular, all the elements on the diagonal of $\Omega_1(\tau)$ are bounded away from zero and from above. All the elements in $\operatorname{Var}(d_{\gamma,j}(\tau))=\Omega_2(\tau)$ are bounded from above by Assumptions (ref) and (ref) but not from below. In particular, we do not exclude that $\Omega_2(\tau)$ is a matrix of zeros. For part (iii), Assumptions (ref)(i) and (ref)(i) imply that $d_{\psi,j}(\tau)$ is uniformly bounded.

Proof of Theorem (ref): Adaptive Asymptotic Distribution

comment$$E[\bar z_j \alpha_j]$$ $$Var(\bar z_j \alpha_j)=E[\bar z_j^2 \alpha_j^2] $$ \begin{align*} \frac{1}{m n^2} \sum_{j = 1}^m Z_{j,l}' X_j (\hat \delta - \delta ) (\hat \delta(\tau') - \delta(\tau') )' X_j' Z_{j,l'}&=\frac{1}{m n^2} \sum_{j = 1}^m\sum_{i=1}^nZ_{ij,l}X_{ij}(\hat \delta(\tau)-\delta(\tau))(\hat\delta(\tau)-\delta(\tau))'X_{ij}'Z_{ij,l}\\ &=\frac{1}{m n^2} \sum_{j = 1}^m\sum_{i=1}^n\sum_{k=1}^K Z_{ij,l}Z_{ij,l'} X_{ij,k}^2 (\hat \delta_k(\tau)-\delta_k(\tau))^2\\ &=\frac{1}{n} \frac{1}{m}\sum_{j = 1}^m\frac{1}{n}\sum_{i=1}^n\sum_{k=1}^K Z_{ij,l}Z_{ij,l'} X_{ij,k}^2 (\hat \delta_k(\tau)-\delta_k(\tau))^2\\ &=\frac{1}{n}O_p\left(\frac{1}{mn}+\frac{Var(\alpha_j(\tau))}{m}\right)\\ &=O_p\left(\frac{1}{mn}\right) \end{align*} where the fourth inequality follows from $\hat\delta(\tau)-\delta(\tau)=O_p\left(\frac{1}{\sqrt{mn}}+\frac{\sqrt{Var(\alpha_j}}{\sqrt m}\right)$ and $X_{ij,k}$ and $Z_{ij,l}$ are bounded. new version - Start with the meat matrix Consider now the term in the middle and insert $\hat u_j(\tau) = \tilde X_j \hat \beta_j(\tau) - X_j \hat \delta (\tau)= \tilde X_j (\hat \beta_j(\tau) - \beta_j(\tau) ) + X_j (\delta(\tau) - \hat \delta(\tau))+ \alpha_j(\tau)$, to obtain, \begin{align*} \frac{1}{m n^2} \sum_{j = 1}^m Z_j' \hat u_j(\tau) \hat u_j(\tau') ' Z_j = & \frac{1}{m n^2} \sum_{j = 1}^m \bigg ( Z_j' \left ( \tilde X_j (\hat \beta_j(\tau) - \beta_j(\tau) ) + X_j (\delta(\tau) - \hat \delta(\tau) )+ \alpha_j(\tau) \right) \\ & \cdot \left ( \tilde X_j (\hat \beta_j(\tau')- \beta_j(\tau') ) + X_j (\delta(\tau') - \hat \delta(\tau') )+ \alpha_j(\tau') \right)' Z_j \bigg ) \\ = & \frac{1}{m n^2} \sum_{j = 1}^m \bigg ( Z_j' \bigg ( \tilde X_j (\hat \beta_j(\tau)- \beta_j(\tau) ) (\hat \beta_j(\tau') - \beta_j(\tau') )' \tilde X_j' + \alpha_j(\tau) \alpha_j(\tau')' \& + X_j (\hat \delta(\tau) - \delta(\tau) ) (\hat \delta(\tau') - \delta(\tau') )' X_j' + X_j (\delta(\tau) - \hat \delta(\tau) ) (\hat \beta_j(\tau') - \beta_j(\tau') )' \tilde X_j' \\ & + \tilde X_j (\hat \beta_j(\tau)- \beta_j(\tau) ) (\delta(\tau') - \hat \delta(\tau') )' X_j' + \alpha_j(\tau) (\hat \beta_j(\tau') - \beta_j(\tau') )' \tilde X_j' \\ & + \tilde X_j (\hat \beta_j(\tau)- \beta_j(\tau) ) \alpha_j(\tau')' + \alpha_j(\tau) (\delta (\tau')- \hat \delta(\tau') )' X_j' \\ & + X_j (\delta(\tau) - \hat \delta(\tau) ) \alpha_j(\tau')' \bigg ) Z_j\bigg ). \end{align*} THE NOTATION HERE IS NOT NICE. ALPHA IS SOMETIMES A SCALAR AND SOMETIMES A VECTOR. Next, we want to show that all but the first two terms converge to zero quickly. We consider each term separately. Let $ \tilde \zeta(k,\tau)= \sqrt{m} \cdot \zeta(k,\tau)$ and $ \zeta(2,\tau)=\frac{1}{\sqrt{mn}}+\frac{1}{\sqrt m}\left\lVert \Omega_{2(L_2 \times L_2)}\right\rVert ^{1/2}$, where the first $M_1$ elements of $\hat \delta$ converge at the $ \zeta (1, \tau) $ and the last $M_2$ elements converge at the $ \zeta (2, \tau) $ rate. Further, recall that $(\hat \beta_j- \beta_j ) = O_p\left ( {n^{-1/2}} \right) $. Consider the first term. By the proof of Lemma (ref)(i), it follows that \begin{align*} \frac{1}{m n^2} \sum_{j = 1}^m Z_j' \tilde X_j (\hat \beta_j(\tau)- \beta_j(\tau) ) & (\hat \beta_j(\tau') - \beta_j(\tau') )' \tilde X_j' \\ = & \frac{1}{m} \sum_{j = 1}^m \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'(\hat \beta_j(\tau') -\beta_j(\tau') ) \right )' \\ = & \mathbb{E} \left [ \left ( \Sigma_{ZXj} \frac{1}{n} \sum_{i = 1}^n \phi_{i,\tau}(\tilde x_{ij}, z_{ij}) \right) \left ( \Sigma_{ZXj} \frac{1}{n} \sum_{i = 1}^n \phi_{i,\tau'}(\tilde x_{ij}, z_{ij}) \right )' \right] \\ &+ o_p \left ( (mn)^{-1} \right ) \\ =& \frac{\Omega_1(\tau, \tau')}{n} + o_p \left ( (mn)^{-1} \right ) . \end{align*} For the second term, we consider each element separately. For some $s$ and $k$ we have that \begin{align} \frac{1}{m} \sum_{j = 1}^m \bar z_{jk} \bar z_{js} ' \alpha_j(\tau)\alpha_j(\tau') = \mathbb{E} [\bar z_{jk} \bar z_{js} \alpha_j(\tau) \alpha_j(\tau')] + O_p \left ( \frac{1}{\sqrt{m}} \right ) || \mathbb{E} [\bar z_{jk} \bar z_{js} \alpha_j(\tau) \alpha_j(\tau') ]|| \end{align} where the first equality follows because, \begin{align*} \sqrt{m}\frac{1}{m}\sum_{j = 1}^m \left [ \mathbb{E} [\bar z_{jk} \bar z_{js} \alpha_j(\tau) \alpha_j(\tau') ]^{-1} \bar z_{jk} \bar z_{zs} ' \alpha_j(\tau)\alpha_j(\tau') -1 \right]= O_p(1), \end{align*} which implies that $\frac{1}{m}\sum_{j = 1}^m \bar z_{jk} \bar z_{zs} ' \alpha_j(\tau)\alpha_j(\tau') -\mathbb{E}[\bar z_{jk} \bar z_{js} \alpha_j(\tau) \alpha_j(\tau')] = O_p \left(\frac{1}{\sqrt{m}} \right) |\mathbb{E}[\bar z_{jk} \bar z_{js} \alpha_j(\tau) \alpha_j(\tau')]| = O_p \left ( \frac{1}{\sqrt{m}} \right )|\Omega_{2, ks}|.$ Thus, different entries of this matrix might converge at a different rate. {\color{red} At the moment, I split the first $M_1$ elements from the remaining $M_2$. We don't want to do that. }For the fast instrument, this matrix is numerically zero. For instrument of the second type, we can write $\frac{1}{m} \sum_{j = 1}^m \bar z_{jk} \bar z_{js}' \alpha_j(\tau)\alpha_j(\tau') = \mathbb{E} [\bar z_{jk} \bar z_{js} \alpha_j(\tau) \alpha_j(\tau')] + o_p\left( \zeta(2,\tau) \right)$ Consider now the third term. We have that \begin{align*} \frac{1}{m n^2} \sum_{j = 1}^m Z_j' X_j (\hat \delta - \delta ) (\hat \delta(\tau') - \delta(\tau') )' X_j' Z_j = & \frac{1}{m} \sum_{j = 1}^m \left( \frac{1}{n} \sum_{i = 1}^n z_{ij} x_{ij}' \right) (\hat \delta(\tau) - \delta(\tau) ) (\hat \delta(\tau') - \delta(\tau') )' \left( \frac{1}{n} \sum_{i = 1}^n x_{ij} z_{ij}' \right) \\ =& \begin{pmatrix} o_p(\zeta(1, \tau)) & o_p(\zeta(1, \tau)) \\ o_p(\zeta(1, \tau)) & o_p(\zeta(2, \tau)) \end{pmatrix} \end{align*} where we exploit the fact that $\frac{1}{n} \sum_{i = 1}^n z_{ij} x_{ij}'$ is lower triangular. [{\color{red} Blaise's sandbox} Consider the element $l,l'$ of the matrix: \begin{align*}\frac{1}{m n^2} \sum_{j = 1}^m Z_{j,l}' X_j (\hat \delta - \delta ) (\hat \delta(\tau') - \delta(\tau') )' X_j' Z_{j,l'}&=\frac{1}{m n^2} \sum_{j = 1}^m\sum_{i=1}^nZ_{ij,l}X_{ij}(\hat \delta(\tau)-\delta(\tau))(\hat\delta(\tau)-\delta(\tau))'X_{ij}'Z_{ij,l}\\ &=\frac{1}{m n^2} \sum_{j = 1}^m\sum_{i=1}^n\sum_{k=1}^K Z_{ij,l}Z_{ij,l'} X_{ij,k}^2 (\hat \delta_k(\tau)-\delta_k(\tau))^2\\ &=\frac{1}{n} \frac{1}{m}\sum_{j = 1}^m\frac{1}{n}\sum_{i=1}^n\sum_{k=1}^K Z_{ij,l}Z_{ij,l'} X_{ij,k}^2 (\hat \delta_k(\tau)-\delta_k(\tau))^2\\ &=\frac{1}{n}O_p\left(\frac{1}{mn}+\frac{Var(\alpha_j(\tau))}{m}\right)\\ &=O_p\left(\frac{1}{mn}\right) \end{align*} where the fourth inequality follows from $\hat\delta(\tau)-\delta(\tau)=O_p\left(\frac{1}{\sqrt{mn}}+\frac{\sqrt{Var(\alpha_j}}{\sqrt m}\right)$ and $X_{ij,k}$ and $Z_{ij,l}$ are bounded. I think that the trick consists in splitting $$O_p(\Omega_{2,ll})=O_p(\operatorname{Var}(\alpha_j))O_p(\operatorname{Var}(\bar z_{j,l})$$ We see that the rate of convergence of a particular element of $\Omega_{2,ll}$ depends on a component that is common to all regressors $\operatorname{Var}(\alpha_j)$ and a component specific to the instrument $\operatorname{Var}(\bar z_{j,l})$. Then, because the constant is part of $z$ and $x$ and it has variance 1, the best possible rate for the whole vector $\hat\delta-\delta$ is $O_p\left(\frac{1}{\sqrt{mn}}\right)+O_p\left(\frac{1}{\sqrt{m}}\right)\sqrt{\operatorname{Var}(\alpha_j)}$. It follows that $$ X_j (\hat \delta - \delta )=o_p\left(\frac{1}{\sqrt{mn}}\right)+O_p\left(\frac{1}{\sqrt{m}}\right)\sqrt{\operatorname{Var}(\alpha_j)}$$ and $$ X_j (\hat \delta - \delta ) (\hat \delta(\tau') - \delta(\tau') )' X_j'=O_p\left(\frac{1}{mn}\right)+O_p\left(\frac{1}{m}\right)\operatorname{Var}(\alpha_j)$$ Then it follows that \begin{align*}\frac{1}{m n^2} \sum_{j = 1}^m Z_{j,l}' X_j (\hat \delta - \delta ) (\hat \delta(\tau') - \delta(\tau') )' X_j' Z_{j,l}&=O_p\left(\frac{1}{mn}\right)+O_p\left(\frac{1}{m}\right)\operatorname{Var}(\alpha_j)\operatorname{Var}(\bar z_{j,l})\\ &=O_p\left(\frac{1}{mn}\right)+O_p\left(\frac{1}{m}\right)\Omega_{2,ll} \end{align*} ] {\color{red} Martina's Sandbox} Take the $n \times n$ matrix $$ X_j (\hat \delta - \delta ) (\hat \delta(\tau') - \delta(\tau') )' X_j'=O_p\left(\frac{1}{mn}\right)+O_p\left(\frac{1}{m}\right)\operatorname{Var}(\alpha_j)$$ Multiply by $1/n \cdot Z_{j,l}$ this must be wrong. $$ \frac{1}{n} Z_{j,l}' X_j (\hat \delta - \delta ) (\hat \delta(\tau') - \delta(\tau') )' X_j' \frac{1}{n} Z_{j,l}=\bar z_{j,l} \left ( O_p\left(\frac{1}{mn}\right)+O_p\left(\frac{1}{m}\right)\operatorname{Var}(\alpha_j) \right ) \bar z_{j,l} $$ The fifth term is just the transpose of the fourth term. Thus we will consider only the fourth one. $$ X_j (\hat \delta - \delta ) (\hat \beta_j - \beta_j) \tilde X_j=O_p\left(\frac{1}{\sqrt{m}n}\right)+O_p\left(\frac{1}{\sqrt{mn}}\right)\sqrt{\operatorname{Var}(\alpha_j)}$$ $$\frac{1}{mn^2} \sum_{j = 1}^m Z_{j,l} X_j (\hat \delta - \delta ) (\hat \beta_j - \beta_j) \tilde X_j Z_{j,l}=O_p\left(\frac{1}{mn}\right)+O_p\left(\frac{1}{m \sqrt{n}}\right)\sqrt{\operatorname{Var}(\alpha_j)} \operatorname{Var}(\bar z_{j,l})$$ \begin{align*} \frac{1}{m n^2} \sum_{j = 1}^m Z_j' X_j & (\delta(\tau) - \hat \delta(\tau) ) (\hat \beta_j(\tau') - \beta_j(\tau') )' \tilde X_j' Z_j \\ =& \frac{1}{m} \sum_{j = 1}^m \left( \frac{1}{n} \sum_{i = 1}^n z_{ij} x_{ij}' \right) (\delta(\tau) - \hat \delta(\tau) ) (\hat \beta_j(\tau') - \beta_j(\tau') )' \left( \frac{1}{n} \sum_{i = 1}^n \tilde x_{ij} z_{ij}' \right) \%= & O_p \left ( \zeta \cdot n^{-1/2} \right ) \\ =& \begin{pmatrix} o_p(\zeta(1, \tau)) & o_p(\zeta(1, \tau)) \\ o_p(\zeta(2, \tau)) & o_p(\zeta(2, \tau)) \end{pmatrix}. \end{align*} For the sixth (and seventh) term(s), we have that \begin{align*} \frac{1}{m n^2} \sum_{j = 1}^m Z_{j,l}' \alpha_j(\tau) (\hat \beta_j(\tau') - \beta_j(\tau') )' \tilde X_j' Z_{j,l} = & \frac{1}{m} \sum_{j = 1}^m \bar z_{j,l} \alpha_j(\tau) (\hat \beta_j(\tau') - \beta_j(\tau) )' \left( \frac{1}{n} \sum_{i = 1}^n \tilde x_{ij} z_{ij,l}' \right) \\ = & O_p\left ( \frac{1}{\sqrt{mn}} \right ) \operatorname{Var}(\bar z_{i,j} \alpha_i) ^{1/2} \end{align*} Finally, for the eighth (and ninth) term(s), it follows that \begin{align*} \frac{1}{m n^2} \sum_{j = 1}^m Z_j' \alpha_j(\tau) (\delta(\tau') - \hat \delta(\tau') )' X_j' Z_j = & \frac{1}{m } \sum_{j = 1}^m \bar z_j \alpha_j(\tau) (\delta(\tau')- \hat \delta(\tau') )'\left( \frac{1}{n} \sum_{i = 1}^n x_{ij} z_{ij}' \right) \\= & = \left ( O_p \left ( \frac{1}{m \sqrt{n}} \right) + O_p \left (\frac{1}{{m}} \right) \operatorname{Var}(\alpha_j)^{1/2} \right ) \operatorname{Var}(\bar z_{j,l} \alpha_j) ^{1/2} \end{align*} Taking all terms into account, including their transpose we find that \begin{align*} \hat \Omega(\tau, \tau') = \frac{\Omega_1(\tau, \tau')}{n} + \Omega_2(\tau, \tau') + \begin{pmatrix} o_p(\zeta(1, \tau)) & o_p(\zeta(2, \tau)) \\ o_p(\zeta(2, \tau)) & o_p(\zeta(2, \tau)) \end{pmatrix} \end{align*} Covariance Matrix Here I have to consider all elements separately and take into account that come elements converge to zero. \begin{align*} \hat G(\tau) \hat \Omega(\tau, \tau') \hat G(\tau') = G(\tau) \Omega(\tau, \tau') G(\tau') + \begin{pmatrix} TBD \end{pmatrix} \end{align*} [The top left must be at least $o_p(1/\sqrt{n})$ and the entire matrix should be $o_p(|| \Omega_2 ||^{1/2})$] \begin{proof}[Proof of Theorem (ref)] Let $\hat G (\tau) = \left( S_{ZX}' \hat W(\tau) S_{ZX} \right) ^{-1} S_{ZX}' \hat W(\tau) $. Starting from the error representation (Lemma (ref)) we have that \begin{align*} \hat \delta(\tau) - \delta(\tau) = & \hat G(\tau) \frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\left( \tilde x_{ij}'\left(\hat \beta_j(\tau)-\beta_j(\tau)\right)+\alpha_j(\tau)\right) .\\ \end{align*} For the first term in the parenthesis, equation ((ref)) implies that $\frac{1}{mn} \sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\tilde x_{ij}' \left(\hat \beta_j(\tau)-\beta_j(\tau)\right) = \frac{1}{N \sqrt{n}}\sum_{j = 1}^ms_j(\tau) + o_p\left(\frac{1}{\sqrt{mn}}\right)$. Therefore, using $ \hat G(\tau) = G (\tau) + o_p(1)$, we obtain {\color{red} we need some condition on $\frac{1}{N \sqrt{n}}\sum_{j = 1}^ms_j(\tau)$. e.g. $\frac{1}{N \sqrt{n}}\sum_{j = 1}^ms_j(\tau) = O(1)$. } From the proof of Lemma (ref), we know that $\frac{1}{N \sqrt T} \sum_{j = 1}^m s_j$ is Donsker and, therefore, also totally bounded. \begin{align*} \hat G(\tau) \frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^n z_{ij}\left( \tilde x_{ij}'\left(\hat \beta_j(\tau)-\beta_j(\tau)\right)\right) = G \frac{1}{N \sqrt{n}}\sum_{j = 1}^ms_j(\tau) + o_p\left(\frac{1}{\sqrt{mn}}\right) \end{align*} For the second term, we have that $\frac{1}{ N}\sum_{j = 1}^m \bar z_j \alpha_j(\tau) = O_p \left ( \frac{ ||\operatorname{Var}(\bar z_j \alpha_j(\tau) ) ||^{1/2} }{\sqrt{m}} \right )$. { \color{red} They have the variance of the entire term with G and $\eta$ in it ($\operatorname{Var}(\eta G \bar z_j \alpha_j)$). I think that it doesn't matter, since the rate of convergence is the same (what about the weight matrix?). But they have little Oh!! They have something like $ || \frac{1}{ N}\sum_{j = 1}^m \bar z_j \alpha_j(\tau) || = o_p \left ( \frac{1}{\sqrt{m}} \right ) \cdot || V_\alpha(\tau) ||^{1/2} $ } Thus, $$ \hat G \frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^n z_{ij} \alpha_j(\tau) = G \frac{1}{mn}\sum_{j = 1}^m\sum_{i = 1}^n z_{ij} \alpha_j(\tau) + O_p \left ( \frac{1}{\sqrt{m}} \right ) \cdot || \operatorname{Var}(\bar z_j \alpha_j(\tau) ) ||^{1/2} $$ Combining yields, {\color{red} For the second term I have big Oh! } \begin{align*} \hat \delta(\tau) - \delta(\tau) = & G(\tau) \cdot \frac{1}{ N}\sum_{j = 1}^m \left [ \frac{1}{\sqrt{n}} s _j(\tau) + \bar z_j \alpha_j(\tau) \right ] + o_p \left (\zeta(\tau) \right ) \end{align*} where $\zeta (\tau) = \frac{1}{\sqrt{mn}} + \frac{1}{\sqrt{m}} \cdot || V_\alpha(\tau) ||^{1/2}$ {\color{red} Here I need to check the definition of $V_\alpha$ because above I use the variance of the moment, below it is the variance of the estimator coming from the second stage noise. } So uniformly in $\tau$, for any $\eta \neq 0$, \begin{align*} \eta' \left ( \hat \delta(\tau) - \delta(\tau) \right) = & \frac{1}{ N}\sum_{j = 1}^m \left ( \eta' \cdot G(\tau) \cdot \left [ \frac{1}{\sqrt{n}} s _j(\tau) + \bar z_j \alpha_j(\tau) \right ] + o_p \left (\zeta(\tau) \right ) \right) \\ = & \frac{1}{ N}\sum_{j = 1}^m \left ( \eta' \cdot \left [ \frac{1}{\sqrt{n}} a_{\beta,j}(\tau) + a_{\alpha,j}(\tau) \right ] + o_p \left (\zeta(\tau) \right ) \right) \\ = & \frac{1}{ N}\sum_{j = 1}^m \left ( \frac{1}{\sqrt{n}} d_{\beta,j}(\tau) + d_{\alpha,j}(\tau) + o_p \left (\zeta(\tau) \right ) \right) \end{align*} where \begin{align*} d_{\beta,j} =& \eta ' G(\tau) s_j(\tau) = \eta ' G(\tau) \Sigma_{ZXj} \frac{1}{\sqrt T } \sum_{i = 1}^n \phi_{i , \tau} \\ d_{\alpha,j} =& \eta ' G(\tau) \bar z_j \alpha_j(\tau) \\ V_{\alpha}(\tau, \tau') =& \frac{1}{m} \sum_{j = 1}^m \mathbb{E}[ d_{\alpha,j} d_{\alpha,j} ] \\ V_{\alpha}(\tau, \tau) =& V_{\alpha}(\tau) = \frac{1}{m} \sum_{j = 1}^m \mathbb{E}[ d_{\alpha,j} d_{\alpha,j} ] = \eta' G(\tau) \mathbb{E} \left [ \bar z_j \bar z_j ' \alpha_j(\tau)^2 \right ] G(\tau) ' \eta \\ V_{\beta,j}(\tau, \tau') =& \frac{1}{m} \sum_{j = 1}^m \cdot \mathbb{E}[ d_{\beta,j} d_{\beta,j} ] \\ V_{\beta,j}(\tau, \tau) =& V_{\beta,j}(\tau) = \eta' G(\tau) \mathbb{E} \left [ \Sigma_{ZXj}\operatorname{Var}(\phi_{i,\tau})\Sigma_{ZXj}'\right ] G(\tau)' \eta \\ \sigma_n^2(\tau, \tau') =& \frac{1}{n} V_{\beta,j}(\tau,\tau') + V_{\alpha,j}(\tau,\tau') \\ \sigma_n^2(\tau) =& \frac{1}{n} V_{\beta,j}(\tau) + V_{\alpha,j}(\tau) \\ \sigma_n^2 (\tau) =& \sigma_n(\tau, \tau) \\ s_{mn}^2 (\tau) = & \frac{1}{m} \sigma^2_n \\ H = & \lim_{N \rightarrow \infty} \frac{\sigma_n^2(\tau, \tau')}{\sigma_n (\tau) \sigma_n (\tau')} \end{align*} To apply Proposition E.1 in Fernandez-Val2022a we need to check that their assumption $E.1$ is satisfied. By the proof of (ref) we have that $\mathbb{E} \left[ d_{\beta,j}(\tau) \right ] = 0$, $\mathbb{E} \left[ d_{\alpha,j}(\tau) \right ] = 0$, and $\mathbb{E} \left[ d_{\beta,j}d_{\alpha,j}(\tau, \tau') \right ] = 0$ for all $\tau, \tau'$. $ V_\alpha(\tau) \geq || \eta ||^2 \lambda_{min} \left( G(\tau) \mathbb{E} \left [ \bar z_j \bar z_j ' \alpha_j(\tau)^2 \right ] G(\tau)\right )$ $$| d_{\beta,j}(\tau) | \leq C || \Sigma_{ZXj} || \ || \frac{1}{\sqrt T } \sum_{i = 1}^n \phi_{i , \tau} || $$ $$| d_{\alpha,j}(\tau) | \leq C || \bar z_j \alpha_j (\tau) || $$ $$ \mathbb{E} \sup_\tau | d_{\beta,1}(\tau) | ^{2+ a} \leq C \sqrt{\mathbb{E} ||\Sigma_{ZXj} ||^{4 + 2a} \mathbb{E} || \frac{1}{\sqrt T} \sum_{i = 1}^n \phi_{i,\tau} (\tilde x_{ij}, z_{ij} ) || ^{4 + 2a} } {\color{red} \text{I need an upper bound} }$$ $$ \mathbb{E} \sup_\tau | \frac{d_{\beta,i(\tau) ^2 }}{V_\beta(\tau) } |^a \leq C \mathbb{E} \sup_\tau \left [ \frac{|| s_j (\tau) ||^2 }{ \lambda_{min} \left( G(\tau) \mathbb{E} \left [ \bar z_j \bar z_j ' \alpha_j(\tau)^2 \right ] G(\tau)\right ) } \right ]^a {\color{red} \text{I need an upper bound} }$$ {\color{red} I don't get how their results show that condition ii is satisfied. They show that both these terms are smaller than C. But they don't show that the sum is smaller than C. I am not sure what C is in this case. I could be either some positive constant or the upper bound on the variance. } If the conditions of assumption E.1 in Fernandez-Val2022a are satisfied we can apply their proposition E.1. Then, it follows directly that \begin{equation} \frac{\eta ' \left (\hat \delta(\cdot) - V_\delta(\cdot) \right ) }{\left [ \eta ' V_\delta(\cdot) \eta \right ]^{1/2}} \rightsquigarrow \mathbb G (\cdot) \end{equation} \end{proof}
proof[Proof of Theorem (ref)] We prove the theorem when Assumption (ref) holds. The proof when instead Assumption (ref) holds is simpler because all coefficients converge at the same rate; it can actually be considered as a special case of the proof below when we set $a_n(\tau)=1$. By Lemma (ref), uniformly in $\tau\in\mathcal T$, \begin{equation*} \hat G(\tau)= G_{mn}(\tau)+ \begin{pmatrix} o_p(1) & o_p \left (\sqrt{a_n(\tau)}\right ) \\ o_p \left (1/\sqrt{a_n(\tau)} \right ) & o_p(1) \end{pmatrix} \end{equation*} where $G_{mn,12}(\tau)=O_p(a_n(\tau))$ and the other elements are $O_p(1)$. Since the rate of convergence may differ across the quantile index and the regressors, we first consider the scalar $\hat\delta_k(\tau)$, which is the $k$-th element of $\hat\delta(\tau)$ for $k\in\{1,\dots,K\}$ and $\tau \in\mathcal{T}$. From Lemma (ref), \begin{equation*} \hat\delta_k(\tau)-\delta_k(\tau)=\hat G_k(\tau)\bar g_{mn}(\hat\delta,\tau)=\hat G_k(\tau)\left(\bar g_{mn}^{(1)}(\hat\delta,\tau)+\bar g_{mn}^{(2)}(\hat\delta,\tau)\right) \end{equation*} where $\hat G_k(\tau)$ is the $k$-th row of $\hat G(\tau)$. From the proof of Lemma (ref)(i), we have \begin{equation*} \bar g_{mn}^{(1)}(\hat\delta,\tau)=\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+o_p\left(\frac{1}{\sqrt{mn}}\right), \end{equation*} uniformly in $\tau\in\mathcal T$. We deal separately with the fast and slow coefficients and show that the result holds for both types of coefficients. For $k\in\{1,\dots,M_1\}$ (fast coefficients), $\hat G_k(\tau)$ is uniformly bounded in probability, which implies that \begin{align*} \hat G_k(\tau)\bar g_{mn}^{(1)}(\hat\delta,\tau)&=\hat G_k(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+o_p\left(\frac{1}{\sqrt{mn}}\right), \end{align*} uniformly in $\tau\in\mathcal T$. In addition, (ref)(i) implies that $\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})=O_p\left(\frac{1}{\sqrt{mn}}\right)$ uniformly in $\tau\in\mathcal T$. For these coefficients, $\underset{\tau\in\mathcal{T}}{\sup}\left\lVert \hat G_k(\tau)-G_{mn,k}(\tau)\right\rVert=o_p\left(1\right)$. This, in turn, implies that \begin{equation*} \left(\hat G_k(\tau)-G_{mn,k}(\tau)\right)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})=o_p\left(\frac{1}{\sqrt{mn}}\right), \end{equation*} uniformly in $\tau\in\mathcal T$. Combining the last two displayed equations, we obtain \begin{align} \hat G_k(\tau)\bar g_{mn}^{(1)}(\hat\delta,\tau)=G_{mn,k}(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+o_p\left(\frac{1}{\sqrt{mn}}\right), \end{align} uniformly in $\tau\in\mathcal T$. For $k\in\{M_1+1,\dots,K\}$ (slow coefficients), by Lemma (ref), uniformly in $\tau$, $\hat G_k(\tau)-G_{mn,k}(\tau)=o_p(1/\sqrt{a_n(\tau)})$ and $G_{mn,k}=O_p(1)$ such that $\hat G_{k}=O_p(1/\sqrt{a_{n}(\tau)})$. In addition, Lemma (ref)(i) implies that $$\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj} \frac{1}{n} \sum_{i = 1}^n \phi_{j,\tau}(\tilde x_{ij},y_{ij})=O_p\left(\frac{1}{\sqrt{mn}}\right).$$ It follows that \begin{align*} \hat G_k(\tau)\bar g_{mn}^{(1)}(\hat\delta,\tau)&=\hat G_k(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+o_p\left(\frac{1}{\sqrt{mna_n(\tau)}}\right), \end{align*} and \begin{equation*} \left(\hat G_k(\tau)-G_{mn,k}(\tau)\right)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})=o_p\left(\frac{1}{\sqrt{mna_n(\tau)}}\right), \end{equation*} uniformly in $\tau\in\mathcal T$. Combining the last two displayed equations, we obtain \begin{align*} \hat G_k(\tau)\bar g_{mn}^{(1)}(\hat\delta,\tau)=G_{mn,k}(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+o_p\left(\frac{1}{\sqrt{mna_n(\tau)}}\right), \end{align*} uniformly in $\tau\in\mathcal T$. Lemma (ref)(ii) implies that \begin{equation}\bar g_{mn,l}^{(2)}(\delta,\tau)=\frac{1}{m}\sum_{j = 1}^m \bar z_{j,l}\alpha_j(\tau)=O_p\left(\frac{1}{\sqrt m}\right)\sqrt{\Omega_{2,ll}(\tau)}\end{equation} uniformly in $\tau\in\mathcal T$, where $g_{mn,l}^{(2)}(\delta,\tau)$ is the $l$-th element of $g_{mn}^{(2)}(\delta,\tau)$, $z_{j,l}$ is the $l$-th element of $z_{j}$, and $\Omega_{2,ll}(\tau)$ it the element in the $l$-th row and $l$-th column of $\Omega_{2}(\tau)$. Let $\hat G_{kl}(\tau)$ and $G_{mn,kl}(\tau)$ be the element in the $k$-th row and $l$-th column of $\hat G(\tau)$ and $G_{mn}(\tau)$, respectively. Note that $\hat G_{kl}(\tau)-G_{mn,kl}(\tau)=o_p(\sqrt{G_{mn,kl}(\tau)})$ for $l\in\{L_1+1,\dots,L\}$ and $k\in\{1,\dots,K\}$. Thus, we have \begin{align} (\hat G_k(\tau)-G_{mn}(\tau)) \bar g_{mn}^{(2)}(\hat\delta,\tau)&=\sum_{l=L_1+1}^{L}(\hat G_{kl}(\tau)-G_{mn,kl}(\tau))\bar g_{mn,l}^{(2)}(\hat \delta,\tau)\nonumber\\ &=\sum_{l=L_1+1}^{L}o_p \left (\sqrt{G_{mn,kl}} \right )O_p\left(\frac{1}{\sqrt m}\right)\sqrt{\Omega_{2,ll}(\tau)}\nonumber\\ &=o_p\left(\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt{m}}\right)\sum_{l=L_1+1}^{L} \sqrt{G_{mn,kl}(\tau))},\nonumber \end{align} uniformly in $\tau\in\mathcal T$. If $k\in\{1,\dots,M_1\}$, define \begin{equation} \zeta(k,\tau)=\frac{1}{\sqrt{mn}}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt m}\sum_{l=L_1 + 1}^L\sqrt{G_{mn,kl}(\tau)}. \end{equation} If $k\in\{M_1+1,\dots,K\}$, define \begin{equation} \zeta(k,\tau)=\frac{1}{\sqrt{mna_n(\tau)}}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt m}\sum_{l=L_1+1}^L\sqrt{G_{mn,kl}(\tau)}. \end{equation} Thus, uniformly in $\tau\in\mathcal T$, \begin{align*} \hat\delta_k(\tau)-\delta_k(\tau)=&\hat G_k(\tau)\bar g^{(1)}_{mn}(\hat\delta,\tau)+\hat G_k(\tau)\bar g^{(2)}_{mn}(\hat\delta,\tau)\\ =&G_{mn,k}(\tau)\left(\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+\frac{1}{m}\sum_{j=1}^m \bar z_{j}\alpha_j(\tau)\right)+o_p\left(\zeta(k,\tau)\right) \end{align*} Hence, we have \begin{equation*} \hat\delta_k(\tau)-\delta_k(\tau)=\sum_{j=1}^m d_j(k,\tau)+o_p\left(\zeta(k,\tau)\right), \end{equation*} where \begin{equation*} d_j(k,\tau)=G_{mn,k}(\tau)\left(\frac{1}{m}\Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+\frac{1}{m}\bar z_{j}\alpha_j(\tau)\right). \end{equation*} Let $D_j$ be the $TK\times 1$ vector $(\operatorname{diag}(\Sigma_{mn}(\tau_1))^{-1/2} d_j(\tau_1),\dots, \operatorname{diag}(\Sigma_{mn}(\tau_T))^{-1/2} d_j(\tau_T))'$ where $d_j(\tau)=(d_j(1,\tau),d_j(2,\tau),\dots,d_j(K,\tau))'$. It follows that \begin{equation*} \operatorname{Var}\left(\sum_{j=1}^m D_j\right)=H_{mn} \end{equation*} Then, from the proof of Lemma (ref) and Assumption (ref), \begin{equation*} H_{mn}^{-1/2}\sum_{j=1}^mD_j\underset{d}{\rightarrow}N(0,I_{TK}). \end{equation*} By Slutsky's theorem, \begin{equation*} H^{-1/2}\sum_{j=1}^mD_j=H_{mn}^{-1/2}\sum_{j=1}^mD_j+o_p(1)\underset{d}{\rightarrow}N(0,I_{TK}) \end{equation*} In the proof of Lemma (ref) we show that $\sum_j \operatorname{diag}(\Sigma_{mn}(\cdot))^{-1/2} d_j(\cdot))$ is asymptotically tight in $\ell^\infty(\mathcal{T})$. It follows that the process $\sum_j \operatorname{diag}(\Sigma_{mn}(\cdot))^{-1/2} d_j(\cdot))$ weakly converges to $\mathbb Z$, a centered Gaussian process with covariance kernel $H(\tau,\tau')$. Finally, we show that $o_p(1)\underset{\tau\in\mathcal T, k\in\{1,\dots,K\}}{\sup}\zeta(k,\tau)\Sigma_{mn,k}(\tau)^{-1/2}=o_p(1)$, where $\Sigma_{mn,k}(\tau)$ is the $(k,k)$ element of $\Sigma_{mn}(\tau)$. We consider separately the fast and slow coefficients For $k\in\{1,\dots,M_1\}$ (fast coefficients), note that $$\Sigma_{mn,k}(\tau)^{1/2}=O_p\left(\frac{1}{\sqrt{mn}} + \frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{n\sqrt{m}}\right)=O_p\left(\frac{1}{\sqrt{mn}}\right),$$ and \begin{align*} \zeta(k,\tau)&=O_p\left(\frac{1}{\sqrt{mn}}+\frac{\sqrt{\operatorname{Var}(\alpha(\tau))}}{\sqrt m}\sum_{l=L_1}^L\sqrt{G_{mn,kl}(\tau)}\right)\\ &=O_p\left(\frac{1}{\sqrt{mn}}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt{m}}\frac{1}{\sqrt{1+\operatorname{Var}(\alpha_j(\tau))n}}\right)\\ &=O_p\left(\frac{1}{\sqrt{mn}}\right), \end{align*} both uniformly in $\tau$. It follows that $o_p(1) \ \underset{\tau\in\mathcal T}{\sup}\ \zeta(k,\tau)\Sigma_{mn,k}(\tau)^{-1/2}=o_p(1)$. For $k\in\{M_1+1,\dots,K\}$ (slow coefficients), note that $$\Sigma_{mn,k}(\tau)^{1/2}=O_p\left(\frac{1}{\sqrt{mn}} + \frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt{m}}\right),$$ and $$\zeta(k,\tau)=\frac{1}{\sqrt{mna_n(\tau)}}+\frac{\sqrt{\operatorname{Var}(\alpha(\tau))}}{\sqrt m}\sum_{l=L_1}^L\sqrt{G_{mn,kl}(\tau)}.$$ For the first term of $\zeta(k,\tau)$, we obtain \begin{align*} \frac{1}{\sqrt{mna_n(\tau)}}\frac{1}{\Sigma_{mn,k}(\tau)^{1/2}}&=O_p\left(\frac{1}{\sqrt{mna_n(\tau)}}\frac{1}{\frac{1}{\sqrt{mn}}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt{m}}}\right)\\ &=O_p\left(\frac{1}{\sqrt{a_n(\tau)}}\frac{1}{1+\sqrt{\operatorname{Var}(\alpha_j(\tau))n}}\right)\\ &=O_p\left(\frac{\sqrt{1+\operatorname{Var}(\alpha_j(\tau))n}}{1+\sqrt{\operatorname{Var}(\alpha_j(\tau))n}}\right)\\ &=O_p(1). \end{align*} \begin{comment} It follows that \begin{align*} \zeta(k,\tau)\Sigma_{mn,k}(\tau)^{-1/2}&= \left(\frac{1}{\sqrt{mna_n(\tau)}}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt m}\right)\frac{1}{\frac{1}{\sqrt{mn}}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt{m}}}\\ &=\frac{1+\sqrt{\operatorname{Var}(\alpha_j(\tau)na_n(\tau)}}{\sqrt{mna_n(\tau)}}\frac{\sqrt{mn}}{1+\sqrt{\operatorname{Var}(\alpha_j(\tau))n}}\\ &=\frac{1+\sqrt{\operatorname{Var}(\alpha_j(\tau)na_n(\tau)}}{\sqrt{a_n(\tau)}\left(1+\sqrt{\operatorname{Var}(\alpha_j(\tau))n}\right)}\\ &=\frac{1+\sqrt{\operatorname{Var}(\alpha_j(\tau)n\frac{1}{1+\operatorname{Var}(\alpha_j(\tau))n}}}{\sqrt{\frac{1}{1+\operatorname{Var}(\alpha_j(\tau))n}}\left(1+\sqrt{\operatorname{Var}(\alpha_j(\tau))n}\right)}\\ &=\frac{\sqrt{\frac{1}{1+\operatorname{Var}(\alpha_j(\tau)n}}\left(\sqrt{1+\operatorname{Var}(\alpha_j(\tau)n}+\sqrt{\operatorname{Var}(\alpha_j(\tau))n}\right)}{\sqrt{\frac{1}{1+\operatorname{Var}(\alpha_j(\tau))n}}\left(1+\sqrt{\operatorname{Var}(\alpha_j(\tau))n}\right)}\\ &=\frac{\sqrt{1+\operatorname{Var}(\alpha_j(\tau))n}}{1+\sqrt{\operatorname{Var}(\alpha_j(\tau))n}}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))n}}{1+\sqrt{\operatorname{Var}(\alpha_j(\tau))n}}\\ =O_p(1) \end{align*} \end{comment} For the second term of $\zeta(k,\tau)$, we obtain \begin{align*} \frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt m}\sum_{l=L_1}^L\sqrt{G_{mn,kl}(\tau)}\Sigma_{mn,k(\tau)}^{-1/2}&=O_p\left(\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt m}\frac{1}{\sqrt{mn}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt m}}\right)\\ &=O_p(1). \end{align*} It follows that, also in this second case, $o_p(1)\ \underset{\tau\in\mathcal T} {\sup}\ \zeta(k,\tau)\Sigma_{mn,k}(\tau)^{-1/2}=o_p(1)$. Hence, uniformly in $\tau$, \begin{align*} \operatorname{diag}(\Sigma_{mn}(\tau))^{-1/2}(\hat\delta(\tau)-\delta(\tau))&=\operatorname{diag}(\Sigma_{mn}(\tau))^{-1/2}\left(\sum_{j=1}^m d_j(\tau)+o_p\left(\zeta(\tau)\right)\right)\\ &=\operatorname{diag}(\Sigma_{mn}(\tau))^{-1/2} \sum_{j=1}^m d_j(\tau)+o_p(1)\\ &\rightsquigarrow \mathbb G(\tau), \end{align*} where $\zeta(\tau)=(\zeta(1,\tau),\dots,\zeta(K,\tau))'$. \begin{comment}- Since the rate of convergence may differ across the quantile index and the regressors, we first consider the scalar $\hat\delta_k(\tau)$, which is the $k$-th element of $\hat\delta(\tau)$ for $k\in\{1,\dots,K\}$ and $\tau \in\mathcal{T}$. From Lemma (ref), \begin{equation*} \hat\delta_k(\tau)-\delta_k(\tau)=\hat G_k(\tau)\bar g_{mn}(\hat\delta,\tau)=\hat G_k(\tau)\left(\bar g_{mn}^{(1)}(\hat\delta,\tau)+\bar g_{mn}^{(2)}(\hat\delta,\tau)\right) \end{equation*} where $\hat G_k(\tau)$ is the $k$-th row of $\hat G(\tau)$. From the proof of Lemma (ref)(i), we have \begin{equation*} \bar g_{mn}^{(1)}(\hat\delta,\tau)=\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+o_p\left(\frac{1}{\sqrt{mn}}\right) \end{equation*} uniformly in $\tau\in\mathcal T$. Together with the uniform boundedness of $\hat G(\tau)$, it implies that \begin{align*} \hat G_k(\tau)\bar g_{mn}^{(1)}(\hat\delta,\tau)&=\hat G_k(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+o_p\left(\frac{1}{\sqrt{mn}}\right) \end{align*} uniformly in $\tau\in\mathcal T$. In addition, Lemma 2(i) implies that $\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})=O_p\left(\frac{1}{\sqrt{mn}}\right)$ uniformly in $\tau\in\mathcal T$. By assumption, $\underset{\tau\in\mathcal{T}}{\sup}\left\lVert \hat G_k(\tau)-G_{mn,k}(\tau)\right\rVert=o_p\left(1\right)$. This, in turn, implies that \begin{equation*} \left(\hat G_k(\tau)-G_{mn,k}(\tau)\right)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})=o_p\left(\frac{1}{\sqrt{mn}}\right) \end{equation*} uniformly in $\tau\in\mathcal T$. Combining the last two displayed equations, we obtain \begin{align} \hat G_k(\tau)\bar g_{mn}^{(1)}(\hat\delta,\tau)=G_{mn,k}(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+o_p\left(\frac{1}{\sqrt{mn}}\right) \end{align} uniformly in $\tau\in\mathcal T$. Lemma (ref)(ii) implies that \begin{equation}\bar g_{mn,l}^{(2)}(\delta,\tau)=\frac{1}{m}\sum_{j = 1}^m \bar z_{j,l}\alpha_j(\tau)=O_p\left(\frac{1}{\sqrt m}\right)\sqrt{\Omega_{2,ll}(\tau)}\end{equation} uniformly in $\tau\in\mathcal T$, where $g_{mn,l}^{(2)}(\delta,\tau)$ is the $l$-th element of $g_{mn}^{(2)}(\delta,\tau)$, $z_{j,l}$ is the $l$-th element of $z_{j}$, and $\Omega_{2,ll}(\tau)$ it the element in the $l$-th row and $l$-th column of $\Omega_{2}(\tau)$. Let $\hat G_{kl}(\tau)$ and $G_{mn,kl}(\tau)$ be the element in the $k$-th row and $l$-th column of $\hat G(\tau)$ and $G_{mn}(\tau)$, respectively. We have \begin{align} (\hat G_k(\tau)-G_{mn}(\tau)) \bar g_{mn}^{(2)}(\hat\delta,\tau)&=\sum_{l=1}^{L}(\hat G_{kl}(\tau)-G_{mn,kl}(\tau))\bar g_{mn,l}^{(2)}(\hat \delta,\tau)\nonumber\\ &=\sum_{l=1}^{L}o_p \left (\sqrt{G_{mn,kl}(\tau)} \right )O_p\left(\frac{1}{\sqrt m}\right)\sqrt{\Omega_{2,ll}(\tau)}\nonumber\\ &=o_p\left(\frac{1}{\sqrt{m}}\right)\sum_{l=1}^{L} \sqrt{G_{mn,kl}(\tau))\Omega_{2,ll}(\tau)}\nonumber \end{align} uniformly in $\tau\in\mathcal T$. \begin{commentP} Focus on the problematic term of the product $(\hat G - G_{nm} ) \bar g_{mn} $: \begin{align*} ( \hat G_{21} - G_{mn21} ) \bar g_{mn} = ( \hat G_{21} - G_{mm 21} ) \bar g_{mn}^{(1)} \\ o_p(a_n^{-1/2})O_p \left ( \frac{1}{\sqrt{mn}} \right ) = o_p \left ( \frac{1}{\sqrt{mn a_n}} \right ) = o_p \left ( \frac{1}{ \sqrt {m } (na_n)^{1/2}} \right ) = o_p \left ( \frac{1}{ \sqrt {m }} ||\Omega_{22} ||^{1/2} \right ) \\o_p \left ( \frac{1}{ \sqrt {m }} ||G_{21,mn}\Omega_{22} ||^{1/2} \right ) \end{align*} Since $n \cdot a_n = || \Omega_{22}||^{-1}$ and $G_{mn21} = O_p(1)$.\\ \end{commentP} Write \begin{equation} \zeta(k,\tau)=\frac{1}{\sqrt{mn}}+\frac{1}{\sqrt m}\left\lVert G_{mn,k}(\tau)\Omega_2(\tau)\right\rVert ^{1/2}. \end{equation} Thus, uniformly in $\tau\in\mathcal T$, \begin{align*} \hat\delta_k(\tau)-\delta_k(\tau)=&\hat G_k(\tau)\bar g^{(1)}_{mn}(\hat\delta,\tau)+\hat G_k(\tau)\bar g^{(2)}_{mn}(\hat\delta,\tau)\\ =&G_{mn,k}(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+o_p\left(\frac{1}{\sqrt{mn}}\right)\\ &+G_{mn,k}(\tau)\frac{1}{m}\sum_{j=1}^m \bar z_{j}\alpha_j(\tau)+o_p\left(\frac{1}{\sqrt{m}}\right) \left\lVert G_{mn,k}(\tau)\Omega_2(\tau)\right\rVert^{1/2}\\ =&G_{mn,k}(\tau)\left(\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+\frac{1}{m}\sum_{j=1}^m \bar z_{j}\alpha_j(\tau)\right)+o_p\left(\zeta(k,\tau)\right) \end{align*} where the second equality follows equations (ref) and (ref) and the third from the definition $\zeta$ in (ref). Thus, we have \begin{equation*} \hat\delta_k(\tau)-\delta_k(\tau)=\sum_{j=1}^m d_j(k,\tau)+o_p\left(\zeta(k,\tau)\right) \end{equation*} where \begin{equation*} d_j(k,\tau)=G_{mn,k}(\tau)\left(\frac{1}{mn}\Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+\frac{1}{m}\bar z_{j}\alpha_j(\tau)\right) \end{equation*} Let $D_j$ be the $TK\times 1$ vector $(\operatorname{diag}(\Sigma_{mn}(\tau_1))^{-1/2} d_j(\tau_1),\dots, \operatorname{diag}(\Sigma_{mn}(\tau_T))^{-1/2} d_j(\tau_T))'$ where $d_j(\tau)=(d_j(1,\tau),d_j(2,\tau),\dots,d_j(K,\tau))'$. It follows that \begin{equation*} \operatorname{Var}\left(\sum_{j=1}^m D_i\right)=H_{mn} \end{equation*} Then, it follows from the proof of Lemma (ref) and Assumption (ref), \begin{equation*} H_{mn}^{-1/2}\sum_{j=1}^mD_i\underset{d}{\rightarrow}N(0,I_{TK}) \end{equation*} By Slutsky's theorem, \begin{equation*} H^{-1/2}\sum_{j=1}^mD_i=H_{mn}^{-1/2}\sum_{j=1}^mD_i+o_p(1)\underset{d}{\rightarrow}N(0,I_{TK}) \end{equation*} In the proof of Lemma (ref) we show that $\sum_j \operatorname{diag}(\Sigma_{mn}(\cdot))^{-1/2} d_j(\cdot))$ is asymptotically tight in $l^\infty(\mathcal{T})$. It follows that the process $\sum_j \operatorname{diag}(\Sigma_{mn}(\cdot))^{-1/2} d_j(\cdot))$ weakly converges to $\mathbb Z$, a centered Gaussian process with covariance kernel $H(\tau,\tau')$. Finally, we show that $o_p(1)\underset{\tau\in\mathcal T, k\in\{1,\dots,K\}}{\sup}\zeta(k,\tau)\Sigma_{mn}(k,\tau)^{-1/2}=o(1)$, where $\Sigma_{mn}(k,\tau)$ is the $(k,k)$ element of $\Sigma_{mn}(\tau)$. We start by showing that \begin{equation} o_p\left(\frac{1}{\sqrt{m}}\sum_{l=1}^{L} \sqrt{G_{mn,kl}(\tau))\Omega_{2,ll}(\tau)} + \frac{1}{\sqrt{mn}}\right) =o_p\left(\frac{1}{\sqrt{m}}\sqrt{G_{mn,k}(\tau)\Omega_2(\tau)G_{mn,k}(\tau)'} + \frac{1}{\sqrt{mn}} \right).\end{equation} First, note that since for all $l \in L_1$ instruments $\Omega_{2,ll} = 0$ \begin{align*} \sum_{l=1}^{L} \sqrt{G_{mn,kl}(\tau))\Omega_{2,ll}(\tau)} = &\sum_{l \in L_2} \sqrt{G_{mn,kl}(\tau))\Omega_{2,ll}(\tau)} \end{align*} \begin{align*} \sqrt{G_{mn,k}(\tau)\Omega_2(\tau)G_{mn,k}(\tau)'} = & \sum_{l \in L_2} \sum_{l' \in L_2} \sqrt{G_{mn,kl}(\tau)\Omega_{2ll'}(\tau)G_{mn,kl'}(\tau)'} \end{align*} Thus, we only need to show that \begin{align*} o_p \left(\frac{1}{\sqrt{m}}\sum_{l\in L_2} \sqrt{G_{mn,kl}(\tau)\Omega_{2,ll}(\tau)} + \frac{1}{\sqrt{mn}} \right) = o_p \left(\frac{1}{\sqrt{m}}\sum_{l \in L_2} \sum_{l' \in L_2} \sqrt{G_{mn,kl}(\tau)\Omega_{2ll'}(\tau)G_{mn,kl'}(\tau)'}+ \frac{1}{\sqrt{mn}} \right) \end{align*} Depending on $k$, there are two possibilities, $G_{mn,kl} = O_p(a_n)$ or $G_{mn,kl}$ is finite and bounded away from zero. If $G_{mn,kl}$ is finite and bounded away from, the previous results follow directly. If $G_{mn,kl} = O_p(a_n)$, we have $$\frac{1}{\sqrt m}\sum_{l \in L_2} \sqrt{G_{mn,kl}(\tau)\Omega_{2,ll}(\tau)} = \frac{1}{\sqrt m} O_p\left ( \sqrt{a_n(\tau) \parallel \Omega_{2,ll}(\tau) \parallel} \right ) = O_p \left (\frac{1}{\sqrt{mn} } \right), $$ where the last equality uses the definition of $a_n$. which implies the equation ((ref)). Then, we have that \begin{equation*} o\left(\frac{1}{\sqrt{mn}}\right)\frac{1}{\inf_{\tau,k}\Sigma_{mn}(k,\tau)^{1/2}}=o\left(\frac{1}{\sqrt{mn}}\right)\frac{1}{\inf_{\tau,k} (G_{mn,k}(\tau)\Omega(\tau)G_{mn,k}(\tau)')^{1/2}}=o(1) \end{equation*} \begin{multline*} o\left(\frac{1}{\sqrt m}\right)\sup_{\tau,k}\left\lVert G_{mn,k}(\tau)\Omega_2(\tau) G_{mn,k}(\tau)'\right\rVert^{1/2}\Sigma_{mn}(k,\tau)^{-1/2} \\ \leq o(1) \sup_{\tau,k}\left(\frac{(G_{mn,k}(\tau)\Omega_2(\tau) G_{mn,k}(\tau)')^{1/2}}{(G_{mn,k}(\tau)\Omega(\tau) G_{mn,k}(\tau)')^{1/2}}\right)=o(1) \end{multline*} Hence, uniformly in $\tau$, \begin{align*} \operatorname{diag}(\Sigma_{mn}(\tau))^{-1/2}(\hat\delta(\tau)-\delta(\tau))&=\operatorname{diag}(\Sigma_{mn}(\tau))^{-1/2}\left(\sum_{j=1}^m d_j(\tau)+o_p\left(\zeta(\tau)\right)\right)\\ &=\operatorname{diag}(\Sigma_{mn}(\tau))^{-1/2} \sum_{j=1}^m d_j(\tau)+o_p(1)\\ &\rightsquigarrow \mathbb Z(\tau) \end{align*} where $\zeta(\tau)=(\zeta(1,\tau),\dots,\zeta(K,\tau))'$. \end{comment}

Proof of Propositions (ref) and (ref): Properties of $\hat\Omega(\tau,\tau')$

Proposition (ref)

proof[Proof of Proposition (ref)] \allowdisplaybreaks We use $\hat u_{ij}(\tau) = \tilde x_{ij}' \hat \beta_j(\tau) - x_{ij}' \hat \delta (\tau)= \tilde x_{ij}' (\hat \beta_j(\tau) - \beta_j(\tau) ) + x_{ij}' (\delta(\tau) - \hat \delta(\tau))+ \alpha_j(\tau)$ to obtain \begin{align*} \hat \Omega(\tau, \tau')&=\frac{1}{m} \sum_{j = 1}^m \left\{\left(\frac{1}{n}\sum_{i=1}^n z_{ij} \hat u_{ij}(\tau)\right) \left(\frac{1}{n}\sum_{i=1}^n z_{ij}\hat u_{ij}(\tau')\right)'\right\}\\ &=\frac{1}{m} \sum_{j = 1}^m \Bigg\{ \left(\frac{1}{n}\sum_{i=1}^n z_{ij} \left[\tilde x_{ij}' (\hat \beta_j(\tau) - \beta_j(\tau) ) + x_{ij}' (\delta(\tau) - \hat \delta(\tau))+ \alpha_j(\tau)\right]\right) \\ & \hphantom{ + \frac{1}{m} \sum_{j = 1}^m \Bigg\\left(\frac{1}{n}\sum_{i=1}^n z_{ij}\left[\tilde x_{ij}' (\hat \beta_j(\tau') - \beta_j(\tau') ) + x_{ij}' (\delta(\tau') - \hat \delta(\tau'))+ \alpha_j(\tau')\right]\right)' \Bigg\}\\ &=\frac{1}{m} \sum_{j = 1}^m \Bigg\{\left ( \frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'(\hat \beta_j(\tau') -\beta_j(\tau') ) \right )'\\ & \hphantom{ + \frac{1}{m} \sum_{j = 1}^m \Bigg\ + \left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\alpha_j(\tau)\right) \left(\frac{1}{n}\sum_{i = 1}^n z_{ij}\alpha_j(\tau')\right)' \\ & \hphantom{ + \frac{1}{m} \sum_{j = 1}^m \Bigg\ +\left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} x_{ij}'(\hat \delta(\tau) -\delta(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} x_{ij}'(\hat \delta(\tau') -\delta(\tau') ) \right )' \\ & \hphantom{ + \frac{1}{m} \sum_{j = 1}^m \Bigg\ -\left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} \tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} x_{ij}'(\hat \delta(\tau') -\delta(\tau') ) \right )'\\ & \hphantom{ + \frac{1}{m} \sum_{j = 1}^m \Bigg\ -\left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} x_{ij}'(\hat \delta(\tau) -\delta(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'(\hat \beta_j(\tau') -\beta_j(\tau') ) \right )' \\ & \hphantom{ + \frac{1}{m} \sum_{j = 1}^m \Bigg\ +\left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} \tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij}\alpha_j(\tau') \right )'\\ & \hphantom{ + \frac{1}{m} \sum_{j = 1}^m \Bigg\ +\left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} \alpha_j(\tau) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} \tilde x_{ij}'(\hat \beta_j(\tau') -\beta_j(\tau') ) \right )'\\ & \hphantom{ + \frac{1}{m} \sum_{j = 1}^m \Bigg\ -\left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} x_{ij}'(\hat \delta(\tau) -\delta(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} \alpha_j(\tau') \right )' \\ & \hphantom{ + \frac{1}{m} \sum_{j = 1}^m \Bigg\ -\left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} \alpha_j(\tau) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij} x_{ij}'(\hat \delta(\tau') -\delta(\tau') ) \right )' \Bigg\}. \end{align*}\allowdisplaybreaks[0] We will show that the first term converges to $\Omega_1(\tau, \tau')/n$, the second term to $\Omega_2(\tau,\tau')$, and the remaining terms vanish at a rate faster than the leading term. By the proof of Lemma (ref)(i), it follows for the first term that \begin{align*} \frac{1}{m} \sum_{j = 1}^m \Bigg( \frac{1}{n}\sum_{i = 1}^n &z_{ij}\tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \Bigg) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ij}\tilde x_{ij}'(\hat \beta_j(\tau') -\beta_j(\tau') ) \right ) ' \\ = & \mathbb{E} \left [ \left ( \Sigma_{ZXj} \frac{1}{n} \sum_{i = 1}^n \phi_{i,\tau}(\tilde x_{ij}, z_{ij}) \right) \left ( \Sigma_{ZXj} \frac{1}{n} \sum_{i = 1}^n \phi_{i,\tau'}(\tilde x_{ij}, z_{ij}) \right )' \right] + o_p \left ( \frac{1}{mn} \right ) \\ =& \frac{\Omega_1(\tau, \tau')}{n} + O_p \left ( \frac{1}{mn} \right ), \end{align*} uniformly in $\tau$. For the second term, we consider each element of the $L\times L$ matrix separately. For $l,l'\in\{1,\dots,L\}$, we have \begin{align*} \frac{1}{m} \sum_{j = 1}^m \left(\frac{1}{n}\sum_{i = 1}^n z_{ijl}\alpha_j(\tau)\right) \left(\frac{1}{n}\sum_{i = 1}^n z_{ijl'}\alpha_j(\tau')\right)'&=\frac{1}{m} \sum_{j = 1}^m \bar z_{jl}\bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau'). \end{align*} If either $l$ or $l'$ (or both) is in $ \{1,\dots,L_1\}$, then $\Omega_{2ll'}(\tau,\tau')=0$ and $\sum_{j = 1}^m \bar z_{jl}\bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau')=0$. Thus, we need only to consider the case that both $\bar z_{jl}$ and $\bar z_{jl'}$ are not equal to $0$ uniformly across all groups. We apply Theorem 9.2 in hansen2022probability. His condition (9.3) is satisfied by the boundedness of $z_{ij}$ in Assumption (ref)(i) and the uniform boundedness of the $4+\varepsilon_C$ moment of $\alpha_j(\tau)$ in Assumption (ref)(i). Condition (9.5) in hansen2022probability is satisfied by the assumption in equation ((ref)). It follows that \begin{equation*} \sqrt m\left(\frac{1}{m} \sum_{j = 1}^m \bar z_{jl}\bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau')-\Omega_{2ll'}(\tau,\tau')\right)\underset{d}{\rightarrow}N(0,C_{l,l'}(\tau,\tau')). \end{equation*} Then, since the condition for Theorem 18.3 in hansen2022probability are satisfied, we have that $\sup_{\tau, \tau'} \left \rVert \frac{1}{m} \sum_{j = 1}^m \bar z_{jl}\bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau') -\Omega_{2ll'}(\tau,\tau')\right \rVert = O_p\left(\frac{\sqrt{C_{l,l'}(\tau,\tau')}}{\sqrt m} \right)$. Let $C_l$ denotes a uniform bound on $|z_{ijl}|$. From the definition of $C_{l,l'}(\tau,\tau')$ we have \begin{equation*} C_{l,l'}(\tau,\tau')\leq C_l^2C_{l'}^2 \operatorname{Var}(\alpha_j(\tau)\alpha_j(\tau'))\leq C_l^2C_{l'}^2\mathbb{E}[\alpha_j(\tau)^2\alpha_j(\tau')^2]\leq C_l^2C_{l'}^2\sqrt{\mathbb{E}[\alpha_j(\tau)^4]\mathbb{E}[\alpha_j(\tau')^4]} \end{equation*} Finally, note that $\frac{\mathbb{E}[\alpha_j(\tau)^4]}{Var(\alpha_j(\tau))^2}$ is bounded. It follows that $\frac{1}{m} \sum_{j = 1}^m \bar z_{jl}\bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau')=\Omega_{2ll'}(\tau,\tau')+O_p\left(\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))\operatorname{Var}(\alpha_j(\tau'))}}{\sqrt m} \right)=\Omega_{2ll'}(\tau,\tau')+O_p\left(\frac{\sqrt{\Omega_{2,ll}(\tau,\tau)\Omega_{2,l'l'}(\tau'\tau')}}{\sqrt m} \right)$. \begin{comment} {\color{red} I think this is wrong. See part in purple!} \color{gray} $$C_{l,l'}(\tau,\tau')\leq \mathbb{E} [\bar z_{jl}^2\bar z_{jl'}^2\alpha_j(\tau)^2\alpha_j(\tau')^2]\leq \sqrt{\mathbb{E} [\bar z_{jl}^2\alpha_j(\tau)^2]\mathbb{E}[\bar z_{jl'}^2\alpha_j(\tau')^2]} = \sqrt{Var(\bar z_{jl}\alpha_j(\tau))Var((\bar z_{jl'}\alpha_j(\tau'))}$$ It follows that \begin{align*} \frac{1}{m} \sum_{j = 1}^m \bar z_{jl}\bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau')& =\Omega_{2ll'}(\tau,\tau')+O_p\left(\frac{\left(Var(\bar z_{jl}\alpha_j(\tau))Var(\bar z_{jl'}\alpha_j(\tau'))\right)^{1/4}}{\sqrt m}\right) \\ & = \Omega_{2ll'}(\tau,\tau') + O_p \left ( \frac{\Omega_{2,ll} ^{1/4}(\tau) \Omega_{2,l'l'} ^{1/4}(\tau')}{\sqrt{m}} \right ) \\ \end{align*} \begin{commentP} \begin{align*} \mathbb{E} [\bar z_{jl}^2\bar z_{jl'}^2\alpha_j(\tau)^2\alpha_j(\tau')^2] &= \operatorname{Cov}(\bar z_{jl}^2\alpha_j(\tau)^2, \bar z_{jl'}^2\alpha_j(\tau')^2) + {\mathbb{E} [\bar z_{jl}^2\alpha_j(\tau)^2]\mathbb{E}[\bar z_{jl'}^2\alpha_j(\tau')^2]} \\ &\leq \sqrt{\mathbb{E} [\bar z_{jl}^4\alpha_j(\tau)^4]\mathbb{E}[\bar z_{jl'}^4\alpha_j(\tau')^4]} + {Var(\bar z_{jl}\alpha_j(\tau))Var((\bar z_{jl'}\alpha_j(\tau'))} \end{align*} where the first term is weakly larger by Lyapunov's inequality. \\ It follows that \begin{align*} \frac{1}{m} \sum_{j = 1}^m \bar z_{jl}\bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau') & = \Omega_{2ll'}(\tau,\tau') + O_p \left ( \frac{\mathbb{E} [\bar z_{jl}^4\alpha_j(\tau)^4]^{1/4}\mathbb{E}[\bar z_{jl'}^4\alpha_j(\tau')^4]^{1/4}} {\sqrt{m}} \right ) \end{align*} \end{commentP} \color{black} \end{comment} \begin{comment} \begin{align*} C_{l,l'}(\tau,\tau') = \operatorname{Var}(\bar z_{jl}\bar z_{jl'}'\alpha_j(\tau)^2) &= \mathbb{E} [\bar z_{jl}'\bar z_{jl'}\bar z_{jl}'\bar z_{jl'}\alpha_j(\tau)^2\alpha_j(\tau')^2]- {\mathbb{E} [\bar z_{jl}' \bar z_{jl'}\alpha_j(\tau)^2]\mathbb{E}[\bar z_{jl}'\bar z_{jl'}\alpha_j(\tau')^2]}' \end{align*} Then note that since the variance is weakly positive, it must be that \begin{align*} O_p\left ( \mathbb{E} [\bar z_{jl}\bar z_{jl'}'\bar z_{jl}\bar z_{jl'}'\alpha_j(\tau)^2\alpha_j(\tau')^2] \right) = & O_p \left( {\mathbb{E} [\bar z_{jl} \bar z_{jl'}'\alpha_j(\tau)^2]\mathbb{E}[\bar z_{jl}\bar z_{jl'}'\alpha_j(\tau')^2]} \right) \\ \leq & O_p \left( {\mathbb{E} [\bar z_{jl} \bar z_{jl}'\alpha_j(\tau)^2]\mathbb{E}[\bar z_{jl'}\bar z_{jl'}'\alpha_j(\tau')^2]} \right) \end{align*} where the last inequality follows by Cauchy-Schwarz Inequality. Hence, \begin{align*} \frac{1}{m} \sum_{j = 1}^m \bar z_{jl}\bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau') = &\Omega_{2ll'}(\tau,\tau') + O_p \left( \frac{{\sqrt{ \Omega_{2,ll}(\tau) \Omega_{2,l'l'}(\tau') } }}{\sqrt{m}} \right) \end{align*} \end{comment} \begin{comment} \color{gray} [By using the entire lower right matrix] For all $L_2$ instruments (for the other, it is zero anyway). \begin{align*} \operatorname{Var}(\bar z_{j}\bar z_j'\alpha_j(\tau)^2) &= \mathbb{E} [\bar z_{j}\bar z_j'\bar z_{j}\bar z_j'\alpha_j(\tau)^2\alpha_j(\tau')^2]- {\mathbb{E} [\bar z_{j} \bar z_j'\alpha_j(\tau)^2]\mathbb{E}[\bar z_{j}\bar z_j'\alpha_j(\tau')^2]}' \end{align*} Then note that since the Covariance matrix is positive semi-definite, it must be that $$O_p\left ( \parallel \mathbb{E} [\bar z_{j}\bar z_j'\bar z_{j}\bar z_j'\alpha_j(\tau)^2\alpha_j(\tau')^2]\parallel \right) = O_p \left(\parallel {\mathbb{E} [\bar z_{j} \bar z_j'\alpha_j(\tau)^2]\mathbb{E}[\bar z_{j}\bar z_j'\alpha_j(\tau')^2]} \parallel\right) $$ I worry about the off-diagonal elements. Then you have a covariance with could be negative. It follows that for the $L_2$ instruments \begin{align*} \frac{1}{m} \sum_{j = 1}^m \bar z_{jl}\bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau') = &\Omega_{2ll'}(\tau,\tau') + O_p \left( \frac{\parallel {\mathbb{E}[\bar z_{j}\bar z_j'\alpha_j(\tau')^2]} \parallel }{\sqrt m}\right) \\ =& \Omega_{2ll'}(\tau,\tau') + o_p \left( {\parallel {\Omega_{2,22}} \parallel }\right) \end{align*} \color{black} \end{comment} For the third term, we also consider each element $l,l'$ of the matrix separately: \begin{multline*} \frac{1}{m} \sum_{j = 1}^m \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{ij}'(\hat \delta(\tau) -\delta(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl'} x_{ij}'(\hat \delta(\tau') -\delta(\tau') ) \right )=\\ \frac{1}{m} \sum_{j = 1}^m \left ( \frac{1}{n}\sum_{i = 1}^n \sum_{k=1}^K z_{ijl} x_{ijk}(\hat \delta_k(\tau) -\delta_k(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n \sum_{k=1}^K z_{ijl'} x_{ijk}(\hat \delta_k(\tau') -\delta_k(\tau') ) \right ). \end{multline*} \begin{comment} We first derive the results for the case when $\dot x_{1ij}$ are included as instruments and, in the case of overidentification, a weighting matrix that satisfies Assumption (ref) is used. Note that these instruments are always valid, and it is efficient to include them. In addition, our efficient weighting matrix satisfies Assumption (ref), see Lemma (ref). It follows that the estimation error of the coefficients on the individual-level variables $x_{1ij}$ are $O_p\left(\frac{1}{\sqrt{mn}}\right)$. In contrast, the estimation error of the coefficients on the group-level variables $x_{2j}$ are $O_p\left(\frac{1}{\sqrt{mn}}\right)+O_p\left(\frac{\sqrt{Var(\alpha_j)}}{\sqrt{m}}\right)$. \end{comment} We split $\sum_{k=1}^K z_{ijl'} x_{ijk}(\hat \delta_k(\tau') -\delta_k(\tau') )$ into the the individual-level and group-level variables. Since the estimation error of the coefficients on the individual-level variables is $O_p(1/\sqrt{mn})$ and $z_{ijl}$ and $x_{ijk}$ are bounded, we obtain \begin{align*} \sum_{k=1}^{K_1} \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{1ijk}(\hat \delta_k(\tau) -\delta_k(\tau) ) & = O_p \left ( \frac{1}{\sqrt {mn} } \right ) , \end{align*} uniformly in $\tau$. For the group-level variables, we obtain \begin{align*} \sum_{k=K_1 + 1}^{K} \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{2jk}(\hat \delta_k(\tau) -\delta_k(\tau) ) & = \sum_{k=K_1+1}^K \bar z_{jl} x_{2jk}(\hat \delta_k(\tau) -\delta_k(\tau) ). \end{align*} It follows that $\sum_{k=K_1 + 1}^{K} \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{jk}(\hat \delta_k(\tau) -\delta_k(\tau) )=0$ if $\bar z_{jl}=0$ for all $j=1,\dots,m$, i.e. if $l\in \{1,\dots,L_1\}$. If $l\in\{L_1+1,\dots,L\}$, uniformly in $\tau$, $\sum_{k=K_1 + 1}^{K} \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{jk}(\hat \delta_k(\tau) -\delta_k(\tau) )=O_p\left(\frac{1}{\sqrt{mn}}+\frac{\sqrt{Var(\alpha_j(\tau))}}{\sqrt m}\right)$. Since $\Omega_{2,ll}(\tau)=Var(\bar z_{jl}\alpha_j(\tau))$, in both cases we can write $$\sum_{k=1}^{K} \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{jk}(\hat \delta_k(\tau) -\delta_k(\tau) )=O_p\left(\frac{1}{\sqrt{mn}}+\frac{\sqrt{\Omega_{2,ll} (\tau)}}{\sqrt m}\right)$$ uniformly in $\tau$. Combining these results, we get uniformly in $\tau, \tau'$ \begin{align} \frac{1}{m} \sum_{j = 1}^m & \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{ij}'(\hat \delta(\tau) -\delta(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl'} x_{ij}'(\hat \delta(\tau') -\delta(\tau') ) \right )\\ & = O_p \left (\frac{\sqrt{ \Omega_{mn,ll}(\tau) \Omega_{mn,l'l'}(\tau')}}{ m} \right ). \end{align} \begin{comment} Keep for later if we want to allow for uniformity in $\bar z_j$ It follows that, for each $K_1+1<k<K$ \begin{align*} & = O_p(\bar z_{jl} ) O_p \left ( \frac{\sqrt{Var(\alpha_j(\tau))}}{\sqrt m} + \frac{1}{\sqrt{mn}} \right ) \\ &= O_p \left ( \frac{\sqrt{ \mathbb{E}[\bar z_{jl}^2] \mathbb{E}[\alpha_j^2(\tau)]}}{\sqrt m} + \frac{\sqrt{\mathbb{E}[\bar z_{jl}^2]}}{\sqrt{mn}} \right ) \\ & = O_p \left ( \frac{\sqrt{ \mathbb{E}[\bar z_{jl}^2 \alpha_j^2(\tau)]}}{\sqrt m} + \frac{\sqrt{\mathbb{E}[\bar z_{jl}^2]}}{\sqrt{mn}} \right ) \\ & = O_p \left ( \frac{\sqrt{ \Omega_{2ll }(\tau)}}{\sqrt m} + \frac{\sqrt{\mathbb{E}[\bar z_{jl}^2]}}{\sqrt{mn}} \right ) \end{align*} where the fourth equality uses $Cov(\bar z_{jl}^2 , \alpha_j^2) = 0$. Let's try the generalization. \begin{align*} \hat\delta_k(\tau)&=\delta_k(\tau)+O_p(\lVert G_{mn,k}\bar g_{mn}(\delta(\tau),\tau)\rVert)\\ &=\delta_k(\tau)+\sum_{l=1}^L O_p(G_{mn,kl}\bar g_{mn,l}(\delta(\tau),\tau))\\ &=\delta_k(\tau)+\sum_{l=1}^L O_p(G_{mn,kl})O_p(\bar g_{mn,l}(\delta(\tau),\tau))\\ &=\delta_k(\tau)+\sum_{l=1}^L O_p(G_{mn,kl})O_p(\bar g_{mn,l}(\delta(\tau),\tau))\\ &=\delta_k(\tau)+O_p\left(\frac{1}{\sqrt{mn}}\right) +\sum_{l=L_1+1}^L O_p(G_{mn,kl})O_p\left(\frac{\sqrt{Var(\alpha)}}{\sqrt m}\right) \end{align*} \end{comment} For the fourth term, similar arguments and the fact that $\sup_\tau \left \rVert \beta_j(\tau)-\beta_j(\tau)\right \rVert=O_p\left(\frac{1}{\sqrt n}\right)$ imply that uniformly in $\tau, \tau'$ \begin{align} \frac{1}{m} \sum_{j = 1}^m & \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl} \tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl'} x_{ij}'(\hat \delta(\tau') -\delta(\tau') ) \right) \\ & = O_p \left (\frac{1}{\sqrt{m}n} + \frac{ \sqrt{\Omega_{2,l'l'}(\tau')}} {\sqrt{mn}} \right ). \end{align} Similarly, for the fifth term \begin{align} \frac{1}{m} \sum_{j = 1}^m & \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl'} x_{ij}'(\hat \delta(\tau') -\delta(\tau') ) \right)\left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl} \tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \right) \\ & = O_p \left (\frac{1}{\sqrt{m}n} + \frac{ \sqrt{\Omega_{2,ll}(\tau)}} {\sqrt{mn}} \right ), \end{align} uniformly in $\tau, \tau'$. Thus, the sum of the fourth and fifth terms is \begin{align*} O_p \left ( \frac{1}{\sqrt{m}n}+\frac{\sqrt{ \Omega_{2,ll}(\tau)}}{ \sqrt{m}\sqrt{n}} + \frac{\sqrt{ \Omega_{2,l'l'}(\tau')}}{ \sqrt{m} \sqrt{n}} \right )=O_p \left (\frac{ \sqrt{\Omega_{mn,ll}(\tau)\Omega_{mn,l'l'}(\tau') }} {\sqrt{m}}\right ),\end{align*} uniformly in $\tau$. For the sixth term, we obtain, uniformly in $\tau, \tau'$ \begin{align*} \frac{1}{m} \sum_{j = 1}^m \Biggl( \frac{1}{n}\sum_{i = 1}^n z_{ijl} & \tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \Biggr) \Biggl( \frac{1}{n}\sum_{i = 1}^n z_{ijl'}\alpha_j(\tau') \Biggr) \\ & = \frac{1}{m} \sum_{j = 1}^m \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl} \tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \right) \bar z_{jl'}\alpha_j(\tau') \\ & = O_p \left ( \frac{ \sqrt{\Omega_{2,l'l'}(\tau') }} {\sqrt{mn}} \right ), \end{align*} and similarly, we can show that the seventh term is $O_p \left ( \frac{ \sqrt{\Omega_{2,ll}(\tau) }} {\sqrt{mn}} \right )$. \begin{comment} This is wrong. This is not fast enough. Try a different approach (exploit that the two terms are asymptotically independent). We apply again Theorem 9.2 in hansen2022probability. [CHECK CONDTIONS]By similar step as in the third term, we find \begin{align} \frac{1}{m} \sum_{j = 1}^m & \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl} \tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl'}'\alpha_j(\tau') \right ) \\ & = \frac{1}{m} \sum_{j = 1}^m \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl} \tilde x_{ij}'(\hat \beta_j(\tau) -\beta_j(\tau) ) \right) \left ( \bar z_{jl'}'\alpha_j(\tau') \right ) = O_p \left ( \frac{ \sqrt{\Omega_{2,l'l'} }} {m\sqrt{n}} \right ) \end{align} \end{comment} For the eighth term, we obtain \begin{align*} \frac{1}{m} \sum_{j = 1}^m \Biggl( \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{ij}'&(\hat \delta(\tau) -\delta(\tau) ) \Biggr) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl'}' \alpha_j(\tau') \right ) \\ & = \frac{1}{m} \sum_{j = 1}^m \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{ij}'(\hat \delta(\tau) -\delta(\tau) ) \right) \bar z_{jl'} \alpha_j(\tau') \\ &= O_p\left(\frac{\sqrt{\Omega_{2,ll}(\tau)\Omega_{2,l'l'}(\tau')}}{m} +\frac{\sqrt{\Omega_{2,l'l'}(\tau')}}{m\sqrt n}\right ) , \end{align*} and the ninth term is $O_p\left(\frac{\sqrt{\Omega_{2,ll}(\tau)\Omega_{2,l'l'}(\tau')}}{m} +\frac{\sqrt{\Omega_{mn,l'l'}(\tau')}}{m\sqrt n}\right )$ such that the sum of the eight and ninth term is $O_p\left(\frac{\sqrt{\Omega_{mn,ll}(\tau)\Omega_{mn,l'l'}(\tau')}}{m}\right)$, where both results are uniformly in $\tau, \tau'$. Combining all the terms, we find that for each $l$, $l'$ entry \begin{align*} \hat \Omega(\tau, \tau') =& \frac{\Omega_{1,ll'}(\tau, \tau')}{n} + \Omega_{2,ll'}(\tau,\tau') + O_p \left ( \frac{\sqrt{\Omega_{mn,ll}(\tau) \Omega_{mn,l'l' }(\tau)}}{\sqrt{m}} \right)\nonumber \\ = &\Omega_{mn,ll'}(\tau, \tau') + o_p \left ( {\sqrt{\Omega_{mn,ll}(\tau) \Omega_{mn,l'l' }(\tau')}} \right), \end{align*} uniformly in $\tau, \tau'$. \begin{comment} \begin{align*} \frac{1}{m} \sum_{j = 1}^m \left ( \frac{1}{n} \sum_{i = 1}^n z_{ijl} \hat u_j(\tau) \right ) \left ( \frac{1}{n} \sum_{i = 1}^n z_{ijl'} \hat u_j(\tau) \right )' = & \frac{\Omega_{1ll'}(\tau, \tau')}{n} + \Omega_{2ll'}(\tau,\tau') + O_p \left ( \frac{\Omega_{2,ll} ^{1/4}(\tau) \Omega_{2,l'l'} ^{1/4}(\tau')}{\sqrt{m}} \right ) \\ & + O_p \left ( \frac{\sqrt{ \Omega_{2,ll}(\tau,\tau')}}{m\sqrt{n}} + \frac{ \sqrt{\Omega_{2,l'l'}(\tau,\tau') }} {m\sqrt{n}}+ \frac{1}{\sqrt{m}n} \right ) \end{align*} \end{comment} \begin{comment} for $l = l'$ we have \begin{align*} \frac{1}{m} \sum_{j = 1}^m \left ( \frac{1}{n} \sum_{i = 1}^n z_{ijl} \hat u_j(\tau) \right ) \left ( \frac{1}{n} \sum_{i = 1}^n z_{ijl} \hat u_j(\tau) \right )' = & \frac{\Omega_{1ll}(\tau, \tau')}{n} + \Omega_{2,ll}(\tau,\tau') + O_p \left ( \frac{\Omega_{2,ll} ^{1/4}(\tau) \Omega_{2,ll} ^{1/4}(\tau')}{\sqrt{m}} \right ) \\ & + O_p \left ( \frac{1}{\sqrt{m}n} \right )\\ & = \frac{\Omega_{1ll}(\tau, \tau')}{n} + \Omega_{2,ll}(\tau,\tau') + o_p(\Omega_{ll}(\tau, \tau')) \end{align*} \end{comment}

Proposition (ref)$'$

When the coefficients of some individual-level variables converge at a slow rate, while those of others converge at a fast rate, the third term in the proof of Proposition (ref) may not be $O_p\left(\frac{\sqrt{\Omega_{mn,ll}(\tau)\Omega_{mn,l'l'}(\tau')}}{m}\right)$. The slowly converging coefficients can introduce an estimation error in the variance of the faster moments that diminishes more slowly than the true value. Proposition (ref)$'$, stated below, provides a more general result that does not assume the coefficients of individual-level variables converge at the $O_p(1/\sqrt{mn})$ rate. Consequently, $\hat\Omega(\tau,\tau')$ is consistent as long as all individual-level variable coefficients converge at the same rate. One example of this is the between estimator for individual-level variables. Another example is the 2SLS estimator applied to the random effects model. In this case, the weighting matrix is full rank, and all coefficients converge at the rate of $1/\sqrt{mn}+\sqrt{\operatorname{Var}(\alpha_j(\tau))}/\sqrt{m}$.

propositionp{(ref)$'$} Let assumptions (ref)-(ref) and (ref)(c) hold. As $m\rightarrow\infty$, for each $l,l'\in \{1,\dots,L\}$ and uniformly in $\tau,\tau'\in\mathcal{T}^2$, \begin{equation*} m^{-1}\sum_{j=1}^m \mathbb{E} \left[\left(\bar z_{jl} \bar z_{jl'}\alpha_j(\tau)\alpha_j(\tau')-\Omega_{2ll'}(\tau,\tau')\right)^2\right]\rightarrow C_{l,l'}(\tau,\tau')<\infty \end{equation*} The estimator used to compute $\hat u_{ij}(\tau)$ satisfies \begin{equation*}\hat\delta(\tau)-\delta(\tau)= O_p\left(\frac{1}{\sqrt{mn}} + \frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt m} \right)\end{equation*} uniformly in $\tau$. Then, for any $ll'$ entry of the $\hat \Omega(\tau, \tau')$ matrix with $l,l' \in \{1, \dots , L \}$ we have \begin{align*} \hat \Omega (\tau, \tau')=\Omega_{mn,ll'}(\tau, \tau') + o_p \left (\frac{1}{n}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt n}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau'))}}{\sqrt n}+\sqrt{\operatorname{Var}(\alpha_j(\tau))\operatorname{Var}(\alpha_j(\tau'))} \right) \end{align*}
proofThe proof follows the same steps as the proof of Proposition (ref), but requires some modifications each time $\hat\delta(\tau)-\delta(\tau)$ is involved. This term is now $O_p\left(\frac{1}{\sqrt{mn}}+\frac{\operatorname{Var}(\alpha_j(\tau))}{\sqrt m} \right)$ instead of $O_p\left(\frac{1}{\sqrt{mn}}+\frac{\Omega_{2}(\tau)}{\sqrt m} \right)$. For the third term, we obtain \begin{align*} \frac{1}{m} \sum_{j = 1}^m & \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl} x_{ij}'(\hat \delta(\tau) -\delta(\tau) ) \right) \left ( \frac{1}{n}\sum_{i = 1}^n z_{ijl'}' x_{ij}'(\hat \delta(\tau') -\delta(\tau') ) \right )\\ & = O_p\left(\frac{1}{mn} + \frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))\operatorname{Var}(\alpha_j(\tau'))}}{m} + \frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt{n }m} + \frac{\sqrt{\operatorname{Var}(\alpha_j(\tau'))}}{\sqrt{n }m}\right). \end{align*} The sum of the fourth and fifth terms is now \begin{align*} O_p\left(\frac{1}{n\sqrt m}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau))}}{\sqrt{n m}}+\frac{\sqrt{\operatorname{Var}(\alpha_j(\tau'))}}{\sqrt{n m}}\right). \end{align*} The sum of the eight and ninth terms is now $$O_p\left(\frac{\sqrt{\Omega_{2,ll}(\tau)}}{m\sqrt n}+\frac{\sqrt{\Omega_{2,l'l'}(\tau')}}{m\sqrt n}+\frac{\sqrt{\operatorname{Var}(\alpha(\tau'))\Omega_{2,ll}(\tau)}}{m}+\frac{\sqrt{\operatorname{Var}(\alpha(\tau))\Omega_{2,l'l'}(\tau')}}{m}\right).$$ The result of the lemma follows.

Proof of Proposition (ref): Adaptive Inference

proof[Proof of Proposition (ref)] We prove the results for $T = 2$ as the proof trivially extends to $T > 2$. We consider the $2K\times 1$ coefficient vector $\hat \delta(\tau, \tau')=(\hat\delta(\tau)',\hat\delta(\tau')')'$ whose asymptotic covariance matrix is \begin{align*} \Sigma_{mn} = \begin{pmatrix} \Sigma_{mn}(\tau) & \Sigma_{mn}(\tau, \tau') \\ \Sigma_{mn}(\tau, \tau') & \Sigma_{mn}( \tau') \end{pmatrix}. \end{align*} This covariance matrix is estimated by \begin{align*} \hat \Sigma = & \frac{1}{m}\begin{pmatrix} \hat G(\tau) & 0\\ 0 & \hat G(\tau') \end{pmatrix} \begin{pmatrix} \hat \Omega(\tau, \tau) & \hat \Omega(\tau,\tau') \\ \hat \Omega(\tau', \tau) & \hat \Omega(\tau') \end{pmatrix} \begin{pmatrix} \hat G(\tau) & 0\\ 0 & \hat G(\tau') \end{pmatrix} ' \\= &\frac{1}{m} \begin{pmatrix} \hat G(\tau) \hat \Omega(\tau) \hat G(\tau) & \hat G(\tau) \hat \Omega(\tau,\tau') \hat G(\tau') \\ \hat G(\tau') \hat \Omega(\tau',\tau) \hat G(\tau) & \hat G(\tau') \hat \Omega(\tau') \hat G(\tau') \end{pmatrix}. \end{align*} For each $k,k'\in\{1,\dots,2K\}$, we will show that \begin{equation*}\hat\Sigma_{kk'}=\Sigma_{mn,kk'}+o_p\left(\sqrt{\Sigma_{mn,kk}\Sigma_{mn,k'k'}}\right).\end{equation*} To simplify the notation, we show this result for the $K\times K$ top-left submatrix of $\Sigma_{mn}$ such that we can drop the dependence on $\tau$. The proof for the other parts is similar. Our analysis relies on two key ingredients: First, by Proposition (ref), for any $ll'$ entry of the $\hat \Omega$ matrix with $l,l' \in \{1, \dots , L \}$, we have \begin{align*} \hat \Omega_{ll'}= \Omega_{mn,ll'} + o_p \left ( \sqrt{ \Omega_{mn,ll} \Omega_{mn,l'l'} } \right ) \end{align*} where $\Omega$ can be split into four submatrices where the dimensions of the top-left part are $L_1\times L_1$ and the bottom-right are $L_2\times L_2$: \begin{equation*} \Omega=\begin{pmatrix} \Omega_{mn,11} & \Omega_{mn,12} \\ \Omega_{mn,21} & \Omega_{mn,22} \end{pmatrix} \end{equation*} such that $\Omega_{mn,11}$, $\Omega_{mn,12}$, and $\Omega_{mn,21}$ are $O_p\left(\frac{1}{n}\right)$ while $\Omega_{mn,22}=O_p\left(\frac{1}{na_n}\right)$. Second, by Lemma (ref), \begin{equation*} \hat G= \begin{pmatrix} G_{mn,11} & G_{mn,12} \\ G_{mn21} & G_{mn,22} \end{pmatrix}+ \begin{pmatrix} o_p\left(1\right) & o_p \left (\sqrt{a_n}\right ) \\ o_p \left (1/\sqrt{a_n} \right) & o_p\left(1\right) \end{pmatrix}, \end{equation*} where $G_{mn,12}=O_p\left(a_n\right)$ and the other elements of $G_{mn}$ are $O_p\left(1\right)$. With this notation, \begin{align*} \Sigma_{mn}&=\frac{1}{m} \begin{pmatrix} G_{mn,11} & G_{mn,12}\\ G_{mn,21} & G_{mn,22} \end{pmatrix} \begin{pmatrix} \Omega_{mn,11} & \Omega_{mn,12}\\ \Omega_{mn,21} & \Omega_{mn,22} \end{pmatrix} \begin{pmatrix} G_{mn,11}' & G_{mn,21}'\\ G_{mn,12}' & G_{mn,22}' \end{pmatrix}\\ &=\begin{pmatrix} \Sigma_{mn,11} & \Sigma_{mn,12}\\ \Sigma_{mn,21} & \Sigma_{mn,22} \end{pmatrix}, \end{align*} where \begin{align*} \Sigma_{mn,11}&= \frac{1}{m} \left (G_{mn,11}\Omega_{mn,11}G_{mn,11}'+G_{mn,11}\Omega_{mn,12}G_{mn,12}'+G_{mn,12}\Omega_{mn,21}G_{mn,11}'+G_{mn,12}\Omega_{mn,22}G_{mn,12}'\right),\\ \Sigma_{mn,12}&=\frac{1}{m} \left (G_{mn,11}\Omega_{mn,11}G_{mn,21}'+G_{mn,11}\Omega_{mn,12}G_{mn,22}'+G_{mn,12}\Omega_{mn,21}G_{mn,21}'+G_{mn,12}\Omega_{mn,22}G_{mn,22}'\right),\\ \Sigma_{mn,21}&=\frac{1}{m} \left (G_{mn,21}\Omega_{mn,11}G_{mn,11}'+G_{mn,21}\Omega_{mn,12}G_{mn,12}'+G_{mn,22}\Omega_{mn,21}G_{mn,11}'+G_{mn,22}\Omega_{mn,22}G_{mn,12}' \right),\\ \Sigma_{mn,22}&=\frac{1}{m} \left (G_{mn,21}\Omega_{mn,11}G_{mn,21}'+G_{mn,21}\Omega_{mn,12}G_{mn,22}'+G_{mn,22}\Omega_{mn,21}G_{mn,21}'+G_{mn,22}\Omega_{mn,22}G_{mn,22}'\right). \end{align*} It follows that $\Sigma_{mn,11}=O_p\left(\frac{1}{mn}\right)$, $\Sigma_{mn,12}=O_p\left(\frac{1}{mn}\right)$, $\Sigma_{mn,21}=O_p\left(\frac{1}{mn}\right)$, and $\Sigma_{mn,22}=O_p\left(\frac{1}{mna_n}\right)$. We consider the estimation error of these four terms separately, each of them being composed of four parts. For the first term, \begin{align*} \hat G_{11}\hat \Omega_{11}\hat G_{11}' &=(G_{mn,11}+o_p\left(1\right))(\Omega_{mn,11}+o_p\left(\Omega_{mn,11}\right))(G_{mn,11}'+o_p\left(1\right))\\ &=G_{mn,11}\Omega_{mn,11}G_{mn,11}'+o_p(\Omega_{mn,11}) \\ &=G_{mn,11}\Omega_{mn,11}G_{mn,11}'+o_p\left(\frac{1}{n}\right),\\ \hat G_{11}\hat \Omega_{12}\hat G_{12}' &=(G_{mn,11}+o_p\left(1\right))\left (\Omega_{mn,12}+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}}\right)\right)\left (G_{mn,12}'+o_p\left(\sqrt{a_n}\right) \right )\\ &=G_{mn,11}\Omega_{mn,12}G_{mn12}'+o_p\left(\Omega_{mn,12}\right)+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}a_n}\right)\\ &=G_{mn,11}\Omega_{mn,12}G_{mn,12}'+o_p\left(\frac{1}{n}\right),\\ \hat G_{12}\hat\Omega_{21}\hat G_{11}'&=(\hat G_{11}\hat \Omega_{12}\hat G_{12}')'\\ &=G_{mn,12}\Omega_{mn,21}G_{mn,11}'+o_p\left(\frac{1}{n}\right),\\ \hat G_{12}\hat \Omega_{22}\hat G_{12}' &=(G_{mn,12}+o_p\left(\sqrt{a_n}\right))(\Omega_{mn,22}+o_p\left(\Omega_{mn,22}\right))(G_{mn,12}'+o_p\left(\sqrt{a_n}\right))\\ &=G_{mn,12}\Omega_{mn,22}G_{mn,12}'+o_p\left(a_n\Omega_{mn,22}\right)\\ &=G_{mn,12}\Omega_{mn,22}G_{mn,12}'+o_p\left(\frac{1}{n}\right). \end{align*} It follows that $\hat\Sigma_{11}=\Sigma_{mn,11}+o_p\left(\frac{1}{mn}\right)=\Sigma_{mn,11}+o_p\left(\Sigma_{mn,11}\right)$. For the second part, \begin{align*} \hat G_{11}\hat \Omega_{11}\hat G_{21}' &=(G_{mn,11}+o_p\left(1\right))\left (\Omega_{mn,11}+o_p\left(\Omega_{mn,11}\right)\right )(G_{mn,21}'+o_p\left(1/\sqrt{a_n}\right))\\ &=G_{mn,11}\Omega_{mn,11}G_{mn,21}'+o_p\left(\Omega_{mn,11}/\sqrt{a_n}\right)\\ &=G_{mn,11}\Omega_{mn,11}G_{mn,21}'+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}}\right),\\ \hat G_{11}\hat \Omega_{12}\hat G_{22}' &=(G_{mn,11}+o_p\left(1\right))\left(\Omega_{mn,12}+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}}\right) \right)(G_{mn,22}'+o_p\left(1\right) )\\ &=G_{mn,11}\Omega_{mn,12}G_{mn,22}'+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}}\right),\\ \hat G_{12}\hat\Omega_{21}\hat G_{21}'&=(G_{mn,12}+o_p\left(\sqrt{a_n}\right))(\Omega_{mn,21}+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}}\right)(G_{mn,21}'+o_p\left(1/\sqrt{a_n}\right))'\\ &=G_{mn,12}\Omega_{mn,21}G_{mn,21}'+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}}\right),\\ \hat G_{12}\hat \Omega_{22}\hat G_{22}' &=(G_{mn,12}+o_p\left(\sqrt{a_n}\right))(\Omega_{mn,22}+o_p(\Omega_{mn,22}))(G_{mn,22}'+o_p\left(1\right))\\ &=G_{mn,12}\Omega_{mn,22}G_{mn,22}'+o_p\left(\sqrt{a_n}\Omega_{mn,22}\right)\\ &=G_{mn,12}\Omega_{mn,22}G_{mn,22}'+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}}\right). \end{align*} It follows that $\hat\Sigma_{12}=\Sigma_{mn,12}+o_p(m^{-1}\sqrt{\Omega_{mn,11}\Omega_{mn,22}})=\Sigma_{mn,12}+o_p(\sqrt{\Sigma_{mn,11}\Sigma_{mn,22}})$. The third part, $\hat\Sigma_{21}$, is the transpose of $\hat\Sigma_{12}$. For the fourth part, \begin{align*} \hat G_{21}\hat \Omega_{11}\hat G_{21}' &=(G_{mn,21}+o_p\left(1/\sqrt{a_n}\right))(\Omega_{mn,11}+o_p\left(\Omega_{mn,11}\right))\left(G_{mn,21}'+o_p\left(1/\sqrt{a_n}\right) \right)\\ &=G_{mn,21}\Omega_{mn,11}G_{mn,21}'+o_p\left(\Omega_{mn,11}/a_n\right)\\ &=G_{mn,21}\Omega_{mn,11}G_{mn,21}'+o_p\left(\Omega_{mn,22}\right),\\ \hat G_{21}\hat \Omega_{12}\hat G_{22}' &=(G_{mn,21}+o_p\left(1/\sqrt{a_n}\right))\left(\Omega_{mn,12}+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}}\right) \right) \left (G_{mn,22}'+o_p\left(1\right) \right )\\ &=G_{mn,21}\Omega_{mn,12}G_{mn,22}'+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}/a_n}\right)\\ &=G_{mn,21}\Omega_{mn,12}G_{mn,22}'+o_p\left(\Omega_{mn,22}\right),\\ \hat G_{22}\hat\Omega_{21}\hat G_{21}'&=(G_{mn,22}+o_p\left(1\right))\Omega_{mn,21}+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}}\right)(G_{mn,21}'+o_p\left(1/\sqrt{a_n}\right))'\\ &=G_{mn,22}\Omega_{mn,21}G_{mn,21}'+o_p\left(\sqrt{\Omega_{mn,11}\Omega_{mn,22}/a_n}\right)\\ &=G_{mn,22}\Omega_{mn,21}G_{mn,21}'+o_p\left(\Omega_{mn,22}\right),\\ \hat G_{22}\hat \Omega_{22}\hat G_{22}' &=(G_{mn,22}+o_p\left(1\right))(\Omega_{mn,22}+o_p\left(\Omega_{mn,22}\right))(G_{mn,22}'+o_p\left(1\right))\\ &=G_{mn,22}\Omega_{mn,22}G_{mn,22}'+o_p\left(\Omega_{mn,22}\right). \end{align*} It follows that $\hat\Sigma_{22}=\Sigma_{mn,22}+o_p\left(m^{-1}\Omega_{mn,22}\right)=\Sigma_{mn,22}+o_p\left(\Sigma_{mn,22}\right)$. The results for these four submatrices imply that \begin{equation*}\hat\Sigma_{kk'}=\Sigma_{mn,kk'}+o_p(\sqrt{\Sigma_{mn,kk}\Sigma_{mn,k'k'}}),\end{equation*} which implies that \begin{equation*} \eta'\hat\Sigma\eta=\eta'\Sigma_{mn}\eta+o_p\left(\eta'\Sigma_{mn}\eta\right). \end{equation*} The proposition follows from Theorem (ref). \begin{comment} Each $k, k'$ entry of this matrix can be written as \begin{align*} \frac{1}{m}\sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau'). \end{align*} By Lemma (ref), uniformly in $\tau\in\mathcal T$, {\color{red}$\hat G(\tau)= G_{mn}(\tau)+ o_p(\sqrt{G_{mn}(\tau)})$} and $G_{mn}(\tau)=O_p(1)$. where $G_{mn,12}(\tau)=O_p(a_n(\tau))$ and the other elements are $O_p(1)$. It follows that $\hat \Sigma_{ll'}=\Sigma_{mn,ll'}+o_p(\sqrt{\Sigma_{mn,ll}\Sigma_{mn,l'l'}})$. It follows that $\hat \Sigma (\tau, \tau') = \hat G(\tau)\hat\Omega(\tau,\tau')\hat G(\tau')'=$ \begin{align*} \begin{pmatrix} G_{11}(\tau)+o_p(1) & G_{12}(\tau)+o_p(\sqrt{a_n(\tau)})\\ G_{21}(\tau)+o_p(1/\sqrt{a_n(\tau)} & \end{pmatrix} \begin{pmatrix} \Omega_{11}(\tau\tau')+o_p(\Omega_{11}(\tau,\tau') & \Omega_{12}+o_p(\sqrt{\Omega_11(\tau,\tau')\Omega_{22}(\tau,\tau')}\\\Omega_{21}+o_p(\sqrt{\Omega_11(\tau,\tau')\Omega_{22}(\tau,\tau')} & \Omega_{22}(\tau,\tau')+o_p(\Omega_22(\tau\tau') \end{pmatrix}\end{align*} \begin{align*} \begin{pmatrix} G_{11}(\tau)+o_p(\sqrt{G_{11}(\tau)}) & G_{12}(\tau)+o_p(\sqrt{G_{12}(\tau)})\\ G_{21}(\tau)+o_p(1/\sqrt{a_n(\tau)}) & G_{22}+o_p(\sqrt{G_{22}(\tau)}) \end{pmatrix} \begin{pmatrix} \Omega_{11}(\tau,\tau')+o_p(\Omega_{11}(\tau,\tau')) & \Omega_{12}+o_p(\sqrt{\Omega_{11}(\tau,\tau')\Omega_{22}(\tau,\tau')}\\\Omega_{21}+o_p(\sqrt{\Omega_{11}(\tau,\tau'))\Omega_{22}(\tau,\tau')}) & \Omega_{22}(\tau,\tau')+o_p(\Omega_{22}(\tau\tau')) \end{pmatrix} \end{align*} Hence, it follows that the $k, k'$ entry of $\frac{1}{m} \hat G(\tau) \hat \Omega(\tau, \tau') \hat G(\tau')$ equals \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') \end{align*} We have to consider each possible combination of $l, l', k,k'$. Out of the $2^4 $ possibility, many are the transpose of each other, therefore, it suffices to consider only one. For instance, for any $l, l'$ the combination $k \in \{1, \dots K_1\}$, $k' \in \{K_1 + 1, \dots , K\}$ is the transpose of $k \in \{K_1 + 1, \dots , K\}$, $k' \in \{1, \dots K_1\} $. A similar argument applies to $l$ and $l'$. Hence, we have to check $10$ possible combinations. We consider separately each of the four possible combination coefficients $k,k'$. For each of these entries, we consider all four possible types of instruments.\\ Part i) $k \in \{1, \dots , K_1\}, k' \in \{1, \dots , K_1\}$. For $l \in \{1, \dots , L_1 \}, l' \in \{1, \dots , L_1 \}$, we have $\underset{\tau\in\mathcal{T}}{\sup}\left\lVert \hat G_{kl}(\tau)-G_{mn,kl}(\tau)\right\rVert=o_p\left(1\right)$,\\ $\underset{\tau'\in\mathcal{T}}{\sup}\left\lVert \hat G_{k'l'}(\tau)-G_{mn,k'l'}(\tau')\right\rVert=o_p\left(1\right)$ and $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\frac{1}{n} \right) $. Hence, \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') + o_p \left ( \frac{1}{mn} \right) \end{align*} L1 L2 For $l = \in \{1, \dots , L_1 \}$, $ l' \in \{L_1 + 1, \dots , L \}$ we have $\underset{\tau\in\mathcal{T}}{\sup}\left\lVert \hat G_{kl}(\tau)-G_{mn,kl}(\tau)\right\rVert=o_p\left(1\right)$, $\underset{\tau'\in\mathcal{T}}{\sup}\left\lVert \hat G_{k'l'}(\tau')-G_{mn,k'l'}(\tau')\right\rVert=o_p\left(\sqrt{G_{mn,k'l'}(\tau')}\right)$, and $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\sqrt{\frac{\operatorname{Var}(\alpha(\tau'))}{n}} \right) $. Further $G_{mn,kl}(\tau) = O_p(1)$ and $G_{mn,kl'}(\tau) = O_p(a_n)$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = & \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\ &+ o_p \left ( \frac{1}{m} \sqrt{\frac{\operatorname{Var}(\alpha(\tau'))}{n}} {\sqrt{ G_{mn, kl'}(\tau')}} \right) \end{align*} Similarly, for $ l' \in \{L_1 + 1, \dots , L \}$, $l' = \in \{1, \dots , L_1 \}$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = & \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau) \&+ o_p \left ( \frac{1}{m} \sqrt{\frac{\operatorname{Var}(\alpha(\tau))}{n}} {\sqrt{G_{mn,kl}(\tau)}} \right) \end{align*} For $l \in \{L_1 + 1, \dots , L \}, l' \in \{L_1 + 1, \dots , L \}$, $\underset{\tau\in\mathcal{T}}{\sup}\left\lVert \hat G_{kl}(\tau)-G_{mn,kl}(\tau)\right\rVert=o_p\left(\sqrt{G_{mn,kl}(\tau)}\right)$, $\underset{\tau\in\mathcal{T}}{\sup}\left\lVert \hat G_{k'l'}(\tau')-G_{mn,k'l'}(\tau')\right\rVert=o_p\left(\sqrt{G_{mn,k'l'}(\tau')}\right)$, and $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\sqrt{\operatorname{Var}(\alpha(\tau)) \operatorname{Var}(\alpha(\tau'))}\right) $ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\+ o_p \left ( \frac{1}{m} \sqrt{G_{mn,kl} (\tau) }\sqrt{G_{mn,k'l'}(\tau') }\sqrt{\operatorname{Var}(\alpha(\tau)) \operatorname{Var}(\alpha(\tau'))} \right) . \end{align*} Noting that for $k = \{1, \dots, K_1\}$, $k' = \{1, \dots, K_1\}$ \begin{align*} \zeta(k, \tau )\zeta(k', \tau' ) = & \frac{1}{mn} + \frac{Var(\alpha(\tau))}{m} \left( \sum_{L_1 +1}^L \sqrt{G_{mn,kl}(\tau)} \right )\left( \sum_{L_1 +1}^L \sqrt{G_{mn,k'l'}(\tau')} \right ) + \frac{1}{ m } \sqrt{\frac{\operatorname{Var}(\alpha(\tau))}{n}} \sum_{L_1 +1}^L \sqrt{G_{mn, kl}(\tau)} \\ &+ \frac{1}{ m } \sqrt{\frac{\operatorname{Var}(\alpha(\tau'))}{n}} \sum_{L_1 +1}^L \sqrt{G_{mn, k'l'}(\tau')} \end{align*} gives that for $k = k' \in \{1, \dots , K_1\}$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\+ o_p \left ( \zeta(k, \tau )\zeta(k, \tau ) \right) \end{align*} Next, let $k = k \in \{1, \dots , K_1\}$ and $k' \in \{K_1+1, \dots , K\}$ For $l = l' \in \{1, \dots , L_1 \}$, we have by Lemma (ref), uniformly in $\tau$, $\hat G_{kl}(\tau)-G_{mn,kl}(\tau)=o_p(1)$, $\hat G_{k'l}(\tau)-G_{mn,k'l}(\tau)=o_p(1/\sqrt{a_n(\tau)})$ where both $G_{mn,k'l}$ and $G_{mn,kl}$ are $O_p(1)$. Further, $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\frac{1}{n} \right)$. Hence, \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') + o_p \left ( \frac{1}{m n \sqrt{a_n}} \right) \end{align*} For $l = \in \{1, \dots , L_1 \}$ and $ l' \in \{L_1 + 1, \dots , L \}$, we have $\hat G_{kl}(\tau)-G_{mn,kl}(\tau)=o_p(1)$, $\hat G_{k'l'}(\tau)-G_{mn,k'l'}(\tau)=o_p(1)$, with $G_{mn,kl} = O_p(1)$ and $G_{mn,k'l'} = O_p(1)$ and $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\sqrt{\frac{ {\operatorname{Var}(\alpha(\tau'))}}{n}} \right)$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\ + o_p \left ( \frac{1}{m } \sqrt{\frac{ {\operatorname{Var}(\alpha(\tau'))}}{n}}\right) \end{align*} For $l = \in \{L_1 + 1, \dots , L \}$ and $ l' \in \{1, \dots , L_1 \}$, we have $\hat G_{kl}(\tau)-G_{mn,kl}(\tau)=o_p\left(\sqrt{G_{kl}(\tau)}\right )$, $\hat G_{k'l'}(\tau)-G_{mn,k'l'}(\tau)=o_p(1 / \sqrt{a_n})$ with $G_{kl}(\tau) = O_p(a_n)$ and $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\sqrt{\frac{ {\operatorname{Var}(\alpha(\tau))}}{n}} \right)$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\ + o_p \left ( \frac{1}{m \sqrt{a_n} } \sqrt{\frac{ {\operatorname{Var}(\alpha(\tau)) }G_{mn,kl}}{n}}\right) \end{align*} For $l = \in \{L_1 + 1, \dots , L \}$ and $ l' \in \{L_1 + 1, \dots , L \}$, we have $\hat G_{kl}(\tau)-G_{mn,kl}(\tau)=o_p\left(\sqrt{G_{kl}(\tau)}\right )$, $\hat G_{k'l'}(\tau)-G_{mn,k'l'}(\tau)=o_p(1)$ with $G_{kl}(\tau) = O_p(a_n)$ and $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\sqrt{ \operatorname{Var}(\alpha(\tau))\operatorname{Var}(\alpha(\tau'))} \right)$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\ + o_p \left ( \frac{1}{m }\sqrt{ \operatorname{Var}(\alpha(\tau))\operatorname{Var}(\alpha(\tau'))}\sqrt{G_{kl}(\tau)} \right) \end{align*} Then noting that \begin{align*} \zeta(k, \tau) \zeta(k', \tau') =& \frac{1}{mn \sqrt{a_n}} + \frac{1}{m\sqrt{n}} \sqrt{\operatorname{Var}(\alpha(\tau')} \sum_{l = L_1 + 1}^L \sqrt{G_{mn, k'l}} + \frac{1}{m \sqrt{n a_n}} \sqrt{Var(\alpha(\tau) } \sum_{l = L_1 + 1}^L \sqrt{G_{mn, kl}} \\ &+ \frac{\sqrt{Var(\alpha(\tau) \operatorname{Var}(\alpha(\tau')}}{m} \sum_{l = L_1 + 1}^L \sqrt{G_{mn, kl}} \sum_{l = L_1 + 1}^L \sqrt{G_{mn, k'l}} \end{align*} It follows that for $k = k \in \{1, \dots , K_1\}$ and $k' \in \{K_1+1, \dots , K\}$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\+ o_p \left ( \zeta(k, \tau )\zeta(k', \tau' ) \right) \end{align*} By the same line of argument it follows that for $k = k \in \{K_1+1, \dots , K\}$ and $k' \in \{1, \dots , K_1\}$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\+ o_p \left ( \zeta(k, \tau )\zeta(k', \tau' ) \right) \end{align*} \\ Hence, it only remains to check the entries for $k \in \{K_1+1, \dots , K\}$, $k' = \in \{K_1+1, \dots , K\}$ For $l l' \in \{1, \dots , L_1 \}, l' \in \{1, \dots , L_1 \}$, we have by Lemma (ref), uniformly in $\tau$, $\hat G_{kl}(\tau)-G_{mn,kl}(\tau)=o_p(1/\sqrt{a_n(\tau))}$, $\hat G_{k'l}(\tau)-G_{mn,k'l}(\tau)=o_p(1/\sqrt{a_n(\tau')})$ where both $G_{mn,k'l}$ and $G_{mn,kl}$ are $O_p(1)$. Further, $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\frac{1}{n} \right)$. Hence, \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\+ o_p \left ( \frac{1}{mn \sqrt{a_n(\tau) a_n(\tau')} } \right) \end{align*} For $l' \in \{1, \dots , L_1 \}$, $l' \in \{L_1 +1, \dots, L \}$ we have by Lemma (ref), uniformly in $\tau$, $\hat G_{kl}(\tau)-G_{mn,kl}(\tau)=o_p(1/\sqrt{a_n(\tau))}$, $\hat G_{k'l'}(\tau)-G_{mn,k'l'}(\tau)=o_p(1)$ where both $G_{mn,kl}$ and $G_{mn,k'l'}$ are $O_p(1)$. Further, $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\sqrt{\frac{\operatorname{Var}(\alpha(\tau')}{n}}\right)$. Hence, \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') + o_p \left ( \frac{1}{m \sqrt{na_n(\tau) }} \sqrt{\operatorname{Var}(\alpha(\tau'))} \right) \end{align*} Similarly, for $l \in \{L_1 +1, \dots, L \}$ and $l' \in \{1, \dots , L_1 \}$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') + o_p \left ( \frac{1}{m \sqrt{na_n(\tau') }} \sqrt{\operatorname{Var}(\alpha(\tau))} \right) \end{align*} Finally, for $l \in \{L_1 +1, \dots, L \}$ and $l' \in \{L_1 +1, \dots, L \}$, we have $\hat G_{kl}(\tau)-G_{mn,kl}(\tau)=o_p(1)$ and $\hat G_{k'l'}(\tau)-G_{mn,k'l'}(\tau)=o_p(1)$ and $\hat \Omega_{ll'}(\tau, \tau') - \Omega_{ll'}(\tau, \tau') = o_p \left (\sqrt{ \operatorname{Var}(\alpha(\tau))\operatorname{Var}(\alpha(\tau'))} \right)$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') + o_p \left ( \frac{1}{m } \sqrt{ \operatorname{Var}(\alpha(\tau))\operatorname{Var}(\alpha(\tau'))}\right) \end{align*} Then noting that for $k \in \{K_1+1, \dots , K\}$, $k' = \in \{K_1+1, \dots , K\}$ \begin{align*} \zeta(k, \tau) \zeta(k', \tau') = & \frac{1}{mn\sqrt{a_n(\tau) a_n(\tau')}} + \frac{\sqrt{\operatorname{Var}(\alpha(\tau))\operatorname{Var}(\alpha(\tau'))}}{m} \left( \sum_{l = L_1}^L \sqrt{G_{mn,kl(\tau) } }\right ) \left (\sum_{l = L_1}^L \sqrt{G_{mn,k'l(\tau') } }\right) \\ & + \frac{1}{m\sqrt{n a_n(\tau) }}\sqrt{\operatorname{Var}(\alpha(\tau'))} \sum_{l = L_1}^L \sqrt{G_{mn,k'l'(\tau') } } \\ & + \frac{1}{m\sqrt{n a_n(\tau') }}\sqrt{\operatorname{Var}(\alpha(\tau))} \sum_{l = L_1}^L \sqrt{G_{mn,kl(\tau) } } \end{align*} gives for the $k, k'$ entry with $k \in \{K_1+1, \dots , K\}$, $k' = \in \{K_1+1, \dots , K\}$ \begin{align*} \frac{1}{m} \sum_{l} \sum_{l'} \hat G_{kl}(\tau) \hat \Omega_{ll'}(\tau, \tau') \hat G_{k'l'}(\tau') = \frac{1}{m}\sum_{l} \sum_{l'} G_{kl}(\tau) \Omega_{ll'}(\tau, \tau') G_{k'l'}(\tau') \\+ o_p \left ( \zeta(k, \tau )\zeta(k', \tau' ) \right) \end{align*} From proposition (ref), we have that the diagonal elements of $\Sigma_{mn}$ are consistently estimated. Differently, the off-diagonal elements are not consistently estimated when it involves the covariance between two coefficients converging at different rates. However, where these off-diagonal terms are not consistently estimated, they are of a lower order than the corresponding diagonal term; so that for any vector $\eta \in \mathbb{R}^{T\times K}$ with $||\eta|| > \varepsilon > 0$ \begin{equation*} \frac{\eta'\hat \Sigma_{mn} \eta }{\eta' \Sigma_{mn} \eta } \xrightarrow{p}1. \end{equation*} From the proof of Theorem (ref), we know that $$ \Sigma_{mn,kk}^{-1/2}\left(\hat \delta_k(\tau, \tau') - \delta_k(\tau, \tau') \right) \xrightarrow{d} N(0,1).$$ Hence, it is straightforward to show that \begin{equation*} \frac{\eta' \left (\hat \delta(\tau, \tau') - \delta(\tau, \tau') \right) }{ \sqrt{\eta' \hat \Sigma_{mn} \eta }} \xrightarrow{d} N(0,1). \end{equation*} \end{comment}
comment\begin{scriptsize} \begin{align*} \eta'\hat H \eta = & {\eta' H \eta + o_P(\eta' \tilde \zeta(\tau, \tau') \tilde \zeta( \tau,\tau') \eta ) } \\ =& \eta' H \eta + \eta' \begin{pmatrix} o_p \left( \sqrt{\parallel G_{mn,1l}(\tau)\Omega_{ll}(\tau) \parallel \parallel \Omega_{l'l'}(\tau) G_{mn,1l'}(\tau)' \parallel }\right) & \dots & o_p \left( \sqrt{\parallel G_{mn,1l}(\tau)\Omega_{ll}(\tau) \parallel \parallel \Omega_{l'l'}(\tau') G_{mn,2Kl'}(\tau')' \parallel }\right) \\ \vdots & \vdots & \vdots\\ o_p \left( \sqrt{ \parallel G_{mn,2Kl}(\tau')\Omega_{ll}(\tau') \parallel \parallel \Omega_{l'l'}(\tau') G_{mn,1l'}(\tau)' \parallel } \right)& \dots & o_p \left( \sqrt{\parallel G_{mn,2Kl}(\tau')\Omega_{ll}(\tau') \parallel \parallel \Omega_{l'l'}(\tau') G_{mn,2Kl'}(\tau')' \parallel} \right) \end{pmatrix}\eta \\ & = \sum_{k} \sum_{k'} \eta_k H_{kk'} \eta_{k'} + \sum_k \sum_{k'} \eta_k o_p \left(\sqrt{ \parallel G_{mn,kl}\Omega_{ll} \parallel \parallel \Omega_{l'l'} G_{mn,k'l'} \parallel } \right) \end{align*} \end{scriptsize} \\ Hence, where $o_p \left ( \frac{\sqrt{\parallel \Omega_{22}(\tau, \tau') \parallel}}{\sqrt n}\right) = o_p \left ( {{\parallel \Omega_{22}(\tau, \tau') \parallel}}\right)$. \begin{commentP} \begin{align*} \hat G \hat \Omega \hat G = \begin{pmatrix} \hat G_{11} \hat \Omega_{11} \hat G_{11}' + \hat G_{12} \hat \Omega_{21} \hat G_{11}' + \hat G_{11} \hat \Omega_{12} \hat G_{12}' + \hat G_{12} \hat \Omega_{22} \hat G_{[12]'} & \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} + \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} + \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} + \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} \\ \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} + \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} + \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} + \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} & \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} + \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} + \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} + \hat G_{[]} \hat \Omega_{[]} \hat G_{[]} \end{pmatrix} \end{align*} Upper left \begin{align*} \hat G_{11} \hat \Omega_{11} \hat G_{11}' + \hat G_{12} \hat \Omega_{21} \hat G_{11}' + \hat G_{11} \hat \Omega_{12} \hat G_{12}' + \hat G_{12} \hat \Omega_{22} \hat G_{[12]'} = \end{align*} \begin{align*} \hat G_{11} \hat \Omega_{11} \hat G_{11}' = G_{mn,11} \Omega_{11} G_{mn,11}' + o_p \left (\frac{1 }{n}\right) \end{align*} \begin{align*} \hat G_{12} \hat \Omega_{21} \hat G_{11}' = G_{mn,12} \Omega_{12} G_{mn,11}' + o_p \left (\frac{1 }{n}\right) \end{align*} \begin{align*} \hat G_{12} \hat \Omega_{22} \hat G_{12}' = G_{mn,12} \Omega_{22} G_{mn,12}'+ o_p \left (\frac{1 }{n}\right) \end{align*} Terms in the lower left \begin{align*} \hat G_{21} \hat \Omega_{11} \hat G_{11}' = G_{mn,21} \Omega_{11} G_{mn,11}' + o_p \left ( \frac{1}{n}\right) \end{align*} \begin{align*} \hat G_{22} \hat \Omega_{21} \hat G_{11}' = G_{mn,22} \Omega_{21} G_{mn,11}' + o_p \left ( \frac{\sqrt{\parallel \Omega_{22} \parallel}}{\sqrt n}\right) \end{align*} \begin{align*} \hat G_{21} \hat \Omega_{12} \hat G_{12}' = G_{mn,21} \Omega_{12} G_{mn,12}' + o_p \left ( \frac{1}{n}\right) \end{align*} \begin{align*} \hat G_{22} \hat \Omega_{22} \hat G_{12'} = G_{mn,22} \Omega_{22} G_{mn,12}' +o_p \left ( \frac{\sqrt{\parallel \Omega_{22} \parallel}}{\sqrt n}\right) \end{align*} Lower right \begin{align*} \hat G_{21} \hat \Omega_{11} \hat G_{21}' = G_{21} \Omega_{11} G_{21}' + o_p \left( \frac{1}{n} \right) \end{align*} \begin{align*} \hat G_{22} \hat \Omega_{21} \hat G_{21}' = G_{22} \Omega_{21} G_{21}' + o_p \left ( \frac{\sqrt{\parallel \Omega_{22} \parallel}}{\sqrt n}\right) \end{align*} \begin{align*} \hat G_{22} \hat \Omega_{21} \hat G_{21}' = G_{22} \Omega_{21} G_{21}' + o_p \left ( \frac{\sqrt{\parallel \Omega_{22} \parallel}}{\sqrt n}\right) \end{align*} \begin{align*} \hat G_{22} \hat \Omega_{22} \hat G_{22}' = G_{22} \Omega_{22} G_{22}' + o_p \left ( {{\parallel \Omega_{22} \parallel}}\right) \end{align*} \end{commentP} Hence, we cannot consistently estimate the covariance between the two coefficients, If the null hypothesis only contains $K_1$ or only $K_2$ type coefficients, the results are trivial. If the hypothesis contains both types of coefficients,
comment\hline Without loss of generality, consider the scalar case (not sure because in the scalar case, then we don't have overidentification). \begin{equation*} \hat G = b(S_{ZW}, \hat W) + f(S_{ZW}, \hat W) \cdot a_n + h(S_{ZW}, \hat W) \cdot a_n^2 \end{equation*} where $b(\cdot)$, $f(\cdot)$ and $h(\cdot)$ are continuous function. Define \begin{equation*} G_{mn} = b(\Sigma_{ZW}, W) + f(\Sigma_{ZW}, W) \cdot a_n + h(\Sigma_{ZW}, W) \cdot a_n^2 \end{equation*} By assumption (ref) we have that $\hat W \xrightarrow{p} W$ and by the law of large number and assumption (ref) $S_{ZX} \xrightarrow{p} \Sigma_{ZX}$. The continous mapping theorem thus implies that $$ f(S_{ZW}, \hat W) = f(\Sigma_{ZW}, W) + o_p(1)$$ $$ h(S_{ZW}, \hat W) = h(\Sigma_{ZW}, W) + o_p(1).$$ $$ b(S_{ZW}, \hat W) = b(\Sigma_{ZW}, W) + o_p(1).$$ Where all the functions are finite. It the follows that \begin{align*} \hat G = & b(S_{ZW}, \hat W) + f(\Sigma_{ZW}, W) \cdot a_n + o_p(1) \cdot a_n + h(\Sigma_{ZW}, W) \cdot a_n^2 + o_p(1) \cdot a_n^2 \\ =& b(S_{ZW}, \hat W) + f(\Sigma_{ZW}, W) \cdot a_n + h(\Sigma_{ZW}, W) \cdot a_n^2 + o_p(a_n) \\ \end{align*} There are two possible scenarios. Either $b(S_{ZX}, \hat W) = b(\Sigma_{ZW}, W) = 0$ or $b(S_{ZX}, \hat W) > \varepsilon > 0$. In the first case $G_{mn}$ is $O_p(a_n)$. While in the second case $G_{mn}$ is $O_p(1)$. \\ Case 1 - $b(S_{ZX}, \hat W) = 0$: \begin{align*} \hat G = & f(\Sigma_{ZW}, W) \cdot a_n + h(\Sigma_{ZW}, W) \cdot a_n^2 + o_p(a_n) \\ =& G_{mn} + o_p(a_n)\\ =& G_{mn} + o_p(G_{mn}) \end{align*} Case 2 - $b(S_{ZX}, \hat W) > \varepsilon > 0$: \begin{align*} \hat G = & b(\Sigma_{ZW}, W) + o_p(1) + f(\Sigma_{ZW}, W) \cdot a_n + h(\Sigma_{ZW}, W) \cdot a_n^2 + o_p(a_n) \\ =& G_{mn} + o_p(1)\\ =& G_{mn} + o_p(G_{mn})\\ \end{align*}
comment\color{gray} Since the rate of convergence may differ across the quantile index and the regressors, we first consider the scalar $\hat\delta_k(\tau)$, which is the $k$-th element of $\hat\delta(\tau)$ for $k\in\{1,\dots,K\}$ and $\tau \in\mathcal{T}$. From Lemma (ref), \begin{equation*} \hat\delta_k(\tau)-\delta_k(\tau)=\hat G_k(\tau)\bar g_{mn}(\hat\delta,\tau)=\hat G_k(\tau)\left(\bar g_{mn}^{(1)}(\hat\delta,\tau)+\bar g_{mn}^{(2)}(\hat\delta,\tau)\right) \end{equation*} where $\hat G_k(\tau)$ is the $k$-th row of $\hat G(\tau)$. From the proof of Lemma (ref)(i), we have \begin{equation*} \bar g_{mn}^{(1)}(\hat\delta,\tau)=\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+o_p\left(\frac{1}{\sqrt{mn}}\right) \end{equation*} uniformly in $\tau\in\mathcal T$. Together with the uniform boundedness of $\hat G(\tau)$, it implies that \begin{align*} \hat G_k(\tau)\bar g_{mn}^{(1)}(\hat\delta,\tau)&=\hat G_k(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})+o_p\left(\frac{1}{\sqrt{mn}}\right) \end{align*} uniformly in $\tau\in\mathcal T$. In addition, Lemma 2(i) implies that $\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})=O_p\left(\frac{1}{\sqrt{mn}}\right)$ uniformly in $\tau\in\mathcal T$. From the proof of Lemma (ref), we have $\underset{\tau\in\mathcal{T}}{\sup}\left\lVert \hat G_k(\tau)-G_k(\tau)\right\rVert=O_p\left(\frac{1}{\sqrt{m}}\right)$. \color{red} DROP. \begin{equation*} \underset{\tau\in\mathcal{T}}{\sup}\left\lVert \left(S_{ZX}'\hat W(\tau) S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau)-\begin{pmatrix}G_{11}(\tau) & 0 \\ G_{21}(\tau) & G_{22}(\tau)\end{pmatrix} \right\rVert=\begin{pmatrix} O_p \left (\frac{a_n^{1/2}}{\sqrt{m}} \right) + O_p\left( \zeta_1 \right) & O_p \left (\frac{a_n^{1/2}}{\sqrt{m}} \right) \\ O_p\left (\frac{1}{\sqrt{m}} \right) & O_p \left (\frac{1}{\sqrt{m}} \right) \end{pmatrix} \end{equation*} Check the term on the top-left! Doesn't matter much. The central element is the top-right. We could only prove that this element it $O_p(a_n / \sqrt{m})$ \color{black} This, in turn, implies that \begin{equation*} \left(\hat G_k(\tau)-G_k(\tau)\right)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})=o_p\left(\frac{1}{\sqrt{mn}}\right) \end{equation*} uniformly in $\tau\in\mathcal T$. Combining the last two displayed equations, we obtain \begin{align} \hat G_k(\tau)\bar g_{mn}^{(1)}(\hat\delta,\tau)=G_k(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+o_p\left(\frac{1}{\sqrt{mn}}\right) \end{align} uniformly in $\tau\in\mathcal T$. Similarly for $\bar g_{mn}^{(2)}(\delta,\tau)$, we have Note that the first $L_1$ of $\bar g_{mn}^{(2)}(\hat\delta,\tau)$ equal to zero. While for the last $L_2$ elements, we have \begin{equation} \bar g_{mn}^{(2)}(\hat\delta,\tau)_{L_2} = O_p \left ( \frac{|| \Omega_{2(L_2 \times L_2)}||^{1/2}}{\sqrt{m}} \right) \end{equation} where $\Omega_{2(L_2 \times L_2)}$ denote the lower-right $L_2 \times L_2$ submatrix of $\Omega_2$. Keeping track of all terms in proof of Lemma (ref) and without assuming that $a_n(\tau) = o_p(1)$, we find that the upper right term of the matrix $\hat G(\tau)$ converges to $G(\tau)_{12}$ at the rate $O_p \left (\frac{a_n(\tau)^{1/2}}{\sqrt{m}} \right)$. \color{red} We don't need to specify that $a_n= o_p(1)$. We need the following result (We need it later anyway): \begin{equation*} \underset{\tau\in\mathcal{T}}{\sup}\left\lVert \left(S_{ZX}'\hat W(\tau) S_{ZX}\right)^{-1}S_{ZX}'\hat W(\tau)-\begin{pmatrix}G_{11}(\tau) & G_{12}(\tau)\\ G_{21}(\tau) & G_{22}(\tau)\end{pmatrix} \right\rVert= o_p(1) \end{equation*} Where $G_{12}(\tau) = O_p(\sqrt{a_n})$ \color{black} \begin{align*} \left(\hat G_k(\tau)-G_k(\tau)\right)\bar g_{mn}^{(2)}(\hat\delta,\tau) = o_p \left (a_k^{1/2} \cdot \frac{|| \Omega_{2(L_2 \times L_2)}||^{1/2}}{\sqrt{m}} \right) \end{align*} uniformly in $\tau\in\mathcal T$ where $a_k$ equals $a_n$ for all $k \leq K_1$ and $1$ otherwise. It implies that \begin{align*} \hat G_k(\tau)\bar g_{mn}^{(2)}(\hat\delta,\tau) = G_k(\tau)\frac{1}{mn}\sum_{j=1}^m\sum_{i=1}^n z_{ij}\alpha_j(\tau) + o_p \left (a_k^{1/2} \cdot \frac{|| \Omega_{2(L_2 \times L_2)}||^{1/2}}{\sqrt{m}} \right) \end{align*} again uniformly in $\tau\in\mathcal T$. Write \begin{equation} \zeta(k,\tau)=\frac{1}{\sqrt{mn}}+\frac{1}{\sqrt m}\left\lVert a_k\cdot \Omega_{2(L_2 \times L_2)}\right\rVert ^{1/2}. \end{equation} Thus, uniformly in $\tau\in\mathcal T$, \begin{align*} \hat\delta_k(\tau)-\delta_k(\tau)=&\hat G_k(\tau)\bar g^{(1)}_{mn}(\hat\delta,\tau)+\hat G_k(\tau)\bar g^{(2)}_{mn}(\hat\delta,\tau)\\ =&G_k(\tau)\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+o_p\left(\frac{1}{\sqrt{mn}}\right)\\ &+G_k(\tau)\frac{1}{m}\sum_{j=1}^m \bar z_{j}\alpha_j(\tau)+ o_p \left (\frac{||a_k \cdot \Omega_{2(L_2 \times L_2)}||^{1/2}}{\sqrt{m}} \right)\\ =&G_k(\tau)\left(\frac{1}{m}\sum_{j = 1}^m \Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+\frac{1}{m}\sum_{j=1}^m \bar z_{j}\alpha_j(\tau)\right)+o_p\left(\zeta(k,\tau)\right) \end{align*} where the second equality follows equations (ref) and (ref) and the third from the definition $\zeta$ in (ref). Thus, we have \begin{equation*} \hat\delta_k(\tau)-\delta_k(\tau)=\sum_{j=1}^m d_j(k,\tau)+o_p\left(\zeta(k,\tau)\right) \end{equation*} where \begin{equation*} d_j(k,\tau)=G_k(\tau)\left(\frac{1}{mn}\Sigma_{ZXj}\left(\frac{1}{n}\sum_{i = 1}^n\phi_{j,\tau}(\tilde x_{ij},y_{ij})\right)+\frac{1}{m}\bar z_{j}\alpha_j(\tau)\right) \end{equation*} Let $D_j$ be the $TK\times 1$ vector $(\operatorname{diag}(\Sigma_{mn}(\tau_1))^{-1/2} d_j(\tau_1),\dots, \operatorname{diag}(\Sigma_{mn}(\tau_T))^{-1/2} d_j(\tau_T))'$ where $d_j(\tau)=(d_j(1,\tau),d_j(2,\tau),\dots,d_j(K,\tau))'$. It follows that \begin{equation*} \operatorname{Var}\left(\sum_{j=1}^m D_i\right)=H_{mn} \end{equation*} Then, it follows from the proof of Lemma (ref) and Assumption (ref) \begin{equation*} H_{mn}^{-1/2}\sum_{j=1}^mD_i\underset{d}{\rightarrow}N(0,I_{TK}) \end{equation*} By Slutsky's theorem, \begin{equation*} H^{-1/2}\sum_{j=1}^mD_i=H_{mn}^{-1/2}\sum_{j=1}^mD_i+o_p(1)\underset{d}{\rightarrow}N(0,I_{TK}) \end{equation*} In the proof of Lemma (ref) we show that $\sum_j \operatorname{diag}(\Sigma_{mn}(\cdot))^{-1/2} d_j(\cdot))$ is asymptotically tight in $l^\infty(\mathcal{T})$. It follows that the process $\sum_j \operatorname{diag}(\Sigma_{mn}(\cdot))^{-1/2} d_j(\cdot))$ weakly converges to $\mathbb Z$, a centered Gaussian process with covariance kernel $H(\tau,\tau')$. Remarks: The rate of convergence $\zeta$ makes sense because it account for two things: (i) degree of heterogeneity in the slow moment $g_2$. It makes sense that the lower right matrix only matters because it is the one associated to the slow instruments but note that $|| \Omega_2 || = || \Omega_{2(L_2 \times L_2)}||$. (ii) It accounts for the relative weights given to the slow moment relative to the fast moments. Finally, we show that $\underset{\tau\in\mathcal T, k\in\{1,\dots,K\}}{\sup}\zeta(k,\tau)\Sigma_{mn}(k,\tau)^{-1/2}=o(1)$, where $\Sigma_{mn}(k,\tau)$ is the $(k,k)$ element of $\Sigma_{mn}(\tau)$. We have \begin{equation*} o\left(\frac{1}{\sqrt{mn}}\right)\frac{1}{\inf_{\tau,k}\Sigma_{mn}(k,\tau)^{1/2}}=\frac{1}{\inf_{\tau,k} (G_k(\tau)\Omega_1(\tau)G_k(\tau)')^{1/2}}o(1)=o(1) \end{equation*} \begin{equation*} o\left(\frac{1}{\sqrt m}\right)\sup_{\tau,k}\left\lVert {a_k}(\tau) \cdot \Omega_2(\tau)_{(L_2 \times L_2)}\right\rVert ^{1/2}\Sigma_{mn}^{-1/2}\leq o(1) \sup_{\tau,k}\left(\frac{ \left ( {a_k}(\tau) \cdot \Omega_2(\tau)_{(L_2 \times L_2)}\right ) ^{1/2}}{(G_k(\tau)\Omega(\tau) G_k(\tau)') ^{1/2}}\right)=o(1) \end{equation*} since all elements in $G_k(\tau)$ are either finite or decrease at the rate $\sqrt{a_k}.$ \color{red} But for $k = 1$ (fast elements) $G_{12} = 0$ so that $G_1 \Omega_2 G_1' = G_{11}' \Omega_{2(L_1\times L_1)}G_{11} = 0$. There is still $\Omega_1/a_n. \\ for $k = 2$ (slow elements) Note that : $G_k(\tau)\Omega(\tau) G_k(\tau)' = G_{11} \Omega_{11} G_{11} +G_{12} \Omega_{12} G_{11} + G_{11} \Omega_{12} G_{21} + G_{12} \Omega_{12} G_{12}$ \color{black} Hence, uniformly in $\tau, \begin{align*} \operatorname{diag}(\Sigma_{mn}(\tau))^{-1/2}(\hat\delta(\tau)-\delta(\tau))&=\operatorname{diag}(\Sigma_{mn}(\tau))^{-1/2}\left(\sum_{j=1}^m d_j(\tau)+o_p\left(\zeta(\tau)\right)\right)\\