Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
93,786 characters · 18 sections · 68 citation commands
A quantile-based nonadditive fixed effects model
\justifying
\doublespacing
Quantile models can be interpreted as random coefficient models in a cross-sectional setting Koenker2005. Consider the standard structural quantile model
with random variable $U_i \sim \textrm{Unif}(0,1)$ and the monotonicity condition that $\boldsymbol{\mathbf{x}}' \boldsymbol{\mathbf{\beta}}(u)$ is increasing in $u$ for all possible non-random $\boldsymbol{\mathbf{x}}$ values. Given the exogeneity assumption $U \protect\mathpalette{\protect\independenT}{\perp} \boldsymbol{\mathbf{X}}$, this random coefficient model leads to linear conditional quantile function
which can be consistently estimated by KoenkerBassett1978. When there is endogeneity, meaning the $\boldsymbol{\mathbf{X}}$ and $U$ are statistically dependent, the conditional quantile slope function $\boldsymbol{\mathbf{\beta}}_{\mathrm{CQF} }(\cdot)$ in (ref) is different from the structural quantile slope function $\boldsymbol{\mathbf{\beta}}(\cdot)$ in (ref). \Citet{ChernozhukovHansen2005} use instruments to identify the structural quantile effect in (ref) in a cross-sectional setting. With panel data, I can relax the independence restriction $U \protect\mathpalette{\protect\independenT}{\perp} \boldsymbol{\mathbf{X}}$ to allow unrestricted dependence between $U$ and $\boldsymbol{\mathbf{X}}$ without using instruments. I extend (ref) to a panel model and aim to identify and estimate the heterogeneous causal effect function $\boldsymbol{\mathbf{\beta}}(\cdot)$.
In this paper, I consider the model
where $U_i \sim \textrm{Unif}(0,1)$ is the individual-specific time-invariant unobserved heterogeneity (i.e., “fixed effect”), $V_{it}$ is the idiosyncratic error which will be discussed soon in (ref), and $U_i$ and the $\boldsymbol{\mathbf{X}}_{it}$ are allowed to be arbitrarily dependent. Due to the endogeneity of $\boldsymbol{\mathbf{X}}_{it}$ and the additive $V_{it}$, the structural function $\boldsymbol{\mathbf{X}}_{it}'\boldsymbol{\mathbf{\beta}}(U_i)$ in (ref) does not generally correspond to any conditional quantile function. Thus, the identification is more complex than showing equivalence to some quantile regression with the right control. It is important to consider the structural models like (ref), because the structural function matters for policy and counterfactual analysis, instead of the conditional quantile function that has only descriptive and predictive meanings without further assumptions.
This model (ref) connects to both the standard fixed effects (FE) model and FE quantile regression (FE-QR) model for the reasons described in detail in (ref). Additionally, similar to HausmanEtAl2021, the $V_{it}$ could be considered as measurement error, and the $\boldsymbol{\mathbf{\beta}}(\tau)$ has quantile-related interpretation as discussed in (ref).
I make four primary contributions. First, I propose a new quantile-based nonadditive fixed effects model (ref). It considers the heterogeneous causal effect as a function of rank variable. I propose the rank variable to be time-stable, for the reasons described later in (ref).
Second, I identify the heterogeneous causal effects in (ref). The structural coefficients measure the causal effects of explanatory variables vector $\boldsymbol{\mathbf{X}}$ on the outcome $Y$ for individuals in the population with rank variable value $U=\tau$. The model deals with endogeneity, meaning the arbitrary dependence between regressors and the unobserved heterogeneity (rank variable). Identification is achieved by assuming endogeneity is due only to fixed effects (not idiosyncratic error). If the stronger monotonicity assumption of ChernozhukovHansen2005 is imposed, then the structural parameter of interest fully describes the corresponding quantile structural functions and the related quantile treatment effects.
Third, I establish uniform consistency of the coefficient function estimators over all individual rank values under a certain rate relation between $n$ and $T$ (the cross-sectional and time series dimensions of the data). I first estimate the individual-specific coefficients, relying on established time series OLS consistency results. Then I implement a sorting approach on a hypothetical outcome at certain $\boldsymbol{\mathbf{x}}^{*}$ with the estimated individual coefficients. Under a certain rate condition between $n$ and $T$, the order of the rank variables can be recovered through the order of the hypothetical outcomes. Therefore, the estimator of coefficient at any specific rank value is determined and consistently estimated.
Finally, I establish the asymptotic normality of the coefficient function estimator uniformly over all rank values over a proper subset of the unit interval $[0,1]$, by applying the functional delta-method to the empirical quantile process.
Although my model is a long panel model, a simulation is provided to show the performance of my estimator in the short-$T$ case (for example, $T=3, 4, \ldots, 10$). The simulation result is compatible with consistency. The simulation also investigates the estimator's performance with various $n/T$ rate relations, which are weaker than the sufficient rate assumed for large-sample theory. I also illustrate the performance of my estimator with various sorting point and compare with the FE-QR estimator and standard FE estimator.
\paragraph{Literature}
My model is a fixed effects (FE) approach because it allows arbitrary dependence between $\boldsymbol{\mathbf{X}}$ and $U$. The fact that $U$ follows a standard uniform distribution is a normalization, rather than a restriction. We can think of $U_i$ as a normalization of the individual fixed effect in the standard FE model (ref), i.e., $U_i=F_\alpha(\alpha_i)$, where $F_\alpha(\cdot)$ is the CDF of $\alpha_i$. As noted by Wooldridge2010, the key difference between the FE approach and correlated random effects (CRE) approach is that the FE approach allows an entirely unspecified relationship between $\alpha_i$ and $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{X}}\mkern-3mu}\mkern3mu_i=(\boldsymbol{\mathbf{X}}_{i1},\ldots, \boldsymbol{\mathbf{X}}_{iT})$, whereas the CRE approach restricts the dependence between $\alpha_i$ and $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{X}}\mkern-3mu}\mkern3mu_i$ in a substantive way, like restricting the distribution of $\alpha_i$ given $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{X}}\mkern-3mu}\mkern3mu_i$. The normalization of $\alpha_i$ alone does not affect the dependence relationship flexibility between $\alpha_i$ and $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{X}}\mkern-3mu}\mkern3mu_i$. Therefore, the arbitrary dependence between $\boldsymbol{\mathbf{X}}$ and $U$ with normalized $U$ still generally belongs to the FE approach.
My model is similar in spirit to the structural random coefficient interpretation of QR Koenker2005, instrumental variables quantile regression ChernozhukovHansen2005,ChernozhukovHansen2006,ChernozhukovHansen2008, and FE-QR Canay2011. In random coefficient form, my model uses the FE approach, rather than the CRE approach that includes Chamberlain1982 for panel models and AbrevayaDahl2008, ArellanoBonhomme2016, and HardingLamarche2017 for panel quantile models. \Citet{GalvaoKato2017} provide a survey about using the FE and CRE approaches in panel quantile models.
Like the standard fixed effects model, my model is a structural model. The FE-QR in Koenker2004 and KatoEtAl2012 are statistical models that summarize conditional distributions of the outcome, but require additional assumptions for the coefficient to have a structural interpretation. In general, my model does not directly connect with these FE-QR models' conditional quantile function (CQF). Only when $V_{it}=0$, $u\mapsto \boldsymbol{\mathbf{x}}'\boldsymbol{\mathbf{\beta}}(u)$ is strictly increasing at all $\boldsymbol{\mathbf{x}}$ values, and $\boldsymbol{\mathbf{X}}$ is exogenous, i.e., $\boldsymbol{\mathbf{X}}$ is independent of $U$, then at time $t$, the cross-sectional $\tau$-CQF of $Y_{it}$ given $\boldsymbol{\mathbf{X}}_{it}=\boldsymbol{\mathbf{x}}$ is $\boldsymbol{\mathbf{x}}'\boldsymbol{\mathbf{\beta}}(\tau)$. My model's object of interest is not the CQF, but rather the structural effects.
Almost every FE-QR model (including the initial Koenker2004 model and later Canay2011) requires a “large-$n$ and large-$T$” framework. Alternatively, the CRE-QR model can work in a short-$T$ framework at the cost of imposing restrictions on the joint distribution between the observed regressors and the unobserved heterogeneity. My model shares the same “large-$n$ and large-$T$” flavor as in the FE-QR literature.
Compared to other random coefficient panel data models using the FE approach, my model focuses on a different estimand of interest. \Citet{GrahamPowell2012} and ArellanoBonhomme2012 contribute a random coefficient panel model using an FE approach in a fixed-$T$ and more general framework that allows multi-dimensional individual fixed effects. \Citet{GrahamPowell2012} focus on the average partial effect; ArellanoBonhomme2012 focus on $\boldsymbol{\mathbf{\beta}}_i$ instead of $\boldsymbol{\mathbf{\beta}}(U_i)$. In addition, ArellanoBonhomme2012 achieve the identification of $\boldsymbol{\mathbf{\beta}}_i$ distribution in a short-$T$ setting by assuming $U_i$ independent of $(V_{i1},\ldots, V_{iT})$ conditional on $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{X}}\mkern-3mu}\mkern3mu_{i}$, using the notation of my (ref), as well as assuming a particular MA model with independent shock on the error term $V_{it}$. My model does not make these two assumptions. Instead, my model identifies the coefficient $\boldsymbol{\mathbf{\beta}}(\cdot)$ as a function of the unobserved rank variable $U$, by making the usual assumptions in quantile models, like monotonicity, continuity, and a scalar rank variable in a large-$T$ setting.
There are some other studies on quantile effects in nonlinear panel data models. \Citet{ChernozhukovEtAl2013panel} has the advantage of considering multiple sources of heterogeneity and works for short-$T$ model, but it is limited to discrete-value regressors. My model works for large-$T$ model, but it can generally apply to both continuous-value and discrete-value regressors. ChernozhukovEtAl2013panel's (ChernozhukovEtAl2013panel) identification and estimation rely on a “time homogeneity” assumption, which is different but non-nested with my model's assumption. \Citet{Powell2022} studies panel quantile regression with nonadditive fixed effects. Both Powell2022's (Powell2022) and my paper's rank variables are not iid over time by construction. \Citet{Powell2022} can estimate consistently in small-$T$ case, but requires instruments and assumes the same strong monotonicity condition as in ChernozhukovHansen2005. Its estimation and computation of standard errors can be numerically challenging, as pointed out by BakerQREGPDStata2016. My model does not resort to instruments and identification is achieved with a weaker monotonicity assumption. The estimation and computation of my method is simple and easy.
Unlike difference-in-differences models that focus on a single binary treatment, my model applies to a variety of structural economic models with a variety of regressors, so my model complements the quantile/distributional difference-in-differences literature that includes the work of AtheyImbens2006, CallawayLi2019, Ishihara2023, and DHaultfoeuilleEtAl2023, among others.
\Citet{GrahamEtAl2018} contribute a “fixed effects” approach in a random coefficient form panel quantile model in the large-$n$ and small-$T$ framework. However, the “fixed effects” approach of GrahamEtAl2018 means the dependence between the regressors $\boldsymbol{\mathbf{X}}$ and the random coefficient $\boldsymbol{\mathbf{\beta}}$ is unrestricted, since they assume coefficient as a function of both $\boldsymbol{\mathbf{X}}$ and $U$, i.e., $\boldsymbol{\mathbf{\beta}}(U, \boldsymbol{\mathbf{X}})$; but they still restrict the conditional distribution of unobserved heterogeneity $U$ given regressors $\boldsymbol{\mathbf{X}}$, specifically $U\mid \boldsymbol{\mathbf{X}} \sim \textrm{Unif}(0,1)$, which means for any subpopulation $\boldsymbol{\mathbf{X}}=\boldsymbol{\mathbf{x}}$, $U \sim \textrm{Unif}(0,1)$, i.e., people do not select $\boldsymbol{\mathbf{X}}$ depending on $U$ and there is no endogeneity between $\boldsymbol{\mathbf{X}}$ and $U$. Alternatively, my model's fixed effects approach assumes the dependence between the regressors $\boldsymbol{\mathbf{X}}$ and the unobserved heterogeneity $U$ is unrestricted, which is in line with the fixed effects definition; and my model assumes the coefficient as a function of the unobserved heterogeneity $\boldsymbol{\mathbf{\beta}}(U)$, which results in the arbitrary dependence between regressors and coefficients. The two approaches are complementary.
\paragraph{Paper structure and notation} (ref) presents the model and identification results. (ref) defines the estimator and shows uniform consistency and the uniform asymptotic normality of the coefficient function estimator. (ref) provides simulation results. (ref) provides empirical illustrations. (ref) concludes.
Acronyms used include those for cumulative distribution function (CDF), conditional quantile function (CQF), data generating process (DGP), fixed effects (FE), ordinary least squares (OLS), mean squared error (MSE), quantile regression (QR), quantile structural function (QSF), and quantile treatment effect (QTE). Notationally, random and non-random vectors are respectively typeset as, e.g., $\boldsymbol{\mathbf{X}}$ and $\boldsymbol{\mathbf{x}}$, while random and non-random scalars are typeset as $X$ and $x$ (with the exception of certain constants like the time series dimension $T$), and random and non-random matrices as $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{X}}\mkern-3mu}\mkern3mu$ and $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{x}}\mkern-3mu}\mkern3mu$; that is, upper case denotes a random variable (like $\boldsymbol{\mathbf{X}}$ or $U$), whereas lower case denotes a non-random realized value (like $\boldsymbol{\mathbf{x}}$ or $u$). The uniform distribution is written $\textrm{Unif}(a,b)$; in some cases this stands for a random variable following such a distribution. Finally, $\rightsquigarrow$ denotes weak convergence.
Consider the structural panel model
where $Y_{it}$ is the dependent variable for individual $i$ at time $t$, $\boldsymbol{\mathbf{X}}_{it}$ is the corresponding $K$-vector of explanatory variables, $U_i$ is the individual-specific unobserved heterogeneity normalized to $U_i\sim\textrm{Unif}(0,1)$, and $V_{it}$ is the idiosyncratic disturbance. The dependence between $U_i$ and $\boldsymbol{\mathbf{X}}_{it}$ is unrestricted. The coefficient is a function of the unobserved heterogeneity. For example, $\boldsymbol{\mathbf{\beta}}(0.5)$ is the structural coefficient vector for an individual with median level unobservable $U_i=0.5$. The coefficient function $\boldsymbol{\mathbf{\beta}}(\cdot) \colon {\mathbb R} \to {\mathbb R}^K $ is deterministic but unknown. The $\boldsymbol{\mathbf{\beta}}(U_i)$ is a random coefficient vector that can vary across individuals, but the source of the randomness is restricted to $U_i$. It captures the heterogeneity in response in unobservables. In addition, another source of randomness comes from the idiosyncratic error $V_{it}$, which requires large $T$ to learn the individual $\boldsymbol{\mathbf{\beta}}(U_i)$ values.
This model connects to both the standard FE model and the FE-QR model, and it has a few interpretations.
First, model (ref) assumes a more general functional form than the standard FE model. Consider the standard FE model
and my model
where $\boldsymbol{\mathbf{X}}_{-1,it}$ is the vector of regressors without the constant term, so $\boldsymbol{\mathbf{X}}_{it}'=(1,\boldsymbol{\mathbf{X}}_{-1,it}')$. Both the $\alpha_i$ in (ref) and $U_i$ in (ref) represent individual fixed effects. The standard FE model assumes additive fixed effects and homogeneous causal effects $\boldsymbol{\mathbf{\beta}}_{-1}$, whereas my model (ref) allows the fixed effects to enter the slope nonseparably and capture the heterogeneous causal effect $\boldsymbol{\mathbf{\beta}}_{-1}(U_i)$ (as a function of the unobserved heterogeneity). When restricting $\boldsymbol{\mathbf{\beta}}(U_i)'=(\beta_0(U_i),\beta_1,\beta_2,\ldots)\equiv(\alpha_i,\boldsymbol{\mathbf{\beta}}'_{-1})$, model (ref) becomes the standard FE model (ref).
Second, model (ref) complements FE-QR models; for example, see Canay2011 and references therein. Consider an additive idiosyncratic error inside the coefficient function,
where $\tilde{V}_{it}$ is the idiosyncratic error. It can be more similar to FE-QR model or my model depending on the relative variation / importance of $U_i$ and $\tilde{V}_{it}$. More specifically, model (ref) is close to the FE-QR model when $\tilde{V}_{it}$ dominates $U_i+\tilde{V}_{it}$. In an extreme case that $U_i=0$, $U_{it}=0+\tilde{V}_{it}$ is iid over $t$. Model (ref) is close to my model when $U_i$ dominates $U_i+\tilde{V}_{it}$, so $U_{it}$ is stable over time for each individual $i$.
More specifically, in model (ref), $U_i$ is the individual time-invariant component of unobserved heterogeneity, which can arbitrarily correlate with regressors $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{X}}\mkern-3mu}\mkern3mu_i\equiv(\boldsymbol{\mathbf{X}}_{i1},\ldots, \boldsymbol{\mathbf{X}}_{iT})$; and $\tilde{V}_{it}$ is the idiosyncratic shock, which is assumed conditional mean zero given $\boldsymbol{\mathbf{X}}_{it}$ and $U_i$, i.e., $\operatorname{E}[\tilde{V}_{it} \mid \boldsymbol{\mathbf{X}}_{it}, U_i]=0$. Consider $U_{it}\coloneqq U_i+\tilde{V}_{it}$ follows a standard uniform distribution on $(0,1)$, and the map $u \mapsto \boldsymbol{\mathbf{x}}'\boldsymbol{\mathbf{\beta}}(u)$ is strictly increasing. Then $U_{it}$ is considered as the rank variable, which represents the rank (quantile) of counterfactual outcome when everyone in the population takes the same $\boldsymbol{\mathbf{x}}$ value.
As noted, an important feature of my model is that the rank variable is stable over time, which is often more economically plausible than the time-iid rank assumption made (often implicitly) in the panel quantile regression literature. The time-iid rank $U_{it}$ means the $U_{it}\sim \textrm{Unif}(0,1)$ is both iid over $i$ and iid over $t$. The time-stable rank $U_{it}$ means the $U_{it} \sim \textrm{Unif}(0,1)$ is iid only over $i$, but has dependence over $t$ for each $i$. The time-stable rank variable idea has been implemented in some studies, including AtheyImbens2006's (AtheyImbens2006) quantile difference-in-differences setting and Powell2022's (Powell2022) extension of instrumental variables quantile regression to panel data with non-additive FE. However, to the best of my knowledge, most papers in the panel quantile studies involving a random coefficient form $\boldsymbol{\mathbf{X}}_{it}'\boldsymbol{\mathbf{\beta}}(U_{it})$ assume a time-iid rank variable, such as Assumption 3.1(a) of Canay2011 and Assumption 2.1(c) of ArellanoBonhomme2016.
The time-stable rank $U_{it}$ is often more economically plausible than the time-iid rank $U_{it}$ in the panel quantile literature. For example, in the return to education example, individual's innate ability is considered as the unobserved rank variable that determines the rank of outcome earnings. If the individual rank is iid over time, it means that a person whose innate ability is at the 90th percentile in the population (high ability person) in this time period could be arbitrarily at the 10th percentile (the same person becomes very low ability) in the next period, which is not empirically realistic.
Third, similar to HausmanEtAl2021, the additive error $V_{it}$ in (ref) can be considered as the measurement error of the outcome variable\footnote{I thank an anonymous reviewer for pointing out this interpretation.}. That is, we observe $Y_{it}=Y_{it}^*+V_{it}$, where the latent outcome $Y_{it}^{*}$ is generated by $Y_{it}^{*}=\boldsymbol{\mathbf{X}}_{it}'\boldsymbol{\mathbf{\beta}}(U_i)$ with the same assumptions about rank variable, monotonicity, and continuity as in the general model (ref). In this case, the $\boldsymbol{\mathbf{\beta}}(\cdot)$ in (ref) can be interpreted in terms of quantile treatment effects and the quantile structural functions for the latent outcome, which will be discussed in detail in (ref).
Next, I introduce a set of formal assumptions for model (ref).
(ref) and the uniform distribution in (ref) are less restrictive than they seem. First, consider a generic non-uniform $W_i$ that generates coefficient vector $\boldsymbol{\mathbf{\gamma}}(W_i)$. If its CDF $F_W(\cdot)$ is continuous, and writing $Q_W(\cdot)$ for its quantile function, then by the probability integral transform $\tilde{U}_i=F_W(W_i)\sim\textrm{Unif}(0,1)$, and we can write $\tilde{\boldsymbol{\mathbf{\beta}}}(u)=\boldsymbol{\mathbf{\gamma}}(Q_W(u))$ so that $Y_{it}=\boldsymbol{\mathbf{X}}_{it}'\boldsymbol{\mathbf{\gamma}}(W_i) = \boldsymbol{\mathbf{X}}_{it}'\tilde{\boldsymbol{\mathbf{\beta}}}(\tilde{U}_i)$. Second, let $Y_i^*\equiv\boldsymbol{\mathbf{x}}^{*\prime}\tilde{\boldsymbol{\mathbf{\beta}}}(\tilde{U}_i)$ with continuous CDF $F_{Y^*}(\cdot)$. Again by the probability integral transform, $U_i\equiv F_{Y^*}(Y_i^*) \sim \textrm{Unif}(0,1)$. By construction, $Y_i^*$ is increasing in $U_i$, and this $U_i$ has the same $\textrm{Unif}(0,1)$ distribution as the $\tilde{U}_i$ above. Define $m(\cdot)$ such that $U_i=m(\tilde{U}_i)$. Then, as long as $m(\cdot)$ is invertible, we can define $\boldsymbol{\mathbf{\beta}}(u) \equiv \tilde{\boldsymbol{\mathbf{\beta}}}(m^{-1}(u))$ and rewrite $Y_{it}=\boldsymbol{\mathbf{X}}_{it}'\tilde{\boldsymbol{\mathbf{\beta}}}(\tilde{U}_i) = \boldsymbol{\mathbf{X}}_{it}'\boldsymbol{\mathbf{\beta}}(U_i)$.
Most importantly, the relationship between $U$ and the possibly endogenous regressors $\boldsymbol{\mathbf{X}}$ is unrestricted.
(ref) is much weaker than the full monotonicity assumption at all possible $\boldsymbol{\mathbf{x}}$ values in the instrumental variables quantile regression model of ChernozhukovHansen2005 (hereafter \citetalias{ChernozhukovHansen2005}). My model only considers monotonicity at a certain $\boldsymbol{\mathbf{x}}^{*}$ value, which suffices for identification, and my model does not impose rank invariance across other $\boldsymbol{\mathbf{x}}$ values, but rather focuses on the rank of potential outcomes with everyone counterfactually taking the certain $\boldsymbol{\mathbf{x}}^*$ value. That said, as described later, the interpretation of my $\boldsymbol{\mathbf{\beta}}(\cdot)$ is much stronger when monotonicity holds across many or all $\boldsymbol{\mathbf{x}}$ values.
The $\boldsymbol{\mathbf{x}}^{*}$ is picked for our research interest. In practice, for example, if $\boldsymbol{\mathbf{X}}=(D, \boldsymbol{\mathbf{Z}})$ contains a binary treatment variable $D$ and some covariates $\boldsymbol{\mathbf{Z}}$, we can let $\boldsymbol{\mathbf{x}}^{*}=(1,\boldsymbol{\mathbf{z}}^*)$ (or $(0,\boldsymbol{\mathbf{z}}^{*})$) if we are interested in the treated (or untreated) counterfactual outcome with everyone taking the covariates value $\boldsymbol{\mathbf{z}}^{*}$. \footnote{ I consider the counterfactual distribution in a similar spirit as ChernozhukovHansen2005 and AbadieEtAl2002. The difference is that their counterfactual distribution is conditional on covariates $\boldsymbol{\mathbf{Z}}=\boldsymbol{\mathbf{z}}$, and considering everyone takes the same treatment status $D=d$ value, whereas I consider the counterfactual distribution that both everyone takes the same treatment status $d^{*}$ and takes the same covariates $\boldsymbol{\mathbf{z}}^{*}$ value. The counterfactual distribution is not conditional on covariates $\boldsymbol{\mathbf{z}}$, but unconditionally and hypothetically that everyone takes $\boldsymbol{\mathbf{z}}^{*}$ value. } If the variable of interest is continuous, then we can let $\boldsymbol{\mathbf{x}}^*$ be the sample average of $\boldsymbol{\mathbf{X}}$, when we are interested in the outcome distribution in a parallel world in which everyone takes the mean value of $\boldsymbol{\mathbf{X}}$. Simulation results indicate the choice of $\boldsymbol{\mathbf{x}}^{*}$ does not matter for estimation as long as the population outcome orderings (over individuals) are identical across various $\boldsymbol{\mathbf{x}}^*$.
(ref) guarantees a unique solution for identification.
(ref) says the individual coefficient $\boldsymbol{\mathbf{\beta}}_i$ is identified as time-dimensional linear projection coefficient, i.e., $ \boldsymbol{\mathbf{\beta}}_i=[\operatorname{E}(\boldsymbol{\mathbf{X}}_{t}\boldsymbol{\mathbf{X}}_{t}')]^{-1}\operatorname{E} \mathopen{}\mathclose\bgroup\originalleft( \boldsymbol{\mathbf{X}}_{t} Y_{t} \aftergroup\egroup\originalright)$. There are various lower-level sufficient conditions for part (iii) like covariance stationarity and bounded second moments. It can also be denoted as $ \boldsymbol{\mathbf{\beta}}_i=[\operatorname{E}(\boldsymbol{\mathbf{X}}_{it}\boldsymbol{\mathbf{X}}_{it}'\mid i)]^{-1}\operatorname{E} \mathopen{}\mathclose\bgroup\originalleft( \boldsymbol{\mathbf{X}}_{it} Y_{it} \mid i \aftergroup\egroup\originalright)$, where the conditional on individual $i$ means treating individual $i$ as fixed.
Note that conditional on individual $i$ implies conditional on $U_i$, but the other way is not true. With $\operatorname{E}[\boldsymbol{\mathbf{X}}_{it}V_{it}\mid U_i]=\boldsymbol{\mathbf{0}}$, we can also derive from the model (ref) to get $\boldsymbol{\mathbf{\beta}}_i=[\operatorname{E}(\boldsymbol{\mathbf{X}}_{it}\boldsymbol{\mathbf{X}}_{it}'\mid U_i)]^{-1}\operatorname{E}(\boldsymbol{\mathbf{X}}_{it}Y_{it}\mid U_i)$. However, since $U_i$ is unobserved, we cannot connect this to the actual estimation method. The estimation in (ref) is conditional on $i$, i.e., treating each individual as fixed and considering its time-dimensional OLS regression.
Note that any additive time-invariant component in $V_{it}$ is absorbed into the intercept term $\boldsymbol{\mathbf{\beta}}_0(U_i)$, so $V_{it}$ does not contain any additive time-invariant component. For example, if $V_{it}=V_{1i}+V_{2t}$, then the model (ref) becomes $ Y_{it}=\underbrace{\beta_0(U_i)+V_{1i}}_{\eqqcolon\gamma_0(U_i) }+\boldsymbol{\mathbf{X}}_{-1, it}'\boldsymbol{\mathbf{\beta}}_{-1}(U_i)+V_{2t} $ with $V_{2t}$ assumed to satisfy (ref).
(ref) guarantees each individual $\boldsymbol{\mathbf{\beta}}_i$ is identified. The identification of individual-specific slopes $\boldsymbol{\mathbf{\beta}}_1, \ldots, \boldsymbol{\mathbf{\beta}}_n$ (from (ref) alone) does not necessarily result in the identification of coefficient functions $\boldsymbol{\mathbf{\beta}}(\cdot)$ without further assumptions.
We are interested in the identification and estimation of the structural coefficient function $\boldsymbol{\mathbf{\beta}}(\cdot)$ in (ref). The model allows for multiple covariates so that in general the coefficient functions $\boldsymbol{\mathbf{\beta}}(\cdot)$ are not necessarily monotone functions. More specifically, we can identify and estimate the individual's slope $\boldsymbol{\mathbf{\beta}}_i \equiv \boldsymbol{\mathbf{\beta}}(U_i)$, but we do not know whether the unobserved rank $U_i$ associated with the observed $\boldsymbol{\mathbf{\beta}}_i$ is high or low. Nor do we know the coefficient function $\boldsymbol{\mathbf{\beta}}(\cdot)$ directly. I use a sorting method in the estimation to recover how the unobserved rank is associated with the observed individual slope, and therefore, reveal the coefficient function $\boldsymbol{\mathbf{\beta}}(\cdot)$. Only in the degenerate case when $X$ is a scalar, the coefficient function $\beta(\cdot)$ has to be monotone by (ref).
Alternatively, the model can be explained by the following thought experiment. Consider there are many counterfactual parallel worlds, forcing everyone to take the $\boldsymbol{\mathbf{x}}$ values in the support of $\boldsymbol{\mathbf{X}}$. \citetalias{ChernozhukovHansen2005}'s rank invariance is to assume the outcome rank is all the same across all parallel worlds. My model (when simplified to have no idiosyncratic error) assumes the outcome rank in one parallel world ($\boldsymbol{\mathbf{x}}^*$) is invariant over time. It does not restrict the outcome rank in other parallel worlds.
The idiosyncratic error $V_{it}$ is similar in spirit to the notion of rank similarity of \citetalias{ChernozhukovHansen2005}. Essentially, they allow some “slippages” in an individual's potential outcome rank as long as they are exogenous. Here, the actual outcome ranking can be time-varying, even at the special $\boldsymbol{\mathbf{x}}^*$ where monotonicity holds, as long as such variations in $V_{it}$ are exogenous.
(ref) establish the identification of the coefficients at any rank $\tau$.
Note that the (ref) results are not conditional probabilities, but unconditional probability considering the potential outcome in a counterfactual parallel world where everyone in the population takes the $\boldsymbol{\mathbf{x}}^{*}$ value.
The parameter $\boldsymbol{\mathbf{\beta}}(\tau)$ has two immediate interpretations. First, from the model functional form (ref), $\boldsymbol{\mathbf{\beta}}(\tau)$ can be interpreted as the ceteris paribus effect of $\boldsymbol{\mathbf{X}}$ on $Y$ for the individual whose rank is $U_i=\tau$. More specifically, for the individual whose rank variable $U_i=\tau$, his structural function is $Y_t=\boldsymbol{\mathbf{X}}_t'\boldsymbol{\mathbf{\beta}}(\tau)+V_t$. If holding $V_t=v$ fixed, and increasing his $\boldsymbol{\mathbf{X}}_t$ from $\boldsymbol{\mathbf{X}}_t=\boldsymbol{\mathbf{x}}_1$ to $\boldsymbol{\mathbf{X}}_t=\boldsymbol{\mathbf{x}}_2$, this individual's outcome $Y_t$ changes by $(\boldsymbol{\mathbf{x}}_2-\boldsymbol{\mathbf{x}}_1)'\boldsymbol{\mathbf{\beta}}(\tau)$. Additionally, if $V_t$ contains some function of $\boldsymbol{\mathbf{X}}_t$ so that it is impossible to hold fixed $V_t$ while increasing $\boldsymbol{\mathbf{X}}_t$, the assumption (ref)(ii) guarantees we can interpret $\boldsymbol{\mathbf{\beta}}(\tau)$ as the marginal effect of $\boldsymbol{\mathbf{X}}$ on the mean of $Y$ for the individual with $U_i=\tau$.
Second, from the identification result in (ref), $\boldsymbol{\mathbf{\beta}}(\tau)$ is interpreted as the effect of $\boldsymbol{\mathbf{X}}$ on the $\tau$-quantile of the counterfactual outcome $Y_{x^{*}}$, provided the monotonicity condition in a neighborhood of $\boldsymbol{\mathbf{x}}^{*}$. That is, $\boldsymbol{\mathbf{\beta}}(\tau)$ measures the effect of $\boldsymbol{\mathbf{X}}$ on the $\tau$-quantile structural function (QSF) $q(\boldsymbol{\mathbf{x}},\tau)$ at $\boldsymbol{\mathbf{x}}=\boldsymbol{\mathbf{x}}^{*}$. The QSF is defined by ImbensNewey2009 as the $\tau$-quantile of counterfactual outcomes when fixing the observed covariate vector at some value $\boldsymbol{\mathbf{x}}$ while letting the unobserved variables vary over their unconditional population distribution. This is effectively the same as the “structural quantile function” in (2.4) of ChernozhukovHansen2008 or the “quantile treatment response function” in (2.2) of ChernozhukovHansen2005, who phrase it more explicitly in terms of potential outcomes. All three papers note that endogeneity prevents estimation of the QSF by conventional quantile regression because then the QSF does not equal the conditional quantile function. In addition, if we counterfactually set everyone's $V_{it}=0$ and $\boldsymbol{\mathbf{X}}_{it}=\boldsymbol{\mathbf{x}}_1$, and then change everyone to $\boldsymbol{\mathbf{X}}_{it}=\boldsymbol{\mathbf{x}}_2$, the population (unconditional) $\tau$-quantile of the counterfactual outcome changes by $(\boldsymbol{\mathbf{x}}_2-\boldsymbol{\mathbf{x}}_1)'\boldsymbol{\mathbf{\beta}}(\tau)$, which measures the $\tau$-quantile treatment effect (QTE), i.e., $\tau\textrm{-QTE} \equiv q(\boldsymbol{\mathbf{x}}_2,\tau)-q(\boldsymbol{\mathbf{x}}_1,\tau) \equiv Q_\tau(Y_{\boldsymbol{\mathbf{x}}_2})-Q_\tau(Y_{\boldsymbol{\mathbf{x}}_1})=(\boldsymbol{\mathbf{x}}_2-\boldsymbol{\mathbf{x}}_1)'\boldsymbol{\mathbf{\beta}}(\tau)$, provided the monotonicity of $\boldsymbol{\mathbf{x}}'\boldsymbol{\mathbf{\beta}}(u)$ in $u$ holds at both $\boldsymbol{\mathbf{x}}_1$ and $\boldsymbol{\mathbf{x}}_2$. Without full rank invariance at every possible $\boldsymbol{\mathbf{x}}$ values, this interpretation is limited to any $(\boldsymbol{\mathbf{x}}_1, \boldsymbol{\mathbf{x}}_2)$ that satisfies the monotonicity condition.
These two interpretations reconcile because the $\tau$-quantile of the counterfactual outcome belongs to the individuals whose $U_i=\tau$ for both $\boldsymbol{\mathbf{X}}_{it}=\boldsymbol{\mathbf{x}}_1$ and $\boldsymbol{\mathbf{X}}_{it}=\boldsymbol{\mathbf{x}}_2$. So the effect on the $\tau$-quantile of counterfactual outcome (i.e., the second interpretation of $\boldsymbol{\mathbf{\beta}}(\tau)$) is identical to the effect on outcome for the $U_i=\tau$ individual (i.e., the first interpretation of $\boldsymbol{\mathbf{\beta}}(\tau)$).
Moreover, inspired by HausmanEtAl2021, we can interpret $V_{it}$ as the measurement error and $Y_{it}^{*}=\boldsymbol{\mathbf{X}}_{it}'\boldsymbol{\mathbf{\beta}}(U_i)$ as the true latent outcome; then the $\boldsymbol{\mathbf{\beta}}(\cdot)$ in my model (ref) has the same interpretation as in a simplified model $Y_{it}=\boldsymbol{\mathbf{X}}_{it}'\boldsymbol{\mathbf{\beta}}(U_i)$. More specifically, $\boldsymbol{\mathbf{\beta}}(\cdot)$ can be interpreted as the marginal effect of $\boldsymbol{\mathbf{X}}$ on the true latent $Y^{*}$ (and mean of $Y$) for the individual whose rank variable value is $U=\tau$.
Note that my model (ref) is a structural model. The parameters of interest are the structural effects. In other words, if (ref) is strengthened to hold in a small neighborhood of $\boldsymbol{\mathbf{x}}^{*}$ and $V=0$, then $\boldsymbol{\mathbf{\beta}}(\tau)$ is interpreted as the effect of $\boldsymbol{\mathbf{X}}$ on the $\tau$-quantile structural function $q_Y(\boldsymbol{\mathbf{x}}^{*}, \tau)$. The QSF is different from the CQF in general, as noted by ImbensNewey2009. \footnote{ ImbensNewey2009 define the quantile structural function and write, “Note that because of the endogeneity of $X$, this [QSF] is in general not equal to the conditional quantile of $g(X,\epsilon)$ conditional on $X=x$, $q_{Y\mid X}(\tau \mid x)$.” } Let $V=0$. If (ref) is strengthened to hold at all $\boldsymbol{\mathbf{x}}$ in the support of $\boldsymbol{\mathbf{X}}$ as in \citetalias{ChernozhukovHansen2005}, then $\boldsymbol{\mathbf{x}}'\boldsymbol{\mathbf{\beta}}(\tau)$ is the $\tau$-QSF $q_Y(\boldsymbol{\mathbf{x}},\tau)$, the same as identified in \citetalias{ChernozhukovHansen2005} using instrumental variables. The differences of the QSF over the values of endogenous regressor yields the quantile treatment effects, i.e., $\tau\textrm{-QTE}=q(\boldsymbol{\mathbf{x}}_2,\tau)-q(\boldsymbol{\mathbf{x}}_1,\tau)$, provided the QSF is identified at both $\boldsymbol{\mathbf{x}}_1$ and $\boldsymbol{\mathbf{x}}_2$.
My structural model coincides with the CQF only in a very special case. If $\boldsymbol{\mathbf{X}}$ is exogenous (i.e., $\boldsymbol{\mathbf{X}}$ and $U$ are independent), $V_{it}=0$, and monotonicity holds at all $\boldsymbol{\mathbf{x}}$, then the $\tau$-quantile of the outcome conditional on $\boldsymbol{\mathbf{X}}_{it}=\boldsymbol{\mathbf{x}}$ is $Q_\tau(Y_{it} \mid \boldsymbol{\mathbf{X}}_{it} = \boldsymbol{\mathbf{x}}) = \underbrace{Q_\tau(\boldsymbol{\mathbf{X}}_{it}'\boldsymbol{\mathbf{\beta}}(U)+\overbrace{V_{it}}^{=0} \mid \boldsymbol{\mathbf{X}}_{it} = \boldsymbol{\mathbf{x}})}_{\boldsymbol{\mathbf{X}} \protect\mathpalette{\protect\independenT}{\perp} U } = \underbrace{Q_{\tau}(\boldsymbol{\mathbf{x}}'\boldsymbol{\mathbf{\beta}}(U))}_{\textrm{by \cref{A2:rank} and monotonicity} } = \boldsymbol{\mathbf{x}}'\boldsymbol{\mathbf{\beta}}(\tau)$ from the structural model.
First, I define the estimated coefficient function at any rank value. Next, I show the uniform consistency and asymptotic normality of the estimated coefficient function as both $T$ and $n$ go to infinity under a certain rate condition.
Define the permutation $\sigma(\cdot)$ to associate the individual “$i$” with its order “$k$”: \footnote{ $U_{n:k}$ represents the $k$th order statistic in a sample of size $n$, i.e., the $k$th-smallest value among $\{ U_1, \ldots, U_n\}$. }
So we can write the coefficient $ \boldsymbol{\mathbf{\beta}} \mathopen{}\mathclose\bgroup\originalleft( U_{n:k} \aftergroup\egroup\originalright) = \boldsymbol{\mathbf{\beta}} \mathopen{}\mathclose\bgroup\originalleft( U_{\sigma(k) } \aftergroup\egroup\originalright)$. Under (ref), the permutation $\sigma(\cdot)$ can be learned from the $Y^*_i$ ordering, i.e.,
because $Y^*_{\sigma(k)}=\boldsymbol{\mathbf{x}}^{*\prime}\boldsymbol{\mathbf{\beta}}(U_{\sigma(k)})=\boldsymbol{\mathbf{x}}^{*\prime}\boldsymbol{\mathbf{\beta}}(U_{n:k})=Y^*_{n:k}$.
Then $\boldsymbol{\mathbf{\beta}} \mathopen{}\mathclose\bgroup\originalleft( U_{n:\lceil \tau n\rceil} \aftergroup\egroup\originalright)$ is the coefficient associated with the $Y^*_{\sigma(\lceil \tau n\rceil)}=Y_{n: \lceil \tau n\rceil}^{*}$, and $Y_{n: \lceil \tau n\rceil}^{*}$ is the $ \lceil \tau n\rceil$-th smallest value among $\{Y_i^{*}\}_{i=1}^n$, where $U_{n:\lceil \tau n\rceil}$ is the $\lceil \tau n\rceil$th order statistic, and $\lceil\tau n \rceil $ denotes the ceiling function of $\tau n$, i.e., the least integer greater than or equal to $\tau n$. Because $0<\tau<1$, $ 1 \le \lceil \tau n \rceil \le n $.
Define the estimated permutation $\hat{\sigma}(\cdot)$ to associate the individual $i$ with its fitted outcome value order $k$:
where the fitted outcome is defined just below in (ref). The estimated permutation $\hat{\sigma} \mathopen{}\mathclose\bgroup\originalleft( k \aftergroup\egroup\originalright)$ depends on the data and the $\boldsymbol{\mathbf{x}}^{*}$ value. Later (ref) shows the estimated order equals the true order with probability approaching one under a certain rate relation between $n$ and $T$, so $\operatorname{P}(\hat\sigma(\cdot)=\sigma(\cdot))\to1$.
Specifically, the estimation of $\boldsymbol{\mathbf{\beta}}(\tau)$ consists of four steps.
I define the estimator in (ref) because $\operatorname{P}(\hat\sigma(\cdot)=\sigma(\cdot))\to1$ from (ref), that $\boldsymbol{\mathbf{\beta}} \mathopen{}\mathclose\bgroup\originalleft( U_{n:\lceil \tau n\rceil} \aftergroup\egroup\originalright)$ is the coefficient associated with the $Y^*_{\sigma(\lceil \tau n\rceil)}=Y_{n: \lceil \tau n\rceil}^{*}$, and that $ U_{n:\lceil \tau n\rceil}\xrightarrow{p} \tau$, as $n\to \infty$.
I first introduce additional assumptions and then show the uniform consistency of the estimated coefficient function.
Let $\hat{\boldsymbol{\mathbf{\beta}}}_T(U_i) $ denote the time series OLS (TS-OLS) estimator for individual $i$'s coefficient $\boldsymbol{\mathbf{\beta}}(U_i) $,
There are many large-sample results (HamiltonText) for the TS-OLS estimator under various assumptions. For example, given (ref) and finite second-order moments of $\boldsymbol{\mathbf{X}}$ and $V$, the TS-OLS estimator is pointwise consistent for each individual. That is, $\hat{\boldsymbol{\mathbf{\beta}}}_T(U_i) = \boldsymbol{\mathbf{\beta}}(U_i) + \mathopen{}\mathclose\bgroup\originalleft( \frac{1}{T} \sum_{t=1}^T \boldsymbol{\mathbf{X}}_{it}\boldsymbol{\mathbf{X}}_{it}' \aftergroup\egroup\originalright)^{-1} \mathopen{}\mathclose\bgroup\originalleft( \frac{1}{T} \sum_{t=1}^T \boldsymbol{\mathbf{X}}_{it}V_{it} \aftergroup\egroup\originalright)\xrightarrow{p} \boldsymbol{\mathbf{\beta}}(U_i) \textrm{ as } T \to \infty , \ \forall i=1, \ldots, n .$
I further assume that the TS-OLS estimator of each individual-specific coefficient is uniformly consistent as $T\to\infty$ and $n\to \infty$.
(ref) implies the pointwise consistency for each individual coefficient. The norm in (ref) is taken as the supremum norm for simplicity, but it can be interpreted as any $L_p$ norm due to the equivalence of norms in finite-dimensional space. (ref) implicitly restricts how fast $T$ must grow with respect to $n$. A sufficient $n/T$ rate relation is given in the uniform consistency result of the estimated coefficient function $\hat{\boldsymbol{\mathbf{\beta}}}(\cdot)$ in (ref).
(ref) is a stronger condition than (ref). With the uniform continuous coefficient function assumed in (ref) and the uniform consistency of the uniform sample quantile function from ShorackWellner1986, I establish the uniform consistency of the coefficient function estimator $\hat{\boldsymbol{\mathbf{\beta}}}(\cdot)$ on $(0,1)$.
The monotonicity condition (ref) assumes the function $g(u)= \boldsymbol{\mathbf{x}}^{*\prime} \boldsymbol{\mathbf{\beta}}(u)$ is strictly increasing on $(0,1)$. (ref) strengthens this slightly, so $g(u)$ has a strictly positive lower bound of its derivative. It prevents the situation that the function $g(u)$ is almost flat on some intervals; for example, $g(u)=\sin(u\pi/2)$ satisfies (ref) because $g'(u)>0$ for all $0<u<1$ but violates (ref) because $\lim_{u\uparrow1}g'(u)=0$. This restriction is not a strong assumption in the sense that it still allows an individual element function of the vector-valued coefficient $\boldsymbol{\mathbf{\beta}}(\cdot)$ to be non-monotone or almost flat.
(ref) helps establish the stochastic order of $\max_{1 \le i \le n} \lVert\hat{\boldsymbol{\mathbf{\beta}}}_{T}(U_i)-\boldsymbol{\mathbf{\beta}}(U_i)\rVert $ in terms of $n$, which then translates to $T$ using a specified $n/T$ rate relation. By Theorem A.7 of LiRacine2007, (ref) implies the individual coefficient TS-OLS convergence rate, i.e., $ \hat{\boldsymbol{\mathbf{\beta}}}_T(U_i)- \boldsymbol{\mathbf{\beta}}(U_i) = O_p \mathopen{}\mathclose\bgroup\originalleft( T^{-\kappa} \aftergroup\egroup\originalright) \textrm{ as } T \to \infty , \ \forall i=1, \ldots, n. $
(ref) shows a sufficient $n/T$ rate relation that guarantees the true order of the $U_i$ can be discovered from the order of $\hat{Y}^*_i$. That is, (ref) implies the estimated permutation is asymptotically the same as the true permutation.
The coefficient estimator defined based on the estimated permutation is uniformly consistent over $(0,1)$.
The $n/T$ rate relation in (ref) and (ref) is a sufficient rate to guarantee uniform consistency. The benchmark rates for FE-QR type of model is $n=o(T^{1/2})$. The simulation section shows the performance of the proposed estimator in an example DGP with various more relaxed $n/T$ rate relations and with small $T$, like $T=3, \ldots, 10$, in some case.
In this subsection, I establish the uniform asymptotic normality of the coefficient function estimator by applying the functional delta-method to the well-known Donsker's theorem, assuming differentiability of the coefficient function.
Donsker's theorem vanderVaartWellner1996 states the asymptotic Gaussianity of the empirical distribution function. Letting $F_n$ be the sequence of the empirical distribution function indexed by sample size $n$, assuming iid sampling from distribution function $F$, then
where $\mathbb{G} $ is the standard Brownian bridge.\footnote{ The standard Brownian bridge process $\mathbb{G} $ on the unit interval $[0,1]$ is mean zero with covariance function $\operatorname{Cov}[ \mathbb{G}(u), \mathbb{G}(v)]=u \wedge v - u v$, where $u \wedge v$ denotes $\min \{ u,v \}$. } The limiting process is mean zero and has covariance function $\operatorname{Cov} \mathopen{}\mathclose\bgroup\originalleft[ \mathbb{G}_F(u), \mathbb{G}_F(v) \aftergroup\egroup\originalright]$ $=$ $F(u) \wedge F(v)- F(u)F(v)$. The symbol $\ell^{\infty}(T)$ denotes the collection of all bounded functions $f\colon T \mapsto {\mathbb R}$ with the norm $\mathopen{}\mathclose\bgroup\originalleft\lVertf\aftergroup\egroup\originalright\rVert_{\infty}=\sup_{t\in T}\lvertf(t)\rvert$.
(ref) is a regularity condition that assumes the differentiability of the coefficient function.
The uniform asymptotic normality of the coefficient function estimator also holds under the same sufficient $n/T$ rate relation for its uniform consistency. The TS-OLS individual coefficient estimates do not exactly equal to the true coefficient parameters, but the order of the rank variables can still be discovered asymptotically under the rate condition as shown in (ref). Under the same rate condition, the TS-OLS estimation error goes away even after scaling by $\sqrt{n}$. Thus, together with applying the fundamental property in (ref) to a standard uniform distribution function, we get the uniform asymptotic normality of the coefficient function estimator
(ref) applies (ref) at a fixed $\tau$ to get the pointwise asymptotic normality of the coefficient estimator at any specific rank.
In this section, first, I present the estimator's performance at various $n/T$ relations. Second, I consider whether the choice of sorting point matters for the estimates. Third, I compare my estimator with Canay2011's (Canay2011) FE-QR estimator. Results can be replicated with my code in R R.core. All codes are available online.\footnote{\url{https://xinliu16.github.io/}}
This section illustrates the performance of the estimator at different $n/T$ rate relations. (ref) provides a theoretical sufficient condition of $n/T$ rate relation to guarantee consistency. The necessary convergence rate can be much smaller than (ref) provides. This section uses simulation to check whether the proposed estimator works well in practice with relatively small $T$.
The data generating process (DGP) is
where the regressor $X_{it}=X_{ind, it}+4+\rho U_i$ with $X_{ind, it} \stackrel{\mathit{iid}}{\sim} N(0,1)$ independent of $(U_i, V_{it})$, and $\rho$ is some non-negative constant that represents the endogeneity level of the DGP; the $+4$ is to shift $X_{ind, it}$ to the right by 4 times of its standard deviation to make more than 99.99% percentage of $X_{it}$ takes positive value in the sample, where the monotonicity condition (of $Y$ in $U$) holds; the intercept function is $\beta_0(u)=u$; the slope function is $\beta_1(u)=u^2$; let $U_i \stackrel{\mathit{iid}}{\sim} \textrm{Unif} [0,1]$, and the idiosyncratic error $V_{it} \stackrel{\mathit{iid}}{\sim} N(\mu, \sigma_V^2)= N(0,1)$.
(ref) reports the mean squared error (MSE) of the slope estimators at $0.25$-, $0.5$-, and $0.75$-quantiles with different $n/T$ rate relations $T=n$, $T=n^{3/4}$, $T=n^{1/2}$, and $T=n^{1/4}$. I let $\rho=1$ to denote the endogeneity and pick the sorting point at $(1, x_1^{*})=(1, 4.5)$. Vertically, the MSE decreases significantly as $n$ increases from $n=100$ to $n=10{,}000$ in all the $n/T$ rate relation columns.\footnote{The blank place is due to out of software memory in R in the large $n$ and large $T$ cases, like $n=T=10{,}000$.} The higher the rate, the quicker MSE drops towards zero. For example, for the $\tau=0.5$ slope estimator, the MSE drops from 0.013 to 0.001, as $n$ increases from 100 to 1000 at rate $T=n$; MSE drops from 0.038 to 0.004, as $n$ increases from $100$ to $2000$ at rate $T=n^{3/4}$; and MSE drops from 0.148 to 0.011, as $n$ increases from 100 to $10{,}000$ at rate $T=n^{1/2}$. Horizontally, at each fixed $n$, the MSE also drops as $T$ increases.
These simulation results are compatible with the consistency of the estimator. The MSE is decreasing towards zero as $n$ increases even with the slow $T=n^{1/4}$ rate, where the $T$ is relatively small (no more than $10$)\footnote{The $T$ ranges from 3 to 10 in the $T=n^{1/4}$ column. Even with the largest $n=10{,}000$, $T=n^{1/4}=10{,}000^{1/4}=10$. } even though theoretically this is not a fixed-$T$ model. This estimator can be applied to many common microeconometric setting.
In this section, I investigate if (and how) the choice of sorting point $x^{*}$ affects the estimates.
Consider the same DGP as in (ref). I consider five various sorting points $x^{*}$ from 2.5 to 6.5 at a increment of 1. Note that the [2.5, 5.5] includes the 6.7th- to 93.3th- percentile of a $N(4,1)$ random variable, and that $\rho U_i \in [0,1]$. So [2.5, 6.5] includes more than 86.6% of possible $X$ values.
Note that if the population order of $Y^{*}=\beta_0(U)+\beta_1(U)x_1^{*}$ at a sorting point $x_1^{*}$ is different from the population order of $Y^{*}=\beta_0(U)+\beta_1(U)x_2^{*}$ at another sorting point $x_2^{*}$, then the sample estimates using $x_1^{*}$ and $x_2^{*}$ are expected to be different. Thus, I consider specifically the situation that when $Y_{x_1^{*}}$ and $Y_{x_2^{*}}$ have the same population order, and see if the sorting points $x_1^{*}$ and $x_2^{*}$ matter for the sample estimates $\hat{\beta}_1(\tau)$.
(ref) shows the biases are almost the same (up to 0.001 difference) using a variety of sorting points $x^{*}$ at all quantile levels ($\tau=0.25, 0.5, 0.75$). It illustrates that the sorting point does not matter for sample estimates, as long as the population orders are identical across various sorting points. (ref) in (ref) includes extra simulation results for the DGP without shifting $X$ and pick various sorting points $x^{*}=0, 0.5, 1, 1.5, 2, 2.5$. That is, half of the unshifted $X$ are negative values and the population orders are different at negative $x^{*}$ from that at positive $x^{*}$ values. The results and conclusion are the same as above that the sample estimates $\hat{\beta}_1(\tau)$ do not rely on the choice of sorting points $x^{*}$, wherever the population orders of $Y^{*}$ are identical across these various $x^{*}$.
(ref) includes extra simulation results with larger sample $(n,T)=(1000,1000)$ and specifically small $T$, such as $(n,T)=(1000, 10)$ case. (ref) shows with large sample size $(n,T)=(1000,1000)$, my estimator performs very well with up to 0.001 bias and less than 0.0005 MSE at all quantile levels $\tau=0.25, 0.5, 0.75$. (ref) shows even in the $T=10$ very short-$T$ case, my estimator has up to 0.005 bias and the estimates do not vary much (up to 0.009 estimates difference) across different $x^{*}$ values.
This section compares my estimator and Canay2011's (Canay2011) FE-QR estimator in various situations.
Consider the DGP that
where the regressor $X_{it}=X_{ind, it}+4+\rho U_i$ with $X_{ind, it} \stackrel{\mathit{iid}}{\sim} N(0,1)$ independent of $(U_i, \tilde{V}_{it})$, and $\rho$ is some non-negative constant that represents the endogeneity level of the DGP; specifically, I consider $\rho=0$ (no endogeneity) and $\rho=1,3,10$ as endogeneity level increases. The intercept function is $\beta_0(u)=u$; the slope function is $\beta_1(u)=u^2$; let $U_i \stackrel{\mathit{iid}}{\sim} \textrm{Unif} (0,1)$, and the idiosyncratic error $\tilde{V}_{it} \stackrel{\mathit{iid}}{\sim} N(0, \sigma_V^2)$. Since the $U_i+\tilde{V}_{it}$ may take value outside $[0,1]$, I normalize it to be on $[0,1]$ to represent the rank variable. That is, I consider the CDF of the random variable $U_i+\tilde{V}_{it}$, i.e., let $ F(U_i+\tilde{V}_{it})\sim \textrm{Unif}(0,1)$ and $\beta_0(\cdot)$ and $\beta_1(\cdot)$ are functions of $F(U_i+\tilde{V}_{it})$.
I take the sorting point at $x^{*}=4$, which is around the center of $X_{it}$, as $X_{ind, it}+4 \sim N(4,1)$ and $\rho U_i \in [0,\rho]$. As shown in (ref), the choice of sorting points $x^{*}$ does not matter for sample estimates, as long as the population orders of $Y^{*}$ are identical across these $x^{*}$ values.
I consider three cases $\sigma_V=0.01$, $0.1$, and $1$, respectively. The third case ($\sigma_V=1$) is towards the Canay FE-QR situation that the time-iid component $\tilde{V}_{it}$ dominates $U_i+\tilde{V}_{it}$. The first case ($\sigma_V=0.01$) is towards the situation that rank variable is time-persistent, i.e., $U_i$ dominates the $U_i+\tilde{V}_{it}$. The second case $\sigma_V=0.1$ is in between and is a time-stable rank variable case.
(ref) reports the bias and MSE of my estimator and Canay2011's (Canay2011) FE-QR estimator. In the first case $\sigma_V=0.01$, my estimator has much smaller bias and smaller MSE than FE-QR estimator at almost all quantile levels ($\tau =0.25, 0.5, 0.75$), especially when the DGP has endogeneity. For example, with high endogeneity $\rho=10$, my estimator has 0.000 bias and 0.001 MSE for $\tau=0.25$ quantile, in contrast to FE-QR estimator has a 0.252 bias and 0.065 MSE.
In the middle case that $\sigma_V=0.1$, my estimator performs comparably to Canay2011's (Canay2011) FE-QR estimator. In the no/low endogeneity case ($\rho=0,1$), my estimator has a smaller bias and MSE than FE-QR estimator. In the high endogeneity case ($\rho=10$), my estimator usually has a smaller bias, but bigger MSE than FE-QR estimator. For example, my estimator has 0.069 bias and 0.034 MSE for $\hat{\beta}_1(0.25)$, in contrast to FE-QR estimator has a 0.118 bias and 0.015 MSE.
The FE-QR estimator performs much better with $\sigma_V=1$ than with smaller $\sigma_V$. Only in cases at $\tau=0.5$ with high endogeneity (like $\rho=10$), my estimator has a smaller bias and smaller MSE than FE-QR estimator. In most other cases ($\tau=0.5$ with mild endogeneity and $\tau=0.25, 0.75$ with various endogeneity), the FE-QR estimator has a much smaller bias and MSE than my estimator.
(ref) shows my model complements FE-QR. The FE-QR performs poorly once $\sigma_V$ is small enough, but it outperforms my approach once $\sigma_V$ is large enough. So in practice we choose between my approach and the FE-QR approach depending on whether we think the $U_i$ or $\tilde{V}_{it}$ is a bigger source of variation.
(ref) shows the $(n,T)=(1000,10)$ results parallel to (ref). The result patterns are the same.
Moreover, my estimator is computationally simple and fast relative to Canay's FE-QR estimator, especially with large $n$ and large $T$. It takes 10 seconds for my estimator to run 500 replications with four endogeneity cases and a sample size of sample $(n,T)=(250, 250)$, but it takes 4 minutes to run the same for FE-QR estimator. It is infeasible to compute Canay's FE-QR estimator with 500 replications and sample size $(n,T)=(500,500)$, whereas it only takes 25 seconds for my estimator in the same setting.
To apply the proposed method to an empirical example, I study the causal effect of a country's oil wealth on its military defense burden using the data from CotetTsui2013. There are many studies debating the effect of oil wealth on political conflict. On one hand, resource scarcity triggers conflict, so resource abundance should mitigate regime instability. On the other hand, political instability in the Middle East and some developing countries with rich oil resources makes people arrive at the “oil-fuels-war” conclusion. The political violence in a regime can be measured in multiple ways, for example, the onset of civil war. Additionally, political conflict may not necessarily lead to regime shift or war. Military defense burden is also considered as a barrier to political conflict. My empirical example studies the effect of oil wealth on military defense spending.
First, I provide point estimates of the effect of oil wealth on military defense spending (as a ratio of GDP) for countries at various rank levels $\tau$. The rank variable represents a country's unobserved general propensity of having military spending. Second, I present bootstrap standard error for the estimated effects. \footnote{ Details of bootstrap algorithm for computing standard error are presented in (ref). }
\Citet{CotetTsui2013} use the standard fixed effect method to study the effect of oil wealth on political violence. They provide several measures of oil wealth, such as oil reserves or oil discoveries, and several measures related to political violence, such as the onset of civil war or the military defense spending as a ratio of GDP. They find that after controlling for the country fixed effect, oil wealth has no effect on civil conflict onset or military spending. I investigate if there is effect heterogeneity across the unobserved military spending rank variable, although the effect might be statistically insignificant as CotetTsui2013 find.
Consider the structural model
The outcome variable is measured as (the log of the percentage points of) the ratio of military defense spending to GDP. The variable of interest is $\log( \textit{OILWEALTH}_{it} )$, the log of dollar-valued inflation-adjusted oil wealth per capita. The parameter $\beta(U_i)$ is the causal effect of oil wealth on military spending. The unobserved country rank variable $U_i$ denotes country $i$'s general propensity to have military spending. The effect can differ across countries with different $U_i$. The covariate vector $\boldsymbol{\mathbf{X}}_{it}$ contains country characteristics including economic growth and population. $V_{it}$ is the idiosyncratic error, which is assumed exogenous so that TS-OLS is consistent. The goal is to estimate $\beta(\tau)$ for various $\tau$ rank.
The data is from CotetTsui2013. My sample includes 45 countries over years 1988--2003. The military defense burden outcome variable is measured as a log of percentage points. For example, if a country's military spending is 3.2% of its GDP, then its outcome variable value is $\log(3.2)$. Oil wealth is measured as log oil value per capita.\footnote{ \Citet{CotetTsui2013} instead use the log of oil wealth per capita divided by $100$. My coefficient estimates can be multiplied by $100$ to compare directly with those of CotetTsui2013. } For example, a country with 6 million dollars per capita oil reserve wealth is counted as $\log( \num[round-mode=none,group-digits=integer]{6000000})=6.8$ as its $\log(\textit{OILWEALTH}_{it})$ value. The effect $\beta(U_i)$ measures the elasticity given a 1% increase in oil wealth per capita, i.e., the percent effect on the percentage points of the military spending to GDP ratio.
The two other covariates included in my application are the country's economic growth and population density. Economic growth is measured as the country's annual GDP growth rate. For example, if a country's annual GDP growth is 3%, then the value of economic growth is $0.03$. The population density is measured by the log of the country's total population.\footnote{ \Citet{CotetTsui2013} use log of the country's population divided by 100 as the population density. }
(ref) presents the coefficient estimates at rank values $\tau=\{0.25, 0.5, 0.75\}$ quantile using sorting point of the sample mean $\boldsymbol{\mathbf{x}}^{*}=1/(nT)\sum_{i=1}^n \sum_{t=1}^T \boldsymbol{\mathbf{X}}_{it}$ including both the regressor of interest $\log( \textit{OILWEALTH})$ and other covariates. The interpretation of the rank variable order is associated with the sorting points. The robustness of estimates using nearby sorting points $\boldsymbol{\mathbf{x}}^*$ is presented in (ref). For comparison, (ref) also presents Canay2011's (Canay2011) FE-QR estimates and the standard FE estimates.
For all three models, the estimated effects are statistically insignificant. This is consistent with the literature CotetTsui2013,FearonLaitin2003. The existing macro-level studies of the oil-fuels-war hypothesis find that the effect is significant in a pooled OLS setting. The inclusion of country fixed effects eliminates the statistical association between oil and civil war, and the statistical significance disappears.
Despite the statistical insignificance, I interpret the estimates economically for illustration. Assuming my model specification, at $\tau=0.5$, the estimated coefficient $-0.18$ means: for a country with median propensity for military spending, a $1\%$ increase in oil wealth per capita is estimated to cause a $0.18\%$ decrease in the percentage points of military spending as a ratio of GDP. Using Canay's FE-QR model specification, the estimate at $\tau=0.5$ is instead $0.04$, which is close to the “mean” effect using the standard FE model. Of course, different model specifications make very different assumptions. The standard FE model assumes a homogeneous slope across units, and Canay's FE-QR model allows slope heterogeneity assuming time-iid rank, whereas my model assumes time-invariant rank. Except in trivial cases, my model and Canay's cannot both be correctly specified, so it is not surprising that the estimates differ economically.
In this paper, I propose a structural quantile-based nonadditive fixed effects panel model with the goal of discovering heterogeneous causal effects while accounting for endogeneity. This model generalizes the standard FE model and complements the FE-QR model. The fixed effects enter the intercept and slopes nonseparably. The causal effect is heterogeneous in the unobserved rank variable, which represents the rank of the counterfactual outcome at certain $\boldsymbol{\mathbf{x}}^{*}$ values that policy-makers are interested in. The rank variable is assumed to be time-stable, which often makes more economic sense than the time-iid rank variable in panel quantile regression literature. I develop several theoretical results, including identification, uniform consistency, and uniform asymptotic normality of the coefficient estimator. Simulation shows the compatibility with consistency even with small-$T$ and at weaker $n/T$ rate relations, that the estimates do not depend on choice of sorting point, and the comparison between my estimator and a popular FE-QR estimator.
I thank the editor, associate editor, and anonymous reviewers for their helpful comments that greatly improved this paper. I'm grateful to David Kaplan for his patient and enthusiastic guidance, encouragements and discussions. I also thank Zack Miller, Shawn Ni, Alyssa Carlson, Rachael Meager, seminar participants at University of Glasgow, University of Amsterdam, Washington State University, University of Connecticut, as well as the participants at 2020 EWMES, 2021 AMES, 2022 MEG, and 2022 SEA for helpful conversations and comments on this project. All errors are mine.
The author reports there are no competing interests to declare.
\iftoggle{SUPPLEMENTAL}{ \singlespacing }
\onehalfspacing