EconBase
← Back to paper

Quantile Random-Coefficient Regression with Interactive Fixed Effects: Heterogeneous Group-Level Policy Evaluation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

72,340 characters · 17 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 456 chars of source]
abstractWe propose a quantile random-coefficient regression with interactive fixed effects to study the effects of group-level policies that are heterogeneous across individuals. Our approach is the first to use a latent factor structure to handle the unobservable heterogeneities in the random coefficient. The asymptotic properties and an inferential method for the policy estimators are established. The model is applied to evaluate the effect of the minimum wage policy on earnings between 1967 and 1980 in the United States. Our results suggest that the minimum wage policy has significant and persistent positive effects on black workers and female workers up to the median. Our results also indicate that the policy helps reduce income disparity up to the median between two groups: black, female workers versus white, male workers. However, the policy is shown to have little effect on narrowing the income gap between low- and high-income workers within the subpopulations. {\it Keywords:} Heterogeneous policy effect, Hierarchical regression, Random coefficient model {\it JEL Codes:} C13, C31, J15, J31, J38.

Introduction

Hierarchical, or nested, data structure is natural in many research fields, including policy evaluation, educational research and meta-analysis, to name a few. In this paper, we are particularly interested in the hierarchical data observed over multiple periods, where the group level consists of a number of cross-sections (such as industries, and states) across time, and the individual level comprises random samples within each group unit. A typical example in survey is the repeated cross-sectional data.

Random-coefficient regression models prevail in modeling hierarchical data, as they not only allow the heterogeneity across groups but also link the individual- and group-level explanations. The development of random-coefficient models in the conditional mean paradigm can be dated back to the 1970s swamy1970efficient,hsiao1975some,borjas1994two,zhang2019identification. However, the above models are misspecified when unobserved heterogeneous individual effects exist, in which case quantile regression serves as an alternative approach since the seminal work by koenker1978regression. Several papers study the quantile random coefficient models, such as kim2011semiparametric, graham2018quantile, pitselis2020multi and chetverikov2016iv.

The aforementioned literature studies the hierarchical data at a single period. These models fail to capture the dynamics across time, and thus, lead to biased estimates when applying to data with multiple periods. Indeed, limited research can be found tailored for the hierarchical data with multiple periods. In this paper, we consider a latent factor structure to flexibly capture the unobserved time and group effects simultaneously. The latent factor structure is prevailingly used in the panel data literature to control the time-varying common shocks that are distinct across groups, which is known as the interactive fixed effects bai2009panel, pesaran2006estimation. Quantile extension of the latent factor structure has recently been studied by ando2020quantile,chen2021quantile. However, to the best of our knowledge, this is the first paper that allows for interactive fixed effects in the quantile random-coefficient model.

To this end, we proposed a quantile random-coefficient model, where the random coefficients are heterogeneous across both time and cross-sections and modelled by a linear regression with interactive fixed effects.\footnote{As cross-sections are the same across time, the group-level regression framework can be viewed as a panel regression.} For this model, we introduce a simple two-step estimation procedure. First, we estimate the quantile regression model, introduced by koenker1978regression, for each group-time pair. Then, we estimate the group-level model on quantile random coefficients using the estimation procedure proposed by bai2009panel. This model is capable of documenting the observed and unobserved heterogeneities at both individual and group levels, as well as the interactions between the two levels. We note that the model is not only of methodological importance but also of empirical interest. We show that the model provides us with a way to study how group-level policies affect the marginal effects differently among observationally identical individuals and uncover policy effects depending on individual observed and unobserved heterogeneities, which is important yet “somewhat neglected” in the policy evaluation koenker2017quantile.

In summary, the contributions of this paper are threefold.

First, from the methodological point of view, we contribute to the literature on quantile random-coefficient models by (i) allowing the random coefficients to be heterogeneous across both groups and time and (ii) characterizing the unobserved group and time heterogeneities within the random coefficients flexibly via interactive fixed effects. A closely related paper to the empirical aspect of our paper is oka2021heterogeneous. They characterize the unobserved time and group heterogeneity via two-way fixed effects in their quantile random coefficients, which can be viewed as a special case of our proposed model. In addition, our paper takes one step further to establish asymptotic results and identification conditions for a number of treatment parameters that quantify the heterogeneous effects of group-level policy. We extend the scope of policy evaluation approaches gobillon2016regional, Arellano2016 by establishing the identification strategy of the distributional group-level policy effect which is heterogeneous across the individuals' characteristics.

Second, from the theoretical perspective, we establish the consistency and limiting distribution of the proposed group-level coefficients estimator when the number of observations per group-time pair $N_{st}$, cross-sections $S$, and time units $T$ all go to infinity simultaneously. The estimation error from the first-step quantile regression imposes a great challenge in deriving the asymptotic expression of the group-level estimators. However, we show that the first-step estimation error is negligible, as long as the relative growth rate between individual-level and group-level sample sizes is sufficiently large.

Last but not least, from the empirical perspective, we apply the model to contribute to the debate on racial income disparity during 1960s and 1970s in the United States freeman1973changes,card1992school,derenoncourt2021minimum. Unlike derenoncourt2021minimum, our method estimates the heterogeneous policy effects, particularly on racial income inequality, under a unified framework without the need for alternating responses and selecting sub-samples. In addition, the interactive fixed effects in our model are suitable for controlling for time-varying unobserved common shocks, such as macroeconomic shocks, which could affect industries differently. Our estimation results support the core conclusion of derenoncourt2021minimum that minimum wage policy helps reduce the racial income gap and provide additional findings that the minimum wage policy has a significant negative impact on between-inequality but little effect on within-inequality.

The paper proceeds as follows. Section (ref) introduces the model and proposes a two-step estimation method. Section (ref) presents the asymptotic properties of the estimators. Section (ref) discusses the identification of treatment effects for group-level policies using our proposed model. In Section (ref), we apply the proposed method to analyze the effect of minimum wage on earnings under the 1966 Labour Standards Act. Section (ref) concludes. Online supplementary material encompasses simulation results, additional empirical analysis, technical assumptions for asymptotic theories and proofs of the theorems presented in the main text.

Model and Estimation

In this section, we first provide the model setup and then introduce the estimation method. Before preceding, we introduce some notations. Throughout the paper, let $\| \cdot \|$ denote the Euclidean norm for vectors and the spectral norm for matrices, that is $\| a\| :=\sqrt{a'a}$ and $\|A\|:= \sup_{a \not = 0}\| A a \| /\| a \|$ for column vector $a$ and matrix $A$. Let $I_p$ denote the $p$-dimensional identity matrix. Let $1\{\cdot\}$ denote the indicator function. Let $e_k$ be a unit column vector having 1 at the $k^\text{th}$ entry and 0 for the others, and the dimension of $e_k$ is allowed to vary according to the context. Let $\text{diag}(\cdot)$ denotes the diagonal matrix, whose diagonal entries are given in the parenthesis. We denote $a \vee b:=\max \{a, b\}$ for scalars $a, b$. Let $C_M$ and $c_M$ be some pre-determined positive real numbers which are independent of the sample.

Model

Given group $s=1, \dots S$ and time $t=1, \dots, T$, let $\{(y_{ist}, z_{ist})\}_{i=1}^{N_{st}}$ be repeated cross-sectional observations of a scalar outcome $y_{ist}$ and a $J \times 1$ regressor vector $z_{ist}$ , which includes a constant 1 if needed, for individual $i$ with the sample size $N_{st}$. We denote supports of $y_{ist}$ and $z_{ist}$ by $\mathcal{Y} \subseteq \mathbb{R}$ and $\mathcal{Z} \subseteq \mathbb{R}^J$, respectively.\footnote{ The supports $\mathcal{Y}$ and $\mathcal{Z}$ can be allowed to depend on group and time, while we suppress the dependency for notational simplicity. } Also, we observe a $K \times 1$ vector of group-level covariates $x_{st}$, including a constant 1 if necessary, whose support is $\mathcal{X} \subseteq \mathbb{R}^{K}$, and a dummy variable $d_{st}$ which takes 1 when some policy is employed in group $s$ and time $t$ and 0 otherwise.

We assume that the repeated cross-sectional observations are randomly sampled within each group-time pair, conditional on group-level information. That is to say, interdependency between individauls within paris and dependency across pairs, characterized by both the observed and unobserved group-level information, are allowed. Specifically, we assume that the $u^{\text{th}}$ quantile of the conditional distribution of $y_{ist}$ is given by

align[align omitted — 304 chars of source]

where $Q_{y_{ist}|z_{ist},\alpha_{st}}(u)$ is the $u^{\text{th}}$ conditional quantile of $y_{ist}$ given $(z_{ist},\alpha_{st})$ for $u \in \mathcal{U} \subseteq (0,1)$, $\alpha_{jst}(u)$ is the scalar random quantile coefficient corresponding to the $j^{\text{th}}$ component of the individaul-level observable $z_{ist}$, $\delta_{jt}(u)$ is a scalar coefficient for the policy effect at time $t$, $\beta_{j}(u)$ is a $K \times 1$ vector of coefficients measuring the marginal effect of the group-level observable $x_{st}$ on $\alpha_{jst}(u)$, $f_{jt}(u)$ is an $r \times 1$ vector of unobservable group-level factors, which are heterogeneous across time $t$, ${\lambda}_{js}(u)$ is the corresponding factor loading vector, and $\eta_{jst}(u)$ is an idiosyncratic group-level error satisfying $\mathbb{E}[\eta_{jst}(u)|d_{st},x_{st},f_{jt}(u),\lambda_{js}(u)] =0$. For notational simplicity, we write $Q_{st}(u | z_{ist}) \equiv Q_{y_{ist}|z_{ist},\alpha_{st}}(u)$ in what follows.

In this paper, we consider a binary group-level policy that is employed at a known time $T_0$ onward with $1 < T_0 < T$. Then, we may write model ((ref)) in the vector form as

equation*[equation* omitted — 115 chars of source]

where $A_{js}(u) := [\alpha_{js1}(u),...,\alpha_{jsT}(u)]'$, $D_s := d_s[e_{T_0},...,e_{T}]$ with $d_s=1$ if group $s$ is treated after $T_0$ and $0$ otherwise, and $e_{k}$ being a unit column vector having $1$ at the $k^\text{th}$ entry and 0 for the others, $\delta_{j} := [{\delta}_{jT_0}(u),...,\delta_{jT}(u)]'$, $X_{s}:=[x_{s1}, \dots, x_{sT}]'$, $F_{j}(u):=[f_{j1}(u), \dots, f_{jT}(u)]'$, and $\eta_{js}(u) := [\eta_{js1}(u),...,\eta_{jsT}(u) ]'$.

It is known that, the quantile coefficient $\alpha_{st}(u)$ can be interpreted as the marginal effect of individual covariates $z_{ist}$ on the $u^{\text{th}}$ quantile of outcome variables. Moreover, when the underlying structural model depends on multi-dimension unobservables, sasaki2015quantile shows that the quantile regression coefficients can be interpreted as the marginal effects of $z_{ist}$ on $y_{ist}$ averaged over the unobserved variables that satisfy mild regularity conditions. Thus, quantile regression coefficients $\alpha_{st}(u)$ can succinctly summarize the marginal effect of observed individual characteristics on the outcome, averaged over individual unobserved heterogeneity among observationally equivalent individuals within each group and time. That said, although model ((ref)) doesn't include an individual fixed effect for each given pair of $(s,t)$, it allows for the unobserved heterogeneity across individuals within the group-time pair.

The key interest of this paper lies in studying how these marginal effects depend on the group-level information, especially the group-level policy. To this end, we impose a linear regression model ((ref)) with interactive fixed effects on each $\alpha_{jst}(u)$. The interactive fixed effects structure ${f}_{jt}(u)' {\lambda}_{js}(u)$ accounts for unobserved group and time effects in a flexible way. For example, it captures time-varying macro shocks ${f}_{jt}(\cdot)$ affecting industry or regions differently via ${\lambda}_{js}(\cdot)$. Also, the two-way fixed effects model is included as a special case if ${f}_{jt}(u)=[1, \nu_{jt}(u)]'$ and ${\lambda}_{js}(u)=[\phi_{js}(u), 1]'$.

As an example to showcase the usefulness of the above hierarchical modelling framework, consider the empirical study in this paper where we aim to quantify the effects of a policy $d_{st}$, which varies at the industry-by-year level, on the distribution of individuals' wages $y_{ist}$ across individuals with innate heterogeneities $z_{ist}$, such as race, gender, etc. A policy may have differential effects on lower wage quantiles for black workers than for white workers; the specification of ((ref)) captures this idea by allowing the researcher to specify the coefficient for the racial indicator as a function of the policy, together with other group-level variables. Furthermore, we show in Section (ref) that our model is capable of quantifying the differentials in the policy effects across different subpopulations and quantiles, which facilitates the analysis of policy effects on inequality measures.

In addition, considering a special case where ((ref)) only applies to the intercept in ((ref)), the group and time fixed effects are additive and no group-level unobservables $\epsilon_{st}(u)$ presents, model ((ref))-((ref)) is reduced to a two-way fixed effects model: $$Q_{y_{ist}|z_{ist},\alpha_{st}}(u) = \delta_{jt}(u) d_{st} + x_{st}'\beta(u) + f_t(u) + \lambda_s(u) + z_{ist}'{\gamma}_{st}(u), $$ which has been used for estimating distributional policy effects in empirical studies angrist2004does,

Estimation

We propose a computationally simple two-step estimation approach for model ((ref))-((ref)). In what follows, we consider a finite set of probability levels $\mathcal{U}$, as our analysis mainly focuses on several quantiles and their spreads. Details of the algorithm for each $u \in \mathcal{U}$ are outlined below.

Step 1: Using the individual-level data $\{(y_{ist},{z}_{ist})\}_{i=1}^{N_{st}}$ for each pair of group and time $(s,t)$ separately, we obtain the estimator $\widehat{{\alpha}}_{st}(u)$ of ${\alpha}_{st}(u)$ as the solution of the following quantile minimization problem:

equation*[equation* omitted — 111 chars of source]

where $\varrho_{u}(v) := (u - 1\{v < 0\})v$ for $v \in \mathbb{R}$.

Step 2: Let $\Lambda_j(u) := [\lambda_{j1}(u),...,\lambda_{jS}(u)]'$. Given a collection of the estimators $\widehat{A}_{js}(u) := [\widehat{{\alpha}}_{js1}(u),...,\widehat{{\alpha}}_{jsT}(u)]'$, we obtain the estimator of $\big ( \delta_{j}(u),\beta_j(u),F_j(u),\Lambda_{j}(u) \big )$ by minimizing the following sum of squared residuals:

align[align omitted — 219 chars of source]

with the normalization condition in Assumption (ref).(i), which ensures the identification of factors and their loadings up to an orthogonal rotation matrix with column sign change bai2009panel. The least squares estimators are obtained using an iterated procedure described below. The superscript $m$ indicates the number of iterations and the converged estimator is represented as { $\big(\widehat{\delta}_{j}(u), \widehat{\beta}_j(u), \widehat{F}_j(u),\widehat{\Lambda}_j(u)\big)$.

enumerate[label=(\roman*)] • Using the LS estimator without the factor components, we obtain an initial estimator $(\widehat{\delta}^{(0)}_{j}(u),\widehat{\beta}^{(0)}_j(u))$, which minimizes $\sum_{s=1}^{S}\|\widehat{A}_{js}(u) - D_{s}{\delta}_{j} - X_{s}'\beta_j\|^2$. • Given $\big(\widehat{\delta}^{(m-1)}_{j}(u),\widehat{\beta}^{(m-1)}_j(u)\big)$ for $m \ge 1$, we obtain $\big(\widehat{F}_j^{(m)}(u), \widehat{\Lambda}_j^{(m)}(u) \big)$ as the solution to $$\min_{(F_j,\Lambda_j)}\text{SSR}_{u} \big ( \widehat{\delta}^{(m-1)}_{j}(u),\widehat{\beta}^{(m-1)}_j(u),F_j,\Lambda_j \big )$$ by applying the principle component analysis (PCA) with the normalization conditions in Assumption (ref).(i). • Given $\big( \widehat{F}_j^{(m)}(u), \widehat{\Lambda}_j^{(m)}(u) \big)$, we obtain $\big(\widehat{\delta}_{j}^{(m)}(u),\widehat{\beta}_{j}^{(m)}(u)\big)$ as the minimizer of the objective function $\text{SSR}_{u} \big ( \delta_{j},\beta_j, \widehat{F}_j^{(m)}(u), \widehat{\Lambda}_j^{(m)}(u) \big ) $. • Repeat (ii)-(iii) until numerical convergence is reached. Specifically, we stop the algorithm if $ \big \| \widehat{{\delta}}^{(m)}_{j}(u) - \widehat{\delta}^{(m-1)}_{j}(u) \big \| \leq 10^{-5} $, $ \big \| \widehat{{\beta}}^{(m)}_j(u) - \widehat{{\beta}}^{(m-1)}_j(u) \big \| \leq 10^{-5} $, and $ \big \|\widehat{F}^{(m)}_j(u)\widehat{\Lambda}^{(m)}_j(u)' - \widehat{F}^{(m-1)}_j(u)\widehat{\Lambda}^{(m-1)}_j(u)' \big \| \leq 10^{-5} $.

}

In practice, as the number of factors $r$ is unknown, we use a popular eigen-ratio criterion in PCA for selection. That is, for each $m$ and $j$, we select the number of factors by minimizing the modified eigen-ratio criterion of casas2021time as follows:

align*[align* omitted — 464 chars of source]

where $r_{\max}$ is a pre-specified integer and $\widehat{\rho}^{(m)}_{j,1}(u),...,\widehat{\rho}^{(m)}_{j,T}(u)$ are the estimated eigenvalues of the $T \times T$ matrix $\widehat{L}_j\big(\widehat{\delta}^{(m-1)}_j(u),\widehat{\beta}^{(m-1)}_j(u)\big)$ in descending order, where

align[align omitted — 208 chars of source]

Given that a relatively large $r_{\max}$ suffices, we set it to the cardinality of the set $\{\widehat{\rho}^{(m)}_{j,r}(u): \widehat{\rho}^{(m)}_{j,r}(u)>T^{-1}\sum_{r=1}^{T}\widehat{\rho}^{(m)}_{j,r}(u), r=1, \dots, T\}$ in the empirical and simulation studies.

The proposed two-step estimation method is tailored for the survey data, with the number of individuals within group-time pairs, groups, and time periods being large. It is computationally less demanding relative to jointly estimating all coefficients for large hierarchical data as it only requires estimating a small number of coefficients each time chetverikov2016iv. However, as the proposed model ((ref))-((ref)) is versatile, alternative one-step estimation methods can also be proposed, especially for the case when the number of individuals per group-time is limited such that the first step of the above algorithm is infeasible.

One major difficulty in the one-step estimation is addressing the omitted-variable bias. Let us consider a simple case where $j=1$. We combine ((ref))-((ref)) and write $Q_{y_{ist}|z_{ist},\alpha_{st}}(u) = z_{ist}d_{st} \delta_{t}(u) + z_{ist}x_{st}'\beta(u) + z_{ist}f_{t}(u)' {\lambda}_{s}(u) + z_{ist}\eta_{st}(u)$.\footnote{We compress the subscript $j$ for all group-level coefficients and idiosyncratic error for notational simplicity.} For this model, one may consider a conventional quantile estimator of koenker1978regression through an iterative algorithm similar to ando2020quantile, which updates $\widehat{\delta}_t(u), \widehat{f}_t(u)$ and $\widehat{\beta}_{s}(u),\widehat{\lambda}_s(u)$ recursively until convergence and then obtains $\widehat{\beta}(u)$ as the average of $\widehat{\beta}_s(u)$. Unfortunately, such an algorithm is biased as it ignores the interaction term between $z_{ist}$ and the group-level idiosyncratic error $\eta_{st}(u)$. However, how to deal with the endogeneity issue in the presence of interactive fixed effects remains an open question for future research.

Asymptotic Properties

In this section, we first introduce the main assumptions and then present asymptotic properties of the proposed estimators. The technical assumptions are provided in Online Supplement (ref).

Assumptions

This subsection provides the necessary assumptions for deriving the asymptotic properties of the recursive estimator along with some detailed explanations. In what follows, we consider the case where the set $\mathcal{U}$ consists of finite points, since our empirical application mainly focuses on multiple quantiles and their spreads, instead of the entire distribution.

assumptionFor each fixed $s,t \geq 1$, \begin{enumerate}[label=(\roman*)] • Individual observations $\{(y_{ist}, z_{ist})\}_{i=1}^{N_{st}}$ are independent and identically distributed (i.i.d.) across $i=1,...,N_{st}$ conditional on $(\alpha_{st},x_{st},d_{st},f_t,\lambda_s)$ . The regressor vector $z_{ist}$ satisfies $\|z_{ist}\|<C_M$ almost surely. • All eigenvalues of $\mathbb{E}[z_{1st}z_{1st}']$ are bounded from below by $c_M>0$. \end{enumerate}

Assumption (ref).(i) imposes a conditional i.i.d. assumption within group-time pairs. It allows for interdependency within group-time pairs that are fully controlled by group-level information, as well as the dependency between individual observations and group-level characteristics. Such an assumption is typically applicable to survey studies characterized by hierarchical data and is also considered in chetverikov2016iv. Assumption (ref).(ii) is a familiar identification condition in regression analysis.

assumptionConsider $(s,t) \in \{1,...,S\} \times \{1,...,T\}$, for all $y \in (z'_{1st}\alpha_{st}(u)-c_M,z'_{1st}\alpha_{st}(u)+c_M)$ for some $c_M>0$, it satisfies \begin{enumerate}[label=(\roman*)] • the conditional density function $g_{st}(y)$ is continuously differentiable with the first-order derivative $g_{st}'(\cdot)$ satisfying $|g_{st}'(y)| \leq C_M$ and $|g_{st}'({z}_{1st}'{\alpha}_{st}(u))| \allowbreak\geq c_M$. • $g_{st}(y) \leq C_M$, and $g_{st}({z}_{1st}'{\alpha}_{st}(u)) \geq c_M$ for some $c_M>0$. \end{enumerate}

Assumptions (ref) is a set of mild regularity conditions that are typically imposed in the quantile regression literature koenker2004quantile,chetverikov2016iv.

assumptionLet $\Lambda_j(u):=[{\lambda}_{j1}(u),...,{\lambda}_{jS}(u) ]'$. For all $(s, t) \in \{1, \dots, S\} \times \{1, \dots, T\}$, \begin{enumerate}[label=(\roman*)] • $T^{-1}F_j(u)'F_j(u) = I_r$ and $S^{-1}\Lambda_j(u)'\Lambda_j(u)$ is a positive-definite diagonal matrix. • $\mathbb{P}(d_{s}=1)$ is bounded away from below by $c_M>0$ and from above by $C_M<1$. • The eigenvalues of $\mathbb{E}[x_{st}x_{st}'] $ are bounded away from zero. \end{enumerate}

Assumption (ref) provides the identification conditions of the regression coefficients, factor and loadings. Specifically, Assumption (ref).(i) guarantees the identification of factor and the loadings up to an orthogonal rotation matrix. Such assumption is standard in the literature of the mean panel factor models bai2009panel,jiang2020recursive and quantile factor models ando2020quantile,chen2021quantile. Assumption (ref).(ii) ensures the identifiability of the policy parameter $\delta_{jt}(u)$, which is standard for policy evaluation hsu2022estimation,noh2023nonparametric. Assumption (ref).(iii) is a conventional assumption in regression analysis bai2009panel to guarantee the identification of regression coefficients $\beta_{j}(u)$.

assumptionFor all $(s,t) \in \{1,...,S\}{\times}\{1,...,T\}$, \begin{enumerate}[label=(\roman*)] • $\mathbb{E}[||{\lambda}_{js}(u)||^4] \leq C_M$. • $\mathbb{E}[\eta_{j s t}(u) | d_{gl}, x_{g l}, {\lambda}_{jg}(u), f_{jl}(u)]=0$ for all $(g,l) \in \{1,...,S\}\times \{1,...,T\}$. • The largest eigenvalue of the $T \times T$ matrix $\mathbb{E}[{\eta}_{js}(u){\eta}_{js}(u)'] $ is bounded uniformly in $s$ and $T$. \end{enumerate}

Assumption (ref).(i) requires standard moment conditions for our analysis. Assumption (ref).(ii) and (iii) impose weak restriction on the correlation among the idiosyncratic error components, group-level regressors and common factors. These assumptions are often imposed in the factor model literature bai2009panel,jiang2020recursive.

assumptionLet $N_{\min} := \min\{N_{st},s=1,...,S,t=1,...,T\}$. As $S,T \rightarrow \infty$, we have (i) $T/S \to \kappa > 0$ and (ii) $(ST)^{3/4}(\ln(N_{\min})/N_{\min})^{1/2} \leq C_M$.

Assumption (ref) controls the diverging rates of the number of groups $S$, time $T$ and individuals per group and time $N_{st}$. Assumption (ref).(i) is standard in the panel factor model literature bai2009panel. Assumption (ref).(ii) requires that the number of individuals per group-time pair grows sufficiently fast as $S$ and $T$ jointly go to infinity, such that the estimation error from the quantile estimation in the first-step is negligible. Compared with Assumption 3 of chetverikov2016iv, Assumption (ref).(ii) imposes a more explicit yet comparable growth rate, which is necessary in analyzing the limiting property of the interactive fixed effects estimator. Although the growth rate seems restrictive, we have shown with both simulation and empirical study that the estimation procedure is valid in practice as long as the minimum number of individuals per group $N_{\min}$ used to perform the first step estimation is comparable to the total number of groups and time ($S \times T$). For example, in the empirical study, we have $N_{\min} = 229$ while $S \times T = 323$, and the estimation algorithm converges within a small number of iterations.

Define the $J \times 1$ vector $K_t(u) := [K_{1t}(u),...,K_{Jt}(u)]'$, whose $j^{\text{th}}$ element is given by $$ K_{jt}(u) := \big(S^{-1}\sum_{s=1}^{S} R_{js}(u)^2\big)^{-1} S^{-1/2} \sum_{s=1}^{S}R_{js}(u)\eta_{jst}(u), $$ where $R_{js}(u) := d_{s} - S^{-1}\sum_{g=1}^{S}\omega_{j,sg}(u)d_g$ with $\omega_{j,sg}(u) := {\lambda}_{jg}(u)' \big( S^{-1}\Lambda_j(u)' \Lambda_j(u) \big)^{-1} {\lambda}_{js}(u)$.

assumptionFor any $u_1,u_2 \in \mathcal{U}$ and $t \geq T_0$, as $S \to \infty$, we have \begin{align*} \bigg [ \begin{array}{c} K_t(u_1) \\[-0.5cm] K_t(u_2) \end{array} \bigg ] \overset{d}{\rightarrow} \mathcal{N} \Bigg( \mathbf{0}, \bigg[ \begin{array}{c} \Sigma_t(u_1,u_1), \Sigma_t(u_1,u_2)\\[-0.5cm] \Sigma_t(u_2,u_1), \Sigma_t(u_2,u_2) \end{array} \bigg] \Bigg ), \end{align*} where $\mathbf{0}$ is a $(2J) \times 1$ vector and $\Sigma_t(u_1,u_2) := \underset{S \to \infty}{\lim} \mathbb{E}[ K_t(u_1) K_t(u_2)']$.

Assumption (ref) is required to derive the joint Central Limit Theorem (CLT) in Theorem (ref) for the convergence of the estimator of the policy parameter, given quantile levels $u_1,u_2 \in \mathcal{U}$. The assumption shares the same idea as Assumption E of bai2009panel.

Asymptotic Results

In this subsection, we present the asymptotic properties of our proposed estimators. To maintain focus on the key parameters of interest, we omit the CLT for $\beta_{j}(u)$, though it can be derived similarly to that for $\delta_{jt}(u)$. Also, for empirical interests, we derive the corresponding consistent estimator of the asymptotic covariance matrix to facilitate the construction of confidence intervals.

theoremSuppose that Assumptions (ref)-(ref) and (ref)-(ref) hold. Then, for any fixed $u\in \mathcal{U}$, $j=1,...,J$ and $m \geq 0$, as $S,T \to \infty$, we have \begin{enumerate}[label=(\roman*)] • $ \sqrt{S} \big ( \widehat{\delta}^{(m)}_{jt}(u) - {\delta}_{jt}(u) \big )= O_P(1)$ \ for each $t \geq T_0$. • $ \sqrt{ST} \big ( \widehat{\beta}^{(m)}_j(u) - {\beta}_j(u) \big ) = O_P(1)$. \end{enumerate}

The time-varying policy effects are estimated for each time period following the policy intervention. The convergence rate of the policy effect estimator depends solely on the group size $S$. On the other hand, the remaining regression coefficients are estimated using the full sample, and their convergence rate depends on $ST$. The convergence rates are slower than those of the conventional one-step quantile estimator, which could typically achieve $\sqrt{STN_{\min}}$. However, as explained at the end of the Section (ref), the existing one-step quantile estimator suffers from the omitted-variable bias.

For the purpose of inference, we then establish the joint CLT of the estimator $(\widehat{\delta}_t(u_1)',\widehat{\delta}_t(u_2)')'$ in the following theorem.

theoremSuppose that Assumptions (ref)-(ref) and (ref)-(ref) hold. Let $\widehat{\delta}_{\cdot t}(u) := [\widehat{\delta}_{1t}(u),...,\widehat{\delta}_{Jt}(u)]'$ denote the estimator of $\delta_{\cdot t}(u)$. Then, for any $u_1,\ u_2 \in \mathcal{U}$ and $t \geq T_0$, we have, as $S,T \to \infty$, \begin{align*} \sqrt{S} \bigg[ \begin{array}{l} \widehat{\delta}_{\cdot t}(u_1) - {\delta}_{\cdot t}(u_1) \\[-0.5cm] \widehat{{\delta}}_{\cdot t}(u_2) - {\delta}_{\cdot t}(u_2) \end{array} \bigg] \overset{d}{\rightarrow} \mathcal{N} \Bigg( \bigg [ \begin{array}{l} B_t(u_1) \\[-0.5cm] B_t(u_2) \end{array} \bigg ] , \bigg[ \begin{array}{c} \Sigma_t(u_1,u_1), \Sigma_t(u_1,u_2)\\[-0.5cm] \Sigma_t(u_2,u_1), \Sigma_t(u_2,u_2) \end{array} \bigg] \Bigg ), \end{align*} where $\Sigma_t(u_1,u_2)$ is defined in Assumption (ref), and $B_t(u) := \underset{S,T \to \infty}{\operatorname*{plim}} [\widetilde{B}_{1t}(u),...,\widetilde{B}_{Jt}(u)]'$ is the bounded asymptotic bias, whose $j^{\text{th}}$ component is given by \begin{align*} \widetilde{B}_{jt}(u) := - \bigg(\frac{1}{S}\sum_{s=1}^{S} R_{js}(u)^2\bigg)^{-1} \frac{1}{S^{3/2}T}\sum_{s,g=1}^{S}d_s \mathbb{E}[{\eta}_{jgt}(u) {\eta}_{jg}(u)'] F_j(u) \bigg(\frac{\Lambda_j(u)'\Lambda_j(u)}{S}\bigg)^{-1}\lambda_{js}(u), \end{align*} for $j=1,...,J$, where $R_{js}(u)$ is defined above Assumption (ref).

In view of this theorem, under general cases, the asymptotic distribution of the recursive estimator $\widehat{\delta}_{jt}(u)$ depends on: (i) the quantiles, (ii) the accuracy of the first-step estimation (i.e., $\widehat{\alpha}_{jst}(u) - \alpha_{jst}(u)$), (iii) the consistency of the initial estimation of the second-step, and (iv) the degenerating estimation error of the regression coefficient, factor and loadings carried over from the iterative steps. Therefore, although the CLT is derived per given time $t$, we still require $S,T \to \infty$ jointly, as the convergence of the policy parameter estimator relies on the convergence of the estimators of the factors and loadings, which only hold when $S,T \to \infty$ jointly.

In addition, we note that the estimation error of the first-step $\widehat{\alpha}_{jst}(u) - \alpha_{jst}(u)$ depends on the sample size of individual observations within each group-time pair ($N_{st}$). Hence, by controlling the relative growth rate between individual-level and group-level sample size (Assumption (ref).(ii)), the asymptotic first-step estimation error becomes negligible in the asymptotic representation of $\widehat{\delta}_{jt}(u)$, and the estimation error from the group-level regression contributed to the asymptotic bias and covariance, similar to bai2009panel. We also note that the growth rate of $N_{\min}$ relative to $S$ and $T$ can be relaxed, but at the expense of more complicated asymptotic expressions.

We now provide a way to estimate the asymptotic bias and covariance given in Theorem (ref), and subsequently, Proposition (ref) below establishes their consistency.

Following bai2009panel, we define $\widehat{B}_t(u) := [\widehat{B}_{1t}(u),...,\widehat{B}_{Jt}(u) ]'$, where, for $j=1,...,J$,

align*[align* omitted — 308 chars of source]

where $\widehat{R}_{js}(u) := d_s -S^{-1}\sum_{g=1}^{S}d_g \widehat{\lambda}_{jg}(u)' \big( S^{-1}\widehat{\Lambda}_j(u)' \widehat{\Lambda}_j(u) \big)^{-1} \widehat{\lambda}_{js}(u)$. We construct an estimator of the asymptotic covariance matrix $\Sigma_t(u_1,u_2)$, denoted as $\widehat{\Sigma}_t(u_1,u_2)$, by its empirical counterpart. $\widehat{\Sigma}_t(u_1,u_2)$ is a $J \times J$ block matrix, whose $(j,k)^{\text{th}}$ block is given by

align*[align* omitted — 279 chars of source]
propositionSuppose that the conditions of Theorem (ref) hold. In addition, we assume that for any fixed $u_1,u_2 \in \mathcal{U}$ and $j,k=1,...,J$, for $t,l \in \{1,...,T\}$ and $s,g \in \{1,...,S\}$, $$\mathbb{E}\big[\eta_{jst}(u_1) \eta_{kgl}(u_2)\big|D_s,D_g,W_s,W_g, {\Lambda}_{j}(u_1), {F}_{j}(u_1),{\Lambda}_{k}(u_2), {F}_{k}(u_2)\big] = 0$$ if $s \neq g$ or $t \neq l$. Then, for any given $u_1, u_2 \in \mathcal{U}$ and $t \geq T_0$, we have (i) $\widehat{B}_t(u) \overset{p}{\rightarrow} B_t(u)$, and (ii) $\widehat{\Sigma}_t(u_1,u_2) \overset{p}{\rightarrow} \Sigma_t(u_1,u_2)$ as $ S,T \rightarrow \infty$.

Proposition (ref) assumes that the idiosyncratic errors are uncorrelated across groups and over time, after conditioning group-level regressors and interactive fixed effects. For the correlated idiosyncratic errors, the analytical expression for the consistent estimator is hard to derived. bai2009panel provides some conjectures for bias-correction and covariance estimators using the partial sample method together with the Newey-West procedure. Recently, yan2023bootstrap propose a wide dependent bootstrap method to consistently estimate the asymptotic covariance matrix when both serial and cross-sectional dependences exist.

Treatment Effects for Group-Level Policy

Measuring the impact of policy interventions is a central interest in economic and social studies. Our proposed modeling framework is capable to serve this purpose. To see this, we consider the true data generating process under the potential outcome framework of rubin1974estimating, and provide identification results for several policy effect parameters.

We consider a binary group-level policy that is employed at a known time $T_{0}$ onward with $1 < T_{0} < T$. Thus, the sample periods can be divided into the before-period ($t < T_{0}$) and after-period ($t \geq T_0$). Also, let $d_{s} = 1$ if group $s$ is treated after $T_{0}$ and 0 otherwise, which implies that $d_{st} = d_s 1\{t \geq T_0\}$.\footnote{ Our analysis and application concentrate on a single policy change event, rather than staggered or sequential policy changes. The latter setup, requiring the causally interpretable estimates framework as discussed in callaway2021difference, sun2021estimating, athey2022design among others, fall outside the scope of our study.} Let $y_{ist}^{1}$ and $y_{ist}^{0}$ denote the individual potential outcomes with and without exposure to the group-level policy ($d_{st} = 1$ and $d_{st} = 0$), respectively. Then, the observed outcome is written as

eqnarray*[eqnarray* omitted — 72 chars of source]

Correspondingly, under treatment status $d_{st} = d \in \{0,1\}$, the $u$\textsuperscript{th} conditional quantile of the potential outcome $y_{ist}^{d}$ is given by

equation[equation omitted — 132 chars of source]

Also, we suppose that the group-level treatment affects the potential conditional quantile through the potential (random) marginal effects $\alpha_{st}^{d}(u)$. That is, we specify the $j^{\text{th}}$ element of $ \alpha_{st}^{d}(u)$ given the treatment status $d_{st} = d$ as

eqnarray[eqnarray omitted — 173 chars of source]

where $\Delta_{jst}(u)$ represents the random policy effects. The correlation between $d_{st}$ and the factor component is unrestricted so that the selection into the treatment can be correlated with factor loadings $\lambda_{js}(u)$. Additionally, the implementation of the policy can be dependent on the aggregate shocks characterized by $f_{jt}(u)$. Below, we first introduce three treatment parameters-of-interest, and then propose the identification results.

Treatment Parameters

For a given probability level $u$ and individual-level regressor $z$, the conditional quantile $Q_{st}^d(u|z)$ is a random variable depending on the policy status and additional group-level characteristics. Arellano2016 propose an average of conditional quantile treatment effect to measure treatment effects in non-linear response models. Extending their idea, we consider the average quantile treatment effect on the treated (AQTT) at time $t \ge T_{0}$, defined as:

eqnarray*[eqnarray* omitted — 111 chars of source]

AQTT can be considered as an extension of the treatment effect measure on unconditional quantiles, which is employed in numerous empirical studies lee1999wage,angrist2004does,bitler2006mean. The extensive body of research on quantile treatment effects, as illustrated by callaway2019quantile,wuthrich2020comparison, typically measures treatment effects as the difference of the quantile functions of responses from the treated and untreated groups. However, the AQTT approach differs by accounting for the heterogeneity in the conditional quantile function $Q_{st}^d(u|z)$ across various groups and time periods.

As an alternative measure, we consider spreads of conditional quantile functions to quantify inequality within and between collections of individuals characterized by the individual-level regressors. Given the policy status $d \in \{0, 1\}$, we fix individual-level regressors $z \in \mathcal{Z}$ and consider two probability levels of interest $u_{1}, u_{2} \in (0,1)$ with $u_{2} > u_{1}$. Then, a within-inequality measure under the policy status $d_{st} = d$ at time $t \ge T_{0}$ is defined as the spread of conditional quantiles: $$ \Delta_{st}^{W, d}(u_{1}, u_{2}|z) := Q_{st}^{d}(u_{2}|z) - Q_{st}^{d}(u_{1}|z). $$ Similarly, we fix individual attributes $z_{1}, z_{2} \in \mathcal{Z}$ and a probability level $u$ to define a between-inequality measure under the policy status $d$ at time $t \ge T_{0}$ as $$ \Delta_{st}^{B, d}(u |z_{1}, z_{2}) := Q_{st}^{d}(u|z_{2}) - Q_{st}^{d}(u|z_{1}). $$

Figure (ref) illustrates the within- and between-inequalities. The within-inequality measures the dispersion of the distribution of the outcome conditional on individual characteristics $z$ by using two conditional quantile functions, whereas the between-inequality measures the distance between two conditional distributions at a certain probability level.

figure[figure omitted — 921 chars of source]

Group-level policies can affect these inequality measures and their impact can be quantified as changes in the inequality measures at time $t$ averaged over treated groups:

eqnarray*[eqnarray* omitted — 409 chars of source]

Identification

We now exhibit conditions, under which model ((ref))-((ref)) on the observed outcome allows for the identification for treatment parameters. Under a similar setup, gobillon2016regional prove the identification of average policy effects using the mean regression model. We make the following assumptions. Throughout the assumptions, we fixed a given $j=1,...,J$ and $u \in \mathcal{U}$, unless stated otherwise.

assumptionFor all $(s, t) \in \{1, \dots, S\} \times \{1, \dots, T\}$ and $j=1,...,J$, \begin{enumerate}[label=(\roman*)] • $\mathbb{E}[\eta_{jst}^d(u)|d_{st}, X_{s}, \lambda_{js}(u), F_{j}(u)] = \mathbb{E}[\eta_{jst}^d(u)|X_{s}, \lambda_{js}(u), F_{j}(u)] = 0$ for $d=0,1$. • $\mathbb{E}[\Delta_{jst}(u)|d_s=1,X_s] = \mathbb{E}[\Delta_{jst}(u)|d_s=1]$. \end{enumerate}

Assumption (ref).(i) requires that the error term for the potential outcome is mean-zero and mean-independent of the treatment status, conditional on group-level observed and unobserved variables. Assumption (ref).(i) implies a type of parallel-trend assumption. For clarification, we first note that the assumption implies that, for $t \geq T_0$, we have $\mathbb{E}[\eta^0_{jst}(u) - \eta^0_{js,T_0-1}(u)|d_{s}=1,X_{s}, \lambda_{js}(u), F_{j}(u)] = \mathbb{E}[\eta^0_{jst}(u) - \eta^0_{js,T_0-1}(u)|d_s=0,X_{s}, \lambda_{js}(u), F_{j}(u)]$. By the iterative law of expectation, this immediately leads to $ \mathbb{E}[\alpha_{jst}^0(u|z) - \alpha_{js,T_0-1}^0(u|z)|d_s = 1] = \mathbb{E}[\alpha_{jst}^0(u|z) - \alpha_{js,T_0-1}^0(u|z)|d_s = 0] $. Essentially, it states that, without treatment, the changes in the marginal effects of any variable $j$ don't depend on whether the individual belongs to a treated group or not on average. Taking the analysis of the minimum wage policy as an example, this assumption suggests that the changes in the marginal effect of gender (or race, education, etc,) over time are the same for individuals in the treated and control industries, supposing that the policy is not introduced.

Assumption (ref).(ii) is a technical assumption that is also considered by gobillon2016regional. It is imposed to fulfil the technical requirement that $\mathbb{E}\big[\Delta_{jst}(u) - \mathbb{E}[\Delta_{jst}|d_s=1] \big| d_{st},x_{st}\big] = 0$. It assumes that the random policy effects are mean-independent of the group-level covariates within the treated group. We note that the same as the case of gobillon2016regional, this assumption is stronger than necessary. However, as the empirical study in this paper is in the absence of $x_{st}$, generalization of this assumption is out of the scope of the current paper. One may generalize the model by interacting covariates with the treatment indicator, and this would substantially weaken this condition caetano2022difference.

In this paper, we do not impose the standard yet restrictive rank invariance imbens2007nonadditive, or less restrictive rank similarity chernozhukov2005iv assumptions on the individual unobserved heterogeneities, which requires an individual's rank in the potential outcome distribution to be the same or has the same probability distribution across treatment status. Instead, our identification results directly rely on the quantile specifications of the potential outcome. However, we do note that, if the rank preservation assumption holds up, AQTT can be additionally interpreted as individual causal effect for (the same) individual at $u^{\text{th}}$ quantile before and after treatment, and quantile treatment effect parameters can be identified accordingly.

The theorem below shows that we can identify the time-varying distributional impact of a group-level policy using model ((ref))-((ref)).

theoremSuppose that Assumptions (ref)-(ref) and (ref) hold. Then, for a given $t \geq T_0$, and $(u, z) \in \mathcal{U} \times \mathcal{Z}$, we have \begin{eqnarray*} \Delta_{t}^{AQTT}(u|z) = z'\delta_{\cdot t}(u), \end{eqnarray*} and, for $u_{1}, u_{2} \in \mathcal{U}$ and $z_{1}, z_{2} \in \mathcal{Z}$, \begin{eqnarray*} \accentset{ .}{\Delta}_{t}^{B}(u| z_{1}, z_{2}) = (z_{2} - z_{1})'\delta_{\cdot t}(u) \ \ \ \ \ \mathrm{and} \ \ \ \ \ \accentset{ .}{\Delta}_{t}^{W}(u_{1}, u_{2}| z) = z'\big (\delta_{\cdot t}(u_{2}) - \delta_{\cdot t}(u_{1}) \big). \end{eqnarray*} Here, $\delta_{\cdot t}(u) := [\delta_{1t}(u),...,\delta_{Jt}(u)]'$, whose $j^{\text{th}}$ element $\delta_{jt} := \mathbb{E}[\Delta_{jst}(u) | d_s = 1]$ can be identified as $ \delta_{jt}(u) = \mathbb{E}[d_{st}\Pi_{st}]^{-1}\mathbb{E}[\Pi_{st}(\alpha_{jst}(u) - f_{jt}(u)'\lambda_{js}(u))] $ with $\Pi_{st} := d_{st} - \mathbb{E}[d_{st}x_{st}']\mathbb{E}[x_{st}x_{st}']^{-1}x_{st} $.

The above result shows that we can identify the treatment effects of group-level policy which are allowed to vary according to individuals' observed and unobserved characteristics. Taking into account the interplay between a group-level policy and individuals' characteristics, our framework can explicitly identify heterogeneous impacts of the policy across individuals sharing the same observed regressors $z$ and also the impact on the within- and between-inequalities among individuals. To simplify the proof, we treat factor and loadings as observed. When they are unobserved, the iterative estimation approach, proposed in Section (ref), can be used to obtain their estimators.

According to the identification Theorem (ref), we define the estimators of the treatment parameters as $\widehat{{\Delta}}^{AQTT}_t(u) := z'\widehat{\delta}_{\cdot t}(u)$, $\widehat{ \accentset{\mbox{\large\bfseries .}}{\Delta}}_{t}{\hspace{-0.1cm}}^{B}(u| z_1,z_2) := (z_2-z_1)'\widehat{\delta}_{\cdot t}(u)$, $\widehat{ \accentset{\mbox{\large\bfseries .}}{\Delta}}_{t}{\hspace{-0.1cm}}^{W}(u_{1}, u_{2}| z) := z'(\widehat{\delta}_{\cdot t}(u_2) - \widehat{\delta}_{\cdot t}(u_1))$. Then, we establish the CLT results for the estimators of the treatment parameters in the next theorem.

theoremUnder the conditions of Theorems (ref) and (ref), for any $u, u_1,\ u_2 \in \mathcal{U}$, $z,z_1,z_2 \in \mathcal{Z}$ and $t \geq T_0$, we have, as $S,T \to \infty$, \begin{align*} \sqrt{S} \Big ( \widehat{{\Delta}}^{AQTT}_t(u|z) - {\Delta}^{AQTT}_t(u|z) \Big ) & \overset{d}{\rightarrow} \mathcal{N} \Big ( z' B_t(u) , z'\Sigma_t(u,u)z \Big ), \\ \sqrt{S} \Big ( \widehat{ \accentset{ .}{\Delta}}_{t}^{B}(u, | z_1,z_2) - \accentset{ .}{\Delta}_{t}^{B}(u| z_1,z_2) \Big ) & \overset{d}{\rightarrow} \mathcal{N} \Big ( \big(z_2-z_1\big)' B_t(u), \sigma_{B,t}^{2}(u, z_{1}, z_{2}) \Big ),\\ \sqrt{S} \Big ( \widehat{ \accentset{ .}{\Delta}}_{t}^{W}(u_{1}, u_{2}| z) - \accentset{ .}{\Delta}_{t}^{W}(u_{1}, u_{2}| z) \Big ) & \overset{d}{\rightarrow} \mathcal{N} \Big ( z'\big( B_t(u_2) - B_t(u_1) \big), \sigma_{W,t}^{2}(u_{1},u_{2}, z) \Big ), \end{align*} where $B_t(u)$ is defined in Theorem (ref), $ \sigma_{B,t}^{2}(u, z_{1}, z_{2}) := (z_2-z_1)' \Sigma_t(u,u)(z_2-z_1) $ and $ \sigma_{W,t}^{2}(u_{1},u_{2}, z) := z'\big(\Sigma_t(u_1,u_1)-\Sigma_t(u_1,u_2)-\Sigma_t(u_2,u_1)+\Sigma_t(u_2,u_2)\big)z $ with $\Sigma_t(u_1,u_2)$ in Assumption (ref).

Recall that Proposition (ref) proposes the consistent estimators of the asymptotic bias and covariance. Using this result with Theorem (ref), it is straightforward to construct the confidence intervals for the time-varying AQTT, changes in between- and within-inequality measures.

Empirical Analysis

Background on Racial Income-Inequality

Racial economic inequalities have persisted in the United States over long periods. Among these inequalities, the income gap between black and white workers is evident. As in Figure (ref), the income gap, measured by the average annual earnings, was around 20-30% for the last two decades, whereas the gap significantly dropped during the late 1960s and early 1970s. The empirical literature has explored factors that narrowed the racial income gap during those periods, including federal anti-discrimination legislation smith1984affirmative and improvements in education smith1977black,card1992school.

figure[figure omitted — 754 chars of source]

Recently, derenoncourt2021minimum put forward a new explanation: the extension of the federal minimum wage to some industries. The 1966 FLAS established a federal minimum wage (effective February 1967) in previously unregulated industries, which employed about $20\%$ of the total workforce in the US and nearly a third of all black workers. They evaluate the minimum wage policy effect on earnings, using a cross-industry difference-in-differences design, in which eight treated and eight control industries were subject to the minimum wage under the 1966 and 1938 FLSA, respectively. Additional information and background about the dataset are provided in the Online Supplement (ref).

Using repeated cross-sections of black and white workers aged between 25 and 55 for years 1961 and 1963--1980, extracted from March CPS,\footnote{ Since the March CPS of year $t$ contains information in calendar year $t-1$, the data source is the 1962, 1964--1981 March CPS. The 1963 March CPS is excluded due to the lack of observations and missing demographic information. } they estimate the following two-way fixed effects mean regression model:

eqnarray[eqnarray omitted — 128 chars of source]

for worker $i$ in industry $s=1, \dots, 16$ and time $t = 1961,1963 \dots, 1980$. Here, $y_{ist}$ is the log annual earnings deflated by annual CPI-U-RS (\$2017)\footnote{ Using March CPS data in the 1960s and early 1970s, we only directly observe annual earnings, but not hourly wages, whereas the CPS contains more detailed individual worker–level information the Bureau of Labour Statistics data. See Section III.B. of derenoncourt2021minimum. }, and $d_{st}$ denotes a dummy variable taking 1 if industrial sector $s$ is subject to the federal minimum wage, and 0 otherwise. Also, $z_{ist}$ is a vector of worker's characteristics, and the unobserved random variables consist of industry fixed effect $\phi_{s}$, time fixed effect $\nu_{t}$ and an idiosyncratic error $\eta_{ist}$. The parameter of interest is $\delta_{t}$, which measures dynamic policy effects.

Their result shows that, after controlling for individual characteristics, the average wage of workers in the newly covered industries is around $5\%$ higher relative to that in control industries in 1967--1980 compared with the pre-period 1961--1966, and the effect of minimum wage reform on workers' log-earning is more than twice as large for black workers as that for white workers on average. In addition to the above regression, they present several regression results to uncover intricate facets of the effects of minimum wage by taking various variables as the dependent variable in ((ref)), including log annual wage or its unconditional quantiles, and also selecting sub-samples based on workers' characteristics.

Model and Practical Implementation

In this paper, we analyze time-varying policy effects from 1967 to 1980 at quantile $u \in \{0.1, 0.3, 0.5, 0.7, 0.9\}$.

We use the same dataset as derenoncourt2021minimum. However, we note that as the Forestry and Fishing industry only contains 35 individual observations per year on average, which could lead to large first-step quantile estimation error, we remove it from the treated industries. Thus, there are seven treated industries and eight controlled industries in the following study. We provide an elementary explanatory data analysis in Appendix (ref). Given the model of derenoncourt2021minimum, we consider the following $u^{\mathrm{th}}$ quantile regression model in ((ref)) with $\alpha_{st}(u) = [\alpha_{1st}(u), \dots, \alpha_{Jst}(u)]'$, where, for $j=1, \dots, J$,

eqnarray[eqnarray omitted — 154 chars of source]

Here, a set of coefficients $\{\delta_{jt}(u)\}_{t=1967}^{1980}$ measures the time-varying policy effect. Model ((ref)) does not involve any individual-level covariates since such information is not available in the March CPS dataset. We incorporate interactive fixed effects into the model to account for the industrial and temporal heterogeneities. The industry-specific loadings can be interpreted as latent industrial factors.

For simplicity of interpretation, we treat some of the original covariates as ordered variables rather than dummy variables for categories. More precisely, $z_{ist}$ includes a constant $1$, dummy variables for race (white/black), gender (male/female) and work type (full-time/part-time), and ordered variables including years of schooling, experience, experience squared, the number of weeks worked in a year, and the number of hours worked in a week.\footnote{ derenoncourt2021minimum use dummy variables to control for the number of weeks worked in a year and the number of hours worked in a week, because hourly wage is not available in the CPS data during the periods of interest.} This selection yields a very similar result to that of the mean-regression in ((ref)) as the one in derenoncourt2021minimum. See Figure (ref). In this empirical study, we are particularly interested in the heterogeneous policy-effects varies between gender and race.

figure[figure omitted — 531 chars of source]

Results Analysis

In Figure (ref), we present time-varying policy effects $\delta_{jt}(u)$ in ((ref)) with $95\%$ confidence intervals for the categorical individual-level covariates. The policy effects for the continuous covariates: education and experience are insignificant across quantiles. That said, after taking the differentiation between gender and race into account, the policy effect is insignificant across the levels of skill (measured by education and experience). For presentation conciseness, those figures are omitted.

Panel (a) reports the effects on the intercept coefficients, which correspond to white, male, full-time workers, which are insignificantly different from zero for most of the estimates. Panel (b) shows statistically significant positive policy effects for black workers in the majority of the years across all quantiles. Especially, the policy effects are most significant at the 0.1th conditional quantile, which is 15--20% (0.15--0.2 log points). Panel (c) presents the estimated policy effects on the female dummy's coefficient. The effects in the late 1970s are positive and significant up to 0.7th quantile with a magnitude of 5%. However, caution is warranted in interpreting the results, which may be an integrated impact of the 1966 Fair Labour Standards Act and two pieces of important legislation that targeted labour market discrimination against women: the Equal Pay Act of 1963 and Title VII of the Civil Rights Act of 1964.\footnote{ The Equal Pay Act of 1963 is a federal law that amends the Fair Labour Standards Act and prohibits wage disparity based on gender. Title VII of The Civil Rights Act of 1964 more broadly prohibits discrimination in employment on the basis of race, colour, religion, national origin, and gender.} Although the gender gap of median wages for full-time, full-year workers was unchanged over the 1960s--1970s blau2017gender, bailey2021changes recently document sharp increases in women's wages relative to men's below median during the 1960s. They underscore the importance of minimum wage policy and the laws to target gender-based workplace discrimination. Our result is consistent with their findings and further suggests long-term positive effects even at 0.5th and 0.7th quantiles, conditional on individual attributes.

figure[figure omitted — 1,037 chars of source]

In Figure (ref), we present estimated policy effects on changes in the within-inequality measure $ \accentset{\mbox{\large\bfseries .}}{\Delta}_{t}^{W}(u_{1}, u_{2}|z)$ in the year $1980$. For quantile pairs $(u_{1}, u_{2}) = (0.1, 0.9)$ or $(0.1, 0.5)$, we measure how much the minimum wage policy changes the conditional quantile spread. In the following analysis, we consider the labour with an average skill level (12 years of education and 20 years of experience), while reporting three pairs of categorical individual attributes, as shown in the horizontal axis. The estimates suggest that the introduction of minimum wage reduces the within-inequality, while the reduction is insignificantly different from 0 for all subpopulations. In addition, we also note that the same conclusion can be drawn when considering less-skilful or more-skilful labours within the three sub-populations considered in Figure (ref), since the estimated $\delta_{jt}(u)$ associated with education and experience are insignificantly different from 0.

figure[figure omitted — 1,089 chars of source]

Figure (ref) reports the policy effects on the changes in the between-inequality measure, $\dot{\Delta}_{t}^{B}(u|z_{1}, z_{2})$, for $t =1980$ and $u \in \{0.1, 0.3, 0.5, 0.7, 0.9\}$. The baseline $z_{2}$ is fixed to include white, male workers and $z_{1}$ changes over the three pairs of categorical attributes as in Figure (ref), while the remaining variables are the same in $z_{1}$ and $z_{2}$.\footnote{According to the identification result in Theorem (ref), $\dot{\Delta}_{t}^{B}(u|z_{1}, z_{2}) = (z_{2} - z_{1})'\delta(u)$, the common values in $z_{1}$ and $z_{2}$ cancel each other and do not affect the conclusions.} Panel (a) suggests significant negative impacts on the between-inequality for black, male workers, compared to white, male workers, with magnitudes 5--20% for all quantiles except the 0.7th conditional quantile. Panel (b) also shows reduction (0.05--0.20) in the between-inequality for white, female workers at the 0.7th conditional quantile and below. Panels (c) plots results for female, black workers, and the estimated changes in the between-inequality range from -0.05 to -0.35 at the 0.7th conditional quantile and below.

figure[figure omitted — 1,194 chars of source]

To illustrate the robustness of the significant policy effects in reducing the between-equality, Figure (ref) plots the changes in between-inequality for black, female workers from 1967 to 1980. In quantiles up to medium, the magnitude of reduction in the between-inequality increases as time increases. In addition, such policy effects are significant after 1970. Similar patterns are also witnessed in the other two subpopulations presented in Figure (ref). We choose not to report those plots for the concise of presentation.

figure[figure omitted — 589 chars of source]

Overall, the results above confirm the findings of derenoncourt2021minimum that the reform was effective in improving the black economic status and reducing the racial income gap. In addition, we provide empirical evidence of a compounded impact of the policy effect in reducing the racial and gender income gap, which leads to a significant reduction in the between-inequality, at least up to the medium.

Conclusion

In this paper, we introduce an estimation method for evaluating the effect of group-level policies under the quantile random-coefficient regression framework with interactive fixed effects. Our method can capture the heterogeneous policy effects through the interaction of policy variables and the individual observed and unobserved characteristics, while controlling the unobserved interactive fixed effects, and provides a straightforward way of identifying the policy effect on inequality measures. The consistency and limiting distribution of the proposed estimators are established. Using our proposed model, we evaluate the effect of the minimum wage policy on earnings between 1967 and 1980 in the United States. Our analysis confirms the findings of derenoncourt2021minimum that the policy helps reduce the racial income gap by improving the black economic status. On top of that, we provide empirical evidence of a compounded policy effect in narrowing the racial and gender gap, which contributes to the significant reduction in the between-inequality.

Acknowledgments

This paper benefited greatly from our discussions with Dukpa Kim. We are also grateful for comments from Martin Huber, Rustam Ibragimov, Artem Prokhorov, and participants at the Center for Econometrics and Business Analytics (CEBA) talk and seminars. We would also like to thank Ellora Derenoncourt and Claire Montialoux for sharing the data and code for the empirical application. Gao and Oka gratefully acknowledge financial support from the Australian Research Council Discovery Programs Scheme under Grant Numbers: DP200102769 and DP190101152, respectively, and Whang thanks financial support from the Korea Bureau of Economic Research and Innovation at the Institute of Economic Research in Seoul National University under Grant 0405-20220046. All errors are our own.

Declaration of Interest Statement

The authors would like to declare that there is no conflict of interest.