Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,571 characters · 16 sections · 95 citation commands
Estimating Heterogeneous Effects: Applications to Labor Economics
\vskip 1cm
\global\long \global\long \global\long \global\long \global\long
Documenting heterogeneity across individuals, firms, or space, has become a central theme in applied economics. In earlier research, heterogeneous parameters used to be treated as a nuisance, and “differenced out” by means of fixed-effects regressions. Increasingly, however, estimating and studying heterogeneous effects has become the main goal of the analysis.
This focus on heterogeneity is enabled by the availability of richer data sets. A leading example is given by administrative data sets that feature many units (such as firms, workers, or neighborhoods), and at the same time provide information about each unit (such as multiple workers in a firm, multiple time periods on a worker, or multiple individuals in a neighborhood).
Settings with heterogeneous parameters are often studied using panel data techniques. However, traditional panel data methods treat units, such as firms or neighborhoods, independently of each other. A growing number of applications instead involve settings where, in order to infer heterogeneous effects, the researcher specifies a research design that compares various units. For example, to estimate firm or neighborhood effects, researchers exploit workers who move between firms or children who move to a new neighborhood, respectively.
We will refer to three leading examples as illustrations. In the first one, kline2022systemic send multiple job applications to several large firms, and estimate firm-specific call-back rates as a function of applicants' characteristics. Documenting differences across firms in call-back rates allows the researchers to study hiring discrimination at the firm level, hence complementing the literature using r\'esum\'e correspondence experiments (e.g., bertrand2017field) by providing firm-level estimates. In this application, discrimination parameters can be estimated independently, firm by firm. Hence, this setting is akin to a traditional panel data or grouped data setting.
In the second example that we analyze, chetty2018impacts2 study how the income at adulthood of children depends on the place where they grew up. To estimate the effects of neighborhoods on income, the researchers exploit mobility of families across neighborhoods. Hence, the effect of neighborhood $j$ on income at adulthood is constructed by comparing the incomes of individuals who grew up in various neighborhoods $j'$. This setting has been studied by a large subsequent literature, including chetty2018impacts, laliberte2021long, bergman2019creating, and aloni2023one.
Our third leading example is the so-called AKM regression framework for matched employer-employee data introduced by abowd1999high. The researchers' goal is to estimate how worker and firm effects contribute to wage dispersion. This setting similarly features comparisons between multiple units, since the differences in wage premia offered by two firms $j$ and $j'$ is informed by workers moving between $j$ and $j'$. The AKM approach has become central to the study of workers and firms, see among others card2013workplace, card2018firms, and song2019firming. Moreover, the methodology in this example has been used to study other questions, such as differences in health care utilization across space inferred from patients moving between cities (finkelstein2016sources), or the impacts of banks and firms on credit growth inferred from banks' loans to multiple firms (amiti2018much).
In these applications, the data sets are complex and the models are high-dimensional. Practitioners need to make a large number of choices for modeling, practical specification, and estimation. Since we have not seen any survey on the econometric analysis of these settings as of yet,\footnote{abowd2008econometric and bonhomme2020econometric survey methods for bipartite networks and matched employer-employee data.} we have decided to write a paper on the topic. Our main goal is to lay out a framework to analyze these settings, and to clarify some of the underlying assumptions and key practical choices. While doing so, we will highlight the economic content behind the main assumptions.
The model that underlies most applications to these settings is a linear normal random coefficients (RC) model. The “random coefficients” refer to the heterogeneous effects that are the focus of the analysis, such as the effects of neighborhoods, or the worker and firm effects. Those coefficients are associated with specific covariates: in chetty2018impacts2 the effect of a neighborhood is simply the coefficient of the exposure to that neighborhood (i.e., of how long the family stayed in that neighborhood), whereas in abowd1999high the effect of firm $j$ is the coefficient of the $j$-th firm indicator. In both cases, the model involves a very large number of such covariates (e.g., many thousands of firm and worker indicators or neighborhood exposures).
The primitive parameters of the RC model are the means and variances of the coefficients (e.g., the neighborhood effects, or the worker and firm effects), as well as the variance of the errors. All of these parameters are potentially functions of all the covariates. They satisfy first and second moment conditions that we present. These moment conditions, which remain valid absent normality, build on moment conditions previously derived for panel data settings (e.g., chamberlain1992efficiency, arellano2012identifying).
While means and variances are useful to answer substantive questions, such as how much dispersion in outcomes is explained by neighborhood, worker, or firm heterogeneity, other important questions require additional information. As we describe, higher-order moments (such as skewness or kurtosis), nonlinear moments, as well as marginal and multivariate distributions, can all be inferred under the normality assumption of the RC model. For example, researchers may be interested in the distribution of neighborhood effects across space, or in the densities of worker and firm effects. In addition, the normal RC model can be used to construct optimal predictors of the effects. Such predictors or “forecasts” are also of considerable interest in various literatures outside our three leading examples, including in the work on teacher quality (e.g., kane2008estimating, chetty2014measuring).
Taking the RC model to data, however, is challenging in the settings that we study, given the large number of covariates (such as neighborhood exposures, or worker and firm effects) and the complex forms of dependence implied by the model. We show that commonly used specifications effectively impose conditional independence assumptions that may be economically restrictive. For example, common strategies in the analysis of neighborhood effects require location choice to be independent of neighborhood heterogeneity conditional on a set of neighborhood-specific covariates. Estimating flexible models of means, covariances, and distributions in these settings is an important yet still relatively unexplored research area.
Lastly, while the normal RC model provides a unified, self-contained framework to estimate heterogeneous parameters, normality and linearity may both be restrictive. In the last part of the paper, we briefly explore how these assumptions could be relaxed.
\paragraph{Relation to the literature.}
The framework we present builds on a vast methodological literature in statistics and econometrics. This includes the statistical literature on mixed models (e.g., jiang2007linear, mcculloch2004generalized), the literature on random coefficients models and correlated random-effects approaches in panel data (e.g., chamberlain1992efficiency, arellano2012identifying) and related panel data work based on decision-theoretic approaches (chamberlain2009decision, chamberlain2016fixed), as well as the empirical Bayes literature (e.g., efron2012large) and the Bayesian and frequentist interpretations of best linear unbiased predictors (BLUP, e.g., robinson1991blup).
It is common in empirical work to focus on the relationship between some (typically scalar) outcomes $y_i$ and some covariates $x_i$ and $z_i$, and to specify a linear regression model of the form
We assume that the researcher has access to multiple observations $i$ about some economic units $j$. The units may be firms, workers, or neighborhoods, depending on the application. Our focus is on settings where $z_i$ contains unit-specific variables, as well as interactions of unit-specific variables with other covariates. Let $p$ denote the number of covariates in $z_i$. Hence, the coefficients $\eta_j$, $j=1,...,p$, are unit-specific “fixed effects”.\footnote{We will refer to the $\eta_j$'s as “fixed effects”, in line with the usual terminology in applied economics. In contrast, in the statistical literature on mixed models $\beta$ are often referred to as “fixed effects”, and $\eta$ as “random effects”. See, e.g., the monograph by jiang2007linear.} Let $q$ denote the number of covariates in $x_i$. We will focus on settings where $p$ is large, and, depending on the application of interest, $q$ may be small or large.
Estimating high-dimensional regressions as in ((ref)) is made possible by the availability of increasingly large and detailed data sets, and by the development of powerful computational methods and software. This enables researchers to “zoom in” on the effects of particular units in the sample.
Under the assumption that $u_i$ are uncorrelated with $x_i$ and $z_i$ in ((ref)), researchers typically estimate $\beta$ and $\eta$ using OLS. This regression delivers parameter estimates $\widehat{\beta}$ and $\widehat{\eta}$. The researcher's goal is then to use the estimates $\widehat{\eta}_1,...,\widehat{\eta}_p$ to learn about the true effects $\eta_1,...,\eta_p$.
We now describe three examples where this setup arises.
Increasingly, researchers estimate regressions where learning about $\eta_j$ requires comparing various units. In that case, the estimate $\widehat{\eta}_j$ for unit $j$ is constructed using the observations from several (and possibly a large number of) other units $j'$. Our next two examples are in this vein.
In the three applications that we have outlined, the researcher's goal is to use the estimates $\widehat{\eta}_1,...,\widehat{\eta}_p$ to learn about the true effects $\eta_1,...,\eta_p$. We now provide examples of quantities of interest taken from the literature.
Throughout the paper we will treat $\eta_1,...,\eta_p$ as random. Our goal will be to estimate features of the joint distribution of $\eta_1,...,\eta_p$, and construct predictors of those effects, without imposing a priori that the $\eta_j$'s are independent of each other or independent of the $x_i$'s and $z_i$'s. When the conditional distribution of $\eta_1,...,\eta_p$ given $x_1,...,x_n,z_1,...,z_n$ is left unrestricted, proceeding in this way does not materially differ from a setup where the $\eta_j$'s are treated as fixed parameters (i.e., as “fixed effects”). The model we will present in Section (ref) will impose restrictions on this conditional distribution, however.
Researchers may have various objectives. A first goal may be to estimate some moments of $\eta_j$. For simplicity we will focus on the case where $\eta_j$ is scalar, although the expressions below are easily adapted to the case of vector-valued $\eta_j$'s.
Moments that are expectations of linear combinations of the $\eta_j$'s can be written as
where $c$ is a $p\times 1$ vector. An example is the mean of the $\eta_j$'s, $$\mathbb{E}\left[\frac{1}{p}\sum_{j=1}^p\eta_j\right].$$
Consider the coefficient in the linear regression of $\eta_j$ on some covariates $W_j$. For example, one may be interested in regressing firm effects on firm size or industry indicators in Example (ref), or in regressing neighborhood effects on average income in the neighborhood in Example (ref). The regression coefficient can be written as $\left(\mathbb{E}\left[W_jW_j'\right]\right)^{-1}\mathbb{E}\left[W_j\eta_j\right]$, which takes the form ((ref)) for $c=\left(\mathbb{E}\left[W_jW_j'\right]\right)^{-1}W_j$.
Moments that are quadratic in $\eta$ can be written as
where $Q$ is a $p\times p$ matrix. For example, the variance of the $\eta_j$'s, $$\limfunc{Var}(\eta_j)=\mathbb{E}\left[\frac{1}{p}\sum_{j=1}^p\left(\eta_j-\frac{1}{p}\sum_{j'=1}^p\eta_{j'}\right)^2\right],$$ can be written in this form for a suitable matrix $Q$. In Example (ref), interest often centers on the variances of worker and firm effects and on the covariance between worker and firm effects, all of which can be written as ((ref)).
Alternatively, one may be interested in the coefficient in the linear regression of some scalar variable $W_j$ on $\eta_j$. For example, in Example (ref) one may be interested in regressing promotion opportunities in a firm on the firm effects. The regression coefficient is $\left(\mathbb{E}\left[\eta_j^2\right]\right)^{-1}\mathbb{E}\left[\eta_jW_j\right]$, which can be written as the ratio between some $m_c$ in ((ref)) and some $v_Q$ in ((ref)), for suitable $c$ vector and $Q$ matrix.
More general, nonlinear moments can be written as
for some function $H:\mathbb{R}^p\rightarrow \mathbb{R}$. For example, one may be interested in the skewness or kurtosis of the $\eta_j$'s, which can be written as ratios of quantities of the form ((ref)) for suitable $H$ functions.
Learning about distributions of effects, beyond their means and variances, is important in all the examples we have mentioned. In Example (ref), kline2022systemic report estimates of the distribution of racial discrimination in hiring across firms. In Example (ref), researchers are often interested in documenting the distribution of neighborhood effects. In Example (ref), an increase in the variance of firm effects, say, has different implications for inequality whether it comes from a deepening of the left tail, an expansion of the right tail, or a symmetric increase in spread.
The weighted cumulative distribution function, for some weights $\omega_j$ that sum up to one, can be written as the following nonlinear moment
One may also be interested in the (weighted) density of the $\eta_j$'s, $f_{\omega}(a)=\frac{\partial F_{\omega}(a)}{\partial a}$. Bivariate counterparts to $F_{\omega}(a)$ and $f_{\omega}(a)$, which reflect the bivariate distribution of workers and firms, are of interest in Example (ref) in order to document sorting patterns along the worker and firm distributions.
Another common goal in applications is to construct predictors of the $\eta_j$'s. The optimal predictor that minimizes the expected sum of squared errors is a set of functions $\varphi_1,...,\varphi_p$ that solves
the solution of which is the set of conditional means
In Example (ref), chetty2018impacts and chetty2018impacts2 construct predictors of neighborhood-specific income effects. bergman2019creating use effect predictors to select the top census tracts for income mobility.\footnote{See gu2023invidious for a broad account of ranking problems and selection of groups of units based on noisy estimates.} A key input to answering these questions is the set of conditional means given by ((ref)).
The data is not directly informative about the $\eta_j$'s, but delivers estimates
where $v_j=\widehat{\eta}_j-\eta_j$ reflects estimation noise. In many applications involving large data sets (large $n$) and many unit-specific parameters (large $p$), the noise $v_j$ is substantial enough to be of practical concern. Inferring $\eta_j$ from $\widehat{\eta}_j$ then requires solving a filtering (or “de-noising”) problem.
It is important to note that directly using the estimates $\widehat{\eta}_j$ in place of $\eta_j$ may be misleading. The noise $v_j$ in ((ref)) reflects the presence of a form of measurement error. When the quantity of interest is nonlinear in $\eta_j$, such as a variance, a higher-order moment, a cumulative distribution function, or a density, the presence of measurement error often leads to unreliable, biased estimates of the quantities of interest.
As an example, suppose the researcher is interested in the variance of the $\eta_j$'s, and that she reports the following “plug in” estimate based on the $\widehat{\eta}_j$'s,
Does a large estimate $\widehat{\limfunc{Var}}(\widehat{\eta}_j)$ indicate that the variance of the true effects $\eta_j$, $\limfunc{Var}(\eta_j)$, is large? Or does the presence of measurement error $v_j$ artificially inflate the dispersion in the estimates $\widehat{\eta}_j$?
As another example, suppose the researcher uses $\widehat{\eta}_j$ as a predictor of $\eta_j$. Since the $p$ parameters ${\eta}_1,...,{\eta}_p$ are estimated in the available sample and $p$ is large (or, alternatively, the noise in ((ref)) is substantial), it is likely that $\widehat{\eta}_j$ “overfits”, in the sense that it reflects too much of the noise $v_j$ and too little of the true effect $\eta_j$. In such cases, one may wish to construct predictors that lead to a lower expected sum of squared errors in ((ref)).
To illustrate these points, consider a simple setting inspired by Example (ref), where $\eta_j$ and $v_j$ in ((ref)) are independent, i.i.d. across $j$, and ${\cal{N}}(\mu_{\eta},\sigma_{\eta}^2)$ and ${\cal{N}}(0,\sigma_{v}^2)$, respectively. In this case, we have
which shows that the variance of the estimates $\widehat{\eta}_j$ is upward-biased for the variance of the true $\eta_j$'s. Moreover, while the expected sum of squared errors of $ \widehat{\eta}_j$ is equal to $p\sigma_v^2$, the expected sum of squared errors of the conditional means
is smaller, equal to $\frac{p\sigma_v^2\sigma_{\eta}^2}{\sigma_{\eta}^2+\sigma_v^2}$. This shows that the “shrunk” quantities $\mathbb{E}(\eta_j\,|\, \widehat{\eta}_j)$ in ((ref)) have a lower expected sum of squared errors than the original estimates $\widehat{\eta}_j$. The difference between the two is greater when the noise variance $\sigma_v^2$ is large relative to the variance $\sigma_{\eta}^2$ of the true effects. Bias-correction methods for variance components based on equations in the spirit of ((ref)), and linear shrinkage predictors akin to ((ref)), are now widespread in applied economics.
This simple example is too stylized to accurately describe the situations in Examples (ref) and (ref), however. In such settings, estimates $\widehat{\eta}_j$ are constructed using observations from other units $j'\neq j$. Hence, the $\widehat{\eta}_j$'s are not independent. This gives rise to more complex forms for the estimation noise $v_j$ in ((ref)), and complicates the way the noise affects the quantities of interest. jochmans2019fixed study how, in settings where the matrix of covariates $z_i$ has a network structure (such as Example (ref), where $z_i$ represent worker and firm employment relationships), the properties of the network, such as how connected it is, affect the precision of the estimates $\widehat{\eta}_j$. Moreover, in the settings of Examples (ref) and (ref), unit-specific parameters $\eta_j$ may not be independent. We next present a framework that applies to an arbitrary matrix of covariates $z_i$ and allows for dependence between units.
In this section we describe a normal Random Coefficients (RC) model, which allows researchers to answer the questions introduced in the previous section.
For the presentation we will remove the term $x_i'\beta$ from equation ((ref)). Depending on the setting, $\beta$ can be estimated using OLS or differenced out, see Remark (ref) below. We then write ((ref)) in vector form, removing $x_i'\beta$ from the equation, as follows,
where $Y$ is an $n\times 1 $ vector with generic element $y_i$, $Z$ is an $n\times p$ matrix with generic row $z_i'$, and $U$ is an $n\times 1 $ vector with generic element $u_i$.
The form of the design matrix $Z$ differs across applications. In Example (ref), $Z$ takes the form $$Z=\left(
\right),$$ where $Z_j$ contains two columns, the first one being a column of $1$'s, and the second one being a column of $0$'s and $1$'s, depending on the race of the applicant. This block-diagonal structure characterizes panel data and grouped data settings.
In Examples (ref) and (ref), $Z$ takes more complex forms. To present Example (ref), let us abstract from the dependence on parental income for simplicity. Then, $z_{ij}$ in ((ref)) is the exposure of $i$ to neighborhood $j$. In every row of the $n\times p $ matrix with elements $z_{ij}$, all but a handful of elements are equal to zero.\footnote{For example, chetty2018impacts2 focus on families that move exactly once, so every row in the matrix has exactly two non-zero elements.} However, the matrix does not have a block-diagonal form. Then, differencing out the origin-and-destination indicators $x_i$ (as explained in Remark (ref) below) leads to the matrix
where $\widetilde{z}_{ij}$ is equal to $z_{ij}$ minus the mean of $z_{i'j}$ for all individuals $i'$ who experience the same neighborhood moves in childhood as individual $i$ (though the time $i'$ and $i$ stay in each neighborhood may differ). $Z$ is sparse, and not block-diagonal.
In Example (ref), $Z$ stacks worker and firm indicators together. For example, with two periods, $K$ workers, and $J$ firms, $Z$ reads
where $f_{kt}^j=1$ if worker $k$ is employed in firm $j$ in period $t$, and $f_{kt}^j=0$ otherwise. Note that $Z$ is not block-diagonal in this case. However, it is typically a sparse matrix since each row has exactly two non-zero elements. The form of $Z$ reflects the network of workers' and firms' employment relationships (andrews2008high, jochmans2019fixed).
Throughout, we assume that $Z$ has full column rank, so $(Z'Z)$ is non-singular. When $p$ is large, this assumption may be restrictive. In Example (ref), ensuring non-singularity requires imposing a normalization on the parameters $\eta_j$ (e.g., that the firm-specific effects sum up to zero), and focusing on a connected component of the firm-worker network (abowd2002computing). Similarly, in Example (ref), a normalization is needed since neighborhood effects are identified relative to the national average (chetty2018impacts2).
The defining assumptions for the normal random coefficients model are as follows.
In part (i) of Assumption (ref) we assume that the error terms $u_i$ in ((ref)) are normally distributed, with zero mean and some $n\times n$ covariance matrix $\Omega(Z)$, independent of $\eta$. In part (ii) we specify a normal model for $\eta$ given $Z$, with mean $\mu(Z)$ (a $p\times 1$ vector) and variance $\Sigma(Z)$ (a $p\times p$ matrix). Hence, we treat the parameters $\eta_j$ as random coefficients, and we specify their conditional distribution given $Z$. This modeling device is often used in panel data and complex data settings. Note that Assumption (ref) implies that the conditional distribution of $Y$ given $Z$ is fully specified, and normal, given the parameters $\Omega(Z)$, $\mu(Z)$, and $\Sigma(Z)$, $$Y\,|\, Z\sim {\cal{N}}\left(Z\mu(Z),Z\Sigma(Z)Z'+\Omega(Z)\right).$$
Assumption (ref) requires the following strict exogeneity assumption: $\mathbb{E}[U\,|\, Z,\eta]=0$. This assumption imposes substantive restrictions on the economic environment. In Example (ref), it requires that the times families spend in every neighborhoods are unrelated to the unobserved determinants of adult outcomes. In Example (ref), it imposes an assumption of so-called “exogenous mobility”, through which workers' decisions to change jobs may be driven by worker and firm effects $\eta_j$ but not by idiosyncratic time-varying shocks $u_i$. Although some authors have attempted to relax strict exogeneity in the setting of Example (ref) (e.g., abowd2019modeling, bonhomme2019distributional), most research to date relying on model ((ref)) makes this assumption.
The concrete specification of $\mu(Z)$, $\Sigma(Z)$, and $\Omega(Z)$ will depend on the application. In panel or grouped data settings, such as in Example (ref), a common assumption is that $(\eta_j,Z_j)$ are independent across $j$, where $Z_j$ denotes the subset of observations $z_i$ that pertain to unit $j$. In that case, $\mu_j(Z)$ depends on $Z$ only through $Z_j$, $\Sigma(Z)$ is diagonal, and $\Sigma_{j,j}(Z)$ only depends on $Z_j$ as well.
However, in Examples (ref) and (ref), it may be more plausible to allow for a rich dependence of $\mu_j(Z)$ and $\Sigma_{j,j'}(Z)$ on the elements of $Z$, and to allow $\Sigma(Z)$ to be a general, non-diagonal symmetric matrix. Indeed, in Example (ref), the de-meaned exposures $\widetilde{z}_{ij'}$ to other neighborhoods $j'\neq j$ in ((ref)) are unlikely to be independent of neighborhood effects $\eta_j$, unless mobility across neighborhoods is unrelated to neighborhood heterogeneity. Likewise, in Example (ref), indicators of firms $j'\neq j$ in ((ref)) will generally correlate with the effect of firm $j$ unless workers' sorting patterns are independent of firm heterogeneity. We will return to specification issues in the next section.
Given Assumption (ref), the OLS estimator of $\eta$ in ((ref)) satisfies
where, denoting the variance-covariance matrix of the OLS estimates $\widehat{\eta}$ as
we have $$V=(Z'Z)^{-1}Z'U\,|\, Z,\eta\sim {\cal{N}}(0,S(Z)).$$
Model ((ref)) under Assumption (ref) is a normal linear mixed model (see, e.g., jiang2007linear and mcculloch2004generalized). As we will review below, normality is not needed to obtain informative moment conditions on $\mu(Z)$, $\Sigma(Z)$, and $\Omega(Z)$, and the following restrictions on first and second moments suffice.
Note that, while Assumption (ref) relaxes normality, it does maintain the strict exogeneity condition $\mathbb{E}[U\,|\, Z,\eta]=0$. If the researcher is only interested in means, variances or covariances of the $\eta_j$'s, or alternatively in coefficients of regressions where $\eta_j$ appears on the left- or right--hand side, then Assumption (ref) can be replaced by the weaker Assumption (ref) that only restricts first and second moments. However, normality is needed to answer questions related to the higher-order and nonlinear moments of the $\eta_j$'s, their distributions, and to construct optimal predictors of the $\eta_j$'s.
Suppose that the normal RC model ((ref)) holds and Assumption (ref) is satisfied, and suppose the covariance and mean functions ${\Omega}(Z)$, ${\mu}(Z)$, and ${\Sigma}(Z)$ are known. Then the model implies closed-form expressions for all the quantities that we mentioned in Section (ref).
For example, the first and second moments $m_c$ in ((ref)) and $v_Q$ in ((ref)) are given, respectively, by
and
The expressions ((ref)) and ((ref)) do not rely on normality, and hold under the weaker Assumption (ref).
Under Assumption (ref), we further can write the nonlinear moment $w_H$ in ((ref)), and the cumulative distribution function $F_{\omega}(a)$ in ((ref)), in closed form as, respectively,
and
where $\mu_j(Z)$ is the $j$-th element of $\mu(Z)$, $\Sigma_{j,j}(Z)$ is the $j$-th diagonal element of $\Sigma(Z)$, and $\Phi$ denotes the standard normal cumulative distribution function. Note that, while the conditional distribution of $\eta\,|\, Z$ is normal under Assumption (ref), the unconditional distribution of $\eta$ is not normal.
Lastly, under Assumption (ref), one can also derive a closed-form expression for the conditional mean of the vector $\eta$ given the data $(Y,Z)$, as
where $S(Z)$ is given by ((ref)), and $$ G(Z)=\left({S}(Z)^{-1}+{\Sigma}(Z)^{-1}\right)^{-1}$$ is the conditional variance of $\eta$ given $(Y,Z)$. The conditional density of $\eta\,|\, Y,Z$ is then
In this section, we discuss several possibilities to specify $\Omega(Z)$, $\mu(Z)$, and $\Sigma(Z)$.
The simplest specification for $\Omega(Z)$ is independent homoskedastic, that is, $$\Omega(Z)=\sigma^2I_n,$$ for a constant variance parameter $\sigma^2$. In Example (ref), andrews2008high rely on this assumption and construct an unbiased estimator of $\sigma^2$ for applications to matched employer-employee data.
However, both homoskedasticity and independence can be restrictive. To relax homoskedasticity, one can introduce covariates $W=(w_i)$ (for example, some functions of the elements of $Z$), and model $$\Omega(Z)=\limfunc{diag}\left(\sigma_{\theta}^2(w_1),...,\sigma_{\theta}^2(w_n)\right),$$ where $\sigma_{\theta}^2(w_i)$ is a parametric function of $w_i$ indexed by some parameter $\theta$.\footnote{While here we assume that $w_i$ is a function of the elements of $Z$, the covariates $w_i$ could also contain additional covariates not functions of $Z$, such as neighborhood characteristics or worker or firm characteristics, depending on the application.} In Example (ref), chetty2018impacts2 assume that $S(Z)$ in ((ref)) is a diagonal matrix, i.e., that the estimates $\widehat{\eta}_j$ are uncorrelated across neighborhoods. Under independence, kline2020leave model $\Omega(Z)$ as a diagonal matrix with unrestricted diagonal elements. They propose a leave-out method that provides unbiased estimates of these diagonal variance elements.
To relax independence, one can model $\Omega(Z)$ as a parametric, matrix-valued, non-diagonal function of the covariates $w_i$. In panel data settings, arellano2012identifying propose parametric ARMA specifications to allow for serial correlation. Relaxing independence has been shown to be important in Example (ref), where assuming serial independence within employment spells is often empirically restrictive.
Lastly, it is worth noting that one cannot leave the matrix $\Omega(Z)$ fully unrestricted while at the same time identifying moments of the effects $\eta_j$. This is because $\eta_j$ and $v_j$ in ((ref)) are both unobserved random variables. As a result, there is an essential trade-off between heterogeneity (the $\eta_j$'s) and the dependence of errors (the matrix $\Omega(Z)$).
Turning now to $\mu(Z)$ and $\Sigma(Z)$, a possible approach in applications is to specify $\mu(Z)$ as a function of some covariates $W$ (e.g., as a linear function of $W$), and to model $\Sigma(Z)$ similarly (e.g., as a constant diagonal matrix). However, the cost of such an approach is that it implicitly imposes a conditional independence assumption that may be economically restrictive. We now illustrate this important point through the help of examples.
Suppose that, in Example (ref), the mean and variance of neighborhood effects are specified as functions of a handful of covariates, for instance some variables measuring the racial and economic composition of the neighborhoods. In that case, the researcher is effectively assuming that the de-meaned neighborhood exposures in $Z$, see ((ref)), are independent of the location-specific effects $\eta_j$ conditional on those covariates. This assumption may be hard to reconcile with an economic model of location choice, and mobility across locations, where families' decisions are in part determined by neighborhood heterogeneity.
Similarly, in Example (ref), assuming that the means, variances, and covariances of worker and firm effects only depend on some worker and firm characteristics restricts job mobility to be independent of the worker and firm effects $\eta_j$ conditional on those characteristics. This assumption may be at odds with economic models of sorting where workers' and firms' decisions are in part determined by worker and firm heterogeneity.
To state the argument formally, let $W$ denote some covariates. Then a specification where $\mu(Z)$ and $\Sigma(Z)$ only depend on $W$ amounts to assuming the following: $$\eta\,|\, Z,W\sim {\cal{N}}(\mu(W),\Sigma(W)),$$ which in particular imposes that: $$Z\text{ and }\eta \text{ are independent given }W.$$ Likewise, assuming that $\mu_j(Z)=\mu_{j}(W_j)$, where $W_j$ denotes the covariates of unit $j$, imposes that: $$\eta_j\text{ are mean of independent } \left(W_{1},...,W_{j-1},W_{j+1},...,W_p\right)\text{ given }W_j.$$
To illustrate that such conditional independence assumptions may be economically restrictive, let us consider two models of firm choice for Example (ref). The first model assumes that workers maximize utility among all firms in a market, period-by-period (card2018firms, lamadon2022imperfect). Suppose that worker $k$'s indirect utility in firm $\ell$ at time $t$ is $$V_{k\ell t}=\rho W_{k \ell t}+\varepsilon_{k\ell t},$$ where $\varepsilon_{k\ell t}$ are i.i.d. type I extreme value preference shocks, independent of log wages $W_{k \ell t}$, and we abstract from non-wage amenities for simplicity. Then the probability that worker $k$ chooses firm $\ell$, given all wages $W=\{W_{k'\ell' t}\}$, is
where ${\cal{M}}(k)$ is the market that $k$ considers when looking for a job. Therefore, the elements in the $Z$ matrix depend on the log wages $W_{k \ell t}$. If, further, log wages are a function of worker heterogeneity $\alpha_k$ and firm heterogeneity $\psi_{\ell}$,\footnote{For example, if log wages are given by the additive specification $W_{k \ell t}=\alpha_k+\psi_{\ell}+U_{k\ell t}$ of abowd1999high, in which case $\Pr\left(\ell\,|\, k, W\right)$ in ((ref)) does not depend on $k$.} then ((ref)) implies that the $\eta_j$'s, which here are the $\alpha_k$'s and $\psi_{\ell}$'s, are not independent of $Z$, even conditional on observed characteristics.
Consider next a dynamic model of workers' mobility across firms, as proposed by lentz2023anatomy (see sorkin2018ranking for a related model). The probability that worker $k$ moves between firms $\ell$ and $\ell'$, conditional on worker heterogeneity $\alpha=\{\alpha_k\}$ and firm heterogeneity $\psi=\{\psi_{\ell}\}$, is specified as
where $\lambda_{k\ell'}$ is the probability that $k$ meets firm $\ell'$, and $\gamma_{k\ell}$ is interpreted as worker $k$'s value of working in firm $\ell$. If the value $\gamma_{k\ell}=\gamma(\alpha_k,\psi_{\ell})$ depends on worker heterogeneity $\alpha_k$ and firm heterogeneity $\psi_{\ell}$, then neither $\alpha_k$ nor $\psi_{\ell}$ are independent of $Z$, even conditional on observed characteristics.
We now discuss several examples of specifications for $\mu(Z) $ and $\Sigma(Z)$ used in practice. In Example (ref), a common approach is to model $\mu_j(Z)$ to be a linear function of some covariates $W_j$, and $\Sigma(Z)$ to be a diagonal matrix, independent of $Z$ and $W$. The model then assumes independence across $j$'s, and rules out dependence between the true effects ($\eta_j$) and location choice and mobility ($Z$) conditional on the covariates. Recently, chen2023empirical proposes an extension of this approach that allows for dependence between the true effects $\eta_j$ and the precision of their estimates $\widehat{\eta}_j$ (as measured by the diagonal elements of $S(Z)$ in ((ref))), while maintaining independence across units.
In Example (ref), woodcock2015match postulates a normal RC model where neither $\mu(Z)$ nor $\Sigma(Z)$ depend on $Z$, and $\Sigma(Z)$ is a diagonal matrix. However, these assumptions impose that workers' sorting patterns, which are encoded in $Z$, do not depend on the worker and firm effects $\eta_j$. bonhomme2023much refine the woodcock2015match model by allowing $\mu(Z)$ and $\Sigma(Z)$ to depend on $Z$. To model the dependence, they cluster firms into a small number of groups using the k-means algorithm based on their wage distributions (as in bonhomme2019distributional). Given this grouping, they allow the means and variances of worker and firm effects to depend on the groups, but not on the worker and firm identities within these groups. Similarly, they allow the covariances in $\Sigma(Z)$, which they do not assume to be diagonal, to depend on the groups. Generalizing such approaches to accommodate structural economic models of workers' mobility across firms is a promising area for future investigation.
In this section we describe various estimation strategies for the parameters of the normal RC model and the quantities of interest introduced in Section (ref).
We start by providing moment conditions on the primitive parameters of the normal RC model: the error variance $\Omega(Z)$, and the mean and variance of the unit-specific effects $\mu(Z)$ and $\Sigma(Z)$.
\paragraph{Moment conditions.}
Let Assumption (ref) hold. Then we have
where the cross-product term is zero since $\mathbb{E}[U\,|\, Z,\eta]=0$. This is a system of $n\times n$ moment conditions. Following arellano2012identifying, we can construct a suitable $n^2\times n^2$ matrix $M(Z)$ that “differences out” the first term on the right-hand side of ((ref)),\footnote{Let $A\otimes B$ denote the Kronecker product between $A$ and $B$, and let $I_{n^2}$ be the $n^2\times n^2$ identity matrix. A possible choice is $M(Z)=I_{n^2}-\left(Z(Z'Z)^{-1}Z'\right)\otimes \left(Z(Z'Z)^{-1}Z'\right)$. The key property is that $M(Z)(Z\otimes Z)=0$, which implies that $M(Z)\limfunc{vec}\left(Z\mathbb{E}\left[\eta\eta'\,|\, Z\right]Z'\right)=M(Z)(Z\otimes Z)\limfunc{vec}\left(\mathbb{E}\left[\eta\eta'\,|\, Z\right]\right)=0$.} to obtain
where the vec operator stacks together the $n$ columns of an $n\times n$ matrix into a $n^2\times 1$ vector. ((ref)) provides a system of conditional moment conditions on the elements of $\Omega(Z)$. Note that, depending on the specification of $\Omega(Z)$, the moment conditions ((ref)) do not necessarily guarantee that the elements of $\Omega(Z)$ are identified. Note also that $M(Z)$ in ((ref)) being singular implies that the system of equations has multiple solutions in the absence of restrictions on $\Omega(Z)$.
Next, taking the conditional mean in ((ref)), we obtain
which provide conditional moment conditions on $\mu(Z)$. These moment conditions depend neither on $\Omega(Z)$ nor on $\Sigma(Z)$.
Lastly, using ((ref)) we obtain
where $S(Z)$ given by ((ref)) is a function of $\Omega(Z)$. This gives the following moment conditions on $\Sigma(Z)$, $\Omega(Z)$, and $\mu(Z)$,
\paragraph{Parameter estimation.}
Given a parametric or semi-parametric specification for $\Omega(Z)$, $\mu(Z)$, and $\Sigma(Z)$, a possible estimation approach is to exploit the moment conditions ((ref))-((ref))-((ref)) using method-of-moments or minimum-distance estimation. This strategy is used in bonhomme2023much in Example (ref), for instance. Let $\theta$ be a parameter vector indexing $\mu_{\theta}(Z)$, $\Sigma_{\theta}(Z)$, and $\Omega_{\theta}(Z)$, possibly adding other conditioning covariates $W$. This estimation step delivers an estimate $\widehat{\theta}$, as well as estimates $\widehat{\mu}(Z)=\mu_{\widehat{\theta}}(Z)$, $\widehat{\Sigma}(Z)=\Sigma_{\widehat{\theta}}(Z)$, and $\widehat{\Omega}(Z)=\Omega_{\widehat{\theta}}(Z)$. An alternative is to perform (quasi-) maximum likelihood estimation, as in the following remark.
We now present three types of estimators for the quantities of interest introduced in Section (ref).
\paragraph{\#1 Bias-corrected fixed-effects estimators.}
Suppose first that the researcher is only interested in estimating linear combinations of the $\eta_j$'s or quadratic forms. In that case, she does not need to estimate $\mu(Z)$ and $\Sigma(Z)$. Indeed, linear combinations $m_c=\mathbb{E}\left[c'\eta\right]$ and, given knowledge of $\Omega(Z)$, quadratic forms $v_Q=\mathbb{E}[{\eta}'Q\eta]$, are nonparametrically identified under Assumption (ref) (without the need for normality), as
and
respectively. By ((ref)), linear combinations of the elements in $\widehat{\eta}$ are unbiased. Moreover, ((ref)) shows that, while quadratic forms in $\widehat{\eta}$ are biased, the bias is a known function of the variance-covariance matrix $S(Z)$ of $\widehat{\eta}$.
Then, a fixed-effects estimator of $m_c$ is
Moreover, given an estimator $\widehat{\Omega}(Z)$ and an associated estimator $\widehat{S}(Z)$ given by
a bias-corrected fixed-effects estimator of $v_Q$ is
Under Assumption (ref), the estimator $\widehat{v}^{\rm FE}_Q$ is unbiased whenever $\widehat{\Omega}(Z)$, and hence $\widehat{S}(Z)$, are themselves unbiased.
In Example (ref), andrews2008high assume that $\Omega(Z)=\sigma^2I_n$. In this case, ((ref)) is equivalent to $$\mathbb{E}\left[\left(I_n-Z(Z'Z)^{-1}Z'\right)YY'\left(I_n-Z(Z'Z)^{-1}Z'\right)\,|\, Z\right]=\sigma^2\left(I_n-Z(Z'Z)^{-1}Z'\right).$$ So, by taking the trace and expectation with respect to $Z$, it follows that
The formula ((ref)) is the well-known degree of freedom correction for variance estimation. In Example (ref), andrews2008high propose the estimators
and
which are unbiased for $\sigma^2$ and $v_Q$, respectively. kline2020leave generalize their approach to the case where $\Omega(Z)$ is a diagonal matrix with unrestricted diagonal elements.
\paragraph{\#2 Model-based estimators.}
Suppose now that the researcher wishes to estimate not only linear combinations and quadratic forms but also other quantities, such as distributions, nonlinear moments, or predictors. In that case, she first needs to produce estimates $\widehat{\Omega}(Z)$, $\widehat{\mu}(Z)$, and $\widehat{\Sigma}(Z)$. Given those, the researcher can produce estimates of all the quantities of interest we listed in Section (ref), as follows:
including estimates of the posterior mean and variance of $\eta$:
where $\widehat{S}(Z)$ is given by ((ref)).
\paragraph{\#3 Posterior estimators.}
To motivate the third type of estimators, consider a generic nonlinear moment of $\eta$, ${w}_H=\mathbb{E}\left[H(\eta)\right]$, and note that, by the law of iterated expectations,
where we have used the expression in ((ref)) of the conditional density of $\eta\,|\, Y,Z$.
Equation ((ref)) motivates the following posterior estimator of $w_H$:
where $\widehat{\mathbb{E}}[\eta\,|\, Y,Z]$ and $\widehat{G}(Z)$ are given by ((ref)) and ((ref)), respectively.
To provide intuition about the posterior estimator, we note that, by a change in variables,
where $\varepsilon\,|\, Y,Z\sim {\cal{N}}(0,I_p)$. Consider two special cases in ((ref)). When the $\widehat{\eta}_j$ are poorly estimated so $\widehat{S}(Z)^{-1}\approx 0$, then ((ref)) becomes
which coincides with the model-based estimator ((ref)). Now, when the $\widehat{\eta}_j$ are well estimated so $\widehat{S}(Z)\approx 0$ (and $S(Z)\approx 0$), then ((ref)) becomes
which means that the posterior estimator recovers the true quantity in this situation, irrespective of whether the normal RC model is well specified or not. This illustrates a robustness property of posterior estimators, as studied by arellano2009robust and bonhomme2022posterior, and provides a motivation for reporting such estimators in practice.
There are several challenges to deriving asymptotic properties for the estimators mentioned in this section, in the type of settings we are focusing on in this paper. A first challenge is the dimensionality of the data and model. Typically, the vector $\mu(Z)$ and the matrices $\Omega(Z)$ and $\Sigma(Z)$ are very large (i.e., both $n$ and $p$ are large). It is therefore necessary to work in a high-dimensional asymptotic regime where both $n$ and $p$ tend to infinity. Moreover, the specifications for $\Omega(Z)$, $\mu(Z)$, and $\Sigma(Z)$ often depend on many parameters. Another challenge is the nature of the matrix $Z$. In Examples (ref) and (ref), $Z$ does not have a block-diagonal structure, which further complicates the asymptotic analysis. Related to this, $\Sigma(Z)$ is often non-diagonal, hence creating complex forms of dependence among observations.
To give a simple example, suppose we are interested in the mean $m_c=\mathbb{E}[\frac{1}{p}\sum_{j=1}^p\eta_j]$. Consistency of $\widehat{m}_c=\frac{1}{p}\sum_{j=1}^p\widehat{\mu}_j(Z)$ for $m_{c}$ is not immediate. To see this, consider first the case where there is no estimation error and one can compute $\widetilde{m}_c=\frac{1}{p}\sum_{j=1}^p\mu_j(Z)$. Consistency of $\widetilde{m}_c$ for $m_{c}$ requires limiting the dependence of the $\mu_{j}(Z)$'s across $j$. For instance, in Example (ref), this requires limiting the dependence of mean exposure effects across neighborhoods, while in Example (ref), this requires limiting the dependence of mean firm effects across different firms, for example, allowing for dependence among “similar” firms only. Next, the addition of estimation error further complicates the argument, since the $\widehat{\mu}_j(Z)$'s are dependent across $j$ conditional on $Z$.
Nevertheless, several results are available in the literature. For the case of fixed-effects estimators of quadratic forms, as in ((ref)), kline2020leave provide conditions for consistency, as well as inference theory, in a setup that assumes $\Omega(Z)$ to be diagonal while leaving $\mu(Z)$ and $\Sigma(Z)$ unrestricted.
The mixed models literature in statistics (as reviewed by, e.g., jiang2017asymptotic) provides a variety of asymptotic results for the quantities of interest we have listed here, including for estimators of posterior means such as ((ref)). However, the models studied in the theoretical literature on mixed models impose specific assumptions on $\Omega(Z)$, $\mu(Z)$, and $\Sigma(Z)$. In particular, a typical assumption is that $\Sigma(Z)$ is a diagonal matrix, as in the so-called “mixed ANOVA” model. We have argued here that such an assumption may be economically restrictive. For example, in Example (ref), $\Sigma(Z)$ is not diagonal whenever the effects of two firms $j$ and $j'$ depend on each other given $Z$, as happens in the model of lentz2023anatomy mentioned in Section (ref).
Under Assumption (ref), which imposes normality, heijmans1986consistent provide conditions for consistency of the maximum likelihood estimator (as in Remark (ref)) in models with general forms of dependence across observations induced by the matrices $\Omega(Z)$ and $\Sigma(Z)$. heijmans1986asymptotic provide a further set of conditions under which the maximum likelihood estimator is asymptotically normal. Importantly, however, while their setup allows for general forms of dependence, it relies on the assumption that the dimension of the model's parameter $\theta$ is kept fixed as the sample size tends to infinity.
To make the model more flexible while keeping the number of parameters to estimate moderate, a possibility is to assume that, while means and variances vary across observations, this variation is driven by a small number of groups. For example, bonhomme2023much assume that the means and variances of worker and firm effects depend on firm groups. Grouping methods, where group membership is estimated using methods akin to the k-means algorithm, are theoretically justified under suitable conditions (bonhomme2015grouped, bonhomme2022discretizing). Assuming a grouped structure reduces the dimension of the model, and can allow for a simpler characterization of asymptotic distributions.
Given the relevance of these high-dimensional, dependent settings for the applied economic literature, more research needs to be done regarding the formal analysis of asymptotic properties and inference methods, both in cases where normality is assumed to hold (as in Assumption (ref)) and in cases where distributions may be non-normal. We next comment on the possibility that normality may be violated.
While the normal random coefficients model provides a simple and powerful tool to estimate heterogeneous effects in high-dimensional settings, it relies on functional forms assumptions that may be empirically restrictive. In this final section of the paper, we briefly explore some strategies that could be used to relax some of those assumptions.
\paragraph{Non-normal effects.}
When it is suspected that the distribution of $\eta\,|\, Z$ is not truly normal, that distribution is sometimes interpreted as a Bayesian prior. This interpretation is central to the empirical Bayes approach, and it is often invoked in applications such as Example (ref). In this perspective, the conditional distribution of $\eta\,|\, Y,Z$ in ((ref)) can be interpreted as the posterior distribution of $\eta$. Linear shrinkage predictors as in ((ref)) possess attractive robustness properties under misspecification (james1961estimation).\footnote{In particular, the James-Stein linear shrinkage estimator in the normal means model achieves asymptotically minimax expected sum-of-squares loss in Euclidean balls (see Theorem 7.48 in wasserman2006all).} Moreover, estimators of nonlinear moments based on posterior distributions (see Subsection (ref)) are less sensitive to misspecification of the normal model compared to estimators based on the prior distribution $\eta\,|\, Z$. However, in many applications the variance of the noise is substantial, and it is still important to correctly model the distribution of $\eta\,|\, Z$.
Suppose that we maintain the normality of $U\,|\, Z$ in Assumption (ref), while leaving the conditional distribution of $\eta\,|\, Z$ unrestricted. Then, ((ref)) becomes a nonparametric deconvolution model with normal errors. Identification and estimation strategies exist in a variety of settings. In Example (ref), kline2022systemic use Efron's deconvolution method (efron2016empirical) to estimate the density of firm-specific discrimination under an independence assumption. chen2023empirical proposes a location-scale model to handle situations where the precision of the estimates $\widehat{\eta}_j$ predicts the true $\eta_j$, assuming independence across units. He shows that this modeling can improve the prediction performance of conditional mean estimates. Relaxing independence across units, thereby fitting the settings of Examples (ref) and (ref), is an important task for future work on these approaches.
\paragraph{Non-normal noise.}
One may also suspect that the distribution of $U\,|\, Z$ is not normal. In applications to Examples (ref) and (ref), the error $V$ in ((ref)) may be approximately normal even when $U\,|\, Z$ is not, due to the fact that $V$ is an estimation error, hence asymptotically normal under standard conditions as the relevant sample size tends to infinity. Approximate normality of $V$ will typically hold in grouped data settings as Example (ref) provided that group sizes be large enough.\footnote{For example, $V$ may be approximately normal even when $Y$ is a vector of binary outcomes. In kline2022systemic, the outcomes $y_i$ are binary call-back indicators, yet estimates $\widehat{\eta}_j$ may still be approximately normal with variance $S(Z)$.} However, the normal approximation need not be accurate in other applications.\footnote{When $U\,|\, Z$ is not normal, the right-hand side in ((ref)) remains unbiased whenever $U$ has zero mean given $Z$ and $\eta$. However, it is no longer the best predictor of $\eta$ in general.}
Dealing with settings where $V$ in ((ref)) is not normal, and $\eta\,|\, Z$ is not normal either, is challenging. This requires using the individual outcome model ((ref)), while allowing for a non-normal distribution of $U\,|\ Z$. In panel data settings, arellano2012identifying show how to identify and estimate the distribution of $U$ under the assumption that it follows a linear independent factor model (see also kotlarski1967characterizing, li1998nonparametric, and bonhomme2010generalized). We are not aware of extensions of such generalized deconvolution approaches to the settings of Examples (ref) and (ref), however.
\paragraph{Nonlinear mean.}
Lastly, an important assumption in the RC model is that the mean outcome is linear in $\eta$, in the sense that
Equation ((ref)) has empirical and economic content. In Example (ref), assuming that neighborhood exposures experienced by children have a separable, constant impact on outcomes may be restrictive (chetty2018impacts2). In Example (ref), assuming away interactions between worker and firm effects in wages implicitly imposes restrictive assumptions on input complementarity, which in turn drives the nature of sorting patterns in many economic models (e.g., becker1973theory, eeckhout2011identifying).
Within the context of Example (ref), bonhomme2019distributional show how to allow for complementarity patterns between firm and worker effects, in a setup where they assume that firm heterogeneity is discrete.\footnote{In that application, the presence of complementarity between worker and firm effects can alternatively be interpreted as reflecting individual heterogeneity in the “treatment effect” of a firm, as studied in hull2018estimating.} More work is needed in this direction.