Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
83,893 characters · 13 sections · 67 citation commands
\thispagestyle{empty}
\hrule
\baselineskip=.88\baselineskip {
In many longitudinal settings, economic theory does not guide practitioners on the type of restrictions that must be imposed to solve the rotational indeterminacy of factor-augmented linear models. We study this problem and offer several novel results on identification using internally generated instruments. We propose a new class of estimators and establish large sample results using recent developments on clustered samples and high-dimensional models. We carry out simulation studies which show that the proposed approaches improve the performance of existing methods on the estimation of unknown factors. Lastly, we consider three empirical applications using administrative data of students clustered in different subjects in elementary school, high school and college.}
{\it JEL: C23; C26; C38; I21; I28} \\ {\it Keywords: Factor Model; Panel Data; Instrumental Variables; Administrative data}
\baselineskip=1.1\baselineskip \hrule \setcounter{page}{0} \setcounter{footnote}{0}
In recent years, there has been an increasing interest in applications of factor models in economics, finance, and psychology. In economics, the identification and estimation of factor models has received substantial attention in a number of areas from macro-finance to labor economics and development bernanke2005measuring,kim2014divorce,oAttanasio2020. Important work has studied the role of cognition, personality traits, and academic motivation on child development fCunha2007,fCunha2008,Borghans02102008,jHeckman13. Factor-augmented regressions as in jStock1999,jStock2002 are known to improve forecasts of macroeconomic time series such as inflation and industrial production. The literature also includes new models for high-dimensional data sets jBai2016, and methods for panels with large cross-sectional ($N$) and time-series ($T$) dimensions, following the influential work by hP06 and jB09. In panel data econometrics, one popular interpretation treats the latent factors as a generalization of traditional fixed effects models harding2014estimating,Chudik2015393,martin2015,moon2017,tAndo2016,aJuodis2018,tAndo2017,mHarding2020.
While the estimators proposed for panels with large dimensions have been widely popular, other methods developed for panels with small, or fixed, $T$ have not been frequently adopted by practitioners conducting empirical academic research. One reason, as mentioned in aJuodis2020 and illustrated in oAttanasio2020 and DelBono2020, is that identification of the factor model requires normalization restrictions that matter for the interpretation of results NBERw22441. In some cases identification is achieved through the use of dedicated measurements, where a priori knowledge is used to associate certain measurements uniquely with specific factors (for example a test can be associated uniquely with a given skill e.g. cunha2021econometrics). One common restriction, labeled “PC3” in BAI2013, normalizes coefficients in the first block of factors. However, a large set of normalizations are available to practitioners when observed measurements per subject do not have a predetermined or natural order. In this paper, we investigate this problem while primarily focusing our analysis on “fixed-$J$” panels, where $J$ denotes a number of clusters or groups (e.g., states, counties, schools, etc. as opposed to time series).
We begin our investigation by introducing a class of estimators that use internally generated instruments. Papers by HEATON2012348, and aJuodis2020, among others, also propose to estimate similar models using internally constructed instruments, an idea that can be traced back to the work of madansky1964instrumental. In contrast to the existing literature, our model is identified based on an alternative non-singular transformation which includes PC3 as a special case. This normalization is convenient for the interpretation of results when economic theory is silent on the type of restrictions that must be imposed to solve the rotational indeterminacy of the factor model. Moreover, we establish large sample results by accommodating the asymptotic theory for clusters developed by bHansen2019.
We then consider adopting multiple non-singular transformations to improve the efficiency of the estimator, and we derive two additional theoretical results on estimation. We propose an estimator considering PC3-type restrictions in fixed-$J$ panels and show that the estimator is consistent and asymptotically normal under standard conditions. However, in applications to high-dimensional data or panel data with a large number of clusters, there is an increasing number of available transformations. The number of instrumental variables can also increase with $J$, creating finite sample bias similar to the one generated by the use of too many instruments jHahn2003, Jerry2008,pBekker1994. Our second development is to address poor finite sample performance by proposing an alternative two-step estimator that accommodates econometric methods for high-dimensional models in a first step \citep*[e.g.,][]{Belloni2012,linton2016,fWind2019}. Although our main focus is on fixed-$J$ panels, we establish the asymptotic distribution theory for multiple transformations and demonstrate to practitioners how to select normalizations out of (possibly) an infinite number of them.
Despite the large body of work on instrumental variables and factor models, this paper develops a new class of estimators that are simple to implement and offer practitioners better performance in small samples. The estimation of slope parameters using instrumental variables is investigated in baing2010, Harding2011197, ahn2013panel, Robertson2015526, aJuodis2020, and NORKUTE2021416, among others. On the other hand, the latent factor structure is estimated in madansky1964instrumental, GH1982, sPudney81, jHeckman1987, and HEATON2012348. In our simulation study, we find that the proposed estimators improve on the performance of existing instrumental variable methods for the estimation of unknown factors.
Lastly, we consider three empirical applications of our method to the estimation of models of educational attainment using administrative data on students. First, we investigate how the distribution of students' abilities at a school district level changes over subsequent years of K12 education. We present evidence on the temporal and geographic variability of educational opportunity across the US using administrative data from over 11,000 school districts. In our second illustration of the approach, we estimate a factor model using administrative data from a higher education institution in Europe DeGiorgi2012. The third application employs data from jAngrist2002 to evaluate the impact of an educational voucher program implemented in Latin America. These examples show intriguing results and highlight the usefulness of our techniques in varied settings in order to identify the strong and weak performers across the unobserved dimensions of academic achievement.
This paper is organized as follows. Section 2 introduces the factor-augmented linear model and the proposed estimator. The section also presents the main theoretical result and discusses the implementation of the estimator. Section 3 investigates estimation under multiple normalizations. Section 4 provides Monte Carlo experiments to investigate the small sample performance of the proposed estimators. Section 5 demonstrates how the approaches can be used in practice by exploring applications using administrative data. Section 6 concludes. Mathematical proofs are offered in the Appendix.
This paper considers the following factor-augmented linear model for $i=1,\hdots,N$ subjects and $j=1,\hdots,J$ clusters:
where $y_{ij} \in \mathbb{R}$ is the $j$-th response variable for subject $i$, $\bm{x}_{ij} \in \mathbb{R}^p$ is a vector of independent variables, $\bm{\beta} \in \mathcal{B} \subseteq \mathbb{R}^p$ is an unknown parameter vector, $\bm{\lambda}_{i} = (\lambda_{i1},\lambda_{i2},\hdots,\lambda_{ir})' \in \mathbb{R}^r$ is a vector of factor loadings, $\bm{f}_{j} = (f_{j1},f_{j2},\hdots,f_{jr})' \in \mathbb{R}^r$ is a vector of latent factors, and $u_{ij}$ is an error term. The number of factors $r$ does not need to be known, as one can determine the number of factors following a number of approaches jB2002,aOnatski2010,gKapetanios2010,sAhn2013,lTrapani2018.
We are interested in the estimation of $\bm{\beta}$ and $\bm{f}_{j}$. For the results in this section, we will fix a subset $A_0$ of groups of interest, and estimate $\left(\bm{f}_{j},j\in A_0\right)$. Throughout this section, even as $J$ diverges, this subset remains fixed. Once estimators of the factors and of $\bm{\beta}$ are available, it is straightforward to construct an estimator for $\bm{\lambda}_{i}$ HEATON2012348,BAI2013. In Section (ref), as an illustration of the approach, we first concentrate our attention on estimation of the factor $\bm{f}_j$, and then we estimate the loading $\bm{\lambda}_i$ for $1 \leq i \leq N$.
Based on equation (ref), consider
where $A_J$ is a set that includes groups that are used to proxy the vector of loadings $\bm{\lambda}_i$, and $B_J$ is a set that includes groups that are used to generate instrumental variables. The number of elements in each set $S$ is denoted by $m_S$, and we require $m_{A_J} \geq r$, and $m_{B_J} \geq r,$ and $\left(A_0 \cup A_J \right) \cap B_J = \emptyset$. We will also require that $A_J \cap A_0 = \emptyset$, although this can be relaxed at the cost of additional notation and subtleties. For instance, we could require $m_{A_J \setminus A_0} \geq r$, and that there is at least one $j \in A_J \setminus A_0$ involved in each of the $r$ averages discussed below. It follows that,
where $\bm{y}_{i A_J}$ is an $m_{A_J} \times 1$ vector of response variables, $\bm{x}_{iA_J} = (\bm{x}_{ij})_{j \in A_J}$ is a $p \times m_{A_J}$ matrix of independent variables, $\bm{f}_{A_J} = (\bm{f}_{j})_{j \in A_J}$ is a $r \times m_{A_J}$ matrix of latent factors, and $\bm{u}_{iA_J}$ is a $m_{A_J} \times 1$ error term.
Let $\bm{D} = \bm{I}_r \otimes \bm{\iota}_{m_r}$, where $\bm{I}_r$ is the identity matrix of dimension $r$, $\bm{\iota}_{m_r}$ is a vector of ones of dimension $m_r = m_{A_J}/r$, and $\otimes$ denotes Kronecker product. We assume, for simplicity, that the number of groups per factor $m_r$ is an integer, because practitioners can always reorder elements after discarding those not in $A_{J} \cup B_{J}$. Let $\mathbb{M} = (\bm{D}' \bm{D})^{-1} \bm{D}'$ be a $r \times m_{A_J}$ matrix that creates $r$ averages of variables considering $m_r$ observations. Multiplying equation (ref) by $\mathbb{M}$, we obtain the following $r$ equations:
where, for instance, $\overline{\bm{y}}_{i A_J} = \mathbb{M} \bm{y}_{i A_J}$ denotes the vector of $r$ possible sample averages considering the elements of the vector $\bm{y}_{i A_J}$. Assuming that the $r \times r$ matrix $\overline{\bm{f}}_{A_J}'$ is invertible, we can solve for $\bm{\lambda}_i$:
Substituting equation (ref) into the augmented factor model (ref), one obtains, for each $j \in A_0$,
where $\bm{\theta}_{j} = \overline{\bm{f}}_{A_J}^{-1} \bm{f}_j$. We emphasize that the parameter depends on the normalization $A_J$ but we omit the dependence to keep the notation simple. By noting that $\bm{\theta}_{j}' \overline{\bm{x}}_{i A_J}' \bm{\beta} = \sum_{k=1}^r \overline{\bm{x}}_{i A_J, k}' \bm{\beta} \theta_{t,k}$, we can write,
where $\overline{\bm{X}}_{i A_J}$ is a vector of $p \times r$ independent variables, the vector $\bm{\gamma}_j = ( \bm{\beta}' \theta_{j,1},..., \bm{\beta}' \theta_{j,r})'$, and $v_{ij} = u_{ij} - \bm{\theta}_{j}' \overline{\bm{u}}_{iA_J}$. Although $\bm{\theta}_j$ and $\bm{\beta}$ could be estimated by standard methods for linear models, the variable in the first term of equation (ref), $\overline{\bm{y}}_{iA_J}$, is endogenous because it is correlated with $\overline{\bm{u}}_{iA_J}$, which appears as part of the error term.
We propose to estimate equation (ref) using internal instruments $\overline{\bm{y}}_{iB_J}$, as well as an expanded set of instruments, as described below. The assumptions imposed below imply that $\overline{\bm{y}}_{iB_J}$ is a strong and valid instrument (see the discussion after the main result). An additional challenge is related to inference because we are interested in estimating simultaneously $\bm{f}_j$ for all $j \in A_{0}$. We proceed by stacking the reduced form equation (ref). To handle the dependence across equations $j \in A_0$ for a given $i$ within the system, we use the asymptotic theory for clusters in bHansen2019.
Recall that the number of equations $m_{A_0}$ is fixed, in the sense that it does not diverge if $J$ does. The system of $m_{A_0}$ equations can be written as:
where $\bm{X}_{iA_{0}} = (\bm{x}_{i1}', \bm{x}_{i2}', ..., \bm{x}_{im_{A_0}}')'$ is a $m_{A_0} \times p$ matrix of exogenous variables, and $\bm{v}_{iA_{0}}$ is a $m_{A_0}$ dimensional vector with typical element $u_{ij} - \bm{\theta}_j' \overline{\bm{u}}_{iA_J}$. The parameter $\bm{\delta} = (\bm{\theta}_{A_{0}}', \bm{\beta}', \bm{\gamma}_{A_{0}}')'$, where $\bm{\theta}_{A_{0}} = (\bm{\theta}_{1}', \bm{\theta}_{2}', \hdots, \bm{\theta}_{m_{A_0}}')'$ and $\bm{\gamma}_{A_{0}} = (\bm{\gamma}_{1}', \bm{\gamma}_{2}', \hdots, \bm{\gamma}_{m_{A_0}}')'$. The total number of parameters in the system of equations (ref) is $k_{A_0} := m_{A_0} r (1+p)+p$.
The Grouped Variable Estimator (GVE) can be obtained as:
where $\bm{Z}_{iA}$ denote a matrix of internally generated instruments. For instance, stacking the instrumental variables analogously, we obtain the instrumental variables
where $\overline{\bm{y}}_{iB_J}$ is a $r$-dimensional vector of individual specific averages. The assumptions we maintain below actually imply a richer set of instruments, namely
The first set of instruments, $\bm{Z}^{(1)}_{i,A}$, leads to a just-identified IV estimator, regardless of (a potentially divergent) $m_{A_J}$. The second set of instruments $\bm{Z}^{(2)}_{i,A}$ is larger, and the number of elements will diverge if $m_{A_J}+m_{B_J}$ diverges.
In a factor model, $\bm{\lambda}_i$ and $\bm{f}_j$ are identified up to a non-singular transformation. To see this, note that the second term in equation (ref) corresponding to the partition $\bm{y}_{iA_0}$ can be written as $\bm{f}_{A_0}' \bm{A} \bm{A}^{-1} \bm{\lambda}_{i}$ for any non-singular $\bm{A}$ matrix of dimension $r \times r$. BAI2013 and bWilliams2020 discuss restrictions imposed to achieve point identification of factors and loadings. One set of restrictions on the $r^2$ free parameters is to normalize the upper $r \times r$ block of a matrix of loadings or factors. Thus, in the case that $m_{A_J}=r$, it is standard to consider $\bm{A} = \bm{f}_{A_J}^{-1}$, which has been used for identification using instrumental variables in HEATON2012348, jHeckman1987, and sPudney81, among others. In these models, the first $r$ factors are normalized to one.
Note that $\bm{\theta}_{j} = \overline{\bm{f}}_{A_J}^{-1} \bm{f}_j$ uses a different non-singular transformation than the one typically considered in the context of instrumental variables. Our transformation for the linear factor model leads to a normalization based on average of factors, which is convenient in terms of interpretation. After the factor model is estimated by (ref), we can employ transformations to uncover a simpler parameter structure. As an illustrative example, we can consider Example (ref). In this case,
showing that a simple reparametrization identifies the relative importance of the first factor. Naturally, there are other non-singular transformations that can be considered including $\bm{A} = \bm{f}_{A_J}^{-1}$, as discussed in the next section.
To think about identification of factors using instrumental variables, it is instructive to consider a special case when $m_r = 1$.
In this section, we establish conditions under which the estimator in (ref) is consistent and asymptotically normal. We will leverage the fact that our estimator can be viewed as a two stage least squares estimator for clustered data, where the cross-section units $i$ are the clusters; the measurements $j$ are observations within a cluster; the dependent variables are $\bm{y}_{iA_0}$; endogenous regressors are $\bm{y}_{iA_J}$; and so on. This allows us to use the asymptotic theory for clustered samples in bHansen2019, in particular their results for two stage least squares estimation in Theorems 8 and 9.
To state our results, define
The proof of Theorem (ref) is presented in Appendix (ref), and consists of verifying the conditions for Theorems 8 and 9 in bHansen2019. Their requirement that the observations from each $i$ are asymptotically negligible (cf. their Assumption 1) for consistency is automatically satisfied, as our panel is balanced by assumption. Moreover, the condition stated in Assumption 2 in bHansen2019 requires that $m_{A_0} / N \to 0$, which is satisfied because $m_{A_0}$ does not grow with $J$. The asymptotic variance $\bm{V}_N$ can be consistently estimated in the usual way (see Theorem 9 in bHansen2019).
The result in Theorem (ref) is obtained considering several standard assumptions. Following Assumption (a), data are generated by model (ref). Assumption (b) guarantees that the instruments are valid by requiring that the error terms in two partitions are not correlated. Assumption (c) is a boundedness condition on the regressors and outcome variable that allows for distributional heterogeneity, and is sufficient for bHansen2019's central limit theorem. Assumption (d) controls the behavior of the $\bm{f}_j$, which is part of the estimand and Assumption (e) asks for sufficient correlation of the instruments with the regressors.
The result of Theorem (ref) holds under $J$ fixed or $J \to \infty$ because the number of parameters $k_{A_0}$ does not depend on $m_{A_J}$ or $m_{B_J}$,\footnote{Recall that we assume that $m_{A_0}$ does not diverge if $J$ does.} and the number of instruments $\bm{Z}^{(1)}_{iA}$ does not diverge with $J$. The estimator in Theorem (ref) uses a fixed number of averages and not all the available internally generated instruments. If $J$ is fixed, it is straightforward to see that the result also holds for $\bm{Z}_{iA} = \bm{Z}^{(2)}_{iA}$ as in equation (ref). The case of $\bm{Z}_{iA}^{(2)}$ and $J \to \infty$ is different, as the number of internal instruments increases with $J$, and therefore, the estimator (ref) faces similar challenges to the ones found in the estimation of high-dimensional models svanderGeer2010, Belloni2012,fWind2019. We investigate the large sample behavior of the estimator in Section (ref).
The GVE estimator we propose have a number of attractive features. First, they are trivial to implement: they are 2SLS estimators with instruments (ref) or (ref) in the linear system (ref). Second, the estimator with instruments $\bm{Z}_{iA}^{(1)}$ has the attractive property that it is consistent without any restrictions on the rate at which $J$ grows with $N$ while also being fixed-$J$ consistent. Third, the estimator combine information from all units in $A_J$ and $B_J$, which are chosen by the researcher and are allowed to diverge. Fourth, existing solutions for handling missing data in 2SLS settings can be used to handle unbalanced panels. Fifth, we can easily accommodate regressors $\bm{x}_{ij}$ correlated with the error term $u_{ij}$ by using external instruments. Below we explore further the performance of our estimation approach in large $J$ settings and provide alternatives with improved finite sample performance.
A drawback of our approach is that it may not incorporate the information in the model efficiently. We offer the following two refinements, leaving careful study of their asymptotic properties for future research.
First, note that our estimators are for the parameters $(\bm{\theta}_{jA_0}',\bm{\beta}',\bm{\gamma}_{jA_0}')$. This overparametrization was chosen for ease of implementation. One could use efficient minimum distance to gain efficiency. Alternatively, we can consider a sequential approach based on a consistent estimator of $\bm{\beta}$ (see Section (ref)).
Second, we could repeat the estimation result for different choices of $A_J$, as long as $A_0$, $A_J$, and $B_J$ satisfy the restrictions outlined in the text above. One would typically set $B_J = \{1,\cdots,J\} \setminus \left( A_0 \cup A_J \right)$. As long as the number of choices of $A_J$ does not diverge, the distributional results can be applied directly.
The issue of multiple normalizations deserves further treatment as there are many situations where economic theory is silent on the type of restrictions imposed to the model. In those situations, practitioners face a possibly large number of normalizations that could be used to eliminate the problem of rotational indeterminacy of the factor model. Considering the first partition as the normalization might be arbitrary, as noted in a series of recent papers \citep*{oAttanasio2020,DelBono2020}. We briefly illustrate this issue using the following example:
Examples (ref) and (ref) illustrate that it not clear a priori whether to normalize based on mathematics, reading, or writing, leading to important practical questions on how to select a normalization and the corresponding partition. In fact, there are $Q_J$ ways of choosing the subset $A_J$, where
The solution we pursue in this section is to simultaneously adopt multiple subsets.
Theorem (ref) establishes conditions under which the GVE that uses one normalization $A_J$ is consistent for the normalized factors and the regression coefficient. In this section, we will assume that $\bm{\beta}$ is known and focus on improved estimation of the factors by using information from multiple normalizations. Define $\bm{R}_i = ( \bm{R}_{iA_0}', \bm{R}_{iA_J}', \bm{R}_{iB_J}')'$ where, for instance, $$\bm{R}_{iA_0} = \bm{y}_{iA_0} - \bm{x}_{iA_0}'\bm{\beta} = \bm{f}_{A_0}' \bm{\lambda}_{i} + \bm{u}_{iA_0},$$ and $\bm{f}_{A_0}$ is a matrix of dimension $r \times m_{A_0}$. While the results of Theorem (ref) hold for general $m_{A_J} \geq r$, we will focus on the special case $m_{A_J} = r$ madansky1964instrumental,sPudney81,jHeckman1987,HEATON2012348,bWilliams2020.
Letting $q = 1, 2, \hdots, Q_J$, for $Q_J$ as in (ref), and $\bm{M}_{iA_0,(q)} = \bm{I}_{m_{A_0}} \otimes \bm{R}_{iA_J,(q)}'$, we can write
where $\bm{\theta}_{A_0,(q)} = [\bm{I}_{m_{A_0}} \otimes \bm{f}_{A_J,(q)}^{-1}] \operatorname{vec}(\bm{f}_{A_0})$. Relative to the parameter $\bm{\theta}_{A_0}$ estimated by the GVE defined in (ref), the parameter $\bm{\theta}_{A_0,(q)}$ in equation (ref) can be based on any non-singular transformation. If $Q_J = 1$ and $m_r=1$, then $\bm{\theta}_{A_0,(q)} = \bm{\theta}_{A_0}$. Moreover, we define
where $W_{q}$ is a matrix of weights. Below, we introduce conditions that cover weighting for models with $J$ fixed (Theorem (ref)) and $J$ increasing to infinity (Theorem (ref)). They allow the use of different weighting schemes, including equal weighting.
Consider the following three illustrative examples.
Then, we estimate (ref) by the weighted grouped variable estimator (WGVE):
where $\widehat{\bm{\theta}}_{A_0,(q)}$ is the GVE defined in (ref) based on partition $q$. The use of weights for the combination of estimators in a linear fashion is naturally not new hP06,linton2016,mHarding2020. If $r=1$ then $Q_J = J - m_{A_0}$, and one could set $W_q = Q_J^{-1} = (J - m_{A_0})^{-1}$, and define the estimator as $\widehat{\bm{\vartheta}}_{A_0} = (J - m_{A_0})^{-1} \sum_{q=1}^{J - m_{A_0}} \widehat{\bm{\theta}}_{A_0,(q)}$, which is similar in spirit to the common correlated effect estimator of \citet*{hP06}. Moreover, the estimator (ref) is similar to the ones investigated by \citet*{linton2016}. For instance, Example 1 in \citet*{linton2016} consider a similar instrumental variable estimator for a simultaneous equation model, and the optimal choice of weights makes a weighted instrumental variable estimator asymptotically equivalent to the classical 2SLS estimator.
We begin by considering a panel data model when $J$ is fixed, and therefore, the number of non-singular transformations, $Q_J$, is constant. Note that $A_J$ and $B_J$ are fixed too, so the number of instruments employed in the first stage and the number of normalizations do not increase. This is the case most relevant in the applications using administrative data presented in Section (ref) and in the recent econometric literature aJuodis2018,aJuodis2020,NORKUTE2021416.
The estimator is a trivial extension of the method discussed in the previous section. In the first step, we obtain $\widehat{\bm{\theta}}_{A_{0},(q)}$ for $q = 1, 2, \hdots, Q_J,$ using the estimator (ref). In the second step, we compute a consistent estimator of $\bm{\vartheta}_{A_0}$ using a linear combination of consistent estimators obtained in the first step, as shown in (ref). As expected, a linear combination of a finite number of consistent and asymptotically normal estimators is consistent and asymptotically normal, as shown in Theorem (ref) below.
Let $\bm{v}_{iA_0,(q)} = \bm{y}_{iA_0} - \bm{M}_{iA_0,(q)} \bm{\theta}_{A_{0},(q)}$, and consider the following definitions:
and, by letting $\bm{\Sigma}_{(q)} := \lim_{N \to \infty} \bm{\Sigma}_{N,(q)}$, define
The next result considers multiple normalizations and builds on Theorem (ref):
Assumptions (i) and (ii) are generalizations of Assumptions (d) and (e) in Theorem (ref). Condition (i) imposes restrictions to generate suitable non-singular transformations across all partitions, and condition (ii) guarantees a well-defined asymptotic distribution across feasible non-singular transformations. Lastly, condition (iii) allows the use of different weighting schemes to improve the performance of the GVE and it is similar to the ones employed in the literature such as \citet*{hP06} and \citet*{linton2016}. We do not consider random weights, but condition (iv) can be easily accommodated to incorporate random matrices as in \citet*{linton2016}.
Optimal weights can be found as minimizers of the asymptotic covariance matrix of the estimator. To see this, write $\mathcal{V}_{A_0}$ as
Let $\bm{\Sigma} = [\bm{\Sigma}_{(lq)}]$ and $\iota_{Q_J}$ be a $Q_J$-dimensional vector of ones. Thus,
It follows that the estimator $\widehat{\bm{\vartheta}}_{A_0}^\ast = \sum_{q=1}^{Q_J} W_{0q}^\ast \widehat{\bm{\theta}}_{A_0,(q)}$ has asymptotic covariance matrix,
In other words, the optimal weighting is proportional to the inverse of the asymptotic covariance matrix of the estimator. The estimation $\mathcal{V}_{A_0}^\ast$ is straightforward and follows the estimation of $\bm{\Sigma}$. See Section 6 in \citet*{linton2016} for specific details.
In the case of panel data models with large $J$, possibly larger than $N$, there are known issues with the approach above. As discussed before, there is a large number of possible non-singular transformations. Moreover, least squares estimation of the regression of endogenous variables on the instruments has poor finite sample properties, and therefore, the GVE estimator is expected to perform poorly in practice. The procedure could suffer from a finite sample bias problem similar to the one investigated in \citet*{Jerry2008} and \citet*{Chao2012}. Therefore, this section investigates the case of large $J$ considering developments in \citet*{Belloni2012} and \citet*{linton2016}, although the large $J$ situation is not common in the analysis of student administrative data (Section (ref)). The procedure in Belloni2012 requires to approximate a large dimensional model by a low-dimensional sub-model. If some instruments are invalid, the procedure can be easily adapted to include the median estimator proposed by \citet*{fWind2019}.
We propose to estimate $\bm{\vartheta}_j$ for all $j \in A_0$ in two main steps, as before. We begin by describing the first step involving the use of instrumental variables. Let $\bm{R}_j = (R_{1j},R_{2j},\hdots,R_{Nj})'$ be an $N$-dimensional vector of dependent variables, $\bm{R}_{A_J,(q)}$ be an $N \times r$ matrix of endogeneous variables, and $\bm{R}_{B_J,(q)}$ be an $N $ by $m_{B_J} = J-m_{A_0}-r$ matrix of internal instruments. Let $\widehat{L}_{il,(q)} := \bm{R}_{i,B_J,(q)}' \widehat{\bm{\pi}}_{l,(q)}$ for $l = 1, 2, \hdots, r$, where $\widehat{\bm{\pi}}_{l,(q)}$ is a Lasso estimator defined as a solution of the following problem:
where the parameter set $\bm{\Pi}_{(q)} \subseteq \mathbb{R}^{(J-m_{A_0}-r)}$ and $\| \bm{b} \|_1$ is the standard $\ell_1$-norm defined as $\| \bm{b} \|_1 = \sum_i | b_i |$ for a generic constant $b_i$. The penalty loadings $\Upsilon_l$ and $\lambda_l$ are selected as in \citet*{Belloni2012}. We then collect the predictions $\hat{L}_{il,(q)}$ for all $1 \leq i \leq N$ and $1 \leq l \leq r$ to obtain a matrix $\widehat{\bm{L}}_{(q)}$ of dimension $N \times r$. In a second stage of the IV approach, we find $\widehat{\bm{\theta}}_{j,(q)}$ as the solution of the following equation:
with $G_{(q)}(\bm{\theta}_{j}) = E(G_{N,(q)}(\bm{\theta}_{j}))$.
The last step includes the following estimator:
where the truncation parameter $Q_J^\ast < Q_J$ for all $J$, and $\widehat{\bm{\theta}}_{j,(q)}$ is a Lasso-type estimator obtained as the solution of (ref).
Before establishing large sample results when the number of normalizations tend to infinity as $N$ and $J$ tend to infinity, we emphasize two conditions that are standard in the literature. The linear model estimated in the first stage of the IV procedure uses Condition AS in \citet*{Belloni2012} to approximate a conditional expectation up to a small non-zero approximation error. Let $L_{il,(q)} := \bm{R}_{i,B_J,(q)}' \bm{\pi}_{0l,(q)} + a_{il,(q)}$, $\max_{1 \leq l \leq r} \| \bm{\pi}_{0l,(q)} \|_0 \leq s_{(q)} = o(N)$, and
for $l = 1, 2, \hdots r$. Condition AS in \citet*{Belloni2012}, reproduced in the previous equations in the context of a factor model, states that at most $s_{(q)}$ variables are needed to approximate well the conditional expectation $L_{il,(q)}$ with a small approximation error. This error is of the same order of magnitude than the estimator error, $\sqrt{s_{(q)}/N}$. Noting that this condition holds for a given $j,q$ pair, the assumption allows us to identify
Moreover, \citet*{bickel2009} and \citet*{Belloni2012} introduce conditions on the Gram matrix $\bm{M} = E ( \bm{R}_{iB_{J},(q)} \bm{R}_{iB_{J},(q)}' )$, which is not positive definite if $m_{B_J} > N$. They propose a notion of “restricted" positive definiteness for vectors in a restricted set. In our model, this set is defined as $\Delta_\mu = \{ \bm{\Psi} \in \mathbb{R}^{J - m_{A_0} - r} : \| \bm{\Psi}_{\mathfrak{J}_{l,(q)}^c} \|_1 \leq \mu \| \bm{\Psi}_{\mathfrak{J}_{l,(q)}} \|_1, \bm{\Psi} \neq 0 \}$. We can then define
We now establish the consistency and asymptotic normality of the estimator introduced in (ref).
Condition (a) extends the condition used in Theorem (ref) to allow the use of weights when there is available a growing number of non-singular transformations, and it is similar to Condition A1 in \citet*{linton2016}. The implication is that truncation parameter $Q_J^\ast$ needs to grow slowly to satisfy the conditions in Lemma 1 in \citet*{linton2016} and are satisfied if $Q_J^\ast$ grows at logarithm rates. The truncation parameter defined in (a) satisfy the condition. Under assumption (b), we can determine the rate of convergence of the Lasso-type estimator in the case of Gaussian models with homocedastic errors. Assumption (c) is needed for the estimation of conditional expectation functions under non Gaussian conditions and heteroskedastic errors, and it is similar to Condition RF as implied by Lemma 3.b in \citet*{Belloni2012}. Conditions (b) and (c), in addition to the conditions on sparsity, are crucial for the consistency of the IV estimator. Assumption (d) is a modified version of a standard condition for uniform convergence of estimators that minimize a criterion function vanderVaart98. The difference is that the condition is imposed on every normalization. Assumption (e) is similar to condition (A.4) in \citet*{linton2016}. These assumptions impose uniformity over $Q_J^\ast$ normalizations. See Lemma 1 in \citet*{linton2016}.
The result in Theorem (ref) is achieved by using a “standardized” binomial coefficient as a truncation parameter, which controls the rate of growth of $Q_J^\ast$ as $J \to \infty$. As long as $r$ is fixed, one can approximate $Q_J$ by $(J-m_{A_J})^r/r!$, and thus, the ratio of $Q_J^\ast \to C$ as $J \to \infty$. This result is important in practice as it indicates that the truncation parameter $Q_J^\ast$ should be determined to minimize computational time as well as to maximize efficiency gains. There are several options for practitioners. One is to set $C=1$ and then potentially investigate the marginal impact of an additional normalization in terms of the standard error of the estimator.
In this section, we conduct an investigation of the performance of the proposed approaches in comparison to existing methods. Using a series of simulation experiments, we report the root mean squared errors of new and existing estimators for different models. We first consider a factor model, and then a factor-augmented model.
We begin by considering a one-factor model similar to the one used in HEATON2012348. Observations are generated from $y_{ij} = \lambda_{1,i} f_{1,j} + u_{ij}$, where the error term $u_{ij} \sim F_u$, and $\lambda_{1,i}$ is drawn as an independent observation from a uniform distribution ranging from 0.5 to 3.5. We generate observations for the factor following the equation $f_{1,j} = 0.8 f_{1,j-1} + \eta_{j}$ for $j = -S + 1, -S + 2, \hdots, 0, 1, \hdots, J$, where $\eta_{j}$ is an i.i.d. random variable distributed as uniform with support ranging from 0 to 1. We set $S = 50$ to minimize the effects of the initial value, $f_{1,-49}=1$. We consider two variations of the model. We first assume that the error term $u_{ij}$ is an i.i.d. Gaussian random variable, and then we assume that $u_{ij}$ is a random variable distributed as $t$-student with 3 degrees of freedom ($t_3$).
Table (ref) presents the root mean squared error (RMSE) for the parameter $\bm{\theta}$. The table shows results from different estimators. We compare our estimators to more traditional approaches such as PCA BAI2013, IV an instrumental variables estimator that uses internal instrumental variables, and LAS an estimator that uses the LASSO instead of the IV estimator. The implementation of the LASSO estimator utilizes the R package hdm as described by chernozhukov2016high. For these estimators, we consider the root mean squared deviations of the estimated normalized factors, $(J^{-1} \sum_{j=1}^J (\hat{\theta}_{j} - \theta_{j})^2)^{1/2}$, where $\theta_j = f_{1,j}/f_{1,1}$. The table also shows the RMSE of the new estimators. GVE denotes the grouped variable estimator as in (ref) using averages of $J/2$ instruments. For the GVE, we define $\theta_j = f_{1,j}/\bar{f}_{1}$, where $\bar{f}_{1}$ is constructed as an average of $J/2 - 1$ factors. Finally, WGVE refers to the weighted estimator that uses all partitions and a LASSO procedure in the first step as in (ref). The RMSE is defined as $(J^{-1} \sum_{j=1}^J ( \hat{\vartheta}_{j} - \vartheta_{j} )^2)^{1/2}$, considering $Q_J^\ast = C = Q_J$. The table shows results for different sample sizes of $N = \{50,100\}$ and $J = \{10,20\}$.
As can be seen from Table (ref), the PCA estimator offers excellent performance among existing methods. The IV estimator only performs marginally better than PCA in models with $t_3$ errors, although the difference in performance disappears when $N=100$ and $J=20$. Furthermore, it is interesting to see, although expected, that LASSO outperforms IV when $J$ is relatively large with respect to $N$. The results reflect the well known issues with IV estimation in factor models, while simultaneously demonstrating the advantages of employing the LASSO regression approach for high-dimensional models. In contrast the performance of the proposed GVEs is excellent and, in general, they offer the smallest RMSE across all variants of the model.
We also investigate the relative performance of the estimator in a factor-augmented panel data model. Following closely hP06, we generate observations based on the following model for $i = 1, 2, 3, \hdots, N$ and $j = -S + 1, \hdots, 0, 1, \hdots, J$:
As in the case of the previous one factor model, the error term in equation (ref) is assumed to be distributed as either Gaussian or $t_3$. The error term in equation (ref) is $\bm{v}_{ij} = (v_{1,ij},v_{2,ij})' \sim \mathcal{N}(0,\bm{I})$. Moreover, we set the parameters of the model to generate an endogenous variable, $x_1$, and an exogenous variable, $x_2$. The parameters in equations (ref) and (ref) are $\beta_1=\beta_2=a_1=1$, $b_1=2$, $c_1=0.5$, $\beta_0=a_2=b_2=c_2=0$, and $\rho=0.8$. Lastly, as before, we set $S = 50$ to minimize the effects of the initial values on the outcome, $f_{1,-49}=1$, and $\eta_j$ is an i.i.d. random variable distributed as uniform $\mathcal{U}[0,1]$.
The focus of this investigation is on the estimation of the latent factor structure in the model, therefore we implement our procedure by first estimating the intercept and slopes of the observed part of equation (ref). We employ two consistent estimators: the estimator for an interactive effects model (IEE) proposed by jB09, and the mean group estimator for the common correlated effects model (CCE) developed by hP06. Moreover, we evaluate the performance of the method in relation to the correlation between endogeneous variables and instruments. In Design 1, we assume that $\lambda_{1,i}$ is an i.i.d. random variable distributed as $\mathcal{U}[0.5,3.5]$, while in Design 2, $\lambda_{1,i}$ is an i.i.d. random variable distributed as $\mathcal{N}(0,1)$. Lastly, in Design 3, we generate $\lambda_{1,i} \sim \mathcal{N}(0,1) $, for $i=1,...,m$ and, following aC2011, $\lambda_{1,i} = 0.5 \times \theta_i / \sum_i \theta_i$ for $i=m+1,...,N$, where $\theta_i \sim \mathcal{U}[0,1]$ and $m = 0.9 N$.
We evaluate the performance of our proposed estimators for a factor-augmented model against standard approaches involving PCA and IV. As previously discussed, in the context of a factor-augmented linear model, we can conceive of the feasible estimation of the factor structure in two steps. First, a consistent estimate of the coefficients on the observed variables is necessary to generate residuals. Second, we apply either PCA, IV, LASSO, or the grouped variable estimators proposed in this paper (GVE and WGVE) to estimate the latent factor structure. In Table (ref), we present the small sample performance of different estimators measured by the RMSE, which is defined as in Table (ref). The results show that the performance of the estimators proposed in this paper leads to significant improvements in terms of RMSE relative to the PCA and IV-type approaches.
In order to demonstrate the usefulness of our methods, we now present three applications to models of educational attainment. Depending on the data and setting, the latent factors estimated by our methods can be interpreted as measures of unobserved teaching quality. Furthermore, we can quantify the distribution of latent ability of the students. First, we investigate how the distribution of latent abilities changes over subsequent years of K12 education. Second, we investigate how the distribution of latent abilities changes over subsequent years of college education, and lastly we investigate the change of the distribution of student ability after the implementation of a voucher program designed to improve educational outcomes. While these examples rely on different data sets and come from different countries, they highlight the usefulness of our techniques in varied settings in order to quantify the unobserved dimensions of student and teacher quality.
In our first example, we present evidence on the temporal and geographic variability of educational opportunity across the US by using administrative data from over 11,000 school districts reardon2019educational. We can model the district-level test scores using the following model, which accounts for the impact of latent school-district and grade-level heterogeneity on educational attainment using a one-factor specification:
Here $y_{ij}$ is the average normalized test score in district $i$ in grade $j$. The model also includes district fixed effects, $\mu_{i}$, and grade fixed effects, $\mu_{j}$. In this model, $\lambda_{i}$ is associated with district educational attainment, and the factor $f_{j}$ is interpreted as measuring educational attainment by grade $j$. Moreover, the term $\lambda_i f_{j}$ represents the interaction between educational attainment in district $i$ and quality of instruction in grade $j$. The value of including these latent terms becomes evident once we consider that high grade teaching quality can have a modest effect on the educational attainment of relatively under-performing districts, although it can dramatically impact performance in over-performing districts.
To estimate the parameters in equation (ref), we use data from the Stanford Education Data Archive (SEDA) for the year 2018 fahle2021stanford. SEDA provides nationally comparable scores for school districts in the U.S. The data set includes information on mathematics and reading tests. Unfortunately, the availability of covariates that vary by grade and district is rather limited and it would reduce the sample size significantly. Thus, we do not include independent variables in the specification but account for district and grade fixed effects which we think will capture most of the relevant variation over a short time horizon.
We first estimate equation (ref) using fixed effects methods separately by subject. In the second stage, we use residuals $R_{ij} = \lambda_{i} f_{j} + u_{ij}$, and we estimate the factors and loadings following the the weighted grouped variable estimator (WGVE) with weights equal to $Q_J^{-1}$. In Table (ref) we report the latent educational attainment for each grade (using grade 3 to normalize the results). We notice that for both mathematics and reading the three different estimators suggest a decreasing trend for higher grades.
In Figure (ref) we plot the estimated model $\lambda$s for each school district grouped by state. The two counties with the lowest loadings for mathematics are Oglala Lakota County, SD (-1.379) and Todd County, SD (-1.367). They are the poorest and third poorest counties in the US respectively. In contrast the districts with the highest loadings are Forsyth County, GA (0.925) and Williamson County, TN (0.8218). Adjusted for cost of living they are some of the wealthiest counties in the US. We also notice the distribution of estimated loadings by state. For Alabama the estimate loadings range from -1.209 to barely above zero at 0.074. In contrast the loadings for Massachusetts range from -0.207 to 0.590. It is worth noting that while Georgia has the district with the highest loading, it also has the 5-th lowest loading for Hancock County, GA (-1.396). The racial disparities between these two counties are particularly striking, and this difference is partially removed by the inclusion of the fixed effects. The population in Forsyth County is close to 90% white and in Hancock County is 84% African-American. The Forsyth school district is well-funded and uses technology extensively including tools that allow parents to monitor student assignments and grades 24 hours a day. (Detailed estimation results for both math and reading scores are available from the authors.)
Figure (ref) plots $\lambda_{i} f_{g}$ and allows us to evaluate the change in the distribution of unobserved district educational attainment by grade and by subject. The distributions appear to change little by grade, indicating the lack of significant differences in district educational achievement that could be attributed to the quality of education at different stages of the K-12 education system.
Next, we use data from administrative records of the economics and finance programs at Bocconi University DeGiorgi2012 in order to estimate the following one-factor model of the effect of class size and socioeconomic class composition on educational attainment:
where $y_{icj}$ is the average grade of student $i$ in a class $c$ at year $j$, and $\bm{d}_{cj}$ is a vector of variables that includes class size, and measures of actual dispersion of gender and income in each class. The vector $\bm{x}_{icj}$ captures observed variables such as gender and income. In this model, $\lambda_{i}$ is associated with student motivation and ability, and the factor $f_{ct}$ is interpreted as measuring teaching quality of the course $c$ taken in year $j$. Moreover, the term $\lambda_i f_{cj}$ represents the interaction between student motivation $\lambda_{i}$ and the quality of the teacher in a class $f_{cj}$. The inclusion of interactive latent factors is considered important since it allows us to account for situations where high teaching quality can have a modest effect on the educational attainment of relatively unmotivated students, although it can dramatically affect performance among strong, motivated students.
The data set captures a rich set of covariates which are included in the specification. It includes information on course grades, background demographic and socioeconomic characteristics such us gender, family income, and pre-enrollment test scores. Additionally, the data set includes information on enrollment year, academic program, number of exams by academic year, official enrollment, official proportion of female students in each class, and official proportion of high income students in each class. We restrict our attention to students who matriculated in the 2000 academic year and took the same non-elective classes in the first three years of the program. See DeGiorgi2012 and harding2014estimating for additional details on the data.
As in DeGiorgi2012, we estimate $\bm{\alpha}$ and $\bm{\beta}$ in equation (ref) using instrumental variables generated by a random assignment of students into classes. Students were assigned to each class by the administration at Bocconi University, and therefore, the random assignment determine the actual class size, percentages of female students in a class and high income students in a class, which are considered to be endogenous variables. The use of the randomized assignment allows for the consistent estimation of the coefficients in equation (ref), satisfying one of the conditions of our approach. In the second stage, we employ residuals $R_{icj} = \lambda_{i} f_{cj} + u_{icj}$, and we estimate the factors and loadings following the procedure described in Sections 2 and 3.
Table (ref) shows the factors $f_{cj}$ estimated using PCA, GVE, and WGVE. Because $m_r=1$, the GVE and IV estimators are identical. While PCA suggests that teacher/course quality $f_{cj}$ does seem to improve linearly over years, the WGVE suggests that the quality of courses improves mainly in the third year of the program. In Figure (ref), we estimate the distribution of $\hat{\lambda}_i \hat{f}_{cj}$ by years in the program. It is interesting to see that the middle and upper tail of the distribution changes over time, and by the third year, the conditional distribution of the average grade becomes more dispersed. This finding suggest that the students who remained in the program became more heterogeneous and the latent abilities of the high-performing students improved over time.
Lastly, we investigate the impact of an educational voucher program that provided opportunities for students to attend private schools. During past decades, numerous educational voucher programs were adopted in the U.S. and Latin America. The empirical literature focused on the evaluation of the effect of the program on observable outcomes jAngrist2002,jAngrist2006,LAMARCHE2011278, but the effect of such programs on latent variables such as cognitive ability of students is unknown. In this example, we illustrate the use of our estimation approach using data from jAngrist2002 concerning a 1991 program in Colombia. The vouchers were assigned using lotteries, and they were renewable as long as the students maintained satisfactory academic progress.
We estimate the following factor-augmented linear panel data model:
where $y_{is}$ is student's $i$ test score in subject $s$ and $d_i$ indicates treatment status (i.e., whether student $i$ won the lottery). The parameter $\alpha$ is the mean treatment effect of the program. The vector of independent variables is denoted by $\bm{x}_{ij}$ and the error term by $u_{is}$. The loading $\lambda_i$ measures the student's intrinsic ability or effort that also drives performance in the three subjects, and the variable $f_s$ is a subject specific effect that impacts student achievement.
We use data on 284 students who took tests in mathematics, reading and writing. These tests were taken three years after the vouchers were distributed. To facilitate the comparison among subjects, the test scores are in standard deviation units. In addition to an indicator variable for whether the student won a voucher, we use the following independent variables: site dummies, strata indicators for whether the student lives in a neighborhood ranked on a scale of 1-6 from poorest to richest, an indicator for whether the interview was done by a house visit since telephones were used in the majority of the interviews, gender, age, and parents' schooling. We also include an indicator for survey form, because jAngrist2002 data also incorporate responses obtained from a pilot survey designed to test questions and interviewing strategies.
Table (ref) presents the factors for Mathematics, Reading and Writing. We estimate separately $f_s$ for students in the control and treatment group, to measure whether these factors differ by treatment status. The table also presents results using the estimation approaches introduced in this paper. The results for Mathematics and Writing in the control group are qualitatively similar when using PCA, GVE or WGVE. In contrast, WGVE estimates significant gains in Mathematics for the treatment group. The results do not seem to suggest improvements in the other subjects resulting from the treatment.
Lastly, to summarize the impact of the voucher program on student achievement, we can evaluate the factor structure in our model of academic achievement. Figure (ref) plots the distribution of student's latent ability by treatment status. The figure reveals that the educational policy implemented in Colombia improved latent cognitive outcomes of low-performing students and high-performing students, while increasing the gap between strong and weak students. We also measure a difference of 0.016 between the values of $\lambda_i$ for students in the treatment and control groups.
In this paper, we investigated the estimation of factor-augmented linear models using internally generated instruments in the spirit of madansky1964instrumental, while addressing challenges such as the potentially large number of equally valid instruments. Given that many normalizations are possible for the identification of such a model, we explore the advantages of creating linear combinations of IV estimators, which leads to efficiency improvements.
While the proposed approach is computationally intensive and identification relies on correctly specifying the dependence between the latent factors and the error term, it nevertheless leads to a simple approach to estimating the latent factors in linear models. Further research may involve relaxing the identification assumptions to more general cases and to the extension of approximate factor models.