EconBase
← Back to paper

Normalizations and misspecification in skill formation models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

124,730 characters · 19 sections · 2 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 970 chars of source]

Abstract

An important class of structural models studies the determinants of skill formation and the optimal timing of interventions. In this paper, I provide new identification results for these models and investigate the effects of seemingly innocuous scale and location restrictions on parameters of interest. To do so, I first characterize the identified set of all parameters without these additional restrictions and show that important policy-relevant parameters are point identified under weaker assumptions than commonly used in the literature. The implications of imposing standard scale and location restrictions depend on how the model is specified, but they generally impact the interpretation of parameters and may affect counterfactuals. Importantly, with the popular CES production function, commonly used scale restrictions fix identified parameters and lead to misspecification. Consequently, simply changing the units of measurements of observed variables might yield ineffective investment strategies and misleading policy recommendations. I show how existing estimators can easily be adapted to solve these issues. As a byproduct, this paper also presents a general and formal definition of when restrictions are truly normalizations.

Introduction

Structural models are key tools of economists to simulate changes in the economic environment and evaluate and design policies. An important class of such models deals with skill and human capital formation. Human capital formation is a main part of many structural models and is an important driver of economic growth and inequality MT:16, making policies that target skill formation particularly vital. This growing literature, originating from the seminal papers of \citeN{CH:08} and \citeN{CHS:10}, estimates production functions of various skills of children, studies how past skills, parental skills, and investments affect future skills, and links the skills to adult outcomes. The results provide valuable insights into determinants of skill formation, timing of investments, and the design of optimal interventions for disadvantaged children.

A major challenge in these models is that skills are not directly observable, lack a natural scale, and can only be approximated through measurements, such as test scores. The early literature has provided sufficient conditions for identification of parametric and nonparametric versions of these complex models, which is often achieved using a two-step approach. First, the distribution of skills is identified from the measurements, and the production function is then identified in the second step (see e.g. \shortciteN{CHS:10}, \citeN{AMNS:17}, and \citeN{AMN:19}). To obtain point identification in the first step, it is necessary to fix the unknown scales and locations of skills. However, it remains unclear whether these restrictions are still required when combined with common parametric assumptions on the production function in the second step.

This paper presents a new identification analysis for skill formation models and investigate the consequences of seemingly innocuous normalizations. Instead of providing sufficient conditions for point identification using multi-step arguments, I start by pooling all parts of the model and characterize the identified set of all parameters without the scale and location restrictions. This approach reveals which parameters are point identified or partially identified and how the additional restrictions affect the identified set. Notably, I show that many critical features of these models are invariant to scale and location restrictions, point-identified without them, and identified under less restrictive assumptions than previously considered. These results apply to both parametric and nonparametric versions of the model.

The exact implications of imposing scale and location restrictions on parameters and counterfactuals depend on the model specification and the object of interest. Specifically, a restriction could be a harmless normalization with one production function, but could impose strong assumptions with a different one. I thus analyze two popular specifications, namely the trans-log and the CES production functions. For the trans-log case, it turns out that standard scale and location restrictions do not impose additional testable restrictions, simply select an element of the identified set, and many objects of interest are invariant to them. However, they still affect production function parameters and certain counterfactuals. Since such restrictions are often arbitrary and tied to the units of measurement of the data, these features can be hard to interpret, difficult to compare across different studies, and potentially mislead policy recommendations. More importantly, for the CES case, I show that commonly restricted scale parameters are in fact identified and setting them to specific values is typically inconsistent with the data. Consequently, with these restrictions, simply changing the units of a skill measure (e.g. from years to months) can impact estimated dynamics, persistence of skills, effects of parental investments, and optimal investment strategies.

In addition to these identification results, I demonstrate how existing estimators can be modified to estimate all identified parameters. While some parameters remain difficult to interpret, the modified estimator enables the calculation of point identified features and counterfactuals that are invariant to scale and location restrictions and to the units of measurement. I illustrate these results through Monte Carlo simulations and an empirical application based on the framework and estimator of \shortciteN{AMN:19}.

A broader takeaway is that researchers should carefully verify whether imposed restrictions are in fact normalizations (as formally defined in Definition (ref)). This is particularly crucial in structural models where identification often proceeds in multiple steps, and supposedly innocuous restrictions in one step may unintentionally affect results in subsequent steps. By focusing on features of the model that are invariant to these restrictions, researchers can ensure that their findings remain interpretable and comparable across different studies.

Literature: Following the influential work of \citeN{CH:08} and \shortciteN{CHS:10}, a growing body of literature has studied the development of latent variables. For example, \citeN{HP:11} study the determinants of children's cognitive and non-cognitive skills using Indian data, \citeN{FK:14} investigate how time allocation affects both cognitive and non-cognitive development using Australian data, \shortciteN{AMN:19} and \shortciteN{AMNS:17} estimate the effects of health and cognition on human capital with data from India and Ethiopia/Peru, respectively, and \shortciteN{ACDMR:19} study which intervention led to gains in cognitive and socio-emotional skills using Colombian data. This literature provides important insights into the nature of persistence, dynamic complementarities, and the optimal targeting of interventions.\footnote{See also \citeN{CH:07}, \citeN{CH:09}, \citeN{Cunha:11}, \citeN{HPS:13}, \citeN{AJ:16}, \citeN{HAP:17}, \citeN{ACM:2022} and references therein.}

In these studies, the measurement system has a factor structure that requires scale and location restrictions to identify the distribution of latent variables. In particular, \citeN{CH:08} study a model where the production function is log-linear. They utilize the two-step identification arguments described above, except that they impose the scale and location restrictions using an adult outcome to obtain well-defined units of measurement (“anchoring”). They also discuss which parameters depend on the anchor or its units of measurement. In this model, I demonstrate that certain policy-relevant parameters, such as optimal investment sequences, may depend on the specific scales and units of measurements of the observed variables. I further derive a set of features that is identified under weaker restrictions on the measures and the production function and without anchoring the skills or fixing their scales. Importantly, I also provide identification results for alternaive production technologies, such as the widely used CES production function, where standard restrictions are unnecessary for identification. In these case, imposing them through an initial period normalization or through anchoring leads to misspecification.

Common restrictions are to fix the scales and locations of each latent factor in each time period by setting parameters in the measurement system.\footnote{For example, \citeN{CH:08} and \shortciteN{CHS:10} set the scale of the first skill measure to 1 in each period - see the discussion around equations (7) -- (8) of \citeN{CH:08} and the first paragraph on page 891 of \shortciteN{CHS:10}. In addition, as discussed in footnote 17 of \citeN{CH:08}, their identification results rely on either setting one of the location parameters to $0$ in each time period or setting the mean of the skill to $0$ in each time period. These restrictions have also been used in subsequent papers such as \shortciteN{AMN:19}. The nonparametric identification results in \shortciteN{CHS:10} impose analogous restrictions in their Assumption (v) of Theorem 2.} If the same values are used across all time periods, Agostinelli and Wiswall (2016a, 2016b, 2024) refer to this as an age-invariance assumption (see Assumption (ref)(a) for a formal definition or Definition 1 of Agostinelli and Wiswall (2024)). In important contributions, Agostinelli and Wiswall (2016a, 2016b, 2024) demonstrate that imposing this assumption can yield a misspecified model when the production function has a known scale and location. They also propose relaxations of production functions with age-invariant measures and show that, even without age-invariance, production function restrictions can yield point identification (see Corollary (ref) below and the related discussion). However, in both cases, they impose scale and location restrictions in the first period, arguing that these are necessary for point identification. While this is true for the trans-log production function used in their application, I show that the scale restriction is not required for the CES production function. Furthermore, I demonstrate for different specifications that while production function parameters, certain counterfactuals, and estimated dynamics may depend on specific scales and units of measurements, many key parameters are invariant to these restrictions, including age-invariance, and are in fact point identified without them.

\nocite{AW:16a}\nocite{AW:16b}\nocite{AW:22}

In independent research, \citeN{DKP:20} show that with a trans-log production function, anchored treatment effects are invariant to scale and location restrictions and are identified without age-invariance. Through simulations, they also show that standard restrictions lead to inconsistent estimated treatment effects with the CES production function. In the trans-log case, these results aligns with part 4 of Theorem (ref) below. While their proof is specific to the trans-log production function with a log-linear measurement system, my results extend to other cases. Additionally, I show that only skill measures in the first period are needed to identify anchored treatment effects, reducing the data requirements and assumptions considerably. I also consider other policy-relevant features. Additionally, for the CES case, I show why standard restrictions can cause misspecification and describe how existing estimators can be adapted to resolve this issue.

Table 1 of \shortciteN{ACM:2022} provides a selective overview of specifications and estimation methods used in well-cited papers in the literature. It shows that with the exception of \shortciteN{CHS:10}, earlier work predominantly employed the Cobb-Douglas specification, likely due to the simplicity of the corresponding estimator. Since \shortciteN{AMN:19}, there has been increasing adoption of the (nested) CES production function, facilitated by their computationally simple two-step estimator (compared to the MLE of \shortciteN{CHS:10}). Next to the papers in this table, recent contributions using this approach include \shortciteN{ANT:23}, \shortciteN{BFHD:24}, and \citeN{GG:23}. The trans-log production function proposed by \citeN{AW:22} is nonnested with the CES specification and can also be estimated using an IV/GMM approach.

The consequences of normalizations have been discussed in various contexts. In factor models, while certain restrictions are necessary for point identification (see e.g. \citeN{AR:56} or \citeN{Madansky:64}), \citeN{Williams:20} shows that certain features, such as variance decompositions, can be identified without them. I combine a factor model with a production function, which provides additional restrictions. Many studies argue and show in specific examples that critical features should not depend on normalizations, see e.g. \citeN{Freyberger:18} and \shortciteN{KSSS:18}. Similar to this paper, but in a very different context, \citeN{AS:14} discuss restrictions that were considered normalizations, but are restrictive assumptions. \shortciteN{KSS:20} show that certain counterfactuals in dynamic discrete choice models are identified, even when the model itself is not. \shortciteN{RWZ:10} define a normalization in vector autoregressive models to pin down unidentified signs; see end of Section (ref) for more details. Matzkin (1994, 2007) discusses several examples of normalizations, some of which are motivated by economic theory. \citeN{Lewbel:19} provides an informal discussion of normalizations, which is conceptually very similar to the formal definition I provide below. When normalizing restrictions are needed for point identification, there are often multiple ways to impose them when estimating the model. Good choices can then yield particularly convenient restrictions on the parameter space (as in \citeN{GL:19}) or even faster rates of convergence (as in \shortciteN{CKK:15}). See also \shortciteN{HWZ:07} for a discussion on estimation with normalizations.

\nocite{Matzkin:94}\nocite{Matzkin:07}

Structure: In Section (ref), I provide a formal definition of a normalization, which to the best my knowledge currently does not exist in the literature, as well as illustrative examples. Section (ref) contains the identification analysis of different parametric skill formation models. A nonparametric version is discussed in Appendix (ref). Sections (ref) and (ref) contain the Monte Carlo simulations and the empirical application, respectively. All proofs are in the appendix.

Normalizations

I begin by providing a formal definition of a normalization, which serves as a basis for the subsequent analysis. I illustrate the definition and potential problems using a probit model and a simple version of the skill formation model.

General definition and illustration

Suppose we have a model where $\tau_0 \in \mathcal{T}$ denotes the true values of the parameters and $\mathcal{T}$ is the parameter space. Here $\tau_0$ could be the coefficients in a regression model or the parameters in a skill formation model. If the model is semiparametric, $\tau_0$ could also contain unknown functions. Let $Z$ contain all observed random variables, such as $Y$ and $X$, with distribution $P(Z)$. For any $\tau \in \mathcal{T}$, the model generates a joint distribution of the data $Z$, denoted by $P(Z,\tau)$. Since the model is assumed to be correctly specified, the true distribution of $Z$ is $P(Z,\tau_{0})$. The model typically contains certain assumptions, such as functional form or independence assumptions, but suppose that so far none of the normalizations are imposed. The identified set for $\tau_{0}$ is $\mathcal{T}_0 = \{ \tau \in \mathcal{T}: P(Z,\tau) = P(Z,\tau_0) \}.$ If $\mathcal{T}_0$ is a singleton, $\tau_0$ is point identified. We say that $\tau_1, \tau_2 \in \mathcal{T}$ are observationally equivalent if they generate the same distribution of the data: $P(Z,\tau_1) = P(Z,\tau_2)$. Let $g(\tau_0)$ be a function of interest, such as a counterfactual. The identified set for $g(\tau_{0})$ is $\mathcal{T}_{g_0} = \{ g(\tau): \tau \in \mathcal{T}_0 \}.$ Notice that $g(\tau_0)$ could be point identified (i.e. $\mathcal{T}_{g_0}$ is a singleton) even if $\tau_0$ is not.

In models with normalizations, $\mathcal{T}_0$ is typically not a singleton. A normalization is a restriction of the form $\tau \in \mathcal{T}_N$, where $\mathcal{T}_N \subseteq \mathcal{T}$ is a known set. Hence, a normalization restricts the feasible values of $\tau$, such as setting an element to $1$. I define a restriction to be a normalization with respect to a function $g(\tau_0)$ if it does not change the identified set of $g(\tau_0)$.

definitionThe restriction $\tau \in \mathcal{T}_N$ is a normalization with respect to $g(\tau_0)$ if $\{g(\tau): \tau \in \mathcal{T}_0 \cap \mathcal{T}_N \} = \{g(\tau): \tau \in \mathcal{T}_0 \}$ for all $\tau_0 \in \mathcal{T}$.

Typically, $\mathcal{T}_0 \cap \mathcal{T}_N$ is a singleton. That is, we achieve point identification with the additional restrictions. Definition (ref) then implies $\{g(\tau_0)\} = \{g(\tau): \tau \in \mathcal{T}_0 \cap \mathcal{T}_N \} = \{g(\tau): \tau \in \mathcal{T}_0 \}$ and thus, that $g(\tau_0)$ is point identified, even without the restriction $\tau \in \mathcal{T}_N$. Since these restrictions are often arbitrary, $\tau_0$ is usually not in $ \mathcal{T}_N$ in which case the restriction $\tau \in \mathcal{T}_N$ is not a normalization with respect to $\tau_0$, but it can be a normalization with respect to particular functions of interest. Moreover, the restriction can be a normalization for some function and not for others. Hence, researchers need to argue that normalizations hold with respect to all functions of interest, such as all counterfactuals. Finally, a normalization cannot impose any additional overidentifying restrictions in the sense that if $\mathcal{T}_0 \neq \emptyset$, then $\mathcal{T}_0 \cap \mathcal{T}_N \neq \emptyset$.

As a simple example, consider the probit model where $Y = \ensuremath{\mathbf{1}}(\beta_{0,1} + \beta_{0,2}X \geq U)$, $var(X) > 0$, $U \mid X \sim N(\mu_0, \sigma_0^2)$ and $\sigma_0^2 > 0$. The true parameter vector is $\tau_0 = (\beta_{0,1}, \beta_{0,2},\mu_0,\sigma_0)'$ and $Z = (Y,X)$. Now notice that $$P(Y=1 \mid X = x) = \Phi \left( \frac{\beta_{0,1}-\mu_0}{\sigma_0} + \frac{\beta_{0,2}}{\sigma_0}x\right),$$ where $\Phi$ denotes the standard normal cdf. Since $var(X) > 0$, $\frac{\beta_{0,1}-\mu_0}{\sigma_0}$ and $\frac{\beta_{0,2}}{\sigma_0}$ are point identified. It is also well known and easy to see that $$\mathcal{T}_0 = \left\{\tau \in \ensuremath{\mathbb{R}}^3 \times \ensuremath{\mathbb{R}}_{>0}: \frac{\beta_{1}-\mu}{\sigma} = \frac{\beta_{0,1}-\mu_0}{\sigma_0} \text{ and } \frac{\beta_{2}}{\sigma}= \frac{\beta_{0,2}}{\sigma_0} \right\}$$ because all values in $\mathcal{T}_0$ imply the same joint distribution of $(Y,X)$.

Since $\tau_0$ is not point identified, it is common to set $\mu = 0$ and $\sigma = 1$. Using the previous notation, this means that $\mathcal{T}_N = \ensuremath{\mathbb{R}}^2 \times 0 \times 1$ and $\mathcal{T}_0 \cap \mathcal{T}_N = \left(\frac{\beta_{0,1}-\mu_0}{\sigma_0} , \frac{\beta_{0,2}}{\sigma_0} , 0, 1\right).$ Clearly, this restriction is not a normalization with respect to $\beta_{0,1}$ or $\beta_{0,2}$, which are typically not objects of interest. In fact, in general $\tau_0 \notin \mathcal{T}_0 \cap \mathcal{T}_N$ unless $\mu_0 = 0$ and $\sigma_0 = 1$. However, this restriction is a normalization with respect to (potentially counterfactual) probabilities $$ P(Y = 1 \mid X = x) = \Phi \left( \frac{\beta_{0,1}-\mu_0}{\sigma_0} + \frac{\beta_{0,2}}{\sigma_0}x\right) $$ or, when $X$ is continuous, marginal effects $\frac{\partial }{\partial x} P(Y = 1 \mid X = x)$. Clearly, these features are point identified even though $\tau_0$ is not.

Normalizations may help to provide useful structural interpretations of other parameters. As an example, suppose there are two covariates, $Y = \ensuremath{\mathbf{1}}(\beta_{0,1} + \beta_{0,2}X_1 + \beta_{0,3}X_2 \geq U)$, and $\beta_{0,2} > 0$ . Instead of setting $\mu = 0$ and $\sigma = 1$, we could impose $\beta_{0,1} = 0$ and $\beta_{0,2} = 1$, which are also normalizations with respect to marginal effects. In addition, we then obtain a model of the form $Y = \ensuremath{\mathbf{1}}( X_1 + \tilde{\beta}_{0,3}X_2 \geq \tilde{U})$ with $ \tilde{\beta}_{0,3} = \beta_{0,3}/\beta_{0,2}$, which can then be interpreted as a relative effect (which is also identified without scale and location restrictions).

In the context of vector autoregressive models, \shortciteN{RWZ:10} define a normalization as a restriction on the parameter space that pins down unidentified signs of coefficients. These restrictions are imposed in addition to other assumptions, such as long run restrictions. Just like above, their restrictions do not impose additional testable assumptions and they can help to provide structural interpretations of certain parameters. Unlike my definition, it is not clear from theirs whether these restrictions are without loss of generality in the sense that they do not affect functions of interest. If they do, one would have to argue why they are reasonable.

Simple skill formation model

As a more involved example, I now discuss a very simple skill formation model, which imposes very restrictive assumptions, but illustrates the previous definition and points out the types of problems that occur in more general models.

Let $\theta_{t}$ denote skills at time $t$ and let $I_t$ be investment at time $t$, where $t = 0,1,\ldots,T$. We are interested in the roles of investment and past skills in the development of future skills. Instead of skills, the data only contains measurements of them, denoted by $Z_{\theta,t,m}$. In this section, I first consider the simplest possible model without measurement error and a single measure $Z_{\theta,t,1}$ that takes the form $Z_{\theta,t,1} = \lambda_{\theta,t,1} \ln \theta_{t}$, where $\{\lambda_{\theta,t,1}\}^T_{t=0}$ are unknown parameters with $\lambda_{\theta,t,1} \neq 0$ for all $t$. In other words, in each period, we observe a scaled version of log-skills. I also assume that there are three periods ($T=2$) and that investment is observed. Finally, for simplicity, I parameterize the marginal distribution of skills in the initial period as $\ln \theta_{0} \sim N(0, s_0^2)$, which implies that $E[Z_{\theta,0,1}] = 0$ and $Var(Z_{\theta,0,1}) = \lambda_{\theta,0,1}^2 s_0^2$.

I consider two simple production functions without unobserved random variables. First suppose skills evolve based on the Cobb-Douglas production function: $$\ln \theta_{t+1} = a_t + \gamma_{1t}\ln \theta_{t} + \gamma_{2t} \ln I_{t}.$$ Define the vector containing the true values of all ten parameters as $$\tau_0 = (\tau_{0,1}, \tau_{0,2}, \ldots, \tau_{0,10}) = ( \lambda_{\theta,0,1}, \lambda_{\theta,1,1}, \lambda_{\theta,2,1}, a_0, a_1, \gamma_{10}, \gamma_{20}, \gamma_{11}, \gamma_{21}, s^2_0 ).$$ Without further assumptions, $\tau_0$ is not point identified. To see the restrictions imposed by the obervables, notice that $\ln \theta_{t} = Z_{\theta,t,1}/ \lambda_{\theta,t,1} $ and therefore $$Z_{\theta,t+1,1} = \lambda_{\theta,t+1,1} a_t + \frac{\lambda_{\theta,t+1,1}}{\lambda_{\theta,t,1}} \gamma_{1t}Z_{\theta,t,1} + \lambda_{\theta,t+1,1}\gamma_{2t} \ln I_{t}.$$ Hence, the joint distribution of $\{Z_{\theta,t+1,1},Z_{\theta,t,1}, I_{t}\}^1_{t=0}$ identifies (a) the scaled intercepts of the production function $ \lambda_{\theta,t+1,1} a_t$, (b) the scaled slope coefficients $\frac{\lambda_{\theta,t+1,1}}{\lambda_{\theta,t,1}} \gamma_{1t}$ and $\lambda_{\theta,t+1,1}\gamma_{2t}$, and (c) the scale variance $ \lambda_{\theta,0,1}^2 s_0^2$. The identified set can be shown to consist of all parameters that imply the same values of these identified features as the true parameters. Formally,

align*[align* omitted — 462 chars of source]

where the three rows correspond to the three sets of features above. To achieve point identification, we could set $ \lambda_{\theta,t,1} = 1$ for all $t$, in which case the intersection of $\mathcal{T}_0$ and the additional restrictions is the singleton $ \big\{ \big( 1, 1, 1, \tau_{0,2} \tau_{0,4}, \tau_{0,3} \tau_{0,5}, \frac{\tau_{0,2}}{\tau_{0,1}} \tau_{0,6}, \tau_{0,2} \tau_{0,7}, \frac{\tau_{0,3}}{\tau_{0,2}} \tau_{0,8}, \tau_{0,3} \tau_{0,9}, \tau_{0,1}^2 \tau_{0,10} \big) \big\}. $

As a simple numerical example, suppose $Z_{\theta,t,1} = 12 \ln \theta_{t}$ and

eqnarray*[eqnarray* omitted — 72 chars of source]

for all $t$, and $\ln \theta_{0} \sim N(0,1) $. Here $\tau_0 = \left(12,12,12,0,0,0.5,0.5,0.5,0.5,1\right)$. If we set $\lambda_{\theta,t,1} = 1$, even though the true value is $12$, we essentially treat $ \ln \tilde{\theta}_{t} \equiv Z_{\theta,t,1} = 12 \ln \theta_{t}$ as the skills and the corresponding production function is $$\ln \tilde{\theta}_{t+1} = 0.5 \ln \tilde{\theta}_{t} + 6 \ln I_t. $$ The intersections of the identified set and set of additional restrictions is then the singleton $\left\{ \left(1,1,1,0,0,0.5,6,0.5,6,144\right) \right\}$ which would be the parameter estimated in practice (instead of $\tau_0$). The coefficients in front of investment are thus hard to interpret. For example, using $\tau_0$ (or setting $\lambda_{\theta,t,1} = 12$) one might conclude that increasing investment by 1% increases skills by $0.5\%$, but with the restriction $\lambda_{\theta,t,1} = 1$ that effect changes to $6\%$. Hence, $\lambda_{\theta,t,1} = 1$ is not a normalization with respect to these parameters. The coefficient in front of log-skills is invariant to the scale restrictions, but only because here $\lambda_{\theta,t,1} = \lambda_{\theta,t+1,1} $ for all $t$. Notice that $Z_{\theta,t,1}$ is simply a scaled version of $\ln \theta_{t} $. If we had an alternative measure with a different scale (say $Z_{\theta,t,2} = \lambda_{\theta,t,2} \ln \theta_{t}$), we would generally obtain different parameters if we replaced $Z_{\theta,t,1}$ with $Z_{\theta,t,2}$. This alternative measure could for example result from changing the units of measurements of $Z_{\theta,t,1}$, such as using years instead of months of education.

Even though the production function parameters are not identified, there are potential interpretations that adapt to the units of measurements. For example, the identified parameter $\lambda_{\theta,t+1,1}\gamma_{2t}$ tells us the effect of a one unit increase in $\ln(I_t)$ on skills, measured in the units of $Z_{\theta,t+1,1}$ (see Section (ref) for a specific example). It can also be shown that

align*[align* omitted — 472 chars of source]

where $i_{t}$ is a fixed level of investment, $F_{\ln\theta_{t+1}}$ is the cdf of $\ln\theta_{t+1}$, and $Q_{\alpha}(\theta_{t})$ is the $\alpha$-quantile of $\theta_{t}$. Since the right hand side only depends on identified parameters, similar to marginal effects in the probit model, we can identify how exogenous changes in someones investment affects her rank in the skill distribution at time $t+1$ for a given skill quantile at time $t$.

As another example, first combine the production functions from two periods, write $$\ln \theta_{2} = a_1 + \gamma_{11}a_0 + \gamma_{11} \gamma_{10}\ln \theta_{0} + \gamma_{11}\gamma_{20} \ln I_{0} + \gamma_{21} \ln I_{1},$$ and for a given investment sequence $(i_0,i_1)$ define $$\ln \theta_{2}(i_0,i_1) = a_1 + \gamma_{11}a_0 + \gamma_{11} \gamma_{10}\ln \theta_{0} + \gamma_{11}\gamma_{20} \ln i_{0} + \gamma_{21} \ln i_{1}$$ We can rewrite this equation in terms of identified features as

align*[align* omitted — 523 chars of source]

showing that we can identify the distribution of $\lambda_{\theta,2,1} \ln \theta_{2}(i_0,i_1)$, which we can interpret as a counterfactual measure for a given investment sequence. Using this result, we can also identify the sequence of investment which maximizes expected log-skills. That is, consider $g(\tau_0) \in \ensuremath{\mathbb{R}}^2$ defined as

align*[align* omitted — 127 chars of source]

where $\mathcal{I}$ is a set of feasible investments. Since the maximizer is invariant to changes in scales and locations of the objective function, $g(\tau_0) = \operatorname*{arg\,max}_{(i_{0},i_{1}) \in \mathcal{I}} E \left[\lambda_{\theta,2,1} \ln \theta_{2}(i_0,i_1) \right] $, showing that $ g(\tau_0)$ is identified. All the identified features above are invariant to the restriction $\lambda_{\theta,t,1} = 1$, implying that $\lambda_{\theta,t,1} = 1$ is a normalization with respect to these features.

Next to primitive parameters, some counterfactuals reported in applications are not be invariant to $\lambda_{\theta,t,1} = 1$ either. As an example, suppose the production function for $\ln \theta_1$ also includes an additive interaction term, $\gamma_{30}\ln \theta_0 \ln I_0$, in which case it can be shown that

align*[align* omitted — 175 chars of source]

for identified parameters $\{\zeta_j\}^5_{j=1}$. Using $ \ln \theta_{0} \sim N(0,s_0^2)$ it follows that $$E[\theta_{2}(i_0,i_1)^{\lambda_{\theta,2,1}}] = \exp\left(\zeta_0 + \zeta_3 \ln i_{0} + \zeta_4 \ln i_{1} + \left( \zeta_1 + \zeta_5 \ln i_{0} \right)^2 s_0^2 \right) $$ but $$E[\theta_{2}(i_0,i_1)] = \exp\left( \frac{1}{\lambda_{\theta,2,1}} \left(\zeta_0 + \zeta_3 \ln i_{0} + \zeta_4 \ln i_{1} + \left( \zeta_1 + \zeta_5 \ln i_{0} \right)^2 \frac{s_0^2}{\lambda_{\theta,2,1}} \right) \right) $$ which implies that generally $\operatorname*{arg\,max}_{(i_{0},i_{1}) \in \mathcal{I}} E \left[ \theta_{2}(i_0,i_1) \right] \neq \operatorname*{arg\,max}_{(i_{0},i_{1}) \in \mathcal{I}} E[\theta_{2}(i_0,i_1)^{\lambda_{\theta,2,1}}] $ Hence, using different scaled versions of log-skills (or different fixed values of $\lambda_{\theta,t,1}$) results in different optimal investment sequences when studying the skill level.

Finally, consider the CES production function \[ \theta_{t+1} = ( \gamma_{1t} \theta_{t}^{\sigma_{t}} + \gamma_{2t} I_{t}^{\sigma_{t}} )^{1/\sigma_{t}} \] The parameter vector is now $\tau_0 = ( \lambda_{\theta,0,1}, \lambda_{\theta,1,1}, \lambda_{\theta,2,1}, \sigma_0, \sigma_1, \gamma_{10}, \gamma_{20}, \gamma_{11}, \gamma_{21}, s^2_0 )$ and we can write the production function in terms of observables as \[ \exp(Z_{\theta,t+1,1}) = ( \gamma_{1t} \exp(Z_{\theta,t,1})^{\sigma_{t}/\lambda_{\theta,t,1}} + \gamma_{2t} I_{t}^{\sigma_{t}} )^{\lambda_{\theta,t+1,1}/\sigma_{t}} \] Due to the functional form restrictions of the CES production function, it can be shown that $\tau_0 $ is point identified without any additional restrictions.

Thus, the commonly imposed scale restrictions ($\lambda_{\theta,t,1} = 1$ for all $t$) lead to misspecification unless the true values are all $1$. Consequently, imposing this restriction yields different conclusions for different scaled versions of log-skills (or different fixed values of $\lambda_{\theta,t,1}$) irrespective of the counterfactual. Similar as in the trans-log case, many important features are identified without additional restrictions and are invariant to scaling of the measures.

These results generalize to more complicated settings as discussed in the next section.

Skill formation models

Model

I now discuss issues arising from normalizations in a general class of skill formation models. As before, $\theta_{t}$ and $I_t$ denote skills and investment at time $t$, respectively. Now neither skills nor investment are directly observed and we denote the observed measurements by $Z_{\theta,t,m}$ and $Z_{I,t,m}$, respectively. Specifically, I consider the model:

eqnarray[eqnarray omitted — 457 chars of source]

The first equation describes the production technology with a production function $f$ that depends on skills and investment at time $t$, a parameter vector $\delta_{t}$, and an unobserved shock $\eta_{\theta,t}$. The second and the third equation describe the measurement system for unobserved (latent) skills $\theta_{t}$ and unobserved investment $I_t$, respectively. Observed investment is a special case with $ \mu_{I,t,m} = 0$, $\lambda_{I,t,m} = 1$, and $\varepsilon_{I,t,m} = 0$ for all $m$ and $t$ in which case $Z_{I,t,m} = \ln I_{t}$.

Next, I introduce two equations to allow for endogenous investment and anchoring at an adult outcome. If investment is exogenous, in the sense that $\eta_{\theta,t}$ is independent of $I_{t}$, then these equations are not needed for the main identification results. That is, let

eqnarray[eqnarray omitted — 239 chars of source]

Here $Y_t$ is parental income (or another exogenous variable that affects investment) and $Q$ is an adult outcome, such as earnings or education. An adult outcome does not necessarily have to be available and we can simply use a skill measure in period $T$ in its place.

In summary, the observed variables are income $\{Y_t\}_{t=0}^{T-1}$, the measures $\{Z_{\theta,t,m}\}_{t=0,\ldots,T, m= 1,2}$ and $\{ Z_{I,t,m} \}_{t=0,\ldots,T-1, m= 1,2}$, and the adult outcome $Q$, but we neither observe skills $\{\theta_{t}\}^{T}_{t=0}$ nor investment $\{I_t\}_{t=0}^{T-1}$. We also do not observe any of the errors/shocks in the five equations. The parameters are $\{\mu_{\theta,t,m},\lambda_{\theta,t,m}\}_{t=0,\ldots,T,m=1,2}$, $\{\mu_{I,t,m},\lambda_{I,t,m}\}_{t=0,\ldots,T-1,m=1,2}$, $\{\delta_t\}^{T-1}_{t=0}$, $\{ \beta_{0t}, \beta_{1t}, \beta_{2t}\}^{T-1}_{t=0}$, and $(\rho_0,\rho_1)$.

In the following analysis, I consider the two most commonly used forms for the production technology in the empirical literature, namely the trans-log production function with

equation[equation omitted — 172 chars of source]

and parameter vector $\delta_{t} = (a_t, \gamma_{1t},\gamma_{2t},\gamma_{3t})$ and the CES production function with

equation[equation omitted — 165 chars of source]

and parameter vector $\delta_{t} = (\gamma_{1t},\gamma_{2t},\sigma_{t},\psi_t)$. When $\gamma_{3t} = 0$, the trans-log reduces to the Cobb-Douglas production function.

I now state several additional assumptions that are common in the literature.

assumption\qquad \begin{enumerate}[(a)] • $\{\{\varepsilon_{\theta,t,m}\}_{t=0,\ldots,T, m= 1,2}, \{\varepsilon_{I,t,m} \}_{t=0,\ldots,T-1, m= 1,2}, \eta_Q\}$ are jointly independent and independent of $\{\{\theta_{t}\}^{T}_{t=0},\{I_t\}_{t=0}^{T-1}\}$ conditional on $\{Y_t\}_{t=0}^{T-1}$. • All random variables have bounded first and second moments. • $E[\varepsilon_{\theta,t,m}] = E[\varepsilon_{I,t,m}] = E[\varepsilon_Q] = 0$ for all $t$ and $m$. • $\lambda_{\theta,t,m}, \lambda_{I,t,m} \neq 0$ for all $t$ and $m$. • For all $t \in \{0,\ldots,T\}$, $cov(\ln \theta_t,\ln I_s) \neq 0$ for some $s \in \{0,\ldots, T-1\}$ or $cov(\ln \theta_t,\ln \theta_s) \neq 0$ for some $s \in \{0, \ldots, T\} \backslash t$ . For all $t \in \{0,1,\ldots,T-1\}$, $cov(\ln I_t,\ln \theta_s) \neq 0$ for some $s \in \{0, \ldots, T\}$ or $cov(\ln I_t,\ln I_s) \neq 0$ for some $s \in \{0, \ldots, T-1\} \backslash t$. • For all $t$ and $m$ the real zeros of the characteristic functions of $\varepsilon_{\theta,t,m}$ are isolated and are distinct from those of its derivatives. Identical conditions hold for the characteristic functions of $\varepsilon_{I,t,m}$ and $\eta_{Q}$. • The support of $(\theta_{t},I_{t},Y_t)$ includes an open ball in $\ensuremath{\mathbb{R}}^3$ for all $t$. • $E[\eta_{I,t} \mid \theta_{t}, Y_t] = 0$ and $E[\eta_{\theta,t} \mid \theta_{t},\eta_{I,t} , Y_t] = \kappa_t \eta_{I,t}$ for all $t$. \end{enumerate}

Part (a) imposes common independence assumptions on the measurement errors. Importantly, $I_t$ and $\theta_t$ are not independent and $I_t$ may be endogenous and contemporaneously correlated with $\eta_{\theta,t}$. Part (b) is a standard restriction, part (c) is needed because all measurement equations contain an intercept, and part (d) ensures that the skills actually affect the measures. Part (e) requires that skills and investment are correlated in some time periods. Sufficient conditions are that $cov(\ln \theta_{t+1}, \ln \theta_t) \neq 0$ and $cov(\ln I_{t+1}, \ln I_t) \neq 0$ for all $t$. Under parts (a) and (d) zero covariances of the latent variables are identified because, for example, $cov(\ln \theta_{t}, \ln \theta_s) = 0$ if and only if $cov( Z_{\theta,t,1}, Z_{\theta,s,1}) = 0$. Notice that I only require two measures in each period. One can drop part (e) by assuming that three measures are available. Part (f) contains weak regularity conditions needed for nonparametric identification of the distributions of skills and investment and that hold for most common distributions. Part (g) is a mild support condition that ensures sufficient variation of $(\theta_{t},I_{t},Y_t)$. Part (h) implies that $Y_t$ can serve as an instrument with identification based on a control function argument, as in \shortciteN{AMN:19}. Linearity of the conditional mean function can be relaxed to allow for more flexible functional forms. Exogenous investment is a special case with $\kappa_{t} = 0$.

Under parts (a)--(f) of Assumption (ref) we get the following result.

lemmaSuppose that parts (a)--(f) of Assumption (ref) hold. Then the joint distribution of $$ \left( \{\mu_{\theta,t,m} + \lambda_{\theta,t,m} \ln \theta_{t}\}_{t=0,\dots, T,m =1,2}, \{ \mu_{I,t,m} + \lambda_{I,t,m} \ln I_{t} \}_{t=0,\dots, T-1,m =1,2}, \rho_0 + \rho_1 \ln \theta_{T} \right) $$ is point identified conditional on $\{Y_t\}_{t=0}^{T-1}$.

The proof follows from an extension of Kotlarski's Lemma EW:12. Lemma (ref) shows that under Assumption (ref), the distribution of linear combinations of log-skills and investments is identified. However, Assumption (ref) does not imply identification of the parameters in equations ((ref))--((ref)), such as $\delta_t$. Below I discuss additional assumptions, which have been used in the literature, to achieve point identification of two sets of parameters: (i) the primitive parameters of the model ((ref))--((ref)): $\{\mu_{\theta,t,m},\lambda_{\theta,t,m}\}_{t=0,\ldots,T,m=1,2}$, $\{\mu_{I,t,m},\lambda_{I,t,m}\}_{t=0,\ldots,T-1,m=1,2}$, $\{\delta_t\}^{T-1}_{t=0}$, $\{ \beta_{0t}, \beta_{1t}, \beta_{2t}\}^{T-1}_{t=0}$, and $(\rho_0,\rho_1)$ and (ii) “policy relevant” parameters, such as how changes in investment or income affect $Q$. I consider various combinations of assumptions and I discuss in what sense restrictions are normalizations.

There are two separate issues concerning identification and normalizations in this model. First, the previous literature has focused on sufficient conditions for point identification, which includes scale and location restrictions. However, it is unclear whether these restrictions are necessary. If they instead restrict identified parameters, the model may be misspecified and estimators are generally inconsistent. Second, even if the restrictions simply select an element of the identified set, it is important to understand which features are invariant to arbitrary scale and location restrictions. Although the implications in the CES case are more interesting, I start with the more transparent trans-log case where only the second issue arises. As pointed out before, whether or not a restriction is a normalization depends on the model, and parameters may be invariant in some settings, but not in others.

Trans-log production function

In this section I consider the trans-log production function given in equation ((ref)). I first introduce additional assumptions that are commonly used in the literature.

assumption$\lambda_{\theta,0,1}=1$ and $\mu_{\theta,0,1} = 0$.
assumption\quad \begin{enumerate}[\label=(a)] • $\lambda_{\theta,t,1}=\lambda_{\theta,t+1,1}$ and $\mu_{\theta,t,1} = \mu_{\theta,t+1,1}$ for all $t = 0, \ldots, T-1$$a_t = 0$ and $ \gamma_{1t} + \gamma_{2t} + \gamma_{3t} = 1$ for all $t=0,\ldots,T-1$. \end{enumerate}
assumption\quad \begin{enumerate}[\label=(a)] • $\lambda_{I,t,1} = 1$ and $\mu_{I,t,1} = 0$ for all $t = 0, \ldots, T-1$. • $\beta_{0t} = 0$ and $\beta_{1t} + \beta_{2t} = 1$ for all $t=0,\ldots,T-1$. \end{enumerate}

Assumption (ref) is usually thought of as a normalization, which is commonly imposed since log-skills are only identified up to scale and location. Here, I impose the restrictions on the first measure, which is without loss of generality. Instead of fixing the intercept and the slope coefficient in equation ((ref)) for $t=0$, one could set $\rho_{0} =0 $ and $\rho_{1}=1$ and thus “anchor” the skills at $Q$. Assumption (ref) anchors the skills at $Z_{\theta,0,1}$, but analogous issues discussed here arise with anchoring at $Q$ (see Section (ref)). Without such an assumption, the parameters are not point identified. Assumption (ref)(a) states that the skill measures are age-invariant (using the terminology of \citeN{AW:22} -- see their Definition 1 and footnote 10). Assumption (ref)(b) imposes restrictions on the technology, which \citeN{AW:22} refer to as a known scale and location assumption in a more general context. Assumption (ref)(a) says that an investment measure is age-invariant. Assumption (ref)(b) states constant return to scale in equation ((ref)), which is a strong assumption and used for point identification without age-invariant investment measures. If investment is observed (i.e. $Z_{I,t,m} = \ln I_t$), Assumption (ref)(a) is automatically satisfied.

I now characterize the identified set of the primitive parameters under Assumption (ref) only. I then discuss point identification under different combinations of Assumptions (ref)--(ref) and show that several policy relevant parameters are invariant to the restrictions in Assumption (ref) and are in fact point identified under Assumption (ref). Finally, I illustrate why Assumption (ref) is generally not a normalization for the primitive as well as some policy relevant parameters.

Identification

Define $\tilde{I}_t = \exp(\mu_{I,t,1})I_{t}^{\lambda_{I,t,1}}$ and $\tilde{\theta}_t = \exp(\mu_{\theta,t,1})\theta_{t}^{\lambda_{\theta,t,1}}$ so that $$\ln \tilde{\theta}_{t} = \mu_{\theta,t,1} + \lambda_{\theta,t,1} \ln \theta_{t} \qquad \text{ and } \qquad \ln \theta_{t} = \frac{\ln \tilde{\theta}_t - \mu_{\theta,t,1}}{\lambda_{\theta,t,1} }.$$ We can then rewrite the production function in terms of $\tilde{\theta}_t$ and $\tilde{I}_t$ because $$\frac{\ln \tilde{\theta}_{t+1} - \mu_{\theta,t+1,1}}{\lambda_{\theta,t+1,1} } = a_t + \gamma_{1t} \frac{\ln \tilde{\theta}_t - \mu_{\theta,t,1}}{\lambda_{\theta,t,1} } + \gamma_{2t} \frac{\ln \tilde{I}_t - \mu_{I,t,1}}{\lambda_{I,t,1} } + \gamma_{3t} \frac{\ln \tilde{\theta}_t - \mu_{\theta,t,1}}{\lambda_{\theta,t,1} }\frac{\ln \tilde{I}_t - \mu_{I,t,1}}{\lambda_{I,t,1} } + \eta_{\theta,t} .$$ After rearranging, we can then rewrite equations ((ref))--((ref)) as

eqnarray[eqnarray omitted — 978 chars of source]

where $$ \tilde{\gamma}_{1t} = \frac{\lambda_{\theta,t+1,1}}{\lambda_{\theta,t,1}} \left( \gamma_{1t} - \frac{\mu_{I,t,1}}{\lambda_{I,t,1}}\gamma_{3t} \right), \;\; \tilde{\gamma}_{2t} = \frac{ \lambda_{\theta,t+1,1} }{\lambda_{I,t,1}} \left( \gamma_{2t} - \frac{\mu_{\theta,t,1}}{\lambda_{\theta,t,1}}\gamma_{3t} \right), \;\; \tilde{\gamma}_{3t} = \frac{\lambda_{\theta,t+1,1}}{\lambda_{\theta,t,1} \lambda_{I,t,1}}\gamma_{3t}$$ and expressions for all other parameters in ((ref))--((ref)) are provided in Appendix (ref). Importantly, by construction, $\tilde{\mu}_{\theta,t,1} = \tilde{\mu}_{I,t,1} = 0$ and $\tilde{\lambda}_{\theta,t,1} = \tilde{\lambda}_{I,t,1} = 1$ and Lemma (ref) implies that the joint distribution of $(\{\ln \tilde{\theta}_{t}\}^T_{t=0}, \{\ln \tilde{I}_{t}\}^T_{t=0})$ is point identified conditional on $\{Y_t\}_{t=0}^{T-1}$, which then yields identification of all parameters in ((ref))--((ref)). These parameters are functions of the primitive parameters in ((ref))--((ref)). The next theorem, which characterizes the identified set, states that the identified set consists of all primitive parameters that imply the same values in ((ref))--((ref)) as the true parameters, which is very similar to the probit model.

theoremSuppose Assumption (ref) holds. \begin{enumerate} • The identified set of $\{\mu_{\theta,t,m},\lambda_{\theta,t,m}\}_{t=0,\ldots,T,m=1,2}$, $\{\mu_{I,t,m},\lambda_{I,t,m}\}_{t=0,\ldots,T-1,m=1,2}$, $(\rho_0,\rho_1)$, $\{ \beta_{0t}, \beta_{1t}, \beta_{2t}\}^{T-1}_{t=0}$, $\{a_t,\gamma_{1t},\gamma_{2t}, \gamma_{3t}\}^{T-1}_{t=0}$ consists of all vectors that yield the same values of $\{\tilde{\mu}_{\theta,t,m},\tilde{\lambda}_{\theta,t,m}\}_{t=0,\ldots,T,m=1,2}$, $\{\tilde{\mu}_{I,t,m},\tilde{\lambda}_{I,t,m}\}_{t=0,\ldots,T-1,m=1,2}$, $(\tilde{\rho}_0,\tilde{\rho}_1)$, $\{ \tilde{\beta}_{0t}, \tilde{\beta}_{1t}, \tilde{\beta}_{2t}\}^{T-1}_{t=0}$, and $\{\tilde{a}_t,\tilde{\gamma}_{1t},\tilde{\gamma}_{2t}, \tilde{\gamma}_{3t}\}^{T-1}_{t=0}$ as the true parameter vectors. • Let $\{\bar{\mu}_{\theta,t,1},\bar{\lambda}_{\theta,t,1}\}_{t=0}^{T}$, $\{\bar{\mu}_{I,t,1},\bar{\lambda}_{I,t,1}\}_{t=0}^{T-1}$, be fixed constants with $\bar{\lambda}_{\theta,t,1},\bar{\lambda}_{I,t,1}\neq0$ for all $t$. If in addition $\{{\mu}_{\theta,t,1},{\lambda}_{\theta,t,1}\}_{t=0}^{T} = \{\bar{\mu}_{\theta,t,1},\bar{\lambda}_{\theta,t,1}\}_{t=0}^{T}$ and $\{{\mu}_{I,t,1},{\lambda}_{I,t,1}\}_{t=0}^{T-1} = \{\bar{\mu}_{I,t,1},\bar{\lambda}_{I,t,1}\}_{t=0}^{T-1}$, then the identified set is a singleton. \end{enumerate}

Part 2 of the theorem shows that the parameters are indeed not point identified under Assumption (ref) only and that the sources of underidentification are the ambiguous scales and locations of skills and investments. For example, without additional assumptions, equations ((ref))--((ref)) and ((ref))--((ref)) are observationally equivalent, and we cannot distinguish between the skills $\theta_t$ and $\tilde{\theta}_t$ and the corresponding production functions. Even if investment was observed (and $\mu_{I,t,1} = 0$ for all $t$), we can then only identify $(\lambda_{\theta,t+1,m}/\lambda_{\theta,t,m})\gamma_{1t}$, but not $\lambda_{\theta,t,m}$ and $\gamma_{1t}$ separately. Hence, we cannot distinguish between changes in the quality of the measurements ($\lambda_{\theta,t+1,m}/\lambda_{\theta,t,m}$) and changes in the technology (${\gamma}_{1t}$). For example, suppose $Z_{\theta,t,m}$ are test scores. We then cannot distinguish between all children getting smarter or tests becoming easier. Similarly, we can at best identify $\gamma_{2t}$ up to scale, even with observed investment.

The following corollary states that all parameters are point identified under additional assumptions. These results are an extension of those in \citeN{AW:22}, who assume that investment is exogenous (in the sense that it is uncorrelated with $ \eta_{\theta,t}$).

corollarySuppose Assumptions (ref) and (ref) hold. Suppose either Assumption (ref)(a) or Assumption (ref)(b) holds. Suppose either Assumption (ref)(a) or Assumption (ref)(b) holds. Then all parameters are point identified.

The corollary also immediately implies that Assumptions (ref)(a) and (ref)(b) together impose additional testable restrictions, which is one of the main contributions of Agostinelli and Wiswall (2016a, 2016b, 2024). Contrarily, as shown in Theorem (ref) in the appendix and illustrated in examples below, if the model is correctly specified and Assumption (ref) holds, then there always exist sets of parameters which are consistent with the data and satisfy Assumptions (ref), (ref), either (ref)(a) or (ref)(b), and either (ref)(a) or (ref)(b). These different sets of assumptions therefore impose no additional restrictions on the distribution of observables. While different sets of assumptions yield point identification and are observationally equivalent, the estimated primitive parameters are usually quite different, as illustrated in Section (ref).

Invariant parameters

I now show that many interesting features are point identified under Assumption (ref) only because they can all be rewritten in terms of the identified parameters in ((ref))--((ref)), as in Section (ref). They include summaries of the productions functions and effects of investment and income on skills and adult outcomes. These features do not constitute an exhaustive list and there may be many others. To state the formal results, recall that $Q_{\alpha}(\theta_t)$ is the $\alpha$ quantile of the skill distribution at time $t$ and $F_{\ln(\theta_{t+1})}(\cdot)$ is the cdf of log-skills at time $t+1$. Define $s_{1t}(\alpha_1,\alpha_2,\alpha_3) = a_t + \gamma_{1t} \ln Q_{\alpha_1}(\theta_t) + \gamma_{2t} \ln Q_{\alpha_2}(I_t) + \gamma_{3t} \ln Q_{\alpha_1}(\theta_t) \ln Q_{\alpha_2}(I_t) + Q_{\alpha_3}(\eta_{\theta,t}) $ which are the log-skills in period $t+1$ for specific quantiles of the inputs in period $t$. Similarly, let $s_{2t}(\alpha_1,\alpha_2,\alpha_3,y) = a_t + \gamma_{1t} \ln Q_{\alpha_1}(\theta_t) + \gamma_{2t} \ln I_t(y) + \gamma_{3t} \ln Q_{\alpha_1}(\theta_t) \ln I_t(y) + Q_{\alpha_2}(\eta_{\theta,t})$, where $\ln I_t(y) = \beta_{0t} + \beta_{1t} \ln Q_{\alpha_1}(\theta_t) + \beta_{2t} \ln y + Q_{\alpha_3}(\eta_{I,t}) $, which are log-skills and log-investment, that depend on a specific exogenously set value of $Y$.

theoremSuppose Assumption (ref) holds. \begin{enumerate} • $F_{\ln \theta_{t+1}}(s(\alpha_1,\alpha_2,\alpha_3) )$ and $\mu_{\theta,t+1,m} + \lambda_{\theta,t+1,m}s_{1t}(\alpha_1,\alpha_2,\alpha_3) + Q_{\alpha_4}(\varepsilon_{\theta,t+1,m})$ are point identified for all $\{\alpha_j\}^4_{j=1}\in (0,1)$. $\frac{\partial \lambda_{\theta,t+1 ,m} \ln \theta_{t+1}}{\partial \lambda_{\theta,t ,m'} \ln \theta_t } \mid_{ I_t = Q_{\alpha_2}(I_t) }$ and $\frac{\partial \lambda_{\theta,t+1 ,m} \ln \theta_{t+1}}{\partial \lambda_{I,t ,m'} \ln I_t } \mid_{ \theta_t = Q_{\alpha_1}(\theta_t) }$ are point identified for all $m,m'$. • $F_{\ln \theta_{t+1}}(s_{2t}(\alpha_1,\alpha_2,\alpha_3,y) )$ and $\mu_{\theta,t+1,m} + \lambda_{\theta,t+1,m}s_{2t}(\alpha_1,\alpha_2,\alpha_3,y) + Q_{\alpha_4}(\varepsilon_{\theta,t+1,m})$ are point identified for all $\{\alpha_j\}^4_{j=1}\in (0,1)$. • $ P\left(Q \leq q \mid \theta_s= Q_{\alpha_1}(\theta_s), \{I_t = Q_{\alpha_{2t}}(I_t) \}^{T-1}_{t=0}, \{\eta_{\theta,t} = Q_{\alpha_{3t}}(\eta_{\theta,t}) \}^{T-1}_{t=s} \right)$ is point identified for all $\alpha_1,\{\alpha_{2t},\alpha_{3t} \}^{T-1}_{t=s} \in (0,1)$. • $ P\left(Q \leq q \mid \theta_s= Q_{\alpha}(\theta_s), \{Y_t = y_t \}^{T-1}_{t=s} \right)$ is point identified for all $\alpha \in (0,1)$. • Suppose Assumption (ref)(a) also holds. Then $\gamma_{1t} + \gamma_{3t}\ln Q_{\alpha}(I_t)$ is point identified for all $\alpha$ and $P(\gamma_{1t} + \gamma_{3t}\ln I_t \leq q)$ is point identified for all $q \in \ensuremath{\mathbb{R}}$. \end{enumerate}

The function $F_{\ln \theta_{t+1}}(a_t + \gamma_{1t} \ln Q_{\alpha_1}(\theta_t) + \gamma_{2t} \ln Q_{\alpha_2}(I_t) + \gamma_{3t} \ln Q_{\alpha_1}(\theta_t) \ln Q_{\alpha_2}(I_t) + Q_{\alpha_3}(\eta_{\theta,t}) )$ measures how changes in skills and investment changes the relative standing in the skill distribution. For example, consider an individual with $\theta_t = Q_{0.1}(\theta_t)$, meaning that the person is at lowest $10\%$ of the skill distribution at time $t$. Then, given investment $I_t = Q_{0.25}(I_t)$ and a median production function shock, $ \eta_{\theta,t} = Q_{0.5}(\eta_{\theta,t})$, $F_{\ln \theta_{t+1}}(s(0.1,0.25,0.5))$ tells us the relative rank (or the quantile) in the skill distribution at time $t+1$. We can then for example vary the investment quantile and analyze how future skill ranks are affected. This feature is identified because I show in the appendix that

eqnarray*[eqnarray* omitted — 597 chars of source]

Thus, one can estimate the model based on ((ref))--((ref)) and calculate the feature using $\tilde{\theta}_{t}$ instead of ${\theta}_{t}$. The estimand then corresponds to the true effect, even if Assumptions (ref)--(ref) do not hold.

Once we know the rank at time $t+1$ and fix investment and production shock quantiles in that period, we can identify the skill rank at time $t+2$. Using recursive arguments, we can identify the relative rank in period $T$, given investment and production shock quantiles in all period and a skill quantile in period $0$. We could then make statements such as: “A person at lowest $10\%$ of the skill distribution in period $0$ would end up at the $30\%$ quantile of the original skill distribution in the final period with a particular investment strategy and median production function shocks.” These statements would allow comparisons of investment strategies, assessing heterogeneous effects, and choosing optimal investments.

Instead of considering ranks of skills, one can interpret the effects in the units of any of the measures. For example, changing investment from $Q_{\alpha_2}(I_t)$ to $Q_{\alpha_2'}(I_t)$ changes skills at time $t+1$ such that skill measure $m$ changes by $\lambda_{\theta,t+1,m}(s_1(\alpha_1,\alpha_2',\alpha_3) - s_1(\alpha_1,\alpha_2,\alpha_3))$ (holding everything else equal). We can also identify that a change in investment, which corresponds to a 1 unit increase in investment measure $m'$ at time $t$, affects skills at time $t+1$ in a way that skill measure $m$ changes by $\frac{\partial \lambda_{\theta,t+1 ,m} \ln \theta_{t+1}}{\partial \lambda_{I,t ,m'} \ln I_t } \mid_{ \theta_t = Q_{\alpha_1}(\theta_t) }$ (see Section (ref) for specific examples).

The second part is very similar, but it also considers exogenous changes in income.

Instead of considering the skill rank in the final period, we could also look at the distribution of the adult outcome $Q$ (or, alternatively, one of the skill measures in the final period). Notice that it is important to condition on the production function shocks, because investment in endogenous. We can either fix a quantile or average over its marginal distribution since we obtain identification for all quantiles. The fourth part focuses on the effect of income on skills. Again, we can point identify averages and features of the distribution such as $\int E\left(Q \mid \theta_s = \theta, \{Y_t = y_t \}^{T-1}_{t=s} \right) f_{\theta_s}(\theta) d\theta$ (which differs from $E\left(Q \mid \{Y_t = y_t \}^{T-1}_{t=s} \right)$ due to a potential dependence between $ \{Y_t = y_t \}^{T-1}_{t=s} $ and skills). Also notice that

align*[align* omitted — 258 chars of source]

Thus, we can identify the sequence $\{Y_t = y_t \}^{T-1}_{t=s}$ that maximizes the conditional expected value of $\ln \theta_T$. For this sequence, we do not necessarily need to observe an adult outcome because we can instead use one of the skill measures in period $T$. To identify these features, one only has to identify the joint distribution of $(Q,Y_1, \ldots, Y_{T-1},\tilde{\theta}_s)$. For example, when $s = 0$, we do not require any skill measures in periods $1,2,\ldots,T$.\footnote{Part 4 of Theorem (ref) also implies identification of $$\int\int\left( \frac{\partial}{\partial y_s} E(Q \mid \theta_s = \theta, \{Y_t = y_t \}^{T-1}_{t=s})\right)f_{\theta_s,Y_s, \ldots, Y_{T-1}}(\theta,y_s, \ldots, y_{T-1}) d\theta dy_s \cdots d y_{T-1} $$ which \shortciteN{DKP:20} refer to as “anchored treatment effects”. \shortciteN{DKP:20} show that these effects are identified without age-invariant measures. My results show that one in fact does not even need any measures of investments or measures of skills in periods $s+1,2,\ldots, T$ (and therefore also no assumptions on the measurement systems, including age-invariance and independence). In addition, my arguments are not specific to the trans-log production function and carry over to other settings.} Importantly objects such as $ \int E\left(\theta_T \mid \theta_s = \theta, \{Y_t = y_t \}^{T-1}_{t=s} \right) f_{\theta_s}(\theta) d\theta$ are not point identified without the scale and location restrictions and are sensitive to the specific values used (see Example (ref) below).

remarkTo summarize the production technology, \citeN{DKP:20} show that standardizing skills can lead to features that are invariant to scale and location restrictions and age-invariance. In particular, they show identification of the distribution of $\left(\frac{\partial \ln \theta_{t+1} }{\partial \ln I_t}\right)/ std(\ln \theta_{t+1}) $. One can then make statements such as “increasing investment by 1%, increases log-skills by $x \times std(\ln \theta_{t+1})$”. Part 1 of Theorem (ref) offers an alternative interpretation in terms of ranks or units of the measures.
remarkInstead of using a two-step approach, where the distribution of a linear combination of skills and investment is identified first, \citeN{AW:22} substitute measures into the production function equation and use IV arguments with exogenous investment (i.e. $I_t$ and $\eta_{\theta,t}$ are independent in ((ref))). Aside from exogenous investment, there are no substantial differences between the required assumptions, as both approaches are based on the joint distribution of the measures. My main contributions are to study the roles of the scale and location restrictions on the parameters of the model, to show which restrictions select an element of the identified set, and to provide features that are invariant to them and are identified without age-invariant measures and restrictions on the production function.

Non-invariant parameters

As shown above, under Assumption (ref), there exist different sets of observationally equivalent parameters that cannot be distinguished using the data. I now illustrate with two examples that the resulting primitive parameters and counterfactuals might differ considerably.

exampleFor simplicity, I assume that investment is observed and exogenous. In this case, Assumption (ref)(a) holds. I first consider a data generating process (DGP) satisfying Assumptions (ref), (ref)(a), and (ref)(b), but not Assumption (ref). I then construct two alternative sets of parameters, both of which are observationally equivalent to the original DGP. One of these sets of parameters satisfies Assumptions (ref), (ref), and (ref)(a) and the other set satisfies Assumptions (ref), (ref), and (ref)(b). Specifically, first assume that \begin{eqnarray*} \ln \theta_{t+1} &=& \frac{1}{2} \ln \theta_{t} + \frac{1}{2} \ln I_t + \eta_{\theta,t} \\ \tilde{Z}_{\theta,t,1} &=& \ln \theta_{t} + \tilde{\varepsilon}_{\theta,t,1} \end{eqnarray*} Here, I set $a_t = \mu_{\theta,t,m} = \rho_{0} = 0$ and focus on the scale restriction in Assumption (ref) only. For brevity, I omit equations for investment, additional measurements, and the adult outcome. Measures often do not have a natural scale. For example, we could divide all test scores by its standard deviation or we could measure education in months rather than years. One would then hope that changing the units of measurement does not affect the economic interpretation of the results. As a specific example, suppose we estimate the model using a scaled version of the measures, namely $Z_{\theta,t,1}= 12 \tilde{Z}_{\theta,t,1}$. Then for all $t$ \begin{eqnarray*} \ln \theta_{t+1} &=& \frac{1}{2} \ln \theta_{t} + \frac{1}{2} \ln I_t + \eta_{\theta,t} \\ Z_{\theta,t,1} &=& 12 \ln \theta_{t} + {\varepsilon}_{\theta,t,1} \end{eqnarray*} where ${\varepsilon}_{\theta,t,1} = 12 \tilde{\varepsilon}_{\theta,t,1} $. Without knowing the true DGP, one would typically impose assumptions that yield point identification when estimating the model. By Corollary (ref), one could use either Assumption (ref)(a) or (ref)(b) next to Assumptions (ref) and (ref). I first construct parameters satisfying Assumptions (ref), (ref), and (ref)(b) using Theorem (ref) in the appendix, which implies that there are $\{\bar{\theta}_{t}\}^T_{t=0}$ such that \begin{eqnarray*} \ln \bar{\theta}_{t+1} &=& \bar{\gamma}_{1t} \ln \bar{\theta}_{t} + (1- \bar{\gamma}_{1t}) \ln I_t + \bar{\eta}_{\theta,t} \\ Z_{\theta,t,1} &=& \bar{\lambda}_{\theta,t,1} \ln \bar{\theta}_{t} + \varepsilon_{\theta,t,1} \end{eqnarray*} where \begin{eqnarray*} (\bar{\gamma}_{10},\bar{\gamma}_{11},\bar{\gamma}_{12},\bar{\gamma}_{13},\bar{\gamma}_{14}, \ldots ) &=& (0.077, 0.351, 0.435, 0.470, 0.485,\ldots) \\ (\bar{\lambda}_{\theta,0,1}, \bar{\lambda}_{\theta,1,1},\bar{\lambda}_{\theta,2,1},\bar{\lambda}_{\theta,3,1},\bar{\lambda}_{\theta,4,1}, \ldots ) &=& (1, 6.5, 9.25, 10.625, 11.3125, \ldots ) \end{eqnarray*} Moreover $\bar{\gamma}_{1t} \rightarrow {\gamma}_{1t} = 1/2$ and $\bar{\lambda}_{\theta,t,1}\rightarrow {\lambda}_{\theta,t,1} = 12$ as $t \rightarrow \infty$. Imposing the restriction $\bar{\lambda}_{\theta,0,1}= 1$ means that we estimate a model with alternative skills $\ln \bar{\theta}_{0} = 12 \ln \theta_{0}$ and that we obtain different parameters and skill distributions. Although these two models suggest very different dynamics, they generate identical measurements. Since the true DGP is unknown, the primitive parameters are therefore hard to interpret. For example the coefficient in front of investment in the original model might be interpreted as: “increasing investment by 1%, increases skills in the next period by 0.5% and the effect is the same for all time periods” (see e.g. \citeN{AW:22} for these interpretations). Contrarily, one might interpret the coefficients in the alternative model with Assumption (ref) as: “increasing investment by $1\%$ in period $0$, increases skills in period $1$ by $0.923\%$ and increasing investment by $1\%$ in period $4$, increases skills in period $5$ by $0.515\%$”, suggesting that investment is more beneficial in early periods. To obtain interpretable parameters, notice that a one unit increase in $\ln(I_t)$ leads to an increase in skills that corresponds to a $\lambda_{\theta,t+1,1}\gamma_{2t} = 6$ units increase in $Z_{\theta,t+1,1}$ or a $1/2$ unit increase in $\tilde{Z}_{\theta,t+1,1}$. This effect does not depend on the set of restrictions and its interpretation adapts to the units of measurement. The consequences of imposing Assumptions (ref), (ref) and (ref)(a) are discussed in Section (ref), which shows that the production function becomes $ \ln \tilde{\theta}_{t+1} = \frac{1}{2} \ln \tilde{\theta}_{t} + 6 \ln I_t + \tilde{\eta}_{\theta,t}$.
exampleI now use a more involved DGP from Section (ref), which is based on the simulations of \shortciteN{AMN:19}.\footnote{In Section (ref) I use a CES production function, as do \shortciteN{AMN:19}, but for this numerical examples I use a trans-log production function that leads to similar observed data.} In this setup $T=2$ and Assumptions (ref), (ref), (ref)(a), and (ref)(a) hold. Importantly, one skill and one investment measure have loadings of $1$. Now suppose $Y_t$ represents income, we can exogenously increase the sum of income of each individual in periods $0$ and $1$ by two standard deviations, and we want to distribute that income optimally across the two periods. Denote the skills in period $2$ as a function of income by $\theta_2(Y_0 + w x, Y_1 + (1-w)x)$, where $Y_0$ and $Y_1$ are the original incomes, $x$ is the additional amount to be distributed, and $w$ is the share invested in period $0$. The left panel of Figure (ref) shows $$\frac{E[\theta_2(Y_0 + w x, Y_1 + (1-w)x) ] - E[\theta_2(Y_0, Y_1)]}{sd(\theta_2(Y_0, Y_1))}$$ as a function of $w$. That is, the $y$-axis shows the increase in the mean measured in standard deviations. The black line shows the effect using the true parameters, leading to an optimal income share of around $32\%$ in period $0$. Next, I multiply all skill measures $Z_{\theta,t,1}$ by $s_{\theta}$ and reestimate the model. Scaling the measures has identical effects to imposing $\lambda_{\theta,0,1} = 1$, when the data is generated with a loading of $s_{\theta}$. The implied effects of increasing income for $s_{\theta} = 2/3$ and $s_{\theta} = 2$ can be seen in the left panel of Figure (ref) as well. Depending on the scale, we obtain inefficient optimal investment choices or erroneous benefits of investment. Intuitively, the reason is that we now maximize $E[\tilde{\theta}_t(Y_0 + w x, Y_1 + (1-w)x)] = E[\theta_t(Y_0 + w x, Y_1 + (1-w)x)^{s_{\theta}}]$ which is not invariant to $s_{\theta}$. While such counterfactuals could be relevant if a welfare function is a function of the skill level, they are not identified without fixing the scales and locations and then depend on the units of measurements of the data. \begin{figure}[t!] \begin{center} \caption{Mean response for different weights} \end{center} \newgeometry{textwidth=16.4cm} { \begin{singlespace} Notes: This figure shows changes in standardized means of the level (left panel) and the log (right panel) of skills for different income transfers and different scales of log-skills. \end{singlespace}} \end{figure} The right panel contains results for the log-skills rather than the level. As discussed below Theorem (ref), in this case, the optimal income sequence does not depend on the scale of the measures, because the production function is sufficiently flexible in log-skills. In addition, dividing by the standard deviation, implies that the objective is invariant to the scale.

Anchoring

Consider equations ((ref))--((ref)) and recall that $\tilde{\mu}_{\theta,t,0} = \tilde{\mu}_{I,t,0} = 0$ and $\tilde{\lambda}_{\theta,t,0} = \tilde{\lambda}_{I,t,0} = 1$, and the parameters in this system of equations are point identified by Theorem (ref). Now define $\tilde{\vartheta}_{t}$ such that $\ln \tilde{\vartheta}_{t} = \tilde{\rho}_0 + \tilde{\rho}_1 \ln \tilde{\theta}_{t}$. We then get

eqnarray*[eqnarray* omitted — 938 chars of source]

\citeN{CH:08} estimate a model for $\ln \tilde{\vartheta}_{t}$ which anchors the skills at $Q$. Their identification strategy is equivalent to using Assumptions (ref) and (ref)(a) and imposing $\rho_{0} = 0$ and $\rho_1 = 1$ instead of Assumption (ref). Anchoring can help with the interpretation of certain parameters of the model. For example, when $E(Q \mid \ln \tilde{\vartheta}_{T} ) = \ln \tilde{\vartheta}_{T} $, then an increase of $\ln \tilde{\vartheta}_{T}$ by one corresponds to a one unit increase in the conditional expectation of $Q$. However, as illustrated in Example (ref), investment or income sequences that maximize the expected levels of skills depend on the units of measurements of the anchor.\footnote{The identification arguments are fundamentally different if the anchoring equations was in levels instead of logs of the skills. That is, if $ Q = {\rho}_{0} + {\rho}_{1} {\theta}_{T} + \eta_Q = {\rho}_{0} + \frac{{\rho}_{1} }{\exp(\mu_{\theta,T,1})^{1/\lambda_{\theta,T,1}} }\tilde{\theta}_T^{1/\lambda_{\theta,T,1}} + \eta_Q $. In this case it can be shown that the joint distribution of $(Q,\tilde{ \theta}_T)$ is identified, which identifies $\lambda_{\theta,T,1}$. Thus, the distribution of $ \theta_T$ is identified up to a scaling factors, which implies that even sequences of income or investment that maximize the level of skills are identified.}

As discussed by \citeN{CH:08}, $\tilde{\gamma}_{1t}$ is invariant to the anchor. However, without Assumption (ref)(a), $\tilde{\gamma}_{1t}$ is typically not equal to $\gamma_{1t}$. The coefficient in front of investment is hard to interpreted because it depends on the specific anchor and its units of measurements (as also noted by \citeN{CH:08}). Moreover, skills can be anchored at the adult outcome, but not investment, and many of the estimated parameters also depend on the units of the investment measures. While anchoring makes most sense under Assumption (ref)(a), in which case the skills in all periods are in the units of $Q$, we could instead use Assumption (ref)(b) to achieve point identification. In this case different adult outcomes or different units of measurements lead to different coefficients that possibly suggest very different dynamics, just as in Example (ref).

CES production function

I now discuss the CES production technology, where

eqnarray[eqnarray omitted — 759 chars of source]

where $\sigma_t \neq 0$, $\gamma_{1t}\neq 0$, and $\gamma_{2t}\neq 0$ for all $t$. The measurement system is linear in $\ln \theta_{t}$ (as in \shortciteN{CHS:10} or \shortciteN{AMN:19}) to ensure that estimated skills are positive.

Identification

Similar to the trans-log case, define $\tilde{\theta}_t = \exp(\mu_{\theta,t,1})\theta_{t}^{\lambda_{\theta,t,1}}$ so that $$\theta_{t} = \exp \left(-\frac{\mu_{\theta,t,1} }{\lambda_{\theta,t,1}}\right)\tilde{\theta}_t^{\frac{1}{\lambda_{\theta,t,1}}}.$$ Similarly, let $\tilde{I}_t = \exp(\mu_{I,t,1})I_{t}^{\lambda_{I,t,1}}$. We can again rewrite the production technology in terms of $\tilde{\theta}_t$ and $\tilde{I}_t$. That is, we can rewrite equations ((ref))--((ref)) to

eqnarray[eqnarray omitted — 1,073 chars of source]

where $$ \tilde{\gamma}_{1t} = \gamma_{1t} \exp\left( \sigma_{t} \left( \frac{ \mu_{\theta,t+1,1} }{ \lambda_{\theta,t+1,1}\psi_t } - \frac{\mu_{\theta,t,1}}{\lambda_{\theta,t,1}} \right) \right) \quad \text{and} \quad \tilde{\gamma}_{2t} = \gamma_{2t} \exp\left( \sigma_{t} \left( \frac{ \mu_{\theta,t+1,1} }{ \lambda_{\theta,t+1,1}\psi_t } - \frac{\mu_{I,t,1}}{\lambda_{I,t,1}} \right) \right) $$ and the other parameters are defined as in the trans-log case. Using the identified joint distribution of $(\{\ln \tilde{\theta}_{t}\}^T_{t=0}, \{\ln \tilde{I}_{t}\}^T_{t=0})$ and the restrictions imposed by the production function, the following theorem characterizes the identified set of the finite dimensional parameters.

theoremSuppose Assumption (ref) holds. \begin{enumerate}[(a)] • The identified set of $\{\mu_{\theta,t,m},\lambda_{\theta,t,m}\}_{t=0,\ldots,T,m=1,2}$, $\{\mu_{I,t,m},\lambda_{I,t,m}\}_{t=0,\ldots,T-1,m=1,2}$, $(\rho_0,\rho_1)$, $\{ \beta_{0t}, \beta_{1t}, \beta_{2t}\}^{T-1}_{t=0}$, $\{ \gamma_{1t},\gamma_{2t}, \sigma_{t}, \psi_t\}^{T-1}_{t=0}$ consists of all vectors that yield the same values of $\{\tilde{\mu}_{\theta,t,m},\tilde{\lambda}_{\theta,t,m}\}_{t=0,\ldots,T,m=1,2}$, $\{\tilde{\mu}_{I,t,m},\tilde{\lambda}_{I,t,m}\}_{t=0,\ldots,T-1,m=1,2}$, $(\tilde{\rho}_0,\tilde{\rho}_1)$, $\{ \tilde{\gamma}_{1t},\tilde{\gamma}_{2t}, \frac{\sigma_{t} }{\lambda_{\theta,t,1}},\frac{\sigma_{t} }{\lambda_{I,t,1}} \}^{T-1}_{t=0}$, $\{\frac{\sigma_{t} \psi_t }{\lambda_{\theta,t+1,1}}\}^{T-1}_{t=0}$, and $\{ \tilde{\beta}_{0t}, \tilde{\beta}_{1t}, \tilde{\beta}_{2t}\}^{T-1}_{t=0}$ as the true parameter vectors. • Let $\{\bar{\mu}_{\theta,t,1} \}_{t=0}^{T}$, $\{\bar{\mu}_{I,t,1} \}_{t=0}^{T-1}$, and $\{\bar{\lambda}_{\theta,t,1} \}_{t=0}^{T}$ be fixed constants with $\bar{\lambda}_{\theta,t,1} \neq 0$ for all $t$. Under the additional restrictions $\{{\mu}_{\theta,t,1} \}_{t=0}^{T} = \{\bar{\mu}_{\theta,t,1} \}_{t=0}^{T}$ and $\{{\mu}_{I,t,1} \}_{t=0}^{T-1} = \{\bar{\mu}_{I,t,1} \}_{t=0}^{T-1}$ and $\{{\lambda}_{\theta,t,1} \}_{t=0}^{T} = \{\bar{\lambda}_{\theta,t,1} \}_{t=0}^{T}$ the identified set is a singleton. \end{enumerate}

An important implications of part (a) is that the fraction $\frac{\lambda_{\theta,t,1}}{\lambda_{I,t,1}} = \frac{\lambda_{\theta,t,1}}{\sigma_t}\frac{\sigma_t}{\lambda_{I,t,1}}$ is point identified for all $t = 1,\ldots, T-1$. Hence, if we restrict $\lambda_{\theta,t,1}$ to a constant, $\lambda_{I,t,1}$ is identified, which is very different to the trans-log case. Intuitively, since skills and investment have the same exponent in the CES production functions, the relative scale is identified by the functional form restrictions.

As before, I now introduce additional assumptions to achieve point identification.

\setcounter{assumptionp}{1}

assumptionp$\lambda_{\theta,0,1}=1$ and $\mu_{\theta,0,1} = 0$.
assumptionp\quad \begin{enumerate}[\label=(a)] • $\mu_{\theta,t,1} = \mu_{\theta,t+1,1}$ for all $t = 0, \ldots, T-1$$\gamma_{1t} + \gamma_{2t} =1 $ for all $t=0,\ldots,T-1$. \end{enumerate}
assumptionp\quad \begin{enumerate}[\label=(a)] • $\mu_{I,t,1} = \mu_{I,t+1,1} = 0$ for all $t = 0, \ldots, T-2$. • $\beta_{0t} = 0$ for all $t=0,\ldots,T-1$. \end{enumerate}
assumptionp\quad \begin{enumerate}[\label=(a)] • $\lambda_{\theta,t,1} = \lambda_{\theta,t+1,1}$ for all $t = 0, \ldots, T-1$. • $\psi_{t} = 1$ for all $t=0,\ldots,T-1$. \end{enumerate}

The following result shows how point identification can be established.

corollarySuppose Assumptions (ref) and (ref) hold. Suppose either Assumption (ref)(a) or Assumption (ref)(b) holds. Suppose either Assumption (ref)(a) or Assumption (ref)(b) holds. Suppose either Assumption (ref)(a) or Assumption (ref)(b) holds. Then all parameters are point identified.

A common restriction with the CES production function is to set $\psi_{t} = 1$ for all $t$ (see e.g. \shortciteN{CHS:10} and \shortciteN{AMN:19}). In this case, Theorem (ref) implies that $\frac{\lambda_{\theta,t+1,1}}{\lambda_{\theta,t,1}}$ is point identified. Hence, assuming age-invariance (i.e. $\lambda_{\theta,t+1,1} = \lambda_{\theta,t,1}$) is not required, and it is in fact testable. Moreover, with $\psi_{t} = 1$, Corollary (ref) implies that the only scale restriction needed is $\lambda_{\theta,0,1} = 1$ (or alternatively $\lambda_{I,0,1}=1$). Nevertheless, it is common practice to set ${\lambda}_{\theta,t,1}= {\lambda}_{I,t,1} = 1$ for all $t$, which are not normalizations, but restrictive assumptions (even if $\psi_t \neq 1$). Setting the scales to different numbers or changing the units of measurement of the data then affects all estimated parameters, optimal investment sequences, and other counterfactuals. Similarly, if the scale of investment is fixed, we cannot anchor the skills at the adult outcome, unless $\rho_1 = 1$, which can only be true for one specific unit of measurements (see Appendix (ref)). The exact consequences of imposing unnecessary restrictions depend on the estimation methods used, and I discuss specific examples in Sections (ref) and (ref).

As shown in Theorem (ref) in the appendix and illustrated in examples below, if the model is correctly specified and Assumption (ref) holds, then there always exist sets of parameters which are consistent with the data and satisfy Assumptions (ref), (ref), either (ref)(a) or (ref)(b), either (ref)(a) or (ref)(b), and either (ref)(a) or (ref)(b). Similar to the trans-log case, different sets of assumptions yield observationally equivalent models with potentially very different primitive parameters.

Invariant parameters

As in the trans-log case, important policy relevant parameters are point identified under Assumption (ref) only, because they can all be written in terms of identified features. Similar to before, now define $s_{1t}(\alpha_1,\alpha_2,\alpha_3) = \left(\gamma_{1t} Q_{\alpha_1}(\theta_t)^{\sigma_t} + \gamma_{2t} Q_{\alpha_2}(I_t)^{\sigma_t} \right)^{\frac{\psi_t}{\sigma_t}}\exp(Q_{\alpha_3}(\eta_{\theta,t})) $ and $s_{2t}(\alpha_1,\alpha_2,\alpha_3,y) = \left(\gamma_{1t} Q_{\alpha_1}(\theta_t)^{\sigma_t} + \gamma_{2t} I_t(y)^{\sigma_t} \right)^{\frac{\psi_t}{\sigma_t}}\exp(Q_{\alpha_2}(\eta_{\theta,t}))$ where $\ln I_t(y) = \beta_{0t} + \beta_{1t} \ln Q_{\alpha_1}(\theta_t) + \beta_{2t} \ln y + Q_{\alpha_3}(\eta_{I,t}) $.

theoremSuppose Assumption (ref) holds. \begin{enumerate} • $F_{\theta_{t+1}}\big(s_{1t}(\alpha_1,\alpha_2,\alpha_3) \big)$ and $\mu_{\theta,t+1,m} + \lambda_{\theta,t+1,m}\ln s_{1t}(\alpha_1,\alpha_2,\alpha_3) + Q_{\alpha_4}(\varepsilon_{\theta,t+1,m})$ are point identified for all $\{\alpha_j\}^4_{j=1}\in (0,1)$. Moreover, $\frac{\partial \lambda_{\theta,t+1 ,m} \ln \theta_{t+1}}{\partial \lambda_{\theta,t ,m'} \ln \theta_t } \mid_{\theta_t = Q_{\alpha_1}(\theta_t), I_t = Q_{\alpha_2}(I_t) }$ and $\frac{\partial \lambda_{\theta,t+1 ,m} \ln \theta_{t+1}}{\partial \lambda_{I,t ,m'} \ln I_t } \mid_{ \theta_t = Q_{\alpha_1}(\theta_t),I_t = Q_{\alpha_2}(I_t) }$ are point identified for all $m,m'$ and $\{\alpha_j\}^2_{j=1}\in (0,1)$. • $F_{ \theta_{t+1}}(s_{2t}(\alpha_1,\alpha_2,\alpha_3,y) )$ and $\mu_{\theta,t+1,m} + \lambda_{\theta,t+1,m} \ln s_{2t}(\alpha_1,\alpha_2,\alpha_3,y) + Q_{\alpha_4}(\varepsilon_{\theta,t+1,m})$ are point identified $\{\alpha_j\}^4_{j=1}\in (0,1)$. • $ P\left(Q \leq q \mid \theta_t= Q_{\alpha_1}(\theta_t), \{I_s = Q_{\alpha_{2s}}(I_s) \}^{T-1}_{s=t}, \{\eta_{\theta,s} = Q_{\alpha_{3s}}(\eta_{\theta,s}) \}^{T-1}_{s=t} \right)$ is point identified for all $\alpha_1,\{\alpha_{2s},\alpha_{3s} \}^{T-1}_{s=t} \in (0,1)$. • $ P\left(Q \leq q \mid \theta_0= Q_{\alpha}(\theta_0), \{Y_t = y_t \}^{T-1}_{t=0} \right)$ is point identified for all $\alpha \in (0,1)$. • If in addition either Assumption (ref)(a) or Assumption (ref)(b) holds, the distributions of $$\frac{\partial \ln \theta_{t+1} }{\partial \ln \theta_{t} } = \frac{\partial }{\partial \ln \theta_{t} } \ln \left(\gamma_{1t} \theta_{t}^{\sigma_t} + \gamma_{2t} I_t^{\sigma_t}\right)^{\frac{1}{\sigma_t}} \; \text{ and } \; \; \frac{\partial \ln \theta_{t+1} }{\partial \ln I_{t} } = \frac{\partial }{\partial \ln I_{t} } \ln \left(\gamma_{1t} \theta_{t}^{\sigma_t} + \gamma_{2t} I_t^{\sigma_t}\right)^{\frac{1}{\sigma_t}} $$ are point identified and $$\left.\frac{\partial \ln \theta_{t+1} }{\partial \ln \theta_{t} } \right|_{ \ln \theta_{t} = Q_{\alpha_1}(\ln \theta_{t}), \ln I_{t} = Q_{\alpha_2}(\ln I_{t})} = \left.\frac{\partial }{\partial \ln \theta_{t} } \ln \left(\gamma_{1t} \theta_{t}^{\sigma_t} + \gamma_{2t} I_t^{\sigma_t}\right)^{\frac{1}{\sigma_t}} \right|_{ \ln \theta_{t} = Q_{\alpha_1}(\ln \theta_{t}), \ln I_{t} = Q_{\alpha_2}(\ln I_{t})} $$ and $$\left.\frac{\partial \ln \theta_{t+1} }{\partial \ln I_{t} } \right|_{ \ln \theta_{t} = Q_{\alpha_1}(\ln \theta_{t}), \ln I_{t} = Q_{\alpha_2}(\ln I_{t})} = \left.\frac{\partial }{\partial \ln I_{t} } \ln \left(\gamma_{1t} \theta_{t}^{\sigma_t} + \gamma_{2t} I_t^{\sigma_t}\right)^{\frac{1}{\sigma_t}} \right|_{ \ln \theta_{t} = Q_{\alpha_1}(\ln \theta_{t}), \ln I_{t} = Q_{\alpha_2}(\ln I_{t})} $$ are point identified for all $\alpha_1, \alpha_2 \in (0,1)$. \end{enumerate}

The features in parts (1)--(4) are analogous to those in the trans-log case and can be used to calculate optimal investment/income strategies and anchored treatment effects, as discussed after Theorem (ref). As opposed to the trans-log case, now the relative scales of skills and investment are point identified. Consequently, elasticities are identified under Assumption (ref) and either age-invariant skill measures or $\psi_t = 1$ only.

Non-invariant parameters

To achieve point identification of all parameters, we need to fix the levels of the logs of skills and investment, e.g. by setting $\mu_{\theta,0,1}$ and $\mu_{0,I,1}$ to $0$. Similar to the trans-log case, these restrictions are generally not normalizations for the primitive parameters. In the example below, I illustrate that with the restriction $\gamma_{1t} + \gamma_{2t} =1$, setting $\mu_{\theta,0,1} = 0$ can imply different dynamics. Here $\mu_{\theta,0,1}$ (and not $\lambda_{\theta,0,1}$) affects the scale because we can identify the distribution of $\tilde{\theta}_{t} = \exp(\mu_{\theta,t,1})\theta_t^{\lambda_{\theta,t,1}} $ and the production function is in levels rather than logs of $\theta_t$.

exampleThe issues with the restriction $\mu_{\theta,0,1} = 0$ in the CES case are analogous to the issues with the restriction $\lambda_{\theta,0,1} = 1$ in the trans-log case discussed in Example (ref). I now consider a numerical example analogous to Example (ref) and focus on the measurement and the production function only. That is, suppose that \begin{eqnarray*} \theta_{t+1} &=& \left(\frac{1}{2} \theta_{t} + \frac{1}{2} I_t\right) \exp(\eta_{\theta,t}) \\ Z_{\theta,t,m} &=& \ln(12) + \ln \theta_{t} + \varepsilon_{\theta,t,m} \end{eqnarray*} Notice that $Z_{\theta,t,m} = \ln (12 \theta_{t}) + \varepsilon_{\theta,t,m}$. Let $\tilde{\theta}_t = 12 \theta_{t}$. Just as in Example (ref), we get \begin{eqnarray*} \tilde{\theta}_{t+1} &=& \left(\tilde{\gamma}_{1t} \tilde{\theta}_{t} + (1- \tilde{\gamma}_{1t}) I_t\right)\exp(\tilde{\eta}_{\theta,t}) \\ Z_{\theta,t,m} &=& \tilde{\mu}_{\theta,t,m} + \ln \tilde{\theta}_{t} + \varepsilon_{\theta,t,m} \end{eqnarray*} where $(\tilde{\gamma}_{10},\tilde{\gamma}_{11},\tilde{\gamma}_{12},\tilde{\gamma}_{13},\ldots ) = (0.077, 0.351, 0.435, 0.470,\ldots) $ and $$(\exp(\tilde{\mu}_{0,\theta,m}),\exp(\tilde{\mu}_{1,\theta,m}),\exp(\tilde{\mu}_{2,\theta,m}),\exp(\tilde{\mu}_{3,\theta,m}), \ldots ) = (1, 6.5, 9.25, 10.625, \ldots ).$$ Just as in Example (ref), these two models are observationally equivalent, but suggest very different dynamics.

Monte Carlo simulations

I use a very similar data generating process as \shortciteN{AMN:19}. In particular, I use $$ \theta_{t+1}= A_{t}\left(\gamma_{t} \theta_{t}^{\sigma_{t}}+\left(1-\gamma_{t}\right) I_t^{\sigma_{t}}\right)^{\frac{1}{\sigma_{t}} }\exp(\eta_{\theta,t}) $$ for $t=1,2$. Allowing for $A_t \neq 1$ is equivalent to not imposing the restriction that the coefficients in front of $\theta_t$ and $I_t$ sum to $1$. Since I want to study income transfers as counterfactuals, I augment the setup of \shortciteN{AMN:19} and add $\ln I_t = \beta_{1t} \ln \theta_t + \beta_{2t} \ln Y_t + \eta_{I,t}$, where $Y_0 = Y_1$. To simulate data, I first draw $(\ln(\theta_{0}),\ln(Y))$ from a normal mixture distribution. Given $(\ln(\theta_{0}),\ln(Y))$ and normally distributed $\eta_{\theta,t}$ and $\eta_{I,t}$ I then generate $I_0$, $\theta_1$, $I_1$, and $\theta_2$ using the model. If $\ln I_t$ was equal to $\ln Y_t$, the setup would be exactly equal to that of \shortciteN{AMN:19} with parameters as in their Table 9, which are based on their empirical results, and I use $\sigma_{0} = \sigma_{1} = -0.5$.\footnote{I use slightly different notation to be consistent with the notation above. Specifically, the periods are $t = 0,1,2$ instead of $t = 1,2,3$, I use $I_t$ instead of $X_t$ for the second latent variable, and I use $\sigma_t$ instead of $\rho_t$ to denote the elasticity of substitution. The distribution of $(\ln(\theta_{0}),\ln(Y))$ is the same as the distribution of $(\ln(\theta_{0}),\ln(X))$ in \shortciteN{AMN:19}.} I deviate slightly from their setting by using the additional investment equation with $\beta_{1t} = 0.1$, $\beta_{2t} = 0.9$, and $\eta_{I,t} \sim N(0,0.1^2)$. I simulate three measures each for $\theta_t$ and $I_t$, which have a factor structure with $\mu_{\theta,t,m} = \mu_{I,t,m} = 0$ for all $m,t$. In addition $\lambda_{\theta,t,1} = \lambda_{I,t,1} = 1$ for all $t$, which imposes the scale restrictions and the age-invariance assumptions. In \shortciteN{AMN:19}, all of these loadings are also “normalized” to $1$ in their estimation procedure. In particular, they first estimate the distribution of the measures and of log-income using a normal mixture model. Then, assuming that $(\ln(\theta_{0}),\ln(\theta_{1}),\ln(\theta_{2}),\ln(I_0),\ln(I_1),\ln(Y))$ also has a normal mixture distribution, they use the factor structure and the restrictions to estimate that distribution.\footnote{Interestingly, one can see from the DGP that $(\ln(\theta_{0}),\ln(\theta_{1}),\ln(\theta_{2}))$ does not have a normal mixture distribution due to the nonlinear production function, but this misspecification bias seems to be relatively small in this simulation setup.} Lastly, they take draws from the distribution and estimate the production function parameters by nonlinear least squares. I implement their estimator and my more flexible approach.

For my approach, I use the estimation procedure explained in Appendix (ref), which simplifies in this setup because investment is exogenous (and thus, $\kappa_t = 0$). Moreover, to focus on the scale restriction only, I set $\mu_{\theta,t,m} = \mu_{I,t,m} = 0$ for all $m$ and $t$. Finally, I impose that the first skill measure and the first investment measure are age-invariant. I then set $\lambda_{I,t,1} = 1$ for all $t$, and estimate $\lambda_{\theta,t,1}$. I therefore impose Assumptions (ref), (ref)(a), and (ref)(a), as well as both parts of Assumption (ref). While only one part of the last assumption is needed (i.e. age-invariance of the measures could be dropped), they are both satisfied in this setup. I then estimate $\lambda_{\theta} = \lambda_{\theta,t,1}$ along with the production function parameters by solving

align*[align* omitted — 444 chars of source]

where $\tilde{\theta}_{t,j}$ and $\tilde{I}_{t,j}$ are draws from the estimated distribution of skills and investment using the estimator of \shortciteN{AMN:19}. Similarly, we can estimate $\beta_{1t}$ and $\beta_{2t}$ from a regression of $\tilde{I}_{t,j}$ on $\tilde{\theta}_{t,j}$ and $Y_{j}$.

In the following, I will investigate the effect of multiplying the skill measures by a single constant $s_{\theta}$ in all periods, which changes the loadings, but not the age-invariance assumption. For example, the first measure, say $\tilde{Z}_{\theta,t,1}$, is generated by $$\tilde{Z}_{\theta,t,1} = \log(\theta_{t}) + \varepsilon_{\theta,t,1}$$ but when I estimate the model, I use $$Z_{\theta,t,1} = s_{\theta} \tilde{Z}_{\theta,t,1} = s_{\theta}\log(\theta_{t}) + s_{\theta}\varepsilon_{\theta,t,1},$$ which is also an age-invariant measure. The estimators still impose that the loadings are equal to 1 ($\lambda_{I,t,1} = 1$ for all $t$ with my estimator and $\lambda_{\theta,t,1} = 1$ and $\lambda_{I,t,1} = 1$ for all $t$ with the estimator of \shortciteN{AMN:19}). Of course, in practice, we do not know the DGP and there is no reason to believe that the true loading is $1$. Ideally, the restriction should be a normalization in which case the results would be invariant to scaling the data or changing its units of measurement. However, Corollary (ref) implies that the estimator of \shortciteN{AMN:19} is inconsistent. I consider the implications for elasticities and counterfactual predictions, which are point identified (as shown in Theorem (ref)) and are invariant to the scaling when using the new estimator.\footnote{These features are also invariant to scaling the investment measures with my estimator and I could have set $\lambda_{\theta,t,1}$ instead of $\lambda_{I,t,1}$ to 1.} I take $s_{\theta} = 1$, which is the correctly specified model, as well as $s_{\theta} = 2/3$ and $s_{\theta} = 2$. I report average estimates from 500 simulated data sets.

To summarize the production function estimates I report $ {\partial \ln(\theta_{1}) }/{\partial \ln(\theta_{0})}$ as a function of quantiles of $\ln(\theta_{0})$ and evaluated at the 25th and 75th percentile of $I_0$ as well as $ {\partial \ln(\theta_{2}) }/{\partial \ln(I_1)}$ as a function of quantiles of $\ln(I_1)$ and evaluated at the 25th and 75th percentile of $\theta_{1}$. Figure (ref) displays these partial derivatives for the true parameters, the 2-step estimator of \shortciteN{AMN:19} with different values of $s_{\theta}$, and the invariant estimator that also estimates the scale. In this and the following figures, the results obtained with the invariant estimator are always almost identical to those with the 2-step estimator and $s_{\theta} = 1$ and very similar to those with the true parameter values. Moreover, the invariant estimator adapts to the scale change. Contrarily, the figure illustrates that small changes in the units of measurements can have a large effect on the results when using the 2-step estimator. One interesting finding is that the 2-step estimator underestimates the partial derivative with respect to $\ln(\theta_0)$ at the 25th percentile of $I_0$ for $s_{\theta} = 2/3$ and overestimates it for $s_{\theta} = 2$.

figure[figure omitted — 664 chars of source]

As a first counterfactual, consider an individual, whose value of $\theta_{0}$ equals $Q_{0.1}(\theta_0)$ and all unobservables are $0$. I then exogenously change the income sequence of that individual and check the implied quantile in the skill distribution in period $2$. Of course, the larger income, the higher the relative rank/quantile in the last period. I report results where a feasible choice of income in each period is a given quantile (that will be varied). Instead of using the feasible choices, I distribute the total income among the two periods to maximize the skill quantile in the last period. The second part of Theorem (ref) implies that these counterfactuals are point identified, can be consistently estimated with the invariant estimator, but the results with the 2-step estimator of \shortciteN{AMN:19} will depend on $s_{\theta}$.

figure[figure omitted — 598 chars of source]

The left panel of Figure (ref) shows the quantile as a function of the feasible income quantile ($0.5$ is the median, etc.). Using the true parameters, we can see that, even for large income, the quantile in period $2$ is at most around $0.26$ and is almost flat below the median. The right panel of Figure (ref) displays the corresponding optimal income shares in period $0$. With the true parameters, income should be similar in both periods. The 2-step estimator with $s_{\theta} =1$ and the invariant estimator (irrespective of the scale) yield similar conclusions for the optimal income sequence, but have a small bias for the estimated quantile in period $2$ (due the approximation error of the joint distribution of the measures/skills as mixtures of normals). Figure (ref) also demonstrates that changing the units of measurement can lead to inefficient income transfers when using the 2-step estimator. For example, with $s_{\theta} = 2/3$ the results suggest that we should mainly invest in $t=0$, and with $s_{\theta} = 2$ it appears that we should mainly invest in $t=1$. Moreover, the estimated gains of income transfers are misleading in this case. The inconsistent estimates in the left panel suggest that large income transfers can increase the quantile to almost $0.5$ in period $2$ when $s_{\theta} = 2/3$. Contrarily, with $s_{\theta} = 2$ we would underestimate the effect. Importantly, the results for the new estimator are invariant to changes in the units of measurements.

figure[figure omitted — 498 chars of source]

Next, I consider how exogenous income changes affect the skill distribution. To do so, I take draws from the estimated joint distribution of income and skills in period $0$ (based on the average of the estimated parameters to get representative results) and consider four counterfactual marginal income distributions. First, I increase everyone's income by two standard deviations in period 0. Second, I increase everyone's income by two standard deviations in period 1. Third, I set income to the median for everyone in both periods. Fourth, I increase income by two standard deviations in both periods, but only if the initial skill and income quantiles are below $0.5$. I set all unobservables to their median values. I report results for the invariant estimator only.

Figure (ref) shows the implied distributions of one of the skill measures in the final period (which could be a test score or an adult outcome in an application). These results depend on and should be interpreted relative to the units of measurements of that measures. Figure (ref) is based on $s_{\theta} = 1$. One can see that increasing income in either of the first two periods has very similar effects and leads to an increase in skills. Since everyone's income increases, everyone is better off. If everyone receives the mean income, the variance of the outcome distribution decreases considerably. If we only increase income for people with low initial skills and income, then predictably only the lower tail of the distribution will be affected.

Empirical illustration

In this section, I illustrate the previous findings by replicating some of the estimates of \shortciteN{AMN:19} and showing how changing the units of measurements of the data (or, equivalently, changing the scale restriction) affects parameter estimates and counterfactuals.

\shortciteN{AMN:19} estimate production functions for cognitive skills and health of children with past skills and health as well as parental investment, health, and skills as inputs. They use the CES production function with $\psi_t = 1$. Theorem 3 and Corollary 2 imply that most scale parameters are identified and one only has to restrict the loading of one latent variable in one time period to 1. \shortciteN{AMN:19} set a loading for each observed factor in each time period to 1. Specifically, they write on page 2520: “One way to meet this condition is to normalize each factor on the same measure every period, assuming that the mapping from measure to factor is invariant with respect to the age of the subject. Fortunately, our data are sufficiently rich that we are able to do this for our model. For child cognitive skills, we always normalize the loading on PPVT to one. Similarly, child health is always normalized on height z scores, investments are normalized on amount spent on books, parental health is normalized on mother’s weight, parental cognitive skills is normalized on mother’s years of schooling, and resources are normalized on income.” These restrictions already take the considerations of Agostinelli and Wiswall (2016a, 2024) into account.

\nocite{AW:16a} \nocite{AW:16b}

I consider two changes of the units of measurements. First, I scale book expenditures used to anchor investment. In the descriptive statistics in their Table 2, \shortciteN{AMN:19} report expenditures in USD, but estimation is based on a standardized version of the expenditures in INR.\footnote{While not discussed in their paper or in the code, their data set contains a standardized book expenditures variable, which turns out to be expenditures in INR, but standardized such that the mean and the variance of the pooled expenditures over all time periods are 0 and 1, respectively.} Since the authors write “we use the same measurements at every age and we normalize on the amount spent on books” (p. 2524) and “our investment measure is measured in monetary units, reflecting cost” (p. 2526), it might not be clear to readers which units the results are based on. Ideally, scaling a measure in this way does not affect the main results, but I show that some of the main results are sensitive to using these standardized expenditures instead of book expenditures in 100 INR.\footnote{This transformation amounts to multiplying the standardized test scores by $7.833$ and adding $4.802$.} Second the estimates in the paper are based on a standardized PPVT test score. I reestimate the model after multiplying all test scores by 3, which implies a variance of the scaled scores that is still well below the variance of the actual PPVT test scores. I then compare some of the estimates of \shortciteN{AMN:19} with the ones obtained using the two scaled versions of the data.\footnote{In the first step, \shortciteN{AMN:19} estimate the joint distribution of skills and then simulate from the estimated distribution to estimate production functions and to calculate counterfactuals. This first step of their code is sensitive to the starting values. To ensure that the results below are based on the same (local) minimum, I simply rescale the simulated variables instead of reestimating that part of the model. }

figure[figure omitted — 1,400 chars of source]

\shortciteN{AMN:19} analyze the marginal product of investment on cognition and write: “When considering the production of cognitive skills (\ldots) the productivity of investments is much higher at younger ages: investments are more able to affect cognition earlier on.” This conclusion is based on the left panel of Figure (ref) below, which corresponds to Figure 4(a) of \shortciteN{AMN:19}.\footnote{When calculating counterfactuals, they exclude the constant in the investment equation and incorrectly order some columns of the covariate matrix. I correct these issues and thus obtain slightly different estimates.} The middle panel of Figure (ref) shows the results based on book expenditures in 100 INR. In this case, the marginal effects are higher at age 8 than at age 5 (and even negative in the latter case).\footnote{I follow \shortciteN{AMN:19} and do not restrict the signs of the coefficients (and many of their specifications result in some negative coefficients as well). } When I rescale test scores instead of book expenditures, the estimated effects are much lower, as can be seen in panel (c).

By Theorem (ref), there are different, invariant ways to interpret the effect of investment on skills. For example, here we can conclude: Fixing all variables at their median values,

enumerate• if we increase someones investment from the 40% quantile to the 60% quantile at age 5, we increase her skills from the 46% quantile to the 51% quantile at age 8, or • if we increase investment by 1% at age 5, we increase skills by 0.004% at age 8, or • if we increase investment s.t. book expenditures increase by 1 std (or 783 INR) at age 5, we increase skills at age 8 s.t. the standardized PPVT score increases by 0.26 std (or 0.23 units).

Although the effect described in interpretation (ii) may appear smaller than that in (i), the magnitudes are in fact consistent with each other. To understand their relationship, note that the logs of the latent variables follow mixtures of normal distributions. Here, the estimated variance of log-investment is substantially larger than that of log-skills. As a result, a 10% increase in investment at the median corresponds to a shift to the 50.18% quantile. In contrast, a 0.04% increase in skills at the median corresponds to a shift to the 50.07% quantile. The relative change in these quantiles ($0.0007/0.0018=0.39$), is comparable to the effect in interpretation (i), where the relative quantile change is $0.05/0.2=0.25$.

figure[figure omitted — 2,526 chars of source]

The authors also study the impact on cognition of an income transfer equal to 25% of mean income in the entire sample. Specifically, they consider the change in the median level of cognition in standard deviation units. The results for the different scales are shown in Figure (ref). The income transfer is made before age 5 in the upper panels and between ages 5 and 8 in the lower panels. The authors write: “In terms of timing, the largest impact is obtained if the transfer takes place when the children are between 5 and 8 for cognition”, which is based on the much larger response in the lower panel at age 12. This conclusion does not hold when using book expenditures in 100 INR, because the effect of investment is then estimated to be negative in panel (e). When scaling test scores instead, the effects are very close to 0 (see panels (c) and (f)).

There are two problems with the counterfactuals reported in Figure (ref) using the specification of \shortciteN{AMN:19}. First, identified parameters are set to 1, which leads to misspecification. Second, this counterfactual does not correspond to those that are invariant to the scale and location restrictions. To fix the second problem, one can focus on the change in the median log of skills rather than the level of skills.

figure[figure omitted — 1,313 chars of source]

The results for the log-skills using the estimator of \shortciteN{AMN:19} are shown in Figure (ref) in Appendix (ref). Since the estimates are based on setting identified parameters to 1, the results are still sensitive to the scales. Figure (ref) in Appendix (ref) shows the estimates corresponding to the counterfactual of \shortciteN{AMN:19}, but based on the estimator that also estimates the identified scales, which also depend on the the scales of the test scores (but not investment). Finally, Figure (ref) displays the counterfactual in terms of log-skill based on the flexible estimator, which is invariant to the units of measurements (and therefore, the other scales are omitted). Based on Figure (ref), we can still conclude that investment between ages 5 and 8 is more beneficial. However, compared to panels (d) of Figure (ref), the new estimates imply much more heterogeneity by wealth. Figure (ref) in Appendix (ref) shows the effects of income transfers on the entire counterfactual distribution of the test scores.

Conclusion

This paper demonstrates that in an important class of skill formation models, seemingly innocuous scale and location restrictions may not function as mere normalizations. Instead, they can constrain identified parameters and influence counterfactuals. The specific implications depend on the feature of interest, the production function, the measurement system, and the estimation method. In a new identification analysis, I pool all restrictions of the model and characterize the identified set without imposing additional scale and location restrictions. Notably, many key features remain invariant to these restrictions, are identified under weaker assumptions, yield robust policy implications, and are comparable across different studies.

Researchers often impose “normalizations” to achieve point identification in models that are otherwise only partially identified. To avoid unintended consequences and potentially misleading conclusions, it is critical to clarify which parameters and features are invariant to these restrictions, ensuring they are genuine normalizations as defined in Definition (ref). While proving such invariance properties theoretically can be challenging, it is typically possible to demonstrate the robustness of key results to these restrictions.

One example closely related to the model considered here are skill formation models that explicitly solve household optimization problems to determine optimal investment functions (see e.g. \citeN{DFW:13} or \citeN{CL:20}). These models, which include a production function, are then estimated using the optimality conditions alongside a measurement system for latent variables. While these models have a different structure compared to the ones I study and a detailed analysis is left for future research, it appears that analogous normalization issues may arise, with consequences that depend on the model specification and the types of counterfactuals examined.

Data availability statement

The data and code underlying this research is available on Zenodo at

\url{https://doi.org/10.5281/zenodo.16437164}.