Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
This text was truncated for display. The citation measures were computed over the complete text.
Heterogeneity, Uncertainty and Learning: Semiparametric Identification and Estimation
abstractWe provide identification results for a broad class of learning models in which continuous outcomes depend on three types of unobservables: known heterogeneity, initially unknown heterogeneity that may be revealed over time, and transitory uncertainty. We consider a common environment where the researcher only has access to a short panel on choices and realized outcomes. We establish identification of the outcome equation parameters and the distribution of the unobservables, under the standard assumption that unknown heterogeneity and uncertainty are normally distributed. We also show that, absent known heterogeneity, the model is identified without making any distributional assumption. We then derive the asymptotic properties of a sieve MLE estimator for the model parameters, and devise a tractable profile likelihood-based estimation procedure. Our estimator exhibits good finite-sample properties. Finally, we illustrate our approach with an application to ability learning in the context of occupational choice. Our results point to substantial ability learning based on realized wages.
pgfscope{abstract}
Introduction
Learning models, in which agents have imperfect information about their environment and update their beliefs over time, are frequently used in economics. These models have received particular interest in various subfields in empirical microeconomics, including labor economics Miller84,AG12,Pastorino15, Hincapie20,Pastorino22, economics of education Arcidiacono04,Zafar11,SS12,stange12,Thomas19,KP21,Proctor22,AAMR25, industrial organization and health ackerberg2003advertising,CS04,CS05,AC05,ChanHamilton06,AJ20. Since the seminal work of EK96, learning models have also been popular in the marketing literature CEK13. However, while learning models are often estimated, much remains to be known about the identification of this important class of models.
In this paper, we provide new semiparametric identification results for a general class of learning models. We consider an environment in which the researcher has access to a short panel of choices and realized outcomes only. As such, our results are widely applicable, including in frequent situations where one does not have access to elicited beliefs data, or to a set of selection-free measurements of unobserved individual heterogeneity. Specifically, throughout our analysis we consider a potential outcome model in which individual $i$'s potential outcome in period $t$ from assignment $d$ is given by
equation[equation omitted — 4,519 chars of source]
Related literatures
Our paper contributes to several strands of the literature. First and foremost, we contribute to the literature that studies the identification of learning models, generally in the context of specific applications AC05,Gong_JMP,Pastorino22,AAMR25. A central distinction from most of the papers in this literature is that we impose only mild restrictions on the choice process. Importantly, we remain agnostic about how choices depend on individual beliefs about $X^{*}_{u,i}$, while allowing these beliefs to depend arbitrarily on past choices and realized outcomes. Particularly relevant for us is the complementary work of Pastorino22, which establishes identification results in a different and non-nested framework of a two-sided learning model in which workers and firms have imperfect information. Key to the identification strategy proposed in that paper is to leverage particular mixture representations of selected one-dimensional outcomes.\footnote{See also recent related work by dPGPS25 which investigates the identification of a two-sided matching model with learning and human capital accumulation. As in our paper, identification of the outcome equations and the distribution of unobserved heterogeneity relies on bruni1985identifiability.} Related mixture representations also play an important role in our analysis.
Our paper also fits into a literature that focuses on the identification of Markovian dynamic discrete choice models in the presence of persistent unobserved heterogeneity HN07,hu2008instrumental,kasahara2009nonparametric,hu2012nonparametric,sasaki2015heterogeneity,HS18,AGL21,bunting22,arellano2017nonlinear. Unlike these papers, we do not impose a Markov structure, since current beliefs and decisions are allowed to depend on the entire history of past outcomes and decisions.\footnote{Although our framework is more general, Bayesian learning models often naturally possess a first order Markov structure. There are several additional significant differences between our paper and the listed literature. Notably, hu2012nonparametric focus on scalar unobserved heterogeneity, whereas the existence of multivariate unobserved heterogeneity is fundamental to our main setting. Beyond this, several of their assumptions may fail to hold in our set-up. For instance, since the support of the latent beliefs is larger than the support of the choices, the requirement that the observed variables be invertible measurements of the latent variables hu2012nonparametric will generally fail to hold.} More broadly, our analysis is related to the literature that deals with the identification of mixture models HKS14, CK16, KL18. In particular, central to our main identification result is the observation that the distribution of current outcomes conditional on the sequence of past choices and outcomes is a mixture of normal distributions.
Since the outcome equation in our model involves interactions between unobserved individual- and time-specific effects, our paper fits into the literature that examines the identification and estimation of panel data models with interactive fixed effects Madansky64,HS87,bai09,freyberger2017non. An important distinction comes from the fact that these papers consider a selection-free environment. In contrast, individual choices, along with associated selection issues that affect the potential outcomes, play a central role in our analysis.
Finally, by applying our framework to examine how imperfect information and learning shape occupational choices and wages, our paper also fits into the literature that highlights the important role of imperfect information in labor market trajectories and outcomes Miller84,AG12,Papageorgiou14,Pastorino15,CPWZ18,GS19,GSS22,AAMR25. A distinctive feature of our approach is that it allows us to remain flexible on how agents sort across occupations and form their beliefs about future earnings. Our identification results allow, in particular, for potential deviations from rational expectations on future outcomes, which recent evidence based on subjective beliefs has shown to be important DGM21,CGSS24.
Organization of the paper
The remainder of the paper is organized as follows. Section (ref) introduces and discusses the set-up of the model. Section (ref) contains our main identification results, both for the general case and for the case of a pure learning model. We discuss in Section (ref) the estimation and inference on the parameters of interest, before turning in Section (ref) to the implementation of our estimator and its finite-sample performances. We illustrate in Section (ref) our approach with an application to ability learning in the context of occupational choice. Section (ref) concludes. The appendix gathers all the proofs, additional material on the variance decompositions, the implementation of our estimator, and further Monte Carlo simulation results. Finally, our estimation method can be implemented using our companion Python package, spmlex, which is available at \hyperlink{https://github.com/pdiegert/spmlex}{https://github.com/pdiegert/spmlex}.
Notation: for a given random variable $A$, we denote by $a$ its realization, $\mathcal{S}(A)$ indicates its support, $F_{A}$ denotes its cumulative distribution function, $q_{\alpha}[A]$ its $\alpha \in [0,1]$ quantile, whereas $f_{A}$ indicates its probability mass or density function. For any sequence $(a_{1},a_{2},\dots,a_{S})$ and $s\le{S}$, we let $a^{s}=(a_{1},a_{2},\dots,a_{s})$. $A\protect\mathpalette{\protect\independenT}{\perp}{B}\mid{C}$ indicates that $A$ and $B$ are statistically independent conditional on $C$. Finally, unless stated otherwise, we suppress the individual subscript $i$ from all random variables in the remainder of the paper.
Set-up
Throughout the paper, we consider a set-up where potential outcomes have an interactive fixed-effect structure of the following form:
equation[equation omitted — 7,486 chars of source]
Identification
We first provide in Subsection (ref) a high-level overview of the underlying reweighting scheme that plays an important role in both of the proposed identification strategies. We then discuss identification in the case with both known and unknown unobserved heterogeneity (Subsection (ref)), before turning to the pure learning case where the only source of permanent unobserved heterogeneity is initially unknown to the agent (Subsection (ref)).
Reweighting strategy
Key to the identification problem analyzed in this paper is how to recover the conditional distributions of potential outcomes (i.e., $f_{Y_t(d_t)|X_{t},X^{*}}$ for each $t$ and $d_t$) and selection probabilities (i.e., $f_{D_t|X_{t},X^{*}_k}$ for each $t$), from the selected population distribution (i.e., $f_{Y^T,D^T, X^T}$) which is directly identified from the data.
We now provide intuition as to how one can leverage the structure imposed on the choice process to address the censored data problem.
To illustrate, consider a simplified version of our model with a binary choice in each period (i.e., $\mathcal{S}(D_{t}) = \{0, 1\}$) and without covariates. Let $D := \prod_{t = 1}^T D_{t}$, $Y := (Y_1, \ldots, Y_T)$ and $Y(1) := (Y_{1}(1), \ldots, Y_{T}(1))$, and focus on identification of the distribution of the potential outcome $Y(1)$. By Bayes' rule, the relationship between the target and censored distributions can be characterized as follows:
align*[align* omitted — 2,420 chars of source]
Known and unknown heterogeneity
This section provides sufficient conditions for identification of the baseline model discussed in Section (ref). We first impose a form of conditional independence on $(\epsilon_{t}(d), D_{t}, X_{t})$.
manualassumption{KL1}
Equation (ref) holds, and for any $t \geq 2$ and $d\in\mathcal{S}(D_{t})$,
\begin{equation*}
F_{\epsilon_{t}(d),D_{t},X_{t}|Y^{t-1},D^{t-1},X^{t-1},X^{*}} = F_{\epsilon_{t}(d)}F_{D_{t}|X^{t},Y^{t-1},D^{t-1},X^{*}_k}F_{X_{t}|Y^{t-1},D^{t-1},X^{t-1}}.
pgfscope{equation*}
Furthermore, for any $d\in\mathcal{S}(D_1)$, $F_{\epsilon_{1}(d),D_{1},X_{1}|X^{*}} = F_{\epsilon_{1}(d)}F_{D_{1}|X_{1},X^{*}_k}F_{X_{1}|X^*}$.
pgfscope{manualassumption}
Assumption (ref) imposes the potential outcome model in Equation (ref) and contains three independence conditions. First, it implies that the additive transitory shock in the outcome equation ($\epsilon_{t}(d)$) is independent of all contemporaneous and lagged variables. This is closely related to the standard fixed effect assumption that dependence in outcomes across periods is due to the latent fixed effect (e.g., freyberger2017non and sasaki2015heterogeneity). However, note that we allow for arbitrary within-period dependence between the additive shocks ($\epsilon_{t}(d)$ and $\epsilon_{t}(\tilde{d})$, for $d\neq \tilde{d}$). Second, the unknown factor ($X^{*}_u$) does not directly affect treatment assignments ($D_{t}$), a natural restriction discussed in Section (ref). Third, we also impose that the transition of the control variables ($X_{t}$) does not directly depend on the time-invariant unobservables ($X^*$). Importantly, this does allow $X_{t}$ to depend on $X^*$ through past choices and outcomes. For instance, in the context of occupational choices, this restriction is implied by the standard assumption that occupation-specific work experiences depend on $X^*$ through past occupational choices (see, e.g., keane1997career, keane1997career).
Our second assumption (ref) imposes that the unknown component of the individual effect is drawn from a multivariate normal distribution, and that the random shock in the outcome equation is normally distributed too. This is a common assumption in the Bayesian learning literature, to which we return in Remark (ref).
\begin{manualassumption}{KL2} For all $(x_1,x_k^*)\in\mathcal{S}(X_1)\times\mathcal{S}(X_k^*)$,
$X^{*}_{u}\mid(X_{1},X^{*}_{k})=(x_1,x_{k}^{*})\sim{N}\left(0,\Sigma_{u}(x_{1})\right)$ and $\forall~d\in\mathcal{S}(D_{t}),~\epsilon_{t}(d)\sim{N}(0,\sigma_{t,d}^{2})$.
pgfscope{manualassumption}
Assumption (ref) implies a Gaussian conjugate posterior distribution for $X_u^*$, which we summarize in Lemma (ref). Importantly, neither this assumption nor Assumption (ref) place any restriction on the dependence between $X_k^*$ and $X_1$.\footnote{Lemma (ref) and our main identification result would go through if one replaces the first part of Assumption (ref) with $X^{*}_{u}\mid(X_{1}=x_{1},X^{*}_{k}=x_{k}^{*})\sim{N}\left(0,\Sigma_{u}(x_1,x_k)\right)$ under appropriate regularity conditions on $x_k\mapsto\Sigma_{u}(x_1,x_k)$, including for each $x_{k}^{*}-\tilde{x}_{k}^{*}>0$, $\Sigma_{u}(x_{1},x_{k}^{*})-\Sigma_{u}(x_{1},\tilde{x}_{k}^{*})$ is positive (or negative) semi-definite. For simplicity, we maintain the stronger Assumption (ref) when establishing identification in Theorem 1 below and in the rest of the paper.} To do so, define $(\mu_{t},\Sigma_{t})$ recursively as follows. First, $(\mu_{1},\Sigma_{1})=(0,\Sigma_{u}(x_{1}))$. Second,
\begin{align*}
\Sigma_{t+1} & =\left(\Sigma_{t}^{-1}+\lambda_{t,d_{t}}^{u}(\lambda_{t,d_{t}}^{u})^{\intercal}\sigma_{t,d_{t}}^{-2}\right)^{-1}, \\
\mu_{t+1} & =\Sigma_{t+1}\left(\Sigma_{t}^{-1}\mu_{t}+\lambda_{t,d_{t}}^{u}\frac{y_{t}-x_{t}^{\intercal}\beta_{t,d_{t}}-x^{*}_{k}\lambda_{t,d_{t}}^{k}}{\sigma_{t,d_{t}}^{2}}\right).
pgfscope{align*}
\begin{lemma}
Let Assumptions (ref) and (ref) hold. Then, for all $t \geq 2$, $X^{*}_{u}$ conditional on $(Y^{t-1},D^{t-1},X^{t},X^{*}_{k})=(y^{t-1},d^{t-1},x^{t},x^{*}_{k})$ is distributed ${N}(\mu_{t},\Sigma_{t})$.
pgfscope{lemma}
Suppose $X^{*}_{u}\in\mathbb{R}^{p}$. Our three remaining assumptions are as follows.
\begin{manualassumption}{KL3}
(A) For some $d\in\mathcal{S}(D_{1})$, the element of $\beta_{1,d}$ associated with the constant term is zero, and $\lambda_{1,d}^{k}=1$. (B) For some $d^{p}\in\mathcal{S}(D^{p})$, $\left(\lambda_{1,d_{1}}^{u}\cdots \lambda_{p,d_p}^{u}\right)=I_{p\times{p}}$.
pgfscope{manualassumption}
Assumption (ref) is a location-scale normalization on finite-dimensional parameters, which reflects the fact that the latent factors are only identified up to location and scale. This type of assumption is standard in interactive fixed effect models.
Finally, we impose in Assumptions (ref) and (ref) below several regularity conditions. We start with Assumption (ref), which places support restrictions on various objects of the model. In what follows, we let $\theta_{1}\coloneqq\left\{\{\beta_{t},\lambda_{t},\sigma_{t}^{2}\}_{t=1}^{T},\Sigma_{u}(x_{1})\right\}\in\Theta_{1}\subset\mathbb{R}^{{\rm dim}\Theta_{1}}$, where $\{\beta_{t},\lambda_{t},\sigma_{t}^{2}\}\coloneqq\{\beta_{t,d},\lambda_{t,d},\sigma_{t,d}^{2}\colon d\in\mathcal{S}(D_t)\}$.
\begin{manualassumption}{KL4}
(A) For each $x_{1}\in\mathcal{S}(X_{1})$, $\Theta_{1}$ is a compact set. (B) $\mathcal{S}(X^{*}_{k})$ is compact. (C) For each $t$ and $d\in\mathcal{S}(D_t)$, $(\lambda_{t,d}^{u})^{\intercal}\Sigma_{t}\lambda_{t,d}^{u}+\sigma_{t,d}^{2}\neq0$, $\sigma_{t,d}^{2}\neq0$ and $\forall~x_1\in\mathcal{S}(X_1)$, $\Sigma_{u}(x_{1})$ is non-singular. (D) For each $y^{t-1},d^{t},x^{t}$ in their support, $\mathcal{S}(X^{*}_{k}\mid({Y}^{t-1},D^{t},X^{t})=(y^{t-1},d^{t},x^{t}))=\mathcal{S}(X_k^*)$ and $Var(X_k^*)\neq0$. (E) For each $t$ and $d\in\mathcal{S}(D_t)$, $E[X_t X_t^\intercal\mid D_t=d]$ is non-singular. (F) For all $t$, $Var(D_{t})\neq0$.
pgfscope{manualassumption}
Part (A) states that the finite-dimensional parameters $\theta_1$ belong to a compact set. Part (B) requires that the known latent factor $X^{*}_{k}$ has compact support. This holds if the distribution of $X^{*}_{k}$ has discrete support, although this clearly applies to a broader set of distributions. We return to this compactness condition in Remark (ref) below. Part (C) requires certain normally distributed random variables to have non-singleton support. Part (D) imposes a rectangular support condition and a nondegeneracy assumption on the distribution of $X^{*}_{k}$. These conditions are typically satisfied in dynamic discrete choice models with unobserved heterogeneity, which generally impose a large support assumption on the random utility shocks. Part (E) imposes that the support of $X_{t}$ conditional on $D_t$ is sufficiently rich. Finally, Part (F) imposes the requirement that the support of the choice variables contain at least two elements.
Next, Assumption (ref) below contains a set of regularity conditions that ensure that the latent individual effect $X^{*}$ alters outcomes sufficiently differently across time and assignments.
\begin{manualassumption}{KL5}
(A) For each $t$ and $d_{t}\in\mathcal{S}(D_t)$ there exist two sequences $(d^{t-1}, \tilde{d}^{t-1})\in\mathcal{S}(D^{t-1})^2$ such that $(\lambda_{t,d_{t}}^{u})^{\intercal}\Sigma_{t}\sum_{s=1}^{t-1}\left(\lambda_{s,d_{s}}^{u}\frac{\lambda_{s,d_{s}}^{k}}{\sigma_{s,d_{s}}^{2}}-\lambda_{s,\tilde{d}_{s}}^{u}\frac{\lambda_{s,\tilde{d}_s}^{k}}{\sigma_{s,\tilde{d}_s}^{2}}\right)\neq0$. (B) For all $t$ and $d_{t}\in\mathcal{S}(D_t)$, $\lambda_{t,d_{t}}^{k}\neq0$. (C) For all $t$ and $d^{t}\in\mathcal{S}(D^t)$, $ \lambda_{t,d_{t}}^{k} - (\lambda_{t,d_{t}}^{u})^{\intercal}\Sigma_{t} \sum_{s=1}^{t-1}\lambda_{s,d_{s}}^{u}\frac{\lambda_{s,d_{s}}^{k}}{\sigma_{s,d_{s}}^{2}}\neq0.$ (D) For all $d^2\in\mathcal{S}(D^2)$, $(\lambda_{2,d_{2}}^{u})^{\intercal}\Sigma_{2}\lambda_{1,d_{1}}^{u}\frac{\lambda_{1,d_{1}}^{k}}{\sigma_{1,d_{1}}^{2}}\neq0$. (E) There exists $\{(d_{2,i},\tilde{d}_{2,i})\in\mathcal{S}{(D_{2})}^2:i=1,2,\dots,p\}$ which satisfy
\begin{align*}
\left(\lambda_{2,d_{2,1}}^{u}\cdots{\lambda_{2,d_{2,p}}^{u}}\right)^{-\intercal}\mathrm{vec}(\lambda_{2,d_{2,1}}^{k},\dots,\lambda_{2,d_{2,p}}^{k})
\neq\left(\lambda_{2,\tilde{d}_{2,1}}^{u}\cdots{\lambda_{2,\tilde{d}_{2,p}}^{u}}\right)^{-\intercal}\mathrm{vec}(\lambda_{2,\tilde{d}_{2,1}}^{k},\dots,\lambda_{2,\tilde{d}_{2,p}}^{k}).
pgfscope{align*}
(F) For all $d^T\in\mathcal{S}(D^T)$, $\{\lambda_{t,d_{t}}^{u}:t=1,\ldots,T\}$ is linearly independent.
pgfscope{manualassumption}
This assumption is fairly mild as it primarily rules out knife-edge cases where the effect of different elements of permanent unobserved heterogeneity is exactly zero.\footnote{This type of assumption is similarly required in latent factor models without selection or learning in order to rule out degeneracies freyberger2017non.} Part (A) requires that the aggregate effect of $X^{*}_{k}$ on outcomes associated with choice $d_{t}$ is different for at least two histories $(d^{t-1},\tilde{d}^{t-1})$. Part (B) assumes that the direct effect of $X^{*}_{k}$ is non-zero in each period and each assignment. Part (C) states that the aggregate effect of $X^{*}_{k}$ on outcomes must be non-zero---that is, that the direct effect $\lambda_{t,d_{t}}^{k}$ is not perfectly offset by the effect mediated through previous choices. Part (D) ensures that there is a non-zero effect of previous choices in $t=2$. Part (E) requires that for $t=2$ the relative effect of known and unknown $X^{*}$ changes across choices. In the special case where $X^{*}_{u}\in\mathbb{R}$ (i.e., $p=1$), the condition reduces to $\frac{\lambda_{2,d_{2}}^{k}}{\lambda_{2,d_{2}}^{u}}\neq\frac{\lambda_{2,\tilde{d}_{2}}^{k}}{\lambda_{2,\tilde{d}_{2}}^{u}}$, i.e., that the ratio of factor loadings varies across some assignments. More generally, for $X^{*}_{u}\in\mathbb{R}^{p}$, this condition implies that, for $t=2$, the set of assignments must contain at least $p+1$ elements. Finally, Part (F) requires that the initially unknown factor affects each outcome via a different linear combination.
We are now in a position to state our main identification result. We denote by $\theta=\left\{\{\beta_{t},\lambda_{t},\sigma_{t},g_t,h_{t}\}_{t=1}^{T},\Sigma_{u},F_{X^{*}_{k},X_1}\right\}\in\Theta$ the model parameters, where $g_t: =dF_{X_{t}|Y^{t-1},D^{t-1},X^{t-1}}$.
\begin{theorem}
Suppose the distribution of $(Y_{t},D_{t},X_{t})_{t=1}^{T}$ is observed for $T={2p}+1$ periods, and that Assumptions (ref)-(ref) hold. Then $\theta$ is point identified.
pgfscope{theorem}
The first step is to show, from Assumptions (ref) and (ref) and Lemma (ref) that $Y_{t}$ is normally distributed conditional on lagged outcomes $Y^{t-1}$, assignments $D^{t}$, covariates $X^{t}$ and the known component of the latent individual effect, $X^{*}_{k}$. This implies that $Y_{t}$ conditional on $(Y^{t-1}, D^{t},X^{t})$ is a Gaussian mixture distribution parameterized by $X^{*}_{k}$. Then under the compact support and non-degeneracy assumptions (Assumptions (ref) (A)-(C)), one can apply a result from bruni1985identifiability to identify the aforementioned mixture distribution up to an affine transformation of $X^{*}_k$. Next, the normalization and regularity assumptions (Assumptions (ref)-(ref)) are used to pin down the affine transformation, leading to identification of the distribution of $(Y^{T},D^{T},X^{T},X^{*}_{k})$. Knowledge of this distribution identifies the components of the model related to the known component of the individual latent effect, namely $\left\{\{\beta_{t},\lambda_{t}^{k},h_{t}\}_{t=1}^{T},F_{X^{*}_{k},X_{1}}\right\}$. The final step is to disentangle the effect of the learned component ($X^{*}_u$) and the idiosyncratic uncertainty ($\epsilon_{t}(d))$ in order to identify $\left\{\{\lambda_{t}^{u},\sigma_{t}^{2}\}_{t=1}^{T},\Sigma_{u}\right\}$. This is done by showing that the joint distribution of $(Y^{T},D^{T},X^{T})$ conditional on $X^{*}_{k}$, suitability weighted by the assignment probabilities, is a normal-weighted mixture of normal distributions. This allows us to identify $\left\{\{\lambda_{t}^{u},\sigma_{t}^{2}\}_{t=1}^{T},\Sigma_{u}\right\}$ from the second moments of the reweighted distribution. We refer the interested reader to Section (ref) for a formal derivation.\footnote{Note that, while we assume for simplicity that $T=2p+1$, extension to a larger horizon $T$ is straightforward. The same applies for the pure learning model considered in Section (ref).}
\begin{remark}[Compact support assumption]
Assumption (ref) (B) imposes that the known component of the latent individual effect has bounded support. In applications, it is common to assume $X^{*}_{k}$ has finite support with known cardinality. Assumption (ref) (B) relaxes this restriction in the sense that the number of support points of $X^{*}_{k}$ need not be known a priori, and indeed may be infinite.\footnote{Compactness is used in particular to apply the Stone-Weierstrass approximation theorem, which plays an important role in the identification proof of bruni1985identifiability.}
pgfscope{remark}
\begin{remark}[Normality of unknown factor]
As summarized in Lemma (ref), an important implication of the normality assumptions (Assumption (ref)) is the resulting normal conjugate prior with a tractable closed form. For this reason, these assumptions are very common in the learning literature. In the context of our analysis though, the key implication of normality is rather to enable identification of the distribution of $Y_{t}\mid\left(Y^{t-1},D^{t},X^{t},X^{*}_{k},\right)$ from variation in the realized outcome $Y_{t}$ only. Namely, under Assumption (ref), the distribution of $Y_{t}\mid\left(Y^{t-1},D^{t},X^{t}\right) $ is a mixture of normal distributions with mixture weights given by the distribution of $X^{*}_{k}\mid\left(Y^{t-1},D^{t},X^{t}\right)$. This allows us to establish identification by leveraging results for mixtures of normal distributions bruni1985identifiability.\footnote{
That identification of the distribution of $X^{*}_{k}$ arises from variation in the scalar outcome variable $Y_{t}$ highlights why we restrict $X^{*}_{k}$ to be a scalar random variable. If $Y_{t}$ was vector-valued instead, then we expect that our arguments would easily extend to allow for a multivariate $X^{*}_{k}$.}
pgfscope{remark}
\begin{remark}[Role of covariates] Inspection of the proof shows that the covariates $X_t$ are not needed to identify the parameters $\theta$, beyond $\{\beta_t:t=1,\ldots,T\}$. In particular, one can easily adapt the proof to establish identification for a more flexible specification where $X_t$ enters the outcome equation through an additive nonparametric shifter. We maintain linearity throughout for estimation precision and to preserve tractability.
pgfscope{remark}
\begin{remark}[Invariance to normalization]
The normalization assumption (Assumption (ref)) is a true normalization in the sense that particular meaningful economic parameters are invariant to the assumption. Specifically, we can show that this is the case of the average and quantile structural functions. To formalize this notion, define $C_{t,d}^{k} \coloneqq X^{*}_{k} \lambda_{t,d}^{k}$, $C_{t,d}^{u} \coloneqq (X^{*}_{u})^{\intercal} \lambda_{t,d}^{u}$ and let $Q_{\alpha}\left[X\right]$ be the $\alpha$-quantile of a random variable $X$. Let $x \in \mathcal{S}(X_{t})$ and define the quantile structural functions associated with the potential outcomes $Y_t(d_t)$ as follows:
\begin{align*}
s_{1,t}(x,\alpha)=&x^{\intercal}\beta_{t,d_t}+Q_{\alpha}[C_{t,d_t}^{k}+C_{t,d_t}^{u}+\epsilon_{t}(d_t)],\\
s_{2,t}(x,\alpha_{1},\alpha_{2},\alpha_{3})=&x^{\intercal}\beta_{t,d_t}+Q_{\alpha_{1}}[C_{t,d_t}^{k}]+Q_{\alpha_{2}}[C_{t,d_t}^{u}]+Q_{\alpha_{3}}[\epsilon_{t}(d_t)],
pgfscope{align*}
and the average structural function as $s_{3,t}(x)=x^{\intercal}\beta_{t,d_t}+\int{u}dF_{C^k_{t,d_t}+C^u_{t,d_t}+\epsilon_{t}(d_t)}(u)$. In Appendix (ref) we prove the following corollary:
\begin{corollary}
Suppose the Assumptions (ref), (ref) and (ref) hold and that for each $(x_1,x_{k}^{*})\in\mathcal{S}(X_1)\times\mathcal{S}(X_k^*)$, $X^{*}_{u}\mid(X_{1},X^{*}_{k})=(x_1,x_{k}^{*})\sim{N}\left(\mu_{u},\Sigma_{u}(x_{1})\right)$ and for all $t$ and $d\in\mathcal{S}(D_t)$, $\epsilon_{t}(d)\sim{N}(c_{t,d},\sigma_{t,d}^{2})$. Furthermore, suppose that for some $d^p\in\mathcal{S}(D^p)$, $(\lambda_{1,d_{1}}^{u}\cdots\lambda_{p,d_{p}}^{u})$ is full rank. Then $s_{1,t}(x,\cdot)$, $s_{2,t}(x,\cdot,\cdot,\cdot)$ and $s_{3,t}(x)$ are identified for all $x$ on the support of $X_{t}$.
pgfscope{corollary}
pgfscope{remark}
Pure learning model
This section considers a special case of the model of Section (ref), in which all components of the latent individual effect are initially unknown to the decision maker ($X^{*}=X^{*}_{u}$). Without needing to distinguish initially known and unknown heterogeneity, a stronger identification result is achieved. In particular, no parametric restrictions on the distribution of the unobservables are required. We establish identification in this model under Assumptions (ref)-(ref) stated below.
manualassumption{L1} For all $t$ and $d\in\mathcal{S}(D_t)$,
$Y_{t}(d)=X_{t}^{\intercal}\beta_{t,d}+(X^{*})^{\intercal}\lambda_{t,d}+\epsilon_{t}(d)$.
For any $t \geq 2$ and $d\in\mathcal{S}(D_t)$,
\begin{equation*}
F_{\epsilon_{t}(d),D_{t},X_{t}|Y^{t-1},D^{t-1},X^{t-1},X^{*}} =
F_{\epsilon_{t}(d)}F_{D_{t}|Y^{t-1},D^{t-1},X^{t}}F_{X_{t}|Y^{t-1},D^{t-1},X^{t-1}}.
pgfscope{equation*}
Furthermore, for any $d\in\mathcal{S}(D_1)$, $F_{\epsilon_{1}(d),D_{1},X_{1}|X^{*}} =
F_{\epsilon_{1}(d)}F_{D_{1}|X_{1}}F_{X_{1}|X^*}.$
pgfscope{manualassumption}
Assumption (ref) adapts Assumption (ref) to reflect that there is no initially known component of unobserved heterogeneity.
\begin{manualassumption}{L2}
(A) The joint density of $(Y,X^{*})$ and $(D,X)$ admits a bounded density with respect to the product measure of the Lebesgue measure on $\mathcal{S}(Y)\times\mathcal{S}(X^{*})$ and some dominating measure on $\mathcal{S}(D)\times\mathcal{S}(X)$. All marginal and conditional densities are bounded. (B) For each $x_1\in\mathcal{S}(X_1)$, $X^{*}\mid{X_1}=x_1$ has full support. (C) For each $t$ and $d\in\mathcal{S}(D_t)$, the characteristic function of $\epsilon_{t}(d)$ is non-vanishing, and $E[\epsilon_{t}]=0$.
pgfscope{manualassumption}
Assumption (ref) substantially weakens Assumption (ref) by replacing the normality assumption with a full support assumption. Let $X^{*} \in \mathbb{R}^p$.
\begin{manualassumption}{L3}
For some $d^p\in\mathcal{S}(D^p)$, (A) $\left(\lambda_{1,d_{1}}\cdots\lambda_{p,d_{p}}\right)=I_{p\times{p}}$ and (B) the element of $\beta_{t,d_t}$ associated with the constant component of $X_t$ is zero.
pgfscope{manualassumption}
\begin{manualassumption}{L4}
(A) For each $(y^{t-1},x^t)\in\mathcal{S}(Y^{t-1},X^t)$, $\Pr(D_t=d\mid Y^{t-1}=y^{t-1},X^{t}=x^t)>0$ for all $d\in\mathcal{S}(D_t)$. (B) For each $x_1\in\mathcal{S}(X_1)$, the variance-covariance matrix of $X^{*}\mid{X_1=x_1}$ is full rank. (C) For each $t$ and $d\in\mathcal{S}(D_t)$, the variance-covariance matrix of $X_t$ conditional on $D_t=d$ is non-singular.
pgfscope{manualassumption}
Assumption (ref) are normalization assumptions, which are standard in interactive fixed effect models. Assumption (ref) (A) is similar to Assumption (ref) (D). It requires that for each history ($y^{t-1},d^{t-1},x^{t}$), some units are assigned to $D_{t}=d_{t}$ for each $d_{t}\in\mathcal{S}(D_{t})$. This assumption is typically satisfied in parametric dynamic discrete choice models (see, e.g., keane1997career, keane1997career and Blundell17, Blundell17 for a survey). At the cost of increased notational burden, this assumption could be weakened to hold for certain sequences of choices only.
\begin{manualassumption}{L5}
For any $d^T\in\mathcal{S}(D^T)$, all $p\times p$ submatrices of $(\lambda_{1,d_1}^{u}\cdots\lambda_{t,d_{t}}^{u})$ are full rank.
pgfscope{manualassumption}
Assumption (ref) is a standard assumption in the interactive fixed-effects literature freyberger2017non. Similarly to Assumption (ref), it rules out degeneracies by ensuring that the outcome in each period $Y_{t}(d_{t})$ depends on a distinct linear combination of $X^{*}_{u}$.
We now define the period $t$ conditional choice probability function as $
{h}_{t}(y^{t-1},d^{t},x^{t})\coloneqq \Pr(D_{t}=d_{t}\mid Y^{t-1}=y^{t-1},D^{t-1}=d^{t-1},X^{t}=x^{t})$. In this pure learning environment, the CCP function does not depend on any latent variable and is thus identified directly from the data. As in Section (ref), our identification result (Theorem (ref) below) does not rely on a particular structure imposed on the belief formation process. However, should there be such structure, our identification result would enable identification of the belief formation process. To illustrate this, consider a situation where agents are rational and Bayesian updaters, and where beliefs about $X^{*}_{u}$ at time $t$ are a known function of the information set and the model parameters. That is, there is a known function $s$ such that beliefs are given by $s(Y^{t-1},D^{t-1},X^{t-1},\theta)$, where $\theta$ are the model parameters. In this case, identification of $\theta$ is sufficient for identification of the beliefs.
We now turn to our identification result. Define $f_{\epsilon_t}=\left\{f_{\epsilon_{t}(d)}\colon{d}\in\mathcal{S}(D_{t})\right\}$. Let the model parameter vector be $\theta=\left\{\{\beta_{t},\lambda_{t},f_{\epsilon_t},g_{t},h_{t}\}_{t=1}^{T},\Sigma_{u},F_{X^{*}_{k},X_1}\right\}\in\Theta$. The following theorem states that the previous conditions are sufficient for point identification of $\theta$.
\begin{theorem}
Suppose the distribution of $(Y_{t},D_{t},X_{t})_{t=1}^{T}$ is observed for $T={2p}+1$ and that Assumptions (ref)-(ref) hold. Then $\theta$ is point identified.
pgfscope{theorem}
Key to this result is a simple but powerful insight, namely that, under Assumption (ref), this pure learning model is a model of selection on observables. That is, although assignment probabilities depend on unobserved beliefs over $X^{*}_{}$, they do not depend on the unobserved factor $X^{*}$ itself. It follows that one can control for beliefs at time $t$ by conditioning on prior outcomes, choices and covariates. This, in turn, allows us to express the joint distribution of $(Y^{t},D^{t},X^{t})$, suitably weighted by the assignment probabilities, as a mixture over the potential outcomes $Y^{t}(d_{t})$, conditional on the latent factor $X^{*}$ and exogenous covariates $X$. From here, the arguments of freyberger2017non yield identification of the mixture and component distributions. See Section (ref) for a formal proof.
\begin{remark}[Auxiliary measurements]
In some cases, additional unselected noisy measurements of known heterogeneity factors are available. This includes, in particular, the Armed Services Vocational Aptitude Battery (ASVAB) ability measures that are available in the National Longitudinal Survey of Youth panels. See, among many others, CHN05, CHS10 and AHMR21. With such auxiliary data, sufficient conditions for identification of the distribution of the latent effect are well known in the literature hu2008instrumental,CHS10. If these conditions are satisfied conditional on each $(Y_{t},D_{t},X_{t})_{t=1}^{T}$, then the joint distribution of $\left((Y_{t},D_{t},X_{t})_{t=1}^{T},X^{*}_{k}\right)$ is identified from the auxiliary measurements. From here, one can redefine $X_{t}$ as $(X_{t},X^{*}_{k})$, and Theorem (ref) then yields distribution-free identification of the model with both known and unknown heterogeneity.
pgfscope{remark}
Estimation
We propose to estimate the model parameters via sieve maximum likelihood. We let $W_{i}=(Y_{i,t},D_{i,t},X_{i,t}\colon t=1,\ldots,{T})$ and $\theta^{*}\in\Theta$ be the true value of the parameters. In the following, we focus on the model of Section (ref) with both known and unknown heterogeneity.\footnote{While we focus on this specification, analogous conditions can be derived for the pure learning model considered in Section (ref).} Under the conditions of Theorem (ref), the log-likelihood contribution of $W_{i}=w$ is given by:
align[align omitted — 4,128 chars of source]
Implementation and Monte Carlo simulations
In this section we show how the sieve MLE estimator introduced in Section (ref) can be tractably implemented, and then perform a Monte Carlo experiment illustrating the good finite sample performance of the estimator.
Implementation
We propose an implementation method combining a profiling approach that exploits the parametric components of our model, with a convenient choice of sieve space. Notice first that by integrating out $X^{*}_u$ in Equation (ref), we obtain $\ell(w; \theta) = \log \int \ell^c(w, x_{k}^{*}; \theta^c)dF_{X^{*}_k \mid X_1}(x_{k}^{*}; x_1)$ with
align*[align* omitted — 4,412 chars of source]
Monte Carlo simulations
Next, we present results from Monte Carlo simulations which illustrate the computational tractability and finite-sample performance of the proposed estimator. We focus here for simplicity on a specification with a parametric assignment model. In Appendix (ref) we consider a specification with a nonparametric assignment model, and show that the estimator achieves similar performance.
The data generating process (DGP) used in the simulations is based on the model in Section (ref) with both known and unknown heterogeneity. We include two time-invariant covariates, $X = (X_1, X_2)$, where $X_1$ has a standard normal distribution and $X_2$ as a Bernoulli distribution with equal weights. We assume that $X_1$ and $X_2$ are independent from each other, and from $X^{*}$.
Assignment probabilities are derived from a model in which agents maximize the following expected utility function,
align*[align* omitted — 337,195 chars of source]