Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,219 characters · 23 sections · 96 citation commands
Continuous permanent unobserved heterogeneity in dynamic discrete choice models
In dynamic discrete choice (DDC) analysis, it is common to use mixture models to control for permanent unobserved heterogeneity. For instance, keane1997career,cameron1998life model the observed distribution of schooling and work decisions as a mixture of individuals with varying unobserved abilities, which differ across occupations.
However, the use of mixture models in DDC analysis has limitations. First, existing identification results restrict the permanent unobserved heterogeneity to be either discrete kasahara2009nonparametric or a scalar random variable hu2012nonparametric. In the schooling and work example, this limitation may mean the mixture model does not capture the full richness of ability types and patterns of comparative advantage across occupations.
Second, identification of mixture DDC models depends on having `enough variation' in agent behaviour kasahara2009nonparametric,hu2012nonparametric, a condition that is typically assumed at a high level. In the context of the schooling and work example, `enough variation' might require that agents with different unobserved abilities respond adequately differently to changes in wages. Concretely, `enough variation' is an injectivity condition. To express the condition formally, let $P_t(a,x,b)$ represent the model-implied probability that an agent chooses action $a$ in period $t$ given observed covariates $x$ and persistent unobserved heterogeneity $b$. The `enough variation' assumption states for any signed measure $\mu$ on the support of persistent unobserved heterogeneity
That is, `enough variation' guarantees that distinct distributions of heterogeneity generate distinct average choice behavior in at least one state. An injectivity condition of this style is imposed in the existing indentification literature.\footnote{Specifically, Equation (ref) generalizes the rank condition assumed in Proposition 1 kasahara2009nonparametric, and is a specialization of Assumption 2 hu2012nonparametric .} Yet, despite the crucial role of the injectivity assumption to identification,\footnote{Under some conditions, injectivity is equivalent to identification. See the discussion of Theorem (ref).} there appear to be few results in the literature on whether it holds in a given DDC model. Gaining insights into the conditions under which injectivity holds is particularly significant given that the assumption, as stated in kasahara2009nonparametric, “is not empirically testable from the observed data.” Moreover, verifying injectivity of an integral operator is known to be a challenging problem in general andrews2017examples.
The main contribution of this paper is to propose a general class of DDC models with permanent unobserved heterogeneity that is both continuous and multivariate, and provide low-level conditions for its identification. Applied to the schooling and work example, the class of DDC models in this paper would allow abilities to vary continuously across individuals and to be occupation-specific. I provide sufficient conditions for point identification of all model parameters, including the distribution of agent types (i.e., the distribution of permanent unobserved heterogeneity) and the type-specific choice model. By establishing low-level conditions for identification, the paper provides affirmation of the injectivity assumption for DDC models, demonstrating that it holds at least within one broad class of DDC models.
The paper contains two main results on identification of multinomial DDC models. The first result (Section (ref)) pertains to DDC models with random coefficients. The second (Section (ref)) relates to DDC models with random intercepts. I also prove several extensions to these main results, encompassing both stationary (i.e., infinite horizon) and non-stationary (i.e., finite horizon) DDC models. Furthermore, I show an important implication of the results under the additional restriction that permanent unobserved heterogeneity is discrete --- an assumption that is standard in applied work. In this case, a key modeling decision is the number of agent types (i.e., the number of support points of permanent unobserved heterogeneity),\footnote{In general, only a lower bound on the number of mixture components is identified (e.g., kasahara2009nonparametric) so identification of finite mixture models requires knowledge of an upper bound (e.g., freyberger2018non). See Section (ref) for discussion.} which may be a challenging decision if there is no theoretical guidance on the number of agent types. My identification results imply a solution to this problem: namely, that the number of agent types is identified if it is assumed to be finite.
Within a standard DDC model in the style of rust1987optimal,magnac2002identifying, the low-level conditions for identification can be broadly categorized into two groups. First, I assume a short panel of observations with some continuous variation in the observed covariates, which is a natural prerequisite for nonparametric identification of a continuous latent distribution. Importantly, the results do not require the covariates to have full support, nor place parametric restrictions on the distribution of the permanent unobserved heterogeneity. Second, restrictions on the model primitives are used to ensure injectivity holds. These restrictions have three components: a distributional assumption on the random utility shock, a functional form assumption on the per-period payoffs, and a relevance condition on the covariates.\footnote{This is a key point of departure from the existing identification literature, which allow for more general DDC models at the expense of imposing injectivity at a high level.} The restrictions have the advantage of being low-level and interpretable. For example, the relevance condition can be interpreted as requiring (at least one) covariate to have a non-zero effect on the agent's utility. Moreover, and notably, many of the restrictions are commonly made in the literature. For example, it is common to make distributional assumptions on the random utility shock and functional form assumptions on the per-period payoffs aguirregabiria2010dynamic. In this way, the results of this paper demonstrate that commonly made assumptions impose structure on DDC models that is useful for proving the (otherwise high-level) injectivity condition.
To implement the identification results, I propose a novel estimation method. Existing DDC estimation methods which focus on the parametric case\footnote{In principle, standard DDC models may be semiparametric in the presence of continuous covariates, however, in practice, continuous covariates are often discretized and treated as such for estimation.} aguirregabiria2002swapping,arcidiacono2011conditional do not apply to the model of this paper, as the distribution of unobserved heterogeneity may be an infinite dimensional parameter of interest. Similarly, the computational complexity of DDC models means that immediately available nonparametric methods (such as sieve likelihood estimation) may be impractical. To address these issues I propose a two-step sieve M-estimator, and show it is consistent for the model parameters. I also propose a computationally convenient sieve space based on heckman1984method. Intuitively, the estimator approximates the possibly continuous distribution of permanent unobserved heterogeneity by a discrete distribution. In this setup, the `fixed grid' of support points of the approximating distribution is a tuning parameter of the sieve estimator. Computationally, this estimator is identical to an estimator for a model with finite types, but instead of the number of support points being an identifying assumption, it is simply a tuning parameter.
I illustrate the theory through a simulation exercise and an empirical application based on the labor supply model of altuug1998effect. In this model, agents value consumption and leisure, deciding each period whether to enter the workforce based on expected wages. My identification results allow individual labor productivity to be continuous and consistently estimated from the labor force participation model. The estimates indicate substantial heterogeneity in labor productivity, with a strongly skewed distribution. A counterfactual exercise measures how wages affect labor force participation across the productivity distribution, revealing a highly varied response.
After discussing related literature, I introduce the model and provide one main identification result (Section (ref)). Section (ref) contains the second main identification result (Section (ref)) and other extensions, including to non-stationary DDC problems. Section (ref) proposes the two-step sieve M-estimator and shows its consistency. Section (ref) contains the simulation exercise, and Section (ref) the application.
\paragraph{Related literature.}
This paper is closely related to the literature on point identification of DDC models with persistent unobserved heterogeneity kasahara2009nonparametric,hu2012nonparametric. These papers use a short panel to identify type-specific conditional choice probabilities and the distribution of unobserved heterogeneity via an eigendecomposition of the observed data. As mentioned earlier, these papers consider persistent unobserved heterogeneity that is either discrete kasahara2009nonparametric or a scalar random variable hu2012nonparametric. Relative to these papers, I allow for permanent unobserved heterogeneity that is both continuous and multivariate. As previously mentioned, another important difference is that I provide low-level conditions for the injectivity condition. On the other hand, their approach allows unobserved heterogeneity to enter the model very flexibly, restricted only by certain high-level assumptions.\footnote{However, it is worth noting that hu2012nonparametric do not allow for identification of permanent unobserved heterogeneity from variation in choice behavior alone. Specifically, hu2012nonparametric Assumption 3(ii) requires variation in the state transition by type. To see this, in their notation let $W_{t}=(Y_{t},X_{t})$ be observed and $X_{t}^{*}=X^{*}$ latent, then their equation (11) becomes
and thus their Assumption 3(ii) which requires $k\left(w_{t}, \bar{w}_{t}, w_{t-1}, \bar{w}_{t-1}, x^{*}\right)$ to vary in $x^{*}$ fails if the state transition $f_{X_{t}|X_{t-1},Y_{t-1},X^{*}}$ does not depend on $X^{*}$. williams2019nonparametric also makes this point.} For example, my assumptions rule out type-specific transition functions (e.g., kasahara2009nonparametric) or unobserved heterogeneity that is first-order Markov (e.g., hu2012nonparametric). See also \textcites{williams2019nonparametric,higgins2023identification} as well as the general review compiani2016using.
Several other papers have analyzed persistent unobserved heterogeneity in DDC models from a partial identification perspective. For instance, aguirregabiria2018sufficient focuses on (point) identification of a subvector of the model parameters, treating permanent unobserved heterogeneity as a nuisance parameter. The related aguirregabiria2021identification considers a DDC model with fixed effects. Some general approaches that allow for set identification include chernozhukov2013average,berry2022compiani. Compared to these papers, I provide conditions for point identification of the DDC model. The paper is also related to the large literature on identification of the distribution of continuous unobserved heterogeneity in binary response models. One stream exploits a linear index and full support covariates, while leaving the distribution of random preference shocks unspecified (ichimura1998maximum,lewbel2000semiparametric,gautier2013nonparametric, among others). Relative to these papers, a DDC model yields a non-linear index with additive parametric preference shocks.
The seminonparametric estimator I propose is based on heckman1984method. Similar `fixed grid' estimators have been analyzed for both the parametric and non-dynamic models fox2011simple,fox2016simple, and are increasingly used in applied work nevo2016usage,illanes2019retirement.
{Notation:} For a random variable $X$, $\mathrm{Supp}(X)$ and $f_X$ denote the support and probability density (or mass) function.
I consider a standard single-agent dynamic discrete choice structural model as described in aguirregabiria2010dynamic. In each period $t=1,\ldots,T=\infty$, a single agent observes a vector of state variables $(S_{t},\epsilon_t)$ and chooses an action $A_{t}$ from a finite set of actions $A \equiv\{0,1, \ldots,J\}$ (with $J>0$) to maximize expected utility. I assume $\epsilon_t=(\epsilon_{t,a}:a\in A)$ is independent of $\left(\epsilon_\tau, {A}_\tau, S_{\tau+1}\right)$ for $\tau< t$, and is identically distributed according to $d F_{\epsilon}(e)=\prod_a dF_{\epsilon_a}(e_a)$. In addition, conditional on $(A_t,S_t)=\left({a}_t, s_t\right)$, $S_{t+1}$ is independent of $\left(\epsilon_\tau, {A}_{\tau-1}, S_{\tau-1}\right)$ for $\tau \leq t$, with probability distribution $d F_s\left(s_{t+1} \mid {a}_t, s_t\right)$. It then follows that $\left(S_{t+1}, \epsilon_{t+1}\right)$ is a Markov process with a probability density that satisfies
The agent has a time-separable utility and discounts future payoffs by $\rho \in[0,1)$, where the period $t$ payoff is $u_t(S_t,\epsilon_t,A_t)$. Under these conditions, the agent's choice in time $t$ satisfies
where $v_t$ is the so-called integrated value function:
In this section I present conditions for identification of the distribution of continuous unobserved heterogeneity within the above model. The first assumption imposes restrictions that are standard for stationary DDC models without permanent unobserved heterogeneity.
Assumption (ref) include standard identifying assumptions for DDC models magnac2002identifying,aguirregabiria2010dynamic, including additive separability of the flow utility, that the discount factor is known, a conditional independence assumption, and the outside good. These assumptions are not innocuous --- for example, norets2014semiparametric show that the choice of outside good may affect predicted counterfactual outcomes. Nevertheless, it is standard to assume the unobserved state variables have a known distribution, of which normal and extreme value type I are common choices. It is also common to assume that $S_t$ lies in a compact set, which helps ensure the integrated value function is a bounded function of $S_t$ rust1987optimal,kristensen2019solving.
The next assumption introduces permanent unobserved heterogeneity into the model as an unobserved state variable.
Assumption (ref)(i) states that permanent unobserved heterogeneity enters the model as an unobserved state variable. The restrictions placed on its distribution are mild. First, it allows the distribution to have uncountable support. Intuitively this means there may be infinitely many types of agents.\footnote{One may replace the probability density function in Assumption (ref)(i) with probability mass function and the subsequent results go through with minor modification. That is, the results allow for the typical assumption of finitely many types as a special case.} Second, there may be arbitrary dependence between the initial state variable and permanent unobserved heterogeneity.
Assumption (ref)(i) further imposes that the dimension of the permanent unobserved heterogeneity is equal to the size of the choice set minus one (i.e., $\dim(\beta)=J$). It also requires that the dimension of the observed state variable equals the dimension of the permanent unobserved heterogeneity plus one (i.e., $k=\dim(\beta)+1$). Combined with part (ii), this implies that the model has $J$ variables with action-specific but agent-homogeneous effects via $\gamma_a$, and one variable with action- and agent-specific effects. It is straightforward, however, to allow for additional state variables with agent-homogeneous effects (i.e., $k \ge \dim(\beta)+1$ and $\dim(\beta)=J$); see Remark (ref) for further discussion.
Parts (ii) and (iii) of Assumption (ref) control how permanent unobserved heterogeneity enters the model. Part (ii) states that the permanent unobserved heterogeneity enters the model as a random coefficient in the per-period payoff. Importantly, the continuous $\beta$ is vector-valued, allowing its effect to differ across different choice alternatives. By making the unit and time subscripts explicit in part (ii), i.e., $$ u(s_{i,t},a_{i,t}) = x_{i,t}^\intercal (\beta_{a,i},\gamma_a^\intercal)^\intercal, $$ we see that $\beta_i = (\beta_{1,i},\ldots,\beta_{J,i})^\intercal$ can be viewed as an action-specific random effect associated with the first element of the state variable. For example, if $\beta_a$ represents an agent's ability in occupation $a\in A$, some agents may be high ability in all occupations, other agents may be high in some occupations and low in others. Part (iii) requires that the transition of the state variable not depend on the unobserved state variable. As explained below (Remark (ref)), this assumption enables conditions on the model primitives to be used for identification.
The next condition (Assumption I2(ref)) imposes that the state variable cannot affect payoffs for each choice in a similar fashion. For example, in the binary choice case ($J=1$), the assumption requires that $\gamma_{1}\neq0\in\mathbb{R}$.
Assumption I2(ref) allows the state transition to be a mixture of an absolutely continuous and discrete random variable, but restricts the probability distribution to be a smooth function of the conditioning state variable. In particular, the component probability density and mass functions must be real analytic functions --- that is, functions that have a convergent power series representation. An example of a state transition satisfying Assumption I2(ref) is a mixture of a mass point at $x_{t+1}=0$ and a truncated normal: $F_{x}(x';x,a)=\pi{1}(x'=0)+(1-\pi)F_{+}(x';x,a)$, where $F_{+}(x';x,a)$ is a truncated normal whose mean and variance are real analytic functions of $(x,a)$. Other examples of real analytic functions include polynomials, the logistic function, trigonometric functions, the Gaussian function, in addition to compositions, products and linear combinations of these functions. This class of functions is known to include good approximators to square-integrable functions chen2007large, and can therefore approximate many density functions arbitrarily well.
Define the conditional choice probability (CCP) function $P(a,x,b)$ to be the model implied probability that $A_t=a$ conditional upon $X_t=x$ and $\beta=b$. The first main theorem states that under the above conditions, the integral operator defined by the CCP function is injective.
The injectivity condition in Theorem (ref) is fundamental to identification of mixture models. To explain, consider the simple case that $\beta$ is independent of $X_t$ and that the CCP function is known.\footnote{Since the state transition is identified directly from the data, given the model specified in Assumptions (ref) and (ref), the CCP function is known if $\gamma=\left\{\gamma_a\in\mathbb{R}^{k-1}:a=1,\ldots,J\right\}$ is known.}. In this case, the data satisfies $\Pr(A_t=a\mid X_t=x)=\int P(a,x,b)dF_\beta(b)$ and the only unknown model parameter is $F_\beta$, the distribution of permanent unobserved heterogeneity. Then, supposing (the interior) of $\mathrm{Supp}(X_t)$ is non-empty and open, the injectivity condition is equivalent to identification of the distribution of unobserved heterogeneity: it states that if two distributions $F_\beta$ and $\tilde{F}_\beta$ are observationally equivalent, i.e.,
for almost every $(a,x)\in A\times\mathrm{Supp}(X_t)$, then the two distributions are the same, i.e., $F_\beta=\tilde{F}_\beta$. More generally, the injectivity condition in Theorem (ref) is an example of the injectivity assumption in the measurement error literature hu2008instrumental, with analogs in the context of DDC models \parencites[Proposition 1]{kasahara2009nonparametric}[Assumption 2]{hu2012nonparametric}.
The proof of Theorem (ref) is provided in Appendix (ref). Before presenting an outline, a few comments are in order.
\paragraph{Overview of proof of Theorem (ref).}
Broadly, the argument has two steps: (i) characterizing injectivity in terms of the approximation properties of the CCP function, and (ii) showing that the CCP function satisfies this property.
The characterization of injectivity is developed in two parts. First, I use real analyticity to effectively expand the set of $x$ used to define injectivity. To explain this part, note that the CCP function inherits the smoothness properties of the utility function $u_t$, the state transition $F_x$, and the idiosyncratic shock $F_\epsilon$ (assumed in (ref)(i), I2(ref), and (ref)(v), respectively). In particular, since these are real analytic, the function $x\mapsto P(a,x,b)$ is also real analytic for each $a\in A$, $b\in\mathrm{Supp}(\beta)$. Under the bounded state variable assumption (Assumption (ref)(vi)), this analyticity extends to $\mathbb{R}^{k}$, as shown formally in Lemma (ref). This allows us to use a straightforward extension of stinchcombe1998consistent Theorem 3.8 (formalized in Lemma (ref)\footnote{A heuristic justification of Lemma (ref) is as follows: if two mixture distributions generate the same observed moment function $g(x)\equiv E[Y\mid X=x]$ on any small open set and $x\mapsto g(x)$ is real analytic, then they would also yield the same observed moment function on the full Euclidean space (assuming the relevant objects are well defined). Thus, for identification purposes, observing a non-empty open set is as informative as observing the Euclidean space. The idea is related to the properties of neural networks with limited weights, e.g., stinchcombe1999neural Theorem 2.3 and references therein.}) to characterize the injectivity condition in Theorem (ref) as
Relative to the injectivity condition in Theorem (ref), equation (ref) may be easier to verify since $\mathbb{R}^k\supset\mathcal{X}$.
For the second part, I show in Lemma (ref) that conditions are satisfied to apply an equivalence result from stinchcombe1998consistent that characterizes condition (ref) in terms of the approximation properties of the set of functions $\{ b\mapsto P(a,x,b):(a,x)\in A \times\mathbb{R}^k\}$. Specifically, that this set is dense in square integrable functions on $\mathrm{Supp}(\beta)$. For intuition of this characterization, consider that in the case that $\beta$ has $R<\infty$ support points, the full (row) rank condition is that the collection of vectors $\{\left(P(a,x,b):b=1,\ldots,R\right):(a,x)\in A \times\mathbb{R}^k\}$ span $\mathbb{R}^{R}$.
The final step of the proof is to show this property, as summarized in Lemma (ref):
To prove Lemma (ref), I adapt methods from the classical neural network literature hornik1989multilayer,hornik1993some. Like hornik1989multilayer, the argument is constructive: for a given target function on $\mathrm{Supp}(\beta)$, I find a linear combination of $b\mapsto P(a,x,b)$ that approximates it arbitrarily well. The key part of the construction is to show that for a particular choice of $x\in\mathbb{R}^k$ and $a=0$, $P(a,x,b)$ can approximate the product of one-dimensional step functions in each component of $b\in\mathbb{R}^J$ (i.e., $\prod_{a=1}^J 1\{b_a > l_a\}$ for $l_1,l_2,\ldots,l_J$). It is in this part that the functional form of $u_t$ (Assumption (ref)(i)), rank condition on $\gamma$ (Assumption (ref)(iv)) and extreme value type I assumption (Assumption (ref)(v)) play key roles -- they enable a theoretical guarantee that variation in $x$ can be used to create the step and shift its location in the $\beta$ space. More concretely, Assumption (ref)(iv) guarantees that the image of $\Gamma\equiv(\gamma_1\gamma_2\ldots\gamma_J)$ is $\mathbb{R}^{\dim(\beta)}$, and Assumptions (ref)(i) and (ref)(v) guarantee the linear structure is relevant. A formal proof is in Section (ref).
To invoke Theorem (ref) for identification of the DDC model, we require the support of the state variable to contain an open set:
Assumption (ref) places restrictions on the support of the observed state variable $X_t\in\mathbb{R}^k$. Part (i) requires that the support of the observed state variable contains an open set. Part (ii) requires that the supports contain $k$ linearly independent elements, a mild rank condition which is standard in linear models. As discussed in Example (ref), Assumption (ref) allows for renewal models like rust1987optimal. However, it rules out lagged dependent variables, that is, when $X_t$ contains the lagged choice $A_{t-1}$. This would rule out, for example, a firm entry problem where the current period's entry decision $A_{t}$ depends on whether the firm is currently active ($A_{t-1}$). In particular, lagged dependent variables contradict Assumption (ref)(ii) since $\mathrm{Supp}(X_4\mid{X_3}=x,A_3=a)$ and $\mathrm{Supp}(X_4\mid{X_3}=x,A_3=\tilde{a})$ are disjoint for $a\neq\tilde{a}$.\footnote{ Although the open set assumption (ref)(i) also rules out purely discrete variables, as discussed in Remark (ref), these can be allowed with minor notational changes. In this case, Assumption (ref)(i) is relaxed but Assumption (ref)(ii) is unchanged. See Section (ref) for a technical statement.} However, unlike some results in the literature, Assumption (ref) does not require that the support be `rectangular'\footnote{For example, this is Assumption 1(c)-(e) used in kasahara2009nonparametric Propositions 1-9 and subsequently relaxed in Propositions 10 and 11.} --- which requires that, starting from any sequence of choices and past state variables, any state can be reached (i.e., for all $t$ and $(a,x)\in \mathrm{Supp}(A_t,X_t)$, $\mathrm{Supp}(X_{t+1}|X_t=x,A_t=a)=\mathrm{Supp}(X_{t+1})=\mathrm{Supp}(X_1)$).
The model parameters are ($F_{x},\gamma,f_{\beta|X_{1}}$): the state transition, the homogeneous payoff parameter, and the conditional distribution of permanent unobserved heterogeneity. As the state transition is identified by direct observation, the following result handles the remaining parameters:
Theorem (ref) is established via a decomposition argument hu2008instrumental,freyberger2018non. The model structure imposed by Assumptions (ref) and (ref) implies the following `factorization equation' representation of the weighted distribution of $(X_{t},A_{t})_{t=1}^{T}$ kasahara2009nonparametric:
which is guaranteed to exist under the support condition in Assumption (ref)(i). In the factorization equation we see the role of Assumption (ref)(iii): since the state transition does not depend on $\beta$, it can be passed through the integral.\footnote{Related homogeneity assumptions can also lead to weighting approaches in other models, such as \textcites[Chapter 21]{hernan2020chapman}[Section 6]{bonhomme2023identification}.} Then, by invoking the injectivity result in Theorem (ref), the representation can be used to express the CCP function $P(a,x,b)$ as the eigenfunction of a particular eigendecomposition.\footnote{This reasoning also suggests that, by directly assuming the injectivity condition in Theorem 1, a related identificaton result may hold under weaker conditions on the model (i.e., weaker versions of Assumptions (ref) and (ref)). See kasahara2009nonparametric, Remark 2.} I then show that the eigendecomposition is unique, which delivers identification of $\gamma \in \mathbb{R}^k$: The argument is related to identification of dynamic discrete choice models without unobserved heterogeneity (e.g., bajari2015identification), with Assumption (ref)(ii) playing a central role. Knowledge of $\gamma$ is then used in combination with the factorization equation and injectivity to identify $f_{\beta|X_1}$. The formal proof is in Section (ref).
In this section I provide identification results for a number of variations on the model in Section (ref). Sections (ref) and (ref) consider finite-horizon environments in which the agent's decision rule may vary across periods. Section (ref) focuses on the case where the terminal period is observed, allowing identification of models with random intercepts. Section (ref) addresses the case where the decision horizon extends beyond the observed sample. It provides two solutions: imposing out-of-sample restrictions or exploiting finite dependence. Section (ref) returns to the infinite-horizon setting and allows for random intercepts under additional assumptions on the transition process. Finally, Section (ref) shows that the number of agent types is identified in models with discrete unobserved heterogeneity.
In many contexts, the agent's decision rule may change between periods: for example, if the agent has a finite time-horizon, or if the state variables are subject to structural breaks. In these cases, it is natural to allow the per-period utility function and state transitions to be non-stationary, i.e., to be time-dependent. In this section I consider a finite horizon dynamic discrete choice model in which the terminal decision period is observed. For example, in a model of retirement from the labor force rust1997social, we may eventually observe all individuals retire. Similarly, in a model of educational attainment, we may observe all individuals reach a terminal state heckman2018returns. By definition, the decision-maker has no strategic influence over future utility flows to consider in the terminal period and thus a different proof strategy is adopted. This argument allows for identification of random intercepts, which was not the case in Section (ref).
I begin by adapting Assumptions (ref) and (ref) to the non-stationary context. In particular, by allowing the flow utility and state transition to be time-dependent.
Assumption (ref) states that permanent unobserved heterogeneity enters the model as a state variable. The restrictions are weaker than those in the infinite horizon model (Assumption (ref)). First, the permanent unobserved heterogeneity can include a random intercept. Second, there may be multiple random coefficients for each option, whereas in Section (ref) the model was limited one action-specific random coefficient (i.e., $p=1$). This relaxation is possible due to the relatively simple structure of the terminal period CCP function. As was the case for the infinite horizon model, the support of permanent unobserved heterogeneity may be finite, but it need not be (see footnote (ref)). Like Assumption I2(ref), Assumption F2(ref) imposes that the state variable cannot affect payoffs for each choice in a similar fashion. Since identification is attained from the terminal period, we place weaker restrictions on the transition $F_{x_t}$ relative to Assumption I2(ref).
To describe the injectivity result for the finite horizon model, denote the CCP function $P_t(a,x,b)$ and let $T$ denote the decision horizon of the agent.
The proof of Theorem (ref) is contained in Section (ref). The proof logic is rather different to Theorem (ref): to show Theorem (ref), I show the implication directly by demonstrating that $\int P_T(a,x,b)d\mu(b)=0$ implies that the induced measure of $P_T(a,x,\beta)$ is zero. The linear utility function and distributional assumption on $F_\epsilon$ are particularly useful for this. The result then follows from masten2018random, Lemma 1.
As for the time stationary model, we require further restrictions on the state variable $X_t$ for identification of the DDC model. First, Assumption (ref) requires there be some continuous variation in $X_T$ after conditioning upon each history of actions and state variables.
To introduce the final assumption, let $\gamma_t=\left\{\gamma_{t,a}:a=1,\ldots,J\right\}$ and define $S_T\equiv\mathrm{Supp}\left(X_{T} \mid A_{T-1}=a_{t-1},X_{T-1}=x_{t-1},\dots,A_1=a_1,X_1=x_1\right)$ and let $E\subset S_{T}\times {A}$, $P_T(a;x,b,\gamma)$ be the model implied probability of $A_T=a$ conditional upon $X_T=x$ evaluated at $\beta=b$ and $\gamma_T=\gamma$, and $\mathcal{L}_\mathcal{A}$ be the set of bounded functions on $\mathcal{A}$. Then define the operator
Denote $(L^{E,\gamma}_{T,\beta})^{-1}$ as the left inverse of $L^{E,\gamma}_{T,\beta}$.
This high-level condition ensures that the parameter $\gamma_{T}$ can be identified without knowledge of the distribution of unobserved heterogeneity. A few comments on Assumption (ref) are in order. First, given Theorem (ref), Assumptions (ref)-(ref) imply that, for any $E$ containing a non-empty open set, $L^{E,\gamma}_{T,\beta}$ is injective so that $L^{E,\gamma,\tilde{E},\tilde{\gamma}}_{T,\beta}$ exists. Second, the condition is stated in terms of observed objects, and thus the operator defined in Assumption (ref) is identified by direct observation. Third, should Assumption (ref) not hold, I show in an appendix (Lemma (ref)) that under Assumptions (ref)-(ref) and a scale restriction on $\gamma_{T}$, that $\gamma_{T}$ and the distribution of unobserved heterogeneity are identified.
Finally, the condition can be related to the high-level necessary conditions for identification of a common parameter in discrete choice panel data given in johnson2004identification,chamberlain2010binary. To describe their result, fix $x\equiv(x_{1},x_{2},\dots,x_{T})$ and for convenience let $A=\{0,1\}$ and $\gamma$ be time-invariant. Let $p(b;x,\gamma)$ be the length $2^{T}$ vector of choice probabilities $\left\{\prod_{t=1}^{T}P_{t}(a_{t},x_{t},b;\gamma):(a_{t})_{t=1}^{T}\in \{0,1\}^{T}\setminus\{0_{T}\}\right\}$ in the $(2^{T}-1)$-dimensional hypercube. johnson2004identification states that the common parameter $\gamma$ will not be identified if the set $\{p(b;x,\gamma):b\in{S}_{\beta}\}$ does not lie in a hyperplane for some $x$. For the static binary choice model with $T=2$, chamberlain2010binary shows that the hyperplane restriction is satisfied if and only if the unobserved state variables are i.i.d. extreme-value type I. Given the remarkable result of chamberlain2010binary, one may conjecture that the $T=2$ dynamic binary choice model does not satisfy johnson2004identification's condition and therefore $\gamma$ is not identified. If this is the case, then $\forall{x_{2}}\in\mathrm{Supp}{(X_{2})}$ and $\gamma\neq\tilde{\gamma}$, there exist some $f_{\beta|X_{1}X_{2}}\neq\tilde{f}_{\beta|X_{1}X_{2}}$ such that
where the distribution of unobserved heterogeneity $f_{\beta|X_{1}X_{2}}$ is allowed to depend on $x_{2}$ as in johnson2004identification,chamberlain2010binary. If the distribution is restricted to be the same for all $x_{2}\in\mathrm{Supp}{(X_{2})}$, the above condition implies that for each $\gamma\neq\tilde{\gamma}$, ${x_{2}}\in\mathrm{Supp}{(X_{2})}$, then there are some $f_{\beta|X_{1}},\tilde{f}_{\beta|X_{1}}$ that satisfy
However, since the distribution of unobserved heterogeneity is required to be the same for all $x_{2}$, there may be some other $\tilde{x}_{2}\in\mathrm{Supp}{(X_{2})}$ such that
Let $E,\tilde{E}$ be neighborhoods of $(x_{2},\tilde{x}_{2})$, respectively. In the proof to Theorem (ref) it is shown that, without knowledge of $f_{\beta|X_{1}}$ or $\tilde{f}_{\beta|X_{1}}$, there does exist such an $\tilde{x}_{2}$ if the operator defined in equation ((ref)) is injective. This can be viewed as a partial converse to johnson2004identification's high-level condition: in that case, without knowledge of $f_{\beta|X_{1}}$ or $\tilde{f}_{\beta|X_{1}}$, one can show there does not exist such an $\tilde{x}_{2}$ if their `rank' condition does not apply. In principle, the logic of Assumption (ref) can be extended to the general discrete choice panel model of johnson2004identification, if the distribution of unobserved heterogeneity is required to be independent of covariates. To state the theorem denote $\gamma=\{\gamma_t:t=1,\ldots,T\}$.
Section (ref) contains the proof of Theorem (ref).
In many empirical settings, the decision horizon of the agent extends beyond the period of observation. For example, a worker's labor force participation decisions may not be observed for their entire working life. This poses an issue for identification since in-sample decisions reflect payoff parameters for both in- and out-of-sample time periods. This section provides two solutions for this issue. The first approach is to impose restrictions on out-of-sample payoffs. Section (ref) adopts this approach and shows that the model without random intercepts is identified.
The second approach is to use a property of the state transition known as `finite dependence', which occurs if multiple sequences of actions leads to the same distribution of the state variable arcidiacono2011practical. Finite dependence limits the number of out-of-sample time periods that affect in-sample decisions. Section (ref) considers a model that exhibits finite dependence, and shows a binary choice model with random coefficients is identified.
For both approaches, I consider a model that satisfies the following condition:
Analagously to the Sections (ref) and (ref), Assumptions (ref) and (ref) are sufficient for injectivity of the integral operator with kernel function $P_t(a,x,b)$.
Let $T$ denote the final observed period and $T_1>T$ denote the final decision period of the agent. Since we do not observe behavior in periods $(T+1,\dots,T_{1})$, the following restriction is placed on out-of-sample behavior:
With these assumptions and a support condition on $X_{t}$ related to Assumption (ref), identification results follows as a Corollary of Theorem (ref). The proof is found in Section (ref).
A DDC model exhibits finite dependence if there are multiple sequences of actions that yield the same distribution over the state variable. Finite dependence is useful for estimation as it allows the continuation value to be expressed in terms of CCPs arcidiacono2011practical. This fact also makes finite dependence useful for identification in models without permanent unobserved heterogeneity, as it reduces the number of periods of out-of-sample behavior that must be assumed known arcidiacono2020identifying.
In this section I show a similar feature is present for models with continuous permanent unobserved heterogeneity. In particular, I assume the transition function exhibits a special case of finite-dependence: the renewal action. The canonical example of renewal is machine replacement, but models of turnover and job matching also display this pattern arcidiacono2020identifying. This idea is formalized in the next assumption, which, in addition to a support condition, is sufficent for identification.
Section (ref) contains the proof to Corollary (ref), whose substance is adapted from the proof of Theorem (ref).
This section considers identification of an infinite-horizon DDC model with random intercepts. It shows point identification can be attained under an additional restriction on the state transition. Specifically, there must be some point in the support of $X_{t}$ for which the state transition is not choice dependent. For instance, the machine replacement model of kasahara2009nonparametric displays this property. Before introducing the restriction on the state transition, the next assumption states that the permanent unobserved heterogeneity enters the model as a random intercept:
The next assumption strengthens Assumption (ref) by requiring the state transition to be constant across choices:
The proof to Corollary (ref) is contained in Section (ref). It follows from the proofs of Theorems (ref) and (ref).
In the existing DDC literature, it is common to assume permanent unobserved heterogeneity is discrete. When this assumption is made, a key parameter is the number of support points of permanent unobserved heterogeneity. In practice, it is common to assume the number of support points is known, although there are methods to identify a lower bound on the number of support points kasahara2009nonparametric,kasahara2014non,kwon2019estimation which have been applied in economics igami2016unobserved. However, in general, these methods can only identify the number of support points if an upper bound is known. This is because there is no guarantee a priori that there is enough variation in the data and structure on the model to to identify any arbitrarily large number of types. Intuitively, the population likelihood may be flat as a mixture component is added, but this may be because the initial likelihood had the true number of mixture components or because the models with and without an additional mixture component are observationally equivalent. Technically, this issue can be resolved by imposing an injectivity condition, i.e., a rank assumption on an unobserved matrix \parencites[Proposition 3]{kasahara2009nonparametric}[Assumption 2.1]{kwon2019estimation}.
The purpose of this section is to show the models of Theorem (ref) and Corollary (ref) satisfy a condition equivalent to kwon2019estimation when the distribution of unobserved heterogeneity is discrete. This means the number of types is identified, without knowledge of an upper bound on the number of types.
The proof to Corollary (ref) is found in Section (ref). The result means that the techniques of kasahara2014non,kwon2019estimation can be used to consistently estimate the number of types should the applied econometrician wish to maintain the standard assumption that permanent unobserved heterogeneity is discrete.\footnote{The model in Corollary (ref) can be directly adapted to the general frameworks of kasahara2014non,kwon2019estimation. See, in particular, kwon2019estimation Equation 2.1 and kasahara2014non Equation 2.} These techniques also give rise to valid hypothesis tests regarding the number of types, including testing the null of type degeneracy (that is, $R=1$). Broadly speaking, these estimators consist of forming a matrix of observed choice probabilities with values of $X_{3}$ varying over the rows, and $X_{2}$ over the columns. Corollary (ref) means that, at the population level, the rank of the matrix equals the true number of types.
This section considers consistent estimation of the model parameters in a short panel. The distribution of $Y\equiv(A_{t},X_{t})_{t=1}^{T}$ can be written as
where $F_{\beta|X_{1}}(b,x_{1})$ is the cumulative distribution function of $\beta$ conditional upon $X_{1}=x_{1}$, $F_{x_{1}}$ is the marginal distribution of $X_{1}$ and the dependence of the CCPs on $(\gamma,F_x)$ is made explicit. I propose two-step sieve M-estimation based on the above expression. The first step consists of estimating the state transitions and marginal distribution of the initial state, $F_{x}=\{F_{x_t}:t=1,\ldots,T\}$. The second step consists of forming the pseudo-likelihood function using the fact that the CCPs $P_{t}$ are known up to the state transition and payoff parameter $(F_{x},\gamma)$, and using sieve M-estimation methods to estimate ($\gamma,F_{\beta|X_{1}}$).
It is of course possible to estimate the model in a single step as a sieve maximum likelihood problem. The advantage of the proposed two-step approach is computational: by treating $F_{x}$ as fixed in the second step, computationally advantageous methods for approximating the value function may be used, such as kristensen2019solving.
Although I show consistency for a general sieve space (Section (ref)), this may be computationally burdensome to implement, since estimation requires computing the CCPs for every point in the support of the sieve. To circumvent this issue, I suggest a `fixed grid' estimator heckman1984method which reduces the computational burden by having a finite number of support points (Section (ref)). Given these results, the practioner's decision to approximate $F_{\beta|X_1}$ by a continuous function or by the `fixed grid' can be viewed as a choice of tuning parameter, rather than an identifying assumption.
In this section, I focus on estimating the cumulative distribution function of $\beta$. While it would be possible to present conditions for consistent estimation of the density function, smoothness restrictions would rule out the possibility that the type distribution has discrete support, which is the standard assumption in the literature. Moreover, focusing on the distribution function of $\beta$ enables the choice of the piecewise constant sieve space described in Section (ref), which has particular computational advantages.
As a final comment, in practice there will be an approximation error in the evaluation of the CCPs. This problem is inherent to dynamic discrete processes with large state spaces, and has received significant attention in the recent literature rust2008dynamic,kristensen2019solving. I assume away the effect of these errors on estimation --- that is, that the approximation error is negligible relative to sampling error. In principle, the results of kristensen2019solving could be used to explicitly consider the effect of value function approximation error on estimation, though I do not pursue this here. Of course, the approximation error can be made arbitrarily small at increased computational cost.
In this section, I briefly outline the two-step sieve M-estimator and present the general consistency result. Denote the true parameters as $\theta_{0}=(F_{x},\gamma,F_{\beta|X_{1}})\in\Theta=\mathcal{F}\times\Gamma\times\mathcal{M}$, where $\mathcal{F}$ is the space of state transitions, $\Gamma\subseteq\mathbb{R}^{\dim\gamma}$, and $\mathcal{M}$ is the space of distribution functions on $\mathrm{Supp}(\beta)$ conditional upon $x\in\mathrm{Supp}(X_1)$. The first step consists of forming a consistent estimator $\hat{F}_{x}$ for the state transition $F_{x}$. Since the state transition is directly observed, standard non-parametric methods are available. For the second step, the log-likelihood contribution of the $i$th observation is
where $P_{t}(a,x,b;\hat{F}_{x},\gamma)$ is the model implied probability of observing choice $a$ in period $t$ conditional upon state $x$ and permanent unobserved heterogeneity $b$, evaluated at the first-step estimate $\hat{F}_{x}$ and candidate parameter $\gamma$. Given a sieve space $\mathcal{M}_{n}$, which approximates $\mathcal{M}$ arbitrarily well for large $n$, the second step estimator is defined as
The following result states that under standard regularity conditions, the estimator is consistent.
The full statement of Theorem (ref) and its proof are contained in Appendix (ref).
In this section I propose a particular choice of sieve which has the advantage of being simple to implement: the first-order monotone spline sieve. This is a popular choice of sieve for seminonparametric models, see for example heckman1984method,chen2007large,fox2016simple. To define the sieve, let $\mathcal{B}_{n}=\{b_j:j=1,\ldots,B(n)\}$ be a set of knots that partition $\mathrm{Supp}(\beta)$ and $\mathcal{X}_{n}=\{\mathcal{X}_{k,n} :k=1,\ldots X(n)\}$ be a partition of $\mathrm{Supp}(X_1)$. The sieve space $\mathcal{M}_{n}$ is defined as follows: {
} where the sets $(\mathcal{B}_{n},\mathcal{X}_{n})$ are tuning parameters. For a given choice of tuning parameters, an element of $\mathcal{M}_{n}$ consists of $X(n)$ piecewise constant (step) functions in $b$, indexed by the partition cells $\mathcal{X}_n$, each such function having jumps of size $P_{j,k}$ at point $b_{j}$. The computational advantages of this sieve are clear: to find the supremum in ((ref)), for each $x_{1}$, the CCP functions need only be evaluated for the values $b_{j}\in\mathcal{B}_{n}$. This would not be the case if the sieve space consisted of functions that were continuous in $b$.
A theoretical advantage of this sieve space is that many of the high-level conditions for consistency are attained as long as the number of knots does not grow too fast. See Appendix (ref) for details.
To implement the estimator, the number and location of grid points must be chosen. For consistency, it is enough that $B(n)X(n)\log(B(n)X(n))=o(n)$ and that the grid points become dense in the support of $(\beta,X_1)$. In principle, convergence rates for this estimator could be derived to determine optimal growth rates for $B(n),X(n)$.
For computation, it may be attractive to use profiling. In particular, to form ($\hat{\gamma},\hat{F}_{\beta|X_{1}}$), fix $\gamma$ and let
For $\mathcal{M}_n$ as in equation (ref), this is a convex optimization problem, with a unique global optimum that can be computed efficiently (e.g., koenker2014convex). The profile estimator is formed as
This section investigates the estimator of Section (ref) in a Monte Carlo simulation. The main goals of this section are twofold: first, to explore the finite sample performance of the estimator; and, second, to provide empirical support for the asymptotic results of Section (ref). I simulate data using a simple labor force participation model based on altuug1998effect, which also acts as a basis for the empirical illustration in Section (ref).
In each period, each individual decides whether or not to enter the labor force, upon observation of the state variable. Thus $A=\{0,1\}$, with $a_{t}=1$ representing an individual decision to enter the labor force at time $t$. The period payoff from entering the labor market depends on the observed state variable $x_t=(x_{t,1},x_{t,2})^\intercal\in\mathbb{R}^2$, the entry-specific shock $\epsilon_{t,1}$, and individual-specific labor productivity $\beta$ as follows:
Following the model of altuug1998effect, $x_{t,1}$ can be interpreted as an average consumption value (see Section (ref) for details) and $x_{t,2}$ is equal to the income of the primary earner in the household. The period payoff from not entering is $\epsilon_{t,0}$. The random preference shock $\epsilon_{t,a}$ is assumed to be distributed extreme value type I and independent across time, choices and agents. Further, the agents' time horizon is assumed to be infinite with exponential discount factor 0.9. In addition, I assume that $\beta$ is independent of $X_1$ and consider three different choices for its distribution. In DGP 1, $\beta$ follows a mixture of three truncated normal distributions:
where $\mathcal{N}_{tr}(\mu,\sigma)$ is the truncated normal distribution with parameters ($\mu,\sigma$), minimum value $0$ and maximum value $50$. In DGPs 2 and 3, I assume $\beta$ follows a uniform distribution on $[0,5]$ and $\{1,2.5,4\}$, respectively. I assume that the first period observed state variable is drawn independently from the uniform distribution on $[0,4]\times[0,4]$, and that $F_x(x'|x,a)=F_1(x'_1|x,a)F_2(x'_2|x,a)$, where $F_1$ and $F_2$ are truncated normal distributions with means $x_1/(a+2)$ and $(x_1+x_2)/(a+2)$ respectively, unit standard deviations and truncated to the interval $[0,4]$. I set $\gamma=2$.
The simulation results are the average of $1{,}000$ i.i.d. datasets $(a_{i,t},x_{i,t}:t=1,\dots,8)_{i=1}^{n}$ drawn from this model.\footnote{In practice, the state space $[0,4]\times[0,4]$ and support of $\beta$ are discretized to solve the model. The discrete state space and support of $\beta$ have 400 and 1{,}000 points of support respectively.} Results are presented for four sample sizes: $n=100,500,1{,}000$, and $10{,}000$. For estimation I choose the number of grid points equal to $4n^{1/4}$ (i.e., $13, 19, 23$ and $40$), which satisfies the rate conditions required for Theorem (ref), and consider a grid of equally spaced points between 0 and 6. For estimation, I assume knowledge of the discount factor, the state transition $F_x$, and impose that the initial state is independent of $\beta$, leaving the unknown parameters as $(\gamma,F_\beta)$, the homogeneous effect of spousal income and the distribution of labor productivity.
Table (ref) presents results for the estimator of $(\gamma,F_\beta)$, in addition to computation times. First consider results for $\gamma$. Here, empirical variance is significantly larger than empirical bias, which diminishes with sample size. Scaled empirical mean squared error is largely flat across sample sizes. In terms of computational burden, the fixed grid estimator takes around 30 seconds to run for the smaller sample sizes, though it takes around 2 minutes for $n=10{,}000$.
Turning to results for the estimation of $F_{\beta}$, both measures of integrated error diminish with sample size.\footnote{Integrated absolute and squared error for simulation run $m$ with estimate $\hat{F}_{\beta,m}$ is $\int|\hat{F}_{\beta,m}(b)-F_{\beta}(b)|db$ and $\int\left(\hat{F}_{\beta,m}(b)-F_{\beta}(b)\right)^{2}db$, respectively.} The number of grid points increases slowly with sample size --- indeed slower than the growth of the number of support points selected by the estimator. For example, in DGP 1 for $n=100$, on average 5.2 points are selected. This increases to 10.1 for the large sample size. This pattern is broadly similar to previous simulation results for a parametric variant of this estimator fox2011simple. The number of support points chosen is similar between DGP 1 and DGP 2, but fewer points are chosen in the DGP with discrete types (DGP 3). Additional simulation results are presented in Appendix (ref).
This section revisits the female labor supply model of altuug1998effect. I combine the life-cycle model of altuug1998effect with the identification results of Section (ref) to estimate the distribution of labor productivity from data on labor force participation and perform a counterfactual exercise to measure how the response to a wage increase varies across the productivity distribution.
altuug1998effect introduces a framework to understand female labor supply that takes into account aggregate shocks and time non-separable preferences. In their model, agents gain utility from consumption and leisure. Under their specification of consumption and Pareto optimality, individual $i$ at time $t$ generates utility from consumption as:
The term $(\eta_{i}\lambda_{t})$ is the shadow value of consumption, which is estimated from data on consumption. The term ($\beta_{i}\omega_{t}\exp(\gamma_{3}'x_{Wit})l_{i,t}$) represents an individual's predicted earnings,\footnote{For clarity, in this section I will denote permanent unobserved heterogeneity as $\beta_i$.} which is equal to the amount of time they spend working conditional on participating, $l_{i,t}$, multiplied by their marginal product. The individual-specific marginal product of labor consists of unobserved aggregate and individual productivity effects ($\omega_{t},\beta_{i}$) in addition to a component that depends on covariates $x_{Wi,t}$. These terms are estimated from the wage equation, which is as follows:
altuug1998effect consider two estimators for the individual-specific productivity $\beta_{i}$. First, they use the fixed effects estimator from the wage equation above. Of course, in the asymptotic framework considered in this paper where $n$ is large but $T$ is fixed, this estimator is subject to the incidental parameters problem and is not consistent in general. For the second estimator, the authors assume that the fixed effect is an unknown function of observables, and then estimate that function non-parametrically. The observed variables consists of demographic data such as race, marital status and education levels. This estimator will be inconsistent if the set of observed variables is misspecified---that is, if individual productivity cannot be written as a function of observed data. The identification results of Section (ref) obviate the need to estimate individual-specific productivity from the wage equation. Instead, $\beta_{i}$ can be interpreted as a random coefficient in the discrete choice model of labor force participation elaborated below.
Suppose the per-period payoff from entering the labor market for individual of type $\beta_i$ is:
with $x_{i,t}=({z}_{i,t},1,\text{hinc}_{i,t},\text{age}_{i,t},\text{kids}_{i,t},\text{educ}_{i,t})$. Here ${z}_{i,t}$ is constructed following the approach of altuug1998effect, that is ${z}_{i,t}=\eta_{i}\lambda_{t}\omega_{t}\exp(\gamma_{3}^\intercal x_{Wi,t}){l}_{i,t}$ where each component is estimated from the consumption/wage regressions described above (see Appendix (ref) for details). The remaining components of $x_{i,t}$ are, respectively, a constant term, annual head-of-household income, an age variable, whether there is a child in the household, and an education variable.\footnote{For simplicity, the age and education variable are dummies indicating whether the individual is over 35 year old and whether they have completed a college degree, respectively. In the DDC model, I assume that college degree status is constant over time (which is true for 97.5% of individuals).}
Relative to the DDC model of participation in altuug1998effect, $\beta_{i}$ is treated as an unobserved random variable. In their model $\beta_{i}$ is replaced by fixed effect estimates and treated as a known constant in their DDC model. Like altuug1998effect, I make the outside good assumption and assume that $\epsilon_{i,t,a}$ is distributed extreme value type I and independent across agents, time and actions. For simplicity, I assume that the agents' time horizon is infinite and that the exponential discount factor is 0.9 and known to the econometrician.
As in altuug1998effect, the labor force participation model is estimated using a subset of data from the PSID. The data construction is described in Appendix (ref), and closely follows the details in altuug1998effect. The final data set contains 3084 individuals, each of whom have between four and ten panel observations, with an average close to eight.
I estimate the model using the two-step estimator described in Section (ref). The first step consists of estimating the state transition $F_x(x'|x,a)$. To simplify this step, I assume that, conditional upon $A=a$, (i) $X'-X$ is independent of $X$ and (ii) the components of $X'-X$ are mutually independent. Then, I estimate the densities $Z'-Z \mid A=a$ and $Hinc'-Hinc \mid A=a$ for each $a=0,1$ via the kernel density estimator with the Gaussian kernel and rule-of-thumb bandwidth.\footnote{The bandwidth is $1.06 \mathrm{std}\left[\sum_{i=1}^{n}\sum_{t=1}^{T_i-1}1\{A_{i,t}=a\}(y_{i,t+1}-y_{i,t})\right] (\sum_{i=1}^{n}\sum_{t=1}^{T_i-1}1\{A_{i,t}=a\})^{-1/5}$, for ${y=Z,Hinc}$ where $\mathrm{std}$ denotes standard deviation and $T_i$ is the panel length of observation $i$.} I note that these restrictions satisfy the real analyticity requirement Assumption (ref)(iv) (more precisely, its generalization in Section (ref)).\footnote{This model has additional state variables with homogeneous effects (i.e., $k>\dim(\beta)+1$ where $k=\dim(X_t)=6$); as discussed in Remark (ref), the conditions of Section (ref) must be adapted accordingly. A formal statement of these conditions is provided in Section (ref).} To see this, observe that, under the above specifications, the state transition cumulative distribution function is $$ \Pr(X'\le x \mid X=x,A=a)=\Phi\left(\frac{z'-z-\mu_{1,a}}{\sigma_{1,a}}\right)\Phi\left(\frac{hinc'-hinc-\mu_{2,a}}{\sigma_{2,a}}\right)h(d',d,a), $$ where $\Phi$ is the standard normal cumulative distribution function, $d=(educ,age,kids)$ are the discrete variables, and $h,\mu_{1,a},\sigma_{1,a},\mu_{2,a},\sigma_{2,a}$ are unknown parameters to be estimated. Thus, for each fixed $(x',d,a)$, the state transition is a bounded real analytic function of $(z,hinc)$ that is supported on $\mathbb{R}^2$. Given this discussion and model assumptions described above, the two sufficient conditions for injectivity are satisfied; then, for identification, I impose the required support condition, which appears plausible given both $Z_{i,t}$ and $Hinc_{i,t}$ are continuous random variables.
The second step requires specifying a sieve space for $\beta_i$. The step-wise constant sieve space of Section (ref) is adopted, with the number and location of the knots as tuning parameters. For simplicity, $\beta_i$ is assumed independent of $X_{i,1}$. Consistent with the simulation design, the number of knots is set to $4n^{1/4} \approx 30$, placed uniformly between 0 and 15. The lower bound of 0 reflects a natural restriction on labor productivity, while the upper bound of 15 is sufficiently large that, for reasonable parameter values, the conditional choice probability is close to 1.
I implement the estimator using the profiling approach described in Section (ref).\footnote{The remaining tuning parameter is the starting value of $\gamma$, which is set as the estimates from the same estimator with five knots, equally spaced between 0 and 15. That estimator is itself initialized with the estimates from the parametric model (i.e., under the assumption that $\beta_i$ is degenerate with unknown support).} The model solutions required in the second step are obtained following kristensen2019solving. Inference is conducted using the standard bootstrap, see Appendix (ref) for evidence on its performance in a simulation exercise. Additional results on the fit of the estimated model are provided in Appendix (ref).
Table (ref) presents point estimates of the finite dimensional parameter $\gamma$ alongside bootstrapped standard errors. Estimates indicate that utility from working increases with education, but decreases with head-of-household income and age. Having children in the household is estimated to have a negligible effect on utility from working.
Figure (ref) presents the estimated distribution of $\beta_i$ from the fixed grid estimator. The estimated distribution has 21 points of support, with mean 3.11, median 3.11, standard deviation 1.35, skewness 2.39 and kurtosis 15.83, indicating substantial heterogeneity in labor productivity.\footnote{For comparison, in a model where $\beta_i$ is assumed to have three unknown points of support and estimated using the method of arcidiacono2003finite, the estimated distribution has mean 2.93, median 2.56, standard deviation 0.91, skewness -0.57 and kurtosis 2.29. See Appendix (ref).}
In this section, I conduct a counterfactual exercise to measure how wages affect labor market participation across the skill distribution. The counterfactual considered is where the agent's expected wage received from working (i.e., under $A_{i,t}=1$) is increased by $x\%$ over its status quo value, for $x=5,10,15,20,25$, holding all else fixed.\footnote{In the model described above, agent $i$'s expected wage from working in period $t$ is $\omega_t\beta_i\exp(\gamma_3^\intercal x_{Wi,t})$.} For each counterfactual wage change of $x\%$, I draw $(\beta^{(x)}_m,X^{(x)}_{m,t},A^{(x)}_{m,t}:t=1,\ldots,T)_{m=1}^{M}$ for $M=1{,}000{,}000$ and $T=5$ from the estimated model,\footnote{Each simulated panel $m=1,2,\ldots,M$ is drawn independently as follows. First, $\beta_m$ is drawn from the estimated distribution $\hat{F}_\beta$ and $X_{m,1}$ is drawn from the empirical distribution of $X_{i,1}$. Then the conditional choice probability $P(1,X_{m,1},\beta_m;\hat{F}_x,\hat{\gamma})$ is computed and used to draw $A_{m,1}$. Next, $X_{m,2}$ is set as $X_{m,1}+\xi_{A_{m,1}}$ where $\xi_a$ is drawn uniformly from the empirical distribution of $X'-X \mid A=a$, with the draw truncated to respect the empirical supports. $A_{m,2}$ and $(X_{m,t},A_{m,t})$ for $t=3,\ldots,T$ are drawn analogously.} and report the average labor market participation rate for six different quantiles of $\beta$.
Table (ref) displays the results of this counterfactual exercise. Each cell displays the average labor market participation for the counterfactual wage increase conditional upon a particular quantile of $\beta$. Specifically, for a $x\%$ wage increase and quantile $q_\alpha \equiv \inf \{c:\hat{F}_\beta(c)\ge \alpha\}$, the table reports $$ \frac{\sum_{m=1}^M\sum_{t=1}^T1\{A^{(x)}_{m,t}=1,\beta^{(x)}_m=q_\alpha\}}{T\sum_{m=1}^M1\{\beta^{(x)}_m=q_\alpha\}}. $$ The table also displays the implied elasticity of quantile-specific labor force participation with respect to wages, based upon the 25% wage increase.\footnote{Specifically, the elasticity is calculated as $\frac{\left(\frac{\sum_{m=1}^M\sum_{t=1}^T1\{A^{(25)}_{m,t}=1,\beta^{(25)}_m=q_\alpha\}}{T\sum_{m=1}^M1\{\beta^{(25)}_m=q_\alpha\}}\right)-\left(\frac{\sum_{m=1}^M\sum_{t=1}^T1\{A^{(0)}_{m,t}=1,\beta^{(0)}_m=q_\alpha\}}{T\sum_{m=1}^M1\{\beta^{(0)}_m=q_\alpha\}}\right)}{\frac{\sum_{m=1}^M\sum_{t=1}^T1\{A^{(0)}_{m,t}=1,\beta^{(0)}_m=q_\alpha\}}{T\sum_{m=1}^M1\{\beta^{(0)}_m=q_\alpha\}}.}.$} For comparison, total (i.e., unconditional) labor force participation is 0.6496, and its elasticity with respect to wages is estimated to be approximately 0.11. Standard errors for the counterfactual estimates are in Table (ref).
Several observations can be made from this counterfactual exercise. First, average labor force participation varies greatly across the distribution of productivity. For instance, it increases from 13% at the first percentile to almost 100% at the 99th percentile.\footnote{In the data, around 14.6% of individuals never work. This percentage is 10.6% in the simulated data with no wage change.} Second, the supply response to a wage increase is much larger at lower skill quantiles: the implied elasticity is 0.30 at the 20th percentile, but only 0.036 at the 80th percentile.
In this paper I show point identification of a broad class of multinomial dynamic discrete choice models with multivariate continuous permanent unobserved heterogeneity. Relative to the existing literature, I allow for permanent unobserved heterogeneity that is both multivariate and continuous, and provide low-level conditions for point identification. My results encompass both finite and infinite horizon models, and do not rely on a full support condition, nor parametric assumptions on the distribution on permanent unobserved heterogeneity.
I propose a seminonparametric estimator for the distribution of continuous permanent unobserved heterogeneity in the style of heckman1984method. The estimator is computationally simple, and coincides with the estimator for a semiparametric model. As a result, the applied econometrician can proceed as they would for discrete permanent unobserved heterogeneity, providing they commit to increasing the number of support points as the sample size grows.
\addcontentsline{toc}{section}{\refname} \printbibliography