EconBase
← Back to paper

Dynamic Discrete-Continuous Choice Models: Identification and Conditional Choice Probability Estimation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

148,150 characters · 29 sections · 132 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

1420 Dynamic Discrete-Continuous Choice Models: Identification and Conditional Choice Probability Estimation

\setstretch{1} {0pt} {0pt}

abstractThis paper develops a general framework for dynamic models in which individuals simultaneously make both discrete and continuous choices. The framework incorporates a wide range of unobserved heterogeneity. I show that such models are nonparametrically identified. Based on constructive identification arguments, I build a novel two-step estimation method in the lineage of hm1993 and am2011 but extended to simultaneous discrete-continuous choice. In the first step, I recover the (type-dependent) optimal choices with an expectation-maximization algorithm and instrumental variable quantile regression. In the second step, I estimate the primitives of the model taking the estimated optimal choices as given. The method is especially attractive for complex dynamic models because it significantly reduces the computational burden associated with their estimation compared to alternative full solution methods. \\ Keywords: Discrete and continuous choice, dynamic model, identification, structural estimation, unobserved heterogeneity. \\

\setstretch{1.30} {6pt} {6pt}

Introduction

Many economic problems involve joint discrete and continuous choices. For example, a firm can decide what to produce and the corresponding sale price crawford2019. Firms also decide whether to register their business and how many workers to hire ulyssea2018. Students select their majors and decide how much effort to exert in their study arcidiaconoetal2019. Consumers decide what to buy and how much to consume dubin1984. In housing, buyers decide on their house size and housing tenure hanemann1984, bckm2013. The buyer of a car selects a model and the mileage of the car bento2009. Individuals decide whether to retire or not and how much they plan to consume accordingly iskhakov2017. Similarly, labor force participation and consumption/savings are joint choices for potential workers altugmiller1998, bcms2016, arellano2017earnings. \\ In all these examples, a rational individual makes both decisions simultaneously. As a result, the discrete choice is endogenous with respect to the continuous choice and vice versa. Taking the labor and consumption problem as the leading example throughout the paper, if an individual works, she consumes differently than if she does not work: she has two different conditional consumption choices. Moreover, her decision to work or not is dependent on these two conditional continuous choices. Unfortunately, the identification of models with simultaneous choices is difficult matzkin2007. Indeed, there is a core observability problem because we only observe the continuous choice made in the selected discrete alternative, and we do not know the counterfactual choices the individual would have made in the other alternatives. Ideally, we would like to recover counterfactual continuous choices using the choices of individuals with similar characteristics but who chose another alternative. However, doing so is not possible if individuals also differ on factors which are unobserved by the econometrician and affect both continuous and discrete choices. In this case, two identical individuals as measured by their observed covariates might still differ along the unobserved dimension. There is likely a problem of selection on unobservables, which prevents the identification of counterfactual continuous choices. To further pursue the example, if a researcher observes that working individuals consume more than unemployed individuals, she cannot identify whether this is because the consumption choice conditional on working is truly higher or because individuals with an unobserved higher taste for consumption select themselves more into working. \\ This paper develops a general framework of dynamic simultaneous discrete-continuous choice models suited for dynamic problems including a wide range of unobserved heterogeneity with transitory period-specific shocks and unobserved permanent types. I show how nonparametric identification of these models can be obtained by combining and extending insights from both, the dynamic discrete choice model literature hm1993, kasaharashimotsu2009, am2011 and the reduced-form literature on quantile treatment effects identification chernozhukovhansen2005, vuongxu2017. Then, building upon the identification, I provide a two-step estimation method for these models. The method is attractive because it yields significant computational gains regarding the estimation of dynamic discrete-continuous choice models, in the lineage of hm1993 for dynamic discrete choice models. \\ The first contribution of this paper is that I provide a constructive proof of the nonparametric identification of a general class of structural dynamic models in which individuals simultaneously make a discrete and a continuous choice. First, I identify the optimal discrete and continuous choice policies directly from the data, and then, taking these policies as given, I identify the primitives of the structural model. \\ The identification of the optimal choices proceeds in two sub-stages, each handling one of the two unobserved endogenous shocks present in the framework. The framework includes both (i) permanent unobserved types, capturing intrinsic latent differences between individuals and (ii) transitory shocks which only affect the individuals in a given time-period. To pursue the previous example, the types capture intrinsic differences in individuals' preferences for consumption and labor, in addition to individual-specific transitory taste shocks to consumption and labor every periods. In the first sub-stage, provided that the panel is long enough (more than $6$ observations for each individuals), I show how to identify the unobserved types from the complete panel of individual choices. To do so, I extend the dynamic discrete choice identification proof of kasaharashimotsu2009 to joint discrete and continuous choices with time-dependent joint densities and lagged dependent variables. Then, given the identified types, I show that the endogeneity of the discrete choice with respect to the transitory shocks can be handled using the lagged discrete choice as an instrumental variable (IV) to nonparametrically identify the optimal continuous and discrete choices. Indeed, provided that there are some switching costs in the discrete choice (e.g., switching costs of changing labor decision), the previous discrete choice affects the current discrete decision. However, in most models, conditional on the current discrete choice (and on the types and other current covariates), the lagged discrete choice has no direct effect on the current continuous choice. Thus, the lagged discrete choice is often a valid instrument, relevant for the current discrete choice (treatment), excluded from the current continuous choice (outcome), and exogenous with respect to the transitory shock. In this way, observable differences in the distribution of the choices due to variations in the instrument can be attributed to unobserved differences in selection, and not to differences in the continuous choices. I show that, paired with restrictions on the effect of unobserved heterogeneity on the continuous choice (monotonicity, rank invariance), the instrument allows us to establish nonparametric identification of the optimal discrete and continuous choices. The proof relates to and extends reduced form results on the nonparametric identification of quantile treatment effects chernozhukovhansen2005, vuongxu2017 with IVs. Indeed, I show that identifying the optimal choices in each period can be framed as identifying the effect of the discrete choice (endogenous treatment) on the continuous choice (outcome). This link is appealing as it grounds the identification of dynamic structural models in the treatment effect literature, making it less reliant on sometimes arbitrary structural assumptions (e.g., timing of the choices, discretization of the continuous choice, or implicit exogeneity assumptions between the choices). \\ Once the optimal choices are identified, I show how to use them to identify the primitives of the structural model. The key lies in linking these choices to the first-order conditions of the true structural model (e.g., the Euler equation determines the optimal continuous choices). Thus, one can reverse engineer the identified choices to identify the true primitives that generated them hm1993, bmm1997, escancianoetal2021. \\ The second contribution of the paper is in terms of estimation. I build a two-step estimation method, similar to hm1993 and am2011, but for discrete and continuous choices. In the first step, one estimates the policies, which I name after Hotz-Miller's CCPs: conditional continuous choices (CCCs) and \textit{conditional choice probabilities} (\textit{CCPs}). This step builds on the identification arguments. The policies are estimated directly from the data without solving the structural model. To account for unobserved types, I use an expectation-maximization (EM) algorithm, in the spirit of am2011. Then, I estimate the CCCs and CCPs building on IV quantile regression (IVQR) literature chernozhukovhansen2006, kaido2021decentralization, using the lagged discrete choice as an instrument which is valid conditional on the estimated types (and covariates). In the second step, one uses the estimated CCCs and CCPs to estimate the structure of the model. More specifically, I exploit the fact that within my framework, the primitives of the model are related to optimal choices through the first-order conditions. Given the estimated optimal choices, one can estimate the primitives of the model that generated these choices by satisfying these optimality conditions. The two-step estimation method is attractive because it yields sizeable computational gains. Typical dynamic discrete \textit{or} continuous choice models are difficult to estimate because they involve solving the theoretical model (either by backward recursion or fixed point algorithms). Dynamic discrete-continuous choice models are even more difficult to estimate because the mixed choices can introduce kinks and non-concavities in the value function iskhakov2017. Given that I can recover the CCCs and CCPs in the first step, I can exploit them to estimate the rest of the model without having to compute the value function or solve the model.\footnote{Since I do not solve for the CCCs and CCPs using an optimization algorithm, there is also no concerns about kinks and non-concavities in the value function that would make the estimation of these optimal choices more complicated. } This yields computational gains comparable to those obtained by hm1993 in the dynamic discrete choice literature, achieving estimation times already hundreds of times faster than the best available alternative iskhakov2017 in a simple toy model (see Section (ref)), and even greater improvements in more complex settings. The gains are so important that they not only reduce the time required to estimate the models, but also make it possible to estimate models that have thus far been deemed computationally intractable. In this respect, my method may facilitate the use of simultaneous discrete-continuous choice models, in particular the estimation of single-agent partial equilibrium life-cycle dynamic models. \\ Overall, the method builds a bridge between more reduced-form policy estimation and dynamic structural models. By enabling the estimation of structural models directly using reduced-form estimates, the method unlocks the possibility of doing counterfactual policy analysis on the basis of reduced-form results. In the leading example of the consumption and labor choice problem, the method described in this paper shows how to estimate standard life-cycle structural models of consumption and labor choices bcms2016 based on reduced-form/semi-structural estimates of optimal labor-specific consumption rules arellano2017earnings. With the structural model deep parameters (e.g., risk aversion), one can run many counterfactual policy analysis, varying tax/subsidies on labor or consumption for example. \\

Related literature. \%\\ There is a vast empirical literature that uses dynamic discrete choice models, for example, in studies of labor market transition and career choice keanewolpin1996, fertility choice ecksteinwolpin1989 and education choice arcidiacono2004. Starting from the bus replacement problem of rust1987, developments have been made regarding the estimation and identification of these models, including hm1993, hmss1994, rust1994, mt2002, aguirregabiriamira2002, aguirregabiriamira2007, kasaharashimotsu2009, am2011, hushum2012, am2019, am2020, abbringdaljord2020, and berry2023instrumental among others. For a survey, see aguirregabiriamira2010 or arcidiaconoellickson2011. \\ Similarly, the literature on dynamic continuous choice models is also voluminous, especially concerning consumption/saving carroll2006 or investment choices hs2010. There are also methods such as bbl2007 that can be applied to either dynamic discrete choice models or dynamic continuous choice models (but not both).\footnote{More precisely, bbl2007 describe problems with discrete or continuous policy functions separately. Extending their estimation techniques to more general Framework with discrete and continuous choices and with several unobservables yielding endogeneity of both choices would require identifying the first stage optimal policies following the approach described in this paper first. } \\ However, many economic problems involve multiple joint decisions, not only one discrete choice or only one continuous choice. For example, labor force participation is very much related to saving decisions. By focusing only on one of these two dimensions and ignoring the other (endogenous) choice, one might be missing something important. Unfortunately, empirical applications of the dynamic discrete-continuous choice framework are less common, as there was no general identification result available. For example, bmm1997 provide identification of such models once the optimal choices are identified but do not directly address the identification of these choices. The existing literature employs several tricks to overcome the problem of selection on unobservables. The most extreme is to assume away the problem by assuming selection on observables only, i.e., conditional on the observed covariates, assume that there is no other unobservable affecting the optimal choices. This is fairly strong, especially in dynamic models where the number of covariates is typically limited. Without ruling out the existence of these unobservables, another common approach is to have implicit or explicit assumptions about the selection process, through assumptions about the relation between the error terms affecting both choices, e.g., independence, measurement errors or known joint distribution dubin1984, hanemann1984, bento2009. Another common technique is to discretize the continuous choice so that the discrete-continuous model can be rewritten as a discrete choice model degrooteverboven2019. This is appealing, as it allows the application of known techniques in the dynamic discrete choice literature. However, discretizing the continuous choice is implicitly equivalent to making an assumption about the selection process via an assumption on the distribution of the additive discrete error terms. Another approach is to resort to timing assumptions which implicitly break the endogeneity of the choices. blevins2014 shows nonparametric identification of dynamic discrete-continuous choice models assuming a specific timing in which the discrete choice takes place before the realization of the nonseparable shocks affecting the continuous choice: hence the selection (discrete choice) does not depend on the nonseparable shock. iskhakov2017 and murphy2018 use similar timing assumptions, which are effectively equivalent to imposing that the discrete choice is exogenous. A more convincing alternative is to allow for endogeneity but reduce the level of unobserved heterogeneity, for example, by including only a finite number of unobserved types bcms2016. My approach is more general, as I allow for a more flexible distribution of unobserved heterogeneity with both period-specific transitory shocks and permanent unobserved types. I handle the complex endogeneity of the discrete and continuous choices by extending techniques from both the dynamic discrete choice literature to identify the types kasaharashimotsu2009, hushum2012, and from the reduced form quantile treatment effect literature to handle the intra-period endogeneity chernozhukovhansen2005. Linking the identification of structural models with nonparametric treatment effect identification results is appealing as these results rely less on sometimes arbitrary structural assumptions (timing, discretization, exogeneity, distribution of the errors, ...). Furthermore, my identification allows to test these assumptions. \\ Most closely related to this paper, contemporaneous work by levy2024identification also addresses the identification of simultaneous discrete-continuous dynamic choice models with rich unobserved heterogeneity in two steps: first they identify the optimal choices, then the model primitives. The main difference between our papers is the manner in which they handle the endogeneity to identify the optimal policies in the first stage. They focus on infinite horizon setups with a stationary environment and address the selection by requiring the existence of a (sequence of) variable(s) such that the probability of selecting some specific discrete alternatives becomes arbitrarily high (tends to one). In practice, however, the existence of such a variable is hard to satisfy in most applications. To provide an analogy with the treatment effect identification literature, their identification arguments are similar to identification-at-infinity arguments, requiring the existence of a "infinitely relevant" instruments, such that the selection probability tends to one. Instead, I only need a weaker standard relevant instrument to address endogeneity and identify the optimal choices. While even standard IVs may be hard to find in the context of dynamic models, I show that in many structural models, the past discrete choice will be a valid instrument as soon as there are nonzero switching costs (conditional on the covariates and the types). This is a relatively mild condition in many applications, and will be testable with my framework. In the special case of levy2024identification's application to retirement and consumption decisions, and more generally in the presence of absorbing states in the discrete choice, our identification arguments coincide. Indeed, retirement is an absorbing state, so the probability of being retired today conditional on being previously retired is one. Consequently, the previous retirement status satisfy their identification-at-infinity condition, and is also an (infinitely) relevant instrument in my case (scenario equivalent to infinitely high switching cost). In fact, I show that when the discrete choice has an absorbing state, my identification arguments are considerably simplified and focussing on individuals who are already in the absorbing state has additional identification power (see Section (ref)). In addition to these, the difference with my paper is that they use pairwise differencing for the estimation of their primitives and require separability of the unobservables in the marginal utilities to do so, while I do not need it. I also take into account auto-correlated shocks through permanent types, making a link with the dynamic discrete choice literature. \\ As already mentioned, this paper builds a general framework that connects the identification of structural models with the reduced form nonparametric identification literature neweypowell2003, chesher2003, newey2007, matzkin2007, matzkin2008, imbensnewey2009, torgovitsky2015, dhaultfoeuillefevrier2015. By casting the optimal choices in the form of a triangular simultaneous system of equations, I show how their identification can be framed as the identification of quantile treatment effects with an endogenous treatment, i.e., the IV quantile regression (IVQR) Framework chernozhukovhansen2005, vuongxu2017, feng2024matching, where the discrete and continuous choices can be understood as the treatment and the outcome, respectively. Moreover, I improve on the existing results of chernozhukovhansen2005 by weakening their relevance condition: instead of their global full rank condition, I show that the identification can be obtained under weaker, testable, and easier to interpret relevance condition. More precisely, I need that the instrument is relevant almost everywhere, except possibly at a finite set of isolated values of the unobservable shocks. Allowing for some isolated points of irrelevance is important, especially in dynamic models where the discrete choice has many alternatives, as these locally irrelevant points may often occur, even in simple models. The reason why I can relax the full rank condition of chernozhukovhansen2005 is that they do not exploit a key property of their quantile model: the fact the continuous choices (outcomes) are strictly increasing in their unobservable shocks (ranks). This monotonicity has extra power in terms of identification. To the best of my knowledge, vuongxu2017 are the only others who also exploit the power of monotonicity to relax the full rank condition of chernozhukovhansen2005 and still identify quantile treatment effects, but only in the context of a binary treatment. I further show that this weaker relevance condition can be expressed as an easy-to-interpret conditions on the conditional choice probabilities (depending on the unobservable shock affecting the continuous choices), which is testable. \\ Similarly, I contribute to the literature on the identification of models with unobserved types kasaharashimotsu2009, hushum2012, higgins2023identification. In particular, I extend the identification of kasaharashimotsu2009 to type-dependent joint discrete-continuous choice densities, where the densities are time-dependent and depend on lagged choices. I also show how to account for covariates for which the transition is deterministic given the choices (e.g., assets), which violates standard assumptions in this literature. \\ For the identification of the primitives of the model given the identified optimal choices, I build upon escancianoetal2021 and bmm1997. I adapt escancianoetal2021 to my framework to identify the marginal utilities and the discount factor from the Euler equations. Then I adapt bmm1997 to identify the remaining primitives (value functions) using these marginal utilities. \\ I also contribute to the literature on fast estimation methods, avoiding the computation of the value function rust1987, hm1993, hmss1994, carroll2006, am2011, iskhakov2017. I provide a faster alternative to indirect inference and the most recent developments of endogenous grid methods iskhakov2017. A timing comparison of the different estimation methods is given in Section (ref). \\

Outline. The Framework contains several building blocks, that we will develop backwards. First, Section (ref) describes the intra-period simultaneous discrete-continuous choice problem, for any given period $t$, and assuming the types are already identified. It also discusses nonparametric identification of the optimal choices within any period. Then, Section (ref) shows the general dynamic models that yields these intra-period problems. Section (ref) shows how to identify the permanent types beforehand. \\ Building on the complete Framework and identification arguments, Section (ref) describes the estimation method and Section (ref) shows the estimator performances, in terms of precision and computational time, using Monte-Carlo simulations of a dynamic life-cycle model of consumption and labor force participation choices. Section (ref) concludes.

The intra-period problem

This section describes the intra-period problem of a dynamic model and its nonparametric identification for any specific period $t$. This serves as a building block and the identification of the dynamic model will then be described in Section (ref). I also proceed conditional on the type, $m$, which should have been identified beforehand (see Section (ref)). I abstract from the period $t$ and type $m$ to simplify the notation. The main text describes the framework with a binary discrete choice, extension and identification with more than $2$ discrete alternatives is in Appendix (ref).

Intra-period Framework

Consider an individual's decision problem with the following timing within a period: {{15pt}

center[center omitted — 489 chars of source]

}

The individual simultaneously selects a discrete action $d$ $\in \mathcal{D} = \{0, 1\}$ and accordingly makes one continuous choice $c_d \in \mathcal{C}_d$, where $\mathcal{C}_d$ is a compact subset of $\mathbb{R}$, to maximize his payoff.\footnote{For a more general discrete choice with $\mathcal{D} = \{0, ..., J\}$, see Appendix (ref).} The decision is made given some state $z \in \mathcal{Z}$ observed by the researcher, as well as two transitory period $t-$specific preference shocks, $\epsilon = (\epsilon_0, \epsilon_1) \in \mathcal{E} \subset \mathbb{R}^2$ and $\eta \in \mathcal{H} \subset \mathbb{R}$. The shocks $\epsilon$ and $\eta$ are realizations of the random variables $\Epsilon = (\Epsilon_0, \Epsilon_1)$ and $\Eta$ and are unobserved by the researcher. The shock $\epsilon$ only affects the discrete choice $d$, while $\eta$ impacts the continuous choice $c$ and the discrete choice. The same $\eta$ impacts the continuous choice decision in both discrete-choice states ($c_0$ and $c_1$), that is, there is rank invariance hsc1997, chernozhukovhansen2005.\footnote{The continuous choices could even represent different variables depending on the discrete option selected: for example, if $d$ represents the choice between working and studying, $c$ might represent the amount of time worked and the effort of the student respectively, hence with possibly different supports. The main restriction is that even if they represent two different choices, these two continuous choices are impacted by the same unobserved shock $\Eta$. }

The payoffs of the individual are given by the function $\mathcal{V}_d(c_d, z, \eta, \epsilon_d)$. The individual simultaneously selects $d$ and $c_d$ to solve:

equation[equation omitted — 94 chars of source]

I require additional assumptions for tractability and identification of the model.

assumption[Additive Separability] The shock $\epsilon$ enters the payoff additively: \begin{align*} \mathcal{V}_d(c_d, z, \eta, \epsilon_d) = \tilde{v}_d(c_d, z, \eta) + \epsilon_d, \quad for d = 0, 1. \end{align*}

The additive separability assumption is common in the discrete choice model literature rust1987, am2011. It applies to $\epsilon$, while $\eta$ can still enter the payoff in a nonseparable manner. A consequence of Assumption (ref) is that the optimal conditional policy functions given $d$, $c_d^*(\cdot)$, will not depend on $\epsilon$:

align*[align* omitted — 189 chars of source]
assumption[Instrument] The state vector contains two kinds of variables, $z = (x, w)$, where $x \in \mathcal{X}$ and $w \in \mathcal{W} = \mathcal{D}$,\footnote{Given the IV I use (past discrete choice), the support of $W$ is the support of $D$. In general the support of $W$ must be larger than the support of $D$. For discrete or continuous $W$, the identification proof follows along the same lines.} and \begin{center} $\tilde{v}_d(c_d, z, \eta)= v_d(c_d, x, \eta) + m_d(x, w, \eta)$, \quad $\text{ for } d=0,1.$ \end{center}

Here, $x$ represents general state variables and $w$ is an `instrument' to recover the optimal conditional policies $c_d^*$. On the one hand, $w$ is excluded from the optimal policies $c_d^*$ since

align*[align* omitted — 174 chars of source]

On the other hand, $w$ might still be relevant and impact the discrete choice $d$.

assumption[Monotonicity] The payoff functions $v_d$ are twice continuously differentiable and \begin{equation*} \frac{\partial^2 v_d(c_d, x, \eta)}{\partial c_d \partial \eta} > 0, \quad for d=0,1. \\ \end{equation*}

Assumption (ref) implies that, given $D=d$ and $X=x$, the conditional optimal policy function $c_d^*(x, \eta)$ is continuously differentiable and strictly increasing in $\eta$. Hence $\eta$ and $c_d^*$ are one-to-one for every $d$ and $x$. This kind of monotonicity condition has been widely used for identification chernozhukovhansen2005, bbl2007, hs2010. In a sense, it means that I only identify monotone effects of the unobserved nonseparable source of heterogeneity, $\eta$. An important limitation of Assumption (ref) is that it requires a nontrivial continuous choice $c_d$ for each discrete alternative $d$. For example, Assumption (ref) is not satisfied in the case where an investor decides whether to invest ($d=1$) or not ($d=0$) and the corresponding investment conditional on investing ($d=1$) hs2010. Indeed, in this case, $c^*_0(x, \eta) = 0$ for all $\eta$ (and $x$), and $c_0^*$ is not strictly increasing in $\eta$. In contrast, Assumption (ref) holds in the case of a discrete choice between portfolios and the corresponding conditional level of investment. \\ Under Assumptions (ref), (ref) and (ref) we obtain the following triangular structure for the reduced-form optimal choices:

align*[align* omitted — 138 chars of source]

This triangular structure links my structural model with the literature on (reduced-form) systems of simultaneous equations chesher2003, matzkin2008, imbensnewey2009 and, more specifically, the related literature on heterogeneous (quantile) treatment effects chernozhukovhansen2005, vuongxu2017. To identify the structure, one needs to first identify the optimal choice functions. To identify them, I need additional assumptions on the shocks.

assumption[Shocks] Conditional on $X=x$, \begin{enumerate*}[label={{(\roman*)}}, ref={$D$(ref)(\roman*)}] • $W$, $\Eta$ and $\Epsilon$ are mutually independent; • $\Eta$ is continuously distributed as $\mathcal{U}(0,1)$; • $\Epsilon$ is continuously distributed with full support; • $\underset{c}{\textrm{max}} \ \tilde{v}_d(c, x, w, \eta) < \infty$ for all $(x, w, \eta, d)$. \end{enumerate*}

The main independence restriction is that $\Eta$ is independent of $W$ given $X$. The identification of $c_d^*$ requires $\Eta$ to have the same distribution, regardless of the realization of $W$. Other than this, the independence assumption is not as restrictive as it may appear. Indeed, note that the additive term $m_d(x, w, \eta)$ can be interpreted in two ways that cannot be separately identified. In Assumption (ref), $m_d$ is an additive part of the payoff $\tilde{v}_d$. However, $m_d$ can also be interpreted as part of a more general additive discrete-choice shock, $\tilde{\epsilon}_d(x, w, \eta) = m_d(x, w, \eta) + \epsilon_d$, in which case $\Epsilon_d$ is the part of the discrete-choice shock that is independent of $\Eta$ and $W$. The continuity of the distribution of $\Eta$ is imposed to obtain smooth conditional distributions of the continuous choices. I cannot identify the distribution of $\Eta$ separately from the utility. Therefore, as is standard in the literature bmm1997, matzkin2003, I normalize $\Eta$ to be uniformly distributed (given $X$). This normalization is innocuous. Formally, I nonparametrically identify the quantiles of the optimal choices and payoffs. Similar to the distribution of $\Eta$, the distribution of $\Epsilon$ is not nonparametrically identified in my setup, but this does not affect the nonparametric identification of the optimal choices nor of the payoff function, $v_d$, as long as $\Epsilon$ has full support. Assumption (ref) is a regularity condition on the functional form ensuring that $0 < \textrm{Pr}(D=d | \Eta=\eta, Z=z) < 1$ for all $(d, \eta, z)$. \\ I need one last (testable) condition for identification.

{

assumption[Instrument Relevance] For every $x \in \mathcal{X}$, \begin{align*} Pr(D=0 | \Eta = \eta, X=x, W=1) \neq Pr(D=0 | \Eta = \eta, X=x, W=0), \end{align*} $\text{ for all } \eta \in \mathcal{H} \backslash \mathcal{K}_x,$ where $\mathcal{K}_x$ is a (possibly empty) finite set containing $K$ values ($K \geq 0$).

}

Identification of the optimal policies requires that the instrument is sufficiently relevant. It needs to be relevant `almost everywhere', but I show that identification still holds even if there is a finite set of values of $\eta$ at which the instrument is not relevant, which could occur if the switching costs vary with $\Eta$. Assumption (ref) yields testable implications for the observed reduced forms distributions of $C$ and $D$. It allows to test whether the structural model is identified, as I discuss in the next section. Finally, note that Assumption (ref), expressed in terms of the conditional choice probabilities, is equivalent to an assumption on the structural functions $m_d$. Indeed,

align*[align* omitted — 317 chars of source]

Since $\underset{c}{\textrm{max}} \ v_1 (c, x, \eta) - \underset{c}{\textrm{max}} \ v_{0} (c, x, \eta)$ is independent of $W$ and since $\Epsilon_d \perp (W, \Eta) | X = x$, we have that:

align*[align* omitted — 199 chars of source]

Summary of the setup. I consider a decision problem where an individual selects $(d, c_d)$ to maximize his payoff:

align*[align* omitted — 124 chars of source]

The setup applies to a wide range of (static and) dynamic discrete-continuous choice models. I provide an example below that will be developed further in Section (ref). \\

Example: Life-cycle model of consumption and labor. \\ Consider a standard dynamic model where individuals choose how much to consume/save and whether to work or not (or to work part time or full time) every period altugmiller1998, bcms2016, arellano2017earnings. The individual simultaneously chooses between working ($d=1$) or not ($d=0$), and how much to consume accordingly, $c_d$. The consumption functions can be thought of as `potential consumptions' (potential outcomes), and are completely flexible functions of the labor decision (treatment). The vector $x$ contains information about the asset, income, education and other individual characteristics (demographics such as the age, gender, marital status, ...). Note that the asset and income may not affect the current period utility directly, but still affect the conditional value functions, $v_d$, indirectly through their impact on the future (see more discussion in Section (ref)). Implicitly here, I omit the unobserved permanent type $m$, which should have been identified beforehand, and could be thought of as another covariate, affecting both the preferences for work and consumption. The shock $\epsilon_d$ represents individual-specific transitory unobserved preferences for work. The shock $\eta$ represents other unobserved transitory shocks of the individual impacting her preference for consumption, and possibly also her preference for work directly. The higher $\eta$ is, the higher $c_d$ for all $d$. In practice, the greatest challenge is to find a good instrument $w$. Fortunately, the previous labor decision could serve as such an IV. Indeed, in the presence of switching costs, e.g., if $m_d(x, w, \eta) > 0$ when $d \neq w$, and $m_d(x, w, \eta) = 0$ when $d = w$, the previous labor decision is relevant for the current one (Assumption (ref)). Conditional on the current decision, on the types which capture intrinsic unobserved characteristics of the individuals, and on covariates which capture their wealth and observed characteristics, the previous decision should have no effect on the current consumption choice: it is excluded. Finally, since $\eta$ is purely transitory and occurring in period $t$, the past labor decisions are independent of it (the past labor still depends on the permanent types). Thus, the previous labor decision has unique properties that makes it a valid IV (relevant, excluded and exogenous) in dynamic models, because its effect on the current consumption is "subsumed" by the effect of the current labor choice. \\

Discussion of simultaneity. This simultaneous choice framework nests the non-simultaneous timings where either the discrete or the continuous choice is made first and is based on expectations about the other choice (and the corresponding shock). These two timings have testable implications for the optimal choices within the simultaneous choice framework:

enumerate[label=(\roman*)] • If the discrete choice is made first (before the realization $\Eta=\eta$ and the continuous choice), then the CCP $\textrm{Pr}(D=d | \Eta=\eta, X=x, W=w)$ does not depend on $\eta$. Indeed, $\eta$ is not yet realized. The discrete choice is only based on expectations about $\eta$ and the corresponding $c^*_d(x, \eta)$. • Conversely, if the continuous choice is made first (before the realization $\Epsilon = \epsilon$ and the discrete choice), then the CCCs $c_d^*(x, \eta)$ do not depend on $d$, i.e., $c_0^*(x, \eta) = c_1^*(x, \eta)$ for all $\eta$ and $x$.

Since I identify the policy functions $c_d^*$ and $\textrm{Pr}(D=d | \Eta=\eta, X=x, W=w)$ in the simultaneous choice framework, I can test the timing of the decisions.

Identification

The unobserved shocks $(\Eta, \Epsilon)$ are independent and identically distributed across individuals. I observe data on the variables $(D, C, X, W)$. I only observe $C=C_0$ if $D=0$ and $C=C_1$ if $D=1$, where $C_0$ and $C_1$ are the potential choices. For all $(x, w, \eta)$ in $\mathcal{X} \times \mathcal{W} \times \mathcal{H}$, I study nonparametric identification of the following objects for $d=0, 1$: the optimal conditional continuous choices (CCCs) $c_d^*(x, \eta)$, the optimal conditional choice probabilities (CCPs) $\textrm{Pr}(D=d | \Eta=\eta, W=w, X=x)$, the indirect payoff functions (taken at the optimal $c$) $\underset{c}{\textrm{max}} \ v_d(c, x, \eta)$ and $m_d(x, w, \eta)$. In this section, I focus on any given value $X=x$ and omit $x$ from the notation in what follows. This is without loss of generality since my assumptions about the distribution of the shocks hold conditional on $X=x$. First, I characterize the reduced forms and constraints imposed by the structural assumptions. Then, I discuss the identification of the optimal policies (CCCs and CCPs) and of the payoffs. \\ In the main text I focus on the case where $D$ is binary. Appendix (ref) discusses identification in the case where $D$ is discrete and takes more than two values.

Reduced forms and constraints

In the data, I observe $(D, C, W)$, where $W$ is exogenous while $C$ and $D$ are endogenous choices. There is a fundamental observability problem, as I only observe one of the two potential choices $C_0$ and $C_1$ depending on the discrete choice selected:

align*[align* omitted — 41 chars of source]

Therefore, from the data, I only recover the distribution of the potential choice $C_d$ conditional on $D=d$ (and $W=w$), that is, $F_{C_d | d, w}(c_d) = \textrm{Pr}(C_d \leq c_d | D=d, W=w)$ for all $d$ and $w$. The functions $F_{C_0 | 1, w}(c_0)$ and $F_{C_1 | 0, w}(c_1)$ are not observed in the data. I also recover the conditional probability of selecting $d$ given $W=w$, that is, $p_{D|w}(d) = \textrm{Pr}(D=d | W=w)$. The data provide the following reduced-form functions, which exhaust all relevant information:

align*[align* omitted — 118 chars of source]

A remark on terminology: in this paper, $\textrm{Pr}(D=d | W=w)$ is part of the reduced form, while $\textrm{Pr}(D=d | \Eta=\eta, W=w)$ is what I call the conditional choice probabilities (CCPs) or selection on unobservables process that I want to identify. This differs from the dynamic discrete choice literature, where $\textrm{Pr}(D=d | W=w)$ are actually called CCPs hm1993, am2011. Here, however, I have simultaneous choices and a nonseparable shock $\eta$ that affects both choices. Thus, the counterparts to the usual CCPs are $\textrm{Pr}(D=d | \Eta=\eta, W=w)$ for all $d$, hence the different terminology. \\ The structural assumptions imply the following constraints on the reduced form.

LemmaUnder Assumptions (ref) and (ref), $F_{C_d | d, w} (\cdot)$: $\mathcal{C}_d \rightarrow [0,1]$ is continuously differentiable and strictly increasing, for $d=0, 1$.
proofAppendix (ref)
LemmaUnder Assumption (ref), \begin{align*} \frac{\partial F_{C_d |d, 1}(c_d) p_{D|1}(d)}{\partial c_d} \neq \frac{\partial F_{C_d |d, 0}(c_d) p_{D|0}(d)}{\partial c_d} \quad for all c_d \in \mathcal{C}_d \backslash \mathcal{K}^{c_d}, for d=0,1, \end{align*} where $\mathcal{K}^{c_d}$ is a (possibly empty) finite set containing $K$ values.
proofAppendix (ref)

Lemmas (ref) and (ref) fully characterize the impact of the structural assumptions on the reduced-form functions. Lemma (ref) is a regularity result on the distributions implied by the structural form. Lemma (ref) provides observable and testable implications of the structural model, specifically of Assumption (ref), on the reduced-form functions. Indeed, in Assumption (ref), $\textrm{Pr}(D=d | \Eta=\eta, W=w)$ is unobserved since $\eta$ is unobserved. However, by monotonicity of the optimal continuous choices, the observed conditional distributions of $C_d$ given $D=d$ are transformations of the unobserved conditional distributions of $\Eta$ given $D=d$. Define the difference

align*[align* omitted — 125 chars of source]

I show that when the instrument is relevant, i.e., when $\textrm{Pr}(D=d | \Eta=\eta, W=1) \neq \textrm{Pr}(D=d | \Eta=\eta, W=0)$, we have $\partial (\textrm{Pr}(\Eta \leq \eta |D=d, W=1) - \textrm{Pr}(\Eta \leq \eta |D=d, W=0))/\partial \eta \neq 0$ and $\partial \Delta F_{C_d}(c_d^*(\eta))/\partial c_d \neq 0$. Now, the functions $\Delta F_{C_d}(c_d)$ and $\partial \Delta F_{C_d}(c_d)/\partial c_d$ are well defined (according to Lemma (ref)) and are directly observable. Therefore, even if we do not observe the conditional distribution of $\Eta$ given $D=d$, we know that if the instrument is sufficiently relevant (Assumption (ref)), Lemma (ref) holds. I use this to test the relevance of the instrument: if the function $\Delta F_{C_d}(c_d)$ is flat over an interval of values $c_d$, then there is a corresponding interval of values $\eta$ where the instrument is not relevant. In this case, the instrument has no impact on the conditional choice probabilities, so the optimal continuous choices are not point identified on this interval of $\eta$.

Identification of conditional continuous choices (CCCs)

As in the literature on continuous choice matzkin2003, bbl2007, hs2010, I would like to exploit the monotonicity assumption to identify the optimal continuous choices. By monotonicity, we have for $d=0,1$,

flalign*Pr(\Eta \leq \ \eta \ | D=d) &= Pr(C_d \leq \ c_d^*(\eta) \ | D=d) &\\ and, hence, by Lemma (ref), \quad & &\\ c_d^*(\eta) &= F_{C_d | d}^{-1}(Pr(\Eta \leq \eta | D=d)). &

Thus, if we knew the distribution of $\Eta$ given $D=d$, we could recover the optimal conditional continuous choices $c_d^*(\eta)$. However, here we only know (by normalization) the unconditional distribution of $\Eta$. The conditional distributions of $\Eta$ given $D=d$ are unobserved. They depend on a selection on unobservables: $\textrm{Pr}(\Eta \leq \eta | D=d) = \textrm{Pr}(D=d | \Eta \leq \eta) \textrm{Pr}(\Eta \leq \eta)/\textrm{Pr}(D=d)$, which precludes the use of inversion. \\ Another way to see the problem is as follows. Since $\Eta$ is exogenous,

align*[align* omitted — 516 chars of source]

If we observed both potential choices $C_0$ and $C_1$ for every individual, irrespective of the discrete choice $d$, then the unconditional distribution of $C_d$, $F_{C_d}$, would be observed for $d=0,1$. Then, knowing that $\Eta$ is uniform, one could exploit monotonicity to recover $c_d^*(\eta)$ by inverting its unconditional distribution: $c_d^*(\eta) = F_{C_d}^{-1}(\textrm{Pr}(\Eta \leq \eta))$. However, we only observe $C_0$ if $D=0$ and $C_1$ if $D=1$. Because of this selection, only the conditional distributions of $C_d$ given $D=d$ are observed, and $c_d^*(\eta)$ is not identified by standard inversion. \\

Identification with the instrument. \\ Instead, to identify $c_d^*(h)$, I use the properties of the instrument (Assumption (ref)) to obtain structural restrictions. We have, for $\eta \in [0,1]$ and $w=0,1$,

align[align omitted — 599 chars of source]

Now, take equation ((ref)) at $w=0$ and $w=1$ to obtain a system of two equations to identify two unknown increasing functions, $c^*_0(\cdot)$ and $c^*_1(\cdot)$. The role of the instrument and Assumption (ref) appears clearly here. First, the exclusion of $w$ from $c_d^*$ is necessary to avoid having four unknown functions $c_d^*(\eta, w)$ (where $d=0,1$ and $w=0,1$) which would not be identified with two equations. Similarly, without a relevant instrument (e.g., if $D \perp W$), $p_{D|0}(d) = p_{D|1}(d)$ and $F_{C_d | d, 0}(c) = F_{C_d | d, 1}(c)$ for $d=0,1$, so the two equations would coincide, giving one equation for two unknown functions. \\

theorem[Identification] For every reduced form compatible with the structural model, there exist unique conditional continuous choice (CCC) functions $c_d(h)$ ($d=0, 1$) that are strictly increasing and satisfy {\tagsleft@true \begin{equation} \eta = F_{C_0 |0, w}(c_0(\eta)) p_{D|w}(0) + F_{C_1 |1, w}(c_1(\eta)) p_{D|w}(1) \quad for \eta \in [0, 1], \ w=0,1. \end{equation} }

The CCC functions are identified if and only if there exist unique functions $c_d(\eta)$, $d=0,1$, that are strictly increasing in $\eta$, satisfy equation ((ref)), and are compatible with the reduced form $R$. \\

proofThe existence of a solution is trivial: the reduced form is compatible with the structural model, so, by construction following equation ((ref)), $c^*_d(\cdot)$ ($d=0,1$) solve ((ref)). \begin{figure}[!t] \begin{subfigure}[b]{0.45\textwidth} \begin{tikzpicture}[scale=1] \begin{axis}[ clip=false, axis lines = middle, ytick=\empty, xtick=\empty, ymin=-3, ymax=5, xmin=0, xmax=1.2, ] \addplot [domain=0.1:0.735, samples=1000, color=black,]{(10*(x-0.1))^0.75}; \addplot [domain=0.35:0.9353, samples=1000, color=black,]{(4*(x-0.5))^2.5}; \node[scale=0.7] at (axis cs:1.25,0) {$c$}; \node[scale=0.7] at (axis cs:1.06,3.5) {$-\Delta F_{C_1}(\cdot)$}; \node[scale=0.7] at (axis cs:0.45,3.5) {$\Delta F_{C_0}(\cdot)$}; \draw[dashed] (axis cs:0.75,0) -- (axis cs:0.75,1); \node[below, scale=0.7] at (axis cs:0.75,-0.10) {$c_1$}; \draw[dashed] (axis cs:0.2,0) -- (axis cs:0.2,1); \node[below, scale=0.7] at (axis cs:0.2,-0.05) {$\tilde{c}_0(c_1)$}; \draw[dashed] (axis cs:0,1) -- (axis cs:0.75,1); \end{axis} \end{tikzpicture} \caption{Strictly monotone $\Delta F_{C_d}(\cdot)$ $(K = 0)$} \end{subfigure} \begin{subfigure}[b]{0.45\textwidth} \begin{tikzpicture}[scale=1] \begin{axis}[ clip=false, axis lines = middle, ytick=\empty, xtick=\empty, ymin=-3, ymax=5, xmin=0, xmax=1.2, ] \addplot [domain=0.1:0.859773, samples=1000, color=black,]{(10*(x-0.1))^0.7*sin(deg(10*(x-0.1)))}; \addplot [domain=0.35:0.957818, samples=1000, color=black,]{(12.5*(x-0.35))^0.7*sin(deg(12.5*(x-0.35)))}; \node[scale=0.7] at (axis cs:1.25,0) {$c$}; \node[scale=0.7] at (axis cs:1.06,3.5) {$-\Delta F_{C_1}(\cdot)$}; \node[scale=0.7] at (axis cs:0.7,3.5) {$\Delta F_{C_0}(\cdot)$}; \draw[dashed] (axis cs:0.87437446,0) -- (axis cs:0.87437446,1); \node[below, scale=0.7] at (axis cs:0.87437446,-0.10) {$c_1$}; \draw[dashed] (axis cs:0.75546807,0) -- (axis cs:0.75546807,1); \node[below, scale=0.7] at (axis cs:0.75546807,-0.05) {$\tilde{c}_0(c_1)$}; \draw[dashed] (axis cs:0,1) -- (axis cs:0.87437446,1); \draw[gray, dashed] (axis cs:0.5583563,0) -- (axis cs:0.5583563,1); \draw[gray, dashed] (axis cs:0.441506,0) -- (axis cs:0.441506,1); \draw[gray, dashed] (axis cs:0.2143825,0) -- (axis cs:0.2143825,1); \draw[gray, dashed] (axis cs:0.36044535,0) -- (axis cs:0.36044535,1); \end{axis} \end{tikzpicture} \caption{Piecewise monotone $\Delta F_{C_d}(\cdot)$ $(K > 0)$} \end{subfigure} \caption{Intuition behind identification} \end{figure} To prove uniqueness, combine the two equations of ((ref)) to give, for $\eta \in \mathcal{H}$, \begin{flalign*} F_{C_0 |0, 0}(c^*_0(\eta)) p_{D|0}(0) + F_{C_1 |1, 0}(c^*_1(\eta)) p_{D|0}(1) &= F_{C_0 |0, 1}(c^*_0(\eta)) p_{D|1}(0) + F_{C_1 |1, 1}(c^*_1(\eta)) p_{D|1}(1) & \\ and hence, after rearranging, \quad \quad \quad \quad \quad \quad \quad & & \\ \quad \Delta F_{C_0}(c^*_0(\eta)) &= - \Delta F_{C_1}(c^*_1(\eta)), & \end{flalign*} where the functions $\Delta F_{C_d}(\cdot)$ are directly observed from the data. Now, even without observing $\eta$, if two conditional choices $\tilde{c_0}$ and $\tilde{c_1}$ correspond to the same unobserved $\eta$, we have $\Delta F_{C_0}(\tilde{c_0}) = - \Delta F_{C_1}(\tilde{c_1})$. That is, \begin{equation} \Delta F_{C_0}(\tilde{c_0}(c_1)) = - \Delta F_{C_1}(c_1) for all c_1 \in c_1^*(\mathcal{H}) = \mathcal{C}_1, \end{equation} where $\tilde{c_0} = c_0^*\circ {c_1^*}^{-1}$ is strictly increasing. The mapping $\tilde{c_0}$ is identified if and only if there exists a unique strictly increasing function solving ((ref)). Notice that the functions $\Delta F_{C_d}(\cdot)$ are transformations (through $c_d^*(\cdot)$) of the same underlying object based on the difference between $\textrm{Pr}(D=0 | \Eta=\eta, W=1) - \textrm{Pr}(D=0 | \Eta = \eta, W=0)$.\footnote{Specifically, as shown in Appendix (ref), for $d=0, 1$, \begin{align*} \Delta F_{C_d}(c) = (-1)^d \int^{(c_d^*)^{-1}(c)}_{0} \Big( Pr(D=0 | \Eta=\eta, W=1) - Pr(D=0|\Eta = \eta, W=0) \Big) d\eta. \end{align*} } Thus $- \Delta F_{C_1}(\cdot)$ is a non-constant shift of $\Delta F_{C_0}(\cdot)$, i.e., both functions go through the same values in the same order. So, in the simple case where the instrument is relevant everywhere (Figure (ref)), the functions $\Delta F_{C_0}(\cdot)$ and $-\Delta F_{C_1}(\cdot)$ are strictly monotone and continuously differentiable, they have the same range and we can invert ((ref)) to get the unique solution: \begin{align*} \tilde{c_0}(c_1) = \Delta F_{C_0}^{-1}\Big(-\Delta F_{C_1}(c_1)\Big) for all c_1 \in \mathcal{C}_1. \end{align*} Now, if the instrument is not relevant at $K$ isolated values of $\eta$ (Figure (ref)), the functions $\Delta F_{C_d}(\cdot)$ are not strictly monotone and not invertible. However, they are piecewise monotone and piecewise invertible. So, using the monotonicity constraint on $\tilde{c_0}(c_1)$, the solution to ((ref)) is also unique in this case. Indeed, for any value $y$ in the range of $\Delta F_{C_0}(\cdot)$, the $k^{th}$ value $c_0$, denoted $c_0^k$, such that $\Delta F_{C_0}(c_0^k) = y$ for the $k^{th}$ time, is the image of the $k^{th}$ value of $c_1$, denoted $c_1^k$, such that $ - \Delta F_{C_1}(c_1^k) = y$ for the $k^{th}$ time. Even though there may exist several solutions (at most $K+1$) such that $y = \Delta F_{C_0}(c_0) = \Delta F_{C_1}(c_1)$ for any given $y$, there is only one $\tilde{c_0}(c_1)$ that is strictly increasing and satisfies ((ref)) for all $c_1$.\footnote{For more details on the proof when $K > 0$, see Appendix (ref).} \\ The only case in which uniqueness does not hold is when the functions $\Delta F_{C_d}(\cdot)$ are flat on some interval, i.e., when the instrument is not relevant on an interval of $\eta$. In this case, $\tilde{c_0}(c_1)$ is partially identified: it is point identified everywhere except on the flat part where there exists an infinite number of solutions satisfying ((ref)). \\ Once we identify $\tilde{c}_0(c_1)$, we recover the corresponding unobserved $\eta$ as \begin{align*} \eta(c_1) = F_{C_0 |0, w}(\tilde{c}_0(c_1)) p_{D|w}(0) + F_{C_1 |1, w}(c_1) p_{D|w}(1) \quad for all c_1 \in \mathcal{C}_1, \text{ for any } w = 0,1. \end{align*} Thus we have a unique increasing solution $(\eta(c_1), \tilde{c}_0(c_1))$ for all $c_1 \in \mathcal{C}_1$. Finally, using strict monotonicity of these functions, we obtain $c_1^*(\eta) = \eta^{-1}(\eta(c_1))$ and $c_0^*(\eta) = \tilde{c_0}(c_1(\eta))$ for all $\eta$. Thus the optimal continuous policy functions $(c^*_0(\eta), c^*_1(\eta))$ for all $\eta \in [0,1]$ are identified as the unique solution to system ((ref)).

One key point in this proof is that identifying assumptions can be relaxed by exploiting the strict monotonicity of $c_d^*(\eta)$. Indeed, even though ((ref)) may have multiple non-monotone solutions, only one solution is strictly monotone. Without monotonicity, I would not obtain general point identification unless I assumed a stronger version of Assumption (ref), for example, assuming $\textrm{Pr}(D=0 | \Eta=\eta, W=1) - \textrm{Pr}(D=0 | \Eta=\eta, W=0) > 0$ for all $\eta$. This would be close to the full rank assumption on the effect of the instrument on the selection in order to identify quantile treatment effects neweypowell2003, chernozhukovhansen2005, chernozhukovhansen2006, chernozhukovhansen2008, which is in fact stronger than necessary. vuongxu2017 also exploit the power of monotonicity to relax chernozhukovhansen2005's full rank condition and still identify binary treatment effects. Their weaker condition remains at a "high level", while I show that it can be easily expressed in terms of the CCPs (Assumption (ref)).

Identification of conditional choice probabilities (CCPs)

Now that the CCCs, $c_d^*(\cdot)$ for $d=0,1$, are identified, identification of the conditional choice probabilities follows readily. Since $\Eta \sim \mathcal{U}(0, 1)$, we have $F_{C_d}(c_d) = {c_d^{*}}^{-1}(c_d)$, with corresponding density $f_{C_d}(c_d) = \partial {c_d^{*}}^{-1}(c_d) / \partial c_d$ for $d=0,1$. The CCPs are identified from

align[align omitted — 261 chars of source]

where $f_{D, C | w}(\cdot)$ is the joint conditional density of $D$ and $C$ given $W=w$, and $f_{C_d | w}(\cdot)$ is the conditional density of $C_d$ given $W=w$. \\ Alternatively, to identify the CCPs notice that, since $c^*_d(.)$ is strictly monotone, one can recover $\Eta$ from observing $(D, C)$ as $\Eta = (c^*_D)^{-1} \big( C \big).$ From there, it is as if $(D, C, W, \Eta)$ was observed from the data. Thus, the CCPs are identified once $\Eta$ is recovered from inverting the CCCs.

Identification of the payoffs

Once the optimal policies are nonparametrically identified, we can use them to identify the remaining primitives of the model. For example, the differences in payoffs between $D=1$ and $D=0$ at the corresponding optimal continuous choice are semi-parametrically identified by the CCPs using hm1993's (hm1993): the result depends on the distribution of $\Epsilon$. Generally, the payoffs can be nonparametrically identified via the CCPs and CCCs, but the identification depends on the model itself, in particular, on how the CCCs relate to the marginal utilities. Typically, in the case of dynamic problems, the marginal utility can be nonparametrically identified using the first order conditions, Euler equations and the CCCs and CCPs.

Dynamic models

In this section, I show how general single-agent (possibly non-stationary) dynamic models can be nonparametrically identified. The main idea is that these dynamic models yield intra-period problems as described in Section (ref), thus the optimal choices (CCCs and CCPs) are identified period by period following Section (ref), and we can then use these optimal choices to identify the primitives of the model bmm1997. Note that all this Framework is implicitly conditional on unobserved types $m$, which are identified beforehand, following Section (ref), and which I abstract from in the notation. The framework is general and nests many life-cycle empirical applications of interest bcms2016, iskhakov2017. I focus on the leading example of a dynamic model of labor and consumption choices.

Dynamic life-cycle model of labor and consumption

Focus on a general dynamic model of labor and consumption choices with a finite horizon ($T < \infty$) (but the arguments also apply when the horizon is infinite).\footnote{In fact, in the case of an infinite horizon with a stationary environment the identification is considerably simplified because the optimal choices are time-independent.} \\ Each period $t$ until $T$, the timing of the individual's problem is the following:

{15pt}

center[center omitted — 792 chars of source]

{21pt}

The current period conditional utility for action $(d_t, c_{dt})$ at time $t$ is given by

equation[equation omitted — 73 chars of source]

In this example, as explained earlier, $c_t$ is consumption and $c_{dt}$ are potential consumption choices, with $c_t = c_{0t} (1-d_t) + c_{1t} d_t$, $d_t$ is the labor decision to work or not, $x_t$ represents all the covariates and $w_t$ is the instrument. The covariates include variables such as age, education and other demographics, impacting current utility. For notational convenience, $x_t$ also includes variables such as assets or income which do not necessarily directly impact preferences but still have an impact on the consumption choice (and labor choice), notably through their transitions. \\ I impose additional assumptions on the current utilities which are necessary (but not sufficient) such that the intra-period problems of this dynamic setup fit into the framework of Section (ref).

assumptionbis{additive}[Additive Separability] The shock $\epsilon_t$ enters the payoff additively \begin{align*} \mathcal{U}_{dt}(c_{dt}, x_t, w_t, \eta_t, \epsilon_t) = \tilde{u}_{dt}(c_{dt}, x_t, w_t, \eta_t) + \epsilon_{dt}, \quad for d_t=0, 1. \end{align*}
assumptionbis{instrument}[Instrument] The instrument $w_t \in \mathcal{W} = \mathcal{D}$ is such that \begin{align*} \tilde{u}_{dt}(c_{dt}, x_t, w_t, \eta_t) = u_{dt}(c_{dt}, x_t, \eta_t) + m_{dt}(x_t, w_t, \eta_t), \quad for d_t=0, 1. \end{align*}
assumptionbis{monotone}[Monotonicity] The conditional current utility functions $u_{dt}$ are twice continuously differentiable and \begin{equation*} \frac{\partial^2 u_{dt}(c_{dt}, x_t, \eta_t)}{\partial c_{dt} \partial \eta_t} > 0, \quad for d_t=0, 1. \end{equation*}

In a dynamic context, the individual chooses $(d_t, c_{dt})$ to maximize her expected discounted sum of current and future payoffs. She discounts the future utilities at a rate $\beta$ and forms rational expectations about the transition probabilities. The transitions from $(x_t, w_t, \epsilon_t, \eta_t)$ and the current choices $(c_t, d_t)$ to $(x_{t+1}, w_{t+1}, \epsilon_{t+1}, \eta_{t+1})$ matter for the choices. In particular, how the current choices impact these transitions is especially important for optimal choice: for example, individuals do not consume all their wealth in a given period because they are forward-looking and want to save for the future. The impacts of the choices on the transitions are often expressed through a budget constraint like $a_{t+1} = (1+r_t) a_t - c_t + y_t d_t,$ where $a_t$ is the value of assets, $r_t$ the return on assets, and $y_t$ is labor income. For now let us be more general and only assume the existence of a general transition density of states and errors, which depend on the choices:

align*[align* omitted — 152 chars of source]

I make additional assumptions on this density for the model to be identified and to fit into the general framework.

assumption[Conditional independence] For all $x_t \in \mathcal{X}$, $w_t \in \mathcal{W}$, $\epsilon_t \in \mathcal{E}$, $\eta_t \in \mathcal{H}$, \begin{align*} &f_{X_{t+1}, W_{t+1}, \Epsilon_{t+1}, \Eta_{t+1} | x_{t}, w_t, c_t, d_t, \epsilon_t, \eta_t}(x_{t+1}, w_{t+1}, \epsilon_{t+1}, \eta_{t+1}) \\ &= f_{X_{t+1}, W_{t+1} | x_t, w_t, c_t, d_t}(x_{t+1}, w_{t+1}) \ f_{\Epsilon} (\epsilon_{t+1}) \ f_{\Eta} (\eta_{t+1}). \end{align*}
assumption[Instrument transition exclusion] For all $x_t \in \mathcal{X}$ and $w_t \in \mathcal{W}$, the current instrument is excluded from the transition density: \begin{align*} f_{X_{t+1}, W_{t+1} | x_t, w_t, c_t, d_t}(x_{t+1}, w_{t+1}) = f_{X_{t+1}, W_{t+1} | x_t, c_t, d_t}(x_{t+1}, w_{t+1}). \end{align*}

I also impose the Assumption (ref) and (ref) contemporaneously (i.e., adapted with index $t$). I do not rewrite them for simplicity of exposition. First, let me show how the intra-period problem of this dynamic model fits into the intra-period problem described in Section (ref), and then discuss further the role of these assumptions. \\ Knowing the transition densities, the individual chooses $(d_t, c_{dt})$ to sequentially maximize her expected discounted sum of payoffs. Let $V_t(z_t) = V_t(x_t, w_t)$ be the (ex ante) value function of this discounted sum of payoffs at the beginning of $t$, just before the shocks $(\epsilon_t, \eta_t)$ are revealed and conditional on behaving according to the optimal decision rule. We have

align*[align* omitted — 231 chars of source]

Given the state variable $z_t$ and choice $(d, c_{dt})$ in period $t$, the expected value function in period $t+1$ is

align*[align* omitted — 176 chars of source]

By the conditional independence (Assumption (ref)) and instrument exclusion from the transition (Assumption (ref)), we can remove $W_t$ from the conditioning variables,

align*[align* omitted — 157 chars of source]

The ex ante value function can be written recursively:

align*[align* omitted — 289 chars of source]

Thus, in each period, after observing $(\Epsilon_t, \Eta_t) = (\epsilon_t, \eta_t)$, the individual chooses $d_t$ and $c_{dt}$ to maximize her expected payoff:

align*[align* omitted — 230 chars of source]

Define the conditional value functions $v_{dt}(\cdot)$ as

align[align omitted — 174 chars of source]

Now the dynamic model yields the same maximization problem as in the general framework of Section (ref). Every period, the individual selects $d_t$ and $c_{dt}$ to solve: \\

align*[align* omitted — 154 chars of source]
Lemma[Dynamic framework] Under Assumptions (ref), (ref), (ref), (ref) and (ref), Assumptions (ref), (ref) and (ref) are satisfied for the conditional value functions defined in equation ((ref)) in the dynamic setup.

Thus, under Assumptions (ref), (ref), (ref), (ref) and (ref), all the assumptions of the general framework of Section (ref) hold (with the shocks and relevance assumptions directly adapted to the dynamic setup). If Assumption (ref) holds for the current utility function, then, by construction, Assumption (ref) will hold for the conditional value functions in Equation ((ref)). Assumptions (ref) and (ref) on the current utility do not translate directly into Assumptions (ref) and (ref) for the conditional value function. One needs additional assumptions about the transitions, i.e., Assumptions (ref) and (ref). \\ Conditional independence assumptions are standard for the identification and empirical tractability of dynamic discrete choice models rust1987, blevins2014. Here, Assumption (ref) implies that the transitions of the state variables are independent of the shocks $(\Epsilon_t, \Eta_t)$. Similarly, the shock transitions are independent of the variables here. There is no time dependence on the shocks, which are iid across periods. Crucially, here, in addition to the standard conditional independence, Assumption (ref) also implies that conditional on $(D_t, C_t, X_t)$, the transitions are independent of the current instrument value $w_t$. In particular, the instrument is excluded from its own transition to future values, conditional on $(D_t, C_t, X_t)$, i.e.,

flalign*&W_{t+1} \perp W_t \ | \ C_t=c_t, D_t=d_t, X_t=x_t &\\ or equivalently \quad \quad \ &f_{W_{t+1} | x_t, w_t, c_t, d_t}(w_{t+1}) = f_{W_{t+1} | x_t, c_t, d_t}(w_{t+1}). &

This excludes the possibility of having time-independent instrument (e.g., $w_t = w$ for all $t$). Assumption (ref) combined with Assumption (ref) will satisfy Assumption (ref) on the conditional value $v_{dt}$ as shown in the computation above. Without the exclusion of the instrument from the transition, $w_t$ could affect the expected future value function and thus enter the conditional value functions, in which case the exclusion restriction of $w_t$ from the payoff (Assumption (ref)) would not be satisfied. \\ Similarly, under Assumption (ref) and the conditional independence of the future from current $\eta_t$ (Assumption (ref)), we have

align*[align* omitted — 323 chars of source]

and the monotonicity of the conditional value functions $v_{dt}(\cdot)$, Assumption (ref), holds. \\

Discussion about the instrument. In many dynamic setups, a convenient instrument that satisfies all the assumptions could be the previous discrete choice, $W_t = D_{t-1}$. In this case, the exclusion from the transition (Assumption (ref)) is likely to be satisfied, because (i) $w_{t+1} = d_t$ in this case, thus conditional on the current $d_t$ choice, $w_{t+1}$ is known irrespective of the value of $w_t$, and (ii) conditional on the current $d_t$ and $x_t$, it is unlikely that $d_{t-1}$ impacts $x_{t+1}$. Moreover, the current exclusion restriction (Assumption (ref)) is also satisfied because conditional on $x_t$, which may, for example, include work experience, it is unlikely that $d_{t-1}$ impacts the current utility $u_{dt}(\cdot)$. Overall, the effect of $d_{t-1}$ on $c_t$ is subsumed by the effect of $d_t$ on $c_t$. Finally, $d_{t-1}$ is a relevant instrument (Assumption (ref)) if there exists an utility switching cost from exiting or entering the workforce, for example.\footnote{$W_t = D_{t-1}$ is relevant if there is some `autocorrelation' (that I interpret as switching costs) in the discrete choice (conditional on the types). This could be driven by autocorrelation in a general $\tilde{\epsilon}_{dt}(x_t, w_t, \eta_t) = m_{dt}(x_t, w_t, \eta_t) + \epsilon_{dt}$ error term. Thus, the assumption about no correlation in $\epsilon_t$ is less restrictive than it seems. } \\ The reason why I do not allow for autocorrelation in $\Eta_t$ is that I recommend to use the past discrete choice as the default instrument. Indeed, with autocorrelated $\Eta_t$, if $W_t = D_{t-1}$, then in the first period $W_1$ and $\Eta_1$ are not independent as they are both correlated with the unobserved $\Eta_{0}$. However, if one can find another instrument satisfying Assumptions (ref), (ref) and (ref) and which do not suffer from this initial period problem, we could include and identify autocorrelation in $\Eta_t$: $f_{\Eta_{t+1} | \eta_t}(\eta_{t+1})$. Because of the exclusion from its own transition (Assumption (ref)), such an instrument is hard to find in dynamic models, and would need to be a purely transitory and unexpected event. Now, with $W_t = D_{t-1}$ I can still allow for autocorrelation in the unobservables by including permanent unobserved types am2011 in the setup (Section (ref)). The types reintroduce time-dependence in the model and attenuate the effect of the conditional independence assumption.

Identification of the dynamic model

First, I show how the transitions, the CCCs and CCPs are identified in the dynamic model. Then, I show how to use them to nonparametrically identify the marginal utility, the discount factor and the conditional payoffs under additional assumptions.

Optimal choices: CCCs and CCPs

Under Lemma (ref), the dynamic framework described in Section (ref) fits into the general framework described in Section (ref). Therefore the CCCs and CCPs are identified period by period from $t \geq 2$ onwards, following the proof developed in Section (ref). The data $\{ D_t, C_{t}, X_t, W_t, t\}_{t=1}^T$ provides the following reduced-form functions:

align*[align* omitted — 292 chars of source]

From these reduced forms, following Section (ref), I identify the CCCs and CCPs

align*[align* omitted — 108 chars of source]

for all $(d_t, \eta_t, x_t, w_t, t) \in \mathcal{D}\times [0, 1] \times \mathcal{X}_t\times\mathcal{W}\times \{2,..., T\}$. Note that identification does not hold for $t=1$ because $W_t=D_{t-1}$ is not available for $t=1$. \\

Special case: Identification of the choices with terminal/absorbing actions. \\ Suppose $D_t = 1$ is a terminal action or an absorbing state. For example, $D_t = 1$ if the individual retires, $D_t=0$ if she stays active. Assuming that an individual cannot go back to working life, the retirement choice is absorbing iskhakov2017, levy2024identification. Now, identification is greatly simplified. Indeed, use $W_t = D_{t-1}$ as the instrument. When $D_{t-1} = 1$, the instrument is `infinitely' relevant: the probability of staying retired is one. Thus by focussing on previously retired individuals ($W_t=1$), equation ((ref)) gives

align*[align* omitted — 136 chars of source]

Since $F_{C_1 | 1, X_t, 1, t}(c)$ is invertible (Lemma (ref)), we recover the continuous choices conditional on being retired as:

align*[align* omitted — 138 chars of source]

It remains to identify the other conditional continuous policy. Take equation ((ref)) at $W_t = 0$, i.e., for individuals who did not select the absorbing state yet. It yields {

flalign*\eta_t =&\quad F_{C_{0t} | 0, x_t, 0}(c^*_{0t}(\eta_t, x_t)) Pr(D_t=0|X_t=x_t, W_t=0) \\ &+ \quad F_{C_{1t} | 1, x_t, 0}(c^*_{1t}(\eta_t, x_t)) Pr(D_t=1|X_t=x_t, W_t=0) &\\ and hence, & &\\ c^*_{0t}(\eta_t, x_t) =& F^{-1}_{C_{0t} | 0, x_t, 0} \left( \frac{\eta_t - F_{C_{1t} | 1, x_t, 0}(c^*_{1t}(\eta_t, x_t)) Pr(D_t=1|X_t=x_t, W_t=0)}{Pr(D_t=0|X_t=x_t, W_t=0)} \right). &

} This identifies the optimal continuous choice of active individuals ($D_t=0$), since all the terms on the right hand side are already known. Once the CCCs are identified, we proceed as previously to identify the CCPs.

Transitions

The transition density $f_{X_{t+1} | X_t=x_t, C_t=c_t, D_t=d_t}(x_{t+1})$ is identified directly from the data by observing the conditional transitions of the variables between consecutive periods $t$ and $t+1$. The transition of the instrument is known by construction if $w_{t+1} = d_t$. With other instruments, it can also be recovered from the data. As is standard in the dynamic choice literature, I assume individuals are rational, so that the observed transition density coincides with the transition densities expected by the individuals. Then, the transitions recovered from the data can be used to build the individual's expectations at each time $t$, and help recover the primitives.

Primitives

Once the CCCs, CCPs and transitions are identified, I build upon existing literature to identify the primitives of the model hm1993, bmm1997, mt2002, escancianoetal2021. I need to introduce additional structure on the covariates' transition and on the current utility function for nonparametric identification of the utility and the discount factor. \\ Budget constraint. The asset transitions are given by the budget constraint\footnote{The budget constraint could be more sophisticated and include taxes, benefits... The effect of income could be a general function $\kappa_{d_t}(y_t)$, where non-working individuals would still receive a part of their income. See the simulations for another example with part-time versus full-time work.}

equation[equation omitted — 71 chars of source]

The asset here plays a different role than the other covariates. Indeed, the transition from $a_t$ to $a_{t+1}$ is directly impacted by the choice $c_t$ through the budget constraint ((ref)). Denote the covariates as $x_t = (\tilde{x}_t, a_t)$ to emphasize the distinct role of the asset.\footnote{Note that $y_t$ and $r_t$ are included in $\tilde{x}_t$, even though, in most applications, they will also be excluded from the current period utility. For notational simplicity and generality, I include them in $\tilde{x}_t$, which enters the current utility and represents all covariates other than $a_t$, i.e., all covariates whose transitions are not impacted by $c_t$ (Assumption (ref)). }

assumption[Asset exclusion] The asset is excluded from the current period utility, i.e., $\partial u_{dt}(c_{dt}, \tilde{x}_t, a_t, \eta_t)/\partial a_t = 0$.
assumption[General covariates transitions] For all $\tilde{x}_t \in \tilde{\mathcal{X}}_t, d_t \in \mathcal{D},$ and $c_t \in \mathcal{C}$, $c_t$ does not impact the transitions of $\tilde{x}_t$ and $w_t$, i.e., \begin{align*} f_{\tilde{X}_{t+1}, W_{t+1} | \tilde{x}_{t}, c_t, d_t}(\tilde{x}_{t+1}, w_{t+1}) = f_{\tilde{X}_{t+1}, W_{t+1} | \tilde{x}_{t}, d_t}(\tilde{x}_{t+1}, w_{t+1}). \end{align*}
assumption[Stationary utility] The current period utility is independent of time, i.e., $u_{dt}(c_{dt}, x_t, \eta_t) = u_d (c_{dt}, x_t, \eta_t)$.
assumption[Monotone utility] The utility is strictly increasing in $c$, i.e., \begin{align*} \frac{\partial u_{dt}(c_{dt}, x_t, \eta_t)}{\partial c_{dt}} > 0 \quad for all (d, c_{dt}, x_t, \eta_t) \in \mathcal{D} \times \mathcal{C}_{dt} \times \mathcal{X}_t \times [0, 1]. \\ \end{align*}
Lemma[Marginal utilities and discount factor] Following escancianoetal2021, under Assumptions (ref)-(ref) and (ref)-(ref), the conditional marginal utilities at the optimal continuous choices, \begin{align*} u'^*_{d}(x_t, \eta_t) = \frac{\partial}{\partial c_{dt}} u_{d}(c_{dt}, x_t, \eta_t) |_{c_{dt} = c^*_{dt}(\eta_t, x_t)}, \end{align*} and the discount factor $\beta$ are nonparametrically point identified by the Euler equation for all $(d, x_t, \eta_t) \in \mathcal{D} \times \mathcal{X}_t \times [0, 1]$.
proofSince $a_t$ is excluded from $u_d(\cdot)$ by Assumption (ref), $u'^*(\cdot)$ only depends on $a_t$ through the optimal CCCs, . Then, given the budget constraint and since $c_t$ only affects the asset transition (Assumption (ref)), the Euler equations are, for $d_t=0,1$, \begin{equation} u'^*_{d_t}(x_t, \eta_t) = \ \beta (1+r_t) \mathbb{E}_t \Big[ u'^*_{D_{t+1}}(X_{t+1}, \Eta_{t+1}) \ \Big| X_t=x_t, C_{t}=c^*_{dt}(\eta_t, x_t), D_t=d_t \Big]. \end{equation} We have a system of two equations with two unknown functions $u'^*_0(\cdot)$ and $u'^*_1(\cdot)$ (and the unknown discount factor $\beta$). Hence the importance of stationarity (Assumption (ref)), since otherwise we would have a different unknown function on each side of the equation. Now, under Assumptions (ref) and (ref), the optimal marginal utilities are positive, \begin{align*} u'^*_d(x_t, \eta_t) > 0 \quad for all d, x_t, \eta_t. \end{align*} Now, Theorem $2$ of escancianoetal2021 shows that the discount factor $\beta$ and the marginal utility functions are nonparametrically globally point identified by the system of Euler equations ((ref)).

Once the marginal utilities are identified, I follow bmm1997 to identify the conditional value functions. Note that even though the marginal utilities are stationary, we still have a non-stationary problem because the conditional value functions are time-dependent with finite horizon.

Lemma[bmm1997] Under Assumptions (ref)-(ref) and (ref)-(ref), the conditional value functions at optimal choices, $v_{dt}(c^*_{dt}(\eta_t, x_t), x_t, \eta_t)$, are identified up to an unknown constant of integration $O_{dt}$ independent from the asset, i.e., \begin{align*} v_{dt}(c^*_{dt}(\eta_t, x_t), x_t, \eta_t) = G_{dt}(\tilde{x}_t, a_t, \eta_t) + O_{dt}(\tilde{x}_t, \eta_t) \quad for all d, x_t, \eta_t, \end{align*} where $G_{dt}$ and $K_{dt}$ are defined in the proof.
proofWe have the first order conditions, holding at optimal CCCs for all $d$: \begin{equation} \frac{\partial}{\partial a_t} v_{dt}(c_{dt}, \tilde{x}_t, a_t, \eta_t) = (1+r_t) \frac{\partial}{\partial c_{dt}} u_{d}(c_{dt}, x_t, \eta_t) \ |_{c_{dt} =c^*_{dt}(\eta_t, \tilde{x}_t, a_t)}. \end{equation} With $v^*_d(\cdot)$ as the conditional value function taken at the optimal continuous choice, we can rewrite the FOC as \begin{align} \forall a_t: \quad \quad \frac{\partial}{\partial a_t} v^*_{dt}(\tilde{x}_t, a_t, \eta_t) &= (1+r_t) \ u'^*_{d}(\tilde{x}_t, a_t, \eta_t). \end{align} Crucially, following Assumption (ref), the asset is excluded from the current period utilities and marginal utilities. The identification strategy relies on this exclusion. Now, integration gives \begin{align*} v^*_{dt}(\tilde{x}_t, a_t, \eta_t) &= \int^{a_t}_0 \ (1+r_t) \ u'^*_{d}(\tilde{x}_t, a, \eta_t) \ da, \end{align*} where the lower bound $0$ is taken arbitrarily. Since $u'^*_d$ is identified, we can identify the optimal conditional value functions nonparametrically as \begin{align*} v^*_{dt}(\tilde{x}_t, a_t, \eta_t) = G_{dt}(\tilde{x}_t, a_t, \eta_t) + O_{dt}(\tilde{x}_t, \eta_t), \end{align*} up to unknown constant of integration $O_{dt}(\tilde{x}_t, \eta_t)$, independent from $a_t$ and depending on the arbitrary lower bound of integration.

Finally, by specifying a distribution for $\Epsilon$ (e.g., generalized extreme value), the differences in the additive terms of the utility, $\Delta m_t(\cdot) = m_{1t}(\cdot) - m_{0t}(\cdot)$, are semi-parametrically identified. Indeed, the difference in total conditional values, $\Delta v_t^*(\cdot) + \Delta m_t(\cdot)$, are identified by the CCPs through hm1993's (hm1993) inversion. Thus, if I impose a normalization of the constant, e.g. $O_{dt} = 0$, $\Delta m_t(x_t, w_t, \eta_t)$ is identified.

Unobserved types

So far, I assumed no autocorrelation in purely transitory $\Eta_t$ (Assumption (ref)), in order to be able to use the previous discrete choice as a relevant instrument ($W_t=D_{t-1}$) without violating the independence between the instrument and $\Eta_t$. However, including only iid transitory period-specific shocks is fairly restrictive in dynamic models, where we often observe serial correlation in the choices. I handle this by including permanent unobserved types into the model, following the standard approach in the dynamic discrete choice literature am2011. These types capture intrinsic latent differences between individuals, while $\Eta$ and $\Epsilon$ are transitory shocks affecting the decisions. In this section, I show how to adapt the identification arguments with unobserved types, by identifying the unobserved types beforehand.

Identification with unobserved types

I assume throughout this section that $W_t = D_{t-1}$.\footnote{If $W_t$ is not $D_{t-1}$ and is a period-$t$ variable, then the identification of unobserved types still holds. It is simplified and only requires $T \geq 3$ time periods in the panel, as in Section 3.1 of kasaharashimotsu2009. } I observe panel data $\{D_t, C_{t}, X_t\}_{t=1}^{T}$ with $T\geq 6$ and $C_t = C_{0t} (1-D_t) + C_{1t} D_t$. The instrument $W_t=D_{t-1}$ is included in the observations of $\{D_t\}_{t=1}^{T}$. Each individual has a time-invariant/permanent type $\mu$ with finite values $m \in \{1, ..., M\}$. The type is unobserved by the researcher. The probability of belonging to type $m$ is $\text{Pr}(\mu=m) = \pi^m$ and is time-invariant and independent of the covariates.\footnote{The setup can be extended to allow for time-varying types (e.g., first-order Markov), time-varying type probabilities, as well as type probabilities that depend on the covariates, using kasaharashimotsu2009 and hushum2012.}

Adaptation of the framework with types

The adjustments to include types are fairly straightforward. Types act similarly to a covariate in $X$, except that it is unobserved by the researcher. The functions $u_{dt}(c_{dt}, x_t, m, \eta_t)$, $m_{dt}(x_t, m, w_t, \eta_t)$, $V_t(x_t, m, w_t)$, and $v_{dt}(c_{dt}, x_t, m, \eta_t)$ are all type-dependent, and I now make this dependence explicit by writing them with an $m$ supperscript as $u_{dt}^m(\cdot), m_{dt}^m(\cdot)$, $V_t^m(\cdot)$ and $v_{dt}^m(\cdot)$. Assumption (ref) now conditions on $X=x$ and $\mu = m$. For simplicity, the covariate transition densities are assumed to remain type-independent: $f_{X_t | x_{t-1}, c_{t-1}, d_{t-1}}(x_t)$ for all $m$.\footnote{Again, this can be relaxed and we can identify type specific transitions $f_t^m(x_t | X_{t-1}=x_{t-1}, C_{t-1}=c_{t-1}, D_{t-1}=d_{t-1})$, following Section $3.2$ of kasaharashimotsu2009.} I only add one assumption on how types enter the model.

assumption[Type-independent Support] Types enter the utilities of the model in a way such that, for all $t, c_t, d_t, x_t$, \begin{align*} f_{D_t, C_t | x_t, d_{t-1}}(d_t, c_{t}) > 0 \iff f^m_{D_t, C_t | x_t, d_{t-1}}(d_t, c_{t}) > 0, for all m=1,...,M. \end{align*}

Here, $f^m_{D_t, C_t | x_t, d_{t-1}}(\cdot)$ is the type-dependent conditional joint density. Assumption (ref) restricts how types enter the utility functions: it must not affect the support of the optimal choices, especially the continuous one. In terms of identification, it means that any possible observation $(d_t, c_{t}, x_t)$ can come from any type $m \in \{1, ..., M\}$.

Identification of the type-dependent conditional joint densities

Given the framework, the joint densities of the choices depend on $X_t=x_t, W_t=d_{t-1}$ and now also depend on the type $\mu=m$. So the reduced form to identify the optimal choices are now type-specific and not directly observable from the data: I need to identify them first to identify the optimal choices following Section (ref) afterwards. \\

Result: If $T \geq 6$, the type probabilities, $\pi^m$, and the type-dependent conditional joint densities, $f^m_{D_t, C_t | x_t, d_{t-1}}(d_t, c_t)$, are identified from observed serial data $\{D_t, C_{t}, X_t\}_{t=1}^{T}$ for all $m, t, d_t, c_{t}, x_t, d_{t-1}$. \\

The idea is to identify the unobserved types by using the identification power of the observed serial correlations of $\{D_t, C_{t}, X_t\}_{t=1}^{T}$. Notice that, except for the first-order autocorrelation between $d_t$ and $d_{t-1}$ (relevance condition), I did not use the observed autocorrelations of the choices to identify the dynamic model before. This is the reason why I can exploit them to identify the type-specific conditional joint densities in a first step, independent of the rest of the identification (which proceeds period by period). Formally, to show the identification of the type-dependent conditional joint densities, I extend the identification proof of kasaharashimotsu2009 to joint choices with both time-dependent conditional choice probabilities and a lagged dependent variable. I also make specific adjustments because my covariates include the value of assets which has a deterministic transition given the choices, violating Assumption $1 (c)$ in kasaharashimotsu2009. The identification proof is given in Appendix (ref).

Identification of the dynamic model with unobserved types

I obtain the type-dependent reduced-form functions from the type-dependent joint choices densities for all $m \in \{1, ..., M\}$

align*[align* omitted — 298 chars of source]

From these $R^m$, following Section (ref), I identify type-dependent CCCs and CCPs

align*[align* omitted — 113 chars of source]

for all $m \in \{1, ..., M\}$ and $(d_t, \eta_t, x_t, w_t, t) \in \mathcal{D}\times [0, 1] \times \mathcal{X}_t\times\mathcal{W}\times \{2,..., T\}$. \\ The transitions are type-independent by assumption, so I identify them directly from the data as before. Then, the identification of the primitives of the dynamic model follows Section (ref), replacing the optimal choices by their type-dependent counterparts, and conditioning everything on the type $m$.

Estimation

I build a two-step estimation procedure. In the first stage, I estimate type-dependent conditional continuous choices (CCCs) and conditional choice probabilities (CCPs). First, I estimate the type probabilities using an expectation-maximization (EM) algorithm aj2003, am2011. Then, I estimate the type-dependent optimal choices given these type probabilities. This step is data-driven and is independent of the structural model specification. In the second stage, I use these estimated optimal policies to estimate the primitives (structural parameters) of the model. To do so, I use the fact that the optimal choices are obtained via the optimality conditions of the model taken at the true parameters: the true parameters are the only parameters that generate these optimal policies, and satisfy the optimality conditions taken at these true policies. The estimation is analogous to that of hm1993, am2011 and hmss1994 but extended to discrete-continuous choice models. \\ The main appeal of this estimation is computational gains. By estimating the optimal choices only once, directly from the data, and taking them as given in the next stage, the computational burden of the estimation is significantly reduced. Indeed, one does not need to solve for the value function or the likelihood for each new set of selected parameters. This allows us to estimate models that were previously computationally intractable. I describe the estimation method in this section, and show the estimator's performance using Monte Carlo simulations in Section (ref).

1st step: conditional choices

Use $W_t = D_{t-1}$ as suggested before. First, I estimate the type-independent covariates transitions directly from the data. Asset transition is known by the budget constraint, instrument transition is known since $w_t = d_{t-1}$, and the other covariate transitions can be estimated using auto-regressive processes of order $1$. This yields the transitions $\hat{f}_{X_t | x_{t-1}, c_{t-1}, d_{t-1}}(x_t)$, where $c_{t-1}$ only affects the asset transition (Assumption (ref)). Then, I estimate the type-dependent reduced forms using an expectation-mazimization (EM) algorithm in the spirit of am2011 (Section (ref)). Using these type-specific probabilities for each individuals, I estimate type-specific CCCs and CCPs building upon the identification arguments (Section (ref)).

EM algorithm for type-dependent reduced forms

Suppose the type-dependent joint densities $f^m_{D_t, C_t | x_t, d_{t-1}}(d_t, c_t, \theta_r)$ are fully parametrized by $\theta_r$.\footnote{Note that this includes the initial period joint density $f_{D_1, C_1, X_1}^m(d_1, c_1, x_1, \theta_{\text{init}})$ when the instrument $D_{0}$ is unobserved.} We want a nonparametric sieve-estimator where the number of parameters in $\theta_r$ increases with the sample size. To estimate $\theta_r$ and $\pi^m$ (the type probabilities) we proceed by iteration, starting from an initial guess $(\theta_r^{(k)}, \pi^{(k)})$, with $k=1$, where $\pi^{(k)} = \{{\pi^{m}}^{(k)}\}_{m=1}^M$. \\

Expectation step. Given the $k^{th}$ guess, the likelihood of observing $\{ d_{it}, c_{it}, x_{it}\}_{t=1}^T$ given the type $m$ for individual $i$ is

align*[align* omitted — 246 chars of source]

Then, the likelihood of observing $\{ d_{it}, c_{it}, x_{it}\}_{t=1}^T$ for $i$, unconditional on type, is

align*[align* omitted — 98 chars of source]

The updated likelihood that individual $i$ belongs to type $m$, denoted $q_i(m)$, is

align*[align* omitted — 93 chars of source]

Given a sample of $N$ individuals, we update $\pi^{(k)}$ to $\pi^{(k+1)}$ for each type $m$ as

align*[align* omitted — 69 chars of source]

Maximization step. Given $q^{(k+1)}= \{q_i^{(k+1)}(m)$ for all $m\}_{i=1}^N$, we can compute the sample likelihood for any $\theta_r$:

align*[align* omitted — 98 chars of source]

We update $\theta_r^{(k)}$ to $\theta_r^{(k+1)}$ by finding the $\theta_r$ which maximizes the log-likelihood

align*[align* omitted — 107 chars of source]

Notice that the empirical conditional joint densities weighted by $q^{(k+1)}$ maximize the log-likelihood am2011. Thus we can directly nonparametrically estimate it, without running any numerical optimization algorithm. \\

EM estimation. Select initial values $(\theta_r^{(1)}, \pi^{(1)})$. For example, randomly assign a type to every individual and estimate the initial joint densities given this guess to obtain the initial values. Starting from these initial values and iterating the expectation and maximization steps, the EM algorithm converges to $(\hat{\theta}_r, \hat{\pi})$, which maximizes the likelihood of the sample. The estimates $(\hat{\theta}_r, \hat{\pi})$ give estimates of the type-dependent reduced forms $\widehat{R}^m$ and provide estimates of the type probabilities of each individual, $\hat{q_i}(m)$, that we use in the next steps.

Type-dependent CCCs and CCPs

Once the type-dependent probabilities $\hat{q_i}(m)$ are estimated, we can use them to estimate the type-dependent optimal choices. \\

Conditional continuous choices (CCCs). We build upon the link between the intra-period problem and the IV-Quantile model of chernozhukovhansen2005 established in Section (ref), and adapt existing IVQR estimation procedures chernozhukovhansen2006, kaido2021decentralization to estimate the CCCs. More precisely, since $W_t=D_{t-1}$ is a valid instrument, we estimate $c_{0t}^m(\Eta_t, x_t)$ and $c_{1t}^m(\Eta_t, x_t)$ for any rank $\Eta_t$ and for all period $t$, covariates $x_t$, and type $m$, by running the weighted IV-Quantile regression of $C_t$ on $D_t$ at each quantile $\Eta_t$, conditional on $X_t=x_t$ and weighted by the estimated type-$m$ probabilities $\hat{q_i}(m)$.\footnote{There are several manners to condition on the covariates $X_t$. Typically, for discrete covariates, we can run separate IV-quantile regressions on each subsamples with $X_t=x_t$, provided that these subsamples contain enough observations. Otherwise, continuous covariates (e.g., the assets) enter additively in the IVQR specification, which effectively restricts the heterogeneity of the effect of $D$ on $C$ with respect to these covariates at each quantile. In theory, we could also split the continuous covariate in subgroups with sufficiently enough observations. Or we could do kernel-based IV quantile regressions to account for these continuous covariates nonparametrically with weights, but this requires a large sample. } This IVQR approach allows to flexibly estimate heterogenous effects of $D$ on $C$ at each quantiles $\Eta$. \\

Conditional choice probabilities (CCPs). Once the CCCs are estimated, one can invert them to estimate the unobserved shock $\eta_{it}$ for every individual, i.e.,

align*[align* omitted — 92 chars of source]

Using these estimated unobserved individual shocks $\hat{\Eta}_{it}$ as a generated covariate, we can directly estimate the type-dependent CCPs, $\textrm{Pr}^m(D_t | \hat{\Eta}_{t}, X_t, W_t)$, using weighted nonparametric kernels or flexible weighted logit/probit regressions of $D_t$ on $X_t, W_t$ and $\hat{\Eta}_t$, weighted by type-$m$ probabilities, $q_i(m)$.

Alternative estimation methods

There are many alternative ways to estimate the optimal choices. For example, the CCCs can be estimated nonparametrically or semi-parametrically by building upon the identification arguments using empirical counterparts of the functions $\Delta F_{C_d}^m(\cdot)$, obtained using the estimated type-dependent reduced forms. This alternative approach has the advantage of working under weaker relevance conditions than the ones imposed by chernozhukovhansen2005, i.e., even with piecewise monotone $\Delta F_{C_d}^m$ functions, as in Figure (ref).

2nd step: structural model

Suppose the primitives can be fully parametrized by $\theta = (\beta, \theta_C, \theta_P)$, where $\theta_C$ characterize the marginal utility with respect to the continuous choice and $\theta_P$ does not.\footnote{More precisely, assume there is a one-to-one mapping between the parameters $\theta$ and the primitives of the model, i.e., each different value of $\theta$ generates different primitives.} In other words, $u^m_d(\cdot, \theta_C)$ is parametrized by $\theta_C$, while $\theta_P$ only impacts the difference $\Delta m_{t}^m(\cdot, \theta_P)$. Denote $\theta^0$ the true parameters that generated the data. \\ Using the nonparametric identification arguments developed previously (Section (ref)), there is a one-to-one mapping between the primitives of the model and the optimal choices (CCCs and CCPs). Each set of parameters $\theta$ characterizing the primitives of the model is associated with distinct optimality conditions (e.g., Euler equations and differences of conditional value functions) which, in turn, yield distinct optimal choices. Consequently, the true CCCs and CCPs which have been consistently estimated directly from the data in the first stage, can only be rationalized by the true value of the parameters, $\theta^0$. In theory, we could estimate the model using standard methods of simulated moments with these CCCs and CCPs as the moments.\footnote{An even more standard approach would be to use moments directly available in the data to estimate the model, e.g., observed quantiles of $C$ given $D$, $X$, and $W$, and estimated probability of selecting $D$ given $X$ and $W$. The key take-away from the identification being that one needs to use moments which depends on the instrument $W$, otherwise the model would not be identified. } A typical method of simulated moment estimator would be as follows: for each value of the parameter $\theta$, one would compute the optimal value function and the corresponding theoretical optimal choices. Then, the estimated $\theta$ would be the set of parameters which make these theoretical optimal choices the closest to the true observed optimal choices moments estimated in the first stage. While theoretically simple, this standard simulated method of moments is impractical for dynamic models. Indeed, even the fastest methods to compute the value functions, namely the endogenous grid method carroll2006, iskhakov2017, is still long, even for only dynamic discrete choice models, and even more so for dynamic discrete and continuous choice models which require additional numerical optimization to solve for the optimal continuous choices. \\ Fortunately, extending what hm1993, hmss1994, am2011 have proposed to estimate dynamic discrete choice models, I propose a faster alternative estimation method that does not require to compute the value function and numerically solve for the optimal choices for each evaluated set of parameters. The key intuition is to directly use the link between the optimality conditions of the model and the optimal choices. Given the known (estimated in the first stage) optimal choices, the first order conditions are only satisfied for the true value of the parameters, $\theta^0$. So, we estimate these true parameters by minimizing the error in the first order conditions where we plugged-in the known optimal choices. The computational difficulty is that these first order conditions involve expectations about the future. In order to compute these, we use forward simulations, as hmss1994 did for dynamic discrete choice models. I split the estimation into two types of first order conditions (i) the Euler equation which determines the CCCs and will allow to estimate $\theta_C^0$ and $\beta^0$, and (ii) the conditional value function comparison which determines the CCPs. The full estimation is described below. \\

Moment selection. Select a set of $S$ moments corresponding to $S$ covariates values: $\{ t^s, \eta^s, x^s, w^s, m^s\}_{s=1}^S$. The CCCs and CCPs have been consistently estimated in the first stage, so these moments corresponds to moments expressed in terms of $C$ and $D$, i.e., $c_d^s = c_{dt^s}^{m^s}(\eta^s, x^s)$ for each $d \in \mathcal{D}$ and $p_d^s = \textrm{Pr}^{m^s}(D_{t^s}=d | \Eta_t = \eta^s, X_t=x^s, W_t = w^s)$. The set of moments needs to be large enough such that there is a one-to-one mapping between the model parameters $\theta$ and all the moments.\footnote{It is possible that two different set of parameters are observationally equivalent locally, for some CCCs and CCPs taken at specific values of the covariates and type. However, if the model is properly parametrized (with no "redundant" parameters), there do not exist two distinct sets of parameters that yield observationally equivalent CCCs and CCPs for every values of $t, X, W,$ and $m$. One cannot test every value of these covariates as moments, but one needs to take sufficiently many different moments such that there is only one optimal set of parameters that generate them.} If the model specification is correct, the optimal choices estimated in the first stage are generated by the true parameters, $\theta^0$. Furthermore, there is a one-to-one mapping between the model and the optimal choices, and these observed moments can only be rationalized by the true $\theta^0=(\beta^0, \theta_C^0, \theta_P^0)$, and no other value of $\theta$. \\

Euler objective, estimation of $(\theta_C, \beta)$. Recall that under the true model $\theta^0$, Euler Equation (ref) holds, i.e., for any moment $s$ with $\{t^s, \eta^s, x^s, w^s, m^s\}$,\footnote{For the Euler equation, the value of the instrument, $w^s$, in the list of moment does not matter because it is excluded from the optimal CCC.} and each $d \in \mathcal{D}$, \\

adjustwidth{-0.5cm}{-0.5cm} \begin{align} & u^{{m^s}'^{**}}_{d}(x^s, \eta^s, \theta_C^0) = \ \beta (1+r_{t^s}) \mathbb{E}_{t^s} \Big[ u^{{m^s}'^{**}}_{D_{t^s+1}}(X_{t^s+1}, \Eta_{t^s+1}, \theta_C^0) \ \Big| X_{t^s}=x^s, C_{dt^s}=c^{m^s}_{dt}(\eta^s, x^s), D_{t^s}=d_t \Big] \nonumber \\ &\overset{def}{\iff} \quad \quad b_{1}(d, t^s, \eta^s, x^s, m^s, \theta_C^0)\ = \ b_{2}(d, t^s, \eta^s, x^s, m^s, \theta_C^0, \beta^0). \end{align}

The functions $u^{{m^s}'^{**}}_{d}(\cdot, \theta_C)$ are the marginal utility taken at the optimal choices, taking the optimal choices as estimated in the first stage, $c^{m}_{dt}(\eta_t, x_t)$, i.e.,

align*[align* omitted — 170 chars of source]

Regardless of the parameters $\theta_C$, the function is evaluated at the true $c^{m}_{dt}(\eta_t, x_t)$ which correspond to the true parameters, $(\theta_C^0, \beta^0)$. This Euler equation uniquely determines the CCCs. Given the nonparametric identification, and the uniqueness of the mapping between the optimal choices and the primitives of the model, there is no alternative set of parameters $\tilde{\theta}_C \neq \theta_C^0$ such that this equation (ref) would hold for all moments $s$. This is because, here we plugged-in the optimal choices of the first stage which correspond to $\theta^0$, and any distinct set of parameters $\theta$ would require different CCCs and CCPs in order to hold for all $s$, due to the uniqueness of the optimum. As a consequence, the idea behind the estimation is to minimize the difference between both sides of the Euler equation, $b_{1}$ and $b_{2}$, taken at the optimal choices estimated in the first stage.\footnote{Equivalently, recall that the marginal utilities are strictly increasing in $C$ by monotonicity. So one could express the Euler equation not in terms of the marginal utilities, but in terms of the optimal $C$ they determine. Then, the objective of the estimator is to find the true values of $\theta_C$ and $\beta$ which minimize the difference between the estimated CCCs in the first stage, and the corresponding theoretical CCCs pinned down by the Euler equation for any moment $s$. } While the left hand side of the Euler equation, $b_{1}(\cdot)$, can be directly estimated consistently for any $\theta_C$ by plugging in the first stage optimal CCCs, the right hand side $b_2(\cdot)$ contains an expectation over the next period optimal marginal utilities. To compute it, we use one-period ahead forward simulations, using the estimated covariates transitions and the next-period type-dependent optimal choices (CCCs and CCPs) estimated in the first stage. Then $(\theta_C, \beta)$ can be estimated by minimizing the sum of squared differences $\widehat{b_{1}}(\cdot) - \widehat{b_{2}}(\cdot)$ over all moments $s$ and alternative $d$, i.e.,

align[align omitted — 283 chars of source]

This `Euler objective' (ref) consistently estimates $\theta_C$ and $\beta$. Indeed, since the optimal choices are consistently estimated in the first stage, the minimum should only be reached at the true values of the primitive parameters $\theta_C^0$ and $\beta^0$ according to the Euler equation (ref). \\

Probability objective, estimation of $\theta_P$. Following a similar intuition, we can semi-parametrically estimate the remaining parameters $\theta_P$ impacting the differences of the additive term, $m_d(\cdot)$ but not the marginal utility with respect to the continuous choice. Given a known distribution of $\Epsilon$, there is a one-to-one mapping between the CCPs and the conditional value functions which are determined by the primitives of the model, $\theta^0$. This is the hm1993's (hm1993) inversion. The only adjustment with respect to hm1993 is that the mapping is with respect to the conditional value functions taken at the optimal continuous choice, denoted $v_{dt}^{m*}(\cdot)$. For example, if $\Epsilon$ is extreme-value type I, for any moment $s$ with with $\{t^s, \eta^s, x^s, w^s, m^s\}$, we know that the CCPs estimated in the first stage satisfy

align[align omitted — 481 chars of source]

where $v_{dt}^{m*}(\cdot, \theta)$ are the conditional values taken at the true optimal choices estimated in the first stage but with parameter $\theta$.\footnote{Attention, as for the optimal marginal utilities, except if $\theta = \theta^0$, these are not the traditional `conditional values functions'. This is because, we plug-in the optimal choices estimated in the first stage (which correspond to the true parameters $\theta^0$) in all the future periods, and not the optimal choices corresponding to $\theta$. The choice of $\theta$ only affects how much utility is derived from these already given optimal choices every period. } Since hm1993's (hm1993) mapping is unique, with the choices estimated in the first stage, (ref) only holds at the optimal value of the parameters, $\theta^0$ for all moments $s$. We use this known link between the CCPs and the primitives of the model to estimate $\theta=(\beta, \theta_C, \theta_P)$. In fact, we take $\widehat{\theta}_C$ and $\widehat{\beta}$ estimated via the Euler equation as given, and use (ref) to estimate the remaining parameters, $\theta_P$.\footnote{In theory, one could estimate all the parameters $\theta$ using only this probability criterium. I recommend to split in two separate estimation steps because the parameters $\theta_C$ and $\beta$ have a larger impact on the Euler equation, and are more precisely estimated using the Euler-criterium. Moreover, the Euler equation estimation is faster because it involves only one-period ahead simulations, while the probability estimation requires forward simulations of the complete remaining life-cycle. } \\ In order to estimate $p^{model}(\cdot, \hat{\beta}, \hat{\theta}_C, \theta_P)$ for any $\theta_P$, we need to compute these conditional value functions at the true optimal choices for any moment $s$. These conditional values contain expectations about the next period value function, and thus about the entire future life-cycle of individuals from time $t^s$ onwards, taking the true optimal choices (estimated in the first stage) as given. In order to estimate these expectations without solving for the value function, I follow the insights of hmss1994 and use forward simulations of the entire remaining life-cycle of individuals.\footnote{Forward simulation details. For any $d \in \mathcal{D}$, recall that the conditional value functions at the optimal choices estimated in the first stage are defined as

align[align omitted — 285 chars of source]

where for any $(t, m, x_t, \eta_t)$, the value functions given the optimal first stage choices are given by

align[align omitted — 444 chars of source]

These value functions at the optimal choices can be estimated by forward simulating the entire life-cycle of individuals from $t$ onwards. For a large number of simulations, $N_S$, we draw new state variables and unobserved $\Eta_\tau$ and $\Epsilon_\tau$ each period, given the previous state variables and choices. Given these new state variables, we use the true optimal choices which have been consistently estimated in the first stage, to draw the discrete and continuous choices of the period. We repeat this process every period until period $T$ is reached (or, if the horizon is infinite, until the discount is so large that the additional period has negligeable impact on the value). Then, given the entire pre-simulated histories of $\{ \{ C^k_\tau, D^k_\tau, X^k_\tau, W^k_\tau, \Eta^k_\tau, \Epsilon^k_\tau \}_{\tau = t+1}^T \}_{k=1}^{N_S}$ for $N_S$ different simulations, for any value of $\theta$ we can compute the corresponding period utility every period, and thus, the value of every simulated life-cycle. Taking the average of these values over the $N_S$ simulations provides an estimate of the value function given $\theta$, described in Equation (ref). Once the value functions are estimated, we can recover the conditional value functions (ref), and then the theoretical CCP, $p^{model}(\cdot, \theta)$ for any parameter $\theta$ and moment $s$. } Thanks to these forward simulation, we can estimate $\widehat{p^{model}}(\cdot, \widehat{\beta}, \widehat{\theta}_C, \theta_P)$ for any $\theta_P$ and for all moments $s$. Then, we estimate $\theta_P$ as the unique set of parameters satisfying the theoretical property (ref). In practice, we estimate $\theta_P$ by minimizing the sum of squared differences over all moments $s$, between the optimal CCPs which have been consistently estimated in the first stage, and their theoretical counterpart in the model, $\widehat{p^{model}}(\cdot, \theta_P)$, i.e.,\footnote{Note that we take the sum over all $d \in \mathcal{D}$ except the reference $d=0$, since all probabilities sum to one so one of the alternative is redundant.}

align[align omitted — 368 chars of source]

This probability objective (ref) consistently estimates the remaining parameters $\theta_P$ as the minimum should only be reached at the true values of $\theta_P^0$ according to (ref). \\

Additional remarks. This second stage estimation is completely independent of the data, the data only affected the estimation of the optimal policies in the first stage. In this sense, it is similar to the method of simulated moments where the data only affects the estimation of the moments. Given the first stage estimated policies, the second stage estimates only depend on the selected set of moments and on the number of forward simulations used to estimate the expectations in the Euler equation and in the conditional value functions. Increasing the number of simulations increases the precision at the cost of increased computational time. \\ Notice also that the CCCs and CCPs estimated for all $t, \Eta_t, X_t, W_t, m$, affect the estimation, and not only the ones at the selected moments. This is because, in the forward simulations, the entire range of covariates and $\Eta$ can be drawn. So the first stage optimal choices need to be well estimated at any $\Eta_t$, even at the tails.

Estimator performance

I illustrate the estimator's performance with Monte Carlo simulations of the estimation of a parametric toy model of simultaneous labor and consumption choices.

Toy model

Period utility. Each period from age $t=1$ to $T$, individuals choose to work full time ($D_t=1$) or part time ($D_t=0$) and to consume ($C_t$). Their period-$t$ utilities are

align*[align* omitted — 77 chars of source]

where $u_{dt}$ is the `main utility' function which depends on the consumption choice, $m_{dt}$ is the `additive part of the utility', which does not depend on the consumption but depends on the instrument (lagged labor choice, $W_t=D_{t-1}$ here), and $\Epsilon_{dt}$ (taking values $\epsilon_{dt}$) are extreme-value type I additive idiosyncratic shocks impacting the preferences for full time work. Both parts of the utility depend on the individuals' time-invariant types $\mu$ taking values $m \in \{1, 2\}$, with $\text{Pr}(\mu=1) = \pi^1$. These types are unobserved by the econometrician, but known by the individuals. \\ We parametrize the main utility, $u(\cdot)$ as a CES utility function

align*[align* omitted — 111 chars of source]

where the $\sigma_d^m$ represents type $m$ and discrete choice $d$-specific risk aversion or intertemporal elasticity of substitution. We allow this main risk aversion to vary with the labor tenure decision. In this model, contrary to standard CES models, the marginal utilities of consumption are heterogenous (even at fixed covariates, $X$, and type, $m$) because they depend on idiosyncratic preference shocks, $\Eta_t$, taking values $\eta_t \in [0, 1]$. As in the general model, $m$ captures the permanent differences (types) between individuals, while $\Eta_t$ captures period-specific transitory idiosyncratic preference shocks. $m$ and $\Eta_t$ are iid for every individuals. Both $m$ and $\Eta_t$ are unobserved by the econometrician. \\ The additive part of the payoff is given by

align*[align* omitted — 139 chars of source]

such that $m_{1t}^m - m_{0t}^m = m_{1t}^m$ represents the additive utility gains (or cost) of choosing to work full time ($D_t=1$) compared to working part time ($D_t=0$). More precisely, we model this cost as a linear function of the unobserved idiosyncratic preference for consumption, $\Eta_t$, where $\gamma^m$ represent the intercept/constant gain of working full time for individuals with $\Eta_t = 0$, while $\alpha^m$ represent the slope of how this gain changes with $\Eta_t$. The parameters $\alpha^m$ captures additional complementarity/substitutability between consumption and labor, in addition to the ones implicitly present in the main utility $u(\cdot)$. Finally, $\omega^m$ represent the utility switching cost (if $\omega^m$ is negative) endured by previously part time workers $(w_t = d_{t-1} = 0$) who switch to full time ($d_t=1$). Again, all the parameters $\gamma^m, \alpha^m, \omega^m$ are type-specific. \\ The model is dynamic and the individuals discount their future utility with a factor $\beta$.\footnote{Instead, we could have specified a type-specific discount factor, $\beta^m$. It would also be identified and precisely estimated.} Thus, the main parameters entering the Euler equations and affecting the consumption choices are $\beta$ and the parameters entering $u_{dt}^m(\cdot)$, i.e., $\theta_C = (\sigma_0^1, \sigma_0^2, \sigma_1^1, \sigma_1^2)$, while the parameters $\theta_P = (\gamma^1, \gamma^2, \alpha^1, \alpha^2, \omega^1, \omega^2)$ only affect the labor supply choices. The parameters $\theta = (\beta, \theta_C, \theta_P)$ describe the primitives of the models and represent the main parameters we want to estimate. \\

Dynamics and covariates transition. The asset, $a_t$, evolves according to the budget constraint

align*[align* omitted — 65 chars of source]

where $y_t$ represents the full-time equivalent yearly income of individuals. Individuals who work full time ($d_t=1$) obtain $y_t$, while individuals who work part time ($d_t=0$) obtain half of it. The income $y_t$ is a random variable, taking only two values for simplicity: $y_L$ and $y_H$, for low and high income, respectively. The income transition is given by

align*[align* omitted — 149 chars of source]

and is directly estimated from the observed data on income. \\ Asset and income are the only two observable covariates, i.e., $X_t = (a_t, y_t)$.\footnote{One could easily complexify this model, by adding more individual characteristics, a more complex income process, type-dependent budget constraint, more realistic pension plans for the retirees... I choose to model only the key features of a standard life-cycle model, as it is sufficient to illustrate the performance of the estimator in terms of precision and computation time. } Even though the utility does not directly depend on asset and income, the optimal consumption and labor choices depend on these through the dynamics of the problem. \\

Retirement. At age $T+1$, the individuals retire for $T^{\text{retire}}$ periods, and then dies. During retirement, they only consume and can no longer work. For simplicity, we specify that every period they obtain the period utility of a part-time working individual with a median $\Eta_t = 0.5$ and without the additive shock $\Epsilon$. They obtain a pension set to $50\%$ of their last full-time equivalent income, $y_T$. There is no bequest motive. The retirement problem has a closed form solution, easily solved for any parameters. \\

Instrument validity. Since this model enter the more general Framework, the previous labor choice, $D_{t-1}$, can be used as an instrument $W_t$ for identification here. Indeed, $W_t=D_{t-1}$ is excluded conditional on $D_t, X_t$ and $m$, and it will also be relevant provided that the switching costs $\omega^m \neq 0$.

figure[figure omitted — 2,062 chars of source]

\newgeometry{top=0.5in, bottom=0.8in, left=1in, right=1in}

table[table omitted — 4,053 chars of source]

\restoregeometry

Monte Carlo simulation results

To assess the performance of the estimator developed in this paper -- denoted DDCC for dynamic discrete-continuous choices -- I run Monte Carlo simulations of the life-cycle model described previously. Each simulation simulates a panel of $N=10,000$ individuals observed for their entire life-cycle of $T=10$ periods. This approximately corresponds to the sample size of real surveys used to estimate life-cycle models (e.g., the PSID), which allows to assess the performance of the estimator in a realistic context. Figure (ref) shows the estimated optimal policies (CCCs and CCPs) in the first stage, for a given type $m$, covariates $X=x$ and instrument $W=w$. Table (ref) shows the corresponding estimates of the deep parameters of the model, $\theta$, using the first stage policies previously estimated. \\

Estimation performances. The DDCC estimates (orange in Figure (ref), column DDCC in Table (ref)) correspond to the estimator described in Section (ref).\footnote{Specification details. For the DDCC method , the CCCs, $c_{0t}^m(\eta, x)$ and $c_{1t}^m(\eta, x)$, are estimated via a weighted (by estimated type-$m$ probabilities) IVQR of $C$ on $D$ given $X$, instrumented by $W$ chernozhukovhansen2006, kaido2021decentralization. The continuous covariates (asset) enters linearly in the specification, while we run a separate IVQR on each subsample of income and age. We estimate the model at each $5\%$ percentiles, from $5\%$ to $95\%$. To extrapolate at the tails, we run a supplemental regression of a shape constrained additive model pya2015shape imposing monotonicity of $C$ with respect to a flexible spline of order $4$ in $\Eta$. From the estimated CCCs, we can recover $\hat{\Eta}$ for every observations. Then, the CCPs are estimated using a flexible weighted (by type-$m$ probabilities) logit regression of $D$ on $\hat{\Eta}$, $X$, and $W$. More precisely, we regress $D$ on a polynomial of order 3 in $\Eta$, on $W$, age and income dummies, and a polynomial of order 3 in assets, and also interactions of the polynomial in $\Eta$ with $W$. Then the primitives of the models are estimated as described in Section (ref) by taking these optimal policies as given in forward simulations of the models to approximate (i) the next period marginal utility expectation (Euler criterion) and (ii) the conditional value functions. } With a balanced panel of $10,000$ individuals observed over $10$ periods, the types are well estimated in the initial EM algorithm and the optimal policies are precisely estimated without bias around the truth (Figure (ref)). Then, the forward simulations approximate well the expected next period marginal utility (Euler criterion) and the conditional value functions (probability criterion), and as a consequence, all the parameters of the models are precisely estimated.\footnote{The probability parameters are more noisy because small changes in these parameters only induce small changes in the observable CCPs.} \\ The `known policies' column in Table (ref) corresponds to an estimation of the second stage using the true first stage policies as if they were known. By comparing the DDCC results to these, we can separate the variance in the estimates caused by the second-stage forward simulation' approximations from the variance due to errors in the first-stage policies that carry over to the second stage. One can see that a large part of the variance is driven by the (unbiased) estimation of the first stage, even though some variance remains purely from the simulation process in the second stage (especially for the probability parameters). Note that the precision of the second stage can be increased by increasing the number of forward simulations, at the cost of increasing computation time as well. \\

Counterfactual exogeneity/sequentiality assumption. The `Counterfactual' estimation (blue in Figure (ref)) corresponds to an estimation obtained by taking the standard empirical approach of (wrongly) assuming the timing that $D$ is decided before $\Eta$ is realized (sequential choices), or, equivalently, that $D$ is exogenous with respect to $C$. As visible in Figure (ref), this exogeneity assumption leads to severely biased CCCs and CCPs estimates. This is because, by assuming that $D$ is independent of $\Eta$, one assumes that the CCPs are flat in $\Eta$ (see Figure (ref)). Consequently, the CCCs, $c_{dt}^m(\eta, x)$, are wrongly estimated by the observed quantiles of $C$ in the subsample $D=d$ (given $X=x$ and type $m$), ignoring the fact that the distributions of $\Eta$ differs in the $D=1$ and $D=0$ subsamples. In other words, the main error with this counterfactual approach is that the CCCs are estimated via simple quantile regression (weighted by estimated type-$m$ probabilities), instead of using the proper (weighted) IV-quantile regression approach to correct for the endogeneity.\footnote{Specification details. The CCCs are estimated via simple weighted (by estimated type-$m$ probabilities) quantile regression of $C$ on $D$ on each subsamples of income and age, and including the asset as a linear covariate in the regression. For the CCP, since $D$ is assumed exogenous, one simply estimate $\textrm{Pr}(D=1 | W, X, m)$ in the data using a weighted logit of $D$ on $W$, age and income dummies, and a polynomial of order $3$ in assets. } \\ These wrong estimates of the first stage policies induce biased estimates of the structural parameters of the model. The bias is especially severe for parameters related to the probability criteria, but we also obtain wrong estimates of the risk aversion and discount factor. This is problematic and means that taking this wrong (but relatively standard) approach can lead to biased counterfactual policy analysis, and wrong policy recommendations. \\

Computational performances. As visible in Table (ref), the estimation with the DDCC estimator of Section (ref) is very fast: it takes about $3$ minutes to estimate the entire model, $143$ seconds to estimate the first stage optimal policies (including the type probabilities) and $45$ seconds for the second stage primitive parameters. Note that these results were obtained using R, a popular language among economists but rarely used for structural estimation due to its slower performance compared to compiled languages like C or Fortran. By enabling rapid estimation in widespread languages like R or Python, the DDCC estimator lowers barriers, making structural modeling accessible for a broader range of applied researchers. \\ To contextualize this performance, I tried to compare the DDCC estimator with the best alternative, i.e., state-of-the-art indirect inference estimation using the endogenous grid method carroll2006, iskhakov2017 to solve for the value functions and optimal policies. The problem is that the EGM estimation is orders of magnitude (at least 200 times) longer: the computation of a single value function of this life-cycle model takes about 8 minutes, and one needs to compute hundreds (if not thousands) of value functions to find the optimal parameters. As a consequence, one value function computation following iskhakov2017 is longer than the entire estimation with my method. Therefore, my method yields sizeable computational gains, even to the point where one can estimate models that would otherwise be considered intractable. This is because the two-step method, even though it introduces a fixed computational cost for the first stage optimal choices, drastically reduces the computational burden by avoiding the computation of value functions for each evaluated set of parameters in the second stage. \footnote{Another advantage of the DDCC estimator with respect to indirect inference with the EGM is that I do not solve numerically for the optimal choices. As a consequence, I do not run into optimization problems and I do not need to smooth potential kinks introduced by the joint discrete-continuous choices, contrary to iskhakov2017 for example. } The more complicated the model, the larger the computational gains.

Conclusion

This paper develops a general class of dynamic discrete-continuous choice models including a wide range of unobserved heterogeneity with transitory shocks and permanent unobserved types. I provide a constructive identification proof for this class of models. Given the identification, I provide a new estimation procedure yielding sizeable computational gains relative to existing alternatives for the estimation of dynamic models. The gains are so large that they should facilitate the practical use of complex dynamic discrete-continuous models in many fields (labor, housing, education, industrial organization, etc.) in the future. \\ This discrete-continuous choice single-agent framework also adapts to stationary infinite horizon (dynamic) games with private information and unobserved market types. The adaptation of the framework is straightforward and similar to how the dynamic discrete choice framework of am2011 adapts to dynamic discrete games. See Appendix (ref) for more details.

{

spacing{0} {0pt}

}

\pagenumbering{arabic}