EconBase
← Back to paper

Identifying Dynamic Discrete Choice Models with Hyperbolic Discounting

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

54,787 characters · 6 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identifying Dynamic Discrete Choice Models with Hyperbolic Discounting

titlepage\begin{abstract} We study identification of dynamic discrete choice models with hyperbolic discounting. We show that the standard discount factor, present bias factor, and instantaneous utility functions for the sophisticated agent are point-identified from observed conditional choice probabilities and transition probabilities in a finite horizon model. The main idea to achieve identification is to exploit variation in the observed conditional choice probabilities over time. We present the estimation method and demonstrate a good performance of the estimator by simulation. \\ \\ \noindentKeywords: Identification, dynamic discrete choice, discount factor, present bias, hyperbolic discounting.\\ \\ \end{abstract} \setcounter{page}{1} \thispagestyle{empty}

\onehalfspacing

Introduction

Dynamic discrete choice (DDC) models have been widely used in applied microeconomics, such as industrial organization, health economics, labor economics, and political economy. The DDC model is an econometric model that has a close link to economic theory, and hence, its main advantage is that the researcher can infer the mechanism of an agent's decision-making from data. Using DDC models, the agent's preferences have been estimated in a broad range of social problems. For example, keane1997career analyze career decision-making, kennan2011effect develop a DDC model on optimal migration, bayer2016dynamic model neighborhood choice, and de2019subsidies infer the agent's incentive to introduce renewable energy technologies.

Many DDC models build on dynamic optimization with exponential discounting, in which an agent discounts future streams of utility exponentially, so that the intertemporal marginal rate of substitution is constant over time. Exponential discounting has been extensively used because it is simple to estimate the model and easy to interpret due to time-consistent preferences.

In real life, however, many cases of decision making do not follow such exponential discounting. strotz1955myopia and thaler1981some provide an example: Most people tend to choose to wait for two apples in one year and a day, rather than one apple in one year, but some people choose to get one apple today rather than two apples tomorrow. This evidence clearly contradicts exponential discounting. thaler1981some provides laboratory results that support this fact. In addition, \citet*{frederick2002time} provide a review of the evidence for hyperbolic discounting, in which the agent has time-inconsistent preferences.

More importantly, in addition to experimental evidence, hyperbolic discounting is motivated in terms of policy implications. This is because the estimation under different types of discounting can potentially yield quite different results. For example, kennan2011effect analyze migration patterns in the United States using a DDC model under exponential discounting. They find that the variables in the moving costs, such as distance, home location, and population have significant effects on moving decisions. However, if we re-estimate the model using hyperbolic discounting, the estimation result can potentially be altered. This change may significantly impact policy decisions, such as subsidies for moving costs. A more detailed explanation can be found in Examples (ref) and (ref) in Section (ref).

To the best of our knowledge, hyperbolic discounting is first modeled by strotz1955myopia. phelps1968second and laibson1997golden later remodel time-inconsistent preferences using a quasi-hyperbolic discounting model, in which the present preferences are represented by the following utility function:

align[align omitted — 131 chars of source]

where each $u_{\tau}$ represents instantaneous utility in period $\tau$, the constant $\delta$ is called the standard discount factor which discounts future streams of utility exponentially, and the constant $\beta$ is called the present bias factor which captures the agent's myopia. The product $\beta \delta^{\tau-t}$ in period $\tau$ captures the degree of discounting future utility in the hyperbolic sense. This quasi-hyperbolic discounting model includes exponential discounting as a special case.\footnote{Throughout, we refer to quasi-hyperbolic discounting simply as hyperbolic discounting for simplicity.}

Little is known about identification results on the primitives of DDC models with hyperbolic discounting. Although some empirical papers (such as fang2009time and paserman2008job) try to estimate DDC models with hyperbolic discounting, they rely on parametric assumptions. Such approaches are sensitive to the functional form and it has been unknown whether DDC models with hyperbolic discounting are identified without parametric assumptions.

This paper models the agent's decision making based on quasi-hyperbolic discounting, following fang2015estimating. We show that the present bias factor, the standard discount factor, and instantaneous utility functions are point-identified from the observed conditional choice probabilities (CCPs) and transition probabilities in the finite-horizon DDC model.

In the finite horizon framework, we achieve identification of the present bias factor and the standard discount factor by exploiting variation in the observed CCPs over time. In the context of identification, because the observed CCPs and the state transition rules are observable from the data, we obtain a closed-form solution of the discount functions (the present bias factor and the standard discount factor) using these observables. The strategy for identification is as follows. First, we get a recursive expression of an agent's perceived long-run value function. Second, in order to solve it, we take the logarithm of the CCP ratio and use the so-called Hotz-Miller inversion. Third, we find pairs of actions that yield the same level of instantaneous utility and impose stationary structures in instantaneous utility functions. Then, all of the unknown variables are expressed in the form of functions of known variables, and hence, the model is identified.

We propose a maximum likelihood estimator to estimate the parameters of the model. Based on the identification results and the proposed estimation methods, we conduct Monte Carlo experiments and demonstrate a good performance of our estimator.

This paper contributes to the literature on the identification of the DDC models. rust1994structural shows the non-identifiability of primitives in the infinite horizon setting. magnac2002identifying show the degree of under-identification. In relation to this paper, they further claim that the standard discount factor is identified using exclusion restriction in the exponential discounting model. kasahara2009nonparametric and hu2012nonparametric provide the nonparametric identification results, taking unobserved heterogeneity into account. arcidiacono2020identifying present the identification results focusing on the short panel setting like our models. Using the exponential discounting model, abbring2020identifying show set-identification results of the standard discount factor. \citet*{an2021dynamic} ease the rational expectation assumption in DDC models and identify the agent's subjective beliefs on state transitions. They study DDC models with exponential discounting and assume that discount functions are known. On the other hand, we identify DDC models including discount functions with hyperbolic discounting, while maintaining the rational expectation assumption. Last but not least, abbring2010identification review the identification results of the DDC models.

This paper is closely related to fang2015estimating, \citet*{abbring2018identifying}, and \citet*{wang2022identification} (WWX hereafter). These papers show identification results in the hyperbolic discounting setting. To the best of our knowledge, fang2015estimating are the first who analyze the identification of the DDC model with hyperbolic discounting. Although their identification results are refuted by abbring2020comment, fang2015estimating provide the basic framework that is useful for our model. \citet*{abbring2018identifying} also consider the identification of DDC models with hyperbolic discounting, but they provide the set identification results. In contrast, we provide the point identification results.

WWX also provide complementary results to ours: they present point identification results for both naive agents and sophisticated agents.\footnote{Agents are said to be sophisticated if they are fully aware of their present bias and naive if not.} Notably, for sophisticated agents, they show that the discounting functions are identified if the final three periods are observed. However, their results crucially depend on the existence of a terminating action. On the other hand, our identification argument does not rely on terminating actions. As a result, our framework can incorporate a wider range of DDC models, including decision processes without terminating actions. For example, the dynamic analysis of the migration decision does not have a terminating action, as it is a life-long decision-making (see Example (ref) in Section (ref)). Although our identification result has an advantage in that it does not require terminating actions, our result requires a larger number of time periods in data. In our framework, the number of time periods required in the data increases linearly as the number of the support of the state variable increases. In contrast, WWX only require the final three periods to be observed. Therefore, there is a trade-off between our result and WWX for identification in terms of the need for terminating actions and the required number of time periods.

The identification approach of our paper and WWX are similar: they both use a variation in CCPs over time and obtain the identification of discount functions by matrix inversion. As a result, they both assume the full rank condition for the matrices related to the transition matrix and CCPs but in different forms. Broadly speaking, this difference in ways of constructing matrices results in the difference in required assumptions.

The structure of this paper is as follows. Section (ref) presents the DDC model with hyperbolic discounting. Section (ref) provides the main identification result. Section (ref) presents the estimation method, and section (ref) shows the results of the simulation. Section (ref) concludes.

Model

In this section, we consider a finite horizon dynamic discrete choice model where an agent has time-inconsistent preferences. Our model is closely related to that of fang2015estimating, but we assume that the time horizon is finite. Our notation broadly follows that of fang2015estimating and \citet*{an2021dynamic}.

In each discrete time period $t=1,2,\ldots,\bar{T}$ ($\bar{T} < \infty$), an agent chooses a discrete action from the set $\mathcal{I} \coloneqq \left\{ 1,2,\ldots,K \right\}$ ($K<\infty$). The agent's instantaneous utility depends on the action they choose and the set of state variables, where $x \in \mathcal{X} \coloneqq \left\{ 1,2,\ldots,J \right\}$ ($J<\infty$) is observed by the econometrician, and $\boldsymbol{\varepsilon} \coloneqq \left( \varepsilon_{1}, \ldots, \varepsilon_{K} \right) \in \mathbb{R}^{K}$ is a vector of choice-specific shocks that are not observed by the econometrician. The agent can observe both the current state variables $x$ and $\boldsymbol{\varepsilon}$, and the agent chooses a current action so as to maximize the expected lifetime utility. Given the current state variables $\left( x, \boldsymbol{\varepsilon} \right)$ and the agent's choice $i$, the state variables for the next period $(x', \boldsymbol{\varepsilon}')$ are realized by the exogenous transition function $f \left( x', \boldsymbol{\varepsilon}' \mid x, \boldsymbol{\varepsilon}, i \right)$. We henceforth use a prime on a state variable when it denotes the state in the next period. Regarding the state transition rule, we pose the following assumption.

asmThe observed and unobserved state variables evolve independently, conditional on $x$ and $i$: \begin{align} f \left( x', \boldsymbol{\varepsilon}' \mid x, \boldsymbol{\varepsilon}, i \right) &= g \left( \boldsymbol{\varepsilon}' \mid x' \right) f \left( x' \mid x, i \right) \nonumber \\ g \left( \boldsymbol{\varepsilon}' \mid x' \right) &= g \left( \boldsymbol{\varepsilon}' \right) \nonumber \end{align}

In Assumption (ref), we assume that the transition rule governs the evolution of the state variables in way of a Markov process given the agent's action. We assume that the transition function is time-invariant, which is common in the literature.

In the agent's dynamic optimization model, we implicitly assume that the agent has rational expectations or perfect expectations, that is, each agent realizes the true transition rule of both observed and unobserved state variables. This assumption is standard in the literature of dynamic discrete choice models (e.g., magnac2002identifying). An important exception is the model of \citet*{an2021dynamic}; they consider the situation in which the agent has subjective beliefs about the law of motion of the observed state variables. Although their identification result provides a new insight into the literature, we stick to the assumption of the agent's perfect expectations.

We now make an assumption on the agent's instantaneous utility.

asmThe instantaneous utility is time-invariant and is given by \begin{align} u_{i}^{*} \left( x, \boldsymbol{\varepsilon} \right) = u_{i} \left( x \right) + \varepsilon_{i}, \nonumber \end{align} for each $i \in \mathcal{I}$.

The additive separability assumption on the agent's instantaneous utility has been widely used in the literature (e.g., rust1987optimal, hotz1993conditional, and rust1994structural). The stationarity assumption in the finite horizon setting can be seen in the model of \citet*{an2021dynamic}.

There are two types of discount functions of the agent. The first one is the standard discount factor $\delta \in (0,1)$, which denotes the agent's long-run discounting in every period. This discounting function exponentially discounts future streams of utility, capturing the time-consistent part of the agent's preferences. The second one is the present-bias factor $\beta \in (0,1)$, which denotes the agent's short-term impatience: it further discounts all the future utility from tomorrow on. Following phelps1968second and laibson1997golden, we represent the agent's intertemporal preferences by

align[align omitted — 162 chars of source]

where $u_{\tau}$ is the agent's instantaneous utility in period $\tau$. According to laibson1997golden, this model is called quasi-hyperbolic discounting, because it approximates the hyperbolic discounting model.

Because the agent has time-inconsistent preferences, we consider the model as the game played by the current agent herself and the future selves. Following fang2015estimating, let $\sigma_{t}: \mathcal{X} \times \mathbb{R}^{K} \to \mathcal{I}$ be a Markovian choice strategy and $\boldsymbol{\sigma}_{t}^{+} \coloneqq \left\{ \sigma_{k} \right\}_{k=t}^{\bar{T}}$ the continuation strategy profile subsequent to period $t$. Given $\boldsymbol{\sigma}_{t}^{+}$, the agent's value function at period $t$ when the realized state variables are $x$ and $\boldsymbol{\varepsilon}$ is recursively written as

align[align omitted — 472 chars of source]

Following o1999doing, o2001choice, and fang2015estimating, we define the perception-perfect strategy profile $\boldsymbol{\sigma}^{*} \coloneqq \left\{ \sigma_{t}^{*} \right\}_{t=1}^{\bar{T}}$ such that

align[align omitted — 347 chars of source]

for all $t$, $x$, and $\boldsymbol{\varepsilon}$ (where $\boldsymbol{\sigma}_{t}^{*+} \coloneqq \left\{ \sigma_{k}^{*} \right\}_{k=t}^{\bar{T}}$). Using (ref) and (ref), we define the agent's perceived long-run value function as

align[align omitted — 218 chars of source]

for all $t$ and $x$.

With the perceived long-run value function, we now define two types of choice-specific value functions to analyze how the agent makes a decision. Following fang2015estimating, the perceived choice-specific long-run value function is defined as

align[align omitted — 192 chars of source]

and the agent's current choice-specific value function is defined as

align[align omitted — 198 chars of source]

The two types of choice-specific value functions differ in two ways. First, in $V_{t,i}\left( x \right)$, the value $s$ periods ahead from $t$ is discounted by $\delta^{s}$, while in $W_{t,i}\left( x \right)$, it is discounted by $\beta\delta^{s}$. Second, the perceived choice-specific long-run value function, $V_{t,i}\left( x \right)$, captures the way of agent's discounting as if the agent ignores their own present-bias (i.e., $\beta=1$). That is, the agent is not aware of their own myopia in discounting. In contrast, the current choice-specific value function, $W_{t,i}\left( x \right)$, correctly captures the agent's discounting behavior including their own present bias.

We impose assumptions on the distribution of unobserved state variables and the conditional choice probability (CCP) $P_{t,i}\left( x \right)$, which is the probability that action $i$ is taken at period $t$ when the state variable $x$ is realized.

asm(a) The unobserved state variable $\boldsymbol{\varepsilon}$ is i.i.d. type 1 extreme value distributed. \\ (b) The CCP $P_{t,i}\left( x \right)$ is known and determined by the current choice-specific value function $W_{t,i}\left( x \right)$.

The extreme value distribution assumption in Assumption (ref) (a) is standard in the literature of dynamic discrete choice models (e.g., Assumption CLOGIT in aguirregabiria2010dynamic). Although there are some other important exceptions on the shock distribution such as the generalized extreme value distribution (e.g., arcidiacono2011conditional) and the unknown distribution (e.g., norets2014semiparametric), we maintain the assumption of type 1 extreme value distribution so as to get analytical convenience. In Assumption (ref) (b), we consider the situation where $N\to\infty$ where $N$ is the number of agents in the data. This makes our identification arguments valid.

Under Assumption (ref) (a) and (b), the (observed) CCP $P_{t,i}\left( x \right)$ is given by

align[align omitted — 390 chars of source]

Before we present the identification result, we consider some empirical examples.

exampleThe choice of housing location has been widely studied. kennan2011effect analyze migration patterns in the United States, using a dynamic discrete choice model. In their model, each agent not only chooses whether or not to stay in their current location, but also chooses where to live if the agent decides to move. The location choices and state variables such as wages, moving costs, nonpecuniary amenity values, and a home premium, affect the agent's instantaneous payoff. kennan2011effect set up a finite-horizon stationary model with exponential discounting. There are no terminating actions, so location choices are not once and for all. kennan2011effect find that the variables in the moving costs, such as distance, home location, and population have significant effects on moving decisions. Their results are based on estimation under exponential discounting, assuming that the standard discount factor is known. The use of hyperbolic discounting and the estimation of discount functions have meaningful implications for policy analysis. This is because the estimation results would be different if we estimate the model under hyperbolic discounting with unknown discount factors. Consequently, the new model and its estimation results could potentially alter policy decisions on subsidies for moving costs, for example.
examplede2019subsidies analyze a subsidy program designed to promote the adoption of solar photovoltaic (PV) systems. The subsidies are provided for future electricity production, not for upfront investment. They develop a DDC model and identify the standard discount factor. In their model, each household makes a discrete choice in each period: which PV alternative to choose or not to adopt. This choice influences both instantaneous utility and future states. State variables include the upfront investment cost of the PV system, electricity cost savings by adoption, and subsidy levels. In this example, adopting one brand of PV systems is a terminating action. de2019subsidies find that households significantly discount future benefits under exponential discounting, which indicates the potential benefit of introducing hyperbolic discounting. This issue is of crucial importance in policy analysis. Under exponential discounting, they find that an upfront investment subsidy program would have been more cost-effective than subsidies for future electricity production. The results of the policy analysis would change significantly if hyperbolic discounting is introduced, thereby affecting the design of the subsidy program by policy makers.

Identification

In this section, we show that, in finite horizon framework, discount functions $\beta$ and $\delta$ are identified using time variation in observed CCPs. Our identification argument is close to that of \citet*{an2021dynamic}. They study DDC models with exponential discounting and relax the rational expectation assumption to identify the agent's subjective beliefs on state transitions, while assuming that discount functions are known. On the other hand, we identify DDC models including discount functions with hyperbolic discounting, while maintaining the rational expectation assumption.

To provide a brief overview of our identification strategy, note that there are three types of value functions, $V_{t}\left( x \right)$, $V_{t,i}\left( x \right)$, and $W_{t,i}\left( x \right)$, as introduced in Section (ref). We exploit the relationships between these value functions and their connection to the observed CCPs, $P_{t,i}\left( x \right)$. By taking time differences and appropriately taking differences from a reference state, we can separate the discounting functions from the observables. Under the assumptions of full rank matrices and the existence of pairs of actions or states that yield the same utility level, both $\beta$ and $\delta$ are individually derived in closed form.

Let $T$ be the last period of the data ($T \le \bar{T}$). First, note that, by (ref) and (ref), we obtain

align[align omitted — 294 chars of source]

which will be useful for identification. By combining equations (ref) and (ref), we have

align[align omitted — 215 chars of source]

Exploiting this relationship as well as equation (ref) and Assumption (ref), we have

align[align omitted — 1,594 chars of source]

By equation (ref), this equation can be rewritten as

align[align omitted — 431 chars of source]

If we take a difference against the reference state $J$, we obtain

align[align omitted — 870 chars of source]

In order to facilitate subsequent arguments, we introduce some notations. Let

align[align omitted — 238 chars of source]

be a $(J-1) \times 1$ vector and define $\log \left( \mathbf{P}_{t,K} \right)$ and $\mathbf{u}_{K}$ similarly. Let

align[align omitted — 280 chars of source]

be a $(J-1) \times (J-1)$ matrix, where $\mathbf{F}_{i} \left( x \right) = \left[ f \left( x'=1 \mid x,i \right), \ldots, f \left( x'=J-1 \mid x,i \right) \right]$ is a $1 \times (J-1)$ vector. Let

align[align omitted — 351 chars of source]

be $(J-1)K \times (J-1)$ matrices, where

align[align omitted — 427 chars of source]

are $(J-1) \times (J-1)$ matrices. We also let

align[align omitted — 166 chars of source]

be a $(J-1) \times (J-1)K$ matrix where

align[align omitted — 326 chars of source]

and let

align[align omitted — 168 chars of source]

be a $(J-1) \times (J-1)K$ matrix where

align[align omitted — 320 chars of source]

Using the above matrices, the right-hand side of equation (ref) can be rewritten as

align[align omitted — 349 chars of source]

Now we consider the CCP ratio to acquire another form of the first difference of the perceived long-run value function vector. For actions $k,\ell \in \mathcal{I}$ and states $x_{1}, x_{2} \in \mathcal{X}$, we define

align[align omitted — 729 chars of source]

where the second equality comes from the result of Proposition 1 in hotz1993conditional and the fourth equality follows from the derivation in Appendix.

In equation (ref), suppose that we have $(J-1)$ pairs of actions $\left( k,\ell \right)$ and states $\left( x_{1}, x_{2} \right)$ and let each pair be $\left( k^{j}, \ell^{j} \right)$ and $\left( x_{1}^{j}, x_{2}^{j} \right)$ for $j = 1,\ldots,J-1$. Then we define the following $\left( J-1 \right) \times \left( J-1 \right)$ matrix:

align[align omitted — 348 chars of source]

In order to get a closed form of the ex-ante value function, we impose the following assumption:

asm(a) There exist at least $J-1$ action pairs $\left( k^{j}, \ell^{j} \right)$ and state pairs $\left( x_{1}^{j}, x_{2}^{j} \right)$ such that: \begin{align} u_{k^{j}} \left( x_{1}^{j} \right) = u_{\ell^{j}} \left( x_{2}^{j} \right), \end{align} with either (i) $k^{j} \ne \ell^{j}$, (ii) $x_{1}^{j} \ne x_{2}^{j}$, or (iii) both, for $j = 1, \ldots, J-1$. \\ (b) the matrix $\tilde{\mathbf{F}}$ is full column rank.

In part (a) of Assumption (ref), the researcher must find $J-1$ pairs of different actions or different states that give the exact same level of instantaneous utility. There are two types of justification for this assumption. First, this assumption includes normalization, which is commonly assumed in the DDC literature. For example, normalization ($u_{K} \left( x \right) = 0$ for all $x$) is imposed in fang2015estimating and abbring2020identifying for DDC models; bajari2015identification and aguirregabiria2020identification for dynamic games. \citet*{an2021dynamic} also assume it for the infinite-horizon two-state DDC model. The reason why Assumption (ref) (a) is more general than normalization is as follows. Our assumption only requires that there exist $J-1$ pairs of states and actions that give the same utility level, and such pairs can be even across states and across actions. On the other hand, normalization requires that all such pairs must be on the reference action $K$. Second, this assumption is easily imposed in the real empirical analysis, because it is directly imposed on the instantaneous utility. abbring2020identifying also impose the similar assumption to ours, in the context of exponential discounting. As they mention in their paper, this type of assumption has superiority to that in magnac2002identifying. This is because magnac2002identifying impose an assumption on the current value function, which is often difficult to get an economic intuition.

The intuition for part (b) of Assumption (ref) is as follows: when $J=2$, this condition reduces to $f \left( x'=1 \mid x_{1}, k \right) \ne f \left( x'=1 \mid x_{2}, \ell \right)$ for some $k, \ell \in \mathcal{I}$ and some $x_{1}, x_{2} \in \mathcal{X}$. Furthermore, when $x_{1} = x_{2}$, this is simplified to $f \left( x'=1 \mid x=1, k \right) \ne f \left( x'=1 \mid x=1, \ell \right)$. This means that, starting from the same state, the transition probability to the identical state must be different for the different action. When $J \ge 3$, this assumption requires that every row vector of $\tilde{\mathbf{F}}$ be linearly independent, which is easily testable. In the literature of identification of DDC models, \citet*{an2021dynamic} introduce a similar assumption for the subjective beliefs for the state transition.

Assumption (ref) makes it possible to express $\mathbf{V}_{t+1}$ in a closed form. When the researcher can find more than $J-1$ action and state pairs, Assumption (ref) (b) ensures that $\tilde{\mathbf{F}}$ has an inverse matrix and the subsequent argument is also valid. Henceforth, we focus on the situation where we have exactly $J-1$ pairs in Assumption (ref) (a) and $\tilde{\mathbf{F}}$ has an inverse matrix $\tilde{\mathbf{F}}^{-1}$. Assumption (ref) and equation (ref) enables us to obtain the following equations:

align[align omitted — 299 chars of source]

where $\mathbf{D}_{t} = \left[ D_{t,k^{1},\ell^{1}} \left( x_{1}^{1}, x_{2}^{1} \right), \ldots, D_{t,k^{J-1},\ell^{J-1}} \left( x_{1}^{J-1}, x_{2}^{J-1} \right) \right]'$ is a $(J-1) \times 1$ vector. Taking the first difference of equation (ref), applying equations (ref) and (ref), and noting that the instantaneous utility is time-invariant by Assumption (ref), we have that, for $t=3,\ldots,T$,

align[align omitted — 594 chars of source]

The right-hand side of equation (ref) can be equivalently expressed as

align[align omitted — 296 chars of source]

where

align[align omitted — 441 chars of source]

is a $3(J-1) \times (T-2)$ matrix where

align[align omitted — 287 chars of source]

and $\Delta \log \left( \mathbf{P}_{K} \right) = \left[ \Delta \log \left( \mathbf{P}_{3,K} \right), \ldots, \Delta \log \left( \mathbf{P}_{T,K} \right) \right]$ is a $(J-1) \times (T-2)$ matrix.

For identification, we make the following assumption:

asm(a) The number of time periods in the data, $T$, satisfies $T \ge 3J-1$. \\ (b) The matrix $\mathbf{A}$ is full row rank.

Assumption (ref) requires the matrix $\mathbf{A}$ to have the right inverse matrix. As a result, both $\frac{1-\beta}{\beta} \mathbf{I}_{J-1}$ and $-\left( \beta \delta \right)^{-1} \mathbf{I}_{J-1}$ in equation (ref) have a closed-form solution and thus identification is made possible. Since both the state transition functions and the observed CCPs are known, Assumption (ref) is empirically testable. Remark (ref) below (at the end of this section) demonstrates how to interpret the matrix $\mathbf{A}$ and the full row rank condition (ref) (b).

If we impose Assumptions (ref) to (ref), we have the following theorem on identification:

thmSuppose that Assumptions (ref) to (ref) hold. Then the discount functions $\beta$ and $\delta$ are identified for $t=1,2,\ldots,T$ with $T \ge 3J-1$. Furthermore, if the utility level of one action for one state is known, then the instantaneous utility functions $u_{i} \left( x \right)$ for all $i \in \mathcal{I}$ and all $x \in \mathcal{X}$ are identified.

The proof for Theorem (ref) is in Appendix. If we impose some degree of stationarity in the finite horizon framework, the present-bias factor and the standard discount factor are identified as a function of the observed CCPs and the transition density. Also, if we set the continuation value at the terminal value to be zero, due to the stationary structure of the instantaneous utility, each instantaneous utility function is identified. Note that our identification strategy is different from the one proposed by fang2015estimating or the one discussed by abbring2020comment.

remWe illustrate the interpretation of matrix $\mathbf{A}$ and Assumption (ref) using a simplified example with $K=2$ (binary action), $J=2$ (two states), and $T=5$ (five periods, which satisfies Assumption (ref) (a)). Let $\pi_{t}\left(x,x^{\prime}\right)$ be the transition probability from state $x$ at period $t$ to state $x^{\prime}$ at period $t+1$. Focusing on the first column of $\mathbf{A}$, some algebra yields: \begin{align} \mathbf{A} = \beta\delta \left[ \begin{array}{ccc} \left[f(1|1,2)-f(1|2,2)\right]\left[V_{4}(1)-V_{4}(2)-V_{3}(1)+V_{3}(2)\right] & * & * \\ \left[\pi_{3}\left(1,1\right) - \pi_{3}\left(2,2\right)\right]\left[V_{4}(1)-V_{4}(2)\right] - \left[\pi_{2}\left(1,1\right) - \pi_{2}\left(2,2\right)\right] \left[ V_{3}(1)-V_{3}(2) \right] & * & * \\ V_{3}(1)-V_{3}(2)-V_{2}(1)+V_{2}(2) & * & * \end{array} \right], \nonumber \end{align} where $*$ represents non-zero entries, defined analogously for $t=4,5$. Assume that $f(1|1,2)-f(1|2,2) \ne 0$ (corresponding to Assumption (ref)(b) in this context). Then the full row rank condition (ref) (b) requires sufficient variation in the value functions $V_{t}(x)$ over time. To see the opposite case, consider the situation where there is no variation in value functions. If $V_{t}(x) = V_{t+1}(x)\ (t=2,3,4)$ for example, then the last row of $\mathbf{A}$ would be a zero row vector. Consequently, matrix $\mathbf{A}$ would not have an inverse, and thus our identification argument would fail. Therefore, Assumption (ref) effectively rules out cases where value functions remain constant over time, ensuring the necessary variation for identification.

Estimation

We have presented the identification results so far and we can directly apply such identification results to estimation. However, as can be seen in Theorem (ref) and its proof, we have to use the right inverse of a matrix, and such a method can be difficult to use when a matrix is near singular. Instead, we present a maximum likelihood estimator to estimate the standard discount factor and the present bias factor.

Let the data have $N$ agents and $T$ periods. Let the data be $\left\{ a_{nt}, x_{nt} \right\}$, where $a_{nt}$ is an action that the agent $n$ chooses at period $t$ and $x_{nt}$ is the corresponding state.

The likelihood function of the data is

align[align omitted — 437 chars of source]

where $\theta_{u}$ is the parameter in the instantaneous utility, $\theta_{f}$ is the parameter in the state transition, and $P_{t} \left( a_{nt} \mid x_{nt} ; \theta_{u}, \theta_{f}, \beta, \delta \right)$ is the CCP of agent $n$ at period $t$. The log-likelihood function of the data is

align[align omitted — 406 chars of source]

so that we can estimate $\theta_{f}$ separately from the other unknown parameters $\theta_{u}$, $\beta$, and $\delta$.

First, we can obtain an estimator $\hat{\theta}_{f}$ of $\theta_{f}$ by maximizing the following objective function:

align[align omitted — 129 chars of source]

Second, after obtaining $\hat{\theta}_{f}$, we estimate agents' CCPs by backward induction. At the terminal period $T$, let the continuation value be zero so that the choice-specific value equals the instantaneous utility. Then, we can date back and obtain the CCPs for all periods. Using the estimated CCPs and $\hat{\theta}_{f}$, we can estimate the unknown parameters $\theta_{u}$, $\beta$, and $\delta$ by maximizing the following function:

align[align omitted — 152 chars of source]

Simulation

We present the results of Monte Carlo experiments based on the identification results and the estimation methods that we have described. We consider a model with one state variable that can take five possible values ($\mathcal{X} = \left\{ 0,1,2,3,4 \right\}$) and with two choices ($\mathcal{I} = \left\{ 1,2 \right\}$) in the finite horizon model ($T=16$, which satisfies Assumption (ref) (a)). We parameterize the instantaneous utility functions as $u_{1} \left( x \right) = \alpha_{0} + \alpha_{1} x$ for all $x \in \mathcal{X}$ and assume that $u_{2} \left( x \right) = 0$ for all $x \in \mathcal{X}$, maintaining Assumption (ref) (a). We set $\alpha_{0} = 0.5$ and $\alpha_{1} = -0.2$. Also, each choice-specific shock $\varepsilon_{i}$ is independently drawn from type 1 extreme value distribution with mean zero. Moreover, each state transition matrix is set so that each element is randomly generated from the uniform distribution over $\left[ 0,1 \right]$ and each row sums to one.

We generate data by the following procedure. First, we set the instantaneous utility functions and the transition probabilities as described above, and the discount functions as below. Second, by backward induction, we calculate the ex-ante value function $V_{t+1} \left( x \right)$, the choice-specific value $W_{t,i} \left( x \right)$, and the conditional choice probability $P_{t,i} \left( x \right)$ for each $x,i$, and $t=1,\ldots,T$. Third, we simulate each agent's action using the transition probabilities and conditional choice probabilities. Given the generated data, we estimate $\alpha_{0}$, $\alpha_{1}$, $\delta$, and $\beta$ by jointly maximizing the likelihood function.

As for the discount functions, we set two sets of the values:

align[align omitted — 183 chars of source]
table[table omitted — 866 chars of source]
table[table omitted — 866 chars of source]

The results of simulation are presented in Tables (ref) and (ref). In each setting of simulation, we use the sample with $N = 2000$, 4000, 8000, and 10000, and standard errors are calculated from 10000 replications. Prior to each estimation, we have confirmed that Assumptions (ref) and (ref) are satisfied so that the matrices we use are full rank. We use nine patterns of initial values of the parameters as follows: initial values of $\alpha_{0}$ and $\alpha_{1}$ are 0.95 times of its each true value and initial values of $\delta$ and $\beta$ are drawn from the set $\left\{ 0.7, 0.8, 0.9 \right\}$.

Our results in Tables (ref) and (ref) suggest that our proposed estimation method does a good job across different set of sample sizes. There are several things to notice. First, the parameters of instantaneous utility functions are precisely estimated with small standard errors. The estimates are stable across different sample sizes. Second, the estimates of discount functions approach the true values as the sample size increases. The standard errors of the discount functions are relatively larger than those of instantaneous utility functions.

Conclusion

We study identification of dynamic discrete choice models with hyperbolic discounting. We show that the standard discount factor, present bias factor, and instantaneous utility functions for the sophisticated agent are point-identified from observed CCPs and transition probabilities in a finite horizon model. The main idea to achieve identification is to exploit variation of the observed CCPs over time. We also propose the estimation method and demonstrate a good performance of our estimator by simulation.

\singlespacing \setlength\bibsep{0pt}

\onehalfspacing