Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
69,615 characters · 14 sections · 66 citation commands
Identifying Present-Biased Discount Functions in Dynamic Discrete Choice Models
\abstract{ We study the identification of dynamic discrete choice models with sophisticated, quasi-hyperbolic time preferences under exclusion restrictions. We consider both standard finite horizon problems and empirically useful infinite horizon ones, which we prove to always have solutions. We reduce identification to finding the present-bias and standard discount factors that solve a system of polynomial equations with coefficients determined by the data and use this to bound the cardinality of the identified set. The discount factors are usually identified, but hard to precisely estimate, because exclusion restrictions do not capture the defining feature of present bias, preference reversals, well. }
\thispagestyle{empty}
The standard experimental approach to measure dynamically-consistent time preferences identifies them from choice responses to variation in future values, holding current payoffs fixed jel02:fredericketal,urminskyzauberman15. Observational studies using dynamic discrete choice models with dynamically-consistent preferences have invoked a similar intuition yaoetal12,Lee2013,chungetal14, bollinger15,aer19:degrooteverboven. qe20:abbringdaljord formalized this intuition by showing that, in such models, exclusion restrictions on utility set identify the geometric discount factor and, nonparametrically, the utility function. We extend qe20:abbringdaljord's approach to dynamic discrete choice models with sophisticated quasi-hyperbolic discounting, which is represented by a {\em present-bias factor} $\beta$ and a {\em standard discount factor} $\delta$. ier15:fangwang introduced such a model, with more general partially-naive quasi-hyperbolic discounting, and used it, under exclusion restrictions on utility, to study present bias in mammography decisions. selfcontrolchan17 and ucb19:mahajanetal used similar models and exclusion restrictions to analyze dynamically-inconsistent time preferences in welfare benefit choices and the demand for insecticide-treated nets.
We first study a finite-horizon model, as in much of the literature ucb19:mahajanetal. For this model, a standard backward recursion argument establishes that a unique solution (up to the resolution of ties) exists. We then extend our analysis to stationary, infinite horizon models, as in ier15:fangwang, which are useful in applications that have no clear decision horizon. We use the specific econometric structure on our infinite-horizon problem to prove that it always admits a pure-strategy solution. We show that the identification analysis of both models can be reduced to finding the discount factors $(\beta,\delta)$ that solve a system of polynomial equations. This system has one economically-interpretable equation for each exclusion restriction, with coefficients determined by the choice and transition probabilities that can be directly estimated from the data. We first focus on the case in which we have two exclusion restrictions, which give two equations to identify $\beta$ and $\delta$. We show that, in this bivariate case, the number of discount factors $(\beta,\delta)$ in the identified set is finite and bounded above by known features of the data, notably the time horizon and the number of states. In turn, each pair $(\beta,\delta)$ of discount factors in the identified set corresponds to a unique (nonparametric) utility function that rationalizes the data.
Our approach leverages the assumption that agents are sophisticated and rationally foresee that they will be present biased in the future. It does not readily extend to the (partially) naive case. ucb19:mahajanetal showed point identification for partially-naive time preferences in a three-period model under a particular set of exclusion restrictions. daljordetal19 established point identification of a fully nonparametric, time-separable discount function in a terminal action problem. We demonstrate that, for a more general set of exclusion restrictions than ucb19:mahajanetal's and more general choice structures and time horizons, but sophisticated time preferences, the number of discount functions in the identified set depends on the number of choice periods (with finite horizons) or states (with infinite horizons).\footnote{ier15:fangwang proposed a proof of identification of partially-naive quasi-hyperbolic time preferences under similar exclusion restrictions to the ones we consider in this paper. ier20:abbringdaljord showed that ier15:fangwang's main identification claim is void--- that it has no implications for identification of the dynamic discrete choice model--- and that its main proof of identification is incorrect and incomplete. selfcontrolchan17 builds on ier15:fangwang's intuition. We emphasize that we do not believe that the incorrect results in ier15:fangwang invalidates the results in selfcontrolchan17. On the contrary, we think our results confirm that the model in selfcontrolchan17, and possibly also the one in ier15:fangwang, are formally identified.} Along the way, we provide a way to concentrate the empirical analysis of such models with nonparametric utility on the present-bias and standard discount factors.
After showing that the quasi-hyperbolic discount function parameters $\beta$ and $\delta$ are formally identified, we note that though their product $\beta \delta$ can be precisely estimated in finite samples, the parameters are likely to be estimated individually at comparably low levels of precision. We show that the poor finite-sample performance follows from the way the parameters $\beta$ and $\delta$ enter largely interchangeably in the identifying moment conditions. We illustrate this point with simulations of a simple three-period model with two exclusion restrictions, which we know to be point identified. We then demonstrate that it also holds in an infinite horizon application to the demand for margarine, which comes with many natural exclusion restrictions.
Our results suggest to more closely follow the experimental literature and develop an identification strategy explicitly around the concept of preference reversals. In thaler81's (thaler81) classic example, subjects who prefer one apple today to two apples tomorrow tend to prefer two apples one year and one day from now to one apple one year from now. Such preference reversals are the defining feature of present-biased time preferences. In empirical work, demand for commitment devices has been viewed as a strategic response to anticipated preference reversals and has been taken as evidence of sophisticated present-bias malmendierdellavigna06. A sophisticated agent may want to lock in her savings to avoid excessive spending by future present-biased selves, that is, to restrict her future choice sets without receiving a current period pay-off. Demonstrated willingness to restrict one's future choice set may be a more promising approach. It has yet not, to our knowledge, been used formally as part of an identification strategy for dynamic discrete choice models.
We first study a finite horizon dynamic discrete choice model in which agents may suffer from sophisticated present-bias. This model is similar to ier15:fangwang's (ier15:fangwang), but is nonstationary, and does not allow for partial naivity. In Section (ref), we study a stationary, infinite horizon version like ier15:fangwang's.
Time is indexed by $t = 1, \hdots, T$; with $T < \infty$. In each period $t$, the agent chooses an action $d_t$ from a finite set $\mathcal{D} = \{1, \hdots, K\}$. Prior to making this choice, the agent draws and observes vectors of state variables $x_t$ and $\epsilon_t = \{\epsilon_{1,t}, \hdots, \epsilon_{K,t}\}$. The observable (to the econometrician) states $x_t$ have finite support $\mathcal{X}$ and evolve as a controlled (by $d_t$) first order Markov process. For notational simplicity only, we take this process to be stationary, with Markov transition distribution $Q_k$ if $k\in{\cal D}$ is chosen. The utility shocks $\epsilon_{k,t}$ are independent from $x_t$ and prior states and choices, over time, and across choices, and have type 1 extreme value distributions.\footnote{Our analysis straightforwardly extends to the case in which the vectors $\epsilon_t$ are independent over time, with {\em known} continuous distributions $G_t$ on a common support $\mathbb{R}^K$. Note that the exact choice of $G_t$, for $t=1,\ldots,T$, does not impose testable restrictions on the type of data that we assume are available in this paper.}
If, in period $t$ and state $x$, the agent chooses $k$, she collects a flow of utility $u_{k,t}(x) + \epsilon_{k,t}$. We normalize $u_{K,t}(x) = 0$ for all $t \in 1, \hdots, T$ and $x \in \mathcal{X}$. This normalization is substantive, but is standard in the literature and cannot be rejected by the type of observational data on choices and states that we will assume available in this paper.\footnote{The way $u_{K,t}$ is normalized affects the model's implied behavioural responses to many, but not all, counterfactual interventions res14:noretstang,qme14:aguirregabirisuzuki,qe18:kalouptsidietal.}
The agent's discount function has two parameters: a non-negative and finite standard discount factor $\delta$ and a present-bias parameter $\beta \in (0,1]$. Since the horizon is finite, we do not require that the discount factor $\delta$ is smaller than one. If $\beta = 1$, the model reduces to one with standard geometric discounting. The present-bias parameter is bounded away from zero to distinguish present-bias from myopia.
Choices in dynamic discrete choice models are regulated by value functions. Since present-biased time preferences are time inconsistent, these value functions do not follow from a standard dynamic program. It is common to think about the values as summarizing the pay-offs to players in a Stackelberg-like game played between selves in different time periods elster85.
Let $\tilde{\sigma}_t: \mathcal{X}\times \mathbb{R}^K \rightarrow \mathcal{D}$ be an arbitrary choice strategy and $\mathbf{\tilde{\sigma}}_t = \{\tilde{\sigma}_\tau\}^T_{\tau = t}$ an arbitrary strategy profile. The agent's current choice specific value function, which regulates the choices, is
for $t<T$, with terminal value $w_{k,T}(x) = u_{k,T}(x)$. The agent trades off current utility versus future values by factor $\beta \delta$, but the stream of all future utilities are discounted geometrically by factor $\delta$ according to the perceived long run value function, which equals
for $t+1<T$, with terminal value $v_{T}(x;\tilde{\mathbf{\sigma}}_{T}) = \mathbb{E}_{\epsilon_T}\left[u_{\tilde{\sigma}_{T}(x, \epsilon_T),T}(x) + \epsilon_{\tilde{\sigma}_{T}(x, \epsilon_T),T}\right]$.
The perceived long run value depends on the current self's perceptions of its future selves' strategies $\tilde{\sigma}_{t+1}$. At the time of decision, each of these future selves have present-biased preferences which are in conflict with the current self's time consistent long run time preferences.
Since the agent is sophisticated, her perceptions of her future strategies are correct in equilibrium. Thus, in a sophisticated intrapersonal equilibrium, her selves use a perception perfect strategy odonoghuerabin99, which is a strategy profile $\mathbf{\sigma}^*_1$ such that each $\sigma^*_t$ is a best response to her perceived future strategy profile ${\mathbf{\sigma}}^*_{t+1}$:
Here, $w_{k,T}(x;{\mathbf{\sigma}}^*_{T+1})$ should be read as $w_{k,T}(x)$.
It is easy to show, by backward induction from time $T$, that a perception perfect strategy exists and is unique, up to the resolution of ties in the decision in (ref). Because $\epsilon_t$ is continuously distributed, the implied equilibrium probabilities
that the agent chooses $k\in{\cal D}$ in state $x\in{\cal X}$ are unique.
Suppose that the data provide the state transition probabilities $Q_1,\ldots,Q_K$; and the conditional choice probabilities $s_{k,t}(x)=\Pr(d_t=k | x_t=x)$ for all $k\in{\cal D}$; $t=1,.\ldots,T$; and $x\in{\cal X}$.\footnote{ In empirical applications, one typically observes draws from the joint process $\{(x_t,d_t); t=1,\ldots,T\}$. Under our model's assumptions, this process is first order Markovian, with a transition distribution that is fully characterized by the state transition and conditional choice probabilities. Of course, this Markov property can be tested and its rejection may point to, for example, persistent unobserved heterogeneity. Such unobserved heterogeneity can be handled in the usual way. Here, we concentrate on what can be learned from dynamic discrete choice data once individual state transition and conditional choice probabilities are identified.} This section studies the extent to which these data uniquely determine the model primitives ${Q}_1,\ldots,{Q}_K$; $\beta$; $\delta$; and $u_{1,t},\ldots,u_{K-1,t}$; $t=1,\ldots,T$. Clearly, the observed state transitions directly identify ${Q}_k$ and thus, because they were assumed rational, the agent's expectations. We therefore focus on the identification of the utility functions $u_{k,t}$ and the discount parameters $\beta$ and $\delta$ from the conditional choice probabilities for given $Q_1,\ldots,Q_K$. To this end, we assume that the observed conditional choice probabilities are generated by a perception perfect strategy of the model:
The choice probabilities only depend on the primitives through the value contrasts $w_{k,t}(x;\mathbf{\sigma}^*_{t+1}) - w_{K,t}(x;\mathbf{\sigma}^*_{t+1})$. In particular, (ref) implies that\footnote{The functional form of the mapping between value contrasts and choice probabilities is specific to the assumption that the $\epsilon_{k,t}$ have independent type 1 extreme value distributions, but the results given here extend to general known $G_t$.}
for all $k\in D/\{K\}$; $t = 1, \hdots, T$; and $x\in{\cal X}$. With the restriction that the choice probabilities add up to one over choices, (ref) gives
With $-\ln (s_{K,t}(x))$ in hand, (ref) determines $s_{k,t}(x)$ from the value contrasts.
Conversely, as in the case without present-bias res93:hotzmiller, using (ref), the current choice specific value contrasts can be uniquely recovered from the observed choice probabilities. Altogether, this implies that we can focus our identification analysis on the question to what extent the discount parameters and utilities are uniquely determined from the value contrasts $w_{k,t}(x;\mathbf{\sigma}^*_{t+1}) - w_{K,t}(x;\mathbf{\sigma}^*_{t+1})$, for given $Q_k$.
It is well known that the dynamic discrete choice model with geometric discounting ($\beta=1$) is not identified (nh94:rust, Lemma 3.3, and ecta02:magnacthesmar, Proposition 2). The underidentification carries over to its generalization with present-bias. Specifically, the following version of ecta02:magnacthesmar's (ecta02:magnacthesmar) Proposition 2 holds.
Theorem (ref) implies that $\beta$ and $\delta$ can only be identified if further data are available or additional assumptions are made. In this paper, we explore identification under exclusion restrictions on the utility functions. Our analysis focuses on the identification of $\beta$ and $\delta$. Theorem (ref) shows that, once $\beta$ and $\delta$ are identified, unique utility functions can be found that rationalize the choice data.
Because ${\cal X}$ is finite--- say it has $J$ elements--- it is convenient to express expectations in matrix notation. To this end, let $\mathbf{v}_t(\mathbf{\sigma}^*_t)$ be a $J\times 1$ vector that stacks the values of $v_t(x;\mathbf{\sigma}^*_t)$, $x\in{\cal X}$, and $\mathbf{Q}_k(x)$ a $1\times J$ vector that stacks the values of $Q_k(x'|x)$, $x'\in{\cal X}$, in corresponding order. Then, (ref) and (ref), with the normalization $u_{K,t}(x)=0$, give
Recall that, given the transition distributions $Q_k$, (ref) contains all information in the choice probabilities about the model's primitives.
We will concentrate the identification analysis on the discount factors by controlling the current period utility $u_{k,t}(x)$ in the right hand side of (ref) with exclusion restrictions and expressing the continuation value in terms of the discount factors and data only. As $\mathbf{Q}_k(x)$ and $\mathbf{Q}_K(x)$ are data, this only requires that we express the perceived long run values $\mathbf{v}_{t+1}(\mathbf{\sigma}^*_{t+1})$ in terms of the discount factors and data. To this end, first substitute (ref) and (ref) into (ref) to get
Next, as we can express the value contrast $w_{k,t+1}-w_{K,t+1}$ in terms of data using (ref), we substract $w_{K,t+1}(x;\mathbf{\sigma}^*_{t+2})$ from the first term in the right hand side of (ref) and add it to the second term, which gives
where
is the McFadden surplus (before observing $\epsilon_{t+1}$) for the choice among $k\in{\cal D}$ with utilities $w_{k,t+1}(x;\mathbf{\sigma}^*_{t+2}) - w_{K,t+1}(x;\mathbf{\sigma}^*_{t+2}) + \epsilon_{k,t+1}$. Under our assumption that $\epsilon_{t+1}$ is extreme value distributed, the right-hand side of (ref) reduces to the right-hand side of (ref), so that $m_{t+1}(x) = -\ln(s_{K,t+1}(x))$ is known from the choice data.\footnote{More generally, given $G$, $m_{t+1}(x)$ is a known function of $w_{k,t+1}(x;\mathbf{\sigma}^*_{t+2}) - w_{K,t+1}(x;\mathbf{\sigma}^*_{t+2})$, $k \in{\cal D}$, and thus, using (ref), of the choice probabilities arcidiaconomiller11.} The term $w_{K,t+1}(x;\mathbf{\sigma}^*_{t+2})$ can be expressed recursively as
Finally, as the expectation over $\epsilon_{t+1}$ in the right hand side of (ref) is effectively an expectation over implied actions $\sigma^*_{t+1}(x, \epsilon_{t+1})$, it can be expressed in terms of the observed choice probabilities using (ref):
Substituting (ref) and (ref) into (ref) gives
where $\overline{\mathbf{Q}}_{t+1}(x) = \sum_{k\in \mathcal{D}}s_{k,t+1}(x)\mathbf{Q}_k(x)$ is the expected state transition probability distribution under strategy $\sigma^*_{t+1}$ in state $x$. This mixture represents an expectation over how the choices of present-biased future selves control future state transitions, choices which are in conflict with the current self's long term preferences.
Define the $J\times J$ matrix of probability mixtures
where $\overline{\mathbf{Q}}_t$ stacks $\overline{\mathbf{Q}}_t(x)$ and $\mathbf{Q}_K$ stacks $\mathbf{Q}_K(x)$. Then, we can write (ref) as a recursive expression for $\mathbf{v}_{t+1}(\mathbf{\sigma}^*_{t+1})$ in vector notation:
Completing the recursion until the end of time $T$ expresses
in terms of the discount factors and data only. Substituting (ref) into (ref) gives
The log choice probability ratio in the left hand side of (ref) measures the observed propensity to choose $k$ over $K$ in state $x$. The right hand side of (ref) explains this observed propensity by the current period's utility difference $u_{k,t}(x)-u_{K,t}(x)=u_{k,t}(x)$ and a difference in continuation values, which is a polynomial in $\beta$ and $\delta$ with coefficients that are fully determined by the choice and transition data. We study identification from variation in these continuation values, under exclusion restrictions on primitive utility that control the effects of variation in the current period's utility. This formalizes the common intuition that holding current period utilities constant, current choice responses to variation in future values are informative about time preferences.
Our identification argument holds for exclusion restrictions on utilities between pairs of time periods, choices, states, or any combinations of the three. To simplify the exposition, we however focus on exclusion restrictions on utilities from one given choice between pairs of states. We primarily focus on the case in which we have two such exclusion restrictions, which is the minimum needed to identify the two unknown discount factors, $\beta$ and $\delta$. In applications, intuition for exclusion restrictions would typically deliver a {\em variable} that affects continuation values, but not the current period's utility. Such a excluded variable would typically imply more than two exclusion restrictions on states, which would further restrict the identified set of discount factors.
So, consider two exclusion restrictions on utility from choice $k\in{\cal D}/\{K\}$, indexed by $a$ and $b$. The first requires that
at time $t_a<T-1$, for states $x_{a,1}, x_{a,2} \in \mathcal{X}$ such that $x_{a,1}\neq x_{a,2}$. The second sets
at time $t_b\leq t_a<T-1$, for states $x_{b,1}, x_{b,2} \in \mathcal{X}$ such that $x_{b,1}\neq x_{b,2}$. Evaluating (ref) at choice $k$ and time $t_a$, differencing between states $x_{a,1}$ and $x_{a,2}$, and using (ref) gives
Similarly, (ref) and (ref) give
The left hand sides of (ref) and (ref) are scalars that are known from the data. They reflect choice differences between the states that figure in each exclusion restrictions and that are known from the data. The right hand sides of (ref) and (ref) are the model's implications for these same choice responses at discount factors $\beta$ and $\delta$. They are bivariate polynomials in $\beta$ and $\delta$ of order $T-t_a$ and $T-t_b$, respectively, with coefficients that are fully determined by the data and do not depend on the unknown utility functions $u_{k,t}$. By Theorem (ref), for any pair of discount factors $(\beta, \delta)$ that solves (ref) and (ref), unique utilities $u_{k,t}$ can be found that rationalize the choice data and, in particular, solve (ref). Therefore, the identified set can be characterized by finding all discount factors $(\beta,\delta)$ that solve the bivariate polynomial equations (ref) and (ref). Solving systems of bivariate polynomial equations is a well-understood problem where we can draw on standard results from algebraic geometry. Before we use these results to formally characterize the identified set, we first give a simple three-period example.
In this simple example, the solutions are easy to find. We however need a more systematic approach for the general case. We can analyze the solutions to the moment conditions using standard results from algebraic geometry. Our exposition builds on coxetal15. The moment conditions (ref) and (ref) can be expressed in terms of their Sylvester matrix. To construct the Sylvester matrix for an arbitrary bivariate system in $\beta$ and $\delta$, we choose either $\beta$ or $\delta$ as a base and treat it as constant without loss of generalization. Suppose the two moment conditions, taking $\delta$ to be the base, are polynomials $q_a$ and $q_b$ in $\beta$, of order $l$ and $m$, respectively, with coefficients $a_0, \hdots, a_l$ and $b_0, \hdots, b_m$, respectively. The Sylvester matrix with base $\delta$ is
an $(l+m) \times (l+m)$ matrix, where the blanks are zeros. The resultant, the determinant of the Sylvester matrix, is a polynomial in the base parameter. From Corollary 4, chapter 3, of coxetal15, each $\delta$ root of the resultant is a root of the moment conditions. This eliminates the base parameter $\delta$, and we can solve for the roots of the remaining parameter, in this case $\beta$. Though we use the resultant to show identification, it also gives an algorithm to solve for the roots of the system: for each value of $\delta$ that sets the resultant to zero, we can solve for the roots of the remaining univariate polynomial in $\gamma$. The literature shows that this algorithm may work well for low order polynomials, but becomes computationally inefficient and unstable when the order exceed e.g. 30. \\ \\ We illustrate the resultant using Example (ref) where we reparametrize the problem to be bivariate in $\gamma = \beta \delta$ and $\delta$. Write the moment conditions in (ref) and (ref) as
where
In this example, we get
The resultant is the determinant $Syl(f_a(\gamma, \delta), f_b(\gamma, \delta))_\gamma$
Setting the resultant to zero gives a unique root
if it exists. Inserting $\delta^*$ in (ref), which is linear in $\beta$, gives the solution, the common roots of the moment conditions.
We see that if $b_{1,1} = 0$, which violates the non-zero condition in (ref), then no root $\delta^*$ exists. This rejects the model: We can not find parameters $\beta$ and $\delta$ that rationalize the data. If $a_1 = 0$, which violates the rank condition in (ref), and $a_0 = 0$, then the resultant is zero for all values of the $\delta$. This implies that the moment conditions have a common factor: it has an infinity of solutions and all identification is lost. The resultant generalizes the rank condition in the geometric case, where the identifying moment condition is a univariate polynomial moment condition, to the hyperbolic case, where the identifying moment conditions are a bivariate set of polynomial moment conditions.
The example illustrates that the common factors of (ref) and (ref) characterize exceptions to the identification of $(\beta,\delta)$ from these moment conditions. We therefore need a definition of common factors. Write the moment conditions (ref) and (ref) as $f_a(\beta,\delta)=0$ and $f_b(\beta,\delta)=0$, respectively, with $f_a$ and $f_b$ $(T-t)$'th order polynomials.
A simple example of a common factor of (ref) and (ref) in the case that their left hand sides are zero (no choice responses) is $h(\beta,\delta)=\delta$.
Bezout's theorem generalizes the fundamental theorem of algebra to multivariate polynomials, see e.g. coxetal15. Theorem (ref) does not guarantee a solution. The zero set may be empty, such as in the example above. In that case, the model is rejected.\footnote{See qe20:abbringdaljord for a discussion of the empirical content of dynamic discrete choice models under exclusion restrictions.}
Except for certain special cases, such as when one moment condition is a multiple of the other, we have not found obvious economic interpretations of common factors. The existence of common factors can however easily be verified on a case-by-case basis by calculating the resultant. In practice, a resultant that is everywhere close to zero signals numerical instability similar to a regression matrix with a high condition number where identification is close to lost to multicollinearity.
Section (ref) analyzes a finite horizon decision model. We will now consider an infinite horizon, stationary version of that model. This is ier15:fangwang's (ier15:fangwang) model, but with sophisticated agents. We will provide a new equilibrium existence result for this model and show that Section (ref)'s identification results extend to it.
Take Section (ref)'s model, but let time be indexed by $t \in\{1, 2, \hdots\}$ and utility from choice $k$ in state $x$ at time equal $u_{k}(x) + \epsilon_{k,t}$, with $u_k$ time invariant. Also, restrict $\delta$ to $[0,1)$ to ensure convergence of the agent's expected discounted utility (recall that $\beta \in (0,1]$).
Because neither the model primitives nor the decision horizon change over time, the resulting decision problem is stationary. Consequently, it is natural for the agent to use a stationary Markov strategy, say $\tilde{\sigma}: \mathcal{X}\times \mathbb{R}^K \rightarrow \mathcal{D}$, to map the current state (but not time itself) into an action in each period. If the agent uses such a strategy $\tilde\sigma$ in all future periods, her current choice-$k$ specific value in state $x$ equals
with perceived long run value
The agent uses a stationary perception perfect strategy $\sigma^*$, which is a best response to employing the same strategy in all future periods:
In turn, the strategy $\sigma^*$ implies stationary conditional choice probabilities
We will exploit the specific econometric structure of our model, in particular the presence of additively separable and independent utility shocks, to demonstrate existence of a stationary perception perfect strategy.\footnote{ier15:fangwang claims, for a model of a partially naive agent that encompasses ours, that “[t]he existence of the perception-perfect strategy profile is shown in Peeters (2004) for the same class of stochastic games.” mu04:peeters, however, assumes a finite state space, so that its results cover neither ier15:fangwang's model nor ours. If anything, the analysis in mu04:peeters suggests that the pure and stationary perfection perfect strategies that we focus on do {\em not} always exist: For its case with finite states, it demonstrates that mixed strategies are needed to ensure existence.} To this end, we return to Section (ref)'s matrix notation, without time subscripts. From (ref), the current choice-$k$ specific values under perceived future use of $\sigma^*$ equal
Using a straightforward adaption of Section (ref)'s derivations, we can rewrite (ref) as \[ \mathbf{v}({\sigma}^*) = -\ln\mathbf{s}_K(\sigma^*)+\mathbf{w}_K(\sigma^*)+\delta(1-\beta)\sum_{k\in{\cal D}}\mathbf{S}_k(\sigma^*){\mathbf{Q}_k}\mathbf{v}({\sigma}^*), \] where $\mathbf{s}_k(\sigma^*)$ is a $J\times 1$ vector that stacks the conditional choice probabilities $s_{k}(x;\sigma^*)$, $x\in{\cal X}$, and $\mathbf{S}_k(\sigma^*)$ is a $J\times J$ diagonal matrix with the same probabilities on its diagonal.\footnote{As in Section (ref), $-\ln\mathbf{s}_K(\sigma^*)$ is a vector of McFadden surpluses. We could denote it with $\mathbf{m}(\sigma^*)$, but it is important for the equilibrium analysis to keep its dependence on $\mathbf{s}_K(\sigma^*)$ explicit.} Substituting (ref) for $k=K$ and rearranging gives
with $\mathbf{I}$ a $J\times J$ identity matrix. Note that $\beta\mathbf{Q}_K+(1-\beta)\sum_{k\in{\cal D}}\mathbf{S}_k(\sigma^*){\mathbf{Q}_k}$ is a stochastic matrix and that $\delta\in[0,1)$, so that the inverse in the right hand side of (ref) exists.
Now stack $\mathbf{s}_1(\sigma^*),\ldots,\mathbf{s}_K(\sigma^*)$ in a $K J\times 1$ vector $\mathbf{s}(\sigma^*)$. Note that $\mathbf{s}(\sigma^*)$ takes values in ${\cal S}\equiv\left\{(\tilde{\mathbf{s}}_1,\ldots,\tilde{\mathbf{s}}_K)\in\mathbb{R}^{K J}\;\vline\; \tilde{\mathbf{s}}_1,\ldots,\tilde{\mathbf{s}}_K\geq 0;~\sum_{k\in{\cal D}}\tilde{\mathbf{s}}_k=\bm{1}\right\}$, with $\bm{1}$ a $K\times 1$ vector of ones. Substituting (ref) in (ref) gives $\mathbf{w}_k(\sigma^*)=\bm{\omega}_k(\mathbf{s}(\sigma^*))$, with $\bm{\omega}_k: {\cal S}\rightarrow\mathbb{R}^J$ such that \[ \bm{\omega}_k(\tilde{\mathbf{s}})= \mathbf{u}_{k} + \beta \delta\mathbf{Q}_k\left[\mathbf{I}-\delta\left(\beta\mathbf{Q}_K+(1-\beta)\sum_{k\in{\cal D}}\tilde{\mathbf{S}}_k{\mathbf{Q}_k}\right)\right]^{-1}\left[-\ln\tilde{\mathbf{s}}_K+\mathbf{u}_K\right]. \] Note that $\bm{\omega}_k$ does not depend on $\sigma^*$, so that $\mathbf{w}_k(\sigma^*)=\bm{\omega}_k(\mathbf{s}(\sigma^*))$ only depends on $\sigma^*$ through $\mathbf{s}(\sigma^*)$. Let $\bm{\omega}_k(x;\mathbf{s}(\sigma^*))$ be the element of $\bm{\omega}_k(\mathbf{s}(\sigma^*))$ corresponding to state $x\in{\cal X}$. For $\sigma^*$ to be a stationary perception perfect strategy, it needs to be a best response to the perception that later choices are made with the probabilities in $\mathbf{s}(\sigma^*)$, as in (ref) with $w_k(x;\sigma^*)=\bm{\omega}_k(x;\mathbf{s}(\sigma^*))$. Such a strategy exists.
Theorem (ref) does not guarantee that the stationary perception perfect strategy $\sigma^*$ is unique, not even up to equivalence of the implied choice probabilities $\mathbf{s}(\sigma^*)$. Specifically, its proof relies on the application of Brouwer's fixed point theorem, which does not rule out that there are multiple fixed points $\mathbf{s}^*$.
Suppose that data allows us to determine conditional choice probabilities $\mathbf{s}$ and state transition probabilities $\mathbf{Q}_1,\ldots,\mathbf{Q}_K$, with
for some stationary perception perfect strategy $\sigma^*$.\footnote{As in the finite horizon case (see Footnote (ref)), we typically observe the process $\{x_t.d_t\}$ over a finite number of periods. Under our model's assumptions, this process is a stationary Markov process. This can be tested. Unobserved heterogeneity, which may now include heterogeneous equilibrium selection, may render the data non-Markovian. Time variation in the parameters may lead to nonstationarity. Standard approaches to deal with heterogeneity can be applied to first identify agent-level choice and state transition probabilities. With these in hand, our results can be applied.} Under this assumption, Section (ref)'s (non-)identification results extend to the infinite horizon model.
First, we establish that, without further restrictions, the model is just identified if we fix the discount factors.
Like Theorem (ref) before, Theorem (ref) implies that further model restrictions are needed to identify $\beta$ and $\delta$. Moreover, it suggests that we concentrate the identification analysis on $\beta$ and $\delta$, as it ensures that we can always find utilities that rationalize the choice probabilities once we have identified those two parameters.
As in Section (ref), we proceed by imposing the minimum of two exclusion restrictions on utility required for identification of $\beta$ and $\delta$: For some $k\in{\cal D}/\{K\}$,
for states $x_{a,1}, x_{a,2}, x_{b,1}, x_{b,2} \in \mathcal{X}$ such that $x_{a,1}\neq x_{a,2}$ and $x_{b,1}\neq x_{b,2}$. We also again set $\mathbf{u}_K=\bm{0}$. A derivation as in Section (ref), based on (ref) and (ref), using (ref), (ref), and (ref) to substitute observed choice probabilties for unknown quantities, and (ref) and (ref) to difference out flow utility, gives
where\footnote{Note that $\mathbf{A}(\beta,\delta)$ is the same matrix as $\mathbf{A}(\tilde{\mathbf{s}})$ in Footnote (ref). We allow this slight abuse of notation, instead of introducing a new symbol for this matrix, to clearly indicate both matrices are the same.}
$\Delta^2_b\ln\mathbf{s}_k$ and $\Delta^2_b\mathbf{Q}_k$ are analogously defined,
and $\mathbf{m}\equiv-\ln\mathbf{s}_K$. As we noted before, the inverse of the $J\times J$ matrix $\mathbf{A}(\beta,\delta)$ exists, for all $\beta$ and $\delta$ in their domains. Using that $A(\beta,\delta)^{-1}=|\mathbf{A}(\beta,\delta)|^{-1}\overline{\mathbf{A}}(\beta,\delta)$, we can rewrite (ref) into
Here, the determinant $|\mathbf{A}(\beta,\delta)|$ is a polynomial of degree $J$ in both $\beta$ and $\delta$ (for a total degree of $2 J$) and the adjugate matrix $\overline{\mathbf{A}}(\beta,\delta)$ contains the cofactors of $\mathbf{A}(\beta,\delta)$, which are polynomials of degree $J-1$ in both $\beta$ and $\delta$. Consequently, the left and right hand sides of (ref) and (ref) are polynomials of degree $J$ in both $\beta$ and $\delta$. Their coefficients depend on the observed choice and state transition probabilities only. So, under an assumption that excludes pathologies, we can mimic the proof of Theorem (ref) to bound the cardinality of the identified set.
Finally, note that the above analysis also implies identification results for the special case with geometric discounting ($\beta=1$). For given $\beta\in[0,1)$, the difference between the left and right hand sides of (ref) is a $J$'th degree polynomial in $\delta$, with coefficients that are known from the data. Provided that this polynomial is nonzero, by the fundamental theorem of algebra, at most $J$ discount factors $\delta$ satisfy (ref) for given $\beta$. So, under exclusion restriction (ref) and assumptions that ensure that the implied polynomial is nonzero, the identified set of the special case with geometric discounting contains at most $J$ elements. This sharpens qe20:abbringdaljord's (qe20:abbringdaljord) Theorem 1 and Corollary 1, which only established that the identified set is discrete, and finite if discount factors near 1 are excluded. For future reference, we explicitly state this result as\footnote{Note that qe20:abbringdaljord denote the geometric discount factor with $\beta$ instead. Also, qe20:abbringdaljord allowed more generally for an exclusion restriction across either states, or choices, or both. Our analysis here directly extends to such a more general exclusion restriction. We keep this implicit for notational simplicity.}
Both Theorems (ref) and (ref) rely on the fact that ${\cal X}$ is finite. In contrast, we conjecture that, under some regularity conditions, qe20:abbringdaljord's (qe20:abbringdaljord) Theorem 1 and Corollary 1 extend to continuous state variables. See blevins2014 for a relevant approach to this case.
Brand loyalty has been studied in static choice models since the 1980's. daljorddubekong2020 notes notes that in these models, a positive coefficient on repeated brand purchases has been interpreted as evidence of brand loyalty. If we leave aside the confounding of persistent heterogeneity and state dependence in these models and accept this interpretation of persistence in brand choice, it still does not explain why consumers are brand loyal.
farrellklemperer07 considers brand loyalty as a classic case of a switching cost that creates a lock-in. Switching costs may be pecuniary, such as terminating a mobile subscription plan, or psychological, such as breaking a habitual brand choice. Using a beckermurphy88 rational addiction like argument, gordonsun15 argues that if consumers are aware that their brand choice creates a lock-in, then forward-looking consumers would take the lock-in into account when choosing among brands. In a consumer packaged goods context, a forward-looking consumer may be wary of choosing a brand with a high average price and little price variation. daljorddubekong2020 tested for forward looking behaviour in a dynamic choice model with brand loyalty and geometric discounting. It found a value of $\delta = 0.41$ on a weekly level, statistically significant, but economically small. In this application, we extend daljorddubekong2020's analysis to allow for present-biased preferences using the identification argument of the previous section applied to the same data.
The data are from the margarine household purchase panel of Nielsen-Kilts HMS that was used in dubeetal2010. [Description and summary statistics of proprietary data suppressed.]
In each period, the household makes a discrete choice $d_{i,t} \in \mathcal{D}$. The choice set has four brands, indexed 1 to 4, and the reference choice $K = 5$ is purchasing some spreadable good other than margarine. The state variables are prices $\bm{p} = [p_1, \hdots, p_4] \in \mathbb{R}^4_+$ and brand loyalty $l \in \mathcal{D}\backslash \{K\}$. The prices are assumed to follow a first-order Markov process that is exogenous to the household's choices. The brand loyalty state $l$ is controlled by the choice and is equal to the last purchase of the inside goods.
The utilities are
This utility function allows for brand loyalty for some products (positive state dependence) and variety seeking in other products (negative state dependence). This specification is more flexible than the parametric counterpart in dubeetal2010 which uses $\mathbbm{1}(k = l)$ as the loyalty state variable and a parametric form that restricts state dependence to be either positive or negative for all products. Substituting in the utilities in (ref), the current choice specific value function has the structure of (ref) and (ref), with choice probabilities given by (ref).
daljorddubekong2020 observes that in the standard dynamic differentiated products model, the utility of purchasing one of the inside goods excludes the prices of rival products, such that
A pair of states $x_{1} = [p_k, \bm{p}_{-k}, l]$ and $x_{2}=[p_k, \bm{p}'_{-k}, l]$ for which the own price and brand loyalty state are constant, but where at least one rivals's product price shifts ($\bm{p}_{-k} \neq \bm{p}'_{-k}$) satisfies the exclusion restrictions in (ref) or (ref) for choice $k$. The discount factor is therefore identified in these models if there are at least two such pairs of states that satisfy the regularity conditions. A theoretical virtue of this identification strategy is that it is always available in the discrete choice models used to analyze CPG demand without further assumptions.
The regularity conditions in Theorem (ref) require that the moment condition corresponding to a given exclusion restriction is non-zero. A necessary condition for a non-zero right hand side of either (ref) or (ref) is
Since prices are assumed to evolve independently of household choices, then unless choice $k$ shifts the next period's loyalty state $l$, it follows that
So if $k = l$, the non-zero condition in (ref) fails and the corresponding moment condition is not informative, though the exclusion restriction holds. We define the set of pairs of states that satisfy necessary conditions for identification
As in daljorddubekong2020, the prices of the four brands are discretized in 30 bins. Along with four brand loyalty states, this gives a total of 120 states. We estimate the parameters $\beta$ and $\delta$ from the moment conditions that can be generated from $\mathcal{X}^{id}$ by minimum distance with quadratic loss. The parameters $\beta$ and $\delta$ are identified if $\mathcal{X}^{id}$ contains at least two unique pairs. In this application, we have 3400.
The geometric discount factor was estimated without restrictions its domain and replicates the result in daljorddubekong2020. [Estimation results of the hyperbolic model on proprietary data suppressed.]
The previous section showed that the exclusion restriction approach that identifies geometric time preferences in dynamic discrete choice models formally extends to present-biased time preferences. The identification argument for the geometric case is intuitive. The moment condition derived from the exclusion restriction captures the choice response to variation in future values for a fixed current utility. The intuition behind the identification of present-bias is less obvious.
The defining feature of present-bias that the identification strategy must capture is preference reversals. A well-known example of preference reversals is Thaler's apples: while most people prefer an apple today to two apples tomorrow, the same people also prefer two apples one year and one day from now to one apple one year from now. Such observed choice contrasts are direct and intuitive evidence of preference reversals that are commonly used in lab studies of time preferences, see e.g. marziliericsonetal15.
Inferring preference reversals from observed choices in dynamic discrete choice models requires utility restrictions. For instance, Thaler's example can be rationalized with non-stationary utilities without invoking preference reversals. Identification of hyperbolic discount functions using exclusion restrictions on utilities relies on evidence of preference reversals similar to Thaler's example, but in a less transparent way. Each moment condition represents the difference in values of a particular choice contrasts that involves a sequence of expected choices. If we fix $\beta = 1$, then $\delta$ is determined from one of the moment conditions. Once $\delta$ is determined, $\mathbf{u}$ is uniquely determined. For a particular set of $\delta$ and $\mathbf{u}$, we can test if these preferences are consistent with the second moment condition, which then has no free parameters. If the second moment condition is not satisfied at $\beta = 1$, it is because we have a preference reversal: a $\delta$ and a $\mathbf{u}$ that rationalizes the choice contrast in the first moment conditions is inconsistent with the revealed preference for the choice contrast in the second moment condition. By allowing $\beta \in (0,1)$ and $\delta$ to simultaneously determine both moment conditions, the observed preference reversal informs the present bias parameter.