EconBase
← Back to paper

Sensitivity Analysis for Dynamic Discrete Choice Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

85,167 characters · 23 sections · 83 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sensitivity Analysis for Dynamic Discrete Choice Models

\fontfamily{ppl}

abstractIn dynamic discrete choice models, some parameters, such as the discount factor, are being fixed instead of being estimated. This paper proposes two sensitivity analysis procedures for dynamic discrete choice models with respect to the fixed parameters. First, I develop a local sensitivity measure that estimates the change in the target parameter for a unit change in the fixed parameter. This measure is fast to compute as it does not require model re-estimation. Second, I propose a global sensitivity analysis procedure that uses model primitives to study the relationship between target parameters and fixed parameters. I show how to apply the sensitivity analysis procedures of this paper through two empirical applications. \noindentKeywords: Dynamic discrete choice models, sensitivity analysis, discount factor

Introduction

In dynamic discrete choice (DDC) models, some structural parameters are often fixed rather than being estimated. A leading example of such parameters is the discount factor, which is nonparametrically unidentified without further restrictions rust1994hdbk, magnacthesmar2002ecta. Apart from DDC models, it is also common to have model parameters being fixed exogenously in calibrating general equilibrium models dawkinsetal2001hoe. In the rest of this paper, the parameters that are fixed in the estimation procedure are referred to as fixed parameters.

The choice of the discount factor in DDC models is usually based on the value used in related papers, some relevant rates of return, or some values larger than 0.9. Researchers may conduct sensitivity analysis by repeating part of the estimation at a few other values of the discount factor. But this can be time-consuming because estimating the full model once can take days or weeks. Hence, the current practice can only offer limited information on how the conclusions are affected by the discount factor due to the high computational cost.

In this paper, I propose new approaches to conduct local and global sensitivity analysis for DDC models with respect to the fixed parameters.

To begin with, I develop a local sensitivity measure that examines the change of the target parameter due to a small change in the fixed parameter. I show that the local sensitivity measure is low-cost to compute because it can be obtained by solving a system of linear equations. Researchers do not need to re-estimate the model in order to compute the local sensitivity measure. In addition, reporting the local sensitivity measure can be more informative than re-estimating the model at a few other values of the fixed parameters because readers can estimate the target parameter at their chosen values of the fixed parameters. I show that the local sensitivity measure can serve as a good local approximation through two empirical applications.

If the literature has some consensus that a certain fixed parameter lies in a tight interval, then local sensitivity analysis can already be informative to approximate how the conclusion changes in such an interval. However, this may not always be the case. For the discount factor, evidence from behavioral economics shows that there can be a lot of variation in the discount factor depending on the context and the sample fredericketal2002jel. A recent study by kongetal2022wp estimates the discount factor for various consumer goods and reports a wide range of discount factors among different products, from 0.357 (for mayonnaise) to 0.999 (for peanut butter). As a result, researchers may be concerned that their conclusion does not hold at another value of the discount factor.

This motivates the global sensitivity analysis in the current paper that examines the structure of DDC models and finds conditions on the model primitives under which the target parameter is monotone in the fixed parameter. Monotonicity can be a useful property because of two computational benefits. First, parameters that are monotone in the fixed parameter are bounded by the endpoints. Second, researchers can easily estimate the breakdown point at which the conclusion changes horowitzmanski1995ecta, klinesantos2013qe, mastenpoirier2020qe. In the current practice of sensitivity analysis, researchers typically re-estimate the model at a few (e.g., three) neighboring values of the fixed parameter used for the main analysis because estimating the model once is costly. For instance, barwickpathak2015rand, fowlieetal2016jpe, and igami2017jpe repeat the estimation using discount factors around the one used for the main analysis, and examine how the parameter estimates change with the discount factor. Although monotonic patterns are usually shown, they might not necessarily generalize to the entire support of the discount factor. Using the discount factor as a leading case of fixed parameters in DDC models, I show that utility can be monotone in the discount factor under some conditions on the transition matrices and conditional choice probabilities. However, counterfactuals are not necessarily monotone in the discount factor, even if utility is monotone in the discount factor. Therefore, I also propose a constrained optimization approach for global sensitivity analysis for more general target parameters and fixed parameters.

As will be discussed in Section (ref), the methodology of this paper is not specific to single-agent DDC models nor the discount factor. The procedures proposed in this paper can be applied to other constrained optimization problems (see problem (ref) ahead) with a unique solution and fixed parameters. Solving the single-agent DDC model using the full solution method is just an example with such a structure. Some other potential economic applications that contain a constrained optimization structure and a subset of parameters being fixed include: dynamic games egesdaletal2015qe, dynamic matching verdierreeling2021restud, chenchoo2022ej, international trade ossa2014aer, and productivity yang2021aejma. See aguirregabiriaetal2021hdbk for a recent comprehensive review that contains many DDC examples related to industrial organization.

Related literature

This paper contributes to several strands of literature in economics. First, it is related to the literature on sensitivity analysis. andrewsetal2017qje and honoreetal2020jae develop measures to analyze the sensitivity of parameter estimates to the moments. More closely related are the papers by iskrev2019jedc and jorgensen2023restat that conduct sensitivity analysis with respect to calibrated parameters. The former paper focuses on Bayesian approaches to macroeconomic models. The latter proposes a local sensitivity measure to study the sensitivity with respect to calibrated parameters when the target parameters are estimated from minimizing an unconstrained optimization problem with a Generalized Method of Moments (GMM) objective. The first contribution of the current paper is to provide computationally attractive tools for conducting local sensitivity analysis for estimators obtained from constrained optimization problems. Fixed-point constraints are common in economic problems to represent equilibrium conditions. The current paper focuses on sensitivity analysis with respect to the fixed parameters and is different from some other recent papers in the sensitivity analysis literature such as armstrongkolesar2021qe which propose confidence intervals robust to local misspecification for overidentified moment condition models, bonhommeweidner2022qe which focus on robustness to misspecification within a larger class of models, and christensenconnault2022wp which examine sensitivity to the distribution of the latent variables.

The global sensitivity analysis section of this paper shares a similar theme as the literature on monotone comparative statics (e.g., topkis1998book). light2021mor is a recent paper that provides conditions for the policy function to be monotone in the discount factor, the parameters in the payoff function, or the transition probability function for Markov decision processes. But light2021mor considers a different model than the one in the current paper and does not consider estimation of model parameters. I also do not impose conditions like increasing differences that are required in light2021mor.

Outline

The rest of the paper is organized as follows. Section (ref) describes the setup and notation. Sections (ref) and (ref) describe the methodology for local and global sensitivity analysis, respectively. Section (ref) contains two empirical applications in which I apply the methodology to the seminal bus engine replacement example in rust1987ecta and to a recent dynamic matching model in chenchoo2022ej. Section (ref) concludes. All proofs can be found in the appendix.

Model

In this paper, the rust1987ecta model is used as the running example to illustrate the local and global sensitivity analysis procedures. In order to introduce the relevant notations, this section starts by describing the canonical single-agent DDC model and common assumptions. Then, I outline some common solution methods for DDC models and their connection with constrained and unconstrained optimization problems that are relevant for the sensitivity analysis procedures in Sections (ref) and (ref).

Notations

Consider a DDC model, where time is indexed by $t = 1, \ldots, T$. In each period $t$, each agent $i \in \cI \equiv \{1, \ldots, N\}$ chooses an action $a_{it} \in \cA \equiv \{0, 1, \ldots, A\}$ to maximize discounted future utility based on the state variables $s_{it}$. The vector of state variables can be decomposed as $s_{it} \equiv (x_{it}, \epsilon_{it})$, where $x_{it} \in \cX \equiv \{1, \ldots, X\}$ is observable by the agents and the researcher and $\epsilon_{it} \equiv (\epsilon_{0it}, \ldots, \epsilon_{Ait}) \in \bR^{A+1}$ is unobservable to the researcher. In addition, assume that $X$ is finite, $\epsilon_{ait}$ is i.i.d. across agents, choices, and states, and that $\epsilon_{it}$ is continuously distributed and has full support over $\bR^{A+1}$.

Each agent chooses the sequence of actions to maximize discounted future utility: \[ \max_{\{a_{it}\}} \bE \left[ \sum^\infty_{t=1} \beta^t \widetilde \pi_i(a_{it}, s_{it};\theta) \right], \] where $\beta \in [0, 1)$ is the discount factor common across agents and $\widetilde\pi_i(a_{it}, s_{it}; \theta)$ is the utility function of choosing action $a_{it}$ at state $s_{it}$ and parameterized by $\theta \in \Theta \subseteq \bR^{d_{\theta}}$.

Single-agent dynamic discrete choice model

In this section, I consider a single-agent stationary DDC model, with the following standard assumptions (see, e.g., hortacsujoo2023bk). I omit the $i$ subscript because there is only one agent.

assu[Additive separability] The utility function can be written as \[ \widetilde\pi(a_{t}, s_{t}; \theta) = \pi(a_{t}, x_{t}; \theta) + \epsilon_{t}, \] where $\pi(a_{t}, x_{t}; \theta)$ is bounded and monotone in $x_{t}$.
assu[i.i.d. error terms] For any $\epsilon_{t}$, $\bP[\epsilon_{t+1}|\epsilon_{t}] = \bP[\epsilon_{t+1}]$.
assu[Conditional independence] Given $(x_{t}, a_{t})$ observed, $x_{t+1} \indep (\epsilon_{t}, \epsilon_{t+1})$.

Let $\overline{V}(x_{t}, \epsilon_{t})$ be the agent's value function. By Bellman's principle of optimality, the value equation can be written as \[ \overline{V}(x_{t}, \epsilon_{t}) = \max_{a \in \cA} \left\{ \pi(a, x_{t}; \theta) + \epsilon_{at} + \beta \bE[\overline{V}(x_{t+1}, \epsilon_{t+1}) | a , x_{t}]\right\}. \] Define the choice-specific value function as \[ v(a, x) \equiv \pi(a, x; \theta) + \beta \bE[\overline{V}(x_{t+1}, \epsilon_{t+1}) | a_{t} = a , x_{t} = x], \] for any $a \in \cA$ and $x \in \cX$. The conditional choice probability (CCP) of choosing action $a \in \cA$ at state $x \in \cX$ is given by \[ \bP[a_{t} = a | x_{t} = x] \equiv \bP\left[\left. a \in \argmax_{j \in \cA} \ \left\{ v(j, x) + \epsilon_{jt} \right\} \right| x_{t} = x \right]. \]

The following assumption on the distributions of the unobservables is standard in the literature.

assu[Distribution of the unobservables] $\epsilon_{at}$ follows a mean-zero type-1 extreme value (T1EV) distribution for each $a \in \cA$ and time $t$.

Under Assumption (ref), the CCP can be written as

equation[equation omitted — 124 chars of source]

for each $a \in \cA$ and $x \in \cX$.

Next, let the state transition be governed by the Markov transition matrix $Q_a$, where the $(x, x')$-entry of $Q_a$ is the probability of transitioning from state $x$ in period $t$ to state $x'$ in period $t+1$ when action $a$ is chosen in period $t$, i.e., $q(x'|x, a) \equiv \bP[x_{t+1} = x' | x_{t} = x, a_{t} = a]$. Let $Q_a(x) \equiv (q(1|x,a), \ldots, q(X|x, a))'$ be the $x$-th row of the transition matrix $Q_a$ for any $a \in \cA$. Define the ex ante value function as $V(x) \equiv \bE[\overline{V}(x, \epsilon_{t})]$, and write $V \equiv (V(1), \ldots, V(X))'$. Then, $V$ satisfies the following fixed-point relationship using Assumption (ref) that the unobservables follow a mean-zero T1EV distribution:

align[align omitted — 141 chars of source]

for any $x \in \cX$. Let $\Psi^V$ represents the Bellman operator on the right hand side of (ref), then the fixed-point relationship (ref) can be summarized as follows

equation[equation omitted — 65 chars of source]

Finally, this section ends with another representation of the flow utility and its connection with CCP. This representation is useful for global sensitivity analysis in Section (ref). Using Lemma 1 of arcidiaconomiller2011ecta, there exists a real-valued function $\psi_a(\cdot)$ such that \[ V(x) = v(x,a) + \psi_a(p(x)), \] for any $a \in \cA$ and $x \in \cX$. Let $\pi_a \equiv (\pi(1,a; \theta), \ldots, \pi(X,a; \theta))'$ be the vector of utility functions at action $a \in \cA$. Following the discussion in kalouptsidietal2021wp, kalouptsidietal2021qe, the vector $\pi_a$ can be expressed as

align[align omitted — 63 chars of source]

for any $a \in \cA \backslash \{A\}$, where

align*[align* omitted — 118 chars of source]

Under Assumption (ref) that the unobservables follow a mean zero T1EV distribution, it follows that $\psi_a(p) = -\log p_a(x)$ for any $a \in \cA$ and $x \in \cX$ (see also hotzmiller1993restud), and $b_a(p)$ can be written as \[ b_a(p) = -A_a \log p_A(x) + \log p_a(x), \] for any $x \in \cX$ and $a \in \cA \backslash \{A\}$.

Solution methods

There are different methods to solve DDC models (see aguirregabiriaetal2021hdbk and hortacsujoo2023bk for details). I briefly outline three common approaches in this section in order to emphasize their structure as constrained or unconstrained optimization problems that would fit into the local sensitivity analysis framework in the next section.

Let $L(\theta, V; \beta)$ be the likelihood function. Here, I introduce $\beta$ as the argument of the likelihood function after the semicolon to indicate that it is a parameter that researchers need to specify in advance, and is fixed throughout the estimation procedure.

The nested fixed-point method (NFXP) by rust1987ecta involves value function iteration and contains two loops. The inner loop takes the parameter $\theta$ as given and finds the fixed point that solves equation (ref). The outer loop finds the parameter $\theta$ that maximizes the likelihood function. sujudd2012ecta show that NFXP is equivalent to solving it by a mathematical program with equilibrium constraints (MPEC):

align[align omitted — 158 chars of source]

Upon convergence, the solution must satisfy the Bellman equations and maximize likelihood. As a result, the system (ref) can be used as a starting point for local sensitivity analysis if the researcher uses NFXP or MPEC to solve the DDC model.

Two-step CCP methods can also be written in a similar manner. Using the hotzmiller1993restud inversion, the ex ante value function can be written as \[ V(x) = v(a, x) - \log \bP[a_{t} = a|x_{t} = x]. \] By a suitable normalization, such as $v(A, x) = 0$ for all $x \in \cX$ hortacsujoo2023bk, the choice-specific value function can be written as

equation[equation omitted — 106 chars of source]

for each $a \in \cA \backslash\{A\}$ and $x \in \cX$. Hence, with a given estimator of the CCP, the ex ante value functions can be estimated via (ref) over all choices and states. With the estimated $V$ and substituting the constraint into the objective, it becomes an unconstrained optimization problem.

aguirregabiriamira2002ecta propose the nested pseudo-likelihood method that iterates on the policy function instead. Their fixed-point equation is written as

equation[equation omitted — 64 chars of source]

where $\Psi^P$ is the policy function operator. The $K$-stage policy iteration estimator takes the estimator of the policy function from the previous stage $\widehat P^{K-1}$ and solves the following problem that updates the policy function via equation (ref):

align*[align* omitted — 130 chars of source]

Running example: The bus engine replacement problem

For the rest of this paper, I provide examples in terms of the seminal bus engine replacement problem in rust1987ecta. In this problem, Harold Zurcher, the manager, observes the bus mileage since the last engine replacement. The bus mileage for each bus $i = 1, \ldots, M$ is denoted by $x_{it}$ and the unobservable state variable is $\epsilon_{it}$. In each period, Zurcher chooses to replace ($a_{it} = 1$) or maintain ($a_{it} = 0$) the bus engine. Assume that the utility functions are the same across the buses. Let $\theta \equiv (\text{MC}, \text{RC})$, $\text{MC}$ be the maintenance cost, $\text{RC}$ be the replacement cost, and $c(x, \text{MC})$ be the cost of maintaining engine at mileage $x = 1, \ldots, X$. The utility function is given by \[ \widetilde \pi (a, x, \epsilon_{it} ; \theta) = \pi(a, x; \theta) + \epsilon_{ait}, \] where $\pi(0, x; \theta) = \text{RC} + c(x, \text{MC})$ and $\pi(1, x; \theta) = c(0, \text{MC})$ for any $x \in \cX$. Here, $Q_0$ and $Q_1$ are the two Markov transition matrices that correspond to the actions that choose to maintain and replace the engine, respectively. Bus mileage is reset to 1 if $a_{it} = 1$. Otherwise, the transition probability of mileage follows a multinomial distribution as below:

itemize• If $x \leq X - 2$, the transition probability is \[ \bP[x_{it+1} = x' | x_{it} = x, a_{it} = 0] = \begin{cases} \phi_1 & ,\ x' = x + 1 \\ \phi_2 & ,\ x' = x + 2 \\ 1 - \phi_1 - \phi_2 & ,\ x' = x \end{cases}. \] • If $x \leq X - 1$, the transition probability is \[ \bP[x_{it+1} = x' | x_{it} = x, a_{it} = 0] = \begin{cases} \phi_1 & ,\ x' = x + 1 \\ 1 - \phi_1 & ,\ x' = x \end{cases}. \]$\bP[x_{it+1} = X | x_{it} = X, a_{it} = 0] = 1$.

The fixed-point relationship for the Bellman equation in this example can be written explicitly as follows \[ V(x) = \log \left\{ \exp[\text{RC} + c(x, \text{MC}) + \beta Q_0(x)'V ] + \exp[c(0, \text{MC}) + \beta Q_1(x)'V ] \right\}, \] for each $x \in \cX$.

Local sensitivity analysis

Let $V \in \cV \subseteq \bR^{d_V}$ be a vector of auxiliary parameters and $\gamma \in \Gamma \subseteq \bR^{d_{\gamma}}$ be a vector of fixed parameters. Assume that a researcher is interested in estimating $\theta$ through the following constrained optimization problem by first fixing the parameter $\gamma$ as follows:

align[align omitted — 173 chars of source]

where $L(\theta, V; \gamma)$ is the criterion function, and the constraint $V = F(\theta, V; \gamma) $ describes some fixed-point relationship that captures the equilibrium constraints.

In terms of the DDC model in Section (ref), $\theta$ is the utility parameter, $V$ is the value function, and $\gamma$ is the discount factor. For the NFXP, $L(\theta, V; \gamma)$ corresponds to the likelihood function, and $V = F(\theta, V; \gamma)$ corresponds to the fixed-point equation (ref) based on Bellman optimality. Here, I allow $\gamma$ to be a vector because researchers might fix multiple parameters. For instance, igami2017jpe calibrates the discount factor, the rate of change of innovation cost, and the number of potential entrants.

Let $(\widehat{\theta}(\gamma), \widehat{V}(\gamma))$ be the solution obtained from solving the constrained optimization problem (ref). Note that the optimal solution has $\gamma$ as an argument because the constrained optimization problem is solved with the pre-specified $\gamma$ that is fixed throughout the estimation procedure.

The following assumptions on the constrained optimization problem (ref) are maintained throughout the paper:

assu\begin{enumerate} • $L$ and $F$ are continuously differentiable in $\theta$, $V$, and $\gamma$ around $(\widehat{\theta}(\gamma), \widehat{V}(\gamma))$. • $(\widehat{\theta}(\gamma), \widehat{V}(\gamma))$ is a regular point and is the unique solution to the optimization problem (ref) and belongs to the interior of $\Theta \times \cV$ for each $\gamma \in \Gamma$. \end{enumerate}

Assumption {\color{ucmaroon}(ref).(ref)} ensures that the derivatives in the local sensitivity measure exist. See rust1988siam and norets2010qe for results on differentiability results related to DDC models. Assumption {\color{ucmaroon}(ref).(ref)} ensures that the first-order condition holds as $(\widehat{\theta}(\gamma), \widehat{V}(\gamma))$ is a regular point bertsekas1999bk and that there is a unique solution to the optimization problem regardless of the value of the fixed parameter $\gamma \in \Gamma$.

Sensitivity measure

The gradient of $\widehat{\theta}(\gamma)$ with respect to $\gamma$, i.e.,

align[align omitted — 98 chars of source]

can be used as a measure of sensitivity. The $(i, j)$-component of the above matrix measures the change in the $i$-th target parameter for a unit change in the $j$-th fixed parameter. Depending on the parameter and the context, the following sensitivity measures may be easier to interpret:

enumerate• The elasticity gives the percentage change in $\widehat{\theta}(\gamma)_i$ for one percentage change in $\gamma_j$. It is defined by \begin{equation} \frac{\partial \widehat{\theta}(\gamma)_i}{\partial \gamma_j} \frac{\gamma_j}{\widehat{\theta}(\gamma)_i}, \end{equation} when $\widehat{\theta}(\gamma)_i, \gamma_j \neq 0$. • The semi-elasticity gives a unit change in $\widehat{\theta}(\gamma)_i$ for a percentage change in $\gamma_j$. It is defined by \begin{equation} \frac{\partial \widehat{\theta}(\gamma)_i}{\partial \gamma_j} \gamma_j, \end{equation} when $\gamma_j \neq 0$.

The sensitivity measure for $\widehat{V}(\gamma)$ can be defined analogously.

I show that computing the local sensitivity measures defined above amounts to solving a linear system of equations. The coefficients and constants in the linear system are evaluated at $(\widehat{\theta}(\gamma), \widehat{V}(\gamma))$ at the original $\gamma$. Hence, there is no need to re-estimate the model at another value of $\gamma$.

The following proposition summarizes the main result of the local sensitivity analysis procedure. For notational simplicity, I write $(\theta^\star, V^\star) \equiv (\widehat{\theta}(\gamma), \widehat{V}(\gamma))$.

propLet Assumption (ref) hold. Denote $\lambda^\star$ as the Lagrange multiplier for the constrained optimization problem (ref), evaluated at the optimal solution. The local sensitivity measure of $({\theta^\star}', {V^\star}', {\lambda^\star}')'$ with respect to $\gamma$, i.e., $\frac{\partial \theta}{\partial \gamma'}$, $\frac{\partial V}{\partial \gamma'}$, and $\frac{\partial \lambda}{\partial \gamma'}$, can be obtained by solving the following system of $(2d_V + d_{\theta})$ equations in $(2d_V + d_{\theta})$ unknowns: \begin{align} \begin{array}{rccrccrccl} A_{\theta, \theta} & \frac{\partial \theta}{\partial \gamma'} & + & A_{\theta, V} & \frac{\partial V}{\partial \gamma'} & - & (F_\theta)' & \frac{\partial \lambda}{\partial \gamma'} & = & - A_{\theta, \gamma} \\ A_{V, \theta} & \frac{\partial \theta}{\partial \gamma'} & + & A_{V, V} & \frac{\partial V}{\partial \gamma'} & + & (I_{d_V} - F_V)' & \frac{\partial \lambda}{\partial \gamma'} & = & - A_{V, \gamma} \\ F_\theta & \frac{\partial \theta}{\partial \gamma'} & + & (F_V - I_{d_V}) & \frac{\partial V}{\partial \gamma'} & & & & = & - F_{\gamma}, \\ \end{array} \end{align} if the system (ref) has a unique solution, where \begin{itemize} • $A_{x,y} \equiv \frac{\partial^2 L}{\partial x \partial y'} - R_{d_x} \frac{\partial \mathrm{vec}[(\frac{\partial F}{\partial x'})']}{\partial y'}$ and $F_x \equiv \frac{\partial F}{\partial x'}$ for $x \in \bR^{d_x}$ and $y \in \bR^{d_y}$. • $\mathrm{vec}(B)$ stacks the columns of the matrix $B \in \bR^{m\times n}$ into a column vector, i.e., $ \begingroup \begingroup\lccode`~=`, \lowercase{\endgroup \edef~{\mathchar\the\mathcode`, \penalty0 \noexpand\hspace{0pt plus 1em}} }\mathcode`,="8000 \mathrm{vec}(B) \equiv (b_{1,1}, \ldots, b_{1,m}, b_{2,1}, \ldots, b_{2,m}, \ldots, b_{n,1}, \ldots, b_{n,m})' \endgroup $. • $R_d \equiv (\lambda_1^\star I_d, \lambda_2^\star I_d, \ldots, \lambda_{d_V}^\star I_d) = {\lambda^\star}' \otimes I_d$, where $\otimes$ denotes Kronecker product, and $I_d$ is the $d\times d$ identity matrix. • All the terms above are evaluated at the optimal solution, e.g., $\frac{\partial^2L}{\partial \theta\partial \theta'} \equiv \left. \frac{\partial^2L}{\partial \theta\partial \theta'} \right|_{(\theta, V) = (\theta^\star, V^\star)}$. \end{itemize}

The proof of Proposition (ref) can be found in the appendix. Note that the quantities required to compute the sensitivity measures are either already computed in the model estimation procedure or are fast to compute. Other quantities that are not immediately available can be obtained analytically or numerically without model re-estimation. If one wishes to compute the coefficient terms by numerical derivatives, model re-estimation is not required because the derivatives are evaluated around the optimal solution for the fixed value of $\gamma$.

As already mentioned in the introduction, Proposition (ref) can be applied to other economic problems with a constrained optimization structure and unique optimum. I show in Section (ref) that Proposition (ref) nests local sensitivity analysis for unconstrained optimization problems. Thus, the framework in this section is not specific to the running example of DDC models, the T1EV assumption, or the discount factor.

Comparing and reporting the local sensitivity measures have two benefits. First, it can be used to compare the sensitivity of the results with respect to the fixed parameters. Researchers can use this to find the fixed parameters that affect the results the most or determine which of the main results are more sensitive to the fixed parameters. Thus, this can also serve as a guide for more extensive sensitivity analysis.

Second, it can be used as a local approximation of the target parameter at another value of the fixed parameter. The empirical applications in Section (ref) conduct sensitivity analysis and examine the performance of local approximation through two empirical applications.

The following example demonstrates how the quantities in Proposition (ref) can be computed in the context of the rust1987ecta model.

egIn this example, I show the analytical expressions for the terms in the rust1987ecta model that are relevant for the linear system in Proposition (ref). The following derivatives have to be evaluated at the optimal solution with the pre-specified discount factor. I follow the notations introduced in Section (ref) and assume the cost function is given by $c(x, \text{MC}) = -\text{MC}x$. In addition, let $x, y, z \in \cX$ and denote $p(x) \equiv \bP[a_{it} = 1 | x_{it} = x]$ for all $x \in \cX$. The likelihood function can be written as \[ L = \sum^M_{i=1} \sum^T_{t=1} \{ a_{it} \log p(x_{it}) + (1 - a_{it}) \log [1 - p(x_{it})] \}. \] The relevant second derivatives for the likelihood function are as follows: \begin{align*} \frac{\partial^2 L}{\partial \theta \partial \theta'} & = -\sum^M_{i=1} \sum^T_{t=1} p(x_{it})[1 - p(x_{it})] \begin{pmatrix} x_{it}^2 & -x_{it} \\ -x_{it} & 1 \end{pmatrix}, \\ \frac{\partial^2 L}{\partial \theta \partial V(y)} & = -\sum^M_{i=1} \sum^T_{t=1} \beta p(x_{it})[1 - p(x_{it})] [q(y|x_{it}, 0) - q(y|x_{it}, 1)] \begin{pmatrix} -x_{it} \\ 1 \end{pmatrix}, \\ \frac{\partial^2 L}{\partial \theta \partial \beta} & = -\sum^M_{i=1} \sum^T_{t=1} \beta p(x_{it})[1 - p(x_{it})] [Q_0(x_{it}) - Q_1(x_{it})]'V \begin{pmatrix} -x_{it} \\ 1 \end{pmatrix}, \\ \frac{\partial^2 L}{\partial V(y) \partial V(z)} & = -\sum^M_{i=1} \sum^T_{t=1} \beta^2 p(x_{it}) [1-p(x_{it})] [q(y|x_{it}, 0) - q(y|x_{it}, 1)] [q(z|x_{it}, 0) - q(z|x_{it}, 1)] , \\ \frac{\partial^2 L}{\partial V(y) \partial \beta} & = -\sum^M_{i=1}\sum^T_{t=1} [q(y|x_{it}, 0) - q(y|x_{it}, 1)] \left\{ [a_{it} - p(x_{it})] + \beta^2 p(x_{it})[1-p(x_{it})] \right\}. \end{align*} The relevant first derivatives for the fixed-point equation are as follows: \begin{align*} \frac{\partial F(x)}{\partial \theta} & = [1 - p(x)] \begin{pmatrix} -x \\ 1 \end{pmatrix}, \\ \frac{\partial F(x)}{\partial V(y)} & = \beta \{ q(y|x, 0)[1 - p(x)] + q(y|x, 1) p(x) \},\\ \frac{\partial F(x)}{\partial \beta} & = Q_0(x)'V [1 - p(x)] + Q_1(x)'V p(x). \end{align*} The relevant second derivatives for the fixed-point equation are as follows: \begin{align*} \frac{\partial^2 F(x)}{\partial \theta \partial \theta'} & = p(x)[1 - p(x)] \begin{pmatrix} x^2 & -x \\ -x & 1 \end{pmatrix}, \\ \frac{\partial^2 F(x)}{\partial \theta \partial V(y)} & = \beta [q(y|x, 0) - q(y|x, 1)] p(x)[1 - p(x)] \begin{pmatrix} -x \\ 1 \end{pmatrix}, \\ \frac{\partial^2 F(x)}{\partial \theta \partial \beta} & = p(x)[1 - p(x)] [Q_0(x) - Q_1(x)]'V \begin{pmatrix} -x \\ 1 \end{pmatrix}, \\ \frac{\partial^2 F(x)}{\partial V(y) \partial V(z)} & = \beta^2 [q(y|x, 0) - q(y|x, 1)][q(z|x, 0) - q(z|x, 1)] p(x) [1 - p(x)],\\ \frac{\partial^2 F(x)}{\partial V(y) \partial \beta} & = q(y|x, 0)[1 - p(x)] + q(y|x, 1)p(x) \\ & \quad + \beta [q(y|x, 0) - q(y|x, 1)] p(x)[1 - p(x)] [Q_0(x) - Q_1(x)]'V. \end{align*}

Sensitivity of counterfactuals

In practice, researchers typically first estimate the parameters and then perform counterfactual analysis. Some common types of counterfactuals include changing the utility parameters (e.g., college subsidy program in keanewolpin1997jpe), changing the transition probabilities (e.g., demand volatility in collardwexler2013ecta), and changing the action and state space (e.g., eliminating patients' actions in crawfordshum2005ecta). See kalouptsidietal2021qe for a detailed discussion on various types of counterfactuals and some identification results for DDC models.

The local sensitivity of the above counterfactuals can be computed based on the estimates from Proposition (ref). In particular, $\frac{\partial \theta}{\partial \gamma'}$ have already been computed. Then, the sensitivity for the counterfactual parameters $\widetilde \theta$ and the counterfactual value function $\widetilde V$ can be obtained as follows.

enumerate• Suppose the counterfactual changes the utility parameters to $\widetilde\theta \equiv H(\theta)$. Then, the sensitivity of the utility parameters under the counterfactual can be computed as \begin{equation} \frac{\partial \widetilde\theta}{\partial \gamma'} = \frac{\partial H}{\partial \theta'} \frac{\partial \theta}{\partial \gamma'}. \end{equation} The elasticity or semi-elasticity measures similar to the ones introduced at the beginning of Section (ref) can also be easily computed for the counterfactuals. They can be computed by replacing $\widehat\theta$ with $\widetilde\theta$ in (ref) and (ref), and using (ref). • For other types of counterfactuals where the utility parameters are unchanged, $\frac{\partial \widetilde\theta}{\partial \gamma'} = \frac{\partial \theta}{\partial \gamma'}$ holds automatically. • Let $\widetilde V = \widetilde F(\widetilde \theta, \widetilde V; \gamma)$ be the updated equilibrium constraints under the counterfactuals. Depending on the counterfactual, there may be a different number of equations. Thus, the sensitivity of $\widetilde V$ can be computed by substituting $\frac{\partial \widetilde\theta}{\partial \gamma'}$, and solving the following system of equations: \[ \frac{\partial \widetilde F}{\partial \theta'} \frac{\partial \widetilde\theta}{\partial \gamma'} + \left( \frac{\partial \widetilde F}{\partial V'} - I_{\widetilde d_V}\right)\frac{\partial \widetilde V}{\partial \gamma'} = - \frac{\partial \widetilde F}{\partial \gamma'}. \]

Other economic applications

While this paper focuses on DDC models, many other economic applications also have a constrained optimization structure as in equation (ref) with some parameters being fixed in the estimation procedure. Some details of the examples mentioned in Section (ref) are as follows.

ossa2014aer studies the effect of tariffs on welfare. The main optimization problem maximizes welfare subject to equilibrium constraints, such as budget constraints and market clearing conditions. Unlike the other applications, the objective is linear in the variables. The calibrated parameter is the elasticity of substitution of industry varieties $\sigma_s$. Sensitivity analysis with respect to calibrated parameters is conducted by resolving the optimization problem with different values of $\sigma_s$.

igami2017jpe studies creative destruction in the hard disk drive industry using a dynamic discrete game model. The model is solved using the nested fixed-point approach. Apart from the discount factor, the rate of change of innovation costs and the number of potential entrants are also fixed in the model. Sensitivity analysis is performed by estimating the full model again at different values of the three fixed parameters.

yang2021aejma studies aggregate productivity losses due to misallocation. The main optimization problem maximizes likelihood subject to firm optimality conditions. There are three calibrated parameters. They are the span of control, sectoral capital share, and conversion factor. Again, sensitivity analysis with respect to calibrated parameters is conducted by resolving the optimization problem at different calibrated parameters.

chenchoo2022ej study dynamic matching problems. Similar to dynamic discrete choice problems, the main optimization problem maximizes likelihood subject to fixed-point constraints. There are additional constraints that characterize matching equilibria. The discount factor is fixed at 0.95.

Connection with unconstrained optimization problems

In this section, I explain how Proposition (ref) is related to unconstrained optimization problems with fixed parameters. This is relevant for solution methods such as hotzmiller1993restud and aguirregabiriamira2002ecta introduced in Section (ref).

Write $V = \overline F(\theta; \gamma)$, so that $V$ can be expressed as a known function of $\theta$ and $\gamma$. As a result, substituting $V$ into the criterion function yields \[ L(\theta, V; \gamma) = L(\theta, \overline F(\theta; \gamma); \gamma) \equiv \overline L(\theta; \gamma). \] Hence, optimization problem (ref) becomes

align[align omitted — 95 chars of source]

Here, the parameter of interest in problem (ref) is $\theta$ and the relevant local sensitivity measure is $\frac{\partial \theta}{\partial \gamma'}$. The following proposition shows how the local sensitivity measure can be computed and the connection with Proposition (ref).

propConsider the optimization problem (ref). Assume that \begin{itemize} • $\overline L$ is continuously differentiable in $\theta$ and $\gamma$ around $\widehat{\theta}(\gamma)$. • $\widehat{\theta}(\gamma)$ is the unique solution to the optimization problem (ref) and belongs to the interior of $\Theta$ for each $\gamma \in \Gamma$. \end{itemize} Then, the following statements hold: \begin{enumerate} • If $\frac{\partial^2 \overline L}{\partial \theta \partial \theta'}$ is invertible, the local sensitivity measure $\frac{\partial \theta}{\partial \gamma'}$ can be obtained by solving \begin{equation} \frac{\partial^2 \overline L}{\partial \theta \partial \theta'} \frac{\partial \theta}{\partial \gamma'} = -\frac{\partial^2 \overline L}{\partial \theta \partial \gamma'}. \end{equation} • In terms of the notations in Proposition (ref), \begin{align*} \frac{\partial^2 \overline L}{\partial \theta \partial \theta'} & = A_{\theta, \theta'} + \frac{\partial^2 L}{\partial \theta \partial V'} \frac{\partial F}{\partial \theta'} + \frac{\partial F}{\partial \theta'}\frac{\partial^2 L}{\partial V\partial \theta'} + \frac{\partial F}{\partial \theta'} \frac{\partial^2 L}{\partial V \partial V'} \frac{\partial F}{\partial \theta'} , \\ \frac{\partial^2 \overline L}{\partial \theta \partial \gamma'} & = A_{\theta, \gamma'} + \frac{\partial^2 L}{\partial \theta \partial V'} \frac{\partial F}{\partial \gamma'} + \frac{\partial F}{\partial \theta'}\frac{\partial^2 L}{\partial V\partial \gamma'} + \frac{\partial F}{\partial \theta'} \frac{\partial^2 L}{\partial V \partial V'} \frac{\partial F}{\partial \gamma'}. \end{align*} \end{enumerate} All the terms above are evaluated at the optimal solution, e.g., $\frac{\partial^2\overline L}{\partial \theta\partial \theta'} \equiv \left. \frac{\partial^2 \overline L}{\partial \theta\partial \theta'} \right|_{\theta = \theta^\star}$.

The assumptions of Proposition (ref) are similar to Assumption (ref) but for unconstrained optimization problems. Part (ref) of Proposition (ref) follows immediately from total differentiating the first-order conditions of optimization problem (ref). Part (ref) of Proposition (ref) shows that the local sensitivity measure for the unconstrained problem (ref) can be expressed using the terms in Proposition (ref) for constrained optimization problem (ref). This shows Proposition (ref) nests unconstrained optimization problems as a special case.

rejorgensen2023restat focuses on the following unconstrained optimization problem with a GMM objective: \begin{equation} \min_{\theta \in \Theta} \ \ g_n(\theta; \gamma)'W_n g_n(\theta; \gamma), \end{equation} where $g_n(\theta; \gamma)$ is some vector-valued functions and $W_n$ is a symmetric positive definite weighting matrix. The unconstrained optimization problem (ref) can be written in terms of optimization problem (ref) based on the preceding discussion. Let $\overline L(\theta; \gamma)$ be the objective function in equation (ref). Since $\frac{\partial \overline L}{\partial \theta} = \left[\frac{\partial g_n(\theta; \gamma)}{\partial \theta'}\right]' W_n g_n(\theta; \gamma) $, the expressions in equation (ref) can be written as \begin{align*} \frac{\partial^2 \overline L}{\partial \theta \partial \theta'} & = \left[\frac{\partial g_n(\theta; \gamma)}{\partial \theta} \right]' W_n \frac{\partial g_n(\theta; \gamma)}{\partial \theta'} + \left[ g_n(\theta; \gamma)' W_n \otimes I_{d_{\theta}}\right] \frac{\partial \mathrm{vec}\{[\frac{\partial g_n(\theta; \gamma)}{\partial \theta} ]' \}}{\partial \theta'}, \\ \frac{\partial^2 \overline L}{\partial \theta \partial \gamma'} & = \left[\frac{\partial g_n(\theta; \gamma)}{\partial \theta} \right]' W_n \frac{\partial g_n(\theta; \gamma)}{\partial \gamma'} + \left[ g_n(\theta; \gamma)' W_n \otimes I_{d_{\theta}}\right] \frac{\partial \mathrm{vec}\{[\frac{\partial g_n(\theta; \gamma)}{\partial \theta} ]' \}}{\partial \gamma'}, \end{align*} and they can be evaluated at the optimal solution of (ref) for a given $\gamma$. The above terms match the expressions in Proposition 1 of jorgensen2023restat. Thus, the local sensitivity measure in Proposition (ref) also nests the unconstrained version in jorgensen2023restat. As discussed in Section (ref), it is common for economic models to be estimated by constrained optimization with equilibrium constraints. The objective function is not necessarily a GMM objective function in empirical applications, as in ossa2014aer.

Global sensitivity analysis

Section (ref) has provided a local sensitivity measure that is easy to compute and can be used as a good local approximation of the target parameter around its optimal value at the given fixed parameter. As mentioned in the introduction, researchers usually conduct sensitivity analysis by re-estimating the model at a few other values of the fixed parameter. Even though monotonic patterns are usually shown, they do not necessarily extend to the entire support of the fixed parameters.

Since the discount factor is the most commonly fixed parameter in DDC models, I start by studying what conditions on model primitives can imply that the flow utility is monotone in the discount factor in DDC models. Then, I focus on the linear-in-parameter utility specification that is common in practice. If the parameters are monotone in the discount factor, it is sufficient to estimate the model at the endpoints of the discount factor to obtain the bounds on the parameter. For more general models, the target parameter may not necessarily be monotone in the fixed parameter. As a result, I propose a constrained optimization approach to compute the bounds of the target parameter over a range of fixed parameters. In the context of DDC models and discount factors, the range of discount factors may be obtained by exclusion restrictions. See magnacthesmar2002ecta and abbringdaljord2020qe for exclusion restrictions and kongetal2022wp for a recent application that estimates the discount factor for various consumer goods.

Monotonicity of utility

In this section, I start from the representation in equation (ref) that connects flow utility and CCP. This is motivated by two-step approaches that estimate CCP in the first step and the utility parameters in the second step.

Assume that the transition matrices $Q_a$ and CCP $p_a$ are available to the researchers for all $a \in \cA$ in equation (ref). In addition, consider the following normalization assumption that is standard in many applications.

assu$\pi_A(x) = 0$ for any $x \in \cX$.

Under Assumptions (ref) and (ref), the utility vector at state $a \in \cA \backslash \{A\}$ is given by

align[align omitted — 150 chars of source]

The following proposition shows the sign of the flow utility with respect to the discount factor.

propLet Assumptions (ref) to (ref) and (ref) hold. The derivative of $\pi_a$ in equation (ref) with respect to $\beta$ is \[ \frac{\partial \pi_a}{\partial \beta} =-[Q_a(I_X - \beta Q_A)^{-1} - (I_X - \beta Q_A)^{-1} Q_a] (I_X - \beta Q_A)^{-1} (-\log p_A), \] for any $a \in \cA \backslash \{A\}$.

The following corollary shows that normalizing the utility to another value gives a similar result.

corLet Assumptions (ref) to (ref) hold and normalize the utility for choice $A$ as $\pi_A = \overline \pi_A$. Then, \[ \frac{\partial \pi_a}{\partial \beta} = -[Q_a(I_X - \beta Q_A)^{-1} - (I_X - \beta Q_A)^{-1}Q_A ] (I_X - \beta Q_A)^{-1} (\overline \pi_A - \log p_A) , \] for any $a \in \cA \backslash \{A\}$.

The above corollary shows that a normalization that is different from the one in Assumption (ref) only replaces $(-\log p_A)$ in Proposition (ref) by $(\overline \pi_A - \log p_A)$ in Corollary (ref). Therefore, the remaining results in this section will maintain Assumption (ref).

Proposition (ref) can be simplified under $\rho$-period finite dependence between actions $a$ and $A$. For examples and discussion on finite dependence, see arcidiaconomiller2011ecta, arcidiaconomiller2019qe and references therein.

propLet Assumptions (ref) to (ref) and (ref) hold. Suppose there exists $a \in \cA \backslash \{A\}$ and $\rho \in \bN$ such that $Q_a(x) Q_A^\rho = Q_A(x)Q_A^\rho$ for all $x \in \cX$. Then, \begin{equation} \frac{\partial \pi_a}{\partial \beta} = -(Q_a- Q_A)(I_X + \beta Q_A + \cdots + \beta^{\rho - 1} Q_A^{\rho - 1})(I_X - \beta Q_A)^{-1} (-\log p_A). \end{equation}

The following corollary shows that if $\rho = 1$, the RHS of equation (ref) does not depend on $\beta$. As a result, the slope of flow utility for each state is constant across $\beta \in [0, 1)$.

corLet Assumptions (ref) to (ref) and (ref) hold. Suppose that there is one-period dependence between choices $a \in \cA \backslash \{A\}$ and $A$ for all states $x\in\cX$. Then, \[ \frac{\partial \pi_a}{\partial \beta} = (-Q_a + Q_A)(-\log p_A), \] for any $a \in \cA \backslash \{A\}$. As a result, for each $a \in \cA \backslash \{A\}$ and $x \in \cX$, $\pi_a(x)$ is either increasing, decreasing, or constant across in $\beta \in [0, 1)$.

One-period finite dependence is common in economics. This holds in models where one of the actions is a renewal or terminal action. A renewal action is an action that resets the states, such as the option to replace the engine in rust1987ecta's bus engine model. A terminal action is a choice where the optimization problem is ended with no more future actions, such as pakesetal2007rand's firm exit model. The following corollary summarizes the conditions for renewal actions.

corLet Assumptions (ref) to (ref) and (ref) hold. Suppose that $p_A(1) > 0$ and action $A$ is a renewal action so that the corresponding transition matrix $Q_A$ is given by \[ Q_A = \begin{pmatrix} 1 & 0 & \cdots & 0 \\ 1 & 0 & \cdots & 0 \\ \vdots & \vdots & \ddots & \vdots \\ 1 & 0 & \cdots & 0 \\ \end{pmatrix}. \] \begin{enumerate} • If $p_A(1) \leq p_A(x)$ for any $x \in \cX $, then $\pi_a$ is nondecreasing in $\beta$ for any $a \in \cA \backslash \{A\}$. • If $p_A(1) \geq p_A(x)$ for any $x \in \cX $, then $\pi_a$ is nonincreasing in $\beta$ for any $a \in \cA \backslash \{A\}$. \end{enumerate}

The condition in Corollary (ref) is easy to verify because it only depends on the CCP of action $A$ and the presence of a renewal action.

While Corollaries (ref) and (ref) provide some convenient conditions under which the utility is monotone in $\beta$ under some special conditions, it is important to note that the utilities can still be monotone in the discount factor whenever the transition probabilities and the CCP are such that the derivative in Proposition (ref) is always positive or negative, even if the conditions in Corollaries (ref) and (ref) do not hold.

Although it is possible to obtain the monotonicity of utility under some additional conditions on the model primitives, characterizing the monotonicity for counterfactuals is more challenging due to the nonlinearity of the model. Let $\widetilde V$ be the vector of counterfactual ex ante value functions. Let $\widetilde p_a$ be the counterfactual CCP, $\widetilde Q_a$ be the counterfactual transition matrix, $\widetilde \pi_a$ be the counterfactual utility at action $a \in \cA$. Assume that $\widetilde \pi_a$ is related to $\pi_a$ through the following affine transformation \[ \widetilde \pi_a = H_a \pi_a + g_a, \] for some $X \times X$ matrix $H_a$ and length $X$ vector $g_a$ that are independent of $\beta$. If $\widetilde \pi_A = 0$, then $\widetilde \pi_a$ can be written as

align[align omitted — 182 chars of source]

for any $a \in \cA \backslash \{A\}$ using (ref). Taking the derivative with respect to $\beta$ on both sides yields

align[align omitted — 519 chars of source]

for any $a \in \cA \backslash \{A\}$, where $\odot$ denotes Hadamard product (entrywise product).

egConsider the rust1987ecta model again, so that $\cA = \{0,1\}$, $A = 1$, and $\widetilde p_1(x) = 1 - \widetilde p_0(x)$ for each $x \in \cX$. Then, equation (ref) can be used to characterize the relationship between counterfactual CCP as follows \begin{align*} H_0 \pi_0 + g_0 = [ I_X + \beta (\widetilde Q_1 - \widetilde Q_0)](-\log \widetilde p_1) + \log \widetilde p_0. \end{align*} Thus, equation (ref) can be written as \[ H_0 \frac{\partial \pi_0}{\partial \beta} = (\widetilde Q_1 - \widetilde Q_0)(-\log \widetilde p_1) - [I_X + \beta(\widetilde Q_1 - \widetilde Q_0)] \left( \frac{1}{\widetilde p_1} \odot \frac{\partial \widetilde p_1}{\partial \beta} \right) - \frac{1}{\widetilde p_0} \odot \frac{\partial \widetilde p_1}{\partial \beta}. \] where $\widetilde{p}_0 \equiv 1_X - \widetilde{p}_1$.

The above shows the relationship between the derivatives of counterfactuals with respect to the discount factor in DDC models is nonlinear and may not be easy to characterize.

reAlthough the above results maintain the T1EV assumption on the unobservables due to Assumption (ref), the same analysis can be applied to other distributions of unobservables. This is because the mapping $\psi_a(p)$ in equation (ref) depends on the distribution of the unobservables and the CCP $p$. The distribution of unobservables is independent of the discount factor, and the CCP is assumed to be known. Hence, $\psi_a(p)$ does not depend on the discount factor. As a result, the results can still be applied by replacing the vector $(-\log p_A)$ by $\psi_A(p)$ for other distributions of the unobservables.

Monotonicity with linear-in-parameters utility

In this section, I connect the results in the previous section to the linear-in-parameters specification of the utility function that is common in empirical applications.

assu[Linear-in-parameters] For each $a \in \cA \backslash \{A\}$, the flow utility at state $x \in \cX$ is given by \[ \pi_a(x) = \varpi_a(x)'\theta, \] where $\theta \in \bR^{d_{\theta}}$ and $\varpi_a(x) \equiv (\varpi_{a,1}(x), \ldots, \varpi_{a,d_{\theta}}(x))'$ are the corresponding coefficients on $\theta$.

Let $\Pi_a$ be the $(X \times d_{\theta})$ matrix that collects the coefficients on $\theta$ over all the states at $a \in \cA \backslash \{A\}$, so the $x$-th row represents $\varpi_a(x)'$ for each $a \in \cA \backslash \{A\}$. Then, the vector of utilities for each $a \in \cA \backslash\{A\}$ can be written as

equation[equation omitted — 59 chars of source]

based on Assumption (ref).

Let $\widehat p_a$ be the estimator of the CCP for action $a \in \cA \backslash \{A\}$. Then, $\widehat \pi_a$ can be estimated via equation (ref). Let $\widehat \pi$ and $\Pi$ be the matrices that stack $\widehat \pi_a$ and $\Pi_a$ over all $a \in \cA \backslash \{A\}$. Then, the utility parameters can be obtained by the following minimum distance problem

equation[equation omitted — 127 chars of source]

where $W$ is a positive definite weighting matrix independent of the discount factor.

The following proposition shows the sensitivity of utility parameters estimated from optimization problem (ref) with respect to the discount factor.

propLet Assumptions (ref) to (ref), (ref) and (ref) hold. Let $\widehat \theta$ be the parameter estimated from problem (ref) and assume that the matrix $\Pi' W\Pi$ is invertible. Then, the derivative of $\widehat \theta$ with respect to $\beta$ is given by \[ \frac{\partial \widehat\theta}{\partial \beta}= (\Pi'W\Pi)^{-1}\Pi'W \frac{\partial \widehat \pi}{\partial \beta}. \]

Proposition (ref) shows that the sensitivity of the utility parameters with respect to $\beta$ depends on $\Pi$ and $W$, in addition to $\frac{\partial \widehat\pi}{\partial \beta}$. Hence, even if $\widehat \pi$ is monotone in $\beta$, the choice of the parameterization and the weighting matrix can also affect whether the utility parameters are monotone in $\beta$. The following corollary shows a special case where the weighting matrix is an identity matrix and when the same parameter does not appear in more than one choice.

corLet Assumptions (ref) to (ref), (ref) and (ref) hold. Suppose that \begin{enumerate} • $\theta$ is partitioned as $\theta = (\theta_0', \ldots, \theta_{A-1}')'$ with $\theta_a \in \bR^{d_a}$ for each $a \in \cA \backslash \{\cA\}$ and $d_{\theta} = \sum_{a\in\cA \backslash \{A\}} d_a$. • $\pi_a = \Pi_a\theta_a$ and $\Pi_a$ has full rank. • $W$ is an identity matrix. \end{enumerate} Then, \[ \frac{\partial \widehat \theta_a}{\partial \beta} = (\Pi_a' \Pi_a)^{-1} \Pi_a' \frac{\partial \widehat \pi_a}{\partial \beta}, \] for each $a \in \cA \backslash \{A\}$.

The above discussion has focused exclusively on the global sensitivity analysis for the discount factor. In practice, researchers may fix some other parameters in the utility function in the estimation procedure. An example is the rate of change of innovation cost that igami2017jpe calibrates in the coefficient for the sunk cost parameter that he estimates. Let $\delta \in \bR$ be another parameter that the researcher calibrates apart from the discount factor. The following proposition shows the sensitivity of the utility parameters with respect to $\delta$ estimated from problem (ref) when part of $\Pi$ contains $\delta$.

propLet Assumptions (ref) to (ref), (ref) and (ref) hold. Suppose that part of the matrix $\Pi$ contains fixed parameter $\delta \in \bR$, the matrix $\Pi'W\Pi$ is invertible, and that $\widehat \theta$ is estimated from problem (ref). Then, the derivative of $\widehat \theta$ with respect to $\delta$ is given by \[ \frac{\partial \widehat \theta}{\partial \delta} = -(\Pi'W\Pi)^{-1} \left[ \Pi' (W' + W) \frac{\partial \Pi}{\partial\delta} \widehat \theta - \left(\frac{\partial \Pi}{\partial \delta}\right)'W\widehat \pi \right]. \]

Whether monotonicity can be established for other fixed parameters depends on the exact structure of the parameterization. This motivates the estimation approach for global sensitivity analysis in the next section.

Other approaches

Researchers can still estimate the bounds on the target parameter over a certain interval of the fixed parameter through a constrained optimization problem. Following Assumption (ref) that the solution is in the interior and that the optimal solution is a regular point, the first-order condition is satisfied at the optimum. Let $\tau(\theta, V; \gamma)$ be the target parameter the researcher wishes to estimate. Then, the bounds on the target parameter over a pre-specified range of fixed parameters $\gamma \in \overline \Gamma \subseteq \Gamma$ can be estimated by the following optimization problem that estimates the bounds subject to the first-order conditions

align[align omitted — 481 chars of source]

If the researcher is interested in estimating the bounds on the target parameter that depends on the counterfactuals with $\widetilde \theta = H(\theta)$, then optimization problem (ref) can be updated as

align*[align* omitted — 556 chars of source]

The computational time by treating $\gamma$ as a variable depends on $\overline{\Gamma}$ and the structure of the problem. The next section explores this using the rust1987ecta model.

egIn terms of the rust1987ecta model, suppose the researcher is interested in the bounds on the target parameter $\tau(\theta, V; \beta)$ over $\beta \in [\underline{\beta}, \overline{\beta}] \equiv \cB$. Then, optimization problem (ref) can be implemented as follows: \begin{align*} \operatornamewithlimits{min/max}_{(MC, RC) \in \Theta, V \in \cV, \beta \in \cB, \lambda \in \bR^{d_V}}\quad & \tau(\theta, V; \beta)\\ s.t. \quad & -\sum^M_{i=1} \sum^T_{t=1} [a_{it} - p(x_{it})] x_{it} - \sum^{d_V}_{j=1} \lambda_j[1 - p(j)](- j )= 0 \\ & - \sum^M_{i=1} \sum^T_{t=1} [a_{it} - p(x_{it})] - \sum^{d_V}_{j=1} \lambda_j[1 - p(j)] = 0 \\ & - \beta \sum^M_{i=1} \sum^T_{t=1} [a_{it} - p(x_{it})] [q(y|x_{it},0) - q(y|x_{it},1)] \\ & \quad + \lambda_y - \beta \sum^{d_V}_{j=1} \lambda_j \{ q(y|j, 0)[1 - p(j)] + q(y|j, 1) p(j) \} \tag*{for all $y \in \cX$} \\ & p(x) = \frac{1}{1 + \exp\{RC - MCx + \beta(Q_0(x)-Q_1(x))' V\}} \tag*{for all $x \in \cX$} \\ & V(x) = \log \left\{ \exp[RC - \text{MC}x + \beta Q_0(x)'V] + \exp[\beta Q_1(x)'V] \right\} \tag*{for all $x \in \cX$} \end{align*} The second and third equations above correspond to the first-order conditions with respect to the two utility parameters. The fourth equation corresponds to the first-order condition with respect to the $x$-th component of the value function. The fifth equation is the equation of the choice probability (ref). The last equation corresponds to the fixed-point equation for the Bellman equation as in (ref). The above problem can be implemented in standard software such as Knitro. Researchers can also easily incorporate multi-start in Knitro to optimize the bounds to search for the global optimum.

Another avenue for conducting global sensitivity analysis is to conduct breakdown analysis horowitzmanski1995ecta, klinesantos2013qe, mastenpoirier2020qe. In terms of the fixed parameters in DDC models, breakdown analysis can be used to find the range of fixed parameters such that a certain conclusion holds. More precisely, suppose that the researcher is interested in knowing whether the target parameter is above a threshold, i.e., $\widehat\tau(\gamma) \geq \tau^\star$. Then, the robust region (RR) is the region such that the conclusion holds and is defined as \[ \text{RR} \equiv \{ \gamma \in \Gamma : \widehat\tau(\gamma) \geq \tau^\star \}, \] and the breakdown frontier is the boundary of the robust region. If one further restricts attention to the discount factor so that $\gamma = \beta$, and that $\widehat \tau(\beta)$ is monotone in $\beta$, the breakdown frontier can be found by the bisection method.

Empirical applications

In this section, I perform sensitivity analysis of two empirical applications with respect to the discount factor. The first application is the seminal bus engine replacement model in rust1987ecta. The second application is a recent dynamic marriage matching model from chenchoo2022ej.

Application 1: Rust (1987)

In this section, I conduct local and global sensitivity analysis for the rust1987ecta model using the data from the group 4 buses. I assume that the utility function is the same as in Section (ref), and I specify the cost function as $c(x, \text{MC}) = -\text{MC} x$.

Local analysis

This section conducts local sensitivity analysis of various parameters with respect to the discount factor. All estimates in this section are obtained by the NFXP. The parameters I consider are as follows:

enumerate• Replacement cost. • Maintenance cost. • Counterfactual conditional choice probability at state 90. • Change in average welfare as defined by $W \equiv \frac{1}{X} \sum^X_{x=1} [\widetilde V(x) - V(x)]$.

For parameters (ref) and (ref) above, the counterfactual I consider is a reduction in maintenance cost of 10%. To analyze the sensitivity of the parameters at different values of the discount factor, I estimate each parameter using the following discount factors $\beta$: 0.85, 0.9, 0.95, 0.99, 0.999, and 0.9999. rust1987ecta uses $\beta = 0.9999$ in the main results. Then, I estimate the local sensitivity measure using the elasticity measure as in (ref) so that the sensitivity measure can be comparable across different target parameters. Using the point estimate and the local sensitivity measure, I approximate the parameter at another discount factor and report the percentage approximation error. The results are reported in Table (ref).

table[table omitted — 3,237 chars of source]

The local sensitivity measure in Table (ref) shows that not all parameters are equally sensitive to the choice of the discount factor. On the other hand, the magnitude of the elasticity is increasing in the discount factor for all four parameters. The last three columns of Table (ref) report the performance when the estimate and the local sensitivity measure obtained at $\beta$ is used for local approximation for the target parameter when the discount factor is changed to $\beta ' = \beta - \Delta\beta$. Except for the change in average welfare when the discount factor is close to 1, local approximation gives a less than 1% approximation error in most cases. An intuition for the larger error in the change in average welfare is that the value function can be interpreted as an infinite sum of the discounted present payoffs.

Table (ref) suggests that researchers should be cautious with the sensitivity of the parameters when the discount factor is large. The discount factor is typically higher when the time between two consecutive periods is lower, or when agents discount the future less. Nevertheless, unless the discount factor is very close to 1, local approximation provides a good approximation of the model parameters without the need to re-estimate the full model.

Global analysis

In this section, I study the global properties of various target parameters. Figure (ref) shows the estimates of the maintenance and replacement cost against the discount factor using NFXP and the minimum distance two-step method as in equation (ref). The two-step method in equation (ref) requires the CCP to be estimated by data. If the data is rich enough, then the CCP can be estimated by the simple frequency estimator. Since not all 90 states can be observed in the data, I estimate the CCP using a simple logistic regression. I regress the observed action on the state and the square of the state. The lines in Figure (ref) are obtained by estimating the utility parameters for each discount factor in the interval $[0, 0.9999]$ with grid points of size 0.0001. The maintenance cost is strictly decreasing in the discount factor, while the replacement cost is strictly increasing in the discount factor.

figure[figure omitted — 176 chars of source]

Figure (ref) shows that the estimated CCP satisfies the conditions of Corollary (ref). On the other hand, the option to replace the bus engine is a renewal action. Hence, the flow utility is monotonically increasing in the discount factor. With the cost function as specified at the beginning of this section, let $\theta = (\text{MC}, \text{RC})$. Then, the matrix of utility coefficients can be written as \[ \Pi_0 =

pmatrix[pmatrix omitted — 72 chars of source]

. \] Applying Corollary (ref), it follows that $\frac{\partial \widetilde{\text{MC}}}{\partial \beta} < 0$ and $\frac{\partial \widetilde{\text{RC}}}{\partial \beta} > 0$.

figure[figure omitted — 164 chars of source]

Although the utility parameters are monotone in the discount factor, the counterfactuals may not necessarily be monotone in the discount factor. Consider again the counterfactual of reducing the maintenance cost by 10%. Figure (ref) shows the counterfactual CCP at three different states. It can be seen that the probabilities can exhibit different properties against the discount factor in Figure (ref). Indeed, this can be seen by total differentiating the counterfactual CCP at state $x$, denoted as $\widetilde{p}(x)$, with respect to the discount factor as follows

equation[equation omitted — 474 chars of source]

for each $x \in \cX$. In equation (ref), (a) is positive. However, (b) and (c) have opposite signs in this model and data. As a result, $\frac{\partial \widetilde p(x)}{\partial \beta}$ is an average of positive and negative terms. Note that $\beta$ also enters the expression in (ref) explicitly. Hence, the CCP is not necessarily monotone in $\beta$. Figure (ref) demonstrates the counterfactual CCPs at three different states. Although the counterfactual CCPs demonstrate different shapes, their range is rather small.

figure[figure omitted — 193 chars of source]

Finally, I explore the performance of using the constrained optimization problem in Example (ref) to find the bounds on the target parameter over a range of discount factors. Table (ref) shows the bounds on the utility parameters and counterfactuals over different intervals of $\beta$. For all rows, I set the lower bound of the interval of $\beta$ as 0.7 with different upper bounds. The computational time is the total time in running the optimization problem for the lower and upper bounds using Knitro. An attractive feature of the constrained optimization approach is that the optimizer would search for the optimal values so researchers do not need to choose the grid points over the support of the discount factor and re-estimate the model at each of the points.

table[table omitted — 799 chars of source]

Application 2: Chen and Choo (2023)

chenchoo2022ej develop a dynamic marriage matching framework that extends the methodology in choo2015ecta. The empirical example in chenchoo2022ej studies how China's one-child policy affects the marriage distribution. The model has a constrained optimization structure, and they estimate their model using the nested fixed-point algorithm. The discount factor is fixed at $\beta = 0.95$ throughout their analysis. I illustrate that the methodology in Section (ref) can be applied in this dynamic matching application, although it contains several additional components when compared to the DDC model described in Section (ref).

I focus on the analysis in Figure 2 of chenchoo2022ej, where they examine how the types of couples affect marital surplus. In their model, individuals differ by education, previous marital status, and age. The education level can be junior high school (JHS), high school (HS), or college (C). Individuals can be previously married and divorced or not married. They report results of ages from 20 to 45 in their Figure 2.

Table 2 of chenchoo2022ej compares the surplus differences $\Pi_1(a_{\text{M}}, p_{\text{M}}) - \Pi_0(a_{\text{M}}, p_{\text{M}}, a_{\text{F}}, e_{\text{F}})$. $\Pi_1(a_{\text{M}}, p_{\text{M}})$ is the marriage surplus for an age $a_{\text{M}}$ male with education level JHS, and martial status $p_{\text{M}}$ and an age 20 female with education level JHS, and previously not married. $\Pi_0(a_{\text{M}}, p_{\text{M}}, a_{\text{F}}, e_{\text{F}})$ is the marriage surplus for a male of the same type as in $\Pi_1(a_{\text{M}}, p_{\text{M}})$ and an age $a_{\text{F}}$ female with education level $e_{\text{F}}$ and previously married and divorced. The support of the parameters they consider are $a_{\text{F}} \in \{20, 21, \ldots, 45\}$, $a_{\text{M}} \in \{25, 30, 35\}$, $e_{\text{F}} \in \{\text{JHS}, \text{HS}, \text{C}\}$, and $p_{\text{M}} \in \{0, 1\}$ that equals 1 if the male is previously married and divorced. chenchoo2022ej find that couples with similar types generally have a similar surplus.

Figure (ref) shows the sensitivity measure of marital surplus differences as in equation (ref) for different types of individuals, arranged in the same format as in Figure 2 of chenchoo2022ej. The terms required in Proposition (ref) are computed using automatic differentiation. The sensitivity measure can be interpreted as the unit change in martial surplus for a unit change in the discount factor $\beta$. It can be seen that the sensitivity of the parameter estimates is not constant across different specifications and is generally increasing in $a_{\text{F}}$.

figure[figure omitted — 266 chars of source]

Finally, I examine the performance of the sensitivity measure as a local approximation. To do this, I re-estimate the full model at other values of the discount factor and estimate the surplus difference as in Figure 2 of chenchoo2022ej. Then, I approximate the surplus difference at another discount factor using the local sensitivity measure estimated at $\beta = 0.95$. Table (ref) reports the summary statistics on the absolute error when local approximation is used to approximate the surplus differences at other discount factors. In the table, $P_x$ refers to the $x$-th percentile. The magnitudes of the absolute errors remain small across different $\beta$ and are mostly less than 1. On the other hand, Table (ref) reports the summary statistics of the absolute percentage error for the approximation. It is important to note that the large percentage errors are due to the actual marital surplus differences being very close to 0. See Appendix Figure (ref) that shows the large percentage errors are all associated with actual marital surplus differences very close to 0. This exercise shows that the local approximation can be accurate in estimating the surplus without estimating the entire model again at another discount factor.

table[table omitted — 1,478 chars of source]
table[table omitted — 1,457 chars of source]

Conclusion

In this paper, I propose two procedures to conduct sensitivity analysis for parameter estimates with respect to fixed parameters in DDC models. First, the local sensitivity measure reports the change in the target parameter estimated from a constrained optimization problem for a unit change in the fixed parameter. This measure is fast to compute, does not require model re-estimation at another value of the fixed parameter, and nests unconstrained estimation problems. The methodology in this paper can be applied to more general estimation problems, and not necessarily DDC models. For global sensitivity analysis, I examine whether target parameters are monotone in the fixed parameters. Using the discount factor as the leading example of fixed parameters in DDC models, I provide conditions under which utility is monotone in the discount factor.

From the empirical examples, I find that the estimates are typically more sensitive when the discount factor is closer to 1. On the other hand, I also show that the local sensitivity measure can serve as a good local approximation to the target parameter without re-estimating the full model. Hence, researchers can report the local sensitivity measure in addition to the point estimates to increase the transparency of structural research.