Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,875 characters · 11 sections · 62 citation commands
Machine Learning for Dynamic Discrete Choice
In empirical work on dynamic models, economists often make specification choices - for example, discretize state space or select covariates for flow utility - in order to achieve computational tractability and precise estimates (e.g. PakesShankman:1984, Pakes:1986, Rust, Ryan:2012). Typically, economists have little intuition about which covariates to select or how to discretize a continuous state variable (PakesLanjouw:1998). Unfortunately, counterfactual predictions of dynamic models are sensitive to specification choices and are difficult to interpret when a model is misspecified. As discussed in athey2017science, there has been recent interest in data-driven model selection based on modern machine learning tools. Moreover, (orthogStructural, doubleml2016, LRSP, AtheyWager) have shown how to leverage these tools into high-quality estimates of causal parameters. Because dynamic models are more challenging to analyze, the model selection in dynamic models remains an open question.
This paper estimates a dynamic model of imperfect competition with a high-dimensional state space. Consider the bus engine replacement model from Rust as a special case. An agent decides whether to replace the bus engine in each period. The agent incurs a fixed cost in the case of replacement and a cost proportional to the current mileage in the case of maintenance. One can imagine that the future bus mileage depends on a vector of observed exogenous characteristics, such as current traffic and weather conditions, encoded in a high-dimensional vector. While these characteristics do not enter into the per-period utility of the owner, they affect the mileage's law of motion in the case of maintenance decision, and hence are taken into account in the agent's optimal renewal policy. Therefore, the expected value of bus ownership depends on a high-dimensional state vector consisting of the current mileage and exogenous characteristics. I am interested in the identified set of the possible values of the cost parameters that are rationalized by the agent's optimal behavior. In addition to the model in Rust, the methods developed here apply to a broad variety of dynamic models: for example, those considered in BBL or Scott:2013, etc. that point$-$ or partially identify their parameters under various assumptions about the agent's optimal behavior.
The main difficulty of this approach is estimating the value function. The plug-in (naive) approach of BBL consists of two steps. In the first step, one estimates the state variable's law of motion and the equilibrium policy function, effectively recovering agents' equilibrium beliefs. In the second step, one estimates the value function by drawing a sequence of states and actions from the conditional distributions estimated in the first step and averaging over multiple simulation draws. If the state variable is high-dimensional, we must regularize the estimate of the first-stage parameter in order to achieve consistency in high dimension. An inherent cost of these methods is bias that converges slower than parametric rate. As a result, first-stage bias carries over into the second stage, resulting in a low-quality estimate of the identified set.
The major challenge of this paper is to overcome transmitting the bias from the first to the second stage. A basic idea, proposed in a point-identified case, is to make the moment equation insensitive, or, formally, Neyman-orthogonal, to the biased estimation of the first-stage parameter (Neyman:1959,doubleml2016). The second idea consists in the use of different samples to estimate the equilibrium beliefs at the first stage and to compute the sample average of the bias correction term at the second stage. Using different samples in the first and the second stages allows me to employ modern machine learning methods to estimate the first stage parameters.
The first contribution of this paper is to derive a Neyman-orthogonal moment for the value function. This equation is presented as the sum of two terms, where the first term is the value function itself and the second term is a bias correction term that makes the sum insensitive (i.e., Neyman-orthogonal) to the first-stage bias. To derive the bias correction term, I characterize the value function as a solution to the recursive (Bellman) equation, which equates the expected value at the current state to the sum of expected immediate payoff and the expected discounted future payoff. This equation can be viewed as a semiparametric moment equation, where the parametric component gives the value function's moment (i.e., weighted average) and the nuisance parameter consists of the first-stage parameters appearing in the value function (i.e., conditional choice probability and state variable's law of motion). The first-stage parameter appears both inside and outside the value function. For example, the conditional probability of a given choice appears as a weight on immediate and discounted future payoffs corresponding to that choice. Applying the implicit function theorem to this equation, I derive the bias correction term and show how to approximate it by simulation. This derivation is novel: this is the first result in the literature that derives the bias correction term for a moment function that is not available in the closed form.
The second contribution of this paper is to extend the general theory of moment inequalities proposed by CHT to allow for moment functions that depend on a first-stage nuisance parameter that can be high-dimensional (e.g., conditional choice probability). I characterize the identified set as the minimizer of the criterion function that penalizes the incorrect sign of the moment inequality (e.g., the sign that contradicts the optimality assumption). I show that, if the moment function is insensitive with respect to the biased estimation of its nuisance parameter at each point of the space of the structural parameter, plugging in the first-stage estimate of the nuisance parameter into the moment function delivers a high-quality sample criterion function. In particular, the estimator of the identified set obtained by inverting the sample criterion function converges at the same rate as if the true value of the first-stage nuisance parameter were known. Furthermore, inferential statistics based on the estimated moment function have a non-degenerate large sample distribution and are used to construct confidence regions for the identified set by subsampling.
This paper leaves a number of open questions. First, this article assumes that the first-stage nuisance parameter is identified. This assumption does not hold for every application (e.g., CilibertoTamer). Second, the high-dimensional state space introduces a wealth of feasible suboptimal Markov policies to choose from for the construction of moment inequalities. In this paper I take the set of chosen alternatives as given, leaving the optimal choice of these alternatives for future research.
The structure of the paper is as follows. Section (ref) gives a brief overview of the results. Section (ref) gives the low-level sufficient conditions for the dynamic discrete choice model in BBL. Section (ref) presents an asymptotic theory for the identified sets defined by the semiparametric moment inequalities.
This paper is built on three lines of work: estimation and inference in partially identified models and Neyman-orthogonal semiparametric estimation.
The first line of research, see e.g. Rosen:2006, CHT, Romano:2006 develops a framework for the estimation and inference of identified sets that are partially identified by moment inequalities. Extending this framework, Kaido:2014 allow the moment function to depend on an identified low-dimensional first-stage parameter in addition to the target. Furthermore, I allow the first-stage parameter to be a high-dimensional vector or a highly complex function and estimate it by modern machine learning methods.
Within the first line of my research, my application is most connected to the estimation of dynamic models of imperfect competition. Specifically, I build on BBL who identifies the structural parameter as a solution to a set of moment inequalities that embody the assumptions about the agent's optimal behavior. This paper proposed a two-stage algorithm to estimate the parameter, where in the first-stage one estimates the state variables's law of motion and equilibrium policy function, and then plugs them into a moment equation derived from the equilibrium assumption. However, to achieve valid inference, BBL imposed parametric restrictions in the first stage. Extending this algorithm, I drop these restrictions and estimate first-stage parameters by machine learning methods.
The second line of research (Hasminskii:1981,Andrews:1994, Newey1994, vdv) is concerned with obtaining a root-$N$ consistent and asymptotically normal estimate for a low-dimensional target parameter in the presence of a nuisance parameter. In this literature, a two-stage statistical procedure is insensitive, or, formally, Neyman-orthogonal, to the estimation error of the first-stage parameter (Neyman:1959). Extending the orthogonality idea from a parametric to semiparametric setup was done in Newey1994, RRZ, Robins. Combining Neyman-orthogonality and sample splitting, doubleml2016 and LRSP incorporated modern machine learning methods to estimate low-dimensional target parameters defined by semiparametric moment equations. Subsequently, the idea has been extended to the case of high-dimensional target parameter in CGST and DenisVas. This paper translates the idea of Neyman-orthogonality from point- to set-identified case.
Within the second line research, my application is most connected to the estimation of dynamic discrete choice models under point-identification (BHKN, BCHN, Arcidiacono:2013, LRSP). Specifically, LRSP introduces high-dimensional state space into a dynamic discrete choice model whose choice set contains a renewal choice and derives a Neyman-orthogonal moment equation for the structural parameter in that model. In this paper, I address the cases that do not have renewal choice property and derive the Neyman-orthogonal moment equation for the value function directly. This result is applicable to both point- and set-identified cases with discrete and continuous choice sets.
I am interested in an identified set defined by moment inequalities
where $D$ is the data vector distributed as $P_D$, $\theta$ is the parameter of interest, and $\eta$ is an identified yet unknown parameter of the distribution $P_D$ whose true value is $\eta_0$. For example, in the bus engine replacement model of (Rust) the parameter of interest, $\theta$, is a vector of operational and replacement costs, the data vector $D$ contains bus mileage and other observed bus characteristics, $\eta$ contains mileage's law of motion and the conditional probabilities of bus replacement, and the inequality restriction comes from the assumption that agent behaves optimally. I allow the state variable, $w$ , to be high dimensional and estimate the parameter $\eta$ by modern machine learning methods.
The examples below demonstrate how a high-dimensional state $w$ may appear in the dynamic discrete choice model.
Consider the setting of Example (ref). I assume that the agent follows a Markov policy $\sigma(w,\epsilon)$ that maps the current state $w \in \mathcal{W} \subset \mathcal{R}^{d_w}$ and the shock vector $\epsilon \in \mathcal{R}^2$ into the action space $\mathcal{A} = \{0,1\}$. The value function $ V(w;\theta;\sigma)$ of a Markov policy $\sigma$ is given by
where $\beta <1$ is a discount factor. This function can also be written recursively:
where the first and second summands show the expected current profit and the future expected discounted value, respectively. Define the choice-specific value function as
where the current deterministic utility is evaluated at the action $a$ and the future discounted value is evaluated at the optimal strategy $\sigma^{*}$, conditional on the current action $a$. The symbols $ (R_0,\mu_0)$ stand for the true values of the cost parameters. The choice $a$ is optimal if and only if its total utility is greater than or equal to the utility of any other choice $a' \in \mathcal{A}$:
Then the optimal Markov policy $\sigma^{*}(w,\epsilon)$ has a cutoff form:
To identify the difference $(v(1,w) - v(0,w))$ I make the following standard assumption (e.g., HotzMiller, Scott:2013) that I will maintain throughout the paper. In general case, the model will involve several players $K$. Denote their choice sets by $\mathcal{A}_k, k\in \{1,2,\dots,K\}$.
As discussed in HotzMiller, the vector of differences of the choice-specific value functions can be expressed as
where $\gamma_0(w)={\mathrm{P}} (a=1|w)$ is the probability of the decision to maintain the engine conditional on the state $w$ which is identified. For expositional purpose, I consider a simple suboptimal Markov policy: the choice of the decision based on the coin toss:
Combining (ref) and (ref), I recognize that the optimal strategy can be viewed as a function of $\gamma$: $$\sigma^{*}(w,\epsilon) = \sigma^{*}(w,\epsilon,\gamma)$$ Therefore, value function $ V(w;\theta;\sigma^{*})$ can be viewed as the function of $\gamma$.
Define the moment function $m(w,\theta,\gamma)$ as the difference of the value function evaluated for the suboptimal and the optimal strategies
Define the identified set $\Theta_I$ as
Because $\sigma^{*}$ is an optimal strategy, the inequality above holds for the true parameter $\theta_0$, and $\Theta_I$ is a valid identified set.
The function $m(w,\theta,\eta)$ given in ((ref)) presents two complications. First, the value function $V(w;\theta;\sigma;\eta)$ depends on the unknown nuisance parameter $\eta$. Second, even if the value of $\eta$ is given, the value function is not readily available in the closed form and must be approximated by simulation. Here I focus on the first complication as if the moment function were readily available, leaving the description of the simulation algorithm for Section (ref).
A possible, though naive estimator of the value function can be constructed as follows. Consider an ideal scenario where the researcher knows mileage's law of motion. Then the choice probability $\gamma$ is the only unknown nuisance parameter that appears in the value function. Given an i.i.d sample $(D_i)_{i=1}^N$ from the law $P_D$, it is split into a main sample $J_2$ and an auxiliary sample $J_1$ of equal size $n =[N/2]$ such that $J_1 \sqcup J_2=\{1,2,\dots,N\}$. After that, the estimator of the value function $V(w;\theta;\sigma^{*};\gamma) $ is constructed as the sample average:
where $\widehat{\gamma}$ is estimated on the auxiliary sample $J_1$. Unfortunately, for some parameter values $\theta \in \Theta$ this estimator has slower than $\sqrt{N}$ convergence:
Therefore, the estimator $\widehat{\Theta}_I$ of the identified set $\Theta_I$ based on the moment function ((ref)) has suboptimal convergence rates.
The source of the slow convergence ((ref)) can be understood through the following decomposition:
The term $\mathbf{a}$ is the centered sample average of the value function $V(w;\theta;\sigma^{*};\gamma_0) $ evaluated at the true value $\gamma_0$ of the choice probability. Due to the sample splitting, the term $\mathbf{c}$ is a centered sample average of conditional on $J_1$ i.i.d random variables and is well-behaved. The term $\mathbf{b}$ stands for the bias of the value function $V(w;\theta;\sigma^{*}; \widehat{\gamma}) $ coming from the biased estimation of the conditional choice probability $\gamma$. This term is responsible for the slow convergence ((ref)).
The divergence of $\mathbf{b}$ comes from the combination of two facts: biased estimation of the first stage and the transmission of the bias from the first to the second stage. Since the conditional choice probability $\gamma(w)$ is a function of a high-dimensional state vector $w$, bias of its machine learning estimate (e.g., $\ell_1$-regularized logistic regression) converges slower than root-$N$. Because the value function $V(w;\theta;\sigma^{*};\gamma) $ is sensitive to this bias, it carries over into the second stage.
\paragraph{Overcoming the Regularization Bias using Orthogonalization.}
To overcome the translation of the first stage bias into the second stage I add the bias correction term to the value function $V(w;\theta;\sigma^{*};\gamma) $ to make it insensitive with respect to the biased estimation of $\gamma$. The new moment function for the value function $V(w;\theta;\sigma^{*};\gamma)$ takes the form
where the function $\Gamma(w;\theta)$ for $\theta = (R,\mu)$ is
Because the bias correction term has zero mean
the new moment function can replace the old one in the definition of the identified set. Moreover, the new function is insensitive to the bias estimation of $\gamma$. As a result, under additional mild regularity conditions, the regularization bias of the choice probability $\widehat{\gamma}(w)$ does not translate into the bias of the moment function:
\paragraph{The Role of Sample Splitting in Preventing Overfitting.} Another key aspect of the proposed analysis is using different samples $J_1$ and $J_2$ for different stages of estimating the value function. Had I used the whole sample to estimate the conditional choice probability $\widehat{\gamma}(w)$, the sample average
would not be a sample average of the i.i.d (or weakly dependent) random observations. The relation between the first stage error $\widehat{\gamma}(w_i) - \gamma_0(w_i)$ and the value of the moment function $ g(D_i, \theta, \gamma_0)$ creates bias, referred to as overfitting bias. To control the overfitting bias in the worst-case scenario:
one must impose complexity constraints on the class of functions $\mathcal{G}$ that are used to estimate the conditional choice probability. While some nonparametric estimators designed for low-dimensional state spaces obey these constraints, some modern machine learning estimators designed for high-dimensional state variables do not. To accommodate machine learning estimators at the first stage, I use different samples.
\paragraph{Bias Correction Terms for Examples (ref) and (ref)}. Suppose a nuisance parameter can be presented as a conditional expectation $\mathbb{E}[U|w]$. Then bias correction term for conditional choice probability takes the form $$ \alpha(D;\theta;\xi) = \Pi (w;\theta) ( U - \mathbb{E} [U|w]),$$ where $\xi$ is an unknown vector-valued function of the state variable $w$. The true value $\xi_0 = \xi_0(\theta)$ consists of the original conditional expectation function $\mathbb{E} [U|w]$ and the function $ \Pi (w;\theta) $
In some point-identified problems the value function $V(w;\theta;\sigma;\eta)$ is only evaluated at the optimal strategy $\sigma^{*}$ and the true parameter value $\theta_0$. Then the bias correction term is evaluated only at the true value $\theta_0$. In particular, in the Examples (ref) and (ref) the function $\Gamma(w,\theta_0)$ can be further simplified as follows:
where we have used ((ref)) in the second line.
I will now introduce the estimator of the identified set $\Theta_I$, leaving the formal definition to ((ref)). Suppose there exists a function $g(D,\theta,\xi)$ that preserves the expectation of the moment function $m(w,\theta,\eta)$
and is insensitive to the biased estimation of its own nuisance parameter $\xi$ around its true value $\xi_0=\xi_0(\theta)$, where $\xi_0(\theta)$ is an identified vector-valued parameter of the distribution $P_D$ for each $\theta \in \Theta$. In many relevant cases such as Example (ref), $\xi_0$ contains the original nuisance parameter $\eta$, but may contain more unknown parameters of the distribution $P_D$.
To estimate the identified set $\Theta_I$ I represent it as the minimizer of the criterion function $Q(\theta,\xi_0)$
where $Q(\theta,\xi_0)$ is
and its sample analog is
My goal is to use different samples in the first and the second stages in order to avoid overfitting. Yet, simple sample splitting has a drawback that only one half of the sample is used for second-stage estimation, which can lead to loss of efficiency in small samples. In order to use the whole sample for the second stage yet keep the sample splitting idea, I use cross-fitting procedure described below.
The contour level $\widehat{c}$, possibly data dependent, is chosen such that $ \widehat{\Theta}_I$ contains the true set $\Theta_I$ with probability approaching one
I establish convergence and inference properties of the contour set estimator $\widehat{\Theta}_I$. The first property is formulated in terms of the convergence rate of $\widehat{\Theta}_I$ to $\Theta_I$ is based on the notion of Hausdorff distance $d_H(\widehat{\Theta}_I,\Theta_I)$:
The set $\widehat{\Theta}_I$ is said to converge to $\Theta_I$ at rate $\epsilon_N$ if the Hausdorff distance between the sets converges at rate $\epsilon_N$:
Under mild regularity conditions it is possible to achieve the nearly efficient rate $\epsilon_N = O(\sqrt{\frac{ \log N}{N}})$.
To conduct inference, I fix a confidence level $\tau \in (0,1)$. A confidence region $C_N(c_{\tau},\widehat{\xi})$ of level $\tau$ is defined as a contour set $C_N(c_{\tau};\widehat{\xi})$ of level $c_{\tau}$ such that $C_N(c_{\tau};\widehat{\xi})$ contains $\Theta_I$ with probability at least $1-\tau$:
I construct a confidence region $ C_N(\widehat{c}_{\tau};\widehat{\xi})$, where $\widehat{c}_{\tau}$ is a consistent estimator of the $\tau$-quantile, denoted $c_{\tau}$, of the inferential statistic:
The consistent estimator of the $\tau$-quantile is chosen by the subsampling algorithm defined below.
Section (ref) establishes the asymptotic validity of the set estimator given in Definition (ref) and subsampling algorithm of Definition (ref)
Consider the dynamic model of strategic interaction from BBL. There are $K$ players, denoted by $\{1,2,\dots, K\}$. Each player $k$ makes a decision $a_k \in \mathcal{A}_k$ from a finite set of discrete alternatives $\mathcal{A}_k$ at time periods $t \in \{0,1,\dots, \infty\}$. In each period $t$ the players commonly observe a vector of state variables $w_t \in \mathcal{W} \subset \mathcal{R}^{d_w}$. Given the state variable $w_t=w$, players choose actions simultaneously. Before choosing his action, each player $k$ observes a vector of private shocks $(\epsilon_k(j))_{j \in \mathcal{A}_k}$ corresponding to each discrete alternative $j$ in his choice set $\mathcal{A}_k$. The transition between states follows a conditional probability distribution: $P(\cdot| a,w) $ conditional on the current state $w$ and the action profile $a=(a_1,a_2,\dots,a_K)$.
I focus on the structural parameter describing the utility of the first player. I assume that his per-period utility function is given by
where $a=(a_1,a_2,\dots,a_K)$ is the profile of the players' actions, $a_1 \in {\mathcal{A}_1}$ is the action of the first player, $w$ is the state variable, and $\theta_0$ is the true value of the structural parameter $\theta$. The per-period utility is presented as the sum of a deterministic component $\tilde{\pi}(a,w;\theta;\zeta)$ and the private shock $\epsilon_1(a_1)$. I allow $\tilde{\pi}(a,w;\theta;\zeta)$ to depend on identified nuisance parameter $\zeta$ whose true value is $\zeta_0$.
I assume that each player follows a pure Markov policy. The pure Markov policy for player one $\sigma_1(w,\epsilon_1):\mathcal{W} \bigtimes \mathcal{E} \rightarrow \mathcal{A}_1$ maps the current state $w$ and the private shock of player one, $\epsilon_1$, into the action space $\mathcal{A}_1$. If the behavior of the players is described by a Markov policy profile $\sigma = (\sigma_1,\sigma_2,\dots,\sigma_K)$, the value function is given by
A strategy profile $\sigma^{*}$ is a Markov perfect equilibrium if each player $k$ prefers its strategy $\sigma^{*}_k$ to all alternative Markov strategies as long as the others follow the equilibrium strategy $\sigma^{*}_{-k}$. That is, the value function $ V(w;\theta_0;\sigma^{*};\eta_0)$ of the first player at the strategy $\sigma^{*}$ is weakly larger than the value function of any other strategy profile $\sigma = (\sigma_1,\sigma_{-1}^{*})$:
where $\sigma = (\sigma_1,\sigma_{-1}^{*})$ consists of a feasible suboptimal alternative for player one, $\sigma_1$, and the equilibrium profile for the other players $\sigma_{-1}^{*}$. I assume that the each observation comes from the same Markov perfect equilibrium $\sigma^{*}$, although I do not provide conditions for the existence of such equilibrium and recognize that there could be many such equilibria.
Define the choice-specific value function of player one $v(a_1,w)$ as the expected present value conditional on the current state $w$ and the choice $a_1$
Then, following the optimal strategy $\sigma^{*}(w,\epsilon)$, the first player chooses $a_1$ if and only if the total utility of $a_1$ is not smaller than the total utility of any other choice $a_1' \in \mathcal{A}_1$
Therefore, the optimal strategy of player one is characterized as
or, equivalently,
As discussed in HotzMiller, under Assumption (ref) the vector of differences of the choice-specific value functions can be expressed as
where ${\mathrm{P}} (a_1|w)$ is the probability of the choice $a_1 \in \mathcal{A}_1$ conditional on the state variable $w$. Finally, as a suboptimal Markov policy $\sigma$, I consider a cutoff-type strategy
where I add a deviation function $\text{dev}(a_1,w)$ for each element $v(a_1,w) - v(1,w)$. As a normalization condition, I set $\text{dev}(1,w)=0 \quad \forall w$. Therefore, the inequality ((ref)) depends on the nuisance parameter
that consists of any original nuisance parameter $\zeta$ that may appear in the utility function ((ref)), the conditional choice probabilities of all players $$\gamma_{jk} (w):= {\mathrm{P}}(a_k=j|w), \quad j \in \mathcal{A}_k, k \in \{1,2,\dots,K\},$$ and the conditional distribution ${\mathrm{P}} (w_{next}|w,a)$.
I construct the identified set $\Theta_I$ using a subset of inequalities implied by the equilibrium definition (i.e, ((ref))). Let $q(w)$ be an $L$-vector of non-negative weighting functions. Let $\text{dev}(a_1,w): \mathcal{A}_1 \bigtimes \mathcal{W} \rightarrow \mathcal{R}^d$ be an $ L$-vector of deviation functions whose coordinate $l \in \{1,2,\dots,L\}$ corresponds to a deviation strategy $\text{dev}_l(a_1,w)$. Define the moment vector-valued function as
whose component $m_l(D,\theta,\eta)$ is equal to the weighted difference $V(w;\theta;\sigma_l;\eta) - V(w;\theta;\sigma^{*};\eta)$: $$ m_l(w,\theta,\eta) :=q_l(w) (V(w;\theta;\sigma_l;\eta) - V(w;\theta;\sigma^{*};\eta)), \quad l \in \{1,2,\dots,L\}$$ evaluated at a strategy profile $\sigma_l = (\sigma_{1,l}, \sigma^{*}_{-1})$. The suboptimal strategy of the first player $\sigma_{1,l}$ is given in ((ref)) with the deviation function $\text{dev}_l(a_1,w)$. The identified set $\Theta_I$ is defined as the collection of parameter values $\theta$ that obey inequality restrictions in expectation
where the expectation is taken with respect to the unconditional distribution of the state $w$.
The value function $V(w;\theta;\sigma;\eta)$ appearing in the moment function ((ref)) is not available in closed form. However, it can be approximated by Monte Carlo simulation as in BBL. In the first stage, one constructs an estimate $\widehat{\eta}$ of the nuisance parameter $\eta$ that is defined in ((ref)). In the second stage, one simulates the sequence of states and shocks from the estimated first stage parameter $\widehat{\eta}$ and averages the realization of the value function over multiple simulation draws. A single simulation draw is given by Algorithm (ref).
In order to evaluate the value function $V(w;\theta;\sigma;\eta)$ at different parameter values $\theta_1 \in \Theta$ and $\theta_2 \in \Theta$, I use the same simulation draws. As long as the number of simulation draws is large enough, the simulation error does not affect the asymptotic properties of the estimator of the identified set $\widehat{\Theta}_I$ as discussed in Pakes:1989.
When the nuisance parameter $\eta$ is high-dimensional and estimated by machine learning, the plug-in estimator of value function is biased. To make the moment function ((ref)) insensitive to first-stage bias, I derive bias correction term. As discussed in Newey1994, the bias correction term for a vector-valued nuisance parameter is equal to the sum of individual terms of individual components. Moreover, the bias correction term for the conditional probability of choice $j$ has the product structure
Furthermore, according to Newey1994, the function $ \Gamma(w;\theta)$ is implicitly defined by the orthogonality condition explained below.
Let $g(D,\theta,\xi)$ be a moment function. Define the Gateaux derivative map $D_r: \Xi \bigtimes \Theta \rightarrow \mathcal{R}^L$ as
for all $r \in [0,1)$, which I assume exists. I also denote the pathwise derivative of the expected moment function at the true value $\xi_0$
The definition of Neyman orthogonality requires that the moment function be insensitive to the biased estimation of $\xi$. This condition is the generalization of the orthogonality condition for point-identified models in doubleml2016 that is required to hold only at the true value $\theta_0$ of identified parameter $\theta$. In contrast to the point-identified case, I require the equality in ((ref)) to hold at each point $\theta$ of the parameter space $\Theta$. In many relevant cases (e.g., if the moment function $ g(D,\theta,\xi)$ is linear in $\theta$), the orthogonality condition ((ref)) on the set $\Theta$ follows from the orthogonality ((ref)) on a finite subset of $\Theta$. Rewriting the orthogonality condition ((ref)) for the bias correction term for the value function gives
I find the function $\Gamma(w;\theta)$ from the recursive definition of the value function ((ref)). In the case of Example (ref) the recursive definition can be rewritten in the unconditional form
where $\gamma_0(w)$ is the true value of the conditional choice probability and $\theta = (R,\mu)$ consists of the replacement and maintenance costs. The unknown function $\gamma(w)$ appears in Equation ((ref)) both {\bf outside} the value function $V(w;\theta;\sigma^{*};\gamma) $ and {\it inside} of this function. I consider a local deviation of the choice probability $\gamma(\cdot)$ from its value $\gamma_0(\cdot)$ for each value of the parameters $R$ and $\mu$. Applying chain rule to the Equation ((ref)) yields a pathwise derivative
where the last equality follows from the logistic distribution of the private shocks (Assumption (ref)).
In what follows I derive the bias correction terms for the other nuisance parameters that are present in the dynamic discrete choice in BBL. Let $q(w)$ be a weighting function. Define
as the expectation of the weighting function $q(w)$ evaluated at the current state $w$ conditional on the future state $w_{next}$. Denote the expectation of the choice $j$ made by player one conditional on the state $w$ by $\gamma_j(w):= \mathbb{E} [1_{\{a_1=j \}}|w]$. Let $$\gamma(w):= (\gamma_2(w),\gamma_3(w),\dots, \gamma_{A_1}(w))$$ be the $A_1-1$ vector of these probabilities. Define the conditional expectation of the current private shock evaluated at the equilibrium strategy $\sigma_1^{*} = \sigma_1^{*}(w,\epsilon,\gamma)$ as
Lemma (ref) gives the bias correction term, $ \alpha^{CCP}_{j}(D;\theta;\gamma_j)$, that makes the value function $V(w;\theta;\sigma^{*};\eta)$ insensitive to the biased estimation of $\gamma_j$.
Both the value function and the bias correction terms presented above are not available in closed form and must be simulated. When the value function is a linear function of $\theta$, the simulation can be simplified. Suppose there exist basis functions $\Psi_1(w;\sigma;\eta), \Psi_2(w;\sigma;\eta) $ such that
is an affine function of $\theta$. Then one can simulate the basis functions $\Psi_1(w;\sigma;\eta), \Psi_2(w;\sigma;\eta) $ instead of simulating the value function for each $\theta \in \Theta$. Lemma (ref) provides the sufficient conditions for the linearity of the value function and the individual bias correction terms provided in Lemmas (ref), (ref), (ref).
Suppose I have a collection of inequality restrictions on an economic model coming from the data structure and/or the assumptions about the data generating process. These restrictions are embodied into a moment function $g(D,\theta,\xi): \mathcal{D} \bigtimes \Theta \bigtimes \Xi \rightarrow \mathcal{R}^L$, where $\theta \in \Theta \subset \mathcal{R}^d$ is the target parameter. In addition to the target parameter $\theta$, the moment function $g(D,\theta,\xi)$ depends on a nuisance parameter $\xi=\xi(\theta)$ whose true value $\xi_0=\xi_0(\theta)$ is an identified parameter of the data distribution $P_D$ and belongs to a convex subset of a normed vector space $ \Xi$. The object of interest is an identified set $\Theta_I$ defined as a collection of parameter values $\theta$ that satisfy the inequality restrictions
at the true value $\xi_0$ of the nuisance parameter.
The identified set $\Theta_I$ is characterized as the minimizer of the criterion function
I assume that the following partial identification condition holds. There exist positive constant $C>0$ and $\delta>0$ such that
This condition states that once $\theta$ is bounded away from $\Theta$, the moment $\mathbb{E} g(D,\theta, \xi_0 )$ is bounded away from $\Theta_I$ by a number that is proportional to the Hausdorff distance $d_H(\theta,\Theta_I)$ from $\theta$ to the identified set $\Theta_I$. This condition ensures that the true moment function $ g(D,\theta, \xi_0) $ distinguishes the boundary of the identified set.
\paragraph{Impact of the First-Stage Estimation.} In the next condition I introduce a sequence of neighborhoods $\Xi_N^{\theta} \subset \Xi^{\theta}$ of $\xi_0(\theta)$ that contain the estimate $\widehat{\xi}(\theta)$ of $\xi_0(\theta)$ w.p. approaching one. As the sample size $N$ increases, the neighborhoods shrink. The quality of the estimation of the first-stage parameter is defined as the speed of shrinkage of the neighborhood $\Xi_N(\theta)$ around $\xi_0(\theta)$. I refer to it as the first-stage rate $g_N$. Finally, I assume that the second-order derivative of the functional $\mathbb{E} g(D,\theta, \xi)$ is well-behaved. Combined with the orthogonality condition and the upper bound on the first-stage rate $g_N$, this assumption ensures that I can ignore the impact of the estimation error $\widehat{\xi}(\theta)-\xi_0(\theta)$ on the second and the higher-order derivatives of the moment function $\mathbb{E} g(D,\theta, \xi)$ with respect to $\xi$ at $\xi_0$.
The next condition requires that the moment function $g(D;\theta;\xi(\theta))$ is sufficiently regular with respect to $\theta$ for each fixed element $\xi \in \Xi$ of the nuisance realization set $\Xi$. I consider a class $$\mathcal{F}_{\xi}:= \{ g(\cdot,\theta, \xi(\theta)), \theta \in \Theta \}$$ and require the uniform covering entropy of this class to be bounded.
Conditions (ref)(1)-(2) are the generalization of the regularity assumption in the point-identified moment problem of doubleml2016. Because $\xi (\theta)$ depends on $\theta$, conditions (ref)(1)-(2) are non-standard and require verification in applications. However, this requirement is mild when the nuisance parameter $\xi (\theta)$ is a linear function of $\theta$. Suppose that the true value of the nuisance parameter $\xi_0(\theta)$ is a linear function of $\theta$
where $\xi_0^{a}(D)$ and $\xi_0^{b}(D)$ are the identified parameters of the distribution $P_D$. Then conditions (ref)(1)-(2) can be reformulated in terms of the nuisance parameters $\{ \xi_0^{a}, \xi_0^{b}\}$ that no longer depend on $\theta$. When ((ref)) holds, Conditions (ref) (1)-(2) are satisfied for many practical cases. In particular, the functions $\xi_0^{a}(D), \xi_0^{b}(D)$ can be estimated by $\ell_1$-regularized methods, random forests, and deep neural nets under plausible assumptions about their structure.
\paragraph{Donsker Property.} Let $\Theta'$ be an open neighborhood of $\Theta$. I require the moment function $g(D_i,\theta,\xi_0)$ to have a Donsker property defined as follows. In the metric space $L^{\infty}(\Theta')$,
where $\Delta(\theta)$ is a mean zero Gaussian process on $\Theta$ with a.s. continuous paths and $\text{Var}(\Delta(\theta))>0$ for each $\theta \in \Theta'$. In addition, the probability space $(\Omega, \mathcal{F}, \mathcal{P})$ is rich enough to support the representation ((ref)).
Theorem (ref) is my first main result. It establishes the sufficient conditions on the moment function to deliver the rate of convergence of $\widehat{\Theta}_I$ to $\Theta_I$. It suggests that the contour level $\widehat{c}$ as small as possible subject to the constraint ((ref)). Setting $\widehat{c} = O_{P} (1)$ subject to ((ref)) delivers the optimal rate, but this choice is infeasible. Setting $\widehat{c} \sim \log N$ delivers a nearly efficient rate.
In many cases it is possible to establish convergence without the requirement ((ref)). This is possible because the criterion function is degenerate ((ref)).
Lemma (ref) is my second main result. It provides the sufficient conditions under which the identified set $\Theta_I$ can be estimated at the fastest possible rate $\epsilon_N = O_{P} (N^{-1/2})$. It requires that the sample criterion function $Q_N(\theta,\xi_0)$ be flat on the (possibly) data-dependent sets $\Theta_N$ that approximate the identified set $\Theta_I$ sufficiently well. If this requirements holds, the contour level $\widehat{c}_0:= \arg \min_{\theta \in \Theta_0} Q_N(\theta,\xi_0)$ delivers the optimal rate for the contour level set ${\cal C}_N(\widehat{c}_0, \xi_0)$ based on the true value of the nuisance parameter (see, e.g. CHT). We show that a modified choice $\widehat{c}$ given in ((ref)) delivers the optimal rate in the presence of the nuisance parameter $\widehat{\xi}$ estimated in the first stage on an auxiliary sample.
\paragraph{Subsampling.} I wish to construct the contour set ${\cal C}(\widehat{c}, \widehat{\xi})$ that has confidence region property ((ref)). To do this, I must find the asymptotic distribution of the inferential statistic $$ {\cal C}_N:= \sup_{\theta \in \Theta_I} Q(\theta; \xi_0)$$ that can be used to estimate the $\tau$-quantile of $ {\cal C}_N$. Define the random variable
where $I(\theta)$ is an $L$-vector of functions $I_l(\theta), l \in \{ 1,\dots,L\})$. The function $I_l(\theta)$ is defined as follows:
CHT show that ${\cal C}_N$ converges to $ \mathcal{C}$ in distribution, and that $ \mathcal{C}$ has a non-degenerate and continuous distribution function. The final requirement for the validity of the subsampling algorithm of Definition (ref) is that the inferential statistic
is well-behaved on the $\epsilon_N$-expansion of the identified set $\Theta_I$. A sufficient condition for this requirement is given below.