EconBase
← Back to paper

Machine Learning for Dynamic Discrete Choice

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

75,875 characters · 11 sections · 62 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Machine Learning for Dynamic Discrete Choice

abstractDynamic discrete choice models often discretize the state vector and restrict its dimension in order to achieve valid inference. I propose a novel two-stage estimator for the set-identified structural parameter that incorporates a high-dimensional state space into the dynamic model of imperfect competition. In the first stage, I estimate the state variable's law of motion and the equilibrium policy function using machine learning tools. In the second stage, I plug the first-stage estimates into a moment inequality and solve for the structural parameter. The moment function is presented as the sum of two components, where the first one expresses the equilibrium assumption and the second one is a bias correction term that makes the sum insensitive (i.e., Neyman-orthogonal) to first-stage bias. The proposed estimator uniformly converges at the root-$N$ rate and I use it to construct confidence regions. The results developed here can be used to incorporate high-dimensional state space into classic dynamic discrete choice models, for example, those considered in Rust, BBL, and Scott:2013.

Introduction

In empirical work on dynamic models, economists often make specification choices - for example, discretize state space or select covariates for flow utility - in order to achieve computational tractability and precise estimates (e.g. PakesShankman:1984, Pakes:1986, Rust, Ryan:2012). Typically, economists have little intuition about which covariates to select or how to discretize a continuous state variable (PakesLanjouw:1998). Unfortunately, counterfactual predictions of dynamic models are sensitive to specification choices and are difficult to interpret when a model is misspecified. As discussed in athey2017science, there has been recent interest in data-driven model selection based on modern machine learning tools. Moreover, (orthogStructural, doubleml2016, LRSP, AtheyWager) have shown how to leverage these tools into high-quality estimates of causal parameters. Because dynamic models are more challenging to analyze, the model selection in dynamic models remains an open question.

This paper estimates a dynamic model of imperfect competition with a high-dimensional state space. Consider the bus engine replacement model from Rust as a special case. An agent decides whether to replace the bus engine in each period. The agent incurs a fixed cost in the case of replacement and a cost proportional to the current mileage in the case of maintenance. One can imagine that the future bus mileage depends on a vector of observed exogenous characteristics, such as current traffic and weather conditions, encoded in a high-dimensional vector. While these characteristics do not enter into the per-period utility of the owner, they affect the mileage's law of motion in the case of maintenance decision, and hence are taken into account in the agent's optimal renewal policy. Therefore, the expected value of bus ownership depends on a high-dimensional state vector consisting of the current mileage and exogenous characteristics. I am interested in the identified set of the possible values of the cost parameters that are rationalized by the agent's optimal behavior. In addition to the model in Rust, the methods developed here apply to a broad variety of dynamic models: for example, those considered in BBL or Scott:2013, etc. that point$-$ or partially identify their parameters under various assumptions about the agent's optimal behavior.

The main difficulty of this approach is estimating the value function. The plug-in (naive) approach of BBL consists of two steps. In the first step, one estimates the state variable's law of motion and the equilibrium policy function, effectively recovering agents' equilibrium beliefs. In the second step, one estimates the value function by drawing a sequence of states and actions from the conditional distributions estimated in the first step and averaging over multiple simulation draws. If the state variable is high-dimensional, we must regularize the estimate of the first-stage parameter in order to achieve consistency in high dimension. An inherent cost of these methods is bias that converges slower than parametric rate. As a result, first-stage bias carries over into the second stage, resulting in a low-quality estimate of the identified set.

The major challenge of this paper is to overcome transmitting the bias from the first to the second stage. A basic idea, proposed in a point-identified case, is to make the moment equation insensitive, or, formally, Neyman-orthogonal, to the biased estimation of the first-stage parameter (Neyman:1959,doubleml2016). The second idea consists in the use of different samples to estimate the equilibrium beliefs at the first stage and to compute the sample average of the bias correction term at the second stage. Using different samples in the first and the second stages allows me to employ modern machine learning methods to estimate the first stage parameters.

The first contribution of this paper is to derive a Neyman-orthogonal moment for the value function. This equation is presented as the sum of two terms, where the first term is the value function itself and the second term is a bias correction term that makes the sum insensitive (i.e., Neyman-orthogonal) to the first-stage bias. To derive the bias correction term, I characterize the value function as a solution to the recursive (Bellman) equation, which equates the expected value at the current state to the sum of expected immediate payoff and the expected discounted future payoff. This equation can be viewed as a semiparametric moment equation, where the parametric component gives the value function's moment (i.e., weighted average) and the nuisance parameter consists of the first-stage parameters appearing in the value function (i.e., conditional choice probability and state variable's law of motion). The first-stage parameter appears both inside and outside the value function. For example, the conditional probability of a given choice appears as a weight on immediate and discounted future payoffs corresponding to that choice. Applying the implicit function theorem to this equation, I derive the bias correction term and show how to approximate it by simulation. This derivation is novel: this is the first result in the literature that derives the bias correction term for a moment function that is not available in the closed form.

The second contribution of this paper is to extend the general theory of moment inequalities proposed by CHT to allow for moment functions that depend on a first-stage nuisance parameter that can be high-dimensional (e.g., conditional choice probability). I characterize the identified set as the minimizer of the criterion function that penalizes the incorrect sign of the moment inequality (e.g., the sign that contradicts the optimality assumption). I show that, if the moment function is insensitive with respect to the biased estimation of its nuisance parameter at each point of the space of the structural parameter, plugging in the first-stage estimate of the nuisance parameter into the moment function delivers a high-quality sample criterion function. In particular, the estimator of the identified set obtained by inverting the sample criterion function converges at the same rate as if the true value of the first-stage nuisance parameter were known. Furthermore, inferential statistics based on the estimated moment function have a non-degenerate large sample distribution and are used to construct confidence regions for the identified set by subsampling.

This paper leaves a number of open questions. First, this article assumes that the first-stage nuisance parameter is identified. This assumption does not hold for every application (e.g., CilibertoTamer). Second, the high-dimensional state space introduces a wealth of feasible suboptimal Markov policies to choose from for the construction of moment inequalities. In this paper I take the set of chosen alternatives as given, leaving the optimal choice of these alternatives for future research.

The structure of the paper is as follows. Section (ref) gives a brief overview of the results. Section (ref) gives the low-level sufficient conditions for the dynamic discrete choice model in BBL. Section (ref) presents an asymptotic theory for the identified sets defined by the semiparametric moment inequalities.

Literature review

This paper is built on three lines of work: estimation and inference in partially identified models and Neyman-orthogonal semiparametric estimation.

The first line of research, see e.g. Rosen:2006, CHT, Romano:2006 develops a framework for the estimation and inference of identified sets that are partially identified by moment inequalities. Extending this framework, Kaido:2014 allow the moment function to depend on an identified low-dimensional first-stage parameter in addition to the target. Furthermore, I allow the first-stage parameter to be a high-dimensional vector or a highly complex function and estimate it by modern machine learning methods.

Within the first line of my research, my application is most connected to the estimation of dynamic models of imperfect competition. Specifically, I build on BBL who identifies the structural parameter as a solution to a set of moment inequalities that embody the assumptions about the agent's optimal behavior. This paper proposed a two-stage algorithm to estimate the parameter, where in the first-stage one estimates the state variables's law of motion and equilibrium policy function, and then plugs them into a moment equation derived from the equilibrium assumption. However, to achieve valid inference, BBL imposed parametric restrictions in the first stage. Extending this algorithm, I drop these restrictions and estimate first-stage parameters by machine learning methods.

The second line of research (Hasminskii:1981,Andrews:1994, Newey1994, vdv) is concerned with obtaining a root-$N$ consistent and asymptotically normal estimate for a low-dimensional target parameter in the presence of a nuisance parameter. In this literature, a two-stage statistical procedure is insensitive, or, formally, Neyman-orthogonal, to the estimation error of the first-stage parameter (Neyman:1959). Extending the orthogonality idea from a parametric to semiparametric setup was done in Newey1994, RRZ, Robins. Combining Neyman-orthogonality and sample splitting, doubleml2016 and LRSP incorporated modern machine learning methods to estimate low-dimensional target parameters defined by semiparametric moment equations. Subsequently, the idea has been extended to the case of high-dimensional target parameter in CGST and DenisVas. This paper translates the idea of Neyman-orthogonality from point- to set-identified case.

Within the second line research, my application is most connected to the estimation of dynamic discrete choice models under point-identification (BHKN, BCHN, Arcidiacono:2013, LRSP). Specifically, LRSP introduces high-dimensional state space into a dynamic discrete choice model whose choice set contains a renewal choice and derives a Neyman-orthogonal moment equation for the structural parameter in that model. In this paper, I address the cases that do not have renewal choice property and derive the Neyman-orthogonal moment equation for the value function directly. This result is applicable to both point- and set-identified cases with discrete and continuous choice sets.

Set-Up and Motivation

I am interested in an identified set defined by moment inequalities

align[align omitted — 101 chars of source]

where $D$ is the data vector distributed as $P_D$, $\theta$ is the parameter of interest, and $\eta$ is an identified yet unknown parameter of the distribution $P_D$ whose true value is $\eta_0$. For example, in the bus engine replacement model of (Rust) the parameter of interest, $\theta$, is a vector of operational and replacement costs, the data vector $D$ contains bus mileage and other observed bus characteristics, $\eta$ contains mileage's law of motion and the conditional probabilities of bus replacement, and the inequality restriction comes from the assumption that agent behaves optimally. I allow the state variable, $w$ , to be high dimensional and estimate the parameter $\eta$ by modern machine learning methods.

The examples below demonstrate how a high-dimensional state $w$ may appear in the dynamic discrete choice model.

example[Engine Replacement Model from Rust with a High-Dimensional State Variable] A single agent makes a binary decision $a \in \mathcal{A} = \{0,1\}$ whether to replace a bus engine in each period $t \in \{1,2,\dots, \infty\}$. His per-period utility function is \begin{align} \pi(a,s,\epsilon) = \begin{cases} -R + \epsilon(0),\quad a=0,\\ -\mu \cdot s + \epsilon(1),\quad a=1, \end{cases} \end{align} where $s \in \mathcal{R}$ is the bus mileage, $R$ is replacement costs, $\mu s$ is operational cost, $a$ is the decision of the agent, and $\epsilon = (\epsilon(0),\epsilon(1))$ is a vector of private shocks associated with each decision. The state variable $w$ consists of the mileage $s$ and additional high-dimensional vector of exogenous variables $x$ (e.g., engine manufacturer characteristics) that I assume do not change with time. After the replacement decision $(a=0)$ the mileage $s_{next}$ resets to $1$ with probability one. After the maintenance decision $(a=1)$, the mileage $s_{next}$ follows a first-order Markov process \begin{align} s_{next} &= \rho_0(w) + e, \quad e \sim N(0,1), \end{align} where $\rho_0(w)$ is the conditional expectation function of the mileage tomorrow $s_{next}$ given the state today $w$ and $e \sim N(0,1)$ is an independent $N(0,1)$ shock. The target parameter $\theta = (R,\mu)$ consists of the replacement and operational cost parameters. The observed data vector consists of the current state, action, and future state (i.e, $D=(w,a,w_{next})$). Under the assumptions discussed below, the unknown yet identified high-dimensional parameters consist of the conditional choice probability $\gamma(w)={\mathrm{P}}(a=1|w)$ and the transition function $\rho(w)$.
example[Entry Game with a Long-Lived and a Short-Lived Player] In each period $t \in \{1,2,\dots, \infty\}$ Apple decides whether to issue a new model of a phone $(a_{Apple}=0)$ or keep the existing one $(a_{Apple}=1)$. In addition to Apple, a short-lived potential entrant (Player 2) decides whether to issue a fake $(a_{P}=1)$ or not $(a_{P}=0)$. Each period Apple faces a new copy of player 2. Player 2's actions do not influence the motion of the state, and he has no dynamic incentives. In each period both players observe a state vector $w = (s,x)$ that consists of the age of the current make $s$ and the vector $x$ of short-lived player's characteristics (e.g., country, information about intellectual property rights protection) that does not change with time (i.e, $x_t=x_0$). In each period both players receive a privately observed shock. Apple's utility function is given by \begin{align} \pi(a,w,\epsilon) = \begin{cases} - R + \delta_1 a_{P} + \epsilon(0), \quad a_{Apple}=0,\\ -\mu \cdot s + \delta_2 a_{P} + \epsilon(1), \quad a_{Apple}=1, \end{cases} \end{align} where $R$ is the fixed cost of replacing the current model with a new one, $-\mu \cdot s$ is the profit from the current make that decays with age, and $\epsilon = (\epsilon(0),\epsilon(1))$ is a vector of Apple's shocks associated with each decision. After the replacement decision $(a_{Apple}=1)$ the age $s_{next}$ resets to $1$ with probability one. After the maintenance decision $(a_{Apple}=2)$ the age $s_{next}$ increases by $1$ with probability $1$ $$s_{next} = s +1.$$ The target parameter $\theta = (R,\mu,\delta_1-\delta_0)$ consists of the cost parameters and the difference between the interaction parameters $\delta_1-\delta_0$. The observed data vector $D=(w,a,w_{next})$ consists of the current state, $w$, action profile $a = (a_{Apple},a_{P})$, and the future state, $w_{next}$. Under the assumptions discussed below, the unknown yet identified high-dimensional parameters consist of the conditional choice probability of both players: $\gamma_A(w)={\mathrm{P}}(a_{Apple}=1|w)$ and $\gamma_P(w):={\mathrm{P}}(a_{P}=1|w)$.

Consider the setting of Example (ref). I assume that the agent follows a Markov policy $\sigma(w,\epsilon)$ that maps the current state $w \in \mathcal{W} \subset \mathcal{R}^{d_w}$ and the shock vector $\epsilon \in \mathcal{R}^2$ into the action space $\mathcal{A} = \{0,1\}$. The value function $ V(w;\theta;\sigma)$ of a Markov policy $\sigma$ is given by

align*[align* omitted — 127 chars of source]

where $\beta <1$ is a discount factor. This function can also be written recursively:

align[align omitted — 220 chars of source]

where the first and second summands show the expected current profit and the future expected discounted value, respectively. Define the choice-specific value function as

align*[align* omitted — 216 chars of source]

where the current deterministic utility is evaluated at the action $a$ and the future discounted value is evaluated at the optimal strategy $\sigma^{*}$, conditional on the current action $a$. The symbols $ (R_0,\mu_0)$ stand for the true values of the cost parameters. The choice $a$ is optimal if and only if its total utility is greater than or equal to the utility of any other choice $a' \in \mathcal{A}$:

align*[align* omitted — 94 chars of source]

Then the optimal Markov policy $\sigma^{*}(w,\epsilon)$ has a cutoff form:

align[align omitted — 188 chars of source]

To identify the difference $(v(1,w) - v(0,w))$ I make the following standard assumption (e.g., HotzMiller, Scott:2013) that I will maintain throughout the paper. In general case, the model will involve several players $K$. Denote their choice sets by $\mathcal{A}_k, k\in \{1,2,\dots,K\}$.

assumption[Independent logit errors] The components of the private shock vector $\epsilon = \{ (\epsilon_k(j))_{j \in \mathcal{A}_k}\}_{k=1}^K \in \mathcal{E}$ are identically and independently distributed with a type $1$ extreme value distribution whose distribution function is equal to $F(t) = \exp(-\exp(-t))$.

As discussed in HotzMiller, the vector of differences of the choice-specific value functions can be expressed as

align[align omitted — 160 chars of source]

where $\gamma_0(w)={\mathrm{P}} (a=1|w)$ is the probability of the decision to maintain the engine conditional on the state $w$ which is identified. For expositional purpose, I consider a simple suboptimal Markov policy: the choice of the decision based on the coin toss:

align[align omitted — 88 chars of source]

Combining (ref) and (ref), I recognize that the optimal strategy can be viewed as a function of $\gamma$: $$\sigma^{*}(w,\epsilon) = \sigma^{*}(w,\epsilon,\gamma)$$ Therefore, value function $ V(w;\theta;\sigma^{*})$ can be viewed as the function of $\gamma$.

Define the moment function $m(w,\theta,\gamma)$ as the difference of the value function evaluated for the suboptimal and the optimal strategies

align[align omitted — 116 chars of source]

Define the identified set $\Theta_I$ as

align*[align* omitted — 77 chars of source]

Because $\sigma^{*}$ is an optimal strategy, the inequality above holds for the true parameter $\theta_0$, and $\Theta_I$ is a valid identified set.

Naive Approach to the Estimation of the Identified Set

The function $m(w,\theta,\eta)$ given in ((ref)) presents two complications. First, the value function $V(w;\theta;\sigma;\eta)$ depends on the unknown nuisance parameter $\eta$. Second, even if the value of $\eta$ is given, the value function is not readily available in the closed form and must be approximated by simulation. Here I focus on the first complication as if the moment function were readily available, leaving the description of the simulation algorithm for Section (ref).

A possible, though naive estimator of the value function can be constructed as follows. Consider an ideal scenario where the researcher knows mileage's law of motion. Then the choice probability $\gamma$ is the only unknown nuisance parameter that appears in the value function. Given an i.i.d sample $(D_i)_{i=1}^N$ from the law $P_D$, it is split into a main sample $J_2$ and an auxiliary sample $J_1$ of equal size $n =[N/2]$ such that $J_1 \sqcup J_2=\{1,2,\dots,N\}$. After that, the estimator of the value function $V(w;\theta;\sigma^{*};\gamma) $ is constructed as the sample average:

align*[align* omitted — 150 chars of source]

where $\widehat{\gamma}$ is estimated on the auxiliary sample $J_1$. Unfortunately, for some parameter values $\theta \in \Theta$ this estimator has slower than $\sqrt{N}$ convergence:

align[align omitted — 167 chars of source]

Therefore, the estimator $\widehat{\Theta}_I$ of the identified set $\Theta_I$ based on the moment function ((ref)) has suboptimal convergence rates.

The source of the slow convergence ((ref)) can be understood through the following decomposition:

align*[align* omitted — 621 chars of source]

The term $\mathbf{a}$ is the centered sample average of the value function $V(w;\theta;\sigma^{*};\gamma_0) $ evaluated at the true value $\gamma_0$ of the choice probability. Due to the sample splitting, the term $\mathbf{c}$ is a centered sample average of conditional on $J_1$ i.i.d random variables and is well-behaved. The term $\mathbf{b}$ stands for the bias of the value function $V(w;\theta;\sigma^{*}; \widehat{\gamma}) $ coming from the biased estimation of the conditional choice probability $\gamma$. This term is responsible for the slow convergence ((ref)).

The divergence of $\mathbf{b}$ comes from the combination of two facts: biased estimation of the first stage and the transmission of the bias from the first to the second stage. Since the conditional choice probability $\gamma(w)$ is a function of a high-dimensional state vector $w$, bias of its machine learning estimate (e.g., $\ell_1$-regularized logistic regression) converges slower than root-$N$. Because the value function $V(w;\theta;\sigma^{*};\gamma) $ is sensitive to this bias, it carries over into the second stage.

\paragraph{Overcoming the Regularization Bias using Orthogonalization.}

To overcome the translation of the first stage bias into the second stage I add the bias correction term to the value function $V(w;\theta;\sigma^{*};\gamma) $ to make it insensitive with respect to the biased estimation of $\gamma$. The new moment function for the value function $V(w;\theta;\sigma^{*};\gamma)$ takes the form

align*[align* omitted — 141 chars of source]

where the function $\Gamma(w;\theta)$ for $\theta = (R,\mu)$ is

align[align omitted — 286 chars of source]

Because the bias correction term has zero mean

align*[align* omitted — 216 chars of source]

the new moment function can replace the old one in the definition of the identified set. Moreover, the new function is insensitive to the bias estimation of $\gamma$. As a result, under additional mild regularity conditions, the regularization bias of the choice probability $\widehat{\gamma}(w)$ does not translate into the bias of the moment function:

align*[align* omitted — 109 chars of source]

\paragraph{The Role of Sample Splitting in Preventing Overfitting.} Another key aspect of the proposed analysis is using different samples $J_1$ and $J_2$ for different stages of estimating the value function. Had I used the whole sample to estimate the conditional choice probability $\widehat{\gamma}(w)$, the sample average

align*[align* omitted — 126 chars of source]

would not be a sample average of the i.i.d (or weakly dependent) random observations. The relation between the first stage error $\widehat{\gamma}(w_i) - \gamma_0(w_i)$ and the value of the moment function $ g(D_i, \theta, \gamma_0)$ creates bias, referred to as overfitting bias. To control the overfitting bias in the worst-case scenario:

align*[align* omitted — 192 chars of source]

one must impose complexity constraints on the class of functions $\mathcal{G}$ that are used to estimate the conditional choice probability. While some nonparametric estimators designed for low-dimensional state spaces obey these constraints, some modern machine learning estimators designed for high-dimensional state variables do not. To accommodate machine learning estimators at the first stage, I use different samples.

\paragraph{Bias Correction Terms for Examples (ref) and (ref)}. Suppose a nuisance parameter can be presented as a conditional expectation $\mathbb{E}[U|w]$. Then bias correction term for conditional choice probability takes the form $$ \alpha(D;\theta;\xi) = \Pi (w;\theta) ( U - \mathbb{E} [U|w]),$$ where $\xi$ is an unknown vector-valued function of the state variable $w$. The true value $\xi_0 = \xi_0(\theta)$ consists of the original conditional expectation function $\mathbb{E} [U|w]$ and the function $ \Pi (w;\theta) $

align[align omitted — 101 chars of source]
remark[Example (ref), continued] Consider the setup in Example (ref). The nuisance parameter $\eta = (\gamma,\rho)$ consists of the conditional choice probability $\gamma$ and the transition function $\rho$ defined in ((ref)). The transition function $\rho$ is present in the value function evaluated for the optimal $\sigma^{*}$ and the suboptimal $\sigma$ strategies. The conditional choice probability enters $V(w;\theta;\sigma^{*};\eta)$ only through the optimal strategy $\sigma^{*}$, described in ((ref))-((ref)). To sum up, the bias correction term for Example (ref) is \begin{align} \alpha(D;\theta;\xi) &= \alpha^{TRANS}_{\sigma}(D;\theta;\xi) - \alpha^{TRANS}_{\sigma^{*}}(D;\theta;\xi) - \alpha^{CCP}(D;\theta;\gamma), \end{align} where e.g., $\alpha^{TRANS}_{\sigma}(D;\theta;\xi)$ is the individual bias correction term for $\rho$ for the case of suboptimal strategy $\sigma$. As discussed above, the bias correction term $\alpha^{CCP}(D;\theta;\gamma)$ that corrects the bias of the conditional choice probability $\gamma_0$ is \begin{align} \alpha^{CCP}(D;\theta;\gamma) &= \frac{1}{1-\beta} \Gamma(w;\theta) (1_{\{ a=1\}} - \gamma(w)), \end{align} where the function $\Gamma(w,\theta)$ is given in ((ref)). The form for the bias correction term for the transition function $\rho$ is the same regardless whether $\sigma$ is optimal or not. Let $\tilde{\sigma} \in \{ \sigma, \sigma^{*}\}$ be a Markov policy. Suppose the state vector $w$ has a stationary distribution. Then the bias correction term of the value function $V(w;\theta;\sigma;\eta)$ for the transition function $\rho(\cdot)$ is equal to: \begin{align} \alpha^{TRANS}_{\tilde{\sigma}}(D;\theta;\xi):= \frac{\beta}{1-\beta} \mathbb{E} [\frac{d V(x;\theta;\tilde{\sigma};\eta_0)}{dx}|_{x = w_{next}}|w,a=1 ] ( s_{next} - \rho(w)), \end{align} where $\xi$ is an unknown vector-valued function of the state variable $w$. Its true value $\xi_0 = \xi_0(\theta)$ consists of the original nuisance parameter $\eta_0$ and the function $\Pi_0(w,\theta)$: $$\Pi_0(w,\theta):= \mathbb{E} [\frac{d V(x;\theta;\sigma;\eta_0)}{dx}|_{x = w_{next}} |w ],$$ which is equal to the expectation of the derivative of the value function with respect to the state variable evaluated for the future state conditional on the current state $w$. The bias correction term for other suboptimal Markov policies is more complicated, but can be derived using the argument of Appendix (ref).
remark[Example (ref), continued] Consider the setup in Example (ref). The nuisance parameter $\eta = (\gamma_A, \gamma_P)$ consists of the conditional choice probabilities of Apple and Player 2. The choice probability $\gamma_A$ enters the value function $V(w;\theta;\sigma^{*};\eta)$ only through the optimal strategy $\sigma^{*}$, described in ((ref))-((ref)). The choice probability $\gamma_P$ enters the value function $V(w;\theta;\tilde{\sigma};\eta), \quad \tilde{\sigma} \in \{ \sigma^{*}, \sigma \}$. To sum up, the bias correction term for Example (ref) is equal to: \begin{align*} \alpha(D;\theta;\eta):= - \alpha^{CCP}_{A}(D;\theta;\gamma) + \alpha^{\sigma}_{P}(D;\theta;\gamma_{P}) - \alpha^{\sigma^{*}}_{P}(D;\theta;\gamma_{P}), \end{align*} where e.g. $ \alpha^{CCP}_{A}(D;\theta;\gamma) $ is the individual bias correction term for $\gamma_A$. Let $\tilde{\sigma} \in \{ \sigma, \sigma^{*}\}$ be a Markov policy. Suppose the state vector $w$ has a stationary distribution. Then the bias correction term of the value function $V(w;\theta;\sigma;\eta)$ for the $\gamma_P$ is equal to: \begin{align*} \alpha^{\sigma}_{P}(D;\theta;\gamma_{P}):= \Gamma^{\tilde{\sigma}}_{P}(w;\theta) (1_{\{ a_{P}=1\}}-\gamma_{P}(w)), \end{align*} where \begin{align*} \Gamma^{\tilde{\sigma}}_P(w;\theta) = \frac{1}{1-\beta}( \delta_0 + (\delta_1 - \delta_0) \gamma^{\tilde{\sigma}}(w)), \end{align*} where $\gamma^{\tilde{\sigma}}(w) = {\mathrm{P}} (\tilde{\sigma}(w,\epsilon) = 1|w)$ is the conditional probability of Apple's decision under the policy $\tilde{\sigma}$. The bias correction term for Apple's conditional choice probability $\gamma_A$ is \begin{align*} \alpha^{CCP}_A(D;\theta;\gamma) &= - \mu s + R + (\delta_1 - \delta_0) \gamma_A(w) + \beta \mathbb{E} [V (w_{next};\theta;\sigma^{*};\eta_0) |w,a=1]\\ &- \beta \mathbb{E} [V (w_{next};\theta;\sigma^{*};\eta_0) |w,a=0]. \end{align*} To sum up, the bias correction term $\alpha(D;\theta;\eta)$ is equal to: \begin{align*} \alpha(D;\theta;\eta):= - \alpha^{CCP}_{A}(D;\theta;\gamma) + \frac{1}{1-\beta} (\delta_1 - \delta_0) (\gamma_{0}^{\sigma}(w) - \gamma_{A,0}(w)) (1_{\{a_{P} =1\}} -\gamma_P(w)). \end{align*}

In some point-identified problems the value function $V(w;\theta;\sigma;\eta)$ is only evaluated at the optimal strategy $\sigma^{*}$ and the true parameter value $\theta_0$. Then the bias correction term is evaluated only at the true value $\theta_0$. In particular, in the Examples (ref) and (ref) the function $\Gamma(w,\theta_0)$ can be further simplified as follows:

align*[align* omitted — 431 chars of source]

where we have used ((ref)) in the second line.

Overview of the Asymptotic Results

I will now introduce the estimator of the identified set $\Theta_I$, leaving the formal definition to ((ref)). Suppose there exists a function $g(D,\theta,\xi)$ that preserves the expectation of the moment function $m(w,\theta,\eta)$

align*[align* omitted — 74 chars of source]

and is insensitive to the biased estimation of its own nuisance parameter $\xi$ around its true value $\xi_0=\xi_0(\theta)$, where $\xi_0(\theta)$ is an identified vector-valued parameter of the distribution $P_D$ for each $\theta \in \Theta$. In many relevant cases such as Example (ref), $\xi_0$ contains the original nuisance parameter $\eta$, but may contain more unknown parameters of the distribution $P_D$.

To estimate the identified set $\Theta_I$ I represent it as the minimizer of the criterion function $Q(\theta,\xi_0)$

align[align omitted — 91 chars of source]

where $Q(\theta,\xi_0)$ is

align*[align* omitted — 78 chars of source]

and its sample analog is

align[align omitted — 133 chars of source]

My goal is to use different samples in the first and the second stages in order to avoid overfitting. Yet, simple sample splitting has a drawback that only one half of the sample is used for second-stage estimation, which can lead to loss of efficiency in small samples. In order to use the whole sample for the second stage yet keep the sample splitting idea, I use cross-fitting procedure described below.

definition[Cross-fitting] \begin{enumerate} • For a random sample of size $N$, denote a $K$-fold random partition of the sample indices $[N]=\{1,2,...,N\}$ by $(J_k)_{k=1}^\mathcal{K} $, where $\mathcal{K}$ is the number of partitions and the sample size of each fold is $n = N/\mathcal{K}$. Also for each $k \in [\mathcal{K}] = \{1,2,...,K\}$ define $J_k^c = \{1,2,...,N\} \setminus J_k$. • For each $k \in [\mathcal{K}]$, construct an estimator $ \widehat{\xi}( V_{i \in J_k^c})$ of the nuisance parameter value $\xi_0$ using only the data from $J_k^c$. For any observation $i \in J_k$, define an estimated signal $\widehat{\xi}_i := \widehat{\xi}( V_{i \in J_k^c})$. \end{enumerate}
definition[Definition of the Set Estimator] Let $\widehat{\xi}= \widehat{\xi}(\theta)$ be the first-stage estimator of the nuisance parameter constructed in Definition (ref). Let the criterion function $Q_N(\theta,\xi)$ be as in ((ref)) and $\widehat{c}$ be a positive number. The estimator $\widehat{\Theta}_I$ of the identified set $\Theta_I$ is chosen as a contour set of the sample criterion function $Q_N(\theta,\xi)$ of level $c$: \begin{align*} \widehat{\Theta}_I:= {\cal C}_N(\widehat{c},\widehat{\xi}):= \{ \theta \in \Theta, \quad N Q_N(\theta,\widehat{\xi}) \leqslant \widehat{c} \}. \end{align*}

The contour level $\widehat{c}$, possibly data dependent, is chosen such that $ \widehat{\Theta}_I$ contains the true set $\Theta_I$ with probability approaching one

align[align omitted — 142 chars of source]

I establish convergence and inference properties of the contour set estimator $\widehat{\Theta}_I$. The first property is formulated in terms of the convergence rate of $\widehat{\Theta}_I$ to $\Theta_I$ is based on the notion of Hausdorff distance $d_H(\widehat{\Theta}_I,\Theta_I)$:

align*[align* omitted — 179 chars of source]

The set $\widehat{\Theta}_I$ is said to converge to $\Theta_I$ at rate $\epsilon_N$ if the Hausdorff distance between the sets converges at rate $\epsilon_N$:

align[align omitted — 95 chars of source]

Under mild regularity conditions it is possible to achieve the nearly efficient rate $\epsilon_N = O(\sqrt{\frac{ \log N}{N}})$.

To conduct inference, I fix a confidence level $\tau \in (0,1)$. A confidence region $C_N(c_{\tau},\widehat{\xi})$ of level $\tau$ is defined as a contour set $C_N(c_{\tau};\widehat{\xi})$ of level $c_{\tau}$ such that $C_N(c_{\tau};\widehat{\xi})$ contains $\Theta_I$ with probability at least $1-\tau$:

align[align omitted — 138 chars of source]

I construct a confidence region $ C_N(\widehat{c}_{\tau};\widehat{\xi})$, where $\widehat{c}_{\tau}$ is a consistent estimator of the $\tau$-quantile, denoted $c_{\tau}$, of the inferential statistic:

align[align omitted — 96 chars of source]

The consistent estimator of the $\tau$-quantile is chosen by the subsampling algorithm defined below.

definition[Subsampling Algorithm] Partition the sample $(D_i)_{i=1}^N$ into $B_N= o(\sqrt{N})$ subsamples $V_j,\quad j \in \{1,2, \dots, B_N\}$ of equal size $b:=[N/B_N]$. Compute the sample criterion function $Q_{j,b}(\theta;\widehat{\xi})=\| \frac{1}{b} \sum_{ i \in V_j} g(D_i, \theta, \widehat{\xi}_i)\|_{+}^2.$ Choose the level $\widehat{c}$ of the order $\widehat{c} \sim \log N$. Report $\widehat{c}_{\tau}$ as the $\tau$-quantile of the sample of statistics $$\{ \sup_{\theta \in {\cal C}_N(\widehat{c},\widehat{\xi})} b Q_{j,b}(\theta,\widehat{\xi}) , j = 1,2, \dots, B_N\}.$$

Section (ref) establishes the asymptotic validity of the set estimator given in Definition (ref) and subsampling algorithm of Definition (ref)

Dynamic Game of Imperfect Information

Consider the dynamic model of strategic interaction from BBL. There are $K$ players, denoted by $\{1,2,\dots, K\}$. Each player $k$ makes a decision $a_k \in \mathcal{A}_k$ from a finite set of discrete alternatives $\mathcal{A}_k$ at time periods $t \in \{0,1,\dots, \infty\}$. In each period $t$ the players commonly observe a vector of state variables $w_t \in \mathcal{W} \subset \mathcal{R}^{d_w}$. Given the state variable $w_t=w$, players choose actions simultaneously. Before choosing his action, each player $k$ observes a vector of private shocks $(\epsilon_k(j))_{j \in \mathcal{A}_k}$ corresponding to each discrete alternative $j$ in his choice set $\mathcal{A}_k$. The transition between states follows a conditional probability distribution: $P(\cdot| a,w) $ conditional on the current state $w$ and the action profile $a=(a_1,a_2,\dots,a_K)$.

I focus on the structural parameter describing the utility of the first player. I assume that his per-period utility function is given by

align[align omitted — 112 chars of source]

where $a=(a_1,a_2,\dots,a_K)$ is the profile of the players' actions, $a_1 \in {\mathcal{A}_1}$ is the action of the first player, $w$ is the state variable, and $\theta_0$ is the true value of the structural parameter $\theta$. The per-period utility is presented as the sum of a deterministic component $\tilde{\pi}(a,w;\theta;\zeta)$ and the private shock $\epsilon_1(a_1)$. I allow $\tilde{\pi}(a,w;\theta;\zeta)$ to depend on identified nuisance parameter $\zeta$ whose true value is $\zeta_0$.

I assume that each player follows a pure Markov policy. The pure Markov policy for player one $\sigma_1(w,\epsilon_1):\mathcal{W} \bigtimes \mathcal{E} \rightarrow \mathcal{A}_1$ maps the current state $w$ and the private shock of player one, $\epsilon_1$, into the action space $\mathcal{A}_1$. If the behavior of the players is described by a Markov policy profile $\sigma = (\sigma_1,\sigma_2,\dots,\sigma_K)$, the value function is given by

align[align omitted — 139 chars of source]

A strategy profile $\sigma^{*}$ is a Markov perfect equilibrium if each player $k$ prefers its strategy $\sigma^{*}_k$ to all alternative Markov strategies as long as the others follow the equilibrium strategy $\sigma^{*}_{-k}$. That is, the value function $ V(w;\theta_0;\sigma^{*};\eta_0)$ of the first player at the strategy $\sigma^{*}$ is weakly larger than the value function of any other strategy profile $\sigma = (\sigma_1,\sigma_{-1}^{*})$:

align[align omitted — 148 chars of source]

where $\sigma = (\sigma_1,\sigma_{-1}^{*})$ consists of a feasible suboptimal alternative for player one, $\sigma_1$, and the equilibrium profile for the other players $\sigma_{-1}^{*}$. I assume that the each observation comes from the same Markov perfect equilibrium $\sigma^{*}$, although I do not provide conditions for the existence of such equilibrium and recognize that there could be many such equilibria.

assumption[Equilibrium Selection] The data are generated by a single Markov perfect equilibrium $\sigma^{*}$.

Casting problem as a moment inequality with a first-stage nuisance parameter

Define the choice-specific value function of player one $v(a_1,w)$ as the expected present value conditional on the current state $w$ and the choice $a_1$

align*[align* omitted — 149 chars of source]

Then, following the optimal strategy $\sigma^{*}(w,\epsilon)$, the first player chooses $a_1$ if and only if the total utility of $a_1$ is not smaller than the total utility of any other choice $a_1' \in \mathcal{A}_1$

align*[align* omitted — 80 chars of source]

Therefore, the optimal strategy of player one is characterized as

align*[align* omitted — 117 chars of source]

or, equivalently,

align*[align* omitted — 120 chars of source]

As discussed in HotzMiller, under Assumption (ref) the vector of differences of the choice-specific value functions can be expressed as

align[align omitted — 113 chars of source]

where ${\mathrm{P}} (a_1|w)$ is the probability of the choice $a_1 \in \mathcal{A}_1$ conditional on the state variable $w$. Finally, as a suboptimal Markov policy $\sigma$, I consider a cutoff-type strategy

align[align omitted — 152 chars of source]

where I add a deviation function $\text{dev}(a_1,w)$ for each element $v(a_1,w) - v(1,w)$. As a normalization condition, I set $\text{dev}(1,w)=0 \quad \forall w$. Therefore, the inequality ((ref)) depends on the nuisance parameter

align[align omitted — 126 chars of source]

that consists of any original nuisance parameter $\zeta$ that may appear in the utility function ((ref)), the conditional choice probabilities of all players $$\gamma_{jk} (w):= {\mathrm{P}}(a_k=j|w), \quad j \in \mathcal{A}_k, k \in \{1,2,\dots,K\},$$ and the conditional distribution ${\mathrm{P}} (w_{next}|w,a)$.

I construct the identified set $\Theta_I$ using a subset of inequalities implied by the equilibrium definition (i.e, ((ref))). Let $q(w)$ be an $L$-vector of non-negative weighting functions. Let $\text{dev}(a_1,w): \mathcal{A}_1 \bigtimes \mathcal{W} \rightarrow \mathcal{R}^d$ be an $ L$-vector of deviation functions whose coordinate $l \in \{1,2,\dots,L\}$ corresponds to a deviation strategy $\text{dev}_l(a_1,w)$. Define the moment vector-valued function as

align[align omitted — 119 chars of source]

whose component $m_l(D,\theta,\eta)$ is equal to the weighted difference $V(w;\theta;\sigma_l;\eta) - V(w;\theta;\sigma^{*};\eta)$: $$ m_l(w,\theta,\eta) :=q_l(w) (V(w;\theta;\sigma_l;\eta) - V(w;\theta;\sigma^{*};\eta)), \quad l \in \{1,2,\dots,L\}$$ evaluated at a strategy profile $\sigma_l = (\sigma_{1,l}, \sigma^{*}_{-1})$. The suboptimal strategy of the first player $\sigma_{1,l}$ is given in ((ref)) with the deviation function $\text{dev}_l(a_1,w)$. The identified set $\Theta_I$ is defined as the collection of parameter values $\theta$ that obey inequality restrictions in expectation

align[align omitted — 120 chars of source]

where the expectation is taken with respect to the unconditional distribution of the state $w$.

Simulation estimator of the value function

The value function $V(w;\theta;\sigma;\eta)$ appearing in the moment function ((ref)) is not available in closed form. However, it can be approximated by Monte Carlo simulation as in BBL. In the first stage, one constructs an estimate $\widehat{\eta}$ of the nuisance parameter $\eta$ that is defined in ((ref)). In the second stage, one simulates the sequence of states and shocks from the estimated first stage parameter $\widehat{\eta}$ and averages the realization of the value function over multiple simulation draws. A single simulation draw is given by Algorithm (ref).

algorithm[algorithm omitted — 1,030 chars of source]

In order to evaluate the value function $V(w;\theta;\sigma;\eta)$ at different parameter values $\theta_1 \in \Theta$ and $\theta_2 \in \Theta$, I use the same simulation draws. As long as the number of simulation draws is large enough, the simulation error does not affect the asymptotic properties of the estimator of the identified set $\widehat{\Theta}_I$ as discussed in Pakes:1989.

Bias Correction Term for the Expected Value Function

When the nuisance parameter $\eta$ is high-dimensional and estimated by machine learning, the plug-in estimator of value function is biased. To make the moment function ((ref)) insensitive to first-stage bias, I derive bias correction term. As discussed in Newey1994, the bias correction term for a vector-valued nuisance parameter is equal to the sum of individual terms of individual components. Moreover, the bias correction term for the conditional probability of choice $j$ has the product structure

align*[align* omitted — 104 chars of source]

Furthermore, according to Newey1994, the function $ \Gamma(w;\theta)$ is implicitly defined by the orthogonality condition explained below.

Let $g(D,\theta,\xi)$ be a moment function. Define the Gateaux derivative map $D_r: \Xi \bigtimes \Theta \rightarrow \mathcal{R}^L$ as

align*[align* omitted — 125 chars of source]

for all $r \in [0,1)$, which I assume exists. I also denote the pathwise derivative of the expected moment function at the true value $\xi_0$

align*[align* omitted — 134 chars of source]
definition[Neyman orthogonality of moment function] The moment function $g(D;\theta;\xi)$ obeys the orthogonality condition at $\xi_0$ with respect to the nuisance realization set $\Xi_N \subset \Xi$ if the pathwise derivative $D_r[\xi - \xi_0]$ exists for all $r \in [0,1)$ and vanishes at $r=0$ for each $\theta \in \Theta$ \begin{align} \partial_{\xi} \mathbb{E} g(D,\theta,\xi_0)[\xi - \xi_0] = 0, \quad \xi \in \Xi, \theta \in \Theta. \end{align}

The definition of Neyman orthogonality requires that the moment function be insensitive to the biased estimation of $\xi$. This condition is the generalization of the orthogonality condition for point-identified models in doubleml2016 that is required to hold only at the true value $\theta_0$ of identified parameter $\theta$. In contrast to the point-identified case, I require the equality in ((ref)) to hold at each point $\theta$ of the parameter space $\Theta$. In many relevant cases (e.g., if the moment function $ g(D,\theta,\xi)$ is linear in $\theta$), the orthogonality condition ((ref)) on the set $\Theta$ follows from the orthogonality ((ref)) on a finite subset of $\Theta$. Rewriting the orthogonality condition ((ref)) for the bias correction term for the value function gives

align*[align* omitted — 152 chars of source]

I find the function $\Gamma(w;\theta)$ from the recursive definition of the value function ((ref)). In the case of Example (ref) the recursive definition can be rewritten in the unconditional form

align[align omitted — 424 chars of source]

where $\gamma_0(w)$ is the true value of the conditional choice probability and $\theta = (R,\mu)$ consists of the replacement and maintenance costs. The unknown function $\gamma(w)$ appears in Equation ((ref)) both {\bf outside} the value function $V(w;\theta;\sigma^{*};\gamma) $ and {\it inside} of this function. I consider a local deviation of the choice probability $\gamma(\cdot)$ from its value $\gamma_0(\cdot)$ for each value of the parameters $R$ and $\mu$. Applying chain rule to the Equation ((ref)) yields a pathwise derivative

align*[align* omitted — 616 chars of source]

where the last equality follows from the logistic distribution of the private shocks (Assumption (ref)).

In what follows I derive the bias correction terms for the other nuisance parameters that are present in the dynamic discrete choice in BBL. Let $q(w)$ be a weighting function. Define

align*[align* omitted — 60 chars of source]

as the expectation of the weighting function $q(w)$ evaluated at the current state $w$ conditional on the future state $w_{next}$. Denote the expectation of the choice $j$ made by player one conditional on the state $w$ by $\gamma_j(w):= \mathbb{E} [1_{\{a_1=j \}}|w]$. Let $$\gamma(w):= (\gamma_2(w),\gamma_3(w),\dots, \gamma_{A_1}(w))$$ be the $A_1-1$ vector of these probabilities. Define the conditional expectation of the current private shock evaluated at the equilibrium strategy $\sigma_1^{*} = \sigma_1^{*}(w,\epsilon,\gamma)$ as

align[align omitted — 122 chars of source]

Lemma (ref) gives the bias correction term, $ \alpha^{CCP}_{j}(D;\theta;\gamma_j)$, that makes the value function $V(w;\theta;\sigma^{*};\eta)$ insensitive to the biased estimation of $\gamma_j$.

lemma[Bias correction term for own conditional choice probability] Suppose the state variable $w$ has a stationary distribution and Assumption (ref) holds. Suppose the conditional probability of each choice $\gamma_j(w), \quad j \in \{2,3,\dots,J\}$ is bounded away from zero and one. Then the bias correction term $\alpha^{CCP}_{j}(D;\theta;\gamma_j)$ is \begin{align} \alpha^{CCP}_{j}(D;\theta;\gamma_j) &= \frac{q(w)}{q(w) - \beta \lambda(w)} \big( \mathbb{E}_{\epsilon_{-1}} [\tilde{\pi} ((j,\sigma^{*}_{-1}(w,\epsilon_{-1}));w;\theta) - \tilde{\pi}(1,\sigma^{*}_{-1}(w,\epsilon_{-1}));w;\theta)] \nonumber \\ &+ \beta \mathbb{E}[ V(w_{next};\theta;\sigma^{*};\eta_0) | w,(j,\sigma^{*}_{-1}(w,\epsilon_{-1}))] \nonumber \\ &- \beta \mathbb{E}[ V(w_{next};\theta;\sigma^{*};\eta_0) | w,(1,\sigma^{*}_{-1}(w,\epsilon_{-1}))] \nonumber \\ &+ \partial_{\gamma_j}PS_{\sigma_1^{*}}(\gamma_0) \big) (1_{a_1=j} - \gamma_j(w)) \end{align} where $\partial_{\gamma_j}PS_{\sigma_1^{*}}(\gamma)$ is the partial derivative of ((ref)) with respect to $\gamma_j$.
lemma[Bias Correction Term for the Law of Motion of the State Variable] Suppose the high-dimensional state variable $w$ has a stationary distribution. Let $a=(a_1,a_2,\dots,a_K)$ be a given profile of actions. Then the bias correction term for the conditional quantile function $Q(u,w,a)$ of level $u$ is \begin{align} \alpha^{TRANS}_{\sigma}(D;\theta;\eta) &= \frac{\beta l(w)}{l(w) - \beta \lambda(w)} \mathbb{E}[ \nabla_{w_{next}} V (w_{next};\theta;\sigma;\eta_0)|w, a] \prod_{k=1}^K {\mathrm{P}} (\sigma_k(w,\epsilon_k)=a_k|w) \\ &\frac{1_{w_{next} \leqslant Q(u,w,j) - u} }{f(w_{next}|w,a=\sigma(w,\epsilon))} \nonumber. \end{align}
lemma[Bias Correction Term for the opponent's conditional choice probability] Suppose the state variable $w$ has a stationary distribution. Let the number of players be $2$ (i.e, $K=2$). Then the bias correction term for the conditional probability of the choice $j_2$ made by my opponent is \begin{align} \nonumber \alpha_{j}^{CCP,op}(D;\theta;\gamma_{j_2}) &= \frac{q(w)}{q(w) - \beta \lambda(w)} \big( \sum_{j_1=1}^{A_1} (\tilde{\pi} ((j_1,j_2),w;\theta) -\tilde{\pi} ((j_1,1),w;\theta )) {\mathrm{P}} (\sigma_1(w,\epsilon_1) = j_1|w) \\ &+ \beta \big[ \sum_{j_1=1}^{A_1} \mathbb{E} [V(w_{next};\theta;\sigma^{*};\eta_0) | w, (j_1,j_2)] - \sum_{j_1=1}^{A_1} \mathbb{E} [V(w_{next};\theta;\sigma^{*};\eta_0) | w, (j_1,1)] \big] \cdot \nonumber \\ &\cdot {\mathrm{P}} (\sigma_1(w,\epsilon_1) = j_1|w)) (1_{a_2 = j_2} - {\mathrm{P}}(a_2=j_2) ). \end{align}

Using Linearity to Reduce Computation

Both the value function and the bias correction terms presented above are not available in closed form and must be simulated. When the value function is a linear function of $\theta$, the simulation can be simplified. Suppose there exist basis functions $\Psi_1(w;\sigma;\eta), \Psi_2(w;\sigma;\eta) $ such that

align[align omitted — 108 chars of source]

is an affine function of $\theta$. Then one can simulate the basis functions $\Psi_1(w;\sigma;\eta), \Psi_2(w;\sigma;\eta) $ instead of simulating the value function for each $\theta \in \Theta$. Lemma (ref) provides the sufficient conditions for the linearity of the value function and the individual bias correction terms provided in Lemmas (ref), (ref), (ref).

lemma[Sufficient Conditions for Linearity] The following conditions hold. (1) The per-period utility function given in ((ref)) is a linear function of $\theta$. (2) The distribution of the private shock for each player is known. Then there exists a vector of basis functions $\Psi(w;\sigma;\eta)$ such that ((ref)) holds. Furthermore, each individual bias correction term given in ((ref)), ((ref)), and ((ref)) is also a linear function $\theta$.

Asymptotic Theory

Suppose I have a collection of inequality restrictions on an economic model coming from the data structure and/or the assumptions about the data generating process. These restrictions are embodied into a moment function $g(D,\theta,\xi): \mathcal{D} \bigtimes \Theta \bigtimes \Xi \rightarrow \mathcal{R}^L$, where $\theta \in \Theta \subset \mathcal{R}^d$ is the target parameter. In addition to the target parameter $\theta$, the moment function $g(D,\theta,\xi)$ depends on a nuisance parameter $\xi=\xi(\theta)$ whose true value $\xi_0=\xi_0(\theta)$ is an identified parameter of the data distribution $P_D$ and belongs to a convex subset of a normed vector space $ \Xi$. The object of interest is an identified set $\Theta_I$ defined as a collection of parameter values $\theta$ that satisfy the inequality restrictions

align[align omitted — 119 chars of source]

at the true value $\xi_0$ of the nuisance parameter.

The identified set $\Theta_I$ is characterized as the minimizer of the criterion function

align*[align* omitted — 145 chars of source]

I assume that the following partial identification condition holds. There exist positive constant $C>0$ and $\delta>0$ such that

align[align omitted — 121 chars of source]

This condition states that once $\theta$ is bounded away from $\Theta$, the moment $\mathbb{E} g(D,\theta, \xi_0 )$ is bounded away from $\Theta_I$ by a number that is proportional to the Hausdorff distance $d_H(\theta,\Theta_I)$ from $\theta$ to the identified set $\Theta_I$. This condition ensures that the true moment function $ g(D,\theta, \xi_0) $ distinguishes the boundary of the identified set.

\paragraph{Impact of the First-Stage Estimation.} In the next condition I introduce a sequence of neighborhoods $\Xi_N^{\theta} \subset \Xi^{\theta}$ of $\xi_0(\theta)$ that contain the estimate $\widehat{\xi}(\theta)$ of $\xi_0(\theta)$ w.p. approaching one. As the sample size $N$ increases, the neighborhoods shrink. The quality of the estimation of the first-stage parameter is defined as the speed of shrinkage of the neighborhood $\Xi_N(\theta)$ around $\xi_0(\theta)$. I refer to it as the first-stage rate $g_N$. Finally, I assume that the second-order derivative of the functional $\mathbb{E} g(D,\theta, \xi)$ is well-behaved. Combined with the orthogonality condition and the upper bound on the first-stage rate $g_N$, this assumption ensures that I can ignore the impact of the estimation error $\widehat{\xi}(\theta)-\xi_0(\theta)$ on the second and the higher-order derivatives of the moment function $\mathbb{E} g(D,\theta, \xi)$ with respect to $\xi$ at $\xi_0$.

condition[Orthogonality] There exists a sequence $\Xi_N^{\theta}$ of subsets of $\Xi^{\theta}$: $\Xi_N^{\theta} \subset \Xi^{\theta}$ such that the following conditions hold. (1) The true value $\xi_0(\theta)$ belongs to $\Xi_N^{\theta}$ for all $N \geqslant 1$. (2) There exists a sequence of numbers $\phi_N=o(1)$ such that with probability at least $1-\phi_N$, $\widehat{\xi}(\cdot)$ belongs to $\Xi:=\bigtimes \Xi^{\theta}$ uniformly over $\theta \in \Theta$. (3) The set $\Xi_N^{\theta}$ shrinks around $\xi(\theta)$ at the following statistical rate $g_N = o(N^{-1/4})$ uniformly over $\theta$: $$ \sup_{\theta \in \Theta} \sup_{\xi(\theta) \in \Xi^{\theta}} \| \xi(\theta) - \xi_0(\theta)\|_{P,2} \leqslant g_N.$$ The moment function $g(D,\theta,\xi)$ obeys the orthogonality condition at $\xi_0$. (5) There exists a sequence $s_N = o(N^{-1/2})$ such that the second Gateaux derivative of the functional $G(\theta,\xi(\theta))$ with respect to $\xi$ at $\xi_0$ is bounded: $$ \sup_{\theta \in \Theta} \sup_{\xi \in \Xi_N^{\theta}} \sup_{r \in [0,1)} \| \partial_r^2 \mathbb{E} g(D,\theta, r(\xi - \xi_0) + \xi_0) \| \leqslant s_N.$$

The next condition requires that the moment function $g(D;\theta;\xi(\theta))$ is sufficiently regular with respect to $\theta$ for each fixed element $\xi \in \Xi$ of the nuisance realization set $\Xi$. I consider a class $$\mathcal{F}_{\xi}:= \{ g(\cdot,\theta, \xi(\theta)), \theta \in \Theta \}$$ and require the uniform covering entropy of this class to be bounded.

condition[Regularity of Moment Function] The following conditions hold. (1) There exists a measurable envelope function $F_{\xi}=F_{\xi}(D)$ that bounds all elements in the function class almost surely $$ \sup_{\theta \in \Theta} | g_l(D,\theta, \xi)| \leqslant F_{\xi} (D) \text{ a.s. }, \quad l \in \{1,2,\dots,L\}.$$ Moreover, the envelope $F_{\xi}$ has a finite $c$-norm for some $c>2$ $\| F_{\xi} \|_{P,c} := \left(\int_{D \in \mathcal{D}} |F_{\xi}(D)|^{c} \right)^{1/c} \leqslant c_1$. (2) There exist finite constants $a$ and $v$ such that the uniform covering entropy of the class $\mathcal{F}_{\xi}$ is bounded \begin{align} \sup_{\tilde{Q} } \log N(\epsilon \| \mathcal{F}_{\xi} \|_{\tilde{Q} ,2}, F_{\xi}, \| \cdot \|_{\tilde{Q} ,2} \leqslant v \log (a/\epsilon), for all 0 < \epsilon \leqslant 1. \end{align} (3) There exists a sequence $r_N'$ obeying $r_N' \log (1/r_N') = o(1)$ that is an upper bound for the following quantity $$ \sup_{\theta \in \Theta} \sup_{\xi \in \Xi^{\theta}} \left ( \mathbb{E} \| g(D, \theta, \xi(\theta)) - g(D, \theta, \xi_0(\theta)) \|^2 \right)^{1/2} \leqslant r_N'.$$

Conditions (ref)(1)-(2) are the generalization of the regularity assumption in the point-identified moment problem of doubleml2016. Because $\xi (\theta)$ depends on $\theta$, conditions (ref)(1)-(2) are non-standard and require verification in applications. However, this requirement is mild when the nuisance parameter $\xi (\theta)$ is a linear function of $\theta$. Suppose that the true value of the nuisance parameter $\xi_0(\theta)$ is a linear function of $\theta$

align[align omitted — 88 chars of source]

where $\xi_0^{a}(D)$ and $\xi_0^{b}(D)$ are the identified parameters of the distribution $P_D$. Then conditions (ref)(1)-(2) can be reformulated in terms of the nuisance parameters $\{ \xi_0^{a}, \xi_0^{b}\}$ that no longer depend on $\theta$. When ((ref)) holds, Conditions (ref) (1)-(2) are satisfied for many practical cases. In particular, the functions $\xi_0^{a}(D), \xi_0^{b}(D)$ can be estimated by $\ell_1$-regularized methods, random forests, and deep neural nets under plausible assumptions about their structure.

\paragraph{Donsker Property.} Let $\Theta'$ be an open neighborhood of $\Theta$. I require the moment function $g(D_i,\theta,\xi_0)$ to have a Donsker property defined as follows. In the metric space $L^{\infty}(\Theta')$,

align[align omitted — 176 chars of source]

where $\Delta(\theta)$ is a mean zero Gaussian process on $\Theta$ with a.s. continuous paths and $\text{Var}(\Delta(\theta))>0$ for each $\theta \in \Theta'$. In addition, the probability space $(\Omega, \mathcal{F}, \mathcal{P})$ is rich enough to support the representation ((ref)).

theorem[Estimation and Inference for Semiparametric Functional Inequalities] Suppose Conditions ((ref)), (ref), (ref), and ((ref)) hold. Let $\widehat{\Theta}_I$ be a contour set estimator of Definition (ref). Let $\widehat{c}$ be such that \begin{align} \widehat{c} \geqslant \sup_{ \theta \in \Theta_I} N Q_N(\theta, \xi_0) \quad w.p. \rightarrow 1 \end{align} holds. Then the Hausdorff distance $d_H(\Theta_I, \widehat{\Theta}_I)$ between the estimated set $\widehat{\Theta}_I$ and $\Theta_I$ converges at rate $O_{P}(\sqrt{(1\vee \widehat{c})/N})$: $d_H(\widehat{\Theta}_I, \Theta_I) = O_{P}(\sqrt{(1\vee \widehat{c})/N})$.

Theorem (ref) is my first main result. It establishes the sufficient conditions on the moment function to deliver the rate of convergence of $\widehat{\Theta}_I$ to $\Theta_I$. It suggests that the contour level $\widehat{c}$ as small as possible subject to the constraint ((ref)). Setting $\widehat{c} = O_{P} (1)$ subject to ((ref)) delivers the optimal rate, but this choice is infeasible. Setting $\widehat{c} \sim \log N$ delivers a nearly efficient rate.

In many cases it is possible to establish convergence without the requirement ((ref)). This is possible because the criterion function is degenerate ((ref)).

definition[Degeneracy] The following conditions hold. (1) There exists a sequence of subsets $\Theta_N$ of $\Theta$, which cannot depend on $\xi$, such that the criterion function $Q_N(\theta, \xi)$ vanishes on these subsets w.p. approaching one. That is, $\forall p > 0$ there exists $N_p$ such that for all $N \geqslant N_P$ $ \inf_{\xi \in \Xi_N} {\mathrm{P}} (Q_N(\theta, \xi) - \inf_{\theta \in \Theta} Q_N(\theta,\xi) = 0 \quad \forall \theta \in \Theta) \geqslant 1-p $. (2) These sets can approximate the identified set $\Theta_I$ in the Hausdorff distance sufficiently well: $d_H(\Theta_N,\Theta_I) \leqslant \epsilon_N$. (3) The sequence $\epsilon_N = O_{P} (N^{-1/2})$.
lemma[Sufficient Conditions for Degeneracy] Suppose Conditions ((ref)), (ref), (ref), and ((ref)) hold. In addition, there exist positive constants $C,M,\delta$ such that \begin{align} \max_{l} \mathbb{E} g_l(D,\theta, \xi_0 ) \leqslant - C (\epsilon \wedge \delta) \quad for all \theta \in \Theta_I^{-\epsilon}, \\ d_H(\Theta_I^{-\epsilon},\Theta_I) \leqslant M \epsilon for all \epsilon \in [0, \delta], \quad l \in \{1,2,\dots, L\} \nonumber \end{align} Then the criterion function $ Q_N(\theta, \xi)$ obeys the degeneracy condition in the sense of Definition (ref). Suppose the contour level $\widehat{c}$ obeys \begin{align} \widehat{c} \geqslant \min_{\theta \in \Theta} Q_N(\theta, \widehat{\xi}) \vee \frac{\log N}{\sqrt{N}} w.p. \rightarrow 1. \end{align} Then the Hausdorff distance $d_H(\Theta_I, \widehat{\Theta}_I)$ converges at rate $O_{P}({N}^{-1/2})$, where $\widehat{\Theta}_I$ is a contour level set as in Definition (ref).

Lemma (ref) is my second main result. It provides the sufficient conditions under which the identified set $\Theta_I$ can be estimated at the fastest possible rate $\epsilon_N = O_{P} (N^{-1/2})$. It requires that the sample criterion function $Q_N(\theta,\xi_0)$ be flat on the (possibly) data-dependent sets $\Theta_N$ that approximate the identified set $\Theta_I$ sufficiently well. If this requirements holds, the contour level $\widehat{c}_0:= \arg \min_{\theta \in \Theta_0} Q_N(\theta,\xi_0)$ delivers the optimal rate for the contour level set ${\cal C}_N(\widehat{c}_0, \xi_0)$ based on the true value of the nuisance parameter (see, e.g. CHT). We show that a modified choice $\widehat{c}$ given in ((ref)) delivers the optimal rate in the presence of the nuisance parameter $\widehat{\xi}$ estimated in the first stage on an auxiliary sample.

\paragraph{Subsampling.} I wish to construct the contour set ${\cal C}(\widehat{c}, \widehat{\xi})$ that has confidence region property ((ref)). To do this, I must find the asymptotic distribution of the inferential statistic $$ {\cal C}_N:= \sup_{\theta \in \Theta_I} Q(\theta; \xi_0)$$ that can be used to estimate the $\tau$-quantile of $ {\cal C}_N$. Define the random variable

align[align omitted — 120 chars of source]

where $I(\theta)$ is an $L$-vector of functions $I_l(\theta), l \in \{ 1,\dots,L\})$. The function $I_l(\theta)$ is defined as follows:

align*[align* omitted — 153 chars of source]

CHT show that ${\cal C}_N$ converges to $ \mathcal{C}$ in distribution, and that $ \mathcal{C}$ has a non-degenerate and continuous distribution function. The final requirement for the validity of the subsampling algorithm of Definition (ref) is that the inferential statistic

align*[align* omitted — 89 chars of source]

is well-behaved on the $\epsilon_N$-expansion of the identified set $\Theta_I$. A sufficient condition for this requirement is given below.

condition[Sufficient Conditions for Subsampling] (1) There exists a constant $C_{\max} < \infty$ such that the moment function is bounded: $\sup_{\theta \in \Theta} | \mathbb{E} g(D,\theta,\xi_0) | \leqslant C_{\max} d_H(\theta, \Theta_I)$. (2) The rates $s_N, r_N'$ and $\epsilon_N$ obey the following bound: $\sqrt{N} (s_N + r_N' \log (1/r_N') + N^{-1/2+1/c} ) \epsilon_N = o(1)$.
theorem[Validity of Subsampling for Moment Inequalities] Suppose Conditions ((ref)), (ref), (ref), ((ref)), (ref) hold. Suppose the number of subsamples satisfies $b = o (\sqrt{N})$, $b \rightarrow \infty$. Let $\tau $ be the desired coverage level. Then (1) the critical value $\widehat{c}$ of Definition (ref) converges in probability to the $\tau$-quantile of $\mathcal{C}$, where $\mathcal{C}$ given in ((ref)) and (2) ${\mathrm{P}} (\Theta_I \subseteq \mathcal{C}_N(\widehat{c}, \widehat{\xi})) = 1 - \tau$.