EconBase
← Back to paper

On the iterated estimation of dynamic discrete choice games

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

106,932 characters · 10 sections · 99 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the iterated estimation of dynamic discrete choice games

\sloppy

\hypersetup{pageanchor=false} \and Jackson Bunting\\Department of Economics\\Duke University\\ [email removed]}}

abstractWe study the first-order asymptotic properties of a class of estimators of the structural parameters in dynamic discrete choice games. We consider $K$-stage policy iteration (PI) estimators, where $K$ denotes the number of policy iterations employed in the estimation. This class nests several estimators proposed in the literature. By considering a “pseudo likelihood” criterion function, our estimator becomes the $K$-PML estimator in aguirregabiria/mira:2002,aguirregabiria/mira:2007. By considering a “minimum distance” criterion function, it defines a new $K$-MD estimator, which is an iterative version of the estimators in pesendorfer/schmidt-dengler:2008 and pakes/ostrovsky/berry:2007. First, we establish that the $K$-PML estimator is consistent and asymptotically normal for any $K \in \mathbb{N}$. This complements findings in aguirregabiria/mira:2007, who focus on $K=1$ and $K$ large enough to induce convergence of the estimator. Furthermore, we show under certain conditions that the asymptotic variance of the $K$-PML estimator can exhibit arbitrary patterns as a function of $K$. Second, we establish that the $K$-MD estimator is consistent and asymptotically normal for any $K \in \mathbb{N}$. For a specific weight matrix, the $K$-MD estimator has the same asymptotic distribution as the $K$-PML estimator. Our main result provides an optimal sequence of weight matrices for the $K$-MD estimator and shows that the optimally weighted $K$-MD estimator has an asymptotic distribution that is invariant to $K$. The invariance result is especially unexpected given the findings in aguirregabiria/mira:2007 for $K$-PML estimators. Our main result implies two new corollaries about the optimal $1$-MD estimator (derived by pesendorfer/schmidt-dengler:2008). First, the optimal $1$-MD estimator is efficient in the class of $K$-MD estimators for all $K \in \mathbb{N}$. In other words, additional policy iterations do not provide first-order efficiency gains relative to the optimal $1$-MD estimator. Second, the optimal $1$-MD estimator is more or equally efficient than any $K$-PML estimator for all $K \in \mathbb{N}$. Finally, the appendix provides appropriate conditions under which the optimal $1$-MD estimator is efficient among regular estimators. \begin{description} • dynamic discrete choice problems, dynamic games, pseudo maximum likelihood estimator, minimum distance estimator, estimation, optimality, efficiency. • \end{description}

\thispagestyle{empty} \setcounter{page}1

\hypersetup{pageanchor=true}

Introduction

This paper investigates the first-order asymptotic properties of a broad class of estimators of the structural parameters in a dynamic discrete choice game, i.e., a dynamic game with discrete actions. Given an i.i.d.\ sample of games $i=1,\dots,n$, we consider the class of $K$-stage policy iteration (PI) estimators, where $K$ denotes the number of policy iterations employed in the estimation. The $K$-stage PI estimator is defined by

equation[equation omitted — 151 chars of source]

where $\alpha $ is the structural parameter of interest with a true value equal to $\alpha^*\in \Theta_{\alpha}\subseteq \mathbb{R}^{d_{\alpha}}$, $\hat{Q}$ is a sample criterion function, and $\hat{P}_{k}$ is the $k$-stage estimator of the conditional choice probabilities (CCPs), which is defined iteratively as follows. The preliminary estimator of the CCPs is denoted by $ \hat{P}_{0}$. One possible choice of $ \hat{P}_{0}$ is the sample frequency estimator of the CCPs, although this is not required. Then, for any $k=1,\dots, K-1$,

equation[equation omitted — 100 chars of source]

where $\Psi$ is the best response CCP mapping of the structural game. Given any set of beliefs $P$, optimal or not, $\Psi(\alpha,P) $ indicates the corresponding optimal CCPs when the structural parameter is $\alpha$. The idea of using iterations to estimate dynamic discrete choice problems was introduced in the seminal contributions of aguirregabiria/mira:2002,aguirregabiria/mira:2007. As argued in aguirregabiria/mira:2007, relative to their non-iterative counterparts, these iterative estimators can be more efficient, have smaller finite sample bias due to a more precise initial estimator of the CCPs, and are robust to inconsistent choices of the initial estimator of the CCPs. Here and throughout this paper, we use “efficiency” to denote first-order asymptotic efficiency, and we say that an estimator is “optimal” within a certain class of estimators if it first-order asymptotically efficient within said class.

Our $K$-stage PI estimator nests most of the estimators proposed in the dynamic discrete choice games literature. By appropriate choice of $\hat{Q}$ and $K$, our $K$-stage PI estimator coincides with the pseudo maximum likelihood (PML) estimator in aguirregabiria/mira:2002,aguirregabiria/mira:2007, the asymptotic least squares estimators in pesendorfer/schmidt-dengler:2008, or the so-called simple estimators in pakes/ostrovsky/berry:2007.

To implement the $K$-stage PI estimator, the researcher must determine the number of policy iterations $K$. This choice poses several related research questions. How should researchers choose $K$? Does it make a difference? If so, what is the optimal choice of $K$? The literature provides arguably incomplete answers to these questions. The main contribution of this paper is to answer these questions. Before describing our results, we review the main related findings in the literature.

aguirregabiria/mira:2002,aguirregabiria/mira:2007 propose $K$-stage PML estimators of the structural parameters in dynamic discrete choice problems. The earlier paper considers single-agent problems whereas the later one generalizes the analysis to multiple-agent problems, i.e., games. In both of these papers, the objective is to maximize the pseudo log-likelihood criterion function $\hat{Q}=\hat{Q}_{PML}$, defined by

equation[equation omitted — 142 chars of source]

In this paper, we refer to the resulting estimator as the $K$-PML estimator. One of the main contributions of aguirregabiria/mira:2002,aguirregabiria/mira:2007 is to study the effect of the number of iterations $K$ on the asymptotic distribution of the $K$-PML estimator.

In single-agent dynamic problems, aguirregabiria/mira:2002 show that the asymptotic distribution of the $K$-PML estimators is invariant to $K$. In other words, any additional round of policy mapping iteration has no first-order effect on the asymptotic distribution. This striking result is a consequence of the so-called “zero Jacobian property”, that naturally occurs when a single agent makes optimal decisions. The zero Jacobian property typically does not hold in dynamic problems with multiple players, as each player makes optimal choices according to their preferences, which may not be aligned with their competitors' preferences. Thus, in dynamic discrete choice games, one might expect the asymptotic distribution of $K$-PML estimators to change with $K$.

In multiple-agent dynamic games, aguirregabiria/mira:2007 show that the asymptotic distribution of the $K$-PML estimators is {\it not} invariant to $K$. They consider two specific choices of $K$. On the one hand, they consider the $1$-PML estimator, which they refer to as the two-step pseudo maximum likelihood (PML) estimator. On the other hand, they propose a sequential nested pseudo likelihood (NPL) algorithm, which consists in increasing $K$ until the $K$-PML estimator converges (i.e., $\hat{\alpha}_{(K-1)-PML}=\hat{\alpha}_{K-PML}$). We refer to the estimator resulting from the convergence the NPL algorithm as the $\infty$-PML estimator.\footnote{aguirregabiria/mira:2007 propose the sequential NPL algorithm as a way of computing their NPL estimator. To define the NPL estimator, we first define the NPL fixed points as all the pairs of $(\hat{\alpha},\hat{P})$ that satisfy two conditions: (a) given $\hat{P}$, $\hat{\alpha}$ maximizes the PML criterion function in Eq.\ (ref) and (b) given $\hat{\alpha}$, $\hat{P}$ is a fixed point of the best response CCP mapping in Eq.\ (ref). Then, the NPL estimator is the $\hat{\alpha}$ of the NPL fixed point that maximizes the PML criterion function.} Under some conditions, aguirregabiria/mira:2007 show that the $1$-PML and $\infty$-PML estimators are consistent and asymptotically normal estimators of $\alpha^* $, i.e.,

align[align omitted — 284 chars of source]

Importantly, under additional conditions, aguirregabiria/mira:2007 show that $\Sigma_{1-PML}-\Sigma_{\infty-PML}$ is positive definite, that is, the $\infty$-PML estimator is more efficient than the $1$-PML estimator. So, although iterations of the policy mapping may be burdensome, they can improve efficiency within the $K$-PML class.

In later work, pesendorfer/schmidt-dengler:2010 indicate that the sequential algorithm used to compute the $\infty$-PML estimator may be inconsistent in certain games with unstable equilibria. The intuition for this is as follows. Recall that the $\infty$-PML estimator is defined as the limit of the $K$-PML estimator when $K$ is increased until its convergence. For a given data sample, sampling error implies that the estimator of the CCPs $\hat{P}_{K}$ differs from the equilibrium CCPs. In an unstable equilibrium, increasing $K\to\infty $ can derail $\hat{P}_{K}$ away from the (unstable) equilibrium CCPs, regardless of the sample size $n$. As a consequence of this, the algorithm used to compute $\infty$-PML may produce inconsistent results in the relevant asymptotic framework in which we first consider $ K\to\infty $ and then consider $n\to\infty$.\footnote{Note that this inconsistency would disappear under an asymptotic framework in which we first consider $n\to\infty$ and then consider $K\to\infty$. Unfortunately, by definition, this asymptotic framework is not the relevant one to analyze the asymptotic properties of the $\infty$-PML.} In this paper, we avoid the problem raised by pesendorfer/schmidt-dengler:2010 because we consider an asymptotic analysis for $K$-stage PI estimators with $K\in \mathbb{N}$ and fixed as $n\to\infty$.

pesendorfer/schmidt-dengler:2008 consider the estimation of dynamic discrete choice games using a class of minimum distance (MD) estimators. Specifically, their objective is to minimize the sample criterion function $\hat{Q}=\hat{Q}_{MD}$, given by

equation*[equation* omitted — 167 chars of source]

where $\hat{W}$ is a weight matrix that converges in probability to a limiting weight matrix $W$. This is a single-stage estimator and, consequently, we refer to it as the $1$-MD estimator. pesendorfer/schmidt-dengler:2008 show that the $1$-MD estimator is a consistent and asymptotically normal estimator of $\alpha $, i.e.,

align*[align* omitted — 132 chars of source]

pesendorfer/schmidt-dengler:2008 show that an appropriate choice of $\hat{W}$ implies that the $1$-MD is asymptotically equivalent to the $1$-PML estimator in aguirregabiria/mira:2007 or the simple estimators in pakes/ostrovsky/berry:2007. Furthermore, pesendorfer/schmidt-dengler:2008 characterize the optimal choice of $W$, denoted by $W_{1-MD}^{\ast }$. In general, $\Sigma_{1-PML}-\Sigma_{1-MD}( W_{1-MD}^{\ast })$ is positive semidefinite, i.e., the optimal $1$-MD estimator is more or equally efficient than the $1$-PML estimator.

In the light of the results in aguirregabiria/mira:2007 and pesendorfer/schmidt-dengler:2008, it is natural to inquire whether an iterated version of the MD estimator could yield efficiency gains relative to the $1$-MD or the $K$-PML estimators. The consideration of iterated MD estimator opens several important research questions. How should we define the iterated version of the MD estimator? Does this strategy result in consistent and asymptotically normal estimators of $\alpha^*$? If so, how should we choose the weight matrix $W$? What about the number of iterations $K$? Finally, does iterating the MD estimator produce efficiency gains as aguirregabiria/mira:2007 find for the $K$-PML estimators? This paper answers these questions.

We now summarize the main findings of our paper. We consider a standard dynamic discrete-choice game as in aguirregabiria/mira:2007 or pesendorfer/schmidt-dengler:2008. In this context, we investigate the asymptotic properties of $K$-PML and $K$-MD estimators.

First, we establish that the $K$-PML estimator is consistent and asymptotically normal for any $K\in \mathbb{N}$. See also aguirregabiria:2004 for related results. This complements findings in aguirregabiria/mira:2007, who focus on $K=1$ and $K$ large enough to induce convergence of the estimator. Under certain conditions, we show that the asymptotic variance of the $K$-PML estimator can exhibit arbitrary patterns as a function of $K$. In particular, depending on the parameters of the dynamic problem, the asymptotic variance could increase, decrease, or even be non-monotonic with $K$.

Second, we also establish that the $K$-MD estimator is consistent and asymptotically normal for any $K\in \mathbb{N}$. This is a novel contribution relative to pesendorfer/schmidt-dengler:2008 or pakes/ostrovsky/berry:2007, who focus on non-iterative $1$-MD estimators. The asymptotic distribution of the $K$-MD estimator depends on the choice of the weight matrix. For a specific weight matrix, the $K$-MD has the same asymptotic distribution as the $K$-PML. We investigate the optimal choice of the weight matrix for the $K$-MD estimator.

Our main result, Theorem (ref), shows that an optimal $K$-MD estimator has an asymptotic distribution that is invariant to $K$. This appears to be a novel result in the literature on PI estimation for games, and it is particularly surprising given the findings in aguirregabiria/mira:2007 for $K$-PML estimators. Our main result implies two important corollaries regarding the optimal $1$-MD estimator (derived by pesendorfer/schmidt-dengler:2008):

enumerate• The optimal $1$-MD estimator is efficient in the class of $K$-MD estimators. In other words, additional policy iterations do not provide efficiency gains relative to the optimal $1$-MD estimator. • The optimal $1$-MD estimator is more or equally efficient than any $K$-PML estimator.

We also show in Section (ref) that, under suitable conditions, the optimal $1$-MD estimator has the same asymptotic distribution as the maximum likelihood estimator (MLE), and so it is efficient within the class of all regular estimator. Finally, we reiterate that our asymptotic analysis focuses on first-order terms and ignores high-order approximation errors. In finite samples, these terms may generate high-order efficiency gains for the optimal $K$-MD as we vary the number of iterations. Besides exploring this in simulations, we study this issue theoretically in Section (ref), where we characterize these high-order terms.

The remainder of the paper is organized as follows. Section (ref) describes the dynamic discrete choice game used in the paper, introduces the PI estimators and the main assumptions, and provides an illustrative example of the econometric model. Section (ref) studies the asymptotic properties of the $K$-PML estimator. Section (ref) introduces the $K$-MD estimation method, relates it to the $K$-PML method, and studies its asymptotic distribution. Section (ref) presents a Monte Carlo simulation, and Section (ref) concludes. The appendix of the paper collects the proofs, intermediate results, and complementary findings.

Setup

This section describes the econometric model, introduces the estimator and the assumptions, and provides an illustrative example.

Econometric model

We consider a standard dynamic discrete-choice game as described in aguirregabiria/mira:2007 or pesendorfer/schmidt-dengler:2008. The game has discrete time $t=1,\ldots ,T\equiv\infty$, and a finite set of players indexed by $j \in J\equiv \{1,\ldots ,|J|\}$. In each period $t$, every player $j$ observes a vector of state variables $s_{jt}$ and chooses an action $a_{jt}$ from a finite and common set of actions $ A\equiv \{0,1,\ldots ,|A|-1\}$ (with $|A|>1$) to maximize his expected discounted utility. The action denoted by $0$ is referred to as the outside option, and we denote $\tilde{A}\equiv \{1,\ldots ,|A|-1\}$. All players choose their action simultaneously and non-cooperatively upon observation of state variables.

The vector of state variables $s_{jt}$ is composed of two subvectors $x_{t}$ and $\epsilon _{jt}$. The subvector $x_{t}\in X\equiv \{1,\dots ,|X|\}$ represents a state variable observed by all other players and the researcher, whereas the subvector $\epsilon _{jt}\in \mathbb{R}^{|A|}$ represents an action-specific state vector only observed by player $j$. We denote $\epsilon _{t} \equiv \{ \epsilon _{jt}:j \in J\} \in \mathbb{R}^{|A| \times |J|}$ and $\vec{a}_{t} \equiv \{ a_{jt}:j \in J\} \in A^{|J|}$.

We assume that $\epsilon _{t}$ is independent of $(\epsilon _{\tau},\vec{a}_{\tau},x_{\tau})$ for $\tau<t$, with density $dF_{\epsilon }(e) = \prod_{j=1}^{J} dF_{\epsilon,j}(e_j)$, where $dF_{\epsilon,j}$ is an absolutely continuous density function. We also assume that $\epsilon _{t}$ has full support and that $ E[\epsilon _{jt}|\epsilon _{jt}\geq e]$ is finite for all $e\in \mathbb{R}^{|A|} $. Conditional on $(\vec{a}_{t},x_{t})$, we assume that $x_{t+1}$ is independent of $(\epsilon _{\tau},\vec{a}_{\tau-1},x_{\tau-1})$ for $\tau\leq t$, and with probability $dF_{x}(x_{t+1}|\vec{a}_{t},x_{t})$. It then follows that $s_{t+1}=(x_{t+1},\epsilon _{t+1})$ is a Markov process with a probability density that satisfies

equation*[equation* omitted — 175 chars of source]

Every player $j$ has a time-separable utility and discounts future payoffs by $\beta^* _{j}\in (0,1)$. The period $t$ payoff is received after every player made their choices and is given by:

equation*[equation* omitted — 97 chars of source]

Following the literature, we assume Markov perfect equilibrium (MPE) as the equilibrium concept for the game. By definition, an MPE is a collection of strategies and beliefs for each player such that each player has: (a) rational beliefs, (b) an optimal strategy given his beliefs and other players' choices, and (c) Markovian strategies. According to pesendorfer/schmidt-dengler:2008, this model has an MPE and it could even have multiple MPEs (e.g., see pesendorfer/schmidt-dengler:2008). We follow aguirregabiria/mira:2007 and assume that data come from one of the MPEs in which every player uses pure strategies.\footnote{As explained in aguirregabiria/mira:2007, this can be rationalized by Harsanyi's Purification Theorem.}

An MPE is a collection of equilibrium strategies and common beliefs. We denote the probability that player $j$ will choose action $a\in A $ given observed state $x$ by $P_{j}^{\ast }(a|x)$, and we denote $P^{\ast }\equiv \{P_{j}^{\ast }(a|x):(j,a,x)\in J\times \tilde{A}\times X\} \in \mathbb{R}^{d_{P}}$ with $d_{P} = |J|\times |\tilde{A}| \times |X|$. Note that beliefs only need to be specified in $\tilde{A}$ for every $(j,x)\in J\times X$, as $P_{j}^{\ast }(0|x)=1-\sum_{a\in \tilde{A}}P_{j}^{\ast }(a|x)$. We denote player $j$'s equilibrium strategy by $\{a_{j}^{\ast }(e,x):(e,x)\in \mathbb{R}^{|A|}\times X\}$, where $a_{j}^{\ast }(e,x)$ denotes player $j$'s optimal choice when the current private shock is $e$ and the observed state is $x$. Given that equilibrium strategies are time-invariant, we can abstract from calendar time for the remainder of the paper, and denote $\vec{a}=\vec{a}_{t}$, $\vec{a}^{\prime }=\vec{a}_{t+1}$, $ x=x_{t}$, and $x^{\prime }=x_{t+1}$.

We use $\theta^{\ast } \in \Theta$ to denote the finite-dimensional parameter vector that collects the model elements $(\{\pi _{j}:j \in J\},\{\beta^* _{j}:j \in J\},dF_{\epsilon },dF_{x})$. Throughout this paper, we split the parameter vector as follows:

align[align omitted — 145 chars of source]

where $\alpha^{\ast } \in \Theta_{\alpha}\subseteq \mathbb{R}^{d_\theta}$ denotes a parameter vector of interest that is estimated iteratively and $g^{\ast } \in \Theta_g\subseteq \mathbb{R}^{d_g}$ denotes a nuisance parameter vector that is estimated directly from the data. We assume $d_\alpha \leq d_P$ and note that, in a typical application, $d_\alpha$ is much smaller than $d_P$. In practice, structural parameters that determine the payoff functions $\{\pi _{j}:j \in J\}$ or the distribution $dF_{\varepsilon}$ usually belong to $\alpha^{\ast } $, while the transition probability density function $dF_{x}$ is typically part of $g^*$.\footnote{Note that the distinction between components of $\theta^*$ is without loss of generality, as one can choose to estimate all parameters iteratively by setting $\theta^* = \alpha^*$. The goal of estimating a nuisance parameter $g^*$ directly from the data is to simplify the computation of the iterative procedure by reducing its dimensionality.}

We now describe a fixed point mapping that characterizes equilibrium beliefs in any MPE. Let $ P=\{P_{j}(a|x):(j,a,x)\in J\times \tilde{A}\times X\}$ denote a set of beliefs, which need not be optimal. Given these beliefs, the ex-ante probability that player $j$ chooses equilibrium action $a$ given observed state $x$ is

equation[equation omitted — 256 chars of source]

where $u_{j}(a,x, \alpha^{\ast },g^{\ast },P)$ denotes player $j$'s conditional choice value function under action $a$, state variable $x$, and with beliefs $P$. In turn,

equation*[equation* omitted — 289 chars of source]

where $\prod_{s\in J\backslash \{j\}}P_{s}(\tilde{a}_{s}|x)$ denotes the beliefs that the remaining players choose $\tilde{a}\equiv \{\tilde{a} _{s}:s\in J\backslash \{j\}\}$ conditional on $x$, and $V_{j}(x;P)$ is player $j$'s ex-ante value function conditional on $x$.\footnote{ The ex-ante value function is the discounted sum of future payoffs in the MPE given $x$ and before players observe shocks and choose actions. It can be computed with the mapping valuation operator defined in aguirregabiria/mira:2007 or pesendorfer/schmidt-dengler:2008.} By stacking up this mapping for all decisions and states $(a,x)\in \tilde{A}\times X$ and all players $j \in J$, we define the probability mapping $\Psi ( \alpha^{\ast },g^{\ast },P)\equiv \{\Psi _{j}(a,x; \alpha^{\ast },g^{\ast },P):(j,a,x)\in J\times \tilde{A} \times X\}$. Given any set of beliefs $P$ (optimal or not), $\Psi ( \alpha^{\ast },g^{\ast },P)$ indicates the corresponding optimal CCPs.

aguirregabiria/mira:2007 and pesendorfer/schmidt-dengler:2008 show that the mapping $\Psi $ fully characterizes equilibrium beliefs $P^{\ast }$ in an MPE. That is, $P^{\ast }$ is an equilibrium belief if and only if

equation[equation omitted — 85 chars of source]

The goal of the paper is to study the problem of inference of $ \alpha^* \in \Theta _{\alpha }$ based on the fixed point equilibrium condition in Eq.\ (ref).

Estimation procedure

The researcher estimates $\theta^{\ast }=(\alpha^{\ast },g^{\ast })\in \Theta \equiv \Theta _{\alpha }\times \Theta _{g}$ using a two-step and $K$-stage PI estimator. For any $K\in \mathbb{N}$, this estimator is defined as follows:

itemize• {\bf Step 1:} Estimate $( g^{\ast },P^{\ast }) $ with preliminary estimators $( \hat{g},\hat{P}_{0})$. • {\bf Step 2:} Estimate $\alpha^{\ast }$ with $\hat{\alpha}_{K}$, computed by the following algorithm. Initialize $k=1$ and then: \begin{itemize} • Compute \begin{equation} \hat{\alpha}_{k} \equiv \underset{{\alpha \in \Theta }_{\alpha }}{\arg \max } {\hat{Q}}_{k}(\alpha ,\hat{g},\hat{P}_{k-1}), \end{equation} where ${\hat{Q}}_{k}:\Theta _{\alpha }\times \Theta _{g}\times \Theta _{P}\to \mathbb{R}$ is the $k$-th step sample objective function. If $k=K$, exit the algorithm. If $k<K$, go to (b). • Estimate $P^{\ast }$ with the $k$-step estimator of the CCPs, given by \begin{equation} \hat{P}_{k} \equiv \Psi ({\hat{\alpha}}_{k},\hat{g},\hat{P}_{k-1}). \end{equation} Then, increase $k$ by one unit and return to (a). \end{itemize}

Throughout this paper, we consider $\alpha^{\ast } $ to be our main parameter of interest, while $g^{\ast }$ is a nuisance parameter. For any $K\in \mathbb{N}$, the two-step and $K$-stage PI estimator of $\alpha^{\ast }$ is given by

equation[equation omitted — 164 chars of source]

and the corresponding estimator of $\theta^{\ast }=(\alpha^{\ast },g^{\ast })$ is $\hat{\theta}_{K} = (\hat{\alpha}_{K}, \hat{g})$.

The algorithm does not specify the first-step estimators $(\hat{g},\hat{P}_{0})$ or the sequence of sample criterion functions $\{{\hat{Q}}_{k}:k\leq K\}$. One possible choice of $ \hat{P}_{0}$ is the sample frequency estimator of the CCPs, although this is not required. Rather than determining these objects now, we restrict them by making assumptions in the next subsection (see Assumptions (ref) and (ref)).

To conclude this subsection, we note that aguirregabiria/mira:2007 consider a version of the algorithm described above in which only preliminary estimator of the CCPs is estimated in Step 1, and the entire parameter vector $(g^*,\alpha^*)$ is estimated (iteratively) in Step 2. This can be considered a special case of our algorithm by using $\alpha$ to denote the entire parameter vector $\theta$.

Assumptions

This section introduces the main assumptions used in our analysis. We note that these conditions are similar to those used in several other papers in the literature, especially aguirregabiria/mira:2007 and pesendorfer/schmidt-dengler:2008. In addition, our simulation results suggest that our assumptions are satisfied in the entry game described in Section (ref).\footnote{The MPEs in our entry game are found numerically, and this complicates verifying these conditions in practice, especially those related to multiplicity of equilibria. However, our simulation evidence does not suggest issues with any of our assumptions.}

As explained in Section (ref), the game has an MPE but this need not be unique. To address this issue, we follow most of the literature, and assume that the researcher observes an i.i.d.\ sample from a single MPE.

assumptionA[(I.i.d.\ data from one MPE)] The data $\{\{(\{a_{j,i}:j \in J\},x_{i},x_{i}^{\prime })\}:i\leq n\}$ are an i.i.d.\ sample from a single MPE. This MPE determines the data generating process (DGP) denoted by $\Pi ^{\ast }\equiv \{\Pi ^{\ast }(\vec{a},x,x^{\prime }):(\vec{a},x,x^{\prime })\in A^{|J|}\times X\times X\}$, where $\Pi ^{\ast }(\vec{a},x,x^{\prime })$ denotes the probability that players choose action $\vec{a}$ and the current state variable evolves from $x$ to $x'$, i.e., \[ \Pi ^{\ast }(\vec{a},x,x')~\equiv ~\Pr [~(\{a_{j,i}:j \in J\},x_{i},x_{i}')=(\vec{a},x,x')~]. \]

See aguirregabiria/mira:2007) for a similar condition. The observations in the i.i.d.\ sample are indexed by $ i=1,\ldots ,n$, which could denote different markets as in aguirregabiria/mira:2007. By Assumption (ref), the data identify the DGP, i.e., $\Pi ^{\ast }(\vec{a},x,x')$ for every $ (\vec{a},x,x')\in A^{|J|}\times X\times X$, which determine the equilibrium CCPs, transition probabilities, and marginal state distribution. For all $(j,\vec{a},x,x')\in J \times A^{|J|}\times X\times X$ with $\vec{a} = (a, \vec{a}_{-j})$, these are denoted by

eqnarray[eqnarray omitted — 565 chars of source]

where ${P}_{j}^{\ast }(a|x)$ denotes the probability that player $j$ will choose action $a$ given that the observed state is $x$, ${ \Lambda }^{\ast }(x'|x,\vec{a})$ denotes the probability that the future state observed state is $x^{\prime }$ given that the current observed state is $x$ and the action vector is $\vec{a}$, and ${m}^{\ast }(x)$ denotes the (unconditional) probability that the current observed state is $x $. Finally, recall that the equilibrium CCPs are $P^{\ast } = \{ P_{j}^{\ast }(a|x):(a,j,x) \in \tilde{A}\times J \times X\}$.

Identification of the CCPs, however, is not sufficient for identification of the parameters of interest. To this end, it is essential to make the following assumption.

assumptionA[(Identification)] $\Psi ( \alpha ,g^{\ast },P^{\ast }) =P^{\ast }$ if and only if $\alpha =\alpha^{\ast }$.

See aguirregabiria/mira:2007 and pesendorfer/schmidt-dengler:2008 for a similar condition. Identification in these models is studied in pesendorfer/schmidt-dengler:2008. In particular, pesendorfer/schmidt-dengler:2008 indicate the maximum number of parameters that could be identified from the model and pesendorfer/schmidt-dengler:2008 provides sufficient conditions for identification.

The $K$-stage PI estimator $\hat{\alpha} _{K}$ is an example of an extremum estimator. The following assumption imposes mild conditions that are typically used to show the asymptotic properties of these estimators.

assumptionA[(Regularity conditions)] Assume the following conditions: \begin{enumerate}[(i)] • $\alpha^{\ast }$ belongs to the interior of $\Theta _{\alpha }$. • $\sup_{\alpha \in \Theta _{\alpha }}|\Psi (\alpha ,\tilde{g},\tilde{P} )-\Psi (\alpha ,g^{\ast },P^{\ast })|=o_{p}(1)$, provided that $(\tilde{g}, \tilde{P})=(g^{\ast },P^{\ast })+o_{p}(1)$. • $\inf_{\alpha \in \Theta _{\alpha }}\Psi _{ajx}(\alpha ,\tilde{g}, \tilde{P})>0$ for all $(a,j,x)\in A\times J\times X$, provided that $(\tilde{ g},\tilde{P})=(g^{\ast },P^{\ast })+o_{p}(1)$. • $\Psi (\alpha ,g,P)$ is twice continuously differentiable in a neighborhood of $(\alpha^{\ast },g^{\ast },P^{\ast })$. We use $\Psi _{\lambda }\equiv \partial \Psi (\alpha^{\ast },g^{\ast },P^{\ast })/\partial \lambda $ for $\lambda \in \{\alpha ,g,P\}$. • $(\mathbf{I}_{d_{P}}-\Psi _{P},-\Psi _{g})\in \mathbb{R}^{d_{P}\times (d_{P}+d_{g})}$ and $\Psi _{\alpha }\in \mathbb{R}^{d_{P}\times d_{\alpha }}$ are full rank matrices. \end{enumerate}

We now comment on these conditions. First, the asymptotic analysis of $\hat{\alpha} _{K}$ follows from the first order condition that results from Eq.\ (ref). Assumption (ref)(i) justifies the use of the first order condition for interior parameter values, and is also required by aguirregabiria/mira:2007 and pesendorfer/schmidt-dengler:2008. Second, standard arguments to establish the consistency of $\hat{\alpha} _{K}$ require that the corresponding sample criterion function converges uniformly to its limit, which can be related to Assumptions (ref)(ii)-(iii). Third, standard argument to prove the asymptotic normality of $\hat{\alpha} _{K}$ are based on a mean value expansion based on the first order condition, which requires second-degree differentiability in Assumption (ref)(iv). We note that this assumption coincides with pesendorfer/schmidt-dengler:2008. Fourth, Assumption (ref)(v) imposes rank conditions on the structure of the dynamic game that also required by pesendorfer/schmidt-dengler:2008. Finally, we note that the validity of Assumptions (ref)(iii)-(v) can be formally tested in empirical applications.

We next introduce assumptions on $(\hat{g},\hat{P}_{0}) $, i.e., the preliminary estimators of $( g^{\ast },P^{\ast }) $. For reference, we first define the sample frequency estimator of the CCPs, given by

equation[equation omitted — 117 chars of source]

with

equation*[equation* omitted — 129 chars of source]

It is not hard to show that

equation*[equation* omitted — 105 chars of source]

where $\Omega _{PP}$ is the block diagonal matrix defined by $\Omega _{PP}\equiv diag\{\Sigma _{jx}:(j,x)\in J\times X\}$ with $ \Sigma _{jx}\equiv (diag\{P_{jx}^{\ast }\}-P_{jx}^{\ast }P_{jx}^{\ast \prime })/m^{\ast }(x)$ and $P_{jx}^{\ast }\equiv \{P_{j}^{\ast }(a|x):a\in \tilde{A} \} $ for all $(j,x)\in J\times X$.

Rather than imposing specific preliminary estimators $( \hat{g},\hat{P} _{0}) $, we entertain two high-level assumptions that restrict the relationship between these and $\hat{P}$.

assumptionA[(Baseline convergence)] $(\hat{P},\hat{P}_{0},\hat{g})$ satisfies the following condition: \begin{equation*} \sqrt{{n}}\left( \begin{array}{c} \hat{P}-P^{\ast } \\ \hat{P}_{0}-P^{\ast } \\ \hat{g}-g^{\ast } \end{array} \right) \overset{d}{\to } N\left( \left( \begin{array}{c} \mathbf{0}_{d_{P}\times 1} \\ \mathbf{0}_{d_{P}\times 1} \\ \mathbf{0}_{d_{g}\times 1} \end{array} \right) ,\left( \begin{array}{ccc} \Omega _{PP} & \Omega _{P0} & \Omega _{Pg} \\ \Omega _{P0}^{\prime } & \Omega _{00} & \Omega _{0g} \\ \Omega _{Pg}^{\prime } & \Omega _{0g}^{\prime } & \Omega _{gg} \end{array} \right) \right). \end{equation*}
assumptionA[(Baseline convergence II)] $(\hat{P},\hat{P}_{0},\hat{g})$ satisfies the following conditions. \begin{enumerate}[(i)] • The asymptotic variance of $(\hat{P},\hat{g})$ is nonsingular. • $( \hat{P},\hat{g})$ is an estimator of $(P^*,g^*)$ that at least as efficient as $((\mathbf{I}_{d_{P}}-M) \hat{P}+M\hat{P}_{0},\hat{g})$ for any $M\in \mathbb{R}^{d_{P}\times d_{P}}$. \end{enumerate}

Assumption (ref) imposes the consistency and joint asymptotic normality of $(\hat{P},\hat{P}_{0},\hat{g})$, which is satisfied by all standard choices for these estimators when the observed states and actions have a finite support. Assumption (ref)(i) is a natural condition that is implicitly imposed by the definition of the optimal weight matrix in pesendorfer/schmidt-dengler:2008. The interpretation of Assumption (ref)(ii) requires additional discussion. The point of a preliminary estimator $(\hat{P}_{0},\hat{g})$ is to approximate $(P^{\ast },g^{\ast })$ without the need for imposing the restrictions from the structural model, since this may be computationally burdensome. If we ignore these restrictions, $\hat{P}$ is the maximum likelihood estimator (MLE) of $ P^{\ast }$ and it is thus an efficient estimator of $P^{\ast }$. As a corollary, $\hat{P}$ is an estimator of $P^*$ that is at least as efficient as $(\mathbf{I}_{d_{P}}-M)\hat{P}+M\hat{P}_{0}$ for any $M\in \mathbb{R}^{d_{P}\times d_{P}}$. Assumption (ref)(ii) essentially requires that this conclusion also applies when these estimators are coupled with $\hat{g}$ as an estimator of $g^*$.

To illustrate Assumptions (ref) and (ref), it is necessary to specify the parameter vector $g^{\ast }$. A typical specification in the literature is the one in pesendorfer/schmidt-dengler:2008, where $g^{\ast }$ is the vector of state transition probabilities, i.e.,

equation[equation omitted — 152 chars of source]

and $\tilde{X}\equiv \{ 2,\ldots ,|X|\} $ is the state space with the first action removed (to avoid redundancy). In this setting, a reasonable specification of $( \hat{P}_{0},\hat{g})$ is to set them equal to their corresponding sample frequency estimators, i.e., $\hat{P}_{0}=\hat{P}$ and $\hat{g}\equiv \{\hat{ g}(\vec{a},x,x^{\prime }):(x^{\prime },\vec{a},x)\in \tilde{X} \times A^{|J|}\times X\}$ with

equation[equation omitted — 233 chars of source]

If we ignore the restrictions of the econometric model, $(\hat{P}_{0},\hat{g})$ is the MLE of $( P^{\ast },g^{\ast }) $, and standard arguments imply Assumptions (ref) and (ref). However, note that one could obtain the same results if were replaced $(\hat{P}_{0},\hat{g})$ with any asymptotically equivalent estimator, i.e., any estimator $(\tilde{P}_{0}, \tilde{g})$ such that $(\tilde{P}_{0},\tilde{g})=(\hat{P}_{0},\hat{g} )+o( n^{-1/2}) $. Examples of asymptotically equivalent estimators would be the ones resulting from the flexible logit estimator (e.g., see arcidiacono/ellickson:2011), the relaxation method in kasahara/shimotsu:2012, or the undersmoothed kernel estimator in grund:1993.

An illustrative example

We illustrate the framework with the two-player version dynamic entry game in aguirregabiria/mira:2007. In each period $t=1,\ldots ,T\equiv\infty$, two firms indexed by $j \in J=\{1,2\}$ simultaneously decide whether to enter or not into the market, upon observation of the state variables. Firm $j$'s decision at time $t$ is $a_{jt}\in A=\{0,1\}$, which takes value one if firm $j$ enters the market at time $t$, and zero otherwise. In each period $t$, the vector of state variables observed by firm $j$ is $s_{jt}=(x_{jt},\epsilon _{jt})$, where $\epsilon _{jt}=(\epsilon _{0jt},\epsilon _{1jt})\in \mathbb{R}^{2}$ represents the privately-observed vector of action-specific state variables and $ x_{jt}=x_{t}\in X\equiv \{ 1,2,3,4\} $ is a publicly-observed state variable that indicates the entry decisions in the previous period, i.e.,

equation*[equation* omitted — 229 chars of source]

We specify the profit function as in aguirregabiria/mira:2007. If firm $j$ enters the market in period $t$, its period profits are

equation[equation omitted — 207 chars of source]

where $\lambda _{RS}^{\ast }$ represents fixed entry profits, $\lambda _{RN}^{\ast }$ represents the effect of a competitor's entry, $\lambda _{FC,j}^{\ast }$ represents a firm-specific fixed cost, and $\lambda _{EC}^{\ast }$ represents the entry cost. On the other hand, if firm $j$ does not enter the market in period $t$, its period profits are

equation*[equation* omitted — 65 chars of source]

Firms discount future profits at a common discount factor $\beta^{\ast }\in (0,1)$. We assume that $\{\epsilon _{ajt}:(a,j,t) \in A \times J \times \mathbb{N}\}$ are i.i.d.\ with standard Gumbel distribution, and so

equation*[equation* omitted — 83 chars of source]

Finally, $x_{t+1}$ is uniquely determined by $ \vec{a}_{t}$ and so,

equation*[equation* omitted — 249 chars of source]

This completes the specification of the econometric model up to $(\lambda _{RN}^{\ast },\lambda _{EC}^{\ast },\lambda _{RS}^{\ast },\lambda _{FC,1}^{\ast },\lambda _{FC,2}^{\ast },\beta^{\ast })$. These parameters are known to the players but not necessarily known to the researcher.

We will use this econometric model to illustrate our theoretical results and for our Monte Carlo simulations. For simplicity, we presume that the researcher knows $(\lambda _{RS}^{\ast },\lambda _{FC,1}^{\ast },\lambda _{FC,2}^{\ast },\beta^{\ast })$, and is interested in estimating $(\lambda _{RN}^{\ast },\lambda _{EC}^{\ast })$. In addition, we assume that these parameters are all estimated iteratively in Step 2, i.e., $ \theta^{\ast }=\alpha^{\ast }\equiv (\lambda _{RN}^{\ast },\lambda _{EC}^{\ast })$. The only task in Step 1 is to estimate $P^{\ast }$ with a preliminary step estimator $\hat{P}_{0}$, i.e., this model has no parameter $g^{\ast }$.

Results for $K$-PML estimation

This section provides formal results for the $K$-PML estimator introduced in aguirregabiria/mira:2002,aguirregabiria/mira:2007 given an arbitrary number of iteration steps $K\in \mathbb{N}$. The $K$-PML estimator is defined by Eq.\ (ref) with the pseudo log-likelihood criterion function, i.e., $\hat{Q}_{K}=\hat{Q}_{PML}$. That is,

itemize• {\bf Step 1:} Estimate $( g^{\ast },P^{\ast }) $ with preliminary step estimators $( \hat{g},\hat{P}_{0}) $. • {\bf Step 2:} Estimate $\alpha^{\ast }$ with $\hat{\alpha}_{K-PML}$, computed by the following algorithm. Initialize $k=1$ and then: \begin{itemize} • Compute \begin{equation*} \hat{\alpha}_{k-PML} \equiv \underset{\alpha \in \Theta _{\alpha }}{\arg \min} \frac{1}{n} \sum_{i=1}^{n}\ln \Psi ( \alpha ,\hat{g},\hat{P}_{k-1}) ( a_{i}|x_{i}) . \end{equation*} If $k=K$, exit the algorithm. If $k<K$, go to (b). • Estimate $P^{\ast }$ with the $k$-step estimator of the CCPs, given by \begin{equation*} \hat{P}_{k} \equiv \Psi ( \hat{\alpha}_{k-PML},\hat{g},\hat{P} _{k-1}) . \end{equation*} Then, increase $k$ by one unit and return to (a). \end{itemize}

As explained in Section (ref), The $K$-PML estimator is the $K$-stage PI estimator introduced by aguirregabiria/mira:2002 for dynamic single-agent problems and aguirregabiria/mira:2007 for dynamic games. aguirregabiria/mira:2007 study the asymptotic behavior of $\hat{\alpha}_{K-PML}$ for two extreme values of $K$: $K=1$ and $K$ large enough to induce the convergence of the estimator. Under some conditions, they show that iterating the $K$-PML estimator until convergence produces efficiency gains.

The results in aguirregabiria/mira:2007 are restricted in two ways. First, they focus on these two extreme values of $K$, without considering other possible values. Second, they restrict attention to the case in which the parameters of the model are estimated in the iterative part of the algorithm (i.e., $\theta = \alpha$). In Theorem (ref), we complement the analysis in aguirregabiria/mira:2007 along these two dimensions. In particular, we derive the asymptotic distribution of the two-step $K$-PML estimator for any finite $K\geq 1$.

theorem[Two-step $K$-PML] Fix $K\in \mathbb{N}$ arbitrarily and assume Assumptions (ref)-(ref). Then, \begin{equation*} \sqrt{n}( \hat{\alpha}_{K-PML}-\alpha^{\ast }) \overset{d}{ \to } N( \mathbf{0}_{d_{\alpha }\times 1},\Sigma _{K-PML}( \hat{P}_{0},\hat{g}) ) , \end{equation*} where \begin{equation*} \Sigma _{K-PML}( \hat{P}_{0}) \equiv \left\{ \begin{array}{c} ( \Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}\times \\ \left[ \begin{array}{c} ( \mathbf{I}_{d_{P}}-\Psi _{P}\Phi _{K,P})^{\prime } \\ -( \Psi _{P}\Phi _{K,0})^{\prime } \\ -\Psi _{g}^{\prime }( \mathbf{I}_{d_{P}}+\Psi _{P}\Phi _{K,g}) ^{\prime } \end{array} \right]^{\prime }\left( \begin{array}{ccc} \Omega _{PP} & \Omega _{P0} & \Omega _{Pg} \\ \Omega _{P0}^{\prime } & \Omega _{00} & \Omega _{0g} \\ \Omega _{Pg}^{\prime } & \Omega _{0g}^{\prime } & \Omega _{gg} \end{array} \right) \left[ \begin{array}{c} ( \mathbf{I}_{d_{P}}-\Psi _{P}\Phi _{K,P})^{\prime } \\ -( \Psi _{P}\Phi _{K,0})^{\prime } \\ -\Psi _{g}^{\prime }( \mathbf{I}_{d_{P}}+\Psi _{P}\Phi _{K,g})^{\prime } \end{array} \right] \\ \times \Omega _{PP}^{-1}\Psi _{\alpha }( \Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}\Psi _{\alpha })^{-1} \end{array} \right\} , \end{equation*} and $\{ \Phi _{k,P}:k\leq K\} $, $\{ \Phi _{k,0}:k\leq K\} $, and $\{ \Phi _{k,g}:k\leq K\} $ are defined as follows. Set $\Phi _{1,P}\equiv \mathbf{0}_{d_{P}\times d_{P}}$, $\Phi _{1,0}\equiv \mathbf{I}_{d_{P}}$, $\Phi _{1,g}\equiv \mathbf{0}_{d_{P}\times d_{P}}$ and, for any $k\leq K-1$, \begin{align} \Phi _{k+1,P}& \equiv ( \mathbf{I}_{d_{P}}-\Psi _{\alpha }( \Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}) \Psi _{P}\Phi _{k,P}+\Psi _{\alpha }( \Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}, \notag \\ \Phi _{k+1,0}& \equiv ( \mathbf{I}_{d_{P}}-\Psi _{\alpha }( \Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}) \Psi _{P}\Phi _{k,0}, \notag \\ \Phi _{k+1,g}& \equiv ( \mathbf{I}_{d_{P}}-\Psi _{\alpha }( \Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}) \Psi _{P}\Phi _{k,g}+( \mathbf{I}_{d_{P}}-\Psi _{\alpha }( \Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }\Omega _{PP}^{-1}) . \end{align}

We make two comments regarding this result. First, Theorem (ref) considers $K\in \mathbb{N}$ and fixed as $n\to \infty$. Because of this, our asymptotic framework is not subject to the criticism raised by pesendorfer/schmidt-dengler:2010. Second, we note that $\Sigma _{K-PML}(\hat{P}_{0},\hat{g})$ can be consistently estimated using consistent estimators of $(\alpha^*,g^*)$ (e.g., $(\hat{\alpha} _{1-PML},\hat{g})$) and the asymptotic variance in Assumption (ref).

Theorem (ref) reveals that the $K$-PML estimator of $\alpha^{\ast }$ is consistent and asymptotically normally distributed for all $K\geq 1$. Thus, the asymptotic mean squared error of the $K$-PML estimator is equal to its asymptotic variance, $\Sigma _{K-PML}( \hat{P}_{0}) $. The goal for the rest of the section is to investigate how this asymptotic variance changes with the number of iterations $K$.

In single-agent dynamic problems, aguirregabiria/mira:2002 show that the so-called zero Jacobian property holds, i.e., $\Psi _{P}=\mathbf{0} _{dP\times dP}$. If we plug in this condition into Theorem (ref), we conclude that the asymptotic variance of the $K$-PML estimator is given by

align*[align* omitted — 372 chars of source]

This expression is invariant to $K$, corresponding to the main finding in aguirregabiria/mira:2002.

In multiple-agent dynamic problems, however, the zero Jacobian property no longer holds. In the special case with $ d_\alpha = d_P$, Theorem (ref) shows that the asymptotic variance of the $K$-PML does not depend on $K$, and is given by

align*[align* omitted — 359 chars of source]

However, most applications will have $ d_\alpha<d_P$. For this more common case, Theorem (ref) reveals that the asymptotic variance of the $K$-PML estimator can be a complicated function of the number of iteration steps $K$. We illustrate this complexity using the example of Section (ref). In this example, the researcher is interested in estimating $(\lambda _{RN}^{\ast },\lambda _{EC}^{\ast })$. For simplicity, we set $\hat{P}_{0}=\hat{P}$. In this context, the asymptotic variance of $\hat{\alpha}_{K-PML}=(\hat{\lambda}_{RN,K-PML},\hat{ \lambda}_{EC,K-PML})$ is given by

equation[equation omitted — 384 chars of source]

where $\{\Phi _{k,P0}:k\leq K\}$ is defined by $\Phi _{k,P0}\equiv \Phi _{k,P}+\Phi _{k,0}$, with $\{\Phi _{k,P}:k\leq K\}$ and $\{\Phi _{k,0}:k\leq K\}$ as in Eq.\ (ref). For any true parameter vector and any $K\in \mathbb{N} $, we can numerically compute Eq.\ (ref). For the exposition, we focus on the asymptotic variance of $\hat{\lambda} _{RN,K-PML}$, which corresponds to the [1,1]-element of $\Sigma _{K-PML}( \hat{P}_0 )$. Figures (ref), (ref), and (ref) show the asymptotic variance of $\hat{ \lambda}_{RN,K-PML}$ as a function of $K$ for $(\lambda _{RN}^{\ast },\lambda _{EC}^{\ast },\lambda _{RS}^{\ast },\lambda _{FC,1}^{\ast },\lambda _{FC,2}^{\ast },\beta^{\ast }) = (2.8,0.8,0.7,0.6,0.4,0.95)$, $(2,1.8,0.2,0.01,0.03,0.95)$, and $ (2.2,1.45,0.45,0.22,0.29,0.95)$, respectively. These figures confirm that, in general, the asymptotic variance of the $K$-PML estimator can decrease, increase, or even fluctuate with the number of iterations $K$. Note that these widely different patterns occur within the same econometric model. Finally, we point out that qualitatively similar results can be obtained for the asymptotic variance of $\hat{\lambda} _{EC,K-PML}$, which corresponds to the [2,2]-element of $\Sigma _{K-PML}( \hat{P}_0 )$.

We view the fact that $\Sigma _{K-PML}( \hat{P}_{0}) $ can change so much with the number of iterations $K$ as a negative feature of the $K$-PML estimator. A researcher who uses the $K$-PML estimator and is interested in efficiency faces difficulties when choosing $K$. Prior to estimation, the researcher cannot be certain regarding the effect of $K$ on the efficiency of the $K$-PML estimator. Additional iterations could help efficiency (as in Figure (ref)) or hurt efficiency (as in Figure (ref)). In principle, the researcher could consistently estimate the asymptotic variance of $K$-PML for each $K$ by plugging in any consistent estimator of the structural parameters (e.g., $\hat{\alpha} _{1-PML}$ and $\hat{g}$).

figure[figure omitted — 397 chars of source]
figure[figure omitted — 397 chars of source]
figure[figure omitted — 402 chars of source]

Results for $K$-MD estimation

In this section, we introduce a new class of $K$-stage PI estimators, referred to as the $K$-MD estimator. We demonstrate that this has several advantages over the $K$-PML estimator. In particular, we show that an optimal $K$-MD estimator dominates the $K$-PML estimator in terms of efficiency, and its asymptotic variance does not change with $K$. In addition, we show that an optimal $K$-MD estimator can be achieved with $K=1$.

For any $K\in \mathbb{N}$, the $K$-MD estimator is defined by Eq.\ (ref) with (negative) minimum distance criterion function:

equation*[equation* omitted — 140 chars of source]

and where $\{ \hat{W}_{k}:k\leq K\} $ is a sequence of positive semidefinite weight matrices. That is,

itemize• Step 1: Estimate $( g^{\ast },P^{\ast }) $ with preliminary step estimators $( \hat{g},\hat{P}_{0}) $. • Step 2: Estimate $\alpha^{\ast }$ with $\hat{\alpha}_{K-MD}$, computed by the following algorithm. Initialize $k=1$ and then: \begin{itemize} • Compute \begin{equation*} \hat{\alpha}_{k-MD} \equiv \underset{\alpha \in \Theta _{\alpha }}{\arg \min } ( \hat{P}-\Psi ( \alpha ,\hat{g},\hat{P}_{k-1}) )^{\prime }\hat{W}_{k}( \hat{P}-\Psi ( \alpha ,\hat{g},\hat{ P}_{k-1}) ) . \end{equation*} If $k=K$, exit the algorithm. If $k<K$, go to (b). • Estimate $P^{\ast }$ with the $k$-step estimator of the CCPs, given by \begin{equation*} \hat{P}_{k} \equiv \Psi (\hat{\alpha}_{k-MD},\hat{g},\hat{P}_{k-1}). \end{equation*} Then, increase $k$ by one unit and return to (a). \end{itemize}

The implementation of the $K$-MD estimator requires several choices: the number of iteration steps $K$ and the associated weight matrices $\{ \hat{W}_{k}:k\leq K\}$. We note that the sequence of weight matrices does not affect the $K$-MD estimator in the special case with $ d_\alpha = d_P$, although, as already mentioned, it is far more common for applications to have $ d_\alpha<d_P$. Also, note that the least squares estimator in pesendorfer/schmidt-dengler:2008 is a particular case of our $1$-MD estimator with $\hat{P}_{0}=\hat{P}$. In this sense, our $K$-MD estimator can be considered as an iterative version of their least squares estimator. The primary goal of this section is to study how to make optimal choices of $K$ and $\{ \hat{W}_{k}:k\leq K\}$.

To establish the asymptotic properties of the $K$-MD estimator, we add the following assumption.

assumptionA[(Weight matrices)] For every $k\leq K$, $\hat{W}_{k}\overset{p}{\to}W_{k}$ and $W_{k} \in \mathbb{R}^{d_P \times d_P}$ is positive definite and symmetric.

The next result derives the asymptotic distribution of the two-step $K$-MD estimator for any $K\in \mathbb{N}$.

theorem[Two-step $K$-MD] Fix $K\in \mathbb{N}$ arbitrarily and assume Assumptions (ref)-(ref) and (ref). Then, \begin{equation*} \sqrt{n}( \hat{\alpha}_{K-MD}-\alpha^{\ast }) \overset{d}{\to} N( \mathbf{0}_{d_{\alpha }\times 1},\Sigma_{K-MD}( \hat{P} _{0},\{ W_{k}:k\leq K\} ) ) , \end{equation*} where \begin{equation*} \Sigma _{K-MD}( \hat{P}_{0},\{ W_{k}:k\leq K\} ) \equiv \left\{ \begin{array}{c} ( \Psi _{\alpha }^{\prime }W_{K}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }W_{K}\times \\ \left[ \begin{array}{c} ( \mathbf{I}_{d_{P}}-\Psi _{P}\Phi _{K,P})^{\prime } \\ -( \Psi _{P}\Phi _{K,0})^{\prime } \\ -\Psi _{g}^{\prime }( \mathbf{I}_{d_{P}}+\Psi _{P}\Phi _{K,g})^{\prime } \end{array} \right]^{\prime }\left( \begin{array}{ccc} \Omega _{PP} & \Omega _{P0} & \Omega _{Pg} \\ \Omega _{P0}^{\prime } & \Omega _{00} & \Omega _{0g} \\ \Omega _{Pg}^{\prime } & \Omega _{0g}^{\prime } & \Omega _{gg} \end{array} \right) \left[ \begin{array}{c} ( \mathbf{I}_{d_{P}}-\Psi _{P}\Phi _{K,P})^{\prime } \\ -( \Psi _{P}\Phi _{K,0})^{\prime } \\ -\Psi _{g}^{\prime }( \mathbf{I}_{d_{P}}+\Psi _{P}\Phi _{K,g})^{\prime } \end{array} \right] \\ \times W_{K}^{\prime }\Psi _{\alpha }( \Psi _{\alpha }^{\prime }W_{K}^{\prime }\Psi _{\alpha })^{-1} \end{array} \right\} , \end{equation*} and $\{ \Phi _{k,0}:k\leq K\} $, $\{ \Phi _{k,P}:k\leq K\} $, and $\{ \Phi _{k,g}:k\leq K\} $ defined as follows. Set $\Phi _{1,P}\equiv \mathbf{0} _{d_{P}\times d_{P}}$, $\Phi _{1,0}\equiv \mathbf{I}_{d_{P}}$, $\Phi _{1,g}\equiv \mathbf{0}_{d_{P}\times d_{P}}$ and, for any $k\leq K-1$, \begin{align} \Phi _{k+1,P}& \equiv ( \mathbf{I}_{d_{P}}-\Psi _{\alpha }( \Psi _{\alpha }^{\prime }W_{k}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }W_{k}) \Psi _{P}\Phi _{k,P}+\Psi _{\alpha }( \Psi _{\alpha }^{\prime }W_{k}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }W_{k}, \notag \\ \Phi _{k+1,0}& \equiv ( \mathbf{I}_{d_{P}}-\Psi _{\alpha }( \Psi _{\alpha }^{\prime }W_{k}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }W_{k}) \Psi _{P}\Phi _{k,0}, \notag \\ \Phi _{k+1,g}& \equiv ( \mathbf{I}_{d_{P}}-\Psi _{\alpha }( \Psi _{\alpha }^{\prime }W_{k}\Psi _{\alpha })^{-1}\Psi _{\alpha }^{\prime }W_{k}) ( \mathbf{I}_{d_{P}}+\Psi _{P}\Phi _{k,g}) . \end{align}

We make several comments about this result. First, as in Theorem (ref), Theorem (ref) considers $K\in \mathbb{N}$ and fixed as $n\to \infty$, and it thus is free from the criticism raised by pesendorfer/schmidt-dengler:2010. Second, we note that $\Sigma_{K-MD}(\hat{P}_{0},\hat{g})$ can be consistently estimated using consistent estimators of $(\alpha^*,g^*)$ (e.g., $(\hat{\alpha} _{1-MD},\hat{g})$) and the asymptotic variance in Assumption (ref). Third, as expected, note that the sequence of weight matrices does not affect the asymptotic distribution of the $K$-MD estimator if $d_\alpha = d_P$. Finally, note that the asymptotic distribution of the $K$-PML estimator coincides with that of the $K$-MD estimator when $W_{k}\equiv \Omega _{PP}^{-1}$ for all $k \leq K$. We record this in the following corollary.

corollary[$K$-PML is a special case of $K$-MD] Fix $K\in \mathbb{N}$ arbitrarily and assume Assumptions (ref)-(ref). The asymptotic distribution of the $K$-PML estimator is a special case of that of the $K$-MD estimator with $ W_{k}\equiv \Omega _{PP}^{-1}$ for all $k\leq K$.

Theorem (ref) reveals that, unless $d_\alpha = d_P$, the asymptotic variance of the $K$-MD estimator is a complicated function of the number of iteration steps $K$ and sequence of limiting weighting matrices $\{ W_{k}:k\leq K\} $. For the remainder of the section, we focus on the typical case in which $d_\alpha <d_P$. A natural question to ask is the following: Is there an optimal way of choosing these parameters? In particular, what is the optimal choice of $K$ and $ \{ W_{k}:k\leq K\} $ that minimizes the asymptotic variance of the $K$-MD estimator? We devote the rest of this section to this question.

As a first approach to this problem, we consider the non-iterative $1$-MD estimator. As shown in pesendorfer/schmidt-dengler:2008, the asymptotic distribution of this estimator is analogous to that of a GMM estimator so we can leverage well-known optimality results. The next result provides a concrete answer regarding the optimal choices of $\hat{P}_{0}$ and $W_{1}$.

theorem[Optimality with $K=1$] Assume Assumptions (ref)-(ref). Let $\hat{\alpha}_{1-MD}^{\ast }$ denote the $1$-MD estimator with $\hat{P}_{0} = \tilde{P}$ that is asymptotically equivalent to $\hat{P}$ in the sense that \begin{equation} \sqrt{n}(\tilde{P}-P^{*}) = \sqrt{n}(\hat{P}-P^{*}) + o_{p}(1), \end{equation} and $W_{1} = W_{1}^{\ast }$ with \begin{align} W_{1}^{\ast } \equiv [ ( \mathbf{I}_{d_{P}}-\Psi _{P}) \Omega _{PP}( \mathbf{I} _{d_{P}}-\Psi _{P}^{\prime }) +\Psi _{g}\Omega _{gg}\Psi _{g}^{\prime } -\Psi _{g}\Omega _{Pg}^{\prime }( \mathbf{I}_{d_{P}}-\Psi _{P}^{\prime }) -( \mathbf{I}_{d_{P}}-\Psi _{P}) \Omega _{Pg}\Psi _{g}^{\prime }]^{-1}. \end{align} Then, \begin{equation*} \sqrt{n}( \hat{\alpha}_{1-MD}^{\ast }-\alpha^{\ast }) \overset{d}{\to} N( \mathbf{0}_{d_{\alpha }\times 1},\Sigma^{\ast }) , \end{equation*} with \begin{align} \Sigma^{\ast } \equiv ( \Psi _{\alpha }^{\prime }[ ( \mathbf{I}_{d_{P}}-\Psi _{P}) \Omega _{PP}( \mathbf{I} _{d_{P}}-\Psi _{P}^{\prime }) + \Psi _{g}\Omega _{gg}\Psi _{g}^{\prime } -\Psi _{g}\Omega _{Pg}^{\prime }( \mathbf{I}_{d_{P}}-\Psi _{P}^{\prime }) -( \mathbf{I}_{d_{P}}-\Psi _{P}) \Omega _{Pg}\Psi _{g}^{\prime } ]^{-1}\Psi _{\alpha })^{-1}. \end{align} Furthermore, $\Sigma _{1-MD}( \hat{P}_{0},W_{1}) -\Sigma^{\ast }$ is positive semidefinite for all $( \hat{P}_{0},W_{1})$, i.e., $\hat{\alpha}_{1-MD}^{\ast }$ is optimal among all $1$-MD estimators that satisfy our assumptions.

Theorem (ref) indicates that using $\hat{P}_{0}$ asymptotically equivalent to $\hat{P}$ and $W_{1}=W_{1}^{\ast }$ produces an optimal $1$-MD estimator. On the one hand, the choice of $\hat{P}_{0}$ is natural since $\hat{P}$ is an optimal preliminary estimator of the CCPs. Given this choice, the asymptotic distribution of the $1$-MD estimator is analogous to that of a standard GMM problem, and the corresponding optimal weight is $W_{1}=W_{1}^{\ast }$. Several comments are in order. First, as one would expect, $W_{1}^{\ast }$ coincides with the optimal weight matrix in the non-iterative analysis in pesendorfer/schmidt-dengler:2008. Second, while Eq.\ (ref) is satisfied by $\hat{P}_{0} = \hat{P}$, this condition can also be achieved by the other estimators of the CCPs mentioned in Section (ref), provided that they are sufficiently flexible. Third, we note that $W_{1}^{\ast } \not=\Omega _{PP}^{-1}$ in general, i.e., the optimal weight matrix does not coincide the one that produces the $1$-PML estimator. In fact, the $1$-PML estimator need not be an optimal $1$-MD estimator, i.e., $\Sigma _{1-MD}( \hat{P}_{0},\Omega _{PP}^{-1}) -\Sigma^{\ast }$ can be positive definite. Finally, we note that $W_{1}^{\ast }$ can be consistently estimated by replacing each component of Eq.\ (ref) by sample analogs. By standard asymptotic arguments, the optimality result in Theorem (ref) extends to the feasible optimal $1$-MD estimator that replaces $W_{1}^{\ast }$ by any consistent estimator. In practice, however, estimating $W_{1}^{\ast }$ can introduce substantial finite sample biases, which can deteriorate the performance of the estimator. For a discussion of these issues, see altonji/segal:1996 and horowitz:1998, among others.

We now move on to the general case with $K\geq 1$. According to Theorem (ref), the asymptotic variance of the $K$-MD estimator depends on the number of iteration steps $K \in \mathbb{N}$, the asymptotic distribution of $\hat{P}_{0}$, and the entire sequence of limiting weight matrices $\{ W_{k}:k\leq K\}$. In this sense, determining an optimal $K$-MD estimator appears to be a complicated task. The next result provides a concrete answer to this problem.

theorem[Invariance and optimality] Fix $K\in \mathbb{N}$ arbitrarily and assume Assumptions (ref)-(ref). Then, \begin{enumerate} • Invariance. Let $\hat{\alpha}_{K-MD}^{\ast }$ denote the $K$-MD estimator with $\hat{P}_{0} = \tilde{P}$ that is asymptotically equivalent to $\hat{P}$ in the sense of Eq.\ (ref), weight matrices $\{ W_{k}:k\leq K-1\} $ for steps $1,\ldots ,K-1$ (if $K>1$), and the corresponding optimal weight matrix in step $K$, which we assume to be well-defined. Then, \begin{equation*} \sqrt{n}( \hat{\alpha}_{K-MD}^{\ast }-\alpha^{\ast }) \overset{d}{\to} N( \mathbf{0}_{d_{\alpha }\times 1},\Sigma^{\ast }) , \end{equation*} where $\Sigma^{\ast }$ is as in Eq.\ (ref). • Optimality. Let $\hat{\alpha}_{K-MD}$ denote the $K$-MD estimator with $\hat{P}_{0}$ and weight matrices $\{ W_{k}:k\leq K\} $. Then, \begin{equation*} \sqrt{n}( \hat{\alpha}_{K-MD}-\alpha^{\ast }) \overset{d}{\to} N( \mathbf{0}_{d_{\alpha }\times 1},\Sigma_{K-MD}( \hat{P} _{0},\{ W_{k}:k\leq K\} ) ). \end{equation*} Furthermore, $\Sigma _{K-MD}(\hat{P}_{0},\{ W_{k}:k\leq K\})-\Sigma^{\ast }$ is positive semidefinite, i.e., $\hat{\alpha}_{K-MD}^{\ast }$ is optimal among all $K$-MD estimators that satisfy our assumptions. \end{enumerate}

Theorem (ref) is the main finding of this paper, and it establishes two central results regarding the optimality of the $K$-MD estimator. We begin by discussing the first one, referred to as “invariance”. The result focuses on a preliminary estimator of the CCPs that is asymptotically equivalent to $\hat{P}$. As explained earlier, this is a natural choice to consider since $\hat{P}$ is an optimal preliminary estimator of the CCPs. Given this choice, the asymptotic variance of the $K$-MD estimator depends on the entire sequence of weight matrices $ \{W_{k}:k\leq K\}$. While the dependence on the first $K-1$ weight matrices is fairly complicated, the dependence on the last weight matrix (i.e., $W_{K}$) resembles that of the weight matrix in a standard GMM problem. Standard GMM results characterize the optimal choice for $W_{K}$, {\it given the sequence of first $K-1$ weight matrices}.\footnote{The result requires this optimal choice for $W_{K}$ to be well defined. For typical choices of the first $K-1$ weight matrices, this additional requirement was not restrictive in our Monte Carlo simulations.} In principle, one might expect that the resulting asymptotic variance depends on the first $K-1$ weight matrices. The “invariance” result reveals that this is not the case. In other words, for $\hat{P}_{0}=\hat{P}$ and an optimal choice of $W_{K}$, the asymptotic distribution of the $K$-MD estimator is invariant to the first $K-1$ weight matrices, or even $K$. Furthermore, the resulting asymptotic distribution coincides with that of the optimal $1$-MD estimator obtained in Theorem (ref).

The “invariance” result is the key to the second result in Theorem (ref), referred to as “optimality”. This second result characterizes the optimal choice of $\hat{P}_{0}$ and $ \{W_{k}:k\leq K\}$ for $K$-MD estimators. The intuition of the result is as follows. First, since $\hat{P}$ is the optimal estimator of CCPs, it is intuitive that optimality requires this choice or possibly something asymptotically equivalent. Second, it is also intuitive that optimality requires an optimal choice of $W_K$, {\it given the sequence of first $K-1$ weight matrices}. At this point, our “invariance” result indicates that the asymptotic distribution does not depend on $K$ or the first $K-1$ weight matrices. From this, we can then conclude that the $K$-MD estimator with $ \hat{P}_{0}$ asymptotically equivalent to $\hat{P}$ and an optimal last weight matrix $W_{K}$ (given any first $K-1$ weight matrices) is efficient among all $K$-MD estimators.

Theorem (ref) implies two important corollaries regarding the optimal $1$-MD estimator discussed in Theorem (ref). The first corollary is that the optimal $1$-MD estimator is efficient in the class of all $K$-MD estimators that satisfy the assumptions of Theorem (ref). In other words, additional policy iterations do not provide efficiency gains relative to the optimal $1$-MD estimator. The intuition behind this result is that the multiple iteration steps of the $K$-MD estimator are merely reprocessing the sample information, i.e., no new information is added in each iteration step. Provided that the criterion function is optimally weighted, the $1$-MD estimator is capable of processing the sample information in an efficient manner.

The second corollary of Theorem (ref) is that the $K$-PML estimator is usually not efficient, and can be feasibly improved upon. In particular, under the assumptions of Theorem (ref), the optimal $1$-MD estimator (i.e.\ with $\hat{P}_{0}$ asymptotically equivalent to $\hat{P}$ and $ W_{1}=W_{1}^{\ast }$) is more or equally efficient than the $K$-PML estimator.\footnote{One special case in which the $K$-PML estimator is as efficient as the optimal $1$-MD estimator is when both: (a) the zero Jacobian property holds (i.e.\ $\Psi_P = {\bf 0}_{dP \times dP}$) and (b) Step 1 of the estimation procedure does not include a preliminary estimator of $g^*$ (i.e.\ $\theta^* = \alpha^*$). It is worthwhile to point out that (a) or (b) taken in isolation would not be sufficient to imply that the $K$-PML estimator is as efficient as the optimal $1$-MD estimator.}

The results up to this point show that the optimal $1$-MD estimator is efficient among all $K$-MD estimators. Given these findings, one might wonder whether the optimal $1$-MD estimator is efficient. Section (ref) in the appendix compares the asymptotic distribution of the optimal $1$-MD estimator with that of the MLE of $\alpha^*$. Under appropriate conditions, we show that these two asymptotic distributions coincide, and so the optimal $1$-MD estimator is indeed efficient among all regular estimators.

We illustrate the results of this section by revising the example of Section (ref) with the parameter values considered in Section (ref). The asymptotic variance of $\hat{\alpha}_{K-MD}=(\hat{\lambda}_{RN,K-MD},\hat{ \lambda}_{EC,K-MD})$ is given by

equation[equation omitted — 355 chars of source]

where $\{\Phi _{k,P0}:k\leq K\}$ is defined by $\Phi _{k,P0}\equiv \Phi _{k,P}+\Phi _{k,0}$, with $\{\Phi _{k,P}:k\leq K\}$ and $\{\Phi _{k,0}:k\leq K\}$ as in Eq.\ (ref). For any true parameter vector and any $K\in \mathbb{N} $, we can numerically compute Eq.\ (ref).

In Section (ref), we considered three specific parameter values of $(\lambda _{RN}^{\ast },\lambda _{EC}^{\ast })$ that produced an asymptotic variance of the $K$-PML estimator of $\lambda_{RN}^{\ast }$ decreased, increased, and fluctuated with $K$. We now compute the optimal $K$-MD estimator of $\lambda_{RN}^*$ for the same parameter values. The results are presented in Figures (ref), (ref), and (ref), respectively.\footnote{According to the “invariance” result in Theorem (ref), there are multiple asymptotically equivalent ways of implementing the optimal $K$-MD estimator. For concreteness, we set the weight matrix optimally in each iteration step.} These graphs illustrate the findings in Theorem (ref). In accordance to the “invariance” result, the asymptotic variance of the optimal $K$-MD estimator does not vary with the number of iterations $K$. Also in accordance to the “optimality” result, the asymptotic variance of the optimal $K$-MD estimator is lower than that of any other $K$-MD estimator. In turn, since the asymptotic distribution of the $K$-PML estimator is a special case of the asymptotic distribution of the $K$-MD estimator, the asymptotic variance of the optimal $K$-MD estimator is lower than that of the $K$-PML estimator for all $K \in \mathbb{N}$. Combining both results, the optimal non-iterative $1$-MD estimator is both computationally convenient and efficient among the estimators under consideration.

figure[figure omitted — 530 chars of source]
figure[figure omitted — 529 chars of source]
figure[figure omitted — 534 chars of source]

Monte Carlo simulations

This section investigates the performance of $K$-PML and $K$-MD estimators in a Monte Carlo simulation. Our main goal in this section is to confirm that our asymptotic analysis can provide a reasonable approximation provided that the size of the sample is sufficiently large. We simulate data using the two-player dynamic entry game described in Section (ref). Recall that this model is specified up to the parameters $(\lambda _{RS}^{\ast },\lambda _{RN}^{\ast },\lambda _{FC,1}^{\ast },\lambda _{FC,2}^{\ast },\lambda _{EC}^{\ast },\beta^{\ast })$. For simplicity, we assume that the researcher knows $(\lambda _{RS}^{\ast },\lambda _{FC,1}^{\ast },\lambda _{FC,2}^{\ast },\beta^{\ast })$ and wants to estimate $\alpha^{\ast }\equiv (\lambda _{RN}^{\ast },\lambda _{EC}^{\ast })$. We consider the three specific parameter values used to illustrate the theoretical results in Sections (ref) and (ref). For each parameter value, we compute equilibrium CCPs numerically by solving the fixed point problem in Eq.\ (ref) up to a small tolerance level.\footnote{We did not encounter evidence of multiple equilibria in our computations.} Derivatives of $\Psi$ were computed numerically, and we note that $\Psi_P$ differs from the zero matrix for all parameter values (i.e., the Zero Jacobian property does not hold).

Our simulation results are the average of $S=10,000$ independent datasets $\{ ( \{ a_{j,i}:j \in J\} ,x_{i},x_{i}^{\prime }) :i\leq n\} $ that are i.i.d.\ distributed according to the econometric model. We show results for sample sizes $n \in \{500,~1,000,~2,000\}$. Compared to typical empirical applications, these sample sizes are admittedly large relative to the state space (e.g., the empirical application in aguirregabiria/mira:2007 has $n=189$ and $d_P=800$), but this choice reflects our asymptotic framework in which the state space is fixed as the number of observations diverges. With these sample sizes, we can verify our findings regarding the behavior of the asymptotic variance with the number of iterations. On the other hand, for smaller sample sizes, we often find that our asymptotic approximation does not provide a good representation of the distribution of these estimators, possibly due to the large finite sample bias in the preliminary estimator of the CCPs.

For the sake of brevity, we show simulation results for the estimation of $\lambda _{RN}^{\ast }$, which was the object of illustration in Sections (ref) and (ref). Qualitatively similar results hold for the estimation of $\lambda _{EC}^{\ast }$, and are available upon request. We use $\hat{P}_0 = \hat{P}$ as the preliminary estimator (recall that there is no $g^*$ in this example).

Table (ref) provides results for the first parameter value, i.e., $(\lambda _{RN}^{\ast },\lambda _{EC}^{\ast },\lambda _{RS}^{\ast },\lambda _{FC,1}^{\ast },\lambda _{FC,2}^{\ast },\beta^{\ast }) = (2.8,0.8,0.7,0.6,0.4,0.95)$. Recall from previous sections that this parameter value produces an asymptotic variance of the $K$-PML estimator of $\lambda_{RN}^{\ast }$ that decreases with $K$ (see Figures (ref) and (ref), repeated at the bottom of Table (ref)).

Let us first focus on the results for the $K$-PML estimator. The simulation results closely resemble the predictions from the asymptotic approximation. As mentioned earlier, this is an expected consequence of using sample sizes that are large relative to the dimensionality of the model. First, the empirical variance and mean squared error are extremely close, indicating that the asymptotic bias is almost negligible. Second, the empirical variance is decreasing with $K$ and is close to the one predicted by our asymptotic analysis. Finally, the computational cost of the $K$-PML estimator is low relative to the optimal $K$-MD estimator, and rises linearly with $K$.

Next, we turn attention to the optimal $K$-MD estimator. Recall that the “invariance” result in Theorem (ref) indicates that there are multiple asymptotically equivalent ways of implementing the optimal $K$-MD estimator. Throughout this section, the optimal $K$-MD estimator is a {\it feasible} estimator of the optimal $K$-MD estimator derived in Theorem (ref) in the last iteration step and the $k$-PML weight matrix in steps $k=1,\dots,K-1$. According to our theoretical results, this feasible optimal $K$-MD estimator is optimal among $K$-MD estimators, has zero asymptotic bias, and has an asymptotic variance that does not change with $K$. For the most part, these predictions are satisfied in our simulations. Once again, this is a consequence of using sample sizes that are large relative to the dimensionality of the model. First, we note that the empirical variance and mean squared error are again extremely close, and so the finite-sample bias is almost negligible. Second, the empirical variance is close to the one predicted by our asymptotic analysis. As predicted by the “optimality” result in Theorem (ref), the feasible optimal $K$-MD estimator is more efficient than the $K$-PML estimator. For most values of $K$ under consideration, the empirical variance of the $K$-MD estimator appears to be invariant to $K$, especially for the larger sample sizes. However, we find that the empirical variance decreases slightly between $K=1$ and $K=2$. Our first-order asymptotic analysis cannot explain this last empirical finding. This anomalous behavior for low values of $K$ is analogous to the one found for the $K$-PML estimator by aguirregabiria/mira:2002 and rationalized by the higher-order analysis in kasahara/shimotsu:2008. In Section (ref), we show that a high-order analysis can explain these anomalous simulation results for the $K$-MD estimator. Finally, the computational cost of the optimal $K$-MD estimator is considerably higher than the $K$-PML estimator. In particular, computing the optimally weighted $1$-MD, $2$-MD, and $3$-MD estimators takes us roughly $33\%$, $75\%$, and $80\%$ more time than computing the $20$-PML estimator, respectively. The reason behind this difference is that the $K$-PML estimator does not require estimating an optimal weight matrix, while the optimal $K$-MD estimator does. This optimal weight matrix for the $K$-MD estimator requires estimating $\Psi_P$ and $\Psi_\alpha$ (the latter only when $K>1$). As noted by an anonymous referee, the computational cost of estimating $\Psi_P$ in large models can be significant, as its dimension grows quadratically with $d_P$.

table[table omitted — 3,941 chars of source]

Table (ref) provides results for the second parameter value, i.e., $(\lambda _{RN}^{\ast },\lambda _{EC}^{\ast },\lambda _{RS}^{\ast },\lambda _{FC,1}^{\ast },\lambda _{FC,2}^{\ast },\beta^{\ast }) = (2,1.8,0.2,0.01,0.03,0.95)$. Recall that this parameter value produced an asymptotic variance of the $K$-PML estimator of $\lambda_{RN}^{\ast }$ that increases with $K$ (see Figures (ref) and (ref), repeated at the bottom of Table (ref)). In turn, Table (ref) provides results for the third parameter value, i.e., $(\lambda _{RN}^{\ast },\lambda _{EC}^{\ast },\lambda _{RS}^{\ast },\lambda _{FC,1}^{\ast },\lambda _{FC,2}^{\ast },\beta^{\ast }) = (2.2,1.45,0.45,0.22,0.29,0.95)$. This parameter value produced an asymptotic variance of the $K$-PML estimator of $\lambda_{RN}^{\ast }$ that wiggles with $K$ (see Figures (ref) and (ref), repeated at the bottom of Table (ref)).

The simulation results for these two parameter values are qualitatively similar to the ones obtained for the first parameter value and, for the most part, support our theoretical conclusions. First, both estimators have very little empirical bias. Second, all the estimators have an empirical variance that is very close to the one predicted by the asymptotic analysis. In particular, the empirical variance of the $K$-PML estimator is increasing in $K$ for the second parameter value and wiggles for the third parameter value. Third, in most cases, the empirical variance of the optimal $K$-MD estimator is lower than that of the $K$-PML estimator. Fourth, the empirical variance of the optimal $K$-MD estimator is invariant to $K$ except for small values of $K$ for which it is decreasing. One notable difference relative to the previous simulation is that the range of iterations over which the empirical variance decreases now extends between $K=1$ and $K=5$. As explained earlier, we attribute this to the high-order analysis that we develop in Section (ref). Finally, the comparison of computational costs is similar to the one described for the first parameter value.

table[table omitted — 3,941 chars of source]
table[table omitted — 3,945 chars of source]

For the sake of comparison, we also compute the one-step MLE described in aguirregabiria/mira:2007, which does not belong to the class of $K$-PML or $K$-MD estimators. The one-step MLE is the result of taking a Newton step in the MLE problem based on an initial estimator of $(\alpha^*,P^*)$ that is consistent and is a fixed point in Eq.\ (ref). This initial estimator serves as a starting point of the Newton step, and is also used to consistently estimate the efficient score and information matrix. aguirregabiria/mira:2007 explain that the one-step MLE has the advantage of being as efficient as the MLE. They propose implementing the one-step MLE with an initial estimator given by their NPL estimator, i.e., the $\infty$-PML estimator. As an approximation to this, we compute the one-step MLE with the $20$-PML estimator as the initial estimator.\footnote{We do not use $K=\infty$ for reasons explained in Section (ref), and 20 is the largest value of $K$ considered in our simulations.} The simulation results for the one-step MLE are provided in Table (ref). These show that the one-step MLE is more efficient than the $K$-PML estimator, and is as efficient as optimal $K$-MD estimator, especially when $K\geq 10$. These findings are in line with our analysis in Section (ref), which reveals that, under appropriate conditions, the optimal $K$-MD estimator is as efficient as the MLE and, consequently, as efficient as the one-step MLE. Finally, the computational costs of estimating the one-step MLE are similar to that of an optimally weighted $20$-MD estimator.

table[table omitted — 2,301 chars of source]

Conclusions

This paper investigates the asymptotic properties of a class of estimators of the structural parameters in dynamic discrete choice games. We consider $K$-stage policy iteration (PI) estimators, where $K$ denotes the number of policy iterations employed in the estimation. This class nests several estimators proposed in the literature. By considering a “maximum likelihood” criterion function, the $K$-stage PI estimator becomes the $K$-PML estimator in aguirregabiria/mira:2002,aguirregabiria/mira:2007. By considering a “minimum distance” criterion function, $K$-stage PI estimator defines a new $K$-MD estimator, which is an iterative version of the estimators in pesendorfer/schmidt-dengler:2008 and pakes/ostrovsky/berry:2007. Since we consider an asymptotic framework with fixed $K \in \mathbb{N}$ as $n\to\infty$, our analysis is not affected by the problems described in pesendorfer/schmidt-dengler:2010.

First, we establish that the $K$-PML estimator is consistent and asymptotically normal for any $K \in \mathbb{N}$. This complements findings in aguirregabiria/mira:2007, who focus on $K=1$ and $K$ large enough to induce convergence of the estimator. Furthermore, we show under certain conditions that the asymptotic variance of the $K$-PML estimator can exhibit arbitrary patterns as a function of $K$. In particular, we show that by changing the parameter values in a typical dynamic discrete choice game, the asymptotic variance of the $K$-PML estimator can increase, decrease, or even be non-monotonic with $K$.

Second, we also establish that the $K$-MD estimator is consistent and asymptotically normal for any $K$. Its asymptotic distribution depends on the choice of the weight matrix. For a specific weight matrix, the $K$-MD estimator has the same asymptotic distribution as the $K$-PML estimator. We investigate the optimal choice of the weight matrix for the $K$-MD estimator. Our main result shows that an optimally weighted $K$-MD estimator has an asymptotic distribution that is invariant to $K$. This appears to be a novel result in the literature on PI estimation for games, and it is particularly surprising given the findings in aguirregabiria/mira:2007 for $K$-PML estimators.

The main result in our paper implies two important corollaries regarding the optimal $1$-MD estimator (derived by pesendorfer/schmidt-dengler:2008). First, the optimal $1$-MD estimator is optimal among all $K$-MD estimators. In other words, additional policy iterations do not provide efficiency gains relative to the optimal $1$-MD estimator. Second, the optimal $1$-MD estimator is more or equally efficient than any $K$-PML estimator for all $K \in \mathbb{N}$. Finally, Section (ref) provides appropriate conditions under which the optimal $1$-MD estimator has the same asymptotic distribution as the MLE, and it is thus efficient among regular estimators.

We explored our theoretical findings in Monte Carlo simulations. Provided that the sample size is large enough, the simulation evidence supports the conclusions of our asymptotic analysis. The $K$-PML and the optimal $K$-MD estimators have negligible empirical bias and have an empirical variance that is very close to the one predicted by the asymptotic analysis. In most cases, the empirical variance of the optimal $K$-MD estimator is lower than that of the $K$-PML estimator. Also, it appears to be invariant to $K$ except for very small values of $K$ for which it is decreasing in $K$. The behavior for low values of $K$ is analogous to the one found by aguirregabiria/mira:2002. Inspired by the analysis in kasahara/shimotsu:2008, Section (ref) studies the high-order properties of the optimal $K$-MD estimator and rationalizes the simulation result for low values of $K$.

There are several topics excluded from this paper, such as allowing for permanent unobserved heterogeneity in the dynamic discrete choice model. In principle, this could be achieved via the developments in aguirregabiria/mira:2007 or arcidiacono/miller:2011. We plan to address this topic in future work.