Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,231 characters · 18 sections · 80 citation commands
Testing homogeneity in dynamic discrete games in finite samples
\linespread{1.5}
In applications of dynamic discrete games, practitioners often assume that the conditional choice probabilities and the state transition probabilities are invariant across time and markets.\footnote{In this paper, we use “market” to denote a cross-sectional unit.} We refer to this as the “homogeneity assumption” in dynamic discrete games. This is a convenient assumption, as it allows the estimation of the model's structural parameters by pooling data from multiple markets and from many time periods.
Despite the widespread use of the homogeneity assumption in dynamic discrete games, it is plausible for this condition to fail in applications. We now provide a few examples. First, a game could suffer from a structural break in the model, which would invalidate the homogeneity assumption across time. Second, markets could be affected by persistent heterogeneity that is observed by the players but not by the researcher (e.g., arcidiacono/miller:2011). This would invalidate the homogeneity assumption across markets. Third and relatedly, there may be multiplicity of equilibria, and different markets could be playing different equilibria. The literature has considered hypothesis testing for the multiplicity of equilibria in games. In particular, depaula/tang:2012 propose a test for the multiplicity of equilibria across markets in static games, while otsu/pesendorfer/takayashi:2016 do this in the context of dynamic games.
In this paper, we propose a hypothesis test for the homogeneity assumption. That is, our test is designed to capture the various possible violations of the homogeneity assumption described in the previous paragraph (both across markets and over time). Our test is implemented via Markov chain Monte Carlo (MCMC) methods, and it is justified by the theory of randomization tests (cf.\ lehmann/romano:2005). While our MCMC test is not a standard randomization test, we establish its validity by coupling it with an underlying randomization test that is valid in finite samples yet computationally infeasible in practically relevant applications. Our contribution is to show that the distribution generated by our MCMC algorithm approximates the underlying randomization test as the number of (user-defined) MCMC draws diverges. In this sense, we interpret our proposed MCMC algorithm as a computationally feasible way to implement a computationally infeasible underlying randomization test. Moreover, our results are {\it valid in finite samples} in the sense that they hold for any fixed and finite number of players, markets, and time periods. This is an important aspect of our contribution, as the datasets used in empirical applications often have a small number of time periods and markets. For example, our empirical application is based on ryan:2012, and has only $n=23$ markets and either $T=9$ or $T=10$ time periods.
Our methodology is especially well suited for our testing problem in dynamic discrete games. As with standard randomization tests, the quality of our test depends on the richness of possible transformations exploited by our MCMC algorithm. There are three aspects of the typical framework in dynamic discrete games that allow us to generate a rich set of such transformations. First, dynamic discrete game models often impose independence across markets and have a Markovian structure. This provides the basis for our randomization-based methodology. Second, the data used in dynamic discrete games are usually either naturally discrete or discretized by the researcher. This is important for our test, as it relies on transformations defined by “conditional” permutations of data, i.e., permutations of the value of the state and the action across markets or time periods for which other related states coincide. Third, dynamic discrete game models often assume that actions are independent conditional on the state. This allows us to simplify the implementation of the transformations that generate our randomization-based methodology. It is worth pointing out that our inference methodology should apply to economic problems beyond the dynamic discrete game setup provided that these three aspects hold.
The econometric framework considered in this paper is arguably very general. It includes the single-agent dynamic discrete choice model (e.g., rust:1987,hotz/miller:1993,hotz/miller/sanders/smith:1994,aguirregabiria/mira:2002) and the Markov equilibrium dynamic game model (e.g., pakes/ostrovsky/berry:2007,aguirregabiria/mira:2007,bajari/benkard/levin:2007,pesendorfer/schmidt-dengler:2008,pesendorfer/schmidt-dengler:2010). Furthermore, it includes the Markov dynamic game model of aguirregabiria/magesan:2020, which allows some players to have biased beliefs. Importantly, our methodology does not impose functional forms on the primitive structure of the dynamic game (e.g., utility functions, state transition functions, etc.). We consider this advantageous, as functional form assumptions are hard to justify solely based on economic arguments, and are thus prone to misspecification.
In a recent paper, otsu/pesendorfer/takayashi:2016 propose several hypothesis tests for dynamic discrete games. Two of their proposals are directly related to the problem considered in our paper.\footnote{The other two testing methodologies are less related to our paper. One test assumes that the state distribution is in its steady state. This condition is not commonly imposed in the literature, and our test does not require it. The other test they propose is based on the frequencies of states conditional on the state distribution in the first period.} Specifically, they consider a method to test the homogeneity across markets of the conditional choice probabilities and the state transition probabilities, under the maintained assumption that these functions are time-homogeneous. Their inference method is based on the bootstrap, and its validity is shown in an asymptotic framework in which the number of time periods $T$ diverges to infinity. However, $T$ is often small in applications. Besides the aforementioned application of ryan:2012 with $T=9$ or $T=10$, we can mention sweeting:2011 with $T=4$, collard-wexler:2013 with $T=24$, and dunne/klimek/roberts/xu:2013 with $T=5$. When $T$ is small, inference methods based on large-$T$ asymptotics, like the one provided by otsu/pesendorfer/takayashi:2016, may yield inaccurate results. In contrast, our methodology is valid for any of the data dimensions, including $T$.
The most critical step of our MCMC algorithm is based on the so-called Euler Algorithm, described in kandel/yossi/unger/winkler:1996. In related work, besag/mondal:2013 uses the Euler Algorithm to test whether a time series of data has a time-homogeneous Markov structure. Relative to this work, our paper incorporates several essential features of dynamic Markov discrete games. First, we recognize that the dataset in a typical dynamic game has information about actions and states. Second, our construction exploits the typical economic structure imposed in dynamic games, such as the conditional independence assumption (i.e., conditional on the current state variable, the current action variable is independent of the past information). Finally, while besag/mondal:2013 mainly focuses on data from a single market (i.e., a single sequence in their setup), our MCMC algorithm exploits the possibility that the data includes observations from multiple markets.\footnote{Section 5 in besag/mondal:2013 briefly describes a few alternative ways to extend their methodology to the case of the multiple sequences of Markov chains (i.e., multiple markets). Their description does not include the formal statistical properties. In this paper, we propose a methodology that differs from theirs, describe its implementation in detail, and prove its validity by connecting it with the theory of randomization tests.} This is a valuable aspect of our contribution, as the datasets used in empirical applications usually include from data multiple markets, e.g., ryan:2012 with $n=23$, sweeting:2011 with $n=102$, collard-wexler:2013 with $n=1,600$, and dunne/klimek/roberts/xu:2013 with $n=639$.
We explore the performance of our hypothesis test in Monte Carlo simulations. Our results show that our method provides excellent size control even in small samples, and can successfully detect relatively small deviations from the homogeneity hypothesis. In these two accounts, our test appears to work favorably in comparison with the bootstrap-based test in otsu/pesendorfer/takayashi:2016. These favorable results appear to extend even when the discrete data have many support points, which is typical in applications. In our empirical example, we investigate the homogeneity of the decisions in the U.S.\ Portland cement industry data used in ryan:2012. This is a key assumption in ryan:2012, as it allows him to pool data from multiple markets to estimate the model's parameters. Unlike otsu/pesendorfer/takayashi:2016's test, our test finds no evidence against the homogeneity hypothesis in the data. We implement our test using the Julia package HomogeneityTestBBU, which is publicly available at the GitHub repository.\footnote{Use Pkg.add("HomogeneityTestBBU") to install the package in Julia. Package documentation is available in \href{https://jacksonbunting.github.io/HomogeneityTestBBU.jl/dev/}{https://jacksonbunting.github.io/HomogeneityTestBBU.jl/dev/}.}
The rest of the paper is organized as follows. Section (ref) describes the dynamic discrete choice model and the hypothesis test. Section (ref) specifies our hypothesis test and its implementation via our MCMC algorithm. Section (ref) establishes the main theoretical results of the paper. The main technical insight is that our hypothesis test is an approximate way of implementing an underlying randomization test that is finite-sample valid yet computationally infeasible. Section (ref) provides the empirical application. In Section (ref), we evaluate the performance of our test in finite samples via Monte Carlo simulations. Section (ref) concludes. The paper's appendix collects all of the proofs, auxiliary results, and computational details related to our proposed MCMC algorithm.
We begin by describing the dynamic discrete game under consideration. We observe the outcome of $n$ markets in which $J$ players choose actions over $T $ time periods. Our setup allows for $J=1$, i.e., single-agent problems, or $J>1$, i.e., multiple-agent games. This paper's inference results are valid for all finite $n$, $ T $, and $ J $.
We consider a setup in which the observed actions and state variables are discretely distributed. This is common in the dynamic discrete choice literature, where the state and action variables are often naturally discrete or discretized by the researcher. For every market $i=1,\ldots ,n$ and period $t=1,\ldots ,T$, let $A_{i,t}$ be the random variable that specifies the actions chosen by the players in market $i$ and period $t$, and let $S_{i,t}$ be the random variable that specifies the state variable of market $i$ and period $t$. We use $\mathcal{S}$ to denote the common support of $S_{i,t}$ and $\mathcal{A}$ to denote the common support of $A_{i,t}$. We define the following $n\times T$ matrices:
In this notation, the data are then given by $$ X~\equiv~(S,A). $$ By definition, the support of $X$ is given by $\mathcal{X} \equiv\mathcal{S}^{nT} \times \mathcal{A}^{nT}$.
The following assumption is standard in much of the literature on dynamic discrete games.
Assumption (ref) has three parts. Assumption (ref)(a) imposes that markets are independently distributed. Assumption (ref)(b) indicates that the observations of state and actions are a Markov process. Assumption (ref)(c) imposes that the current actions are independent of past information once we condition on the current state. Assumptions (ref)(b)-(c) are high-level restrictions that are typically imposed on the equilibrium strategies used by the players. In particular, they follow from the assumption that players use Markov strategies (e.g., maskin/tirole:2001), as assumed in pakes/ostrovsky/berry:2007,aguirregabiria/mira:2007,bajari/benkard/levin:2007,pesendorfer/schmidt-dengler:2008. These conditions are imposed even in models in which the players' beliefs are allowed to be out of equilibrium, i.e., do not coincide with the true equilibrium probabilities (e.g., aguirregabiria/magesan:2020). Finally, we clarify that Assumption (ref) refers to {\it observed} states and actions. As such, it may fail if there are state or action variables that are unobserved to the researcher and influence the distribution of their observed counterparts.
Assumption (ref) is the only maintained assumption to study the validity of our hypothesis test. Notably, we do not impose functional forms on the primitive structure of the dynamic game, such as the utility functions, the state transition functions, or the discount factors. As explained earlier, we view this as a virtue of our methodology, as these restrictions can be hard to justify based only on economic arguments. Relatedly, while one could improve the statistical power of our test by exploiting functional form restrictions on the primitives of the dynamic game, they inevitably carry the risk of producing invalid inference when these are misspecified.
We now introduce the necessary notation to express our hypothesis of interest. We use $\sigma_{i,t}$ to denote the conditional choice probability for market $i$ and period $t$, i.e., for every $(s,a)\in\mathcal{S\times A}$, $$ \sigma_{i,t}(a|s)~\equiv ~P(A_{i,t}=a|S_{i,t}=s). $$ We use $f_{i,t+1}$ to denote the state transition probability from period $t$ to $t+1$ for market $i$, i.e., for every $(s,a,s^{\prime})\in\mathcal{S}\times\mathcal{A}\times\mathcal{S}$ , $$ f_{i,t+1}(s^{\prime}|a,s)~\equiv ~P(S_{i,t+1}=s^{\prime}|(S_{i,t},A_{i,t})=(s,a)). $$ Finally, we use $m_{i}(s)$ to denote the marginal state distribution for market $i$ in period 1, i.e., for every $s\in\mathcal{S}$, $$ m_{i}(s)~\equiv ~P(S_{i,1}=s). $$ With this notation in place, we specify our hypothesis testing problem in the next section.
Our goal is to test whether the “homogeneity assumption” holds in the data, i.e., whether the conditional choice probabilities and state transition probabilities are homogeneous across time and markets. That is,
Note that $H_{0}$ in (ref) represents two types of homogeneity: time and market homogeneity, and involves two functions: conditional choice probabilities and state transition probabilities. In this sense, our hypothesis test evaluates four homogeneity conditions: time homogeneity of the conditional choice probabilities, market homogeneity of the conditional choice probabilities, time homogeneity of the state transition probabilities, and market homogeneity of the state transition probabilities. Under Assumption (ref), a rejection of $H_{0}$ would be indicative that one or more of these homogeneity conditions is violated, suggesting it is not appropriate to pool data across markets and time periods.
As discussed in the introduction, there may be many possible reasons for the homogeneity assumption to fail and, in general, it may be difficult to distinguish among the possible reasons. For example, given the recent literature on separately identifying equilibrium selection and market-specific permanent unobserved heterogeneity aguirregabiria/mira:2019,luo/xiao/xiao:2022, it may be difficult to distinguish between these two possible causes for failure of the homogeneity assumption. Nevertheless, in certain applications, one may feel comfortable that some of the conditions are satisfied and should be part of our maintained assumptions. For example, in a given application, one may be confident that the conditional choice probability and state transition probability are time-homogeneous. Then, one could reinterpret $H_{0}$ as testing the market homogeneity of the conditional choice probabilities and state transition probabilities. Also, if one is confident that market and time homogeneity holds for some subsets of the time periods (e.g., before and after a policy change), then $H_{0}$ may be reinterpreted as testing homogeneity of the conditional choice probabilities and state transition probabilities across the subsets of periods.
Under Assumption (ref) and $H_{0}$, Lemma (ref) in the appendix shows that the likelihood of the data ${X}=({S},{A})$ evaluated at any realization $\tilde{X}=(\tilde{S},\tilde{A})\in\mathcal{X}$ is as follows:
This expression reveals that the markets are independently distributed (Assumption (ref)(a)), but they are not necessarily identically distributed because $m_i(\cdot)$ may depend on $i$. That is, even though the conditional choice probabilities and state transition probabilities are homogeneous under $H_{0}$, markets can still be heterogeneous due to differences in their initial state values. This is a desired feature in our testing problem, as the dynamic discrete choice literature usually allows the initial state distribution to be market-specific.
We conclude the section with an observation about the type of economic models considered in this paper. From an econometric viewpoint, our goal is to evaluate the homogeneity assumption (i.e., $H_{0}$ in (ref)) using discrete data that satisfies Assumption (ref). The discreteness of the data is necessary for our MCMC algorithm to find data transformations that can deliver non-trivial power. We motivated this problem using dynamic discrete choice games because they are an important class of models ideally suited to this econometric framework. However, it is worth highlighting that our methodology applies to any other discrete panel-data model that satisfies Assumption (ref).\footnote{We thank an anonymous referee for this observation.} To our knowledge, the main roadblock to applying our test beyond dynamic discrete games and single-agent problems is the requirement that the data be discrete.
In this paper, we propose to reject $H_{0}$ in (ref) whenever the significance level $\alpha$ is larger than or equal to our $p$-value, which we denote by $\hat{p}_{K}$. That is,
In turn, our $p$-value $\hat{p}_{K}$ is the result of constructing $K$ transformations of the data via our MCMC algorithm, which is specified in Section (ref). This MCMC algorithm produces $K$ sequential transformations of the data $X$, denoted by $(X^{(1)},\ldots,X^{(K)})$. Our $p$-value is then computed as follows
where $\tau:\mathcal{X} \to \mathbb{R}$ denotes the test statistic designed to detect departures from $H_{0}$ in the data.
One notable feature of our hypothesis test is that its validity does not depend on the choice of the test statistic (see Theorem (ref)). However, the power of our test depends on this choice. Example (ref) below specifies several test statistics considered in the related literature and describes the types of heterogeneity they are designed to detect. In practice, the choice of the test statistic should be guided by the type of heterogeneity one is most interested in detecting.\footnote{While our test is valid for any test statistic, this does not extend to when multiple instances of our test are implemented with several test statistics. This can create a multiple-testing problem, thereby invalidating the conclusions.}
Our MCMC algorithm requires some notation. Let $I= (I_1,I_2)$ denote an arbitrary pair of markets $I_1$ and $I_2$ in the data, i.e., $I_1,I_2 \in \{1,2,\dots,n\}$. We allow for $I_1=I_2$. We use $\mathcal{I}$ to denote the collection of all such pairs of markets, i.e., $|\mathcal{I}|=n^2$. We also define several sets.
In words, $R_{S}(I,\breve{S})$ is the set of all state configurations that result from permuting the state data $\breve{S}$ subject to conditions (a)-(c), which we now interpret. First, condition (a) indicates that the initial value of the state variable must remain unchanged across markets. The reason behind this restriction is that our framework does not restrict the initial state distribution (i.e., $\{m_i(\tilde{S}_{i,1})\}_{i=1}^{n}$ in (ref)). In turn, conditions (b)-(c) imply that the aggregate state transition frequencies across all markets $i=1,\dots,n$ must remain constant. This restriction is achieved by requiring the state transition frequencies to remain invariant for each market $i \not\in \{I_1,I_2\}$ (by condition (b)) and on aggregate for markets $i \in \{I_1,I_2\}$ (by condition (c)). The main reason behind breaking an aggregate restriction into conditions (b) and (c) is computational tractability. Under Assumption (ref) and $H_{0}$, conditions (a)-(c) imply that each state configuration in $R_{S}(I,\breve{S})$ has the same value of the likelihood function, provided that it is paired with a suitable action configuration. These suitable action configurations are precisely those in next definition.
By definition, $R_{A}(\tilde{S},(\breve{S},\breve{A}))$ is the set of action configurations that result from permuting the action data $\breve{A}$ subject to conditions (a)-(b), which we explain next. Condition (a) implies that the aggregate state and action transition frequencies across all markets $i=1,\dots,n$ remain constant. Condition (b) imposes an analogous requirement for the terminal period. Under Assumption (ref) and $H_{0}$, these restrictions imply that the hypothetical data $(\breve{S},\breve{A})$ has the same likelihood as the state configuration $\tilde{S}$ paired with any action configuration in $R_{A}(\tilde{S},(\breve{S},\breve{A}))$.
Before explaining how $R_{S}(I,\breve{S})$ and $R_{A}(\tilde{S},(\breve{S},\breve{A}))$ are used in our MCMC algorithm, we illustrate their computation in a relatively simple example. While the conditions in Definitions (ref) and (ref) are not conceptually complicated, the example reveals that computing these sets explicitly requires thoughtful consideration, even in a relatively simple case.
Having introduced and illustrated Definitions (ref) and (ref), we now specify our MCMC algorithm.
At each step $k=2,\dots,K$, our MCMC algorithm randomly permutes actions and states in the data. By construction, the algorithm implies the following transition probabilities for all $k=2,\dots,K$, $X^{(1)},\ldots,X^{(k-1)}\in\mathcal{X}$, $I\in\mathcal{I}$, and $\tilde{X}= (\tilde{S},\tilde{A})\in\mathcal{X}$,
We note that (ref) and (ref) are well defined, as both denominators can be shown to be positive.
Each iteration of the MCMC algorithm (ref) involves three steps. Step 1 is computationally and conceptually straightforward. Steps 2 and 3 require randomly drawing state and action configurations uniformly over the sets $ R_{S}(I^{(k)},S^{(k-1)})$ and $ R_{A}(S^{(k)},X^{(k-1)})$, respectively. As we argued in the context of Example (ref), these sets may be difficult to enumerate even for simple data configurations. Importantly, our MCMC algorithm does not require us to enumerate these sets, but rather sample from them uniformly. The remainder of this section provides an overview of how we implement Steps 2 and 3. We defer to Section (ref) for details.
Step 2 requires sampling $S^{(k)}$ uniformly from the set $ R_{S}(I^{(k)},S^{(k-1)})$. Given a pair of markets $I^{(k)}$ and state data in $S^{(k-1)}$, the restrictions considered in $ R_{S}(I^{(k)},S^{(k-1)})$ are relatively hard to implement. To construct a feasible implementation of step 2, we crucially rely on the Euler Algorithm (see kandel/yossi/unger/winkler:1996,besag/mondal:2013 for details). In particular, when both markets in $I^{(k)}$ are equal (i.e., $I^{(k)}=(i,i)$ for $i=1,2,\dots,n$), step 2 can be implemented by applying the Euler Algorithm for each market. Our marginal contribution in step 2 is to extend the Euler Algorithm to the case where the markets in $I^{(k)}$ differ. Our proposal is to concatenate the state information from both markets in $I^{(k)}$ and repeatedly apply the Euler Algorithm until two conditions hold: the initial state is the same in each market (i.e., $\tilde{S}_{i,1}=\breve{S}_{i,1}$) and the $T^{\text{th}}$ state in the first market satisfies $\tilde{S}_{I_1,T}\in\{\breve{S}_{I_1,T},\breve{S}_{I_2,T}\}$. The properties of the Euler Algorithm ensure that these two conditions imply the resulting chain belongs to $R_{S}(I^{(k)},S^{(k-1)})$. Importantly, these two conditions are far simpler to verify than the conditions in Definition (ref), which is one reason that our implementation is computationally feasible whereas the enumeration approach is not.\footnote{Not only are conditions simpler to verify, they may be verified without needing to complete the full $2\times T$ length chain. For example, if the candidate $\tilde{S}_{I_1,T}$ is not an element of $\{\breve{S}_{I_1,T},\breve{S}_{I_2,T}\}$, one may stop after constructing a $T$ length chain. Indeed, it is sometimes possible to verify that the second condition fails when the chain being constructed is of length $1<t<T$.} Relative to the market-by-market version of the algorithm, our modification typically generates a much larger set of data permutations, which tends to improve the power properties of our hypothesis test. We provide additional information about step 2 of the MCMC algorithm in Section (ref) of the appendix, where we specify the original Euler Algorithm (Algorithm (ref)) and our modification (Algorithm (ref)), and we formally show that the latter exactly implements step 2 (see Lemma (ref)). Algorithm (ref) draws $S^{(k)}$ uniformly from $ R_{S}(I^{(k)},S^{(k-1)})$, because we repeatedly sample uniformly from a superset of $R_{S}(I^{(k)},S^{(k-1)})$ until the realization belongs to $ R_{S}(I^{(k)},S^{(k-1)})$.
Step 3 requires sampling $A^{(k)}$ uniformly from the set $R_{A}(S^{(k)},X^{(k-1)})$. Given data in $X^{(k-1)}$ and state data in $S^{(k)}$, the restrictions considered in $R_{A}(S^{(k)},X^{(k-1)})$ are relatively easy to impose (compared to those in $ R_{S}(I^{(k)},S^{(k-1)})$). As a consequence, step 3 is computationally light. All we need to do is to permute the action data in $A^{(k-1)}$ subject to the simple restrictions in $R_{A}(S^{(k)},X^{(k-1)})$. Further details of step 3 are provided in Section (ref) of the appendix, where we specify an algorithm (Algorithm (ref)) and we prove that it implements step 3 (see Lemma (ref)).
We open this section with the main theoretical result of this paper.
Theorem (ref) establishes that the proposed test in (ref) controls size as the length of the MCMC draws $K$ diverges. We remark that $K$ is under the control of the researcher, who can increase $K$ to guarantee the convergence in (ref). Remarkably, Theorem (ref) holds regardless of the number of markets $n$, time periods $T$, and players $J$, which remain constant in our analysis. In addition, and as promised in Section (ref), this result also holds irrespective of the specific choice of test statistic $\tau(X)$ used in the construction of the $p$-value in (ref). Finally, we note that the inequality in equation (ref) could be turned into equality by changing (ref) to a random decision rule whenever $\phi_K(X)=\alpha$. We decided against this modification for the sake of simplicity.
An important practical consideration is how one should choose the number of MCMC draws $K$ in a given application. According to Theorem (ref), the size control of our test is guaranteed as $K$ diverges. The main drawback of increasing $K$ is the additional computation burden of implementing our test. In this sense, we recommend choosing $K$ as large as computationally possible. However, our theoretical results and practical experience can be combined to provide a more concrete recommendation regarding $K$. First, our theoretical results in later sections establish that the $p$-value in (ref) used to implement our test converges as $K$ diverges.\footnote{In particular, see Lemma (ref) and the related result in (ref), which are building blocks of Theorem (ref).} Second, our experience from the empirical application and the Monte Carlo simulations suggests that the outcome of our test tends to become stable for a sufficiently large $K$. In conclusion, we recommend considering large values of $K$ (as large as computationally possible) and deciding on a value for which the test decision appears to become stable.
The key insight behind Theorem (ref) is the connection between our hypothesis test and the literature on randomization tests (see lehmann/romano:2005). In particular, Theorem (ref) follows from showing that the $p$-value in (ref) approximates the $p$-value of an underlying randomization test for $H_{0}$ in (ref) that is computationally infeasible. Recall that randomization tests enjoy validity in finite samples under suitable conditions. This explains why Theorem (ref) does not require the number of markets $n$, time periods $T$, or players $J$ to grow.
The remainder of this section develops the connection between our hypothesis test and the underlying randomization test for $H_{0}$. It is organized as follows. Section (ref) provides an alternative representation of the likelihood of the data under Assumption (ref) and $H_{0}$. This result allows us to define a sufficient statistic of the data under these conditions, denoted by $U(X)$. Section (ref) relates our MCMC algorithm to a transformation group of the data, $\mathbf{G}$, which does not change the value of the sufficient statistic $U(X)$. Section (ref) defines the underlying randomization test for $H_{0}$ based on the transformation group $\mathbf{G}$, and argues that it is both finite-sample valid and computationally infeasible. Finally, Section (ref) shows that our MCMC-based test in (ref) can successfully approximate the underlying randomization test as the number of MCMC draws diverges.
The next result provides an alternative representation of the likelihood of the data under Assumption (ref) and $H_{0}$ in (ref).
From this result, we can deduce the following corollary.
Corollary (ref) implies that, under Assumption (ref) and $H_{0}$, a transformation of the data that maintains $U(X)$ will not change the value of the likelihood. This observation provides the basis of the underlying randomization test.
In this section, we show that our MCMC algorithm is a transformation group of $\cal{X}$ that preserves the value of $U(X)$. See lehmann/romano:2005 for the definition of the notion of a transformation group.
Our proposed MCMC algorithm can be understood as an iteration of transformations to the data $X$. In particular, $X^{(1)}=X$ is the identity transformation, $X^{(2)}$ follows from applying Steps 1-3 to $X$, $X^{(3)}$ follows from applying Steps 1-3 twice to $X^{(2)}$, and so forth. More formally, each iteration of our MCMC algorithm applies a transformation from a particular transformation group. To define this properly, we first require the following definition.
Lemma (ref) in the appendix shows that ${\bf G}(I)$ is a transformation group. By Definition (ref), ${\bf G}(I)$ is the transformation group representation of Steps 2-3 of our MCMC algorithm. Given a randomly chosen pair of markets $I^{(k)}$ in step 1, steps 2-3 obtain the next element of the Markov chain $X^{(k)} = (S^{(k)},A^{(k)})$ by selecting a randomly chosen element of $\{g(X^{(k-1)})\colon{g}\in{\bf G}(I^{(k)})\}$. In this sense, Steps 2-3 of our MCMC algorithm are a specific way of choosing a particular transformation in ${\bf G}(I^{(k)})$.
By the description in the previous paragraph, our MCMC algorithm applies a randomly chosen transformation in ${\bf G}(I)$ for random pairs of markets $I$, and iteratively applies them to the data. These iterative transformations are related to the set that we define next.
See Example (ref) for an illustration of $\mathbf{G}$. The next result states that $\mathbf{G}$ is a transformation group with desirable properties.
The properties shown in Lemma (ref) imply that we can use ${\bf G}$ to define a valid randomization test. We do this in Section (ref).
Following lehmann/romano:2005, we can use the transformation group ${\bf G}$ to define the underlying randomization test. This test rejects $H_{0}$ in (ref) whenever the significance level $\alpha$ is larger than or equal to the randomization $p$-value, which we denote by $\hat{p}$. That is,
where
By the arguments in lehmann/romano:2005, the randomization test in (ref) is finite-sample valid. We record this in the next result.
The finite-sample validity in Lemma (ref) makes the randomization test in (ref) an excellent candidate for testing $H_{0}$. Unfortunately, this randomization test is not computationally feasible in practice. This is because the test requires working with the transformation group $\mathbf{G}$, which is typically impossible to enumerate in practice. We illustrate this in Example (ref) in the appendix, where we enumerate $\mathbf{G}$ in two simple examples with $n=2$ markets, $T=2$ time periods, and binary actions and states. Given the challenges presented even by these very simple cases, it is not hard to imagine that $\mathbf{G}$ is computationally impossible to enumerate in realistic data settings.
In the randomization testing literature, it is not uncommon to work with a huge transformation group $\mathbf{G}$. As lehmann/romano:2005 explains, one can still implement a random version of the test in (ref) by drawing randomly from $\mathbf{G}$ {\it in a uniform fashion}. This point is routinely exploited in standard settings to construct tests based on permutations or sign changes. However, to the best of our knowledge, there is no known feasible way of obtaining such random draws in the current context without fully enumerating $\mathbf{G}$.
The previous paragraphs explain why the underlying randomization test in (ref) is computationally infeasible and, thus, we cannot directly exploit its finite-sample validity. The main technical insight of our paper is that our hypothesis test in (ref) is an approximate way of implementing the computationally infeasible underlying randomization test in (ref). In particular, the following section formally states that our MCMC-based $p$-value in (ref) approximates the underlying $p$-value in (ref) as the length of the MCMC diverges.
Our main theoretical result is Theorem (ref), which shows that the test in (ref) controls size as the number of MCMC draws $K$ diverges to infinity. The following lemma provides the fundamental ingredient to prove this result.
Lemma (ref) shows that, as the number of MCMC draws diverges, the conditional distribution based on the MCMC algorithm converges to the conditional distribution of the computationally infeasible underlying randomization test described in Section (ref). It is worth noting that Lemma (ref) considers $K\to\infty$ while the complexity of the underlying randomization test, characterized by $|\mathbf{G}|$, stays constant. While we do not derive formal results, we expect that, as $|\mathbf{G}|$ increases, a larger number of MCMC draws $K$ is required to achieve a specific level of approximation. For related discussions on diagnosing convergence in MCMC algorithms, see robert/casella:2004.
By applying Lemma (ref) with $t=\tau(X)$, we can deduce that the $p$-value in (ref) approximates the $p$-value in (ref) as the number of MCMC draws $K$ diverges. That is, conditional on $X$,
By combining this observation with the finite-sample validity of the underlying randomization test in (ref) (Lemma (ref)), it follows that our proposed MCMC-test becomes valid as the number of MCMC draws $K$ diverges. This argument provides the intuition behind Theorem (ref), and why it holds regardless of the number of markets $n$, time periods $ T $, and players $ J $.
Our analysis in this paper focuses on the properties of our test under the null hypothesis. While analyzing our test's power properties is very desirable, we consider this to be a formidable task within our finite-sample setting. The rejection rate of our test depends on the specification of the dynamic discrete game (i.e., conditional choice probabilities, state transition probabilities, and marginal distributions specified under the alternative hypothesis). To the best of our knowledge, obtaining general results for all possible specifications of the dynamic discrete game in finite samples is impossible. An alternative to the finite-sample power analysis would be to consider asymptotic power results with a diverging number of markets $n$, time periods $T$, support points $|\mathcal{X}|$, or all three. Note that the current results in this paper are finite sample valid, and do not require such an asymptotic framework. The behavior of our test is expected to vary with the specific asymptotic framework under consideration. While such asymptotic power results have been developed for some randomization tests in the literature (e.g., see lehmann/romano:2005), these do not apply to approximate randomization tests like the one proposed in this paper. Developing this extension of our results seems out of the scope of our current contribution. As an admittedly imperfect substitute for a general power analysis, Section (ref) of our paper explores the power properties of our proposed test in several empirically relevant economic models. All of our simulation evidence suggests that our test has desirable power properties.
In this section, we revisit the application in ryan:2012, as studied in otsu/pesendorfer/takayashi:2016. ryan:2012 considers a dynamic discrete game to study the welfare costs of the 1990 Amendments to the Clean Air Act on the U.S.\ Portland cement industry. He develops a dynamic oligopoly game based on ericson/pakes:1995, and estimates it using the two-stage method developed by bajari/benkard/levin:2007. This method's first stage is to estimate optimal entry, exit, and investment decisions as a function of production capacity, and it relies on the assumption that markets are homogeneous. Our hypothesis test can be used to investigate the validity of this assumption.
We use the same data as in otsu/pesendorfer/takayashi:2016. For each year in 1980-1998 and 23 geographically separated U.S.\ markets, we observe the sum of the production capacities for all the firms in that market. Table (ref) provides summary statistics of this aggregate production capacity before and after the 1990 Amendments, and Figure (ref) provides the corresponding histogram.
These data represent the result of the firms' optimal entry, exit, and investment decisions in the dynamic game estimated by ryan:2012. We follow otsu/pesendorfer/takayashi:2016 and discretize the market production capacity into 50 bins with equal intervals of 250 thousand tons each (0-250 thousand tons, 250-500 thousand tons, and so on). For each $i=1,\ldots ,n=23$ and year $t=1,\ldots ,19$, we use $A_{i,t}\in\mathcal{A}=\{ 1,\ldots ,50\} $ to denote the production capacity bin. The state variable in any market is the previous period's action, i.e.,
and so $S_{i,t}\in\mathcal{S}=\{ 1,\ldots ,50\} $. We note that (ref) implies that the state transition probabilities are homogeneous (given by $f_{i,t+1}(s^{\prime}|a,s)=1\{s^{\prime}=a\}$), and so $H_{0}$ in (ref) is equivalent to the homogeneity of the conditional choice probabilities.
Following ryan:2012 and otsu/pesendorfer/takayashi:2016, we allow the 1990 Amendments to affect the decision of the firms. We then test the homogeneity of the conditional choice probabilities for two subsets of data: before and after 1990. That is, we test the following hypotheses:
We note that the before and after 1990 samples used to test the hypotheses in (ref) and (ref) have a relatively small number of time periods ($T=10$ and $T=9$ for (ref) and (ref), respectively) and markets (in both cases, $n=23$). This represents an ideal scenario for our proposed test, as its validity does not rely on either one of these dimensions diverging. We also note that both of these dimensions are smaller than the support of the data, i.e., $|\mathcal{S}| =|\mathcal{A}|=50$.
Table (ref) shows the results of applying our procedure to test the hypotheses in (ref) and (ref). Following the literature, we use the test statistics in (ref). As explained in Example (ref), these test statistics compare market-specific conditional choice probabilities with their pooled counterpart and are thus specifically designed to detect heterogeneity across markets. This objective seems appropriate for this empirical application, as ryan:2012's methodology relies on the homogeneity of the data before and after the 1990 Amendments.\footnote{It is relevant to note that these test statistics would be largely ineffective in detecting the presence of a structural break (such as the 1990 Amendments) if the break impacts equally all markets in the economy. This happens because the market-specific conditional choice probabilities would average over time and coincide with their pooled counterpart.} At a significance level of $ \alpha =5\%$, we do not reject the homogeneity of the conditional choice probabilities. Our tests were implemented with $K=50,000$, but our hypothesis testing decision (i.e., non-rejection for standard significance levels) remains invariant for any $K>10,000$. Using a standard desktop computer, our Julia package completed our test with $K=50,000$ in 4.8 minutes for the subsample before 1990 and 1.9 minutes for the subsample after 1990. The computation time increases linearly with $K$.
For contrast, Table (ref) also shows the results of bootstrap-based test proposed by otsu/pesendorfer/takayashi:2016. As opposed to our test, their methods reject the hypothesis of homogeneity of the conditional choice probabilities in the sample prior to 1990. Since both tests rely on the same test statistic, these differences are entirely driven by the differences in the $p$-values. Table (ref) reveals that our test and the one proposed by otsu/pesendorfer/takayashi:2016 can produce different conclusions. This is also clearly shown in our Monte Carlo simulations in Section (ref). It is natural to inquire which hypothesis test is correct about the homogeneity of the sample before 1990. Of course, this is impossible to determine with certainty in an empirical application. However, we consider that our Monte Carlo evidence in Section (ref) may shed light on this matter. These simulations suggest that in data settings similar to those in the empirical application (i.e., with $|\mathcal{S}|$ large relative to $n$ and $T$), our test controls size adequately while the test by otsu/pesendorfer/takayashi:2016 can suffer from overrejection.
In this section, we explore the performance of our proposed test in Monte Carlo simulations. We consider three simulation designs. Our first design is based on the duopoly entry game in pesendorfer/schmidt-dengler:2008. Our second design is based on our empirical application in Section (ref). Our third simulation is based on a dynamic single-agent human capital formation model in keane/wolpin:1997. Our three designs offer a comprehensive description of the finite sample behavior of our proposed methodology, as they exhibit significant differences in crucial aspects of the dynamic discrete problem, including the number of support points, the number of players, and the nature of their strategic interaction.
This Monte Carlo design is also used by otsu/pesendorfer/takayashi:2016, which follows from the duopoly entry game in pesendorfer/schmidt-dengler:2008. The simulated data are generated by two oligopolistic firms deciding whether to enter or not into $n$ markets, and over $T$ time periods. This dynamic game has multiple equilibria, which we exploit to generate departures from the homogeneity assumption.
In each period $t=1,\dots,T$ and market $i=1,\dots,n$, there are four possible actions in this game: $A_{i,t}=1$ denotes that neither firm entered the market, $A_{i,t}=2$ denotes that only firm 2 enters, $A_{i,t}=3$ denotes that only firm 1 enters, and $A_{i,t}=4$ denotes that both firms enter. This implies that $\mathcal{A}=\{ 1,2,3,4\}$. As in the empirical application, the state variable in any market is the previous period's action (i.e., (ref) holds). This implies that $\mathcal{S}=\{ 1,2,3,4\}$ and that the state transition probabilities are homogeneous (and given by $f_{i,t+1}(s^{\prime}|a,s)~=~1\{s^{\prime}=a\}$). As a consequence, $H_{0}$ in (ref) is equivalent to
The data produced by this game is a matrix $X=(S,A)\in\mathcal{X}$ constructed exactly as in otsu/pesendorfer/takayashi:2016. We simulate data from a mixture of two data-generating processes: DGP 1 and DGP 2. They represent Markov perfect equilibria of the dynamic game, which differ in the conditional choice probabilities $\sigma(a|s)$. The matrices of conditional choice probabilities in DGP 1 and DGP 2 are $$ \left(
\right) and \left(
\right), $$ respectively, where the index of the column indicates the state $s \in \mathcal{S}=\{ 1,2,3,4\}$, and the index of the row indicates the value of the action $a \in \mathcal{A}=\{ 1,2,3,4\}$. Each market is sampled independently. Market $i=1,\dots,n$ behaves according to DGP 1 with probability $\lambda$ and DGP 2 with probability $1-\lambda$. Therefore, $\lambda\in[0,1] $ represents the average proportion of markets in DGP 1. Each market is initialized with a state equal to $1$, and we simulate the corresponding action according to the conditional choice probabilities. This, in turn, determines the next period's state according to \eqref{eq:trivialstate transition probability_app}, i.e., $S_{i,t+1}=A_{i,t}$. We then proceed iteratively until we have simulated $T+100$ periods for each market. The first 100 periods are discarded, producing a sample of $T$ periods for $n$ markets, which are then observed by the researcher.
For each simulated data, we implement our proposed test in (ref) with $K=20,000$. We consider simulations with $n\in\{ 20,40,80,160\} $, $T\in\{ 5,10,20,40,80\} $, and $\lambda\in\{ 1,0,0.5,0.9\} $. As explained earlier, $\lambda $ represents the proportion of markets that are in DGP 1. If $\lambda =1$ or $\lambda =0$, all markets are sampled from the same distribution, and so the conditional choice probabilities are homogeneous across markets, i.e., $H_{0}$ holds. In turn, if $\lambda =0.5$ or $\lambda =0.9$, each data is composed of markets from both distributions, and so the conditional choice probabilities are not homogeneous across markets, i.e., $H_{0}$ fails. Note that $ \lambda =0.5$ generates data in which both distributions are equally represented, and so the heterogeneity in the conditional choice probabilities is more salient. On the other hand, the case with $\lambda =0.9$ produces data with a vast majority of markets in DGP 1, and so the heterogeneity in the conditional choice probabilities is harder to detect. For each simulation design, we compute rejection rates based on $2,000$ independently simulated datasets.
The results from the Monte Carlo simulation are shown in Table (ref) for $\lambda\in\{0,1\}$ and Table (ref) for $\lambda\in\{0.5,0.9\}$, respectively. For the sake of comparison, we also include the results from the test proposed by otsu/pesendorfer/takayashi:2016. Their test compares the same test statistics in (ref) with critical values based on the bootstrap. As mentioned earlier, they show the validity of their test in an asymptotic framework with $T\to\infty $ and $n$ fixed. In contrast, our main result in Theorem (ref) is valid for any finite $n$ and $T$.
Table (ref) reveals that our test achieves relatively good size control for all values of time periods and market sizes under consideration. The table shows the result of running 80 hypothesis tests for different data configurations that satisfy $H_{0}$ (four market sizes, five time periods, two test statistics, and two distributions). Across these 80 numbers, our proposed test has an average rejection rate of 5.1%, a standard deviation of 0.04%, and a range of 4.1% to 6.35%. We note that Theorem (ref) implies that our test should not produce overrejection as $K$ becomes large, but it is silent about the possibility of underrejection. Table (ref) reveals that our test does not seem to suffer from underrejection in these simulations. For otsu/pesendorfer/takayashi:2016's test, the average rejection rate is also 5.1%, but with a standard deviation is 2.2% and a range of 0.6% to 13.5%. We note that the larger rejection rates occur in simulations with $T=5$, which is not unexpected for a test whose validity is proven in an asymptotic framework in which $T$ diverges.
Table (ref) explores the performance of these tests for data configurations that do not satisfy $H_{0}$. We begin by explaining the results that are common to both hypothesis tests. First, recall that $\lambda$ denotes the proportion of the $n$ markets in the data that are in DGP 1. As $\lambda$ becomes closer to either zero or one, the data are increasingly coming from a single distribution, making the departure from the $H_{0}$ harder to detect. Second, as the number of markets $n$ grows, the inference methods gain more evidence of the presence of multiplicity, resulting in higher rejection rates. The same phenomenon occurs as the number of time periods $T$ increases. Third, we find that the hypothesis tests implemented with $\tau_2(X)$ tend to produce higher rejection rates than those implemented with $\tau_1(X)$. This finding appears consistent with the large $T$ optimality result in otsu/pesendorfer/takayashi:2016. We now turn to compare rejection rates between the two tests. In most simulation designs, our test appears to have a higher or equal rejection rate than otsu/pesendorfer/takayashi:2016's test. The few exceptions occur in designs with $n=20$ and $T\in \{5,10\}$, which correspond to designs in which otsu/pesendorfer/takayashi:2016's test overrejects under $H_0$. This suggests that any power advantage of their test relative to ours may disappear when considering a size-corrected version.
In this subsection, we explore the performance of our test in two DGPs related to the empirical application in Section (ref). The first data-generating process (DGP 1) satisfies $H_{0}$ in (ref), and the second one (DGP 2) does not. DGP 1 represents a discretized version of the pre-1990 Amendments data in the empirical application (i.e., $t\leq T_0 \equiv 9$), and is generated as follows. First, we discretize the data into $|\mathcal{S}|$ evenly spaced bins, which we denote by $\{\tilde{S}_{i,t}:i=1,\dots,n,~t=1,\dots,T\}$. As in the empirical application, the state variable in any market is the previous period's action (i.e., (ref) holds). For each $i=1,\dots,n$, we simulate $S_{i,1}$ independently from the pre-1990 Amendments discretized distribution, i.e., for all $s \in \mathcal{S}=\{1,\dots,|\mathcal{S}|\}$,
Second, for each $i=1,\dots,n$ and $t=1,\dots, T-1$, we simulate $A_{i,t}$ independently across markets according to the pre-1990 Amendments choice probabilities, i.e., for all $s,a \in \mathcal{S}=\{1,\dots,|\mathcal{S}|\}$,
where $S_{i,t}=A_{i,t-1}$ for all $i=1,\dots,n$ and $t=2,\dots, T$. Since the production capacity in each market and time period is drawn according to the market- and time-homogeneous conditional choice probabilities in (ref), DGP 1 satisfies $H_{0}$.
DGP 2 represents an economy in which half of the markets are negatively impacted by the 1990 Amendments, and is generated as follows. In the pre-Amendments periods (i.e., $t\leq T_0 \equiv 9$), DGP 2 coincides exactly with DGP 1. In the post-Amendments periods (i.e., $t> T_0$), the data is independently generated across markets in the following fashion. For markets with even index $i$ (i.e., $i=2,4,\dots,22$), the production level is distributed as in the pre-1990 Amendments periods (i.e., as in (ref)). For markets with odd index $i$ (i.e., $i=1,3,\dots,23$), the production level is uniformly chosen to be weakly lower, i.e., for all $s,a \in \mathcal{S}=\{1,\dots,|\mathcal{S}|\}$,
That is, markets with an even index $i$ are unaffected by the 1990 Amendments, while markets with an odd index $i$ are negatively affected. As in DGP 1, the state variable in any market is the previous period's action (i.e., (ref) holds). The structural change caused by the 1990 Amendments implies that DGP 2 does not satisfy $H_{0}$.
For each simulated data, we implement our proposed test in (ref) with $K=20,000$. We consider simulations with $n=23 $, $T=19$, $T_0 =9$, and $|\mathcal{S}|\in \{5, 10, 15, \dots, 80\}$. The first three parameters are those in the empirical application, which has $|\mathcal{S}|=50$ bins. For each simulation design, we compute rejection rates based on $2{,}000$ independently simulated datasets.
The results from the Monte Carlo simulations are presented in Table (ref). We include results for our test and the one proposed by otsu/pesendorfer/takayashi:2016 with bootstrap-based $p$-values (see their Section 5 for details). We first describe results under DGP 1, i.e., when $H_{0}$ holds. Our test achieves good size control for all discretizations under consideration. Across the 10 hypothesis tests that satisfy $H_{0}$ (five discretizations and two test statistics), our test has an average rejection rate of 5.4%, with a standard deviation of 0.3%, and a range of 4.9% to 5.9%. These numbers also reveal that our test does not exhibit underrejection. On the other hand, otsu/pesendorfer/takayashi:2016's test suffers from overrejection, and this problem tends to exacerbate as $|\mathcal{S}|$ increases. For instance, when $|\mathcal{S}|$ is as in the empirical application (i.e., $|\mathcal{S}|=50$), their test has a rejection rate of 23.1% for $\tau_1(X)$ and 26% for $\tau_2(X)$, more than 4 times higher than the nominal size of $\alpha=5\%$. This issue may be explained by the fact that their validity result relies on $T \to \infty$, and these simulations only have $T=19$, which is smaller than $|\mathcal{S}| \in \{5, 10, 15, \dots, 80\}$.
We now turn to the simulations under DGP 2, i.e., when $H_0$ fails. The results show that our test has nontrivial power for all values of $|\mathcal{S}| \in\{5, 10, 15, \dots, 80\}$. If particular, when $|\mathcal{S}|$ is as in the empirical application (i.e., $|\mathcal{S}|=50$), our test has a rejection rate of 36% for $\tau_1(X)$ and 21.6% for $\tau_2(X)$, which are considerably larger than the nominal size of $\alpha=5\%$. As one may expect, the power of our test tends to decrease with $|\mathcal{S}|$. This is because the power of our test is based on permutations with common state values, which become increasingly rare as $|\mathcal{S}|$ grows. Also noteworthy is that, for $|\mathcal{S}|>20$, our test implemented with $\tau_1(X)$ has more power than when implemented with $\tau_2(X)$, which is an opposite pattern to that in the previous Monte Carlo simulations. Finally, we recognize that otsu/pesendorfer/takayashi:2016's test achieves much higher rejection rates, but these occur in the context of overrejection under the null hypothesis.
We now consider Monte Carlo simulations based on the human capital formation single-agent model in keane/wolpin:1997. In this model, individuals choose an occupation each period throughout their working life. A distinctive feature of this model is that each individual has permanent unobserved heterogeneity, representing “innate talents” that are unobserved by the econometrician. We use this aspect of the model to generate departures from the homogeneity assumption.
We simulate datasets with $n=100$ individuals choosing among occupations over $T=10$ time periods. The choice of $T=10$ is inspired by the data used for keane/wolpin:1997's structural estimation. In each period $t=1,\dots,T$, individual $i=1,\dots,n$ chooses between home production, white-collar work, blue-collar work, schooling, and military work, which we denote as $A_{i,t}\in\mathcal{A}=\{1,2,3,4,5\}$, respectively. The state variable for individual $i$ in period $t$, denoted as $S_{i,t}$, encodes the experience vector in each occupation, i.e., $$S_{i,t} ~=~ \Big(\sum\nolimits_{s<t}1\{A_{i,s}=a\}:a \in \mathcal{A}\Big).$$ By definition, $S_{i,t+1}$ is a deterministic function of $A_{i,t}$ and $S_{i,t}$, and so the state transition probability is homogeneous. As a consequence, $H_{0}$ in (ref) is equivalent to
We draw the data $X=(S,A)\in\mathcal{X}$ as a mixture of three agent types: type 1, type 2, and type 3, each of which is motivated by keane/wolpin:1997. Type 1 represents a baseline individual with $\sigma(a|s) = \hat{\sigma}(a|s)$, where $\hat{\sigma}(a|s)$ denotes the empirical counterpart computed from the pooled NLSY79 sample across all individuals and time periods. In the pooled sample, $S_{i,t}$ has 417 support points. Type 2 represents an individual with innate talent for white-collar work, resulting in $\sigma(2|s)= \min\{1,\max\{5\hat\sigma(2|s),1/2\}\}$ and all other choice probabilities scaled appropriately, i.e., $\sigma(a|s)=\hat\sigma(a|s)/(1-\sigma(2|s))$ for $a\neq 2$. Finally, type 3 is the analog of type 2 but for blue-collar work, i.e., $\sigma(3|s)= \min\{1,\max\{5\hat\sigma(3|s),1/2\}\}$ and $\sigma(a|s)=\hat\sigma(a|s)/(1-\sigma(3|s))$ for $a\neq 3$.
We simulate independent datasets characterized by the parameter \(\lambda = (\lambda_1, \lambda_2)\). In each dataset, individuals are independently drawn, and are of type 1 with probability $1 - \lambda_1 - \lambda_2$, type 2 with probability $\lambda_{1}$, and type 3 with probability $\lambda_{2}$. We simulate datasets from two DGPs. The first DGP uses $\lambda = (0,0)$, which produces a homogeneous sample composed of individuals of type 1, i.e., $H_0$ in (ref) holds. The second DGP uses $\lambda=(0.230, 0.556)$, which generates a sample with unobserved heterogeneity, i.e., $H_0$ in (ref) fails. The values in $\lambda=(0.230, 0.556)$ correspond to the empirical frequencies estimated in keane/wolpin:1997.
The Monte Carlo results are shown in Table (ref). We include results for our test and the one proposed by otsu/pesendorfer/takayashi:2016 with bootstrap-based $p$-values. Under $H_0$, our test exhibits relatively good size control, with perhaps a slight tendency to overreject. On the other hand, otsu/pesendorfer/takayashi:2016's test suffers from considerable overrejection, which may be explained by the fact that the current empirical setting with $T=10$ cannot be well represented by their asymptotic results as $T\to \infty$. Under $H_1$, our test exhibits small yet non-trivial power. As in our first design, our test implemented with $\tau_2(X)$ has more power than when implemented with $\tau_1(X)$. Given their results under $H_0$, we do not dwell on the performance of otsu/pesendorfer/takayashi:2016's test under $H_1$.
This paper proposes a hypothesis test for the “homogeneity assumption” in dynamic discrete games. Our test is implemented by an MCMC algorithm and does not rely on functional forms imposed by the researcher. We show that our test is valid as the (user-defined) number of MCMC draws diverges, regardless of the number of markets and time periods in the data. This result contrasts with that of available methods in the literature, which require the number of time periods to diverge. We establish our validity result by showing that our proposed test is an MCMC approximation to a computationally infeasible underlying randomization test, which is valid in finite samples. Our Monte Carlo simulations reveal that our test has an excellent performance in finite samples, both in terms of size control and power.