Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,226 characters · 21 sections · 85 citation commands
Robust Minimum Distance Inference in Structural Models
Minimum distance inference is commonly used in structural econometric models, with applications varying from dynamic and static games of imperfect information (see, e.g., pesendorfer2008asymptotic), dynamic stochastic general equilibrium (DSGE) models (see, e.g., rotemberg1997optimization, amato2003estimation, and christiano2005nominal), to any economic model estimated by simulated-based methods such as Simulated Method of Moments (SMM) or Indirect Inference (II) (see, e.g., nakamura2018high). Two key assumptions in the literature are that structural parameters are point-identified and the corresponding Jacobian matrix is of full column rank. In this paper, we relax these two conditions and propose inference for a structural parameter of interest $\beta$ that is robust to lack of under-identification of other nuisance structural parameter $\alpha$. Moreover, it is not necessary to have prior knowledge of the degree of identification of the vector $\alpha$. There is substantial theoretical and empirical evidence that structural parameters in econometric models are generally not identified, see the references below. However, the level of underidentification is often unknown to the researcher. We propose an inference method that is robust to such knowledge and also to singularities in the optimal weighting matrix. Therefore, the proposed method provides a simple robust alternative to the standard econometric practice. Notably, our approach does not require the identification of the $\beta$ parameter either, as the restricted model is used to compute the test statistic following a Lagrange multiplier or score-type approach. As a result, this testing procedure is particularly well-suited to make inferences on calibrated (fixed) parameters in structural macroeconomic and microeconomic applications. We refer to this application of our method as testing the validity of calibrated parameters.
In structural models estimated by minimum distance methods the degree of identification of the structural parameters is often unknown. This is so because of a lack of an analytic solution for the Jacobian (see, e.g., mcfadden1989method on SMM or gourieroux1993indirect on II), or because the rank of the Jacobian depends on an unknown population parameter (see, e.g, pesendorfer2008asymptotic on Dynamic Games, or kalouptsidi2021identification on counterfactuals in Dynamic Discrete Choice models). In this paper, we present an inference method for a parameter of interest that is robust to the lack of identification of the nuisance parameter, both under the null and under the alternative. Importantly, our method does not require any prior knowledge about the unknown degree of identification of the model, which is unknown but identified from data. Additionally, our proposed method has the secondary benefit of providing an estimation of the degree of identification of the model.
We describe our setting. Let $\theta\in\Theta\subset\mathbb{R}^{m}$ be a reduced-form parameter which is identified from the data distribution $F_{0}$ as $\theta_{0}$. Let $\alpha\in\mathcal{A}\subset\mathbb{R}^{q}$ and $\beta\in\mathcal{B} \subset\mathbb{R}^{p}$ be structural parameters. Structural parameters are related to reduced-form parameters through the mapping $g : \Theta \times \mathcal{A} \times \mathcal{B} \xrightarrow{} \Theta$ such that
This paper proposes a test for the hypothesis
where $\beta_{0}$ is a known fixed value, and we allow for the nuisance structural parameter $\alpha$ to be partially identified under the null, i.e., the set \[ \mathcal{A}_{0}:=\left\{ \alpha\in\mathcal{A}\subset\mathbb{R}^{q}:\theta _{0}=g(\theta_{0},\alpha,\beta_{0})\right\} \] is not necessarily a singleton. We assume that when the null hypothesis is true there exists an $\alpha_{0} \in \mathcal{A}_{0}$, i.e, $\mathcal{A}_{0} $ is non-empty. We also assume there exists an asymptotically normal estimator for $\theta_{0},$ say $\widehat{\theta},$ satisfying \[ \sqrt{n}\left( \widehat{\theta}-\theta_{0}\right) \rightarrow_{d}N\left( 0,\Sigma\right) , \] where $\Sigma$ might be a singular matrix. Consider a standard minimum distance (MD) test statistic
where $\widehat{W}$ is a consistent estimator for a positive semi-definite symmetry non-stochastic matrix $W.$ Assume $g(\theta,\alpha,\beta_{0})$ is twice continuously differentiable at $\theta_{0}\in\Theta\ $and $\alpha_{0} \in\mathcal{A}_{0},$ and define the matrices $\nabla_{\theta}g_{0}:=\partial g(\theta_{0},\alpha_{0},\beta_{0})/\partial\theta^{\prime}$ and $\nabla _{\alpha}g_{0}:=\partial g(\theta_{0},\alpha_{0},\beta_{0})/\partial \alpha^{\prime}.$ Let $r_{\alpha}$ and $r_{\Sigma}$ denote the ranks of $\nabla_{\alpha}g_{0}$ and $\Sigma$, respectively. We assume the rank of the variance and covariance matrix is larger than the rank of the Jacobian, i.e., $r_{\Sigma}>r_{\alpha}$. Under the regularity conditions below, $r_{\alpha}$ does not depend on the point $\alpha _{0}\in\mathcal{A}_{0}$ where it is evaluated. Let $I_{m}$ be the $m$-dimensional identity matrix. The main result of the paper shows that under regularity conditions and $H_{0},$ if $W=\left[ (I_{m}-\nabla_{\theta }g_{0})\Sigma(I_{m}-\nabla_{\theta}g_{0})^{\prime}\right] ^{\dagger}$, where $A^{\dagger}$ denotes the Moore-Penrose generalized inverse of the matrix $A$, then
where $\chi_{d}^{2}$ denotes a chi-squared distribution with $d$ degrees of freedom. This result includes the case where $\alpha_{0}$ is point identified and $r_{\alpha}=q$ (see gourieroux1995testing ). The more general case discussed here allows for lack of identification of $\alpha_{0}$ and $r_{\alpha}<q.$ The convergence in ((ref)) suggests a feasible $\tau\%$ nominal test with critical region $\widehat{F}(\beta_{0})>\chi_{\widehat{r}_{\Sigma}-\widehat{r}_{\alpha},1-\tau}^{2},$ where $\widehat{r}_{\Sigma}-\widehat{r}_{\alpha}$ is a consistent estimator of $r_{\Sigma}-r_{\alpha}$ and $\chi_{d,\tau}^{2}$ denotes the $\tau -$quantile of $\chi_{d}^{2}.$ The proposed method applies general results on the second-order differentiability properties of the mapping \[ v(\theta):=\min_{\alpha\in\mathcal{A}}\left( \theta-g(\theta,\alpha,\beta _{0})\right) ^{\prime}W\left( \theta-g(\theta,\alpha,\beta_{0})\right) , \] at $\theta=\theta_{0},$ see shapiro1986asymptotic. When $\alpha_{0}$ is a regular point (see definition below) and other regularity conditions hold, $v(\theta)$ is twice continuously differentiable and it can be locally approximated by the projection norm onto a tangent linear space of dimension $r_{\alpha}.$ That is, the $q-$dimensional manifold \[ \Theta_{0}:=\left\{ \theta\in\Theta\subset\mathbb{R}^{m}:\theta=g(\theta ,\alpha,\beta_{0}), \hspace{0.25cm} \textit{for some} \hspace{0.25cm}\alpha\in\mathcal{A}\subset\mathbb{R}^{q}\right\} \] has a tangent space of dimension $r_{\alpha}$ at $\theta_{0},$ which is generated by the columns of $\nabla_{\alpha}g_{0}.$ This result does not require identification of $\alpha_{0}$ or full column rank of $\nabla_{\alpha}g_{0}.$ The asymptotic distribution of $\widehat{F}(\beta_{0})$ then follows from a standard application of the Delta-Method, after noticing that $\widehat {F}(\beta_{0})=v(\widehat{\theta})+o_{P}(1)$.
One important novelty of our method is that it does not require any prior knowledge about the degree of identification of the model, i.e, $r_{\alpha}$, as long as $r_{\Sigma}>r_{\alpha}$. Such lack of knowledge about the degree of identification arises whenever the Jacobian matrix depends on the population parameter or there is no analytic solution for it. In SMM, for instance, no analytical solutions of the Jacobian are available, though it can be approximated. In structural models such as Dynamic Games or Dynamic Discrete Choice models, a closed-form solution for $r_{\alpha}$ is available, but it does depend on the population value of $\theta_{0}$, which is unknown. We estimate the rank of $\Sigma$ and $\nabla_{\alpha}g_{0}$ using a hard thresholding approach.
A prototypical application of our results is for testing the validity of calibrated parameters in structural models. This refers to testing for $H_{0}:\beta=\beta_{0}$ for calibrated values $\beta_{0}$ of a structural parameter. “Validity” corresponds to a lack of rejection by our robust test. Another application of our results could be for the construction of robust confidence intervals for a structural parameter of interest by inverting our robust test. We evaluate the finite sample performance of our tests (size and power) in the context of two sets of Monte Carlo experiments, one for a static Bayesian entry game and another for the New-Keynesian model in nakamura2018high. The Monte Carlo results confirm the robustness of our method, relative to traditional non-robust tests based on asymptotic normality under point-identification (t-tests), and show that estimation of the level of under-identification does not have an impact on the performance of the robust test.
We apply our robust method to study both the validity of the calibrated parameters and the inference on the slope of the Phillips curve and the information effect in a simplified version of nakamura2018high. We find their calibration to be valid according to our criteria, and obtain robust confidence intervals for the slope of the Phillips curve and the information effect, which are not only shorter than those obtained by the classical t-test but also shorter than those suggested by the nonparametric bootstrap. The identification-robust inference seems to be well motivated in this application given that the Jacobian for nuisance structural parameters is of reduced rank and the near singularity of the asymptotic variance matrix of reduced form parameters.
The paper is organized as follows: after this introduction and a short literature review, Section (ref) discusses several examples of structural econometric models where our results apply. Section (ref) provides the asymptotic null distribution and power theory for the proposed test. Section (ref) discusses the implementation. Section (ref) investigates the finite sample performance of the proposed methods, while Section (ref) illustrates our method in the application of nakamura2018high. Finally, Section (ref) concludes. The main proofs of Section (ref) are gathered into (ref).
This paper proposes an identification-robust inference method for parameters of interest allowing for nuisance parameters to be unidentified both under the null and under the alternative hypothesis, without prior knowledge of the model's identification degree.
Our problem can be placed in the literature of identification-robust inference of a parameter of interest when the nuisance parameters are not identified under the null. Some examples of this literature are davies1977hypothesis, engle1984wald, andrews1994optimal, hansen1996inference, stinchcombe1998consistent, conniffe2001score. Such literature indexes the test statistic by the nuisance parameter, and then, finds the supremum of all possible test statistics. The distribution of the supremum test statistic is not standard, and they use Monte Carlo or resampling methods to find critical values for the test. Our method leads to critical values that are known up to the degrees of freedom, which can be estimated from data without resampling. Our environment allows the model to be partially identified. Inference on partially identified models has been treated by horowitz1998censoring, horowitz2000nonparametric, chernozhukov2002inference, imbens2004confidence, romano2008inference, beresteanu2008asymptotic, chernozhukov2007estimation, bugni2010bootstrap, bugni2016comparison, and many others. As a common trait, such literature has focused both on constructing confidence intervals that contain the whole identified set and developing confidence intervals for the true value. Most of the methods rely on subsampling or bootstrap to carry out inference.
This paper also relates to the literature on testing with a singular information matrix, such as satorra1989alternative, satorra1992asymptotic, andrews1987asymptotic, ravikumar2000robust or dufour2016rank, among many others. Our paper mainly differs from this literature in the assumption of point identification of the nuisance parameter.
Inference under weak identification is also related to our problem. Some papers in this literature are stock2000gmm,andrews2012estimation, andrews2013maximum,andrews2015maximum, andrews2016conditional, andrews2016geometric, han2019estimation, andrews2019identification; or lee2022robust. kleibergen2005testing assumes the full column rank of the Jacobian for the nuisance parameter. andrews2016geometric gives finite sample bounds on the distribution of minimum statistics used in minimum distance models. They assume the full column rank of the Jacobian. Moreover, in their framework the model $g(\cdot)$ does not depend on the reduced-form parameter $\theta_{0}$. andrews2019identification develop a singular weighting matrix robust test. They also assume identification of the nuisance parameter. andrews2012estimation develops an estimation and inference method that is robust to lack of identification (and thus weak identification) of nuisance parameters. There are two main differences with our paper. First, to use their method, knowledge about which parameters can be identified is required, while we are agnostic about the identification problem. Second, the objective function does not depend on the nuisance parameter when it is not identified, while we allow the objective function to depend on the nuisance parameter even if it is not identified. han2019estimation develops estimation and inference in a GMM framework, with a Jacobian that is not full column rank. han2019estimation works on the andrews2012estimation environment, and therefore, prior knowledge of the identification problem is required. Moreover, the objective function cannot depend on unidentified nuisance parameters.
antoine2022identification develop an identification-robust inference of a structural parameter of interest using a minimum distance estimation method. The main difference with our paper is that we allow for under-identification of the nuisance parameter. antoine2022identification also relies on bootstrap methods to find the critical values of the test. We avoid using bootstrap, making our procedure computationally less expensive for a large variety of complex structural models.
The proposed method is motivated by the application of dynamic games by pesendorfer2008asymptotic. These authors suggested the MD estimator
where $\widehat{\theta}$ is a consistent estimator for the ex-ante choice probabilities $\theta_{0}$ and $g$ is the best response mapping, linking the structural parameter $\alpha_{0}$ to the choice probability $\theta_{0}$. In equilibrium $\theta_{0}=g(\theta_{0},\alpha_{0}).$ Much of the literature assumes $\alpha_{0}$ is uniquely identified from the equilibrium conditions. However, there is theoretical evidence that these models are not identified. Non-identification results have been shown in the seminal paper by rust1994structural and magnac2002identifying for single-agent models and by pesendorfer2008asymptotic for multiple-agent models. Here, we present a methodology that can be applied to partially identified models, and therefore, has wider applicability than existing procedures. pesendorfer2008asymptotic show how many popular estimators fall under the class ((ref)). Specifically, they show that the moment estimator of hotz1993conditional or the pseudo-maximum likelihood estimator of aguirregabiria2002identification are asymptotically equivalent to estimators in the class ((ref)) for specific choices of the plim of $\hat{W}$, $W.$ Here, we allow these choices as special cases, while permitting the structural parameters to be under-identified.
In this application $\theta_{0}$ is a vector of impulse response functions obtained from a Structural Vector Autoregression model, and $g(\cdot)$ is the link between structural parameters and reduced-form parameters. MD estimation in this context has been proposed by rotemberg1997optimization, amato2003estimation, and christiano2005nominal , among many others. In this application, $g$ does not depend on $\theta_{0}$.
Simulated-based methods such as SMM and II also fit our environment, after an appropriate modification that we give below in Section (ref). Moreover, due to the nature of their problem, the degree of identification of the model is unknown. Furthermore, simulated-based methods are computationally demanding, making bootstrap and subsampling inference not feasible computationally. Therefore, our method addresses two main problems in this literature. It is agnostic about the degree of identification, while it is also computationally feasible.
This method is introduced in the seminal papers by mcfadden1989method, pakes1989simulation, duffie1990simulated, and lee1991simulation.
Assume we observe a sample of random variables $X_{i}$ of size $n$, where $X_{i}$ follows a distribution $ F(\cdot)$. Moreover, suppose we have a structural economic model, indexed by parameters $(\alpha,\beta)$ that generates a simulated random variable $\tilde{X}_{j}(\alpha,\beta) \sim F_{\alpha,\beta}$. For a given measurable function:
where $\mathcal{X}$ is the domain of $X_{i}$, let
The reduced-form parameter $\theta_{0}$ is estimated using the sample analog in ((ref)). In SMM, the moment $g(\cdot)$ is defined by
which is not analytically known. We use Monte Carlo, and compute
where $\{x_{j}(\alpha,\beta)\}$ are $B$ random draws from $F_{\alpha,\beta}$. The SMM minimizes a quadratic objective function that depends on the distance between the estimated reduced-form parameter ((ref)) and the simulated moment condition ((ref)), i.e.,
where $\widehat{W}$ will be the estimated optimal weighting matrix of this minimum distance problem. By choosing $B$ significantly large, one shows that the numerical approximation of ((ref)) has no impact on the asymptotic distribution of ((ref)),
and hence, SMM fits our framework with a $g(\cdot)$ given by ((ref)) that does not depend on $\theta$.
This section investigates the asymptotic null distribution of $\widehat {F}(\beta_{0}).$ The first condition requires the asymptotic normality of the reduced-form estimator. The test statistic was defined in equation ((ref)) as
These assumptions are standard in the literature on MD inference. We say the point $\alpha_{0}\in\mathcal{A}$ is locally regular if belongs to the interior of $\mathcal{A}$ and the Jacobian matrix $\nabla_{\alpha}g(\alpha):=\partial g(\theta_{0},\alpha,\beta_{0})/\partial\alpha^{\prime}$ has the same rank as $\nabla_{\alpha}g_{0}\equiv\nabla_{\alpha}g(\alpha_{0})$, say $r_{\alpha},$ for every $\alpha$ in a neighborhood of $\alpha_{0}.$ We say the point $\alpha_{0} \in\mathcal{A}$ is regular if it is locally regular and there exist neighborhoods $\mathcal{U}$ and $\mathcal{V}$ of $\alpha_{0}$ and $\theta_{0},$ respectively, such that $\Theta_{0}\cap\mathcal{V}=g(\theta_{0},\mathcal{U} ,\beta_{0}),$ where $\Theta_{0} :=\left\{ \theta\in\Theta\subset\mathbb{R}^{m}:\theta=g(\theta ,\alpha,\beta_{0}), \hspace{0.25cm} \textit{for some} \hspace{0.25cm}\alpha\in\mathcal{A}\subset\mathbb{R}^{q}\right\} $.
Assumption (ref)$(i)$ is standard and it can be relaxed at the cost of introducing further assumptions on $g.$ A subset $S$ of a topological space $\mathcal{X}$ is said to be connected whenever $S$ cannot be expressed as the disjoint union of two non-empty open subsets of $\mathcal{X}$. When Conditions (ref)$(i)$, (ref)$(ii)$ and (ref)$(iii)$ do not hold, our test controls the size, although it might become conservative, see Remark (ref). Assumptions (ref) and (ref) are standard in the literature. Assumption (ref) is a local identification condition for reduced form parameters.
If $W=\left((I_{m}-\nabla_{\theta}g_{0})\Sigma (I_{m}-\nabla_{\theta}g_{0})^{\prime}\right)^{\dagger}$, then the limiting distribution depends on the data-generating process only through the rank of the matrix $\nabla_{\alpha}g_{0}.$ In this case, it is natural to consider the sample analog \[ \widehat{r}_{\alpha}=rank\left[ \nabla_{\alpha}g(\widehat{\theta},\widehat{\alpha },\beta_{0})\right] , \] where $\widehat{\alpha}$ solves the optimization problem
and $\lambda_{n}$ is a penalization parameter such that $\lambda_{n}\downarrow0$.
Furthermore, to satisfy Assumption (ref) whenever $W=\left((I_{m}-\nabla_{\theta}g_{0})\Sigma (I_{m}-\nabla_{\theta}g_{0})^{\prime}\right)^{\dagger}$, an estimator based on the analog principle of $W$ might be inconsistent. This is because the Moore-Penrose g-inverse is not continuous. stewart1969continuity gives necessary and sufficient conditions to estimate consistently a Moore-Penrose g-inverse. A detailed discussion on how to construct a consistent estimator of $W$ is provided in Section (ref).
Henceforth, to save notation we denote $d \equiv r_{\Sigma}-r_{\alpha}$ and $\widehat{d} \equiv \widehat{r}_{\Sigma}-\widehat{r}_{\alpha}$. Consider the oracle test $\tilde{\phi}_{\tau}=1(\widehat{F}(\beta_{0} )>\chi_{d,1-\tau}^{2}),$ the one that uses the true degree's of freedom $d$, where $1(A)$ denotes the indicator function of the event $A$ (1 if $A$ holds, and 0 otherwise). Define the asymptotic power function $\pi_{\tau}(\beta)\equiv \lim_{n\rightarrow\infty }\mathbb{P}_{\beta}\left( \tilde{\phi}_{\tau}=1\right) $, where $\mathbb{P} _{\beta}$ denotes the underlying probability when the true parameter is $\beta.$ Call feasible test $\hat{\phi}_{\tau} = 1(\widehat{F}(\beta_{0} )>\chi_{\widehat{d}}^{2})$, i.e, the test using the estimated degree of freedom $\widehat{d}$.
The next theorem justifies the proposed feasible test. Define the set \[ \mathcal{B}_{1}:=\left\{ \beta_{1}\in\mathcal{B}\subset\mathbb{R}^{p} :\min_{\alpha\in\mathcal{A}}\left( g(\theta_{0},\alpha_{1},\beta_{1})-g(\theta _{0},\alpha,\beta_{0})\right)^{\prime}W\left( g(\theta_{0},\alpha_{1},\beta_{1})-g(\theta _{0},\alpha,\beta_{0})\right) >0\right\}. \] where $\alpha_{1} \equiv \alpha_{1}(\beta_{1})$ is the minimum norm solution of the problem
Theorem (ref) shows that the test have power if differences between $g(\theta_{1},\alpha_{1},\beta_{1})$ and $g(\theta_{0},\alpha_{0},\beta_{0})$ can be detected when using $||x||_{W} = \left(x'Wx\right)^{\frac{1} {2}}$ as seminorm, i.e, $||g(\theta_{1},\alpha_{1},\beta_{1})-g(\theta_{0},\alpha_{0},\beta_{0})||_{W} > 0$.
The next results give sufficient conditions to have local power under Pittman alternatives. Let $\beta_{1n} = \beta_{0} + \frac{\delta}{\sqrt{n}}$, and with some abuse of notation denote the asymptotic local power function by $\pi_{\tau}(\delta)\equiv \lim_{n\rightarrow\infty }\mathbb{P}_{\beta_{1n}}\left( \tilde{\phi}_{\tau}=1\right) $. We call $\delta$ the "direction" of the alternative.
Define the Jacobian of the parameter of interest as $\nabla _{\beta}g_{0}:=\partial g(\theta_{0},\alpha_{0},\beta_{0})/\partial \beta^{\prime}. $
This section will illustrate the test algorithm, with additional details provided in the following subsections.
In this subsection, we motivate the simple hard-thresholding method we use to estimate the rank of a matrix. Take $\widehat{M}$ and $M$ as general positive semi-definite symmetric matrices of dimension $m\times m$, with $\widehat{M} \xrightarrow{p} M$, and $r_{M} = Rank(M)$. The estimator for the rank of $M$ is:
where $\widehat{\lambda}_{i}$ is the $i$th eigenvalue of $\widehat{M}$ and $0<b<1$ is a tuning parameter. Previous literature such as lutkepohl1997modified or dufour2016rank has used $b= \frac{1}{3}$, but we instead use $b=0.99$. Our choice is motivated by the better performance of our test in the Monte Carlo exercise. We have tuned other values of $b>\frac{1}{2}$ and obtained similar results. Future research will investigate cross validated methods for the choice of $b$.
with $s=Rank(\Omega) \leq m^{2}.$
Instead of using spectral cut-off methods, a common alternative is to use a sequential testing procedure to estimate the rank. Statistical tests, such as those proposed by robin2000tests or kleibergen2005testing can be employed for this purpose. However, this approach requires a consistent estimator of the asymptotic variance matrix, denoted as $\Omega$, of $\sqrt{n}\left(Vec(\widehat{M})-Vec(M) \right)$. Obtaining such an estimator, $\hat{\Omega}$, might be challenging in our context, particularly when a closed-form solution for the matrix $M$ is unavailable. On the other hand, spectral cut-off methods, such as those used by lutkepohl1997modified, dufour2016rank or lee2022robust among others, do not require estimating $\Omega$. Our approach differs from theirs in that we exploit the fact that the estimate $\widehat{\lambda}_{i}$ converges faster than $\sqrt{n}$ to the population eigenvalue $\lambda_{i}$ whenever $\lambda_{i} = 0$.\\ The following Theorem is taken from robin2000tests, see the Appendix for a proof.
To sum up, we illustrate the algorithm for the estimation of the rank of both $\Sigma$ and $\nabla_{\alpha}g_{0}$.
In this subsection, we discuss how to implement a consistent estimate of the optimal weighting matrix
The main problem in consistently estimating $W$ is that Moore-Penrose g-inverses are not continuous. An estimator based on the analog principle of $W$ might be inconsistent. Define $$A \equiv (I_{m}-\nabla_{\theta}g_{0})\Sigma (I_{m}-\nabla_{\theta}g_{0})^{\prime},$$ $$\hat{A} \equiv (I_{m}-\hat{\nabla}_{\theta}g)\widehat{\Sigma} (I_{m}-\hat{\nabla}_{\theta}g)^{\prime}.$$
stewart1969continuity gives necessary and sufficient conditions to estimate consistently a Moore-Penrose g-inverse. Such conditions are: if $\hat{A} \xrightarrow{p} A$, then
To ensure that such a condition holds, we truncate the eigenvalues of $\widehat{A}$. First, use the spectral decomposition
Let $\hat{p}_{j}$ be the diagonal elements of $\hat{P}$, i.e, the eigenvalues. Then truncate the eigenvalues that are smaller than $\tfrac{1}{n^{b}}$, $\tilde{p}_{j} = \hat{p}_{j}\mathbbm{1}{ \{\hat{p}_{j}>\frac{1}{n^{b}}\}}$. Define $\Tilde{P}$ as the diagonal matrix composed by the $jth$ different elements $\tilde{p}_{j}$, $\Tilde{P}^{\dagger}$ as the diagonal matrix with the $jth$ element given by $1/\tilde{p}_{j}$ when $\tilde{p}_{j}\neq0$, and zero otherwise, and compute
Then, set
Following the results of Theorem (ref) $Pr\left(Rank(\tilde{A}) = Rank(A)\right)\xrightarrow{p}1$. Hence, the estimator $\widehat{W}$ satisfies stewart1969continuity necessary and sufficient conditions for
In the following subsections, we examine the finite sample properties of our method using Monte Carlo in two different applications. First, in a Static Bayesian Game proposed by pesendorfer2008asymptotic. Second, in a setting that resembles the New-Keynesian model proposed by nakamura2018high.
In this section, we demonstrate our methodology using an example from pesendorfer2008asymptotic. We examine the properties of our test in finite samples, including size and power. The example we consider is a static Bayesian Game involving two firms that must decide whether or not to enter a market simultaneously. The payoffs for Firm 1 and Firm 2 are as follows: $$\pi_{1}\left(a_{1}, a_{2}, s\right)=
$$
$$\pi_{2}\left(a_{1}, a_{2}, s\right)=
$$ The variable $s \in \{s_{1},s_{2},s_{3}\}$ is the state of the economy, with $s_{1}<s_{2}<s_{3}$. The utility for a firm $i$ with state $s$ and other firm's action $j$, is defined by:
where $\epsilon_{i}$ is a random unobserved variable distributed as a standard normal. Firm $i$th decides to be active whenever
Firm $i$ has expectations about the probability of firm $j$ being active. In equilibrium, such expectations will be equal to the actual probability of firm $j$ choosing to be active, i.e
and
where $\theta_{j,s}$ is the probability of choosing to be active for a firm $j$ given state $s$. The mapping $g$ is the CDF of a standard normal distribution. Our parameter of interest will be $\beta_{0}$, i.e, the coefficient of firm 1's payoff associated with being active when the other is not. The remaining parameters are considered nuisance parameters.
In this example, we have six moment conditions (three states for every single firm) and four parameters, three of which are nuisance parameters. The Jacobian of the nuisance parameter is
Where the density of the standard normal, denoted by $\dot{g}_{i,s}$, is evaluated at $ (1-\theta_{i,s}) \cdot s\cdot \beta + \theta_{i,s}\cdot\alpha_{1} \cdot s$ for $s \in \{s_{1},s_{2},s_{3}\}$ and firm $i$. It is worth noting that when the choice probabilities of firm one are one-half for all possible states of the economy, i.e., $\theta_{1,s_{1}}=\theta_{1s_{2}}=\theta_{1,s_{3}}=\frac{1}{2}$, the Jacobian's rank becomes deficient.
This simple example illustrates how the degree of identification might depend on population values of the reduced form parameter, which are unknown.
We test the hypothesis:
The value of $\beta_{0}$ is set to $\beta_{0} = 1.5$ in the identified case (case 1), and $\beta_{0} = 0.3$ in the not identified case (case 2). We compare three procedures: an Oracle test (Oracle) that assumes knowledge of the degrees of freedom $d$, our robust test (Robust) that estimates $d$, and the t-test (T-test) based on the asymptotic normality of the estimate of the structural parameter $\beta_{0}$ under point-identification. A summary of the performance of the different tests in terms of size for the identified case is presented in Table (ref). The findings indicate that rank estimation has a minimal impact on the procedure's overall level of uncertainty.
Table (ref) considers the unidentified case, the T-test presents large size distortions with a larger size than $5\%$. In contrast, the Robust and Oracle tests control the size. Again, the estimation of the rank does not affect the performance of the Robust test.
In the following figures, we report the power of the different tests for the $\textit{Case 2}$ of not identification under the null. In this example, the parameter of interest $\beta$ is identified under any possible alternative.
As a consequence of $\beta$ being identified under the alternative, the t-test has non-trivial power. Yet, the size is not controlled. The Robust and the Oracle tests yield similar results in terms of power, suggesting that the estimation of the degrees of freedom does not have a big impact in finite samples.
In this section, we carry out a Monte Carlo simulation based on nakamura2018high. Our analysis consists of two parts. Firstly, we assess the finite sample properties by testing the validity of the calibrated (fixed) parameters in terms of both the size and power of our method. Secondly, we evaluate the finite sample properties of conducting inference on two structural parameters of interest, namely the Phillips Curve slope and the Information Effect. For these structural parameters, we compare the size and power obtained from our robust inference method with those from standard inference methods.
nakamura2018high provides compelling evidence of the non-neutrality of monetary policy. They accomplish this by identifying the information effect of Federal Reserve announcements on economic fundamental beliefs. Specifically, nakamura2018high utilizes data on unexpected changes in interest rates occurring within 30 minutes of a Fed announcement. In addition, they construct and estimate a structural model using the simulated method of moments. Their structural model involves the estimation of $m=33$ reduced-form parameters and $q=5$ structural parameters. The mapping function, $g(\cdot)$, has no closed-form solution, and it is therefore estimated numerically as in Section (ref).
Given that nakamura2018high uses two different databases to estimate the reduced-form parameters, the matrix of variance and covariance $\Sigma$ is not identified.
To illustrate our methodology, we employ a subset of moments and parameters, estimating $m=8$ reduced-form parameters and $q=2$ structural parameters. The latter corresponds to the slope of the Phillips Curve and the information effect of monetary policy. On the other hand, nakamura2018high calibrate eight parameters in their study, while we calibrate $k=11$ parameters. Notably, the reduced-form parameters are estimated from a single database, enabling the identification of the variance and covariance matrix of the estimated parameters.
As demonstrated in Section (ref), the estimated rank of the Jacobian matrix is deficient, and the asymptotic variance matrix of the reduced-form parameters, $\widehat{\Sigma}$, is close to being singular. These limitations motivate the use of robust inference techniques. It is worth noting that throughout the Monte Carlo exercise, we employ the estimated $\hat{\Sigma}$ from Section (ref) as the true one. Table (ref) presents the structural parameter values for different data-generating processes.
We refer to calibration validity as not rejecting the calibrated value with our robust test. In other words, we test if $\beta=\beta_{0}$, where $\beta_{0}$ are the calibrated values.
In this subsection, we test the validity of the calibration of the modified version of nakamura2018high explained above. The vector $\beta$ is composed by all calibrated (fixed) parameters $$\beta = \left( \rho, \eta, \omega, \gamma, \phi, \pi, \delta, \sigma, \rho_{1}, \rho_{2}, b \right)'.$$
The nuisance parameters are $\alpha = \left(\kappa \zeta, \psi \right)'$. There is no closed form solution for the mapping $g: \mathcal{B}\times\mathcal{A} \xrightarrow{} \Theta$, with $\mathcal{B} \subset \mathbb{R}^{11}$ and $\mathcal{A} \subset \mathbb{R}^{2}$. We use numerical methods to estimate $g(\cdot)$. We test the null hypothesis:
where $\beta_{0}$ are the fixed parameter values shown in Table (ref).
In the replication of the application, the model is identified, albeit weakly. This is the reason why, in Table (ref), it can be observed that using the oracle test (which utilizes the true degree of freedom) yields a size different from 5%. The empirical size of the Robust test is satisfactory. To illustrate the power of the test, we maintain all elements of $\beta$ constant except for the shock to the inflation target, denoted as $\pi$. We focus on this alternative hypothesis since $\pi$ contributes the most to the deviation with maximum local power, as can be seen in Table (ref) in Section (ref).
In this subsection, we assume our previous calibration is valid. Moreover, the vector of structural parameters of interest is $\beta = \kappa \zeta$, and the nuisance parameter is $\alpha = \psi$. We test the hypothesis:
where $\kappa\zeta_{0}$ are the structural values presented in Table (ref) that correspond to every different case of study.
Under the Null hypothesis, both the t-test and the Oracle test do not exhibit a significance level of 5%. This deviation from the expected significance level can be attributed to the weak identification of the model and the near singularity of $\Sigma$ (covariance matrix). In contrast, the robust test has a satisfactory size performance. Figure (ref) reports the results on power. The t-test has low power, while that of the Robust and Oracle tests is high for deviations from the right of the null hypothesis. In contrast to the Static Bayesian Game case, the model is unidentified under any alternative hypothesis. Furthermore, the covariance matrix $\Sigma$ is nearly singular. Consequently, the t-test fails to control the significance level. The Robust test has non-trivial power for any alternative hypothesis. Nevertheless, the power function is not symmetric around the null. Specifically, smaller alternative hypotheses relative to the null hypothesis yield lower power compared to larger alternatives. This would suggest that the derivative of the mapping $g$ with respect to $\kappa \zeta$ is getting closer to 0 as $\kappa \zeta$ goes to 0.
In this subsection, the vector of structural parameters of interest is $\beta = \psi $, and the nuisance parameter is $\alpha = \kappa\zeta$. Whereas the rest of the structural parameters are assumed to be valid. We test the hypothesis:
where $\psi_{0}$ are the structural values presented in Table (ref) that correspond to every different case of study. In the following Table (ref) and Figure (ref) we report the results for this case. A remarkable difference relative to the previous section is that the Robust test has trivial power against any alternative, due to the fact that there exist observational equivalent models for any alternative $\beta_{1} \in \mathcal{B}$, i.e, $\forall$ $\beta_{1} \in \mathcal{B}$ $\beta_{1} \notin \mathcal{B}_{1}$, check Theorem (ref) for further details.
In this particular application, it seems that the Information Effect is not identified. This lack of identification could explain why there is trivial power in the Robust Test.
In this section, we employ our robust inference method to re-analyze the study by nakamura2018high and compare our results to those obtained through standard methods. In particular, first, we check the validity of the calibrated parameters in the main paper using our methodology. Second, we compare the inferences obtained by inverting our test with those obtained by standard inference methods.
To evaluate the robustness of their approach, we estimate the rank of the Jacobian of the model evaluated at the estimated parameters, obtaining $\hat{r}_{\alpha}=Rank(\nabla_{\alpha}g(\widehat{\alpha}))=4<q=5$, indicating potential identification issues.
Unfortunately, in nakamura2018high the asymptotic variance matrix of reduced form estimators is not identifiable due to the use of at least two different databases. To circumvent this issue, we opted to employ a subset of the estimated reduced form parameters, ensuring the identification of $\Sigma$. Our estimation procedure involves computing $m=8$ reduced-form parameters and $q=2$ structural parameters, the latter of which represent the slope of the Phillips Curve and the information effect of monetary policy. However, even though $\Sigma$ is identified in this simplified version of the problem, it is nearly singular. Table (ref) shows that the estimated rank of this matrix is $Rank(\hat{\Sigma})=5$. As a result, we advocate for identification-robust and singular-weighting-matrix-robust inference techniques in this application. The null hypothesis is
where $\beta_{0} = \left( \rho, \eta, \omega, \gamma, \phi, \pi, \delta, \sigma, \rho_{1}, \rho_{2}, b \right)'$ are the values fixed in Table (ref).
Table (ref) shows that the null hypothesis is not rejected by our robust test. Indicating that calibrated values are valid. Moreover, it is worth noting that the model is not identified, as the estimated value of $\widehat{r}_{\alpha} = Rank(\nabla_{\alpha}g(\widehat{\alpha}))$ is only one, which falls short of the required value of $q=2$. We conducted an empirical analysis of the local power and determined that the dimension of the local nontrivial power space, denoted by $\mathcal{B}_{\tau}$, is eight. This means that out of 11 parameters, only deviations in three parameters can lead to the rejection of the test with local nontrivial power.
Table (ref) illustrates the individual contributions of each calibrated parameter towards the direction of maximum local power, as discussed in Remark (ref). The top three contributors to the maximum local power are the inflation target shock, the endogenous feedback in the Taylor rule, and the first autoregressive root of the monetary shock.
Table (ref) presents a summary of the estimation and inference results for the structural parameters. In comparison, the confidence intervals (CI) based on the t-test are wider than those obtained using the bootstrap or the robust test. Moreover, the CI using the robust test appears to be shorter. CI bootstrap is the method that nakamura2018high uses, and it is obtained using non-parametric bootstrap to estimate the finite sample distribution of the estimates. Notably, the estimated information effect is approximately zero, indicating that nothing of the real effects of monetary policy stem from revisions in agents' fundamental economic beliefs. However, the estimation's confidence intervals are wide, primarily due to the exclusive use of inflation-related reduced-form parameters.
In this paper, we propose inference for structural parameters based on minimum distance that is robust to identification problems of nuisance parameters. Moreover, to implement this method is not necessary to have prior knowledge about the identification degree of the model. Last but not least, this inference method is computationally fast. Our robust inference methods can be applied to test the validity of calibrated (fixed) parameters in structural models under weak identification of nuisance parameters. They can be also applied to obtain identification-robust confidence intervals by inverting our robust test. Therefore, this method is suitable for a large class of applications in structural models commonly employed in Macroeconomics and Microeconomics.