EconBase
← Back to paper

Inference on the New Keynesian Phillips Curve with Very Many Instrumental Variables

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

62,326 characters · 10 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference on the New Keynesian Phillips Curve with Very Many Instrumental Variables

abstractLimited-information inference on New Keynesian Phillips Curves (NKPCs) and other single-equation macroeconomic relations is characterised by weak and high-dimensional instrumental variables (IVs). Beyond the efficiency concerns previously raised in the literature, I show by simulation that ad-hoc selection procedures can lead to substantial biases in post-selection inference. I propose a Sup Score test that remains valid under dependent data, arbitrarily weak identification, and a number of IVs that increases exponentially with the sample size. Conducting inference on a standard NKPC with 359 IVs and 179 observations, I find substantially wider confidence sets than those commonly found.\\

Introduction

Instrumental variable (IV) methods are often used to conduct limited-information inference on (structural) single-equation macroeconomic relations that describe the dependence of a scalar variable on a set of covariates. Examples of such macroeconomic relations include New Keynesian Phillips Curves (NKPCs), Euler equations, and Taylor rules. IV-based limited-information inference on such macroeconomic relations has arguably proven popular because there is no requirement that parts of the model other than the specified relation itself be necessarily true to conduct valid inference. In virtually all applications, the relation is assumed to contain an additive error term that is shown (e.g., by the assumption of Rational Expectations (RE)) or primitively assumed to be uncorrelated with predetermined variables excluded from the specified relation. This makes any predetermined variable a valid IV.

As documented extensively in the existing literature, using IVs to conduct limited-information inference on such macroeconomic relations often runs into issues related to weak identification. This occurs when the variation in the IVs is only able to explain a small portion of the variation of the endogenous variables.\footnote{For the case of NKPCs, see Kapetanios:2015fp, Mirza:2014wd, Mavroeidis:2014ge, Kleibergen:2009do, Dufour:2006bh, Ma:2002wl, for the case of Euler equations see Ascari:2019ch, Kleibergen:2005ug, Yogo:2004vm, Stock:2000vt, for the case of Taylor rules see Mirza:2014wd, Mavroeidis:2010bv.} This problem is especially pronounced when the analysis is restricted to using only a few variables to forecast the endogenous variables, a restriction that arises when using IV methods that treat the number of IVs as fixed relative to the sample size. Since any predetermined variable is a valid (if not very informative) IV, this naturally raises the question of which IVs to choose out of the very many available ones.

The limited literature that seeks to formally address the high dimensionality of the available IVs in such macroeconomic settings is primarily motivated by the potential inefficiency of using IVs selected in an ad-hoc way Berriel:2019vp, Bayar:2018er, Berriel:2016hj, Mirza:2014wd, Kapetanios:2015fp. Through simulations and/or empirical applications, these studies find smaller confidence sets than the ones implied by IVs traditionally used in the past. Although this evidence is certainly suggestive, it should be noted that formal efficiency claims rely on conditions that are not easily verifiable in practice.\footnote{For instance, factor-based approaches to reduce the dimensionality of the IVs likely work well if there is a factor structure, and if whatever explains most of the variation in the IVs, also explains (a good portion of) the variation of the endogenous variables. While the former may be made plausible through certain tests, the latter remains an assumption the researcher has to make. Similarly, a LASSO-based selection of IVs works well only under the assumption that the relation between the endogenous variables and the candidate IVs is sufficiently sparse.}

Rather than being motivated by such efficiency concerns, this paper revisits the question of high-dimensional limited-information inference because some types of formal or intuitive regularisation can lead to invalid inference, even if weak-IV robust methods are used after regularisation. This is due to what Chernozhukov:2015iq call the `endogeneity bias', which arises when variables are selected on the basis of their in-sample correlation with a model's error terms.

The first contribution of this paper consists in illustrating how improper selection of IVs can lead to invalid inference in the context of limited-information inference on a standard NKPC. I do this by extending the simulations in Mavroeidis:2014ge to the more realistic case where the econometrician does not have oracle knowledge on which IVs are the relevant ones, but rather has to choose amongst the very many available IVs. I consider different IV selection techniques, and show that several of them result in substantially invalid inference. The example of the NKPC is chosen for the sake of concreteness and due to its popularity in the literature. The same concerns extend to any of the many cases in Macroeconomics where a given structural equation can be estimated with very many valid IVs.

As a second contribution, I propose a Sup Score test to conduct IV-based limited-information inference on single-equation Macroeconomic relations. Contrarily to other approaches in the literature, this statistic requires no assumption on the factor structure of the IVs, nor does it make any sparsity-type assumption that requires only a few of the very many IVs to be relevant, while allowing for a number of IVs that increases exponentially with the sample size. This test directly contributes to the (very) many weak IVs literature predominantly restricted to the cross-sectional case (see Mikusheva:2020uo, Crudu:2020bk, Belloni:2012kw, Anatolyev:2010kk), and can find application well beyond the example of NKPCs considered in this paper.

The third contribution consists in applying the selection procedures considered in the simulation section and the Sup Score test to conduct IV-based limited-information inference on a standard hybrid NKPC with 359 IVs on a sample of 179 observations. I find that both the IVs selected and the confidence set implied by the selection procedure that yields the worst size distortion in the simulations are similar to the IVs selected and the confidence set implied by the IVs traditionally used in the past. This suggests that the results previously reported in the literature may suffer from endogeneity bias, and that they hence may undercover the true parameter values. By contrast, the confidence sets implied by the Sup Score test are considerably wider.

Notation. For any real number $a$, $\floor*{a}$ indicates the smallest integer $b$ such that $b \leq a$. For any two real numbers $c$ and $d$, $c \lesssim d$ if $c$ is smaller than or equal to $d$ up to a universal positive constant. The remaining notation follows standard conventions.

Organisation of the paper. Section (ref) introduces the model considered in this paper. Section (ref) outlines the methods used in this paper to conduct inference in the context of very many IVs. Section (ref) provides simulation-based evidence on the size and power of these methods. Section (ref) revisits inference on the US NKPC using very many IVs. Section (ref) concludes.

Model

The structural equation I consider is the hybrid NKPC of Gali:1999tx,

equation[equation omitted — 141 chars of source]

where $\pi_t$ is the inflation rate, $s_t$ is the forcing variable, and $\lambda$, $c$, $\gamma_{f}$, and $\gamma_b$ are parameters of the model. $u_t$ is an unobserved disturbance term, which can be interpreted as a measurement error, or as a shock to inflation, such as a cost-push shock.

The identifying moment conditions can be derived within the framework of Generalised Instrumental Variable (GIV) estimation. In this approach, realised one-period-ahead inflation is substituted in for expected inflation. This means that Equation (ref) can be re-written as

equation*[equation* omitted — 215 chars of source]

If it is further assumed that $\mathbb{E}_{t-1}[u_t] = 0$, the assumption of RE gives rise to the moment conditions

equation*[equation* omitted — 79 chars of source]

for any $k\times 1$ vector of predetermined variables ${Z}_t$. Due to the very large number of predetermined time series available, the dimension of $Z_t$ is comparable to or larger than the number of observations, $T$.

It should be noted that the example of NKPCs (including the particular specification chosen), and the assumption of RE are not central to two of the contributions of this paper. The same concerns relating to the endogeneity bias persist, and the same Sup Score test proposed below remains valid for the broad class of models defined by single-equation relations of the type

equation[equation omitted — 90 chars of source]

and moment equations given by

equation[equation omitted — 84 chars of source]

where $ y$ is a $T\times 1$ vector, $g$ is a known real-valued function, $ Y$ is a $T\times p_1$ matrix of endogenous covariates, $ X$ is a $T\times p_2$ matrix of exogenous covariates, $Z$ is a $T\times k$ matrix of variables such that $k \geq p_1 + p_2$, $\theta$ is a $(p_1 + p_2)\times 1$ vector of coefficients, $p_1$ and $p_2$ are both fixed, and $ \varepsilon$ is a $T\times 1$ vector of error terms. This setup encompasses many popular applications in Macroeconomics, where $k$ is of the same magnitude or even larger than $T$, such as limited-information inference on NKPCs, Euler equations, and Taylor rules.

In particular, the NKPC considered in Equation (ref) can be mapped into the more general model in Equation (ref) as follows. Since the NKPC is linear, the exogenous (predetermined) variables can be partialled out. Hence, $y = M_X\pi$, $Y = M_X[s \text{ } \pi_{+1}]$, $M_X = I - X(X'X)^{-1}X'$, $ X = [{1}_{T\times 1} \text{ } \pi_{-1}]$, $g( Y, X, \theta) = Y\theta$, $\theta = [\lambda, \gamma_f]'$, $\varepsilon = M_X\epsilon$, $Z = M_X\tilde{Z}$, $\tilde{Z}$ is a $T\times (k-2)$ matrix of excluded IVs, $s, \pi_{+1}$, $\pi_{-1}$, and $\epsilon$ are the $T\times 1$ stacked vectors of $s_t$, $\pi_{t+1}$, $\pi_{t-1}$, and $\epsilon_t$, respectively.

Methodology

For all methods considered in this paper, confidence sets are constructed by inverting statistics that test the hypothesis

equation[equation omitted — 117 chars of source]

The $(1-\alpha)$ confidence set can be constructed by collecting the values of $\theta_0$ for which the null hypothesis in Equation (ref) is not rejected at the $\alpha$ level of significance. For convenience, define ${\varepsilon}_0 \equiv y - g( Y, X, \theta_0)$.

Post-Selection Low-Dimensional Inference

Most of the existing literature that conducts inference on relations of the form presented in Equation (ref) using moment conditions of the type shown in Equation (ref) has employed methods that require the IVs to be low-dimensional. In the presence of very many IVs, these approaches can be seen as a two-step procedure. First, the IVs are selected. Second, a low-dimensional (weak-identification robust) method is applied with the selected IVs. The first step is usually not made explicit, and is often not given any attention, which makes it impossible to model this step accurately. In Section (ref), I consider three different selection procedures that reasonably cover (in terms of their deleterious effect on subsequent inference) the range of selection procedures used in the previous literature. These are random selection, `crude thresholding', and LASSO. In Section (ref), I outline the $S$ statistic of Stock:2000vt, which forms the post-selection inferential method common to all three selection procedures considered in this paper.

Before proceeding, it is helpful to gain some intuition as to why IV selection may lead to invalid IVs. For simplicity, suppose that all variables are endogenous (or that the model is linear and that the exogenous covariates have been partialled out). Consider the following projection (`first stage')

equation*[equation* omitted — 33 chars of source]

where $\zeta$ is a $k\times 1$ vector of coefficients and $v$ is a $T\times 1$ vector of error terms. Consider the case of no identification at all, $\zeta = 0$, and a selection procedure that selects the IVs that are most highly correlated with the endogenous variables, $Y$. This amounts to selecting those IVs that are most highly correlated in-sample with the first-stage error term. By the endogeneity of the system, this means that those IVs most correlated with the error term, $\varepsilon$, will be selected, so that conditional on selection, the IVs are no longer valid. This phenomenon carries over more broadly to cases of weak (but non-zero) identification as discussed in Hansen:2014ie.

The Stock:2000vt $S$ Statistic

In this paper, the GMM-based $S$ statistic of Stock:2000vt will be used for low-dimensional post-selection inference.\footnote{More powerful and computationally intensive (GMM-based) weak-identification robust methods could be used instead of the $S$ statistic (see Mirza:2014wd, Kleibergen:2009do). Considering them instead of the $S$ statistic does not qualitatively affect the results of the simulations, while increasing their computational burden substantively. Furthermore, Mavroeidis:2014ge state that amongst the different specifications for the NKPC they consider, the confidence sets implied by these more powerful methods are similar to the ones implied by the $S$ statistic.} Letting $k_s \geq p_1 + p_2$ denote the number of IVs selected, the $S$ statistic is given by $T$ times the value of the continuously updated GMM objective function given by

equation[equation omitted — 118 chars of source]

where $\varepsilon_T(\theta_0) = T^{-1}\sum_{t = 1}^TZ_{st}\varepsilon_{0t}$, $Z_{st}$ is the $k_s\times 1$ vector containing the IVs selected, and $W_T(\theta_0)$ is the continuously updated $k_s\times k_s$ weight matrix that is a consistent estimator of the covariance matrix of the moment conditions of the selected IVs as in Kleibergen:2009do, Stock:2000vt. Throughout, I use the heteroscedasticity and autocorrelation consistent (HAC) estimator of Newey:1987ua. Under the null hypothesis in Equation (ref) and the regularity conditions discussed in Stock:2000vt, this statistic is asymptotically $\chi^2_{k_s}$. Whenever $p_2 \neq 0$ (i.e., there are exogenous covariates in the relation), the exogenous covariates can be concentrated out, to yield the concentrated $S$ statistic as in Stock:2000vt.\footnote{In both the simulations and the empirical application below, the constant and the one-period lagged inflation are concentrated out.}

The $S$ statistic further recommends itself in this context because it allows for a straightforward test of the exclusion restrictions of the IVs. It may be hoped that any substantial bias caused by improper selection may be flagged in the form of a low $p$-value for the test of the null hypothesis that the IVs selected, $Z_s$, are uncorrelated with the structural error term, $\varepsilon$. To investigate this possibility further, in the simulations, I also evaluate the weak-identification robust Hansen test. This is given by the minimum value of the $S$ statistic in Equation (ref). Without making an assumption of strong identification, this statistic is asymptotically bounded by a $\chi^2_{k_s-p_2}$ distribution Mavroeidis:2014ge, which provides a weak-identification robust critical value for the test of the overidentifying restrictions of the IVs selected.

Ad-Hoc Selection of Instrumental Variables

Conducting inference with the $S$ statistic requires selecting a sufficiently small subset of $k_s$ IVs from the available $k$ IVs.\footnote{An often-used rule of thumb is to select $k_s$ to be of the order of magnitude of $T^{1/3}$. This rate result is motivated by the results in Andrews:2007bl, and Newey:2009fs, who show that this rate condition is sufficient for the case of independent data. Recently, fully weak-identification robust AR-type statistics have been developed that allow for the number of IVs to be of the order of magnitude of $T$ Mikusheva:2020uo, Crudu:2020bk, Anatolyev:2010kk. However, all of these approaches treat the IVs as fixed, and are hence not applicable in the context of time series. } In most of the empirical studies on IV-based limited-information inference on macroeconomic relations, no explicit reason is given for choosing the $k_s$ IVs that are subsequently used for analysis. Often, the choice of IVs is simply motivated with reference to previous studies that used those IVs. It is hence impossible to model the choice of IVs of the previous literature accurately in a simulation exercise. As an (imperfect) approximation, I consider the following three selection procedures.

The first selection procedure involves randomly selecting $k_s$ IVs out of the $k$ available IVs. Since the selection of IVs is not informed by the data itself, this selection procedure is guaranteed to not violate the identifying moment conditions.

The second selection procedure I consider will be referred to as crude thresholding. This involves first computing $p_1$ separate $k\times 1$ vectors containing the sample correlations between the endogenous variables and all the candidate IVs, sorting the IVs in descending order of correlation, and constructing the vector of IVs for post-selection inference by taking the union of the first $\floor*{k_s/p_1}$ entries in each of the vectors. By selecting the variables based on in-sample correlations, this selection procedure is likely to break the exclusion restriction of the IVs selected. Although (to my knowledge) this crude thresholding has not been applied to IV-based limited-information inference, more sophisticated versions of thresholding have been considered in the past (e.g., Mirza:2014wd and Bayar:2018er).\footnote{The hard thresholding in Mirza:2014wd and Bayar:2018er is not applicable in high-dimensional contexts, since OLS is infeasible when there are more variables than observations.}

The first two selection procedures (random selection and crude thresholding) arguably cover the extremes in terms of the effects IV selection can have on the validity of the IVs. Random selection provides the selection ideal, since it leaves the identifying moment conditions completely unaffected. However, particularly with reference to the traditional IVs often considered in the literature, it seems unlikely that random selection (over the very many available predetermined macroeconomic time series) led to choosing proximate lags of the endogenous variables as IVs. Indeed, given the persistence of most macroeconomic time series (and hence of the endogenous variables in any given application), it seems plausible that at least part of the motivation for considering proximate lags of the endogenous variables as IVs stems from their ability to usefully explain some of their in-sample variation. Suggestive evidence for this type of selection is also given by the fact that the IVs selected by crude thresholding in the empirical application in Section (ref) show substantial overlap with these traditional IVs.\footnote{See also the ranking of IVs based on $t$-values in Mirza:2014wd.} Therefore, it seems likely that random selection and crude thresholding provide a suggestive lower and upper bound on the selection-induced bias that could underlie existing empirical applications.

The third selection procedure I consider is a LASSO-based selection of IVs. This is motivated by the recent increase in popularity of penalisation-based approaches to the (very) many IV problem (see Hansen:2014ie, Belloni:2012kw, Ng:2011eq). Furthermore, LASSO-based approaches to IV selection have also been applied to the case of NKPCs in Berriel:2019vp, Berriel:2016hj. Here, IVs are selected by solving a LASSO optimisation problem of the following form for each endogenous variable

equation*[equation* omitted — 210 chars of source]

where $Y_{rt}$ is the element in position $t$ of the $T\times 1$ vector $Y_r$ given by the $r^{th}$ column of $Y$, $\zeta_r$ for $r = 1, \dots, p_1$ is a $k\times 1$ vector, and $\Lambda_r > 0$ for $r = 1, \dots, p_1$ are scalar penalty parameters that are set such that $\hat{\zeta}_r$ has $\floor*{k_s/p_1}$ elements. The IVs selected are given by the IVs that have at least one corresponding non-zero entry in at least one of $\hat{\zeta}_r$ for $r = 1, \dots, p_1$.

A High-Dimensional Sup Score Test for Dependent Data

Although the interplay between weak identification and high-dimensional IVs has recently received some attention (see Hansen:2014ie), none of the currently available approaches are both robust to arbitrarily weak identifcation and applicable in a time-series context. Indeed, to the best of my knowledge, the only approach that is formally robust to arbitrarily weak identification in the presence of very many IVs is the Sup Score test of Belloni:2012kw. The Sup Score test of Belloni:2012kw, however, treats the IVs as fixed, and is hence not applicable in time-series contexts. In this section, I propose a Sup Score test that remains valid under high-dimensional dependent data using recent results of Zhang:2018hy, Zhang:2014vm.

The Sup Score statistic I propose is given by

equation[equation omitted — 150 chars of source]

This can be seen as a non-studentised version of the Belloni:2012kw Sup Score statistic, which in turn can be interpreted as an extension to high dimensions of the Anderson:1949tx (AR) statistic. It also bears some resemblance to the non-studentised AR statistic proposed by Horowitz:2018to.

The critical values for the test statistic in Equation (ref) are computed using a block bootstrap. Let $l_T \equiv \floor*{T / b_T}$, where $b_T$ is the block length. Define the block sums

equation*[equation* omitted — 217 chars of source]

where $\left\{\mkern 1.5mu\overline{\mkern-1.5muZ'\varepsilon_0\mkern-1.5mu}\mkern 1.5mu\right\}_j$ is the $j^{th}$ element of the $k\times 1$ vector $\frac{1}{T}\sum_{t = 1}^TZ_t\varepsilon_t$. Consider the bootstrap statistic given by

equation*[equation* omitted — 136 chars of source]

where $\{e_t\}$ is a sequence of i.i.d. $\mathcal{N}[0, 1]$ random variables. The critical value for a test of size $\alpha$ of Equation (ref) is given by

equation*[equation* omitted — 167 chars of source]

The decision rule for testing the null hypothesis in Equation (ref) at the $\alpha$ level of significance is given by

equation*[equation* omitted — 68 chars of source]

I now turn to conditions that are sufficient to ensure that the test described above has correct size.

assumption$\left.\right.$ \begin{enumerate}[label = \roman*.] • $Z_t\varepsilon_t$ is a stationary time series that allows for the causal representation $Z_{t}\varepsilon_t = \mathcal{G}(\dots, u_{t-1}, u_t)$ for some measurable function $\mathcal G$, where $u_t$ are a sequence of mean-zero i.i.d. random variables. Furthermore, assume that $Z_{tj}\varepsilon_t = \mathcal{G}_j(\dots, u_{t-1}, u_t)$ for all $j = 1, \dots, k$, where $\mathcal{G}_j$ is the $j$th component of the map $\mathcal G$. • $\mathbb{E}[Z_t\varepsilon_t] = 0$, $\mathbb{E}[Z_{tj}^2\varepsilon_t^2] > 0$, and $\mathbb{E}[Z_{tj}^4\varepsilon_t^4] < \infty$ for all $j = 1, \dots, k$. • $k \lesssim \text{exp}(T^b)$, $b_T \lesssim T^{\tilde{b}}$ for $b < 1/15$, $4\tilde{b} + 7b < 1$, $\tilde{b} - 2b > 0$. • $\underset{1 \leq j, h \leq k}{\text{ max }}\sum_{l = -\infty}^\infty |l| \mathbb{E}[|Z_{tj}\varepsilon_{t}Z_{t+l, h}\varepsilon_{t+l}|] = O(T^{\breve{b}})$, $\breve{b} < \tilde{b} - 2b$. • $\mathbb{E}[|\mathcal{G}_j(\dots, u_{t-1}, u_t) - \mathcal{G}_j(\dots, u^*_{-1}, u^*_0, u_1, \dots, u_t)|^q] \leq C \rho^t$, for some $0 <\rho < 1$, and some positive constant $C$, where $q \geq 4$, and $\{u^*_t\}$ are i.i.d. copies of $\{u_t\}$. \end{enumerate}

Assumption (ref).(ref) requires the product of the IVs and the error terms to be stationary, and have some causal representation. Assumption (ref).(ref) makes weak assumptions on the moments of the data, and includes the identifying moment condition. In practice, I standardise the IVs in-sample to ensure that the test is invariant to the scaling of IVs. Assumption (ref).(ref) bounds the degree of high dimensionality permitted and the size of the block bootstraps. Although the restriction on the dimensionality ($b < 1/15$) is stronger than the ones usually encountered in the independent case (see Deng:2020wk, Belloni:2012kw), it still allows for very many IVs compared to the sample size. Assumption (ref).(ref) imposes restrictions on the correlation of the product of the IVs with the error term across different points in time. Assumption (ref).(ref) imposes a (uniform) Geometric Moment Contraction (GMC) restriction on the product of the IVs and the error terms as in Wang:2019ta. The GMC requires that the process under consideration have a sufficiently `short memory'. Processes that obey such a condition include (under suitable assumptions) standard linear processes (e.g., standard vector autoregressions and Volterra processes) as well as several nonlinear processes (e.g., autoregressive models with conditional heteroscedasticity, random coefficient autoregressive models, and exponential autoregressive models). I refer to Wang:2019ta, Zhang:2018hy, Chen:2016fh, Zhang:2014vm, Wu:2005uh, Hsing:2004gma and the references therein for a discussion of the different processes that obey such a condition.

No assumption on the first stage (i.e., the relationship between $Y$ and $Z$) has to be made. This means that the proposed Sup Score test is uniformly valid over all (finite) values of the coefficient on the IVs in the first stage (including arbitrarily weak identification). This also means that no restriction on the factor or sparsity structure of the first stage has to be imposed. The lack of assumptions on the first stage also implies that the Sup Score test does not suffer from any `missing IV problem' (see also Dufour:2009bu).

Whether these conditions are satisfied in any given macroeconomic application depends on the error terms (i.e., the structural equation), and on the properties of the excluded IVs. Example (ref) shows that under suitable assumptions on the error term that encompass, amongst others, some popular assumptions made in the literature on NKPCs (e.g., Dufour:2006bh), it is only required that the IVs satisfy a GMC condition. This is attractive in the context of limited-information inference, since the researcher only has to assume that the IVs belong to one of the many processes that have been shown to obey such a condition, without having to take a stance on the particular process.

exampleAssume that $\varepsilon_t$ is i.i.d. across $t$, $\mathbb{E}[\varepsilon_t] = 0$, $\mathbb{E}[\varepsilon_t^2] > 0$, and $\mathbb{E}[\varepsilon_t^4] < \infty$. Assume that $Z_t$ is a stationary time series and allows for the causal representation $Z_t = \mathcal{F}(\dots, v_{t-1}, v_t), Z_{tj} = \mathcal{F}_j(\dots, v_{t-1}, v_t)$ for some measurable function $\mathcal F$, where $v_t$ are a sequence of mean-zero i.i.d. random variables (independent of $\varepsilon_t$). Assume further that $\mathbb{E}[Z_{tj}^2] >0$ and $\mathbb{E}[Z_{tj}^4] < \infty$ for all $j = 1, \dots, k$. Assume that the conditions on the dimensionality of the IV problem in Assumption (ref).(ref) are satisfied. Further, assume that $Z_t$ satisfies: \begin{equation} \mathbb{E}[|Z_{tj} - \mathcal{F}_j(\dots, v^*_{-1}, v_0^*, v_1, \dots, v_t)|^4] < \tilde{C}{\tilde{\rho}}^{t} \end{equation} where $\{v^*_t\}$ are i.i.d. copies of $\{v_t\}$, $\tilde{C}$ is some constant, and $0 <\tilde{\rho} < 1)$. Then the conditions in Assumption (ref). hold. \begin{proof} See Appendix (ref). \end{proof}

I now state the main theoretical result of this paper, which ensures that the approach proposed controls the size of the test.\footnote{It should be noted, however, that--similarly to other sup-based test statistics, such as in Chernozhukov:2018ema, Belloni:2012kw--the above approach is not efficient. This is to be expected, given the weak assumptions made on (the structure of) the IVs. The (finite-sample) power properties of the above approach will be investigated in the simulation section below. The results show that it has non-trivial power.}

theoremUnder Assumption (ref). and the null hypothesis in Equation (ref), \begin{equation*} \underset{T\to \infty}{ lim }\mathbb{P}(Reject H_0) \leq \alpha. \end{equation*}
proofSee Appendix (ref).

Theorem (ref) makes it possible to construct confidence sets by inverting the test as outlined above.

Simulations

The simulations presented in this section serve a twofold purpose. First, I use the simulations to study how improper selection of IVs can lead to problematic post-selection inference. Second, I use the simulations to illustrate the asymptotic validity of the Sup Score test established in the section above, as well as its finite-sample power properties. Taken together, the simulations hence motivate and further justify applying the Sup Score test proposed in Section (ref) in practice.\footnote{Due to the focus of this paper on the bias introduced by IV selection, I do not consider the factor-based approaches of Kapetanios:2015fp, Mirza:2014wd. The substantial biases caused by the improper selection of a small number of IVs can also serve to motivate the use of such factor methods. However, the factor-based GMM approach in Mirza:2014wd seems to treat the number of IVs as fixed (and does not provide formal conditions for validity), while the factor AR statistic of Kapetanios:2015fp is only applicable in a high-dimensional context if a sufficiently strong factor structure is assumed. In contrast, the Sup Score test proposed in this paper remains valid in high-dimensional contexts regardless of the factor or sparsity structure of the IVs.}

The simulations in this paper are based on the approach in Mavroeidis:2014ge. The central difference is that rather than modelling the econometrician as having perfect knowledge of the relevant IVs, and incorrectly employing methods that are not robust to weak identification, I model the econometrician as using exclusively weak-identification robust methods, but not knowing which IVs correspond to the truly relevant ones. Given the extensive literature that pointed out that NKPCs can suffer from weak identification, this setup seems closer to the estimation problem that an econometrician is likely to face.

I base my simulations on the simplest possible specification considered in Mavroeidis:2014ge. This involves imposing the restriction $\gamma_b + \gamma_f = 1$ (which is known to the econometrician) and setting $c = 0$ (which is not known to the econometrician), so that the NKPC can be re-written as

equation[equation omitted — 134 chars of source]

I embed this NKPC into a dynamic system by specifying that the reduced-form dynamics of the forcing variable and inflation follow a VAR model given by

equation[equation omitted — 380 chars of source]

where

equation*[equation* omitted — 292 chars of source]

and $f_t$ is a scalar factor variable. All coefficients except for $a_{11}, a_{12}$, and $a_{13}$ have to be calibrated. The coefficients $a_{11}, a_{12}$, and $a_{13}$ are backed out of the NKPC based on the Anderson:1985wr algorithm.

High dimensionality of the IVs is introduced by specifying that there exists an $m\times 1$ vector of variables $Q_t$ that follow the process given by

equation[equation omitted — 124 chars of source]

where $\xi$ is an $m\times 1$ vector of factor loadings.

The econometrician conducts inference on $\lambda$ and $\gamma_f$ within the GIV and RE framework,

equation[equation omitted — 206 chars of source]

where the variables are defined as in Section (ref) and Section (ref). The econometrician does not observe the factor itself, but only observes the forcing variable and inflation, as well as the $m$ variables in $Q_t$. In this setup, $Z_{st}$ is a subset of the available IVs given by $Z_t = [1, \pi_{t-1}, s_{t-1}, Q_{t-1}']'$ that always includes a constant (since it is specified in the structural equation the econometrician considers).\footnote{For the case of post-selection inference based on the $S$ statistic, the constant is concentrated out. For the case of the Sup Score statistic, it is partialled out.}

This setup recommends itself for two reasons. First, it constitutes a minimal departure from popular simulations in the existing literature. This ensures that any reported results are not an artefact of a particularly uncharitable setup.\footnote{It is, for instance, straightforward to include the variables in $Q_t$ directly in the reduced-form VARs. However, the results from such a DGP are very sensitive to the particular calibration of the parameters chosen.} Second, it ensures the existence of a sufficiently small set of (excluded) `oracle IVs' (given by $\pi_{t-1}, s_{t-1}, f_{t-1}$) without necessarily imposing a sparse setup on the observed first-stage projection (although it can be imposed by setting $a_{13} = a_{23} = \omega_{13} = \omega_{31} = \omega_{23} = \omega_{32} = 0$ or simply $\xi = 0$).\footnote{Ensuring a sufficiently sparse set of `oracle IVs' further motivates considering only a single lag of a single factor.} The former is desirable because it allows for a comparison of the different ad-hoc inference procedures relative to the most efficient approach (conditional on using the $S$ statistic). The latter is desirable because it allows for a more general (and perhaps more realistic, see Giannone:2018uv) approach to modelling the first stage. This setup is able to achieve both a sparse (unobserved) oracle first stage and an observed first stage that is not necessarily sparse because the elements of $Q_{t-1}$ that have a non-zero corresponding entry in $\xi$ will contain some relevant variation for identification, due to the dependence of the endogenous variables $[\pi_{t+1} - \pi_{t-1}, s_t]'$ on $f_{t-1}$. Based on the setup above, it is possible to derive two different concentration parameters ($\mu^2_O$ and $\mu^2_E$) that reflect the strength of identification in the sparse unobserved oracle first stage and the observed first stage. The details are given in Appendix (ref).

The calibrations are as follows. Throughout, I set $\gamma_f = 0.8$, $\lambda = 0.05$, $T = 100$, and $\omega_{11} = 0.07$, $\omega_{12} = \omega_{21} = 0.03$, $\omega_{22} = 0.7$, $\omega_{13} = \omega_{31} = \omega_{23} = \omega_{32} = 0$, and $\omega_{33} = 0.4$ (see also Mavroeidis:2014ge). The results are not sensitive to this choice of covariance matrix, and this setup makes it possible to create a perfectly sparse observed first stage by setting $a_{23} = 0$. For simplicity, I set $a_{31} = a_{32} = 0$, so that the factor structure follows an autoregressive process with coefficient given by $a_{33} = 0.7$ (the results do no change appreciably if this is relaxed or a different choice for $a_{33}$ is considered). I set $m = 200$ and $\xi_{q} = \tau(-1)^q\log\left((q+1)^2/mq\right)$ for $q = 1, \dots, m$ and $\tau = 0.05$. This is meant to provide a deterministic calibration that balances positive and negative, as well as large and small coefficients. The small value chosen for $\tau$ ensures that the information on the factor contained in the observed variables is sufficiently diluted, and that there is some interesting variation in the informational content of the unobserved oracle first stage and the one actually observed.\footnote{I refer to Appendix (ref) for more details on this. The derivations also show that choosing small values of $\tau$ has a similar effect to choosing a larger term for the variance of the errors in Equation (ref).} The results are unaffected by different choices of $\xi$ or $\tau$. For all selection procedures, I force the selection of $k_s = 4$ IVs to ensure that the first stage is not overfitted, which again ensures that any distortions in inference are attributable to the selection step itself. For the $S$ statistic, I set the lag-length for the Newey:1987ua HAC variance estimator to 4. For the Sup Score test proposed above, I set the block size to $b_T = 4$ and the bootstrap replications to $500$. I allow $a_{21}$, $a_{22}$, and $a_{23}$ to take on different values. The coefficient $a_{23}$ controls how informative the factor is in predicting the endogenous variables, and by extension how informative the variables in $Q_{t-1}$ are.

Table (ref) shows the size of the $S$ and Sup Score statistic following the different selection procedures outlined above for a test with nominal size 10%. The calibrations chosen ensure that a broad range of identification strength and sparsity structures are considered. The first panel for $a_{23}$ corresponds to the perfectly sparse first stage where none of the variables in $Q_{t-1}$ are informative IVs. As a consequence, the concentration parameter of the unobserved oracle first stage is the same as the one that is observed. The second and third panel increase the dependence of the two endogenous variables on the unobserved factor. Due to the dense calibration of $\xi$, this means that all of the IVs observed by the econometrician are at least somewhat informative. Since the variables in $Q_t$ contain noisy information on the unobserved factor, the concentration parameter in the observed first stage will now be lower than the concentration parameter of the unobserved oracle first stage. As expected, the oracle IVs yield correct, if somewhat conservative, size. Since random selection does not make use of any correlations present in the actual data, the $S$ statistic with randomly selected IVs also yields correct size. The results for crude thresholding and the LASSO suggest that in all cases size is not controlled, although the distortions appear to be somewhat milder for the LASSO. The results for the Sup Score test proposed in this paper show that the test controls for size regardless of the DGP considered.

Table (ref) also reports the rejection frequency of a two-step approach that tests the null hypothesis at a given level of significance only if the robust test of overidentifying restrictions fails to reject the hypothesis of exogeneity for the IVs selected at that level of significance. This is a very conservative approach. Indeed, when faced with evidence that the selected IVs may be endogenous, rather than abandoning the analysis altogether, it seems more likely that the econometrician will proceed to select other IVs, potentially worsening the endogeneity bias. Even in this conservative approach, crude thresholding fails to control for size. LASSO selection followed by this two-step approach appears to control for size. These results suggest that while the test of overidentifying restrictions can help mitigate some of the endogeneity bias introduced by improper selection, it is unable to fully remove it.

Figure (ref) shows the power of the different approaches. I present the results for the case where $a_{21} = a_{22} = a_{23} = 0.450$. The results are similar for other calibrations. The map traced out by the oracle IVs corresponds to the most powerful procedure possible (conditional on exclusively using the $S$ statistic) that controls for size. The results show that randomly selecting IVs yields no power. This is unsurprising, given that in this setup the first stage is sparse, so that random selection predominantly selects not very informative IVs. The power heatmaps for crude thresholding and the LASSO have a similar shape to the oracle heatmaps. However, for certain parts of the parameter space considered, the rejection frequency of these procedures is substantially higher than the one of the oracle test. Conditional on using the same test post-selection, both crude thresholding and the LASSO can be at most as powerful as the test that directly uses the oracle IVs. Therefore, this excess rejection frequency is spurious, which in practice would translate to small confidence sets. The power heatmaps for the Sup Score test show that the Sup Score test has non-trivial power.

table[table omitted — 4,309 chars of source]
figure[figure omitted — 1,217 chars of source]

Empirical Application

Data for the empirical part of this paper is taken from FRED. I use the non-farm labour share as transformed in Gali:1999tx as the forcing variable. I use the inflation rate implied by the GDP deflator. I consider the period 1974Q2-2018Q4, and include 90 variables aimed to reflect different parts of the US economy based on the list in McCracken:2015ct with four lags each, transforming them as recommended therein.\footnote{I do not include all the variables listed in McCracken:2015ct since they are not all available over a sufficiently long period of time.} This yields 179 observations with 359 IVs. Appendix (ref) contains a detailed description of the data.

A natural question to ask is whether considering these 359 IVs is enough to dispel concerns about potential endogeneity biases. Though being more than any number of IVs previously considered in the literature, there are certainly more valid IVs (i.e., additional predetermined variables). However, a substantial endogeneity bias caused by the selection of these 359 IVs would emerge only if the variables were included in the list of McCracken:2015ct based on their correlation with the endogenous variables of this application. This seems very unlikely.

For all ad-hoc selection procedures, I limit the number of IVs selected to four, to ensure that overfitting is not a concern, and set the lag-length for the Newey:1987ua HAC variance estimator to 4. The results do not change appreciably when other values are chosen. The confidence sets yielded by the traditional IVs and the ad-hoc selection procedures are shown in Figure (ref), and the corresponding IVs are listed in Table (ref). Mirroring the results in Section (ref), the confidence set resulting from random selection is extremely wide, and suggests that the hybrid NKPC is essentially unidentified. The confidence set from applying the LASSO is smaller than the one implied by random selection, but it also does not exclude that the coefficient on expected inflation is in fact equal to zero.

table[table omitted — 1,336 chars of source]
figure[figure omitted — 1,140 chars of source]

The confidence sets implied by traditional IVs and crude thresholding are qualitatively very similar. This similarity is explained by the fact that the IVs chosen by crude thresholding are very similar to the traditional IVs, as shown in Table (ref). This suggests two things. First, it suggests that the ad-hoc selection procedures used in this paper may in fact provide a reasonable approximation to the approach taken for selecting IVs in the past literature. Second, given that in the simulation exercise in Section (ref) crude thresholding yields the worst size distortions of the selection procedures considered, this result suggests that the confidence sets reported in the previous literature are likely to suffer from at least some distortion due to endogeneity bias. In particular, it suggests that the process of trying to find `strong' IVs may have led to an undercovering of the true parameter values.

The confidence sets of the Sup Score test proposed in this paper for different block lengths (4, 6, 8, and 10) are shown in Figure (ref). For all block lengths considered (the results do not appear to be sensitive to the choice of block length), the confidence sets are smaller than the ones yielded by the random selection approach, but wider than for the traditional, crude thresholding, and LASSO approach. This is likely due to a combination of the incorrect size of the latter approaches, and the low power of the Sup Score test documented in Section (ref). The results suggest that while certain parts of the parameter space considered can be rejected at the 10% level of significance, neither $\lambda$ nor $\gamma_f$ are found to be different from zero for all values of the parameter space considered.

figure[figure omitted — 1,051 chars of source]

The Sup Score test can also help shed some light on what the most relevant IVs are, since there is a (likely) unique IV that maximises the Sup Score statistic in Equation (ref) for every null hypothesis being tested. I record the identity of the IV maximising the Sup Score statistic for each null hypothesis tested, and report the results in Table (ref) and Figure (ref). Table (ref) shows the identity of the IVs maximising the Sup Score statistic. Figure (ref) shows in which part of the parameter space the different IVs maximise the Sup Score statistic. Two things stand out. First, the IVs that feature prominently in Table (ref) (i.e., IVs selected by the ad-hoc procedures) also tend to feature in the set of IVs maximising the Sup Score statistic (i.e., lags of PRS85006173 and CES3000000008). The reason why the confidence sets are larger for the Sup Score test is in part due to the Sup Score test being able to account for the very many other valid IVs that these variables were chosen from. Second, there are some IVs that maximise the Sup Score statistic that do not feature in Table (ref), such as the three-period lagged Housing Starts in Northeast Census Region (e.g., HOUSTNE.-3). This relates to the predominant motivation for wanting to consider very many IVs: the truly relevant IVs can often be `exotic', in the sense that intuition alone would not point to their relevance.

table[table omitted — 1,145 chars of source]
figure[figure omitted — 346 chars of source]

Conclusion

IV-based limited-information estimation of single equations has become increasingly popular in Macroeconomics over the last 20 years. Using a simulation exercise based on NKPCs, I showed that selecting IVs in ad-hoc ways (random selection, crude thresholding, and LASSO) can invalidate them, thus yielding invalid inference even if tests with desirable properties (such as robustness to weak identification) are used post-selection. To address this issue, I propose a Sup Score test that remains valid for high-dimensional IVs and for time series data. In the same simulation exercise that showed that ad-hoc selection procedures can lead to invalid inference, this statistic yielded correct size and reasonable power. Finally, I applied the Sup Score test to conduct inference on the US NKPC with 359 IVs on a sample size of 179 observations. The results showed that the confidence sets implied by the Sup Score test are substantially wider than the ones of all other approaches. The simulation results and the empirical application point to the importance of developing further high-dimensional IV methods with good power properties that remain valid under dependence and arbitrarily weak identification.

\printbibliography