EconBase
← Back to paper

Inferential Theory for Granular Instrumental Variables in High Dimensions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

133,855 characters · 10 sections · 86 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inferential Theory for Granular Instrumental Variables in High Dimensions

refsegment\begin{titlepage} \thispagestyle{empty} \begin{abstract} {0.0cm} \singlespacing The Granular Instrumental Variables (GIV) methodology exploits panels with factor error structures to construct instruments to estimate structural time series models with endogeneity even after controlling for latent factors. We extend the GIV methodology in several dimensions. First, we extend the identification procedure to a large $N$ and large $T$ framework, which depends on the asymptotic Herfindahl index of the size distribution of $N$ cross-sectional units. Second, we treat both the factors and loadings as unknown and show that the sampling error in the estimated instrument and factors is negligible when considering the limiting distribution of the structural parameters. Third, we show that the sampling error in the high-dimensional precision matrix is negligible in our estimation algorithm. Fourth, we overidentify the structural parameters with additional constructed instruments, which leads to efficiency gains. Monte Carlo evidence is presented to support our asymptotic theory and application to the global crude oil market leads to new results. \end{abstract}\thispagestyle{empty} {Keywords:\,\relax }{Interactive effects, Factor error structure, Simultaneity, Power-law tails, Asymptotic Herfindahl index, Global crude oil market, Supply and demand elasticities, Precision matrix.} JEL Classifications: C26, C36, C38, C46, C55 \end{titlepage}

Introduction

{0.5cm}

In the absence of randomized control trials, finding valid and strong instruments to circumvent unobserved confounders is a very challenging task. The Granular Instrumental Variables, hereafter GIV, methodology that gabaix2020granular propose, establishes a systematic way to construct instruments from suitably weighted idiosyncratic shocks, from observational datasets and use them as instruments for aggregate endogenous variables.

Constructing instruments. There are some existing methodologies which seek to eliminate the need to find an instrument. A leading example is the arellano1991some framework in the context of estimating the speed of adjustment or state dependence parameters using dynamic panel data models with fixed effects, in which higher order lags of the dependent variable serve as instruments for the included lags of the dependent variable. The bartik1991benefits methodology (aka shift-share estimators) where instruments are constructed from identities involving the (endogenous) explanatory variable whose shift component is interacted with shares. The rigobon2003identification setting exploits the existence of structural breaks in the conditional heteroskedasticity regime, which is common place in many applications of interest. This allows one to bring a system with less equations than unknowns to a just identified system with as many equations as unknowns. The bai2010instrumental methodology lays out a panel simultaneous equations model (similar to the model analyzed in this paper) where the estimated (strong) factors can be used as instrumental variables under certain conditions. We will return to the bai2010instrumental methodology when we overidentify the structural parameters of interest as it is inspired by their framework. The vast methodological refinements cited within the papers referenced above are not listed here for brevity.

Microeconomic (granular) origins of aggregate fluctuations. How can idiosyncratic shocks be relevant for endogenous aggregate variables? The literature on "granularity" traces back to historic debates in macroeconomics; no attempt to fully catalog this debate is made here, rather a concise summary is offered. long1983real demonstrate that in a multisector stochastic neoclassical growth model, sectoral shocks (as opposed to aggregate shocks) can potentially lead to GDP fluctuations. Intuitively, complex production processes form sectoral linkages which in turn provide a transmission mechanism of shocks across sectors. Subsequently, horvath2000sectoral and dupor1999aggregation debate whether sectoral shocks decay according to $\frac{1}{\sqrt{N}}$ as the central limit theorem would suggest. gabaix2011granular provides an initial theoretical solution to the debate by showing that when the firm size distribution is heavy tailed, the central limit theorem does not apply and sectoral volatility decays much slower than $\frac{1}{\sqrt{N}}$. gabaix2011granular coins this mechanism as the so-called "granular" hypothesis, in which the economy is composed of incompressible grains as opposed to infinitesimally small micro units. acemoglu2012network formulates a network approach to demonstrate that sectoral idiosyncratic shocks generate non-negligible aggregate volatility when there exists sufficient asymmetry in the input-output relationships. pesaran2020econometric build off of the theoretical approach of acemoglu2012network and develop econometric theory to measure the degree of network dominance and in their application they find some evidence of sector-specific shock propagation albeit not overwhelmingly strong for the US input-output accounts data over the period 1972-2002. More empirical evidence for such propagation mechanism is presented in gatti2005new, canals2007trade, koren2007volatility, blank2009shocks, malevergne2009professor, yan2011role, gabaix2011granular, carvalho2013great, schiaffi2013granularity, acemoglu2017microeconomic, jannati2017geographic and lera2017quantification.

GIV, Gabaix and Koijen (2021). In an econometric framework, GK illustrate that when the market under consideration is sufficiently concentrated, then one can use the collection of idiosyncratic shocks to individual micro units, at each time period $t$, as an instrument for endogenous aggregate variables. The instrumental relevance follows heuristically from the paragraphs above. The exogeneity condition, as in any instrumental variables procedure, requires assumptions on unobserved random variables. However it should be noted that the exogeneity condition exploited in this framework is a relatively mild assumption that is often made in factor models (e.g. bai2002determining) for identification purposes. The insight and contribution of GK opens the doors to a wide possibility of ways in which one can continue building on the promising new GIV methodology.

Contributions of this paper. Our contributions to the GIV methodology are primarily focused on the underlying econometric issues. First, we naturally extend GK's identification procedure to a large $N$ and large $T$ framework (GK formally introduced GIV for a fixed $N$ and large $T$) by establishing and restricting the asymptotic behavior of the Herfindahl index for large $N$ markets as a function of the tail index of the size distribution. Given the large $N$ and large $T$ framework, we treat both the factors and loadings as unknown and allow the idiosyncratic error term to be weakly cross-sectionally correlated.\footnote{GK treat the factor loadings as known and extract the factors via period-by-period cross-sectional regressions. While they advocate extraction of latent factors via principal components analysis when loadings are unknown, they abstract away from the corresponding sampling error. We will show that the sampling error is indeed negligible.} As such, from our preliminary stage, we extract not only the estimated factors but also the estimated loadings via principal components analysis, PCA hereafter, or depending on the generality of the model ($k_x \neq 0$ in our notation from Section (ref)), we use the iterative OLS-PCA method of bai2009panel. Second, we show that the sampling error in the estimated instrument and estimated factors is negligible when considering the limiting distribution of the structural parameters of interest; that is, the estimator is robust to the latent factor structure. Moreover, the exogeneity requirement for one of the structural parameters generally depends on a potentially high dimensional precision matrix (the inverse of the covariance matrix). Third, we show that the sampling error in the high dimensional precision matrix is negligible in our iterative estimation algorithm for said structural parameter. Fourth, we overidentify the structural parameters which leads to efficiency gains. This leads to new and improved results in our empirical application of GIV to the global crude oil markets. Monte Carlo evidence is presented to confirm the finite sample behavior of our estimators are well approximated by the asymptotic distributions. We label our refinement to the GIV methodology as Feasible Granular Instrumental Variables or FGIV for short. Finally, an empirical application of the estimation methods to estimate demand and supply elasticities of the global crude oil markets are presented to demonstrate the estimation procedures.

Notation. We distinguish vectors and matrices from scalars by making an object bold. Let $\{X_{it}, i=1,\dots,N; t=1,\dots,T\}$ be a double index process of random variables where $N$ denotes the number of cross-sectional units and $T$ denotes the number of time periods. We frequently stack across $i$, in which we obtain $\underset{N \times 1}{\boldsymbol{X}_{\cdot t}}:=

pmatrix[pmatrix omitted — 36 chars of source]

'$. Similarly, if we stack across $t$ we obtain $\underset{T \times 1}{\boldsymbol{X}_{i \cdot}} :=

pmatrix[pmatrix omitted — 37 chars of source]

'$. When $\boldsymbol{X}_{it}$ is itself a vector, say of dimension $k$, then we obtain a matrix when we stack across $i$ or $t$, e.g. $\underset{N \times k}{\boldsymbol{X}_{\cdot t}}$ or $\underset{T \times k}{\boldsymbol{X}_{ i \cdot}}$. Define $X_{\boldsymbol{w}t}$ as the cross-sectionally weighted average of $X_{it}$, that is $X_{\boldsymbol{w}t}:= \boldsymbol{w}'X_{\cdot t}=\sum^N_{i=1} w_i X_{it}$. Common weights, $\underset{N \times 1}{\boldsymbol{w}}=(w_i)$, used frequently throughout the paper are (1) the precision weights, $\boldsymbol{E}:= \frac{\boldsymbol{\Sigma}^{-1}_u \boldsymbol{\iota}}{\boldsymbol{\iota}'\boldsymbol{\Sigma}^{-1}_u\boldsymbol{\iota}}$ where $\underset{N \times N}{\boldsymbol{\Sigma}_u}:= \mathbbm{E}(\boldsymbol{u}_{\cdot t}\boldsymbol{u}_{\cdot t}')$ is the covariance matrix of the idiosyncratic error term, $u_{it}$, $\underset{N \times 1}{\boldsymbol{\iota}}$ is a vector of ones and (2) the share weights, which we simply refer to as size weights, $\boldsymbol{S}:=

pmatrix[pmatrix omitted — 32 chars of source]

'$. Let $\widetilde{X}_{it}=X_{it}-\bar{X}_t$, where $\bar{X}_t=\frac{1}{N}\sum_{i=1}^N X_{it}$, denote a cross-sectionally demeaned variable. Unless otherwise specified, we denote the $L^2$-norm as $||\cdot||$ or sometimes explicitly as $||\cdot||_2$, the $L^1$-norm as $|| \cdot ||_1$ and the Frobenius norm as $||\cdot||_F$; if another norm is used, it will be explicitly noted. Given a square matrix $\boldsymbol{A}$, let $\gamma_{max}(\boldsymbol{A})$ denote the maximum eigenvalue of $\boldsymbol{A}$. Joint convergence of $N$ and $T$ will be denoted as $(N,T) \overset{j}\rightarrow \infty$ without any restriction on the relative rates; whenever restrictions on relative rates of convergence are imposed, it will be explicitly noted. The expression $\overset{p}{\rightarrow}$ denotes convergence in probability while $\overset{d}{\rightarrow}$ denotes convergence in distribution. The equation $\boldsymbol{y} = \mathcal{O}_p(\boldsymbol{x})$ states that the vector of random variables $\boldsymbol{y}$ is at most of order $\boldsymbol{x}$ in probability. The equation $a=\Theta_p(b)$ states that $a$ is stochastically bounded by $b$ and $b$ is stochastically bounded by $a$, hence $a$ and $b$ rise jointly proportionally.

Model

A general formulation of the model examined in this paper is given in the following panel simultaneous equations model with factor error structure

align*[align* omitted — 214 chars of source]

where $\boldsymbol{y}_{it}=

pmatrix[pmatrix omitted — 43 chars of source]

'$ is a $G \times 1$ vector of dependent variables, $\boldsymbol{x}_{it}=

pmatrix[pmatrix omitted — 44 chars of source]

'$ is a $k_x \times 1$ vector of strictly exogenous variables (which can be arbitrarily correlated with the common factors, $\boldsymbol{F}_t$, and/or the loadings, $\boldsymbol{\Lambda}_i$), $\boldsymbol{a}_t=

pmatrix[pmatrix omitted — 42 chars of source]

'$ is a $k_a \times 1$ vector of potentially endogenous aggregate variables, $\boldsymbol{v}_{it}$ is a $G \times 1$ vector of composite error terms which admit a low-rank plus sparse (factor structure) error decomposition, where $\boldsymbol{\Lambda}_i$ is an $r \times G$ matrix of latent factor loadings and $\boldsymbol{F}_t$ is an $r \times 1$ vector of latent factors.

In our exposition, we focus on the canonical setting of estimating the supply and demand elasticities in the global crude oil market, so we set the dimension of $G=2$ for supply and demand variables respectively. We take $k_x=0$ for ease of exposition but we present a general estimation algorithm for when $k_x \neq 0$. Moreover, we assume that only one of the $G=2$ variables has a panel structure, whereas the other variable is an aggregate time series. The main results extend relatively naturally to the case where both variables have a panel model. That is, $\boldsymbol{y}_{it}=

pmatrix[pmatrix omitted — 29 chars of source]

'$ where $d_t$ is the log change of aggregate crude oil consumption and $y_{it}$ is the log change of country $i$'s crude oil production, $a_t=p_t$, with $k_a =1$, is the log change of real crude oil price (where we deflate the nominal oil price with the U.S. general price in3dex).\footnote{One may wonder why $p_{t}$ is not disaggregated; in fact the $p_t$ we use can be considered as the weighted average of country specific real oil prices (in changes). As shown in \citet{mohaddes2016country}, for a proper global analysis, deflating the nominal oil price in U.S. dollars by the U.S. price index is generally theoretically invalid unless the law of one price holds universally. Namely, let $P_{it}$ denote the general price index faced by country $i$, $E_{it}$ denotes country $i$'s exchange rate measured as units of country $i$'s currency per U.S. dollar, $p_{it}$ denote country specific log of real oil prices and $\tilde{p}_t$ denotes nominal oil prices in U.S. dollars, if $E_{it}P_{US,t}=P_{it} \,\, \forall i$; then it follows that $\sum_{i=1}^N w_i p_{it}=\tilde{p}_t + \sum_{i=1}^N w_i log(E_{it}/P_{it})=\tilde{p}_t + \sum_{i=1}^N w_i log(1/P_{US,t})=\tilde{p}_t-p_{US,t} := p_t$. As it turns out, $p_t=\tilde{p}_t-p_{US,t}$ is an appropriate approximation as documented in \citet{mohaddes2016country} for their long run analysis, in the sense that it respects the long-run equilibrium relationships. We assume it is an appropriate approximation for our short-run analysis.} Given our stylizations the coefficient matrix $\boldsymbol{C}$ and composite error, $\boldsymbol{v}_{it}$ becomes

align*[align* omitted — 348 chars of source]

where the coefficients $\phi^d$ and $\phi^s$ denote the crude oil demand and supply elasticities, respectively, and $\boldsymbol{\eta}_t, \boldsymbol{\lambda}_i$ are $r \times 1$ vectors of latent factors and latent loadings, respectively. Our stylized simultaneous equations model takes the simple form \setcounter{equation}{0}

align[align omitted — 144 chars of source]

The global market clearing condition is given by $y_{St}=d_t$, where $y_{St}:=\boldsymbol{S}'\boldsymbol{y}_{\cdot t} =\sum_{i=1}^N S_{i} y_{it} $, $\boldsymbol{S}$ is the $N \times 1$ vector of shares that are normalized such that $\sum_{i=1}^N S_i =1$ and $i$ and $t$ take the values $i=1,\dots,N$ and $t=1,\dots,T$, respectively.\footnote{As oil is a storable good, one could easily allow oil prices to adjust to the gap between supply and demand, e.g. as in mohaddes2016country. This introduces more complex notations without adding any substance to the main points of the paper.} Making use of the global market clearing condition we see that

align[align omitted — 142 chars of source]

which makes the simultaneity clear, e.g., that prices are composed of size-weighted idiosyncratic shocks, aggregate supply shocks and the demand shock. The objective of the GIV methodology is to extract the idiosyncratic shocks and use them as instruments for price.

Demand estimation in the case of uniform loadings ($\boldsymbol{\lambda}_i = \boldsymbol{\lambda} \,\forall \, i$). To momentarily fix ideas, it is helpful to consider a major simplification when constructing the instrument. Suppose that the loadings are uniform, $\boldsymbol{\lambda}_i=\boldsymbol{\lambda} \, \forall i$. Then, the instrument$, z_t$, can be formed as

align*[align* omitted — 338 chars of source]

where $\boldsymbol{\Gamma} := \boldsymbol{S}- \boldsymbol{\iota}/N$ is an $N \times 1$ random vector such that $\boldsymbol{\iota}'\boldsymbol{\Gamma}=\sum_{i=1}^N \Gamma_i=0$, by construction. $\Gamma_i$ is random because we assume the shares follow a fat-tailed distribution, see (ref). Identification and estimation of demand by GIV requires that

align[align omitted — 156 chars of source]

(ref) is our exogeneity condition and (ref) gives $\mathbbm{E}(z_tp_t)\neq 0$, relevance. A sufficient condition for the moment condition in (ref) to be zero is $\mathbbm{E}(u_{it}\varepsilon_t|\boldsymbol{\Gamma})=0$, which effectively requires that conditional on size, $u_{it}$ and $\varepsilon_t$ are uncorrelated. Given relevance, exogeneity implies the following demand elasticity estimator $\widehat{\phi}^d=\frac{\sum_td_tz_t}{\sum_t p_tz_t}$. Intuitively, $z_t$ places larger weights on the idiosyncratic shocks to larger oil producers, these granular shocks will shift the supply curve while keeping the aggregate demand curve fixed since demand responds to these shocks only through their affects on prices. This allows for consistent estimation of the demand elasticity. The uniform loadings assumption in this case tremendously facilitate the analysis. Uniform loadings allow one to construct the instrument, as in (ref), from observables. \color{black} In practice, uniform loadings are quite restrictive and we subsequently relax this assumption. However, before moving on to the general case, we also illustrate supply estimation under simplifying assumptions to fix ideas. \color{black}

Supply estimation in the case of uniform loadings and $u_{it} \,\,\boldsymbol{ i.i.d.}$ Continuing on with the uniform loadings case, remarkably, GK show that one can use the same instrument, $z_t$, to also estimate the supply elasticity using a cross-sectionally aggregated supply equation. Now, GK further assume that $u_{it}$ are $i.i.d.$, $\mathbbm{E}(\boldsymbol{u}_{\cdot t}\boldsymbol{u}_{\cdot t}'):= \boldsymbol{\Sigma}_u=\sigma^2_u \boldsymbol{I}_N$, where $\boldsymbol{u}_{\cdot t}:=

pmatrix[pmatrix omitted — 37 chars of source]

'$ and $\boldsymbol{I}_N$ is the identity matrix and define the $N \times 1$ precision weight vector $\boldsymbol{E} := \frac{\boldsymbol{\Sigma}^{-1}_u\boldsymbol{\iota}}{\boldsymbol{\iota}'\boldsymbol{\Sigma}_u^{-1}\boldsymbol{\iota}}$ which reduces to $\boldsymbol{\iota}/N$ when $u_{it}$ are $i.i.d.$ across $i$. Aggregation of the supply equation is performed using the vector $\boldsymbol{E}$, we have that $y_{Et} =\phi^sp_t + \boldsymbol{\lambda}'\boldsymbol{\eta}_t + u_{Et}.$ Identification and estimation of supply by GIV requires that the instrument satisfies exogeneity with respect to the \textit{composite} error term\footnote{\color{black}In the general case to follow, we estimate the factors and thus only exploit $\mathbbm{E}(u_{Et}z_t)=0$ to estimate $\phi^s$.\color{red}}

align[align omitted — 97 chars of source]

The first term in (ref) has similar interpretation as in (ref), i.e., size-weighted idiosyncratic supply shocks are uncorrelated with the aggregate supply component, $\boldsymbol{\lambda}'\boldsymbol{\eta}_t$. Miraculously, the second term is exactly zero

align*[align* omitted — 427 chars of source]

\footnotetext{One may wonder then why this particular form of $\boldsymbol{\Gamma}$ was selected. Appealing to Proposition 3 in GK, they establish that $\boldsymbol{\Gamma}=\boldsymbol{S}-\boldsymbol{E},$ for this example, turns out to be the optimal weight vector, amongst the class of weights which sum-to-zero. $\boldsymbol{\Gamma}$ is optimal in the sense that it minimizes the asymptotic variance of the structural parameters.} The moment condition (ref) is zero due to independence of $\Gamma_i$ and $u_{it}$ by assumption and the sum-to-zero property of $\boldsymbol{\Gamma}$. For identification with large $N$, we assume size to follow a power law in tail (see (ref)), thus $\Gamma_i$ is stochastic and assumed to be independent of $u_{it}$.\footnote{For a fixed $N$, it is not required to assume independence of $\Gamma_i$ and $u_{it}$ because $\Gamma_i$ can be treated as constant and (ref) is zero solely by virtue of the fact that $\boldsymbol{\iota}'\boldsymbol{\Gamma}=0$.} So again, we have $\mathbbm{E}(z_tp_t)\neq 0$ and for this simplified example, we avoid the need to estimate the factor structure since (i) due to uniform loadings, $z_t$ is constructed from observables and (ii) $z_t$ is uncorrelated with the composite error term. If either of (i) or (ii) fails to hold, estimation of the factor structure becomes a preliminary step, as in our general procedure. Nevertheless, (ref) leads to the following simple supply elasticity estimator $\widehat{\phi}^s=\frac{\sum_ty_{Et}z_t}{\sum_tp_t z_t}$. The intuition here is that, again, $z_t$ places larger weights on the idiosyncratic shocks to larger oil producers, these granular shocks keep the simple average (or more generally precision-weighted, i.e., weighted heavily towards more stable oil producers) supply curve fixed. That is, on average, precision-weighted supply responds to these granular shocks only through their effects on prices (due to $\mathbbm{E}((\boldsymbol{\lambda}'\boldsymbol{\eta}_t+u_{Et})z_t)=0$) and at the same time since smaller oil producers take as given price changes caused by these granular shocks, it will shift their \color{black}supply \color{black} curves which enables consistent estimation of the supply elasticity.

\color{black}Discussion.\color{black} In the case of uniform loadings and $u_{it}$ $i.i.d.$, the vector $\boldsymbol{E}$ and the instrument are constructed from observables, the large sample properties of $\widehat{\phi}^s$ and $\widehat{\phi}^d$ only entail fixed $N$, large $T$ asymptotics for which GK have laid out. In general, however, the cross-section will need to be exploited to estimate $\boldsymbol{E}$ since one can not know if $u_{it}$ are $i.i.d.$ across $i$. Indeed, the factors typically take care of a substantial portion of the cross-sectional correlations but it is prudent to allow for cross correlations in $u_{it}$ since the exogeneity condition for estimation of the supply elasticity heavily exploits the structure of $\boldsymbol{\Sigma}_u$. Therefore, it will be important to generally allow for some weak cross correlations in $\boldsymbol{\Sigma}_u$, which our algorithm accommodates, as discussed in Section (ref) and Section (ref).

\color{black}Moreover, \color{black} although homogeneous loadings was only an abstraction to illustrate the instrument, GK advocate the use of $y_{\Gamma t}=y_{St}-\frac{1}{N}\sum_i y_{it}$ in practice even when the loadings are not uniform. In the general heterogeneous loadings case, their instrument becomes

align[align omitted — 115 chars of source]

They label this instrument with a capital case convention, to distinguish it because it is no longer solely composed of weighted idiosyncratic shocks, $u_{\Gamma t}$, as the $\boldsymbol{\lambda}_{\Gamma}'\boldsymbol{\eta}_t$ term is contaminating the instrument. However, this clever formulation is possible because they advocate estimation of the factors in practice, which they augment to their structural equations, thereby controlling for the second term which can potentially make their moment conditions different from zero.

Feasible Granular Instrumental Variables

Homogeneous loadings are overly restrictive but relaxing this can be easily accommodated in practice via PCA or iterative OLS-PCA methods, e.g., bai2003inferential or bai2009panel in a preliminary stage to construct an estimate of the instrument.\footnote{For our theory, we assume a balanced panel. However, in the case of unbalanced panels with data missing at random (which is beyond the scope of this paper) one can instead use the bai2015unbalanced method or bai2021matrix method to estimate the factor structure and the instrument. \color{black}In the more realistic case where data are not missing at random, one can use the methods developed in xiong2019large. Remark?\color{black}} Although in GK's asymptotic theory they assume homogeneous loadings and that the instrument is exogenous with respect to the composite error, which circumvents the need to estimate the factor structure, they indeed advocate augmenting their structural equations with estimated factors either via period-by-period cross sectional regressions when the loadings are known or via PCA in the case of non-parametric (unknown) loadings. GK abstract away from the sampling error in suggesting the use of augmented factors, which only vanishes for both large $N$ and $T$. bai2006confidence and greenaway2012asymptotic have developed the asymptotic distribution for structural parameters in factor augmented regressions in time series and panel models respectively. In this paper, a variant of their corresponding result is established in showing the sampling error from estimating the high dimensional precision matrix, the factors, as well as the instrument is negligible in the asymptotic distribution of the structural parameters.

The general heterogeneous loadings case and $u_{it}$ non-$\boldsymbol{i.i.d.}$ Now we formulate the estimation approach in the general case, which makes much heavier use of the cross-section. When we cross-sectionally demean the supply equation and stack across $i$ we obtain (recall $\widetilde{\boldsymbol{X}}$ denotes a generic demeaned variate)

align[align omitted — 155 chars of source]

which is estimable with vanilla PCA when the factor structure is strong.\footnote{Strong factors in the sense that $\boldsymbol{\Lambda}'\boldsymbol{\Lambda}/N \overset{p}{\rightarrow} \boldsymbol{\Sigma}_{\Lambda}>0$; thus we assume the factors are strong/pervasive in the sense that a significant fraction of cross-sectional units are affected by their presence. Consistent estimation of weak factors is beyond the scope of this paper, see for example onatski2012asymptotics, bailey2016exponent or freyaldenhoven2021factor for suitable conditions for which it is possible. Even when estimable, their convergence rates are slower relative to estimates of strong factors, e.g., see bai2021approximate. This will generally require modifications to the limiting distributions we derive in this paper.} Letting $\boldsymbol{Q}=(\boldsymbol{I}_N-\boldsymbol{\widetilde{\Lambda}(\widetilde{\Lambda}'\widetilde{\Lambda})^{-1}\widetilde{\Lambda}}')$, then $\boldsymbol{Q}\boldsymbol{\widetilde{y}}_{\cdot t}=\boldsymbol{Q}\boldsymbol{\widetilde{u}}_{\cdot t}$, completely purges the process of the common factors through the loading space. Premultiplying the share weights gives the instrument

align[align omitted — 232 chars of source]

where $\boldsymbol{\Gamma}:= \boldsymbol{QS}$ is unknown because $\boldsymbol{Q}$ is unknown, but $\boldsymbol{Q}$ is easily estimated from data. Once we have $\boldsymbol{\widehat{Q}}$, \color{black}which just replaces $\widetilde{\Lambda}$ with $\widehat{\widetilde{\Lambda}}$ \color{black}, we form $\widehat{z}_t=\boldsymbol{S}'\boldsymbol{\widehat{Q}}\boldsymbol{\widetilde{y}}_{\cdot t}$ from observables. Importantly, when $\boldsymbol{\lambda}_i=\boldsymbol{\lambda} \, \forall i$, then $\boldsymbol{\Gamma}=(\boldsymbol{I}_N-\boldsymbol{\widetilde{\Lambda}(\widetilde{\Lambda}'\widetilde{\Lambda})^{-1}\widetilde{\Lambda}}')\boldsymbol{S}=\boldsymbol{S}-\boldsymbol{\iota} /N$ as in the previous case with homogenous loadings. This gives rise to a more general demand elasticity estimator

align[align omitted — 157 chars of source]

In Section (ref), we show that the demand elasticity can be estimated as if the infeasible instrument, $z_t$, is used.

In the case of the supply elasticity, the estimator will additionally depend on the estimated (potentially high dimensional) precision matrix. That is, $\widehat{\phi}^s=\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{\widehat{\Sigma}}^{-1}_u)$. This creates the need to jointly estimate $\boldsymbol{\widehat{\Sigma}}^{-1}_u$ to form $\boldsymbol{\widehat{E}}$ in order to aggregate the panel to estimate $\widehat{\phi}^s$. We propose a simple iterative procedure and show that the supply elasticity can be estimated as if the infeasible precision matrix, $\boldsymbol{\Sigma}_u^{-1}$, and instrument, $z_t$, were used. More specifically, let $y_{Et} =\phi^s p_t +\boldsymbol{\lambda}_E'\boldsymbol{\eta}_t +u_{Et} :=\boldsymbol{f}_t' \boldsymbol{\theta}^s + u_{Et},$ where $\boldsymbol{\theta}^s=

pmatrix[pmatrix omitted — 45 chars of source]

'$ and $\boldsymbol{f}_t=

pmatrix[pmatrix omitted — 39 chars of source]

'$ are $(1+r) \times 1$ vectors. The remarkable result $\mathbbm{E}(z_t u_{Et})=0$, shown in \eqref{babymagic} for the previous simple example with homogeneous loadings, continues to hold in this setting as well, with $z_t=\boldsymbol{S}'\boldsymbol{Q}\boldsymbol{\tilde{y}_{\cdot t}}=\boldsymbol{S}'\boldsymbol{Q}\boldsymbol{\tilde{u}_{\cdot t}}$ and $\boldsymbol{\Gamma}=\boldsymbol{Q}\boldsymbol{S}$ (recall that $\boldsymbol{\iota}'\boldsymbol{\Gamma}=0$)

align*[align* omitted — 823 chars of source]

So we have that (where the estimated factors self-instrument)

align[align omitted — 288 chars of source]

However, given our interest lies in inference for $\phi^s$, it is useful to stack over $t$, $\boldsymbol{y}_{E} = \boldsymbol{p} \,\phi^s + \boldsymbol{\eta} \,\boldsymbol{\lambda}_E + \boldsymbol{u}_{E},$ where $\boldsymbol{y}_{E}, \boldsymbol{p},$ and $\boldsymbol{u}_{E}$ are $T \times1$ vectors and $\boldsymbol{y}_{\widehat{E}}$ is the feasible counterpart of $\boldsymbol{y}_{E}$. Let $\boldsymbol{M}_{\widehat{\eta}}=(\boldsymbol{I}_T-\boldsymbol{\widehat{\eta}(\widehat{\eta}'\widehat{\eta})^{-1}\widehat{\eta}}')$, then it follows from standard partitioned regression results that

align[align omitted — 313 chars of source]

As $\boldsymbol{\widehat{\Sigma}}_u^{-1}$ depends on $\widehat{\phi}^s$, (ref) generally requires an iterative estimation procedure. To that end, note that if $\phi^s$ were known, $y_{it}-p_t\phi^s=\boldsymbol{\lambda}_i'\boldsymbol{\eta}_t+u_{it}$ follows an approximate factor structure. Thus, a covariance estimator, $\boldsymbol{\widehat{\Sigma}}_u$, for the idiosyncratic part can be obtained following fan2013large by applying thresholding to the eigenvalue decomposition, $\frac{1}{T}\sum_{t=1}^T(\boldsymbol{y}_{\cdot t}-\boldsymbol{\iota}p_t\phi^s)(\boldsymbol{y}_{\cdot t}-\boldsymbol{\iota}p_t\phi^s)'=\sum_{i=1}^N \gamma_i \boldsymbol{\xi}_i\boldsymbol{\xi}_i',$ where $\gamma_i$ and $\boldsymbol{\xi}_i$ are the eigenvalues (sorted in decreasing order) and corresponding eigenvectors, respectively. More specifically, if $\phi^s$ were known, we have

align[align omitted — 255 chars of source]

where $\boldsymbol{\widehat{\Sigma}}_u^{\mathbfcal{{T}}}=\sum_{i=r+1}^N \widehat{\gamma}_i \boldsymbol{\widehat{\xi}}_i\boldsymbol{\widehat{\xi}}_i'=(\widehat{\sigma}^{\mathcal{T}}_{u,ij})_{N \times N}$,

align[align omitted — 169 chars of source]

and $h_{ij}(\cdot)$ is a generalized shrinkage function of antoniadis2001regularization.\footnote{Examples of $h_{ij}(\cdot)$ include hard thresholding $h_{ij}(x)=x\mathbbm{1}(|x|\geq \tau_{ij})$ and soft thresholding $h_{ij}(x)=\text{sgn}(x)(|x| - \tau_{ij})_{+}$. The entry dependent threshold, $\tau_{ij}>0$, can be defined as $C\omega_T \sqrt{\widehat{\alpha}_{ij}}$, where $\widehat{\alpha}_{ij}=\frac{1}{T}\sum_{t=1}^T(\widehat{u}_{it}\widehat{u}_{jt}-\widehat{\sigma}_{u,ij})^2$, $\widehat{\sigma}_{u,ij}=\frac{1}{T}\sum_{t=1}^T\widehat{u}_{it}\widehat{u}_{jt}$ and $\widehat{u}_{it}=y_{it}-\phi^sp_t-\boldsymbol{\widehat{\lambda}}_i'\boldsymbol{\widehat{\eta}}_t$ for some predetermined decreasing sequence $\omega_T>0$ and $C>0$. The choice of $C$ can be data driven; fan2013large choose $C$ through multifold cross-validation to maintain positive definiteness of $\boldsymbol{\widehat{\Sigma}}_u^{\mathbfcal{{T}}}(C)$. In our algorithm below, we make use of the R package for POET, written by the authors fan2013large.} Of course, $\phi^s$ can not be known as it requires an estimate of $\boldsymbol{\Sigma}_u^{-1}$. Thus, we now address joint estimation of $\phi^s$ and $\boldsymbol{\Sigma}_u^{-1}$ in what follows and subsequently establish that the sampling error in $\boldsymbol{\widehat{E}}$ is negligible given some regularity conditions. The iterative procedure is summarized in Algorithm (ref) presented below.

center[center omitted — 1,162 chars of source]

When $r$ is unknown, one can augment Step 1 and estimate $r$ using a procedure as in bai2002determining, onatski2010determining or ahn2013eigenvalue; we use the $ER$ and $GR$ methods of ahn2013eigenvalue (hereafter AH). For more details of the $ER$ and $GR$ methods, see Section (ref) of the Supplementary Appendix.

FGIV algorithm accommodating cross-section specific covariates. When $k_x \neq 0$ then the demeaning transformation from (ref) results in $\boldsymbol{\widetilde{y}}_{\cdot t}=\boldsymbol{\widetilde{\Lambda}}\boldsymbol{\eta}_t+\boldsymbol{\widetilde{x}}_{\cdot t}\boldsymbol{\beta}+\boldsymbol{\widetilde{u}}_{\cdot t},$ where $\boldsymbol{\widetilde{x}}_{\cdot t}$ is an $N \times k_x$ matrix, which leaves $\underset{k_x \times 1}{\boldsymbol{\beta}}$ as an additional parameter to estimate. $\boldsymbol{\beta}$ can be easily estimated by adapting the procedure of bai2017inferences, which is generalizing bai2009panel, to handle endogeneity of prices even after controlling for latent common factors. More specifically,

align[align omitted — 630 chars of source]

since (ref) follows a factor structure, the $T \times r$ factor matrix, $\boldsymbol{\eta}(\boldsymbol{\beta},\boldsymbol{\Sigma}_u^{-1})$, can be estimated using the principal components estimator whose columns are the eigenvectors corresponding to the largest $r$ eigenvalues of the $T \times T$ matrix $(\boldsymbol{\widetilde{y}}_{\cdot \cdot}-\boldsymbol{\widetilde{x}}_{\cdot \cdot}(\boldsymbol{\beta}))\boldsymbol{\Sigma}_u^{-1}(\boldsymbol{\widetilde{y}}_{\cdot \cdot}-\boldsymbol{\widetilde{x}}_{\cdot \cdot}(\boldsymbol{\beta}))'$, where the $T \times N$ matrix $\boldsymbol{\widetilde{x}}_{\cdot \cdot}(\boldsymbol{\beta}):=

pmatrix[pmatrix omitted — 133 chars of source]

$ and $\boldsymbol{\widetilde{\Lambda}}(\boldsymbol{\beta},\boldsymbol{\Sigma}_u^{-1})=\frac{1}{T}\sum_{t=1}^T(\boldsymbol{\widetilde{y}}_{\cdot t}-\boldsymbol{\widetilde{x}}_{\cdot t}\boldsymbol{\beta})\boldsymbol{\eta}_t' (\boldsymbol{\beta},\boldsymbol{\Sigma}_u^{-1}).$ Thus, to deal with general (strictly exogenous) covariates, $\boldsymbol{x}_{it}$, Algorithm (ref) can be applied.

center[center omitted — 2,064 chars of source]

When $r$ is unknown, one can augment Step 2 and iteratively estimate $r$ using the $ER$ and $GR$ methods of ahn2013eigenvalue.

The main takeaway is that when both $(N,T)$ are large, one can generalize the GIV estimators proposed by GK along different dimensions; here we accommodate latent heterogeneous loadings, latent factors and latent precision matrix (e.g., $u_{it}$ can be weakly cross-correlated and heteroskedastic). As mentioned earlier, we call the proposed estimators of the elasticities in (ref), Algorithm (ref) and Algorithm (ref) as FGIV estimators.

assumpRIn principle, the theory for the estimators proposed in this paper allows for $N \gg T$. This case is relevant in many empirical settings (e.g., empirical industrial organization and finance). However, it may be beneficial to avoid estimating the precision matrix for cases where $N \ll T$ (e.g., empirical macro). But, as (ref) and (ref) show, to have a valid instrument for which the moment equation is exactly zero, we must specify $\boldsymbol{\Sigma}_u$ correctly. This is the primary motivation for estimating the general precision matrix in Algorithms (ref) and (ref). In order to avoid estimating the precision matrix, we must assume (potentially erroneously) $u_{it}$ are cross-sectionally independent. We now analyze the consequences of making this assumption when in fact $u_{it}$ are cross-sectionally correlated. Suppose we erroneously assume cross-sectional independence, then the vector $\boldsymbol{E}$ reduces to $\boldsymbol{\iota}/N$ and we end up with the following moment equation \begin{align*} \mathbbm{E}(u_{Et}z_t )&=\mathbbm{E}\left(\boldsymbol{E}'\boldsymbol{u}_{\cdot t} \boldsymbol{\tilde{u}}_{\cdot t}'\boldsymbol{\Gamma}\right)=\mathbbm{E}(\boldsymbol{E}'\boldsymbol{u}_{\cdot t}(\boldsymbol{u}_{\cdot t}-\bar{u}_{t}\boldsymbol{\iota})'\boldsymbol{\Gamma}) \\ &=\mathbbm{E}(\boldsymbol{E}'\mathbbm{E}(\boldsymbol{u}_{\cdot t}\boldsymbol{u}_{\cdot t}'|\boldsymbol{\Gamma})\boldsymbol{\Gamma}) -\mathbbm{E}(\boldsymbol{E}'\boldsymbol{u}_{\cdot t} \boldsymbol{\iota}'\boldsymbol{\Gamma})\bar{u}_{t} =\frac{1}{N}\mathbbm{E}(\boldsymbol{\iota}'\boldsymbol{\Sigma}_u\boldsymbol{\Gamma})-0= o(1). \addtocounter{equation}{1}\tag{\theequation} \end{align*} Hence, $z_t$ is not a valid instrument in the traditional sense because we allow $\mathbbm{E}(u_{Et}z_t )\neq 0$ for any given sample. Nevertheless, this moment converges to zero for large $N$. Indeed, the moment satisfies $\mathbbm{E}(u_{Et}z_t )=o\left(1\right)$ under our regularity assumptions, and thus, $z_t$ is asymptotically a valid instrument.\footnote{It can be shown that $\boldsymbol{\iota}'\boldsymbol{\Sigma}_u\boldsymbol{\Gamma} =\boldsymbol{\iota}'\boldsymbol{\Sigma}_u\boldsymbol{QS} \leq \boldsymbol{\iota}'\boldsymbol{\Sigma}_u\boldsymbol{S} \gamma_{max}(\boldsymbol{Q}) =\sum_{i,j}\sigma_{u,ij} S_j \leq \left(\sum_{i,j} \sigma^2_{u,ij} \right)^{1/2} ||\boldsymbol{S}||^2_2 = || \boldsymbol{\Sigma}_u ||_{F} || \boldsymbol{S} ||^2_2 \leq || \boldsymbol{\Sigma}_u ||_{1} || \boldsymbol{S} ||^2_2 \leq \mathcal{O}(m_N) \Theta_p(1) = o(N)\Theta_p(1)=o(N)$, where $m_N$ is defined in (ref) and $m_N=o(N)$.} This insight reveals that this moment is approaching zero, hence it may prove to be beneficial to aggregate the panel, $y_{it}$, using weights $\boldsymbol{\iota}/N$ regardless of the covariance structure. The immediate implication is that $\widehat{\phi}^s=\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{I}_N)$, so there is no need for an algorithmic estimation procedure, the simple analytical formula for the supply elasticity estimator with potentially misspecified covariance structure for $u_{it}$ is given by $\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{I}_N) = \frac{\boldsymbol{\widehat{z}}^{\,'}\,\boldsymbol{M}_{\widehat{\eta}}\,\boldsymbol{\bar{y}}}{\boldsymbol{\widehat{z}}^{\,'} \,\boldsymbol{M}_{\widehat{\eta}}\,\boldsymbol{p}},$ where $\boldsymbol{\bar{y}}$ stacks $\bar{y}_t = \frac{1}{N}\sum_{i=1}^N y_{it}$ for each $t=1,\dots, T$; this estimator is essentially Step 2 and Step 3 of Algorithm (ref). Asymptotically, it holds that $\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{\widehat{\Sigma}}_u^{-1}) = \widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{I}_N) + o_p(1).$ However, regarding performance in finite samples, when $u_{it}$ are not $i.i.d.$ and when $N \ll T$, it is not clear ex-ante if $\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{I}_N)$ will outperform $\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{\widehat{\Sigma}}_u^{-1})$. When $N \gg T$ one would expect ex-ante that $\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{I}_N)$ will be less efficient than $\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{\widehat{\Sigma}}_u^{-1})$ since the former is not optimally weighting the observations, whereas the latter is. When $u_{it}$ are indeed $i.i.d.$ we would expect $\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{I}_N)$ to perform better.\footnote{In unreported simulations where $N \ll T$ and $u_{it}$ are non-i.i.d., we find that $\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{\widehat{\Sigma}}_u^{-1})$ typically has a smaller bias than $\widehat{\phi}^s(\boldsymbol{\widehat{z}},\boldsymbol{I}_N)$ (in absolute terms, the bias of both estimators are very small) but with a slightly larger variance.}

Efficient GMM Estimation: Factor-Augmented FGIV

We now proceed to overidentify the elasticities, which yields overidentified FGIV estimators. We will refer to the overidentified FGIV estimators simply as efficient GMM estimators and the just identified FGIV estimators simply as FGIV estimators. It will be of interest to practitioners to see if overidentification is possible for the supply and demand equations. In this section, we show that the system is indeed overidentified to varying degrees for the supply and demand equations.

Demand. It is common practice to assume uncorrelated aggregate supply and aggregate demand shocks, that is $\mathbbm{E}(\boldsymbol{\eta}_t \varepsilon_t)=0$. When we are willing to entertain this, then our supply factors, estimated via principal components, serve as valid instruments in estimation of the demand elasticity, rendering an overidentified parameter. In fact, the theory for using principal components as instruments was laid out in bai2010instrumental under strong instrument asymptotics, as well as kapetanios2010factor under many/weak instrument asymptotics. In the remainder of this section, we let the GIV be denoted as $z_{t,GIV}:= z_t$ to distinguish it from the full instrument vector we introduce with upper case conventions. Our full instrument matrix for the demand equation is $\underset{T \times (1+r)}{\boldsymbol{Z}_d} :=

pmatrix[pmatrix omitted — 55 chars of source]

$ with $\mathbbm{E}(\boldsymbol{Z}_{dt}\varepsilon_t)=0$; $\boldsymbol{Z}_{dt}$ simply augments factors to be used as instruments. Making use of the $(1 + r) \times 1$ dimensional moment condition, the efficient GMM demand elasticity estimator is defined as

align[align omitted — 518 chars of source]

where $\underset{ (1+r) \times (1+r)}{\boldsymbol{W}_d}$ is an arbitrary positive definite weight matrix, but is optimally set as $\boldsymbol{\widehat{W}}_d=\boldsymbol{\widehat{\Omega}}_d^{-1}$, where $\boldsymbol{\widehat{\Omega}}_d=\frac{1}{T}\sum_{t=1}^T \boldsymbol{\widehat{Z}}_{dt}\boldsymbol{\widehat{Z}}_{dt}'(d_t-p_t\widehat{\phi}^d_{2SLS})^2$. It is clear that (ref) nests the FGIV estimator for the demand elasticity as a special case. In this sense, $\widehat{\phi}^d_{GMM}$ will be robust to scenarios where $ z_t$ is weaker.

Supply. In the same vein, the supply elasticity can always be overidentified given our identifying assumptions because $\mathbbm{E}(\varepsilon_t u_{Et})=0$ and thus $\varepsilon_t$ can serve as an additional instrument. To estimate the entire parameter vector for the supply equation, let $\underset{T \times (2+r)}{\boldsymbol{Z}_s}:=

pmatrix[pmatrix omitted — 82 chars of source]

$, where the augmented factors self-instrument as they are part of the supply equation. Then $\boldsymbol{y}_E=\boldsymbol{f}\boldsymbol{\theta}^s +\boldsymbol{u}_E$ and recall $\boldsymbol{\theta}^s=

pmatrix[pmatrix omitted — 45 chars of source]

'$ and $\boldsymbol{f}_t=

pmatrix[pmatrix omitted — 39 chars of source]

'$ are $(1+r) \times 1$ vectors and the matrix $\boldsymbol{f}$ is $T \times (1+r)$, which stacks $\boldsymbol{f}_t$. We have $\mathbbm{E}(\boldsymbol{Z}_{st} \boldsymbol{u}_{Et})=0$; hence, making use of the $(2+r)\times 1$ dimensional moment conditions, the efficient GMM supply elasticity estimator is defined as

align[align omitted — 567 chars of source]

where $\underset{ (2+r) \times (2+r)}{\boldsymbol{W}_s}$ is an arbitrary positive definite weight matrix, but is also optimally set as $\boldsymbol{\widehat{W}}_s=\boldsymbol{\widehat{\Omega}}_s^{-1}$, where $\boldsymbol{\widehat{\Omega}}_s=\frac{1}{T}\sum_{t=1}^T \boldsymbol{\widehat{Z}}_{st}\boldsymbol{\widehat{Z}}_{st}'(y_{\widehat{E}t}-\widehat{\boldsymbol{f}}'_t\boldsymbol{\widehat{\theta}}^s_{GMM})^2$.\footnote{In the case of the demand elasticity estimator in (ref) we use 2SLS residuals to construct $\boldsymbol{\widehat{\Omega}}_d$. However, we implement (ref) via Algorithm (ref) which, by iteration, renders the residuals used to construct $\boldsymbol{\widehat{\Omega}}_s$ to be GMM residuals.} It is clear that (ref) nests the FGIV estimator for the supply equation as a special case. As in the just identified case in (ref), $\widehat{\boldsymbol{\theta}}^s_{GMM}$ in (ref) depends on $\widehat{\boldsymbol{\Sigma}}_u^{-1}$, hence, will generally require an iterative estimation procedure. Algorithm (ref) below generalizes Algorithm (ref) by extending the joint estimation of the supply elasticity estimator and the precision matrix to the overidentified case for when $k_x=0$. In view of Algorithm (ref), Algorithm (ref) can be further extended to the case when $k_x > 0$, but we omit the details for brevity.

center[center omitted — 1,752 chars of source]

In addition to efficiency gains, the efficient GMM estimators exhibit superior finite sample properties and are also robust to the GIV itself being a weak instrument. We illustrate these points in greater detail in (ref) and Section (ref).

The intuition for the overidentified estimators can be seen from observing the reduced form equation for (equilibrium) prices, $p_t=\frac{1}{\phi^d - \phi^s} \left(u_{St} +\boldsymbol{\lambda}_S' \boldsymbol{\eta}_t -\varepsilon_t \right)$. Clearly $\mathbbm{E}(p_t\boldsymbol{\eta}_t) \neq 0$ and $\mathbbm{E}(p_t\varepsilon_t) \neq 0$ and so instrumental relevancy is established. Thus, we are effectively back to the classical approach of finding exogenous supply shifters, in this case $\varepsilon_t$, to estimate the supply elasticity and finding exogenous demand shifters, in this case $\boldsymbol{\eta}_t$, to estimate the demand elasticity. With the exception that these shifters, $\boldsymbol{\eta}_t$ and $\varepsilon_t$ are unobserved. In what follows, we show that estimating $\boldsymbol{\eta}_t$ and $\varepsilon_t$ has a negligible effect on the limiting distributions of the estimators of demand and supply elasticities, respectively.

Assumptions

Below we lay out the assumptions needed to derive our main results. (ref), (ref) and (ref) are standard in the literature; see, for example, bai2003inferential, fan2013large and bai2017inferences, \color{black}but are relevant for a thorough understanding of the subsequent theorems\color{black}. Whereas, (ref) parts ii.) and iii.) are new so we provide more details.

assump[Factor Error Structure] The composite error term in (ref) is assumed to admit an (approximate) factor structure representation $v_{it} := \boldsymbol{\lambda}'_i \boldsymbol{\eta}_t + u_{it},$ where $\boldsymbol{\eta}_t=\begin{pmatrix}\eta_{1t} & \dots & \eta_{rt}\end{pmatrix}'$ is an $r \times 1$ vector of latent common factors and $\boldsymbol{\lambda}_i=\begin{pmatrix} \lambda_{1i} & \dots & \lambda_{ri}\end{pmatrix}'$ is an $r \times 1$ vector of latent factor loadings. We assume the factors are pervasive in the sense that $\boldsymbol{\Lambda}'\boldsymbol{\Lambda}/N$ converges to some $r \times r$ positive definite matrix.
assump(Strict Stationarity, Exponential Tails & Strong Mixing) \newline $\textbf{(A2i.)} \,\, \{\boldsymbol{\eta}_t, u_{i t}, \varepsilon_t\}_{t \geq 1}$ is strictly stationary and each with a zero mean. \newline $\textbf{(A2ii.)} \,\,\exists \,\, c_1, c_2 > 0$ with $\gamma_{\emph{min}}(\boldsymbol{\Sigma}_u) > c_2$, $\underset{j \leq N}{\emph{max}} || \gamma_j || < c_1$, $c_2 < \gamma_{\emph{min}}(\mathbbm{cov}(\boldsymbol{\eta}_t)) \leq \gamma_{\emph{max}}(\mathbbm{cov}(\boldsymbol{\eta}_t)) < c_1.$\newline $\textbf{(A2iii.)} \,\,$Exponential tail: $\exists \,\, r_1, r_2 > 0$ and $b_1, b_2 >0$, such that for any $s >0$, $i \leq N$ and $j \leq r$, $\mathbbm{P}(|u_{it}|>s) \leq \emph{exp}(-(s/b_1)^{r_1})),$ and $\mathbbm{P}(|\eta_{t,j}|>s) \leq \emph{exp}(-(s/b_2)^{r_2})).$ \newline $\textbf{(A2iv.)} \,\,$Strong Mixing: $\exists \,\, r_3, \,C > 0$ $\forall \,\, T >0, \,\, r_1^{-1} + r_{2}^{-1} + r_{3}^{-1} > 1$, $\underset{A \in \mathcal{F}_{-\infty}^0, \,\, B \in \mathcal{F}_{T}^{\infty}}{\emph{sup}} |\mathbbm{P}(A)\mathbbm{P}(B)-\mathbbm{P}(AB)| < \emph{exp}(-CT^{r_3}),$ where $\mathcal{F}_{-\infty}^0$ and $\mathcal{F}_{T}^{\infty}$ denote the $\sigma$-algebras generated by $\{(\boldsymbol{\eta}_t, u_{it}, \varepsilon_t) : t<0\}$ and $\{(\boldsymbol{\eta}_t, u_{it}, \varepsilon_t) : t>T\}$ respectively.
assump[Sparsity on $\boldsymbol{\Sigma}_u$] Let $\boldsymbol{\Sigma}_u =(\sigma_{u,ij})$, for some $q \in [0,1/2),$ define \begin{align} m_N=\underset{i \leq N}{max}\sum_{j=1}^N |\sigma_{u,ij}|^q. \end{align} We require that there is $q \in [0,1/2)$ such that $m_N \omega_{N,T}^{1-q}=o(1)$, where $\omega_{N,T}=\sqrt{\frac{\emph{log}(N)}{T}} +\frac{1}{\sqrt{N}}$. \newline

\setcounter{assumpB}{3}

assumpB(Identification by GIV)\newline $\textbf{(A4i.)} \,\,\, \mathbbm{E} (z_t u_{Et})=\mathbbm{E} (z_t \varepsilon_t) =\mathbbm{E} (\boldsymbol{Z}_{st} u_{Et})=\mathbbm{E} (\boldsymbol{Z}_{dt} \varepsilon_t) = 0$. \newline (A4ii.) The sizes $\mathscr{S}_1,\dots,\mathscr{S}_N$ are drawn $i.i.d.$ from an arbitrary distribution for which the tail of the size distribution (i.e. above some threshold) follows a power law, with tail index, $\mu>0$ \begin{align*} \mathbbm{P}(\mathscr{S}>s) = cs^{-{\mu}}. \nonumber \end{align*} The tail index $\mu$ determines the probability of observing extreme values. We assume that $\mathscr{S}_i$ is independent of $u_{it}$. \newline (A4iii.) Suppose the sizes are ordered in decreasing fashion as such: $\mathscr{S}_{(1)} \geq \mathscr{S}_{(2)} \geq \dots \geq \mathscr{S}_{(N-1)} \geq \mathscr{S}_{(N)} ,$ and we partition the cross-section as, $\mathcal{N}_{dominant}:= \{1,\dots,N_1\}$ and $\mathcal{N}_{fringe}:= \{N_1+1,\dots,N\}$ such that $\mathcal{N}_{dominant}\cup \mathcal{N}_{fringe}:=\mathcal{N}_{full}$. Let $S_i=\frac{\mathscr{S}_i}{\sum_{j=1}^N\mathscr{S}_j}$ denote the normalized shares such that $\sum_i S_i =1$. We assume $\forall \, i \,\in \mathcal{N}_{dominant}$, $\,S_i=\Theta_p(1)$, and $\forall \, i \in \mathcal{N}_{fringe},$ $S_{i}=\mathcal{O}_p\left(\frac{1}{N}\right)$. We further assume that the cardinality of the dominant units is fixed as $N \rightarrow \infty$, that is, $|\mathcal{N}_{dominant}|=N_1$ and $N_1$ does not rise with $N$ while the cardinality of the fringe grows with $N$, $|\mathcal{N}_{fringe}|=N-N_1 \rightarrow \infty$ as $N \rightarrow \infty$. \newline
assumpRThe first condition gives us instrumental exogeneity for the FGIV and efficient GMM estimators. The second condition allows for instrumental relevance in the extension of a large N framework. An important implication of the second condition is that the Herfindahl index, $h_{N,\mu}$, has the following asymptotic property \begin{align*} \sqrt{h_{N,\mu}}=\left|\big| \boldsymbol{S} \big|\right|_2 = \begin{cases} \Theta_p\left( 1\right) for \mu \in (0,1), \\ \mathcal{O}_p\left( g_{N,\mu}\right) for \mu \in [1,2), \end{cases} \end{align*} with $\mathcal{O}_p\left( g_{N,\mu}\right) \gg1/\sqrt{N}$. The variance of the just identified estimators is inversely proportional to the Herfindahl index, that is $\mathbbm{V}(\widehat{\phi}^j_{FGIV})=\mathcal{O}(h_{N,\mu}^{-1})$ for $j=s,d$, reflecting the fact that the more concentrated the market, the more precise the GIV methodology will be and also reflecting the fact that if the Herfindahl converges to zero in the limit, the variance will diverge.\footnote{The derivation of the asymptotic behavior of $h_{N,\mu}$ can be found in Supplementary Appendix (ref).} However, if $\mu$ is slightly greater than 1, theoretically identification breaks down for large $N$ but in any finite sample the GIV could be relevant (precisely due to $\mathcal{O}_p\left( g_{N,\mu}\right) \gg1\sqrt{N}$). Nevertheless, we rule this case out for the purpose of asymptotic inference.\footnote{For more details on instrumental relevance for large $N$, see Section (ref).} Note, that the third condition is consistent with $\mu \in (0,1)$, but is slightly stronger. The third condition is also a generalization of the so-called "granular" weights in the panel data literature, say $\underset{{N \times 1}}{\boldsymbol{w}}$, which are typically assumed to satisfy $||\boldsymbol{w}||_2=\mathcal{O}\left(\frac{1}{\sqrt{N}}\right)$ and $\frac{w_i}{||\boldsymbol{w}||_2}=\mathcal{O}\left(\frac{1}{\sqrt{N}}\right) \,\, \forall \, i$. The third condition allows the share vector to be partitioned into a dominant part and a fringe part. That is, $\boldsymbol{S}=\begin{pmatrix} \boldsymbol{S}'_{d} & \boldsymbol{S}'_f \end{pmatrix}'$ where $\boldsymbol{S}_{d}$ is $N_1 \times 1$, is the dominant part and $\boldsymbol{S}_f$ is $N_2 \times 1$, is the fringe part; with $N_1+N_2=N$, the key being that $N_1(N)=N_1$ is fixed while $N_2(N) \rightarrow \infty$ as $N \rightarrow \infty$. This assumption can be empirically justified in concentrated markets, see Section (ref) as an example; as well as mathematically justified, see logan1973limit.
assumpRTaking the variance of the equilibrium price process (assuming the covariances to be zero for simplicity) we obtain $\mathbbm{V}(p_t) =\frac{1}{(\phi^d-\phi^s)^2}(\mathbbm{V}(u_{St})+\mathbbm{V}(\boldsymbol{\lambda}_S'\boldsymbol{\eta}_t)+\mathbbm{V}(\varepsilon_t))= \Theta(1),$ where the last equality follows by the second and the third conditions in (ref), details can be found in Lemma (ref) in the Appendix. Without these conditions, one would obtain the unsatisfactory result that $\mathbbm{V}(p_t)=\mathcal{O}(N)$, that is, the variance of the price process is unbounded for each $t$ as $N\rightarrow \infty$. Effectively, (ref) allows the coexistence of a finite number of dominant units, in terms of size, whose cardinality can not grow with $N$, while at the same time allowing for a bounded variance for the aggregate endogenous variable $p_t$.

Limiting Distributions

In this section, we first present the limiting distributions of the FGIV elasticity estimators, corresponding to (ref) and (ref) with Algorithm (ref). We then move on to the limiting distributions of the efficient GMM elasticity estimators, corresponding to (ref) and (ref) with Algorithm (ref).

Just identified demand elasticity. The just identified demand elasticity estimator in (ref) is given by

align*[align* omitted — 256 chars of source]

Hence,

align*[align* omitted — 307 chars of source]

From above, it is apparent we need to show $\frac{1}{T}\sum_{t=1}^T(\widehat{z}_t-z_t)\varepsilon_t =\frac{1}{T}\sum_{t=1}^T \boldsymbol{S}'(\boldsymbol{\widehat{Q}}-\boldsymbol{Q})\boldsymbol{\tilde{y}_{\cdot t}} \varepsilon_t=o_p(1)$ and $\frac{1}{T}\sum_{t=1}^T(\widehat{z}_t-z_t)p_t =\frac{1}{T}\sum_{t=1}^T \boldsymbol{S}'(\boldsymbol{\widehat{Q}}-\boldsymbol{Q})\boldsymbol{\tilde{y}}_{\cdot t} p_t =o_p(1)$. Indeed, we show in Lemma (ref), in the Appendix, that

align[align omitted — 598 chars of source]

where $C_{NT} := \text{min} \{\sqrt{N},\sqrt{T}\}$. The terms in (ref) and (ref) are $o_p(1)$ when $N<T$ without any restrictions; however, when $T < N$ we require the mild restriction that $N/T^2 \rightarrow 0$, i.e., if $T<N$, $T^2$ does not grow too slowly relative to $N$. Thus, making use of (ref) and (ref) we obtain

align[align omitted — 228 chars of source]

The order of the sampling error generally relies, in part, on the order of the Herfindahl. The order of the Herfindahl, in turn, critically depends on $\mu$, the tail index of the size distribution (\color{black} see (ref)\color{black}). Results on the order of the Herfindahl as a function of the tail index parameter $\mu$ entails a total of six possible cases. The results can be found in Table (ref) of Supplementary Appendix (ref). However, for inference, we require $\mu \in (0,1)$ (regularly varying tails) or $\mu \rightarrow 0$ (slowly varying tails) as discussed in detail in the previous section's remarks. Given this, even after pinning down the order of the Herfindahl, the panel dimensions can distinguish more cases as seen above. Nevertheless, as (ref), (ref) and (ref) indicate, for consistency we have the following result: \setcounter{theorem}{0}

theorem[Consistency of $\widehat{\phi}^d$] Under Assumptions 1-4, as $(N,T) \overset{j}{\rightarrow} \infty$, we have that when $N\geq T$ and $N/T^2 \rightarrow 0$ or when $N<T$ \begin{align} \widehat{\phi}^d-\phi^d \overset{p}{\rightarrow} 0. \end{align}

All proofs are deferred tot the appendix. Now, multiplying (ref) by $\sqrt{T}$

equation[equation omitted — 361 chars of source]

We can state the following result for the limiting distribution:

theorem[Limiting distribution for $\widehat{\phi}^d$] Under Assumptions 1-4 as $(N,T) \overset{j}{\rightarrow} \infty$, we have that when $N\geq T$, $N/T^{3/2} \rightarrow 0$ and $\sqrt{T}/N \rightarrow 0$; or when $N<T$ only $\sqrt{T}/N \rightarrow 0$ \begin{align} \sqrt{T} (\widehat{\phi}^d-\phi^d) \overset{d}{\rightarrow} \mathcal{N}\left(0, \mathbbm{v}_d\right), \end{align} where $\mathbbm{v}_d :=\mathbbm{m}_{zp}^{-2}\,\mathbbm{v}_{z\varepsilon}$, $\mathbbm{v}_{z\varepsilon}:=\mathbbm{E}(z^2_t\varepsilon^2_t)$ and $\mathbbm{m}_{zp} :=\mathbbm{E}(z_tp_t)$.

$\mathbbm{v}_{z\varepsilon}$ can be consistently estimated with

align[align omitted — 446 chars of source]

where HC and HAC denote heteroskedasticity-consistent and heteroskedasticity and autocorrelation consistent estimators, respectively. Hence $\frac{\sqrt{T}(\widehat{\phi}^d-\phi^d)}{\widehat{\mathbbm{v}}_d^{1/2}} \sim \text{t}_{df} \overset{d}{\rightarrow} \mathcal{N}(0,1)$, where $\widehat{\mathbbm{v}}_d^{1/2}=\widehat{\mathbbm{m}}_{zp}^{-2}\widehat{\mathbbm{v}}_{z\varepsilon}^{1/2}$, with $\widehat{\mathbbm{m}}_{zp}=T^{-1}\sum_{t=1}^T \widehat{z}_t p_t$ also consistent for $\mathbbm{m}_{zp}$. We will see in Section (ref) that the asymptotic theory provides good approximations to the finite sample distribution.

assumpRAs in GK, we express $\mathbbm{v}_d$ as inversely related to the Herfindahl, $h_{N,\mu}$, as claimed in (ref), for insights on the role of market concentration on precision of the GIV. Assuming conditional homoskedasticity of $\varepsilon_t$ and homoskedasticity of $u_{it}$, we have that \begin{align*} \mathbbm{v}_{z\varepsilon} &=\mathbbm{E}(z_t^2\varepsilon_t^2)=\sigma^2_{\varepsilon}\cdot\sigma^2_{\tilde{u}} \cdot \mathbbm{E}(\boldsymbol{S}'\boldsymbol{Q}\boldsymbol{S}). \addtocounter{equation}{1}\tag{\theequation} \end{align*} If $\boldsymbol{\lambda}_i=\boldsymbol{\lambda} \,\, \forall i$, then there is no need to purge the factor structure through the loading space. That is, a simple cross-sectional demeaning transformation will suffice, $\boldsymbol{Q}=(\boldsymbol{I}_N-\boldsymbol{\tilde{\Lambda}(\tilde{\Lambda}'\tilde{\Lambda})^{-1}\tilde{\Lambda'}})=(\boldsymbol{I}_N-\boldsymbol{\frac{\iota\iota'}{N}})$. We can simplify equation (ref) to (where we make use of the normalization that $\boldsymbol{S}'\boldsymbol{\iota}=1$) \begin{align*} \mathbbm{v}_{z\varepsilon} &=\sigma^2_{\varepsilon}\cdot\sigma^2_{\tilde{u}} \cdot \left(\mathbbm{E}(\boldsymbol{S'S})-\frac{1}{N}\right) =\sigma^2_{\varepsilon}\cdot \underbrace{\sigma^2_{\tilde{u}} \cdot \left(\mathbbm{E}(h_{N,\mu})-\frac{1}{N}\right)}_{$\mathbbm{E}(z_t^2)$}, \end{align*} whereas, $\mathbbm{m}_{zp}=\mathbbm{E}(p_tz_t) \propto \mathbbm{E}(z_t^2)$. Hence, \begin{align} \mathbbm{v}_d &\propto \dfrac{\sigma^2_{\varepsilon} \cdot \sigma^2_{\tilde{u}} \cdot \left(\mathbbm{E}(h_{N,\mu})-\frac{1}{N}\right) }{\left[\sigma^2_{\tilde{u}} \cdot \left(\mathbbm{E}(h_{N,\mu})-\frac{1}{N}\right)\right]^2} =\dfrac{\sigma^2_{\varepsilon} }{\sigma^2_{\tilde{u}} \cdot \left(\mathbbm{E}(h_{N,\mu})-\frac{1}{N}\right)}. \end{align} Thus, the more concentrated the market, the more precise the estimator. See Section (ref) for a more general treatment.

Just identified supply elasticity. For the just identified supply elasticity estimator in (ref), upon convergence of Algorithm (ref), we have that

align[align omitted — 596 chars of source]

We can write the scalars $\widehat{A}$, $\widehat{B}$ and $\widehat{C}$ as follows

align*[align* omitted — 1,329 chars of source]

It is shown in Lemma (ref) of the Appendix that the terms $a_i, b_i, c_j$ are $o_p(1)$ for $i=1,2$; $j=1,2,3$, such that

align[align omitted — 218 chars of source]

We can now state the following result:

theorem[Consistency of $\widehat{\phi}^s$] Under Assumptions 1-4, as $(N,T) \overset{j}{\rightarrow} \infty$, we have \begin{align} \widehat{\phi}^s - \phi^s \overset{p}{\rightarrow} 0. \end{align}

Now, multiplying (ref) by $\sqrt{T}$

equation[equation omitted — 593 chars of source]

we can state the following result for the limiting distribution:

theorem[Limiting distribution for $\widehat{\phi}^s$] Under Assumptions 1-4, as $(N,T) \overset{j}{\rightarrow} \infty$, we have that when $N\geq T$, $N/T^{3/2} \rightarrow 0$ and $\sqrt{T}/N \rightarrow 0$; or when $N < T$ only $\sqrt{T}/N \rightarrow 0$ \begin{align} \sqrt{T} (\widehat{\phi}^s-\phi^s) \overset{d}{\rightarrow}\mathcal{N}\left(0, \mathbbm{v}_s\right), \end{align} where $\mathbbm{v}_s:=\mathbbm{m}_{z\tilde{p}}^{-2}\,\mathbbm{v}_{zu}$, $\mathbbm{v}_{zu}:=\mathbbm{E}(z_t^{\,2} \,(\boldsymbol{M}_{\eta}\,\boldsymbol{u}_{E})_{t}^{\,2})$ and $\mathbbm{m}_{z\tilde{p}} =\mathbbm{E}(z_t(\boldsymbol{M}_{{\eta}}\,\boldsymbol{p})_{t})$.

$\mathbbm{v}_{zu}$ can be consistently estimated with

align[align omitted — 635 chars of source]

Hence $ \frac{\sqrt{T}(\widehat{\phi}^{\,s}-\phi^s)}{\widehat{\mathbbm{v}}_{s}^{\,\,1/2}} \sim \text{t}_{df} \overset{d}{\rightarrow} \mathcal{N}(0,1)$, where $\widehat{\mathbbm{v}}_s^{1/2}=\widehat{\mathbbm{m}}_{z\tilde{p}}^{-2}\widehat{\mathbbm{v}}_{zu}$, with $\widehat{\mathbbm{m}}_{z\tilde{p}}=T^{-1}\sum_{t=1}^T\widehat{z}_t(\boldsymbol{M}_{\widehat{\eta}}\,\boldsymbol{p})_{t}$ consistent for $\mathbbm{m}_{z\tilde{p}}$. We will see in Section (ref) that asymptotic theory provides good approximations to the finite sample distribution.

Overidentified demand elasticity. For the overidentified demand elasticity estimator in (ref), recall the estimated instrument matrix consists of $\boldsymbol{\widehat{Z}}_d=

pmatrix[pmatrix omitted — 75 chars of source]

$. For the strong factors, $\boldsymbol{\widehat{\eta}}$, estimated via PCA, \citet{bai2010instrumental} showed that the generated regressors problem of \citet{pagan1984econometric} does not arise when both $N$ and $T$ are large. Thus, the sampling error in $\boldsymbol{\widehat{\eta}}$ is negligible in consideration of the limiting distribution of the overidentified demand elasticity estimate. In the previous section, we established that estimation of $\boldsymbol{\widehat{z}}_{GIV}$ is also negligible under regularity, we can then use standard asymptotic theory to also obtain asymptotic normality of the efficient GMM estimator in the case of demand since

align[align omitted — 602 chars of source]

We can now state the following theorem:

theorem[Limiting distribution for $\widehat{\phi}^d_{GMM}$] Under Assumptions 1-4, with $\mathbbm{E}(\varepsilon_t\boldsymbol{\eta}_t)=0 \, \forall t$, as $(N,T) \overset{j}{\rightarrow} \infty$, we have that when $N\geq T$, $N/T^{3/2} \rightarrow 0$ and $\sqrt{T}/N \rightarrow 0$; or when $N < T$ only $\sqrt{T}/N \rightarrow 0$ \begin{align} \sqrt{T} (\widehat{\phi}^d_{GMM}-\phi^d) \overset{d}{\rightarrow} \mathcal{N}\left(0, \mathbbm{V}(\widehat{\phi}^d_{GMM})\right), \end{align} where \begin{align} \mathbbm{V}(\widehat{\phi}^d_{GMM})&=\left(\boldsymbol{m}'_{{Z}_{d}p}\, \boldsymbol{\Omega}_d^{-1} \boldsymbol{m}_{{Z}_{d}p}\right)^{-1}, \end{align} with $\boldsymbol{m}_{{Z}_{d}p}=\mathbbm{E}(\boldsymbol{Z}_{dt}\,p_t)$ and $\boldsymbol{\Omega}_d=\emph{plim }T^{-1}\sum_{t=1}^T\boldsymbol{\widehat{Z}}_{dt}\boldsymbol{\widehat{Z}}_{dt}^{'}(d_t-p_t\widehat{\phi}^d_{2SLS})^2$.

$\mathbbm{V}(\widehat{\phi}^d_{GMM})$ can be consistently estimated using $2SLS$ residuals with

align[align omitted — 251 chars of source]

where $\boldsymbol{\widehat{\Omega}}_d=T^{-1}\sum_{t=1}^T\boldsymbol{\widehat{Z}}_{dt}\boldsymbol{\widehat{Z}}_{dt}^{\,'}\,(d_t-p_t\widehat{\phi}^d_{2SLS})^2$. It is well known that $\mathbbm{V}(\widehat{\phi}^d_{GMM})$ attains the semiparametric efficiency bound, as shown by chamberlain1987asymptotic, which reduces to (ref) in the linear model. Standard overidentification tests can be carried out since

align[align omitted — 292 chars of source]

where the degrees of freedom is given by $df_d=(1+r)-k_d=r$ and $k_d=1$ is the number of endogenous regressors. We will see that simulation evidence shows that the size of the $J$-test is near the nominal size when the true $r$ is used and when $rmax > r$ factors are used; which is important in empirical work when $r$ is typically estimated and it is generally known that an overestimate of $r$ is preferred in order to prevent an effect akin to omitted variable bias, see moon2015linear who formalize this notion.

Overidentified supply elasticity. The full instrument matrix for the overidentified supply elasticity estimator in (ref) consists of $\boldsymbol{\widehat{Z}}_s=

pmatrix[pmatrix omitted — 110 chars of source]

$, (recall the factors self instrument here, as they are part of the supply equation). We show that the sampling error in $\widehat{\boldsymbol{Z}}_s$ is indeed negligible. This is again due to both large $N$ and $T$. As a result, $\underset{(2+r) \times (2+r)}{\boldsymbol{\Omega}_s}=plim \,T^{-1}\sum_{t=1}^T\boldsymbol{\widehat{Z}}_{st}\boldsymbol{\widehat{Z}}^{'}_{st} (y_{\widehat{E}t}-\boldsymbol{\widehat{f}}_t'\boldsymbol{\widehat{\theta}}^s_{GMM})^2$ is sufficient when constructing the efficient weighting matrix, even though it does not take the sampling error in our estimate of ${\phi}^d$ into account (since $\boldsymbol{\widehat{\varepsilon}}=\boldsymbol{\varepsilon}(\widehat{\phi}_{GMM}^d)$). That is

align[align omitted — 608 chars of source]

We can now state the following theorem:

theorem[Limiting distribution for $\boldsymbol{\widehat{\theta}}^s_{GMM}$] Under Assumptions 1-4, as $(N,T) \overset{j}{\rightarrow} \infty$, we have that when $N\geq T$, $N/T^{3/2} \rightarrow 0$ and $\sqrt{T}/N \rightarrow 0$; or when $N < T$ only $\sqrt{T}/N \rightarrow 0$ \begin{align} \sqrt{T} (\boldsymbol{\widehat{\theta}}^s_{GMM}-\boldsymbol{\theta}^s)&\overset{d}{\rightarrow} \mathcal{N}\left(\boldsymbol{0}, \mathbbm{V}(\boldsymbol{\widehat{\theta}}^s_{GMM})\right), \end{align} where \begin{align} \mathbbm{V}(\boldsymbol{\widehat{\theta}}^s_{GMM})&=\left(\boldsymbol{m}'_{{Z}_{s}f}\, \boldsymbol{\Omega}_s^{-1} \boldsymbol{m}_{{Z}_{s}f}\right)^{-1}, \end{align} with $\boldsymbol{m}_{{Z}_{s}f}=\mathbbm{E}(\boldsymbol{Z}_{st}\boldsymbol{f}'_t)$ and $\boldsymbol{\Omega}_s=\emph{plim }T^{-1}\sum_{t=1}^T\boldsymbol{\widehat{Z}}_{st}\boldsymbol{\widehat{Z}}_{st}^{'}(y_{\widehat{E}t}-\boldsymbol{\widehat{f}}_t'\boldsymbol{\widehat{\theta}}^s_{GMM})^2$.

$\mathbbm{V}(\boldsymbol{\widehat{\theta}}^s_{GMM})$ can be consistently estimated using $GMM$ residuals with

align[align omitted — 285 chars of source]

where $\boldsymbol{\widehat{\Omega}}_s=T^{-1}\sum_{t=1}^T\boldsymbol{\widehat{Z}}_{st}\boldsymbol{\widehat{Z}}_{st}^{\,'}\,(y_{\widehat{E}t}-\widehat{\boldsymbol{f}}_t'\boldsymbol{\widehat{\theta}}^s_{GMM})^2$. Just as in the case of the overidentified demand elasticity estimator, (ref) achieves the semiparametric efficiency bound. Overidentification tests can be carried out since

align[align omitted — 328 chars of source]

where the degrees of freedom are given by $df_s=2-k_s=2-1=1$ and $k_s=1$ is the number of endogenous regressors for the supply equation.

assumpRThe asymptotic distribution of the FGIV and efficient GMM estimators of demand and supply elasticities are established. However, the finite sample moments of these estimators, are unbounded to different degrees. The extensive literature on the classic simultaneous equations model has documented this result in many forms, see mariano1973approximations, hatanaka1973existence, sawa1972finite, takeuchi1970exact, ullah1974exact, sargan1978existence, fuller1977some and hillier1981exact. A complete representation of the above results was given by kinal1980existence. Kinal's result for 2SLS states that, if the dependent variable, explanatory variables and instruments are jointly normal, then $\mathbbm{E}||\widehat{\phi}^j_{2SLS}||^m<\infty$ for $m < \ell_j - k_j +1$, $j=d,s$, where $\ell_j$ is the number of instruments and $k_j$ is the number of endogenous regressors. Thus, the FGIV estimators for supply and demand exhibit no bounded absolute moments since $\ell_j=k_j=1$, $j=d,s$. Whereas, the efficient GMM estimators exhibits $\ell_d-k_d=(1+r)-1=r$ bounded absolute moments in finite samples for the case of demand and $\ell_s-k_s=2-1=1$ bounded absolute moment in finite samples for the case of supply. Hence, the efficient GMM elasticity estimators (overidentified FGIV estimators) exhibit superior finite sample properties relative to their (just identified) FGIV counterparts. Of course, in general, just identified instrumental variables estimators (with strong instruments) exhibit nice properties asymptotically.

Weak Instruments

The classical weak instruments framework introduced by staiger1997instrumental has its analog in this framework. Interestingly, here the "weak" aspect is partially linked to the Herfindahl index without making the usual local-to-zero assumption as in staiger1997instrumental. Moreover, the traditional notion of local-to-zero with $\frac{1}{\sqrt{T}}$ scaling which matches the rate of convergence of the estimator need not necessarily apply here for weak instruments to arise. More specifically, the locality to zero can be expressed as decaying functions of $N$, except in the case of $\mu \in (0,1)$, which we require for inference under our maintained strong instruments assumption; whereas the rate of convergence is at the $\sqrt{T}$ rate. To make things more clear, it is useful to see the reduced form, equilibrium price equation again. Recall from (ref), we have that $p_t = \dfrac{1}{\phi^d - \phi^s} \left(u_{St} + \boldsymbol{\lambda}'_S \boldsymbol{\eta}_t -\varepsilon_t \right)$. Thus, it is clear that for finite $N$, $\mathbbm{Cov}(p_t, z_{t}) > 0$, which automatically renders the GIV as relevant. However, for large $N$, writing $z_{t}=\boldsymbol{S}'\boldsymbol{Q} \boldsymbol{\tilde{u}}_{\cdot t}$ we observe that

align[align omitted — 382 chars of source]

where $\boldsymbol{\Sigma}_{\tilde{u}}:=\mathbbm{E}(\boldsymbol{\tilde{u}}_{\cdot t}\boldsymbol{\tilde{u}}_{\cdot t}')$. The term inside the expectation can be simplified to

align*[align* omitted — 647 chars of source]

where we make use of $\gamma_{max}(\boldsymbol{\Sigma}_{u}) =\mathcal{O}\left(1\right)$ and the fact that a symmetric idempotent matrix, such as $\boldsymbol{Q}$, has eigenvalues of $0$ or $1$ and so $\gamma_{max}(\boldsymbol{Q})=1$. Taken together, (ref) and (ref) imply

align[align omitted — 116 chars of source]

As such, only when we are in a tail regime indexed by $\mu \in (0,1)$ do we avoid the locality to zero. For example, when $\mu > 2$, we have that $\mathbbm{V}(\boldsymbol{S}'\boldsymbol{Q} \boldsymbol{\tilde{u}}_{\cdot t}) \leq \mathcal{O}\left(1\right)\mathbbm{E}(h_{N, \, \mu>2})=\mathcal{O}\left(\frac{1}{N}\right)$. As a result, when $\mu >2$, we have that $z_{t}=\boldsymbol{S}'\boldsymbol{Q}\boldsymbol{\tilde{u}}_{\cdot t}=\mathcal{O}_p \left(\frac{1}{\sqrt{N}} \right)$, so our equilibrium price equation simplifies to the following large $N$ representation

align[align omitted — 225 chars of source]

This would render the GIV as very weak since $\mathbbm{Cov}(p_t, u_{St}) =\mathbbm{Cov}(p_t, z_{t}) = \mathcal{O} \left(\frac{1}{N}\right)$. Note that the $\mathbbm{Cov}(p_t, u_{St})$ and $\mathbbm{Cov}(p_t, z_{t})$ are of the same order precisely because $\gamma_{max}(\boldsymbol{Q})=1$. (ref) is effectively the relationship that was exploited by mohaddes2016country, who assumed the so-called "granular" weights of order $\mathcal{O}\left(\frac{1}{N}\right)$ and used this weak correlation for large $N$ to ultimately deduce that prices can be treated as weakly exogenous.\footnote{"Granular" has a different definition in the panel data literature, which is referring to properties of weights and heuristically, rules out the existence of dominant units, see e.g., mohaddes2016country and (ref). On the contrary, our usage of the term "granular" follows gabaix2011granular and is essentially referring to the existence of dominant cross-sectional units, see Section (ref).}

Consider the well documented and empirically relevant case where $\mu$ is just above $1$ (Zipf's law corresponds to $\mu=1$); when $\mu \in (1,2)$ we have $h_{N, \, \mu \in (1,2)}=\mathcal{O}_p\left(1/(N^{2-\frac{2}{\mu}})\right)$. So,

align[align omitted — 179 chars of source]

Therefore, even though $\mathbbm{Cov}(p_t, z_{t}) = \mathcal{O}\left(1/(N^{2-\frac{2}{\mu}})\right)$, it is in fact decaying to zero so slowly for $\mu$ near 1, that this potentially corresponds to a highly relevant instrument in any finite sample. That is, $\mathbbm{Cov}(p_t, z_{t}) = \mathcal{O}\left(1/(N^{2-\frac{2}{\mu}})\right)$, is potentially consistent with $z_t$ accounting for large fractions of aggregate variation, see gabaix2011granular.

However, the case we theoretically entertain, for consistency and valid asymptotic inference, requires $\mu \in (0,1)$, which in conjunction with the additional regularity assumptions, renders $\mathbbm{Cov}(p_t, z_{t}) =\Theta(1)$ even as $N \rightarrow \infty$.

Rothemberg Representations. Moreover, to further assess the likelihood of weak instruments, we can analyze the efficient GMM estimator of the demand elasticity which uses both the GIV and the factors as instruments and for comparison we can analyze the just identified demand elasticity estimator which uses only the GIV as an instrument. We analyze the overidentified case with conditional homoskedasticity (assuming only for remainder of this section). Define the projection matrix $\boldsymbol{P}_{Z_d} = \boldsymbol{Z}_d(\boldsymbol{Z}_d'\boldsymbol{Z}_d)^{-1}\boldsymbol{Z}_d'$, the 2SLS estimator takes the form

align[align omitted — 186 chars of source]

Write the structural and reduced form equations as

align[align omitted — 165 chars of source]

where $\underset{T\times (1+r)}{\boldsymbol{z}}=

pmatrix[pmatrix omitted — 50 chars of source]

$, $\underset{(1+r) \times 1}{\boldsymbol{\pi}}=

pmatrix[pmatrix omitted — 91 chars of source]

'$ and $\underset{T \times 1}{\boldsymbol{v}}=\frac{1}{\phi^d-\phi^s}\cdot \boldsymbol{\varepsilon}$.

assumpRThe difference between $\boldsymbol{z}$ in (ref) and our actual instrument, $\boldsymbol{Z}_d$, boils down to the difference between their first columns, $\boldsymbol{Z}_d[\cdot,1]=\boldsymbol{z}_{GIV}=\boldsymbol{\tilde{u}}_{\cdot \cdot}\boldsymbol{Q}\boldsymbol{S}$ and $\boldsymbol{z}[\cdot,1]=\boldsymbol{u}_S=\boldsymbol{u}_{\cdot \cdot}\boldsymbol{S}$. In the case of the demand equation, $\boldsymbol{u}_S$ is ideal, whereas $\boldsymbol{z}_{GIV}$ is a proxy. The reason the proxy is used is simply due to a simpler theoretical exposition than a direct estimate for the ideal. Indeed $\boldsymbol{z}_{GIV}$ is in fact a good proxy. For example, the correlation between $\boldsymbol{u}_S$ and $\boldsymbol{z}_{GIV}$ is over 90% regardless of the complexity of our DGP in Monte Carlo simulations even for small configurations of $(N,T)$. Moreover, in the case of the supply equation, $\boldsymbol{u}_S$ is no longer valid, whereas $\boldsymbol{z}_{GIV}$ is; see (ref). Mathematically, $\boldsymbol{S}'\boldsymbol{u}_{\cdot t}-\boldsymbol{S}'\boldsymbol{Q}\boldsymbol{\tilde{u}}_{\cdot t} = \boldsymbol{S}'\boldsymbol{P}_{\tilde{\Lambda}}\boldsymbol{u}_{\cdot t}=\boldsymbol{S}'\boldsymbol{P}_{\tilde{\Lambda}}\boldsymbol{P}_{\tilde{\Lambda}}\boldsymbol{u}_{\cdot t}$ for each $t$, where $\boldsymbol{P}_{\tilde{\Lambda}}$ is the symmetric and idempotent projection matrix in the demeaned loading space. Hence, $\boldsymbol{S}'\boldsymbol{P}_{\tilde{\Lambda}}\boldsymbol{P}_{\tilde{\Lambda}}\boldsymbol{u}_{\cdot t}$ is zero when the loadings and the share vector are asymptotically uncorrelated and/or the loadings and idiosyncratic errors are asymptotically uncorrelated; which explains why our simulations exhibit near perfect correlation.

Then, it follows from rothenberg1984approximating that (ref) has the following illustrative representation

align[align omitted — 203 chars of source]

where $X=\boldsymbol{\pi}'\boldsymbol{z}' \boldsymbol{P}_{Z_d}\boldsymbol{\varepsilon} /(\sigma^2_{\varepsilon} \boldsymbol{\pi}'\boldsymbol{z}'\boldsymbol{z}\boldsymbol{\pi})^{\frac{1}{2}}$ and $Y=\boldsymbol{\pi}'\boldsymbol{z}' \boldsymbol{P}_{Z_d} \boldsymbol{v} /(\sigma^2_{v} \boldsymbol{\pi}'\boldsymbol{z}'\boldsymbol{z}\boldsymbol{\pi})^{\frac{1}{2}}$ are bivariate standard normal variates with correlation coefficient $\rho$. The random variable $\omega_1=\boldsymbol{v}' \boldsymbol{P}_{Z_d}\boldsymbol{\varepsilon} /(\sigma^2_{\varepsilon}\sigma^2_{v})^{\frac{1}{2}}$ has mean equal to $\text{rank}(P_{Z_d})\rho=(r+1)\rho$ and variance equal to $(r+1)(1+\rho^2)$. The random variable $\omega_2=\boldsymbol{v}'\boldsymbol{P}_{Z_d} \boldsymbol{v} /\sigma^2_{v}$ has mean equal to $\text{rank}(\boldsymbol{P}_{Z_d})=(r+1)$ and variance equal to $2(r+1)$. Finally, $\mu_{d,GMM}$ is the square root of the so-called concentration parameter $\mu_{d,GMM}^2=\boldsymbol{\pi}'\boldsymbol{z}'\boldsymbol{z}\boldsymbol{\pi}/\sigma^2_{v}$ for the demand equation. $\mu_{d,GMM}$ plays the role of $\sqrt{T}$, that is, when $\mu_{d,GMM}$ is large, $\mu_{d,GMM}(\widehat{\phi}^d_{GMM}-\phi^d)$ is well approximated by a $\mathcal{N}(0,1)$ variate. Large values of $\mu_{d,GMM}$ are consistent with large values of $T$, i.e., our typical large sample approximations. However, large values of $\mu_{d,GMM}$ are also consistent with small values of $\sigma^2_v$, regardless of the value of $T$, i.e., small-$\sigma$ asymptotics, as introduced originally by kadane1971comparison. More insights can be gained by simplifying the concentration parameter for the demand elasticity

align*[align* omitted — 604 chars of source]

where the approximation is due to ignoring the terms involving $\boldsymbol{\eta}'\boldsymbol{u}_S$, which are zero only in expectation. (ref) is very intuitive, if the proportion of the volatility in the GIV and size-weighted common components dominate the volatility of the demand shocks, so that the ratio in (ref) is large, then the concentration parameter $\mu_{d,GMM}$ will be large and one should expect good approximations to the finite sampling distributions.

On the other hand, when only the GIV is used as an instrument, if we redefine $\underset{T\times 1}{\boldsymbol{z}}=\boldsymbol{u}_S$, $\underset{1 \times 1}{\boldsymbol{\pi}}=\frac{1}{\phi^d-\phi^s}$ and $\underset{T \times 1}{\boldsymbol{v}}=\frac{1}{\phi^d-\phi^s}\cdot (\boldsymbol{\varepsilon}+\boldsymbol{\eta}\boldsymbol{\lambda}_S)$ from (ref) and simply follow the logic above through (ref), we arrive at the following concentration parameter for the FGIV estimator

align[align omitted — 163 chars of source]

Thus, by inspection of (ref) against (ref) we can see that in the case of the just identified FGIV estimator, we would need the volatility of just the GIV to drive up the ratio of the concentration parameter, $\mu^2_{d,GIV}$, and the size-weighted common component would be working against us (in the denominator), in this case, instead of working for us as in $\mu^2_{d,GMM}$.

Although the literature on granularity has demonstrated that idiosyncratic shocks alone can be quite volatile, in this context, we advocate starting with the efficient GMM estimators, since the $J$-test is well sized as illustrated with simulation evidence, because the efficient GMM estimators can exhibit substantially improved finite sample properties relative to the just identified estimators and are less likely to suffer from weak instrument issues as well.

Monte Carlo

We simulate the following panel simultaneous equations system with latent factor structure that was analyzed in the theoretical sections:

align*[align* omitted — 399 chars of source]

We consider two sets of simulated experiments. In Design 1, we let $u_{it}$ be $i.i.d.$ to establish a set of baseline results. In Design 2, we allow for sparse cross-sectional dependence in $u_{it}$. In addition, in unreported simulations, we simulate $z_{t,GIV}$ to be a weak instrument to illustrate that the efficient GMM estimators are robust to this as they optimally shift their weights away from this point of weakness, whereas the just identified estimators of GK and our FGIV will substantially deteriorate in their performances.

Design 1 - $u_{it} \,\, \boldsymbol{i.i.d.}$ case. We set $\phi^s=0.1$ and $\phi^d=-0.3$. We draw the supply factors and loadings as, $\underset{T\times r}{\boldsymbol{\eta}} \overset{i.i.d.}{\sim} \mathcal{N}(0,\boldsymbol{I}_r)$ and $\underset{N\times r}{\boldsymbol{\Lambda}} \overset{i.i.d.}{\sim} \mathcal{N}(0,\sigma^2_{\Lambda}\boldsymbol{I}_N)$, respectively, with $r=2$.\footnote{The results do not change significantly if we draw the loadings from a Uniform distribution with non-zero mean.} We draw the idiosyncratic supply shocks as $\underset{T\times N}{\boldsymbol{u}} \overset{i.i.d.}{\sim} \mathcal{N}(0,\sigma^2_u\boldsymbol{I}_N\otimes \boldsymbol{I}_{T})$ and aggregate demand shocks as $\underset{T\times 1}{\boldsymbol{\varepsilon}} \overset{i.i.d.}{\sim} \mathcal{N}(0,\sigma^2_{\varepsilon}\boldsymbol{I}_T)$.

Design 2 - $u_{it}$ non-$\boldsymbol{i.i.d.}$ case. Everything is identical to Design 1, except that we no longer set $\boldsymbol{\Sigma}_u=\sigma^2_u \boldsymbol{I}_N$ for each $t=1,\dots, T$. We generate a non-diagonal banded covariance matrix; as such, it satisfies the sparsity requirement from (ref).\footnote{In unreported simulations, we find that the results do not change significantly if we generate a dense, non-diagonal $\boldsymbol{\Sigma}_u$, such as one arising from a cross-sectional AR(p) process. This is because although the cross-sectional AR(p) generates a dense matrix, it is not too dense since the off-diagonals decay exponentially fast to 0 as $|i-j| \rightarrow \infty$.} We consider the following banded idiosyncratic covariance matrix with cross-sectional dependence and heteroskedasticity

align[align omitted — 172 chars of source]

with bandwidth $k=3$ and $\sigma^2_{u,i}$ are drawn from $\mathcal{U}[0.5,1]$.

Target parameterizations. The variance of the price process takes the following form $\mathbbm{V}(p_t)=c\cdot \left( \mathbbm{V}(u_{St}) +\mathbbm{V}(\boldsymbol{\lambda}'_S\boldsymbol{\eta}_t) +\mathbbm{V}(\varepsilon_t) \right)$, where $c=\dfrac{1}{(\phi^d-\phi^s)^2}$. This conveniently allows us to parameterize the relative volatilities of the various components of equilibrium prices. We parameterize the individual variances, $\sigma^2_u$ (for Design 1), $\sigma^2_{\Lambda}$ and $\sigma^2_{\varepsilon}$ such that $\psi_{u} := \dfrac{\mathbbm{V}\left(\sqrt{c}\cdot u_{St}\right)}{\mathbbm{V}(p_t)} \in (0.15, 0.35),$ $\psi_{u+\eta} := \dfrac{\mathbbm{V}\left(\sqrt{c}\cdot (u_{St}+\boldsymbol{\lambda}'_S\boldsymbol{\eta}_t)\right)}{\mathbbm{V}(p_t)} \in (0.45, 0.65),$ and $\psi_{u+\varepsilon} := \dfrac{\mathbbm{V}\left(\sqrt{c}\cdot (u_{St}+\varepsilon_t)\right)}{\mathbbm{V}(p_t)} \in (0.45, 0.65).$\footnote{The interval for $\psi_u$ is consistent with the literature on granularity, which has documented the proportion of aggregate fluctuations traced back to idiosyncratic shocks falling in this specified range.} In Design 1, we achieve an average across simulations of $\bar{\psi}_u \approx 0.23$, $\bar{\psi}_{u+\eta} \approx 0.58$ and $\bar{\psi}_{u+\varepsilon} \approx 0.65$ which implies that $\bar{\psi}_{\eta}=0.35$ and $\bar{\psi}_{\varepsilon}=0.42$. That is, the idiosyncratic shocks are not the dominating force in terms of observed price volatility; however, their granular role is still substantial enough to draw inferences from when used as instruments. In Design 2, we achieve an average across simulations of $\bar{\psi}_u \approx 0.27$, $\bar{\psi}_{u+\eta} \approx 0.64$ and $\bar{\psi}_{u+\varepsilon} \approx 0.63$.

Let $\widehat{\phi}^j(m)$, $j=d,s$, denote the estimate in the $m^{th}$ monte carlo repetition, $m=1,...,M$. We report the monte carlo bias: $\text{Bias}(\widehat{\phi}^j)= \left(\dfrac{1}{M} \sum_{m=1}^M \widehat{\phi}^j(m)-\phi^j \right)$ for $j=d,s;$ and square root of the monte carlo MSE: $\text{RMSE}(\widehat{\phi}^j)= \sqrt{\dfrac{1}{M}\sum_{m=1}^M(\widehat{\phi}^j-\phi^j)^2}$ for $j=s,d.$ Additionally we report the size of the $t$-test for all estimators and size of the $J$-test for the efficient GMM estimators. The results are reported in Table (ref) (Design 1) and Table (ref) (Design 2). In Table (ref), we multiply the bias by 100 because all the estimators perform quite well in this ideal setting. For nearly each configuration of $(N,T)$ the GMM estimators perform the best, in terms of bias and RMSE, as the theory suggests. Importantly, the $t$-test and $J$-test is well sized even when $rmax=3 > r=2$ factors are used. In Table (ref), we report the bias as is, and we find that for a given configuration of $(N,T)$ the bias is two orders of magnitude larger in Design 2 relative to Design 1. Nevertheless, as the theory would suggest, the efficient GMM estimators perform the best in terms of bias and RMSE. There are some size distortions for the supply side estimators but the distortions are decreasing in $N$ for a given $T$.

Application to Global Crude Oil Markets

The data construction follows the recent literature: kilian2009not, caldara2019oil, baumeister2019structural (hereafter BH) and GK. The following is a breakdown of the raw variables collected for Jan. 1985 - Dec. 2015 $(T=372$ months$)$: monthly oil production for $N=22$ countries from the U.S. Energy Information Administration (hereafter EIA); world oil production from the U.S. EIA; monthly oil prices based on the refiner acquisition cost of imported crude oil from the U.S. EIA; U.S. CPI from the St. Louis FRED database; monthly change in inventories from BH; monthly industrial production index from BH. The CPI is used to deflate nominal oil prices to arrive at the real price of oil, which is highly non-stationary. Following the aforementioned literature, we take the logarithm of the real price of oil series and then take first differences. We apply the same transformation to the monthly oil production for each country. These transformations render the production and price series stationary as confirmed by a host of Dickey-Fuller tests. For ensuring the tail index of the size-distribution, $\mu$, is in the region the theory requires, we provide visual evidence along with 6 estimates of $\mu$ that all fall beneath 1, see Table (ref). Also, see Figure (ref), Figure (ref) & Figure (ref).

Let $y_{it}$ denote the log difference of the oil supply for country $i$ at time $t$ and $p_t$ denote the log difference of the real price of oil. Following GK, we estimate an OPEC factor using information on the cross-section of countries (i.e., known loadings). To that end, let $o_{it}$ denote a dummy variable equal to 1 if country $i$ is an OPEC member at time $t$ and note that $o_{it}=o_i$ for most $i$, with the exception of Gabon and Ecuador in our sample. Finally, $\boldsymbol{c}_{t-1}$ denotes a $4 \times 1$ vector containing: lagged $p_t$, lagged world supply growth, lagged change in inventories, and lagged growth in industrial production. The system is given as follows

align[align omitted — 301 chars of source]

where we lose the observation $t=1$ due to differencing. The cross-sectionally demeaned supply equation is given by the approximate factor model,

align[align omitted — 224 chars of source]

where $\tilde{e}_{it} :=\boldsymbol{\tilde{\lambda}}_i'\boldsymbol{\eta}_t + \tilde{u}_{it}$. Note that (ref) implies we can obtain the OPEC factor, $\eta_{OPEC,t}$, via cross-sectional regression, for each $t>1$, that is $\widehat{\eta}_{OPEC,t}=( \boldsymbol{\tilde{o}}_{\cdot t}' \boldsymbol{\tilde{o}}_{\cdot t})^{-1} \boldsymbol{\tilde{o}}_{\cdot t}'\boldsymbol{\tilde{y}}_{\cdot t}$. Hence, in our preliminary stage, we extract $\widehat{\eta}_{OPEC,t}$ and then run PCA on $\widehat{\tilde{e}}_{it}=\tilde{y}_{it}-\widehat{\eta}_{OPEC,t}$ to extract the latent demeaned loadings and latent factors. Define $\boldsymbol{y}^*_{\cdot t} := \boldsymbol{\tilde{y}}_{\cdot t}-\boldsymbol{\tilde{x}}_{\cdot t}\widehat{\eta}_{OPEC,t}$, then we purge the latent factors via $\boldsymbol{Q}$ as in the main text: $\boldsymbol{{Q}}\,\boldsymbol{y}^*_{\cdot t}= \boldsymbol{{Q}}\,\boldsymbol{\tilde{u}}_{\cdot t}$. However, when forming the GIV, there is a minor difference that we have time-varying size-weights, so we no longer construct $\boldsymbol{z}_{GIV}$ with a time-invariant share vector $S_i$, but rather we weight each idiosyncratic component at time $t$ with its corresponding share from time $t-1$ to avoid endogeneity issues arising from contemporaneous weighting

align[align omitted — 350 chars of source]

Besides these modifications from the stylized model in the theory, we estimate the elasticities using the estimators outlined in the main text. The number of factors, $r$, is estimated via the AH procedures (as outlined in the Supplementary Appendix (ref)). The $ER$ method of AH estimated $\widehat{r}_{ER}=1$, while the $GR$ method estimated $\widehat{r}_{GR}=3$; with $k_{max}=10$. To be safe, we take $\widehat{r}=\widehat{r}_{GR}+1$.

Supply results. The results for the supply elasticity are presented in Table (ref). In Table (ref), the 2nd column displays GK's results. The instrument GK use is given by (ref) and their dependent variable is simply the cross-sectional average of the log difference of oil supply (i.e., $\boldsymbol{E}=\boldsymbol{\iota}/N$). Our results are in columns 3 and 4. In contrast to (ref), the instrument we use in column 3 purges the common factors through the loading space. The instrument we use in column 4 also adds an estimate of the unobserved aggregate demand shocks, $\widehat{\varepsilon}_t$, to our FGIV. Moreover, the dependent variable we use is weighted using the estimated precision vector $\boldsymbol{\widehat{E}}$, which allows for cross-sectional correlations and heteroskedasticity in $u_{it}$. These differences lead to significantly different results. Columns 2 and 3, which attempt to use only GIVs as instruments, both lead to weak instruments as indicated by the first-stage $F$-statistics less than the rule of thumb, 10. Nevertheless, the FGIV supply elasticity estimate (0.016) from column 3 (estimated via Algorithm (ref)) is roughly one third that of GK's (0.044). Whereas, our efficient GMM supply elasticity estimate (0.005) from column 4 (estimated via Algorithm (ref)) is highly significant at the 1% level. Additionally, our results reveal that using estimates of unobserved aggregate demand shocks as supply instruments indeed renders a strong instrument as indicated by the first-stage $F$-stat of 14.33 in column 4. Moreover, the $p$-value for the $J$-statistic (0.11) fails to reject the null hypothesis of a valid model. An $F$-stat greater than 10, coupled with a small $J$-statistic provides statistical evidence in favor of our efficient GMM point estimate for the supply elasticity.

Demand results. Turning now to the demand elasticity in Table (ref), the dependent variable GK use in this case is the same as the one we use. However, the instruments are different. Column 2 displays GK's demand elasticity (-0.463), again using the instrument as in (ref). Column 3 presents our result when using only the FGIV as an instrument (-.0009), which is roughly 400 times smaller than GK's estimate. Columns 4 through 7 sequentially add principal components to the instrument vector for our efficient GMM estimator from column 3 until 4 principal components are used. Here we find that none of the models yield first-stage $F$-statistics greater than 10. It is reassuring, however, that the $J$-statistic for columns 4 through 7 all fail to reject the null of a valid model. Lastly, column 8 presents the bai2010instrumental estimator which only includes the four principal components but not the FGIV as instruments, nearly all statistics remain unchanged except that inclusion of the FGIV increases the $t$-stat by about 25%.

Taken together, our empirical results suggest that supply shocks, whether they be aggregate or idiosyncratic supply shocks, albeit valid, do not serve as strong instruments for estimation of the demand elasticity. Whereas, aggregate demand shocks indeed seem to be a strong source of exogenous variation to tease out the supply elasticity.

Concluding Remarks

In this paper, we have further developed the GIV methodology introduced by gabaix2020granular, which takes advantage of panel data to construct instruments for estimation of structural time series regression models that involve endogenous regressors. This paper focuses on the underlying econometric issues involved in developing FGIV in a large $N$ and large $T$ framework where the loadings are treated as unknown parameters to be estimated before constructing the FGIV instrument. We further demonstrate that the sampling error arising from estimating the instrument, factors and a high dimensional precision matrix does not affect the limiting distribution for the structural parameters of interest. We also overidentify the structural parameters, which leads to new and improved results in the crude oil markets application and demonstrate that the $J$-test is well sized with simulation evidence. Our Monte Carlo study illustrates that our estimators and algorithms exhibit desirable performance with the finite sample distributions being well approximated by the asymptotic distributions.

More fruitful areas of research would be empirical applications of the theoretical results derived in this paper. Interesting theoretical extensions would be to allow for random slope coefficients with correlated heterogeneity, the presence of weak factors and unbalanced panels with data not missing at random. We are currently pursuing the dynamic panel data extension, as well as adapting the GIV methodology for unit-specific endogenous variables.

table[table omitted — 4,083 chars of source]
table[table omitted — 4,102 chars of source]
table[table omitted — 848 chars of source]
table[table omitted — 1,580 chars of source]
table[table omitted — 2,166 chars of source]