EconBase
← Back to paper

Granular Instrumental Variables in Large Panels: Identification and Inference Across Strong, Nearly Weak, and Weak GIV

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

100,837 characters

Granular Instrumental Variables in Large Panels: Identification and Inference Across Strong, Nearly Weak, and Weak GIV



\addtocontents{toc}{\protect\setcounter{tocdepth}{-10}}

\begin{titlepage}
    \maketitle
    \thispagestyle{empty}

\begin{abstract}
    I develop the asymptotic theory of instrument strength for Granular Instrumental Variables (GIV) in large panels with both $N$ and $T$ growing. The strength of the GIV depends on the presence of dominant units. I formalise what dominance means and characterise three regimes of instrument strength. When a few units dominate the aggregate, the instrument is strong. The GIV estimator is consistent and asymptotically normal at the standard $\sqrt{T}$ rate. When large units stand out but do not dominate, the instrument weakens. But I show that the parameter of interest remains recoverable. The GIV estimator remains consistent and asymptotically normal, now at a rate slower than $\sqrt{T}$. When units are comparable in size and none stands out, the instrument is weak in the standard sense. The GIV estimator is inconsistent and has a non-standard distribution. Wald inference is reliable only outside the weak regime. When the instrument is weak, I recommend Anderson--Rubin confidence sets. In practice, the instrument must be constructed in a first stage. I show that the feasible estimator attains the same rate, but its asymptotic variance picks up an additional term from the first-stage estimation. Valid inference must use standard errors that account for this term. I apply the GIV estimator with the correct standard errors to recover the short-run demand elasticities of three commodities: refined copper, crude oil, and natural gas.

\end{abstract}




    \textit{Keywords:} Granular instrumental variables, Weak Instruments, Factor models, Power Law.
\end{titlepage}

\section{Introduction} \label{sec:introduction}

Many questions in economics require estimating structural relationships
between aggregate variables, for instance, how asset prices respond to changes in
aggregate demand, or how exchange rates react to capital flows. A central challenge is
endogeneity. The aggregate regressor is correlated with the structural error.
\citet{Gabaix2024GranularVariables} proposed Granular Instrumental Variables (GIV) as a
solution. The key insight is that when a few large units (dominant firms, banks, or
funds) disproportionately drive the aggregate, their idiosyncratic shocks can serve as
instruments. Several studies apply this idea to asset markets, bank lending, and sovereign risk.
However, the existing theory assumes a fixed number of units. This paper extends GIV to large panels while focusing on a core question: what is granularity and how much granularity is enough?

Specifically, \citet{Gabaix2024GranularVariables} (GK hereafter) consider the canonical supply-demand system
\begin{align*}
    d_t    &= \phi_d\, p_t + \varepsilon_t, \\
    y_{it} &= \phi_s\, p_t + \lambda_i' F_t + u_{it},
\end{align*}
where $d_t$ is the change in aggregate demand for a commodity, $p_t$ the change in market-clearing price, and $y_{it}$ the change in supply of unit $i$. The unit-level supply follows a panel data model with interactive fixed effects, where unit-specific loadings $\lambda_i$ interact with common time factors $F_t$. Market clearing, $d_t = \sum_{i=1}^N S_i y_{it}$, ties the two equations together, with $S_i$ the long-run market share of unit $i$. The object of interest is the demand elasticity $\phi_d$. The endogeneity problem is that $p_t$ responds to the demand shock $\varepsilon_t$.

For known shares, $S$, and the demeaning matrix, $D_N = I_N - \frac{\iota \iota'}{N}$, the Granular Instrumental Variable $z_t = S' D_N u_t$ is the optimal instrument for $\phi_d$.

As illustrated by \citet{Gabaix2024GranularVariables} and further explained by Gopalan, Nagasawa, and Renault (GNR hereafter), the strength of the instrument depends on the granularity of the setting. That is, some of the shares, $S_{i},i=1,...,N$ have to be different from $\frac{1}{N}$.
GK call this a granular setting, and hence the name of the instrument. GK only consider the case where $N$ is fixed and $T \to \infty$.

This asymptotic framework is not appropriate for many empirical applications. For instance, \citet{Aldasoro2023TheMarkets} has $N = 21$. \citet{Chodorow-Reich2021AssetInsulators,Galaasen2020GranularRisk,Ma2021ExpectationsLending} have $N > 100$. In these cases, a more appropriate asymptotic framework is one with both $N$ and $T \to \infty$. This necessitates extending the theory of GIV to large panels.

But as $N \to \infty$, granularity requires a careful asymptotic treatment. Granularity is a property of the cross section as it arises from the behavior of $\left( S_{i}\right) _{1\leq i\leq N}$. When $N$ is fixed, in the simplest case, we just need $S_i$ to be different from $\frac{1}{N}$ for some $i$ for the instrument to be valid. But as $N \to \infty$, the instrument can weaken if the cross-section is not concentrated enough. Thus, as we extend the theory of GIV to large panels, we need to formalize the idea of granularity and identify how it affects instrument strength.

This paper extends the theory of Granular Instrumental Variables to large panels $(N,T \to \infty)$ with a formal treatment of granularity, making three contributions. First, I formalize granularity by modeling unit sizes as draws from a power-law distribution. The tail index $\mu$ measures how much the largest units stand out, and serves as a single
sufficient statistic for instrument strength. Second, I characterize how instrument
strength varies with $\mu$ and the $N/T$ trajectory, identifying three regimes with distinct implications for
consistency, convergence rates, and inference. Third, I provide the correct asymptotic theory for the feasible GIV estimator which I construct from estimated rather than known idiosyncratic shocks. I then illustrate the theory empirically, using GIV to estimate the short-run demand elasticities of three major commodities (copper, crude oil, and natural gas).

Dominant units (granularity) arise naturally when unit sizes follow a power-law distribution. This is well documented across very different empirical settings such as firm sizes, city sizes, and bank assets \citep{Gabaix2009PowerFinance,Axtell2001ZipfSizes}.
Hence I model that the size of the unit, $\s_i$ comes from a power-law distribution. That is $\p(\s_i > s_i) = c s_i^{-\mu}$, $\mu >0$. From the observed sizes, we construct the shares as $S_i = \frac{\s_i}{\sum_{j=1}^N \s_j}$.

This tail index governs the strength of the instrument. Together with the $N/T$ trajectory, it delivers three regimes.

When a few units dominate the aggregate, the instrument is strong. This is the case $\mu \in (0,1)$, where the tail is heavy enough that a handful of units make up a non-vanishing share of the aggregate. The GIV estimator is consistent and asymptotically normal at the standard $\sqrt{T}$ rate. This is the classical strong instrument regime.

When large units stand out but do not dominate, the instrument weakens without breaking. This is the case $\mu > 1$ with $N/T \to 0$. The largest units are still big enough that their idiosyncratic shocks survive aggregation, but their influence fades as $N$ grows. Identification is nearly weak in the sense of \citet{Antoine2021GMMIdentification}. The GIV estimator remains consistent and asymptotically normal, now at the slower rate $\sqrt{T}/N^{\delta}$, where $\delta = \min(1 - 1/\mu, 1/2)$. The cross-sectional signal is diluted, but enough time periods recover it.

When units are comparable in size and none stands out, the instrument is weak. This is the case $\mu > 2$ with $N/T \to c$. No unit is systematically larger than the rest, so the granular variation the instrument relies on vanishes, and time-series information no longer compensates for the cross-sectional dilution. The estimator is inconsistent and identification is weak in the sense of \citet{Staiger1997InstrumentalInstruments}.

For inference, Wald confidence intervals are reliable outside the weak regime, that is when $\mu < 2$. When $\mu > 2$, the instrument can be arbitrarily weak. Its strength carries no guaranteed lower bound as $N$ and $T$ grow, so the normal approximation is unreliable in finite samples. I therefore recommend inverting the Anderson--Rubin test of \citet{Anderson1949EstimatorsEquations}. Its $\chi^2_1$ limiting distribution holds no matter how weak the instrument, so the recommendation remains valid regardless of the $N/T$ trajectory and covers the weak regime as a special case.

The results above assume that the idiosyncratic shocks $u_{it}$ entering the GIV are
known. In practice, we do not observe them and must estimate them by removing the factor structure from the panel. When $N$ is fixed, we can consistently estimate only the common factors. Consistency also requires that the idiosyncratic shocks $u_{it}$ are homoskedastic. With
$N,T \to \infty$, however, we consistently recover both factors and loadings. I show
that the feasible GIV, constructed from estimated residuals, attains the same
convergence rate as the infeasible instrument across all regimes, subject to the
additional growth restriction $\sqrt{T}/N \to 0$. The asymptotic variance is not the
same. The first-stage estimation contributes an additional term at the same order as the
infeasible variance, so inference must use standard errors that account for this
generated-regressor contribution.

In the empirical application, copper and natural gas fall into the strong instrument regime at $95\%$ confidence, while crude oil extends into the nearly weak regime. GIV corrects the biased OLS estimates and leads
to economically plausible negative demand elasticities: $-0.135$ for copper,
$-0.109$ for crude oil, and $-0.056$ for natural gas.

Finally, in simulation exercises calibrated to the panel data from copper and crude oil, I study the sensitivity of the GIV estimates to granularity. As expected, the strong regime has very low bias, tight confidence sets, and the correct coverage. In the weak regime with $\mu > 2$, we observe substantial bias together with confidence intervals that explode in length, so that the Wald interval over-covers rather than attaining its nominal level.

I organize the rest of the paper as follows. In Section~\ref{sec:three_regimes}, I set up the model and formalize granularity through power-law size distributions. I present the main asymptotic results when the factor structure is observed and when it is unobserved in Sections \ref{section:infeasible giv} and \ref{section:feasible giv} respectively. In Section~\ref{sec:empirical}, I apply the theory to estimate short-run demand elasticities for major commodities. I study the small-sample behavior of the estimator in Section~\ref{sec:simulation} through Monte Carlo simulations and conclude in Section~\ref{sec:conclusion}. I close this section by placing the paper in the context of the related literature.

\subsection{Related Literature}

This paper contributes to several strands of the literature.

\textbf{GIV theory.} \citet{Gabaix2024GranularVariables} introduce the GIV framework under fixed $N$ and $T \to \infty$, establishing consistency, asymptotic normality, and the theory for the estimator.

\citet{Banafti2022InferentialDimensions} extend GIV to high dimensions,
allowing $N$ to grow with $T$, and derive the asymptotic distribution of
the feasible estimator under the same power law assumption I impose in
Assumption \ref{ass:size_of_firm_power_law}. They restrict attention to
the strong instrument regime ($\mu < 1$), and the asymptotic distribution
they obtain differs from the one I derive in this paper. Their derivation
requires an additional condition---their Assumption 4(iii)---on the
behavior of shares as $N \to \infty$. In Section \ref{subsec:banfti},
I show that this condition is incompatible with the power law assumption:
the two cannot simultaneously hold. Their proof therefore breaks down
in the power law setting, and their asymptotic distribution result does
not apply. Establishing the asymptotic distribution of the GIV estimator
in large panels, with both $N$ and $T$ tending to infinity, therefore
remains an open problem and this paper resolves that.

\citet{Qian2023Heterogeneity-robustInstruments} constructs heterogeneity-robust granular instruments that remain valid when the structural parameter varies across a fixed number of units. \citet{Baumeister2023AVariables} propose a full-information approach that jointly estimates the factor structure and structural parameter under parametric assumptions. This method becomes untenable as $N \to \infty$.

On the empirical side, GIV has been applied to asset markets \citep{Chodorow-Reich2021AssetInsulators}, bank credit risk \citep{Galaasen2020GranularRisk}, bank lending \citep{Ma2021ExpectationsLending}, sovereign bonds \citep{Aldasoro2023TheMarkets}, exchange rates \citep{Hau2022GlobalRates}, and monetary policy \citep{Holm-Hadulla2024GranularPolicy}. These applications involve $N$ ranging from 21 to well over 100, underscoring the need for a large-panel theory.

\textbf{Factor models.} The feasible GIV requires estimating idiosyncratic shocks by removing the common factor structure. \citet{Bai2003InferentialDimensions} establishes the convergence rates for factors and loadings estimated by principal components when $N, T \to \infty$. I use these results to show that the estimation error affects the asymptotics of the GIV.  We can determine the number of common factors using information criteria such as those of \citet{Bai2002DeterminingModels} or \citet{Ahn2013EigenvalueFactors}.

\textbf{Power laws.} Power-law size distributions are well documented in firm sizes \citep{Axtell2001ZipfSizes}, city sizes, and financial returns \citep{Gabaix2009PowerFinance}. I build on this regularity. The same power-law tail that drives granularity also determines whether the GIV is a strong or nearly weak instrument.

\textbf{Instrument strength.} \citet{Staiger1997InstrumentalInstruments} formalize weak instruments by modeling the first-stage coefficient as local to zero. In my setting, weakness is structural rather than local-to-zero. It is comparable to instrument weakness in large markets for differentiated products as in \citet{Armstrong2016LargeSupply}. \citet{Antoine2021GMMIdentification} provide a nearly weak identification framework where identification strength vanishes, but slowly enough for consistency at a rate slower than $\sqrt{T}$. My nearly weak regime ($\mu > 1$) maps directly onto their framework. The local-to-zero asymptotics of \citet{Staiger1997InstrumentalInstruments} emerge as a special case under $N/T \to c$ with $\mu > 2$. For inference when $\mu > 2$, I construct Anderson--Rubin confidence sets \citep{Anderson1949EstimatorsEquations} that remain valid regardless of instrument strength.


\section{Three Regimes of Instrument Strength} \label{sec:three_regimes}

This section studies when the Granular Instrumental Variable (GIV) is strong
enough to identify an aggregate demand elasticity. The instrument $z_t$ is a
share-weighted average of unit-level supply shocks. It is valid by assumption
(Assumption~\ref{ass:weak stationarity and idiosyncrasy} delivers exogeneity),
but its \emph{relevance} is not guaranteed. Informativeness depends on the
presence of dominant units. For enough individuals, the share $S_i$ must sit far
from $1/N$.

Proposition~\ref{prop:weakness of infeasible instrument} formalises this. The
instrument stays fixed only when unit sizes are drawn from a fat-tailed
distribution. Otherwise it decays at a rate set by the tails. Weakness is
therefore not something I impose through a local-to-zero parameterization. It
arises structurally, from the heavy-tailed distribution of individual sizes.

The rate of decay sorts the design into three regimes. When the instrument does
not decay, it is strong. When it decays slower than $\sqrt{N}$, it is nearly
weak. When it decays at the $\sqrt{N}$ rate, it is weak in the classical sense of
\citet{Staiger1997InstrumentalInstruments}. These regimes govern whether the
elasticity can be estimated consistently, at what rate, and whether standard
inference stays reliable. I develop those consequences for the known factor
structure in Section~\ref{section:infeasible giv} and for the unknown factor
structure in Section~\ref{section:feasible giv}. I begin by stating the model and
the assumptions behind Proposition~\ref{prop:weakness of infeasible instrument}.


\subsection{Model} \label{subsec:model}

I study the estimation of aggregate demand elasticities for a commodity. At each date $t = 1, \ldots, T$, change in aggregate demand $d_t$ is governed by the structural equation
\begin{equation}
    d_t    = \phi_d\, p_t + X_t^d + \varepsilon_t \label{eq:demand}
\end{equation}
where $p_t$ is the change in market-clearing price, $\phi_d$ is the demand elasticity, and $X_t^d$ are observed controls uncorrelated with $\varepsilon_t$. The structural parameter of interest is the demand elasticity, $\phi_d$.

Further, we observe changes in individual level supply/production of the commodity, $y_t = \left( y_{it} \right)_{1 \leq i \leq N}$. The individual supply is governed by the structural equations
\begin{align}
    y_{it} &= \phi_s\, p_t + X_{it}^y + \lambda_i' F_t + u_{it}, \label{eq:supply}
\end{align}
where $\phi_s$ is the supply elasticity, and $X_{it}^y$ are observed controls uncorrelated with both $\varepsilon_t$ and $u_t$. The unit-level supply in \eqref{eq:supply} follows a panel data model with interactive fixed effects (common shocks $F_t$ that load heterogeneously across units through $\lambda_i$). The factors $F_t$, the loadings $\lambda_i$, and the idiosyncratic shocks $u_{it}$ are all unobserved. Stacking across units,
\begin{equation*}
    y_t = e_N \phi_s p_t + X_t^y + \Lambda F_t + u_t,
\end{equation*}
where $e_N$ is the $N$-vector of ones.

Market clearing links the two equations: the aggregate change in supply equals the change in demand, so $d_t = \sum_{i=1}^N S_i y_{it} \defeq y_{St}$, where $S_i$ is the equilibrium market share of unit $i$. We assume that the equilibrium market shares are determined by a different structural process, so that at the frequency of interest $S_i$ is independent of all changes in demand and supply. The market share is linked to the individual sizes, $\s_i$ as
 $S_i = \frac{\s_i}{\sum_j \s_j}$.


For ease of exposition, I present the theory without accounting for the observed controls. When such controls are present, we can partial them out; mutatis mutandis, the theory applies equally to the resulting residualized variables by the Frisch-Waugh-Lovell theorem. See Appendix \ref{appendix_sec_controls} for details.

The model is characterized by the following assumption.
\begin{Ass} \label{ass:weak stationarity and idiosyncrasy}
The multivariate time series $\left(p_{t},y_{t}^{\prime },u_{t}^{\prime }, \varepsilon_t' \right) ^{\prime }$\ is a weakly stationary process with finite second moments. The vector $u_{t}$\ of error terms has a zero mean and is idiosyncratic in the sense that $\e[F_t u_t] = 0$, and
\begin{equation}
E[(y_{St}-\phi_d p_{t})u_{t}] =0  \label{eq:conditional_moment_restriction}
\end{equation}



     where $e_N$ is the N-dimensional vector of ones. Without loss of generality, we assume that the first common factor in the interactive fixed effects structure of the error is a time fixed effect:
\[
\lambda _{i}^{\prime }F_{t}+u_{it}=F_{1t}+\sum_{k=2}^{r}\lambda
_{ik}F_{kt}+u_{it}
\]
In matrix form:
\[
\Lambda =\left[
\begin{array}{ccc}
\Lambda ^{1} & ... & \Lambda ^{r}
\end{array}
\right],\quad \Lambda ^{1}=e_{N}.
\]
Without loss of generality, we also assume that the columns $
\Lambda ^{k},k=2,...,r,$ are orthogonal to the first column, that is that they have a zero mean. Call the $N \times r-1$ matrix formed by dropping the first column, $\tilde{\Lambda}$. Define $\tilde F_t = (F_{2t}, \ldots, F_{rt})'$
\end{Ass}

By Assumption \ref{ass:weak stationarity and idiosyncrasy}, $u_t$ is a valid instrument for the estimation of $\phi_d$. The moment condition delivers exclusion; whether $u_t$ is \emph{relevant} enough to identify $\phi_d$ as $N \to \infty$ is the subject of Proposition~\ref{prop:weakness of infeasible instrument}.

We do not observe $u_t$. In small panels, we can consistently estimate $\Lambda$ (see GK and GNR). As $T \to \infty $ (infeasible), we can extract from data, $M_{\Lambda}y_t = M_{\Lambda}u_t$, where
\begin{align*}
    P_{\Lambda } &= \Lambda \left( \Lambda ^{\prime }\Lambda \right) ^{-1}\Lambda' \enspace \text{and,} \\
    M_{\Lambda } &= I_N - P_{\Lambda }
\end{align*}

GK and GNR show that the optimal GIV in the case of linear conditional expectation is given by $z_t = S' M_{\Lambda} y_t = S' M_{\Lambda} u_t$. In large panels ($N$ and $T$ large), we can go further and estimate both the common factors and the factor loadings. We can demean (which kills $F_{1t}$ via $D_N e_N = 0$) and estimate consistent $\hat\Lambda, \hat F$ by PCA. We then subtract the estimated common component $\hat C_t = \hat{\Lambda} \hat{F}_t$ from the observed data to get $\hat{u}_t = D_N y_t - \hat{C}_t$. As $T \to \infty$ (infeasible), we have
\begin{equation*}
    D_N y_t - \tilde{C}_t = D_N u_t
\end{equation*}
where $\tilde{C}_t= \tilde{\Lambda} \tilde{F}_t$ and $D_N = I_N - \frac{e_N e_N'}{N}$ is the demeaning matrix with $e_N$ being the $N$-dimensional vector of ones. For any $N \times K$ matrix $X$, I write $\bar X \defeq D_N X$ for its demeaned version. Large panels allow consistent estimation of common factors, leading to the moment restriction
\begin{equation}
E[(y_{St}-\phi_d p_{t}) (D_N y_t - C_t)] =0  \label{eq:conditional_moment_aggregate}
\end{equation}
The Granular Instrumental variable associated with the above moment condition is
\begin{equation}
    z_t = S' (D_N y_t - C_t)= S'D_N u_t
\end{equation}
In the shorthand just introduced, the infeasible instrument is $z_t = S' \bar u_t$. In the general case, we do not directly observe $C_t$, and it needs to be estimated. For clarity, I will first develop the asymptotic theory for the infeasible instrument before presenting the theory for the feasible one.
To account for granular settings, we assume that the share vector, $S$ is random and follows the power law. We similarly make suitable assumptions on the other cross-sectional variable, the factor loadings. We assume that the cross-section $(\s_i, \tilde\Lambda_i)$ is i.i.d., sizes have power-law tails, loadings have bounded fourth moments, and both are independent of the time-series shocks.

\begin{Ass} \label{ass:size_of_firm_power_law}
    The absolute sizes of individual units, $\s_i$ satisfy the following conditions:
    \begin{enumerate}
        \item The absolute sizes of individual units, $\s_i$ are drawn from an arbitrary distribution whose tail follows a power law.
    That is, the probability that it is above a fixed threshold, $s_i$ is given by
    \begin{equation*}
        \p(\s_i > s_i) = c s_i^{-\mu}
    \end{equation*}
    \item The absolute sizes are independent of all time series shocks, namely $u_t, F_t$, and $\varepsilon_t$.
    \begin{equation*}
        \s_i \perp (u_t, F_t, \varepsilon_t) \quad \forall i,t
    \end{equation*}
    \item The absolute sizes are independent of the factor loadings in the shocks, i.e., $\s_i \perp \tilde{\Blambda}_i$ and the $r$-dimensional vector, $(\s_i, \tilde{\Blambda}_i')'$ is independent across $i$, and identically distributed.
    \item The factor loadings are such that they are independent of the time series shocks. That is
    \begin{equation*}
        \tilde{\Blambda}_i \perp (u_t, F_t, \varepsilon_t) \quad \forall i,t
    \end{equation*}
    The tails are also bounded. That is, $\e \| \tilde{\Blambda}_i \|^4 < \infty$.
    \end{enumerate}
\end{Ass}
Independence of the absolute sizes is not a restrictive assumption. The individual size is set in the long-term equilibrium. The shocks are all short-term in nature, and do not affect the long-term equilibrium. The same applies to the joint i.i.d.\ assumption of the vector, $(\s_i, \tilde{\Blambda}_i')'$. When loadings are treated as non-random (as is common), the i.i.d. clause reduces to i.i.d. sizes. The final assumption is standard in the factor literature when the factor loadings are random \citep{Bai2003InferentialDimensions}.
As sizes are observed in equilibrium, we can construct the individual shares as
\begin{equation*}
    S_i = \frac{\s_i}{\sum_{j=1}^N \s_j}
\end{equation*}

\begin{Ass} \label{ass:bounded eigenvalues}
    The eigenvalues of the variance of the idiosyncratic errors are bounded above and bounded away from zero. That is, defining $\e[u_t u_t'] = \Omega$, there exist constants $0 < \underline\lambda \leq K < \infty$, independent of $N$, such that
    \begin{equation*}
        \underline\lambda \;\leq\; \gamma_{\text{min}}(\Omega) \;\leq\; \gamma_{\text{max}}(\Omega) \;\leq\; K,
    \end{equation*}
    where $\gamma_{\text{min}}(\cdot)$ and $\gamma_{\text{max}}(\cdot)$ denote the smallest and largest eigenvalue operators, respectively.
\end{Ass}

Under these assumptions, we will see how the tail index affects the granularity of the setting and hence the instrument strength.

\subsection{Structural Origin of Weakness}
In large panels, weakness arises due to the behavior of the tail index of the size variable. I formally state that idea in the first Proposition. \cite{Banafti2022InferentialDimensions} had formally stated the result for $\mu \in (0,1)$. I extend it to all values of $\mu$, following \citet{Gabaix2011TheFluctuations}.

\begin{Prop} \label{prop:weakness of infeasible instrument}
    Suppose Assumptions~\ref{ass:weak stationarity and idiosyncrasy}, \ref{ass:size_of_firm_power_law}, and \ref{ass:bounded eigenvalues} hold. Then
    \begin{equation*}
        z_t =
        \begin{cases}
            O_\p(1) & \mu \in (0,1), \\
            O_\p\!\left(\frac{1}{N^{1 - \frac{1}{\mu}}}\right) & \mu \in (1,2), \\
            O_\p\!\left(\frac{1}{\sqrt{N}}\right) & \mu > 2.
        \end{cases}
    \end{equation*}
\end{Prop}
\begin{proof}
Proof in Appendix \ref{app:behavior of herfindahl}.
\end{proof}

Proposition \ref{prop:weakness of infeasible instrument} delineates three regimes of instrument strength, governed jointly by the tail index $\mu$ and the relative growth of $N$ and $T$.

For $\mu \in (0,1)$, the instrument does not decay. Identification is strong, and this corresponds to the strong instrument dynamics of \citet{Gabaix2024GranularVariables}.

For $\mu > 1$, the instrument vanishes as $N \to \infty$. For $\mu \in (1,2)$, $z_t = O_{\p}(\frac{1}{N^{1 - 1/\mu}})$. For $\mu > 2$, $z_t = O_{\p}(\frac{1}{\sqrt{N}})$. Under $N/T \to 0$, instrument strength accumulates fast enough to identify the structural parameter and conduct inference. The estimator is consistent and asymptotically normal at the slower rate $\sqrt{T}/N^{\delta}$, where $\delta = \min(1 - 1/\mu, 1/2)$. Following \citet{Antoine2021GMMIdentification}, I call this nearly weak identification.

When $\mu > 2$ and $N/T \to c > 0$, time-series information no longer overtakes the cross-sectional dilution of the instrument. The estimator is inconsistent and identification is weak in the sense of \citet{Staiger1997InstrumentalInstruments}.

Within the nearly weak regime, $\mu = 2$ is a boundary for the concentration parameter. For $\mu \in (1,2)$, the concentration parameter has a polynomial floor in $N$. For $\mu > 2$, the floor is only slowly divergent and can grow arbitrarily slowly along admissible $(N,T)$ sequences. The Gaussian approximation is therefore reliable for $\mu < 2$ but not for $\mu > 2$, where I recommend Anderson--Rubin confidence sets. The same Anderson--Rubin procedure remains valid under the weak identification regime, so the recommendation handles both $\mu > 2$ subcases at once.

Hence weakness arises directly from the structure of the setting, similar to \citet{Armstrong2016LargeSupply}. The heavy-tailed concentration of shares delivers an idea of weakness without imposing the local-to-zero assumption of \citet{Staiger1997InstrumentalInstruments}. That assumption emerges only as a special case under $N/T \to c$ with $\mu > 2$.

We will first consider the infeasible estimation of the structural parameters. The estimation is infeasible as we assume that the factor structure is available to us. This helps us fix the basic ideas. And in the subsequent section, we will deal with the feasible estimation where we need to estimate the factor structure. Throughout, define $\delta = \min(1 - 1/\mu, 1/2)$, so that $\delta = 1 - 1/\mu \in (0, 1/2)$ when $\mu \in (1,2)$ and $\delta = 1/2$ when $\mu > 2$. Proposition \ref{prop:weakness of infeasible instrument} can then be written compactly as $z_t = O_{\p}(N^{-\delta})$ for $\mu > 1$.





\section{GIV with Known Factor Structure} \label{section:infeasible giv}

In this section, I assume that we know the factor structure, specifically the factor loadings. This is primarily for exposition but includes some settings of practical interest. One such case is when we have only time fixed effect, that is, $\Lambda = e_N$. In this case, we can perfectly recover the demeaned idiosyncratic shocks to construct the instrument.
\begin{equation*}
    y_t = e_N \phi_s p_t + e_N F_t  + u_t
\end{equation*}
$D_N y_t = D_N u_t$ perfectly recover the demeaned idiosyncratic shocks.

Another case is when we do not directly observe the loadings, but we have a parametric form, $\tilde{\lambda}_i = X_i \dot{\lambda}$ where $\tilde{\lambda}_i$ and $X_i$ are $r-1$ dimensional vectors and $\dot{\lambda}$ is a $r \times r$ matrix which is invariant across $i$. This implies $\tilde{\lambda} = X \dot{\lambda}$. We observe the vector of characteristics, $X$. In this case,
\begin{align*}
    y_t &= e_N \phi_s p_t + \Lambda F_t + u_t \\
    \quad D_N y_t &= \tilde{\Lambda} \tilde{F}_t + D_N u_t = X \dot{\lambda} \tilde{F}_t + D_N u_t \\
    M_X D_N y_t &= M_X D_N u_t
\end{align*}
Thus, $M_X D_N y_t$ recovers the demeaned idiosyncratic shocks and the optimal instrument is $S'M_X D_N u_t$.

However, most of the empirical examples of GIV assumes a more general factor structure. This requires estimation of the factor structure in a first stage before we construct the instrument. See Section \ref{section:feasible giv} for the analysis of this general case.

From the demeaned idiosyncratic shocks, we construct the infeasible instrument as $z_t = S'(D_N y_t - \tilde{C}_t)$. Proposition \ref{prop:weakness of infeasible instrument} gives the behavior of the infeasible instrument for different values of the tail index of the size variable. I find three regimes. When $\mu \in (0,1)$, identification is strong and the estimator is $\sqrt{T}$-consistent and asymptotically normal. When $\mu > 1$ and $N/T \to 0$, identification is nearly weak in the sense of \citet{Antoine2021GMMIdentification}. The estimator is consistent and asymptotically normal at the slower rate $\sqrt{T}/N^{\delta}$, where $\delta = \min(1 - 1/\mu, 1/2)$. When $\mu > 2$ and $N/T \to c$, identification is weak in the sense of \citet{Staiger1997InstrumentalInstruments} and the estimator is inconsistent. For inference, Wald confidence intervals are reliable when $\mu < 2$. For $\mu > 2$, I recommend Anderson--Rubin confidence sets, which remain valid regardless of the $N/T$ trajectory.

\subsection{Assumptions}

To establish these results, we require assumptions on the idiosyncratic shocks. The large panel setting allows us to accommodate a richer structure for the time series and cross-sectional dependence than the fixed-$N$ (small-panel) case, which requires i.i.d.\ samples with no cross correlation.

\begin{Ass}(Time and Cross-Sectional Dependence and Heteroskedasticity) \label{ass:LN_time and cross sectional dependence}
There exists a positive constant, $M < \infty$ such that for all $N$ and $T$,
\begin{enumerate}
    \item $\e[u_{it}] = 0$, and $\e |u_{it}|^4 \leq M$
    \item For every $i,j$ and $t$, $\e[u_{it} u_{jt}]$ is bounded.
    Define $\gamma(i,j) = \e \left[ u_i' u_j/T \right] = \e \left[ \frac{1}{T} \sum_{t=1}^T u_{it} u_{jt} \right]$, for all $1\leq i,j \leq N$, and
    \begin{align*}
       &a. \quad  \sum_{i=1}^N | \gamma(i,j) |  \leq M \\
       &b. \quad N^{-1}\sum_{i=1}^N \sum_{j=1}^N | \gamma(i,j) | \leq M
    \end{align*}
    \item Let $\e[u_{it}^2 u_{is}^2] = \tau_{st}$. For all $i$,
    \begin{align*}
        T^{-1} \sum_{t=1}^T \sum_{s=1}^T \tau_{st} \leq M
    \end{align*}
    \item For every $i,j$,
    \begin{equation*}
        \e \left| T^{-\frac{1}{2}} \sum_{t=1}^T \bigg( u_{it} u_{jt} - \e[u_{it} u_{jt}]  \bigg) \right|^4 \leq M
    \end{equation*}
    \item (Uniform Rosenthal-type moment bound.) There exists a constant $C_q < \infty$, independent of $N$, such that for any deterministic $\{a_j\}_{j=1}^N \subset \mathbb{R}$ and $q = 8 + 2\pi$,
    \begin{equation*}
        \e\!\left[ \Big| \sum_{j=1}^N a_j u_{jt} \Big|^{q} \right] \;\leq\; C_q \left[ \Big( \sum_{j=1}^N a_j^2 \, \e u_{jt}^2 \Big)^{q/2} + \sum_{j=1}^N |a_j|^q \, \e |u_{jt}|^q \right].
    \end{equation*}
    This holds, in particular, when $\{u_{jt}\}_{j=1}^N$ are cross-sectionally independent \citep{Rosenthal1970OnVariables}.
    \item (Cross-sectional weak dependence of the products $u_{jt} \varepsilon_t$ and $u_{jt} F_t$.) For the structural shock $\varepsilon_t$ and the common factors $F_t$, let $v_{jt} = u_{jt} \varepsilon_t$ and $w_{jt} = u_{jt} F_t$. Then
    \begin{equation*}
        \frac{1}{NT} \sum_{j=1}^N \sum_{k=1}^N \sum_{t=1}^T \sum_{s=1}^T \big| \cov(v_{jt}, v_{ks}) \big| \leq M, \qquad \frac{1}{NT} \sum_{j=1}^N \sum_{k=1}^N \sum_{t=1}^T \sum_{s=1}^T \big\| \cov(w_{jt}, w_{ks}) \big\| \leq M.
    \end{equation*}
    These are the joint analogue of part 2(b), imposing absolute summability of the covariance arrays of $u_{jt} \varepsilon_t$ and $u_{jt} F_t$ across both time and the cross-section. They are conditions on the covariances themselves and do not require $\varepsilon_t$ or $F_t$ to be independent of $\{u_{jt}\}$. Because $\varepsilon_t$ and $F_t$ are common across units, they add no cross-unit linkage of their own, so the conditions only rule out cross-sectional comovement in the supply shocks driven by $\varepsilon_t$ or $F_t$ strong enough to break the summability.
\end{enumerate}
\end{Ass}

These are standard in the factor literature and relax the independence restriction in the shorter panel GIV literature.

\begin{Ass}(Strong Mixing and Higher Moments)\label{ass:mixing for stationary time series}
The multi-dimensional time series, $\{(F_t', u_t', \varepsilon_t) \}$ is weakly stationary and is a strong mixing sequence of size $-(\frac{2 + \pi}{\pi})$, where $\pi > 0$. In other words, define the sigma algebra,
\begin{equation*}
    \mathcal{F}_{a,b} = \sigma \{ (F_t', u_t', \varepsilon_t); a \leq t \leq b \}
\end{equation*}
and
\begin{equation*}
    \alpha(h) \defeq \sup_{A \in \mathcal{F}_{-\infty, 0}, B \in \mathcal{F}_{h, \infty}} |\p(A \cap B) - \p(A) \p(B) |
\end{equation*}
We have, for some $\pi > 0$,
\begin{equation*}
    \sum_{h=1}^{\infty} \alpha(h)^{\frac{\pi}{2 + \pi}} < \infty
\end{equation*}

Further we assume that the following higher moments exist: $\e \| F_t \|^{8 + 2\pi}, \e | \varepsilon_t|^{8 + 2\pi}$, and $\e | u_{jt}|^{8 + 2\pi}$ for all $j$.

\end{Ass}

With the dependence structure in place, I now turn to the behavior of the infeasible GIV estimator. Proposition~\ref{prop:weakness of infeasible instrument} identifies three regimes governed by the tail index $\mu$ and the $N/T$ trajectory. I analyze them in turn, beginning with the strong regime ($\mu \in (0,1)$), where the Herfindahl does not vanish and the estimator is $\sqrt{T}$-consistent and asymptotically normal.

\subsection{Strong Regime: $\mu \in (0,1)$}

Recall the aggregate demand equation of interest is
\begin{equation*}
    y_{St} = \phi_d p_t + \varepsilon_t
\end{equation*}

The moment condition for the estimation of the demand parameter is $\e[(y_{St} - \phi_d p_t )z_t] = 0$. Thus the GIV estimator of the demand (aggregate) structural parameter is
\begin{align*}
    \hat{\phi_d} &= \frac{z' y_S}{z' p} = \phi_d + \frac{z' \varepsilon}{z' p} \\
    \hat{\phi}_d - \phi_d &= \frac{z' \varepsilon}{z' p} = \frac{\displaystyle \frac{1}{T} \sum_t z_t \varepsilon_t}{\displaystyle \frac{1}{T} \sum_t z_t p_t}
\end{align*}

I will now formally state the results for the strong regime. In the small-panel GIV framework of \citet{Gabaix2024GranularVariables}, $N$ is fixed, the shares and factor loadings are treated as constants, and consistency and asymptotic normality follow from standard IV arguments. Moving to the large-panel setting introduces four complications. First, $N \to \infty$, so the instrument $z_t = S'\bar{u}_t$ is a growing weighted sum whose behavior depends on the concentration of the shares. Second, the shares $S$ are now random, drawn from a power-law distribution, so the Herfindahl $S'S$ is itself a random variable whose order must be established. Third, the factor loadings $\tilde{\Blambda}_i$ are random, and the share-weighted loading $S'\tilde{\Lambda}$ must be shown not to contaminate the instrument. Fourth, the central limit theorem for the numerator $T^{-1/2}\sum_t z_t \varepsilon_t$ is itself non-standard. The summand combines a share-weighted cross-sectional sum with time-series dependence, and its moments must be bounded uniformly in $N$.

When $\mu \in (0,1)$, all four complications are resolved by the heavy tails. Proposition~\ref{prop:weakness of infeasible instrument} shows that $S'S = O_{\p}(1)$, so the instrument does not degenerate as $N$ grows. Proposition~\ref{prop:behavior of S times lambda} shows that $S'\tilde{\Lambda} = O_{\p}(1)$ as well, so the common-factor contamination remains bounded. Corollary~\ref{corollary:central limit theorem} delivers the CLT for the numerator $T^{-1/2}\sum_t z_t \varepsilon_t$. The heavy-tail concentration of $S$ is what keeps its moments bounded uniformly in $N$. Under these conditions, the GIV estimator retains $\sqrt{T}$-consistency and asymptotic normality.

\begin{Theorem} \label{theorem:strong_consistency_aggregate}
    Suppose Assumptions \ref{ass:weak stationarity and idiosyncrasy} to \ref{ass:mixing for stationary time series} hold with $\mu \in (0,1)$. Then, conditional on $S$, for almost every realization of the shares, the GIV estimator for the aggregate structural parameter is consistent and asymptotically normal,
    \begin{equation*}
       \frac{\Gamma_{zp}}{\sqrt{V_{z \varepsilon}(S)}} \cdot \sqrt{T} [\hat{\phi}_d - \phi_d] \xrightarrow{d} \n(0,  1)
    \end{equation*}
     where $\Gamma_{zp}$ is the conditional probability limit of $T^{-1}\sum_t z_t p_t$ given $S$, with explicit form $\Gamma_{zp} = \frac{S' \Sigma_u S}{\phi_d - \phi_s},$ and $V_{z\varepsilon}(S) = \lim_{T \to \infty} \frac{1}{T} \sum_{s=1}^T \sum_{t=1}^T \e[z_t z_s \varepsilon_t \varepsilon_s \mid S]$. These limits exist and $\Gamma_{zp} \neq 0$ for almost every realization of $S$, provided $\phi_d \neq \phi_s$.
\end{Theorem}
\begin{proof}
    Proof in Appendix \ref{app:other proofs}. The proof proceeds by showing that the denominator $T^{-1}\sum_t z_t p_t$ converges to a nonzero limit, while the numerator $T^{-1/2}\sum_t z_t \varepsilon_t$ satisfies a CLT under the mixing conditions of Assumption \ref{ass:mixing for stationary time series}.
\end{proof}

\subsection{Nearly Weak Identification: $\mu > 1$}

I now turn to the case $\mu > 1$. Define $\delta = \min(1 - 1/\mu, 1/2)$. The heavy-tail concentration of the shares is weaker than in the strong regime, and Proposition~\ref{prop:weakness of infeasible instrument} pins down by how much. The Herfindahl satisfies $S'S = O_{\p}(N^{-2\delta})$, so the instrument $z_t = S'\bar{u}_t$ no longer has order one. It dilutes as $N$ grows. Only the rescaled instrument $N^{\delta} z_t$ has a non-degenerate limit. Under $N/T \to 0$, time-series information accumulates fast enough that the GIV estimator remains consistent and asymptotically normal at the slower rate $\sqrt{T}/N^{\delta}$. Following \citet{Antoine2021GMMIdentification}, I call this nearly weak identification.

Two of the four complications from the strong regime resolve exactly as before. The share-weighted loading $S'\tilde{\Lambda}$ remains controlled by Proposition~\ref{prop:behavior of S times lambda}, and the CLT for the numerator goes through under the same mixing conditions. The other two now carry an explicit $N$-dependence inherited from the dilution of the instrument. This is what slows the rate of convergence and forces the requirement $N/T \to 0$. Time-series information must accumulate fast enough to overcome the cross-sectional dilution.

The boundary $\mu = 2$ separates two qualitatively different dilution patterns. For $\mu \in (1,2)$, the dilution exponent $\delta = 1 - 1/\mu$ lies in $(0, 1/2)$ and the Herfindahl is governed by a stable law. For $\mu > 2$, the size variable has finite second moment, so the Herfindahl is governed by the law of large numbers rather than a stable law, and the dilution exponent caps at $\delta = 1/2$. The asymptotic statement of the theorem below covers both cases at once. What changes across the boundary is the behavior of the concentration parameter, which in turn drives the choice between Wald and Anderson--Rubin inference.

\begin{Theorem} \label{theorem:weak_consistency_aggregate} \label{theorem:semi_strong_consistency_aggregate}
    Suppose Assumptions \ref{ass:weak stationarity and idiosyncrasy} to \ref{ass:mixing for stationary time series} hold with $\mu > 1$ and $N/T \to 0$. Let $\delta = \min(1 - 1/\mu, 1/2)$, so that $\delta = 1 - 1/\mu \in (0, 1/2)$ when $\mu \in (1,2)$ and $\delta = 1/2$ when $\mu > 2$. Then, conditional on $S$, for almost every realization of the shares, the GIV estimator for the aggregate structural parameter is consistent and asymptotically normal,
    \begin{equation*}
       \frac{\Gamma_{zp}}{\sqrt{V_{z \varepsilon}(S)}} \cdot \sqrt{T}\, [\hat{\phi}_d - \phi_d] \xrightarrow{d} \n(0,  1),
    \end{equation*}
    where $\Gamma_{zp}$ is the conditional probability limit of $\frac{1}{T}\sum_t z_t p_t$ given $S$, with explicit form $\Gamma_{zp} = \frac{S'\Sigma_u S}{\phi_d - \phi_s}$, and $V_{z\varepsilon}(S) = \lim_{T \to \infty} \frac{1}{T}\sum_{s=1}^T\sum_{t=1}^T \e[z_t z_s \varepsilon_t \varepsilon_s \mid S]$. These limits exist and $\Gamma_{zp} \neq 0$ for almost every realization of $S$, provided $\phi_d \neq \phi_s$. The studentization satisfies $\Gamma_{zp}/\sqrt{V_{z\varepsilon}(S)} = O_{\p}(N^{-\delta})$, so the standardized statistic vanishes at the rate $\sqrt{T}/N^{\delta}$.

    If $\mu > 2$ and $N/T \to c > 0$, then $\hat{\phi}_d$ is inconsistent and identification is weak in the sense of \citet{Staiger1997InstrumentalInstruments}.
\end{Theorem}
\begin{proof}
    Proof in Appendix~\ref{app:other proofs}. The argument parallels the strong regime, with the rate now $N$-dependent. By Proposition~\ref{prop:weakness of infeasible instrument}, both $\Gamma_{zp}$ and $V_{z\varepsilon}(S)$ are $O_{\p}(N^{-2\delta})$. Combining numerator and denominator gives $\hat\phi_d - \phi_d = O_{\p}(N^{\delta}/\sqrt{T})$. The CLT for the numerator follows from Corollary~\ref{corollary:central limit theorem} under the mixing of Assumption~\ref{ass:mixing for stationary time series}. Consistency requires $N^{2\delta}/T \to 0$, which is implied by $N/T \to 0$ since $2\delta \leq 1$. The $\mu > 2$ failure mode under $N/T \to c$ falls out of the same algebra. There, $\hat\phi_d - \phi_d = O_{\p}(\sqrt{N/T}) = O_{\p}(1)$, so the estimator is inconsistent.
\end{proof}

\begin{remark}[Three regimes] \label{rem:three_regimes_aggregate}
Theorems~\ref{theorem:strong_consistency_aggregate} and~\ref{theorem:weak_consistency_aggregate} together encompass three regimes of instrument strength, governed by the tail index $\mu$ and the $N/T$ trajectory.
\begin{enumerate}
    \item \textbf{Strong identification} ($\mu \in (0,1)$). The instrument does not dilute, and Theorem~\ref{theorem:strong_consistency_aggregate} gives $\sqrt{T}$-consistency and asymptotic normality.
    \item \textbf{Nearly weak identification} ($\mu > 1$ with $N/T \to 0$). The instrument dilutes at rate $N^{-\delta}$ with $\delta = \min(1 - 1/\mu, 1/2)$, but $T$ outgrows $N$ fast enough to preserve identification. Theorem~\ref{theorem:weak_consistency_aggregate} gives consistency and asymptotic normality at the slower rate $\sqrt{T}/N^{\delta}$.
    \item \textbf{Weak identification} ($\mu > 2$ with $N/T \to c > 0$). $T$ no longer outgrows $N$ and the concentration parameter is $O_{\p}(1)$. The estimator is inconsistent, in the sense of \citet{Staiger1997InstrumentalInstruments}. This is the failure-mode statement at the end of Theorem~\ref{theorem:weak_consistency_aggregate}.
\end{enumerate}
\end{remark}

\subsubsection{Inference} \label{subsec:infeasible_AR}

The choice between Wald and Anderson--Rubin inference depends on the behavior of the concentration parameter. The concentration parameter $\kappa^2_{\text{conc}}$ measures the signal-to-noise ratio of the first-stage moment \citep{Stock2002AMoments}. For the GIV estimator,
\begin{equation*}
    \kappa^2_{\text{conc}} \;=\; \frac{T \, \Gamma_{zp}^2}{V_{z\varepsilon}(S)},
\end{equation*}
and the asymptotic standardization $\kappa_{\text{conc}}(\hat\phi_d - \phi_d) \xrightarrow{d} \n(0,1)$ summarizes the rate at which the estimator concentrates on $\phi_d$. The Gaussian approximation is reliable in finite samples when $\kappa^2_{\text{conc}}$ is large. The first-stage $F$-statistic is its sample analog.

Parametrize any admissible sequence as $T = N \cdot f(N)$ with $f(N) \to \infty$. Then
\begin{equation*}
    \kappa^2_{\text{conc}} \;=\; \begin{cases}
        N^{1-2\delta} \cdot f(N) & \mu \in (1,2), \\
        f(N) & \mu > 2.
    \end{cases}
\end{equation*}
For $\mu \in (1,2)$ the polynomial floor $N^{1-2\delta}$ with $1 - 2\delta > 0$ does not depend on $f$. So $\kappa^2_{\text{conc}}$ clears any fixed weak-instrument threshold for moderate $N$. For $\mu > 2$ the floor is just $f(N)$, which can grow arbitrarily slowly. The boundary $\mu = 2$ is precisely where the polynomial floor disappears.

\paragraph{Wald inference for $\mu \in (1,2)$.} The polynomial floor in $\kappa^2_{\text{conc}}$ keeps Wald inference reliable. The practitioner does not need to know $\delta$. The asymptotic variance is
\begin{equation*}
    \widehat{\mathrm{Avar}}(\hat{\phi}_d) \;=\; \frac{1}{T}\cdot\frac{\hat{V}_{z\varepsilon}}{\hat{\Gamma}_{zp}^{2}},
\end{equation*}
where $\hat{\Gamma}_{zp}$ and $\hat{V}_{z\varepsilon}$ are sample analogs of $\Gamma_{zp}$ and $V_{z\varepsilon}(S)$. The studentization in Theorem~\ref{theorem:weak_consistency_aggregate} absorbs the rate automatically, so $\delta$ never enters the formula. This parallels GMM under near-weak identification \citep{Antoine2021GMMIdentification}. Standard Wald confidence intervals remain valid, even though the rate is slower than in the strong regime.

\paragraph{Anderson--Rubin inference for $\mu > 2$.} The polynomial floor disappears. The concentration parameter can grow arbitrarily slowly, so its realized value in finite samples need not be large. The estimator has poor finite-sample performance even though it is asymptotically normal. I recommend inverting the Anderson--Rubin (AR) test of \citet{Anderson1949EstimatorsEquations} to construct the confidence interval.

Anderson--Rubin avoids the dependence on $\kappa^2_{\text{conc}}$ entirely. At hypothesized value $\phi_0$, the sample moment is
\begin{equation*}
    g_T(\phi_0) \;=\; \frac{1}{T}\sum_{t=1}^T z_t (y_{St} - \phi_0 p_t),
\end{equation*}
which under $H_0: \phi_d = \phi_0$ reduces to $T^{-1}\sum_t z_t \varepsilon_t$ and, by Theorem~\ref{theorem:central limit theorem}, satisfies $\sqrt{T}\,g_T(\phi_0)/\sqrt{V_{z\varepsilon}(S)} \xrightarrow{d} \n(0,1)$. The AR statistic and the corresponding $1-\alpha$ confidence set are
\begin{equation*}
    AR_T(\phi_0) \;=\; \frac{T \cdot g_T(\phi_0)^2}{\hat V_{z\varepsilon}(S)} \;\xrightarrow{d}\; \chi^2_1, \qquad
    \mathcal{C}^{AR}_{1-\alpha} \;=\; \big\{\phi_0 \in \mathbb{R} : AR_T(\phi_0) \leq \chi^2_{1, 1-\alpha}\big\}.
\end{equation*}
The $\chi^2_1$ limit holds uniformly across $\mu > 2$ and $N/T \to 0$. It requires no condition on the rate at which $\kappa^2_{\text{conc}} \to \infty$. It needs only a consistent estimator of $V_{z\varepsilon}(S)$, the variance of the moment.

The same AR procedure remains valid when $N/T \to c$ and identification is weak in the Staiger--Stock sense. AR does not depend on consistency of $\hat\phi_d$. It uses only the sample moment evaluated at the hypothesized value, so it covers the failure mode without modification.

\section{Feasible GIV with unknown Factor Structure} \label{section:feasible giv}

Section \ref{section:infeasible giv} took the factor structure as known. In practice it is not, and we replace the infeasible common component $\tilde C_t$ with the principal-component estimate $\hat C_t$ of \citet{Bai2003InferentialDimensions}. This places GIV in the constructed-regressor setting: the first-stage estimation error propagates to the GIV moment. Additionally, consistent inference now requires $\sqrt{T}/N \to 0$. This section formalises the effects of the first stage estimation and re-establishes the convergence results of Section \ref{section:infeasible giv} under the feasible instrument.


We are interested in estimating the structural parameter $\phi_d$ in:
\begin{align*}
    y_{St} = \phi_d p_t + \varepsilon_t
\end{align*}
We need to construct the instrument from the supply side equation
\begin{equation*}
    y_t = e_N \phi_s p_t + \Lambda F_t + u_t
\end{equation*}

From the previous definitions, $\Lambda = [\boldsymbol{1}_{N} \enspace \tilde{\Lambda} ]$ and $F_t = [F_{1t},\, \tilde F_t']'$,
\begin{equation*}
    D_N y_t = \tilde{\Lambda} \tilde{F}_t + D_N u_t
\end{equation*}

Thus $D_Ny_t$ has a factor structure. In large panels, we can consistently estimate both $\tilde{\Lambda}$ and $\tilde{F}$ \citep{Bai2003InferentialDimensions}. As $T > N$, we estimate $\hat{\Lambda}$ using principal components. The first order condition of the PCA objective concentrates out $\hat{F} = \Bar{Y} \hat{\Lambda}/N$.

From these two estimators, we have the estimator for the common component $\tilde{C}_{it} = \tilde{\Blambda}_i' \tilde{F}_t$ as $\hat{C}_{it} = \hat{\Blambda}_i' \hat{F}_t$. Call the corresponding vector, $\hat{C}_t$. From this estimator, we can construct another feasible instrument:
\begin{equation*}
    \hat{z}_t' = S'[D_N y_t - \hat{C}_t] \defeq S' \hat{u}_t
\end{equation*}
where $\hat{u}_t = D_N y_t - \hat{C}_t$ is the estimated residual. \citet{Gabaix2024GranularVariables} proposed a different form of the instrument.
\begin{equation*}
    \hat{z}^{\text{GK}}_t = S' M_{\hat{\Lambda}} D_N y_t
\end{equation*}

The two formulations are equivalent: the PCA first-order condition $\hat F_t = \hat\Lambda' D_N y_t / N$ implies $\hat C_t = (I - M_{\hat\Lambda}) D_N y_t$, so $\hat z_t = \hat z_t^{\text{GK}}$ \footnote{Explicitly, $\hat u_t = D_N y_t - \hat\Lambda \hat F_t = [I - \hat\Lambda \hat\Lambda'/N] D_N y_t = [I - \hat\Lambda(\hat\Lambda'\hat\Lambda)^{-1}\hat\Lambda'] D_N y_t = M_{\hat\Lambda} D_N y_t$, where the second equality uses the PCA normalization $\hat\Lambda'\hat\Lambda/N = I_{r-1}$.}. I work with $\hat C_t$ rather than $M_{\hat\Lambda} D_N y_t$ in the asymptotic analysis. $M_{\hat\Lambda} = \hat\Lambda(\hat\Lambda'\hat\Lambda)^{-1}\hat\Lambda'$ is a non-linear product of the estimator $\hat\Lambda$. Hence its asymptotic linear form involves derivatives of non-linear transformations of the estimator and is mathematically complex. $\hat C_t$ admits a much cleaner asymptotic linear expansion derived in Appendix~\ref{app:estimation of common components}. The instrument is therefore
\begin{equation*}
    \hat{z}_t = S'[D_N y_t - \hat{C}_t].
\end{equation*}


In this section, we will see that the estimation of the factor structure in the first stage places an additional condition on the rates of convergence of $N$ and $T$. Consistency and asymptotic normality of the estimates of the structural parameters using the infeasible instrument for the nearly weak regime in Theorem \ref{theorem:weak_consistency_aggregate} require $\frac{N}{T} \to 0$. But when we use the feasible instrument after estimation of the factor structure, consistency and asymptotic normality require an additional restriction on the rates of $N$ and $T$, which is that $\frac{\sqrt{T}}{N} \to 0$.

Similar to the previous section, I state the results separately for the strong and nearly weak regimes. For $\mu > 2$, I provide Anderson--Rubin confidence sets. The estimation of the factor structure requires a number of additional assumptions on the factor structure which I state in the next sub-section.

\subsection{Assumptions} \label{subsec:assumptions}
\begin{Ass}[Strong Factor Structure and Distinct Eigenvalues] \label{ass:LN_strong factor structure}
The factor structure is strong. That is,
\begin{enumerate}
    \item The factor structure is strong. $\e \| \tilde{F}_t \|^4 \leq M < \infty $ and $T^{-1} \sum_{t=1}^T \Tilde{F}_t \Tilde{F}_t' \xrightarrow{p} \Sigma_{\Tilde{F}}$, for some $r-1 \times r-1$ positive definite matrix, $\Sigma_{\Tilde{F}}$.
    \item $ \| \tilde{\Blambda}_j \| \leq \Bar{\Blambda} < \infty$ and $ \e \| \tilde{\Blambda}_j \|^4 \leq M < \infty$ for every $j$. Each factor has a non-trivial contribution on the variance of $Y_t$. That is, $N^{-1} \sum_{i=1}^N \tilde{\Blambda}_i \tilde{\Blambda}_i' \xrightarrow{p} \Sigma_{\Tilde{\Lambda}}$,, for some positive definite $r-1 \times r-1$ matrix $\Sigma_{{\tilde{\Lambda}}}$. $\tilde{\Blambda}_i$ is independent of $u_t$ and $\tilde{F}_t$ for all $i$ and $t$.
    \item The eigenvalues of the $r-1 \times r-1$ matrix $\Sigma_{\tilde{\Lambda}} \cdot \Sigma_{\Tilde{F}}$ are distinct
\end{enumerate}

\end{Ass}


\begin{Ass}(Weak conditional dependence between factors and idiosyncratic errors) \label{ass:weak dep between agg shocks and idio errors}
There exists an $M < \infty$, such that
    \begin{equation*}
      \e \left(   \frac{1}{N} \sum_{i=1}^N \left \| \frac{1}{\sqrt{T}} \sum_{t=1}^T \tilde{F}_t u_{it} \right \|^2 \right) \leq M
    \end{equation*}
\end{Ass}

\begin{Ass}(Moments) \label{ass:LN_moments and CLT}
There exists an $M < \infty$, such that for all $N$ and $T$,
    \begin{enumerate}
        \item For each $i$,
        \begin{equation*}
            \e \left \| \frac{1}{\sqrt{NT}} \sum_{j=1}^N \sum_{t=1}^T \tilde{\Blambda}_j \left( u_{it} u_{jt} - \e[u_{it} u_{jt}]  \right)   \right \|^2 \leq M
        \end{equation*}
        \item The $r \times r$ matrix satisfies
        \begin{equation*}
            \e \left \| \frac{1}{\sqrt{NT}} \sum_{j=1}^N \sum_{t=1}^T \tilde{\Blambda}_j \tilde{F}_t' u_{jt}  \right \|^2 \leq M
        \end{equation*}
    \end{enumerate}

\end{Ass}






\subsection{Effect of Estimating the Factor Structure}
The proofs adapt \citet{Bai2003InferentialDimensions} to the GIV setting:
Appendix~\ref{app:estimation of factor loadings} treats the factor
loadings, Appendix~\ref{app:estimation of common factors} the factors,
and Appendix~\ref{app:estimation of common components} combines the two
to obtain an influence-function expansion of $\hat C_t - \tilde C_t$.
I quote that expansion here and trace it through the GIV moment.

From \eqref{eq:app_influence function of estimator}, we can see that the difference between the estimated common component and the true common component is
\begin{equation} \label{eq:influence function of estimator}
    \hat{C}_t - \tilde{C}_t = \tilde{F}_t' \left[ \frac{\tilde{F}' \tilde{F}}{T} \right]^{-1} \frac{1}{T} \sum_{m=1}^T \tilde{F}_m \Bar{u}_{m} + \tilde{\Lambda} \left[ \frac{\tilde{\Lambda}' \tilde{\Lambda}}{N} \right]^{-1} \frac{1}{N} \sum_{j = 1}^N \tilde{\Blambda}_j \Bar{u}_{jt} + O_{\p} \left( \frac{1}{N} \right)
\end{equation}

The two leading terms in \eqref{eq:influence function of estimator} have
distinct origins and behave differently in $N$ and $T$. The first term,
$\tilde F_t' (\tilde F'\tilde F/T)^{-1} (1/T)\sum_m \tilde F_m \bar u_m$,
is the error transmitted from estimating the loadings $\tilde\Lambda$. It
involves a $T$-direction sample average between the factors and the
idiosyncratic shocks, reflecting that $\hat\Lambda$ is identified from
time-series variation. The second term,
$\tilde\Lambda (\tilde\Lambda'\tilde\Lambda/N)^{-1} (1/N)\sum_j
\tilde\Blambda_j \bar u_{jt}$, is the error from estimating the factors
$\tilde F$. It is a loading-weighted cross-section average of the shocks at
$t$, reflecting that $\hat F_t$ is identified from cross-sectional variation.

Notation: $\bar u_m$ is the $N$-vector of demeaned shocks at time $m$
(so $\bar u_m = D_N u_m$), with $j$-th entry $\bar u_{jm}$.

The feasible estimator is $\hat{z}_t = S'[D_N y_t - \hat{C}_t] = z_t - S'[\hat{C}_t - \tilde{C}_t]$. The difference between the estimator and the true value is
\begin{align*}
    \hat{\phi}_d - \phi_d = \frac{\sum_{t=1}^T \hat{z}_t \varepsilon_t}{ \sum_{t=1}^T \hat{z}_t p_{t}} = \frac{\sum_{t=1}^T z_t \varepsilon_t - \sum_{t=1}^T S'[\hat{C}_t - \tilde{C}_t] \varepsilon_t}{\sum_{t=1}^T z_t p_{t} - \sum_{t=1}^T S'[\hat{C}_t - \tilde{C}_t] p_{t}}
\end{align*}
Compared to the infeasible case, we need to analyse the additional terms in the numerator and denominator. By Lemma \ref{lemma:numerator},
\begin{align*}
    \frac{1}{\sqrt{T}} \sum_{t=1}^T S' \big( \hat{C}_t - \tilde{C}_t \big) \varepsilon_t
    &= \frac{1}{T} \sum_{t=1}^T \tilde{F}_t' \varepsilon_t \cdot \left[ \frac{\tilde{F}' \tilde{F}}{T} \right]^{-1} \frac{1}{\sqrt{T}} \sum_{m=1}^T \tilde{F}_m \Bar{u}_{Sm}   \\
    & \enspace +  S' \tilde{\Lambda} \left[ \frac{\tilde{\Lambda}' \tilde{\Lambda}}{N} \right]^{-1} \cdot \frac{1}{\sqrt{T}} \sum_{t=1}^T \left[ \frac{1}{N} \sum_{j=1}^N \tilde{\Blambda}_j \Bar{u}_{jt} \right] \varepsilon_t + O_{\p} \left( \frac{\sqrt{T}}{N} \right)  \\
     &= \frac{1}{T} \sum_{t=1}^T \tilde{F}_t' \varepsilon_t \cdot \left[ \frac{\tilde{F}' \tilde{F}}{T} \right]^{-1} \frac{1}{\sqrt{T}} \sum_{m=1}^T \tilde{F}_m \Bar{u}_{Sm}   \\
    & \enspace + O_{\p}\left( \frac{1}{\sqrt{N}} \right) + O_{\p} \left( \frac{\sqrt{T}}{N} \right)
\end{align*}


By Lemma \ref{lemma:denominator},
\begin{equation*}
    \frac{1}{T} \sum_{t=1}^T S' \big( \hat{C}_t - \tilde{C}_t \big) p_{t}
    = O_{\p} \left( \frac{1}{\sqrt{N}} \right)
\end{equation*}

In every regime, the first-stage estimation contributes an additional term to the asymptotic distribution. This term enters at the same order as the corresponding infeasible quantity. The new term is $O_{\p}(1)$ in the strong regime ($\mu \in (0,1)$) and $O_{\p}(N^{-\delta})$ in the nearly weak regime ($\mu > 1$), where $\delta = \min(1 - 1/\mu, 1/2)$. The rate of convergence of the feasible estimator therefore coincides with that of the infeasible estimator from Section~\ref{section:infeasible giv}. The rates are $\sqrt{T}$ and $\sqrt{T}/N^{\delta}$ respectively. Only the asymptotic variance changes, picking up an additive contribution from the first stage.

For $\mu > 2$, the concentration parameter of the feasible estimator still has only a slowly-divergent floor, just as in the infeasible case. For the same reasons given in Section~\ref{section:infeasible giv}, I therefore recommend Anderson--Rubin confidence sets there. Now we can formally state the results on consistency and asymptotic normality of the feasible GIV estimator.

\subsection{Feasible GIV in Strong Regime: $\mu \in (0,1)$}

In the strong regime, the heavy-tail concentration of the shares delivers $\sqrt{T}$-consistency and asymptotic normality, just as in Theorem~\ref{theorem:strong_consistency_aggregate}. The first-stage estimation contributes an $O_{\p}(1)$ term to the asymptotic distribution, which leaves the rate unchanged but affects the asymptotic variance. I formally state this result in Theorem~\ref{theorem:feasible_strong_consistency_aggregate}

\begin{Theorem} \label{theorem:feasible_strong_consistency_aggregate}
    Suppose Assumptions \ref{ass:weak stationarity and idiosyncrasy} to \ref{ass:LN_moments and CLT} hold with $\mu \in (0,1)$ and $\frac{\sqrt{T}}{N} \to 0$. Then, conditional on $S$, for almost every realization of the shares, the GIV estimator for the aggregate structural parameter is consistent and asymptotically normal,
    \begin{equation*}
       \frac{\Gamma_{zp}}{\sqrt{V_{z \bar{\varepsilon}}(S)}} \cdot \sqrt{T} [\hat{\phi}_d - \phi_d] \xrightarrow{d} \n(0,  1)
    \end{equation*}
     where $\Gamma_{zp}$ is the conditional probability limit of $T^{-1}\sum_t z_t p_t$ given $S$, with explicit form $\Gamma_{zp} = \frac{S' \Sigma_u S}{\phi_d - \phi_s},$ and $V_{z\bar{\varepsilon}}(S) = \lim_{T \to \infty} \frac{1}{T} \sum_{s=1}^T \sum_{t=1}^T \e[z_t z_s \bar{\varepsilon_t} \bar{\varepsilon_s} \mid S]$ where $\bar{\varepsilon}_t = \varepsilon_t - \e[\tilde{F}_t' \varepsilon_t] \Sigma_{\tilde{F}}^{-1} \tilde{F}_t $. These limits exist and $\Gamma_{zp} \neq 0$ for almost every realization of $S$, provided $\phi_d \neq \phi_s$.
\end{Theorem}
\begin{proof}
    Proof in Appendix \ref{subsec:proof of feasible_theorems}. The proof follows the proof of Theorem \ref{theorem:strong_consistency_aggregate} with the additional first-stage term characterized by Lemmas~\ref{lemma:numerator} and~\ref{lemma:denominator}.
\end{proof}

\subsubsection{Comparison with \citet{Banafti2022InferentialDimensions}} \label{subsec:banfti}

\citet{Banafti2022InferentialDimensions} also study large-panel GIV in the
strong-instrument case. Their Theorems 2 and 4 conclude that the first-stage estimation of the instrument has no impact on the asymptotic variance of the GIV estimator. My
Theorem~\ref{theorem:feasible_strong_consistency_aggregate} reaches the
opposite conclusion. I show that the first-stage estimation has a first-order contribution to the asymptotic variance of the GIV estimator.

The difference arises due to two reasons. The first is an assumption they impose. I show that this assumption is not compatible with the power law setting and hence needs to be relaxed. The second is the order of one term in their Lemma 2. They show that this term is insignificant in the limit. But I show that this result relies on very strict assumptions which even \citet{Banafti2022InferentialDimensions} do not formally impose.

The first point is their Assumption 4(iii). This assumption is in addition to the power law assumption, identical to our Assumption \ref{ass:size_of_firm_power_law}. However, I show that this Assumption 4(iii) is not compatible with the power law setting. That is, with the individual sizes $\s_i$ following the power law, Assumption 4(iii) is not possible.



The second point is regarding the order of a term in the asymptotic form of the estimator. This term vanishes only under very strict assumptions. In the general case, this term adds to the asymptotic variance of the estimator. Hence this term directly drives the difference. I develop both points in Appendix~\ref{appendix:banafti}.




\subsection{Feasible GIV in Nearly Weak Identification: $\mu > 1$}

In the nearly weak regime, the feasible estimator inherits the rate $\sqrt{T}/N^{\delta}$ of Theorem~\ref{theorem:weak_consistency_aggregate}, where $\delta = \min(1 - 1/\mu, 1/2)$. The first-stage estimation contributes an $O_{\p}(N^{-\delta})$ term to the asymptotic distribution. This matches the order of the infeasible quantities and leaves the rate unchanged. I formally state this result in Theorem~\ref{theorem:feasible_weak_consistency_aggregate}.

\begin{Theorem} \label{theorem:feasible_weak_consistency_aggregate} \label{theorem:feasible_semi_strong_consistency_aggregate}
    Suppose Assumptions \ref{ass:weak stationarity and idiosyncrasy} to \ref{ass:LN_moments and CLT} hold with $\mu > 1$, $N/T \to 0$, and $\frac{\sqrt{T}}{N} \to 0$. Let $\delta = \min(1 - 1/\mu, 1/2)$, so that $\delta = 1 - 1/\mu \in (0, 1/2)$ when $\mu \in (1,2)$ and $\delta = 1/2$ when $\mu > 2$. Then, conditional on $S$, for almost every realization of the shares, the GIV estimator for the aggregate structural parameter is consistent and asymptotically normal,
    \begin{equation*}
       \frac{\Gamma_{zp}}{\sqrt{V_{z \bar{\varepsilon}}(S)}} \cdot \sqrt{T}\, [\hat{\phi}_d - \phi_d] \xrightarrow{d} \n(0, 1),
    \end{equation*}
    where $\Gamma_{zp}$ is the conditional probability limit of $\frac{1}{T}\sum_t z_t p_t$ given $S$, with explicit form $\Gamma_{zp} = \frac{S'\Sigma_u S}{\phi_d - \phi_s}$, and $V_{z \bar{\varepsilon}}(S) = \lim_{T \to \infty} \frac{1}{T}\sum_{s=1}^T \sum_{t=1}^T \e[z_t z_s \bar{\varepsilon}_t \bar{\varepsilon}_s \mid S]$ with $\bar{\varepsilon}_t = \varepsilon_t - \e[\tilde{F}_t' \varepsilon_t] \Sigma_{\tilde{F}}^{-1} \tilde{F}_t$. These limits exist and $\Gamma_{zp} \neq 0$ for almost every realization of $S$, provided $\phi_d \neq \phi_s$. The studentization satisfies $\Gamma_{zp}/\sqrt{V_{z\bar{\varepsilon}}(S)} = O_{\p}(N^{-\delta})$, so the standardized statistic vanishes at the rate $\sqrt{T}/N^{\delta}$.

    If $\mu > 2$ and $N/T \to c > 0$, then $\hat{\phi}_d$ is inconsistent and identification is weak in the sense of \citet{Staiger1997InstrumentalInstruments}.
\end{Theorem}
\begin{proof}
    Proof in Appendix \ref{subsec:proof of feasible_theorems}. The argument follows Theorem~\ref{theorem:weak_consistency_aggregate}, with the additional first-stage term characterized by Lemmas~\ref{lemma:numerator} and~\ref{lemma:denominator}.
\end{proof}

\begin{remark}[Three regimes, feasible] \label{rem:three_regimes_feasible}
The trichotomy of Remark~\ref{rem:three_regimes_aggregate} carries over to the feasible estimator. Theorem~\ref{theorem:feasible_strong_consistency_aggregate} covers the strong regime, Theorem~\ref{theorem:feasible_weak_consistency_aggregate} covers the nearly weak regime under $N/T \to 0$, and the same theorem's failure-mode clause states the weak identification regime under $N/T \to c$ with $\mu > 2$. The additional growth restriction $\sqrt{T}/N \to 0$ from estimating the factor structure applies uniformly across the three regimes.
\end{remark}

We need a consistent estimator of the asymptotic variance for inference. I propose a Heteroskedasticity and Auto-Correlation (HAC) consistent estimator for the asymptotic variance. Define the regime-specific rescaling
\begin{equation*}
    a_N = \begin{cases} 1 & \mu \in (0,1), \\ N^{\delta},\ \delta = \min(1 - \tfrac{1}{\mu}, \tfrac{1}{2}) & \mu > 1. \end{cases}
\end{equation*}
Now define the estimator of asymptotic variance
\begin{equation} \label{eq:HAC estimator_disaggregated_variance}
    \hat{V}^H_{{z} \bar{\varepsilon}} = \frac{a_N^2}{T} \sum_{t=1}^T \hat{z}_t^2 \tilde{\varepsilon}_t^2 + \frac{2 a_N^2}{T} \sum_{s=1}^{b_T} w(\frac{s}{b_T}) \sum_{t = s+1}^T \hat{z}_t \hat{z}_s \tilde{\varepsilon}_t \tilde{\varepsilon}_s
\end{equation}

where $b_T$ is the bandwidth and $w(x)$ is a kernel function, $w: \mathbb{R}^+ \to [0,1]$ such that $w(x) = 0$ for $x>1$ and $w(0) = 1$. $\tilde{\varepsilon}_t = \hat{\varepsilon}_t - \frac{1}{T} \sum_t \hat{F}'_t  \hat{\varepsilon}_t \cdot \hat{\Sigma}_{\tilde{F}}^{-1} \cdot  \hat{F}_t $, where $\hat{\varepsilon}_t = d_t - \hat{\phi}_d p_{t}$ and $\hat{\Sigma}_{\tilde{F}} = \left[ \frac{\hat{F}' \hat{F}}{T} \right]$.

\begin{Prop} \label{prop:consistency of HAC_aggregate}
    Suppose Assumptions \ref{ass:weak stationarity and idiosyncrasy} to \ref{ass:LN_moments and CLT} hold,  $\frac{N}{T} \xrightarrow{} 0$, and $\frac{\sqrt{T}}{N} \xrightarrow{} 0$. Then conditional on $\s$,
    \begin{equation*}
        \hat{V}^H_{{z} \bar{\varepsilon}} - V_{z \bar{\varepsilon}}(S) \xrightarrow{p} 0.
    \end{equation*}
\end{Prop}
\begin{proof}
    See Appendix \ref{subsec:proof or HAC_aggregate}
\end{proof}


\subsubsection{Inference}

The inference choice mirrors Section~\ref{subsec:infeasible_AR}. The concentration parameter of the feasible estimator has a polynomial floor in $N$ when $\mu \in (1,2)$ and only a slowly-divergent floor when $\mu > 2$. Wald is reliable for the former. I recommend Anderson--Rubin for the latter.

\paragraph{Wald inference for $\mu \in (1,2)$.} The polynomial floor keeps Wald inference reliable. The asymptotic variance is given by the HAC estimator of Proposition~\ref{prop:consistency of HAC_aggregate}, which absorbs the rate $\delta$ automatically through the rescaling $a_N$. Standard Wald confidence intervals are valid.

\paragraph{Anderson--Rubin inference for $\mu > 2$.} The concentration parameter of the feasible estimator can grow arbitrarily slowly, so its realized value in finite samples need not be large. The estimator therefore has poor finite-sample performance even though it is asymptotically normal. I recommend inverting the Anderson--Rubin (AR) test of \citet{Anderson1949EstimatorsEquations}.

The Anderson--Rubin test is evaluated at the null, so it never uses the GIV estimate $\hat{\phi}_d$. Two features follow, and together they make the $N/T$ condition irrelevant. First, under $H_0: \phi_d = \phi_d^0$, the residual $d_t - \phi_d^0 p_t = \varepsilon_t$ is exact, so no first-stage estimation of $\phi_d$ enters the variance estimator. Second, the statistic is self-normalising, and the rate $a_N$ cancels between the numerator and the variance estimator, so the test refers neither to the rate $\delta$ nor to $N/T$. The only condition that survives is $\sqrt{T}/N \to 0$, from estimating the factor structure in $\hat{z}_t$.

Consider the null hypothesis, $H_0: \phi_d = \phi_d^0$. Define the null-imposed residual $\varepsilon_t(\phi_d^0) = d_t - \phi_d^0 p_t$ and its factor-projected version $\tilde{\varepsilon}_t(\phi_d^0) = \varepsilon_t(\phi_d^0) - \frac{1}{T} \sum_t \hat{F}_t' \varepsilon_t(\phi_d^0) \cdot \hat{\Sigma}_{\tilde{F}}^{-1} \hat{F}_t$. Define the Anderson-Rubin statistic \citep{Anderson1949EstimatorsEquations},
\begin{equation} \label{eq:AR statistic_aggregate}
    \text{AR}(\phi_d^0) =  \frac{N}{T} \left( \sum_t \hat{z}_t (y_{St} - \phi_d^0 p_{t}) \right)^2 \frac{1}{\hat{V}^H_{z \bar{\varepsilon}}(\phi_d^0)}
\end{equation}

where $\hat{V}^H_{z \bar{\varepsilon}}(\phi_d^0)$ is the HAC estimator defined in \eqref{eq:HAC estimator_disaggregated_variance}, specialised to $\mu > 2$ so that $a_N^2 = N$, and built from the null-imposed residual,
\begin{equation*}
    \hat{V}^H_{{z} \bar{\varepsilon}}(\phi_d^0) = \frac{N}{T} \sum_{t=1}^T \hat{z}_t^2 \tilde{\varepsilon}_t(\phi_d^0)^2 + \frac{2 N}{T} \sum_{s=1}^{b_T} w(\frac{s}{b_T}) \sum_{t = s+1}^T \hat{z}_t \hat{z}_s \tilde{\varepsilon}_t(\phi_d^0) \tilde{\varepsilon}_s(\phi_d^0)
\end{equation*}

The $N/T$ in the numerator cancels the $a_N^2/T = N/T$ in $\hat{V}^H_{z \bar{\varepsilon}}(\phi_d^0)$, so the statistic is self-normalised and carries no reference to the rate. It is consistent for the conditional asymptotic variance $V_{z \bar{\varepsilon}}(S)$, with $\bar{\varepsilon}_t = \varepsilon_t - \e[\tilde{F}_t' \varepsilon_t] \Sigma_{\tilde{F}}^{-1} \tilde{F}_t$, by the argument of Proposition~\ref{prop:consistency of HAC_aggregate} with the residual reduction now exact. Inversion of this test yields a confidence region of the correct size.

\begin{Theorem} \label{th:weakness robust test_aggregate}
    Suppose Assumptions \ref{ass:weak stationarity and idiosyncrasy} to \ref{ass:LN_moments and CLT} hold with $\mu > 2$ and $\frac{\sqrt{T}}{N} \to 0$. Then, for any trajectory of $N/T$, under the null $H_0: \phi_d = \phi_d^0$,
\begin{equation*}
    \text{AR}(\phi_d^0) \xrightarrow{d} \chi^2_1.
\end{equation*}
\end{Theorem}
\begin{proof}
    In Appendix~\ref{subsec:proof of AR_aggregate}. The proof works at the null, so the GIV estimate never enters and no restriction on $N/T$ is required.
\end{proof}

Theorem~\ref{th:weakness robust test_aggregate} places no condition on $N/T$. It therefore covers the weak-identification case $N/T \to c > 0$, where $\hat{\phi}_d$ is inconsistent. The test does not depend on consistency of $\hat{\phi}_d$. It uses only the moment and the variance, both evaluated at the hypothesised value.

\section{Empirical Application}
\label{sec:empirical}

We apply the GIV estimator to three commodity markets---refined copper, crude
oil, and natural gas---to estimate their respective price elasticities of
demand.  Each market provides a global supply panel whose country-level
idiosyncratic shocks serve as the GIV instrument.

\subsection{Model}
\label{sec:model}

\paragraph{Demand equation.}
The equation of interest is
\begin{equation}
  y_{St} \;=\; \phi_d\, p_t \;+\; X_t'\beta \;+\; \varepsilon_t,
\end{equation}
where $y_{St}$ is the year-on-year growth rate of aggregate demand,
$p_t$ is the year-on-year growth rate of the real commodity price,
$X_t$ is a vector of observable common controls, and $\varepsilon_t$ is an aggregate
demand shock.  The parameter of interest is $\phi_d$, the price elasticity of
demand.

OLS estimation of~\eqref{eq:demand} is biased because $p_t$ and
$\varepsilon_t$ are correlated. Aggregate demand expansions raise both quantities and
prices simultaneously, attenuating the estimated elasticity toward zero, or
even reversing its sign when supply shocks are small relative to demand
shocks.

\paragraph{Supply panel and GIV instrument.}
We observe the changes to supply at individual country level. Let
 $y_{it}$ denote country~$i$'s change in supply in period~$t$.  The change in supply follows the structural equation:
\begin{equation}
  \label{eq:panel}
  y_{it} \;=\; \phi_s p_t + \lambda_i' F_t \;+\; u_{it},
\end{equation}
where $F_t$ is an $r$-vector of common factors (spanning price and aggregate
demand shocks) and $u_{it}$ is an idiosyncratic supply shock. The feasible
GIV instrument is the share-weighted idiosyncratic shock,
\begin{equation}
  \label{eq:giv}
  \hat{z}_t \;=\; \sum_{i=1}^N S_{i}\,\hat{u}_{it} = S'(D_N y_t - C_t)
\end{equation}
where $S_i$ is the long-term market share of country~$i$. In our dataset, we construct $S_i$ by calculating the share for every $t$ and averaging across all time periods. I construct all growth rates as year-on-year midpoint growth.
\begin{equation} \label{eq:yoygrowth}
  g_{i,t} = \frac{Y_{i,t} - Y_{i,t-12}}{\tfrac{1}{2}(Y_{i,t} + Y_{i,t-12})},
\end{equation}

\paragraph{Factor selection.}
The number of factors $r$ is selected from the data using the eigenvalue
ratio~(ER) criterion of \citet{Ahn2013EigenvalueFactors}. The \citet{Bai2002DeterminingModels} information criteria are not used because their $O(\ln N / N)$ penalty is calibrated for large~$N$ and proves too weak to discriminate at the cross-section sizes ($N = 21$--$29$) encountered here.


\subsection{Data}
See Appendix \ref{appendix_data} for more comments on the data construction.
\subsubsection{Copper}

We use monthly country-level data from Bloomberg covering January 2009 to
December~2025 ($T = 204$ months).  The supply panel for the GIV instrument
consists of refined copper supply across $N = 29$ countries.  Each panel includes a rest-of-world residual. The price series is the LME spot copper price (monthly average), deflated by U.S.\ CPI rebased to 2015~$= 100$.

Post transformation, we have $T = 192$ estimation periods. The covariate matrix~$X_t$ includes an intercept, one lag of aggregate refined demand growth, and the trade-weighted
U.S.\ dollar index.

\subsubsection{Crude Oil}

We use monthly EIA International Energy Statistics data on crude oil
production from January~1973 to November~2025 ($T = 635$ months).  The
USSR and its successor states are treated as a single continuous unit (Former
USSR series pre-1992; sum of 15 successor states post-1991).  After dropping
countries with any zero or missing observation, the panel has $N = 21$ units
(20 countries plus rest of world).  The year 2020 is excluded from estimation
to avoid contamination from the COVID-19 production collapse.

The primary price series is a splice of the FRED OILPRICE series (January
1946--August~2024) and WTI (September 2024--December 2025), deflated by U.S.\
CPI rebased to 2015~$= 100$.  Post transformation, we have $T = 611$ estimation periods. Aggregate demand is constructed by aggregate production growth adjusted for inventory changes. The covariate matrix~$X_t$ includes an intercept and two lags of aggregate production growth.

\subsubsection{Natural Gas}

Monthly country-level natural gas production data are obtained from the JODI
Gas Database, covering January~2010 to November~2025 ($T = 191$ months). The panel contains $N = 27$ countries.  The year 2020 is excluded from
estimation to avoid contamination from the COVID-19 demand collapse.

The price series is the Henry Hub Natural Gas Spot Price (dollars per million
British thermal units), deflated by U.S.\ CPI rebased to
2015~$= 100$.  Post transformation, we have $T = 156$ periods.  The covariate matrix~$X_t$ includes an intercept, eleven lags of aggregate production growth, and the growth rate of the real WTI crude oil price.  The eleven lags are motivated by strong seasonality in natural gas markets, where winter heating demand drives pronounced annual cycles in both quantities and prices.  The real oil price is included because natural gas and oil are partial substitutes in
power generation and industrial use, making oil price variation a relevant
demand shifter.

\subsection{Results}
\label{sec:results}

\subsubsection{Granularity}

We assume that the cross-sectional size distribution follows a power law, $\Pr(S \geq s) \propto s^{-\mu}$. Table~\ref{tab:power_law_combined} reports Pareto exponent $\mu$ estimates for the supply panels using three methods: Hill~(1975) MLE, log-rank OLS, and the bias-corrected
Gabaix--Ibragimov regression \citep{Gabaix2011RankExponents}.


\begin{table}[H]
\centering
\caption{Pareto Tail Exponent Estimates}
\label{tab:power_law_combined}
\small
\begin{tabular}{lcccccc}
\toprule
Method & \multicolumn{2}{c}{Copper} & \multicolumn{2}{c}{Crude Oil} & \multicolumn{2}{c}{Natural Gas} \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5}\cmidrule(lr){6-7}
  & $\hat{\mu}$ & 95\% CI & $\hat{\mu}$ & 95\% CI & $\hat{\mu}$ & 95\% CI \\
\midrule
Hill (1975) MLE & 0.22 & [0.21, 0.76] & 0.38 & [0.34, 0.77] & 0.20 & [0.17, 0.31] \\
Naive OLS & 0.51 & [0.39, 0.64] & 0.58 & [0.44, 0.71] & 0.29 & [0.24, 0.34] \\
Gabaix--Ibragimov (2011) & 0.58 & [0.28, 0.87] & 0.65 & [0.26, 1.05] & 0.32 & [0.15, 0.49] \\
\bottomrule
\end{tabular}

\vspace{4pt}
\raggedright\footnotesize
\textit{Notes:} Copper: Refined Supply panel, $N = 29$. Crude Oil: Production (EIA) panel, $N = 21$. Natural Gas: Production panel, $N = 27$. Hill (1975) MLE: $\hat{\mu} = N / \sum_i \log(S_i / S_{\min})$ with nonparametric bootstrap 95\% CI (5{,}000 replications, resampling with replacement). Naive OLS: log-rank regression $\log(\text{rank}) = a - \mu \log(S)$ with OLS standard errors. Gabaix--Ibragimov (2011): shifted log-rank regression $\log(\text{rank} - 1/2) = a - \mu \log(S)$ with analytic SE $= \sqrt{2/N}\,\hat{\mu}$.
\end{table}



\paragraph{Regime classification.} The estimates partition the three
commodities cleanly. Copper has $\hat\mu_{\mathrm{GI}} = 0.58$ with
confidence interval $[0.28, 0.87]$, and natural gas has
$\hat\mu_{\mathrm{GI}} = 0.32$ with $[0.15, 0.49]$. Both lie firmly in
the strong regime $\mu < 1$ of Section~\ref{sec:three_regimes}. Crude
oil has $\hat\mu_{\mathrm{GI}} = 0.65$ but a wider interval,
$[0.26, 1.05]$, that reaches into the nearly weak region
$\mu \in (1,2)$.
Theorems~\ref{theorem:strong_consistency_aggregate} and~\ref{theorem:weak_consistency_aggregate} establish
consistency and asymptotic normality for the entire range
$\mu \in (0,2)$, so the point estimates and Wald confidence intervals
reported below remain valid for crude oil under either reading of $\mu$.
We can comfortably rule out $\mu > 2$ for all three commodities.

All point estimates are below~1 for all three commodities, consistent with
the original GIV validity condition~($\mu < 1$).  However, the
Gabaix--Ibragimov confidence interval for crude oil contains~1. This is precisely the situation my extended theory is designed to cover. Even if the concentration is not extreme enough to satisfy $\mu < 1$ with certainty, we can safely use the point estimates and conduct inference for $\mu \in (0,\,2)$.

Figure~\ref{fig:zipf} plots log-rank against log-share for all three panels.
The near-linear relationship over the full support confirms that a Pareto
distribution is a reasonable description of the size distribution in each
market, with the estimated Gabaix--Ibragimov slope $\hat\mu$ shown as the fitted
line.

\begin{figure}[H]
  \centering
  \includegraphics[width=\linewidth]{figures/zipf_plot.pdf}
  \caption{Log-rank versus log-share for the supply panels used in GIV
    estimation.  Each point is one country's time-averaged market share.
    The red line shows the Gabaix--Ibragimov fitted slope~$\hat\mu_{\mathrm{GI}}$.}
  \label{fig:zipf}
\end{figure}

\subsubsection{Demand Elasticities}

Table~\ref{tab:elasticity_combined} presents the main elasticity estimates.


\begin{table}[H]
\centering
\caption{Price Elasticity of Demand: OLS and GIV Estimates}
\label{tab:elasticity_combined}
\footnotesize
\begin{tabular}{lccc}
\toprule
Estimator & Copper & Crude Oil & Natural Gas \\
  & Refined Supply ($N=29$) & Production (EIA) ($N=21$) & Production ($N=27$) \\
\midrule
OLS & -0.0512 (0.0226) & 0.0041 (0.0033) & 0.0010 (0.0046) \\
GIV --- infeasible SE & -0.1351 (0.0502) & -0.1092 (0.0538) & -0.0556 (0.0338) \\
GIV --- feasible SE & -0.1351 (0.0513) & -0.1092 (0.0548) & -0.0556 (0.0255) \\
\bottomrule
\end{tabular}

\vspace{4pt}
\raggedright\footnotesize
\textit{Notes:} Standard errors in parentheses. OLS is endogenous (common demand shocks bias the estimate; attenuation toward zero or sign reversal is expected). Feasible SE corrects for estimation error in the estimation of the GIV instrument; infeasible SE treats the instrument as fixed. All standard errors: HAC (Newey--West, Bartlett kernel). Copper: $r = 1$ factors, Crude Oil: $r = 2$ factors, Natural Gas: $r = 2$ factors.
\end{table}


The OLS estimates starkly illustrate the endogeneity problem.  For copper,
OLS yields $-0.051$, roughly one-third of the GIV estimate in magnitude.  For
crude oil and natural gas, OLS is positive. The raw price-quantity
covariance is driven by demand shocks that raise both price and the aggregate,
producing a positive spurious correlation.  GIV corrects all three estimates
to economically plausible negative demand elasticities: $-0.135$ for copper,
$-0.109$ for crude oil, and $-0.056$ for natural gas.

For copper and crude oil, the difference between the infeasible and feasible
GIV standard errors is small (at most two basis points), confirming that the
PCA estimation step adds little additional uncertainty in practice.  For
natural gas, the feasible standard error is surprisingly \emph{smaller} than
the infeasible one.  This occurs because the first-stage correction term that
projects $\varepsilon_t$ onto the estimated factors, namely
$\frac{1}{T}\sum_t \hat{F}_t'\hat{\varepsilon}_t \cdot \hat\Sigma_{\tilde F}^{-1} \hat{F}_t$,
can be negatively correlated with the uncorrected IV
residuals, reducing the total variance of the corrected moment conditions.
This further illustrates the importance of accounting for the estimation error in the construction of the instrument.

All three commodities display inelastic demand.  The estimated elasticities
($|\hat\phi| < 0.15$) are consistent with the industrial nature of these commodities. All have limited short-run substitutes, making quantity responses to price changes modest.  The estimates are also broadly
in line with the existing literature, which typically reports short-run
elasticities in the range $[-0.05,\,-0.25]$ for crude oil \citep{Baumeister2019StructuralShocks}, $[-0.05,\,-0.20]$ for natural gas \citep{Auffhammer2018NaturalBills,Labandeira2017ADemand}, and $[-0.07,\,-0.10]$ for copper \citep{Shojaeddini2025UnderstandingMinerals,Lanz2013SubglobalIndustry}.



\section{Simulation} \label{sec:simulation}

I assess the finite sample performance of the GIV estimator under different regimes of instrument strength via Monte Carlo simulation. The factor structure is calibrated to two empirical commodity panels: Copper and Crude Oil, while all other components of the data generating process are simulated. Running the experiment on both commodities allows us to examine how the estimator performs under different cross-section and time-series dimensions: Copper provides $N = 29$ countries over $T = 192$ months, while Crude Oil provides $N = 21$ countries over $T = 611$ months.

\subsection{Data Generating Process}

The structural equation of interest is the aggregate demand for a commodity,
\begin{equation}\label{eq:structural}
    d_t = y_{St} = \phi_d p_t + \varepsilon_t,
\end{equation}
The supply panel follows the factor model
\begin{equation}\label{eq:factor_model}
    y_t = \iota_N F_t^1 + \tilde{\Lambda} \tilde{F}_t + u_t,
\end{equation}
where $F_t^1$ is a common factor, $\tilde{\Lambda}$ is the $N \times r - 1$ matrix of demeaned factor loadings, $\tilde{F}_t$ is the $r-1 \times 1$ vector of factors, and $u_t$ is the $N \times 1$ vector of idiosyncratic shocks.

The demeaned factor loadings $\tilde{\Lambda}$ and the factors $\tilde{F}_t$ are extracted from each empirical estimation and held fixed across all Monte Carlo replications.

All remaining quantities are drawn independently each replication. The common factor and idiosyncratic shocks are drawn from normal distributions and scaled to be comparable to the empirical factor:
\begin{equation}\label{eq:draws}
    F_t^1 \sim \mathcal{N}(0, \sigma_F^2), \qquad
    u_{it} \sim \mathcal{N}(0, \sigma_u^2),
\end{equation}
where $\sigma_F = \text{std}(\hat{F})$ matches the standard deviation of the empirical factor series, and $\sigma_u = \text{std}(\hat{u}_t)$ matches the pooled standard deviation of the empirical PCA estimation. The structural error $\varepsilon_t$ is drawn with $\text{std}(\varepsilon_t) = \sigma_F$ and $\text{corr}(\varepsilon_t, \tilde{F}_t) = 0.8$.

\subsubsection{Calibration}

For the Copper DGP, the refined copper supply panel consists of year-on-year midpoint growth rates for $N = 29$ countries over $T = 192$ months. The ER criterion yields $r-1 = 1$ factor. The calibration scales are $\sigma_F = 0.109$ and $\sigma_u = 0.171$, and the true elasticity is set to $\phi_d = -0.135$ (the empirical GIV estimate).

For the Crude Oil DGP, the crude oil production panel consists of year-on-year midpoint growth rates for $N = 21$ countries over $T = 611$ months (COVID year 2020 dropped). The ER criterion yields $r-1 = 2$ factors. The calibration scales are $\sigma_F = 0.090$ and $\sigma_u = 0.142$, and the true elasticity is set to $\phi_d = -0.110$ (the empirical GIV estimate of $-0.109$, rounded to two decimals).

\subsubsection{Individual sizes and Granularity}

Individual sizes are drawn from the power distribution $\mathbb{P}(s_i > s) = c \, s^{-\mu}$ via the inverse CDF transform $s_i = (1 - U_i)^{-1/\mu}$ with $U_i \sim \mathcal{U}[0,1]$, and normalized so that $\sum_{i=1}^N S_i = 1$. Recall that the Pareto exponent $\mu$ determines the market concentration, and consequently the strength of the instrument. We consider $\mu \in \{0.3, 0.5, 0.8, 1.2, 1.4, 1.8, 2.5, 3.5, 6.0\}$, with $B = 5{,}000$ replications per value.

From this DGP, we construct the observational data using $y_t = \iota_N F_t^1 + \tilde{\Lambda} \tilde{F}_t + u_t$, $y_{St} = S' y_t$, and $p_t = \frac{y_{St} - \varepsilon_t}{\phi_d}$. The price is endogenous. Specifically, $\text{Cov}(p_t, \varepsilon_t) = -\sigma_\varepsilon^2 / \phi_d > 0$ (since $\phi_d < 0$), inducing an upward OLS bias.

\subsubsection{Augmenting the cross section}

The Crude Oil panel has only $N = 21$ countries, which limits the scope for studying the estimator's behavior as $N$ grows. To investigate the effect of a larger cross-section while preserving the empirical factor structure, we augment $N$ by resampling rows of $\tilde{\Lambda}$ with replacement. Specifically, for a target $N_{\text{sim}} > N$, we draw $N_{\text{sim}}$ rows from the $21 \times r$ empirical loading matrix uniformly with replacement. The factors $\tilde{F}_t$, calibration scales $\sigma_F$ and $\sigma_u$, and the true elasticity $\phi_d$ remain unchanged. Each replication still draws fresh $u_{it}$, $F_t^1$, $\varepsilon_t$, and $S$, so the augmented countries are distinguished by their idiosyncratic shocks and shares even when they share a loading vector. I report results for $N \in \{21, 50, 100\}$.

\subsection{Results}

Tables~\ref{tab:mc_table1} and~\ref{tab:mc_table2} organise the simulation
evidence by the Pareto exponent $\mu$, which Proposition~\ref{prop:weakness of infeasible instrument}
maps directly onto the regimes of instrument strength. The CI Length
column reports the standard $t$-based 95\% interval with White
heteroskedasticity-robust standard errors. The Anderson--Rubin confidence
sets recommended in Section~\ref{subsec:infeasible_AR} and its feasible
counterpart are not used here. We let the $t$-based interval reveal where
it breaks down.


\begin{table}[htbp]
\centering
\caption{Monte Carlo Simulation Results: Empirical Commodity Panels}
\label{tab:mc_table1}
\begin{tabular}{ccccccc}
\toprule
$\mu$ & Median $\hat{\phi}_d$ & Bias & RMSE & Coverage & CI Length & Median $F$ \\
\midrule
\multicolumn{7}{l}{\textit{Panel A: Copper ($N=29$, $T=192$, $\phi_d=-0.135$)}} \\
\midrule
0.3 & -0.1351 & -0.0004 & 0.0069 & 0.948 & 0.0203 & 124.0 \\
0.5 & -0.1351 & -0.0008 & 0.0097 & 0.947 & 0.0258 & 78.2 \\
0.8 & -0.1351 & -0.0003 & 0.0740 & 0.949 & 0.0399 & 34.0 \\
1.2 & -0.1349 & -0.0039 & 0.1491 & 0.947 & 0.0674 & 12.5 \\
1.4 & -0.1347 & -0.0066 & 0.2756 & 0.954 & 0.0803 & 8.7 \\
1.8 & -0.1334 & -0.0187 & 0.7069 & 0.949 & 0.1130 & 4.6 \\
2.5 & -0.1302 & -0.1731 & 8.1220 & 0.958 & 0.1769 & 2.1 \\
3.5 & -0.1222 & -0.0022 & 1.7957 & 0.960 & 0.2688 & 1.0 \\
6.0 & -0.1065 & 0.0043 & 1.4744 & 0.967 & 0.4138 & 0.5 \\
\midrule
\multicolumn{7}{l}{\textit{Panel B: Crude Oil ($N=21$, $T=611$, $\phi_d=-0.110$)}} \\
\midrule
0.3 & -0.1100 & -0.0001 & 0.0035 & 0.946 & 0.0091 & 433.8 \\
0.5 & -0.1100 & -0.0001 & 0.0047 & 0.946 & 0.0120 & 246.6 \\
0.8 & -0.1100 & -0.0003 & 0.0065 & 0.947 & 0.0185 & 106.5 \\
1.2 & -0.1100 & -0.0031 & 0.1412 & 0.953 & 0.0300 & 39.2 \\
1.4 & -0.1099 & -0.0019 & 0.0355 & 0.958 & 0.0358 & 27.5 \\
1.8 & -0.1099 & -0.0026 & 0.0422 & 0.953 & 0.0486 & 14.8 \\
2.5 & -0.1104 & -0.0022 & 0.1976 & 0.957 & 0.0742 & 6.7 \\
3.5 & -0.1086 & -1.4437 & 101.1734 & 0.959 & 0.1117 & 3.1 \\
6.0 & -0.0999 & 0.0190 & 1.9230 & 0.969 & 0.2058 & 1.0 \\
\bottomrule
\end{tabular}

\vspace{4pt}
\raggedright\footnotesize
\textit{Notes:} Shares drawn from power law $\mathbb{P}(s_i > s) = c\,s^{-\mu}$. Structural error $\varepsilon_t$ correlated with the dominant common factor at $\rho=0.8$ by construction. Coverage is the fraction of replications where the true value lies in the 95\% confidence interval. Standard errors: White heteroskedasticity-robust. $F$ is the first-stage $F$-statistic.
\end{table}



\begin{table}[htbp]
\centering
\caption{Monte Carlo Simulation Results: Augmented Cross-Section (Crude Oil)}
\label{tab:mc_table2}
\begin{tabular}{ccccccc}
\toprule
$\mu$ & Median $\hat{\phi}_d$ & Bias & RMSE & Coverage & CI Length & Median $F$ \\
\midrule
\multicolumn{7}{l}{\textit{Panel A: Crude Oil ($N=50$, $T=611$, $\phi_d=-0.110$)}} \\
\midrule
0.3 & -0.1099 & 0.0000 & 0.0029 & 0.940 & 0.0086 & 476.7 \\
0.5 & -0.1100 & -0.0001 & 0.0037 & 0.949 & 0.0112 & 285.3 \\
0.8 & -0.1100 & -0.0003 & 0.0060 & 0.945 & 0.0178 & 114.1 \\
1.2 & -0.1099 & -0.0007 & 0.0124 & 0.950 & 0.0323 & 37.3 \\
1.4 & -0.1098 & -0.0020 & 0.0450 & 0.954 & 0.0400 & 24.7 \\
1.8 & -0.1096 & 0.0397 & 3.0934 & 0.945 & 0.0571 & 12.1 \\
2.5 & -0.1101 & -0.0256 & 0.7741 & 0.953 & 0.0923 & 5.0 \\
3.5 & -0.1061 & -0.0771 & 5.1400 & 0.961 & 0.1453 & 2.1 \\
6.0 & -0.0958 & 0.0220 & 1.1442 & 0.974 & 0.2781 & 0.7 \\
\midrule
\multicolumn{7}{l}{\textit{Panel B: Crude Oil ($N=100$, $T=611$, $\phi_d=-0.110$)}} \\
\midrule
0.3 & -0.1099 & 0.0000 & 0.0025 & 0.947 & 0.0083 & 512.0 \\
0.5 & -0.1100 & -0.0001 & 0.0033 & 0.952 & 0.0104 & 326.8 \\
0.8 & -0.1100 & -0.0003 & 0.0059 & 0.948 & 0.0174 & 121.3 \\
1.2 & -0.1096 & -0.0016 & 0.0340 & 0.958 & 0.0338 & 33.3 \\
1.4 & -0.1102 & -0.0018 & 0.0600 & 0.950 & 0.0445 & 19.5 \\
1.8 & -0.1100 & -0.0022 & 0.2524 & 0.955 & 0.0675 & 8.6 \\
2.5 & -0.1078 & -0.0155 & 0.7632 & 0.956 & 0.1153 & 3.1 \\
3.5 & -0.1020 & 0.0166 & 1.4702 & 0.970 & 0.1846 & 1.3 \\
6.0 & -0.0901 & 0.0306 & 1.7110 & 0.980 & 0.3196 & 0.6 \\
\bottomrule
\end{tabular}

\vspace{4pt}
\raggedright\footnotesize
\textit{Notes:} Shares drawn from power law $\mathbb{P}(s_i > s) = c\,s^{-\mu}$. Structural error $\varepsilon_t$ correlated with the dominant common factor at $\rho=0.8$ by construction. Coverage is the fraction of replications where the true value lies in the 95\% confidence interval. Standard errors: White heteroskedasticity-robust. $F$ is the first-stage $F$-statistic.
\end{table}


\paragraph{Strong regime ($\mu < 1$).} The strong-instrument theory predicts
standard inference, and that is what we observe. For Copper, the median
$F$-statistic is large throughout, falling from $124$ at $\mu = 0.3$ to $78$
and $34$ at $\mu = 0.5$ and $0.8$ as shares become less concentrated. The
median $\hat\phi_d$ matches the true value to three decimals and bias is at
most $0.0008$. RMSE is below $0.01$ at $\mu \in \{0.3, 0.5\}$; at $\mu = 0.8$
it rises to $0.074$ as a few extreme draws enter the second moment, even
though the median and bias remain on target. Crude Oil is sharper, with
median $F$ between $107$ and $434$ and RMSE below $0.01$ throughout. Coverage
of the $t$-based interval is at the nominal $0.95$ in both panels.

\paragraph{Nearly weak regime ($\mu \in (1,2)$).} The estimator continues
to recover the true elasticity. The median estimate stays at the true
value to three decimals and coverage of the $t$-based interval is at
nominal levels in both calibrations. The point-estimate moments begin to
reflect occasional extreme draws---Copper RMSE rises from $0.15$ at
$\mu = 1.2$ to $0.71$ at $\mu = 1.8$---but the centre of the distribution
remains on target. The first-stage $F$ is the quantity that signals
weakness, and it is sensitive to $T$. Copper at $\mu = 1.4$ has $F = 8.7$,
below the conventional Stock--Yogo cutoff of 10, while Crude Oil at the
same $\mu$ has $F = 27.5$ because $T$ is three times larger.
Theorem~\ref{theorem:weak_consistency_aggregate} guarantees that standard normal inference remains valid, and the simulations confirm it.

\paragraph{$\mu > 2$ and the case for Anderson--Rubin.}
For $\mu > 2$ the regime is no longer pinned down by the tail index alone; it
is the $N/T$ trajectory that separates nearly weak from weak identification.
The two panels of Table~\ref{tab:mc_table1} hold $N$ fixed at $29$ and $21$
while $T$ is large ($192$ and $611$), so $N/T \to 0$ and the design sits in
the nearly weak regime, where the estimator is consistent but the
concentration parameter has only a slowly-divergent floor. The genuinely weak
case $N/T \to c > 0$, in which the estimator is inconsistent, is approached in
Table~\ref{tab:mc_table2} by raising $N$ at fixed $T$.
Standard inference begins to break down here. The median $F$ falls to
$2.1$ at $\mu = 2.5$, $1.0$ at $\mu = 3.5$, and $0.5$ at $\mu = 6.0$ in
Copper, and to $6.7$, $3.1$, and $1.0$ in Crude Oil, at or below the
conventional Stock--Yogo cutoff of~$10$. The median estimate drifts away
from the true elasticity, reaching $-0.122$ for Copper and $-0.109$ for
Crude Oil at $\mu = 3.5$ and $-0.107$ and $-0.100$ at $\mu = 6.0$,
against true values of $-0.135$ and $-0.110$. RMSE is dominated by
extreme draws and peaks at $8.1$ for Copper at $\mu = 2.5$ and $101$
for Crude Oil at $\mu = 3.5$. The clearest symptom of the breakdown is
the $t$-based confidence interval: its length grows by more than an
order of magnitude relative to the strong regime, from about $0.02$ to
$0.41$ in Copper and $0.009$ to $0.21$ in Crude Oil between $\mu = 0.3$
and $\mu = 6.0$. Coverage of the $t$-based interval departs from
nominal, drifting to $0.967$ in Copper and $0.969$ in Crude Oil at
$\mu = 6.0$. Each of these is a consequence of the
slowly-divergent concentration parameter floor.

As Section~\ref{subsec:infeasible_AR} shows, the Wald interval scales
with $1/\kappa^2_{\text{conc}}$, which is exactly why its length
explodes and its coverage goes wrong in this regime. The
Anderson-Rubin statistic avoids this dependence entirely. Its pivotal
$\chi^2_1$ limit holds along any admissible $(N,T)$ sequence with
$\mu > 2$ and requires no lower bound on $\kappa^2_{\text{conc}}$. The
simulations therefore reproduce the asymmetry the theory predicts. Wald
inference suffices for $\mu < 2$ and Anderson-Rubin is the appropriate
tool for $\mu > 2$.

\paragraph{Augmenting the cross-section (Table~\ref{tab:mc_table2}).} The
augmented Crude Oil experiments separate the role of $N$ from the role of
$T$ and confirm the rate-based predictions of the theory regime by regime.
In the strong regime, performance is essentially invariant to $N$. RMSE at
$\mu = 0.3$ moves only from $0.0035$ at $N = 21$ to $0.0029$ at $N = 50$
and $0.0025$ at $N = 100$. In the nearly weak regime with
$\mu \in (1,2)$, consistency requires $N/T \to 0$. Increasing $N$ at
fixed $T$ therefore degrades performance, and indeed RMSE at $\mu = 1.4$
rises from $0.036$ at $N = 21$ to $0.045$ at $N = 50$ and $0.060$ at
$N = 100$. For $\mu > 2$ the rate is $\sqrt{T/N}$, so larger $N$ at
fixed $T$ is doubly costly. Raising $N$ at fixed $T$ also lifts $N/T$ away
from zero, moving the design out of the nearly weak regime and toward the
weak regime $N/T \to c > 0$, where the estimator is inconsistent. Consistent
with this, CI length grows uniformly with
$N$ and the median $F$ falls---at $\mu = 6.0$ it declines from $1.0$ at
$N = 21$ to $0.6$ at $N = 100$---while RMSE remains large and
tail-dominated throughout, reaching $1.7$ at $N = 100$.




\section{Conclusion} \label{sec:conclusion}

This paper extends Granular Instrumental Variables to large panels with $N, T \to \infty$ and shows how the granularity of the cross section governs instrument strength. Under a power-law assumption on unit sizes, the tail index $\mu$ together with the $N/T$ trajectory determines the asymptotic behavior of the estimator. Three regimes emerge. The strong regime $\mu \in (0,1)$ recovers the classical $\sqrt{T}$ rate of \citet{Gabaix2024GranularVariables}. The nearly weak regime $\mu > 1$ with $N/T \to 0$, following \citet{Antoine2021GMMIdentification}, delivers consistency and asymptotic normality at the slower rate $\sqrt{T}/N^{\delta}$, where $\delta = \min(1 - 1/\mu, 1/2)$. The weak regime $\mu > 2$ with $N/T \to c$ corresponds to \citet{Staiger1997InstrumentalInstruments}-style local-to-zero asymptotics, and the estimator is inconsistent. For inference, Wald is reliable when $\mu < 2$. For $\mu > 2$, I recommend Anderson--Rubin confidence sets, which remain valid regardless of the $N/T$ trajectory.

In practice the GIV instrument is built from estimated rather than known idiosyncratic shocks. Under the additional growth restriction $\sqrt{T}/N \to 0$, the feasible estimator attains the same convergence rate as the infeasible one, but its asymptotic variance is different. The first-stage estimation contributes an additional term that enters at the same order as the infeasible variance, so the formulas for the standard error are not the same. Valid inference therefore requires standard errors that explicitly account for the first-stage estimation error, and I provide a HAC-consistent variance estimator that does so.

I apply the GIV estimator to estimate short-run demand elasticities for refined copper, crude oil, and natural gas. All three markets have estimated tail indices below one, with crude oil's confidence interval reaching into the nearly weak region. The estimated elasticities are $-0.135$, $-0.109$, and $-0.056$ respectively, all consistent with the inelastic short-run response that industrial commodities are known for.



\bibliography{ref}