Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
100,837 characters · 33 sections · 67 citation commands
Granular Instrumental Variables in Large Panels: Identification and Inference Across Strong, Nearly Weak, and Weak GIV
\addtocontents{toc}{\setcounter{tocdepth}{-10}}
Many questions in economics require estimating structural relationships between aggregate variables, for instance, how asset prices respond to changes in aggregate demand, or how exchange rates react to capital flows. A central challenge is endogeneity. The aggregate regressor is correlated with the structural error. Gabaix2024GranularVariables proposed Granular Instrumental Variables (GIV) as a solution. The key insight is that when a few large units (dominant firms, banks, or funds) disproportionately drive the aggregate, their idiosyncratic shocks can serve as instruments. Several studies apply this idea to asset markets, bank lending, and sovereign risk. However, the existing theory assumes a fixed number of units. This paper extends GIV to large panels while focusing on a core question: what is granularity and how much granularity is enough?
Specifically, Gabaix2024GranularVariables (GK hereafter) consider the canonical supply-demand system
where $d_t$ is the change in aggregate demand for a commodity, $p_t$ the change in market-clearing price, and $y_{it}$ the change in supply of unit $i$. The unit-level supply follows a panel data model with interactive fixed effects, where unit-specific loadings $\lambda_i$ interact with common time factors $F_t$. Market clearing, $d_t = \sum_{i=1}^N S_i y_{it}$, ties the two equations together, with $S_i$ the long-run market share of unit $i$. The object of interest is the demand elasticity $\phi_d$. The endogeneity problem is that $p_t$ responds to the demand shock $\varepsilon_t$.
For known shares, $S$, and the demeaning matrix, $D_N = I_N - \frac{\iota \iota'}{N}$, the Granular Instrumental Variable $z_t = S' D_N u_t$ is the optimal instrument for $\phi_d$.
As illustrated by Gabaix2024GranularVariables and further explained by Gopalan, Nagasawa, and Renault (GNR hereafter), the strength of the instrument depends on the granularity of the setting. That is, some of the shares, $S_{i},i=1,...,N$ have to be different from $\frac{1}{N}$. GK call this a granular setting, and hence the name of the instrument. GK only consider the case where $N$ is fixed and $T \to \infty$.
This asymptotic framework is not appropriate for many empirical applications. For instance, Aldasoro2023TheMarkets has $N = 21$. Chodorow-Reich2021AssetInsulators,Galaasen2020GranularRisk,Ma2021ExpectationsLending have $N > 100$. In these cases, a more appropriate asymptotic framework is one with both $N$ and $T \to \infty$. This necessitates extending the theory of GIV to large panels.
But as $N \to \infty$, granularity requires a careful asymptotic treatment. Granularity is a property of the cross section as it arises from the behavior of $\left( S_{i}\right) _{1\leq i\leq N}$. When $N$ is fixed, in the simplest case, we just need $S_i$ to be different from $\frac{1}{N}$ for some $i$ for the instrument to be valid. But as $N \to \infty$, the instrument can weaken if the cross-section is not concentrated enough. Thus, as we extend the theory of GIV to large panels, we need to formalize the idea of granularity and identify how it affects instrument strength.
This paper extends the theory of Granular Instrumental Variables to large panels $(N,T \to \infty)$ with a formal treatment of granularity, making three contributions. First, I formalize granularity by modeling unit sizes as draws from a power-law distribution. The tail index $\mu$ measures how much the largest units stand out, and serves as a single sufficient statistic for instrument strength. Second, I characterize how instrument strength varies with $\mu$ and the $N/T$ trajectory, identifying three regimes with distinct implications for consistency, convergence rates, and inference. Third, I provide the correct asymptotic theory for the feasible GIV estimator which I construct from estimated rather than known idiosyncratic shocks. I then illustrate the theory empirically, using GIV to estimate the short-run demand elasticities of three major commodities (copper, crude oil, and natural gas).
Dominant units (granularity) arise naturally when unit sizes follow a power-law distribution. This is well documented across very different empirical settings such as firm sizes, city sizes, and bank assets Gabaix2009PowerFinance,Axtell2001ZipfSizes. Hence I model that the size of the unit, $\s_i$ comes from a power-law distribution. That is $\p(\s_i > s_i) = c s_i^{-\mu}$, $\mu >0$. From the observed sizes, we construct the shares as $S_i = \frac{\s_i}{\sum_{j=1}^N \s_j}$.
This tail index governs the strength of the instrument. Together with the $N/T$ trajectory, it delivers three regimes.
When a few units dominate the aggregate, the instrument is strong. This is the case $\mu \in (0,1)$, where the tail is heavy enough that a handful of units make up a non-vanishing share of the aggregate. The GIV estimator is consistent and asymptotically normal at the standard $\sqrt{T}$ rate. This is the classical strong instrument regime.
When large units stand out but do not dominate, the instrument weakens without breaking. This is the case $\mu > 1$ with $N/T \to 0$. The largest units are still big enough that their idiosyncratic shocks survive aggregation, but their influence fades as $N$ grows. Identification is nearly weak in the sense of Antoine2021GMMIdentification. The GIV estimator remains consistent and asymptotically normal, now at the slower rate $\sqrt{T}/N^{\delta}$, where $\delta = \min(1 - 1/\mu, 1/2)$. The cross-sectional signal is diluted, but enough time periods recover it.
When units are comparable in size and none stands out, the instrument is weak. This is the case $\mu > 2$ with $N/T \to c$. No unit is systematically larger than the rest, so the granular variation the instrument relies on vanishes, and time-series information no longer compensates for the cross-sectional dilution. The estimator is inconsistent and identification is weak in the sense of Staiger1997InstrumentalInstruments.
For inference, Wald confidence intervals are reliable outside the weak regime, that is when $\mu < 2$. When $\mu > 2$, the instrument can be arbitrarily weak. Its strength carries no guaranteed lower bound as $N$ and $T$ grow, so the normal approximation is unreliable in finite samples. I therefore recommend inverting the Anderson--Rubin test of Anderson1949EstimatorsEquations. Its $\chi^2_1$ limiting distribution holds no matter how weak the instrument, so the recommendation remains valid regardless of the $N/T$ trajectory and covers the weak regime as a special case.
The results above assume that the idiosyncratic shocks $u_{it}$ entering the GIV are known. In practice, we do not observe them and must estimate them by removing the factor structure from the panel. When $N$ is fixed, we can consistently estimate only the common factors. Consistency also requires that the idiosyncratic shocks $u_{it}$ are homoskedastic. With $N,T \to \infty$, however, we consistently recover both factors and loadings. I show that the feasible GIV, constructed from estimated residuals, attains the same convergence rate as the infeasible instrument across all regimes, subject to the additional growth restriction $\sqrt{T}/N \to 0$. The asymptotic variance is not the same. The first-stage estimation contributes an additional term at the same order as the infeasible variance, so inference must use standard errors that account for this generated-regressor contribution.
In the empirical application, copper and natural gas fall into the strong instrument regime at $95\%$ confidence, while crude oil extends into the nearly weak regime. GIV corrects the biased OLS estimates and leads to economically plausible negative demand elasticities: $-0.135$ for copper, $-0.109$ for crude oil, and $-0.056$ for natural gas.
Finally, in simulation exercises calibrated to the panel data from copper and crude oil, I study the sensitivity of the GIV estimates to granularity. As expected, the strong regime has very low bias, tight confidence sets, and the correct coverage. In the weak regime with $\mu > 2$, we observe substantial bias together with confidence intervals that explode in length, so that the Wald interval over-covers rather than attaining its nominal level.
I organize the rest of the paper as follows. In Section (ref), I set up the model and formalize granularity through power-law size distributions. I present the main asymptotic results when the factor structure is observed and when it is unobserved in Sections (ref) and (ref) respectively. In Section (ref), I apply the theory to estimate short-run demand elasticities for major commodities. I study the small-sample behavior of the estimator in Section (ref) through Monte Carlo simulations and conclude in Section (ref). I close this section by placing the paper in the context of the related literature.
This paper contributes to several strands of the literature.
GIV theory. Gabaix2024GranularVariables introduce the GIV framework under fixed $N$ and $T \to \infty$, establishing consistency, asymptotic normality, and the theory for the estimator.
Banafti2022InferentialDimensions extend GIV to high dimensions, allowing $N$ to grow with $T$, and derive the asymptotic distribution of the feasible estimator under the same power law assumption I impose in Assumption (ref). They restrict attention to the strong instrument regime ($\mu < 1$), and the asymptotic distribution they obtain differs from the one I derive in this paper. Their derivation requires an additional condition---their Assumption 4(iii)---on the behavior of shares as $N \to \infty$. In Section (ref), I show that this condition is incompatible with the power law assumption: the two cannot simultaneously hold. Their proof therefore breaks down in the power law setting, and their asymptotic distribution result does not apply. Establishing the asymptotic distribution of the GIV estimator in large panels, with both $N$ and $T$ tending to infinity, therefore remains an open problem and this paper resolves that.
Qian2023Heterogeneity-robustInstruments constructs heterogeneity-robust granular instruments that remain valid when the structural parameter varies across a fixed number of units. Baumeister2023AVariables propose a full-information approach that jointly estimates the factor structure and structural parameter under parametric assumptions. This method becomes untenable as $N \to \infty$.
On the empirical side, GIV has been applied to asset markets Chodorow-Reich2021AssetInsulators, bank credit risk Galaasen2020GranularRisk, bank lending Ma2021ExpectationsLending, sovereign bonds Aldasoro2023TheMarkets, exchange rates Hau2022GlobalRates, and monetary policy Holm-Hadulla2024GranularPolicy. These applications involve $N$ ranging from 21 to well over 100, underscoring the need for a large-panel theory.
Factor models. The feasible GIV requires estimating idiosyncratic shocks by removing the common factor structure. Bai2003InferentialDimensions establishes the convergence rates for factors and loadings estimated by principal components when $N, T \to \infty$. I use these results to show that the estimation error affects the asymptotics of the GIV. We can determine the number of common factors using information criteria such as those of Bai2002DeterminingModels or Ahn2013EigenvalueFactors.
Power laws. Power-law size distributions are well documented in firm sizes Axtell2001ZipfSizes, city sizes, and financial returns Gabaix2009PowerFinance. I build on this regularity. The same power-law tail that drives granularity also determines whether the GIV is a strong or nearly weak instrument.
Instrument strength. Staiger1997InstrumentalInstruments formalize weak instruments by modeling the first-stage coefficient as local to zero. In my setting, weakness is structural rather than local-to-zero. It is comparable to instrument weakness in large markets for differentiated products as in Armstrong2016LargeSupply. Antoine2021GMMIdentification provide a nearly weak identification framework where identification strength vanishes, but slowly enough for consistency at a rate slower than $\sqrt{T}$. My nearly weak regime ($\mu > 1$) maps directly onto their framework. The local-to-zero asymptotics of Staiger1997InstrumentalInstruments emerge as a special case under $N/T \to c$ with $\mu > 2$. For inference when $\mu > 2$, I construct Anderson--Rubin confidence sets Anderson1949EstimatorsEquations that remain valid regardless of instrument strength.
This section studies when the Granular Instrumental Variable (GIV) is strong enough to identify an aggregate demand elasticity. The instrument $z_t$ is a share-weighted average of unit-level supply shocks. It is valid by assumption (Assumption (ref) delivers exogeneity), but its relevance is not guaranteed. Informativeness depends on the presence of dominant units. For enough individuals, the share $S_i$ must sit far from $1/N$.
Proposition (ref) formalises this. The instrument stays fixed only when unit sizes are drawn from a fat-tailed distribution. Otherwise it decays at a rate set by the tails. Weakness is therefore not something I impose through a local-to-zero parameterization. It arises structurally, from the heavy-tailed distribution of individual sizes.
The rate of decay sorts the design into three regimes. When the instrument does not decay, it is strong. When it decays slower than $\sqrt{N}$, it is nearly weak. When it decays at the $\sqrt{N}$ rate, it is weak in the classical sense of Staiger1997InstrumentalInstruments. These regimes govern whether the elasticity can be estimated consistently, at what rate, and whether standard inference stays reliable. I develop those consequences for the known factor structure in Section (ref) and for the unknown factor structure in Section (ref). I begin by stating the model and the assumptions behind Proposition (ref).
I study the estimation of aggregate demand elasticities for a commodity. At each date $t = 1, \ldots, T$, change in aggregate demand $d_t$ is governed by the structural equation
where $p_t$ is the change in market-clearing price, $\phi_d$ is the demand elasticity, and $X_t^d$ are observed controls uncorrelated with $\varepsilon_t$. The structural parameter of interest is the demand elasticity, $\phi_d$.
Further, we observe changes in individual level supply/production of the commodity, $y_t = \left( y_{it} \right)_{1 \leq i \leq N}$. The individual supply is governed by the structural equations
where $\phi_s$ is the supply elasticity, and $X_{it}^y$ are observed controls uncorrelated with both $\varepsilon_t$ and $u_t$. The unit-level supply in (ref) follows a panel data model with interactive fixed effects (common shocks $F_t$ that load heterogeneously across units through $\lambda_i$). The factors $F_t$, the loadings $\lambda_i$, and the idiosyncratic shocks $u_{it}$ are all unobserved. Stacking across units,
where $e_N$ is the $N$-vector of ones.
Market clearing links the two equations: the aggregate change in supply equals the change in demand, so $d_t = \sum_{i=1}^N S_i y_{it} \defeq y_{St}$, where $S_i$ is the equilibrium market share of unit $i$. We assume that the equilibrium market shares are determined by a different structural process, so that at the frequency of interest $S_i$ is independent of all changes in demand and supply. The market share is linked to the individual sizes, $\s_i$ as $S_i = \frac{\s_i}{\sum_j \s_j}$.
For ease of exposition, I present the theory without accounting for the observed controls. When such controls are present, we can partial them out; mutatis mutandis, the theory applies equally to the resulting residualized variables by the Frisch-Waugh-Lovell theorem. See Appendix (ref) for details.
The model is characterized by the following assumption.
By Assumption (ref), $u_t$ is a valid instrument for the estimation of $\phi_d$. The moment condition delivers exclusion; whether $u_t$ is relevant enough to identify $\phi_d$ as $N \to \infty$ is the subject of Proposition (ref).
We do not observe $u_t$. In small panels, we can consistently estimate $\Lambda$ (see GK and GNR). As $T \to \infty $ (infeasible), we can extract from data, $M_{\Lambda}y_t = M_{\Lambda}u_t$, where
GK and GNR show that the optimal GIV in the case of linear conditional expectation is given by $z_t = S' M_{\Lambda} y_t = S' M_{\Lambda} u_t$. In large panels ($N$ and $T$ large), we can go further and estimate both the common factors and the factor loadings. We can demean (which kills $F_{1t}$ via $D_N e_N = 0$) and estimate consistent $\hat\Lambda, \hat F$ by PCA. We then subtract the estimated common component $\hat C_t = \hat{\Lambda} \hat{F}_t$ from the observed data to get $\hat{u}_t = D_N y_t - \hat{C}_t$. As $T \to \infty$ (infeasible), we have
where $\tilde{C}_t= \tilde{\Lambda} \tilde{F}_t$ and $D_N = I_N - \frac{e_N e_N'}{N}$ is the demeaning matrix with $e_N$ being the $N$-dimensional vector of ones. For any $N \times K$ matrix $X$, I write $\bar X \defeq D_N X$ for its demeaned version. Large panels allow consistent estimation of common factors, leading to the moment restriction
The Granular Instrumental variable associated with the above moment condition is
In the shorthand just introduced, the infeasible instrument is $z_t = S' \bar u_t$. In the general case, we do not directly observe $C_t$, and it needs to be estimated. For clarity, I will first develop the asymptotic theory for the infeasible instrument before presenting the theory for the feasible one. To account for granular settings, we assume that the share vector, $S$ is random and follows the power law. We similarly make suitable assumptions on the other cross-sectional variable, the factor loadings. We assume that the cross-section $(\s_i, \tilde\Lambda_i)$ is i.i.d., sizes have power-law tails, loadings have bounded fourth moments, and both are independent of the time-series shocks.
Independence of the absolute sizes is not a restrictive assumption. The individual size is set in the long-term equilibrium. The shocks are all short-term in nature, and do not affect the long-term equilibrium. The same applies to the joint i.i.d.\ assumption of the vector, $(\s_i, \tilde{\Blambda}_i')'$. When loadings are treated as non-random (as is common), the i.i.d. clause reduces to i.i.d. sizes. The final assumption is standard in the factor literature when the factor loadings are random Bai2003InferentialDimensions. As sizes are observed in equilibrium, we can construct the individual shares as
Under these assumptions, we will see how the tail index affects the granularity of the setting and hence the instrument strength.
In large panels, weakness arises due to the behavior of the tail index of the size variable. I formally state that idea in the first Proposition. Banafti2022InferentialDimensions had formally stated the result for $\mu \in (0,1)$. I extend it to all values of $\mu$, following Gabaix2011TheFluctuations.
Proposition (ref) delineates three regimes of instrument strength, governed jointly by the tail index $\mu$ and the relative growth of $N$ and $T$.
For $\mu \in (0,1)$, the instrument does not decay. Identification is strong, and this corresponds to the strong instrument dynamics of Gabaix2024GranularVariables.
For $\mu > 1$, the instrument vanishes as $N \to \infty$. For $\mu \in (1,2)$, $z_t = O_{\p}(\frac{1}{N^{1 - 1/\mu}})$. For $\mu > 2$, $z_t = O_{\p}(\frac{1}{\sqrt{N}})$. Under $N/T \to 0$, instrument strength accumulates fast enough to identify the structural parameter and conduct inference. The estimator is consistent and asymptotically normal at the slower rate $\sqrt{T}/N^{\delta}$, where $\delta = \min(1 - 1/\mu, 1/2)$. Following Antoine2021GMMIdentification, I call this nearly weak identification.
When $\mu > 2$ and $N/T \to c > 0$, time-series information no longer overtakes the cross-sectional dilution of the instrument. The estimator is inconsistent and identification is weak in the sense of Staiger1997InstrumentalInstruments.
Within the nearly weak regime, $\mu = 2$ is a boundary for the concentration parameter. For $\mu \in (1,2)$, the concentration parameter has a polynomial floor in $N$. For $\mu > 2$, the floor is only slowly divergent and can grow arbitrarily slowly along admissible $(N,T)$ sequences. The Gaussian approximation is therefore reliable for $\mu < 2$ but not for $\mu > 2$, where I recommend Anderson--Rubin confidence sets. The same Anderson--Rubin procedure remains valid under the weak identification regime, so the recommendation handles both $\mu > 2$ subcases at once.
Hence weakness arises directly from the structure of the setting, similar to Armstrong2016LargeSupply. The heavy-tailed concentration of shares delivers an idea of weakness without imposing the local-to-zero assumption of Staiger1997InstrumentalInstruments. That assumption emerges only as a special case under $N/T \to c$ with $\mu > 2$.
We will first consider the infeasible estimation of the structural parameters. The estimation is infeasible as we assume that the factor structure is available to us. This helps us fix the basic ideas. And in the subsequent section, we will deal with the feasible estimation where we need to estimate the factor structure. Throughout, define $\delta = \min(1 - 1/\mu, 1/2)$, so that $\delta = 1 - 1/\mu \in (0, 1/2)$ when $\mu \in (1,2)$ and $\delta = 1/2$ when $\mu > 2$. Proposition (ref) can then be written compactly as $z_t = O_{\p}(N^{-\delta})$ for $\mu > 1$.
In this section, I assume that we know the factor structure, specifically the factor loadings. This is primarily for exposition but includes some settings of practical interest. One such case is when we have only time fixed effect, that is, $\Lambda = e_N$. In this case, we can perfectly recover the demeaned idiosyncratic shocks to construct the instrument.
$D_N y_t = D_N u_t$ perfectly recover the demeaned idiosyncratic shocks.
Another case is when we do not directly observe the loadings, but we have a parametric form, $\tilde{\lambda}_i = X_i \dot{\lambda}$ where $\tilde{\lambda}_i$ and $X_i$ are $r-1$ dimensional vectors and $\dot{\lambda}$ is a $r \times r$ matrix which is invariant across $i$. This implies $\tilde{\lambda} = X \dot{\lambda}$. We observe the vector of characteristics, $X$. In this case,
Thus, $M_X D_N y_t$ recovers the demeaned idiosyncratic shocks and the optimal instrument is $S'M_X D_N u_t$.
However, most of the empirical examples of GIV assumes a more general factor structure. This requires estimation of the factor structure in a first stage before we construct the instrument. See Section (ref) for the analysis of this general case.
From the demeaned idiosyncratic shocks, we construct the infeasible instrument as $z_t = S'(D_N y_t - \tilde{C}_t)$. Proposition (ref) gives the behavior of the infeasible instrument for different values of the tail index of the size variable. I find three regimes. When $\mu \in (0,1)$, identification is strong and the estimator is $\sqrt{T}$-consistent and asymptotically normal. When $\mu > 1$ and $N/T \to 0$, identification is nearly weak in the sense of Antoine2021GMMIdentification. The estimator is consistent and asymptotically normal at the slower rate $\sqrt{T}/N^{\delta}$, where $\delta = \min(1 - 1/\mu, 1/2)$. When $\mu > 2$ and $N/T \to c$, identification is weak in the sense of Staiger1997InstrumentalInstruments and the estimator is inconsistent. For inference, Wald confidence intervals are reliable when $\mu < 2$. For $\mu > 2$, I recommend Anderson--Rubin confidence sets, which remain valid regardless of the $N/T$ trajectory.
To establish these results, we require assumptions on the idiosyncratic shocks. The large panel setting allows us to accommodate a richer structure for the time series and cross-sectional dependence than the fixed-$N$ (small-panel) case, which requires i.i.d.\ samples with no cross correlation.
These are standard in the factor literature and relax the independence restriction in the shorter panel GIV literature.
With the dependence structure in place, I now turn to the behavior of the infeasible GIV estimator. Proposition (ref) identifies three regimes governed by the tail index $\mu$ and the $N/T$ trajectory. I analyze them in turn, beginning with the strong regime ($\mu \in (0,1)$), where the Herfindahl does not vanish and the estimator is $\sqrt{T}$-consistent and asymptotically normal.
Recall the aggregate demand equation of interest is
The moment condition for the estimation of the demand parameter is $\e[(y_{St} - \phi_d p_t )z_t] = 0$. Thus the GIV estimator of the demand (aggregate) structural parameter is
I will now formally state the results for the strong regime. In the small-panel GIV framework of Gabaix2024GranularVariables, $N$ is fixed, the shares and factor loadings are treated as constants, and consistency and asymptotic normality follow from standard IV arguments. Moving to the large-panel setting introduces four complications. First, $N \to \infty$, so the instrument $z_t = S'\bar{u}_t$ is a growing weighted sum whose behavior depends on the concentration of the shares. Second, the shares $S$ are now random, drawn from a power-law distribution, so the Herfindahl $S'S$ is itself a random variable whose order must be established. Third, the factor loadings $\tilde{\Blambda}_i$ are random, and the share-weighted loading $S'\tilde{\Lambda}$ must be shown not to contaminate the instrument. Fourth, the central limit theorem for the numerator $T^{-1/2}\sum_t z_t \varepsilon_t$ is itself non-standard. The summand combines a share-weighted cross-sectional sum with time-series dependence, and its moments must be bounded uniformly in $N$.
When $\mu \in (0,1)$, all four complications are resolved by the heavy tails. Proposition (ref) shows that $S'S = O_{\p}(1)$, so the instrument does not degenerate as $N$ grows. Proposition (ref) shows that $S'\tilde{\Lambda} = O_{\p}(1)$ as well, so the common-factor contamination remains bounded. Corollary (ref) delivers the CLT for the numerator $T^{-1/2}\sum_t z_t \varepsilon_t$. The heavy-tail concentration of $S$ is what keeps its moments bounded uniformly in $N$. Under these conditions, the GIV estimator retains $\sqrt{T}$-consistency and asymptotic normality.
I now turn to the case $\mu > 1$. Define $\delta = \min(1 - 1/\mu, 1/2)$. The heavy-tail concentration of the shares is weaker than in the strong regime, and Proposition (ref) pins down by how much. The Herfindahl satisfies $S'S = O_{\p}(N^{-2\delta})$, so the instrument $z_t = S'\bar{u}_t$ no longer has order one. It dilutes as $N$ grows. Only the rescaled instrument $N^{\delta} z_t$ has a non-degenerate limit. Under $N/T \to 0$, time-series information accumulates fast enough that the GIV estimator remains consistent and asymptotically normal at the slower rate $\sqrt{T}/N^{\delta}$. Following Antoine2021GMMIdentification, I call this nearly weak identification.
Two of the four complications from the strong regime resolve exactly as before. The share-weighted loading $S'\tilde{\Lambda}$ remains controlled by Proposition (ref), and the CLT for the numerator goes through under the same mixing conditions. The other two now carry an explicit $N$-dependence inherited from the dilution of the instrument. This is what slows the rate of convergence and forces the requirement $N/T \to 0$. Time-series information must accumulate fast enough to overcome the cross-sectional dilution.
The boundary $\mu = 2$ separates two qualitatively different dilution patterns. For $\mu \in (1,2)$, the dilution exponent $\delta = 1 - 1/\mu$ lies in $(0, 1/2)$ and the Herfindahl is governed by a stable law. For $\mu > 2$, the size variable has finite second moment, so the Herfindahl is governed by the law of large numbers rather than a stable law, and the dilution exponent caps at $\delta = 1/2$. The asymptotic statement of the theorem below covers both cases at once. What changes across the boundary is the behavior of the concentration parameter, which in turn drives the choice between Wald and Anderson--Rubin inference.
The choice between Wald and Anderson--Rubin inference depends on the behavior of the concentration parameter. The concentration parameter $\kappa^2_{\text{conc}}$ measures the signal-to-noise ratio of the first-stage moment Stock2002AMoments. For the GIV estimator,
and the asymptotic standardization $\kappa_{\text{conc}}(\hat\phi_d - \phi_d) \xrightarrow{d} \n(0,1)$ summarizes the rate at which the estimator concentrates on $\phi_d$. The Gaussian approximation is reliable in finite samples when $\kappa^2_{\text{conc}}$ is large. The first-stage $F$-statistic is its sample analog.
Parametrize any admissible sequence as $T = N \cdot f(N)$ with $f(N) \to \infty$. Then
For $\mu \in (1,2)$ the polynomial floor $N^{1-2\delta}$ with $1 - 2\delta > 0$ does not depend on $f$. So $\kappa^2_{\text{conc}}$ clears any fixed weak-instrument threshold for moderate $N$. For $\mu > 2$ the floor is just $f(N)$, which can grow arbitrarily slowly. The boundary $\mu = 2$ is precisely where the polynomial floor disappears.
\paragraph{Wald inference for $\mu \in (1,2)$.} The polynomial floor in $\kappa^2_{\text{conc}}$ keeps Wald inference reliable. The practitioner does not need to know $\delta$. The asymptotic variance is
where $\hat{\Gamma}_{zp}$ and $\hat{V}_{z\varepsilon}$ are sample analogs of $\Gamma_{zp}$ and $V_{z\varepsilon}(S)$. The studentization in Theorem (ref) absorbs the rate automatically, so $\delta$ never enters the formula. This parallels GMM under near-weak identification Antoine2021GMMIdentification. Standard Wald confidence intervals remain valid, even though the rate is slower than in the strong regime.
\paragraph{Anderson--Rubin inference for $\mu > 2$.} The polynomial floor disappears. The concentration parameter can grow arbitrarily slowly, so its realized value in finite samples need not be large. The estimator has poor finite-sample performance even though it is asymptotically normal. I recommend inverting the Anderson--Rubin (AR) test of Anderson1949EstimatorsEquations to construct the confidence interval.
Anderson--Rubin avoids the dependence on $\kappa^2_{\text{conc}}$ entirely. At hypothesized value $\phi_0$, the sample moment is
which under $H_0: \phi_d = \phi_0$ reduces to $T^{-1}\sum_t z_t \varepsilon_t$ and, by Theorem (ref), satisfies $\sqrt{T}\,g_T(\phi_0)/\sqrt{V_{z\varepsilon}(S)} \xrightarrow{d} \n(0,1)$. The AR statistic and the corresponding $1-\alpha$ confidence set are
The $\chi^2_1$ limit holds uniformly across $\mu > 2$ and $N/T \to 0$. It requires no condition on the rate at which $\kappa^2_{\text{conc}} \to \infty$. It needs only a consistent estimator of $V_{z\varepsilon}(S)$, the variance of the moment.
The same AR procedure remains valid when $N/T \to c$ and identification is weak in the Staiger--Stock sense. AR does not depend on consistency of $\hat\phi_d$. It uses only the sample moment evaluated at the hypothesized value, so it covers the failure mode without modification.
Section (ref) took the factor structure as known. In practice it is not, and we replace the infeasible common component $\tilde C_t$ with the principal-component estimate $\hat C_t$ of Bai2003InferentialDimensions. This places GIV in the constructed-regressor setting: the first-stage estimation error propagates to the GIV moment. Additionally, consistent inference now requires $\sqrt{T}/N \to 0$. This section formalises the effects of the first stage estimation and re-establishes the convergence results of Section (ref) under the feasible instrument.
We are interested in estimating the structural parameter $\phi_d$ in:
We need to construct the instrument from the supply side equation
From the previous definitions, $\Lambda = [\boldsymbol{1}_{N} \enspace \tilde{\Lambda} ]$ and $F_t = [F_{1t},\, \tilde F_t']'$,
Thus $D_Ny_t$ has a factor structure. In large panels, we can consistently estimate both $\tilde{\Lambda}$ and $\tilde{F}$ Bai2003InferentialDimensions. As $T > N$, we estimate $\hat{\Lambda}$ using principal components. The first order condition of the PCA objective concentrates out $\hat{F} = \Bar{Y} \hat{\Lambda}/N$.
From these two estimators, we have the estimator for the common component $\tilde{C}_{it} = \tilde{\Blambda}_i' \tilde{F}_t$ as $\hat{C}_{it} = \hat{\Blambda}_i' \hat{F}_t$. Call the corresponding vector, $\hat{C}_t$. From this estimator, we can construct another feasible instrument:
where $\hat{u}_t = D_N y_t - \hat{C}_t$ is the estimated residual. Gabaix2024GranularVariables proposed a different form of the instrument.
The two formulations are equivalent: the PCA first-order condition $\hat F_t = \hat\Lambda' D_N y_t / N$ implies $\hat C_t = (I - M_{\hat\Lambda}) D_N y_t$, so $\hat z_t = \hat z_t^{\text{GK}}$ \footnote{Explicitly, $\hat u_t = D_N y_t - \hat\Lambda \hat F_t = [I - \hat\Lambda \hat\Lambda'/N] D_N y_t = [I - \hat\Lambda(\hat\Lambda'\hat\Lambda)^{-1}\hat\Lambda'] D_N y_t = M_{\hat\Lambda} D_N y_t$, where the second equality uses the PCA normalization $\hat\Lambda'\hat\Lambda/N = I_{r-1}$.}. I work with $\hat C_t$ rather than $M_{\hat\Lambda} D_N y_t$ in the asymptotic analysis. $M_{\hat\Lambda} = \hat\Lambda(\hat\Lambda'\hat\Lambda)^{-1}\hat\Lambda'$ is a non-linear product of the estimator $\hat\Lambda$. Hence its asymptotic linear form involves derivatives of non-linear transformations of the estimator and is mathematically complex. $\hat C_t$ admits a much cleaner asymptotic linear expansion derived in Appendix (ref). The instrument is therefore
In this section, we will see that the estimation of the factor structure in the first stage places an additional condition on the rates of convergence of $N$ and $T$. Consistency and asymptotic normality of the estimates of the structural parameters using the infeasible instrument for the nearly weak regime in Theorem (ref) require $\frac{N}{T} \to 0$. But when we use the feasible instrument after estimation of the factor structure, consistency and asymptotic normality require an additional restriction on the rates of $N$ and $T$, which is that $\frac{\sqrt{T}}{N} \to 0$.
Similar to the previous section, I state the results separately for the strong and nearly weak regimes. For $\mu > 2$, I provide Anderson--Rubin confidence sets. The estimation of the factor structure requires a number of additional assumptions on the factor structure which I state in the next sub-section.
The proofs adapt Bai2003InferentialDimensions to the GIV setting: Appendix (ref) treats the factor loadings, Appendix (ref) the factors, and Appendix (ref) combines the two to obtain an influence-function expansion of $\hat C_t - \tilde C_t$. I quote that expansion here and trace it through the GIV moment.
From (ref), we can see that the difference between the estimated common component and the true common component is
The two leading terms in (ref) have distinct origins and behave differently in $N$ and $T$. The first term, $\tilde F_t' (\tilde F'\tilde F/T)^{-1} (1/T)\sum_m \tilde F_m \bar u_m$, is the error transmitted from estimating the loadings $\tilde\Lambda$. It involves a $T$-direction sample average between the factors and the idiosyncratic shocks, reflecting that $\hat\Lambda$ is identified from time-series variation. The second term, $\tilde\Lambda (\tilde\Lambda'\tilde\Lambda/N)^{-1} (1/N)\sum_j \tilde\Blambda_j \bar u_{jt}$, is the error from estimating the factors $\tilde F$. It is a loading-weighted cross-section average of the shocks at $t$, reflecting that $\hat F_t$ is identified from cross-sectional variation.
Notation: $\bar u_m$ is the $N$-vector of demeaned shocks at time $m$ (so $\bar u_m = D_N u_m$), with $j$-th entry $\bar u_{jm}$.
The feasible estimator is $\hat{z}_t = S'[D_N y_t - \hat{C}_t] = z_t - S'[\hat{C}_t - \tilde{C}_t]$. The difference between the estimator and the true value is
Compared to the infeasible case, we need to analyse the additional terms in the numerator and denominator. By Lemma (ref),
By Lemma (ref),
In every regime, the first-stage estimation contributes an additional term to the asymptotic distribution. This term enters at the same order as the corresponding infeasible quantity. The new term is $O_{\p}(1)$ in the strong regime ($\mu \in (0,1)$) and $O_{\p}(N^{-\delta})$ in the nearly weak regime ($\mu > 1$), where $\delta = \min(1 - 1/\mu, 1/2)$. The rate of convergence of the feasible estimator therefore coincides with that of the infeasible estimator from Section (ref). The rates are $\sqrt{T}$ and $\sqrt{T}/N^{\delta}$ respectively. Only the asymptotic variance changes, picking up an additive contribution from the first stage.
For $\mu > 2$, the concentration parameter of the feasible estimator still has only a slowly-divergent floor, just as in the infeasible case. For the same reasons given in Section (ref), I therefore recommend Anderson--Rubin confidence sets there. Now we can formally state the results on consistency and asymptotic normality of the feasible GIV estimator.
In the strong regime, the heavy-tail concentration of the shares delivers $\sqrt{T}$-consistency and asymptotic normality, just as in Theorem (ref). The first-stage estimation contributes an $O_{\p}(1)$ term to the asymptotic distribution, which leaves the rate unchanged but affects the asymptotic variance. I formally state this result in Theorem (ref)
Banafti2022InferentialDimensions also study large-panel GIV in the strong-instrument case. Their Theorems 2 and 4 conclude that the first-stage estimation of the instrument has no impact on the asymptotic variance of the GIV estimator. My Theorem (ref) reaches the opposite conclusion. I show that the first-stage estimation has a first-order contribution to the asymptotic variance of the GIV estimator.
The difference arises due to two reasons. The first is an assumption they impose. I show that this assumption is not compatible with the power law setting and hence needs to be relaxed. The second is the order of one term in their Lemma 2. They show that this term is insignificant in the limit. But I show that this result relies on very strict assumptions which even Banafti2022InferentialDimensions do not formally impose.
The first point is their Assumption 4(iii). This assumption is in addition to the power law assumption, identical to our Assumption (ref). However, I show that this Assumption 4(iii) is not compatible with the power law setting. That is, with the individual sizes $\s_i$ following the power law, Assumption 4(iii) is not possible.
The second point is regarding the order of a term in the asymptotic form of the estimator. This term vanishes only under very strict assumptions. In the general case, this term adds to the asymptotic variance of the estimator. Hence this term directly drives the difference. I develop both points in Appendix (ref).
In the nearly weak regime, the feasible estimator inherits the rate $\sqrt{T}/N^{\delta}$ of Theorem (ref), where $\delta = \min(1 - 1/\mu, 1/2)$. The first-stage estimation contributes an $O_{\p}(N^{-\delta})$ term to the asymptotic distribution. This matches the order of the infeasible quantities and leaves the rate unchanged. I formally state this result in Theorem (ref).
We need a consistent estimator of the asymptotic variance for inference. I propose a Heteroskedasticity and Auto-Correlation (HAC) consistent estimator for the asymptotic variance. Define the regime-specific rescaling
Now define the estimator of asymptotic variance
where $b_T$ is the bandwidth and $w(x)$ is a kernel function, $w: \mathbb{R}^+ \to [0,1]$ such that $w(x) = 0$ for $x>1$ and $w(0) = 1$. $\tilde{\varepsilon}_t = \hat{\varepsilon}_t - \frac{1}{T} \sum_t \hat{F}'_t \hat{\varepsilon}_t \cdot \hat{\Sigma}_{\tilde{F}}^{-1} \cdot \hat{F}_t $, where $\hat{\varepsilon}_t = d_t - \hat{\phi}_d p_{t}$ and $\hat{\Sigma}_{\tilde{F}} = \left[ \frac{\hat{F}' \hat{F}}{T} \right]$.
The inference choice mirrors Section (ref). The concentration parameter of the feasible estimator has a polynomial floor in $N$ when $\mu \in (1,2)$ and only a slowly-divergent floor when $\mu > 2$. Wald is reliable for the former. I recommend Anderson--Rubin for the latter.
\paragraph{Wald inference for $\mu \in (1,2)$.} The polynomial floor keeps Wald inference reliable. The asymptotic variance is given by the HAC estimator of Proposition (ref), which absorbs the rate $\delta$ automatically through the rescaling $a_N$. Standard Wald confidence intervals are valid.
\paragraph{Anderson--Rubin inference for $\mu > 2$.} The concentration parameter of the feasible estimator can grow arbitrarily slowly, so its realized value in finite samples need not be large. The estimator therefore has poor finite-sample performance even though it is asymptotically normal. I recommend inverting the Anderson--Rubin (AR) test of Anderson1949EstimatorsEquations.
The Anderson--Rubin test is evaluated at the null, so it never uses the GIV estimate $\hat{\phi}_d$. Two features follow, and together they make the $N/T$ condition irrelevant. First, under $H_0: \phi_d = \phi_d^0$, the residual $d_t - \phi_d^0 p_t = \varepsilon_t$ is exact, so no first-stage estimation of $\phi_d$ enters the variance estimator. Second, the statistic is self-normalising, and the rate $a_N$ cancels between the numerator and the variance estimator, so the test refers neither to the rate $\delta$ nor to $N/T$. The only condition that survives is $\sqrt{T}/N \to 0$, from estimating the factor structure in $\hat{z}_t$.
Consider the null hypothesis, $H_0: \phi_d = \phi_d^0$. Define the null-imposed residual $\varepsilon_t(\phi_d^0) = d_t - \phi_d^0 p_t$ and its factor-projected version $\tilde{\varepsilon}_t(\phi_d^0) = \varepsilon_t(\phi_d^0) - \frac{1}{T} \sum_t \hat{F}_t' \varepsilon_t(\phi_d^0) \cdot \hat{\Sigma}_{\tilde{F}}^{-1} \hat{F}_t$. Define the Anderson-Rubin statistic Anderson1949EstimatorsEquations,
where $\hat{V}^H_{z \bar{\varepsilon}}(\phi_d^0)$ is the HAC estimator defined in (ref), specialised to $\mu > 2$ so that $a_N^2 = N$, and built from the null-imposed residual,
The $N/T$ in the numerator cancels the $a_N^2/T = N/T$ in $\hat{V}^H_{z \bar{\varepsilon}}(\phi_d^0)$, so the statistic is self-normalised and carries no reference to the rate. It is consistent for the conditional asymptotic variance $V_{z \bar{\varepsilon}}(S)$, with $\bar{\varepsilon}_t = \varepsilon_t - \e[\tilde{F}_t' \varepsilon_t] \Sigma_{\tilde{F}}^{-1} \tilde{F}_t$, by the argument of Proposition (ref) with the residual reduction now exact. Inversion of this test yields a confidence region of the correct size.
Theorem (ref) places no condition on $N/T$. It therefore covers the weak-identification case $N/T \to c > 0$, where $\hat{\phi}_d$ is inconsistent. The test does not depend on consistency of $\hat{\phi}_d$. It uses only the moment and the variance, both evaluated at the hypothesised value.
We apply the GIV estimator to three commodity markets---refined copper, crude oil, and natural gas---to estimate their respective price elasticities of demand. Each market provides a global supply panel whose country-level idiosyncratic shocks serve as the GIV instrument.
\paragraph{Demand equation.} The equation of interest is
where $y_{St}$ is the year-on-year growth rate of aggregate demand, $p_t$ is the year-on-year growth rate of the real commodity price, $X_t$ is a vector of observable common controls, and $\varepsilon_t$ is an aggregate demand shock. The parameter of interest is $\phi_d$, the price elasticity of demand.
OLS estimation of (ref) is biased because $p_t$ and $\varepsilon_t$ are correlated. Aggregate demand expansions raise both quantities and prices simultaneously, attenuating the estimated elasticity toward zero, or even reversing its sign when supply shocks are small relative to demand shocks.
\paragraph{Supply panel and GIV instrument.} We observe the changes to supply at individual country level. Let $y_{it}$ denote country $i$'s change in supply in period $t$. The change in supply follows the structural equation:
where $F_t$ is an $r$-vector of common factors (spanning price and aggregate demand shocks) and $u_{it}$ is an idiosyncratic supply shock. The feasible GIV instrument is the share-weighted idiosyncratic shock,
where $S_i$ is the long-term market share of country $i$. In our dataset, we construct $S_i$ by calculating the share for every $t$ and averaging across all time periods. I construct all growth rates as year-on-year midpoint growth.
\paragraph{Factor selection.} The number of factors $r$ is selected from the data using the eigenvalue ratio (ER) criterion of Ahn2013EigenvalueFactors. The Bai2002DeterminingModels information criteria are not used because their $O(\ln N / N)$ penalty is calibrated for large $N$ and proves too weak to discriminate at the cross-section sizes ($N = 21$--$29$) encountered here.
See Appendix (ref) for more comments on the data construction.
We use monthly country-level data from Bloomberg covering January 2009 to December 2025 ($T = 204$ months). The supply panel for the GIV instrument consists of refined copper supply across $N = 29$ countries. Each panel includes a rest-of-world residual. The price series is the LME spot copper price (monthly average), deflated by U.S.\ CPI rebased to 2015 $= 100$.
Post transformation, we have $T = 192$ estimation periods. The covariate matrix $X_t$ includes an intercept, one lag of aggregate refined demand growth, and the trade-weighted U.S.\ dollar index.
We use monthly EIA International Energy Statistics data on crude oil production from January 1973 to November 2025 ($T = 635$ months). The USSR and its successor states are treated as a single continuous unit (Former USSR series pre-1992; sum of 15 successor states post-1991). After dropping countries with any zero or missing observation, the panel has $N = 21$ units (20 countries plus rest of world). The year 2020 is excluded from estimation to avoid contamination from the COVID-19 production collapse.
The primary price series is a splice of the FRED OILPRICE series (January 1946--August 2024) and WTI (September 2024--December 2025), deflated by U.S.\ CPI rebased to 2015 $= 100$. Post transformation, we have $T = 611$ estimation periods. Aggregate demand is constructed by aggregate production growth adjusted for inventory changes. The covariate matrix $X_t$ includes an intercept and two lags of aggregate production growth.
Monthly country-level natural gas production data are obtained from the JODI Gas Database, covering January 2010 to November 2025 ($T = 191$ months). The panel contains $N = 27$ countries. The year 2020 is excluded from estimation to avoid contamination from the COVID-19 demand collapse.
The price series is the Henry Hub Natural Gas Spot Price (dollars per million British thermal units), deflated by U.S.\ CPI rebased to 2015 $= 100$. Post transformation, we have $T = 156$ periods. The covariate matrix $X_t$ includes an intercept, eleven lags of aggregate production growth, and the growth rate of the real WTI crude oil price. The eleven lags are motivated by strong seasonality in natural gas markets, where winter heating demand drives pronounced annual cycles in both quantities and prices. The real oil price is included because natural gas and oil are partial substitutes in power generation and industrial use, making oil price variation a relevant demand shifter.
We assume that the cross-sectional size distribution follows a power law, $\Pr(S \geq s) \propto s^{-\mu}$. Table (ref) reports Pareto exponent $\mu$ estimates for the supply panels using three methods: Hill (1975) MLE, log-rank OLS, and the bias-corrected Gabaix--Ibragimov regression Gabaix2011RankExponents.
\paragraph{Regime classification.} The estimates partition the three commodities cleanly. Copper has $\hat\mu_{\mathrm{GI}} = 0.58$ with confidence interval $[0.28, 0.87]$, and natural gas has $\hat\mu_{\mathrm{GI}} = 0.32$ with $[0.15, 0.49]$. Both lie firmly in the strong regime $\mu < 1$ of Section (ref). Crude oil has $\hat\mu_{\mathrm{GI}} = 0.65$ but a wider interval, $[0.26, 1.05]$, that reaches into the nearly weak region $\mu \in (1,2)$. Theorems (ref) and (ref) establish consistency and asymptotic normality for the entire range $\mu \in (0,2)$, so the point estimates and Wald confidence intervals reported below remain valid for crude oil under either reading of $\mu$. We can comfortably rule out $\mu > 2$ for all three commodities.
All point estimates are below 1 for all three commodities, consistent with the original GIV validity condition ($\mu < 1$). However, the Gabaix--Ibragimov confidence interval for crude oil contains 1. This is precisely the situation my extended theory is designed to cover. Even if the concentration is not extreme enough to satisfy $\mu < 1$ with certainty, we can safely use the point estimates and conduct inference for $\mu \in (0,\,2)$.
Figure (ref) plots log-rank against log-share for all three panels. The near-linear relationship over the full support confirms that a Pareto distribution is a reasonable description of the size distribution in each market, with the estimated Gabaix--Ibragimov slope $\hat\mu$ shown as the fitted line.
Table (ref) presents the main elasticity estimates.
The OLS estimates starkly illustrate the endogeneity problem. For copper, OLS yields $-0.051$, roughly one-third of the GIV estimate in magnitude. For crude oil and natural gas, OLS is positive. The raw price-quantity covariance is driven by demand shocks that raise both price and the aggregate, producing a positive spurious correlation. GIV corrects all three estimates to economically plausible negative demand elasticities: $-0.135$ for copper, $-0.109$ for crude oil, and $-0.056$ for natural gas.
For copper and crude oil, the difference between the infeasible and feasible GIV standard errors is small (at most two basis points), confirming that the PCA estimation step adds little additional uncertainty in practice. For natural gas, the feasible standard error is surprisingly smaller than the infeasible one. This occurs because the first-stage correction term that projects $\varepsilon_t$ onto the estimated factors, namely $\frac{1}{T}\sum_t \hat{F}_t'\hat{\varepsilon}_t \cdot \hat\Sigma_{\tilde F}^{-1} \hat{F}_t$, can be negatively correlated with the uncorrected IV residuals, reducing the total variance of the corrected moment conditions. This further illustrates the importance of accounting for the estimation error in the construction of the instrument.
All three commodities display inelastic demand. The estimated elasticities ($|\hat\phi| < 0.15$) are consistent with the industrial nature of these commodities. All have limited short-run substitutes, making quantity responses to price changes modest. The estimates are also broadly in line with the existing literature, which typically reports short-run elasticities in the range $[-0.05,\,-0.25]$ for crude oil Baumeister2019StructuralShocks, $[-0.05,\,-0.20]$ for natural gas Auffhammer2018NaturalBills,Labandeira2017ADemand, and $[-0.07,\,-0.10]$ for copper Shojaeddini2025UnderstandingMinerals,Lanz2013SubglobalIndustry.
I assess the finite sample performance of the GIV estimator under different regimes of instrument strength via Monte Carlo simulation. The factor structure is calibrated to two empirical commodity panels: Copper and Crude Oil, while all other components of the data generating process are simulated. Running the experiment on both commodities allows us to examine how the estimator performs under different cross-section and time-series dimensions: Copper provides $N = 29$ countries over $T = 192$ months, while Crude Oil provides $N = 21$ countries over $T = 611$ months.
The structural equation of interest is the aggregate demand for a commodity,
The supply panel follows the factor model
where $F_t^1$ is a common factor, $\tilde{\Lambda}$ is the $N \times r - 1$ matrix of demeaned factor loadings, $\tilde{F}_t$ is the $r-1 \times 1$ vector of factors, and $u_t$ is the $N \times 1$ vector of idiosyncratic shocks.
The demeaned factor loadings $\tilde{\Lambda}$ and the factors $\tilde{F}_t$ are extracted from each empirical estimation and held fixed across all Monte Carlo replications.
All remaining quantities are drawn independently each replication. The common factor and idiosyncratic shocks are drawn from normal distributions and scaled to be comparable to the empirical factor:
where $\sigma_F = \text{std}(\hat{F})$ matches the standard deviation of the empirical factor series, and $\sigma_u = \text{std}(\hat{u}_t)$ matches the pooled standard deviation of the empirical PCA estimation. The structural error $\varepsilon_t$ is drawn with $\text{std}(\varepsilon_t) = \sigma_F$ and $\text{corr}(\varepsilon_t, \tilde{F}_t) = 0.8$.
For the Copper DGP, the refined copper supply panel consists of year-on-year midpoint growth rates for $N = 29$ countries over $T = 192$ months. The ER criterion yields $r-1 = 1$ factor. The calibration scales are $\sigma_F = 0.109$ and $\sigma_u = 0.171$, and the true elasticity is set to $\phi_d = -0.135$ (the empirical GIV estimate).
For the Crude Oil DGP, the crude oil production panel consists of year-on-year midpoint growth rates for $N = 21$ countries over $T = 611$ months (COVID year 2020 dropped). The ER criterion yields $r-1 = 2$ factors. The calibration scales are $\sigma_F = 0.090$ and $\sigma_u = 0.142$, and the true elasticity is set to $\phi_d = -0.110$ (the empirical GIV estimate of $-0.109$, rounded to two decimals).
Individual sizes are drawn from the power distribution $\mathbb{P}(s_i > s) = c \, s^{-\mu}$ via the inverse CDF transform $s_i = (1 - U_i)^{-1/\mu}$ with $U_i \sim \mathcal{U}[0,1]$, and normalized so that $\sum_{i=1}^N S_i = 1$. Recall that the Pareto exponent $\mu$ determines the market concentration, and consequently the strength of the instrument. We consider $\mu \in \{0.3, 0.5, 0.8, 1.2, 1.4, 1.8, 2.5, 3.5, 6.0\}$, with $B = 5{,}000$ replications per value.
From this DGP, we construct the observational data using $y_t = \iota_N F_t^1 + \tilde{\Lambda} \tilde{F}_t + u_t$, $y_{St} = S' y_t$, and $p_t = \frac{y_{St} - \varepsilon_t}{\phi_d}$. The price is endogenous. Specifically, $\text{Cov}(p_t, \varepsilon_t) = -\sigma_\varepsilon^2 / \phi_d > 0$ (since $\phi_d < 0$), inducing an upward OLS bias.
The Crude Oil panel has only $N = 21$ countries, which limits the scope for studying the estimator's behavior as $N$ grows. To investigate the effect of a larger cross-section while preserving the empirical factor structure, we augment $N$ by resampling rows of $\tilde{\Lambda}$ with replacement. Specifically, for a target $N_{\text{sim}} > N$, we draw $N_{\text{sim}}$ rows from the $21 \times r$ empirical loading matrix uniformly with replacement. The factors $\tilde{F}_t$, calibration scales $\sigma_F$ and $\sigma_u$, and the true elasticity $\phi_d$ remain unchanged. Each replication still draws fresh $u_{it}$, $F_t^1$, $\varepsilon_t$, and $S$, so the augmented countries are distinguished by their idiosyncratic shocks and shares even when they share a loading vector. I report results for $N \in \{21, 50, 100\}$.
Tables (ref) and (ref) organise the simulation evidence by the Pareto exponent $\mu$, which Proposition (ref) maps directly onto the regimes of instrument strength. The CI Length column reports the standard $t$-based 95% interval with White heteroskedasticity-robust standard errors. The Anderson--Rubin confidence sets recommended in Section (ref) and its feasible counterpart are not used here. We let the $t$-based interval reveal where it breaks down.
\paragraph{Strong regime ($\mu < 1$).} The strong-instrument theory predicts standard inference, and that is what we observe. For Copper, the median $F$-statistic is large throughout, falling from $124$ at $\mu = 0.3$ to $78$ and $34$ at $\mu = 0.5$ and $0.8$ as shares become less concentrated. The median $\hat\phi_d$ matches the true value to three decimals and bias is at most $0.0008$. RMSE is below $0.01$ at $\mu \in \{0.3, 0.5\}$; at $\mu = 0.8$ it rises to $0.074$ as a few extreme draws enter the second moment, even though the median and bias remain on target. Crude Oil is sharper, with median $F$ between $107$ and $434$ and RMSE below $0.01$ throughout. Coverage of the $t$-based interval is at the nominal $0.95$ in both panels.
\paragraph{Nearly weak regime ($\mu \in (1,2)$).} The estimator continues to recover the true elasticity. The median estimate stays at the true value to three decimals and coverage of the $t$-based interval is at nominal levels in both calibrations. The point-estimate moments begin to reflect occasional extreme draws---Copper RMSE rises from $0.15$ at $\mu = 1.2$ to $0.71$ at $\mu = 1.8$---but the centre of the distribution remains on target. The first-stage $F$ is the quantity that signals weakness, and it is sensitive to $T$. Copper at $\mu = 1.4$ has $F = 8.7$, below the conventional Stock--Yogo cutoff of 10, while Crude Oil at the same $\mu$ has $F = 27.5$ because $T$ is three times larger. Theorem (ref) guarantees that standard normal inference remains valid, and the simulations confirm it.
\paragraph{$\mu > 2$ and the case for Anderson--Rubin.} For $\mu > 2$ the regime is no longer pinned down by the tail index alone; it is the $N/T$ trajectory that separates nearly weak from weak identification. The two panels of Table (ref) hold $N$ fixed at $29$ and $21$ while $T$ is large ($192$ and $611$), so $N/T \to 0$ and the design sits in the nearly weak regime, where the estimator is consistent but the concentration parameter has only a slowly-divergent floor. The genuinely weak case $N/T \to c > 0$, in which the estimator is inconsistent, is approached in Table (ref) by raising $N$ at fixed $T$. Standard inference begins to break down here. The median $F$ falls to $2.1$ at $\mu = 2.5$, $1.0$ at $\mu = 3.5$, and $0.5$ at $\mu = 6.0$ in Copper, and to $6.7$, $3.1$, and $1.0$ in Crude Oil, at or below the conventional Stock--Yogo cutoff of $10$. The median estimate drifts away from the true elasticity, reaching $-0.122$ for Copper and $-0.109$ for Crude Oil at $\mu = 3.5$ and $-0.107$ and $-0.100$ at $\mu = 6.0$, against true values of $-0.135$ and $-0.110$. RMSE is dominated by extreme draws and peaks at $8.1$ for Copper at $\mu = 2.5$ and $101$ for Crude Oil at $\mu = 3.5$. The clearest symptom of the breakdown is the $t$-based confidence interval: its length grows by more than an order of magnitude relative to the strong regime, from about $0.02$ to $0.41$ in Copper and $0.009$ to $0.21$ in Crude Oil between $\mu = 0.3$ and $\mu = 6.0$. Coverage of the $t$-based interval departs from nominal, drifting to $0.967$ in Copper and $0.969$ in Crude Oil at $\mu = 6.0$. Each of these is a consequence of the slowly-divergent concentration parameter floor.
As Section (ref) shows, the Wald interval scales with $1/\kappa^2_{\text{conc}}$, which is exactly why its length explodes and its coverage goes wrong in this regime. The Anderson-Rubin statistic avoids this dependence entirely. Its pivotal $\chi^2_1$ limit holds along any admissible $(N,T)$ sequence with $\mu > 2$ and requires no lower bound on $\kappa^2_{\text{conc}}$. The simulations therefore reproduce the asymmetry the theory predicts. Wald inference suffices for $\mu < 2$ and Anderson-Rubin is the appropriate tool for $\mu > 2$.
\paragraph{Augmenting the cross-section (Table (ref)).} The augmented Crude Oil experiments separate the role of $N$ from the role of $T$ and confirm the rate-based predictions of the theory regime by regime. In the strong regime, performance is essentially invariant to $N$. RMSE at $\mu = 0.3$ moves only from $0.0035$ at $N = 21$ to $0.0029$ at $N = 50$ and $0.0025$ at $N = 100$. In the nearly weak regime with $\mu \in (1,2)$, consistency requires $N/T \to 0$. Increasing $N$ at fixed $T$ therefore degrades performance, and indeed RMSE at $\mu = 1.4$ rises from $0.036$ at $N = 21$ to $0.045$ at $N = 50$ and $0.060$ at $N = 100$. For $\mu > 2$ the rate is $\sqrt{T/N}$, so larger $N$ at fixed $T$ is doubly costly. Raising $N$ at fixed $T$ also lifts $N/T$ away from zero, moving the design out of the nearly weak regime and toward the weak regime $N/T \to c > 0$, where the estimator is inconsistent. Consistent with this, CI length grows uniformly with $N$ and the median $F$ falls---at $\mu = 6.0$ it declines from $1.0$ at $N = 21$ to $0.6$ at $N = 100$---while RMSE remains large and tail-dominated throughout, reaching $1.7$ at $N = 100$.
This paper extends Granular Instrumental Variables to large panels with $N, T \to \infty$ and shows how the granularity of the cross section governs instrument strength. Under a power-law assumption on unit sizes, the tail index $\mu$ together with the $N/T$ trajectory determines the asymptotic behavior of the estimator. Three regimes emerge. The strong regime $\mu \in (0,1)$ recovers the classical $\sqrt{T}$ rate of Gabaix2024GranularVariables. The nearly weak regime $\mu > 1$ with $N/T \to 0$, following Antoine2021GMMIdentification, delivers consistency and asymptotic normality at the slower rate $\sqrt{T}/N^{\delta}$, where $\delta = \min(1 - 1/\mu, 1/2)$. The weak regime $\mu > 2$ with $N/T \to c$ corresponds to Staiger1997InstrumentalInstruments-style local-to-zero asymptotics, and the estimator is inconsistent. For inference, Wald is reliable when $\mu < 2$. For $\mu > 2$, I recommend Anderson--Rubin confidence sets, which remain valid regardless of the $N/T$ trajectory.
In practice the GIV instrument is built from estimated rather than known idiosyncratic shocks. Under the additional growth restriction $\sqrt{T}/N \to 0$, the feasible estimator attains the same convergence rate as the infeasible one, but its asymptotic variance is different. The first-stage estimation contributes an additional term that enters at the same order as the infeasible variance, so the formulas for the standard error are not the same. Valid inference therefore requires standard errors that explicitly account for the first-stage estimation error, and I provide a HAC-consistent variance estimator that does so.
I apply the GIV estimator to estimate short-run demand elasticities for refined copper, crude oil, and natural gas. All three markets have estimated tail indices below one, with crude oil's confidence interval reaching into the nearly weak region. The estimated elasticities are $-0.135$, $-0.109$, and $-0.056$ respectively, all consistent with the inelastic short-run response that industrial commodities are known for.