Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,599 characters · 11 sections · 7 citation commands
Testing Finite Moment Conditions for the Consistency and the Root-N Asymptotic Normality of the GMM and M Estimators
Many estimators of interest in economic analysis are GMM or M estimators. They include, but are not limited to, the ordinary least squares (OLS) estimators, the generalized least squares (GLS) estimators, the quasi maximum likelihood estimators (QMLE), and the two stage least squares (2SLS) estimators. Under random sampling, the consistency of these estimators is usually established via the weak law of large numbers (WLLN), which requires a finite first moment of the score. The root-n asymptotic normality of these estimators is usually established via Lindeberg-L\'evy Central Limit Theorem (CLT), which requires a finite second moment of the score. These conditions are usually taken for granted by the authors of empirical economics papers that report point estimates and their standard errors, report confidence interval, and/or conduct hypothesis testing based on the limit normal distribution.
However, these assumptions are not necessarily plausibly satisfied in applications. There are a couple of possible scenarios in which the score may not have finite first and second moments. First, some dependent variables (e.g., infant birth weight\footnote{See \citet*[][Section 6.2]{ChernozhukovFernandezVal2011} for example.} and murder rate,\footnote{See \citet*[][Appendix A]{GandhiLuShi2017} for example.} as well as income, wealth, and stock returns) are reported to exhibit heavy tailed distributions, and their many outliers may contribute to heavy tailed distributions of residuals. Second, suppose that a dependent variable is the logarithm of a variable, as is the case with demand analysis under differentiated products markets. It is a common empirical practice to replace zeros by an infinitesimal value to avoid the logarithm of zero, but this operation may result in a heavy tailed distribution of the residuals for those observations with originally non-zero values of the dependent variable -- we will later show a case in point based on actual empirical data.
In doubt about the assumptions of the finite first and second moments of the score in certain applications, we naturally desire to have a method of testing these conditions for the consistency and the root-n asymptotic normality. This article is motivated by this objective, and we therefore propose a method of testing these assumptions. With our proposed test, researchers can assess whether the bounded moment conditions for the consistency and the root-n asymptotic normality are satisfied for past and future empirical studies. In the event where the test supports the consistency and the root-n asymptotic normality for a selected study, the test result will reinforce the credibility of the scientific conclusions reported by that study. On the other hand, in the event where the test rejects the finite moment conditions for consistency or the root-n asymptotic normality for a selected study, researchers would like to substitute alternative robust methods -- see the related literature ahead. In this way, our proposed method of testing the finite moment conditions is expected to contribute to enhancing the credibility of past and future empirical economic studies.
Our proposed method is based on extreme value theory, and works in the following simple manner. Consider a test of finite $r$-th moment of the score -- set $r=1$ (respectively, $r=2$) for a test of the consistency (respectively, the root-n asymptotic normality). First, sort the $r$-th power of the norm of the estimated score in descending order. Second, pick the largest $k$ of these order statistics and self-normalize them. We show that these self-normalized statistics asymptotically follow a known joint distribution up to the unknown tail index parameter. A sub-unit (respectively, super-unit) value of this tail index parameter indicates a finite (respectively, infinite) $r$-th moment of the score. Lastly, using these dichotomous characteristics, we construct a likelihood ratio test based on the limit joint distribution of the self-normalized statistics. We establish a nearly uniform size control property of the proposed test over the set of data generating processes compatible with the null hypothesis of a finite $r$-th moment of the score.\footnote{See Section (ref) ahead for a precise description of the near uniformity.} This near uniformity property of the test is attractive since researchers do not ex ante know or do not want to fix the true distribution of the $r$-th power of the norm of the score under the composite null hypothesis in consideration.
Simulation studies support the theoretical result of the nearly uniform size control property. Applying the proposed method of testing to the widely used market share data from Dominick's Finer Foods retail chain, we find that the common ad hoc treatment of zero market shares by adding an infinitesimal positive value results in a failure of the consistency and the root-n asymptotic normality. This failure results from the fact that inclusion of logs of these infinitesimal numbers (i.e., large negative values) induces a heavy-tailed distribution of the regression residuals for observations with originally non-zero market shares.
{\bf Relation to the Literature:} We are not aware of any existing paper that develops a test of the finite moment condition for the consistency or the root-n asymptotic normality of the GMM or M estimators, as we do in this paper. A different but related topic is a set of tools to test non- and weak-identification \citep*[e.g.,][]{Wright2003,StockYogo2005,InoueRossi2011,SandersonWindmeijer2016}. These are related to our framework on one hand because non- and weak-identification also results in a failure of the canonical consistency and the root-n asymptotic normality, and therefore these testing methods serve for related objectives. On the other hand, these are different from our framework because the non- and weak-identification concerns about non- and weak-invertibility of the expected gradient of the score,\footnote{For general matrix rank tests, see e.g., GillLewbel1992,CraggDonald1996,CraggDonald1997,RobinSmith2000,KleibergenPaap2006,CambaMendezKapetanios2009,AlSadoon2017.} whereas the issue of our concern is instead about the finiteness of moments of the score as the conditions for the WLLN and CLT. In this sense, the purposes of our method of test are different from those of the preceding methods of tests of non- and weak-identification, while they indeed play complementary roles.
Also related is the paper by \citet*{ShaoYuYu2001} that proposes a test of finite variance. On the one hand, our test of the root-n asymptotic normality is also based on the test of finite second moments, similarly to \citet*{ShaoYuYu2001}. On the other hand, our objective of testing the asymptotic normality for the GMM and M estimators requires to take into account that the score is not directly observed in data, but has to be estimated via the GMM or M estimation. With these similarities and differences, our proposed method also contributes to this existing literature on testing finite moments by allowing for generated data.
For scalar locations and single equation models, an infinite first or second moment of the score is often imputed to outliers. Recognizing this issue, \citet*{Edgeworth1887} proposes to use the absolute loss instead of the square loss for a robust estimation of the equation parameters. This idea later extends and generalizes to other robust methods based on the check losses \citep*{KoenkerBassett1978} and the Huber loss \citep*{Huber1992}. While we propose a test of the finite second moment of a norm of the score for the root-n asymptotic normality of GMM and M estimators in general, there are existing papers that establish limit distribution theories (which are not necessarily root-n or normal) without requiring the finite second moment condition in these frameworks \citep*[e.g.,][]{DavisResnick1985,DavisResnick1986,DavisKnightLiu1992,HillProkhorov2016}. In the event where our test fails to support the finite moment conditions, a researcher can resort to one of these alternative robust methods instead of relying on the standard methods of inference based on the consistency and the root-n asymptotic normality of the GMM and M estimators.
We in particular highlight the case of demand estimation in differentiated products markets with market share data as a motivating example, where the dependent variable in linear models is often defined as the logarithm of a variable that may occasionally take the value of zero. Researchers sometimes substitute infinitesimal positive values for the zero in order to avoid the logarithm of zero, but this practice may also entail infinite first and second moments of the score -- see our empirical application in Section (ref) ahead. It is not the logarithm of these infinitesimal numbers per se that act as outliers, but they induce a heavy-tailed distribution of residuals of observations with originally non-zero market shares. In light of these unfavorable test results for the common ad hoc practice, we suggest that a researcher may in stead want to resort to alternative robust methods such as \citet*{GandhiLuShi2017} for demand analysis with zero market shares. A similar suggestion applies to gravity analysis of international trade, where the logarithm of zero is ubiquitous -- a researcher may want to resort to alternative robust methods such as \citet*{SilvaTenreyro2006} for gravity analysis with zero trade flows.
Finally, our method is based on recent developments in extreme value theory. We refer readers to \citet*{HaanFerreira2006} for a very comprehensive review of this subject. In particular, our inference approach is based on fixed-$k$ asymptotics, and takes advantage of and extends the technique developed by \citet*{MullerWang2017} and \citet*{Mueller2020}. The fixed-$k$ approach is useful in practice, because the asymptotic size control is valid for any predetermined fixed number $k$, unlike traditional increasing-$k$ approaches that require a sequence of changing tuning parameters as the sample size grows for which a sensible choice rule is difficult to obtain in small samples. The fixed-$k$ approach also allows for robustness against errors in preliminary estimation -- this type of robustness, benefiting from the fixed-tuning parameter setup, has been similarly explored in other contexts in the existing literature, e.g., the fixed-$b$ asymptotic inference under heteroskedasticity and autocorrelation proposed by KieferVogelsang2005 and the robust inference in kernel estimations proposed by CattaneoCrumpJansson2014. In constructing our likelihood ratio test, we take advantage of the computational algorithm developed by \citet*{ElliottMullerWatson2015}.
In this section, we introduce the general frameworks of the GMM and M estimators for which we propose tests of the finite moment conditions for the consistency and the root-n asymptotic normality. A concrete empirical example will follow after the presentation of the general frameworks.
{\bf M-Estimation:} Consider the class of estimators defined by
where the criterion function $Q_n$ takes the form of $ \hat Q_n(\theta) = n^{-1}\sum_{i=1}^n g_i(\theta). $ Under regularity conditions for this class, the influence function representation takes the form of
where $ \hat H_n(\theta) = n^{-1}\sum_{i=1}^n D_\theta^2 g_i(\theta) $ and $ g_i'(\theta) = \nabla_\theta g_i(\theta). $ The consistency of $\hat\theta$ (via the weak law of large numbers) requires $ E\left[\left\| g_i'(\theta_0) \right\|\right] < \infty. $ Likewise, the asymptotic normality of $\sqrt{n}\left(\hat\theta - \theta_0\right)$ (via multivariate Lindeberg-L\'evy CLT) requires $ E\left[\left\| g_i'(\theta_0) \right\|^2\right] < \infty. $ Common examples include the following two classes of estimators:
In this paper, we propose tests of the null hypothesis: $ E\left[ A_i^1(\theta_0) \right] < \infty, $ the condition that is required for establishing the consistency of $\hat\theta$; and the null hypothesis: $ E\left[ A_i^2(\theta_0) \right] < \infty, $ the condition that is required for establishing the asymptotic normality of $\sqrt{n}\left(\hat\theta - \theta_0\right)$.
${}$\\ {\bf GMM:} Next, consider the class of estimators defined by
where the criterion function $\hat{Q}_n$ takes the form of $ \hat Q_n(\theta) = \left[n^{-1} \sum_{i=1}^n g_i(\theta) \right]^\intercal \hat W \left[n^{-1} \sum_{i=1}^n g_i(\theta) \right]. $ Under regularity conditions for this class, the influence function representation takes the form of
where $ \hat G_n(\theta) = n^{-1} \sum_{i=1}^n \nabla_\theta g_i(\theta) $ and $ \hat W \stackrel{p}{\rightarrow} W_{0}. $ The consistency of $\hat\theta$ (via the weak law of large numbers) requires $ E\left[\left\| g_i(\theta_0) \right\|\right] < \infty. $ Likewise, the asymptotic normality of $\sqrt{n}\left(\hat\theta - \theta_0 \right)$ (via multivariate Lindeberg-L\'evy CLT) requires $ E\left[\left\| g_i(\theta_0) \right\|^2\right] < \infty. $ A common example is:
Similarly to the M-estimation case, we propose tests of the null hypothesis: $ E\left[ A_i^1(\theta_0) \right] < \infty, $ which is required for establishing the consistency of $\hat\theta$; and the null hypothesis: $ E\left[ A_i^2(\theta_0) \right] < \infty, $ which is required for establishing the asymptotic normality of $\sqrt{n}\left(\hat\theta - \theta_0\right)$.
In applications, $A_i^r(\theta_0)$ may not have a finite moment in the presence of outliers in the dependent variable. Outliers may be innate in data in some applications. In other applications, outliers may be produced as artifacts of ad hoc procedures taken by researchers. In the empirical application to be presented in Section (ref) ahead, we highlight this point in the context of the demand analysis under the following setup.
This section presents an overview of the procedure of our proposed test. Formal theoretical justifications for why this method works will be presented in Section (ref).
First, consider a non-negative random variable $A_i^r(\theta_{0})$, examples of which are introduced in Section (ref). Whether its moment is finite is fully determined by its right-tail behavior. Specifically, suppose that there is a single parameter $\xi \ge 0$ so that $n^{\xi}$ characterizes the order of magnitude of the sample maximum of $A_i^r(\theta_{0})$. Then, by extreme value theory, the finiteness of the moment of $A_i^r(\theta_{0})$ is equivalent to the condition that $\xi < 1$. To obtain a compact null space, we set $\varepsilon$ to be some small slackness parameter, say 0.01.\footnote{Another possible treatment is to switch the null and alternative hypotheses so that $H_{0}$ is $\xi \in [1, \bar{\xi}$]. However, deriving the asymptotic behavior of our test will be challenging under such $H_{0}$, if possible at all, since the estimator of $\theta_{0}$ could behave poorly.} Then, we write the two competing hypotheses:
where $\bar{\xi}$ is the upper bound of the parameter space that includes all empirically relevant values of $\xi$. We set $\bar{\xi}=2$ in later sections.
Since the limiting density (see ((ref)) below) and the power function are continuous in $\xi$, setting the slackness parameter $\varepsilon$ is innocuous and without loss of generality unlike the seemingly related treatments in the contexts of the (near) unit root or weak-/non-identification. In practice, we can set $\varepsilon$ arbitrarily close to zero (and $\bar{\xi}$ arbitrarily large). In this sense, the near uniformity of our proposed test means controlling size uniformly over $\xi \in \lbrack 0,1-\varepsilon ]$, which can be arbitrarily close to $\lbrack 0,1)$. In other words, it guarantees the size control over all distributions of $A_i^r(\theta_{0})$ that possess a finite $1+\delta$ moment for $\delta < \varepsilon/(1-\varepsilon)$. As recently studied by Mueller2020, the asymptotic normality may provide a very poor approximation for both the classic t-statistic and the bootstrap if the second moment is finite but the third moment is not. If one sets $r=2$, then rejecting $H_0$ indicates that the asymptotic normality of $\sqrt{n}\left(\hat\theta-\theta_0\right)$ is not reliable.
Second, in order to test the above hypothesis, we would like to observe large values of $A_i^r(\theta_{0})$, but it is of course unobserved. To overcome this issue, we make use of a consistent estimator $\hat\theta$ of $\theta_0$. Let $ A_{(1)}^r(\hat{\theta})\geq A_{(2)}^r(\hat{\theta})\geq \ldots \geq A_{(n)}^r(\hat{\theta}) $ denote the order statistics by sorting $\{A_i^r(\hat\theta)\}_{i=1}^n$ in the descending order. For a pre-determined integer $k \ge 3$, collect the $k$ order statistics as
By extreme value theory again, the joint distribution of the largest order statistics asymptotically approaches a well-defined parametric joint distribution that is fully characterized by the location, the scale, and the scalar parameter $\xi$. Therefore, if we conduct location and scale normalization by considering the statistics
then $\mathbf{A}_{\ast}^r(\hat{\theta})$ converges in distribution to a limiting random vector $\mathbf{V}_{\ast}$, whose density $f_{\mathbf{V}_{\ast}}$ is fully characterized by $\xi$ and is invariant to location, scale, and the order $r$ -- see ((ref)) ahead for the formula of $f_{\mathbf{V}_{\ast}}$. The estimation error in $\hat\theta$ is asymptotically negligible since it is of a smaller order of magnitude than the largest order statistics of $A_i^r(\theta_{0})$ under the null hypothesis.
Now, the limiting testing problem has become straightforward: we construct a test based on a random draw of $\mathbf{V}_{\ast}$ from its parametric density $f_{\mathbf{V}_{\ast}}$ about the only unknown scalar parameter $\xi$. When the null and alternative hypotheses are both simple, the optimal solution is known to be the Neyman-Pearson test, where large values of the likelihood ratio statistic reject the null hypothesis. Therefore, we transform the null and alternative hypotheses of ((ref)) into simple ones by considering weighted average likelihoods, and our proposed test rejects the null hypothesis that $A_i^r(\theta_0)$ has a finite moment if
where $W\left( \cdot \right)$ denotes a weight chosen to reflect the importance of rejecting different alternatives, and $\Lambda \left( \cdot \right) $ is some pre-determined weight defined on the null space. The critical value is subsumed in $\Lambda \left( \cdot \right) $ so that the likelihood ratio on the left-hand side is compared with one on the right-hand side. More details about this test are presented in the following section.
We close the preview with some heuristic discussions of the asymptotic property of the new test - formal discussions will follow in Section (ref). Figure (ref) plots oracle rejection probabilities of the test with $\mathbf{V}_{\ast }$ generated from the limiting distribution $f_{\mathbf{V}_{\ast }}$ with various values of $\xi$ and the nominal size of $0.05$. The plots are based on simulations with 10000 iterations. (This is a power prescription and is different from Monte Carlo simulation studies. Full-blown Monte Carlo simulation studies with concrete econometric models and small samples will be conducted and presented in Section (ref) to evaluate the finite sample performance.) Observe that the rejection probabilities for $\xi \in [0,1-\varepsilon]$ are uniformly dominated by the nominal size, $0.05$. In other words, the test has a size control property for nearly all distributions with tail index less than one, where the slackness $\varepsilon$ can be made arbitrarily small. This uniformity property is useful in practice because researchers do not ex ante know or do not want to fix the true value of $\xi$ in econometric models of their interest in the composite null hypothesis of consideration.
The following section formally presents the asymptotic theory for the test.
We now present a formal theory to guarantee that our proposed test works in large samples. For the convenience of exposition, we first introduce additional notations and some definitions.
Denote by $D_{i}$ the $i$-th observation so that we can write $A_{i}^r\left( \theta \right) = A^r\left( \theta ;D_{i}\right)$. For example, $D_i = (X_i^\intercal,Y_i)^\intercal$ in the context of the OLS, and $D_i = (X_i^\intercal,Z_i^\intercal,Y_i)^\intercal$ in the context of the 2SLS presented in Section (ref). Let $F_{A^r\left( \theta \right) }$ denote the cumulative distribution function (CDF) of $A_{i}^r\left( \theta \right)$ and $\theta _{0}$ denote the (pseudo-) true value of $\theta $. Let $B_{\eta _{n}}\left( \theta _{0}\right) $ denote an open ball centered at $\theta _{0}$ with radius $\eta _{n}\rightarrow 0$. Let $f_{A^r\left( \theta _{0}\right) }$ and $Q_{A^r\left( \theta _{0}\right) }$ be the probability density function (PDF) and quantile function of $A_{i}^r\left( \theta _{0}\right) $, respectively.
We say that a distribution $F$ is within the domain of attraction of the extreme value distribution, denoted by $F \in \mathcal{D}\left( G_{\xi}\right)$, if there exist sequences of constants $a_n$ and $b_n$ such that for every $v$, $$ \lim_{n \rightarrow \infty} F^{n}(a_n v + b_n) = G_\xi(v) $$ holds, where $$ G_\xi(v) =
$$ is referred to as the generalized extreme value distribution. This condition, characterizing the tail shape of the underlying distribution, is mild and satisfied by many commonly used distributions. In particular, the case with $\xi>0$ covers distributions with a regularly varying tail, including, for example, Pareto, Student-t, and F distributions. The case with $\xi=0$ covers thin tailed distributions, such as the Gaussian family. The case with $\xi<0$ covers distributions with bounded right end-point, such as the uniform distribution. See \citet*[][Ch.1]{HaanFerreira2006} for a comprehensive review. With these notations and definitions, we impose the following regularity conditions to prove the uniform size control property of our proposed test.
Condition (ref) (i) requires random sampling from some fixed population distribution. It rules out the trivial case where the location and the scale of the data diverge with the sample size. As a consequence, the existence of the moment only depends on the tail heaviness of the underlying distribution, which should remain invariant to location and scale shifts.
Condition (ref) (ii) requires that the distribution of $A^r(\theta_0)$ falls in the domain of attraction of the extreme value distribution, and that it has an unbounded support.\footnote{The finiteness of $\mathbb{E}[A_i^r(\theta_0)]$ is trivially satisfied if $\xi$ is negative.} This domain of attraction assumption bridges the finiteness of moments and the tail heaviness. In particular, the finiteness of the first moment of $A_i^r(\theta_0)$ is equivalent to the condition that $\xi$ is less than than $1$ \citep*[cf.][Ch.5.3.1]{HaanFerreira2006}.\footnote{Besides, if $A_i^{1}(\theta_0)$ satisfies Condition (ref) (ii) with the tail index $\xi>0$, then $A_i^{2}(\theta_0)$ satisfies it with the tail index $2\xi$. See, for example, \citet*[][Proposition B.1.9]{HaanFerreira2006}.} Therefore, our hypothesis testing problem can be written as in ((ref)).
The first part of Condition (ref) (iii) requires that the estimator $\hat\theta$ is consistent for $\theta_0$. Note that this in general only requires the identification and finite first moments of the score. For the case of testing the finite first moment condition for consistency (i.e., the case of setting $r=1$), this assumption is satisfied under the null hypothesis in ((ref)). For the case of testing the finite second moment condition for root-n asymptotic normality (i.e., the case of setting $r=2$), this assumption is satisfied even under a range of alternative hypotheses as well as under the null hypothesis. In any case, the fact that this consistency condition is satisfied under the null hypothesis allows us to establish a size control property based on this condition.
The second part of Condition (ref) (iii) requires that the gradient of $A_i^r(\theta)$ grows not too fast as the sample size increases. Since this last piece of the condition is a high-level statement, it will be useful to consider stronger lower level sufficient conditions in a specific example. Since we focus on the case of the GMM in our application, we look at a lower level condition in the context of the GMM.
${}$\\ {\bf Discussion of Condition (ref) (iii) -- Case of GMM:} Consider the case of setting $r=2$ for testing the root-n asymptotic normality. Recall that we define $ A_{i}^2\left( \theta \right) =\left( Y_{i}-X_{i}^{\intercal }\theta\right) ^{\intercal }Z_{i}^{\intercal }Z_{i}\left( Y_{i}-X_{i}^{\intercal}\theta \right).$ Thus,
The triangle inequality and Cauchy-Schwartz inequality yield
Condition (ref) (ii) implies that $ \sup_{i}\left\vert \left\vert A_{i}^2\left(\theta _{0}\right) \right\vert \right\vert \rightarrow \infty $ so that $ \left( \sup_{i}\left\vert \left\vert A_{i}^2\left( \theta _{0}\right) \right\vert \right\vert \right) ^{1/2} $ is of a smaller order than $ \sup_{i}\left\vert \left\vert A_{i}^2\left( \theta _{0}\right) \right\vert \right\vert. $ Therefore, a sufficient condition for Condition (ref) (iii) is that $ \sup_{i}\left\vert \left\vert X_{i}Z_{i}^{\intercal }\right\vert \right\vert $ is of a smaller order than $ \left( \sup_{i}\left\vert \left\vert A_{i}^2\left( \theta _{0}\right) \right\vert \right\vert \right) ^{1/2}. $ An even stronger sufficient condition for this sufficient condition is that $X_{i}$ and $Z_{i}$ have bounded supports, which is satisfied by the typical applications in the demand analysis, in particular the one that we consider in our empirical application in Section (ref). $\square$ \\$$
As discussed earlier, the domain of attraction assumption as in Condition (ref) (ii) allows for the hypothesis testing problem to be written as in ((ref)).
Recall the following notations from Section (ref). Let $ A_{(1)}^r(\hat{\theta})\geq A_{(2)}^r(\hat{\theta})\geq \ldots \geq A_{(n)}^r(\hat{\theta}) $ denote the order statistics by sorting $\{A_i^r(\hat\theta)\}_{i=1}^n$ in the descending order, and
The following lemma shows that these order statistics asymptotically follow the joint extreme value distribution.
A proof is provided in Appendix (ref).
Since $a_{n}$ and $b_{n}$ are unknown, we would like to eliminate them in constructing feasible test statistics. We do so by constructing the self-normalized statistic
By the continuous mapping theorem, change of variables, and Lemma (ref), we obtain
where the density function $f_{\mathbf{V}_{\ast }}$ of the limit observation $\mathbf{V}_\ast$ is given by
and $v_{\ast i}$ denotes the $i$-th component of $\mathbf{v}_{\ast }$. With this density function, we construct the likelihood ratio test
where $W\left( \cdot \right)$ denotes a weight chosen to reflect the importance of rejecting different alternatives\footnote{We set $W$ to be the uniform distribution on $[0,1-\varepsilon]$ with $\varepsilon=0.01$ in later sections.} and $\Lambda \left( \cdot \right) $ is some pre-determined weight that subsumes the critical value and transforms the composite null space into a simple one. Besides, such $\Lambda$ is referred to as the least favorable distribution \citep*[e.g.,][Ch.3.8]{LehmannRomano2005} that guarantees the uniform size control. \citet*{ElliottMullerWatson2015} develop a generic algorithm to numerically construct $\Lambda$, which we adapt to our setting -- see Appendix (ref) for details.
The following theorem establishes the asymptotic uniform size control of our test ((ref)), which is the main theoretical result of this article. Let $\mathbb{E}_{\xi }\left[\; \cdot \; \right]$ denote the expectation with respect to the density ((ref)) with the parameter value $\xi$.
A proof is provided in Appendix (ref).
A few remarks are in order about this new test. First, the fixed-$k$ asymptotic design leads to the desired uniform size control property as stated in the theorem. This feature provides a practical advantage because a researcher does not ex ante know or does not want to fix the true distribution under the composite null hypothesis. Second, the fixed-$k$ design is useful not only for the uniform size control but also for the robustness of the test against estimation error in $\hat\theta$. This is because the fixed-$k$ design allows for the estimation error, $\hat{\theta}-\theta_0$, to be $o_p(1)$ and hence to be dominated by the infeasible largest order statistics of $\{ A_i^r({\theta_0})\}$. This type of robustness, benefiting from the fixed-tuning parameter setup, has been similarly explored in other contexts in the existing literature. See, for example, the fixed-$b$ asymptotic inference under heteroskedasticity and autocorrelation proposed by KieferVogelsang2005 and the robust inference in kernel estimations proposed by CattaneoCrumpJansson2014.
While we focus on the fixed-$k$ asymptotic design for these practical and theoretical advantages, it is also possible to analyze an increasing-$k$ asymptotic design. Specifically, if we consider the asymptotic setting where $k \rightarrow \infty$ and $k/n \rightarrow 0$, then $\xi$ can be consistently estimated -- many such estimators exist in the statistics literature HaanFerreira2006. These estimators are usually asymptotically normal with the root-$k$ convergence rate. One can thereby build a confidence interval for $\xi$ to conduct a test of $H_0$ for $r=2$ that is consistent against fixed alternatives. Furthermore, we can always numerically examine the asymptotic power of the test as in Figure (ref) presented in Section (ref), which shows that the test has good power as long as $k$ is not too small. With this device, a practitioner can implement preliminary power analysis in the choice of $k$ before actually conducting the test, and this feature is practically more useful in finite samples.
Using Monte Carlo simulations, we demonstrate that the proposed test has the claimed uniform size control property. We consider two of the most popular econometric models, namely the linear regression model and the linear IV model, for data generating designs. For each of these two designs, we consider the test of the consistency (by setting $r=1$) and the test of the root-n asymptotic normality (by setting $r=2$).
First, consider the linear regression model:
where $\theta = (\theta_1,\theta_2)^\intercal = (1,1)^\intercal$. The independent variable is generated according to $X_i \sim N(0,1)$. The error $U_i$ is generated according to the zero symmetric Pareto distribution with tail index $\xi_U$, independently from $X_i$. We vary the value of $\xi_U$ across sets of simulations. The test is based on the $r$-th moment of the score: $ A_i^r(\theta) = (1 + X_i^2)^{r/2} \cdot |Y_i - \theta_1 - \theta_2 X_i |^r $ for $r=1,2$. Since we do not know $\theta_0$, we replace $\theta_0$ by the OLS $\hat\theta$. We thus use $k$ order statistics of
to construct our test, following the procedure outlined in Section (ref). We set $r=1$ for the test of the consistency of $\hat\theta$, and $r=2$ for the test of the asymptotic normality of $\sqrt{n}\left(\hat\theta-\theta_0\right)$.
Second, consider the linear IV model:
where $\theta = (\theta_1,\theta_2)^\intercal = (1,1)^\intercal$ and $\pi = (\pi_1,\pi_2)^\intercal = (1,1)^\intercal$. The instrument is generated according to $Z_i \sim N(0,1)$ independently of the tri-variate error components $(U_i,V_i,R_i)^\intercal$. The heavy tailed part of the error $U_i$ is generated according to the zero symmetric Pareto distribution with tail index $\xi_U$, independently from $(V_i,R_i)^\intercal$. We vary the value of $\xi_U$ across sets of simulations. The endogenous part of the error components $(V_i,R_i)^\intercal$ is generated according to $(V_i,R_i)^\intercal \sim N(\vec{0},\Sigma)$ where $\Sigma = (1,0.5;0.5,1)$. The test is based on the $r$-th moment of the score: $ A_i^r(\theta) = (1 + Z_i^2)^{r/2} \cdot |Y_i - \theta_1 - \theta_2 X_i |^r $ for $r=1,2$. Since we do not know $\theta_0$, we replace $\theta_0$ by the IV estimator $\hat\theta$. We thus use $k$ order statistics of
to construct our test, following the procedure outlined in Section (ref). We set $r=1$ for the test of the consistency of $\hat\theta$, and $r=2$ for the test of the asymptotic normality of $\sqrt{n}\left(\hat\theta-\theta_0\right)$.
For each of the linear regression model and the linear IV model introduced above, we experiment with sample sizes of $n= 10^4$, $10^5$ and $10^6$, which are similar to the sample size that we actually encounter in our empirical application in Section (ref). For testing the finite first moment condition for the consistency of $\hat\theta$, we experiment with the tail index values of $\xi_U= 0.19$, $0.39$, $0.59$, $0.79$, $0.99$, $1.19$, $1.39$, $1.59$, $1.79$ and $1.99$. Note that $\xi_U \in \{0.19, 0.39, 0.59, 0.79, 0.99\}$ satisfy the condition for the consistency, but $\xi_U \in \{1.19, 1.39, 1.59, 1.79, 1.99\}$ fail to satisfy it. For testing the finite second moment condition for the asymptotic normality of $\sqrt{n}\left(\hat\theta - \theta_0\right)$, we experiment with the tail index values of $\xi_U= 0.09$, $0.19$, $0.29$, $0.39$, $0.49$, $0.59$, $0.69$, $0.79$, $0.89$ and $0.99$. Note that $\xi_U \in \{0.09, 0.19, 0.29, 0.39, 0.49\}$ satisfy the condition for the asymptotic normality, but $\xi_U \in \{0.59, 0.69, 0.79, 0.89, 0.99\}$ fail to satisfy it. We also experiment with various numbers $k= 50$, $100$, and $200$ of order statistics for construction of the test. Each set of simulations consists of 5000 Monte Carlo iterations.
Table (ref) shows Monte Carlo simulation results of testing the finite first moment condition for the consistency of $\hat\theta$ in (A) the linear regression model and (B) the linear IV model. In both of the two panels, (A) and (B), we can see that the simulated rejection probabilities are dominated by the nominal size 0.05 for all of $\xi_U \in \{0.19, 0.39, 0.59, 0.79\}$ in the null region, and those are approximately the same as the nominal size 0.05 near the boundary, i.e., $\xi_U = 0.99$, of the null region. These results support the uniform size control property of the test that is established in Theorem (ref) as the main result of this paper. The condition for the consistency holds for any of $\xi_U < 1$, but a researcher does not ex ante know or does not want to fix which exact value $\xi_U$ takes for a specific application under the composite null hypothesis in consideration. For this reason, this nearly uniform size control property is important in practice.
Table (ref) shows Monte Carlo simulation results of testing the finite second moment condition for the asymptotic normality of $\sqrt{n}\left(\hat\theta - \theta_0\right)$ in (A) the linear regression model and (B) the linear IV model. The findings here are very similar to those in the consistency test presented above. Namely, in both of the two panels, (A) and (B), we can see that the simulated rejection probabilities are dominated by the nominal size 0.05 for all of $\xi_U \in \{0.09, 0.19, 0.29, 0.39\}$ in the null region, and those are approximately the same as the nominal size 0.05 near the boundary, i.e., $\xi_U = 0.49$, of the null region. Again, these results support the nearly uniform size control property of the test that is established in Theorem (ref).
In this section, we present an empirical application of the proposed test procedure. Recall the framework of demand estimation in differentiated products markets introduced in Example (ref). The dependent variable is defined by the logarithm of the market share of a product relative to that of an outside product. In rich data sets, we often encounter zero empirical market shares. Since the logarithm of zero is undefined, empirical practitioners often use ad hoc procedures to deal with observations with zero market share. One common way is to simply remove observations with zero empirical market shares. Another common way is to replace zeros with a small positive value. Both of these two ad hoc treatments result in biased estimates in general, as demonstrated through Monte Carlo simulation studies by \citet*{GandhiLuShi2017}. In implementing the second approach, empirical researchers often substitute infinitesimal positive values $\Delta$ for zeros, perhaps in efforts to mitigate such biases. In this paper, we show that substitution of infinitesimal positive values $\Delta$ in fact results in pathetic asymptotic behaviors of the estimator. Specifically, such an ad hoc estimator fails the root-n asymptotic normality, as we reject the finite second moment condition of the score. Furthermore, such an estimator is not even likely to converge in probability to a possibly biased pseudo-true target either, as we reject the finite first moment condition of the score too. These results follow because the introduction of the a huge negative number (as the logarithm of an infinitesimal number) turns some of the observations with originally non-zero shares into outliers, as we will carefully illustrate ahead after presenting the test results.
Following preceding papers on market analysis, we use scanner data from the Dominick's Finer Foods (DFF) retail chain.\footnote{We thank James M. Kilts Center, University of Chicago Booth School of Business for allowing us to use this data set. It is available at https://www.chicagobooth.edu/research/kilts/datasets/dominicks.} The unit of observation is defined by the product of UPC (universal product code), store, and week. Our analysis, as described below, follows that of \citet*{GandhiLuShi2017}. We focus on the product category of canned tuna. Empirical market shares are constructed by using quantity sales and the number of customers who visited the store in the week. Control variables include the price, UPC fixed effects, and a time trend. We instrument the possibly endogenous prices by the wholesale costs, which are calculated by inverting the gross margin.
The number of observations is approximately $10^6$, similar to the sample sizes considered in our Monte Carlo simulation studies in Section (ref). This feature of the data allows us to use a reasonably large number $k$ of order statistics to enhance the power of our proposed test. Among this large number of observations, approximately 44% of the observations are recorded to have zero empirical market share. The smallest non-zero empirical market share is approximately $10^{-5}$. Therefore, it is sensible to replace the zero empirical market share by an infinitesimal positive number $\Delta$ that is no larger than $10^{-5}$. In our analysis, therefore, we consider the following numbers to replace zero: $\Delta=$ $10^{-5}$, $10^{-6}$, ..., $10^{-19}$, $10^{-20}$.
Table (ref) summarizes the p-values of testing the finite first moment condition for the consistency. Similarly, Table (ref) summarizes the p-values of testing the finite second moment condition for the root-n asymptotic normality. For the sake of transparency, we show results for various numbers of $k$ ranging from 1000 to 5000. Before discussing these results, first note that small numbers $k$ of order statistics in general entail short power. In view of Figure (ref), we can see that $k$ ranging from 1000 to 5000 yields very strong powers of the test. Furthermore, note also that the number $k=1000$ corresponds to only 0.1 percent of the whole sample, so that the extreme value approximation should perform well. With these in mind, observe that the results reported in Tables (ref) and (ref) suggest that we start to reject the null hypothesis of a finite first and second moments when $k$ is larger than 1000. The rejection of the finite second moment conditions (Table (ref)) implies that the root-n asymptotic normality of the demand estimator may perform poorly if we conduct the ad hoc practice of replacing the zero empirical market share by any of the infinitesimal positive values $\Delta=$ $10^{-5}$, $10^{-6}$, ..., $10^{-19}$, $10^{-20}$. Furthermore, the rejection of the finite first moment conditions (Table (ref)) implies that such an ad hoc estimator may not even converge in probability to a possibly biased pseudo-true target.
While the test rejects the null hypotheses of finite moments of $A_i^1(\theta_0)$ and $A_i^2(\theta_0)$, a natural question is why the ad hoc procedure of adding a small constant to the zero market share causes the heavy tailed distributions of $A_i^1(\theta_0)$ and $A_i^2(\theta_0)$. Since the logarithm of a small constant is finite anyway, it appears to only produce a 44% point mass of absolutely very large yet finite constants. As such, these small numbers do not seem to contribute to heavy tails by themselves. To see what is going on behind our test rejecting the null hypotheses, we display eight scatter plots in Figures (ref) and (ref). Figure (ref) displays plots of (A) $\log$(share) on $A^1(\hat\theta)$ for $\Delta=10^{-5}$; (B) $\log$(share) on $A^2(\hat\theta)$ for $\Delta=10^{-5}$; (C) $\log$(share) on $A^1(\hat\theta)$ for $\Delta=10^{-10}$; and (D) $\log$(share) on $A^2(\hat\theta)$ for $\Delta=10^{-10}$. Figure (ref) displays plots of (A) $\log$(share) on $A^1(\hat\theta)$ for $\Delta=10^{-15}$; (B) $\log$(share) on $A^2(\hat\theta)$ for $\Delta=10^{-15}$; (C) $\log$(share) on $A^1(\hat\theta)$ for $\Delta=10^{-20}$; and (D) $\log$(share) on $A^2(\hat\theta)$ for $\Delta=10^{-20}$. Those observations above the top 0.0001-quantile of $A^1(\hat\theta)$ and $A^2(\hat\theta)$ are marked by black crosses, while all else are marked by gray dots.
In each of the panels in Figures (ref) and (ref), note that the observations with originally zero market share appear on the horizontal line at the vertical level of $\log(\Delta)$. As $\Delta$ becomes smaller, these lines move downward and they tend to behave as observations with an absolutely large $Y$ value. However, these observations with originally zero shares are not necessarily outliers by themselves because as many as 44% of the observations exist on this line. Instead, many of the outliers (i.e., observations marked by the black crosses) stem from the group of observations with originally non-zero market shares. Furthermore, the horizontal distances between those marked by the black crosses and the major cluster of observations marked by gray dots widen as $\Delta$ becomes smaller, i.e., the horizontal spread is the smallest in Figure (ref) (A)--(B) and the largest in Figure (ref) (C)--(D). This pattern implies that, while most of the 44% of observations with originally zero market share are not outliers by themselves despite the isolated levels of $\log(\Delta)$, smaller values of $\Delta$ are turning some of the observations with originally non-zero market shares into outliers to larger extents.
We conclude this section by discussing the implications of our test results and practical suggestions in light of them. Rich market share data often include zero empirical market shares. Since the logarithm of zero is undefined, empirical researchers often employ the ad hoc practice of replacing the zero by an infinitesimal positive value $\Delta$. This practice has already been known to incur biased estimates \citep*[see][]{GandhiLuShi2017}, but can also result in a failure of the root-n asymptotic normality in addition. Furthermore, the ad hoc estimator is not even guaranteed to converge in probability to a possibly biased pseudo-true target either. As such, both the point estimates and their standard errors are incredible. Empirical researchers may, therefore, want to resort to alternative methods that are robust against zero market shares, such as the method proposed by \citet*{GandhiLuShi2017}.
Many empirical studies in economics rely on the GMM and M estimators including, but not limited to, the OLS, GLS, QMLE, and 2SLS. Furthermore, they usually rely on the consistency and the root-n asymptotic normality of these estimators when drawing scientific conclusions via statistical inference. Although the conditions for the consistency and the root-n asymptotic normality are usually taken for granted as such, they may not be always plausibly satisfied. In this light, this paper proposes a method of testing the hypothesis of finite first and second moments of scores, which serve as key conditions of the consistency and the root-n asymptotic normality, respectively.
There are two desired properties of our proposed test in practice. First, unlike other approaches in extreme value theory that require a sequence of tuning parameter values that change as the sample size grows, our test is valid for any predetermined fixed number $k$ of order statistics to be used to construct the test. This is a useful property in practice because it relieves researchers from worrying about a `valid' data driven choice of tuning parameters for the purpose of size control. Second, our test has a nearly uniform size control property over the set of data generating processes for which the asymptotic normality holds. This nearly uniform size control property is useful in practice, because researchers usually do not ex ante know or do not want to fix the true tail index $\xi$ in the applications of their interest under the composite null hypothesis in consideration. Monte Carlo simulation studies indeed support this theoretical property for two of the most commonly used econometric frameworks, namely the linear regression model and the linear IV model.
A failure of the consistency and the root-n asymptotic normality may be caused by the following two cases among others. First, some dependent variables (e.g., wealth, infant birth weight, murder rate) are reported to exhibit heavy tailed distributions, and they can induce infinite first and second moments of the score of an estimator. Second, when a dependent variable is the logarithm of a variable, practitioners sometimes employ an ad hoc procedure of replacing zeros by infinitesimal values. This practice can lead to heavy tailed distribution of the residuals for observations with originally non-zero market shares. In our empirical application, we highlighted the latter case. Using scanner data from the Dominick's Finer Foods (DFF) retail chain, we reject the consistency and the root-n asymptotic normality for demand estimators based on such an ad hoc practice.
Finally, we conclude this paper by remarking that the test can be used to enhance the quality and credibility of past and future empirical studies. On one hand, if our test supports the finite moment conditions for consistency and the root-n asymptotic normality for a selected empirical work, then the test result reinforces the credibility of scientific conclusions reported by that work. On the other hand, if our test fails to support the finite moment conditions for a selected empirical work, then a researcher may want to consider one of the alternative robust approaches for more credible empirical research.