Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
126,317 characters · 18 sections · 107 citation commands
Testing Firm Conduct
{6pt} {6pt} \setlist{itemsep=0.1pt,topsep=1.0pt} {0.7\baselineskip plus 0.1\baselineskip minus 0.1\baselineskip} {10pt plus 2pt minus 2pt}
Firm conduct is fundamental to industrial organization (IO), either as the object of interest (e.g., detecting collusion) or as part of a model for policy evaluation (e.g., environmental regulation). The true model of conduct is often unknown. Researchers may choose to test a set of candidate models motivated by theory.
In an ideal testing setting, the researcher compares markups implied by a model of conduct to the true markups. As true markups are rarely observed, bh14 provide a falsifiable restriction for a model of conduct which requires instruments. In the standard parametric differentiated products environment, we show that a model is falsified by the instruments if its predicted markups (markups projected on instruments) are different from the predicted markups for the true model. This intuition connects testing with unobserved true markups to the ideal setting while also highlighting the role of instruments in distinguishing conduct. In practice, a researcher must implement a test by encoding the falsifiable restriction into statistical hypotheses. In this paper, we elucidate how the form of the hypotheses and the strength of the instruments affect inference. Based on our findings, we provide methods to ensure valid testing of conduct.\looseness=-1
The IO literature uses both model selection and model assessment procedures to test conduct. These differ in the hypotheses they formulate from the falsifiable restriction. Model selection compares the relative fit of two competing models. Model assessment checks the absolute fit of a given model. We show that this distinction affects inference under misspecification of either demand or cost. Specifically, the model selection test in rv02 (RV) distinguishes the model for which predicted markups are closer to the truth. In this precise sense, RV is “robust to misspecification”. Instead, with misspecification, model assessment tests reject the true model in large samples.
As misspecification is likely, our results establish the importance of using RV to test firm conduct. However, this test can suffer from degeneracy, defined as zero asymptotic variance of the difference in lack of fit between models rv02. With degeneracy, the asymptotic null distribution of the RV test statistic is no longer standard normal. Despite the precise definition above, the economic causes and inferential effects of degeneracy remain opaque. Thus, researchers often ignore degeneracy when testing firm conduct. We show that assuming away degeneracy amounts to imposing that at least one of the models is falsified by the instruments. Degeneracy obtains if either the markups for the true model are indistinguishable from the markups of the candidate models, or the instruments are uncorrelated with markups.
To shed light on the inferential consequences of degeneracy, we define a novel weak instruments for testing asymptotic framework adapted from ss97. Under this framework, we show that the asymptotic null distribution of the RV test statistic is skewed and has a non-zero mean. As skewness declines in the number of instruments while the magnitude of the mean increases, the resulting size distortions are non-monotone in the number of instruments. With one instrument or many instruments, large size distortions are possible. With two to nine instruments, we find that size distortions above 2.5% are impossible.\looseness=-1
Our characterization of degeneracy allows us to develop a novel diagnostic for weak instruments, which aids researchers in drawing proper conclusions when testing conduct with RV. In the spirit of sy05 and olea13, our diagnostic uses an effective $F$-statistic. However, the $F$-statistic is formed from two auxiliary regressions as opposed to a single first stage. Like sy05, we show that instruments can be diagnosed as weak based on worst-case size. With one or many instruments, the proposed $F$-statistic needs to be large to conclude that the instruments are strong for size; we provide the appropriate critical values. A rule of thumb for more than nine instruments is that the $F$-statistic must exceed twice the number of instruments in excess of nine to contain the size distortions of RV below 2.5%.\looseness = -1
A distinguishing feature of our weak instruments framework is that power is a salient concern. In fact, the maximal power of the RV test against either model of conduct is strictly less than one, even in large samples. The attainable power of the test can differ across the competing models and is lowered by misspecification. Thus, diagnosing whether the instruments are strong for power is crucial. The same $F$-statistic used to detect size distortions is also informative about the maximal power of the RV test. However, the critical values to compare the $F$-statistic against are different. For power, the critical values are monotonically declining in the number of instruments, making low power the primary concern with two to nine instruments. In sum, researchers no longer need to assume away degeneracy; instead, they can interpret the results of RV through the lens of our $F$-statistic.\looseness=-1
Up to this point, the discussion presumes that researchers precommit to a single set of instruments. bh14 show that many sources of exogenous variation exist for testing conduct. Pooling all sources of variation into one set of instruments may obscure the degree of misspecification while also diluting instrument strength. Instead, we suggest that researchers accumulate evidence by running separate RV tests using each of these sources. We propose a conservative procedure whereby a researcher concludes for a set of models when all strong instruments support them.
In an empirical application, we revisit the setting of v07 and test five models of vertical conduct in the market for yogurt, including models of double marginalization and two-part tariffs. The application illustrates the empirical relevance of our results for inference on conduct with misspecification and weak instruments. Inspection of the price-cost margins implied by different models only allows us to rule out one model. To obtain sharper results, we then perform model selection with RV. Commonly used sets of instruments are weak for testing as measured by our diagnostic. When the RV test is implemented with these weak instruments, it has essentially no power. This illustrates the importance of using our diagnostic to assess instrument strength in terms of both size and power when interpreting the results of the RV test. \looseness=-1
Using our procedure to accumulate evidence from different sources of variation, we conclude for a model in which manufacturers set retail prices. All strong instruments reject the other models. This application speaks to an important debate over vertical conduct in consumer packaged goods industries. Several applied papers assume a model of two-part tariffs where manufacturers set retail prices n01,mw17. Our results support this assumption. \looseness=-1
This paper develops tools relevant to a broad literature seeking to understand firm conduct in the context of structural models of demand and supply. Focusing on articles that pursue a testing approach, collusion is a prominent application p83,s85,b87,glv92,v96,gm98,n01,s20. Other important applications include common ownership bcs20, vertical conduct v07,bd10,g13, z21, price discrimination ddf19, price versus quantity setting fl95, and non-profit behavior dmr20. Outside of IO, labor market conduct is a recent application rs21. \looseness=-1
This paper is also related to econometric work on the testing of non-nested hypotheses pw01. We build on the insights of the econometrics literature that performs inference under misspecification and highlights the importance of model selection procedures w82,v89, hi03, k03, mo12,ls20. Two recent contributions, s15 and sw17, modify the likelihood based test in v89 to correct size distortions under degeneracy. In our GMM setting, we show that power, not size, is the salient concern. Furthermore, by connecting degeneracy to instrument strength, our work is related to the econometrics literature on inference under weak instruments ass19.\looseness=-1
The paper proceeds as follows. Section (ref) describes the environment: a general model of firm conduct. Section (ref) formalizes our notion of falsifiability when true markups are unobserved. Section (ref) explores the effect of hypothesis formulation on inference, contrasting model selection and assessment approaches under misspecification. Section (ref) connects degeneracy of RV to instrument strength, characterizes the effect of weak instruments on inference, and introduces a diagnostic for weak instruments to report alongside the RV test. Section (ref) provides a procedure to accumulate evidence across different sets of instruments. Section (ref) develops our empirical application: testing models of vertical conduct in the retail market for yogurt. Section (ref) concludes. Proofs are found in Appendix (ref). \looseness=-1
We consider testing models of firm conduct using data on a set of products $\mathcal{J}$ offered by firms across a set of markets $\mathcal{T}$. For each product and market combination $(j,t)$, the researcher observes the price $\boldsymbol{p}_{jt}$, market share $\boldsymbol{s}_{jt}$, a vector of product characteristics $\boldsymbol{x}_{jt}$ that affects demand, and a vector of cost shifters $\textbf{w}_{jt}$ that affects the product's marginal cost. For any variable $\boldsymbol{y}_{jt}$, denote $\boldsymbol{y}_t$ as the vector of values in market $t$. We assume that, for all markets $t$, the demand system is $\boldsymbol{s}_{t} = \mathbcal{s}\big(\boldsymbol{p}_t,\boldsymbol{x}_t,\boldsymbol\xi_t,\theta^D_0\big)$, where $\boldsymbol{\xi}_t$ is a vector of unobserved product characteristics, and $\theta^D_0$ is the true vector of demand parameters.\looseness=-1
The equilibrium in market $t$ is characterized by a system of first order conditions arising from the firms' profit maximization problems:
where $\boldsymbol{\Delta}_{0t} = \boldsymbol\Delta_0\big(\boldsymbol{p}_t,\boldsymbol{s}_t,\theta^D_0\big)$ is the true vector of markups in market $t$ and $\boldsymbol{c}_{0t}$ is the true vector of marginal costs. Following bh14 we assume cost has the separable form $ \boldsymbol{c}_{0jt} = \bar{\boldsymbol{c}}(\boldsymbol{q}_{jt},\textbf{w}_{jt}) + \omega_{0jt},$ where $\boldsymbol{q}_{jt}$ is quantity and $\omega_{0jt}$ is an unobserved shock. To speak directly to the leading case in the applied literature, we assume marginal costs are constant in $\boldsymbol{q}_{jt}$, and maintain $E[\bar{\boldsymbol{c}}(\textbf{w}_{jt}) \omega_{0jt}]=0$. However, the results in the paper apply to the important case of non-constant marginal cost, as shown in Appendix (ref).\looseness=-1
The researcher can formulate alternative models of conduct, obtain an estimate $\hat \theta^D$ of the demand parameters, and compute estimates of markups $\hat {\boldsymbol{\Delta}}_{mt} = {\boldsymbol{\Delta}}_{m}\big(\boldsymbol p_t, \boldsymbol s_t, \hat \theta^D\big)$ under each model $m$. For clarity, we abstract away from the demand estimation step and treat $\boldsymbol{\Delta}_{mt} = {\boldsymbol{\Delta}}_{m}\big(\boldsymbol p_t, \boldsymbol s_t, \theta^D\big)$ as data, where $\theta^D =\plim \hat \theta^D $.\footnote{When demand is estimated in a preliminary step, the variance of the test statistics presented in Section (ref) needs to be adjusted. The necessary adjustments are in Appendix (ref).} We focus on the case of two candidate models, $m=1,2$, and defer a discussion of more than two models to Section (ref). To simplify notation, we replace the $jt$ index with $i$ for a generic observation. We suppress the $i$ index when referring to a vector or matrix that stacks all $n$ observations in the sample. Our framework is general, and depending on the choice of $\boldsymbol \Delta_1$ and $\boldsymbol \Delta_2$ allows us to test many models of conduct found in the literature. Canonical examples include the nature of vertical relationships, whether firms compete in prices or quantities, collusion, intra-firm internalization, common ownership and nonprofit conduct.\footnote{In important applications mw17,bcs20, markups are functions of a parameter $\kappa$ ($\boldsymbol{\Delta}_m=\boldsymbol{\Delta}(\kappa_m)$). Researchers may investigate conduct by either estimating $\kappa$ or testing. ms21 provide a comparison of testing and estimation approaches in this setting.}
Throughout the paper, we consider the possibility that the researcher may misspecify demand or cost, or specify two models of conduct (e.g., Bertrand or collusion) which do not match the truth (e.g., Cournot). In these cases, $\boldsymbol{\Delta}_0$ does not coincide with the markups implied by either candidate model. We show that misspecification along any of these dimensions has consequences for testing, contrasting it to the case where $\boldsymbol{\Delta}_0 = \boldsymbol{\Delta}_1$.
Another important consideration for testing conduct is whether markups for the true model $ \boldsymbol{\Delta}_{0}$ are observed. In an ideal testing environment, the researcher observes not only markups implied by the two candidate models, but also the true markups $\boldsymbol{\Delta}_{0}$ (or equivalently marginal costs). However, $\boldsymbol\Delta_0$ is unobserved in most empirical applications, and we focus on this case in what follows. Testing models thus requires instruments for the endogenous markups $\boldsymbol{\Delta}_1$ and $\boldsymbol{\Delta}_2$. We maintain that the researcher constructs instruments $\boldsymbol z$, such that the following exclusion restriction holds:
This assumption requires that the instruments are {exogenous for testing}, and therefore uncorrelated with the unobserved cost shifters for the true model. Assumption (ref) describes the case where a researcher uses a single set of instruments. From bh14, any exogenous variation which moves the residual marginal revenue curve for at least one firm can serve as a valid instrument. These include variation in the set of rival firms and rival products, own and rival product characteristics, rival cost, and market demographics. In Section (ref), we discuss how researchers can separately use these sources of variation to test firm conduct, without needing to precommit to any of them.
The following assumption introduces regularity conditions that are maintained throughout the paper and used to derive the properties of the tests discussed in Section (ref).
Part (i) is a standard assumption for cross-sectional data. Extending parts of the analysis to allow for dependent data is straightforward and discussed in Appendix (ref). Part (ii) excludes cases where the two competing models of conduct have identical markups and cases where the instruments $\boldsymbol z$ are linearly dependent with the cost shifters $\textbf{w}$. Part (iii) is a standard regularity condition allowing us to establish asymptotic approximations as $n \rightarrow \infty$.\looseness=-1
We further maintain that marginal costs are a linear function of observable cost shifters $\textbf{w}$ and $\omega_0$, so that $\boldsymbol{c}_{0} = \textbf{w} \tau + \omega_{0}$, where $\tau$ is defined by the orthogonality condition $E[\textbf{w}_i \omega_{0i}]=0$. This restriction allows us to eliminate the cost shifters $\textbf{w}$ from the model, which is akin to the thought experiment of keeping the observable part of marginal cost constant across markets and products. For any variable $\boldsymbol y$, we therefore define the residualized variable $y = \boldsymbol y - \textbf{w}E[\textbf{w}'\textbf{w}]^{-1}E[\textbf{w}'\boldsymbol y]$ and its sample analog as $\hat y = \boldsymbol y - \textbf{w}(\textbf{w}'\textbf{w})^{-1}\textbf{w}'\boldsymbol y$.
The following section discusses the essential role of the instruments $\boldsymbol z$ in distinguishing between different models of conduct. For this discussion, a key role is played by the part of residualized markups $\Delta_m$ that are predicted by $z$:\looseness=-1
and its sample analog $\hat \Delta^z_m = \hat z \hat \Gamma_m$ where $\hat \Gamma_m = (\hat z'\hat z)^{-1} \hat z'\hat \Delta_{m}$. bcs20 highlight the importance of modeling non-linearities both in the cost function and the predicted markups. Our linearity assumptions are not restrictive insofar as $\textbf{w}$ and $\boldsymbol{z}$ are constructed flexibly from exogenous variables in the data. When stating theoretical results, the distinction between population and sample counterparts matters, but for building intuition there is no need to separate the two. We refer to both $\Delta^z_m$ and $\hat \Delta^z_m$ as predicted markups for model $m$. \looseness=-1
We begin by reexamining the conditions under which models of conduct are falsified. Models are characterized by their markups $\Delta_m$. In the ideal setting where true markups are observed, a model is falsified if the markups implied by model $m$ differ from the true markups with positive probability, or $E\big[(\Delta_{0i} - \Delta_{mi})^2\big] \neq 0$. Instead, when true markups are unobserved, b82 shows that researchers need to rely on a set of excluded instruments to distinguish any wrong model from the true one. \looseness=-1
In our setting, instruments provide a benchmark for distinguishing models through the moment condition in Assumption (ref), $E[z_i\omega_{0i}] = 0$. For each model $m$, the analog of this condition is $E[z_i(p_i-\Delta_{mi})]= 0$, where $p_i-\Delta_{mi}$ is the residualized marginal revenue under model $m$. Thus, to falsify model $m$, the correlation between the instruments and the residualized marginal revenue implied by model $m$ must be different from zero. This is in line with the result in bh14 that valid instruments need to alter the marginal revenue faced by at least one firm to distinguish conduct.
However, it is not apparent from the restriction $E[z_i(p_i-\Delta_{mi})]= 0$ how the instruments distinguish model $m$ from the truth based on their key economic feature, markups. Notice that under Assumption (ref) the covariance between residualized price and the instrument is equal to the covariance between the residualized unobserved true markup and the instrument, or $E[z_ip_i] = E[z_i\Delta_{0i}]$. This equation highlights the role of instruments: they recover from prices a feature of the unobserved true markups. Thus testing relies on the comparison between $E[z_i\Delta_{0i}]$ and $E[z_i\Delta_{mi}]$. If we rescale the moments by the variance in $z$, we can restate the falsifiable restriction for model $m$ in terms of the mean squared error (MSE) in predicted markups, which we formally connect to bh14 in the following lemma.\footnote{Our environment is an example of Case 2 discussed in Section 6 of bh14.} \looseness=-1
The lemma establishes an analog to testing with observed true markups. When $\Delta_0$ is observed, a model $m$ is falsified if its markups differ from the truth with positive probability. Here, $\Delta_0$ is unobserved and falsifying model $m$ requires the markups predicted by the instruments to differ, i.e., $\Delta_{0i}^z \neq \Delta_{mi}^z$ with positive probability. Therefore testing when markups are unobserved still relies on differences in economic features between model $m$ and the true model, insofar as these differences result in different correlations with the instruments. Thus, falsifiability in our environment is a joint feature of a pair of models and a set of instruments. \looseness=-1
Moreover, a consequence of the lemma is that the sources of exogenous variation discussed in bh14 permit testing in our context. These sources of variation not only include demand rotators considered in b82, but also own and rival product characteristics, rival cost shifters, and market demographics.\footnote{We discuss how to separately use these sources of variation in Section 6.} In addition to being exogenous, Lemma (ref) shows that instruments need to be relevant for testing in order to falsify a wrong model of conduct. In particular, a model $m$ can be falsified only if the instruments are correlated with, and therefore generate non-zero predicted markups for, at least one of $\Delta_0$ and $\Delta_m$. We illustrate this point in an example.\looseness=-1
Example 1: Consider the canonical example in b82 of distinguishing monopoly and perfect competition.\footnote{b82 also allows for non-constant marginal cost, which we consider in Appendix (ref).\looseness=-1 } Notice that, under perfect competition, both markups and predicted markups are zero. Thus, falsifying perfect competition when data are generated under monopoly (or vice versa) requires that the instruments generate non-zero monopoly predicted markups. This occurs whenever variation in the instruments induces variation in the monopoly markups. Given that these markups are a function of market shares and prices, the sources of variation in bh14 typically suffice.
While Equation (ref) is a falsifiable restriction in the population, performing valid inference on conduct in a finite sample requires two steps. First, we need to encode the falsifiable restriction into hypotheses. Second, we need strong instruments to falsify the wrong model. We turn to these problems in the next two sections.\looseness=-1
To test amongst alternative models of firm conduct in a finite sample, researchers need to choose a testing procedure, four of which have been used in the IO literature.\footnote{E.g., bcs20 use an RV test, bd19 use an Anderson-Rubin test to supplement an estimation exercise, mw17 use an estimation based test, and v07 uses a Cox test. All these procedures accommodate instruments and do not require specifying the full likelihood as was done in earlier literature b87,glv92. } As discussed below, these can be classified as model assessment or model selection tests based on how each formalizes the null hypothesis. In this section, we present the standard formulation of RV, a model selection test, and the ar49 test (AR), a model assessment test. We focus on AR as its properties in our environment are representative of the three model assessment tests used in IO to test conduct.\footnote{In Appendix (ref), we show that the other model assessment procedures have similar properties to AR.} We relate the hypotheses of these tests to our falsifiable restriction in Lemma (ref). Then, we contrast the statistical properties of RV and AR, allowing us to formalize the exact sense in which RV is robust to misspecification.\looseness=-1
Rivers-Vuong Test (RV): A prominent approach to testing non-nested hypotheses was developed in v89 and then extended to models defined by moment conditions in rv02. The null hypothesis for the test is that the two competing models of conduct have the same fit,
where ${Q}_m$ is a population measure for lack of fit in model $m$. Relative to this null, we define two alternative hypotheses corresponding to cases of better fit of one of the two models:
With this formulation of the null and alternative hypotheses, the statistical problem is to determine which of the two models has the best fit, or equivalently, the smallest lack of fit.
We define lack of fit via a GMM objective function, a standard choice for models with endogeneity. Thus, ${Q}_m = {g}_m' W {g}_m$ where ${g}_m= E[z_i(p_i-\Delta_{mi})]$ and $ W=E[z_i z_i']^{-1}$ is a positive definite weight matrix.\footnote{This weight matrix allows us to interpret ${Q}_m$ in terms of Euclidean distance between predicted markups for model $m$ and the truth, directly implementing the MSE of predicted markups in Lemma (ref).\looseness=-1} The sample analog of ${Q}_m$ is $\hat Q_m=\hat g_m'\hat W \hat g_m$ where $\hat g_m = n^{-1} \hat z'(\hat p - \hat \Delta_m)$ and $\hat W = n (\hat z'\hat z)^{-1}$.
For the GMM measure of fit, the RV test statistic is then
where $\hat\sigma^2_\text{RV}$ is an estimator for the asymptotic variance of the scaled difference in the measures of fit appearing in the numerator of the test statistic. We denote this asymptotic variance by $\sigma^2_\text{RV}$. Throughout, we let $\hat \sigma^2_\text{RV}$ be a delta-method variance estimator that takes into account the randomness in both $\hat W$ and $\hat g_m$. Specifically, this variance estimator takes the form
where $\hat V_{\ell k}^\text{RV}$ is an estimator of the covariance between $\sqrt{n}\hat W^{1/2}\hat g_\ell$ and $\sqrt{n}\hat W^{1/2}\hat g_k$. Our proposed $\hat V^\text{RV}_{\ell k}$ is given by $\hat V^\text{RV}_{\ell k} = n^{-1} \sum_{i=1}^n \hat \psi_{\ell i} \hat \psi_{ki}'$ where
This variance estimator is transparent and easy to implement. Adjustments to $\hat\psi_{mi}$ and/or $\hat V^\text{RV}_{\ell k}$ can also accommodate initial demand estimation and clustering; see Appendix C.\footnote{An alternative way of estimating this variance would be by bootstrapping, which can be costly especially when demand has to be re-estimated in each bootstrap sample.}\looseness=-1
The test statistic $T^\text{RV}$ is standard normal under the null as long as $\sigma^2_\text{RV}>0$. The RV test therefore rejects the null of equal fit at level $\alpha \in (0,1)$ whenever $\abs{ T^\text{RV} }$ exceeds the $(1-\alpha/2)$-th quantile of a standard normal distribution. If instead $\sigma^2_\text{RV} = 0$, the RV test is said to be degenerate. In the rest of this section, we maintain non-degeneracy.
While Assumption (ref) is often maintained in practice, severe inferential problems may occur when $\sigma^2_\text{RV} = 0$. These problems include large size distortions and little to no power throughout the parameter space. Thus, it is essential to understand degeneracy and diagnose the inferential problems it can cause. We return to these issues in Section (ref).
Anderson-Rubin Test (AR): In this approach, the researcher writes down the following equation for each of the two models $m$:
where $\pi_m$ is defined by the orthogonality condition $E[z e_m]=0$. She then performs the test of the null hypothesis that $\pi_m=0$ with a Wald test. This procedure is similar to an ar49 testing procedure. For this reason, we refer to this procedure as AR. Formally, for each model $m$, we define the null and alternative hypotheses:
For the true model, $\pi_m$ is equal to zero since the dependent variable in Equation (ref) is equal to $\omega_0$ which is uncorrelated with $z$ under Assumption (ref). \looseness=-1
We define the AR test statistic for model $m$ as:
where $\hat \pi_m$ is the OLS estimator of $\pi_m$ in Equation (ref) and $\hat V^\text{AR}_{mm}$ is White's heteroskedasticity-robust variance estimator. This variance estimator is $\hat V^\text{AR}_{\ell k} = n^{-1} \sum_{i=1}^n \hat \phi_{\ell i} \hat \phi_{ki}'$ where $\hat \phi_{mi} = \hat W \hat z_i\big( \hat p_i - \hat \Delta_{mi} - \hat z_i'\hat \pi_m\big)$. Under the null hypothesis corresponding to model $m$, the large sample distribution of the test statistic $T^\text{AR}_m$ is a (central) $\chi^2_{d_z}$ distribution and the AR test rejects the corresponding null at level $\alpha$ when $T^\text{AR}_m$ exceeds the $(1-\alpha)$-th quantile of this distribution.\looseness=-1
We now show that the null hypotheses of both tests can be reexpressed in terms of our falsifiable restriction in Lemma (ref).\looseness=-1
While it may seem puzzling that the hypotheses depend on a feature of the unobservable $\Delta_0,$ recall that $\Delta_0^z$ is identified by the observable covariance between $p$ and $z$. The formulation of the hypotheses in Proposition (ref) shows the connection between the statistical procedures used in applied work and the key economic implications of the models of conduct. AR and RV implement Equation (ref) through their null hypotheses, but they do so in distinct ways.
AR forms hypotheses for each model directly from Equation (ref), separately evaluating whether each model is falsified by the instruments. From Proposition (ref), the null of the AR test asserts that the MSE of predicted markups for model $m$ is zero, while the alternative is that the MSE is positive. Thus, the hypotheses depend on the absolute fit of the model measured in terms of predicted markups. In fact, we show in Appendix (ref) that other procedures used in the IO literature to distinguish models of conduct share the same null as AR. All these model assessment tests may reject both models if they both have poor absolute fit.\looseness=-1
As opposed to checking the falsifiable restriction in Lemma (ref) for each model, one could pursue a model selection approach by comparing the relative fit of the models. Proposition (ref) shows that the RV test compares the MSE of predicted markups for model 1 to the MSE for model 2. The null of the RV test asserts that these are equal. Meanwhile, the alternative hypotheses assert that the relative fit of either model 1 or model 2 is superior. If the RV test rejects, it will never reject both models, but only the one whose predicted markups are farther from the true predicted markups.
The previous section showed that our analog to the falsifiable restriction in bh14 can be used to perform either model selection or model assessment, depending on how the null is formed from the moment in Equation (ref). In this section, we explore the implications that these two formulations of the null have on inference. Crucially, as $\Delta^{z}_m$ is a function of demand parameters and is residualized with respect to cost shifters, these implications depend on whether demand or cost are misspecified.\footnote{Nonparametric estimation of demand is possible c20, yet researchers often rely on parametric estimates. While these can be good approximations, some misspecification is likely.} To provide an overview, we first contrast the performance of AR and RV in the presence of a fixed amount of misspecification for either markups or costs. Cost misspecification can be fully understood as a form of markup misspecification. We then compare AR and RV when the level of markup misspecification is local to zero.\looseness=-1
Global Markup Misspecification: When we allow markups to be misspecified by a fixed amount, important differences in the performance of the tests arise:\looseness=-1
It is instructive to interpret the lemma in the special case where model 1 is the true model. If demand elasticities are correctly specified, then $\Delta^z_1 = \Delta^z_0$. Further suppose that model 2, a wrong model of conduct, can be falsified by the instruments. For AR, the null hypothesis for model 1 is satisfied while the null hypothesis for model 2 is not. Thus, without misspecification, the researcher can learn the true model of conduct with a model assessment approach. However, it is more likely in practice that markups are misspecified such that $E\big[(\Delta^z_{0i}-\Delta^z_{1i})^2\big]\neq 0$. Regardless of the degree of misspecification, Lemma (ref) then shows that AR rejects the true model in large samples, and also generically rejects model 2.\footnote{Appendix (ref) shows that similar results obtain for other model assessment tests.} While the researcher learns that the predicted markups implied by the two models are not correct, the test gives no indication on conduct.\looseness=-1
By contrast, RV rejects in favor of the true model in large samples, regardless of misspecification, as long as $E\big[(\Delta^z_{0i}-\Delta^z_{1i})^2\big]< E\big[(\Delta^z_{0i}-\Delta^z_{2i})^2\big]$. This gives precise meaning to the oft repeated statement: RV is “robust to misspecification.” If misspecification is not too severe such that $\Delta^z_0$ is closer to $\Delta^z_1$ than to $\Delta^z_2$, RV concludes for the true model of conduct.\looseness=-1
Finally, consider the scenario where markups are misspecified and neither model is true. AR rejects any candidate model in large samples. Conversely, RV points in the direction of the model that appears closer to the truth in terms of predicted markups. If the ultimate goal is to learn the true model of conduct as opposed to the true markups, model selection is appropriate under global markup misspecification while model assessment is not.
Example 1 - continued: Consider again the case of distinguishing perfect competition from the true model of monopoly, now using market demographics as instruments. Suppose that the researcher misspecifies the demand model, for instance by estimating a mixed logit model that omits a significant interaction between demographics and product characteristics. Let $\Delta_0$ be monopoly markups with the true demand system, and $\Delta_1$ and $\Delta_2$ be the monopoly and perfect competition markups, respectively, with the misspecified demand system. Thus, $\Delta_2$ and $\Delta^z_2$ are both zero. Because substitution patterns are misspecified, the degree to which market demographics affect $\Delta_0$ and $\Delta_1$ is different. AR then rejects monopoly. Instead, as long as the MSE of $\Delta_1^z$ is smaller than the variance of $\Delta^z_0$, or $E\big[(\Delta^z_{0i}-\Delta^z_{1i})^2\big]< E\big[(\Delta^z_{0i})^2\big]$, RV concludes in favor of monopoly.
Global Cost Misspecification: In addition to demand being misspecified, it is also possible that a researcher misspecifies marginal cost. Here we show that testing with misspecified marginal costs can be reexpressed as testing with misspecifed markups so that the results in the previous section apply. As a leading example, we consider the case where the researcher specifies $\textbf{w}_\textbf{a}$ which are a subset of $\textbf{w}$. This could happen in practice because the researcher does not observe all the variables that determine marginal cost or does not specify those variables flexibly enough in constructing $\textbf{w}_\textbf{a}$.\footnote{Under a mild exogeneity condition, the results here extend to the case where $\textbf{w}\neq \textbf{w}_\textbf{a}$.}
To perform testing with misspecified costs, the researcher would residualize $\boldsymbol p$, $\boldsymbol{\Delta}_1$, $\boldsymbol{\Delta}_2$ and $\boldsymbol z$ with respect to $\textbf{w}_\textbf{a}$ instead of $\textbf{w}$. Let $y^\textbf{a}$ denote a generic variable $\boldsymbol {y}$ residualized with respect to $\textbf{w}_\textbf{a}$. Thus, with cost misspecification, both RV and AR depend on the moment\looseness=-1
where $\Breve{\Delta}^\textbf{a}_{mi}={\Delta}^\textbf{a}_{mi}-\text{w}^\textbf{a}\tau$ and $\text{w}^\textbf{a}$ is $\textbf{w}$ residualized with respect to $\textbf{w}_\textbf{a}$. When price is residualized with respect to cost shifters $\textbf{w}_\textbf{a}$, the true cost shifters are not fully controlled for and the fit of model $m$ depends on the distance between $\Delta^{\textbf{a}z}_0$ and $\Breve\Delta^{\textbf{a}z}_m$. For example, suppose model 1 is the true model and demand is correctly specified so that $\boldsymbol\Delta_0 = \boldsymbol\Delta_1$. Model 1 is still falsified as $\Breve{\Delta}^\textbf{a}_{1}={\Delta}^\textbf{a}_{0} - \text{w}^\textbf{a}\tau$. Thus, when performing testing with misspecified cost, it is as if the researcher performs testing with markups that have been misspecified by $-\text{w}^\textbf{a}\tau$. We formalize the implications of cost misspecification on testing in the following lemma:
Thus the effects of cost misspecification can be fully understood as markup misspecification, so we focus on markup misspecification in what follows. \looseness=-1
Local Markup Misspecification: To more fully understand the role of misspecification on inference of conduct, we would like to contrast the performance of AR and RV in finite samples. However, it is not feasible to characterize the exact finite sample distribution of AR and RV under our maintained assumptions. Instead, we can approximate the finite sample distribution of each test by considering local misspecification, i.e., a sequence of candidate models that converge to the null space at an appropriate rate. As model assessment and model selection procedures have different nulls, we define distinct local alternatives for RV and AR based on $\Gamma_m$. For model assessment, local misspecification is characterized in terms of the absolute degrees of misspecification for each model:\looseness=-1
By contrast, local alternatives for model selection are in terms of the relative degree of misspecification between the two models:
Under the local alternatives in Equations (ref) and (ref), we approximate the finite sample distribution of AR and RV with misspecification in the following proposition. To facilitate a characterization in terms of predicted markups, we define stable versions of predicted markups under either of the two local alternatives considered: $\Delta_{mi}^{\text{RV},z} = n^{1/4}\Delta_{mi}^{z}$ and $\Delta_{mi}^{\text{AR},z} = n^{1/2}\Delta_{mi}^{z}$. We also introduce an assumption of homoskedastic errors, which in this section serves to simplify the distribution of the AR statistic:\looseness=-1
The intuition developed in this section does not otherwise rely on Assumption (ref).\looseness=-1
From Proposition (ref), both test statistics follow a non-central distribution. However, the non-centrality term differs for the two tests because of the formulation of their null hypotheses. For AR, the non-centrality for model $m$ is the ratio of the MSE of predicted markups to the noise given by $\sigma^2_{m}$. Alternatively, the noncentrality term for RV depends on the ratio of the difference in MSE for the two models to the noise. Thus, how one formulates the hypotheses from Equation (ref) also affects inference on conduct in finite samples.
We illustrate the relationship between hypothesis formulation and inference in Figure (ref). Each panel represents, for either model selection or model assessment, and for either a high or low noise environment, the outcome of testing in the coordinate system of MSE in predicted markups $\big(E\big[(\Delta^z_{0i}-\Delta^z_{1i})^2\big],E\big[(\Delta^z_{0i}-\Delta^z_{2i})^2\big]\big)$. Regions with horizontal lines indicate concluding for model 1, while regions with vertical shading indicate concluding for model 2. \looseness=-1
Panels A and C correspond to testing with RV in a high and low noise environment respectively. As the noise declines from A to C, the noncentrality term for RV, whose denominator depends on the noise, increases. Hence the shaded regions expand towards the null space and the RV test becomes more conclusive in favor of a model of conduct. Conversely, Panels B and D correspond to testing with AR in a high and a low noise environment. As the noise decreases from Panel B to D, the noncentrality term of AR, whose denominator depends on the noise, increases and the shaded regions approach the two axes. Thus, as the noise decreases, AR rejects both models with higher probability. If the degree of misspecification is low, the probability RV concludes in favor of the true model increases as the noise decreases. Instead, AR only concludes in favor of the true model with sufficient noise.\looseness=-1
An analogy may be useful to summarize our discussion in this section. Model selection compares the relative fit of two candidate models and asks whether a “preponderance of the evidence” suggests that one model fits better than the other. Meanwhile, model assessment uses a higher standard of evidence, asking whether a model can be falsified “beyond any reasonable doubt.” While we may want to be able to conclude in favor of a model of conduct beyond any reasonable doubt, this is not a realistic goal in the presence of misspecification. If we lower the evidentiary standard, we can still learn about the true nature of firm conduct. Hence, in the next section we focus on the RV test. However, to this point we have assumed $\sigma^2_\text{RV} > 0$ and thereby assumed away degeneracy. We address this threat to inference with the RV test in the next section.
Having established the desirable properties of RV under misspecification, we now revisit Assumption (ref). First, we connect degeneracy to our falsifiable restriction in Lemma (ref) and show that maintaining Assumption (ref) is equivalent to ex ante imposing that at least one of the models is falsified by the instruments. To explore the consequences of such an assumption, we define a novel weak instruments for testing asymptotic framework adapted from ss97 and for which degeneracy occurs. We use the weak instrument asymptotics to show that degeneracy can cause size distortions and low power in finite samples. To help researchers interpret the frequency with which the RV test makes errors, we propose a diagnostic in the spirit of sy05. This proposed diagnostic is a scaled $F$-statistic computed from two first stage regressions and researchers can use it to gauge the extent to which inferential problems are a concern.\looseness=-1
We first characterize when the RV test is degenerate in our setting. Since $\sigma^2_\text{RV}$ is the asymptotic variance of $\sqrt{n}(\hat Q_1-\hat Q_2)$, it follows that Assumption (ref) fails to be satisfied whenever $\hat Q_1-\hat Q_2=o_p\big( 1/\sqrt{n} \big)$ rv02. In the following proposition, we reinterpret this condition through the lens of our falsifiable restriction.
The proposition shows that when $\sigma^2_\text{RV}=0$, neither model is falsified by the instruments. Thus, Assumption (ref) is equivalent to assuming the falsifiable restriction in Equation (ref) is violated for at least one model.
Such a characterization permits us to better understand degeneracy. Consider two extreme cases where instruments are weak: (i) the instruments are uncorrelated with $\Delta_0$, $\Delta_1$, and $\Delta_2$ such that $z$ is irrelevant for testing of either model, and (ii) models 0, 1, and 2 imply similar markups such that $\Delta_1$ and $\Delta_2$ overlap with $\Delta_0$. Much of the econometrics literature focuses on (ii) as it considers degeneracy in the maximum likelihood framework of v89. As RV generalizes the v89 test to a GMM framework, degeneracy is a broader problem that encompasses instrument strength. We illustrate these ideas in the following two examples that correspond respectively to cases (i) and (ii).\looseness=-1
Example 2: Consider an industry where firms compete across many local markets, but charge uniform prices across all markets. Suppose a researcher wants to distinguish a model of uniform Bertrand pricing ($m=1$) and a model of uniform monopoly pricing ($m=2$). Let $m=1$ be the true model and assume demand and cost are correctly specified so that $\Delta_0 = \Delta_1$. The researcher forms instruments from local variation in rival cost shifters. If the number of markets is large, the contribution of any one market to the firm-wide pricing decision is negligible. Thus, the local variation leveraged by the instruments becomes weakly correlated with $\Delta_0$, $\Delta_1$, and $\Delta_2$, resulting in degeneracy.\looseness=-1
Example 3: Consider three models of simple “rule of thumb” pricing, where markups are a fixed fraction of cost. Suppose that the true model implies markups $\boldsymbol \Delta_0 = \boldsymbol c_0,$ and models 1 and 2 correspond to $\boldsymbol \Delta_1 = 0.5 \boldsymbol c_0$ and $\boldsymbol \Delta_2 = 2 \boldsymbol c_0$ respectively. Given that $\boldsymbol c_0 = \textbf{w}\tau + \omega_0$, the residualized markups are $\Delta_0 = \omega_0$, $\Delta_1 = 0.5\omega_0$ and $\Delta_2 = 2\omega_0$. As the instruments are uncorrelated with $\omega_0$, they are also uncorrelated with the residualized markups for all three models, and predicted markups are therefore zero. From the perspective of the instruments, both model 1 and model 2 overlap with the true model, and degeneracy obtains for any choice of $z$ satisfying Assumption (ref).
As shown in the examples, degeneracy can occur in standard economic environments. It is therefore important to understand the consequences of violating Assumption (ref). To do so, we connect degeneracy to the formulation of the null hypothesis of RV. From Proposition (ref), degeneracy occurs as a special case of the null of RV. Intuitively, when degeneracy occurs, there is not enough information to falsify either model in the population. Thus, both models have perfect fit. Figure (ref) illustrates this point by representing both the null space and the space of degeneracy in the coordinate system of MSE of predicted markups $\big(E\big[(\Delta^z_{0i}-\Delta^z_{1i})^2\big],E\big[(\Delta^z_{0i}-\Delta^z_{2i})^2\big]\big)$. While the null hypothesis of RV is satisfied along the full 45-degree line, degeneracy only occurs at the origin.\footnote{If $\Delta_0 = \Delta_1$, the graph shrinks to the $y$-axis and degeneracy arises whenever the null of RV is satisfied. This special case is in line with hp11, who note RV is degenerate if both models are true.}\looseness=-1
As degeneracy is a special case of the null, maintaining Assumption (ref) has no consequences for size control if the RV test reliably fails to reject the null under degeneracy. However, we show that degeneracy can cause size distortions and a substantial loss of power close to the null. To make this point, we recast degeneracy as a problem of weak instruments.\looseness=-1
Proposition (ref) shows that degeneracy arises when the predicted markups across models 0, 1 and 2 are indistinguishable. Given the definition of predicted markups in Equation (ref), this implies that the projection coefficients from the regression of markups on the instruments: $\Gamma_0$, $\Gamma_1$, and $\Gamma_2$ are also indistinguishable. Thus, we can rewrite Proposition (ref) as follows:
Degeneracy is characterized by $\Gamma_0-\Gamma_m$ being zero for both $m=1$ and $m=2$. Thus when models are fixed and $\Gamma_0-\Gamma_m$ is constant in the sample size, degeneracy is a problem of irrelevant instruments.\looseness=-1
To better capture the finite sample performance of the test when the instruments are nearly irrelevant, it is useful to conduct analysis allowing $\Gamma_0-\Gamma_m$ to change with the sample size. Thus, we forgo the classical approach to asymptotic analysis where the models are fixed as the sample size goes to infinity. Instead, we now adapt ss97's asymptotic framework of weak instruments in the following assumption:\looseness=-1
Here, the projection coefficients $\Gamma_0-\Gamma_m$ change with the sample size and are local to zero which enables the asymptotic analysis in the next subsection. This approach is technically similar to the analysis of local misspecification conducted in Proposition (ref). However, it does not impose Assumption (ref). Instead, Assumption (ref) implies that $\sigma^2_\text{RV}$ is zero so that degeneracy obtains. Thus, in the next subsection, we use weak instrument asymptotics to clarify the effect of degeneracy on inference.
We now use Assumption (ref) to show that RV has inferential problems under degeneracy and to provide a diagnostic for instrument strength in the spirit of sy05. The diagnostic relies on formulating an $F$-statistic that can be constructed from the data. An appropriate choice is the scaled $F$-statistic for testing the joint null hypotheses of the AR model assessment approach for the two models. The motivation behind this statistic is Corollary (ref). Note that $\Gamma_0-\Gamma_m = E[z_iz_i']^{-1}E[z_i(p_i-\Delta_{mi})] = \pi_m$, the parameter being tested in AR. Thus, instruments are weak for testing if both $\pi_1$ and $\pi_2$ are zero, and degeneracy occurs when the null hypotheses of the AR test for both models, $H_{0,1}^\text{AR}$ and $H_{0,2}^\text{AR}$, are satisfied.\looseness=-1
A benefit of relying on an $F$-statistic to construct a single diagnostic for the strength of the instruments is that its asymptotic null distribution is known. However, it is more informative to scale the $F$-statistic by $1-\hat \rho^2$ where $\hat \rho^2$ is the squared empirical correlation between $e_{1i}-e_{2i}$ and $e_{1i}+e_{2i}$, where $e_m$ is the error in the regression of $p-\Delta_m$ on $z$ used to estimate $\pi_m$. Expressed formulaically, our proposed $F$-statistic is then
While maintaining homoskedasticty as in Assumption (ref), we will describe how $F$ can be used to diagnose the quality of inferences made based on the RV test. In the language of olea13, ours is an effective $F$-statistic as it relies on heteroskedasticity-robust variance estimators.\footnote{$F$ is closely related to the likelihood ratio statistic for the test of $\pi_1=\pi_2=0$. However, the likelihood ratio statistic does not scale by $1-\hat \rho^2$ nor does it use heteroskedasticity-robust variance estimators as $F$ does.\looseness=-1} For this reason, we expect that $F$ remains useful to diagnose weak instruments outside of homoskedastic settings. For simulations that support this expectation in the standard IV case, we refer to ass19.
In the following proposition, we characterize the joint distribution of the RV statistic and our $F$. As our goal is to learn about inference and to provide a diagnostic for size and power, we only need to consider when the RV test rejects, not the specific direction. Thus, we derive the asymptotic distribution of the absolute value of $T^\text{RV}$ in the proposition. This result forms the foundation for interpretation of $F$ in conjunction with the RV statistic. We use the notation $\boldsymbol e_1$ to denote the first basis vector $\boldsymbol e_1 = (1,0,\dots,0)' \in \mathbb{R}^{d_z}$.\footnote{Proposition (ref) introduces objects with plus and minus subscripts, as these objects are sums and differences of rotated versions of $W^{1/2}g_1$ and $W^{1/2}g_2$ and their estimators. The role of these objects is discussed after the proposition, while we defer a full definition to Appendix (ref) to keep the discussion concise.}
The proposition shows that the asymptotic distribution of $T^\text{RV}$ and $F$ in the presence of weak instruments depends on $\rho$ and two non-negative nuisance parameters, $\mu_-$ and $\mu_+$, whose magnitudes are tied to whether $H_0^\text{RV}$ holds, and to whether $H_{0,1}^\text{AR}$ and $H_{0,2}^\text{AR}$ hold, respectively. Specifically, the null of RV corresponds to $\mu_-=0$. Furthermore, the proposition sheds light on the effects that degeneracy has on inference for RV. Unlike the standard asymptotic result, the RV test statistic converges to a non-normal distribution in the presence of weak instruments. For compact notation, let this non-normal limit distribution be described by the variable $T^\text{RV}_\infty = \Psi_-' \Psi_+/\!\left( \norm{ \Psi_-}^2 + \norm{ \Psi_+}^2 + 2\rho \Psi_-' \Psi_+ \right)^{1/2}$. Under the null, the numerator of $T^\text{RV}_\infty$ is the product of $\Psi_-$, a normal random variable centered at 0, and $\Psi_+$, a normal random variable centered at $\mu_+ \ge 0$. When $\rho \neq 0$, the distribution of this product is not centered at zero and is skewed, both of which may contribute to size distortions.
Alternatives to the RV null are characterized by $\mu_- \in (0,\mu_+]$. For a given value of $\rho$, maximal power is attained when $\mu_-=\mu_+$. This maximal power is strictly below one for any finite $\mu_+$, so that the test is not consistent under weak instruments. Actual power will often be less than the envelope, as $\mu_{-} = \mu_{+}$ generally only occurs with no misspecification. Furthermore, the lack of symmetry in the distribution of the RV test statistic when $\rho \neq 0$ leads to different levels of maximal power for each model. \looseness=-1
Ideally, one could estimate the parameters and then use the distribution of the RV statistic under weak instruments asymptotics to quantify the distortions to size and the maximal power that can be attained. However, this is not viable since $\mu_{-}$, $\mu_{+}$, and the sign of $\rho$ are not consistently estimable. Instead, we adapt the approach of sy05 and develop a diagnostic to determine whether $\mu_+$ is sufficiently large to ensure control of the highest possible size distortions. Given the threat of low power, we develop a similar diagnostic to ensure a lower bound on the maximal power for both models.\looseness=-1
One might wonder if robust methods from the IV literature would be preferable when instruments are weak. For example, AR is commonly described as being robust to weak instruments in the context of IV estimation. Note that while AR maintains the correct size under weak instruments, this is of limited usefulness for inference with misspecification since neither null is satisfied. Furthermore, tests proposed in k02 and m03 do not immediately apply to our setting. The econometrics literature has also developed modifications of the v89 test statistic that seek to control size under degeneracy s15, sw17. While these may be adaptable to our setting, the benefits of size control may come at the cost of lower power. As we show in the next section, power as opposed to size is the main concern with a moderate number of instruments.\looseness=-1
To implement our diagnostic for weak instruments, we need to define a target for reliable inference. Motivated by the practical considerations of size and power, we provide two such targets: a worst-case size $r^s$ exceeding the nominal level of the RV test ($\alpha = 0.05$) and a maximal power $r^p$. Then, we construct separate critical values for each of these targets. A researcher can choose to diagnose whether instruments are weak based on size, power, or ideally both by comparing $F$ to the appropriate critical value. We construct the critical values based on size and power in turn.
Diagnostic Based On Maximal Size: We first consider the case where the researcher wants to understand whether the RV test has asymptotic size no larger than $r^s$ where $r^s \in (\alpha,1)$. For each value of $\rho$, we then follow sy05 in denoting the values of $\mu_+$ that lead to a size above $r^s$ as corresponding to weak instruments for size:
The role of $F$, when viewed through the lens of size control, is to determine whether it is exceedingly unlikely that the true value of $\mu_+$ corresponds to weak instruments for size for any value of $\rho$. Using the distributional approximation to $F$ in Proposition (ref) and the standard burden of a five percent probability to denote an exceedingly unlikely event, we say that the instruments are strong for size whenever $F$ exceeds
where $\chi^2_{df,.95}(nc)$ denotes the upper $95$th percentile of a non-central $\chi^2$-distribution with degrees of freedom $df$ and non-centrality parameter $nc$. Note that $cv^s$ will vary by the number of instruments $d_z$ and the tolerated test level $r^s$.
Diagnostic Based on Power Envelope: For interpretation of the RV test, particularly when the test fails to reject, it is important to understand the maximal power that the test can attain. By considering rejection probabilities when $\mu_-= \mu_+$ and linking these probabilities to values of $F$, it is also possible to let the data inform us about the power potential of the test. To do so we consider an ex ante desired target of maximal power $r^p$ and define weak instruments for power as the values of $\mu_+$ that lead to maximal power less than $r^p$:\looseness=-1
We determine the strength of the instruments by considering the power envelope for the RV test for any value of $\rho$. This is to ensure that the power against both models exceeds $r^p$ for any value of $\rho$. Again using the distributional approximation to $F$ in Proposition (ref), we say that the instruments are strong for power if $F$ is larger than
It is important to stress that the event $F > cv^p$ expresses that the maximal power against both models is above $r^p$ with high probability for the worst-case $\rho$. The power that the test attains in a given application depends both on $\rho$ and on the degree of misspecification. With misspecification, the actual power of the test will be smaller than the envelope for the worst-case $\rho$. However, for values of $\rho$ different than the worst-case, the power may exceed $r^p$ for one of the models. In this way, our diagnostic for power is informative about the RV test when the null is not rejected.\looseness=-1
\noindentComputing Critical Values: To compute $cv^s$ for a given $(d_z,r^s)$, we numerically solve for $\sup {\cal S}(\rho; r^s)$. The symmetry of the problem implies that the probability used to define ${\cal S}(\rho; r^s)$ does not depend on the sign of $\rho$ so we only need to consider $\rho \in [0,1)$. Thus, we consider a grid of values for $\rho$ from 0 to 1 at steps of 0.01. For each value of $\rho$, we find $\sup {\cal S}(\rho;r^s)$ numerically, by considering a large grid for $\mu_+$ that extends from zero to 80. To compute $cv^p$ for a given $(d_z,r^p)$, we use the same procedure as for size, but for $\mu_{-} = \mu_{+}$ instead of $\mu_{-} = 0$ and $\rho\in(-1,1)$ as the problem is no longer symmetric.\looseness=-1
Discussion of the Diagnostic: To diagnose whether instruments are weak for size or power, a researcher would compute $F$ and compare it to the relevant critical value. Table (ref) reports the critical values used to diagnose whether instruments are weak in terms of size (Panel A) or power (Panel B). These critical values explicitly depend on both the number of instruments $d_z$ and a target for reliable inference.\footnote{Additionally, their use requires residualizing the variables with respect to $\textbf{w}$ and forming $Q_m$ with the 2SLS weight matrix $W = E[z_iz_i']$, as assumed throughout the paper.} The table reports critical values for up to 30 instruments. For size, we consider targets of worst-case size $r^s \in \{ 0.075,\ 0.10,\ 0.125\}$. For power, we consider targets of maximal power $r^p \in \{ 0.95,\ 0.75,\ 0.50\}$.
Suppose a researcher wanting to diagnose whether instruments are weak based on size has fifteen instruments and measures $F = 6$. Given a target worst-case size of 0.10, the critical value in Panel A is 5.0. Since $F$ exceeds $cv^s$, the researcher concludes that instruments are strong in the sense that size is no larger than $0.10$ with at least 95 percent confidence. Instead, for a target of 0.075, the critical value is 10.8. In this case, $F < cv^s$ and the researcher cannot conclude that the instruments are strong for size. Thus, the interpretation of our diagnostic for weak instruments based on size is analogous to the interpretation that one draws for standard IV when using an $F$-statistic and sy05 critical values.\looseness=-1
If the researcher also wants to diagnose whether instruments are weak based on power, she can compare $F$ to the relevant critical value in Panel B. For fifteen instruments and a target maximal power of 0.75, the critical value is again 5.0. Since $F = 6$, the researcher can conclude that instruments are strong in the sense that the maximal power the test could obtain exceeds 0.75 with at least 95 percent confidence. Instead, for a target maximal power of 0.95, the critical value is 6.3. In this case, $F < cv^p$ and the researcher cannot conclude that the instruments are strong for power.
The columns of Panels A and B in Table (ref) are sorted in terms of increasing maximal type I (Panel A) and type II errors (Panel B). Unsurprisingly, the critical values decrease with the target error as larger $F$-statistics are required to conclude for smaller type I and II errors. Inspection of the columns are useful to understand when size distortions and low power are relevant threats to inference. The RV test statistic has a skewed distribution whose mean is not zero. The effect of skewness on size is largest with one instrument, so in Panel A, the critical value is large when $d_z=1$. As the effect of skewness on size decreases in $d_z$, there are no size distortions exceeding 0.025 with 2-9 instruments. Meanwhile, the effect of the mean on size is increasing in $d_z$, and becomes relevant when $d_z$ exceeds 9. Thus the critical values are monotonically increasing from 10 to 30 instruments. Inspection of the first column of Panel A also suggests a simple rule of thumb for diagnosing weak instruments in terms of maximal size equal to 0.075. For $d_z > 9$, instruments are strong if $F > 2(d_z-9)$. Alternatively, for power, the critical values are monotonically decreasing in the number of instruments. Taken together, the critical values indicate that (except for the case of one instrument) low maximal power is the main concern when testing with a few instruments, while size distortions are the main concern when testing with many instruments. \looseness=-1
To illustrate the usefulness of our $F$-statistic, consider an example where the researcher has two instruments and computes an RV test statistic $T^\text{RV} = 0.54$. For a target size of 0.075, the critical value is zero and there are no size distortions above 0.025. Thus, low power is the only salient concern. If the $F$-statistic is below 10.4 which is the critical value for target maximal power of 0.5, then the researcher can conclude rejection was very unlikely in this setting even if the null is violated. Suppose instead that $T^\text{RV} = 5.54$ and $F = 10$. This occurrence may seem pathological as the test rejects the null although our diagnostic detects low power. However, recall that the critical value of 10.4 is computed to ensure that maximal power against both models exceeds 0.5 for all values of $\rho$. Thus, power against one of the models could be higher than 0.5 even though $F=10$. In other words, when power is the salient concern, our $F$-statistic is necessary to interpret no rejection. Likewise, when size is a concern, our $F$-statistic is necessary to interpret rejections of the null.
Up to this point, we have considered the case where the researcher has one set of instruments they will use for testing two candidate models. Indeed, if the researcher chooses their instruments for testing once-and-for-all based on intuition, the procedure for testing conduct is straightforward: run the RV test and then inspect whether the instruments pass the diagnostic for strength. In practice, several sets of instruments may be available to the researcher. Furthermore, in many settings including our application, a researcher wants to test more than two models. In the next section, we discuss how an applied researcher can perform RV testing on multiple models with multiple sets of instruments while using the $F$-statistic to guide inference.
bh14 show that multiple sources of exogenous variation in marginal revenue can be used to construct instruments for testing conduct. As mentioned in Section (ref), these typically include demand rotators, own and rival product characteristics, rival cost shifters, and market demographics. A researcher wanting to exploit variation from all available sources in her application faces two major decisions. First, should she run one RV test with a single pooled set of instruments or should she keep the sources of variation separate and run multiple RV tests? Second, which functional form should the researcher use to construct instruments from her chosen sources? The latter point is addressed in bcs20, who consider efficiency in the spirit of c87. In this section, we focus instead on the first consideration.
Based on the results in Sections (ref) and (ref), there are two main reasons a researcher may want to keep the sources of variation separate. First, drawing inference on conduct by pooling sources of variation can conceal the severity of misspecification. As seen in Section (ref), the RV test concludes for the model with the lower MSE of predicted markups. With misspecification, strong instruments constructed from economically different sources of variation (e.g., demand shifters versus rival cost shifters) could conclude for different models. By keeping the sources of variation separate and running multiple RV tests, a researcher can observe such conflicting evidence. Instead, a single RV test run with pooled instruments could conclude for one model, obscuring the severity of misspecification. Below, Example 4 provides an economic setting where misspecifying models generates conflicting evidence.\looseness=-1
Second, pooling variation may have adverse consequences for the strength of the resulting instrument set, which occurs in our empirical application (see Appendix (ref)). For example, if some sources of variation on their own yield weak instruments for power, combining these with strong instruments dilutes the power of the strong instruments, manifesting itself in a lower $F$-statistic. Furthermore, Panel A of Table (ref) shows that the combined set of instruments faces a larger critical value for size. Thus, if pooling across sources creates many instruments, size distortions can undermine inference on firm conduct.
Example 4: Consider firms that compete across many local markets but charge uniform Bertrand prices. In each period, firms offer the same products in all local markets. Suppose a researcher specifies two incorrect models: perfect competition ($m=1$) and local Bertrand pricing ($m=2$). The researcher constructs two sets of instruments, one from local variation in rival cost and the other from local variation in rival product characteristics. As in Example 2, local variation in cost is weakly correlated with true markups so that $\Delta^z_0 = \Delta^z_1 = 0$ while $\Delta^z_2\neq 0$. Thus, with cost instruments, the researcher concludes for perfect competition. Instead, local variation in product characteristics similarly moves both uniform and local Bertrand markups as firms add and drop products in all markets. Under most formulations of demand and cost, the researcher concludes for local Bertrand competition. Because both models are misspecified, the two sets of instruments generate conflicting evidence.\looseness=-1
Accumulating Evidence: Researchers who want to keep their sources of variation separate need to aggregate information across multiple RV tests. We suggest that a researcher can conclude for a model insofar as there is no conflicting evidence across sets of instruments and all the strong instruments support it. Continuing the legal analogy made in Section (ref), we have adopted a preponderance of the evidence standard by using model selection. However, we may not want to rely on a single piece of evidence to convict, nor would we want to rely on weak evidence. To achieve the two aims above, we propose a conservative approach that utilizes both the RV test and the $F$-statistic.
Suppose we want to test a set of two models $M=\{1,2\}$ using $L$ sets of instruments. In a preliminary step we run separate RV tests with each instrument set $z_\ell$ and denote the model confidence set (MCS) $M^*_\ell$ as the set of models that are not rejected.\footnote{Thus, $M^*_\ell = \{1,2\}$ if the RV null is not rejected and $M^*_\ell = \{1\}$ if the RV null is rejected in favor of a superior fit of model 1.} Our goal is to generate $M^*$, an MCS which aggregates evidence from all $M^*_\ell$. Our approach, illustrated in Figure (ref), proceeds in two steps. In step 1, the researcher needs to check that the evidence coming from the $L$ sets of instruments is not in conflict. We say that evidence arising from $L$ RV tests is not in conflict if, for every pair of ($M^*_\ell$, $M^*_{\ell'}$), one is a weak subset of the other. In step 2 we form $M^*$ based on step 1. If the evidence is in conflict, the researcher concludes $M^* = \{1,2\}$. If the evidence is not in conflict, we first set $M^*$ equal to the smallest MCS $M^*_\ell$, and then take the union with all MCS for which the instruments are strong based on the $F$-statistic. \looseness=-1
To illustrate the rationale behind our approach, we consider a few examples. In each, we use $L = 2$ sets of 2-9 instruments, so that there are no size distortions above 0.025 and power is the salient concern. First, we illustrate the importance of step 1. If the researcher had computed $M^*_1 = \{1\}$ and $M^*_2 = \{2\}$, then the instruments $z_1$ suggest model 2 can be rejected in favor of superior fit of model 1 while $z_2$ suggest the exact opposite. As $M^*_1 \not\subseteq M^*_2$ and $M^*_2 \not\subseteq M^*_1$ we say the evidence is in conflict. Hence, misspecification is severe and the researcher should let $M^* = \{1,2\}$, in line with the conservative spirit of the procedure.
Suppose now that $M^*_1 = \{1\}$ and $M^*_2 = \{1,2\}$. Because $M^*_1 \subset M^*_2$, there is no conflicting evidence found in step 1. In step 2 we initialize $M^* = \{1\}$, the smallest MCS. By doing so, we use the information that $z_1$ reject model 2 regardless of the power potential diagnosed by the $F$-statistic. We then only add model 2 to $M^*$ if the $F$-statistic suggests that instruments $z_2$ are strong for power. If $z_2$ are weak, then not rejecting the null is likely a consequence of low power and not informative about firm conduct.
Extension to More than Two Models: In many settings, including our application, a researcher may want to test a set $M$ of more than two models. To accumulate evidence across sets of instruments using the procedure in Figure (ref), we need to define $M^*_\ell$ for each of the $L$ instrument sets. We adopt the procedure of hln11 to construct each $M^*_\ell$. This procedure initializes the $M^*_\ell$ to $M$, and then checks in each iteration whether the model of worst fit according to MSE of predicted markups can be excluded. This occurs if the largest RV test statistic in magnitude across all pairs of models in $M^*_\ell$ exceeds the $(1-\alpha)$-th quantile of its asymptotic null distribution.\footnote{This quantile can be simulated by drawing from the asymptotic null distribution, see Appendix (ref).} When no model can be excluded, the procedure stops. If there are only two models, this procedure coincides with the RV test as discussed above. As shown in hln11, ${M}^*_\ell$ controls the familywise error rate as it contains the model(s) with the best fit with probability at least $1-\alpha$ in large samples. Moreover, every other model with strictly worse fit is excluded from $M^*_\ell$ with probability approaching one.\footnote{Under no degeneracy, $M^*$ is guaranteed to contain the true model with probability at least $1-\alpha$, as each $M^*_\ell$ has the same property and $M^*$ is the union of these model confidence sets.} \looseness=-1
To illustrate the construction of ${M}^*_\ell$, suppose a researcher wants to test candidate models $m=1,2,3$. For a given set of instruments $z_\ell$, the MCS procedure computes three RV test statistics $T^\text{RV}_{m,m'} = {\sqrt{n}(\hat{Q}_{m\phantom{\!'}}-\hat{Q}_{m'})}/{\hat\sigma_{\text{RV},mm'}}$, one for each distinct pair of models. Suppose $T^\text{RV}_{1,2} = 5.34$, $T^\text{RV}_{1,3} = 4.35$, and $T^\text{RV}_{2,3} = 0.32$. If $T^\text{RV}_{1,2}$, the largest test statistic in magnitude, exceeds the critical value for the max of three RV test statistics, then model 1 is excluded from $M^*_\ell$. In the next iteration, only models 2 and 3 remain, so the only relevant RV test statistic is $T^\text{RV}_{2,3} = 0.32$. As the null of equal fit cannot be rejected, $M^*_\ell = \{2,3\}$.
We revisit the empirical setting of v07. She investigates the vertical relationship of yogurt manufacturers and supermarkets by testing different models of vertical conduct.\footnote{v07 uses a Cox test which is a model assessment procedure with similar properties to AR, as shown in Appendix (ref).} This setting is ideal to illustrate our results as theory suggests a rich set of models and the data is used in many applications.
Our main source of data is the IRI Academic Dataset for 2010 bkm08. This dataset contains weekly price and quantity data for UPCs sold in a sample of stores in the United States. We define a market as a retail store-quarter and approximate the market size with a measure of the traffic in each store, derived from the store-level revenue information from IRI. We drop the 5% of stores for which this approximation results in an unrealistic outside share below 50%. \looseness=-1
We further restrict attention to UPCs labelled as “yogurt” in the IRI data and focus on the most commonly purchased sizes: 6, 16, 24 and 32 ounces. Similar to v07, we define a product as a brand-fat content-flavor-size combination, where flavor is either plain or other and fat content is either light (less than 4.5% fat content) or whole. We further standardize package sizes by measuring quantity in six ounce servings. Based on market shares, we exclude niche firms for which their total inside share in every market is below five percent. We drop products from markets for which their inside share is below 0.1 percent. Our final dataset has 205,123 observations for 5,034 markets corresponding to 1,309 stores.\looseness=-1
We supplement our main dataset with county level demographics from the Census Bureau's PUMS database which we match to the DMAs in the IRI data. We draw 1,000 households for each DMA and record standardized household income and age of the head of the household. We exclude households with income lower than \$12,000 or bigger than \$1 million. We also obtain quarterly data on regional diesel prices from the US Energy Information Administration. With these prices, we measure transportation costs as average fuel cost times distance between a store and manufacturing plant.\footnote{We thank Xinrong Zhu for generously sharing manufacturer plant locations used in z21.} We summarize the main variables for our analysis in Table (ref).\looseness=-1
To perform testing, we need to estimate demand and construct the markups implied by each candidate model of conduct.
Demand Model: Our model of demand follows v07 in adopting the framework from blp95. Each consumer $i$ receives utility from product $j$ in market $t$ according to the indirect utility:\looseness=-1
where $\boldsymbol{x}_j$ includes package size, dummy variables for low fat yogurt and for plain yogurt, and the log of the number of flavors offered in the market to capture differences in shelf space across stores. $\boldsymbol{p}_{jt}$ is the price of product $j$ in market $t$, and $\boldsymbol{\xi}_{t}$, $\boldsymbol{\xi}_{s}$, and $\boldsymbol{\xi}_{b(j)}$ denote fixed effects for the quarter, store, and brand producing product $j$ respectively. $\boldsymbol{\xi}_{jt}$ and $\boldsymbol{\epsilon}_{ijt}$ are unobservable shocks at the product-market and the individual product market level, respectively. Finally, consumer preferences for characteristics ($\beta^x_i$) and price ($\beta^p_i$) vary with individual level income and age of the head of household:
where $\bar\beta^p$ and $\bar\beta^x$ represent the mean taste, $\boldsymbol{D}_i$ denotes demographics, while $\tilde\beta^p$ and $\tilde \beta^x$ measure how preferences change with $\boldsymbol{D}_i$.
To close the model we make additional standard assumptions. We normalize consumer $i$'s utility from the outside option as $\boldsymbol{u}_{i0t} = \boldsymbol{\epsilon}_{i0t}$. The shocks $\boldsymbol{\epsilon}_{ijt}$ and $\boldsymbol{\epsilon}_{i0t}$ are assumed to be distributed i.i.d. Type I extreme value. Assuming that each consumer purchases one unit of the good that gives her the highest utility from the set of available products $\mathcal{J}_{t}$, the market share of product $j$ in market $t$ takes the following form:\looseness=-1
Identification and Estimation: Demand estimation and testing can either be performed sequentially, in which demand estimation is a preliminary step, or simultaneously by stacking the demand and supply moments. Following v07, we adopt a sequential approach which is simpler computationally while illustrating the empirical relevance of the findings in Sections (ref), (ref), and (ref).\looseness=-1
The demand model is identified under the assumption that demand shocks $\boldsymbol{\xi}_{jt}$ are orthogonal to a vector of demand instruments. By shifting supply, transportation costs help to identify the parameters $\bar \beta^p$, $\tilde \beta^p$, and $\tilde \beta^x$. Following gh19, we use variation in mean demographics across DMAs as a source of identifying variation by interacting them with both fuel cost and product characteristics. We estimate demand as in blp95 using PyBLP cg19.\looseness=-1
Results: Results for demand estimation are reported in Table (ref). As a reference, we report estimates of a standard logit model of demand in Columns 1 and 2. In Column 1, the logit model is estimated via OLS. In Column 2, we use transportation cost as an instrument for price and estimate the model via 2SLS. When comparing OLS and 2SLS estimates, we see a large reduction in the price coefficient, indicative of endogenity not controlled for by the fixed effects. Column 3 reports estimates of the full demand model which generates elasticities comparable to those obtained in v07.
Models of Conduct: We consider five models of vertical conduct from v07.\footnote{v07 also considers retailer collusion and vertically integrated monopoly. As we do not observe all retailers in a geographic market, we cannot test those models.} A full description of the models is in Appendix (ref).
Given our demand estimates, we compute implied markups $\boldsymbol{\Delta}_m$ for each model $m$. We specify marginal cost as a linear function of observed shifters and an unobserved shock. We include in $\textbf{w}$ an estimate of the transportation cost for each manufacturer-store pair and dummies for quarter, brand and city.
Inspection of Implied Markups and Costs: Economic restrictions on price-cost margins ${\boldsymbol{\Delta}_m}/{\boldsymbol{p}}$ (PCM) and estimates of cost parameters $\tau$ may be used to learn about conduct, and are complementary to formal testing. For every model, we estimate $\tau$ by regressing implied marginal cost on the transportation cost and fixed effects. The coefficient of transportation cost is positive for all models, consistent with intuition. Thus, no model can be ruled out based on estimates of $\tau$.\looseness=-1
Figure (ref) reports the distributions of PCM for all models. Compared to Table 7 in v07, our PCM are qualitatively similar both in terms of median and standard deviations, and have the same ranking across models. While distributions of PCM are reasonable for models 1 to 4, model 5 implies PCM that are greater than 1 (and thus negative marginal cost) for $32$ percent of observations. We rule out model 5 based on the figure alone. However, discriminating between models 1 to 4 requires our more rigorous procedure.\footnote{Including model 5 in our testing procedure does not change our results as it is always rejected.}\looseness=-1
Instruments: Instruments must first be exogenous for testing. Following bh14, several sources of variation may be used to construct exogenous instruments. These include: (i) both observed and unobserved characteristics of other products, (ii) own observed product characteristics (excluded from cost), (iii) the number of other firms and products, (iv) rival cost shifters, and (v) market level demographics. Instruments must also be relevant for testing. Lemma (ref) shows that differences in predicted markups across models distinguish conduct. To distinguish models 1 and 2 we thus need to differentially move downstream markups, while to distinguish 1, 3, 4, and 5 we need to differentially move upstream markups. Theoretically, for every pair of models, variation in sources (i)--(v) move upstream and downstream markups for at least one model, making them plausibly relevant.
We then need to form instruments from the exogenous and plausibly relevant sources of variation. We consider four instrument choices constructed from these sources that are standard in estimating demand gh19 and have been used in testing conduct bcs20.
We first leverage sources of variation (i)-(iii) by considering two sets of BLP instruments: the instruments proposed in blp95 (BLP95) and the differentiation instruments proposed in gh19 (DIFF). These instruments have been shown to perform well in applications of demand estimation. As they leverage variation in product charactetistics and move markups, they are appropriate choices in our setting. For product-market $jt$, let $O_{jt}$ be the set of products other than $j$ sold by the firm that produces $j$, and let $R_{jt}$ be the set of products produced by rival firms. For product characteristics $\boldsymbol{x}$, the instruments are:\looseness=-1 \[\boldsymbol{z}^BLP95_{jt}=\left[
\right]\] \[\boldsymbol{z}^DIFF_{jt}=\left[
\right]\] where $\boldsymbol{d}_{jkt}\equiv \boldsymbol{x}_{kt}- \boldsymbol{x}_{jt}$ and $sd(\boldsymbol{d})$ is the vector of standard deviations of the pairwise differences across markets for each characteristic.\footnote{Following c12, c17, and bcs20, we perform RV testing with the leading principal components of each of the sets of instruments. We choose the number of principal components corresponding to 95% of the total variance, yielding two BLP95 instruments and five DIFF instruments. The results below do not qualitatively depend on our choice of principal components.} To form instruments from rival cost shifters, we average transportation costs of rival firms' products (COST). Finally, gh19 suggests that variation in demographics can be leveraged for demand estimation by interacting market level moments with product characteristics. Given the heterogeneity in consumer preferences in our demand system, we interact mean income with light and mean age with size and light to construct our fourth set of instruments (DEMO).\looseness = -1
AR Test: We first perform the AR test with the BLP95 instruments. Table (ref) reports test statistics obtained for each pair of models. The results illustrate Propositions (ref) and (ref): AR rejects all models when testing with a large sample.
RV Test: We perform RV tests using BLP95, DIFF, COST, and DEMO instruments. Following Section (ref), we keep the instrument sets separate and construct model confidence sets using the procedure of hln11. We report the results in Table (ref).\footnote{The results are computed with the Python package pyRVtest available on GitHub dmss_code. The package, portable to a wide range of applications, seamlessly integrates with PyBLP cg19 to import results of demand estimation. A researcher needs only specify the models they want to test, the instruments and the cost shifters, and the package outputs all the information in Table (ref). The variance estimators developed in this paper enable fast computation of all elements in that table, even in large datasets and with flexible demand systems.} \looseness=-1
To ease the reader in to the results, we begin by explaining Panel A in depth. The first three columns give the pairwise RV test statistics for all pairs of models. For each pair, a value above $1.96$ indicates rejection of the null of equal fit in favor of the column model. Instead, a value below $-1.96$ corresponds to rejection in favor of the row model. The second three columns give all the pairwise $F$-statistics. Finally, the last column reports the MCS $p$-values. In Panel A, the MCS contains only model 2 corresponding to zero retail margins; the MCS $p$-value for the other three models is below 0.05, our chosen level. The $F$-statistics show that the BLP95 instruments are strong for testing: there are no size distortions above 0.025 with two instruments and the critical value for target maximal power of 0.95 is 18.9, as seen in Table (ref). If a researcher precommitted to the BLP95 instruments for testing, Panel A shows the results that would obtain.
Panels B-D report test results in the same format as Panel A for the other three sets of instruments. Results vary markedly across panels. While the MCS in Panel B contains only model 2, coinciding with the MCS in Panel A, the MCS in Panels C and D contain additional models. Inspection of the pairwise $F$-statistics shows that the failure to reject models in Panels C and D is due to the COST and DIFF instruments having low power. For instance, the five DIFF instruments in Panel D, while strong for size, are weak for testing at a target maximal power of 0.50 for all pairs of models, as the critical value is 6.2. Given that the diagnostic is based on maximal power, the realized power could be considerably lower than 0.5. Similarly, the single rival COST instrument in Panel C is weak for testing: for several pairs of models, the instrument is weak for size at a target of 0.125 and weak for power at a target of 0.50. Given the null is not rejected in these cases, power is the salient concern. \looseness=-1
The diagnostic enhances the interpretation of the RV test results in Table (ref). Had the researcher precommitted to DIFF or COST instruments, the conclusions one could draw on firm conduct would not be informative. Because it is hard, in this context, to precommit to any one set of instruments, we suggest the researcher accumulates evidence across instrument sets.\footnote{Alternatively, we could pool all instrument sets. Appendix (ref) shows that doing so dilutes instrument power, resulting in lower $F$-statistics and a larger MCS.} To do so, we implement the procedure in Figure (ref). In step 1, we check for conflicting evidence. As all MCS for each set of instruments are nested, there is no conflicting evidence in this setting. Thus, in step 2 we initially set $M^*=\{2\},$ which is the smallest MCS arising from BLP95 and DEMO instruments. As DIFF IVs and COST IVs are not strong for all pairs of models, there is no addition to be made to $M^*$. Thus, the evidence accumulated across the four sets of instruments supports concluding for model 2.\looseness=-1
Main Findings: This application highlights the practical importance of allowing for misspecification and degeneracy when testing conduct. First, by formulating hypotheses to perform model selection, RV offers interpretable results in the presence of misspecification. Instead, AR rejects all models in our large sample. Second, instruments are weak in a standard testing environment, affecting inference. When RV is run with the DIFF or COST instruments, it has little to no power in this application. Thus, assuming at least one of the models is testable is not innocuous. Our diagnostic distinguishes between weak and strong instruments, allowing the researcher to assess whether inference is valid. Finally, by not having to precommit to a choice of instruments, our procedure for accumulating evidence allows researchers to draw sharp conclusions on firm conduct in this setting.\looseness=-1
In addition to illustrating our results, this application speaks to how prices are set in consumer packaged goods industries. Unlike v07 who concludes for the zero wholesale margin model, only a model where manufacturers set retail prices is supported by our testing procedure. Our finding is important for the broader literature studying conduct in markets for consumer packaged goods as it supports the common assumption that manufacturers set retail prices n01,mw17.\looseness=-1
In this paper, we discuss inference in an empirical environment encountered often by IO economists: testing models of firm conduct. Starting from the falsifiable restriction in bh14, we study the effect of formulating hypotheses and choosing instruments on inference. Formulating hypotheses to perform model selection allows the researcher to learn the true nature of firm conduct in the presence of misspecification. Alternative approaches based on model assessment instead will reject the true model of conduct if noise is sufficiently low. Given that misspecification is likely in practice, we focus on the RV test.\looseness=-1
However, the RV test suffers from degeneracy when instruments are weak for testing. Based on this characterization, we outline the inferential problems caused by degeneracy and provide a diagnostic. The diagnostic relies on an $F$-statistic which is easy to compute, and can inform the researcher about the presence of size distortions or low maximal power. We also show how to aggregate evidence across different sets of instruments, while using the $F$-statistic to draw sharp conclusions.\looseness=-1
An empirical application testing vertical models of conduct v07 highlights the importance of our results. We find that AR rejects all models of conduct. This illustrates the importance of allowing for misspecification and adopting a model selection approach. Four sets of exogenous and plausibly relevant instruments exist in this setting. Two of these are weak, as diagnosed by our $F$-statistic. Adopting our procedure for accumulating evidence across RV tests with separate instrument sets, we conclude for a single model in which manufacturers set retail prices.