Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,474 characters · 11 sections · 41 citation commands
Testing and Identifying Substitution and Complementarity Patterns
Substitution/complementarity relationships between goods have been studied in various applications such as online news versus print newspapers, digital books versus traditional books, and cigarettes versus e-cigarettes. The relationship plays a crucial role in consumers' decisions; therefore, understanding substitution patterns is important for predicting demand for a good and analyzing the welfare effects of, for example, a merger of two companies or the introduction of a new good \citep*{petrin2002, goolsbee2004, gentzkow2007}.
The standard multinomial choice models typically assume that consumers can buy only one good at a time, which rules out complementarity by assumption. However, even some goods traditionally perceived as substitutes are shown to be complements in different contexts. For example, zhao2019 suggests that cigarettes and e-cigarettes could be complements, and grzybowski2008 demonstrate the complementarity between telephone calls and messages.
With these motivations in mind, I propose a semiparametric panel multinomial choice model with fixed effects to study substitution/complementarity patterns. This model allows consumers to purchase two goods simultaneously, accommodating the possibility that the two goods are either substitutes or complements. The model also permits heterogeneous complementarity relationships through observed characteristics.
Identifying substitution/complementarity patterns with bundles involves several challenges. First, the demand for one good includes consumers who buy this good alone and those who buy a bundle. Therefore, a large demand for one good could come from consumers' high utility for this good, or its complementarity with another good, or both. We need to disentangle the two sources to identify the complementarity relationship. Second, the purchase of two goods together may be due to either the goods' complementarity or the unobserved correlation between consumers' preferences over the two goods. For example, consumers may buy a variety of organic goods because of their preferences over organic goods instead of the complementarity between these goods. Distinguishing the complementarity relationship and the correlation between consumers' tastes for goods could be challenging since they are both unobserved and can affect consumers' decisions simultaneously.
To tackle these challenges, my paper exploits a conditional stationarity assumption about preference shocks over time, which enables us to use intertemporal variation in conditional choice probabilities for identification. I first provide an approach to test the substitution/complementarity relationship between goods and then derive the sharp identified set for the model parameters.
The testing approach exploits the relationship between the demand for two goods with the covariate indices of those goods under a substitution or complementarity relationship between goods. This approach does not require the estimation of any model parameter. Instead, it constructs conditional moment inequalities that only depend on observed variables. Therefore, testing complementarity can be conducted by directly testing these moment inequalities.
To derive the sharp identified set, the paper uses intertemporal comparisons of conditional choice probabilities to obtain identifying restrictions on the model parameters. The result not only identifies the sign of the complementarity but also bound its value. The sharpness of the identification results is established, indicating that all available information from the data for the model parameters has been exhausted. Additionally, the paper shows that under large support conditions for the covariates and a linear specification of the complementarity, point identification can be achieved.
For estimation and inference, I characterize conditional moment inequalities based on the identification results and apply the approach in shi2018. In Monte Carlo simulations, I compare the finite sample performance of the method in this paper to that of a parametric method, which assumes a parametric distribution over the error terms and a linear model for fixed effects. The simulation results show that the semiparametric method in this paper has the advantage of performing robustly over different DGP designs, whereas the parametric estimator is vulnerable to misspecifications.
I explore several extensions to the model. Firstly, I consider the possibility of unobserved heterogeneity in the complementarity relationship, which allows for differences in complementarity across individuals that are not captured by covariates. By exploiting variation in demand for two goods, I characterize partial identification results for the fraction of people for whom the two goods are complements. In the online Appendix, I also extend the model to include more than two goods and establish partial identification results for model parameters under additional assumptions. Additionally, I extend my approach to incorporate nonseparable utility functions as well as cross-sectional models.
This paper contributes to the literature studying substitution/complementarity patterns in a discrete choice model with bundles. gentzkow2007 develops a discrete choice model that allows for complementarity to study the substitution relationship between online news and print newspapers. His paper is flexible about substitution patterns and also admits correlations between error terms across choices. Compared to gentzkow2007, my paper does not reply on parametric assumptions, and also allows for flexible dependence structures between covariates and fixed effects. dunker2015 and iaria2020 allow for endogeneity and provide identification results of models with bundles by extending the classic BLP approach in berry1995. Their methods use demand inversion and parametric distributions over error terms, and they address endogeneity using instrumental variables. monardo2021 studies a more flexible model for the inverse demand and uses an instrument to construct moment conditions. My paper mainly exploits intertemporal variation in panel data to derive sharp identification results and the method does not require demand inversion or instrumental variables.
There are several papers that allow for unknown distributions of error terms to study substitution patterns. fox2017 study semiparametric identification of a discrete choice model with bundles under a large support assumption and exogenous covariates. allen2022 allow for unobserved heterogeneity in the complementarity and provide partial identification for the fraction of people for whom the two goods are complements. My paper focuses on heterogeneous complementarity through observed covariates and provide sharp identification with bounded support of covariates. In an extension of this paper, I allow for unobserved heterogeneous complementarity while relaxing the exogenous covariates assumption and the exclusion restriction in allen2022.
My paper is also related to a large body of literature on panel multinomial choice models with fixed effects. chamberlain1980 provides a conditional fixed effect logit estimator for the panel multinomial choice model under a logistic distribution over disturbances. manski1987 as well as honore2002 relax the logistic distribution assumption and study semiparametric identification of a binary choice model. manski1987 uses a maximum score approach that relies on a group stationarity assumption, and honore2002 exploit the idea of a special regressor to identify the panel binary choice model. honore2020, khan2020, and dobronyi2021 study identification of dynamic binary choice models.
pakes2019 and \citet*{shi2018} study a panel multinomial choice model. pakes2019 derive sharp identification of the model by characterizing conditional moment inequalities, while \citet*{shi2018} use cyclic monotonicity for identification and estimation. khan2021 provide inference methods for multinomial choice models. gao2020 relax the separable utility function assumption in the previous papers and study a class of nonseparable utility functions. The aforementioned papers focus on the identification of own price coefficients rather than substitution patterns between different goods. My paper builds on this literature to allow for bundles in the panel multinomial choice model and characterizes sharp identification for substitution patterns.
The rest of this paper is organized as follows. Section (ref) introduces the panel multinomial choice model with bundles. Section (ref) provides testable implications for the complementarity between two goods. Section (ref) characterizes the sharp identified set for the model parameters and provides sufficient conditions for point identification. Section (ref) develops conditional moment inequalities, and Section (ref) examines the finite sample performance via Monte Carlo simulations. Section (ref) studies an extension. Section (ref) concludes.
This section presents a panel multinomial choice model allowing for bundles. Consider a short-panel structure: let $i\in \mathcal{I}$ denote consumers and $t\leq T$ denote time periods where the length of the panel $T\geq 2$ is fixed. Since this paper focuses on substitution patterns between two goods, I consider the case of two goods: $\{A,B\}$.\footnote{Appendix (ref) studies the case of more than two goods and establishes partial identification results under additional assumptions.} Instead of assuming that consumers can buy either only good $A$ or only good $B$, this model allows consumers to purchase goods $A$ and $B$ simultaneously and focuses on identifying the complementarity between goods.
The choice set for consumers is $\mathcal{C}=\{A, B, AB, O\}$, where $A$ (or $B$) denotes purchasing only good $A$ (or $B$), $AB$ denotes purchasing $A$ and $B$ simultaneously, and $O$ denotes the outside option. I assume that consumers buy at most one unit of each good, and they select the choice yielding the highest utility in their choice set.
Let $u_{ijt}$ denote consumer $i$'s utility from consuming choice $j\in \mathcal{C}$ at time $t$. Following gentzkow2007, let $\Gamma_{it}=(u_{iABt}-u_{iAt})-(u_{iBt}-u_{iOt})$ denote the incremental utility of consuming good $B$ when good $A$ is also consumed. The utility $u_{ijt}$ is specified as\footnote{In Appendix (ref), I discuss how this utility specification connects to a conventional discrete choice model. In Appendix (ref), I also investigate a class of nonseparable utility functions.}
Here $X_{ijt}\in \mathbbm{R}^{d_x}$ denotes a vector of observed characteristics, which may include consumer $i$'s characteristics (e.g., income), product $j$'s characteristics (e.g., price), and the interaction terms between them; $\alpha_{ij}\in \mathbbm{R}$ denotes an unobserved individual-specific fixed effect for product $j$ that does not change over time, such as consumers' loyalty to a brand; $\epsilon_{ijt}\in \mathbbm{R}$ denotes an unobserved and time-varying shock that affects consumers' utility over time; and $\beta_0\in \mathbbm{R}^{d_x}$ is a finite-dimensional unknown parameter vector. Without loss of generality, we normalize the utility of the outside option to zero so the utilities of the remaining choices are defined relative to the utility of the outside option.
The sign of the incremental utility $\Gamma_{it}$ represents the complementarity relationship between the goods. Two goods are considered complements if their combined utility is greater than the sum of their individual utilities ($\Gamma_{it}> 0$), and substitutes if the reverse is true ($\Gamma_{it}< 0$). This definition is equivalent to an alternative definition of substitution patterns using aggregate demand, as discussed in Section (ref). In this model, $\Gamma_{it}$ can be either positive, negative, or zero, which allows for the possibility that two goods can be either substitutes or complements. I provide identification results for both the coefficient $\beta_0$ and the complementarity $\Gamma_{it}$.
In addition to the covariate $X_{ijt}$, I assume that consumer $i$'s choice at time $t$ is observed which is denoted as $Y_{it}\in \mathcal{C}$. Consumers select the choice with the highest utility, implying
To simplify analysis, I assume that tie outcomes between choices happen with a zero probability, eliminating the need to account for such situations. This assumption is satisfied as long as one of the error terms is continuously distributed. Even if ties happen with nonzero probability, all identification results hold as long as the tie-breaking rule is fixed over time.
The main objective of this paper is to identify the substitution/complementarity relationship and the coefficient $\beta_0$ from consumers' choices $Y_{it}$ and covariates $X_{ijt}$.
Next, I introduce some assumptions about the model.
Assumption (ref) requires the complementarity to only depend on observed covariates $Z_{i}$. The function $\Gamma$ is flexible and can be unknown to researchers, which could admit rich complementarity patterns. The covariate $Z_{i}$ may include consumers' characteristics such as income and age so that Assumption (ref) allows for heterogeneous complementarity relationships through observed characteristics. I assume the covariate $Z_{i}$ to be fixed over time for any $t$. If the covariate changes over time, then the same analysis can be conducted conditional on the same value of the covariate over time: $Z_{is}=Z_{it}=z$.
One restriction of Assumption (ref) is that it excludes unobserved heterogeneity in the complementarity and assumes the same complementarity relationship for individuals with the same characteristic $Z_i=z$. This structure allows us to not only identify the sign of the complementarity $\Gamma(z)$ but also bound the magnitude of the complementarity $\Gamma(z)$ given $Z_i=z$. In Section (ref), I consider an extension that allows for unobserved heterogeneity in the complementarity $\Gamma_{it}$, where we can only obtain weaker results and partially identify the distribution of the sign of $\Gamma_{it}$.
The exclusion assumption requires that there exists one variable that only influences the utility for good $A$ or $B$ but not the complementarity between the two goods. One example of this variable is the price of good $A$ or $B$, which affects the utility of a single good but may not influence the complementarity between the two goods. The sign of the coefficient for $X^{*}_{it}$ can still be unknown to researchers. Moreover, this assumption does not restrict the covariate $Z_i$; any variable affecting the complementarity is allowed to influence the utility of a single good.
The last assumption is the stationarity condition for the distribution of the unobserved shocks. Let $X_{it}=(X_{iAt}, X_{iBt}), \alpha_i=(\alpha_{iA}, \alpha_{iB})$, and $\epsilon_{it}=(\epsilon_{iAt}, \epsilon_{iBt})$ collect covariates, fixed effects, and error terms of the two goods.
This assumption is a multinomial extension of the conditional homogeneity assumption in manski1987. It is commonly used in the literature on panel multinomial choice models, including pakes2019 and shi2018, who study identification of the coefficient $\beta_0$ under this assumption. Assumption (ref) restricts the conditional distribution of $\epsilon_{it}$ to be stationary over time, but it allows the error term $\epsilon_{it}$ to be dependent across choices and over time. In addition, it does not impose any distributional restrictions on the unobserved term $\epsilon_{it}$. Therefore, the standard logit/probit models and i.i.d. assumption of the error term can be nested in Assumption (ref).
One crucial feature of Assumption (ref) is that it can accommodate endogenous covariates by allowing for arbitrary dependence structures between the fixed effects $\alpha_{ij}$ and the covariates $X_{it}$. Endogeneity is important in demand estimation because the price of a product could potentially depend on the unobserved heterogeneity of the product, such as the quality of the product or consumers' taste for the product. chesher2013 and berry2014identification provide more detailed discussions about the importance of allowing endogeneity in demand estimation.
Of course, Assumption (ref) does impose some restrictions. For example, it excludes some dependence structures between $\epsilon_{it}$ and the covariate $X_{it}$. Consider that if $\epsilon_{it}$ only depends on $X_{it}$ for any period $t$, then $\epsilon_{is}$ may have a different distribution than $\epsilon_{it}$ when $X_{is}$ and $X_{it}$ take different values. However, some dependence structures between $\epsilon_{it}$ and $X_{it}$ are still allowed in Assumption (ref): for example, if $\epsilon_{it}$ depends on covariates in a time-invariant form such as $\frac{1}{T}\sum_{t=1}^T X_{it}'\beta_0$, Assumption (ref) still holds.
This section discusses the relationship between two different definitions of substitution patterns. My paper uses the sign of $\Gamma(z)$ to represent the substitution relationship between two goods, which captures the incremental utility from consuming the bundle compared to consuming a single good. As shown in Lemma (ref), this turns out to be equivalent to an alternative definition that was previously used in the literature.
The alternative definition of substitution patterns centers on how the demand for good $A$ (or $B$) is affected by an increase in the price of good $B$ (or $A$). The two goods are substitutes if the demand for good $A$ increases, complements if it decreases, and independent if the demand does not change. Let $p_{jt}$ denote the price of good $j$ whose coefficient is nonzero, and let $\tilde{X}_{it}=X_{it} \setminus \{p_{Bt}\} $ denote the remaining covariates in $X_{it}$ excluding the price of good $B$. I fix all other covariates $\tilde{X}_{is}=\tilde{X}_{it}=\tilde{x}$ over time and compare the conditional demand for good $A$ under different prices $p_{Bs}\neq p_{Bt}$ of good $B$. The demand for good $A$ comes from two sources: individuals who purchase only good $A$ and those who purchase the bundle $AB$. Let $D_{\ell}=\{\ell, AB\}$ collect all choices containing good $\ell \in\{A, B\}$. Let $\text{sign}(x)=\mathbbm{1}\{x>0\}-\mathbbm{1}\{x<0\}$ denote the sign function.
The substitution pattern $s_{AB}(z)$ conditional on the covariate $Z_i=z$ is defined as
The value of $s_{AB}(z) \in\{-1,0,1\}$ represents the complementarity relationship between goods $A$ and $B$. For consumers with the covariate $Z_i=z$, the two goods are substitutes if $s_{AB}(z) =1$, independent if $s_{AB}(z) =0$, and complements if $s_{AB}(z) =-1$. Under the aforementioned Assumptions (ref)-(ref), the value of $s_{AB}(z)$ is the same defined by any two periods $s\neq t$, and it is independent of other variables except $z$ since the complementarity term $\Gamma(z)$ depends only on $z$; therefore, $s_{AB}(z)$ is written as a function of only $z$.
It is often difficult to study substitution patterns directly from the definition of $s_{AB}(z)$. The term $s_{AB}(z)$ uses only variation in prices and requires the fixing of all other covariates. This may not be feasible since the other covariates may change simultaneously with prices or the covariates may include time-varying variables such as time dummies. In addition, variation in prices may not be available in some scenarios in which the prices of products are constant over time. Moreover, as the definition $s_{AB}(z)$ involves conditional choice probabilities, directly estimating $s_{AB}(z)$ may perform poorly, especially when the dimension of covariates is large.
The next lemma establishes the relationship between $s_{AB}(z)$ and the incremental utility $\Gamma(z)$.
Lemma (ref) shows that $s_{AB}(z)$ always has the opposite sign of the incremental utility term $\Gamma(z)$. This lemma implies that the sign of $s_{AB}(z)$ can be learned if the sign of the incremental utility $\Gamma(z)$ is identified. Therefore, identifying the complementarity term $\Gamma(z)$ is sufficient for studying substitution patterns defined by $s_{AB}(z)$.
To illustrate the intuition of Lemma (ref), I will focus on the case in which the incremental utility is positive, i.e., $\Gamma(z)> 0$. In this case, consumers with a small utility from a single good will still purchase the bundle since they can obtain additional positive utility from consuming the two goods together. When the price of good $B$ increases such that the utility of the bundle decreases, some consumers will switch from buying the bundle to buying the outside option since their utility from a single good is small. Therefore, the demand for good $A$ decreases, which implies $s_{AB}(z)\leq 0$.
A similar result to Lemma (ref) is shown in gentzkow2007 with cross-sectional data. The difference is that gentzkow2007 requires an independence condition between unobserved error terms and observed covariates, so his results do not apply to the case with endogenous covariates. My paper leverages the stationarity assumption, which allows for endogenous covariates. Therefore, Lemma (ref) shows that even with endogenous covariates, the relationship between the two definitions of substitution patterns ($\Gamma(z)s_{AB}(z)\leq 0$) still holds by exploiting intertemporal variation in covariates.
Before introducing the identification results, this section develops a method to test the complementarity relationship. This approach does not involve the estimation of any model parameters and directly constructs testable implications that only depend on observed variables.
I will focus on testing the complementarity between the two goods; the analysis for testing the substitutability is similar, which can be found in Appendix (ref). For consumers with covariate $Z_i=z$, testing the complementarity relationship between goods is equivalent to testing $\Gamma(z)\geq 0$. The null hypothesis $H_0$ and alternative hypothesis $H_1$ are given as
The main idea of testing the complementarity is that changing the covariate index $X_{iAt}'\beta_0$ of good $A$ will affect the demand for good $B$ in different directions, depending on their substitution/complementarity relationships. If the two goods are complements (under $H_0$), an increase in the covariate index of good $A$ will encourage consumers to purchase the bundle and thus increase the demand for good $B$. However, if the two goods are substitutes (under $H_1$), people will switch to choosing only good $A$ such that the demand for good $B$ decreases. Therefore, we can reject $H_0$ when a decrease in the demand for good $B$ is observed in this scenario.
However, the sign of the covariate index $X_{ijt}'\beta_0$ is unknown since it involves the unknown parameter $\beta_0$. Therefore, the first step of testing the complementarity is to learn the sign of covariate indices of two goods using observed variables $(X_{it}, Z_i, Y_{it})$.
Let $P_{t}(\{j\} \mid x_s, x_t, z)$ denote the probability of choosing $j\in \mathcal{C}$ at time $t$ conditional on covariates $(X_{is}, X_{it})=(x_s, x_t)$ and $Z_i=z$, given as
The conditional choice probability depends only on observed variables, which can be identified from data. Since the conditional distribution of error terms is the same over time (as per Assumption (ref)), any variation in choice probabilities can only arise from changes in covariate indices. From the variation in observed choice probabilities, we can infer the signs of the covariate indices and construct testable implications.
Let $\xi_{s, t}^{1}(x_{s}, x_{t}\mid z)$ denote an indicator for increasing probabilities of all choices $j\in \{A, B, AB\}$ conditional on $(X_{is}, X_{it})=(x_s, x_t)$ and given $Z_i=z$:
When an increase in conditional probabilities of all choices $j\in \{A, B, AB\}$ is observed, it implies that the covariate indices of both goods increase:
The above relationship exploits variation in conditional probabilities of multiple choices to identify the signs of covariate indices of both goods. If we observe an increase in the probability of only one choice, such as good A, it could be due to either an improvement in good A or a decline in good B, and we are unable to distinguish between the two scenarios. In contrast, if we observe an increase in the probabilities of all choices $\{A, B, AB\}$, it can only occur when both goods improve, enabling us to determine the signs of the variation in the covariate indices for both goods.
Now we are ready to establish testable conditions for the null hypothesis. With the null hypothesis of the two goods being complements, an increase in the covariate indices of both goods would result in an increase in demand for the two goods. This relationship generates testable implications for the null hypothesis. If a decrease in demand for either good is observed, it suggests that the two goods are substitutes and the null hypothesis is rejected.
Recall that $D_{\ell}=\{\ell, AB\}$ for $\ell \in \{A, B\}$, and the conditional demand of good $\ell$ is expressed as \[\Pr(Y_{it}\in D_{\ell}\mid x_{s}, x_{t}, z)=E[\mathbbm{1}\{Y_{it}\in D_{\ell}\}\mid x_{s}, x_{t}, z].\]
Proposition (ref) provides testable implications for the null hypothesis $H_0$ by characterizing conditional moment restrictions that depend only on observed variables. Therefore, the null hypothesis can be tested by directly testing the above conditional moment inequalities. The conditions in the proposition apply to any two periods and any pair of covariates, allowing us to test for complementarity by exploiting the variation in covariates over time.
Besides testing the sign of $\Gamma(z)$, we may also be interested in the value of the complementarity $\Gamma(z)$ as well as the utility coefficient $\beta_0$. This section establishes sharp identification results for the parameter $\theta_0=(\beta_0, \Gamma)$. The observed variables include the covariates $(X_{it}, Z_i)$ and consumers' choices $Y_{it}\in \mathcal{C}$ in each period. Since only the relative utility between choices matters for consumers' decisions, the parameter can be only identified up to a constant. Therefore, I normalize the first element $\theta^1$ of the parameter $\theta$ to be one for the following analysis: $\Theta=\{\theta: |\theta^{1}|=1\}$.
For any subset $K\subset \mathcal{C}$, let $P_{t}(K\mid x_s, x_t, z)$ denote the conditional probability of $Y_{it}\in K$ at time $t$ given covariates $(X_{is}, X_{it})=(x_s, x_t)$ and $Z_i=z$; that is, the probability of existing one choice in the set $K$ generating the highest utility among all choices:
When $K=\{j\}$ is a singleton, this reduces to the conditional choice probability (CCP) of $j$. The main idea of the identification analysis is to derive identifying restrictions of the true parameter $\theta_0$ from intertemporal variation in conditional choice probabilities across two different periods. All parameters satisfying those identifying restrictions form an identified set for the true parameter.
Let $\delta_{\ell t}=x_{\ell t}'\beta_0$ denote the covariate index for good $\ell \in \{A, B\}$ given $X_{i\ell t}=x_{\ell t}$. Let $\delta_{ABt}=\delta_{At}+\delta_{Bt} $ and $\delta_{Ot}=0$ denote the covariate indices for bundle $AB$ and the outside option, respectively. Let $\Delta_{s,t} \delta_{j}=\delta_{js}-\delta_{jt}$ denote the change in the covariate index for choice $j\in \mathcal{C}$ between periods $s$ and $t$.
In models assuming that consumers can buy only one good at a time, two goods can only be substitutes. Since the complementarity relationship is known, the only unknown factor affecting conditional choice probabilities is variation in covariate indices of all choices. My paper allows for the possibility that two goods can be either substitutes ($\Gamma(z)<0$) or complements ($\Gamma(z)>0$) and the complementarity relationship is unknown. Therefore, two unknown sources are affecting conditional choice probabilities in this paper: one is changes in covariate indices and the other is the complementarity relationship between the two goods. Distinguishing between the two sources poses a challenge for the identification analysis.
The following proposition characterizes the identifying restrictions for the parameter $\theta_0$ under Assumptions (ref)-(ref). Let $C_1 \lor C_2$ mean that either condition $C_1$ or $C_2$ holds or both hold, and let $C_1 \land C_2$ mean that both $C_1$ and $C_2$ hold.
Proposition (ref) characterizes identification restrictions for the parameter $\theta_0$ from comparisons of conditional choice probabilities across two periods that can be identified from data. The identifying restrictions for $\theta_0$ in Proposition (ref) are free from unobserved terms, such as the fixed effects $\alpha_i$ and the error term $\epsilon_{it}$. As the results hold for any fixed length $T$ of panel data, we can utilize variation in conditional choice probabilities for any two periods to identify $\theta_0$ and take intersections of the identified sets. Later I will formulate conditional moment inequalities based on the identifying restrictions in Proposition (ref), which can be used to conduct estimation and inference for the parameter $\theta_0$.
Condition (ref) in Proposition (ref) contains the identifying restrictions for the coefficient $\beta_0$. The intuition of this result is described as follows: if the conditional probability of selecting choice $j$ increases, then it is impossible that choice $j$ becomes worse (in terms of the covariate index) compared to all other choices. Therefore, it can be inferred that the covariate index for choice $j$ should increase relative to at least one other choice.
The remaining two conditions in Proposition (ref) provide novel identification results for the complementarity parameter $\Gamma(z)$. Condition (ref) identifies the sign of the complementarity $\Gamma(z)$ and bounds its absolute value by comparing the conditional demand of the two goods over time. Condition (ref) establishes both lower and upper bounds for the complementarity $\Gamma(z)$ using the sum of probabilities of two different choices over two periods. Next, I will explain the intuition behind the two conditions.
Condition (ref) mainly exploits the idea that the relationship between the covariate index of one good and the demand for the other good is different under different complementarity relationships between the two goods. When the two goods are complements, an increase in the covariate index of good $A$ will incentivize consumers to buy the bundle $AB$, resulting in an increase in demand for good $B$. So if a decrease in demand for either good is observed, the two goods must be substitutes $(\Gamma(z)< 0)$. Similarly, when the covariate index for good $A$ decreases and the covariate index for good $B$ increases along with observing an increase in the demand for good $A$, the two goods are identified as complements $(\Gamma(z)>0)$.
Condition (ref) in Proposition (ref) can bound the value of the complementarity $\Gamma(z)$ by looking at the sum of conditional probabilities of two choices over time. For example, when the sum of the conditional probabilities of buying two goods together and neither good is large, then a lower bound for the complementarity $\Gamma(z)$ is established. The intuition is that when good $A$ becomes less attractive, consumers will switch from buying two goods together to buying neither good if the two goods are complements, but they will switch to buying only good $B$ if the two goods are substitutes. Therefore, if a large probability of buying two goods together and neither good is observed, we can infer the complementarity between the two goods and provide a lower bound for $\Gamma(z)$. Similarly, if the sum of the conditional probabilities of buying a single good is large, then the upper bound for the value of $\Gamma(z)$ can be obtained.
The identifying restrictions (ref)-(ref) in Proposition (ref) characterize an identified set $\Theta_I$ for $\theta_0$, which is defined as
Theorem (ref) shows that conditions (ref)-(ref) have exhausted all available information from the observed data for the parameter $\theta_0$. The proof of the sharpness is conducted through direct construction. For any parameter in the identified set $\Theta_I$, I construct an underlying DGP that satisfies Assumptions (ref)-(ref) and matches the observed conditional choice probabilities, which shows the sharpness of the identified set $\Theta_I$. The main challenge of the construction is that the unknown DGP involves conditional distributions over the whole space of unobserved error terms, which are infinite dimensional.
I address this difficulty by constructing “choice sets," which are collections of unobserved terms such that a single choice is selected conditional on covariates. It is sufficient to focus on constructing the distributions on the choice sets because their distributions determine the observed choice probabilities. The number of choice sets is finite due to the finite number of choices; accordingly, I only need to assign probabilities on the finite number of sets, which simplifies the construction. Then the paper shows that for any parameter in the identified set $\Theta_I$, there exists a conditional distribution on the choice sets that satisfies the assumptions and generates the observed choice probabilities. The construction of the probabilities on the choice sets depends on the sign of the complementarity $\Gamma(z)$ as well as the covariate index $\Delta_{s, t} \delta_j$, which is discussed in detail in Appendix (ref).
Some discussion of Theorem (ref) is in order. First, similar to pakes2019, the identification analysis in Proposition (ref) and the sharpness result for $\Theta_I$ can be extended to a more general model, $u_{i\ell t}=f(X_{i\ell t}, \beta_0)+g(\alpha_{i\ell}, \epsilon_{i\ell t})$, where $f$ is a known function up to a finite-dimensional parameter $\beta_0$ and $g$ is a function that can be unknown to econometrician. This utility function allows for infinite dimensional fixed effects and idiosyncratic shocks as well as admits arbitrary interactions between them. Appendix (ref) also discusses a class of utility functions that can be nonseparable between covariates and unobserved fixed effects/error terms. Second, the identified set $\Theta_I$ employs only marginal choice probabilities at each period yet it is shown to be sharp. Therefore, joint choice probabilities over different periods do not provide any additional information for the parameter $\theta_0$. Moreover, the sharpness result exploits the correlations of error terms across choices and over time. So if one is willing to impose additional assumptions on the dependence structure (e.g., i.i.d.), then the identified set $\Theta_I$ could be further tightened.
This section studies the conditions under which the model parameters can be point identified up to scale. The analysis depends on the specification of the additional utility term $\Gamma(Z_i)$. I focus on point identification under a linear specification of the complementarity: $\Gamma(Z_i)=Z_i'\gamma_0$.
For simplicity of notation, I consider a two-period model ($T=2$) to illustrate the idea.\footnote{With more than two periods, it is straightforward that point identification can be achieved when there exists a pair of periods that satisfy Assumptions (ref)-(ref).} Let $\Delta X_{i\ell }= X_{i\ell 2}-X_{i\ell 1}$ denote the change in observed covariates for consumer $i$ and good $\ell\in \{A, B\}$ over the two periods, and let $\Delta X_i=(\Delta X_{iA}, \Delta X_{iB})$ collect the changes in covariates for the two goods. I use a superscript $k$ to denote the $k$th element of a vector, e.g., $\Delta X_{iA}^{k}$ represents the $k$th element of the vector $\Delta X_{iA}$.
I first introduce sufficient conditions for point identification of the coefficient $\beta_0$. The coefficient $\beta_0$ can only be point identified up to scale since multiplying consumers' utilities of all choices by a positive constant will not change observed choices.
Assumption (ref) eliminates the uninformative scenario where conditional choice probabilities remain unchanged despite changes in the covariate indices over time. Assumption (ref) is a support condition on the covariate $\Delta X_i$. It requires at least one covariate for each good to have large support, while the support of the remaining covariates is unrestricted. The large support condition guarantees that there is sufficient variation in the covariate over time such that the true parameter can be distinguished from any other candidate parameters.
Under these assumptions, $\beta_0$ can be point identified by using the first identifying restriction (condition (ref)) in Proposition (ref). For any parameter $b$ such that $b\neq k\beta_0$, $\forall k>0$, the large support condition in Assumption (ref) implies that there exists one value $\Delta x_{\ell}$ of the covariate such that the covariate index $\Delta x_{\ell}' \beta$ has different signs under the true parameter $\beta_0$ and the candidate parameter $b$. The conditional choice probabilities then change in different directions under $\beta_0$ and $b$ so that the parameter $\beta_0$ is identified. For example, suppose that the covariate index satisfies $\Delta x_{\ell}' \beta_0 >0$ and $\Delta x_{\ell}' b <0$ for any $\ell \in \{A, B\}$. Then under Assumption (ref), the conditional choice probability of buying bundle $AB$ will strictly increase under the true parameter $\beta_0$, but strictly decrease under the parameter $b$. Therefore, $\beta_0$ is identified.
Since $\beta_0$ is point identified, the sign of the covariate index $\Delta X_{ij}'\beta_0$ is also identified. Next, I present the conditions for point identification (up to scale) of the complementarity parameter $\gamma_0$.
Similar to Assumption (ref), this assumption requires a large support restriction on the covariate $Z_i$. Based on condition (ref) in Proposition (ref), the sign of the complementarity $Z_i'\gamma_0$ is identified from intertemporal variation in the conditional demand for the two goods. Since the sign of $Z_i'\gamma_0$ is identified, the parameter $\gamma_0$ is also point identified up to scale under the large support assumption. The analysis for $\gamma_0$ is similar to the coefficient $\beta_0$. For any candidate parameter $\tilde{\gamma}$ such that $\tilde{\gamma}\neq k\gamma_0$, $\forall k>0$, Assumption (ref) ensures that there exists some value of the covariate $Z_i$ such that the sign of the complementarity $Z_i'\gamma$ is different under the true parameter $\gamma_0$ and the candidate parameter $\tilde{\gamma}$. Thus, the parameter $\gamma_0$ can be point identified.
The identified set $\Theta_I$ characterized by conditions (ref)-(ref) in Proposition (ref) is abstract and it is a challenging task to check whether every candidate parameter satisfies all of the identifying conditions. This section develops an alternative characterization of the identified set $\Theta_I$ by constructing conditional moment inequalities of the parameter. Based on this characterization, the literature has developed many methods to do inference for conditional moment inequalities (e.g., andrews2013, chernozhukov2013, and armstrong2015).
The identification conditions (ref)-(ref) in Proposition (ref) share a similar structure, which involves deriving restrictions for the parameter $\theta_0$ through intertemporal comparisons of conditional choice probabilities. I focus on the first condition (ref) in Proposition (ref) to describe the idea of constructing conditional moment inequalities. Let $W_{ist}=(X_{is}, X_{it}, Z_i)$ collect all of the covariates at the two periods $(s, t)$, and let $w_{st}=(x_{s}, x_t, z)$ denote one realization of the covariate $W_{ist}$.
Condition (ref) exploits comparisons of the conditional probability of a single choice $j\in \mathcal{C}$ to derive restrictions for the parameter. Let $\lambda_{s, t}^{j}(w_{st}, \theta)$ denote the indicator index of the identifying restriction in condition (ref), defined as
Condition (ref) derives the identifying restriction $\lambda_{s,t}^{j}$ from a positive variation in the conditional probability of selecting choice $j$ over time:
The contraposition of the above condition is presented as follows: if the identifying restriction $\lambda^{j}_{s,t}$ does not hold, then the variation in the conditional probability of selecting choice $j$ is nonpositive.
Plugging into the definition of the conditional choice probability $P_t(\{j\}\mid w_{st})=E[\mathbbm{1}\{Y_{it}=j \} \mid W_{ist}=w_{st}]$, the above condition leads to the following conditional moment inequality for any $w_{st}$,
The above conditional moment inequality holds since either the binary index holds $\lambda^{j}_{s,t}(w_{st}, \theta_0)=1$ so that the moment function $g_{s,t}^{j}$ is zero or the binary index does not hold $\lambda_{s,t}(w_{st}, \theta_0)=0$ implying that the function $g_{s,t}^{j}$ is nonpositive. I provide an equivalent characterization to condition (ref) using conditional moment inequalities. The characterization for conditions (ref)-(ref) can be constructed similarly.
Condition (ref) derives restrictions of the parameter from comparisons of the demand for good $\ell \in \{A, B\}$. The indicator $\lambda_{s,t}^{D_{\ell}}(w_{st}, \theta)$ of the identifying restriction in condition (ref) is defined as follows, let $\ell_{-1} \in \{A, B\}$ and $\ell_{-1} \neq \ell$,
From comparisons of the demand for good $\ell\in \{A, B\}$, the conditional moment inequality can be constructed as follows:
Condition (ref) derives lower and upper bounds for the complementarity $\Gamma(z)$ from the sum of conditional probabilities of two choices over two different periods. The binary indices of the identifying restrictions in condition (ref) are defined as
Similarly, the conditional moment inequalities are constructed as follows based on condition (ref) in Proposition (ref):
I have developed conditional moment inequalities that are equivalent to the identifying conditions (ref)-(ref) in Proposition (ref). Let $g_{s, t}=(\{g_{s,t}^{j}\}_{j\in \mathcal{C}}, g_{s,t}^{D_{A}}, g_{s, t}^{D_B}, g_{s,t}^{L}, g_{s, t}^{U})'$ denote a vector of all conditional moment functions. The identified set $\Theta_I$ is characterized by the set of parameters satisfying the conditional moment inequalities.
Proposition (ref) characterizes the identified set using conditional moment inequalities. For estimation and inference, we need to specific a parametric form for the function $\Gamma$ (such as a linear function) so that the dimension of the unknown parameter is finite. Then, the estimation and inference for the parameter can be conducted using methods in the literature developed for conditional moment inequalities.
This section will compare the method in this paper with a parametric method via Monte Carlo simulation. The parametric method imposes distributional assumptions over $\epsilon_{ijt}$ and a linear structure over $\alpha_{ij}$, which will be described in detail later. The simulation results demonstrate that misspecifications in either parametric distributions or dependence structures between covariates and fixed effects lead to misleading estimators for the complementarity parameter.
I study a linear specification of the complementarity $\Gamma(Z_i)=Z_i \gamma_0$, and look at a two period model $T=2$. Section (ref) has established point identification results under large support of covariates, so this section focuses on the case where the parameter is point identified. More simulation results are presented in Appendix (ref) regarding the performance of the point estimator under longer panels ($T\geq 2$) and the set estimator with bounded support of covariates.
I implement the criterion function approach in shi2018 for estimation. The criterion function can be developed as follows based on conditional moment inequalities in Proposition (ref):
where $||x||_1=\sum_j |x^j|$ for a vector $x=(x^1,...,x^j)'$.
Similar to shi2018, a two-step estimator is developed based on the above criterion function. The first step estimates the conditional choice probability $P_{t}(\{j\} \mid w_{st})$ using a nonparametric estimator $\hat{P}_{t}(\{j\} \mid w_{st})$. I use a single layer artificial neural network estimator and the asymptotic property of this estimator has been established in chen1999. The neural network estimator is computationally easy to implement and there is a readily used package for the estimator (bischl2016). Let $\hat{g}_{s, t}$ denote the estimated moment function that replaces the conditional choice probability $P_{t}(\{j\} \mid w_{st})$ with its estimator $\hat{P}_{t}(\{j\} \mid w_{st})$, then the sample objective function $\hat{\Omega}_N(\theta)$ is constructed as follows:
The second-step estimator for the parameter is obtained by minimizing the sample objective function $\hat{\Omega}_N$. Since the parameter $\beta_0$ and $\gamma_0$ can only be point identified up to a scale, the first element of the two parameters is normalized as one. This normalization is also used for other approaches to compare these methods fairly.
To better evaluate the performance of the two-step estimator (Two-Step Est.) in this paper, I implement a parametric estimator (Parametric Est.) using the method of simulated moments for comparison. For this parametric estimator, the error terms $\epsilon_{ijt}$ are assumed to follow a standard Gumbel distribution, independent across choices and periods, and also independent of all covariates. I allow the fixed effects $\alpha_{ij}$ to depend on covariates through a linear specification: $\alpha_{ij}=\eta_0+\bar{X}_{ij}'\eta_1+v_{ij}$, where $\bar{X}_{ij}=\frac{1}{T} \sum_{t} X_{ij t}$ denote the average covariates and $v_{ij}\sim \mathcal{N}(0, 1)$ follows standard normal distribution and independent of all covariates. This parametric estimator is $\sqrt{N}$ consistent when its assumptions are all correct, while could be inconsistent if either the parametric distribution or the linear model of the fixed effects is misspecified.
For the coefficient $\beta_0$, I also evaluate the performance of two other estimators for comparison that do not allow for the purchase of bundles $\Gamma_{it}=-\infty$. One estimator is Chamberlain's conditional fixed-effect logit estimator (FE Logit Est.). This estimator assumes $\epsilon_{ijt}$ to follow standard Gumbel distribution while leaving the distribution of the fixed effects $\alpha_{ij}$ unrestricted. The other estimator is the semiparametric estimator (Semi. Est.) which is developed under the stationarity assumption but assumes no bundles. Therefore, this estimator only uses conditional choice probabilities of $\{A, B, O\}$ to identify the coefficient $\beta_0$.
Now I describe the simulation setup. Let $d_x=2$ and $d_z=2$ denote the dimension of the covariates $X_{it}$ and $Z_i$ respectively. In each simulation, $X_{i\ell t}$ is drawn from the normal distribution $\mathcal{N}(0, d_x)$, independently across choices $\ell\in \{A, B\}$ and time $t\leq T$. Let the first element of $Z_i$ be drawn from $\mathcal{N}(2, 2)$ and the second element from $\mathcal{N}(0, 1)$. The true parameters are set as: $\beta_0=\gamma_0=(1,1)$.
I study four different designs of the error terms $\epsilon_{ijt}$ and fixed effects $\alpha_{ij}$. The first design considers the correct specification for the parametric estimator: $\epsilon_{ij}$ follows a Gumbel distribution and the fixed effects are specified as $\alpha_{ij}=\bar{X}_{ij}'\beta_0/2+ v_{ij}$. In the second design, the error term $\epsilon_{it}$ follows a bivariate normal distribution with the correlation $\rho=-0.7$. So the parametric distribution of $\epsilon_{it}$ is misspecified in this design. In the third design, I allow the fixed effects $\alpha_{ij}$ to depend on the covariates of the other good in a non-additive form: $\alpha_{ij}=(\bar{X}_{ij}/2-\bar{X}_{ik})'\beta_0*(1+v_{ij})$ for $j\in \{A, B\}$ and $k\neq j \in \{A, B\}$. In this design, the parametric estimator assumes a wrong model for the fixed effects $\alpha_{ij}$. The last design combines the second and third design, which considers both a misspecified distribution of $\epsilon_{it}$ and a misspecified model of $\alpha_{ij}$ for the parametric estimator. The following summarizes the four designs:
For the above four designs, I compare different approaches by reporting their root mean-squared error (rMSE) and mean of absolute deviation (MAD). For the complementarity parameter $\gamma_0$, I also report the probability of estimating the sign of substitution patterns incorrectly (Err) of the two estimators, defined as \[\text{Err}=E | \mathop{\rm sign} (Z_{i} \gamma_0)-\mathop{\rm sign}(Z_{i} \hat{\gamma} ) |.\]
I study three different sample sizes $N=\{1000, 2000, 4000\}$ and set the simulation repetitions to $B=500$. The performance of the estimator for $\gamma_0$ and $\beta_0$ is displayed in Table (ref) and in Table (ref), respectively.
Table (ref) compares the performance of the two-step estimator with the parametric estimator for the complementarity parameter $\gamma_0$. The parametric estimator performs better only when its assumptions are correctly specified (design 1), but has a worse performance under misspecifications (designs 2-4) especially when both the parametric distribution and the model of fixed effects are misspecified. The two-step estimator has uniform performance in all four designs, showing its advantage of performing robustly under different designs of parametric distributions and models of fixed effects. Moreover, as the sample size increases, the deviation and bias of the two-step estimator both shrink significantly. However, the bias of the parametric estimator does not decrease as the sample increases in designs 2-4, which shows the inconsistency of this estimator under misspecifications.
Table (ref) compares the performance of the two-step estimator with three other estimators described before for the coefficient $\beta_0$ under the four designs. Similarly, the two-step estimator performs uniformly better than the other three estimators in designs 2-4, and the difference becomes more significant as the sample size increases. To summarize, the results in Table (ref) and Table (ref) demonstrate the advantage of the two-step estimator in performing robustly with respect to different parametric distributions or specifications of dependence structures between covariates and fixed effects.
The previous sections focused on the case where heterogeneity in the complementarity $\Gamma_{it}$ only comes from observed covariates: $\Gamma_{it}=\Gamma(Z_{i})$. This section accommodates unobserved heterogeneity in the complementarity across individuals by allowing $\Gamma_{it}$ to be different for each individual regardless of their covariates. The sign of $\Gamma_{it}$ captures the heterogeneous complementarity relationship among the two goods for each individual. Therefore, I focus on identifying the distribution of the sign of $\Gamma_{it}$ which represents the fraction of people for whom the two goods are complements or substitutes.
Next, I introduce some assumptions on the complementarity and error terms.
Assumption (ref) is similar to Assumption (ref), except it also assumes a stationarity condition for the complementarity $\Gamma_{it}$. Under the assumption that the complementarity $\Gamma_{it}$ only depends on observed covariates (Assumption (ref)), Assumption (ref) degenerates to the stationarity condition in Assumption (ref) since $\Gamma_{it}$ is a constant conditional on the covariate.
Assumption (ref) only requires that the distribution of $\Gamma_{it}$ remains stationary over time, but it still allows the complementarity $\Gamma_{it}$ for each individual to vary over time. Moreover, this assumption does not restrict the dependence between the complementarity $\Gamma_{it}$ with the unobserved terms, including the fixed effects $\alpha_{i}$ and the error terms $\epsilon_{it}$.
Let $X_i=(X_{it})_{t=1}^T$ collect the covariate of all time periods.
Assumption (ref) assumes the independence between the complementarity $\Gamma_{it}$ and the vector of covariates for all periods. Under this condition, variation in all covariates can be used to identify the distribution of the complementarity $\Pr(\Gamma_{it}\geq 0)$. This assumption can be relaxed to accommodate the situation where there is a subset of covariates that are independent of $\Gamma_{it}$ while other covariates can be correlated with $\Gamma_{it}$. In such a scenario, the analysis can be conducted conditional on the covariates that are potentially correlated with the complementarity.
Under the above assumptions, I establish the identification of the fraction of individuals for whom the two goods are complements, denoted as $\eta=\Pr(\Gamma_{it}\geq 0)$. According to Assumption (ref), the distribution of $\Gamma_{it}$ is stationary over time so that $\eta$ does not depend on $t$. The identification result for $\Pr(\Gamma_{it}< 0)$ can be directly derived using the formula $\Pr(\Gamma_{it}<0)=1-\Pr(\Gamma_{it}\geq0)$ so it is omitted here.
The intuition of the identification strategy for $\eta$ is described as follows. The conditional demand for one good can be expressed as a mixture of two groups: people for whom the two goods are complements ($\Gamma_{it}\geq 0$) and people for whom the two goods are substitutes ($\Gamma_{it}<0$). When the covariate index of good $A$ increases, it will affect the demand for good $B$ for the two groups of people in different directions. This relationship can help identify the fraction of the two groups.
Similar to Section (ref), the first step is to derive the sign of variation in covariate indices $(\Delta_{s, t} \delta_{A}, \Delta_{s, t} \delta_{B})$ from conditional choice probabilities. Let $\xi^{1}_{s,t} (x_s, x_t)$ and $\xi^{2}_{s,t} (x_s, x_t)$ be defined as
From the variation in observed choice probabilities, the sign of variation in covariate indices of two goods can be identified. When an increase in probabilities of all choices $\{A, B, AB\}$ is observed, it can be inferred that the covariate indices of both goods should increase. Similarly, an increase in probabilities of all choices except choice $B$ suggests that the covariate index of good $A$ increases and that of good $B$ decreases.
Let $\mathcal{X}^{1}_{s,t}=\{ (x_s, x_t)\mid \xi^{1}_{s,t} (x_s, x_t) =1\} $ and $\mathcal{X}^{2}_{s,t}=\{ (x_s, x_t)\mid \xi^{2}_{s,t} (x_s, x_t) =1\} $ denote the collection of covariates such that $\xi^{1}_{s,t} (x_s, x_t) =1$ and $\xi^{2}_{s,t} (x_s, x_t)=1$ respectively, which implies
Given the sign of covariate indices, using variation in the demand for two goods can identify the fraction of people for whom the two goods are complements $\eta=\Pr(\Gamma_{it}\geq 0)$. To convey the idea, I first consider that the covariate indices for goods $A$ and $B$ both increase: $(x_s, x_t)\in \mathcal{X}^{1}_{s,t}$. In this scenario, the demand for the two goods would increase for people for whom the two goods are complements, but may decline for people for whom the two goods are substitutes. Therefore, a decline in demand for either of the two goods in data can only come from people with $\Gamma_{it}<0$, which can help establish a lower bound for the fraction of people with $\Gamma_{it}<0$ and thus an upper bound for the fraction of people with $\Gamma_{it}\geq 0$. Similarly, a lower bound for $\eta=\Pr(\Gamma_{it}\geq 0)$ can be provided when covariates satisfy $(x_s, x_t)\in \mathcal{X}^{2}_{s,t}$.
The next proposition characterizes the partial identification results for $\eta=\Pr(\Gamma_{it}\geq 0)$.
Proposition (ref) establishes both lower and upper bounds for $\eta$ by exploiting variation in the demand for the two goods under different sets of covariate indices. According to the definition of the lower and upper bounds, it is always true that $L_{\eta}\geq 0, U_{\eta}\leq 1$ and the bounds use variation in all values of covariates over any two periods. The range of the bounds depends on variation in conditional demand of the two goods, and larger variation leads to tighter bounds. The result in Proposition (ref) also provides testable implications for Assumptions (ref)-(ref) since it implies that the upper bound should be no smaller than the lower bound: $U_{\eta}\geq L_{\eta}$. The proof for Proposition (ref) is provided in Appendix (ref).
The prior work by allen2022 also studies latent complementarity and provides bounds for the fraction of the population for whom the two goods are complements with cross-sectional data. They exploit an exclusion restriction and an independence assumption between the covariates and unobserved terms. My method complements their approach by considering panel data setting and using intertemporal variation over time. I allow covariates to be arbitrarily dependent with unobserved fixed effects. Moreover, I do not impose exclusion restrictions and can still partially identify $\eta$ when covariates of both goods change simultaneously.
This paper uses a panel multinomial choice model with bundles to study the substitution and complementarity relationship between goods. The model imposes no parametric assumptions on the idiosyncratic error terms and allows for endogeneity by admitting flexible dependence structures between observed covariates and unobserved fixed effects. I provide testable implications for the substitution and complementarity relationship, and derive the sharp identification set for the model parameters.
The primary identification strategy is to derive identifying restrictions on unknown parameters through intertemporal variation in conditional choice probabilities that are identified from data. I construct conditional moment inequalities to characterize the identified set for estimation and inference. The method in the paper is shown via Monte Carlo simulations to perform more robustly than the parametric method concerning different specifications of fixed effects and distributions of error terms. In the extension, the paper allows for unobserved heterogeneity in the complementarity and establishes partial identification results. In the online Appendix, I also study the case of more than two goods, a class of nonseparable utility functions, as well as cross-sectional models.
This paper focuses on a static panel multinomial choice model where consumers' utility of goods only depends on characteristics in the same periods. It would be interesting to investigate how to identify the complementarity relationship in a dynamic model, where consumers' choices may also depend on past choices. The analysis would be more complicated as one has to disentangle the effect from previous choices and the current complementarity. In addition, it could also be worthwhile to explore how to identify the complementarity and utility coefficients with heterogeneous and unknown choice sets.