The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
108,322 characters
Identification of Semiparametric Panel Multinomial Choice Models with Infinite-Dimensional Fixed Effects
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long\global\long\global\long\global\long
\global\long\global\long
\global\long
\global\long
\global\long
\global\long\global\long\global\long\global\long\global\long
\global\long
\global\long\def\mc#1{\mathscr{#1}}
\global\long\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }
\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }
\global\long\def\turd#1{\frac{#1}{3}}
\global\long
\global\long\def\sand#1{\left\lceil #1\right\vert }
\global\long\def\wich#1{\left\vert #1\right\rfloor }
\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }
\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\inprod#1{\left\langle #1\right\rangle }
\global\long\def\ol#1{\overline{#1}}
\global\long\def\ul#1{\underline{#1}}
\global\long\def\td#1{\tilde{#1}}
\global\long\def\bs#1{\boldsymbol{#1}}
\global\long
\global\long
\global\long
\global\long
\global\long\global\long
\title{Identification of Semiparametric Panel Multinomial\\
Choice Models with Infinite-Dimensional Fixed Effects\thanks{Researcher(s)\textquoteright{} own analyses calculated (or derived)
based in part on data from Nielsen Consumer LLC and marketing databases
provided through the NielsenIQ Datasets at the Kilts Center for Marketing
Data Center at The University of Chicago Booth School of Business.
The conclusions drawn from the NielsenIQ data are those of the researcher(s)
and do not reflect the views of NielsenIQ. NielsenIQ is not responsible
for, had no role in, and was not involved in analyzing and preparing
the results reported herein.} \thanks{We thank Xiaohong Chen, Peter Phillips, and Phil Haile for their invaluable
advice and encouragement. We thank three anonymous referees for their
comments that significantly improved the paper. We thank Don Andrews,
Isaiah Andrews, Tim Armstrong, Xu Cheng, Tim Christensen, Ben Connault,
Francis Diebold, Bo Honor\'{e}, Joel Horowitz, Yuichi Kitamura, Patrick
Kline, Lixiong Li, Yuan Liao, Charles Manski, Aviv Nevo, Matt Seo,
Xiaoxia Shi, Frank Schorfheide, Elie Tamer, Ed Vytlacil, Rui Wang,
Sheng Xu and participants at various seminars and conferences for
helpful comments. We thank Wenli Lyu and Chuyue Tian for excellent
research assistance.}}
\author{Wayne Yuan Gao\thanks{Gao: Department of Economics, University of Pennsylvania, \protect[email removed].}
~and Ming Li\thanks{Li: Department of Economics and Risk Management Institute, National
University of Singapore, \protect[email removed].}}
\maketitle
\begin{abstract}
\noindent This paper proposes a robust method for semiparametric identification
and estimation in panel multinomial choice models, where we allow
for infinite-dimensional fixed effects that enter into consumer utilities
in an additively nonseparable way, thus incorporating rich forms of
unobserved heterogeneity. Our identification strategy exploits multivariate
monotonicity in parametric indices, and uses the logical contraposition
of an intertemporal inequality on choice probabilities to obtain identifying
restrictions. We provide a consistent estimation procedure, and demonstrate
the practical advantages of our method with Monte Carlo simulations
and an empirical illustration on popcorn sales with the NielsenIQ
data.
\end{abstract}
\newpage{}
\setlength{\abovedisplayskip}{3pt}
\setlength{\belowdisplayskip}{3pt}
\section{\label{sec:Intro}Introduction}
\noindent This paper proposes a method for semiparametric identification
and estimation in panel multinomial choice models, where we allow
for infinite-dimensional fixed effects that enter into consumer utilities
in an additively nonseparable manner. The proposed method also applies
more widely beyond panel multinomial choice models, and can be adapted
to a wide range of models characterized by \emph{multi-index single-crossing
conditions}, which we introduce later in this paper.
To fix ideas, we start with the following panel multinomial choice
model:
\[
y_{ijt}=\mathbf{\mathbbm1}\left\{ u\left(X_{ijt}^{'}\beta_{0},A_{ij},\epsilon_{ijt}\right)\geq\max_{k\in\left\{ 1,\ldots,J\right\} }u\left(X_{ikt}^{'}\beta_{0},A_{ik},\epsilon_{ikt}\right)\right\} ,
\]
where agent $\mathit{i}$'s utility from a candidate product $j$
at time $t$, represented by $u(X_{ijt}^{'}\beta_{0},A_{ij},\epsilon_{ijt})$,
is taken to be a function of three components. The first is a linear
index $X_{ijt}^{'}\beta_{0}$ of observable characteristics $X_{ijt}$,
which contains a finite-dimensional parameter of interest $\beta_{0}$
we will identify and estimate. The second term $A_{ij}$ is an infinite-dimensional
fixed effect that can be heterogeneous across each agent-product combination.
We emphasize that $X_{ijt}$ and $A_{ij}$ can be arbitrarily dependent.
The last term $\epsilon_{ijt}$ is an idiosyncratic time-varying error term
of arbitrary dimension. The three components are then aggregated by
an unknown utility function $u$ in an additively nonseparable way,
with the only restriction being that each agent's utility $u(X_{ijt}^{'}\beta_{0},A_{ij},\epsilon_{ijt})$
is\emph{ increasing} in its first argument. Each agent then chooses
a certain product in a given time period, represented by $y_{ijt}=1$,
if and only if this product gives her the highest utility among all
available products.
The infinite dimensionality of the terms $u$, $A_{ij}$, and $\epsilon_{ijt}$,
together with the model\textquoteright s additively non-separable
interaction structure, jointly generates a rich class of unobserved
heterogeneity. Across each agent-product combination $ij$, we are
effectively allowing for nonparametric variations in agent utilities.
Such variation proxies for the effects of complicated unobserved factors
that influence choice behavior, such as brand loyalty, subtle flavors,
and unique styles of products. In addition, we work with a nonparametric
time homogeneity assumption on the error terms $\epsilon_{ijt}$ that restricts
$\epsilon_{ijt}$ and $\epsilon_{ijs}$ from two periods $t$ and $s$ to have
the same marginal distribution given the observed covariates from
the two periods. Apart from this, we impose no parametric restrictions
on the distribution of $\epsilon_{ijt}$ and no additional restrictions
on its dependence across time $t$ and products $j$. In particular,
the fully unrestricted dependence of $\epsilon_{ijt}$ across products $j$
allows our framework to remain robust to the well-known ``Blue-Bus/Red-Bus''
problem and related pathologies that arise in many standard multinomial
choice models, as discussed in \citet*{berry2007pure}.\footnote{See the discussion after Assumption \ref{assu:EpsDist} in Section
\ref{subsec:PMC_Setup} for more details.}
The generality of our framework nests many semiparametric (and parametric)
panel multinomial choice models with scalar fixed effects, scalar
error terms, and varying degrees of additive separability, including
the following standard specification:
\[
y_{ijt}=\mathbf{\mathbbm1}\left\{ X_{ijt}^{'}\beta_{0}+A_{ij}+\epsilon_{ijt}>\max_{k\in\left\{ 1,\ldots,J\right\} }\left(X_{ikt}^{'}\beta_{0}+A_{ik}+\epsilon_{ikt}\right)\right\} .
\]
Relative to existing work, our framework accommodates both the infinite
dimensionality of unobserved heterogeneity and non--additive separability
in agent utilities, under a standard time-homogeneity assumption on
the idiosyncratic error term that is widely used in the literature.
Our identification strategy leverages multivariate monotonicity in
contrapositive form. The intuition is straightforward: if the choice
probability of a given product (or subset of products) strictly \textit{increases}
from one period to the next, then it \textit{cannot} be that this
product (or all products in the subset) becomes \textit{worse} while
all other products become \textit{better} over the two periods. By
applying this contraposition to a carefully constructed inequality
in conditional choice probabilities, we obtain an identifying restriction
on the index values that is free of all infinite-dimensional nuisance
parameters. We further show that, in a two-period setting, the identified
set obtained by aggregating these restrictions across all product
subsets is sharp.
Based on our identification result, we provide consistent two-step
set (or point) estimators, together with a computational algorithm
adapted to the technical challenges of our framework. The first stage
takes the form of a standard nonparametric regression, where we estimate
a collection of intertemporal differences in conditional choice probabilities.
In the second stage, we numerically minimize our sample criterion
function with the first-stage estimates plugged in. A highlight of
our computational procedure is the adoption of a spherical-coordinate
reparameterization of our criterion functions in terms of \emph{angles},
which enables us to exploit a combination of topological, geometric
and computational advantages. A simulation study is conducted to analyze
the finite-sample performance of our method and the adequacy of our
computational procedure for practical implementation.
We also provide an empirical illustration of our procedure, where
we use the NielsenIQ data on popcorn sales in the United States to
explore the effects of marketing promotion. The results show that
our procedure produces estimates that conform well with economic intuition.
For example, we find that special in-store displays boost sales not
only through a direct promotion effect but also through the attenuation
of consumer price sensitivity. Intuitively, marketing managers are
more likely to promote products for which they know consumers are
more sensitive to price and promotion. Hence, the average effective
price sensitivities of promoted products tend to be larger than those
not promoted due to the selection effect. Given the non-additive nature
of such selection effects, estimators based on additive separability
will be biased. In contrast, our method is robust to such confounding
effects, thus producing more sensible estimates.
The validity of our identification strategy, as well as our estimation
procedure, relies solely on monotonicity in an index structure, and
thus extends naturally beyond panel multinomial choice models. We
also introduce the \emph{multi-index single-crossing (MISC)} condition
framework, a general econometric framework under which our method
can be applied. This framework encompasses the key ingredients of
a large class of models, such as binary choice models with awareness,
binary choice with endogeneity, dyadic network formation, bilateral
matching, and endogenous censoring.
We acknowledge several limitations of our identification approach.
First, our current model setup does not allow for time-varying endogeneity
between observed covariates $X_{ijt}$ and the error term $\epsilon_{ijt}$,
which effectively rules out the inclusion of contemporaneously endogenous
covariates and/or lagged outcomes. See a subsequent paper by \citet{gao2023identification}
for a weaker version of the time homogeneity assumption that can be
exploited for identification in panel multinomial choice models with
endogenous and dynamic covariates. Second, our current approach does
not allow for random coefficients on time-varying covariates. That
said, rich forms of time-invariant taste heterogeneity have been absorbed
into the nonparametric fixed effects under our current setting. Third,
in the short-panel setting, our current approach effectively ``differences
out'' the unobserved individual fixed effects $A_{i}$. As a result,
it cannot identify counterfactual parameters that depend on the distribution
of $A_{i}$. However, in long panels, our approach can be adapted
to identify counterfactual parameters, which we discuss in more detail
in Appendix \ref{sec:Counterfactual-Analysis}.
\medskip{}
This paper builds upon and contributes to a large literature on semiparametric
(and parametric) discrete choice models, dating back to \citet*{mcfadden1974conditional}
and \citet*{manski1975maximum}, and more specifically to the line
of literature on panel multinomial choice models. Our work is most
closely related to \citet*{pakes2016moment}, who also exploit weak
monotonicity and time homogeneity, but restrict the effect of unobserved
heterogeneity to be a scalar index that is additively separable from
the index of observable characteristics. \citet*{shi2017estimating}
exploit cyclical monotonicity of \emph{vector}-valued functions in
a fully additive panel multinomial choice model. \citet*{khan2021inference}
consider another additive model, but utilize the subsample of observations
with time-invariant covariates along \emph{all products but one} so
as to leverage univariate monotonicity. \citet*{honore2000panel}
also exploit univariate monotonicity when certain covariates across
two periods are equal in a dynamic panel setting. \citet*{chernozhukov2017nonseparable}
consider a model with an additive effect under an ``on-the-diagonal''
restriction (i.e., when covariates at two different time periods coincide).
By allowing non-additiveness in the specification of utility functions
and infinite-dimensional fixed effects, our method is different from
and thus complementary to those proposed in these aforementioned papers.
This paper is also connected to the literature on the nonparametric
identification and estimation in discrete choice models (e.g., \citet{berry2014identification,compiani2022market}).
That literature assumes monotonicity restrictions of the demand functions
in product-market specific parametric indices to invert the demand
system. This is particularly useful as it leads to a system of equations
with only one unobservable per equation, from which the unobservable
product-market specific demand shifter can be successfully constructed.
Our paper considers a different model, but also leverages monotonicity
restrictions in parametric indices to facilitate identification and
estimation of the structural parameters. In both cases, the index
assumption restricts the amount of unobserved heterogeneity in preferences
for the variables included in the index.
More broadly, our work connects to the semiparametric literature on
the identification and estimation of models characterized by monotonicity
in a single parametric index. A related class of estimators that leverage
univariate monotonicity, known as \emph{maximum score }or\emph{ rank-order
estimator}s, dates back to a series of important contributions by
\citet*{manski1975maximum,manski1985semiparametric,manski1987semiparametric},
and is further investigated in \citet*{han1987non}, \citet*{horowitz1992smoothed},
\citet*{abrevaya2000rank}, \citet*{honore2002semiparametric}, \citet*{fox2007semiparametric},
and \citet*{yan2019semiparametric}.\footnote{We clarify that \citet{manski1975maximum}, \citet*{fox2007semiparametric}
and \citet*{yan2019semiparametric} consider multinomial choice models
with multiple parametric indices, but focus on settings where the
``multi-index'' problem can be reduced to leverage single-variate
monotonicity (or a single-variate rank-order property).} Despite the similar reliance on monotonicity, the\emph{ multi-index}
nature of our model, and more importantly the multivariate monotonicity
condition that we leverage, induce key differences from the \textit{single-index}
(and univariate monotonicity) setting, leading to a substantially
different estimation method relative to rank-order estimators.
\medskip{}
The rest of this paper is organized as follows. Section \ref{sec:PMC}
introduces our main model specifications and assumptions. Section
\ref{subsec:ID} presents our key identification strategy. In Section
\ref{sec:S_EstComp}, we provide consistent estimators along with
a computational procedure to implement it. Section \ref{sec:Ext_MMIM}
discusses the generalization of our method to the framework of multi-index
single-crossing conditions. Sections \ref{sec:Sim} and \ref{sec:Emp}
switch back to our main panel multinomial choice model, for which
we provide a simulation study and an empirical illustration using
the NielsenIQ data. We conclude in Section \ref{sec:Conclusion}.
\section{\label{sec:PMC}Panel Multinomial Choice Model}
\subsection{\label{subsec:PMC_Setup}Model and Assumptions}
In this section, we present a semiparametric panel multinomial choice
model featuring infinite-dimensional unobserved heterogeneity and
flexible forms of non-separability, which serves as the main framework
for illustrating our identification and estimation method.
Specifically, we consider the following model, in which individual
$i$ chooses product $j$ at time $t$ if and only if $i$ prefers
product $j$ to all other alternatives at time $t$:
\begin{align}
y_{ijt} & =\mathbf{\mathbbm1}\left\{ u\left(X_{ijt}^{'}\beta_{0},A_{ij},\epsilon_{ijt}\right)>\max_{k\in\left\{ 1,\ldots,J\right\} \backslash\left\{ j\right\} }u\left(X_{ikt}^{'}\beta_{0},A_{ik},\epsilon_{ikt}\right)\right\} \label{eq:Model_PMC}
\end{align}
where:
\begin{itemize}
\item $i\in\{1,\ldots,N\}$ denotes $N$ \textit{individuals}.
\item $j\in{\cal J}:=\{1,\ldots,J\}$ denotes the set of $J$ choice alternatives,
with \textit{products} indexed by $1,\ldots,J$. Throughout this paper,
we treat the number of products $J$ as fixed.
\item $t\in\{1,\ldots,T\}$ denotes the $T\geq2$ time periods. In this
paper, we consider a short-panel setting in which $T$ is fixed.
\item $X_{ijt}$ is an $\mathbb{R}^{D}$-valued vector of observable characteristics
specific to each agent--product--time tuple $ijt$. These may include,
for example, buyer characteristics such as income, product characteristics
such as price and promotion status, as well as interaction and higher-order
terms of these variables.
\item $y_{ijt}$ is an observable binary variable, with $y_{ijt}=1$ indicating
that buyer $i$ chooses product $j$ at time $t$, and $y_{ijt}=0$
indicating otherwise.
\item $\beta_{0}\in\mathbb{R}^{D}$ is the finite-dimensional parameter of interest.
\item $A_{ij}$ represents an $ij$-specific time-invariant unobserved heterogeneity
term of arbitrary dimension, which we refer to as the $ij$-specific
\textit{fixed effect}.
\item $\epsilon_{ijt}$ is an $ijt$-specific unobserved error term of arbitrary
dimension, which captures time-idiosyncratic utility shocks to product
$j$ for agent $i$ at time $t$.
\item $u$ is an unknown function, interpreted as a \textit{utility function}
that aggregates the parametric index $X_{ijt}^{'}\beta_{0}$, the fixed
effect $A_{ij}$, and the error term $\epsilon_{ijt}$ into a scalar representing
agent $i$'s utility from choosing product $j$ at time $t$.
\end{itemize}
We first present our main modeling assumptions and then discuss these
assumptions in conjunction with our model specification \eqref{eq:Model_PMC}.
To economize on notation, we refer to the collection of variables
concatenated along product and time dimensions: ${\bf X}_{it}=(X_{ijt})_{j=1}^{J}$,
${\bf X}_{i}=({\bf X}_{it})_{t=1}^{T}$, ${\bf A}_{i}=(A_{ij})_{j=1}^{J}$,
$\boldsymbol{\epsilon}_{it}=(\epsilon_{ijt})_{j=1}^{J}$, and $\boldsymbol{\epsilon}_{i}=(\boldsymbol{\epsilon}_{it})_{t=1}^{T}$.
We also write $\delta_{ijt}=X_{ijt}^{'}\beta_{0}$ to denote the parametric
index.
\begin{assumption}[Cross-Sectional Random Sampling]
\label{assu:RandSamp} $({\bf Y}_{i},{\bf X}_{i},{\bf A}_{i},\boldsymbol{\epsilon}_{i})$
is i.i.d. across $i\in\left\{ 1,\ldots,N\right\} $ with $N\to\infty$.
\end{assumption}
\noindent Assumption \ref{assu:RandSamp} is a standard assumption.\footnote{It is worth noting that we have not yet imposed any explicit restrictions
on the structure of the spaces in which the arbitrary-dimensional
random elements ${\bf A}_{i}$ and $\bs{\epsilon}_{i}$ are defined. However,
implicit in our model specification and in Assumption \ref{assu:RandSamp}
is the requirement that $({\bf Y}_{i},{\bf X}_{i},{\bf A}_{i},{\bf \epsilon}_{i})$
be well-defined as random elements---that is, measurable functions---on
a sufficiently rich probability space $(\Omega,\mathscr{F},\mathbb{P})$.} Recall that the number of time periods $T$ is held fixed, and we
focus on a short panel setting with cross-sectional asymptotics.
\begin{assumption}[Monotonicity in the Index]
\label{assu:PMC_Mono} For every realization of $(A_{ij},\epsilon_{ijt})$,
the mapping $\tilde{\delta}\longmapsto u(\tilde{\delta},A_{ij},\epsilon_{ijt})$
is weakly increasing in the scalar-valued argument $\tilde{\delta}$.
\end{assumption}
\noindent Essentially, Assumption \ref{assu:PMC_Mono} states that
$u(\delta_{ijt},A_{ij},\epsilon_{ijt})$, the utility of individual $i$ from
choosing product $j$ at time $t$, is weakly increasing\footnote{It should be clarified that increasingness is without loss of generality
given monotonicity, since the index $\delta_{ijt}=X_{ijt}^{'}\beta_{0}$
contains an unknown parameter with unrestricted signs.} in the index $\delta_{ijt}$. Given the index structure, monotonicity
itself is a relatively mild assumption. In the standard panel multinomial
choice model with scalar-valued $A_{ij}$ and $\epsilon_{ij}$ along with
additive $u$, i.e., $u(\delta_{ijt},A_{ij},\epsilon_{ijt})=\delta_{ijt}+A_{ij}+\epsilon_{ijt}$,
Assumption \ref{assu:PMC_Mono} is trivially satisfied (with strictness)
by construction.
In a way, Assumption \ref{assu:PMC_Mono} endows the index $\delta_{ijt}$
and the parameter $\beta_{0}$ with economic interpretations. Under Assumption
\ref{assu:PMC_Mono}, $\delta_{ijt}$ may be considered as a quality measure
of the match between agent $i$ and product $j$ based on their observable
characteristics at time $t$, inducing an interpretation of $\beta_{0}$
as representing how a certain change in a linear combination of observable
characteristics may increase utilities for \textit{all} agents from
a certain product $j$, \emph{ceteris paribus}. Hence, without Assumption
\ref{assu:PMC_Mono}, it would be hard to interpret $\beta_{0}$.
Moreover, the monotonicity restriction is imposed on $\delta_{ijt}$,
but not directly on any specific observable characteristics in $X_{ijt}$:
quadratic or higher-order polynomial terms, and other functions of
observable characteristics can be included in $X_{ijt}$ whenever
appropriate.
\begin{assumption}[Pairwise Time Homogeneity]
\label{assu:EpsDist}The distributions of $\boldsymbol{\epsilon}_{it}$
and $\boldsymbol{\epsilon}_{is}$ conditional on $({\bf X}_{it},{\bf X}_{is},{\bf A}_{i})$
across any pair of periods $t\neq s\in\{1,\ldots,T\}$ satisfy
\[
\rest{\boldsymbol{\epsilon}_{it}}\left({\bf X}_{it},{\bf X}_{is},{\bf A}_{i}\right)\rest{\sim\boldsymbol{\epsilon}_{is}}\left({\bf X}_{it},{\bf X}_{is},{\bf A}_{i}\right).
\]
\end{assumption}
\noindent Assumption \ref{assu:EpsDist}, a multinomial extension
of the group homogeneity assumption in \citet*{manski1987semiparametric},
is also a standard assumption in the literature on panel multinomial
choice models, such as in \citet*{chernozhukov2017nonseparable},
\citet*{shi2017estimating}, and \citet*{pakes2016moment}.\footnote{In particular, \citet*{pakes2016moment} investigate the following
panel multinomial choice model:
\begin{equation}
y_{ijt}=\mathbf{\mathbbm1}\left\{ g_{j}\left(X_{ijt},\beta_{0}\right)+f_{j}\left(A_{ij},\epsilon_{ijt}\right)>\max_{k\neq j}g_{k}\left(X_{ikt},\beta_{0}\right)+f_{k}\left(A_{ik},\epsilon_{ikt}\right)\right\} ,\label{eq:Model_PakesPorter}
\end{equation}
where the function $g_{j}$ produces a potentially nonlinear parametric
index and $f_{j}$ aggregates fixed effects and idiosyncratic errors
into a scalar value in a nonseparable way, while additive separability
between the observable covariate index $g_{j}(X_{ijt},\beta_{0})$ and
the unobserved heterogeneity index $f_{j}(A_{ij},\epsilon_{ijt})$ is still
maintained. Moreover, although the dimensions of $A_{ij}$ and $\epsilon_{ijt}$
are not restricted in \citet*{pakes2016moment}, their joint effect
is effectively summarized by a single scalar index $f_{j}(A_{ij},\epsilon_{ijt})$.
We reiterate that our model \eqref{eq:Model_PMC} not only incorporates
infinite-dimensionality in unobserved heterogeneity as captured by
$A_{ij}$ and $\epsilon_{ijt}$, but also allows such heterogeneity to enter
into agent utility functions in a fully \textit{nonseparable }way.} As shown in the next subsection, Assumption \ref{assu:EpsDist} is
the key condition underlying our identification strategy. Essentially,
we rely on Assumption \ref{assu:EpsDist} to link (a particular class
of) intertemporal changes in conditional choice probabilities between
two periods $s$ and $t$ to intertemporal changes in the parametric
indices of all products, $(\delta_{ijt}-\delta_{ijs})_{j\in{\cal J}}$, thereby
yielding identifying restrictions for the parameter $\beta_{0}$. Note
that Assumption \ref{assu:EpsDist} concerns only the marginal distributions
of $\boldsymbol{\epsilon}_{it}$ in different periods, and we make no explicit
assumptions about the serial dependence between $\boldsymbol{\epsilon}_{is}$
and $\boldsymbol{\epsilon}_{it}$, nor do we impose any restrictions on
the dependence structure of $A_{ij}$ and $\epsilon_{ijt}$ across products
$j\in{\cal J}$.
It is worth emphasizing that the absence of restrictions on the cross-product
dependence structure of $\epsilon_{ijt}$ in our setup also enables our
choice model to circumvent the well-known ``Blue-Bus/Red-Bus'' problem\footnote{See, for example, \citet*{mcfadden1974conditional} and \citet{train2009discrete}
for descriptions of the Independence of Irrelevant Alternatives property
and the ``Blue-Bus/Red-Bus'' problem.} that arises in discrete choice models with additive errors that are
assumed to be independent across products. Specifically, consider
a standard multinomial logit model: if we arbitrarily create a new
dummy product by replicating an existing one and giving it a different
name, then under various (mixed) logit models a new independent copy
of logit error is drawn for this new dummy product, which results
in a strict increase in the consumers' indirect utilities, even though
the new product is a simple duplicate of an existing product. This
``Blue-Bus/Red-Bus'' phenomenon also implies that consumers' indirect
utilities would diverge to infinity if we keep adding such dummy new
products, and that consumers will choose one of these duplicate products
with increasingly higher probabilities, both of which are unrealistic.
This problem, as well as other conceptual issues with independent
additive errors, has been well discussed, say, in \citet{berry2007pure}.
We emphasize that our current model does \emph{not} lead to the ``Blue-Bus/Red-Bus''
problem, since we allow errors to be arbitrarily correlated across
products. Hence, when duplicate products are created, the errors of
duplicate products are allowed to be perfectly correlated (as they
should be), so that adding duplicate products with different name
labels does not blow up the indirect utility of the consumer, nor
does it make the consumer more likely to choose one of the duplicates
relative to any set of remaining products.
That said, Assumption \ref{assu:EpsDist} does entail certain limitations
for the model. First, the assumption effectively rules out time-varying
endogeneity\footnote{Note that time-invariant endogeneity can be incorporated through the
unrestricted dependence between $\epsilon_{ijt}$ and the fixed effect $A_{ij}$,
which can be arbitrarily correlated with ${\bf X}_{i}$ (in time-invariant
manners).}: for example, if ${\bf X}_{it}\neq{\bf X}_{is}$, the conditional
distribution of $\boldsymbol{\epsilon}_{it}$ is nevertheless assumed to
be the same as that of $\boldsymbol{\epsilon}_{is}$, which is likely to
be violated if the distribution of $\boldsymbol{\epsilon}_{it}$ covaries
with that of ${\bf X}_{it}$. This is admittedly a limitation of the
approach, though it is common to this line of literature that utilizes
Assumption \ref{assu:EpsDist} as cited before. Nonetheless, subsequent
work by \citet*{gao2023identification} has proposed a weakened notion
of the stationarity/homogeneity assumption that is only imposed on
``exogenous covariates'', and has proposed an approach for partial
identification, albeit in a slightly different setting with additive
scalar-valued fixed effects. Relatedly, \citet{li2024identificationestimationtimevaryingendogenous}
proposes a correlated random coefficient linear panel model in which
regressors can be correlated with time-varying, individual-specific
random coefficients, thereby accommodating time-varying endogeneity
in the covariates. Second, a further limitation of Assumption \ref{assu:EpsDist}
is that it rules out random coefficients, a modeling device that was
popularized by \citet*{blp1995io} due to its ability to generate
rich substitution patterns among products with multi-dimensional observable
characteristics. However, the flexibility afforded by our general
fixed effect specification can incorporate arbitrarily complicated
substitution patterns with respect to \emph{time-invariant} components
of observed and unobserved product characteristics. Our infinite-dimensional
fixed-effect approach is thus more suitable to panel-data settings
where researchers are more interested in incorporating an arbitrarily
complicated form of time-invariant heterogeneity across agent-product
pairs.
Beyond Assumption \ref{assu:EpsDist}, another limitation of the paper
is that we treat the distribution of unobserved heterogeneity (i.e.,
the fixed effects) as a nuisance parameter and focus on the identification
of $\beta_{0}$. Many counterfactual parameters require knowledge of
the distribution of the relevant unobserved heterogeneity terms, which
is not identified given that the fixed effects are ``differenced out''
under our current approach in a short-panel setting. To this end,
we discuss in Appendix \ref{sec:Counterfactual-Analysis} how to use
the estimated $\beta_{0}$ in a long panel setting ($T\to\infty$) to
perform counterfactual analysis. The idea is when a long panel is
available, we can consistently estimate for each individual certain
parameters of interest by using the observations only from that individual,
which effectively controls for the unobserved ${\bf A}_{i}$.
Furthermore, there are interesting economic questions that can be
addressed based on the knowledge of $\beta_{0}$, and we discuss two
of them here.\footnote{We thank an anonymous referee for the suggestions here.}
First, consider research questions that focus on the existence and
direction of an effect, which can often be determined by the relative
magnitudes of effects from covariates as captured by $\beta_{0}$. For
instance, one may ask whether advertising by a competitor within the
same product category exerts a positive or negative influence on the
demand for a focal brand. Theoretical considerations offer plausible
arguments for both substitution and category-expansion effects \citep*{simon1980shape,narayanan2004return},
rendering the sign of the net impact an empirical matter of interest.
Second, the proposed approach may also be applied to model specification
testing (or as a robustness check) for models with more restrictive
specifications on the fixed effects and errors. Specifically, by comparing
the estimate of $\beta_{0}$ obtained from our model with that from a
more standard model (say, multinomial logit with fixed effects), one
can assess whether the parametric restrictions, particularly those
pertaining to unobserved heterogeneity, are supported by the data.
This aligns with the broader econometric literature on specification
testing (e.g., \citet*{hausman1978specification,vuong1989likelihood}),
and can be particularly useful for evaluating the validity of assumptions
regarding the distributional structure or functional form of heterogeneity.\footnote{\citet*{otaspecification} develop a specification test of parametric
binary choice models via the maximum score estimator. Given that our
approach generalizes the maximum score idea to settings with multivariate
monotonicity and infinite-dimensional fixed effects, extending the
specification testing idea of \citet*{otaspecification} to our multivariate
monotone panel setting can be a promising avenue for future work.}
\subsection{\label{subsec:ID}Key Identification Strategy}
In this section, we present our main semiparametric identification
result for model \eqref{eq:Model_PMC} under Assumptions \ref{assu:PMC_Mono}
and \ref{assu:EpsDist}, and detail our key identification strategy,
which exploits multivariate monotonicity in the presence of additive
non-separability and nonparametric fixed effects.
To start, fix any subset of products $\tilde{{\cal J}}\subseteq{\cal J}$, a pair
of time periods $t\neq s\in\{1,\ldots,T\}$ and a generic realization
of observable covariates in the two periods $t$ and $s$, i.e., $({\bf X}_{it},{\bf X}_{is})=({\bf x}_{t},{\bf x}_{s})\in\text{Supp}({\bf X}_{it},{\bf X}_{is})$.
For simpler notation, we write ${\bf X}_{i,ts}=({\bf X}_{it},{\bf X}_{is})$,
${\bf x}_{ts}=({\bf x}_{t},{\bf x}_{s})$, $\delta_{jt}=x_{jt}^{'}\beta_{0}$,
where $x_{jt}$ denotes the $j$-th column of ${\bf x}_{t}$.
For each individual $i$, consider the intertemporal change in the
probability of that individual choosing any product $j\in\tilde{{\cal J}}$ across
periods $t$ and $s$, conditional on $({\bf X}_{i,ts},{\bf A}_{i})$,
the joint realization of the fixed effect and the observable covariates
in both periods $t$ and $s$. Formally, writing $y_{i\tilde{{\cal J}} t}=\sum_{j\in\tilde{{\cal J}}}y_{ijt}$,
we have
\begin{align}
& \mathbb{E}\left[\rest{y_{i\tilde{{\cal J}} t}-y_{i\tilde{{\cal J}} s}}{\bf X}_{i,ts}={\bf x}_{ts},{\bf A}_{i}\right]\nonumber \\
= & \int\mathbf{\mathbbm1}\left\{ \max_{j\in\tilde{{\cal J}}}u\left(\delta_{jt},A_{ij},\epsilon_{ij{\color{red}t}}\right)>\max_{k\notin\tilde{{\cal J}}}u\left(\delta_{kt},\,A_{ik},\epsilon_{ik{\color{red}t}}\right)\right\} \text{d}{\cal \mathbb{P}}\left(\rest{\bs{\epsilon}_{i{\color{red}t}}}{\bf X}_{i,ts}={\bf x}_{ts},{\bf A}_{i}\right)\nonumber \\
& -\int\mathbf{\mathbbm1}\left\{ \max_{j\in\tilde{{\cal J}}}u\left(\delta_{js},A_{ij},\epsilon_{ij{\color{red}s}}\right)>\max_{k\notin\tilde{{\cal J}}}u\left(\delta_{ks},\,A_{ik},\epsilon_{ik{\color{red}s}}\right)\right\} \text{d}{\cal \mathbb{P}}\left(\rest{\bs{\epsilon}_{i{\color{red}s}}}{\bf X}_{i,ts}={\bf x}_{ts},{\bf A}_{i}\right)\nonumber \\
= & \int\left[\begin{array}{c}
\mathbf{\mathbbm1}\left\{ \max_{j\in\tilde{{\cal J}}}u\left(\delta_{jt},A_{ij},\tilde{\epsilon}_{j}\right)>\max_{k\notin\tilde{{\cal J}}}u\left(\delta_{kt},A_{ik},\tilde{\epsilon}_{k}\right)\right\} \\
-\mathbf{\mathbbm1}\left\{ \max_{j\in\tilde{{\cal J}}}u\left(\delta_{js},A_{ij},\tilde{\epsilon}_{j}\right)>\max_{k\notin\tilde{{\cal J}}}u\left(\delta_{ks},A_{ik},\tilde{\epsilon}_{k}\right)\right\}
\end{array}\right]\text{d}{\cal \mathbb{P}}\left(\rest{\tilde{\bs{\epsilon}}}{\bf X}_{i,ts}={\bf x}_{ts},{\bf A}_{i}\right),\label{eq:dyXA}
\end{align}
where the last equality follows from Assumption \ref{assu:EpsDist}
(pairwise time homogeneity):
\[
\rest{\bs{\epsilon}_{i{\color{red}s}}\sim\bs{\epsilon}_{i{\color{red}t}}\sim\tilde{\bs{\epsilon}}}{\bf X}_{i,ts}={\bf x}_{ts},{\bf A}_{i}
\]
in which $\tilde{\bs{\epsilon}}$ is a name-holder random element with the
shared marginal distribution of $\bs{\epsilon}_{i{\color{red}t}}$ and $\bs{\epsilon}_{i{\color{red}s}}$
given $({\bf x}_{ts},{\bf A}_{i})$.
Since $u$ is weakly increasing in its index argument, the choice
indicator
\[
\mathbf{\mathbbm1}\left\{ \max_{j\in\tilde{{\cal J}}}u\left(\delta_{jt},A_{ij},\tilde{\epsilon}_{j}\right)>\max_{k\notin\tilde{{\cal J}}}u\left(\delta_{kt},A_{ik},\tilde{\epsilon}_{k}\right)\right\}
\]
is weakly increasing in the vector $(\delta_{jt})_{j\in{\cal \tilde{{\cal J}}}}$ and
decreasing in the remaining vector $(\delta_{kt})_{k\notin{\cal \tilde{J}}}$
for every possible realization of $\tilde{\bs{\epsilon}}$ and ${\bf A}_{i}$.
Hence, if
\begin{equation}
\delta_{jt}\leq\delta_{js},\ \forall j\in\tilde{{\cal J}}\quad\text{and}\quad\delta_{kt}\geq\delta_{ks}\ \forall k\notin\tilde{{\cal J}},\label{eq:dXb_jk}
\end{equation}
then, for every possible realization of $\tilde{\bs{\epsilon}}$ and ${\bf A}_{i}$,
we have
\begin{equation}
\left[\begin{array}{c}
\mathbf{\mathbbm1}\left\{ \max_{j\in\tilde{{\cal J}}}u\left(\delta_{jt},A_{ij},\tilde{\epsilon}_{j}\right)>\max_{k\notin\tilde{{\cal J}}}u\left(\delta_{kt},A_{ik},\tilde{\epsilon}_{k}\right)\right\} \\
-\mathbf{\mathbbm1}\left\{ \max_{j\in\tilde{{\cal J}}}u\left(\delta_{js},A_{ij},\tilde{\epsilon}_{j}\right)>\max_{k\notin\tilde{{\cal J}}}u\left(\delta_{ks},A_{ik},\tilde{\epsilon}_{k}\right)\right\}
\end{array}\right]\leq0,\label{eq:dyXAe>0}
\end{equation}
and consequently
\begin{equation}
\mathbb{E}\left[\rest{y_{i\tilde{{\cal J}} t}-y_{i\tilde{{\cal J}} s}}{\bf X}_{i,ts}={\bf x}_{ts},{\bf A}_{i}\right]\leq0\quad\text{for every possible realization of }{\bf A}_{i}.\label{eq:dyXA>0}
\end{equation}
Now, consider the observable intertemporal change in conditional choice
probabilities:
\begin{align}
\gamma_{\tilde{{\cal J}},ts}\left({\bf x}_{ts}\right) & :=\mathbb{E}\left[\rest{y_{i\tilde{{\cal J}} t}-y_{i\tilde{{\cal J}} s}}{\bf X}_{i,ts}={\bf x}_{ts}\right].\label{eq:dyX}
\end{align}
Then, whenever \eqref{eq:dXb_jk} holds, by \eqref{eq:dyXA>0} we
can deduce that
\begin{align*}
\gamma_{\tilde{{\cal J}},ts}\left({\bf x}_{ts}\right) & =\int\underset{\leq0}{\underbrace{\mathbb{E}\left[\rest{y_{i\tilde{{\cal J}} t}-y_{i\tilde{{\cal J}} s}}{\bf X}_{i,ts}={\bf x}_{ts},{\bf A}_{i}\right]}}\text{d}\mathbb{P}\left(\rest{{\bf A}_{i}}{\bf X}_{i,ts}={\bf x}_{ts}\right)\leq0.
\end{align*}
In summary, since \eqref{eq:dXb_jk} implies inequality \eqref{eq:dyXAe>0}
for every possible realization of $\tilde{\bs{\epsilon}}$ and ${\bf A}_{i}$,
this inequality will be preserved after $\tilde{\bs{\epsilon}}$ and ${\bf A}_{i}$
are integrated out \emph{cross-sectionally} with respect to the conditional
distribution $\mathbb{P}\left(\rest{\tilde{\bs{\epsilon}},{\bf A}_{i}}{\bf X}_{i,ts}={\bf x}_{ts}\right)$,
regardless of how complicated this unknown conditional distribution
may be.
The next proposition formalizes the identification strategy described
above, which produces an identifying restriction on the parameter
$\beta_{0}.$
\begin{prop}[Key Identifying Restrictions]
\label{prop:ID_rest} Under model \eqref{eq:Model_PMC} and Assumptions
\ref{assu:RandSamp}--\ref{assu:EpsDist},
\begin{equation}
\gamma_{\tilde{{\cal J}},ts}\left({\bf x}_{ts}\right)>0\ \text{\ensuremath{\Rightarrow}}\ \textup{NOT}\ \left\{ \left(x_{jt}-x_{js}\right)^{'}\beta_{0}\leq0,\ \forall j\in\tilde{{\cal J}}\ \mathrm{and}\ \left(x_{kt}-x_{ks}\right)^{'}\beta_{0}\geq0\ \forall k\notin\tilde{{\cal J}}\right\} \label{eq:ID_rest}
\end{equation}
for any (ordered) pair of time periods $(t,s)$ with $t\neq s\in\{1,\ldots,T\}$,
any subset of products $\tilde{{\cal J}}\subseteq{\cal J},$ and any realization
of observables ${\bf x}_{ts}\in\text{Supp}({\bf X}_{i,ts})$.
\end{prop}
\noindent Proposition \ref{prop:ID_rest} establishes an identifying
restriction on $\beta_{0}$ that is free of all unknown nonparametric
heterogeneity terms---$u$, ${\bf A}$, and $\bm{\epsilon}$---and holds
in the presence of additive non-separability and nonparametric fixed
effects. Proposition \ref{prop:ID_rest} is also intuitive: if the
total market share of the products in $\tilde{{\cal J}}$ increases between two
periods, then it cannot be the case that the indices of products in
$\tilde{{\cal J}}$ have all (weakly) worsened while those of products not in $\tilde{{\cal J}}$
have all (weakly) improved.
\begin{thm}[Identified Set]
\label{thm:SetID} Let $B_{0}$ be the set of all $\beta\in\mathbb{R}^{D}$
such that \eqref{eq:ID_rest} holds with $\beta$ in lieu of $\beta_{0}$,
for almost all ${\bf x}_{ts}\in\text{Supp}({\bf X}_{i,ts})$, all $t\neq s\in\{1,\ldots,T\}$,
and all $\tilde{{\cal J}}\subseteq{\cal J}$. Then, under model \eqref{eq:Model_PMC}
and Assumptions \ref{assu:RandSamp}--\ref{assu:EpsDist}, $\beta_{0}\in B_{0}.$
\end{thm}
\noindent We refer to $B_{0}$ as the \emph{identified set}. In Appendix
\ref{sec:App_PID}, we provide sufficient conditions for point identification
of $\beta_{0}$ up to a scale normalization. The assumptions we impose
are similar to those used for point identification in the maximum-score
literature, such as \citet*{manski1985semiparametric}, and in related
work on panel multinomial choice models, such as \citet*{shi2017estimating}
and \citet*{khan2021inference}.
Relative to the well-known maximum-score criterion function studied
by \citet{manski1985semiparametric,manski1987semiparametric} under
univariate monotonicity, our criterion function is non-standard. This
non-standardness arises from a key distinction between multivariate
and univariate monotonicity. To see this more clearly, consider the
special case of a \emph{single-index} setting $(J=1)$\footnote{This arises naturally in binomial choice models with the characteristics
of the outside option set to be zero. In this case, even though there
are nominally two choice alternatives, choice behavior is completely
determined by a single index based on the characteristics of the non-default
option.}, in which case the following equivalence relationship holds given
the \emph{univariate }monotonicity in the index:
\begin{equation}
\left\{ \gamma\left({\bf x}_{ts}\right)>0\right\} \ \Leftrightarrow\ \left\{ \left(x_{t}-x_{s}\right)^{'}\beta>0\right\} ,\label{eq:Equiv_MaxScore}
\end{equation}
Such an ``if-and-only-if'' relationship is a unique feature of the
single-index setting that \emph{cannot} be generalized to the multi-index
setting with $J\geq2$, as the right-hand side of \eqref{eq:ID_rest},
\[
\mathrm{NOT}\ \left\{ \left(x_{jt}-x_{js}\right)^{'}\beta_{0}\leq0,\ \forall j\in\tilde{{\cal J}}\ \mathrm{and}\ \left(x_{kt}-x_{ks}\right)^{'}\beta_{0}\geq0\ \forall k\notin\tilde{{\cal J}}\right\} ,
\]
\emph{does not} imply $\gamma_{\tilde{{\cal J}},ts}\left({\bf x}_{ts}\right)\geq0$
in the converse direction. This breaks the ``if-and-only-if'' relationship
that the maximum-score criterion function in \citet{manski1985semiparametric,manski1987semiparametric}
is built upon. Thus, the maximum-score estimator does not generalize
to multi-index settings. The lack of ``if-and-only-if'' relationship
in the multi-index setting leads to a key difference in the criterion
functions, and consequently a different estimation approach. Importantly,
while the original maximum score criterion and estimator cannot be
generalized to multi-index settings, our procedure can be applied
under a general econometric framework characterized by \emph{multi-index
single-crossing} conditions, which we introduce in Section \ref{sec:Ext_MMIM}.
\subsection{\label{subsec:Sharp}Two-Period Sharpness}
We now establish the sharpness of the identified set $B_{0}$ in a
two-period setting. This sharpness result can be interpreted as the\emph{
pairwise sharpness} of the identifying restrictions in \eqref{eq:ID_rest}:
for fixed periods $s$ and $t$ such that $s<t$, the inequality restrictions
in \eqref{eq:ID_rest} for $(s,t)$ and $(t,s)$ exhaust all the identifying
information available from the model (its specification and assumptions)
and from the distribution of the observable data in periods $s$ and
$t$.
\begin{thm}[Pairwise Sharpness]
\label{thm:Id_sharp} Under model \eqref{eq:Model_PMC} and Assumptions
\ref{assu:RandSamp}--\ref{assu:EpsDist}, $B_{0}$ is sharp for
$T=2$.
\end{thm}
\noindent The proof of Theorem \ref{thm:Id_sharp} exploits and generalizes
a corresponding result in \citet*{pakes2016moment}. Specifically,
\citet*{pakes2016moment} consider a specification where the utility
index $X_{ijt}^{'}\beta_{0}$ and an unobserved heterogeneity index $\lambda\left(A_{ijt},\epsilon_{ijt}\right)$
are additively separable, and establish the sharpness of their identification
result by showing the existence of a nonnegative solution to a system
of linear equations. Here, we consider a more general setup without
requiring additive separability and propose a correspondingly more
general identification argument. Nevertheless, we show that the sharpness
of our identification result under our more general setup can be reduced
to the nonnegative solvability of the same system of linear equations
in \citet*{pakes2016moment}. Hence, by the result in \citet*{pakes2016moment},
our identification result is sharp.
Admittedly, Theorem \ref{thm:Id_sharp} only establishes sharpness
in a two-period setting; however, it does not directly imply ``all-period''
sharpness for $T\geq3$. While the existence of an observationally
equivalent latent error distribution can be established for any realization
${\bf X}_{i,ts}$ and any pair of periods $\left(t,s\right)$ as in
the proof of Theorem \ref{thm:Id_sharp}, to establish the stronger
``all-period'' sharpness result, we would need to show in addition
that there exists an all-period joint distribution of latent errors
that matches all-period observed joint conditional choice probabilities
and satisfies the pairwise time homogeneity assumption. This appears
to be a technically cumbersome exercise, given that the pairwise time-homogeneity
assumption enters as an implicit aggregate restriction on the two-period
error distributions (with all other periods aggregated out). We thus
do not pursue ``all-period'' sharpness here, and only present the
sharpness result above as in \citet*{pakes2016moment}, which also
focuses on two-period (pairwise) sharpness.
\section{\label{sec:S_EstComp}Estimation and Computation}
\subsection{\label{subsec:Criterion}Formulation of Population Criterion Function}
We now propose a population criterion function that encodes the identifying
information in Proposition \ref{prop:ID_rest}. We represent the right-hand
side of \eqref{eq:ID_rest} in Boolean algebra by
\begin{align}
\lambda_{\tilde{{\cal J}}}\left({\bf x}_{ts};\beta\right) & :=\prod_{k=1}^{J}\mathbf{\mathbbm1}\left\{ \left(-1\right)^{\mathbf{\mathbbm1}\left\{ k\in\tilde{{\cal J}}\right\} }\left(x_{kt}-x_{ks}\right)^{'}{\color{blue}\beta}\geq0\right\} ,\label{eq:lambda}
\end{align}
where $\left(-1\right)^{\mathbf{\mathbbm1}\left\{ k\in\tilde{{\cal J}}\right\} }$ takes the
value $-1$ for $k\in\tilde{{\cal J}}$ and $1$ for $k\notin\tilde{{\cal J}}$. Therefore,
Proposition \ref{prop:ID_rest} can be written algebraically as: $\gamma_{\tilde{{\cal J}},t,s}\left({\bf x}_{ts}\right)>0$
implies $\lambda_{\tilde{{\cal J}}}\left({\bf x}_{ts};\beta_{0}\right)=0$ for any ${\bf x}_{ts}\in\text{Supp}\left({\bf X}_{i,ts}\right)$.
We now define the following criterion function by taking a cross-sectional
expectation over the random realization of ${\bf X}_{i,ts}$ and aggregating
over all subsets $\tilde{{\cal J}}\subseteq{\cal J}$:
\begin{align}
Q_{t,s}\left(\beta\right) & :=\sum_{\tilde{{\cal J}}\subseteq{\cal J}}\mathbb{E}\left[\mathbf{\mathbbm1}\left\{ \gamma_{\tilde{{\cal J}},t,s}\left({\bf X}_{i,ts}\right)>0\right\} \lambda_{\tilde{{\cal J}}}\left({\bf X}_{i,ts};\beta\right)\right],\label{eq:Q_ts}
\end{align}
which is nonnegative and minimized to zero at $\beta_{0}$. Without normalization
and further assumptions for point identification, there could be multiple
values of $\beta$ that minimize $Q_{t,s}$ to zero.
More generally, fix any function $G:\mathbb{R}\to\mathbb{R}$ that is \emph{one-sided
sign preserving}, i.e., $G\left(z\right)>0$ for $z>0$ and $G\left(z\right)=0$
for $z\leq0$. For example, we can choose $G\left(z\right)=\left[z\right]_{+}$
where $\left[z\right]_{+}$ is the positive part function. Then, we
define $Q_{t,s}^{G}$ as
\begin{align}
Q_{t,s}^{G}\left(\beta\right) & :=\sum_{\tilde{{\cal J}}\subseteq{\cal J}}\mathbb{E}\left[G\left(\gamma_{\tilde{{\cal J}},t,s}\left({\bf X}_{i,ts}\right)\right)\lambda_{\tilde{{\cal J}}}\left({\bf X}_{i,ts};\beta\right)\right],\label{eq:Q_G_jts}
\end{align}
which is also minimized to zero at $\beta_{0}$. The sign-preserving
function $G$, if further set to be monotone, continuous, or bounded,
serves as a \emph{smoothing} function that can improve the finite-sample
performance of our estimators. We provide more discussions on function
$G$ in the next section, when we construct estimators based on the
sample analog of the population criterion function defined here.
$Q_{t,s}^{G}$ above is defined for a fixed pair of periods $\left(t,s\right)$,
but in practice we may utilize the information across all pairs of
periods by defining the aggregated criterion function:
\begin{equation}
Q^{G}\left(\beta\right):=\sum_{t\neq s}^{T}Q_{t,s}^{G}\left(\beta\right),\quad\text{for any }\beta\in\mathbb{R}^{D}.\label{eq:PMC_BigQ}
\end{equation}
For notational simplicity, we suppress $G$ in $Q_{t,s}^{G}$ and
$Q^{G}$ in the rest of this paper.
\subsection{\label{subsec:Est}Two-Step Semiparametric Estimation}
We construct our estimator as a semiparametric two-step M-estimator
based on \eqref{eq:PMC_BigQ}. The first stage of our procedure is
concerned with nonparametrically estimating the intertemporal differences
in conditional choice probabilities of the following form:
\[
\gamma_{\tilde{{\cal J}},t,s}\left({\bf x}_{ts}\right)=\sum_{j\in\tilde{{\cal J}}}\gamma_{j,t,s}\left({\bf x}_{ts}\right),
\]
where $\gamma_{j,t,s}\left({\bf x}_{ts}\right)=\mathbb{E}\left[\rest{y_{ijt}-y_{ijs}}{\bf X}_{i,ts}={\bf x}_{ts}\right]$
can be separately estimated for each product $j\in{\cal J}.$\footnote{In practice, we only need to estimate $\gamma_{j,t,s}$ for $\left(J-1\right)$
products and $\frac{1}{2}T\left(T-1\right)$ \emph{ordered} pairs
of periods. The former is because conditional choice probabilities
must sum to one across all $J$ products. Hence, the estimator for
the last product from the other $\left(J-1\right)$ estimates can
be directly derived by $\gamma_{J,t,s}=-\sum_{j=1}^{J-1}\gamma_{j,t,s}$.
The latter is because $\gamma_{j,t,s}=-\gamma_{j,s,t}$ by construction, so
we may estimate it for either $\left(t,s\right)$ or $\left(s,t\right)$
pair. Notice, however, that each ordered pair $\left(t,s\right)$
or $\left(s,t\right)$ provides complementary identifying information,
as $\lambda\left({\bf X}_{i,ts};\beta\right)$ and $\lambda\left({\bf X}_{i,st};\beta\right)$
do not admit such kind of deterministic relationships.} We note that the first stage estimation includes the observable characteristics
of all products $J$. For example, when $J=3$ and $D=3$, there are
$3\times3\times2=18$ variables in the conditioned set of $\gamma$.
Given the potentially large number of regressors, one may want to
use neural networks \citep{bach2017breaking,chen1999improved} or
penalized sieves \citep{chen_2013} for the first-step estimation.
Given the first-stage estimators $\hat{\gamma}_{j,t,s}$ and the smoothing
function $G$, in the second stage we numerically compute minimizers
of the sample criterion function,
\begin{align*}
\hat{Q}\left(\beta\right):=\sum_{t\neq s}^{T}\hat{Q}_{\tilde{{\cal J}},t,s}\left(\beta\right), & \text{ where }\hat{Q}_{t,s}\left(\beta\right):=\frac{1}{N}\sum_{i=1}^{N}\sum_{\tilde{{\cal J}}\subseteq{\cal J}}G\left(\hat{\gamma}_{\tilde{{\cal J}},t,s}\left({\bf X}_{i,ts}\right)\right)\lambda_{\tilde{{\cal J}}}\left({\bf X}_{i,ts};\beta\right).
\end{align*}
It is worth noting that while $\hat{Q}_{t,s}(\beta)$ is defined as a
summation over all $2^{J}$ possible subsets $\tilde{{\cal J}}\subseteq{\cal J}$,
computationally there is no need to fully evaluate $Q_{\tilde{{\cal J}},t,s}\left(\beta\right)$
for each possible $\tilde{{\cal J}}\subseteq{\cal J}$ under a given $\beta$ when
the knife-edge cases of $(X_{ijt}-X_{ijs})^{'}\beta=0$ are ignorable.\footnote{There are at least two reasons why the ``knife-edge'' events of
the form $(X_{ijt}-X_{ijs})^{'}\beta=0$ should be ignored. First, $(X_{ijt}-X_{ijs})^{'}\beta=0$
technically occurs with probability zero for all $\beta$ provided that
$X_{ijt}\neq X_{ijs}$ almost surely, which is a natural assumption
given that we require $X_{ijt}$ be time-varying. Second, for programming
reasons, it is often practically necessary to ignore knife-edge strict
equalities of continuously valued variables, since such equalities
are extremely sensitive to unavoidable numerical errors induced by
the ``machine epsilon.''} This is because, as long as $\lambda_{\tilde{{\cal J}}}\left({\bf X}_{i,ts};\beta\right)=0$,
the contribution from the $\tilde{{\cal J}}$-summand would be zero. However, a
careful inspection of $\lambda_{\tilde{{\cal J}}}$ reveals that $\lambda_{\tilde{{\cal J}}}\left({\bf X}_{i,ts};\beta\right)=1$
only if $\tilde{{\cal J}}=\left\{ j\in{\cal J}:\left(X_{ijt}-X_{ijs}\right)^{'}\beta\leq0\right\} $.
Hence, in practical implementation we may simply compute
\[
\hat{Q}_{t,s}\left(\beta\right):=\frac{1}{N}\sum_{i=1}^{N}G\left(\sum_{j\in{\cal J}}\hat{\gamma}_{j,t,s}\left({\bf X}_{i,ts}\right)\mathbf{\mathbbm1}\left\{ \left(X_{ijt}-X_{ijs}\right)^{'}\beta\leq0\right\} \right).
\]
In addition, the scale of $\beta_{0}$ is not identified since $\lambda_{j}\left({\bf X}_{i,ts};\beta\right)$
consists of indicator functions of the form $\mathbf{\mathbbm1}\left\{ \left(X_{ijt}-X_{ijs}\right)^{'}\beta\geq0\right\} $.
Hence, we impose the scale normalization $\beta_{0}\in\mathbb{\mathbb{S}}^{D-1}:=\left\{ v\in\mathbb{R}^{D}:\norm v=1\right\} $.
Following \citet*{chernozhukov2007estimation}, we define the set
estimator by
\begin{equation}
\hat{B}_{\hat{c}}:=\left\{ \beta\in\mathbb{\mathbb{S}}^{D-1}:\ \hat{Q}\left(\beta\right)\leq\min_{\tilde{\beta}\in\mathbb{\mathbb{S}}^{D-1}}\hat{Q}\left(\tilde{\beta}\right)+\hat{c}\right\} \label{eq:Theta_hat_CHT07}
\end{equation}
with $\hat{c}:=O_{p}\left(c_{N}\log N\right)$.
We now introduce assumptions for establishing the consistency of $\hat{B}_{\hat{c}}$.
\begin{assumption}[First-Stage Estimation]
\label{assu:FS_Conv} For any $\left(j,t,s\right)$ tuple:
\begin{itemize}
\item[(i)] $\gamma_{j,t,s}\in\Gamma$, and $\mathbb{P}\left(\hat{\gamma}_{j,t,s}\in\Gamma\right)\to1$,
with $\Gamma$ being a $\mathbb{P}$-Donsker class of functions in $L_{2}\left({\bf X}\right)$.
\item[(ii)] $\norm{\hat{\gamma}_{j,t,s}-\gamma_{j,t,s}}_{2}:=\sqrt{\int\left(\hat{\gamma}_{j,t,s}\left({\bf X}_{i,ts}\right)-\gamma_{j,t,s}\left({\bf X}_{i,ts}\right)\right)^{2}\mathrm{d}\mathbb{P}\left({\bf X}_{i,ts}\right)}=O_{p}\left(c_{N}\right)$
\textup{with $c_{N}\searrow0$.}
\end{itemize}
\end{assumption}
\noindent Through Assumption \ref{assu:FS_Conv} we take as given
the large set of theoretical results on nonparametric regression in
the literature. Many kernel-based and sieve-based methods have been
developed, with their properties demonstrated under various sets of
conditions. See \citet*{wasserman2006all} and \citet*{chen2007sieve}
for more comprehensive surveys.
\begin{assumption}[Nice Smoothing Function]
\label{assu:NiceG} The one-sided sign-preserving function $G:\mathbb{R}\to\mathbb{R}_{+}$
is Lipschitz continuous with a finite Lipschitz constant.
\end{assumption}
\noindent Assumption \ref{assu:NiceG} is stronger than necessary
for consistency per se given that our identification result is valid
with any choice of the one-sided sign-preserving function $G$, nevertheless
we take $G$ to be Lipschitz to simplify the proof.
To state the next assumption, we decompose each row (corresponding
to each product) of ${\bf x}_{t}-{\bf x}_{s}$ as the product of its
norm and its \emph{direction}, i.e., ${\bf x}_{jt}-{\bf x}_{js}\equiv r_{j}\left({\bf x}_{t}-{\bf x}_{s}\right)v_{j}\left({\bf x}_{t}-{\bf x}_{s}\right)$,
where $r_{j}\left({\bf x}_{t}-{\bf x}_{s}\right):=\norm{{\bf x}_{jt}-{\bf x}_{js}}$,
and $v_{j}\left({\bf x}_{jt}-{\bf x}_{js}\right):=\left({\bf x}_{jt}-{\bf x}_{js}\right)/\norm{{\bf x}_{jt}-{\bf x}_{js}}$
if ${\bf x}_{jt}\neq{\bf x}_{js}$ while $v_{j}\left({\bf x}_{jt}-{\bf x}_{js}\right):={\bf 0}$
if ${\bf x}_{jt}={\bf x}_{js}$.
\begin{assumption}[Continuous Distribution of Directions]
\label{assu:NoMass} The marginal distribution of $v_{j}\left({\bf X}_{it}-{\bf X}_{is}\right)$
has no mass point except possibly at ${\bf 0}$ and is not supported
on any proper linear subspace of $\mathbb{R}^{D}$ for each $\left(j,t,s\right)$
tuple.
\end{assumption}
\noindent Assumption \ref{assu:NoMass} ensures the continuity of
the population criterion function. We note that Assumption \ref{assu:NoMass}
is mild: it essentially requires that the \emph{directions} of intertemporal
differences in observable characteristics are continuously distributed
on their own supports. In particular, this allows all but one dimensions
of observable characteristics to be discrete.
With the above assumptions imposed, we now establish the consistency
of our set estimator $\hat{B}_{\hat{c}}$, using the results in \citet*{chernozhukov2007estimation}.
\begin{thm}[Consistency]
\label{thm:Consistency}Under Assumptions \ref{assu:RandSamp}--\ref{assu:NoMass},
the set estimator $\hat{B}_{\hat{c}}$ is consistent in Hausdorff
distance: $d_{H}\left(\hat{B}_{\hat{c}},B_{0}\right)=o_{p}\left(1\right)$,
where $d_{H}\left(\hat{B}_{\hat{c}},B_{0}\right)=\max\left\{ \sup_{\beta\in\hat{B}_{\hat{c}}}\inf_{\tilde{\beta}\in B_{0}}\norm{\beta-\tilde{\beta}},\sup_{\beta\in B_{0}}\inf_{\tilde{\beta}\in\hat{B}_{\hat{c}}}\norm{\beta-\tilde{\beta}}\right\} .$
Furthermore, if $\beta_{0}$ is point-identified on $\mathbb{\mathbb{S}}^{D-1}$, $\norm{\hat{\beta}-\beta_{0}}=o_{p}\left(1\right)$
for any $\hat{\beta}\in\hat{B}_{\widehat{c}=0}$.
\end{thm}
\subsection{\label{subsubsec:Comp}Computation}
We now explain how we implement the semiparametric two-step estimation
procedure proposed above. Since the model's criterion function is
possibly non-convex, standard gradient-based optimizers are susceptible
to converging to local minima. To address this, we employ a multi-stage
adaptive-grid search algorithm that explores the parameter space more
robustly and aims to locate the global minimizer of the objective
function. The code for our computation algorithm is publicly available
on GitHub.\footnote{https://github.com/mingliecon/GL\_PMC}
\subsubsection*{Choice of the Smoothing Function $G$}
Besides the requirement of Lipschitz continuity in Assumption \ref{assu:NiceG},
in practice we take $G$ to be bounded from above by setting $G\left(z\right)=2\Phi\left(\left[z\right]_{+}\right)-1$,
where $\Phi$ is the standard normal CDF. We now motivate our choice
of $G$.
Recall that our identification strategy is based on the logical implication
of the event $\gamma_{\tilde{{\cal J}},t,s}\left({\bf x}_{ts}\right)>0$. Thus, for
identification purposes we are only interested in $\mathbf{\mathbbm1}\left\{ \gamma_{\tilde{{\cal J}},t,s}({\bf x}_{ts})>0\right\} $,
i.e., whether the event $\gamma_{\tilde{{\cal J}},t,s}\left({\bf x}_{ts}\right)>0$
occurs, but not in the exact magnitude of $\gamma_{\tilde{{\cal J}},t,s}\left({\bf x}_{ts}\right)$.
However, when $\gamma_{\tilde{{\cal J}},t,s}\left({\bf x}_{ts}\right)$ is close to
zero, the estimator $\hat{\gamma}_{\tilde{{\cal J}},t,s}\left({\bf x}_{ts}\right)$
is relatively more likely to have the wrong sign, so that the plug-in
estimator $\mathbf{\mathbbm1}\left\{ \hat{\gamma}_{\tilde{{\cal J}},t,s}\left({\bf x}_{ts}\right)>0\right\} $
may induce a large error of magnitude $1$. Hence, the smoothing by
$G$ helps down-weight the observations when $\hat{\gamma}_{\tilde{{\cal J}},t,s}\left({\bf x}_{ts}\right)$
is close to zero and shrinks the magnitude of possible errors.
On the other hand, when $\gamma_{\tilde{{\cal J}},t,s}({\bf x}_{ts})$ is positive
and large so that $\mathbf{\mathbbm1}\left\{ \gamma_{\tilde{{\cal J}},t,s}({\bf x}_{ts})>0\right\} $
can be estimated well, the magnitude of $\gamma_{\tilde{{\cal J}},t,s}({\bf x}_{ts})$
itself does not provide additional identifying information. By setting
$G$ to be bounded from above, we dampen the influence of large values
of $\gamma_{\tilde{{\cal J}},t,s}({\bf x}_{ts})$, so that the numerical minimization
of $\hat{Q}$ is less sensitive to potentially large but redundant
variations in $\hat{\gamma}_{\tilde{{\cal J}},t,s}({\bf x}_{ts})$.
\subsubsection*{Angle-Space Reparameterization of $\protect\mathbb{\mathbb{S}}^{D-1}$}
To minimize $\hat{Q}(\beta)$ over $\beta\in\mathbb{\mathbb{S}}^{D-1}$, we work with a reparameterization
of $\mathbb{\mathbb{S}}^{D-1}$ with $D-1$ angles in spherical coordinates.\footnote{The idea and the motivation for using the angle-space reparameterization
can also be found in \citet*{manski1986operational}, who however
use only one angle parameter.} Specifically, define the angle space $\Theta$ by
\begin{equation}
\Theta:=\left[-\pi,\pi\right)\times\left[-\frac{\pi}{2},\frac{\pi}{2}\right]^{D-2},\label{eq:Theta}
\end{equation}
and the transformation $\theta\longmapsto\beta(\theta)$ by standard spherical
coordinate transformation. We now instead solve the optimization of
$\hat{Q}(\beta(\theta))$ over $\Theta$, which we further equip with its natural
geodesic metric $\rho_{\Theta}\left(\theta,\tilde{\theta}\right):=\arccos\left(\beta\left(\theta\right){}^{'}\beta\left(\tilde{\theta}\right)\right)$.
Note that $\rho_{\Theta}\left(\theta,\tilde{\theta}\right)$ is strongly equivalent\footnote{Two metrics $d_{1}$ and $d_{2}$ defined on some nonempty set $X$
are\emph{ strongly equivalent} if and only if there exist positive
constants $c_{1}$ and $c_{2}$ such that $c_{1}d_{1}\left(x,y\right)\leq d_{2}\left(x,y\right)\leq c_{2}d_{1}\left(x,y\right)$
for every $x,y\in X$.} to the (imported) Euclidean distance $\norm{\beta\left(\theta\right)-\beta\left(\tilde{\theta}\right)}$.
This reparameterization $\left(\Theta,\rho_{\Theta}\right)$ enables us to
exploit the compactness and convexity of the parameter space $\Theta=\left[-\pi,\pi\right)\times\left[-\frac{\pi}{2},\frac{\pi}{2}\right]^{D-2}$,
which takes the form of a hyper-rectangle. First, $\left(\Theta,\rho_{\Theta}\right)$
preserves all topological structures of the unit sphere, and particularly
inherits the compactness of $\left(\mathbb{\mathbb{S}}^{D-1},\norm{\cdot}\right)$, automatically
satisfying the compactness condition usually imposed for extremum
estimation and making it numerically feasible to initiate a grid on
the whole parameter space. Second, while the unit sphere $\mathbb{\mathbb{S}}^{D-1}$
is not convex, the new parameter space $\Theta$ becomes convex algebraically,
making it computationally easy to define bisection points in the parameter
space. Third, $\left(\Theta,\rho_{\Theta}\right)$ preserves the geometric
structures of the sphere, including, for instance, the obvious observation
that $-\pi$ and $\pi$ in the first coordinate of $\Theta$ should be
treated as exactly the same point, or more rigorously, $\rho_{\Theta}\left(\left(\pi-\epsilon,\theta_{2},\ldots,\theta_{D-1}\right),\left(-\pi,\theta_{2},\ldots,\theta_{D-1}\right)\right)\to0$
as $\epsilon\to0$. This seemingly trivial property is nevertheless important
in defining and interpreting whether certain parameter estimates converge
asymptotically or not.
\subsubsection*{An Adaptive-Grid Algorithm}
With the angle reparameterization, we seek to numerically compute
a conservative rectangular enclosure of $\arg\min\hat{Q}\left(\theta\right)$,
deploying a bisection-style\emph{ }grid-search algorithm that recursively
shrinks and refines an \emph{adaptive grid} to any pre-chosen precision
(as defined by $\rho_{\Theta}$). Unlike gradient-based local optimization
algorithms, our adaptive grid algorithm handles the built-in discreteness
in our sample criterion function, whose derivative is zero almost
everywhere, while still maintaining global coverage over the entire
parameter space. While a brute-force global search algorithm is the
safest choice when the dimension of the product characteristics $D$
is relatively small, our adaptive-grid algorithm runs significantly
faster. The essential structure of our algorithm is laid out as follows.
\medskip{}
Step 1: Initialize a global grid $\Theta^{\left(1\right)}$ of some chosen
size $M_{0}^{D-1}$ on $\Theta$.
Step 2: Compute $\hat{Q}\left(\theta\right)$ for each $\theta\in\Theta^{\left(1\right)}$,
and select all points in $\Theta^{\left(1\right)}$ with a criterion value
below the $\alpha$th-quantile in $\hat{Q}\left(\Theta^{\left(1\right)}\right):=\left\{ \hat{Q}\left(\theta\right):\theta\in\Theta^{\left(1\right)}\right\} $
into
\begin{equation}
\ul{\Theta}^{\left(1\right)}:=\left\{ \theta\in\Theta^{\left(1\right)}:\ \hat{Q}\left(\theta\right)\leq\mathrm{quantile}_{\alpha}\left(\hat{Q}\left(\Theta^{\left(1\right)}\right)\right)\right\} .\label{eq:Theta_ul1}
\end{equation}
Step 3: Take the enclosing rectangle of $\ul{\Theta}^{\left(1\right)}$,
by defining $\ul{\theta}_{d}^{\left(1\right)}:=\mathrm{min}^{*}\ul{\Theta}_{d}^{\left(1\right)}$
and $\ol{\theta}_{d}^{\left(1\right)}:=\mathrm{max}^{*}\ul{\Theta}_{d}^{\left(1\right)},$
where $\ul{\Theta}_{d}^{\left(1\right)}:=\left\{ \theta_{d}:\theta\in\ul{\Theta}^{\left(1\right)}\right\} $
for each $d=1,\ldots,D-1$ and the operator $\mathrm{min}^{*}$ and
$\mathrm{max}^{*}$ have standard definitions of $\min$ and $\max$
except for the first dimension $d=1$. For the first dimension, it
is necessary to account for the underlying spherical geometry and
the periodicity of angles, i.e. $\theta_{1}+2\pi\equiv\theta_{1}$ and in
particular $-\pi\equiv\pi$. This, however, is largely a programming
nuisance: whenever $\ul{\Theta}_{1}^{\left(1\right)}\subsetneq\Theta_{1}^{\left(1\right)}$
crosses over at $-\pi$ and $\pi$, we can add $2\pi$ to every $\theta_{1}\in\ul{\Theta}_{1}^{\left(1\right)}$
and obtain lower and upper bounds of $\ul{\Theta}_{1}^{\left(1\right)}+2\pi$,
as illustrated in Figure \ref{fig:Algo}.
Step 4: We initialize a refined grid $\Theta^{\left(2\right)}$ on $\ol{\ul{\Theta}}^{\left(1\right)}:=\times_{d=1}^{D-1}\left[\ul{\theta}_{d}^{\left(1\right)},\ol{\theta}_{d}^{\left(1\right)}\right]$
of size $M_{0}^{D-1}$.
Step 5: Iterate until refinement stops (falls below a certain numerical
precision).
\medskip{}
Note that the above is simply a sketch of our algorithm: see Appendix
\ref{sec:app_GridSearch} and the documentation on GitHub for more
implementation details.\footnote{Our algorithm relies heavily on the compactness and convexity of the
angle space $\Theta$. Compactness allows us to start with a global grid
over the whole parameter space for initial evaluations of the sample
criterion function. At each step of recursion, the convexity of $\Theta$
enables us to conveniently refine the grid by separately cutting each
coordinate of $\ol{\ul{\Theta}}^{\left(m\right)}$ into smaller pieces
through simple division.} To be conservative, we add in buffers at each step of refinement,
keep track of both outer and inner boundaries of the lower-quantile
set $\ul{\Theta}^{\left(m\right)}$, and make sure that the minimizers
of the criterion functions at all computed points are indeed enclosed
by the set returned in the end. We find the current algorithm to be
conservative and to perform well in our simulations.
This multi-stage approach is designed to balance computational feasibility
with a robust search. The initial coarse search efficiently discards
large, suboptimal regions of the parameter space, while the subsequent
refinement and boundary identification stages provide a high-precision
estimate in the most promising area. However, the algorithm's performance
is inherently tied to the selection of tuning parameters, particularly
the initial grid size \texttt{M\_Step} and the quantile used for pruning,
which must be chosen carefully to ensure the global minimum is not
discarded prematurely. Furthermore, it can get computationally intensive
when the dimension of $\beta$ is high. We find that running 1,000
simulations for $D=3$ usually takes a few hours on modern computers,
while for $D=4$ it may take up to a day.
\section{\label{sec:Ext_MMIM}General Econometric Framework of Multi-Index
Single-Crossing Conditions}
Our key identification strategy, and consequently the associated estimation
method, apply more widely beyond panel multinomial choice models.
We now introduce a general econometric framework defined by \emph{multi-index
single-crossing} (MISC) conditions, and show how our proposed methods
can be exploited in a wide range of models nested under the MISC condition
framework.
Formally, let $\left(y_{i},X_{i}\right)_{i=1}^{n}$ be a random sample
of data with $X_{i}$ distributed on the support ${\cal X}\subseteq\mathbb{R}^{d_{x}}$
and $y_{i}$ distributed on ${\cal {\cal Y}}\subseteq\mathbb{R}^{d_{y}}$.
Let $h_{0}:{\cal X}\to\mathbb{R}$ denote a functional of the conditional
distribution of $y_{i}$ given $X_{i}$ that is directly identified
from data. For each of $j=1,\ldots,J\in\mathbb{N}$, let $\phi_{j}:{\cal X}\to\mathbb{R}^{d_{\theta_{j}}}$
be some known transformation of $X_{i}$, and define $W_{ij}:=\phi_{j}\left(X_{i}\right)$
with $W_{i}:=\left(W_{i1},\ldots,W_{iJ}\right)$. Let $\theta_{0j}\in\Theta_{j}\subseteq\mathbb{R}^{d_{\theta_{j}}}$
be an unknown finite-dimensional parameter and write $\theta_{0}:=\left(\theta_{01}^{'},\ldots,\theta_{0J}^{'}\right)^{'}\in\Theta:=\times_{j=1}^{J}\Theta_{j}$.
\begin{defn}[\emph{Multi-Index Single-Crossing Condition}]
\label{assu:Assum_Mono} We say that $\left(h_{0},\theta_{0}\right)$
satisfy the (\emph{weak})\emph{ multi-index single-crossing condition
}if, for any realization $x\in{\cal X}$ and $w=\phi\left(x\right)$,
\begin{align}
w_{j}^{'}\theta_{0j}\geq0,\ \forall j=1,\ldots,J\quad & \Rightarrow\quad h_{0}\left(x\right)\geq0,\nonumber \\
w_{j}^{'}\theta_{0j}\leq0,\ \forall j=1,\ldots,J\quad & \Rightarrow\quad h_{0}\left(x\right)\leq0.\label{eq:MISC}
\end{align}
The condition is said to be strict if the inequalities on the right-hand
side of \eqref{eq:MISC} are strict.
\end{defn}
\noindent In words, the MISC condition states that if all the $J$
parametric indices $w_{1}^{'}\theta_{01}$, $w_{2}^{'}\theta_{02}$, $\ldots$,
and $w_{J}^{'}\theta_{0J}$ are (weakly) positive, then the functional
$h_{0}$ must be (weakly) positive; if the $J$ indices are all zero,
then $h_{0}$ must be zero; if the $J$ indices are all negative,
then $h_{0}$ must be negative. Essentially, the MISC condition provides
a parsimonious way to semiparametrically model how the multiple economic
factors jointly affect a certain statistic of the relevant economic
outcome. The MISC condition basically requires that, if all the relevant
factors reach certain thresholds, then the outcome statistics must
also reach certain thresholds. Such requirements are often easy to
obtain in an economic or econometric model: while multiple factors
in an economic model may interact with each other in potentially complicated
manners and there might be many configurations of the factors that
lead to ambiguous theoretical predictions, there are also often simple
configurations that we understand reasonably well. Hence, the MISC
condition imposes only mild requirements on the underlying economic
or econometric model for the problem and thus provides a general framework
for semiparametric econometric analysis, in which most of the modeling
ingredients can be left nonparametric except for the parametric indices
that capture different economic factors in the problem.
Clearly, the panel multinomial choice model considered in previous
sections falls under the MISC condition framework. Specifically, focusing
on a pair of time periods $\left(t,s\right)$ and a particular product
$j_{0}$ for illustration, define $\theta_{0j}:=\beta_{0}$, $h_{0}\left({\bf X}_{i}\right):=\gamma_{j_{0},ts}\left({\bf X}_{i}\right)$,
$W_{ij_{0}}:=X_{ij_{0}t}-X_{ij_{0}s}$ and $W_{ij}:=-\left(X_{ijt}-X_{ijs}\right)$
for $j\neq j_{0}$. Then, the MISC condition \eqref{eq:MISC} is satisfied
under model \eqref{eq:Model_PMC} and Assumptions \ref{assu:RandSamp}--\ref{assu:EpsDist}.
We now provide a few more examples of models nested in the MISC condition
framework.
\begin{example}[Binary Choice with Awareness]
\label{exa:Bin_Aware} Consider the following binary choice model
\[
y_{i}=\mathbf{\mathbbm1}\left\{ X_{i1}^{'}\theta_{01}\geq u_{i}\right\} \cdot\mathbf{\mathbbm1}\left\{ X_{i2}^{'}\theta_{02}\geq v_{i}\right\}
\]
where $y_{i}$ denotes whether consumer $i$ purchases a certain product
or not, $X_{i1}$ denotes a vector of covariates that influences the
consumer's utility from a product, and $X_{i2}$ denotes a vector
of covariates that affects the consumer's awareness of the product
(e.g., advertising). Here, we have $J=2$, $X_{i}:=\left(X_{i1},X_{i2}\right)$,
$W_{i1}:=X_{i1}$, and $W_{i2}:=X_{i2}$. Define $h_{0}\left(x\right):=\mathbb{E}\left[\rest{y_{i}}X_{i}=x\right]-\frac{1}{4}.$
Then, under the conditional median restrictions $\text{med}\left(\rest{u_{i}}X_{i}\right)=\text{med}\left(\rest{v_{i}}X_{i}\right)=0$
and the conditional independence restriction $\rest{u_{i}\perp v_{i}}X_{i}$,
it is true that
\begin{align*}
X_{i1}^{'}\theta_{01}>0,\ X_{i2}^{'}\theta_{02}>0 & \quad\Rightarrow\quad h_{0}\left(X_{i}\right)>0,\\
X_{i1}^{'}\theta_{01}<0,\ X_{i2}^{'}\theta_{02}<0 & \quad\Rightarrow\quad h_{0}\left(X_{i}\right)<0,
\end{align*}
satisfying the MISC condition.
\end{example}
\begin{example}[Binary Choice with Endogeneity]
\label{exa:Bin_Choice_Endo} Consider the binary choice model
\[
Y_{i}=\mathbf{\mathbbm1}\left\{ W_{i}^{'}\beta_{0}\geq\epsilon_{i}\right\} ,
\]
and let one component of $W_{i}$, say, $W_{i1}$ be endogenous. Suppose
that there exists a vector of instrumental variables $Z_{i}$ and
define $\xi_{i}:=W_{i1}-Z_{i}^{'}\gamma_{0}$ as the residual from the
reduced-form linear projection of $W_{i1}$ on $Z_{i}$. Assume that
the endogeneity between $\epsilon_{i}$ and $W_{i1}$ is captured by the
following control function
\[
\text{med}\left(\rest{\epsilon_{i}}Z_{i},\xi_{i}\right)=\lambda\left(\alpha_{0}\xi_{i}\right),
\]
where $\lambda$ is an unknown increasing function with location normalization
$\lambda\left(0\right)=0$, and the sign parameter $\alpha_{0}\in\left\{ -1,1\right\} $
controls the direction of the monotonicity. The above can be viewed
as an adaptation of the binary choice model that combines the conditional
median restriction in \citet*{manski1975maximum} with the control
function approach in \citet*{blundell2004endogeneity}: here we only
impose the control function restriction on the conditional median
instead of the whole distribution as in \citet*{blundell2004endogeneity}.
Then, writing $\ol Z_{i}:=\left(W_{i1},Z_{i}\right)$ and $\ol{\gamma}_{0}:=\left(-\alpha_{0},\alpha_{0}\gamma_{0}\right)^{'}$,
we have
\begin{align*}
W_{i}^{'}\beta_{0}>0,\ \ol Z_{i}^{'}\ol{\gamma}_{0}>0\quad\Rightarrow\quad\mathbb{E}\left[\rest{Y_{i}}W_{i},Z_{i}\right] & >\frac{1}{2}
\end{align*}
and its ``$<$'' counterpart, which can be viewed as a MISC condition
with $K=2,$ $h_{0}\left(W_{i},Z_{i}\right):=\mathbb{E}\left[\rest{Y_{i}-\frac{1}{2}}W_{i},Z_{i}\right]$,
$\phi_{1}\left(W_{i},Z_{i}\right):=W_{i}$, $\phi_{2}\left(W_{i},Z_{i}\right):=\left(W_{i1},Z_{i}\right)$,
$\theta_{0,1}:=\beta_{0}$, and $\theta_{0,2}:=\ol{\gamma}_{0}$.
\end{example}
\begin{example}[Dyadic Network Formation]
\label{exa:NetForm} Consider the dyadic network formation model
of \citet*{gao2023logical}, which extends \citet{graham2017econometric}
to a semiparametric setting:
\begin{align*}
\mathbb{E}\left[\rest{y_{ij}}X_{i},X_{j},A_{i},A_{j}\right]\ =\ & \psi\left(w\left(X_{i},X_{j}\right)^{'}\theta_{0},A_{i},A_{j}\right).
\end{align*}
Here $y_{ij}$ is a binary outcome indicating whether individuals
$i$ and $j$ are linked in an undirected network, $X_{i}$ and $X_{j}$
are the individuals' observable covariates, $w(X_{i},X_{j})$ is a
known pairwise transformation of individual covariates (with the leading
example being $w_{h}\left(X_{i},X_{j}\right):=\left|X_{i,h}-X_{j,h}\right|$
for each coordinate $h=1,\ldots,d_{x}$), $A_{i}$ and $A_{j}$ are
unobserved individual degree heterogeneity terms, and $\psi:\mathbb{R}^{3}\to\mathbb{R}$
is an unknown function assumed to be increasing in all its three arguments.
Specifically, fixing a particular pair of individuals $(\ol i,\ol j)$
and two realizations $\ol x,\ul x$ of $X_{i}$, it can be shown that,
with
\[
\ol w:=w\left(x_{\ol j},\ol x\right)-w\left(x_{\ol i},\ol x\right),\quad\ul w:=w\left(x_{\ol i},\ul x\right)-w\left(x_{\ol j},\ul x\right),
\]
and
\begin{align*}
h_{0}\left(\ol x,\ul x\right):= & \max\left(0,\mathbb{E}\left[\rest{y_{\ol ik}-y_{\ol jk}}X_{k}=\ol x\right]\right)\mathbb{E}\left[\rest{y_{\ol ik}-y_{\ol jk}}X_{k}=\ul x\right]\\
& -\max\left(0,\mathbb{E}\left[\rest{y_{\ol jk}-y_{\ol ik}}X_{k}=\ol x\right]\right)\mathbb{E}\left[\rest{y_{\ol jk}-y_{\ol ik}}X_{k}=\ul x\right]
\end{align*}
the weak MISC condition is satisfied under mild conditions:
\begin{align*}
\ol w^{'}\theta_{0}>0,\ \ul w^{'}\theta_{0}>0\quad & \Rightarrow\quad h_{0}\left(\ol x,\ul x\right)\geq0,\\
\ol w^{'}\theta_{0}<0,\ \ul w^{'}\theta_{0}<0\quad & \Rightarrow\quad h_{0}\left(\ol x,\ul x\right)\leq0.
\end{align*}
\end{example}
\begin{example}[Censored Monotone Transformation Model with Endogeneity]
\label{exa:CensorEndo} The approach proposed in Example \ref{exa:Bin_Choice_Endo}
above can also be adapted to the following censored monotone transformation
model with endogeneity:
\[
Y_{i}=\max\left\{ \phi\left(W_{i}^{'}\beta_{0},\epsilon_{i}\right),0\right\} ,
\]
where one component of the observed covariates, $W_{i1}$, is endogenous,
and $\phi$ is an unknown bivariate increasing function. This model
generalizes the usual censored regression model $Y_{i}=\max\left\{ W_{i}^{'}\beta_{0}+\epsilon_{i},0\right\} $,
say, in \citet*{blundell2007censored}, by incorporating a flexible
unknown monotone transformation $\phi$ with non-additive error term.
Since $\beta_{0}$, $\epsilon_{i}$ and $\phi$ are all unknown, one may normalize
$\phi\left(0,0\right)=0$. By the equivariance of (conditional) quantiles
under monotone transformations, we have
\[
\text{med}\left(\rest{Y_{i}}W_{i},Z_{i}\right)=\max\left\{ \phi\left(W_{i}^{'}\beta_{0},\text{med}\left(\rest{\epsilon_{i}}W_{i},Z_{i}\right)\right),0\right\} .
\]
Similar to Example \ref{exa:Bin_Choice_Endo}, define $\xi_{i}:=W_{i1}-Z_{i}^{'}\gamma_{0}$
as the residual from the reduced-form linear projection of $W_{i1}$
on the instrumental variables $Z_{i}$, and assume that the endogeneity
between $\epsilon_{i}$ and $W_{i1}$ is captured by the control function
$\text{med}\left(\rest{\epsilon_{i}}W_{i},Z_{i}\right)=\text{med}\left(\rest{\epsilon_{i}}Z_{i},\xi_{i}\right)=\lambda\left(\alpha_{0}\xi_{i}\right)$,
where $\lambda$ is an increasing function with normalization $\lambda\left(0\right)=0$.
Writing $\ol Z_{i}:=\left(W_{i1},Z_{i}\right)$ and $\ol{\gamma}_{0}:=\left(\alpha_{0},-\alpha_{0}\gamma_{0}\right)^{'}$,
we have
\begin{align*}
W_{i}^{'}\beta_{0}>0,\ \ol Z_{i}^{'}\ol{\gamma}_{0}>0\quad\Rightarrow\quad\text{med}\left(\rest{Y_{i}}W_{i},Z_{i}\right) & >0\text{ and}\\
W_{i}^{'}\beta_{0}\leq0,\ \ol Z_{i}^{'}\ol{\gamma}_{0}\leq0\quad\Rightarrow\quad\text{med}\left(\rest{Y_{i}}W_{i},Z_{i}\right) & =0,
\end{align*}
which can be viewed as a MISC condition with a weak ``$\leq$''
side and $h_{0}\left(X_{i}\right):=\text{med}\left(\rest{Y_{i}}W_{i},Z_{i}\right)$
given by the conditional median function.
\end{example}
Based on the MISC conditions \eqref{eq:MISC}, we can again obtain
identifying restrictions by taking their logical contrapositions,
which can be encoded algebraically in a similar way as in \eqref{prop:ID_rest}
and \eqref{eq:Q_G_jts}. Specifically, let $G$ be a one-sided sign-preserving
function as in \eqref{eq:Q_G_jts} and define
\[
\lambda\left(W_{i};\theta\right):=\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ W_{ij}^{'}\theta_{j}\leq0\right\} .
\]
\begin{prop}
\label{prop:ID_Gen} Under condition \eqref{eq:MISC}, we have
\begin{align*}
h_{0}\left(X_{i}\right)>0 & \ \Rightarrow\ \text{NOT}\ \left\{ W_{ij}^{'}\theta_{j}\leq0\ \forall j\right\} ,\\
h_{0}\left(X_{i}\right)<0 & \ \Rightarrow\ \text{NOT}\ \left\{ W_{ij}^{'}\theta_{j}\geq0\ \forall j\right\} .
\end{align*}
Furthermore, with $Q\left(\theta\right):=Q_{+}\left(\theta\right)+Q_{-}\left(\theta\right)$
where
\begin{align*}
Q_{+}\left(\theta\right) & :=\mathbb{E}\left[G\left(h_{0}\left(X_{i}\right)\right)\lambda\left(W_{i};\theta\right)\right]\ \text{ and }\ Q_{-}\left(\theta\right):=\mathbb{E}\left[G\left(-h_{0}\left(X_{i}\right)\right)\lambda\left(-W_{i};\theta\right)\right],
\end{align*}
we have $Q\left(\theta\right)\geq Q\left(\theta_{0}\right)=0$.
\end{prop}
\noindent Proposition \ref{prop:ID_Gen} generalizes Theorem \ref{thm:SetID}.
Notice that Proposition \ref{prop:ID_Gen} applies to all functionals
$h_{0}$ of the conditional distribution of $y_{i}$ given ${\bf X}_{i}$
that satisfy the MISC conditions.
One could also proceed with the two-step estimation procedure described
in Section \ref{sec:S_EstComp}. Given a first-stage nonparametric
estimator $\hat{h}$ of $h_{0}$, we can estimate $\theta_{0}$ (or the
identified set) by minimizing the sample criterion $\hat{Q}\left(\theta\right):=\hat{Q}_{+}\left(\theta\right)+\hat{Q}_{-}\left(\theta\right)$
with
\[
\hat{Q}_{+}\left(\theta\right):=\frac{1}{n}\sum_{i=1}^{n}G\left(\hat{h}\left(X_{i}\right)\right)\lambda\left(W_{i};\theta\right)\ \text{ and }\ \widehat{Q}_{-}\left(\theta\right):=\frac{1}{n}\sum_{i=1}^{n}G\left(-\hat{h}\left(X_{i}\right)\right)\lambda\left(-W_{i};\theta\right).
\]
\section{\label{sec:Sim}Simulation}
We now switch back to the panel multinomial choice model introduced
in Section \ref{sec:PMC} and examine the finite sample performance
of our proposed estimator. For each DGP, we run $M=1,000$ simulations
of model \eqref{eq:Model_PMC} with the following utility specification:
\[
u\left(X_{ijt}^{'}\beta_{0},\,A_{ij},\,\epsilon_{ijt}\right)=A_{i0}\left(X_{ijt}^{'}\beta_{0}+A_{ij}\right)+\epsilon_{ijt},
\]
in which $A_{i0}$ is an unobserved scale fixed effect that captures
agent-level heteroskedasticity in utilities, and $A_{ij}$ is an unobserved
location shifter specific to each agent-product pair. The ability
to deal with nonlinear dependence caused by unobserved fixed effects
in a relatively robust way is a distinctive feature of our method
compared with existing approaches. To allow for such dependence, we
generate correlation between the observable characteristics ${\bf X}_{i}$
and the fixed effects ${\bf A}_{i}$ via a latent variable $Z$. We
draw $Z_{i}\sim\mathcal{N}$$\left(0,1\right)$ and let $A_{i2}=\left[Z_{i}\right]_{+}$.
We construct $X_{ijt,2}=W_{ijt}+Z_{i}$ with $W_{ijt}\sim\mathcal{N}\left(0,2J\right)$.
Thus, $X$ and $A$ are correlated via $Z$. The DGPs for the rest
of ${\bf A}$ and ${\bf X}$ are: $A_{i0}\sim\mathcal{U}\left[2,2.5\right]$,
$A_{i1}\equiv0$, $A_{ij}\sim\mathcal{U}\left[-0.25,0.25\right]$
for $j\geq3$, $X_{ijt,1}\sim\mathcal{U}\left[-1,1\right]$, $X_{ijt,d}\sim\mathcal{N}\left(0,1\right)$
for $d\geq3$. Furthermore, we set $\ol{\beta}_{0}=\left(2,1,\ldots,1\right)^{'}\in\mathbb{R}^{D}$
and $\beta_{0}=\ol{\beta}_{0}/\norm{\ol{\beta}_{0}}$, and draw $\epsilon_{ijt}\sim TIEV\left(0,1\right)$.
To summarize, for each of the $M=1,000$ simulations we first generate
$\left(\beta_{0},{\bf X}_{it},{\bf A}_{i},\boldsymbol{\epsilon}_{it}\right)$
for all $(i,t)$ pairs. Then, we calculate the individual choice ${\bf Y}$
matrix according to model \eqref{eq:Model_PMC}. Next, we compute
$\hat{\beta}$ from the simulated observable data of $\left({\bf X},{\bf Y}\right)$.
To obtain $\hat{\beta}$, we first use nonparametric regression with
second-order polynomial basis functions with $\ell_{1}$-regularization
and 10-fold cross validation to estimate $\gamma$. Then, we apply
the adaptive-grid algorithm detailed in Section \ref{subsubsec:Comp}.
Finally, we assess how well $\hat{\beta}$ performs compared with the
true value $\beta_{0}$.\footnote{In Appendix \ref{sec:App_ADDSIM}, we provide additional simulation
results. Specifically, we first present a graphical illustration of
the identified set $B_{0}$ based on the population criterion \eqref{eq:PMC_BigQ}.
Second, we inspect how our estimator performs without point identification.
Third, we vary $\left(D,J,T\right)$ to examine how robust our method
is against various simulation specifications. Lastly, we include a
simulation illustration of the robustness of our approach to the ``Blue-Bus/Red-Bus''
problem.}
\subsubsection*{Baseline Results}
For the baseline configuration, we set $N=10,000,\ D=3,\ J=3,\text{ and }T=2$.
Since in this case the conditions for the point identification are
satisfied, any point from the argmin set $\hat{B}_{b}:=\arg\min_{\beta\in\mathbb{\mathbb{S}}^{D-1}}\hat{Q}_{b}\left(\beta\right)$
is a consistent estimator of $\beta_{0}$ for each round of simulation
$b=1,\ldots,M$. Specifically, we define
\[
\hat{\beta}_{b,d}^{u}:=\max\hat{B}_{b,d},\quad\hat{\beta}_{b,d}^{l}:=\min\hat{B}_{b,d},\quad\text{and}\quad\hat{\beta}_{b,d}^{m}:=\frac{1}{2}\left(\hat{\beta}_{b,d}^{u}+\hat{\beta}_{b,d}^{l}\right),
\]
where $\hat{\beta}_{b,d}^{u}$, $\hat{\beta}_{b,d}^{l}$, and $\hat{\beta}_{b,d}^{m}$
represent the maximum, minimum, and middle point along dimension $d$
for each round of simulation $b$ of the argmin set $\hat{B}$, respectively.
\begin{table}
\caption{Baseline Performance\label{tab:Baseline-Estimation-Performance}}
\bigskip{}
\noindent \centering{}
\begin{tabular}{ccccc}
\toprule
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$ & $\beta_{0}=\left(0.82,0.41,0.41\right)^{'}$ & $\hat{\beta}_{1}$ & $\hat{\beta}_{2}$ & $\hat{\beta}_{3}$\tabularnewline
\midrule
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$mid bias & $\frac{1}{M}\sum_{b=1}^{M}\left(\hat{\beta}_{b,d}^{m}-\beta_{0,d}\right)$ & -0.0005 & -0.0003 & -0.0034\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$upper bias & $\frac{1}{M}\sum_{b=1}^{M}\left(\hat{\beta}_{b,d}^{u}-\beta_{0,d}\right)$ & 0.0075 & 0.0080 & 0.0083\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$lower bias & $\frac{1}{M}\sum_{b=1}^{M}\left(\hat{\beta}_{b,d}^{l}-\beta_{0,d}\right)$ & -0.0085 & -0.0086 & -0.0150\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$mean(u$-$l) & $\frac{1}{M}\sum_{b=1}^{M}\left(\hat{\beta}_{b,d}^{u}-\hat{\beta}_{b,d}^{l}\right)$ & 0.0160 & 0.0166 & 0.0233\tabularnewline
standard deviation & $\sqrt{\frac{1}{M}\sum_{b=1}^{M}\left(\hat{\beta}_{b,d}^{m}-\ol{\hat{\beta}_{d}^{m}}\right)^{2}}$ & 0.0299 & 0.0308 & 0.0431\tabularnewline
root MSE (by coordinate) & \textcolor{black}{$\left(\frac{1}{M}\sum_{b=1}^{M}\left(\hat{\beta}_{b,d}^{m}-\beta_{0,d}\right)^{2}\right)^{1/2}$} & 0.0288 & 0.0296 & 0.0417\tabularnewline
\midrule
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$root
MSE (whole vector)} & \textcolor{black}{$\left(\frac{1}{M}\sum_{b=1}^{M}\norm{\hat{\beta}_{b}^{m}-\beta_{0}}^{2}\right)^{1/2}$} & \multicolumn{3}{c}{\textcolor{black}{0.0587}}\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\begin{array}{c}
\text{mean norm}\\
\text{deviations (MND) }
\end{array}$ & \multirow{1}{*}{$\frac{1}{M}\sum_{b=1}^{M}\norm{\hat{\beta}_{b}^{m}-\beta_{0}}$} & \multicolumn{3}{c}{0.0511}\tabularnewline
\bottomrule
\end{tabular}
\end{table}
Table \ref{tab:Baseline-Estimation-Performance} summarizes our baseline
results. In the first row we use the middle point $\hat{\beta}^{m}$
along each dimension of $\hat{B}$ to calculate the bias. The biases
are very small across all three dimensions with a magnitude between
-0.0034 and -0.0005. The next two rows show the biases in estimating
$\beta_{0,d}$ using $\hat{\beta}_{d}^{u}$ and $\hat{\beta}_{d}^{l}$ respectively,
which are again close to zero. The fourth row reports the average
widths of the set $\hat{B}$ along each dimension. These widths are
small relative to the magnitude of $\beta_{0}$. The fifth and sixth
rows summarize the standard deviation and rMSE for each coordinate
of $\widehat{\beta}^{m}$. In the second part of Table \ref{tab:Baseline-Estimation-Performance},
we report the vector rMSE and MND based on $\hat{\beta}^{m}$, and the
results suggest that our method performs well.
\subsubsection*{Results Varying $N$}
Next, we vary $N$ while maintaining $D=3,\ J=3,\ \text{and }T=2$
to assess how our method performs under different sample sizes. In
addition to the baseline setup with $N=10,000$, we calculate mean
absolute deviation (MAD), average size of the estimated set, rMSE,
and MND for $N=4,000$ and $N=1,000$. Results are summarized in Table
\ref{tab:PerformanceVaryingN}.
\begin{table}
\caption{Performance under Varying $N$\label{tab:PerformanceVaryingN}}
\bigskip{}
\noindent \centering{}
\begin{tabular}{ccccc}
\toprule
\multirow{2}{*}{} & \multirow{2}{*}{$\sum_{d}\left|\text{bias}_{d}\right|$} & \multirow{2}{*}{$\sum_{d}\text{mean(u-l)}_{d}$} & \multirow{2}{*}{\textcolor{black}{$\text{rMSE}$}} & \multirow{2}{*}{MND}\tabularnewline
& & & & \tabularnewline
\midrule
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$$N=10,000$ & 0.0042 & 0.0560 & 0.0587 & 0.0511\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$$N=\ 4,000$ & 0.0136 & 0.0865 & \textcolor{black}{0.0742} & 0.0650\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$$N=\ 1,000$ & 0.0606 & 0.1664 & \textcolor{black}{0.1369} & 0.1159\tabularnewline
\midrule
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$ & $\left(\dfrac{N}{1,000}\right)^{1/2}$ & $\left(\dfrac{N}{1,000}\right)^{1/3}$ & $\dfrac{\text{rMSE}_{1000}}{\text{rMSE}_{N}}$ & $\dfrac{\text{MND}_{1000}}{\text{MND}_{N}}$\tabularnewline
\midrule
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$\textcolor{blue}{${\color{black}N=10,000}$} & 3.16 & 2.15 & 2.33 & 2.27\tabularnewline
$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$$N=\ 4,000$ & 2.00 & 1.59 & 1.84 & 1.78\tabularnewline
\bottomrule
\end{tabular}
\end{table}
Table \ref{tab:PerformanceVaryingN} provides numerical evidence that
a larger $N$ helps with overall performance. The sum of absolute
bias decreases from 0.0606 to 0.0042 when $N$ increases from $1,000$
to $10,000$. The average size of the estimated sets, rMSE, and MND
follow a similar pattern. Notably, even with a relatively small $N=1,000$,
the results remain informative and reasonably accurate, with the rMSE
and MND equal to 0.1369 and 0.1159, respectively. We note that $T$
is set to 2 here, which is the minimum required for our method to
work. Since our method can extract information from each of the $T\left(T-1\right)$
ordered pairs of time periods, a larger $T$ would generally improve
the performance of our estimators. Appendix \ref{sec:App_ADDSIM}
presents additional simulation results for a larger $T$.
Finally, we numerically investigate the speed of convergence when
we increase $N$ from $1,000$ to $4,000$ and $10,000$ in the second
part of Table \ref{tab:PerformanceVaryingN}. Compared with the case
of $N_{0}=1,000$, the relative ratios of rMSE are 1.84 for $N=4,000$
and 2.33 for $N=10,000$, both of which lie between $\left(N/N_{0}\right)^{1/3}$
and $\left(N/N_{0}\right)^{1/2}$. A similar pattern is also observed
for calculations based on MND. These results suggest that our estimator
converges at a rate slower than $N^{-1/2}$ but faster than $N^{-1/3}$.
\section{\label{sec:Emp}Empirical Application}
\subsection{\label{subsec:Data}Data and Methodology}
We now present an empirical application of the panel multinomial choice
model and our proposed estimation method, using NielsenIQ Retail Scanner
Data on popcorn sales to examine the effects of display promotions.
The data contain weekly store-level information on prices, sales,
and display promotion status, collected from approximately 35,000
participating retail stores with point-of-sale systems across the
United States.
We focus on popcorn among the wide array of products for two reasons.
First, popcorn purchases are more likely to be impulsive, with limited
intertemporal planning. Second, popcorn exhibits substantial variation
in in-store display promotions, allowing us to estimate how special
displays influence consumers\textquoteright{} purchase decisions.
We aggregate store-level observations to the designated market area
(DMA) level ($N=205$) for 2015. We focus on the top three brands
by market share, pool the remaining brands into a fourth product (``all
other products''), and include an outside option of ``no purchase.''
We compute market shares as the dependent variable for each of the
$J=5$ alternatives---the three leading brands, the ``all other products''
category, and the outside option. The observed product characteristics
include price, display-promotion status, and their interaction.\footnote{We define $\text{Price}_{cjt}$ as the weighted-average unit price
of all UPCs of the brand $j$ in DMA $c$ during week $t$. The dataset
includes two promotion indicators: display and feature. Given their
similarity, we construct $\text{Promo}_{cjt}$ as $($feature$\lor$display$)_{cjt}$.
The interaction term $\text{Price}_{cjt}\times\text{Promo}_{cjt}$
is included in $X$ to allow price sensitivity (elasticity) to vary
under promotion.} Notationally, $c$ denotes each of the $N=205$ DMAs, $j$ represents
each of the $J=5$ brands, and $t$ indexes the $T=52$ weeks in 2015.
The summary statistics of these variables are provided in Table \ref{tab:Empirical-Application:-Summary}.
\begin{table}
\caption{Empirical Application: Summary Statistics\label{tab:Empirical-Application:-Summary}}
\bigskip{}
\noindent \centering{}
\begin{tabular}{ccccc}
\toprule
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$} & mean & s.d. & min & max\tabularnewline
\midrule
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\text{DMA-level\ Market\ Share}$
$s_{cjt}$ & $25.06\%$ & $21.59\%$ & $0.08\%$ & $96.69\%$\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\text{Price}$$_{cjt}$ & 0.4924 & 0.1803 & 0.1094 & 1.3587\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\text{Promo}_{cjt}$ & 0.0282 & 0.0377 & 0.0000 & 0.5000\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\text{Price}{}_{cjt}$
$\times$ $\text{Promo}_{cjt}$ & 0.0136 & 0.0203 & 0.0000 & 0.4505\tabularnewline
\bottomrule
\end{tabular}
\end{table}
Since the data are at the DMA level, while our approach was originally
developed for individual-level data, we now describe how to adapt
the method to the DMA-level setting. We treat the observed DMA-level
market shares $s_{cjt}$ as noisy measurements\footnote{Alternatively, with market-level data, one could treat the observed
$s_{cjt}$ as a sufficiently good approximation of $\mathbb{E}\left[\rest{y_{cjt}}{\bf X}_{ct},{\bf A}_{c}\right]$
as in Section 6.1 of \citet*{shi2017estimating}, in which case our
first-stage nonparametric regression is \emph{no longer} required.
We did not pursue this approach for two reasons: First, we do not
wish to impose the assumption that the observed market shares are
measured with negligible errors. Second, we intend this empirical
exercise as an illustration of our two-stage procedure, and thus focus
on a setting where the first-stage nonparametric regression is required.} of $\mathbb{E}\left[\rest{y_{cjt}}{\bf X}_{ct},{\bf A}_{c}\right]$, i.e.,
\[
s_{cjt}=\mathbb{E}\left[\rest{y_{cjt}}{\bf X}_{ct},{\bf A}_{c}\right]+u_{cjt},\quad\text{with }\mathbb{E}\left[\rest{u_{c}}{\bf X}_{c},{\bf A}_{c}\right]=0.
\]
Then, we use $s_{cjt}$ to nonparametrically estimate the following
intertemporal difference:
\[
\mathbb{E}\left[\rest{s_{cjt}-s_{cjs}}{\bf X}_{c,ts}\right]=\int\left(\mathbb{E}\left[\rest{y_{cjt}}{\bf X}_{ct},{\bf A}_{c}\right]-\mathbb{E}\left[\rest{y_{cjs}}{\bf X}_{cs},{\bf A}_{c}\right]\right)d\mathbb{P}\left(\rest{{\bf A}_{c}}{\bf X}_{c,ts}\right).
\]
Specifically, we nonparametrically regress $\left(s_{cjt}-s_{cjs}\right)$
on the second-order polynomial basis functions of $\mathbf{X}_{c,ts}$
with $\ell_{1}$-regularization and 10-fold cross-validation to obtain
an estimator $\hat{\gamma}_{j}$ of $\gamma_{j}\left(\overline{{\bf X}},\underline{{\bf X}}\right):=\mathbb{E}\left[\rest{s_{cjt}-s_{cjs}}{\bf X}_{c,ts}=\left(\overline{{\bf X}},\underline{{\bf X}}\right)\right]$.
Finally, we plug $\hat{\gamma}$ into our second-stage algorithm and compute
the (approximate) argmin set $\hat{B}_{\hat{c}}$.
\subsection{Results and Discussion\label{subsec:emp-Results-and-Discussion}}
We report our estimation results in Table \ref{tab:Comp}. $\hat{\beta}_{\hat{c}}^{m}:=\frac{1}{2}\left(\hat{\beta}_{\hat{c}}^{l}+\hat{\beta}_{\hat{c}}^{u}\right)$
corresponds to the middle point of the (approximate) argmin set $\hat{B}_{\hat{c}}$
using our method. We show both the exact argmin set ($\hat{c}=0$)
and the approximate argmin set with $\hat{c}=0.1\times N^{-\frac{1}{4}}\log\left(N\right)\approx0.14$
for $N=205$. The estimated coefficients for Price (negative) and
Promo (positive) are economically intuitive.
The most interesting result is the positive estimated coefficient
on the interaction term $\text{Price}{}_{cjt}$ $\times$ $\text{Promo}_{cjt}$.
An intuitive explanation for the positive sign is that by displaying
certain products in front rows, consumers no longer see their price
tags adjacent to those of their competitors, and thus become less
price-sensitive for these specially promoted products.
\begin{table}
\caption{Empirical Illustration: Comparison of Results\label{tab:Comp}}
\bigskip{}
\noindent \centering{}
\begin{tabular}{cccccccc}
\toprule
\multirow{2}{*}{\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}} & \multirow{2}{*}{$\hat{\beta}_{\hat{c}=0}^{m}$} & \multirow{2}{*}{$\hat{\beta}_{\hat{c}=0.14}^{m}$} & \multirow{2}{*}{$\hat{\beta}^{CyclicMono}$} & \multirow{2}{*}{$\hat{\beta}^{OLS}$} & \multirow{2}{*}{$\hat{\beta}^{OLS-FE}$} & \multirow{2}{*}{$\hat{\beta}^{MLogit-FE}$} & \multirow{2}{*}{$\hat{\beta}^{RCLM}$}\tabularnewline
& & & & & & & \tabularnewline
\midrule
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\text{Price}$$_{cjt}$ & -0.9351 & -0.9283 & $\text{-}0.3781$ & $0.0240$ & $\text{-}0.3807$ & $\text{-}0.6249$ & -0.9705\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\text{Promo}_{cjt}$ & 0.1793 & 0.1912 & $\text{-}0.0567$ & $0.5760$ & $0.5976$ & $0.5881$ & 0.2157\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\text{Price}{}_{cjt}$
$\times$ $\text{Promo}_{cjt}$ & 0.2687 & 0.2505 & $0.9240$ & $\text{-}0.8171$ & $\text{-}0.7057$ & $\text{-}0.5135$ & 0.1078\tabularnewline
\bottomrule
\end{tabular}
\end{table}
Furthermore, we compare our $\hat{\beta}^{m}$ with the estimates obtained
through four other methods, i.e., Cyclic Monotonicity (CM) based on
\citet*{shi2017estimating}\footnote{We use 2-week cycles for all available weeks in the data for the CM
method.}, OLS, OLS with scalar-valued fixed effects (OLS-FE), the multinomial
logit with fixed effects (MLogit-FE), and the random coefficients
logit model (RCLM)\footnote{See Appendix \ref{subsec:Blue-Bus/Red-Bus} for the details of the
RCLM estimator.}. Results (normalized to $\mathbb{\mathbb{S}}^{D-1}$) are summarized in Table \ref{tab:Comp}.
The OLS estimator for $\text{Price}$ is a positive 0.0240, which
is counterintuitive. Moreover, displaying the product in the front
rows of the store likely makes consumers less price-sensitive, suggesting
a positive coefficient on Price$\times$Promo. However, the estimated
coefficients for the interaction term using OLS, OLS-FE, and MLogit-FE
are all negative. Next, the CM-based estimator for the coefficient
of Promo is negative at -0.0567, whereas the estimated coefficient
on $\text{Price}\times\text{Promo}$ is a large positive 0.9240. While
the aggregate effect of Promo is likely to be positive for most prices
observed in the data, it makes the coefficient of Price positive for
those promoted products (i.e., $\widehat{\beta}_{Price}+\text{Promo}\times\widehat{\beta}_{Price\times Promo}>0$
when Promo = 1). Finally, the estimates from RCLM have the same sign
as our method. Nonetheless, it reports a smaller estimated coefficient
for the interaction term $\text{Price}{}_{cjt}\times\text{Promo}_{cjt}$,
making the effect from Promo on alleviating price sensitivity less
significant.
We view the contrast between our findings and those from alternative
methods as empirical evidence that, by accommodating more flexible
forms of unobserved heterogeneity---via high-dimensional fixed effects
that enter consumers\textquoteright{} utility functions in an additively
nonseparable manner---our approach yields more economically plausible
results.
\subsection{A Possible Explanation via Monte-Carlo Simulations}
In this subsection, we provide a possible explanation for the empirical
findings reported in Table \ref{tab:Comp} through simulation analysis.
Recall that ``Promo'' indicates whether a product receives increased
in-store exposure by being highlighted by the store. We argue that
the negative estimates on $\text{Price}{}_{ijt}$ $\times$ $\text{Promo}_{ijt}$
reported for traditional methods in Table \ref{tab:Comp} likely arise
from a positive correlation between display promotions and an unobserved
index of price sensitivity.
Specifically, suppose the utility function is
\begin{equation}
u_{ijt}=A_{ij}\times\left(X_{ijt}^{'}\beta_{0}\right)+\epsilon_{ijt},\label{eq:mcemp_uijt}
\end{equation}
where $X_{ijt}$ contains Price, Promo, and Price$\times$Promo, $A_{ij}$
is the $ij$-specific fixed effect which may capture index sensitivity
(which can be thought as inversely related to unobserved brand loyalty),
and $\epsilon_{ijt}$ is the exogenous random shock. Suppose $A_{ij}$
and Promo$_{ijt}$ are positively correlated, which is reasonable
because marketing managers with their expertise are more likely to
promote products to which consumers are more price- and promotion-sensitive.
Thus, traditional estimation methods based on linearity would be unable
to detect such a pattern and wrongly attribute the effect on price
elasticities from $A_{ij}$ to Promo.
To provide some numerical evidence of the claim, we run the following
Monte Carlo simulation. We set $\beta_{0}=\left(-4,2,2\right)^{'}$,
$Z_{ij}\sim\mathcal{U}\left[0,1\right]$, $A_{ij}=Z_{ij}+1$, and
$\epsilon_{ijt}\sim TIEV\left(0,1\right)$. For the $X_{ijt}$ vector,
we draw $X_{ijt,1}\sim\mathcal{U}\left[0,4\right]\text{ and }W_{ijt}\sim\mathcal{U}\left[0,1\right],$
and let $X_{ijt,2}=\left(1-\alpha\right)\times W_{ijt}+\alpha\times Z_{ij}$
and $X_{ijt,3}=X_{ijt,1}\times X_{ijt,2}$. We emphasize that $X_{ijt,2}$
(Promo) is positively correlated with $A_{ij}$ through $Z_{ij}$,
with $\alpha$ measuring the strength of the correlation. We consider
three values of $\alpha$: $0.15,\ 0.3,\text{ and }0.5$.
We run 1,000 simulations for each of the five methods in Table \ref{tab:Comp}
to estimate $\beta_{0}$. To replicate the data structure of the empirical
exercise, we set $N=205$, $D=3$, $J=4$, and $T=10$. We report
in Table \ref{tab:MCEmp} the percentage of simulations that the corresponding
method produces correct signs for all coordinates of $X_{ijt}$.
\begin{table}
\caption{Percentage of Correct Signs of Estimated Coefficients\label{tab:MCEmp}}
\bigskip{}
\noindent \centering{}
\begin{tabular}{>{\centering}p{1.8cm}>{\centering}p{1.8cm}>{\centering}p{1.8cm}>{\centering}p{1.8cm}>{\centering}p{1.8cm}>{\centering}p{1.8cm}>{\centering}p{1.8cm}}
\toprule
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}$\alpha$ & $\hat{\beta}^{m}$ & $\hat{\beta}^{CyclicMono}$ & $\hat{\beta}^{OLS}$ & $\hat{\beta}^{OLS-FE}$ & $\hat{\beta}^{MLogit-FE}$ & $\hat{\beta}^{RCLM}$\tabularnewline
\midrule
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}0.15 & 91.50\% & 0.00\% & 0.00\% & 0.00\% & 28.00\% & 0.00\%\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}0.30 & 85.90\% & 0.00\% & 0.00\% & 0.00\% & 0.20\% & 0.00\%\tabularnewline
\textcolor{black}{$\phantom{\frac{\frac{1}{1}}{\frac{1}{1}}}$}0.50 & 74.20\% & 0.00\% & 0.00\% & 0.00\% & 0.00\% & 0.00\%\tabularnewline
\bottomrule
\end{tabular}
\end{table}
The percentages of simulations where our proposed method produces
correct signs for all coordinates of $X_{ijt}$ for $\alpha=0.15,\ 0.3,$
and $0.5$ are 91.50\%, 85.90\%, and 74.20\%, respectively. The accuracy
of the estimator is negatively affected by the correlation between
$X_{ijt,2}$ (Promo) and $A_{ij}$ (multiplicative fixed effect).
In contrast, none of the other methods in Table \ref{tab:MCEmp} generates
correct signs as ours does. The alternative models, owing to their
additively separable structure,\footnote{We note that the CM method requires $A_{ij}$ entering the utility
function linearly, which is violated in \eqref{eq:mcemp_uijt}.} may overlook the positive dependence between Promo and the multiplicative
fixed effect $A_{ij}$, which can bias the resulting estimates.\footnote{Notably, RCLM produces \textquotedblleft wrong signs\textquotedblright{}
in this simulation exercise, even though it yields the expected signs
in the empirical application in Section \ref{subsec:emp-Results-and-Discussion}.
A plausible interpretation is that, while RCLM is more flexible than,
for example, MLogit-FE, the unknown selection effect in this dataset
may be insufficiently strong to cause RCLM to fail, yet strong enough
for MLogit-FE and related methods to do so. This observation highlights
the potential value of our method as a robustness-check tool.}
Intuitively, since products with larger $A_{ij}$ are more likely
to be promoted $\left(X_{ijt,2}=1\right)$ by the selection of marketing
managers, the average effective price sensitivity of promoted products
tends to be greater in magnitude than that of non-promoted products.
This drives those estimators that ignore such selection effects to
produce a negative coefficient on the interaction term. In contrast,
our method handles such \emph{non-additive} dependence between observable
characteristics and unobserved fixed effects well, illustrating its
robustness in these models.
\section{\label{sec:Conclusion}Conclusion}
\noindent This paper develops a method for semiparametric identification
and estimation in panel multinomial choice models that feature infinite-dimensional
fixed effects and nonadditive utility, thereby accommodating rich
forms of unobserved heterogeneity. We also introduce a general identification
strategy based on multivariate monotonicity of parametric indices,
applicable to a broad class of econometric models defined by the MISC
conditions. In addition, we present a computational algorithm that
leverages angle-space reparameterization and adaptive-grid search,
which prove effective given the nonstandard criterion function implied
by our identifying restrictions.
Future research could investigate how our approach might be applied
and adapted to other specific microeconometric models within the MISC
framework. Along this line, \citet*{gao2023logical}---a companion
paper to the present study---illustrates how the approach proposed
here can be adapted to the context of dyadic network formation models.
\citet{gao2023identification} propose a method for addressing endogeneity
in discrete choice models under an adapted time-homogeneity condition.
In ongoing work, we are investigating how to combine techniques developed
in this line of work to analyze strategic network formation models
with endogenous covariates.
\bibliographystyle{ecta}
\phantomsection\addcontentsline{toc}{section}{\refname}\bibliography{GL_PMC1}
\newpage{}