EconBase
← Back to paper

Conditional-Moment Estimation and Inference in the BLP Model

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

82,073 characters

Conditional-Moment Estimation and Inference in the BLP Model







\maketitle
\doublespacing

\begin{abstract}
The random-coefficient demand model of Berry, Levinsohn, and Pakes (1995) is commonly estimated by the generalized method of moments (GMM), using an unconditional moment restriction with a fixed set of instruments. Identification of the model, however, rests on a conditional moment restriction. The two are not equivalent: the unconditional restriction may admit additional parameter values. We construct a counterexample in which the model is identified by the conditional restriction yet standard GMM is not, even with the optimal instrument. Building directly on the identifying restriction, we propose a two-step estimator, following \cite{ai2003efficient}, that first estimates the relevant conditional expectations nonparametrically and then selects the structural parameters by a conditional-variance-weighted minimum-distance criterion; standard GMM is recovered as the special case of a linear projection onto finitely many instruments. We establish $\sqrt T$-asymptotic normality for the proposed estimator, and we develop the theory for both kernel and series implementations of the first stage. The two implementations share a common limiting distribution, attaining the semiparametric efficiency bound. Simulation evidence illustrates the consequences of the identification gap and demonstrates that the proposed estimator outperforms standard GMM in finite samples.

\end{abstract}

\thispagestyle{empty}
\clearpage


\setcounter{page}{1}

\section{Introduction}
{
In demand estimation for differentiated products, the \citet*{berry1995automobile} or BLP approach has become the workhorse framework of empirical industrial organization (IO). \cite{nevo2000mergers,nevo2001measuring}, for instance, examines competition and merger effects in the ready-to-eat cereal market, \cite{petrin2002quantifying} measures consumer welfare generated by the minivan's entry, \cite{berry1999voluntary} assess how voluntary export restraints reshaped the U.S. auto market, and \cite{goldberg2001evolution} document price dispersion across European car markets, among many other applications.

However, estimation and identification in this literature rest on different moment conditions. The standard BLP estimation approach uses the generalized method of moments (GMM), which imposes an unconditional moment restriction on a fixed set of instruments. The unconditional moment using optimal-instrument constructions or feasible approximations to them is commonly applied in practice (see, e.g., \cite{berry2021foundations}). By contrast, identification is assumed on the conditional moments. The conditional moments provide basis for the optimal instrument and imply the unconditional moments, but estimation based on the unconditional moments only exploits a necessary condition. It has been pointed out by the seminal work of \citet{dominguez2004consistent} that this gap can be consequential, leading to potential failure of identification by unconditional moments even when the instrument is optimal. Whether it is consequential in the BLP model, however, has not been examined, and the tension is not widely recognized in the applied literature.

In this paper, we directly show that this identification gap can arise in the BLP model. Specifically, we construct an example in which the conditional moment restriction point-identifies the structural parameters, whereas the population GMM unconditional moments with the optimal instrument admit an additional solution. As such, it raises caution for the common BLP practice of using GMM with the unconditional moment restriction, even with the optimal instrument. Motivated by this finding, we propose a semiparametric estimator that directly implements the conditional moment restriction.

The answer is not a generic feature of all conditional-moment models. It is well known that the conditional restriction implies the unconditional one, so GMM imposes a necessary condition for identification; what is not immediate is whether that necessary condition is also sufficient in the BLP model. In a linear model, a full-rank collection of linearly independent unconditional moments with dimension equal to the parameters delivers identification \citep{dominguez2004consistent}. This paper's contribution is twofold. First, it provides an explicit example showing that the BLP model can fail to be identified by the unconditional moments. The example demonstrates how the nonlinearity of the BLP inverse can generate additional solutions to the population GMM moments. Second, the paper establishes properties of estimation and inference by directly building on the conditional moments under BLP-specific conditions.

To make the point concrete, we consider a simplified BLP model with $J=2$ products, where the market share of product $j$ in market $t$ is given by
\begin{align*}
    s_{jt} = \int_{-\infty}^{+\infty} \frac{\exp\left(p_{jt}\alpha_0 + p_{jt} \sigma_0 v  \right)}{1 + \sum_{i=1}^2 \exp\left( p_{it}\alpha_0 + p_{it} \sigma_0 v \right) } \phi (v) dv.
\end{align*}
Here, $p_{jt}$ is the price of product $j$ in market $t$, $\sigma_0$ is the standard deviation of the random coefficient, $\alpha_0$ is the price coefficient, $v$ captures consumer heterogeneity, and $\phi(\cdot)$ is the probability density function of the standard normal distribution. The parameters of interest are $(\sigma_0,\alpha_0)$, and an instrument $z_{jt}$ is also observed alongside $s_{jt}$ and $p_{jt}$.

For any candidate $\sigma$, not necessarily equal to $\sigma_0$, the share system can be inverted for the mean utilities, which delivers the conditional moment restriction $\mathbb{E}\left[y_j \left(\sigma, p_{1t}, p_{2t}\right) - \alpha p_{jt} \mid z_{jt} \right] =0$. Two features of the notation are worth flagging. The inverse depends on $\sigma$ but not on $\alpha$. And because there is no unobserved characteristic here, the shares are a deterministic function of the two prices at the true parameter, $s_t=s(p_{1t},p_{2t})$; substituting them out gives the inverse as $y_j(\sigma,p_{1t},p_{2t})$, suppressing the shares that it in fact depends on. The following example shows that the unconditional moment built on the optimal instrument can fail to identify $(\sigma_0,\alpha_0)$ even though the conditional restriction identifies them.
\begin{example}\label{eg:counter}
Let $J=2$. Set $\sigma_0 =1$, $\alpha_0 =-2$. Assume $p_{1t}$ and $p_{2t}$ follow the same discrete distribution: $\Pr\left(p_{jt} = -4\right) = 0.6$, $\Pr\left(p_{jt}=2\right) = 0.3$, $\Pr\left(p_{jt}=3\right) = 0.1$. Further assume $z_{jt} = p_{jt}$, i.e., the price is exogenous. The optimal instrument based on $z_{jt}$ is \(z_{jt}^\ast = \left(p_{jt}, \mathbb{E}\left[\frac{\partial y_j}{\partial \sigma }\left(\sigma_0, p_{1t}, p_{2t}\right) \mid p_{jt} \right] \right)'\).
The unconditional moment $\mathbb{E}\left[\left(y_j \left(\sigma, p_{1t}, p_{2t}\right) - \alpha p_{jt}\right) z_{jt}^\ast\right]=0$ obtains identification of $(\sigma_0, \alpha_0)$ iff $g_2\left(\sigma\right) = 0$ implies $\sigma = \sigma_0$, where
\begin{align*}
    g_2 \left(\sigma \right)&= \mathbb{E}\left[y_j\left(\sigma, p_{1t}, p_{2t}\right)  \mathbb{E}\left[\frac{\partial y_j}{\partial \sigma }\left(\sigma_0, p_{1t}, p_{2t}\right) \mid p_{jt} \right] \right] \\
    &\quad - \frac{\mathbb{E}\left[ y_j\left(\sigma, p_{1t}, p_{2t}\right) p_{jt}  \right]}{ \mathbb{E}\left[p_{jt}^2\right]} \mathbb{E}\left[ p_{jt} \mathbb{E}\left[\frac{\partial y_j}{\partial \sigma }\left(\sigma_0, p_{1t}, p_{2t}\right) \mid p_{jt} \right]\right].
\end{align*}
Although the close-form expression of $g_2\left(\sigma\right)$ is not available, Figure \ref{fig:example:optimal_IV} plots the numerical evaluation.
\begin{figure}[t]
    \begin{subfigure}{0.5\textwidth}
        \centering
        \includegraphics[width=0.7\linewidth]{Optimal_IV_Moment_7.pdf}
        \caption{$g_2\left(\sigma\right)$ with Optimal IV}
        \label{fig:example:optimal_IV}
    \end{subfigure}
    \begin{subfigure}{0.5\textwidth}
        \centering
        \includegraphics[width=0.7\linewidth]{Conditional_Moments_7.pdf}
        \caption{Two Functions of $\sigma$ from the Conditional Moments}
        \label{fig:example:conditional}
    \end{subfigure}
\end{figure}
The function has a second zero for $g_2\left(\sigma\right)$ in addition to $\sigma_0 =1$, so idenfication based on the unconditional moment fails.

By contrast, the conditional moment $\mathbb{E}\left[y_j \left(\sigma, p_{1t}, p_{2t}\right) - \alpha p_{jt} \mid z_{jt} \right] =0$ identifies $(\sigma_0, \alpha_0)$. Indeed, $(\sigma_0, \alpha_0)$ is identified if an only if $\sigma_0$ is the unique zero for the following two functions jointly:
\begin{align*}
    \mathbb{E}\left[ y_1 \left(\sigma, p_{1t}, p_{2t} \right) \mid p_{1t} = -4 \right] + 2 \mathbb{E}\left[ y_1 \left(\sigma, p_{1t}, p_{2t} \right) \mid p_{1t} = 2 \right] &=0, \; \text{and}\\
    3\mathbb{E}\left[ y_1 \left(\sigma, p_{1t}, p_{2t} \right) \mid p_{1t} = 2 \right] - 2 \mathbb{E}\left[ y_1 \left(\sigma, p_{1t}, p_{2t} \right) \mid p_{1t} = 3 \right] &=0.
\end{align*}
Figure \ref{fig:example:conditional} plots their numerical evaluations. Their only common zero is $\sigma_0 =1$, so the conditional moment restriction identifies the true parameters.
\end{example}

This paper proposes a semiparametric approach, built directly on the identification assumption, for estimation and inference in the BLP model. The approach estimates the conditional expectations and conditional variance-covariance matrix nonparametrically, and then chooses the parameters from a weighted minimum distance method. It has two advantages. First, estimation and identification rest on the same condition: it imposes no auxiliary moment restriction distinct from the identifying one, thereby avoiding the failure exposed by Example \ref{eg:counter}.
Second, it is flexible in the first stage. Any nonparametric estimator satisfying suitable regularity conditions can be used to estimate the conditional expectation function, including machine-learning methods.
The standard GMM approach is nested as a special case of this nonparametric procedure.
When the conditional expectation is estimated by a linear projection onto a fixed, finite set of instruments, the method reduces to standard GMM, as noted by \cite{ai2003efficient}, while sieve estimation of the conditional expectation corresponds to GMM with a growing set of instruments.

We establish consistency and asymptotic normality for the proposed estimator, which delivers valid inference.
We consider two nonparametric implementations, sieve estimation and kernel estimation, and present simulation evidence that the estimator outperforms the standard GMM approach in finite samples when the identification of the unconditional moments is unclear. As a practical takeaway, when the conditional moment restriction is assumed, we recommend estimation directly from the that restriction for the BLP model, rather than using it to construct optimal instruments and then estimating from the unconditional moments those instruments generate.
}



\paragraph{Related literature.}
This paper connects to several strands of the literature on the identification and estimation of differentiated-products demand.

{Our starting point is the literature on the nonparametric identification of demand from aggregate data. \citet{berry2014identification} use a completeness-type condition on the instruments---of the form $\mathbb{E}[f(s_t,p_t)\mid z_t,x_t]=0$ if and only if $f=0$---to identify demand from a conditional moment restriction. We retain a parametric random-coefficients specification and assume that the conditional restriction identifies the structural parameters. Our concern is whether the fixed, finite collection of unconditional moments used in conventional GMM preserves that identification. We show through a counterexample that it need not, even when the moments are formed with the optimal instrument, and develop an estimator that directly implements the identifying conditional restriction.}

A large literature studies the estimation of BLP-type models by the generalized method of moments and, in particular, the choice of instruments used to form the unconditional moments. \cite{chamberlain1987asymptotic} characterizes the efficient instruments for a conditional moment model, and a subsequent literature develops feasible approximations to them: \cite{berry1999voluntary} and \cite{reynaert2014improving} document the gains from approximating the optimal instruments. \cite{gandhi2019measuring} propose differentiation instruments that can improve the approximating instruments. \cite{conlon2020best,conlon2023incorporating} compare and implement optimal instrument sets in practice, and \cite{donald2009choosing} study how many approximating instruments to include. Our contribution is complementary but distinct: this literature seeks better \emph{finite} instrument sets within the unconditional-moment GMM framework, whereas we show that the selected unconditional moments need not preserve identification by the conditional moment restriction, even when the selected instruments are optimal in the sense of efficiency.

Methodologically, we build on the literature on sieve estimation of conditional moment restrictions, in particular the sieve minimum-distance framework of \cite{ai2003efficient}. Although that framework is not developed for the BLP model, it applies directly once demand is written as a conditional moment restriction in $\theta$, and it nests standard GMM as the special case in which the conditioning information is projected onto a fixed, finite set of instruments. Our estimator can be viewed as an instance of this framework tailored to the BLP inverse, with the number of approximating terms allowed to grow with the sample.

{This paper's approach is related to but differs from the literature using the integrated conditional moment method; see, e.g., \citet{bierens1982consistent,lavergne2013smooth,antoine2014conditional,escanciano2018simple}. These papers replace the conditional moment restriction with an equivalent continuum of unconditional moment restrictions, and under suitable conditions the same construction could conceptually be applied to the BLP model. Yet to our knowledge it has not been discussed in the previous literature, and is beyond the scope of this paper. Establishing the corresponding theory for integrated conditional-moment estimators in the BLP setting and comparing their statistical and computational properties with our approach are directions for future research.}


Our approach is distinct from the recent literature that estimates BLP-type demand allowing a nonparametric distribution of consumer heterogeneity. \cite{wang2023sieve} proposes a sieve BLP estimator that nonparametrically recovers the distribution of random coefficients, and \cite{fox2012random} and \cite{fox2016nonparametric} establish nonparametric identification of that distribution under support and regularity conditions that differ from the standard BLP environment. \cite{lu2023semi} develop a semiparametric estimator and recover the unknown distribution of random coefficients but are under the large-$J$ asymptotic framework. We instead retain the parametric random-coefficient specification and target $(\sigma,\alpha)$ directly with $J$ fixed; our object of interest is the estimation--identification gap rather than the flexibility of the heterogeneity distribution. Finally, our contribution differs from work that accelerates the computation of the standard BLP estimator, such as the MPEC reformulation of \cite{dube2012improving} and the accelerated iteration of \cite{lee2015computationally}: those papers compute a given estimator more efficiently, whereas we study a different population criterion and estimation strategy. \cite{salanie2022fast} construct tractable approximations based on low-order Taylor expansions of the demand inverse. Our approach instead retains the original demand inverse and estimates the relevant conditional expectations nonparametrically.

\paragraph{Organization.}
The remainder of this paper is structured as follows.
Section 2 discusses the framework and estimation.
Section 3 presents the general theory for estimation and inference.
Section 4 summarizes the Monte Carlo simulation results, and Section 5 concludes.
All proofs and lemmas are collected in the Appendix.

\paragraph{Notations.}
$\lVertA\rVert:=\sqrt{tr(A'A)}$.
$\mathbb{E}_T:= \frac{1}{T}\sum_{t=1}^T$ and $\mathbb{E}_J:= \frac{1}{J}\sum_{j=1}^J$.
$a_n\asymp b_n$ means $a_n=O(b_n)$ and $b_n=O(a_n)$.
$a_n\lesssim b_n$ means $a_n=O(b_n)$.
$\lVerta\rVert = (a'a)^{1/2}$ as the Euclidean norm of a vector $a$.
$\lVerta\rVert_\infty=\max_i |a_i|$ as the sup-norm of a vector $a$.
$\lVertA\rVert_{op}=\sup_{\lVertb\rVert=1}\lVertAb\rVert=\max\{|\lambda_{\max}(A)|,|\lambda_{\max}(-A)|\}$ as the operator norm of a matrix $A$.
$\lVertf(z)\rVert_{L^2}=\mathbb{E}[f(z)^2]^{1/2}$ as the $L^2$ norm of a function $f(z)$.

\section{Model and Estimation}
\subsection{Model}\label{sec:model}
We adopt the random-coefficients discrete-choice model of \cite{berry1995automobile}.
Consider a sequence of markets indexed by $t=1,\dots,T$.\footnote{The definition of a market is application-specific: with cross-sectional data markets are typically geographic areas, with panel data they are time periods, and in many cases a combination of the two.}
In each market a unit mass of consumers chooses among $J$ inside products, $j=1,\dots,J$, or an outside option $j=0$.
Product $j$ in market $t$ is summarized by the triple $(x_{jt}',p_{jt},\xi_{jt})$: a vector of $d_x$ exogenous characteristics $x_{jt}\in\mathbb{R}^{d_x}$, a scalar price $p_{jt}\in\mathbb{R}$, and a scalar characteristic $\xi_{jt}\in\mathbb{R}$ that is observed by consumers and firms but not by the econometrician.
Because prices are determined in equilibrium, they are correlated with $\xi_{jt}$ and are treated as endogenous.
Identification therefore relies on a vector of instruments $z_{jt}\in\mathbb{R}^{d_z}$ satisfying the conditional moment restriction
\begin{equation}\label{eq:exog}
    \mathbb{E}[\xi_{jt}\mid z_{jt}]=0,
    \qquad
    \sup_{1\le j\le J}\mathbb{E}[\xi_{jt}^{2}\mid z_{jt}]<C<\infty,
\end{equation}
where the exogenous characteristics $x_{jt}$ are themselves included among the instruments $z_{jt}$.

Given this market structure, we turn to the specification of preferences.
The indirect utility that consumer $i$ derives from product $j$ in market $t$ is
\begin{equation}\label{eq:utility}
    u_{ijt}=\underbrace{x_{jt}'\beta_0+p_{jt}\alpha_0+\xi_{jt}}_{\textstyle y_{jt}}+x_{jt}^{(1)\prime}\beta_i+\epsilon_{ijt},
\end{equation}
and the utility of the outside option is normalized to $u_{i0t}=\epsilon_{i0t}$.
The idiosyncratic terms $\epsilon_{ijt}$ are independent and identically distributed across $i$, $j$, and $t$, following the Type-I extreme-value distribution.
The taste for the characteristic $x_{jt}^{(1)}$ is consumer-specific and $\beta_i$ is a random variable drawn from a distribution $F(\,\cdot\,;\sigma_0)$ known up to a finite-dimensional parameter $\sigma_0\in\mathbb{R}^{d_\sigma}$.
The bracketed term $y_{jt}=x_{jt}'\beta_0+p_{jt}\alpha_0+\xi_{jt}$ is the \emph{mean utility}: the linear term of product $j$ in market $t$.

We collect the parameters in $\theta_0=(\sigma_0,\alpha_0,\beta_0')'\in\Theta$, where $\Theta$ is compact and $\theta_0$ lies in its interior.
Throughout, $\beta_i$, $\epsilon_{ijt}$, and $(\xi_{jt},x_{jt},p_{jt})$ are mutually independent, and $(\beta_i,\epsilon_{ijt})$ are i.i.d.\ across consumers.

Aggregating these preferences over consumers delivers the market shares predicted by the model.
Consumer $i$ purchases the product that maximizes her utility, so her choice of product $j$ is described by
\begin{equation*}
    d_{ijt}=
    \begin{cases}
        1 & \text{if } u_{ijt}>u_{ikt}\text{ for all }k\neq j,\\
        0 & \text{otherwise.}
    \end{cases}
\end{equation*}
Integrating the individual choices over the extreme-value shocks and the distribution of $\beta_i$ yields the market share of product $j$,
\begin{align}
    s_{jt}
    &=f_{js}(x^{(1)}_t,y_t,\sigma_0)
    =\int d_{ijt}\,\mathrm{d}F(\epsilon_{i0t},\dots,\epsilon_{iJt})\,\mathrm{d}F(\beta_i;\sigma_0)\notag\\
    &=\int\frac{\exp(y_{jt}+x_{jt}^{(1)\prime}\beta_i)}{1+\sum_{k=1}^{J}\exp(y_{kt}+x_{kt}^{(1)\prime}\beta_i)}\,\mathrm{d}F(\beta_i;\sigma_0),
    \label{eq:share}
\end{align}
where the second line follows from integrating the Type-I extreme-value shocks analytically to obtain the familiar logit kernel and then averaging over the random coefficient $\beta_i$.
Stacking the products, the vector of shares in market $t$ is
\begin{equation}\label{eq:share-vec}
    s_t=f_s(x^{(1)}_t,y_t,\sigma_0)
    :=\bigl(f_{1s}(x^{(1)}_t,y_t,\sigma_0),\dots,f_{Js}(x^{(1)}_t,y_t,\sigma_0)\bigr)',
\end{equation}
with market-level vectors $y_t=(y_{1t},\dots,y_{Jt})'$, $x^{(1)}_t=(x_{1t}^{(1)},\dots,x_{Jt}^{(1)})'$, and $\xi_t=(\xi_{1t},\dots,\xi_{Jt})'$ defined analogously.
Equation~\eqref{eq:share} is the random-coefficients logit share, and \eqref{eq:share-vec} defines the map from parameters and characteristics to shares.

The share system just derived can be inverted to recover the mean utilities, which is the property underlying our estimator.
\cite{berry1994estimating} and \cite{berry1995automobile} establish that this map can be inverted in the mean utilities.
For any fixed $\sigma$, any characteristics vector $x^{(1)}_t$, and any share vector $s_t$ in the interior of the unit simplex, there exists a unique $y\in\mathbb{R}^{J}$ solving $s_t=f_s(x^{(1)}_t,y,\sigma)$; that is, $f_s(x^{(1)}_t,\cdot\,,\sigma)$ is invertible.\footnote{Equivalently, \cite{berry1994estimating} shows that for given $x^{(1)}_t$, $s_t$, $x_t=(x_{1t}',\dots,x_{Jt}')'$, and $(\beta,\alpha,\sigma)$ there is a unique $\xi\in\mathbb{R}^{J}$ solving $s_t=f_s(x^{(1)}_t,x_t\beta+p_t\alpha+\xi,\sigma)$. Uniqueness of $y$ is the special case obtained by setting $\beta=0$ and $\alpha=0$.}
Denote the inverse by $f_s^{-1}(s_t,x^{(1)}_t,\sigma)\in\mathbb{R}^{J}$, and define the imputed mean utilities and demand shocks
\begin{align}
    y(\sigma,w_t)&:=f_s^{-1}(s_t,x^{(1)}_t,\sigma),\label{eq:inv}\\
    \xi(\theta,w_t)&:=y(\sigma,w_t)-x_t'\beta-p_t\alpha,\label{eq:xi}
\end{align}
where $w_t$ denotes all the data in market $t$, $\theta=(\sigma,\alpha,\beta')'$ and $x_t'\beta$ denotes the vector with $j$-th element $x_{jt}'\beta$.
Let $y_{j}(\sigma,w_t)$ and $\xi_{j}(\theta,w_t)$ be the $j$-th elements of $y(\sigma,w_t)$ and $\xi(\theta,w_t)$.
By construction these imputations recover the truth at $\theta_0$: $y_{j}(\sigma_0,w_t)=y_{j t}$ and $\xi_{j}(\theta_0,w_t)=\xi_{j t}$.

The identification assumption (the conditional moment restriction) is
\begin{align}
    \mu_{y,j}(z_{jt};\sigma)-\alpha\,\mu_{p,j}(z_{jt})-x_{jt}'\beta=0 \quad &\text{for all } j\; a.s.\text{ iff}\quad \theta=\theta_0 \label{eq:id1}\\
        \qquad\Longleftrightarrow\qquad& \nonumber\\
    \mathbb{E}_J\mathbb{E}\Bigl[\mu_{y,j}(z_{jt};\sigma)-\alpha\,\mu_{p,j}(z_{jt})-x_{jt}'\beta\Bigr]^2=0 \quad &\text{iff}\quad \theta=\theta_0.\label{eq:id2}
\end{align}
Restriction~\eqref{eq:id2} is the population relationship on which our estimator is built.


\subsection{Estimation}

The estimation proceeds in two steps.
\begin{enumerate}
    \item[Step 1:] For each candidate $\sigma$, estimate the conditional expectations $\mu_{y,j}(z_{jt};\sigma)=\mathbb{E}[y_{j}(\sigma,w_t)\mid z_{jt}]$ and $\mu_{p,j}(z_{jt})=\mathbb{E}[p_{jt}\mid z_{jt}]$ nonparametrically, yielding $\hat\mu_{y,j}(z_{jt};\sigma)$ and $\hat\mu_{p,j}(z_{jt})$.
    \item[Step 2:] Choose $\theta=(\sigma,\alpha,\beta)$ to minimize the sample criterion
    \begin{equation}\label{eq:est}
        \hat{\theta}=\arg\min_{\theta\in\Theta}\sum_{j,t}\big(\hat\mu_{y,j}(z_{jt};\sigma)-\alpha\,\hat\mu_{p,j}(z_{jt})-x_{jt}'\beta\big)^2\,\hat\Lambda_j(z_{jt}),
    \end{equation}
    where $\hat\Lambda_j$ is an estimated weight, constructed below from a nonparametric estimate of the conditional variance $\Sigma_{j}(z_{jt})=\mathbb{E}[\xi_{j}^2(\theta_0,w_t)\mid z_{jt}]$ and a trimming factor that discards instruments lying close to the boundary of their support.
\end{enumerate}
The two steps mirror the population identifying restriction: Step~1 replaces the unknown conditional expectations $\mu_{y,j}$ and $\mu_{p,j}$ with nonparametric estimates, and Step~2 chooses $\theta$ so that the resulting residual $\hat\mu_{y,j}(z_{jt};\sigma)-\alpha\,\hat\mu_{p,j}(z_{jt})-x_{jt}'\beta$---the sample analogue of $\mathbb{E}[\xi_{j}(\theta,w_t)\mid z_{jt}]$---is as close to zero as possible.
The population weight targeted by $\hat\Lambda_j$ is $\Lambda_j(z_{jt})=\Sigma_j(z_{jt})^{-1}$, the conditional-variance weighting familiar from generalized least squares: it downweights observations with noisier unobserved product characteristics and, as we show in Section~\ref{sec:theory}, delivers the efficient estimator within this class. The theory there in fact accommodates any weight bounded away from zero and infinity; the unweighted criterion used in the preliminary step below is the case $\Lambda_j\equiv1$.
We now describe each step in detail.

\paragraph*{Step 1.}
The conditional expectations can be estimated by any nonparametric smoother; we consider two standard choices, kernel and series estimation.
A useful feature common to both is that the smoothing weights depend only on the instruments $z_{jt}$ and not on $\sigma$, so that only the regressand $y_{j}(\sigma,w_t)$ changes as $\sigma$ varies.
Re-estimating $\hat\mu_{y,j}(z_{jt};\sigma)$ across values of $\sigma$ therefore requires no re-computation of the weights, which makes the outer optimization in \eqref{eq:est} computationally light.

\emph{Kernel estimation.}
Using a kernel $K_h(\cdot)$ with bandwidth $h$, the local-constant (Nadaraya--Watson) estimators are
\begin{subequations}\label{est:kernel}
\begin{align}
    \hat\mu_{y,j}(z_{jt};\sigma)&=\frac{\sum_{l=1}^{T}K_h(z_{jl}-z_{jt})\,y_{j}(\sigma,w_{l})}{\sum_{l=1}^{T}K_h(z_{jl}-z_{jt})},\\[4pt]
    \hat\mu_{p,j}(z_{jt})&=\frac{\sum_{l=1}^{T}K_h(z_{jl}-z_{jt})\,p_{jl}}{\sum_{l=1}^{T}K_h(z_{jl}-z_{jt})},
\end{align}
\end{subequations}

where the sums run over markets $l=1,\dots,T$ and the estimators are computed product by product.
If the conditional distributions of $y_{j}(\sigma,w_t)$ and $p_{jt}$ given $z_{jt}$ do not vary across products---for instance when products are exchangeable---the regression functions $\mu_{y,j}(\cdot\,;\sigma)$ and $\mu_{p,j}(\cdot)$ are common across $j$, and pooling over products yields a more efficient estimator by enlarging the effective sample:
\begin{align*}
    \hat\mu_{y,j}(z_{jt};\sigma)&=\frac{\sum_{l=1}^{T}\sum_{j'=1}^{J}K_h(z_{j'l}-z_{jt})\,y_{j'}(\sigma,w_l)}{\sum_{l=1}^{T}\sum_{j'=1}^{J}K_h(z_{j'l}-z_{jt})},\\[4pt]
    \hat\mu_{p,j}(z_{jt})&=\frac{\sum_{l=1}^{T}\sum_{j'=1}^{J}K_h(z_{j'l}-z_{jt})\,p_{j'l}}{\sum_{l=1}^{T}\sum_{j'=1}^{J}K_h(z_{j'l}-z_{jt})}.
\end{align*}

\emph{Series estimation.}
Let $\{\phi_k(z)\}_{k=1}^{K}$ be a set of basis functions and collect them in $\phi^{K}(z_{jt})=(\phi_1(z_{jt}),\dots,\phi_{K}(z_{jt}))'$, with the number of terms $K=K_n$ growing with the sample size.
The series estimators are the fitted values of the least-squares projections onto these bases,
\begin{subequations}\label{est:series}
\begin{align}
    \hat\mu_{y,j}(z_{jt};\sigma)&=\phi^{K}(z_{jt})'\hat{\Pi}_{y,j}(\sigma),\\
    \hat\mu_{p,j}(z_{jt})&=\phi^{K}(z_{jt})'\hat{\Pi}_{p,j},
\end{align}
\end{subequations}

where $\hat{\Pi}_{y,j}(\sigma)$ and $\hat{\Pi}_{p,j}$ are the coefficient vectors from regressing $y_{j}(\sigma,w_t)$ and $p_{jt}$, respectively, on $\phi^{K}(z_{jt})$.
Because the projection is linear in the regressand, the estimated coefficients are
\begin{align*}
    \hat{\Pi}_{y,j}(\sigma)&=\big(\mathbb{E}_T\phi^{K}(z_{jt})\phi^{K}(z_{jt})'\big)^{-1}\mathbb{E}_T\phi^{K}(z_{jt})\,y_{j}(\sigma,w_t),\\
    \hat{\Pi}_{p,j}&=\big(\mathbb{E}_T\phi^{K}(z_{jt})\phi^{K}(z_{jt})'\big)^{-1}\mathbb{E}_T\phi^{K}(z_{jt})\,p_{jt}.
\end{align*}
So the design matrix is inverted once and only the second factor is updated as $\sigma$ changes---the same computational saving noted for the kernel estimator.
As with the kernel case, if the regression functions are common across products---so that $\mu_y(\cdot\,;\sigma)$ and $\mu_p(\cdot)$ do not depend on $j$---the coefficients may be estimated by stacking all products in a single least-squares regression:
\begin{align*}
    \hat{\Pi}_{y,j}(\sigma)&=\Big(\mathbb{E}_J\mathbb{E}_T\phi^{K}(z_{jt})\phi^{K}(z_{jt})'\Big)^{-1}\mathbb{E}_J\mathbb{E}_T\phi^{K}(z_{jt})\,y_{j}(\sigma,w_t),\\[4pt]
    \hat{\Pi}_{p,j}&=\Big(\mathbb{E}_J\mathbb{E}_T\phi^{K}(z_{jt})\phi^{K}(z_{jt})'\Big)^{-1}\mathbb{E}_J\mathbb{E}_T\phi^{K}(z_{jt})\,p_{jt},
\end{align*}
and the fitted values $\hat\mu_{y,j}(z_{jt};\sigma)=\phi^{K}(z_{jt})'\hat{\Pi}_{y,j}(\sigma)$ and $\hat\mu_{p,j}(z_{jt})=\phi^{K}(z_{jt})'\hat{\Pi}_{p,j}$ are formed as before.
By contrast, the product-specific estimator restricts each sum to a single $j$, replacing $\sum_{j=1}^{J}\sum_{t=1}^{T}$ with $\sum_{t=1}^{T}$ and yielding a separate coefficient vector $\hat{\Pi}_{y,j}(\sigma)$ for each product.

\paragraph*{Step 2.}
Given the first-step estimates, the criterion in \eqref{eq:est} is minimized over $\theta=(\sigma,\alpha,\beta)$.
The residual is linear in the finite-dimensional parameters $(\alpha,\beta)$ and enters $\sigma$ only through $\hat\mu_{y,j}(\cdot\,;\sigma)$.
Hence, for any fixed $\sigma$, the inner minimization over $(\alpha,\beta)$ is a weighted least-squares problem with the closed-form solution
\begin{equation*}
    \big(\hat\alpha(\sigma),\hat\beta(\sigma)\big)
    =\Big(\textstyle\sum_{j,t}\tilde{x}_{jt}\tilde{x}_{jt}'\,\hat\Lambda_j(z_{jt})\Big)^{-1}
    \sum_{j,t}\tilde{x}_{jt}\,\hat\mu_{y,j}(z_{jt};\sigma)\,\hat\Lambda_j(z_{jt}),
    \qquad
    \tilde{x}_{jt}=\big(\hat\mu_{p,j}(z_{jt}),\,x_{jt}'\big)',
\end{equation*}
so the optimization reduces to a one-dimensional search over $\sigma$ (or $d_\sigma$-dimensional when $\sigma$ is a vector), with $(\alpha,\beta)$ concentrated out.
This profiling substantially lightens the numerical burden relative to a joint search over all of $\theta$.


\paragraph*{The estimated weight.}
The weight factors into a variance estimate and a trimming indicator,
\begin{equation}\label{eq:weight-factor}
    \hat\Lambda_j(z)=\hat\Sigma_j(z)^{-1}\,\tau_T(z),
\end{equation}
and the two factors do different jobs. The first is obtained in a preliminary estimation: one solves \eqref{eq:est} with $\hat\Lambda_j(z_{jt})= \tau_T(z_{jt})$ (unweighted), forms the residuals $\hat\xi_{jt}=y_j(\tilde\sigma,w_t)-\tilde\alpha \ p_{jt}-x_{jt}'\tilde\beta$, where $\tilde\sigma,\ \tilde\alpha,\ \tilde\beta$ are the preliminary estimates, and estimates $\Sigma_j$ by smoothing $\hat\xi_{jt}^2$ on $z_{jt}$ with the same nonparametric method used in Step~1, denoted by $\hat\Sigma_j(z)$.
The criterion \eqref{eq:est} is then re-minimized with the estimated weights. The second factor, $\tau_T$, takes the value zero for instruments close to the boundary of their support and one elsewhere for kernel estimation, while for series estimation it takes the value one for all $z_{jt}$. It is there because $\hat\Sigma_j$ and the first-step estimates are nonparametric fits: near the boundary they are computed from few effective observations, and dropping those observations is what allows the uniform control required for inference to be imposed on the interior region alone, which is crucial in the kernel estimation. The discarded region shrinks with the sample, so the trimming does not affect the limiting distribution. Both factors depend on the smoother, and we take the two in turn. Throughout, $\mathcal Z$ denotes the support of $z_{jt}$ and
\[
    d(z,\partial\mathcal Z)=\inf_{z'\in\partial\mathcal Z}\|z-z'\|_\infty
\]
is the sup-norm distance from $z$ to the boundary of $\mathcal Z$, which for $z$ in the interior of $\mathcal Z$ is the distance to the boundary $\partial\mathcal Z$.

\emph{Kernel.} The conditional variance is the Nadaraya--Watson fit of the squared preliminary residuals,
\[
    \hat\Sigma_j(z)=\frac{\sum_{l=1}^{T}K_h(z_{jl}-z)\,\hat\xi_{jl}^{\,2}}{\sum_{l=1}^{T}K_h(z_{jl}-z)} .
\]
Let $R<\infty$ be the radius of the support of the kernel, so that $K(u)=0$ whenever $\|u\|_\infty>R$ and $K_h(\cdot-z)$ is supported on the window $\{z':\|z'-z\|_\infty\le Rh\}$. The trimming is then
\[
    \tau_T(z)=\mathbf{1}\{d(z,\partial\mathcal Z)\ge Rh\},
\]
which retains exactly those points whose smoothing window lies inside $\mathcal Z$. Since the discarded collar has width $Rh\to0$, the fraction of observations it removes vanishes.

\emph{Series.} The conditional variance is the least-squares projection of the squared preliminary residuals on the same dictionary,
\[
    \hat\Sigma_j(z)=\phi^{K}(z)'\hat\Pi_{\xi^2,j},
    \qquad
    \hat\Pi_{\xi^2,j}=\big(\mathbb{E}_T\phi^{K}(z_{jt})\phi^{K}(z_{jt})'\big)^{-1}\mathbb{E}_T\phi^{K}(z_{jt})\,\hat\xi_{jt}^{\,2},
\]
formed with the same design matrix as in Step~1, so no additional inversion is required. Here no trimming is needed and one sets $\tau_T\equiv1$, so that $\hat\Lambda_j=\hat\Sigma_j^{-1}$: a series fit is a global projection rather than a local average, so it has no boundary window to truncate.

Finally, the smoothing parameters---the bandwidth $h$ for the kernel estimator and the number of basis terms $K$ for the series estimator---govern the bias--variance trade-off of the first step; the rate conditions they must satisfy for valid inference on $\theta$ are given in Section~\ref{sec:theory}.


\section{Asymptotic Theory}\label{sec:theory}

















\subsection{General Theory}\label{sec:theory-general}

We first fix notation common to both results.  Consider a compact parameter
space $\Theta$ and support set of $z_{jt}$ as $\mathcal Z$.  We maintain the condition that the linear characteristics
are their own projection onto the instruments,
$\mathbb E[x_{jt}\mid z_{jt}]=x_{jt}$.  Under this condition the identifying
restriction takes the form
\begin{equation}\label{eq:theory-id}
  \mu_{y,j}(z_{jt};\sigma_0)=\alpha_0\,\mu_{p,j}(z_{jt})+x_{jt}'\beta_0.
\end{equation}

The asymptotic analysis requires the derivatives of $\mu_{y,j}$ in the
nonlinear parameter.  We denote the first and second derivatives by
\[
  \dot\mu_{y,j}(z;\sigma)=\tfrac{\partial}{\partial\sigma}\mu_{y,j}(z;\sigma)
  =\mathbb E\Bigl[\tfrac{\partial}{\partial\sigma}y_j(\sigma,w)\Bigm| z\Bigr],
  \qquad
  \ddot\mu_{y,j}(z;\sigma)=\tfrac{\partial^2}{\partial\sigma^2}\mu_{y,j}(z;\sigma)
  =\mathbb E\Bigl[\tfrac{\partial^2}{\partial\sigma^2}y_j(\sigma,w)\Bigm| z\Bigr].
\]
The corresponding nonparametric estimators of $\mu_{y,j}(z;\sigma)$,
$\dot\mu_{y,j}(z;\sigma)$ and $\ddot\mu_{y,j}(z;\sigma)$ are written
$\hat\mu_{y,j}(z;\sigma)$, $\hat{\dot\mu}_{y,j}(z;\sigma)$ and
$\hat{\ddot\mu}_{y,j}(z;\sigma)$, respectively.  It is convenient to
abbreviate the direction,
\[
  \Gamma_{jt}\;:=\;\bigl(\dot\mu_{y,j}(z_{jt};\sigma_0),\,-\mu_{p,j}(z_{jt}),\,-x_{jt}\bigr)'
  \;=\;\nabla_\theta\, r_{jt}(\theta_0),
\]
with $r_{jt}(\theta)$ the residual term, defined as
$r_{jt}(\theta)=\mu_{y,j}(z_{jt};\sigma)-\alpha\,\mu_{p,j}(z_{jt})-x_{jt}'\beta$ for the population residual; under \eqref{eq:theory-id}, $r_{jt}(\theta_0)=0$ for every $(j,t)$.

\paragraph{Weighting and trimming.} The criterion (\ref{eq:est}) is built from two
further objects.  The first is a population weight
$\omega_j:\mathcal Z\to(0,\infty)$, $j=1,\dots,J$, together with an estimator
$\hat\omega_j$ of it.  We do \emph{not} require $\omega_j$ to be the
conditional variance $\Sigma_j$: any weight bounded away from zero and
infinity is admissible, $\omega_j\equiv1$ giving the unweighted estimator and
$\omega_j=\Sigma_j$ the efficient one (under suitable conditions).  The second is a trimming indicator.
Letting $\varepsilon_T\to0$ be a nonrandom sequence, put
\[
  \mathcal Z_T\;:=\;\bigl\{z\in\mathcal Z:\ d(z,\partial\mathcal Z)\ \ge\ \varepsilon_T\bigr\},
  \qquad
  \tau_T(z)\;:=\;\mathbf 1\bigl\{d(z,\partial\mathcal Z)\ \ge\ \varepsilon_T\bigr\},
\]
so that $\tau_T(z)=0$ for $z$ within $\varepsilon_T$ of the boundary and
$\tau_T(z)=1$ for $z$ in the interior region $\mathcal Z_T$, and write
\[
  p_T\;:=\;\max_{1\le j\le J}\Pr\bigl(\tau_T(z_{jt})=0\bigr)
\]
for the probability mass discarded.  The leading case, used for the kernel
implementation in Section \ref{sec:theory-kernel}, is $\varepsilon_T=Rh$, where $h$ is the
bandwidth and $R$ is the radius of the support of the kernel, so that for
every retained $z$ the smoothing window $\{z':\|z'-z\|_{\infty}\le Rh\}$ lies entirely
inside $\mathcal Z$; the series implementation in Section \ref{sec:theory-series} needs no
trimming and takes $\varepsilon_T=0$, i.e.\ $\tau_T\equiv1$ and $p_T=0$.
Finally, collect the weights as they enter the criterion,
\[
  \Lambda_j(z)\;:=\;\omega_j(z)^{-1},
  \qquad
  \hat\Lambda_j(z)\;:=\;\hat\omega_j(z)^{-1}\,\tau_T(z),
  \qquad
  a_{jt}\;:=\;\Lambda_j(z_{jt})\,\Gamma_{jt},
\]
so that $\Lambda_j$ is the infeasible target, $\hat\Lambda_j$ the feasible
trimmed weight, and $\hat\Lambda_j(z)=0$ whenever $z\notin\mathcal Z_T$.

It is convenient to treat these conditional-mean functions as an
infinite-dimensional nuisance parameter, which we collect in
$\eta=(\eta_p,\eta_y,\dot\eta_y,\eta_\Lambda)$, with true value
$\eta_0=(\mu_p,\mu_y,\dot\mu_y,\Lambda)$, where $\mu_p=(\mu_{p,1},\dots,\mu_{p,J})$,
$\mu_y=(\mu_{y,1},\dots,\mu_{y,J})$, $\dot\mu_y=(\dot\mu_{y,1},\dots,\dot\mu_{y,J})$
and $\Lambda=(\Lambda_1,\dots,\Lambda_J)$; the last coordinate is the weight
as it enters the criterion, so the trimming is carried inside the nuisance
parameter.  Define the moment function
\begin{equation}\label{eq:moment-fn}
  g(w_t,\theta,\eta)=\mathbb E_J\left[
    \bigl(\eta_{y,j}(z_{jt};\sigma)-\alpha\,\eta_{p,j}(z_{jt})-x_{jt}'\beta\bigr)\,
    \eta_{\Lambda,j}(z_{jt})
    \begin{pmatrix}\dot\eta_{y,j}(z_{jt};\sigma)\\[2pt] -\eta_{p,j}(z_{jt})\\[2pt] -x_{jt}\end{pmatrix}
  \right],
\end{equation}
The estimator in (\ref{eq:est}) is then a solution to the sample first-order
condition
\[
  \mathbb E_T\,g(w_t,\hat\theta,\hat\eta)=0,
  \qquad
  \hat\eta=(\hat\mu_p,\hat\mu_y,\hat{\dot\mu}_y,\hat\Lambda).
\]



\begin{condition}\label{cond:regularity}
\begin{enumerate}
\item\label{cond:independence} (Independence) The markets $t=1,\dots,T$ are independent and
identically distributed draws of
$w_t=(z_{jt},x_{jt},p_{jt},s_{jt},\xi_{jt})_{j=1}^J$, with $z_{jt}\supset x_{jt}$.
\item\label{cond:cons-reg} (Regressors and parameter space) $\|\xi_{jt}\|_\infty$,
$\|x_{jt}\|_\infty$, $\|z_{jt}\|_\infty\le C$ and $|p_{jt}|\le C$ almost
surely; $\Theta$ is compact and $\theta_0$ lies in its interior.
\item\label{cond:cons-weight} (Weighting) There exist constants $0<c\le C<\infty$ such that
$c\le\Lambda_j(z)=\omega_j(z)^{-1}\le C$ for all $z$ and $j$.
\item\label{cond:cons-id} (Identification) $\theta_0$ uniquely minimizes the population criterion
$Q(\theta)=\mathbb E_J\mathbb E\bigl[r_{jt}(\theta)^2\,\Lambda_j(z_{jt})\bigr]$.
\item\label{cond:cons-bound} (Smoothness) $\mu_{y,j},\mu_{p,j}$ are bounded uniformly and continuous
in $(\sigma,z)$ and $z$ respectively for each $j$.
\end{enumerate}
\end{condition}

\begin{condition}\label{cond:consistency}
\begin{enumerate}
\item\label{cond:cons-rate} (First-stage consistency) Uniformly over $\sigma\in\Theta_\sigma$,
\[
  \sup_{\sigma\in\Theta_\sigma}\mathbb{E}_J\mathbb{E}_T\Big[\big|\hat\mu_{y,j}(z_{jt};\sigma)-\mu_{y,j}(z_{jt};\sigma)\big|^2\tau_T(z_{jt})\Big]
  +\mathbb{E}_J\mathbb{E}_T\Big[\big|\hat\mu_{p,j}(z_{jt})-\mu_{p,j}(z_{jt})\big|^2\tau_T(z_{jt})\Big]=o_p(1).
\]
\item\label{cond:cons-uniform} (Uniform convergence) Uniformly over $z\in\mathcal Z_T$ for each $j$:
\[
  \sup_{z\in\mathcal Z_T}\bigl|\hat\omega_j(z)^{-1}-\omega_j(z)^{-1}\bigr|=o_p(1),
  \qquad\text{equivalently}\qquad
  \sup_{z\in\mathcal Z}\bigl|\hat\Lambda_j(z)-\Lambda_j(z)\tau_T(z)\bigr|=o_p(1).
\]
\item\label{cond:cons-trim} (Trimming) The sequence $\varepsilon_T$ is nonrandom with
$\varepsilon_T\to0$, and $p_T=\max_j\Pr(\tau_T(z_{jt})=0)\to0$.
\end{enumerate}
\end{condition}

\begin{remark}
    Conditions~\ref{cond:regularity} and \ref{cond:consistency} are the primitives behind uniform convergence of the sample criterion to its population analogue, which is what drives consistency of $\hat\theta$.

Condition~\ref{cond:regularity} concerns the data and the population problem. Part~\eqref{cond:independence} treats markets as i.i.d.\ draws; dependence across products \emph{within} a market is unrestricted, since all population objects average over $j$.
Parts~\eqref{cond:cons-reg} and \eqref{cond:cons-weight} are mild boundedness conditions: bounded product characteristics, unobserved characteristics, and instruments on a compact parameter space, and a weight bounded away from zero and infinity so that $\Lambda_j$ neither degenerates nor explodes. Part~\eqref{cond:cons-weight} does not tie the weight to the conditional variance.
Part~\eqref{cond:cons-id} is the identification assumption \eqref{eq:id1}, restated as a unique minimizer. Since $\Lambda_j$ is bounded above and away from zero by Part~\eqref{cond:cons-weight}, $Q(\theta)=0$ if and only if $r_{jt}(\theta)=0$ almost surely for every $j$. Part~\eqref{cond:cons-bound} keeps the conditional means bounded and continuous.

Condition~\ref{cond:consistency} concerns the first stage. Part~\eqref{cond:cons-rate} asks only that the estimated conditional means converge in mean square, uniformly over the nonlinear parameter $\sigma$; this is a high-level requirement, verified for the kernel and series estimators in Sections~\ref{sec:theory-kernel} and \ref{sec:theory-series}. Part~\eqref{cond:cons-uniform} asks the same of the estimated weight, in sup norm, but only on the retained region $\mathcal Z_T$: this is what the trimming buys, since it is near $\partial\mathcal Z$ that a nonparametric weight is estimated from fewest effective observations and is least reliable.
Part~\eqref{cond:cons-trim} requires only that the discarded region be asymptotically negligible; no rate is imposed. It is implied by the geometry rather than assumed on top of it: if $\mathcal Z$ is compact with a Lipschitz boundary (``sufficiently regular''), and $z_{jt}$ has a density function $f_j$ with $\sup_z f_j(z)\le\bar f<\infty$, then $Leb\{z:d(z,\partial\mathcal Z)<\varepsilon\}\le C\varepsilon$ and hence $p_T=O(\varepsilon_T)$, so Part~\eqref{cond:cons-trim} holds for any $\varepsilon_T\to0$, and in particular for the kernel choice $\varepsilon_T=Rh$; the series implementation takes $\tau_T\equiv1$ and $p_T=0$. Together, Parts~\eqref{cond:cons-uniform} and \eqref{cond:cons-trim} and Condition~\ref{cond:regularity}\eqref{cond:cons-weight} give $0\le\hat\Lambda_j(z)\lesssim C$ uniformly on $\mathcal Z$ with probability approaching one, the boundedness used throughout the proofs.
\end{remark}

\begin{theorem}\label{thm:consistency}
Under Condition~\ref{cond:regularity} and \ref{cond:consistency}, $\hat\theta\xrightarrow{p}\theta_0$.
\end{theorem}


\begin{condition}\label{cond:y-smooth}
$\mu_{y,j}(z;\sigma)$ is bounded and three times continuously differentiable
in $\sigma$ on $\Theta_\sigma$ for each $z$, i.e., $\sup_{z\in\mathcal{Z}}|\partial^l_{\sigma} \mu_{y,j}(z;\sigma)|\le C$ for $l=0,1,2,3$ and some constant $C>0$.
\end{condition}

\begin{condition}\label{cond:normality}
\begin{enumerate}
\item\label{cond:deriv} The derivative estimators are the derivatives of the first-stage
estimator, $\hat{\dot\mu}_{y,j}(z;\sigma)=\partial_\sigma\hat\mu_{y,j}(z;\sigma)$
and $\hat{\ddot\mu}_{y,j}(z;\sigma)=\partial^2_\sigma\hat\mu_{y,j}(z;\sigma)$.
\item\label{cond:norm-deriv} (Derivative estimators consistency)
\[
  \sup_{\sigma\in\Theta_\sigma}\mathbb E_J\mathbb E_T\Big[\bigl|\hat{\dot\mu}_{y,j}(z_{jt};\sigma)-\dot\mu_{y,j}(z_{jt};\sigma)\bigr|^2\tau_T(z_{jt})\Big]
  +\sup_{\sigma\in\Theta_\sigma}\mathbb E_J\mathbb E_T\Big[\bigl|\hat{\ddot\mu}_{y,j}(z_{jt};\sigma)-\ddot\mu_{y,j}(z_{jt};\sigma)\bigr|^2\tau_T(z_{jt})\Big]=o_p(1).
\]
\item\label{cond:norm-msr} (Mean-square rate) The first-stage estimators converge in mean square
faster than $T^{-1/4}$:
\begin{align*}
    &\mathbb{E}_J\mathbb{E}_T\Big[\bigl|\hat\mu_{y,j}(z_{jt};\sigma_0)-\mu_{y,j}(z_{jt};\sigma_0)\bigr|^2\tau_T(z_{jt})\Big]
    +\mathbb{E}_J\mathbb{E}_T\Big[\bigl|\hat{\dot\mu}_{y,j}(z_{jt};\sigma_0)-\dot\mu_{y,j}(z_{jt};\sigma_0)\bigr|^2\tau_T(z_{jt})\Big]\\
    & +\;\mathbb{E}_J\mathbb{E}_T\Big[\bigl|\hat\mu_{p,j}(z_{jt})-\mu_{p,j}(z_{jt})\bigr|^2\tau_T(z_{jt})\Big]
    +\mathbb E_J\mathbb E_T\Bigl[\bigl(\hat\omega_j(z_{jt})^{-1}-\omega_j(z_{jt})^{-1}\bigr)^2\tau_T(z_{jt})\Bigr]=o_p(T^{-1/2}).
\end{align*}
The last term is equivalent to $\mathbb E_J\mathbb E_T(\hat\Lambda_j-\Lambda_j\tau_T)^2$.
\item\label{cond:norm-linear} (Asymptotic linearity) There is a market-level function $\delta(\cdot)$ with $\mathbb{E}[\delta(w_t)]=0$ and $\mathbb{E}\|\delta(w_t)\|^2<\infty$ such that
\[
  \mathbb{E}_J\mathbb{E}_T\Big[\big\{(\hat\mu_{y,j}-\mu_{y,j})-\alpha_0(\hat\mu_{p,j}-\mu_{p,j})\big\}\,
  \Lambda_j(z_{jt})\,\tau_T(z_{jt})\,\Gamma_{jt}\Big]
  =\mathbb{E}_T\delta(w_t)+o_p(T^{-1/2}),
\]
all first-stage objects being evaluated at $\sigma_0$.
\end{enumerate}
\end{condition}

\begin{remark}
Condition~\ref{cond:y-smooth} is mild: three-times differentiability of $\mu_{y,j}(z;\sigma)$ in $\sigma$ comes from smoothness of the inverse function $y_{j}(\sigma,w)$, which is inherited from smoothness of the random-coefficient density $f(b;\sigma)$ in $\sigma$.

Condition~\ref{cond:normality} strengthens Condition~\ref{cond:consistency} from consistency to $\sqrt T$-inference, and is likewise verified for the kernel and series estimators in Sections~\ref{sec:theory-kernel} and \ref{sec:theory-series}.
Part~\eqref{cond:deriv} is a definitional requirement: the derivative estimators are the analytic $\sigma$-derivatives of the first-stage fit, not separately smoothed objects; this is what enables analysis on $\partial_\sigma \hat\mu_{y,j}$.
Part~\eqref{cond:norm-deriv} then asks these derivative estimators to be mean-square consistent uniformly over $\sigma$; this is an analogy to the Condition~\ref{cond:consistency}\eqref{cond:cons-rate}, but for derivatives, which are concerned in asymptotic normality.
Parts~\eqref{cond:norm-msr} and \eqref{cond:norm-linear} are the heart of the normality argument, and both exploit the degeneracy of the moment at the truth, $r_{jt}(\theta_0)=0$, which holds for every admissible weight. Part~\eqref{cond:norm-msr} imposes the familiar $T^{-1/4}$ root-mean-square rate on every first-stage object, including the weight---on the retained region only, matching Condition~\ref{cond:consistency}\eqref{cond:cons-uniform}; this is precisely what makes the second-order (quadratic) remainder in the expansion of the moment negligible, since that remainder is a product of two first-stage errors each of order $o_p(T^{-1/4})$.
Part~\eqref{cond:norm-linear} is the high-level asymptotic-linearity requirement that isolates the surviving first-order term; the function $\delta$ is the influence function of $\hat\theta$ (up to the Jacobian factor $-M^{-1}$), whose explicit form is derived for each estimator in Sections~\ref{sec:theory-kernel} and \ref{sec:theory-series}.

\end{remark}

\begin{theorem}\label{thm:normality}
Suppose Condition~\ref{cond:regularity}, \ref{cond:consistency}, \ref{cond:y-smooth}, \ref{cond:normality} hold and that
\[
  M=\mathbb E_J\,\mathbb E\left[\Lambda_j(z_{jt})\,\Gamma_{jt}\Gamma_{jt}'\right]
   =\mathbb E_J\,\mathbb E\left[\omega_j(z_{jt})^{-1}
   \begin{pmatrix}\dot\mu_{y,j}(z_{jt};\sigma_0)\\[2pt] -\mu_{p,j}(z_{jt})\\[2pt] -x_{jt}\end{pmatrix}
   \bigl(\dot\mu_{y,j}(z_{jt};\sigma_0),\,-\mu_{p,j}(z_{jt}),\,-x_{jt}'\bigr)\right]
\]
is nonsingular. Then
\[
  \sqrt T\,(\hat\theta-\theta_0)\;\xrightarrow{\ d\ }\;N\bigl(0,\;M^{-1}V_\delta(M^{-1})'\bigr),
  \qquad V_\delta=\operatorname{Var}\bigl(\delta(w_t)\bigr),
\]
with $\delta$ as in Condition~\ref{cond:normality}\eqref{cond:norm-linear}.
\end{theorem}

\begin{remark}
The estimator is $\sqrt T$-consistent and asymptotically normal, with the
sandwich form $M^{-1}V_\delta(M^{-1})'$ standard for a two-step estimator:
$M$ is the Jacobian of the moment and $V_\delta$ the variance of its
influence function.  Nonsingularity of $M$ is the local identification
requirement that the moment be sensitive to $\theta$ in every direction at
$\theta_0$.  Two features of the statement are worth emphasizing.  First, the
weight enters only through $M$ and $V_\delta$: the argument is unchanged for
any admissible $\omega_j$. In the kernel and series implementations
$\delta(w_t)=\mathbb E_J[\xi_{jt}a_{jt}]$ with $a_{jt}=\Lambda_j(z_{jt})\Gamma_{jt}$,
giving $V_\delta=\mathbb E[(\mathbb E_J a_{jt}\xi_{jt})(\mathbb E_J a_{jt}\xi_{jt})']$.
Second, the trimming leaves no trace: both $M$ and $V_\delta$ are the
untrimmed population objects.
Efficiency is a property of the particular weight $\omega_j=\Sigma_j$: when
$\xi_{jt}$ are independent across $j$ conditional on $z_t$, a matrix
Cauchy--Schwarz inequality gives
$M^{-1}V_\delta(M^{-1})'\succeq J^{-1}\bigl(\mathbb E_J\mathbb E[\Sigma_j^{-1}\Gamma_{jt}\Gamma_{jt}']\bigr)^{-1}$
for every admissible weight, with equality at $\omega_j=\Sigma_j$, in which
case $V_\delta=M/J$ and the asymptotic variance collapses to $M^{-1}/J$, the
semiparametric efficiency bound for the conditional-moment restriction \eqref{eq:theory-id}.
Without conditional independence across $j$ the efficient weight becomes more complicated.
\end{remark}

\subsection{Theory for Kernel Estimation}\label{sec:theory-kernel}
This section provides primitive conditions for the kernel estimator in~\eqref{est:kernel} to satisfy the high-level conditions of Section~\ref{sec:theory-general}.
It is convenient to work with the kernel \emph{numerator}, the density-weighted regression function, and to keep the density estimate separate. For a generic regressand $\psi_{jt}(\sigma)$ (below, $\psi_{jt}(\sigma)\in\{y_j(\sigma,w_t),\ \partial_\sigma y_j(\sigma,w_t),\ \partial^2_\sigma y_j(\sigma,w_t),\ p_{jt}\}$, all of which are uniformly bounded by Condition~\ref{cond:regularity} \eqref{cond:cons-reg} and Condition~\ref{cond:kernel}\eqref{cond:kernel-psi}), define
\[
    m_{\psi,j}(z;\sigma)=\mu_{\psi,j}(z;\sigma)\,f_j(z),
    \qquad
    \mu_{\psi,j}(z;\sigma)=\mathbb{E}[\psi_{jt}(\sigma)\mid z_{jt}=z],
\]
with estimators
\[
    \hat m_{\psi,j}(z;\sigma)=\mathbb{E}_T\big[K_h(z_{jt}-z)\,\psi_{jt}(\sigma)\big],
    \qquad
    \hat f_j(z)=\mathbb{E}_T\big[K_h(z_{jt}-z)\big],
    \qquad
    \hat\mu_{\psi,j}(z;\sigma)=\frac{\hat m_{\psi,j}(z;\sigma)}{\hat f_j(z)}.
\]
Specializing $\psi$ recovers the objects of Section~\ref{sec:theory-general}: writing $m_{y,j}=\mu_{y,j}f_j$, $\dot m_{y,j}=\dot\mu_{y,j}f_j$, $\ddot m_{y,j}=\ddot\mu_{y,j}f_j$, and $m_{p,j}=\mu_{p,j}f_j$, the first-stage estimators are $\hat\mu_{y,j}=\hat m_{y,j}/\hat f_j$ and $\hat\mu_{p,j}=\hat m_{p,j}/\hat f_j$, and the derivative estimators are $\hat{\dot\mu}_{y,j}=\hat{\dot m}_{y,j}/\hat f_j=\partial_{\sigma}\hat\mu_{y,j}$ and $\hat{\ddot\mu}_{y,j}=\hat{\ddot m}_{y,j}/\hat f_j=\partial^2_{\sigma}\hat\mu_{y,j}$, obtained by replacing $\psi=y_j(\sigma)$ with its $\sigma$-derivatives in the numerator (the denominator $\hat f_j$ does not depend on $\sigma$).

\paragraph{Weighting.} Section~\ref{sec:theory-general} leaves the weight $\omega_j$ free. From here on we specialize to
\[
    \omega_j(z)=\Sigma_j(z)=\mathbb{E}[\xi_{jt}^2\mid z_{jt}=z],
    \qquad\text{so that}\qquad
    \Lambda_j(z)=\Sigma_j(z)^{-1},
    \qquad
    \hat\Lambda_j(z)=\hat\Sigma_j(z)^{-1}\tau_T(z),
\]
with $\hat\Sigma_j$ the kernel estimator of $\Sigma_j$ obtained by smoothing the squared preliminary residuals which is constructed using the preliminary estimator $\tilde\theta$. With an abuse of notation, the preliminary estimator $\tilde\theta$ that supplies those residuals is the case $\omega_j\equiv1$, for which $\Lambda_j\equiv1$ and $\hat\Lambda_j=\tau_T$.

\paragraph{Trimming.} The trimming introduced in Section~\ref{sec:theory-general} is instantiated here, and the kernel dictates the choice.  Let $R<\infty$ be such that $\operatorname{supp}K\subseteq[-R,R]^{d_z}$ (Condition~\ref{cond:kernel}\eqref{cond:kernel-smoothness-bw}), write $d(z,\partial\mathcal Z)=\inf_{z'\in\partial\mathcal Z}\|z-z'\|_\infty$ for the sup-norm distance from $z$ to the boundary $\partial\mathcal Z$ and set
\begin{equation}\label{eq:kernel-trim}
    \varepsilon_T=Rh,
    \qquad
    \mathcal Z_T=\big\{z\in\mathcal Z:\ d(z,\partial\mathcal Z)\ge Rh\big\},
    \qquad
    \tau_T(z)=\mathbf{1}\big\{d(z,\partial\mathcal Z)\ge Rh\big\}.
\end{equation}
Because $K_h(\cdot-z)$ is supported on $\{z':\|z'-z\|_\infty\le Rh\}$, the definition says exactly that the smoothing window of a retained point lies entirely inside $\mathcal Z$:
\begin{equation}\label{eq:window-inside}
    z\in\mathcal Z_T
    \iff
    \big\{z':\|z'-z\|_\infty\le Rh\big\}\subseteq\mathcal Z .
\end{equation}
Since $h=h_T\to0$ is nonrandom, so is $\tau_T$, as Condition~\ref{cond:consistency}\eqref{cond:cons-trim} requires.

\begin{condition}\label{cond:kernel}
    Let $\mathcal Z$ denote the support of $z_{jt}$.
\begin{enumerate}
    \item\label{cond:kernel-boundedness} $\mathcal Z$ is a compact set with Lipschitz boundary. $\bar f\ge\sup_{z\in\mathcal Z}f_j(z)\ge\inf_{z\in\mathcal Z}f_j(z)\ge\underline f>0$, where $f_j$ is the density of $z_{jt}$. $0<\underline c\le \Sigma_j(z)^{-1}\le \bar c<\infty$ for all $z\in\mathcal Z$ and $j=1,\dots,J$.
    \item\label{cond:kernel-smoothness-bw} The kernel $K:\mathbb{R}^{d_z}\to\mathbb{R}$ is a Lipschitz function. It has compact support, $\operatorname{supp}K\subseteq[-R,R]^{d_z}$ for some $R<\infty$, and order $l$, i.e.
    \[
        \int K(u)\,du=1,\qquad \int K(u)\,u^{m}\,du=0 \ \ \text{for } 1\le m<l,
    \]
    and $K_h(u)=h^{-d_z}K(u/h)$ with bandwidth $h=h_T\to0$ and $u\in\mathbb{R}^{d_z}$.
    The kernel has order $l>d_z$, and the bandwidth is $h=T^{-q}$ with $\tfrac{1}{2l}<q<\tfrac{1}{2d_z}$.
    \item\label{cond:kernel-psi}
    For each $\psi_{jt}(\sigma) \in \bigl\{y_j(\sigma, w_t),\ \partial_\sigma y_j(\sigma, w_t),\ \partial_\sigma^2 y_j(\sigma, w_t)\bigr\}$, we have $\sup_{w_t} |\psi_{jt}(\sigma_0)| < C$, and $\psi_{jt}(\sigma)$ is Lipschitz continuous in $\sigma$; that is, there exists a constant $L > 0$ such that
\[
    |\psi_{jt}(\sigma_1) - \psi_{jt}(\sigma_2)| \le L\,|\sigma_1 - \sigma_2|, \qquad \forall\, \sigma_1, \sigma_2 \in \Theta_\sigma.
\]
Moreover, for each $\psi_{jt}(\sigma) \in \bigl\{y_j(\sigma, w_t),\ \partial_\sigma y_j(\sigma, w_t),\ \partial_\sigma^2 y_j(\sigma, w_t)\bigr\}$, the map $m_{\psi,j}(\cdot\,;\sigma)$ is $l$-times continuously differentiable in $z$, with derivatives bounded uniformly over $\Theta_\sigma \times \mathcal{Z}$:
\[
    \sup_{\sigma \in \Theta_\sigma,\ z \in \mathcal{Z},\ k \le l} \bigl\|\partial_z^k m_{\psi,j}(z; \sigma)\bigr\| \le C.
\]
Finally, for each $\psi(z) \in \{\mu_{p,j}(z),\ f_j(z),\ \Sigma_j(z)\}$, the map $\psi(\cdot)$ is $l$-times continuously differentiable in $z$, with derivatives bounded uniformly over $\mathcal{Z}$:
\[
    \sup_{z \in \mathcal{Z},\ k \le l} \bigl\|\partial_z^k \psi(z)\bigr\| \le C.
\]
\end{enumerate}
\end{condition}

\begin{remark}
Condition~\ref{cond:kernel} collects the standard primitives for uniform kernel rates, together with what the trimming requires.

Part~\eqref{cond:kernel-boundedness} restricts attention to a compact support on which the instrument density is bounded away from zero. This lets the ratio $\hat\mu_{\psi,j}=\hat m_{\psi,j}/\hat f_j$ inherit the numerator's rate on $\mathcal Z_T$, since there $\hat f_j\ge\underline f/2$ with probability approaching one. It also bounds the mass discarded by the trimming: for a compact set with Lipschitz boundary the collar satisfies $Leb\{z:d(z,\partial\mathcal Z)<\varepsilon\}\le C\varepsilon$, so with $f_j\le\bar f$,
\[
    p_T=\max_j\Pr\big(\tau_T(z_{jt})=0\big)\le \bar f\,C R\,h=O(h)\longrightarrow0,
\]
which verifies Condition~\ref{cond:consistency}\eqref{cond:cons-trim}.

Part~\eqref{cond:kernel-smoothness-bw} is the higher-order-kernel construction: a compactly supported, Lipschitz kernel of order $l$ yields a bias of order $h^{l}$. The kernel order and bandwidth are chosen so that the uniform rate $\rho_T=(\log T/Th^{d_z})^{1/2}+h^{l}$ is $o(T^{-1/4})$, in fact fast enough that the estimated weight $\hat\Sigma_{j}$ also converges at $o_p(T^{-1/4})$, and keep the bias term introduced by the first stage estimation negligible; this is what the tightened window $\tfrac{1}{2l}<q<\tfrac{1}{2d_z}$ secures. The requirement $l>d_z$ keeps this interval nonempty, so a large instrument dimension $d_z$ calls for a high-order kernel---the standard cost of dimensionality here. The support radius $R$ enters no rate; it is named only to fix the trimming in \eqref{eq:kernel-trim}, and the choice $\varepsilon_T=Rh$ is what makes the  bias legitimate. The regression functions are smooth only \emph{inside} $\mathcal Z$: by part~\eqref{cond:kernel-boundedness} the density is bounded away from zero on $\mathcal Z$ and is zero outside it, so $f_j$ and hence $m_{\psi,j}=\mu_{\psi,j}f_j$ jump at $\partial\mathcal Z$. For $z$ within $Rh$ of the boundary the change of variables in the bias step integrates the kernel over the truncated region $\{v:z+vh\in\mathcal Z\}$, on which $\int K(v)\,dv\ne1$ and the higher-order moments no longer vanish; the local-constant estimator then has larger bias, and $\rho_T=o(T^{-1/4})$ fails. By \eqref{eq:window-inside} the trimming removes exactly those points and nothing more.

Part~\eqref{cond:kernel-psi} imposes boundedness and Lipschitz continuity in $\sigma$ on each regressand, and $l$-fold smoothness in $z$ on the density-weighted targets $m_{\psi,j}$, the density $f_j$, and the variance weight $\Sigma_{j}$. The Lipschitz-in-$\sigma$ property supplies the uniform-over-$\Theta_\sigma$ boundedness, while the $z$-smoothness of $\Sigma_{j}$ is what ensures the kernel estimator of $\Sigma_{j}$ inherits the same rate. Note that this implies $\mu_{\psi,j}$ is $l$-fold smooth in $z$ as well, since $\mu_{\psi,j}=m_{\psi,j}/f_j$ and $f_j$ is bounded away from zero on $\mathcal Z$.
\end{remark}

The limiting distribution of the kernel estimator then follows from Theorem~\ref{thm:normality}.
\begin{theorem}[Asymptotic normality for kernel estimator]\label{thm:kernel-normality}
Suppose Conditions~\ref{cond:regularity}, \ref{cond:y-smooth} and \ref{cond:kernel} hold, and that $M=\mathbb E_J\,\mathbb E\left[\Sigma_j(z_{jt})^{-1}\,\Gamma_{jt}\Gamma_{jt}'\right]$ is nonsingular. Then the kernel estimator $\hat\theta$ satisfies
\[
    \sqrt T\,(\hat\theta-\theta_0)\xrightarrow{d}N\!\big(0,M^{-1}V\ (M^{-1})'\big),
\]
with $V=\mathbb{E}[(\mathbb{E}_J \xi_{jt}\Sigma_j(z_{jt})^{-1}\Gamma_{jt})(\mathbb{E}_J \xi_{jt}\Sigma_j(z_{jt})^{-1}\Gamma_{jt})']$.
\end{theorem}


\begin{remark}
Theorem~\ref{thm:kernel-normality} shows the kernel estimator is asymptotically normal at the rate $T^{-1/2}$, despite the nonparametric first stage. Because the weight equals the conditional variance $\Sigma_{j}(z_{jt})=\mathbb{E}[\xi_{jt}^2\mid z_{jt}]$, under a stronger condition that $\xi_{jt}$ are mean-zero and independent across $j$ conditional on $z_{t}$, the sandwich collapses to $V=M/J$, and the asymptotic variance simplifies to $M^{-1}/J$. Then the estimator attains the semiparametric efficiency bound for the conditional-moment restriction \eqref{eq:theory-id}. The first-stage rate requirements are entirely absorbed into the bandwidth and kernel choice of Condition~\ref{cond:kernel}\eqref{cond:kernel-smoothness-bw}: it removes the first-stage bias, so no bias term enters the limit and the distribution is centered at zero. The only price is the high-order kernel needed when $d_z$ is large.
\end{remark}


\subsection{Theory for Series Estimation}\label{sec:theory-series}

We now verify the requirements of Section~\ref{sec:theory-general} for a series first stage. Fix a product $j$; the estimator projects the regressand onto a $K$-dimensional dictionary. Let $\phi^K(z)=(\phi_1(z),\dots,\phi_K(z))'$ be a vector of basis functions. For a generic regressand $\psi_{jt}(\sigma)\in\{y_j(\sigma,w_t),\ \partial_\sigma y_j(\sigma,w_t),\ \partial^2_\sigma y_j(\sigma,w_t),\ p_{jt}\}$, the least-squares series estimator of the conditional mean $\mu_{\psi,j}(z;\sigma)=\mathbb{E}[\psi_{jt}(\sigma)\mid z_{jt}=z]$ is
\[
    \hat\mu_{\psi,j}(z;\sigma)=\phi^K(z)'\hat\Pi_{\psi,j}(\sigma),
    \qquad
    \hat\Pi_{\psi,j}(\sigma)=\hat Q_j^{-1}\,\mathbb{E}_T\!\big[\phi^K(z_{jt})\,\psi_{jt}(\sigma)\big],
\]
where $\hat Q_j=\mathbb{E}_T[\phi^K(z_{jt})\phi^K(z_{jt})']$ is the empirical Gram matrix and $Q_j=\mathbb{E}[\phi^K(z_{jt})\phi^K(z_{jt})']$ its population counterpart. Specializing $\psi$ recovers the four first-stage objects of Section~\ref{sec:theory-general}: $\hat\mu_{y,j}$, the derivative estimators $\hat{\dot\mu}_{y,j}$ and $\hat{\ddot\mu}_{y,j}$—obtained by replacing $\psi=y_j(\sigma,w_t)$ with its $\sigma$-derivatives in the regressand, the design $\hat Q_j$ being free of $\sigma$—and $\hat\mu_{p,j}$. Write $e^{\psi}_{jt}(\sigma)=\psi_{jt}(\sigma)-\mu_{\psi,j}(z_{jt};\sigma)$ for the projection residual, which satisfies $\mathbb{E}[e^{\psi}_{jt}(\sigma)\mid z_{jt}]=0$. Finally, let
\[
    \zeta_0(K)=\sup_{z\in\mathcal Z}\big\|\phi^K(z)\big\|
\]
denote the envelope of the basis.

\paragraph{Trimming.} The trimming function is trivial here: $\tau_T(z)\equiv1$ for all $z\in\mathcal Z$, so that $\mathcal Z_T=\mathcal Z$. The series estimator is defined on the entire support of the instruments, and no trimming is needed.



\begin{condition}\label{cond:series}
\begin{enumerate}
    \item\label{cond:series-gram} \textup{(Design)} $\mathcal Z$ is compact, and the eigenvalues of $Q_j$ are bounded above and away from zero uniformly in $j$: $0<\underline{q}\le\lambda_{\min}(Q_j)\le\lambda_{\max}(Q_j)\le\overline{q}<\infty$. The basis satisfies the envelope bound $\zeta_0(K)<\infty$ for each $K$.  $0<\underline c\le \Sigma_j(z)^{-1}\le \bar c<\infty$ for all $z\in\mathcal Z$ and $j=1,\dots,J$.
    \item\label{cond:series-approx} \textup{(Approximation)} There is $\gamma>0$ such that, for each\\
     $\psi\in\{y_j(\sigma,w_t),\ \partial_\sigma y_j(\sigma,w_t),\ \partial^2_\sigma y_j(\sigma,w_t),\ p_{jt},\ x_{jt}',\ \xi_{jt}^2, \Sigma_{j}(z_{jt})^{-1}, \Sigma_{j}(z_{jt})^{-1}\cdot \partial_\sigma y_{j}(\sigma,w_t),\Sigma_{j}(z_{jt})^{-1}\cdot p_{jt}, \Sigma_{j}(z_{jt})^{-1}\cdot x_{jt}'\}$, there exist coefficient vectors $\Pi^{K}_{\psi,j}(\sigma)$ with
    \[
        \sup_{\sigma\in\Theta_\sigma}\ \sup_{z\in\mathcal Z}\big|\mu_{\psi,j}(z;\sigma)-\phi^K(z)'\Pi^{K}_{\psi,j}(\sigma)\big|=O(K^{-\gamma}).
    \]
    \item\label{cond:series-Lipschitz} \textup{(Lipschitz)} For each $\psi\in\{y_j(\sigma,w_t),\ \partial_\sigma y_j(\sigma,w_t),\ \partial^2_\sigma y_j(\sigma,w_t)\}$, $\sup_{w_t}|\psi_{jt}(\sigma_0)|<C$, and there exists a constant $L>0$ that does not depend on data $z_{jt}$ such that $|\psi_{jt}(\sigma_1)-\psi_{jt}(\sigma_2)|\le L|\sigma_1-\sigma_2|$ for all $\sigma_1,\sigma_2\in\Theta_\sigma$.
    \item\label{cond:series-rate} \textup{(Rates)} As $T\to\infty$, $K=K_T\to\infty$ with
    \[
        \sqrt T K^{-\gamma}=o(1),
        \qquad
        \frac{\zeta_0(K) K \log (K\vee T)}{\sqrt{T}}=o(1).
    \]
\end{enumerate}
\end{condition}

\begin{remark}
Condition~\ref{cond:series} is the series counterpart of the kernel primitives in Condition~\ref{cond:kernel} and follows the framework of \citet{newey1997convergence}.
Part~\eqref{cond:series-gram} restricts the dictionary so that the population design $Q_j$ is well-conditioned. Together with the envelope $\zeta_0(K)$, this controls the estimation error of the empirical Gram matrix, and hence ensures $\hat Q_j$ is invertible with probability approaching one, so that the linear projection is well defined. It is the series analogue of the kernel requirement that the instrument density be bounded away from zero.
Part~\eqref{cond:series-approx} is the sieve approximation requirement, imposing a common rate $K^{-\gamma}$ on the four first-stage regression functions $\mu_{y,j},\dot\mu_{y,j},\ddot\mu_{y,j},\mu_{p,j}$ and, in addition, on $x_{jt}$, the conditional variance $\Sigma_{j}(z_{jt})=\mathbb{E}[\xi_{jt}^2\mid z_{jt}]$, its inverse, the multiplications $\Sigma_j^{-1}\cdot \partial_\sigma y_{j}(\sigma,w_t)$, $\Sigma_j^{-1}\cdot p_{jt}$, and $\Sigma_j^{-1}\cdot x_{jt}'$. The last two are needed because $x_{jt}$ and $\Sigma$ enter the moment and because the feasible weight is itself a series fit of squared residuals. The rate $\gamma$ reflects the smoothness of these functions relative to the basis; for $s$-times differentiable functions of a $d_z$-dimensional instrument, $\gamma=s/d_z$ with polynomial or spline dictionaries. Particularly, if the function is analytic and univariate, the approximation error diminishes at $\exp(-K)$, with polynomial dictionaries.
Part~\eqref{cond:series-Lipschitz} is the series analogue of the kernel Lipschitz-in-$\sigma$ requirement, which is needed to control the uniformity of the first-stage convergence over $\Theta_\sigma$.
Part~\eqref{cond:series-rate} collects the two operative rate restrictions. The first, $\sqrt T\,K^{-\gamma}=o(1)$, requires the dictionary must be rich enough that the approximation bias is negligible relative to $T^{-1/2}$, so no bias term enters the limit—the series analogue of $h^{l}=o(T^{-1/2})$. The second, $\zeta_0(K)K\log (K\vee T)/\sqrt T=o(1)$, is the binding upper restriction on how fast $K$ may grow; it controls the estimation error of $\hat Q_j^{-1}$ and, through it, the second-order remainder in the linearization, playing the role that $q<1/(2d_z)$ plays for the kernel. As in the kernel case, the crude rate $T^{-1/4}$ suffices only because the moment is degenerate at the truth ($r_{jt}(\theta_0)=0$), so the error enters at the second order.
\end{remark}



\begin{theorem}[Asymptotic normality, series estimator]\label{thm:series-normality}
Suppose Conditions~\ref{cond:regularity}, \ref{cond:y-smooth} and \ref{cond:series} hold, and that $M=\mathbb E_J\,\mathbb E\left[\Sigma_j(z_{jt})^{-1}\,\Gamma_{jt}\Gamma_{jt}'\right]$ is nonsingular. Then the series estimator $\hat\theta$ satisfies
\[
    \sqrt T\,(\hat\theta-\theta_0)\xrightarrow{d}N\!\big(0,M^{-1}V\ (M^{-1})'\big),
\]
with $V=\mathbb{E}[(\mathbb{E}_J \xi_{jt}\Sigma_j(z_{jt})^{-1}\Gamma_{jt})(\mathbb{E}_J \xi_{jt}\Sigma_j(z_{jt})^{-1}\Gamma_{jt})']$.
\end{theorem}

\begin{remark}
The series estimator attains the same limiting distribution $N(0,M^{-1}V(M^{-1})')$---same Jacobian $M$ and same influence function---as the kernel estimator of Theorem~\ref{thm:kernel-normality}. The reason is that the influence function does not depend on which smoother produced it. In both implementations the projection of the first-stage error onto $a(z_{jt})$ collapses to $\mathbb{E}_T\mathbb{E}_J[\xi_{jt}a_{jt}]$---Lemma~\ref{lemma:kernel-linear} for the kernel and Lemma~\ref{lemma:series-linear} for the series---while the smoother-specific pieces, the bias $h^{l}$ or $K^{-\gamma}$ and the quadratic remainders, are held below $T^{-1/2}$ by the rate restrictions in Conditions~\ref{cond:kernel}\eqref{cond:kernel-smoothness-bw} and \ref{cond:series}\eqref{cond:series-rate}. The choice between the two is therefore governed by finite-sample and computational considerations rather than by efficiency.
\end{remark}


\section{Monte Carlo Simulation}
In this section, we provide simulation evidence to show performance of estimators across scenarios.

\begin{table}[t]
    \centering
    \begin{threeparttable}
        \caption{Simulation Results $\lambda = 0$}
        \label{tab:simulation-results-0}

        \setlength{\tabcolsep}{6pt}
        
        \begin{tabular}{lrrrrrrrr}
            \toprule
            Parameter
            & True
            & Bias
            & Abs.\ Bias
            & St.\ Err.
            & RMSE
            & Min
            & Max
            & Rej.\ Rate \\
            \midrule

            \multicolumn{9}{l}{\textit{Kernel Estimator}} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& -0.026
& 0.066
& 0.108
& 0.110
& 0.600
& 1.300
& 0.060 \\

$\alpha$
& -1.000
& 0.017
& 0.046
& 0.070
& 0.071
& -1.174
& -0.727
& 0.060 \\


\addlinespace

\multicolumn{9}{l}{\textit{Series Estimator (B-spline, df $= 5$)}} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& -0.006
& 0.050
& 0.077
& 0.076
& 0.800
& 1.100
& 0.040 \\

$\alpha$
& -1.000
& 0.003
& 0.037
& 0.051
& 0.051
& -1.090
& -0.879
& 0.040 \\


\addlinespace

\multicolumn{9}{l}{
    \textit{GMM Estimator 1: Infeasible Optimal IV}
} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& 2.462
& 2.514
& 2.944
& 3.815
& 0.800
& 10.000
& 0.200 \\

$\alpha$
& -1.000
& -1.738
& 1.774
& 2.077
& 2.692
& -7.401
& -0.863
& 0.180 \\


\addlinespace

\multicolumn{9}{l}{
    \textit{GMM Estimator 2: Feasible Optimal IV}
} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& 0.916
& 1.008
& 2.088
& 2.261
& 0.600
& 9.200
& 0.120 \\

$\alpha$
& -1.000
& -0.647
& 0.708
& 1.464
& 1.587
& -6.566
& -0.754
& 0.120 \\

            \bottomrule
        \end{tabular}

        \begin{tablenotes}
            \footnotesize
            \item \textit{Notes:}
            Abs.\ Bias denotes mean absolute bias.
            Rej.\ Rate denotes the empirical rejection rate.
            GMM Estimator 1 uses the true conditional expectation
            in the optimal instrument.
            GMM Estimator 2 uses the feasible approximation in
            \citet{berry1999voluntary}.
        \end{tablenotes}
    \end{threeparttable}
\end{table}


\begin{table}[t]
    \centering
    \begin{threeparttable}
        \caption{Simulation Results $\lambda = 0.3$}
        \label{tab:simulation-results-0.3}

        \setlength{\tabcolsep}{6pt}
        
        \begin{tabular}{lrrrrrrrr}
            \toprule
            Parameter
            & True
            & Bias
            & Abs.\ Bias
            & St.\ Err.
            & RMSE
            & Min
            & Max
            & Rej.\ Rate \\
            \midrule

\multicolumn{9}{l}{\textit{Kernel Estimator}} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& -0.084
& 0.124
& 0.239
& 0.251
& 0.100
& 1.500
& 0.100 \\

$\alpha$
& -1.000
& 0.080
& 0.111
& 0.320
& 0.327
& -1.319
& 1.152
& 0.020 \\


\addlinespace

\multicolumn{9}{l}{\textit{Series Estimator (B-spline, df $= 5$)}} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& -0.004
& 0.040
& 0.073
& 0.072
& 0.800
& 1.200
& 0.060 \\

$\alpha$
& -1.000
& 0.003
& 0.032
& 0.050
& 0.049
& -1.127
& -0.861
& 0.060 \\


\addlinespace

\multicolumn{9}{l}{
    \textit{GMM Estimator 1: Infeasible Optimal IV}
} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& 0.396
& 0.436
& 1.560
& 1.594
& 0.800
& 8.200
& 0.060 \\

$\alpha$
& -1.000
& -0.272
& 0.303
& 1.077
& 1.100
& -6.019
& -0.863
& 0.060 \\


\addlinespace

\multicolumn{9}{l}{
    \textit{GMM Estimator 2: Feasible Optimal IV}
} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& 0.470
& 0.578
& 1.923
& 1.961
& 0.100
& 9.800
& 0.060 \\

$\alpha$
& -1.000
& -0.333
& 0.400
& 1.347
& 1.375
& -7.138
& -0.585
& 0.060 \\

            \bottomrule
        \end{tabular}

        \begin{tablenotes}
            \footnotesize
            \item \textit{Notes:}
            Abs.\ Bias denotes mean absolute bias.
            Rej.\ Rate denotes the empirical rejection rate.
            GMM Estimator 1 uses the true conditional expectation
            in the optimal instrument.
            GMM Estimator 2 uses the feasible approximation in
            \citet{berry1999voluntary}.
        \end{tablenotes}
    \end{threeparttable}
\end{table}


\begin{table}[t]
    \centering
    \begin{threeparttable}
        \caption{Simulation Results $\lambda = 1$}
        \label{tab:simulation-results-1}

        \setlength{\tabcolsep}{6pt}
        
        \begin{tabular}{lrrrrrrrr}
            \toprule
            Parameter
            & True
            & Bias
            & Abs.\ Bias
            & St.\ Err.
            & RMSE
            & Min
            & Max
            & Rej.\ Rate \\
            \midrule

\multicolumn{9}{l}{\textit{Kernel Estimator}} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& -0.024
& 0.128
& 0.292
& 0.291
& 0.100
& 2.400
& 0.060 \\

$\alpha$
& -1.000
& 0.039
& 0.095
& 0.287
& 0.286
& -1.893
& 0.750
& 0.040 \\


\addlinespace

\multicolumn{9}{l}{\textit{Series Estimator (B-spline, df $= 5$)}} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& -0.012
& 0.032
& 0.063
& 0.063
& 0.800
& 1.100
& 0.040 \\

$\alpha$
& -1.000
& 0.007
& 0.025
& 0.038
& 0.038
& -1.075
& -0.867
& 0.060 \\


\addlinespace

\multicolumn{9}{l}{
    \textit{GMM Estimator 1: Infeasible Optimal IV}
} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& -0.004
& 0.040
& 0.067
& 0.066
& 0.800
& 1.100
& 0.020 \\

$\alpha$
& -1.000
& 0.002
& 0.028
& 0.040
& 0.040
& -1.076
& -0.867
& 0.020 \\


\addlinespace

\multicolumn{9}{l}{
    \textit{GMM Estimator 2: Feasible Optimal IV}
} \\
\cmidrule(lr){1-9}

$\sigma$
& 1.000
& 0.004
& 0.084
& 0.194
& 0.192
& 0.100
& 1.800
& 0.040 \\

$\alpha$
& -1.000
& -0.006
& 0.052
& 0.109
& 0.108
& -1.527
& -0.585
& 0.040 \\


            \bottomrule
        \end{tabular}

        \begin{tablenotes}
            \footnotesize
            \item \textit{Notes:}
            Abs.\ Bias denotes mean absolute bias.
            Rej.\ Rate denotes the empirical rejection rate.
            GMM Estimator 1 uses the true conditional expectation
            in the optimal instrument.
            GMM Estimator 2 uses the feasible approximation in
            \citet{berry1999voluntary}.
        \end{tablenotes}
    \end{threeparttable}
\end{table}

\subsection{Simulation Design}

In the simulations, we illustrate the finite-sample performance of different estimators when the true data-generating process resembles Example \ref{eg:counter}. We draw samples of size $T=500$ with $J=2$. The mean utility takes the form $y_{jt} = p_{jt} \alpha_0 + \xi_{jt}$. The random coefficient for $p_{jt}$ follows $N(0,\sigma_0^2)$. We set the true parameters $\alpha_0 = -1$ and $\sigma_0 = 1$. We draw $\xi_{jt} \sim \text{U}[-1,1]$, independent of the instrument $z_{jt}$, with $p_{jt} = z_{jt} + 0.3 \xi_{jt}$. The instrument $z_{j,t}$ is generated from a mixed distribution: denoting $F_z \coloneqq \frac{5}{11} N(-2,0.0625) + \frac{1}{11} N(1, 0.0625) + \frac{5}{11} N(2,0.0625)$, we draw $z_{1t} \sim F_z$ and $z_{2t} \sim \lambda U[-2,2] + (1-\lambda)F_z$. We consider three configurations with $\lambda=0$, $0.3$, and $1$. Notice that when $\lambda=1$, the distribution of $z_{jt}$ is discrete, similar to the distribution of $p_{jt}$ in Example \ref{eg:counter} with no endogeneity. Our setup therefore illustrates performance of estimators under various generalization of Example \ref{eg:counter}.

For each of the three configurations, we implement estimation of $(\sigma_0, \alpha_0)$ with a kernel estimator, a series estimator, and two GMM estimators. The kernel and series estimators are based on the conditional moment. The two GMM estimators are based on the unconditional moment with respectively the infeasible optimal instrument and a feasible optimal instrument. We implement the estimation in each of the $50$ simulations.

More specifically, we construct the kernel estimator using the sixth-order Epanechnikov kernel, and construct the series estimator using B-splines. For the first GMM estimation, we adopt the infeasible optimal instrument as
\begin{align*}
    \left( \mathbb{E}\left[ x_{jt} \mid z_{jt} \right], \mathbb{E}\left[\frac{\partial y_j}{\partial \sigma} \left( \sigma_0, x_{1t}, x_{2t}\right) \mid z_{jt}  \right]   \right)'= \left( z_{jt}, \mathbb{E}\left[\frac{\partial y_j}{\partial \sigma} \left( \sigma_0, x_{1t}, x_{2t}\right) \mid z_{jt}  \right]   \right)',
\end{align*}
for $j=1,2$. Here we approximate the conditional expectation by averaging over $200$ draws for each grid point of $z_{jt}$ and interpolate the conditional expectation between grid points. For the second GMM estimation, we follow the feasible approximation to the optimal instrument suggested by \citet{berry1999voluntary}





\subsection{Results}
Table \ref{tab:simulation-results-0} summarizes the results under $\lambda=0$. In this case, $z_{jt} \sim F_z$ and thus the scenario is similar to Example \ref{eg:counter}. The kernel and series estimators exhibit small biases and low root mean square errors (RMSEs), indicating good performance in estimating the parameters. In contrast, the GMM estimators show systematically larger biases and RMSEs, suggesting problem of identification of unconditional moments with the optimal instrument, both for the infeasible and the feasible optimal instrument. The empirical rejection rates for the kernel and series estimators are also close to the nominal level, while the GMM estimators display invalid inference at the true parameter values.

Table \ref{tab:simulation-results-0.3} summarizes the results under $\lambda=0.3$, with $z_{2t}$ having a probability of $0.3$ for taking values from $U[-2,2]$. Although this setup is more different from Example \ref{eg:counter}, larger biases are still displayed in GMM estimators. The kernel and series estimators maintain smaller bias in this case. Under $T=500$ across $50$ Monte Carlo draws, the rejection rate of the kernel estimator displays unstability, but the bias and RMSE is still much smaller than GMM estimators. The series estimator achieves both smaller bias and currect rejection rate.

Table \ref{tab:simulation-results-1} summarizes the results under $\lambda=1$. Here $z_{2t}\sim U[-2,2]$. With the distribution of $z_{jt}$ far from Example \ref{eg:counter}, here the systematically larger biases in GMM estimators disappear. All four estimators display good performance in biases and rejection rates. This serves as an important baseline check, showing the source of the GMM bias is mainly the theoretical property, rather than the computational implementation.

Overall, the simulation results are consistent with our theory for the conditional moment methods, illustrating cases where the unconditional methods are limited while the conditional moment methods still apply.



\section{Conclusion}\label{sec:conclusion}
This paper revisits estimation and inference in the random-coefficients demand model of \cite{berry1995automobile} through its identifying assumption. We begin from the observation that the conditional moment restriction $\mathbb{E}[\xi_{jt}(\theta)\mid z_{jt}]=0$ on which the prevalent identification rests is not equivalent to the unconditional restriction $\mathbb{E}[\xi_{jt}(\theta)\,z_{jt}]=0$ that standard GMM exploits: the unconditional restriction may admit additional parameter values. We give a counterexample in which the model is identified by the conditional restriction yet standard GMM fails to recover $\theta_0$. This gap motivates an estimator built directly on the conditional restriction.

The proposed estimator is a two-step procedure. The first step estimates the relevant conditional expectations nonparametrically; the second chooses the structural parameters to minimize a conditional-variance-weighted sum of squared conditional-moment residuals. Standard GMM arises as the special case in which the conditional expectation is replaced by a linear projection onto a fixed, finite set of instruments.

Our main results establish that this estimator is $\sqrt T$-consistent and asymptotically normal, and we provide two concrete implementations of the first stage—kernel and series—each with primitive conditions under which the limiting distribution holds. Three features of the theory are worth emphasizing.
First, the two implementations share the same limiting distribution and the same influence function, which depends only on the unobserved product characteristics $\xi_{jt}$. This is not because the first-stage error washes out. With the true conditional expectations the residual is identically zero and the criterion is minimized at $\theta_0$ in every sample, so the first-stage error is in fact the whole source of the limiting variance; what the theory establishes is that it enters only through its projection onto the weighted direction $a_{jt}=\Sigma_j(z_{jt})^{-1}\Gamma_{jt}$.
Second, because the moment is degenerate at the truth, the conditional expectation may be estimated with a slow nonparametric rate without disturbing the first-order distribution, so a simple plug-in approach suffices.
Third, with the weight set to the inverse conditional variance, the estimator is efficient for the conditional-moment problem under the additional assumption of within-market independence. The bandwidth and sieve-dimension requirements needed for these conclusions are collected in transparent rate conditions, and the same plug-in variance estimator delivers standard errors for both implementations.

Several extensions are natural. The framework accommodates any first-stage nonparametric estimator satisfying the high-level rate and linearity conditions, so alternatives such as penalized or machine-learning first stages could be analyzed within the same structure. The analysis also takes the market shares as observed; incorporating the sampling and simulation error studied by \cite{freyberger2015asymptotic} into the present conditional-moment framework is left for future work.





\bibliography{bib.bib}