Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Kernel Minimum Distance Estimation and Testing with Conditional Moment Restrictions: A Unified Framework
bibunit\begin{spacing}{1.2}
\begin{abstract}
We propose a unified Kernel Minimum Distance (KMD) framework for estimating and testing models defined by conditional moment restrictions. By embedding conditional moments into a Reproducing Kernel Hilbert Space (RKHS), we construct a closed-form $V$-statistic objective function that quantifies the distance from the restrictions. We establish the $\sqrt{n}$-consistency and asymptotic normality of the associated minimum distance estimator. Within this framework, the minimized objective function naturally yields a consistent omnibus specification test. Unlike projection-based methods that require auxiliary nonparametric estimation for Neyman orthogonalization, our test inherently captures the estimation effect via a projected kernel structure. We derive asymptotic properties of the test statistics under the null hypothesis, the alternative hypothesis, and a sequence of local alternatives converging to the null at the parametric rate $n^{-1/2}$. The validity of a computationally simple multiplier bootstrap is established to facilitate inference. Simulation results demonstrate robust finite-sample performance, and the framework is illustrated by analyzing Engel curves using UK Family Expenditure Survey data.
\noindentKeywords: Kernel Minimum Distance, Conditional Moment Restrictions, Reproducing Kernel Hilbert Space
\noindentJEL Classification Number: C12; C13; C14; C26.
\end{abstract}
\end{spacing}
\thispagestyle{empty}
\section{Introduction}
Conditional moment restrictions (CMRs) serve as a fundamental framework in econometric analysis, encompassing a wide range of empirical settings, including instrumental variable estimation, nonlinear regression, and structural modeling. In these applications, economic theory typically postulates a parametric structural relationship that implies moment conditions of the form:
\begin{equation}
\mathbb{E}\left[\rho\left(Z,\theta_0\right)|X\right]=0 \quad almost surely (a.s.),
\end{equation}
for some unknown true parameter $\theta_0 \in \Theta$. Here, $Z \in \mathcal{Z}$ and $X \in \mathcal{X}$ denote the observable variables, and $\rho: \mathcal{Z} \times \Theta \to \mathbb{R}^q$ is a pre-specified generalized residual function known up to the parameter $\theta \in \Theta \subset \mathbb{R}^p$, which encodes the structural restrictions imposed by the model. Consequently, empirical researchers face two closely related tasks: consistently estimating the structural parameter $\theta_0$ and assessing whether the assumed structural model is correctly specified. Without the latter, the economic interpretation of the estimated parameters remains questionable.
The standard approach to estimation under conditional moment restrictions is the Generalized Method of Moments (GMM; hansen1982large), which approximates the conditional restriction by a finite vector of unconditional moments. While computationally convenient and widely used, this finite reduction may lead to fundamental identification problems. When the conditioning variables have support with infinite cardinality, the conditional restriction implies an infinite collection of unconditional moment conditions, and satisfying only a finite subset does not ensure global identification in nonlinear models, even when optimal instruments are employed dominguez2004consistent. As a consequence, GMM estimators may be inconsistent even when the underlying conditional restriction is correctly specified. A closely related limitation concerns specification testing. The classical $J$-test sargan1958estimation,hansen1982large is not omnibus and is consistent only against violations of the selected moment conditions newey1985generalized.
To restore global identification, the literature has evolved to exploit the infinite information inherent in conditional restrictions. Prominent contributions in this vein include sieve-based methods that approximate conditional moments using an expanding sequence of basis functions donald2003empirical and kernel-based approaches that rely on local smoothing kitamura2004empirical. Closely related to the latter is lavergne2013smooth, who develop a smooth minimum distance estimator and uniform-in-bandwidth inference for conditional estimating equations.
While many of these methods are theoretically capable of achieving semiparametric efficiency, their implementation often involves choosing a dimension, bandwidth, or regularization parameter to manage the bias-variance trade-off. Consequently, implementation may entail carefully selecting such objects; for instance, determining the optimal number of basis functions in sieve estimation requires complex data-driven procedures donald2009choosing. A distinct strategy leverages the Integrated Conditional Moment (ICM) approach, pioneered by bierens1982consistent. This framework rests on the fundamental result that the original conditional restriction is equivalent to satisfying a continuum of unconditional moment restrictions, obtained by integrating the conditional moments with respect to a weighting function. This equivalence has been extensively developed for consistent estimation carrasco2000generalization, dominguez2004consistent and omnibus specification testing bierens1990consistent, bierens1997asymptotic, delgado2006consistent, escanciano2006consistent, dominguez2015simple. However, despite their theoretical appeal, the practical implementation of ICM methods remains challenging. The test statistics are typically constructed by integrating empirical processes over a continuum of indices or weighting functions, a step that may lack an analytic solution. Closed-form expressions exist only for restrictive choices of weighting functions bierens1982consistent, escanciano2006consistent, otherwise necessitating computationally intensive numerical integration. Furthermore, because the asymptotic distributions of these omnibus statistics are often non-standard and data-dependent, obtaining valid critical values for many of these procedures requires complex adjustments or computationally intensive resampling methods, which may limit their routine use in applied econometric work.
In the spirit of the ICM approach, we propose a unified Kernel Minimum Distance (KMD) framework for both the estimation and specification testing of models defined by conditional moment restrictions. Rather than working with a pre-specified indexed family of transformations or weighting functions, KMD takes the relevant test functions to be the unit ball of a vector-valued Reproducing Kernel Hilbert Space (RKHS), so that the criterion measures the largest normalized violation of the conditional moment restriction over this RKHS class. The kernel's reproducing property then turns this infinite-dimensional criterion into a closed-form pairwise $V$-statistic objective function. Thus, the kernel plays two roles: it determines the class of moment violations against which the conditional restriction is assessed, and it delivers a tractable sample objective for estimation and testing. We establish the $\sqrt n$-consistency and asymptotic normality of the resulting estimator and discuss a one-step Newton--Raphson update toward semiparametric efficiency. Building directly on the same estimation framework, we construct a consistent omnibus specification test based on the minimized objective function. Under the null hypothesis, the limiting distribution of the test statistic automatically accounts for the first-order estimation effect via a projected-kernel structure, yielding an automatic Neyman-type orthogonalization without the need to construct a separate orthogonalized testing process. We further establish consistency against fixed alternatives, examine power against sequences of local alternatives converging to the null at the parametric rate $n^{-1/2}$, and establish the validity of a simple multiplier bootstrap procedure that is computationally efficient and avoids repeated parameter re-estimation or resampling observations.
Our framework is related to two strands of work: RKHS-based measures of distributional and moment discrepancies, and minimum-distance approaches to conditional moment restrictions in econometrics. In statistics and machine learning, the use of RKHS norms to measure discrepancies is closely related to the maximum mean discrepancy of gretton2006kernel,gretton2012kernel; for a comprehensive review of kernel mean embeddings and their applications, see muandet2017kernel. For conditional moment restrictions, on the estimation front, muandet2020kernel formulate the Maximum Moment Restriction (MMR) principle, and zhang2023instrumental recently apply a related kernel maximum-moment loss to instrumental variable regression and establish basic asymptotic properties. Their framework is primarily formulated as a learning problem: the emphasis is on constructing and optimizing a kernel-based empirical risk for IV regression, with particular attention to algorithmic implementation and predictive performance. We complement this literature by tailoring the estimator to structural inference and providing valid asymptotic inference based on consistent variance estimation.
The closest econometric precursor on the estimation side is the smooth minimum distance approach of lavergne2013smooth. Their criterion is derived from local smoothing of conditional estimating equations, and one of their key insights is that consistent estimation of the structural parameter does not require consistent nonparametric estimation of the local conditional moment itself. Our criterion is closely related in continuous Euclidean settings with translation-invariant smoothing kernels, but it starts from a different primitive object: the RKHS norm of an embedded conditional moment. This formulation makes explicit the role of the kernel in measuring moment violations over the conditioning space and allows kernels to be specified directly for the variables being conditioned on. For example, when the conditioning variables contain both continuous covariates and discrete indicators, as is common in econometric practice, one may combine a Gaussian kernel for the former with a categorical kernel for the latter.
On the testing front, existing literature similarly harnesses the geometry of RKHS to aggregate conditional moment restrictions. muandet2020kernel primarily focus on regression settings where the structural parameter $\theta_0$ is assumed known. Relatedly, sancetta2022testing develops RKHS score-type tests for functional subspace restrictions in the presence of high- or infinite-dimensional nuisance parameters, using estimated projections of test functions to remove nuisance effects. escanciano2024gaussian extends the RKHS/Gaussian-process approach to the composite hypothesis setting but adopts a decoupling strategy based on an exogenous estimator. This estimator-agnostic framework allows for estimators with slower convergence rates and dispenses with the requirement of an asymptotic linear Bahadur representation. While this flexibility is advantageous for high-dimensional nuisance parameters, it can require additional orthogonalization steps in standard structural applications. Specifically, achieving such robustness may require an auxiliary estimate of the relevant conditional score or nuisance projection whenever the gradient is not measurable with respect to the $\sigma$-algebra generated by the conditioning covariates -- a common feature in models with endogeneity, such as instrumental variable regressions. In contrast, our unified KMD framework defines the test statistic directly as the minimized value of the estimation objective function. In this setting, the parameter estimation uncertainty is captured by the statistic's geometry through a projected kernel derived from the estimator's first-order conditions, thereby enabling asymptotic inference without constructing a separate orthogonalized testing process.
Conceptually, our approach aligns with the unified paradigm established by dominguez2004consistent and dominguez2015simple. These works elegantly link estimation and testing by interpreting the minimized value of an integrated moment objective function as an omnibus specification test. However, their reliance on indicator-weighted empirical processes renders the methodology susceptible to the “curse of dimensionality” and leads to rapid power loss as the number of covariates increases. We advance this unified framework by substituting indicator functions with reproducing kernels. By defining the estimator $\widehat{\theta}_n$ as the minimizer of the KMD objective, we retain the computational simplicity of their framework, particularly in nonlinear and endogenous settings, thereby yielding a statistic that can be interpreted as a generalized Sargan test. Crucially, this kernel-based formulation leverages geometric adaptability to significantly enhance the robustness of both estimation and testing in multivariate settings.
The remainder of the paper is organized as follows. Section (ref) presents the identification strategy via an RKHS embedding, defines the KMD estimator, establishes its asymptotic normality, develops asymptotic inference for individual parameters and for linear and nonlinear restrictions, and discusses a two-step update that attains semiparametric efficiency. Section (ref) develops an omnibus specification test based on the minimized objective function, derives its limiting distribution under the null hypothesis, establishes consistency against fixed alternatives, and analyzes power against local alternatives. Section (ref) introduces a computationally efficient multiplier bootstrap procedure and establishes its asymptotic validity. Section (ref) investigates the finite-sample performance of the proposed framework through extensive Monte Carlo simulations. Section (ref) illustrates the method's empirical relevance by revisiting the structural analysis of Engel curves. Section (ref) concludes. Proofs of the main results and additional numerical and empirical analyses are provided in the Online Supplementary Material.
Notation. Throughout the paper, $\|\cdot\|$ denotes the Euclidean norm for vectors and the Frobenius norm for matrices. The superscript $\prime$ denotes the transpose of a vector or matrix. $\mathbf{I}_q$ represents the $q \times q$ identity matrix. We denote convergence in probability and in distribution by $\xrightarrow{p}$ and $\xrightarrow{d}$, respectively. The stochastic order symbols $O_p(\cdot)$ and $o_p(\cdot)$ are defined in the standard sense.
\section{Estimation}
\subsection{Identification Strategy via RKHS Embedding}
In this section, we consider the estimation of the structural parameter $\theta_0$ defined by the conditional moment restrictions in (ref). In the spirit of the ICM approach, we embed the conditional moment restrictions into a vector-valued RKHS. By the Moore--Aronszajn theorem, every symmetric positive definite kernel $k$ on $\mathcal X$ induces a unique Hilbert space $\mathcal H_k$ of real-valued functions on $\mathcal X$. Its reproducing property, $h(x)=\langle h,k(\cdot,x)\rangle_{\mathcal H_k}$ turns evaluation at $x$ into an inner product. This property allows function-indexed moment restrictions to be represented through kernel evaluations. We do not require the unknown conditional mean $x\mapsto \mathbb E[\rho(Z,\theta)| X=x]$ to belong to the RKHS. Rather, the RKHS provides the class of test functions and the geometry used to measure violations of the conditional moment restriction. Identification rests on the equivalence between the conditional restriction (ref) and the vanishing norm of the corresponding moment embedding within the RKHS. To formalize this framework, we introduce the following assumptions.
\begin{assumption}
\quad
\begin{itemize}
{0pt}
{0pt}
{0pt}
• The observed data $\{W_i\}_{i=1}^n$ is an i.i.d. sample of a random vector $W$. The vector $W_i$ contains subvectors $Z_i$ and $X_i$, which are not necessarily disjoint.
• The function $\rho(z, \theta): \mathcal{Z} \times \Theta \to \mathbb{R}^q$ is continuous at each $\theta \in \Theta$ with probability one, and satisfies $\mathbb{E}[\sup_{\theta \in \Theta}\|\rho(Z, \theta)\|^2]<\infty$.
• The parameter space $\Theta \subset \mathbb{R}^p$ is compact. There exists a unique $\theta_0 \in \Theta$ such that $\mathbb{E}[\rho(Z, \theta_0)| X]=0$ a.s.
• The conditioning space $\mathcal X$ is a separable metric space equipped with its Borel $\sigma$-field. The kernel $k:\mathcal X\times\mathcal X\to\mathbb R$ is symmetric, continuous, bounded, and positive definite. Moreover, $k$ is strictly positive definite over finite signed Borel measures: for every nonzero finite signed Borel measure $\nu$ on $\mathcal X$,
$$
\iint_{\mathcal X\times\mathcal X}
k(x,\widetilde x)\,d\nu(x)\,d\nu(\widetilde x)>0 .
$$
We define the $q\times q$ matrix kernel as $K(x_1,x_2)\coloneqq k(x_1,x_2)\cdot \mathbf I_q$.
\end{itemize}
\end{assumption}
Assumptions (ref)(i) and (ii) are standard regularity conditions. Assumption (ref)(iii) imposes global identification of the structural parameter. Assumption (ref)(iv) imposes injectivity of the RKHS embedding at the level of finite signed measures,\footnote{This requirement is stronger than the usual characteristic-kernel property, which concerns injectivity of the kernel mean embedding over probability measures. See sriperumbudur2011universality for detailed characterizations.} which is the relevant requirement for conditional moment restrictions. Indeed, for each component of the generalized residual, the map $A\mapsto E[\rho_l(Z,\theta)\mathbf 1\{X\in A\}]$ defines a finite signed measure on $\mathcal X$. Thus, the vanishing of the RKHS embedding recovers the conditional moment restriction only if the kernel embedding is injective for such signed measures. This formulation is not tied to a Euclidean conditioning space or to the existence of a Lebesgue density for $X$. In the standard
continuous-covariate setting, the condition is satisfied by many familiar bounded translation-invariant kernels whose spectral measure has full support, including Gaussian, Laplace, and Mat\'ern kernels. In settings with discrete or mixed conditioning variables, one may instead use kernels defined directly on the relevant support, for instance, products of a continuous kernel and a strictly positive-definite categorical kernel. Hence Assumption (ref)(iv) accommodates the types of continuous, discrete, and mixed conditioning variables commonly encountered in econometric applications.
Let $\mathcal H_K$ denote the vector-valued RKHS associated with the matrix kernel $K$ defined in Assumption (ref)(iv).\footnote{Since the matrix kernel has a diagonal structure
$K(x,\widetilde x)=k(x,\widetilde x)\mathbf I_q$, the vector-valued RKHS
$\mathcal H_K$ can be identified with the Hilbert direct sum of $q$ copies
of the scalar RKHS $\mathcal H_k$ generated by $k$. Equivalently,
$\mathcal H_K
=
\left\{
h=(h_1,\ldots,h_q)^\prime:
h_l\in\mathcal H_k,\ l=1,\ldots,q
\right\}$, with inner product
$\langle h,g\rangle_{\mathcal H_K}
=
\sum_{l=1}^q \langle h_l,g_l\rangle_{\mathcal H_k}$.
Thus each element $h\in\mathcal H_K$ is a vector-valued function
$h:\mathcal X\to\mathbb R^q$.} We next relate the KMD construction to the traditional ICM approach. The ICM approach converts a conditional moment restriction into a collection of unconditional moment restrictions indexed by test functions of the conditioning variables. Specifically, if $\mathbb E[\rho(Z,\theta)\mid X]=0$ a.s., then for every admissible vector-valued test function $h$,
$$
\mathbb E[\rho(Z,\theta)^\prime h(X)]
=
\mathbb E\!\left[
\mathbb E[\rho(Z,\theta)\mid X]^\prime h(X)
\right]
=0.
$$
The converse implication requires the test-function class to be sufficiently rich. Our KMD formulation implements this idea by taking the admissible test-function class to be the unit ball of $\mathcal H_K$. For each fixed $\theta\in\Theta$, define the linear functional $\mathcal L_\theta:\mathcal H_K\to\mathbb R$ by $\mathcal L_\theta h\coloneqq \mathbb E[\rho(Z,\theta)^\prime h(X)]$. Assumptions (ref)(ii) and (iv) guarantee that $\mathcal L_\theta$ is bounded. Its operator norm is
$$
\|\mathcal L_\theta\|_{\mathrm{op}}
\coloneqq
\sup_{\|h\|_{\mathcal H_K}\le 1}
|\mathcal L_\theta h|
=
\sup_{\|h\|_{\mathcal H_K}\le 1}
\left|
\mathbb E[\rho(Z,\theta)^\prime h(X)]
\right|.
$$
Thus, $\|\mathcal L_\theta\|_{\mathrm{op}}$ measures the largest normalized unconditional moment violation over the RKHS class. When the conditional moment restriction holds at $\theta$, $\mathcal L_\theta$ is the zero functional and $\|\mathcal L_\theta\|_{\mathrm{op}}=0$. Conversely, the implication from $\|\mathcal L_\theta\|_{\mathrm{op}}=0$ back to the conditional moment restriction relies on the richness of the RKHS class, ensured here by the signed-measure injectivity condition in Assumption (ref)(iv).
By the Riesz Representation Theorem, there exists a unique element $\mu(\theta)\in\mathcal H_K$, referred to as the conditional moment embedding (CME) muandet2017kernel,muandet2020kernel, such that $\mathcal L_\theta h=\langle h,\mu(\theta)\rangle_{\mathcal H_K}$ for all $h\in\mathcal H_K$. Consequently, $\|\mathcal L_\theta\|_{\mathrm{op}}=\|\mu(\theta)\|_{\mathcal H_K}$. The following lemma formalizes the identification content of this RKHS representation.
\begin{lemma}
Suppose Assumption (ref) holds. Then, for any $\theta\in\Theta$,
$$
\mathbb{E}[\rho(Z,\theta)\mid X]=0 \;\text{a.s.}
\iff
\|\mu(\theta)\|_{\mathcal H_K}=0.
$$
\end{lemma}
Based on Lemma (ref), we define the population objective function as the squared RKHS norm $Q(\theta)\coloneqq\|\mu(\theta)\|_{\mathcal H_K}^2$. To obtain a computable expression for this criterion, note that the reproducing property implies, for every $h\in\mathcal H_K$,
\begin{equation}
\mathcal L_\theta h
=
\mathbb E[\rho(Z,\theta)^\prime h(X)]
=
\mathbb E\!\left[
\left\langle h,K(\cdot,X)\rho(Z,\theta)
\right\rangle_{\mathcal H_K}
\right]
=
\left\langle
h,\mathbb E[K(\cdot,X)\rho(Z,\theta)]
\right\rangle_{\mathcal H_K}.
\end{equation}
By the uniqueness of the Riesz representer, this identifies $\mu(\theta)=\mathbb E[K(\cdot,X)\rho(Z,\theta)]$. Hence, the RKHS norm admits the closed-form double expectation
\begin{align}
Q(\theta)
&=
\left\langle
\mathbb E[K(\cdot,X)\rho(Z,\theta)],
\mathbb E[K(\cdot,\widetilde X)\rho(\widetilde Z,\theta)]
\right\rangle_{\mathcal H_K}
\nonumber \\
&=
\mathbb E\!\left[
\rho(Z,\theta)^\prime K(X,\widetilde X)\rho(\widetilde Z,\theta)
\right],
\end{align}
where $(Z,X)$ and $(\widetilde Z,\widetilde X)$ are independent copies from the population distribution.
Specifically, for a continuous translation-invariant scalar kernel on
\(\mathbb R^d\), \(k(x,\widetilde x)=\varphi(x-\widetilde x)\), the
connection with characteristic-function-based ICM criteria can be made
explicit. By Bochner's theorem, there exists a finite nonnegative Borel
measure \(\Lambda\) on \(\mathbb R^d\) such that
\[
k(x,\widetilde x)
=
\int_{\mathbb R^d}
\exp\!\left\{\mathrm{i}\omega^\prime(x-\widetilde x)\right\}
d\Lambda(\omega).
\]
Substituting this representation into (ref) and
applying Fubini's theorem and the independence of the two copies yields
\[
Q(\theta)
=
\int_{\mathbb R^d}
\left\|
\mathbb E\!\left[
\rho(Z,\theta)\exp(\mathrm{i}\omega^\prime X)
\right]
\right\|^{2}
d\Lambda(\omega).
\]
Thus, in this special case, KMD aggregates a continuum of
characteristic-function moments, with the kernel determining how different
frequencies are weighted through the spectral measure \(\Lambda\). This
provides a direct link to Bierens-type ICM criteria and to the Fourier-moment
representation underlying the SMD construction of
lavergne2013smooth; see also muandet2020kernel. The
representation is interpretive rather than computational:
\(Q(\theta)\) is evaluated directly through the closed-form pairwise
expression in (ref), without numerical integration
over \(\omega\).
Regardless of this special-case representation,
\(Q(\theta)=\|\mu(\theta)\|_{\mathcal H_K}^{2}\geq 0\), and
Lemma (ref) establishes that equality holds if and
only if
\(\mathbb E[\rho(Z,\theta)\mid X]=0\) a.s. By the global identification
condition in Assumption (ref)(iii), this occurs if and only if
\(\theta=\theta_0\). Hence, the true parameter is uniquely identified as
the global minimizer of the population objective function. This directly
motivates our estimation strategy, which minimizes the sample analog of
(ref).
\subsection{The KMD Estimator}
To implement the estimation, we construct the sample counterpart of the population objective $Q(\theta)$ using a $V$-statistic formulation:
\begin{equation*}
\widehat{Q}_n(\theta) = \frac{1}{n^2}\sum_{i=1}^{n}\sum_{j=1}^{n}\rho(Z_i,\theta)^{\prime} K(X_i,X_j)\rho(Z_j,\theta).
\end{equation*}
The Kernel Minimum Distance (KMD) estimator is then defined as:
\begin{equation*}
\widehat{\theta}_n = \arg\min_{\theta \in \Theta} \widehat{Q}_n(\theta).
\end{equation*}
\begin{remark}
We use the \(V\)-statistic criterion \(\widehat Q_n(\theta)\) rather than the off-diagonal \(U\)-statistic because it preserves the minimum-distance geometry. Specifically, \(\widehat Q_n(\theta)=\|\widehat\mu_n(\theta)\|_{\mathcal H_K}^2\), where \(\widehat\mu_n(\theta):=n^{-1}\sum_{i=1}^n K(\cdot,X_i)\rho(Z_i,\theta)\), so \(\widehat Q_n(\theta)\) is the squared RKHS norm of the empirical embedding. Since the kernel matrix is positive semidefinite, \(\widehat Q_n(\theta)\ge0\) globally, preserving the metric interpretation and improving the finite-sample optimization geometry. Off-diagonal \(U\)-statistic criteria are natural in smooth minimum-distance and local-smoothing approaches, such as lavergne2013smooth; our use of the \(V\)-statistic reflects a finite-sample geometric consideration rather than an asymptotic distinction. In particular, because \(K-\operatorname{diag}(K)\) need not be positive semidefinite, the off-diagonal criterion need not define a squared norm and may exhibit negative regions or boundary minima in finite samples. The diagonal bias is only \(O(n^{-1})\) and vanishes asymptotically, so consistency of \(\widehat\theta_n\) is unaffected; see Section (ref) of the Online Supplementary Material for further discussion.
\end{remark}
The following lemma formally establishes this consistency result.
\begin{lemma}
Under Assumptions (ref), $\widehat{\theta}_n \xrightarrow{p} \theta_0$.
\end{lemma}
To establish the asymptotic normality of $\widehat{\theta}_n$, we impose the following regularity conditions necessary to validate the Taylor expansion of the first-order condition and to ensure the validity of the resulting asymptotic linear representation.
\begin{assumption}
\quad
\begin{itemize}
{0pt}
{0pt}
{0pt}
• $\theta_0\in\operatorname{int}\left(\Theta\right)$.
• $\rho(z, \theta)$ is twice continuously differentiable with respect to $\theta$ in an open neighborhood $\mathcal{N}\left(\theta_0\right)$ of $\theta_0$, with probability one. Let $g(z, \theta) = \partial \rho(z, \theta)/\partial \theta^{\prime}$ denote the score function with dimension $q \times p$. $\mathbb{E}\left(\sup _{\theta \in \mathcal{N}\left(\theta_0\right)}\|g(Z, \theta)\|^2\right)<\infty$, $\mathbb{E}\left[\left\| \rho\left(Z, \theta_0\right)\right\|^4\right]<\infty$, $\mathbb{E}\left[\left\| g\left(Z, \theta_0\right)\right\|^4\right]<\infty$ and the Hessian matrix $\mathbb{E}\left(\sup _{\theta \in \mathcal{N}\left(\theta_0\right)} \|\nabla_{\theta \theta^{\prime}}^2 \rho_l\left(Z, \theta\right) \|^2\right)<\infty$, for $l=1, \cdots q$, where $\rho_l\left(\cdot\right)$ denotes the $l$-th element in $\rho\left(\cdot\right)$.
• The $p \times p$ matrix $\Delta\left(\theta_0\right) = \mathbb{E}\left[g(Z, \theta_0)' K(X, \tilde{X}) g(\tilde{Z}, \theta_0)\right]$ is non-singular.
\end{itemize}
\end{assumption}
Assumption (ref) outlines the regularity conditions for the asymptotic representation. Condition (i) requires $\theta_0$ to be an interior point to validate the expansion, while condition (iii) guarantees the invertibility of the population Hessian. In condition (ii), the moment bounds involving the supremum are specifically imposed to establish the uniform convergence of the sample Hessian matrix $\widehat{H}_n(\theta)$ via the Uniform Weak Law of Large Numbers (UWLLN) for $V$-statistics. The finite fourth moments for $\rho(Z, \theta_0)$ and $g(Z, \theta_0)$ serve as sufficient conditions to ensure $\mathbb{E}[\|g(Z, \theta_0)'\rho(Z, \theta_0)\|^2] < \infty$, which enables the H\'{a}jek projection of the $V$-statistic score vector $\nabla_\theta \widehat{Q}_n(\theta_0)$. We explicitly retain these separate fourth-moment conditions to maintain a unified framework with the specification testing procedure, which relies on an analogous $V$-statistic formulation (see Section (ref) of the Online Supplementary Material for a detailed discussion of $U$- versus $V$-statistics for testing). Notably, these moment requirements are slightly stronger than standard second-moment conditions; this is a direct consequence of employing the $V$-statistic estimator, which necessitates control over the stochastic behavior of the diagonal terms ($i=j$). With these regularity conditions in place, the following lemma establishes the asymptotic linear representation of $\widehat{\theta}_n$.
\begin{lemma}
Under Assumptions (ref) and (ref), the kernel minimum distance estimator $\widehat\theta$ admits the asymptotic linear representation
$$
\sqrt{n}\left(\widehat{\theta}_n-\theta_0\right)=-\Delta\left(\theta_0\right)^{-1} \frac{1}{\sqrt{n}} \sum_{i=1}^n G\left(X_i, \theta_0\right)^\prime \rho\left(Z_i, \theta_0\right)+o_p(1),
$$
where $G\left(X_i, \theta_0\right)=\mathbb{E}\left[ K\left(X_i, X_j\right)g\left(Z_j, \theta_0\right)|X_i\right]_{q \times p}$.
\end{lemma}
Applying the Central Limit Theorem to Lemma (ref) yields:
\begin{theorem}
Under Assumptions (ref) and (ref),
$$
\sqrt{n}\left(\widehat{\theta}_n-\theta_0\right) \xrightarrow{d} \mathcal{N}\left(0, \Delta \left(\theta_0\right)^{-1}\Omega\left(\theta_0\right) \Delta\left(\theta_0\right)^{-1}\right),
$$
where $\Omega\left(\theta_0\right)=\mathbb{E}\left[G\left(X, \theta_0\right)^\prime \rho\left(Z, \theta_0\right) \rho\left(Z, \theta_0\right)^{\prime} G\left(X, \theta_0\right)\right]_{p\times p}$.
\end{theorem}
For statistical inference, a consistent estimator of the asymptotic covariance matrix $\Sigma \coloneqq \Delta(\theta_0)^{-1}\Omega(\theta_0) \Delta(\theta_0)^{-1}$ is obtained via the plug-in sample analog $\widehat{\Sigma}_n = \widehat{\Delta}_n^{-1} \widehat{\Omega}_n \widehat{\Delta}_n^{-1}$, with components
\begin{equation}
\widehat{\Delta}_n = \frac{1}{n^2} \sum_{i=1}^n\sum_{j=1}^n \widehat{g}_i' K(X_i, X_j) \widehat{g}_j,\quad \widehat{\Omega}_n = \frac{1}{n} \sum_{i=1}^n \widehat{G}_n(X_i)^\prime\widehat{\rho}_i \widehat{\rho}_i^\prime \widehat{G}_n(X_i),
\end{equation}
where $\widehat{\rho}_i \coloneqq \rho(Z_i, \widehat{\theta}_n)$, $\widehat{g}_i \coloneqq g(Z_i, \widehat{\theta}_n)$, and $\widehat{G}_n(X_i) \coloneqq n^{-1} \sum_{j=1}^n K(X_i, X_j)\widehat{g}_j$.
\begin{proposition}
Under Assumptions (ref) and (ref), $\widehat{\Sigma}_n \xrightarrow{p} \Sigma$ as $n \to \infty$.
\end{proposition}
Proposition (ref) provides a basis for standard asymptotic inference on smooth functionals of $\theta_0$. Pointwise confidence intervals for individual components follow immediately from the asymptotic normality of $\widehat{\theta}_n$ and the consistency of $\widehat{\Sigma}_n$. For a full-row-rank matrix $R\in\mathbb{R}^{q\times p}$ and vector $r\in\mathbb{R}^q$, the linear restriction $\mathbb{H}_0:R\theta_0=r$ may be tested using the feasible Wald statistic $W_n^{L}=(R\widehat{\theta}_n-r)^\prime (R\widehat{\Sigma}_nR^\prime/n)^{-1}(R\widehat{\theta}_n-r)$. Under $\mathbb{H}_0$, we have $W_n^{L}\xrightarrow{d}\chi_q^2$. More generally, let $h:\mathbb{R}^p\to\mathbb{R}^q$ be continuously differentiable, with Jacobian $\nabla h(\theta)$ of full row rank at $\theta_0$. For the nonlinear restriction $\mathbb{H}_0:h(\theta_0)=0$, the feasible Wald statistic is $W_n^{NL}=h(\widehat{\theta}_n)^\prime (\nabla h(\widehat{\theta}_n)\widehat{\Sigma}_n\nabla h(\widehat{\theta}_n)^\prime/n)^{-1} h(\widehat{\theta}_n)$, and $W_n^{NL}\xrightarrow{d}\chi_q^2$ under $\mathbb{H}_0$. These standard Wald statistics concern restrictions on \(\theta_0\), which should be distinguished from the omnibus specification test of the conditional moment restrictions developed in Section (ref); see, e.g., lavergne2013smooth for SMD-based tests of parameter restrictions.
Finally, although the baseline KMD estimator is generally not semiparametrically efficient, it admits a two-step efficient refinement in the spirit of dominguez2004consistent. In particular, starting from the $\sqrt{n}$-consistent estimator $\widehat{\theta}_n$, one may take a single Newton--Raphson step toward the feasible efficient GMM objective constructed from nonparametric estimators of the relevant conditional moments; see newey1990efficient and robinson1991best. However, there is no free lunch in achieving semiparametric efficiency: this refinement requires additional nonparametric estimation, together with the attendant smoothing choices, and thus sacrifices some of the simplicity and robustness of the KMD procedure. We therefore defer the formal discussion to Section (ref) of the Online Supplementary Material.
\section{Specification Testing}
\subsection{Test Statistic}
Following the identification and estimation of the structural parameters discussed in the preceding section, we now turn to the specification test of the conditional moment restrictions. Specifically, we evaluate the validity of the model by testing the null hypothesis of correct specification against the omnibus alternative:
\begin{equation}
\mathbb{H}_0: (ref) holds \quad \text{v.s.}\quad\mathbb{H}_1: \mathbb{P}\left(\mathbb{E}\left[\rho(Z, \theta) \mid X\right]=0\right)<1 \text{ for all }\theta \in \Theta.
\end{equation}
To construct the test statistic, we utilize the minimized value of the KMD objective function. The test statistic $\widehat{T}_n$ is defined as the scaled $V$-statistic evaluated at $\widehat{\theta}_n$:
$$
\widehat{T}_n=n\widehat{Q}_n(\widehat{\theta}_n)=\frac{1}{n}\sum_{i=1}^n\sum_{j=1}^n\rho(Z_i,\widehat\theta_n)^{\prime} K(X_i,X_j)\rho(Z_j,\widehat\theta_n).
$$
In what follows, we derive the asymptotic distribution of $\widehat{T}_n$ under the null hypothesis $\mathbb{H}_0$, establish its consistency under the fixed alternative $\mathbb{H}_1$, and analyze its local power properties under the sequence of local alternatives $\mathbb{H}_{1n}$.
\subsection{Behavior Under the Null $\mathbb{H}_0$}
We first establish an asymptotic representation of $\widehat{T}_n$ under the null hypothesis, showing that the estimation effect is asymptotically equivalent to replacing the reproducing kernel with its projection.
\begin{lemma}
Suppose Assumptions (ref) and (ref) hold. Under $\mathbb{H}_0$, the test statistic $\widehat{T}_n$ has the following asymptotic representation:
$$
\widehat{T}_n = \frac{1}{n}\sum_{i=1}^n\sum_{j=1}^n\rho\left(Z_i,\theta_0\right)^{\prime} K^p(X_i,X_j)\rho\left(Z_j,\theta_0\right)+o_p(1),
$$
where $K^p(X_i,X_j)=K(X_i,X_j)-G(X_i,\theta_0)\Delta(\theta_0)^{-1} G(X_j,\theta_0)^{\prime}$.
\end{lemma}
\begin{remark}
The projected kernel $K^p$ induces a projection geometry within the vector-valued RKHS $\mathcal{H}_K$ onto the orthogonal complement of the score subspace $\mathcal{S} = \operatorname{span}\{\psi_l\}_{l=1}^p$, where $\psi_l(\cdot) = \mathbb{E}[K(\cdot, X) g_l(Z, \theta_0)]$. As detailed in the Online Supplementary Material, for any function $h(\cdot) = \sum_{i=1}^n K(\cdot, x_i)c_i$ with arbitrary coefficients $\{c_i\}_{i=1}^n \subset \mathbb{R}^q$, the quadratic form of $K^p$ satisfies the isometry:
\begin{equation}
\sum_{i=1}^n \sum_{j=1}^n c_i^{\prime} K^p(x_i, x_j) c_j = \left\| (I - P_{\mathcal{S}}) h \right\|_K^2 \ge 0,
\end{equation}
where $P_{\mathcal{S}}$ denotes the orthogonal projection operator onto $\mathcal{S}$. This identity establishes that $K^p$ is positive semidefinite.
In particular, the asymptotic representation of $\widehat{T}_n$ in Lemma (ref) corresponds to the specific case where the coefficients are the normalized residuals, $c_i = n^{-1/2}\rho(Z_i, \theta_0)$. By capturing the component of the moment conditions orthogonal to the score functions $g_l$, this asymptotic representation effectively eliminates the first-order estimation effect.
\end{remark}
Building on this representation, the following theorem derives the limiting distribution of $\widehat{T}_n$ under the null, which exhibits the standard spectral form for degenerate $V$-statistics.
\begin{theorem}
Suppose Assumptions (ref) and (ref) hold. Then, under $\mathbb{H}_0$:
$$
\widehat{T}_n \xrightarrow{d} \sum_{k=1}^{\infty} \lambda_kW_k^2,
$$
where $\left\{W_k\right\}_{k=1}^\infty$ are i.i.d. $\mathcal{N}(0,1)$ and $\left\{\lambda_k\right\}_{k=1}^{\infty}$ are eigenvalues of the operator $\mathcal{A}$ defined on the function space $L_2\left(\mathcal{Z}\times\mathcal{X},\mathbb{P}_{XZ}\right)$ as
\begin{equation}
\mathcal{A} \psi(s)=\int_{\mathcal{Z}\times\mathcal{X}} f_p\left(s, \tilde s, \theta_0\right) \psi\left(\tilde s\right) d \mathbb{P}_{XZ}\left(\tilde s\right),\quad s\in\mathcal{Z}\times\mathcal{X},\psi\in L_2\left(\mathcal{Z}\times\mathcal{X},\mathbb{P}_{XZ}\right),
\end{equation}
where $f_p\left(s, s^{\prime}, \theta_0\right)=\rho\left(z, \theta_0\right)^{\prime}K^p\left(x, x^{\prime}\right)\rho\left(z^{\prime}, \theta_0\right)$.
\end{theorem}
\begin{remark}
The spectral representation $\sum_{k=1}^{\infty} \lambda_k W_k^2$ derived in Theorem (ref) characterizes the test statistic $\widehat{T}_n$ as a \textit{Generalized Sargan Test} for a continuum of moment conditions. In the classical GMM framework with a fixed number of moments $m$ and parameters $p$, Hansen's $J$-statistic converges to a $\chi^2_{m-p}$ distribution, assessing the validity of the $m-p$ overidentifying restrictions. Our framework corresponds to the limiting case $m \to \infty$. Specifically, $\widehat{T}_n$ asymptotically isolates the component of the embedded moment conditions lying in the orthogonal complement of the score subspace, thereby assessing the structural validity of the continuum of overidentifying restrictions beyond the $p$ degrees of freedom consumed by parameter estimation. This interpretation parallels the unified framework of dominguez2015simple. While they address the estimation effect by projecting the empirical process, our framework relies on the projection kernel $K^p$.
\end{remark}
The test rejects $\mathbb{H}_0$ at significance level $\alpha$ if $\widehat{T}_n > c_{1-\alpha}$, where $c_{1-\alpha}$ is the $(1-\alpha)$-quantile of the limiting distribution defined in Theorem (ref). However, since the eigenvalues $\{\lambda_k\}_{k=1}^\infty$ depend on the kernel and the unknown data-generating process, this distribution is non-pivotal. In Section (ref), we introduce a computationally simple multiplier bootstrap to obtain feasible critical values.
\subsection{Behavior Under the Fixed Alternative $\mathbb{H}_1$}
We now establish the test's consistency against fixed alternatives by showing that $\widehat{T}_n$ diverges under global misspecification. This requires the following modified regularity conditions, under which a pseudo-true parameter $\theta_1$ is defined as the unique minimizer of the population objective function.
\begin{assumption}
Suppose Assumption (ref) (i), (ii), and (iv) hold, and (iii) is replaced by: The parameter space $\Theta\subset \mathbb{R}^p$ is compact. $Q\left(\theta\right)$ is uniquely minimized at a pseudo true value $\theta_1\in\Theta$ such that $Q(\theta_1) > 0$.
\end{assumption}
\begin{assumption}
Suppose Assumption (ref) holds with $\theta_0$ replaced by $\theta_1$. Furthermore, the generalized Hessian matrix under misspecification,
$$
\widetilde{\Delta}(\theta_1) = \mathbb{E}\left\{k(X_i,X_j)\left[g(Z_i,\theta_1)^{\prime}g(Z_j,\theta_1) + \sum_{l=1}^q\rho_l(Z_j,\theta_1)\nabla_{\theta\theta^{\prime}}^2\rho_l(Z_i,\theta_1)\right]\right\},
$$
is non-singular.
\end{assumption}
Analogous to Lemma (ref), the following result establishes the asymptotic representation of $\widehat{\theta}_n$ around the pseudo-true value $\theta_1$.
\begin{lemma}
Under $\mathbb{H}_1$, suppose Assumptions (ref) and (ref) hold, then
$$
\sqrt{n}\left(\widehat{\theta}_n-\theta_1\right)=-\widetilde\Delta\left(\theta_1\right)^{-1} \frac{1}{\sqrt{n}} \sum_{i=1}^n \left[G\left(X_i, \theta_1\right)^\prime \rho\left(Z_i, \theta_1\right)+g(Z_i,\theta_1)'M(X_i,\theta_1)\right]+o_p(1),
$$
where $G\left(X_i, \theta\right)$ is defined in Lemma (ref), and $M\left(X_i, \theta\right)=\mathbb{E}\left[K\left(X_i, X_j\right) \rho\left(Z_j, \theta_0\right)|X_i\right]_{q \times 1}$.
\end{lemma}
In contrast to the degenerate behavior under $\mathbb{H}_0$, the objective function $\widehat{Q}_n(\widehat{\theta}_n)$ under $\mathbb{H}_1$ behaves asymptotically as a non-degenerate $V$-statistic, yielding $\sqrt{n}-$asymptotic normality centered at the positive constant $Q(\theta_1)$.
\begin{theorem}
Suppose Assumptions (ref) and (ref) hold. Under the fixed alternative $\mathbb{H}_1$:
\begin{equation*}
\sqrt{n}\left(\widehat{Q}_n(\widehat{\theta}_n) - Q(\theta_1)\right) \xrightarrow{d} \mathcal{N}\left(0, \sigma^2(\theta_1)\right),
\end{equation*}
where $\sigma^2(\theta_1) = 4\operatorname{Var}\left[\rho(Z,\theta_1)^{\prime}M(X,\theta_1)\right]$.
\end{theorem}
Since $Q(\theta_1)>0$, Theorem (ref) implies that $\widehat{T}_n=n\widehat{Q}_n(\widehat{\theta}_n)\to\infty$. Hence, the test is consistent against any fixed alternative.
\subsection{Behavior Under the Local Alternative $\mathbb{H}_{1n}$}
We further examine the asymptotic power under a sequence of Pitman local alternatives converging at the parametric rate $n^{-1/2}$. Let $\mathbb{P}$ denote the null measure satisfying $\mathbb{E}_{\mathbb{P}}[\rho(Z, \theta_0) \mid X] = 0$ a.s., and $\mathbb{P}^{(n)}$ denote the sequence of alternative measures. We specify the form of these local perturbations as follows.
\begin{assumption}
Under $\mathbb{H}_{1n}$, the observations $(X,Z)$ follow a sequence of probability measures $\mathbb{P}^{(n)}$ such that $d\mathbb{P}^{(n)}/d\mathbb{P}(x,z)=1+n^{-1/2}h(x,z)$, where $\mathbb{E}_{\mathbb{P}}[h(X,Z)^4]<\infty$ and $\mathbb{E}_{\mathbb{P}}[h(X,Z)\mid X]=0$ a.s..
\end{assumption}
Applying the change of measure formula under Assumption (ref), we obtain the specific Pitman alternative:
\begin{equation}
\mathbb{H}_{1n}: \quad \mathbb{E}_{\mathbb{P}^{(n)}}\left[\rho(Z, \theta_0) \mid X\right] = \frac{\delta(X)}{\sqrt{n}}.
\end{equation}
Here, the drift function is defined as $\delta(X) \coloneqq \mathbb{E}_{\mathbb{P}}[\rho(Z, \theta_0)h(X, Z)|X]$, which is square-integrable due to the fourth finite moment conditions on $\rho$ and $h$. This local drift induces an asymptotic bias shift in the estimator, as formalized in the following linear representation.
\begin{comment}
\begin{equation*}
\mathbb{E}_{\mathbb{P}^{(n)}}[\rho(Z, \theta_0) \mid X]
= \frac{\mathbb{E}_{\mathbb{P}}\left[\rho(Z, \theta_0) \frac{d\mathbb{P}^{(n)}}{d\mathbb{P}} \bigg| X\right]}{\mathbb{E}_{\mathbb{P}}\left[\frac{d\mathbb{P}^{(n)}}{d\mathbb{P}} \bigg| X\right]}
= \frac{\mathbb{E}_{\mathbb{P}}\left[\rho(Z, \theta_0)\left(1 + \frac{1}{\sqrt{n}}h(X, Z)\right) \bigg| X\right]}{\mathbb{E}_{\mathbb{P}}\left[1 + \frac{1}{\sqrt{n}}h(X, Z) \bigg| X\right]}.
\end{equation*}
Define $\delta(X)=\mathbb{E}_{\mathbb{P}}\left[\rho(Z, \theta_0)h(X, Z) \mid X\right]$, we obtain the specific Pitman alternative:
\begin{equation}
\mathbb{H}_{1n}:\quad\mathbb{E}_{\mathbb{P}^{(n)}}[\rho(Z, \theta_0) \mid X] = \frac{\delta(X)}{\sqrt{n}}.
\end{equation}
Note that the drift function $\delta(X)$ is square-integrable with respect to the marginal probability measure of $X$ derived from $\mathbb{P}$. This follows from the fourth-moment conditions on $\rho$ and $s$ via Hölder's inequality.
\end{comment}
\begin{lemma}
Under the sequence of local alternatives $\mathbb{H}_{1n}$ satisfying Assumption (ref), and assuming Assumptions (ref)(ii)--(iv) and (ref) hold under the null measure $\mathbb{P}$, the estimator satisfies the asymptotic linear representation:
$$
\sqrt{n}\left(\widehat{\theta}_n-\theta_0\right) = -\Delta\left(\theta_0\right)^{-1} \frac{1}{\sqrt{n}} \sum_{i=1}^n G\left(X_i, \theta_0\right)^\prime \epsilon_i - \Delta\left(\theta_0\right)^{-1} B_\delta(\theta_0) + o_p(1),
$$
where $\epsilon_i \coloneqq \rho(Z_i, \theta_0) - n^{-1/2}\delta(X_i)$ represents the centered error term under $\mathbb{H}_{1n}$, and $B_\delta(\theta_0) = \mathbb{E}_{\mathbb{P}}\left[G(X_i, \theta_0)^\prime\delta(X_i)\right]$ denotes the asymptotic bias shift.
\end{lemma}
Building on this linear representation, the following theorem establishes the asymptotic distribution of $\widehat{T}_n$ under local alternatives.
\begin{theorem}
Suppose Assumptions for Lemma (ref) hold. Then, under $\mathbb{H}_{1n}$, the test statistic converges in distribution as:
$$
\widehat{T}_n \xrightarrow{d} \sum_{k=1}^{\infty}\lambda_k\left(W_k+c_k\right)^2,
$$
where $\{W_k\}_{k=1}^{\infty}$ are i.i.d. $\mathcal{N}(0,1)$, and $\{\lambda_k\}_{k=1}^{\infty}$ and $\{\varphi_k(\cdot)\}_{k=1}^{\infty}$ are the eigenvalues and orthonormal eigenfunctions, respectively, of the integral operator $\mathcal{A}$ defined in Theorem (ref). The non-centrality parameters are given by $c_k = \mathbb{E}_{\mathbb{P}}[\varphi_k(Z, X) h(Z, X)]$.
\end{theorem}
\begin{remark}
Theorem (ref) characterizes the limiting distribution as an infinite sum of weighted non-central $\chi^2_1$ variables. This result situates the proposed statistic within the established framework of ICM tests, exhibiting a spectral structure analogous to that derived by bierens1997asymptotic for local power analysis.
Its asymptotic local power is completely determined by the non-centrality parameters $\{c_k\}_{k=1}^\infty$. In particular, local power is trivial, in the sense that it converges to size $\alpha$, if and only if $c_k=0$ for all $k$ with positive eigenvalues. As shown in the proof of Theorem (ref), this is equivalent to
$$
\sum_{k=1}^\infty \lambda_k c_k^2
= \mathbb{E}_{\mathbb{P}}\!\left[\delta(X)^\prime K^p(X,\tilde X;\theta_0)\delta(\tilde X)\right]
=0.
$$
Since the projected kernel operator is positive semi-definite, this holds if and only if $\delta(X)$ lies in its null space $\mathbb{P}$-a.s., that is, $\mathbb{E}_{\mathbb{P}}[K^p(X,\tilde X;\theta_0)\delta(\tilde X)\mid X]=0$. By Remark (ref), this means that the local drift $\delta(X)$ belongs to the score subspace $\mathcal S$. Provided $K$ is ISPD, trivial power obtains if and only if $\delta(X)=\mathbb{E}_{\mathbb P}[g(Z,\theta_0)\mid X]\mathbf v$ $(\mathbb P\text{-a.s.})$ for some $\mathbf v\in\mathbb R^p$.
Intuitively, trivial power arises precisely when the local drift mimics the score function. In that case, it is asymptotically indistinguishable from a local shift in $\theta_0$ along the direction $\mathbf v$, and is therefore absorbed by the estimator $\widehat{\theta}_n$. This is a consequence of the preliminary estimation step rather than a deficiency of the test itself.
\end{remark}
\begin{comment}
$$
\sum_{k=1}^{\infty} \lambda_k (W_k + c_k)^2 \ \stackrel{d}{=} \ \underbrace{\sum_{k=1}^{\infty} \lambda_k W_k^2}_{Y^{(p)}} + \underbrace{2\sum_{k=1}^{\infty} \lambda_k c_k W_k}_{Z^{(p)}} + \underbrace{\sum_{k=1}^{\infty} \lambda_k c_k^2}_{C_\delta^p}.
$$
Here, $Y^{(p)}$ corresponds to the asymptotic null distribution, $C_\delta^p$ is the deterministic drift energy, and $Z^{(p)} \sim \mathcal{N}(0, 4\sum \lambda_k^2 c_k^2)$ is a zero-mean Gaussian component capturing the cross-interaction. Notably, the quadratic component $Y^{(p)}$ and the linear component $Z^{(p)}$ are asymptotically uncorrelated, but not necessarily independent.
\end{comment}
\begin{comment}
\begin{theorem}
Under the local alternative $\mathbb{H}_{1n}$, suppose Assumption (ref) holds. Further, assume that Assumptions (ref)(ii)--(iv) and (ref) are satisfied with respect to the null measure $\mathbb{P}$. Then,
$$
\widehat{T}_n(\widehat{\theta}_n) \xrightarrow{d} Y^{(p)} + Z^{(p)} + C_\delta^p,
$$
where $Y^{(p)} = \sum_{k=1}^{\infty} \lambda_kW_k^2$ is the asymptotic null distribution defined in Theorem (ref), $Z^{(p)}$ is a zero-mean Gaussian random variable, $Z^{(p)} \sim \mathcal{N}(0, \tau^2)$, and $C_\delta^p$ is a non-negative constant drift term. The variance $\tau^2$ and the drift $C_\delta^p$ are given by
\begin{align*}
&\tau^2 = 4 \mathrm{Var}_{\mathbb{P}}\left\{\rho(Z_i, \theta_0)^{\prime} \mathbb{E}\left[K^p(X_i, X_j; \theta_0) \delta(X_j) \Big| X_i\right]\right\}, \\
&C_\delta^p = \mathbb{E}_{\mathbb{P}}\left[\delta(X_i)^{\prime} K^p(X_i, X_j; \theta_0) \delta(X_j)\right].
\end{align*}
\end{theorem}
\end{comment}
\section{Multiplier Bootstrap}
The asymptotic null distribution of the test statistic $\widehat{T}_n$ is generally non-pivotal, as it depends on the underlying data-generating process via the operator $\mathcal{A}$. To circumvent this issue, we introduce an easy-to-implement multiplier bootstrap procedure to approximate the critical values. Motivated by the asymptotic representation established in Lemma (ref), the bootstrap procedure is implemented as follows:
\begin{enumerate}
• Generate a sequence of i.i.d. random variables $\left\{V_i\right\}_{i=1}^n$ with zero mean, unit variance and finite fourth moment, independent of the original samples $\left\{X_i,Z_i\right\}_{i=1}^n$; e.g., Rademacher random variable, standard normal random variable, or Bernoulli random variable with $ P(v=1-\kappa)=\kappa/\sqrt{5} $ and $ P(v=\kappa) = 1-\kappa/\sqrt{5} $, where $ \kappa = (\sqrt{5}+1)/2 $ mammen1993bootstrap.
• Compute the multiplier bootstrap statistic
\begin{equation*}
\widehat{T}^{*}_n = \frac{1}{n}\sum_{i=1}^n\sum_{j=1}^nV_iV_j\rho(Z_i,\widehat{\theta}_n)^{\prime}\widehat{K}_n^p(X_i,X_j,\widehat{\theta}_n)\rho(Z_j,\widehat{\theta}_n),
\end{equation*}
where $\widehat{K}^{p}_n(X_i,X_j,\widehat{\theta}_n) \coloneq K(X_i,X_j) - \widehat{G}_n(X_i)\widehat{\Delta}_n^{-1}\widehat{G}_n(X_j)^{\prime}$, with $\widehat{G}_n$ and $\widehat{\Delta}_n$ defined as in (ref).
• Repeat Steps 1 and 2 $B$ times to obtain a sequence of bootstrap statistics $\left\{\widehat{T}_{n,b}^*\right\}_{b=1}^B$.
• For a given significance level $\alpha \in (0,1)$, compute the $(1-\alpha)$-th quantile of $\left\{\widehat{T}_{n,b}^*\right\}_{b=1}^B$, denoted by $c_{n,\alpha}^*$. The null hypothesis is rejected if $\widehat{T}_n > c_{n,\alpha}^*$.
\end{enumerate}
To establish the asymptotic validity of this procedure, the following lemma first provides an asymptotic representation of the bootstrap statistic $\widehat{T}^{*}_n$, replacing the estimated quantities with their population counterparts.
\begin{lemma}
Let $\theta^*$ be the (pseudo) true value under different hypotheses, i.e., $\theta^*=\theta_0$ under $\mathbb{H}_0$ or $\mathbb{H}_{1n}$ and $\theta^*=\theta_1$ under $\mathbb{H}_1$. Under Assumptions (ref) and (ref) ( or Assumptions (ref) and (ref) for $\theta^*=\theta_1$), the multiplier bootstrap statistic has the following asymptotic representation:
\begin{equation*}
\widehat{T}^{*}_n =\frac{1}{n}\sum_{i=1}^n\sum_{j=1}^nV_iV_j\rho(Z_i,{\theta}^*)^{\prime}{K}^p(X_i,X_j,{\theta}^*)\rho(Z_j,{\theta}^*)+o_p(1).
\end{equation*}
\end{lemma}
Denote by “ $\xrightarrow{d,*}$ in probability” the weak convergence in probability under the bootstrap law, that is, conditional on the original sample $\left\{X_i,Z_i\right\}_{i=1}^n$ and let $\mathbb{P}^*$ be the conditional probability related to $\left\{V_i\right\}_{i=1}^n$ given $\left\{X_i,Z_i\right\}_{i=1}^n$. The next theorem establishes the asymptotic validity of the proposed multiplier bootstrap procedure.
\begin{theorem}
Suppose Assumptions for Lemma (ref) hold. Then,
\begin{itemize}
{0pt}
{0pt}
{0pt}
• Under $\mathbb{H}_0$ in (ref) or $\mathbb{H}_{1n}$ in (ref), we have
$$
\widehat{T}_n^*\xrightarrow{d,*}\sum_{k=1}^{\infty}\lambda_kW_k^2 \text{ in probability},
$$
where $\sum_{k=1}^{\infty}\lambda_kW_k^2$ is defined in Theorem (ref).
• Under $\mathbb{H}_1$ in (ref), for every $\varepsilon>0$, there exists a positive $A$ such that
$$
\mathbb{P}^*\left(\widehat{T}_n^*\geq A\right)\leq \varepsilon +o_{p}(1).
$$
\end{itemize}
\end{theorem}
On one hand, under $\mathbb{H}_0$, both $\widehat{T}_n^*$ and $\widehat{T}_n$ converge in distribution to $\sum_{k=1}^{\infty}\lambda_kW_k^2$, as defined in Theorem (ref). Hence, a test based on the bootstrapped critical value would yield an asymptotically correct level for $\widehat{T}_n$. On the other hand, under $\mathbb{H}_1$, by Theorem (ref), $\widehat{T}_n(\widehat{\theta}_n)$ diverges to infinity with probability approaching one, whereas $\widehat{T}_n^*$ is stochastically bounded conditional on the original sample. This implies that the test based on the bootstrapped critical value is consistent against all such fixed alternatives. Moreover, Theorem (ref) combined with Theorem (ref) also implies that our tests implemented through the multiplier bootstrap preserve the asymptotic local power properties of $\widehat{T}_n$ under $\mathbb{H}_{1n}$. Therefore, the multiplier bootstrap procedure is asymptotically valid.
\section{Simulation Study}
This section evaluates the finite-sample properties of the proposed KMD framework across three data-generating processes of increasing complexity: the Linear Regression Model, the Box--Cox Transformation Model, and Instrumental Variable Models. We assess estimation accuracy using bias and RMSE, with additional coverage rate results available in Section (ref) of the Online Supplementary Material. For the specification test, we examine empirical rejection rates to validate size and power; we also provide a benchmark comparison with dominguez2015simple in Section (ref) of the Online Supplementary Material.
Across all designs, we consider sample sizes $n \in \{100,200\}$ for estimation and extend the analysis to $n=400$ for specification testing. We vary the covariate dimension $p \in \{3,5,10\}$ to investigate performance in higher-dimensional settings. Except for the mixed IV design with discrete instruments, where the kernel is specified separately, KMD uses a Gaussian kernel with the median-heuristic bandwidth. All results are based on 1,000 Monte Carlo replications, with test critical values obtained from 999 bootstrap repetitions.
\subsection{The Linear Regression Model}
We begin with a linear regression benchmark under homoskedastic and heteroskedastic disturbances. We compare the KMD estimator with OLS, feasible weighted least squares, and infeasible optimal GMM, and study the corresponding KMD specification test under several basic alternatives. This benchmark mainly serves as a sanity check: the KMD estimator is essentially unbiased, with moderate efficiency loss relative to the parametric benchmarks, and the KMD test shows good size control and strong finite-sample power. To conserve space, the design and full results are reported in Section (ref) of the Online Supplementary Material. We now turn to nonlinear designs that more directly highlight the comparative advantages of the proposed method.
\subsection{The Box--Cox Transformation Model}
We next extend the simulation to the generalized Box--Cox transformation model. Our design draws upon the setups in shin2008semiparametric and dominguez2015simple, governed by the structural equation:
\begin{equation*}
y_i^{(\lambda_0)} = \alpha_0 + X_i^{\prime} \beta_0 + \epsilon_i, \quad i = 1, \dots, n.
\end{equation*}
To accommodate outcomes with non-positive support without error truncation, we employ the generalized transformation family of bickel1981analysis:
\begin{equation*}
y_i^{(\lambda)} =
\begin{cases}
\dfrac{|y_i|^\lambda \operatorname{sgn}(y_i) - 1}{\lambda} & \text{if } \lambda \neq 0, \\
\ln(y_i) & \text{if } \lambda = 0.
\end{cases}
\end{equation*}
We fix the intercept at $\alpha_0 = 0.5$ and evaluate performance across $\lambda_0 \in \{0, 0.5, 1\}$. The slope parameter is set to $\beta_0=(1,\dots,1,-1,\dots,-1)^{\prime}$, with the first $\lceil p/2\rceil$ components equal to $1$, and the regressors $X_i$ are generated from a multivariate normal distribution $\mathcal{N}(0,\Sigma_X)$ with Toeplitz covariance structure $\Sigma_{X,jk}=0.5^{|j-k|}$. We consider both homoskedastic and heteroskedastic disturbances, with $\epsilon_i \sim \mathcal{N}(0,1)$ in the former case and $\epsilon_i=\exp(0.25X_{i,1})\nu_i$, $\nu_i \sim \mathcal{N}(0,1)$, in the latter.
To assess estimation performance, we compare the KMD estimator with two feasible consistent alternatives: the nonlinear two-stage least squares (NL2S) estimator of amemiya1981comparison and the minimum distance estimator of shin2008semiparametric. For the log-linear case ($\lambda_0=0$), we also report the infeasible Oracle GMM (GMM$^*$), based on the optimal instrument $A^*(X_i)=\sigma^{-2}(X_i)\bigl(-1,\,-X_i^\prime,\,[(\alpha_0+X_i^\prime\beta_0)^2+\sigma^2(X_i)]/2\bigr)^\prime$.
\begin{table}[htbp]
\caption{Estimation Performance of Box--Cox Model Parameters ($\lambda=0$)}
\begin{tabular*}{\textwidth}{
@{\extracolsep{\fill}}
c c c
*{4}{S[table-format=-1.3]}
*{4}{S[table-format=1.3]}
@
}
\toprule
& & & \multicolumn{4}{c}{Bias} & \multicolumn{4}{c}{RMSE} \\
\cmidrule(lr){4-7} \cmidrule(lr){8-11}
{$n$} & {$p$} & {Param.} & {NL2S} & {Shin-MD} & {GMM$^*$} & {KMD} & {NL2S} & {Shin-MD} & {GMM$^*$} & {KMD} \\
\midrule
\multicolumn{11}{l}{\textit{Panel A: Homoskedastic Errors}} \\
\addlinespace[0.2em]
\multirow{9}{*}{$100$} & \multirow{3}{*}{$3$}
& $\alpha$ & 0.003 & -0.003 & 0.009 & -0.013 & 0.206 & 0.163 & 0.162 & 0.150 \\
& & $\beta$ & 0.018 & 0.007 & 0.005 & 0.008 & 0.145 & 0.160 & 0.133 & 0.134 \\
& & $\lambda$ & -0.004 & -0.007 & 0.002 & -0.013 & 0.100 & 0.072 & 0.071 & 0.066 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$5$}
& $\alpha$ & -0.007 & -0.006 & 0.005 & -0.007 & 0.189 & 0.187 & 0.141 & 0.143 \\
& & $\beta$ & 0.009 & 0.019 & 0.011 & 0.013 & 0.134 & 0.188 & 0.130 & 0.134 \\
& & $\lambda$ & -0.005 & -0.003 & -0.000 & -0.005 & 0.047 & 0.037 & 0.028 & 0.028 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$10$}
& $\alpha$ & -0.001 & -0.019 & 0.001 & 0.000 & 0.168 & 0.240 & 0.140 & 0.138 \\
& & $\beta$ & 0.015 & 0.015 & 0.017 & 0.015 & 0.139 & 0.206 & 0.138 & 0.139 \\
& & $\lambda$ & -0.000 & -0.003 & 0.000 & -0.000 & 0.014 & 0.012 & 0.009 & 0.009 \\
\addlinespace[0.5em]
\multirow{9}{*}{$200$} & \multirow{3}{*}{$3$}
& $\alpha$ & -0.004 & -0.002 & 0.002 & -0.010 & 0.141 & 0.138 & 0.107 & 0.115 \\
& & $\beta$ & 0.008 & 0.007 & 0.008 & 0.012 & 0.095 & 0.113 & 0.090 & 0.094 \\
& & $\lambda$ & -0.004 & -0.003 & 0.000 & -0.007 & 0.066 & 0.061 & 0.043 & 0.048 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$5$}
& $\alpha$ & -0.004 & -0.000 & 0.001 & -0.006 & 0.125 & 0.127 & 0.097 & 0.097 \\
& & $\beta$ & 0.005 & 0.010 & 0.004 & 0.004 & 0.092 & 0.131 & 0.090 & 0.093 \\
& & $\lambda$ & -0.002 & -0.000 & -0.000 & -0.002 & 0.030 & 0.027 & 0.019 & 0.019 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$10$}
& $\alpha$ & 0.001 & -0.016 & -0.004 & -0.000 & 0.114 & 0.238 & 0.088 & 0.090 \\
& & $\beta$ & 0.012 & 0.013 & 0.010 & 0.012 & 0.093 & 0.174 & 0.092 & 0.095 \\
& & $\lambda$ & 0.000 & -0.002 & -0.000 & -0.000 & 0.010 & 0.010 & 0.006 & 0.006 \\
\midrule
\multicolumn{11}{l}{\textit{Panel B: Heteroskedastic Errors}} \\
\addlinespace[0.2em]
\multirow{9}{*}{$100$} & \multirow{3}{*}{$3$}
& $\alpha$ & -0.020 & -0.061 & 0.004 & -0.051 & 0.210 & 0.166 & 0.152 & 0.160 \\
& & $\beta$ & 0.012 & 0.018 & 0.006 & 0.036 & 0.148 & 0.149 & 0.122 & 0.138 \\
& & $\lambda$ & -0.014 & -0.039 & 0.002 & -0.034 & 0.103 & 0.078 & 0.065 & 0.074 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$5$}
& $\alpha$ & -0.028 & -0.049 & -0.000 & -0.033 & 0.191 & 0.183 & 0.140 & 0.149 \\
& & $\beta$ & 0.013 & 0.017 & 0.014 & 0.016 & 0.141 & 0.175 & 0.126 & 0.141 \\
& & $\lambda$ & -0.008 & -0.016 & 0.001 & -0.011 & 0.048 & 0.038 & 0.027 & 0.031 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$10$}
& $\alpha$ & -0.024 & -0.060 & 0.009 & -0.013 & 0.181 & 0.250 & 0.135 & 0.143 \\
& & $\beta$ & 0.017 & 0.028 & 0.016 & 0.018 & 0.147 & 0.197 & 0.131 & 0.148 \\
& & $\lambda$ & -0.003 & -0.006 & 0.001 & -0.002 & 0.015 & 0.013 & 0.009 & 0.010 \\
\addlinespace[0.5em]
\multirow{9}{*}{$200$} & \multirow{3}{*}{$3$}
& $\alpha$ & -0.003 & -0.037 & -0.003 & -0.032 & 0.158 & 0.135 & 0.104 & 0.118 \\
& & $\beta$ & 0.011 & 0.010 & 0.006 & 0.018 & 0.100 & 0.104 & 0.082 & 0.096 \\
& & $\lambda$ & -0.003 & -0.023 & -0.001 & -0.020 & 0.075 & 0.063 & 0.042 & 0.053 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$5$}
& $\alpha$ & -0.014 & -0.030 & -0.000 & -0.022 & 0.131 & 0.124 & 0.092 & 0.100 \\
& & $\beta$ & 0.008 & 0.016 & 0.009 & 0.015 & 0.098 & 0.119 & 0.084 & 0.098 \\
& & $\lambda$ & -0.005 & -0.010 & -0.001 & -0.008 & 0.034 & 0.027 & 0.017 & 0.022 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$10$}
& $\alpha$ & -0.012 & -0.051 & 0.004 & -0.010 & 0.122 & 0.223 & 0.091 & 0.096 \\
& & $\beta$ & 0.008 & 0.018 & 0.007 & 0.009 & 0.099 & 0.159 & 0.088 & 0.100 \\
& & $\lambda$ & -0.001 & -0.005 & 0.000 & -0.001 & 0.011 & 0.010 & 0.006 & 0.006 \\
\bottomrule
\end{tabular*}
\begin{minipage}{\textwidth}
\textit{Note:} The bias for $\alpha$ and $\lambda$ is the raw bias, while the bias for $\beta$ is the $L_2$ norm of the bias vector.
\end{minipage}
\end{table}
\begin{table}[htbp]
\caption{Estimation Performance of Box--Cox Model Parameters ($\lambda=0.5$)}
{0pt}
\begin{tabular*}{\textwidth}{
@{\extracolsep{\fill}}
c c c
*{3}{S[table-format=-1.3]}
*{3}{S[table-format=1.3]}
@
}
\toprule
& & & \multicolumn{3}{c}{Bias} & \multicolumn{3}{c}{RMSE} \\
\cmidrule(lr){4-6} \cmidrule(lr){7-9}
{$n$} & {$p$} & {Param.} & {NL2S} & {Shin-MD} & {KMD} & {NL2S} & {Shin-MD} & {KMD} \\
\midrule
\multicolumn{9}{l}{\textit{Panel A: Homoskedastic Errors}} \\
\addlinespace[0.2em]
\multirow{9}{*}{$100$} & \multirow{3}{*}{$3$}
& $\alpha$ & 0.019 & 0.065 & 0.009 & 0.215 & 0.232 & 0.153 \\
& & $\beta$ & 0.026 & 0.176 & 0.010 & 0.133 & 0.475 & 0.130 \\
& & $\lambda$ & 0.014 & -0.018 & 0.005 & 0.109 & 0.210 & 0.066 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$5$}
& $\alpha$ & -0.007 & 0.052 & -0.002 & 0.224 & 0.273 & 0.162 \\
& & $\beta$ & 0.028 & 0.074 & 0.014 & 0.134 & 0.323 & 0.133 \\
& & $\lambda$ & 0.001 & -0.001 & 0.001 & 0.068 & 0.124 & 0.040 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$10$}
& $\alpha$ & -0.077 & -0.074 & -0.049 & 0.246 & 0.271 & 0.168 \\
& & $\beta$ & 0.022 & 0.038 & 0.021 & 0.140 & 0.214 & 0.139 \\
& & $\lambda$ & -0.015 & -0.016 & -0.011 & 0.044 & 0.040 & 0.026 \\
\addlinespace[0.5em]
\multirow{9}{*}{$200$} & \multirow{3}{*}{$3$}
& $\alpha$ & 0.008 & 0.066 & 0.005 & 0.154 & 0.219 & 0.118 \\
& & $\beta$ & 0.019 & 0.198 & 0.008 & 0.093 & 0.470 & 0.095 \\
& & $\lambda$ & 0.007 & -0.025 & 0.003 & 0.077 & 0.204 & 0.051 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$5$}
& $\alpha$ & -0.008 & 0.081 & -0.004 & 0.162 & 0.318 & 0.113 \\
& & $\beta$ & 0.013 & 0.173 & 0.006 & 0.093 & 0.422 & 0.095 \\
& & $\lambda$ & -0.000 & -0.019 & -0.000 & 0.050 & 0.161 & 0.029 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$10$}
& $\alpha$ & -0.056 & -0.063 & -0.037 & 0.197 & 0.250 & 0.129 \\
& & $\beta$ & 0.007 & 0.023 & 0.007 & 0.095 & 0.187 & 0.095 \\
& & $\lambda$ & -0.010 & -0.014 & -0.007 & 0.036 & 0.044 & 0.020 \\
\addlinespace[0.5em]
\multicolumn{9}{l}{\textit{Panel B: Heteroskedastic Errors}} \\
\addlinespace[0.2em]
\multirow{9}{*}{$100$} & \multirow{3}{*}{$3$}
& $\alpha$ & -0.001 & 0.025 & -0.021 & 0.225 & 0.227 & 0.156 \\
& & $\beta$ & 0.037 & 0.154 & 0.010 & 0.145 & 0.459 & 0.139 \\
& & $\lambda$ & 0.009 & -0.038 & -0.009 & 0.125 & 0.191 & 0.069 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$5$}
& $\alpha$ & -0.028 & 0.023 & -0.029 & 0.232 & 0.302 & 0.167 \\
& & $\beta$ & 0.020 & 0.099 & 0.005 & 0.143 & 0.360 & 0.142 \\
& & $\lambda$ & -0.006 & -0.024 & -0.010 & 0.072 & 0.135 & 0.044 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$10$}
& $\alpha$ & -0.098 & -0.090 & -0.066 & 0.281 & 0.373 & 0.187 \\
& & $\beta$ & 0.020 & 0.061 & 0.021 & 0.145 & 0.238 & 0.145 \\
& & $\lambda$ & -0.018 & -0.026 & -0.014 & 0.050 & 0.068 & 0.030 \\
\addlinespace[0.5em]
\multirow{9}{*}{$200$} & \multirow{3}{*}{$3$}
& $\alpha$ & 0.013 & 0.033 & -0.007 & 0.171 & 0.189 & 0.117 \\
& & $\beta$ & 0.027 & 0.186 & 0.010 & 0.099 & 0.463 & 0.097 \\
& & $\lambda$ & 0.011 & -0.046 & -0.005 & 0.093 & 0.205 & 0.053 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$5$}
& $\alpha$ & -0.006 & 0.057 & -0.008 & 0.188 & 0.313 & 0.117 \\
& & $\beta$ & 0.015 & 0.155 & 0.009 & 0.098 & 0.397 & 0.098 \\
& & $\lambda$ & 0.002 & -0.025 & -0.002 & 0.061 & 0.151 & 0.031 \\
\addlinespace[0.2em]
& \multirow{3}{*}{$10$}
& $\alpha$ & -0.084 & -0.097 & -0.049 & 0.221 & 0.235 & 0.133 \\
& & $\beta$ & 0.017 & 0.033 & 0.014 & 0.101 & 0.161 & 0.100 \\
& & $\lambda$ & -0.015 & -0.018 & -0.009 & 0.041 & 0.028 & 0.022 \\
\bottomrule
\end{tabular*}
\begin{minipage}{\textwidth}
\textit{Note:} The bias for $\alpha$ and $\lambda$ is the raw bias, while the bias for $\beta$ is the $L_2$ norm of the bias vector.
\end{minipage}
\end{table}
Tables (ref) and (ref) summarize the finite-sample estimation results for the Box--Cox transformation model under $\lambda_0=0$ and $\lambda_0=0.5$, respectively.\footnote{To conserve space, estimation results for $\lambda_0=1$ are relegated to Section (ref) of the Online Supplementary Material. These results exhibit patterns similar to the $\lambda_0=0.5$ case.} For the log-linear specification ($\lambda_0=0$, Table (ref)), the proposed KMD estimator demonstrates robust performance across all considered designs. In terms of consistency, KMD exhibits negligible bias, comparable to the infeasible Oracle GMM. Regarding efficiency, the KMD estimator consistently yields lower RMSEs than the feasible consistent alternatives, NL2S and Shin-MD, in most scenarios. Under homoskedasticity (Panel A), the KMD estimator performs remarkably well, exhibiting precision levels that closely approach the semiparametric efficiency bound. Furthermore, under heteroskedasticity (Panel B), where GMM$^*$ explicitly utilizes the true conditional variance, the unweighted KMD estimator remains highly competitive, delivering precision levels comparable to the theoretical efficiency bound.
The advantages of the KMD framework are most pronounced in the specification with $\lambda_0=0.5$ (Table (ref)). Across the considered designs, the KMD estimator consistently yields the lowest finite-sample bias and RMSE. Notably, KMD significantly outperforms the Shin-MD estimator, with the performance gap likely attributable to the covariate dimensionality. Whereas shin2008semiparametric relies on indicator-weighted cumulative moments (e.g., $\mathbb{I}(X_i \leq X_j)$) that are vulnerable to data sparsity in higher-dimensional spaces, the KMD approach employs a continuous kernel metric that adapts effectively to the multivariate geometry. Consequently, the KMD estimator maintains high precision without substantial deterioration as the dimension $p$ increases, offering a robust alternative for generalized regression models in nonlinear settings.
We next test the validity of the Box--Cox specification:
\begin{equation*}
\mathbb{H}_0: \mathbb{E}[Y^{(\lambda_0)} - (\alpha_0 + X^{\prime}\beta_0) \mid X] = 0
\quad \text{a.s. for some } \theta_0 \in \Theta.
\end{equation*}
Focusing on the log-linear case ($\lambda_0=0$), we define the residual as $\rho(Z, \theta) = \ln Y - (\alpha + X^{\prime}\beta)$.
To evaluate power, we generate data under the alternatives by introducing an additive drift $f(X_i)$ into the latent index: $Y_i = \exp(\alpha_0 + X_i^{\prime}\beta_0 + f(X_i) + \epsilon_i)$. We consider three representative forms of misspecification:
\begin{itemize}
{0pt}
{0pt}
{0pt}
• $\mathbb{H}_{1a}$ (Quadratic): $f(X_i)=\delta X_{i,1}^{2}$, representing smooth global nonlinearity.
• $\mathbb{H}_{1b}$ (Discontinuous): $f(X_i)=\delta \cdot \mathbb{I}(|X_{i,1}|\le 0.5)$, representing a localized structural break or bump that is often harder to detect.
• $\mathbb{H}_{1c}$ (Interaction): $f(X_i)=\delta X_{i,1}X_{i,2}$, capturing omitted interaction effects.
\end{itemize}
The perturbation magnitude is set to $\delta=0.5$ for $\mathbb{H}_{1a}$ and $\mathbb{H}_{1c}$, and to $\delta=1.5$ for $\mathbb{H}_{1b}$.
\begin{table}[htbp]
\caption{Empirical Rejection Rates for Box--Cox Specification Tests ($\alpha=0.05$)}
\newcolumntype{Y}{>{\arraybackslash}X}
\begin{tabularx}{\textwidth}{cc YYY YYY}
\toprule
& & \multicolumn{3}{c}{Homoskedastic Errors} & \multicolumn{3}{c}{Heteroskedastic Errors} \\
\cmidrule(lr){3-5} \cmidrule(lr){6-8}
$n$ & DGP & $p=3$ & $p=5$ & $p=10$ & $p=3$ & $p=5$ & $p=10$ \\
\midrule
\multirow{4}{*}{$100$}
& Size ($\mathbb{H}_0$) & $0.046$ & $0.058$ & $0.039$ & $0.060$ & $0.054$ & $0.047$ \\
& Power ($\mathbb{H}_{1a}$: Quad.) & $0.619$ & $0.651$ & $0.549$ & $0.674$ & $0.671$ & $0.489$ \\
& Power ($\mathbb{H}_{1b}$: Disc.) & $0.821$ & $0.545$ & $0.227$ & $0.745$ & $0.492$ & $0.231$ \\
& Power ($\mathbb{H}_{1c}$: Inter.) & $0.640$ & $0.587$ & $0.540$ & $0.677$ & $0.587$ & $0.495$ \\
\addlinespace[0.5em]
\multirow{4}{*}{$200$}
& Size ($\mathbb{H}_0$) & $0.056$ & $0.059$ & $0.035$ & $0.052$ & $0.046$ & $0.045$ \\
& Power ($\mathbb{H}_{1a}$: Quad.) & $0.953$ & $0.984$ & $0.935$ & $0.968$ & $0.979$ & $0.905$ \\
& Power ($\mathbb{H}_{1b}$: Disc.) & $0.986$ & $0.908$ & $0.574$ & $0.928$ & $0.798$ & $0.482$ \\
& Power ($\mathbb{H}_{1c}$: Inter.) & $0.930$ & $0.949$ & $0.931$ & $0.973$ & $0.940$ & $0.883$ \\
\addlinespace[0.5em]
\multirow{4}{*}{$400$}
& Size ($\mathbb{H}_0$) & $0.046$ & $0.047$ & $0.043$ & $0.064$ & $0.053$ & $0.050$ \\
& Power ($\mathbb{H}_{1a}$: Quad.) & $1.000$ & $1.000$ & $1.000$ & $1.000$ & $1.000$ & $1.000$ \\
& Power ($\mathbb{H}_{1b}$: Disc.) & $1.000$ & $0.999$ & $0.954$ & $0.996$ & $0.995$ & $0.928$ \\
& Power ($\mathbb{H}_{1c}$: Inter.) & $1.000$ & $1.000$ & $0.999$ & $1.000$ & $0.998$ & $0.999$ \\
\bottomrule
\end{tabularx}
\end{table}
Table (ref) summarizes the empirical rejection rates for the Box--Cox specification test. Under the null hypothesis, the test maintains accurate size control across all considered dimensions and error structures, with rejection rates close to the nominal level. Regarding power, we observe varying sensitivity to covariate dimensionality. For the smooth alternatives (Quadratic $\mathbb{H}_{1a}$ and Interaction $\mathbb{H}_{1c}$), the empirical power remains relatively robust as the dimension increases from $p=3$ to $p=10$. In contrast, the discontinuous alternative ($\mathbb{H}_{1b}$) exhibits a more substantial reduction in power in higher-dimensional settings, reflecting the challenge of detecting local irregularities in sparse spaces. However, these distortions vanish with larger sample sizes; rejection rates for all alternatives converge to 1 as $n$ increases to 400, confirming the test's consistency.
\subsection{Linear and Nonlinear Instrumental Variable Models}
Finally, we evaluate the finite-sample performance of the KMD estimator in the presence of endogeneity. We consider three DGPs with nonlinear reduced forms, including two continuous-instrument designs and one mixed-instrument design.
For Designs 1--2, we generate a $p$-dimensional continuous instrument vector
$Z_i=(Z_{i1},\ldots,Z_{ip})'$, where the components $Z_{ij}\sim \mathcal{U}[0,3]$ are mutually independent. For Design 3, the conditioning vector is $W_i=(C_i',D_i')'$, where $C_i=(C_{i1},\ldots,C_{ip_c})'$ contains independent $\mathcal{U}[0,3]$ continuous components and $D_i=(D_{i1},\ldots,D_{ip_d})'$ is binary, as is common in empirical applications. Let $\Lambda(a)=1/(1+\exp(-a))$. Conditional on $C_i$, we generate
$$
D_{im}\sim \operatorname{Bernoulli}(\pi_i),
\qquad
\pi_i=\Lambda\{0.75(C_{i1}-1.5)\},
\qquad m=1,\ldots,p_d,
$$
which induces dependence between the continuous and discrete components. We set $(p,p_c,p_d)=(3,2,1),(5,3,2),(10,5,5)$, with $p=p_c+p_d$. Endogeneity is induced through correlated latent errors:
$$
\begin{pmatrix}\eta_i \\ v_i\end{pmatrix}
\sim
\mathcal{N}\left(
\begin{pmatrix}0\\0\end{pmatrix},
\begin{pmatrix}1&0.8\\0.8&1\end{pmatrix}
\right),
\qquad
u_i=v_i,\qquad
\epsilon_i=\sigma_i\eta_i .
$$
We report both homoskedastic and heteroskedastic designs. In the homoskedastic case, $\sigma_i=1$. In the heteroskedastic case, $\sigma_i=\exp\{0.25(Z_{i1}-1.5)\}$ for Designs 1--2, while for Design 3, $\sigma_i=\exp\{0.25(C_{i1}-1.5)+0.15A_i\}$, with $A_i$ introduced below. We set $\theta_0=0.5$ and specify the three designs as follows:
\begin{itemize}
{0pt}
{0pt}
{0pt}
• \textit{Design 1: Linear Structural Equation.}
$$
Y_i = \theta_0 X_i + \epsilon_i,
\qquad
X_i = Z_{i1} + 0.5 Z_{i2}^{2} + u_i .
$$
• \textit{Design 2: Nonlinear Structural Equation.}
$$
Y_i = \exp(1+\theta_0 X_i) + \epsilon_i,
\qquad
X_i = 0.2 Z_{i1} + \cos(Z_{i2}Z_{i3}) + u_i .
$$
• \textit{Design 3: Nonlinear Structural Equation with Mixed Instruments.}
$$
Y_i = \exp(1+\theta_0 X_i) + \epsilon_i,
\qquad
X_i = 0.2C_{i1}+0.5\cos(C_{i1}C_{i2})+A_i+B_i\cos(C_{i1}C_{i2})+u_i,
$$
where $A_i=p_d^{-1/2}\sum_{m=1}^{p_d}(D_{im}-\pi_i)$ and $B_i=p_d^{-1/2}\sum_{m=1}^{p_d}(-1)^{m+1}(D_{im}-\pi_i)$.
\end{itemize}
For estimation, we benchmark KMD against alternative methods categorized by their handling of endogeneity and nonlinearity. A useful feature of our reproducing kernel formulation is that the kernel can be tailored to the support of the conditioning variables. We use a Gaussian kernel with the median-heuristic bandwidth in Designs 1--2. In the mixed-data Design 3, KMD uses a product kernel $K(W_i,W_j)=K_C(C_i,C_j)K_D(D_i,D_j)$, where $K_C$ is Gaussian and $K_D$ is a coordinatewise soft categorical kernel with $k(d,d')=\mathbf{1}\{d=d'\}+0.5\,\mathbf{1}\{d\neq d'\}$.
\begin{itemize}
{0pt}
{0pt}
{0pt}
• \textit{Plug-in 2SLS:} We consider an oracle plug-in estimator based on the true conditional mean $\mathbb{E}[X| Z]$. While valid in linear designs, plug-in two-stage procedures are generally inconsistent in nonlinear structural models because of the forbidden-regression problem wooldridge2010econometric.
• \textit{GMM Approaches:} We report \textit{Oracle GMM}, based on DGP-specific moment functions, and \textit{Sieve-GMM}, based on polynomial sieve moments with the efficient GMM weighting matrix.
• \textit{Smooth Minimum Distance (SMD):} Following lavergne2013smooth, we include a smooth minimum distance benchmark based on the off-diagonal $U$-statistic criterion. We report two bandwidth choices: a fixed bandwidth $h=1$, denoted SMD-h1, and a dimension-dependent smoothing bandwidth $ h_{\dim}=n^{-1/(p+4)}$, denoted SMD-hdim.
• \textit{DML-OS:} To handle higher-dimensional nuisance components, we also report a cross-fitted orthogonal-score benchmark, motivated by chernozhukov2018double, and implemented with random forests.\footnote{Implementation details for the DML-OS are provided in Section (ref) of the Online Supplementary Material.}
\end{itemize}
\begin{table}[htbp]
\caption{Estimation Performance of IV Estimators (Bias and RMSE)}
{0pt}
\begin{tabular*}{\textwidth}{
@{\extracolsep{\fill}}
c c l
*{7}{S[table-format=-2.3]}
@
}
\toprule
& & & \multicolumn{7}{c}{Estimators} \\
\cmidrule(lr){4-10}
{$n$} & {$p$} & {Metric$^{\dagger}$} & {2SLS$^*$} & {GMM$^*$} & {Sieve-GMM} & {SMD-h1} & {SMD-hdim} & {DML-OS} & {KMD} \\
\midrule
\multicolumn{10}{l}{\textit{Panel A: Linear Structural Equation}} \\
\addlinespace[0.2em]
\multirow{6}{*}{$100$}
& \multirow{2}{*}{$3$} & Bias & 0.092 & 0.053 & 0.266 & -0.081 & -0.077 & -0.105 & 0.120 \\
& & RMSE & 2.976 & 3.065 & 3.221 & 3.295 & 3.363 & 3.064 & 3.272 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$5$} & Bias & -0.064 & -0.086 & 0.443 & -0.318 & -0.377 & -0.282 & -0.079 \\
& & RMSE & 2.928 & 3.083 & 3.524 & 3.381 & 3.824 & 3.057 & 3.237 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$10$} & Bias & 0.105 & 0.148 & 1.144 & -0.124 & -0.470 & -0.077 & 0.198 \\
& & RMSE & 2.932 & 3.060 & 4.121 & 3.997 & 6.149 & 3.071 & 3.317 \\
\addlinespace[0.5em]
\multirow{6}{*}{$200$}
& \multirow{2}{*}{$3$} & Bias & 0.040 & 0.037 & 0.136 & -0.125 & -0.118 & -0.072 & -0.025 \\
& & RMSE & 2.106 & 2.125 & 2.188 & 2.295 & 2.318 & 2.123 & 2.266 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$5$} & Bias & 0.062 & 0.069 & 0.335 & -0.072 & -0.039 & -0.058 & 0.005 \\
& & RMSE & 2.167 & 2.194 & 2.366 & 2.389 & 2.524 & 2.230 & 2.389 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$10$} & Bias & -0.014 & -0.007 & 0.552 & -0.166 & -0.220 & -0.133 & -0.064 \\
& & RMSE & 2.102 & 2.126 & 2.458 & 2.747 & 3.966 & 2.147 & 2.408 \\
\midrule
\multicolumn{10}{l}{\textit{Panel B: Nonlinear (Exponential) Structural Equation}} \\
\addlinespace[0.2em]
\multirow{6}{*}{$100$}
& \multirow{2}{*}{$3$} & Bias & 9.537 & -0.091 & 0.503 & -0.585 & -0.562 & -0.530 & 0.001 \\
& & RMSE & 10.248 & 1.656 & 1.850 & 2.289 & 2.395 & 2.310 & 1.930 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$5$} & Bias & 9.382 & -0.130 & 0.892 & -0.743 & -1.174 & -0.612 & -0.022 \\
& & RMSE & 10.036 & 1.621 & 1.903 & 2.572 & 4.805 & 2.238 & 1.946 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$10$} & Bias & 9.766 & -0.005 & 1.766 & -1.986 & -19.627 & -0.819 & 0.083 \\
& & RMSE & 10.429 & 1.643 & 2.550 & 12.529 & 69.540 & 6.210 & 2.030 \\
\addlinespace[0.5em]
\multirow{6}{*}{$200$}
& \multirow{2}{*}{$3$} & Bias & 9.654 & -0.034 & 0.266 & -0.271 & -0.275 & -0.248 & -0.010 \\
& & RMSE & 9.990 & 1.153 & 1.297 & 1.486 & 1.482 & 1.445 & 1.395 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$5$} & Bias & 9.689 & 0.002 & 0.514 & -0.260 & -0.349 & -0.246 & 0.044 \\
& & RMSE & 10.024 & 1.078 & 1.268 & 1.557 & 1.918 & 1.389 & 1.383 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$10$} & Bias & 9.930 & 0.056 & 1.123 & -0.249 & -3.873 & -0.189 & 0.102 \\
& & RMSE & 10.239 & 1.068 & 1.649 & 1.923 & 23.944 & 1.385 & 1.348 \\
\midrule
\multicolumn{10}{l}{\textit{Panel C: Nonlinear (Exponential) Structural Equation with Mixed Covariates}} \\
\addlinespace[0.2em]
\multirow{6}{*}{$100$}
& \multirow{2}{*}{$3$} & Bias & 7.281 & 0.048 & 0.261 & -0.629 & -0.621 & -0.509 & 0.094 \\
& & RMSE & 8.064 & 1.516 & 1.713 & 2.293 & 2.312 & 2.069 & 1.852 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$5$} & Bias & 9.343 & 0.221 & 0.818 & -0.759 & -1.414 & -0.943 & 0.359 \\
& & RMSE & 10.405 & 1.738 & 1.968 & 2.835 & 8.190 & 6.091 & 1.926 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$10$} & Bias & 8.372 & 0.218 & 1.501 & -3.600 & -27.061 & -1.254 & 1.196 \\
& & RMSE & 10.172 & 1.773 & 2.372 & 18.851 & 85.758 & 8.697 & 1.936 \\
\addlinespace[0.5em]
\multirow{6}{*}{$200$}
& \multirow{2}{*}{$3$} & Bias & 7.389 & -0.019 & 0.133 & -0.288 & -0.302 & -0.256 & 0.046 \\
& & RMSE & 7.790 & 1.033 & 1.143 & 1.380 & 1.340 & 1.243 & 1.292 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$5$} & Bias & 9.469 & 0.133 & 0.410 & -0.349 & -0.478 & -0.319 & 0.209 \\
& & RMSE & 10.017 & 1.232 & 1.357 & 1.682 & 2.076 & 1.587 & 1.431 \\
\addlinespace[0.3em]
& \multirow{2}{*}{$10$} & Bias & 8.263 & 0.118 & 0.896 & -0.422 & -4.658 & -0.229 & 0.764 \\
& & RMSE & 9.122 & 1.141 & 1.467 & 2.273 & 28.903 & 1.576 & 1.367 \\
\bottomrule
\end{tabular*}
\begin{minipage}{\textwidth}
\textit{Note:} $^{\dagger}$ All values are multiplied by $100$ for readability.
\end{minipage}
\end{table}
Table (ref) reports the estimation results for the IV designs, with all values multiplied by 100 for readability. In the linear structural setting (Panel A), KMD performs comparably to the oracle IV benchmarks. Its bias is small across all sample sizes and dimensions, and its RMSE is of the same order as oracle 2SLS, oracle GMM, and DML-OS. As the dimension increases, KMD remains stable, whereas the sieve- and smoothing-based alternatives exhibit somewhat larger RMSEs across several designs.
Panels B and C show that the advantages of KMD become more visible in nonlinear IV settings. The oracle plug-in 2SLS estimator exhibits large biases in both panels, reflecting the forbidden-regression problem in nonlinear structural models, while oracle GMM serves as an infeasible benchmark based on correctly specified moment functions. Among the feasible alternatives, KMD exhibits low bias and a stable RMSE across sample sizes and dimensions. In Panel B, KMD is broadly comparable to the strongest feasible competitors at low dimensions and becomes more stable as $p$ increases. The SMD estimator with a fixed bandwidth performs reasonably well in several cases; however, the dimension-dependent smoothing bandwidth deteriorates markedly in higher-dimensional designs, illustrating the finite-sample curse of dimensionality faced by local-smoothing-based criteria. In Panel C, the mixed-data design further highlights this stability: KMD remains well behaved when the conditioning variables include both continuous and discrete instruments, whereas DML-OS and the SMD estimators exhibit larger finite-sample variability in several higher-dimensional mixed designs.
We conclude by evaluating the specification test in instrumental variable models. Let $W_i=Z_i$ in Designs 1--2 and $W_i=(C_i',D_i')'$ in Design 3. The null hypothesis asserts the validity of the conditional moment restriction:
\begin{equation*}
\mathbb{H}_0: \exists \theta_0 \in \Theta \text{ such that } \mathbb{E}[\rho(Y_i,X_i,\theta_0) \mid W_i] = 0 \text{ a.s.,}
\end{equation*}
where $\rho(Y_i,X_i,\theta)=Y_i-\theta X_i$ in Design I and $\rho(Y_i,X_i,\theta)=Y_i-\exp(1+\theta X_i)$ in Designs II--III. To evaluate power, we introduce perturbations of magnitude $\delta=0.1$ that target distinct misspecifications across the three designs. For the linear framework, we generate data under a quadratic alternative to introduce functional-form misspecification: $Y_i=\theta_0X_i+\delta X_i^2+\epsilon_i$. For the continuous nonlinear framework, we simulate an exclusion-type violation by allowing the instrument $Z_{i1}$ to enter the structural equation: $Y_i=\exp(1+\theta_0X_i+\delta Z_{i1})+\epsilon_i$. For the mixed-data nonlinear framework, we perturb the structural equation through the mixed component $A_i$: $Y_i=\exp(1+\theta_0X_i+\delta A_i)+\epsilon_i$.
\begin{table}[htbp]
\caption{Empirical Rejection Rates for IV Specification Tests ($\alpha=0.05$)}
\newcolumntype{Y}{>{\arraybackslash}X}
\begin{tabularx}{\textwidth}{cc YYY YYY}
\toprule
& & \multicolumn{3}{c}{Homoskedastic Errors}
& \multicolumn{3}{c}{Heteroskedastic Errors} \\
\cmidrule(lr){3-5} \cmidrule(lr){6-8}
$n$ & DGP
& $d=3$ & $d=5$ & $d=10$
& $d=3$ & $d=5$ & $d=10$ \\
\midrule
\multicolumn{8}{l}{\textit{Panel A: Linear structural model with continuous instruments}} \\
\multirow{2}{*}{$100$}
& Size ($\mathbb H_0$)
& $0.043$ & $0.040$ & $0.037$
& $0.040$ & $0.052$ & $0.031$ \\
& Power ($\mathbb H_1$)
& $0.653$ & $0.518$ & $0.313$
& $0.620$ & $0.456$ & $0.253$ \\
\addlinespace[0.3em]
\multirow{2}{*}{$200$}
& Size ($\mathbb H_0$)
& $0.052$ & $0.047$ & $0.040$
& $0.057$ & $0.040$ & $0.043$ \\
& Power ($\mathbb H_1$)
& $0.931$ & $0.885$ & $0.723$
& $0.901$ & $0.845$ & $0.686$ \\
\addlinespace[0.3em]
\multirow{2}{*}{$400$}
& Size ($\mathbb H_0$)
& $0.056$ & $0.039$ & $0.045$
& $0.048$ & $0.045$ & $0.044$ \\
& Power ($\mathbb H_1$)
& $0.999$ & $0.997$ & $0.974$
& $0.999$ & $0.997$ & $0.978$ \\
\addlinespace[0.7em]
\multicolumn{8}{l}{\textit{Panel B: Nonlinear structural model with continuous instruments}} \\
\multirow{2}{*}{$100$}
& Size ($\mathbb H_0$)
& $0.050$ & $0.046$ & $0.032$
& $0.048$ & $0.031$ & $0.033$ \\
& Power ($\mathbb H_1$)
& $0.817$ & $0.713$ & $0.508$
& $0.756$ & $0.639$ & $0.429$ \\
\addlinespace[0.3em]
\multirow{2}{*}{$200$}
& Size ($\mathbb H_0$)
& $0.056$ & $0.057$ & $0.040$
& $0.054$ & $0.047$ & $0.035$ \\
& Power ($\mathbb H_1$)
& $0.982$ & $0.960$ & $0.890$
& $0.979$ & $0.951$ & $0.858$ \\
\addlinespace[0.3em]
\multirow{2}{*}{$400$}
& Size ($\mathbb H_0$)
& $0.051$ & $0.058$ & $0.042$
& $0.048$ & $0.055$ & $0.037$ \\
& Power ($\mathbb H_1$)
& $0.998$ & $1.000$ & $0.997$
& $1.000$ & $1.000$ & $0.994$ \\
\addlinespace[0.7em]
\multicolumn{8}{l}{\textit{Panel C: Mixed nonlinear structural model with continuous and discrete instruments}} \\
\multirow{2}{*}{$100$}
& Size ($\mathbb H_0$)
& $0.053$ & $0.048$ & $0.032$
& $0.044$ & $0.041$ & $0.034$ \\
& Power ($\mathbb H_1$)
& $0.486$ & $0.559$ & $0.355$
& $0.409$ & $0.482$ & $0.296$ \\
\addlinespace[0.3em]
\multirow{2}{*}{$200$}
& Size ($\mathbb H_0$)
& $0.053$ & $0.048$ & $0.042$
& $0.055$ & $0.054$ & $0.032$ \\
& Power ($\mathbb H_1$)
& $0.856$ & $0.914$ & $0.761$
& $0.863$ & $0.883$ & $0.730$ \\
\addlinespace[0.3em]
\multirow{2}{*}{$400$}
& Size ($\mathbb H_0$)
& $0.046$ & $0.054$ & $0.042$
& $0.048$ & $0.052$ & $0.033$ \\
& Power ($\mathbb H_1$)
& $0.994$ & $0.999$ & $0.991$
& $0.999$ & $0.997$ & $0.992$ \\
\bottomrule
\end{tabularx}
\end{table}
\FloatBarrier
Table (ref) reports the empirical rejection rates for the IV specification tests under both homoskedastic and heteroskedastic errors. Under the null, the test controls size well across all three designs, with rejection rates generally close to the nominal 5% level; the mild undersizing observed in some high-dimensional cases is small and does not worsen under heteroskedasticity. Under the alternatives, the test shows strong power against the three types of misspecification considered: functional-form misspecification in Panel A, exclusion-type violations in Panel B, and mixed-data nonlinear misspecification in Panel C. Power is lower in smaller samples and tends to decrease with the dimension of the conditioning variables, but it increases rapidly with $n$. By $n=400$, rejection rates are close to one across all panels and error designs, indicating that the test remains effective in detecting violations of the IV conditional moment restrictions.
In summary, the simulation results underscore the KMD framework's universality and robustness across both estimation and testing tasks. In terms of estimation, the proposed estimator exhibits consistent stability across diverse settings, demonstrating distinct advantages particularly in scenarios characterized by nonlinearity and endogeneity. Regarding specification testing, the KMD statistic maintains accurate size control and satisfactory power, whereas the alternative test of dominguez2004consistent suffers significantly more from dimensionality; see Section (ref) of the Online Supplementary Material for a detailed comparison.
\section{Empirical Application}
To illustrate the practical relevance of the proposed KMD framework, we revisit the Engel curve application of blundell2007semi using their pre-processed 1995 UK Family Expenditure Survey (FES) sample. The data consist of 1,655 households, restricted to married couples with an employed head aged 25 to 55, and are treated as i.i.d. following blundell2007semi. The sample contains budget shares for seven commodity groups ($w_{il}$ for $l=1,\dots,7$: food, catering, alcohol, fuel, motor, fares, and leisure), together with log total expenditure ($\ln x_i$), log gross earnings ($\ln v_i$), and the number of children ($k_i$). The data are publicly available as the \texttt{Engel95} dataset in the R package \texttt{np}.
Following blundell2007semi, we consider the Engel curve specification
\begin{equation}
w_{il} = h_l(u_i) + \gamma_l k_i + \epsilon_{il}, \quad l = 1, \dots, 7,
\end{equation}
where $u_i=\ln x_i-\theta k_i$ denotes log effective total expenditure, $\theta$ is a common equivalence-scale parameter, and $\gamma_l$ captures commodity-specific demographic effects. Whereas blundell2007semi leave $h_l(\cdot)$ unrestricted and estimate it nonparametrically, our interest lies in testing economically motivated parametric restrictions on Engel curves. This is relevant because empirical demand analysis often relies on parsimonious functional forms, even though their adequacy is ultimately an empirical question. We therefore consider two leading specifications for $h_l(\cdot)$: the linear Working--Leser form working1943statistical,leser1963forms, $h_l(u_i)=\alpha_l+\beta_l u_i$, and the QUAIDS form banks1997quadratic, $h_l(u_i)=\alpha_l+\beta_l u_i+\lambda_l u_i^2$.\footnote{In this cross-sectional setting, households face common prices, so the price-index terms in the original QUAIDS specification are absorbed into the intercept $\alpha_l$.}
To address the endogeneity of total expenditure, we use the log of gross earnings as an instrument. Consequently, the null hypothesis of correct model specification is formulated as the following CMR:
\begin{equation}
\mathbb{H}_0: \mathbb{E}[{\epsilon}_{i} \mid {Z}_i] = \mathbb{E}[{w}_{i} - {h}(u_i; \theta) - {\Gamma} k_i \mid {Z}_i] = {0},
\end{equation}
where ${Z}_i$ collects the instrument ($\ln v_i$) and demographic controls ($k_i$). The vector ${w}_i$ contains budget shares, ${h}(\cdot)$ denotes the vector of shape functions, and ${\Gamma} = (\gamma_1, \dots, \gamma_7)^\top$ represents the vector of demographic coefficients. Based on this framework, we test two canonical specifications: the linear \textit{Working--Leser} form\footnote{Note that in the linear specification $h_l(u_i) = \alpha_l + \beta_l(\ln x_i - \theta k_i)$, the equivalence scale parameter $\theta$ is not separately identified as it is collinear with the direct demographic effect $\gamma_l$. In this case, $\theta$ is absorbed into the reduced-form coefficient on $k_i$.}, and the \textit{Quadratic Almost Ideal Demand System (QUAIDS)}.
We implement the KMD test using a Gaussian kernel, with the bandwidth selected via the median heuristic and critical values obtained from $B=2999$ multiplier bootstrap replications. The results decisively favor the quadratic specification: the linear Working--Leser form is rejected at the 5% level ($p=0.046$), indicating a misspecified functional form. In contrast, the quadratic QUAIDS specification fails to reject the null hypothesis ($p=0.703$), confirming that the inclusion of the quadratic term is sufficient to capture the structural nonlinearities in the demand system.
Given that the QUAIDS specification is not rejected by the data, we report the parameter estimates from this preferred model. We first identify the common equivalence scale parameter shared across the demand system, which we estimate as $\hat{\theta}_1 = 0.049$ (standard error $0.103$). While the positive sign is consistent with theoretical expectations regarding household composition costs, the estimate is not statistically significant in this sample. In light of this, we perform a robustness check by re-estimating the model and repeating the specification tests under the restriction $\theta=0$. The results, detailed in Section (ref) of the Online Supplementary Material, indicate that our primary findings regarding the functional form specification remain robust to the exclusion of the equivalence scale parameter. Table (ref) summarizes the remaining commodity-specific coefficients and their corresponding standard errors.
\begin{table}[htbp]
\caption{Parameter Estimates for the Quadratic (QUAIDS) Specification}
{0pt}
\begin{tabular*}{\textwidth}{
@{\extracolsep{\fill}}
l
c
c
c
c
}
\toprule
Commodity & $\alpha_l$ & $\beta_l$ & $\lambda_l$ & $\gamma_l$ \\
\midrule
Food & $-0.319$ $(0.640)$ & $\phantom{-}0.259$ $(0.232)$ & $-0.031$ $(0.021)$ & $\phantom{-}0.048$ $(0.009)$ \\
Catering & $-0.312$ $(0.437)$ & $\phantom{-}0.138$ $(0.160)$ & $-0.012$ $(0.015)$ & $-0.005$ $(0.003)$ \\
Alcohol & $-0.524$ $(0.453)$ & $\phantom{-}0.234$ $(0.164)$ & $-0.023$ $(0.015)$ & $-0.023$ $(0.004)$ \\
Fuel & $\phantom{-}1.055$ $(0.267)$ & $-0.330$ $(0.097)$ & $\phantom{-}0.027$ $(0.009)$ & $\phantom{-}0.009$ $(0.005)$ \\
Motor & $-2.041$ $(0.728)$ & $\phantom{-}0.842$ $(0.264)$ & $-0.080$ $(0.024)$ & $-0.020$ $(0.006)$ \\
Fares & $\phantom{-}0.822$ $(0.376)$ & $-0.310$ $(0.138)$ & $\phantom{-}0.030$ $(0.013)$ & $-0.007$ $(0.003)$ \\
Leisure & $\phantom{-}1.265$ $(1.030)$ & $-0.564$ $(0.380)$ & $\phantom{-}0.065$ $(0.035)$ & $-0.010$ $(0.015)$ \\
\bottomrule
\end{tabular*}
\begin{minipage}{\textwidth}
\textit{Notes:} The table reports the commodity-specific coefficients with standard errors in parentheses. The parameters correspond to the specification $w_{il} = \alpha_l + \beta_l u_i + \lambda_l u_i^2 + \gamma_l k_i + \epsilon_{il}$.
\end{minipage}
\end{table}
Inspection of Table (ref) indicates that the standard errors for the individual shape parameters $\beta_l$ and $\lambda_l$ are relatively large in some equations. This pattern likely reflects the structural multicollinearity between the linear regressor $u_i$ and its quadratic counterpart $u_i^2$. To assess the influence of total expenditure, we further conduct equation-specific Wald tests for the joint null hypothesis $\mathbb{H}_0: \beta_l = \lambda_l = 0$, utilizing the established asymptotic normality of the KMD estimator. The results, summarized in Table (ref), show a decisive rejection of the null hypothesis at the 1% significance level for six out of seven commodities. This finding confirms that, despite inflated standard errors for individual coefficients, the expenditure terms are jointly significant in characterizing the demand system. The sole exception is catering, where the expenditure elasticity does not appear statistically distinguishable from zero.
\begin{table}[htbp]
\caption{Wald Tests for Joint Significance of Expenditure Terms ($\mathbb{H}_0: \beta_l = \lambda_l = 0$)}
{0pt}
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}}l ccccccc}
\toprule
& Food & {Catering} & {Alcohol} & {Fuel} & {Motor} & {Fares} & {Leisure} \\
\midrule
$p$-value & $0.000^{***}$ & $0.435$ & $0.006^{**}$ & $0.000^{***}$ & $0.000^{***}$ & $0.008^{**}$ & $0.000^{***}$ \\
\bottomrule
\end{tabular*}
\begin{minipage}{\textwidth}
\textit{Note:} Significance levels: $^{*} p<0.05$, $^{**} p<0.01$, $^{***} p<0.001$.
\end{minipage}
\end{table}
\begin{figure*}[htbp]
\caption{Estimated QUAIDS Engel curves for seven commodity groups. Grey dots represent observed household shares, while solid blue-black lines indicate the fitted quadratic specification. Sub-labels (i) through (vii) refer to food, fuel, catering, alcohol, fares, motor, and leisure, respectively.}
\end{figure*}
\FloatBarrier
Finally, Figure (ref) illustrates the economic implications of the estimated model by plotting the predicted Engel curves over the support of the data. The recovered shapes are generally consistent with established demand theory. The curves for food and fuel exhibit a predominant downward trend consistent with Engel's Law, although the fuel curve suggests a mild reversal at the upper tail of the expenditure distribution. In contrast, the curves for motor and alcohol exhibit a distinct inverted-U-shaped trajectory, capturing the “rank-3” flexibility inherent in the QUAIDS specification. Notably, this non-monotonic pattern for alcohol is consistent with the empirical findings of banks1997quadratic. Furthermore, leisure exhibits a convex, accelerating upward trend, indicative of a luxury good. Collectively, these heterogeneous patterns highlight the limitations of the linear form and support adopting the quadratic specification.
\section{Conclusion}
This paper establishes a unified Kernel Minimum Distance (KMD) framework for estimating and testing models defined by conditional moment restrictions. By embedding conditional moments into a Reproducing Kernel Hilbert Space, we construct a tractable, closed-form $V$-statistic objective function. The proposed framework delivers a $\sqrt{n}$-consistent, asymptotically normal estimator and yields a naturally associated omnibus specification test based directly on the minimized objective value. We establish the asymptotic properties of our proposed tests under the null hypothesis, the fixed alternative, and a sequence of local alternatives converging to the null at the parametric rate $n^{-1/2}$. A distinguishing feature of our approach is that the estimation effect is inherently captured through a projected kernel structure, thereby obviating the need for auxiliary orthogonalization. Supported by a computationally efficient multiplier bootstrap, the proposed framework is practically implementable; extensive simulation results demonstrate robust performance in both estimation and testing across diverse settings.
\putbib
bibunit