EconBase
← Back to paper

Cressie Read Power Divergence for Moment-Based Estimation: Hyperparameter and Finite Sample Behavior

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

92,240 characters · 10 sections · 21 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Cressie–Read Power Divergence for Moment-Based Estimation: Hyperparameter and Finite-Sample Behavior

abstractWe study Cressie–Read power divergence (CRPD) estimation for moment-based models, focusing on finite-sample behavior. While generalized empirical likelihood estimators, dual to CRPD, are known to outperform generalized method of moments estimators in small to moderate samples, the power parameter is typically chosen arbitrarily by the researcher, serving mainly as an index. We interpret it as a hyperparameter that determines the loss function and governs the learning procedure, shaping the curvature of the objective and influencing finite-sample performance. Using second-order asymptotics, we show that it affects both the structural estimator and the associated Lagrange multipliers, governing robustness, bias, and sensitivity to sampling variation. Monte Carlo simulations illustrate how estimator performance varies with the choice of the power parameter and underlying distributional features, with implications for second-order bias and coverage distortion. An empirical illustration based on Owen (2001)’s classical example highlights the practical relevance of tuning the power parameter.\\ Keywords: Moment-based estimation; Cressie–Read power divergence; Generalized empirical likelihood; Finite-sample behavior; Hyperparameter.\\ JEL codes: C12, C13, C14.

Introduction

First-order asymptotic analysis has been widely adopted for its simplicity and analytical tractability. Its effectiveness hinges on the availability of sufficient information relative to the complexity of the estimation problem. In many empirical settings, this requirement is not well met. In some cases, the available data are inherently limited by the nature of the problem.\footnote{For example, when observations are costly, rare, or constrained by institutional or temporal factors—so that the sample size cannot be meaningfully increased} Even when the nominal sample size is moderately large, the effective information per parameter may still be limited, particularly in settings where the number of moments or parameters is large relative to the sample size. As a result, estimation problems may behave more like small-sample regimes, where first-order approximations provide a poor guide to finite-sample performance Huber1973.

These considerations motivate the use of Cressie--Read power divergence (CRPD) estimators CressieandRead1984 in moment-based estimation. CRPD estimators are closely connected to the generalized empirical likelihood (GEL) class through a well-known duality. The GEL literature has extensively documented favorable finite-sample properties relative to GMM (see, e.g., Owen1988; QinandLawless1994; HansenHeatonandYaron1996; KitamuraandStutzer1997; ImbensSpadyandJohnson1998; NeweyandSmith2004; Schennach2007), particularly in small to moderate samples. Through this duality, CRPD inherits these advantages, offering improved robustness by incorporating the finite-sample geometry of the moment conditions into the estimation procedure.

A key feature of the CRPD is the power parameter, which indexes a family of divergence functions. Different values of the power parameter label distinct members of the divergence family, with canonical choices corresponding to empirical likelihood, exponential tilting, or continuous updating. Typically, the power parameter is chosen arbitrarily by the researcher. Or, because the resulting estimators are first-order asymptotically equivalent under correct specification (see, e.g., NeweyandSmith2004), the choice of the power parameter is generally regarded as asymptotically irrelevant.

In this paper, we show that the power parameter $\gamma$ should not be selected arbitrarily, but instead can be interpreted as a learning hyperparameter in the sense of modern machine learning (ML) for two reasons. First, $\gamma$ determines the loss function itself and therefore governs the curvature of the learning objective. This elevates $\gamma$ beyond a conventional tuning constant to a parameter that shapes how the estimation problem is solved. We show that this choice is non-arbitrary: the CRPD criterion is well behaved as a function of $\gamma$ and admits a unique minimizer, implying that the curvature of the objective is endogenously determined rather than imposed by convention.

Specifically, within the CRPD framework, $\gamma$ controls the curvature of the penalty used to enforce moment conditions through empirical reweighting. Estimation begins with uniform weights assigned to all observations and then minimally perturbs these weights so that the sample moments satisfy the model’s restrictions. The power parameter governs how such perturbations are penalized.\footnote{Negative values of $\gamma$ induce aggressive down-weighting of observations associated with large moment residuals, enhancing robustness in the presence of heavy tails or outliers. Values of $\gamma$ near zero yield logarithmic sensitivity, providing a smooth benchmark, while positive values generate gentler reweighting schemes that retain greater influence from all observations but may sacrifice robustness.}

Second, we show that $\gamma$ does not enter the first-order asymptotic expansion, but plays a substantive role in the second-order behavior of both the structural estimator and the associated Lagrange multipliers. First-order asymptotics govern identification and consistency, whereas second-order terms determine finite-sample properties, including bias, robustness, and sensitivity to sampling variation (see, e.g., Rothenberg1984; NeweyandSmith2004; HuberandRonchetti2009). Because $\gamma$ does not affect first-order identification, it cannot be interpreted as a structural parameter. Its role is instead to govern how learning proceeds in finite samples through its effect on higher-order behavior.

From this perspective, $\gamma$ plays a role analogous to a hyperparameter: it does not alter the population target, but influences how the estimator responds to finite-sample features of the data. This interpretation connects the CRPD framework to ML ideas—such as tuning and model selection—while remaining within a standard moment-based econometric setting. In particular, $\gamma$ can be viewed as a continuous parameter that governs the curvature of the objective and thereby the robustness and sensitivity of the estimator. Different values of $\gamma$ lead to systematically different finite-sample behavior within a class of first-order equivalent estimators.

Bringing this perspective to the CRPD framework yields several practical advantages. Rather than fixing $\gamma$ a priori, it can be selected in a data-driven manner, for example using cross-validation or out-of-sample criteria. This allows the estimator to adapt to the empirical environment and mitigates distortions arising from unreliable first-order approximations. In this sense, $\gamma$ serves as a continuous control over finite-sample learning behavior, enabling the researcher to tailor robustness and sensitivity to the features of the data.

Our interpretation of $\gamma$ as a learning hyperparameter admits a useful analogy to optimization procedures. In particular, it parallels the role of the step size (or learning rate) in stochastic gradient descent (SGD), where different choices can lead to the same asymptotic limit while producing different finite-sample behavior (see, e.g., ToulisandAiroldi2017). Although the step size does not affect the asymptotic target under standard conditions, it plays an important role in determining convergence speed, stability, and sensitivity to noise. Similarly, in the CRPD framework, the choice of $\gamma$ leaves first-order identification and consistency unchanged, but governs second-order properties such as robustness, bias, and sensitivity to sampling variation. This analogy highlights a broader principle: parameters that are asymptotically irrelevant may nonetheless be important for finite-sample behavior.

More broadly, interpreting the CRPD power parameter as governing the curvature of the objective connects naturally to the literature on loss design in statistical estimation. Classical asymptotic theory organizes estimators into the $M$-, $R$-, and $L$-estimation frameworks Serfling1980. While empirical likelihood and the broader CRPD/GEL family are not directly formulated within this taxonomy—being defined through constrained divergence minimization over probability weights—profiling out the weights and associated multipliers yields a representation analogous to an $M$-estimator. This connection allows their asymptotic behavior to be analyzed using standard $M$-estimation tools.

Within this perspective, our framework is most closely related to the robust $M$-estimation literature Huber1964,HuberandRonchetti2009, where robustness is achieved through the design of loss functions whose curvature governs the influence of extreme observations. In that setting, tuning constants regulate finite-sample robustness and efficiency without corresponding to structural parameters of the data-generating process. Our framework extends this idea to moment-based estimation: rather than modifying residual-level losses, the CRPD power parameter shapes the curvature of a divergence-based objective that enforces moment conditions through empirical reweighting. As in robust $M$-estimation, $\gamma$ is not identified by population moments but plays a central role in determining finite-sample behavior.

Our analysis is closely related to NeweyandSmith2004, who develop higher-order asymptotic theory to compare the finite-sample performance of GMM and GEL estimators under fixed divergence choices. Their analysis treats the divergence as given and studies how different members of the GEL class affect higher-order behavior of the structural parameter. Accordingly, the power parameter enters only insofar as it indexes different estimators within the GEL family. While we adopt the same higher-order expansion framework, our focus is different: rather than comparing estimators conditional on a fixed divergence, we study the role of the divergence parameter itself. We show that $\gamma$ affects finite-sample behavior without altering population identification. In particular, while first-order results do not depend on $\gamma$, it enters the second-order behavior of both the estimator and the associated Lagrange multipliers, thereby governing robustness and constraint enforcement in finite samples.

Our work is also related to Schennach2007, who develops the exponentially tilted empirical likelihood (ETEL) framework within a divergence-based class indexed by a power parameter. While our setting differs, we share the perspective of treating the power parameter as a continuum-valued index linking different estimators. In contrast to Schennach2007, who focuses on misspecified moment conditions, we maintain correct specification and show that the power parameter matters through second-order effects, shaping robustness and sensitivity even when estimators are first-order asymptotically equivalent.

The rest of the paper is organized as follows. In Section 2, we illustrate how the power parameter shapes the loss geometry. In Section 3, we develop the moment-based CRPD estimator and show how the structural parameter and the Lagrange multipliers depend on the power parameter at the second order, affecting finite-sample properties. In Section 4, we present Monte Carlo simulations. In Section 5, we discuss the implications of our proposed framework. In Section 6, we present an empirical application. In Section 7, we conclude.

The power parameter and the loss geometry

Let $\boldsymbol{\pi} = (\pi_1, \ldots, \pi_n)$ denote the model-implied probability vector, where $\pi_i > 0$ for all $i = 1, \ldots, n$ and $\sum_{i=1}^n \pi_i = 1$. Let $\mathbf{q} = (q_1, \ldots, q_n)$ denote a reference probability vector against which $\boldsymbol{\pi}$ is compared. For $\gamma\in\mathbb{R}\setminus\{-1,0\}$, the CRPD criterion CressieandRead1984 is defined as \[ \mathscr I_\gamma(\boldsymbol{\pi},\mathbf{q}) = \frac{1}{\gamma(\gamma+1)} \sum_{i=1}^n \pi_i\left[\left(\frac{\pi_i}{q_i}\right)^\gamma-1\right]. \]

Observe that $\gamma$ governs the curvature of the divergence penalty and therefore shapes how the estimator trades off fidelity to the reference measure against satisfaction of the moment conditions. For large positive values of $\gamma$, the divergence grows rapidly for small deviations, tightly constraining the optimized weights $\pi_i$ around the reference distribution unless the sample moments strongly compel otherwise. As $\gamma$ approaches zero, the divergence converges to the Kullback--Leibler form, allowing moderate deviations at relatively low cost. When $\gamma$ is negative, the loss function becomes flatter around the center of the distribution while penalizing extreme distortions more sharply, effectively down-weighting observations that generate large moment residuals.

These changes in curvature have direct implications for estimation. In particular, the choice of $\gamma$ determines how the estimator responds to empirical features such as heavy tails, skewness, or localized moment violations. Smaller or negative values of $\gamma$ induce robustness by limiting the influence of extreme observations, while larger values enforce tighter adherence to the reference distribution. Consequently, different values of $\gamma$ correspond to different second-order sensitivities of the estimator to local misspecification.

In particular, as the reference distribution, we adopt the uniform distribution over the observed data points. That is, we take $\mathbf{q}:=\mathbf{i}/n$ because by the Glivenko--Cantelli theorem, the empirical distribution—which assigns equal mass to each realized sample point—converges uniformly to the true data-generating distribution. Hence, uniform weighting over the sample points (not over the population support) provides a consistent estimator of the true distribution.

The CRPD between $\boldsymbol{\pi}$ and $\mathbf{i}/n$ is defined as

equation[equation omitted — 206 chars of source]

with continuous extensions at $\gamma=0$ and $\gamma=-1$. Up to an additive constant, $\mathscr{I}_{\gamma}(\boldsymbol{\pi},\mathbf{i}/n)$ admits the equivalent representations

equation*[equation* omitted — 322 chars of source]

where the case $\gamma=0$ corresponds to the exponential tilting (ET) divergence, and the case $\gamma=-1$ corresponds to the empirical likelihood (EL) divergence.

Our objective is to formalize the interpretation of $\gamma$ as a parameter governing the geometry of the learning objective, rather than as a fixed modeling choice imposed ex ante. Establishing existence and uniqueness of $\gamma$ within the CRPD family is therefore essential to ensure that the learning objective is well defined and non-arbitrary.

A defining feature of the CRPD framework is that different values of $\gamma$ correspond to distinct loss functions, each characterized by a different curvature and sensitivity to deviations between the empirical distribution and the reference distribution. In this sense, $\gamma$ indexes a family of admissible learning objectives, with each value implying a different robustness profile, rather than acting as a tuning constant chosen independently of the estimation problem.

Observe that $\gamma$ determines the form of the implied loss function. For instance, when the underlying measurements are approximately Gaussian, the squared-error ($\ell_2$) loss is optimal in the sense of maximum likelihood. Within the CRPD family, this case corresponds to $\gamma = 1$, since

equation*[equation* omitted — 347 chars of source]

which coincides with the squared Euclidean distance between the empirical weights and the uniform distribution. Other values of $\gamma$ induce alternative loss geometries, emphasizing different aspects of deviation and robustness to misspecification as summarized in Table (ref).

table[table omitted — 1,047 chars of source]

Interpreting the CRPD objective as a loss function clarifies why $\gamma$ should be regarded as a legitimate parameter of the learning problem rather than a fixed modeling convention. While conventional applications treat $\gamma$ as exogenously specified, such a practice implicitly fixes the curvature of the learning objective independently of the estimation problem. In contrast, our framework allows the curvature parameter $\gamma$ to be treated as an object of analysis, whose role can be formally characterized.

Moment-based CRPD estimation

We consider CRPD estimation under moment conditions. Let $\{Z_i\}_{i=1}^n$ be an i.i.d.\ sample drawn from a probability distribution $P$ on $(\mathcal X,\mathcal A)$, where $Z_i=(Y_i,X_i)$ collects the observed outcome and covariates for unit $i$. Let $Z=(Y,X)$ denote a generic draw from $P$. Let $g:\mathcal X\times\Theta\to\mathbb R^q$ be a measurable moment function. We assume that the true parameter $\theta_0\in\Theta$ satisfies the population moment condition $ E\bigl[g(Z,\theta_0)\bigr]=0. $ Let $g_i(\theta):=g(Z_i,\theta)\in\mathbb{R}^q$ and define $ \bar g(\theta):=\frac{1}{n}\sum_{i=1}^n g_i(\theta), \qquad \hat\Omega(\theta):=\frac{1}{n}\sum_{i=1}^n g_i(\theta)g_i(\theta)^{\prime}. $ Let $G_i(\theta):=\partial g_i(\theta)/\partial\theta^{\prime}$, and define $ \bar G(\theta):=\frac{1}{n}\sum_{i=1}^n G_i(\theta), \; G_0:=E[G(Z,\theta_0)], \; \Omega_0:=E[g(Z,\theta_0)g(Z,\theta_0)^{\prime}]. $ Let $H_i(\theta)$ denote the collection of second derivatives of $g_i(\theta)$ w.r.t.\ $\theta$ (i.e.\ a $q\times p\times p$ tensor), and assume $E[\sup_{\theta\in\Theta}\|H(Z,\theta)\|]<\infty$.

Moment-based CRPD estimator

Fix $\gamma\in\Gamma$ with $\gamma\neq 0$ and $\gamma\neq -1$. Define the CRPD estimator:

equation*[equation* omitted — 268 chars of source]

The Lagrange function is

equation[equation omitted — 467 chars of source]

The first derivative of (ref) with respect to $\pi_{i}$ is

equation[equation omitted — 318 chars of source]

The first-order condition (FOC) is obtained by setting (ref) equal to zero. Observe

equation*[equation* omitted — 135 chars of source]

Rearranging terms,

equation*[equation* omitted — 121 chars of source]

Multiplying $\gamma$,

equation*[equation* omitted — 109 chars of source]

Thus,

equation[equation omitted — 143 chars of source]

Before proceeding to the asymptotic analysis, it is useful to characterize the population solution of the CRPD problem under correct specification. The following lemma establishes the benchmark values of the parameters and multipliers around which the subsequent expansions are developed.

lemma(Population solution under correct specification) Under correct specification, the solution to the CRPD problem lies in a neighborhood of \[ \theta=\theta_{0}, \qquad \lambda=0, \qquad \pi_{i}=1/n, \qquad \delta_{0}=-\frac{1}{\gamma+1}. \]

Proof. See Appendix (ref).

Lemma (ref) identifies the population reference point for the CRPD optimization problem. Under correct specification, the population solution corresponds to uniform weights $\pi_i = 1/n$ with $\lambda = 0$ and $\delta_0 = -1/(\gamma + 1)$. This characterization provides the baseline around which the sample estimators and Lagrange multipliers are expanded in the subsequent analysis.

Asymptotic analysis and finite sample properties

For each $\theta$, the pair $(\lambda^{\prime},\delta)$ is determined by the two constraints

equation[equation omitted — 147 chars of source]

Define the stacked system \[ \Psi_n(\theta,\lambda,\delta) :=

pmatrix[pmatrix omitted — 84 chars of source]

=

pmatrix[pmatrix omitted — 138 chars of source]

, \] where $w_i(\theta,\lambda,\delta):= \Big(\frac{1}{\gamma+1}-\gamma\delta-\gamma\lambda^{\prime}g_i(\theta)\Big)^{1/\gamma}$, so that (ref) is equivalent to $\Psi_n(\theta,\lambda,\delta)=0$.

For each $\theta$, let $(\hat\lambda(\theta),\hat\delta(\theta))$ denote a solution of $\Psi_n(\theta,\lambda,\delta)=0$ (existence/uniqueness shown below) and define $\hat{\boldsymbol{\pi}}(\theta)=\boldsymbol{\pi}(\theta,\hat\lambda(\theta),\hat\delta(\theta))$. Define the profiled objective \[ L_n(\theta):=\mathscr{I}_{\gamma}\big(\hat{\boldsymbol{\pi}}(\theta),\textbf{i}/n\big), \qquad \hat\theta\in\arg\min_{\theta\in\Theta}L_n(\theta). \]

We assume the followings.

assumption(Compactness) $\Theta$ is compact and $\theta_0\in\mathrm{int}(\Theta)$.
assumption(Smoothness and envelope) $g(z,\theta)$ is continuously differentiable in $\theta$ for $P$-a.e.\ $z$, with Jacobian $G(z,\theta):=\partial g(z,\theta)/\partial\theta^{\prime}$. Moreover, \[ E\Big[\sup_{\theta\in\Theta}\|g(Z,\theta)\|^2\Big]<\infty, \qquad E\Big[\sup_{\theta\in\Theta}\|G(Z,\theta)\|\Big]<\infty. \]
assumption(Identification) $E[g(Z,\theta_0)]=0$ and $E[g(Z,\theta)]\neq 0$ for $\theta\neq\theta_0$. Let \[ G_0:=E[G(Z,\theta_0)],\qquad \Omega_0:=E[g(Z,\theta_0)g(Z,\theta_0)^{\prime}], \] and assume $G_0$ has full column rank and $\Omega_0$ is positive definite.
assumption(Interior feasibility / positivity) There exists a neighborhood $\mathcal{N}$ of $\theta_0$ and a constant $\kappa>0$ such that, with probability approaching one, for each $\theta\in\mathcal{N}$ there exist $(\lambda,\delta)$ with $\|\lambda\|+\|\delta-\delta_0\|\le \kappa$ satisfying \[ \frac{1}{\gamma+1}-\gamma\delta-\gamma\lambda^{\prime}g_i(\theta)\ge \kappa \quad\text{for all } i, \] and the constraints (ref).

Assumption (ref) ensures that the parameter space is bounded and closed, which facilitates uniform convergence arguments and guarantees the existence of minimizers of the profiled criterion. The interior condition $\theta_0\in\mathrm{int}(\Theta)$ rules out boundary issues and permits local expansions around the true parameter. Assumption (ref) provides the regularity required for both uniform laws of large numbers and Taylor expansions. Continuous differentiability of $g(z,\theta)$ in $\theta$ ensures that the multiplier system and the profiled objective are smooth functions of $\theta$. The envelope conditions guarantee integrable bounds that allow uniform convergence of sample moments and control stochastic equicontinuity. Assumption (ref) ensures that the population moment condition $E[g(Z,\theta)]=0$ uniquely characterizes $\theta_0$. The full column rank of $G_0$ guarantees local identification and invertibility of the Jacobian in the multiplier system, while the positive definiteness of $\Omega_0$ ensures nondegeneracy of the moment covariance matrix. Together, these conditions imply both consistency and asymptotic normality of the estimator. Assumption (ref) is a technical regularity condition ensuring that the implied weights are well-defined and that the multiplier system remains smooth in a neighborhood of $\theta_0$. The uniform positivity bound prevents the implied probabilities from approaching the boundary of the simplex. This guarantees interior solutions, avoids singularities in the CRPD power transformation, and ensures that the mapping $(\theta,\lambda,\delta)\mapsto w_i(\theta,\lambda,\delta)$ is continuously differentiable on the relevant region.

theorem(Uniform convergence of the profiled objective) Fix the hyperparameter $\gamma$. Under Assumptions 1--4, there exists a deterministic function $L(\theta)$ such that \[ \sup_{\theta\in\Theta}\bigl|L_n(\theta)-L(\theta)\bigr| \xrightarrow{p}0, \] and $L(\theta)$ is uniquely minimized at $\theta_0$.

Proof. See Appendix (ref).

Theorem (ref) shows that the profiled sample objective $L_n(\theta)$ converges uniformly to a deterministic population criterion $L(\theta)$. This result ensures that the finite-sample objective approximates a well-defined population problem. Because $L(\theta)$ is uniquely minimized at $\theta_0$, the minimizer of the sample objective will concentrate near the true parameter as the sample size increases. This property forms the basis for establishing consistency of the CRPD estimator in the subsequent results.

To analyze the profiled criterion $L_n(\theta)$, we must first ensure that the Lagrange multipliers solving the constraint system are well defined. In particular, we require that for $\theta$ in a neighborhood of the true value $\theta_0$, the system of first-order conditions admits a unique solution for the multipliers. The following lemma establishes the local existence and uniqueness of $(\hat\lambda(\theta),\hat\delta(\theta))$. $(\hat\lambda(\theta),\hat\delta(\theta))$.

lemma[Local existence and uniqueness of multipliers] Under Assumptions (ref),(ref),(ref),(ref), there exists a neighborhood $\mathcal{N}$ of $\theta_0$ such that, with probability approaching one, for every $\theta\in\mathcal{N}$ the system $\Psi_n(\theta,\lambda,\delta)=0$ has a unique solution $(\hat\lambda(\theta),\hat\delta(\theta))$ satisfying $\|\hat\lambda(\theta)\|+\|\hat\delta(\theta)-\delta_0\|\le \kappa$.

Proof. See Appendix (ref).

We can now characterize the large-sample behavior of the estimator obtained from minimizing $L_n(\theta)$. The following theorem establishes the consistency of $\hat{\theta}$.

theorem[Consistency of $\hat\theta$] Under Assumptions (ref),(ref),(ref), and Theorem (ref), any measurable selection $\hat\theta\in\arg\min_{\theta\in\Theta}L_n(\theta)$ satisfies $ \hat\theta\xrightarrow{p}\theta_0. $

Proof. See Appendix (ref).

The consistency of $\hat{\theta}$ further implies the consistency of the associated Lagrange multipliers. The following corollary formalizes this result.

corollary[Consistency of multipliers] Suppose $\hat{\theta} \xrightarrow{p} \theta_0$. For each $\theta$ in a neighborhood of $\theta_0$, let $(\lambda(\theta),\delta(\theta))$ denote the unique solution to $ \Psi(\theta,\lambda,\delta)=0, $ and assume that the mapping $ \theta \mapsto (\lambda(\theta),\delta(\theta)) $ is continuous at $\theta_0$. Define $ (\hat{\lambda},\hat{\delta}) := (\lambda(\hat{\theta}),\delta(\hat{\theta})). $ Then $ (\hat{\lambda},\hat{\delta}) \xrightarrow{p} (\lambda_0,\delta_0), $ where $(\lambda_0,\delta_0) = (\lambda(\theta_0),\delta(\theta_0))$.

Proof. See Appendix (ref)

Having established the consistency of the estimator and the associated multipliers, we next characterize their asymptotic behavior. The following result provides the first-order expansion of $\hat{\lambda}$ and the convergence rate of $\hat{\delta}$.

theorem[First-order expansion for $\hat\lambda(\hat\theta)$ and rate for $\hat\delta(\hat\theta)$] Under Assumptions (ref),(ref),(ref),(ref) and Theorem (ref), let $\hat\lambda:=\hat\lambda(\hat\theta)$ and $\hat\delta:=\hat\delta(\hat\theta)$. Then, under correct specification, \[ \hat\lambda = \hat\Omega(\theta_0)^{-1}\bar g(\theta_0) +o_p(n^{-1/2}), \qquad \hat\delta-\delta_0 = O_p(n^{-1}), \] and \[ \sqrt n\,\hat\lambda \ \xrightarrow{d}\ N(0,\Omega_0^{-1}). \]

Proof. See Appendix (ref).

While Theorem (ref) establishes the convergence rate of $\hat{\delta}$, its limiting distribution requires a finer expansion. The following theorem characterizes the nondegenerate limit of the probability multiplier.

theorem[Nondegenerate limit for the probability multiplier] Assume the conditions of Theorem (ref). Suppose additionally that a third-order remainder control holds so that the quadratic approximation in $t_i(\theta)$ is valid at the $n^{-1}$ scale. Then \[ n(\hat\delta-\delta_0) = -\frac{\gamma+1}{2}\,\Big(\sqrt n\,\bar g(\theta_0)\Big)^{\prime}\Omega_0^{-1}\Big(\sqrt n\,\bar g(\theta_0)\Big) +o_p(1),\] and hence \[ n(\hat\delta-\delta_0)\ \xrightarrow{d}\ -\frac{\gamma+1}{2}\,\chi^{2}_{q}, \] where $\chi^{2}_{q}$ denotes a chi-square random variable with $q$ degrees of freedom.

Proof. See Appendix (ref).

With the expansions for the multipliers established, we now derive the asymptotic distribution of the estimator.

theorem[Asymptotic normality of $\hat\theta$] Under Assumptions (ref),(ref),(ref),(ref) and Theorem (ref). Then \[ \sqrt n\,(\hat\theta-\theta_0) = -(G_0^{\prime}\Omega_0^{-1}G_0)^{-1}G_0^{\prime}\Omega_0^{-1}\sqrt n\,\bar g(\theta_0) +o_p(1), \] and hence \[ \sqrt n\,(\hat\theta-\theta_0)\ \xrightarrow{d}\ N\!\left(0,\ (G_0^{\prime}\Omega_0^{-1}G_0)^{-1}\right). \]

Proof. See Appendix (ref).

To facilitate the second-order analysis, we first derive a second-order expansion of the moment equation implied by the CRPD constraints.

lemma[Second-order expansion of the moment equation] Let $(\hat\theta,\hat\lambda,\hat\delta)$ satisfy the CRPD constraints: \[ \frac{1}{n}\sum_{i=1}^n w_i(\hat\theta,\hat\lambda,\hat\delta)=1, \qquad \frac{1}{n}\sum_{i=1}^n w_i(\hat\theta,\hat\lambda,\hat\delta)\,g_i(\hat\theta)=0. \] Then, uniformly on events of probability approaching one, \begin{align} 0 &= \bar g(\hat\theta) -\hat\Omega(\hat\theta)\hat\lambda +\frac{1-\gamma}{2}\cdot \frac{1}{n}\sum_{i=1}^n \bigl(t_i(\hat\theta,\hat\lambda,\hat\delta)\bigr)^2\,g_i(\hat\theta) +o_p(n^{-1}). \end{align}

Proof. See Appendix (ref).

Using the second-order expansion of the moment equation in Lemma (ref), we can now refine the first-order result for the multiplier. The following theorem provides the second-order expansion of $\hat{\lambda}$.

theorem[Second-order expansion for $\hat\lambda$] Under the conditions of Lemma (ref) and the first-order expansion \[ \hat\lambda = \hat\Omega(\theta_0)^{-1}\bar g(\theta_0) + o_p(n^{-1/2}), \] we have \begin{equation} \hat\lambda = \hat\Omega(\theta_0)^{-1}\bar g(\theta_0) \;+\; \frac{1}{n}\,B_{\lambda,n}(\gamma) \;+\; o_p(n^{-1}), \end{equation} where the sample-dependent second-order correction is \begin{equation} B_{\lambda,n}(\gamma) := \frac{1-\gamma}{2}\, \hat\Omega(\theta_0)^{-1} \Bigg[ \frac{1}{n}\sum_{i=1}^n \Bigl( (\hat\Omega(\theta_0)^{-1}\sqrt n\,\bar g(\theta_0))^{\prime} g_i(\theta_0) \Bigr)^2 g_i(\theta_0) \Bigg], \end{equation} and $B_{\lambda,n}(\gamma)=O_p(1)$.

Proof. See Appendix (ref).

Finally, combining the first-order asymptotic representation of $\hat{\theta}$ with the second-order expansion of the multiplier $\hat{\lambda}$, we can now derive the second-order expansion of the CRPD estimator.

theorem(Second-order expansion for $\hat\theta$) Assume the conditions of Theorems (ref) and (ref). Then the CRPD estimator $\hat{\theta}$ admits the expansion \begin{equation*} \hat{\theta}-\theta_0 = \big(G_0^{\prime}\Omega_0^{-1}G_0\big)^{-1} G_0^{\prime}\Omega_0^{-1}\bar g(\hat{\theta}) + \frac{1}{n}B_{\theta,n}(\gamma) + o_p(n^{-1}), \end{equation*} where $\bar g(\theta)=\frac{1}{n}\sum_{i=1}^n g_i(\theta)$ and $ B_{\theta,n}(\gamma) = \big(G_0^{\prime}\Omega_0^{-1}G_0\big)^{-1}\Delta_n(\gamma)=O_p(1)$, with $\Delta_n(\gamma)=O_p(1)$ collects all order-$n^{-1}$ terms arising from the second-order expansion of the Lagrange multiplier, the curvature term involving $\bar H(\theta_0)$, and plug-in effects from replacing sample matrices with their population limits.

Proof. See Appendix (ref).

Our asymptotic analysis characterizes the large-sample behavior of the CRPD estimator and its associated multipliers. The estimator $\hat{\theta}$ is consistent and asymptotically normal, sharing the same first-order limit as standard moment-based estimators. The power parameter $\gamma$ does not affect the leading asymptotic distribution, but it appears in the order-$n^{-1}$ bias term arising from the second-order expansions of the multipliers and the estimator. Consequently, the choice of $\gamma$ influences finite-sample performance while leaving first-order asymptotic efficiency unchanged.

Monte Carlo simulations

We conduct Monte Carlo simulations with 1{,}000 replications. Each replication uses sample sizes $n\in\{25,50,100\}$ to represent small and moderate sample settings commonly encountered in empirical applications. The objective is to evaluate finite-sample behavior across different tail environments and to examine how the CRPD power parameter $\gamma$ affects higher-order performance while preserving first-order asymptotic properties.

We consider estimation of the parameter vector $\theta=(\mu,\sigma^2)$ using moment conditions based on the first two central moments. In addition, we impose a third moment condition on skewness. Specifically, the moment conditions are \[ \mathrm{E}[X-\mu]=0,\qquad \mathrm{E}\!\left[(X-\mu)^2-\sigma^2\right]=0,\qquad \mathrm{E}\!\left[(X-\mu)^3\right]=0. \] The third moment restriction is a valid population moment under the data-generating processes considered below. Because the Student-$t$ and normal distributions are symmetric, the population third central moment equals zero. Thus, the third restriction does not introduce an additional structural parameter, but it renders the system overidentified. This avoids the degeneracy that can arise when too few moment conditions are imposed, under which the CRPD objective may admit a trivial solution with uniform weights.

Data are generated from symmetric $t$ and normal distributions. Let $Z_i \sim t_{\nu}$ with mean zero and degrees of freedom $\nu \in \{5,15\}$. The variance of the $t$ distribution is $\nu/(\nu-2)$, while the variance of the normal distribution is set to 1. Smaller values of $\nu$ correspond to heavier tails, whereas larger values approximate the normal distribution. This allows us to evaluate estimator performance across a range of tail behaviors.

We evaluate the estimator over a grid $\gamma \in [-1,1]$ with step size $0.25$ to illustrate how finite-sample performance varies with the choice of $\gamma$. For each replication and each candidate value of $\gamma$, the CRPD estimator is computed by searching over grids for $(\mu,\sigma^2)$ that cover a data-driven region around the sample mean and variance. For every candidate triple $(\mu,\sigma^2,\gamma)$, we solve for the associated multipliers $(\lambda,\delta)$ and the implied optimal weights $\{\pi_i\}$, and evaluate the CRPD objective. Each replication yields the estimators $\hat{\mu}(\gamma)$ and $\hat{\sigma}^2(\gamma)$, together with the associated Lagrange multipliers and implied weights.

We evaluate simulation performance using both higher-order and first-order measures. Finite-sample performance is evaluated using bias, mean squared error (MSE), and coverage distortion. Bias is defined as the average of $\hat{\mu}-\mu_0$ across replications. MSE denotes the mean squared error of $\hat{\mu}$.\footnote{Although MSE decomposes as $\mathrm{Var}(\hat{\theta})+\mathrm{Bias}(\hat{\theta})^{2}$, the leading variance term is governed by the first-order asymptotic approximation and is invariant to $\gamma$ in our setting. Differences in MSE across values of $\gamma$ therefore arise primarily through bias, which is driven by higher-order terms. For this reason, MSE is interpreted here as a finite-sample (higher-order) performance measure.} Coverage distortion is defined as the deviation of the empirical coverage probability from the nominal 95% level.\footnote{Coverage distortion measures the deviation of the empirical coverage probability from the nominal confidence level. Because first-order asymptotic theory guarantees correct coverage only in the limit, deviations from the nominal level arise from higher-order terms in finite samples. Coverage distortion therefore reflects higher-order accuracy of the estimator and the associated standard error.} Bias, MSE, and coverage distortion are interpreted as higher-order performance measures because they reflect finite-sample deviations from the first-order asymptotic approximation.

First-order behavior is assessed by comparing the empirical standard deviation of $\hat{\mu}$ with its average estimated standard error. The empirical SD is the standard deviation of $\hat{\mu}$ across replications, and the mean SE is the average of the estimated standard errors. The SD--SE ratio denotes the ratio of the empirical SD to the mean SE. The empirical SD, mean SE, and their ratio summarize first-order behavior by capturing the variability of the estimator and the accuracy of the first-order standard error approximation.

The simulation results for small samples are reported in Tables (ref), (ref), and (ref) for $n \in \{25,50\}$ and Tables S.1, S.2 in Supplement, and the results for a larger sample size of $n=100$ are reported in Tables S.3 and S.4 in Supplement. Table (ref) summarizes both first- and second-order performance metrics. We observe that the higher-order measures (bias, MSE, and coverage distortion) and the first-order diagnostic SD--SE ratio vary with the value of $\gamma$. This indicates that when the sample size is small, both first- and higher-order measures fluctuate. The first-order measures vary because the sample size has not yet reached the asymptotic regime, while the higher-order measures vary because finite-sample effects remain non-negligible when $n$ is small. These findings suggest that finite-sample considerations are important and that estimator performance depends on the choice of $\gamma$.

The optimal value of $\gamma$ depends on the tail behavior of the data-generating process (DGP). Under heavier tails (e.g., $t_{5}$), negative values of $\gamma$ tend to produce smaller MSE. As the tails become lighter and approach the normal distribution, positive values of $\gamma$ tend to yield smaller MSE and smaller coverage distortions. With an appropriate choice of $\gamma$, coverage improves as $n$ moves toward the asymptotic regime. In particular, coverage distortion decreases as the sample size increases from $n=25$ to $n=50$. For example, under the $t_{5}$ distribution, $\gamma=-1$ yields the lowest MSE when $n=25$, with coverage distortion $-0.1090$ and an SD--SE ratio of $1.3841$. When $n=50$, the lowest MSE occurs at $\gamma=-0.75$, with coverage distortion $-0.0373$ and an SD--SE ratio of $1.1584$, both of which are smaller than their counterparts at $n=25$.

Figure (ref) plots second-order bias and coverage distortion for DGPs given by a $t_{5}$ distribution and a normal distribution with $n=50$.\footnote{We use $n=50$ in this figure to ensure that first-order asymptotic approximations are reasonably accurate. When $n=25$, the sample size appears too small for first-order behavior to be well approximated. At $n=50$, first-order effects are more stable, while second-order effects remain non-negligible.} The filled circle indicates the value of $\gamma$ that minimizes the absolute value of the metric. For the $t_{5}$ DGP, this optimal $\gamma$ tends to be negative for both bias and coverage distortion, reflecting the need to downweight outliers under heavy tails. In contrast, for the normal DGP, the optimal $\gamma$ is positive across both metrics. This pattern suggests that $\gamma$ acts as a data-dependent tuning parameter, adjusting to distributional features to improve finite-sample performance.

Tables (ref), S.2, and (ref) report the estimates of $\mu$, the associated Lagrange multipliers, and summary statistics of the estimated weights $\boldsymbol{\pi}$. Because the model is overidentified through the additional skewness restriction, the Lagrange multipliers associated with the moment conditions are generally nonzero. Consequently, the estimated weights $\boldsymbol{\pi}$ do not collapse to the uniform distribution. Depending on the characteristics of the data and the value of $\gamma$, the multipliers $\lambda$ and $\delta$ can occasionally take relatively large values, reflecting the degree of adjustment in the probability weights required to satisfy the moment restrictions under the chosen divergence parameter $\gamma$.

These findings are consistent with the theoretical results developed earlier. In particular, the first-order asymptotic distribution of the CRPD estimator does not depend on the value of $\gamma$, while higher-order terms governing finite-sample bias and coverage distortion may vary with $\gamma$. The simulation evidence reflects this structure: measures related to first-order behavior, such as the empirical SD and mean SE, remain relatively stable across different values of $\gamma$, whereas higher-order measures such as bias, MSE, and coverage distortion display noticeable variation. This pattern confirms that the choice of $\gamma$ primarily influences higher-order finite-sample properties rather than the leading asymptotic distribution.

The dependence of the optimal $\gamma$ on tail behavior can be interpreted through the way the divergence parameter shapes the implied probability weights. Negative values of $\gamma$ tend to assign relatively more weight to observations with larger moment residuals, which can help stabilize estimation in heavy-tailed environments. In contrast, positive values of $\gamma$ penalize large residuals more strongly, leading to more concentrated weights that perform better when the underlying distribution is closer to normal. As a result, the value of $\gamma$ effectively controls how aggressively the estimator adjusts the implied probabilities in response to deviations from the moment conditions.

Overall, the simulation results highlight the importance of the divergence parameter in finite samples. Although all values of $\gamma$ share the same first-order asymptotic properties, finite-sample performance can differ substantially depending on the choice of $\gamma$ and the tail behavior of the underlying distribution. This suggests that selecting $\gamma$ with attention to finite-sample considerations can lead to meaningful improvements in estimation accuracy and coverage performance.

A caveat is that the relationship between MSE and $\gamma$ is generally non-monotonic. Consequently, the value of $\gamma$ that minimizes MSE may occur at an interior point rather than at the boundary of the admissible range. This highlights the importance of evaluating the estimator over a sufficiently wide window of $\gamma$ values; restricting the search to a narrow range could fail to identify the value of $\gamma$ that delivers the best finite-sample performance.

Discussion

In classical regularization problems, a hyperparameter is chosen to balance bias and variance, typically by minimizing MSE. For an estimator indexed by a tuning parameter $\zeta$, one often has \[ \mathrm{MSE}(\zeta) = \mathrm{Var}(\hat{\theta}_\zeta) + \mathrm{Bias}(\hat{\theta}_\zeta)^2. \] The optimal value of $\zeta$ therefore reflects a bias--variance tradeoff. In the CRPD framework, however, the role of the power parameter $\gamma$ is structurally different.

The first-order asymptotic distribution of $\hat{\theta}$ does not depend on $\gamma$: \[ \sqrt{n}(\hat{\theta}-\theta_0) \;\xrightarrow{d}\; N(0,V), \] so the leading variance component in the MSE, $V/n$, is invariant to $\gamma$.

table[table omitted — 7,402 chars of source]
figure[figure omitted — 841 chars of source]
landscape\begin{table}[] \caption{Simulation results ($t_{5}$)} \resizebox{1.35\textheight}{!}{ \begin{tabular}{cccccccccccccc} \hline \multirow{2}{*}{$n$} & \multirow{2}{*}{$\gamma$} & \multirow{2}{*}{$\hat{\theta}$} & \multicolumn{4}{c}{Lagrange multipliers} & \multicolumn{6}{c}{$\boldsymbol{\pi}$} \\ \cline{4-13} & & & $\lambda_{1}$ & $\lambda_{2}$ & $\lambda_{3}$ & $\delta$ & Minimum & Q1 & Median & Mean & Q3 & Maximum \\ \hline 25 & -1.00 & -0.0174 & -0.0648 & -0.0797 & 0.1092 & -1.00E+04 & 0.0208 & 0.0337 & 0.0374 & 0.0400 & 0.0449 & 0.0642 \\ & & (0.3034) & (1.2644) & (1.3803) & (2.0990) & (1.42E-05) & (0.0122) & (0.0051) & (0.0043) & (3.88E-09) & (0.0042) & (0.0336) \\ & -0.75 & -0.0354 & -2.92E+10 & -2.01E+10 & 1.09E+10 & -1.70E+09 & 0.0192 & 0.0331 & 0.0367 & 0.0388 & 0.0437 & 0.0600 \\ & & (0.3283) & (5.15E+11) & (3.44E+11) & (1.85E+11) & (3.46E+10) & (0.0128) & (0.0075) & (0.0076) & (0.0067) & (0.0087) & (0.0344) \\ & -0.50 & -0.2125 & -3.38E+12 & -1.91E+12 & 8.70E+11 & -2.47E+11 & 0.0124 & 0.0258 & 0.0288 & 0.0304 & 0.0345 & 0.0489 \\ & & (0.4823) & (9.04E+13) & (4.52E+13) & (1.78E+13) & (6.65E+12) & (0.0129) & (0.0152) & (0.0167) & (0.0171) & (0.0198) & (0.0459) \\ & -0.25 & -0.4166 & -1.66E+12 & -1.87E+12 & 1.14E+12 & -2.04E+11 & 0.0088 & 0.0178 & 0.0197 & 0.0211 & 0.0231 & 0.0398 \\ & & (0.5203) & (1.53E+13) & (1.82E+13) & (1.09E+13) & (2.01E+12) & (0.0126) & (0.0177) & (0.0194) & (0.0200) & (0.0227) & (0.0723) \\ & 0.00 & -0.1347 & 4.51E+01 & -8.70E+01 & 2.70E+01 & -8.82E+01 & 0.0155 & 0.0296 & 0.0326 & 3.30E+74 & 0.0370 & 8.24E+75 \\ & & (0.4425) & (3.40E+02) & (5.26E+02) & (1.51E+02) & -6.68E+02 & (0.0138) & (0.0140) & (0.0151) & -1.04E+76 & (0.0172) & (2.61E+77) \\ & 0.25 & -0.0007 & 0.0070 & -0.0190 & -0.0024 & -0.8618 & 0.0164 & 0.0352 & 0.0388 & 0.0400 & 0.0446 & 0.0598 \\ & & (0.3331) & (0.4426) & (0.0952) & (0.1307) & (0.1263) & (0.0137) & (0.0053) & (0.0041) & (5.43E-10) & (0.0038) & (0.0354) \\ & 0.50 & -0.0205 & 0.0271 & -0.0188 & -0.0036 & -0.7522 & 0.0160 & 0.0347 & 0.0385 & 0.0400 & 0.0447 & 0.0612 \\ & & (0.3754) & (0.4858) & (0.0874) & (0.1254) & (0.1810) & (0.0139) & (0.0064) & (0.0045) & (2.24E-10) & (0.0038) & (0.0367) \\ & 0.75 & -0.0436 & 0.0498 & -0.0173 & -0.0040 & -0.6803 & 0.0157 & 0.0341 & 0.0385 & 0.0400 & 0.0452 & 0.0610 \\ & & (0.3929) & (0.5223) & (0.0872) & (0.1217) & (0.2200) & (0.0141) & (0.0075) & (0.0044) & (1.26E-10) & (0.0042) & (0.0326) \\ & 1.00 & -0.0685 & 0.0919 & -0.0200 & -0.0051 & -0.6494 & 0.0157 & 0.0331 & 0.0383 & 0.0400 & 0.0460 & 0.0627 \\ & & (0.4215) & (0.6059) & (0.0800) & (0.1294) & (0.2983) & (0.0142) & (0.0084) & (0.0045) & (1.44E-10) & (0.0052) & (0.0327) \\ \hline 50 & -1.00 & -0.0020 & 0.0018 & -0.0009 & -1.50E-05 & -1.00E+04 & 0.0101 & 0.0180 & 0.0195 & 0.0200 & 0.0219 & 0.0306 \\ & & (0.1906) & (0.2734) & (0.0112) & (0.0843) & (3.15E-06) & (0.0057) & (0.0016) & (0.0009) & (1.31E-11) & (0.0015) & (0.0086) \\ & -0.75 & -0.0034 & -1.10E+05 & -9.25E+04 & 5.64E+04 & -1.87E+03 & 0.0096 & 0.0182 & 0.0196 & 0.0200 & 0.0218 & 0.0298 \\ & & (0.1898) & (3.46E+06) & (2.92E+06) & (1.78E+06) & (5.90E+04) & (0.0059) & (0.0017) & (0.0012) & (0.0006) & (0.0017) & (0.0130) \\ & -0.50 & -0.0730 & -1.09E+12 & -1.09E+12 & 4.96E+11 & -1.45E+11 & 0.0079 & 0.0169 & 0.0182 & 0.0185 & 0.0202 & 0.0269 \\ & & (0.3283) & (2.26E+13) & (2.49E+13) & (1.05E+13) & (3.10E+12) & (0.0062) & (0.0051) & (0.0053) & (0.0053) & (0.0060) & (0.0102) \\ & -0.25 & -0.3466 & -3.81E+12 & -2.51E+12 & 1.38E+12 & -2.32E+11 & 0.0054 & 0.0122 & 0.0130 & 0.0131 & 0.0141 & 0.0195 \\ & & (0.5162) & (5.27E+13) & (3.10E+13) & (1.57E+13) & (3.10E+12) & (0.0064) & (0.0089) & (0.0095) & (0.0095) & (0.0103) & (0.0214) \\ & 0.00 & -0.0412 & -0.2243 & -1.1852 & 1.0296 & -1.2787 & 0.0083 & 0.0181 & 0.0192 & 0.0192 & 0.0207 & 0.0257 \\ & & (0.2638) & (16.4336) & (12.3150) & (7.5537) & (10.2533) & (0.0065) & (0.0038) & (0.0039) & (0.0038) & (0.0043) & (0.0076) \\ & 0.25 & -0.0267 & 0.0396 & -0.0063 & -0.0049 & -0.8422 & 0.0082 & 0.0183 & 0.0196 & 0.0200 & 0.0216 & 0.0304 \\ & & (0.2910) & (0.2994) & (0.0276) & (0.0613) & (0.1181) & (0.0067) & (0.0025) & (0.0017) & (6.68E-11) & (0.0015) & (0.0228) \\ & 0.50 & -0.0380 & 0.0395 & -0.0088 & -0.0029 & -0.7349 & 0.0081 & 0.0180 & 0.0194 & 0.0200 & 0.0217 & 0.0313 \\ & & (0.3263) & (0.3666) & (0.0354) & (0.0601) & (0.2089) & (0.0069) & (0.0032) & (0.0021) & (3.04E-11) & (0.0016) & (0.0253) \\ & 0.75 & -0.0427 & 0.0445 & -0.0101 & -0.0015 & -0.6584 & 0.0176 & 0.0194 & 0.0200 & 0.0220 & 0.0307 & 0.0070 \\ & & (0.3374) & (0.4073) & (0.0435) & (0.0626) & (0.2283) & (0.0070) & (0.0037) & (0.0020) & (1.71E-11) & (0.0021) & (0.0197) \\ & 1.00 & -0.0565 & 0.0677 & -0.0158 & -0.0019 & -0.6172 & 0.0172 & 0.0194 & 0.0200 & 0.0225 & 0.0310 & 0.0071 \\ & & (0.3611) & (0.4643) & (0.0696) & (0.0629) & (0.2600) & (0.0071) & (0.0042) & (0.0019) & (8.12E-11) & (0.0028) & (0.0174) \\ \hline \multicolumn{13}{l}{ \begin{minipage}{26cm} Note: Results are based on 1{,}000 Monte Carlo replications. The data-generating process follows a $t$ distribution with 5 degrees of freedom and mean zero. Empirical standard deviations are reported in parentheses. $E$ denotes scientific notation. $\nu$ denotes the degrees of freedom, $n$ the number of observations, and $\gamma$ the power parameter. $\hat{\theta}$ is the estimate of the structural parameter $\theta$. $\lambda_{1}, \lambda_{2}, \lambda_{3}$ are the Lagrange multipliers associated with the moment conditions, and $\delta$ is the Lagrange multiplier associated with the probability constraint. $\boldsymbol{\pi}$ denotes the estimated weights on the observations. Q1 and Q3 denote the first and third quartiles, respectively. \end{minipage}} \end{tabular}} \end{table}
landscape\begin{table}[] \caption{Simulation results (Normal)} \resizebox{1.35\textheight}{!}{ \begin{tabular}{ccccccccccccc} \hline \multirow{2}{*}{$n$} & \multirow{2}{*}{$\gamma$} & \multirow{2}{*}{$\hat{\theta}$} & \multicolumn{4}{c}{Lagrange multipliers} & \multicolumn{6}{c}{$\boldsymbol{\pi}$} \\ \cline{4-13} & & & $\lambda_{1}$ & $\lambda_{2}$ & $\lambda_{3}$ & $\delta$ & Minimum & Q1 & Median & Mean & Q3 & Maximum \\ \hline 25 & -1.00 & -0.0180 & -0.0522 & -0.0309 & 0.0812 & -1.00E+04 & 0.0270 & 0.0339 & 0.0376 & 0.0400 & 0.0447 & 0.0588 \\ & & (0.2317) & (0.6503) & (0.3998) & (0.7750) & (1.16E-05) & (0.0096) & (0.0049) & (0.0038) & (3.79E-09) & (0.0035) & (0.0295) \\ & -0.75 & -0.0416 & -2.82E+09 & -2.47E+09 & 1.64E+09 & -1.23E+08 & 0.0252 & 0.0327 & 0.0363 & 0.0382 & 0.0429 & 0.0540 \\ & & (0.2602) & (6.35E+10) & (5.11E+10) & (3.00E+10) & (2.83E+09) & (0.0111) & (0.0082) & (0.0086) & (0.0083) & (0.0100) & (0.0314) \\ & -0.50 & -0.2743 & -1.99E+12 & -2.46E+12 & 1.59E+12 & -1.44E+11 & 0.0141 & 0.0205 & 0.0231 & 0.0247 & 0.0281 & 0.0386 \\ & & (0.4291) & (3.77E+13) & (4.33E+13) & (2.56E+13) & (2.39E+12) & (0.0140) & (0.0167) & (0.0186) & (0.0194) & (0.0225) & (0.0463) \\ & -0.25 & -0.4780 & -2.45E+12 & -2.87E+12 & 2.31E+12 & -2.09E+11 & 0.0075 & 0.0123 & 0.0138 & 0.0148 & 0.0163 & 0.0263 \\ & & (0.4192) & (1.74E+13) & (1.91E+13) & (1.64E+13) & (1.73E+12) & (0.0124) & (0.0167) & (0.0186) & (0.0193) & (0.0220) & (0.0578) \\ & 0.00 & -0.1594 & 8.49E+01 & -1.88E+02 & 6.93E+01 & -1.36E+02 & 0.0203 & 0.0288 & 0.0317 & 5.35E+11 & 0.0366 & 1.34E+13 \\ & & (0.3896) & (5.99E+02) & (9.78E+02) & (3.47E+02) & (8.92E+02) & (0.0146) & (0.0141) & (0.0153) & (1.69E+13) & (0.0178) & (4.23E+14) \\ & 0.25 & -0.0003 & 0.0042 & -0.0036 & -0.0017 & -0.8362 & 0.0233 & 0.0349 & 0.0387 & 0.0400 & 0.0453 & 0.0529 \\ & & (0.2250) & (0.4652) & (0.0241) & (0.2268) & (0.0528) & (0.0130) & (0.0042) & (0.0029) & (6.60E-10) & (0.0038) & (0.0133) \\ & 0.50 & 0.0032 & 0.0119 & -0.0049 & -0.0062 & -0.7097 & 0.0227 & 0.0351 & 0.0388 & 0.0400 & 0.0452 & 0.0525 \\ & & (0.2317) & (0.4597) & (0.0264) & (0.1966) & (0.0692) & (0.0134) & (0.0044) & (0.0029) & (3.11E-10) & (0.0037) & (0.0140) \\ & 0.75 & 0.0036 & 0.0150 & -0.0033 & -0.0049 & -0.6222 & 0.0230 & 0.0350 & 0.0389 & 0.0400 & 0.0453 & 0.0523 \\ & & (0.2300) & (0.4732) & (0.0235) & (0.1815) & (0.0921) & (0.0136) & (0.0048) & (0.0028) & (1.63E-10) & (0.0039) & (0.0143) \\ & 1.00 & 0.0018 & 0.0153 & -0.0046 & -0.0034 & -0.5637 & 0.0221 & 0.0347 & 0.0390 & 0.0400 & 0.0456 & 0.0527 \\ & & (0.2423) & (0.5024) & (0.0329) & (0.1842) & (0.1226) & (0.0137) & (0.0052) & (0.0028) & (6.90E-11) & (0.0044) & (0.0150) \\ \hline 50 & -1.00 & -0.0060 & -0.0155 & 4.60E-05 & 0.0059 & -1.00E+04 & 0.0134 & 0.0179 & 0.0194 & 0.0200 & 0.0220 & 0.0272 \\ & & (0.1478) & (0.3186) & (0.0038) & (0.1322) & (1.95E-06) & (0.0046) & (0.0015) & (0.0009) & (9.24E-13) & (0.0013) & (0.0061) \\ & -0.75 & -0.0071 & -3.23E+06 & -2.92E+06 & 2.01E+06 & -8.57E+04 & 0.0130 & 0.0178 & 0.0193 & 0.0199 & 0.0219 & 0.0268 \\ & & (0.1602) & (5.49E+07) & (5.06E+07) & (3.50E+07) & (1.48E+06) & (0.0050) & (0.0021) & (0.0018) & (0.0017) & (0.0023) & (0.0062) \\ & -0.50 & -0.1240 & -1.57E+12 & -1.34E+12 & 1.02E+12 & -7.21E+10 & 0.0105 & 0.0151 & 0.0164 & 0.0168 & 0.0186 & 0.0231 \\ & & (0.3453) & (4.53E+13) & (3.84E+13) & (2.96E+13) & (1.97E+12) & (0.0066) & (0.0067) & (0.0072) & (0.0073) & (0.0082) & (0.0175) \\ & -0.25 & -0.4980 & -2.33E+12 & -2.54E+12 & 1.93E+12 & -1.89E+11 & 0.0047 & 0.0076 & 0.0083 & 0.0084 & 0.0093 & 0.0117 \\ & & (0.4615) & (1.85E+13) & (2.07E+13) & (1.57E+13) & (2.01E+12) & (0.0066) & (0.0090) & (0.0097) & (0.0099) & (0.0109) & (0.0135) \\ & 0.00 & -0.0493 & 0.5854 & -12.4907 & 7.1143 & -4.4559 & 0.0117 & 0.0172 & 0.0185 & 0.0188 & 0.0206 & 0.0241 \\ & & (0.2509) & (110.3248) & (182.4158) & (57.5511) & (150.0305) & (0.0062) & (0.0045) & (0.0047) & (0.0047) & (0.0053) & (0.0074) \\ & 0.25 & -0.0017 & -0.0130 & -0.0007 & 0.0062 & -0.8165 & 0.0120 & 0.0183 & 0.0197 & 0.0200 & 0.0219 & 0.0253 \\ & & (0.1491) & (0.2887) & (0.0077) & (0.1187) & (0.0243) & (0.0062) & (0.0014) & (0.0007) & (9.23E-11) & (0.0014) & (0.0044) \\ & 0.50 & -0.0021 & -0.0094 & -0.0028 & 0.0050 & -0.6869 & 0.0117 & 0.0183 & 0.0198 & 0.0200 & 0.0219 & 0.0251 \\ & & (0.1480) & (0.2914) & (0.0162) & (0.1112) & (0.0342) & (0.0065) & (0.0016) & (0.0008) & (4.47E-11) & (0.0015) & (0.0045) \\ & 0.75 & 0.0013 & -0.0018 & -0.0042 & 0.0019 & -0.5990 & 0.0114 & 0.0182 & 0.0197 & 0.0200 & 0.0220 & 0.0252 \\ & & (0.1563) & (0.3178) & (0.0249) & (0.1080) & (0.0591) & (0.0067) & (0.0019) & (0.0009) & (4.64E-11) & (0.0017) & (0.0057) \\ & 1.00 & -0.0089 & 0.0166 & -0.0063 & -0.0025 & -0.5346 & 0.0111 & 0.0181 & 0.0198 & 0.0200 & 0.0220 & 0.0252 \\ & & (0.1661) & (0.3275) & (0.0460) & (0.1001) & (0.0924) & (0.0068) & (0.0021) & (0.0008) & (3.36E-11) & (0.0018) & (0.0070) \\ \hline \multicolumn{13}{l}{ \begin{minipage}{26cm} Note: Results are based on 1{,}000 Monte Carlo replications. The data-generating process follows a normal distribution with mean zero and variance one. Empirical standard deviations are reported in parentheses. $E$ denotes scientific notation. $n$ the number of observations, and $\gamma$ the power parameter. $\hat{\theta}$ is the estimate of the structural parameter $\theta$. $\lambda_{1}, \lambda_{2}, \lambda_{3}$ are the Lagrange multipliers associated with the moment conditions, and $\delta$ is the Lagrange multiplier associated with the probability constraint. $\boldsymbol{\pi}$ denotes the estimated weights on the observations. Q1 and Q3 denote the first and third quartiles, respectively. \end{minipage}} \end{tabular}} \end{table}

The influence of $\gamma$ appears only in higher-order terms. A higher-order expansion yields \[ \hat{\theta}-\theta_0 = \frac{B(\gamma)}{n} + \frac{1}{\sqrt{n}}Z + o(n^{-1}), \] which implies \[ \mathrm{MSE}_n(\gamma) = \frac{V}{n} + \frac{B(\gamma)^2}{n^2} + o(n^{-2}). \] Hence differences in MSE across values of $\gamma$ arise from second-order bias rather than from changes in asymptotic efficiency. Selecting $\gamma$ to minimize MSE (or its empirical proxy via cross-validation) therefore amounts to choosing the loss geometry that minimizes finite-sample distortion while leaving first-order efficiency unchanged.

Given this interpretation, $\gamma$ can be selected in a data-driven manner, for example through cross-validation. The procedure separates hyperparameter selection from structural estimation. For each candidate $\gamma$, the parameters $(\theta,\lambda,\delta)$ are estimated subject to the moment and probability constraints, while $\gamma$ determines the curvature of the CRPD objective. Each candidate value of $\gamma$ is then evaluated using a $K$-fold cross-validation criterion based on out-of-sample moment instability, which serves as a proxy for finite-sample risk. The final estimator is obtained by refitting the model on the full sample at the selected value $\hat{\gamma}$. The pseudo-algorithm is presented in Algorithm (ref).

As a caveat, note that because the CRPD divergence $\mathscr I_\gamma(\cdot)$ varies with $\gamma$, its numerical values are not directly comparable across candidate values of $\gamma$. Each value of $\gamma$ defines a different divergence and therefore a different loss geometry. Consequently, the minimized value of the CRPD objective reflects the fit of the model under that particular loss function rather than a common scale across specifications. Comparing these minimized values across $\gamma$ would therefore be misleading, as the underlying optimization problems differ. Cross-validation therefore evaluates $\gamma$ using a common out-of-sample moment criterion, selecting the value that yields the most stable moment fit.

Empirical Application

To illustrate our proposed method in a simple empirical setting, we use the dataset from Owen (2001, Table 3.2, p. 44), which reports the pounds of milk produced and the number of days milked in 1936 for $n=22$ dairy cows. Suppose we want to estimate the mean daily milk yield $\theta$.

Table (ref) reports descriptive statistics for milk production per day (mpd). The sample consists of 22 observations. The average milk production per day is 12.43 units, with a standard deviation of 3.08, indicating moderate dispersion across cows. The minimum and maximum values are 7.55 and 18.66, respectively, suggesting substantial variation in daily milk productivity within the sample. Overall, the distribution of mpd reflects heterogeneity in production levels across cows.

algorithm[algorithm omitted — 1,868 chars of source]
table[table omitted — 603 chars of source]

\paragraph{Empirical specification.} Let $X_i=\textit{mpd}_i$ denote milk production per day for cow $i$. Suppose our parameter of interest is the population mean $ \mu = \mathrm{E}[X_i]. $ The identifying moment condition is therefore $ \mathrm{E}[X_i-\mu]=0. $ However, with a single moment condition the CRPD estimator collapses to the trivial uniform-weight solution, since the sample mean satisfies the moment restriction exactly.

To generate an overidentified system, we introduce an instrument moment based on the observable variable $\textit{Days}_i$. Specifically, we impose $ \mathrm{E}\big[(X_i-\mu)\,Days_i\big]=0. $ This restriction implies that deviations of daily milk production from the population mean should not systematically vary with the number of observed days. A Pearson correlation test yields a $p$-value of $0.32$, indicating that the correlation is not statistically significant and that the moment condition is consistent with the data. Thus, we consider the moment system \[ g(X_i,\mu)=

pmatrix[pmatrix omitted — 41 chars of source]

, \] which provides two moment conditions for the single parameter $\mu$, yielding an overidentified model ($q=2>p=1$) suitable for CRPD estimation.

\paragraph{Data-driven selection of $\gamma$.} The CRPD estimator depends on the power parameter $\gamma$, which determines the curvature of the divergence and therefore the implied weighting scheme. To select $\gamma$ in a data-driven way, we implement five-fold cross-validation over a grid $ \gamma \in [-10,10] $ with step size $0.001$. For each candidate value of $\gamma$, the estimator is computed on the training folds, and the out-of-sample risk is evaluated on the validation fold using the squared prediction loss $ \frac{1}{n_v}\sum_{i\in V}(X_i-\hat{\mu})^2, $ where $V$ denotes the validation sample and $n_v$ its size. The cross-validated estimate $\hat{\gamma}$ is chosen as the value minimizing the average validation risk across folds. The final CRPD estimator of $\mu$ is then obtained by re-estimating the model using the full sample at $\gamma=\hat{\gamma}$.

\paragraph{Estimation results.} Table (ref) reports the estimation results across several specifications of the CRPD power parameter. When $\gamma$ is selected via cross-validation, the procedure yields $\hat{\gamma}=-0.654$ and $\hat{\theta}=12.8245$ lb/day. The table also illustrates the potential risks of selecting a specific value of $\gamma$ a priori: extreme values of the power parameter can substantially alter the estimator. For example, when $\gamma=-1.5$ or $\gamma=-2$, the estimated mean shifts noticeably and the divergence criterion increases sharply. This behavior indicates that highly negative values of $\gamma$ produce a loss geometry that requires substantial distortion of the empirical distribution in order to satisfy the moment conditions. In contrast, for moderate values of $\gamma$ the estimator remains stable, suggesting that the data are broadly compatible with the moment restrictions without excessive reweighting.

While the point estimate $\hat{\theta}$ remains nearly invariant across moderate values of $\gamma$ up to the fourth decimal place, the implied probability weights $\hat{\boldsymbol{\pi}}(\gamma)$ differ even when $\hat{\theta}$ values appear similar (for example, between the pre-specified choices $\gamma=-1$, $\gamma=0$, and the cross-validated $\hat{\gamma}=-0.654$). Figure (ref) illustrates these differences in the estimated weights. The dashed line denotes the uniform weight benchmark $1/n$. When $\gamma$ is close to zero (ET), the weights cluster tightly around this benchmark, indicating that the moment conditions are nearly satisfied under equal weighting. As $\gamma$ becomes more negative (EL-like), the weight distribution becomes more dispersed, assigning greater influence to a subset of observations in order to balance the moment equations. The cross-validated estimate $\hat{\gamma}=-0.654$ lies between these regimes and produces an intermediate weighting scheme.

Notably, the estimated probability weights $\hat{\boldsymbol{\pi}}(\gamma)=\{\hat{\pi}_i(\gamma)\}_{i=1}^n$ provide useful diagnostic information about how the estimator uses the data. Because the estimator can be written as $ \hat{\theta}(\gamma)=\sum_{i=1}^n \hat{\pi}_i(\gamma)X_i, $ the distribution of the weights reveals how influence is allocated across observations. Weights close to the uniform benchmark $1/n$ indicate that the moment restrictions are broadly compatible with the empirical distribution and that the estimator effectively averages information across the sample. In contrast, a more dispersed or heavy-tailed weight distribution indicates that the estimator relies disproportionately on certain observations in order to satisfy the moment conditions.

Such patterns provide practical insight into the structure of the data: unusually large weights may highlight observations that carry strong identifying information or exert substantial influence on the moment equations, while systematic deviations from uniform weighting may signal heterogeneity or tension between the model restrictions and the observed sample. In some cases, clusters of observations with similar weights may suggest latent subgroup structure, indicating that the global moment restriction interacts differently with different parts of the data.

Overall, the results illustrate why $\gamma$ should be interpreted as a continuum parameter and selected in a data-driven manner. Different values of $\gamma$ correspond to different ways of reconciling the sample with the imposed moment restrictions through reweighting of the empirical distribution. Even when the point estimate $\hat{\theta}$ appears insensitive to $\gamma$, the distribution of the estimated weights can differ substantially, revealing meaningful differences in the implied weighting schemes and hence in the finite-sample geometry of the estimator. Selecting $\gamma$ adaptively helps avoid extreme weighting schemes that may introduce non-negligible second-order bias in $\hat{\theta}$ or lead to disproportionate influence of a small subset of observations.

table[table omitted — 1,113 chars of source]
figure[figure omitted — 916 chars of source]

Conclusion

This paper develops the Cressie--Read power divergence (CRPD) framework for moment-based estimation with a focus on finite-sample behavior, interpreting the power parameter $\gamma$ as a hyperparameter. The CRPD family is indexed by $\gamma$, which has traditionally been treated as a researcher-chosen tuning index, with common benchmark values including empirical likelihood ($\gamma=-1$) and exponential tilting ($\gamma=0$). We argue that $\gamma$ is more naturally understood as a parameter governing the curvature of the divergence objective. Through this channel, it shapes the geometry of the implied weighting scheme and can materially affect finite-sample performance, even though it does not alter population identification.

This interpretation clarifies how divergence-based moment estimation can be adapted to the empirical environment. The choice of $\gamma$ controls how the estimator reweights observations in response to moment residuals, thereby affecting robustness and sensitivity to sampling variation. At the same time, the moment-based framework preserves structural transparency: the economic content remains encoded in explicitly stated moment restrictions, ensuring interpretability of the resulting estimates. Treating $\gamma$ as a hyperparameter thus provides a systematic way to tailor finite-sample behavior while maintaining the discipline of the underlying econometric model.

An important direction for future research is to study CRPD estimation and the role of $\gamma$ under moment misspecification. Under misspecification, the first-order equivalence across members of the CRPD class generally breaks down (see, e.g., Schennach2007), so $\gamma$ may influence not only higher-order behavior but also the leading asymptotic approximation. Characterizing the pseudo-true targets, robustness properties, and optimal data-driven selection of $\gamma$ in misspecified environments, and their implications for inference, would further strengthen the ML interpretation of CRPD.

Acknowledgements

We thank Xiaofeng Shao (Washington University in St. Louis) and the audience at the Econometrics Brown Bag Seminar at the University of Illinois Urbana–Champaign for helpful comments. All remaining errors are our own.