Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
86,340 characters · 12 sections · 87 citation commands
From dense grids to valid inference: Accounting for regularization bias in nonparametric random coefficient models
{\setstretch{.8}
}
\hrule { $^a$ University of Groningen, Nettelbosje 2, 9747 AE Groningen, The Netherlands.}\\ { Correspondence to: [email removed]. (M.D. Pen).}
\onehalfspacing
Random-coefficients (RC) models are widely used in economics to accommodate heterogeneous preferences across economic agents. In these models, individual-level parameters vary according to an unknown distribution that researchers aim to recover from the data. The objects of central interest in applications are typically functionals of the RC distribution, such as willingness-to-pay measures, demand elasticities, or predicted outcomes under counterfactual policies. This paper develops bias-aware inference for such functionals when the RC distribution is estimated nonparametrically using the penalized fixed-grid estimator of heiss2022. We characterize the sampling distribution of the penalized plug-in estimator and construct confidence intervals that incorporate a feasible upper bound for the bias induced by regularization. These inferential needs arise in applications across environmental economics faure2021, public economics simonov2022, industrial organization dubois2020, and health economics heiss2021.
Empirical implementations of RC models commonly impose a parametric distribution on unobserved preference heterogeneity, most often a normal or log-normal distribution. Such assumptions are convenient and supported by standard software, but they restrict the shape and tails of the heterogeneity distribution. When these restrictions are misspecified, they can substantially affect the economic quantities computed from the model. For example, miravete2022 show how parametric restrictions on heterogeneity constrain the curvature of demand functions and the range of estimated elasticities, while compiani2022 emphasizes that counterfactual predictions can be sensitive to the assumed distribution of heterogeneous preferences. These concerns motivate the use of nonparametric estimators, which estimate the RC distribution under weaker ex ante restrictions and accommodate a wider range of feasible distributional shapes.
A computationally tractable nonparametric approach is the fixed-grid estimator introduced by bajari2007 and further developed by fox2011. It approximates the unknown RC distribution by a discrete distribution over prespecified support points and estimates the probability mass assigned to each point. fox2016 provide convergence results for the estimated distribution, and the approach has been used in empirical applications such as nevo2016 and blundell2020. Other flexible approaches include the likelihood-based estimators of train2008 and train2016, based on discrete mixtures and approximations using polynomials, splines, or step functions, and the Bayesian hierarchical models of rossi2005 and burda2008, based on countably infinite mixtures of normals. Bayesian methods provide posterior inference for features of the RC distribution, although their implementation typically relies on Markov chain Monte Carlo and can be computationally demanding in large discrete-choice applications. Our analysis focuses on the distinct inferential issues created by regularization in the fixed-grid approach.
An important trade-off researchers face when applying the fixed-grid estimator is the density of the grid. Within an appropriately chosen support domain, a finer grid can reduce the approximation error generated by representing a flexible distribution with finitely many support points. On a dense grid, however, choice probabilities evaluated at nearby support points become difficult to distinguish, producing severe multicollinearity and unstable estimates of the probability weights. heiss2022 analyze support recovery in this setting and propose adding an $\ell_2$-penalty to stabilize estimation. The penalty can make substantially denser grids feasible, but it also introduces regularization bias into plug-in estimates of the functionals of interest. Relatedly, escanciano2023 shows that some features of unobserved heterogeneity, including quantiles in relevant settings, can be irregularly identified. Thus, a coarse grid can leave non-negligible approximation bias, whereas a dense grid may call for regularization whose inferential consequences cannot be ignored.
This trade-off has a direct implication for inference. Because the penalty changes the criterion, the penalized plug-in estimator is centered around the functional evaluated at a penalized pseudo-true parameter rather than the corresponding unpenalized fixed-grid target. An interval that accounts only for sampling uncertainty, therefore, need not cover the latter target. Some prominent fixed-grid applications use bootstrap-based inference for functionals obtained from counterfactual simulations nevo2016, blundell2020. In the penalized setting studied here, however, resampling the estimator without an explicit bias allowance need not account for the shift between the penalized and unpenalized targets. More generally, fang2018 show that standard bootstrap procedures can fail for directionally differentiable transformations. A related literature develops inference for estimators subject to inequality or boundary constraints andrews1999estimation, fang2023inference, fan2023wald. Those methods address important features of the unpenalized fixed-grid problem and can apply to linear functionals in some settings, but they do not directly address the additional displacement induced by the $\ell_2$-penalty.
We address this problem by adapting the sieve-inference framework of chen2014sieve and chen2015sieve to the penalized fixed-grid estimator of heiss2022. The support-recovery result in heiss2022 motivates conducting the local analysis on the active set of nonzero probability weights. We complement that result with active-set stability conditions under which the penalized and unpenalized fixed-grid targets share the same active face of the simplex and local perturbations remain on that face with probability approaching one. Under these conditions, binding nonnegativity constraints do not affect the first-order local expansion through varying zero-weight directions. Within this framework, we derive the asymptotic distribution of the plug-in estimator around the functional evaluated at the penalized pseudo-true parameter, use a sieve Riesz representer to estimate sampling uncertainty, and construct a feasible upper bound on the difference between the penalized and unpenalized fixed-grid functionals.
Combining these components yields confidence intervals that are asymptotically conservative for the fixed-grid best approximation target, under the stated conditions, even when regularization bias is not negligible relative to sampling uncertainty. The general construction applies to a class of smooth scalar functionals that includes mean random coefficients, willingness-to-pay measures, elasticities, welfare changes, and counterfactual choice probabilities. It does not require the unpenalized design matrix to be nonsingular and can accommodate irregular functionals whose corresponding Riesz representer's norm may diverge. We also provide a simpler direct bootstrap procedure for regular functionals. In contrast to approaches that remove first-order regularization effects through debiasing or orthogonalization chernozhukov2018, chernozhukov2022, we retain the penalized estimator and incorporate a bound on its bias into the interval. The contribution is therefore a bias-aware confidence-interval construction for an existing estimator: it makes explicit both the stabilizing role of regularization and its cost for inference.
Monte Carlo simulations illustrate the distinct roles of grid approximation and regularization. In the designs examined, a coarse grid leaves consequential approximation bias over the sample sizes considered and produces substantial undercoverage, while our proposed method performs better along a denser grid.
In reported designs, the bias-aware intervals are only moderately wider or even narrower on average than the corresponding intervals based on the unpenalized fixed-grid estimator. This finite-design comparison illustrates that a reduction in sampling variance can partly offset the additional allowance for regularization bias. In an application to the travel-mode data of meijer2006, the nonparametric specifications place more probability mass at large absolute values in the Value-of-Time (VoT) distribution and imply larger mean VoT magnitudes than the parametric log-normal benchmark. The bias-aware intervals remain informative and, in this application, are only moderately wider than the corresponding unpenalized fixed-grid intervals. These empirical differences illustrate the sensitivity of the reported functionals to distributional assumptions.
The remainder of the paper is organized as follows. Section (ref) introduces the RC logit model, the functionals considered in the analysis, and the penalized and unpenalized fixed-grid estimators. Section (ref) derives the asymptotic theory and constructs the confidence-conservative intervals. Section (ref) reports the Monte Carlo study, and Section (ref) presents the empirical application. Section (ref) concludes.
This section reviews the RC model and common functionals frequently considered in empirical applications, which are the primary objects of inference in this paper. It then revisits the nonparametric fixed-grid estimator of fox2011 and the regularized extension of heiss2022, which are used for plug-in estimators for these functionals. Finally, Subsection (ref) discusses the trade-off between approximation and regularization bias, thereby motivating the confidence-conservative inference procedure developed in the following sections.
Suppose that each decision maker $i$ faces a finite set of mutually exclusive alternatives $j=0, \ldots, J$, where $j=0$ denotes a possible outside option. For each decision maker $i$ and inside alternative $j=1, \ldots, J$, the researcher observes $y_{i,j}$, where $y_{i,j}=1$ if decision maker $i$ chooses alternative $j$ and $0$ otherwise, together with a $K$-dimensional vector of covariates $x_{i,j}$.\footnote{We focus on cross-sectional data, while extensions of the nonparametric fixed-grid estimator to panel data are discussed in fox2011.}
In line with the theoretical results derived in Section (ref), we consider a model that allows for a combination of fixed and random coefficients.\footnote{Although assuming some coefficients to be homogeneous across individuals reduces the model's flexibility, we consider specifications with both fixed and random coefficients because the fixed-grid estimator presented below suffers from the curse of dimensionality, making it computationally feasible only for models with relatively few random coefficients in practice.} Let $\delta_0$ denote the $d_\delta$-dimensional vector of true fixed coefficients and $\beta$ the $d_\beta$-dimensional vector of random coefficients with true distribution $F_0$. Furthermore, let $x_i=(x_{i,1}',\ldots,x_{i,J}')'$ denote the matrix of stacked covariates across inside alternatives. Integrating over the distribution of the RCs yields the unconditional choice probability
where $g_j(\cdot)$ is a known function determined by the assumed discrete choice model. An example widely applied in the literature is the RC logit model, in which $g_j(\cdot)$ is given by the logit kernel, \[ g_j(x_i, \delta_0, \beta) = \frac{\exp(x_{i,j,1}'\delta_0 + x_{i,j,2}'\beta)} {1+\sum_{l=1}^{J}\exp(x_{i,l,1}'\delta_0 + x_{i,l,2}'\beta)}, \] where $x_{i,j, 1}$ and $x_{i,j,2}$ denote the subvectors of product characteristics associated with the fixed and random coefficients, respectively.\footnote{For a detailed treatment of RC logit models, see train2009.}
The researcher's objective is to estimate the unknown parameters $\delta_0$ and $F_0$ from the observed choice data. In empirical applications, however, researchers are typically not directly interested in these estimates, but rather in functionals of the unknown model parameters, such as willingness-to-pay measures, own- and cross-price elasticities, or counterfactual predictions. A broad class of economically meaningful quantities of interest can be written as average functionals of the form
where $s(\cdot)$ is a known measurable function and $G$ denotes the population distribution of $x$. In several examples, the outer integral with respect to $G$ drops out. Otherwise, we treat $G$ as known and ignore the additional estimation error arising from estimating $G$.
A common approach in applications of RC models is to restrict $F_0$ to a prespecified parametric family, such as the normal or log-normal family. Although computationally convenient, such a restriction may misrepresent the underlying preference heterogeneity and thereby distort estimates of the functionals of interest.
To avoid parametric shape restrictions, bajari2007 and fox2011 propose a computationally tractable nonparametric estimator of the RC distribution. Let $R_N$ denote the number of fixed grid points, and let $\mathcal{B}_{R_N}\equiv\{\beta_1,\ldots,\beta_{R_N}\}\subset\mathbb{R}^{d_\beta}$ denote the grid specified by the researcher before estimation. The probability mass $\theta_r$ associated with grid point $\beta_r$, $r=1,\ldots,R_N$, is estimated from the data. The estimator approximates $F_0$ by the discrete distribution $F(\cdot)=\sum_{r=1}^{R_N}\theta_r\mathbf{1}\{\beta_r\leq\cdot\}$.\footnote{Unless otherwise stated, vector inequalities such as “$\leq$” are understood componentwise.} As $R_N$ increases with $N$, a suitably chosen grid can provide increasingly accurate approximations to $F_0$.\footnote{{fox2016 show that distributions supported on the grid can approximate any RC distribution arbitrarily closely in the weak topology as $R_N\rightarrow\infty$ and the grid becomes dense in the coefficient space.}}
To estimate the probability weights, consider the linear probability approximation that arises from replacing $F_0$ by a discrete distribution supported on the finite grid $\mathcal{B}_{R_N}$ heiss2022
where $\epsilon_{i,j}=y_{i,j}-P_j(x_i)$ is the linear probability error. Building on (ref), heiss2022 propose the penalized estimator (referred to as the ENet estimator in later discussions)
where $\theta=(\theta_1,\ldots,\theta_{R_N})'$, $y_i=(y_{i,1},\ldots,y_{i,J})'$, and $g(x_i, \delta,\beta_r)=(g_1(x_i,\delta, \beta_r),\ldots,g_J(x_i,\delta, \beta_r))'$ for all $r=1,\ldots,R_N$. The $J\times J$ matrix $\Sigma(x_i)^{-1}$ is a positive-definite weighting matrix (e.g., heiss2022 use the identity matrix).\footnote{For the construction of a variance-based weighting matrix, see Section 5 of fox2011.} The tuning parameter $\mu_N\geq0$ controls the strength of the $\ell_2$-penalty. Setting $\mu_N=0$ recovers the estimator of fox2011 (referred to as the FKRB estimator in later discussions), whereas larger values imply stronger regularization. heiss2022 recommend selecting $\mu_N$ by cross-validation. For models with both fixed and random coefficients, heiss2022 implement (ref) with an iterative algorithm that alternates between updating the probability weights through constrained least squares and updating the fixed coefficients through a weighted conditional logit likelihood. We use this algorithm in the empirical application; Appendix B of heiss2022 provides details.
The estimated RC distribution is thus obtained from the estimated probability weights $\widehat{\theta}$ as
Given the fixed-grid estimator $(\widehat{\delta}, \widehat{F})$ and the object of interest in the form of Equation ((ref)), Section (ref) derives a bias-aware inference approach based on the plug-in estimator $\tau(\widehat{\delta}, \widehat{F})$.
Although the estimator described in the previous subsection is computationally attractive, its statistical properties are governed by a fundamental trade-off between approximation bias and regularization bias. On the one hand, the grid should be sufficiently dense so that the discrete approximation accurately represents the unknown RC distribution. On the other hand, increasing the number of grid points makes the estimation problem increasingly ill-conditioned, requiring regularization.
As the grid becomes denser, the resulting design matrix becomes increasingly collinear, because, e.g., neighboring grid points generate very similar choice probabilities and thus similar regressors in the linear probability model in Equation ((ref)), rendering the constrained least squares problem ill-conditioned. This issue has been documented for the fixed-grid estimator by fox2016, nevo2016, and heiss2022, and more generally for structural models with nonparametric unobserved heterogeneity by escanciano2023.
A natural way to address this issue is through regularization. In the context of the fixed-grid estimator, regularization permits the use of substantially denser grids, thereby reducing sieve approximation error and improving estimation of the underlying RC distribution. Specifically, heiss2022 augment the estimator of fox2011 with the $\ell_2$ penalty in Equation ((ref)), while escanciano2023 develop a related regularized estimator.
However, regularization comes at the cost of inducing regularization bias. Consequently, researchers face an inherent trade-off. A coarse grid reduces ill-conditioning but may induce substantial approximation bias because the support is too sparse to accurately represent the true distribution. Conversely, a dense grid reduces approximation bias but requires stronger regularization, thereby increasing regularization bias. The following section develops a valid inference for average functionals of the form in Equation ((ref)) that accounts for the regularization bias of the penalized fixed-grid estimator in Equation ((ref)).
This section derives conservative confidence intervals for plug-in estimators of functionals, including the average functionals in (ref). To accommodate the near-singular design matrices that arise with dense grids nevo2016,fox2016, we use the penalized fixed-grid estimator of heiss2022. The main result establishes asymptotic normality around a penalized pseudo-true parameter. We then bound the regularization bias induced by the $\ell_2$-penalty in (ref) and incorporate that bound into the confidence interval.
The analysis builds on the sieve-inference frameworks of chen2014sieve and chen2015sieve, adapted to the fixed-grid estimator of fox2011 and the penalized extension of heiss2022. The central idea is to linearize the functional locally around the penalized pseudo-true parameter and characterize its first-order term with a sieve Riesz representer. This yields an asymptotically normal representation, an explicit asymptotic variance, and a tractable regularization-bias bound.
\@startsection{paragraph}{4}{\z@} {3.25ex \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Notation.} Let $\operatorname{clsp}(A)$ denote the closed linear span of $A$ under the norm in (ref). For positive deterministic sequences $(a_N)$ and $(b_N)$, write $a_N\lesssim b_N$ if $a_N/b_N\leq C$ for some finite constant $C$ and all sufficiently large $N$; write $a_N\asymp b_N$ if both $a_N\lesssim b_N$ and $b_N\lesssim a_N$. For positive random sequences $(U_N)$ and positive deterministic sequences $(V_N)$, write $U_N=O_p(V_N)$ if $\lim_{M\to\infty}\limsup_{N\to\infty}P(U_N/V_N>M)=0$, and write $U_N=o_p(V_N)$ if $\lim_{N\to\infty}P(U_N/V_N>\varepsilon)=0$ for every $\varepsilon>0$. We omit the subscript “$p$” for nonrandom sequences. Let $|A|$ denote the cardinality of a set $A$, and let $\mathbb{R}_+=[0,\infty)$. For a vector $a\in\mathbb{R}^K$, let $a'$ denote its transpose and $\|a\|_e=(a'a)^{1/2}$. For a matrix $A$, let $\|A\|_e=\{\operatorname{tr}(A'A)\}^{1/2}$ denote the Frobenius norm. For a map $\phi:\boldsymbol{B}\to\boldsymbol{H}$ between Banach spaces, define \[ D\phi(\alpha)[v]\equiv\left.\frac{\partial}{\partial t}\phi(\alpha+tv)\right|_{t=0}. \] Finally, let $E(\cdot)$ denote expectation/expected value, and $\bar{E}_N\{h(z)\}\equiv N^{-1}\sum_{i=1}^N\{h(z_i)-E[h(z_i)]\}$ denote the centered empirical process indexed by $h$ for a sample $z=(z_1,\cdots, z_N)$.
Let $\mathcal{B}\subset\mathbb{R}^{d_\beta}$ be the compact support of the RCs, and let $\mathcal{F}$ be the class of distributions supported on $\mathcal{B}$. The parameter space is $\mathcal{A}\equiv\mathcal{C}\times\mathcal{F}$, with \[ \alpha=(\delta,F),\qquad\delta\in\mathcal{C}\subset\mathbb{R}^{d_\delta},\qquad F\in\mathcal{F}. \] Let $\alpha_0=(\delta_0,F_0)$ denote the true parameter.
For each sample size $N$, let $\mathcal{B}_N=\{\beta_1,\ldots,\beta_{R_N}\}\subset\mathcal{B}$ denote the fixed grid of $R_N$ support points. Define the simplex \[ \mathcal{D}_N=\{\theta\in[0,1]^{R_N}:\iota_{R_N}'\theta=1\}, \] where $\iota_k$ denotes the $k$-vector of ones. For $\theta\in\mathcal{D}_N$, let \[ F_\theta(\cdot) =\theta'\mathbf{1}_N(\cdot) =\sum_{r=1}^{R_N}\theta_r\mathbf{1}\{\beta_r\leq\cdot\}, \] where $\mathbf{1}_N(\cdot)$ stacks $\{\mathbf{1}(\beta_r\leq\cdot):r=1,\ldots,R_N\}$. The sieve parameter space is \[ \mathcal{A}_N=\{(\delta,F_\theta):\delta\in\mathcal{C},\ \theta\in\mathcal{D}_N\}. \]
For each $\alpha\in\mathcal A_N$, we write $\alpha=(\delta_\alpha,F_{\theta_\alpha})$ to distinguish from the corresponding components of the sieve Reisz representer introduced below. The notation $F_{\theta_\alpha}$ emphasizes that the distribution component of $\alpha$ is determined by the probability weights $\theta_\alpha$, which is subject to the simplex constraint, assigned to the fixed grid points.
For $\alpha=(\delta_\alpha,F_{\theta_\alpha})\in\mathcal{A}_N$, define the strong norm \[ \|\alpha\|_s\equiv\|\delta_\alpha\|_e+\|\theta_\alpha\|_e, \]
and define the unpenalized weighted least-squares criterion \[ \ell_N^0(z_i,\alpha)\equiv\frac12\rho(z_i,\alpha)'\Sigma(x_i)^{-1}\rho(z_i,\alpha), \] with $\rho(z_i,\alpha)\equiv y_i-P(x_i,\alpha)$, where $z_i=(x_i,y_i)$ are i.i.d., $P(x_i,\alpha)$ is the model-implied vector of choice probabilities, and $\Sigma(x)$ is positive definite with eigenvalues uniformly bounded above and away from zero. Moreover, let \[ Q_N^0(\alpha)\equiv E[\ell_N^0(z,\alpha)], \qquad \ell_N^\mu(z,\alpha)\equiv\ell_N^0(z,\alpha)+\frac{\mu_N}{2}\|\theta_\alpha\|_e^2, \qquad Q_N^\mu(\alpha)\equiv E[\ell_N^\mu(z,\alpha)]. \] define the unpenalized and penalized population objective function, respectively. In comparison to (ref), we omit the constant factor $1/J$, which has no effect once the penalty parameter is rescaled accordingly.
When $\mu_N=0$, the population criterion may have multiple minimizers because a dense grid can produce a singular design matrix. When $\mu_N>0$, we assume that the penalized population criterion has a unique minimizer.
Let $\alpha_N^0$ denote a best fixed-grid approximation to the population target,\footnote{When $F_0$ has a density, one common benchmark assigns to each grid point the probability of an associated grid cell. If $F_0$ is discrete and its mass points are included in the grid, the approximation error can vanish.} and let $\alpha_N^\mu$ denote the penalized pseudo-true parameter: \[ \alpha_N^0\in\arg\min_{\alpha\in\mathcal{A}_N}Q_N^0(\alpha), \qquad \alpha_N^\mu=\arg\min_{\alpha\in\mathcal{A}_N}Q_N^\mu(\alpha). \]
The sample analogue of $Q_N^\mu(\alpha)$ is \[ \widehat Q_N^\mu(\alpha)=\frac1N\sum_{i=1}^N\ell_N^\mu(z_i,\alpha) \] and the corresponding estimator is defined by \[ \widehat{\alpha}_N^\mu\in\arg\min_{\alpha\in\mathcal{A}_N}\widehat Q_N^\mu(\alpha). \qquad \] The unpenalized sample objective and estimator are obtained by setting $\mu_N=0$ and are denoted by $\widehat Q_N^0$ and $\widehat{\alpha}_N^0$, respectively.
Since some probability weights may be zero at the penalized optimum, we define the following restricted sieve spaces. For any $S\subseteq\{1,\ldots,R_N\}$, let \[ \mathcal{A}_N(S)\equiv\{(\delta,F_{\theta})\in\mathcal{A}_N:\theta_r=0\text{ for }r\notin S\}. \] The active set of the penalized pseudo-true parameter is \[ S_N=\{r:\theta_{\alpha_N^\mu,r}>0\}, \] where $\theta_{\alpha_N^\mu,r}$ is the $r$th entry of penalized probability weights, $\theta_{\alpha_N^\mu}$ such that $\alpha_N^\mu=(\delta_{\alpha_N^\mu}, F_{\theta_{\alpha_N^\mu}})$.
The subsequent analysis is local to this active set. For a positive sequence $\xi_N$, define the local neighborhood
The positive sequence $\xi_N$ typically converges to zero. Lemma (ref) in Appendix (ref) provides one admissible rate. To characterize local perturbations around the penalized pseudo-true parameter, define the local direction space \[ \mathcal{V}_N^\mu \equiv \operatorname{clsp}\bigl(\mathcal{N}_N^\mu-\{\alpha_N^\mu\}\bigr). \]
Assumption (ref) requires the best approximation and the penalized target to share an active face of the simplex and ensures that the estimator admits the local perturbations used below. Theorem 1 in heiss2022 establishes support recovery by the penalized sample estimator relative to the best fixed-grid weights under sufficient conditions. Lemma (ref) in Appendix (ref) gives related sufficient conditions for Assumption (ref).
These conditions permit $\mu_N=0$ when the unpenalized population criterion has sufficient local curvature. When the design is nearly singular, a positive $\ell_2$-penalty provides the curvature required for local identification, albeit around a biased target. The simulation results below illustrate this trade-off.
Following the local quadratic expansion used in chen2014sieve, consider directions $v_1,v_2\in\mathcal{V}_N^\mu$ and define the inner product
The induced norm is
Under $\mu_N>0$ and the maintained regularity conditions, this norm remains well defined even when the design matrix is near singular. We also assume that differentiation and expectation can be interchanged through the second order, so that
where
Equipped with the inner product in Equation ((ref)), $\mathcal{V}_N^\mu$ is a finite-dimensional Hilbert space. Its elements respect the active-set restriction: directions outside $S_N$ have zero grid-weight components. Thus, any $v\in\mathcal{V}_N^\mu$ can be represented as \[ v=(\delta_v,\theta_v'\mathbf{1}_{S_N}(\cdot)), \] where $\delta_v \in \mathbb{R}^{d_\delta}, \theta_v\in\mathbb{R}^{|S_N|}$ and $\mathbf{1}_{S_N}(\cdot)$ stacks the indicator functions $\{\mathbf{1}(\beta_r\le\cdot)\}_{r\in S_N}$. Moreover, for any $\alpha\in\mathcal{N}_N^\mu$ and $v\in\mathcal{V}_N^\mu$, sufficiently small local perturbations $\alpha+tv$ remain in $\mathcal{N}_N^\mu$ whenever the path stays in the relative interior. Thus, by restricting to the interior point within $\mathcal{N}_N^\mu$, we avoid the boundary issue raised in constrained optimizations.
Having characterized the local geometry of the penalized estimator, we now turn to the scalar functional $\tau(\alpha)$, which is the object of primary interest. For a scalar functional $\tau:\mathcal{A}_N\to\mathbb{R}$, define the directional derivative at the penalized target in the direction $v\in\mathcal{V}_N^\mu$ by \[ D\tau(\alpha_N^\mu)[v]\equiv\left.\frac{\partial}{\partial t}\tau\!\left(\alpha_N^\mu+tv\right)\right|_{t=0}. \] If $D\tau(\alpha_N^\mu)[\cdot]$ is linear on $\mathcal{V}_N^\mu$, then finite-dimensionality of $\mathcal{V}_N^\mu$ implies the existence of a sieve Riesz representer $v_N^\tau\in\mathcal{V}_N^\mu$ such that
Let $\gamma=(\delta',\theta')'\in\mathbb{R}^{d_{\bm{\delta}}}\times\mathbb{R}^{|S_N|}$, and define
The Riesz representer solving Equation ((ref)) is \[ v_N^\tau=(\delta^\tau,\theta^{\tau'}\mathbf{1}_{S_N}(\cdot)), \qquad(\delta^{\tau'},\theta^{\tau'})'=({H}_{N}^\mu)^{-1}{D}_{N}^\mu, \]
whose derivation follows the same steps as the solution of Equation ((ref)) and is therefore omitted. Example (ref) gives the corresponding expression for the mean WTP functional introduced in Example (ref).
\\
For any $v=(\delta_v,\theta_v'\mathbf{1}_{S_N}(\cdot))\in\mathcal V_N^\mu$, define the regularized score norm \[ \|v\|_{sd,\mu}^2\equiv\operatorname{Var}\!\left(D\ell_N^\mu(z,\alpha_N^\mu)[v]\right)+\mu_N\|\theta_v\|_e^2. \] For the Riesz representer, define the score variance \[ \sigma_N^2\equiv\operatorname{Var}\!\left(D\ell_N^\mu(z,\alpha_N^\mu)[v_N^\tau]\right). \] and, without loss of generality, we only consider $\sigma_N>0$.
Assumption (ref) requires the local curvature of the penalized population criterion and the sampling variation of its first-order derivative (plus the $\ell_2$-penalty term) to be uniformly comparable over the sieve direction space. It is a regularized analog of Assumption 3.2 in chen2014sieve. The additional deterministic $\ell_2$ component in $\|v\|_{sd,\mu}$ is induced by the $\ell_2$-penalty term.
To establish asymptotic normality around the penalized pseudo-true value and derive asymptotically conservative confidence intervals for the best fixed-grid target, we impose local smoothness conditions on the criterion function and the functional. These conditions parallel those in chen2014sieve, but here they are centered at $\alpha_N^\mu$ for each fixed grid and penalty. Their formal statements are collected in the appendix.
Theorem (ref) establishes asymptotic normality of the penalized plug-in estimator around the penalized pseudo-true value $\tau(\alpha_N^\mu)$. The following corollary combines this result with a bound on the regularization bias to obtain an asymptotically conservative confidence interval for the unpenalized fixed-grid target $\tau(\alpha_N^0)$.
Corollary (ref) separates sampling uncertainty, measured by $\sigma_N$, from regularization bias, bounded by $\|v_N^\tau\|_\mu\|\alpha_N^0-\alpha_N^\mu\|_\mu$. A larger $\mu_N$ may reduce sampling variance but generally increases regularization bias. For a fixed penalty, a denser grid may reduce approximation bias while increasing ill-conditioning and sampling uncertainty. The coverage statement concerns the best fixed-grid target $\tau(\alpha_N^0)$. It extends to $\tau(\alpha_0)$ under an additional approximation condition such as negligible approximation bias $|\tau(\alpha_N^0)-\tau(\alpha_0)|=o(N^{-1/2}\sigma_N)$. The theorem also allows irregular functionals, for which $\sigma_N$ and $\|v_N^\tau\|_\mu$ may diverge. The following subsections construct feasible estimators of the sampling and bias components.
We first estimate the sieve Riesz representer by its empirical counterpart and then use it to construct feasible sampling-variance and bias terms. Let \[ \widehat S_N=\{r:\theta_{\widehat{\alpha}_N^\mu,r}>0\}, \qquad \widehat{\mathcal{V}}_N^\mu= \operatorname{clsp}\bigl(\mathcal{A}_N(\widehat S_N)-\{\widehat\alpha_N^\mu\}\bigr). \] For $v_1,v_2\in\widehat{\mathcal{V}}_N^\mu$, define the empirical inner product \[ \langle v_1,v_2\rangle_N\equiv\frac1N\sum_{i=1}^N D^2\ell_N^\mu(z_i,\widehat\alpha_N^\mu)[v_1,v_2], \] with induced norm $\|\cdot\|_N$. The empirical directional derivative is \[ D\tau(\widehat{\alpha}_N^\mu)[v]\equiv\left.\frac{\partial}{\partial t}\tau\!\left(\widehat{\alpha}_N^\mu+tv\right)\right|_{t=0}. \] The empirical sieve Riesz representer $\widehat v_N^\tau$ solves
Lemma (ref) gives its closed form.
The difference between $\alpha_N^0$ and $\alpha_N^\mu$ is driven by the $\ell_2$ penalty applied to the probability weights. Lemma (ref) bounds this difference in the pseudo-norm $\|\cdot\|_\mu$. The corresponding difference in the strong norm $\|\cdot\|_s$ need not be small.
If $\epsilon_N^*\xi_N=o(N^{-1/2})$ (see, for example, Assumption 5.2 of chen2014sieve), Corollary (ref) and Lemmas (ref)--(ref) yield the feasible half-length \[ \widehat c_N(q) =\Phi^{-1}(1-q/2)N^{-1/2}\widehat\sigma_N (\sqrt{\mu_N}+N^{-1/2}) (1+N^{-1/2}\xi_N^{-1}) \|\widehat v_N^\tau\|_\mu. \] Thus, $[\tau(\widehat{\alpha}_N^\mu)-\widehat c_N(q), \tau(\widehat{\alpha}_N^\mu)+\widehat c_N(q)]$ is a feasible conservative interval; Corollary (ref) gives the formal statement. It does not require the unpenalized design matrix to be nonsingular and can accommodate irregular functionals. Its implementation, however, requires the estimated sieve Riesz representer and $\widehat\sigma_N^2$. The next subsection gives a simpler bootstrap procedure for regular functionals.
This subsection provides a direct bootstrap procedure for regular functionals. We call $\tau(\cdot)$ regular at $\alpha_N^\mu$ when the Riesz representer of $D\tau(\alpha_N^\mu)[\cdot]$ is uniformly bounded, as specified in Assumption (ref). Boundedness permits a nonstudentized bootstrap approximation without the need to compute any variance estimator (otherwise, the nonstudentized bootstrapped distribution may not serve as a good approximation, see also, e.g., Theorem 5.2.(2) in chen2015sieve).
Assumption (ref) and Lemma (ref) imply that any $b_N>0$ satisfying $\sqrt{\mu_N}=o(b_N)$ bounds $|\tau(\alpha_N^0)-\tau(\alpha_N^\mu)|$ up to an $o(N^{-1/2})$ remainder under the conditions stated below. A sharper option is available when the functional places little weight on directions whose curvature is generated primarily by the penalty, and we discuss one such case in Assumption (ref).
Let $v_N^\tau =(\delta^\tau,\theta^{\tau'}\mathbf{1}_{S_N}(\cdot))$ be the corresponding Riesz representer of $D\tau(\alpha_N^\mu)[\cdot]$, $\gamma_\tau=(\delta^{\tau'},\theta^{\tau'})'=(H_N^\mu)^{-1}D_N^\mu,$ and $H_N^\mu=\sum_iw_{N,i}h_{N,i}h_{N,i}'$ be the spectral decomposition, with eigenvalues $w_{N,i}$ in descending order. The next assumption requires $D_N^\mu$ to place negligible weight on the lower spectral band $H_N^\mu$, that is, on directions whose curvature is close to the $\ell_2$-induced lower edge $\mu_N$.
In the well-conditioned case, all eigenvalues of $H_N^\mu$ are uniformly bounded away from zero. The lower spectral band is then empty for all sufficiently large $N$ whenever $\mu_N=o(1)$, so Assumption (ref) holds automatically. More generally, the assumption ensures that the sampling variance $\sigma_N^2$ is sufficiently large relative to the $\ell_2$ component $\mu_N\|\theta^\tau\|_e^2$, leading to the following result.
Lemma (ref) bounds the regularization bias in the functional up to an asymptotically negligible remainder by $\sqrt{2\mu_N}$ times its asymptotic standard deviation. Without calculating the Riesz representer case by case, we show below that $\sigma_N$ can be consistently estimated by bootstrap under conditions stated in Corollary (ref), and the result also yields a feasible bias adjustment without requiring direct estimation of the distance between the penalized and unpenalized pseudo-true parameters.
We use the multinomial bootstrap to estimate $\sigma_N$. In each bootstrap iteration, observations are reweighted by random weights $(\omega_{n,N})_{n=1}^N$ drawn from the multinomial distribution such that \[ (\omega_{n,N})_{n=1}^N \sim \operatorname{Multinomial} \left(N; N^{-1},\ldots,N^{-1}\right). \] See, for example, Assumption Boot.2 in chen2015sieve. For a realization $\omega=(\omega_{1,N},\ldots,\omega_{N,N})'$, define \[ \widehat Q_N^{\mu,B}(\alpha)= \frac1N\sum_{n=1}^N \omega_{n,N}\ell^0_N(z_n,\alpha) +\frac{\mu_N}{2}\|\theta_\alpha\|_e^2. \] Let $\widehat\alpha_N^{\mu,B}$ minimize $\widehat Q_N^{\mu,B}$ over $\mathcal{A}_N$, and let $s^B$ denote the conditional bootstrap standard deviation of $\sqrt{N}\{\tau(\widehat{\alpha}_N^{\mu,B}) -\tau(\widehat{\alpha}_N^\mu)\}$.
Next, we demonstrate the performance of the proposed intervals in the above corollary in the following two sections.
We conduct a Monte Carlo study to evaluate the finite-sample performance of the proposed confidence-conservative intervals for different average functionals. The data are generated from an RC logit model in which each individual $i$ chooses among $J=3$ mutually exclusive alternatives and an outside option. We use a linear utility function $u_{i,j}=x_{i,j,1}\beta_{i,1}+x_{i,j,2}\beta_{i,2}+\varepsilon_{i,j}$ to model the utility individual $i$ derives from alternative $j$, where $x_{i,j,k}$ are observed covariates, $\beta_i=(\beta_{i,1},\beta_{i,2})'$ are individual-specific RCs whose distribution we aim to estimate from the data, and $\varepsilon_{i,j}$ are unobserved error terms. Under the assumption that $\varepsilon_{i,j}$ are i.i.d.\ type-I extreme value, the model implies multinomial logit choice probabilities conditional on the individual-specific coefficients. We draw the covariates $x_{i,j,k}$ independently from $\mathcal{U}[0,3]$ across individuals, alternatives, and attributes $k\in\{1,2\}$.
To assess whether the proposed intervals attain nominal coverage in settings where standard parametric approaches fail, the RCs are drawn from a two-component mixture of bivariate normal distributions,
We estimate this RC distribution using both the unpenalized fixed-grid estimator (FKRB) and the penalized fixed-grid estimator (ENet). We consider sample sizes $N\in\{1000, \, 10000\}$ and repeat each design $M=1000$ times. Because the true RC distribution is continuous and bimodal, it requires a relatively dense grid to accurately approximate the underlying distribution. We therefore estimate the model using a uniform grid with $R_N\in\{25,49,81,289\}$ grid points over the domain $[-8,-0.1]\times[0.1,8]$ which covers the support of the true distribution with coverage probability close to one.\footnote{As a robustness check, we also consider a substantially larger grid domain, $[-10,0]\times[0,10]$, placing many grid points in regions with near-zero probability. This is close to the special scenario, placing grids outside the support, used to verify Assumption (ref) in Appendix (ref). The corresponding results are reported in Appendix (ref). This additional robustness check leads to similar conclusions as the baseline specification: the ENet estimator continues to outperform FKRB, and denser grids generally improve finite-sample performance.} As the grid becomes denser, however, neighboring support points generate increasingly similar choice probabilities, leading to severe multicollinearity in the fixed-grid regression. This design, therefore, provides a natural setting to evaluate whether regularization improves both estimation and inference. Figure (ref) in Appendix (ref) illustrates the grid for $R_N=25$ and $R_N=81$ support points together with contour lines of the \st{true} mixture distribution.
For the fixed-grid estimators, we use the weighting matrix as suggested by Section 5 of fox2011. For the ENet estimator, the regularization parameter is selected separately in each estimation step by five-fold cross-validation over 101 candidate values generated with the R package glmnet, including $\mu_N=0$ to allow for the unpenalized FKRB estimator.
For each Monte Carlo replication, we compute plug-in estimates of two functionals: the mean of the first RC, \[ \tau_{\beta_1,0} = \int \beta_1\,dF_0(\beta), \] and the own-price elasticity evaluated at the mean covariate vector $\bar{x}=(1.5,1.5)$, \[ \tau_{\eta,0} = \int \left\{ \frac{\bar{x}_{j,k}}{P_j(\bar{x})} \frac{\partial g_j(\bar{x},\beta)}{\partial x_{j,k}} \right\} dF_0(\beta). \] Confidence intervals for FKRB and ENet estimators are calculated using the results in Corollary (ref) with $\mu_N=0$ for the FKRB estimator. We use a multinomial block bootstrap with $B=500$ bootstrap replications to estimate $s^B$.
As a benchmark, we additionally estimate a parametric RC logit model that assumes a correlated multivariate normal distribution for the RCs. Confidence intervals for the parametric estimator are constructed using a percentile bootstrap with the same resampling design. The parametric model is estimated using the R package logitr logitr with $S=250$ simulation draws. All confidence intervals are constructed at the 95% confidence level.
We evaluate the finite-sample performance in terms of the average absolute relative bias of the plug-in estimators and the empirical coverage of the confidence intervals. Let $\tau_{\beta_1,0}$ and $\tau_{\eta,0}$ denote the true population values of the corresponding functionals implied by the continuous mixture distribution, and let $\hat{\tau}^{(m)}$ denote the corresponding plug-in estimator in Monte Carlo replication $m=1,\ldots,M$. The average absolute relative bias is computed as \[ \widehat{\mathrm{Bias}} = \frac{1}{M} \sum_{m=1}^{M} \left| \frac{\hat{\tau}^{(m)}-\tau_0}{\tau_0} \right|, \] whereas the empirical coverage is computed as \[ \widehat{\mathrm{Cov}} = \frac{1}{M} \sum_{m=1}^{M} \mathbbm{1} \left\{ \tau_0 \in CI^{(m)} \right\}, \] where $CI^{(m)}$ denotes the confidence interval obtained in replication $m$. Since the objective in practice is to recover the true population functionals of the underlying continuous RC distribution rather than the pseudo-true functionals implied by a finite-grid approximation, both bias and coverage are evaluated relative to the average functionals evaluated at the true RC distribution.\\
Figure (ref) plots the kernel density estimates of the standardized plug-in estimators for the own-elasticity at the mean against the standard normal distribution. The corresponding results for the mean RC are reported in Figure (ref) in Appendix (ref). For the fixed-grid estimators, and in particular for the ENet estimator, the standardized statistics are centered close to zero and resemble the shape of a standard normal distribution, especially for $N=10000$. In contrast, the standardized NMXL estimator is clearly shifted away from zero. This indicates that imposing an incorrect parametric distribution for the RCs leads to substantial finite-sample bias in the estimated functionals.
Table (ref) reports the average relative bias, the empirical coverage of nominal 95% confidence intervals, and the average standard errors across 1000 Monte Carlo replications for the mean RC, \(\tau_{\beta_1,0}\), and the own-elasticity at the mean, \(\tau_{\eta,0}\). Consistent with the severe bias of the parametric estimator displayed in Figure (ref), the corresponding confidence intervals substantially undercover the true value. The problem becomes even more pronounced as the sample size increases because the standard errors shrink while the specification bias remains unchanged. These findings illustrate that imposing an incorrect parametric distribution for the RCs can lead to severely misleading inferential conclusions, highlighting the importance of nonparametric estimators.
Compared with the misspecified parametric estimator, both nonparametric estimators lead to substantially lower bias and more reliable inference. However, their performance depends critically on the density of the grid. With a coarse grid ($R=25$), approximation bias remains substantial, resulting in confidence intervals that still undercover, particularly for the mean RC. As the grid becomes denser, the ENet estimator becomes increasingly accurate, and the corresponding confidence intervals achieve coverage close to the desired nominal 95% level for both functionals for $R=289$ -- while the confidence intervals for the mean RC still exhibit slight undercoverage (0.936 for $N=1000$ and 0.929 for $N=10000$), the corresponding intervals for the own-elasticity attain coverage essentially equal to the nominal level. Hence, these results demonstrate that sufficiently dense grids are essential for an accurate approximation and , therefore, reliable inference about the true parameter when applying the fixed-grid estimator.
The comparison with the FKRB estimator highlights the importance of regularization. While FKRB also reduces the bias relative to the parametric estimator, its bias does not decrease systematically as the grid becomes denser. Consequently, the corresponding confidence intervals continue to undercover for both functionals across all sample sizes. This behavior is consistent with the ill-conditioned problem discussed in Section (ref). As the grid becomes denser, neighboring grid points become increasingly collinear, causing the FKRB estimator to suffer from selection inconsistency and inaccurate estimation of the probability weights: the estimator assigns positive probability mass to grid points outside the true distribution's support while setting weights within the true support to zero, preventing it from fully exploiting the improved approximation afforded by a denser grid heiss2022. In contrast, the additional \(\ell_2\)-penalty stabilizes the ENet estimator, leading to a coverage close to the 95% level when the regularization bias is properly handled.
Across all simulation designs, the ENet estimator consistently exhibits smaller standard errors than the FKRB estimator. Moreover, the precision of the ENet estimator improves as the grid becomes denser, whereas the opposite pattern is observed for FKRB. This demonstrates that the additional \(\ell_2\)-regularization effectively mitigates the multicollinearity induced by dense grids
Finally, Figure (ref) reports the average 95% confidence intervals across Monte Carlo replicates for the mean RC (left panel) and the own-elasticity (right panel). As the comparison with the FKRB and parametric estimators reveals, the bias correction does not lead to overly conservative confidence intervals. Across all simulation designs, the ENet intervals are consistently narrower than the corresponding FKRB intervals and only moderately wider than the parametric NMXL intervals, despite the substantially greater flexibility of the nonparametric specification. Moreover, the ENet confidence intervals become substantially narrower as the sample size increases, reflecting that the estimator becomes more efficient with increasing sample size. In addition, the intervals also decrease slightly in width as the grid becomes denser, indicating that the additional $\ell_2$-regularization successfully controls the multicollinearity induced by dense grids while allowing the estimator to benefit from the improved approximation of the RC distribution.
To illustrate the proposed inference procedure, we revisit the stated-preference travel-choice data analyzed by meijer2006.\footnote{The data are available via the Journal of Applied Econometrics \href{https://journaldata.zbw.eu/dataset/measuring-welfare-effects-in-models-with-random-coefficients}{Data Archive}.} Using survey data on travel-mode choices, they study travelers’ value of time (VoT) -- their willingness to pay for reduced travel time -- under normal, log-normal, and gamma specifications of the RC distribution in RC logit models. The VoT corresponds to the willingness-to-pay functional in Example (ref), with the RC for travel time in the numerator and that for ticket fare in the denominator. They find that the implied VoT distribution is highly sensitive to the assumed parametric specification, making this dataset well suited for illustrating the benefits of our nonparametric inference procedure.
The data originate from a stated-preference survey conducted for the Dutch railway company NS in 1987, in which respondents repeatedly chose between two hypothetical train connections that differed in fare (Euro), travel time (minutes), number of interchanges (0--4), and comfort level ($\{0,1,2\}$, where $2$ denotes the lowest comfort). The sample comprises 235 respondents who completed, on average, 12.5 choice tasks, yielding 2929 observed choices. To keep the empirical illustration simple, we treat these choice tasks as independent cross-sectional observations and abstract from the panel structure.\footnote{Extensions of the fixed-grid estimator to panel data are straightforward using Section 5 of fox2011. Moreover, we verified that accounting for the panel structure has little effect on the corresponding parametric estimates.}
We specify respondent $i$'s utility from train connection $j\in\{1,2\}$ as
where $\beta_i=(\beta_{i,\text{fare}},\beta_{i,\text{time}})'$ contains the RCs on fare and travel time, $\delta=(\delta_{\text{int}},\delta_{\text{comf}})'$ contains fixed coefficients on interchanges and comfort, and $\varepsilon_{i,j}$ are i.i.d.\ Type I extreme value errors. In contrast to meijer2006, we treat only the fare and travel-time coefficients as random. Allowing these two coefficients to vary jointly enables us to flexibly estimate the VoT distribution -- our primary object of interest -- while keeping the RC vector two-dimensional enables to specify a sufficiently dense grid.\footnote{The total number of grid points grows exponentially with the number of RCs. Allowing additional coefficients to be random would make a sufficiently dense grid infeasible given the available sample size.} To assess the implications of restricting heterogeneity to the fare and travel-time coefficients, we compare the parametric log-normal benchmark used in our analysis with the corresponding four-dimensional log-normal specification considered by meijer2006. The resulting estimates are very similar, supporting the use of two RCs in both our parametric benchmark and nonparametric specifications.
We estimate the joint RC distribution using the penalized and unpenalized fixed-grid estimators with $R\in\{25,81,289\}$ uniformly spaced support points. The grid domain is determined from preliminary estimates of a parametric bivariate normal model. For each $k\in\{\text{fare},\text{time}\}$, the lower bound is $\underline{\beta}_k=\widehat{\mu}_k-4\widehat{\sigma}_k$, where $\widehat{\mu}_k$ and $\widehat{\sigma}_k$ denote the estimated mean and standard deviation under the parametric normal model. The upper bounds are $\overline{\beta}_{\text{fare}}=-0.1$ and $\overline{\beta}_{\text{time}}=0$, imposing negative marginal utility of fare and travel time.
The model is estimated using the iterative algorithm of heiss2022, which alternates between updating the probability weights via a linear probability model estimated with constrained least squares and updating the fixed coefficients via a weighted conditional logit model until convergence. In line with Section (ref), we follow fox2011 for the construction of the weighting matrix. The ENet tuning parameter is selected by five-fold cross-validation over 101 candidate values generated by glmnet, including zero. We construct 95% confidence intervals for the mean RCs and the implied mean VoT using the proposed confidence-conservative procedure with a multinomial block bootstrap based on 500 replications.
As a benchmark, we estimate a parametric model in which the fare and travel-time coefficients follow a correlated bivariate log-normal distribution. We focus on the log-normal specification because it provided the best fit among the parametric models considered by meijer2006 while ensuring a finite mean VoT. The model is estimated by simulated maximum likelihood using 250 Halton draws per respondent. Since the ratio of two jointly log-normal variables is itself log-normal, we calculate the mean VoT from its closed-form mean using the estimated mean vector and covariance matrix of the underlying bivariate log-normal distribution.\footnote{Since the coefficients are jointly log-normal, the logarithm of their ratio is normally distributed. Let $\mu_{\text{time}}$ and $\mu_{\text{fare}}$ denote the means of the log-coefficients and $\sigma_{\text{time}}^2$, $\sigma_{\text{fare}}^2$, and $\sigma_{\text{tf}}$ the corresponding variances and covariance. The mean VoT is then given by $\mathbb{E}[VoT]=-\exp\!\left((\mu_{\text{time}}-\mu_{\text{fare}})+1/2(\sigma_{\text{time}}^2+\sigma_{\text{fare}}^2-2\sigma_{\text{tf}}) \right)$.} The corresponding confidence intervals are obtained via a block bootstrap with 500 replications.
Table (ref) reports the estimated parameters and implied mean VoT estimates for the parametric benchmark, FKRB, and ENet estimators. For the nonparametric fixed-grid estimators, we report the results for \(R=289\), which provides the best fit to the data in terms of the log-likelihood among the considered grid specifications. The corresponding results for \(R=25\) and \(R=81\) are reported in Table (ref) in Appendix (ref).\footnote{The results show that the estimated means and standard deviations are relatively robust across grid specifications for both FKRB and ENet, while the absolute mean VoT decreases as $R$ increases, with more pronounced differences for FKRB than for ENet.}
The results reveal substantial differences between the parametric and nonparametric estimates of the RC distribution and the implied mean VoT estimates. The differences are relatively small for the mean fare coefficient, although the parametric specification implies substantially greater dispersion in fare sensitivities among respondents. In contrast, the differences are much more pronounced for the travel-time coefficient: both fixed-grid estimators imply substantially stronger preferences for travel-time reductions and markedly greater heterogeneity across respondents. As a consequence, the implied mean VoT for a one-hour reduction in travel time is considerably larger under the nonparametric specifications. While the parametric benchmark implies a mean VoT of \(9.36\) EUR per hour, the corresponding estimates equal \(16.25\) EUR for FKRB and \(25.77\) EUR for ENet.\footnote{meijer2006 report a mean VoT estimate of \(9.53\) EUR per hour for a specification in which all four coefficients are assumed to follow a correlated log-normal distribution. This similarity suggests that restricting heterogeneity to fare and travel time has only a minor effect on the implied mean VoT.}
The confidence intervals provide further information about the precision of the estimates across estimators. All reported coefficients and mean VoT estimates are statistically different from zero at the \(5\%\) level. The intervals under the nonparametric estimators are generally wider than those under the log-normal specification, reflecting the usual efficiency advantage of parametric estimation. Importantly, despite accounting for regularization bias, the ENet intervals are only moderately wider than the FKRB intervals because regularization reduces sampling variance. The proposed intervals therefore remain informative rather than being excessively conservative. Moreover, the log-normal and ENet point estimates of mean VoT lie outside each other's confidence intervals. Although this does not constitute a formal test of equality, together with the simulation evidence that parametric misspecification can generate persistent bias, it indicates that the VoT estimates are sensitive to the imposed distributional assumptions.
Figure (ref) compares the VoT distributions implied by the parametric benchmark and the two nonparametric fixed-grid estimators for $R=289$ grid points to further motivate the difference in the estimated mean VoT between the three estimators. The plot groups the VoT grid points into ten equally spaced bins, where the reported mass corresponds to the sum of probability weights associated with all grid points inside the respective bin. In addition, we zoom in on the left tail of the distribution to examine those more closely and compare the parametric benchmark against the nonparametric estimators. Although all three estimators allocate considerable mass near zero, the estimated log-normal distribution decays much more rapidly as the VoT increases, whereas FKRB and ENet retain substantial mass at larger VoT values. Among the nonparametric estimators, FKRB produces a sparser distribution than ENet, as is particularly visible in the left-tail zoom. These thicker tails explain the differences in the estimated mean VoT reported in Table (ref). Figure (ref) illustrates that the results are robust across different numbers of grid points.
To better understand the sources of the differences in the estimated VoT distributions, Figure (ref) displays the underlying joint distributions of the fare and travel-time coefficients. Similarly to the VoT distribution in Figure (ref), the heatmap is constructed using eight equally spaced bins per dimension. We compare the parametric log-normal distribution with FKRB and ENet for $R=289$ grid points.
Under the parametric specification, the coefficients are more strongly positively correlated, so small travel-time coefficients tend to be accompanied by small fare coefficients, leading to a large share of travelers with small VoT values. In contrast, FKRB and ENet recover a non-negligible share of travelers with small fare and large travel-time coefficients, and hence larger VoT values in absolute terms. Consistent with this pattern, the estimated correlations are $0.560$, $0.413$, and $0.291$ for Log-N, FKRB, and ENet, respectively (see Table (ref) in Appendix (ref)). These differences explain the thicker VoT tails and larger mean VoT estimates reported in Table (ref). Figure (ref) shows that the qualitative findings are robust across different numbers of grid points.
Nonparametric estimators of RC models are increasingly used to avoid restrictive distributional assumptions about unobserved preference heterogeneity. Among these methods, the fixed-grid estimator of bajari2007 and fox2011 is particularly attractive because it is computationally tractable and easy to implement. Accurately approximating the RC distribution requires a sufficiently dense grid, which, however, generates severe multicollinearity and unstable estimates of the probability weights. The penalized fixed-grid estimator of heiss2022 addresses this problem through $\ell_2$-regularization, but the resulting gains in stability come at the cost of regularization bias.
This paper develops bias-aware inference for economically relevant functionals based on the penalized fixed-grid estimator. We derive the asymptotic distribution of the penalized plug-in estimator and a feasible upper bound on the resulting regularization bias. Combining these components yields confidence intervals that remain valid when regularization bias is non-negligible. The procedure applies to scalar functionals commonly reported in empirical applications, including mean RCs, elasticities, willingness-to-pay measures, and predicted choice probabilities.
The Monte Carlo simulations demonstrate the importance of addressing both regularization and sieve approximation. When the grid is sufficiently dense, the proposed intervals achieve coverage close to the nominal level while remaining informative in finite samples. In contrast, a coarse grid generates persistent approximation bias and substantial undercoverage, even as the sample size increases. The simulations therefore show that regularization is important for making estimation on dense grids sufficiently stable, while the proposed bias adjustment is needed to account for the bias introduced by that regularization. In addition, our travel-mode application shows that the nonparametric estimators imply substantially thicker tails of the VoT distribution and considerably larger mean VoT estimates than the parametric log-normal specification. At the same time, the bias-aware ENet intervals are only moderately wider or even narrower than the FKRB intervals, indicating that accounting for regularization bias does not result in excessively conservative inference.
Several extensions provide promising directions for future research. First, the simulation results highlight the importance of the grid domain and resolution. Developing data-driven methods that jointly select the location and number of grid points would be valuable. And second, the proposed intervals account for regularization bias but require the sieve approximation bias to be sufficiently small for inference on the true population functional. An interesting extension would be to construct confidence intervals that explicitly account for sieve approximation bias, particularly in high-dimensional applications where the exponential growth of the grid limits a dense resolution in each dimension.