Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,865 characters · 17 sections · 25 citation commands
Informativeness under Model Uncertainty: Shadow Prices and Ridge Penalties
\onehalfspacing
Empirical research often involves competing theories under multiple model restrictions. Importance of accommodating this model uncertainty is gaining overdue attention. We focus on a setting with potentially invalid or approximate restrictions on a target component of models, $m_1$, say, with possibly further nuisance components. Our goal is to assess degrees of compatibility with the data evidence, priced by shadow prices, avoiding suboptimal model selection strategies on $m_1$.
This perspective aligns with semiparametric and modern machine learning frameworks under sparsity. We view the model as consisting of a target component $m_1$, a nuisance component $m_2$, managed by Machine Learning, and an error term (see, e.g., Chernozhukovetal2018). After accounting for the nuisance component $m_2$, inference reduces to the finite-dimensional target $m_1$, on which economically meaningful, theory-driven restrictions are imposed. This is a generalization of partially linear models and General Linear Models (GLM). Methods for robust and debiased ML, prior to inference on $m_1$ abound. See DrukkerandLiu2022 for a recent review and algorithms, highlighting the centrality of Neyman Orthogonalization and computational time and efficiency.
Empirical decisions are sensitive to weak or misspecified restrictions. This is the case whether restrictions produce greater sparsity or enlarge predictive sets by irrelavant variables; see Gospodinov, Kan, and Robbati (Gospodinovetal2014)-GKR. Fundamental tradeoff, imposing restrictions may introduce bias, while discarding them sacrifices efficiency, is exacerbated by possibility of limiting properties of statistics, see GKR (Gospodinovetal2014) and MaasoumiandPhillips1982. This tension mirrors the bias--variance tradeoff inherent in modern regularization and shrinkage methods.
We propose a unified framework for estimation and the assessment of informativeness (empirical relevance) under multiple restrictions and model uncertainty. Rather than selecting a single model, our approach allows for controlled degrees of misspecification and lets the data determine the empirical importance of each restriction. Specifically, the framework delivers consistent estimation in the presence of multiple candidate restrictions while simultaneously quantifying the empirical relevance of each restriction for model fit. In this sense, we move from model selection toward a structured form of model averaging or mixed estimation (of $m_1$) that better reflects practical empirical work.
Initial assessment of restrictions based on their empirical relevance precedes traditional testing. This effectively prioritizes informative restrictions and downweights or discards those that are weak or severely misspecified. The approach is consistent with the spirit of modern double machine learning, where nuisance components are flexibly controlled and inference focuses on a low-dimensional target. By incorporating a data-driven assessment of restrictions, the framework reduces the multiplicity burden while maintaining a fixed error rate, thereby improving the power of subsequent inference.
At the estimation stage, we adopt a bi-level (profiling) algorithm. The inner problem solves a Lagrangian constrained optimization for a given tolerance level governing restriction misspecification, yielding the constrained estimator and its associated shadow price. The outer problem selects the tolerance parameter by optimizing a Stein-type risk criterion, allowing the estimator to adapt to the empirical environment and optimally balance regularization bias and variance. Importantly, this approach does not require prior knowledge of which restrictions are valid or informative and is robust to weak or misspecified restrictions.
We further refine estimation by correcting a residual centering bias. Although the optimized tolerance parameter achieves an optimal bias--variance tradeoff, it does not eliminate the regularization bias induced by the constraint. We therefore introduce a debiasing step based on a local characterization of this bias. From the Karush--Kuhn--Tucker conditions, the deviation of the pseudo-true parameter from the unconstrained target is proportional to the Lagrange multiplier, which captures the local tightness of the constraint. This leads to a simple correction that removes the leading distortion and recenters the estimator without affecting its first-order variance.
Building on this estimation framework, we assess the empirical relevance of restrictions through individual shadow prices (ISP), which decompose the global shadow price into restriction-specific components. These ISP measure the marginal gain from relaxing each restriction and provide a coherent and comparable metric of empirical relevance. We then propose a plateau rule to separate signal from noise by detecting a structural break in the ordered ISP sequence.
We establish consistency and asymptotic normality of the constrained estimator with optimized misspecification tolerance and its debiased counterpart, and characterize the asymptotic behavior of the ISP. Notably, the constrained optimization approach yields an estimator that is robust to restriction misspecification, as the tolerance parameter allows the data to determine the extent to which restrictions are enforced. The asymptotic analysis formalizes this property: the optimized tolerance converges to a pseudo-true value governing the bias--variance tradeoff, so that the constrained estimator targets a pseudo-true parameter reflecting both sampling variation and misspecification. While this does not eliminate regularization bias, the debiasing step corrects the leading distortion and restores convergence to the structural target.
Simulation results demonstrate strong finite-sample performance: the proposed estimator is consistent, analytical inference attains near-nominal coverage comparable to the bootstrap, and the plateau rule effectively separates signal from noise. The empirical application highlights a key insight: statistical validity and empirical relevance are distinct. A restriction may be rejected by standard tests yet remain empirically irrelevant, as relaxing it does not improve model fit. This highlights the value of the proposed approach in identifying which restrictions matter for estimation.
The rest of the paper is organized as follows. In Section 1.1, we discuss the relation of our framework to the existing literature. In Section 2, we present our theory framework. In Section 3, we develop the asymptotic theory. In Section 4, we develop bootstrap-based inference as an extension. In Section 5, we report Monte Carlo simulation results. In Section 6, we provide an empirical application. In Section 7, we conclude. Additional discussions are provided in the Supplement.
\paragraph{Restrictions from economic theory.} We consider a priori restrictions on the parameter space and through statistical distributions (e.g., moment restrictions). Direct restrictions may be imposed in restricted estimation, and implied restrictions may be tested by diagnostic testing.
For example, in macroeconomics and growth, structural approaches such as Solow1956 and Swan1956, the human-capital-augmented growth regressions of MankiwRomerandWeil1992, and the spatial growth model of ErturandKoch2007 incorporate cross-parameter restrictions implied by theory directly into the regression specification. Similarly, in microeconomics and industrial organization, rational expectations models HansenandSargent1980 and auction models based on Bayesian Nash equilibrium GuerrePerrigneandVuong2000 embed equilibrium restrictions within the estimating equations.
Other “restrictions” are implied and treated as testable implications of an empirical model. In this case, estimation proceeds without imposing the full set of theoretical restrictions, and the restrictions are evaluated ex post using the data. Examples include present-value tests based on vector autoregressions CampbellandShiller1987, revealed-preference and shape-restriction tests in demand analysis Blundell2005, and asset pricing tests that assess zero-alpha conditions or bounds on risk aversion FamaandMacBeth1973,MehraandPrescott1985.
We provide a unified, data-driven method to rank and select restrictions prior to formal validity testing. The proposed constrained optimization framework accommodates a broad class of linear and nonlinear restrictions.
\paragraph{Model uncertainty, ambiguity, and misspecification.} Model uncertainty arises when multiple candidate restrictions are considered as if they were (approximately) valid. Constrained models may be within a limited ball of deviating candidate models (model ambiguity), or globally misspecified. Metrics and different rates for deviations are necessary. It has been increasingly recognized that shrinkage methods, with or without pretesting interpretation, provide an actionable approach to formalizing degrees of misspecification.
Early work by Maasoumi1978 proposed combining estimators under model uncertainty using information from test statistics. More recent literature on model ambiguity and sensitivity (e.g., BonhommeandWeider2022, ChristensenandConnault2023) treats such uncertainty as intrinsic, allowing models or distributions to vary within neighborhoods of a benchmark specification. These approaches quantify the sensitivity of conclusions to misspecification, often through a notion of misspecification size. For instance, ChristensenandConnault2023 derive counterfactual bounds by optimizing over divergence-based neighborhoods. Our framework builds on this perspective but differs in a key dimension. Rather than varying distributions or moment conditions directly, we parameterize misspecification through a tolerance level $c$ on restriction violations and select it endogenously. This shifts the focus from evaluating robustness over a fixed neighborhood to optimizing the degree of deviation from imposed restrictions.
A context is provided by the Hansen--Jagannathan distance (GKR, Gospodinovetal2013,Gospodinovetal2016), which provides a geometric measure of misspecification. While the HJ-distance treats misspecification as an ex post phenomenon, our framework treats it as a decision variable. Moreover, Lagrange multipliers are central in our analysis and are interpreted as ISP, which quantify the marginal contribution of each restriction. A bound on a constraining hyper sphere (or elipsoid) which represents tolerance for model deviations (see below for definition of “c”), is optimized in the spirit of Stein-Like estimators, for an explicit bias--variance tradeoff, extending existing robustness frameworks to an optimization-based approach to model uncertainty and misspecification.
In contrast to local misspecification as the key feature of model ambiguity, global or other considerations of misspecification raise the question of model selection and model averaging. In the ML context, LASSO is an example of model selection, whereas Ridge is a form of model averaging. Model selection entails biases and generally requires presumption of a reference as “the true DGP”. This is known to be suboptimal.\footnote{Bayesian methods are also generally model selection procedures, but may be extended to model averaging.} Model averaging can avoid these shortcomings at the cost of adopting an oracle model average object; see GospodinovandMaasoumi2021.
\paragraph{Inference-based insights for regularization methods in machine learning.} Our framework is closely related to machine learning literature through the duality between constrained optimization and penalized estimation HoerlandKennard1970,Tibshirani1996,ZouandHastie2005. This duality implies that the estimator can be viewed as solving a constrained problem in which deviations from imposed restrictions are controlled within a tolerance level, thereby allowing for restriction misspecification and enhancing robustness. Importantly, while the machine learning literature typically addresses regularization bias through tuning parameters, our approach treats it as a statistical optimization problem, using a Stein-type risk criterion for selecting the tolerance parameter and a debiasing correction based on analytic theory. Optimization techniques for choice of “tuning parameters”, such as by Bayesian Information Criterion (BIC), compete with Cross Validation (CV) and other ad hoc methods.
Our framework extends these methods to an inferential setting and contributes to the growing literature on inference in machine learning. While inference tools for machine learning models, such as post hoc decomposition methods like SHAP LundbergandLee2017, quantify the contribution of individual covariates to predictions from a fixed model, our approach develops statistical decision rules for assessing the relative empirical support for restrictions (that may be false). This is consistent with our preference to avoid assumption of a “true DGP”.
Let \(g(\theta) \in \mathbb{R}^q\) denote a vector function of the parameter \(\theta \in \mathbb{R}^p\), and let \(\Sigma \in \mathbb{R}^{q \times q}\) be a known positive definite weighting matrix. We represent the restriction system through the quadratic form $ h(\theta)=g(\theta)^{\prime}\Sigma^{-1}g(\theta).$
We introduce a finite upper bound $c_0 > 0$ on $h(\theta)$, from a candidate set \(\mathcal C \subset (0,c_0]\) of tolerance levels. For each \(c \in \mathcal C\), the scalar \(c\) determines how tightly the restrictions are enforced: small values impose strong adherence, while larger values allow greater deviation and thus provide robustness to misspecification. Rather than fixing \(c\) a priori, we select it from the data using a Stein-type analytical risk criterion. These restrictions represent many test statistics and penalization functions, including the canonical Ridge. The scalar $c$ resembles the critical level of frequentist testing procedures, but has counterparts in Bayesian prior interpretations.
Our framework is cast in a general M-estimation setting. Let $ \phi_n(\theta)=\frac{1}{n}\sum_{i=1}^n \phi(Z_i,\theta)$, $\phi(\theta)=\mathbb{E}[\phi(Z_{i},\theta)]$, and $\phi^0=\mathbb{E}[\phi(Z_i,\theta^{0})], $ where \(Z_i=(y_i,x_i)\) are i.i.d.\ observations, and \(\phi(\cdot)\) is a loss function. For each \(c \in \mathcal C\), define the constrained estimator
The associated Lagrangian is
with first-order condition $ \nabla_\theta \phi_n(\theta) + 2\lambda \nabla_\theta g(\theta)^{\prime}\Sigma^{-1}g(\theta) = 0. $ Thus, for each \(c\), the pair \((\hat{\theta}(c),\hat{\lambda}(c))\) is determined jointly from the Karush--Kuhn--Tucker (KKT) system. The multiplier \(\hat{\lambda}(c)\) measures the marginal gain from relaxing the restriction and admits a natural interpretation as a shadow price.
The outer step selects the tolerance level via
yielding the optimized estimator $ \hat{\theta}=\hat{\theta}(\hat c),\; \hat{\lambda}=\hat{\lambda}(\hat c). $ The procedure is therefore bi-level: the inner problem computes \(\hat{\theta}(c)\), while the outer problem selects \(c\) to balance regularization bias and variance.
\paragraph{Population characterization.} Let \((\theta^*(c),\lambda^*(c))\) denote the population KKT solution. If \(\lambda^*(c)>0\), the constraint binds and affects the solution; if \(\lambda^*(c)=0\), the constraint is slack and the estimator coincides locally with the unconstrained solution. In general, \(\theta^*(c)\) differs from the true parameter and represents a pseudo-true value.
\paragraph{Soft restrictions via constrained optimization.} This framework departs from exact-restriction regression by treating theory as approximate rather than binding. Instead of imposing $g(\theta)=0$, we allow deviations through the constraint $g(\theta)^{\prime}\Sigma^{-1}g(\theta)\leq c$, letting the data determine the extent of enforcement. This preserves feasibility under misspecification and introduces a bias--variance tradeoff governed by $c$, where smaller values approximate exact restrictions and larger values approach the unrestricted estimator. The Lagrange multiplier provides a shadow-price interpretation, measuring the marginal gain in fit from relaxing the restriction.
\paragraph{Economic interpretation of ISP.} The global shadow price is given by $\partial \mathcal{L}/\partial c = -\lambda$, which measures the marginal reduction in the objective from relaxing the tolerance level. The ISP decompose this quantity across restrictions via directional derivatives: \[ \mathrm{ISP}_j(c) = 2\lambda(c)\,\mathrm{sign}(g_j)\,[\Sigma^{-1}g]_j. \] Each ISP captures the marginal gain in fit from relaxing the $j$th restriction. Their magnitudes therefore provide a natural measure of empirical relevance and offer a principled basis for ranking restrictions.
We develop the theory in several steps. First, we show that on the active region the KKT system can be locally profiled with respect to the tolerance parameter \(c\), so that the constrained estimator and its associated shadow price vary smoothly with \(c\). Next, treating \(c\) as fixed, we derive the inner asymptotic theory for the constrained estimator, the global shadow price, and the ISP vector. We then analyze the outer problem that selects \(c\) by minimizing a Stein-type analytical risk proxy and characterize the resulting debiased optimized estimator. Finally, we introduce a plateau rule for inference on the ISP vector, which separates signal from noise using a bootstrap procedure.
We first study the local profiling map with respect to the tolerance parameter \(c\). Because smooth profiling requires the equality form of the KKT system, the result is stated on the active region \(\mathcal C_{\mathrm{active}}\), where the quadratic restriction binds and the multiplier is strictly positive. This profiling result is therefore conditional on local empirical relevance of the aggregate restriction, not an assumption imposed on the whole model. For each \(c\in\mathcal C\), define the sample KKT map
and the population KKT map $ \Phi(\theta,\lambda;c) :=
. $ Under the binding regime, the population solution \((\theta^*(c),\lambda^*(c))\) satisfies $ \Phi(\theta^*(c),\lambda^*(c);c)=0, \; \lambda^*(c)>0, $ and the sample one \((\hat\theta(c),\hat\lambda(c))\) satisfies $ \Phi_n(\hat\theta(c),\hat\lambda(c);c)=0, \; \hat\lambda(c)>0. $
Proof. See Appendix (ref).
Proof. See Appendix (ref).
Proof. See Appendix (ref).
We now treat \(c\) as fixed and derive the inner asymptotic theory for the constrained estimator, the global shadow price, and the ISP vector. The theory is stated without imposing a global positive-multiplier assumption. Instead, whenever a result requires a binding quadratic restriction, this is stated explicitly by restricting attention to \(c\in\mathcal C_{\mathrm{active}}\).
Assumption (ref) ensures that the feasible set has a nonempty interior and that the KKT conditions characterize the constrained optimum. Assumption (ref) ensures existence and uniqueness of the population constrained minimizer and justifies a second-order expansion of the Lagrangian. Assumption (ref) guarantees local regularity of the constraint surface and identification of the KKT system. Assumption (ref) ensures strong duality and validity of the KKT characterization. Definition (ref) distinguishes the active and inactive regimes. On the active region, the shadow-price system is nondegenerate and admits a smooth first-order expansion. On the inactive region, the quadratic restriction is slack and the constrained estimator reduces locally to the unconstrained regime. Accordingly, the asymptotic results below for \((\hat\theta(c),\hat\lambda(c))\) and the ISP vector are stated for fixed \(c\in\mathcal C_{\mathrm{active}}\). Assumption (ref) provides the stochastic equicontinuity and central limit theorem required for a first-order expansion of the score around \(\theta^*(c)\). The uniform laws of large numbers for \(\phi_n(\theta)\) and its derivatives guarantee local stability of the objective and its curvature, allowing a valid Taylor expansion in a neighborhood of \(\theta^*(c)\).
For each fixed \(c\in\mathcal C_{\mathrm{active}}\), define $ a(c)\equiv \nabla_\theta h(\theta^*(c)) = 2\,G(\theta^*(c))^{\prime}\Sigma^{-1}g(\theta^*(c)), $ and the population Hessian of the Lagrangian, $ A(c)\equiv \nabla^2_{\theta\theta}\mathcal L(\theta^*(c),\lambda^*(c);c) = H(\theta^*(c))+\lambda^*(c)\nabla^2_{\theta\theta}h(\theta^*(c)), $ with $H(\theta^*(c))\equiv \nabla^2_{\theta\theta}\phi(\theta^*(c))$. Here \(a(c)\) is the gradient of the binding restriction and therefore the normal vector to the constraint surface \(\{\theta:h(\theta)=c\}\) at \(\theta^*(c)\). As a normal vector, \(a(c)\) identifies the direction in parameter space along which movement is locally forbidden when the constraint binds: moving in the direction of \(a(c)\) or \(-a(c)\) changes the value of \(h(\theta)\) to first order and hence relaxes or tightens the restriction. In contrast, perturbations orthogonal to \(a(c)\) lie in the tangent space of the constraint surface and preserve feasibility to first order.
The matrix \(A(c)\) is the curvature matrix of the population Lagrangian with respect to \(\theta\), evaluated at the KKT solution. Since \(\theta^*(c)\) is assumed to be a local minimizer of the constrained loss, standard second-order sufficient conditions imply that \(A(c)\) is positive definite, ensuring local uniqueness and identification of the solution. Under these conditions, \(A(c)^{-1}\) exists and is positive definite, and therefore \(a(c)^{\prime}A(c)^{-1}a(c)>0\) for any nonzero \(a(c)\).
Proof. See Appendix (ref).
Proof. See Appendix (ref).
The matrix \(M(c)\) can be interpreted as a projected inverse Hessian. It modifies the unconstrained influence matrix \(A(c)^{-1}\) by removing variation in the direction that would locally violate the binding constraint. This direction is identified by the constraint gradient \(a(c)=\nabla_\theta h(\theta^*(c))\), which is normal to the constraint surface. Variation along \(a(c)\) would move the estimator off the feasible set and is therefore ruled out by the binding restriction. The projection implemented by \(M(c)\) eliminates this forbidden component, ensuring that the asymptotic variation of \(\hat{\theta}(c)\) lies entirely in directions that preserve feasibility to first order.
For each \(c\in\mathcal C_{\mathrm{active}}\), define the population ISP: $ISP_j^*(c) := 2\lambda^*(c)\,\mathrm{sign}\!\big(g_j(\theta^*(c))\big)\, \big[\Sigma^{-1}g(\theta^*(c))\big]_j, $ and the sample ISP estimator $ \widehat{ISP}_j(c) := 2\hat\lambda(c)\,\mathrm{sign}\!\big(g_j(\hat\theta(c))\big)\, \big[\Sigma^{-1}g(\hat\theta(c))\big]_j. $ Stack $ \widehat{ISP}(c)\equiv(\widehat{ISP}_1(c),\dots,\widehat{ISP}_q(c))^{\prime}, \; ISP^*(c)\equiv(ISP_1^*(c),\dots,ISP_q^*(c))^{\prime}. $
Assumption (ref) guarantees that the sign of each relevant restriction is locally stable in a neighborhood of the population target \(\theta^*(c)\). Hence, with probability approaching one, the sign operator is locally constant and does not affect first-order delta-method expansions of the ISP. If \(g_j(\theta^*(c))=0\) for some \(j\), the sign map is non-differentiable at \(\theta^*(c)\), and stochastic sign switching may occur in \(n^{-1/2}\) neighborhoods, leading to nonregular asymptotics that require separate analysis. Note that Assumption (ref) is imposed at $\theta^*(c)$, rather than at $\theta^0$. Since ISP is evaluated at a tolerance level $c$ that optimizes the bias--variance tradeoff and is therefore not chosen to recover $\theta^0$, a restriction need not hold exactly at the pseudo-true constrained target $\theta^*(c)$, even if it holds exactly at the true data-generating parameter $\theta^0$.
Proof. See Appendix (ref).
\noindentProof. See Appendix (ref).
To formalize ranking in terms of magnitudes, let \(r^*(c)\) denote the permutation that sorts population ISP magnitudes in weakly decreasing order:
\noindentProof. See Appendix (ref).
\paragraph{Plateau cutoff \(\hat m(c)\).} Building on the ordered ISP magnitudes, we define a point estimator of the boundary between noise and signal restrictions. Let $ |\widehat{ISP}_{(1)}(c)| \le \cdots \le |\widehat{ISP}_{(q)}(c)| $ denote the ordered absolute ISPs, where the ordering is now from smallest to largest so that the lower tail corresponds to candidate noise restrictions. For each candidate cutoff \(m\in\{m_0,\dots,q-1\}\), define $ \mathcal G_1(m)=\{1,\dots,m\}, \; \mathcal G_2(m)=\{m+1,\dots,q\}, $ where \(\mathcal G_1(m)\) is interpreted as a candidate plateau and \(\mathcal G_2(m)\) as the remaining signal region. We consider the contrast \[ \Delta_m(c) = \frac{1}{m}\sum_{\ell=1}^m |\widehat{ISP}_{(\ell)}(c)| - \frac{1}{q-m}\sum_{\ell=m+1}^q |\widehat{ISP}_{(\ell)}(c)|. \] To assess whether a structural break occurs at \(m\), we test \[ H_0:\ \Delta_m(c)=0, \] using the Wald-type statistic $ B_m(c) = n\,\Delta_m(c)^2 \big/ \widehat{\mathrm{Var}}(\Delta_m(c)), $ where \(\widehat{\mathrm{Var}}(\Delta_m(c))\) is obtained from \(\widehat{\Sigma}_{ISP}(c)\) via a linear contrast.
To ensure that the lower block behaves like a plateau, we also impose a within-block homogeneity screen. For each candidate \(m\), we test \[ H_0:\ |\widehat{ISP}_{(1)}(c)| = \cdots = |\widehat{ISP}_{(m)}(c)|, \] using a joint Wald test based on pairwise contrasts within \(\mathcal G_1(m)\). Let \(\mathcal M_{\mathrm{adm}}\) denote the set of candidate cutoffs that pass this homogeneity screen. We then define the point estimator $ \hat m(c) = \arg\max_{m\in\mathcal M_{\mathrm{adm}}} B_m(c), $ with the maximization carried out over all feasible candidates if no \(m\) passes the screen. The estimated plateau is given by \(\{1,\dots,\hat m(c)\}\), and restrictions beyond \(\hat m(c)\) are classified as signal.
\paragraph{Population counterpart.} Let $ |ISP_{(1)}^*(c)| \le \cdots \le |ISP_{(q)}^*(c)| $ denote the ordered population ISP magnitudes, and define $ \Delta_m^*(c) = \frac{1}{m}\sum_{\ell=1}^m |ISP_{(\ell)}^*(c)| - \frac{1}{q-m}\sum_{\ell=m+1}^q |ISP_{(\ell)}^*(c)|. $
Assumption (ref) states that the ordered population ISP sequence admits a well-defined structural break separating a lower-magnitude noise region from a higher-magnitude signal region, in the sense that the population break criterion has a unique maximizer.
\noindentProof. See Appendix (ref).
We now turn to the outer problem and endogenize the tolerance level by selecting \(c\) through an analytical Stein-type risk criterion. Let \(\theta^{0}\) denote the unconstrained target parameter defined by $ \nabla_\theta \phi^0(\theta^{0})=0. $ For each \(c\in\mathcal C_{\mathrm{active}}\), define the population risk $ R(c):=\mathbb E\big[\|\hat\theta(c)-\theta^{0}\|^2\big]. $ Using the decomposition \[ \hat\theta(c)-\theta^{0} = \big(\hat\theta(c)-\theta^*(c)\big) + \big(\theta^*(c)-\theta^{0}\big), \] the leading terms of the risk are a variance component and a regularization-bias component.
Proof. See Appendix (ref).
Proposition (ref) implies that the leading population risk approximation is
Motivated by (ref), define the feasible analytical proxy
where $ \widehat W(c) = \frac{1}{n}\operatorname{tr}\big(\widehat V_1(c)\big), $ and $ \widehat B_{\mathrm{lin}}(c) = \hat{\lambda}(c)^2\, \hat a^{\prime}\hat H^{-1\prime}\hat H^{-1}\hat a. $ Here \(\tilde\theta\) is the unconstrained target estimator $ \tilde\theta = \arg\min_{\theta}\phi_n(\theta), $ and $ \hat a=\nabla_\theta h(\tilde\theta), \; \hat H=\nabla_{\theta\theta}^2\phi_n(\tilde\theta). $
The tolerance level is selected as $ \hat c \in \arg\min_{c \in \mathcal C_{\mathrm{active}}} \widehat R_{\mathrm{lin}}(c). $ Let the pseudo-true optimal tolerance be $ c^{\dagger} = \arg\min_{c \in \mathcal C_{\mathrm{active}}} R_{\mathrm{lin}}(c). \label{eq:c_dagger_def} $
Proof. See Appendix (ref).
Proof. See Appendix (ref).
The optimization over \(c\) via the Stein-type risk criterion selects a tolerance level that optimally balances regularization bias and variance, thereby determining the degree of bias that is acceptable in minimizing overall estimation risk. Consequently, the estimator \(\hat\theta(\hat c)\) is tuned to achieve an optimal bias--variance tradeoff.
Nevertheless, this optimization operates at a global level and does not, in general, eliminate the bias in the centering of the estimator. While the optimized tolerance parameter controls the bias--variance tradeoff, it does not remove the regularization bias induced by the constraint. A subsequent debiasing step is therefore required to eliminate the remaining bias and restore correct centering for inference. In particular, even at the optimal tolerance \(c^\dagger\), we typically have \(\theta^*(c^\dagger)\neq \theta^{0}\) whenever the quadratic restriction remains locally binding.
\paragraph{Local characterization of the regularization bias.} Proposition (ref) shows that the leading regularization bias is of order \(O(\lambda^*(c))\), with direction determined by the curvature of the objective and the gradient of the constraint. Thus, the Lagrange multiplier governs both the strength and direction of the distortion induced by the constraint. Since $ \theta^*(c)-\theta^{0} = O\big(\lambda^*(c)\big), $ the leading bias is proportional to \(\lambda^*(c)\), which captures the local tightness of the constraint. This characterization operates at the level of the parameter \(\theta\), providing a tractable approximation to the residual bias after optimization over \(c\), and thereby enables a simple debiasing correction that recenters the estimator toward \(\theta^{0}\) without affecting its first-order variance.
Motivated by (ref), define the population bias proxy $ b^*(c):=\lambda^*(c)H_0^{-1}a_0. $ Its feasible analogue is $ \hat b(c):=\hat\lambda(c)\,\hat H^{-1}\hat a, \; \hat a=\nabla_\theta h(\tilde\theta), \; \hat H=\nabla_{\theta\theta}^2\phi_n(\tilde\theta), $ where \(\tilde\theta\) is the unconstrained estimator. The debiased optimized estimator is then
The sign convention in (ref) follows from $ \theta^*(c)-\theta^{0}\approx -\,b^*(c), $ so that adding \(\hat b(\hat c)\) removes the leading \(O(\lambda)\) regularization bias.
Proof. See Appendix (ref).
As an extension, we develop a bootstrap-based implementation of the inference procedure. The bootstrap is formally justified for the smooth objects in the system, namely \((\hat\theta(c),\hat\lambda(c))\) and the ISP vector. It is then used to propagate sampling variability through the plateau-selection rule and to construct an empirical distribution of the cutoff. This is especially important because the separation condition in Assumption (ref) is not directly verifiable in practice and may be weak when several restrictions have similar empirical effects.
Let \(\hat c\) denote the optimized tolerance selected by the outer optimization step, and let \((\hat\theta(\hat c),\hat\lambda(\hat c))\) be the corresponding constrained estimator and shadow price from the inner problem. Define the residual $ \hat u_i(\hat c)=y_i-x_i^{\prime}\hat\theta(\hat c), $ evaluated at the constrained estimator associated with the selected tolerance. For each bootstrap replication \(b=1,\dots,B\), draw i.i.d.\ multipliers \(\{w_i^{(b)}\}_{i=1}^n\) such that $ \mathbb E[w_i]=0,\; \mathbb E[w_i^2]=1,\; \mathbb E\!\left[|w_i|^{2+\nu}\right]<\infty \ \text{for some }\nu>0, $ independently of the data. Construct pseudo-outcomes by perturbing the residual component while holding regressors fixed: \[ y_i^{(b)}=x_i^{\prime}\hat\theta(\hat c)+\hat u_i(\hat c)\,w_i^{(b)}. \]
For each bootstrap sample, we then re-implement the full estimation procedure. First, using \(\{(y_i^{(b)},x_i)\}_{i=1}^n\), we solve the outer optimization problem to obtain a bootstrap-selected tolerance \(\hat c^{(b)}\). Second, conditional on \(\hat c^{(b)}\), we solve the corresponding inner constrained problem to obtain \((\hat\theta^{(b)}(\hat c^{(b)}),\hat\lambda^{(b)}(\hat c^{(b)}))\): \[ (\hat\theta^{(b)}(\hat c^{(b)}),\hat\lambda^{(b)}(\hat c^{(b)})) \in \arg\min_{\theta\in\mathbb R^p}\ \phi_n^{(b)}(\theta) \quad \text{s.t.} \quad h(\theta)\le \hat c^{(b)}, \] where \(\phi_n^{(b)}(\theta)\) denotes the bootstrap loss and \(h(\theta)=g(\theta)^{\prime}\Sigma^{-1}g(\theta)\). Thus, each bootstrap replication incorporates both inner-loop estimation uncertainty and outer-loop selection uncertainty.
The ISP in replication \(b\) is evaluated at the bootstrap constrained estimator, namely \[ \widehat{\mathrm{ISP}}^{(b)} = 2\hat{\lambda}^{(b)}(\hat c^{(b)})\, \mathrm{diag}\!\Big(\mathrm{sign}\big(g(\hat{\theta}^{(b)}(\hat c^{(b)}))\big)\Big)\, \Sigma^{-1}g(\hat{\theta}^{(b)}(\hat c^{(b)})). \] We emphasize that ISP is evaluated at the constrained estimator within the inner problem, rather than at the debiased estimator, since otherwise \(g(\theta)\) may be driven artificially close to zero and remove the variation needed for relevance assessment.
Assumption (ref) is standard for bootstrap validity of smooth \(M\)-estimators. Given the KKT-based asymptotic linear representation, bootstrap validity reduces to verifying that the bootstrap score reproduces the same first-order linear structure, which we impose as a high-level condition. In the optimized procedure, bootstrap samples are centered at the estimator evaluated at the selected tolerance \(\hat c\), and the full estimation routine is repeated in each replication.
Proof. See Supplement (ref).
Proof. See Supplement (ref).
Theorems (ref)--(ref) justify the bootstrap for the smooth inner-level objects at a fixed tolerance \(c\). In practice, however, we implement the bootstrap after the optimal tolerance \(\hat c\) has been selected and then re-run the full outer-inner estimation procedure within each bootstrap replication. This yields bootstrap draws \((\hat c^{(b)},\hat\theta^{(b)}(\hat c^{(b)}),\hat\lambda^{(b)}(\hat c^{(b)}),\widehat{ISP}^{(b)})\), which incorporate both selection uncertainty and estimation uncertainty. Accordingly, the fixed-\(c\) bootstrap theory serves as the inner-level justification, while the practical algorithm propagates this uncertainty through the optimized choice of \(\hat c\).
As discussed above, bootstrap implementation of the plateau rule is particularly important because the separation condition in Assumption (ref) is not directly verifiable in practice and may be weak when several restrictions have similar empirical effects. In such cases, small perturbations in the data can alter the ordering of ISP magnitudes and shift the location of the cutoff. The bootstrap addresses this issue by repeatedly perturbing the sample and recomputing the entire plateau-selection procedure, thereby revealing the sensitivity of the cutoff to sampling variability. To this end, we apply the plateau rule within each bootstrap replication and use the resulting distribution to characterize cutoff uncertainty.
For each bootstrap replication \(b=1,\dots,B\), let $ |\widehat{ISP}^{(b)}_{(1)}(c)| \le \cdots \le |\widehat{ISP}^{(b)}_{(q)}(c)| $ denote the ordered bootstrap ISP magnitudes. Repeating the same structural break and homogeneity-screening procedure yields a bootstrap cutoff $ \hat m^{(b)}(c). $ The collection $ \bigl\{\hat m^{(b)}(c)\bigr\}_{b=1}^B $ thus defines an empirical distribution of plateau cutoffs.
This bootstrap distribution serves as the primary object of inference for plateau determination. It provides an operational measure of empirical separability in the ISP sequence: when ISP magnitudes are well separated, the bootstrap cutoffs concentrate around a stable value; when separation is weak, the distribution of \(\hat m^{(b)}(c)\) becomes diffuse, indicating ambiguity in the plateau boundary. In this way, the bootstrap translates an otherwise untestable separation condition into an observable pattern of variability.
In practice, we summarize the bootstrap distribution by its mean and standard deviation, \[ \bar m_B(c)=\frac{1}{B}\sum_{b=1}^B \hat m^{(b)}(c), \qquad s_B(c)=\left(\frac{1}{B-1}\sum_{b=1}^B\big(\hat m^{(b)}(c)-\bar m_B(c)\big)^2\right)^{1/2}, \] and use these summaries to construct an uncertainty-adjusted cutoff. In particular, we define the uncertainty-adjusted plateau boundary as \(\bar m_B(c) + s_B(c)\), which expands the plateau when cutoff selection is unstable across bootstrap replications. This provides a conservative and data-driven rule for separating noise from signal.
Accordingly, our practical inference is based on the distribution of \(\hat m^{(b)}(c)\) rather than solely on the point estimate \(\hat m(c)\). While the consistency result for \(\hat m(c)\) remains a useful theoretical benchmark under Assumption (ref), the bootstrap cutoff distribution offers a more relevant and robust tool for applied implementation.
In this section, we study the finite-sample performance of our proposed method through Monte Carlo simulations. We set the number of iterations to 1{,}000 and the sample size to $n = 1{,}000$. The parameter vector has dimension $p+1=11$, consisting of an intercept and $p=10$ slope coefficients.\footnote{The intercept is included to capture the baseline level of the outcome variable but is excluded from the shrinkage restriction, as penalizing it lacks substantive interpretation and would shift the model’s location rather than regulate structural relationships among slope parameters.}
\paragraph{DGP.} Let $X = [\mathbf{1}, X_{\text{raw}}] \in \mathbb{R}^{n \times (p+1)}$ denote the regressor matrix, where $\mathbf{1}$ is an $n$-vector of ones and $X_{\text{raw}} = Z \Sigma_X^{1/2} \in \mathbb{R}^{n \times p}$ contains the slope covariates. Here, $Z \in \mathbb{R}^{n \times p}$ has i.i.d.\ standard normal entries, $Z_{ij} \overset{\text{i.i.d.}}{\sim} \mathcal{N}(0,1)$, and $\Sigma_X \in \mathbb{R}^{p \times p}$ follows an AR(1) structure, $(\Sigma_X)_{jl} = \rho^{|j-l|}$ for $\rho \in (0,1)$, inducing cross-covariate dependence. Throughout, we set $\rho = 0.8$ to generate strong correlation among the $p=10$ slope variables.
The outcome is generated as $y = X \theta^{0} + \varepsilon$, where $\theta^{0} \in \mathbb{R}^{p+1}$ is the true parameter vector and $\varepsilon \sim \mathcal{N}(0, \sigma_\varepsilon^{2} I_n)$. The disturbance variance is calibrated to achieve a target signal-to-noise ratio of approximately one, $\mathrm{Var}(X\theta^{0}) / \mathrm{Var}(\varepsilon) \approx 1$, corresponding to $R^{2} \approx 0.5$, so that the signal is dominant but the noise remains non-negligible.
\paragraph{Scenarios and parameters.} We consider the least squares loss $\phi_n(\theta) = \frac{1}{2n}\,\lVert y - X\theta \rVert^2$ and study three cases. In Case 1, only structural restrictions are imposed. In Cases 2 and 3, we additionally impose the restriction $\theta_1 + \theta_2 + \theta_3 + \theta_4 = 0$, which is correctly specified in Case 2 and misspecified in Case 3.
The parameter vector $\theta \in \mathbb{R}^{11}$ consists of an intercept and ten slope coefficients. In Case 1, $ \theta^{0}_{\text{Case 1}} = (0.1, 0.3, 0, -0.5, 0, 0.3, 0, 0.4, 0, -0.2, 0)^{\prime}. $ In Case 2, $ \theta^{0}_{\text{Case 2}} = (0.1, 0.3, 0.2, -0.5, 0, 0.3, 0, 0.4, 0, -0.2, 0)^{\prime}, $ so that the imposed restriction holds at the population level, as $ \theta_1^{0} + \theta_2^{0} + \theta_3^{0} + \theta_4^{0} = 0.3 + 0.2 - 0.5 + 0.0 = 0. $ In Case 3, $ \theta^{0}_{\text{Case 3}} = (0.1, 0.3, 0, -0.5, 0, 0.3, 0, 0.4, 0, -0.2, 0)^{\prime}, $ for which the restriction is violated, since $\theta_1^{0} + \theta_2^{0} + \theta_3^{0} + \theta_4^{0} = 0.3 + 0.0 - 0.5 + 0.0 = -0.2 \neq 0. $
We adopt the credibility matrix $\Sigma = I$, assigning equal weight to all restrictions and reflecting the absence of prior preference or differential confidence, thereby capturing model uncertainty. This specification, combined with restrictions imposed only on the structural parameters, coincides with the standard ridge formulation. The tolerance parameter is constrained to $c \in (0, c_0]$ with $c_0 = 1$, and is selected by minimizing the Stein-type risk criterion.
\paragraph{Results.} The results show several desirable properties of our proposed framework. First, the estimator consistently recovers the true data-generating parameters under model uncertainty, regardless of whether the imposed restrictions hold, as shown in Table (ref). This reflects that the regularization bias induced by potentially misspecified restrictions is effectively controlled through the optimized tolerance and subsequent debiasing step. Second, inference achieves the intended nominal coverage, with empirical coverage rates close to the 95% level and comparable to those obtained from bootstrap-based procedures (Table (ref)). Third, the plateau is clearly identified by the shaded region in Figure (ref), which corresponds to the uncertainty-adjusted cutoff. The restrictions classified within this plateau are marked with * in Table (ref). Notably, these restrictions are associated with true parameter values equal to zero and exhibit little empirical relevance for model fit. When an additional restriction holds at the population level (Case 2), it is correctly included in the plateau. In contrast, when the restriction is violated at the population level (Case 3), it falls outside the plateau, indicating that relaxing the restriction improves model fit. Thus, the plateau rule effectively distinguishes between empirically irrelevant and relevant restrictions.
In this section, we apply our proposed estimation method to the textbook Solow growth model Solow1956, Swan1956. The canonical empirical specification includes the logarithm of the saving rate and the logarithm of the effective population growth rate, capturing the accumulation and dilution channels of capital in steady state.
The Solow model implies a cross-parameter restriction that the coefficients on saving and effective population growth are equal in magnitude but opposite in sign. Specifically,
which yields the restriction $ \theta_s + \theta_n = 0. $ While this restriction captures a first-order implication of the theory, it relies on strong assumptions such as homogeneity and common adjustment dynamics. In practice, growth processes may exhibit nonlinearities and heterogeneous dynamics DurlaufandJohnson1995, Temple1999, Quah1996, DurlaufandQuah1999. To account for such model uncertainty, we consider structured nonlinear relaxations of the Solow relation:
which allow for curvature while preserving qualitative predictions. The inclusion of polynomial terms—particularly up to the cubic case—is intended for illustrative purposes, providing a simple and tractable way to capture potential nonlinearities rather than a comprehensive specification of the underlying growth process.
In specifying the set of candidate restrictions, we focus on economically meaningful hypotheses. In particular, we do not impose restrictions such as $\theta_s=0$ or $\theta_n=0$, which are inconsistent with the qualitative implications of the Solow model. Instead, we evaluate the empirical relevance of the linear Solow restriction and its nonlinear relaxations. To reflect equal prior credibility across these restrictions, we set $\Sigma = I_3$.
The estimation results are reported in Table (ref). Our approach imposes soft rather than exact restrictions, allowing the data to determine the extent to which each restriction is enforced. The selected tolerance parameter is $\hat{c} = 379.7468$, with a Stein-type risk estimate of $0.0711$, decomposed into a negligible bias proxy ($0.0003$) and a dominant variance component ($0.0708$). Consistent with this, the estimated coefficients coincide with the unrestricted benchmark, and the $R^2$ remains unchanged at $0.6272$, indicating no gain from relaxing the restriction.
At the same time, the Wald test rejects the null hypothesis $\theta_s + \theta_n = 0$ (p-value $=0.0380$), indicating that the restriction is statistically invalid. However, the ISP-based plateau rule classifies it as empirically irrelevant. This illustrates a key distinction: statistical validity concerns whether a restriction holds exactly, whereas empirical relevance concerns whether its violation matters for estimation. In this application, deviations from the Solow restriction occur in a direction that is effectively flat for the objective function, so relaxing it does not improve model fit.
In contrast, the cubic restriction $\theta_s - (-\theta_n)^3 = 0$ lies outside the plateau, indicating that it is relatively more informative than the other restrictions. This suggests that nonlinear features capture aspects of the data that are not reflected in the linear Solow restriction.
Notably, these findings highlight that empirical relevance is restriction-specific. A theoretically motivated restriction may be statistically invalid yet practically inconsequential, while alternative structured relaxations can provide relatively more informative features. Our method adapts accordingly, imposing restrictions only when they improve the bias--variance tradeoff and otherwise defaulting to the unrestricted benchmark without introducing unnecessary distortion.
This paper develops a framework for estimation and empirical relevance inference under model uncertainty with multiple restrictions. By combining Lagrangian-constrained optimization with a data-driven tolerance parameter, the estimator adapts to the bias--variance tradeoff, while a debiasing step restores valid inference. The proposed individual shadow prices (ISP) and plateau rule provide a practical way to assess which restrictions matter for improving model fit.
The framework can be naturally interpreted within a semiparametric perspective. The object of interest is a finite-dimensional target component, while potentially complex nuisance components are treated as given. This structure is compatible with modern machine learning methods, which can be used to flexibly estimate nuisance components without affecting inference on the target parameter.
Our theoretical results establish consistency and asymptotic normality, and the empirical application illustrates how the method distinguishes between statistical validity and empirical relevance. In particular, restrictions that are rejected by standard tests need not be useful for improving estimation, while alternative structured relaxations can capture relatively more informative features.
An important direction for future research is to extend this framework to high-dimensional settings by integrating it with double/debiased machine learning (DML) Chernozhukovetal2018. In such settings, nuisance components can be flexibly estimated using machine learning methods, while orthogonalization and cross-fitting ensure that inference on the finite-dimensional target remains valid. Embedding our approach within a DML framework would allow theory-driven restrictions to be imposed on the target component while accommodating high-dimensional controls, thereby preserving valid inference under model uncertainty. More broadly, this integration would enable scalable evaluation of a large set of candidate restrictions in complex, high-dimensional environments.
We thank Zheng Fang, Florian Gunsilius (Emory University), Nikolay Gospodinov (Federal Reserve Bank of Atlanta), David Drukker (Clemson University), and the audience at the Georgia Econometrics Workshop for helpful comments. All remaining errors are our own.