arXiv 6 Jun 2026 · Econometrics
arXiv:2606.07984 · PDF · DOI · OpenAlex · Extracted main text
This study investigates a statistical property of Lagrange multipliers in constrained Maximum Likelihood Estimation (MLE) and Least Squares (LS) problems from the perspective of numerical optimization. Building on large-sample theory, we show that the associated Lagrange multipliers converge to zero as the sample size increases, provided the distribution is correctly specified in MLE or the residuals are normally distributed in LS. Although this asymptotic behavior has long been recognized in statistics, it has received little explicit attention in numerical optimization and has rarely been exploited in algorithmic design. Importantly, the insight extends beyond classical low-dimensional settings: even in modern high-dimensional applications, such as deep learning, where the number of parameters may exceed the sample size, the same reasoning applies provided the generalization performance is good. This observation has two main implications. First, many constrained optimization algorithms, including the Augmented Lagrangian Method, Sequential Quadratic Programming, and Interior Point methods, require initial values for the multipliers, and choosing zero is statistically justified. Numerical experiments for constrained regressions and dynamic discrete choice model estimations support this implication by showing that initializing multipliers at zero usually lead to stable and efficient performance. Second, penalty-based approaches that convert constrained problems into unconstrained ones can perform well when the true multipliers are small. This helps explain why penalty-based methods often perform well in practice.
appendix boundary found by appendix_command · 81% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Lu, Lu and Pestourie, Raphael and Yao, Wenjie and Wang, Zhicheng and… (2021) Physics-informed neural networks with hard constraints for inverse design | 1.000 | 5 | 3 | 100% |
| 2 | Na, Sen and Mahoney, Michael (2025) Statistical inference of constrained stochastic optimization via sketched sequential quadratic programming | 0.928 | 5 | 3 | 80% |
| 3 | Aitchison, John and Silvey, SD (1958) Maximum-likelihood estimation of parameters subject to restraints | 0.811 | 4 | 2 | 100% |
| 4 | Rust, John (1987) Optimal replacement of GMC bus engines: An empirical model of Harold Zurcher | 0.737 | 4 | 3 | 50% |
| 5 | Gourieroux, Christian and Monfort, Alain (1995) Statistics and econometric models | 0.644 | 2 | 2 | 100% |
| 6 | Su, Che-Lin and Judd, Kenneth L (2012) Constrained optimization approaches to estimation of structural models | 0.644 | 2 | 2 | 100% |
| 7 | Aguirregabiria, Victor and Mira, Pedro (2002) Swapping the nested fixed point algorithm: A class of estimators for discrete Markov decision models | 0.405 | 1 | 1 | 100% |
| 8 | Basir, Shamsulhaq and Senocak, Inanc (2023) An adaptive augmented lagrangian method for training physics and equality constrained artificial neural networks | 0.405 | 1 | 1 | 100% |
| 9 | Berahas, Albert S and Curtis, Frank E and Robinson, Daniel and Zhou,… (2021) Sequential quadratic optimization for nonlinear equality constrained stochastic optimization | 0.405 | 1 | 1 | 100% |
| 10 | Berahas, Albert S and Curtis, Frank E and O’Neill, Michael J and Rob… (2024) A stochastic sequential quadratic optimization algorithm for nonlinear-equality-constrained optimization with rank-deficient Jac… | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 30 scored citations.