Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
69,387 characters · 16 sections · 50 citation commands
Orthogonality conditions for convex regression
\captionsetup[figure]{labelfont={bf},labelformat={default},labelsep=period,name={Fig.}} \captionsetup[table]{labelfont={bf},labelformat={default},labelsep=period,name={Table}}
\thispagestyle{empty} \setcounter{page}{1} \setcounter{footnote}{0} \pagenumbering{arabic} \baselineskip 20pt
Shape-constrained nonparametric regression avoids restrictive assumptions about the functional form of the regression function by building upon convexity Hildreth1954 and/or monotonicity constraints Brunk1955, brunk1958. Convexity and monotonicity constraints are particularly relevant in the microeconomic applications where the theory implies certain monotonicity and convexity/concavity properties for many functions of interest Afriat1967, Afriat1972, Varian1982, Varian1984. For example, the cost function of a firm must be monotonic increasing and convex with respect to the input prices. The recent developments in the convex regression enable researchers to impose concavity or convexity constraints implied by the theory to estimate the functions of interest without any parametric functional form assumptions. Since the development of an explicit piecewise linear characterization by Kuosmanen2008, convex regression has attracted growing interest in econometrics, statistics, operations research, machine learning, and related fields Seijo2011, Lim2012, Hannah2013, Mazumder2019, Yagi2018, Bertsimas2020. Extensions of the convex regression to the quantile and expectile estimation have proved useful for estimating marginal abatement costs of pollution emissions Kuosmanen2020b, Dai2023, Dai2025.
In the linear regression, the random error term is assumed to be uncorrelated with the explanatory variables. In the generalized method of moments (GMM) interpretation of the linear regression, this property is called the population orthogonality condition. It is easy to show that the residuals of the ordinary least squares (OLS) estimator satisfy the corresponding sample orthogonality condition, that is, OLS residuals do not correlate with any of the regressors by construction. Essential statistical properties such as unbiasedness and consistency of the OLS estimator critically depend on the orthogonality condition. If the orthogonality condition does not hold, the problem is referred to as endogeneity. The standard econometric approach is then to find instrumental variables (IVs) that are highly correlated with the regressors but uncorrelated with the error term.
The orthogonality conditions for convex regression remain unknown. While the statistical consistency of the convex regression has been proved assuming the standard orthogonality condition, we do not know how convex regression behaves if regressors are endogenous. Consequently, practitioners are unable to evaluate the suitability of convex regression in applications where explanatory variables may be endogenous, and they lack the means to address potential endogeneity bias. The contributions of this paper are three-fold:
The first contribution of this paper is to formally analyze the sample orthogonality for convex regression. Applying Lagrange duality, we establish the sample orthogonality conditions for the most essential variants of convex regression, including additive and multiplicative formulations of the regression model, with and without monotonicity and homogeneity constraints. The results of this paper are relevant to practitioners who need to understand what kind of exogeneity assumptions are required for different variants of convex regression.
The second contribution of this paper is to propose a hybrid IV control function approach integrated into convex regression. The classic two-stage least squares (2SLS) approach can be adapted to correct the endogeneity bias in convex regression, but the nonparametric IV estimation may lead to poor statistical performance due to its ill-posedness Chetverikov2017. Moreover, it is difficult to estimate convex regression with endogeneity (a pure nonparametric regression model) because the 2SLS approach is not directly transferable to a nonlinear or nonparametric setting Yatchew2003. Alternatively, the control function is commonly used to address the simultaneity bias Olley1996, Levinsohn2003, Ackerberg2015 or the measurement error in inputs Collard2016, Dong2021 when estimating the production function. In the spirit of fully utilizing the IV and control function, we propose a hybrid IV control function-based convex regression to address the potential endogeneity problem. The consistency of the proposed approach is established.
Our third contribution is to conduct three Monte Carlo experiments in the context of the classic Cobb–Douglas production function. The first experiment examines the finite sample performance of OLS and convex regression estimators under endogeneity. The second experiment compares the finite sample performance of convex regression and convex regression with the 2SLS and IV control function approaches, where the monotonicity is relaxed. The third experiment tests whether the performance of convex regression with monotonicity improves or deteriorates when the IV and IV control function approaches are introduced. An empirical application to the Chilean manufacturing production data illustrates whether the hybrid IV control function approach empirically makes a difference and whether monotonicity should be avoided in convex regression.
We further shed new light on the question: if the sample orthogonality does not hold, then how does the performance of convex regression with IV control function? In the existing literature, there is little evidence of whether the performance should improve or deteriorate after using the IV control function or even the 2SLS approach. Our simulation findings suggest that the convex regression with monotonicity constraints (i.e., the sample orthogonality condition cannot be held) can improve its performance after using the IV control function approach in the case of endogeneity.
The rest of this paper is organized as follows. Section (ref) briefly introduces the additive and multiplicative convex regression. Section (ref) examines the sample orthogonality conditions in the cases with and without monotonicity. The two-stage IV control function-based convex regression is developed in Section (ref). Sections (ref) and (ref) implement the Monte Carlo simulations and empirical application to Chilean manufacturing data. Section (ref) concludes the paper with suggested future avenues. The formal proofs and additional tables and figures are attached in the Appendix.
Consider the following multivariate additive model
where $y_i$ is the dependent variable, $\bx_i$ is a $k$-dimensional vector of explanatory variables, and $\varepsilon_i$ is a random error term with zero mean and a constant and finite variance. $f:\real_+^k \rightarrow \real$ is assumed to be a (homogeneous) concave/convex regression function with an unknown functional form Kuosmanen2008, Seijo2011, Kuosmanen2012c.
Building on the results by Afriat1967 (Afriat1967, Afriat1972), Kuosmanen2008 develops the first operational solution to the multivariate convex regression problem.\footnote{ In this paper, we focus on a concave $f$, but a convex $f$ can be modeled by reversing the sign of the Afriat inequalities Afriat1967, Afriat1972. We use the term “convex regression” because in both cases $f$ is a support of a convex set, whereas concave sets do not exist. See Liao2024 for the case of convex $f$. } Specifically, we consider the least squares estimator of multivariate convex regression
where concavity (ref) is imposed throughout the paper, but monotonicity (ref) and homogeneity (ref) are optional.
Kuosmanen2008 also proves that the optimal solution to the infinite-dimensional optimization problem (ref) is always equivalent to the optimal solution to the following quadratic programming problem with linear constraints (also referred to as the convex nonparametric least squares estimator)
where the first set of constraints (ref) restates the regression equation (ref) in terms of a piecewise linear approximation of the true but unknown regression function $f$. Kuosmanen2008 has proved that without any loss of generality, the estimated regression function can be parametrized as a piecewise linear function and characterized by the coefficients $\alpha$ and $\bbeta$.\footnote{ It is worth noting that there exist alternative parametrizations of concavity (e.g., used by Seijo2011); the advantage of the applied parametrization in this paper is that we can easily impose monotonicity and linear homogeneity as well. } The second set of constraints (ref) enforces the concavity of the piecewise linear regression function (reversing the sign of the inequality imposes convexity). The third set of constraints (ref) enables the regression function to be monotonic increasing (reversing the sign makes the function monotonic decreasing). The fourth constraint set in (ref) enforces linear homogeneity of the regression function.
\setcounter{equation}{3}
While the literature on convex regression focuses almost exclusively on the additive model introduced above, in economic applications, it is often more natural to posit a multiplicative model (Kuosmanen2012c, Kuosmanen2012c)
Taking the natural logarithm of both sides yields
where the multiplicative formulation is not just an inconsequential data transformation, because $\ln{f(\bx_i)} \neq f(\ln({\bx_i}))$ and hence the orthogonality conditions of the multiplicative formulation deserve a separate treatment.
Accordingly, the multiplicative convex regression estimator is formulated as
where the first set of constraints (ref) restates the logarithm of the multiplicative function model (ref), and the other three constraints (ref)-(ref) are also used to guarantee the concavity, monotonicity, and linear homogeneity of the regression function $f$, respectively.
In this section, we systematically analyze the sample orthogonality conditions of convex regression in various cases, including the additive and multiplicative models with and without monotonicity.
We first consider the additive convex regression problem (ref) without monotonic and linearly homogeneous assumptions (i.e., excluding constraints (ref) and (ref)). The Lagrangian function of convex regression without monotonicity and linear homogeneity is stated as
where $\lambda_{ih} \ge 0$ and $\mu_i$ are the Lagrangian multipliers (shadow prices) of the concavity constraint (ref) and regression equation (ref). Problem (ref) is convex, and the optimal solution exists and is unique. Furthermore, the optimal solution of problem (ref) satisfies the first-order conditions, which are the first-order derivatives of the Lagrangian function $\Ls$,
where $e_i$ is the estimates of $\varepsilon_i$.\footnote{ In order to keep the symbols easy to read, we avoid using other symbols for the estimated values of other variables such as $\alpha$ and $\bbeta$. } Furthermore, the optimal solution necessarily satisfies the complementary slackness conditions (i.e., the Karush-Kuhn-Tucker conditions),
Conditions (ref)-(ref) make a foundation to prove several properties of convex regression.
Seijo2011 apply Moreau's decomposition theorem to prove an equivalent result $\sum\limits_{i=1}^n y_i=\sum\limits_{i=1}^n \hat{y}_i$. The proof of Lemma (ref) is an intuitive direct proof based on the first-order conditions of convex regression. This strategy also enables us to prove the following Theorem (ref).
This result is important for understanding the underlying statistical assumptions of the convex regression estimator. Recall that the OLS estimator of the linear regression model satisfies exactly the same sample orthogonality condition. The corresponding population orthogonality condition (also referred to as the exogeneity condition) is
In other words, the random error term $\varepsilon$ is assumed to be uncorrelated with regressors $\bx$. Exactly the same exogeneity condition applies to convex regression. Hence, the convex regression estimator is subject to endogeneity problems similar to those of the OLS estimator. To address such endogeneity problems, note that the convex regression estimator can be understood as a GMM estimator when the population orthogonality condition (ref) holds. The convex regression estimator is also a maximum likelihood estimator if $\varepsilon_i$ is normally distributed, as already noted by Hildreth1954. Therefore, it is possible to apply an IV in the context of convex regression, stating the population orthogonality condition (ref) in terms of the instrument $z$ that correlates with $x$ but is uncorrelated with $\varepsilon$ (we revisit this in detail below).
We know from Seijo and Sen (Seijo2011, Lemma 2.4, part ii) the following weaker result,
Theorem (ref) implies this result, but the converse is not true. Note that the population orthogonality corresponding to Seijo and Sen's result is
which is the standard exogeneity condition in parametric nonlinear regression. Interestingly, Theorem (ref) demonstrates that convex regression also satisfies the stronger sample orthogonality condition of linear regression. It is obviously critically important to know exactly what kind of assumptions must be made before applying the estimator.
Considering next the multiplicative convex regression problem (ref) without monotonic and linear homogeneous assumptions (i.e., excluding constraints (ref) and (ref)), we have the following Lagrangian function
Denoting $\hat{y}_i = \alpha_i + \bbeta_i^\prime\bx_i$ and differentiating, we have the first-order condition (ref) and the following
Conditions (ref) and (ref) imply
In this case, it is easy to confirm that we cannot conclude a result similar to Lemma (ref). That is, the sum of residuals in the optimal solution to (ref) is not necessarily zero. However, a close conclusion is stated as
Lemma (ref) is interesting as it builds a framework to study the properties of a multiplicative error model using the ratio of regressors to the fitted value. Based on Lemma (ref), we have the following Theorem (ref) to document the sample orthogonality condition for the multiplicative model.
The corresponding population orthogonality condition is
Therefore, it is not necessary to assume that $\bx$ are uncorrelated with the random error term $\varepsilon$, but rather, we need to assume that ratios $\bx/f(\bx)$ do not correlate with $\varepsilon$. This result can guide one with the choice of good instruments and help to circumvent certain specific types of endogeneity problems (e.g., the simultaneity problem by Marschak1944 stressed by Olley1996 among others). For instance, if low productivity firms need to use more inputs $\bx$, but $f(\bx)$ increases by the same proportion, then this would not cause the potential simultaneity problem.
Suppose the regression function $f$ is concave and monotonic increasing (i.e., excluding constraints (ref) and (ref) from additive model (ref) and multiplicative model (ref), respectively). Convex regression implements monotonicity by imposing a sign constraint for the slope coefficients $\bbeta_i$.
In the case of the additive model subject to concavity and monotonicity, the Lagrangian function $\Ls$ becomes
where ${\boldeta}_i \ge 0$ are shadow prices of the monotonicity constraints.
Differentiating, we obtain the first-order optimality conditions (ref), (ref), and the following
Similarly, the first-order conditions for the multiplicative model subject to concavity and monotonicity are summarized as (ref), (ref), and the following
In this case, conditions (ref) and (ref) imply that $\sum\limits_{i=1}^n \dfrac{e_i}{\hat{y}_i}=0$, thus Lemma (ref) is valid.
In both additive and multiplicative models, if at least for one observation the monotonicity constraint is active (i.e., $\mu_{ij}>0$), then the sample orthogonality condition for convex regression must fail (Theorem (ref)).
The corresponding population orthogonality conditions for convex regression are not necessarily held. Specifically, for the additive model subject to concavity and monotonicity, we have $$\E[\varepsilon_i\bx_{ik}]\le0 \,\ \forall k, $$ and for the multiplicative model subject to concavity and monotonicity, there is $$\E\left[\varepsilon_i\dfrac{\bx}{f(\bx)}\right]\le \mathbf{0}.$$
In fact, the sample orthogonality condition holds if and only if the monotonicity constraint is redundant for all observations. However, it is unlikely in practice that the monotonicity constraints do not bind for any single observation. Even if the true function is monotonic, there are usually some violations due to the random error $\varepsilon$. When monotonicity constraints are binding, the residuals negatively correlate with the explanatory variables. This might suggest that convex regression with monotonicity constraints can accommodate specific types of endogeneity. However, the sum of shadow prices $\sum\limits_{i=1}^n\eta_{ik}$ does not seem to have any obvious econometric meaning, and it is difficult to suggest the exact population orthogonality condition.
If the sample orthogonality does not hold, then what does it imply for the assumptions regarding the true $\varepsilon$? Do we implicitly assume that $\text{Cov}(x, \varepsilon) < 0$? Is that enough? Further, what happens if the usual population orthogonality $\text{Cov}(x, \varepsilon) = 0$ does hold, but one imposes monotonicity? Does the violation of the sample orthogonality cause bias in this case? If yes, what is the direction of bias? If the sample orthogonality is violated, would using IVs to address endogeneity improve the performance of the convex regression estimator or rather make things worse? In Section (ref), we shed new light on these questions through the following Monte Carlo simulations and an empirical illustration.
Let us further assume the regression function $f$ to be concave, monotonic increasing, and linearly homogeneous in both additive and multiplicative models (i.e., problems (ref) and (ref)).
While the linear homogeneity is considered in both convex regression problems, the sample orthogonality conditions will not be altered (see Theorem (ref)).
In the additive model, at the optimal solution to problem (ref), the sum of residuals is not necessarily equal to zero, i.e., $\sum\limits_{i=1}^n e_i \neq 0$ and consequently $\sum\limits_{i=1}^n y_i \neq \sum\limits_{i=1}^n \hat{y}_i$. Unsurprisingly, this is analogous to the OLS regression through the origin (without an intercept), in which the sum of residuals is not usually zero Eisenhauer2003.
Similarly, the sum of residuals is not necessarily zero in the multiplicative model (ref), but the orthogonality condition can be held if input variables are scaled by the fitted values (i.e., $\sum\limits_{i=1}^ne_i\dfrac{x_{ik}}{\hat{y}_i}=0 \,\ \forall k$) in the case of multiplicative model subject linear homogeneity and concave.
The validity of the orthogonality condition is mainly due to the first-order condition of the slope coefficients $\bbeta$, and imposing a sign constraint on the intercept coefficient $\alpha$ does not violate the orthogonality condition. Therefore, imposing a sign constraint on the intercept parameter $\alpha$ does not violate the orthogonality condition.
Suppose that in the nonparametric additive model (ref), one of the explanatory variables, $x_1$, is endogenous, meaning either
This implies that $x_1$ is correlated with the random error term $\varepsilon$, or the conditional expectation of $y$ given $\bx$ deviates from the true underlying function $f(\bx)$. That is, the nonparametric estimator $\hat{f}(\bx)$ is inconsistent, $ \hat{f}(\bx) \overset{p}{\centernot\longrightarrow} f(\bx)$.
The classic remedy to such endogeneity bias is to utilize instrumental variables that are highly correlated with the regressors but uncorrelated with the disturbance term. In the present context of convex regression, good instruments $\bz$ should satisfy the relevant population orthogonality conditions identified in the previous section. If such instruments are available, the classic 2SLS approach can be easily adapted to the convex regression as follows. In the first stage, we regress the endogenous regressor $x_e$ on instruments $\bz$ and other exogenous variables using linear regression. In the second stage, we replace the endogenous $x_e$ by the predicted $\hat{x}_e$ obtained in the first stage.
Inspired by Yatchew2003 and Chetverikov2017, we also consider an alternative hybrid IV control function approach to alleviate the endogeneity bias in the convex regression. In this approach, the instruments are introduced by means of a control function and the two-stage estimation strategy is then applied to obtain consistent estimates in the presence of endogeneity.
Suppose there exist instruments $\bz$ such that
where $\E(\zeta | \bz) = 0$ and $\E(\varepsilon | \bz) = 0$. For instance, when estimating the regression function with measurement error in inputs, we can instrument the capital stock (i.e., endogenous variable $x_1$) with (lagged) investment (i.e., exogenous variable $z$; an alternative measure of capital) Collard2016.
We notice that in earlier studies, the control function is usually a single variable that, when added to a regression, can reduce the endogeneity of a policy variable, but in modern econometrics, the valid control function generally relies on the availability of one or more instruments Wooldridge2015. In this sense, the control function approach inherits the spirit of the IV approach but can overcome the ill-posed inverse problem in nonparametric regression. We thus practically consider an instrument and other exogenous variables as the valid control function, $\bz$, in the empirical illustration.
Suppose further that $\E(\varepsilon |x_1, \zeta) = \lambda \zeta$ and hence we have $\varepsilon = \lambda \zeta + \nu$. We need to regress (ref) using OLS to obtain residuals $\zeta$ and then rewrite the nonparametric additive model (ref) to the following semi-nonparametric model\footnote{ The same control function approach can also be applied to the nonparametric multiplicative model (ref) to address the endogeneity problem. }
where $\E(\nu | \bx, \hat{\zeta}) = 0$ and $\text{Var}(\nu | \bx, \hat{\zeta})=\sigma^2(\bx, \hat{\zeta})$. Such a two-stage estimation to correct endogeneity is also sometimes called a residual inclusion method in the literature Terza2008, Amsler2016. Compared to the 2SLS approach, the hybrid IV control function approach can not only capture the impact of measurement error in the input but also control the unobserved simultaneous bias.
To estimate $\lambda$ and $f$ in (ref), we can solve the following semi-nonparametric convex regression model
where $\hat{\zeta}_i$ is the estimated residuals by the OLS regression (ref). We consider the concave case of the regression function in the additive semi-nonparametric convex regression model (ref). The monotonicity constraints are eliminated to maintain the sample orthogonality condition, as suggested by Theorems (ref) and (ref). Similar to convex regression (ref), the first set of constraints in (ref) is the reformulation of semi-nonparametric regression (ref), and the second set of constraints guarantees the regression function $f$ to be concave.
The estimated $\hat{\alpha}_i$ and $\hat{\bbeta}_i$ in (ref) vary for each observation, but $\hat{\lambda}$ is common to all observations. Further, the semi-nonparametric convex regression model (ref) is a restricted special case of convex regression (ref), in which $\hat{\zeta} \subseteq \bx$ with $\lambda_i = \lambda_h$. The estimated $\hat{\lambda}$ in (ref) has been proved consistent, unbiased, asymptotically efficient, and converges at the rate of $O(n^{-1/2})$ Johnson2011, Johnson2012a.
Based on the OLS regression and convex regression, the consistency of the nonparametric regression function $f$ in the proposed hybrid IV control function approach can be stated as
Alternatively, we can utilize the double residual method to estimate $f$ and $\lambda$ separately Robinson1988, or even resort to the first-order differencing approach in the univariate case Yatchew2003. In addition to the parametric OLS regression, in the first stage (ref), we can also use nonparametric methods such as kernel regression and local polynomial method to estimate $\zeta$ consistently. By contrast to our linear control function, Rodseth2025 propose a fully nonparametric approach to model the control function. Nevertheless, further comparison of linear and nonparametric modeling of control functions is warranted for future investigations.
Another appealing feature of the proposed hybrid IV control function approach is that we can simply use the $t$- or $F$-test to examine whether $x_1$ is exogenous after the semi-nonparametric convex regression (ref) is applied (i.e., the null hypothesis is $H_0: \hat{\lambda} = 0$). That is, if $\hat{\lambda}$ significantly departs from zero, then the explanatory variable $x_1$ is endogenous.
In this section, we perform a Monte Carlo study to compare the finite sample performance of convex regression and OLS under endogeneity. We have three objectives in designing the following simulations. First, we investigate how large the endogeneity bias of OLS is compared to the functional form-related specification error. Second, we explore whether the 2SLS and IV control function approaches can help convex regression without monotonicity to mitigate the impact of endogeneity in various scenarios. Third, we examine whether the 2SLS and IV control function approaches can improve the performance of convex regression with monotonicity under endogeneity.
We consider the following two data-generating processes (DGPs) to generate inputs $\bx$ and output $y$ (see, e.g., Cordero2015, Rodseth2025).
where the exogenous $x_1$ and/or $x_2$ are independently drawn from a uniform distribution, $x_1/x_2 \sim U[5, 50]$, while the endogenous $x_e$ and random error $\varepsilon$ are generated according to the specific setting in each of the experiments below.
In all experiments, we estimate the OLS regression using Python/scikit-learn and convex regression using Python/pyStoNED Dai2024 with the off-the-shelf solver KNITRO (13.2). Each scenario is duplicated 1000 times to compute the root mean square error (RMSE) and bias statistics for all estimators in terms of the production function estimation. All simulations are run on Finland's high-performance computing cluster Puhti with Xeon @2.2 GHz processors, 2 CPUs, and 20 GB of RAM per task.
The endogeneity problem of OLS is well understood: if the basic assumption $\E(\varepsilon_i x_{ik})=0, \exists k$ fails, then the standard OLS estimator is biased and inconsistent. However, there is little simulation evidence on how large the endogeneity bias is compared to the specification error, for example, due to the mis-specified parametric functional form.\footnote{ The true model would be in logs, so the misspecification is due to the fact that the OLS regression is in levels. In the first experiment, both OLS regressions---specified in logarithmic and level forms---are employed to assess whether they are subject to endogeneity. } Therefore, we design the first experiment to examine the finite sample performance of OLS and convex regression estimators under endogeneity.
Similar to Mutter2013, we generate the random error $\varepsilon$ as
where the parameter $\rho \in [0, 1]$ is a prespecified correlation coefficient determining the degree of endogeneity between $x_e$ and $\varepsilon$, the random variable $W \sim N(0, 1)$, and $\text{std}(x_e)$ is the standardized endogenous variable $x_e$ with a mean of zero and a standard deviation of one. The endogenous $x_e$ is independently drawn from a uniform distribution, $x_e \sim U[5, 50]$.
In experiment 1, we consider 72 scenarios with different numbers of observations $n \in \{50, 100, 200, 400\}$, different magnitudes of endogeneity $\rho \in \{0, 0.45, 0.9\}$, different noise levels $\sigma_\varepsilon \in \{0.5, 1, 2\}$, and different numbers of inputs $k \in \{2, 3\}$. The standard OLS regression and multiplicative convex regression (ref) are applied to the designed DGPs for estimating production functions. The RMSE statistics of the alternative estimators are reported in Tables (ref) (the bias statistics are available in Table (ref) in the Appendix).
The correctly specified OLS regression (logs) yields the smallest RMSE in the case of no endogeneity in all considered scenarios and hence outperforms the multiplicative convex regression, which conforms with the findings in, e.g., Kuosmanen2008 and Tsionas2022b. However, as the degree of endogeneity increases (i.e., $\rho \rightarrow 1$), the performance of the OLS regression (logs) quickly deteriorates (see Table (ref) for $\rho=0.45$), and the RMSE of the OLS regression (logs) is notably higher than that of convex regression, irrespective of monotonicity constraints, particularly when the error standard deviation ($\sigma_\varepsilon$) is high. Yet, the performance of convex regression is relatively robust to the changes in endogeneity level. This suggests that the impact of endogeneity is more pronounced for the OLS than the convex regression estimator. Note that the misspecified OLS regression (levels) exhibits the highest RMSE across all scenarios, implying that the biases due to model misspecification outweigh those stemming from endogeneity.
When comparing the difference in RMSE between OLS and convex regression, we observe that the gap narrows in the absence of endogeneity but widens largely when endogeneity is introduced to DGPs. This observation remains unchanged across all noise levels. For instance, in the scenarios with $\sigma_\varepsilon=1$ and $k=2$, the absolute difference in RMSE between OLS (logs) and convex regression (conc) is 1.97, 4.43, and 18.04 with three given $\rho$, respectively (see Table (ref) for $\rho=0.45$). Furthermore, the RMSE results indicate that, even when alternative functional forms are considered (e.g., OLS in logs), OLS estimators still perform notably worse than convex regression methods, especially when $\rho$ is large. This suggests that, under high endogeneity, the bias from endogeneity in OLS may outweigh the bias arising from functional form misspecification.
Across all DGPs and under varying levels of noise and endogeneity, the convex regression model imposing both concavity and monotonicity constraints (conc+mon) systematically outperforms its counterpart that imposes only concavity (conc) in terms of RMSE. The performance gains from incorporating monotonicity are particularly pronounced when the variance of the random error term is large, highlighting the usefulness of monotonicity constraints. Even under conditions of severe endogeneity, the conc+mon specification yields marginally lower RMSEs, suggesting that the additional structural constraint enhances the estimator's robustness without introducing substantial bias or overfitting (cf., Table (ref)).
Several additional insights emerge from Tables (ref) and (ref). First, the bias associated with OLS (logs) is always positive, as expected, due to the non-linear nature of the logarithmic transformation. Second, for both OLS and convex regression estimators, both RMSE and bias increase with the level of noise, indicating a deterioration in estimation accuracy as the signal-to-noise ratio worsens. Third, as the dimensionality of the input space increases, the RMSE of convex regression estimators tends to rise, supporting the findings from prior work such as Dai2023 and Dai2023c, which highlight the curse of dimensionality of nonparametric estimators. Finally, in the presence of endogeneity (e.g., $\rho = 0.9$), the bias of convex regression estimators remains substantially lower than that of OLS, further reinforcing the result that convex regression exhibits greater robustness to endogeneity bias.
We further investigate the performance of both estimators for different sample sizes and for a wide range of variances of random error $\sigma_\varepsilon$ in Table (ref). In the case of moderate endogeneity, we observe that the larger $n$ for either $k=2$ or $k=3$, the better performance in terms of RMSE and bias for all methods. The average standard deviations of RMSE clearly decrease as $n$ increases. Again, the results in Table (ref) demonstrate that as $\sigma_\varepsilon$ decreases, the performance of the estimator gets better.
As mentioned in Section (ref), convex regression subject to concavity can satisfy the sample orthogonality condition. A valid instrument is then expected to improve the performance of such a convex regression estimator in the presence of endogeneity. We thus design the second experiment to compare the finite sample performance of convex regression and convex regression with the 2SLS and IV control function approaches, where the monotonicity constraint in all estimators is relaxed.
We draw the endogenous $x_e$ from the following multiplicative model $$x_{ei} = \tilde{\rho} x_{zi} \cdot \exp(\nu_i),$$ where $x_z$ is an instrument to be used in the control function and 2SLS approaches, and the parameter $\tilde{\rho}$ captures the strength of correlation between $x_z$ and $x_e$. For the individual $\varepsilon_i$ and $\nu_i$, we assume $$\left(
\right) \sim N(0, \Sigma); \Sigma = \left(
\right).$$
Since $\text{Cov}(\varepsilon, \nu) \neq 0$, it follows that $\text{Cov}(x_e, \varepsilon) \neq 0$, confirming that $x_e$ is an endogenous variable. Given that $\text{Cov}(x_z, \nu) = 0$, $x_z$ can serve as a valid instrument for $x_e$ and is randomly generated from a uniform distribution, $x_z \sim U[5, 50]$. We further set $\sigma_\varepsilon = \sigma_\nu = \sigma$ uniformly across all considered cases. In the simulations, we choose $n \in \{50, 100, 200, 400\}$, $\rho \in \{0, 0.45, 0.9\}$, $\tilde{\rho} \in \{0.45, 0.9\}$, $\sigma \in \{1, 2, 5\}$ and $k \in \{2, 3\}$. The 2SLS approach and the hybrid IV control function approach introduced in Section (ref) are applied.
Fig. (ref) depicts the RMSE statistic of three different convex regression estimators with $\sigma=2$ and $k=3$. See also the scenarios with $\sigma=1$ and $k=3$ in Fig. (ref). From the demonstrated figures, we observe that in the case of no endogeneity ($\rho=0$), convex regression generally performs best in three estimators, but the performance difference between convex regression and convex regression with IV control function is quite small, particularly in the large sample (e.g., $n=400$). Regarding the IV-based convex regression, its performance deteriorates significantly when the sample size exceeds 100.
After introducing a correlation between the input and the random error (i.e., $\rho>0$) to the data space, the RMSE results of convex regression clearly indicate the presence of an endogeneity bias, similar to the findings in Tables (ref) and (ref). While convex regression subject to concavity can satisfy the sample orthogonality condition, it is also not immune to endogeneity bias. As shown in Fig. (ref), the IV-based and IV control function-based convex regression estimators can both be used to address the potential endogeneity problem, but the latter is more effective. Specifically, the RMSEs of convex regression with IV and IV control function approach are less than those of convex regression in the cases of $\rho=0.45$ and $\rho=0.9$, and the RMSE of IV-based convex regression is also higher than that of IV control convex regression. Furthermore, the strong instrument does not appear to reduce RMSE values significantly in both IV-related approaches.
We next evaluate the performance of estimators from the bias perspective. Overall, the bias results presented in Fig. (ref) align with the RMSE findings shown in Fig. (ref). In scenarios where $n=200$ or $n=400$, the bias of the convex regression with the IV approach is negative, while the biases of both convex regression and convex regression with the IV control function approach are positive across all considered scenarios, as expected. However, bias should always be positive in the designed multiplicative error structure. Furthermore, for the large sample size, the bias of the IV control function approach is closer to zero compared to the 2SLS approach. Therefore, we would stress that the proposed hybrid IV control function approach is more suitable to address the endogeneity problem in the context of nonparametric convex regression.
Section (ref) shows that the sample orthogonality condition cannot hold in the case of monotonicity, and the exact population orthogonality condition is also unknown. We thus design experiment 3 to test if the performance of convex regression with monotonicity improves or deteriorates when the IV and IV control function approaches are introduced. The same DGPs and scenario settings as Experiment 2 are considered.
Fig. (ref) plots the RMSE statistic of three convex regression estimators in the case of monotonicity for the low noise and three-input setting (see Fig. (ref) for $k=2$). When the correlation coefficient $\rho=0$, convex regression outperforms the other two IV-based convex regression methods, as indicated by its lower RMSE. When there is endogeneity between $x_e$ and $\varepsilon$ in the DGPs, the IV control function convex regression produces the smallest RMSEs while accounting for monotonicity. The 2SLS-based convex regression performs better than convex regression in cases of severe endogeneity in $x_e$ ($\rho=0.9$) or more inputs (Fig. (ref)). Overall, while the sample orthogonality condition cannot be satisfied in the case of monotonicity, the IV-based or the IV control function-based convex regression can, to some extent, help to address the endogeneity problem, but the performance of the IV control function approach is more robust.
A further comparison of Figs. (ref) and (ref) provides evidence that imposing monotonicity constraints in convex regression slightly decreases RMSEs and improves its performance across all scenarios, irrespective of whether IV or IV control function approaches are applied. This finding persists across various degrees of endogeneity and sample sizes, and also confirms the result discussed in Table (ref). While imposing monotonicity constraints in convex regression theoretically violates the sample orthogonality condition, it often leads to improved performance in terms of RMSE. The better performance can be explained by several factors. When the DGP follows a monotonic and convex form, such as the Cobb-Douglas function, imposing monotonicity constraints brings the estimator closer to the true structure and reduces misspecification risk. In finite samples, the bias from violating orthogonality is often limited, while the variance reduction from the constraint can lead to better overall accuracy.
The existing studies typically rely on a control function to address endogeneity concerns in estimating the production function Olley1996, Levinsohn2003, Ackerberg2015. It is thus worth examining whether the hybrid IV control function approach empirically makes a difference in convex regression. Furthermore, excluding the monotonicity assumption from convex regression can guarantee its sample orthogonality conditions. It is also interesting to investigate whether monotonicity is necessary in the endogeneity correction.
To explore the above two empirical objectives, we apply the proposed hybrid IV control function to the Chilean manufacturing production data, covering all firms with more than 10 employees from 1979--1996. Such census data conducted by Chile's National Institute of Statistics has been widely used in works such as Levinsohn2003, Gandhi2020, and Tsionas2022b.
We restrict our estimation sample to a subset of five manufacturing industries in the years 1995 and 1996, with ISIC codes 311 (Food Products), 321 (Textiles), 322 (Apparel), 331 (Wood Products), and 381 (Metals). For each industry, the sample data include the real gross output (measured as deflated revenues), labor (a weighted sum of blue- and white-collar workers), capital (measured by the perpetual inventory method), intermediate inputs (a sum of raw materials, energy, and service expenditures), and investment. A more detailed description of the data and the selected variables is available in Gandhi2020.
We start by testing if the endogeneity problem exists in production function estimation with Chilean manufacturing data. In doing so, we use the lagged investment as the instrument for capital input in the hybrid IV control function approach. Given the potential heterogeneity among manufacturing industries, we conduct separate estimations for each industry. The endogeneity test results in Table (ref) indicate that capital input is an endogenous variable for all industries, at least at the 90% significance level, regardless of whether monotonicity is relaxed. To correct the endogeneity bias in production function estimation, we need the hybrid IV control function approach.
We next report the median coefficients across observations for each industry in Table (ref), where the 25th and 75th percentiles of the coefficients are shown in parentheses. For a thorough comparison, we discuss the estimates by the convex regression with monotonicity, the 2SLS-based convex regression with monotonicity, and the IV control function-based convex regression with and without monotonicity in Table (ref).
Compared to convex regression, the 2SLS-based convex regression leads to higher capital coefficients in all industries. For instance, the capital coefficients increase from a median of 0.004 to 0.012 for the Food Product industry and 0.002 to 0.01 for the Wood Products industry, respectively. This suggests that instrumenting capital input with lagged investment may yield a higher capital estimate, conforming with the findings in Collard2016, even though a parametric estimation is employed.
After controlling for the impacts of simultaneity bias and mismeasurement in input, we again observe the higher median estimates for capital in all industries, regardless of whether monotonicity is considered. Likewise, for instance, the capital coefficients increase from a median of 0.004 to 0.006 (the IV control function approach with monotonicity) for the Food Product industry and 0.002 to 0.003 for the Wood Products industry, respectively. Furthermore, the capital median coefficients estimated by convex regression without monotonicity are larger than those by convex regression with monotonicity. The higher capital estimates may suggest that the IV control function is more effective in the absence of monotonicity, providing support for Theorems (ref) and (ref).
Regarding the labor estimate, we find that the median value decreases in all industries but Apparel, which seems to offset the increase in capital after using the IV control function approach or even the 2SLS estimation. For example, the median labor coefficient decreases from 0.029 by convex regression with monotonicity to 0.027 by the IV control function-based convex regression without monotonicity in the Food Products industry.
This paper reviews and explains the sample orthogonality condition for the most standard specifications of convex regression. The orthogonality condition is critically important for understanding possible sources of endogeneity bias. The approach applied in this paper---leveraging the Lagrangian duality theory and orthogonality analysis---offers a potential framework for testing endogeneity of the nonparametric shape-constrained regression estimators.
Our analytical results indicate that the validity of the sample orthogonality condition in convex regression is not determined by the presence of a homogeneity constraint. The validity of the sample orthogonality condition depends solely on the first-order condition of the slope coefficients of the explanatory variables (i.e., $\bbeta_i$) and not on the intercept parameter (i.e., $\alpha_i$). Thus, a sign constraint on the intercepts of the regression hyperplanes does not affect the sample orthogonality condition of convex regression.
Imposing a sign constraint on the slope coefficients alters the sample orthogonality condition. Thus, a convex regression problem that includes a monotonicity condition violates the orthogonality condition. The covariance between the regression residuals and the explanatory variables is negative. Intuitively, this means that for higher values of the independent variables, more observation points fall below the estimated regression function.
The results for the convex regression with a multiplicative error term indicate that the usual sample orthogonality condition does not hold. However, when the explanatory variables are normalized by the fitted value (i.e., $\bx/f(\bx)$), then the sample orthogonality condition holds. Analogous to the additive model, the special case of the sample orthogonality condition is valid for convex regression with or without a linear homogeneity constraint. By imposing a monotonicity constraint, the sample orthogonality condition does not hold in the multiplicative model either.
Since the convex regression is not immune to endogeneity shown by the sample orthogonality condition analysis, we propose a hybrid IV control function approach to alleviate the endogeneity bias. The simulation results confirm the superiority of the proposed approach in endogeneity correction in comparison with the traditional 2SLS approach. The monotonicity is helpful in improving the accuracy of regression function estimation, even though it could introduce bias. The empirical application also suggests that the hybrid IV control function approach makes a difference in convex regression to mitigate the impact of endogeneity and that imposing monotonicity constraints can be useful in practice.
The authors acknowledge CSC – IT Center for Science, Finland, for computational resources.
\baselineskip 12pt
\baselineskip 20pt