EconBase
← Back to paper

Semiparametric Distribution Regression with Instruments and Monotonicity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

34,692 characters · 11 sections · 36 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric Distribution Regression with Instruments and Monotonicity

\newtheorem{theorem}{Theorem} \newtheorem{definition}{Definition} \newtheorem{lemma}{Lemma} \newtheorem{assumption}{Assumption} \theoremstyle{definition} \newtheorem{example}{Example} \newtheorem{remark}{Remark} \parindent 0cm

{

abstractThis paper proposes IV-based estimators for the semiparametric distribution regression model in the presence of an endogenous regressor, which are based on an extension of IV probit estimators. We discuss the causal interpretation of the estimators and two methods (monotone rearrangement and isotonic regression) to ensure a monotonically increasing distribution function. Asymptotic properties and simulation evidence are provided. An application to wage equations reveals statistically significant and heterogeneous differences to the inconsistent OLS-based estimator.

JEL Classification: C26, J30

Keywords: Control function, endogeneity, isotonic regression, wage equations }

\doublespacing

Introduction

The semiparametric distribution regression (DR) model introduced by foresiperacchi:1995 has become a popular model for conditional distributions if other quantities than only the conditional expectation are of interest. An important feature of this model is that no distribution assumptions on the response are made, e.g. $Y$ is not assumed to be normally distributed, conditionally on covariates. At the same time, the model provides interpretable functional forms between the regressors and the outcomes, while estimating the conditional response distribution semi-parametrically. From the estimated distribution function, quantiles could be directly obtained by inversion.

One typical application are conditional wage distributions, where upper or lower quantiles are supposed to be modelled. chernozhukov:2013 and rothewied:2013 show that the DR model might be better suited than quantile regression for handling certain characteristics of wage data such as genuine point masses in the distribution of wages, nonlinearities around the minimum wage and rounding effects. The appealing property is that e.g. censoring points do not have to be included ex ante as in the case of censored quantile regression, but are detected by the estimation itself. chernozhukov:2013 show how the model can be used for estimating counterfactual distributions, rothewied:2020 propose a method for estimating conditional densities and quantile partial effects in this model. See also koenker:2013 for a comparison of quantile and distribution regression.

A restriction of the literature up to now is that the regressors are assumed to be exogenous. For example, rothewied:2013 consider a version of Mincer's earnings function by explaining the logarithmic wage with the years of education and the years of experience among others, not taking into account that, for example, the years of education might be an endogenous regressor. This does not mean that the DR estimates in such approaches are not useful. They do estimate conditional distribution functions consistently, but there is no control for unobserved confounders. A particular value of the years of education is correlated to some degree of ability or motivation of employee, so that one gets the distribution only for a subset of the population. The novelty of the present approach is the control for confounders, so that we get a clearer picture of the population.

There are some recent papers on DR estimation with endogenous regressors. sanchez:2020 consider DR estimation based on instrumental variables, but they use parametric models based on splines among others. chernozhukov:2022 discuss semiparametric DR models in the context of sample selection.

The present paper proposes IV-based estimators for the semiparametric DR model. Taking into account that the DR model is fitted by pointwise estimators of simple binary outcome models, we adapt consistent estimators for binary outcome models with endogenous regressors. On the one hand, we consider maximum likelihood estimation, which is asymptotically efficient, on the other hand, we propose a computationally better tractable three-step estimator. For both estimators, consistency and convergence to Gaussian limit processes are proved. As these estimator are unconstrained, monotonicity is not guaranteed. We discuss two methods for enforcing monotonicity in a second step, monotone rearrangement and isotonic regression.

In the following, we first present the model (Section 2), then the estimation procedures including a causal interpretation and asymptotic results (Section 3). Afterwards. we consider the monotonizing methods (Section 4) and some simulation evidence (Section 5). An application to a Mincer-type wage regression (Section 5) demonstrates the importance in empirical practice to take endogeneity into account for estimating DR models and to use the new method. Section 6 makes some suggestions for future research.

Model

Consider an outcome variable $Y$ and regressors $X_1,\ldots,X_k$. In the semiparametric DR model, the conditional distribution function of $Y$ given the set of regressors $X$ is modelled by $F_{Y|X}(y|x) = \Lambda(x'\beta(y))$ for some link function $\Lambda$ such as the distribution function of the standard normal distribution $\Phi$ and some function $\beta(y)$. Although the function $\Lambda$ must be chosen in advance, in this note, the model is called semiparametric: Usually, in the literature, no explicit restrictions on $\beta(y)$ such as continuity are imposed and there is a parameter for every $y$. Anyway, it is clear that $\beta(y)$ has to be chosen such that $F_{Y|X}(y|x)$ fulfills the properties of a conditional distribution function if the model is correctly specified. Given $x$, the function must be monotonically increasing and converge to $1$ and $0$ for $y \rightarrow \infty$ and $y \rightarrow -\infty$, respectively.

In the simple linear model $Y = 1 + X + U$, where $U$ is distributed with distribution function $\Lambda$ and independent of $X$, it holds $F_{Y|X}(y|x) = \Lambda(y-1-x)$ such that $\beta(y) = (y-1,-1)$ and the monotonicity condition is fulfilled. If, for example, $Y$ describes earnings of employees and there is minimum wage at some value $y^*$, the function $\beta(y)$ might contain a discontinuity point at $y^*$.

Based on an i.i.d. sample of length $n$, $F_{Y|X}(y|x)$ can be consistently estimated by maximum likelihood estimation similarly as a probit model would be estimated, for example. The estimation is performed separately for each $y$ and requires that the regressors are exogenous. To be precise, one introduces the indicator functions $I_y := 1\{Y \leq y\}$ with $E(I_y|X=x) = \Lambda(x'\beta(y))$. The model can be interpreted as a latent variable model with $I_y = 1$ if $I_y^* := X'\beta(y) \geq U$ and $I_y = 0$ otherwise. The random variable $U$ is distributed with distribution function $\Lambda$ and exogeneity means that $X$ is independent of $U$.

Estimation Procedure

There are different IV-based approaches for estimating binary outcome models if $X$ and $U$ are not independent. We present two adaptions for the case of estimating such models separately for each $y$. The first one is based on the maximum likelihood estimator, which is asymptotically efficient for fixed $y$. As this estimator is computationally demanding, we then propose a three-step estimator. This is an adaption of the two-step estimator introduced by riversvuong:1988, which, for fixed $y$, is also asymptotically efficient in the case of just identified models.

Maximum Likelihood Estimator

The original maximum likelihood estimator was brought forward and disussed by amemiya:1978, newey:1987, riversvuong:1988, is explained in detail in wooldridge:2002, Section 15.7.2 and hansen:2022, Section 25.12., and is implemented in Stata (command ivprobit). The estimator is applicable if $\Lambda$ is equal to the distribution function of a standard normal distribution $\Phi$ and if the endogenous regressors are continuously distributed. We focus on the case of one endogenous regressor. Using the notation from the literature, $X$ denotes a $k$-dimensional vector with exogenous regressors, $Y_2$ is scalar, endogenous and continuously distributed and $Z$ is a $l$-dimensional vector with exogenous instruments.

In our situation, the goal is to estimate $E_V(I_y|X=x,Y_2=y_2) := P_V(Y \leq y| X=x,Y_2=y_2) := E_V(I_y|X=x,Y_2=y_2,V=v)$, i.e. we first control for possible confounders (such as ability/motivation) and integrate these confounders out afterwards. Roughly spoken, $Y_2$ is made exogenous by this integration. Standard probit would be a suitable estimator for the conditional distribution function $E(I_y|X=x,Y_2=y_2)$.

The assumption is that $E_V(I_y|X=x,Y_2=y_2) = P_V(I_y^* \geq 0|X=x,Y_2=y_2)$ with

eqnarray[eqnarray omitted — 124 chars of source]

Here, the random variables $U(y)$ and $V$ are jointly normally distributed conditionally on $X$ and $Z$,

equation*[equation* omitted — 207 chars of source]

Note that the variance of $U$ is set to $1$ because an additional parameter would not be identified. With the explanations above, $P_V(Y \leq y | X=x,Y_2=y_2) = \Phi(x'\beta_1(y) + y_2 \beta_2(y))$.

remarkTo get more intuition on this, note that one can write $U(y) = \rho V + \epsilon(y)$, where $\epsilon \sim \mathcal{N}(0,\sigma_\epsilon^2)$ is independent from $V$, $\rho := \frac{\sigma_{12}}{\sigma_2^2}$ and $\sigma_\epsilon^2 := 1-\rho^2\sigma_2^2$. Then we have $I_y^* = X'\beta_1(y) + Y_2 \beta_2(y) + \rho V + \epsilon(v)$ and $E(I_y^*|X=x,Y_2=y_2,V) = x'\beta_1(y) + y_2 \beta_2(y) + \rho V$- The expectation of the latter term with respect to $V$ is $x'\beta_1(y) + y_2 \beta_2(y)$ because $E(V)=0$. In contrast to this, assuming joint normality of $(U(y),V,X,Z)$, $U(y) = \psi Y_2 + C(y)$ for $\psi = \frac{\sigma_{12}}{\sigma_{Y_2}^2} = \frac{\rho \sigma_2}{\sigma_{Y_2}^2}$, where $C(y)$ is independent of $Y_2$ and $\sigma_{Y_2}^2 = Var(Y_2)$. Then, similarly as in li:2022, \begin{equation*} P(Y \leq y | X=x,Y_2=y_2) = \Phi\left(\frac{x'\beta_1(y) + \left(\beta_2(y) + \psi \right) y_2}{\sqrt{1-\tau^2}}\right) \end{equation*} with $\tau = \frac{\sigma_{12}}{\sigma_{Y_2}} = \frac{\rho \sigma_2}{\sigma_{Y_2}}$. This is the quantity that standard probit would estimate, but this is not the quantity we are interested in. $\Box$

This means that consistent estimation of the parameters in (ref) leads to consistent estimation of $E_V(I_y|X=x,Y_2=y_2)$ by the continuous mapping theorem.

For estimation purposes, we use the i.i.d. sample $(I_{y,i},Y_{2,i},X_i,Z_i)$ with $I_{y,i} = 1\{Y_i \leq y\}$. Similarly as in hansen:2022, the likelihood is derived by factorizing the joint density of $I_y$ and $Y_2$. The log-likelihood is then essentially the sum of the standard regression and the standard probit log-likelihood. It is given as $L_y(\theta(y)) = \sum_{i=1}^n L_{y,i}(\theta(y))$ with the parameter vector\footnote{Note that $\sigma_{12}$ and $\sigma_\epsilon^2$ can be calculated from the other parameters.} $\theta(y) := (\beta_1(y),\beta_2(y),\gamma_1,\gamma_2,\rho,\sigma_2^2)$ and

eqnarray*[eqnarray* omitted — 357 chars of source]

It holds $\mu_{y,i}(\theta(y)) = X_i'\beta_1(y) + Y_{2,i} \beta_2(y) + \rho(Y_{2,i} - X_i'\gamma_1 - Z_i'\gamma_2 )$ and $\sigma_\epsilon = \sqrt{1-\rho^2\sigma_2^2}$.

For each $y$, the parameter estimator can be equivalently calculated by maximizing the likelihood function or by minimizing some norm of the score function. Thus, the estimator falls into the framework of Z-estimators analyzed in chernozhukov:2013 and one can derive consistency and asymptotic normality, both pointwisely and uniformly in $y$. The estimator for the conditional distribution function\footnote{For better readability, we do not write $F_V$ in the following, which would be the coherent notation.} $F_{Y|X,Y_2}(y|x,y_2) := \Phi(x'\beta_1(y) + y_2 \beta_2(y))$ is given by $\hat F_{ML,Y|X,Y_2}(y|x,y_2) := \Phi(x'\hat \beta_1(y) + y_2 \hat \beta_2(y))$. This is a continuous transformation and the limit results carry over by means of the functional delta method. Then, under some additional assumptions as described in the Appendix, we obtain

theoremLet Assumption 1 be fulfilled. Then it holds that \begin{equation} \sqrt{n}\left(\hat F_{ML,Y|X,Y_2}(\cdot|x,y_2)- F_{Y|X,Y_2}(\cdot|x,y_2) \right) \end{equation} converges to a Gaussian process $\mathbb{G}_1(\cdot)$ on a compact subinterval of $\mathbb{R}$.

The proof of this theorem shows that the limit process depends both on the limit properties of $\hat \theta(y)$ and the shape of the function $F_{Y|X,Y_2}(y|x,y_2)$.

Three Step Approach

The two step estimator by riversvuong:1988 is explained in detail in wooldridge:2002, Section 15.7.2. The estimator requires a non-trivial adjustment to our situation, however, because it does not directly estimate the parameters $\beta_1(y)$ and $\beta_2(y)$ consistently. The numerical calculations later on are performed with this estimator because it works much faster and the problem of boundary solutions does not appear. The maximum likelihood estimator might be used in special situations, in which one prefers to estimate the conditional distribution function only for a single point $y$, say.

The setup is similar to the one in the former subsection and the estimator is based on the decomposition $U(y) = \rho V + \epsilon(y)$ which leads to the equation $$I_y^* = X'\beta_1(y) + Y_2 \beta_2(y) + V \frac{\rho}{\sigma_2} + \epsilon(y).$$ The error term $\epsilon(v)$ is independent of $X$, $Y_2$ and $V$ and is $N(0,1-\rho^2)$-distributed. This means that, for fixed $y$, standard probit estimation would consistently estimate the parameters $\tilde{\beta}_1(y):= \frac{\beta_1(y)}{\sqrt{1-\rho^2}}, \tilde{\beta}_2(y) := \frac{\beta_2(y)}{\sqrt{1-\rho^2}}, \tilde{\rho} := \frac{\rho}{\sigma_2\sqrt{1-\rho^2}}$ under the same assumptions as discussed after Assumption 1 in the Appendix. As $V$ is not observable, this term is replaced with the residuals of an OLS regression of $Y_2$ on $X$ and $Z$.

For fixed $y$, this is the two step estimator by riversvuong:1988. In the case of just identified models (one instrument for the endogenous regressor), this estimated is even numerically equal to the maximum likelihood estimator for $\tilde{\beta}_1(y),\tilde{\beta}_2(y),\tilde{\rho}$, so that Theorem (ref) can be directly applied to this.\footnote{This is also true for the AGLS estimator from amemiya:1978, which is implemented in the R-package ivprobit.}

The parameter $\rho$ is not known, so that these estimators cannot be used directly. However, they can be used to consistently estimate the conditional expectation $$E(I_y|X=x,Y_2=y_2,V=v) = \Phi(x'\tilde{\beta}_1(y) + y_2 \tilde{\beta}_2(y) + v \tilde{\rho}).$$ As $E(I_y|X=x,Y_2=y_2) = E_V(I_y|X=x,Y_2=y_2,V)$ by the law of iterated expectations, a consistent estimator for $E'(I_y|X=x,Y_2=y_2)$ is given by $$\hat F_{Y|X,Y_2}(y|x,y_2) := \frac{1}{n} \sum_{i=1}^n \Phi(x'\widehat{\tilde{\beta}_1(y)} + y_2 \widehat{\tilde{\beta}_2(y)} + V_i \widehat{\tilde{\rho}}),$$ where $V_i$ are the residuals of an OLS regression of $Y_{2i}$ on $X_i$ and $Z_i$, $i=1,\ldots,n$.

theoremLet Assumption 1 be fulfilled and consider the case of a just identified model. Then it holds that \begin{equation} \sqrt{n}\left(\hat F_{Y|X,Y_2}(\cdot|x,y_2)- F_{Y|X,Y_2}(\cdot|x,y_2) \right) \end{equation} converges to a Gaussian process $\mathbb{G}_2(\cdot)$ on a compact subinterval of $\mathbb{R}$.

Due to the discretization in the estimation, the estimator $\hat F_{Y|X,Y_2}(y|x,y_2)$ can attain at most $n$ different values for fixed $X$ and $Y_2$. The differences arise at the different outcomes $Y_i$, so that it is reasonable to evaluate the estimated distribution function at all $Y_i$, if computationally feasible.

Monotonicity

While the proposed estimators from the last section are consistent under appropriate assumptions, there is no reason to assume that the estimated conditional distribution functions are monotonically increasing in $y$ in finite samples. This might be a drawback for interpretation purposes, e.g. if the estimators are used for calculating conditional quantiles and it turns out that the estimated $90\%$-quantile is smaller than the estimated $80\%$-quantile. We discuss two methods to fix this, monotone rearrangement as well as isotonic regression.\footnote{foresiperacchi:1995 discuss in their Section 2.1 some other possibilities to get monotonous estimators of the distribution function, but do not elaborate on them in more detail.} While the former has well-known asymptotic properties, the latter is computationally more appealing. A simulation study reveals that both approaches share similar properties in terms of the mean squared error.

Monotone Rearrangement

chernozhukov:2010 propose a monotone rearrangement approach, mainly for quantile regression in order to ensure that estimated conditional quantiles do not cross. As discussed in chernozhukov:2013, this approach can also be applied to distributional regression. It is based on the identity

equation[equation omitted — 116 chars of source]

so that in a first step the conditional quantile function needs to be estimated, before it is appropriately integrated. This leads to the estimator $$\tilde F_{Y|X,Y_2}(y|x,y_2) = \int_0^1 \mathbf{1}\{\hat Q_{Y|X,Y_2}(u|x,y_2) \leq y\}du$$ with the estimated conditional quantile function\footnote{A researcher only interested in conditional quantiles could of course directly use this estimator.} $\hat Q_{Y|X,Y_2}(u|x,y_2) = \mathsf{inf}_y \{\hat F_{Y|X,Y_2}(y|x,y_2) \geq u \}$. The asymptotic properties of this estimator are well understood. As discussed in chernozhukov:2010, given a result like Theorem (ref) from the last section, the convergence rate (in our case $\sqrt{n}$) carries over due to the Hadamard differentiability of the operator from (ref) and an application of the functional delta method. Moreover, it is possible to estimate the limit process by a bootstrap approximation.

Isotonic Regression

An alternative to the monotone rearrangement is the application of an isotonic regression, which can be applied directly on the functional estimator. This estimation procedure is discussed in barlow:1972 and robertson:1988, for example. By construction, the estimated distribution function only changes its value at the observed $Y_1,\ldots,Y_n$ and is constant between these points. The idea is to replace the points $\hat F_i := \Phi(x'\hat \beta_1(Y_i) + y_2 \hat \beta_2(Y_i))$ by points $\tilde{\tilde{F_i}}$ that are close to $\hat F_i$, but fulfill the monotonicity restriction. This means that one solves the quadratic minimization problem

equation*[equation* omitted — 129 chars of source]

under the constraint

equation[equation omitted — 105 chars of source]

The problem can be solved numerically with the {\it pool adjacent violators algorithm}, an implementation in software packages such as R (command isoreg in the package stats) is available. The computational complexity for given $n$ is $O(n)$ for already sorted data, see best:1990. So, the approach is less complex than the monotone arrangement, where the quantile function has to be used and where integrals have to be solved. A potential drawback is the tendency to obtain flat functions, which leads to a bias in finite samples, if the true distribution function is strictly increasing.

Having obtained a monotonically increasing distribution function for the points $Y_1,\ldots,Y_n$, forecasts for other values of $y$ might be obtained by linear interpolation, for example. Also conditional quantiles can be calculated in this way.

If (ref) already holds for the $\hat F_i, i=1,\ldots,n,$ $\sum_{i=1}^n \left(\tilde{\tilde{F_i}} - \hat F_i \right)^2$ is equal to $0$. So, it is intuitive that the monotonized estimator is consistent if the true conditional distribution function is monotonically increasing and the estimated distribution function is uniformly consistent (over $y$).

In other contexts, the convergence rate of isotonic regression is smaller than $\sqrt{n}$, for example $n^{1/3}$ in abrevaya:2005. In these cases, the standard bootstrap (drawing with replacement) might behave erratic, see patra:2018. Monte Carlo evidence in the following subsection suggests that such 5problems should not expected the present context, at least not for the setting considered in the empirical application. The intuition is that in our case, the isotonic regression is just a finite-sample correction in second step of an estimator which asymptotically fulfills the monotonicity restriction.

Simulations

We simulate from the model $Y^* = \mathsf{max}(2,\tilde Y)$ and

eqnarray*[eqnarray* omitted — 59 chars of source]

where $X$ and $Z$ are i.i.d. $N(0,1)$-distributed and $(U,V)$ is bivariate normally distributed with zero mean and covariance matrix $

pmatrix[pmatrix omitted — 35 chars of source]

$. Here, $X$ represents the exogenous regressor, $Y_2$ the endogenous regressor, which is correlated with $U$, and $Z$ the exogenous instrument. We consider a censored $\tilde Y$, which mimics the application of modelling wages with a minimum wage and which shall highlight the appealing property of distributional regression of detecting such censoring points. In this case, $F_{Y|X,Y_2}(y|x,y_2) = \Phi(y-1-x-y_2)$ for $y \geq 2$ and $0$ elsewhere. We fix $\rho=0.7$ and calculate $\tilde F_{Y|X,Y_2}(y|x,y_2)$ (monotone rearrangement) and $\tilde{\tilde{F}}_{Y|X,Y_2}(y|x,y_2)$ (isotonic regression) for $x=y_2=1$ and $x=y_2=2$. As $E(X)=0$ and $E(Y_2)=1$, $X$ and $Y_2$ are further away from their expectations in the latter case. The grid points for $y$ are equidistant in the interval $[1,5]$ with $50$ grid points in total. For the rearrangement, the quantile levels are equidistant in the interval $[0.01,0.99]$ with $99$ grid points in total. To mimic the setting of the empirical application, the sample sizes are $n=100,200,400$. For each case, $1000$ Monte Carlo replications are performed. The results are compared with the standard probit estimates that ignore the endogeneity.

As we are concerned with uniform convergence to the true function (see Theorem (ref)), we consider the average squared bias, the average variance and the average MSE of $\tilde F_{Y|X,Y_2}(y|x,y_2)$ and $\tilde{\tilde{F}}_{Y|X,Y_2}(y|x,y_2)$ over the grid of $41$ y-values. Tables (ref) shows the results.

center[center omitted — 44 chars of source]
table[table omitted — 1,402 chars of source]

With the IV approach, the average MSE is dominated by the variance and is similar for both procedures with a slight advantage for the isotonic regression. Bias, variance and MSE are slightly higher if $x$ and $y_2$ are further away from their expectations and halve when the sample size is doubled. This suggests that, in this setup, the convergence rate of both estimators is $\sqrt{n}$. The variance of the OLS approach also halves with doubled sample size and slightly exceeds that of the IV approach for $x=y_2=1$, but is considerably biased as expected. So, its MSE is much higher than that of the IV approach.

Application to Wage Equations

We revisit wage data from mroz:1987 with $n=428$ individuals, who were working in 1975, and estimate a Mincer-type regression to estimate the returns of education. To be precise, the logarithmic hourly wage is explained by the years of education and the years of working experience (the latter both linearly and quadratically). The variable years of education is assumed to be endogenous as it might be correlated with unobserved variables such as ability or motivation. While there might be some correlation with the years of working experience as well, we assume that other influences are more relevant in that case, so that we assume this variable to be exogenous.

A possible instrument for the years of education is the years of education of the mother. In this dataset, the first stage $F$-statistic is given by approximately $75$ so that the instrument can be assumed to be sufficiently strong. See wooldridge:2016 for some discussion why this model might be reasonable.

First, we estimate a simple linear model with OLS and with IV:

equation*[equation* omitted — 110 chars of source]

Table (ref) shows the estimated coefficients as well as the estimated conditional expectations for the $10\%, 50\%$ and $90\%$ quantiles of educ and exper, respectively (10 and 4, 12 and 12, 16 and 24). This way, the expected wages are calculated for three groups of employees, the low-educated/low-experienced, the middle-educated/middle-experienced and the high-educated/high-experienced.

center[center omitted — 44 chars of source]
table[table omitted — 849 chars of source]

Similarly as in other studies with this type of instrument, the IV estimate for educ is smaller than the OLS estimate (while the standard error is larger). The intuition is that both the years of education and an unobserved variable which measures ability and/or motivation are positively correlated with the wage, compare also the discussion in breitung:2022. The conditional expectations increase if higher values of educ and exper are considered. Interestingly, the results for OLS and IV are similar for the $50\%$ quantiles. For the $10\%$ quantiles, the IV estimate is larger than the OLS estimate, for the $90\%$ quantile, the IV estimate is smaller. There seems to be a tendency that the variability in terms of the regressor values is lower for the IV estimation. These results will be confirmed and extended by the DR analysis.

Figure (ref) shows the estimated conditional distribution functions for both OLS and IV, again for the $10\%, 50\%$ and $90\%$ quantiles of educ and exper. The estimated distribution functions are evaluated at all outcomes $Y_i$. In all cases, the monotonized version based on isotonic regression discussed in the last section is considered. For higher values of educ and exper, the distribution functions are shifted more and more to the right. For the $50\%$ quantile, the two functions are rather similar. For the $10\%$ quantile, the IV curve generally lies to the right of the OLS curve, where the largest differences are visible for values of $log(wage)$ between $0.5$ and $1$ as well as around $0$. For the $90\%$ quantile, the IV curve generally lies to the left with the largest differences between $1.5$ and $2$ and the maximal difference is slightly larger than for the $10\%$ quantile.

center[center omitted — 47 chars of source]
figure[figure omitted — 413 chars of source]

For completeness, Figure (ref) shows the estimated DR curve for the $90\%$ quantiles without monotonization, illustrating why it makes sense to add the monotonizing step.

center[center omitted — 59 chars of source]
figure[figure omitted — 267 chars of source]

To give more evidence about the difference between OLS and IV estimation, pointwise confidence bounds for the differences of the conditional distribution functions are calculated and plotted in Figure (ref). This is done by bootstrap, i.e. by drawing with replacement $B=200$ times from the individuals. For each $y$, the confidence interval to the level of significance $90\%$ is calculated. This yields a Hausman-type statistical test for the relevance of the IV approach: If $0$ is not contained in the interval, one can conclude that the two estimators of the distribution functions are statistically significantly different. Assuming that the instrument is exogenous and correlated with the endogenous regressor, the IV-based estimator is then the only valid one.

center[center omitted — 47 chars of source]
figure[figure omitted — 490 chars of source]

The confidence bounds essentially confirm the analysis from Figure (ref). For the $10\%$ as well as for the $90\%$ quantile, the bounds do not contain $0$ for some subsets of the ranges of $log(wage)$ described above. For the $50\%$ quantile, the confidence bounds are smaller than these for the $10\%$ and $90\%$ quantile, which is the expected behavior from the simulations and from Table (ref)). The value $0$ lies outside the bounds for $y$ slightly smaller than $1$. Summed up, the message is that IV-estimation of DR models does make a difference compared to OLS-estimation.

Conclusion and Outlook

The paper has proposed a new consistent estimator for the semiparametric DR model which allows for endogenous regressors and where monotonicity is enforced. The method is easy to implement and should be appealing to practitioners. Apparently, the proposed procedure only works well if the instruments are sufficiently strong. To circumvent the problem of choosing appropriate instruments, it might be an idea for future research to adapt the procedure proposed by breitung:2022 for linear regression models to DR models. Here, rank-based transformations of non-normal regressors are used as additional regressors and no external instruments are necessary to obtain consistent parameter estimators. Other tasks for future research would be analytical results for the isotonic regression and a framework for discrete endogenous variables. The latter could be done along the lines of wooldridge:2002, Section 15.7.3, but would be computationally harder because there would no simple more step procedure available.