EconBase
← Back to paper

Practically significant differences between conditional distribution functions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,312 characters · 5 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Practically significant differences between conditional distribution functions

abstractIn the framework of semiparametric distribution regression, we consider the problem of comparing the conditional distribution functions corresponding to two samples. In contrast to testing for exact equality, we are interested in the (null) hypothesis that the $L^2$ distance between the conditional distribution functions does not exceed a certain threshold in absolute value. The consideration of these hypotheses is motivated by the observation that in applications, it is rare, and perhaps impossible, that a null hypothesis of exact equality is satisfied and that the real question of interest is to detect a practically significant deviation between the two conditional distribution functions. The consideration of a composite null hypothesis makes the testing problem challenging, and in this paper we develop a pivotal test for such hypotheses. Our approach is based on self-normalization and therefore requires neither the estimation of (complicated) variances nor bootstrap approximations. We derive the asymptotic limit distribution of the (appropriately normalized) test statistic and show consistency under local alternatives. A simulation study and an application to German SOEP data reveal the usefulness of the method.\\ Keywords: Distribution regression, empirical processes, self-normalization\\ JEL: C12, J31

Introduction

\setcounter{equation}{0}

In recent years, distribution regression has become a popular approach to model the entire conditional distribution function of an outcome (e.g. income of individuals) given covariates (e.g. the years of education or the years of working experience). For many real-world questions in empirical research this technique provides a very flexible and much more informative inference tool than methods with a direct focus on the conditional mean because of its ability to address aspects such as variability, asymmetry, tail behavior, or multimodality, which play a critical role in contexts involving inequality, risk or discrimination. This allows for richer inference, greater interpretability, and more realistic representations of heterogeneity across populations. For example, chernozhukov:2013 used this approach for modeling counterfactual income distributions, delgado:2022 applied it in duration analysis, wang:2023 on insurance data, spady:2025 on the gender wage gap, while chernozhukov:2020, sanchez:2020, wied:2024 and chernozhukov:2025 considered aspects of endogeneity in these models. Indeed, there are various approaches for distribution regression in the literature; a recent survey on this and on possible applications is given by kneib:2023.

In this paper, we are concerned with the comparison of two conditional distribution functions in the framework of the semiparametric distribution regression introduced by foresiperacchi:1995. We choose this particular model due to the combination of flexibility and interpretability, noting that our methodology can also be used for other models. The problem of comparing conditional distribution functions has recently been addressed by hulei:2024 using a different (i.e. the conformal prediction) framework. These authors state two important applications for this test problem: On the one hand, the equality of the training and test distribution in machine learning can be tested. On the other hand, the equality of conditional distributions under different experimental conditions in causal inference can be investigated. Earlier papers about this topic include zhou:2017, who proposed smooth tests for the equality of distributions. General specification testing in conditional distribution models was considered by andrews:1997, RotheWied2013 and troster:2021 among others.

A common feature of all work in this area consists in the fact that it has its focus on testing the exact equality

align[align omitted — 78 chars of source]

of the conditional distribution functions $F^1_{Y|X}$ and $ F^2_{Y|X}$ corresponding to the two samples. In this paper we take a different point of view on the problem of comparing conditional distribution functions. Our work is motivated by the fact that in many cases it is very unlikely that the conditional distribution functions exactly coincide over the full domain of interest. We thus address a concern which was already mentioned in berger1987, who argued that it {\it “is rare, and perhaps impossible, to have a null hypothesis that can be exactly modeled by a parameter being exactly $0$”} (here the difference of the conditional distribution functions over the domain of interest). A similar point was raised by tukey1991 who mentioned in the context of multiple comparisons of means that {\it$\ldots $ All we know about the world teaches us that the effects of A and B are always different - in some decimal place - for any A and B. Thus asking “Are the effects different?” is foolish. $\ldots $” } In the context of comparing two conditional distribution functions this means that we might not be interested in testing the exact equality as formulated in (ref) as we do not believe that this hypothesis can be exactly satisfied (at least not on a microscopic scale). Instead it might be more reasonable to investigate if the deviation between $ F^1_{Y|X}$ and $F^2_{Y|X} $ is in some sense “small” or not. Therefore we propose to test the hypotheses

align[align omitted — 136 chars of source]

where $d$ is a metric on the set of distribution functions (we will later specify a particular norm), and $\Delta > 0$ is a given threshold that determines when the deviation between the conditional distribution functions is considered as {\it practically significant} {\it or relevant}. Note that we obtain the hypotheses in (ref) from (ref) for $\Delta =0$, but in the present paper we are not interested in arbitrary small deviations between the conditional distribution functions and therefore we restrict ourselves to the case $\Delta >0$ throughout this paper. Thus by testing the hypotheses (ref) we are trying to answer the question: are the differences between the two conditional distribution functions economically significant? In causal inference, one would answer the question if there is a {\it relevant treatment effect}. \\ Hypotheses of this form are often called {\it relevant} hypotheses. The interchanged hypotheses of this type, that is $ H_0: d > \Delta ~\text{ versus }~ H_1: d \leq \Delta$, have found considerable attention in the biostatistics literature for comparing finite dimensional parameters wellek2010testing, but - despite their importance - relevant hypotheses have not been studied intensively in the econometric literature. A few exceptions are the works of see detteschumann:2024 and kuttadette who considered these hypotheses in the context of testing the equivalence of pre-trends in difference-in-differences estimation and for validating the homogeneity of the slopes in high-dimensional panel models, respectively.

An essential ingredient in implementing any test for the hypotheses (ref) is the specification of the threshold $\Delta$. This choice is case specific, depends sensitively on the particular problem under consideration and has to be carefully discussed for each application. We argue that such a specification is possible in many cases. For example, in the empirical application below, we consider the case of changes in the German income distribution between $2013$ and $2020$. A relevant change in the conditional income distribution here means that the incomes increased stronger compared to inflation, which leads to a natural choice of $\Delta$. We consider the hypothetical case, in which the model structure remains the same and the incomes increased to exactly the same extent as the consumer prices. If we use a log-linear model, an increase of the incomes of, say, 7%, implies a shift of the intercept of $0.07$. Then, we calculate the distance between these two conditional distribution functions for some fixed values of the regressors, which yields the value $\Delta$. In order to address model uncertainty at this stage, it would be possible to calculate $\Delta$ for several models and check if the null hypothesis is rejected for the minimum of these values of $\Delta$.

Moreover, for applications, where a specification of $\Delta$ is difficult, we will develop a method for determining a threshold from the data, which can serve as measure of evidence for similarity between the two conditional distribution functions with a controlled type I error $\alpha$ (see Remark (ref) for more details).

In the present paper we will develop a pivotal test for the hypotheses (ref) where the metric $d$ is given by the $L^2$-norm. We consider the semiparametric distribution regression model, which has been proposed by foresiperacchi:1995 for the conditional distribution functions and is introduced in Section (ref). The basic idea consists of using a suitable estimator of the $L^2$-norm of the difference between the conditional distribution functions and reject for large values of this estimate. However, deriving critical values either by asymptotic theory or by bootstrap under the composite null hypothesis in (ref) is nontrivial. In particular the asymptotic distribution depends on nuisance parameters, which are difficult to estimate in practice. To develop a pivotal test for the hypotheses in (ref) we propose in Section (ref) a self-normalization approach which is based on the weak convergence of a stochastic process estimating the difference $F^{1} _{Y|X}(\cdot |x) - F^{2} _{Y|X}(\cdot |x) $ of the conditional distribution functions and is used to normalize the difference between the (squared) $L^2$ norm of the estimated difference and the (squared) $L^2$-norm of the true difference of the conditional distribution functions. As a consequence the limiting distribution is free of nuisance parameters and no bootstrap approximation or variance is required to calculate quantiles for a decision rule. We prove that the resulting test is consistent, has asymptotic level $\alpha$ and can detect local alternatives at a parametric rate. Moreover, for applications, where a specification of the threshold $\Delta$ is difficult, we develop a method for determining a threshold in the hypothesis (ref) from the data, which can serve as measure of evidence for similarity between the two conditional distribution functions with a controlled type I error $\alpha$ (see Remark (ref) for more details). In the same section we also propose an asymptotically pivotal confidence interval for the distance between the two conditional distribution functions and discuss extensions to testing for practically significant (or relevant) endogeneity. In Section (ref), we study the finite sample properties of the new test by means of a simulation study, and in Section (ref) we illustrate the potential of our approach in an empirical application. Remarks on possible extensions to detect relevant endogeneity and some discussions conclude the paper.

Semiparametric distribution regression

\setcounter{equation}{0}

We use semiparametric distribution regression as introduced by foresiperacchi:1995 to model the dependence between a real valued outcome variable $Y$ and a $p$-dimensional predictor $X$. Here, the conditional distribution function of $Y$ given the set of regressors $X$ is modeled by

align[align omitted — 71 chars of source]

for some (known) link function $\Lambda$ (such as the normal distribution) and some function $\beta(y)$. Although the function $\Lambda$ must be chosen in advance, this model is called semiparametric in the literature. Moreover, in the finite sample study in Section (ref), we demonstrate that our approach is not very sensitive with respect to a misspecification of the link function.

A nice feature of model (ref) is that it allows for potentially different parameters for each argument $y$ of the conditional distribution function and that the same time it is easy to interpret because of the linear model structure, which is matched with some link function. Inverting the conditional distribution function leads to models for the conditional quantiles. A potential advantage of this approach compared to standard quantile regression lies in the ability to automatically incorporate jump points. This might be interesting for data sets, where roundings or phenomena such as thresholds for limited-earning jobs are present. Usually, in the literature, no explicit restrictions on $\beta(y)$ such as continuity are imposed. However, it is clear that $\beta(y)$ has to be chosen such that $F_{Y|X}(y|x)$ fulfills the properties of a conditional distribution function if the model is correctly specified. Given $x$, this function must be right-continuous, monotonically increasing and converge to $1$ and $0$ for $y \rightarrow \infty$ and $y \rightarrow -\infty$, respectively. This way, conditional quantiles can be obtained without facing the quantile crossing problem. As pointed out in wied:2024, $Y$ may be discretely distributed even if the link function $\Lambda$ is continuous and there is no one-to-one connection between the distribution of $Y$ and the link function.

For an i.i.d. sample of observations, the conditional distribution function $F_{Y|X}(y|x)$ (more precisely, the corresponding parameter) can be consistently estimated by a Z-estimator similarly as a probit model would be estimated, for example, where the estimation is performed separately for each $y$. To be precise, one introduces the indicator functions $I_y := 1\{Y \leq y\}$ with $E(I_y|X=x) = F_{Y|X}(y|x) = \Lambda(x^\top \beta(y))$. The model can be interpreted as a latent variable model with $I_y = 1$ if $X^\top \beta(y) + U \geq 0$ and $I_y = 0$ otherwise, where $U$ is an exogenous random variable (meaning that $X$ is independent of $U$) with distribution function $\Lambda$.

Relevant differences between conditional distribution functions

\setcounter{equation}{0}

We will develop a pivotal test for a relevant difference between two conditional distribution functions as formulated in equation (ref) in the common two-sample problem (for example, comparing the distributions of populations in two different years). For $\ell=1,2$ let $(X_1^{\ell},Y_1^{\ell}),\ldots,(X_{n_{\ell}}^{\ell},Y_{n_{\ell}}^{\ell})$ denote two independent samples of i.i.d observations, where $Y_i{^\ell}\in\mathbb{R}$ and $X_i{^\ell}\in\mathbb{R}^p$. For the conditional distribution function of $Y^{\ell} $ given $X^{\ell} =x$ we consider the model

align[align omitted — 88 chars of source]

where $\beta_\ell (y) $ denotes a $p$-dimensional parameter (depending on $y$) and $\Lambda$ is a given distribution function. For a given threshold $\Delta > 0$ we are interested in the relevant hypotheses

eqnarray[eqnarray omitted — 403 chars of source]

where $x$ is a fixed covariate,

equation*[equation* omitted — 43 chars of source]

denotes the (squared) $L^2$-norm of the function $f$ and $I\subset \mathbb{R}$ is a given interval of interest for the response $Y$. The maximum likelihood estimator described in the previous section is denoted by $\hat \beta_\ell (y) $ for each sample $(\ell=1,2)$. For our approach it is also necessary to consider the corresponding estimators for the samples $(X_1^{\ell},Y_1^{\ell}),\ldots,(X_{\lfloor n_\ell t \rfloor}^{\ell},Y_{\lfloor n_\ell t \rfloor}^{\ell})$ of the first $\lfloor n_\ell t \rfloor$ observations in each sample, where $t\in [0,1] $ is a fixed constant. We denote these estimators by $\hat \beta ^{\ell}(t,y)$ and note that $\hat \beta ^{\ell}(y)= \hat \beta ^{\ell}(1,y)$ ($\ell =1,2$). Finally, we define $$ \hat F^\ell _{Y|X}(t,y|x) := \Lambda(x^\top \hat \beta^\ell(t,y)) $$ as the sequential estimator of the conditional distribution function $F^\ell_{Y|X}(y|x)$ from the sample $(X_1^{\ell},Y_1^{\ell}),\ldots,(X_{\lfloor n_\ell t \rfloor}^{\ell},Y_{\lfloor n_\ell t \rfloor}^{\ell})$ ($\ell =1,2$). For constructing the test, we then define the statistics

equation[equation omitted — 75 chars of source]

and

equation[equation omitted — 139 chars of source]

where $\epsilon >0 $ is a constant used to achieve numerical stability and

equation[equation omitted — 96 chars of source]

denotes the difference between the sequential estimates of the conditional distribution functions. Note that (under standard assumptions) $\hat\Delta(t,y|x)$ estimates

align[align omitted — 75 chars of source]

consistently, and consequently the statistic $\hat T_n$ defined in (ref) is a consistent estimate of $$ \int_I\Delta^2 (y|x) dy = \| F^1_{Y|X}(\cdot |x) - F^2_{Y|X}(\cdot |x) \|_2^2. $$ Therefore, we propose to reject the null hypothesis in (ref), whenever

equation[equation omitted — 79 chars of source]

where $q_{1-\alpha}$ is the $(1-\alpha)$-quantile of the distribution of the random variable

equation[equation omitted — 141 chars of source]

and $\mathbb{B}$ is a standard Brownian motion. Note that the quantiles of this distribution can easily be obtained by simulation. For example, for $\epsilon = 0.1$, the $0.95$-quantile of the distribution of $\mathbb{W}$ is given by $1.7546$. We also emphasize that the rejection probabilities of the test (ref) are not very sensitive with respect to choice of $\epsilon$, as this parameter appears in the quantile and the statistic $\hat V_n$; see Remark (ref) for more details and Table (ref) in Section (ref) for some empirical evidence.

Under the following assumptions, which are similar to the assumptions on the conditional distribution regression model considered in chernozhukov:2013, we can show that the test (ref) is consistent and has asymptotic level $\alpha$.

assumption{\rm For $\ell=1,2$, the following holds: \begin{enumerate} • The two populations are independent from each other. • The conditional distribution function takes the form $F^\ell_{Y|X}(\boldsymbol{\cdot}|x) = \Lambda(x^\top \beta_\ell(y))$, where $\Lambda$ is such that, for fixed $y$, $\beta_\ell(y)$ is consistently estimated by the Z-estimator mentioned in Section 2. • $I \subset \mathbb{R}$ is either a compact interval or a finite set. In the former case, the conditional density function $f^\ell_{Y|X}(\boldsymbol{\cdot}|x)$ exists, is uniformly bounded and uniformly continuous. • $E||X^\ell||^2 < \infty$ and the minimum eigenvalue of the $p\times p$-matrix \begin{equation*} E\left(X^\ell X^{\ell^\top} \frac{\lambda(X^\ell \beta_\ell(y))^2}{\Lambda(X^\ell \beta_\ell(y)) [1-\Lambda(X^\ell \beta_\ell(y))]} \right) \end{equation*} is bounded away from zero uniformly with respect to $y \in I $, where $\lambda$ is the derivative of $\Lambda$. \end{enumerate} }
theoremIf Assumption (ref) holds, and $n_1,n_2\rightarrow\infty$, such that $\tfrac{n_1}{n_1+n_2}\rightarrow c\in(0,1)$, the decision rule (ref) defines a consistent asymptotic level $\alpha$ test. In particular, \begin{eqnarray*} \lim_{ \substack{n_1,n_2\rightarrow\infty \\ {n_1}/({n_1+n_2})\rightarrow c\in(0,1)}} \mathbb{P}\big(\hat T_n > \Delta^2 + q_{1-\alpha} \hat V_n\big) = \begin{cases} 0, & if\ \int_I \hat\Delta^2(y|x) dy < \Delta^2 \\ \alpha, & if\ \int_I \hat\Delta^2(y|x) dy = \Delta^2 and \tau^2>0 \\ 1, & if\ \int_I \hat\Delta^2(y|x) dy > \Delta^2 . \end{cases} \end{eqnarray*}
proofThe statement essentially follows from the weak convergence of the stochastic process \begin{align} \left\{\sqrt{n}\big(\hat\Delta(t,y|x)- \Delta(y|x)\big)\right\}_{(t,y)\in [\epsilon,1]\times I} \leadsto \left\{\mathbb{H}(t,y)\right\}_{(t,y)\in [\epsilon,1]\times I} \end{align} in the space $\ell^\infty ([\epsilon,1]\times I ) $ of all bounded functions on the set $[\epsilon,1]\times I, $ where $\hat\Delta(t,y|x)$ and $\Delta(y|x)$ are defined in (ref) and (ref), respectively, and $\left\{\mathbb{H}(t,y)\right\}_{(t,y)\in [\epsilon,1]\times I}$ is a two-dimensional centered Gaussian process with covariance structure $$ {\rm Cov} \big(\mathbb{H}(t_1,y_1),\mathbb{H}(t_2,y_2)\big)= {t_1 \land t_2 \over t_1t_2} \cdot H(y_1,y_2) $$ for some function $H$ (here we do not reflect the dependence of $\mathbb{H}$ and $H$ on $x$ in the notation). This result corresponds to Theorem (ref) in the Appendix and is established through several carefully detailed steps provided therein (for the definition of the function $H$, see equation (ref)). Granted with the weak convergence in (ref) we then proceed as follows. We first rewrite the statistic $\hat T_n$ in (ref) as \begin{align} \nonumber \hat T_n - \int_I \Delta^2(y|x) dy &= \int_I \big(\hat\Delta(1,y|x) - \Delta(y|x)\big)^2 dy + 2 \int_I \Delta(y|x)\big(\hat\Delta(1,y|x) - \Delta(y|x)\big) dy \\ &= \mathbb{D}_n(1) + \int_I \big(\hat\Delta(1,y|x) - \Delta(y|x)\big)^2 dy \nonumber \\ &= \mathbb{D}_n(1) + o_{\mathbb{P}} \Big ( {1 \over \sqrt{n}} \Big ) , \end{align} where the process $ \{ \mathbb{D}_n(t) \}_{t\in [\epsilon,1]}$ is defined by \begin{equation*} \mathbb{D}_n(t) = 2 \int_I \Delta(y|x)\big(\hat\Delta(t,y|x) - \Delta(y|x)\big) dy . \end{equation*} By (ref) and the continuous mapping theorem, this stochastic process converges weakly in $\ell^\infty [\epsilon,1]$, that is \begin{equation} \big\{\sqrt{n}\ \mathbb{D}_n(t)\big\}_{t\in [\epsilon,1]} \leadsto \left\{ 2 \int_I \Delta(y|x) \mathbb{H}(t,y) dy\right\}_{t\in [\epsilon,1]}, \end{equation} where $\mathbb{H}$ is the centered Gaussian process in (ref). Note that the estimate (ref) is a direct consequence of the weak convergence (ref) in $\ell^\infty [\epsilon,1]$. Obviously, the process on the right hand side of (ref) is a Gaussian process as well, and its covariance structure is given by \begin{equation} {t_1 \land t_2 \over t_1t_2} \tau^2, \end{equation} where \begin{equation} \tau^2 = 4 \int_I\int_I \Delta(y_1|x)\Delta(y_2|x)H(y_1,y_2)dy_1 dy_2 \end{equation} and $H$ is defined in (ref). Consequently it follows from (ref) and (ref) that \begin{equation} \big\{\hat{\mathbb{D}}_n(t)\big\}_{t\in [\epsilon,1]} \leadsto \left\{ {\tau \over t} \mathbb{B}(t)\right\}_{t\in [\epsilon,1]} \end{equation} in $\ell^\infty [\epsilon,1]$, where $\mathbb{B}$ is a standard Brownian motion. The statistic $\hat V_n^2$ in (ref) is used as a self-normalizing term, and we obtain by similar arguments the representation \begin{eqnarray} \hat V_n^2&=& \int_\epsilon^1 \Big\{\int_I\big(\hat\Delta^2(t,y|x) - \Delta^2(y|x)\big)dy - \int_I \big(\hat\Delta^2(1,y|x)- \Delta^2(y|x)\big)dy\Big\}^2 dt \nonumber\\ &=& \int_\epsilon^1 \Big\{\int_I\big(\hat\Delta(t,y|x) - \Delta(y|x)\big)^2 dy +2\int_I \Delta(y|x)\big(\hat\Delta(t,y|x) - \Delta(y|x)\big) dy \nonumber\\ &-& \int_I \big(\hat\Delta(1,y|x) - \Delta(y|x)\big)^2 dy -2\int_I \Delta(y|x)\big(\hat\Delta(1,y|x) - \Delta(y|x)\big) dy \Big\}^2 dt \nonumber \\ &=& \int_\epsilon^1 \big \{ \mathbb{D}_n(t) - \mathbb{D}_n(1) \} ^2 dt + o_{\mathbb{P}} \Big ( {1 \over \sqrt{n}} \Big ) . \end{eqnarray} Combining (ref) with the representations (ref) and (ref) it follows from the continuous mapping theorem that \begin{equation} \sqrt{n}\bigg( \hat T_n - \int_I \Delta^2(y|x) dy, \hat V_n\bigg) \overset{\mathcal D}{\longrightarrow} \tau\left(\mathbb{B}(1), \Big[\int_\epsilon^1 \big(\mathbb{B}(t)/t-\mathbb{B}(1) \big)^2 dt \Big]^{\tfrac{1}{2}} \right ). \end{equation} Consequently, a further application of the continuous mapping theorem gives (provided that $\tau>0$) \begin{equation} \frac{\hat T_n - \int_I \Delta^2(y|x) dy}{\hat V_n} \overset{\mathcal D}{\longrightarrow} \mathbb{W}, \end{equation} where the random variable $\mathbb{W}$ is defined in (ref). We will now use this result to prove the statement regarding the rejection probabilities in Theorem (ref) noting that a simple calculation gives \begin{eqnarray} \mathbb{P}\big(\hat T_n > \Delta^2 + q_{1-\alpha} \hat V_n\big) = \mathbb{P}\Big (\frac{\hat T_n - \int_I \Delta^2(y|x) dy}{\hat V_n} > q_{1-\alpha} + \frac{\Delta^2 - \int_I \Delta^2(y|x) dy}{\hat V_n}\Big ). \end{eqnarray} It follows from (ref) that $\hat V_n \overset{\mathbb P}{\longrightarrow} 0$, which implies that \begin{eqnarray*} \frac{\Delta^2 - \int_I \Delta^2(y|x) dy}{\hat V_n} \overset{\mathbb P}{\longrightarrow} \begin{cases} \infty & if\ \int_I \hat\Delta^2(y|x) dy < \Delta^2 \\ 0 & if\ \int_I \hat\Delta^2(y|x) dy = \Delta^2 \\ -\infty & if\ \int_I \hat\Delta^2(y|x) dy > \Delta^2 \end{cases} . \end{eqnarray*} Therefore, the assertion of Theorem (ref) is a consequence of the weak convergence (ref) and the representation (ref).

The test (ref) can detect local alternatives converging to the null hypothesis at a rate $1/\sqrt{n}$. Compared to testing the exact equality $F^1_{Y|X}(\boldsymbol{\cdot}|x) = F^2_{Y|X}(\boldsymbol{\cdot}|x)$ of the conditional distribution functions, where local alternatives can be defined in the form $F^1_{Y|X}(\boldsymbol{\cdot}|x) - F^2_{Y|X} (\boldsymbol{\cdot}|x)= {1 \over \sqrt{n} } h(y|x)$ (for some function $h$), there are several possibilities to define local alternatives for the hypotheses (ref) with $\Delta >0 $ such that

eqnarray[eqnarray omitted — 187 chars of source]

for some constant $\delta >0$. To be specific we consider the testing problem (ref) and assume a local alternative of the form

equation[equation omitted — 164 chars of source]

where $\delta >0$ is some constant and $ g(\cdot |x)$ and $h(\cdot |x)$ are functions satisfying $$ \| g(\cdot |x) \|^2 = \Delta^2~,~~\int_I g(y |x) h(y |x) dy = 1 $$ (note that a possible choice is $h=g /\Delta^2$, where $g$ is an arbitrary function with $L^2$-norm $\Delta$). By a straightforward calculation it follows that in this case (ref) holds and the following result shows that the test (ref) has non-trivial power under these local alternatives. The proof follows by a careful inspection of the proof of Theorem (ref) and the arguments given in the online appendix.

theoremLet Assumption (ref) be satisfied and assume that $\tau^2>0 $ and $n_1,n_2\rightarrow\infty$, such that $\tfrac{n_1}{n_1+n_2}\rightarrow c\in(0,1)$. Under the local alternatives of the form (ref) the decision rule (ref) satisfies \begin{eqnarray} \nonumber \lim_{ \substack{n_1,n_2\rightarrow\infty \\ {n_1}/({n_1+n_2})\rightarrow c\in(0,1)}} \mathbb{P}\big(\hat T_n > \Delta^2 + q_{1-\alpha} \hat V_n\big) &=& \mathbb{P} \Big ( \mathbb{W} > q_{1-\alpha } - {\delta \over \big[\int_\epsilon^1 \big(\mathbb{B}(t)/t-\mathbb{B}(1) \big)^2 dt \big]^{{1}/{2}}} \Big ) \\ &>& \mathbb{P} \big ( \mathbb{W} > q_{1-\alpha } \big ) = \alpha. \end{eqnarray}
proofBy an inspection of the proof of Theorem (ref) it is easy to see that the weak convergence in (ref) also holds under the local alternatives (ref). Consequently, the first equality follows by similar arguments as given in the proof of Theorem (ref). Moreover, the random variable $\int_\epsilon^1 \big(\mathbb{B}(t)/t-\mathbb{B}(1) \big)^2 dt$ has a Lebesgue density and is therefore positive with probability equal to $1$. The same property implies that the distribution function of the random variable $ \mathbb{W} $ in (ref) is strictly increasing, which proves the strict inequality in the second line of (ref).
remark{\rm Note that the hypotheses in (ref) and (ref) are nested. Therefore, it follows that the rejection of the null hypothesis (ref) by the test (ref) for $\Delta= \Delta_0$ also yields (asymptotically) rejection of $H_{0}$ for all $\Delta\leq \Delta_0$. By the sequential rejection principle, we may simultaneously test the hypotheses in (ref) and (ref) for different $\Delta \geq 0$ starting at $\Delta = 0$ and increasing $\Delta $ to find the minimum value of $\Delta $, say \begin{align} \hat \Delta_\alpha:= \min \Big \{ \big \{ 0 \} \cup \big \{\Delta \ge 0 \,| \, \hat T_n > \Delta^2 + q_{1-\alpha} \hat V_n \big \} \Big \} = \big \{ ( T_n - q_{1-\alpha} \hat V_n ) \lor 0 \big \}^{1/2} \end{align} for which $H_0$ in (ref) is not rejected. The quantity $\hat \Delta_\alpha $ could be interpreted as a measure of evidence against the null hypothesis in (ref) (note that the null hypothesis is accepted for all thresholds $ \Delta \geq \hat \Delta_\alpha $ and rejected for $ \Delta < \hat \Delta_\alpha $). This means that higher values of $\hat \Delta_\alpha$ yield stronger evidence against the null hypothesis of a small difference. In this sense, the question of a reasonable choice of the threshold $\Delta$ to test the relevant hypotheses may be postponed until after seeing the data. Moreover, if one is not directly interested in testing, a pivotal confidence interval for the $L^2$-distance between the conditional distributions can be constructed. To be precise, note that it follows from the proof of Theorem (ref) that $$ \lim_{n \to \infty } \mathbb{P} \Big ( \frac{\hat T_n - \int_I \Delta^2(y|x) dy}{\hat V_n} \leq q_{1- \alpha} \Big ) = 1- \alpha~, $$ and therefore an asymptotic one-sided $(1-\alpha)$-confidence interval for the deviation $\| \Delta^2(\cdot |x) \|_2 $ between the two conditional distribution functions is given by \begin{align*} [\hat \Delta_\alpha,\infty)= \Big [ \big \{ ( \hat T_n - q_{1-\alpha} \hat V_n ) \lor 0 \big \}^{1/2} ,\infty \Big ) . \end{align*} }
remark{\rm Note that the test (ref) depends on the specification of the constant $\epsilon> 0$, which has been introduced in the definition of the statistic $\hat V_n$ to achieve numerical stability. We give a heuristic argument that the test is not very sensitive with respect to this choice. To see this, we define the random variable \begin{equation} \mathbb{V}_\epsilon =\Big( \int_\epsilon^1 \big(\mathbb{B}(t)/t-\mathbb{B}(1)\big)^2 dt \Big)^{1/2} , \end{equation} corresponding to the denominator in (ref), $M^2 = \int_I \Delta^2 (y|x) dy $ as the squared $L^2$-norm of the difference of the conditional distribution functions and make the dependence of the quantile $q_{1-\alpha}$ on the distribution of $\mathbb{W} = \mathbb{B}(1)/ \mathbb{V}_\epsilon $ more explicit using the notation $q_{1-\alpha} (\mathbb{W}) $. With these notations we obtain for the probability of rejection \begin{align} \mathbb{P}\big( \hat T_n > \Delta^2 + q_{1-\alpha} (\mathbb{W} ) \hat V_n \big) & \approx \mathbb{P}\Big( \mathbb{B}(1) > \frac{\sqrt{n} (\Delta - M^2)}{\tau} + \mathbb{V }_\epsilon \cdot q_{1-\alpha}(\mathbb{B}(1)/ \mathbb{V}_\epsilon ) \Big) , \end{align} where we have used the weak convergence in (ref) for the approximation of the probabilities. Observe that in the last expression only the quantity $\mathbb{V }_\epsilon \cdot q_{1-\alpha}( \mathbb{B}(1)/\mathbb{V}_\epsilon)$ depends on the constant $\epsilon$, which enters in the definition of the random variable $\mathbb{V}_\epsilon$. However, for fixed $v>0 $ we have $ {v }\cdot q_{1-\alpha}( \mathbb{B}(1)/v) = q_{1-\alpha}( \mathbb{B}(1) )$, which gives a heuristic explanation why the probability in (ref) is not very sensitive with respect to the choice of $\epsilon$ (which was also observed empirically, see Table (ref) in Section (ref)). Note that the same argument applies if the Lebesgue measure in the definition of the statistic $\hat V_n$ is replaced by a different measure. }
remark{\rm (a) It follows from its proof that the statement of Theorem (ref) is also valid under less restrictive assumptions, for example under time series assumptions. More precisely, Theorem (ref) remains true and the decision rule (ref) defines a pivotal, consistent and asymptotic level $\alpha$-test for the hypotheses (ref), whenever the weak convergence in (ref) can be established. Now a careful inspection of the proofs in the Appendix (in particular the proof of Theorem (ref) and (ref)) shows that such results can be obtained under appropriate mixing Bradley.2007, physical dependence Wu2005 or $m$-approximability HrmannKok conditions (however a proof of a statement corresponding to Theorem (ref) would change substantially). A similar comment can be made regarding the independence assumption on the two samples. (b) The limit distribution of the statistic in (ref) is not pivotal if $\Delta (y|x) \equiv 0 $. To see this, note that it follows by the weak convergence (ref), the definition of $T_n$ and $\hat V_n$ in (ref) and (ref), respectively, and the continuous mapping theorem that \begin{align} \frac{\hat T_n }{\hat V_n} \overset{\mathcal D}{\longrightarrow} { \int_I \mathbb{H}^2(1,y) dy \over \big [ \int_\epsilon^1 \big\{\int_I\mathbb{H}^2(t,y) dy - \int_I \mathbb{H}^2(1,y) dy\big\}^2 dt \big]^{1/2}} \end{align} where $\mathbb{H}$ denotes the Gaussian process on $[\epsilon,1]\times I $ with covariance structure defined in equation (ref) of the Appendix. Consequently, the distribution of the right hand side of (ref) is not pivotal and its quantiles cannot be used for a decision rule which rejects the null hypothesis (ref) with $\Delta=0$ for large values of ${\hat T_n }$. (c) Note that in the formulation of the hypothesis pair, we have fixed the value of the predictor $x$. This is particularly relevant for our empirical application, where we compare the tests' results for different values of $x$. The approach could easily be extended to an aggregation over the regressors. }
remark{\rm The test approach can in principle also be used for detecting “relevant endogeneity”, compare chernozhukov:2020, sanchez:2020, wied:2024 and chernozhukov:2025 for results about endogeneity in distribution regression models. wied:2024 considers another conditional distribution function, in which potential endogeneity of one regressor $Y_2$ is addressed and an instrument $Z$ is available. This function is defined as $$ F^V_{Y|X,Y_2}(y|x,y_2) = \int P(I_y^* \geq 0|X=x,Y_2=y_2,V=v) d F_V(v) $$ with \begin{eqnarray*} I_y^* &=& X'\beta_1(y) + Y_2 \beta_2(y) + U(y) \\ Y_2 &=& X'\gamma_1 + Z'\gamma_2 + V. \end{eqnarray*} Here, $U(y)$ and $V$ are latent variables with mean zero, for which there exists a decomposition \begin{equation*} U(y) = \alpha_1 V + \alpha_2 \epsilon(y). \end{equation*} In this one-sample setting, one could test for relevant endogeneity with the hypothesis pair \begin{eqnarray*} &H_0:& || F_{Y|X,Y_2}(y|x,y_2) - F^V_{Y|X,Y_2}(y|x,y_2) || \leq \Delta \\ &H_1:& || F_{Y|X,Y_2}(y|x,y_2) - F^V_{Y|X,Y_2}(y|x,y_2) || > \Delta. \end{eqnarray*} The structure of the test statistic would be the same, using appropriate estimators of $F_{Y|X,Y_2}(y|x,y_2)$ and $F^V_{Y|X,Y_2}(y|x,y_2)$. While an estimator for the first conditional distribution function is obtained similarly as described in Section (ref), the estimation of the function $F^V_{Y|X,Y_2}(y|x,y_2)$ requires the use of a control function and an additional average step, as described in wied:2024. }

Monte Carlo simulations

\setcounter{equation}{0}

In this section, we demonstrate the application of the proposed testing procedure (ref) and present finite sample evidence. Throughout this section, we assume the same sample size for both groups, i.e. $n=n_1=n_2$. We consider two different scenarios, where the first one is given by a simple linear model and the second one is inspired by the application discussed in Section (ref). For both scenarios, we will consider different sample sizes and vary the underlying distributions. In addition, we will use one scenario to investigate the robustness of the procedure with respect to misspecifcation of the link function $\Lambda$. Further, in order to investigate the effect of the parameter $\epsilon$ in definition of the self-normalizing statistic (ref), we will simulate the type I error rates of the test for three different choices of $\epsilon \in \{0.05,0.1,0.2\}$, in all scenarios under consideration. All simulations have been run using R Version 4.4.3 with a total number of $1000$ simulation runs for each configuration, where the significance level was chosen as $\alpha=0.05$.

In Scenario 1 we assume that

equation[equation omitted — 75 chars of source]

where for both groups $\ell=1,2$, the error term $U_\ell\sim \mathcal{N}(0,\sigma^2)$ is independent of the regressor $X_{\ell 2}\sim \mathcal{N}(0,1)$, such that $F^\ell_{Y|X}(y|x) = \Lambda(x^\top \beta_\ell(y))$, where $x^\top=(1,x_{12}) $. The link function $\Lambda$ is the distribution function of a normal distribution with mean $0$ and variance $\sigma^2$ such that $\beta_1(y)=(y-1,-1)$, $\beta_2(y)=(y-1.7,-1)$. For a given $x^\top =(1,1)$ we obtain, depending on the choice of $\sigma^2$, the following underlying true values

equation[equation omitted — 242 chars of source]

for $y \in I=[-1,5]$.

Figure (ref) displays the empirical rejection probabilities of the test (ref) for the three choices of the variance $\sigma^2$ and three different sample sizes $n=200,500,1000$ in dependence of the threshold $\Delta^2$, where the parameter $\epsilon $ in the self-normalizing statistic is chosen as $\epsilon =0.1$. For all three sub-figures, the red vertical dashed line indicates the true underlying value of the squared $L^2$-distance between the conditional distribution functions given in (ref) and thus the margin of the null hypothesis in (ref). Therefore, the empirical rejection probabilities on the right of this line correspond to the situation under the null hypothesis and thus display simulated type I error rates, whereas the values on the left of this line correspond to the alternative and hence represent the simulated power. We observe that all simulated type I error rates are close to or well below the chosen significance level of $\alpha=0.05$, approaching $0$ for larger choices of the threshold $\Delta^2$.

Next we investigate the sensitivity of the test (ref) with respect to the choice of the parameter $\epsilon$ in the self-normalizing statistic (ref). We concentrate on the margin $|| F^1_{Y|X}(\boldsymbol{\cdot}|x) - F^2_{Y|X}(\boldsymbol{\cdot}|x) ||_2 = \Delta $ of the hypothesis (ref) and display in Table (ref) the simulated type I error of the test (ref) for three different choices of $\epsilon \in \{0.05,0.1,0.2\}$. We observe that the choice of $\epsilon$ does not influence the performance of the test, as the results are qualitatively the same across all three configurations under consideration.

Finally, considering the situation under the alternative, we note that the simulated power approaches $1$ if the variance under consideration is sufficiently small or the sample size large. These findings underline the theoretical properties of the test stated in Theorem (ref).

figure[figure omitted — 609 chars of source]

We continue investigating the robustness of the test procedure. For this purpose we assume that the true underlying link function is given by the distribution function of a logistic distribution with location parameter $0$ and scale parameter $1$, whereas estimation is performed under the assumption of a standard normal distribution. For this scenario the true underlying value of the squared $L^2$-distance between the conditional distribution functions is given by $$\| F^1_{Y|X}(\cdot |(1,1)) - F^2_{Y|X}(\cdot | (1,1)) ||^2_2= 0.080, $$ where $F^1$ and $F^2$ now denote the distribution functions of the standard logistic distribution. The functions $F^1_{Y|X}(\cdot |(1,1)) $ and $F^2_{Y|X}(\cdot |(1,1))$ are shown in the left panel of Figure (ref), while the right panel shows the simulated rejection probabilities of the test (ref). We observe that even in this situation of misspecification, the test keeps its nominal level for all three sample sizes under consideration while still achieving reasonable power, approaching a maximum of $0.8$ for the largest sample size of $n=1000$.

figure[figure omitted — 824 chars of source]

In a second scenario we now consider a more complex setting which is inspired by the real data application presented in Section (ref). Precisely we assume that

eqnarray[eqnarray omitted — 140 chars of source]

where the regressors $(X_{\ell 2},X_{\ell 2})^\top $ are bivariate normal distributed, that is $$(X_{\ell 2},X_{\ell 3})\sim \mathcal{N}_2(0,\Sigma), with \Sigma= \sigma^2 \cdot

pmatrix[pmatrix omitted — 33 chars of source]

, $$ $\sigma\in\{0.5,1,2\}$, $X_{\ell 4}$ is binomial distributed, that is $X_{\ell 4}\sim \Bin(n_\ell,0.5),$ and the errors $U_\ell$ are independent centered normal distributed with variance $\sigma^2$. As a consequence, the link function $\Lambda$ is the distribution function of a normal distribution with mean $0$ and variance $\sigma^2$. For a given $x=(1,1,1,1,1)$ we obtain the following underlying true values

equation[equation omitted — 254 chars of source]

for $y \in I=[2,8]$ and we denote this scenario by Scenario 2a. A visualization of the true underlying conditional distribution functions for $\sigma=1$ is given in the upper left panel of Figure (ref). Compared to model (ref), the model (ref) is more complex and the estimation of the corresponding parameters requires larger sample sizes to obtain reasonable results. Motivated by the size of the real data set in Section (ref), we now consider sample sizes given by $n=1000,2000,5000,10000$ for the simulations.

In the upper part of Figure (ref) we display the empirical rejection probabilities for the medium variance of $\sigma^2=1$ and the large variance of $\sigma^2=4$, respectively. Again, the vertical red line indicates the margin of the null hypothesis in (ref). We observe that, as expected, the power increases with increasing sample sizes, while the approximation of the level is very precise for all sample sizes under consideration. We conclude that the test is able to detect relevant differences with a high degree of accuracy, even when the underlying models are complex and involve the estimation of many parameters.

figure[figure omitted — 1,145 chars of source]

Finally, we consider the model

align[align omitted — 84 chars of source]

where $Y_1$ and $Y_2$ are defined in (ref), keeping the distributions of the regressors and the error terms as stated above. This scenario is denoted by Scenario 2b and describes the presence of a jump with moderate size in the distribution function (see the lower left panel of Figure (ref)). Such a phenomenon might occur in income data (see wied:2024),and violates Assumption 1.3. Our simulations show that this violation is not problematic in the types of applications we are interested in. For a given $x=(1,1,1,1,1)$ we now obtain the following underlying true values

equation[equation omitted — 252 chars of source]

for $y \in I=[2,8]$. The lower row of Figure (ref) shows the empirical rejection probabilities for the medium variance of $\sigma^2=1$ (middle column) and the high variance of $\sigma^2=4$ (right column). Compared to Scenario 2a the results are almost exactly the same and the differences are barely visible. The simulated type I error rates of the test (ref) at the margin of the hypothesis (ref) for Scenarios 2a and 2b are shown in second and third row of Table (ref), where we consider different choices $\epsilon \in \{0.05,0.1,0.2\}$ for the constant $\epsilon$ in the self-normalizing statistic (ref). As in Scenario 1, these results demonstrate that the choice of $\epsilon$ only slightly influences the performance of the test. Furthermore, investigating the effect of the jump point in the distribution function, i.e. comparing Scenario 2a and Scenario 2b, we observe that differences are only visible for the configurations with $\sigma^2=1$ and $\sigma^2=4$. However, even in these cases the results are very similar and the differences are negligible.

table[table omitted — 2,444 chars of source]

Application to German income data

We apply our new test for detecting relevant changes in the conditional income distribution of German employees based on micro data from the German SOEP (soep:2022) between the years 2013 and 2020. The dependent variable is given by $log(income)^\ell$, $\ell \in \{2013,2020\}$, while the linear predictor is given by Mincer's earning equation

equation*[equation* omitted — 167 chars of source]

This means that we model the expression

equation[equation omitted — 133 chars of source]

where $\Phi$ denotes the distribution function of the standard normal distribution. We compare the conditional distribution functions of the logarithmic monthly income given the covariates {\it years of education}, {\it years of working experience}, {\it squared years of working experiences}, {\it part time job yes/no} in the years 2013 and 2020. The sample sizes are given by $16031$ for the year 2020 and $18262$ for the year 2013. We choose this particular time interval to simplify the interpretation of the results: there was a rise in the salary for a {\it Minijob} in 2013 and the Covid pandemic led to more substantial changes in the income structure after 2020.

The goal is to answer the question in which sense the incomes have increased statistically significantly more than the consumer prices. Economists are typically interested in how the link between incomes (or wages) and inflation looks like. For example, gordon:1988 showed empirically that the link is less strong than expected, while jorda:2023 show that the link might have become stronger after the pandemic in 2020. In fact, for our data, there is some empirical evidence that incomes increased to a higher extent than consumer prices, with the implication that there are other important drivers of inflation. In general, incomes increased by $17\%$ (sozialpolitik:2025), while inflation was around $7.4\%$ (destatis:2025), as Table (ref) shows. However, these numbers do not yield information about changes in the whole (conditional) distribution of income.

table[table omitted — 373 chars of source]

For the application of the test (ref) in the present context, we have to specify a value of the threshold $\Delta$. Our approach considers the $L^2$-distance between the two conditional distribution functions (on the interval $I= [2,10]$) corresponding to the years $2013$ and $2020$ (for a specified set of regressor values) under the assumption that the change of the incomes is exactly the same as the change of the consumer prices. Under this assumption, all coefficients in (ref) remain the same except the intercept, which increases by $0.074$. We are interested in the results for three cases: $x^\top =(q^{edu}_{0.1}, q^{exper}_{0.1}, m )$, $x^\top=(q^{edu}_{0.5}, q^{exper}_{0.5}, m^{partt} )$ and $x^\top=(q^{edu}_{0.9}, q^{exper}_{0.9}, m^{partt} )$, where $q^{edu}_{\alpha}$ and $q^{exper}_{\alpha}$ denote the empirical $\alpha$-quantile of the regressors {\it educ} and {\it exper}, respectively, and $ m^{partt} )$ denotes the empirical mean of the component {\it partt}. For the three different cases, the values for $\Delta^2$ are chosen as $0.001$, $0.0008$ and $0.0009$, respectively.

Let us first discuss the calculation of the test results. We consider $\epsilon=0.1$, and for this choice the quantile $q_{0.95}$ is given by $1.7546$. The values of the self-normalizing statistic $\hat V_n$ in (ref) are given by $0.0223$, $0.0236$, $0.0095$ (for the regressor quantiles of $10\%$, $50\%$ and $90\%$). For the different values of $\Delta^2$, we then obtain $$ \hat T_n=0.048 > 0.0419 = 0.001 + 1.7546 \cdot 0.0223$$ for the $10\%$-quantile ($p$-value 0.031), $$ \hat T_n=0.028 < 0.0422 = 0.0008 + 1.7546 \cdot 0.0236 $$ for the $50\%$-quantile ($p$-value 0.12) and $$ \hat T_n=0.028 > 0.0176 = 0.0009 + 1.7546 \cdot 0.0095 $$ for the $90\%$-quantile ($p$-value 0.012).

This means that, for the lower and upper quantile of the regressors, but not for the median, there is a statistically significant relevant change. Figure (ref) illustrates the results. The upper panel shows the estimated conditional distribution functions for the three quantiles and the two years. The lower panel shows the $p$-values of the test (ref) for the different cases as a function of $\Delta^2$. This curve increases in $\Delta^2$, which is obvious given the construction of the test. For the lower and upper quantile of the regressors, the $p$-values are below $0.05$ for some values of $\Delta^2$. For the median, this is never the case. The estimates of $\hat \Delta_\alpha$ in (ref), which yield measures of evidence against the null hypothesis, are given by $0.0097$ for the lower quantile, $0.0129$ for the upper quantile and $0$ for the median. This suggests that the “more extreme” employees faced a stronger increase in income compared to the “moderate” ones.

figure[figure omitted — 811 chars of source]

{\bf Acknowledgements} This work was supported by TRR 391 Spatio-temporal Statistics for the Transition of Energy and Transport (Project number 520388526) funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation).