Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,504 characters · 11 sections · 77 citation commands
Distribution Regression in Duration Analysis: an Application to Unemployment Spells
\setcounter{equation}{0}
Existing semiparametric duration models can be broadly classified into two groups: those based on the conditional hazard and those based on the quantile regression. The former includes the proportional hazard model (Cox1972, Cox1975), the proportional odds model (Clayton1976, Bennett1983, Murphy1997), and the accelerated failure time model (Kalbfleisch1980); see Guo2014 for an overview. In these models, the conditional hazard identifies the conditional cumulative distribution expressed in terms of the error's marginal distribution of a transformed failure time regression model (see, e.g., Hothorn2014). Censored quantile regression proposals are relatively more recent, see, e.g., Ying1995, Honore2002, Portnoy2003, Peng2008, and Wang2009c.
Conditional hazard and quantile regression models are alternative modeling strategies with advantages and drawbacks. Classical conditional hazard specifications impose identification conditions that are difficult to justify in some circumstances. For instance, the proportional hazard specification rules out important forms of heterogeneity, see, e.g., Portnoy2003 and the discussion in Section (ref). On the other hand, models based on quantile regression specifications avoid this problem, but impose that the underlying conditional cumulative distribution of duration time is absolutely continuous, which rules out other sampling schemes such as discrete outcomes.
The distribution regression approach was proposed by Foresi1995 and formalized by Chernozhukov2013a; see also Rothe2013a, Rothe2016a, Chernozhukov2019 and Chernozhukov2018a for later developments. One attractive feature of this modeling strategy is that marginal effects associated with distribution regression models are more flexible than those obtained from classical conditional hazard specifications, as they can accommodate richer forms of heterogeneity. In contrast to quantile regression, distribution regression can accommodate discrete, continuous and mixed duration data in a unified manner under fairly weak regularity conditions. Chernozhukov2013a show that distribution regression encompasses the Cox1972 model as a special case and represents a useful alternative to quantile regression.
The distribution regression model specifies the distribution of the duration outcome variable by means of a generalized linear model with known link function and nonparametric varying coefficient depending on duration time. When data is uncensored, the varying coefficients at each duration value are consistently estimated by binary regressions. This method is also valid using censored data with known censoring time, which is an unlikely situation in duration studies, e.g., unemployment duration spells. When duration and censoring times are potentially unobserved, the standard binary regression approach lead to inconsistent estimators.
In this article, we propose a weighted binary regression procedure using Kaplan1958 weights. The varying coefficients are identified by means of a set of moment restrictions of the duration time and covariates' joint distribution. These moments are consistently estimated by their corresponding Kaplan-Meier integrals when the duration time is subjected to censoring; if censoring is not an issue, the Kaplan-Meier integrals reduce to the sample analogue of the moment restrictions. The resulting estimator admits a representation as a non-degenerate U-statistic of order three, whose H\'{a}jek projection is asymptotically distributed as a normal with an asymptotic variance that can be consistently estimated from the data. The resulting inference procedures on the coefficients do not rely on choosing tuning parameters such as bandwidths, imposing smoothness conditions on the censoring random variable, or using truncation arguments.
We establish the consistency and finite-dimensional distribution convergence of the varying coefficients under weak regularity conditions. Under more restrictive assumptions we justify the convergence in distribution of the estimated coefficient function as a random element of a suitable metric space. We also justify bootstrap-based inference procedures.
In order to justify the asymptotic inferences based on our proposed method, we impose restrictions on the joint distribution of the duration time, censoring time and covariates. Stute1993, Stute1996a provided strong consistency and a central limit theorem for Kaplan-Meier integrals that estimate moments of (partially observed) durations time and covariates under the following identification conditions: (a) the duration and censoring times are independent, and (b) the event of being censored is conditionally independent of the covariates given the duration time. Stute shows that the two conditions suffice to identify the joint distribution of duration time and covariates from the observed censored data, and, hence, to identify corresponding moments of these conditional distributions. Of course, the identification of alternative models may require alternative regularity conditions. For instance, identification of the Cox1972's hazard model, Aalen1980's additive hazard model, or censored quantile regressions models (Portnoy2003) only requires that duration and censoring times are independent given covariates. However, as we mention in Section 2, these models usually exclude important types of heterogeneity, or restrict their attention to absolutely continuous duration data. Alternatively, one can potentially estimate the moment equations using inverse probability weighting (IPW) methods, as suggested by Robins1992, Robins1995a and later extended by Wooldridge2007a to general missing data problems. Although this path allows one to only assume that duration and censoring times are independent given covariates, it also requires that the conditional distribution of the censoring times given covariates is (uniformly) bounded away from one, which rules out discrete distributions and the most popular distribution functions used in survival analysis such as exponential, Gamma, log-normal and Weibull, among others. If one is not comfortable with these additional restrictions, truncation/trimming arguments need to be carefully introduced, and the choice of tuning parameters becomes a first-order concern. By adopting the Kaplan-Meier integrals approach, we avoid these alternative restrictions, but rely on conditions (a) and (b) above.
We illustrate the relevance of our proposal by assessing how changes in unemployment insurance benefits affect, on average, the distribution of unemployment duration, using data from the Survey of Income and Program Participation for the period 1985-2000. We find that, by allowing the distributional marginal effects to vary over the unemployment spell, our proposed method can reveal interesting insights when compared to traditional hazard models as those used by Chetty2008. For instance, our results suggest a non-monotone marginal effect of an increase in unemployment insurance on the unemployment duration distribution. Such a finding is in sharp contrast with those obtained using classical proportional hazard models. On the other hand, our results agree with Chetty2008 in that increases in unemployment insurance have larger effects on liquidity-constrained workers, suggesting that unemployment insurance affects unemployment duration not only through a moral hazard channel but also through a liquidity effect channel.
The rest of the article is organized as follows. Section (ref) introduces the basic notation and motivates distribution regression models using duration data as an alternative to classical conditional hazard modeling. Section (ref) describes our estimation procedure, whereas Section (ref) introduces regularity conditions needed to justify inferences on the distribution regression varying coefficients. All the results are proved in the Supplementary Appendix. Application of the results to different contexts are placed in Section (ref). Section (ref) briefly summarizes the results of the Monte Carlo simulations -- detailed discussion is presented in the Supplementary Appendix\footnote{The Supplementary Appendix is available at \url{https://pedrohcgs.github.io/files/Delgado_GarciaSuaza_SantAnna_2021_EJ-Appendix.pdf}}. Finally, we apply the proposed techniques to investigate the effect of unemployment benefits on unemployment duration in Section (ref). All our replication files are available at \url{https://github.com/pedrohcgs/KMDR-replication}.
\setcounter{equation}{0}
Consider the $\mathbb{R}^{1+k}-valued$ random vector $\left( T,X\right) $ defined on $\left( \Omega ,\mathcal{F},\mathbb{P}\right) ,$ where $T$ is the duration outcome of interest and $X$ is a $k-$dimensional vector of time-invariant covariates, with supports $\mathcal{T}$ and $\mathcal{X},$ respectively. Henceforth, let $\mathbf{X} = (1,X')'$ and $\mathbf{x} = (1,x')'$. We assume that the conditional cumulative distribution function (CDF) of $T$ given $X$ follows the distribution regression (DR) model,
where $\theta _{0}\left( \cdot \right) =\left( \varphi \left( \cdot \right) ,\beta _{0}\left( \cdot \right) ^{\prime }\right) ^{\prime }\mapsto \Theta \subseteq \mathbb{R}^{k+1}$ is a vector of nonparametric functions, and $\Lambda$ is a known link function.
Chernozhukov2013a point out that the DR model is a flexible alternative to classical duration models as it allows coefficients to vary with duration time. In particular, they show that it nests the Cox1972 proportional hazard (PH) model. Model ((ref)) also generalizes other traditional models commonly adopted in duration analysis such as Kalbfleisch1980 accelerated failure time (AFT) model, and Clayton1976 proportional odds (PO) model.
We note that all the aforementioned classical duration models are special cases of the linear transformation model
where $\beta _{0}\in \mathbb{R}^{k}$ is a vector of unknown parameters, $ \varphi $ is a (potentially unknown) monotonically increasing transformation function, $\Lambda ^{-1}$ is the $\Lambda ^{\prime }s$ quantile function, and $U$ is uniformly distributed in $\left[ 0,1\right] $ independently of $X$. For instance, $(a)$ the PH model corresponds to model ((ref)) with $\Lambda $ the complementary log-log (cloglog) link function, $\Lambda \left(u\right) =1-\exp \left( -\exp \left( u\right) \right) $; $(b)$ the AFT model corresponds to ((ref)) with $\Lambda $ commonly assumed to be the link function of a log-logistic or of a Weibull distribution; and $(c)$ the PO model corresponds to ((ref)) with $\Lambda $ the logistic link function, $ \Lambda \left( u\right) =\exp \left( u\right) \left/ \left( 1+\exp \left( u\right) \right) \right. $; see, e.g. Doksum1990, Cheng1995 and Hothorn2014. Indeed, the CDF associated with ((ref)) is given by
Therefore, ((ref)) is a natural generalization of ((ref)) that allows all the slope coefficients varying with $t$. Other duration models with varying slope coefficients are also special cases of ((ref)). For instance, Aalen1980 semi-parametric additive hazard model reduces to ( (ref)) with $\Lambda \left( u\right) =1-\exp \left( -u\right) $. The non-proportional odds model considered by McCullagh1980 and Armstrong1989 can also be expressed in terms of the DR specification ((ref)) when $\Lambda $ is the logistic link function.
We conclude this section by emphasizing that allowing for varying slope coefficients is particularly important to capture richer heterogeneous effects of covariates across the distribution of the duration outcome. Indeed, traditional duration specifications, such as the PH, PO or AFT models, implicitly impose that all covariates affect the CDF of $T$ in a proportional, and monotonic form. If $\Lambda $ is differentiable with Lebesgue density $\lambda$, partial effects for these models have the form
Thus, the sign of the marginal effect of any $X$'s component is not allowed to vary with $t$, which may be too restrictive in some applications. For instance, in the context of the empirical application in Section (ref), ((ref)) does not allow a non-monotonic effect, e.g. U-shaped, ruling out non-stationary search models in the spirit of VandenBerg1990. The DR model ((ref)) bypass such limitations since
which sign is allowed to vary with $t$. We view this added flexibility as an attractive feature of the DR modeling approach.
\setcounter{equation}{0}
The main challenge when dealing with duration data is that the outcome of interest $T$ is subject to censoring according to a variable $C$. As so, inferences must be based on observations $\left\{ Y_{i},X_{i},\delta _{i}\right\} _{i=1}^{n}$, with $Y_{i}=\min \left( T_{i},C_{i}\right) $ the realized (potentially censored) outcome, $ \delta _{i}=1_{\left\{ T_{i}\leq C_{i}\right\} }$ indicates whether or not $T$ is observed, and $\left\{ T_{i},C_{i},X_{i}\right\} _{i=1}^{n}$ are independent and identically distributed ($iid$) as $\left( T,C,X\right)$. Henceforth, $1_{\left\{ A\right\} }$ is the indicator function of the event $A$.
Suppose for the moment that $\delta _{i}=1,$ i.e., $T_{i}=Y_{i},$ for all $i=1,\dots ,n$. In this case, the joint $\left( T,X'\right)'$ CDF, $F\left( t,x\right)$, is consistently estimated by its sample version $ \tilde{F}_{n}\left( t,x\right) =n^{-1}\sum_{i=1}^{n}1_{\left\{ Y_{i}\leq y,X_{i}\leq x \right\} }$, where inequalities are coordinate-wise. Therefore, following Foresi1995 and Chernozhukov2013a, $\theta _{0}\left( t\right) $ can be consistently estimated by $\tilde{\theta}_{n}(t)$, the maximizer of the conditional likelihood function of $\left\{ 1_{\left\{ Y_{i}\leq t\right\} },X_{i}\right\} _{i=1}^{n}$
with
However, when $T$ is subject to right-censoring, i.e., $T\neq Y$, $\tilde{F}_{n}$ is no longer a consistent estimator of $F$ and, hence, neither is $\tilde{\theta}_{n}\left( t\right) $.
Since $T$ is not always observed, $F$ and $\theta (t)$ must be identified from the joint distribution of the observed data on $\left( Y,X,\delta \right)$. However, this is not always possible without additional information. Henceforth, for any random variable $\xi $, which can be $T,~C$ or $Y$, define $F_{\xi }(t)\equiv \mathbb{P}\left( \xi \leq t\right) $ and $\tau _{\xi }\equiv\inf \left( t:\mathbb{P}\left( \xi \leq t\right) =1\right) $, and let $A$ denote the (possibly empty) set of $F_{T}$ jumps. Also, for any generic function $g$, $ g(t-)\equiv\lim_{s\uparrow t}g(s)$. If no information about $T$ beyond $\tau _{Y}$ is available from the data, identification of
is the best that one can hope for; see, e.g., Tsiatis1975, Tsiatis1981, and Stute1993. In order to identify $F_{\ast }$ we impose the following assumptions.
Assumption (ref) is introduced to identify $F_{\ast }\left( \cdot ,\infty ,...,\infty \right) $ (see Tsiatis1975, Tsiatis1981 for discussion), whereas Assumption (ref) is introduced by Stute1993 and allows us to incorporate covariates into the analysis. Taken together, these two assumptions imply that covariates should have no effect on the probability of being censored once $T$ is known. Of course, Assumptions (ref) and (ref) hold if $C$ is independent of $\left( T,X\right)$, but can also hold under more general circumstances. See Stute1993, Stute1996a, Stute1999 for a discussion on Assumptions (ref) and and (ref), and the identification of $F_*$ and its moments. Identification can be achieved under alternative conditions, in the context of alternative conditional survival models, or alternative estimation procedures. See Introduction for further comments.
In order to identify $\theta_{0}(t)$, we ensure that $F=F_{\ast }$ by imposing the following condition; see Stute1999 for a similar assumption in the case of nonlinear regression with randomly censored outcomes.
Thus, $\theta _{0}\left( t\right) $ is identified as the parameter value that maximizes $Q\left( \theta, t \right) $ under standard conditions in binary regression, e.g., Amemiya1985, or more recently, VanderVaart1998.
Henceforth, $Y_{1:n}\leq ...\leq Y_{n:n}$ are the ordered $Y-values$, where ties within outcomes $T$ or within censoring times $C$ are ordered arbitrarily and ties among $T$ and $C$ are treated as if the former precedes the latter. For observations $\left\{ \xi _{i}\right\} _{i=1}^{n}$ of a random variable $\xi $, which may be $\delta $ or $X$, $\xi _{\left[ i:n \right] }$ is the $i-th$ $\xi -$concomitant of the order statistics $\left\{Y_{i:n}\right\} _{i=1}^{n},$ i.e., $\xi _{\left[ i:n\right] }=\xi _{j}$ if $Y_{i:n}=Y_{j}$.
In the absence of covariates, $F_{\ast T}(t)=F_{\ast }\left( t,\infty ,...,\infty \right) $ is consistently estimated by the Kaplan1958 product-limit estimator,
By noticing that the jumps of $\hat{F}_{Tn}$ at $Y_{i:n}$ are
we can express $\hat{F}_{Tn}$ in an additive form as
When covariates are present, the natural $F_{\ast }\left( t,x\right) $ estimator is
see, Stute1993, Stute1996a. Stute1993 showed that, under Assumptions (ref)-(ref),
Thus, ((ref)) is indeed a natural candidate to replace $\tilde{F}_{n}$ in ( (ref)) under censoring. This suggests estimating $Q_{\ast }\left( t,\theta \right) $ by the Kaplan-Meier integral,
which is a weighted version of $\tilde{Q}_{n}\left( \theta ,t\right)$.
The Kaplan-Meier distribution regression (KMDR) estimator of $\theta _{0}(t)$ is then given by
Notice that, to compute the KMDR estimators, one simply needs to run a weighted binary regression, where the weights are the (random) Kaplan-Meier weights $W_{in}$. In the absence of censoring, i.e., $Y_{i}=T_{i},$ $\delta _{i}=1$, $i=1,..,n$, we have that the Kaplan-Meier weights $W_{in}=n^{-1}$, and $\hat{\theta} _{n}\left( t\right) $ reduces to $\tilde{\theta}_{n}\left( t\right) .$
\setcounter{equation}{0}
In this section, we present the asymptotic properties of our proposed KMDR estimator $\hat{\theta}_{n}(t)$ for $\theta _{0}\left( t\right)$.
In order to establish consistency, we impose the following fairly weak assumption which is compatible with discrete, continuous and mixed duration outcomes.
Next theorem, like any other result in the paper, is proved in the Supplemental Appendix. The proof applies the arguments in Example 5.40 in VanderVaart1998 to show that
In order to provide the finite dimensional asymptotic distribution of $\hat{\theta}_{n}(t),$ we need the following assumption.
The restriction on $\Lambda$ in Assumption (ref) is a classical assumption, which can be found, for instance, in Amemiya1985 Assumption 9.2.1. The assumption is satisfied by the normal and logistic distribution (probit and logit), but also for many other distributions. In the general case, Assumption (ref) essentially rules out extremes $t$, since, in these cases, $\Lambda$ may be near zero or one. This assumption can be relaxed at the cost of more involved proofs.
The score function is given by
where
The conditional Fisher information that the binary variable $1_{\left\{ T\leq t\right\} }$ contains about $\theta _{0}(t)$ given $X$ is $\mathcal{I}_{0}\left( t\right) =\mathcal{I}\left( \theta _{0}\left( t\right) ,t\right)$, where
Notice that, under our assumptions, $\theta \mapsto \mathcal{I}\left( \theta ,\cdot \right) $ is continuously differentiable.
We first pay attention to the asymptotic distribution of the score $\hat{\Psi }_{n}^{0}\left( t\right) =\hat{\Psi}_{n}\left( \theta _{0}(t),t\right) .$ Applying Stute1995a, Stute1996a lemmata, $\hat{\Psi}_{n}\left( \theta ,t\right) $ can be expressed as an $U-$statistic of order three with H\'{a}jek projection $\hat{U}_{n}^{0}(t)=\hat{U}_{n}(\theta _{0}(t),t)$, where
and for any function $\left( y,x\right) \mapsto \varphi \left( y,x\right) $,
where $F_{Y}^{0}(t) =\mathbb{P}\left( Y\leq t,\delta =0\right) \text{ and } F^{1}(t,x)=\mathbb{P}\left( Y\leq t,X\leq x,\delta =1\right)$.
To derive the asymptotic normality of the score, we need the following extra moment conditions.
Assumption (ref)$\left( a\right)$ guarantees that the variance of the leading term in $\zeta _{\psi _{\theta _{0}\left( t\right) t}}$ is bounded, and implies that $\mathbb{E}\left\Vert \zeta _{\psi _{\theta t}}(Y,X,\delta )\right\Vert^{2}$ is finite. The bias of the Kaplan-Meier integral is not necessarily $o\left( n^{-1/2}\right) $ for any integrand function, and may decrease to zero at a polynomial rate depending on the degree of censoring, which is characterized by the function $S$. Assumption (ref)$\left( b\right)$ on $\psi _{\theta_{0}(t)t}$ guarantees that the $\hat{\Psi}_{n}^{0}\left( t\right) $ bias is of order $o\left( n^{-1/2}\right) $; see Stute1994a, Chen1996, and Chen1997.
Let $\left\{ Z\left( t\right) \right\}_{t=t_{1}}^{t_{m}}$ be a $k+1$ Gaussian random vector with zero mean and covariances
with $j,m=1,...,k+1$.
With Assumption (ref), we show that $\hat{\Psi}_{n}^{0}\left( t\right) $ is asymptotically equivalent to a $U-statistic$ of order 3 with H\'{a}jek projection $\hat{U}_{n}^{0}\left( t\right)$. Thus, $\sqrt{n}\hat{U}_{n}^{0}$ is asymptotically normal with finite dimensional covariance $\Omega _{0}\left(j,m\right)$. Theorem (ref) below follows by applying Theorem 5.21 in VanderVaart1998.
In order to conduct inferences on $\theta_0(\cdot)$, $\Omega _{0}\left( j,m\right) $ is estimated by
where
and $\mathcal{I}_{0}\left( t\right) $ is estimated by
where $\hat{\psi}_{i:n}(t) = \psi _{\hat{\theta}_{n}(t)t}\left( Y_{i:n},X_{[i:n]}\right)$\footnote{Alternatively, one can estimate $-\mathcal{I}_{0}\left( t\right) $ using the Hessian of $\hat{Q}_n(\hat{\theta}_n,t)$. We follow this path in our simulations.}.
We can also justify inferences with the assistance of a multiplier bootstrap technique in the spirit of Theorem 2.3 in Stute2000 and Theorem 4 in SantAnna2016a, using an external resample of $\left\{ \hat{\zeta}_{i}(t)\right\} _{i=1}^{n}.$ Once we generate $iid$ random numbers $\left\{ V_{i}\right\} _{i=1}^{n}$ independently of the data with mean $0$, variance $1$, and finite third moment, the “resampled” $\left\{ \hat{\zeta}_{i}^{\ast }(t)\right\} _{i=1}^{n}$, with $\hat{\zeta}_{i}^{\ast }(t)=\hat{\zeta} _{i}(t)V_{i}$, forms a basis to compute the bootstrap analog of $\hat{\theta}_{n}(t)$ based on the asymptotic linearization ((ref)),
where
is the bootstrap version of $\hat{U}_{n}^{0}\left( t\right)$.
The theorem follows by checking the conditional version of the Linderberg-Levy conditions to show that, with probability 1, $\hat{U}_{n}^{\ast }\left( t\right) $ converges in distribution to $Z$ under the bootstrap law, i.e., conditional on the data, with probability 1.
Next, we strengthen the pointwise results in Theorems (ref) and (ref) to hold uniformly in $t$. Towards this end, let the space $\ell ^{\infty }\left( \mathcal{G}\right) ^{q}$ be the set of all $q\times 1$ vectors of bounded functions on $\mathcal{G}$, that is, all the functions' vectors $g:u\mapsto \mathbb{R}$ such that $\sup_{u\in \mathcal{G}}\left\Vert g\left( u\right) \right\Vert <\infty .$ $\mathcal{G}$ can be either an Euclidean or a functional space. We interpret the multivariate empirical process indexed by functions in $\mathcal{G}$ as a random element in the metric space $\ell ^{\infty }\left( \mathcal{G}\right) ^{q}$ endowed with the sup-norm. The empirical process $\hat{\Psi}_{n}$ is indexed by $\left( \theta ,t\right) \in \Theta \times \mathcal{T}$; in this case $\mathcal{G} =\Theta \times \mathcal{T}_{0},$ where $\mathcal{T}_{0}$ is the compact interval of $\mathcal{T}$ of interest. But $\hat{\Psi}_{n}$ can also be interpreted as an empirical process indexed by the class $\mathcal{F}$ of functions $\psi _{\theta t}:\mathcal{X\times }\mathcal{T\times }\left[ 0,1 \right] \mapsto \mathbb{R}^{1+k};$ in this case $\mathcal{G}=\mathcal{F}$. The random process $\hat{\Psi}_{n}$ is viewed as a random element of the metric space $\ell ^{\infty }\left( \Theta \times \mathcal{T}_{0}\right) ^{k+1}$ or $\ell ^{\infty }\left( \mathcal{F}\right) ^{k+1},$ where $ \mathcal{T}_{0} \in \mathcal{T}$ is the compact subset of $\mathcal{T}$ we are interested in.
In order to establish the weak convergence of $\sqrt{n}\left( \hat{\theta} _{n}-\theta _{0}\right) $ as a random element of the metric space $\ell ^{\infty }\left( \mathcal{T}_{0}\right) ^{k+1}$, we impose the following additional assumption, which is also assumed in Chernozhukov2013a.
Notice that $f_{\left. T\right\vert X}\left( \left. t\right\vert X\right) = \dot{\theta}_{0}^{\prime }\left( t\right) \cdot \boldsymbol{X\cdot }\lambda \left( \theta _{0}^{\prime }\left( t\right) \boldsymbol{X}\right) $ a.s., where $\dot{\theta}_{0}\left( t\right) =d\theta _{0}\left( t\right) /dt.$ Therefore, assuming that $f_{\left. T\right\vert X}$ is a proper probability density may not be innocuous. In practice, $\mathcal{T}_0$ should be bounded by the minimum and maximum value of the observed uncensored duration.
Let $\mathbb{Z}$ be a centered tight Gaussian element of $\ell ^{\infty }\left( \mathcal{T}_{0}\right) ^{k+1}$ with matrix of variance and covariance functions
with $\Omega _{0}\left( t_{1},t_{2}\right)$ defined in ((ref)).
This functional CLT is proved extending Chernozhukov2013a (CFM henceforth) strategy to our setup. To this end, we first show, using Stute1995a, Stute1996a lemmata, that the score function $\hat{\Psi}_{n}$ and $U_n$ are asymptotically equivalent in the metric space $\ell^{\infty}(\Theta, \mathcal{T}_0)^{k+1}$, i.e., $$\sup_{\theta \in \Theta ,t\in \mathcal{T}_{0}}\left\Vert \hat{\Psi}_{n}\left( \theta ,t\right) -U_{n}\left( \theta ,t\right) \right\Vert=o_{p}\left( n^{-1/2}\right), $$ where $U_{n}$ is a $U-process$ with H\'{a}jek projection $\hat{U}_{n}$. Then, we show that $$\sup_{\theta \in \Theta ,t\in \mathcal{T}_{0}}\left\Vert U_{n}\left( \theta ,t\right) -\hat{U}_{n}\left(\theta ,t\right) \right\Vert =o_{p}\left( n^{-1/2}\right) $$ applying asymptotic results for U-process in Arcones1993, Arcones1995. Since the class $\mathcal{E}$ of functions $\left( y,x,d\right) \mapsto \zeta _{\psi _{\theta t}}\left( y,x,d\right) $ is Donsker, $$\left\{ \sqrt{n}\left( \hat{U} _{n}\left( \theta ,t\right) -\mathbb{E}\left[\zeta _{\psi _{\theta t}}\left( Y,X,\delta \right) \right]\right) :\theta \in \Theta ,t\in \mathcal{T}_{0}\right\} $$ converges weakly in $\ell ^{\infty }\left( \Theta \times \mathcal{T}_{0}\right) ^{k+1}$. Then, noticing that Assumptions (ref) - (ref) imply Assumption DR in CFM, the functional $\theta \mapsto \Psi _{0}\left( \theta ,t\right) $, with $\Psi _{0}\left( \theta ,t\right) = \mathbb{E}\left[\zeta _{\psi _{\theta t}}\left( Y,X,\delta \right)\right]$, is continuously differentiable for each $t\in \mathcal{T}_{0},$ with $\left. \left. d\Psi _{0}\left( \theta ,t\right) \right/ d\theta ^{\prime }\right\rfloor _{\theta =\theta _{0}\left( t\right) }= - \mathcal{I}_{0}\left( t\right) ,$ which is invertible. Then, according to CFM's Lemma E.3,
with $\sup_{t\in \mathcal{T}_{0}}\left\Vert \hat{r}_{n}\left( t\right) \right\Vert =o_{p}\left( 1\right)$. This result and $\mathcal{E}$'s Donskerness prove the theorem.
A bootstrap version of this theorem follows by applying the multiplier CLT (see sections 2.6 and 3.6 in VanderVaart1996).
In next section we discuss applications of these results to test restrictions on $\theta _{0}$, as well as to make inferences on counterfactual distributions and on average distribution marginal effects.
\setcounter{equation}{0}
A natural class of hypotheses to be tested is those of the form
for some compact subset $\mathcal{T}_{0}$ of $\mathcal{T},$ where $\theta \mapsto \Phi \left( \theta \right) \in \mathbb{R}^{q}$, $q\leq k+1,$ is a continuously differentiable map with derivatives $\dot{\Phi}\left( \theta \right) =\partial \Phi \left( \theta \right) /\partial \theta ^{\prime },$ such that $rank\left( \dot{\Phi}\left( \theta \right) \right) =q$ for all $ \theta $ in a neighborhood of $\theta _{0}\left( t\right)$. For instance, when $\Phi \left( \theta\right) = \theta$, the null and alternative reads
which is a significance test for varying coefficients. We can also test that a linear combination of coefficients is satisfied, e.g., $\Phi(\theta) = R \theta - r$, where $R$ and $r$ are known and $rank(R)=k+1$.
A natural statistic is,
By the delta-method, and applying Theorem (ref), uniformly in $t\in \mathcal{T}_{0},$
and under $H_{0},$
Since the $\omega _{\infty }^{\prime }s$ distribution depends on $\theta _{0}\left( \cdot \right) $ and other unknown features of the underlying data generating process, analytical critical values seems unfeasible, at least with this level of generality. However, critical values can be estimated using the bootstrapped estimator of $\theta _{0}\left( t\right)$
Applying Theorem (ref), uniformly in $t\in \mathcal{T}_{0},$ for almost every sample, it follows that
Therefore, we can use the bootstrap test statistic
By Theorem (ref), under $H_{0}$, with probability 1, $\hat{\omega}_{n}^{\ast }$ converges in distribution under the bootstrap law to $\omega _{\infty }$ . Under $H_{1},$ $\hat{\omega}_{n}^{\ast }=O\left( 1\right) $ with probability 1.
We now describe a practical bootstrap algorithm to conduct such types of tests.
An interesting application of this type of test is to testing constancy of some functional of $\theta _{0}$. For instance, one can set $$\Phi \left( \theta \right) =\left[ \boldsymbol{0}\vdots I_{K} \right] \theta$$ to test that all the slope coefficients are constant. A test statistic for this hypothesis is
with $\bar{\hat{\Phi}}_n = \int_{s\in \mathcal{T}_0} {\Phi \left( \hat{\theta}_{n}\left( s\right) \right)} \Psi(ds)/\int_{s\in \mathcal{T}_0} \Psi(ds)$, $\Psi$ being a known probability measure.\footnote{In practice, one can set $\Psi$ to be equal to the empirical distribution of the censored outcome $Y$, though, in this case, the bootstrap procedure needs to be adjusted to account for this additional source of randomness. See Section (ref) for related results.} Bootstrap critical values can be obtained using the bootstrapped statistic,
with $\bar{\hat{\Phi}}_n^* = \int_{s\in \mathcal{T}_0} {\Phi \left( \hat{\theta}_{n}^*\left( s\right) \right)} \Psi(ds)/\int_{s\in \mathcal{T}_0} \Psi(ds)$.
Theorems (ref) and (ref) can also be applied to make inferences on $F_{\left.T\right\vert X}\left( \left. t\right\vert x\right) $. The DR $F_{\left. T\right\vert X}\left( \left. t\right\vert x\right) ^{\prime }s$ estimator is
with bootstrap analog,
The asymptotic distribution can be obtained using the delta-method. Applying Theorem (ref), we have that
where $\mathcal{A}_{0}=\mathcal{T}_{0}\times \mathcal{X}_{0}$ is a compact subset of $\mathcal{T}_{0}\times \mathcal{X}_{0}$. Likewise, applying Theorem (ref), with probability 1,
Given two samples $\{Y_i^{(j)}, \delta_i^{(j)}, X_i^{(j)} \}_{i=1}^{n_{j}},~ j=1,2$, under some standard overlap conditions, we can use the conditional CDF estimator computed using sample 1, $\widehat{F}^{(1)}_{T|X,n}$, to compute the marginal counterfactual CDF of population 1 with respect to population 2, $$\widehat{F}^{(1,2)}_{T,n} (t) = \int_\mathcal{X} \widehat{F}^{(1)}_{T|X,n}(t|x) \Tilde{F}^{(2)}_{X,n}(dx) = \dfrac{1}{n_2}\sum_{i=1}^{n_2} \left( \widehat{F}^{(1)}_{T|X,n}(t|X_i^{(2)}\right).$$
Although $\hat{\theta}_{n}\left( t\right) $ provides useful information about the direction/sign of the effect of changes in $X$ on $F_{T|X}\left( t|X\right) ,$ in general, $\hat{\theta}_{n}(t)$ may not have a clear economic interpretation. In this section, we argue that this potential limitation can be easily avoided by focusing on the average distribution marginal effects (ADME) of $X$,\footnote{ To avoid cumbersome notation, we consider the case where the covariates $X$ are continuous, and enter the distribution regression model linearly. The asymptotic validity of all our results do not rely on this simplification.}
The ADME is the distributional analogue of the popular average partial effects. From ((ref)), it is clear that he natural estimator for $\eta _{0}\left( t\right)$ is
In what follows, we derive the asymptotic properties of $\hat{\eta}_{n}\left( t\right)$. As in Theorem (ref), we first show that, uniformly in $t \in \mathcal{T}_0$,
where
with $\zeta _{i}(t)$ is as defined in ((ref)), with $\varphi = \psi _{\theta t}$, $\dot{\lambda}\left( u\right) =d\lambda (u)/du$, $H\equiv \left[ 0_{k},I_{k}\right] $, $0_{k}$ the $ k\times 1$ vector of zero, and $I_{k}$ the $k$-dimensional identity matrix.
Let $\mathbb{Z}_{ADME}$ be a centered tight Gaussian element of $\ell ^{\infty }\left( \mathcal{T}_{0}\right) ^{k}$ with matrix of variance and covariance functions
Consider the bootstrapped estimator for the ADME
where $iid$ random numbers $\left\{ V_{i}\right\} _{i=1}^{n}$ are $iid$ random variables with mean $0$ and variance $1$, generated independently of the sample, and
where $\hat{\zeta} _{i}(t)$ is as in ((ref)) and $\widehat{\mathcal{I}}_{n}\left( t\right)$ is as in ((ref)).
We now describe a practical bootstrap algorithm to compute simultaneous confidence intervals for the ADME associated with a given covariate $X_{1}$
\setcounter{equation}{0}
In this section, we briefly summarize the simulation results discussed in Section (ref) of the Supplementary Appendix. In short, we compare the finite sample performance of the proposed KMDR estimators with those based on the proportional hazard (PH) model and on the proportional odds (PO) model. We consider three different data generating processes (DGPs): one DGP that satisfies the PH but not the PO assumption, one DGP that satisfies the PO but not the PH assumption, and one DGP where both PH and PO assumptions are violated, though it admits a DR specification with time-varying slope coefficient. We consider different levels of random censoring, and compare the KMDR, PH and PO models based on the conditional CDF, $F_{\left. T\right\vert X}$, and the ADME as defined in ((ref)). Both functionals have a clear economic interpretation.
Overall, the simulation results highlight that our proposed KMDR estimators performs nearly as well as the PH and PO model when these models are correctly specified, since KMDR nests both PH and PO specifications. On the other hand, when PH and PO models are misspecified, the proposed KMDR estimators performs better than these other popular specifications. Such gains are especially noticeable when one is interested in the ADME$\left( t\right) $; see Table (ref). \afterpage{
} When comparing DR specifications with different link functions, we notice that estimators that use the cloglog link function tend to be more robust against model misspecifications than those based on the logit link function. This is perhaps because the cloglog link function is asymmetric and adapts better to the non-central parts of the distribution where there are many zeros (left tail) or many ones (right tail). Thus, we recommend favoring the cloglog specification in detriment of the logit one, though a more formal discussion about such choices are beyond the scope of this paper.
\setcounter{equation}{0}
One of the main concerns of the design of unemployment insurance policies is their adverse effect on unemployment duration. The prevailing view of the economics literature is that increasing unemployment insurance (UI) benefits leads to higher unemployment duration driven by a moral hazard effect: higher UI increases the agent's reservation wage and reduces the incentive to job search, see, e.g., Krueger2002a and the references therein. Given that moral hazard leads to a reduction of social welfare, this argument has been used against increases in UI benefits.
In a seminal paper, Chetty2008 challenges the traditional view that the link between unemployment benefits and duration is only because of moral hazard. He shows, among many other things, that distortions cause by UI on search behavior are mostly due to a \textquotedblleft liquidity effect\textquotedblright . In simple terms, UI benefits provide cash-in-hand that allows liquidity constrained agents to equalize the marginal utility of consumption when employed and unemployed. Such a liquidity effect reduces the pressure to find a new job, leading to longer unemployment spells. However, in contrast to the moral hazard effect, the liquidity effect is a socially beneficial response to the correction of market failures. Thus, if one finds support in favor of liquidity effects, increases in UI benefits may lead to improvements in total welfare.
In this section, we provide additional evidence of the existence of this liquidity effect by comparing the effect of UI on household that are liquidity constrained with those that are unconstrained. In contrast to Chetty2008, we do not rely on the Cox proportional hazard model but rather use our proposed KMDR tools. In this context, the Cox hazard model might be too rigid as does not allow for the effect of UI benefits to vary depending on whether a worker is starting their unemployment claim or they have been unemployed for a while. As we argued before, our proposed tools allow for this richer types of heterogeneity.
As in Chetty2008, our data comes from the Survey of Income and Program Participation (SIPP) for the period spanning 1985-2000. Each SIPP panel surveys households at four-month intervals for two-four years, collecting information on household and individual characteristics, as well as employment status. The sample consists of prime-age males who have experienced job separation and report to be job seekers, are not on temporary layoff, have at least three months of work history in the survey and took up unemployment insurance benefits within one month after job loss. These restrictions leave 4,529 unemployment spells in the sample, 21.3% of those being right-censored. Unemployment durations are measured in weeks while individuals' UI benefits are measured using the two-step imputation method described in Chetty2008. For further details about the data, see Chetty2008.
To analyze the effect of UI benefits on unemployment duration, we estimate the DR model
where $\Psi \left( u\right) =1-\exp \left( -\exp \left( u\right) \right) $ is the complementary log-log link function, $UI$ is the worker's weekly unemployment insurance benefits, and $Z$ is a vector of controls including the worker's age, years of education, marital status dummy, logged pre-unemployment annual wage, and total wealth. To control for local labor market conditions and systematic differences in risk and performance across sector types, $Z$ also includes the state average unemployment rate and dummies for industry. We note that Chetty2008 considers additional controls, including state, year and occupation fixed effects, resulting in a specification with almost 90 unknown parameters. For this reason, we adopt a more parsimonious specification described above. Let $\theta \left( t\right) =\left( \alpha (t),\beta (t),\gamma (t)^{^{\prime }}\right) ^{\prime }$, and $\mathbf{X=}\left( 1,\ln UI,Z^{\prime }\right) ^{\prime }$.
Our main goal is to understand the effect of changes in UI on the probability of one finding a job in $t$ weeks, where $t=2,3,\dots 50.$ Although the sign of $\beta \left( t\right) $ indicates the direction of such a change, its magnitude may not have a straightforward economic interpretation. Thus, we focus on the ADME$\left( t\right) $ of\ $\ln UI$,
which is easy to interpret. For instance, an estimate of $ADME_{lnUI}(t)$ equal to 0.1 suggests that, on average, raising unemployment benefits by one percent increases the probability of finding a job in $t$ weeks or less by 0.1 percentage point. As discussed in Theorem (ref), we can estimate $ ADME_{\ln UI}\left( t\right) $ by
where $\psi \left( u\right) =d\Psi \left( u\right) /du$, and $\hat{\theta} _{n}\left( t\right) =\left( \hat{\alpha}_{n}(t),\hat{\beta}_{n}(t),\hat{ \gamma}_{n}(t)^{^{\prime }}\right) ^{\prime }$ is the KMDR estimator of $ \theta \left( t\right) $. When $\Psi $ is the cloglog link function, $\psi \left( u\right) =\exp \left( u\right) \exp \left( -\exp \left( u\right) \right) $.
In what follows, we multiply the estimators of ADME$_{\ln UI}\left( t\right) $ by 10, which is interpreted as the effect of a 10% increase in UI benefits. Figure (ref) reports the estimates for the full sample (solid line), together with the 90 percent bootstrapped pointwise, and simultaneous confidence intervals (dark and light shaded area, respectively) computed using Algorithm (ref).
The result reveals interesting effects. On average, a 10% increase in UI benefits appears to have no effect on the probability of a worker finding a job in the first seven weeks of the unemployment spell. Nonetheless, Figure (ref) shows that a change in UI is associated with a reduction of the probability of finding a job in the first $t=\left\{ 8,\dots ,50\right\} $ weeks. Such an effect seems to be monotone until week 18, where a 10% increase in UI is associated with a two-percentage-point decrease in the average probability of finding a job until that week. After week 18, the effect of an increase in UI benefits on employment probabilities seems to weaken but remains statistically significant at the 10% level, except for $ t=\left\{ 34,35,39\right\} $. Note that the bootstrap simultaneous confidence interval is slightly wider than the pointwise one. However, it is important to mention that the bootstrap uniform confidence interval is designed to contain the entire true path of the ADME$_{\ln UI}\left( t\right) $ 90% of the time, which is in sharp contrast to the bootstrap pointwise confidence interval. This highlights the practical appeal of using simultaneous instead of pointwise inference procedures to better quantify the overall uncertainty in the estimation of all ADMEs. \afterpage{
} Overall, Figure (ref) shows that although an increase in UI benefits is close to zero at the beginning of the unemployment spell, they have an U-shaped effect on the unemployment duration distribution. Interestingly, estimates of ADME$_{\ln UI}\left( t\right) $ based on the Cox PH model with the same set of covariates as in ((ref)) suggests that the effect of UI on unemployment duration distribution is monotone for $t \in [0,50]$. However, once the proportional hazard specification is tested using either Grambsch1994 procedure or our bootstrap-based testing procedure for the null hypothesis that all $\hat{\theta}_{n}(t), t=\{2,\dots,50\}$ are constant in $t$, as described in Section (ref), the null of proportionality is rejected at the usual significance levels\footnote{The p-value associated with our proposed test statistics is 0.004.}, implying that, indeed, a proportional hazard model may not be appropriate for this application. Our KMDR model does not rely on such an assumption.
Although the results in Figure (ref) show that, on average, changes in UI have a U-shaped effect on the unemployment duration distribution, the analysis remains silent about the liquidity effects. To shed light on the importance of the liquidity effect relative to the moral hazard effect, Chetty2008 argues that one can compare the response to an increase in UI benefits of workers who are not financially constrained with those who are constrained. Given that unconstrained workers have the ability to smooth consumption during unemployment, liquidity effects are absent and UI benefits lengthen unemployment duration only via moral hazard effects for these subgroup of individuals. To pursue this logic, we follow Chetty2008 and use two proxy measures of liquidity constraint: liquid net wealth at the time of job loss (\textquotedblleft net wealth\textquotedblright ) and an indicator for having to make a mortgage payment. Chetty2008 argues that workers with higher net wealth are less sensitive to UI benefit levels because they are less likely to be financially constrained. Similarly, workers that have to make mortgage payments before job loss have less ability to smooth consumption during unemployment because they are unlikely to sell their homes during the unemployment spell, whereas renters can adjust faster. To assess the role of liquidity effects, we divide the entire sample into four subsamples: $(a)$ workers with net wealth below the median, $(b)$ workers with net wealth above the median, $\left( c\right) $ workers with a mortgage, and $\left( d\right) $ workers without a mortgage. For each subsample, we estimate ((ref)) using the KMDR approach, its corresponding estimator of ADME$_{\ln UI}\left( t\right) $, $\widehat{ADME}_{n,\ln UI}\left( t\right) $ as in ((ref)), and construct 90% bootstrap pointwise and simultaneous confidence intervals.
Figure (ref) reports the average effects of a 10% increase in UI benefits on the unemployment duration distribution in each of the four subsamples. The solid lines are the point estimates and the dark and light shaded areas are the 90% bootstrap pointwise and simultaneous confidence intervals, respectively.
The results show interesting heterogeneity of UI benefits effects with respect to liquidity constraint proxies. For those workers with net wealth above the median, an increase in UI benefits has no statistically significant effect on the unemployment duration distribution, except for $ t=\left\{ 17,18\right\} $. Thus, for those workers who are not constrained (and, therefore, for whom the liquidity effect is approximately zero), the moral hazard effect seems to be close to zero. On the other hand, for those workers with net wealth below the median, an increase in UI benefits is associated with lower probabilities of finding a job. The conclusion using mortgage as a proxy for liquidity constraint is qualitatively the same. Note that our analysis suggests that the ADME of an increase of UI is non-monotone across the unemployment duration distribution, highlighting the flexibility of the DR approach. Indeed, using our proposed test for all slope coefficients being constant, the proportional hazard specification is rejected at the 10% significance levels in each subsample.
In other to further highlight that the effect of UI benefits on unemployment duration is different depending on whether a worker is likely to be liquidity constrained or not, in Figure (ref) we plot the difference of the estimated ADME between liquidity constrained and not-liquidity constrained workers, together with their pointwise and simultaneous confidence bands. Panel (a) suggests that UI benefits have a more negative effect on workers with net wealth below the median than on wealthier workers, and that this difference is statistically significant at the 10% level. Panel (b) reveals a qualitatively similar pattern when one compared workers with mortgage with those without mortgage, though the magnitude of this difference is less pronounced (but we can still reject the null hypothesis that the effects are the same, at the 10% significance level.)
Taken together, our results provide suggestive evidence that UI benefits have a non-monotone effect on the unemployment duration distribution and that such an effect varies whether workers are likely to be liquidity constrained or not. More precisely, our results suggest that the effect of UI on unemployment duration is larger for liquidity constrained workers. Through the lens of the results of Chetty2008, our findings suggest that an increase in UI benefits affects unemployment duration not only through moral hazard but also because of a \textquotedblleft liquidity effect.\textquotedblright
We are most grateful to the editor and two referees for their constructive comments which have lead to a improved paper. Research funded by Ministerio Econom\'{\i}a y Competitividad (Spain), ECO2-17-86675-P, and by Comunidad de Madrid (Spain), MadEco-CM S2015/HUM-3444. Andr\'{e}s Garc\'{\i}a-Suaza was supported by the Colombia Científica-Alianza EFI Research Program, with code 60185 and contract number FP44842-220-2018, funded by The World Bank through the call Scientific Ecosystems, managed by the Colombian Ministry of Science, Technology and Innovation. Part of this article was written when Pedro H. C. Sant'Anna was visiting the Cowles Foundation at Yale University, whose hospitality is gratefully acknowledged.