Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
42,989 characters · 7 sections · 49 citation commands
Linear Programming Approach to Nonparametric Inference under Shape Restrictions: with an Application to Regression Kink Designs
Nonparametric inference under shape restrictions is often computationally demanding. For instance, inference based on test inversion would require a grid search over a high-dimensional sieve parameter space. In this paper, we propose a computationally attractive method for nonparametric inference about regression functions under shape restrictions. Notably, our method can be implemented via a linear programming, despite the complicated nature of nonparametric inference under shape restrictions.
In many applications, economic structures often motivate shape restrictions, and such restrictions may contribute to delivering more informative statistical inference about the economic structure and causal effects. We highlight a case in point in the context of the regression kink design nielsen2010estimating,card2015inference,dong2016jump. Estimation and inference in the RKD rely on derivative estimators of nonparametric regression functions, which typically suffer from slow convergence rates and thus may lead to wide confidence intervals. On the other hand, there are often natural and economically motivated restrictions in the levels and slopes of the regression function to the left and/or right of the kink location, and they can contribute to shrinking the lengths of the confidence interval. In the context of the regression discontinuity design, armstrong2015adaptive and babii2019isotonic suggest usage of shape restrictions with related motivations. The benefits of shape restrictions may well be even greater for the RKD than for the regression discontinuity design due to the slower convergence rates of the RKD estimators.
We are far from the first to study the problem of nonparametric inference under shape restrictions. dumbgen2003optimal, cai2013adaptive, armstrong2015adaptive, chernozhukov2015constrained, horowitz2017nonparametric, chen2018shape, freyberger2018inference, mogstad2018using, fang2019general, and zhu2020inference, among others, propose various approaches to nonparametric inference under shape restrictions. See chetverikov2018econometrics and the journal issue edited by samworth2018special for a comprehensive review of the related literature. We advance the frontier of this literature by providing a computationally attractive approach. Specifically, we provide a novel method of constructing confidence bands/regions/intervals whose boundaries can be fully characterized as solutions to linear programs.
This paper is closely related to freyberger2015identification, who have considered a linear programming approach to inference under shape restrictions. Specifically, they propose a linear programming approach to inference about linear functionals of finite-dimensional parameters, where the parameter values are the values of the regression function evaluated at finite support points.\footnote{fang2020inference also propose a linear programming approach to inference for a growing number of linear systems, although their focus is different from nonparametric regression functions under shape restrictions as in this paper.} On the other hand, as acknowledged in freyberger2015identification, “[t]he use of shape restrictions with continuously distributed variables is beyond the scope of” their paper. We contribute to this literature by accommodating (discretely or continuously) infinite-dimensional parameters. This extended framework allows for analysis of nonparametric regressions with infinitely supported (discrete or continuous) regressors, which are relevant to many applications including the regression discontinuity and kink designs among others.
Our proposed inference procedure works as follows. First, we use the sieve approximation chen2007large of the nonparametric regression function. We then construct a supremum test statistic as a linear function of the sieve parameters, compute its critical value by applying chernozhukov2017, and then translate their relation into an inequality constraint. Subject to this inequality constraint, together with the additional linear-in-sieve-parameter inequality constraints stemming from shape restrictions, we find the lower (respectively, upper) bound of the confidence band/interval by the minimizing (respectively, maximizing) the sieve representation with respect to the sieve parameters. In the final step, we inflate the bounds by a sieve approximation error bound similarly to armstrong2018finite,armstrong2020simple, noack2019bias, schennach2020bias, and katosasakiura2020.
The rest of this paper is organized as follows. Section (ref) presents the model and an overview of the proposed procedure. Section (ref) presents the size control. Section (ref) describes the procedure when we are interested in a finite-dimensional linear feature of the regression function. Section (ref) presents an application of the RKD, with detailed implementation procedures tailored to this application. In an empirical application, we demonstrate that shape restrictions can shrink the lengths of the confidence interval. Section (ref) concludes. Mathematical proofs and simulation analysis are collected in the appendix.
Throughout this paper, we assume that a data set $\{(Y_i,X_i^T):i=1,\dots,n\}$ consists of i.i.d. random vectors following the law of $(Y,X^T)$, where $Y$ is a real-value random variable and $X$ is a finite-dimensional random variable with the support $\mathcal{X}\subset\mathbb{R}^{\dim X}$. Let $E_n$ denote the sample mean, that is, $E_n[f(Y,X^T)]\equiv\frac{1}{n}\sum_{i=1}^nf(Y_i,X_i^T)$ for any measurable function $f$.
In this paper, we are interested in a linear feature of the unknown mean regression function $g_0(x)\equiv E[Y\mid X=x]$, so that the parameter of interest can be written as $$ \theta_0\equiv\mathbf{A}_0 g_0 $$ for a known linear operator $\mathbf{A}_0$. We assume this parameter $\theta_0$ to be a function from some set $\mathcal{W}_0$ into $\mathbb{R}$, which allows $\theta_0$ to be a scalar, a vector, or a function from $\mathcal{X}$ into $\mathbb{R}$. For example, when $\mathbf{A}_0$ is the identity function, the parameter of interest is the conditional mean function $g_0$ itself. Other examples for $\theta_0$ include $g_0(x)$ for a given point $x$, the integral $\int g_0(x)d\mu(x)$, and the derivative $\partial g_0(x)/\partial x_j$, among others. In Section (ref), we discuss how we can tailor the procedure to the case when $\theta_0$ is finite dimensional.
The objective of this paper is to construct a confidence region for $\theta_0$ under the shape restrictions
for a known linear operator $\mathbf{A}_1$.\footnote{In this paper, the shape restriction does not have any improvement in the identification analysis, because $g_0$ is identified over $\mathcal{X}$ and therefore $\theta_0$ is identified.} We are going to construct a confidence region ${CR}_{\theta}$ for $\theta_0$ satisfying the following two properties: (i) the boundaries of ${CR}_{\theta}$ are the set of solutions to linear programming problems; and (ii) ${CR}_{\theta}$ controls the asymptotic size under the shape restriction.
We approximate $g_0$ by a linear combination of $k$ functions $p_1,\ldots,p_k$ on $\mathcal{X}$.\footnote{Recall that $\mathcal{X}$ is the support of $X$. We assume $k\geq 2$, which guarantees $\log k\geq 0$.} These $k$ functions are denoted by $$ p_{1:k}\equiv (p_1,\ldots,p_k)^T. $$ We can consider the linear regression of $Y$ on $p_{1:k}(X)$, and the population coefficient vector for this regression is $$ \bar{\beta}\equiv E[p_{1:k}(X)p_{1:k}(X)^T]^{-1}E\left[p_{1:k}(X)Y\right]. $$ With these definitions and notations, we make the following assumption about error bounds for the approximation of $g_0$ by $p_{1:k}^T\bar{\beta}$.
This assumptions plays the role of restricting the function class where $g_0$ resides, similarly to katosasakiura2020 in the spirit of the honest inference approach armstrong2018finite,armstrong2020simple and the bias bound approach schennach2020bias.\footnote{We allow $k$, $\delta_0$ and $\delta_1$ to be a function of $n$. We do not require $k\rightarrow\infty $ as $n\rightarrow \infty$ but it is allowed. In Assumption (ref), we bound the biases coming from the approximation of $g_0$ by $p_{1:k}^T\bar{\beta}$ by known $\delta_0$ and $\delta_1$. Without accounting for such approximation bounds, conventional methods would set $\delta_0\rightarrow 0$ and $\delta_1\rightarrow 0$ as $n\rightarrow 0$ in light of that the bias asymptotically vanishes with undersmoothing. That said, by Assumption (ref), we take this honest or bias bound approach in this paper for the sake of generality, with the special case of undersmoothing leading to the conventional approach in particular.}
For a generic value $\beta\in\mathbb{R}^k$, we can implement a hypothesis testing for the null hypothesis $H_0: \bar\beta=\beta$ against the alternative hypothesis $H_1: \bar\beta\ne\beta$ as follows. In this hypothesis testing problem, we aim to detect a violation of the null hypothesis $$ H_0: E[p_{1:k}(X)(Y-p_{1:k}(X)^T\beta)]=0, $$ which is equivalent to $\bar\beta=\beta$ under the invertibility of $E[p_{1:k}(X)p_{1:k}(X)^T]$. We can estimate the left hand side of the above equation by $E_n[p_{1:k}(X)(Y-p_{1:k}(X)^T{\beta})]$ and its asymptotic variance (under $H_0$) by $E_n[\hat{\omega}\hat{\omega}^T]$, where $$ \hat{\omega}\equiv p_{1:k}(X)(Y-p_{1:k}(X)^T E_n\left[p_{1:k}(X)p_{1:k}(X)^T\right]^{-1}E_n\left[p_{1:k}(X)Y\right]). $$ Note that $\hat{\omega}$ estimates $\omega\equiv p_{1:k}(X)(Y-p_{1:k}(X)^T\bar{\beta})$. With these estimates, we consider the test statistic $$ \left\|E_n[\hat{\omega}\hat{\omega}^T]^{-1/2}E_n[p_{1:k}(X)(Y-p_{1:k}(X)^T{\beta})]\right\|_{\infty}. $$ To obtain a critical value, we apply the multiplier bootstrap by calculating the $(1-\alpha)$ quantile, denoted by ${cv}$, of $$ \left\|E_n[\hat{\omega}\hat{\omega}^T]^{-1/2}E_n[\eta\hat{\omega}] \right\|_{\infty} $$ conditional on the data set, where $\eta_1,\ldots,\eta_n$ are independent Rademacher multiplier random variables that are independent of the data. Note that the critical value ${cv}$ does not depend on a specific value of $\beta$, which enables us to construct a confidence region characterized by linear inequalities for $\beta$.
We can construct a confidence region for $\theta_0$ based on the test inversion. Using the test statistic and the critical value, we can define a confidence region for $\theta_0$, denoted by ${CR}_{\theta}$. Namely, $CR_{\theta}$ is the set of $\theta$ satisfying the following linear constraints for some $\beta\in\mathbb{R}^k$:
where $[\mathbf{A}_0p_{1:k}^T](w_1)\beta\equiv [\mathbf{A}_0(p_{1:k}^T\beta)](w_1)$ and $[\mathbf{A}_1p_{1:k}^T](w_1)\beta\equiv [\mathbf{A}_1(p_{1:k}^T\beta)](w_1)$.
In the definition of ${CR}_{\theta}$, we have three types of linear constraints. First, (ref) comes from the hypothesis test for $H_0: \bar\beta=\beta$. Second, (ref) controls the approximation error between $\mathbf{A}_0p_{1:k}^T\bar\beta$ and $\theta_0$ under (ref) in Assumption (ref). Third, (ref) uses the knowledge that the shape restriction (ref) holds for true $g_0$, together with (ref) in Assumption (ref). This confidence region could be more informative than that without the shape-restriction inequalities in (ref).
For every value $w_0 \in \mathcal{W}_0$, the following theorem states that the projection of ${CR}_{\theta}$ to $\theta_0(w_0)$ can be computed by solving two linear programming problems. A proof is provided in Appendix (ref).
Therefore, the boundary points are the solutions to linear programs.
For the asymptotic size control, we are going to impose the following assumptions. Let $b>0$, $q \in [4,\infty), \nu\in (2,\infty)$ be some constants and let $B_n\ge 1$ denote a sequence of finite constants that may possibly diverge to infinity. Consider the following assumption.
Assumption (ref) (a) implies Condition A.2 in Assumption belloni2015some. It imposes a restriction to rule out overly strong co-linearity among $p_1,\dots,p_k$. Assumptions (ref) (b)-(ii) and (ref) (b)-(iii) correspond to Conditions (M.1), (M.2) and (E.2) in chernozhukov2017. It requires that the polynomial moments of the maximal component of normalized $\omega$ will not be growing too fast, as well as it imposes conditions that dictate how fast the number of basis functions can grow. The maximum is allowed to be growing at a rate of $O(n^{a})$ for some $a$ between zero and one. Assumption (ref) (c) covers Conditions A.3-A.5 in belloni2015some as well as rate conditions in the statement of their Theorem 4.6. Assumption (ref) (c)-(i) requires the residual to have a finite $\nu$-th moment for some $\nu>2$. Assumptions (ref) (c)-(ii) and (ref) (c)-(iii) impose bounds on the approximation errors of $g_0$ using $p_1,\dots,p_k$, as well as restrictions on the size of basis functions, measured by the Euclidean norm and the Lipschitz constant. Assumption (ref) (c)-(iv) imposes some more constraints on the relative growth rates of the approximation errors, the size and number of basis functions. Notice that it does not require the approximation errors to be diminishing asymptotically, and hence does not require undersmoothing.
The following theorem states the asymptotic size control for ${CR}_{\theta} $ as a confidence region for $\theta_0$. A proof is provided in Appendix (ref).
With some additional notations and rate conditions, it is possible to strengthen the statement of Theorem (ref) to hold uniformly over a set of data generating processes. This is due to the fact that key theoretical building blocks in the proof of Theorem (ref) -- i.e. the anti-concentration inequality in chernozhukov2015comparison, the high-dimensional central limit theorem of chernozhukov/chetverikov/kato:2017, and Rudelson's concentration inequality belloni2015some -- all provide non-asymptotic bounds with constants only depending on a few key features of the model such as $b,q$ and $\nu$.
When the parameter of interest $\theta_0$ is finite dimensional, we can directly test $\mathbf{A}_0[p_{1:k}^T\bar\beta]=\theta$ for a generic value of $\theta$, instead of testing $\bar\beta=\beta$ as in Section (ref). In the current section, we describe the inference procedure when $\theta_0$ is a finite-dimensional column vector.
For a generic value $\theta$, we consider the null hypothesis $H_0: A_{0,k}\bar\beta=\theta$ and the alternative hypothesis $H_1: A_{0,k}\bar\beta\ne\theta$, where $A_{0,k}$ is the matrix defined by $A_{0,k}\beta=\mathbf{A}_0[p_{1:k}^T\beta]$ for every $k\times 1$ vector $\beta$. Based on the definition of $\bar\beta$, we aim to measure the violation of the null hypothesis $$ H_0: A_{0,k}E[p_{1:k}(X)p_{1:k}(X)^T]^{-1}E\left[p_{1:k}(X)Y\right]=\theta. $$ We can estimate the left hand side by $A_{0,k}E_n\left[p_{1:k}(X)p_{1:k}(X)^T\right]^{-1}E_n\left[p_{1:k}(X)Y\right]$ and its the asymptotic variance under $H_0$ by $$ \hat{V}\equiv A_{0,k}E_n\left[p_{1:k}(X)p_{1:k}(X)^T\right]^{-1}E_n[\hat{\omega}\hat{\omega}^T]E_n\left[p_{1:k}(X)p_{1:k}(X)^T\right]^{-1}A_{0,k}^T. $$ With these estimators, we consider the test statistic $$ \left\| \hat{V}^{-1/2}(A_{0,k}E_n\left[p_{1:k}(X)p_{1:k}(X)^T\right]^{-1}E_n\left[p_{1:k}(X)Y\right]-\theta) \right\|_{\infty}. $$ To obtain its critical value, we apply the multiplier bootstrap and compute the $(1-\alpha)$ quantile, denoted by $\widehat{cv}$, of $$ \left\| \hat{V}^{-1/2}A_{0,k}E_n\left[p_{1:k}(X)p_{1:k}(X)^T\right]^{-1}E_n[\eta\hat{\omega}] \right\|_{\infty} $$ conditional on the data set, where $\eta_1,\ldots,\eta_n$ are independent Rademacher multiplier random variables that are independent of the data.
A confidence region for $\theta_0$ can be constructed based on the test inversion. In this setup, we can construct a confidence region for $\theta_0$, $\widehat{CR}_{\theta}$, by collecting all $\theta$'s satisfying the following linear constraints for some $\beta\in\mathbb{R}^k$:
$$ |[\mathbf{A}_0p_{1:k}^T](w_0)-\theta(w_0)|\leq\delta_0(w_0)\mbox{ for every }w_0\in\mathcal{W}_0,\mbox{ and } $$
For every value $w_0 \in \mathcal{W}_0$, we can compute the projection of $\widehat{CR}_{\theta}$ to $\theta_0(w_0)$ by solving two linear programming problems w.r.t. $\beta$: $$ \mbox{ minimize }[A_{0,k}\beta](w_0)-\delta_0(w_0)\mbox{ over $\beta$ subject to } \eqref{eq:sample_error_control_proj} \ \& \ \eqref{eq:additiona_restriction_proj}, $$ and $$ \mbox{ maximize }[A_{0,k}\beta](w_0)+\delta_0(w_0)\mbox{ over $\beta$ subject to } \eqref{eq:sample_error_control_proj} \ \& \ \eqref{eq:additiona_restriction_proj}. $$ In other words, the projection is the closed interval $$ \left[\ \underset{\text{s.t.}\ \eqref{eq:sample_error_control_proj}\&\eqref{eq:additiona_restriction_proj}}{\min_{\beta}} [A_{0,k}\beta](w_0)-\delta_0(w_0),\ \ \underset{\text{s.t.}\ \eqref{eq:sample_error_control_proj}\&\eqref{eq:additiona_restriction_proj}}{\max_{\beta}} [A_{0,k}\beta](w_0)+\delta_0(w_0)\ \right]. $$
Formal theoretical properties of the the confidence interval constructed by this procedure follow from analogous arguments to those in Sections (ref) and (ref). In the application presented in the following section, the parameter $\theta_0$ of interest is a scalar (and finite dimensional in particular) and we therefore adopt this approach to constructing its confidence interval.
In this section, we present an application of our proposed method to the regression kink design (RKD). Since the regression kink design is based on estimates of slopes as opposed to levels, statistical inference based on nonparametric estimates often entails slow convergence rates and thus wide confidence intervals. To mitigate this adverse feature of the regression kink design, we propose to impose shape restrictions that are motivated by the underlying economic structures.
To introduce the RKD, consider the structure $$ Y=Y(T,X,U)\mbox{ and }T= T(X), $$ where $Y$ denotes the outcome variable, $T$ denotes the treatment variable, $X$ denotes the running variable, and $U$ denotes the random vector of unobserved characteristics. A researcher is often interested in the partial effect $\partial Y(T,X,U) / \partial T$ of the treatment variable on the outcome variable. Since the unobserved characteristics $U$ are generally correlated with the running variable $X$ and thus with the treatment $T=T(X)$, one would need to exploit exogenous variations in the treatment variable in order to identify this partial effect. If the treatment policy function $T(\cdot)$ exhibits a `kink' at a known point $\bar{x}$, then this shape restriction can be exploited to induce local exogenous variations in the treatment variable $T$ as well, so that the partial effect of interest may be identified. This approach of the so-called regression kink design (RKD) was proposed by nielsen2010estimating and card2015inference -- see dong2016jump for the case of a binary treatment, and see chiang2019causal and chen2020quantile for heterogeneous treatment effects.
Suppose that a researcher is interested in conducting inference for the average partial effect $h^1(\bar{x})\equiv E\left[\left. {\partial Y(T,X,U)}/{\partial T} \right\vert X=\bar{x}\right]$ at the kink point $\bar{x}$. Under regularity conditions, we can obtain the following decomposition of the derivative $g_0'(X)$ of $g_0(x) = E[Y|X=x]$:
If $T'(\cdot)$ is discontinuous (i.e., $T(\cdot)$ is kinked) at $\bar{x}$ while each of $h^1$, $h^2$ and $h^3$ is continuous at $\bar{x}$, then this decomposition implies that the partial effect of interest at $\bar{x}$ can be identified by $$ h^1(\bar{x}) = \frac{\lim_{x\downarrow\bar{x}} g_0'(x) - \lim_{x\uparrow\bar{x}} g_0'(x)}{\lim_{x\downarrow\bar{x}} T'(x) - \lim_{x\uparrow\bar{x}} T'(x)}, $$ cf. nielsen2010estimating,card2015inference. We can represent the parameter of interest via $h^1(\bar{x})=\mathbf{A}_0 g_0$, using a linear operator $\mathbf{A}_0$ defined by
Even though $g_0$ is unknown, the operator $\mathbf{A}_0$ is known since $T(\cdot)$ is a known function. In this case, $\mathcal{W}_0=\{\bar{x}\}$, and the parameter of interest $\theta_0=\mathbf{A}_0 g_0$ is a scalar.
Although $\theta_0$ is nonparametrically estimable, an estimator based on slopes of nonparametric regression functions usually suffers from slow rates of convergence, and thus it may not provide an informative confidence interval. If an economic structure motivates shape restrictions, then imposing such restrictions may conceivably contribute to shrinking the length of the confidence interval. With this motivation, in Section (ref), we demonstrate how shape restrictions help in conducting statistical inference in the analysis of of unemployment insurance (UI).
Unemployment insurance (UI) benefits play important roles in supporting consumption smoothing under the risk of unemployment. A potential drawback of the UI benefits is the moral hazard effects, that is, the UI benefits may discourage unemployed workers from looking for jobs, leading to elongated unemployment durations and thus economic inefficiency. Identifying and estimating these moral hazard effects have been of research interest in labor economics. landais2015assessing suggests to exploit the non-smooth UI benefit schedule as detailed below, and thus to use the regression kink design to identify the effects of UI benefits on the duration of unemployment. Applying this identification strategy to the data of the Continuous Wage and Benefit History Project moffitt1985effect, landais2015assessing finds that there are positive effects of the UI benefit amounts on the duration of unemployment, even after controlling for unobserved source of endogenous selection of the duration that may be correlated with the pre-unemployment income and thus the benefit amount. chiang2019causal further investigate heterogeneous effects of the UI benefit amount on the duration by using the quantile regression kink design.
landais2015assessing considers the following empirical framework of assessing the welfare effects of unemployment benefits. The outcome $Y$ of interest is the duration of unemployment. Upon becoming unemployed, an individual can apply for UI and receives a weekly benefit amount of $T=T(X)$, where $X$ is the highest quarterly earning in the last four completed calendar quarters prior to the date of the UI claim. The partial effect $\partial Y(T,X,U) / \partial T$ measures the moral hazard effect of the UI benefits on the duration of unemployment in this setting. Since the unobserved characteristics $U$ contain cognitive and non-cognitive skills of the individual, such as attitudes toward work, that are generally correlated with the labor income $X$ received prior to the unemployment, one would need exogenous variations in the treatment variable in order to identify this moral hazard effect.
As in landais2015assessing, we can exploit the fact that the UI benefits policy $T(\cdot)$ exhibits a kinked shape. In particular, the UI schedule in the state of Louisiana is linear in $X$ with a constant $t\equiv 1/25$ of proportionality up to a fixed ceiling $t_{\max}$. (Note that the unit of $X$ is U.S. dollars per quarter, whereas the unit of $T(X)$ is U.S. dollars per week. Therefore, this constant of proportionality implies that the UI benefit amount is approximately a half of the prior earnings.) The maximum UI benefit amount is $\bar{t}=$ \$183 during the period between September 1981 and September 1982, and $\bar{t}=$ \$205 during the period between September 1982 and December 1983. In short, the UI benefits policy takes the form of $$ T(x) =
$$ and $T$ is thus kinked at $\bar{x}= t_{\max}/t$. Individuals can continue to receive the benefits determined by this formula as far as they remain unemployed up to the maximum duration of 28 weeks.
We construct a data set by following the data construction in landais2015assessing and chiang2019causal. We focus on the observations in Louisiana. The sample size of the original data is 9,008 for the period between September 1981 and September 1982, and 16,463 for the period between September 1982 and December 1983. Since we are interested in the information around the kink location $\bar{x}$, for simplicity, we focus on the (sub-)sample of the observations in the interval $X\in[\bar{x}-5000,\bar{x}+5000]$. The resultant sample size is 8,677 for the period between September 1981 and September 1982, and the resultant sample size is 15,763 for the period between September 1982 and December 1983.
In this empirical application, we can consider a few shape restrictions on the unknown conditional mean function $g_0(x)=E[Y\mid X=x]$. First of all, to impose the continuity of $g_0$ at $\bar{x}$, we can use the shape restriction
This restriction is not redundant when we use difference sieves for the left of $\bar{x}$ and the right of $\bar{x}$. Moreover, it may be reasonable to assume that $h^2$ and $h^3$ are both non-increasing. Specifically, the direct effect $h^2$ is non-increasing if formerly higher-income earner can find the next job more quickly than formerly lower-income earners on average. The endogenous effect $h^3$ is non-increasing if individuals with higher abilities can find the next job more quickly than those with lower abilities on average. Since $T(\cdot)$ is a constant function to the right of the kink location in this application, this assumption together with the decomposition (ref) implies that the reduced form $g_0$ is non-increasing to the right of the kink location $\bar{x}$. This consideration leads to the slope restriction
In the notations in Section (ref), we can summarize the shape restrictions (ref) and (ref) as
where $\mathcal{W}_1=\{-2,-1\}\cup\{w_1:w_1>\bar{x}\}$ and $$ [\mathbf{A}_1 g] (w_1) =
$$
Now, we outline the concrete implementation procedure to exploit these shape restrictions (ref), for inference about the causal parameter $\theta_0 = \mathbf{A}_0 g_0$ defined in (ref). For every even natural number $k$, we use the basis functions $$ p_{1:k}=(\ell_{L,0},\ell_{R,0},\cdots,\ell_{L,k/2-1},\ell_{R,k/2-1}), $$ where $\left(\ell_{L,0},\ell_{L,1},\cdots,\ell_{L,k/2-1}\right)$ are the first $k/2$ terms of an orthonormal basis for $L^2([\bar{x}-5000,\bar{x}])$ and $\left(\ell_{R,0},\ell_{R,1},\cdots,\ell_{R,k/2-1}\right)$ are the first $k/2$ terms of an orthonormal basis for $L^2([\bar{x},\bar{x}+5000])$. We use the shifted Legendre bases in the empirical application in this subsection as well as in the simulation studies in Section (ref). We follow Section (ref) to construct the $(1-\alpha)$-level confidence interval for $\theta_0$ subject to the shape constraint (ref), where we restrict $\mathcal{W}_1=\{-2,-1\}\cup\{\xi_1,\dots,\xi_l\}$ with 99 equally spaced grid points $\{\xi_1,\dots,\xi_l\} \subset (\bar{x},\bar{x}+5000)$. The following algorithm provides a step-by-step procedure of the construction.
Table (ref) summarizes the results for the statistical inference about the marginal effects of UI benefits on unemployment duration in Louisiana, based on the above algorithm. Displayed are the 95% confidence intervals and their lengths for each of the period between September 1981 and September 1982 (top panel) and the period between September 1982 and December 1983 (bottom panel). We use the largest sieve dimension $k=12$ among those that were used in our simulation studies presented in Appendix (ref). (The shape restrictions do not bind for the cases of $k=4$ or $k=8$. It is possibly because the current sample sizes are much larger than those used in our simulation studies.) For the UI benefit amount $T(X)$, we use two alternative measures. One is the amount of UI benefits claimed (left half of each panel) and the other is the amount of UI benefits actually paid (right half of each panel) by following the prior work. That said, these two alternative measures provide almost the same results, and therefore our discussions below apply to the results based on both of the two measures.
The reported confidence intervals contain the point estimates reported in the prior work by landais2015assessing. That said, the econometric specifications are different, and results are thus hard to compare. Our results based on no shape restriction are effectively what we would get from the standard method with running the fifth-degree polynomial regressions on each side of the left and right of $\overline{x}$. In contrast, landais2015assessing uses the polynomials of degree one, i.e., the linear specification, for the main estimation results reported in his Table 2. Due to the greater flexibility of our econometric specification, our method naturally incurs wider confidence intervals, but we demonstrate that shape restrictions will contribute to providing more informative results.
Our confidence interval includes the zero for the period between September 1981 and September 1982 (the first panel of Figure (ref)) if no shape restriction is imposed, i.e., if the conventional approach is taken. However, in this panel (for the period between September 1981 and September 198), shape restrictions (ref) shrink the confidence intervals. (Although these shrunken confidence intervals have their lower bounds approximately 0.000, note that we do not directly impose a sign restriction on the causal effects per se, in the shape restrictions (ref). See our discussions above (ref) for motivations of these shape restrictions.) On the other hand, the confidence intervals are already informative for the period between September 1981 and September 1982 even without any shape restriction, and imposing shape restrictions (ref) therefore will not contribute to shrinking the confidence intervals. These results thus demonstrate one case in which shape restrictions contribute to enhancing the informativeness of statistical inference, and another case in which they do not.
Nonparametric inference under shape restrictions can demand high computational burdens, e.g., a grid search over a high-dimensional sieve parameter space. In this paper, we provide a novel method of constructing confidence bands/intervals for nonparametric regression functions under shape constraints. The proposed method can be implemented via a linear programming, and it thus relieves the conventional computationally burdens. A usage of this new method is illustrated with an application to the regression kink design. Inference in the regression kink design often suffers from wide confidence intervals due to the slow convergence rates of nonparametric derivative estimators. If economic models and structures motivate shape restrictions, then these restrictions may contribute to shrinking the confidence interval. We demonstrate this point with real data for an analysis of the causal effects of unemployment insurance benefits on unemployment durations. Specifically, for analysis of the effects of unemployment insurance benefits on the unemployment duration, the shape restrictions motivated by non-increasing direct effects and non-increasing endogenous effects drastically shrink the confidence interval of causal effects.