EconBase
← Back to paper

Testing Shape Restrictions with Continuous Treatment: A Transformation Model Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

46,478 characters · 10 sections · 38 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing Shape Restrictions with Continuous Treatment: A Transformation Model Approach

abstractWe propose tests for the convexity/linearity/concavity of a transformation of the dependent variable in a semiparametric transformation model. These tests can be used to verify monotonicity of the treatment effect, or, equivalently, concavity/convexity of the outcome with respect to the treatment, in (quasi-)experimental settings. Our procedure does not require estimation of the transformation or the distribution of the error terms. The statistic takes the form of a U statistic or a localised U statistic, and we show that critical values can be obtained by bootstrapping. In our application we test the convexity of loan demand with respect to the interest rate using experimental data from South Africa. \\ \\ JEL: C12, C21, C14 \newline Keywords: Shape restrictions, Transformation model, Bootstrap, U statistic, Treatment effects

Introduction

In this paper we consider testing if a treatment has a diminishing or increasing effect on the outcome. This is often of interest on top of the question if there is an effect at all or what sign it has. For example, often it is natural to expect that demand will be decreasing in price. However, it is less clear if increasing the price will have a larger effect at lower or higher price levels. Our test can address this question without fully estimating the demand relationship.

Let $X$ denote the vector of treatment and control variables. As estimating the effect of $X$ on the outcome $Y$ nonparametrically with non-trivial number of treatments and controls would suffer from a curse of dimensionality, we impose a single-index structure and assume that the nonlinearity of the treatment effect comes from a nonlinear transformation of $Y$. In other words, consider a transformation model of the form:

gather[gather omitted — 50 chars of source]

where $Y$ is a scalar dependent variable, $X$ is a vector of $q$ nondegenerate explanatory variables, $\beta_0$ is a vector of coefficients belonging to a compact set $\Theta_{\beta} \subset \mathbb{R}^q$, $T(\cdot)$ is an increasing function and $\varepsilon$ is an unobserved error term with distribution $F$ that is independent of $X$. In order to identify the model we need a location normalisation: e.g. $T(0)=0$, $E(\varepsilon)=0$ or $Me(\varepsilon)=0$; and a scale normalisation: e.g. $|\beta_{0,1}|=1$. The benefit of using the transformation model compared to a standard single-index model (i.e. $Y=T(X'\beta_0) + \varepsilon$), besides the fact that it facilitates our testing approach, is that the transformation model allows the treatment effect of $X_{k}$ to depend on the values of both observed and unobserved characteristics (note that $Y=T^{-1}(X'\beta_0+\varepsilon)$) so can be seen as a simple way of introducing heterogenous treatment effects that vary with unobservables.

The main objective of this article is to develop a practically appealing test to determine whether the transformation function $T(\cdot)$ is concave/linear/convex. The examples below illustrate the importance of testing curvature of the transformation.

exmp(Experiments with continuous treatment) Using normalization $E(\varepsilon) = 0$ and $Q_{\alpha}(\varepsilon) = 0$ where $Q_{\alpha}$ denotes the $\alpha$ quantile, we have, respectively: \begin{align*} \frac{\partial^{2} E(Y|X)}{\partial X_{k}^{2}} &= - \beta_{0,k}^{2} E\left[ \frac{T”(T^{-1}(X'\beta_0+\varepsilon))}{T'(T^{-1}(X'\beta_0+\varepsilon))^{3}}\bigg| X\right] \qquad (mean regression) \\ \intertext{and} \frac{\partial^{2} Q_{\alpha}(Y|X)}{\partial X_{k}^{2}} &= -\beta_{0,k}^{2} \frac{T”(T^{-1}(X'\beta_0))} {T'(T^{-1}(X'\beta_0))^{3}} \qquad (quantile regression) \end{align*} As $T'(\cdot)>0$ the curvature of the mean/quantile effect depends on $T''(\cdot)$. Thus, the sign of the second derivative $T''(\cdot)$ determines if the effect of the treatment $X$ on the $\alpha$ quantile of $Y$ is concave or convex in $X$. For example, if a company randomises marketing spending in different markets, testing for concavity would answer the question if marketing spending has diminishing returns (e.g. on mean revenue). Test of curvature has also application in experimental studies of demand elasticities (e.g. jessoe_rapson14, hainmueller_et_al15, karlan_zinman19) where it can be used to verify if demand is concave, linear or convex. Finally, when the transformation is linear, the treatment effect does not depend on the observed ($X$) and unobserved ($\varepsilon$) heterogeneity. Thus, the test of linearity of $T$ can be seen as a test of treatment effect heterogeneity.
exmp(Duration models: testing hazard monotonicity) Let $\lambda(\cdot)$ and $\Lambda(\cdot)$ denote baseline hazard and integrated baseline hazard, respectively. In a duration model: $T(Y) = \log \Lambda (Y)$, and \begin{gather*} T”(Y) = \frac{\lambda'(Y)}{\Lambda(Y)}-\left(\frac{\lambda(Y)}{\Lambda(Y)}\right)^{2}. \end{gather*} Hence, rejecting concavity of $T(\cdot)$ (i.e. $T''(\cdot)<0$) implies that the baseline hazard is non-decreasing ($\lambda'(\cdot)\geq 0$). One can, thus, use the test of concavity of the transformation as a test for monotonicity of the baseline hazard, or in other words, as a test for positive duration dependence. In the economic context, one may be interested in detecting non-monotonicity of unemployment exit rate due to unemployment benefit exhaustion effects (see card_et_al07b for discussion).

Beyond these examples our procedure can also be used for specification search, i.e. determining if one should use a concave or convex transformation, and to test the curvature in wage regressions, e.g. if the effect of education or experience is concave, or the curvature of the marginal utility (or profit) function in hedonic models (see ekeland_et_al04).

Our test statistic simply compares triples of $Y$'s corresponding to equally spaced index values $X'\beta_0$, thus it does not require estimation of the transformation function $T$ or the distribution of $\varepsilon$. We only require an estimator of $\beta_0$ and, basically, the symmetry of the distribution of the error terms. Estimating $\beta_0$ can be done, for example, by using the maximum rank correlation estimator, han87, or semiparametric least squares, ichimura93.

We propose both a global test that has power to detect globally convex or concave functions and leads to asymptotic normal critical values, as well as a more general test that detects local deviations from linearity, i.e. has power against alternatives that are both convex and concave on the domain of $T$. Our tests do not have a pivotal asymptotic distribution but we show that the critical values can be obtained by parametric bootstrap.

The cost of our test’s easy implementation is that the power properties of our local test are complex, and the test may not be consistent against some relevant alternatives. Nevertheless, our application demonstrates the usefulness of our testing approach in detecting nonconvexities in the demand for loans as a function of the interest rate.

Our statistic resembles the approach in abrevaya_jiang05 who test curvature in a nonparametric regression model. However, unlike their approach our test does not suffer from the curse of dimensionality due to the single index structure of the regression part and allows the marginal effect of covariates to vary with the error term (though in a manner restricted by the single-index). Also the details of the derivation of the asymptotic distribution are different due to presence of estimated $\beta_0$ in our statistic and somewhat distinct approach to obtaining power against general alternatives. These traits are shared by abrevaya_et_al10, who test monotonicity in a generalised regression model with endogeneity, though a distinguishing feature of our work is that we formally show validity of bootstrap for obtaining critical values.

Related literature includes tests for the sign of the treatment effect, see e.g. kline16, and testing for treatment effect heterogeneity (abadie02, crump_et_al08, santanna21, chernozhukov_et_al23). Unlike the latter papers, our approach allows the treatment effect to vary with (independent) unobserved heterogeneity at the cost of imposing much more structure on the treatment effect model and the covariates. Testing curvature of the transformation can also be seen as a generalisation of specification testing in neumeyer_et_al16 and szydlowski20. The idea of using curvature of the integrated hazard to test monotonicity of the baseline hazard has been utilised by hall_van_keilegom05. Testing shape restrictions in a nonparametric regression model has been considered by ghosal_et_al00, gutknecht16, chetvertikov19 and komarova_hidalgo23, among others. Similarly to this paper, chen_kato20 propose a bootstrap procedure for approximating the supremum of a local U-statistic.

The article is organized as follows. Section (ref) discusses the idea behind the testing procedure informally and the formal results are postponed till Sections (ref)-(ref). Sections (ref)-(ref) contain Monte Carlo results and our application to loan demand. Proofs, besides the proof of the main proposition, are located in the Appendix, which also contains some additional MC simulations.

Main idea

Figure (ref) portrays the intuition behind our test. The transformation plotted in the figure is concave and we display three ordered data points in the figure, $(Y_{i},Y_{j},Y_{k})$, for which the transformation function is equally spaced, i.e. $T(Y_{k})-T(Y_{j})=T(Y_{j})-T(Y_{i})$.

figure[figure omitted — 988 chars of source]

Concavity of $T(\cdot)$ implies that $Y_{k}-Y_{j}>Y_{j}-Y_{i}$. Note that $T(Y_{k})-T(Y_{j})=T(Y_{j})-T(Y_{i})$ is equivalent to $(X_{k}-X_{j})'\beta_0 +\varepsilon_{k}-\varepsilon_{j}= (X_{j}-X_{i})'\beta_0 +\varepsilon_{j}-\varepsilon_{i}$. Hence, “on average” equally spaced $T(Y)$'s mean equally spaced index values $X'\beta_0$. Therefore, we can detect deviations from concavity by considering the following criterion:

gather*[gather* omitted — 178 chars of source]

where $X_{ji} \equiv X_{j}-X_{i}$.

In order to make this criterion operational with continuous distribution of $X'\beta_0$ (which is required for identification) we need to replace the last indicator function with a smooth kernel $K_{h}(\cdot) = h^{-1}K(\cdot/h)$ and $\beta_0$ with its estimator $\hat{\beta}$:

gather*[gather* omitted — 189 chars of source]

Deviations from convexity can be detected in a similar fashion. Finally, we can combine both criterion functions to detect deviations from linearity:

gather[gather omitted — 177 chars of source]

Here, very negative values of the objective function signify convexity and large positive values mark concavity. Outside testing, one may use the above measure itself to characterise an “average” curvature of function $T$, e.g. large positive values suggest that the function is predominantly concave.

Note that our test requires only one-dimensional kernel smoothing, thus it does not suffer from the curse of dimensionality. Also, unlike the specification test in szydlowski20 it does not require estimation of the transformation function, a computationally intense task itself.

Formal definition and asymptotic theory

Global test

As the probability limit of the criterion functions introduced in the previous section depends on the extent of concavity, but these functions are asymptotically centred at known values under linearity, we form our procedure as a test of:

align*[align* omitted — 295 chars of source]

depending on the question of interest. We will use the test statistic

gather*[gather* omitted — 35 chars of source]

and reject the null hypothesis at level $\alpha$ if $S_{n}<c_{\alpha}, |S_{n}|>c_{1-\alpha/2}$ or $S_{n}>c_{1-\alpha}$ depending on $H_{0}$, respectively, where $c_{\alpha}$ denotes an $\alpha$ quantile from an appropriate asymptotic distribution.

Define $\varepsilon_{ji}=\varepsilon_j - \varepsilon_i$. Proposition (ref) justifies our testing strategy. We need the following symmetry assumption:

customassmpt{SYM} Define $f_{\Delta \varepsilon}(\cdot,\cdot)$ to be the joint distribution of $(\varepsilon_{ji}, \varepsilon_{kj})$. Assume that $\{(X_{i},Y_{i})\}_{i=1}^n$ are i.i.d and that $f_{\Delta \varepsilon}(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2)= f_{\Delta \varepsilon}(\varepsilon^{\Delta}_2,\varepsilon^{\Delta}_1)$ for all $(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2)$.
propositionUnder Assumption (ref): $U_{n} \to^{p} \theta$ as $n\to \infty$, where: (i) $\theta\geq 0$ if $T(\cdot)$ is globally concave, (ii) $\theta = 0$ if $T(\cdot)$ is globally linear, (iii) $\theta\leq 0$ if $T(\cdot)$ is globally convex.
proofLet $f_{\xi\xi}(\cdot,\cdot)$ denote the joint distribution of $(X_{ji}'\beta_0, X_{kj}'\beta_0)$. By standard arguments: \begin{gather*} U_{n} \to^{p} \quad \int_{-\infty}^{+\infty} E[\mathbbm{1}\{Y_{i}<Y_{j}<Y_{k}\}sgn(Y_{k}-2Y_{j}+Y_{i})|X_{kj}'\beta_0=X_{ji}'\beta_0=\xi] f_{\xi\xi}(\xi,\xi)d\xi \end{gather*} and we can write: \begin{align*} E[\mathbbm{1} & \{Y_{i}<Y_{j}<Y_{k}\} sgn(Y_{k}-2Y_{j}+Y_{i})|X_{kj}'\beta_0 = X_{ji}'\beta_0=\xi] = \\ &= E[sgn(Y_{k}-2Y_{j}+Y_{i})|Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0=X_{ji}'\beta_0 =\xi]P(Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0=X_{ji}'\beta_0=\xi) \\ & \equiv \tilde{\theta} P(Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0=X_{ji}'\beta_0=\xi) \end{align*} Further: \begin{align*} \tilde{\theta} &= P(Y_{k}-2Y_{j}+Y_{i}>0|Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0 =X_{ji}'\beta_0=\xi) \\ & \qquad \qquad \qquad \qquad - P(Y_{k}-2Y_{j}+Y_{i}<0|Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0=X_{ji}'\beta_0=\xi) \\ & \equiv (1) - (2) \end{align*} and the sign of the probability limit of $U_{n}$ is determined by the sign of $\tilde{\theta}$ (for all $\xi$). For simplicity let $\Xi$ denote the conditioning event in the probabilities above and note that this event is equivalent to $\{\varepsilon_{ji}>-\xi, \varepsilon_{kj}>-\xi, X_{kj}'\beta_0=X_{ji}'\beta_0=\xi\}$. (i) We will show that $(1)\geq (2)$ under concavity of $T$ (i.e. convexity of $T^{-1}$). Denote $a=X_{j}'\beta_0+\varepsilon_{j}$, $b=X_{i}'\beta_0+\varepsilon_{i}$ and $\Delta = \xi + \varepsilon_{kj}$. Note that the conditioning event $Y_{i}<Y_{j}<Y_{k}$ implies that $\Delta>0, a-b>0$ by monotonicity of $T$. We can rewrite the event $Y_{k}-2Y_{j}+Y_{i}>0$ as: \begin{gather*} T^{-1}(a+\Delta)-T^{-1}(a)>T^{-1}(a)-T^{-1}(b) \end{gather*} Conditional on $\Xi(\xi)$ this event is implied by $\varepsilon_{kj}>\varepsilon_{ji}$. To see that observe that the latter event is implied by $\Delta>a-b$, which under convexity of $T^{-1}$ (i.e. concavity of $T$) gives the desired result. Thus, we have: \begin{gather*} P(Y_{k}-2Y_{j}+Y_{i}>0|\Xi) \geq P(\varepsilon_{kj}>\varepsilon_{ji}|\Xi). \end{gather*} On the other hand, $Y_{k}-2Y_{j}+Y_{i}<0 \iff T^{-1}(a+\Delta)-T^{-1}(a)<T^{-1}(a)-T^{-1}(b)$, which under concavity implies $\varepsilon_{kj}< \varepsilon_{ji}$ as we need $\Delta<a-b$ for this event to occur and by definition $\Delta = a-b +\varepsilon_{kj}-\varepsilon_{ji}$. Hence: \begin{gather*} P(Y_{k}-2Y_{j}+Y_{i}<0|\Xi) \leq P(\varepsilon_{kj}<\varepsilon_{ji}|\Xi). \end{gather*} which implies $\tilde{\theta} \geq 2 P(\varepsilon_{kj}>\varepsilon_{ji}|\Xi) - 1$ and in order to show that $\tilde{\theta} \geq 0$ we need $P(\varepsilon_{kj}<\varepsilon_{ji}|\Xi)=0.5$ but that follows from symmetry of $f_{\Delta \varepsilon}(\cdot,\cdot)$ along the 45$^{\circ}$ line. (ii) It is enough to note that if $T$ is linear $P(Y_{k}-2Y_{j}+Y_{i}<0|\Xi) = P(\varepsilon_{kj}<\varepsilon_{ji}|\Xi)$, which implies $\tilde{\theta}=0$. (iii) This part follows from an argument mirroring the one in (i).
remarkBy direct calculation $f_{\Delta \varepsilon}(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2) = \int f(\varepsilon - \varepsilon^{\Delta}_1)f(\varepsilon)f(\varepsilon + \varepsilon^{\Delta}_2) d\varepsilon$. Therefore, $f_{\Delta \varepsilon}(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2)$ is symmetric along the 45$^{\circ}$ line if the probability density function of $\varepsilon$, $f$, is symmetric around zero. Note that symmetry of the error term distribution is assumed in abrevaya_jiang05.
remarkIf we defined our test statistic using two independent pairs of observations, namely: \begin{gather*} \tilde{S}_{n}=\frac{\sqrt{n} }{n(n-1)(n-2)(n-3)}\sum_{i\neq j \neq k \neq l} \mathbbm{1}\{Y_{i}<Y_{j}<Y_{k}<Y_{m} \} sgn(Y_{m}-Y_{k}-Y_{j}+Y_{i})K_{h}\left((X_{mk}-X_{ji})'\hat{\beta}\right) \end{gather*} the requirement for symmetry of $f_{\Delta \varepsilon}(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2)$ could potentially be dropped. This would, however, come at the increased computational cost as we have a 4-th order U statistic now. Also as we require four observations on $Y$ with approximately equally spaced index values $X'\beta_0$ instead of triples in our original statistic, the test based on $\tilde{S}_{n}$ is likely to have lower finite sample power than our baseline test.

In order to obtain critical value for our tests we will assume that the model under the null hypothesis is linear. This is the worst-case $H_{0}$ for testing concavity/convexity as any small local deviation from linearity violates the hypothesis. This can also be seen from the proof of Proposition (ref). In order to derive the asymptotic distribution of our statistic we make the following assumptions.

assumption\begin{enumerate}[(a)] • The kernel $K(\cdot)$ is a bounded, nonnegative, symmetric, twice continuously differentiable function with support on $[-1,1]$ and uniformly bounded derivatives satisfying: \begin{enumerate}[(i)] • $\int K(s) ds = 1$, • $\int s^{2} K(s) ds < \infty$. • Let $\mathcal{K}$ be the antiderivative of $K$. We have: \begin{gather*} \int \int K(s_1)K(s_2) [\mathcal{K} (2s_1-s_2) + 2 \mathcal{K} (2s_2-s_1)] s_1 d s_1 d s_2 > 0 \end{gather*} \end{enumerate} • $h (\log n)^5 \to 0$ and $n h^3 \to \infty$ as $n\to \infty$. • Conditional on the remaining regressors, the distribution of the first element of $X$ is absolutely continuous with respect to the Lebesgue measure, with bounded and twice continuously differentiable density and uniformly bounded second derivatives. Each element of $X$ has a finite fourth moment. • The density of $\varepsilon$ is bounded and twice continuously differentiable, the derivatives are uniformly bounded. • The estimator of $\beta_0$, $\hat{\beta}$, satisfies: \begin{gather*} \hat{\beta}-\beta_0 = \frac{1}{n} \sum_{i=1}^{n}\Omega(X_{i}, Y_i) + o_{p}(n^{-1/2}). \end{gather*} where $E[\Omega(X_{i}, Y_i)]=0$ and each element of $\Omega$ has a finite fourth moment and is continuous in the second argument. \end{enumerate}

Assumption (ref)(ref)(iii) is nonstandard. It is not used for our global test, $S_n$, in this section but is needed for our local test in the next section to have power. Technically, the condition implies that the mean of our local statistic is negative (see proof of Theorem (ref)). This assumption is satisfied by frequently used kernels like standard normal or Epanechnikov (with integral taking values 0.0072 and 0.0177, respectively). The bandwidth rate condition in Assumption (ref)(ref) is rather weak and allows standard “rule-of-thumb” bandwidth choice $h \sim n^{-1/5}$. Assumption (ref)(ref) requires $\hat{\beta}$ to be consistent under both the null and alternative hypothesis and is satisfied by a wide range of estimators including Han's MRC or Ichimura's semiparametric least squares (see e.g. appendix in szydlowski20 for discussion).

Let $\xi_{ji}$ denote the index $X_{ji}'\beta_0$ and $f_{\xi|\xi}$ denote the density of $\xi_{kj}$ given $\xi_{ji}$. Define:

align*[align* omitted — 672 chars of source]

and let $\mu_{2}$ denote the derivative of $\mu$ with respect to the second argument.

theoremIf Assumption (ref) holds, we have: \begin{gather*} \sqrt{n}(U_{n}-\theta) \to^{d} N(0,E[\psi_{i}^{2}]) \end{gather*} where $\psi_{i}= E[\delta(Y_{i},X_{i},\xi_{ji},\xi_{ji})|Y_{i},X_{i}]- E[\mu_{2}(\xi_{ji},\xi_{ji})]\Omega(X_{i}, Y_i)$.

As linearity is the boundary case for testing $H_{0}$: $T(\cdot)$ is concave, or $H_{0}$: $T(\cdot)$ is convex, we will reject concavity if $S_{n}<c_{\alpha}$ and reject convexity if $S_{n}>c_{1-\alpha}$, where $c_{\alpha}$ denotes the $\alpha$ quantile from the normal asymptotic distribution. As typical with U-statistics, estimating the variance of the asymptotic distribution of $S_{n}$ is difficult. On the one hand, the plug-in estimator will involve estimating derivatives of conditional moments and distributions, which requires delicate choices of bandwidths and, essentially, calculation of higher order U-statistics. On the other hand, using the standardisation approach in ghosal_et_al00 will lead to a 5-th order U-statistic, which would be difficult to calculate with sample sizes typically encountered in applications.\footnote{Proceeding as there would involve calculating: \scriptsize

align*[align* omitted — 422 chars of source]

where $\tilde{\mu}_{2}$ is an estimator of $-E[\mu_{2}(\xi_{ji},\xi_{ji})]$ (which itself is a 3-rd order U statistic). } Thus, in Section (ref) we resort to bootstrap for calculating the critical value as bootstrapping involves only repeated calculation of a 3-rd order U-statistic, $U_{n}$, which we find computationally easier than the aforementioned methods.

The global test introduced in this section only has power against global deviations from concavity/linearity/convexity and does not have power if the function is both convex and concave on different parts of the domain. Thus, it can be used as a first check -- for example, if the test rejects concavity, the researcher concludes that the function cannot be globally concave. Failure to reject would, then, mean that one has to consider our local test described in the next section, in order to verify if indeed the function is globally concave.

Local test

The main idea of the local test is to consider only triples of the kind portrayed in Figure (ref) local to a point $y$ in the domain of the transformation function. In other words, we will check if the transformation function is concave/linear/convex locally around $y$. Local concavity may be of interest by itself. However, usually we are interested in verifying if the treatment effect is concave on the whole domain, thus we will take the minimum of the local statistics at different points $y$ to run the test.

Formally, define:

align*[align* omitted — 254 chars of source]

where in practice the bandwidth used for $(Y_{i}-y)$ can be different than for $(X_{kj} - X_{ji})'\hat{\beta}$ but, in order to simplify exposition and mathematical arguments, we assume that both bandwidths are of the same order and denote both by $h$. Now our local test statistic for testing concavity is defined as:

gather*[gather* omitted — 73 chars of source]

where $\mathcal{Y}$ is a compact set, and the statistic for testing convexity is defined with $\sup$ replacing $\inf$ above. Linearity can be tested by replacing $U_{n}(y)$ with its absolute value. For the rest of the article we concentrate on $S_{n}^{conc}$ as results for testing convexity and linearity follow by very similar arguments.

Intuitively, low values of $S_{n}^{conc}$ show that there is a large deviation from concavity at some point $y$ and, hence, the treatment effects are not accelerating on the whole domain of the outcome.\footnote{Note that by the formulas in Example 1 concavity of $T$ is equivalent to convexity of the outcome in the treatment.} Therefore, the null hypothesis of concavity would be rejected if $S_{n}^{conc} < c_{\alpha}^{conc}$ where $c_{\alpha}^{conc}$ is an appropriate quantile from the asymptotic approximation to the distribution of our statistic.

In order to obtain $c_{\alpha}^{conc}$ one could imagine proceeding as in ghosal_et_al00: approximate the standardised U-statistic process $\sqrt{n} U_{n}(y)/\sigma_{n}(y)$, where $\sigma_{n}(y)$ is the estimator of the asymptotic variance, by a Gaussian process and then apply the extreme value theory in order to derive the distribution of the infimum. However, pursuing that approach is difficult in our setup for two reasons: 1) under $H_0$ we still have $E[U_n(y)]=O(h)$ rather than $o(h)$ therein, 2) as discussed above, estimating $\sigma_{n}(y)$ is computationally expensive. Instead we propose to use bootstrap to approximate $c_{\alpha}^{conc}$.

Define:

align*[align* omitted — 519 chars of source]

The following asymptotic approximation will be useful in justifying our bootstrap procedure.

theoremIf Assumption (ref) holds, then: \begin{gather*} \sup_{y \in \mathcal{Y}}\left|U_{n}(y) - E[U_n(y)] - \frac{1}{n} \sum_{i=1}^{n} \phi_{i,n}(y)\right| = o_{p}((nh)^{-1/2}) \end{gather*}

An interesting consequence of this result is that the asymptotic distribution of the statistic does not depend on the estimation of $\beta_0$ as long as Assumption (ref)(ref) is satisfied, unlike the global test (cf. Theorem (ref)).

Bootstrap critical values

Our bootstrap procedure for obtaining the critical value for the global or local test is as follows:

enumerate• Estimate $\beta_0$ (e.g. by Han's MRC) and calculate the residuals $\hat{\varepsilon}_{i} = Y_{i} - X_{i}'\hat{\beta}$. • For the global test: draw a random sample $\{v_i\}_{i=1}^{n}$ from a two point distribution with $P(v_{i}=-1)=P(v_{i}=1)=1/2$ and define $\varepsilon_{i}^{*} =v_{i} \hat{\varepsilon}_{i}$. For the local test: (a) draw $\{\varepsilon_i^*\}_{i=1}^n$ with replacement from $\{\hat{\varepsilon}_i\}_{i=1}^n$ or (b) follow the same procedure as for the global test. Then, generate $Y_{i}^*$ by: \begin{gather*} Y_{i}^{*} = X_{i}'\hat{\beta} + \varepsilon_i^*. \end{gather*} • Estimate $\beta_0$ (e.g. by Han's MRC) using the bootstrap sample. Let the resulting estimate be denoted by $\beta^*$. • Calculate the statistic $S_{n}$ or $S_{n}^{conc}$ on the bootstrap sample using $\beta^*$ instead of $\hat{\beta}$. Denote the resulting bootstrap statistics by $S_{n}^{*}$ and $S_{n}^{conc,*}$. • Obtain the empirical distribution of $S_{n}^*$ and $S_{n}^{conc,*}$ by repeating steps 1-4 many times. Calculate the $\alpha$ quantiles of the empirical distribution of $S_{n}^*$ and $S_{n}^{conc,*}$ and denote them by $c_{\alpha}^{*}$ and $c_{\alpha}^{conc,*}$, respectively.

Step two imposes the null hypothesis of linearity on the bootstrap sample so it “recentres” the bootstrap statistic on the linear case. Furthermore, sampling from a symmetric two point distribution in the wild bootstrap imposes symmetry of the error distribution in the bootstrap sample, thus implying that Assumption (ref) is satisfied in that sample. Note that the jacknife multiplier bootstrap (JMB) of chen_kato20 will not work for the local test here (without modifications) as we have $E[U_n(y)]=O(h)$ even under the null so the JMB statistic will not be centred correctly. We assume an equivalent of Assumption (ref)(ref) for the bootstrap estimator $\beta^{*}$:

assumption• The estimator of $\beta^*$ satisfies: $\beta^* - \hat{\beta} = \frac{1}{n} \sum_{i=1}^{n}\Omega(X_{i}, Y_i^*) + o_{p}(n^{-1/2})$, where $E[\Omega(X_{i}, Y_i^*)]=0$ has zero mean and each element of $\Omega$ has a finite fourth moment and is continuous in the second argument.

This assumption is satisfied in our leading case, when $\beta^*$ is estimated by MRC (see subbotin07).

theoremIf Assumptions (ref)-(ref) hold and $T(\cdot)$ is linear: \begin{gather*} \lim_{n\to \infty} P(S_{n}^{conc} \leq c_{\alpha}^{conc,*}) = \alpha \end{gather*} If, additionally, Assumption (ref) holds, we have: \begin{gather*} \lim_{n\to \infty} P\left(\sqrt{n}U_n \leq c_{\alpha}^{*}\right) = \alpha . \end{gather*} (and equivalent result holds for a test of linearity or convexity).

Theorem (ref) implies, for example, that we can reject global concavity of $T$, i.e. increasing treatment effect, when $S_{n}^{conc}<c_{\alpha}^{conc,*}$, and reject global linearity when $|S_{n}^{conc}|>c_{1-\alpha/2}^{conc,*}$. We prove validity of bootstrap in unconditional probability, which is a weaker result than the standard result conditional on the sample (see e.g. cavaliere_georgiev20 for discussion), as the fact that $E[U_n(y)] = O(h)$ under $H_0$ and the need to use parametric bootstrap complicates the analysis. Note that we only require the symmetry of the error distribution for the global test. The next theorem characterises alternatives for which our tests are consistent.

Assume that $c^{conc,*}_{\alpha}$ is generated using method (a) in step 2 above (see Appendix (ref) for the analysis with the wild bootstrap in (b)). Let $f_{xb}$ denote the pdf of $X_i'\beta_0$ and $\mathbf{1} = [1 \quad 1 \quad 1]'$. Further, let $f_{\boldsymbol{\varepsilon}}$ denote the joint distribution of $(\varepsilon_i,\varepsilon_j,\varepsilon_k)$ and $\nabla f_{\boldsymbol{\varepsilon}}$ denote its gradient and define $f_{\boldsymbol{\tilde{\varepsilon}}}$, $\nabla f_{\boldsymbol{\tilde{\varepsilon}}}$, similarly for $\tilde{\varepsilon} = Y -X'\beta_0$.

theoremIf Assumptions (ref)-(ref) and Assumption (ref) hold, and $T(\cdot)$ is globally convex: \begin{gather*} \lim_{n\to \infty} P\left(\sqrt{n}U_n \leq c_{\alpha}^{*}\right) = 1. \end{gather*} Moreover, if Assumptions (ref)-(ref) hold, $T(\cdot)$ is locally convex at some $y$ and the following condition holds: \scriptsize \begin{align} &\sup_{y \in \mathcal{Y}} \int \int \nabla f_{\boldsymbol{\tilde{\varepsilon}}} \begin{pmatrix} y -\zeta +\xi \\ y -\zeta \\ y -\zeta -\xi \end{pmatrix}' \mathbf{1} f_{xb}(\zeta-\xi)f_{xb}(\zeta)f_{xb}(\zeta+\xi) d \zeta d\xi \nonumber \\ & \quad - \sup_{y \in \mathcal{Y}} \int \int \left\{ T'(y)^4 \nabla f_{\boldsymbol{\varepsilon}} \begin{pmatrix} T(y) -\zeta +\xi \\ T(y) -\zeta \\ T(y) -\zeta -\xi \end{pmatrix} ' \mathbf{1} + 3T”(y)T'(y)^2 f_{\boldsymbol{\varepsilon}} \begin{pmatrix} T(y) -\zeta +\xi \\ T(y) -\zeta \\ T(y) -\zeta -\xi \end{pmatrix} \right\} f_{xb}(\zeta-\xi)f_{xb}(\zeta)f_{xb}(\zeta+\xi) d \zeta d\xi < 0, \end{align} then we have: \begin{gather*} \lim_{n\to \infty} P(S_{n}^{conc} \leq c_{\alpha}^{conc,*}) = 1 \end{gather*} (and equivalent result holds for a test of linearity or convexity).

The condition in (ref) is technical and it is not straightforward to find general sufficient conditions for it to hold as it depends not only on the shape of $T$ but also the distributions of $\epsilon$ and $X$. However, we can make the following observations.

remarkFor testing linearity (vel testing treatment effect heterogeneity), the condition would only require the expression in (ref) to be nonzero and, thus, should be satisfied generically for non-linear $T(y)$'s, and the test should be consistent against (almost) all nonlinear alternatives.
remarkTheorem (ref) shows the cost of using a simple bootstrap test that avoids estimation of the transformation function, in particular the fact that our bootstrap does not correctly approximate the error distribution $f$ but rather distribution of OLS residuals, $\tilde{\varepsilon}$. The latter distribution depends on the distribution of $X$ and on $T$, complicating the power properties of the test. In principle, one could try to estimate $T'$ and $\nabla f_{\boldsymbol{\varepsilon}}$ and approximate the term: \scriptsize \begin{gather*} \sup_{y \in \mathcal{Y}} \int \int T'(y)^4 \nabla f_{\boldsymbol{\varepsilon}} \begin{pmatrix} T(y) -\zeta +\xi \\ T(y) -\zeta \\ T(y) -\zeta -\xi \end{pmatrix} ' \mathbf{1} f_{xb}(\zeta-\xi)f_{xb}(\zeta)f_{xb}(\zeta+\xi) d \zeta d\xi \end{gather*} by a bootstrap procedure, making sure that condition (ref) would be satisfied. However, this approach would substantially complicate the calculation of the critical value as estimators of $T$ and $f$ in the transformation model often involve non-smooth criterion functions and do not readily offer estimates of the derivatives (e.g. chen02). Thus, for practical reasons we prefer our current approach, even though it carries a cost in terms of power of the tests for concavity and convexity.
remarkAt an additional computational cost one could estimate $T$ and use semiparametric residuals: $\hat{\varepsilon}_i = \hat{T}(Y_i) - X_i'\hat{\beta}$ in our bootstrap. We look at the latter case in Appendix (ref) where we show that the finite sample performance of such test does not dominate our local test.

Finally, let us note that the computation of the local statistic involves evaluating a third order U-statistic at different points $y$ and, hence, takes significantly longer to compute than the global statistic. In practice, we recommend running the global test first and then proceed with the local test if the null hypothesis cannot be rejected in the first step.

Monte Carlo simulations

The data is generated from the following five models:

align*[align* omitted — 280 chars of source]

where we draw $X$ and $\varepsilon$ from the standard normal distribution. Note that D0 is the worst-case model in the null hypothesis, D1 imposes concavity of the transformation, D3 convexity and D2 and D4 are neither concave or convex.

figure[figure omitted — 129 chars of source]

We run 1000 Monte Carlo replications. We use Gaussian kernel functions and rule-of-thumb bandwidths for both $(X_{kj}-X_{ji})'\hat{\beta}$ and $Y_{i}$, namely $h=1.06\hat{\sigma}n^{-1/5}$, where $\hat{\sigma}$ is a sample standard deviation of $(X_{kj}-X_{ji})'\hat{\beta}$ or $Y_{i}$.\footnote{The choice of bandwidth for $Y_i$ seems trickier as some transformations $T^{-1}$ make the distribution of $Y_i$ severly skewed. We have also tried using cross-validation (ucv in R) and plug-in cross-validation (bcv in R) with similar results.} In order to calculate the local statistic $S_{n}^{conc}$ we use a grid of values for $y$: -2:0.25:2, and take a minimum over the grid. The number of bootstrap replications used to calculate the critical value is 500 and we consider three sample sizes: $n=100, 250$ and $500$.

table[table omitted — 960 chars of source]

Table (ref) contains the results of the Monte Carlo simulations. We concentrate on testing concavity, as the results for the linearity and convexity tests, are very similar. The rejection probabilities in the linear case (D0) are close to the nominal level for both the global and the local test and both tests have perfect power against a globally convex alternative in D3.

As predicted, the global test has low power against D2 and D4, for which the transformation function is concave on part of the domain and convex on the other part. In this case deviations from concavity, as measured by our global statistic, cancel with positive values of the statistic obtained for the part of the domain where the function is concave, resulting in the global statistic taking values close to zero, just as for the linear case.

For designs D2 and D4, our local test significantly improves over the global test with almost perfect detection of non-concavity in D2 for $n = 250$. The power against D4 increases slower with $n$ than D2 but we still reject H0 with close to certainty with $n=500$. This is in line with the intuition that the local test will concentrate on the region of the largest violation of concavity instead of averaging over the measures of concavity for different regions.

Finally, note that, unlike for the global test, sampling from a symmetric distribution is not required in the local test bootstrap, so it is a question of finite sample performance if we would prefer the regular parametric bootstrap (part (a) in Step 2) or the symmetric wild parametric bootstrap in part (b). Comparing the last two panels in Table (ref) shows that both bootstrap tests have the right level but the symmetric wild bootstrap test has lower power for our designs (see D2 and D4).

Application: Curvature of loan demand

We use data from karlan_zinman08, who ran randomised trials with a for-profit consumer lender in South Africa targeting high-risk consumer loan market. The lender randomised individual interest rate direct mail offers , conditional on the client’s risk category. karlan_zinman08 found that the loan demand curves are downward sloping. We will investigate if the demand curves are convex or, in other words, if increases in interest rate have a stronger negative effect on loan demand the lower the rate.

figure[figure omitted — 129 chars of source]

The dependent variable is the amount borrowed (in rands) at the offered interest rate.\footnote{We abstract from selectivity issues here. See karlan_zinman08 for discussion.} The experiment was ran in three mailer waves over four months, thus as in karlan_zinman08 we include the risk category and wave dummies as controls in $X$. The data contains 2325 observations. As the first step, we estimated the transformation function in our model using the estimator in chen02 and plotted it in Figure (ref) together with a smoothed version.\footnote{This is for illustrative purposes only. Note that the estimator in chen02 does not impose monotonicity, so a full exercise of estimating $T$ should include additional step of monotonising the estimate. We want to stress that our testing procedure does not rely on any estimator of $T$.} The figure suggests that the transformation is not far from being concave on most of its domain, maybe besides the values of loan size below 1000 rands, which implies convex demand curve by formulas in Example 1.

table[table omitted — 964 chars of source]

In order to formally test our conjectures, we first apply our global test to verify if the demand function is concave, linear or convex. As Table (ref) shows, global test rejects linearity and concavity of demand. Thus, in the second step we apply the local test to the null hypothesis of convexity and, in fact, reject the null hypothesis, concluding that the loan demand is not globally convex in the interest rate. To shed more light on where the non-convexity may come from, we re-ran the test on the sample excluding loans lower than 1000 rands and found that convexity is not rejected on this sub-sample. Therefore, overall we confirm that loan demand is mostly convex in interest rate, besides very small loans sector of the market.

Conclusion

Our application demonstrates usefulness of our testing procedures in recovering monotonicity of treatment effects. Particular appeal of the procedures described in this article comes from the fact that they avoid estimation of the transformation function and only require OLS estimation of the vector of coefficients $\beta_0$. Thus, they are relatively easy to implement. Additionally, computation of the third order U-statistic involved in our tests can be done efficiently by sorting the data by $Y$ first -- this reduces computational complexity to $O(nlog(n)+n(n-2)/2)$ from $O(n^{3})$ for a straightforward triple loop through the observations.

Appendix