EconBase
← Back to paper

Testing Shape Restrictions with Continuous Treatment: A Transformation Model Approach

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

46,478 characters

Testing Shape Restrictions with Continuous Treatment: A Transformation Model Approach


\maketitle

\begin{abstract}
We propose tests for the convexity/linearity/concavity of a transformation of the dependent variable in a semiparametric transformation model. These tests can be used to verify monotonicity of the treatment effect, or, equivalently, concavity/convexity of the outcome with respect to the treatment, in (quasi-)experimental settings. Our procedure does not require estimation of the transformation or the distribution of the error terms. The statistic takes the form of a U statistic or a localised U statistic, and we show that critical values can be obtained by bootstrapping. In our application we test the convexity of loan demand with respect to the interest rate using experimental data from South Africa. \\
\vspace{2cm}\\
JEL: C12, C21, C14
\newline
Keywords: Shape restrictions, Transformation model, Bootstrap, U statistic, Treatment effects
\end{abstract}

\section{Introduction} \label{intro}

In this paper we consider testing if a treatment has a diminishing or increasing effect on the outcome. This is often of interest on top of the question if there is an effect at all or what sign it has. For example, often it is natural to expect that demand will be decreasing in price. However, it is less clear if increasing the price will have a larger effect at lower or higher price levels. Our test can address this question without fully estimating the demand relationship.

Let $X$ denote the vector of treatment and control variables. As estimating the effect of $X$ on the outcome $Y$ nonparametrically  with non-trivial number of treatments and controls would suffer from a curse of dimensionality, we impose a single-index structure and assume that the nonlinearity of the treatment effect comes from a nonlinear transformation of $Y$. In other words, consider a transformation model of the form:
\begin{gather}
T(Y)=X'\beta_0+\varepsilon \label{1}
\end{gather}
where $Y$ is a scalar dependent variable, $X$ is a vector of $q$ nondegenerate explanatory variables, $\beta_0$ is a vector of coefficients belonging to a compact set $\Theta_{\beta} \subset \mathbb{R}^q$, $T(\cdot)$ is an increasing function and $\varepsilon$ is an unobserved error term with distribution $F$ that is independent of $X$. In order to identify the model we need a location normalisation: e.g. $T(0)=0$, $E(\varepsilon)=0$ or $Me(\varepsilon)=0$; and a scale normalisation: e.g. $|\beta_{0,1}|=1$. The benefit of using the transformation model compared to a standard single-index model (i.e. $Y=T(X'\beta_0) + \varepsilon$), besides the fact that it facilitates our testing approach, is that the transformation model allows the treatment effect of $X_{k}$ to depend on the values of both observed and unobserved characteristics (note that $Y=T^{-1}(X'\beta_0+\varepsilon)$) so can be seen as a simple way of introducing heterogenous treatment effects that vary with unobservables.

The main objective of this article is to develop a practically appealing test to determine whether the transformation function $T(\cdot)$ is concave/linear/convex. The examples below illustrate the importance of testing curvature of the transformation.
\begin{exmp}\emph{\textbf{(Experiments with continuous treatment)}}
Using normalization $E(\varepsilon) = 0$ and $Q_{\alpha}(\varepsilon) = 0$ where $Q_{\alpha}$ denotes the $\alpha$ quantile, we have, respectively:
\begin{align*}
\frac{\partial^{2} E(Y|X)}{\partial X_{k}^{2}} &= - \beta_{0,k}^{2} E\left[ \frac{T''(T^{-1}(X'\beta_0+\varepsilon))}{T'(T^{-1}(X'\beta_0+\varepsilon))^{3}}\bigg| X\right] \qquad \text{(mean regression)} \\
\intertext{and}
\frac{\partial^{2} Q_{\alpha}(Y|X)}{\partial X_{k}^{2}} &= -\beta_{0,k}^{2} \frac{T''(T^{-1}(X'\beta_0))} {T'(T^{-1}(X'\beta_0))^{3}}
\qquad \text{(quantile regression)}
\end{align*}
As $T'(\cdot)>0$ the curvature of the mean/quantile effect depends on $T''(\cdot)$. Thus, the sign of the second derivative $T''(\cdot)$ determines if the effect of the treatment $X$ on the $\alpha$ quantile of $Y$ is concave or convex in $X$. For example, if a company randomises marketing spending in different markets, testing for concavity would answer the question if marketing spending has diminishing returns (e.g. on mean revenue).  Test of curvature has also application in experimental studies of demand elasticities (e.g. \cite{jessoe_rapson14}, \cite{hainmueller_et_al15}, \cite{karlan_zinman19}) where it can be used to verify if demand is concave, linear or convex.

Finally, when the transformation is linear, the treatment effect does not depend on the observed ($X$) and unobserved ($\varepsilon$) heterogeneity. Thus, the test of linearity of $T$ can be seen as a test of treatment effect heterogeneity.
\end{exmp}

\begin{exmp} \emph{\textbf{(Duration models: testing hazard monotonicity)}}
Let $\lambda(\cdot)$ and $\Lambda(\cdot)$ denote baseline hazard and integrated baseline hazard, respectively. In a duration model: $T(Y) = \log \Lambda (Y)$, and
\begin{gather*}
T''(Y) =  \frac{\lambda'(Y)}{\Lambda(Y)}-\left(\frac{\lambda(Y)}{\Lambda(Y)}\right)^{2}.
\end{gather*}
Hence, rejecting concavity of $T(\cdot)$ (i.e. $T''(\cdot)<0$) implies that the baseline hazard is non-decreasing ($\lambda'(\cdot)\geq 0$). One can, thus, use the test of concavity of the transformation as a test for monotonicity of the baseline hazard, or in other words, as a test for positive duration dependence. In the economic context, one may be interested in detecting non-monotonicity of unemployment exit rate due to unemployment benefit exhaustion effects (see \cite{card_et_al07b} for discussion).
\end{exmp}

Beyond these examples our procedure can also be used for specification search, i.e. determining if one should use a concave or convex transformation, and to test the curvature in wage regressions, e.g. if the effect of education or experience is concave, or the curvature of the marginal utility (or profit) function in hedonic models (see \cite{ekeland_et_al04}).

Our test statistic simply compares triples of $Y$'s corresponding to equally spaced index values $X'\beta_0$, thus it does not require estimation of the transformation function $T$ or the distribution of $\varepsilon$. We only require an estimator of $\beta_0$ and, basically, the symmetry of the distribution of the error terms. Estimating $\beta_0$ can be done, for example, by using the maximum rank correlation estimator, \cite{han87}, or semiparametric least squares, \cite{ichimura93}.

We propose both a global test that has power to detect globally convex or concave functions and leads to asymptotic normal critical values, as well as a more general test that detects local deviations from linearity, i.e. has power against alternatives that are both convex and concave on the domain of $T$. Our tests do not have a pivotal asymptotic distribution but we show that the critical values can be obtained by parametric bootstrap.

The cost of our test’s easy implementation is that the power properties of our local test are complex, and the test may not be consistent against some relevant alternatives. Nevertheless, our application demonstrates the usefulness of our testing approach in detecting nonconvexities in the demand for loans as a function of the interest rate.

Our statistic resembles the approach in \cite{abrevaya_jiang05} who test curvature in a nonparametric regression model. However, unlike their approach our test does not suffer from the curse of dimensionality due to the single index structure of the regression part and allows the marginal effect of covariates to vary with the error term (though in a manner restricted by the single-index). Also the details of the derivation of the asymptotic distribution are different due to presence of estimated $\beta_0$ in our statistic and somewhat distinct approach to obtaining power against general alternatives. These traits are shared by \cite{abrevaya_et_al10}, who test monotonicity in a generalised regression model with endogeneity, though a distinguishing feature of our work is that we formally show validity of bootstrap for obtaining   critical values.

Related literature includes tests for the sign of the treatment effect, see e.g. \cite{kline16}, and testing for treatment effect heterogeneity (\cite{abadie02}, \cite{crump_et_al08}, \cite{santanna21}, \cite{chernozhukov_et_al23}). Unlike the latter papers, our approach allows the treatment effect to vary with (independent) unobserved heterogeneity at the cost of imposing much more structure on the treatment effect model and the covariates. Testing curvature of the transformation can also be seen as a generalisation of specification testing in \cite{neumeyer_et_al16} and \cite{szydlowski20}. The idea of using curvature of the integrated hazard to test monotonicity of the baseline hazard has been utilised by \cite{hall_van_keilegom05}. Testing shape restrictions in a nonparametric regression model has been considered by \cite{ghosal_et_al00}, \cite{gutknecht16}, \cite{chetvertikov19} and \cite{komarova_hidalgo23}, among others. Similarly to this paper, \cite{chen_kato20} propose a bootstrap procedure for approximating the supremum of a local U-statistic.

The article is organized as follows. Section \ref{mainid} discusses the idea behind the testing procedure informally and the formal results are postponed till Sections \ref{formal}-\ref{sec:boot}. Sections \ref{MC}-\ref{appl} contain Monte Carlo results and our application to loan demand.
Proofs, besides the proof of the main proposition, are located in the Appendix, which also contains some additional MC simulations.


\section{Main idea} \label{mainid}

Figure \ref{fig:intui} portrays the intuition behind our test. The transformation plotted in the figure is concave and we display three ordered data points in the figure, $(Y_{i},Y_{j},Y_{k})$, for which the transformation function is equally spaced, i.e. $T(Y_{k})-T(Y_{j})=T(Y_{j})-T(Y_{i})$.
\begin{figure}[ht]
\caption{Testing concavity}
\begin{center}
\begin{tikzpicture}[scale=1.75]
\draw[->] (-2,-1)--(1.5,-1) node[right]{$Y$};
\draw[->] (-1,-2)--(-1,1.5) node[above]{$T(Y)$};
\draw[scale=1, domain=-1:1.5, smooth, variable=\x, blue, thick] plot ({\x}, {ln(\x+1.12)});
\node[black!60!green] at (-1,-0.5) [left] {$T(Y_{i})=X_{i}'\beta+\varepsilon_{i} $};
\draw[dashed, black!60!green] (-1.05,-0.5) -- (-0.5134693,-0.5)--(-0.5134693,-1.05);
\node[black!60!green] at (-0.5134693,-1) [below] {$Y_{i}$};
\node[black!60!green] at (-1,0) [left] {$T(Y_{j})=X_{j}'\beta+\varepsilon_{j} $};
\draw[dashed, black!60!green] (-1.05,0) -- (-0.12,0)--(-0.12,-1.05);
\node[black!60!green] at (-0.12,-1) [below] {$Y_{j}$};
\node[black!60!green] at (-1,0.5) [left] {$T(Y_{k})=X_{k}'\beta+\varepsilon_{k} $};
\draw[dashed, black!60!green] (-1.05,0.5) -- (0.5287213,0.5)--(0.5287213,-1.05);
\node[black!60!green] at (0.5287213,-1) [below] {$Y_{k}$};
\end{tikzpicture}
\end{center}
\label{fig:intui}
\end{figure}

Concavity of $T(\cdot)$ implies that $Y_{k}-Y_{j}>Y_{j}-Y_{i}$. Note that $T(Y_{k})-T(Y_{j})=T(Y_{j})-T(Y_{i})$ is equivalent to $(X_{k}-X_{j})'\beta_0 +\varepsilon_{k}-\varepsilon_{j}= (X_{j}-X_{i})'\beta_0 +\varepsilon_{j}-\varepsilon_{i}$. Hence, ``on average'' equally spaced $T(Y)$'s mean equally spaced index values $X'\beta_0$. Therefore, we can detect deviations from concavity by considering the following criterion:
\begin{gather*}
-\frac{1}{n(n-1)(n-2)}\sum_{i\neq j \neq k} \mathbbm{1}\{Y_{i}<Y_{j}<Y_{k} \} \mathbbm{1}\{Y_{k}-Y_{j}<Y_{j}-Y_{i} \}\mathbbm{1}\{X_{kj}'\beta_0 = X_{ji}'\beta_0\}
\end{gather*}
where $X_{ji} \equiv X_{j}-X_{i}$.

In order to make this criterion operational with continuous distribution of $X'\beta_0$ (which is required for identification) we need to replace the last indicator function with a smooth kernel $K_{h}(\cdot) = h^{-1}K(\cdot/h)$ and $\beta_0$ with its estimator $\hat{\beta}$:
\begin{gather*}
-\frac{1}{n(n-1)(n-2)}\sum_{i\neq j \neq k} \mathbbm{1}\{Y_{i}<Y_{j}<Y_{k} \} \mathbbm{1}\{Y_{k}-Y_{j}<Y_{j}-Y_{i} \}K_{h}\left(X_{kj}'\hat{\beta} - X_{ji}'\hat{\beta}\right)
\end{gather*}
Deviations from convexity can be detected in a similar fashion.
Finally, we can combine both criterion functions to detect deviations from linearity:
\begin{gather}
U_{n}=\frac{1}{n(n-1)(n-2)}\sum_{i\neq j \neq k} \mathbbm{1}\{Y_{i}<Y_{j}<Y_{k} \} sgn(Y_{k}-2Y_{j}+Y_{i})K_{h}\left(X_{kj}'\hat{\beta} - X_{ji}'\hat{\beta}\right)
\end{gather}
Here, very negative values of the objective function signify convexity and large positive values mark concavity. Outside testing, one may use the above measure itself to characterise an ``average'' curvature of function $T$, e.g. large positive values suggest that the function is predominantly concave.

Note that our test requires only one-dimensional kernel smoothing, thus it does not suffer from the curse of dimensionality. Also, unlike the specification test in \cite{szydlowski20} it does not require estimation of the transformation function, a computationally intense task itself.

\section{Formal definition and asymptotic theory}\label{formal}

\subsection{Global test}

As the probability limit of the criterion functions introduced in the previous section depends on the extent of concavity, but these functions are asymptotically centred at known values under linearity, we form our procedure as a test of:
\begin{align*}
H_0:
\begin{cases}
\text{$T(\cdot)$ is concave} \\
\text{$T(\cdot)$ is linear} \\
\text{$T(\cdot)$ is convex} \\
\end{cases}
 \qquad vs \qquad
H_{A}:
\begin{cases}
\text{$T(\cdot)$ is non-concave} \\
\text{$T(\cdot)$ is non-linear} \\
\text{$T(\cdot)$ is non-convex} \\
\end{cases}
\end{align*}
depending on the question of interest.  We will use the test statistic
\begin{gather*}
S_{n}=\sqrt{n} U_{n}
\end{gather*}
and reject the null hypothesis at level $\alpha$ if $S_{n}<c_{\alpha}, |S_{n}|>c_{1-\alpha/2}$ or $S_{n}>c_{1-\alpha}$ depending on $H_{0}$, respectively, where $c_{\alpha}$ denotes an $\alpha$ quantile from an appropriate asymptotic distribution.

Define $\varepsilon_{ji}=\varepsilon_j - \varepsilon_i$. Proposition \ref{prop:plim} justifies our testing strategy. We need the following symmetry assumption:

\begin{customassmpt}{SYM}\label{A:SYM}
Define $f_{\Delta \varepsilon}(\cdot,\cdot)$ to be the joint distribution of $(\varepsilon_{ji}, \varepsilon_{kj})$. Assume that $\{(X_{i},Y_{i})\}_{i=1}^n$ are i.i.d and that $f_{\Delta \varepsilon}(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2)= f_{\Delta \varepsilon}(\varepsilon^{\Delta}_2,\varepsilon^{\Delta}_1)$ for all $(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2)$.
\end{customassmpt}

\begin{proposition} \label{prop:plim}
Under Assumption \ref{A:SYM}:  $U_{n} \to^{p} \theta$ as $n\to \infty$, where:
(i) $\theta\geq 0$ if $T(\cdot)$ is globally concave,
(ii) $\theta = 0$ if $T(\cdot)$ is globally linear,
(iii) $\theta\leq 0$ if $T(\cdot)$ is globally convex.
\end{proposition}

\begin{proof}
Let $f_{\xi\xi}(\cdot,\cdot)$ denote the joint distribution of $(X_{ji}'\beta_0, X_{kj}'\beta_0)$.  By
standard arguments:
\begin{gather*}
U_{n} \to^{p}  \quad \int_{-\infty}^{+\infty} E[\mathbbm{1}\{Y_{i}<Y_{j}<Y_{k}\}sgn(Y_{k}-2Y_{j}+Y_{i})|X_{kj}'\beta_0=X_{ji}'\beta_0=\xi] f_{\xi\xi}(\xi,\xi)d\xi
\end{gather*}
and we can write:
\begin{align*}
E[\mathbbm{1} & \{Y_{i}<Y_{j}<Y_{k}\}  sgn(Y_{k}-2Y_{j}+Y_{i})|X_{kj}'\beta_0 = X_{ji}'\beta_0=\xi] = \\
&=  E[sgn(Y_{k}-2Y_{j}+Y_{i})|Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0=X_{ji}'\beta_0 =\xi]P(Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0=X_{ji}'\beta_0=\xi) \\
& \equiv \tilde{\theta} P(Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0=X_{ji}'\beta_0=\xi)
\end{align*}
Further:
\begin{align*}
\tilde{\theta} &= P(Y_{k}-2Y_{j}+Y_{i}>0|Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0 =X_{ji}'\beta_0=\xi) \\
& \qquad \qquad \qquad \qquad  - P(Y_{k}-2Y_{j}+Y_{i}<0|Y_{i}<Y_{j}<Y_{k}, X_{kj}'\beta_0=X_{ji}'\beta_0=\xi) \\
& \equiv (1) - (2)
\end{align*}
and the sign of the probability limit of $U_{n}$ is determined by the sign of $\tilde{\theta}$ (for all $\xi$). For simplicity let $\Xi$ denote the conditioning event in the probabilities above and note that this event is equivalent to $\{\varepsilon_{ji}>-\xi, \varepsilon_{kj}>-\xi, X_{kj}'\beta_0=X_{ji}'\beta_0=\xi\}$.

(i) We will show that $(1)\geq (2)$ under concavity of $T$ (i.e. convexity of $T^{-1}$).
Denote $a=X_{j}'\beta_0+\varepsilon_{j}$, $b=X_{i}'\beta_0+\varepsilon_{i}$ and $\Delta = \xi + \varepsilon_{kj}$. Note that the conditioning event $Y_{i}<Y_{j}<Y_{k}$ implies that $\Delta>0, a-b>0$ by monotonicity of $T$. We can rewrite the event $Y_{k}-2Y_{j}+Y_{i}>0$ as:
\begin{gather*}
T^{-1}(a+\Delta)-T^{-1}(a)>T^{-1}(a)-T^{-1}(b)
\end{gather*}
Conditional on $\Xi(\xi)$ this event is implied by $\varepsilon_{kj}>\varepsilon_{ji}$. To see that observe that the latter event is implied by $\Delta>a-b$, which under convexity of $T^{-1}$ (i.e. concavity of $T$)  gives the desired result.
Thus, we have:
\begin{gather*}
P(Y_{k}-2Y_{j}+Y_{i}>0|\Xi) \geq P(\varepsilon_{kj}>\varepsilon_{ji}|\Xi).
\end{gather*}
On the other hand, $Y_{k}-2Y_{j}+Y_{i}<0 \iff T^{-1}(a+\Delta)-T^{-1}(a)<T^{-1}(a)-T^{-1}(b)$, which under concavity implies $\varepsilon_{kj}< \varepsilon_{ji}$ as we need $\Delta<a-b$ for this event to occur and by definition $\Delta = a-b +\varepsilon_{kj}-\varepsilon_{ji}$. Hence:
\begin{gather*}
P(Y_{k}-2Y_{j}+Y_{i}<0|\Xi) \leq P(\varepsilon_{kj}<\varepsilon_{ji}|\Xi).
\end{gather*}
which implies $\tilde{\theta} \geq 2 P(\varepsilon_{kj}>\varepsilon_{ji}|\Xi) - 1$ and in order to show that $\tilde{\theta} \geq 0$ we need $P(\varepsilon_{kj}<\varepsilon_{ji}|\Xi)=0.5$ but that follows from symmetry of $f_{\Delta \varepsilon}(\cdot,\cdot)$ along the 45$^{\circ}$ line.

(ii) It is enough to note that if $T$ is linear $P(Y_{k}-2Y_{j}+Y_{i}<0|\Xi) = P(\varepsilon_{kj}<\varepsilon_{ji}|\Xi)$, which implies $\tilde{\theta}=0$.

(iii) This part follows from an argument mirroring the one in (i).
\end{proof}

\begin{remark}
By direct calculation $f_{\Delta \varepsilon}(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2) = \int f(\varepsilon - \varepsilon^{\Delta}_1)f(\varepsilon)f(\varepsilon + \varepsilon^{\Delta}_2) d\varepsilon$. Therefore, $f_{\Delta \varepsilon}(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2)$ is symmetric along the 45$^{\circ}$ line if the probability density function of $\varepsilon$, $f$, is symmetric around zero.
Note that symmetry of the error term distribution is assumed in \cite{abrevaya_jiang05}.
\end{remark}

\begin{remark}
If we defined our test statistic using two independent pairs of observations, namely:
\small
\begin{gather*}
\tilde{S}_{n}=\frac{\sqrt{n} }{n(n-1)(n-2)(n-3)}\sum_{i\neq j \neq k \neq l} \mathbbm{1}\{Y_{i}<Y_{j}<Y_{k}<Y_{m} \} sgn(Y_{m}-Y_{k}-Y_{j}+Y_{i})K_{h}\left((X_{mk}-X_{ji})'\hat{\beta}\right)
\end{gather*}
\normalsize
the requirement for symmetry of $f_{\Delta \varepsilon}(\varepsilon^{\Delta}_1,\varepsilon^{\Delta}_2)$ could potentially be dropped. This would, however, come at the increased computational cost as we have a 4-th order U statistic now. Also as we require four observations on $Y$ with approximately equally spaced index values $X'\beta_0$ instead of triples in our original statistic, the test based on $\tilde{S}_{n}$ is likely to have lower finite sample power than our baseline test.
\end{remark}

In order to obtain critical value for our tests we will assume that the model under the null hypothesis is linear. This is the worst-case $H_{0}$ for testing concavity/convexity as any small local deviation from linearity violates the hypothesis. This can also be seen from the proof of Proposition \ref{prop:plim}. In order to derive the asymptotic distribution of our statistic we make the following assumptions.

\begin{assumption}
\label{A1}
\begin{enumerate}[(a)]
\item The kernel $K(\cdot)$ is a bounded, nonnegative, symmetric, twice continuously differentiable function with support on $[-1,1]$ and uniformly bounded derivatives satisfying: \label{A1a}
\begin{enumerate}[(i)]
\item $\int K(s) ds = 1$,
\item $\int s^{2} K(s) ds < \infty$.
\item  \label{A1aiii} Let $\mathcal{K}$ be the antiderivative of $K$. We have:
\begin{gather*}
\int \int K(s_1)K(s_2) [\mathcal{K} (2s_1-s_2) + 2 \mathcal{K} (2s_2-s_1)] s_1 d s_1 d s_2 > 0
\end{gather*}
\end{enumerate}
\item $h (\log n)^5 \to 0$ and
$n h^3 \to \infty$ as $n\to \infty$. \label{A1b}
\item Conditional on the remaining regressors, the distribution of the first element of $X$ is absolutely continuous with respect to the Lebesgue measure, with bounded and twice continuously differentiable density and uniformly bounded second derivatives. Each element of $X$ has a finite fourth moment. \label{A1c}
\item The density of $\varepsilon$ is bounded and twice continuously differentiable, the derivatives are uniformly bounded.
\item The estimator of $\beta_0$, $\hat{\beta}$, satisfies: \label{A1d}
\begin{gather*}
\hat{\beta}-\beta_0 = \frac{1}{n} \sum_{i=1}^{n}\Omega(X_{i}, Y_i) + o_{p}(n^{-1/2}).
\end{gather*}
where
$E[\Omega(X_{i}, Y_i)]=0$ and each element of $\Omega$ has a finite fourth moment and is continuous in the second argument.
\end{enumerate}
\end{assumption}

Assumption \ref{A1}\eqref{A1a}(iii) is nonstandard. It is not used for our global test, $S_n$, in this section but is needed for our local test in the next section to have power. Technically, the condition implies that the mean of our local statistic is negative (see proof of Theorem \ref{thm:bootpwr}). This assumption is satisfied by frequently used kernels like standard normal or Epanechnikov (with integral taking values 0.0072 and 0.0177, respectively). The bandwidth rate condition in Assumption \ref{A1}\eqref{A1b} is rather weak and allows standard ``rule-of-thumb'' bandwidth choice $h \sim n^{-1/5}$. Assumption \ref{A1}\eqref{A1d} requires $\hat{\beta}$ to be consistent
under both the null and alternative hypothesis and is satisfied by a wide range of estimators including Han's MRC or Ichimura's semiparametric least squares (see e.g. appendix in \cite{szydlowski20} for discussion).

Let $\xi_{ji}$ denote the index $X_{ji}'\beta_0$ and $f_{\xi|\xi}$ denote the density of $\xi_{kj}$ given $\xi_{ji}$. Define:
\begin{align*}
H(Y_{i},X_{i},\xi_{ji},\xi_{kj}) &= 2E[\mathbbm{1}\{Y_{i}<Y_{j}<Y_{k}\}sgn(Y_{k}-2Y_{j}+Y_{i})|Y_{i},X_{i},\xi_{ji},\xi_{kj}]\\
& \quad+2E[\mathbbm{1}\{Y_{j}<Y_{i}<Y_{k}\}sgn(Y_{k}-2Y_{i}+Y_{j})|Y_{i},X_{i},\xi_{ji},\xi_{kj}]\\
& \quad +2E[\mathbbm{1}\{Y_{j}<Y_{k}<Y_{i}\}sgn(Y_{i}-2Y_{k}+Y_{j})|Y_{i},X_{i},\xi_{ji},\xi_{kj}]\\
\delta(Y_{i},X_{i},\xi_{ji},\xi_{kj}) &= 3 H(Y_{i},X_{i},\xi_{ji},\xi_{kj})f_{\xi|\xi}(\xi_{kj}|\xi_{ji}=(X_{j}-X_{i})'\beta_0) \\
G(\xi_{ji},\xi_{kj}) &= 6E[\mathbbm{1}\{Y_{i}<Y_{j}<Y_{k}\}sgn(Y_{k}-2Y_{j}+Y_{i})(X_{kj}-X_{ji})'|\xi_{ji},\xi_{kj}]\\
\mu(\xi_{ji},\xi_{kj}) &=  4G(\xi_{ji},\xi_{kj})f_{\xi|\xi}(\xi_{kj}|\xi_{ji})
\end{align*}
and let $\mu_{2}$ denote the derivative of $\mu$ with respect to the second argument.

\begin{theorem}\label{thm:global}
If Assumption \ref{A1} holds,
we have:
\begin{gather*}
\sqrt{n}(U_{n}-\theta) \to^{d} N(0,E[\psi_{i}^{2}])
\end{gather*}
where $\psi_{i}= E[\delta(Y_{i},X_{i},\xi_{ji},\xi_{ji})|Y_{i},X_{i}]- E[\mu_{2}(\xi_{ji},\xi_{ji})]\Omega(X_{i}, Y_i)$.
\end{theorem}

As linearity is the boundary case for testing $H_{0}$: $T(\cdot)$ is concave, or $H_{0}$: $T(\cdot)$ is convex, we will reject concavity if $S_{n}<c_{\alpha}$ and reject convexity if $S_{n}>c_{1-\alpha}$, where $c_{\alpha}$ denotes the $\alpha$ quantile from the normal asymptotic distribution. As typical with U-statistics, estimating the variance of the asymptotic distribution of $S_{n}$ is difficult. On the one hand, the plug-in estimator will involve estimating derivatives of conditional moments and distributions, which requires delicate choices of bandwidths and, essentially, calculation of higher order U-statistics. On the other hand, using the standardisation approach in \cite{ghosal_et_al00} will lead to a 5-th order U-statistic, which would be difficult to calculate with sample sizes typically encountered in applications.\footnote{Proceeding as there would involve calculating:
\scriptsize
\begin{align*}
\hat{\sigma}^{2} = \frac{1}{n(n-1)(n-2)(n-3)(n-4)}\sum_{i\neq j \neq k \neq l \neq m} &(\mathbbm{1}\{Y_{i}<Y_{j}<Y_{k}\}sgn(Y_{k}-2Y_{j}+Y_{i})K_{h}\left(X_{kj}'\hat{\beta} - X_{ji}'\hat{\beta}\right)+\tilde{\mu}_{2}\Omega_{i})\\
&\times (\mathbbm{1}\{Y_{i}<Y_{l}<Y_{m}\}sgn(Y_{m}-2Y_{l}+Y_{i})K_{h}\left(X_{ml}'\hat{\beta} - X_{li}'\hat{\beta}\right)+\tilde{\mu}_{2}\Omega_{i})\\
& + \text{symmetric terms},
\end{align*}
\footnotesize
where $\tilde{\mu}_{2}$ is an estimator of $-E[\mu_{2}(\xi_{ji},\xi_{ji})]$ (which itself is a 3-rd order U statistic).
}
Thus, in Section \ref{sec:boot} we resort to bootstrap for calculating the critical value as bootstrapping involves only repeated calculation of a 3-rd order U-statistic, $U_{n}$, which we find computationally easier than the aforementioned methods.

The global test introduced in this section only has power against global deviations from concavity/linearity/convexity and does not have power if the function is both convex and concave on different parts of the domain. Thus, it can be used as a first check -- for example, if the test rejects concavity, the researcher concludes that the function cannot be globally concave. Failure to reject would, then, mean that one has to consider our local test described in the next section, in order to verify if indeed the function is globally concave.

\subsection{Local test}
The main idea of the local test is to consider only triples of the kind portrayed in Figure \ref{fig:intui} local to a point $y$ in the domain of the transformation function. In other words, we will check if the transformation function is concave/linear/convex locally around $y$. Local concavity may be of interest by itself. However, usually we are interested in verifying if the treatment effect is concave on the whole domain, thus we will take the minimum of the local statistics at different points $y$ to run the test.

Formally, define:
\begin{align*}
U_{n}(y) = \frac{1}{n(n-1)(n-2)}\sum_{i\neq j \neq k} \mathbbm{1}\{Y_{i}<Y_{j}<Y_{k} \} & sgn(Y_{k}-2Y_{j}+Y_{i})  K_{h}(Y_{i}-y)K_{h}(Y_{j}-y)\\
& \quad \quad \quad \quad \times K_{h}(Y_{k}-y)K_{h}\left((X_{kj} - X_{ji})'\hat{\beta}\right)
\end{align*}
where in practice the bandwidth used for $(Y_{i}-y)$ can be different than for $(X_{kj} - X_{ji})'\hat{\beta}$ but, in order to simplify exposition and mathematical arguments, we assume that both bandwidths are of the same order and denote both by $h$. Now our local test statistic for testing concavity is defined as:
\begin{gather*}
S_{n}^{conc} = \inf_{y \in \mathcal{Y}} \sqrt{nh} U_{n}(y)
\end{gather*}
where $\mathcal{Y}$ is a compact set, and the statistic for testing convexity is defined with $\sup$ replacing $\inf$ above. Linearity can be tested by replacing $U_{n}(y)$ with its absolute value. For the rest of the article we concentrate on $S_{n}^{conc}$ as results for testing convexity and linearity follow by very similar arguments.

Intuitively, low values of $S_{n}^{conc}$ show that there is a large deviation from concavity at some point $y$ and, hence, the treatment effects are not accelerating on the whole domain of the outcome.\footnote{Note that by the formulas in Example 1 concavity of $T$ is equivalent to convexity of the outcome in the treatment.} Therefore, the null hypothesis of concavity would be rejected if $S_{n}^{conc} < c_{\alpha}^{conc}$ where $c_{\alpha}^{conc}$ is an appropriate quantile from the asymptotic approximation to the distribution of our statistic.

In order to obtain $c_{\alpha}^{conc}$ one could imagine proceeding as in \cite{ghosal_et_al00}: approximate the standardised U-statistic process $\sqrt{n} U_{n}(y)/\sigma_{n}(y)$, where $\sigma_{n}(y)$ is the estimator of the asymptotic variance, by a Gaussian process and then apply the extreme value theory in order to derive the distribution of the infimum. However, pursuing that approach is difficult in our setup for two reasons: 1) under $H_0$ we still have $E[U_n(y)]=O(h)$ rather than $o(h)$ therein, 2) as discussed above, estimating $\sigma_{n}(y)$ is computationally expensive.
Instead we propose to use bootstrap to approximate $c_{\alpha}^{conc}$.

Define:
\footnotesize
\begin{align*}
\phi_{i,n}(y) =& 6(E[ \mathbbm{1}\{Y_{i}<Y_{j}<Y_{k} \} sgn(Y_{k}-2Y_{j}+Y_{i})K_{h}(Y_{i}-y)K_{h}(Y_{j}-y)K_{h}(Y_{k}-y)K_{h}\left((X_{kj} - X_{ji})'\beta_0\right)|Y_{i},X_{i}]\\
&+E[ \mathbbm{1}\{Y_{j}<Y_{i}<Y_{k} \} sgn(Y_{k}-2Y_{i}+Y_{j})K_{h}(Y_{i}-y)K_{h}(Y_{j}-y)K_{h}(Y_{k}-y)K_{h}\left((X_{ki} - X_{ij})'\beta_0\right)|Y_{i},X_{i}]\\
&+E[ \mathbbm{1}\{Y_{j}<Y_{k}<Y_{i} \} sgn(Y_{i}-2Y_{k}+Y_{j})K_{h}(Y_{i}-y)K_{h}(Y_{j}-y)K_{h}(Y_{k}-y)K_{h}\left((X_{ik} - X_{kj})'\beta_0\right)|Y_{i},X_{i}]).
\end{align*}
\normalsize
The following asymptotic approximation will be useful in justifying our bootstrap procedure.

\begin{theorem}\label{thm:local}
If Assumption \ref{A1} holds, then:
\begin{gather*}
\sup_{y \in \mathcal{Y}}\left|U_{n}(y)  - E[U_n(y)] - \frac{1}{n} \sum_{i=1}^{n} \phi_{i,n}(y)\right| = o_{p}((nh)^{-1/2})
\end{gather*}
\end{theorem}

An interesting consequence of this result is that
the asymptotic distribution of the statistic does not depend on the estimation of $\beta_0$ as long as Assumption \ref{A1}\eqref{A1d} is satisfied,  unlike the global test (cf. Theorem \ref{thm:global}).

\section{Bootstrap critical values}\label{sec:boot}

Our bootstrap procedure for obtaining the critical value for the global or local test is as follows:
\begin{enumerate}
\item Estimate $\beta_0$ (e.g. by Han's MRC) and calculate the residuals $\hat{\varepsilon}_{i} = Y_{i} - X_{i}'\hat{\beta}$.
\item For the global test: draw a random sample $\{v_i\}_{i=1}^{n}$ from a two point distribution with $P(v_{i}=-1)=P(v_{i}=1)=1/2$ and define $\varepsilon_{i}^{*} =v_{i} \hat{\varepsilon}_{i}$. For the local test: (a) draw $\{\varepsilon_i^*\}_{i=1}^n$ with replacement from $\{\hat{\varepsilon}_i\}_{i=1}^n$ or (b) follow the same procedure as for the global test. Then, generate $Y_{i}^*$ by: \label{boot2}
\begin{gather*}
Y_{i}^{*} = X_{i}'\hat{\beta} + \varepsilon_i^*.
\end{gather*}
 \item Estimate $\beta_0$ (e.g. by Han's MRC) using the bootstrap sample. Let the resulting estimate be denoted by $\beta^*$.
 \item Calculate the statistic $S_{n}$ or $S_{n}^{conc}$ on the bootstrap sample using $\beta^*$ instead of $\hat{\beta}$. Denote the resulting bootstrap statistics by $S_{n}^{*}$ and $S_{n}^{conc,*}$.
 \item Obtain the empirical distribution of $S_{n}^*$ and $S_{n}^{conc,*}$ by repeating steps 1-4 many times. Calculate the $\alpha$ quantiles of the empirical distribution of $S_{n}^*$ and $S_{n}^{conc,*}$ and denote them by $c_{\alpha}^{*}$ and $c_{\alpha}^{conc,*}$, respectively.
\end{enumerate}

Step two imposes the null hypothesis of linearity on the bootstrap sample so it ``recentres'' the bootstrap statistic on the linear case. Furthermore, sampling from a symmetric two point distribution in the wild bootstrap imposes symmetry of the error distribution in the bootstrap sample, thus implying that Assumption \ref{A:SYM} is satisfied in that sample.
Note that the jacknife multiplier bootstrap (JMB) of \cite{chen_kato20} will not work for the local test here (without modifications) as we have $E[U_n(y)]=O(h)$ even under the null so the JMB statistic will not be centred correctly.
We assume an equivalent of Assumption \ref{A1}\eqref{A1d} for the bootstrap estimator $\beta^{*}$:
\begin{assumption}
\item The estimator of $\beta^*$ satisfies:
$\beta^* - \hat{\beta} = \frac{1}{n} \sum_{i=1}^{n}\Omega(X_{i}, Y_i^*) + o_{p}(n^{-1/2})$,
where
$E[\Omega(X_{i}, Y_i^*)]=0$ has zero mean and each element of $\Omega$ has a finite fourth moment and is continuous in the second argument.
\label{BA1d}
\end{assumption}
This assumption is satisfied in our leading case, when $\beta^*$ is estimated by MRC (see \cite{subbotin07}).

\begin{theorem}\label{thm:boot}
If Assumptions \ref{A1}-\ref{BA1d} hold
and $T(\cdot)$ is linear:
\begin{gather*}
\lim_{n\to \infty} P(S_{n}^{conc} \leq c_{\alpha}^{conc,*}) = \alpha
\end{gather*}
If, additionally, Assumption \ref{A:SYM} holds, we have:
\begin{gather*}
\lim_{n\to \infty} P\left(\sqrt{n}U_n \leq c_{\alpha}^{*}\right) = \alpha .
\end{gather*}
(and equivalent result holds for a test of linearity or convexity).
\end{theorem}

Theorem \ref{thm:boot} implies, for example, that we can reject global concavity of $T$, i.e. increasing treatment effect, when $S_{n}^{conc}<c_{\alpha}^{conc,*}$, and reject global linearity when $|S_{n}^{conc}|>c_{1-\alpha/2}^{conc,*}$. We prove validity of bootstrap in unconditional probability, which is a weaker result than the standard result conditional on the sample (see e.g. \cite{cavaliere_georgiev20} for discussion), as the fact that $E[U_n(y)] = O(h)$ under $H_0$ and the need to use parametric bootstrap complicates the analysis. Note that we only require the symmetry of the error distribution for the global test. The next theorem characterises alternatives for which our tests are consistent.

Assume that $c^{conc,*}_{\alpha}$ is generated using method (a) in step 2 above (see Appendix \ref{app:wildboot} for the analysis with the wild bootstrap in (b)). Let $f_{xb}$ denote the pdf of $X_i'\beta_0$
and $\mathbf{1} = [1 \quad  1 \quad  1]'$. Further, let $f_{\boldsymbol{\varepsilon}}$ denote the joint distribution of $(\varepsilon_i,\varepsilon_j,\varepsilon_k)$ and $\nabla f_{\boldsymbol{\varepsilon}}$ denote its gradient and define $f_{\boldsymbol{\tilde{\varepsilon}}}$,  $\nabla f_{\boldsymbol{\tilde{\varepsilon}}}$, similarly for $\tilde{\varepsilon} = Y -X'\beta_0$.

\begin{theorem}\label{thm:bootpwr}
If Assumptions \ref{A1}-\ref{BA1d} and Assumption \ref{A:SYM} hold,
and $T(\cdot)$ is globally convex:
\begin{gather*}
\lim_{n\to \infty} P\left(\sqrt{n}U_n \leq c_{\alpha}^{*}\right) = 1.
\end{gather*}
Moreover, if Assumptions \ref{A1}-\ref{BA1d} hold, $T(\cdot)$ is locally convex at some $y$ and the following condition holds:
\scriptsize
\begin{align}
 &\sup_{y  \in \mathcal{Y}} \int \int \nabla f_{\boldsymbol{\tilde{\varepsilon}}} \begin{pmatrix} y -\zeta +\xi \\ y -\zeta  \\ y -\zeta -\xi \end{pmatrix}' \mathbf{1}  f_{xb}(\zeta-\xi)f_{xb}(\zeta)f_{xb}(\zeta+\xi) d \zeta d\xi
  \nonumber \\
& \quad - \sup_{y \in \mathcal{Y}} \int \int   \left\{ T'(y)^4 \nabla f_{\boldsymbol{\varepsilon}}   \begin{pmatrix} T(y) -\zeta +\xi \\ T(y) -\zeta  \\ T(y) -\zeta -\xi  \end{pmatrix}  ' \mathbf{1} + 3T''(y)T'(y)^2 f_{\boldsymbol{\varepsilon}} \begin{pmatrix} T(y) -\zeta +\xi \\ T(y) -\zeta   \\ T(y) -\zeta -\xi  \end{pmatrix} \right\}
 f_{xb}(\zeta-\xi)f_{xb}(\zeta)f_{xb}(\zeta+\xi) d \zeta d\xi
< 0,  \label{eq:pwrcond}
\end{align}
\normalsize
then we have:
\begin{gather*}
\lim_{n\to \infty} P(S_{n}^{conc} \leq c_{\alpha}^{conc,*}) = 1
\end{gather*}
(and equivalent result holds for a test of linearity or convexity).
\end{theorem}

The condition in \eqref{eq:pwrcond} is technical and it is not straightforward to find general sufficient conditions for it to hold as it depends not only on the shape of $T$  but also the distributions of $\epsilon$ and $X$.
However, we can make the following observations.

\begin{remark}
For testing linearity (\emph{vel} testing treatment effect heterogeneity), the condition would only require the expression in \eqref{eq:pwrcond} to be nonzero and, thus, should be satisfied generically for non-linear $T(y)$'s, and the test should be consistent against (almost) all nonlinear alternatives.
\end{remark}

\begin{remark}
Theorem \ref{thm:bootpwr} shows the cost of using a simple bootstrap test that avoids estimation of the transformation function, in particular the fact that our bootstrap does not correctly approximate the error distribution $f$ but rather distribution of OLS residuals, $\tilde{\varepsilon}$. The latter distribution depends on the distribution of $X$ and on $T$, complicating the power properties of the test. In principle, one could try to estimate $T'$ and $\nabla f_{\boldsymbol{\varepsilon}}$ and approximate the term:

\scriptsize
\begin{gather*}
 \sup_{y \in \mathcal{Y}} \int \int  T'(y)^4 \nabla f_{\boldsymbol{\varepsilon}}   \begin{pmatrix} T(y) -\zeta +\xi \\ T(y) -\zeta  \\ T(y) -\zeta -\xi  \end{pmatrix}  ' \mathbf{1} f_{xb}(\zeta-\xi)f_{xb}(\zeta)f_{xb}(\zeta+\xi) d \zeta d\xi
\end{gather*}
\normalsize
by a bootstrap procedure, making sure that condition \eqref{eq:pwrcond} would be satisfied.  However, this approach would substantially complicate the calculation of the critical value as estimators of $T$ and $f$ in the transformation model often involve non-smooth criterion functions and do not readily offer estimates of the derivatives (e.g. \cite{chen02}).
Thus, for practical reasons we prefer our current approach, even though it carries a cost in terms of power of the tests for concavity and convexity.
\end{remark}

\begin{remark}
At an additional computational cost one could estimate $T$ and use semiparametric residuals: $\hat{\varepsilon}_i = \hat{T}(Y_i) - X_i'\hat{\beta}$ in our bootstrap. We look at the latter case in Appendix \ref{app:MCnpar} where we show that the finite sample performance of such test does not dominate our local test.
\end{remark}

Finally, let us note that the computation of the local statistic involves evaluating a third order U-statistic at different points $y$ and, hence, takes significantly longer to compute than the global statistic. In practice, we recommend running the global test first and then proceed with the local test if the null hypothesis cannot be rejected in the first step.

\section{Monte Carlo simulations}\label{MC}
The data is generated from the following five models:
\begin{align*}
Y &= X + \varepsilon & \text{(D0)}\\
\log(Y + 2.12) - \log(2.12) &= X + \varepsilon & \text{(D1)}\\
\frac{1}{13}\sinh(2Y) &= X + \varepsilon & \text{(D2)}\\
\log(2.12) - \log(2.12-Y) &= X + \varepsilon & \text{(D3)} \\
5 \tanh(0.5Y)& =  X + \varepsilon & \text{(D4)}
\end{align*}
where
we draw $X$ and $\varepsilon$ from the standard normal distribution. Note that D0 is the worst-case model in the null hypothesis, D1 imposes concavity of the transformation, D3 convexity and D2 and D4 are neither concave or convex.

\begin{figure}[th]
\centering
\caption{Monte Carlo designs}
\includegraphics[width=0.7\textwidth]{mcdesign3.png}
\label{MCdesign1}
\end{figure}

We run 1000 Monte Carlo replications. We use Gaussian kernel functions and rule-of-thumb bandwidths for both $(X_{kj}-X_{ji})'\hat{\beta}$ and $Y_{i}$, namely $h=1.06\hat{\sigma}n^{-1/5}$,
where $\hat{\sigma}$ is a sample standard deviation of $(X_{kj}-X_{ji})'\hat{\beta}$ or $Y_{i}$.\footnote{The choice of bandwidth for $Y_i$ seems trickier as some transformations $T^{-1}$ make the distribution of $Y_i$ severly skewed. We have also tried using cross-validation (\emph{ucv} in R) and plug-in cross-validation (\emph{bcv} in R) with similar results.}
In order to calculate the local statistic $S_{n}^{conc}$ we use a grid of values for $y$: -2:0.25:2, and take a minimum over the grid. The number of bootstrap replications used to calculate the critical value is 500 and we consider three sample sizes: $n=100, 250$ and $500$.

\begin{table}[ht]
	\centering
	\footnotesize
\begin{tabular}{ccccccccccc}
\toprule
	&		&	\multicolumn{3}{c}{Global test}					&	\multicolumn{3}{c}{Local test}					&	\multicolumn{3}{c}{Local test - wild bootstrap}					\\
	&		&	$n=100$	&	$n=250$	&	$n=500$	&	$n=100$	&	$n=250$	&	$n=500$	&	$n=100$	&	$n=250$	&	$n=500$	 \\  \midrule
H0 true	&	D0	&	0.047	&	0.051	&	0.062	&	0.056	&	0.048	&	0.053	&	0.060	&	0.063	&	0.053	 \\
H0 true	&	D1	&	0.000	&	0.000	&	0.000	&	0.000	&	0.000	&	0.000	&	0.000	&	0.000	&	0.000	 \\
H0 false	&	D2	&	0.105	&	0.122	&	0.099	&	0.868	&	0.994	&	1.000	&	0.717	&	0.953	&	0.995	 \\
H0 false	&	D3	&	1.000	&	1.000	&	1.000	&	1.000	&	1.000	&	1.000	&	1.000	&	1.000	&	1.000	 \\
H0 false	&	D4	&	0.079	&	0.073	&	0.092	&	0.878	&	0.932	&	0.938	&	0.615	&	0.657	&	0.677	 \\
\bottomrule
\end{tabular}
		\caption{Test of concavity, rejection probabilities, 5\% level}
	\label{tab:MC1}
		\floatfoot{Note: 1000 Monte Carlo simulations, 500 bootstrap replications.}
\end{table}
\normalsize

Table \ref{tab:MC1} contains the results of the Monte Carlo simulations. We concentrate on testing concavity, as the results for the linearity and convexity tests, are very similar. The rejection probabilities in the linear case (D0) are close to the nominal level for both the global and the local test and both tests have perfect power against a globally convex alternative in D3.

As predicted, the global test has low power against D2 and D4, for which the transformation function is concave on part of the domain and convex on the other part. In this case deviations from concavity, as measured by our global statistic, cancel with positive values of the statistic obtained for the part of the domain where the function is concave, resulting in the global statistic taking values close to zero, just as for the linear case.

For designs D2 and D4, our local test significantly improves over the global test with almost perfect detection of non-concavity in D2 for $n = 250$. The power against D4 increases slower with $n$ than D2 but we still reject H0 with close to certainty with $n=500$. This is in line with the intuition that the local test will concentrate on the region of the largest violation of concavity instead of averaging over the measures of concavity for different regions.

Finally, note that, unlike for the global test, sampling from a symmetric distribution is not required in the local test bootstrap, so it is a question of finite sample performance if we would prefer the regular parametric bootstrap (part (a) in Step 2) or the symmetric wild parametric bootstrap in part (b). Comparing the last two panels in Table \ref{tab:MC1} shows that both bootstrap tests have the right level but the symmetric wild bootstrap test has lower power for our designs (see D2 and D4).

\section{Application: Curvature of loan demand}\label{appl}

We use data from \cite{karlan_zinman08}, who ran randomised trials with a for-profit consumer lender in South Africa targeting high-risk consumer loan market. The lender randomised individual interest rate direct mail offers
, conditional on the client’s risk category. \cite{karlan_zinman08} found that the loan demand curves are downward sloping. We will investigate if the demand curves are convex or, in other words, if increases in interest rate have a stronger negative effect on loan demand the lower the rate.

\begin{figure}[h]
\centering
\caption{Estimated transformation}
\includegraphics[width=0.6\textwidth]{Tplot.eps}
\label{fig:Tappl}
\end{figure}

The dependent variable is the amount borrowed (in rands) at the offered interest rate.\footnote{We abstract from selectivity issues here. See \cite{karlan_zinman08} for discussion.} The experiment was ran in three mailer waves over four months, thus as in \cite{karlan_zinman08} we include the risk category and wave dummies as controls in $X$. The data contains 2325 observations. As the first step, we estimated the transformation function in our model using the estimator in \cite{chen02} and plotted it in Figure \ref{fig:Tappl} together with a smoothed version.\footnote{This is for illustrative purposes only. Note that the estimator in \cite{chen02} does not impose monotonicity, so a full exercise of estimating $T$ should include additional step of monotonising the estimate. We want to stress that our testing procedure does not rely on any estimator of $T$.} The figure suggests that the transformation is not far from being concave on most of its domain, maybe besides the values of loan size below 1000 rands, which implies convex demand curve by formulas in Example 1.


\begin{table}[h]
	\centering
	\small
\begin{tabular}{c|ccc|cccccc}
\toprule
	&	\multicolumn{3}{c}{Global test}							&	\multicolumn{3}{|c}{Local test}							&	\multicolumn{3}{c}{Local test, $loan \geq 1000$} 					\\
	&				&	\multicolumn{2}{c|}{Reject $H_0$?}			&				&	\multicolumn{2}{c}{Reject $H_0$?}			&		&	\multicolumn{2}{c}{Reject $H_0$?}			\\
$H_0$	&	Statistic			&	5\%	&	10\%	&	Statistic			&	5\%	&	10\%	&	Statistic	&	5\%	&	10\%	 \\ \midrule
convexity	&	0.26			&	No	&	No	&	$-0.27 \times 10^{-4}$			&	Yes	&	Yes	&	$0.19 \times 10^{-4}$	&	No	&	No	 \\
linearity	&	0.26			&	Yes	&	Yes	&				&		&		&		&		&		 \\
concavity	&	0.26			&	Yes	&	Yes	&				&		&		&		&		&		 \\
\bottomrule
\end{tabular}
\caption{Testing curvature of loan demand}
\label{tab:appl}
\floatfoot{Note: For the local test we report the statistic multiplied by $n(n-1)(n-2)$. We choose bandwidth using cross-validation. 500 bootstrap replications. Data source: \cite{karlan_zinman08}, replication files.}
\end{table}

In order to formally test our conjectures, we first apply our global test to verify if the demand function is concave, linear or convex.  As Table \ref{tab:appl} shows, global test rejects linearity and concavity of demand. Thus, in the second step we apply the local test to the null hypothesis of convexity and, in fact, reject the null hypothesis, concluding that the loan demand is not globally convex in the interest rate.
To shed more light on where the non-convexity may come from, we re-ran the test on the sample excluding loans lower than 1000 rands and found that convexity is not rejected on this sub-sample. Therefore, overall we confirm that loan demand is mostly convex in interest rate, besides very small loans sector of the market.

\section{Conclusion}
Our application demonstrates usefulness of our testing procedures in recovering monotonicity of treatment effects. Particular appeal of the procedures described in this article comes from the fact that they avoid estimation of the transformation function and only require OLS estimation of the vector of coefficients $\beta_0$. Thus, they are relatively easy to implement. Additionally, computation of the third order U-statistic involved in our tests can be done efficiently by sorting the data by $Y$ first -- this reduces computational complexity to $O(nlog(n)+n(n-2)/2)$ from $O(n^{3})$ for a straightforward triple loop through the observations.

\newpage
\section*{Appendix}