Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
46,478 characters · 10 sections · 38 citation commands
Testing Shape Restrictions with Continuous Treatment: A Transformation Model Approach
In this paper we consider testing if a treatment has a diminishing or increasing effect on the outcome. This is often of interest on top of the question if there is an effect at all or what sign it has. For example, often it is natural to expect that demand will be decreasing in price. However, it is less clear if increasing the price will have a larger effect at lower or higher price levels. Our test can address this question without fully estimating the demand relationship.
Let $X$ denote the vector of treatment and control variables. As estimating the effect of $X$ on the outcome $Y$ nonparametrically with non-trivial number of treatments and controls would suffer from a curse of dimensionality, we impose a single-index structure and assume that the nonlinearity of the treatment effect comes from a nonlinear transformation of $Y$. In other words, consider a transformation model of the form:
where $Y$ is a scalar dependent variable, $X$ is a vector of $q$ nondegenerate explanatory variables, $\beta_0$ is a vector of coefficients belonging to a compact set $\Theta_{\beta} \subset \mathbb{R}^q$, $T(\cdot)$ is an increasing function and $\varepsilon$ is an unobserved error term with distribution $F$ that is independent of $X$. In order to identify the model we need a location normalisation: e.g. $T(0)=0$, $E(\varepsilon)=0$ or $Me(\varepsilon)=0$; and a scale normalisation: e.g. $|\beta_{0,1}|=1$. The benefit of using the transformation model compared to a standard single-index model (i.e. $Y=T(X'\beta_0) + \varepsilon$), besides the fact that it facilitates our testing approach, is that the transformation model allows the treatment effect of $X_{k}$ to depend on the values of both observed and unobserved characteristics (note that $Y=T^{-1}(X'\beta_0+\varepsilon)$) so can be seen as a simple way of introducing heterogenous treatment effects that vary with unobservables.
The main objective of this article is to develop a practically appealing test to determine whether the transformation function $T(\cdot)$ is concave/linear/convex. The examples below illustrate the importance of testing curvature of the transformation.
Beyond these examples our procedure can also be used for specification search, i.e. determining if one should use a concave or convex transformation, and to test the curvature in wage regressions, e.g. if the effect of education or experience is concave, or the curvature of the marginal utility (or profit) function in hedonic models (see ekeland_et_al04).
Our test statistic simply compares triples of $Y$'s corresponding to equally spaced index values $X'\beta_0$, thus it does not require estimation of the transformation function $T$ or the distribution of $\varepsilon$. We only require an estimator of $\beta_0$ and, basically, the symmetry of the distribution of the error terms. Estimating $\beta_0$ can be done, for example, by using the maximum rank correlation estimator, han87, or semiparametric least squares, ichimura93.
We propose both a global test that has power to detect globally convex or concave functions and leads to asymptotic normal critical values, as well as a more general test that detects local deviations from linearity, i.e. has power against alternatives that are both convex and concave on the domain of $T$. Our tests do not have a pivotal asymptotic distribution but we show that the critical values can be obtained by parametric bootstrap.
The cost of our test’s easy implementation is that the power properties of our local test are complex, and the test may not be consistent against some relevant alternatives. Nevertheless, our application demonstrates the usefulness of our testing approach in detecting nonconvexities in the demand for loans as a function of the interest rate.
Our statistic resembles the approach in abrevaya_jiang05 who test curvature in a nonparametric regression model. However, unlike their approach our test does not suffer from the curse of dimensionality due to the single index structure of the regression part and allows the marginal effect of covariates to vary with the error term (though in a manner restricted by the single-index). Also the details of the derivation of the asymptotic distribution are different due to presence of estimated $\beta_0$ in our statistic and somewhat distinct approach to obtaining power against general alternatives. These traits are shared by abrevaya_et_al10, who test monotonicity in a generalised regression model with endogeneity, though a distinguishing feature of our work is that we formally show validity of bootstrap for obtaining critical values.
Related literature includes tests for the sign of the treatment effect, see e.g. kline16, and testing for treatment effect heterogeneity (abadie02, crump_et_al08, santanna21, chernozhukov_et_al23). Unlike the latter papers, our approach allows the treatment effect to vary with (independent) unobserved heterogeneity at the cost of imposing much more structure on the treatment effect model and the covariates. Testing curvature of the transformation can also be seen as a generalisation of specification testing in neumeyer_et_al16 and szydlowski20. The idea of using curvature of the integrated hazard to test monotonicity of the baseline hazard has been utilised by hall_van_keilegom05. Testing shape restrictions in a nonparametric regression model has been considered by ghosal_et_al00, gutknecht16, chetvertikov19 and komarova_hidalgo23, among others. Similarly to this paper, chen_kato20 propose a bootstrap procedure for approximating the supremum of a local U-statistic.
The article is organized as follows. Section (ref) discusses the idea behind the testing procedure informally and the formal results are postponed till Sections (ref)-(ref). Sections (ref)-(ref) contain Monte Carlo results and our application to loan demand. Proofs, besides the proof of the main proposition, are located in the Appendix, which also contains some additional MC simulations.
Figure (ref) portrays the intuition behind our test. The transformation plotted in the figure is concave and we display three ordered data points in the figure, $(Y_{i},Y_{j},Y_{k})$, for which the transformation function is equally spaced, i.e. $T(Y_{k})-T(Y_{j})=T(Y_{j})-T(Y_{i})$.
Concavity of $T(\cdot)$ implies that $Y_{k}-Y_{j}>Y_{j}-Y_{i}$. Note that $T(Y_{k})-T(Y_{j})=T(Y_{j})-T(Y_{i})$ is equivalent to $(X_{k}-X_{j})'\beta_0 +\varepsilon_{k}-\varepsilon_{j}= (X_{j}-X_{i})'\beta_0 +\varepsilon_{j}-\varepsilon_{i}$. Hence, “on average” equally spaced $T(Y)$'s mean equally spaced index values $X'\beta_0$. Therefore, we can detect deviations from concavity by considering the following criterion:
where $X_{ji} \equiv X_{j}-X_{i}$.
In order to make this criterion operational with continuous distribution of $X'\beta_0$ (which is required for identification) we need to replace the last indicator function with a smooth kernel $K_{h}(\cdot) = h^{-1}K(\cdot/h)$ and $\beta_0$ with its estimator $\hat{\beta}$:
Deviations from convexity can be detected in a similar fashion. Finally, we can combine both criterion functions to detect deviations from linearity:
Here, very negative values of the objective function signify convexity and large positive values mark concavity. Outside testing, one may use the above measure itself to characterise an “average” curvature of function $T$, e.g. large positive values suggest that the function is predominantly concave.
Note that our test requires only one-dimensional kernel smoothing, thus it does not suffer from the curse of dimensionality. Also, unlike the specification test in szydlowski20 it does not require estimation of the transformation function, a computationally intense task itself.
As the probability limit of the criterion functions introduced in the previous section depends on the extent of concavity, but these functions are asymptotically centred at known values under linearity, we form our procedure as a test of:
depending on the question of interest. We will use the test statistic
and reject the null hypothesis at level $\alpha$ if $S_{n}<c_{\alpha}, |S_{n}|>c_{1-\alpha/2}$ or $S_{n}>c_{1-\alpha}$ depending on $H_{0}$, respectively, where $c_{\alpha}$ denotes an $\alpha$ quantile from an appropriate asymptotic distribution.
Define $\varepsilon_{ji}=\varepsilon_j - \varepsilon_i$. Proposition (ref) justifies our testing strategy. We need the following symmetry assumption:
In order to obtain critical value for our tests we will assume that the model under the null hypothesis is linear. This is the worst-case $H_{0}$ for testing concavity/convexity as any small local deviation from linearity violates the hypothesis. This can also be seen from the proof of Proposition (ref). In order to derive the asymptotic distribution of our statistic we make the following assumptions.
Assumption (ref)(ref)(iii) is nonstandard. It is not used for our global test, $S_n$, in this section but is needed for our local test in the next section to have power. Technically, the condition implies that the mean of our local statistic is negative (see proof of Theorem (ref)). This assumption is satisfied by frequently used kernels like standard normal or Epanechnikov (with integral taking values 0.0072 and 0.0177, respectively). The bandwidth rate condition in Assumption (ref)(ref) is rather weak and allows standard “rule-of-thumb” bandwidth choice $h \sim n^{-1/5}$. Assumption (ref)(ref) requires $\hat{\beta}$ to be consistent under both the null and alternative hypothesis and is satisfied by a wide range of estimators including Han's MRC or Ichimura's semiparametric least squares (see e.g. appendix in szydlowski20 for discussion).
Let $\xi_{ji}$ denote the index $X_{ji}'\beta_0$ and $f_{\xi|\xi}$ denote the density of $\xi_{kj}$ given $\xi_{ji}$. Define:
and let $\mu_{2}$ denote the derivative of $\mu$ with respect to the second argument.
As linearity is the boundary case for testing $H_{0}$: $T(\cdot)$ is concave, or $H_{0}$: $T(\cdot)$ is convex, we will reject concavity if $S_{n}<c_{\alpha}$ and reject convexity if $S_{n}>c_{1-\alpha}$, where $c_{\alpha}$ denotes the $\alpha$ quantile from the normal asymptotic distribution. As typical with U-statistics, estimating the variance of the asymptotic distribution of $S_{n}$ is difficult. On the one hand, the plug-in estimator will involve estimating derivatives of conditional moments and distributions, which requires delicate choices of bandwidths and, essentially, calculation of higher order U-statistics. On the other hand, using the standardisation approach in ghosal_et_al00 will lead to a 5-th order U-statistic, which would be difficult to calculate with sample sizes typically encountered in applications.\footnote{Proceeding as there would involve calculating: \scriptsize
where $\tilde{\mu}_{2}$ is an estimator of $-E[\mu_{2}(\xi_{ji},\xi_{ji})]$ (which itself is a 3-rd order U statistic). } Thus, in Section (ref) we resort to bootstrap for calculating the critical value as bootstrapping involves only repeated calculation of a 3-rd order U-statistic, $U_{n}$, which we find computationally easier than the aforementioned methods.
The global test introduced in this section only has power against global deviations from concavity/linearity/convexity and does not have power if the function is both convex and concave on different parts of the domain. Thus, it can be used as a first check -- for example, if the test rejects concavity, the researcher concludes that the function cannot be globally concave. Failure to reject would, then, mean that one has to consider our local test described in the next section, in order to verify if indeed the function is globally concave.
The main idea of the local test is to consider only triples of the kind portrayed in Figure (ref) local to a point $y$ in the domain of the transformation function. In other words, we will check if the transformation function is concave/linear/convex locally around $y$. Local concavity may be of interest by itself. However, usually we are interested in verifying if the treatment effect is concave on the whole domain, thus we will take the minimum of the local statistics at different points $y$ to run the test.
Formally, define:
where in practice the bandwidth used for $(Y_{i}-y)$ can be different than for $(X_{kj} - X_{ji})'\hat{\beta}$ but, in order to simplify exposition and mathematical arguments, we assume that both bandwidths are of the same order and denote both by $h$. Now our local test statistic for testing concavity is defined as:
where $\mathcal{Y}$ is a compact set, and the statistic for testing convexity is defined with $\sup$ replacing $\inf$ above. Linearity can be tested by replacing $U_{n}(y)$ with its absolute value. For the rest of the article we concentrate on $S_{n}^{conc}$ as results for testing convexity and linearity follow by very similar arguments.
Intuitively, low values of $S_{n}^{conc}$ show that there is a large deviation from concavity at some point $y$ and, hence, the treatment effects are not accelerating on the whole domain of the outcome.\footnote{Note that by the formulas in Example 1 concavity of $T$ is equivalent to convexity of the outcome in the treatment.} Therefore, the null hypothesis of concavity would be rejected if $S_{n}^{conc} < c_{\alpha}^{conc}$ where $c_{\alpha}^{conc}$ is an appropriate quantile from the asymptotic approximation to the distribution of our statistic.
In order to obtain $c_{\alpha}^{conc}$ one could imagine proceeding as in ghosal_et_al00: approximate the standardised U-statistic process $\sqrt{n} U_{n}(y)/\sigma_{n}(y)$, where $\sigma_{n}(y)$ is the estimator of the asymptotic variance, by a Gaussian process and then apply the extreme value theory in order to derive the distribution of the infimum. However, pursuing that approach is difficult in our setup for two reasons: 1) under $H_0$ we still have $E[U_n(y)]=O(h)$ rather than $o(h)$ therein, 2) as discussed above, estimating $\sigma_{n}(y)$ is computationally expensive. Instead we propose to use bootstrap to approximate $c_{\alpha}^{conc}$.
Define:
The following asymptotic approximation will be useful in justifying our bootstrap procedure.
An interesting consequence of this result is that the asymptotic distribution of the statistic does not depend on the estimation of $\beta_0$ as long as Assumption (ref)(ref) is satisfied, unlike the global test (cf. Theorem (ref)).
Our bootstrap procedure for obtaining the critical value for the global or local test is as follows:
Step two imposes the null hypothesis of linearity on the bootstrap sample so it “recentres” the bootstrap statistic on the linear case. Furthermore, sampling from a symmetric two point distribution in the wild bootstrap imposes symmetry of the error distribution in the bootstrap sample, thus implying that Assumption (ref) is satisfied in that sample. Note that the jacknife multiplier bootstrap (JMB) of chen_kato20 will not work for the local test here (without modifications) as we have $E[U_n(y)]=O(h)$ even under the null so the JMB statistic will not be centred correctly. We assume an equivalent of Assumption (ref)(ref) for the bootstrap estimator $\beta^{*}$:
This assumption is satisfied in our leading case, when $\beta^*$ is estimated by MRC (see subbotin07).
Theorem (ref) implies, for example, that we can reject global concavity of $T$, i.e. increasing treatment effect, when $S_{n}^{conc}<c_{\alpha}^{conc,*}$, and reject global linearity when $|S_{n}^{conc}|>c_{1-\alpha/2}^{conc,*}$. We prove validity of bootstrap in unconditional probability, which is a weaker result than the standard result conditional on the sample (see e.g. cavaliere_georgiev20 for discussion), as the fact that $E[U_n(y)] = O(h)$ under $H_0$ and the need to use parametric bootstrap complicates the analysis. Note that we only require the symmetry of the error distribution for the global test. The next theorem characterises alternatives for which our tests are consistent.
Assume that $c^{conc,*}_{\alpha}$ is generated using method (a) in step 2 above (see Appendix (ref) for the analysis with the wild bootstrap in (b)). Let $f_{xb}$ denote the pdf of $X_i'\beta_0$ and $\mathbf{1} = [1 \quad 1 \quad 1]'$. Further, let $f_{\boldsymbol{\varepsilon}}$ denote the joint distribution of $(\varepsilon_i,\varepsilon_j,\varepsilon_k)$ and $\nabla f_{\boldsymbol{\varepsilon}}$ denote its gradient and define $f_{\boldsymbol{\tilde{\varepsilon}}}$, $\nabla f_{\boldsymbol{\tilde{\varepsilon}}}$, similarly for $\tilde{\varepsilon} = Y -X'\beta_0$.
The condition in (ref) is technical and it is not straightforward to find general sufficient conditions for it to hold as it depends not only on the shape of $T$ but also the distributions of $\epsilon$ and $X$. However, we can make the following observations.
Finally, let us note that the computation of the local statistic involves evaluating a third order U-statistic at different points $y$ and, hence, takes significantly longer to compute than the global statistic. In practice, we recommend running the global test first and then proceed with the local test if the null hypothesis cannot be rejected in the first step.
The data is generated from the following five models:
where we draw $X$ and $\varepsilon$ from the standard normal distribution. Note that D0 is the worst-case model in the null hypothesis, D1 imposes concavity of the transformation, D3 convexity and D2 and D4 are neither concave or convex.
We run 1000 Monte Carlo replications. We use Gaussian kernel functions and rule-of-thumb bandwidths for both $(X_{kj}-X_{ji})'\hat{\beta}$ and $Y_{i}$, namely $h=1.06\hat{\sigma}n^{-1/5}$, where $\hat{\sigma}$ is a sample standard deviation of $(X_{kj}-X_{ji})'\hat{\beta}$ or $Y_{i}$.\footnote{The choice of bandwidth for $Y_i$ seems trickier as some transformations $T^{-1}$ make the distribution of $Y_i$ severly skewed. We have also tried using cross-validation (ucv in R) and plug-in cross-validation (bcv in R) with similar results.} In order to calculate the local statistic $S_{n}^{conc}$ we use a grid of values for $y$: -2:0.25:2, and take a minimum over the grid. The number of bootstrap replications used to calculate the critical value is 500 and we consider three sample sizes: $n=100, 250$ and $500$.
Table (ref) contains the results of the Monte Carlo simulations. We concentrate on testing concavity, as the results for the linearity and convexity tests, are very similar. The rejection probabilities in the linear case (D0) are close to the nominal level for both the global and the local test and both tests have perfect power against a globally convex alternative in D3.
As predicted, the global test has low power against D2 and D4, for which the transformation function is concave on part of the domain and convex on the other part. In this case deviations from concavity, as measured by our global statistic, cancel with positive values of the statistic obtained for the part of the domain where the function is concave, resulting in the global statistic taking values close to zero, just as for the linear case.
For designs D2 and D4, our local test significantly improves over the global test with almost perfect detection of non-concavity in D2 for $n = 250$. The power against D4 increases slower with $n$ than D2 but we still reject H0 with close to certainty with $n=500$. This is in line with the intuition that the local test will concentrate on the region of the largest violation of concavity instead of averaging over the measures of concavity for different regions.
Finally, note that, unlike for the global test, sampling from a symmetric distribution is not required in the local test bootstrap, so it is a question of finite sample performance if we would prefer the regular parametric bootstrap (part (a) in Step 2) or the symmetric wild parametric bootstrap in part (b). Comparing the last two panels in Table (ref) shows that both bootstrap tests have the right level but the symmetric wild bootstrap test has lower power for our designs (see D2 and D4).
We use data from karlan_zinman08, who ran randomised trials with a for-profit consumer lender in South Africa targeting high-risk consumer loan market. The lender randomised individual interest rate direct mail offers , conditional on the client’s risk category. karlan_zinman08 found that the loan demand curves are downward sloping. We will investigate if the demand curves are convex or, in other words, if increases in interest rate have a stronger negative effect on loan demand the lower the rate.
The dependent variable is the amount borrowed (in rands) at the offered interest rate.\footnote{We abstract from selectivity issues here. See karlan_zinman08 for discussion.} The experiment was ran in three mailer waves over four months, thus as in karlan_zinman08 we include the risk category and wave dummies as controls in $X$. The data contains 2325 observations. As the first step, we estimated the transformation function in our model using the estimator in chen02 and plotted it in Figure (ref) together with a smoothed version.\footnote{This is for illustrative purposes only. Note that the estimator in chen02 does not impose monotonicity, so a full exercise of estimating $T$ should include additional step of monotonising the estimate. We want to stress that our testing procedure does not rely on any estimator of $T$.} The figure suggests that the transformation is not far from being concave on most of its domain, maybe besides the values of loan size below 1000 rands, which implies convex demand curve by formulas in Example 1.
In order to formally test our conjectures, we first apply our global test to verify if the demand function is concave, linear or convex. As Table (ref) shows, global test rejects linearity and concavity of demand. Thus, in the second step we apply the local test to the null hypothesis of convexity and, in fact, reject the null hypothesis, concluding that the loan demand is not globally convex in the interest rate. To shed more light on where the non-convexity may come from, we re-ran the test on the sample excluding loans lower than 1000 rands and found that convexity is not rejected on this sub-sample. Therefore, overall we confirm that loan demand is mostly convex in interest rate, besides very small loans sector of the market.
Our application demonstrates usefulness of our testing procedures in recovering monotonicity of treatment effects. Particular appeal of the procedures described in this article comes from the fact that they avoid estimation of the transformation function and only require OLS estimation of the vector of coefficients $\beta_0$. Thus, they are relatively easy to implement. Additionally, computation of the third order U-statistic involved in our tests can be done efficiently by sorting the data by $Y$ first -- this reduces computational complexity to $O(nlog(n)+n(n-2)/2)$ from $O(n^{3})$ for a straightforward triple loop through the observations.