Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,715 characters · 16 sections · 70 citation commands
Nonparametric Estimation and Inference in Economic and Psychological Experiments
\address{Universit� dell'Insubria, Stony Brook University, Universit� Ca' Foscari}
The aim of this paper is to provide a statistical theory useful for the nonparametric analysis of laboratory experiments in economics and psychology.
In the typical experiment we have in mind, there are $n$ subjects who are administered $T$ tasks. Task $t$ is characterized by $X_{t}$, a $d$-dimensional stimuli-vector that is the same for each subject $i$, for $i=1,\dots,n$. The response or choice of subject $i$ in task $t$ is denoted by $Y_{it}$. We suppose that the Data Generating Process (DGP) of the set of answers $Y_{it}$ ($i=1,...,n$, $t=1,...,T$) has a deterministic component represented by a nonparametric function $f\left(X_{t}\right)$ of the stimuli, and a stochastic component $\varepsilon_{it}$ arising from individual error terms. Hence, we have the system:
The function $f:\mathcal{X}\mapsto\mathbb{R}$ can be interpreted as a deterministic theory that maps every value of the vector of stimuli $X$ to a space of real valued responses.
Equation (ref) resembles a framework already discussed in experimental economics by hey2005, which involves a nonstandard econometric system, and encompasses several models arising in economics and psychologica experiments. To the best of our knowledge, however, a complete statistical analysis of this system has yet to be conducted. For instance, the vector $X_{t}$ can represent the prizes and the probabilities of a lottery presented in an economic experiment in which subjects are asked to give the certainty equivalent of the lottery. In this case, $f\left(\cdot\right)$ is the function used by subjects to evaluate the lotteries.\footnote{The present approach can be equally applied, though it may not be the most efficient, to situations in which $X_{t}$ is defined as $X_{t}=\left(a_{t},b_{t}\right)$, for two lotteries $a_{t}$ and $b_{t}$ presented to subjects in pairwise choice experiments, and $Y_{it}$ is simply the choice (coded in some way) of subject $i$.} In psychophysical experiments, the vector $X_{t}$ can be thought to represent stimuli such as light, sound, weight, distance for which subjects are asked to assess in pairwise comparisons the relative magnitude. In this case, $f\left(\cdot\right)$ is the scale used by subjects to measure the stimuli. The fact that the explanatory variables are the same across individuals has a double justification: first of all, in many empirical studies the choice of the stimuli is so difficult that it is not possible to conduct it for each individual; second, when the regressors are the same across individuals the estimation problem is more difficult (in the sense that we have less information from the variation in the independent variables).
In statistics and econometrics, this model can be cast in the well-known and extensively studied framework of the nonparametric regression model liracine2007,tsybakov2009. There are, however, several distinctive features of this model that make its analysis different and, in some respects, more challenging than the standard nonparametric regression. First, the realizations of the stimuli are not random observations from an underlying statistical distribution, but are chosen by the experimenter: the more complex the function one wishes to estimate, the richer the support of the data one needs to achieve consistency. Second, the statistical approach proposed in this work does not impose any specific restrictions on the structure of the error terms. In particular, it seems important that even if (ref) holds true, the error terms for different individuals should be allowed to be differently distributed, provided $\mathbb{E}\left(\varepsilon_{it}\right)=0$. Among other things, this means that we allow different individuals to have different degrees of precision. This seems especially important in economics and psychology, when a researcher may have little to no knowledge about a theory that explains randomness in the responses. Also, individual variances may contain very persistent components, and, therefore, consistent estimation of an individual-specific response function may be unfeasible.\footnote{A similar statistical framework has been studied by stanis1998. While they also allow for the stimuli-vector to be time-varying, they suppose that the error term is a white noise. }
Our goal in this paper is to study the nonparametric estimation of the function $f\left(\cdot\right)$ in (ref) using the method of sieves newey1997,jong2002,chen2007,belloni2014,chenchristensen2013. The function of interest is approximated by a finite linear combination of some known basis functions (e.g., power series, regression splines, trigonometric polynomials), which effectively reduce the estimation problem to a finite number of parameters. The weights in the function approximation can be estimated through linear regression supposing that individuals and answers across individuals are independent. The number of approximating terms diverges to infinity with the sample size. We show that this estimator of the function $f\left(\cdot\right)$ is consistent, and we provide the convergence rate for this nonparametric estimator. We show that the convergence rate depends on the number of tasks ($T$), the number of individuals ($n$), and the number of basis functions used to approximate $f\left(\cdot\right)$. Our convergence rate, however, also depends on the properties of the average covariance matrix across individuals for a given task $t$. Heuristically, it thus explicitly takes into account the precision of the subjects in answering the questions and/or selecting specific choices.
We also provide asymptotic normality results for both linear and nonlinear functionals of the nonparametric estimator which are useful to obtain the asymptotic properties of the Wald test in this framework chen2012b. We derive the properties of the latter both in the case where the number of constraints is finite (parametric restrictions), which gives the standard $\chi^{2}$ distribution under the null; and when the number of constraints diverges to infinity along with other asymptotic parameters (a normal distribution). Lastly, we investigate what happens when the average variance matrix appearing in the previous tests is replaced by an estimator. We believe inference is an essential part of our statistical theory, as it allows us to test specific behavioral assumptions.
hey2005 points out that underlying system (ref) is the idea that the theory under investigation is deterministic, but that people apply the theory with noise. Such an approach, which is sometimes referred to as Thurstonian or Fechnerian, underlies for example the investigations conducted by falmagne1976, orme1994, buschena2000, and blavatskyy2007. Alternatively, other authors (including camerer1994, loomes1995, loomes2002, myung2005) tend to interpret individual behavior in experiments (and possibly in the real world) as inherently stochastic, in the sense that while the theories remain deterministic, their predictions are not because of the imprecision of people to know and to use the same specification of the theory every time it is required.\footnote{For example, describing the philosophy behind the approach with reference to preference theories, loomes2005 argues that the approach “rather than supposing an individual to have a single true preference function to which white noise is added, ... treats imprecision as if an individual's preference consist of a set of such functions. Thus to say that a particular individual behaves according to a certain `core' theory is to say that the individuals' preferences can be represented by some functions, all of which are consistent with the theory; but that on any particular occasion, the individual act as if she picks one of those functions at random from the set, applies it to the decision at hand, then replaces that function before picking again at random in order to address the next decision” (p. 306). Antecedents of this approach can be found in becker1963. It is also important to emphasize that this approach is still very different from the case in which the `core' theory itself would be made inherently stochastic --- as for example advocated by luce1997.} The distinction between the two approaches, however, though quite interesting philosophically, is of practical relevance only when either of the following two circumstances applies. The first is when the dependence of the answers $Y_{it}$ from the stimuli $X_{t}$ (according to function $f\left(\cdot\right)$) is parametric. In this case, the question of the two approaches turns into a fundamental question about whether the parameters to be estimated can be interpreted in deterministic terms or as random variables (as for example in the Bayesian approach pursued by karabatsos2005, and myung2005). The second applies to the specific restrictions on the errors terms which may be required by the statistical procedure used to analyze the experimental data (see for example ballinger1997, for the discussions of several restrictions often imposed for the empirical analyses of data from decision theory experiments).
The present statistical approach is unaffected by both circumstances, so that it can be viewed to encompass both philosophies. First of all, a notable feature of the present approach is that the dependence of the answers $Y_{it}$ from the stimuli $X_{t}$ is left nonparametric. Nonparametric dependence has the advantage that theoretical and/or behavioral properties of interest can be estimated and tested without the mediation of parametric restrictions, which (for the reasons just exposed) may not be unambiguously interpreted. Furthermore, nonparametric dependence is natural if one wishes to fit the experimental data without imposing any restrictions on behavior. Such `unrestricted' model could be a useful benchmark against which to compare any structural model indicated by specific theories.\footnote{See bernasconi2008,bernasconi2010a, for applications of such an approach in regards to experimental investigations of, respectively, psychophysical measurement theories and decisions theories.}
As mentioned above, we also allow for the possibility that the precision of answers of the same individual varies across different questions. This is important because various previous studies have emphasized forms of heterogeneity occurring both at levels of individuals and of different experimental tasks ballinger1997,buschena2000,carbone2000,blavatskyy2007,butler2007.
Finally, we should note that a long debated dispute in economic and psychological experiments is whether the analysis of the individual responses $Y_{it}$ should be conducted for the aggregate of the individuals or individual by individual. The analysis presented here is primarily thought for the former case. In particular, depending on the degree of heterogeneity and precision of the experimental sample and of the theory one would like to test, large values of $n$ and/or of $T$ may have different impacts on the consistency of our estimator of the function $f\left(\cdot\right)$ and its derivatives.\footnote{An additional reason to prefer an aggregate analysis is that an experimenter may decide to assign different values of the stimuli to different subjects. Assuming that the assignment mechanism is random, the aggregate model would allow to approximate the function $f\left(\cdot\right)$ on a richer support. In this situation, the convergence rates presented in this paper constitute a worst-case scenario.} If, however, one believes that the aggregate analysis cannot be carried forward because all individuals are characterized by different functions $f_{i}\left(X_{t}\right)$,\footnote{In psychology, the risks of averaging across individuals when they are characterized by different functions has been stressed, e.g., in skinner_reinforcement_1958, yost_variability_1981 and Bernasconi-Seri-JMP-2016.} then our results apply verbatim, simply taking $n=1$ and letting the number of tasks $T$ to diverge. In this case, our analysis can be seen as an extension of the results in newey1997 and belloni2014 to the case of deterministic regressors. Consistency is then guaranteed only under more stringent conditions on the variance of each individual error term.
The model in (ref) can also be interpreted as a panel data specification in which the covariates vary only with $t$. We do not pursue this interpretation further in this work, but we notice that su2012 have considered a panel data model with factor structure in the error term in which the function $f\left(\cdot\right)$ is allowed to vary across individuals, with $n,T\rightarrow\infty$. Notice, however, that time series data have a natural ordering that can be used in the asymptotic analysis, while our data do not possess such ordering. We show below that our convergence rates are amenable to some of their results, upon additional restrictions on the variance of the error term.
The present statistical approach is also suitable for various extensions which we indicate in the conclusions.
We recall that the data generating process is modeled as follows:
where $i=1,\dots,n$ denotes the individual (or respondent) and $t=1,\dots,T$ denotes a specific task. The dependent variable $Y\in\mathbb{R}$ and the vector of independent variables $X\in\mathcal{X}$, where $\mathcal{X}$ is taken to be a compact subset of $\mathbb{R}^{d}$.\footnote{Taking $\mathcal{X}$ to be compact does not appear to be a strong restriction in this setting, as points are chosen by the experimenter possibly within a bounded interval. In the development of our theory, this assumption could be relaxed by substantially modifying our method of proof chenchristensen2013.} In the following, we will suppose that the function $f\left(\cdot\right)$ belongs to a space $\mathcal{F}$ that will not be specified explicitly: when in the discussion of the results we will suppose that $f\left(\cdot\right)$ has at least $s$ continuous derivatives, it is intended that $\mathcal{F}$ will coincide with the Sobolev space $\mathcal{W}_{s,\infty}=\left\{ f:\left|f\right|_{s}<\infty\right\} $.
This statistical models fits several experimental and quasi-experimental frameworks.
We now rewrite the model in equation (ref) using matrix notations. We form the $T\times1$-vectors $Y_{i}=\left[Y_{i1},\dots,Y_{iT}\right]^{\prime}$ and $\varepsilon_{i}=\left[\varepsilon_{i1},\dots,\varepsilon_{iT}\right]^{\prime}$. We suppose that $\varepsilon_{i}$ has a distribution with mean $0$ and variance $\Sigma_{i}$, for every $i=1,\dots,n$. We further define the $\left(T\times d\right)$-matrix $\mathbf{X}$ obtained stacking the vectors $\lbrace X_{t},t=1,\dots,T\rbrace$. Finally, \[ Y_{i}=f\left(\mathbf{X}\right)+\varepsilon_{i},\quad\quad i=1,\dots,n, \] where the function $f\left(\cdot\right)$ is supposed to apply row-wise to the matrix $\mathbf{X}$. We make the following assumption about the vector of errors $\varepsilon_{i}$.
There are not noteworthy details in this assumption. We take the error terms to be uncorrelated across individuals $i$ and we impose some regularity conditions on the covariance matrix, which is otherwise left unspecified.
We now structure the statistical model for the whole data. We build the $\left(nT\times1\right)-$vectors $\mathbf{Y}$ and $\varepsilon$ by stacking respectively the $n$ vectors $Y_{1},\dots,Y_{n}$ and $\varepsilon_{1},\dots,\varepsilon_{n}$.
We finally have:
where $e$ is a $\left(n\times1\right)$-vector of ones. $\varepsilon$ is then a vector with mean $0$ and variance $\bigoplus_{i=1}^{n}\Sigma_{i}$, where $\bigoplus$ denotes the direct sum of matrices. That is, $\bigoplus_{i=1}^{n}\Sigma_{i}=\textrm{diag}\left(\Sigma_{1},\dots,\Sigma_{n}\right)$.
For estimation and inference, we take an approximation of $f\left(\cdot\right)$ using a linear combination of basis functions in $\mathcal{X}$. Thus, at $X_{t}=x$, we take \[ f_{P}\left(x\right)=\psi_{P}\left(x\right)\beta, \] where $\psi_{P}\left(x\right)=\left[\psi_{1,P}\left(x\right),\dots,\psi_{P,P}\left(x\right)\right]$ is a $1\times P$ vector of given basis functions and $\beta$ a $P\times1$ vector of unknown coefficients, with $P\rightarrow\infty$, with $n,T$. When $d$, the dimension of the vector of stimuli, is larger than $1$, then can be taken as a tensor product basis of total dimension $P$. We also denote as \[ \mathbf{\Psi}=
. \] the $\left(T\times P\right)$-matrix that stacks the approximating bases at every point $\lbrace X_{t},t=1,\dots,T\rbrace$. Then, we finally have:
where \[ U=e\otimes\left(f(\mathbf{X})-\mathbf{\Psi}\beta\right)+\varepsilon. \] The true value of the parameter $\beta$, which we denote $\beta_{0}$ is taken to satisfy \[ \mathbb{E}_{i}\left[\left(e\otimes\mathbf{\Psi}\right)^{\prime}\left(\mathbf{Y}-e\otimes\mathbf{\Psi}\beta_{0}\right)\right]=0, \] where the expectation is taken with respect to the distribution of $Y$ for each individual $i$. That is, \[ \beta_{0}=e\otimes\left[\mathbf{\Psi}^{\prime}\mathbf{\Psi}\right]^{-1}\mathbf{\Psi}^{\prime}\mathbb{E}_{i}\left[\mathbf{Y}\right]. \] The estimation of the model is performed by least squares, under the hypothesis that $\mathbb{E}(U)=0$. Our sieve-based least-squares estimator is therefore given by:
where $\otimes$ denotes a matrix Kronecker product and we denote \[ f_{P,0}(\mathbf{X})\triangleq\mathbf{\Psi}\beta_{0},\text{ and }\hat{f}_{P}(\mathbf{X})\triangleq\mathbf{\Psi}\hat{\beta}. \] In some instances, we may omit the argument of the function, and simply use the notations $\hat{f}_{P}$ and $f_{P,0}$. Notice that the estimator (and the model itself) could be simply written as \[ \hat{\beta}=\left(\mathbf{\Psi}^{\prime}\mathbf{\Psi}\right)^{-1}\mathbf{\Psi}^{\prime}\bar{\mathbf{Y}} \]
where $\bar{\mathbf{Y}}$ is a $T\times1$vector of average individual responses for a given task $t=1,\dots,T$. Our model could be then treated as any other nonparametric regression model with correlated errors. However, we believe this interpretation defies the entire purpose of our analysis. As a matter of fact, we would like emphasize the idea that the number of individuals performing the same task is essential when we cannot make standard regularity assumptions about the error term.
We first define the norms:
The norm $\left|\cdot\right|_{s}$ is what is sometimes written as $\left\Vert \cdot\right\Vert _{s,\infty}$. We do not need to consider more general weighted Sobolev norms gallant1987, since the functions we consider have nonstochastic arguments.
Throughout the paper we will also need the following quantities. For every $s\geq0$: \[ \zeta_{s}\left(P\right)\triangleq\max_{\left|\lambda\right|\le s}\sup_{x\in\mathcal{X}}\left\Vert \partial^{\lambda}\psi_{P}\left(x\right)\right\Vert _{F}, \] where $\Vert\cdot\Vert_{F}$ denotes the Frobenius norm. For every integer $s\geq0$, we define: \[ N_{P}\triangleq\left|f-f_{P,0}\right|_{s}. \] Denote $\lambda_{nT}\triangleq\lambda_{\max}\left(\bar{\Sigma}\right)$ and $\tau_{nT}\triangleq\mathrm{tr}\left(\bar{\Sigma}\right)$ to be the largest eigenvalue and the trace, respectively, of the average covariance matrix of the errors $\bar{\Sigma}\triangleq\frac{1}{n}\sum_{i=1}^{n}\Sigma_{i}$.
The following assumption is needed to derive, together with the previous definitions, a uniform upper bound for the sieve estimator $\hat{f}_{P}$.
The following assumption is needed to obtain consistency of the sieve estimator.
Assumption (ref) (i) restricts the asymptotic behavior of the matrix of design points $\mathbf{\Psi}^{\prime}\mathbf{\Psi}$. Notice that this assumption is not explicitly needed, e.g., in newey1997, jong2002 and belloni2014 as they deal with stochastic regressors and therefore this is guaranteed by an appeal to a Law of Large Numbers. In particular, Assumption (ref) implies that the eigenvalues of $\frac{\mathbf{\Psi}^{\prime}\mathbf{\Psi}}{T}$ converge to the eigenvalues of $Q_{P}$, for a fixed $P$.
newey1997 and belloni2014 derive rates of convergence in probability for $\frac{\mathbf{\Psi}^{\prime}\mathbf{\Psi}}{T}$ towards the fixed matrix $Q_{P}$. Newey's result implies that $\mathbb{E}\left\Vert \frac{\mathbf{\Psi}^{\prime}\mathbf{\Psi}}{T}-Q_{P}\right\Vert _{F}^{2}\le\frac{\zeta_{0}^{2}\left(P\right)P}{T}$. However, we cannot use directly this result since our regressors are supposed to be deterministic. However, reasoning as in reimer1997, we can see that for any probability measure $\inf_{\left\{ X_{t}\right\} _{t=1}^{T}}\left\Vert \frac{\mathbf{\Psi}^{\prime}\mathbf{\Psi}}{T}-Q_{P}\right\Vert _{F}^{2}\le\mathbb{E}\left\Vert \frac{\mathbf{\Psi}^{\prime}\mathbf{\Psi}}{T}-Q_{P}\right\Vert _{F}^{2}$, so that it is possible to find a point-set $\left\{ X_{t}\right\} _{t=1}^{T}$ such that $\left\Vert \frac{\mathbf{\Psi}^{\prime}\mathbf{\Psi}}{T}-Q_{P}\right\Vert _{F}\le\zeta_{0}\left(P\right)\sqrt{\frac{P}{T}}$.\footnote{Note that, from belloni2014 (see their Section 6.1 and Theorem 4.6), one can infer that $\left\Vert \frac{\mathbf{\Psi}^{\prime}\mathbf{\Psi}}{T}-Q_{P}\right\Vert _{F}\lesssim_{\mathbb{P}}\zeta_{0}\left(P\right)\cdot\sqrt{\frac{P\ln P}{T}}$ from which one gets the weaker bound $\inf_{\left\{ X_{t}\right\} _{t=1}^{T}}\left\Vert \frac{\mathbf{\Psi}^{\prime}\mathbf{\Psi}}{T}-Q_{P}\right\Vert _{F}^{2}\lesssim\zeta_{0}\left(P\right)\cdot\sqrt{\frac{P\ln P}{T}}$.} Better convergence rates can be obtained in special cases (see the discussion after Theorem (ref)).
Assumptions (ref) (ii) and (ref) (ii) are used to bound the approximation error and to define a uniform upper bound on the derivative of the vector of basis functions, as measured through the Sobolev norm, both of which are standard assumptions in the sieve literature. When the function $f\left(\cdot\right)$ is taken to be $c$ times continuously differentiable, we can take $N_{P}=O\left(P^{-c/d}\right)$ newey1997,huang2003,chen2007,belloni2014.
Assumption (ref) (i) restricts the behavior of the largest eigenvalue and the trace of the average covariance matrix of the errors $\bar{\Sigma}$. Depending on the structure of the covariance matrix, these quantities may or may not be uniformly bounded away from infinity.
Under the previous assumptions, it is possible to derive an upper bound for the convergence rate of $\hat{f}_{P}$ to $f$.
Assumption (ref) ensures that the upper bound of Theorem (ref) is $o\left(1\right)$ as $n,T\rightarrow\infty$. Notice that, if $n$ is taken to be finite and $\lambda_{nT}$ and $\tau_{nT}$ are uniformly bounded away from infinity, the upper bound of the variance in Theorem (ref) is the same as in newey1997.
The terms $\lambda_{nT}$ and $\tau_{nT}$ that enter the formulas have a behavior that can be clarified in some cases of interest.
A first case arises when the answers for a given individual are supposed to be uncorrelated to each other (even if heteroskedastic), so that $\Sigma_{i}=\textrm{diag}\left(\sigma_{i1}^{2},\dots,\sigma_{iT}^{2}\right)$. Define $\sigma_{nT}^{2}=\max_{1\le t\le T}\left(\frac{1}{n}\sum_{i=1}^{n}\sigma_{it}^{2}\right)$. In this case, $\lambda_{nT}=\sigma_{nT}^{2}$, and $\tau_{nT}\le T\sigma_{nT}^{2}$, so that the bound in Theorem (ref) yields: \[ \left|\hat{f}_{P}-f\right|_{s}=O_{\mathbb{P}}\left(\zeta_{s}\left(P\right)\left(\frac{P}{nT}\right)^{1/2}\sigma_{nT}+\zeta_{s}\left(P\right)N_{P}\right). \] If $\sigma_{nT}^{2}$ is bounded, the bound does not make any difference between $T$ and $n$. Assumption (ref) (i) requires that $T\rightarrow\infty$, so that also $n=1$ is sufficient to ensure consistency, provided that $P/T\rightarrow0$. In this case our estimator reduces to the one in newey1997. Despite our assumptions are not comparable with the ones in newey1997, as we consider the case of deterministic regressors, ours are generally weaker as we require uncorrelatedness and boundedness of the variances of the errors whereas he requires independence and identical distribution. Lack of dependence between tasks and boundedness of the maximum variance allow one to estimate consistently individual response functions. The hypothesis that the average covariance matrix of the errors $\bar{\Sigma}$ is diagonal could be in principle tested. However, we have not been able to find in the literature a test valid under our conditions, i.e. for possibly non identically distributed error vectors of increasing dimension. We leave the development of such a test to future work.
A second case of interest arises when the errors for every individual have a factor structure, i.e. every error term $\varepsilon_{it}$ can be written as $\varepsilon_{it}=\nu_{i}+\upsilon_{it}$ where $\mathbb{V}\left(\nu_{i}\right)=\sigma_{\nu}^{2}$, $\mathbb{V}\left(\upsilon_{it}\right)=\sigma_{\upsilon}^{2}$, $\textrm{Cov}\left(\upsilon_{it},\upsilon_{i^{\prime}t}\right)=0$ for $i\neq i^{\prime}$, and $\textrm{Cov}\left(\nu_{i},\upsilon_{it}\right)=0$, for all $i$. This means that the matrix $\Sigma_{i}$ can be written as \[ \Sigma_{i}=\sigma_{\nu}^{2}ee^{\prime}+\sigma_{\upsilon}^{2}I_{T},\text{ and }\bar{\Sigma}=\sigma_{\nu}^{2}ee^{\prime}+\sigma_{\upsilon}^{2}I_{T}. \] According to Lemma 2.1 in magnus1982, $\lambda_{nT}=\sigma_{\nu}^{2}T+\sigma_{\upsilon}^{2}$, and we obtain $\lambda_{nT}\asymp\tau_{nT}\asymp T$. The bound in Theorem (ref) implies that
\[ \left|\hat{f}_{P}-f\right|_{s}=O_{\mathbb{P}}\left(\zeta_{s}\left(P\right)\cdot\left(n^{-1/2}+N_{P}\right)\right). \] If the number of tasks $T$ is held fixed and the number of respondents is allowed to diverge to infinity, one could strengthen Assumption (ref) (i) to allow for the design matrix to be nonsingular for any finite value of $P$ and $T$ such that $P\leq T$. However, the finiteness of $P$ implies that the bias component does not disappear asymptotically and thus the estimator is not consistent. Similarly, if the number of individuals is held fixed and $T\rightarrow\infty$, while the bias vanishes, the variance does not disappear asymptotically, and again the estimator is not consistent. In our framework, however, letting both $n,T\rightarrow\infty$ yields a consistent estimator of $f\left(\cdot\right)$. Parametric rates of convergence can be achieved for $s=0,$ when $N_{P}=O\left(n^{-1/2}\right)$, i.e. when both $T$ and $P$ diverge sufficiently fast. Nonparametric rates for specific choices of the sieve basis are discussed in the Appendix.
To conclude this section, we briefly discuss the possibility of augmenting the model to include subject characteristics, which we denote by $Z_{i}$. In most experiments, these characteristics are inherently discrete (e.g., age, treatment group, gender, etc.). Assume for simplicity that the $j-$th element of the vector $Z_{i}$ can take values $\lbrace0,1,\dots,L_{j}\rbrace$ with strictly positive probability. Therefore, every element $Z_{ij}$ of $Z_{i}$ can be decomposed into $L_{j}$ dummy variables, each one taking value $1$ if $Z_{ij}=l$, and $0$ otherwise, for $l=1,\dots,L_{j}$. Without loss of generality, we can thus define $Z_{i}\in\lbrace0,1\rbrace^{q}$, where $q\geq0$ is a positive integer, and $Z_{i}$ can also include arbitrary interactions between the observed individual characteristics. For such a binary random vector, we impose that the joint function $f\left(X_{t},Z_{i}\right)$ is such that $f\left(X_{t},Z_{i}\right)=f_{0}\left(X_{t}\right)$, whenever all the elements of $Z_{i}$ are equal to $0$. Hence, we can write any joint function $f\left(X_{t},Z_{i}\right)=f_{0}\left(X_{t}\right)+Z_{i}^{\prime}f_{1}\left(X_{t}\right)$, which follows from the fact that its value changes in $Z_{ij}$ only when $Z_{ij}$ is equal to $1$, for $j=1,\dots,q$. This finally implies the following statistical model
\[ Y_{it}=f_{0}\left(X_{t}\right)+Z_{i}^{\prime}f_{1}\left(X_{t}\right)+\varepsilon_{it}, \] where the unknown functional coefficients depend on $X_{t}.$ This nonparametric regression can be cast as a varying coefficient model hastie1993,fan1999,fan2005. In this flexible semiparametric framework, we can include an arbitrary number of individual specific covariates without incurring the curse of dimensionality. Letting $\tilde{Z_{i}}=\left(1,Z_{i}\right)$, the vector $f=\left(f_{0},f_{1}\right)$ satisfies the following system of moment restrictions
\[ \mathbb{E}\left(\tilde{Z_{i}}Y_{it}\left|X_{t}=x\right.\right)=\mathbb{E}\left(\tilde{Z_{i}}\tilde{Z}_{i}^{\prime}\left|X_{t}=x\right.\right)f\left(x\right)=\mathbb{E}\left(\tilde{Z_{i}}\tilde{Z}_{i}^{\prime}\right)f\left(x\right), \] where the last equality follows from the fact that the vector $X_{t}$ should be determined independently of individual's characteristics, and therefore all the moments of the distribution of $Z_{i}$ are independent of $X_{t}$. For identification, we only require the additional condition that the matrix $\mathbb{E}\left(\tilde{Z_{i}}\tilde{Z}_{i}^{\prime}\right)$ is full rank. A nonparametric sieve estimator of the functional coefficients can be obtained by simply replacing the function $f$ with some finite dimensional approximation on a space of basis functions, and the unknown population moments of $Z_{i}$ with their sample counterpart. Estimation of this model is equivalent to splitting the sample in $2^{q}$ subsamples, and estimate the unknown regression functions for each one of these subsamples. However, joint estimation of the vector of coefficients is naturally more efficient, because it uses the entire sample size. Hence, the resulting sieve estimators inherit the same properties as above fanzhang2008.
In the following we investigate the asymptotic normality of functionals of our nonparametric estimator. These are useful to study the properties of classical statistical tests. We focus here on the properties of the Wald test and we provide its bias and the rates of convergence to its asymptotic distribution: we choose this strategy to be able to evaluate accurately the interplay between the different asymptotic parameters appearing below. We do not discuss whether it is possible to estimate these functionals at $\sqrt{nT}$-rate. Arguably, one could extend the results in newey1997 to our setting to provide such results.
In the following, we consider estimation of a functional of the function $f$. As in andrews1991, we write this functional as $\Gamma:\mathcal{F}\rightarrow\mathbb{R}^{R}$, where $R>0$ denotes the number of restrictions. Note that $\Gamma$ is allowed to depend on $n$ and $T$. We provide conditions for asymptotic normality of the quantity \[ W_{nT}=A_{nT}\left(\Gamma\left(\hat{f}_{P}\right)-\Gamma\left(f\right)\right) \] where $A_{nT}=V_{nT}^{-1/2}$ and $V_{nT}=\mathbb{V}\left(\Gamma\left(\hat{f}_{P}\right)\right)$. First of all we consider the case in which $\Gamma$ is linear, then we will move to the nonlinear case.
We consider both the case in which $R$ is fixed and the case in which $R$ diverges with $n$ and $T$. The case when $R$ is allowed to increase with $n$ and $T$ will be particularly useful to derive the asymptotic properties of Wald tests. We suppose that when applied to $f_{P}=\mathbf{\Psi}\beta$ the functional $\Gamma$ yields: \[ \Gamma\left(f_{P}\right)=\Gamma\left(\mathbf{\Psi}\beta\right)=\widetilde{\mathbf{\Psi}}\beta \] where $\widetilde{\mathbf{\Psi}}$ can take different values according to the linear functional $\Gamma$. A full list of examples is in andrews1991, but some very simple instances are the following ones:
All of these examples can also be considered in vector form, as in andrews1991.
The variance of the linear functional applied to an estimated function, namely $\Gamma\left(\hat{f}_{P}\right)$, is given by:
where $\overline{\Sigma}=n^{-1}\sum_{i=1}^{n}\Sigma_{i}$. Provided $V_{nT}$ is symmetric positive definite, let $A_{nT}=V_{nT}^{-1/2}$ be the symmetric positive definite square root of the inverse of $V_{nT}$.
We decompose $W_{nT}$ in two parts: an error term \[ \overline{W}_{nT}\triangleq W_{nT}-\mathbb{E}W_{nT}=A_{nT}\left(\Gamma\left(\hat{f}_{P}\right)-\mathbb{E}\Gamma\left(\hat{f}_{P}\right)\right)=A_{nT}\widetilde{\mathbf{\Psi}}\left(\widehat{\beta}-\mathbb{E}\widehat{\beta}\right) \] and a bias term \[ \mathbb{E}W_{nT}=A_{nT}\left(\widetilde{\mathbf{\Psi}}\mathbb{E}\widehat{\beta}-\Gamma\left(f\right)\right). \] We provide conditions under which the first term converges in distribution to a $R-$dimensional standard normal vector and the second term tends to a null vector. The vector $W_{nT}$ enters in the formula of the Wald test for the hypothesis $\mathsf{H}_{0}:\Gamma(f)=\Gamma_{0}$:
In the following we will consider a standardized version of the statistic, namely $W=\frac{W^{\ast}-R}{\sqrt{2R}}$. This is particularly useful when the number of constraints $R$ is allowed to increase with $n$ and $T$, as we show below.
We make the following assumptions.
Now we pass to consider asymptotic normality of nonlinear functionals, also denoted as $\Gamma:\mathcal{F}\rightarrow\mathbb{R}^{R}$. Our treatment extends the one in Theorem 2 in newey1997 to the multivariate case. We suppose that $\Gamma$ possesses a directional derivative $\Gamma^{\prime}$ enjoying the properties stated in Assumption (ref) (ii). When the second argument of $\Gamma^{\prime}$ is the true function $f$, we simply write $\Gamma^{\prime}\left(\cdot,f\right)=\Gamma^{\prime}\left(\cdot\right)$. In case of nonlinear functionals, we replace Assumption (ref) (ii) with the following.
The variance $V_{nT}$ and its square root $A_{nT}$ entering the statement of the theorem have exactly the same definition as above, with $\widetilde{\mathbf{\Psi}}$ defined as in Assumption (ref) (i). In this case the decomposition of $W_{nT}$ into an error and a bias term has a parallel in the decomposition into a linear part and a remainder part (that is not strictly speaking a bias since it is random and depends upon $\widehat{\beta}$). The linear part provides the asymptotic distribution and the remainder has to be bounded adequately.
In this section we provide some results about the effect of replacing $\overline{\Sigma}$ with an estimator $\widehat{\overline{\Sigma}}$. Wald-type tests require the replacement to occur in the test statistics but have standard asymptotic distributions: in this case, we provide a bound on the absolute error of the replacement in $\overline{W}_{nT}^{\prime}\overline{W}_{nT}$.
In some cases, it can be of interest to estimate the conditional variance as a function of observable characteristics of the stimuli and/or the individuals butler2007. We will investigate this alternative procedure rather informally. Suppose that the following model holds: \[ \mathbb{V}\left(\varepsilon_{it}\right)=\mathbb{E}\varepsilon_{it}^{2}=h\left(\tilde{X}_{it}\right),\quad\quad i=1,\dots,n,\quad t=1,\dots,T, \] where the vector $\tilde{X}_{it}$ shares some of the features of $X_{t}$ and it may also coincides with it. The latter implies, among other things, that the variance is the same across individuals but different across questions.
\[ \widehat{U}_{it}^{2}=h\left(\widetilde{X}_{it}\right)+\nu_{it},\quad\quad i=1,\dots,n,\quad t=1,\dots,T. \] If the objective is to estimate the entire structure of the matrix $\overline{\Sigma}$ as a function $\tilde{X_{it}}$, we are left with the problem of estimating covariances. A potential solution is to suppose that errors are equicorrelated, in which case the correlations can be estimated from the standardized residuals. We do not pursue this topic here.
We now apply our estimation procedure to two simple examples introduced above in Economics and Psychology. In both examples, we estimate the unknown function nonparametrically, and then we use the Wald statistics to test meaningful restrictions either on the function itself or on its derivatives.
The function is approximated using tensor products of Legendre polynomials, which make the estimation and testing procedures straightforward and intuitive. We denote as $L_{j}\left(x\right)$, the Legendre polynomial of order $j$, with $j=0,1,2,\dots$. The order of the polynomial is chosen by cross-validation hansen2012.
We now use the results in Examples (ref), (ref), (ref), (ref) and (ref) to analyze the data of an economic experiment. We employ a classical experimental design to elicit the preference of an individual in choice under uncertainty. The elicitation procedure is known as the probability equivalence method and is dual to the certainty equivalence method. It works as follows. In a sequence of pairwise comparisons $t$ for $t=1,\dots,T$, an individual is asked to state the probability $p_{t}$ that would make the individual indifferent between receiving the sure amount of money $CE_{t}$ or a lottery giving the monetary prize $X_{t}$ with probability $p_{t}$ and $0$ otherwise.
Thus, in the probability equivalence method $CE_{t}$ and $X_{t}$ are the stimuli and $p_{t}$ the response, whereas in the certainty equivalence method $X_{t}$ and $p_{t}$ are the stimuli and $CE_{t}$ the response. We also remark that, though both methods are in principle capable to elicit the preferences of the individual, since their early applications it is known that both methods can lead to various types of inconsistencies. For this reasons, various proposals of revising the basic elicitation procedures have been made in the literature in attempt to control for the inconsistencies (references and discussion in, e.g., hershey1985probability,wakker1996eliciting,abdellaoui2000parameter).
Both elicitation procedures are nevertheless very simple and are useful for the purpose of illustrating our nonparametric method of estimation and inference. We in particular conducted the probability equivalence experiment with 98 participants ($n=98$). Each of them gave the responses $p_{t}$'s to 100 questions ($T=100$). The 100 questions employed monetary prizes $X_{t}$ distributed uniformly between a minimum of 15 Euro and a maximum of 66 Euro, with the sure amount $CE_{t}$ varying between a minimum of 4 Euro and a maximum of 57 Euro. The experiment was run individually and was computerized: questions were presented sequentially to each individual on a computer screen, with the order of the questions randomized independently for each participant. Each participant had the opportunity to reconsider the decision to each individual questions several times before confirming it; but once confirmed, the computer moved the participant to another question and previous choices couldn't be any longer revised or accessed. At the end of each individual experiment one question was randomly selected for each participant and each participant was paid according to the choice he or she made in the selected question. In particular, participants were incentivized to give correct answers by the use of the standard becker1964measuring payment method.
A summary of the distribution of $X_{t}$, $CE_{t}$, and the sample mean of $p_{t}$ across individuals is reported in the Table (ref).
With these data, we first estimate the following fully nonparametric regression model: \[ p_{it}=f\left(\ln\left(CE_{t}\right),\ln\left(X_{t}\right)\right)+\varepsilon_{it}, \] using a tensor product of polynomials of order $2$ in $\ln\left(CE_{t}\right)$ and polynomials of order $3$ in $\ln\left(X_{t}\right)$. Figure (ref) depicts the nonparametric estimator of the regression function $g$.
To implement the testing procedure, we slightly undersmooth compared to the estimation above, and take a cubic polynomial in $\ln CE_{t}$ and a quartic polynomial in $\ln X_{t}$ (see Theorem (ref)). Therefore, we have a total of $R=18$ restrictions, but only 16 of them are linearly independent. As $K=3$ and $J=4$, these restrictions are: \[ \beta_{1,3}=\beta_{2,2}=\beta_{2,3}=\beta_{3,1}=\beta_{3,2}=\beta_{3,3}=\beta_{4,0}=\beta_{4,1}=\beta_{4,2}=\beta_{4,3}=0 \] and: \[ \beta_{0,1}+\beta_{1,0}=\beta_{0,2}-\beta_{2,0}=\beta_{0,3}+\beta_{3,0}=\beta_{1,2}+\beta_{2,1}=3\beta_{0,2}+\beta_{1,1}=5\beta_{0,3}+\beta_{1,2}=0. \] The value of the test statistic is $557.865$, which leads to a rejection of the null hypothesis, with a $p-$value strictly lower than $0.01$. This implies that the usual power specification of the utility function may not be justified in our setting, and such an assumption could lead to an inconsistent estimation of the probability weighting function.
Now we consider an experimental application related to the model in Examples (ref), (ref), (ref), (ref), (ref) and (ref). The only difference with respect to the examples is that we use Legendre polynomials, but this has no impact on the conditions. In the experiment, $n=69$ individuals were asked to give their estimates of ratios of distances of pairs of Italian cities from a reference city. Participants were undergraduate students in economics from the University of Insubria in Italy. We presented to the subjects 10 pairs of Italian cities and we asked them to estimate the ratio of their distances with respect to Milan: the $T=10$ pairs were given by all the possible combinations out of the five cities Turin, Venice, Rome, Naples and Palermo. The range of the stimuli goes from 124 to 885 km and the range of the real distance ratios from 2 to 7.137. We asked participants, first, to state for each comparison which of the two items they thought was larger, and then to quantify the relative dominance of the two items, i.e. how many times the city that they considered more distant from Milan was, according to them, actually more distant from Milan than the city they considered less distant. All the experiments were performed in a random order. The data were already analyzed, with different aims and techniques, in Bernasconi-Choirat-Seri-MS-2010,Bernasconi-Choirat-Seri-JMP-2011,Bernasconi-Choirat-Seri-EJOR-2014.
A summary of the magnitude of the stimuli and of the ratio reported by the individual is given in Table (ref).
A nonparametric series estimator of the function $f$ of Example (ref) is reported in Figure (ref). We take quadratic polynomials in both $\ln X_{1t}$ and $\ln X_{2t}$.
We test Stevens' model (see Examples (ref) and (ref)), in which $f$ is a linear function of the difference of the log between the two stimuli. That is, our null hypothesis is: \[ H_{0}:f\left(\ln X_{1},\ln X_{2}\right)=\kappa\cdot\left(\ln X_{1}-\ln X_{2}\right), \] for some real $\kappa$. The test for this hypothesis has been considered in Examples (ref), (ref) and (ref). Under the null hypothesis, the intercept of the model is equal to zero; the two slopes sum up to zero; and all other higher-order coefficients are equal to zero. We therefore tests $R=8$ restrictions on the vector of estimated parameters. We remark that the number of individuals $n$ is much larger than the number of questions $T$, thus providing some support for the conditions in Example (ref).
In this case, we do not undersmooth to implement our testing procedure, since the quadratic tensor product basis already exhausts the degrees of freedom in our model. The value of the test statistic is $1781.218$, which leads to a rejection of the null hypothesis, with a $p-$value strictly lower than $0.01$.
We present in this paper a set of statistical tools useful for the analysis of some experimental data when the researcher aims at estimating and testing the average individual response function nonparametrically. In lack of a theory of errors in either economic or psychophysical experiments, we allow our regression errors to be correlated within individual tasks, and to feature heteroskedasticity between different individuals. In particular, and differently from a large body of literature on nonparametric regressions, we do not assume that variance of the error term is uniformly bounded away from infinity. This approach makes the tools we suggest robust to several structures of the error covariance matrix, for which theory often provides scarce or contradictory information. We finally point out that our results can be considered a worst-case scenario. That is, we only consider the case when the same tasks are submitted to all individuals. We conjecture that better asymptotic properties could be obtained if one allows for different individuals to perform different tasks. We defer the analysis of such a case to further research.