Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
24,793 characters · 9 sections · 20 citation commands
Testing the Drift-Diffusion Model
The {\em drift diffusion model \ \/}(DDM) is a model of sequential sampling with diffusion (Brownian) signals, where the decision maker accumulates evidence until the process hits a stopping boundary, and then stops and chooses the alternative that corresponds to that boundary. This model has been widely used in psychology, neuroeconomics, and neuroscience to explain the observed patterns of choice and response times in a range of binary choice decision problems. One class of papers study \textquotedblleft perception tasks\textquotedblright with an objectively correct answer (e.g. \textquotedblleft are more of the dots on the screen moving left or moving right?\textquotedblright; here the drift of the process is related to which choice is objectively correct Ratcliff-review,ShadlenKiani13. The other class of papers study \textquotedblleft consumption tasks\textquotedblright\ such as \textquotedblleft which of these snacks would you rather eat?\textquotedblright; here the drift is related to the relative appeal of the alternatives FehrRangel,Roeetal,ClitheroRangel13,Krajbichetal10,Krajbichetal11,Krajbichetal12,Milosavljevicetal10,Krbahafe15,Reutskajaetal11 .
The simplest version of the DDM\ assumes that the stopping boundaries are constant over time Wald47,Stone60,Edwards65,Ratcliff78. More recently a number of papers use non-constant boundaries to better fit the data, and in particular the observed correlation between response times and choice accuracy, i.e., that correct responses are faster than error responses Luce86,Milosavljevic10,Drugowitsch12,FSS2018.
Constant stopping boundaries is the optimal solution for perception tasks where the volatility of the signals and the flow cost of sampling are both constant, and the prior belief is that the drift of the diffusion has only two possible values, depending on which decision is correct. Even with constant volatility and costs, non-constant boundaries are optimal for other priors. FSS2018 characterize the optimal boundaries for the consumption task: the decision maker is uncertain about the utility of each choice, with independent normal priors on the value of each option. Drugowitsch12 show how to computationally derive the optimal boundaries for the perception task: the signal coherence varies from trial to trial, so some decision problems are harder than others.
This paper provides a statistical test for DDM's with general boundaries. We first prove a characterization theorem: we find a condition on choice probabilities that is satisfied if and only if the choice probabilities are generated by some DDM. Moreover, we show that the drift and the boundary are uniquely identified. We then use our condition to nonparametrically estimate the drift and the boundary and construct a test statistic based on finite samples.
Recent related work on DDM includes Drugowitsch12 who conducted a Bayesian estimation of a collapsing boundary model and FSS2018 who conducted a maximum likelihood estimation. hawkins2015revisiting estimate collapsing boundaries in a parametric class, allowing for a random nondecision time at the start. Chiongetal18 estimate a version of DDM with constant boundaries but random starting point of the signal accumulation process; Ratcliff02 estimates a similar model where other parameters are made random. Baldassi partially characterize DDM with constant boundary.\footnote{ They ignore the issue of correlation between response times choices by looking only at marginal distributions, which makes their conditions necessary but not sufficient.}
Other work on DDM-like models includes the decision field theory of BusemeyerTownsend,BT93,DFT allows the signal process to be mean-reverting. Alosetal18 and ES17 study models where response time is a deterministic function of the utility difference. HebertWoodford,Woodford14,CheMierendorff,Liangetal,LiangMu,Zhong19 study dynamic costly optimal information acquisition.
Let $X$ be the universe of alternatives (actions) and $T=\mathbb{R}_{+}$ be time. For every pair of objects $\{x, y\}$ the analyst observes pairwise stochastic choices and decision times. In the limit as the sample size grows large, the analyst will have access to the joint distribution over which object is chosen and at which time a choice is made. We denote by $F^{xy}(t)$ the probability that the agent makes a choice by time $t$, and let $ p^{xy}(t) $ the probability that the agent picks $x$ conditional on stopping at time $t $. Throughout, we restrict attention to cases where $F$ has full support and no atoms at time $0$, so that $F(0)=0$, and we assume that $F$ is strictly increasing with $\lim_{t\rightarrow \infty} F(t)=1$. These restrictions imply the agent never stops immediately, that there is a positive probability of stopping in every time interval, and that the agent always eventually stops. We call $(p^{xy},F^{xy})$, the {\em stochastic choice function\/}.
An immediate restriction on the stochastic choice function is that the choices of the agent are unaffected by which object we consider to be the first and which object we consider to be the second. This is formally equivalent to
Without loss of generality we only consider stochastic choice functions which satisfy this restriction. We also assume that each option is chosen with positive probability $0<p^{xy}(t)<1$ for all $t$.
Given $(p^{xy},F^{xy})$ we define the choice imbalance at each time $t$ to be
This is the Kullback-Leibler divergence (or relative entropy) between the Binomial distribution of the agent's time $t$ choice $P(t)=(p(t),1-p(t))$ and the permuted choice distribution $Q(t)=(1-p(t),p(t))$. As the Kullback-Leibler divergence is a statistical measure of the similarity between distributions $I(t)$ captures the imbalance of the agent's choice at time $t$. Note that $I=0$ means that both choices are equally likely; $ I=\infty $ when $p$ equals $0$ or $1$, and that $I$ is symmetric about 0.5. We define $\bar{I}^{xy}$ to be the average choice imbalance,
and we define $\bar{T}^{xy}$ to be the average decision time,
and define $\bar{p}^{xy}$ to be the average choice probability$,$
and assume that all of these integrals exist. Finally, we relabel objects as needed so that the first object is chosen weakly more often, i.e. $\bar{p} ^{xy}\geq 0.5$ for all $x,y$.
The drift diffusion model (DDM) is commonly used to explain the stochastic choice data in neuroscience and psychology. The two main ingredients of a DDM are the stimulus process $Z_{t}$ and a time-dependent stopping boundary $ b(t)$. In the DDM representation, the stimulus process $Z_{t}$ is a Brownian motion with drift $\delta $ and volatility $\alpha $:
where $B_{t}$ is a standard Brownian motion, so in particular $Z_{0}=0$. Define the hitting time $\tau $
i.e., the first time the absolute value of the process $Z_{t}$ hits the boundary $b$. Let $F^{\ast }(t;\delta ,b,\alpha ):=\mathbb{P}\left[ \tau \leq t\right]$ be the distribution of the stopping time $\tau $. Likewise, let $p^{\ast }(t;\delta ,b,\alpha )$ be the conditional choice probability induced by (ref) and (ref) and a decision rule that chooses $ x$ if $Z_{\tau }=b(\tau )$ and $y$ if $Z_{\tau }=-b(\tau )$.
Our goal in this paper is to determine which data is consistent with a DDM representation, and when it is, when the representation is unique. When the drift $\delta=0$, each alternative will be chosen half of the time regardless of the shape of the boundary, so we will exclude this case going forward.
The original formulation of the DDM was for \textquotedblleft perception tasks\textquotedblright\ where the drift $\delta $ is either $+1$ or $-1$ depending on which decision is correct; more generally there can be a distinct drift $\delta ^{xy}$ for each pair $x,y$. In consumption-choice problems (otherwise known as value-based problems, see, e.g., Milosavljevic10) it is natural to assume that the net drift $\delta ^{xy}$ is the difference between two signals, an $x$-signal with drift $u(x)$ equal to the utility of $x$ and a $y$-signal with drift $u(y)$ equal to the utility of $y$, so that $\delta ^{xy}=u(x)-u(y).$ This imposes some consistency conditions that we discuss below.
Note that this definition requires that the data from all of the menus $ \{x,y\}$ is generated with the {\em same\/} boundary function $b$. This corresponds to cases where the agent treats each decision problem as a random draw from a fixed environment.\footnote{ In an optimal stopping model, the shape of the boundary is determined by the agent's prior over these draws.} We are interested in characterizing which stochastic choice functions admits a DDM representation. The following result follows immediately from rescaling $\delta $ and $b$.
We will thus without loss of generality only consider the DDM model where we normalized $\alpha =1$. We write $p^{\ast }(t,\delta ,b)$ and $F^{\ast }(t,\delta ,b)$ as short-hands for $p^{\ast }(t,\delta ,b,1)$ and $F^{\ast }(t,\delta ,b,1)$.
Given a stochastic choice function $(p^{xy},F^{xy})$, define the {\em revealed drift\/}
When the revealed drift is is non zero, we define the {\em revealed boundary \/} as
The revealed drift is high for a pair $x,y$ whenever the agent either makes very imbalanced choices or decides quickly, and low for choices that are slow and close to 50-50. Over time the boundary at time $t$ follows the log-odds ratio of the agent's choice at time $t$ which is zero whenever the agent's choice is balanced and and increases in the imbalance of the agent's choice. The revealed boundary is smaller for pairs with a larger revealed drift. In the knife-edge case when the revealed drift is 0 the revealed boundary is not defined and our results do not apply.
Theorem (ref) below says that if the true data generating process is a DDM, then the revealed drift and boundary will exactly match the true parameters. Moreover, Theorem (ref) allows us to test whether the true data generating process is indeed a DDM.
Our first result characterizes the DDM for a fixed pair $x, y\in X$.
Thus, the stochastic choice function $(p^{xy},F^{xy})$ is consistent with DDM whenever the observed distribution of stopping times $F^{xy}$ equals to the distribution of hitting times generated by the revealed drift $\tilde{ \delta}^{xy}$ and revealed boundary $\tilde{b}^{xy}$. Theorem (ref) shows that the revealed drift and boundary are the unique candidate for a DDM representation. It thus allows us to identify the parameters of the DDM model directly from choice data. This permits the model to be calibrated to the data without computing the likelihood function, which requires computationally costly Monte-Carlo simulations. More substantially, as Theorem (ref) connects the primitives of the model directly to data it allows us to better understand their economic meaning. The drift in the DDM model is a measure of how imbalanced and quick the agent's choices are and the shape of the boundary follows the imbalance of the agent's choices over time. We hope that this interpretation makes the empirical content of the parameters of DDM model more transparent and the model thus more useful.
Note that this theorem shows that the distribution of stopping times contains additional information that is not captured by the mean. For example, choice data where $p^{xy}(t)$ and $\bar{T}^{xy}$ are any 2 given constants is only consistent with one possible distribution of stopping times $F^{xy}$\ \ However a test based only on the mean choice probability and mean stopping time will accept any model that matches those two numbers, and in particular regardless of $F^{xy}$ the data is consistent with a constant stopping boundary. (See Baldassi).
Our next result extends the characterization to all pairs $x, y\in X$.
Thus, in addition to satisfying the condition from Theorem 1 pairwise, we have two additional consistency conditions imposed across pairs. Condition (ii) follows from our assumption that the agent uses the same stopping boundary in every menu. Condition (iii) comes from the assumption that the drift in a given menu depends on the difference of utilities, that is $ \delta ^{xy}=u(x)-u(y)$.\footnote{ The proof of the theorem follows from Theorem (ref) and the Sincov functional equation, see, e.g., Aczel66.}
The idea for the test is based on Theorem (ref), which requires that the observed distribution of stopping times matches the distribution induced by the revealed boundary $\tilde{b}$ and drift $\tilde{d}$. We first describe a nonparametric estimator of $\tilde{b}$ and $\tilde{\delta}$ based on a finite data set. Next, we show how to test the distribution matching condition. This test could be extended to multiple-alternatives settings along the lines of Theorem (ref), but we do not do so here.
Suppose that we have a fixed pair $x,y\in X$. Define
Each data point consists of the time $\tau _{i}$ at which the choice is made and the choice $\gamma _{i}$ made at time $\tau _{i}.$
The unknown features of the DDM model are the drift $\delta $ and the boundary $b(t).$\ We use estimators based on equations (ref) and (ref) that identify the revealed drift and boundary. Both of them depend on the choice probability, so we first give an estimator of that. Here $p^{xy}\left( t\right) :=\Pr (\gamma _{i}=1|\tau _{i}=t)$ is the probability of choice $x$ conditional on the choice being made at $t$.
The nonparametric estimator we construct is a spline regression: that is, a least squares regression of $\gamma _{i}$ on approximating functions of $ \tau _{i}.$ For simplicity, we use a linear probability estimator of $ p^{xy}(t)$.\footnote{ We reserve consideration of other estimators of the choice probability to future work, including logit or probit with a series approximation inside the logit or probit CDF.}
We first transform $\tau _{i}$ to the unit interval.\footnote{ In DDM models where $b(t)$ does not reach zero, there is no uniform bound on realized decision times $\tau _{i}$. Because $\tau _{i}$ is the conditioning variable (i.e. regressor) in the choice probability, it is important to allow for an unbounded regressor.} For this purpose let $G(t)$ be a CDF of a positive random variable with PDF that is positive on $(0,\infty ).$ Consider
Because $G_{i}$ lies in the unit interval we can use standard series estimation to estimate $p^{xy}(t)$. We consider regression spline estimation of $p^{xy}(t)$. For this purpose let
be a $B$-spline vector, say for evenly spaced knots on $(0,1)$. Let $\hat{ \beta}$ be OLS coefficients from regressing $\gamma _{i}$ on $ q_{i}^{K}=q^{K}\left( G_{i}\right) $. The choice probability estimator we consider is
We give conditions for this estimator to be consistent and have other important large sample properties in Assumptions 2 and 3 to follow.
We can estimate the drift $\delta $ by plugging in $\hat{p}(t)$ for $ p^{xy}(t)$ in formula (ref) and replacing expectations with sample averages. Let
The estimator of $\delta $ is then
The estimator of the boundary $b\left( t\right) $ is obtained by plugging in $\hat{\delta}$ and $\hat{p}(t)$ in the expression of equation (ref), giving
We now have to test whether the observed distribution of stopping times matches the one induced by the revealed drift and boundary. We do this by comparing sample moments of functions of the decision time with estimators of the moments that predicted by the model. To describe such a test let $ m_{J}(\tau )=(m_{1J}(\tau ),...,m_{JJ}(\tau ))^{\prime }$ be a vector of functions of $\tau $. Examples include indicator functions for intervals and B-splines in $G(\tau ).$ The sample average vector will be $\bar{m} =\sum_{i=1}^{n}m_{J}(\tau _{i})/n$.\footnote{ The Kolmogorov--Smirnoff test uses indicator functions but instead of the the average of $m$ it takes the supremum. The Cramer--von Mises test takes the sum of squares. We look at the average of $m$ because the target cdf we are comparing with is not fixed, but involves estimates of the boundary and drift, see Newey94.} We use simulation to obtain model prediction. To describe the simulated predictions, let $\{B_{t}^{1},...,B_{t}^{S}\}$ be $S$ independent copies of Browning motion and
A moment vector predicted by the model would be $\hat{m}_{S}= \sum_{s=1}^{S}m_{J}(\hat{\tau}_{s})/S.$ A test of the model can be based on comparing $\bar{m}$ and $\hat{m}.$ Let $\hat{V}$ be a consistent estimator of the asymptotic variance of $\sqrt{n}(\bar{m}-\hat{m}_{S})$ when the model is correctly specified, as we will describe below. A test statistic can be formed as
The model would be rejected if $\hat{A}$ exceeds the critical value of a $ \chi ^{2}(J)$ distribution. If $J$ is allowed to grow with $n$ and the $ m_{J}(\tau )$ is allowed to grow in dimension and richness as $n$ grows then this approach will test all the restrictions implied by DDM as $n$ grows. In Appendix A we describe the construction of $\hat{V}.$
In formulating conditions for the asymptotic distribution of this test we will let $m_{jJ}(\tau ),$ $(j=1,...,J)$ be indicator functions for disjoint intervals. Let $\tau _{jJ}=G^{-1}(j/(J+1)),$ $(j=0,...,J),$ $\tau _{J+1,J}=\infty .$ Consider
The test based on these functions will be based on comparing empirical probabilities of intervals with those predicted by the model. The normalization of multiplying by $\sqrt{J+1}$ is convenient in making the second moment of these functions of the same magnitude for different values of $J$. Note that we have left out the indicator for the interval $ (0,1/(J+1)).$ We have done this to account for the fact that the estimator the drift parameter uses some information about $\tau _{i}$, so that we are not able to test all of the implications of the DDM for the distribution of $ \tau _{i}.$ As usual we can only test overidentifying restrictions.
We derive results under the following conditions:
This assumption is equivalent to the ratio of the pdf of $\tau _{i}$ to $ dG(t)/dt$ being bounded and bounded away from zero. It is straightforward to weaken this condition to allow it to only hold on compact, connected interval that is a subset of $(0,1),$ if we assume the $b(t)$ is constant on known intervals near $0$ and where $\tau $ is large.
We also make a smoothness assumption on the boundary function.
This condition requires that the derivatives of $b(t)$ go to zero in the tails of the distribution of $\tau _{i}$ as fast as the pdf of $G(t)$ does. We also require that the drift parameter be nonzero.
We need to add other conditions about the smoothness of CDF of $\tau _{i}$ as a function of the drift $\delta $ and the boundary and about rates of growth of $J$ and $K$. The involve much notation, so we state them in Assumption (ref) in Appendix (ref).
We can now state the following result on the limiting distribution of $\hat{A }.$