Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,168 characters · 27 sections · 82 citation commands
How Flexible is that Functional Form? Quantifying the Restrictiveness of Theories
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\inprod#1{\left\langle #1\right\rangle } \global\long\def\ol#1{\overline{#1}} \global\long\def\ul#1{#1} \global\long\def\td#1{\tilde{#1}} \global\long\global\long\global\long\global\long\global\long \sloppy
\thispagestyle{empty}
\setcounter{page}{1}
If a parametric model fits the data well, is it because the model captures structure specific to the observed data, or because the model is so flexible that it would fit almost all conceivable data? This paper provides a quantitative measure of restrictiveness that can distinguish between these two explanations, and is easy to compute in a variety of applications.
Our approach for evaluating the restrictiveness of a model is to generate synthetic data sets, and evaluate how well the model fits this synthetic data. Some models have known properties, for example Cumulative Prospect Theory requires that certainty equivalents for lotteries respect first-order stochastic dominance. For these models, the relevant question may not be whether the model is restrictive at all, but instead how much content it has beyond these known constraints. We define the eligible data to be those data sets that satisfy specified background constraints, and measure a model's restrictiveness by its (normalized) average error across the eligible data.
We complement the evaluation of restrictiveness, which is based solely on synthetic data, with an evaluation of the model's performance on actual data, using the completeness measure proposed in FKLM. Restrictiveness and completeness provide complementary perspectives, and define a Pareto frontier where models that rule out more regularities, yet capture the regularities that are present in real data, are preferred.\footnote{These are not the only considerations that matter for evaluating models, and we do not speak to other important concerns such as parameter estimation and causal inference. Nevertheless, these two measures may be relevant to those problems as well: If a model can fit almost any data set, then its good fit to a specific real data set does not necessarily mean that the model is the “right" model.}
Section (ref) provides axioms for our restrictiveness measure to clarify its theoretical properties. The main axioms require that the measure is homogeneous in the unit scale used to quantify model error, and that the measure has a linearity property as the background constraints are varied. An additional “symmetry" axiom requires that the model's ability to approximate different synthetic data sets has the same effect on the restrictiveness measure. Dropping this axiom returns a broader class of restrictiveness measures, where instead of averaging across synthetic data sets, the data sets are weighted by an analyst's prior. We develop estimators for both the restrictiveness and completeness measures in Section (ref), and establish their asymptotic properties so that users can compute confidence intervals.
A key feature of our restrictiveness measure is that is computable without the guidance of theoretical results about the model's implications or empirical content. This differentiates restrictiveness from measures such as the model's VC dimension, or its hit-rate and accuracy-rate as defined in Selten.\footnote{There are representation theorems for many non-parametric theories of individual choice, and some analytic results for the sets of equilibria in games, but we are unaware of representation theorems for most functional forms that are commonly used in applied work.} (Section (ref) reviews the related literature and relates it to our work.) The measure's tractability makes it easy to apply to a variety of contexts, as we demonstrate by applying it to models from three economic domains: (1) predicting certainty equivalents for binary lotteries (where we evaluate Cumulative Prospect Theory and Disappointment Aversion); (2) predicting initial play in matrix games (where we evaluate the Poisson Cognitive Hierarchy Model (PCHM), Logit PCHM, and Logit Level-1); and (3) predicting takeup of microfinance in Indian villages (where we evaluate linear regression models based on economically-motivated regressors, and a structural model of diffusion).\footnote{In addition to these applications, Schwaninger uses our restrictiveness measure to evaluate models of bargaining with inequity aversion, EllisKarizOzbay uses it to evaluate models of consumer demand from budget sets, and ba2023over uses it to evaluate models of reaction to information.} The first two settings use data from the lab, our third application uses field data. In each of these domains, these measures reveal new insights about the models we examine, which we now summarize:
Application 1: Certainty Equivalents. We evaluate two models on a set of binary lotteries from Bruhin: a popular three-parameter specification of Cumulative Prospect Theory CPT, henceforth CPT, and a two-parameter specification of Disappointment Aversion gul1991, henceforth DA. We find that CPT performs strikingly well on the Bruhin data, achieving a completeness of 95%, while DA's completeness is only 27%.
One explanation for this finding is that CPT is a much better model of risk preferences than DA. Another possibility is that CPT is simply more flexible. We thus evaluate the restrictiveness of the two models, where our background constraints are that the synthetic average certainty equivalents must lie within the range of the lotteries' possible payoffs, and must respect first-order stochastic dominance (FOSD). We find that CPT is indeed substantially less restrictive than DA: CPT performs better than DA not only on the real data set but also on the other eligible data sets. This tells us that FOSD constitutes a large part of the empirical content of CPT on the domain of binary lotteries, while DA imposes substantial additional restrictions.\footnote{DA's low completeness suggests that these restrictions are not supported by the experimental data.}
Besides comparing distinct models such as CPT and DA, restrictiveness and completeness can be compared across nested models to reveal the role played by specific parameters. Adding a parameter always at least weakly increases completeness and decreases restrictiveness, but some parameters achieve greater improvements in completeness for the same decrease in restrictiveness. We find that several parameters lead to large drops in restrictiveness in return for only marginal improvements in completeness, suggesting that these parameters may add flexibility in the wrong directions. The CPT parameter that governs the curvature of the probability weighting function, however, achieves a large improvement in completeness compared to the flexibility it adds, so this parameter seems to capture an important part of risk preferences. Indeed, it is the curvature of the probability weighting function that has played a key role in many of the applications of CPT to financial data (e.g., barberis2008stocks and green2012initial).
Application 2: Initial Play in Games. Next, we evaluate three models on a set of $3 \times 3$ matrix games from FudenbergLiang: the Poisson Cognitive Hierarchy Model, or PCHM CamererHoChong04; Logit PCHM LeytonBrownWright, which allows for logistic best replies in the PCHM; and Logit Level-1, which models the distribution of play as a logistic best reply to the uniform distribution. We impose the background constraint that strictly dominant actions are played at least as often as if by chance (i.e. with probability at least $1/3$) and that strictly dominated actions are played with probability no more than $1/3.$ We find that all three models are highly restrictive relative to these constraints, which shows that the constraints on the frequency of strictly dominated and strictly dominant strategies are a very small part of their empirical content. The restrictiveness of Logit PCHM and Logit Level-1 is nearly identical, although Logit PCHM has two parameters while Logit Level-1 has one.
Application 3: Diffusion on a Social Network. Finally, we consider the prediction of microfinance takeup rates in the set of Indian villages studied by banerjee2013diffusion, banerjee2019using, and compare the performance of OLS regression on various economically-motivated regressors with that of an economically-motivated partially linear model built upon “network gossip centrality.” Here we find that the partially linear model is dominated by a simple OLS model based on the average eigenvector centrality of leaders: the latter has higher restrictiveness and higher completeness.
Besides these specific findings about each of these economic domains, our analyses make the high-level point that it is not sufficient to count parameters to understand a model's restrictiveness. Even with just 3 parameters, CPT is not very restrictive on the domain of binary lotteries, and models with different numbers of parameters (such as Logit PCHM and Logit Level-1) turn out to be similarly restrictive. These comparisons are not obvious from the functional forms, but are easy to discover with our restrictiveness measure.
Before formally defining our measure, we use a simple example to illustrate it. Suppose there is a binary covariate $x \in \{x_0,x_1\}$ and an outcome variable $y \in [0,1]$. A data set is an observed outcome for each covariate value, i.e., a point in $\mathbb{R}^2$, and the eligible data $\mathcal{F}$ is collection of possible data sets, i.e. a subset of $\mathbb{R}^2$. A model is also a subset of $\mathbb{R}^2$. The model explains a data set exactly if the data set is an element of the model.
Figure (ref) considers eligible data $[0,1]^{2}$ and depicts three models. Model A includes all of $[0,1]^{2}$, and thus can exactly explain any (eligible) data set. Model B includes all data sets $(y_0,y_1)$ satisfying $y_1 > y_0$, and so can only explain data sets where the outcome is higher at covariate $x_1$ than at $x_0$. Model C discretizes the data into a grid and includes every other element of the grid.
One way of evaluating the restrictiveness of these models is the fraction of eligible data sets that they can fit exactly Selten. But evaluating restrictiveness in this way obscures important differences between models such as B and C. Both exactly explain 50% of the data yet Model B appears to impose a more substantive restriction.
Our restrictiveness measure instead takes as given a measure of how well a model approximates the data. This makes it computationally straightforward to estimate in applications even when we have very little analytical guidance about the model's predictions (in contrast to the Selten measure, which requires determining exact fit). To estimate restrictiveness, we uniformly sample over all eligible data, evaluate the model's average approximation error to the realized datasets, and compare it to average approximation error of a benchmark model. For example, if we use Euclidean distance as our measure of approximation error (as in Figure (ref)), and the constant model $\{(1/2,1/2)\}$ as the benchmark, the restrictiveness of Model B is numerically estimated to be about $0.30$, while the restrictiveness of Model C is approximately $0.02$. Thus Model B is substantially more restrictive by our measure.
Section (ref) formally defines our measure of restrictiveness. Section (ref) reviews the measure of completeness from FKLM. Section (ref) combines these concepts with the idea of a Pareto frontier of models that are undominated in completeness and restrictiveness. Section (ref) further discusses the interpretation of our restrictiveness measure. Section (ref) describes the relationship to the literature.
Our starting point is a data set of observations of $(X,Y)$, where $X$ is a covariate vector and $Y \in \mathcal{Y}$ is an outcome, with $\mathcal{Y}$ a compact subset of a finite-dimensional Euclidean space. We use $\mathcal{X}$ to denote the set of covariate vectors, and $P_X$ to denote the marginal distribution of $X$. We assume that $\mathcal{X}$ is finite, and $P_X$ is chosen by or known to the researcher.\footnote{In laboratory experiments the set of features and their relative frequencies are chosen by the experimenter while in field experiments these are chosen by Nature, but in either case we treat them as known.} A prediction rule is a function $f: \mathcal{X} \rightarrow \mathcal{Y}$. We denote the set of all such functions by $\ol{\mathcal{F}} \equiv \mathcal{Y}^{|\mathcal{X}|}$, and endow it with the usual topology.
We take as a primitive a discrepancy function $d: \ol{\mathcal{F}} \times \ol{\mathcal{F}} \rightarrow \mathbb{R}_+$ where $d(f,f')$ measures how different the two prediction rules $f$ and $f'$ are. For example, if $Y$ is a vector in $\mathbb{R}^{n}$, a natural choice for $d$ is the expected mean-squared distance between the predictions (with respect to $P_X$), and if $Y$ is a distribution a natural choice for $d$ is the expected KL-divergence (again with respect to $P_X$). We allow for functions $d$ that are not distances (such as KL-divergence), but require that $d(f,f')=0$ if and only if $f=f'$. We also assume that $d$ is uniformly bounded, and that $d(\cdot,f)$ and $d(f,\cdot)$ are continuous almost everywhere for each $f\in\ol{{\cal F}}$.\footnote{Given that ${\cal Y}$ is assumed to be bounded, the uniform boundedness of $d$ is a very weak requirement. The only reason that we allow for discontinuity in $d$ is to accommodate the case of $\mathbf{1}\{f=f'\},$ the discrepancy function used in Selten. We recommend in Appendix B that practitioners use a continuous discrepancy function $d$.}
We will evaluate the restrictiveness of a parametric model $\mathcal{F}_{\Theta}:=\{f_{\theta}\}_{\theta\in\Theta}\subseteq \overline{\mathcal{F}}$, where the prediction rules $f_{\theta}$ depend continuously on a parameter $\theta$ from a compact set $\Theta$.\footnote{Because $\mathcal{X}$ is assumed to be finite, $\Theta$ can viewed as a subset of a finite-dimensional Euclidean space without loss of generality.} Restrictiveness is defined relative to a compact set of “eligible” rules $\mathcal{F} \subseteq \ol{\mathcal{F}}$ that reflect any constraints the model is known to have. For example, if a model is known to imply that choices respect first-order stochastic dominance, we can define $\mathcal{F}$ to be all rules with this property, and measure the model's additional restrictiveness beyond this. In general, the eligible set $\mathcal{F}$ consists of all prediction rules that satisfy user-specified background constraints, where the special case of $\mathcal{F} = \ol{\mathcal{F}}$ corresponds to the question of whether $\mathcal{F}_{\Theta}$ imposes any restrictions at all.
We define the restrictiveness of a model to be its expected discrepancy to a prediction rule $f$ drawn uniformly at random from the eligible set, normalized with respect to the expected discrepancy of a baseline prediction rule $f_{\text{base}}$. The baseline prediction rule is chosen to suit the setting, and we interpret its performance as a lower bound that any sensible model should outperform.\footnote{For example, in our application to predicting initial play in games, we define the baseline prediction rule to be a uniform distribution over actions. Note that while the choice of baseline affects the value of restrictiveness, it does not affect the comparative restrictiveness of two models on the same domain.}
Normalizing with respect to a baseline has several advantages: First, it makes our measure invariant to affine rescalings of the units of discrepancy. Second, whenever $f_{\text{base}}$ is chosen from $\mathcal{F}_{\Theta}$, restrictiveness ranges from $0$ to 1. A model with $r=0$ is completely unrestrictive, while a model with $r=1$ fits synthetic data no better than the baseline prediction rule does. If a model performs well on real data and is also highly restrictive, then its good performance occurs not simply because the model can fit any data, but because it precisely identifies regularities in real behavior.
The ratio in ((ref)) is well-defined as long as the denominator exceed zero, so we will impose this an assumption going forward:
Section (ref) provides axioms for the restrictiveness measure, which help to clarify the measure's theoretical properties.
While restrictive models are desirable holding all else equal, a restrictive model is not useful if it poorly fits real data. To evaluate model fit to real data, we use the completeness measure introduced in FKLM. This takes as a primitive a loss function $l: \mathcal{Y} \times \mathcal{Y} \rightarrow \mathbb{R}_+$, which is assumed to be continuous. Let $P_{Y|X}$ denote the distribution of $Y$ given $X$, and $P:=(P_X, P_{Y|X})$ denote the joint distribution of $X$ and $Y$. The prediction rule that minimizes expected loss on the real data is given by \[f^* \in \operatorname*{arg\,min}_{f \in \overline{\mathcal{F}}} e_{P}(f)\] where \[e_{P}(f):=\mathbb{E}_{P}\left[l(f(X),Y)\right] \quad \forall f\in \ol{\cal F}.\] For example, if $\mathcal{X}$ is a set of lotteries, $\mathcal{Y}$ are subjects' reported certainty equivalents for each lottery, and $l$ is squared error, then $f^*$ takes each lottery into its average certainty equivalent across subjects. If $\mathcal{X}$ is a set of payoff matrices, $\mathcal{Y}$ is the set of distributions over actions, and $l(Y,Y')$ is Kullback-Leibler divergence from $Y'$ to $Y$, then $f^*$ maps each game to the corresponding distribution over actions.
By construction, $\kappa$ lies within the unit interval. A model with $\kappa=1$ matches the true $f^{*}$ exactly, while a model with $\kappa=0$ is no better at matching $f^{*}$ than the baseline prediction rule $f_{\text{base}}$. In the special case where discrepancy is the expected mean-squared distance $d(f,f') = \mathbb{E}_{P_X}[(f(X)-f'(X))^2]$ and the baseline prediction rule is constant at the expectation of $Y$, $f_{\text{base}} = \mathbb{E}_{P}[Y]$, completeness specializes to the familiar (population) definition of $R^2$, but completeness is applicable more generally.
We report both restrictiveness $r$ and completeness $\kappa$ for each application that we consider. Completeness is defined using the loss function $l$, while restrictiveness is defined using the discrepancy function $d$. When the discrepancy function $d$ and the loss function $l$ are “paired” in the sense of Online Appendix (ref),\footnote{Loosely speaking, being paired means that $d(f,f^*)$ is the difference between the error of $f$ and the error of the best mapping $f^*$.} then $\kappa(\mathcal{F}_{\Theta}) = 1 - r(\mathcal{F}_{\Theta},\overline{\mathcal{F}})$, so that completeness is the complement of the restrictiveness of model $\mathcal{F}_{\Theta}$ with respect to the (unconstrained) eligible set $\overline{\mathcal{F}}$. Our first and third application use mean-squared error as the loss function and expected squared distance as the discrepancy function; our second application uses negative log-likelihood as the loss function and expected KL divergence as the discrepancy function. Both are examples of paired functions.
Our restrictiveness and completeness measures generate a “Pareto frontier” consisting of models that are undominated in the sense that none of the other models considered are simultaneously more restrictive and more complete. Although this is a very partial order, it has bite in our Application (ref) (see Figure (ref)), as well as in the work of EllisKarizOzbay.
Unlike in typical economic problems, the Pareto frontier here need not be concave, so the preferred model may not maximize a weighted sum of the two scores. For example, the frontier might consist of 3 points with scores (3/4,1/4), (1/3,1/3), and (1/4,3/4), and the analyst might prefer the model with scores $1/3$ each. Of course, given the estimated parameter values of two models on the actual and hypothetical data sets, one could make predictions by taking pointwise combinations of the two model's predictions, which would mechanically lead to a weakly concave frontier of undominated models, but it seems hard to interpret this exercise.
While it is natural to prefer undominated models to dominated ones, it is less obvious how to aggregate the two measures to pick a preferred model, as the tradeoff between the measures is context-specific and also a matter of taste. Nevertheless, when two models have completeness-restrictiveness values that cannot be Pareto-ranked, one can consider the size of the improvement in completeness relative to the size of the reduction in restrictiveness. In Section (ref) we show that adding an “elevation” parameter to a Cumulative Prospect Theory specification leads to a large drop in restrictiveness in return for only a small gain in completeness, while the parameter that governs the curvature of the probability weighting function leads to a sizeable improvement in completeness with only a small reduction in restrictiveness. We take this to mean that the curvature parameter plays a more important role in capturing risk preferences.\footnote{\citet*{ba2023over} conduct a similar exercise to compare two models which are not Pareto-ranked.}
\paragraph{Context dependence.} Restrictiveness is context-specific, in the sense that it depends on the set of feature vectors $\mathcal{X}$ and the outcome to be predicted. For example, we show that the restrictiveness of Cumulative Prospect Theory depends on the support size of the lotteries that are considered. Evaluating the restrictiveness of a model across contexts can reveal that it is very restrictive for one kind of prediction problem but unrestrictive for others. An interesting direction for followup work would be to develop a measure of restrictiveness that takes into account how restrictive a model is across different contexts. For example, we might consider one model to be “generally more restrictive” than a second model if the distribution of restrictiveness values for the first model first-order stochastically dominates the distribution for the latter, as we find in Section (ref).
\paragraph{Choosing the eligible set.} The restrictiveness of a model is measured with respect to a specific eligible set $\mathcal{F} \subseteq \overline{\mathcal{F}}$, which is chosen based on what is known about the model. In Application 3, we investigate the restrictiveness of a structural model of network diffusion for predicting takeup of microfinance. Since there is relatively little known about the empirical content of this model, we define the eligible set to include all possible takeup rates, and study whether the model placed any restrictions at all. In contrast, the model of interest in Application 1, Cumulative Prospect Theory, implies that any lottery that first order stochastically dominates another must have a higher certainty equivalent. So we place this restriction on the eligible set, and see how much additional restrictiveness the model imposes.
In general, there is not a single correct choice of eligible set. While we focus on comparing the restrictiveness of models with respect to a given eligible set, an interesting complementary exercise is to fix a model and compare its restrictiveness relative to different eligible sets, as in Sections (ref) and (ref).
\paragraph{Why the uniform distribution?} Section (ref), which develops and axiomatizes a broader class of restrictiveness measures, provides an axiom that pins down the uniform distribution. Besides this axiom, there are many reasons to prefer the uniform distribution. First, once the eligible set is specified, the uniform distribution on this set is pinned down (under our assumptions that $\mathcal{X}$ is finite and $\mathcal{Y}$ is a subset of finite-dimensional Euclidean space). This reduces the number of primitives to be chosen, and helps prevent cherry-picking with respect to the distribution on $\mathcal{F}$. Second, the uniform distribution is computationally easy to implement, even for eligible sets $\mathcal{F}$ with potentially complicated structures.\footnote{For example, in our application to prediction of certainty equivalents, we build monotonicity with respect to FOSD into our definition of $\mathcal{F}$, and it is straightforward to sample uniformly from $\mathcal{F}$ by first sampling from a larger space without the monotonicity constraints, and then only keeping the draws that satisfy the monotonicity constraints. In contrast, non-uniform weightings over $\mathcal{F}$ require additional specification of how exactly $\mathcal{F}$ is parametrized, making the dependence of restrictiveness on $\mathcal{F}$ less transparent.} Finally, our use of the uniform distribution follows up on Becker's proposal of the uniform distribution over budget-exhausting bundles as a model of irrational consumer behavior, and parallels Selten's use of area (see Section (ref)).
\paragraph{Why are more restrictive models better?}
Our paper takes the perspective that restrictiveness is inherently desirable: if two models have the same level of predictive accuracy, we should prefer the one that imposes more restrictions to the more flexible alternative. A potential reason for this preference is that models are often meant to capture behavior in related but not-identical domains. Given enough data, models that are very unrestrictive will fit any specific data set well, but may do so by learning idiosyncratic details of those datasets that do not in fact transfer across settings. In contrast, if a highly specific and structured model happens to fit a data set well, this may generate more confidence that the model's structure extends to other settings.\footnote{AFLW compare the transfer performance of highly flexible black box models with less flexible economic models in a setting similar to our Application 1, and find that the black box models transfer more poorly.}
Our restrictiveness measure generalizes the notion of “observational restrictiveness" introduced in Koopmans, where a model is observationally restrictive if the distributions permitted by the model are a proper subset of the distributions that would otherwise be possible.\footnote{As Koopmans points out, a special case of an observationally restrictive specification is an overidentifying restriction. See e.g. sargan1958estimation, hausman1978specification, hansen1982large, and chen2018overidentification for econometric tests of overidentification.} A model that is not observationally restrictive can perfectly match all data and so has $r=0$. Our restrictiveness measure allows us to quantify just how restrictive a model is.
Section 2 already discussed Selten's measure of flexibility, and showed how its use of exact instead of approximate fit can lead to very different conclusions than ours. The Selten measure has been applied by BeattyCrawford, Hey, and HarlessCamerer, and BlowBrowningCrawford among others, to understand the restrictiveness of nonparametric economic models. It is typically difficult to determine whether a parametric model can exactly fit a given data set without the guidance of prior analytical results, while our measure is easy to compute in a variety of applications.\footnote{BeattyCrawford analytically derives the set of budget shares that are consistent with GARP, and HarlessCamerer uses results about generalized expected utility theories to determine whether choices between specially chosen pairs of lotteries (for example, lotteries sharing a common ratio of outcome probabilities) are consistent with those theories. But we do not know how to analytically determine the predictions that are consistent with PCHM or the structural model of microfinance takeup in Application 3.}
In considering approximate rather than exact fit, our approach is related to papers that measure the distribution of the Afriat index Choietal,Polisson.\footnote{Choietal and Polisson relax the implications of expected utility maximization using Afriat's “efficiency index” as an analog of our loss function. They compare the distribution of the efficiency indices of the actual subjects with its counterpart in randomly generated data.} These approaches are motivated by the testing of rationality of choices; our aim here is to show that similar techniques can be applied to a substantially broader class of models. BeattyCrawford propose an alternative “smoothed out" version of Selten's measure for the revealed preference setting that resembles restrictiveness, except that it does not allow for restrictions on the eligible data and normalizes by reference to a worst case.\footnote{Another approach for model selection that does not require exact fit is ClippelRozen's suggestion to select models by comparing the ratio of the likelihood of observing the real data under the specified model to the likelihood under a uniform distribution over all possible models.}
Our use of synthetic data to evaluate restrictiveness is similar to the use of simulated data to evaluate the power of a hypothesis test, as in Bronars and andreoni2013power. Their power measures are based on particular specifications of the alternative hypothesis, while we focus on an aggregate measure over a class of “alternative hypotheses.” Moreover, because our objective is to measure the content of a model's restrictions and not hypothesis testing, we use approximate rather than exact fit.
Our measure is related to various measures from computer science, statistics, and econometrics, but differs in a few key ways. First, compared to classic measures for the complexity of function classes, such as VC dimension, Rademacher complexity, and metric entropy, our measure can be computed without analytical results about the empirical content of the estimated model.
Second, compared to measures such as empirical Rademacher complexity, AIC, and BIC, which are often used for model selection, our restrictiveness measure does not depend on the observed data and is not indexed to sample size.\footnote{We could loosely interpret our restrictiveness measure as analogous to a limiting case of Rademacher complexity for large samples, where we use the discrepancy function $d$, rather than correlation, to measure the model's ability to fit the synthetic data.} This reflects a difference in objectives: A primary goal of model selection is to avoid overfitting a complex model to a finite (and small) quantity of data, while our objective is to provide a measure of restrictiveness that does not depend on the quantity of data used to estimate it.\footnote{Specifically, our measure does not depend on the number of observations $(x,y)$ in the data or on the values of the $y$'s, though it does depend on the feature set $\mathcal{X}$.} Relatedly, while previous metrics aggregate a notion of completeness with some notion of restrictiveness,\footnote{For example, the AIC combines the log-likelihood, which is about fitness to real data (corresponding to “completeness") and the number of parameters, which is about the flexibility of the model without reference to real data (corresponding to “restrictiveness") in an additive way} we trace the associated Pareto frontier (see Section (ref)).
This section provides an axiomatixation for the un-normalized version of the restrictiveness measure (i.e., the numerator of ((ref))), which we call approximation error. Readers primarily interested in applications of the measure can skip ahead to the next section.
We endow the set $\overline{\mathcal{F}}$ with the Lebesgue $\sigma$-algebra and a $\sigma$-finite measure $\mu$, which can be interpreted as the analyst's prior. An approximation error $e$ takes as input the model $\mathcal{F}_{\Theta}\subseteq \overline{\mathcal{F}}$, a compact set of eligible prediction rules $\mathcal{F}\subseteq \overline{\mathcal{F}}$, and a discrepancy function $d$. The quantity $e(\mathcal{F}_{\Theta},\mathcal{F},d)$ is interpreted as the approximation error of the model $\mathcal{F}_{\Theta}$ to the eligible set $\mathcal{F}$, where the quality of the approximation is measured using $d$. We would like for this approximation error function to satisfy the following axioms. First, approximation error should always be nonnegative.
Second, if one model is better able to approximate every eligible prediction rule than another, the first model has lower approximation error.
Third, any linear rescaling of the units of $d$ is inherited by the approximation error, and a linear rescaling of the discrepancy between a model $\mathcal{F}_{\Theta}$ to each prediction rule $f$ leads to the same value of approximation error as rescaling the units of the discrepancy $d$.
Fourth, consider constraining the set of eligible prediction rules $\mathcal{F}$ to a subset $\mathcal{F}_1$ or its complement $\mathcal{F}_2$. The ex post approximation errors of a model $\mathcal{F}_{\Theta}$ with respect to either of these new eligible sets is, respectively, $e(\mathcal{F}_{\Theta},\mathcal{F}_1, d)$ or $e(\mathcal{F}_{\Theta},\mathcal{F}_2, d)$. The subsequent axiom says that the ex ante approximation error $e(\mathcal{F}_\Theta,\mathcal{F},d)$ is a convex combination of the ex post approximation errors, where each ex post subset contributes to the ex ante approximation error in proportion to its measure.
Finally, permuting the various discrepancies between the model and the eligible prediction rules $f$ does not affect the overall approximation error. This reflects a “principle of indifference" over the eligible prediction rules.
Our restrictiveness measure assumes ((ref)), and normalizes the approximation error of model $\mathcal{F}$ relative to the approximation error of the baseline $f_{\text{base}}$.
We now discuss how to implement our approach in practice. Recall that we restrict $\mathcal{X}$ to be finite, so $\overline{\mathcal{F}}$ is finite-dimensional.
\paragraph{Computing Restrictiveness}
The following is an algorithm for computing $r$: Sample $M$ times independently from a uniform distribution on the eligible set $\mathcal{F}$. For each sampled $f_m\in{\cal F}$, compute $d(\mathcal{F}_{\Theta},f_m)$ and $d(f_{\text{base}},f_m)$. Then $$\hat{r}_M := \frac{\frac{1}{M}\sum_{m=1}^{M}d(\mathcal{F}_{\Theta},f_m)}{\frac{1}{M}\sum_{m=1}^{M}d(f_{\text{base}},f_m)}$$ is an estimator for restrictiveness $r = r(\mathcal{F}_{\Theta},\mathcal{F})$. In principle, the number of simulations we run, $M$, can be arbitrarily large, so $\hat{r}_{M}$ can be made arbitrarily close to $r$. Moreover, it is straightforward to obtain the formula for the asymptotic standard error of the simulated $r$, based on which confidence intervals can be constructed.\footnote{Under Assumption (ref), $\sqrt{M}\left(\hat{r}_{M}-r\right)/\hat{\sigma}_{\hat{r}}\overset{d}{\longrightarrow}{\cal N}\left(0,1\right)$, where the asymptotic variance estimator $\hat{\sigma}_{\hat{r}}^2$ is defined by $\hat{\sigma}_{\hat{r}}^{2}:=\left[\hat{\sigma}_{{\cal G}}^{2}-2\hat{r}\hat{\sigma}_{{\cal G},f_{\text{base}}}+\hat{r}^{2}\hat{\sigma}_{f_{\text{base}}}^{2}\right]/\left[\left(\frac{1}{M}\sum_{m=1}^{M}d(f_{\text{base}},f_m)\right)^{2}\right]$, with $\hat{\sigma}_{{\cal G}}^{2}$ being the sample variance of $d({\cal G},f_m)$, $\hat{\sigma}_{f_{\text{base}}}^{2}$ the sample variance of $d(f_{\text{base}},f_m)$, and $\hat{\sigma}_{{\cal G},f_{\text{base}}}^{2}$ the sample covariance of $d({\cal G},f_m)$ and $d(f_{\text{base}},f_m)$, across $m=1,...,M$. We note that the standard error here simply measures the approximation error of $r$ based on a finite number of simulations and do not reflect randomness in experimental data.}
\paragraph{Estimating Completeness}
Suppose that the analyst has access to a finite sample of data $\left\{ Z_{i}:=\left(X_{i},Y_{i}\right)\right\} _{i=1}^{N}$ drawn from the unknown true distribution $P^{*}$. To estimate completeness, which is defined based on the loss function $l$ introduced in Section (ref), we use $K$-fold cross-validation to estimate the out-of-sample prediction error of the model.
(Our applications make the standard choice of $K=10$.) Specifically, we randomly divide ${\bf Z}_{N}=(Z_1, \dots, Z_N)$ into $K$ (approximately) equal-sized groups. To simplify notation, assume that $J_{N}=\frac{N}{K}$ is an integer. Let $k\left(i\right)$ denote the group number of observation $Z_{i}$, and fix an arbitrary set of maps $\widetilde{\cal F}$. In the $k$-th fold of cross-validation, we will use the observations in group $k$ for testing and the remaining observations for training.
For each group $k=1,...,K$, define $\hat{f}^{-k} :=\arg\min_{f\in \widetilde{{\cal F}}}\frac{1}{N-J_{N}}\sum_{k\left(i\right)\neq k}l(f,Z_i)$ to be the minimizer in $\widetilde{\cal F}$ on the $k$-th training set (i.e., all observations outside of group $k$), and $\hat{e}_{k} :=\frac{1}{J_{N}}\sum_{k\left(i\right)=k}l\left(\hat{f}^{-k},Z_i\right)$ to be the out-of-sample error on the $k$-th test set. Then the average test error across the $K$ folds, $\hat{e}_{CV}\left(\widetilde{{\cal F}}\right) :=\frac{1}{K}\sum_{k=1}^{K}\hat{e}_{k}$, is an estimator for the unobservable expected error of the best prediction rule from class $\widetilde{\mathcal{F}}$. Setting $\widetilde{\cal F}$ to be $\overline{\mathcal{F}}$, ${\cal G}$, or $\{f_{\text{base}}\}$, we can compute $\hat{e}_{CV}\left(\overline{\mathcal{F}} \right)$, $\hat{e}_{CV}\left({\cal G}\right)$ and $\hat{e}_{CV}\left(f_{\text{base}}\right)$ from the data, leading to the following estimator for $\kappa$: \[ \hat{\kappa}= 1 - \frac{\hat{e}_{CV}\left({\cal G}\right)-\hat{e}_{CV}\left({{\cal F}^*}\right)}{\hat{e}_{CV}\left(f_{\text{base}}\right)-\hat{e}_{CV}\left({{\cal F}^*}\right)}. \]
It is crucial that the denominator in $\hat{\kappa}$ does not vanish asymptotically, so we impose the following assumption:
This assumption says that the baseline prediction rule performs strictly worse in expectation than the best prediction rule so there is some room for a model to do better. We show that $\hat{\kappa}$ is asymptotically normal by adapting Proposition 5 in austern2020asymptotics.
Our first application is to the prediction of certainty equivalents for a set of 25 binary lotteries from Bruhin. Each lottery is described as a tuple $x=(\overline{z},\underline{z},p)$, where $\overline{z}>\underline{z} \geq 0$ are the possible prizes, and $p$ is the probability of the larger prize. Each observation consists of a lottery and a reported certainty equivalent by a given subject, so we can describe the feature space $\mathcal{X}$ by the 25 lottery tuples $(\overline{z},\underline{z},p)$ in the Bruhin data, and the outcome space by $\mathcal{Y}=\mathbb{R}$. Note that the residual uncertainty in $Y$ conditional on $X$ reflects heterogeneity in certainty equivalents reported across subjects for the same lottery.
We predict the average certainty equivalent (over subjects) for each lottery in this data set. A prediction rule for this problem is any function $f: \mathcal{X} \rightarrow \mathbb{R}$ from the 25 lotteries to their average certainty equivalents, and the discrepancy between two mappings is defined to be their average mean-squared distance $d(f,f') = \frac{1}{\vert \mathcal{X} \vert} \sum_{x \in \mathcal{X}} (f(x)-f'(x))^2.$
We evaluate the restrictiveness and completeness of two economic models. First we consider a three-parameter version of Cumulative Prospect Theory indexed by $\theta=(\alpha,\gamma,\delta)$, which specifies a utility $ w(p)v(\overline{z}) + (1-w(p)) v(\underline{z})$ for each lottery $(\overline{z},\underline{z},p)$, where
The predicted certainty equivalent of a binary lottery is then given by $f_\theta(\overline{z}, \underline{z},p) =v^{-1}\left( w(p)v(\overline{z}) + (1-w(p)) v(\underline{z})\right).$ Following the literature, we restrict $\alpha,\gamma \in [0,1]$, and $\delta\geq 0$. We specify $\mathcal{F}$ as the set of all such functions $f_{\theta}$ with parameters $\theta$ in this range, and refer to this model simply as CPT. As a baseline, we consider the function $f_{\text{base}}$ that maps each lottery into its expected value, corresponding to $\alpha=\gamma=\delta=1$.
Second, we consider the Disappointment Aversion model of gul1991, using a parametrization proposed in routledge2010generalized with the parameters $\lambda=(\alpha,\eta)$, where $\alpha \in [0,1]$ and $\eta >-1$.\footnote{To facilitate comparison with CPT, we depart slightly from routledge2010generalized by imposing the functional form $v(z)=z^\alpha$ instead of $v(z) = z^\alpha/\alpha$.} The value function for money is the same as in ((ref)), but the probability weighting function is given instead by $\widetilde{w}(p) = \frac{p}{1+(1-p)\eta}.$ There are two parameters: $\alpha$ again reflects the curvature of the utility function, while $\eta>0$ corresponds to “disappointment aversion,” i.e. aversion to realizations of the lottery that are worse than its certainty equivalent. Here the predicted certainty equivalent is $f_\lambda(\overline{z},\underline{z},p)=v^{-1}(\widetilde{w}(p)v(\overline{z}) + (1-\widetilde{w}(p))v(\underline{z})).$ We specify $\mathcal{F}_{\Lambda}$ as the set of all such functions and refer to this model as DA. Again, we use expected value as the baseline prediction, which corresponds to $\alpha=1$ and $\eta=0$ in DA.
We evaluate completeness using mean-squared error as the loss function, i.e., if the reported certainty equivalent is $y$ when the model predicts $\hat{y}$, the loss in that observation is $(\hat{y}-y)^2$.\footnote{This loss function is paired to the average mean-squared discrepancy function we used for measuring restrictiveness, see Appendix (ref) for details.} CPT achieves a striking out-of-sample performance for predicting certainty equivalents in the Bruhin data: it is 95% complete.\footnote{FKLM reports a similar finding for a sample of gain-domain and loss-domain lotteries.} Thus, the model achieves almost all of the possible improvement in prediction accuracy over the baseline.\footnote{This finding is consistent with PeysakhovichNaecker's result that CPT approximates the predictive performance of lasso regression trained on a high-dimensional set of features.} In contrast, DA is only 27% complete on the same data. One explanation is that CPT more precisely captures the observed risk preferences in the data than DA, but another possibility is that CPT is flexible enough to mimic most functions from binary lotteries to certainty equivalents, while DA imposes more substantial restrictions. These explanations have very different implications for how to interpret CPT's empirical success compared to DA's.
To distinguish between these explanations, we now compute the restrictiveness of the two models. We define the eligible set to be all prediction rules satisfying the following criteria:
Constraint (i) requires that the certainty equivalent is within the range of the possible payoffs, while (ii) is equivalent to monotonicity with respect to first-order stochastic dominance.\footnote{The CDF of a binary lottery with $\ol{z}>\ul{z}$ and $0<p<1$ is $F(z) = (1-p)\mathbf{1}\{ \ul{z} \leq z < \ol{z}\} + \mathbf{1}\{z\geq \ol{z}\}$, which is weakly decreasing in $(\ol{z},\ol{z},p)$ for all $z$, so $(\ol{z},\ol{z},p)$ FOSD $(\ol{z}',\ol{z}',p')$ if and only if $(\ol{z},\ol{z},p) \gneqq (\ol{z}',\ol{z}',p').$ There are many pairs of lotteries in the Bruhin lottery data that can be compared via (ii), so these conditions are not vacuous.}
Table (ref) reports the completeness and restrictiveness of both models.
The restrictiveness of CPT is $0.28$, so on average CPT's approximation error is about one fourth of the error of the expected value. DA is more restrictive, with an average approximation error almost one half of the error of the baseline. Thus the two models are not directly comparable: CPT performs substantially better for predicting the real data, but would have performed well out-of-sample given sufficient data from almost any underlying data-generating process that respects first-order stochastic dominance. DA rules out more behaviors that satisfy first-order stochastic dominance, but in doing so is unable to well approximate the actual Bruhin data.
In addition to comparing models such as CPT and DA, our approach can be used to learn more about the role played by specific parameters. Adding a parameter must at least weakly decrease restrictiveness and increase completeness, but we find that parameters can differ substantially in their effectiveness in trading off between these two goals. We also show that models with the same number of parameters can have very different levels of restrictiveness, and thus a simple parameter count is substantively less informative than our measure.
Specifically, we consider alternative specifications of CPT and DA with fewer parameters. Some of these specifications have been studied in the literature: CPT($\alpha,\gamma$), with $\delta=1$, is used in karmarkar1979\footnote{This specification with weighting function $w(p)=\frac{p^\gamma}{p^\gamma+(1-p)^\gamma}$ is very similar to one used in CPT, where the weighting function was $w(p) = \frac{p^\gamma}{p^\gamma+(1-p)^\gamma)^{1/\gamma}}$.}; CPT($\gamma,\delta$), with $\alpha=1$, corresponds to a risk-neutral CPT agent whose utility over money is $u(z)=z$ but exhibits nonlinear probability weighting; CPT($\alpha$), with $\delta=\gamma=1$, corresponds to an Expected Utility decision-maker whose utility function is as given in ((ref)), and is also equivalent to DA($\alpha$).\footnote{See the survey fehrdudaepper for further discussion of these different parametric forms, and others which have been used in the literature.} The model CPT($\gamma$), with $\alpha=\delta=1$, and CPT($\delta$), with $\alpha=\gamma=1$ have not been studied in the literature, but we report them for comparison. We also consider DA($\eta$) as in gul1991, with $\alpha=1$, which corresponds to a disappointment-averse decision maker whose utility is linear in money.
Figure (ref) plots restrictiveness and completeness for these alternative specifications, which reveals that some specifications fall in the interior of the restrictiveness-completeness Pareto frontier introduced in Section (ref): Each of CPT($\alpha,\delta$) and DA($\alpha,\eta$) are dominated, in the sense that another model is simultaneously more complete and also more restrictive.\footnote{Each of CPT($\alpha,\delta$) and DA($\alpha,\eta$) is less complete and less restrictive than the single parameter model CPT($\gamma$), and these differences are statistically significant. (See also Table (ref) in Online Appendix (ref).)} The figure also reveals substantial dispersion in the restrictiveness of these specifications (ranging from $r=0.28$ to $r=0.92$), even though all of the specifications use only a small number of parameters. This observation emphasizes the distinction between our method and a simple parameter count.
By looking more specifically at how restrictiveness and completeness vary across two nested specifications, we can better understand the role that any specific parameter plays. Figure (ref) shows that the different parameters for probability weighting are not equally effective. Adding the parameter $\delta$, which governs the elevation of the probability weighting curve, to any specification of CPT leads to a large drop in restrictiveness in return for only a small gain in completeness. We find a similar result for the “disappointment aversion" parameter $\eta$ in DA, which barely improves upon the completeness of DA$(\alpha)$, but leads to a substantial drop in restrictiveness. In contrast, the parameter $\gamma$, which governs the curvature of the probability weighting function, appears to play an important role in capturing risk preferences: Adding $\gamma$ to any CPT specification leads to a sizeable improvement in completeness at the cost of a modest reduction in restrictiveness. This supports previous findings that probability distortions play an important role in fitting experimental and field data SnowbergWolfers,fehrdudaepper,Donoghue.
We show that the qualitative findings in this section are robust to certain natural changes in the eligible set and the feature set. Together with the robustness check in Section (ref), these results also speak to the sensitivity of the restrictiveness measure in general: although the measure will typically vary with these specifications, it may not be very sensitive in practice for many economic models of interest.
\paragraph{Different distribution over the eligible set.} The uniform distribution is the same as $\mbox{beta}(1,1)$, so to test the sensitivity of the restrictiveness measure we consider nearby $\mbox{beta}(a,b)$ distributions with parameters $(a,b)$ sampled from a uniform distribution over $[0.9,1.1]\times [0.9,1.1]$. For each $(a,b)$ pair, we generate certainty equivalents from a $\mbox{beta}(a,b)$ distribution over the prize range, again keeping only those functions $f$ that satisfy FOSD. Over 100 such distributions $\mbox{beta}(a,b)$, the average restrictiveness is 0.29, with a minimum value of 0.27 and a maximum value of 0.32.
\paragraph{Different eligible set.} Next, we compute the restrictiveness of CPT$(\alpha,\delta,\gamma$) with respect to an eligible set that imposes the range restriction in (i) but drops the FOSD restrictions in (ii). The model's errors are substantially higher when we drop FOSD (increasing from 63.75 to 102.41), but so are the errors of the Expected Value benchmark. The relative performance of CPT$(\alpha,\delta,\gamma)$ compared to the expected-value baseline is nearly identical regardless of whether or not we impose FOSD: the model's restrictiveness relative to this larger eligible set is 0.29 (compared to 0.28 relative to the original eligible set).
\paragraph{Other sets of binary lotteries.} In our main analysis, the feature space $\mathcal{X}$ consisted of 25 binary lotteries from Bruhin data. Below we report the restrictiveness of CPT$(\alpha,\gamma,\delta)$ and DA$(\alpha,\eta)$ with respect to alternative sets of binary lotteries, drawn from five additional papers (see Appendix (ref) for details). Figure (ref) shows the CDF of restrictiveness values across these lotteries (including the Bruhin lotteries) for both models. We find that CPT is not very restrictive on any of these sets of lotteries, and that the distribution of restrictiveness for DA first-order stochastically dominates that of CPT.
\paragraph{Lotteries over the loss domain.} On 25 binary lotteries over the loss domain from Bruhin, the 3-parameter specification of CPT indexed to $(\beta,\gamma, \delta)$ predicts the certainty equivalent $v^{-1}\left((1-w(1-p))\cdot v(\overline{z}) + w(1-p) \cdot v(\underline{z})\right)$ for each lottery $(\overline{z},\underline{z},p)$, where $v(z)= -((-z)^{\beta})$ and $ w(p)= (\delta p^\gamma)/(\delta p^\gamma + (1-p)^\gamma) $. The restrictiveness of CPT on these lotteries is 0.31, with a standard error of 0.02.
\paragraph{Lotteries with larger supports.} Finally, we evaluate the restrictiveness of CPT($\alpha,\delta,\gamma$) on gains-domain lotteries with more than two possible outcomes. For each lottery $(z_1,...,z_n; p_1,...p_n)$, where $0\leq z_1 <... < z_n $, the predicted certainty equivalent is $$v^{-1} \left( \sum_{i} u(x_i) \left[ w \left( \sum_{k=1}^i p_k \right) - w \left( \sum_{k=1}^{i-1} p_k \right) \right]\right),$$ where for $i=1$ we define $ \sum_{k=1}^{0} p_{k}=0$, and $v$ and $w$ have the same functional forms as used above. On 18 three-outcome gain-domain lotteries from BernheimSprenger, the restrictiveness of CPT is 0.57, with a standard error of 0.02. Thus CPT is about twice as restrictive for certainty equivalents on three-outcome lotteries as it is on binary lotteries. On a set of 10 six-outcome lotteries from fudenbergpuri, the restrictiveness of CPT is $0.83$, with a standard error of 0.01. These results suggest that CPT is more restrictive on lotteries with larger supports.
Our second application is to predicting the distribution of initial play in games. Here the feature space $\mathcal{X}$ consists of the 466 unique $3 \times 3$ payoff matrices from FudenbergLiang.\footnote{These data are an aggregate of three data sets: the first is a meta data set of play in 86 games, collected from six experimental game theory papers in LeytonBrownWright; the second is a data set of play in 200 games with randomly generated payoffs, which were gathered on MTurk for FudenbergLiang; the third is a data set of play in 200 games that were “algorithmically designed" for a certain model (level 1 with risk aversion) to perform poorly, again from FudenbergLiang.} The outcome space is the set $\mathcal{Y}=\Delta(\{a_1,a_2,a_3\})$ of distributions of row player actions chosen by the participants in the experiments. The analyst seeks to predict this distribution for each game.
For any two prediction rules $f$ and $f'$, we define $d(f,f')$ to be the average Kullback-Liebler divergence between the predicted distributions: $d(f,f') = \frac{1}{466} \sum_{x \in \mathcal{X}} D(f(x) \| f'(x))$, where $D$ denotes the Kullback-Leibler divergence.
We consider three economic models: The Poisson Cognitive Hierarchy Model (PCHM) of CamererHoChong04, the Level-1 model with logistic best replies (henceforth Logit Level-1), and the PCHM with logistic best replies (henceforth Logit PCHM). The PCHM supposes that there is a distribution over players of differing levels of sophistication: The level-0 player randomizes uniformly over his available actions, the level-1 player best responds to level-0 play StahlWilson94,StahlWilson95,Nagel; and for $k\geq 2$, level-$k$ players best respond to a perceived distribution
over (lower) opponent levels, where $\pi_{\tau}$ is the Poisson distribution with rate parameter $\tau \geq 0$. The parameter $\tau$ is the single parameter of the model.
The Logit Level-1 prediction is defined as follows. For each row player action $a_i$, let $\overline{u}(a_{i})$ be the expected payoff of $a_i$ when the column player uses a uniform distribution. The predicted frequency with which $a_i$ is played is $\exp\left(\lambda \cdot \overline{u}(a_i)\right)/\sum_{i=1}^3 \exp\left(\lambda \cdot \overline{u}(a_i)\right)$, where the logit parameter $\lambda \in \mathbb{R}_+$ is the single parameter of the model.
The Logit PCHM (see e.g. LeytonBrownWright) replaces the assumption of exact maximization in the PCHM with a logit best response. That is, the level-0 player chooses $f_0=(1/3,1/3,1/3)$ as in the PCHM, but we recursively construct the distribution of play for higher levels as follows. For each $k\geq 1$, define \[v_k(a_i) = \sum_{h=0}^{k-1} p_k(h,\tau) \left(\sum_{j=1}^3 f_{h}(a_{j}) u(a_i,a_j)\right)\] to be the expected payoff of action $a_i$ against a player whose type is distributed according to $p_k(\cdot, \tau)$, where $p_k(h,\tau)$ is as given in ((ref)). The distribution of play for a level-$k$ player is then $f_k(a_i)= \exp(\lambda \cdot v_k(a_i))/\sum_{j=1}^3 \exp(\lambda \cdot v_k(a_j))$, where $\lambda \in \mathbb{R}_+$ is a logit parameter. We aggregate across levels using a Poisson distribution with rate parameter $\tau \in \mathbb{R}_+$ to yield the predicted distribution of play.
Finally, we define the baseline prediction rule $f_{\text{base}}$ to predict uniform play in every game $x$. This prediction rule is nested in all three models.\footnote{Let $\tau=0$ in the PCHM or Logit PCHM, and let $\lambda=0$ in Logit Level-1.}
We evaluate completeness using negative log-loss as the loss function, i.e., if the chosen action is $a_i$ when the model predicts distribution $(p_1, p_2,p_3)$, the loss in that observation is $-\log(p_i)$.\footnote{This loss function is paired to the Kullback-Leibler discrepancy function we used for measuring restrictiveness, see Appendix (ref) for details.} The models PCHM, Logit Level-1, and Logit PCHM are 43.6%, 72.7%, and 72.9% complete. Thus, as observed in a related study by LeytonBrownWright, Logit PCHM provides much better predictions of the distribution of play than the baseline PCHM does. Perhaps surprisingly, almost all of Logit PCHM's improved performance can be obtained by simply adding the logit parameter to the Level-1 model; the further improvement from allowing for multiple levels of sophistication is negligible.\footnote{FudenbergLiang found that the Level-1 model provides a good prediction of the modal action, but this does not imply that Logit Level-1 will perform well in predicting the full distribution of play. The fact that it does further suggests that initial play in many of these experiments is rather unstrategic.}
We turn now to evaluating the restrictiveness of these models. We have relatively little understanding about their empirical content, but we do know that they all imply that if an action is strictly dominated, then the frequency with which it is chosen does not exceed 1/3, and that if an action is strictly dominant, then the frequency with which it is chosen is at least $1/3$. We define the eligible set to be all prediction rules that satisfy these conditions.\footnote{In our data, the median frequency of a strictly dominated action is 0.03, and the highest frequency is 0.35; the median frequency for a strictly dominant action is 0.86, and the lowest frequency is 0.69. Payoff maximization implies that dominant strategies should have probability 1 and dominated strategies have probability 0, but this is inconsistent with observed play in most game theory experiments.}
All three models are very restrictive relative to this eligible set: Logit Level-1's restrictiveness is $0.970$, PCHM's restrictiveness is $0.992$, and Logit PCHM's restrictiveness is 0.971. Since the models' completeness ranges from 0.436 to 0.729, they are much better predictors of the real data than of the synthetic data. Table (ref) reports completeness and restrictiveness measures for the models. We find that Logit Level-1 and Logit PCHM are substantially more complete than PCHM and only slightly less restrictive, but none of the models is dominated by another. Moreover, Logit Level-1 and Logit PCHM are almost identical in terms of completeness and restrictiveness, even though the parametric forms of the two models are not evidently related.\footnote{No value of $\tau$ in the PCHM yields the Level-1 model, so Logit Level-1 is not nested within Logit PCHM.}
Finally, as a robustness check, we consider strengthening the background constraints imposed on the eligible set $\mathcal{F}$. For each $t\in [0,0.3)$, we define the eligible set $\mathcal{F}(t)$ to include all prediction rules $f$ that satisfy the following conditions: (1) If an action is strictly dominated, then the frequency with which it is chosen does not exceed $1/3-t$; (2) If an action is strictly dominant, then the frequency with which it is chosen is at least $1/3+t$. The constraint imposed by these conditions increases in $t$, and $t=0$ returns our original specification of $\mathcal{F}$. We find that across choices of $t\in [0,0.3)$, the restrictivenesses of PCHM, Logit PCHM, and Logit Level-1 do not fall below 0.89 (see Table (ref) below). This tells us that constraints on the frequency of strictly dominated and strictly dominant strategies are a very small part of the empirical content of these models.
Our final application is to the prediction of microfinance takeup rates following diffusion of information in social networks. We use data from a study by banerjee2013diffusion, in which certain “leaders” in 43 villages in Karnatka, India were given information about a microfinance program, and takeup of the program was then tracked.\footnote{In 2007, the microfinance institution Bharatha Swamukti Samsthe invited leaders within each village to an information meeting, and asked the leaders to spread the information. The data set contains the resulting microfinance takeup rate for each village and some measures of social connections between households.}
For each village $i$, let $y_{i}$ be the average takeup rate among non-leader households.\footnote{This is the outcome variable that banerjee2013diffusion focus on.} Our goal is to predict $y_{i}$ given the observed characteristics $X_{i}$ of village $i$. Specifically, a village configuration $X_{i}:=\left(N_{i},A_{i},L_{i}\right)$ consists of a set $N_{i}$ of villagers, an $n_{i}\times n_{i}$ adjacency matrix $A_{i}$ that represents the measured social network, and the set $L_{i}$ of leaders in village $i$. The feature space $\mathcal{X}$ is the collection of 43 village configurations, and prediction rules are maps $f:{\cal X}\to\left[0,1\right]$ that from village configurations to the takeup rate among non-leaders. There are no obvious a priori restrictions on the takeup rates, so we set $\mathcal{F}$ to be the set $\left[0,1\right]^{43}$ of all possible prediction rules from ${\cal X}$ to $\left[0,1\right]$. We set the discrepancy function as $d(f,g):=\frac{1}{43}\sum_{i=1}^{43} (f(x_i)-g(x_i))^2$ and the loss function as $l(f(x),y) := (f(x)-y)^2.$
The first parametric models we consider are OLS regressions with various subsets of the following eight network statistics as regressors: (1) average eigenvector centrality of leaders; (2) average degree centrality of leaders; (3) average degree centrality of all villagers; (4) average betweenness centrality of leaders; (5) clustering coefficient of village network; (6) average path length in village network; (7) proportion of connected (non-isolated) villagers; (8) proportion of leaders.
We compute the restrictiveness and completeness of a sequence of OLS models by incrementally adding the regressors listed above. We set the baseline as OLS regression on a constant, which is a special case of all the linear models we consider. With the loss function $l(f(x),y):=(y-f(x))^2$, an estimator of completeness (computed based on in-sample errors without the use of cross validations) reduces to the R squared of the OLS regression.\footnote{Recall that the R-squared of an OLS regression is defined by $R^2:= 1 - SSR/SST$, where $SSR := \sum_i (y_i - x_i'\hat{\beta})^2$ is the expected loss under an OLS regression model and $SST := \sum_i (y_i - \ol{y})^2$ is the expected loss under a constant model.}
We also consider a partially linear model built upon the “network gossip centrality” described in banerjee2019using. To do this, we model each non-leader household's takeup probability as a function of its position in the village. We define the “hearing matrix” of village $i$ by $H_{i}\left(\theta_{1}\right):=\sum_{t=1}^{T}\theta_{1}^{t}A_{i}^{t}$, where $T$ is some given number of time periods for information diffusion.\footnote{$\left(\sum_{t=1}^{T}A_{i}^{t}\right)_{jk}$ counts the number of paths from $j$ to $k$ of length up to $T$. We set $T = 5$ following banerjee2019using.} With $\theta_{1}=1$, the $jk$-th entry of $H_{i}\left(1\right)$ can be interpreted as the expected number of times villager $k$ hears a piece of information that originates from villager $j$ within $T$ periods of time. The parameter $\theta_{0}\in\left(0,1\right)$ discounts longer paths of diffusion. For each non-leader $k$ in village $i$, we define $x_{i,k}\left(\theta_{1}\right):=\sum_{j\in L_{i}}\left(H_{i}\left(\theta_{1}\right)\right)_{jk}$ as the “network gossip centrality” of non-leader $k$, which counts the (discounted) sum of number of paths from the leaders of village $i$ to non-leader $k$. Next, we model the takeup probability of non-leader $k$ as function of $k$'s “network gossip centrality” based on a logistic model $p_{i,j}\left(\theta_{0},\theta_{1}\right):=\frac{\exp\left(\theta_{0}+x_{i,j}\left(\theta_{1}\right)\right)}{1+\exp\left(\theta_{0}+x_{i,j}\left(\theta_{1}\right)\right)}$, where $\theta_{0}$ is a location parameter.\footnote{Note that we do not include a scale parameter here, since if present, it will be absorbed into $\theta_{1}$.} The expected village-level takeup rate among non-leaders can then be derived as the average $p_{i,j}(\theta_0,\theta_1)$ among non-leaders. To allow additional flexibility, and to nest the naive constant model as a special case, we introduce two additional linear parameters $(\theta_2,\theta_3)$, and set: $f_{i}\left(\theta\right):=\theta_{2}+\theta_{3}\cdot\frac{1}{\left|N_{i}\backslash L_{i}\right|}\sum_{j\notin L_{i}}p_{ij}\left(\theta_{0},\theta_{1}\right)$. This model is very stylized; our purpose is to illustrate how our algorithmic approach can be used to evaluate the restrictiveness of a structural model whose flexibility is otherwise difficult to gauge.
Table (ref) reports the restrictiveness and completeness of the models described above.\footnote{Table (ref) displays the restrictiveness of the linear models based on M = $10000$ simulations, while restrictiveness for the partially linear models is computed using $M = 100$ simulations. Completeness for all models is computed based the real data with $N = 43$ villages.} The panel “Linear Models" contains results about a sequence of linear models, with a new regressor added to the OLS regression in each row.\footnote{We add the regressors sequentially according to the ordering above, and omit many other different orderings of the same set of regressors, since the regressions in Table (ref) suffice to illustrate our main point.} For example, the row “+ Degree Centrality” corresponds to an OLS regression of takeup rates on a constant, the leaders' average eigenvector, and the leaders' average degree centrality.
The numerical results for linear models are as expected: as more regressors are added the model becomes more flexible, so restrictiveness decreases while completeness increases. While restrictiveness seems to be decreasing at an approximately linear rate starting from the second regression, the corresponding increases in completeness appear less uniform, and in particular, completeness barely changes when we add the regressor “average path length in the village.” Note that this does not mean that this additional regressor approximately lies in the linear span of all previously included regressors, since we do observe a nontrivial reduction in restrictiveness from the addition of this regressor: New regressors eventually barely improve fit to the data, but they continue to decrease restrictiveness.
A priori it is unclear how restrictive the partially linear model is. It turns out that its restrictiveness is very high, 0.94, suggesting that the individual-level modeling of takeup rates as a function of network gossip centrality imposes substantial restrictions across village configurations. However, this model's completeness is only 0.07, so it does not capture much of the variation in village takeup rates.
This four-parameter partially linear model is dominated by the simple linear model with a constant and the average eigenvector centrality of leaders as the single regressor: the latter has both higher restrictiveness (0.9762 \textgreater 0.9408) and higher completeness (0.2577 \textgreater 0.0674). This shows that even a detailed, structured, and economically-motivated model may turn out to be more flexible than a simple linear model, and that the added flexibility need not help it fit real data.
When a theory fits the data well, it matters whether this is because the theory captures important regularities in the data, or whether the theory is so flexible that it can explain any behavior at all. We provide a practical, algorithmic approach for evaluating the restrictiveness of a theory, and demonstrate that it reveals new insights into models from two economic domains. The method is easily applied to models across diverse domains.
As highly flexible machine learning methods become more popular in economics, economic theory is distinguished in part by the structure it imposes on behaviors. We view these restrictions as an important part of the value added by economic theory, so it is natural to ask how restrictive economic models are compared to the highly flexible approaches used in machine learning. Our restrictiveness measure offers a way to quantify this.