EconBase
← Back to paper

Second-order Inductive Inference: an axiomatic approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

62,331 characters · 4 sections · 26 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Second-order Inductive Inference: an axiomatic approach

\pagenumbering{gobble}

\pagestyle{fancy} \thispagestyle{plain}

abstractConsider a predictor who ranks eventualities on the basis of past cases: for instance a search engine ranking webpages given past searches. Resampling past cases leads to different rankings and the extraction of deeper information. Yet a rich database, with sufficiently diverse rankings, is often beyond reach. Inexperience demands either “on the fly” learning-by-doing or prudence: the arrival of a novel case does not force (i) a revision of current rankings, (ii) dogmatism towards new rankings, or (iii) intransitivity. For this higher-order framework of inductive inference, we derive a suitably unique numerical representation of these rankings via a matrix on eventualities $\times$ cases and describe a robust test of prudence. Applications include: the success/failure of startups; the veracity of fake news; and novel conditions for the existence of a yield curve that is robustly arbitrage-free.

{11.5cm} \epigraph{From the past, the present acts prudently, lest it spoil future action.}{Titian: Allegory of Prudence }

Introduction

Experience is the basis of prediction. Yet outside of stylised settings, even the most experienced forecasters do not claim access to “the full model” that is closed with respect to the relevant states of the world.

exampleThe canonical large-world setting is that of the global financial markets. In 2019, all models ommited details of COVID-19. Similarly, in 2007, a clear description of the sub-prime mortgage crisis was beyond reach. The central role of simulation and bootstrap methods in empirical finance points to a prevalence of inductive reasoning.\footnote{Consider Cowles-Forecasting, White-Data_snooping, FF-Luck_vs_skill and HL-Lucky_factors.} That is, to forecasting on the basis of past cases as opposed to a full description of future states. At the same time, market makers need to set prices that are robust to changes that open the door to exploitation via arbitrage.

\pagenumbering{arabic} \setcounter{page}{2} Incomplete models are a key motivation for recent axiomatic updates to the standard Bayesian framework such as “reverse Bayesianism” of KV_Reverse_Bayes. This and other work (KV-Awareness_of_U, HR-Knowledge_of_U and GKMQT_Robust_experiments) on unawareness and robustness in the state-space sense provide part of our inspiration for the present upgrade to the axiomatic foundations of inductive inference in GS_Inductive_inference. By taking rankings as primitive, we extend $\textup{[GS]}$\ to accommodate a forward-looking version of the second-order induction of AG-Second-order_induction.

In the present framework, the basic building blocks of the model are observations or (synonymously) past cases. A given past case may be empirical or theoretical, and the predictor's model is naturally bounded in size and scope by the predictor's experience. Our contribution is to extend the framework of $\textup{[GS]}$\ to model less experienced predictors that are prudent. That is to agents that are able to proactively engage in second-order induction: by “pulling themselves up by the bootstraps” and looking at how their model extends to novel cases. Thus allowing them to survive their initial phase of inexperience.

In the remainder of this section, we informally introduce: the model of (ref); the axioms and matrix representation of (ref); and the applications to which we return in (ref). Proofs of main results appear in appendices (ref) to (ref).

\paragraph{Synopsis of model and results with examples.} The predictor is endowed with a qualitative plausibility ranking of eventualities given her current database $D^{\star}$ of past cases. Moreover, the same is true for every finite resampling of $D^{\star}$. We identify conditions on the resulting family of (ordinal) rankings for the existence of a suitably unique real-valued matrix $\mathbf v$ on eventualities$\times$cases that represents the information in these rankings. The form of this representation is linear on databases (\ie\ additive over cases) and separable on eventualities, so that for every database $D$, eventuality $y$ is more likely than $x$ if, and only if,

linenomath*\begin{equation} \sum_{c\,\in D} \mathbf v(x,c) \leq \sum_{c\,\in D} \mathbf v(y,c). \end{equation}

The similarity weight $\mathbf{v}(x,c)$ is the degree of support that case $c$ lends to $x$.

exampleAs a canonical example, consider predicting the slope coefficient $\beta_{i}$ from the regression of asset $i$'s returns on the market portfolio FM-Two_pass. Then values $\beta_{i}$ are eventualities and $D^{*}$ is the current sample of past returns. In this setting, $\mathbf v$ is an empirical log-likelihood function that we generate via a generalised notion of bootstrapping of cases in $D^{\star}$.

Key to the contribution of $\textup{[GS]}$\ is an endogenous notion of case types: a partition of cases according to the marginal information they contribute to a given database. This marginal contribution is measured in terms of the impact on the rankings of eventualities.\footnote{This notion of case type is therefore close to Quine's “perceptual similarity” (see (ref)).} Case types are analogous to states in the sense that they form the model's dimensions: the lower the dimension the less the experience.

example[Search Engine Results Page, SERP] Advertising aside, when users conduct a web search, the search engine compiles a ranking $\mathbin{\preceq}_{D^{\star}}$ of (web)pages $x, y, z, \dots$ on the basis of its database $D^{\star }$ of past cases: searches of past users plus feedback from subsequent clicks. Resampling yields other databases $D$ and other rankings. At one extreme, past cases may be so similar to one another that the same plausibility ranking arises regardless of how the data is resampled. At the other, past cases may be sufficiently rich that resampling yields every feasible plausibility ranking of eventualities.

A rich set of past cases is at the heart of the diversity axiom of $\textup{[GS]}$--which we refer to as 4-diversity. This restricts the model to predictors whose current data is sufficiently rich that a resampling exercise generates all $4!=24 $ strict (\ie\ total) rankings of every subset of four eventualities. In this paper, we accommodate less inexperienced predictors by only replacing 4-diversity\ with conditional-\textit{2}-diversity. Given the other basic axioms of $\textup{[GS]}$, \textup{conditional-\textit{2}-diversity}\ turns out to be equivalent to requiring that, for every three distinct eventualities $x$, $y$ and $z$, resampling generates at least $3$ of the $3! = 6$ possible distinct strict rankings of this triple. \textup{Conditional-\textit{2}-diversity}\ is minimal in the sense that, in its absence, we lose both existence and uniqueness of the similarity representation (see (ref)).

To compensate for a lack of experience, we introduce a subtle and more flexible notion of cases that allows us to capture the predictor's potential awareness of her limited experience. Formally, related notions in the literature go by the name of unforeseen consequences in GQ-Surprises and shadow propositions in the setting dynamic awareness in HP-Dynamic_awareness. Via a content-free case $\mathfrak f $, the predictor can explore the impact on her model of the arrival of a novel case type. In effect, this involves meta-analysis of how her similarity function will evolve over time. Hence the reference to second-order induction.

The prudent predictor ensures her model is robust to the arrival of novel case types. She ensures that, when a novel case arrives, she can accept the ranking it generates without finding herself in the potentially costly position of generating intransitive rankings when she combines past and novel cases.

example[Second-order inductive inference] Consider a search-engine startup seeking to establish itself in the face of incumbents with the experience of Google. The start-up engages in second order induction when it is learning the similarity function itself. This may include learning the values $\mathbf{v}(x,c)$ of (ref), but it may also involve costly updates of the model structure “on the fly”, \eg\ redefining case types, rankings, \etc. The startup is prudent if, ex ante, it structures its model to ensure that it is relatively costless to extend to novel case types$:$ once they arrive.

Prudence is only worthwhile when revisions of the predictor's model `on the fly', once the novel case arrives, is costly. Consider “zero-day attacks” in the setting of cyber security. \Citet{Hota_et_al-Cyber_security} highlight the essence of time when a novel attack on a computer network arrives; and that such attacks are novel precisely because cyber-security experts have already built in solutions to known vulnerabilities. Also, tradeoffs between time, cost and learning are nowhere more important than in finance. For a bond-market setting, we are able to provide a formal equivalence between prudence and arbitrage pricing in (ref). In (ref), we also discuss: other applications; empirical evidence linking intransitivity, memories and novelty; and connections with the literature on second-order induction in more detail.

Model

Following $\textup{[GS]}$, let the nonempty set $X$ denote the conceivable eventualities of the present prediction problem and let $ \operatorname{rel}(X)$ denote the set of binary relations or rankings on $X$. For instance, for a search engine, we identify webpage $x$ with the eventuality “page $x$ is the desired webpage”. The predictor is equipped with her current memory ${C^\star}$: the union of a finite set of past cases ${D^\star}$ and a variable or free case $\mathfrak f$. The cases in ${D^\star}$ collectively represent the forecaster's relevant observations or experience. Our first and most fundamental modification of the primitives of $\textup{[GS]}$ is the inclusion of $\mathfrak f$ in the current memory ${C^\star}$.

remark*On a computer, a natural implementation of this setup is the following. Take every case $c \in {C^\star} $ to consist of a pair $p \times m:$ a pointer $p$ that references a memory location and the memory content $m$. Each $c\in {D^\star}$ is identical to a case in the setting of $\textup{[GS]}$. But, for $c = \mathfrak f$, there is no meaningful memory content, so $m$ is “empty” or assigned an arbitrary null value. From another perspective, cases in $ {D^\star}$ are constant (of arity zero) whereas $\mathfrak f$ is a variable (of positive arity). For every nonempty subsample $ D \subseteq {D^\star} $, the predictor has sufficient information to determine a well-defined ranking $\mathbin{\preceq}_ D $ in $ \operatorname{rel}(X)$. In contrast, since $\mathfrak f = \acute{p} \times \acute{m}$ has no meaningful memory content, $\mathbin{\preceq}_\mathfrak f$ is indeterminate and a free variable in $\operatorname{rel}(X)$. The prudent predictor gains a better understanding of her current model by assigning a ranking to $\acute{m}$ and exploring the extensions of (ref).

Like $\textup{[GS]}$, we accommodate a forecaster that goes beyond her current memory and includes hypothetical cases ${\mathds C}$ that she may not have experienced, but which, through reasoning, interpolation or resampling, she can clearly describe. These hypothetical cases are formally constant, like members of ${D^\star}$. With case resampling and subsampling from the literature on bootstrapping in mind, let \[{\mathds D}\defeq \left\{ D\subseteq {\mathds C}: \mathbin{\sharp}\hskip1pt D<\infty\right\} \] denote the set of (finite) determinate or constant databases. (These are referred to as memories in $\textup{[GS]}$.) Like ${D^\star}$, each $D\in {\mathds D} $ contains no copies of $\mathfrak f$.

Let $[\mathfrak f] $ denote a set of copies of $\mathfrak f $. Finally, let ${\mathds C^{\mathfrak f}} \defeq {\mathds C} \cup [\mathfrak f]$ and let ${\mathfrak C}$ denote a member of $\{{\mathds C}, {\mathds C^{\mathfrak f}}\}$. Let ${\mathds D^{\mathfrak f}} $ denote the corresponding set of all finite subsets of $ {\mathds C^{\mathfrak f}} $ and take

linenomath*\begin{equation*} {\mathfrak D} = \left\{ \begin{array}{ll} {\mathds D} & if, and only if, ${\mathfrak C}= {\mathds C}$, and\\ {\mathds D^{\mathfrak f}} &otherwise. \end{array}\right. \end{equation*}

The predictor is endowed with a well-defined plausibility ranking $\mathbin{\preceq}_D $ in $ \operatorname{rel}(X)$ for each $D$ in ${\mathds D}$. Denote the symmetric part by $\simeq _ D$ and asymmetric part by $ \mathbin{\prec} _ D $. In a minor departure from $\textup{[GS]}$, the primitive of our model is a point in $\operatorname{rel}(X)^{{\mathds D}}$

linenomath*\[\mathbin{\preceq}_{{\mathds D}} \defeq \langle\mathbin{\preceq}_{D}:D\in {\mathds D}\rangle\,.\footnote{The present approach is equivalent to taking $\mathbin{\preceq}_{{\mathds D}} = \left\{D \times \mathbin{\preceq}_{D}: D \in {\mathds D}\right\}$. This way we maintain pairwise distinctness of $\mathbin{\preceq}_{C}=\mathbin{\preceq}_{D}$ such that $C \neq D$. I thank Maxwell B. Stinchcombe for bringing this point to my attention.} \]

For each $C$ in $ {\mathds D^{\mathfrak f}} \bs {\mathds D} $, the fact that for some $ c \in [ \mathfrak f ] $, $ c \in C $ means that $\mathbin{\preceq}_{C} $ is indeterminate, free variable in $\operatorname{rel}(X)$. Although, in isolation each such $ \mathbin{\preceq}_{C} $ is free, when the axioms we introduce hold, the potential values of the variable $ \mathbin{\preceq} _ {\mathds D^{\mathfrak f}} \defeq \langle \mathbin{\preceq} _ C : C \in {\mathds D^{\mathfrak f}} \rangle $ are constrained by the current values of the constant $\mathbin{\preceq}_{\mathds D}$.

\paragraph{Case types.} As in $\textup{[GS]}$, two past cases $ c , d \in {\mathds C} $ are of the same case type if, and only if, the marginal information of $ c $ is everywhere equal to the marginal information of $ d $. Formally, $ c \sim ^{ \star } d $ if, and only if, for every $ D \in {\mathds D} $ such that $ c , d \notin D $, $ \mathbin{\preceq} _ { D \cup \{ c \} } = \mathbin{\preceq} _ { D \cup \{ d \} }$. By observation 1 of $\textup{[GS]}$, $ \sim ^{ \star }$ is an equivalence relation on $ {\mathds C} $ and as its collection of equivalence classes generate a partition ${\mathds {T}}$ of ${\mathds C}$.

We extend $ \sim ^{ \star } $ to $ {\mathds C^{\mathfrak f}} $ by taking $ [ \mathfrak f ]$ to be an equivalence class of its own, so that, for every $ c \in {\mathds C} $, $ c \nsim ^{ \star } \mathfrak f $. We let ${\mathds{T} ^ \mathfrak f }$ denote the corresponding partition of ${\mathds C^{\mathfrak f}}$.

Like $\textup{[GS]}$, we also extend $\sim^{\star}$ to ${\mathds D^{\mathfrak f}}$ by treating databases that contain the same number of each case type as equivalent. That is, $C \sim ^{\star } D$ if, and only if, for every $t \in {\mathds{T} ^ \mathfrak f }$, the numbers $\mathbin{\sharp}\hskip1pt (C \cap t)$ and $\mathbin{\sharp}\hskip1pt (D \cap t) $ of that case type coincide. To enable a translation of each database to counting vectors $t \mapsto \mathbin{\sharp}\hskip1pt (D \cap t) $, we impose a

assumption*For every $ t \in {\mathds{T} ^ \mathfrak f }$, there are infinitely many cases in $ t $.

Our key definition is the following.

definition$\mathbin{\mc R} \defeq \langle \mathbin{\mc R}_{D}: D\in {\mathfrak D} \rangle $ is an extension, and in particular a $Y$-extension, of $ \preceq_{{\mathds D}}$ if, for some nonempty $ Y \subseteq X $, the following all hold$:$ \begin{enumerate} • for every $ D\in {\mathfrak D} $, $\mathrel{\mc R}_{D} $ belongs to $ \operatorname{rel} (Y)$, • for every $ D \in {\mathds D}$ and every $x,y\in Y$, $x \mathrel{\mc R}_{D} y $ if, and only if, $x \preceq_{D} y$, • for every $D\in {\mathfrak D}$ and every $c,d \in {\mathfrak C}\bs D$, if $ c \sim ^ \star d $ then $ \mathbin{\mc R} _ { D \cup \{c\} } = \mathbin{\mc R} _ { D \cup \{d\}}$. \end{enumerate} An extension $\mathrel{\mc R}_{{\mathfrak D}}$ is proper if ${\mathfrak D} = {\mathds D^{\mathfrak f}}$ and improper otherwise.

By assigning rankings to the free case $\mathfrak f$, (potential) extensions of $\preceq_{{\mathds D}}$ simulate the arrival of novel information. Part (ref) of the definition of an extension ensures that, for every proper $Y$-extension $\mathrel{\mc R}$ and every $D\in {\mathds D^{\mathfrak f}}$, $\mathrel{\mc R}_{D}$ is a well-defined binary relation on $Y$. For every extension $\mathrel{\mc R}_{{\mathfrak D}}$ and every $D \in {\mathfrak D}$, let $ \mathrel{\mc I}_{D}$ and $ \mathrel{\mc P}_{D}$ respectively denote the symmetric and asymmetric parts of $\mathrel{\mc R}_{D}$.

Part (ref) of the definition implies that, for every $Y$-extension $\mathrel{\mc R}$ and every $D\in {\mathds D}$, $ \mathrel{\mc R}$ is simply the restriction $\mathbin{\preceq}_{D}\cap Y^{2} $ of $\preceq_{D}$ to $Y$. We therefore refer to part (ref) of the definition of an extension as the preservation or nonrevision condition.

Part (ref) of the definition ensures that, for proper extensions $\mathrel{\mc R} $, the partition ${\mathds{T} ^ \mathfrak f }$ of case types generated by $\sim^{\star}$ is at least as fine the partition generated by the equivalence relation generated by $ \mathrel{\mc R}$. Two cases $ c,d \in {\mathds C^{\mathfrak f}} $ are equivalent with respect to $ \mathrel{\mc R} $, written $ c \sim^{\mathbin{\mc R}} d $, if, for every $ D \in {\mathds D} $ such that $ c,d \notin D $, $ \mathbin{\mc R} _ { D \cup \{c\} } = \mathbin{\mc R} _ { D \cup\{d\} }$. This notion allows us to partition the set of proper extensions as follows.

definition*A proper extension $\mathrel{\mc R} $ is either regular or novel. It is novel whenever $[\mathfrak f]$ is a distinct equivalence class of $\sim^{{\mathrel{\mc R}}}$, so that, for every $ c \in {\mathds C} $, $ c \nsim ^ {\mathrel{\mc R}} \mathfrak f $.

For novel extensions, $\mathfrak f $ mimmicks the potential arrival of new information. Yet novel extensions need not feature qualitatively new rankings (\ie\ rankings that do not feature in $\preceq_{{\mathds D}}$). In (ref), we show that this is because a novel case type is characterised by the quantitative notion of a similarity weight.

For every regular extension $\mathrel{\mc R}$, there exists $c \in {\mathds C}$ such that $c \sim^{\mathbin{\mc R}} \mathfrak f$. Thus, there are as many regular extensions as there are past case types (\ie\ $\mathbin{\sharp}\hskip1pt {\mathds {T}}$). Yet every regular $Y$-extension $\mathrel{\mc R}$ is equivalent to the unique improper $Y$-extension $ \langle \mathbin{\preceq}_{D} \cap Y^{2}: D \in {\mathds D}\rangle$ in the sense of

observationFor every regular $Y$-extension $\mathrel{\mc R}$ and improper $Y$-extension $\mathrel{\acute{\mathrel{\mathcal R}}}$, for every $ C \in {\mathds D^{\mathfrak f}} $, there exists $ D \in {\mathds D} $ such that $C \sim^{\mathbin{\mc R}} D$ and $\mathbin{\mc R}_{C} = \mathbin{\acute{\mathbin{\mathcal R}}}_{D}$.\footnote{This means that there is a canonical embedding of $\left\{C \times \mathbin{\mc R}_{C}: C \in {\mathds D^{\mathfrak f}}\right\}$ in $\{D \times \mathbin{\acute{\mathbin{\mathcal R}}}_{D}: D \in {\mathds D}\}$. The converse embedding follows from the nonrevision condition of (ref).}
proof[Proof of (ref)] Fix $Y\subseteq X$ nonempty and $\mathrel{\acute{\mathrel{\mathcal R}}}$ regular. \Wlog, take $ C \in {\mathds D^{\mathfrak f}} \bs {\mathds D} $, so that $ C $ contains at least one copy of $ \mathfrak f $. For any $ c \in C \cap [ \mathfrak f ] $, the fact that $ \mathrel{\acute{\mathrel{\mathcal R}}} $ is regular implies that $ c \sim^{\mathbin{\acute{\mathbin{\mathcal R}}}} c _ 1 $ for some $ c _ 1 \in {\mathds C} $. The richness assumption ensures that we may choose $ c _ 1 $ from the complement of $ C $. Then, since neither $c$ nor $c_{1}$ belong to $ C _ 1 \defeq C \bs \{c\} $, $c \sim^{\mathbin{\acute{\mathbin{\mathcal R}}}} c_{1}$ implies $ \mathbin{\acute{\mathbin{\mathcal R}}}_{C} = \mathbin{\acute{\mathbin{\mathcal R}}} _ { C _ 1 \cup \{c_{1}\} }$. If $ c $ is the unique member of $ C \cap [ \mathfrak f ]$, then the proof is complete. Otherwise, using the fact that $ C $ is finite, we may proceed by induction until we obtain a set $ C _ n $ such that $ C _ n \cap [\mathfrak f ] $ is empty and $ D \defeq C _ n \cup \{ c _ 1 , \dots , c _ n \} $ belongs to $ {\mathds D} $. Part (ref) of (ref) then implies $ \mathbin{\acute{\mathbin{\mathcal R}}} _ { D } = \mathbin{\preceq} _ { D } \cap Y^{2}$, so that, since $\mathrel{\mc R}$ is improper, $\mathbin{\acute{\mathbin{\mathcal R}}}_{D}=\mathbin{\mc R}_{D}$. Finally, since $C \sim ^{\mathbin{\acute{\mathbin{\mathcal R}}}} D$, $ \mathbin{\acute{\mathbin{\mathcal R}}} _ { C } = \mathbin{\acute{\mathbin{\mathcal R}}}_{D} $, as required.

Axioms and main theorem

\paragraph{The basic axioms\hskip-10pt} of $\textup{[GS]}$, which we rewrite in terms of extensions, are the following. In each of these axioms, $\mathrel{\mc R} $ is an arbitrary $Y$-extension of $\preceq_{{\mathds D}}$.

taggedblank{$\textup{A}0$}[Transitivity axiom for $\mathrel{\mc R}$] For every $D \in {\mathfrak D} $, $ \mathrel{\mc R}_{D} $ is transitive.
taggedblank{$\textup{A}1$}[Completeness axiom for $\mathrel{\mc R}$] For every $ D \in {\mathfrak D} $, $\mathrel{\mc R}_{D} $ is complete.
taggedblank{$\textup{A}2$}[Combination axiom for $\mathrel{\mc R}$] For every disjoint $ C, D\in {\mathfrak D} $ and every $x,y\in Y$, if $ x \mathrel{\mc R}_{C} y $ and $ x \mathrel{\mc R}_{ D } y $, then $ x \mathrel{\mc R}_{ C \cup D} y $$;$ and if $x \mathrel{\mc P}_{C} y$ and $x \mathrel{\mc R}_{D} y$, then $x\mathrel{\mc P} _{C \cup D } y$.
taggedblank{$\textup{A}3$}[Archimedean axiom for $\mathrel{\mc R}$] For every disjoint $ C,D \in {\mathfrak D} $ and every $x,y\in Y $, if $x \mathrel{\mc P}_{D} y$, then there exists $ k \in \mbb Z_{\text{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\srcsize}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}$+\mkern-2mu+$}} $ such that, for every pairwise disjoint collection $\left\{ D _ j : \text{ $ D _ j \sim ^ {\mathrel{\mc R}} D $ and $ C \cap D _ j =\emptyset$}\right\} _ 1 ^ k$ in ${\mathfrak D}$, $ x \mathrel{\mc P}_{ C \cup D_{1} \cup \cdots \cup D_{k} } y.$

\paragraph{The diversity axioms\hskip-10pt} that now follow require that ${\mathds D}$ is sufficiently rich to support $Y$-extensions $\mathrel{\mc R}$ with a variety of total orderings: \ie\ complete, transitive and antisymmetric ($x \mathrel{\mc R}_{D} y $ and $y \mathrel{\mc R}_{D} x$ implies $x = y$). Let $ \textup{total} ( \mathrel{\mc R} )$ denote the set $ \left\{ R : \text{for some $ C \in {\mathfrak D} $, $ R = \mathbin{\mc R} _{ C }$ is total}\right\}$ of of total orders that feature in $ \mathrel{\mc R} $. For $ k = 4 $, the following axiom is a restatement of the diversity axiom of $\textup{[GS]}$.

taggeddiv{\hskip-5pt}[$k$-Diversity axiom] For every $ Y \subseteq X $ of cardinality $ n= 2, \dots, k $, every regular $ Y$-extension $\mathrel{\mc R}$ of $ \mathbin{\preceq} _{ {\mathds D} }$ is such that $\mathbin{\sharp}\hskip1pt \textup{total}(\mathrel{\mc R}) = n!$\hskip.5pt.

We say $k$-diversity holds on $Z$ if the axiom holds with $Z$ in the place of $X$. By lowering the bar for the required number of total orders, the following axiom substantively weakens 3-diversity\ and a fortiori 4-diversity.

taggedblank{$\textup{A}4$$'$}[Partial 3-diversity] For every $ Y\subseteq X $ with cardinality $ n = 2 $ or $ 3 $, every regular $X$-extension $ \mathrel{\mc R} $ of $ \preceq _{ {\mathds D} }$ is such that $ \mathbin{\sharp}\hskip1pt \textup{total} ( \mathrel{\mc R} ) \geq n $.

(ref) of the appendix provides an example of $\preceq_{{\mathds D}}$ satisfying (ref)–(ref), but not 3-diversity. Partial-3-diversity\ is clearly stronger than 2-diversity\ and, moreover, it plays the dual role of guaranteeing uniqueness of the representation and allowing us to avoid restrictions on the cardinality of $ X $. (ref) shows that \textup{partial-\textit{3}-diversity}\ is the weakest axiom with these properties. Moreover, the observation below shows that, in our setting, \textup{partial-\textit{3}-diversity}\ is equivalent to

taggedblank{A$4$}[Conditional-2-diversity] For every three distinct elements $ x , y , z \in X $, within one of the sets $ \{ D' : x \prec _{D ' } y \}$ and $ \{ D' : y \prec_{D ' } x \} $ there exists both $C$ and $D$ such that $ z \prec _{ C } x $ and $ x \prec _{ D } z $. If $ \mathbin{\sharp}\hskip1pt X = 2 $, then 2-diversity\ holds on $X$.
theoremEnd[link to proof]{observation} For $\preceq_{{\mathds D}}$ satisfying (ref)--(ref), conditional-2-diversity\ and partial-3-diversity\ are equivalent.
proofEndThis follows directly from (ref) and the construction of $\preceq_{\mathds J}$.

\paragraph{The prudence axiom \hskip-10pt} that follows is our final requirement. It is distinguished by the fact that it imposes structure on novel extensions. As we will see in the proof of the main theorem (see (ref)), novel extensions are characterised by a cardinal notion. In contrast, the notions of testworthiness and perturbation that we now introduce are ordinal in nature.

definition*A proper extension $ \mathrel{\mc R} $ of $\mathbin{\preceq}_{{\mathds D}}$ is testworthy if it satisfies (ref)--(ref) and, for some $D\in {\mathds D}$ such that $\mathrel{\mc R}_{D}$ is total, $ \mathbin{\mc R} _{ \mathfrak f }$ is the inverse of $\mathbin{\mc R}_{D}$.\footnote{Recall that the inverse $ \mathrel{\mc R} _{ D } ^{ - 1 }$ of $ \mathrel{\mc R} _{ D }$ satisfies $ x \mathrel{\mc R} _{ D } ^{ - 1 } y $ if, and only if, $ y \mathrel{\mc R} _{ D } x$.}

We now introduce perturbations. The motivation is that, for any given extension $\mathrel{\mc R}$, the predictor knows $\mathrel{\mc R}_{{\mathds D}}$ and freely chooses $\mathrel{\mc R}_{\mathfrak f}$, but the rankings at other databases are somewhat arbitrary. For although the rankings the predictor associates with members of ${\mathds D^{\mathfrak f}}\bs ({\mathds D} \cup \{\mathfrak f\}) $ are constrained, what matters is not the actual rankings, but rather their potential for consistency with the axioms.

definition*Let $\mathrel{\mc R}$ and $\mathrel{\acute{\mathrel{\mathcal R}}}$ be extensions of $\mathbin{\preceq}_{{\mathds D}}$. $\mathrel{\acute{\mathrel{\mathcal R}}}$ is a perturbation of $\mathrel{\mc R}$ if $\mathbin{\acute{\mathbin{\mathcal R}}}_{\mathfrak f } = \mathbin{\mc R}_{\mathfrak f}$. Moreover, $\mathrel{\acute{\mathrel{\mathcal R}}}$ is a nondogmatic perturbation if $\mathbin{\sharp}\hskip1pt \textup{total} (\mathbin{\acute{\mathbin{\mathcal R}}}) \geq \mathbin{\sharp}\hskip1pt \textup{total} ( \mathbin{\mc R})$.

We are interested in perturbations that are nondogmatic because may reveal intransitivities that the predictor's inexperience conceals. The prudent predictor may exploit them and revise her model before new cases arrive.

prudence*For every $Y\subseteq X$ with cardinality $ 3$ or $ 4 $, every testworthy $ Y $-extension of $ \mathbin{\preceq} _{ {\mathds D} }$ that is novel has a nondogmatic perturbation that satisfies (ref)–(ref).

Given that the extensions we consider are nonrevisionistic, it is natural to ask whether (ref)--(ref) are superfluous in the presence of 4-prudence. Our response is twofold. Firstly, in practice, we would expect (ref)--(ref) to hold more frequently than 4-prudence\ which is more cognitively demanding. Secondly, as we will see in the proof of the main theorem, when $ {\mathds {T}} $ is infinite, for some $ Y \subseteq X $, the set of testworthy $ Y $-extensions that are novel may be empty. For every such $Y$, the following observation confirms that 4-diversity\ holds on $Y$.

theoremEnd[link to proof]{observation}[on testworthy extensions] Let $\preceq_{{\mathds D}}$ satisfy (ref)–(ref). For every $Y\subseteq X$ of cardinality $ 3$ or $ 4$, the set of testworthy $Y$-extensions is nonempty. If, for some $Y$, every testworthy $Y$-extension is regular, then 4-diversity\ holds on $Y$.
proofEndRespectively, these two statements follow via (ref) and (ref).

It is also natural to ask whether 4-prudence\ is simply requiring that, on $Y$ such that $\preceq_{\mathds J}$ fails to satisfy 4-diversity, there exists a testworthy $Y$-extension that is novel and satisfies 4-diversity. In the proof of the theorem that now follows, we show that this is not the case.

\paragraph{The main theorem\hskip-5pt} that now follows involves real-valued function $ \mathbf{v} $ on the product $X \times {\mathds C} $. We view $\mathbf{v}$ as a matrix and $ \mathbf{v} ( x , \cdot )$ as one of its rows. The matrix $ \mathbf{v} $ is a representation of $ \preceq _{ {\mathds D} }$ whenever it satisfies

linenomath*\begin{equation}\notag \left\{ \begin{array}{l} for every $ x , y \in X$ and every $ D \in {\mathds D} $,\\ x \preceq _ D y \quad if, and only if, \quad \sum _ { \, c \,\in\, D} \mathbf{v} ( x , c ) \leq \sum _ {\, c \,\in\, D } \mathbf{v} ( y , c ) . \end{array}\right. \end{equation}

The matrix $\mathbf{v}$ respects case equivalence (with respect to $\preceq_{{\mathds D}}$) if, for every $c,d\in {\mathds C}$, $c \sim^{\star} d$ if, and only if, the columns $\mathbf{v}(\cdot,c)$ and $ \mathbf{v}(\cdot,d)$ are equal.

theorem[Part I, Existence] Let there be given $X$, ${\mathds C^{\mathfrak f}}$, $\mathbin{\preceq}_ {\mathds D}$ and associated extensions, as above, such that the richness condition holds. Then (ref) and (ref) are equivalent. \begin{enumerate}[label=((ref).\roman*)] • (ref)--(ref) and 4-prudence\ hold for $\preceq_{{\mathds D}}$. • There exists a matrix $ \mathbf{v} : X \times {\mathds C} \rightarrow \R $ satisfying (ref) and (ref)$\,:$ \begin{enumerate}[label=((ref).\alph*)] • $ \mathbf{v} $ is a representation of $ \preceq _ { {\mathds D} }$ that respects case equivalence$\,;$ • no row of $\mathbf{v}$ is dominated by any other row, and for every three distinct elements $x,y, z \in X$ and $\lambda \in \R$, $\mathbf{v} (x, \cdot) \neq \lambda \mathbf{v}(y,\cdot) + (1-\lambda) \mathbf{v}(z,\cdot)$\,.\footnote{Observe that $\mathbf{v}(x,\cdot)- \mathbf{v}(z,\cdot)$ and $\mathbf{v}(y,\cdot)-\mathbf{v}(z,\cdot)$ are noncollinear if, and only if, the affine independence condition of (ref) holds.} \end{enumerate} \end{enumerate}

Our uniqueness result is identical to that of $\textup{[GS]}$. \setcounter{theorem}{0}

theorem[Part II, Uniqueness] If (ref) $[$or (ref)$]$ holds, then the matrix $ \mathbf{v} $ is unique in the following sense$\,:$ for every other matrix $ \mathbf{u} : X \times {\mathds C} \rightarrow \R $ that represents $\mathbin{\preceq}_{{\mathds D}}$, there is a scalar $ \lambda > 0 $ and a matrix $ \beta : X \times {\mathds C} \rightarrow \R$ with identical rows (\ie\ with constant columns) such that $ \mathbf{u} = \lambda \mathbf{v} + \beta$.

The proof of (ref) appears in online appendix (ref). It relies upon a translation from the abstract database/memory set up of the model to the setting of rational vectors similar to $\textup{[GS]}$. We show that the translated theorem (ref) (see (ref)) is equivalent to (ref). The proof of (ref) is then the subject of online appendix (ref).

\paragraph{We appeal to the following corollary}\hskip-7pt when we connect the present framework to the concept of arbitrage in finance. Central to this connection is

definitionFor $ Y \in 2 ^ { X } $, the matrix $ v^{ {(\cdot,\cdot)} } : Y^{2}\times {\mathds C} \rightarrow \R $ satisfies the Jacobi identity whenever, for every $ x , y , z \in Y $, the rows satisfy $ v ^{ {(x,z)} } = v ^{ {(x, y)} } + v ^{ {(y,z)} }$.

In (ref), of (ref), we show that, for a given extension $\mathrel{\mc R}$, (ref)--(ref) and 2-diversity\ yield a pairwise representation $v^{{(\cdot,\cdot)}}$ of $ \mathrel{\mc R} $.\footnote{That is, for every $x,y\in X$ and $D \in {\mathds D}$, $x \mathbin{\mc R}_{D} y$ if, and only if $ \sum_{c \,\in D}v^{{(x, y)}}(c) \geq 0$.}

theoremEnd[link to proof]{corollary}[a characterisation of prudence] Let the number of case types be finite and let $\preceq_{{\mathds D}}$ satisfy (ref). Then $\preceq_{{\mathds D}}$ satisfies 4-prudence\ if, and only if, $\preceq_{{\mathds D}}$ has a pairwise representation $v^{{(\cdot,\cdot)}}$ that satisfies the Jacobi identity. Moreover, for every other pairwise representation $u^{{(\cdot,\cdot)}} $, there exists $\lambda >0$, such that $u^{{(\cdot,\cdot)}} = \lambda v^{{(\cdot,\cdot)}}$.
proofEndThis follows from (ref), (ref) and the fact that, via (ref), 4-prudence\ implies (ref)--(ref) when the number of case types is finite.

Discussion

We begin by restating the existence part of the main theorem of $\textup{[GS]}$.

theorem*[Existence] Let there be given $X$, $ {\mathds C} $ and $\mathbin{\preceq}_ {\mathds D}$, as above, such that the richness condition holds. Then (ref) and (ref) are equivalent. \begin{enumerate}[label=(\roman*)] • (ref)--(ref) and 4-diversity hold for $\preceq_{{\mathds D}}$. • There exists a matrix $ \mathbf{v} : X \times {\mathds C} \rightarrow \R $ satisfying (ref) and (ref)$\,:$ \begin{enumerate}[label=(\alph*)] • $ \mathbf{v} $ is a representation of $ \preceq _ { {\mathds D} }$ that respects case equivalence$\,;$ • if $\mathbin{\sharp}\hskip1pt X < 4$, then no row is dominated by an affine combination of the other rows, and for every four distinct elements $x,y,z,w \in X$ and every $\lambda , \mu, \theta \in \R$ such that $ \lambda +\mu + \theta = 1$, $\mathbf{v}(x,\cdot ) \not \leq \lambda \mathbf{v}(y,\cdot )+\mu \mathbf{v}(z,\cdot)+ \theta \mathbf{v}(w,\cdot)$. \end{enumerate} \end{enumerate}

Although diversity axioms play an important technical role, they are not obviously behavioural. Instead, diversity axioms impose restrictions on what is beyond the predictor's control and on what is central to inductive inference: experience. Our main contention is that $ {C^\star} $ may not be so rich as to support $\mathbin{\preceq}_{{\mathds D}}$ satisfying 4-diversity. That is to say, there may exist $ Y \subseteq X $ such that $ \mathbin{\sharp}\hskip1pt Y = 4$, and such that the data is insufficiently rich to support all $ 4 ! = 24 $ strict rankings. A casual comparison of condition (ref) and (ref) confirms that the present framework achieves the main purpose of accommodating the less experienced.

For the remainder of this discussion, we take both $X$ and the set ${\mathds {T}}$ of case types to be finite and of cardinality $m$ and $n$ respectively. Via case equivalence, for any $\preceq_{{\mathds D}}$ (satisfying (ref) or (ref)) we may efficiently summarise rankings using a real-valued $m\times n$ matrix $\mathbf{v}$ on the product of $X$ and case types $ {\mathds {T}}$.

\paragraph{Comparing the complexity of conditional-2-diversity\ and 4-diversity,}\hskip-8pt in the presence of the other axioms, provides a measure of the value of experience. As an estimate, we compare condition (ref) with (ref) of Gilboa and Schmeidler's theorem. Verifying (ref) involves checking $n$ affine dominance constraints: one for each case type. This is well-known to be equivalent to the complexity of linear programming with real variables DR-Linear_programming. In the absence of knowledge regarding the sparsity of $\mathbf v$, the fastest algorithm for achieving this is of order $n^{3}$ LS_Linear_programming. Since this holds for every subset of four distinct eventualities, checking (ref) is of order $\binom{m}{4}n^{3}$. In contrast, verifying $\mathbf{v}(x,\cdot )\not \leq \mathbf{v}(y,\cdot)$ takes at most $n$ steps for each of the $\binom{m}{2}$ subsets of $2$ distinct elements. Likewise, for every three distinct elements $x,y,z$ in $X$, checking for noncollinearity of two vectors takes at most $n$ steps. Thus, a na\"{i}ve algorithm for checking (ref) is of order $\binom{m}{3}n$. Even for $m = 4 $ and $n = 4$, the difference is stark: $\binom{m}{3}n = 16$ versus $ \binom{m}{4} n^3 = 64$. (The threshold $4$ is important as we show in (ref) below.)

\paragraph{A robust test of prudence\hskip-7pt} is available precisely when (ref)--(ref) hold and 4-diversity\ fails. For, when $n<\infty$, the existence of a similarity representation satisfying (ref) is equivalent to 4-prudence. The test is robust in that, generically, predictors that fail to check for the arrival of new cases also fail to satisfy (ref). This is because, the row differences $v^{(x, y)} = - \mathbf{v}(x,\cdot) + \mathbf{v}(y,\cdot)$, the Jacobi identity $v^{{(x,z)}} = v^{{(x, y)}} + v^{{(y,z)}}$ holds for every $x,y,z \in X$. This in turn implies that the hyperplanes $H^{\{x,y\}}, H^{\{y,z\}}$ and $H^{\{x,z\}}$, to which the row differences are normal, are congruent. Congruence implies the hyperplanes are not in general position. (When $n = 2$, the same argument can be made in terms of extensions: see (ref) of the proof of (ref).) In other words, there is a zero (conditional) probability of the predictor striking lucky and appearing to be prudent when in fact they are not engaging in this form of second order inductive inference.

\paragraph{Novelty, memory and transitivity.} An early test and taxonomy of intransitive behaviour is due to Weinstein-Intransitivity,Weinstein-Transitivity. Weinstein points out that intransivity can sometimes be rational in complex situations. Interestingly, he shows that young are people significantly less transitive in their choices. He also points out that the law itself is designed to accommodate irresponsible under-age decisions. In the psychology literature, BR-Novelty_and_intransitivity show that presenting novel objects is more likely to trigger intransitive choices.

More recently, Enkavi-Hippocampal_dependence provide evidence that people with a specific form of memory impairment (lesions in the hippocampus of the brain) are significantly more likely to violate transitivity in pairwise choices of chocolate bars: even though they rank numbers transitively.\footnote{The hippocampus is associated with learning and memory. \Citet{Hassabis-Hippocampal_dependence} and Schacter-Hippocampal_dependence present evidence showing that the hippocampus plays an important role in imagining future experiences on the basis of past ones. \Citet{Enkavi-Hippocampal_dependence} go further by showing that it plays a role in the value-based decision making framework of Rangel-Value-based_neurobiology.} Whilst this literature does not yet offer a direct test of the present model, it supports the case-based framework of constructing preference as well as our premise that violations of transitivity are often driven by novelty, or equivalently impaired memory.

In line with the case-based approach, the experimental evidence of Enkavi-Hippocampal_dependence suggests that agents are constructing their preferences on the basis of past experience. Moreover, it seems natural to interpret agents with hippocampal impairment as inexperienced predictors. The fact that impaired agents then make intransitive decisions is very much in line with what our model predicts as they are, in effect, facing a novel situation and are required to construct their preferences on the fly. It appears that impaired agents are also failing to be prudent, though it is not obvious that chocolates warrant the additional neural computation that accompanies 4-prudence.

remark*A closer look at the relationship between the proportion $\rho$ of hippocampal impairment and the percentage $\sigma$ of intransitive choices Enkavi-Hippocampal_dependence suggests another interpretation. For $\rho $ above $ \frac{1}{4}$, $\sigma $ is above $ 20\%$$:$ twice as high as it is for $\rho < \frac{1}{4}$. Memories and the rankings they generate are latent variables to the observable $\rho$ and this threshold is where conditional-2-diversity\ fails to hold and our model breaks down. The $16$ cases in the data are thus partitioned into three groups$:$ two that satisfy 3-diversity\ ($\rho < 0.05 $ and $\sigma< 5\%$)$;$ $12$ intermediate cases that satisfy conditional-2-diversity\ ($0.05 \leq \rho \leq 0.25 $ and $5\%\leq \sigma \leq 10\%$)$;$ and $2$ severe cases that fail to satisfy \textup{conditional-\textit{2}-diversity}\ ($0.25 < \rho$ and $10\% < \sigma$).

\paragraph{The success or failure of startups. \hskip-7pt} Inexperience raises significant barriers to entry. Overcoming these barriers is either the result of making mistakes and learning by doing “on the fly” or the result of being prudent. Which form of second-order induction bears out in practice will depend on many factors.

The following proposition confirms our thesis that experienced predictors (\ie\ those that satisfy 4-diversity) have indeed encountered a high number of case types. As the main theorem of $\textup{[GS]}$\ shows, experienced predictors have no need for the additional structure of extensions. Unless the prediction problem changes (\eg\ new eventualities become relevant), they have no need to engage in second order induction. This saving in cognitive effort is the prize that experience confers.

theoremEnd{proposition}[experience and case types] If $\mathbin{\preceq}_{{\mathds D}}$ satisfies (ref)--(ref) and 4-diversity, then $ n \geq \min \{4, m\} $, and, for every $Y \subseteq X$ of cardinality $m'$ and regular $Y$-extension $\mathrel{\mc R} $, the number $n'$ of equivalence classes of $\sim^{\mathrel{\mc R}}$ satisfies $ n' \geq \min \{4,m'\}$.
proofEndVia (ref), (ref)--(ref) and 4-diversity\ hold for $\preceq_{{\mathds D}}$ if and only if (ref)--(ref) and 4-diversity\ hold for $\preceq_{\mathds J}$. Let $Y\subseteq X$ be of cardinality $m' = 1, 2, 3$ or $4$ and let $\mathrel{\mc R}$ be a regular $Y$-extension. Via (ref), there exists a pairwise representation $v^{{(\cdot,\cdot)}} $. For $m' = 1$, $n' = 1 $ because $\mathrel{\mc R}_{J}$ is constant on ${\mathds {J}^{\mathfrak f}}$. For $m' = 2$, $n' \geq 2$, since via part (ref) of (ref), $G^{{(x, y)}}$ and $G^{{(y, x)}}$ are both nonempty. By way of contradiction, first suppose $n' = 2$ and $m' \geq 3$. Via (ref) of appendix (ref), $\textup{total}(\mathrel{\mc R}) \leq 4$. In contrast, 4-diversity\ requires $\textup{total}(\mathrel{\mc R}) = 6$. The remaining case is where $n' = 3$ and $m' \geq 4$. If the rank $\mathbf r$ of $v^{{(\cdot,\cdot)}}$ satisfies $\mathbf r \geq 3$, then the kernel $A^{Y}$ of $v^{{(\cdot,\cdot)}}$ is zero-dimensional. Then $0$ is the unique element of $ A^{Y}$. Thus, the positive kernel $A^{Y}_{\text{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\srcsize}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}$+\mkern-2mu+$}}$ of $v^{{(\cdot,\cdot)}}$ is empty. Then Zaslavski's theorem implies that $\textup{total}(\mathrel{\mc R})< 4!$, so that \textit{4}-\textup{diversity}\ fails to hold. If $\mathbf r \leq 2$, then an application of the rank version Zaslavski's theorem (in particular (ref) with $\acute{\mathbf r } = \mathbf r = 2$) yields \begin{linenomath*} \[ \lvert \mc G_{\text{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\srcsize}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}$+\mkern-2mu+$}} \rvert \leq 1 - 6 + 15 - 20 + 15 + 6 + 1 = 12. \] \end{linenomath*} Thus, once again \textit{4}-\textup{diversity}\ fails to hold. Thus $n' \geq \min \{4, m'\}$, as required. Finally, since $Y\subseteq X$, $m\geq m'$, and, via part (ref) of (ref), $n \geq n'$.

In contrast, conditional-2-diversity\ implies no restrictions on the cardinality of ${\mathds {T}}$ beyond $n \geq 2$ and this is also a virtue of 2-diversity.

\paragraph{When is prudence worth the trouble?} The simple answer to this question is: when revising a model “on the fly”, once a novel case arrives, is costly. The following is our main example of such a setting.

exampleConsider a fair market maker of zero-coupon (treasury) bonds.\footnote{Similar to a fair insurer, the fair market maker sets the market spread to zero.} The compound-interest formula for the accumulation process of such a bond is \begin{linenomath*} \[a^{{(x, y)}}= \left(1+r^{\{x,y\}}\right)^{-x + y}, \] \end{linenomath*} where $r^{\{x,y\}}$ is the implied yield on a forward contract that accrues interest between dates $x$ and $y$. If $ x$ is later than $y$, then the contract is to sell, and the market maker pays this yield, so that $r^{\{y,x\}} = r^{\{x,y\}}$. Let $X \subseteq \R_{\text{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\srcsize}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}$+$}}$ index a suitable sequence of trading dates with $0 \in X$ being the spot date. It is well known that the (normalised) spot bond price $x \mapsto b(x) = \left(1+r^{\{0,x\}}\right)^{-x} $ is arbitrage-free if, and only if, the log-accumulation process satisfies \begin{linenomath*} \begin{equation} for every $x,y,z \in X$,\quad \log a^{{(x, y)}} = \log a^{{(x,z)}} + \log a^{{(z,y)}} \,.\footnote{To see this, suppose that, for some $x< y$, she sets $ a^{{(x, y)}} < a^{{(x,z)}}a^{{(z,y)}}$. Another trader would do well to sell the forward contract ${(x, y)}$, buy the spot contract ${(x,z)}$ and sell the spot contract ${(z,y)}$. A risk-free arbitrage opportunity is also available if the reverse inequality holds.} \end{equation} \end{linenomath*} This no-arbitrage condition is a special case of the Jacobi identity. We now explain how a market maker might infer the accumulation process from past cases. Let each case $c \in D^{\star}$ consist of market-relevant data from a given time interval (a block of time periods) in the past. These blocks are chosen so that, for any finite resampling $D $ of cases from $D^{\star}$, the sequence of cases that makes up $ D $ is exchangeable. The free case $\mathfrak f $ has no additional structure beyond that of (ref). Next, for every finite resampling $D$ and every date $x$ and $y$, let $x \preceq_{D} y$ if, and only if, in answer to the question “At which date will the price be higher?” the market maker finds that $y$ is more plausible than $x$. For the following corollary, we introduce the empirical implied yield function that maps triples $x \times y \times c $ to $ r^{\{x,y\}}_{c} \in \R$. This function is characterised by three conditions: for time intervals of length zero, the yield is zero; fair pricing; and case equivalence. These are respectively formalised as follows: for every $x,y\in X$ and every $c , d\in {\mathds C}$, $ r^{\{x,x\}}_{c} = 0$; $r^{\{y,x\}}_{c} = r^{\{x,y\}}_{c}$; and $c \sim^{\star} d$ if, and only if, $r^{\{x,y\}}_{c} = r^{\{x,y\}}_{d}$. The empirical bond price function $B : X \times {\mathds D} \rightarrow \R$ maps every pair $x \times D$ to \[ B(x,D) = \prod_{c\, \in D}\left( 1 + r^{\{0,x\}}_{c} \right)^{-\frac{x}{ \lvert D \rvert}}. \] When $D^{\star}$ belongs to ${\mathds D}$, the number of case types is finite, and we have \begin{corollary} Let $\preceq_{{\mathds D}}$ satisfy (ref). Then $\preceq_{{\mathds D}}$ satisfies 4-prudence\ if, and only if, there exists empirical implied yield and empirical bond price functions, such that \begin{linenomath*} \begin{equation}\tag{$*$} \left\{ \begin{array}{l} for every $ x , y \in X$ and every $ D \in {\mathds D} $,\\ x \preceq_{D} y \quad if, and only if,\quad B(x,D) \leq B(y,D). \end{array}\right. \end{equation} \end{linenomath*} Moreover, for every $D\in {\mathds D}$, the bond price $B(\cdot,D)$ is arbitrage-free. \end{corollary} \begin{proof}[Proof of (ref)] Via (ref), there exists a pairwise Jacobi representation $v^{{(\cdot,\cdot)}}: X^{2} \times {\mathds C} \rightarrow \R$ such that, for every $c,d\in {\mathds C}$, $v^{{(\cdot,\cdot)}}(c) = v^{{(\cdot,\cdot)}}(d)$ if, and only if, $c \sim^{\star} d$. Moreover, $v^{{(\cdot,\cdot)}}$ is unique upto multiplication by a positive scalar. Then, for every $x, y, z \in X$, $v^{(x,x)}(\cdot) = 0$, $v^{{(y, x)}} = - v^{{(x, y)}}$, and $v^{{(x, y)}} = v^{{(x,z)}} + v^{{(z,y)}}$. Recalling that $0 \in X$, for every $x\in X$ and $D \in {\mathds D}$, let \begin{linenomath*} \begin{equation} \textstyle B(x,D) = \exp\left( -\frac{1}{\lvert D \rvert} \sum_{c\,\in D} v^{(x,0)}_{c} \right). \end{equation} \end{linenomath*} For the proof of (ref), recall that $x \preceq_{D} y$ if, and only if, $\sum_{c\,\in D} v^{{(x, y)}}(c) \geq 0$. Since $- v^{(y,0)}_{c} = v^{(0,y)}_{c}$ and, via the Jacobi identity, $v^{(x,0)} + v^{(0,y)} = v^{{(x, y)}}$, we have: \begin{linenomath*} \[\textstyle - \log B(x,D) + \log B(y,D) = \frac{1}{\lvert D\rvert} \sum_{c\,\in D} \left(v^{(x,0)}_{c} - v^{(y,0)}_{c}\right) = \frac{1}{\lvert D\rvert} \sum_{c\,\in D} v^{{(x, y)}}_{c}. \] \end{linenomath*} It remains for us to confirm that the bond price is a suitable function of the empirical yield function. For every $x,y \in X$ and $c \in {\mathds C}$, $ \log a^{{(x, y)}}_{c} = -v^{{(x, y)}}(c)$: so that, as the solution to $ (y - x )\log (1+r^{\{x,y\}}_{c}) = -v^{{(x, y)}}_{c} = v^{{(y, x)}}_{c}$, for $x \neq y$, \begin{linenomath*} \begin{equation} 1+ r^{\{x,y\}}_{c} = \exp\left(\frac{v^{{(y, x)}}_{c}}{y-x}\right) = \exp\left( \frac{v^{{(x, y)}}_{c}}{x-y}\right) = 1+r^{\{y,x\}}_{c}. \end{equation} \end{linenomath*} We therefore observe that $r^{\{x,y\}}_{c} = r^{\{y,x\}}_{c}$. For $x = y$, $v^{{(x, y)}}_{c}= 0$ ensures that we can take $r^{\{x,x\}}_{c} = 0$. Finally, note that for $c \sim^{\star} d$, the property $r^{\{x,y\}}_{c} = r^{\{x,y\}}_{c}$ is inherited from $v^{{(x, y)}}_{c} = v^{{(x, y)}}_{d}$, so that we have an empirical implied yield function. The fact that, for every $D$, $B(\cdot, D)$ is arbitrage-free follows by virtue of the fact that $v^{{(\cdot,\cdot)}}$ satisfies the Jacobi identity. \end{proof} In the present setting, because we are modelling a normalised bond price, we obtain a stronger uniqueness result relative to part II of (ref). If $\tilde B$ is another function that satisfies the present corollary, then, for every $D \in {\mathds D}$, the spot price $\tilde B(0,D) = 1$. Thus, via (ref) and (ref), for some $\lambda >0$, $\tilde B = \textup{e}^{-\lambda} B$. We now point out an interesting implication of conditional-\textit{2}-diversity\ the recent prevalence of negative interest rates. \Wlog, fix $x<y$. Then, given \textit{4}-\textup{prudence}, \textit{2}-\textup{diversity}\ implies that there exists $c,d \in {\mathds C}$ such that $v^{{(x, y)}}(c) < 0 < v^{{(x, y)}}(d)$. This is equivalent to $v^{{(y, x)}}(d) < 0 < v^{{(y, x)}}(c)$, and, via (ref), $r^{\{x,y\}}_{d} < 0 < r^{\{x,y\}}_{c}$. That is, \textit{2}-\textup{diversity}\ requires that the market maker's data is rich enough to contain at least one case where the yield between date $x$ and $y$ is negative (as well as one where it is positive). \textup{Conditional-\textit{2}-diversity}\ extends this notion to require that $r^{\{x,y\}}_{D} < 0 < r^{\{x,y\}}_{C}$ for some $C$ and $D$ such that $r^{\{x,z\}}_{C}\cdot r^{\{x,z\}}_{D} >0$.

\paragraph{Discussion of second-order induction.}\vskip-8pt The market maker of (ref) engages in second-order induction when she acts prudently. She reflects on her model by checking that the basic axioms of $\textup{[GS]}$\ will continue to hold when a novel case arrives. By way of contrast, suppose the bond price of the market maker is such that $\preceq_{{\mathds D}}$ is consistent with the basic axioms, but not 4-prudence. Then when a novel case arrives, she may be exposed to arbitrage and need to respecify her entire model “on the fly”. Such a step corresponds to the intermittent respecification of her similarity weighting function $\mathbf{v}(x,c)$, that AG-Second-order_induction describe. In AG-Second-order_induction, the “leave-one-out” technique of cross-validating the model by omitting a case of each type is intuitively and operationally close to our inclusion of the free case $\mathfrak f$. The difference is that by allowing $\mathfrak f $ more degrees of freedom, our market maker can study novel extensions and peer into the future through the lens of her current model. She can exploit the intervals of time inbetween the arrival of novel cases by continuously engaging in second-order induction.

Through an example, we now show that the present framework provides the flexibility to accommodate second-order induction without sacrificing the computational or normative advantages that additive similarity functions provide.

example*[second-order induction, $\textup{[GS]}$, p.12] Let $c$ denote a case where Mary chooses restaurant $x$ over restaurant $y$. In the absence of any further information, it is tempting to assume some similarity between John and Mary. The predictor then finds it plausible that John prefers $x$ to $y$ given $\{c\}$. A separate database $D$ contains no choices between $x$ and $y$. Thus, in the absence of further information, $x$ and $y$ appear equally likely based on $D$. Additivity of the similarity function (or (ref))) implies it is plausible that John prefers $x$ to $y$ given $\{c\} \cup D$. The violation of (ref) arises when a more careful examination of the contents of $D$ reveals many choices between other pairs of restaurants where John and Mary consistently differ.

Quine's notion of perceptual similarity Quine-Roots_of_reference offers a check on the predictor's inference about John's choice given $\{c\}$. John and Mary may just as well be two drivers passing through an intersection at different times. Although their situations are broadly speaking very similar, if one faces a red light and the other a green light, their responses will differ. In the restaurant setting, some pivotal information is omitted from $c$. Observing that the evidence in $c$ in favour of $x$ over $y$ is somewhat weak, a prudent predictor instead recasts $\{c\}$ as a database $C$ that combines past observations with copies of the pivotal novel case $\mathfrak f$. With a more refined model, that explicitly allows for omitted variables, the predictor can check to see if her model extends to higher dimensions without violating the basic axioms.

remarkObserve that predictors that only fail to satisfy (ref) can, with some additional regularity conditions, still be represented by a nonlinear function $u: X \times {\mathds D} \rightarrow \R $ such that for every $D \in {\mathds D}$ and every $x,y \in X $, $x \preceq_{D} y$ if, and only if, $ u(x,D)\leq u(y,D)$ OCallaghan-Parametric_continuity. But predictors that satisfy (ref)--(ref) but not 4-prudence\ can also be represented by such a function: because $\preceq_{{\mathds D}}$ is complete and transitive for each $D \in {\mathds D}$. The present framework allows us to disentangle the latter kind of predictor from those who, for good reason, fail to satisfy the combination axiom. (See $\textup{[GS]}$\ for examples of such reasons.)

\paragraph{On the veracity of false news.} How should a predictor check whether her model consistently extends to higher dimensions (when novel cases arrive)? If our model is a guide then, the most useful rankings that she might wish to examine are those that are far from her own. This is because our definition of testworthy extensions involves assigning to the novel case $\mathfrak f$ the inverse of some total ranking $\preceq_{D}$. This may offer some rationale for why information that differs from our own is intrinsically valuable. Testworthy extensions play a vital role in taming the complexity of our proof. It seems plausible that something similar may be at play when agents encounter radically different information from their own on social media: even if it is fake. This may help to explain why false news is significantly more veracious than real news online Vosoughi-Roy-Aral-Veracity. The fact that real news is typically closer to what we have observed in the past means that it is of less value to the prudent predictor that finds it costly to imagine worlds that are far from her own.

\makeatletter \def\@seccntformat#1{Appendix\,\csname the#1\endcsname.\quad} \makeatother