Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Second-order Inductive Inference: an axiomatic approach
\pagenumbering{gobble}
\pagestyle{fancy}
\thispagestyle{plain}
abstractConsider a predictor who ranks eventualities on the basis of past cases: for
instance a search engine ranking webpages given past searches. Resampling past
cases leads to different rankings and the extraction of deeper
information. Yet a rich database, with sufficiently diverse rankings, is often
beyond reach. Inexperience demands either “on the fly” learning-by-doing or
prudence: the arrival of a novel case does not force (i) a revision of
current rankings, (ii) dogmatism towards new rankings, or (iii)
intransitivity.
For this higher-order framework of inductive inference, we derive a suitably
unique numerical representation of these rankings via a matrix on
eventualities $\times$ cases and describe a robust test of
prudence. Applications include: the success/failure of startups; the veracity
of fake news; and novel conditions for the existence of a yield curve that is
robustly arbitrage-free.
{11.5cm}
\epigraph{From the past, the present acts prudently, lest it spoil future
action.}{Titian: Allegory of Prudence
}
Introduction
Experience is the basis of prediction. Yet outside of stylised settings, even
the most experienced forecasters do not claim access to “the full
model” that is closed with respect to the relevant states of the
world.
exampleThe canonical large-world setting is that of the global financial markets. In
2019, all models ommited details of COVID-19. Similarly, in 2007, a clear
description of the sub-prime mortgage crisis was beyond reach. The central
role of simulation and bootstrap methods in empirical finance points to a
prevalence of inductive reasoning.\footnote{Consider
Cowles-Forecasting, White-Data_snooping,
FF-Luck_vs_skill and HL-Lucky_factors.} That is, to
forecasting on the basis of past cases as opposed to a full description of
future states. At the same time, market makers need to set prices that
are robust to changes that open the door to exploitation via arbitrage.
\pagenumbering{arabic} \setcounter{page}{2} Incomplete models are a key
motivation for recent axiomatic updates to the standard Bayesian framework such
as “reverse Bayesianism” of KV_Reverse_Bayes. This and other work
(KV-Awareness_of_U, HR-Knowledge_of_U and
GKMQT_Robust_experiments) on unawareness and robustness in the
state-space sense provide part of our inspiration for the present upgrade to the
axiomatic foundations of inductive inference in GS_Inductive_inference. By taking rankings as primitive, we extend
$\textup{[GS]}$\ to accommodate a forward-looking version of the second-order induction of
AG-Second-order_induction.
In the present framework, the basic building blocks of the model are
observations or (synonymously) past cases. A given past case may be empirical or
theoretical, and the predictor's model is naturally bounded in size and scope by
the predictor's experience. Our contribution is to extend the framework of
$\textup{[GS]}$\ to model less experienced predictors that are prudent. That is to agents
that are able to proactively engage in second-order induction: by “pulling
themselves up by the bootstraps” and looking at how their model extends to
novel cases. Thus allowing them to survive their initial phase of inexperience.
In the remainder of this section, we informally introduce: the model of
(ref); the axioms and matrix representation of
(ref); and the applications to which we return in
(ref). Proofs of main results appear in appendices
(ref) to (ref).
\paragraph{Synopsis of model and results with examples.} The predictor is endowed with a
qualitative plausibility ranking of eventualities given her current database
$D^{\star}$ of past cases. Moreover, the same is true for every finite
resampling of $D^{\star}$. We identify conditions on the resulting family of
(ordinal) rankings for the existence of a suitably unique
real-valued matrix $\mathbf v$ on eventualities$\times$cases that represents the
information in these rankings. The form of this representation is linear on
databases (\ie\ additive over cases) and separable on eventualities, so that for
every database $D$, eventuality $y$ is
more likely than $x$ if, and only if,
linenomath*\begin{equation}
\sum_{c\,\in D} \mathbf v(x,c) \leq \sum_{c\,\in D} \mathbf v(y,c).
\end{equation}
The similarity weight $\mathbf{v}(x,c)$ is the degree of support that case $c$
lends to $x$.
exampleAs a canonical example, consider predicting the slope coefficient
$\beta_{i}$ from the regression of asset $i$'s returns on the market portfolio
FM-Two_pass. Then values $\beta_{i}$ are
eventualities and $D^{*}$ is the current sample of past returns. In this
setting, $\mathbf v$ is an empirical log-likelihood function that we generate
via a generalised notion of bootstrapping of cases in $D^{\star}$.
Key to the contribution of $\textup{[GS]}$\ is an endogenous notion of case types: a
partition of cases according to the marginal information they contribute to a
given database. This marginal contribution is measured in terms of the impact on
the rankings of eventualities.\footnote{This notion of case type is therefore
close to Quine's “perceptual similarity” (see (ref)).} Case
types are analogous to states in the sense that they form the model's
dimensions: the lower the dimension the less the experience.
example[Search Engine Results Page, SERP]
Advertising aside, when users conduct a web search, the search engine compiles
a ranking $\mathbin{\preceq}_{D^{\star}}$ of (web)pages $x, y, z, \dots$ on the basis
of its database $D^{\star }$ of past cases: searches of past users plus
feedback from subsequent clicks. Resampling yields other databases $D$ and
other rankings. At one extreme, past cases may be so similar to one another
that the same plausibility ranking arises regardless of how the data is
resampled. At the other, past cases may be sufficiently rich that resampling
yields every feasible plausibility ranking of eventualities.
A rich set of past cases is at the heart of the diversity axiom of $\textup{[GS]}$--which
we refer to as 4-diversity. This restricts the model to predictors whose current
data is sufficiently rich that a resampling exercise generates all
$4!=24 $ strict (\ie\ total) rankings of every subset of four
eventualities. In this paper, we accommodate less inexperienced predictors by
only replacing 4-diversity\ with conditional-\textit{2}-diversity. Given the other basic axioms of
$\textup{[GS]}$, \textup{conditional-\textit{2}-diversity}\ turns out to be equivalent to requiring that, for every
three distinct eventualities $x$, $y$ and $z$, resampling generates at least $3$
of the $3! = 6$ possible distinct strict rankings of this triple. \textup{Conditional-\textit{2}-diversity}\
is minimal in the sense that, in its absence, we lose both existence and
uniqueness of the similarity representation (see (ref)).
To compensate for a lack of experience, we introduce a subtle and more flexible
notion of cases that allows us to capture the predictor's potential awareness of
her limited experience. Formally, related notions in the literature go by the
name of unforeseen consequences in GQ-Surprises and shadow propositions
in the setting dynamic awareness in HP-Dynamic_awareness. Via a
content-free case $\mathfrak f $, the predictor can explore the impact on her model
of the arrival of a novel case type. In effect, this involves meta-analysis of
how her similarity function will evolve over time. Hence the reference to
second-order induction.
The prudent predictor ensures her model is robust to the arrival of novel case
types. She ensures that, when a novel case arrives, she can accept the ranking
it generates without finding herself in the potentially costly position of
generating intransitive rankings when she combines past and novel cases.
example[Second-order inductive inference]
Consider a search-engine startup seeking to establish itself in the face of
incumbents with the experience of Google. The start-up engages in second order
induction when it is learning the similarity function itself. This may include
learning the values $\mathbf{v}(x,c)$ of (ref), but it may
also involve costly updates of the model structure “on the fly”, \eg\
redefining case types, rankings, \etc. The startup is prudent if, ex
ante, it structures its model to ensure that it is relatively costless to
extend to novel case types$:$ once they arrive.
Prudence is only worthwhile when revisions of the predictor's model `on the fly',
once the novel case arrives, is costly.
Consider “zero-day attacks” in the setting of cyber security.
\Citet{Hota_et_al-Cyber_security} highlight the essence of time when a novel
attack on a computer network arrives; and that such attacks are novel precisely
because cyber-security experts have already built in solutions to known
vulnerabilities. Also, tradeoffs between time, cost and learning are nowhere
more important than in finance. For a bond-market setting, we are able to
provide a formal equivalence between prudence and arbitrage pricing in
(ref). In (ref), we also discuss: other
applications; empirical evidence linking intransitivity, memories and novelty;
and connections with the literature on second-order induction in more detail.
Model
Following $\textup{[GS]}$, let the nonempty set $X$ denote the conceivable
eventualities of the present prediction problem and let $ \operatorname{rel}(X)$
denote the set of binary relations or rankings on $X$. For instance, for a
search engine, we identify webpage $x$ with the eventuality “page $x$ is the
desired webpage”. The predictor is equipped with her current memory
${C^\star}$: the union of a finite set of past cases ${D^\star}$ and a
variable or free case $\mathfrak f$. The cases in ${D^\star}$ collectively
represent the forecaster's relevant observations or experience. Our first and
most fundamental modification of the primitives of $\textup{[GS]}$ is the inclusion of
$\mathfrak f$ in the current memory
${C^\star}$.
remark*On a computer, a natural implementation of this setup is the
following. Take every case $c \in {C^\star} $ to consist of a pair
$p \times m:$ a pointer $p$ that references a memory location and the memory
content $m$. Each $c\in {D^\star}$ is identical to a case in the setting of
$\textup{[GS]}$. But, for $c = \mathfrak f$, there is no meaningful memory content, so $m$ is
“empty” or assigned an arbitrary null value. From another perspective,
cases in $ {D^\star}$ are constant (of arity zero) whereas $\mathfrak f$ is a variable
(of positive arity).
For every nonempty subsample $ D \subseteq {D^\star} $, the predictor has
sufficient information to determine a well-defined ranking $\mathbin{\preceq}_ D $ in
$ \operatorname{rel}(X)$. In contrast, since $\mathfrak f = \acute{p} \times \acute{m}$
has no meaningful memory content, $\mathbin{\preceq}_\mathfrak f$ is indeterminate and a
free variable in $\operatorname{rel}(X)$. The prudent predictor gains a better
understanding of her current model by assigning a ranking to $\acute{m}$ and
exploring the extensions of (ref).
Like $\textup{[GS]}$, we accommodate a forecaster that goes beyond her current memory and
includes hypothetical cases
${\mathds C}$ that she may not have experienced, but which, through reasoning,
interpolation or resampling, she can clearly describe. These hypothetical cases
are formally constant, like members of ${D^\star}$.
With case resampling and subsampling from the literature on bootstrapping
in mind, let
\[{\mathds D}\defeq \left\{ D\subseteq {\mathds C}: \mathbin{\sharp}\hskip1pt D<\infty\right\} \] denote
the set of (finite) determinate or constant databases. (These are
referred to as memories in $\textup{[GS]}$.) Like ${D^\star}$, each $D\in {\mathds D} $
contains no copies of $\mathfrak f$.
Let $[\mathfrak f] $ denote a set of copies of $\mathfrak f $. Finally, let
${\mathds C^{\mathfrak f}} \defeq {\mathds C} \cup [\mathfrak f]$ and let ${\mathfrak C}$ denote a member of
$\{{\mathds C}, {\mathds C^{\mathfrak f}}\}$. Let ${\mathds D^{\mathfrak f}} $ denote the corresponding set of all finite
subsets of $ {\mathds C^{\mathfrak f}} $ and take
linenomath*\begin{equation*}
{\mathfrak D} = \left\{
\begin{array}{ll}
{\mathds D} & if, and only if, ${\mathfrak C}= {\mathds C}$, and\\
{\mathds D^{\mathfrak f}} &otherwise.
\end{array}\right.
\end{equation*}
The predictor is endowed with a well-defined plausibility ranking $\mathbin{\preceq}_D $
in $ \operatorname{rel}(X)$ for each $D$ in ${\mathds D}$. Denote the symmetric part by
$\simeq _ D$ and asymmetric part by $ \mathbin{\prec} _ D $. In a minor departure from
$\textup{[GS]}$, the primitive of our model is a point in $\operatorname{rel}(X)^{{\mathds D}}$
linenomath*\[\mathbin{\preceq}_{{\mathds D}} \defeq \langle\mathbin{\preceq}_{D}:D\in {\mathds D}\rangle\,.\footnote{The
present approach is equivalent to taking
$\mathbin{\preceq}_{{\mathds D}} = \left\{D \times \mathbin{\preceq}_{D}: D \in
{\mathds D}\right\}$. This way we maintain pairwise distinctness of
$\mathbin{\preceq}_{C}=\mathbin{\preceq}_{D}$ such that $C \neq
D$. I thank Maxwell B. Stinchcombe for bringing this point to my attention.}
\]
For each $C$ in $ {\mathds D^{\mathfrak f}} \bs {\mathds D} $, the fact that for some
$ c \in [ \mathfrak f ] $, $ c \in C $ means that $\mathbin{\preceq}_{C} $ is indeterminate,
free variable in $\operatorname{rel}(X)$. Although, in isolation each such
$ \mathbin{\preceq}_{C} $ is free, when the axioms we introduce hold, the potential
values of the variable
$ \mathbin{\preceq} _ {\mathds D^{\mathfrak f}} \defeq \langle \mathbin{\preceq} _ C : C \in {\mathds D^{\mathfrak f}} \rangle $ are
constrained by the current values of the constant $\mathbin{\preceq}_{\mathds D}$.
\paragraph{Case types.}
As in $\textup{[GS]}$, two past cases $ c , d \in {\mathds C} $ are of the same case type
if, and only if, the marginal information of $ c $ is everywhere equal to the
marginal information of $ d $. Formally, $ c \sim ^{ \star } d $ if, and only
if, for every $ D \in {\mathds D} $ such that $ c , d \notin D $,
$ \mathbin{\preceq} _ { D \cup \{ c \} } = \mathbin{\preceq} _ { D \cup \{ d \} }$. By
observation 1 of $\textup{[GS]}$, $ \sim ^{ \star }$ is an equivalence relation on
$ {\mathds C} $ and as its collection of equivalence classes generate a partition
${\mathds {T}}$ of ${\mathds C}$.
We extend $ \sim ^{ \star } $ to $ {\mathds C^{\mathfrak f}} $ by taking $ [ \mathfrak f ]$ to be an
equivalence class of its own, so that, for every $ c \in {\mathds C} $,
$ c \nsim ^{ \star } \mathfrak f $. We let ${\mathds{T} ^ \mathfrak f }$ denote the corresponding
partition of ${\mathds C^{\mathfrak f}}$.
Like $\textup{[GS]}$, we also extend $\sim^{\star}$ to ${\mathds D^{\mathfrak f}}$ by treating databases that
contain the same number of each case type as equivalent. That is,
$C \sim ^{\star } D$ if, and only if, for every $t \in {\mathds{T} ^ \mathfrak f }$, the numbers
$\mathbin{\sharp}\hskip1pt (C \cap t)$ and $\mathbin{\sharp}\hskip1pt (D \cap t) $ of that case type coincide. To
enable a translation of each database to counting vectors
$t \mapsto \mathbin{\sharp}\hskip1pt (D \cap t) $, we impose a
assumption*For every $ t \in {\mathds{T} ^ \mathfrak f }$, there are infinitely many cases in $ t $.
Our key definition is the following.
definition$\mathbin{\mc R} \defeq \langle \mathbin{\mc R}_{D}: D\in {\mathfrak D} \rangle $
is an extension, and in particular a $Y$-extension, of $ \preceq_{{\mathds D}}$ if, for some nonempty
$ Y \subseteq X $, the following all hold$:$
\begin{enumerate}
• for every $ D\in {\mathfrak D} $, $\mathrel{\mc R}_{D} $ belongs
to $ \operatorname{rel} (Y)$,
• for every $ D \in {\mathds D}$ and every $x,y\in Y$,
$x \mathrel{\mc R}_{D} y $ if, and only if, $x \preceq_{D} y$,
• for every $D\in {\mathfrak D}$ and
every $c,d \in {\mathfrak C}\bs D$, if $ c \sim ^ \star d $ then
$ \mathbin{\mc R} _ { D \cup \{c\} } = \mathbin{\mc R} _ { D \cup \{d\}}$.
\end{enumerate}
An extension $\mathrel{\mc R}_{{\mathfrak D}}$ is proper if ${\mathfrak D} = {\mathds D^{\mathfrak f}}$ and
improper otherwise.
By assigning rankings to the free case $\mathfrak f$, (potential) extensions of
$\preceq_{{\mathds D}}$ simulate the arrival of novel information. Part
(ref) of the definition of an extension ensures that, for every
proper $Y$-extension $\mathrel{\mc R}$ and every $D\in {\mathds D^{\mathfrak f}}$, $\mathrel{\mc R}_{D}$ is a
well-defined binary relation on $Y$. For every extension $\mathrel{\mc R}_{{\mathfrak D}}$ and
every $D \in {\mathfrak D}$, let $ \mathrel{\mc I}_{D}$ and $ \mathrel{\mc P}_{D}$ respectively denote the
symmetric and asymmetric parts of $\mathrel{\mc R}_{D}$.
Part (ref) of the definition implies that, for every
$Y$-extension $\mathrel{\mc R}$ and every $D\in {\mathds D}$, $ \mathrel{\mc R}$ is simply the restriction
$\mathbin{\preceq}_{D}\cap Y^{2} $ of $\preceq_{D}$ to $Y$. We therefore refer to part
(ref) of the definition of an extension as the preservation or
nonrevision condition.
Part (ref) of the definition ensures that, for proper extensions
$\mathrel{\mc R} $, the partition ${\mathds{T} ^ \mathfrak f }$ of case types generated by $\sim^{\star}$ is at
least as fine the partition generated by the equivalence relation generated by
$ \mathrel{\mc R}$. Two cases $ c,d \in {\mathds C^{\mathfrak f}} $ are equivalent with respect to
$ \mathrel{\mc R} $, written $ c \sim^{\mathbin{\mc R}} d $, if, for every $ D \in {\mathds D} $ such
that $ c,d \notin D $, $ \mathbin{\mc R} _ { D \cup \{c\} } = \mathbin{\mc R} _ { D \cup\{d\} }$.
This notion allows us to partition the set of proper extensions as follows.
definition*A proper extension $\mathrel{\mc R} $ is either regular or novel. It is
novel whenever $[\mathfrak f]$ is a distinct equivalence class of $\sim^{{\mathrel{\mc R}}}$,
so that, for every $ c \in {\mathds C} $, $ c \nsim ^ {\mathrel{\mc R}} \mathfrak f $.
For novel extensions, $\mathfrak f $ mimmicks the potential arrival of new
information. Yet novel extensions need not feature qualitatively new rankings
(\ie\ rankings that do not feature in $\preceq_{{\mathds D}}$). In (ref),
we show that this is because a novel case type is characterised by the
quantitative notion of a similarity weight.
For every regular extension $\mathrel{\mc R}$, there exists $c \in {\mathds C}$ such that
$c \sim^{\mathbin{\mc R}} \mathfrak f$. Thus, there are as many regular extensions as there are
past case types (\ie\ $\mathbin{\sharp}\hskip1pt {\mathds {T}}$). Yet every regular $Y$-extension $\mathrel{\mc R}$
is equivalent to the unique improper $Y$-extension
$ \langle \mathbin{\preceq}_{D} \cap Y^{2}: D \in {\mathds D}\rangle$ in the sense of
observationFor every regular $Y$-extension $\mathrel{\mc R}$ and improper $Y$-extension $\mathrel{\acute{\mathrel{\mathcal R}}}$, for
every $ C \in {\mathds D^{\mathfrak f}} $, there exists $ D \in {\mathds D} $ such that
$C \sim^{\mathbin{\mc R}} D$ and $\mathbin{\mc R}_{C} = \mathbin{\acute{\mathbin{\mathcal R}}}_{D}$.\footnote{This means that
there is a canonical embedding of
$\left\{C \times \mathbin{\mc R}_{C}: C \in {\mathds D^{\mathfrak f}}\right\}$ in
$\{D \times \mathbin{\acute{\mathbin{\mathcal R}}}_{D}: D \in {\mathds D}\}$. The converse embedding follows from
the nonrevision condition of (ref).}
proof[Proof of (ref)]
Fix $Y\subseteq X$ nonempty and $\mathrel{\acute{\mathrel{\mathcal R}}}$ regular. \Wlog, take $ C \in {\mathds D^{\mathfrak f}}
\bs {\mathds D} $, so that $ C $ contains at least one copy of $ \mathfrak f
$. For any $ c \in C \cap [ \mathfrak f ] $, the fact that $ \mathrel{\acute{\mathrel{\mathcal R}}}
$ is regular implies that $ c \sim^{\mathbin{\acute{\mathbin{\mathcal R}}}} c _ 1 $ for some $ c _ 1 \in {\mathds C}
$. The richness assumption ensures that we may choose $ c _ 1
$ from the complement of $ C $. Then, since neither $c$ nor
$c_{1}$ belong to $ C _ 1 \defeq C \bs \{c\} $, $c \sim^{\mathbin{\acute{\mathbin{\mathcal R}}}}
c_{1}$ implies $ \mathbin{\acute{\mathbin{\mathcal R}}}_{C} = \mathbin{\acute{\mathbin{\mathcal R}}} _ { C _ 1 \cup \{c_{1}\} }$. If $ c
$ is the unique member of $ C \cap [ \mathfrak f
]$, then the proof is complete. Otherwise, using the fact that $ C
$ is finite, we may proceed by induction until we obtain a set $ C _ n
$ such that $ C _ n \cap [\mathfrak f ] $ is empty and $ D \defeq C _ n \cup \{ c _
1 , \dots , c _ n \} $ belongs to $ {\mathds D}
$. Part (ref) of (ref) then implies $ \mathbin{\acute{\mathbin{\mathcal R}}} _
{ D } = \mathbin{\preceq} _ { D } \cap Y^{2}$, so that, since
$\mathrel{\mc R}$ is improper, $\mathbin{\acute{\mathbin{\mathcal R}}}_{D}=\mathbin{\mc R}_{D}$. Finally, since $C \sim ^{\mathbin{\acute{\mathbin{\mathcal R}}}}
D$, $ \mathbin{\acute{\mathbin{\mathcal R}}} _ { C } = \mathbin{\acute{\mathbin{\mathcal R}}}_{D} $, as required.
Axioms and main theorem
\paragraph{The basic axioms\hskip-10pt} of $\textup{[GS]}$, which we rewrite in terms of
extensions, are the following. In each of these axioms, $\mathrel{\mc R} $ is an arbitrary
$Y$-extension of $\preceq_{{\mathds D}}$.
taggedblank{$\textup{A}0$}[Transitivity axiom for $\mathrel{\mc R}$]
For every $D \in {\mathfrak D} $, $ \mathrel{\mc R}_{D} $ is transitive.
taggedblank{$\textup{A}1$}[Completeness axiom for $\mathrel{\mc R}$]
For every $ D \in {\mathfrak D} $, $\mathrel{\mc R}_{D} $ is complete.
taggedblank{$\textup{A}2$}[Combination axiom for $\mathrel{\mc R}$]
For every disjoint $ C, D\in {\mathfrak D} $ and every $x,y\in Y$, if
$ x \mathrel{\mc R}_{C} y $ and $ x \mathrel{\mc R}_{ D } y $, then $ x \mathrel{\mc R}_{ C \cup D} y
$$;$ and if $x \mathrel{\mc P}_{C} y$ and $x \mathrel{\mc R}_{D} y$, then $x\mathrel{\mc P} _{C \cup D } y$.
taggedblank{$\textup{A}3$}[Archimedean axiom for $\mathrel{\mc R}$]
For every disjoint $ C,D \in {\mathfrak D} $ and every $x,y\in Y $, if
$x \mathrel{\mc P}_{D} y$, then there exists $ k \in \mbb Z_{\text{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\srcsize}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}$+\mkern-2mu+$}} $ such that, for every
pairwise disjoint collection $\left\{ D _ j : \text{ $ D _ j \sim ^ {\mathrel{\mc R}} D $
and $ C \cap D _ j =\emptyset$}\right\} _ 1 ^ k$ in ${\mathfrak D}$,
$ x \mathrel{\mc P}_{ C \cup D_{1} \cup \cdots \cup D_{k} } y.$
\paragraph{The diversity axioms\hskip-10pt} that now follow require that ${\mathds D}$
is sufficiently rich to support $Y$-extensions $\mathrel{\mc R}$ with a variety of
total orderings: \ie\ complete, transitive and antisymmetric
($x \mathrel{\mc R}_{D} y $ and $y \mathrel{\mc R}_{D} x$ implies $x = y$).
Let $ \textup{total} ( \mathrel{\mc R} )$ denote the set $ \left\{ R : \text{for some
$ C \in {\mathfrak D} $, $ R = \mathbin{\mc R} _{ C }$ is total}\right\}$ of of total orders that
feature in $ \mathrel{\mc R} $. For $ k = 4 $, the following axiom is a restatement of the
diversity axiom of $\textup{[GS]}$.
taggeddiv{\hskip-5pt}[$k$-Diversity axiom]
For every $ Y \subseteq X $ of cardinality $ n= 2, \dots, k $, every regular
$ Y$-extension $\mathrel{\mc R}$ of $ \mathbin{\preceq} _{ {\mathds D} }$ is such that
$\mathbin{\sharp}\hskip1pt \textup{total}(\mathrel{\mc R}) = n!$\hskip.5pt.
We say $k$-diversity holds on $Z$ if the axiom holds with $Z$ in the
place of $X$. By lowering the bar for the required number of total orders, the
following axiom substantively weakens 3-diversity\ and a fortiori 4-diversity.
taggedblank{$\textup{A}4$$'$}[Partial 3-diversity]
For every $ Y\subseteq X $ with cardinality $ n = 2 $ or $ 3 $, every regular
$X$-extension $ \mathrel{\mc R} $ of $ \preceq _{ {\mathds D} }$ is such that $ \mathbin{\sharp}\hskip1pt
\textup{total} ( \mathrel{\mc R} ) \geq n $.
(ref) of the appendix provides an example of $\preceq_{{\mathds D}}$
satisfying (ref)–(ref), but not 3-diversity. Partial-3-diversity\ is clearly
stronger than 2-diversity\ and, moreover, it plays the dual role of guaranteeing
uniqueness of the representation and allowing us to avoid restrictions on the
cardinality of $ X $. (ref) shows that \textup{partial-\textit{3}-diversity}\ is the
weakest axiom with these properties. Moreover, the observation below shows
that, in our setting, \textup{partial-\textit{3}-diversity}\ is equivalent to
taggedblank{A$4$}[Conditional-2-diversity]
For every three distinct elements $ x , y , z \in X $, within one of the sets
$ \{ D' : x \prec _{D ' } y \}$ and $ \{ D' : y \prec_{D ' } x \} $ there
exists both $C$ and $D$ such that $ z \prec _{ C } x $ and
$ x \prec _{ D } z $. If $ \mathbin{\sharp}\hskip1pt X = 2 $, then 2-diversity\ holds on $X$.
theoremEnd[link to proof]{observation}
For $\preceq_{{\mathds D}}$ satisfying (ref)--(ref), conditional-2-diversity\ and
partial-3-diversity\ are equivalent.
proofEndThis follows directly from (ref) and the construction of
$\preceq_{\mathds J}$.
\paragraph{The prudence axiom \hskip-10pt} that follows is our final requirement. It
is distinguished by the fact that it imposes structure on novel extensions. As
we will see in the proof of the main theorem (see (ref)), novel
extensions are characterised by a cardinal notion. In contrast, the notions of
testworthiness and perturbation that we now introduce are ordinal in nature.
definition*A proper extension $ \mathrel{\mc R} $ of $\mathbin{\preceq}_{{\mathds D}}$ is testworthy if it satisfies
(ref)--(ref) and, for some $D\in {\mathds D}$ such that $\mathrel{\mc R}_{D}$
is total, $ \mathbin{\mc R} _{ \mathfrak f }$ is the inverse of $\mathbin{\mc R}_{D}$.\footnote{Recall
that the inverse $ \mathrel{\mc R} _{ D } ^{ - 1 }$ of $ \mathrel{\mc R} _{ D }$ satisfies
$ x \mathrel{\mc R} _{ D } ^{ - 1 } y $ if, and only if, $ y \mathrel{\mc R} _{ D } x$.}
We now introduce perturbations. The motivation is that, for any
given extension $\mathrel{\mc R}$, the predictor knows $\mathrel{\mc R}_{{\mathds D}}$ and freely chooses
$\mathrel{\mc R}_{\mathfrak f}$, but the rankings at other databases are somewhat arbitrary. For
although the rankings the predictor associates with members of
${\mathds D^{\mathfrak f}}\bs ({\mathds D} \cup \{\mathfrak f\}) $ are constrained, what
matters is not the actual rankings, but rather their potential for consistency
with the axioms.
definition*Let $\mathrel{\mc R}$ and $\mathrel{\acute{\mathrel{\mathcal R}}}$ be extensions of $\mathbin{\preceq}_{{\mathds D}}$. $\mathrel{\acute{\mathrel{\mathcal R}}}$ is a
perturbation of $\mathrel{\mc R}$ if $\mathbin{\acute{\mathbin{\mathcal R}}}_{\mathfrak f } =
\mathbin{\mc R}_{\mathfrak f}$. Moreover, $\mathrel{\acute{\mathrel{\mathcal R}}}$ is a nondogmatic perturbation if
$\mathbin{\sharp}\hskip1pt \textup{total} (\mathbin{\acute{\mathbin{\mathcal R}}}) \geq \mathbin{\sharp}\hskip1pt \textup{total} ( \mathbin{\mc R})$.
We are interested in perturbations that are nondogmatic because may reveal
intransitivities that the predictor's inexperience conceals.
The prudent predictor may exploit them and revise her model before new cases arrive.
prudence*For every $Y\subseteq X$ with cardinality $ 3$ or $ 4 $, every testworthy
$ Y $-extension of $ \mathbin{\preceq} _{ {\mathds D} }$ that is novel has a
nondogmatic perturbation that satisfies (ref)–(ref).
Given that the extensions we consider are nonrevisionistic, it is natural to ask
whether (ref)--(ref) are superfluous in the presence of 4-prudence. Our response
is twofold. Firstly, in practice, we would expect (ref)--(ref) to hold more
frequently than 4-prudence\ which is more cognitively demanding. Secondly, as we
will see in the proof of the main theorem, when $ {\mathds {T}} $ is infinite, for some
$ Y \subseteq X $, the set of testworthy $ Y $-extensions that are novel may be
empty. For every such $Y$, the following observation confirms that 4-diversity\
holds on $Y$.
theoremEnd[link to proof]{observation}[on testworthy
extensions]
Let $\preceq_{{\mathds D}}$ satisfy (ref)–(ref). For every $Y\subseteq X$ of
cardinality $ 3$ or $ 4$, the set of testworthy $Y$-extensions is
nonempty. If, for some $Y$, every testworthy $Y$-extension is regular, then
4-diversity\ holds on $Y$.
proofEndRespectively, these two statements follow via (ref)
and (ref).
It is also natural to ask whether 4-prudence\ is simply requiring that, on $Y$ such that
$\preceq_{\mathds J}$ fails to satisfy 4-diversity, there exists a testworthy
$Y$-extension that is novel and satisfies 4-diversity. In the proof of the
theorem that now follows, we show that this is not the case.
\paragraph{The main theorem\hskip-5pt} that now follows involves real-valued
function $ \mathbf{v} $ on the product $X \times {\mathds C} $. We view $\mathbf{v}$
as a matrix and $ \mathbf{v} ( x , \cdot )$ as one of its rows. The matrix
$ \mathbf{v} $ is a representation of $ \preceq _{ {\mathds D} }$ whenever it
satisfies
linenomath*\begin{equation}\notag
\left\{
\begin{array}{l}
for every $ x , y \in X$ and every $ D \in {\mathds D} $,\\
x \preceq _ D y \quad if, and only if, \quad \sum _ { \, c \,\in\, D} \mathbf{v} ( x
, c ) \leq \sum _ {\, c \,\in\, D } \mathbf{v} ( y , c ) .
\end{array}\right.
\end{equation}
The matrix $\mathbf{v}$ respects case equivalence (with respect to
$\preceq_{{\mathds D}}$) if, for every $c,d\in {\mathds C}$, $c \sim^{\star} d$ if,
and only if, the columns $\mathbf{v}(\cdot,c)$ and $ \mathbf{v}(\cdot,d)$ are equal.
theorem[Part I, Existence]
Let there be given $X$, ${\mathds C^{\mathfrak f}}$, $\mathbin{\preceq}_ {\mathds D}$ and
associated extensions, as above, such that the richness condition holds. Then
(ref) and (ref) are equivalent.
\begin{enumerate}[label=((ref).\roman*)]
• (ref)--(ref) and 4-prudence\ hold for $\preceq_{{\mathds D}}$.
• There exists a matrix
$ \mathbf{v} : X \times {\mathds C} \rightarrow \R $ satisfying (ref) and (ref)$\,:$
\begin{enumerate}[label=((ref).\alph*)]
• $ \mathbf{v} $ is a representation of
$ \preceq _ { {\mathds D} }$ that respects case equivalence$\,;$
• no row of $\mathbf{v}$ is dominated by
any other row, and for every three distinct elements $x,y, z \in X$ and
$\lambda \in \R$, $\mathbf{v} (x, \cdot) \neq \lambda
\mathbf{v}(y,\cdot) + (1-\lambda) \mathbf{v}(z,\cdot)$\,.\footnote{Observe that
$\mathbf{v}(x,\cdot)- \mathbf{v}(z,\cdot)$ and
$\mathbf{v}(y,\cdot)-\mathbf{v}(z,\cdot)$ are noncollinear if, and only if,
the affine independence condition of (ref) holds.}
\end{enumerate}
\end{enumerate}
Our uniqueness result is identical to that of $\textup{[GS]}$. \setcounter{theorem}{0}
theorem[Part II, Uniqueness]
If (ref) $[$or (ref)$]$ holds, then the matrix
$ \mathbf{v} $ is unique in the following sense$\,:$ for every other matrix
$ \mathbf{u} : X \times {\mathds C} \rightarrow \R $ that represents
$\mathbin{\preceq}_{{\mathds D}}$, there is a scalar $ \lambda > 0 $ and a matrix
$ \beta : X \times {\mathds C} \rightarrow \R$ with identical rows (\ie\ with
constant columns) such that $ \mathbf{u} = \lambda \mathbf{v} + \beta$.
The proof of (ref) appears in online appendix (ref). It
relies upon a translation from the abstract database/memory set up of the model
to the setting of rational vectors similar to $\textup{[GS]}$. We show that the translated
theorem (ref) (see (ref)) is equivalent to
(ref). The proof of (ref) is then the subject of online
appendix (ref).
\paragraph{We appeal to the following corollary}\hskip-7pt
when we connect the present framework to the concept of arbitrage in finance. Central to this connection is
definitionFor $ Y \in 2 ^ { X } $, the matrix
$ v^{ {(\cdot,\cdot)} } : Y^{2}\times {\mathds C} \rightarrow \R $ satisfies the Jacobi
identity whenever, for every $ x , y , z \in Y $, the rows satisfy
$ v ^{ {(x,z)} } = v ^{ {(x, y)} } + v ^{ {(y,z)} }$.
In (ref), of (ref), we show that, for a given
extension $\mathrel{\mc R}$, (ref)--(ref) and 2-diversity\ yield a pairwise
representation $v^{{(\cdot,\cdot)}}$ of $ \mathrel{\mc R} $.\footnote{That is, for every $x,y\in X$
and $D \in {\mathds D}$, $x \mathbin{\mc R}_{D} y$ if, and only if
$ \sum_{c \,\in D}v^{{(x, y)}}(c) \geq 0$.}
theoremEnd[link to proof]{corollary}[a characterisation of
prudence]
Let the number of case types be finite and let $\preceq_{{\mathds D}}$ satisfy
(ref). Then $\preceq_{{\mathds D}}$ satisfies 4-prudence\ if, and only if,
$\preceq_{{\mathds D}}$ has a pairwise representation $v^{{(\cdot,\cdot)}}$ that satisfies the
Jacobi identity. Moreover, for every other pairwise representation
$u^{{(\cdot,\cdot)}} $, there exists $\lambda >0$, such that $u^{{(\cdot,\cdot)}} = \lambda v^{{(\cdot,\cdot)}}$.
proofEndThis follows from (ref), (ref) and the fact
that, via (ref), 4-prudence\ implies (ref)--(ref) when
the number of case types is finite.
Discussion
We begin by restating the existence part of the main theorem of $\textup{[GS]}$.
theorem*[Existence]
Let there be given $X$, $ {\mathds C} $ and $\mathbin{\preceq}_ {\mathds D}$, as above, such that the
richness condition holds. Then (ref) and (ref) are equivalent.
\begin{enumerate}[label=(\roman*)]
• (ref)--(ref) and 4-diversity hold for $\preceq_{{\mathds D}}$.
• There exists a matrix
$ \mathbf{v} : X \times {\mathds C} \rightarrow \R $ satisfying (ref) and (ref)$\,:$
\begin{enumerate}[label=(\alph*)]
• $ \mathbf{v} $ is a representation of $ \preceq _ { {\mathds D} }$ that respects case equivalence$\,;$
• if $\mathbin{\sharp}\hskip1pt X < 4$, then no row is dominated by an
affine combination of the other rows, and for every four distinct elements
$x,y,z,w \in X$ and every $\lambda , \mu, \theta \in \R$ such that
$ \lambda +\mu + \theta = 1$,
$\mathbf{v}(x,\cdot ) \not \leq \lambda \mathbf{v}(y,\cdot )+\mu
\mathbf{v}(z,\cdot)+ \theta \mathbf{v}(w,\cdot)$.
\end{enumerate}
\end{enumerate}
Although diversity axioms play an important technical role, they are not
obviously behavioural. Instead, diversity axioms impose restrictions on what is
beyond the predictor's control and on what is central to inductive inference:
experience. Our main contention is that $ {C^\star} $ may not be so rich as to
support $\mathbin{\preceq}_{{\mathds D}}$ satisfying 4-diversity. That is to say, there may exist
$ Y \subseteq X $ such that $ \mathbin{\sharp}\hskip1pt Y = 4$, and such that the data is
insufficiently rich to support all $ 4 ! = 24 $ strict rankings. A casual
comparison of condition (ref) and (ref) confirms that the
present framework achieves the main purpose of accommodating the less
experienced.
For the remainder of this discussion, we take both $X$ and the set ${\mathds {T}}$ of
case types to be finite and of cardinality $m$ and $n$ respectively. Via case
equivalence, for any $\preceq_{{\mathds D}}$ (satisfying (ref) or
(ref)) we may efficiently summarise rankings using a real-valued
$m\times n$ matrix $\mathbf{v}$ on the product of $X$ and case types $ {\mathds {T}}$.
\paragraph{Comparing the complexity of conditional-2-diversity\ and 4-diversity,}\hskip-8pt in
the presence of the other axioms, provides a measure of the value of
experience. As an estimate, we compare condition (ref) with
(ref) of Gilboa and Schmeidler's theorem. Verifying (ref)
involves checking $n$ affine dominance constraints: one for each case type. This
is well-known to be equivalent to the complexity of linear programming with real
variables DR-Linear_programming. In the absence of knowledge
regarding the sparsity of $\mathbf v$, the fastest algorithm for achieving this
is of order $n^{3}$ LS_Linear_programming. Since this holds for
every subset of four distinct eventualities, checking (ref) is of
order $\binom{m}{4}n^{3}$. In contrast, verifying
$\mathbf{v}(x,\cdot )\not \leq \mathbf{v}(y,\cdot)$ takes at most $n$ steps for
each of the $\binom{m}{2}$ subsets of $2$ distinct elements. Likewise, for every
three distinct elements $x,y,z$ in $X$, checking for noncollinearity of two
vectors takes at most $n$ steps. Thus, a na\"{i}ve algorithm for checking
(ref) is of order $\binom{m}{3}n$. Even for $m = 4 $ and $n = 4$, the
difference is stark: $\binom{m}{3}n = 16$ versus $ \binom{m}{4} n^3 = 64$. (The
threshold $4$ is important as we show in (ref) below.)
\paragraph{A robust test of prudence\hskip-7pt} is available precisely when
(ref)--(ref) hold and 4-diversity\ fails. For, when $n<\infty$, the existence
of a similarity representation satisfying (ref) is equivalent to
4-prudence. The test is robust in that, generically, predictors that fail to check
for the arrival of new cases also fail to satisfy (ref). This is
because, the row differences
$v^{(x, y)} = - \mathbf{v}(x,\cdot) + \mathbf{v}(y,\cdot)$, the Jacobi identity
$v^{{(x,z)}} = v^{{(x, y)}} + v^{{(y,z)}}$ holds for every $x,y,z \in X$. This in turn
implies that the hyperplanes $H^{\{x,y\}}, H^{\{y,z\}}$ and $H^{\{x,z\}}$, to
which the row differences are normal, are congruent. Congruence implies the
hyperplanes are not in general position. (When $n = 2$, the same argument can be
made in terms of extensions: see (ref) of the proof of
(ref).) In other words, there is a zero (conditional) probability of
the predictor striking lucky and appearing to be prudent when in fact they are
not engaging in this form of second order inductive inference.
\paragraph{Novelty, memory and transitivity.} An early test and taxonomy of
intransitive behaviour is due to
Weinstein-Intransitivity,Weinstein-Transitivity. Weinstein points out
that intransivity can sometimes be rational in complex situations.
Interestingly, he shows that young are people significantly less transitive in their
choices. He also points out that the law itself is designed to accommodate
irresponsible under-age decisions. In the psychology literature,
BR-Novelty_and_intransitivity show that presenting novel objects is more
likely to trigger intransitive choices.
More recently, Enkavi-Hippocampal_dependence provide evidence that
people with a specific form of memory impairment (lesions in the hippocampus of
the brain)
are significantly more likely to violate transitivity in pairwise choices of
chocolate bars: even though they rank numbers transitively.\footnote{The
hippocampus is associated with learning and
memory. \Citet{Hassabis-Hippocampal_dependence} and
Schacter-Hippocampal_dependence present evidence showing that the
hippocampus plays an important role in imagining future experiences on the
basis of past ones. \Citet{Enkavi-Hippocampal_dependence} go further by
showing that it plays a role in the value-based decision making framework of
Rangel-Value-based_neurobiology.} Whilst this literature does not yet
offer a direct test of the present model, it supports the case-based framework
of constructing preference as well as our premise that violations of
transitivity are often driven by novelty, or equivalently impaired memory.
In line with the case-based approach, the experimental evidence of
Enkavi-Hippocampal_dependence suggests that agents are constructing
their preferences on the basis of past experience. Moreover, it seems natural to
interpret agents with hippocampal impairment as inexperienced predictors. The
fact that impaired agents then make intransitive decisions is very much in line
with what our model predicts as they are, in effect, facing a novel situation
and are required to construct their preferences on the fly. It appears that
impaired agents are also failing to be prudent, though it is not obvious that
chocolates warrant the additional neural computation that accompanies 4-prudence.
remark*A closer look at the relationship between the proportion $\rho$ of hippocampal
impairment and the percentage $\sigma$ of intransitive choices Enkavi-Hippocampal_dependence suggests another interpretation.
For $\rho $ above $ \frac{1}{4}$, $\sigma $ is above $ 20\%$$:$ twice as high as
it is for $\rho < \frac{1}{4}$. Memories and the rankings they generate are
latent variables to the observable $\rho$ and this threshold is where
conditional-2-diversity\ fails to hold and our model breaks down. The $16$ cases in the
data are thus partitioned into three groups$:$ two that satisfy 3-diversity\
($\rho < 0.05 $ and $\sigma< 5\%$)$;$ $12$ intermediate cases that satisfy
conditional-2-diversity\ ($0.05 \leq \rho \leq 0.25 $ and $5\%\leq \sigma \leq 10\%$)$;$ and
$2$ severe cases that fail to satisfy \textup{conditional-\textit{2}-diversity}\ ($0.25 < \rho$ and
$10\% < \sigma$).
\paragraph{The success or failure of startups. \hskip-7pt} Inexperience raises
significant barriers to entry. Overcoming these barriers is either the result of
making mistakes and learning by doing “on the fly” or the result of being
prudent. Which form of second-order induction bears out in practice will depend
on many factors.
The following proposition confirms our thesis that experienced predictors (\ie\
those that satisfy 4-diversity) have indeed encountered a high number of case
types. As the main theorem of $\textup{[GS]}$\ shows, experienced predictors have no need
for the additional structure of extensions. Unless the prediction problem
changes (\eg\ new eventualities become relevant), they have no need to engage in second order induction. This saving in
cognitive effort is the prize that experience confers.
theoremEnd{proposition}[experience and case types]
If $\mathbin{\preceq}_{{\mathds D}}$ satisfies (ref)--(ref) and 4-diversity, then
$ n \geq \min \{4, m\} $, and, for every $Y \subseteq X$ of cardinality $m'$
and regular $Y$-extension $\mathrel{\mc R} $, the number $n'$ of equivalence classes of
$\sim^{\mathrel{\mc R}}$ satisfies $ n' \geq \min \{4,m'\}$.
proofEndVia (ref), (ref)--(ref) and 4-diversity\ hold for
$\preceq_{{\mathds D}}$ if and only if (ref)--(ref) and 4-diversity\ hold for
$\preceq_{\mathds J}$. Let $Y\subseteq X$ be of cardinality $m' = 1, 2, 3$ or
$4$ and let $\mathrel{\mc R}$ be a regular $Y$-extension. Via (ref), there
exists a pairwise representation $v^{{(\cdot,\cdot)}} $. For $m' = 1$, $n' = 1 $ because
$\mathrel{\mc R}_{J}$ is constant on ${\mathds {J}^{\mathfrak f}}$. For $m' = 2$, $n' \geq 2$, since via
part (ref) of (ref), $G^{{(x, y)}}$ and $G^{{(y, x)}}$ are both
nonempty.
By way of contradiction, first suppose $n' = 2$ and $m' \geq 3$. Via
(ref) of appendix (ref),
$\textup{total}(\mathrel{\mc R}) \leq 4$. In contrast, 4-diversity\ requires $\textup{total}(\mathrel{\mc R}) = 6$.
The remaining case is where $n' = 3$ and $m' \geq 4$. If the rank
$\mathbf r$ of $v^{{(\cdot,\cdot)}}$ satisfies $\mathbf r \geq 3$, then the kernel
$A^{Y}$ of $v^{{(\cdot,\cdot)}}$ is zero-dimensional. Then $0$ is the unique element of
$ A^{Y}$. Thus, the positive kernel $A^{Y}_{\text{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\srcsize}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}$+\mkern-2mu+$}}$ of $v^{{(\cdot,\cdot)}}$ is
empty. Then Zaslavski's theorem implies that $\textup{total}(\mathrel{\mc R})< 4!$, so that
\textit{4}-\textup{diversity}\ fails to hold. If $\mathbf r \leq 2$, then an application of the
rank version Zaslavski's theorem (in particular (ref) with
$\acute{\mathbf r } = \mathbf r = 2$) yields
\begin{linenomath*}
\[ \lvert \mc G_{\text{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\srcsize}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}$+\mkern-2mu+$}} \rvert \leq 1 - 6 + 15 - 20 + 15 + 6 + 1 =
12. \]
\end{linenomath*}
Thus, once again \textit{4}-\textup{diversity}\ fails to hold. Thus $n' \geq \min \{4, m'\}$, as
required. Finally, since $Y\subseteq X$, $m\geq m'$, and, via part
(ref) of (ref), $n \geq n'$.
In contrast, conditional-2-diversity\ implies no restrictions on the cardinality of
${\mathds {T}}$ beyond $n \geq 2$ and this is also a virtue of 2-diversity.
\paragraph{When is prudence worth the trouble?}
The simple answer to this question is: when revising a model “on the fly”,
once a novel case arrives, is costly. The following is our main example of such
a setting.
exampleConsider a fair market maker of zero-coupon (treasury) bonds.\footnote{Similar
to a fair insurer, the fair market maker sets the market spread to zero.}
The compound-interest formula for the accumulation process of such a bond
is
\begin{linenomath*}
\[a^{{(x, y)}}= \left(1+r^{\{x,y\}}\right)^{-x + y}, \]
\end{linenomath*}
where $r^{\{x,y\}}$ is the implied yield on a forward contract that accrues
interest between dates $x$ and $y$. If $ x$ is later than $y$, then the
contract is to sell, and the market maker pays this yield, so that
$r^{\{y,x\}} = r^{\{x,y\}}$. Let $X \subseteq \R_{\text{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\@setfontsize{\srcsize}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}}{3pt}{3pt}$+$}}$ index a suitable
sequence of trading dates with $0 \in X$ being the spot date. It is well
known that the (normalised) spot bond price
$x \mapsto b(x) = \left(1+r^{\{0,x\}}\right)^{-x} $ is arbitrage-free if, and
only if, the log-accumulation process satisfies
\begin{linenomath*}
\begin{equation}
for every $x,y,z \in X$,\quad \log a^{{(x, y)}} = \log a^{{(x,z)}} + \log a^{{(z,y)}}
\,.\footnote{To see this, suppose that, for some
$x< y$, she sets $ a^{{(x, y)}} < a^{{(x,z)}}a^{{(z,y)}}$. Another trader would do well
to sell the forward contract ${(x, y)}$, buy the spot contract ${(x,z)}$ and sell the
spot contract ${(z,y)}$. A risk-free arbitrage opportunity is also
available if the reverse inequality holds.}
\end{equation}
\end{linenomath*}
This no-arbitrage condition is a special case of the Jacobi
identity. We now explain how a market maker might infer the
accumulation process from past cases.
Let each case $c \in D^{\star}$ consist of market-relevant data from a given
time interval (a block of time periods) in the past. These blocks are chosen so
that, for any finite resampling $D $ of cases from $D^{\star}$, the sequence of
cases that makes up $ D $ is exchangeable. The free case $\mathfrak f $ has no
additional structure beyond that of
(ref).
Next, for every finite resampling $D$ and every date $x$ and $y$, let
$x \preceq_{D} y$ if, and only if, in answer to the question “At which date
will the price be higher?” the market maker finds that $y$ is more plausible
than $x$.
For the following corollary, we introduce the empirical implied yield function
that maps triples $x \times y \times c $ to $ r^{\{x,y\}}_{c} \in \R$. This
function is characterised by three conditions: for time intervals of length
zero, the yield is zero; fair pricing; and case equivalence. These are
respectively formalised as follows: for every $x,y\in X$ and every
$c , d\in {\mathds C}$, $ r^{\{x,x\}}_{c} = 0$; $r^{\{y,x\}}_{c} = r^{\{x,y\}}_{c}$;
and $c \sim^{\star} d$ if, and only if, $r^{\{x,y\}}_{c} = r^{\{x,y\}}_{d}$. The
empirical bond price function $B : X \times {\mathds D} \rightarrow \R$ maps every
pair $x \times D$ to
\[ B(x,D) = \prod_{c\, \in D}\left( 1 + r^{\{0,x\}}_{c}
\right)^{-\frac{x}{ \lvert D \rvert}}.
\]
When $D^{\star}$ belongs to ${\mathds D}$, the number of case types is finite, and we have
\begin{corollary}
Let $\preceq_{{\mathds D}}$ satisfy (ref). Then $\preceq_{{\mathds D}}$ satisfies
4-prudence\ if, and only if, there exists empirical implied yield and
empirical bond price functions, such that
\begin{linenomath*}
\begin{equation}\tag{$*$}
\left\{
\begin{array}{l}
for every $ x , y \in X$ and every $ D \in {\mathds D} $,\\
x \preceq_{D} y \quad if, and only if,\quad B(x,D) \leq B(y,D).
\end{array}\right.
\end{equation}
\end{linenomath*}
Moreover, for every $D\in {\mathds D}$, the bond price
$B(\cdot,D)$ is arbitrage-free.
\end{corollary}
\begin{proof}[Proof of (ref)]
Via (ref), there exists a pairwise Jacobi representation
$v^{{(\cdot,\cdot)}}: X^{2} \times {\mathds C} \rightarrow \R$ such that, for every
$c,d\in {\mathds C}$, $v^{{(\cdot,\cdot)}}(c) = v^{{(\cdot,\cdot)}}(d)$ if, and only if, $c \sim^{\star}
d$. Moreover, $v^{{(\cdot,\cdot)}}$ is unique upto multiplication by a positive
scalar. Then, for every $x, y, z \in X$, $v^{(x,x)}(\cdot) = 0$,
$v^{{(y, x)}} = - v^{{(x, y)}}$, and $v^{{(x, y)}} = v^{{(x,z)}} + v^{{(z,y)}}$.
Recalling that $0 \in X$, for every $x\in X$ and $D \in {\mathds D}$, let
\begin{linenomath*}
\begin{equation}
\textstyle B(x,D) = \exp\left( -\frac{1}{\lvert D \rvert} \sum_{c\,\in D} v^{(x,0)}_{c}
\right).
\end{equation}
\end{linenomath*}
For the proof of (ref), recall that $x \preceq_{D} y$ if,
and only if, $\sum_{c\,\in D} v^{{(x, y)}}(c) \geq 0$. Since
$- v^{(y,0)}_{c} = v^{(0,y)}_{c}$ and, via the Jacobi identity,
$v^{(x,0)} + v^{(0,y)} = v^{{(x, y)}}$, we have:
\begin{linenomath*}
\[\textstyle
- \log B(x,D) + \log B(y,D) = \frac{1}{\lvert D\rvert} \sum_{c\,\in D}
\left(v^{(x,0)}_{c} - v^{(y,0)}_{c}\right) = \frac{1}{\lvert D\rvert}
\sum_{c\,\in D} v^{{(x, y)}}_{c}.
\]
\end{linenomath*}
It remains for us to confirm that the bond price is a suitable function of the
empirical yield function. For every $x,y \in X$ and $c \in {\mathds C}$,
$ \log a^{{(x, y)}}_{c} = -v^{{(x, y)}}(c)$: so that, as the solution to
$ (y - x )\log (1+r^{\{x,y\}}_{c}) = -v^{{(x, y)}}_{c} = v^{{(y, x)}}_{c}$, for
$x \neq y$,
\begin{linenomath*}
\begin{equation}
1+ r^{\{x,y\}}_{c} = \exp\left(\frac{v^{{(y, x)}}_{c}}{y-x}\right) =
\exp\left( \frac{v^{{(x, y)}}_{c}}{x-y}\right) = 1+r^{\{y,x\}}_{c}.
\end{equation}
\end{linenomath*}
We therefore observe that $r^{\{x,y\}}_{c} = r^{\{y,x\}}_{c}$. For $x = y$,
$v^{{(x, y)}}_{c}= 0$ ensures that we can take $r^{\{x,x\}}_{c} = 0$. Finally, note
that for $c \sim^{\star} d$, the property $r^{\{x,y\}}_{c} = r^{\{x,y\}}_{c}$ is
inherited from $v^{{(x, y)}}_{c} = v^{{(x, y)}}_{d}$, so that we have an empirical implied
yield function.
The fact that, for every $D$, $B(\cdot, D)$ is arbitrage-free follows by virtue
of the fact that $v^{{(\cdot,\cdot)}}$ satisfies the Jacobi identity.
\end{proof}
In the present setting, because we are modelling a normalised bond price, we
obtain a stronger uniqueness result relative to part II of (ref). If
$\tilde B$ is another function that satisfies the present corollary, then, for
every $D \in {\mathds D}$, the spot price $\tilde B(0,D) = 1$. Thus, via
(ref) and (ref), for some $\lambda >0$,
$\tilde B = \textup{e}^{-\lambda} B$.
We now point out an interesting implication of conditional-\textit{2}-diversity\ the recent
prevalence of negative interest rates. \Wlog, fix $x<y$. Then, given \textit{4}-\textup{prudence},
\textit{2}-\textup{diversity}\ implies that there exists $c,d \in {\mathds C}$ such that
$v^{{(x, y)}}(c) < 0 < v^{{(x, y)}}(d)$. This is equivalent to
$v^{{(y, x)}}(d) < 0 < v^{{(y, x)}}(c)$, and, via (ref),
$r^{\{x,y\}}_{d} < 0 < r^{\{x,y\}}_{c}$. That is, \textit{2}-\textup{diversity}\ requires that the
market maker's data is rich enough to contain at least one case where the yield
between date $x$ and $y$ is negative (as well as one where it is
positive).
\textup{Conditional-\textit{2}-diversity}\ extends this notion to require that
$r^{\{x,y\}}_{D} < 0 < r^{\{x,y\}}_{C}$ for some $C$ and $D$ such that
$r^{\{x,z\}}_{C}\cdot r^{\{x,z\}}_{D} >0$.
\paragraph{Discussion of second-order induction.}\vskip-8pt
The market maker of (ref) engages in second-order induction when
she acts prudently. She reflects on her model by checking that the basic
axioms of $\textup{[GS]}$\ will continue to hold when a novel case arrives. By way of
contrast, suppose the bond price of the market maker is such that
$\preceq_{{\mathds D}}$ is consistent with the basic axioms, but not 4-prudence. Then
when a novel case arrives, she may be exposed to arbitrage and need to
respecify her entire model “on the fly”. Such a step corresponds to the
intermittent respecification of her similarity weighting function
$\mathbf{v}(x,c)$, that AG-Second-order_induction
describe. In AG-Second-order_induction, the “leave-one-out”
technique of cross-validating the model by omitting a case of each type is
intuitively and operationally close to our inclusion of the free case
$\mathfrak f$. The difference is that by allowing $\mathfrak f $ more degrees of
freedom, our market maker can study novel extensions and peer into the future
through the lens of her current model. She can exploit the intervals of time
inbetween the arrival of novel cases by continuously engaging in second-order
induction.
Through an example, we now show that the present framework provides the
flexibility to accommodate second-order induction without sacrificing the
computational or normative advantages that additive similarity functions
provide.
example*[second-order induction, $\textup{[GS]}$, p.12] Let
$c$ denote a case where Mary chooses restaurant $x$ over restaurant $y$. In
the absence of any further information, it is tempting to assume some
similarity between John and Mary. The predictor then finds it plausible that
John prefers $x$ to $y$ given $\{c\}$. A separate database $D$ contains no
choices between $x$ and $y$. Thus, in the absence of further information,
$x$ and $y$ appear equally likely based on $D$. Additivity of the similarity
function (or (ref))) implies it is plausible that John prefers $x$ to $y$
given $\{c\} \cup D$. The violation of (ref) arises when a more careful
examination of the contents of $D$ reveals many choices between other pairs
of restaurants where John and Mary consistently differ.
Quine's notion of perceptual similarity Quine-Roots_of_reference offers
a check on the predictor's inference about John's choice given $\{c\}$. John and
Mary may just as well be two drivers passing through an intersection at
different times. Although their situations are broadly speaking very similar, if
one faces a red light and the other a green light, their responses will
differ. In the restaurant setting, some pivotal information is omitted from $c$.
Observing that the evidence in $c$ in favour of $x$ over $y$ is somewhat weak, a
prudent predictor instead recasts $\{c\}$ as a database $C$ that combines past
observations with copies of the pivotal novel case $\mathfrak f$. With a more
refined model, that explicitly allows for omitted variables, the predictor can
check to see if her model extends to higher dimensions without violating the
basic axioms.
remarkObserve that predictors that only fail to satisfy (ref) can, with some
additional regularity conditions, still be represented by a nonlinear function
$u: X \times {\mathds D} \rightarrow \R $ such that for every $D \in {\mathds D}$ and every
$x,y \in X $, $x \preceq_{D} y$ if, and only if, $ u(x,D)\leq u(y,D)$
OCallaghan-Parametric_continuity. But predictors that satisfy
(ref)--(ref) but not 4-prudence\ can also be represented by such a function:
because $\preceq_{{\mathds D}}$ is complete and transitive for each $D \in {\mathds D}$. The
present framework allows us to disentangle the latter kind of predictor from
those who, for good reason, fail to satisfy the combination axiom. (See $\textup{[GS]}$\
for examples of such reasons.)
\paragraph{On the veracity of false news.} How should a predictor check whether
her model consistently extends to higher dimensions (when novel cases arrive)?
If our model is a guide then, the most useful rankings that she might wish to
examine are those that are far from her own. This is because our definition of
testworthy extensions involves assigning to the novel case $\mathfrak f$ the inverse
of some total ranking $\preceq_{D}$. This may offer some rationale for why
information that differs from our own is intrinsically valuable. Testworthy
extensions play a vital role in taming the complexity of our proof. It seems
plausible that something similar may be at play when agents encounter radically
different information from their own on social media: even if it is
fake. This may help to explain why false news is significantly more veracious
than real news online Vosoughi-Roy-Aral-Veracity. The fact that real
news is typically closer to what we have observed in the past means that it is
of less value to the prudent predictor that finds it costly to imagine worlds
that are far from her own.
\makeatletter
\def\@seccntformat#1{Appendix\,\csname the#1\endcsname.\quad}
\makeatother