EconBase
← Back to paper

Modelplasticity and Abductive Decision Making

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

58,664 characters · 18 sections · 35 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Modelplasticity and Abductive Decision Making$^*$ footnote-1

}

tabularx[tabularx omitted — 91 chars of source]

}

abstract`All models are wrong but some are useful' (George box1979robustness). But, how to find those useful ones starting from an imperfect model? How to make informed data-driven decisions equipped with an imperfect model? These fundamental questions appear to be pervasive in virtually all empirical fields---including economics, finance, marketing, healthcare, climate change, defense planning, and operations research. This article presents a modern approach (builds on two core ideas: abductive thinking and density-sharpening principle) and practical guidelines to tackle these issues in a systematic manner.

\noindentKeywords: Abductive Decision-making; Model Risk Management; The Uncertainty Principle; Density-Sharpening Principle; Creation of New Knowledge; Quantile Decision Analysis.

\vskip1em

The Uncertainty Principle

How to make decisions under uncertainty? Decision-making under uncertainty relies mainly on how efficiently we can extract useful knowledge from the data that were previously unknown to the decision-maker\footnote{“Anything that gives us new knowledge gives us an opportunity to be more rational” – Herbert Simon}. C. R. Rao, in his rao1996 article\footnote{The existence of this article is barely known.} on `Uncertainty, Statistics, and Creation of New Knowledge' provided an exquisite description of the mechanics of decision-making under uncertainty using a simple logical formula:

\beq {-1.25em} {1.5em} \fbox{Uncertainty of knowledge} + \fbox{Knowledge of uncertainty}\, =\, \fbox{Usable knowledge.} \eeq

A decision analyst confronts data $X_1,\ldots, X_n$ equipped with a tentative (imprecise and uncertain) probabilistic model $f_0(x)$ of the underlying phenomena. The challenge then boils down to effectively using the misspecified model $f_0(x)$ to learn from data and to apply that knowledge for informed decision-making. Rao's uncertainty principle suggests the following three-staged approach, which we call the `model-building triad':

\vskip.75em

Stage 1. Model Elicitation. The first step of decision making is formulating a probability model of the phenomena of interest, from which economic agents derive their initial expectations. A simple parametrized model-0 $f_0(x)$ is usually formed based on either gut instinct or the scientific context of the investigation. The uncertainty of $f_0(x)$ arises due to the lack of perfect knowledge\footnote{what keynes1937general called `uncertain knowledge.'} about the underlying probability law. Accordingly, the modeler has to start the analysis by acknowledging the uncertainty of the initial knowledge model $f_0(x)$.

\vskip.45em

Stage 2. Model Uncertainty Quantification. Before making decisions based on the provisional model $f_0(x)$, it is crucial to investigate its uncertainty (blind spots) in light of the new data. It's always a good practice to inspect expert opinions based on hard empirical facts by asking\footnote{Those who ignore experts' knowledge and only trust data are empirical-fools. Those who ignore data and only trust their gut-instinct are emotional-fools tversky1974. Expert decision-makers always use empirically-guided intuition by appropriately combining both data and available knowledge.}: what's new in the data that can't be explained by the assumed model? Discovering surprising and previously unknown facts can prompt decision makers to consider other alternative actions.

\vskip.45em

Stage 3. Model Rectification and Risk Management. Finally, we incorporate the learned uncertainty into the uncertain model $f_0(x)$ to produce a rectified model for making empirically-guided informed decisions. It is important to sharpen the “judgment component” (intuition based on past experiences) in light of the new data before it gets outdated.

\vskip.75em

The purpose of this article is to describe a general statistical theory that permits us to implement this three-staged model-building procedure for data analysis and decision-making.

Learning with Imperfect Model

\vskip1em

quote{`All analysts approach data with preconceptions. The data never speak for themselves. Sometimes preconceptions are encoded in precise models. Sometimes they are just intuitions that analysts seek to confirm and solidify. A central question is how to revise these preconceptions in the light of new evidence.'} \begin{flushright} {\rm --- heckman2017abducting} \end{flushright}

Empirical scientific inquiry typically starts with a simple yet believable model of reality (model-0) and aims to sharpen existing knowledge by gathering new observations. \vskip.25em We observe a random sample $X_1,\ldots,X_n \mathrel{\dot\sim} F_0$. By “$\mathrel{\dot\sim}$” we mean that $F_0$ is an `approximately correct' structured provisional model for $X$ that is given to us by subject-matter experts. We like to extract new knowledge from the data by smartly leveraging existing knowledge\footnote{Model amendment principle: the starting model $f_0(x)$ is incomplete but not useless. It contains valuable background knowledge. Rather than throwing this vital information, we want to build a model by smartly taking clues from it. The goal is to amend model-0, not to abandon it completely.} that is encoded in the initial approximate model $f_0(x)$.\footnote{As for notation: by $F_0(x)$, we denote the cumulative distribution function (cdf) of the starting model-0; $f_0(x)$ is the probability density function (pdf) and quantile function is denoted by $Q_0(u)$. The expectation with respect to $f_0(x)$ will be abbreviated as $\Ex_0$.} \vskip.25em

Creating knowledge-guided statistical models. The core mechanism of our process involves: (i) inspecting whether the structured provisional model-0 is still a good fit in light of fresh data; (ii) if not, then we like to know what's new in the data that cannot be tackled by the current model; and, finally, (iii) repair the current misspecified model in order to cope with the new reality. However, the question remains as to how can we design an inference machine that can offer these successively fine-grained insights? To address this question, we will describe a new statistical model building principle, called the `density-sharpening principle.'

A Dyadic Model

We introduce a dyadic model with two interrelated subsystems that accommodates the decision maker's concern for misspecification of the starting expert-guided model.

defn[Dyadic model] $X$ be a general (discrete, continuous, or mixed) random variable with true unknown density $f(x)$ and cdf $F(x)$. Let $f_0(x)$ represents a simple approximate model for $X$ with cdf $F_0(x)$, whose support includes the support of $f(x)$.\footnote{For dealing with truly zero-probability events see coletti2002book} Then the following dyadic density decomposition formula holds: \beq {1.4em} {1.4em} f(x)\,=\,f_0(x)\,d\big(F_0(x);F_0,F\big), \eeq here $d(u;F_0,F)$ is defined as \beq {1.4em} {1.4em} d(u;F_0,F)= \dfrac{f(Q_0(u))}{f_0(Q_0(u))}, \,0<u<1,\eeq \vskip.2em where $Q_0(u)=\inf\{x: F_0(x) \ge u\}$ for $0<u<1$ is the quantile function. The function $d(u;F_0,F)$ is called `comparison density' because it compares the initial model-0 $f_0(x)$ with the true $f(x)$ and it integrates to one: \[\int _0^1 d(u;F_0,F)\dd u \,=\, \int_x d(F_0(x);F_0,F) \dd F_0(x) \,=\,\int_x \big(f(x)/f_0(x)\big) \dd F_0(x)\,=\, 1. ~~\] \vskip.1em However, we will interpret the $d$-function as the density-sharpening function (DSF), since it plays the role of “sharpening” the initial model-0 to hedge against its potential misspecification. To simplify the notation, $d(F_0(x);F_0,F)$ of eq. (ref) will be abbreviated as $d_0(x)$.

A few remarks on density-sharpening law: \vskip.35em 1. The model building mechanism of Definition 1 provides a statistical process of transforming and refining a crude initial model into a useful one for better decision-making. \vskip.45em 2. Note that if $d(u;F_0, F) \neq 1$, i.e., if $d(u;F_0, F)$ deviates from uniform distribution then change of probability assignment is needed to embrace the current scenario. The density sharpening mechanism of (ref) prescribes how to revise the old probability assignments in light of new evidence. \vskip.45em 3. Similar to Rao's uncertainty law (ref), we can also write down a simple logical equation that captures the essence of the density-sharpening based model building principle (def. (ref)): \beq {2em} {1.85em} \fbox{Misspecified model-0} \,\,\times\, \,\fbox{Knowledge of misspecification}\, =\, \fbox{Upgraded model-1} \eeq Interpretation of the components: the first component is the starting imprecise model $f_0(x)$, coming from expert knowledge. The second component $d_0(x)$ is the quality-assurer of the model that manages the risk of misspecification of the initial $f_0(x)$. $d_0(x)$ \textit{sharpens} the decision-makers initial mental model by extracting knowledge from data that is previously unknown, which justifies its name---density sharpening function (DSF). Finally, the model-0 is “stretched” by $d_0(x)$ following eq. (ref) (only when the ideal scenario is different from the expected one) to incorporate the newly discovered information into the revised model. The class of $d$-sharp distributions turns the uncertain knowledge-distribution $f_0(x)$ into a usable distribution by properly sharpening using $d_0(x)$. Also, see Supp. A2, where a comparison between the traditional Bayes' rule and the density-sharpening-based multiplicative model update rule is presented.

Comparison Coding

The density-sharpening law provides a mechanism of building a model $f(x)$ for the data $X_1,\ldots, X_n$ by inheriting knowledge from the assumed working model $f_0(x)$. To apply the formula (ref), we need to estimate $d_0(x)$ from data.\footnote{To keep the theory of estimation simple, we will mainly focus on the $X$ continuous case. A detailed account for the discrete case can be found in D21discrete.} And we call this learning process `comparison coding' because $d_0(x)$ codes how surprising the current situation is in light of the model-0 by contrasting expectations with reality.

\vskip.25em Since the density-sharpening function $d_0(x):=d(F_0(x);F_0,F)$ is a function of $F_0(x)$, we can approximate it by a linear combination of polynomials that are function of $F_0(x)$ and orthonormal with respect to the base-model $f_0(x)$. One such orthonormal system is the LP-family of polynomials D20copula, D21discrete, Deep17LPMode, which can be constructed as follows. For an arbitrary continuous $F_0$, define the first-order LP-basis function as standardized $F_0(x)$: \beq {1.2em} {1.2em} T_1(x;F_0)\,=\,\sqrt{12} \big\{F_0(x) - 1/2\big\}. \eeq Note that $\Ex_0(T_1(X;F_0))=0$ and $\Var_0(T_1(X;F_0))=1$. Next, apply Gram-Schmidt procedure on powers of the first-order LP-basis functions $\{T_1^2(x;F_0), T_1^3(x;F_0),\ldots\}$ to construct a higher-order LP orthogonal system $T_j(x;F_0)$: \bea T_2(x;F_0) &=& \sqrt{5} \big\{ 6 F^2_0(x) - 6 F_0(x) + 1 \big\} \\ T_3(x;F_0) &=& \sqrt{7} \big\{ 20 F^3_0(x) - 30F^2_0(x) + 12F_0(x) -1 \big\} \\ T_4(x;F_0) &=& \sqrt{9} \big\{ 70F^4_0(x) - 140F^3_0(x) + 90F^2_0(x) -20F_0(x) +1 \big\}, \eea and so on. Compute these polynomials by performing the Gram-Schmidt process numerically, which can be done using readily available computer packages like R or python.

defn[Comparison coding] Expand comparison density in the LP-orthogonal series \beq d_0(x) := d(F_0(x);F_0,F)\,=\,1+\sum\nolimits_j \LP[j;F_0,F] \,T_j(x;F_0). \eeq To estimate the unknown LP-Fourier coefficient, note that: { \addtolength\abovedisplayskip{-.1\baselineskip} \addtolength\belowdisplayskip{-.3\baselineskip} \bea \LP[j;F_0,F]&=& \int T_j(x;F_0) d_0(x) f_0(x) \dd x\\ \nonumber &=& \int T_j(x;F_0) f(x) \dd x \\ \nonumber &=& \Ex_F\big[ T_j(X;F_0)\big]. \eea }

Replacing $\LP[j;F_0,F]$ with its plug-in estimator in (ref) we get \beq {1.25em} {1em} \wtd_0(x)\,=\,1+\sum\nolimits_j \tLP[j;F_0,F] \,T_j(x;F_0), \eeq where \beq \tLP[j;F_0,F] \,=\, \Ex_{\wtF}[T_j(X;F_0)] \,=\, \frac{1}{n} \sum_{i=1}^n T_j(x_i;F_0).\eeq

Although (ref) provides a robust nonparametric comparison-coding procedure, it has one drawback: the estimated $\wtd$ may be unsmooth due to the presence of a large number of small noisy LP-coefficients. To avoid unnecessary ripples in $\wtd$, we need to isolate the small number of non-zero LP-coefficients. Our denoising strategy goes as follows D21peirce: sort the empirical $\tLP[j;F_0,F]$ in descending order based on their absolute value and compute the penalized ordered sum of squares. This Ordered PENalization scheme will be referred as OPEN model-selection method: \beq OPEN(m) = Sum of squares of top $m$ LP coefficients - \dfrac{\gamma_n}{n}m. \eeq Throughout, we use AIC penalty with $\gamma_n=2$. Find the $m$ that maximizes the ${\rm OPEN}(m)$. Store the selected indices $j$ in the set $\cJ$. The OPEN-smoothed LP-coefficients will be denoted by $\widehat{\LP}_j$. Finally, return the following smoothed estimate: \beq {1.15em} {1.15em} \whd_0(x)\,=\,1+\sum\nolimits_{j \in \cJ} \hLP[j;F_0,F] \,T_j(x;F_0). \eeq

rem[The scientific value of sparse $d$] The DSF $d_0(x)$ is the bridge between the theoretical world (idealized model) and the empirical world (real observations). A meaningful way to measure the simplicity of a model is the number of “new” statistical parameters that it contains beyond the given scientific parameters---that is, the parsimony (number of parameters) of $d_0(x)$.\footnote{It selects only a handful of reasonable (rival) hypotheses out of a vast collection of possibilities. By `reasonable,' we mean hypotheses with a high OPEN($m$) score, which balances complexity (number of parameters of $d_0$) and accuracy (of explaining the surprising phenomenon).} A sparse $\whd_0$ provides an intelligent and parsimonious way to elaborate the model-0 (not an indiscriminate, brute-force elaboration) to produce a `sophisticatedly simple' model. Simplicity is vital to make the model usable and interpretable by decision-makers, who like to understand how to change the initial model to explain the data.

A Deep Dive into Model Uncertainty

Understanding the deficiency of the current model is an essential part of the process of iterative model building and refinement: Have we overlooked something? Where are our knowledge gaps? This section provides a comprehensive understanding and exploratory tool for representing and assessing potential model misspecifications. Also, see Supp. note A1, which explains the distinctions between parametric and nonparametric model uncertainties.

figure[figure omitted — 857 chars of source]

Graphical Exploration of Model Uncertainty

exampleConsider the following scenario: Fig. (ref) displays the data that a physicist just collected from an experiment. The blue curve is the physics-informed background distribution $f_0(x)$, which, in this case, is an exponential distribution with $\la_0=25$, and the red curve is the true unknown probability distribution. The physicist is mainly interested in knowing whether there is any new physics hidden in the data, i.e., anything new in the data that was overlooked by existing theory. If so, what is it? How does the theory ($f_0$) relate to practice? This will help the physicist to come up with some scientific explanations and potential alternative theories. The Shape of Uncertainty. The researcher ran the density-sharpening algorithm of the previous section with $m=10$, and the resulting $\whd_0(x)$ is displayed in the right of Fig. (ref) as a function of $F_0(x)$. Few conclusions: (i) {\bf Model appraisal}: The non-uniformity of $\whd$ tells us that the “shape of the data” is inconsistent with the presumed model-0. (ii) {\bf Model amendment}: The shape of $\whd$ also informs the scientist about the nature of deficiency of the old model---i.e., what are the most worrisome aspects of the presumed model? In this example, the most consequential unanticipated pattern is the presence of a prominent `bump' (excess mass) around $F_0^{-1}(0.63) \approx 24.85$, which might be indicative of new physics. This newly discovered pattern can now be used to improve the background exponential model.
rem[Visual explanatory decision-aiding tool] One of the unique abilities of our exploratory learning is its ability to generate explanations on why and how the model-0 is incomplete\footnote{Explanation-based statistical reasoning is at the core of abductive inference, as discussed later.}. Thus, the graph of $\whd(u;F_0,F)$ explicitly addresses decision-makers model misspecification concerns. It digs into the observations to uncover the “blind spots” of the current model that can ultimately drive discovery (locating novel hypotheses) and better decisions.

Measure of Model Uncertainty

A general measure of the degree of model misspecification is defined using the Csisz{\'a}r information divergence class.

defnFor $\psi: [0, \infty) \mapsto \bbR$ a convex function with $\psi(1)=0$, define the Csisz{\'a}r class of statistical divergence measure between $F$ and $F_0$: \beq I_{\psi}(F,F_0)\,=\, \int_{-\infty}^\infty \psi\left( \frac{f(x)}{f_0(x)} \right) f_0(x) \dd x\eeq We prefer to represent it in terms of density-sharpening function as follows: \bea I_{\psi}(F,F_0)&=& \int_{-\infty}^\infty \psi \circ d(F_0(x);F_0,F) \dd F_0(x) \nonumber \\ &=& \int_0^1 \psi \circ d(u;F_0,F) \dd u, where u=F_0(x). \eea

One can recover popular divergence measures by appropriately choosing the $\psi$-function:

itemize[itemsep=8pt] • KL-divergence: $\psi(x)=x\log(x)$; $I_{{\rm KL}}(F,F_0)=\int d \log d$. • Total variation divergence: $\psi(x)=|x-1|$; $I_{{\rm TV}}(F,F_0)=\int |d - 1|$. • Squared Hellinger distance: $\psi(x)=(1-\sqrt{x})^2$; $I_{{\rm H}}(F,F_0)=\int (1 - \sqrt{d})^2$. • $\chi^2$-divergence: $\psi(x)=(x-1)^2$; $I_{\chi^2}(F,F_0)=\int (d - 1)^2 = \int d^2 - 1$.

\vskip.4em

One can quickly estimate the $\chi^2$-model misspecification index by expressing it in terms of LP-Fourier coefficients (applying Parseval's identity to equation (ref)): \beq {1.5em} {1.5em} I_{\chi^2}(F,F_0)= \int d^2 - 1 = \sum_{j=1}^m \big| \LP[j;F_0,F] \big|^2. \eeq $I_{\chi^2}(F,F_0)$ quantifies the uncertainty of the preliminary model $f_0(x)$ in light of the given data---i.e., whether $f_0(x)$ is catastrophically wrong or slightly wrong. Estimate it by plugging the empirical LP-coefficients (ref) into (ref). Since, under $H_0: F=F_0$, the sample LP-coefficients have the following limiting null distribution (see Theorem 2 of Deep17LPMode): \[ \setlength{\abovedisplayskip}{1.5em} \setlength{\belowdisplayskip}{1.5em} \sqrt{n} \tLP[j,F_0,F] \xrightarrow[]{d} \cN(0,1),~~ \textrm{i.i.d for all}~ j,\] $n \widetilde{I}_{\chi^2}(F,F_0)$ follows $\chi^2_m$ under null. One can use this to compute the $p$-value. Applying this measure to example 1, we get a $p$-value of practically zero---indicating that the background exponential model is badly damaged and should be repaired before making a decision.

remIt is interesting to contrast our theory with hansen2022risk, keeping in mind that their $m(x)$ is exactly our sharpening function $d_0(x)$. This reinforces our belief that our theory can be applied broadly to econometrics and decision-making under uncertainty.

$d$-Sharp Models

defn$\DS(F_0,m)$ stands for {\bf D}ensity-{\bf S}harpening of $f_0(x)$ using $m$-term LP-series approximated $d_0(x)$, given by: \beq {1.3em} {1.5em} f(x)\,=\,f_0(x)\Big[ 1\,+\, \sum_{j=1}^m \LP[j;F_0,F]\, T_j(x;F_0)\Big], \eeq obtained by replacing (ref) into (ref). $\DS(F_0,m)$ generates a relevant class of plausible models in the neighbourhood of the postulated $f_0(x)$ that are worthy of consideration.

A few additional points on density-sharpening: \vskip.14em

1. The $\DS(F_0,m)$-based density-sharpening principle provides a mechanism for exploring data by exploiting the uncertain background knowledge model. It starts with data and an approximate model $f_0(x)$---and produces a more refined picture of reality following (ref).

\vskip.24em

2. The process of density-sharpening suitably `stretches' the theory-informed model to create a class of robust empirico-scientific models. Moreover, it shows how new models are born out of data-driven mutation of pre-existing ones.

\vskip.24em

3. The truncation point $m$ indicates the radius of the neighborhood around the elicited $f_0(x)$ to create permissible models. $\DS(F_0,m)$ models with higher $m$ entertain alternative models of higher complexity. However, to maintain conceptual appeal and interpretability, it is advisable to focus on the vicinity of $f_0$ by choosing an $m$ that is not too large. Substituting the smooth estimates $\hLP[j;F_0,F]$ of eq. (ref) into the formula (ref), we get the most economical model (among competing alternatives around $f_0$) that best explains the empirical surprise.\footnote{It brings our theory close to Gilbert Harman's “Inference to the best explanation” idea; see harman1965. This is an area that merits further research.}

\vskip.24em 4. It provides an architecture of an `intelligent agent' that simultaneously possesses the ability to: learn (what's new can we learn from the data), reason (how to explain the surprising empirical findings), and plan (how to self-modify to adapt in the new situations).

example[Glomerular filtration data] We are given glomerular filtration rates\footnote{Glomerular filtration rate (GFR) measures how much blood is filtered through the kidney to remove excess wastes and fluids. Low gfr value indicates that the kidneys are not functioning as well as they should.} for 211 kidney patients. The experiment was done at Dr. Bryan Myers' Nephrology research lab at Stanford University. The dataset was previously analyzed in efron2016computer. \vskip.33em The blue curve on the left plot of Fig. (ref) shows the best-fitted lognormal (LN) distribution. We start our analysis by asking whether the parametric LN model needs to be refined to fit the data. The middle panel displays the density-sharpening function, which provides insights into the nature of misspecification of the LN model: the peak and the tails of the initial LN distribution need repairing; LN underestimates the peak and neglects the presence of heavier tails. The repaired LN model (displayed on right-hand side of Fig. (ref)) is given by \beq {.7em} {.7em} \hf(x)\,=\, f_0(x) \big[ 1 \,+\, 0.18 T_4(x;F_0) \big], \eeq where $f_0(x)$ is ${\rm LN}(\mu_0,\sigma_0)$, with $\mu_0=4$ and $\sigma_0=0.24$. The part in the square bracket comes from $d_0(x)$, which provides recommendations on how to suitably elaborate the LN-model to capture the unexplained shape. The point of this example was to show how the density-sharpening principle (DSP) allows an analyst to explicitly perform model formulation, fitting, checking, and repairing---all seamlessly combined into one workflow.

It is interesting to compare our $d$-sharp LN-model (the red curve) with the seven-parameter exponential family fit shown in Fig. 5.7 of efron2016computer. The most noticeable difference lies in the right tail. Efron's seven-parameter exponential family model shows weird spikes on the extreme-right tail. The main reason for this is that it is based on polynomials of raw $x$: ($x,x^2,\ldots,x^7$), which are not robust. That is to say, these traditional bases are unbounded and highly sensitive to `large' data points. In contrast, our LP-polynomials are functions of $F_0(x)$, not raw $x$, and thus robust by design. The other operational difference between our approach and Efron's exponential family approach is that we model the “gap” between lognormal and the data, which is often far easier to approximate nonparametrically (only required one parameter, see eq. (ref)) than modeling the data from scratch.\footnote{There is an easy way to see that: compare the shapes of the histograms of the left two plots of Fig. (ref).}

figure[figure omitted — 700 chars of source]

Modelplasticity and Abductive Inference Machine

quoteNot the smallest advance can be made in knowledge beyond the stage of vacant staring, without making an abduction at every step. {\rm --- C. S. peirce1901proper}

{\bf Modelplasticity}---Models ability to modify and adapt itself in response to new data. The density-sharpening principle enables the model to develop new shapes in the face of change.

{\bf Density-sharpening and model evolution}. Modeling is a continual process, not a one-time data-fitting exercise. The density sharpening mechanism allows us to combine new observations with a priory expected model to generate new insights, as depicted in Fig. (ref).

center[center omitted — 1,439 chars of source]

{\bf Statistical law of model evolution}. Density-sharpening supports this dynamic process of recursive model upgrading: $f_{k}(x) = f_{k-1}(x) \,d_{k-1}(x)$, for $k=1,2,\ldots$, by allowing the model to constantly evolve and reshape itself with fresh sets of data---going from a simple approximate model to a much more mature, accurate model of reality.

\vskip.2em {\bf Abduction and creation of new knowledge}. Abduction is the creative part of an inferential process that aims at producing new theories from data. It builds upon what we know to discover new facts about nature. Abductive learning is concerned with the following questions: What new can we learn from the data? How to change the prior hypothetical model to explain the current situation? Which alternative classes of models are worthy of being entertained?

quote{`Does statistics help in the search for an alternative hypothesis? There is no codified statistical methodology for this purpose. Text books on statistics do not discuss either in general terms or through examples how to elicit clues from data to formulate an alternative hypothesis or theory when a given hypothesis is rejected.'} \begin{flushright} {\rm --- C. R. rao2001statistics} \end{flushright}

Charles Sanders Peirce (1837–1914) was the pioneer of abductive reasoning; see stigler1978 and D21peirce for more details on the Peircean view of statistical modeling. The goal of Abductive Inference Machine\footnote{They are not traditional pattern recognition (or matching) engine, they are pattern discovery engine.} or AIM is to provide a learning framework that endows a model with this ability to learn, grow and change with new information. \vskip.25em

remThe density-sharpening process plays an essential role for abductive inference, which provides the computational machinery for generating novel hypotheses with explanatory merit and selecting specific ones for further examinations.
rem[Abductive inference $\neq$ Hypothesis testing] Any scientific inquiry begins with observations and some initial hypotheses. Classical statistical inference develops tools to test the validity of the null model in light of the data. Since all scientific theories are incomplete, accepting or rejecting a particular hypothesis is a pointless exercise. The real question is not whether the null hypothesis is true or false. The real question is: how far is the reality from the postulated model? In which direction(s) should we search to find a better model? Density-sharpening law provides a process of progressive refinement of yesterday's hypothesis.

Attention Mechanism

quoteWe often neglect how we get rid of the things that are less important...And oftentimes, I think that’s a more efficient way of dealing with information. \begin{flushright} --- {\rm Duje Tadin\footnote{Jordana Cepelewicz (2019) To Pay Attention, the Brain Uses Filters, Not a Spotlight Quanta Magazine, https://www.quantamagazine.org/to-pay-attention-the-brain-uses-filters-not-a-spotlight-20190924.}} \end{flushright}

Attention is the prerequisite of gaining new knowledge. Intelligent learners have the ability to quickly infer where to focus attention to gain knowledge. In our modeling framework $d_0(x)$ draws analyst's attention quickly and efficiently to the new informative part by suppressing boring details; verify it from the graphs of $d_0(x)$ in Figs. (ref) and (ref). It acts as a `gating mechanism' that filters out the new interesting (surprising) aspects of the data, and ignores the dull and unsurprising part---thereby sharpening the model's intelligence by guiding where to pay attention for information processing.

quote“The whole function of the brain is summed up in: error-correction” \vskip.24em --- {\rm W. Ross Ashby}, English psychiatrist and a pioneer in cybernetics.
remIn the brain, a dedicated circuit (or system) performs information-filtering similar to what $d_0(x)$ does for our dyadic model. The existence of such a brain circuit was first hypothesized by Francis crick1984function---he called it `The Searchlight Hypothesis.' Since then, significant progress has been made to hunt down the brain region, what is now called basal ganglia, that suppresses irrelevant inputs. For more details see halassa2017 and gu2021com. Basal ganglia help us focus on what's important and tune out the rest. The mechanics of our model-building mimic the brain's cognitive process that uses existing knowledge to sieve out the new information for correcting the error (sharpening) of the earlier mental model.

Decision-Making with Imperfect Model

quoteHow should a decision maker acknowledge model misspecification in a way that guides the use of purposefully simplified models sensibly? \begin{flushright} {\rm --- larscerreia2020} \end{flushright}

This section demonstrates how practicing abductive inference based on the density-sharpening principle can enable better decision-making in highly uncertain environments.

Abductive Model of Decision Making

Abduction is the process of generating and revising a model before choosing the optimal action. An abducer makes decisions in a dynamic uncertain environment by allowing for potential model misspecification.\footnote{The importance of model uncertainty in economics, finance, and business is beautifully illustrated in hansen2014, although from a different perspective.} Abductive decision-making is about knowing when to change course and how to change it.

\vskip.35em How can a decision-maker abduct? The mechanics of abductive decision-making consist of three steps: (i) generating a set of plausible alternative models based on new evidence; (ii) constructing a `robust' model (by choosing the least favorable alternative model or by averaging the alternative models with proper weights); and (iii) selecting an action that maximizes expected utility under the newly revised model. Two modes of abductive decision-making under uncertainty are presented below.

\vskip.35em

{\bf Notation}. A decision-maker (DM) has to take an action $a$ from the set of available actions $\mathbb{A}=\{a_1,\ldots,a_q\}$ based on observed outcome $X_1,\ldots,X_n$ from an unknown probability distribution $f(x)$, representing some natural or social phenomenon. The DM selects the optimal action that minimizes expected loss (or risk) under the assumed model-0: \beq \ha_0 := \argm_{a \in \Abb} \int L_a(x) \dd F_0(x), \eeq where $f_0(x)$ is the DM's posited probability distribution over outcomes. However, as an abducer, the DM is completely aware that the uncertainty about the outcomes may not be fully captured by a single, rigidly-defined probability distribution $f_0(x)$ and thus wants to choose the best decision by accommodating the uncertainty of model-0.

\vskip.35em {\bf Decision making based on density sharpening principle}. To account for the imperfect nature of model-0, the most natural thing to do is to work with an enlarged class of plausible distributions around the vaguely acceptable $f_0(x)$: \beq \Gamma_M = \big\{ f: \, f \in \DS(F_0,m), m \le M \big\} \eeq within a certain reasonable neighbourhood, say $M=10$. We like to use this enlarged class of distributions $\Gamma_M$ for robust decision-making. Two such strategies are discussed below.

\vskip.35em {\bf Method 1.} A cautious DM selects an action by its expected loss under the least favourable distribution within the set $\Gamma_M$: \beq \breve{f}_{a,M}\, =\, \argsup_{F \in \Gamma_M} \int L_a(x) \dd F(x). \eeq We call this an abductive-minimax procedure. Our proposal is partly inspired by the `local-minimax' idea of hansen2001, hansen2001robust.

\vskip.35em {\bf Method 2.} We now describe another robust decision-making procedure that takes into account the uncertainty in the analyst's elicited probability model of future states. Two key concepts are: bootstrap model averaging and action-profile function. \vskip.25em Step 1. We use bootstrap to explore $f \in \Gamma_M$ in an intelligent way. Draw $n$ samples with replacement from the original data. Denote the bootstrap empirical cdf as $\wtF_*^{(1)}$. Perform density-sharpening algorithm based on $\wtF_*^{(1)}$, and denote the estimated $d$-sharp model as $f_*^{(1)}$.

Step 2. Use $f_*^{(1)}$ to select the best action from the given set of $q$-actions $\{a_1,\ldots, a_q\}$. Denote the selected action as $a_*^{(1)}$.

Step 3. Repeat steps 1-2, $B$ times (say $B=1000$ times). And return:

itemize[itemsep=4pt] • The sample bootstrap distribution $p_A$ of optimal actions $\{a_*^{(1)}, \ldots, a_*^{(B)}\}$---which we call the action profile of the decision problem. • Bootstrap systematically generates probable alternative models $\{f_*^{(1)}(x),\ldots, f_*^{(B)}(x)\}$ from the class $\Gamma_M$ that can explain the data. Compute bootstrap model averaged distribution\footnote{This is also known as bagging breiman96bagging or bootstrap smoothing efron2014est.}: \beq \bar f(x)\,=\, \frac{1}{B}\sum_{j=1}^B f_*^{(j)}(x). \eeq By averaging all plausible alternatives, $\bar f(x)$ becomes robust to model uncertainty. In this strategy, the policymaker does not have to put his/her complete faith in a single alternative distribution to assign probabilities. Bootstrap density exploration generates and weights different alternatives (from the class $\Gamma_m$) appropriately to create a realistic model. Fig. (ref) shows the bootstrap-generated densities for the gfr data of example 2. The light blue curves are the plausible alternative models, and the dark blue is the averaged density that takes into account all likely scenarios. It's worth contrasting $\bar f(x)$ with more traditional parametric uncertainty-based Bayesian predictive density; see supplementary A1.
figure[figure omitted — 526 chars of source]
remTwo remarks: (i) Rather than assuming that alternative probable models are given to us a priori, we use density-sharpening-based abductive inference method to synthesize them, allowing us to account for a much broader range of model uncertainty than is possible with conventional approaches like Bayesian model averaging; see Supp. note A4. (ii) Furthermore, in our scheme, bootstrap provides a way to automatically estimate the posterior probability of accepting each synthesized alternative model (light blue curves in Fig. (ref)), eliminating the need for arbitrary subjective probabilities over models. To compute the $\bar f$, bootstrap-derived model posterior probabilities are used to weigh the various alternatives from the density-set $\Gamma_M$; see Supp. note A3.

Step 4. This bootstrap scheme can also be used to approximately compute the least favourable distribution, defined in (ref): \[ \breve{\breve{f}}_{a,M} = \argmax_j \int L_a(x) \dd F_{*}^{(j)}(x).\] The decision-maker can use this estimated model to carry out the proposed abductive-minimax procedure (method 1).

Step 5. {\bf Robust procedure}\footnote{Our philosophy of robustness is in complete agreement with huber77robust, who advocated distributional robustness: “one would like to make sure that methods work well not only at the [idealized parametric] model itself, but also in a neighborhood of it.’’}: A pragmatic\footnote{Pragmatism is the logic of abduction.} decision-maker chooses an action (or ranks the actions) that minimizes expected loss (or maximizes the expected utility) with respect to the averaged-distribution: $\ha_{{\rm robust}} := \argm_{a \in \Abb} \int L_a(x) \dd \bar{F}(x)$. Our strategy prescribes action that is robust across a wide range of plausible alternative models. It could be especially powerful for dealing with “deep uncertainty” in making robust policies. For a comprehensive overview on this subject, see marchau2019decision.

Step 6. Quantifying the `robustness' of the action (or decision rule): How much does the optimal action change when a model is selected from a reasonable neighborhood of the assumed initial opinion, i.e. from the $\Gamma_m$ class? The shape of the action profile distribution can be used to determine how robust the optimal action is to model perturbation. In particular, the entropy of the action profile distribution can be used to assess the robustness (or stability) of the inference to potential model misspecification:

\beq {\rm Entropy}[p_A] \,=\,- \sum_{i=1}^q p_A(i) \,\log p_A(i) = \,- \sum_{i=1}^q \Pr(A=a_i) \,\log \Pr(A=a_i). \eeq Uniform probability over possible actions yields maximum uncertainty---indicating that the decision is highly non-robust (sensitive) to model misspecification.

Quantile Decision Analysis

Until now, we have assumed experts can precisely formulate their opinion in a probabilistic form $f_0(x)$. However, for complex real-world decision-making problems, experts might only have incomplete information about the uncertainty distribution of the target variable. A decision-analyst often elicit their partial knowledge about an uncertain quantity as a set of quantile-probability (QP) pairs $\{x_i, F(x_i)\}$, for $i=1,\ldots,\ell$. The job of an analyst is to find a simple, flexible, and parameterizable density that honors the assessed percentiles. \vskip.3em

\vskip.25em

{\bf Model ambiguity due to incomplete information}. The task of eliciting an expert's probability distribution from a small set of QP pairs is a vital yet nascent topic in decision analysis; see powley2013quantile, keelin2011quantile, hadlock2017quantile. In this section, we present an algorithm called Q2D (stands for quantile to distribution) that provides a systematic approach to deduce a reliable expert distribution from $\ell$ arbitrary QP-specifications.

\vskip.35em {\bf Probability-gap Approximation}. The main theoretical idea behind Q2D algorithm: Recall our $\DS(F_0,m)$ model \beq f(x)\,=\,f_0(x)\Big[ 1\,+\, \sum_{j=1}^m \LP[j;F_0,F]\, T_j(x;F_0)\Big] \eeq Integrating from minus infinity to $x$ on both sides, we have \[ \int_{-\infty}^x ( f(z) - f_0(z) ) \dd z = \sum_{j=1}^m \LP[j;F_0,F] \int_{-\infty}^x S_j(F_0(z)) \dd F_0(z),\] where $S_j(u)=T_j(Q_0(u); F_0)$ is defined over the unit interval $[0,1]$ and $Q_0(u)$ is the quantile function of the distribution $f_0$. This leads to \beq F(x) - F_0(x)\,=\, \sum_{j=1}^m \LP[j;F_0,F] \int_0^{F_0(x)} S_j(u) \dd u. \eeq

Given a set of arbitrary $\ell$ quantile-probability data $(x_i,F(x_i)),$ for $i=1,\ldots,\ell$, we can rewrite (ref) compactly as a matrix equation \beq v = S_0 \bbe \eeq where $v_i=F(x_i) - F_0(x_i)$, $\be_i=\LP_j$, and $S_0 \in \cR^{\ell\times m}$, $S_0[i,j]=\int_0^{F_0(x_i)} S_j(u)$. The desired parameters are ${\bm \be}=(\be_1,\ldots,\be_m)$, where $\be_j$ is shorthand for $\LP[j;F_0,F]$.

For $m \le \ell$, we can uniquely estimate $\bbe$ using the least-square method \beq \widetilde{\bbe}\,=\,\miniz_{\bbe} \| v - S_0 \bbe \|^2 \, =\, (S_0^TS_0)^{-1}S_0^{T}v. \eeq For large $\ell$ (say, $\ell\ge 5$), a better, more stable estimate can be found through regularization \beq \widehat{\bbe}\,=\,\miniz_{\bbe} \| v - S_0 \bbe \|^2 + \la \|\bbe\|_1 \eeq where $\| \cdot\|_p$ is the $\ell_p$ norm, and $\la>0$ is the regularization parameter. The lasso lasso1996 penalized $\widehat{\bbe}$ yields a sparse estimate and counters over-fitting. This penalized estimate provides a tradeoff between accuracy and interpretability.\footnote{Note that due to regularization, Eq. (ref) can even tackle cases with $m > \ell$. In such scenarios, the OLS is ill-posed, with an infinite number of solutions.} Finally, plug the estimated LP-Fourier coefficients $\be_j$ into the primary equation (ref) to get the expert distribution.

remThe expert quantile specifications should not be viewed as a `gold standard'---they are nothing but a preliminary guess (prone to errors of judgment or hindsight bias) whose purpose is to steer the analyst in the right direction\footnote{winkler1967 emphasized that the expert does not have some `true' density function waiting to be elicited, only a `satisficing' initial distribution that the policymaker is `content to live with at a particular moment of time.'}. For that reason, we recommend the smoothed (denoised) regularized $\widehat{\bbe}$ over the naive $\widetilde{\bbe}$, since it makes little sense to find an exact fit to the noisy QP-data.
figure[figure omitted — 625 chars of source]
example[Bimodal Distribution] We are given the following quantile judgments: \begin{table}[h] \vskip.3em \begin{tabular}{l | ccccccc} Quantile: $x_i$ & -3.40 & -2.53 & -1.20 & 0 & 2.0& 2.83& 3.60 \\[.4em] \hline & \\[\dimexpr-\normalbaselineskip+.7em] Probability: $F(x_i)$ & 0.04 &0.15& 0.39& 0.50 &0.75 &0.90& 0.97 \end{tabular} \end{table} In our Q2D algorithm, we choose $F_0$ (an initial approximate shape) to be normal distribution. To estimate the parameters $\mu_0$ and $\sigma_0$ of the normal distribution, note that the quantile function $Q(u) \approx \mu_0 + \sigma_0 \Phi^{-1}(u)$. Thus one can quickly get a rough estimate by simply performing a linear regression\footnote{This technique will work for any location-scale family $f_0(x)$, e.g. normal, Laplace, logistic, etc.} on $(\Phi^{-1}(u_i), Q(u_i))$; see Fig. (ref). The estimated normal distribution is shown in the right panel, along with the Q2D-estimated density.
example[U.S. Navy data] Fig. (ref) shows a histogram of 122 repair times (in hours) for a component of a U.S. Navy weapons system. The dataset was analyzed in law2011select. Imagine that for privacy and other reasons, we do not have access to the full data. The goal is to infer a probability distribution that faithfully represents the following quantiles: \begin{table}[h] \vskip.3em \begin{tabular}{l | ccccc} Quantile: $x_i$ & 0.12 & 1.30& 3.00& 7.00& 26.17 \\[.4em] \hline & \\[\dimexpr-\normalbaselineskip+.7em] Probability: $F(x_i)$ & 0.01 & 0.20& 0.50& 0.80& 0.99 \end{tabular} \end{table} We start with exponential distribution as our initial guess, which is often taken as a `default' distribution (model-0) in reliability analysis. For $X \sim {\rm Exp}(\la)$, we have \[{\rm Median}(X)\, =\, \la \ln(2),~~{\rm where}~ \la=\Ex(X).~~\] From the quantile table we get $\hat \la = 3/\ln(2) = 4.32.$ Next, we apply the Q2D algorithm to derive the LP-parameters with $f_0={\rm Exp}(4.32)$. The resulting density sharpening function and the final $d$-sharp exponential are shown in Fig. (ref). The red curve on the right plot shows an excellent fit to the data, which was derived by the Q2D algorithm simply by utilizing the five quantile-probability pairs.
figure[figure omitted — 767 chars of source]

Decision-making based on Multiple Experts

High-stakes decision-making (say, COVID-19 pandemic or climate change) is often based on multiple experts' opinions instead of putting all bets on a single rigidly-defined probability model. The challenge is to aid data-driven decision-making by appropriately combining several experts' models. We describe one possible way to build a `consensus committee model' that can be used as a possible model-0 within an abductive decision-making framework.

\vskip.4em

{\bf Learning from multiple expert distributions}. Given $k$ competing probability models $\{f_{01},\ldots, f_{0k}\}$, which may differ markedly in shape, define the following model-weights: \beq Relevance weight: w_\ell = \dfrac{1}{1 + \sum_j |\LP_{j|\ell}|^2}, {\rm for} \ell=1,\ldots,k \eeq where $\LP_{j|\ell}$ is the LP-Fourier coefficients of the $\ell$-th model: \beq d_{0\ell}(x) := d(F_{0\ell}(x); F_{0\ell},\wtF)\,=\, 1+\sum_j \LP_{j|\ell} T_j(x;F_{0\ell}). \eeq Note that the relevance weight for the $\ell$-th model is always $0 < w_\ell \le 1$, and \[ w_\ell=1 ~~~\text{if and only if}~~~\LP_{j|\ell}=0,~\forall~ j.~~ \] $\LP_{j|\ell}=0$ for all $j$ when $f_{0\ell}$ fully explain the data and there is no need to sharpen it further (i.e., $d_{0\ell}=1$). In that sense, $w_\ell$'s are data-driven weights (which will keep changing as we get more and more fresh data), computed based on the degree of agreement between the observed data and expert model $f_{0\ell}$. Define mixture expert distribution as

\beq f^0_{\rm mix} (x) = \sum_{\ell=1}^k \pi_\ell\, f_{0\ell} (x), \eeq where $\pi_\ell = w_\ell/\sum_\ell w_\ell$. This model serves two purposes: it tries to resolve conflicting opinions based on data and at the same time encourages one to include as much diverse information as possible.

Additional layer of model uncertainty. If the analyst believes that the correct model might not be among the collection of models being considered, then use the combined expert model $f^0_{\rm mix} (x)$ as a model-0 in the subsequent density-sharpening-based learning and abductive decision-making process.

Model Management Science

How should an analyst use imperfect models to learn from data?\footnote{The challenge of learning from uncertain knowledge is also a fundamental issue in the development of intelligent systems.} What should be the output of such an analysis that can ultimately aid informed decision-making? We address these questions by introducing a general inferential framework for statistical learning and decision-making under uncertainty---which builds on two core ideas: abductive thinking and density-sharpening principle. Some of the defining features of our approach for data analysis, scientific discovery, and decision-making are highlighted below:

$\bullet$ Data analysis and science of model management: No model is perfect, irrespective of how cunningly it is designed. The central problem of statistical model developmental process is to understand how a relatively simple model can evolve into a more complex and mature one in the presence of a new data environment. The principle of density-sharpening assists this model evolution process (thereby helping empirical scientists to abduct): by abductively generating explanations on why the presumed model-0 is unfit for the data [playing the role of a quality inspector] and also providing recommendations on how to fix the misspecification issues [serving as a policy adviser] in order to make better decisions in new circumstances.

\vskip.77em $\bullet$ Discovery and creation of new knowledge: Abductive data analysts are less interested in testing a particular working model. They are mainly interested in conceptual innovation: discovering new pursuitworthy hypotheses based on surprising empirical evidence.\footnote{A largely unexplored topic relative to the vast literature on hypothesis testing. As noted by George E. P. box2001dis: “Much of what we have been doing is adequate for testing but not adequate for discovery.”} The density-sharpening function $d(u;F_0,F)$ picks out `what's new' in the data beyond the current scientific knowledge encoded in $f_0(x)$, thereby helping the scientist to uncover new unexpected knowledge from the data using graphical tools. The density-sharpening principle (DSP) provides a learning mechanism that isolates the `known' from the `unknown' and allows us to focus on the newfound pattern in the data, which is the basis for knowledge-creation\footnote{Curious readers are invited to read the paper “Nobel Turing Challenge: creating the engine for scientific discovery” by Hiroaki kitano2021nobel, where he argued that the single-most-important mission of AI is to accelerate scientific discovery.}.

\vskip.77em $\bullet$ Abductive inference and decision-making: The proposed theory of abductive decision-making tackles model uncertainty induced by imprecise, ambiguous, and incomplete knowledge about the underlying probabilistic structure. An abductive-decision support system automatically discovers and explicitly articulates the possible alternatives to the analysts, which forces them to rethink their choices before taking impulsive action; see Supp. note A5. This style of empirical reasoning and adaptive decision-making could be especially valuable in situations where strategic planners need to take quick action in the face of uncertainty, equipped with approximate subject-matter knowledge.

Code and Data availability

All the datasets and R-code written for the analysis are available upon request to the author.

Ethical Statement

The Author declares that there is no conflict of interest/competing interests.

Supplementary material

For more discussion on how our abductive statistical approach compares to the more traditional Bayesian statistical approach for model misspecification, robustness, and decision-making, see the Supplementary section.

\setcounter{page}{1} \setcounter{equation}{0}