Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,664 characters · 18 sections · 35 citation commands
Modelplasticity and Abductive Decision Making$^*$ footnote-1
}
}
\noindentKeywords: Abductive Decision-making; Model Risk Management; The Uncertainty Principle; Density-Sharpening Principle; Creation of New Knowledge; Quantile Decision Analysis.
\vskip1em
How to make decisions under uncertainty? Decision-making under uncertainty relies mainly on how efficiently we can extract useful knowledge from the data that were previously unknown to the decision-maker\footnote{“Anything that gives us new knowledge gives us an opportunity to be more rational” – Herbert Simon}. C. R. Rao, in his rao1996 article\footnote{The existence of this article is barely known.} on `Uncertainty, Statistics, and Creation of New Knowledge' provided an exquisite description of the mechanics of decision-making under uncertainty using a simple logical formula:
\beq {-1.25em} {1.5em} \fbox{Uncertainty of knowledge} + \fbox{Knowledge of uncertainty}\, =\, \fbox{Usable knowledge.} \eeq
A decision analyst confronts data $X_1,\ldots, X_n$ equipped with a tentative (imprecise and uncertain) probabilistic model $f_0(x)$ of the underlying phenomena. The challenge then boils down to effectively using the misspecified model $f_0(x)$ to learn from data and to apply that knowledge for informed decision-making. Rao's uncertainty principle suggests the following three-staged approach, which we call the `model-building triad':
\vskip.75em
Stage 1. Model Elicitation. The first step of decision making is formulating a probability model of the phenomena of interest, from which economic agents derive their initial expectations. A simple parametrized model-0 $f_0(x)$ is usually formed based on either gut instinct or the scientific context of the investigation. The uncertainty of $f_0(x)$ arises due to the lack of perfect knowledge\footnote{what keynes1937general called `uncertain knowledge.'} about the underlying probability law. Accordingly, the modeler has to start the analysis by acknowledging the uncertainty of the initial knowledge model $f_0(x)$.
\vskip.45em
Stage 2. Model Uncertainty Quantification. Before making decisions based on the provisional model $f_0(x)$, it is crucial to investigate its uncertainty (blind spots) in light of the new data. It's always a good practice to inspect expert opinions based on hard empirical facts by asking\footnote{Those who ignore experts' knowledge and only trust data are empirical-fools. Those who ignore data and only trust their gut-instinct are emotional-fools tversky1974. Expert decision-makers always use empirically-guided intuition by appropriately combining both data and available knowledge.}: what's new in the data that can't be explained by the assumed model? Discovering surprising and previously unknown facts can prompt decision makers to consider other alternative actions.
\vskip.45em
Stage 3. Model Rectification and Risk Management. Finally, we incorporate the learned uncertainty into the uncertain model $f_0(x)$ to produce a rectified model for making empirically-guided informed decisions. It is important to sharpen the “judgment component” (intuition based on past experiences) in light of the new data before it gets outdated.
\vskip.75em
The purpose of this article is to describe a general statistical theory that permits us to implement this three-staged model-building procedure for data analysis and decision-making.
\vskip1em
Empirical scientific inquiry typically starts with a simple yet believable model of reality (model-0) and aims to sharpen existing knowledge by gathering new observations. \vskip.25em We observe a random sample $X_1,\ldots,X_n \mathrel{\dot\sim} F_0$. By “$\mathrel{\dot\sim}$” we mean that $F_0$ is an `approximately correct' structured provisional model for $X$ that is given to us by subject-matter experts. We like to extract new knowledge from the data by smartly leveraging existing knowledge\footnote{Model amendment principle: the starting model $f_0(x)$ is incomplete but not useless. It contains valuable background knowledge. Rather than throwing this vital information, we want to build a model by smartly taking clues from it. The goal is to amend model-0, not to abandon it completely.} that is encoded in the initial approximate model $f_0(x)$.\footnote{As for notation: by $F_0(x)$, we denote the cumulative distribution function (cdf) of the starting model-0; $f_0(x)$ is the probability density function (pdf) and quantile function is denoted by $Q_0(u)$. The expectation with respect to $f_0(x)$ will be abbreviated as $\Ex_0$.} \vskip.25em
Creating knowledge-guided statistical models. The core mechanism of our process involves: (i) inspecting whether the structured provisional model-0 is still a good fit in light of fresh data; (ii) if not, then we like to know what's new in the data that cannot be tackled by the current model; and, finally, (iii) repair the current misspecified model in order to cope with the new reality. However, the question remains as to how can we design an inference machine that can offer these successively fine-grained insights? To address this question, we will describe a new statistical model building principle, called the `density-sharpening principle.'
We introduce a dyadic model with two interrelated subsystems that accommodates the decision maker's concern for misspecification of the starting expert-guided model.
A few remarks on density-sharpening law: \vskip.35em 1. The model building mechanism of Definition 1 provides a statistical process of transforming and refining a crude initial model into a useful one for better decision-making. \vskip.45em 2. Note that if $d(u;F_0, F) \neq 1$, i.e., if $d(u;F_0, F)$ deviates from uniform distribution then change of probability assignment is needed to embrace the current scenario. The density sharpening mechanism of (ref) prescribes how to revise the old probability assignments in light of new evidence. \vskip.45em 3. Similar to Rao's uncertainty law (ref), we can also write down a simple logical equation that captures the essence of the density-sharpening based model building principle (def. (ref)): \beq {2em} {1.85em} \fbox{Misspecified model-0} \,\,\times\, \,\fbox{Knowledge of misspecification}\, =\, \fbox{Upgraded model-1} \eeq Interpretation of the components: the first component is the starting imprecise model $f_0(x)$, coming from expert knowledge. The second component $d_0(x)$ is the quality-assurer of the model that manages the risk of misspecification of the initial $f_0(x)$. $d_0(x)$ \textit{sharpens} the decision-makers initial mental model by extracting knowledge from data that is previously unknown, which justifies its name---density sharpening function (DSF). Finally, the model-0 is “stretched” by $d_0(x)$ following eq. (ref) (only when the ideal scenario is different from the expected one) to incorporate the newly discovered information into the revised model. The class of $d$-sharp distributions turns the uncertain knowledge-distribution $f_0(x)$ into a usable distribution by properly sharpening using $d_0(x)$. Also, see Supp. A2, where a comparison between the traditional Bayes' rule and the density-sharpening-based multiplicative model update rule is presented.
The density-sharpening law provides a mechanism of building a model $f(x)$ for the data $X_1,\ldots, X_n$ by inheriting knowledge from the assumed working model $f_0(x)$. To apply the formula (ref), we need to estimate $d_0(x)$ from data.\footnote{To keep the theory of estimation simple, we will mainly focus on the $X$ continuous case. A detailed account for the discrete case can be found in D21discrete.} And we call this learning process `comparison coding' because $d_0(x)$ codes how surprising the current situation is in light of the model-0 by contrasting expectations with reality.
\vskip.25em Since the density-sharpening function $d_0(x):=d(F_0(x);F_0,F)$ is a function of $F_0(x)$, we can approximate it by a linear combination of polynomials that are function of $F_0(x)$ and orthonormal with respect to the base-model $f_0(x)$. One such orthonormal system is the LP-family of polynomials D20copula, D21discrete, Deep17LPMode, which can be constructed as follows. For an arbitrary continuous $F_0$, define the first-order LP-basis function as standardized $F_0(x)$: \beq {1.2em} {1.2em} T_1(x;F_0)\,=\,\sqrt{12} \big\{F_0(x) - 1/2\big\}. \eeq Note that $\Ex_0(T_1(X;F_0))=0$ and $\Var_0(T_1(X;F_0))=1$. Next, apply Gram-Schmidt procedure on powers of the first-order LP-basis functions $\{T_1^2(x;F_0), T_1^3(x;F_0),\ldots\}$ to construct a higher-order LP orthogonal system $T_j(x;F_0)$: \bea T_2(x;F_0) &=& \sqrt{5} \big\{ 6 F^2_0(x) - 6 F_0(x) + 1 \big\} \\ T_3(x;F_0) &=& \sqrt{7} \big\{ 20 F^3_0(x) - 30F^2_0(x) + 12F_0(x) -1 \big\} \\ T_4(x;F_0) &=& \sqrt{9} \big\{ 70F^4_0(x) - 140F^3_0(x) + 90F^2_0(x) -20F_0(x) +1 \big\}, \eea and so on. Compute these polynomials by performing the Gram-Schmidt process numerically, which can be done using readily available computer packages like R or python.
Replacing $\LP[j;F_0,F]$ with its plug-in estimator in (ref) we get \beq {1.25em} {1em} \wtd_0(x)\,=\,1+\sum\nolimits_j \tLP[j;F_0,F] \,T_j(x;F_0), \eeq where \beq \tLP[j;F_0,F] \,=\, \Ex_{\wtF}[T_j(X;F_0)] \,=\, \frac{1}{n} \sum_{i=1}^n T_j(x_i;F_0).\eeq
Although (ref) provides a robust nonparametric comparison-coding procedure, it has one drawback: the estimated $\wtd$ may be unsmooth due to the presence of a large number of small noisy LP-coefficients. To avoid unnecessary ripples in $\wtd$, we need to isolate the small number of non-zero LP-coefficients. Our denoising strategy goes as follows D21peirce: sort the empirical $\tLP[j;F_0,F]$ in descending order based on their absolute value and compute the penalized ordered sum of squares. This Ordered PENalization scheme will be referred as OPEN model-selection method: \beq OPEN(m) = Sum of squares of top $m$ LP coefficients - \dfrac{\gamma_n}{n}m. \eeq Throughout, we use AIC penalty with $\gamma_n=2$. Find the $m$ that maximizes the ${\rm OPEN}(m)$. Store the selected indices $j$ in the set $\cJ$. The OPEN-smoothed LP-coefficients will be denoted by $\widehat{\LP}_j$. Finally, return the following smoothed estimate: \beq {1.15em} {1.15em} \whd_0(x)\,=\,1+\sum\nolimits_{j \in \cJ} \hLP[j;F_0,F] \,T_j(x;F_0). \eeq
Understanding the deficiency of the current model is an essential part of the process of iterative model building and refinement: Have we overlooked something? Where are our knowledge gaps? This section provides a comprehensive understanding and exploratory tool for representing and assessing potential model misspecifications. Also, see Supp. note A1, which explains the distinctions between parametric and nonparametric model uncertainties.
A general measure of the degree of model misspecification is defined using the Csisz{\'a}r information divergence class.
One can recover popular divergence measures by appropriately choosing the $\psi$-function:
\vskip.4em
One can quickly estimate the $\chi^2$-model misspecification index by expressing it in terms of LP-Fourier coefficients (applying Parseval's identity to equation (ref)): \beq {1.5em} {1.5em} I_{\chi^2}(F,F_0)= \int d^2 - 1 = \sum_{j=1}^m \big| \LP[j;F_0,F] \big|^2. \eeq $I_{\chi^2}(F,F_0)$ quantifies the uncertainty of the preliminary model $f_0(x)$ in light of the given data---i.e., whether $f_0(x)$ is catastrophically wrong or slightly wrong. Estimate it by plugging the empirical LP-coefficients (ref) into (ref). Since, under $H_0: F=F_0$, the sample LP-coefficients have the following limiting null distribution (see Theorem 2 of Deep17LPMode): \[ \setlength{\abovedisplayskip}{1.5em} \setlength{\belowdisplayskip}{1.5em} \sqrt{n} \tLP[j,F_0,F] \xrightarrow[]{d} \cN(0,1),~~ \textrm{i.i.d for all}~ j,\] $n \widetilde{I}_{\chi^2}(F,F_0)$ follows $\chi^2_m$ under null. One can use this to compute the $p$-value. Applying this measure to example 1, we get a $p$-value of practically zero---indicating that the background exponential model is badly damaged and should be repaired before making a decision.
A few additional points on density-sharpening: \vskip.14em
1. The $\DS(F_0,m)$-based density-sharpening principle provides a mechanism for exploring data by exploiting the uncertain background knowledge model. It starts with data and an approximate model $f_0(x)$---and produces a more refined picture of reality following (ref).
\vskip.24em
2. The process of density-sharpening suitably `stretches' the theory-informed model to create a class of robust empirico-scientific models. Moreover, it shows how new models are born out of data-driven mutation of pre-existing ones.
\vskip.24em
3. The truncation point $m$ indicates the radius of the neighborhood around the elicited $f_0(x)$ to create permissible models. $\DS(F_0,m)$ models with higher $m$ entertain alternative models of higher complexity. However, to maintain conceptual appeal and interpretability, it is advisable to focus on the vicinity of $f_0$ by choosing an $m$ that is not too large. Substituting the smooth estimates $\hLP[j;F_0,F]$ of eq. (ref) into the formula (ref), we get the most economical model (among competing alternatives around $f_0$) that best explains the empirical surprise.\footnote{It brings our theory close to Gilbert Harman's “Inference to the best explanation” idea; see harman1965. This is an area that merits further research.}
\vskip.24em 4. It provides an architecture of an `intelligent agent' that simultaneously possesses the ability to: learn (what's new can we learn from the data), reason (how to explain the surprising empirical findings), and plan (how to self-modify to adapt in the new situations).
It is interesting to compare our $d$-sharp LN-model (the red curve) with the seven-parameter exponential family fit shown in Fig. 5.7 of efron2016computer. The most noticeable difference lies in the right tail. Efron's seven-parameter exponential family model shows weird spikes on the extreme-right tail. The main reason for this is that it is based on polynomials of raw $x$: ($x,x^2,\ldots,x^7$), which are not robust. That is to say, these traditional bases are unbounded and highly sensitive to `large' data points. In contrast, our LP-polynomials are functions of $F_0(x)$, not raw $x$, and thus robust by design. The other operational difference between our approach and Efron's exponential family approach is that we model the “gap” between lognormal and the data, which is often far easier to approximate nonparametrically (only required one parameter, see eq. (ref)) than modeling the data from scratch.\footnote{There is an easy way to see that: compare the shapes of the histograms of the left two plots of Fig. (ref).}
{\bf Modelplasticity}---Models ability to modify and adapt itself in response to new data. The density-sharpening principle enables the model to develop new shapes in the face of change.
{\bf Density-sharpening and model evolution}. Modeling is a continual process, not a one-time data-fitting exercise. The density sharpening mechanism allows us to combine new observations with a priory expected model to generate new insights, as depicted in Fig. (ref).
{\bf Statistical law of model evolution}. Density-sharpening supports this dynamic process of recursive model upgrading: $f_{k}(x) = f_{k-1}(x) \,d_{k-1}(x)$, for $k=1,2,\ldots$, by allowing the model to constantly evolve and reshape itself with fresh sets of data---going from a simple approximate model to a much more mature, accurate model of reality.
\vskip.2em {\bf Abduction and creation of new knowledge}. Abduction is the creative part of an inferential process that aims at producing new theories from data. It builds upon what we know to discover new facts about nature. Abductive learning is concerned with the following questions: What new can we learn from the data? How to change the prior hypothetical model to explain the current situation? Which alternative classes of models are worthy of being entertained?
Charles Sanders Peirce (1837–1914) was the pioneer of abductive reasoning; see stigler1978 and D21peirce for more details on the Peircean view of statistical modeling. The goal of Abductive Inference Machine\footnote{They are not traditional pattern recognition (or matching) engine, they are pattern discovery engine.} or AIM is to provide a learning framework that endows a model with this ability to learn, grow and change with new information. \vskip.25em
Attention is the prerequisite of gaining new knowledge. Intelligent learners have the ability to quickly infer where to focus attention to gain knowledge. In our modeling framework $d_0(x)$ draws analyst's attention quickly and efficiently to the new informative part by suppressing boring details; verify it from the graphs of $d_0(x)$ in Figs. (ref) and (ref). It acts as a `gating mechanism' that filters out the new interesting (surprising) aspects of the data, and ignores the dull and unsurprising part---thereby sharpening the model's intelligence by guiding where to pay attention for information processing.
This section demonstrates how practicing abductive inference based on the density-sharpening principle can enable better decision-making in highly uncertain environments.
Abduction is the process of generating and revising a model before choosing the optimal action. An abducer makes decisions in a dynamic uncertain environment by allowing for potential model misspecification.\footnote{The importance of model uncertainty in economics, finance, and business is beautifully illustrated in hansen2014, although from a different perspective.} Abductive decision-making is about knowing when to change course and how to change it.
\vskip.35em How can a decision-maker abduct? The mechanics of abductive decision-making consist of three steps: (i) generating a set of plausible alternative models based on new evidence; (ii) constructing a `robust' model (by choosing the least favorable alternative model or by averaging the alternative models with proper weights); and (iii) selecting an action that maximizes expected utility under the newly revised model. Two modes of abductive decision-making under uncertainty are presented below.
\vskip.35em
{\bf Notation}. A decision-maker (DM) has to take an action $a$ from the set of available actions $\mathbb{A}=\{a_1,\ldots,a_q\}$ based on observed outcome $X_1,\ldots,X_n$ from an unknown probability distribution $f(x)$, representing some natural or social phenomenon. The DM selects the optimal action that minimizes expected loss (or risk) under the assumed model-0: \beq \ha_0 := \argm_{a \in \Abb} \int L_a(x) \dd F_0(x), \eeq where $f_0(x)$ is the DM's posited probability distribution over outcomes. However, as an abducer, the DM is completely aware that the uncertainty about the outcomes may not be fully captured by a single, rigidly-defined probability distribution $f_0(x)$ and thus wants to choose the best decision by accommodating the uncertainty of model-0.
\vskip.35em {\bf Decision making based on density sharpening principle}. To account for the imperfect nature of model-0, the most natural thing to do is to work with an enlarged class of plausible distributions around the vaguely acceptable $f_0(x)$: \beq \Gamma_M = \big\{ f: \, f \in \DS(F_0,m), m \le M \big\} \eeq within a certain reasonable neighbourhood, say $M=10$. We like to use this enlarged class of distributions $\Gamma_M$ for robust decision-making. Two such strategies are discussed below.
\vskip.35em {\bf Method 1.} A cautious DM selects an action by its expected loss under the least favourable distribution within the set $\Gamma_M$: \beq \breve{f}_{a,M}\, =\, \argsup_{F \in \Gamma_M} \int L_a(x) \dd F(x). \eeq We call this an abductive-minimax procedure. Our proposal is partly inspired by the `local-minimax' idea of hansen2001, hansen2001robust.
\vskip.35em {\bf Method 2.} We now describe another robust decision-making procedure that takes into account the uncertainty in the analyst's elicited probability model of future states. Two key concepts are: bootstrap model averaging and action-profile function. \vskip.25em Step 1. We use bootstrap to explore $f \in \Gamma_M$ in an intelligent way. Draw $n$ samples with replacement from the original data. Denote the bootstrap empirical cdf as $\wtF_*^{(1)}$. Perform density-sharpening algorithm based on $\wtF_*^{(1)}$, and denote the estimated $d$-sharp model as $f_*^{(1)}$.
Step 2. Use $f_*^{(1)}$ to select the best action from the given set of $q$-actions $\{a_1,\ldots, a_q\}$. Denote the selected action as $a_*^{(1)}$.
Step 3. Repeat steps 1-2, $B$ times (say $B=1000$ times). And return:
Step 4. This bootstrap scheme can also be used to approximately compute the least favourable distribution, defined in (ref): \[ \breve{\breve{f}}_{a,M} = \argmax_j \int L_a(x) \dd F_{*}^{(j)}(x).\] The decision-maker can use this estimated model to carry out the proposed abductive-minimax procedure (method 1).
Step 5. {\bf Robust procedure}\footnote{Our philosophy of robustness is in complete agreement with huber77robust, who advocated distributional robustness: “one would like to make sure that methods work well not only at the [idealized parametric] model itself, but also in a neighborhood of it.’’}: A pragmatic\footnote{Pragmatism is the logic of abduction.} decision-maker chooses an action (or ranks the actions) that minimizes expected loss (or maximizes the expected utility) with respect to the averaged-distribution: $\ha_{{\rm robust}} := \argm_{a \in \Abb} \int L_a(x) \dd \bar{F}(x)$. Our strategy prescribes action that is robust across a wide range of plausible alternative models. It could be especially powerful for dealing with “deep uncertainty” in making robust policies. For a comprehensive overview on this subject, see marchau2019decision.
Step 6. Quantifying the `robustness' of the action (or decision rule): How much does the optimal action change when a model is selected from a reasonable neighborhood of the assumed initial opinion, i.e. from the $\Gamma_m$ class? The shape of the action profile distribution can be used to determine how robust the optimal action is to model perturbation. In particular, the entropy of the action profile distribution can be used to assess the robustness (or stability) of the inference to potential model misspecification:
\beq {\rm Entropy}[p_A] \,=\,- \sum_{i=1}^q p_A(i) \,\log p_A(i) = \,- \sum_{i=1}^q \Pr(A=a_i) \,\log \Pr(A=a_i). \eeq Uniform probability over possible actions yields maximum uncertainty---indicating that the decision is highly non-robust (sensitive) to model misspecification.
Until now, we have assumed experts can precisely formulate their opinion in a probabilistic form $f_0(x)$. However, for complex real-world decision-making problems, experts might only have incomplete information about the uncertainty distribution of the target variable. A decision-analyst often elicit their partial knowledge about an uncertain quantity as a set of quantile-probability (QP) pairs $\{x_i, F(x_i)\}$, for $i=1,\ldots,\ell$. The job of an analyst is to find a simple, flexible, and parameterizable density that honors the assessed percentiles. \vskip.3em
\vskip.25em
{\bf Model ambiguity due to incomplete information}. The task of eliciting an expert's probability distribution from a small set of QP pairs is a vital yet nascent topic in decision analysis; see powley2013quantile, keelin2011quantile, hadlock2017quantile. In this section, we present an algorithm called Q2D (stands for quantile to distribution) that provides a systematic approach to deduce a reliable expert distribution from $\ell$ arbitrary QP-specifications.
\vskip.35em {\bf Probability-gap Approximation}. The main theoretical idea behind Q2D algorithm: Recall our $\DS(F_0,m)$ model \beq f(x)\,=\,f_0(x)\Big[ 1\,+\, \sum_{j=1}^m \LP[j;F_0,F]\, T_j(x;F_0)\Big] \eeq Integrating from minus infinity to $x$ on both sides, we have \[ \int_{-\infty}^x ( f(z) - f_0(z) ) \dd z = \sum_{j=1}^m \LP[j;F_0,F] \int_{-\infty}^x S_j(F_0(z)) \dd F_0(z),\] where $S_j(u)=T_j(Q_0(u); F_0)$ is defined over the unit interval $[0,1]$ and $Q_0(u)$ is the quantile function of the distribution $f_0$. This leads to \beq F(x) - F_0(x)\,=\, \sum_{j=1}^m \LP[j;F_0,F] \int_0^{F_0(x)} S_j(u) \dd u. \eeq
Given a set of arbitrary $\ell$ quantile-probability data $(x_i,F(x_i)),$ for $i=1,\ldots,\ell$, we can rewrite (ref) compactly as a matrix equation \beq v = S_0 \bbe \eeq where $v_i=F(x_i) - F_0(x_i)$, $\be_i=\LP_j$, and $S_0 \in \cR^{\ell\times m}$, $S_0[i,j]=\int_0^{F_0(x_i)} S_j(u)$. The desired parameters are ${\bm \be}=(\be_1,\ldots,\be_m)$, where $\be_j$ is shorthand for $\LP[j;F_0,F]$.
For $m \le \ell$, we can uniquely estimate $\bbe$ using the least-square method \beq \widetilde{\bbe}\,=\,\miniz_{\bbe} \| v - S_0 \bbe \|^2 \, =\, (S_0^TS_0)^{-1}S_0^{T}v. \eeq For large $\ell$ (say, $\ell\ge 5$), a better, more stable estimate can be found through regularization \beq \widehat{\bbe}\,=\,\miniz_{\bbe} \| v - S_0 \bbe \|^2 + \la \|\bbe\|_1 \eeq where $\| \cdot\|_p$ is the $\ell_p$ norm, and $\la>0$ is the regularization parameter. The lasso lasso1996 penalized $\widehat{\bbe}$ yields a sparse estimate and counters over-fitting. This penalized estimate provides a tradeoff between accuracy and interpretability.\footnote{Note that due to regularization, Eq. (ref) can even tackle cases with $m > \ell$. In such scenarios, the OLS is ill-posed, with an infinite number of solutions.} Finally, plug the estimated LP-Fourier coefficients $\be_j$ into the primary equation (ref) to get the expert distribution.
High-stakes decision-making (say, COVID-19 pandemic or climate change) is often based on multiple experts' opinions instead of putting all bets on a single rigidly-defined probability model. The challenge is to aid data-driven decision-making by appropriately combining several experts' models. We describe one possible way to build a `consensus committee model' that can be used as a possible model-0 within an abductive decision-making framework.
\vskip.4em
{\bf Learning from multiple expert distributions}. Given $k$ competing probability models $\{f_{01},\ldots, f_{0k}\}$, which may differ markedly in shape, define the following model-weights: \beq Relevance weight: w_\ell = \dfrac{1}{1 + \sum_j |\LP_{j|\ell}|^2}, {\rm for} \ell=1,\ldots,k \eeq where $\LP_{j|\ell}$ is the LP-Fourier coefficients of the $\ell$-th model: \beq d_{0\ell}(x) := d(F_{0\ell}(x); F_{0\ell},\wtF)\,=\, 1+\sum_j \LP_{j|\ell} T_j(x;F_{0\ell}). \eeq Note that the relevance weight for the $\ell$-th model is always $0 < w_\ell \le 1$, and \[ w_\ell=1 ~~~\text{if and only if}~~~\LP_{j|\ell}=0,~\forall~ j.~~ \] $\LP_{j|\ell}=0$ for all $j$ when $f_{0\ell}$ fully explain the data and there is no need to sharpen it further (i.e., $d_{0\ell}=1$). In that sense, $w_\ell$'s are data-driven weights (which will keep changing as we get more and more fresh data), computed based on the degree of agreement between the observed data and expert model $f_{0\ell}$. Define mixture expert distribution as
\beq f^0_{\rm mix} (x) = \sum_{\ell=1}^k \pi_\ell\, f_{0\ell} (x), \eeq where $\pi_\ell = w_\ell/\sum_\ell w_\ell$. This model serves two purposes: it tries to resolve conflicting opinions based on data and at the same time encourages one to include as much diverse information as possible.
Additional layer of model uncertainty. If the analyst believes that the correct model might not be among the collection of models being considered, then use the combined expert model $f^0_{\rm mix} (x)$ as a model-0 in the subsequent density-sharpening-based learning and abductive decision-making process.
How should an analyst use imperfect models to learn from data?\footnote{The challenge of learning from uncertain knowledge is also a fundamental issue in the development of intelligent systems.} What should be the output of such an analysis that can ultimately aid informed decision-making? We address these questions by introducing a general inferential framework for statistical learning and decision-making under uncertainty---which builds on two core ideas: abductive thinking and density-sharpening principle. Some of the defining features of our approach for data analysis, scientific discovery, and decision-making are highlighted below:
$\bullet$ Data analysis and science of model management: No model is perfect, irrespective of how cunningly it is designed. The central problem of statistical model developmental process is to understand how a relatively simple model can evolve into a more complex and mature one in the presence of a new data environment. The principle of density-sharpening assists this model evolution process (thereby helping empirical scientists to abduct): by abductively generating explanations on why the presumed model-0 is unfit for the data [playing the role of a quality inspector] and also providing recommendations on how to fix the misspecification issues [serving as a policy adviser] in order to make better decisions in new circumstances.
\vskip.77em $\bullet$ Discovery and creation of new knowledge: Abductive data analysts are less interested in testing a particular working model. They are mainly interested in conceptual innovation: discovering new pursuitworthy hypotheses based on surprising empirical evidence.\footnote{A largely unexplored topic relative to the vast literature on hypothesis testing. As noted by George E. P. box2001dis: “Much of what we have been doing is adequate for testing but not adequate for discovery.”} The density-sharpening function $d(u;F_0,F)$ picks out `what's new' in the data beyond the current scientific knowledge encoded in $f_0(x)$, thereby helping the scientist to uncover new unexpected knowledge from the data using graphical tools. The density-sharpening principle (DSP) provides a learning mechanism that isolates the `known' from the `unknown' and allows us to focus on the newfound pattern in the data, which is the basis for knowledge-creation\footnote{Curious readers are invited to read the paper “Nobel Turing Challenge: creating the engine for scientific discovery” by Hiroaki kitano2021nobel, where he argued that the single-most-important mission of AI is to accelerate scientific discovery.}.
\vskip.77em $\bullet$ Abductive inference and decision-making: The proposed theory of abductive decision-making tackles model uncertainty induced by imprecise, ambiguous, and incomplete knowledge about the underlying probabilistic structure. An abductive-decision support system automatically discovers and explicitly articulates the possible alternatives to the analysts, which forces them to rethink their choices before taking impulsive action; see Supp. note A5. This style of empirical reasoning and adaptive decision-making could be especially valuable in situations where strategic planners need to take quick action in the face of uncertainty, equipped with approximate subject-matter knowledge.
All the datasets and R-code written for the analysis are available upon request to the author.
The Author declares that there is no conflict of interest/competing interests.
For more discussion on how our abductive statistical approach compares to the more traditional Bayesian statistical approach for model misspecification, robustness, and decision-making, see the Supplementary section.
\setcounter{page}{1} \setcounter{equation}{0}