Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
112,738 characters · 35 sections · 49 citation commands
InClass Nets: Independent Classifier Networks for Nonparametric Estimation of Conditional Independence Mixture Models and Unsupervised Classification
In many fields of science one encounters multivariate models which consist of several distinct sub-populations or components, say $C$ in number. Each component $i\in\{1,\dots,C\}$ has its own characteristic probability density function $f^{(i)}(\mathcal{X})$ of the relevant multi-dimensional variable ${\mathcal{X}}$. Such models are referred to as multivariate finite mixture models McLaghlan2019, and the probability density of ${\mathcal{X}}$ under such a model is given by
where the non-negative weights $w_i$ parameterize the mixing proportions of the individual components. An important special case of these finite mixture models is that of the conditional independence multivariate finite mixture models\footnote{In the literature, these models are also referred to as finite mixtures of product measures teicher1967.} chauveau2015,Zhu2016, which for brevity we will simply refer to as conditional independence mixture models (CIMMs). Under this special case, the variable ${\mathcal{X}}$ is parameterized using $V$ “variates” as ${\mathcal{X}} \equiv (x_1,\dots,x_V)$ such that for each component $i$, the density function $f^{(i)}$ factorizes into a product of distributions for the individual variates as
so that (ref) becomes
Here $f^{(i)}_v(x_v)$ is the unit-normalized probability density of $x_v$ within component $i$---the top index $(i)$ denotes the component and the bottom index $v$ denotes the variate. In our treatment, the individual variates $x_v$ are themselves allowed to be multi-dimensional. In other words, $V$ is not the dimensionality of the data, but rather the number of groups the attributes in ${\mathcal{X}}$ can be partitioned into so that they are (conditionally) independent of each other within the given component $i$ a datapoint belongs to. This is similar to the treatment, for example, in Hall2005,chauveau2015,CHAUVEAU20161. In particular, the technique we develop below will be applicable in situations where the variates $x_v$ are high-dimensional (dim($x_v$) $\gg 1$). Unless otherwise stated, henceforth a “mixture model” shall refer to the conditional independence mixture model of (ref).
Conditional independence mixture models have applications in situations where the correlations and dependence between different variables in the data are explained in terms of a latent or hidden confounding variable which influences the observed variables. This is referred to as Latent Structure Analysis (LSA) lazarsfeld_henry_1968,Andersen1982. In particular, when the confounding variable is discrete or categorical, it can be interpreted as representing the class a given datapoint belongs to. Such models are referred to as latent class models and their study and analysis is referred to as Latent Class Analysis (LCA) Clogg1995.
The connection between CIMMs and LCMs can be seen in a straightforward manner as follows. We can sample a datapoint as per the mixture model in (ref) by first generating the component index $i\in\{1,\dots,C\}$ as per the multinomial probability distribution induced by the weights $w_i$, and then sampling $(x_1,\dots,x_V)$ as per the distribution $f^{(i)}({\mathcal{X}})$ within component $i$. Now, the component index $i$ can be interpreted as the latent variable that explains the dependence between the random variates $\{x_1,\dots,x_V\}$ in the mixture.
CIMMs, LCA, and mixture models in general have applications in a wide range of fields, including econometrics Compiani2016,SpecialIssue2017, social sciences Vermunt2003,VermuntMagidson2004,Porcu2017,Petersen2019, bioinformatics Yu2010, astronomy and astrophysics Nemec1991,Bovy:2011xm,Lee2012,Melchior:2016asy,Kuhn2017,Necib:2018iwb,Jones:2019njd, high energy physics Stepanek:2015jqa,Cranmer:2015bka,Cranmer:2016swd, and many others.
Estimation of a CIMM is simply the process of estimating the weights $w_i$ and functions $f^{(i)}_v$ (assuming the number of components $C$ is known) from a dataset sampled from the joint distribution ${\mathcal{P}}({\mathcal{X}})$ of the variates under the mixture model, see (ref).
Under parametric estimation of mixture models (conditionally independent or otherwise), one assumes that each of the distributions $f^{(i)}({\mathcal{X}})$ is from an appropriately chosen parametric class of distributions, e.g. multivariate Gaussians. The choice of the class of functions assumed to contain the true $f^{(i)}({\mathcal{X}})$ is informed by practitioner's prior knowledge of the problem at hand. This reduces the problem of estimating the mixture to the more tractable problem of estimating the weights $w_i$ and the parameter values corresponding to the true $f^{(i)}({\mathcal{X}})$. This weight and parameter estimation from the data is typically approached as a maximum likelihood estimation problem doi:10.1111/insr.12315, often tackled using the expectation--maximization algorithm 10.2307/2984875.
Semi-parametric estimation of mixture models has been studied in many works, including doi:10.1080/00949650214669,doi:10.1198/jcgs.2009.07175,chauveau2015,Xiang2019. In this paper, however, we are interested in the nonparametric estimation of conditional independence mixture models, i.e., no parametric forms will be assumed for the functions $f^{(i)}_v(x_v)$. Nonparametric estimation has been addressed by several works recently Hall2003,Hall2005,doi:10.1198/jcgs.2009.07175,JSSv032i06,10.5555/3023638.3023695,Levine2011,Zhu2016,Kasahara2014,CHAUVEAU20161,Zheng2019. In this paper, we introduce a novel machine-learning-based-approach, called the InClass nets technique, to split the dataset into its different components in a nonparametric way. This splitting naturally leads to the estimation of the mixture model. In order to perform the splitting, the InClass nets technique directly exploits the fact that the variates $x_v$ are mutually independent of each other within each component. The biggest advantage offered by our machine-learning-based-technique over existing approaches is the possibility of tackling situations where the individual variates $x_v$ are high-dimensional. This opens up the possibility of using CIMMs for hitherto unfeasible applications.
The estimation of mixture models is closely tied to the concept of identifiability of mixture models. A statistical model is said to be identifiable if it is theoretically possible to estimate the model (i.e., uniquely identify the parameters and functions that describe it) based on an infinite dataset sampled from it. A model will not be identifiable if two or more parameterizations of the model are observationally indistinguishable under those circumstances.
In the situations where the CIMM is identifiable, our technique estimates the true $w_i$ and $f^{(i)}_v$. On the other hand, when the model is not identifiable, our technique will yield one of the parameterizations that best fits the available data. In \hyperref[sec:identifiability]{Section \ref*{sec:identifiability}}, we provide some new results on the (nonparametric) identifiability of bivariate ($V=2$) conditional independence mixture models, to supplement existing results on nonparametric mixture model identifiability teicher1967,yakowitz1968,Gyllenberg1994,Hall2003,Hall2005,Elmore2005,Allman2009,Kasahara2009,kovtun_akushevich_yashin_2014,Tahmasebi2018.
In this paper, we approach the estimation of CIMMs as a classification problem---classifying the datapoints in a given dataset into the different categories will lead to the estimation of the mixture model in a straightforward way.
This classification needs to be performed in an unsupervised manner since the dataset being analyzed does not contain labels for the component each datapoint belongs to. In this way, our method has connections to unsupervised clustering techniques like k-means clustering 1056489 and other density-based clustering techniques 10.5555/3001460.3001507. However, our approach does not rely on the different components being spatially clustered to perform the classification.
The intuition behind our method can be understood as follows: In supervised classification, the target class labels associated with the training datapoints serve as the supervisory signal for training the classifier. In the absence of target labels, a quantity that is dependent on or shares mutual information with, the (unavailable) target label can be used as the supervisory signal. Now, lets say we are training a classifier that bases its decision or output only on the first variate $x_1$. The other variates $\{x_2,\dots,x_V\}$ can serve as the supervisory signal, since they contain information regarding the component $i$ the datapoint belongs to. Our approach can be thought of as training $V$ classifiers, one for each of the $V$ variates, with each classifier relying on the other $V-1$ variates to act as the supervisory signal for training.
There are a few ways of interpreting and actualizing this intuition 10.1145/956750.956764,10.5555/647235.720080,ji2019invariant. For example, in ji2019invariant, neural networks were trained without supervision to classify images, using a training dataset consisting of pairs of images, where the images in a given pair are from the same category. In the InClass nets approach developed in this paper, the neural network architecture and training cost functions we develop will primarily be geared towards estimating conditional independence mixture models, which, as shown below in \hyperref[sec:MNIST]{Section \ref*{sec:MNIST}}, can nevertheless be used for classification similar to the technique of ji2019invariant. In \hyperref[sec:multilabel]{Section \ref*{sec:multilabel}}, we also discuss a straightforward extension of the technique in ji2019invariant to handle n-tuples of data, where the components of the n-tuple could be from different sample spaces (as opposed to pairs of images from the same sample space of images).
The idea of using a quantity that shares information with the true labels as the supervisory signal has been employed previously in weakly supervised classification techniques like “Learning from Label Proportions” (LLP) 10.1145/1390156.1390254,NIPS2014_5453,Yu2014OnLW and “Classification Without Labels” (CWoLa) Metodiev:2017vrx. LLP and CWola learn to distinguish between different classes of datapoints, using multiple mixed datasets which differ in the mixing proportions of the classes---the identity of the mixed dataset a given datapoint belongs to serves as the supervisory signal, since it contains information about the class the datapoint belongs to (due to the mixing proportions in different mixtures being different). While CWola and LLP are not fully unsupervised techniques (since they still require a label indicating which mixture a training datapoint belongs to), they are applicable even in situations where the distribution of the feature ${\mathcal{X}}$ within a given class $i$ does not factorize as $\displaystyle\prod_{v=1}^V\,f^{(i)}_v$.
The idea of separating a mixture into its components using a mutual-information-based technique is similar in spirit to Independent Component Analysis (ICA) hyvarinen2001independent. However, ICA solves a signal separation problem where multiple mixtures with different mixing weights for the components are provided---this is different from the problem of separating data from a single CIMM into its components.
Mixture models have applications in data analysis in high energy physics Stepanek:2015jqa,Cranmer:2015bka,Cranmer:2016swd, where datasets are mixtures of “events” (datapoints) produced under different “processes” (categories). The ${}_s\mathcal{P}lot$ technique Pivk:2004ty, which is popular in data analysis in high energy physics, is used to analyze bivariate conditional independence mixture models where the distribution of one of the variables (referred to as the discriminating variable) is known a priori. In such situations, the ${}_s\mathcal{P}lot$ technique can estimate the distribution of the other variable, referred to as the control variable. On the other hand, the InClass nets approach introduced in this paper is capable of estimating the mixture model without any knowledge of the distributions of any of the variables. In \hyperref[subsec:prior_knowledge]{Section \ref*{subsec:prior_knowledge}}, we describe how the InClass nets approach can be modified to incorporate additional information about the distributions of some of the variates.
For the purpose of nonparametric estimation of conditional independence mixture models, we introduce a new neural network architecture which we shall call “Independent Classifier networks” or “InClass nets” for short. Under InClass nets, the $V$ variates $\{x_1,\dots,x_V\}$ of the input ${\mathcal{X}}$ are fed into $V$ independent neural networks---one variate for each independent network. Each of the $V$ networks returns a multi-class classifier output. More explicitly, for each $v\in \{1,\dots,V\}$, the $v$-th classifier network returns a vector $\left(\eta^{(1)}_v(x_v),\dots,\eta^{(C)}_v(x_v)\right)$, whose $i$-th component can roughly be interpreted as the probability that a datapoint belongs to category $i$, conditional only on its $x_v$ value.
The outputs of the independent classifiers are constrained to obey
possibly using the softmax output layer\footnote{For the case of $C=2$, one can also simply use a one-dimensional output layer constrained to be in $[0,1]$, with $(\texttt{output}, 1-\texttt{output})$ serving as $(\eta^{(1)}, \eta^{(2)})$.} Goodfellow-et-al-2016 as follows:
where the $z^{(i)}$-s are the inputs\footnote{This is assuming that the output layer only performs the softmax operation. If softmax is used as an activation function of the final layer, then the $z^{(i)}$-s represent the outputs of the layer before applying the activation function.} to the final (output) layer of the corresponding network. The $\texttt{softmax}$ function is defined as
\hyperref[fig:inclass]{Figure \ref*{fig:inclass}} illustrates the basic architecture of InClass nets. In the next few sections, we will build the framework for estimating conditional independence mixture models using InClass nets.
Recall that the variates $x_v$ can be multi-dimensional. In particular, InClass nets can handle high-dimensional data types like images. The choice of architecture for the individual classifier networks can be influenced by the nature of the input data the classifier will handle.
For the purposes of this paper, we have restricted the output dimensionality of the individual classifiers to be the same (equal to $C$). We have also restricted the inputs $\{x_1,\dots,x_V\}$ of the individual classifiers to form a non-overlapping partition of the features or attributes in ${\mathcal{X}}$, which is in line with the structure of conditional independence mixture models. However, InClass nets can have applications outside mixture model estimation as well, and for those purposes it may be appropriate to lift the above restrictions. For example, InClass nets can be used to perform unsupervised multi-label classification, where the outputs of different classifiers correspond to different labels. In this case, the different classifiers can have different output dimensionalities and the inputs to these networks can also potentially have overlapping features. In \hyperref[sec:multilabel]{Section \ref*{sec:multilabel}} we will briefly indicate how the multi-label variant of InClass nets can be trained to perform unsupervised classification by maximizing the mutual information between the classifier outputs.
In this section we will show how conditional independence mixture models can be parametrized using InClass nets. The parametrization will be done using the Independent Pseudo Classifiers representation of mixture models which will be introduced in \hyperref[subsec:ipcr]{Section \ref*{subsec:ipcr}}. But as a useful lead-up, let us first introduce the Constrained Independent Classifiers representation.
The mixture model in (ref) is completely specified by the mixture weights $w_i$ and the distributions $f^{(i)}_v$. Recall that they satisfy
The goal of this paper is to develop a machine learning technique to fit a mixture model to the given data {\bf in an agnostic, nonparametric, manner}. In other words, we will estimate the weights $w_i$ and the distributions $f^{(i)}_v$, without assuming, a priori, any parameterized forms (like Gaussians, exponentials, etc.) for $f^{(i)}_v$. We will approach this as an unsupervised multi-class classification problem---classifying the data into different components will automatically result in an estimation of the mixture model\footnote{Directly modeling the distributions $f^{(i)}_v$ is possible using generative networks, but classifier outputs are more robust quantities, e.g., they are invariant under invertible transformations of the $x_v$-s, and are typically easier to learn in machine learning.}. To this end, let us rewrite the mixture model distribution in terms of the marginal distributions ${\mathcal{P}}_{\!v}(x_v)$ of the individual variates and multi-class classifiers $\alpha^{(i)}_v(x_v)$ given by
${\mathcal{P}}_{\!v}$ is the probability density of the $v$-th variate in the full mixture and can be directly accessed from a dataset sampled from ${\mathcal{P}}$. $\alpha^{(i)}_v(x_v)$ can be interpreted as the probability that an observed datapoint is from component $i$ conditional on the value of $x_v$. The vector function $\left(\alpha^{(1)}_v(x_v),\dots,\alpha^{(C)}_v(x_v)\right)$ can be interpreted as the output of a multi-class “classifier” that returns the probability of a datapoint ${\mathcal{X}}$ to have come from the different components based only on the $v$-th variate. At this point, one might already notice an emerging connection with InClass nets, which we shall crucially exploit below. The marginals density functions ${\mathcal{P}}_{\!v}$ and the multi-class classifiers $\alpha^{(i)}_v$ satisfy
where the integrals in (ref) are simply equal to the weight $w_i$ of the $i$-th component. There is a one-to-one map\footnote{This is not a statement on the identifiability of conditional independence mixture models. Identifiability of mixture models will be briefly discussed in \hyperref[sec:identifiability]{Section \ref*{sec:identifiability}}.} from the description of the mixture model in terms of the $w_i$-s and $f^{(i)}_v$-s satisfying (ref) to the description in terms of ${\mathcal{P}}_{\!v}$-s and $\alpha^{(i)}_v$-s satisfying (ref). This can be seen from the existence of the inverse transform shown below:
where $E_{{\mathcal{P}}}[\,\cdots]$ represents the expectation value of $\,\cdots$ under the model. The probability density of ${\mathcal{X}}$ under the corresponding mixture model is given by
As mentioned earlier, the marginal distributions ${\mathcal{P}}_{\!v}$ can be directly estimated from the data. The functions $\alpha^{(i)}_v$ can potentially be modeled using InClass nets. The only hurdle is that while the outputs $\alpha^{(i)}_v$ of the $V$ neural networks can be constrained to obey (ref) and (ref) using the softmax output layer (as seen in (ref)), the constraint in (ref) in general will {\bf not} be satisfied by independent classifiers. We will handle this difficulty next in \hyperref[subsec:ipcr]{Section \ref*{subsec:ipcr}}. We will refer to the description in terms of ${\mathcal{P}}_{\!v}$-s and $\alpha^{(i)}_v$-s satisfying ((ref)--(ref)) as the Constrained Independent Classifiers (CIC) representation of the mixture model.
To accommodate the fact that independent classifiers will not obey the constraint (ref) of the Constrained Independent Classifiers representation, we introduce the Independent Pseudo Classifiers (IPC) representation in terms of pseudo marginals ${\mathcal{Q}}_{v}$ and pseudo classifiers $\beta^{(i)}_v$ which only satisfy the equivalents of constraints ((ref)--(ref)):
The mixture weights under the IPC representation are given by
where the unnormalized weights $\tilde{w}_i$-s are given by
where $\varphi^{(i)}_v \equiv E_{{\mathcal{Q}}}\left[\beta^{(i)}_v\right]$ represents the expectation value of $\beta^{(i)}_v$ under the distribution ${\mathcal{Q}}_{v}$. In (ref) we have used the geometric mean\footnote{It is also possible to use other mean functions, including generalized means, instead of the geometric mean (with appropriate modifications to other relevant quantities and expressions).} $\tilde{w_i}$ of the mixture weights $\varphi^{(i)}_v$ “proposed” by the individual pseudo classifiers for component $i$ as its actual mixture weight $w_i$ in the mixture model, after an appropriate scaling to make the weights add up to 1 across all components. We will refer to $\varphi^{(i)}_v$ as the pseudo weight of component $i$ corresponding to variate $v$. In analogy to (ref), we write the probability density function for the mixture model in the IPC representation as
where the $\tilde{w}_i$-s can be written in terms of ${\mathcal{Q}}_{v}$-s and $\beta^{(i)}_v$-s using (ref). We will now make the following observations relevant to our goal of fitting a mixture model to data using InClass nets:
Observation (4) means that in order to fit a mixture model to data, we can restrict ourselves to IPC representations of the mixture models with the pseudo marginals set to the marginals of the data. The only remaining unknowns in the IPC representation are the pseudo classifiers $\beta^{(i)}_v(x_v)$ which we can parameterize using an InClass net, identifying $\beta^{(i)}_v$ with the network output $\eta^{(i)}_v$. In the next section we will develop the technique to fit a mixture model parameterized with an InClass net to a given dataset.
In this section we will construct a cost function which can be used to train InClass nets to fit mixture models to the given data. Let the data to which we want to fit a mixture model be sampled from the true underlying distribution ${\mathcal{P}}^\ast({\mathcal{X}})$ with true marginals ${\mathcal{P}}^\ast_{\!v}(x_v)$. As per observation (4) in the previous section, we restrict our attention to IPC representations with ${\mathcal{Q}}_{v}\equiv {\mathcal{P}}^\ast_{\!v}$. Using (ref) and (ref), we can write the probability density of ${\mathcal{X}}$ under this restricted class of mixture models as
where $E_{{\mathcal{P}}^\ast}$ refers to the expectation value under the true distribution of the data. The best-fitting ${\mathcal{P}}$ can be estimated by minimizing the Kullback--Leibler (KL) divergence from ${\mathcal{P}}$ to ${\mathcal{P}}^\ast$ given by
Note that minimizing the KL divergence (over some class of distributions) is equivalent to, and commonly known in some disciplines as, maximizing the likelihood in the large statistics limit. Using the expression for ${\mathcal{P}}$ from (ref), we can rewrite (ref) as
where $C^\ast(x_1,\dots,x_V)$ is the total correlation 5392532,garner_1962 of the $V$ variates in the data which is one of the generalizations of mutual information to more than two variables. It is given by the KL divergence from the product distribution $\prod\limits_v {\mathcal{P}}^\ast_{\!v}(x_v)$ to the joint distribution ${\mathcal{P}}^\ast({\mathcal{X}})$ as
Note that the $C^\ast$ term in (ref) is independent of the state of the InClass net under consideration. This means that the second term in (ref) can be used as a cost function for the network to minimize in order to minimize the KL divergence, and hence fit the mixture model parameterized by the InClass net to the data. Noting the similarity between the two terms in (ref) and drawing inspiration from the naming of “cross entropy”, we introduce the “negative cross total correlation” cost function (neg_ctc_cost) defined as
Note that $\beta^{(i)}_v$ are functions of the corresponding input variate $x_v$. Despite the complicated appearance, this cost function provides a viable approach to learning the underlying mixture model from data. Let us make the following observations in the context of training InClass nets using this cost function, with outputs $\eta^{(i)}_v$ of the network identified with $\beta^{(i)}_v$.
After training the InClass net, ((ref)--(ref)) and (ref) can be used to extract the fitted model, with the pseudo marginals ${\mathcal{Q}}_v$ set to the true marginals ${\mathcal{P}}^\ast_v$. The classifiers $\alpha^{(i)}_v(x_v)$ can be extracted from the pseudo classifiers $\beta^{(i)}_v(x_v)$ using (ref). If one is interested in classifying the individual datapoints based on the full information ${\mathcal{X}}$, an aggregate classifier can be constructed, based on (ref), as
Note that if there is a mismatch between the model learned by the InClass net and the true distribution the data is sampled from, then classifying the data using the aggregate classifier will not necessarily lead to components within which the $x_v$-s are independent.
When analyzing real data with conditional independence mixture models, a common difficulty is the identification of a suitable partitioning of the attributes of ${\mathcal{X}}$ into variates $x_v$ so that the distribution within each component would factorize to a good approximation. In this sense, a higher number of (conditionally independent) variates represents stronger assumptions about the underlying model. This makes the bivariate case ($V=2$) extremely important. The bivariate case is also difficult from an identifiability point of view---data distributed according to a conditional independence bivariate mixture model, in general, will not uniquely identify the model, since several different mixture models can lead to the same overall probability density ${\mathcal{P}}({\mathcal{X}})$. In \hyperref[sec:identifiability]{Section \ref*{sec:identifiability}}, we will present some new results on the identifiability of conditional independence bivariate mixture models. In particular, we will provide the conditions under which bivariate mixture models are identifiable.
Despite being the most difficult case in terms of identifiability, the bivariate case lets us gain some useful intuition, as demonstrated with several examples in \hyperref[sec:raindancesvi]{Section \ref*{sec:raindancesvi}} below. But first, in preparation for \hyperref[sec:raindancesvi]{Section \ref*{sec:raindancesvi}}, let us summarize the results from the previous sections for the bivariate case, in the order in which a typical analysis might use them.
The expressions from the previous sections become easier to follow if we explicitly write out the two variates, thus avoiding the product notation. To this end, let us simplify the notation by giving names $x$ and $y$ to our two variates $x_1$ and $x_2$, resulting in
Under this notation, the conditional independence mixture model of (ref) becomes simply
Noting that total correlation is a generalization of mutual information for more than two variables, we will refer to the negative cross total correlation cost function of (ref) in the bivariate special case as the “negative cross mutual information” cost function (neg_cmi_cost). Under our new notation, it is given by
where, as before, $E_{{\mathcal{P}}^\ast}$ represents the expectation over the true distribution ${\mathcal{P}}^\ast(x, y)$ from which the data is sampled.
After training the InClass net, the trained $\beta^{(i)}_x$ and $\beta^{(i)}_y$ cannot directly be interpreted as classifiers based on $x$ and $y$ since they may correspond to different mixture weights. In order to extract the learned mixture model (and the corresponding classifiers), we can first estimate the pseudo weights $\varphi^{(i)}_x$ and $\varphi^{(i)}_y$ from the data as
Now, using (ref) and (ref), the marginals and classifiers for the model represented by the InClass net can be constructed as
From (ref) and (ref), the component weights of the learned model are given by
and from (ref), the distributions $f^{(i)}_x$ and $f^{(i)}_y$ within each component are given by
The corresponding joint distribution is given by
Note that after estimating ${\mathcal{P}}^\ast_{\!x}$, ${\mathcal{P}}^\ast_{\!y}$, $\varphi^{(i)}_x$, and $\varphi^{(i)}_y$ from the dataset, the mixture model can be read off directly from the InClass net using (ref) and (ref).
From (ref), the aggregate classifier that classifies the individual datapoints based on the full information $(x,y)$ is given by
We provide a public, tensorflow-based abadi2016tensorflow, implementation of InClass nets as a Python 3 package called \href{https://prasanthcakewalk.gitlab.io/raindancesvi}{RainDancesVI} rd6. The package provides routines for wrapping the classifier networks of individual variates into InClass nets and cost functions for training them. It also provides utilities for extracting the model learned by the network post-training. In this section we will demonstrate the working of InClass nets using several toy examples analyzed using RainDancesVI.
In the first example, we consider the mixture of two independent bivariate Gaussians. In the first component, $x$ and $y$ are both (independently) normally distributed with mean $-1$ and standard deviation $1.5$. The second component is identical, except $x$ and $y$ both have mean $+1$. The mixture weights are taken to be $w_1 = 0.4, w_2 = 0.6$. \hyperref[tab:bivariate_gaussian]{Table \ref*{tab:bivariate_gaussian}} summarizes the mixture model specification and \hyperref[fig:bivariate_gaussian_comps]{Figure \ref*{fig:bivariate_gaussian_comps}} shows the normalized joint distributions of $(x,y)$ under each of the two components as heatmaps. \hyperref[fig:bivariate_gaussian_combined]{Figure \ref*{fig:bivariate_gaussian_combined}} shows the normalized joint distribution of $(x,y)$ under the mixture model and our InClass net will estimate the mixture model based on data generated as per this distribution.
The classifier networks $\beta^{(i)}_x$ and $\beta^{(i)}_y$ were constructed using keras with the tensorflow backend. The neural networks are distinct, but have identical architectures. The networks are fairly simple, consisting of three sequential dense layers of 32 nodes using the rectified linear unit (ReLU) 10.5555/3104322.3104425 activation function. The output layer is a dense layer with 2 nodes (since $C=2$), with the softmax activation function. The individual classifier networks were then wrapped into an InClass net using the RainDancesVI package. The resulting network has a total of 4,484 trainable parameters.
We trained the InClass net to minimize the negative cross mutual information cost function, using 100,000 datapoints sampled from the distribution depicted in \hyperref[fig:bivariate_gaussian_combined]{Figure \ref*{fig:bivariate_gaussian_combined}}. The optimization was performed for 15 epochs with the Adam adam optimizer (with default hyperparameters) using a batch size of 50. After training the network, we used the same dataset to estimate the pseudo weights $\varphi^{(i)}_{x}$ and $\varphi^{(i)}_{y}$ and the mixture weights $w_i$ using (ref) and (ref). Note that the estimation of mixture models can only be performed up to permutations of the components indexed by $i$. For clarity of the presentation, unless otherwise stated, the components of the true mixture model will be matched with the respective closest candidates from the machine-learned components. The results of the estimation of the mixture weights are summarized in \hyperref[tab:bivariate_gaussian_weights]{Table \ref*{tab:bivariate_gaussian_weights}}, which demonstrates an excellent agreement between the true and estimated values.
The solid red curves in \hyperref[fig:bivariate_gaussian_classifiers]{Figure \ref*{fig:bivariate_gaussian_classifiers}} depict the classifiers $\alpha^{(i)}_{x}(x)$ (left panel) and $\alpha^{(i)}_{y}(y)$ (right panel) learned by the network---they are extracted from $\beta^{(i)}_{x}$ and $\beta^{(i)}_{y}$ with the help of (ref). For comparison, the true classifiers based on the exact functional forms of the component distributions are also shown as green dash-dot curves. The red solid lines and the green dash-dot lines almost coincide, which validates our method.
Next, we used (ref) to estimate the distributions $f^{(i)}_x$ and $f^{(i)}_y$. The resulting distributions are shown with red solid lines in the left and right panels of \hyperref[fig:bivariate_gaussian_dists]{Figure \ref*{fig:bivariate_gaussian_dists}}, respectively. In applying (ref), for simplicity we used the exact expressions for the marginal distributions of $x$ and $y$ in the mixture. In a typical example, the exact expressions for the marginals will not be available, but can be easily estimated from the data, say using a histogram or kernel density estimation rosenblatt1956,parzen1962. In \hyperref[fig:bivariate_gaussian_dists]{Figure \ref*{fig:bivariate_gaussian_dists}}, we also show the true $f^{(i)}_x$ and $f^{(i)}_y$ as green dash-dot curves. The good agreement between the true $w_i$, $f^{(i)}_{x}$, and $f^{(i)}_{y}$ and their estimates shown in \hyperref[tab:bivariate_gaussian_weights]{Table \ref*{tab:bivariate_gaussian_weights}} and \hyperref[fig:bivariate_gaussian_dists]{Figure \ref*{fig:bivariate_gaussian_dists}}, demonstrates that the InClass net has successfully estimated the mixture model. Finally, we use (ref) to estimate the aggregate classifier $\alpha^{(i)}_\text{aggregate}(x,y)$ which is shown as a heatmap in the left panel of \hyperref[fig:bivariate_gaussian_aggregate]{Figure \ref*{fig:bivariate_gaussian_aggregate}}. For comparison, in the right panel we show the true aggregate classifier based on the exact functional forms of the component distributions $f^{(i)}(x,y)$. As expected, the two heatmaps are in very good agreement.
Now we will look at a toy example which was instrumental in the conception and development of the InClass nets technique, see Figures (ref) and (ref). \hyperref[fig:checkerboard_combined]{Figure \ref*{fig:checkerboard_combined}} shows the joint distribution of $(x,y)$ for a “checkerboard” mixture under which the datapoints are uniformly distributed on the bright squares of a $4\times4$ checkerboard spanning the region $0\leq x, y < 4$, while the dark squares have zero density. For concreteness, the vertical (horizontal) boundaries between cells are assigned to the cell on the right (top). It is easy to see that $x$ and $y$ are, individually, uniformly distributed between $0$ and $4$. It can also be seen that $x$ and $y$, despite being uncorrelated, are not mutually independent in the mixture, since x lies within $[0, 1) \cup [2, 3)$ if and only if $y$ does as well.
As shown in \hyperref[fig:checkerboard_comps]{Figure \ref*{fig:checkerboard_comps}}, the checkerboard mixture can be separated into two equally weighted components within which $x$ and $y$ are mutually independent. Under the first component, $x$ and $y$ both lie within $[0, 1) \cup [2, 3)$, and under the second component $x$ and $y$ both lie within $[1, 2) \cup [3, 4)$. Note that each of these components has four spatially disconnected regions---the classification cannot be achieved using spatial clustering techniques. This example also naturally evokes the intuition of the variates $x$ and $y$ serving as each other's supervisory signal, since the value of either $x$ or $y$ uniquely determines the component the datapoint belongs to.
Let us now analyze this toy example using an InClass net. All the details of the network training process are identical to the analysis of the example in \hyperref[subsec:bivariate_gaussian]{Section \ref*{subsec:bivariate_gaussian}}, including the network architectures, the size of the training dataset, the choice of optimizer, batch size and epoch count. The estimated mixture weights are $w_1=0.501, w_2=0.499$, which is in excellent agreement with their true values of $w_1 = w_2 = 0.5$. \hyperref[fig:checkerboard_dists]{Figure \ref*{fig:checkerboard_dists}} shows, in solid red curves, the distributions of $x$ (left panel) and $y$ (right panel) under the first component learned by the network, using the same procedure as in \hyperref[subsec:bivariate_gaussian]{Section \ref*{subsec:bivariate_gaussian}}. We only show the first component in this figure for the sake of clarity---the second component fills the gaps in the univariate distributions of $x$ and $y$ so that $w_1\,f^{(1)}_{x} + w_2\,f^{(2)}_{x}$ and $w_1\,f^{(1)}_{y} + w_2\,f^{(2)}_{y}$ are constant. For comparison, the true “rectangular wave” distributions are also shown as green dash-dot curves, which are also seen to agree with the estimates.
The biggest advantage offered by a machine learning based technique over existing non-machine-learning techniques for nonparametric mixture model estimation is the possibility of tackling high-dimensional data. As a proof of concept, in this section we will train an InClass net to classify images of handwritten digits from the MNIST database lecun2010mnist, with the classes corresponding to the digits $0\text{--}9$. With this example, we will focus more on the data classification aspect of this paper than the mixture model estimation.
The MNIST dataset contains $28\text{px}\times28\text{px}$ grayscale images of handwritten digits. Each image also has an associated label indicating the digit contained in the image. We will construct a bivariate mixture model out of the MNIST dataset, where each datapoint is a pair of images. A single datapoint of the dataset will be sampled by first choosing a class between $0\text{--}9$ uniformly at random, and then sampling two images\footnote{Alternatively, one can generate image pairs by sampling the first image from the MNIST dataset, and applying a random transformation on the sampled image to get the second, as in Xiang2019.} containing that digit uniformly from the MNIST dataset (with replacement). This gives us a bivariate conditional independence mixture model with 10 classes of equal mixture weights---note that within each component (or class), the two images are mutually independent of each other. \hyperref[fig:digits]{Figure \ref*{fig:digits}} illustrates the kind of data the InClass net will see, with 5 randomly chosen datapoints from the mixture model (one in each column).
For analyzing this dataset, instead of creating two different classifiers for the variates, we use the same neural network for classifying both $x$ and $y$. Viewed differently, the networks classifying $x$ and $y$ are identical in architecture and share their weights as well, and only differ in the input (output) they receive (return). The network uses a sequential architecture and contains, in order, a layer to flatten the $28\times28$ image data, 3 dense layers each with the ReLU activation function and 32 nodes, and finally a dense layer with the softmax activation function and 10 nodes (since $C=10$). This network has a total of 27,562 trainable parameters.
In principle, an InClass net can be trained without supervision to distinguish the digits. However, considering the large number of input dimensions and classes, without supervision, our network is expected to have difficulties “discovering” new classes in the data, and will end up in bad local minima of the cost function. We will discuss some ways of overcoming this difficulty in \hyperref[appendix:gradient]{\ref*{appendix:gradient}}.
In this example, we addressed this issue by taking a semi-supervised approach: We “seeded” the classes (digits) in the network by performing a supervised training over a small dataset with noisy labels. For this purpose, we used a training dataset containing 2,000 images. The noisy label associated with each image matches its true label with probability $0.6$, and matches one of the other 9 incorrect labels (chosen uniformly) with probability $0.4$. The network was trained using the categorical cross-entropy loss function with the Adam optimizer for 30 epochs (batch size 20).
After this pre-training, we trained the network further using our neg_ctc_cost function on 100,000 pairs of images from our mixture model\footnote{The 200,000 images were all sampled with replacament from a set of 60,000 total images in the MNIST training dataset---repetitions will occur within the dataset.}. 10% of the 100,000 datapoints were set aside as a validation dataset to monitor the evolution of the network performance, though no hyperparameter optimization was actively performed using the validation data. The training was done using the Adam optimizer with a batch size of 100 for 20 epochs.
Finally, we evaluated the performance of the classifier on a testing dataset of 10,000 single images unseen by the network (either during training or during validation). The performance is illustrated as a confusion matrix in the left panel of \hyperref[fig:confusion]{Figure \ref*{fig:confusion}}. Each row of the confusion matrix shows the output of the network averaged over test images containing a given digit (true label), both as a heatmap and as numerical values within each cell of the matrix. Recall that our network output, for each image, is 10-dimensional and can be interpreted as the probabilities assigned by the network to the different classes. Because the classes were pre-seeded into the network in a supervised manner, they matched with the true classes without requiring any manual reassignment.
For comparison, we also show the confusion matrix of the network after the supervised pre-training performed on the noisily labelled data in the right panel of \hyperref[fig:confusion]{Figure \ref*{fig:confusion}}. Note that the training that resulted in the performance improvement from the right panel to the left panel was completely unsupervised. We will discuss this semi-supervised training approach in the context of real world applications in \hyperref[subsec:semisupervised]{Section \ref*{subsec:semisupervised}}.
In this example, we will demonstrate that the InClass nets technique works for the estimation of mixture models with more than two variates as well. We consider the mixture of four independent trivariate Gaussians, with the third variate denoted by $z$. \hyperref[tab:four_gaussians]{Table \ref*{tab:four_gaussians}} summarizes the mixture model specification.
The classifier networks $\beta^{(i)}_x$, $\beta^{(i)}_y$, and $\beta^{(i)}_z$ have a similar architecture to the classifier architectures used in \hyperref[subsec:bivariate_gaussian]{Section \ref*{subsec:bivariate_gaussian}}, except that the output layer has 4 nodes, since $C=4$. The InClass net constructed out of the classifiers has a total of 6,924 trainable parameters. We trained the InClass net using 1,000,000 datapoints for 15 epochs with a batch size of 500 (the other details of the training process remained the same as in \hyperref[subsec:bivariate_gaussian]{Section \ref*{subsec:bivariate_gaussian}}), and estimated the mixture model. The estimated mixture weights, shown in the last column of \hyperref[tab:four_gaussians]{Table \ref*{tab:four_gaussians}}, are in good agreement with the true weights of the components (second column in \hyperref[tab:four_gaussians]{Table \ref*{tab:four_gaussians}}). The estimated distributions $f^{(i)}_x$, $f^{(i)}_y$, and $f^{(i)}_z$ of the variates $x$, $y$, and $z$, respectively, are shown in \hyperref[fig:four_gaussians_dists]{Figure \ref*{fig:four_gaussians_dists}} as red solid curves, along with the true distributions depicted as green dash-dot curves. In all twelve cases (4 components $\times$ 3 variates) we observe good agreement between the true and estimated distribution.
Identifiability of a statistical model is concerned with whether the parameters and functions that describe the model are uniquely identifiable from an infinite sample of datapoints produced from the model. The identifiability of mixture models is an important concept, especially in the context of using the techniques introduced in this paper for science applications. For the sake of precision, in this section we will distinguish between a statistical model and an instance of a statistical model as follows:
There are two related notions of identifiability in the literature. An instance of a statistical model is said to be identifiable if it is observationally distinguishable from every other instance of the same model, i.e., if no other instance leads to an equivalent\footnote{In this context, two distributions are considered equivalent if they are equal almost surely.} probability distribution of the observed data. A statistical model is said to be identifiable if every instance of the model is observationally distingushable from every other instance.
For a given statistical model, the definition of what it means to estimate the model is usually chosen to be practically useful. In the context of estimating nonparametric CIMMs, the definition allows for the following “leniencies”:
Even with these leniencies, in general, nonparametric CIMMs are not identifiable. The results in the literature are usually concerned with the identifiability of specific instances of CIMMs---they provide conditions under which instances of CIMMs are identifiable Hall2003,Hall2005,Elmore2005,Allman2009,Tahmasebi2018.
The $(V=1, C\geq 2)$ case (univariate) is always unidentifiable nonparametrically. For the $(V\geq 3, C=2)$ case, Hall2003 provided certain regularity conditions under which instances of CIMMs are identifiable, and Allman2009 generalized the result to the $(V\geq 3, C\geq 2)$ case. The result from Allman2009 states that an instance of a CIMM with $V\geq 3$ and $C\geq 2$ is identifiable if the functions $\left\{f^{(1)}_v,\dots, f^{(C)}_v\right\}$ are linearly independent, for all $v=1,\dots,V$.
This leaves the bivariate case $(V=2)$, which is the main focus of this section. Ref. Hall2003 showed that in the $(V=2, C=2)$ case, instances of nonparameteric CIMMs are not identifiable in general. In particular it was shown that for any instance of a two component bivariate nonparametric CIMM, there exists a two-parameter family of instances which leads to the same distribution of the observed variables $(x, y)$. The authors also noted that non-negativity conditions will introduce constraints on the allowed values for the two parameters. Extending this result from Hall2003, we derive the following two theorems which provide a sufficient and a (different) necessary condition for instances of nonparametric CIMMs with $(V=2, C\geq 2)$ to be identifiable. The two conditions coincide for the $C=2$ case. We relegate the proof of the theorems to \hyperref[appendix:proof]{\ref*{appendix:proof}}.
The essential supremum can be thought of as an adaptation of the notion of supremum of a function, allowing for ignoring the behaviour of the function over regions with a total probability measure\footnote{It is understood that the relevant probability measure in (ref) and (ref) is the one that corresponds to the mixture model itself.} of zero. These conditions can roughly be interpreted as follows: The sufficient condition (ref) will be satisfied if, for every component $i$ and variate $x$ or $y$, there exists some region in the phase space of the variate where component $i$ completely dominates the mixture, i.e., all the datapoints in that region are from component $i$. The necessary condition (ref) will be satisfied if, for every pair of components $i\neq j$ and variate $x$ or $y$, there exists some region in the phase space of the variate where component $i$ completely dominates the mixture of components $i$ and $j$.
Let us now revisit the examples considered earlier in \hyperref[sec:raindancesvi]{Section \ref*{sec:raindancesvi}} from the point of view of identifiablilty. Considering the successful estimation of mixture models and/or classifier training in those examples, we can expect them to be identifiable. For the mixture of two independent bivariate Gaussians, \hyperref[fig:bivariate_gaussian_classifiers]{Figure \ref*{fig:bivariate_gaussian_classifiers}} shows how, for both $x$ and $y$, the true and reconstructed classifier output for the first (second) component approaches $1$ for increasingly negative (positive) values. This ensures that the sufficient condition for identifiability (ref) is satisfied. Similarly, for the checkerboard mixture, by construction, there are regions in $x$ and $y$ which contain points from only component 1 or only component 2, see \hyperref[fig:checkerboard_comps]{Figure \ref*{fig:checkerboard_comps}}.
Recall that in our treatment, the individual variates $x$ and $y$ are themselves allowed to be multi-dimensional. In the special case of one-dimensional variates $x$ and $y$, typically a component will only dominate the mixture in either the left tail or the right tail of the other components. This means that for most natural examples with one-dimensional $x$ and $y$, it is unlikely for the mixture to be identifiable for more than two components. On the other hand, this limitation does not apply to higher dimensional variates $x$ and $y$ which our InClass nets specialize in. For instance, the sufficient condition (ref) for the mixture model constructed out of the MNIST dataset becomes: “For every digit $d$, there must exist some region in the space of images, within which the images look unmistakably like the digit $d$”. This condition is naturally expected to be satisfied, considering the reliability of good handwritten communication. \vskip 2mm \noindentReduced identifiability due to limited statistics. The unique estimation of mixture model instances guaranteed by \hyperref[th:sufficient]{Theorem \ref*{th:sufficient}} can only be achieved with an infinite dataset. There will always be an uncertainty associated with estimation performed using finite datasets Matchev:2020jqz. The level of this uncertainty is related to (among other things) how close the conditions (ref) and (ref) are to being satisfied within the region of sample space covered sufficiently by the finite dataset at hand. In this sense, the result in \hyperref[th:sufficient]{Theorem \ref*{th:sufficient}} is useful from a practical point of view. To illustrate this, we repeated the two Gaussians example from \hyperref[subsec:bivariate_gaussian]{Section \ref*{subsec:bivariate_gaussian}}, with the same setup, but with much fewer datapoints, namely 5,000 instead of 100,000. With fewer datapoints, the dataset is less likely to probe the tails of the $x$ and $y$ distributions, where a single component dominates. As expected, this time the estimation of the weights is slightly worse --- we obtained $w_1=0.44$ and $w_2=0.56$, to be compared with the true values of $w_1=0.4$ and $w_2=0.6$. The results for the classifiers $\alpha^{(i)}_{x}(x)$ and $\alpha^{(i)}_{y}(y)$ and for the component distributions $f^{(i)}_x$ and $f^{(i)}_y$ are shown in Figures (ref) and (ref), respectively. Comparing to the analogous high statistics Figures (ref) and (ref), we see that the estimation has generally succeeded (after all, the model was identifiable), but is not perfect and suffers from statistical uncertainties.
\vskip 2mm \noindentUnidentifiable situations. When a CIMM instance is not identifiable, our technique will yield one of the parameterizations (weights and functions) that best fits the available data. Note that the unidentifiability of an instance of a nonparametric CIMM is not a weakness of our InClass nets approach, but rather a statement on the impossibility of the task of unique estimation.
In this section, we will discuss some considerations which might be relevant in the context of science applications of the InClass nets technique introduced in this paper.
An important aspect of estimating a model (or equivalently, fitting a model to the available data), is providing an uncertainty on the estimate. Uncertainties in parametric estimation are conceptually straightforward---they correspond to the (possibly correlated) uncertainties in the estimated values of the parameters. The corresponding approach in the context of nonparameteric models (which allow arbitrary functions), would be to treat either the neural network outputs $\beta^{(i)}_v$ or the estimated distributions $f^{(i)}_v$ as Gaussian processes 10.5555/1162254. For a Gaussian process $G(\texttt{input})$, the value of $G$ at any finite set of $\texttt{input}$ values is taken to be randomly distributed according to a multivariate normal distribution. This allows us to assign uncertainty estimates on the value of $G$ at individual $\texttt{input}$ points, while also accounting for the correlations in the uncertainties between the values at different $\texttt{input}$ points. It has been shown that Gaussian processes can be modeled using Bayesian neural networks with wide layers neal_priors_1996,Lee2018DeepNN,g.2018gaussian. Using wide Bayesian neural networks as the individual classifiers of the InClass net, one can obtain robust uncertainties on the estimated mixture model. In many scenarios, one is simply interested in visualizing a band of uncertainty around the estimated $\beta^{(i)}_v$-s, $\alpha^{(i)}_v$-s, or $f^{(i)}_v$-s and even narrow Bayesian neural networks may be sufficient for this purpose.
Note that the Gaussian process approach will work when the instance of the CIMM at hand is identifiable, and the uncertainty in the estimated model arises only from the finiteness of the dataset being analyzed. It is presently unclear whether Bayesian neural networks can capture the degree(s) of freedom in the model specification which are introduced by the unidentifiability of the CIMM instance.
In many applications, one does not a priori know the number of components in the CIMM Kasahara2014,mbakop. In other situations, the assumption that the distribution of the data can be written as a CIMM may not necessarily be valid. In such situations, by training different InClass nets with different values of $C$, one may be able to a) verify the validity of the conditional independence assumption, and b) estimate $C$.
Note that increasing the number of components increases the fitting ability of a CIMM. More concretely, every CIMM instance with $C$ components can be thought of as a CIMM instance with $C'>C$ components (with $C'-C$ additional zero-weight components). As a result, an InClass net with more components should strictly perform better (in terms of the minimum cost value achieved), up to network training deficiencies and statistical fluctuations due to the finiteness of the training dataset. However, the improvement (in the minimum cost achieved) resulting from increasing $C$ is expected to diminish beyond a certain point.
In particular, if the true probability distribution of the data ${\mathcal{P}}^\ast({\mathcal{X}})$ can be modeled as a CIMM, then there exists a minimum number of components $C_\text{min}$ required to express ${\mathcal{P}}^\ast({\mathcal{X}})$ in the form
Increasing $C$ from $1$ to $C_\text{min}$ will show an improvement in $\texttt{neg\_ctc\_cost}$ value, but beyond $C_\text{min}$, the performance is expected to saturate. This feature, if observed, can simultaneously a) confirm that the data is consistent with the conditional independence assumption, and b) provide an estimate of $C_\text{min}$. Note that the $C_\text{min}$ value identified in this way is only an estimate---inferring the presence of a component with a small mixing weight, or the presence of two components with very similar distributions $f^{(i)}({\mathcal{X}})$ may be statistically limited by the amount of data available. If the actual number of components is a priori unknown, then the $C_\text{min}$ estimate can serve as an Occam's razor estimate of $C$.
On the other hand, if such a sharp saturation of network performance is not observed at a particular value of $C$, and the saturation is more gradual, this could be a sign of a Latent Factor Model---the underlying latent variable that explains the dependence of the different variates could be continuous instead of being the discrete category label $i$.
When estimating $C_\text{min}$ using the method described in \hyperref[subsec:C_estimation]{Section \ref*{subsec:C_estimation}}, one relies on observing a saturation in the value of the minimum cost achieved. However, such a saturation could also result from deficiencies in the architecture and/or training of the network. It it therefore useful to have an estimate of the minimum possible $\texttt{neg\_ctc\_cost}$ achievable by the best fitting model. Recall from (ref), that
where $\mathrm{KL}\left[{\mathcal{P}}^\ast~\big|\big|~{\mathcal{P}} \right]$is the KL divergence from the distribution represented by the InClass net ${\mathcal{P}}$ to the true distribution ${\mathcal{P}}^\ast$ and $C^\ast(x_1,\dots,x_V)$ is the total correlation of the variates under the true distribution. Since, the KL divergence is manifestly non-negative and equals 0 only when ${\mathcal{P}}^\ast$ is equivalent to ${\mathcal{P}}$, we have the following inequality
where the equality is achieved when ${\mathcal{P}}^\ast$ matches ${\mathcal{P}}$ almost surely. Thus, the negative total correlation $-C^\ast$ provides a (theoretically achievable) lower-bound on the negative cross total correlation $\texttt{neg\_ctc\_cost}$. From (ref) the total correlation $C^\ast$ is given by
where ${\mathcal{P}}_{\!v}^\ast$ represents the marginal distribution of $x_v$ in the data. For low dimensional data, $C^\ast$ can be estimated directly using this formula, after first estimating the distributions ${\mathcal{P}}^\ast({\mathcal{X}})$ and ${\mathcal{P}}_{\!v}^\ast(x_v)$.
Alternatively, for both low and high dimensional data, one can estimate $C^\ast$ using supervised machine learning as follows. Let the distribution ${\mathcal{Q}}^\ast({\mathcal{X}})$ be defined as
Note that $C^\ast$ is simply the Kullback--Leibler divergence $\mathrm{KL}\left[{\mathcal{P}}^\ast~\big|\big|~{\mathcal{Q}}^\ast \right]$ from ${\mathcal{Q}}^\ast$ to ${\mathcal{P}}^\ast$. One can produce datapoints as per the distribution ${\mathcal{Q}}^\ast({\mathcal{X}})$ by independently sampling the variates $x_1,\dots,x_V$ from the available dataset. This gives us two datasets: the original one distributed as per ${\mathcal{P}}^\ast$ and a resampled one distributed as per ${\mathcal{Q}}^\ast$. One can train a machine in a supervised manner to distinguish between these two datasets and estimate the KL divergence, and hence $C^\ast$, from the trained classifier.
In some applications, one may have additional prior knowledge about the mixture model, beyond the conditional independence assumption. It may be possible to incorporate this knowledge into the InClass net directly. For example, in the MNIST image classification example considered in \hyperref[sec:MNIST]{Section \ref*{sec:MNIST}}, we used the information that the variates $x$ and $y$ are both images of digits to use the same classifier neural network for both variates.
As a different example, if the distribution of a given variate $x_v$ is known under a given component $i$, then the value of $\beta^{(i)}_v$ can be set to $f^{(i)}_v(x_v) / {\mathcal{P}}_{\!v}^\ast(x_v)$ up to a multiplicative weight factor which will constitute a single, trainable parameter. For the special case where the distribution of a given variate $x_v$ is known under every component, the classifier for the $v$-th variate can be parameterized by only the mixture weights of the components---in this way the InClass nets technique can be applied in the situations where the ${}_s\mathcal{P}lots$ technique is currently being used in high energy physics.
As another example, consider the case where the weights of the different components are a priori known (but not the distributions of the variates within the components). Then an extra term can be added to the cost function to force the mixture weights $w_i$ estimated by the InClass net towards the true known weights $w_i^\text{true}$. One possible form of the extra term is inspired by the cross entropy:
where $\lambda$ is a parameter that controls the relative importance of the new term in the cost function. The additional term could be added either at the beginning, or after training the network for a few epochs (and identifying the map from the true component indices to the learned component indices). The additional term may be particularly useful in estimating unidentifiable CIMMs, where the additional knowledge of the mixture weights could help rule out the observationally indistinguishable “fake” CIMM instances.
In this section we will discuss some potential variations and extensions of the InClass nets technique introduced in this paper. A detailed exploration of these ideas is beyond the scope of this work.
For the cost function $\texttt{neg\_ctc\_cost}$ from (ref), we can define a surrogate cost function $\texttt{unnorm\_neg\_ctc\_cost}$ as
This surrogate cost function is an alternative cost function whose minimization will also lead to the minimization of the $\texttt{neg\_ctc\_cost}$. More concretely, if the true distribution ${\mathcal{P}}^\ast$ does correspond to a conditional independence mixture model, then the surrogate cost function will be minimized only when the network outputs (pseudo classifiers) $\beta^{(i)}_v$ match the classifiers $\alpha^{(i)}_v$ that correspond to a best fitting CIMM instance. This can be proved as follows: From (ref) and (ref), we have
In (ref), we have used the inequality of arithmetic and geometric means and in step (ref), we have used the constraint $\displaystyle\sum_{i=1}^C~\beta^{(i)}_v(x_v) = 1$ satisfied by the neural network outputs. Note that setting the pseudo classifiers $\beta^{(i)}_v$ to be equal to the classifiers $\alpha^{(i)}$ corresponding to a best fitting CIMM instance both a) minimizes $\texttt{neg\_ctc\_cost}$, and b) satisfies the condition for equality in (ref). This completes the proof that the unnorm_neg_ctc_cost is a suggogate cost function for the neg_ctc_cost when the data does correspond to some CIMM. The bivariate special case unnorm_neg_cmi_cost which is a surrogate cost function for the neg_cmi_cost can be explictly written as
These surrogate cost functions are also implemented in the RainDancesVI package.
Recall that the dataset at hand could be consistent with multiple CIMM instances, either due to the unidentifiability of the instance or due to the finiteness of the dataset. This allows us to add additional properties for the learned model to satisfy. For example, depending on the application at hand, one might be interested in reducing the number of dominant (high mixture weight) components in the model identified by the network. This can be encouraged by adding (to the cost function) additional regularization terms like
where $\lambda$ is a positive constant. If one is interested in more evenly weighted components, the same regularizer terms can be used with $\lambda$ set to a negative value.
The InClass nets architecture introduced in this paper can have more general data mining applications beyond the estimation of CIMMs. If the datapoints in a dataset are comprised of the (possibly multi-dimensional) variates $x_1,\dots,x_V$, the joint distribution of these variates may be understandable in terms of the classes of datapoints within the dataset, even if it does not fall under a CIMM. Furthermore, the set of classes corresponding to one variate need not necessarily be the same as the set of classes corresponding to another. In the literature, the existence of different sets of classes within the dataset falls under the realm of multi-label classification.
For example, consider a dataset containing paired data: each datapoint contains the identities of a book and a movie liked by a person. The working assumption could be that there exists a classification of books and a (different) classification of movies, such that the class of books liked by a person is related to the class of movies liked by the same person. In such cases, it may be possible to simultaneously train a book and a movie classifier using the InClass nets architecture, by simply maximizing the mutual information between the classes predicted by the network.
To this end, we define the “negative total correlation function” $\texttt{neg\_tc\_cost}$ and its bivariate special case “negative mutual information” cost function $\texttt{neg\_mi\_cost}$ as
where the outputs of the InClass net $\eta^{(i)}_v$ are directly interpreted as the classifier output $\alpha^{(i)}_v$, and $C_v$ is the number of classes for the classifier corresponding to the $v$-th variate---note that the $C_v$-s need not all be equal. We point out that the neg_tc_cost of (ref) is a generalization of the cost function used in ji2019invariant for the case where a) there can be more than two variates in the data, b) the classifiers $\alpha^{(i)}_v$ are not necessarily the same for different variates, and c) the number of classes $C_v$ could be different for different variates. Also, in this formulation, it is not required that the inputs $x_v$ to the different classifiers have only non-overlapping attributes of the datapoint.
In the MNIST image classification example considered in \hyperref[sec:MNIST]{Section \ref*{sec:MNIST}}, we seeded the categories into the classifier network via supervised learning using a small, noisily labeled dataset. After the categories were seeded in, we used the unsupervised training of the InClass nets technique to further train the network.
This strategy has straightforward applications in semi-supervised learning scenarios where only a subset of the datapoints in the training dataset are labeled. For example, in the training of neural networks to perform medical diagnosis 2005.07377, generating labeled datasets requires manual annotation by experts, and only a small number of labeled samples may be available. On the other hand, a large number of unlabeled samples are typically available for training purposes. If, say, two different aspects (or variates) of the medical records are expected to only be weakly dependent on each other, but a confounding factor like the presence or absence of a disease can influence both variates, then we can train a neural network to perform the diagnosis leveraging both the labeled and unlabeled datasets. A hybrid cost function that incorporates a supervised classification cost function (for the labeled datapoints), as well an unsupervised cost function introduced in this paper (for the unlabeled datapoints) may be appropriate for the task.
Note that the medical diagnosis example considered here will not strictly be a conditional independence mixture model. For instance, in addition to the presence or absence of the disease, the severity of a particular case is also likely to influence the medical record. It may be possible to accommodate this particular effect by having multiple labels for different severity levels. Despite not strictly being an example of conditional indepedence mixture model, training using the $\texttt{neg\_ctc\_cost}$ or the $\texttt{neg\_tc\_cost}$ can still potentially yield useful diagnostic tools.
In this paper we introduced a novel approach for the nonparameteric estimation of conditional independence mixture models defined by (ref). In this approach, the estimation of a CIMM is treated as a multi-class classification problem, which we solve with machine-learning methods. The main results of the paper are as follows.
As discussed in Sections (ref) and (ref), the InClass nets technique has many potential applications beyond the narrow focus of CIMM. Specifically, the use of machine learning opens new avenues for addressing old-standing problems in nonparametric statistics.
The authors would like to thank M. Lisanti and A. Roman for useful discussions. The work of PS was supported in part by the University of Florida CLAS Dissertation Fellowship (funded by the Charles Vincent and Heidi Cole McLaughlin Endowment) and the Institute of Fundamental Theory Fellowship. This work was supported in part by the United States Department of Energy under Grant No. DE-SC0010296.
The code and data that support the findings of this study are openly available at the following URL: \url{https://gitlab.com/prasanthcakewalk/code-and-data-availability/}.