Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
119,603 characters · 15 sections · 90 citation commands
Handling Sparse Non-negative Data in Finance
Data with non-negative outcome (dependent) variables are ubiquitous in applications across a diverse range of disciplines, including finance coles2006managerial,hirshleifer2012overconfident,akey2021limits, economics hausman1984econometric,card2001estimating, health services dow2003choosing, epidemiology holfold1980rates,bohning1999zero, political science king1989event, and sociology boulton2018analyzing.
Specific examples include discrete count data, such as the number of corporate defaults and patents, and continuous count-like data, such as health expenditures and sales revenue. A popular approach to modeling non-negative data, particularly in finance and economics, is regressing the log transformation of the outcome variable (often shifted by a positive constant such as 1) on explanatory variables.
Despite the popularity of this log-linear regression approach, there is increasing evidence that it may not be the best approach to modeling non-negative data. For example, cohn2022count,chen2024logs point out the lack of meaningful interpretation of log-linear estimates, while king1988statistical,silva2006log,silva2011further,mullahy2022transform demonstrate the fragility of log-linear regression in the presence of heteroskedasticity or sparsity, two prominent features of count and non-negative data. As a solution, recent works have advocated for the use of pseudo maximum likelihood (PML) estimation gourieroux1984pseudoa based on count models, most notably the Poisson regression wooldridge1999quasi.\footnote{In the econometrics literature, some works use the term “quasi-maximum likelihood” white1982maximum,wooldridge1999quasi. In this paper, we follow the terminology of gourieroux1984pseudoa and silva2006log and use “pseudo maximum likelihood”.} Compared to log-linear regression, Poisson regression possesses several desirable properties that facilitate its application in practice, such as interpretability, robustness, computational efficiency, and the ability to accommodate separable group fixed effects and instrumental variables windmeijer1997endogeneity,mullahy1997instrumental,correia2020fast,cohn2022count,chang2024inferring. These considerations have led to a shift towards Poisson regression in economics and finance sautner2023firm,addoum2023temperature, hollingsworth2024gift.
Importantly, estimates of model parameters can change significantly when switching from a log-linear to a Poisson specification. For example, cohn2022count find
Their findings highlight the importance of selecting an appropriate regression model in empirical research, including corporate finance studies on topics such as innovation he2013dark,fang2014does and industrial pollution akey2021limits,xu2022financial. More broadly, a shift from log-linear to Poisson regression may carry significant implications for empirical work across finance and economics. This naturally raises a fundamental question: should Poisson regression be the {\it default choice} when analyzing non-negative, count-like data?
Existing works have only partially addressed this question by focusing on {\it either} the heteroskedasticity {\it or} the sparsity of the data. For example, the efficiency properties of Poisson and other PML estimators under heteroskedasticity are well-known and depend on the dispersion of data gourieroux1984pseudob,cameron2013regression. On the other hand, when data exhibits high levels of sparsity, silva2006log,silva2011further provide simulation evidence on the robustness of Poisson PML compared to log-linear regression and other PML estimators. Alternatives such as zero-inflated lambert1992zero and hurdle models mullahy1986specification for sparse non-negative data have also been studied.
However, as is well-known to empirical researchers and also illustrated by our examples in (ref), non-negative data in many finance and economics applications frequently exhibit both sparsity and heteroskedasticity. Yet, few studies address these characteristics simultaneously, and it remains unclear whether Poisson regression should be used as a rule of thumb when handling sparse non-negative data. This major gap motivates the central question of our paper: How can we effectively model non-negative data when both sparsity and heteroskedasticity are prominent?
We address this question by developing a novel unified framework that integrates classical pseudo maximum likelihood methods, including Poisson regression, which so far have often been applied in an ad hoc manner. As a main contribution of this paper, we show that the optimal modeling choice depends on the relative prominence of each feature in the data, and could diverge considerably from the Poisson estimator. Furthermore, we provide a simple data-driven procedure that empirical researchers can employ to determine the optimal choice.
The key component of our framework is a novel class of moment estimators that includes classical PML estimators for non-negative data, such as Poisson and gamma, as special cases. These estimators place heterogeneous weights on observations in the estimating equations. On one hand, heteroskedasticity encourages larger weights on samples with smaller conditional means to improve efficiency. On the other hand, the presence of many zeros discourages such weighting to reduce bias. As a result, there exists an estimator that optimally balances the fundamental bias-variance trade-off created by these two aspects of the data. As illustrated in (ref) and (ref), the optimal estimator is sensitive to the level of sparsity and heteroskedasticity, and can improve significantly upon the Poisson PML estimator (see (ref) for full details). Notably, although non-linear least squares (NLS) has largely been avoided due to its inefficiency under heteroskedasticity, our simulation results suggest that it could be preferable when heteroskedasticity is mild compared to sparsity. Moreover, this trade-off is a {\it unique} feature of data with sparse non-negative outcome variables, setting it apart from the conventional bias-variance trade-off in statistics and machine learning, which is typically driven by model complexity.
We apply our framework to simulated and real finance datasets, including those related to corporate defaults, credit ratings, corporate patents hirshleifer2012overconfident, and real estate investments bekkerman2023effect. We find that the estimator recommended by our approach varies significantly with the dataset, and can improve upon the Poisson regression estimator. For instance, an estimator close to the Poisson regression estimator yields the best fit when modeling industry-level defaults. In contrast, for credit ratings and corporate patents data, the optimal estimator recommended by our framework differs from and outperforms Poisson significantly, while for real estate investment data our framework recommends the gamma PML estimator. Across these datasets, the selected estimator can reduce the out-of-sample root mean squared error (RMSE) by 50% to 90% compared to the Poisson PML. It can also provide a better description of the sparsity level in the data compared to Poisson regression, as illustrated in (ref). Moreover, the choice of the estimator can result in substantially different parameter estimates and economic conclusions ((ref) in (ref)). These findings challenge the practice of applying a single regression model across all contexts involving non-negative dependent variables, and underscore the importance of evaluating model suitability for the specific dataset at hand.
The rest of the paper is organized as follows. In (ref), we provide the econometric background on non-negative data modeling and introduce the formal setup and key concepts of the paper. In (ref), we propose a class of moment estimators that we call generalized pseudo maximum likelihood estimators. We characterize the bias and variance of these estimators in the presence of heteroskedasticity and excess sparsity, and propose a cross-validation procedure to select the optimal generalized PML estimator. In (ref), we conduct extensive numerical experiments to study the bias-variance trade-off behavior of generalized PML estimators under excess sparsity and heteroskedasticity. In (ref), we apply our framework to major datasets in finance.
Data with count or non-negative dependent variables often exhibits two prominent features: heteroskedasticity and sparsity. We begin by providing econometric background on these characteristics. To To contextualize the discussion, we present several relevant regression models that help motivate our framework and main results.
One of the most popular approaches to modeling data with non-negative dependent variables is log-linear regression, i.e., assuming that the logarithm of the response variable $Y \in \mathbb{R}^+$ is linear in the covariates $X\in \mathbb{R}^d$. To deal with zero values of $Y$, this approach typically uses the transformation $\log(1+Y)$, resulting in the specification
Prominent applications of log-linear regression include the modeling of trade data and panel data with non-negative outcomes, such as earnings card2001estimating, and count and count-like data in finance, such as number of corporate patents granted hirshleifer2012overconfident and firms' toxic waste release volumes akey2021limits. However, previous works have raised some important drawbacks of the log-linear regression approach. For example, cohn2022count demonstrate that log-linear models do not always produce estimates with economically meaningful interpretations, and can even result in the wrong sign in expectation. chen2024logs similarly show that treatment effect estimates based on log-linear transformations are not invariant under change of units and therefore should not be interpreted as percentages. More importantly, as discussed in previous works, log-linear regression is not robust against heteroskedastic silva2006log,cohn2022count and sparse data silva2010existence,silva2011further. Similar issues with other non-linear transformations of non-negative outcome variables, such as the inverse hyperbolic sine (IHS), have been studied by mullahy1998much, manning1998logged, ai2000standard, and mullahy2022transform.
As a more robust alternative, previous works have recommended using pseudo maximum likelihood (PML) estimation white1982maximum,gourieroux1984pseudoa,gourieroux1984pseudob,wooldridge1999quasi. PML has been widely applied to model data when the true likelihood is difficult to specify, e.g., for generalized linear models mccullagh2019generalized. The Poisson PML (PPML) is a popular choice in practice for modeling non-negative outcomes, especially count data cameron2013regression. When $Y\in\mathbb{Z}^+$ is integer valued, a natural modeling approach is to assume that $Y$ conditional on $X$ has a Poisson distribution with rate parameter $\lambda(X,\theta)$ that is log-linear in $X$ with parameter $\theta$, i.e., $\lambda(X,\theta)=\exp(\theta^{T}X)$. The likelihood is given by
The maximum likelihood estimator (MLE) $\hat{\theta}$ of $\theta$, given $n$ independent and identically distributed (i.i.d.) samples $\{y_{i},x_{i}\}_{i=1}^{n}$, is then obtained as the maximizer of the log-likelihood
modulo constants. When the true distribution of $Y$ given $X$ is not Poisson, the PML approach continues to maximize (ref). A seminal result of gourieroux1984pseudob implies that, under standard assumptions, even when $Y \in \mathbb{R}^+$ is not Poisson distributed or even integer valued, the PML estimator which maximizes (ref) is still consistent as long as the conditional mean of $Y$ given $X$ is log-linear, i.e., $\mathbb{E}(Y \mid X) =\exp({\theta^{T}X})$. Importantly, the consistency of Poisson PML holds even if the data is heteroskedastic, i.e., $\text{Var}(Y\mid X)$ depends on $X$. In contrast, log-linear regression is only consistent when the data is homoskedastic.\footnote{The crucial distinction between PPML and log-linear regression lies in the ordering of expectation and logarithm operations. Whereas log-linear regression assumes that the conditional expectation of the log of $Y$ (or $Y+1$) is linear in $X$, i.e., $\mathbb{E}(\log Y\mid X) =\beta^{T}X$, PPML assumes that the log of the conditional expectation of $Y$ is linear in $X$, i.e., $\log(\mathbb{E}(Y\mid X)) =\theta^{T}X$. The implications of this distinction, including the robustness of PPML to heteroskedasticity, are discussed extensively in silva2006log.}
{Heteroskedasticity} arises naturally when the outcome variable is constrained to be non-negative. Intuitively, when the conditional mean $\mathbb{E}\left(Y\mid X\right) \geq 0$ is large, the residual $Y-\mathbb{E}\left(Y\mid X\right)$ can have larger variance without violating the non-negativity requirement $Y\geq0$. However, when $\mathbb{E}\left(Y\mid X\right)$ is close to zero, the residual needs to have smaller variance in order to satisfy $Y\geq0$. The intrinsic nature of heteroskedasticity in non-negative data can also be justified when economic theory prescribes a model of $\mathbb{E}(Y\mid X)$ with a multiplicative form, e.g., $Y=\exp(\theta^{T}X)\cdot\eta$ with $\mathbb{E}(\eta\mid X)=1$. Examples include the gravity models of international trade anderson1979theoretical,anderson2003gravity. Expressing $Y$ in the standard additive form $Y = \exp(\theta^{T}X) + \varepsilon$ with $\mathbb{E}(\varepsilon\mid X)=0$ yields $\varepsilon = (\eta - 1)\cdot \exp(\theta^T X)$. The conditional variance $\text{Var}(\varepsilon\mid X)$ will in general depend on $X$, even if the conditional variance $\text{Var}(\eta\mid X)$ does not, resulting in heteroskedasticity.
Although Poisson pseudo maximum likelihood is robust to heteroskedasticity, it is only efficient when the conditional variance of $Y$ is equal to the conditional mean of $Y$, i.e., ${\rm Var}(Y\mid X) =\mathbb{E}(Y\mid X)$, often called equi-dispersion. When the conditional variance of $Y$ does not scale linearly with the conditional mean, giving rise to the so-called over-dispersion and under-dispersion phenomena, there exist other consistent PML estimators that are more efficient. For example, the gamma PML (GPML) is efficient when ${\rm Var}(Y\mid X)=\mathbb{E}^2(Y\mid X)$, while the closely related negative binomial PML is efficient when ${\rm Var}(Y\mid X) = \mathbb{E}(Y\mid X) + a \mathbb{E}^ 2(Y\mid X)$ for $a>0$. The non-linear least squares (NLS) estimator can also be viewed as a PML estimator, and is efficient when ${\rm Var}(Y\mid X)$ is constant. In this paper, we use $\alpha\geq 0$ to index heteroskedasticity in the specification
It may appear that under heteroskedasticity, one should always use the estimator that best captures the conditional variance as a function of the conditional mean, by determining the value of $\alpha$. Previous works have proposed these types of methods to determine the dispersion using regression-based tests park1966estimation,cameron1990regression,manning2001estimating. However, that is only half of the story. When the data contains many observations whose outcomes are zero, efficiency alone may not be sufficient to justify the use of a particular estimator. For example, simulation studies in previous works and our paper find that with sparse outcomes resulting from asymmetric censoring, Poisson PML can outperform the gamma PML in terms of mean squared error, even when gamma PML is most efficient. This phenomenon reveals the importance of sparsity when modeling non-negative data, and is a main motivation for the current paper. We therefore discuss sparsity next.
Non-negative count-like data often have distributions that are right-skewed and exhibit sparsity, i.e., having many outcomes with value zero. Such sparsity may arise from features of the data-generating process, including the rarity of the underlying events or the structural properties of networks. For instance, financial networks that capture contractual relationships or interactions between entities are typically sparse and display a core-periphery architecture craig2014interbank. Sparse outcomes also arise naturally when the data is censored near zero. For example, in international trade, measurements of large trade volumes are usually more accurate than those of small trade volumes frankel1993trade, which are often rounded down to zero. Previous studies have proposed various methods to explain and model sparsity in non-negative data. We now focus on two approaches that are particularly relevant to our framework, beginning with a modification of the Poisson process.
Suppose that conditional on explanatory variables $X$, the non-negative outcome variable $Y$ follows a Poisson distribution with mean parameter $\exp(\theta_0^T X)$. Then $\mathbb{P}(Y=0 \mid X) = \exp(-\exp(\theta_0^T X))$. However, in practice, the fraction of zeros observed in data is often not consistent with the predictions of this model, which in a large sample is close to $\int_X \mathbb{P}(Y=0 \mid X) dX$. In this case, one needs to modify the model to account for the discrepancy. A popular proposal is the two-part or hurdle model mullahy1986specification,king1989event,mullahy2022transform. It assumes a two-part generating process of $Y$ given $X$. An independent binomial random event determines whether $Y$ is set to zero. If not, the realization of $Y$ is determined according to a truncated distribution that takes on strictly positive values. In the simplest case, instead of assuming $ \mathbb{P}(Y=0 \mid X) = \exp(-\exp(\theta_0^T X))$ as in a Poisson model, we assume \[ \mathbb{P}(Y=0 \mid X) = \exp(-\beta\exp(\theta_0^T X)), \] which can be understood as the probability of a Poisson variable with mean parameter $\beta\exp(\theta_0^T X)$ being equal to 0, while the distribution of $Y>0$ follows a truncated Poisson with mean parameter $\exp(\theta_0^T X)$, i.e.,
When $\beta=1$, we recover the original Poisson model. When $\beta<1$, this modified data generating process will result in more sparsity than the standard Poisson model, and vice versa for $\beta>1$. More generally, we can use non-Poisson based distributions for both the binomial zero event variable and the positive outcome variable. This hurdle model has been widely used in practice for non-negative data with high levels of sparsity.
We can also understand the data generating process in the hurdle model equivalently as the following two-step censoring model. For example, when $\beta<1$, in the first step, a variable $Y \geq 0$ is drawn according to a Poisson distribution with mean parameter $\exp(\theta_0^T X)$. If $Y>0$, then with probability
$Y$ is censored to $0$. We can verify that the resulting variable is zero with probability $\exp(-\beta\exp(\theta_0^T X))$, and has the same distribution as the truncated Poisson when it is positive. This censoring process is closely related to the zero-inflation model of lambert1992zero, which is another common approach to handle excess zeros. In lambert1992zero, a standard Poisson model is generated with mean parameter $\exp(\theta_0^T X)$. Then with probability \[\frac{1}{1+\exp(\beta\theta_{0}^{T}X)},\] the observation is censored to zero. The main difference with the hurdle model is in the censoring probabilities, which is modeled by a logistic function for the zero-inflated Poisson. See also zorn1998analytic for a discussion of the hurdle model and zero-inflation model.
Importantly, in both the hurdle and the zero-inflation models, the probability of observing zero depends on $\exp(\theta_{0}^{T}X)$ (hence on $X$). For example, in the zero-inflation model with $\beta>0$, the censoring probability $\frac{1}{1+\exp(\beta\theta_{0}^{T}X)}$ decreases in $\exp(\theta_{0}^{T}X)$, the conditional mean. Similarly, in the hurdle model, the probability $\exp(-\beta\exp(\theta_0^T X))$ of observing a zero outcome variable decreases rapidly in $\exp(\theta_0^T X)$. This feature aligns with how excess zeros appear in applications. In survey and census data, a small outcome or response variable is often more likely to be rounded down to zero than a large outcome, which is less prone to measurement errors frankel1997regional. Similarly, count variables of rare events are much more likely to be zero when its conditional mean is close to zero. As we will demonstrate in this paper, this {asymmetric sparsity} feature of non-negative data is important for its modeling, as it calls for the use of estimators that may otherwise be inefficient under heteroskedasticity. For example, with sparsity generated by rounding near zero, silva2006log,silva2011further demonstrate in simulations that Poisson PML is generally more robust than the gamma PML estimator, even under over-dispersion with ${\rm Var}(Y\mid X) =\mathbb{E}^2(Y\mid X)$, when gamma PML is the most efficient PML. Similarly, our findings in (ref) suggest that non-linear least squares can be more robust than the Poisson PML when sparsity is strong, even under equi-dispersion when Poisson is most efficient.
These observations have received relatively less attention than the efficiency properties of PML estimators, but they raise some important questions on PML estimators, including Poisson regression. Although their behaviors with respect to heteroskedasticity and sparsity have been studied separately before, they are less clear when both features are significant. When modeling data with non-negative outcomes, understanding the interactions between these features can help empirical researchers decide which regression approach to use that best handles the sparse non-negative data at hand.
In this paper, we provide a systematic study of the interplays between these two prominent features of non-negative data. We use the model (ref) to capture heteroskedasticity indexed by $\alpha$. To capture excess levels of sparsity, we use the following censoring model:
where the probability $P(X,\theta_0)$ takes one of the following asymmetric forms:
The probabilities specified in (ref) generalize those used in the hurdle and zero-inflation models, with $\beta$ controlling the imbalance in the likelihood of zero between samples with small and large conditional means $\exp(\theta_{0}^{T}X)$. The additional parameter $\tau$ controls the range of $\exp(\theta_{0}^{T}X)$ near zero for which the probability of zero outcomes is significant. A larger $\tau$ leads to a lower level of overall sparsity. As we will see, this probability model induces biases for a family of estimators indexed by $\kappa$, with biases generally increasing and variances decreasing in $\kappa$. The optimal estimator therefore balances this bias-variance trade-off in the presence of excess sparsity and heteroskedasticity.
In this section, we build on the preceding discussions and propose a novel family of estimators, which encompasses existing PML estimators such as Poisson and gamma PML. This family of single parameter estimators place heterogeneous weights on samples with varying values of conditional means, which enable them to balance the relative magnitudes of heteroskedasticity and sparsity. We show that these two salient features of non-negative data create a bias-variance trade-off under the models discussed in the previous section. In particular, the optimal choice of estimator varies significantly depending on the data. To enable empirical researchers to leverage these estimators, we also propose a simple cross-validation method to select the estimator that best fits their data, instead of resorting to a default method, such as the Poisson PML.
PML estimators for non-negative data are generally constructed based on likelihoods of qualitatively distinct distributions such as the Poisson and gamma distributions. In this paper, we focus on another perspective that reveals their closer connections through the moment conditions. gourieroux1984pseudob show that the following PML estimators are all consistent for $\theta_0$ when i.i.d. data is generated from the model $Y=\exp(\theta_0^{T}X)+\varepsilon$:
Note that the first order conditions of the above PML estimators are all of the form
NLS, Poisson, and gamma PML correspond to $c=0$ and $\kappa=1,0,-1$, respectively, while negative binomial corresponds to $c=1/b$ and $\kappa=-1$. Therefore, these PML estimators can be understood as solving a system of equations that balance the observed outcomes $y_i$ with the expected outcomes $\exp(\theta^{T}x_{i})$, weighted by the covariates $x_i$ and an estimator-specific term $(c+\exp(\theta^{T}x_{i}))^{\kappa}$ for different values of $c$ and $\kappa$.
The moment equations (ref) provide some intuition on the behavior of PML estimators with respect to heteroskedasticity and sparsity. For example, the well-known inefficiency of NLS under strong heteroskedasticity can be attributed to the weights $(c+\exp(\theta^{T}x_{i}))^{\kappa}=\exp(\theta^{T}x_{i})$ it places on samples: samples with larger (estimated) conditional means $\exp(\theta^{T}x_{i})$ receive larger weights in (ref). However, such samples also tend to be noisier under heteroskedasticity, when ${\rm Var}(Y\mid X)$ increases with $\mathbb{E}(Y\mid X)$. In contrast, the gamma PML estimator, which uses the weights $\exp(-\theta^{T}x_i)$, places higher weight on samples with smaller conditional means, which are less noisy, leading to efficiency gains under strong heteroskedasticity. The Poisson PML has uniform weights $(c+\exp(\theta^{T}x_{i}))^{\kappa}\equiv 1$ on samples, and can therefore be seen as a mid-point between gamma and NLS. In short, under stronger heteroskedasticity in (ref), i.e., larger $\alpha$ in ${\rm Var}(Y\mid X)=\mathbb{E}^\alpha(Y\mid X)$, PMLs which place smaller weights on samples with larger conditional means, i.e., smaller $\kappa$ in (ref), tend to be more efficient and thus preferable.
However, the opposite is true for bias. When data exhibits exceptional levels of sparsity than predicted by a standard PML model (as in (ref)), the corresponding PML estimator may suffer from bias. Moreover, when samples with smaller conditional means are more likely to be zero, the bias is more {severe} for estimators that place higher weights on such samples. Take for example the model (ref) with censoring probability \[P(X,\theta_0)=\frac{1}{1+ (\tau\exp(\theta_{0}^{T}X))^\beta}\] in (ref). Under such a model, samples with smaller $\exp(\theta_0^{T}x_i)$ become less reliable as they are more likely to be censored. Consequently, a PML estimator that places larger weights on such samples, i.e., having a smaller $\kappa$ in (ref), suffers from more severe bias. This explains the observation by some existing works that when sparsity is strong, Poisson ($\kappa=0$) tends to have smaller bias than gamma PML ($\kappa=-1$).
The above discussions suggest that there is an intrinsic trade-off when applying PML methods to data with non-negative outcomes. On one hand, heteroskedasticity necessitates smaller weights on larger (noisier) samples to improve efficiency. On the other hand, sparsity encourages smaller weights on smaller (more bias-inducing) samples to reduce bias. In the former case, PMLs with small $\kappa$ such as gamma PML and negative binomial are generally preferred, while in the latter case PMLs with larger $\kappa$ such as NLS are generally preferred. silva2006log argue that the Poisson PML is a “reasonable compromise” between NLS and gamma PML when both heteroskedasticity and excess sparsity are present. However, it is clear that depending on the relative magnitude of the two features, there may be other estimators that can better balance the bias and variance. Our work is precisely motivated by this intuition.
We propose the following general class of $Z$-estimators amemiya1985advanced with estimating equations
This class of estimators includes as special cases the standard PML estimators. When $c=0$ and $\kappa=0$, we recover Poisson PML. When $c=0, \kappa=1$, we recover NLS. When $c=0, \kappa=-1$ we recover gamma PML. When $c=1,\kappa=-1$, we recover negative binomial PML. For non-integer values of $\kappa$, the resulting estimator does not arise from pseudo maximum likelihood estimation based on a particular distribution, but their estimating equations generalize those of the standard PML estimators. For this reason, we refer to them as generalized PML estimators\footnote{When $\kappa>0$, the estimators can also be understood as M-estimators that maximize a non-concave sample average objective, especially when some covariates are strictly positive or negative. See (ref).}. We primarily focus on the case with $c=0$, resulting in the estimating equations
Our main message is that, depending on the magnitudes of heteroskedasticity and excess sparsity, generalized PML estimators with different (potentially non-integer) values of $\kappa$ is preferable. In particular, Poisson regression ($\kappa=0$) does not always optimally balance the bias-variance trade-off created by the two features. To illustrate the magnitude of improvements over Poisson regression, we provide the RMSEs of the optimal estimator vs. the Poisson regression estimator in our simulation studies in (ref).
In the rest of the section, we study the asymptotic properties of generalized PML estimators, and provide practical guidance on how to select $\kappa$ in practice.
In this subsection, we provide formulae for the bias and asymptotic variance for the estimators defined by (ref). Under the censoring model (ref), the bias does not have an exact formula in general. Instead, we provide an approximation based on Taylor expansions.
In particular, in the absence of censoring, i.e., $ P(\theta_{0},x)\equiv 0$, so that $b\equiv 0$ and (ref) implies that all generalized PML estimators are consistent estimators of $\theta_0$. Moreover, the bias does not depend on $\alpha$. The quality of the approximation formula in (ref) is an important consideration in practice. Intuitively, since the Taylor expansions used in the derivation of (ref) rely on the small magnitude of $\exp(\theta_{0}^{T}x)$, the formula (ref) provides a good approximation of the bias when the censoring probability and $\kappa$ result in integrals with small densities on $x$ with large $\exp(\theta_{0}^{T}x)$. In (ref), we provide simulation evidence that the approximation formula we derive is close to the true bias in a wide range of settings.
Next, we compute the asymptotic variance of the generalized PML estimators, taking into account the bias that results from the censoring model (ref).
The proof of (ref) relies on asymptotic theory of method of moment estimators newey1994large. The uniform boundedness assumptions in (ref) and (ref) are standard and ensure bounded moments and the use of uniform law of large numbers. In the case of no censoring and consistent estimation of $\theta_0$, the asymptotic variance reduces to
which implies that for heteroskedasticity indexed by $\alpha$ in (ref), the most efficient generalized PML estimator is the one with $\kappa=1-\alpha$, a fact known in existing works for Poisson ($\kappa=0$), gamma ($\kappa=-1$), and NLS ($\kappa=1$).
As discussed in (ref), the generalized PML estimators defined by the moment equations (ref) balance the bias-variance trade-off created by heteroskedasticity and excess sparsity. The asymptotic properties of generalized PML estimators established in this section can help provide quantitative descriptions of this trade-off. For example, when $\alpha=0$, $x$ is uniform on $[0,1]$, and $P(\theta_0,x)$ is a threshold censoring model, both the bias and variance of NLS are upper bounded by those of the Poisson PML. As $\alpha$ increases from 0 to 2, the biases do not change, while the variance of NLS increases dramatically relative to that of the Poisson PML (see (ref)). These behaviors lead to a particular value of $\overline{\alpha}\in(0,2)$ where the RMSE of NLS starts to exceed that of the Poisson. Therefore, below heteroskedasticity level $\overline{\alpha}$, NLS should in fact be preferred than Poisson, even though NLS may have larger variance. We illustrate this phenomenon in (ref), based on the same simulation design as that used in (ref) and (ref). We see that the bias of estimators decreases with $\kappa$, as the corresponding estimator places less weight on samples with small conditional means. Meanwhile, as $\alpha$ increases, estimators with smaller $\kappa$ becomes more efficient. The trade-off is manifested through the observation that below $\overline{\alpha}\approx 1.4$, the RMSE of NLS is the smallest, despite Poisson and gamma being more efficient.
More generally, for a particular level of heteroskedasticity and sparsity, there exists a $\kappa$ that optimally balances the bias-variance trade-off. If we consider the two dimensional plane where the x axis is $\alpha$, which quantifies heteroskedasticity, and the y axis is $\tau$, which quantifies the sparsity level, we observe a phase transition, where the optimal $\kappa$ changes depending on the particular region of the plane. In (ref), we demonstrate through numerical experiments that such a phase transition indeed occurs ((ref)).
The bias-variance trade-off considered in this paper has important implications on how one should handle data with non-negative outcomes in practice. Although Poisson regression is now a popular choice, our work calls for a more principled approach to select the generalized PML estimator with optimal $\kappa$ given the magnitudes of heteroskedasticity and excess sparsity of the data at hand. These characteristics can in principle be determined from data, using for example statistical tests park1966estimation,mullahy1986specification. In order to then determine the optimal $\kappa$, a natural idea is to leverage the asymptotic results developed in this section to assess the bias-variance trade-off. However, there are several challenges to this approach. First, it requires determining the parameters $\alpha$ in the heteroskedasticity model (ref) and $\beta,\tau$ in the censoring model. Second, the formulae for bias and asymptotic variance depend on $\theta_0$, whose estimator $\hat \theta$ based on (ref) is biased under a censoring model, which may affect the accuracy of the bias and variance estimates. In this work, we propose an alternative, simpler method to select $\kappa$, which we describe next.
In practice, we can determine the optimal $\kappa$ using a standard $k$-fold cross-validation. More precisely, consider a dataset consisting of $n$ i.i.d. samples $\{y_{i},x_{i}\}_{i=1}^{n}$ where $y_i\geq0$. We split the dataset randomly into $k$ folds $F_1\cup \dots\cup F_k=\{y_{i},x_{i}\}_{i=1}^{n}$ with roughly equal sizes. For each fold $j=1,\dots,k$, we construct generalized PML estimators $\hat \theta^{(j)}_\kappa$ with $\kappa$ in a grid on $[-b,b]$ using data from the complement $\cup_{j'\neq j}F_{j'}$ of $F_j$, and compute the out-of-sample MSE \[e_{\kappa}^{(j)}=\frac{1}{|F_{j}|}\sum_{(x_{i},y_{i})\in F_{j}}(y_{i}-\exp(x_{i}^{T}\hat{\theta}_{\kappa}^{(j)}))^{2}.\] We then average the MSEs across all folds, i.e., $e_\kappa = \frac{1}{k}\sum_{j=1}^k e_{\kappa}^{(j)}$, and select the $\kappa$ that minimizes $e_\kappa$. In (ref), we apply this procedure to four datasets from finance applications, and demonstrate that the optimal $\kappa$ can vary significantly across applications ((ref)).
In this section, we verify our theoretical findings and demonstrate the prevalence of the bias-variance trade-off for data with non-negative outcomes.
Our simulation design follows the setup in silva2006log, combined with the censoring model introduced in this paper. First, we generate a 2-dimensional covariate $X$ with its first coordinate $X_1$ drawn from a standard normal distribution and second coordinate $X_2$ drawn from the uniform distribution on $[0,1]$. Given $X$, the outcome variable is generated according to the model (ref):
where the parameter of interest $\theta_0=[1,1]$ and the multiplicative error $\eta_i$ is log-normal, i.e., $\log \eta_i$ is normal with zero mean. The variance of $\eta_i$ is set to be $\exp^{\alpha-2}(\theta_{0}^{T}X_{i})$, so that the conditional variance of $\exp(\theta_{0}^{T}X_{i})\cdot\eta_{i}$ is equal to $\exp^{\alpha}(\theta_{0}^{T}X_{i})$. The censoring probability $P(\theta_{0},X_{i})$ is given by \[P(X_i,\theta_0) =\frac{1}{(1+ (\tau\exp(\theta_{0}^{T}X_{i}))^\beta)},\] which can be viewed as continuous approximations to the step function. Compared to the double exponential specification in (ref), the power decay specification provides a heavier tail, but their qualitative behaviors with respect to $\beta$ and $\tau$ are similar. $\beta$ regulates the contrast in censoring probabilities between samples with small or large conditional means. As $\beta$ increases from 1, samples with $\tau \exp(\theta_{0}^{T}X_{i})<1$ become increasingly susceptible to censoring, while samples with $\tau \exp(\theta_{0}^{T}X_{i})\geq1$ become less affected. $\tau$ regulates the point of “discontinuity”. As $\tau$ increases from 0, the range of $\exp( \theta_{0}^{T}X_{i})$ for which censoring is significant becomes smaller and more concentrated around 0, resulting in overall less sparsity. For some constant $b\geq 1$, we consider the family of generalized PML estimators defined by (ref):
and study their bias and variance. A practical consideration is how to obtain the generalized PML estimators for $\kappa \notin \{-1,0,1\}$. In (ref), we propose an approach by solving optimization problems whose first order conditions recover the estimating equations (ref).
In this part of the simulation studies, we investigate the various manifestations of the bias-variance trade-off, by holding subsets of the tuple $(\alpha,\tau,\beta,\kappa)$ in our framework fixed and varying the rest.
First, in (ref) we plot the bias, standard deviation, and root mean squared error (RMSE) of generalized PML estimators as a function of the degree of heteroskedasticity $\alpha\in[0,2]$, for fixed $\beta=2, \tau=2$. We find that bias is decreasing in $\kappa$, and is essentially flat in $\alpha$, which is expected since the bias approximation formula does not depend on $\alpha$. Moreover, the standard deviation increases with $\alpha$ for most estimators, but those with larger $\kappa$ experience faster increase of the standard deviation, as they are less efficient under heteroskedasticity. As a result, when $\alpha$ grows, for the first coordinate of $\theta_0$, the standard deviation of NLS starts to be dominated by that of the Poisson at $\alpha\approx 0.6$, which in turn is lower bounded by the standard deviation of the generalized PML estimator with $\kappa=-0.5$ starting $\alpha\approx 1.35$ and lower bounded by the gamma PML starting $\alpha\approx 1.5$. Interestingly, the standard deviation of gamma PML ($\kappa=-1$) is larger than that of $\kappa=-0.5$ even for $\alpha=2$, when in the absence of asymmetric censoring gamma PML is most efficient. This is due to censoring having an effect on the asymptotic variance.
If we ignore the bias from asymmetric censoring, we might conclude that Poisson PML is preferable for heteroskedasticity levels $\alpha \in [0.6, 1.35]$-a rule of thumb that is, in fact, often recommended in practice. However, the RMSE comparisons in (ref) suggest that the NLS in fact dominates the Poisson (and gamma) for all $\alpha \leq 1.35$, due to its smaller bias. Moreover, the efficiency gain of gamma PML and $\kappa=-0.5$ for large $\alpha$ is overwhelmed by their large biases, making the Poisson PML the best option for $\alpha>1.35$. Similar observations hold for the second coordinate of $\theta_0$. These results suggest that in the presence of sparsity, conventional consensus on the superiority of Poisson PML may be incomplete and misguided, and that the bias plays an important role in a more systematic analysis. Our simulation results in (ref) suggest that, under the common asymmetric censoring model that results in excess sparsity, the RMSE of NLS continues to dominate that of the Poisson PML beyond $\alpha=1$, with a transition threshold of $\alpha\approx 1.35$ to $1.5$.
Next, in (ref), we plot the bias, variance, and RMSE of generalized PML estimators as a function of $\tau$ under the censoring probability $\frac{1}{1+(\tau\exp(\theta_{0}^{T}x))^\beta}$ with $\beta=2$ and $\alpha=1$. Recall that increasing $\tau$ reduces overall sparsity of the data, so we should expect the bias of estimators to decrease, which is indeed the case. Moreover, the bias plots show that NLS is relatively insensitive to the level of sparsity in the data, whereas PML estimators with negative $\kappa$ are quite sensitive to small values of $\tau$, when censoring is strong and samples with small conditional means are less reliable. The variance of Poisson is the smallest among all generalized PML estimators, since $\alpha=1$. The RMSE plots again demonstrate the bias-variance trade-off, with NLS having the smallest RMSE for $\tau\leq 4$, while Poisson starts to have smaller RMSE for $\tau>4$, when the bias from sparsity becomes less dominant. These results again suggest that in the presence of excess sparsity, contrary to popular belief, the Poisson PML is not necessarily the best choice, even when it is expected to be the most efficient. The optimal choice depends on the degree of censoring.
Our next set of simulation results in (ref) reports the bias, standard deviation, and RMSE of generalized PML estimators for $\kappa \in [-1,1]$, under censoring probability $\frac{1}{1+(\tau\exp(\theta_{0}^{T}x))^\beta}$ with $\tau=1$, $\beta=2$, and $\alpha\in [0,1,2]$. We find that the bias decreases as $\kappa$ increases from $-1$ to 1, as the corresponding generalized PML estimators place increasing weights on larger samples thus reducing bias. Similarly, the standard deviation of estimators decreases then increases with $\kappa$, resulting in a choice of $\kappa$ with the smallest variance. However, this variance-minimizing $\kappa$ generally differs from the one in the uncensored case. For example, when $\alpha=2$, we know that the gamma PML ($\kappa=-1)$ is the most efficient estimator in the absence of censoring. However, in the presence of asymmetric censoring, the standard deviation is minimized at around $\kappa\approx-0.6$ for both coordinates of $\theta_0$. This observation is in line with our theoretical result on the asymptotic variance, which depends on the censoring parameters. Lastly, the RMSE plots confirm our intuition and quantitative results on the bias-variance trade-off. There is an optimal $\kappa$ with the smallest RMSE, which depends on the degree of heteroskedasticity $\alpha$ and sparsity $\tau$. For example, although Poisson PML ($\kappa=0$) is widely used practice especially when heteroskedasticity is not negligible (e.g., $\alpha=1$), in the presence of moderate censoring ($\tau=1$ and $\beta=2$), NLS ($\kappa=1$) in fact has lower RMSE and should be preferred.
We summarize the preceding three sets of simulation results in the phase transition plot in (ref). We see that, as heteroskedasticity decreases or sparsity increases, the optimal $\kappa$ decreases, and vice versa. Therefore, we advocate for a more systematic modeling of non-negative dependent variables based on assessing the relative importance of the two aspects of data. In order to determine $\alpha$ and $\tau$, one could consider performing tests along the lines of park1966estimation,mullahy1986specification or cameron1990regression. In this paper, we employ the simpler cross-validation procedure proposed in (ref) to select $\kappa$ directly, and apply it to a collection of datasets in finance to demonstrate its usefulness.
In this section, we illustrate the appeal of the proposed generalized PML approach using four distinct datasets: corporate defaults, credit ratings, corporate patents, and real estate investments. The first two applications focus on modeling default rates and credit ratings across industries defined by Moody's 35 Industry Categories\footnote{\url{https://www.moodys.com/sites/products/ProductAttachments/MCO_35%20Industry%20Categories.pdf}}. In the third application, we analyze the number of corporate patents granted, using the dataset from cohn2022count, who replicate the study by hirshleifer2012overconfident on the impact of overconfident CEOs on corporate innovation. In the fourth application, we model the number of real estate permits using the dataset provided by bekkerman2023effect. We demonstrate that, using the simple cross validation procedure based on out-of-sample MSE, one can determine the generalized PML estimator that best fits the data. Moreover, this optimal choice is highly data-dependent. In some cases, it is close to the Poisson regression estimator recommended in the literature; for other applications, it is very different from Poisson and significantly outperforms it. We provide some intuitive evidence on the different levels of heteroskedasticity in the datasets. Our results in this section therefore demonstrate that, instead of always using a particular method like the Poisson to model count-like data, one should employ a more principled approach to model such data that takes into account both heteroskedasticity and sparsity.
We first describe the datasets used in this section. We construct our datasets on corporate default and credit rating leveraging Moody's Default and Recovery Database (DRD), which provides a detailed history of firm defaults and ratings across a broad collection of industries. To construct the default outcome variable, we aggregate the total number of default events across all firms in each of the Moody's 35 Industry Categories for each month between January 2019 and December 2023. Each observation consists of the monthly default number for a particular industry category and particular month-year, which is a sparse non-negative variable. Default events as defined in the Moody's database include bankruptcy, missed payments, distressed exchange, and others. (ref) provides a histogram of the number of monthly default events. The majority of industry-year/month pairs have zero defaults, as expected. A large outlier of 35 events is recorded for February 2023 in the banking industry. These events correspond to the series of bank failures and bankruptcies in early 2023. The ratings outcome is constructed in a similar way: we compute the monthly number of firms in each of Moody's Industry Categories that are rated investment grade, i.e., Baa3 or higher. We then normalize the counts by their standard deviations in each month. Distribution of this non-negative outcome variable is also described in the histogram in (ref).
To construct the predictors for default rates, we follow the setup in duffie2007multi and consider the following macroeconomic and firm-specific factors:
The T-bill data is obtained from the Federal Reserve Bank of St. Louis, the trailing 1-year return on the S&P 500 index is obtained from CRSP. Monthly industry average of trailing 1-year stock return is obtained from CRSP/Compustat. To calculate firm's distance to default, we combine several data sources available on WRDS and apply the iterative algorithm proposed by vassalou2004default. In order to match stock returns data from CRSP with the Moody's 35 Industry Categories, we use the 3-digit NAICS code in the Moody's dataset and the first 3 digits of the 6-digit NAICS code in the CRSP data. Similarly, for the credit rating outcome, we use the following predictors constructed from CRSP/Compustat data that have been found to have predictive power hirk2022corporate:
The dataset on corporate patents is provided by cohn2022count who replicate the work of hirshleifer2012overconfident. It consists of the number of patents granted to a firm as the dependent variable, as well as firm characteristics such as stock returns, sales, institutional holdings, and CEO overconfidence. The dataset on real estate investments is constructed by bekkerman2023effect using data from Airbnb, and consists of the number of residential permits in each zipcode area in a particular month as the outcome variable, along with a vector of time-varying zipcode characteristics from the American Community Survey, such as median household income, population size, percentage of people aged between 25 to 60 with a bachelor's degree, and employment rate.
In this section, we investigate the performance of generalized PML estimators. To this end, we split the samples 80%--20% randomly into train and test sets. We estimate generalized PML estimators with varying $\kappa$ on the train set, and report the test set RMSEs averaged over 2000 random splittings of the data. When $\kappa\neq -1,0,1$, the corresponding generalized PML estimator is obtained by solving (ref) using the scipy.optimize module in Python. The average out of sample RMSEs for default rates are summarized in (ref). We observe that for the corporate default data, the RMSE is optimized at $\kappa = 0.10$, with an average value of 0.435. Therefore, for the purpose of predicting corporate defaults in this setting, the conventional choice of Poisson PML estimator with $\kappa=0$ is close to the optimal choice. However, this is not always the case for other datasets and applications, as demonstrated by the results in (ref) for the three other datasets. For credit ratings and corporate patents, estimators with $\kappa>0$ are recommended with out-of-sample RMSEs about 50% smaller than Poisson PML; for residential permits, the gamma PML is recommended with an out-of-sample RMSE that is about 90% smaller than Poisson PML.
It is therefore crucial to determine the optimal generalized PML estimator in a principled manner, e.g., using the cross validation procedure we propose. Moreover, for the default data, the RMSEs of generalized PML estimators with $\kappa \geq 0$ tend to outperform those with $\kappa < 0$. This provides evidence that the excess sparsity feature of the default data dominates efficiency considerations, leading us to favor estimators with smaller weights on samples with small conditional means. On the other hand, the results for credit rating prediction, summarized in (ref), suggest that estimators with $\kappa > 0$ have a significantly better fit of the credit rating data compared to the Poisson or other generalized PML estimators. This demonstrates that, although estimators with $\kappa > 0$ like the NLS are generally considered inefficient when modeling non-negative outcome data, they could still be preferable in some cases. Lastly, in (ref) we provide the point estimates and bootstrapped standard errors of some of the relevant variables from each dataset. We find that the optimal generalized PML estimate could differ significantly from the Poisson estimate, leading to qualitatively different economic conclusions. These results highlight the importance of appropriately choosing the regression model for non-negative sparse data instead of using a single model by default.
Our results in this section demonstrate that the out-of-sample RMSE can be used as a general cross validation metric to determine the optimal $\kappa$ when modeling data with non-negative outcomes. For example, when estimating a gravity model on trade data, we can generate a plot similar to that in (ref), and use it to assess the relative performance of the family estimators indexed by $\kappa \in \mathbb{R}$. In settings with strong heteroskedasticity, an estimator with $\kappa \leq 0$ should generally be preferred, while in settings with excess sparsity, an estimator with $\kappa \geq 0$ will be selected. We conclude by noting that other procedures to select $\kappa$ may be more appropriate depending on the modeling task. For example, we may determine the optimal $\kappa$ based on the bias-variance trade-off analysis in this paper. To determine the degree $\alpha$ of heteroskedasticity, we can employ hypothesis tests discussed in manning2001estimating. To determine the degree $\beta$ of asymmetric censoring, we can consider estimating a zero-inflated or hurdle model. We leave a systematic study of this proposal to future works.
In this paper, we show that existing approaches to modeling non-negative count and count-like data in finance can perform poorly on many datasets that jointly exhibit heteroskedasticity and sparsity.
To address this limitation, we develop a unified framework for quantifying the bias-variance trade-off arising from these two features. Within this framework, we introduce a novel class of estimators---generalized pseudo maximum likelihood (PML) estimators---that encompasses classical methods such as Poisson PML and nonlinear least squares (NLS). To guide empirical implementation, we propose a simple cross-validation procedure that allows to select the estimator best suited to the structure of their specific dataset.
We validate the effectiveness and practical appeal of our approach through extensive numerical studies. Notably, while non-linear least squares (NLS) and other generalized PML estimators with $\kappa > 0$ are often viewed as inefficient under heteroskedasticity and over-dispersion---and therefore typically avoided in practice---we find that, in empirically relevant settings with pronounced sparsity, these estimators can outperform traditional alternatives such as Poisson regression, provided the associated optimization problems are properly addressed. Our work therefore calls for a more systematic assessment of count-like data in finance and provides a principled approach to the modeling of such data in practice.