Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
87,931 characters · 21 sections · 0 citation commands
Summary We propose an asymptotic theory for distribution forecasting from the log normal chain-ladder model. The theory overcomes the difficulty of convoluting log normal variables and takes estimation error into account. The results differ from that of the over-dispersed Poisson model and from the chain-ladder based bootstrap. We embed the log normal chain-ladder model in a class of infinitely divisible distributions called the generalized log normal chain-ladder model. The asymptotic theory uses small $\sigma$ asymptotics where the dimension of the reserving triangle is kept fixed while the standard deviation is assumed to decrease. The resulting asymptotic forecast distributions follow t distributions. The theory is supported by simulations and an empirical application.
Keywords chain-ladder, infinitely divisibility, over-dispersed Poisson, bootstrap, log-normal.
Reserving in general insurance usually relies on chain-ladder-type methods. The most popular method is the traditional chain-ladder. A contender is the log-normal chain-ladder, which we study here. Both methods have proved to be valuable for point forecasting. In practice, distribution forecasting is needed too. For the standard chain-ladder there are presently three methods available. Mack (1999) has suggested a method for recursive calculation of standard errors of the forecasts, but without proposing an actual forecast distribution. The bootstrap method of England and Verrall (1999) and England (2002) is commonly used, but it does not always produce satisfactory results. Recently, Harnau and Nielsen (2017) have developed an asymptotic theory for the chain-ladder in which the idea of a over-dispersed Poisson framework is embedded in a formal model. This was done through a class of infinitely divisible distributions and a new Central Limit Theorem. An asymptotic theory provides an analytic tool for evaluating the distribution of forecast errors and building inferential procedures and specification tests for the model. Here we adapt the infinitely divisible framework of Harnau and Nielsen (2017) to the log-normal chain-ladder and present an asymptotic theory for the distribution forecasts and model evaluation. Thereby, asymptotic distribution forecasts and model evaluation tools are now available for two different models, which together cover a wide range of reserving triangles.
The data consists of a reserving triangle of aggregate amounts that have been paid with some delay in respect to portfolios of insurances. Table (ref) provides an example. The objective of reserving is to forecast liabilities that have occurred but have not yet been settled or even recorded. The reserve is an estimate of these liabilities. Thus, the problem is to forecast the lower reserving triangle and then add these forecasts up to get the reserve. The traditional chain-ladder provides a point forecast for the reserve.
The chain-ladder is maximum likelihood in a Poisson model. This is useful for estimation and point forecasting. Mart{\' i}nez Miranda, Nielsen and Nielsen (2015) have developed a theory for inference and distribution forecasting in such a Poisson model in order to analyze and forecast incidences of mesothelioma. However, this is not of much use for the reserving problem because the data is nearly always severely over-dispersed. The over-dispersion arises because each entry in the paid triangle is the aggregate amount paid out to an unknown number of claims of different severity. It is common to interpret this as a compound Poisson variable, see Beard, Pentik{\" a}inen and Pesonen (1984, \S 3.2). Compound Poisson variables are indeed over-dispersed in the sense that the variance to mean ratio is larger than unity. They are, however, difficult to analyze and even harder to convolute. England and Verrall (1999) and England (2002) developed a bootstrap to address this issue. This often works, but it is known to give unsatisfactory results in some situations. The model underlying the bootstrap is not fully described, so it is hard to show formally when the bootstrap is valid and to generalize it to other situations, including the log-normal chain-ladder.
{2pt}
{6pt}
The infinitely divisible framework of Harnau and Nielsen (2017) provides a plausible over-dispersed Poisson model and framework for distribution forecasting with the traditional chain-ladder. It utilizes that the compound Poisson distribution is infinitely divisible. If the mean of each entry in the paid triangle is large, then the skewness of compound Poisson variable is small and a Central Limit Theorem applies. Thus, keeping the dimension of the triangle fixed, while letting the mean increase, the reserving triangle is asymptotically normal with mean and variance estimated by the chain-ladder. Since the dimension is fixed we then arrive at an asymptotic theory that matches the traditional theory for analysis of variance (anova) developed by Fisher in the 1920s. If the over-dispersion is unity and therefore known as in the Poisson model of Mart{\' i}nez Miranda, Nielsen and Nielsen (2015) then inference is asymptotically $\chi^2$ and distribution forecasts are normal. When the over-dispersion is estimated as appropriate for reserving data then we arrive at inference that is asymptotically $\mathsf{F}$ and distribution forecasts that are asymptotically $\mathsf{t}$. The chain-ladder bootstrap could potentially be analyzed within this framework, but this is yet to be done.
When it comes to the log-normal model the situation is different. The log-normal model has apparently been suggested by Taylor in 1979, and then analyzed by for instance Kremer (1982), Renshaw (1989), Verrall (1991, 1994), Doray (1996) and England and Verrall (2002). The main difference to the over-dispersed Poisson model is that the mean-variance ratio is constant across the triangle in that model, while the mean-standard deviation ratio is constant in the log-normal model. Therefore the tails of distributions are expected to be different, which may matter in distribution forecasting.
Estimation is easy in the log-normal model. It is done by least squares from the log triangle. Recently, Kuang, Nielsen and Nielsen (2015) have provided exact expressions for all estimators along with a set of associated development factors. Least squares theory provides a distribution theory for the estimators and for inference. However, the reserving problem is to make forecasts of reserves that are measured on the original scale. Each entry in the original scale is log-normally distributed. While there are expressions for such log-normal distributions it is unclear how to incorporate estimation uncertainty, let alone convolute such variable to get the reserve.
The infinitely divisible theory provides a solution also for the log-normal model. Thorin (1977) showed that the log-normal distribution is infinitely divisible. First of all, this indicates that the log normal variables actually have an interpretation as compound sums of claims. Secondly, the framework of Harnau and Nielsen (2017) and their Central Limit Theorem apply, albeit with subtle differences. In the over-dispersed Poisson model the mean of each entry is taken to be large in the asymptotic theory, whereas for generalized log-normal model we will let the variance be small in the asymptotic theory. In both cases the mean-dispersion ratio is then small. In this paper we will exploit that infinitely divisible theory to provide an asymptotic theory for the log-normal distribution forecasts.
We also discuss specification tests for the log-normal model. Mis-specification can appear both in the mean and the variance of the log-normal variables. The mean could for instance have an omitted calendar effect. Thus, we study the extended chain-ladder model discussed by Zehnwirth (1994), Barnett and Zehnwirth (2000), and Kuang, Nielsen and Nielsen (2008a,b,2011). The variance could be different in subgroups of the triangle as pointed out by Hertig (1985). Barlett (1937) proposed a test for this problem. Recently, Harnau (2017) has adapted that test to the traditional chain-ladder. We extend this to the generalized log-normal model.
We illustrate the new methods using a casualty reserving triangle from XL Group (2017) as shown in Table (ref). The triangle is for US casualty and includes gross paid and reported loss and allocated loss adjustment expense in 1000 USD.
We conduct a simulation study where the data generating process matches the XL Group data in Table (ref). We find that that the asymptotic results give good approximations in finite samples. The asymptotic will work even better if the mean-dispersion ratio is larger. The generalized log-normal model is also compared with the over-dispersed Poisson model and the England (2002) bootstrap. The bootstrap is found not to work very well by an order of magnitude for this log-normal data generating process. The over-dispersed Poisson model works better although it is dominated by the generalized log-normal model.
In \S(ref) we review the well known log-normal models for reserving. In \S(ref) we set up the asymptotic generalized log-normal model based on the infinitely divisible framework. We check that the log-normal model is embedded in this class and show that the results for inference in the log-normal model caries over to the generalized log-normal model. We also derive distribution forecasts. We apply the results to the XL Group data in \S(ref), while \S(ref) provides the simulation study. Finally, we discuss directions for future research in \S(ref). All proofs of theorems are provided in an Appendix.
A competitor to the chain-ladder is the log-normal model. In this model the log of the data is normal so that parameters can be estimated by ordinary least squares. We review the log-normal model by describing the structure of the data, the model, statistical analysis, point forecasts and extension by a calendar effect.
Consider a standard incremental insurance run-off triangle of dimension $k$. Each entry is denoted $Y_{ij}$ so that $i$ is the origin year, which can be accident year, policy year or year of account, while $j$ is the development year. Collectively we have data $\mathbf{Y}=\{ Y_{ij} , \forall i,j \in \mathcal{I} \}$, where $\mathcal{I}$ is the triangular index set
Let $n=k(k+1)/2$ be the number of observations in the triangle $\mathcal{I}$. One could allow more general index sets, see Kuang, Nielsen and Nielsen (2008a), for instance to allow for situations where some accidents are fully run-off or only recent calendar years are available. We are interested in forecasting the lower triangle with index set
In the log-normal model the log claims have expectation given by the linear predictor
The predictor $\mu_{ij}$ is composed of a an accident year effect $\alpha_i$, a development year effect $\beta_j$ and an overall level $\delta$. The model is then defined as follows.
The parametrisation presented in ((ref)) does not identify the distribution. It is common to identify the parameters by setting, for instance, $\delta=0$ and $\sum_{j=1}^k \beta_j = 0$. Such an ad hoc identification makes it difficult to extrapolate the model beyond the square composed of the upper triangle $\mathcal{I}$ and the lower triangle $\mathcal{J}$ and it is not amenable to the subsequent asymptotic analysis. Thus, we switch to the canonical parametrisation of Kuang, Nielsen and Nielsen (2009, 2015) so that the model becomes a regular exponential family with freely varying parameters. The canonical parameter is
where $\Delta \alpha_{i}=\alpha _{i}-\alpha _{i-1}$ is the relative accident year effect and $\Delta \beta _{j}=\beta _{j}-\beta _{j-1}$ is the relative development year effect, while $\mu _{11}$ is the overall level. The length of $\xi$ is denoted $p$, which is $p=2k-1$ with the chain-ladder structure. We can then write
with the convention that empty sums are zero and $X_{ij}\in\mathbb{R}^p$ is the design vector
where the indicator function $1_{(m\le i)}$ is unity if $m\le i$ and zero otherwise.
The log observations $y_{ij}=\log Y_{ij}$ have a normal log likelihood given by
Stacking the observations $y_{ij}=\log Y_{ij}$ and the row vectors $X_{ij}'$ then gives an observation vector $y$ and a design matrix $X$ and a model equation of the form
The least squares estimator for $\xi$ and the residuals are then
while the variance $\omega^2$ is estimated by
Kuang, Nielsen and Nielsen (2015) derive explicit expressions for each coordinate of the canonical parameter and they provide an interpretation in terms of so-called geometric development factors.
Standard least squares theory provides a distribution theory for the estimators, see for instance Hendry and Nielsen (2007), so that
Individual components of $\hat\xi$ will also be normal. Standardizing those components and replacing $\omega^2$ by the estimate $s^2$ gives the $\mathsf{t}$-statistic, which is $\mathsf{t}_{n-p}$ distributed.
We may be interested in testing linear restrictions on $\xi$. This can be done using $\mathsf{F}$-tests. For instance, the hypothesis that all $\Delta\alpha$ parameters are zero would indicate that the policy year effect is constant over time. Such restrictions can be formulated as $\xi=H \zeta$ for some known matrix $H\in\mathbb{R}^{p\times p_H}$ and a parameter vector $\zeta\in \mathbb{R}^{p_H}$. In the example of zero $\Delta\alpha$'s the $H$ matrix would select the remaining parameters, the $\mu_{11}$ and the $\Delta\beta_j$s. We then get a restricted design matrix $X_H=XH$ and a model equation of the form $y=X_H\zeta + \varepsilon$. We then get estimators $$ \hat\zeta = (X_H'X_H)^{-1} X_H'y, \qquad s_H^2 = \frac{RSS_H}{n-p_H} , $$ where the residual sum of squares $RSS_H = \sum_{i,j\in\mathcal{I}} \hat\varepsilon_{H,ij}^2 $ is formed from the residuals $ \hat\varepsilon_{H,ij} = y_{ij} - X_{H,ij}'\hat\zeta $ as before. The hypothesis can be tested by $\mathsf{F}$-statistic
We may also be interested in affine restrictions. For instance, the hypothesis that all $\Delta\alpha$ parameters are known corresponds the hypothesis of known values of relative ultimates. This may be of interest in an Bornhuetter-Ferguson context, see Margraf and Nielsen (2018). This is analyzed by restricted least squares which also leads to $\mathsf{t}$ and $\mathsf{F}$ statistics.
In practice we will want to forecast the variables $Y_{ij}$ on the original scale. Since $y_{ij}$ is $\mathsf{N}(\mu_{ij},\omega^2)$ then $Y_{ij}=\exp (y_{ij} )$ is log-normally distributed with mean $\exp (\mu_{ij}+\omega^2/2)$. Thus, the point forecast for the lower triangle $\mathcal{J}$, as well as the predictor for the upper triangle $\mathcal{I}$, can be formed as
We will also be interested in distribution forecasting. However, the log-normal model has the drawback that it is a non-trivial problem to characterize the joint distribution of the variables on the original scale. Renshaw (1989) provides expressions for the covariance matrix of the variables on the original scale, but a further non-trivial step would be needed to characterize the joint distribution. Once it comes to distribution forecasting we would also need to take the estimation error into account. This does not make the problem easier. We will circumvent these issues by exploiting the infinitely divisible setup of Harnau and Nielsen (2017).
It is common to extend the chain-ladder parametrization with a calendar effect, so that linear predictor in ((ref)) becomes
where $i+j-1$ is the calendar year corresponding to accident year $i$ and development year $j$. This model has been suggesting in insurance by Zehnwirth (1994). Similar models have been used in a variety of displines under the name of age-period-cohort models, where age, period and cohort are our development, calendar and policy year. The model has an identification problems. The canonical parameter solution of Kuang, Nielsen and Nielsen (2008a) is to write $ \mu _{ij,apc} = X_{ij,apc}' \xi_{apc} $ where, with $h(i,s)=\max (i-s+1,0)$, we have
The dimension of these vectors is $p_{apc}=3k-3$.
This model can be analyzed by the same methods as above. Stack the design vectors $X_{ij,apc}'$ to a design matrix $X_{apc}$ and regress $y$ on $X_{apc}$ to get an estimator $\xi_{apc}$ of the form ((ref)) along with a residual sum of squares $RSS_{apc}$ and a variance estimator $s^2_{apc}$ The significance of the calendar effect can be tested using an $\mathsf{F}$-statistic as in ((ref)), where $\xi$ and $p$ now correspond to the extended model, while $\zeta$ and $p_H$ correspond to the chain-ladder specification.
When it comes to forecasting it is necessary to extrapolate the calendar effect. This has to be done with some care due to identication problem, see Kuang, Nielsen and Nielsen (2008b, 2011).
The log-normal distribution is infinitely divisible as shown by Thorin (1977). We can therefore formulate a class of infinitely divisible distributions encompassing the log-normal. We will refer to this class of distributions as the generalized log-normal chain-ladder model. In the analysis we exploit the setup of Harnau and Nielsen (2017) to provide distribution forecasts for the generalized log-normal model.
The infinitely divisible setup of Harnau and Nielsen (2017, \S 3.7) encompasses the log-normal model. Recall that a distribution $D$ is infinitely divisible, if for any $m\in\mathbb{N}$, there are independent, identically distributed random variables $X_1, \dots ,X_m$ such that $\sum_{\ell=1}^m X_\ell$ has distribution $D$. The log-normal distribution is infinitely divisible as shown by Thorin (1977). This matches the fact that the paid amounts are aggregates of number of payments. In our data analysis we neither know the number nor the severities of the payments. Due to the infinite divisibility the log-normal distribution can therefore be a good choice for modelling aggregate payments.
We will need two assumptions. The first assumption is about a general infinite divisible setup. The second assumption gives more specific details on the log-normal setup.
We have the following Central Limit Theorem for non-negative, infinitely divisible distributions with vanishing skewness. This is different from the standard Lindeberg-L{\'e}vy Central Limit Theorem for averages of independent, identically distributed variables, but proved in a similar fashion by analyzing characteristic function and exploiting the L{\'e}vy-Kintchine formula for infinitely divisible distributions.
We need some more specific assumptions for the log-normal setup. When describing the predictor we write $\mu_{ij}=X_{ij}'\xi$ to indicate that any linear structure is allowed as long as $\xi$ is freely varying when estimating in the statistical model. This could be the chain-ladder structure as in ((ref)), ((ref)) or an extended chain-ladder model with a calendar effect.
We check that the log-normal model set out in Assumption (ref) is indeed of the generalized log-normal model.
A first consequence of the generalized log-normal model is that Theorem (ref) provides an asymptotic theory for the claims on the original scale. We now check that we have a normal theory for the log claims. The proof applies the delta method. Theorem (ref) is useful in deriving the inference in Theorem (ref) and estimation error for forecasts in Theorem Theorem (ref) in later sections.
We will need to reformulate the Central Limit Theorem (ref) slightly. The issue is that the generalized log-normal model leaves the variance of the variable unspecified in finite sample, so that the Central Limit Theorem is difficult to manipulate directly. Theorem (ref) is useful in deriving the process error for forecasts in Theorem (ref) later.
We check that the inferential results for the log-normal model, described in \S(ref), carry over to the generalized log-normal model. First, we consider the asymptotic distribution of estimators and then the properties of $\mathsf{F}$-statistics for inference.
We can derive inference for of the estimator $\hat\xi$ using asymptotic $\mathsf{t}$ distribution. The proof follows Theorem (ref) and the Continuous Mapping Theorem.
We can also make inference using asymptotic $\mathsf{F}$-statistics, mirroring the $\mathsf{F}$-statistic ((ref)) from the classical normal model. The proof is similar to Theorem 4 of Harnau and Nielsen (2017).
The aim is to predict a sum of elements in the lower triangle, that could be the overall sum, which is the total reserve; or it could be row sums or diagonal sums giving a cash flow. We denote such sums by ${Y}_{\mathcal{A}} = \sum_{(i,j)\in\mathcal{A}} {Y}_{ij}$ for some subset $\mathcal{A}\in\mathcal{J}$. The point forecasts for a single entry are $\hat Y_{ij}=\exp (X_{ij}'\hat\xi + s^2/2)$ as given in ((ref)), while the overall point forecast is
To find the forecast error we expand
which we will sum over $\mathcal{A}$. This is sometimes called the forecast taxonomy. This expansion gives some insight into the asymptotic forecast distribution, although the detailed proof will be left to the appendix. The first term in ((ref)) is the process error. When extending Theorem (ref) to the lower triangle $\mathcal{J}$ we will get
where
The second term in ((ref)) is the estimation error for the canonical parameter $\xi$. From Theorem (ref) we will be able to derive
where
The third term in ((ref)) vanishes asymptotically. We will estimate $\omega^2$ by $s^2$, which turns the asymptotic normal distributions into $\mathsf{t}$-distribution. The process error and the estimation error are asymptotically independent as they are based on independent variables for the upper and lower triangle, $\mathcal{J}$ and $\mathcal{I}$. We can describe the asymptotic forecast error as follows.
Thus, the distribution forecast is
Specification tests for the log-normal model can be carried out by allowing a richer structure for the predictor or for the variance. We have already seen how the generalized log-normal chain-ladder model can be tested against the extended chain-ladder model using an asymptotic $\mathsf{F}$-test. We can test whether the variance is constant across the upper triangle by adopting the Bartlett (1937) test. Recently, Harnau (2017) has shown how to do model specification tests for the over-dispersed Poisson model. Here we will adapt the Bartlett test to the log-normal chain-ladder. It should be noted that one can of course also allow a richer structure for the predictor and the variance simultaneously following the principles outlined here.
Suppose the triangle $\mathcal{I}$ can be divided into two or more groups as indicated in Figure (ref). Thus, the index set $\mathcal{I}$ is divided into disjoint sets $\mathcal{I}_\ell$ for $\ell=1,\dots ,m$. We then set up a log-normal chain-ladder seperately for each group. Note that the full canonical parameter vector $\xi$ may not be identified on the subsets. As we will only be interested in the fit of the models we can ad hoc identify $\xi$ by dropping sufficiently many columns of the design matrix $X$. This gives us a parameter $\xi_\ell$ and a design vector $X_{ij\ell}$ for each subset $\mathcal{I}_\ell$ and a predictor $\mu_{ij\ell}=X_{ij\ell}'\xi_\ell$. Thus the model for each group is that $y_{ij\ell}$ is $\mathsf{N}(\mu_{ij\ell},\omega_\ell^2)$. Let $p_\ell$ denote the dimension of these vectors, while $n_\ell$ is the number of elements in $\mathcal{I}_\ell$ giving the degrees of freedom $df_\ell=n_\ell-p_\ell$.
When fitting the log-normal chain-ladder seperately to each group we get estimators $\hat\xi_\ell$ and predictors $\hat\mu_{ij\ell}=X_{ij\ell}'\hat\xi_\ell$. From this we can compute the residual sum of squares and variance estimators as
If there are only two subsets then we have two choices of tests available. The first test is a simple $\mathsf{F}$-test for the hypothesis that $\omega_1=\omega_2$. In the log-normal model this is
In the generalized log-normal the $\mathsf{F}$-distribution can be shown to be valid asymptotically. Harnau (2017) has proved this for the over-dispersed Poisson model using an infinitely divisible setup. That proof extends to the generalized log-normal setup following the ideas of the proofs of the above theorems. We can then construct a two sided test. Choosing a 5% level this test rejects when $F^\omega$ is either smaller than the 2.5% quantile or larger than the 97.5% quantile of the $\mathsf{F}_{n_2-p_2,n_1-p_1}$-distribution.
The second test is known as Bartlett's test and applies to any number of groups. Thus, suppose we have $m$ groups and want to test $\omega_1=\cdots =\omega_m$. In the exact log-normal case then $s_1^2,\dots ,s_m^2$ are independent scaled $\chi^2$ variables. Bartlett found the likelihood for this $\chi^2$ model. Under the hypothesis the common variance is estimated by
while the likelihood ratio test statistic for the hypothesis is
The exact distribution of the likelihood ratio test statistic depends on the degrees of freedom of the groups, but not on their ordering. No analytic expression is known. However, Bartlett showed that this distribution is very well approximated by a scaled $\chi^2$-distribution. That is
The factor $C$ is known as the Bartlett correction factor. Formally, the approximation is a second order expansion which is valid when the small group is large, so that $\min_\ell df_\ell$ is large. However, the approximation works exceptionelly well in very small samples; see the simulations by Harnau (2017). Once again the Bartlett test ((ref)) will be applicable in the generalized log-normal model, which can be proved by following the proof of Harnau (2017).
In practice, we can fit seperate log-normal models to each group, that is $y_{ij\ell}$ is assumed $\mathsf{N}(\mu_{ij\ell},\omega_\ell^2)$. If the Bartlett test does not reject the hypothesis of common variance we then arrive at a model where $y_{ij\ell}$ is assumed $\mathsf{N}(\mu_{ij\ell},\omega^2)$. This model can be estimated by a single regression where the design matrix is block diagonal, $X^m=\mathrm{diag}(X_1,X_2,\dots ,X_m)$ of dimension $p_\cdot=\sum_{\ell=1}^m p_\ell$. We then compare the models with design matrices $X^m$ and the original $X$ of the maintained model through an $\mathsf{F}$-test.
We apply the theory to the insurance run-off triangle shown in Table (ref). All R (2017) code is given in the supplementary material. We use the R packages apc, see Nielsen (2015) and ChainLadder, see Gesmann et. al. (2015). First, we apply the proposed inference and estimation procedures to the data. This is followed first by distribution forecast and then by an analysis of the model specification.
We apply the log-normal model to the data and consider three nested parametrizations:
Table (ref) shows an analysis of variance. This conforms with the exact distribution theory in \S(ref) and the asymptotic distribution theory in Theorems (ref), (ref) in \S(ref).
First, we test the chain-ladder model (ac for age-cohort) against the extended chain-ladder model (apc for age-period-cohort) with $\mathsf{p}=0.984$. The chain-ladder hypothesis is clearly not rejected at a conventional $5\%$ test level. Next, we test the further restriction (ad for age-drift) that the row differences are constant, that is $\Delta^2\alpha_i=0$. We get $\mathsf{p}=0.000$ and $\mathsf{p}=0.000$ when testing against the apc and ac models respectively. This suggests that a further reduction of the model is not supported. In summary, the analysis of variance indicates that it is adequate to proceed with a chain-ladder specification and thereby ignore calendar effects.
Table (ref) shows the estimated parameters for the log-normal model with chain-ladder structure (ac). We report standard errors $se_t$ following Theorem (ref). They are the same for $\Delta\alpha$ and $\Delta\beta$ due to symetry of $(X'X)^{-1}$ at the diagonal. These follow a $\mathsf{t}$-distribution with $n-p=171$ degrees of freedom, since the triangle has dimension $k=20$ and $n=k(k+1)/2=210$ and $p=2k-1=39$. The corresponding two-sided 95% critical values are 1.97. We also report the degrees of freedom corrected estimate, $s^2$, for $\omega^2$. We see that many of the development year effects $\Delta\beta$, in particular $\Delta\beta_2$, are significant. The first few development year effects are positive, which matches the increases seen in first few columns of the data in Table (ref). At the same time many the accident year effects $\Delta\alpha$ are not individually significant, although they are jointly significant as seen in Table (ref). The signs of the $\Delta\alpha$'s match the relative increase or decrease of the amounts seen in the rows of Table (ref).
In Appendix (ref) we present a further Table (ref) with estimates. These are the estimated parameters for the log-normal model with an extended chain-ladder structure (apc) as in \S(ref). These will be used for the simulation study. The $\Delta^2\gamma$-coefficients measure the calendar effect and are restricted to zero in the chain-ladder model.
Table (ref) shows forecasts of reserves for the US casualty data in different accident years, i.e.\ the row sums in the lower triangle $\mathcal{J}$. We report results from the generalized log-normal chain-ladder model (GLN), the over-dispersed Poisson chain-ladder (ODP) and England (2002) bootstrap (BS). For each method, we present a point forecast of the reserve, the standard error over point forecast ($se$/Res) and the 1 in 200 over point forecast values (99.5%/Res).
For the generalized log-normal chain-ladder model we use the asymptotic distribution forecast in ((ref)). For the over-dispersed Poisson model we use the asymptotic distribution forecasts from Harnau and Nielsen (2017, equation 11). For the bootstrap we use the ChainLadder package by Gesmann et al (2005), based on the method described in England (2002). We apply $10^5$ bootstrap draws using the gamma option.
Table (ref) shows that the over-dispersed Poisson forecasts are similar to the bootstrap. Their point forecasts are smaller than that of the generalized log-normal model. This is in part due to the additional factor $\exp(s^2/2)=\exp(0.169/2)=1.088$ in the generalized log-normal point forecast. The difference seems large compared to the authors' experience with other data. It is possibly due to the relatively large dimension of the triangle, so that there are more degrees of freedom to pick up differences between the over-dispersed Poisson and the generalized log-normal models.
The standard error and 99.5% quantiles over reserve ratios are generally lower and less variable for the generalized log-normal chain-ladder model. This is especially pronounced for early accident years and the latest accident year.
{-7pt}
{6pt}
Figure (ref) shows the trends of the reserve and standard error and 99.5% quantile over reserve ratios for the three methods. The point forecast trends are similar for models, showing an increasing trend with accident year as expected. The ratios are seen to be flatter for the generalized log-normal model. This is related to the assumption of the generalized log-normal chain-ladder model that standard deviation to mean ratio is constant across the entries, while the variance to mean ratio is assumed constant for the over-dispersed Poisson model and the bootstrap.
To check the robustness of the model we apply the distribution forecasting recursively. Thus, we apply the distribution forecast to subsets of the triangle.
In this way, Table (ref) shows standard error and 99.5% over reserve ratios. It has 9 panels, where the rows are for the asymptotic generalized log-normal model, the over-dispersed Poisson model and the bootstrap, respectively. In the first column we show the ratios for the last 5 accident years based on the full triangle. These numbers are the same as those in Table (ref). In the second column we omit the last diagonal of the data triangle to get a $k-1=19$ dimensional triangle. We then forecast the last 5 accident years relative to that triangle. In the third column we omit the last two diagonals of the data triangle to get a $k-2=18$ dimensional triangle.
We see that the generalized log-normal forecasts are stable for all years. The over-dispersed Poisson and bootstrap forecasts are less stable in the latest accident year. This is possibly because of instability in the corners of the data triangle shown in Table (ref), that may be dampened when taking logs. Alternatively, it could be attributed to a better fit of the log-normal model across the entire triangle. We will explore the model specification using formal tests in the next section.
We now apply the specification test outlined in \S(ref) for the log-normal model and in Harnau (2017) for the over-dispersed Poisson model. For the tests we split the data triangle of Table (ref) as outlined in Figure (ref):
For each split we estimate a chain-ladder structure separately for each sub-group. We then compute the Bartlett test statistic $LR^\omega/C$ from ((ref)) for a common variance across groups. Given a common variance we also compute an $F$-statistic for common chain-ladder structure in the mean.
For each of the generalized log-normal and over-dispersed Poisson model we are conducting 6 tests. When chosing the size of each individual test, that is the probability of falsely rejecting the hypothesis, we would have to keep in mind the overall size of rejecting any of the hypotheses. If the test statistics were independent and the individual tests were conducted at level $p$ the overall size would be $1-(1-p)^6\approx 6p$ by binomial expansion, see also Hendry and Nielsen (2007, \S 9.5). Thus, if the individual tests are conducted at a 1% level we would expect the overall size to be about 5%. At present we have no theory for a more formal calculation of the joint size of the tests.
Starting with the log-normal model we see that there is only moderate evidence against model. The worst cases are that variance differs across the (a) split and the chain-ladder structure differs across the (b) split. In contrast, the over-dispersed Poisson model is rejected by all 6 tests.
In Theorems (ref) and (ref) we presented asymptotic results for inference and distribution forecasting. We now apply simulation to investigate the quality of these asymptotic approximations.
We assess the finite sample performance of the $F$-tests proposed in Theorem (ref) and applied in Table (ref). We simulate under the null hypothesis of a chain-ladder specification, ac, as well as under the alternative hypothesis of an extended chain-ladder specification, apc. We choose the distribution to be log-normal so, to be specific, we actually illustrate the well-known exact distribution theory for regression analysis. Theorem (ref) also applies for infinitely divisible distributions that are not log-normal but satisfy Assumptions (ref) and (ref). Such infinitely divisible distributions are, however, not easily generated. The real point of the simulations is therefore to illustrate the small variance asymptotics in Theorem (ref) by showing that power increases with shrinking variance.
The data generating processes are constructed from the US casualty data as follows. We consider a $k=20$ dimensional triangle. We assume that the variables $Y_{ij}$ in the upper triangle $\mathcal{I}$ are independent log-normal distributed, so that $y_{ij}=\log(Y_{ij})$ is normal with mean $\mu_{ij}$ and variance $\sigma^2$. Under the null hypothesis of a chain-ladder specification, $\mathsf{H}_{ac}$, then $\mu_{ij}$ is defined from ((ref)) where the parameters $\mu_{ij}$ are chosen to match those of Table (ref). We also choose $\sigma^2$ to match the estimate $s^2$ from Table (ref), but multiplied by a factor $v^2$ where $v$ is chosen as $2, 1, 1/2$ to capture the small-variance asymptotics. Under the alternative, we apply the extended chain-ladder specification $\mathsf{H}_{apc}$ where the parameters are chosen to match those of Table (ref). In all cases we draw $10^5$ repetitions.
We note that the $\mathsf{F}(18,153)$-distribution is exact under the null hypothesis, since we are operating on the log-scale and simulate normal variables so that standard regression theory applies. Indeed, Table (ref) shows that simulated size (type I error) is correct apart from Monte Carlo standard error. We check this for at the 1%, 5% and 10% level for $v=2, 1, 1/2$.
Under the alternative we simulate power (unity minus type II error). The exact distribution is a non-central $\mathsf{F}$-distribution. The simulations show that the power increases for shrinking variance $v^2\omega^2$ and for increasing level (type I error) of the test.
We can also illustrate the increasing power with shrinking variance through the following analytic example. Suppose we consider variables $Z_1,\dots ,Z_n$ that are independent $\mathsf{N}(\mu,\omega^2)$-distributed. Then the parameters are estimated by $\hat\mu=\bar{Z}$ and $s^2=(n-1)^{-1}\sum_{i=1}^n (Z_i - \bar{Z})^2$. The $\mathsf{t}$-statistic for $\mu=0$ has the expansion $$ \frac{\hat\mu - 0}{\sqrt{s^2/(n-1)}} = \frac{\hat\mu -\mu}{\sqrt{s^2/(n-1)}} + \frac{ \mu - 0}{\sqrt{s^2/(n-1)}} . $$ The first term is $\mathsf{t}$ distributed with $(n-1)$ degrees of freedom regardless of the value of $\mu$. The second term is zero under the hypothesis $\mu=0$. Under the alternative $\mu\ne 0$ the second term is non-zero and measures non-centrality so that the overall $\mathsf{t}$-statistic is non-central $\mathsf{t}$. In standard asymptotic theory $n$ is large so that for fixed $\mu$, $\omega$ then $s^2$ is consistent for $\omega^2$ and the second term is close to $\mu/\surd{\omega^2 / (n-1)}=(\mu/\omega)\surd (n-1)$. Due to the $(n-1)$-factor the non-centrality diverges, so that the power increases to unity and the test is consistent. In the small variance asymptotics $\omega^2$ shrinks to zero while $n$ is fixed. Then $s^2$ vanishes, see Theorem (ref), and the non-centrality diverges in a similar way even though $n$ is fixed.
We assess the finite sample performance of the asymptotic distribution forecasts proposed in Theorem (ref) and applied in Table (ref). These asymptotic distribution forecasts are compared to the over-dispersed Poisson forecast of Harnau and Nielsen (2017) and the bootstrap of England and Verrall (1999) and England (2002). Two different log-normal chain-ladder data generating processes are used. First, we apply the estimates from the US casualty data so that the parameters are chosen to match those of Table (ref). As before the variance $\omega^2$ is multiplied by a factor $v^2$ where $v=2,1,1/2$. We have seen that the over-dispersed Poisson model is poor for this data set and we will expect the generalized log-normal distribution forecasts to be superior. Secondly, we obtain similar estimates for the Taylor and Ashe (1983) data, see also Harnau and Nielsen (2017, Table 1). For those data the generalized log-normal model and the over-dispersed Poisson model provide equally good fits so that the different distributions forecasts should be more similar in performance.
We first compare the asymptotic distribution forecast from Theorem (ref) with the exact forecast distribution. This is done by simulating log-normal chain-ladder for both the upper and the lower triangles, $\mathcal{I}$ and $\mathcal{J}$. The true forecast error distribution is then based on $Y_\mathcal{A}-\tilde{Y}_\mathcal{A}$, where $Y_\mathcal{A}$ is computed from the simulated lower triangle $\mathcal{J}$ while $\tilde{Y}_\mathcal{A}$ is the log-normal point forecast computed from the upper triangle data $\mathcal{I}$. We compute the true forecast error $Y_\mathcal{A}-\tilde{Y}_\mathcal{A}$ for each simulation draw and report mean, standard error and quantiles of the draws. This is done for the entire reserve, so that $\mathcal{A}=\mathcal{J}$. The asymptotic theory in Theorem (ref) provides a $\mathsf{t}$-approximation, so that for each draw of the upper triangle $\mathcal{I}$, we also compute mean, standard error and quantiles from the $\mathsf{t}$-approximations and report averages over the draws.
The first panel of Table (ref) compares the simulated actual forecast distribution, $true^{GLN}$, with the simulated $\mathsf{t}$-approximations, $t^{GLN}$. We see that with shrinking variance factor $v$ then the overall forecast distribution becomes less variable and the $\mathsf{t}$-approximation becomes relatively better. The $\mathsf{t}$-approximation is symmetric and does not fully capture the asymmetry of the actual distribution. We note that the performance of the $\mathsf{t}$-approximation is better in the upper tail than the lower tail, which is beneficial when we are interested in 99.5% value at risk.
The second panel of Table (ref) shows the performance of the traditional chain-ladder. Since the data are log-normal we expect the chain-ladder to perform poorly. We apply the asymptotic theory of Harnau and Nielsen (2017) and the bootstrap of England and Verrall (1999) and England (2002) as implemented by Gesmann et al.\ (2015) The results are generated as before with the difference that the point forecasts are based on the traditional chain-ladder, while the data remain log-normal. The actual forecast errors, $true^{ODP}$ are similar to the previous actual errors $true^{GLN}$, particular in the right tail of the distribution. The asymptotic distribution approximation, $t^{ODP}$, and the bootstrap approximation, $BS$, do not provide the same quality of approximations as $t^{GLN}$ did for $true^{GLN}$. For large $v=2$ the bootstrap is very poor, possibly because of resampling of large residuals arising from the mis-specification.
We also simulate the root mean square forecast error for the three methods. For the log-normal asymptotic distribution approximation this is computed as follows. We first find mean, standard deviation and quantiles of the infeasible reserve based on the draws of the lower triangle $\mathcal{J}$. This is the true forecast distribution. For each draw of the upper triangle $\mathcal{I}$ we then compute mean, standard deviation and quantiles of the asymptotic distribution forecast ((ref)) and subtract the mean, standard deviation and quantiles, respectively, of the true forecast distribution. We square, take average across the draws of the upper triangle $\mathcal{I}$, and then the take the square root. Similar calculations are done for the over-dispersed approximation and the bootstrap.
The third panel of Table (ref) shows the root mean square forecast errors. We see that the generalized log-normal distribution approximation is superior in all cases and that the bootstrap can be very poor if $v$ is not small.
In Table (ref) we repeat the simulation exercise for the Taylor and Ashe (1983) data. For these data we repeated the empirical exercise of \S(ref), although we do not report the results here. We found that the generalized log-normal chain-ladder and the over-dispersed chain-ladder appear to give equally good fit, so that we will expect less difference between the methods in this case. We suspect that this arises because of two features in the data. The Taylor and Ashe triangle has a smaller dimension of $k=10$ and there is less difference between the accident year parameters, see also Harnau and Nielsen (2017, Table 2). As before we simulate a log-normal distribution with parameters equal to the estimates from the data.
Table (ref) shows that the three methods perform similarly. In this discussion we focus on the root mean square error for the 99.5% quantile which is perhaps of most practical interest. For large $v=2$ and $v=1$ the over-dispersed Poisson method actually dominates the generalized log-normal model even though the data are generated to be log-normal. For a smaller $v=1/2$ the asymptotic approximation for the generalized log-normal beats that of the over-dispersed model slightly. However, the bootstrap appears to be best for $v=1$ and $v=1/2$.
We have presented a new method for distribution forecasting of general insurance reserves in terms of the generalized log-normal model. The forecasts are done under the asymptotic framework which allows users to draw inferences and make model selections easily. This gives an alternative to the traditional chain-ladder where we have the commonly used bootstrap method developed by England and Verrall (1999) and England (2002) along with the recent asymptotic theory of Harnau and Nielsen (2017).
The actuary will have to choose whether the traditional or the normal chain ladder or a third method should be used for a given reserving triangle. In some situations the normal chain ladder will be better than the traditional chain latter as shown in our empirical data analysis and simulation study. In addition, we have considered a number of London market datasets. We compared the standard error over mean forecast trends by year of account with the actuaries' selected volatilities and found that the generalized log-normal trends are more in line with the actuaries selected trends than the over-dispersed Poisson model.
The generalized log-normal model distribution forecasts developed here could also improve the actuarial process for a corporation. The log-normal is also often used in simulating attritional reserve risk for capital modelling. At present this is some times combined with the bootstrap method for the traditional chain ladder. This can result in inconsistencies often between reserving and capital modelling.
A limitation of the log-normal model is that it only fits positive incremental values, while in real life some values can be negative due to reinsurance recoveries, salvage or other data issues such as mis-allocation between classes of business or currencies. In these cases judgements are required and further research must look at how to provide statistical tools to overcome such a limitation.
There is also scope to develop a more advanced model selection process than the model specification tests discussed here. This will give actuaries a statistical basis to select one model over another rather than just eye-balling a distribution fit on a graph. Testing constancy of the dispersion as presented here for the log normal chain ladder and by Harnau (2017) for the traditional chain ladder is a beginning of that research agenda.
The bootstrap method has become popular in recent decades. This is because it usually produces distributions that appear reasonable and it is a simulation based technique which is favoured by many actuaries. A deeper understanding of the bootstrap method can be developed so that it allows model selections and extensions to generate reserve forecasts under other distributions than the over-dispersed Poisson.