The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
65,559 characters
Hidden Group Time Profiles: Heterogeneous Drawdown Behaviours in Retirement
\title{\vspace{-10mm}
Hidden Group Time Profiles: \\Heterogeneous Drawdown Behaviours in Retirement
\vspace{-4mm}}
\author{Igor Balnozan
\footnote{University of New South Wales, UNSW Sydney, NSW 2052, Australia.}
\newcounter{UNSW}
\setcounter{UNSW}{\value{footnote}}
\footnote{Correspondence: [email removed] / West Lobby Level 4,
UNSW Business School Building, UNSW Sydney, NSW 2052, Australia.}, Denzil G. Fiebig
\footnotemark[\value{UNSW}]
,
Anthony Asher
\footnotemark[\value{UNSW}]
, Robert Kohn
\footnotemark[\value{UNSW}]
, Scott A. Sisson
\footnotemark[\value{UNSW}]
}
\date{Version: 19 March 2025\vspace{-4mm}}
\maketitle
\begin{abstract}
\begin{onehalfspace}
This article investigates retirement decumulation behaviours using the Grouped Fixed-Effects (GFE) estimator applied to Australian panel data on drawdowns from phased withdrawal retirement income products. Behaviours exhibited by the distinct latent groups identified suggest that retirees may adopt simple heuristics determining how they draw down their accumulated wealth. Two extensions to the original GFE methodology are proposed: a latent group label-matching procedure which broadens bootstrap inference to include the time profile estimates, and a modified estimation procedure for models with time-invariant additive fixed effects estimated using unbalanced data.
\end{onehalfspace}
\end{abstract}
\vfill
\begin{singlespace}
{\footnotesize
\textbf{Keywords}
\vspace{-4mm}
Panel data, discrete heterogeneity, microeconomics, retirement
\textbf{JEL codes}
\vspace{-4mm}
C51, D14, G40
\textbf{Acknowledgements}
\vspace{-4mm}
Balnozan is grateful for the support provided by the Commonwealth Government of Australia through the provision of an Australian Government Research Training Program Scholarship, and to State Super for provision of the State Super Academic Scholarship.
Balnozan, Kohn and Sisson are partially supported by the Australian Research Council through the Australian Centre of Excellence in Mathematical and Statistical Frontiers (ACEMS; CE140100049).
We
thank Plan For Life, Actuaries \& Researchers, who collected, cleaned and allowed
us to analyse the data used in this research; this data capture forms
one part of a broader survey into retirement incomes, commissioned
by the Institute of Actuaries of Australia.
This research includes computations using the computational cluster
Katana supported by Research Technology Services at UNSW Sydney. Katana DOI: 10.26190/669x-a286
\textbf{Conflict of interest statement}
\vspace{-4mm}
The authors have no conflict of interest to declare.
}
\end{singlespace}
\pagebreak
\doublespacing
\section{\label{sec:Introduction}Introduction}
The importance and prevalence of preference heterogeneity is a pervasive
feature of microeconomic modelling. Such heterogeneity
is usually unobserved, remaining unexplained after controlling
for the observable characteristics of individuals. Where the effects of unobservables have clear economic interpretations, methods of identifying latent groups of individuals that share common behaviours can provide insights into economic phenomena.
There is a need to improve understanding of the behaviours of retirees who fund their consumption by drawing down savings accumulated in personal Defined
Contribution accounts. Such understanding can inform appropriate financial advice and product development.
Existing theoretical work by \citet{bateman2008choices}
gives reason to expect distinct behavioural groups, where group behaviours correspond to strategies that individuals employ when accessing their retirement savings using phased withdrawal products.
This article investigates drawdown behaviours in phased withdrawal
retirement income products by studying latent time profiles estimated using the Grouped Fixed-Effects (GFE) estimator of \citet{bonhomme2015grouped}. Of primary interest is testing for the presence of multiple behavioural groups in the data, and characterising any groups found. Furthermore, in this application, using the GFE estimator also allows testing whether retirees were heterogeneous
in their responses to the Global Financial
Crisis and retirement incomes policy changes, demonstrating the value of the GFE estimator to such event studies.
In the presence of group-level time-varying
unobservable heterogeneity, using the standard two-way fixed-effects (2WFE) model defined in Section \ref{sec:Methodology} generally gives biased covariate effect estimates
and unrepresentative time effect estimates. This general result is consistent with the findings of the present study: estimates for the time effects in the standard 2WFE model, which allows for only one set of time effects, obscure the different latent group-level
effects which suggest retirees may adopt simple, distinct heuristics when
determining how they draw down their accumulated wealth during
retirement. The importance of this comparison with the 2WFE model results motivates our modification to the GFE estimator described in Section \ref{sec:Methodology}.
A key novelty in this application of the GFE estimator is
the
focus
on performing statistical inference on the estimated effects of the latent
heterogeneity. To robustify the GFE estimator in this scenario,
this article proposes
an extension to the bootstrap method
outlined in the Online Supplement to \citet{bonhomme2015grouped}. While
these authors describe a procedure for using the bootstrap to obtain standard
errors for covariate effects, a group label switching problem prevents finding
standard errors for the effects of group-level time-varying unobservable
heterogeneity. A method to match
group labels across independent GFE estimations is required so that the bootstrap can
also explore uncertainty in the latent heterogeneity effect estimates.
Related to this, the fixed-$T$ variance estimate formula given in the
Online Supplement to \citet{bonhomme2015grouped} provides an alternative approach
to performing inference on the effects of the latent heterogeneity, where $T$ is the maximum number of observations on each unit in the sample.
The proposed label-matching method also allows studying the
properties of the standard errors for the latent group effects derived
from this analytical formula, by observing their distribution across
a large number of simulated datasets. This is useful in determining the
potential applicability of the analytical formula for obtaining standard
errors of the latent group effects in finite samples, explored in the Online Supplement to this article.
The rest of the paper is organised as follows. Section \ref{sec:Related-Literature} summarises the literature to which this article contributes. Section \ref{sec:Superannuation-Drawdowns-Dataset} describes the available data for our application to retirement decumulation behaviours. Section \ref{sec:Methodology} outlines the GFE estimator from \citet{bonhomme2015grouped}, and explains our proposed extensions. Section \ref{sec:Main-Results} presents the main results from applying the GFE estimator to the available data. Section \ref{sec:Discussion} concludes by examining the implications for retirement incomes research and policy.
This article has an Online Supplement which provides robustness checks, a simulation exercise, descriptive analysis of the data, and further details on the second proposed methodological extension.
\section{\label{sec:Related-Literature}Related literature}
\subsection{\label{subsec:Drawdown-Behaviours-in}Drawdown behaviours in phased withdrawal retirement income products}
Recent
work
studying drawdown behaviours builds on \citet{horneff2008following}. These authors compare three rate-based drawdown strategies against a level guaranteed
lifetime annuity, using a model of retiree utility that incorporates
stochastic interest returns and retiree lifetimes. The first strategy
draws a fixed proportion of the account balance each year; the
second determines a terminal time horizon $T$ and draws a proportion $1/T$ of the account balance
in the first year, $1/(T-1)$ in the second year, and so on, continuing until
the $T^{th}$ year of the plan when all the remaining funds are drawn
down; the third draws a proportion of the balance that updates each year based on the surviving retiree's expected remaining lifetime.
\citet{bateman2008choices} extend
\citet{horneff2008following}, motivated by newly legislated minimum annual drawdown rates
for a phased withdrawal retirement income product in Australia known
as the account-based pension. Table \ref{tab:Age-based-Minimum-Drawdown-Rates}
details these age-based rates, which start at 4\% of the account balance
for retirees under age 65, and increase as a step function of age
to a maximum of 14\% for retirees aged 95 or above; these are specified in
Schedule 7 of the Superannuation Industry (Supervision) Regulations
1994.
The rates, when applied to a retiree's account balance at the
start of the relevant financial year, give the dollar amount the retiree
must draw down from their account during that financial year.
\citet{bateman2008choices} use a stochastic lifecycle
model to compare five drawdown strategies. Three are based on \citet{horneff2008following}; the remaining two are based on legislated
minimum drawdown rates in Australia: always drawing at the newly legislated
minimum drawdown rates effective from 1 July 2007;\footnote{Superannuation Industry (Supervision) Amendment Regulations 2007 (No.~1) Schedule 3.}
and drawing at the minima previously in effect. \citet{bateman2008choices} also use their
model to derive the implied optimal drawdown plan,
and examine the sensitivity of these results to changes in the model
assumptions. The fixed-rate strategy
they consider draws a fixed proportion of
the remaining account balance each year, where this fixed proportion is
the annuity payout rate the retiree would have received in
the market at the time of writing by purchasing a guaranteed lifetime
annuity using their account balance at retirement. The authors find that while the newly legislated minima outperform the second and third rate-based strategies of \citet{horneff2008following} described earlier, in most cases, simulated utility is increased by initially drawing at a fixed rate higher than the minimum, then switching to the new minimum drawdown schedule once the minimum rate rises higher than the fixed rate.
Based on this theoretical literature alone, researchers can expect, \emph{a priori}, that analysing drawdowns data will reveal groups
of retirees exhibiting distinct drawdown behaviours based
on simple rules, or heuristics. The present article uses recent advances in panel
data methods to empirically investigate whether such groups exist, and if so, what behaviours they display.
Section \ref{subsec:Capturing-Time-varying-Unobserve} summarises recent methodological progress in panel
data econometrics.
\subsection{\label{subsec:Capturing-Time-varying-Unobserve}Capturing time-varying unobserved heterogeneity}
\citet{bonhomme2015grouped} present the Grouped Fixed-Effects (GFE)
estimator and apply it to a linear panel data model that allows for group-level time-varying unobserved heterogeneity and unknown group membership. Importantly, their method allows for correlation between the latent group effects and the observed regressors.
Being able to control for, and estimate, the varying latent group time profiles is the key advantage of using this model over the traditional fixed-effects model, which captures
only time-constant unit-level unobservable heterogeneity.
Factor-model approaches to capturing unobserved heterogeneity \citep[e.g.,][]{bai2009panel, su2018identifying} provide more flexible specifications for the individual-specific heterogeneity. These models assume the presence of common, but unobserved, time-varying factors to which individuals respond heterogeneously.
\citet{bonhomme2015grouped} show that their linear panel model, to which they apply the GFE estimator, is a special type of latent-factor model, where the latent factors are group-specific time effects and the factor loadings are group membership indicator variables. In applications where group time profiles are of economic interest, having this interpretation of the time-varying unobservable heterogeneity is an advantage of using the GFE estimator over a factor-model approach. Furthermore, compared to competing latent-factor model methods, the GFE estimator converges faster to the true parameter values with $T$, and so has better finite-sample performance \citep{bonhomme2015grouped}. This makes GFE the preferred estimator when analysing a typical microeconomic panel, which may have a large number of units $N$, but rarely more than a moderate length $T$.
A related research area concerns finite
mixture models, which can incorporate unobservable heterogeneity when the heterogeneity is constant over time. Finite mixture models are parsimonious solutions to modelling unobserved
heterogeneity in populations, wherein data from latent subpopulations
can be drawn from different distributions, and the true allocation
of specific individuals to subpopulations is unobserved. \citet{deb2013finite} develop finite mixture models with
time-constant fixed effects. Critically, these models cannot specify the likelihood of being in
a particular group as a function of the covariates. By contrast, one
of the primary advantages of using the GFE estimator over a standard 2WFE model estimator is the ability to control for correlation between latent group effects and the
included covariates. Mixture-of-experts
models \citep[e.g.,][]{jacobs1991adaptive, jordan1994hierarchical} generalise mixture models by
parametrically specifying a link between group identity and covariate
values; however, these models do not control for unit-level fixed effects
in panel data.
The discussion above suggests that
the GFE
estimator
is the most appropriate for the present application, because: the available dataset only has
a moderate panel length $T$; a review of the theoretical drawdowns literature
suggests that there may be a finite number
of distinct behavioural groups in the population, consistent with
the GFE assumption that there are a finite number of latent group time profiles; the time profiles recovered
by the GFE estimator are likely to have economically meaningful interpretations.
\section{\label{sec:Superannuation-Drawdowns-Dataset}Data on superannuation drawdowns}
The available superannuation dataset contains information on drawdowns from account-based
pensions (ABPs), a phased withdrawal retirement income product
in Australia. Using ABPs, retirees generate personal income streams or receive
lump-sum payments by drawing down their accumulated savings during
retirement. Throughout, the balance remains invested in financial
markets based on each retiree's chosen combination of safe and risky
exposures.
Following the Global Financial Crisis (GFC), the Australian government introduced
temporary minimum drawdown rate reductions as per Schedule
7 of the Superannuation Industry (Supervision) Regulations 1994. These resulted in `concessional' minimum drawdown
rates, which came into effect for the financial year ended 30 June
2009. For financial years 2009--2011 inclusive, the reduction
was 50\%, so that any retiree's minimum drawdown rate over this period
was halved relative to the standard schedule in Table \ref{tab:Age-based-Minimum-Drawdown-Rates}.
For financial years 2012 and 2013, the reduction decreased to 25\%,
meaning the minimum drawdown rates were three-quarters of those given
in the table. For subsequent financial years, the rates returned
to their nominal values.
To generate an income stream from an ABP, retirees can elect a payment
amount and a frequency, for example \$1000 monthly, which the superannuation
fund follows in paying the retiree from their account balance.
This type of drawdown, which is specified by the beginning
of a given financial year, is referred to here as a `regular' drawdown, and is the primary
object of interest in this study. These regular drawdowns are distinct
from `ad-hoc' drawdowns, referring to any additional
lump-sum payments requested by the retiree during the financial year.
A rigorous analysis of ad-hoc drawdowns is crucial
in understanding retirees' needs for flexibility and insurance, but
is not the focus of this paper.
The dependent variable studied here is $y_{it}:=\ln {\rm DR}_{it}$; ${\rm DR}_{it}$ is the rate of drawdown for retiree $i$ over financial year $t$, which is a ratio of the regular drawdown amount over the account balance: ${\rm DR}_{it}:={\rm DA}_{it}/{\rm AB}_{it}$. The dependent variable is the logarithm of this rate because the raw rate variable is strictly positive and right-skewed. Appendix \ref{sec:p1-appendix} gives further details on the relationship between the drawdown rate and the account balance, as well as histograms of the dependent variable and the raw rate variable.
Qualitatively,
one interpretation of the drawdown rate is the speed at which retirees deplete
their accumulated savings throughout retirement. In ABPs, retirees
are exposed to the risk of outliving their savings and their retirement income becoming reliant
on non-superannuation assets or taxpayer-funded transfer payments known as the age pension.
Policy and product design should support retiree needs
for regular income, and so understanding the drawdown behaviours that
manifest in this flexible withdrawal product has the potential to
inform policymakers, financial advisors and superannuation fund trustees.
The superannuation dataset contains $N=9516$ individuals,
combining source data from multiple large industry and retail superannuation funds. Due to the small number of funds in the sample whose
data permitted determining the regular drawdown rate, the economic results
in this paper may not be representative of the Australian superannuation
system as a whole.
The data capture window spans the financial years 2004 to 2015 inclusive, for $T=12$ annual observations. The main results use a balanced subset of each superannuation fund's data, removing individuals with missing observations.
However, as the data capture window is not the same for each fund, when combining these balanced subsets, the resulting dataset is unbalanced.
The Online Supplement shows that the results are robust to using a fully balanced subsample of all the available data.
For 2.4\% of the records in the sample, missing distinctions between regular and ad-hoc drawdown amounts can be imputed based on observed movements in the corresponding total drawdown amounts. The imputation method involves calculating, for each individual--time combination, the individual's average monthly total drawdown amount in that financial year. Comparing these averages with the exact monthly total drawdown amounts, if an individual's exact drawdown amount is less than or equal to 105\% of the average drawdown amount, the exact drawdown amount is treated as a regular drawdown, and the corresponding ad-hoc drawdown amount for that month is zero. Otherwise, if an individual's exact drawdown amount is more than 105\% of the average drawdown amount, the regular drawdown amount for that period is set to the computed average, and the excess is treated as an ad-hoc drawdown in that month. In periods where the individual makes an ad-hoc drawdown, this method may underestimate the ad-hoc drawdown amount.
Individuals who die during the sample observation period are dropped from the data.
The superannuation dataset is derived from administrative data, and so has limited demographic
information on members. Table \ref{tab:agg-time-invariant-variables}
provides summary statistics for characteristics that vary across individuals but not time; Table \ref{tab:agg-time-varying-variables} summarises the time-varying variables in the dataset. The two covariates considered are the minimum
drawdown rate for each individual at each time point, and their account
balance at the beginning of the respective financial year, both on the log scale.
Thus the model, with regular drawdown rate on the log scale as the
dependent variable, estimates the elasticity of the regular drawdown
rate with respect to the minima and account balances. The remainder
of the variation in log regular drawdown rates is attributed to the
latent group effects estimated by the GFE procedure, and residual noise.
The lack of demographic information, particularly health and marital
status, is also a data-based motivation for using GFE as an estimation
procedure. These unobserved characteristics are a potential source of omitted variable bias; they are possibly correlated with the dependent variable as well as the included covariates. Because relevant characteristics such as health and marital status may vary over time, controlling for time-constant
unobserved heterogeneity by using a standard fixed-effects model may not remove the biasing effect of all relevant unobservables. For this reason, a model allowing for sources of time-varying unobserved heterogeneity,
such as the linear panel model to which the GFE estimator is applied, may be more appropriate; the results in
Section \ref{subsec:Common-Parameter-Estimates} provide evidence
for this proposition.
\section{\label{sec:Methodology}Methodology}
\subsection{\label{subsec:Grouped-Fixed-Effects}The Grouped Fixed-Effects (GFE) Estimator}
The GFE estimator introduced by \citet{bonhomme2015grouped} considers a linear model of the
form
\begin{equation}
y_{it}=x_{it}'\theta+\alpha_{g_{i}t}+v_{it};\label{eq:GFE-model-vanilla}
\end{equation}
in our application $i=1,2,...,N$ indexes individual retirees; $t=1,2,...,T$ indexes financial
years; $y_{it}$ is the dependent variable; $x_{it}$ is a vector
of covariates; $\theta$ are the partial effects of covariates
on the dependent variable after controlling for group-level time profiles;
$g_{i}=1,2,...,G$ identifies the group membership for unit
$i$, where $G$ is the chosen number of latent groups in the sample; $\alpha_{g_{i}t}$ is a term representing time-varying,
group-specific unobserved heterogeneity---these are time-varying model intercepts that can differ across individuals, depending on the value of the individual's group identifier, $g_i$; $v_{it}$ represents
the residual effect of all other unobserved determinants of the dependent
variable.
For the estimator consistency and asymptotic Normality results of \cite{bonhomme2015grouped} to apply, the distribution of the errors, $v_{it}$, must have finite fourth moment and be uncorrelated with the covariates, $x_{it}$; errors can be correlated across cross-sectional units if there is a finite limit on the magnitude of the correlation; the unit-level time-sequences of errors can be autocorrelated if the sequences are strongly mixing processes; the tails of the error distribution must exhibit a faster-than-polynomial decay rate. The time profile values, $\alpha_{gt}$, may be correlated with $x_{it}$ but not $v_{it}$. For the complete list of assumptions required for the GFE estimator, see \cite{bonhomme2015grouped}, pp.~1156--1159.
The linear model in (\ref{eq:GFE-model-vanilla}) can be extended to include time-constant, individual-specific
fixed effects, $c_{i}$, so that
\begin{equation}
y_{it}=x_{it}'\theta+c_{i}+\alpha_{g_{i}t}+v_{it},\label{eq:GFE-model-with-c_i}
\end{equation}
because applying the within transformation (centering the variables around individual-specific means) reduces (\ref{eq:GFE-model-with-c_i}) to the form of (\ref{eq:GFE-model-vanilla}). To see
this, for any variable $z_{it}$, define $\dot{z}_{it}:=z_{it}-\bar{z}_{i}$, where $\bar{z}_{i}:=T^{-1}\sum_{t=1}^{T}z_{it}$; then,
\begin{align}
\dot{y}_{it} =\dot{x}_{it}'\theta+\dot{\alpha}_{g_{i}t}+\dot{v}_{it},\label{eq:GFE-model-with-c_i-within-transformed}
\end{align}
which has the same form as (\ref{eq:GFE-model-vanilla}). In (\ref{eq:GFE-model-with-c_i-within-transformed}), the time profiles, $\dot{\alpha}_{g_{i}t}$, have a different economic interpretation to the $\alpha_{g_{i}t}$ from (\ref{eq:GFE-model-vanilla}); Section \ref{subsec:time-profs-in-trans-model} explains the treatment and interpretation of the group time profile estimators in the transformed model (\ref{eq:GFE-model-with-c_i-within-transformed}).
Throughout this paper, `GFE model' refers to (\ref{eq:GFE-model-with-c_i}). A
useful way to consider the GFE model is as a generalisation of the 2WFE model,
\begin{equation}
y_{it}=x_{it}'\theta+c_{i}+\alpha_{t}+v_{it},\label{eq:standard-two-way-fixed-effects-model}
\end{equation}
where the GFE model allows distinct time profiles for $G$ groups, with group membership unobserved.
\citet{bonhomme2015grouped} state the assumptions needed for large-$N$, large-$T$ consistency and asymptotic normality of the GFE estimator, in
models both with, and without, time-invariant fixed effects. Moreover, the asymptotics show that the estimator converges in distribution even when $T$ grows substantially more slowly than $N$. For any fixed value of $T$, however, the $\theta$ and $\alpha$ estimators are root-$N$ consistent not for the true population parameters, but for `pseudo-true' parameter values, and although these pseudo-true values may differ from the true population values for any fixed $T$, any difference vanishes rapidly as $T$ increases. Specifically, the pseudo-true and true values may differ when group classification is imperfect; because of this possibility, the authors derive a fixed-$T$ consistent variance estimate formula which takes into account potential group misclassification, providing a practical alternative to the large-$T$ variance estimate formula \citep[][p.~1161]{bonhomme2015grouped}. To examine the finite-$T$ properties of the GFE estimator, \citet{bonhomme2015grouped}
use a Monte Carlo exercise with simulated datasets that match the $N=90$,
$T=7$ panel used in their empirical application.
In the present study, the superannuation dataset has large $N=9516$ but moderate $T=12$; the Online Supplement finds that the GFE estimator and the fixed-$T$ consistent variance estimator perform well on simulated datasets of this size.
There are two problems in using the standard 2WFE model when there is group-level time-varying unobservable heterogeneity,
$\alpha_{g_{i}t}$, correlated with observable characteristics,
$x_{it}$: first, estimates of the covariate effects, $\theta$, may suffer
from omitted variable bias;
second, the group time
profiles, which in many applications have interesting economic
interpretations, remain hidden.
The 2WFE
model is a special case of the GFE model with $G=1$; hence, the GFE model is an appealing alternative to the 2WFE
model in applications where there is the possibility
of unobservable time-varying heterogeneity in addition to time-invariant unobservable heterogeneity. While
determining the precise number of groups, $G$, lacks a general
solution, estimating the GFE model can provide evidence for
whether only controlling for one set of time effects, as in the 2WFE
model, is adequate; Section \ref{subsec:Common-Parameter-Estimates} discusses this issue.
The GFE estimator for (\ref{eq:GFE-model-vanilla}) minimises the sum of squared residuals, giving
\begin{equation}
(\widehat{\theta},\widehat{\alpha},\widehat{\gamma})=\operatorname*{arg\,min}_{(\theta,\alpha,\gamma)\in\mathit{\Theta}\times\mathcal{A}^{GT}\times\mathit{\Gamma}_{G}}\sum_{i=1}^{N}\sum_{t=1}^{T}(y_{it}-x_{it}'\theta-\alpha_{g_{i}t})^{2},\label{eq:GFE-estimator-definition}
\end{equation}
where the vector $\gamma=(g_{1},g_{2},...,g_{N})$ defines the
grouping of each of the $N$ units into one of the $G$ groups; $\mathit{\Gamma}_{G}$
is the set of all possible groupings of $N$ units into $G$ or fewer
groups; $\mathit{\Theta}$ is a subset of $\mathbb{R}^{k}$, the $k$-dimensional real space, where $k$ is
the number of covariates or equivalently the dimension of $x_{it}$;
$\mathcal{A}$ is a subset of $\mathbb{R}$; the $\alpha_{gt}$ parameters belong to $\mathcal{A}^{GT}$.
\citet{bonhomme2015grouped} present an iterative algorithm for estimating (\ref{eq:GFE-estimator-definition}). The
algorithm initialises values for the grouping vector $\gamma$ and
the model parameters $(\theta,\alpha)$; it then alternates between the
following two steps until convergence: 1) a grouping update step, which
allocates each unit to the group minimising the sum of squared
residuals given the most recent estimate of the model parameters $\theta$
and $\alpha$; 2) a parameter update step, which estimates the parameters $(\theta,\alpha)$
conditioning on the most recent estimate of the grouping vector $\gamma$.
Estimating the GFE model does not require a balanced panel; the algorithm can be adjusted to run even when the sample contains unit--period combinations with missing data.
However, when estimating the GFE model with unbalanced data, preserving the equivalence between the 2WFE model (\ref{eq:standard-two-way-fixed-effects-model}) and the GFE model (\ref{eq:GFE-model-with-c_i}) with $G = 1$ requires an adjustment to the original GFE estimator.
The original GFE estimator for (\ref{eq:GFE-model-with-c_i}) with $G=1$ gives the same estimates as applying the within transformation to the data and running a least-squares regression on the time-demeaned data, including $T$ time dummy variables and no constant term.
In balanced samples, this returns the same numerical results as standard implementations of the 2WFE
model: for example \texttt{Stata}'s
\texttt{xtreg} command with the \texttt{fe} option and time dummy variables as covariates.
However, in unbalanced panels, the results differ.
With unbalanced data, obtaining the results from standard implementations of the 2WFE model requires time-demeaning the time dummy variables at the unit level. Hence, this article proposes and utilises a modified GFE estimator that for $G=1$ recovers precisely the same estimates as a standard 2WFE estimator even when the data is unbalanced and the model includes time-invariant unit-level unobservable heterogeneity; see Section \ref{subsec:Extension:-Alternative-Estimation-Method}. When the panel is balanced, or when there are no time-invariant unit-level unobservables, the results from this method are identical to the unmodified GFE estimation results. However, panel data models in microeconomic applications generally control for time-invariant unit-level unobservable heterogeneity.
Hence, this is a useful extension when applying the GFE estimator to microeconomic applications more broadly.
To examine the sensitivity of the present results to using the modified algorithm, the Online Supplement provides a robustness check using the unmodified GFE estimation procedure in place of the proposed modified algorithm.
\citet{bonhomme2015grouped} explain that in the absence of covariates, their algorithm outlined above reduces to $k$-means clustering. Similarly to standard implementations of $k$-means clustering
algorithms, the results depend on the starting values. Running the algorithm multiple times with randomly generated starting values increases the likelihood of finding
solutions corresponding to smaller values of the objective function in
(\ref{eq:GFE-estimator-definition}).
In the present study, the algorithm is initialised by drawing starting values for each covariate effect from independent
Gaussian distributions, centred at the coefficient estimates from a 2WFE
regression fit to the data,
and
with standard deviations equal to the magnitude of these coefficient
estimates. Results for the application
take the most optimal solution across 1000 independent runs of the
algorithm using randomly drawn starting values to initialise
the parameters. The Online Supplement also tests the sensitivity of the results to the number
of starting values used.
To facilitate inference on the group time profiles, we implement the fixed-$T$
variance estimate formula in the Online Supplement to \citet{bonhomme2015grouped}, as well as an extension to their non-parametric bootstrap method.
Notably, approximating the variability in time profile estimates
using the non-parametric bootstrap method described by \citet{bonhomme2015grouped}
requires solving a label-switching problem; see Section \ref{subsec:Extension:-Bootstrapping-Time}.
\subsection{\label{subsec:time-profs-in-trans-model}Time profiles in the transformed model}
Within-transforming the data to
remove the effect of any time-constant unobservable heterogeneity
results in time-demeaned group time profiles, $\dot{\alpha}_{gt}$. As the $\dot{\alpha}_{gt}$ contain only
information about changes in the group effects over time relative to their mean, and no information
about the absolute level of the effect, the estimates of the $\dot{\alpha}_{gt}$ are modified in
order to obtain the desired economic interpretations. All
estimated time profiles are shifted to begin at a value of 0 at $t=1$,
so that the interpretations of the values are as changes relative
to the first time period. This is analogous to estimating a set of
$T-1$ coefficients for the time effects in the standard 2WFE
model, where each estimate is interpreted as a change
relative to the omitted reference period.
This study uses two methods to estimate the uncertainty in the shifted time-demeaned group time profiles,
defined as $\tilde{\dot{\alpha}}_{gt}:=\dot{\alpha}_{gt}-\dot{\alpha}_{g,1}$,
for all $g$ and $t$.
First, Normal-approximation 95\% confidence intervals are constructed for the $\tilde{\dot{\alpha}}_{gt}$, using variance estimates which are computed from the variance--covariance matrix components for the time-demeaned time profiles, $Var(\dot{\alpha}_{gt})$, $Var(\dot{\alpha}_{g,1})$ and $Cov(\dot{\alpha}_{gt},\dot{\alpha}_{g,1})$; these objects are obtained by applying the fixed-$T$ variance estimate
formula after estimating the GFE model on the within-transformed
data. Second, these confidence intervals are compared to the corresponding intervals obtained from the bootstrap, using the proposed extension for matching group labels across bootstrap replications.
A comparison of the time profile standard errors
derived from the fixed-$T$ variance estimate formula versus simulated standard
errors from a Monte Carlo exercise shows that the fixed-$T$ variance estimate formula performs well on simulated data calibrated
to the superannuation dataset; see the Online Supplement for details. Moreover, in the present study, confidence intervals
constructed from the bootstrap procedure results are broadly similar to those derived from the fixed-$T$ variance estimate formula, implying practically equivalent
inference on the estimated parameters. For smaller datasets, where
asymptotic results are less likely to provide accurate standard
errors, and where the computational cost of running the bootstrap
is lower, it may be preferable to use the bootstrap. However, the results
suggest that in larger datasets, using the fixed-$T$ variance estimate formula may be sufficiently accurate while keeping the
computational burden low compared to using the bootstrap.
\subsection{\label{subsec:selecting-G}Selecting $G$}
\citet{bonhomme2015grouped}
propose an information criterion that correctly selects the number
of groups
when $N$ and $T$ are both large and similar in magnitude; however, for large-$N$ moderate-$T$ panels,
like the superannuation dataset, this criterion may overestimate the true value of $G$ \citep{bonhomme2015grouped}.
An alternative method to choose the number of groups, which \citet{bonhomme2015grouped}
employ in their empirical application, uses the fact that underestimating
$G$ leads to a type of omitted-variable bias in the $\theta$ parameters, to the extent that the unobserved effects are correlated with the included covariates.
Conversely, overestimating $G$ does not
bias the $\theta$ estimates, although it does increase the number of parameters
to estimate and the parameter space of possible groupings, resulting
in noisier estimates for all model parameters.
Taken together, these properties suggest that observing a plot of
how the coefficient estimates change as $G$ increases
can reveal the point at which the coefficients have
been de-biased relative to the $G=1$ case. Visually, the coefficients
might vary significantly between $G=1$ and some higher value, $G^\star$, after
which they may appear to stabilise, albeit becoming noisier as $G$
continues to increase. One can then select $G^\star$ as the number of groups suggested by the data.\footnote{This method would not apply to a model with group-specific covariate effects.} This is how \citet{bonhomme2015grouped}
decide on $G=4$ in their application, rather than $G=10$ as suggested by the information
criterion---which, given the dimensions of their panel, may
overestimate the true number of groups.
For applications that focus only on estimating unbiased effects of the covariates, the above suggests choosing the number of groups by selecting the smallest value of $G$ that places the covariate effect estimates close to their stable values.
However, the present application is interested in the time profiles for all latent groups in the sample.
It is possible that stopping
at the first value of $G$ that seems to successfully de-bias the
regression coefficients will result in at least one estimated group whose
time profile is an average over distinct time profiles of constituent subgroups.
For the purposes of obtaining unbiased regression coefficients alone,
separating out these mixed groups by further increasing $G$ may unnecessarily increase the complexity of the model. However, identifying all distinct behavioural groups with
economically meaningful interpretations may require searching beyond the smallest $G$ that de-biases the regression coefficients.
On the other hand, overestimating $G$ can result in the separation of a particular time profile into multiple biased representations of itself, where the difference between the representations is spuriously driven by noise \citep{bonhomme2015grouped}. This means that the more similar a pair of time profiles appear, the more sceptical the researcher should be in regarding them as distinct effects.
Hence,
this study posits a selection rule that picks the greater of two values
of $G$: the smallest $G$ that appears to result in unbiased estimates
of the regression coefficients; or the last value of $G$ for which
the marginal change from $G-1$ to $G$ groups finds an economically
important behaviour, sufficiently distinct from all others and exhibited by a non-trivial portion of the sample.
Section \ref{sec:Main-Results} discusses this selection process alongside the main results.
\subsection{\label{subsec:Extension:-Bootstrapping-Time}Extension 1: Bootstrapping time profile estimates}
\citet{bonhomme2015grouped} address how to obtain bootstrap standard errors for the covariate effects, but not the time profiles. As the covariate effects are unambiguously labelled across multiple
runs of the GFE procedure, estimating their standard errors by comparing
estimates across a large number
of bootstrap replications is standard.
By contrast, group-level results from the procedure are unique only up to a relabelling of the groups, and the group labels are determined at random during estimation. This means that the same economic group may receive different labels across multiple independent estimation runs, e.g., when performing the bootstrap. Hence, estimating the standard
errors of group-level time effects by comparing time profile estimates across bootstrap replications requires a method to match labels across runs.
The bootstrap procedure outlined in the Online Supplement to \citet{bonhomme2015grouped}
involves running the GFE procedure
$B$ times, each time on a different bootstrap replicate dataset.
To create the bootstrap
replicate datasets, all $N$ units in the original
data are sampled with replacement $N$ times, taking all observations corresponding
to a given unit into the replicate dataset when that unit is sampled.
Each of the $B$ model runs produces a set of coefficient estimates
for the model covariates, as well as an estimated grouping and a set
of time profiles for each group.
To match group labels across replications, we propose a
label-matching procedure,
similar to the solution \citet{hofmans2015added}
use for the analogous problem in $k$-means clustering; see the Online Supplement for details.
This method matches time profiles across different
bootstrap replications in the main results presented in Section \ref{sec:Main-Results},
as well as across different simulated datasets and bootstrap replications
in the simulation exercise reported in the Online Supplement. In general, it could be seen as an issue that time profile bootstrap results are contingent on some method to match labels, because, in principle, the bootstrap results could prove sensitive to the choice of label-matching method. The results here show that in this application, the bootstrap and analytical formula results are similar, and so little is gained by using the bootstrap over the analytical formula. Thus, we did not pursue this question further, but this general problem may be of interest to future research.
While deriving standard errors for the shifted time-demeaned group time
profiles, $\tilde{\dot{\alpha}}_{gt}$, using the fixed-$T$ variance estimate formula requires combining elements from the variance--covariance matrix of the time-demeaned time profile
estimators, $\dot{\alpha}_{gt}$, as described in Section \ref{subsec:time-profs-in-trans-model}, the bootstrap procedure can directly estimate the variability of the $\tilde{\dot{\alpha}}_{gt}$ terms.
\subsection{\label{subsec:Extension:-Alternative-Estimation-Method}Extension 2: Alternative estimation method for unbalanced data}
Consider the model in (\ref{eq:standard-two-way-fixed-effects-model}), which, after time-demeaning, has the form
\begin{equation}
\dot{y}_{it}=\dot{x}_{it}'\theta+\dot{\alpha}_{t}+\dot{v}_{it}.\label{eq:2WFE-using-GFE-notation}
\end{equation}
The primary estimands of interest are the $T-1$ relative effects $\tilde{{\dot{\alpha}}}_{t}=\dot{\alpha}_{t}-\dot{\alpha}_{1}$ for $t=2,3,...,T$. Since $\dot{\alpha}_{t}:=\alpha_{t}-\bar{\alpha}$, where $\bar{\alpha}=T^{-1}\sum_{t=1}^{T}\alpha_{t}$, the set of $T$ values of $\dot{\alpha}_{t}$ satisfy $\sum_{t=1}^{T}\dot{\alpha}_{t}=0$.
In balanced panels, the following three methods obtain identical estimates for the $\tilde{{\dot{\alpha}}}_{t}$:\footnote{This claim is justified in the Online Supplement.}
\begin{enumerate}
\item
regress $\dot{y}_{it}$ on $\dot{x}_{it}$ and $T$ time dummy variables with no constant term to estimate all $T$ values for $\dot{\alpha}_{t}$ as the coefficients of the time dummies; then compute $\tilde{{\dot{\alpha}}}_{t}=\dot{\alpha}_{t}-\dot{\alpha}_{1}$ for $t=2,3,...,T$;
\item
regress $\dot{y}_{it}$ on $\dot{x}_{it}$ and $T-1$ time dummy variables excluding the dummy variable for the first time period, including a constant term, after which the $T-1$ desired estimates of $\tilde{{\dot{\alpha}}}_{t}$ are the estimated coefficients of the included $T-1$ time dummy variables;
\item
first, time-demean the time dummy variables corresponding to periods $2,3,...,T$; then, regress $\dot{y}_{it}$ on $\dot{x}_{it}$ and the $T-1$ time-demeaned time dummy variables, not including a constant term, after which the $T-1$ desired estimates of $\tilde{{\dot{\alpha}}}_{t}$ are the estimated coefficients of the included $T-1$ time dummy variables.
\end{enumerate}
When the panel is unbalanced, the choice of estimation strategy requires further consideration. With unbalanced data, methods 1 and 2 described above for obtaining estimates of $\tilde{{\dot{\alpha}}}_{t}$ no longer guarantee that the corresponding
values of $\dot{\alpha}_{t}$ sum to zero; a nonzero sum contradicts a property of the model given by (\ref{eq:2WFE-using-GFE-notation}). Conversely, method 3 obtains the desired estimates while ensuring that the $T$ estimates of $\dot{\alpha}_{t}$, derived from the $T-1$ estimates for $\tilde{{\dot{\alpha}}}_{t}$,
sum to 0, consistent with the model.
Using the GFE estimator implementation provided by \citet{bonhomme2015grouped} with $G=1$ on unbalanced data gives estimates for the time effects in line with those produced by methods 1 and 2, whereas standard implementations of the 2WFE
model give the same estimates as method 3. This motivates our modification to the GFE methodology, which aligns the $G=1$ results with the standard 2WFE
model estimates, regardless of whether the panel is balanced or not: in the modified procedure's parameter update step, the algorithm interacts the group identity variables with $T-1$ time-demeaned time dummy variables, analogous to using method 3 for $G=1$.\footnote{The Online Supplement provides further details of this modification, as well as results using the unmodified GFE estimation procedure to compare with the results presented here.}
\section{\label{sec:Main-Results}Main results}
\subsection{\label{subsec:Common-Parameter-Estimates}Covariate effects}
Figure \ref{fig:covrt-ests-vs-G-for-G-1-to-large-sdsm2} plots the covariate effect estimates versus $G$ for the GFE model fitted to the superannuation drawdowns dataset.
The values of $G$ on the horizontal axis correspond to the number of latent groups specified; the vertical axis denotes partial effect values for the two included covariates: log minimum drawdown rate and
log account balance. The lines connect covariate effect estimates, while the
shaded regions around these correspond to 95\% confidence intervals
constructed using standard errors derived from the fixed-$T$ variance
estimate formula in the Online Supplement to \citet{bonhomme2015grouped}.
As the dependent variable and covariates are on the log scale,
the covariate estimates represent the elasticity
of regular drawdown rates to changes in account balances and the minimum
drawdown rates. Section \ref{subsec:Choice-of-G} provides the economic interpretation for the
corresponding estimates after choosing
the value of $G$.
A critical feature of Figure \ref{fig:covrt-ests-vs-G-for-G-1-to-large-sdsm2}
is that as $G$ increases, changes in the log account balance covariate effect estimates are large relative to the confidence intervals for the first few values of $G$. After $G=7$, the point estimates stabilise and
successive 95\% confidence intervals show significant overlap, suggesting that from
$G=7$, the GFE procedure removes the bias arising from correlation
between group-level latent time effects and the included covariates.
Using the proposed modified GFE model estimation procedure, the point estimates corresponding to $G=1$ are identical to the results from a standard 2WFE
estimation on the drawdowns dataset.
Hence, Figure \ref{fig:covrt-ests-vs-G-for-G-1-to-large-sdsm2} implies that the estimators using a more traditional analysis including only one time profile are biased compared to the estimators in the GFE model for $G\geq7$. The Online Supplement shows that this result holds
in the fully balanced case, where both the modified and unmodified GFE procedures give results identical to the 2WFE model when $G=1$.
\subsection{\label{subsec:Choice-of-G}Choice of $G$}
Figure \ref{fig:Geq4-to-9-time-profile-plots} plots the
time profiles estimated for $G=4,5,...,9$, all shifted to begin
at a $y$-axis value of 0, and omitting confidence interval bounds
for clarity. The figure suggests that while increases in $G$ initially produce new and economically distinct time profiles, at $G=7$, incremental moves to a larger value of $G$ split existing time profiles into highly similar representations of the same trajectory. This is consistent with a theoretical result from \citet{bonhomme2015grouped} that implies overestimating $G$ can create spurious copies of the true time profiles, differing by random noise.
Hence, the number of latent groups to use
in the model is set to $G=7$. Section \ref{subsec:Time-Profile-Estimates}
explains the economic interpretations of the resulting time profile
estimates.
Selecting $G=7$ groups corresponds to estimated covariate effects of 0.144
and $-$0.147 for the log minimum drawdown rate and log account balance,
respectively. These are statistically significant at the usual levels, e.g. 5\%, with respective standard errors of 0.0251 and 0.0140.
As the dependent variable and covariates are on the log scale,
the covariate effects represent partial elasticities of the regular drawdown rate
with respect to the minimum drawdown rate and account balances. For
example, consider an increase of 0.1 in the log minimum drawdown
rate, which translates to a proportional increase in the level of
the minimum drawdown rate of approximately 10.5\%. This increase in the log minimum drawdown
rate is expected to raise regular
drawdown rates proportionally by an estimated $e^{0.0144}-1=1.45$\%
on average, holding account balance equal and controlling for group-level
time-varying heterogeneity. A similar
proportional increase in a retiree's account balance is expected to effect
a proportional decrease of $1.46$\% ($1-e^{-0.0147}$) in their regular drawdown
rate on average, holding the minimum drawdown rate constant and controlling
for group effects.
\subsection{\label{subsec:Time-Profile-Estimates}Time profiles}
Figure \ref{fig:fixed-T-variance-formula-results-for-time-profile-plot-for-Geq7-model} shows
the time profiles estimated using the selected value of $G=7$ with element-wise confidence interval bounds constructed using standard
errors derived from the fixed-$T$ variance estimate formula.
Figure \ref{fig:bootstrap-B-eq-1000-results-for-the-G-eq-7-model-where-the-bootrep-datasets-are-drawn-from-the-actual-sample}
presents the equivalent plot with confidence interval bounds computed
from the empirical 2.5 and 97.5 percentiles of the time profile estimates
from 1000 bootstrap replications.\footnote{Each bootstrap replication uses distinct sets of randomly generated starting values for the estimation algorithm.} While qualitatively similar overall to the intervals constructed using standard errors derived from the fixed-$T$ variance estimate formula, a comparison highlights some differences.
In particular,
the intervals are wider around financial years 2011--2013 for group 6;
the interval is wider for financial year 2010 for group 4;
the intervals are tighter for group 5 throughout;
the interval is wider in the terminal financial year for group 3; and
the point estimates from the original sample tend towards the lower end of the 95\% bootstrap confidence intervals for groups 1, 2 and 7.
However, due to the estimated effect magnitudes, these numerical differences in the sets of confidence intervals do not meaningfully change inferential conclusions regarding the estimates or their economic interpretation.
The values of the time profiles represent behavioural effects on the proportional
changes in a retiree's regular drawdown rate, relative to their own
rate of drawdown in 2004.
As
the dependent variable is on the log scale, a value of, say, $0.5$
on the $y$-axis represents a proportional change in the regular drawdown
rate of $e^{0.5}\approx1.65$, or roughly a 65\% increase in the drawdown
rate, relative to 2004 levels.
Figure \ref{fig:Time-Profiles-Geq1model-CIs-from-formula} shows the time profile estimates for the $G=1$ model estimated using the modified GFE procedure, with the $y$-axis scale set to the same scale used
for the $G=7$ model group time profile plot (Figure \ref{fig:fixed-T-variance-formula-results-for-time-profile-plot-for-Geq7-model}).
Comparing these shows that a model controlling for only one set of time effects
fails to capture the shape and magnitude of any of the time profiles
in the seven-group model. Visually, the resulting single time profile
appears to `average over' the seven markedly distinct time profiles.
Hence, not only are
the covariate effect estimates biased when $G=1$,
but the
single time profile is entirely unrepresentative of the latent group-level
time profiles.
\subsection{\label{subsec:Characterising-Groups}Behavioural interpretations}
The following refers to the group time profiles in Figure
\ref{fig:fixed-T-variance-formula-results-for-time-profile-plot-for-Geq7-model}.
All the time profiles exhibit similar, slowly rising
trajectories until the 2008 financial year, after which
they diverge in statistically and economically significant ways.
Consider group 5, whose time profile describes a rapid reduction
in drawdown rates between financial years 2008 and 2010, and then
a gradual return towards pre-reduction levels. The timing of these movements follows closely the progression
of concessional minimum drawdown rates (see Section \ref{sec:Superannuation-Drawdowns-Dataset}).
This implies that members of group 5 were the most responsive to
the changing minimum drawdown rates over this period, compared to
the rest of the sample. Figure \ref{fig:g5-log-rdr-vs-fy} confirms that many individuals in group 5 show a step-like pattern
in their log regular drawdown rate series during the second half of
the observation period. This pattern is suggestive of the new step-function
schedule for minimum drawdown rates which came into effect 1 July
2007.
Groups 1 and 2 display time profiles rising steadily following financial
year 2008 for the remainder of the sample. Group 7 follows a similar
rising trend with a reduced magnitude for the first few financial years
after 2008, before stabilising for the remainder of the observation
window. The panel plots in Figure \ref{fig:some-g-log-rda-and-log-ab-vs-fy} suggest a tendency
for individuals in these groups to draw constant dollar amounts, while
their account balances gradually decline over time.
The corresponding plots for group 3 show a similar preference for constant dollar amounts, but after financial year 2013, individuals in this group suffer
a rapid decline in both the amounts drawn down and account balances.
Many members of group 6 undergo a similar evolution, characterised
by mostly constant drawdowns initially and a subsequent reduction in
the level in later years. The magnitudes of the downwards revisions appear smaller
for this group, and occur earlier for many retirees. Moreover, while
group 3 members continue to revise down over successive
financial years, it appears that many individuals in group 6 revise
down once and then continue to draw at the new, reduced level.
The group 7 plots show gradual downward revisions in the amounts year on year.
Thus, groups 1, 2, 3, 6 and 7 appear to be similar in that members draw mostly constant dollar amounts
over time; however, these groups seem to differ in how many members
make a downwards revision in the level of their drawdown amounts, and whether this
revision is once-off or the beginning of a downward trend. Crucially, the timing of downwards revisions often appears
to align with periods where account balances begin falling at an accelerated
rate. Hence, the heuristic of drawing constant amounts over time
may be adversely impacting retiree financial security by contributing
to a premature exhaustion of account balances.
Figure \ref{fig:fixed-T-variance-formula-results-for-time-profile-plot-for-Geq7-model} shows that the time profile for group 4 closely follows
the zero line before financial year 2011.
A time profile with all values near zero suggests the model for the group's data is well approximated by $y_{it}=x_{it}'\theta+c_{i}+v_{it}$.
Correspondingly, a possible interpretation for group 4's behaviour is that before financial year 2011, they were constantly tuning their drawdown rates as they progressively faced higher minimum
drawdown rates and as their account balances changed. This can potentially
represent a group of `engaged' retirees, who regularly occupy themselves
with determining their desired rate of drawdown; however, it is not possible to further investigate this hypothesis using this dataset.
Group 4's
time profile drops significantly below 0 after financial year 2010.
These movements, while of a smaller magnitude compared to the time profile values of other groups, are still economically significant; they suggest a behavioural response, after netting out the effect of covariates, of reducing drawdown rates relative to earlier levels from 2011 onwards.
The Online Supplement provides summary statistics and descriptive plots for all seven groups. Based on these, some of the ways in which the groups differ in terms of observable characteristics are:
the proportion of males in groups 1 to 3 are 62\%, 62\% and 69\% respectively, while the sample on aggregate is 56\% male;
the members of groups 1 and 3 tend to be older, with median retirees aged 81 at 31 Dec 2015;
members in groups 4, 5 and 6 have the youngest median retirees, aged 78 at 31 Dec 2015;
group 5 members have the highest median risk appetite, a variable defined in the Online Supplement as a summary measure of the relative sizes of derived equity returns within individual accounts compared to the S\&P/ASX 200 market index;
group 7 members have the lowest median risk appetite;
group 7 members make ad-hoc drawdowns least frequently, in roughly 3\% of person--years observed, followed by group 1, for 5\% of person--years observed---for all other groups this frequency was 9--11\%;
in years where they make ad-hoc drawdowns, group 3 members tend to draw down the largest proportions of their account balance using ad-hoc drawdowns over the course of a year---on average, these ad-hoc drawdowns over the year amount to 35\% of their account balance at the start of the respective financial year.
\subsection{Prior expectations}
Section \ref{sec:Introduction} presents this study's prior expectations
for finding at least two types of strategies in the data: drawing constant dollar amounts, and following closely
the minimum drawdown rates. Most of the group behaviours found under the
seven-group assumption can be interpreted as cases of these two hypothesised behaviours. Moreover,
a restricted model assuming only two groups already begins to show
evidence for the existence of both of these groups.
Figure \ref{fig:Time-Profiles-Geq2-model-CIs-from-formula}
shows the time profiles for the two-group model. As in Figure \ref{fig:Time-Profiles-Geq1model-CIs-from-formula},
the $y$-axis scale allows for direct comparison of the estimated
magnitudes with the $G=7$ model results.
The time profile of the first group appears similar to those for the groups in the seven-group model that seem to target
constant dollar amounts for the regular drawdowns; the second profile shows a dip comparable to the group
that seems to follow the minimum drawdown rate.
The Online Supplement also provides descriptive panel
plots for these two groups. These plots are broadly consistent with
the interpretation that the first group often attempts to hold drawdown
amounts constant, although with downward revisions in the amount common,
as well as rapid declines in account balance towards the end of the
period; the second group mainly makes decisions regarding
their drawdown rate. This analysis, alongside the observation that moving to a
two-group model already removes much of the bias in the covariate
effect estimates, suggests that accounting for both rate-based and
amount-based strategies seems to be the most important step in controlling
for unobservable heterogeneity in the data.
\section{\label{sec:Discussion}Discussion}
\subsection{\label{subsec:Retirement-Incomes}Retirement incomes}
A key implication of our work for retirement incomes research and policy is that there is now a statistical basis for empirical results that indicate different behavioural stances explain much of the variation in the drawdowns from phased withdrawal retirement income products. The significant change in covariate effect estimates and the emergence of distinct time profiles, as the number of groups assumed by the GFE estimator increases, is evidence for the presence of multiple behavioural groups in the data. The magnitudes of the estimated group time profile values suggest a sizeable contribution of this behavioural component to the observed drawdowns.
Retirees need to trade-off current against future consumption in the face of investment, longevity, and other, background risks, including inflation rates and health expenditure \citep{yang2009impact}. Retirement fund trustees and financial advisors can use the findings from this research when aiding retirees who are navigating the trade-offs in apportioning their accumulated retirement savings between consumption, income-generating assets and precautionary assets. The results do not find clear evidence for a group following the optimal heuristic strategy from \citet{bateman2008choices}, which is to initially draw at a fixed rate until the rising minima require drawing larger proportions. Two-thirds of the sample belong to the groups exhibiting a preference for drawing down a constant dollar amount over time. We do not have evidence as to whether the initial level is optimal, but two of the groups targeting level income streams had their account balances begin a rapid decline following the GFC; retirees with this preference who are unable to respond to periods of low or negative investment returns by curtailing their income streams may be at risk of prematurely depleting their account balances. 14\% of the sample belong to the group who most closely follow the next-most optimal heuristic strategy, which is to always draw at the legislated minimum rates.
Those providing financial advice may wish to convey these findings to retirees in the form of cautionary tales, or as part of a standard set of recommendations presented to retirees. Practitioners and policymakers may also wish to consider these findings when designing retirement income products and policy to align with the Comprehensive Income Products for Retirement framework, which places an emphasis on generating income streams without compromising the ability to retain superannuation savings until late into retirement by including some life annuity products \citep{Treas2016}. Financial advisors may benefit retirees by predicting their individual needs for income and precautionary savings later in life, monitoring their actual experience relative to this forecast, and recommending a reduction in expenditure before poorer experience exhausts their accounts.
The present study clarifies how a robust behavioural analysis
of drawdowns data is vital for informing the various stakeholders in the
retirement incomes space, including retirees, the government, and
the financial services industry. However, Section \ref{sec:Superannuation-Drawdowns-Dataset}
notes that despite having data from multiple superannuation
funds, due to the small number of funds in the estimation
sample, the results may not describe the population of
Australian retirees in general. In particular, using the current dataset
may overlook behavioural patterns present
in funds not included in the sample.
\subsection{\label{subsec:GFE-in-Behavioural}GFE in behavioural microeconomics and event studies}
We believe that the GFE model is valuable for research in behavioural microeconomics and event studies. While
the present paper focuses on the behavioural interpretations
of the latent group effects, the superannuation data
covers a period in which multiple common macroeconomic and policy
shocks affect all members in the sample. Hence, the superannuation application
also invites an event-study interpretation, although it is impossible
to disentangle the effects of several common shocks occurring
in close proximity. These include the Global Financial Crisis, the introduction of a new
schedule of minimum drawdown rates for account-based pensions, and
temporary tweaks made to the minimum drawdown rates. Taken together,
however, the time profile plots presented in Figure \ref{fig:fixed-T-variance-formula-results-for-time-profile-plot-for-Geq7-model}
clearly show how otherwise homogeneous-looking trends diverge radically,
plausibly driven by one or more of these shocks. Hence, the results
broadly show how researchers can utilise the same methodology
in an event-study framework.
In the superannuation application, the GFE assumption of a group structure to the time-varying
unobservable heterogeneity is tenable because it was \emph{a priori} expected that an important source
of heterogeneity in the drawdowns arises from the latent motivations
that drive individuals to choose one of a finite number of drawdown
strategies. In general, other behavioural economics studies
with panel data may have analogous motivations to assume
that there is some true, finite number of groups in the population, which the GFE estimator attempts to characterise. This prior assumption may
be important as it is not known precisely how the estimator behaves when a group structure only approximates, rather
than accurately captures, the nature of the unobservable heterogeneity.
For a discussion on approximating more general forms
of heterogeneity, see e.g. \citet{bonhomme2021discretizing}.