Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
66,083 characters · 32 sections · 0 citation commands
Time-to-Event Estimation with Unreliably Reported Events in Medicare Health Plan Payment
\setstretch{1}
Time-to-event (TTE) analyses typically aim to compare differences between groups drawn from two or more populations in a randomized clinical trial or observational study. This may include plotting events over time using either Kaplan-Meier or cumulative incidence curves. Most commonly, hazard ratios are used to estimate the magnitude of the effect between groups\textsuperscript{1--3}. For example, a review of 66 clinical trials with TTE primary outcomes across four major medical journals found that 80% of trials reported a hazard ratio as a main finding, while only 21% of studies reported any alternative approaches\textsuperscript{4}.
However, hazard ratios rely on a proportional hazards assumption, which is often unrealistic in health data\textsuperscript{1,4,5}. Hazard ratios also have an unintuitive interpretation for nonstatistical audiences and are a relative value with unclear meaning for decision-making\textsuperscript{2,6}. Less frequently, a difference or ratio of survival probabilities at a single time point (e.g., mean or median) is reported and compared across groups. These types of measures also have limitations in practice, including that they omit most of the data and may not be estimable in certain scenarios\textsuperscript{2,7}.
Furthermore, alternate approaches are needed if competing risks, or when there is more than one mutually exclusive outcome, are present. Despite being pervasive in health studies, competing risks are frequently ignored, which can result in biased estimates of the primary outcome\textsuperscript{8--10}. Additional important considerations in TTE analyses include the choice of monitoring period (i.e., the time window analyzed between a pre-specified origin and end time) and comparison group.
Although first proposed in 1949\textsuperscript{11}, restricted mean survival time (RMST) was revisited for TTE estimation in more contemporary literature\textsuperscript{5,6,12--14}. RMST is defined as the area under the survival curve of time to an event for a single monitoring period. It can be interpreted as the mean time to event for all study participants followed in that monitoring period\textsuperscript{6}. RMST addresses many of the issues of more popular comparison approaches. With large enough sample sizes, it is estimable nonparametrically, does not require proportional hazards, and is censoring independent\textsuperscript{6}. Recent methods development has focused on expansions for adaptive and group sequential trials\textsuperscript{15,16}, estimating RMST with varying end times\textsuperscript{7,17,18}, and permutation testing for small sample sizes\textsuperscript{19--22}.
In TTE analyses with a single outcome, the area above the survival curve in one monitoring period (or, the end time minus the RMST) is the restricted mean time lost (RMTL)\textsuperscript{23}. RMTL has a number of appealing features, including that it can straightforwardly be extended to correspond to the area below cause- or event-specific cumulative incidence curves in competing risk settings\textsuperscript{10,23,24}. In such scenarios, an event-specific RMTL corresponds to the area below its cumulative incidence curve and can be interpreted as the mean time without that particular event in a monitoring period\textsuperscript{10}, or a summary of both how many and when events occur in that time. Differences in RMTL between two groups (for example a treatment and control or other comparison group) can also be used to quantify effects\textsuperscript{25}.
Our motivating application focuses on the evaluation of incident diagnostic coding in Medicare. Medicare is a federally funded insurance program administered by the Centers for Medicare and Medicaid Services (CMS), supporting over 69 million Americans who are aged 65 years and older or chronically disabled\textsuperscript{26}. Beneficiaries choose between receiving coverage through Traditional Medicare (TM) or Medicare Advantage (MA). TM beneficiaries' care is generally paid directly by CMS for each service provided (often referred to as fee-for-service). In contrast, more than half of all Medicare beneficiaries are part of the MA program and receive health insurance from private insurers who are paid by CMS for each beneficiary's anticipated health needs.
We examine the Medicare risk adjustment algorithm that determines prospective health plan payments in MA based on binary demographic and health diagnostic variables using ordinary least squares linear regression. The outcome of this algorithm is a risk score, which is used to adjust a benchmark payment amount per beneficiary. Importantly, insurers retain the amount of money paid by CMS regardless of what care their beneficiaries actually receive\textsuperscript{27,28}.
There are 115 variables that report the purported presence or absence of certain health conditions, termed hierarchical condition categories (HCCs), in the current risk adjustment formula\textsuperscript{29}. Although HCCs overall correspond to dozens of distinct health conditions, certain subsets of these variables also represent severity levels within the same condition and are therefore billed in a mutually exclusive manner. For example, an insurer can only be paid for coding one of Pancreas Transplant Status, Diabetes with Severe Acute Complications, Diabetes with Chronic Complications, or Diabetes with Glycemic, Unspecified, or No Complications\textsuperscript{28}. We propose that such sets of HCCs can be considered conceptually analogous to competing risks in TTE analyses because only one HCC can be recorded and paid for at a time, where the amount paid to insurers increases with severity level.
Evaluation of coding in MA is of great interest to health policy researchers and policymakers. The structure of the risk adjustment formula (called the CMS-HCC formula) as an ordinary least squares regression with only positive coefficients incentivizes insurers to code for as many diagnoses as possible to increase profits\textsuperscript{30--32}. Insurers making their beneficiaries appear sicker than they are by coding for more diagnoses or more severe diagnoses than necessary is referred to as upcoding. We focus on two types of upcoding in this paper. The first we call severity-based upcoding, as beneficiaries who have a lower-severity version of an HCC are instead coded with a more severe HCC of that health condition. The second we refer to as any-available upcoding, where any beneficiaries previously not coded with a given HCC could be upcoded, potentially fraudulently. When upcoding has not been concretely confirmed (as is most often the case in real-world settings) we reflect this by describing it as possible upcoding.
Assessment of upcoding in MA is usually done by comparing frequency of coding of beneficiaries, or coding intensity, to coding of similar health conditions in TM, with slight variation on inclusion and exclusion criteria used. However, TM is known to have underreporting (i.e., undercoding) of health conditions\textsuperscript{28,33}. This type of unreliable reporting is important to account for as measures of differences between the programs will likely be inflated otherwise. Alternative reference data besides TM diagnoses have been explored in prior work but remain underutilized, including mortality data\textsuperscript{34} and prescription drug utilization\textsuperscript{35}. More recently, a set of measures were proposed to distinguish new (incident) versus continuing (persistent) coding for individual HCCs\textsuperscript{36}.
In 2025, high MA coding intensity--which likely includes upcoding--is estimated to cost CMS \$40 billion in unnecessary spending to private insurers without clear benefit to MA beneficiaries\textsuperscript{32}. In some past cases, CMS or the United States Department of Justice have determined that certain behaviors are upcoding from whistleblower reports of fraud\textsuperscript{37,38} or by manually auditing claims and electronic health records\textsuperscript{39,40}, neither of which are scalable approaches. This is a substantial problem, as most major MA insurers have been accused of fraud related to their billing practices by whistleblowers or the United States Department of Justice\textsuperscript{38,39,41}.
We expand RMTL estimation approaches for use in an impactful but understudied health policy application: evaluating incident HCC coding in MA following CMS-HCC formula updates. Here, individual HCCs and sets of mutually exclusive HCCs corresponding to severity levels of a single health condition are our events of interest, where an event occurs when an incident diagnosis is reported by an insurer. Besides being a unique application of TTE methods in itself, our approach expands on prior RMTL-based methodological work by proposing estimators for such analysis that can account for underreporting of events in TM, which, to our knowledge, has not been previously considered. Further, we propose a novel estimation approach for identifying one type of possible severity-based upcoding.
Our contributions also include the creation of an R package to simulate baseline HCC diagnoses as well as the ability to upcode or undercode these data. Our package simulates realistic co-occurring HCCs based on older American adults' self-reported health conditions from a large national survey. Self-reported data are not impacted by coding incentives to the same degree as billing claims or electronic health record data\textsuperscript{42--44}, and enable analysis of incident coding without underreporting. Our package additionally allows users to modify these baseline data by either underreporting existing baseline diagnoses (realistic to TM data) or upcoding specific HCCs. Labeled upcoding data did not previously exist, so these simulated data are broadly useful for methods development in health policy.
We outline our statistical approaches for TTE estimation in this section. In our setting, TTE and time to reported event are equivalent because we do not observe each coding event directly. However, since our estimates are over an entire monitoring period, we do not consider potential delays in reporting events to have notable impact. In addition, multiple events may be reported at each of several time points within a monitoring period (e.g., a monitoring period could be a calendar year where each time interval is three months). There also may be more than one monitoring period where reporting occurs. Given this, our goals are to both (1) evaluate reporting of incident events within one monitoring period and (2) compare incident event reporting across sequential monitoring periods. We focus on estimating differences in time to incident coding behaviors, as absolute measures are especially useful in health policy contexts\textsuperscript{45} and new behaviors following policy changes are as well.
An event (reported coding of either a single HCC or a single member of a set of HCCs corresponding to severity levels of one health condition) is considered incident if it was not reported in prior monitoring periods. The reporting of an incident event is represented by a vector \(S\), with \(s \ge 1\) possible mutually exclusive subtype events. \(S\) encodes which of these events occurs first, analogous to competing risks. These are referred to as competing events and examined in a event-specific manner. When \(s > 1\), the possible values \(S\) can take are written \(s \in \{1, ..., k\}\) in order of increasing severity. Incident reporting of these events is observed in a set of two independent groups labeled by \(g\), where \(g=0\) is a comparison group for \(g=1\), over \(m \ge 2\) pre-specified monitoring periods.
For a given monitoring period, group, and subtype event, \(T\) is the true time to incident reporting for that event only. In addition, \(C\) is the time to censoring not due to any competing event. We observe \(Y= \text{min}(T, C)\) and know whether censoring or reporting of the event happened first, which is denoted by \(\Delta = I(T \le C)\). Our observed data are therefore of the form \(\{(Y_1, S_1\Delta_1),...,(Y_n, S_n\Delta_n)\}\) for the \(n\) total events reported within the monitoring period. It is possible that multiple incident events could be reported at the same time. Ordered discrete event reporting times are \(t_1 < t_2 < ... < \tau\), where \(\tau\) is the end time of the monitoring period. A single time of event reporting in a monitoring period is denoted \(t_i\). Finally, when comparing across groups or monitoring periods we add a respective \(g\) or \(m\) subscript to the notation above. Sequential monitoring periods are written as \(m-1\) and \(m\).
We use a set of reference events, or HCCs distinct from the HCCs examined for incident reporting, to estimate underreporting. In theory, reference HCCs reported in one monitoring period should continue to be reported at equal rates in subsequent monitoring periods. However, in practice some reference events (and events in \(g=0\) more broadly) may be underreported. Reference events are denoted as a vector \(S^*\) of \(h\) distinct HCCs, which can take values \(s^* \in \{1, ..., h\}\). None of these reference events have competing events. The persistence of a given reference event \(s^*\) in monitoring period \(m\), meaning the proportion of individuals coded with \(s^*\) in monitoring period \(m\) who were previously coded with \(s^*\) in monitoring period \(m-1\), is \(q_{s^*,m}\) in line with prior work\textsuperscript{36}.
Before describing our estimands for incident event reporting, we first describe some key components. Within a group \(g\) and monitoring period \(m\), \(F(t) = P(T \le t)\) is the cumulative distribution function of \(T\) for all events beginning in that monitoring period. The overall hazard \(\lambda (t) = P(T = t \mid T > t)\) can be defined in terms of \(F(t)\) as \(\lambda (t) = (F(t) - F(t-1))/(1-F(t-1))\). In addition, \(\overline{F}(t) = 1 - F(t) = P(T > t) = \prod_{t_i \le t} (1 - \lambda(t_i))\), which is equivalent to the overall survival function in standard TTE analyses\textsuperscript{9}.
The event-specific hazard is analogous to a cause-specific hazard in TTE analyses with competing risks. For event \(s\) within group \(g\) and monitoring period \(m\), the event-specific hazard at \(t_i\) is \(\lambda_s (t_i) = P(T = t_i, S = s \mid T \ge t_i)\). Then, the corresponding event-specific cumulative incidence is \(F_s(t) = P(T \le t, S = s) = \sum_{t_i \le t} \overline{F}(t_i)\lambda_s(t_i) = \sum_{t_i \le t} \theta (t_i)\)\textsuperscript{9}. Thus, the mean time without event \(s\) is \(\mu_s(\tau) = \sum_{t_i < \tau} (t_{i+1} - t_i)F_s(t_i)\)\textsuperscript{10}. If we are comparing across sequential monitoring periods or groups, we write \(\mu_{s,m}(\tau)\) to specify the monitoring period \(m\) and \(\mu_{s,g}(\tau)\) to specify group \(g\).
Right censoring is assumed to be non-informative except for when an individual is censored due to a competing event. When an individual is recorded as having an event \(s\) that has any competing events (e.g., if \(s > 1\)), they leave the risk set to be coded for any other competing event. Both \(g=1\) and \(g=0\) groups are expected have events recorded at all equivalent reporting times within the monitoring period. We also impose that the overall sample is fixed across all monitoring periods.
The \(s^* \ge 1\) reference events occur in \(g=0\) only. So, in the comparison group \(g=0\) only and two sequential monitoring periods \(m-1\) and \(m\), our estimand for the underreporting proportion \(\epsilon\) is one minus the average persistence of all reference events, or \(\epsilon = 1 - \frac{1}{h} \sum q_{s^*,m}\). Multiple pairs of sequential monitoring periods are distinguished by adding a subscript, \(\epsilon_m\), where \(m\) denotes the second monitoring period in a pair.
We are first interested in estimating the difference in incident reporting for event \(s\) across the two groups within one monitoring period. So, our estimand is \(\psi = \mu_{s,g=1}(\tau) - \mu_{s,g=0}(\tau)\). In a monitoring period where \(\epsilon\) is known, the estimand is modified to account for underreporting by shifting the event-specific cumulative incidence curve in the comparison group \(g=0\) by \(\epsilon\), or \(\mu^*_{s,g=0}(\tau) = \sum_{t_i < \tau} (t_{i+1} - t_i)(F_s(t_i) + \epsilon)\). This is written as \(\psi^* = \mu_{s,g=1}(\tau) - \mu^*_{s,g=0}(\tau)\).
Across the two groups in two sequential monitoring periods, the reported difference in mean time without event that does not account for underreporting is \(\psi_{M} = \psi_m - \psi_{m-1}\). \(\psi_m\) and \(\psi_{m-1}\) are the same as \(\psi\) but with an additional subscript to specify the monitoring period. This estimand can also be adjusted to account for underreporting when \(\epsilon_m\) and \(\epsilon_{m-1}\) are known, becoming \(\psi_M^* = \psi^*_m - \psi^*_{m-1}\).
Possible severity-based upcoding can also be estimated within a monitoring period. Here, \(s \ge 2\), where the least severe subtype corresponds to \(s = 1\) and the most severe subtype to \(s = k\). In one monitoring period, possible severity-based upcoding across groups is \(\omega = \omega_{g=1} - \omega_{g=0}\), where each \(\omega_{g} = \mu_{s=k,g}(\tau) - \mu_{s=1,g}(\tau)\). This is equivalent to comparing the difference in incident reporting of the most severe event versus the least severe event across groups. If such reporting is higher in \(g=1\) compared with \(g=0\) (e.g., \(\omega > 0\)), then severity-based upcoding may be occurring.
In order to introduce the estimator of \(\mu_s(\tau)\), we first describe the estimator for the event-specific hazard at \(t_i\): \(\hat{\lambda}_s(t_i) = d_{s,i}/r_i\), where \(d_{s,i}\) is the count of event \(s\) reported at \(t_i\) and \(r_i\) is the number of individuals at risk at \(t_i\), meaning the number of individuals who have not yet been recorded as having \(s\) or any competing event by \(t_i\). \(d_i = \sum d_{s,i}\) is the count of all competing events reported at \(t_i\). We also have the event-specific cumulative incidence estimator: \(\hat{F}_s(t) = \sum_{t_i \le t} \hat{\theta}(t_i)\), with \(\hat{\theta}(t_i) = \widehat{\overline{F}}(t_i) \times \hat{\lambda}_s(t_i)\), where \(\widehat{\overline{F}}(t) = \prod_{t_i \le t}(1 - \hat{\lambda}(t_i)) = \prod_{t_i \le t}(1 - (d_i/r_i))\) is the Kaplan-Meier estimator for \(\overline{F}(t) = \prod_{t_i \le t}(1 - \lambda(t_i))\)\textsuperscript{9,46}.
The variance for \(\hat{F}_s(t)\) is \(\widehat{\text{var}}(\hat{F}_s(t)) = \sum_{t_l \le t} \widehat{\text{var}}(\hat{\theta}(t_l)) + 2 \sum_{t_l < t} \sum_{t_l < t_i \le t} \widehat{\text{cov}}(\hat{\theta}(t_l), \hat{\theta}(t_i))\), and covariance is \(\widehat{\text{cov}}(\hat{F}_s(t), \hat{F}_s(u)) = \widehat{\text{var}}(\hat{F}_s(t)) + \sum_{t_l \le t} \sum_{t < t_i \le u} \widehat{\text{cov}}(\hat{\theta}(t_l), \hat{\theta}(t_i))\). Here \(t_l\) indicates an event reporting time prior to \(t_i\) and \(u\) is an arbitrary time distinct from \(t\). \(\hat{\theta}(t_i)\) has variance \(\widehat{\text{var}}(\widehat{\theta}(t_i))= (\widehat{\theta}(t_i))^2 \times ((r_i - d_{si})/(d_{si}r_i) + \sum_{t_l < t_i} (d_l/(r_l(r_l - d_l))))\) and covariance \(\widehat{\text{cov}}(\hat{\theta} (t_i), \hat{\theta} (t_l)) = \hat{\theta} (t_i) \hat{\theta} (t_l)(-(1/r_i) + \sum_{t_l < t_i} (d_l/(r_l(r_l - d_l))))\)\textsuperscript{10}, where \(d_l = \sum d_{s,l}\) is the reported count of events of any subtype at \(t_l\), \(d_{s,l}\) is analogous to \(d_{s,i}\), and \(r_l\) is the count of beneficiaries at risk at \(t_l\).
We can now define the estimator of \(\mu_s(\tau)\) based on prior literature\textsuperscript{10}, as this is a component of the novel estimators we develop next: \(\hat{\mu}_s(\tau) = \sum_{t_i < \tau} (t_{i+1} - t_i)\hat{F}_s(t_i)\). This estimator has variance \(\widehat{\text{var}}(\hat{\mu}_s(\tau)) = \sum_{t_i < \tau} (t_{i + 1} - t_i)^2 \widehat{\text{var}}(\hat{F}_s(t_i)) + 2 \sum_{t_i < \tau} \sum_{t_l < t_i} (t_{i+1} - t_i)(t_{l+1} - t_l) \widehat{\text{cov}}(\hat{F}_s(t_i), \hat{F}_s(t_l))\).
For each reported reference event \(s^*\) in \(g=0\) we first compute persistence, or the proportion of individuals reported as having \(s^*\) in monitoring period \(m\) who where also reported as having \(s^*\) in monitoring period \(m-1\). This is then averaged across all reference events to obtain \(\hat{\epsilon} = 1 - \frac{1}{h} \sum \hat{q}_{s^*,m}\).
We propose that the difference in mean time without an event, \(\psi\), is estimated across groups as \(\widehat{\psi} = \hat{\mu}_{s,g=1}(\tau) - \hat{\mu}_{s,g=0}(\tau)\) with variance \(\widehat{\text{var}}(\widehat{\psi}) = \widehat{\text{var}}(\hat{\mu}_{s,g=1}(\tau)) + \widehat{\text{var}}(\hat{\mu}_{s,g=0}(\tau))\), as groups are assumed to be independent. To adjust for underreporting, \(\hat{\mu}_{s,g=0}(\tau) = \sum_{t_i < \tau} (t_{i+1} - t_i) \hat{F}_{s}(t_i)\) is modified to \(\hat{\mu}_{s,g=0}(\tau) = \sum_{t_i < \tau} (t_{i+1} - t_i) (\hat{F}_{s}(t_i) + \epsilon)\) in the estimator for \(\widehat{\psi}\), which does not change the variance. The underreporting-adjusted estimator is denoted \(\hat{\psi}^*\).
Across sequential monitoring periods, \(\psi_M\) is estimated as \(\hat{\psi}_M = \hat{\psi}_m - \hat{\psi}_{m-1}\) with variance \(\widehat{\text{var}}(\widehat{\psi}_M) = \widehat{\text{var}}(\widehat{\psi}_m) + \widehat{\text{var}}(\widehat{\psi}_{m-1})\). Although the components of this estimator come from the same fixed population, the correlation between component estimates is assumed to be negligible because we are comparing differences in RMTL across independent groups estimated in disjoint time intervals. Therefore, covariance is zero. When an underreporting estimate is known for both monitoring periods, this estimator becomes \(\hat{\psi}_M^* = \hat{\psi}^*_m - \hat{\psi}^*_{m-1}\), and variance remains unchanged.
Our estimator of possible severity-based upcoding across groups, \(\omega\), is given by: \(\widehat{\omega} = \widehat{\omega}_{g=1} - \widehat{\omega}_{g=0}\), where each \(\widehat{\omega}_{g} = \hat{\mu}_{s=k}(\tau) - \hat{\mu}_{s=1}(\tau)\) within one monitoring period. The variance is \(\widehat{\text{var}}(\widehat{\omega}) = \widehat{\text{var}}(\widehat{\omega}_{g=1}) + \widehat{\text{var}}(\widehat{\omega}_{g=0})\), where \(\widehat{\text{var}}(\widehat{\omega_g}) = \widehat{\text{var}}(\hat{\mu}_{s=k}(\tau)) + \widehat{\text{var}}(\hat{\mu}_{s=1}(\tau)) - 2\widehat{\text{cov}}(\hat{\mu}_{s=k}(\tau), \hat{\mu}_{s=1}(\tau))\) and \(\widehat{\text{cov}}(\hat{\mu}_{s=k}(\tau), \hat{\mu}_{s=1}(\tau)) = \sum_{i} \sum_j \widehat{\text{cov}}(\hat{F}_k(t_{i-1}), \hat{F}_1(t_{j-1}))\) for distinct event times \(t_i\) and \(t_j\). Here, \(t_i\) corresponds to event reporting increments in the cumulative incidence function for severity level \(k\), while \(t_j\) corresponds to the equivalent for severity level \(1\). Although event reporting times are the same, we write these using separate variables because covariance is estimated between all pairs of increments for these two severity levels.
Upcoding in Medicare has been inconsistently examined in the medical and economics literature for several decades\textsuperscript{31}. One reason for this is that it is difficult to definitively identify upcoding, particularly given that researchers and policymakers typically only have access to national Medicare claims data without beneficiaries' corresponding electronic health records or other data. There are also a number of barriers to the development of upcoding estimation approaches. First, gaining access to individual-level Medicare data is a time-consuming process that is inaccessible to many researchers. Second, there are limited national data resources describing co-occurring health conditions in older adults that are free of coding incentives\textsuperscript{42--44}. Third, labeled upcoding data to evaluate estimators are not available to researchers. This also means that comparing proposed methods is challenging, as the data used to develop or evaluate such methods are often not able to be shared by researchers.
We developed the open source upcoding R package (\url{https://github.com/StanfordHPDS/upcoding}) to help address many of these issues. The package enables simulation of longitudinal coding data for a Medicare-eligible population as well as more reproducible evaluation of approaches for evaluating HCC coding. Features include:
We describe this functionality in further detail in the subsections below. A brief tutorial on specific package functions is available in the package's Github repository.
To simulate realistic co-occurring HCCs not influenced by coding incentives, co-occurring self-reported health conditions from participants aged 65 years and older (i.e., Medicare eligible) were extracted from the National Institutes of Health's All of Us study\textsuperscript{48}. We used self-reported survey data to obtain co-occurring health conditions because other national datasets (e.g., Medicare billing claims) report diagnoses for billing purposes and therefore may be impacted by coding incentives\textsuperscript{42--44}. The All of Us study was designed to enroll a million participants across the United States, focusing especially on groups historically underrepresented in clinical and biomedical research\textsuperscript{48} and included questions that overlapped with many HCCs in the current version of CMS-HCC (Version 28, or V28).
V28 HCCs (listed in the Supporting Information) were manually mapped to Systemized Nomenclature of Medicine-Clinical Terms (SNOMED-CT) concept identifiers for survey questions (using version 7 of All of Us data; see this project's Github repository for mapping). Then, survey responses to all available SNOMED-CT concepts were queried from the All of Us database. Included surveys, available respondent sociodemographic characteristics, and coverage of V28 HCCs are described in the Supporting Information.
For each surveyed person, self-reported co-occurring V28 HCCs were extracted. This was then summarized into a table of unique co-occurring HCC sets and a count of the number of All of Us survey respondents that reported having these diagnoses. In line with the All of Us Data and Statistics Dissemination Policy, only sets of co-occurring HCCs with more than 21 respondents were exported from All of Us's platform (the Researcher Workbench) to be used in analyses. This both omits all sets of 20 or fewer respondents and also excludes several sets of co-occurring HCCs with more than 20 respondents. Ultimately, baseline data were simulated by sampling with replacement from these sets of co-occurring HCCs, where the respondent count is used to weight sampling. Users also have the option to use alternate sets of co-occurring V28 HCCs for baseline data sampling if they prefer.
The package enables users to simulate two types of upcoding realistic to MA: any-available or severity-based, although the latter can only occur if an HCC has competing events. This upcoding is randomly split across a user-specified number of time points. Even though there is inherent right censoring in upcoded data (because only a proportion of all simulated individuals will be upcoded or coded at all), the package provides users with the option to include additional right censoring. This is meant to be representative of loss to follow up or death, both of which are known issues in following a population aged 65 years and older over time\textsuperscript{47}. Users specify the proportion of loss to follow up they want to include, and a randomly selected set of rows (e.g., individuals) are right censored over each time point. Once someone is lost to follow up, they cannot be coded for any HCC in subsequent time periods.
The package also separately enables undercoding. Here, users specify a proportion of all coded diagnoses (from simulated baseline data) to randomly remove across the entire data set. This occurs on a dataset-wide level because undercoding is a systemic issue in TM\textsuperscript{28,33} and so we do not assume that it impacts specific V28 HCCs disproportionately.
Our simulation study demonstrates how the estimators we propose can be used to monitor Medicare coding behaviors. To do this, different upcoding and underreporting scenarios for the estimator introduced in Section (ref) were compared to an estimator currently used by policymakers, defined later in Section (ref). Degrees of upcoding and undercoding were constructed to align with current estimates of these issues in MA and TM from the literature. In addition, the temporal structure was implemented to be analogous to quarterly reporting for the length of time (around two years) that a given risk adjustment formula version is typically in place. Each individual scenario was replicated 1000 times.
In each replicate, two sets of 1,000,000 observations of baseline data were independently simulated using our upcoding R package. For each dataset, the columns are the V28 HCCs (listed in the Supporting Information). Both any-available and severity-based upcoding were implemented in the first baseline data set (the MA-like data). For the former type of upcoding, an HCC without any competing HCCs was upcoded. As an illustration, we use HCC238 (Specified Heart Arrhythmias). For the latter, an HCC that has lower severity HCCs was upcoded, specifically HCC125 (Dementia, Severe). HCC125's less severe HCCs are HCC126 (Dementia, Moderate) and HCC127 (Dementia, Mild or Unspecified). The scenarios implemented were as follows:
Scenario 1: Upcoding of an HCC that lacks competing events (HCC238) with varying underreporting in the comparison group. Any observation not previously coded with HCC238 was eligible to be upcoded. Baseline MA data were separately upcoded to varying degrees (20%, 25%, 30%) sequentially within each monitoring period. The comparison group (i.e., TM data) was simulated by first undercoding the baseline TM data to varying levels (0%, 5%, 10%, 15%) and then upcoding HCC238 analogously but to a lower amount (5%) per monitoring period.
Scenario 2: Upcoding of an HCC with lower-severity competing events (HCC125) with varying underreporting in the comparison group. Only observations previously coded with the lower severity HCCs (HCC126 or HCC127) were eligible to be upcoded. Baseline MA data were separately upcoded to varying degrees (20%, 25%, 30%) sequentially within each monitoring period. The comparison group was simulated by first undercoding the baseline TM data to varying levels (0%, 5%, 10%, 15%) and then upcoding HCC125 analogously but to a lower amount (5%) per monitoring period.
We compared our estimators to a coding intensity estimator similar to the Demographic Estimate of Coding Intensity (DECI), which is widely used by Medicare policymakers\textsuperscript{34,49}. DECI assumes complete data (e.g., no censoring) and has the following formula:
\[
\] Here, the numerator is a risk score estimated with all HCCs included in the CMS-HCC formula, while the demographic-only risk adjustment risk score is estimated using only the demographic variables in that formula. Importantly, this also means that the estimand targeted by DECI is different from that of our proposed estimators. Further, as we did not have CMS-HCC risk score coefficients or demographic variables available in our simulated data, our comparator estimator was more precisely DECI-like and we defined it as: \[
\] Our \(\text{DECI}^\dagger\) comparator is estimated in the same simulated data described earlier in this section, but relies on counts of all HCCs at the end of each monitoring period rather than time to incident coding of individual HCCs. Although DECI has a different target estimand than our estimators and ignores censoring, comparing our simulation results to the DECI-like estimator \(\text{DECI}^\dagger\) is useful as it illustrates how our estimators can be a complement to existing practice.
Cumulative incidence of coding for scenario 1 given 20% any-available upcoding in MA and 5% any-available upcoding in TM after all four degrees of undercoding is shown in Figure (ref). As expected, within each monitoring period the cumulative incidence of the upcoded group was higher than the comparison group, which was only upcoded 5%. We also saw that the impact of undercoding on cumulative incidence estimates was limited. Given that undercoding occurs across the entire set of 115 HCCs, it was less likely to notably influence estimates for any single HCC. Lastly, we observed that the gaps between upcoded and comparison groups become smaller in sequential monitoring periods, which is a consequence of the fixed sample we imposed across all monitoring periods. As monitoring periods increased, the number of individuals available to upcode decreased. Results plots for additional degrees of upcoding (25%, 30%) and Scenario 2 upcoding (severity-based upcoding for HCC125) are in the Supporting Information.
Estimates corresponding to the Section (ref) estimators for upcoding of HCC238 are presented in Figure (ref), which examines the difference in time without reported incident HCC238 coding between MA and TM within each monitoring period. Examining the first monitoring period (M1), estimates clearly increase with the degree of upcoding. Similarly, for monitoring period 2 (M2) alone, estimates also increase with the degree of upcoding.
In addition, undercoding has a limited effect on estimates. However, researchers also have the option to adjust for underreporting using the estimator proposed in Section (ref). As the sample is fixed across monitoring periods, estimates across monitoring periods within a degree of upcoding and undercoding decrease, analogously to Figure (ref). The gap between M1 and M2 increases with degree of upcoding as well, because more aggressive incident upcoding in M1 limits the availability of individuals for incident upcoding in M2.
This suggests that within a monitoring period, researchers could compare estimates of different HCCs to obtain a ranked list of HCCs to study further, where HCCs with the highest estimates have the most incident coding. In addition, in this scenario, being in TM appeared to have a “protective effect” against being coded with HCC238, although supplemental analyses would be needed to account for possible confounding and other issues. Additional results plots for severity-based upcoding of HCC125 are available in the Supporting Information.
Figure (ref) shows the average \(\text{DECI}^\dagger\) estimate across simulation replicates by monitoring period as well as upcoding and undercoding degree. As degree of upcoding increases, these estimates increase negligibly both for individual monitoring periods and across sequential monitoring periods. This is intuitive--only two of the 115 HCCs in the MA risk adjustment algorithm are being upcoded--but it suggests a limitation of DECI in identifying upcoding occurring in a minority of HCCs. Regardless of upcoding degree, the \(\text{DECI}^\dagger\) estimate increasing in line with the degree of TM undercoding indicates that it may be more sensitive to undercoding in TM. Especially since we know undercoding is prevalent in TM\textsuperscript{28,33}, this suggests that DECI may be misrepresenting the overall spending gap between MA and TM due to its undercoding sensitivity.
We proposed a set of estimators that extend RMTL methodology for evaluating time to incident coding by private insurers in Medicare. Given the timeline of risk adjustment formula updates, our approach realistically presupposes that reported coding is examined over multiple monitoring periods, each of which has several time points where new events are reported. Novel estimators were introduced to evaluate differences in time-to-reporting both within monitoring periods and across monitoring periods, including an estimator for severity-based upcoding. Our approach also included an adjustment for possible underreporting, which is a known major issue in TM data\textsuperscript{28,33}. Finally, we developed an open source R package that enables users to simulate co-occurring HCCs free of coding incentives and similar to those reported by individuals residing in the United States eligible for Medicare. Users can also undercode or upcode this data over time.
Our simulated data were upcoded to degrees aligned to those reported in literature\textsuperscript{32}, and we found in our simulation results that our estimators were able to recover differences in upcoding both within and across monitoring periods while showing limited sensitivity to undercoding. DECI-like estimates of the same data were a useful complement, but were limited in that these comparator estimates were very sensitive to undercoding and did not vary when upcoding occurred in a minority of HCCs. Therefore, our estimators show considerable promise as tools to help evaluate reported coding patterns over time. These estimators can be used as a first step to identify potentially upcoded HCCs while making more realistic assumptions than DECI. Follow up analyses could include examining coding patterns within specific insurers, providers, or beneficiaries.
This work has a number of limitations and areas for future development. First, several assumptions could be further relaxed. This includes our assumptions that the population is fixed across both comparison groups and monitoring periods and that estimates in sequential monitoring periods across groups are approximately independent. Upcoding estimators could be expanded to additional types of upcoding and to correct for underreporting. Covariates could also be incorporated into estimation, which would be useful for addressing issues like confounding. There is also more heterogeneity and missingness in real-world claims data than was included in our simulations. Self-reported diagnoses--as we use in our simulated baseline data--also have recall bias and other issues that we do not address here.
Since the MA program's inception, high intensity of private insurer coding, including possible upcoding, is estimated to have cost the federal government and taxpayers \$224 billion dollars without clear benefit to beneficiaries\textsuperscript{32}. Thus, estimating differences in coding between MA and TM or more and less severe competing HCCs has the potential to help policymakers locate issues with specific HCCs earlier and at scale. This could suggest areas for improvement in the risk adjustment formula and potentially save significant funds in the Medicare program. Our work has provided both novel estimators and novel simulated data tailored to addressing these important policy considerations.
This research was funded by National Institutes of Health Director's Pioneer Award DP1LM014278 and a Stanford Interdisciplinary Graduate Fellowship. We thank Malcolm Barrett, Gabriela Basel, Lizzie Kumar, and Neha Srivathsa for their valuable insights and contributions to code performance and review. We gratefully acknowledge All of Us participants for their contributions, without whom this research would not have been possible. We also thank the National Institutes of Health's All of Us Research Program (\href{http://127.0.0.1:19869/\#0}{https://allofus.nih.gov/}) for making available the participant survey data included in this study. We thank Stanford University and Stanford Research Computing for providing additional computational resources, including support for computing performed on the Sherlock cluster.
The simulated data for this study can be generated using code in the project repository: \url{https://github.com/StanfordHPDS/tte_estimation_medicare}. Additional upcoding, undercoding, and baseline data can be simulated using the upcoding package (\url{https://github.com/StanfordHPDS/upcoding}). In line with program policies, non-summarized version 7 All of Us survey data used to derive baseline co-occurring HCCs can be accessed by authorized users via the All of Us Researcher Workbench only (\url{https://www.researchallofus.org/data-tools/workbench/}) and requires Registered Tier access. Code that was used for this project in the Researcher Workbench is in the above project repository.
Additional study summary information and results can be found online in the Supporting Information.
\addcontentsline{toc}{section}{References}
\phantomsection