Medicare star ratings (health policy)

The content on this page was written by AI under human supervision.

Medicare rates private health plans, hospitals and patients’ experience of care with one to five stars, and for Medicare Advantage plans the stars set bonus payments worth billions of dollars a year. Six of these rating programs share one stated rule for placing the star boundaries: minimize the total squared distance of the scores from their bin centers. The paper behind this page computes that minimum exactly and finds that the published star bins often miss it.

Stars and money

The Centers for Medicare & Medicaid Services (CMS) publishes several star ratings; the paper examines six. Each Medicare Advantage contractthe unit in which Medicare’s private health plans are rated and paid; the paper identifies contracts only by CMS contract number and names no plan or hospital receives an overall rating each year, and a contract rated 4.0 stars or above earns a bonus under the Quality Bonus Program. The benchmark it is paid against rises by five percent per enrollee the following year; a contract at 3.5 gets no increase. These bonuses are estimated to total $13.4 billion for 2026. The Overall Hospital Quality Star Rating gives each hospital a single star rating on Medicare’s Care Compare site. Four patient-experience surveys of the CAHPS family (Consumer Assessment of Healthcare Providers and Systems) assign further stars to hospitals, home health agencies, hospices and dialysis facilities. Only the Medicare Advantage rating affects payments directly; the other five inform patients’ choices.

The Medicare Advantage ratings are also the subject of litigation, in which the paper takes no side. A 2024 ruling in a suit brought by SCAN Health Plan led CMS to recalculate that year’s ratings nationwide, and SCAN and Alignment Healthcare filed new suits in July 2026. The closest precedent for examining the clustering step itself is Bo Bayles’s 2017 finding that CMS’s hospital code stopped before its iteration had converged, which changed hundreds of ratings. Neither that work nor anything the author found in the literature, court dockets or public code tested the published cut points against the stated criterion.

The criterion and the algorithms

All six programs group the scores, placed on a line, into five bins, the star levels, so that the scores lie as close as possible to their bin centers; the regulations call the quantity to be minimized the within-cluster sum of squares,

$$\mathrm{SSQ} \;=\; \sum_{b=1}^{k} \sum_{i \in B_b} \bigl(x_i - \bar{x}_b\bigr)^2 ,$$

where $x_1\le\dots\le x_n$ are the sorted scores, $B_1,\dots,B_k$ the bins ($k=5$) and $\bar{x}_b$ the mean score in bin $b$; each term is one score’s squared distance from its own bin’s mean. For Medicare Advantage the criterion is regulation text: 42 CFR 422.162(a) states that “The Star Ratings levels are assigned to the clusters that minimize the within-cluster sum of squares,” a sentence CMS’s annual Technical Notes repeat. A separate provision, 422.166(a)(2)(i), prescribes the method instead, “mean resampling with the hierarchal clustering.” The hospital methodology prescribes k-means clustering, whose objective is the same sum, and the four survey programs prescribe Ward’s hierarchical clustering with the same stated aim.

Neither Ward’s method nor k-means is guaranteed to reach the minimum. Ward’s method starts with every score in its own cluster and repeatedly merges the two clusters whose merger raises SSQ least; merging clusters $a$ and $b$, holding $n_a$ and $n_b$ scores with means $\bar{x}_a$ and $\bar{x}_b$, raises it by

$$\Delta(a,b) \;=\; \frac{n_a\, n_b}{n_a+n_b}\,\bigl(\bar{x}_a-\bar{x}_b\bigr)^2 ,$$

and merging stops at five clusters. A merge is never undone, even when the best final grouping would separate the scores it joined. K-means, used for hospitals, starts instead from initial bin centers fixed by the code, moves a bin boundary whenever a single move lowers SSQ, and stops at a local optimum that depends on where it started.

For data of arbitrary dimension, minimizing SSQ is NP-hard, meaning that no known fast algorithm is guaranteed to find the minimum, so fast heuristics are used instead. On a line, an optimal grouping consists of consecutive runs of sorted scores, and the only question is where to put the $k-1$ cut points among the $n-1$ gaps. In 1958 Walter D. Fisher showed how to search those choices without repeating work, by dynamic programming. Writing $D_m(i)$ for the smallest SSQ achievable when the first $i$ sorted scores fill $m$ bins,

$$D_m(i) \;=\; \min_{m\le j\le i}\bigl\{\, D_{m-1}(j-1) + \mathrm{SSQ}(x_j,\dots,x_i) \,\bigr\} ,$$

where $\mathrm{SSQ}(x_j,\dots,x_i)$ is the sum of squares of a single bin holding scores $j$ through $i$. In words, the best split of $i$ scores into $m$ bins is the best split of a shorter initial run into $m-1$ bins plus one final bin, minimized over where that bin begins. The final value, $D_k(n)$, is the exact minimum, and a variant lists every grouping attaining it, which shows whether the optimum is unique. The paper runs this and every comparison in exact rational arithmetic.

Animated cartoon. Twenty-one black dots, toy scores, sit on a horizontal line labeled scores on a line. Ward’s method runs first: translucent slate capsules mark the current clusters, the two clusters about to merge are outlined in orange, and a counter at the top right falls from 21 clusters to 5 while the running SSQ rises to 438.0. A slate ruler then appears above the dots, divided into five numbered star bins with solid edges dropping to the line and SSQ 438.0 printed at its right. Below the dots a violet ruler rises, the five bins of the exact optimum from Fisher’s dynamic program, with dashed edges and SSQ 351.7. Two of the four dashed edges lie to the left of their solid partners; the four dots caught between a solid edge and its dashed partner turn orange, with the note that those 4 scores move up a star. The picture fades back to the bare dots and repeats.

The two computations on 21 made-up scores (a toy example drawn for this page, not CMS data; bins are numbered 1 to 5 from low scores to high). Ward’s method merges the pair of clusters whose union raises SSQ least, one merge at a time and never undoing one, until five remain; the slate ruler above the dots shows the resulting star bins, with solid edges, and their SSQ. Fisher’s dynamic program then finds the grouping with the smallest SSQ of all, the violet ruler below, with dashed edges. Here the greedy merges stop at SSQ 438.0 against a minimum of 351.7, two of the four optimal edges lie to the left of Ward’s, and the four scores between a solid edge and its dashed partner (orange) sit one bin higher under the optimum. The figures below draw CMS’s published bins and the exact optimum in the same way.

The hospital ratings

The hospital program is the most open of the six: CMS publishes the code (in the statistical languages SAS for rating years 2021–2025 and R for 2026), the input files and the stars, and no step is randomized. Run unmodified, the 2026 R package matches the published Care Compare star for all 3,182 hospitals that have a star in both its output and the published file. The SAS packages, with one proprietary k-means routine reimplemented from its documentation, reproduce every published star of 2021–2025. (Code and inputs come from public GitHub mirrors of CMS’s releases.) Hospitals are binned within peer groupspeer group 3, 4 or 5: hospitals reporting enough measures to be scored in three, four or all five of the method’s measure groups; in 2026 the groups held 177, 749 and 2,277 hospitals. In 2026 the exact optimum on the same scores assigns a different star to 205 of the 3,182 displayed hospitals, 6.4 percent, every one of them higher. All twelve published cut points that year lie above their optimal counterparts, so the hospitals that differ are those caught between a published cut point and its optimal partner (shown below for one peer group).

Hospital peer group 3 of the 2026 Overall Hospital Quality Star Rating, 177 hospitals, each drawn as a black dot at its standardized summary score along a horizontal axis running from about minus 2.8 to plus 1.4; the vertical scatter of the dots is jitter and carries no meaning. Above the dots a gray ruler divided into five numbered boxes shows the star bins CMS published (labeled CMS shipped bins in the figure), with solid vertical lines dropping from its four bin edges and a thick tick in each box at the bin's mean score; its SSQ, 8.0613, is printed at the right. Below the dots a blue ruler shows the five bins of the exact optimum (labeled certified-optimal bins in the figure), with dashed vertical edges and its SSQ, 7.6847, at the right. Each dashed blue edge lies to the left of the corresponding solid gray edge. A small bar chart underneath compares the number of hospitals in each star bin under the two groupings: the optimum has fewer hospitals at one, two and three stars and more at four and five.

One of the three hospital peer groups in 2026 (177 hospitals), each hospital a dot at its summary score; the vertical spread is jitter only. The gray ruler above shows the five star bins CMS published, with solid edges; the blue ruler below shows the bins of the exact optimum of the stated criterion, with dashed edges (the figure’s own labels are “CMS shipped bins” and “certified-optimal bins”). Thick ticks mark each bin’s mean score, and each grouping’s SSQ is printed at the right: 8.0613 published, 7.6847 optimal. Every published edge lies to the right of its optimal partner, so the 55 hospitals between a solid line and its dashed partner receive one star fewer than the optimum assigns; the bar chart beneath shows the bin populations shifting upward. Across all three peer groups 213 of 3,203 rated hospitals differ at the clustering step, 208 after a safety cap that holds some hospitals to four stars, and 205 among the 3,182 with a displayed star. (Figure 1 of the paper.)

The 2026 k-means iteration does converge, after two steps in each peer group. The published grouping is therefore a genuine local optimum, fixed by the code’s starting centers, whose SSQ still exceeds the minimum; in the 2017 case the iteration had stopped early. Over the six rating years the optimum differs from the published grouping in 1,146 of 18,493 stars issued and is unique in all eighteen peer-group computations. The paper adds that the criterion is nearly flat near the optimum: in 2026 most of the 213 hospitals that differ at the clustering step regain their published star under some grouping nearer the optimum than CMS’s. The direction of the differences varies by year (in 2024, 15 up and 258 down), and the paper presents the all-upward pattern of 2026 as a fact about that year, not a rule.

The Medicare Advantage cut points

In Medicare Advantage, CMS clusters contract scores separately for each quality measure, so every measure in every star year has its own cut points, the thresholds that determine whether an 87 percent screening rate is worth three stars or four. The paper calls one such calculation a computation; star years 2024–2026 contain 123. Under the Technical Notes, outlying scores beyond a quartile-based fence are trimmed and a SAS procedure assigns the remaining contracts at random to ten equal groups. Ten times, one group is set aside and the rest are clustered by Ward’s method, and the ten values of each cut point are averaged; this is the regulation’s “mean resampling,” a delete-one-group jackknifea resampling scheme that recomputes a statistic with one block of the data left out at a time and averages the results. CMS then applies a guardrailCMS’s term for the cap on how far a measure’s cut point may move from its prior-year value; in the worked example below the window is the prior-year cut point plus or minus 0.04 capping each cut point’s movement from the prior year, rounds, and aggregates the resulting measure stars into each contract’s overall rating.

CMS does not publish which scores entered each computation, so the paper rebuilds each clustering sample from the public score files under the documented exclusion rules. The two fence values the Notes print for every computation serve as a check: the rules alone reproduce them in 82 of the 123 computations, and for the rest the paper fits score profiles for unpublished contracts. On each sample it runs the clustering step both ways, Ward’s method and Fisher’s dynamic program, before any resampling, guardrail or rounding. The Ward grouping’s SSQ is strictly above the minimum in 104 of the 123 computations (104 to 106 across three reconstructions of the samples, and still 103 when every fitted profile is dropped) and equal to it in 16; three are degenerate (all scores identical, or too few distinct values). In all 120 non-degenerate computations the optimal grouping is unique. The gaps, plotted below, span seven orders of magnitude. When the contracts are rescored against the optimal cut points, 9,809 of the 47,752 contract-measure stars assigned by these computations change, about one in five.

Three side-by-side panels headed 2024, 2025 and 2026. In each, blue dots plot, for every cut-point computation of that Medicare Advantage star year, the SSQ of CMS's Ward grouping minus the exact minimum, on a logarithmic vertical axis running from ten to the minus four up past a thousand, with the computations sorted along the horizontal axis from the largest gap to the smallest. The dots descend from a plateau near a few hundred down through one and below; labels read max 1813.0 and 36/40 off optimum for 2024, max 1430.7 and 33/40 off optimum for 2025, and max 1658.1 and 35/43 off optimum for 2026. Below an axis break, open gray circles at a row labeled 0 (exact) mark the computations whose gap is exactly zero (19 in all across the three years).

For each of the 123 cut-point computations of star years 2024–2026, the amount by which the SSQ of CMS’s Ward grouping exceeds the exact minimum on the same trimmed sample, one panel per year, sorted by size and drawn on a logarithmic axis because the gaps run from below 0.0002 to above 1,800 in each measure’s squared score units. Filled dots are the 104 computations that miss the minimum (36 of 40, 33 of 40 and 35 of 43 by year); open circles below the axis break are the 19 with gap exactly zero, the 16 that attain the minimum and three degenerate ones. The comparison is made before resampling, the guardrail and rounding, on clustering samples reconstructed from the documented rules with fitted score profiles. In most computations of every year the prescribed clustering step does not reach the minimum of the criterion stated in the program’s own rules. (Figure 2 of the paper.)

Across the bonus line

A different cut point changes a payment only if it moves a contract’s overall rating across 4.0. The paper runs the full-sample optimal cut points through the later published stages (guardrail, rounding, aggregation) and compares the resulting ratings with the stars CMS assigned, under seven stated conventions. One convention holds the improvement and survey measures at their published stars, since their inputs are unpublished or not clustered; the count of changed ratings is therefore conditional on those stars and could be larger or smaller. Sixty contract-years change bonus status. Of these, 34 are valued in dollars (31 distinct contracts; the 34 are among 1,724 published contract-year ratings), sixteen crossing upward and eighteen downward; the other 26 have no bonus program, no published overall rating, or a status the paper could not determine. At $400–600 per enrollee-year, a band the paper calls a reference value and not a measurement, the bonus payments that depend on these 34 changes total an estimated gross $0.9 to $1.4 billion over payment years 2025 through 2027. At the $1,146 million midpoint, $130.3 million comes from the upward changes and $1,015.3 million from the downward ones; the two are not netted. These are estimates, the paper stresses, and imply nothing about the conduct of any contract or of CMS.

The dollar total is concentrated: one contract-year, H0524 in star year 2024, accounts for 59.2 percent of the valued enrollment (1,355,332 enrollee-years, $542–814 million), and its published rating is only about 0.030 above the 3.75 rounding edge where the displayed star changes. Contracts move both ways because the optimal cut points are neither uniformly harder nor uniformly easier than CMS’s. The guardrail then determines which differences reach the stars (figure below): where the cap binds it clamps the cut point identically under both methods, and elsewhere the difference passes through. Of the 60 changes, 33 involve a clamped measure, and without the guardrail 18 would disappear, all upward.

Two panels, one above the other, (a) raw resampled cut-point means and (b) after guardrails, for the Medicare Advantage measure Complaints about the Health Plan in star year 2026. The horizontal axis lists the four star boundaries, 5|4, 4|3, 3|2 and 2|1; the vertical axis is complaints per 1,000 members from 0 to 1.4, lower being better. At each boundary a dark horizontal line marks the 2025 published cut point and a gray band around it marks the guardrail window of plus or minus 0.04. Blue circles show the resampled cut-point means from Ward's method with the documented tie rule and green diamonds those from the exact dynamic-programming optimum. In panel (a) both markers lie below the window at 4|3, 3|2 and 2|1, with a note that the clamp binds there, and inside the window at 5|4, labeled 0.108 / 0.119. In panel (b) the markers at 4|3, 3|2 and 2|1 have been moved to the lower edge of their windows and coincide, each labeled clamped, while at 5|4 they remain distinct, labeled 0.11 / 0.12, free: difference survives, with a note that the 0.01 difference at 5|4 moves 18 contracts up one measure star.

One Medicare Advantage measure, Complaints about the Health Plan (a Part C, or medical-plan, measure; star year 2026), followed through the guardrail. At each of the four star boundaries the dark line is the prior-year published cut point and the gray band the guardrail window around it, plus or minus 0.04. Blue circles are the resampled cut-point means from Ward’s method and green diamonds those from the exact optimum (labeled “certified DP optimum” in the figure). In the upper panel, before the guardrail, both methods fall below the window at three boundaries; in the lower panel, after it, those three are clamped to the window’s edge identically for both methods, and only the 5|4 boundary keeps its 0.01 difference (0.11 against 0.12), which in the illustrative group assignment drawn here moves 18 contracts up one measure star. The clamp acts identically at every assignment; the surviving difference and the count of 18 are specific to this one. Where the guardrail binds it removes the difference between the two methods, and where it does not, the difference reaches the stars. (Figure 3 of the paper.)

The unpublished input order

Two things missing from the published files keep CMS’s own resampled computation from being rerun: the exact set of scores that entered each computation, and the full specification of its one random step. The Technical Notes print the SAS call that assigns contracts to the ten groups, seed included (8675309, which “allows for future replication of the randomization process”). The assignment also depends on the unpublished order in which the contracts were read (and, in two of the years, on which random-number generator ran). The Notes do not spell out the assignment rule itself; the paper states the rule it assumes and, under that assumption, recomputes every contract’s bonus status at 28 sampled input orders. Under different orders, 118 contract-years fall on different sides of the bonus line, so the published files leave their status undetermined (a statement about the files, the paper adds, not about any contract).

The paper expresses the dependence on the input order in dollars, as the expected gross bonus payments, over one randomly drawn order, of the contracts whose status under that order differs from their usual one. With Ward’s method inside the ten groups, as CMS runs it, the expectation is $787.2 million; with the exact optimum inside the same groups it is $98.4 million, smaller by a factor of eight (each figure measures one method’s own variation with the order, not the difference between the methods, which at any single order still disagree on the bonus status of 26 to 47 valued contract-years). Publishing one production run’s clustering input file (each computation’s scores in the order read) or its group assignment would settle every such case; for the two earlier years the generator is also needed. The input order alone would not suffice, since each clustering sample is a reconstruction that can be checked against the published fence values but is not fixed by them.

The patient surveys

No payment depends on the survey stars. Two of the four programs publish the clustered scores; the home health and hospice surveys publish only stars and cannot be checked. For the hospital survey (HCAHPS), the paper examines five quarterly refreshes (April 2025 to May 2026) and the eight measures starred in all five. The published grouping misses the minimum in 38 of the 40 refresh-measure cells, and regrouping to the optimum moves 17,484 hospital-measure stars, in both directions. However, Ward’s method run on the public file recovers only one of the eight published groupings on the May 2026 refresh. If CMS runs the Ward step as the notes describe it, its inputs are therefore not exactly the file’s scores, and these divergences mix input differences with the algorithm’s shortfall. Patients see one summary star per hospital, built from the per-measure stars by a published rule; applied to the optimal per-measure stars, that rule changes the summary star of 413 of 3,176 hospitals (13.0 percent) on the May 2026 refresh.

For the dialysis survey (ICH CAHPS) the algorithm can be tested on its own. The survey’s April 2026 file contains the clustered quantity itself, a linearized 0–100 score for each of 2,594 facilities on six measures. Ward’s method applied to that file reproduces five of the six published groupings identically and the sixth up to ties, which confirms the inputs. None of the six published groupings attains the minimum, and each optimum is unique. Regrouping moves 4,061 facility-measure stars, about 26 percent of the 15,564 in the file, from 3 percent of facilities on one measure to 40 percent on each of the two shown below. On every measure the moves run one way only, up on five and down on one.

Two stacked panels for the dialysis-facility survey, ICH CAHPS, April 2026 file, each showing 2,594 facilities as black dots at their published linearized scores on a horizontal axis from about 45 to 100; because scores are integers the dots form vertical columns, and the vertical spread is jitter. Upper panel: Nephrologists' communication and caring (n = 2,594; 1,042 up), with a gray ruler of the five published bins above the dots, SSQ 10329.0, and a blue ruler of the exact-optimum bins below, SSQ 9116.5; every dashed blue edge lies left of its solid gray partner. Lower panel: Rating of the nephrologist (n = 2,594; 1,042 down), published SSQ 7129.8 against optimal SSQ 6224.7; here the lowest dashed blue edge coincides with its solid gray partner and the other three lie to the right of theirs. Thick ticks in each box mark bin means.

The two most-affected measures of the dialysis-facility survey (ICH CAHPS, April 2026 file), drawn as in the hospital figure: each of 2,594 facilities is a dot at its published score, the gray ruler above shows the published star bins and the blue ruler below those of the exact optimum, with each grouping’s SSQ at the right. On nephrologists’ communication and caring (top) every optimal edge lies below its published partner and 1,042 facilities, 40 percent, move up; on rating of the nephrologist (bottom) the lowest edge is unchanged, the other three optimal edges lie above their partners, and 1,042 facilities move down. The optimal SSQ is strictly lower on both (9,116.5 against 10,329.0; 6,224.7 against 7,129.8). Here the published scores are confirmed to be CMS’s inputs, so the whole difference is the algorithm’s. (Figure 4 of the paper.)

Scope and open questions

The paper attributes no error to anyone. Because both algorithms are deterministic, a published computation can be reproduced exactly without revealing how far it is from the optimum, and that distance could persist for years without a mistake by anyone. Which of the two Medicare Advantage provisions controls is a question for the courts, and the paper leaves it aside. If the method is read as the rule, the paper measures how far the method falls short of the objective; if the objective is read as the rule, it measures payments that may have been misdirected. The result concerns specific contracts, whose bonus status differs under the two computations, rather than any aggregate tendency to rate too generously or too harshly.

The paper lists its own limitations: the dollar figures depend on the seven conventions, the per-enrollee band and a handful of large contracts. The Medicare Advantage clustering samples are reconstructions, and some of the valued changes depend on the fitted score profiles in them. The hospital comparison rests on mirrored code, and the two survey comparisons take the technical notes’ description of the inputs as given; for the hospital survey, CMS’s inputs are evidently not exactly the public file’s scores. The court rulings that have altered the ratings so far each turned on one element of the methodology; none reached the question measured here. CMS’s rule for 2027 provides for displaying de-identified sample data “needed to replicate the cut point methodology,” though the author could not perform that replication from the documents published so far. A proposal to replace clustering with percentile cut points was deferred, so clustering remains the method in force.

The paper

Supplementary material

The paper’s materials are hosted on this site. The two files below rerun the Medicare Advantage cut-point comparison with Python alone; the full record package behind the paper is being re-issued.

References

Code of Federal Regulations, 42 C.F.R. § 422.162: Medicare Advantage Quality Rating System and 42 C.F.R. § 422.166: Calculation of Star Ratings, ecfr.gov (accessed 2026)the regulation text: the sum-of-squares criterion in 422.162(a), the prescribed method in 422.166(a)(2)(i)
Centers for Medicare & Medicaid Services, Medicare 2026 Part C & D Star Ratings Technical Notes, technical report (2025); the 2024 and 2025 editions likewiserestate the criterion and document the trimming, the ten-group resampling and its seed, and the guardrail
W. D. Fisher, On grouping for maximum homogeneity, J. Am. Stat. Assoc. 53 (1958) 789the dynamic program that finds the exact minimum for scores on a line
D. Aloise, A. Deshpande, P. Hansen and P. Popat, NP-hardness of Euclidean sum-of-squares clustering, Mach. Learn. 75 (2009) 245minimizing the criterion is NP-hard for data of arbitrary dimension
H. Wang and M. Song, Ckmeans.1d.dp: Optimal k-means clustering in one dimension by dynamic programming, The R Journal 3 (2011) 29Fisher’s algorithm in standard statistical software, an R package
B. Bayles, The fault is not in our stars, but in ourselves, web post, bbayles.com (2017)found that CMS’s 2017 hospital k-means code stopped before converging, changing hundreds of ratings
R.-H. Huang, rstarating: CMS Hospital Compare Star Rating SAS Pack Replica, software repository, github.com/huangrh/rstarating (2017)reimplemented the 2017 hospital code and put the affected share near one quarter
M. E. Barclay, M. Dixon-Woods and G. Lyratzopoulos, Concordance of hospital ranks and category ratings using the current technical specification of US hospital star ratings and reasonable alternative specifications, JAMA Health Forum 3 (2022) e221006varied earlier steps of the 2021 hospital computation, clustering held fixed; about half of hospitals changed star
H. Rogers and M. Smith, Recalculating Medicare Advantage: Potential SCAN and Elevance ruling implications for MA stakeholders, white paper, Milliman (2024)replicates CMS’s published Medicare Advantage computation and runs counterfactuals for the 2024 SCAN ruling
M. Darden and I. M. McCarthy, The star treatment: Estimating the impact of star ratings on Medicare Advantage enrollments, J. Hum. Resour. 50 (2015) 980the published stars as data: effects of the bonus threshold on payments and enrollment
A. Anderson and M. K. Meiselbach, Fluctuating Star Ratings and Medicare Advantage bonuses, JAMA Health Forum 6 (2025) e254398year-to-year instability of the published ratings and the bonuses that follow them
B. Vatter, Quality disclosure and regulation: Scoring design in Medicare Advantage, Econometrica 93 (2025) 959the design of Medicare Advantage scoring rules, taking the published computation as given

← back to the web summaries