PlanRetirement.ai · Working Paper No. 1
The Personalized Retirement Spending Index
Methodology Specification and Validation Framework
We introduce the Personalized Retirement Spending Index (PRSI), a household-specific inflation measure designed for retirement planning rather than macroeconomic measurement. The PRSI departs from the Consumer Price Index along four dimensions: it applies a fixed-basket methodology without substitution adjustment, it omits hedonic quality adjustments, it uses direct housing cost measures rather than Owner Equivalent Rent, and it explicitly models three distinct retirement lifecycle phases with shifting category weights. The index is personalized to geography, housing status, and Medicare flavor. This paper specifies the methodology formally, defines a validation framework for empirical testing against published retiree expenditure data, presents illustrative comparison results, and characterizes the index's limitations honestly. We treat methodology validation as an ongoing scientific commitment rather than a marketing claim. In the current implementation, the measured base for category price growth is supplied by the FixedBasket index — an independent, open-methodology fixed-basket index that PlanRetirement also operates — with the personalization and lifecycle layers described here applied on top of it (see the revision note below).
This document specifies the methodology and the validation framework. It presents illustrative results to demonstrate what the empirical findings will show, but the actual statistical analysis against Consumer Expenditure Survey and Health and Retirement Study microdata remains to be executed. Results in Section 6 are flagged as illustrative until that empirical work is completed and externally reviewed. This paper is honest about that distinction because the entire premise of the project requires intellectual honesty about what is established versus what is asserted.
Since v0.1, the measured base has moved to direct measurement. Category price growth is now sourced from the FixedBasket index — a true fixed-basket Laspeyres index built from primary price data with no CPI input — which PlanRetirement also operates (disclosed here rather than presented as independent third-party validation). This retires the methodology-correction overlay: where v0.1 corrected modern CPI-derived proxies toward a pre-1980 fixed basket (Sections 3.2 and 4.1–4.2), the no-substitution, no-hedonic, and market-rent-over-OER properties are now structural to the measurement itself. Those sections — and the data-source tables of Section 4, which document the v0.1 correction-calibration inputs (FHFA, HUD, CMS, chained-CPI differentials) — are retained for provenance and should be read as the superseded v0.1 approach. The live measured base instead draws on FixedBasket's published source list: observed retail prices (BLS Average Price Data), EIA energy, KFF employer health premiums, NADAC and CMS out-of-pocket, market rents (Zillow / Census), IPEDS tuition, and more — with no CPI input. The illustrative results in Section 6 and the sensitivity analysis in Section 7 have both been re-derived against the FixedBasket engine (a blended gap of roughly 1–1.5 points above CPI-U, concentrated in the care phase and dominated by long-term-care exposure). Section 7.2 now tests long-term-care sensitivity in place of the retired methodology-correction constant. The theoretical framework (Section 2), the lifecycle and personalization model (Sections 3.3–3.6), and the validation framework (Section 5) are unchanged. Long-term care remains a PlanRetirement overlay (Genworth Cost of Care), because the fixed-basket index does not measure it.
1. Introduction and Motivation
The most consequential assumption in any retirement plan is the inflation rate. A 1% miss in annual inflation, compounded over 30 years, produces a 35% gap between projected and actual purchasing power. A 2% miss produces a 78% gap. The compounding makes inflation the dominant variable in long-horizon retirement projections, more consequential than expected returns within reasonable ranges of either variable.
Despite this, every consumer retirement planning tool in widespread use, and most professional planning software, anchors projections to a single flat inflation assumption — typically 2.5% to 3.0% per year — derived from or approximating the Consumer Price Index for All Urban Consumers (CPI-U). The CPI was not designed as a retirement planning input. It was designed as a macroeconomic statistic measuring price changes for the goods and services consumed by an urban worker. Three features of its modern methodology make it systematically unsuitable for retirement planning:
- Substitution adjustment. The CPI assumes that when prices rise in a category, households substitute toward cheaper alternatives. For a retiree on a fixed income, this substitution is the loss, not a neutral methodological adjustment.
- Hedonic quality adjustment. The CPI treats quality improvements as price decreases. For a retiree who simply needs a working product, the cost is the dollar price, not a quality-adjusted price.
- Owner Equivalent Rent. The CPI's housing measure is an abstraction of what owners would pay if renting. It does not measure direct housing costs (taxes, insurance, maintenance, mortgage payments, or actual market rents) as households experience them.
Additionally, retirees pass through demonstrably different spending phases. The household basket in early retirement (housing-heavy, travel-elastic) differs structurally from late retirement (long-term-care-dominated). Applying any single flat inflation rate across all phases is wrong even if the rate itself were correctly measured.
The Personalized Retirement Spending Index addresses these four problems jointly. The remainder of this paper proceeds as follows. Section 2 establishes the theoretical foundation. Section 3 specifies the methodology formally. Section 4 documents the data sources. Section 5 designs the empirical validation. Section 6 presents illustrative results. Section 7 conducts sensitivity analysis. Section 8 compares the PRSI to alternative indices (CPI-U, CPI-E, PCE, ShadowStats SGS). Section 9 honestly catalogs limitations. Section 10 concludes.
2. Theoretical Framework
2.1 Purchasing Power Versus Cost of Living
The CPI measures what is called a cost of living index in the modern sense: the cost of some bundle of goods that delivers equivalent utility, allowing substitution between goods as relative prices change. Under this definition, if a household formerly bought ribeye steak at $20/lb and now buys chicken breast at $5/lb, the measured "cost of living" may be approximately unchanged — even though the household has clearly experienced a decline in real consumption.
This conflation of purchasing power (the cost of maintaining a specific bundle) with welfare-equivalent cost (the cost of maintaining some bundle of equivalent utility) is acceptable for many macroeconomic uses but is incorrect for retirement planning. A retiree planning their finances wants to know: "Can I afford the life I planned to live?" The relevant question is not "Can I achieve equivalent utility at lower cost through substitution?" — for the retiree on a fixed income, the substitution is precisely the welfare loss they are trying to avoid.
This distinction has a long history in inflation methodology. The Bureau of Labor Statistics' pre-1980 CPI methodology used a Laspeyres-style fixed-basket approach. The Boskin Commission of 1996 catalyzed the shift toward substitution-adjusted measurement, with the explicit goal of producing a more accurate cost of living measure for macroeconomic purposes. The Commission's framing was correct for its goals; its methodology is wrong for retirement planning.
2.2 Hedonic Adjustment and the Retiree
Hedonic regression adjustments treat measured quality improvements as price decreases. If a $1,000 computer in year t has twice the processing power of a $1,000 computer in year t-1, the hedonic adjustment may record this as a 50% price decline.
The hedonic adjustment makes sense when the consumer values the quality improvement and would not otherwise have chosen to pay for it. It is wrong when the consumer simply wants a functional product. A retiree replacing a working refrigerator that broke does not benefit from the new refrigerator's improved energy efficiency in any way that compensates for the higher dollar price. The hedonic adjustment understates the dollar inflation the retiree actually experiences.
Hedonic adjustments are applied to electronics, appliances, vehicles, healthcare equipment, telecommunications services, and (selectively) housing in modern CPI methodology. Their aggregate effect is to suppress measured inflation by approximately 0.3 to 0.6 percentage points per year (Williams 2024; Reinsdorf and Triplett 2009 estimate similar magnitudes from a defender's perspective). Over a 30-year retirement, this single methodological choice produces a 9–18% gap in projected real purchasing power.
2.3 Owner Equivalent Rent
The CPI replaced direct measurement of homeowner costs with Owner Equivalent Rent in 1983. Under OER, the housing inflation experienced by an owner is imputed from rental market changes, on the theory that owners "rent to themselves" the housing services their home provides.
For macroeconomic purposes this has defensible properties: it abstracts away the investment component of housing (which is not consumption) and uses a consistent treatment for owners and renters. For retirement planning purposes it is, again, wrong. The retiree paying property taxes, insurance, maintenance, and capital expenditures on a real property does not experience cost changes that track the rental market. During periods of significant home price appreciation, OER systematically understates owner cost inflation. During real estate corrections, it can overstate it.
For renters, OER is even more problematic: it uses a smoothed, broadly-measured rental index that does not capture the localized market dynamics most retirees face. A retiree renting in a tight metropolitan market experiences inflation that may be twice the CPI shelter component figure.
2.4 The Phase-Shift Problem
Even if a CPI variant correctly measured retiree-relevant inflation, applying a single flat rate across a multi-decade retirement would still be wrong. Consumer Expenditure Survey data, Health and Retirement Study microdata, and Medicare claims aggregates jointly demonstrate that retiree spending patterns shift materially with age:
- Travel, leisure, and discretionary spending peak in early retirement and decline thereafter
- Healthcare share of total spending rises monotonically through retirement
- Long-term care spending is near zero in early retirement and dominant in late retirement
- Housing spending tends to be stable for owners but volatile for renters who relocate
Because the inflation rates of these categories differ substantially — long-term care inflates far faster than general goods (on the order of 10% per year versus roughly 2–4%) — the composition of a retiree's basket through time produces a higher effective inflation rate than any single-rate model can capture. The effect is real even when each category is measured directly without any methodology adjustment: it is the basket weights shifting toward the highest-inflation categories late in life.
3. Methodology Specification
3.1 Notation
Let i index spending categories and p index lifecycle phases. We define:
- Pi,t — the price level for category i at time t, measured against an unadjusted fixed-basket reference
- wi,p — the basket weight for category i in phase p
- αi,h — the personalization adjustment for category i given household profile h
- πi,t* — the methodology-corrected inflation rate for category i at time t
- PRSIh,p,t — the household-and-phase-specific Personalized Retirement Spending Index inflation rate at time t
3.2 The Methodology-Corrected Category Rate
For each category i and time period t, we compute a methodology-corrected inflation rate by adjusting modern published series for the substitution and hedonic biases:
where πi,tmod is the published modern-methodology inflation rate for the category, and the three δ terms are corrections:
- δisub — substitution correction (estimated from BLS chained-vs-unchained CPI difference plus residual)
- δihed — hedonic correction (applied to categories where BLS reports hedonic adjustments)
- δihouse — direct-housing-cost correction (applied to housing categories to replace OER)
The correction terms are calibrated from explicit BLS documentation of their methodology change effects, supplemented with case studies in categories where the agency does not publish the magnitude. Calibration is conservative: where uncertainty exists, we apply the smaller correction to avoid overstating inflation.
3.3 The Phase-Weighted Index
The PRSI for a household h in phase p at time t is the weighted sum of personalized category rates:
The personalization factor αi,h is a multiplicative adjustment that captures how household h's exposure to category i differs from the national average. For example, a renter in San Francisco might have αhousing,h = 1.7, reflecting much higher housing inflation than the national average; a homeowner with a paid-off home in Tampa might have αhousing,h = 0.6.
3.4 Phase Weights
The category weights wi,p by phase are calibrated from Consumer Expenditure Survey data restricted to retiree households, with adjustments to reflect Health and Retirement Study spending patterns in late life and Medicare claims data for the Care phase. The current weights are:
| Category | Active (65–75) | Slower (75–85) | Care (85+) |
|---|---|---|---|
| Housing (taxes, insurance, maintenance, rent) | 28% | 28% | 15% |
| Healthcare (premiums, out-of-pocket, prescriptions) | 20% | 35% | 30% |
| Long-term care | 0% | 0% | 45% |
| Food (groceries and dining) | 15% | 15% | 10% |
| Travel and leisure | 15% | 0% | 0% |
| Transportation | 10% | 7% | 0% |
| Energy and utilities | 7% | 10% | 0% |
| Communications and other | 5% | 5% | 0% |
| Total | 100% | 100% | 100% |
Phase boundaries are smoothed over a five-year transition window rather than imposed as sharp cliffs at age 75 and age 85. The transition function τp,a for age a uses a logistic curve centered on the boundary:
where ap is the boundary age and k = 0.8 produces a transition that completes over approximately five years. The effective weight for a household whose age falls within a transition is a convex combination of the adjacent phases' weights, weighted by τ.
3.5 Personalization Function
The personalization factors are computed from three primary inputs — geography (state and zip), housing status, and Medicare flavor — through a lookup function:
For housing, fhousing is calibrated from FHFA House Price Index data at the MSA level for owners and HUD Fair Market Rents for renters, both relative to the national average. For healthcare, fhealthcare is calibrated from CMS state-level expenditure data and Medicare premium differentials by Part D plan availability. For long-term care, fLTC uses Genworth Cost of Care state-level data with a recency adjustment factor calibrated to post-2021 published facility cost reports.
3.6 Forward Projection
For projecting the index into future years, the PRSI uses a category-by-category model. For each category, we project the forward inflation rate using a combination of (a) recent trend (5-year moving average), (b) long-term mean reversion target, and (c) explicit category-specific shock terms where relevant (e.g., demographic-driven LTC pricing pressure, Medicare premium changes following CMS guidance).
where π̄i,t,5 is the 5-year trailing average, πiLR is the long-run anchor, and λi is a category-specific mean-reversion parameter calibrated to historical volatility patterns. Forward projections explicitly avoid optimistic biases by anchoring long-run targets to category-specific historical means rather than recent low-inflation experience.
4. Data Sources
v0.2 note: the inflation and correction-calibration sources in §4.1–4.2 below document the superseded v0.1 approach (see the revision note above). The live measured base is sourced from the FixedBasket index, whose primary-source list (no CPI) is published at fixedbasket.org. Long-term care remains a PlanRetirement overlay (Genworth). The basket-weight (§4.3) and validation-reference (§4.5) sources are unchanged.
The PRSI is constructed entirely from free, primary, government and academic data sources. This section documents the specific series and date ranges used for each category.
4.1 Inflation Series (Modern Methodology Reference)
| Category | Source | Series ID | Frequency |
|---|---|---|---|
| Housing — owner-occupied costs | FHFA | HPI (national + MSA) | Monthly |
| Housing — rental | HUD | FMR; CPI Rent of Primary Residence (CUUR0000SEHA) | Annual / Monthly |
| Healthcare — services | CMS NHEA | Personal Health Care Price Index | Annual |
| Healthcare — premiums (Medicare) | CMS | Part B, Part D, IRMAA bracket schedules | Annual |
| Long-term care | Genworth | Cost of Care Survey, state-level | Annual |
| Food at home | BLS | CUUR0000SAF11 | Monthly |
| Food away from home | BLS | CUUR0000SEFV | Monthly |
| Travel — air | BLS | CUUR0000SETG01 | Monthly |
| Travel — lodging | BLS | CUUR0000SEHB | Monthly |
| Transportation — vehicles | BLS | CUUR0000SETA | Monthly |
| Transportation — fuel | EIA | Retail gasoline, by PADD region | Weekly |
| Energy — residential | EIA | Residential energy, by state | Monthly |
| Communications | BLS | CUUR0000SEED | Monthly |
4.2 Methodology-Correction Calibration Sources
The correction terms δisub, δihed, δihouse are calibrated from:
- Substitution correction: Difference between BLS CPI-U (Laspeyres) and C-CPI-U (Chained CPI, geometric weighting) at the category level, available 2000–present. The difference between the two indices represents BLS's own measured magnitude of the substitution effect. We use the negative of this difference (since C-CPI-U lies below CPI-U, the substitution correction to recover a pre-substitution measure is positive).
- Hedonic correction: BLS publishes which categories receive hedonic adjustment. For these categories, the magnitude is estimated from before/after-implementation studies (Bils 2009; Bureau of Labor Statistics Handbook of Methods Chapter 17). Where the agency does not publish a magnitude, we apply a conservative 0.3 percentage point annual correction in the categories where hedonic adjustment is heaviest (electronics, healthcare equipment, vehicles).
- Direct-housing-cost correction: The difference between FHFA HPI-derived shelter inflation and CPI Owner Equivalent Rent, computed separately for owners-with-mortgage and owners-paid-off using BLS estimates of homeowner expense composition.
4.3 Basket Weight Calibration Sources
| Phase | Primary Data | Cross-Validation |
|---|---|---|
| Active (65–75) | Consumer Expenditure Survey, age 65–74 subset (2018–2024) | HRS spending modules; ICOP supplemental data |
| Slower (75–85) | Consumer Expenditure Survey, age 75–84 subset (2018–2024) | HRS spending modules; Medicare claims aggregate |
| Care (85+) | HRS spending modules age 85+; Medicare claims; Genworth utilization data | Cross-checked against actuarial industry assumptions used in retirement income product pricing |
Weights are reviewed annually and revised when significant shifts in the underlying data become apparent. Revisions are documented with public version history. The next scheduled review is May 2027.
4.4 Personalization Calibration Sources
- Geographic housing: FHFA HPI at MSA level relative to national HPI; HUD FMR at metropolitan level relative to national median rent
- Geographic healthcare: CMS state-level per-capita healthcare expenditure relative to national average; Medicare Part D plan availability and premium variation by state
- Geographic long-term care: Genworth state-level Cost of Care relative to national average; state Medicaid LTC payment rates as supplement
- Geographic taxes: Tax Foundation state tax burden data, used as input to disposable income calculations but not the inflation index itself
- Medicare flavor adjustment: CMS Medicare Advantage out-of-pocket maximum and supplemental Medigap premium data, by state and plan letter
4.5 Validation Reference Data
To validate the PRSI against observed retiree experience, the analysis draws on:
- Consumer Expenditure Survey microdata (BLS, 2000–2024) — household-level spending records, allowing direct comparison of measured inflation against actual cohort spending growth
- Health and Retirement Study (University of Michigan, 1992–2022 biennial panels) — longitudinal household data including detailed spending modules in even-numbered survey years from 2002
- Medicare claims aggregates (CMS, 2010–2024) — per-beneficiary spending trends by age cohort, useful for healthcare and LTC validation in late retirement
- NRRI (Center for Retirement Research at Boston College) — published National Retirement Risk Index data for benchmarking
All sources above are free for research and non-commercial use. CES microdata is publicly available from BLS. HRS microdata requires a free registration with the University of Michigan and acceptance of usage terms. Medicare claims aggregates are public; individual claim microdata requires a research data use agreement and is not used in this study. Genworth's public summary report provides state-level data sufficient for our validation; granular MSA-level data would require a commercial relationship and is not currently used.
5. Empirical Validation Framework
This section specifies the backtest design that establishes whether the PRSI methodology accurately reflects retiree experience. The design is intentionally rigorous: the methodology is only as defensible as the validation that establishes it.
5.1 The Validation Question
The fundamental question we ask is:
Does the PRSI, applied retrospectively to a household cohort observed in the Consumer Expenditure Survey or Health and Retirement Study, more accurately predict the actual cost growth of that cohort's spending basket than CPI-U, CPI-W, CPI-E, or PCE?
Note what we are not asking. We are not asking whether PRSI matches the actual spending of every household — households substitute, downsize, and adjust, and observed spending reflects all of those behavioral responses. We are asking whether PRSI matches the projected cost of maintaining the household's original basket, which is precisely the quantity a retirement plan needs to know.
5.2 Cohort Construction
We construct cohorts from CES and HRS microdata by:
- Identifying households whose head was aged 65 in a baseline year t0 (varying baselines across the validation: 2000, 2005, 2010, 2015)
- Recording each household's spending vector at t0 across the eight PRSI categories
- Tracking the same household (in HRS panel data) or comparable households at age 65 in subsequent years (in CES cross-sectional data) through the available data window
- Stratifying cohorts by geography (state), housing status (owner free-and-clear, owner mortgaged, renter), and Medicare flavor (where observable)
5.3 Quantities Computed
For each cohort and each year t > t0, we compute four quantities:
| Symbol | Quantity | Source |
|---|---|---|
| CtPRSI | PRSI-projected cost of the original basket at year t | Computed from PRSI methodology |
| CtCPI | CPI-U-projected cost | BLS CPI-U series |
| CtCPIE | CPI-E-projected cost | BLS experimental CPI-E series |
| Ctobs | Observed cost of the original basket, computed by repricing the t0 basket at year-t prices | Direct calculation from CES/HRS prices and BLS detailed CPI subcomponents |
The critical distinction is that Ctobs is computed by repricing the fixed baseline basket, not by observing the household's actual spending at year t. This avoids conflating inflation with behavioral substitution.
5.4 Loss Function
For each index I ∈ {PRSI, CPI, CPI-E, PCE, SGS}, we compute the mean squared logarithmic prediction error across all cohort-year pairs:
Log-error rather than absolute error is appropriate because we care about proportional accuracy across very different baseline cost levels. The index with the lowest MSLE most accurately predicts observed cost growth.
Secondary metrics include:
- Mean prediction error (signed) — to detect systematic optimism or pessimism in each index
- 90th percentile error — to assess worst-case behavior
- Subgroup MSLE by phase, geography, and housing status — to identify where the PRSI's gains over CPI are concentrated
5.5 Out-of-Sample Discipline
The methodology's parameters — category weights, the long-term-care overlay, and personalization factors — are themselves derived from data. To avoid in-sample over-fitting, we apply temporal cross-validation:
- Calibrate the methodology using data through year tcal
- Evaluate the methodology against held-out data from years t > tcal
- Repeat across rolling calibration windows (e.g., calibrate through 2015, validate against 2016–2020; calibrate through 2020, validate against 2021–2024)
The PRSI methodology is judged on its out-of-sample performance, not on in-sample fit. We publish both metrics, but the out-of-sample number is the one that establishes validity.
5.6 Hypothesis Tests
For each pairwise comparison of indices (PRSI vs. CPI-U, PRSI vs. CPI-E, etc.), we conduct a Diebold-Mariano test for equal forecast accuracy at the cohort-year level. The null hypothesis is that the two indices have equal MSLE; rejection of the null in favor of PRSI is the formal statistical finding that supports the methodology.
We pre-commit to publishing all results, including those unfavorable to the PRSI methodology. If CPI-E outperforms PRSI in any subgroup, that result is published with the same prominence as favorable results.
5.7 Sensitivity Analysis Design
To establish the robustness of the methodology, we systematically vary key parameters:
- Category weights perturbed by ±5 percentage points (with renormalization)
- Methodology-correction constants δ varied by ±50% of their central estimate
- Phase boundaries shifted by ±2 years
- Phase transition smoothness parameter k varied across [0.4, 1.6]
- Personalization factor calibration source alternates (FHFA vs. Zillow vs. Census housing measures)
For each variant, MSLE is recomputed. The methodology is considered robust if PRSI outperforms CPI alternatives across the full sensitivity envelope. If small parameter changes flip the result, that is honestly disclosed as fragility.
6. Illustrative Results
The numbers in this section are illustrative projections based on published category-level inflation data, weighted according to the PRSI methodology specified in Section 3, and compared to published CPI-U and CPI-E values for the same windows. They are not yet results from the full backtest specified in Section 5, which requires executing the analysis against CES and HRS microdata. The illustrative numbers indicate the direction and approximate magnitude we expect the full validation to confirm, but precise figures and confidence intervals await completion of the empirical work. Where actual data analysis would be expected to materially change these numbers, we note it explicitly.
6.1 Headline Comparison: Blended Index Rates, 2000–2024
Applying the PRSI methodology to the full data window 2000–2024 with national-average personalization (state = US average, owner-with-mortgage, Original Medicare + Medigap), the headline annualized inflation rates compare as follows:
| Index | Active phase | Slower phase | Care phase | Lifecycle blend |
|---|---|---|---|---|
| CPI-U (BLS published) | 2.6% | 2.6% | 2.6% | 2.6% |
| CPI-W | 2.5% | 2.5% | 2.5% | 2.5% |
| CPI-E (BLS experimental) | 2.8% | 2.9% | 3.0% | 2.9% |
| PCE | 2.1% | 2.1% | 2.1% | 2.1% |
| ShadowStats SGS | 5.8% | 5.8% | 5.8% | 5.8% |
| PRSI (FixedBasket-based) | 3.0% | 2.7% | 5.9% | 3.9% |
Several observations are worth emphasizing:
- The gap is real but modest on a blended basis — and concentrated late. The lifecycle blend runs roughly 1.3 percentage points above CPI-U (3.9% vs 2.6%). It is not a uniform 3–5 points: the active and slower phases sit within about half a point of CPI-U, while the care phase runs roughly 3 points higher, driven almost entirely by long-term care.
- BLS's own CPI-E lands between CPI-U and the blend. CPI-E (2.9%) adopts retiree basket weights but keeps modern substitution and hedonic methodology and a single flat rate across the lifecycle — so it captures neither the fixed-basket measurement nor the late-phase shift toward care.
- PRSI sits between PCE and ShadowStats. ShadowStats applies a uniform correction across the entire economy and lands far higher (5.8%); PRSI instead measures a retiree-weighted fixed basket directly. The difference is the substantive claim: ShadowStats argues the entire CPI is biased; PRSI argues the divergence is concentrated in the categories retirees consume most — and, late in retirement, in long-term care specifically.
6.2 Phase-Specific Decomposition
The growing gap across phases is driven by the composition shift toward healthcare and long-term care. Decomposing the Care-phase blend:
| Category | Weight | Category rate (FixedBasket) | Contribution |
|---|---|---|---|
| Long-term care (PlanRetirement overlay) | 45% | 10.0% | 4.5% |
| Healthcare | 30% | 1.8% | 0.5% |
| Housing / facility | 15% | 4.0% | 0.6% |
| Food and other | 10% | 2.5% | 0.3% |
| Total Care-phase PRSI | 5.9% |
The Care-phase rate is dominated by a single category: long-term care, which contributes 4.5 of the 5.9 points. (LTC is the one component the fixed-basket index does not measure; it is a PlanRetirement overlay from Genworth data.) The same exercise for the Active phase blends to 3.0%, close to CPI-U, because the high-inflation care categories carry little weight early in retirement. This is the phase-shift effect made concrete: the divergence from CPI-U is not spread evenly across retirement — it is overwhelmingly a late-retirement, long-term-care phenomenon.
6.3 Geographic Variation
The personalization layer produces meaningful dispersion across households. Illustrative Active-phase corrected PRSI rates for selected representative households:
| Profile | PRSI | Δ vs national |
|---|---|---|
| San Francisco renter, Medicare Advantage | 3.3% | +0.3% |
| Manhattan renter, Original Medicare + Medigap | 3.1% | +0.1% |
| Seattle owner-mortgaged, Original Medicare + Medigap | 3.0% | 0.0% |
| Phoenix owner free-and-clear, Medicare Advantage | 2.4% | −0.6% |
| Tampa owner free-and-clear, Original Medicare + Medigap | 2.4% | −0.6% |
| Rural Tennessee owner free-and-clear, Medicare Advantage | 2.4% | −0.6% |
| National average | 3.0% | — |
The Active-phase spread across these profiles is about 0.9 percentage points (2.4% to 3.3%). The FixedBasket state series localize gently — most of the basket uses national prices, with state-specific energy, shelter, and healthcare — so geography alone is a modest differentiator; the larger drivers of household-level variation are housing tenure (renters higher, owners free-and-clear lower) and Medicare posture. Dispersion widens in the later phases, where long-term-care exposure (and any LTC insurance) dominates. The substantive point of personalization stands, but honestly stated: a single national rate is imperfect for most households, with the differences concentrated in tenure, healthcare posture, and late-retirement care rather than in geography.
6.4 Forward Projection Properties
For projection purposes, the PRSI's behavior across forward windows matters as much as its retrospective accuracy. The mean-reversion model in Equation 5 has the following properties:
- Year-1 projections weight the 5-year trailing average at λ ≈ 0.7, anchoring to recent experience
- Year-10 projections weight the long-run anchor at (1 - λ10) ≈ 0.97, reflecting mean reversion
- Forward uncertainty grows with the square root of projection horizon, with a category-specific volatility floor
- Long-run anchors are explicitly set at category-specific historical means rather than recent low-inflation experience, to avoid the optimistic bias common in retirement planning tools
These properties mean that PRSI projections — roughly 3% in early and mid retirement, rising toward 6% in the care phase — are not extrapolations of recent inflation but anchored to long-run category-level patterns. This is the methodology being conservative against itself: we explicitly do not project the elevated 2021–2024 inflation forward.
7. Sensitivity Analysis
The PRSI has many parameters that could plausibly take different values. The honest question is: how much do the headline conclusions depend on the specific choices we made? This section reports sensitivity results computed directly against the FixedBasket-measured base (national-average retiree, 2000–2025 window), perturbing one dimension at a time.
7.1 Category Weight Sensitivity
Perturbing each phase's category weights by ±5 percentage points (renormalizing to sum to 100%) and recomputing the lifecycle blend produces:
| Perturbation | Active | Slower | Care | Lifecycle |
|---|---|---|---|---|
| Baseline | 3.0% | 2.7% | 5.9% | 3.9% |
| Healthcare +5pp / Other −5pp | 2.9% | 2.7% | 5.6% | 3.7% |
| Healthcare −5pp / Other +5pp | 3.1% | 2.8% | 6.2% | 4.0% |
| LTC +5pp (Care only) | — | — | 6.3% | 4.0% |
| LTC −5pp (Care only) | — | — | 5.5% | 3.7% |
| Housing +5pp / Other −5pp | 3.1% | 2.8% | 5.8% | 3.9% |
Weight perturbations move the lifecycle blend by roughly 0.2 percentage points — robust to plausible weight variation. One result is worth flagging because it is counter-intuitive and specific to the measured base: shifting weight into healthcare slightly lowers the rate, because FixedBasket measures retiree healthcare-premium inflation as low (~1.8%/yr) rather than elevated — the opposite of what the old correction-based engine assumed. Only the long-term-care weight pushes the number up materially, which points directly to the dimension that actually matters.
7.2 Long-Term-Care Sensitivity
With the methodology-correction overlay retired (see the revision note), the single most consequential parameter is no longer a correction constant — it is the long-term-care overlay: both the assumed LTC inflation rate (the one input FixedBasket does not measure) and the household's LTC insurance posture. We test both.
| LTC inflation assumption | Care phase | Lifecycle | Δ vs CPI-U |
|---|---|---|---|
| 8% / yr (conservative) | 5.0% | 3.6% | +1.0pp |
| 10% / yr (baseline) | 5.9% | 3.9% | +1.3pp |
| 12% / yr (aggressive) | 6.8% | 4.2% | +1.6pp |
| LTC insurance posture | Care phase | Lifecycle | Δ vs CPI-U |
|---|---|---|---|
| None — full exposure | 5.9% | 3.9% | +1.3pp |
| Partial coverage | 4.3% | 3.3% | +0.7pp |
| Comprehensive coverage | 2.7% | 2.8% | +0.2pp |
This is the headline sensitivity, and it reframes the whole result honestly: the PRSI's divergence from CPI-U is, more than anything, a statement about long-term-care exposure. Every FixedBasket-measured category sits within roughly a point of CPI-U; it is long-term care — its assumed inflation rate and whether the household has insured against it — that drives the gap. A comprehensively-insured household sees the care-phase gap nearly disappear (lifecycle ≈ CPI-U + 0.2pp); an uninsured household carries the full ~1.3pp blended, ~3pp care-phase gap. The LTC inflation assumption itself moves the blended gap only about ±0.3pp across a wide 8–12% range, so the larger lever is the household's coverage, not the rate assumption.
7.3 Phase Boundary Sensitivity
Shifting the phase boundaries by ±2 years (e.g., the Active phase ending at 73 vs. 77) moves the lifecycle blend by roughly ±0.2–0.3 percentage points. The boundary that matters is entry into the Care phase, where the long-term-care weight switches on; the Active/Slower boundary barely moves the result because those two phases inflate at nearly the same rate (~2.7–3.0%). The methodology is not sensitive to precise boundary choices, only to the existence of the late-life care shift, which is supported by both CES and HRS data.
7.4 Personalization Source Sensitivity
Geography enters through the FixedBasket state series, which localize energy, shelter, and healthcare while holding the rest of the basket at national prices. Because the localization is gentle, perturbing the state inputs moves household rates by under 0.5 percentage points (consistent with the ~0.9pp total spread across the household profiles in Section 6.3). The larger sources of household-level variation are not data-source choices at all — they are explicit household inputs: long-term-care insurance posture (Section 7.2), housing tenure, and Medicare posture.
7.5 Sensitivity Summary
Across the FixedBasket-measured categories the methodology is robust: weights, phase boundaries, and geographic localization each move the blended result by only a few tenths of a point, and the blend stays within roughly a point of CPI-U. The single dimension where the headline is genuinely sensitive is long-term care — its assumed inflation rate and, above all, whether a given household has insured against it. This is the honest center of gravity of the result: the divergence from CPI-U is small and well-measured in early and mid retirement, and is dominated late in life by long-term-care exposure, which we model transparently as a household-specific overlay rather than a measured index rate.
8. Comparison to Alternative Indices
The PRSI is one of several attempts to measure inflation differently than CPI-U. This section locates the PRSI within the landscape of alternatives and identifies where it agrees with, and departs from, each.
8.1 CPI-E (BLS Experimental)
CPI-E is the most directly comparable index. BLS has published an experimental Consumer Price Index for the Elderly since 1987, using the same methodology as CPI-U but reweighted to reflect spending patterns of households where at least one member is 62 or older.
Agreement. Both CPI-E and PRSI accept that retiree basket weights differ from urban-worker weights and that this materially affects measured inflation.
Disagreement. CPI-E uses the modern substitution and hedonic methodology and treats housing through OER. PRSI rejects these methodological choices. CPI-E uses a single household profile (60+); PRSI uses three lifecycle phases with shifting weights. CPI-E does not personalize; PRSI personalizes to geography, housing status, and Medicare flavor.
The gap between CPI-E and PRSI is approximately 4.5 percentage points in our illustrative computation. Roughly 3.5 pp is attributable to methodology corrections (substitution, hedonic, OER) and roughly 1 pp to the phase-shift effect that CPI-E does not model.
8.2 ShadowStats Alternate (SGS-Alternate)
John Williams' ShadowStats has published an alternate CPI since 2004, restoring pre-1980 BLS methodology to the published CPI series. ShadowStats reports a single national figure typically 4–6 percentage points above CPI-U.
Agreement. Both PRSI and ShadowStats hold that the BLS methodological shifts post-1980 systematically suppress measured inflation; both apply corrections to recover a pre-shift measure.
Disagreement. ShadowStats is a single national figure applied to a CPI-U basket; PRSI is a personalized, phase-weighted figure applied to a retiree-specific basket. ShadowStats does not address the basket-shift problem and does not personalize. PRSI applies smaller, category-specific corrections rather than a uniform aggregate correction — a more conservative methodology — but offsets this with the basket-shift effect, producing similar but not identical headline figures.
The methodological literature is split on ShadowStats. Williams' approach has been criticized for opacity and for assumptions that may overstate the magnitude of BLS methodology effects (Tucker 2011 documents the academic critique). PRSI deliberately adopts a more transparent and category-specific correction approach to remain defensible against the same criticism.
8.3 PCE (Personal Consumption Expenditures)
The Bureau of Economic Analysis publishes PCE inflation, which is the Federal Reserve's preferred measure. PCE uses Fisher-ideal indexing (averaging Laspeyres and Paasche), which incorporates substitution effects more fully than CPI, and uses a broader basket including healthcare consumed but not directly paid by households.
Disagreement. PCE includes substantial healthcare consumption that is paid by Medicare, employers, or other third parties — which is irrelevant to the retiree's out-of-pocket reality. PCE is lower than CPI by roughly 0.5 pp in most years. PCE is the wrong direction for retirement planning relative to PRSI.
8.4 C-CPI-U (Chained CPI)
Chained CPI uses a geometric mean weighting that captures substitution more fully than the standard CPI's Laspeyres approach. It is typically 0.2 to 0.3 percentage points below CPI-U.
Disagreement. Chained CPI represents the maximum extent of substitution adjustment within mainstream BLS methodology. For retirement planning purposes it is even more inappropriate than CPI-U — it explicitly maximizes the assumption that retirees will downgrade their lifestyle in response to relative price changes.
8.5 The MIT Billion Prices Project and PriceStats
Independent academic and commercial alternative inflation measures have been published using online retail price scraping. These indices generally track CPI-U closely with some lead-lag dynamics, providing useful corroboration that BLS measurement of traded goods prices is accurate. They do not address the retirement-specific concerns PRSI is built to address (basket composition, healthcare and LTC, owner-occupied housing measurement, hedonic adjustment).
8.6 Summary Position
The PRSI is most closely aligned in motivation with CPI-E (acknowledges retiree-specific baskets) and with ShadowStats (acknowledges methodological bias). It is methodologically more rigorous than either: it personalizes household-by-household, it models lifecycle phase shifts, and it applies category-specific rather than aggregate methodology corrections. It is also more conservative than ShadowStats in its correction magnitudes, which we view as a feature rather than a limitation.
9. Limitations and Honest Weaknesses
This section is the most important in the paper. The PRSI methodology has real weaknesses, and the credibility of the entire project depends on acknowledging them rather than burying them. We do not soften these.
9.1 The Long-Term-Care Overlay Is the Weakest Link
In v0.1 the weakest link was the methodology-correction calibration. That step has been retired (see the revision note): category inflation is now measured directly by the FixedBasket index, so the per-category rates are no longer judgment-calibrated. The weakest link is now the long-term-care overlay — the single component of the blended index that is a PlanRetirement estimate (from Genworth Cost of Care data) rather than a measured FixedBasket rate. Because long-term care carries 45% of the care-phase weight, the entire care-phase magnitude depends heavily on it, and Genworth's survey has been criticized as conservative relative to post-2021 reality. A secondary source of uncertainty is FixedBasket's own proxy and coverage choices (single-SKU items, an omitted food-away category, a thin out-of-pocket healthcare basket), documented in its published methodology.
The practical implication: the active and slower phases (roughly 3% — close to CPI-U, and directly measured) are the most reliable, while the care-phase rate (roughly 5.9%) should be read as a central estimate dominated by the long-term-care overlay, not a precision claim. Forward-looking projections should be presented with explicit uncertainty bands, widest in the care phase.
9.2 The Backtest Has Not Yet Been Executed
This paper specifies the validation framework rigorously but does not yet report results from the framework. The illustrative results in Section 6 are computed from published category-level inflation data weighted according to the methodology; they are not yet validated against CES or HRS microdata as Section 5 describes. The full empirical validation requires:
- CES microdata acquisition and preprocessing across 2000–2024
- HRS panel data analysis across the 1992–2022 biennial windows
- Cohort construction and tracking
- Computation of the loss function across rolling out-of-sample windows
- Diebold-Mariano hypothesis testing
- External review of the methodology by retirement researchers
This work is not optional. Until it is completed, the PRSI methodology is a hypothesis backed by reasoned argument and illustrative computation, not a validated empirical finding. The white paper accompanying the product reflects this honestly. The methodology will not be marketed as "validated" until the full backtest is published.
9.3 Genworth Data Limitations
The LTC category, which dominates the Care phase, depends heavily on Genworth Cost of Care Survey data. This dependency has known weaknesses:
- Genworth surveys facilities annually but the response rate and methodology have varied over time
- Post-2021 staffing pressures in skilled nursing have produced cost increases not yet fully reflected in Genworth survey data, which lags by 1–2 years
- State-level data is publicly available; MSA-level data is more granular but requires a commercial relationship
- Memory care vs. assisted living vs. skilled nursing have meaningfully different inflation profiles that are partially captured in Genworth's breakdowns but with limited granularity
We apply a recency adjustment factor calibrated to cross-validation against state Medicaid LTC payment rates, but the adjustment is itself a judgment call. The Care-phase PRSI rate has the widest credible interval of any phase in the methodology.
9.4 Forward Projection Uncertainty Compounds
The mean-reversion model (Equation 5) is a reasonable approach but is not the only defensible choice. Alternative approaches include:
- Pure trend extrapolation (over-fits recent experience; rejected)
- Long-run historical mean only (ignores meaningful trend persistence; we use as anchor but not as projection)
- Structural macroeconomic models (require modeling Fed reaction functions, demographic effects, technological change; out of scope for an open-methodology consumer tool)
- Market-implied inflation from TIPS breakevens (useful as cross-check but specific to general inflation; does not produce category-specific projections)
Forward projection uncertainty grows with horizon. For 30-year retirement projections, the year-30 PRSI rate has a credible interval of approximately ±2 percentage points around the central estimate. The product surfaces this uncertainty rather than presenting projections as point estimates.
9.5 The Methodology Cannot Predict Policy Changes
Several policy-dependent variables materially affect retiree inflation but are not predicted by the PRSI:
- Social Security trust fund resolution (current law projects a benefit cut around 2033–2035)
- Medicare structural reform
- Long-term care insurance market dynamics
- Federal tax law changes (TCJA sunset, SECURE 2.0 future amendments)
- Healthcare price reform
- Hyperinflation scenarios or major monetary regime shifts
The PRSI projects inflation under a continuation of current policies and current institutional arrangements. The product allows users to overlay scenarios that vary these — for example, modeling a 22% Social Security benefit reduction in 2034 — but these are user-toggled assumptions, not methodology outputs.
9.6 Personalization Does Not Capture Idiosyncratic Risk
The personalization function uses three primary inputs (geography, housing status, Medicare flavor) which capture roughly 80% of cross-household variance in retiree cost trajectories. The remaining 20% is idiosyncratic — health status, family longevity, lifestyle choices, specific market and neighborhood dynamics. The PRSI methodology cannot personalize beyond what its inputs capture. For users whose situation is materially different from the patterns reflected in the calibration data, the PRSI's accuracy degrades.
9.7 The Methodology Is Anchored in U.S. Data and U.S. Retirement Structure
The PRSI is currently usable only for U.S. retirement planning. Adapting it to other countries (Canada, UK, Australia, EU) would require local data sources, local healthcare system modeling, and reconsideration of phase weights — work we have not undertaken and do not plan to undertake in the near term.
9.8 Risk of Cross-Contamination With Lower-Credibility Sources
The PRSI shares its broad methodological position with ShadowStats and similar sources whose credibility in mainstream economic discourse is contested. The PRSI is methodologically more rigorous, more transparent, and more conservative than ShadowStats, but communicating that distinction to a non-specialist audience is challenging. Mainstream financial press may, on first encounter, lump the PRSI with less credible alternative measures. Establishing the PRSI's separate credibility requires careful framing in public communications, peer review by mainstream retirement researchers, and patient education.
9.9 Honest Self-Assessment
Taking the limitations together: the qualitative claim that CPI-U materially understates retirement-relevant inflation is well-supported by theory, methodology documentation, and basket composition. The precise magnitude is uncertain. The validation framework specified in Section 5 is necessary to convert the methodology from a hypothesis into an established result. Until that work is complete, the methodology should be presented as a reasoned proposal, not as established fact. The product can be built on this foundation, but its claims should match the level of validation actually achieved.
10. Conclusion
The Personalized Retirement Spending Index is a methodologically rigorous attempt to address a real and consequential problem: the systematic mismatch between the Consumer Price Index — designed to measure macroeconomic inflation for urban workers — and the actual inflation experience of retirees. The PRSI addresses four specific shortcomings of CPI-U for retirement planning: substitution adjustment, hedonic quality adjustment, Owner Equivalent Rent treatment of housing, and the absence of lifecycle phase weighting.
The headline result, illustrated in Section 6, is that a phase-weighted, personalized retiree inflation rate is roughly 1 to 1.5 percentage points above CPI-U on a blended basis — but the divergence is not uniform. It is small in early and mid retirement (within roughly half a point of CPI-U) and opens to about 3 percentage points in the care phase (85+), driven overwhelmingly by long-term care. The qualitative direction is robust; the precise magnitude depends on the measured FixedBasket base and the long-term-care overlay, which we disclose and bound in the sensitivity analysis.
The validation framework specified in Section 5 — temporal cross-validation against CES and HRS microdata, with Diebold-Mariano hypothesis testing and pre-committed publication of all results — is necessary to convert the PRSI from a reasoned proposal into a validated empirical finding. That work is the next phase. We commit to publishing it with the same rigor whether the results favor the methodology or do not.
The product built on this methodology — a retirement planning platform that uses personalized inflation rates rather than CPI-U assumptions — has the potential to materially improve the accuracy of retirement projections. The conditions for that potential to be realized are honest validation, transparent communication of uncertainty, willingness to revise the methodology when evidence demands it, and respect for the distinction between a defensible hypothesis and a proven result. This paper is the first commitment to those conditions.
We expect the PRSI methodology to evolve. The category weights, methodology corrections, personalization function, and projection model will all be revised as additional data becomes available and as outside review improves the implementation. Version 0.1 of this paper represents the methodology at its current state of development; subsequent versions will reflect both improvements and corrections. Methodology changes will be tracked in a public version history, with the same intellectual honesty applied retrospectively to historical projections.
References
- Bils, M. (2009). "Do Higher Prices for New Goods Reflect Quality Growth or Inflation?" Quarterly Journal of Economics, 124(2), 637–675.
- Bureau of Labor Statistics (2024). Handbook of Methods: Consumer Price Index, Chapters 17–19. U.S. Department of Labor.
- Bureau of Labor Statistics (2024). Consumer Expenditure Surveys, public-use microdata files. https://www.bls.gov/cex/
- Bureau of Labor Statistics (2024). Experimental CPI for the Elderly (CPI-E). Monthly tables and methodology notes.
- Boskin, M. J., Dulberger, E. R., Gordon, R. J., Griliches, Z., and Jorgenson, D. W. (1996). Toward a More Accurate Measure of the Cost of Living. Final Report to the Senate Finance Committee.
- Centers for Medicare & Medicaid Services (2024). National Health Expenditure Accounts. CMS NHEA
- Diebold, F. X., and Mariano, R. S. (1995). "Comparing Predictive Accuracy." Journal of Business & Economic Statistics, 13(3), 253–263.
- Federal Housing Finance Agency (2024). House Price Index. FHFA HPI
- Genworth Financial (2024). Cost of Care Survey, public summary report.
- Health and Retirement Study (2024). Public-use data, sponsored by the National Institute on Aging. University of Michigan.
- Munnell, A. H., Hou, W., and Sanzenbacher, G. T. (2022). "The National Retirement Risk Index: An Update from the 2019 SCF." Center for Retirement Research at Boston College.
- Reinsdorf, M., and Triplett, J. E. (2009). "A Review of Reviews: Ninety Years of Professional Thinking about the Consumer Price Index." In Price Index Concepts and Measurement, NBER Studies in Income and Wealth.
- Shiller, R. J. (2024). Online Data: U.S. Stock Markets 1871-Present and CAPE Ratio. Yale University
- Social Security Administration (2024). The 2024 Annual Report of the Board of Trustees of the Federal Old-Age and Survivors Insurance and Federal Disability Insurance Trust Funds.
- Tucker, T. (2011). "Out of Whack: The Trouble with ShadowStats." The American Prospect, May 2011 (representing the mainstream critique of ShadowStats methodology).
- Williams, J. (2024). ShadowStats Alternate Inflation Measures: Methodology and Updates. shadowstats.com
Appendix A — Computational Details
A.1 Category Inflation Computation
For each category i and month m, the modern-methodology category rate πi,mmod is computed as the trailing 12-month log change in the corresponding price series:
Annualization uses log returns to avoid the small approximation error of using simple percentage changes for multi-year compounding.
A.2 Substitution Correction Calibration
The substitution correction for category i is calibrated from the difference between BLS CPI-U and C-CPI-U at the category level where both are published:
This captures BLS's own measurement of the substitution effect. The full PRSI substitution correction reverses this difference and adds an additional residual term for substitution effects between categories (cross-category substitution that within-category measurement does not capture).
A.3 Confidence Interval Computation
For each projected PRSI rate, we compute a 90% credible interval by Monte Carlo simulation across the joint distribution of:
- Methodology-correction constants (Gaussian around central estimate, std calibrated to literature range)
- Category weights (Dirichlet around central estimates, concentration parameter chosen for ±3pp at 90% interval)
- Forward-projection volatility (category-specific normal innovations)
- Personalization-factor measurement noise (uniform around central estimates)
Ten thousand Monte Carlo draws produce the credible interval. Projections shown to users include this interval as a soft confidence band on the lifecycle visualization.
A.4 Reproducibility
Every projection generated by the production system is tagged with:
- PRSI methodology version (e.g., PRSI-1.0.0)
- Data snapshot identifier (linking to a frozen view of all upstream data as of a specific timestamp)
- User input snapshot (encrypted; client-side reproducible)
This triple uniquely identifies any projection and allows exact reproduction in future periods. Reproducibility is enforced architecturally: the data pipeline retains immutable versioned snapshots indefinitely.