Back to Research

The Methods Audit

Audit period: July 2026 refresh cycle. Engine build at completion: August 2026.

Our methodology page makes a promise: if a tool only ever produces one answer, it isn't a model; it's an argument. This page is that promise kept. During a routine refresh of the July 2026 study, a systematic audit of the calculation engine found four defects. All four were fixed, the corrected engine was re-validated against the Social Security Administration's published replacement-rate benchmarks, and the study's conclusions changed as a result.

We publish this because a reader deciding whether to trust the model deserves to know exactly how it failed, how it was fixed, and how the answer moved. Everything below describes the state of the model as of July 2026. It is a record of a point in time, not a certification that the current or any future build is defect-free.

The four defects (as found, July 2026)

1. Nominal earnings were frozen at 2009 levels. The earnings engine built each worker's career from a historical income table (1964–2009) and clamped every later year to the 2009 value — producing zero nominal wage growth after 2009. A worker aged 60 in 2060 was assigned exactly the same nominal dollars as one aged 60 in 2040. This understated both payroll contributions and Social Security earnings records, and did so unevenly across birth cohorts.

2. AIME was computed without SSA wage indexing. The engine averaged raw nominal top-35 earnings with no wage indexing, while still projecting the PIA bend points forward at the assumed wage-growth rate. The two legs of the SSA's formula are designed to cancel; with one leg frozen and the indexing step absent, they did not. Younger cohorts' earnings histories collapsed below the first bend point, so moderate earners were treated as poverty-level earners and granted the 90% replacement rate. This erased Social Security's progressivity for younger cohorts and manufactured a spurious "the advantage grows for younger generations" result.

3. A hidden preset field silently overrode demographic inputs. Restoring a saved scenario applied the demographic bundle of its stored preset over the user's explicit selections. This produced an earlier "education is inert" finding — an artifact. With the bug fixed, education is the largest single demographic effect on benefit levels (though nearly neutral on the DC-vs-SS ratio, because it moves both sides together).

4. The Bonds tier was miscalibrated between the two engines. The deterministic engine compounded future bond years at a stated 6.6% average, while the Monte Carlo engine bootstrapped a stored series whose geometric mean was 3.92% — so the deterministic result sat above its own simulation's 90th percentile, and the app's two measures reached opposite conclusions for the same scenario. The stored bond series was itself inconsistent with the real U.S. Aggregate record and was corrected. After the fix the two engines agreed within about 2–5%, with the deterministic path sitting just above the Monte Carlo median, inside the P10–P90 band.

Deliberate assumption changes made during the audit

  • Wage growth (NAWI) lowered from 4.5% to 3.6%, implying about 1.1% real wage growth — matching the SSA's intermediate assumption.
  • Social Security benefits reported gross of Medicare premiums on both sides (the prior study netted Part B out of the SS side).
  • Expense ratios applied per tier (0.00%–0.61%).
  • Return rates now reflect each tier's geometric mean, the correct measure for long-run compounding, disclosed in-app and user-editable.

How the corrected engine was validated (July 2026)

  1. SSA replacement-rate benchmarks. The corrected model reproduced the published curve: low earner ≈ 58% (SSA ~55%), medium ≈ 40% (SSA ~40%), high/max ≈ 29% (SSA 27–33%).
  2. Cross-cohort invariance. At every income level tested, replacement rates varied ≤ 2.4 percentage points across birth years 1960–2000 — the same real earner now receives the same treatment regardless of generation.
  3. Wage-growth consistency. Same-age earnings across cohorts differed by exactly the assumed NAWI rate (ratio 2.028 over 20 years at 3.6%, matched to three decimals).
  4. Engine agreement. Deterministic versus Monte Carlo median came within ~2–5% (previously as much as 2.14× apart).
  5. Wage-growth sensitivity re-examined. An earlier draft reported the DC-vs-SS ratio was identical at 3.6% and 4.5% wage growth and treated this as a coherence signal. That sensitivity run was later found to be defective: on a consistent engine the ratio is not insensitive to wage growth — higher wage growth raises SS benefits (faster bend-point indexing) and lowers the DC compounding advantage, so the two legs move together and the ratio shifts materially. The headline ratio uses the SSA intermediate NAWI rate (3.6%); see Corrections & Retractions.

What did not change

The core structural claims survived, sharpened: the equity cliff is real (and now binary in the Monte Carlo data); the DC advantage rises with income, mirroring Social Security's progressivity; and the inheritable-balance asymmetry (a DC account can leave an estate; Social Security leaves nothing) remains — now honestly conditional on the investment posture.

For the specific claims we withdrew, see Corrections & Retractions.


Audit v2 (August 2026)

In August 2026 we commissioned a second independent audit against the live production app. The first audit tested the calculation engine; this one targeted everything around it: the withdrawal modes behind our headline claims, the PDF and share exports, the "Ask the Model" assistant, married-household edge cases, and the app's own regression defenses. It was run black-box against the production site across five build cycles, with every fix re-verified live by value.

The defects (as found, August 2026)

  1. The spousal top-up ignored claiming-age rules. For a single-earner couple, a spouse claiming the spousal benefit at 62 received the full 50% top-up (SSA reduces it to ~65% of that), and a spouse claiming at 70 appeared to receive delayed-retirement credits (SSA pays none on spousal benefits). The first half was a real engine defect — the reduction factor was keyed to the wrong age. The second half turned out to be a display defect: the first-year benefit was being deflated by the wrong calendar year, leaving up to three years of cost-of-living adjustments undeflated and masquerading as credits. Both fixed and verified: the claiming-age curve is now reduced-below-FRA and flat-above-FRA, matching SSA's rules.

  2. Scenario state could silently diverge from the controls. Editing one input could revert another; preset scenarios could hold a value while the visible control showed a different one; selecting a preset silently reset assumptions. In the worst case, the number on screen belonged to a scenario the controls no longer described. Fixed: all controls are now bound to a single scenario state, preset application is total and visible, and an automated guard edits every control and asserts the saved scenario matches the screen — including the controls that weren't touched.

  3. The PDF report mislabeled its own inputs. Most seriously, it printed "Trust Fund Shortfall: Disabled" on reports whose benefit figures demonstrably included the modeled 2033 cut — telling readers they were looking at scheduled benefits when they were looking at reduced ones. It also printed a vestigial "$0 peak income" field and miscounted household earners. All fields are now bound to the live scenario and verified in both shortfall modes.

  4. "Ask the Model" invented numbers. Asked for the model's benefit for a described person, the assistant produced confident dollar figures that the engine does not produce — off by +4.5% in one test case and −10.6% in another, in opposite directions, while claiming the numbers came "from the app's modeling engine." They did not; the assistant has no connection to the engine. It now declines to quote benefit figures for hypotheticals, says why, and routes you to the simulator — including when pressed for "just a ballpark."

  5. The survivor benefit skipped the 2033 cut. In benefit-cut mode, every post-2033 payment was reduced — except the survivor benefit, which paid the full scheduled amount (the discrepancy was exact: survivor rows equaled pre-death rows divided by the cut factor, to the dollar). Fixed: the shortfall reduction now applies to survivor benefits identically.

  6. Married preset cards briefly applied as single households. A regression introduced by the state-layer fix (#2) and caught in the same audit cycle: the married preset cards applied demographics but dropped the spouse. Fixed and added to the automated guard.

  7. Share links dropped the sender's assumptions. A shared link carried the scenario and strategy but silently recomputed under default assumptions (wage growth, inflation, display mode, career span) — and opening one overwrote the visitor's own saved settings. Fixed: share links now display a visible banner on open ("This link uses the sender's scenario with default assumptions for wage growth, inflation, and display settings") and no longer overwrite the visitor's persisted scenario without an explicit confirmation prompt.

Worth stating plainly: every engine-side defect this audit found — the spousal top-up, the survivor cut bypass — overstated Social Security, not the DC alternative. The errors ran against this site's thesis, which is exactly what you'd expect from a model that isn't tuned to its conclusion, and exactly why we publish audits either way.

How the corrected build was validated (August 2026)

  • The July golden-master values reproduced exactly on every one of the five builds tested — the four SSA reconciliation cases never moved.
  • The claiming-age factors (worker and spousal), the 1.5×/2.0× household identities, and the real/nominal conversion identities were re-verified by value on the final build.
  • The Social Security formula was independently re-implemented and checked: replacement rates decline smoothly and progressively with income, and the benefit flattens exactly at the taxable maximum.
  • The education effect — a college profile raising a near-retirement worker's computed benefit at the same current income — was specifically investigated and validated as legitimate: it corresponds to a back-inferred college/HS career premium of about 1.41×, inside the observed U.S. range of 1.6–1.9×. It is real modeling, not a bug, and the implied career average is now disclosed next to the education control.
  • The audit itself made and published five corrections to its own preliminary findings, under the same standard this page applies to the model.

What changed in our own defenses. The July audit's lesson was to validate against the SSA. The August audit's lesson was subtler: twice, an internal diagnostic passed while users saw the bug, because the diagnostic tested a function directly rather than the path the app actually runs. Every guard test now asserts on what a user can observe — the saved scenario, the rendered value, the assistant's actual reply — and the pre-publish gate runs the four SSA cases, the spousal claiming curve, the household identities, and the state-integrity sweep on every build.

What did not change. The bend-point formula, wage indexing, contribution mechanics, market-history data, Monte Carlo method, and every assumption default are unchanged from the corrected July engine. The headline conclusions — the equity cliff, the all-fixed-income loss to Social Security from both the equal-income and equal-wealth directions — reproduced live throughout, at every income tier tested, deterministic and Monte Carlo.

As with the July audit: this is a record of a point in time, not a certification that the current or any future build is defect-free.


About these figures — please read. Every number on this page describes the Is Social Security Worth It? model as of the August 2026 build. These are model estimates, not predictions, advice, or guarantees, and they describe the model on that date — not a promise about future versions. The model is revised over time and figures may change in later builds. Past performance does not guarantee future results. This is an educational tool; your actual benefit is set by the SSA at ssa.gov/myaccount. Per-cell data: DCvsSSv3audit.csv.