Research · September 2026

How bad would Y2K actually have been?

The world spent somewhere between $300 and $600 billion preventing the Millennium Bug, and then spent twenty-six years arguing about whether it needed to. Nobody ever wrote down the counterfactual. So I did.

Method: seven parallel research agents working primary sources (OMB, GAO, NERC, NRC, FDA, the Senate Special Committee, the Federal Reserve), then a hand-built pathway model calibrated against measured infrastructure failures with published mortality studies. Every figure below is traced to a primary document or explicitly flagged as unverified. Search-engine summaries were treated as zero evidence.

The answer, up front

Counterfactual: near-zero remediation. No awareness campaign from 1993, no programs, no testing, no contingency planning, no extra rollover staffing. Log scales.

Global direct economic damage, first 12 months

~$400 billion

50% CI: $150bn to $900bn (1999 USD). About 1.2% of 1999 world GDP.

actually spent: $300-600bn
$400bn
$50bn
$100bn
$500bn
$1tn
$2tn

50% credible interval median estimate what was actually spent preventing it

Global excess deaths, first 12 months

~1,800

50% CI: 300 to 9,000. Actual documented Y2K deaths in the real, remediated world: zero.

1,800
100
1,000
10,000
100,000

The right tail runs past 100,000 and is driven by two scenarios, neither inside the 50% interval: a multi-week winter grid failure, and a false nuclear early-warning signal.

The finding that should bother everyone is not that Y2K was a hoax. It wasn't. The finding is that the median damage from doing nothing and the actual cost of the prevention program land in the same range, which makes the program's net value genuinely ambiguous, and in twenty-six years nobody has done the arithmetic that would settle it.

1 What was actually done

US federal government spend was $8.38 billion across FY1996 to FY2000, the terminal figure in OMB's eleventh and final Quarterly Report to Congress. That report's footnote 2 matters more than the number: it explicitly excludes "the costs of upgrades or replacements that would otherwise occur as part of the normal systems life cycle." The federal figure is genuinely incremental.

US economy-wide spend was about $100 billion, or $365 per resident, peaking near $30 billion a year in 1998 and 1999, per the Commerce Department's Economics and Statistics Administration. Unlike OMB, Commerce does not separate incremental from accelerated spend. It names the problem and declines to quantify it: some of the spending "may involve 'shifting forward' new, productive software and hardware investments which would have occurred eventually." That unquantified share is the biggest hole in every cost-benefit claim anyone has ever made about Y2K, including mine.

Globally, Gartner's $300bn to $600bn range is the most-cited figure and checks out against two independent primary sources. Capers Jones of Software Productivity Research put it at over $1.6 trillion, and the Senate report explains the divergence precisely: Jones included over $300 billion for litigation and damages, Gartner did not. Hold that thought.

The institutions

Executive Order 13073 created the President's Council on Year 2000 Conversion under John Koskinen in February 1998, running 25-plus sector working groups with 250-plus partner organizations. Two statutes did the legal work. The Year 2000 Information and Readiness Disclosure Act shielded good-faith readiness disclosures so firms would share status instead of lawyering up. The Y2K Act capped punitive damages in Y2K civil actions.

What the regulators actually concluded

These documents are the most useful thing in the whole file, because they contain the regulators' own risk language rather than press paraphrase.

Revealed preference

The Federal Reserve's own behavior is the cleanest instrument available. Currency in circulation ran $535.7bn in August 1999, $595.97bn in December, and back to $566.1bn by February 2000: a real precautionary build of about $60 billion. The Fed also stood up a Special Liquidity Facility and a Standby Financing Facility that auctioned $481 billion in repo-call options.

None were exercised. Peak daily borrowing on the liquidity facility was $1.2 billion. The central bank built a half-trillion-dollar backstop and the market used roughly a quarter of one percent of it.

2 What actually broke

The catalogue of verified rollover failures is short, and its severity distribution is the single most important fact in this analysis.

FailureConsequenceDuration
NRO satellite ground station, Fort BelvoirDegraded intelligence processing, ran on backup~2 days
Medicare claims processing50,475 claims from 872 submitters dated 1900 or 2099, all originating in providers' own unpatched systemsReprocessed
FAA Low Level Wind Shear Alert SystemFailed at 8 sites. GAO noted it "could have affected aviation operations if weather conditions had been severe"~2 hours
FAA Kavouras weather displayReported the year as 2010, rejected National Weather Service data, broke 13 flight service stationsHours
Japanese nuclear plants (four sites)Radiation and meteorological monitoring, control-rod position display, prefectural data feeds10 min to 5 hrs
Sheffield NHS pathology labDown's syndrome screening miscalculated maternal age. 150+ women given wrong risk results, 4 affected births, 2 terminations~5 months
Delaware racetrack slot machinesOfflineOne evening

Notice the pattern. Every verified failure produced wrong output or lost monitoring. Not one produced loss of control of a physical process. The Sheffield case is the most serious documented harm anywhere in the record, and its mechanism is date arithmetic handing a clinician a wrong number, not a machine misbehaving. Koskinen's own contemporaneous tally was "about 90 glitches and failures around the world."

No Y2K-attributable death has ever been documented, anywhere. That survived deliberate adversarial searching.

The counter-evidence, which is real

Martyn Thomas, who led Deloitte's international Y2K practice, lists control-system defects found in testing: the UK Rapier missile system would have failed to fire; a Swedish reactor's computers shut the reactor down when clocks were advanced to 1 January 2000; BP Exploration found an oil-platform error its consultants said justified the entire program; the US Coast Guard found bugs that would have caused loss of ship steering control and fire-system failure. These are genuine existence proofs that safety-critical software carried real date defects, and they are the strongest argument for the remediation program, because testing is exactly what found them.

But read the Swedish reactor case carefully. The failure mode was an unplanned safe shutdown. That is fail-safe design working as designed, and it generalizes. Safety-critical systems are overwhelmingly built to fail toward the safe state, so a date bug in one produces an outage rather than a catastrophe. Which means the loss-of-life pathway does not run through spectacular single-system failures. It runs through aggregate loss of service.

3 The natural experiments

Low-spend countries

The skeptics' favourite argument, and the data is messier than either side admits. Gartner sorted countries into four tiers by predicted share of organizations suffering at least one mission-critical failure. Tier 1 at 15% included the US and UK. Tier 4, at 66%, included China and Russia. That is a falsifiable prediction from the era's most-cited authority, and it was wrong everywhere.

But the Meta Group spending table in the same Senate report undercuts the simple story. Russia's estimated remediation cost was 7.3% of 1996 GDP, the highest share in the entire twenty-country table, above the US at 2.5%. Italy came in at 2.8%, essentially identical to the US. South Korea at 4.7%. Only China, at 0.5% of GDP, genuinely matches the "spent nothing and shrugged" story. Italy, the country skeptics name most often, was not actually a low-spend country by this measure, though it was far behind on execution.

The two camps agree on the fact and disagree on the mechanism, and neither has tested their explanation. Thomas argues late starters free-rode on remediation tools and supply-chain fixes paid for by early movers. The economist John Quiggin rebuts this directly: free-riding "seems inconsistent with an account in which detailed checking of vast numbers of individual devices was crucial," and if it worked, Australia should have copied Italy. He also dismisses the "less computerized" defense as "absurd in relation to OECD countries like Italy." This is the crux of the entire debate and it remains unresolved in the literature.

The control group nobody measured

NFIB's 1998 survey found roughly 40% of US small businesses were aware of their Y2K exposure and planned no action, covering an estimated 330,000 computer-dependent firms. A large, real, deliberately unremediated population. What happened to them?

The President's Council's own final report admits it does not know: "The sheer number of these companies, over 23 million, and the absence of regular reporting relationships... also made it difficult to determine how many actually experienced Y2K difficulties after the date change." Nobody measured the control group. That is the central evidentiary failure of the entire episode.

Cambridge, and the GPS rollovers

Ross Anderson published the one bottom-up empirical study by a skeptic, two weeks before the rollover. He audited Cambridge University across medical departments, labs, catering, accommodation and an aerial photography unit: 7,000 employees, £300m turnover. Total spend was £20K central plus £47K on upgrades, with about £100K of noncompliant embedded systems found, "most of these were medical devices." His finding: "none of the bugs found so far would have had a serious impact on business operations even if they had been ignored until January 2000." His conclusion: "The average small business owner, who is doing absolutely nothing about the bug, appears to be acting quite rationally."

The cleanest natural experiment is also the most underused. GPS week-number rollovers are real, unavoidable, date-related failures in a safety-critical global system, with a known population of non-compliant receivers and no possibility of a "the fix worked" explanation. The April 2019 rollover produced: Honeywell flight management software failures causing flight cancellations in China and a KLM delay, because technicians had not applied the patch; the NYCWiN municipal wireless network crashing; and Vaisala weather-balloon groundstations suspending launches for up to two weeks. That is what genuinely unremediated date rollover looks like in a modern system. Real, concrete, disruptive, bounded, fixed afterward.

The confound that contaminates all of it

A large share of critical infrastructure was not running normally at the moment of rollover, regardless of remediation status. In a peer-reviewed survey of New York City emergency department directors, 87% of hospitals activated their disaster plan in advance, 74% ran a live mock Y2K drill, and 90% stood up an Incident Command System. Airlines cancelled a large share of New Year's flights. Gaseous diffusion plants planned to shut down entirely.

So "nothing happened" at the highest-stakes institutions is evidence about remediation plus extraordinary manual fallback, combined. Public data cannot separate the two.

4 Who put a number on it, and how they scored

This calibrates how much to trust anyone's counterfactual, mine included.

ForecasterThe predictionWhat happenedVerdict
Ed Yardeni
Chief economist, Deutsche Bank Securities
"I believe there is a 70% chance of such a worldwide recession, which could last 12 months starting in January 2000" (March 1999). Severity benchmark: the 1974-75 oil-shock recession. Recanted 3 January 2000: "As it appears now, there will be no global recession as a result of Y2K." Miss
Capers Jones
Software Productivity Research
Over $300 billion in Y2K litigation and damages, included in his $1.6tn global total. GAO counted 18 cases invoking the Y2K Act, out of 42 total raising Y2K issues. Miss
Gartner Group Country risk tiers: 15% of organizations failing in the US and UK, 66% in China and Russia. No material difference in outcome by tier. China, the genuine low-spender, was fine. Miss
NERC "Credible worst case": area blackout. "Probable": 10-15% generation loss, partial SCADA loss. No grid impact. But the ceiling they named was already far below the public fear. Calibrated
Ross Anderson "Lots of things may break, but few of them will actually matter much" (December 1999), from a bottom-up audit. Correct, and correct for the stated reason, published before the event. Hit

Yardeni's 70% deserves a second look, because it is routinely misread. His forecast was conditional on the remediation that was actually happening. He was not forecasting the do-nothing world. He was forecasting ours, and he was badly wrong about it. That fact should pull any counterfactual estimate downward, hard, which is why his implied number sits near my 90th percentile rather than my median.

5 The counterfactual

Two counterfactuals give wildly different answers, and conflating them is how the public argument stays stuck.

Almost nobody serious argues A would have been fine. The live argument is about B.

Base rates

Converging estimates of the share of date-handling systems carrying a real defect: research cited by Thomas found ~5% of embedded systems failing compliance tests, manufacturing ~15%, sophisticated systems 50-80%. FDA found 8% of biomedical manufacturers reporting date problems, and GAO's independent audit found ~12% of products noncompliant. The President's Council noted that "over 70 percent of companies reported they had experienced Y2K glitches" before the rollover.

Working figure: 5% to 15% of date-handling systems carried a real defect, with severity heavily skewed toward wrong output rather than loss of control.

For scale: the 2024 CrowdStrike outage hit about 8.5 million machines, roughly 1% of Windows endpoints, for one to two days, and cost $5.4 billion in direct losses to the US Fortune 500 alone. Counterfactual A is a five to fifteen times larger blast radius with a much longer manual-remediation tail, but with far less of it landing on systems whose failure stops a business cold.

The economic estimate

My median of $400bn sits between two anchors. Yardeni's implied number for a 1974-75-severity global recession is roughly 3% of GDP, about $1 trillion, and belongs near my 90th percentile given how the forecast scored. At the other end, IDC's own contemporaneous counterfactual was $237bn in global lost revenue avoided by $308bn of spend, meaning IDC scored the program as a $71bn net loss, from an interested party with every incentive to inflate the benefit.

The uncomfortable arithmetic

Under counterfactual A you avoid $300bn to $600bn in remediation. Against a median $400bn in damage, the median net cost of having done nothing is approximately zero. If the true damage sits toward my lower bound, doing nothing was cheaper.

This is not a fringe position. It is what IDC concluded at the time, with their own numbers.

One caveat cuts the other way, and Finkelstein, the loudest skeptic, is the one who makes it: "I believe that most organisations that spent heavily on Y2K were justified in doing so! Not because of Y2K itself but rather because an organisation that does not know which are its business critical systems... is headed for trouble regardless of date." A meaningful share of the spend bought system inventories, documentation and legacy knowledge with independent value. Nobody has ever measured how much.

The mortality estimate

Built by pathway, because the pathways have very different evidence quality.

Global excess deaths by pathway, counterfactual A

Bars show the 50% credible interval, the vertical rule shows the median. Log scale.

Healthcare delivery100 to 3,000 · med. 600
Grid, heat, water50 to 8,000 · med. 700
Secondary and behavioral100 to 2,000 · med. 400
Transport10 to 300 · med. 50
10
100
1,000
10,000

Healthcare is the best-identified pathway. Neprash, McGlave and Nikpay (American Economic Journal: Economic Policy, 2026) use a quasi-experimental event study of hospital ransomware attacks against Medicare claims and find hospital volume falls 17-25% in the attack week and in-hospital mortality rises 34-38% among patients already admitted. Applied to a US inpatient census of ~500,000 with ~2% baseline in-hospital mortality, a week of degradation at 30% of hospitals gives roughly 1,000 excess US deaths. Sheffield proves the mechanism reaches clinical decisions.

But WannaCry is the negative control, and it matters: a real national NHS IT failure costing £92 million with 19,000 cancelled appointments, and the National Audit Office attributed no deaths to it. The pathway is real but not automatic.

Grid is the fattest tail and the weakest evidence. Rate anchors: Texas winter storm Uri, 11 million affected for ~3 days, produced about 700 excess deaths by excess-mortality methods against 246 officially counted, roughly 64 deaths per million per three-day winter outage. The 2003 Northeast blackout, 50 million affected for ~2 days in summer, produced about 90 excess New York City deaths. Hurricane Maria, 3.4 million people with a median 84-day outage, produced 2,975 to 4,645 excess deaths depending on method.

A January rollover is the worst possible date for this, since winter outage kills faster per day than summer outage. If 100 million people in the northern hemisphere lost power for four days at ~15 deaths per million per day, that pathway alone gives ~6,000 deaths.

So why is my median only 700? Because I put the probability of a multi-day outage affecting more than 10 million people in a developed country at roughly 20-30%, on four pieces of evidence: NERC's own credible worst case was area blackout, not systemic collapse; generation is overwhelmingly mechanical; protection relays are largely analog or use relative rather than absolute time; and grid operation is human-in-the-loop with manual fallback by design. The discriminating test is China, which genuinely under-invested at 0.5% of GDP and had no grid failure. That is a real test on exactly the pathway that dominates the estimate, and it came back negative.

The tail outside the interval

Two scenarios sit outside the 50% CI and deserve naming. A two-to-four-week grid failure in a cold region of 10 million-plus people gives tens of thousands of deaths on Maria-rate scaling; I put this in the low single-digit percent.

And the one nobody talks about. The US and Russia thought a false nuclear early-warning signal credible enough to put Russian officers in a room in Colorado Springs for a month. Probability very low, magnitude in the millions. It is the single strongest argument that the program was worth the money regardless of the expected-value arithmetic, because you do not buy insurance against the median.

6 What I actually think

The defect was real and the engineering work was legitimate. Thomas is right that "it was a hoax" is an ignorant position. The Rapier missile system, the Coast Guard steering bugs, FDA's thousand-plus noncompliant device models and the Sheffield screening failure are not imaginary. Peter de Jager, who raised the alarm in 1993, holds up well: "We had a problem. For the most part, we fixed it. The notion that nothing happened is somewhat ludicrous."

The scale of the response was very probably larger than the risk justified, and this was knowable at the time. Anderson published a bottom-up audit saying so in December 1999, with data. Finkelstein named the incentive structure: IT managers using Y2K "as a loophole" to fund overdue legacy maintenance, consultants who "recognised that potential consequences of Y2K had been exaggerated but felt that, on balance, the panic was good for business," and a press for whom the story was "too good to check." Quiggin adds the mechanism that best explains the country pattern: in common-law jurisdictions, fix-on-failure carried personal blame risk while over-remediation carried none, which is why high-spend countries cluster in the English-speaking world rather than tracking actual exposure.

Nobody can settle this, because the control group was never measured. The President's Council said so in its own final report. Twenty-six years on, the two camps still agree on the facts and disagree on the mechanism, and neither has run the comparison that would discriminate between them. Australia spent roughly $12 billion and produced a seventeen-page self-congratulatory final report with no cost-benefit analysis. That is the actual scandal, and it is a procedural one.

The transferable lesson is not "prepare more" or "panic less." It is Quiggin's: institutionalize adversarial forecast review and mandate ex-post program evaluation, so that next time a large preventive expenditure is proposed, somebody is paid to argue the other side and somebody is required to check the score afterward. Y2K produced neither. Koskinen saw the bill coming twenty years early:

"Perhaps one thing about Y2K was that we did it too well... So when a new, urgent crisis arises, there is a risk that people might say, 'They said that last time about Y2K, but that really never was a problem.'"

John Koskinen, chair of the President's Council on Year 2000 Conversion

7 Confidence and limits

Verified at primary source: all OMB, GAO, Commerce, Senate, NERC, NRC, FDA and FAA figures above; Yardeni's 70% and his recantation; the GAO litigation count; Federal Reserve currency and liquidity-facility figures; the Neprash, Anderson & Bell, Kishore and GW Milken mortality studies; the Thomas, Finkelstein, Anderson and Quiggin retrospectives.

Could not verify, and did not rely on: Gartner's 10/35/55 failure-timing split; the "50 billion embedded chips" figure, which traces to no Gartner document and should be treated as folklore; Yardeni's alleged downward revision to 45%, likely a conflation with his 2019 tariff call; Lloyd's-specific Y2K exposure numbers; Japan's ministry-level incident counts.

The biggest weakness in my own estimate is the grid pathway, which dominates the tail and rests on a probability judgment rather than data. If you think NERC's credible worst case was optimistic, or that the China comparison is confounded by low computerization, the mortality median moves up by a factor of three to five and the interval widens with it.

Second weakness: every figure is in 1999 dollars, and the economic estimates mix direct-loss and direct-plus-indirect methods across analogues. Uri shows how much that matters: $4.3bn on a narrow value-of-lost-load basis, $80-130bn on the Dallas Fed's direct-plus-indirect basis, or $197-296bn on Perryman's full multiplier, for the identical storm. A fiftyfold spread from method alone.

Sources

Primary government. OMB, 11th Quarterly Report on Y2K Conversion (Dec 1999) · Commerce ESA, The Economics of Y2K and the Impact on the United States (Nov 1999) · Senate Special Committee, Investigating the Impact of the Year 2000 Problem (S.Prt. 106-10) · GAO/AIMD-00-290, Lessons Learned · GAO/T-AIMD-00-26, Biomedical Equipment · GAO/T-AIMD-00-27, Nuclear Power · GAO/AIMD-99-114, Electric Power · GAO/GGD-00-196R, Y2K Act court cases · Federal Reserve series CURRCIR · FRBNY, Current Issues in Economics and Finance 6(15), Dec 2000.

Retrospectives. Martyn Thomas, The Guardian (Dec 2019) · Thomas, Gresham College lecture (Apr 2017) · Anthony Finkelstein, Computing & Control Engineering Journal 11(4), 2000 · Ross Anderson, The Millennium Bug: Reasons Not to Panic (Dec 1999) · John Quiggin, The Y2K Scare: Causes, Costs and Cures, University of Queensland (2004) · Kevin Kliesen, FRB St Louis Review 85, 2003.

Calibration studies. Neprash, McGlave & Nikpay, AEJ: Economic Policy 18(1), 2026 · Anderson & Bell, Epidemiology 23(2), 2012 · Kishore et al., NEJM, 2018 · Santos-Burgoa et al., GW Milken Institute, 2018 · Texas DSHS final report, Dec 2021 · FRB Dallas (Golding, Kumar & Mertens), Apr 2021 · Parametrix CrowdStrike report, 2024 · UK DHSC WannaCry cost report, Oct 2018.