A half-joking interface for a serious uncertainty · v1.2.0

How it is computed

Model version
1.2.0
Released
19 Aug 2026
METR data cutoff
2026-05-08
Default anchor
Claude Opus 4.6 · 12.0 h

The short answer: from one number. Everything this service shows is derived from METR's autonomous task horizon — the length of work a model completes unaided. The default anchor is 12.0 hours at 50% success (Claude Opus 4.6), from METR data with a cutoff of 2026-05-08. Which measurement to anchor on is itself a choice, and the singularity page lets you change it — including to the frontier point METR considers unreliable.

From one measurement to the outputs

The same four kinds of number used across the site, in the order they feed into each other. The chain is short on purpose: if it were long, nobody could check it.

  1. MEASURED METR’s 50% time horizon: the length of task a frontier model finishes unaided, half the time. One number, with a source, a date and a confidence interval. Selectable on the singularity page.
  2. EXTRAPOLATED The doubling time carries that measurement forward. It continues an observed trend and is therefore still tied to data — but a continuation is not an observation, and nothing guarantees the trend holds.
  3. ASSUMED Everything else: transfer coefficients to other domains, deployment lags, rung weights, mitigation ceilings, trigger multipliers, misuse and control-failure pressure. Author’s judgement, every one of them on a slider. This is the group that carries the load, which is why the sensitivity table on the front page exists.
  4. OUTPUT Dates, cumulative probability curves and the clock. Derived quantities — they inherit every weakness above and add none of their own.

The load sits almost entirely in the third group. That is the argument for exposing it as sliders rather than burying it in code.

UNVERIFIED One dataset does not belong to the chain: the country table is a prototype that has not been reconciled with any source. It is hidden behind a button, labelled, and affects nothing unless you switch it on.

Extrapolation

D_eff = doubling time × friction × Π(trigger multipliers)
D(u) = D_eff × r^(u / Y), r = 1 + bend/100, Y = days in a year
log₂H(t) = log₂H₀ + Y × (1 − r^(−τ/Y)) / (D_eff × ln r), τ = t − t₀ in days

At bend zero the second expression collapses to the first divided by D_eff — a straight line on a log scale, which is what this page assumed until version 1.2.0. A positive bend makes the sum converge: the horizon approaches a ceiling of Y / (D_eff × ln r) doublings above the anchor, and any threshold above that ceiling is never reached in any year. Try it and watch the global risk go up — a slower trend keeps the window of vulnerability open longer rather than closing it.

The line is anchored at the METR point you selected, not fitted to history. Two consequences look like bugs and are not: the line diverges from the early METR points, and raising friction pushes future dates to the right while pushing already-passed ones further into the past.

The computation runs on the logarithm of the horizon rather than the horizon itself. This is not aesthetics: with the minimum doubling time and two accelerating triggers, 2^1590 by 2100 overflows a double, and the user would see NaN instead of a date.

Dates for the breakdown rows

date = date(coefficient × threshold × reliability) + lag × group multiplier

The difficulty coefficient is how many times longer a chain of reasoning a domain demands relative to software engineering, which is what METR measures. The deployment lag is the years for hardware, capital, trust and regulators after the capability technically exists. Raising the bar from 50% to 80% success multiplies the required horizon by 5: per METR, the 80% horizon is roughly five times shorter.

The rationale for each coefficient sits inside its row in the Singularity section — expand the row.

The ladder of catastrophes

λᵢ(t) = (malice·wᵢ + control failure·uᵢ) · cᵢ(t) · d(t) · (1 − mitigation·eᵢ) · aᵢ(t)

Four multipliers, each of which you set by eye. A product of such multipliers is accurate to an order of magnitude at best — every decimal place in these dates is a polite lie told by the interface.

The rungs are nested

A rung's curve is the probability of an event at that level or worse. The bottom rung therefore always sits above the top one: a global catastrophe by definition also clears "≥1,000 dead". In the prototype the rungs were computed independently, and under some assumptions the model claimed extinction was likelier than a local incident. That has been fixed.

The window of vulnerability

The aᵢ(t) multiplier damps risk after capability plateaus: the world learns to live with what has already happened. Without it every curve runs to 100% and the model stops meaning anything. Setting the window slider to 100 years effectively switches the decay off.

The count starts today

Integration begins in the current year, not a fixed one. Otherwise, a few years from now the service would be accumulating risk over years already survived, and every curve would creep upward from the mere passage of time.

The Doomsday Clock

position = min(1, log₁₀(P / 0.001) / 3), P = P(global) by 2100
minutes to midnight = max(0.2; 15 × (1 − position))

The dial is logarithmic: every tenfold increase in probability moves the hand 5 minutes closer to midnight. It used to be linear (15 × (1 − P)), which put the entire range people actually argue over — one per cent to fifty — inside three minutes of dial, so the main visual on the site barely responded to changing your mind. Fifteen minutes now means 0.1% or less; midnight means certainty.

The hand position is not the probability. On the old linear scale the two coincided, and a caption on the front page went on deriving one from the other after the switch — it announced a 76.9% risk where the model computed 20.1%. The probability is now carried explicitly from the model core and the interface never reconstructs it. The scale is invented here and bears no relation to the real Doomsday Clock of the Bulletin of the Atomic Scientists beyond the visual nod.

Red lines

This service deals in casualty figures. That imposes obligations, and they are kept:

  • no personal predictions and no geotargeted forecasts;
  • no specific company, laboratory or country is named as responsible for future events;
  • the scenarios on the catastrophe cards are described at the level of mechanism — "cascading failure" — with none of the operational detail anyone could read as instructions;
  • triggers are phrased as observable criteria, not as calls to action or predictions.

And the main thing

This apparatus does not predict the future and cannot. An event that has never occurred has no training sample, no validation, and no way to tell a good model from a pretty one. All that happens here is that assumptions get laid out so you can see which one carries the whole load. If the control-failure slider moves the global date by forty years and the doubling-time slider moves it by five, then the argument is about the first one. That is the entire value; the dates are a by-product.

Version history

Only changes that move the numbers or change what they mean are listed here. Layout, wording and infrastructure live in the commit log, not on this page. Each entry records the METR cutoff and the default anchor it computed from, because a date on this site means nothing without them.

1.2.0 19 Aug 2026 · METR data through 2026-05-08 · anchor Claude Opus 4.6 · 12.0 h

  • DATA All METR values re-read from the published measurement file instead of being transcribed by hand. The previous anchor was recorded as 320 minutes where METR reports 293.0, and the historical points mixed two measurement rounds. Fifteen points, one round, no approximations.
  • FEATURE The anchor is now a choice rather than a constant: four METR measurements, the data cutoff, the confidence interval and the source shown next to it. The default is the most recent measurement METR still stands behind; the frontier point is one button away, with a warning.
  • MODEL Severity by level now reports the probability of an event at exactly that level and nothing worse — the difference between adjacent cumulative rungs. The previous figure ignored competing risk and could not be summed across levels.
  • FIX The clock caption derived probability backwards from the hand position, which stopped being valid when the dial went logarithmic. It announced 76.9% where the model computed 20.1%. The probability now comes from the model core directly.
  • MODEL Baseline horizon doubling time 131 → 129 days, matching METR’s point estimate for the trend since 2023.
  • MODEL Trend bend: the doubling time can now change year on year, so a straight line on a log scale is a setting rather than an assumption baked into the code. A positive bend makes the horizon converge on a ceiling, and rows above it are never reached at all. It also shows something counter-intuitive — slowing the trend raises global risk by 2100, because the window of vulnerability stays open longer instead of closing.
  • FEATURE A survey-calibrated preset that differs from the author’s baseline by exactly one slider, and each preset button now shows the global risk it produces.

1.1.0 19 Aug 2026 · METR data through 2026-01-29 · anchor Claude Opus 4.5 · 320 min

  • MODEL The clock dial became logarithmic: five minutes of hand per order of magnitude. On the linear scale the entire range people argue over sat inside three minutes of dial.
  • FEATURE Sensitivity analysis on the front page: every assumption nudged by a fifth of its range, sorted by how far it moves global risk. This is the output the service actually exists for.
  • FEATURE Provenance labels — measured, extrapolated, assumed, unverified — on every slider and every dataset.
  • MODEL Sliders renamed to what they are: pressures and strengths on an arbitrary scale, not probabilities. “Probability of control failure — 20%” read as a claim about the world; it is a multiplier.

1.0.0 17 Aug 2026 · METR data through 2026-01-29 · anchor Claude Opus 4.5 · 320 min

  • MODEL First public model. Catastrophe rungs are nested — each curve means “an event at this level or worse” — and integration starts from the current year rather than a fixed one.
  • MODEL The horizon is extrapolated in log₂ space. At the fastest settings the linear form overflows a double before 2100.

Openness

The code is MIT; the model constants and texts are CC BY 4.0. A model whose assumptions cannot be inspected has no claim to being taken seriously, so every coefficient lives in one versioned file and changes through a pull request with a stated rationale.