The date is not the result. The sensitivity is.
This is an interactive stress test for AI futures, not a forecasting service. It starts from one measured quantity — how long a task frontier AI can finish on its own — extrapolates it, and then asks you for everything the extrapolation cannot answer: how fast capability gets deployed, how often it gets misused, whether a failure can be rolled back, how well any of it can be mitigated.
- It starts from measured dataMETR measures the length of software, ML and security tasks that frontier models complete unaided. That number is external, dated and sourced.
- You supply the assumptionsEverything after that measurement is judgement: transfer to other domains, deployment lags, misuse, control failure, mitigation. All of it sits on sliders instead of being buried in the code.
- The point is how much the answer movesMove one slider and a date can jump by decades; move another and it barely twitches. Which sliders matter is the actual output of this service. The dates are a by-product.
Most arguments about AI risk look like arguments about dates. They are almost always arguments about two or three hidden assumptions. This is a place to find out which ones.
The default scenario is the author’s, not a consensus. It puts global catastrophe risk near 20% by 2100, whereas surveys of AI researchers cluster around 5–10%. Treat it as one contestable position among several, and move the sliders.
To singularity
5 years 116 days
Date: December 13, 2031. The model-implied date when AI crosses the selected autonomous-task threshold, unaided, across 50% of 30 categories of activity. Not a claim about consciousness or general intelligence. Already crossed: 13%.
To the first catastrophe of any level
15 years 58 days
Date: October 16, 2041. Median date of the first event at the local level or worse. This is not the end of the world — it is the bottom rung of the ladder.
The totally unofficial AI Risk Clock
3 minutes 29 seconds
to midnight. The model puts global catastrophe by 2100 at 20.1%. The dial is logarithmic: five minutes of hand per order of magnitude of probability, so the range people actually argue over is visible instead of bunched up against the top. The hand position is not the probability — read the percentage. Not affiliated with the real Doomsday Clock of the Bulletin of the Atomic Scientists.
Serious
What actually drives this result?
Each assumption is nudged by a fifth of its range in both directions, with everything else held where you left it, and the bar shows how far global catastrophe risk by 2100 moves. Long bar means the argument is about that assumption; short bar means it is not worth having.
| Assumption | Effect | global catastrophe risk by 2100 |
|---|---|---|
| Control-failure pressure | ±14 pp | |
| Window of vulnerability | ±12 pp | |
| Time to deployment saturation | ±3.1 pp | |
| Mitigation strength | ±2 pp | |
| Misuse pressure | ±2 pp | |
| Real-world friction | ±1.1 pp | |
| Horizon doubling time | ±1.1 pp | |
| Trend bend | ±0.9 pp | |
| Current wiring into critical systems | ±0.8 pp | |
| Singularity threshold | below 0.1 pp |
What this is
- A model, not a forecast. One extrapolation — METR’s autonomous task horizon — turned into two countdowns and a three-rung ladder of catastrophe.
- Probabilities by year. Other p(doom) calculators give a single number with no time axis; here every level has a curve from today to 2100.
- Three scales, told apart. A thousand deaths and the end of the species are not the same event at different odds, and the model refuses to average them.
- Recomputed from your assumptions. Every constant that carries weight is a slider, and each default says where it came from.
- Open. Code under MIT, model constants under CC BY, every coefficient in one versioned file you can argue with through a pull request.
What this is not
- Not a prediction. An event that has never happened has no training sample, no validation, and no way to tell a good model from a pretty one.
- Not a position. The service does not argue that AI is dangerous or that it is safe. It hands you the apparatus and stays out of the conclusion.
- Not a date. If a slider moves the answer by forty years, the date was never the answer — its sensitivity is.
- Not a measurement. Difficulty coefficients, deployment lags and rung weights are expert judgement, and the weakest of them is labelled as such.
- Not advice. No personal or geographic predictions, no survival guidance, and no named company, laboratory or country blamed for anything.
Every number on this site is one of four kinds:
- MEASURED
- An external observation with a source and a date.
- EXTRAPOLATED
- A mathematical continuation of an observed trend, not an observation.
- ASSUMED
- An expert judgement by the author of the model. Contestable by design.
- UNVERIFIED
- Draft data that has not been reconciled with any source. Treat as illustrative.