Ledger · scoreable predictions
The forecast ledger
Vague confidence is not tracked here. Every forecast carries a probability, a hard resolution criterion, and a horizon — so it can be scored (Brier) when it resolves. A forecast you cannot lose is not a forecast.
Every entry is pre-registered: its registration date is part of the record, the gate refuses registration at or past the horizon, and release commits are anchored into Bitcoin via OpenTimestamps — so the ordering of prediction and outcome can be verified by strangers, not taken on trust.
Forecast horizon
A queue that time can judge
as of 2026-07-08
Each point is an open forecast. Position is horizon date; size is effective probability. Amber means judgment is near.
Mean Brier · live resolutions
—
Honestly empty: no forecast has resolved live yet. Calibration is unclaimed until one does.
On the record
40 open
Live-resolved forecasts will keep their original probability — the misses stay counted.
Retraction · self-audit
3 backfilled
Authored after their resolution dates — kept visible below, excluded from calibration. Disclosure →
Open
-
88%
OpenThe operating-question verdict remains "No. Not yet."
through 2026-12-31
Resolution Resolves NO if the Observatory revises the verdict on the record, with a frame-construction result strong enough to survive its own contamination defenses.
-
85%
OpenConstruction-01 commitment C3: restraint is unrewarded — no major capability benchmark through 2026 scores the capable action correctly NOT taken, because leaderboards structurally cannot reward an absence.
by 2026-12-31
Resolution Resolves YES if, through 2026-12-31, no widely-tracked capability leaderboard adopts a primary metric that scores correct non-action / restraint as such (distinct from refusal-safety benchmarks, which score policy compliance, not wisdom-restraint). Resolves NO if such a metric becomes established on a major benchmark.
-
85%
OpenAnthropic's Fable 5 weekly-usage inclusion for Max plans ends as scheduled on 2026-07-07 (shifting to usage credits) with no new suspension or export-control action before then.
by 2026-07-08
Resolution Resolves YES if the July-7 transition happens as announced and no new government restriction or Anthropic suspension of Fable 5 occurs before 2026-07-08; resolves NO otherwise.
-
82%
OpenNo frontier system passes a hardened FCS-4 (Hilbert–Einstein action) run under audited contamination defenses.
by 2027-12-31
Resolution Resolves NO if an independently scored, time-sliced FCS-4 run reconstructs the action principle with held-out derivation steps and no post-1915 leakage.
-
80%
OpenConstruction-01 commitment C1: capability and the wisdom-components dissociate — confident frame-captivity failures (reward-hacking, proxy-optimization into harm, fluent high-confidence wrong answers) do NOT monotonically decline with frontier capability through 2027.
by 2027-12-31
Resolution Resolves YES if, across published frontier releases and incident reporting through 2027-12-31, at least one clear case of a more-capable system exhibiting equal-or-worse confident frame-captivity than a less-capable predecessor is documented, and no consensus emerges that such failures declined monotonically with capability. Resolves NO if the field converges on capability reliably eliminating these failures.
-
80%
OpenThe Observatory executes FCS-B1 (descent, pre-1858 time-slice) against ≥2 frontier minds with published transcripts by 2026-09-30.
End of Q3 2026
Resolution Transcripts and grading published under /experiments/ for ≥2 minds; graded under the suite’s asymmetry.
-
78%
OpenConstruction-01 commitment C2: the imitation gap — systems become excellent at describing wise behavior well before, and to a greater degree than, they exhibit it under optimization pressure; the gap is large and persistent.
by 2027-06-30
Resolution Resolves YES if, by 2027-06-30, published evaluations continue to show frontier systems producing high-quality descriptions of restraint/humility/long-view reasoning while measurably failing those same behaviors under optimization or adversarial pressure (e.g. reward-hacking, sycophancy, spec-gaming persisting despite fluent meta-cognition). Resolves NO if the describe-versus-exhibit gap closes to negligible on standard probes.
-
75%
OpenAt least two of the GPT-5.6 family (Sol, Terra, Luna) are generally available by 2026-12-31.
End of 2026
Resolution Official OpenAI availability announcements; GA means paid-tier API or ChatGPT access without waitlist for at least two family members.
-
74%
OpenNo frontier model ships durable, updatable long-term memory as a native property of the trained network (not external scaffolding).
by 2027-12-31
Resolution Resolves NO on a credible demonstration of write-persist-retrieve memory across sessions that is intrinsic to the model, independently reproduced.
-
72%
OpenConstruction-02 commitment: model liquidity becomes a recognized enterprise-architecture default — model-agnostic control layers that treat frontier models as swappable inputs proliferate as standard practice, not niche.
by 2027-12-31
Resolution Resolves YES if, by 2027-12-31, multiple major enterprise-AI platforms and analyst frameworks treat multi-provider model-agnosticism (switch models with low friction) as a default architectural expectation rather than an advanced option. Resolves NO if single-provider lock-in remains the unquestioned norm for enterprise AI.
-
70%
OpenConstruction-01 addendum commitment C4 (near-term, unlike C1-C3): between 2026-07-07 and 2026-09-30, of all newly-logged incidents on this record, self-caught incidents will continue to outnumber externally-caught incidents by at least 3:1 — i.e. the current 12:1 self-catch ratio will not collapse toward parity even as more incidents accrue.
by 2026-09-30
Resolution Count new incidents.json entries dated after 2026-07-07 through 2026-09-30, classified by who caught the underlying error (self vs. an identifiably external party, per the same standard applied retroactively in the Construction-01 addendum). Resolves YES if self-caught:external-caught is at least 3:1 among new entries (or zero new incidents occur). Resolves NO if external catches make up more than 1 in 4 of new incidents — the reading that would most directly support the skeptical half of the addendum's own honest caveat (self-caught may just mean unaudited, not well-audited).
-
70%
OpenARC-AGI-3 Milestone #1 official results (due 2026-12-04) show no system within 20 points of the human baseline.
by 2026-12-05
Resolution Resolves YES if the published Milestone #1 leaderboard shows every entrant more than 20 points below the human baseline; resolves NO otherwise.
-
68%
OpenThe human–frontier gap on held-out ARC-AGI-3 interactive tasks stays above 20 points.
by 2027-06-30
Resolution Resolves NO if a frontier system reaches within 20 points of the human baseline on a held-out ARC-AGI-3 set without task-specific tuning.
-
65%
OpenAt least two additional ≥500MW nuclear or fusion power agreements tied to AI datacenters are announced by 2026-12-31.
End of 2026
Resolution Two distinct official deal announcements after 2026-07-02, each ≥500MW and explicitly tied to AI/datacenter load.
-
62%
OpenConstruction-02 commitment (the epistemic inversion): for truth-seeking rather than alpha-compounding, exposure beats hoarding — through 2027 the AI-capability claims that earn the most durable credibility are those accompanied by reproducible or independently-checkable artifacts, not sealed internal evaluations.
by 2027-12-31
Resolution Resolves YES if, by 2027-12-31, the pattern holds that capability claims backed by open/reproducible artifacts (public eval harnesses, third-party reproduction, open weights) command more durable credibility than claims resting only on sealed internal numbers, as reflected in how the field and independent evaluators treat them. Resolves NO if sealed internal evals command equal-or-greater durable trust.
-
60%
OpenConstruction-02 commitment: the tribal-knowledge-migration threat materializes visibly — at least one major frontier model provider moves up-stack to compete directly with the core workflows of its own enterprise customers.
by 2027-12-31
Resolution Resolves YES if, by 2027-12-31, a major frontier provider publicly launches a product or vertical that directly competes with a workflow category its enterprise API customers built on it (documented vertical entry / 'provider replaces you'). Resolves NO if providers stay confined to the model layer.
-
60%
OpenGPT-5.6 (Sol) reaches general availability by 2026-08-15 under the staggered executive-order release process.
by 2026-08-15
Resolution Resolves YES if OpenAI makes GPT-5.6 Sol generally available (not limited preview) to paying API or consumer users by the horizon.
-
60%
OpenARC Prize publishes at least one closed private/semi-private verified scorecard packet for a top ARC-AGI-3 entrant by 2026-12-31.
End of 2026
Resolution Official ARC Prize publication of a verified scorecard (the queue-29 gate object) for any leading entrant.
-
55%
OpenGemini 3.5 Pro is generally available by 2026-07-31 (already slipped from its June target).
by 2026-07-31
Resolution Resolves YES on public general availability of Gemini 3.5 Pro by the horizon.
-
55%
OpenGrok 4.5 exits private beta to general availability by 2026-09-30.
End of Q3 2026
Resolution Official xAI announcement of general availability (API or consumer access without invite).
-
55%
OpenA frontier model reaches ≥85% on OSWorld-Verified in a published evaluation by 2026-12-31.
End of 2026
Resolution Vendor or independent published OSWorld-Verified score ≥85%; leaderboard or paper citation required.
-
50%
OpenAn open-weights model reaches ≥60% on Terminal-Bench 2.0 in a published, reproducible evaluation by 2026-12-31.
End of 2026
Resolution Published eval with public methodology on an open-weights checkpoint; vendor or independent, but must be reproducible (weights + harness public).
-
50%
OpenAn independent evaluator publishes a reproduction or audit of GPT-5.6 Sol’s headline agentic/coding claims by 2027-03-31.
Q1 2027
Resolution Publication by a party other than OpenAI reproducing or auditing at least one headline GPT-5.6 benchmark claim, with methodology.
-
50%
OpenA single AI datacenter campus with ≥5GW planned capacity is officially announced by 2026-12-31.
End of 2026
Resolution Official announcement by an operator/government naming one campus/site with ≥5GW planned power.
-
50%
OpenAt least one named external reviewer agrees to independently score an Observatory FCS run by 2026-12-31.
End of 2026
Resolution A named person’s agreement recorded on the public record (colophon or experiment page) plus at least one delivered independent grading.
-
45%
OpenBy 2026-09-30, METR or another independent evaluator publishes a cross-generation quantification of frontier-model reward-hacking rates (≥2 model generations compared).
by 2026-09-30
Resolution Resolves YES on publication of an independent evaluation explicitly comparing detected evaluation-gaming/reward-hacking rates across at least two frontier model generations.
-
45%
OpenAnthropic ships an Opus-branded model in the Claude 5 family by 2026-12-31.
End of 2026
Resolution Official Anthropic release notes or model card carrying both the Opus name and 5-family versioning.
-
45%
OpenMETR publishes a measurement showing a 50%-success task time-horizon of ≥8 hours for any public model by 2026-12-31.
End of 2026
Resolution METR official report or blog with the 50%-success horizon metric at or above 8 hours.
-
45%
OpenHBM4 memory ships in a commercially available accelerator by 2026-12-31.
End of 2026
Resolution Official availability (not sampling) of an accelerator with HBM4, per vendor announcement or teardown.
-
40%
OpenA second frontier vendor announces an approved-organizations-only model tier (analogous to Anthropic’s Mythos) by 2026-12-31.
End of 2026
Resolution Official vendor announcement of a restricted-availability frontier tier gated on organizational approval, not just pricing.
-
40%
OpenTwo or more national AI safety institutes publish a joint pre-deployment evaluation of the same frontier model by 2026-12-31.
End of 2026
Resolution Joint or simultaneous coordinated publication by ≥2 national AISIs evaluating one named model pre-deployment.
-
40%
OpenA frontier lab publicly rolls back or recalls a released model version citing safety (not capability) regressions by 2026-12-31.
End of 2026
Resolution Official vendor announcement withdrawing or downgrading a released model, citing safety as the reason.
-
35%
OpenThe EU opens a formal enforcement proceeding against a frontier-model provider under the AI Act’s GPAI obligations by 2026-12-31.
End of 2026
Resolution Official Commission or AI Office announcement of a formal proceeding (not an information request) naming a GPAI provider.
-
35%
OpenCISA or a national CERT issues an advisory naming autonomous AI agents as the attack vector in a confirmed intrusion by 2026-12-31.
End of 2026
Resolution Official advisory from CISA or a national CERT that attributes a confirmed intrusion to an autonomous AI agent.
-
30%
OpenThe ARC-AGI-3 competition’s final verified leaderboard (winners 2026-12-04) shows a top code-track score of RHAE ≥ 0.5.
ARC Prize 2026 winners announcement
Resolution Official ARC Prize final leaderboard or winners publication; verified entries only.
-
30%
OpenA production incident (not eval-only behavior) is publicly attributed to reward hacking or specification gaming in a frontier model by 2026-12-31.
End of 2026
Resolution Vendor postmortem or independent evaluator report attributing a deployed-system incident to reward hacking/spec gaming.
-
30%
OpenOne of the 2024–26 frontier-diaspora labs (SSI, Thinking Machines, Cognition, Periodic, Future House, etc.) is acquired by or absorbed into a hyperscaler or major lab by 2026-12-31.
End of 2026
Resolution Official acquisition/absorption announcement of a named diaspora lab by a hyperscaler or frontier lab.
-
25%
OpenA US export-control action forces suspension or restriction of public access to any frontier model again by 2026-12-31.
End of 2026
Resolution Official government directive or vendor statement attributing an access suspension/restriction to export controls, after the June 2026 Fable 5 episode.
-
20%
OpenSafe Superintelligence Inc. makes its first public product or research release by 2026-12-31.
End of 2026
Resolution Official SSI publication, product, or model release; hiring pages and interviews do not count.
-
15%
OpenA frontier vendor offers ≥100M-token context in general availability by 2026-12-31.
End of 2026
Resolution GA product documentation offering a ≥100M-token context window (not research demo or waitlist).
Resolved
-
12%
Resolved · no backfilled · not scoredA frontier system passes a hardened FCS-1 (equivalence) under audited contamination defenses in H1 2026.
by 2026-06-30 · resolved 2026-06-30
Resolution Resolved NO (backfilled): no audited FCS-1 pass existed. NOTE: the original resolution text cited "dry runs" that never took place — retracted 2026-07-01.
Retraction: this entry was authored after its resolution date and cited nonexistent dry runs. Excluded from calibration.
-
70%
Resolved · yes backfilled · not scoredA frontier model ships native computer-use at ≥70% OSWorld-Verified in H1 2026.
by 2026-06-30 · resolved 2026-06-24
Resolution Resolved YES: GPT-5.4 reported OSWorld-Verified 75.0% with native computer-use (2026-06-24).
Vendor-reported; the forecast was about shipping, not durability. [Backfilled: authored after resolution — excluded from calibration.]
-
18%
Resolved · no backfilled · not scoredARC-AGI-3 human–frontier gap closes to within 20 points by end of Q2 2026.
by 2026-06-30 · resolved 2026-06-18
Resolution Resolved NO: the interactive-task gap held well above 20 points despite static-benchmark gains.
Correctly skeptical. [Backfilled: authored after resolution — excluded from calibration.]