Independent benchmark evidence
The frontier moved. Here is the evidence.
GPT-5.6 Sol (pro, max) currently leads the composite dataset at 162.1 ECI. Individual benchmarks show uneven progress: hard reasoning, autonomous task duration, and evaluation durability remain distinct constraints rather than one finish line.
- Current composite frontier
- 162.1 ECI
- Observed leader
- GPT-5.6 Sol (pro, max)
- Evidence basis
- 6 public measures
Epoch Capabilities Index
Frontier index
A stitched general-capability scale built from more than 50 benchmarks.
Useful for comparing the frontier over long periods, but not a probability or countdown to AGI.
Inspect source dataOptional interpretation
Ask GPT-OSS to explain this signal
The model receives the selected public record, not an open web prompt. It cannot alter the score.
01 / Public record
Signals shaping the frontier
Each row is the latest observed leader for one measure. Expand it to see the prior-frontier comparison and source.
GPT-5.6 Sol (pro, max) defines the observed frontier index frontier 162.1 on Epoch Capabilities Index. Useful for comparing the frontier over long periods, but not a probability or countdown to AGI. +1.1 ECI High confidence
GPT-5.6 Sol (pro, max), reported by OpenAI, improved the observed frontier by +1.1 ECI relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.
GPT-5.6 Sol (Max) defines the observed generalization frontier 92.5% on ARC-AGI-2. Strong performance can reveal flexible reasoning, while benchmark saturation can weaken the signal. +2.5 pp Medium confidence
GPT-5.6 Sol (Max), reported by OpenAI, improved the observed frontier by +2.5 pp relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.
GPT-5.5 Pro pre-release (high) defines the observed reasoning frontier 52.4% on FrontierMath. Higher scores show stronger hard-problem solving; they do not establish broad reliability. +0.7 pp Medium confidence
GPT-5.5 Pro pre-release (high), reported by OpenAI, improved the observed frontier by +0.7 pp relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.
Claude Opus 4.7 (max) defines the observed software work frontier 83.5% on SWE-bench Verified. Measures bounded repository tasks, not unattended ownership of production systems. +4.8 pp Medium confidence
Claude Opus 4.7 (max), reported by Anthropic, improved the observed frontier by +4.8 pp relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.
Gemini 3.1 Pro Preview defines the observed autonomy frontier 1.5 hr on METR task horizon (80%). Longer horizons matter, but a benchmark task is not the same as safe, persistent agency. +20 min Medium confidence
Gemini 3.1 Pro Preview, reported by Google DeepMind, improved the observed frontier by +20 min relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.
video-SALMONN 2+ defines the observed multimodal frontier 79.7% on Video-MME. Captures one form of visual-language understanding, not embodied world modelling. +4.7 pp Medium confidence
video-SALMONN 2+, reported by ByteDance, improved the observed frontier by +4.7 pp relative to the previous dated leader in this dataset. Scores can change as benchmark maintainers add or revise evaluations.
03 / Method
Evidence-weighted, not a countdown.
Road to AGI displays the published Epoch Capabilities Index alongside individual benchmark frontiers. It does not convert those measurements into an AGI probability, arrival date, or hidden proprietary score.
What the index does
- 01 Use dated benchmark records and preserve the original unit.
- 02 Show the frontier and the previous frontier so movement is inspectable.
- 03 Keep capability evidence separate from unresolved reliability questions.
- 04 Link every displayed signal to the underlying public source.
Update cycle
The public archive is checked every day at 04:17 UTC. Automation rebuilds the committed evidence snapshot, and a connected Cloudflare Pages project deploys the refreshed result from the production branch.
Where the model stops
Groq-hosted GPT-OSS is optional and only explains already-selected evidence. It does not choose sources, alter scores, or determine the frontier.