METR · Time Horizon 1.1
Which software tasks can it complete?
METR gives AI agents programming, machine learning and cybersecurity tasks. Difficulty is expressed as the time a human expert would need: the hours in the chart are not the AI’s running time.
Tasks the agent is predicted to complete 5 times out of 10. Longer bars mean harder tasks.
* Estimates above 16 hours are unreliable; bars stop at 16 h. Selection of 10 models, ordered by release date. Same TH 1.1 benchmark.
Uncertainty and exact values
Estimates have uncertainty: the ranges below are the intervals supplied by METR. Model and agent tools are evaluated together.
GPT-4 · 2023-03-14
2 min – 8 min
Claude 3.5 Sonnet (Jun 2024) · 2024-06-20
5 min – 22 min
Claude 3.7 Sonnet · 2025-02-24
33 min – 1 h 44 min
o3 · 2025-04-16
1 h 15 min – 3 h 11 min
Claude Opus 4.5 · 2025-11-24
2 h 42 min – 10 h 24 min
Claude Opus 4.6 · 2026-02-05
5 h 17 min – 60 h 34 min
GPT-5.3 Codex · 2026-02-05
3 h 15 min – 13 h 36 min
Gemini 3.1 Pro · 2026-02-19
3 h 54 min – 11 h 35 min
GPT-5.4 · 2026-03-05
3 h 7 min – 12 h 49 min
Claude Mythos Preview (early) · 2026-04-07
8 h 29 min – 55 h 4 min
METR warns that measurements above 16 hours are unreliable with its current suite. These percentages do not measure proximity to AGI.
METR source and method ↗Page updated 8 May 2026 · accessed 4 Sep 2026

