6 articles
METR's latest evaluation puts OpenAI's GPT-5.3-Codex at a 50% task success horizon of around 6.5 hours on software assignments, with a 95% confidence interval spanning 3 to 17 hours.
GPT-5.2 has successfully operated without interruption for seven consecutive days, building excitement around upcoming METR performance benchmarks that could showcase dramatic improvements in AI task efficiency.
METR's latest research reveals AI model time horizons are accelerating beyond exponential growth patterns, signaling rapid advances in how long systems can maintain coherent reasoning across complex tasks.
New METR benchmark data exposes dramatic differences in how fast leading AI models complete complex tasks, with some systems taking 2.6 times longer than others despite similar capabilities.
METR evaluation data shows Claude Opus 4.5 achieving a long median task horizon, while maintaining a much shorter high-reliability threshold.
A METR benchmark chart reveals Claude Opus 4.5 achieving over 4-hour task completion times, significantly outperforming OpenAI's latest models on complex software engineering challenges.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy