5 articles
A new open-source benchmark evaluating real-world AI automation ranks OpenAI's GPT-5.3 Codex first among more than two dozen models, scoring 97.8% across 23 practical OpenClaw tasks.
OpenAI's latest GPT-5.3 Codex model has claimed the top spot on the WeirdML benchmark with a 79.3% score, edging out Claude Opus 4.6 while keeping costs surprisingly low at just $2.35 per run.
GPT-5.3-Codex achieved 86% accuracy on IBench benchmark testing, significantly ahead of Gemini 3.1 Pro's 69% score. The performance gap highlights growing differences between top-tier language models in coding and reasoning tasks.
Lovable has integrated GPT-5.3-Codex for its most complex tasks, reporting the model is significantly stronger than GPT-5.2 and 3-4 times more token-efficient - a shift that could reshape how AI tools handle demanding technical workloads.
METR's latest evaluation puts OpenAI's GPT-5.3-Codex at a 50% task success horizon of around 6.5 hours on software assignments, with a 95% confidence interval spanning 3 to 17 hours.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy