3 articles
A refreshed LiveBench leaderboard now ranks Claude 4.5 Opus as the top-performing AI model with a 76.20 global average. The benchmark was redesigned to better reflect real-world AI behavior and prevent gaming.
China's new open-source coding model IQuest-Coder has outperformed GPT-5.1 and Claude Sonnet 4.5 on several major coding benchmarks, hitting 81.4% on SWE-Bench Verified and 81.1% on LiveCodeBench v6—all while running on just 40 billion parameters.
A new "Peer Arena" benchmark puts AI models head-to-head in debates where they vote on each other's responses. Anthropic's Claude 4.5 took the top spot by rating, but OpenAI's GPT-5.2 showed the highest tendency to vote for itself among all tested systems.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy