12 articles
New evaluations show GLM-5.1 closing in on Claude Opus 4.6 in coding performance, highlighting a shrinking gap between open and proprietary AI models.
OpenAI's GPT-5.4 Design Skill shows limited gains in DesignArena rankings. Claude Opus 4.6 remains the top-performing model with a significant lead.
Claude Opus 4.6 tops the MRCR v2 long-context benchmark with a 78.3% match ratio at 1M tokens, a massive leap from the previous 18.5%.
Anthropic's Claude Opus 4.6 has secured the top position in Search Arena rankings, scoring 1255 points and outperforming major competitors including Grok and GPT models.
Claude Opus 4.6 and other advanced AI models are hitting task completion milestones faster than anyone expected. New METR data shows capability gains accelerating well beyond earlier projections, with doubling time shrinking from seven months to roughly four.
MiniMax M2.5 has overtaken Claude Opus 4.6 in total tokens generated this week, hitting 3.07 trillion according to the latest leaderboard data. The shift puts real-world usage front and center over benchmark scores.
New benchmark data shows Claude Opus 4.6 degrades more slowly over long-duration tasks compared to rival models, signaling a shift toward valuing sustained performance over raw speed.
New evaluation data shows Claude Opus 4.6 achieving a roughly 14.5-hour task completion horizon at a 50% success rate. The finding highlights accelerated AI capability growth in complex software tasks.
New SWE-ReBench evaluation reveals Claude Opus 4.6 topping the leaderboard at 51.7%, while Qwen3-Coder-Next delivers competitive 40% performance despite smaller parameter count.
Perplexity's Deep Research feature now runs on Anthropic's Opus 4.6 model, achieving 81.9% benchmark performance and rolling out to Max and Pro subscribers.
A developer built a complete video editing app using Claude Agent SDK and Opus 4.6, generating approximately 10,000 lines of code that runs entirely on local machines.
Anthropic just dropped Claude Opus 4.6 with massive improvements in reasoning, coding, and information retrieval. The model's dominating long-context benchmarks with a 1M token window.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy