7 articles
New research on artificial general intelligence suggests optimization pressures can push AI systems to exploit unmeasured constraints. The findings raise real questions about alignment and verification as AI capabilities keep growing.
Fresh benchmark data reveals Claude Opus 4.5 leading the pack in long-duration AI task performance, sustaining successful execution for nearly five hours at 50% success probability.
Hyperbrowser previewed its AI Research Comparison tool by analyzing Cursor against Google's Antigravity IDE, showing how the platform generates instant side-by-side evaluations with scoring dashboards and visual summaries.
Anthropic has sharply reduced the price of Claude Opus 4.5; the cost per token has fallen by about two thirds because the model now runs on more efficient code and stronger infrastructure.
A leaked Epoch AI schedule hints Opus 4.5 could drop today, fueling speculation about whether it can beat Google's Gemini 3 Pro and shift the balance in the AI race.
Nano Banana 2 and Gemini 3.0 are launching next week, with early tests suggesting major capability improvements. The timing puts fresh competitive pressure on OpenAI and Anthropic in the race for AI model leadership.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy