28 articles
A new AI system, SWE-Vision, is outperforming leading models like GPT-5.2 and Seed-2.0-Pro across multiple visual reasoning benchmarks, highlighting advances in code-driven image analysis.
AI trained on 700K citation pairs predicts high-impact research and beats GPT-5.2 in scientific forecasting.
Yuan Lab introduced Yuan3.0 Ultra, a trillion-parameter multimodal MoE model designed for retrieval and reasoning tasks. The model achieved leading scores across several RAG and multimodal benchmarks.
OpenAI's latest GPT-5.2-chat model has broken into the Arena.ai Text Arena top five rankings, scoring 1478 points and marking a significant 40-point jump over its predecessor.
Zhipu's GLM-5 just landed third place on the DesignArena leaderboard, pulling in a 1353 Elo score across 888K user votes — beating out GPT-5.2 and Gemini 3 Pro while keeping costs well below most frontier rivals.
OpenAI's GPT-5.2 has uncovered a previously unknown gluon interaction that physicists later confirmed, marking a significant breakthrough in theoretical physics research.
Grok delivered 95% accuracy on the challenging AIME 2026 math benchmark while costing just $0.06 per inference—nearly matching GPT-5.2's performance at a fraction of the price.
OpenAI's GPT-5.2 just crushed the FrontierMath benchmark, landing near the top across multiple math-focused tests. The results show serious progress in advanced reasoning capabilities.
GPT-5.2 has successfully operated without interruption for seven consecutive days, building excitement around upcoming METR performance benchmarks that could showcase dramatic improvements in AI task efficiency.
GPT-5.2 smashed through the 33% accuracy barrier on LiveCodeBench Pro (Hard) way earlier than anyone expected, leaving 2030 predictions in the dust and sparking fresh debates about where AI is really headed.
AI has hit a major milestone by solving a longstanding Erdős problem in number theory, showing that machine intelligence can now tackle complex mathematical proofs independently.
The newest Artificial Analysis Intelligence Index reveals performance rankings across 10 major AI benchmarks, with Gemini 3 Pro Preview (High) and GPT-5.2 Thinking X-High tied at the top spot.
New session data reveals GPT-5.2-Codex responds significantly faster in medium reasoning mode, with median latency of just 4.9 seconds compared to 29.7 seconds in extra-high mode.
A new "Peer Arena" benchmark puts AI models head-to-head in debates where they vote on each other's responses. Anthropic's Claude 4.5 took the top spot by rating, but OpenAI's GPT-5.2 showed the highest tendency to vote for itself among all tested systems.
One developer reported consuming half a billion AI tokens over three weeks while running extended workflows through Roocode, now expanding to connect with Anthropic's Claude Max and OpenAI's GPT-5.2 for even larger-scale automation.
OpenAI's GPT-5.2 reaches its highest benchmark position yet in no-reasoning tasks, ranking 14th, while GPT-5.1-high maintains the top spot at 8th place among reasoning-focused models in the latest LMArena data.
A new benchmark result puts GPT-5.2 in the spotlight, while version tracking reveals how far apart major AI platforms have drifted in their development cycles.
Gemini 3 Flash, Gemini 3 Pro, and GPT-5.2 now cluster at the top of major AI benchmarks, with several tests hitting 90%+ scores that signal performance saturation.
OpenAI's latest benchmark shows GPT-5.2 Thinking now matches or outperforms human experts across most real-world professional work, completing tasks 10x faster at 1% of the cost.
Google's Gemini 3 Flash now matches the scores of GPT-5.2 and Claude 4.5 Sonnet on Vending-Bench 2 plus it clearly surpasses the results of all earlier model versions.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy