28 articles
Fresh test results show that the best known AI programs solve science problems at very different speeds. OpenAI's latest model finished first leaving Gemini besides Claude behind.
The Sansa benchmark reveals major differences in content filtering across leading AI models, with GPT-5.2 showing the highest restriction levels among systems tested.
OpenAI's latest GPT-5.2 xhigh model completely failed a cutting-edge physics reasoning test, scoring zero on the CritPt benchmark designed to evaluate expert-level theoretical physics capabilities.
GPT-5.2 secured 16th position in the Dubesors LLM Benchmark with a 73.5% overall score, landing in the middle tier among over 250 evaluated large language models.
OpenAI's GPT-5.2 Thinking shows major accuracy improvements across 256k token contexts in new MRCRv2 benchmarks, dramatically reducing "context rot" and opening doors for more reliable enterprise AI agents.
OpenAI's latest documentation reveals GPT-5.2 dramatically cuts factual errors compared to earlier versions, marking a major leap in AI reliability.
OpenAI is expected to release its next frontier model, GPT-5.2, with new Image-2 systems likely to follow. Prediction markets show elevated confidence ahead of the anticipated announcement.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy