54 articles
A researcher reports that GPT-5 proposed the central idea developed in a new Quantum Field Theory paper published in Physics Letters B. Screenshots show GPT-5 outlining a key conceptual step using the Tomonaga–Schwinger formalism.
DeepSeek v3.2 appears with a sparse attention system that cuts computational costs and still matches the performance of top AI models. This marks a clear move toward greater efficiency in the design of large language models.
DeepSeek just dropped two new AI models, with V3.2 Speciale beating GPT-5 High on key benchmarks while staying completely open-source. The release also shows strong performance against Gemini 3.0 Pro at way lower costs.
Sam Altman revealed a major shift in AI development, with models advancing from seconds-long tasks to 5-hour workflows in GPT-5. The next goal is enabling AI systems to operate across months or even years using integrated enterprise context.
OpenAI's upcoming GPT-5 model is demonstrating capabilities far beyond prior AI systems, including generating new mathematics, identifying obscure research, and proposing publishable scientific experiments. Early observations suggest a major shift in how scientific discovery may scale.
New benchmark data from DesignArena.ai shows OpenAI's GPT-5 and GPT-5.1 outperforming top Claude models across multiple ELO categories, with GPT-5.1 claiming the platform's #1 overall ranking.
OpenAI's GPT-5.1 release is facing backlash for delivering only incremental improvements, with critics suggesting the company is holding back stronger internal models to appeal to mainstream users rather than maximizing performance.
Cursor has released two new GPT-5 Codex variants — FAST and High FAST — that code twice as fast but cost double, targeting developers who value speed over savings.
Kimi-K2 Thinking scored 42.1% on the WeirdML benchmark, making it China's top open-source model and outperforming both Claude 4.1 Opus and Grok-4
A new AI model delivers GPT-5-level performance while slashing costs by up to 95% and running 150% faster, potentially reshaping the economics of enterprise AI deployment.
AI researchers are highlighting growing differences between U.S. and Chinese AGI development. OpenAI's GPT-5 Pro is matching top benchmark results while analysts map out potential AGI timelines around 2032.
Sam Altman claims GPT-5 is showing the first hints of generating genuinely new scientific ideas, while GPT-6 could deliver a breakthrough comparable to the leap from GPT-3 to GPT-4. If true, AI may be shifting from a research tool to an active participant in scientific discovery.
Google DeepMind's Gemini 3.0 might drop between November 12-18, right when several Gemini 2.x models officially retire on November 18.
A new Nature study reveals that GPT-5, despite sounding more confident and coherent than ever, still gets more than half of difficult medical cases wrong. The research highlights a troubling gap between how well AI can talk and how well it can actually think—raising serious safety concerns for healthcare applications.
A practical test shows how lightweight Haiku 4.5 beats GPT-5 in document parsing accuracy, challenging assumptions about AI capability and model size.
Verdent hit 76.1% on single-try and 81.2% on three-try runs in SWE-Bench Verified, beating GPT-5 and Claude Code. The results show Verdent leads in practical, reliable AI coding.
Epoch AI's new Capabilities Index shows that more training compute directly translates to better model performance, with GPT-5 and Grok 4 leading the pack
GPT-5 Pro independently suggested dupilumab—a drug already approved for eczema and asthma—as a potential treatment for FPIES, a rare food allergy. This matched real clinical findings from Dr. Oral Alpan.
OpenAI has rolled out Aardvark, a GPT-5-powered security agent that automatically hunts down and fixes code vulnerabilities. Currently in private beta, it's a big leap forward in AI-powered cybersecurity.
An academic paper references GPT-5o, showing accuracy between 77% and 89% on software reasoning tests, with one benchmark completed in just 138 seconds.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy