54 articles
Stanford-Princeton MedOS beats GPT-5, Gemini 3 Pro on clinical benchmarks and deploys into real hospital workflows with XR glasses and collaborative robotics.
GPT-5.4 scores 77.7% on WeirdML, placing just behind Claude Opus 4.6. Results reveal how cost, token use, and code efficiency now define AI model rankings.
New estimates suggest OpenAI's GPT-5 generated roughly $6 billion in revenue but likely produced a net loss after operating expenses and revenue-sharing costs were included.
A new academic study examining "shadow APIs" finds that third-party services claiming to offer access to frontier AI models may deliver significantly weaker performance than official APIs, with accuracy gaps reaching more than 47%.
Researchers from Fudan, University of Washington, and NUS introduced AdaReasoner, a reinforcement-learning driven visual reasoning model showing 24.9% improvement over baseline with strong performance on visual tool-using benchmarks.
ByteDance and research partners launched NL2Repo-Bench to test if AI can autonomously create complete software repositories. Top models are struggling, with pass rates stuck below 40%.
New financial analysis reveals GPT-5 achieved 45% gross margins on $6.1 billion in revenue but recorded a $1.9 billion net loss after operating costs. The numbers expose the massive spending required to scale cutting-edge AI models.
A new multi-model architecture outperformed GPT-5 on advanced reasoning tasks while running 2.5x faster and cutting costs by 70%, proving that smart coordination beats raw model size.
QwenLong-L1.5 can reason across documents containing up to 4 million tokens—equivalent to about 100 novels—matching top long-context AI models like GPT-5 on benchmark evaluations.
A fresh research paper unveils DiffThinker, a diffusion-powered AI system that tackles visual reasoning challenges by creating images rather than text. Researchers say it crushed GPT-5 and Gemini-3-Flash across seven benchmark tests.
The AI landscape in 2025 saw major breakthroughs from DeepSeek R1, Claude 4, Grok 4, GPT-5 and Opus 4.1, with the industry rapidly moving toward "agent-first" platforms designed for real-world action.
OpenAI just released a GPT-5 Prompting Guide with 9 structured rules to help users craft better, clearer prompts. With AI adoption heating up and names like NVDA staying front and center, the guide's designed to make working with advanced AI smoother and more predictable.
For the first time ever, an AI has cracked an open math problem completely on its own, delivering a verified proof without any human help or hints.
Fresh research unveils PaCoRe, a parallel reasoning system that smartly scales computing power during testing. Using this approach, an 8-billion parameter model outscored GPT-5 on a challenging math benchmark.
New FrontierMath data from EpochAI reveals Chinese open-weight models are roughly seven months behind frontier AI systems in mathematical reasoning tasks, with the performance gap widening significantly on the most difficult problems.
A viral radar chart comparing GPT-5 and GPT-4 reveals massive jumps in language, math, and reasoning abilities, fueling claims that AI has already blown past human experts in key intellectual areas.
OpenAI's GPT-5 achieved a 79-fold improvement in molecular cloning efficiency during real-world wet-lab experiments with Red Queen Bio.
Fresh SimpleBench results reveal GPT-5.2 trailing far behind competitors in common-sense reasoning tests, with Gemini 3 Pro Preview taking the crown.
Google's Gemini 3 Pro just claimed the top spot on the newly launched FACTS Benchmark Suite, beating out every major AI model when it comes to accuracy across grounding, search, knowledge recall, and multimodal reasoning.
DeepSeek's new V3.2 and V3.2-Speciale models deliver breakthrough performance in reasoning and coding benchmarks, with optimized vLLM support now available for production deployment.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy