7 articles
Grok 4.20 records the lowest hallucination rate at 22%, beating Claude 4.5 Haiku, MiniMax V2 Pro, and GLM-5 in factual accuracy.
China's AI models are closing the gap with US leaders and already surpass them in open-source performance - a shift that could reshape global AI adoption.
MiniMax M2.7 debuts with self-evolving training, 30% gains, and benchmark scores rivaling Claude and GPT-5
A new multimodal benchmark reveals that most leading AI systems struggle with context-driven reasoning, relying on memorization rather than adaptive thinking.
Grok 4.20 Beta (500B) claimed the #1 ranking on Search Arena with Style Control enabled and #2 without it. The 500-billion-parameter model reportedly outperformed several trillion-parameter competitors.
MMDeepResearch-Bench from OSU, Amazon, and UMich tests how well AI research agents build citation-rich reports. Early results from 25 models reveal serious gaps in connecting visual evidence to written claims.
AI coding assistants are everywhere now, but how well do they actually handle real software engineering work? A new research framework called SwingArena is putting that question to the test by throwing AI models into actual developer workflows - complete with bug fixes, code reviews, and continuous integration pipelines.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy