16 articles
New evaluations show GLM-5.1 closing in on Claude Opus 4.6 in coding performance, highlighting a shrinking gap between open and proprietary AI models.
A new benchmark, MME-Emotion, evaluates emotional intelligence in AI using thousands of video samples. Results show current models still struggle with emotion recognition and reasoning.
Google's Gemini 3.1 Flash Live Preview ranks second in a key speech reasoning benchmark. The model introduces configurable thinking levels balancing performance and latency.
xAI's Grok Imagine video model ranked first in Video Arena, Image-to-Video Arena, and Video Editing Arena on DesignArena leaderboards.
A new framework tests AI embedding models across four memory types, revealing that bigger models don't always win on long-horizon tasks.
A new MADQA study shows Gemini 3 Pro matching human accuracy at 82.2%, but humans and AI still solve different problems in completely different ways.
Innovator-VL hits 85.6 on AI2D and 65.1 on MolParse using under 5M training samples - a lean, reproducible approach to multimodal AI.
GPT-5.4 has crossed the human baseline in computer-use tasks, posting a 75.0% success rate across browser interaction and OSWorld-Verified evaluations.
Cognition's early SWE-1.6 language model preview shows improved reasoning performance on the SWE-Bench Pro coding benchmark, achieving a 51.7% score while maintaining fast inference speed.
A coalition of 56 researchers just released VBVR, the world's largest video reasoning benchmark. Top AI systems hit only 54% accuracy - while humans cleared 97%.
The Reagent framework was introduced to improve AI agent reasoning by providing detailed critiques of each step - delivering measurable performance gains across multiple reasoning benchmarks, including 43.7% on GAIA.
MMDeepResearch-Bench from OSU, Amazon, and UMich tests how well AI research agents build citation-rich reports. Early results from 25 models reveal serious gaps in connecting visual evidence to written claims.
A new "Bullshit Benchmark" tests how AI models handle illogical questions, revealing that many still generate confident answers to meaningless prompts instead of rejecting them outright.
ByteDance and research partners launched NL2Repo-Bench to test if AI can autonomously create complete software repositories. Top models are struggling, with pass rates stuck below 40%.
A GPT-5-powered agent paired with Opus 4.5 scored 72.6% on the OSWorld benchmark, showcasing significant advances in AI systems that can handle real-world computer tasks autonomously.
Fresh benchmark data reveals GPT-5.1 cuts intelligence costs by 300× compared to o3-preview while maintaining strong performance at 72.8% on ARC-AGI-1. Sam Altman admits he underestimated how fast AI intelligence would become affordable.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy