7 articles
OpenAI's GPT-5.4 Design Skill shows limited gains in DesignArena rankings. Claude Opus 4.6 remains the top-performing model with a significant lead.
GPT-5.4 mini scores 72.1% on OSWorld-Verified, beating Claude Haiku 4.5's 50.7% in efficiency benchmarks.
Cursor launched CursorBench to score AI models on agentic coding tasks, measuring both intelligence and token efficiency across 10+ leading models.
Benchmark charts suggest OpenAI's GPT-5.4 is among the more knowledgeable models tested — yet it also records one of the highest hallucination rates. The contrast has sparked debate about how well AI benchmark scores reflect real-world reliability.
GPT-5.4 has crossed the human baseline in computer-use tasks, posting a 75.0% success rate across browser interaction and OSWorld-Verified evaluations.
OpenAI has launched GPT-5.4 Thinking and GPT-5.4 Pro across ChatGPT, the API, and Codex, posting standout scores on reasoning, coding, and computer-use benchmarks - and taking direct aim at rivals like Claude Opus 4.6 and Gemini 3.1 Pro.
Screenshots from public repository edits suggest GPT-5.4 could introduce a 2 million token context window, full-resolution vision, and faster response tiers. Community prediction markets already assign high odds to an imminent release.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy