21 articles
Multiple flagship AI models are rolling out across leading platforms in the coming weeks. The packed release schedule signals how development timelines are tightening across the industry.
China's new open-source coding model IQuest-Coder has outperformed GPT-5.1 and Claude Sonnet 4.5 on several major coding benchmarks, hitting 81.4% on SWE-Bench Verified and 81.1% on LiveCodeBench v6—all while running on just 40 billion parameters.
GLM-4.7 has emerged as the top-performing open-source AI model by surpassing GPT-5.1 in Vending-Bench2 testing, demonstrating superior capabilities in extended task management and consistent performance over a 350-day simulation period.
Fresh benchmark data reveals Claude Opus 4.5 holds up better than GPT-5.1-Codex-Max when tasks stretch into multi-hour territory, with the performance gap widening as duration increases.
Google's Gemini 3 Flash now matches the scores of GPT-5.2 and Claude 4.5 Sonnet on Vending-Bench 2 plus it clearly surpasses the results of all earlier model versions.
Microsoft is rolling out GPT-5.1 for Smart Mode and building a new Reminders feature. The move puts the company in the race for stickier AI assistants as platforms bet on simple productivity tools that keep users coming back.
OpenAI launched ChatGPT for Teachers, a secure GPT-5.1 workspace with unlimited messages, file uploads, search tools, and connectors for verified U.S. educators—free through June 2027 as AI adoption accelerates in K–12 classrooms.
OpenAI just dropped GPT-5.1 Instant and GPT-5.1 Thinking, giving users the ability to tweak tone and communication style. The update tackles earlier complaints about responses feeling too cookie-cutter and aims to make conversations flow more naturally.
OpenAI rolled out GPT-5.1-Codex-Max, a cutting-edge agentic coding model built for extended software development workflows. The release shows substantial accuracy improvements and better efficiency than previous versions.
OpenAI's latest model sets a new benchmark for sustained autonomous operation, reaching the longest task duration recorded by METR while staying within established safety boundaries.
OpenAI has unveiled GPT-5.1-Codex-Max, a significant upgrade to its Codex coding model that's built to tackle larger, long-running development projects with improved speed and intelligence compared to earlier versions.
Fresh benchmark data reveals GPT-5.1 cuts intelligence costs by 300× compared to o3-preview while maintaining strong performance at 72.8% on ARC-AGI-1. Sam Altman admits he underestimated how fast AI intelligence would become affordable.
OpenAI's GPT-5.1 brings adaptive reasoning to the table, letting the model decide how much thinking time each task deserves. It speeds through straightforward questions while dedicating serious processing power to harder challenges.
A developer test revealed GPT-5.1 needed almost two minutes to read a 500-line markdown file in one setup, while processing the same file in seconds using a different configuration. The findings have triggered debate about the model's performance reliability.
OpenAI's GPT-5.1 release is facing backlash for delivering only incremental improvements, with critics suggesting the company is holding back stronger internal models to appeal to mainstream users rather than maximizing performance.
OpenAI has released GPT-5.1 to the API, delivering adaptive reasoning, major speed improvements, and new developer tools. The upgrade reduces token usage, enhances coding performance, and expands prompt-caching capabilities.
GPT-5.1 now follows instructions more accurately and respects custom formatting preferences, including the option to avoid em-dashes when specified.
A new AI model believed to be GPT-5.1 takes the lead in creative writing as crypto mining faces regulatory challenges.
Fresh code snippets from OpenAI's Enterprise RBAC system hint that three GPT-5.1 models are on the way. Developer screenshots show internal descriptions for GPT-5.1, GPT-5.1 Reasoning, and GPT-5.1 Pro.
A leaked code snippet referencing "GPT-5.1 Thinking" has ignited speculation about OpenAI's next-generation reasoning model, though the company has yet to confirm any such development.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy