1841 articles
GLM-4.7 has emerged as the top-performing open-source AI model by surpassing GPT-5.1 in Vending-Bench2 testing, demonstrating superior capabilities in extended task management and consistent performance over a 350-day simulation period.
Tencent's new Vision-Language Model introduces 4D reasoning that beats existing video understanding systems by over 20%, achieving 58.9% accuracy while maintaining strong general video comprehension.
NVIDIA unveils Nemotron 3 Nano, a 30-billion parameter AI model that activates only 3 billion parameters per pass, achieving 3.3x higher throughput than competitors while handling 1-million token contexts for advanced agentic reasoning tasks.
A new "Peer Arena" benchmark puts AI models head-to-head in debates where they vote on each other's responses. Anthropic's Claude 4.5 took the top spot by rating, but OpenAI's GPT-5.2 showed the highest tendency to vote for itself among all tested systems.
Microsoft CEO Satya Nadella has personally stepped in to tackle Copilot's reliability problems, joining a 100-engineer Teams channel and running weekly fix-it sessions to address user complaints across Microsoft 365.
Grok Code just claimed the top spot on the Kilo Agentic Model Leaderboard after logging 35.6 billion tokens in real developer usage—more than quadruple its nearest competitor.
Leading AI companies are paying interns engineer-level salaries and providing dedicated compute budgets as competition for early-career AI talent intensifies.
Tesla's ramping up its humanoid robot program in a big way—110 new job openings just went live for Optimus development, while Elon Musk eyes early 2026 for a production-ready prototype and aims to kick off mass manufacturing by year's end.
Latest data reveals ChatGPT's massive advantage in daily active users over Google Gemini, though market dynamics show shifting competitive landscape in AI platforms.
A breakthrough reinforcement learning technique called GTR-Turbo is transforming how vision-language models are trained. By creating self-contained teacher models from merged checkpoints, the method achieves double-digit accuracy gains while cutting training time in half and reducing compute expenses by 60%.
AgiBOT's Genie 2 humanoid robot just pulled off something impressive at PIA Smart's factory: assembling flexible components in 15 seconds flat with a 99% success rate — beating human workers by three seconds.
StepFun has released Step-DeepResearch, a deep-research AI agent that scores 61.4% on the Scale AI Research Rubrics benchmark with a 32B model, positioning it near OpenAI and Alphabet's Gemini DeepResearch while emphasizing lower cost.
2025 marked a turning point for AI as systems from OpenAI and Alphabet (GOOGL) reached gold-medal performance in global math and coding competitions, while agent-based tools moved from labs into everyday use.
MiniMax's language model achieves sub-second response times running on local GPU hardware, with dashboard data showing 918ms latency and 97.5 tokens per second throughput.
The AI landscape in 2025 saw major breakthroughs from DeepSeek R1, Claude 4, Grok 4, GPT-5 and Opus 4.1, with the industry rapidly moving toward "agent-first" platforms designed for real-world action.
Cyber attackers are changing tactics, embedding malicious activity inside trusted software, AI chatbots, Docker environments and Android devices as traditional standalone malware gives way to infrastructure-based attacks.
A critical vulnerability in LangChain Core allows attackers to extract secrets and manipulate AI output through prompt injection, earning CVE-2025-68664 a severe 9.3 CVSS rating.
One developer reported consuming half a billion AI tokens over three weeks while running extended workflows through Roocode, now expanding to connect with Anthropic's Claude Max and OpenAI's GPT-5.2 for even larger-scale automation.
MiniMax's latest M2.1 model shows significant improvements in coding performance while facing setbacks in math and spatial reasoning, revealing the trade-offs in rapid AI development.
Samsung's upcoming Exynos 2600 processor debuts AMD's RDNA4 architecture in mobile chips, featuring the Xclipse 960 GPU that promises double the computing power and 50% better ray tracing than its predecessor.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy