1841 articles
Z.AI's GLM-5 demonstrates impressive coding capabilities, surpassing GPT-5.1-Codex and Sonnet 4.5 in specialized developer benchmarks.
Mistral has launched the open-weight Ministral 3 family in three sizes: 14B, 8B, and 3B parameters. These models leverage cascade distillation technology to compete with larger alternatives while demanding significantly less computational power.
Unitree showcased humanoid robots performing actual assembly work in its factory, demonstrating major advances in robotic dexterity and close-range vision for manufacturing applications.
A small team's TinyFish web agent topped a major web agent benchmark, outscoring Google, OpenAI, and Anthropic systems by over 20 percentage points in real-world web task performance.
Independent SWE-rebench testing exposes major performance gaps between AI coding models, with Chinese releases underperforming earlier claims while frontier models cluster around 50%.
Nvidia's GB300 GPUs deliver breakthrough performance running DeepSeek models, with benchmarks showing up to 20x throughput improvements over previous-generation hardware in AI inference workloads.
Anthropic landed $30 billion in fresh capital and announced $14 billion in annual run-rate revenue, with explosive growth in its Claude Code enterprise product.
Google's Gemini Deep Think system now verifies and repairs mathematical proofs iteratively, achieving approximately 90% accuracy on the IMO-ProofBench Advanced benchmark through its Aletheia agent.
MiniMax introduced its open-source M2.5 model with high benchmark scores across coding, search, and tool-calling tasks.
An innovative open-source humanoid robot prototype proves that advanced home robotics doesn't need a luxury price tag—just $10,000 in parts and some engineering know-how.
A recent comparison revealed xAI's Grok delivering instant responses to breaking information while competing AI systems lag behind with outdated data, highlighting significant speed differences in real-time AI capabilities.
Researchers released a lightweight open-source AI agent framework that compresses complex agent architecture into a compact Python system.
A real-world coding benchmark revealed large performance differences between major AI coding models. GLM-5 lagged competitors in task completion and response speed despite strong theoretical rankings.
Google's Gemini 3 Deep Think achieved 84.6% on the ARC-AGI-2 reasoning benchmark, marking a significant leap in AI capabilities and intensifying the race among tech giants.
Moonshot AI just dropped a model that can juggle multiple AI agents at the same time. The system handles coding, research, browsing, and fact-checking all in parallel.
Elon Musk predicts AI will create executable binaries by the end of 2026, completely skipping traditional coding and compilers.
Grok has rocketed from an unranked early-stage launch to become the world's third most popular generative AI tool in just one year.
Ant Open Source launched LLaDA2.1 Flash, a language diffusion model that generates tokens in parallel and delivers peak speeds of 892 tokens per second.
The OpenClaw project hit 145,000 GitHub stars as a new setup guide explains safe deployment and architecture.
A new research paper introduces a structured world model for AI agents. The concept focuses on prediction and planning rather than scaling model size.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy