1659 articles
A viral comparison shows Grok delivering a direct "No" answer to a controversial sports question, while ChatGPT, Gemini, and Claude provided nuanced responses instead of a simple yes-or-no reply.
A breakthrough training method called Co-rewarding helps AI language models think better and learn more reliably without needing human-verified answers, showing real improvements on mathematical problem-solving tests.
A new AI communication framework called Vision Wormhole proposes using visual channels as a universal messaging layer between models - aiming to improve bandwidth and interoperability compared with traditional token exchanges.
Google Trends data from late February shows a sharp but short-lived surge in Anthropic searches following a U.S. agency usage ban report, while OpenAI held steady. Both brands ended the week with an equal average score of 26 - revealing how social media noise can distort real market signals.
Grok's weekend traffic decline was the lowest among major generative AI tools, according to traffic data shared online. Competing platforms like ChatGPT and DeepSeek saw steeper drops.
A new prompting technique is reshaping how developers optimize reasoning in GPT-4, Claude, and Gemini - without retraining the models.
DeepSeek is preparing to launch its V4 multimodal model in mid-February, optimized for Chinese AI chips rather than Nvidia hardware. The move highlights shifting infrastructure strategies in the global AI landscape.
Alibaba Group and Beijing University of Posts and Telecommunications have introduced SpatialGenEval, a benchmark for spatial reasoning in text-to-image models. Evaluation of 23 leading systems reveals persistent object placement gaps, though targeted fine-tuning shows measurable gains.
DEXFORCE Robotics has deployed its DexForce W1 Pro (Gen2) humanoid robot inside a Shenzhen community convenience store, where it autonomously prepares customized meals, heats them, and delivers them to customers without human help.
AITECH data shows that 79% of enterprises are now deploying AI agents operationally, not just experimentally. The trend signals a structural shift in how organizations implement autonomous AI.
New data reveals China's electricity generation surged 74% from 2014-2024, while U.S. power grew just 6%. The widening gap raises questions about energy infrastructure's role in AI development and compute capacity expansion.
Grok Imagine has entered the Top 3 of the Video Editing Arena leaderboard, posting the highest overall win rate of 66% and reaching 1253 Elo. The 480p model now trails the top Elo position by just three points.
Anthropic's Claude Opus 4.6 has secured the top position in Search Arena rankings, scoring 1255 points and outperforming major competitors including Grok and GPT models.
Epoch AI research reveals AI training efficiency improved through compute scaling, reaching 21,400× total gains with estimates of nearly 10× annual improvements driven by scale-dependent algorithmic innovations.
A new independent benchmark reveals CodeRabbit leading AI code review tools with 51.30% F1 score, while Gemini places third at 49.70%. The evaluation tracks both controlled tests and real-world bug fixes in open-source projects.
Grok reached 385.7 million combined desktop and mobile web visits in January 2026. Most traffic now comes directly rather than through embedded platform access.
AI training compute efficiency has been improving several times per year, according to Epoch AI analysis. Multiple studies confirm substantial gains - though uncertainty ranges remain wide across different methodologies.
Anthropic's Claude has more than doubled its weekly traffic in just six weeks, reaching nearly 79 million visits by mid-February 2025. While ChatGPT and Gemini still dominate in scale, Claude is outpacing both on growth rate by a wide margin.
NVIDIA has unveiled EgoScale, a foundation model that trains humanoid robots to perform real-world tasks using as few as 100 demonstrations - slashing the data requirements that have long made teleoperation impractical at scale.
GPT-5.3-Codex achieved 86% accuracy on IBench benchmark testing, significantly ahead of Gemini 3.1 Pro's 69% score. The performance gap highlights growing differences between top-tier language models in coding and reasoning tasks.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy