1841 articles
A viral comparison shows Grok delivering a direct "No" answer to a controversial sports question, while ChatGPT, Gemini, and Claude provided nuanced responses instead of a simple yes-or-no reply.
A breakthrough training method called Co-rewarding helps AI language models think better and learn more reliably without needing human-verified answers, showing real improvements on mathematical problem-solving tests.
A new AI communication framework called Vision Wormhole proposes using visual channels as a universal messaging layer between models - aiming to improve bandwidth and interoperability compared with traditional token exchanges.
Google Trends data from late February shows a sharp but short-lived surge in Anthropic searches following a U.S. agency usage ban report, while OpenAI held steady. Both brands ended the week with an equal average score of 26 - revealing how social media noise can distort real market signals.
A new prompting technique is reshaping how developers optimize reasoning in GPT-4, Claude, and Gemini - without retraining the models.
Perplexity AI launched "Perplexity Computer," a workflow workspace that combines multi-model orchestration with app connectors, exclusive to its Max subscription. Early adopters say the right setup unlocks a genuine autonomous assistant.
Zhipu AI and Tsinghua University released GLM-5, a 744 billion parameter open-weights model that hits state-of-the-art marks in coding, reasoning, and autonomous agentic tasks — placing it squarely alongside leading proprietary systems.
DeepSeek is preparing to launch its V4 multimodal model in mid-February, optimized for Chinese AI chips rather than Nvidia hardware. The move highlights shifting infrastructure strategies in the global AI landscape.
AITECH data shows that 79% of enterprises are now deploying AI agents operationally, not just experimentally. The trend signals a structural shift in how organizations implement autonomous AI.
OpenAI has raised $110 billion at a $730 billion pre-money valuation, with major backing from SoftBank, NVIDIA (NVDA), and Amazon (AMZN). The company reports 900 million weekly ChatGPT users and an expanded strategic compute partnership.
Perplexity launched two new embedding model families - pplx-embed-v1 and pplx-embed-context-v1 - claiming strong benchmark performance and significantly higher storage efficiency compared to competitors like Gemini. The models are fully open-source and optimized for web-scale retrieval.
Microsoft (MSFT) previewed Copilot Tasks, a new AI agent that interprets plain-language requests and executes them on its own in the cloud. The feature is currently available to a limited group of early testers.
Alibaba introduced MobilityBench, a large-scale benchmark for testing LLM route-planning agents using 100,000 real Amap mobility queries across 22 countries. The platform includes a deterministic API-replay sandbox to ensure reproducible evaluation.
Anthropic's Claude Opus 4.6 has secured the top position in Search Arena rankings, scoring 1255 points and outperforming major competitors including Grok and GPT models.
Epoch AI research reveals AI training efficiency improved through compute scaling, reaching 21,400× total gains with estimates of nearly 10× annual improvements driven by scale-dependent algorithmic innovations.
A new independent benchmark reveals CodeRabbit leading AI code review tools with 51.30% F1 score, while Gemini places third at 49.70%. The evaluation tracks both controlled tests and real-world bug fixes in open-source projects.
Grok reached 385.7 million combined desktop and mobile web visits in January 2026. Most traffic now comes directly rather than through embedded platform access.
AI training compute efficiency has been improving several times per year, according to Epoch AI analysis. Multiple studies confirm substantial gains - though uncertainty ranges remain wide across different methodologies.
Anthropic's Claude has more than doubled its weekly traffic in just six weeks, reaching nearly 79 million visits by mid-February 2025. While ChatGPT and Gemini still dominate in scale, Claude is outpacing both on growth rate by a wide margin.
Google's Nano Banana 2 (Gemini 3.1 Flash Image Preview) has reached the top of the Artificial Analysis Text-to-Image leaderboard, delivering leading performance at roughly half the price of competing models.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy