1841 articles
Meta researchers introduced a structured reasoning method that forces AI systems to trace code step by step, significantly improving accuracy when verifying software updates without executing the code.
Anthropic rolled out a free interactive prompt engineering course with hands-on Claude notebooks that's already racked up more than 31,900 stars on GitHub, showing massive developer interest.
Cognition's early SWE-1.6 language model preview shows improved reasoning performance on the SWE-Bench Pro coding benchmark, achieving a 51.7% score while maintaining fast inference speed.
Google DeepMind has launched Gemini 3.1 Flash-Lite, claiming faster throughput and lower pricing than Gemini 2.5 Flash. The model adds adjustable reasoning levels designed to handle complex tasks more cost-efficiently.
MIT researchers discovered that large language models often degrade in quality when they process their own previous responses, a phenomenon called "context pollution" that could reshape how AI systems are built and deployed.
xAI has launched Grok 4.20 Beta 2 with improved instruction following, reduced capability hallucinations, better LaTeX support, and stronger multimodal performance - continuing Grok's fast-paced development cycle.
New web traffic data shows Claude.ai accelerating sharply in the second half of February 2026, closing in on competitors Grok and DeepSeek. The trend highlights shifting user engagement among leading generative AI platforms.
Screenshots from public repository edits suggest GPT-5.4 could introduce a 2 million token context window, full-resolution vision, and faster response tiers. Community prediction markets already assign high odds to an imminent release.
Xiaomi's humanoid robots completed a continuous 3-hour assembly trial in its EV plant, reaching 90.2% installation accuracy and signaling the company's push toward full factory automation within five years.
Anthropic's Claude AI services experienced multiple partial outages this month, raising questions about reliability as the company scales rapidly. While uptime metrics remain above 98%, clustered disruptions suggest potential infrastructure strain.
OpenAI's latest GPT-5.3 Codex model has claimed the top spot on the WeirdML benchmark with a 79.3% score, edging out Claude Opus 4.6 while keeping costs surprisingly low at just $2.35 per run.
Alibaba's Qwen 3.5 Small Model Series arrives with benchmark results that punch well above its weight class. The 9B variant scored 90.0 in math and reached the mid-80s to low-90s across reasoning tests, making a strong case for compact, efficient AI deployment.
A coalition of 56 researchers just released VBVR, the world's largest video reasoning benchmark. Top AI systems hit only 54% accuracy - while humans cleared 97%.
A novel automated translation framework using test-time compute scaling has demonstrated markedly higher quality output, with LLM judges preferring its translations four times more often than standard approaches.
Similarweb data shows AI tools are becoming the dominant starting point for consumer discovery, while search engines regain relevance in comparison and purchase stages. The shift marks changing patterns in online decision journeys.
Grok 4.20 Beta (500B) claimed the #1 ranking on Search Arena with Style Control enabled and #2 without it. The 500-billion-parameter model reportedly outperformed several trillion-parameter competitors.
MBZUAI's MediX-R1 framework advances medical AI with reinforcement learning, achieving 73.6% on clinical benchmarks from ~51K examples. The development highlights increasing AI capability across modalities including imaging and free-form answer generation.
AI development is shifting from static RAG (Retrieval-Augmented Generation) workflows toward dynamic memory-enabled agents with read-write capabilities. Persistent memory systems now allow continual learning, personalization, and context retention across interactions.
Anthropic's valuation has exploded from $4 billion to $380 billion, prompting questions about how public market participants can gain exposure. A breakdown highlights key tech companies positioned around Anthropic's growth.
Nvidia (NVDA) has teamed up with major global telecom firms to develop AI-native 6G wireless networks - open, secure, and intelligent infrastructure built to handle tomorrow's connectivity demands well beyond what 5G can offer today.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy