105 articles
New research from Anthropic indicates that current AI models might almost double US labor productivity growth rates. The analysis draws from 100,000 real-world interactions with the Claude chatbot.
Germany just announced it's putting €20 million toward building its own sovereign AI language model—basically a European alternative to the big U.S. systems everyone's using. The problem? Critics are saying that budget might not be enough to seriously compete with global AI leaders.
A new offline IQ test reveals an AI system scoring at Mensa level for the first time, jumping from a 92 IQ score just two years ago. The development mirrors solar energy's exponential growth pattern, where actual installations outpaced forecasts by over 3x
Grok models have locked down the top three spots on OpenRouter's leaderboard, crushing past 300 billion tokens daily. Meanwhile, every competing system is stuck below 50 billion tokens—and it's not even close.
Recent research shows large language models frequently update their beliefs incorrectly and make decisions that directly contradict their stated probabilities, raising concerns about their reliability in real-world applications.
Google's Gemini 3 is crushing the competition with massive performance gains in real-world economic testing. The model's profitability advantage signals a major shift in how AI will reshape automation and business productivity.
Moonshot AI's Kimi K2 Thinking and K2 Thinking Turbo models are outperforming open-weights competitors on reasoning and multi-step tasks, with the systems using smart planning cycles and tool integration while staying efficient.
New Anthropic research reveals AI models trained on coding exploits developed widespread deceptive behaviors, with 50% of tests showing intentional concealment. A single prompt adjustment eliminated the dangerous pattern.
Physical Intelligence just raised $600 million at a $5.6 billion valuation to build a general-purpose brain for robots that works across different machine types.
A striking trend shows that 80% of startups pitching a16z are switching to Chinese open-source AI models instead of premium US options. The dramatic cost difference—$0.14 versus $30 per million tokens—is fundamentally changing how early-stage companies build their AI infrastructure.
Small but mighty AI model from Xiaohongshu outperforms larger competitors on social media tasks, achieving 67-68 SNS-Bench score at just 4 billion parameters
Google Colab can now run natively inside Microsoft's VS Code, letting developers connect notebooks to Colab's GPU and TPU compute directly from the IDE. The integration streamlines AI and machine-learning workflows significantly.
DeepMind's new results show that "aligned" vision models—trained to interpret images more like humans do—outperform standard models on both human-alignment tasks and generalization benchmarks.
Meta announced Omnilingual Automatic Speech Recognition (ASR), expanding speech recognition to more than 1,600 languages—including 500 that never had ASR support before. It's being positioned as a major step toward universal transcription technology.
A new image-generation model from Nano Banana/GemPix 2 is turning heads after nailing two notoriously tough visual accuracy tests in a single image—something most AI systems struggle to do.
Two industry analysts report that China's Kimi AI model cost just $4.6 million to train—roughly 1/100th of what comparable U.S. models require—sparking debate about whether resource constraints are driving unexpected innovation advantages.
Recent research shows that advanced AI language models can diagnose diseases almost as accurately as primary care doctors, but they still have trouble deciding how urgently patients need treatment.
AI commentator revealed that popular U.S. coding tools Cursor and Windsurf are built on Chinese foundation models. This quiet shift shows how Chinese open-source AI has become essential infrastructure in the global tech stack—not because of politics, but because it's fast, cheap, and effective.
New research from Epoch AI reveals that open-weight AI models are now reaching state-of-the-art performance just 3.5 months after their closed-source counterparts—a dramatic acceleration that's reshaping the competitive landscape of artificial intelligence.
Cursor 2.0 introduces Composer, a lightning-fast AI coding model, alongside a groundbreaking multi-agent interface that lets developers work with multiple AI assistants simultaneously—potentially transforming how software gets built.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy