1841 articles
DeepSeek's new "DualPath" paper challenges the standard prefill-centric approach to KV-cache loading, reporting up to 1.87x higher offline throughput and up to 1.96x more agent runs per second in online serving, driven by smarter scheduling and infrastructure design.
Nvidia shattered expectations with record quarterly revenue of $68.13 billion and adjusted EPS of $1.62 - beating Wall Street forecasts. Data center and networking segments led the charge, highlighting surging AI infrastructure demand.
A 27-billion-parameter model is closing the gap on the Artificial Intelligence Index, challenging larger and more established AI systems across 10 benchmarks.
NVIDIA has unveiled EgoScale, a foundation model that trains humanoid robots to perform real-world tasks using as few as 100 demonstrations - slashing the data requirements that have long made teleoperation impractical at scale.
A new AI framework called CineTrans dramatically improves scene transition quality in AI-generated videos by training on 250,000 film clips and applying a novel masking technique. The result is output that finally starts to feel like something a human editor might actually approve.
Grok-4.20 Beta 1 (500B) has claimed the top spot in Search Arena rankings while securing #4 in Text Arena. The model currently outperforms several major competitors in the latest benchmark results.
New benchmark data show GLM-5 holding its own against top AI models in some evaluations while trailing in others - a reminder that no single score tells the whole story.
Google's Gemini 3.1 Pro is reportedly experiencing a sharp throughput decline on Google Vertex, falling to 50 TPS from 92 TPS at launch. Users are reporting slower response times and inconsistent performance across deployments.
Wells Fargo's latest infrastructure outlook shows U.S. data center construction starts surging 67% year over year, with power project activity up 68% and North American utility capex forecasts revised upward to roughly $222 billion in 2026.
New research on artificial general intelligence suggests optimization pressures can push AI systems to exploit unmeasured constraints. The findings raise real questions about alignment and verification as AI capabilities keep growing.
Cursor's autonomous cloud agents now generate nearly 1 in 3 merged pull requests internally, running full dev environments end-to-end — from building features to submitting production-ready PRs without human involvement.
MiniMax released the M2.5 model as fully open source, claiming performance close to Claude Opus at a fraction of the cost. With an 80.2% SWE-bench Verified score and speeds three times faster than leading AI systems, it's a serious new contender in the open-source AI space.
A new "Bullshit Benchmark" tests how AI models handle illogical questions, revealing that many still generate confident answers to meaningless prompts instead of rejecting them outright.
Tencent-Hunyuan rolled out GradLoc, a framework that hunts down and fixes token-level training crashes in large reasoning models. The system uses binary search diagnostics paired with adaptive layerwise clipping to keep training stable.
Lovable has integrated GPT-5.3-Codex for its most complex tasks, reporting the model is significantly stronger than GPT-5.2 and 3-4 times more token-efficient - a shift that could reshape how AI tools handle demanding technical workloads.
A new NBER paper reports generative AI significantly reduces the productivity gap between workers with high and low education. An experiment involving 1,174 adults shows AI closes roughly three quarters of the initial performance difference in an incentivized task.
Claude Opus 4.6 and other advanced AI models are hitting task completion milestones faster than anyone expected. New METR data shows capability gains accelerating well beyond earlier projections, with doubling time shrinking from seven months to roughly four.
Alibaba's Qwen 3.5 Medium Model Series delivers benchmark scores that rival much larger AI systems, proving you don't always need the biggest model to get top-tier results.
Anthropic has accused DeepSeek, Moonshot AI, and MiniMax of creating over 24,000 fraudulent accounts and running more than 16 million interactions to extract capabilities from Claude, framing it as both an IP violation and a national security concern.
China's enterprise AI consumption skyrocketed to 37 trillion daily tokens in H2 2025 - a 263% jump from the first half. Open-source models like Qwen, Doubao, and DeepSeek now dominate the market, flipping the script on closed-source systems.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy