1659 articles
A new AI framework called CineTrans dramatically improves scene transition quality in AI-generated videos by training on 250,000 film clips and applying a novel masking technique. The result is output that finally starts to feel like something a human editor might actually approve.
Grok-4.20 Beta 1 (500B) has claimed the top spot in Search Arena rankings while securing #4 in Text Arena. The model currently outperforms several major competitors in the latest benchmark results.
Researchers from Fudan, University of Washington, and NUS introduced AdaReasoner, a reinforcement-learning driven visual reasoning model showing 24.9% improvement over baseline with strong performance on visual tool-using benchmarks.
Cursor's autonomous cloud agents now generate nearly 1 in 3 merged pull requests internally, running full dev environments end-to-end — from building features to submitting production-ready PRs without human involvement.
MiniMax released the M2.5 model as fully open source, claiming performance close to Claude Opus at a fraction of the cost. With an 80.2% SWE-bench Verified score and speeds three times faster than leading AI systems, it's a serious new contender in the open-source AI space.
Google's Gemini 3.1 Pro teams up with NotebookLM to create research-heavy workflows built on verified sources, shifting focus from pure benchmark numbers to practical context handling.
MMDeepResearch-Bench from OSU, Amazon, and UMich tests how well AI research agents build citation-rich reports. Early results from 25 models reveal serious gaps in connecting visual evidence to written claims.
A new "Bullshit Benchmark" tests how AI models handle illogical questions, revealing that many still generate confident answers to meaningless prompts instead of rejecting them outright.
Tencent-Hunyuan rolled out GradLoc, a framework that hunts down and fixes token-level training crashes in large reasoning models. The system uses binary search diagnostics paired with adaptive layerwise clipping to keep training stable.
Anthropic has accused DeepSeek, Moonshot AI, and MiniMax of creating over 24,000 fraudulent accounts and running more than 16 million interactions to extract capabilities from Claude, framing it as both an IP violation and a national security concern.
China's enterprise AI consumption skyrocketed to 37 trillion daily tokens in H2 2025 - a 263% jump from the first half. Open-source models like Qwen, Doubao, and DeepSeek now dominate the market, flipping the script on closed-source systems.
New employment data reveal a surprising trend: the jobs most exposed to AI language models have actually grown faster since ChatGPT launched in 2022, while least-exposed roles have lagged behind.
Researchers at Communication University of China have introduced LaGoVAD, an AI model that uses everyday language to detect security anomalies in video footage, achieving breakthrough zero-shot performance across seven industry benchmarks.
OpenAI's latest GPT-5.2-chat model has broken into the Arena.ai Text Arena top five rankings, scoring 1478 points and marking a significant 40-point jump over its predecessor.
Researchers from Tsinghua University and Zhejiang Lab launched WorldArena, a groundbreaking benchmark that tests whether AI world models can actually help robots perform real tasks - not just generate pretty videos.
JPMorgan Asset Management highlights that advanced AI agents now handle much longer tasks with improved reliability, but many models still have high rates of incorrect answers. The research raises trust concerns as AI usage grows across industries.
NanoClaw has launched as a lightweight, container-isolated personal Claude assistant with agent swarm support and external I/O capabilities, targeting developers who want secure, customizable AI workflows.
Chinese robotics developer X-Humanoid unveiled Embodied TienKung 3.0, an open hardware and software humanoid robot platform designed to cut redundant engineering and speed up real-world deployment.
ByteDance and research partners launched NL2Repo-Bench to test if AI can autonomously create complete software repositories. Top models are struggling, with pass rates stuck below 40%.
New Similarweb audience data reveals ChatGPT users are the most loyal among leading Gen AI platforms, while Claude and Grok attract more exploratory, multi-platform visitors.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy