1841 articles
Meta researchers introduced a structured reasoning method that significantly improves how large language models verify software code without executing it.
PaddlePaddle's FlashMaskV4 is a new attention masking framework built on FlashAttention-4 architecture. It targets transformer efficiency and flexible masking for large-scale, long-context AI workloads.
DeepSeek's latest V4Lite checkpoint is showing improved benchmark performance, particularly in math and coding, as competition in the AI model landscape continues to intensify.
A new open-source benchmark evaluating real-world AI automation ranks OpenAI's GPT-5.3 Codex first among more than two dozen models, scoring 97.8% across 23 practical OpenClaw tasks.
New estimates suggest OpenAI's GPT-5 generated roughly $6 billion in revenue but likely produced a net loss after operating expenses and revenue-sharing costs were included.
Anthropic partnered with Mozilla to test whether its Claude AI model could detect security vulnerabilities in Firefox. The system identified 22 bugs in two weeks, including 14 classified as high-severity.
OpenSandbox is an open-source execution environment for AI agents, offering secure sandboxes for running code, interacting with interfaces, and training machine learning models.
Alphabet's Google is expanding the financial analysis capabilities of its Gemini AI platform. The system can now generate detailed stock research, interpret valuation metrics, and analyze market sentiment.
Allen AI introduced OLMo Hybrid, a 7B open-weight language model that combines attention and recurrent layers. The hybrid architecture improves efficiency while maintaining strong benchmark performance, with long-context scores jumping from 70.9% to 85.0%.
A new Anthropic report analyzing labor-market data identifies which occupations face the highest exposure to AI automation, comparing theoretical AI capability with real-world usage across job categories.
Citadel Securities suggests generative AI will mirror past technology cycles, scaling rapidly before infrastructure costs - compute, energy, and data centers - put a ceiling on growth.
OpenAI's GPT-5.4 and GPT-5.4 Pro posted strong results on the ARC-AGI-2 benchmark, with the models achieving scores of 74.0% and 83.3% respectively while demonstrating a cost-performance trade-off.
Researchers from Huawei and partner institutes introduced CLI-Gym, a framework designed to automatically generate command-line troubleshooting tasks for AI agents. The system created a dataset of 1,655 tasks and helped the LiberCoder model reach 46.1% on Terminal-Bench.
Pruna AI's P-Video has topped a new Artificial Analysis comparison as the fastest and cheapest AI video model available via public API — generating a 720p five-second clip in around 10 seconds while undercutting rivals on price.
OpenAI has launched GPT-5.4 Thinking and GPT-5.4 Pro across ChatGPT, the API, and Codex, posting standout scores on reasoning, coding, and computer-use benchmarks - and taking direct aim at rivals like Claude Opus 4.6 and Gemini 3.1 Pro.
ModelScope releases Step-3.5 Flash large language model as open source with its SteptronOSS training framework, showing strong performance across coding, reasoning, and agent benchmarks.
Microsoft's rumored Windows 12 could introduce deeper AI integration through Copilot and a redesigned interface. Reports suggest some advanced AI features may be offered through a subscription model.
Anthropic's Claude Opus 4.6 reportedly solved a long-standing mathematical conjecture by legendary computer scientist Donald Knuth. The breakthrough highlights growing interest in advanced AI reasoning systems in scientific research.
New marketing data from 50 websites shows AI platforms like ChatGPT and Gemini generate higher customer lifetime value than traditional search engines, even as Google retains its lead in raw traffic volume.
Noble Machines has unveiled an industrial humanoid robot already working inside a Fortune Global 500 facility, built by a team from SpaceX, NASA, and Apple to handle hazardous and physically intensive tasks.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy