1659 articles
Benchmark charts suggest OpenAI's GPT-5.4 is among the more knowledgeable models tested — yet it also records one of the highest hallucination rates. The contrast has sparked debate about how well AI benchmark scores reflect real-world reliability.
Researchers from Microsoft and The Chinese University of Hong Kong have introduced EmotionThinker, a speech AI system that explains how it detects emotions in audio. The model uses reinforcement learning and prosody analysis to produce interpretable reasoning.
Google released its open-source Agent Development Kit built for Gemini 3.1 Flash-Lite, enabling developers to build AI agents with persistent memory that runs 24/7 in the background.
OpenSandbox is an open-source execution environment for AI agents, offering secure sandboxes for running code, interacting with interfaces, and training machine learning models.
Allen AI introduced OLMo Hybrid, a 7B open-weight language model that combines attention and recurrent layers. The hybrid architecture improves efficiency while maintaining strong benchmark performance, with long-context scores jumping from 70.9% to 85.0%.
A new Anthropic report analyzing labor-market data identifies which occupations face the highest exposure to AI automation, comparing theoretical AI capability with real-world usage across job categories.
Citadel Securities suggests generative AI will mirror past technology cycles, scaling rapidly before infrastructure costs - compute, energy, and data centers - put a ceiling on growth.
Researchers introduced VDR-Bench, a benchmark designed to test whether AI systems genuinely analyze images or rely on textual shortcuts. The framework aims to improve multimodal reasoning by forcing models to search images iteratively for hidden visual clues.
A new academic study examining "shadow APIs" finds that third-party services claiming to offer access to frontier AI models may deliver significantly weaker performance than official APIs, with accuracy gaps reaching more than 47%.
GPT-5.4 has crossed the human baseline in computer-use tasks, posting a 75.0% success rate across browser interaction and OSWorld-Verified evaluations.
Researchers from Huawei and partner institutes introduced CLI-Gym, a framework designed to automatically generate command-line troubleshooting tasks for AI agents. The system created a dataset of 1,655 tasks and helped the LiberCoder model reach 46.1% on Terminal-Bench.
xAI's Grok Imagine Video has claimed the top spot on the Arena Image-to-Video leaderboard, outperforming multiple versions of Google's Veo 3.1 based on more than 49,000 user votes.
Researchers from CAS and Langboat Technology introduced LightRetriever, a new LLM-based retrieval architecture designed to dramatically accelerate query inference. The system delivers up to 1000x faster query encoding while maintaining most of the original model performance.
Aheadform has introduced the F1 half-humanoid robot designed for companionship and social interaction. The system can dynamically switch between teacher, learning assistant, companion, and therapist roles depending on context.
Yuan Lab introduced Yuan3.0 Ultra, a trillion-parameter multimodal MoE model designed for retrieval and reasoning tasks. The model achieved leading scores across several RAG and multimodal benchmarks.
Meta FAIR and NYU researchers introduced new findings on native multimodal pretraining. The research explores how unified architectures can combine visual understanding, generation, and language reasoning in a single model.
ModelScope releases Step-3.5 Flash large language model as open source with its SteptronOSS training framework, showing strong performance across coding, reasoning, and agent benchmarks.
Anthropic's Claude Opus 4.6 reportedly solved a long-standing mathematical conjecture by legendary computer scientist Donald Knuth. The breakthrough highlights growing interest in advanced AI reasoning systems in scientific research.
Microsoft Research Asia introduced UniG2U-Bench, a benchmark designed to test whether generation improves understanding in unified multimodal AI models. The framework evaluates systems across seven reasoning regimes and 30 subtasks.
Noble Machines has unveiled an industrial humanoid robot already working inside a Fortune Global 500 facility, built by a team from SpaceX, NASA, and Apple to handle hazardous and physically intensive tasks.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy