63 articles
Researchers introduced OPUS, a dynamic data selection method for LLM pre-training that delivers up to 8x computation reduction while boosting benchmark accuracy by 2.2% on average.
Huawei and Shanghai Jiao Tong University introduced HyperOffload, a compiler-assisted framework that reduces LLM peak device memory usage by up to 26%. The solution rethinks how data moves across AI hardware - maintaining performance while solving one of the field's core infrastructure challenges.
A new study shows AI agents can autonomously create advanced jailbreak attacks, outperforming over 30 human-designed methods - and hitting up to 100% success rates against specific models.
A new 3B-parameter model from SII-GAIR shows performance comparable to 7B models, highlighting growing efficiency in AI development.
ByteDance Seed researchers introduced behavioral calibration, a reinforcement learning method that reduces hallucinations in LLMs. The approach enables models to recognize uncertainty and avoid incorrect answers
A new study introduces T-MAP, a trajectory-aware method exposing critical security flaws in LLM agents operating through multi-step tool interactions in real-world AI environments.
A new speculative perception framework for agentic multimodal AI delivers speed gains of up to 3.35x while pushing benchmark accuracy from 81% to 84%.
A new AI framework boosts LLM accuracy by 70% and cuts token usage by 39% through smarter reasoning allocation.
New research shows AI agent groups fail to reach agreement even in simple tasks, challenging multi-agent coordination assumptions.
ReBalance cuts LLM token use by 47% with a training-free method that dynamically balances reasoning depth using real-time confidence signals.
Aligned LLMs diverge from real human decisions. Base models outperform aligned versions 10:1 in strategic scenarios.
Moonshot AI's Attention Residuals replace fixed layer accumulation with dynamic selection, improving LLM reasoning and memory efficiency.
A new study of 500 keywords finds that Google's top-ranked brands dominate AI-generated answers, revealing a tight link between search visibility and AI mentions.
New MIT CSAIL research explores how large language models build reasoning skills - and whether massive human-text datasets are really necessary.
A lightweight token-embedding trick gives large language models a serious capacity upgrade without piling on the compute costs.
Meta and Yale researchers show reasoning-based judges cut reward hacking in RL alignment pipelines, but challenges remain in non-verifiable training settings.
TAPPA framework reveals predictable attention patterns in LLMs, enabling faster AI via smarter KV cache compression.
A new synthetic dataset called Chimera shows that compact language models can rival far larger systems in complex reasoning tasks. Training on just 9,225 curated problems allowed a 4B model to approach the performance of models up to 50 times its size.
Researchers introduced LightMem, a memory architecture designed to make large language models faster and more efficient. Early results show improved accuracy while dramatically reducing token usage and API calls.
Meta researchers introduced a structured reasoning method that significantly improves how large language models verify software code without executing it.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy