63 articles
Researchers from Caltech, Stanford, and Carleton College have published one of the first comprehensive surveys on why large language models still break down during basic reasoning tasks, introducing a unified framework that categorizes both reasoning types and their root causes.
Researchers from CAS and Langboat Technology introduced LightRetriever, a new LLM-based retrieval architecture designed to dramatically accelerate query inference. The system delivers up to 1000x faster query encoding while maintaining most of the original model performance.
New marketing data from 50 websites shows AI platforms like ChatGPT and Gemini generate higher customer lifetime value than traditional search engines, even as Google retains its lead in raw traffic volume.
MIT researchers discovered that large language models often degrade in quality when they process their own previous responses, a phenomenon called "context pollution" that could reshape how AI systems are built and deployed.
A novel automated translation framework using test-time compute scaling has demonstrated markedly higher quality output, with LLM judges preferring its translations four times more often than standard approaches.
Alibaba introduced MobilityBench, a large-scale benchmark for testing LLM route-planning agents using 100,000 real Amap mobility queries across 22 countries. The platform includes a deterministic API-replay sandbox to ensure reproducible evaluation.
Google Research reveals that simply repeating prompts can dramatically improve LLM accuracy across Gemini, GPT, Claude, and DeepSeek models - with one benchmark jumping from 21% to 97% - while adding zero latency or output bloat.
Researchers have developed a diffusion-based technique that trains on internal language model activations to stabilize AI outputs while allowing precise behavior steering, marking a breakthrough in controllable AI systems.
LM Studio's latest update brings Anthropic compatibility to local AI models. Developers can now run Claude Code privately using GGUF and MLX models right from their terminals.
Researchers introduce a virtual computer environment enabling AI models to solve unfamiliar tasks using tools and code. Benchmarks show consistent improvements across multiple domains without extra training.
Fennec, a newly unveiled large language model, is making waves with its massive 1 million token context window and pricing that undercuts leading systems by 50%, while claiming superior benchmark performance.
SecureShell has been featured in LangChain's Community Spotlight as a zero-trust security layer that validates and blocks risky shell commands before LLM agents can execute them.
OpenMed has kicked off an 8-day open-source dataset release series, dropping 225K medical reasoning samples to help train next-gen healthcare AI models.
A fresh early-access chapter on LLM self-refinement just dropped, pushing inference-time scaling techniques way beyond basic self-consistency and voting. The update brings an iterative critique-and-refine loop with fully functional code implementations.
Researchers have introduced AP2O-Coder, a new training framework that reduces compilation and runtime errors in AI-generated code by analyzing and correcting specific failure types through targeted learning.
A new research method replaces autoregressive speculative decoding with discrete diffusion to accelerate large language model inference. Tests show substantial speedups while preserving identical output quality.
A new academic study reveals how AI agents can be secretly tricked into burning through massive computing resources while still giving you the right answers. It exposes a critical blind spot in how these systems interact with external tools.
MARSHAL, a new framework from Tsinghua University and partners, improves LLM performance by 28.7% through self-play. It strengthens multi-agent reasoning and boosts results on key benchmarks.
Token costs for top-tier large language models are dropping 900x per year, while mid-tier models see 40x reductions and basic models fall 9x annually, making AI accessible at unprecedented scale.
Why upstream data quality is becoming the backbone of scalable machine learning and LLM systems—and what production-grade pipelines actually look like.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy