1 article
Anthropic's Opus 4.5 model achieved 95 % on CORE-Bench Hard. This test is hard. It checks if an AI can run a full real world scientific study without help.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy