2 articles
GLM-4.7 just claimed the top spot among open-weights models on the GDPval-AA benchmark with a 1224 ELO score, while MiniMax's M2.1 version also showed solid gains. The benchmark tests how well AI models handle actual work tasks instead of artificial tests.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy