1 article
Mistral's Devstral-Small-2 achieves impressive local inference speeds on Apple M3 Ultra using MLX, with 6-bit quantization proving to be the sweet spot for balanced performance.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy