Local inference speedup
Unsloth v0.1.806-beta: 2x faster Qwen3.8-Flash and GLM-5.3-Flash via MTP
Unsloth's new beta runs Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with multi-token prediction enabled by default, bundled with 170+ improvements across training, chat, hardware support, and APIs in its local inference and fine-tuning studio. It is the most directly actionable speedup this week for anyone serving the leading open-weight models on their own rig.
