Major runtime release
vLLM 0.29.0 makes Model Runner V2 the default for all models
The latest vLLM release turns Model Runner V2 into the default everywhere and adds support for Hy4-preview, Qwen3.8-Flash-Next, GraniteSWA/MoeSWA and NemotronH_Omni_Reasoning_V3. It also ships major Kimi K3 and DeepSeek V4 kernel optimizations plus expanded speculative decoding (DSpark, MTP, sparse MLA). For anyone serving open-weight models locally, this is the most consequential runtime update of the day.




