What vllm-project/vllm shipped
Written by FoxPlug from public releases; not affiliated with vLLM. An automatic summary of the public release, pull request and commit data of github.com/vllm-project/vllm. vLLM did not write it and does not use or endorse FoxPlug. Every line links to the public change it describes.
Get a weekly update like this for your product, free
Week of September 21, 2026
What shipped
- Fused small-batch DSv4.1 O-projection kernels on SM100/SM103 to reduce four separate kernel calls to two. Pull request #58634
- Bumped FlashKDA to keep recurrent state in fp32, fixing accumulated rounding errors on long prefills. Pull request #58846
- Added support for running DCP target models with non-DCP DSpark drafts. Pull request #56723
- Fused DSV4.1 MoE finalize operation into TP all-reduce and mHC boundary. Pull request #58586
- Indexed expert mapping lookups in RoutedExperts.load_weights to fix quadratic scaling with expert count. Pull request #58720
- Gated per-request multimodal processor kwargs to prevent untrusted overrides from exhausting memory. Pull request #58830
- Isolated registry tests that need fresh process to avoid 15s interpreter/torch/vllm import overhead on AMD CI. Pull request #58316
- Added sharding-aware NCCL M2N weight-transfer backend for distributed training on GPUs with different weight layouts. Pull request #51520
Changelog entry
- Frontend: switched to oss-harmony for Python dependency to eliminate runtime tiktoken encoding downloads Pull request #55128
- KV Connector: retry Mooncake bootstrap registration on timeout with internal 3-attempt limit Pull request #58919
- ROCm: fall back to default GEMM for CPU tensors in quantization dispatch Pull request #58923
- MoE: enabled fused MiniMax2 routing with non-unit routed scaling factors Pull request #58880
- Weight cache daemon: added /health endpoint reporting starting/ready/failed states Pull request #58552
- DSv4.1: fused WO-A with inverse RoPE and MXFP8 quantization on SM100/SM103 Pull request #58634
- FlashKDA: updated to keep recurrent state in fp32 for improved long-context accuracy Pull request #58846
- MoE: indexed expert mapping lookups to reduce weight loading from quadratic to linear complexity Pull request #58720
- Frontend: rejected untrusted per-request multimodal processor kwargs to prevent memory exhaustion Pull request #58830
- CI: isolated registry tests needing fresh process to avoid 15s import overhead on AMD systems Pull request #58316
- RL: added sharding-aware NCCL M2N weight-transfer backend for distributed training Pull request #51520
vLLM updates: MoE routing fusion, DSv4.1 kernel optimization, FlashKDA fp32 state preservation, expert mapping indexing, sharding-aware weight transfer, and multimodal security hardening.
vLLM's latest updates focus on performance and reliability: fused MoE routing for non-unit scaling factors, optimized DSv4.1 O-projection kernels, fixed accumulated rounding errors in long prefills, indexed expert mapping to fix quadratic scaling, enabled mixed DCP/non-DCP pipelines, and hardened multimodal processor security. CI improvements reduce overhead on AMD systems.