Independent project. Not a U.S. government website.

USASI

Release

vLLM 0.31.0 released with a preload command for faster engine restarts

Event: · Published:

On October 5, 2026 the vLLM project published version 0.31.0 on GitHub. The release notes highlight a new `vllm preload` command that launches a weight-cache daemon to keep post-quantized weights in GPU memory across engine restarts, and experimental engine snapshots (`vllm snapshot create/restore`) that use CRIU to restore a fully initialized TP1 engine. Under security, the notes say per-request `mm_processor_kwargs` and `media_io_kwargs` are now rejected unless `--trust-request-mm-kwargs` is set. The release also lists breaking changes, including the removal of `tokenizer_mode="slow"`, and performance work for DeepSeek-V4.1-Flash. USASI has not tested any of these changes.1

Sources

  1. 1.
    vLLM v0.31.0 (GitHub release notes) (external site: github.com)

    vLLM project · Release notes · published Oct 5, 2026 · accessed Oct 6, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project