vLLM LLM inference server releases
new versionvLLM LLM inference server is tracked under self-hosted apps & servers. Its latest version, 0.29.0, shipped on Sep 9, 2026.
- Latest version
- 0.29.0
- Released
- Category
- Self-hosted apps & servers
- Last checked
How to update vLLM LLM inference server
- Read the release notes first — self-hosted projects break config and schemas between versions more often than appliances do.
- Docker: pull the new image tag and recreate the container (`docker compose pull && docker compose up -d`).
- Package or bare metal: update through your distribution's package manager, or replace the binary from the releases page.
- Back up the config and database directory first. For most projects a downgrade is not supported once the schema has migrated.
Always download firmware from the manufacturer or project itself. Firmwarely links out; it never hosts firmware.
What changed
v0.29.0 Highlights This release features 594 commits from 277 contributors (91 new)! Model Runner V2 is now the default for all models (#53183), completing the rollout that began with pooling models (#48290). MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-s…
Version history
| 0.29.0 |
FAQ
What is the latest version of vLLM LLM inference server?
Version 0.29.0, released Sep 9, 2026. This page is refreshed every 30 minutes from the manufacturer's release page.
How do I update vLLM LLM inference server?
See the step-by-step above. Pull the new image or package and restart the service. Read the release notes and back up the data directory first; most projects can't downgrade once the database has migrated.
How does Firmwarely know when there's a new version?
Every 30 minutes we read the manufacturer's official release page for the vLLM LLM inference server and record the version, date and changelog. Subscribers watching this device get one email a day when it changed.
Is this an official vLLM page?
No. Firmwarely is independent. Project and product names belong to their owners; always install releases from the project's own source.
Get alerts for vLLM LLM inference server
One email when a new or critical firmware ships. Free for up to 3 devices.