Allocation of Resources Without Limits or Throttling Affecting vllm package, versions [,0.26.0)


Severity

Recommended
0.0
medium
0
10

CVSS assessment by Snyk's Security Team. Learn more

Threat Intelligence

Exploit Maturity
Proof of Concept
EPSS
0.34% (28th percentile)

Do your applications use this vulnerable package?

In a few clicks we can analyze your entire application and see what components are vulnerable in your application, and suggest you quick fixes.

Test your applications

Snyk Learn

Learn about Allocation of Resources Without Limits or Throttling vulnerabilities in an interactive lesson.

Start learning
  • Snyk IDSNYK-PYTHON-VLLM-18912241
  • published18 Aug 2026
  • disclosed18 Aug 2026
  • creditRex Liu

Introduced: 18 Aug 2026

NewCVE-2026-71486  (opens in a new tab)
CWE-770  (opens in a new tab)

How to fix?

Upgrade vllm to version 0.26.0 or higher.

Overview

vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

Affected versions of this package are vulnerable to Allocation of Resources Without Limits or Throttling via the derender endpoints (derender_chat_response and derender_completion_response in vllm/entrypoints/scale_out/derender/serving.py) due to missing resource bounds validation on caller-supplied token structures. An authenticated remote attacker can send requests containing arbitrarily large generate_responses payloads — with oversized token_ids, logprobs.content, top_logprobs, or prompt_logprobs arrays — causing unbounded tokenizer.decode() and parser invocations that exhaust CPU and memory resources on the server.

CVSS Base Scores

version 4.0
version 3.1