Incorrect Calculation of Buffer Size Affecting vllm-openai-cuda-12.9 package, versions <0.28.0-r0


Severity

Recommended
low

Based on default assessment until relevant scores are available.

Threat Intelligence

EPSS
0.43% (35th percentile)

Do your applications use this vulnerable package?

In a few clicks we can analyze your entire application and see what components are vulnerable in your application, and suggest you quick fixes.

Test your applications
  • Snyk IDSNYK-CHAINGUARDLATEST-VLLMOPENAICUDA129-19538070
  • published5 Sept 2026
  • disclosed12 May 2026

Introduced: 12 May 2026

CVE-2026-44223  (opens in a new tab)
CWE-131  (opens in a new tab)
CWE-704  (opens in a new tab)

How to fix?

Upgrade Chainguard vllm-openai-cuda-12.9 to version 0.28.0-r0 or higher.

NVD Description

Note: Versions mentioned in the description apply only to the upstream vllm-openai-cuda-12.9 package and not the vllm-openai-cuda-12.9 package as distributed by Chainguard. See How to fix? for Chainguard relevant fixed versions and status.

vLLM is an inference and serving engine for large language models (LLMs). From to before 0.20.0, the extract_hidden_states speculative decoding proposer in vLLM returns a tensor with an incorrect shape after the first decode step, causing a RuntimeError that crashes the EngineCore process. The crash is triggered when any request in the batch uses sampling penalty parameters (repetition_penalty, frequency_penalty, or presence_penalty). A single request with a penalty parameter (e.g., "repetition_penalty": 1.1) is sufficient to crash the server. This vulnerability is fixed in 0.20.0.