Arbitrary Code Injection Affecting tritonserver-backend-vllm-cuda-13.0 package, versions <25.11-r12


Severity

Recommended
low

Based on default assessment until relevant scores are available.

Threat Intelligence

EPSS
0.91% (58th percentile)

Do your applications use this vulnerable package?

In a few clicks we can analyze your entire application and see what components are vulnerable in your application, and suggest you quick fixes.

Test your applications
  • Snyk IDSNYK-CHAINGUARDLATEST-TRITONSERVERBACKENDVLLMCUDA130-19235327
  • published25 Aug 2026
  • disclosed22 Jun 2026

Introduced: 22 Jun 2026

CVE-2026-41523  (opens in a new tab)
CWE-94  (opens in a new tab)
CWE-617  (opens in a new tab)

How to fix?

Upgrade Chainguard tritonserver-backend-vllm-cuda-13.0 to version 25.11-r12 or higher.

NVD Description

Note: Versions mentioned in the description apply only to the upstream tritonserver-backend-vllm-cuda-13.0 package and not the tritonserver-backend-vllm-cuda-13.0 package as distributed by Chainguard. See How to fix? for Chainguard relevant fixed versions and status.

vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.0, an assert-based security check in vLLM's activation function loading allows any unauthenticated attacker to achieve arbitrary code execution on the server by publishing a malicious HuggingFace model, when vLLM runs in Python optimized mode (python -O or PYTHONOPTIMIZE=1). This vulnerability is fixed in 0.22.0.