vllm@0.31.0

A high-throughput and memory-efficient inference and serving engine for LLMs

  • latest version

    0.31.0

  • first published

    3 years ago

  • latest version published

    2 hours ago

  • licenses detected

  • Direct Vulnerabilities

    Known vulnerabilities in the vllm package. This does not include vulnerabilities belonging to this package’s dependencies.

    Fix vulnerabilities automatically

    Snyk's AI Trust Platform automatically finds the best upgrade path and integrates with your development workflows. Secure your code at zero cost.

    Fix for free
    VulnerabilityVulnerable Version
    • H
    Improper Handling of Exceptional Conditions

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Improper Handling of Exceptional Conditions in mooncake_connector.py, the Mooncake KV connector's receiver loop fails to report remote KV load failures back to the scheduler. When a remote KV cache transfer returns an error or raises an exception, the failure is only logged and the affected request's block IDs are never marked invalid, leaving the scheduler unaware that the transfer did not complete. This causes the scheduler to stall waiting for results that will never arrive, resulting in a crash or hang that denies service to all pending requests.

    How to fix Improper Handling of Exceptional Conditions?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • H
    Improper Check or Handling of Exceptional Conditions

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Improper Check or Handling of Exceptional Conditions via incomplete kv_transfer_params dictionary entries in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. An attacker can send requests with missing keys in the kv_transfer_params dictionary to trigger an uncaught KeyError in EngineCore scheduling, causing the decode engine to terminate and making all routed requests fail until a manual restart is performed.

    How to fix Improper Check or Handling of Exceptional Conditions?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • H
    Allocation of Resources Without Limits or Throttling

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Allocation of Resources Without Limits or Throttling via unbounded P2P KV offloading sessions in OffloadingConnector. Attackers can supply arbitrary remote host and port values in kv_transfer_params to create unreachable peer sessions that retain ZeroMQ sockets until the context quota is exhausted, causing an uncaught ZMQError that crashes EngineCore and stops all inference.

    Note: This is only exploitable when OffloadingConnector is configured with TieringOffloadingSpec and a peer-to-peer secondary tier.

    How to fix Allocation of Resources Without Limits or Throttling?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • H
    Memory Allocation with Excessive Size Value

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Memory Allocation with Excessive Size Value via the tp_size parameter in kv_transfer_params on OpenAI-compatible completion endpoints, which is not validated before use. An unauthenticated attacker can supply an arbitrary tp_size value in prefill/decode disaggregated deployments, causing unbounded memory allocation that exhausts available memory and triggers a kernel OOM-kill of the decode worker process.

    How to fix Memory Allocation with Excessive Size Value?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • H
    Reachable Assertion

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Reachable Assertion via an assertion failure in NixlBaseConnectorWorker._apply_prefix_caching, which fails to properly validate block counts across multi-prompt completion requests in prefill/decode disaggregated deployments. An attacker can submit completion requests with multiple prompts of varying lengths, triggering the assertion failure and causing the decode worker to terminate and become unavailable until restarted.

    How to fix Reachable Assertion?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • M
    Improper Handling of Exceptional Conditions

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Improper Handling of Exceptional Conditions in mooncake_connector.py via the Mooncake KV transfer connector, when a remote KV cache load fails during receive_kv_from_single_worker, the failure is not reported back to the scheduler. Instead of propagating the error so the scheduler can fail or recompute the affected request, the connector silently discards the failure, leaving the scheduler unaware that the transfer did not complete. This causes the affected request to stall indefinitely, resulting in a denial of service for that inference request. An unauthenticated remote attacker can trigger this condition by causing a KV transfer error over the network.

    How to fix Improper Handling of Exceptional Conditions?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • L
    Improper Validation of Array Index

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Improper Validation of Array Index via SamplingParams.update_from_tokenizer(), which fails to validate bad_words token indices against the model's generation output width. An authenticated attacker can supply out-of-bounds token indices that corrupt the logits memory of concurrent requests, causing other in-flight HTTP requests to receive incorrect tokens in their responses.

    How to fix Improper Validation of Array Index?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • M
    Out-of-bounds Write

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Out-of-bounds Write via the Triton _bincount_kernel, where prompt token IDs index the penalty prompt-presence bitset without bounds checking against the vocabulary size. An attacker can submit multimodal audio requests containing tokens equal to the vocabulary size, triggering out-of-bounds writes that corrupt concurrent requests' sampler state and alter repetition penalty behavior.

    How to fix Out-of-bounds Write?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • H
    Missing Release of Memory after Effective Lifetime

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Missing Release of Memory after Effective Lifetime via the decode-worker process. An attacker can cause memory exhaustion by submitting requests with max_tokens=0, leading to unbounded memory usage until the worker restarts.

    How to fix Missing Release of Memory after Effective Lifetime?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • M
    Inefficient Algorithmic Complexity

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Inefficient Algorithmic Complexity in the thinking_budget_state.py file. An attacker can cause excessive resource consumption by submitting specially crafted input that triggers inefficient processing.

    How to fix Inefficient Algorithmic Complexity?

    A fix was pushed into the master branch but not yet published.

    [0.26.0,)
    • M
    Denial of Service (DoS)

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Denial of Service (DoS) through the MoRIIOWrapper._handle_release_message process. An attacker can cause excessive resource usage by manipulating the request_id or kv_transfer_params arguments remotely.

    How to fix Denial of Service (DoS)?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • M
    Allocation of Resources Without Limits or Throttling

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Allocation of Resources Without Limits or Throttling in the Jinja Template Rendering of the /v1/chat/completions endpoint when manipulating the chat_template argument. An attacker can cause excessive resource consumption by submitting crafted input to this argument.

    How to fix Allocation of Resources Without Limits or Throttling?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • M
    Improper Resource Shutdown or Release

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Improper Resource Shutdown or Release via the TiktokenTokenizer::new function in the tiktoken vocab File Handler. An attacker can cause the application to become unresponsive or crash by providing specially crafted input to this function. This is only exploitable if the attacker has local access to the system.

    How to fix Improper Resource Shutdown or Release?

    A fix was pushed into the master branch but not yet published.

    [0.1.0,)
    • H
    Directory Traversal

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Directory Traversal via the hardcoded trust_remote_code parameter in the model loading process. An attacker can execute arbitrary code by supplying a malicious HuggingFace model repository.

    How to fix Directory Traversal?

    There is no fixed version for vllm.

    [0,)
    • M
    Interpretation Conflict

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Interpretation Conflict in the image processing pipeline. An attacker can cause the model to interpret images differently from human expectations by supplying images with manipulated EXIF orientation or PNG tRNS transparency, potentially leading to misclassification or unintended model behavior.

    How to fix Interpretation Conflict?

    A fix was pushed into the master branch but not yet published.

    [0.11.0,)
    • H
    Improper Validation of Specified Type of Input

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Improper Validation of Specified Type of Input due to improper validation of the temperature parameter while sampling. An attacker can cause the inference worker to crash or exhibit undefined behavior by supplying non-finite float values such as NaN or Infinity, which bypass validation and propagate to GPU kernels.

    How to fix Improper Validation of Specified Type of Input?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • H
    Improper Handling of Highly Compressed Data (Data Amplification)

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Improper Handling of Highly Compressed Data (Data Amplification) through the audio.py file. An attacker can cause excessive memory consumption by uploading a specially crafted compressed audio file that decompresses to a very large size, leading to resource exhaustion and potential service disruption.

    How to fix Improper Handling of Highly Compressed Data (Data Amplification)?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • M
    Insertion of Sensitive Information into Log File

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Insertion of Sensitive Information into Log File in the error handling process for certain API and WebSocket routes, where unsanitized exception messages containing sensitive memory addresses are returned in response bodies. An attacker can obtain internal memory address information by submitting malformed image data or triggering exceptions that cause object representations to be included in error messages.

    Note: This issue remains due to an incomplete fix for CVE-2026-22778.

    How to fix Insertion of Sensitive Information into Log File?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • L
    Incorrect Conversion between Numeric Types

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Incorrect Conversion between Numeric Types in the ggml_dequantize, ggml_mul_mat_vec_a8, ggml_mul_mat_a8, and ggml_moe_a8 functions when tensor dimensions are truncated due to an integer overflow. An attacker can access residual GPU memory contents from previous inference requests by supplying a specially crafted model file with tensor dimensions whose product exceeds the maximum value of a 32-bit integer.

    Note: This is only exploitable if the deployment is multi-tenant and loads attacker-controlled GGUF model files.

    How to fix Incorrect Conversion between Numeric Types?

    A fix was pushed into the master branch but not yet published.

    [0.5.5,)
    • M
    Improper Resource Shutdown or Release

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Improper Resource Shutdown or Release via the OpenAI-compatible Serving Path component. An attacker can cause the service to become unavailable by sending specially crafted requests remotely.

    How to fix Improper Resource Shutdown or Release?

    There is no fixed version for vllm.

    [0,)
    • M
    Use of Uninitialized Resource

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Use of Uninitialized Resource via the has_mamba_layers function in the KV Block Handler. An attacker can cause unintended behavior by leaking data between sessions.

    How to fix Use of Uninitialized Resource?

    A fix was pushed into the master branch but not yet published.

    [0,)
    • M
    Server-side Request Forgery (SSRF)

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Server-side Request Forgery (SSRF) via the download_bytes_from_url function. An attacker can cause the server to make arbitrary HTTP or HTTPS requests to internal or external resources by supplying a crafted file_url value in batch input JSON, potentially accessing sensitive internal services or causing denial of service.

    How to fix Server-side Request Forgery (SSRF)?

    A fix was pushed into the master branch but not yet published.

    [0.16.0,)
    • H
    Deserialization of Untrusted Data

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Deserialization of Untrusted Data via the SUB ZeroMQ socket, where the deserialization is performed using the unsafe pickle library. An attacker on the same cluster can execute arbitrary code on the remote machine by sending maliciously crafted deserialized payloads.

    Note The V0 engine is off by default since v0.8.0, and the V1 engine is not affected. Due to the V0 engine's deprecated status and the invasive nature of a fix, the developers recommend ensuring a secure network environment if the V0 engine with multi-host tensor parallelism is still in use.

    How to fix Deserialization of Untrusted Data?

    There is no fixed version for vllm.

    [0.5.2,)
    • H
    Deserialization of Untrusted Data

    vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs

    Affected versions of this package are vulnerable to Deserialization of Untrusted Data in the MessageQueue.dequeue() API function. An attacker can execute arbitrary code by sending a malicious payload to the message queue.

    How to fix Deserialization of Untrusted Data?

    There is no fixed version for vllm.

    [0,)