The probability is the direct output of the EPSS model, and conveys an overall sense of the threat of exploitation in the wild. The percentile measures the EPSS probability relative to all known EPSS scores. Note: This data is updated daily, relying on the latest available EPSS model version. Check out the EPSS documentation for more details.
In a few clicks we can analyze your entire application and see what components are vulnerable in your application, and suggest you quick fixes.
Test your applicationsUpgrade vllm to version 0.29.0 or higher.
vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs
Affected versions of this package are vulnerable to Improper Check for Unusual or Exceptional Conditions via the disaggregated serving endpoint /inference/v1/generate in vllm/entrypoints/serve/disagg/serving.py, which builds a multimodal EngineInput directly from caller-supplied token_ids without checking their length against model_config.max_model_len. When the request includes a features (multimodal) payload, GenerateRequest.token_ids in vllm/entrypoints/serve/disagg/protocol.py is not validated against the model's maximum context length. For multimodal processors where skip_prompt_length_check returns True (including Nemotron Parse, Whisper, and FireRedLID), InputProcessor._validate_prompt_len() returns immediately for both encoder and decoder prompts, allowing an overlong prompt to become an EngineCoreRequest and reach the worker input-batch copy into a fixed max_model_len-wide NumPy row. A remote attacker can submit an overlong token_ids list to trigger a NumPy broadcast failure in the worker, causing a denial of service.
Note: This is only exploitable when the deployment uses a multimodal model configuration whose processor sets skip_prompt_length_check=True (e.g. Nemotron Parse, Whisper, or FireRedLID).