Логотип exploitDog
Консоль
Логотип exploitDog

exploitDog

redhat логотип

CVE-2026-44223

Опубликовано: 12 мая 2026
Источник: redhat
CVSS3: 7.5
EPSS Низкий

Описание

vLLM is an inference and serving engine for large language models (LLMs). From 0.18.0 to before 0.20.0, the extract_hidden_states speculative decoding proposer in vLLM returns a tensor with an incorrect shape after the first decode step, causing a RuntimeError that crashes the EngineCore process. The crash is triggered when any request in the batch uses sampling penalty parameters (repetition_penalty, frequency_penalty, or presence_penalty). A single request with a penalty parameter (e.g., "repetition_penalty": 1.1) is sufficient to crash the server. This vulnerability is fixed in 0.20.0.

A flaw was found in vLLM, an inference and serving engine for large language models (LLMs). The extract_hidden_states speculative decoding proposer returns a tensor with an incorrect shape after the first decode step. This can be triggered by a remote attacker sending a request that uses sampling penalty parameters, such as repetition_penalty, frequency_penalty, or presence_penalty. Successful exploitation leads to a RuntimeError that crashes the EngineCore process, resulting in a denial of service (DoS).

Отчет

This Important denial of service flaw in vLLM, as used in Red Hat AI Inference Server and Red Hat OpenShift AI, allows a remote attacker to crash the EngineCore process. By sending a request with specific sampling penalty parameters, an attacker can trigger an incorrect tensor shape, leading to a service disruption for affected AI inference workloads.

Меры по смягчению последствий

Mitigation for this issue is either not available or the currently available options do not meet the Red Hat Product Security criteria comprising ease of use and deployment, applicability to widespread installation base, or stability.

Затронутые пакеты

ПлатформаПакетСостояниеРекомендацияРелиз
Red Hat AI Inference Serverrhaiis/vllm-cpu-rhel9Will not fix
Red Hat AI Inference Serverrhaiis/vllm-cuda-rhel9Affected
Red Hat AI Inference Serverrhaiis/vllm-neuron-rhel9Will not fix
Red Hat AI Inference Serverrhaiis/vllm-rocm-rhel9Affected
Red Hat AI Inference Serverrhaiis/vllm-spyre-rhel9Affected
Red Hat AI Inference Serverrhaiis/vllm-tpu-rhel9Will not fix
Red Hat AI Inference Serverrhaii/vllm-cpu-rhel9Affected
Red Hat AI Inference Serverrhaii/vllm-cuda-rhel9Affected
Red Hat AI Inference Serverrhaii/vllm-gaudi-rhel9Affected
Red Hat AI Inference Serverrhaii/vllm-neuron-rhel9Affected

Показывать по

Дополнительная информация

Статус:

Important
Дефект:
CWE-130
https://bugzilla.redhat.com/show_bug.cgi?id=2476827vllm: vLLM: Denial of Service via malformed tensor shape in speculative decoding

EPSS

Процентиль: 29%
0.00367
Низкий

7.5 High

CVSS3

Связанные уязвимости

CVSS3: 6.5
nvd
3 месяца назад

vLLM is an inference and serving engine for large language models (LLMs). From 0.18.0 to before 0.20.0, the extract_hidden_states speculative decoding proposer in vLLM returns a tensor with an incorrect shape after the first decode step, causing a RuntimeError that crashes the EngineCore process. The crash is triggered when any request in the batch uses sampling penalty parameters (repetition_penalty, frequency_penalty, or presence_penalty). A single request with a penalty parameter (e.g., "repetition_penalty": 1.1) is sufficient to crash the server. This vulnerability is fixed in 0.20.0.

CVSS3: 6.5
debian
3 месяца назад

vLLM is an inference and serving engine for large language models (LLM ...

CVSS3: 6.5
github
3 месяца назад

vLLM: extract_hidden_states speculative decoding crashes server on any request with penalty parameters

EPSS

Процентиль: 29%
0.00367
Низкий

7.5 High

CVSS3