Логотип exploitDog
Консоль
Логотип exploitDog

exploitDog

redhat логотип

CVE-2026-53923

Опубликовано: 22 июн. 2026
Источник: redhat
CVSS3: 4.3

Описание

vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor processing. The output tensor is allocated at full size via torch::empty (uninitialized memory), but the dequantize CUDA kernel processes only a truncated number of elements. The unfilled portion of the output tensor retains whatever was previously in GPU memory. In multi-tenant inference deployments, this residual GPU memory may contain tensor data from other users' inference requests, constituting information disclosure. This vulnerability is fixed in 0.23.1rc0.

A flaw was found in vLLM. Integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels leads to partial tensor processing. This results in the output tensor retaining previously used GPU memory, which, in multi-tenant inference deployments, can expose sensitive tensor data from other users' requests. This constitutes an information disclosure vulnerability.

Отчет

Red Hat rates this issue as having Low impact for Red Hat AI products. The upstream issue is limited information disclosure via integer truncation in vLLM sampling parameters. Red Hat OpenShift AI, Red Hat AI Inference Server, and Red Hat Enterprise Linux AI images are not considered affected because untrusted clients cannot control the vulnerable parameters in supported deployment models.

Меры по смягчению последствий

No mitigation is required for unaffected deployments. Restrict untrusted access to inference APIs as a general hardening measure.

Затронутые пакеты

ПлатформаПакетСостояниеРекомендацияРелиз
Red Hat AI Inference Serverrhaiis/vllm-cpu-rhel9Not affected
Red Hat AI Inference Serverrhaiis/vllm-cuda-rhel9Not affected
Red Hat AI Inference Serverrhaiis/vllm-neuron-rhel9Not affected
Red Hat AI Inference Serverrhaiis/vllm-rocm-rhel9Not affected
Red Hat AI Inference Serverrhaiis/vllm-spyre-rhel9Not affected
Red Hat AI Inference Serverrhaiis/vllm-tpu-rhel9Not affected
Red Hat AI Inference Serverrhaii/vllm-cpu-rhel9Not affected
Red Hat AI Inference Serverrhaii/vllm-cuda-rhel9Not affected
Red Hat AI Inference Serverrhaii/vllm-gaudi-rhel9Not affected
Red Hat AI Inference Serverrhaii/vllm-neuron-rhel9Not affected

Показывать по

Дополнительная информация

Статус:

Low
Дефект:
CWE-824
https://bugzilla.redhat.com/show_bug.cgi?id=2491579vllm: vLLM: Information disclosure via integer truncation

4.3 Medium

CVSS3

Связанные уязвимости

CVSS3: 7.5
nvd
около 1 месяца назад

vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor processing. The output tensor is allocated at full size via torch::empty (uninitialized memory), but the dequantize CUDA kernel processes only a truncated number of elements. The unfilled portion of the output tensor retains whatever was previously in GPU memory. In multi-tenant inference deployments, this residual GPU memory may contain tensor data from other users' inference requests, constituting information disclosure. This vulnerability is fixed in 0.23.1rc0.

CVSS3: 7.5
debian
около 1 месяца назад

vLLM is an inference and serving engine for large language models (LLM ...

CVSS3: 7.5
github
около 2 месяцев назад

vLLM: GGUF dequantize kernel int truncation exposes uninitialized GPU memory in multi-tenant serving

4.3 Medium

CVSS3