Логотип exploitDog
Консоль
Логотип exploitDog

exploitDog

redhat логотип

CVE-2026-34760

Опубликовано: 02 апр. 2026
Источник: redhat
CVSS3: 5.9

Описание

vLLM is an inference and serving engine for large language models (LLMs). From version 0.5.5 to before version 0.18.0, Librosa defaults to using numpy.mean for mono downmixing (to_mono), while the international standard ITU-R BS.775-4 specifies a weighted downmixing algorithm. This discrepancy results in inconsistency between audio heard by humans (e.g., through headphones/regular speakers) and audio processed by AI models (Which infra via Librosa, such as vllm, transformer). This issue has been patched in version 0.18.0.

A flaw was found in Librosa, a software library used by artificial intelligence (AI) models like vLLM for processing audio. The library's method for converting stereo audio to mono differs from international standards, causing AI models to interpret audio differently than humans. This inconsistency could allow an attacker to provide specially crafted audio, leading to the AI model processing the data incorrectly. This could result in a high impact on the integrity of the AI system's operations.

Отчет

The vLLM inference and serving engine for LLM models uses a flawed version of librosa, which downmix audios to mono while international standards mandates it should be used a weight downmixing algorithm. When vLLM is processing an audio file this may result in discrepancy between what's processed by the AI model being served and what humans are able to hear. An attacker may leverage that by crafting a multichannel audio tracker with Low-Frequency Effects interference, hiding content that cannot be heard by the user but may mask critical audio features that are essential to speech recognition. This may result in integrity or safety issues as it may eventually by-pass content moderation or voice recognition based authentication systems. Red Hat Product Security team has rated this vulnerability as having a Moderate security impact, as crafting a valid malicious audio file is considered a high complexity task.

Меры по смягчению последствий

Mitigation for this issue is either not available or the currently available options do not meet the Red Hat Product Security criteria comprising ease of use and deployment, applicability to widespread installation base, or stability.

Затронутые пакеты

ПлатформаПакетСостояниеРекомендацияРелиз
Red Hat AI Inference Serverrhaiis-preview/vllm-cuda-rhel9Fix deferred
Red Hat AI Inference Serverrhaiis/vllm-cpu-rhel9Fix deferred
Red Hat AI Inference Serverrhaiis/vllm-cuda-rhel9Fix deferred
Red Hat AI Inference Serverrhaiis/vllm-neuron-rhel9Fix deferred
Red Hat AI Inference Serverrhaiis/vllm-rocm-rhel9Fix deferred
Red Hat AI Inference Serverrhaiis/vllm-spyre-rhel9Fix deferred
Red Hat AI Inference Serverrhaiis/vllm-tpu-rhel9Fix deferred
Red Hat Enterprise Linux AI (RHEL AI) 3rhelai3/bootc-aws-cuda-rhel9Fix deferred
Red Hat Enterprise Linux AI (RHEL AI) 3rhelai3/bootc-azure-cuda-rhel9Fix deferred
Red Hat Enterprise Linux AI (RHEL AI) 3rhelai3/bootc-azure-rocm-rhel9Fix deferred

Показывать по

Дополнительная информация

Статус:

Moderate
Дефект:
CWE-358
https://bugzilla.redhat.com/show_bug.cgi?id=2454645vLLM: Librosa: numpy: Librosa: AI model data integrity impact due to audio processing discrepancy

5.9 Medium

CVSS3

Связанные уязвимости

CVSS3: 5.9
nvd
4 месяца назад

vLLM is an inference and serving engine for large language models (LLMs). From version 0.5.5 to before version 0.18.0, Librosa defaults to using numpy.mean for mono downmixing (to_mono), while the international standard ITU-R BS.775-4 specifies a weighted downmixing algorithm. This discrepancy results in inconsistency between audio heard by humans (e.g., through headphones/regular speakers) and audio processed by AI models (Which infra via Librosa, such as vllm, transformer). This issue has been patched in version 0.18.0.

CVSS3: 5.9
debian
4 месяца назад

vLLM is an inference and serving engine for large language models (LLM ...

CVSS3: 5.9
github
25 дней назад

vLLM: Processing differential in multi-channel audio downmixing enables hidden-input/moderation bypass for audio models

5.9 Medium

CVSS3