vllm-project/vllm
Documented errors, page 3 of 6. Back to vllm-project/vllm
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| invalid --allowed-methods value | validation | error | configuration, cors, http, rust, vllm, startup |
| invalid --allowed-origins value | validation | error | cors, configuration, http-header, server, rust |
| num_speculative_tokens was provided but without speculative… | exception | error | speculative-decoding, config, cli, api-misuse |
| Sleep mode allocator is not available on platform | exception | error | vllm, sleep-mode, platform, allocator, cuda, xpu |
| tenant_id must be a string or null, got | exception | error | mooncake, config, type-error |
| numa_bind_cpus ranges must be ascending, but got | validation | error | numa, cpu-affinity, configuration, parallelism |
| offload_num_in_group | validation | error | vllm, config, offloading, kv-cache, validation |
| ` ` only provides a unified parser; the same reasoning… | exception | error | reasoning-parser, tool-parser, unified-parser, configuration, rust |
| only applicable when profiler is set to 'proton | validation | error | profiling, proton, configuration |
| server info value must serialize | panic | critical | panic, serialization, serde, server-info, rust, vllm |
| Hybrid KV cache manager was explicitly enabled but is not… | validation | error | kv-cache, hybrid-attention, kv-connector, startup-config |
| mori currently only support arch gfx942 and gfx950 | exception | error | vllm, rocm, mori, expert-parallel, gpu-arch |
| Can't determine cudagraph shapes that are both a multiple of | validation | error | |
| --use-replayssm requires --mamba-backend triton | validation | error | vllm, config, mamba, triton, replayssm |
| Cannot use --renderer-num-workers > 1 with the multimodal… | validation | error | |
| engine-core error | exception | error | wrapper, engine-core, error-chaining, rust |
| max_num_seqs ( ) exceeds available Mamba cache blocks ( )… | validation | error | |
| {pooling_type} | validation | error | pooling, invalid-value, configuration |
| Unsupported type for size | exception | error | mooncake, config, type-error, parsing |
| xpumem allocator extension is not available | exception | error | vllm, xpu, extension, vllm-xpu-kernels, sleep-mode |
| FlexKV is not installed. Please install it to use… | exception | error | kv-transfer, import, flexkv, dependencies, vllm |
| nvfp4 KV cache is not supported with MLA (Multi-head Latent… | validation | error | kv-cache, nvfp4, mla, deepseek, quantization |
| Wheel metadata missing path | exception | error | python, packaging, wheels, metadata, vllm |
| ZMQ runtime task failed | exception | error | rust, tokio, zeromq, tasks, panic |
| Arctic Inference is required for suffix decoding. Install… | exception | error | speculative-decoding, suffix-decoding, dependency, import-error |
| Compilation mode cannot be NO_COMPILATION | exception | error | torch-compile, configuration, vllm |
| MoRIIO KV cache block size mismatch for layer | exception | error | kv-cache, model-config, hybrid-attention, kv-transfer |
| torch_shm is known to fail without… | validation | error | multimodal, shared-memory, environment, multiprocessing |
| Backend error | exception | error | nixl, eplb, elastic-ep, dtype-mismatch, vllm |
| Malformed zmq_address | exception | error | zmq, kv-transfer, config, validation |
| must be non-negative or -1, got | validation | error | rust, validation, logprobs, vllm |
| requested of , which is greater than max allowed | validation | error | rust, validation, logprobs, limits, vllm |
| Stochastic rounding for Mamba cache requires the SSM cache… | validation | error | mamba, cache-dtype, stochastic-rounding, startup-config |
| The `_qutlass_C` extension is not loaded. Make sure your… | exception | error | quantization, extension-loading, cuda, environment |
| tok_pooling_type is not set; it should be resolved by… | validation | error | pooling, internal-contract, api-misuse |
| Key ' ' already exists in the storage. | validation | error | api-misuse, key-conflict, shared-memory |
| Offline data parallel mode is not supported/useful for… | validation | error | offline-inference, data-parallel, environment-variables, dense-model |
| target_model_config must be present for mtp | exception | error | speculative-decoding, mtp, config, api-misuse |
| engine-core output dispatcher closed | exception | critical | zmq, dispatcher, engine-core, process-crash, rust |
| engine input registration timed out after | exception | error | rust, startup, registration, timeout, network |
| No compatible wheel found for | exception | error | python, packaging, wheels, network, vllm |
| Async scheduling is not compatible with ROCm DeepEP… | validation | error | rocm, deepep, async-scheduling, dbo, moe |
| Expected num_speculative_tokens to be greater than zero | exception | error | speculative-decoding, num-speculative-tokens, range-validation, pydantic |
| Unknown all2all backend | validation | error | vllm, distributed, all2all, config, typo |
| `a` and `b` must be on the same device. | exception | error | quantization, mx-format, device-mismatch, validation, multi-gpu |
| unexpected frame! | exception | error | zmq, handshake, protocol-mismatch, kv-transfer |
| Could not read enough frames from video file | exception | error | video, multimodal, opencv, data-quality, validation |
| only dim 0 all-gatherv is supported | exception | error | xpu, distributed, unsupported-operation |
| request : request_id has no embedded peer zmq_address and… | exception | error | kv-transfer, config, routing, worker, disaggregated-prefill |
| suffix_decoding_max_cached_requests= | exception | error | speculative-decoding, suffix-decoding, cache, range-validation |
| The optimized moe_wna16_gemm kernel is only available on… | exception | error | moe, quantization, weight-only, cuda-only, platform-support |
| failed to build vLLM ZMQ runtime | panic | critical | panic, tokio, resource-exhaustion, zmq, rust |
| OpenTelemetry is not available. Unable to configure… | validation | error | vllm, config, observability, opentelemetry, dependencies |
| Model Runner V2 does not yet support | validation | error | model-runner, feature-support, startup-config |
| VLLM_ROCM_QUICK_REDUCE_MIN_SIZE_BYTES_MB must be… | validation | error | rocm, allreduce, env-var, config |
| Cannot merge nested option | exception | error | vllm-recipes, argv-parsing, config-merge |
| Invalid syntax ' ' for custom op, must be 'all', 'none'… | validation | error | configuration, custom-ops, vllm |
| ❌ line( ) | validation | warning | pre-commit, lint, config, ast, tooling |
| --mamba-block-size can only be set with… | validation | error | mamba, prefix-caching, block-size, startup-config |
| Unsupported node from codegen | exception | error | compilation, torch-compile, fx-graph, codegen, internal |
| invalid --allowed-headers value | validation | error | configuration, cors, http-headers, rust, vllm, startup |
| mtp_layer_types must have one entry per MTP layer: got | exception | error | speculative-decoding, mtp, checkpoint, config |
| Unknown runtime environment | exception | error | build, environment-detection, setup-py, platform |
| Attention backend 'XFORMERS' has been removed (See PR… | validation | error | vllm, config, attention-backend, multimodal, migration |
| Expected recipe argv to start with: ['vllm', 'serve'… | exception | error | vllm-recipes, argv-parsing, cli |
| Stochastic rounding for Mamba cache with triton backend… | validation | error | |
| Could not open video file | exception | error | video, multimodal, opencv, file-io, validation |
| tokenizer is missing reasoning delimiter token | exception | error | tokenizer, reasoning-parser, configuration, rust |
| unexpected engine id in startup handshake: expected | exception | error | rust, handshake, identity, configuration, duplicate |
| Unrecognized distributed executor backend | validation | error | parallelism, executor, type-validation, api-misuse |
| chat_template.json does not contain a valid template | exception | error | rust, chat-template, json, validation, transformers |
| 'mm_shm_cache_max_object_size_mb' should only be set when… | validation | error | vllm, config, multimodal, cache, validation |
| Model has no usable hardware JSON paths. | exception | error | vllm-recipes, api-discovery, data-quality |
| synthetic_acceptance_rates must have length | exception | error | speculative-decoding, config, validation |
| tool parser parsing failed | exception | warning | tool-parser, streaming, model-output, rust |
| Do not combine a positional recipe JSON source with… | exception | error | vllm-recipes, cli-usage, argument-validation |
| ec_transfer_config must be set to create a connector | validation | error | configuration, ec-transfer, factory |
| failed to parse `RUST_LOG` | panic | critical | rust, tracing, environment, startup, vllm |
| cuda_stream other than the current stream is not supported | validation | error | ray, cuda-stream, pipeline-parallel, distributed |
| EC connect must not be None | validation | error | configuration, ec-transfer, factory |
| Async scheduling is not compatible with… | validation | error | async-scheduling, speculative-decoding, startup-config |
| Cannot find CMake executable | exception | error | python, build, cmake, setup, vllm |
| duplicate tool name | validation | error | rust, tools, validation, chat, duplicate |
| Elastic EP is not compatible with data_parallel_external_lb… | validation | error | elastic-ep, load-balancing, data-parallel, not-implemented, configuration |
| Failed to find the NIXL wheel after building it. | exception | error | nixl, build, wheels, ubuntu, tooling |
| invalid method , must be 'quest' or 'abs_max | exception | error | quantization, mx-format, validation, api-misuse |
| kv_connector_module_path cannot be an empty string. | validation | error | kv-transfer, config, validation, vllm |
| torch.xpu.memory MemPool APIs are not available (need… | exception | error | vllm, xpu, torch, mempool, version-mismatch |
| Connector ' ' is already registered. | validation | error | ec-transfer, plugin-registry, duplicate-registration |
| Invalid HTTP URL: A valid HTTP URL must have scheme 'http'… | validation | error | vllm, url, validation, network, config |
| kv_transfer_config must be set to create a connector | validation | error | kv-transfer, config, vllm, connector-factory |
| allgather is not supported | exception | error | ray, pipeline-parallel, collective, not-implemented |
| Cannot set both `pooling_type` and `seq_pooling_type` | validation | error | pooling, configuration, conflicting-fields |
| Failed to find the repaired NIXL wheel. | exception | error | nixl, auditwheel, build, ubuntu, tooling |
| max_logprobs must be non-negative or -1, got | validation | error | configuration, logprobs, validation, rust, vllm, startup |
| Recipe JSON does not contain an `argv` field | exception | error | vllm-recipes, schema-drift, argv-parsing |
| tokenizer is missing unified parser token | exception | error | tokenizer, unified-parser, configuration, rust |
| ` ` does not support async scheduling yet. | validation | error | async-scheduling, executor-backend, distributed |
| when both logprobs and logprob_token_ids are set, logprobs… | validation | error | rust, validation, logprobs, vllm |
| Currently, async scheduling is only supported with… | validation | error | async-scheduling, speculative-decoding, startup-config |