vllm-project/vllm
Documented errors, page 4 of 6. Back to vllm-project/vllm
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| engine start index does not fit usize | validation | error | configuration, transport, bootstrap, validation, rust, vllm, startup |
| 'mm_encoder_fp8_scale_path' and… | validation | error | vllm, config, multimodal, fp8, quantization, validation |
| request_id does not embed a peer zmq_address and… | exception | error | kv-transfer, config, routing, disaggregated-prefill |
| torch.xpu.memory is not available | exception | error | vllm, xpu, torch, sleep-mode, environment |
| VllmConfig does not have enough slots to schedule a token… | validation | error | scheduler, speculative-decoding, batched-tokens |
| VLLM_ROCM_QUICK_REDUCE_MIN_SIZE_BYTES_MB must be less than… | validation | error | rocm, allreduce, env-var, config |
| capture_torch_profiler is only applicable when profiler is… | validation | error | profiling, cuda-graph, configuration |
| synthetic_acceptance_length must be in | exception | error | speculative-decoding, config, validation, math |
| Fast prefill optimization for KV sharing is not compatible… | validation | error | kv-cache, kv-sharing, eagle, speculative-decoding |
| Fault tolerance requires a single API server process… | validation | error | fault-tolerance, api-server, configuration, parallelism |
| FP8 scale file not found | exception | error | vllm, config, multimodal, fp8, file-not-found, paths |
| Sleep-mode backend ' ' is not supported on this platform. | validation | error | vllm, sleep-mode, platform, hardware, driver |
| Number of experts in the model must be greater than 0 when… | validation | error | parallelism, moe, config, startup |
| target_model_config must be present for dspark | exception | error | speculative-decoding, dspark, config, api-misuse |
| The model is an hybrid without a layers_block_type or an… | validation | error | hybrid-model, hf-config, config, startup |
| use_inductor_graph_partition is only supported with… | validation | error | configuration, version-compat, torch, vllm |
| --use-replayssm does not support speculative decoding | validation | error | vllm, config, replayssm, speculative-decoding, mamba |
| --ssl-ca-certs is required when --ssl-cert-reqs is | validation | error | configuration, tls, mtls, security, rust, vllm, startup |
| unknown Mooncake mode | validation | error | mooncake, config, validation |
| Unsupported data type for serialization | validation | error | serialization, api-misuse, type-error |
| Unsupported platform, please use CUDA, ROCm, or CPU. | exception | error | build, requirements, platform, setup-py |
| Adaptive verification only supported with DSpark | exception | error | speculative-decoding, dspark, adaptive-verification, invalid-combination |
| embedded mode requires global_segment_size > 0 | exception | error | mooncake, config, rdma, validation |
| Endpoint not ready after | exception | error | nixl, eplb, elastic-ep, metadata-mismatch, vllm |
| local_tp_rank must be in | exception | error | tensor-parallel, kv-transfer, validation |
| Invalid numeric value | exception | error | mooncake, config, parsing |
| Serialized object size | validation | error | configuration, size-limit, serialization |
| Unexpected frame! | exception | error | zmq, handshake, protocol-mismatch, kv-transfer |
| must be in , got | validation | error | rust, validation, sampling-params, vllm |
| Rank not initialized | exception | error | kv-transfer, hf3fs, metadata-server, initialization, vllm |
| TP sizes must be positive | exception | error | tensor-parallel, kv-transfer, config, validation |
| expected str or QuantKey, got | validation | error | quantization, type-validation, pydantic, api-misuse |
| varies across layers and has no whole-model value: . Only… | validation | error | model-arch, heterogeneous-layers, internal-api |
| Invalid compilation mode | validation | error | configuration, compilation-mode, vllm |
| Could not determine Python executable. Please provide it… | exception | error | tooling, cmake, environment, interactive-prompt |
| offload_prefetch_step | validation | error | vllm, config, offloading, prefetch, validation |
| requested logprob_token_ids of length | validation | error | rust, validation, logprobs, vllm |
| torch_profiler_dir must be set when profiler is 'torch' | validation | error | profiling, torch-profiler, configuration |
| utility call ` ` returned an invalid result (call_id= ) | exception | error | serialization, version-skew, utility-call, rust |
| collect_detailed_traces requires `--otlp-traces-endpoint`… | validation | error | vllm, config, observability, tracing, validation |
| mm_tensor_ipc='torch_shm' is not supported with… | validation | error | multimodal, ipc, parallelism, config |
| torch_profiler_dir is only applicable when profiler is set… | validation | error | profiling, torch-profiler, configuration |
| Invalid format | exception | error | mooncake, config, parsing |
| remote tp_size must be a multiple of local tp_size for… | exception | error | tensor-parallel, kv-transfer, heterogeneous-tp, config |
| Model has no rendered recommended_command in the Recipes… | exception | error | vllm-recipes, api-discovery, schema-validation |
| Parent directory for FP8 scale save path not found | exception | error | vllm, config, multimodal, fp8, file-not-found, paths |
| torch.xpu.memory.XPUPluggableAllocator is not available | exception | error | vllm, xpu, torch, version-mismatch, allocator |
| Ping failed after retries | exception | critical | network, zmq, kv-transfer, connectivity, retry |
| synthetic_acceptance_rates must be non-increasing, got | exception | error | speculative-decoding, config, validation, math |
| Expected 128 bytes for ncclUniqueId, got | validation | error | nccl, serialization, distributed, validation |
| The model type does not support float16. Reason | validation | error | dtype, float16, model-support, config |
| local tp_size must be a multiple of remote tp_size for… | exception | error | tensor-parallel, kv-transfer, heterogeneous-tp, config |
| ` ` input is not supported by this model | validation | error | nixl, eplb, initialization, vllm |
| Unsupported file type | exception | warning | pre-commit, lint, spdx, license, tooling |
| No recipe model matched | exception | error | cli, recipes, search, tooling |
| request output stream for | exception | error | streaming, request-lifecycle, engine-core, rust |
| tool_choice requires at least one available tool | validation | error | rust, tool-choice, validation, chat |
| dcp_comm_backend='a2a' requires… | validation | error | context-parallelism, communication-backend, configuration |
| dspark_draft_topk is only supported by Qwen3DSparkModel | exception | error | speculative-decoding, dspark, topk, architecture-check |
| Expected model after 'vllm serve', got | exception | error | vllm-recipes, argv-parsing, cli |
| Unsupported task: Supported tasks | exception | error | pooling, task-config, config |
| Class not found in | exception | error | kv-transfer, import, external-connector, vllm |
| reducescatter is not supported | exception | error | ray, pipeline-parallel, collective, not-implemented |
| `a` must have at least 1 dimension. | exception | error | quantization, mx-format, fp4, validation, shape |
| asymmetric int8 activation quantization is unsupported on… | exception | error | quantization, int8, xpu, intel-gpu, platform-support |
| Recipe `env` must be an object, got | exception | error | vllm-recipes, env, schema-validation |
| failed to render jinja template | exception | error | rust, jinja, chat-template, renderer, minijinja |
| max_num_batched_tokens | exception | error | scheduler, config, context-length, validation |
| numa_bind_cpus entries must use numactl CPU list syntax… | validation | error | vllm, config, numa, cpuset, regex, validation |
| Decode context parallelism for GQA/MQA requires… | validation | error | parallelism, decode-context-parallel, gqa, config |
| engine-core client is closed | exception | error | lifecycle, shutdown, api-misuse, rust |
| hf3fs_fuse.io is not available. Please install the… | exception | error | kv-transfer, hf3fs, import, dependencies, vllm |
| The Inkling checkpoint does not contain MTP weights | exception | error | speculative-decoding, mtp, checkpoint, config |
| Unexpected level value | exception | error | profiler, cli, argument-validation, tooling |
| Cannot set both `pooling_type` and `tok_pooling_type` | validation | error | pooling, configuration, conflicting-fields |
| request ` ` is already in flight | validation | error | request-id, duplicate, api-misuse, rust |
| Unknown KV cache group kind | validation | error | configuration, attention-backend, kv-cache, vllm |
| handshake failed, unexpected msg type | exception | error | zmq, handshake, protocol-mismatch, kv-transfer |
| method='custom_class' requires 'model' to contain the… | exception | error | speculative-decoding, custom-class, config, api-misuse |
| --ssl-cert-reqs must be 0 (none), 1 (optional), or 2… | validation | error | configuration, tls, ssl, security, rust, vllm, startup |
| Interactive input is unavailable. Pass --model and… | exception | error | cli, interactive-prompt, ci, tooling |
| standalone-store mode requires global_segment_size == 0 | exception | error | mooncake, config, validation |
| TP sizes and total_num_kv_heads must be positive | exception | error | tensor-parallel, kv-transfer, model-config, validation |
| Unable to use nsight profiling unless workers run with Ray. | validation | error | profiling, nsight, ray, configuration |
| When PCP is enabled, DCP must be disabled, span the PCP… | validation | error | context-parallelism, parallelism, configuration |
| last dim of `a` must be divisible by 32, got | exception | error | quantization, mx-format, fp4, validation, shape, alignment |
| did not return a model list. | exception | error | network, recipes, api, json, tooling |
| Sleep-mode backend ' ' is already registered. | validation | error | vllm, sleep-mode, plugin, registry, duplicate |
| Unexpected socket type | exception | error | zmq, internal-api, validation |
| Unsupported dtype : should be one of int8, uint8, int32… | validation | error | nccl, dtype, distributed, validation |
| connected engine range | validation | error | configuration, transport, bootstrap, data-parallel, rust, vllm, startup |
| numa_bind_nodes and numa_bind_cpus require numa_bind=True. | validation | error | numa, cpu-affinity, configuration |
| consumer tp_size must be a multiple of producer tp_size for… | exception | error | tensor-parallel, kv-transfer, ack, heterogeneous-tp, config |
| logit_sigma cannot be 0 (division by zero) | validation | error | pooling, classification, calibration, configuration |
| numa_bind_nodes must not be empty. | validation | error | vllm, config, numa, parallel, validation |
| data parallel size ( ) exceeds the two-byte engine identity… | validation | error | configuration, data-parallel, limits, rust, vllm, startup |
| got per-layer configs for a model with layers | validation | error | model-arch, heterogeneous-layers, checkpoint, internal-api |
| Hardware recipe JSON does not contain a usable `strategy`… | exception | error | cli, recipes, json, validation, tooling |
| Invalid initialization parameters | http | error | kv-transfer, hf3fs, metadata-server, http-400, validation, vllm |
| chrome_trace output requires proton_data='trace' | validation | error | profiling, proton, configuration |