vllm-project/vllm
Documented errors, page 5 of 6. Back to vllm-project/vllm
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| dspark_draft_topk is only supported by DSpark | exception | error | speculative-decoding, dspark, config, invalid-combination |
| Strategy is not available for this model/hardware. Available | exception | error | cli, recipes, strategy, validation, tooling |
| Configuration error | exception | error | nixl, eplb, initialization, vllm |
| Invalid MoRIIO backend | exception | error | moriio, kv-transfer, config, rdma, validation |
| local_buffer_size must be > 0 | exception | error | mooncake, config, validation |
| MoRIIO is not available. Please ensure the 'mori' package… | exception | critical | dependencies, mori, kv-transfer, environment |
| kv_transfer_config must be set for KVConnectorBase_V1 | validation | error | kv-transfer, config, base-class, vllm |
| log_balancedness_interval must be greater than 0. | validation | error | vllm, config, eplb, logging, validation |
| No selectable items found. | exception | error | cli, recipes, search, tooling |
| Recipe JSON must be a JSON object. | exception | error | vllm-recipes, json-shape, input-validation |
| The fused grouped_topk kernel is only available on CUDA… | exception | error | moe, topk, routing, cuda-only, platform-support |
| Pipeline parallelism is not supported for this model… | validation | error | pipeline-parallelism, model-support, config, startup |
| Connector ' ' is not registered. | validation | error | kv-transfer, registry, config, vllm |
| Unsupported op | validation | error | nccl, reduce-op, distributed, validation |
| Async EPLB is only supported with the default policy. | validation | error | vllm, config, eplb, moe, load-balancing, validation |
| Inkling LoRA requires INKLING_MULTIMEM_AR=0: the Lamport… | validation | error | |
| suffix_decoding_max_spec_factor= | exception | error | speculative-decoding, suffix-decoding, range-validation, config |
| utility call ` ` failed (call_id= ) | exception | error | rpc, utility-call, engine-core, rust |
| chat template is required but none was configured | exception | error | eplb, pynccl, initialization, vllm |
| messagepack value decode failed | exception | error | rust, deserialization, messagepack, corruption |
| The model is not multimodal. | validation | error | multimodal, api-misuse, config |
| data parallel rank is not connected to this frontend… | validation | error | data-parallel, configuration, engine-core, rust |
| Model has no per-hardware renderings in… | exception | error | vllm-recipes, api-discovery, hardware |
| unsupported multimodal content | validation | error | nixl, eplb, device-mismatch, vllm |
| Lock must be provided for readers. | validation | error | configuration, api-misuse, shared-memory |
| long_prefill_token_threshold | exception | error | scheduler, config, validation |
| PCP does not support data parallelism yet. | validation | error | context-parallelism, data-parallel, parallelism, configuration |
| Prefill context parallelism requires Model Runner V2… | validation | error | context-parallelism, model-runner, environment |
| tool_choice function | validation | error | rust, tool-choice, validation, typo, chat |
| chat template error | exception | error | eplb, elastic-ep, stateless, backend-selection, vllm |
| Unknown dtype | validation | error | dtype, hf-config, head, validation |
| chat request must contain at least one message | exception | error | eplb, pynccl, dtype, quantization, vllm |
| numa_bind_cpus entries must not be empty. | validation | error | vllm, config, numa, cpuset, parsing, validation |
| The Proton profiler requires CUDA graphs to be disabled… | validation | error | profiling, proton, cuda-graphs, startup-config |
| compile_cache_save_format must be 'binary' or 'unpacked'… | validation | error | configuration, compile-cache, vllm |
| max_num_batched_tokens | exception | error | scheduler, config, validation |
| PyYAML is required. Install it with: pip install pyyaml | exception | error | dependencies, pyyaml, cli, tooling |
| suffix_decoding_min_token_prob= | exception | error | speculative-decoding, suffix-decoding, probability, range-validation |
| Hardware is not available for this model. Available | exception | error | cli, recipes, hardware, validation, tooling |
| parsing is not available for model | exception | error | kv-events, zmq, endpoint, config, vllm |
| numa_bind_nodes must contain non-negative integers. | validation | error | vllm, config, numa, parallel, validation |
| is missing its value | exception | error | vllm-recipes, argv-parsing, parallelism |
| data parallel size must be at least 1 | validation | error | configuration, data-parallel, validation, rust, vllm, startup |
| VLLM_ROCM_QUICK_REDUCE_QUANTIZATION_MIN_SIZE_KB must be… | validation | error | rocm, allreduce, quantization, env-var |
| The environment variable 'MOONCAKE_CONFIG_PATH' is not set. | exception | error | mooncake, environment, config, startup |
| CUDA is required for this proof. | exception | error | cuda, tooling, gpus, numerics |
| generate request ` ` has an empty prompt_token_ids | exception | error | validation, prompt, api-misuse, rust |
| failed to parse chat_template.json | exception | error | rust, json, chat-template, corruption, download |
| num_redundant_experts is set to | validation | error | eplb, moe, configuration |
| Elastic EP with async EPLB requires the NIXL package… | validation | error | elastic-ep, eplb, nixl, dependencies, configuration |
| -O is missing its value | exception | error | vllm-recipes, argv-parsing, optimization-level |
| allowed_token_ids should not be empty | validation | error | rust, validation, token-ids, vllm |
| CUDAGraphMode. is not supported with backend (support: )… | validation | error | |
| CUDAGraphMode. is not supported with backend (support: ) … | validation | error | |
| engine count must be at least 1 | validation | error | configuration, transport, bootstrap, data-parallel, rust, vllm, startup |
| numa_bind_cpus must not be empty. | validation | error | vllm, config, numa, cpuset, validation |
| The Proton profiler currently supports NVIDIA CUDA only | validation | error | profiling, proton, platform-support, cuda |
| IO error | exception | critical | nixl, eplb, rdma, transfer-failure, vllm |
| At most (s) may be provided in one prompt. | validation | error | eplb, backend-selection, invalid-argument, vllm |
| parsing is disabled by frontend configuration | exception | error | kv-events, publisher, registry, config, vllm |
| Model Runner V2 requires Triton. | validation | error | triton, model-runner, environment, dependencies |
| output requires proton_data='tree | validation | error | profiling, proton, configuration |
| Elastic EP is only supported with enable_eplb=True. | validation | error | elastic-ep, eplb, moe, configuration |
| utility call ` ` closed unexpectedly (call_id= ) | exception | error | rpc, shutdown-race, utility-call, rust |
| data_parallel_rank ( ) must be in the range | validation | error | data-parallel, multi-node, launcher, configuration |
| Expert parallelism load balancing is only supported on CUDA… | validation | error | eplb, moe, platform-support, configuration |
| Hardware recipe JSON `alternatives` must be an object when… | exception | error | cli, recipes, json, validation, tooling |
| Connector name is not set in KVTransferConfig | validation | error | kv-transfer, config, vllm |
| enable_expert_parallel must be True to use EPLB. | validation | error | eplb, expert-parallel, moe, configuration |
| Unsupported sleep-mode backend | validation | error | vllm, sleep-mode, config, registry, typo |
| `--decode-context-parallel-size | validation | error | parallelism, decode-context-parallel, gqa, config |
| Total number of attention heads | validation | error | tensor-parallelism, attention, config, startup |
| The model's number of query heads per KV head | validation | error | parallelism, decode-context-parallel, gqa, config |
| data_parallel_external_lb can only be set when… | validation | error | data-parallel, load-balancing, configuration, parallelism |
| synthetic_acceptance_rates entries must be in [0, 1], got | exception | error | speculative-decoding, config, validation, math |
| data_parallel_size_local | validation | error | data-parallel, configuration, multi-node, parallelism |
| engine start index + engine count overflows | validation | error | configuration, transport, bootstrap, overflow, rust, vllm, startup |
| Quantization method is an override but is has not been… | validation | error | |
| EPLB requires tensor, prefill-context, or data parallelism… | validation | error | eplb, expert-parallel, moe, parallelism, configuration |
| cannot continue the final message when the last message is… | exception | error | eplb, pynccl, nccl, vllm |
| Invalid request format: need 'rank' and 'keys' | http | error | kv-transfer, hf3fs, metadata-server, http-400, validation, vllm |
| Quantization method specified in the model config | validation | error | |
| a model must have at least one layer | validation | error | model-arch, internal-api, validation |
| Image generation should not fail | panic | error | eplb, pynccl, nccl, device-mismatch, vllm |
| proton_profiler_dir must be set when profiler is 'proton' | validation | error | profiling, proton, configuration |
| tp_size= must be divisible by dcp_size= . | validation | error | context-parallelism, tensor-parallel, parallelism, configuration |
| Elastic EP is not supported with pipeline parallelism | validation | error | elastic-ep, pipeline-parallel, moe, configuration |
| Mamba SSU algorithm selection is only supported with the… | validation | error | |
| tokenizer must be a string, got | validation | error | |
| Stochastic rounding for Mamba cache is only supported on… | validation | error | |
| Failed to infer device type, please set the environment… | validation | error | |
| EPLB communicator 'pynccl' requested but expert weights… | exception | error | |
| Size cannot be empty. | exception | error | mooncake, config, parsing |
| max_model_len must be a positive integer, got | validation | error | |
| Video pruning method | validation | error | |
| custom_ops cannot both enable and disable the same… | validation | error | |
| Embedding models do not support `--runner | validation | error | |
| This model does not support `--runner pooling`. You can… | validation | error | |
| No compilation mode is set. This method should only be… | validation | error | |
| multimodal input is not supported by this chat renderer | validation | error | nixl, eplb, dependency, backend-selection, vllm |