ErrLookup › vllm-project/vllm
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs · Python · 2,614 source files
Analyzed at c794754062 on 2026-08-14. 543 documented errors.
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| tokenize endpoint unavailable | exception | error | moriio, zmq, protocol-version, kv-transfer, concurrency, vllm |
| Tokenizer error | exception | critical | moriio, rdma, timeout, kv-transfer, vllm |
| JSON error | exception | critical | moriio, rdma, kv-transfer, memory-registration, vllm |
| HTTP request failed | exception | error | moriio, kv-transfer, validation, numpy, vllm |
| KV connector is incompatible with… | validation | error | kv-transfer, allocator, environment, disaggregation |
| DeepSeek V4 does not support PIECEWISE CUDA graphs with… | validation | error | cuda-graph, deepseek, model-runner, config |
| Request is not in _unfinished_requests, but it is scheduled… | exception | error | lmcache, scheduler, state-inconsistency, kv-transfer |
| messagepack encode failed for | exception | error | rust, serialization, messagepack, protocol |
| Cannot run the multi-modal processor on | validation | error | vllm, config, multimodal, device, gpu, oom, disaggregation |
| cannot use in-process coordinator with bootstrapped… | panic | critical | rust, config, coordinator, transport, startup |
| --enable-return-routed-experts is incompatible with KV… | validation | error | moe, expert-routing, kv-transfer, disaggregation |
| Mooncake Transfer Engine initialization failed. | exception | critical | mooncake, rdma, initialization, config, kv-transfer |
| --enable-return-routed-experts is incompatible with… | validation | error | moe, pipeline-parallelism, expert-routing, config |
| No KV cache tensors were registered with Mooncake. | exception | critical | mooncake, kv-cache, registration, kv-transfer |
| Either vllm_config must be provided, or all of… | validation | error | lmcache, config, api-misuse, kv-transfer |
| synthetic_acceptance_rates / synthetic_acceptance_length… | exception | error | speculative-decoding, config, validation |
| unexpected startup handshake message | exception | error | rust, handshake, protocol, validation, version-skew |
| parser ` ` is not registered | exception | error | rust, parser, configuration, registry, structured-output |
| Confirmation failed | http | error | hf3fs, metadata-server, http-500, state-corruption, kv-transfer |
| vLLM failed to compile the model. The most likely reason… | exception | error | compilation, torch-compile, cache, vllm |
| unsupported auxiliary frame(s): expected 1 frame, got | exception | error | rust, zeromq, multimodal, aux-frames, limits |
| Target and draft model should have the same vocabulary… | validation | error | speculative-decoding, tokenizer, config, validation |
| The quantization method | validation | error | quantization, gpu-capability, hardware, config |
| sampling distribution replay requires… | validation | error | sampling, logprobs, config, validation |
| Mooncake preferred_segment override must be a non-empty… | validation | error | mooncake, rdma, config, validation |
| `min_tokens` must be less than or equal to `max_tokens`… | validation | error | rust, validation, sampling-params, vllm |
| text request stream ` ` closed before terminal output | exception | error | rust, streaming, engine-core, vllm |
| is not a valid config field | validation | error | config, validation, typo, api-misuse |
| Unsupported new_block_ids type | exception | error | lmcache, version-mismatch, type-validation, scheduler, kv-transfer |
| managed Python headless engine exited unexpectedly with… | exception | critical | rust, engine, process-exit, oom, cuda, supervisor |
| this model's maximum context length is | validation | error | request-validation, context-length, prompt, tokenizer, api, rust, vllm |
| LMCacheMPConnector only works without hybrid kv cache… | exception | error | lmcache, hybrid-kv-cache, config, kv-transfer, cli-flag |
| Engine ID mismatch for dp_rank= | http | error | mooncake, bootstrap-server, engine-id, dp-rank, http-400, kv-transfer |
| MLA only works with naive serde mode.. | validation | error | lmcache, mla, config, serde, kv-transfer |
| Invalid request format: need 'rank' and 'confirmations' | http | error | hf3fs, metadata-server, http-400, request-validation, kv-transfer |
| Override for must be a mapping or , got | validation | error | config, validation, type-mismatch, api-misuse |
| sampling distribution replay does not support speculative… | validation | error | sampling, speculative-decoding, config, incompatibility |
| invalid structured outputs params | validation | error | rust, structured-outputs, validation, request-building |
| Mooncake batch memory registration failed. | exception | critical | mooncake, rdma, memory-registration, gpu, kv-transfer |
| coordinator requires a Python-compatible two-byte engine… | validation | error | rust, coordinator, identity, configuration |
| Unsupported speculative method | exception | error | speculative-decoding, config, not-implemented, method-dispatch |
| disable_any_whitespace is only supported for xgrammar and… | validation | error | structured-output, grammar, config, validation |
| received a positional argument of type , but no parameter… | exception | error | torch-compile, decorators, type-checking, vllm |
| Source code has changed since the last compilation… | exception | error | compilation, cache-invalidation, source-tracking, vllm |
| sampling distribution replay does not support diffusion… | validation | error | sampling, diffusion, config, incompatibility |
| use_heterogeneous_vocab only works with method='draft_model' | exception | error | speculative-decoding, config, tokenizer, validation |
| World size ( ) is larger than the number of available GPUs… | validation | critical | parallelism, gpu, configuration, startup |
| Mooncake is not available | exception | critical | mooncake, import-error, dependencies, rdma, kv-transfer |
| MooncakeStoreConnector does not support | validation | error | mooncake, kv-transfer, hybrid-attention, mamba, context-parallel, config |
| call_module is not allowed for codegen target | exception | error | compilation, torch-compile, fx-graph, codegen, internal |
| module has no attribute | exception | error | vllm, lazy-import, pep562, attributeerror, api-surface |
| No precompiled vllm wheel found for architecture | exception | error | build, wheels, precompiled, architecture, setup-py |
| PostGradPassManager can not be kept in CompilationConfig. | exception | error | compilation, torch-compile, configuration, api-misuse |
| unsupported field ` ` in | validation | error | rust, validation, unsupported-feature, request-building |
| CUDART error | exception | error | cuda, distributed, native-bindings, runtime |
| Numerics check failed for case | exception | error | python, benchmark, helion, numerics, testing |
| Selected model has no JSON API path. | exception | error | vllm-recipes, api-discovery, schema-validation |
| Failed to connect to metadata server | exception | critical | hf3fs, metadata-server, network, connection-failure, kv-transfer |
| sampling distribution replay does not support custom logits… | validation | error | sampling, logits-processors, config, incompatibility |
| Not enough space in the data buffer, try calling free_buf()… | exception | error | shared-memory, out-of-memory, distributed |
| engine core reported fatal failure | exception | critical | rust, engine-crash, lifecycle, zeromq |
| prompt_lookup_min= must be <= prompt_lookup_max= | exception | error | speculative-decoding, ngram, config, validation |
| The compiled artifact is not serializable. This usually… | exception | error | compilation, cache, serialization, torch-compile, version-skew |
| disable_additional_properties is only supported for the… | validation | error | structured-output, grammar, json-schema, config |
| Hidden-states block-size mismatch: derived | exception | error | kv-transfer, block-size, hybrid-kv-cache, vllm |
| Assigning / modifying buffers of nn.Module during forward… | exception | error | cudagraph, bytecode-hook, silent-corruption, vllm |
| communicator is incompatible with async EPLB due to NCCL… | validation | error | vllm, config, eplb, nccl, moe, validation |
| is not supported for quantization method . Supported dtypes | validation | error | quantization, dtype, config, validation |
| --enable-return-routed-experts is incompatible with context… | validation | error | moe, context-parallelism, expert-routing, config |
| Unknown KVConnectorRole | exception | error | lmcache, enum, version-mismatch, config, kv-transfer |
| No device communicator found | validation | error | elastic-ep, distributed, initialization |
| OpenAI server shut down unexpectedly without error | exception | error | rust, server, shutdown, lifecycle, supervisor |
| Partial-tail offloads for one request must share a boundary | exception | error | mooncake, kv-transfer, internal-invariant, hybrid-attention |
| failed to build request runtime | panic | critical | panic, tokio, runtime, resources, limits, rust, vllm, startup |
| `thinking_token_budget` must be a non-negative integer or… | validation | error | rust, validation, thinking, sampling-params, vllm |
| use_heterogeneous_vocab currently only supports greedy… | exception | error | speculative-decoding, config, sampling, validation |
| aot_compile is not supported by the current configuration… | exception | error | aot-compile, version-compat, torch-compile, vllm |
| padded_n is not supported with TRTLLM 8x4 scale layout. | exception | error | quantization, nvfp4, trtllm, gpu, validation |
| cannot combine `--headless` with… | validation | error | rust, cli, configuration, headless, argument-conflict |
| sampling distribution replay requires Model Runner V2 | validation | error | sampling, model-runner, config, validation |
| Invalid wheel filename format | exception | error | rocm, packaging, wheel, pep427 |
| combined parser is constructed from split parser instances | exception | error | unified-parser, api-misuse, constructor, rust |
| Request is not in _unfinished_requests | exception | error | mooncake, kv-transfer, scheduler, race-condition, internal-bug |
| engine control channel closed unexpectedly | exception | critical | zmq, engine-core, ipc, process-crash, rust |
| tokenizer error | exception | error | tokenizer, huggingface, model-loading, network, rust, vllm |
| Flashinfer allreduce is not supported for multi-node… | validation | error | distributed, flashinfer, allreduce, multi-node, config |
| tool call stream state is inconsistent | exception | error | rust, tool-calls, streaming, invariant, internal |
| unknown quantization name | validation | error | quantization, config, validation, pydantic |
| No dynamic dimensions found in the forward method of | exception | error | torch-compile, decorators, dynamic-shapes, vllm |
| Could not uniquely identify the extract-hidden-states KV… | exception | error | kv-transfer, hidden-states, kv-cache-groups, vllm |
| Expected but got for | exception | error | compilation, cache, version-skew, torch-compile |
| Data for address:id ' : ' has been modified or is invalid. | validation | error | shared-memory, stale-handle, race-condition |
| Timed out waiting for EC mmap file to reach | exception | critical | ec-transfer, timeout, filesystem, startup |
| unexpected output on main dispatcher path | exception | error | rust, dispatcher, output-routing, version-skew, protocol |
| 'tensor_parallel_size' is not a valid argument in the… | exception | error | speculative-decoding, config, tensor-parallelism, renamed-field |
| received pp_rank > 0 handshake metadata but does not… | validation | error | kv-transfer, pipeline-parallelism, handshake, vllm |
| CUDA graph capturing detected at an inappropriate time… | exception | error | cudagraph, runtime-monitor, cuda, vllm |
| Either prompt_lookup_max or prompt_lookup_min must be… | exception | error | speculative-decoding, ngram, config, dead-code |
| cannot be other value than 1 or target model… | exception | error | speculative-decoding, tensor-parallelism, draft-model, config |
| Insufficient space in | exception | error | shared-memory, docker, distributed, resource-limits |