sgl-project/sglang
Documented errors, page 18 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| num_heads // num_epi_subtiles must be divisible by 4 (FMA… | validation | error | cutedsl, kernel-config, shape-validation, attention |
| --prefill-only-disable-kv-cache is incompatible with… | validation | error | sglang, context-parallelism, prefill, kv-cache, server-args |
| q must be contiguous | validation | error | contiguity, cuda, sparse-attention |
| return_hidden_states_mode must be one of: None, 'last', or… | validation | error | server-args, hidden-states, cli-validation, sglang |
| schema_ is required for json_schema response format request. | validation | error | json-schema, response-format, validation |
| t2va takes no conditioning inputs; pick another task | validation | error | comfyui, sgldiffusion, minimax-h3, input-validation |
| Timed out waiting for ACK from FlexKV layerwise worker | exception | critical | flexkv, timeout, eventfd, worker-hang |
| total_consumer_count must be positive | validation | error | cuda-ipc, config-validation, multimodal-transport |
| Unknown sparse algorithm | validation | error | sparse-attention, factory, algorithm-name, value-error |
| Unsupported activation | exception | error | laguna, activation, config-validation |
| Unsupported socket type | validation | error | zmq, network, socket |
| Unsupported sol_attn dense_backend= | validation | error | sol-attn, config, invalid-value, enum |
| ' ' is a reserved speculative algorithm name; cannot be… | validation | error | speculative-decoding, plugin-api, name-collision |
| Cannot pad RoPE freqs of length | validation | error | z-image, rope, shape-validation |
| Cannot release inactive | error_code | error | multimodal, memory-pool, double-free, lease-lifecycle |
| Cannot update weights while the server is sleeping. Call… | error_code | error | sleep-wake, weight-update, rl-workflow |
| Expected audio_latent shape [B, T, C], got shape= | validation | error | joy-echo, audio, latent, tensor-shape, memory-slot |
| Expected one Kimi-K3 image span for each image | validation | error | multimodal, kimi-k3, deferred-preprocessing, span-mismatch |
| Expected scheduler.sigmas to be a tensor for LTX-2. | exception | error | ltx-2, scheduler, sigmas, type-validation |
| FA4 path does not support non-consecutive batch indices or… | exception | error | flash-attention, fa4, batch-indices, left-padding, unsupported-operation |
| Invalid causality_axis | validation | error | python, value-error, invalid-argument, ltx-2-audio, causal-conv |
| Invalid decoupled draft scheduler rid | validation | error | sglang, speculative-decoding, decoupled, rid-parsing, validation |
| Invalid simulate_acc_method | validation | error | speculative-decoding, validation, configuration |
| Latents must be provided | validation | error | hunyuan3d, pipeline-order, missing-state, latents |
| LoRA with name does not exist. Loaded LoRAs | validation | error | lora, registry, not-found, unload, sglang |
| max_new_tokens must be at least 0, got | validation | error | sampling-params, validation, max-new-tokens, sglang |
| `mixed_qkv` must be contiguous in the last dim. | validation | error | fla, fused-recurrent, contiguity, stride-check |
| No processor registered for architecture | validation | error | sglang, multimodal, unsupported-model, version-mismatch |
| Online MXFP4 quantization for MoE is only supported on AMD… | exception | critical | quantization, mxfp4, aiter, rocm, dependency-missing |
| PD decode DCP requires --disaggregation-transfer-backend… | validation | error | sglang, pd-disaggregation, dcp, transfer-backend, config-validation |
| q must have shape (s_q, h_q, d_qk), got | validation | error | rank-validation, shape-validation, sparse-mla, q8kv8 |
| 'Req' object has no attribute 'sampling_params' | validation | error | attribute-access, req, serialization |
| RMSNorm expected hidden size | exception | error | rmsnorm, shape-mismatch, model-architecture |
| Tag mismatch: expected CMD_STORE_COMPLETE, got | exception | critical | flexkv, pipeline-parallel, protocol-mismatch, distributed |
| The params dtype must be float16, but got | validation | error | marlin, dtype, float16, config-validation |
| Unsupported modality for EPD preprocessing | validation | error | modality, dispatch, input-validation, multimodal |
| v_pool shape must match k_pool shape, got | validation | error | shape-validation, kv-cache, metal, rope, sgl-kernel |
| attention-backend='nsa' is deprecated; use 'dsa' instead… | console | warning | deprecation, attention-backend, nsa, dsa, migration |
| Comfy W4A8 layer needs a 2D packed weight, got | validation | error | quantization, shape-mismatch, w4a8, safetensors |
| Crusoe API key required. Pass api_key= or set… | validation | error | frontend, api-key, missing-env-var, crusoe, sglang |
| DFLASH requires explicit layer ids for aux hidden capture. | validation | error | dflash, speculative-decoding, hidden-states, qwen3 |
| DSV4 target and draft pools must share paged SWA geometry… | validation | error | disaggregation, deepseek-v4, page-size, sliding-window |
| embed_dim must be divisible by num_heads | validation | error | siglip2, vision-config, divisibility |
| For multimodal input processing do not set… | validation | error | sglang, config, multimodal, batch-encode |
| head_dim mismatch: unified_kv= | validation | error | attention, head-dim, shape-validation, kv-cache |
| `initial_state` must be 4D | validation | error | pytorch, tensor-shape, replayssm, state-management |
| Invalid backend: . Must be one of | validation | error | config, enum-validation, backend-selection |
| . is required | validation | error | input-validation, multimodal, missing-field |
| kv-canary: launch_plan_entries_kernel requires… | validation | error | kv-canary, plan-entries, argument-validation |
| kv-canary: speculative_num_draft_tokens must be… | exception | error | kv-canary, validation, speculative-decoding, value-error |
| kv_scales shape does not match expected ( , ) | validation | error | fp8, kv-cache, shape-validation, scales |
| MiniCPM SALA does not support hierarchical cache | validation | error | minicpm, hierarchical-cache, hicache, sglang |
| out must be a contiguous tensor with the expected shape… | validation | error | out-buffer, shape-mismatch, allocation, ulysses |
| PD state transfer does not support TP-mismatched non-MLA… | exception | error | disaggregation, swa, tensor-parallel, heterogeneous-tp, unsupported-feature |
| Please install mooncake by following the instructions at… | exception | error | mooncake, import, missing-dependency, installation |
| --quantization nvfp4_online supports only… | validation | error | sglang, nvfp4, moe-runner-backend, quantization, config-conflict |
| Regular expression is not supported in the VertexAI backend. | console | warning | regex, vertexai, structured-output, constraint-dropped, warning |
| SANA-WM refiner decoding expects decoded video shaped (B… | validation | error | sana-wm, refiner, vae, decode-shape, valueerror |
| scalar_type_id doesn't exists. | validation | error | scalar-type, registry, version-mismatch, quantization |
| Serialized kitchen_int8 layer | validation | critical | quantization, convrot, group-size, checkpoint |
| state index length mismatch: prefill= , dst= | exception | critical | pd-disagg, state-index, kv-corruption-guard |
| The layout of q is not supported | exception | error | cuda, flash-attention, memory-layout, sm100 |
| unknown action keys ; allowed keys are | validation | error | validation, action-string, whitelist, sana-wm |
| Unknown visual_type | validation | error | visual, dispatch, input-validation, version-skew |
| Unsupported layout for models with head_dim != v_head_dim | validation | error | sglang, hicache, layout, kv-cache |
| Weight input_size_per_partition = | validation | error | marlin, tensor-parallel, group-size, shape-validation |
| `b` must have shape [B, HV] with HV= | exception | error | kda, helion, shape-validation, packed-layout |
| component_attention_backends must be a dict or a… | validation | error | config, type-error, attention-backend |
| D= must be divisible by GROUP_SIZE= | validation | error | fp8, head-dim, kv-cache, shape-validation |
| DSpark could not resolve speculative_num_draft_tokens; set… | validation | error | speculative-decoding, dspark, missing-argument, num-draft-tokens |
| Either spatial_upsample or temporal_upsample must be True | validation | error | upsampler, config, init, value-error |
| FA4 does not support updating KV cache in-place. | validation | error | flash-attention, fa4, kv-cache, unsupported-operation |
| FusedScaleResidualNormScaleShift cuda not available, using… | console | warning | cuda, kernel-fallback, layernorm, performance, shape-constraint |
| GGUFConfig must be constructed from a GGUF checkpoint | validation | error | gguf, quantization, config, unsupported-operation |
| H3 conditioning projection expects width | exception | critical | minimax-h3, conditioning-projection, forward, width-mismatch |
| Hidden size must be divisible by num_attention_heads | exception | critical | config-validation, attention-heads, init-time, joyimage |
| Kimi-K3 MLA V projection must remain GGUF Q2_K | exception | error | gguf, quantization, weight-loading, kimi, mla |
| LoRA adapter ' ' was requested, but LoRA is not enabled… | validation | error | lora, adapter, server-args, configuration |
| LTX-2 SP time-sharding for packed token latents currently… | validation | error | sglang, ltx-2, sequence-parallelism, patch-size, video, latent-sharding |
| mm_content_hashes has | validation | error | multimodal, artifact-cache, kimi-k3, input-validation |
| MXFP8 KV cache does not support DCP KV masks. | exception | error | kv-cache, mxfp8, dcp-mask, not-implemented |
| No pre-tokenizer regex known for tokenizer.ggml.pre= | exception | error | gguf, tokenizer, version-mismatch |
| num_draft_layers must be positive, got | validation | error | speculative-decoding, dflash, config-validation |
| num_fused_shared_experts > 1 | validation | critical | moe, shared-experts, cuda, platform-limit, bailing |
| Query heads not divisible by KV heads | validation | error | quest, sparse-attention, gqa, head-mismatch, value-error |
| Unknown separator style | validation | error | chat-template, json-template, sep-style, enum-key |
| unsupported tar material URI | validation | error | minimax-h3, tar-uri, scheme, material-io |
| Attention backend ' ' is not supported by this attention… | validation | error | attention-backend, fail-closed, config-validation, multimodal |
| block_token_tags must cover the rank-local packed sequence | validation | error | minimax-h3, token-tags, sequence-parallelism |
| capture layout needs 1 <= num_slots <= num_tokens, got… | validation | error | speculative-decoding, cuda-graph, validation |
| Completion template is not a built-in template name or a… | error_code | error | completion-template, template-manager, file-not-found, server-args |
| Expected RGB video with trailing channel dim 3, got shape= | validation | error | joy-echo, video, rgb, channels, tensor-shape |
| Hunyuan3D SD2.1 UNet requires two ResNet layers and one… | validation | error | stable-diffusion, unet, config-validation, hunyuan3d |
| kda_prefill is the inference forward path: cp_context, and… | exception | error | kda, training-vs-inference, not-implemented, invalid-argument |
| state payload requires transitions | validation | error | realtime, control-events, schema, validation |
| max_tokens must be positive | validation | error | anthropic, max-tokens, validation, request-validation |
| MiniMax H3 Ulysses size must be positive. | validation | critical | sequence-parallel, ulysses, config, validation |
| Missing required argument for SparseVideoGen2Attention | validation | error | kwargs, metadata, attention-backend, validation |
| moe_a2a_backend='pplx' only supports low-latency mode; set… | validation | error | pplx, deepep-mode, moe, a2a-backend, server-args |
| must have shape , got | validation | error | shape-validation, output-buffer, sparse-mla, fp8 |