sgl-project/sglang
Documented errors, page 4 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| Action endpoint is not implemented for | validation | error | action-endpoint, not-implemented, model-support, dispatch |
| allowed media domains cannot be empty | validation | error | validation, media, security, config |
| AttentionOffsetCache should not store data | error_code | error | mlx, kv-cache, api-misuse |
| eqlen with B>1 and T % | exception | error | kda, linear-attention, shape-mismatch, not-implemented |
| experimental_sgl_marlin LoRA requires… | validation | error | lora, marlin, virtual-experts, startup-validation, experimental, sglang |
| f"The class must implement the 'embedding' method, see… | exception | error | quantization, embedding, not-implemented, constructor |
| language_model does not support set_embed_and_head(). | exception | error | speculative-decoding, kimi, attribute-error, model-loading |
| MiniMax H3 SGLang backend only supports… | validation | error | minimax-h3, video-generation, output-mode, validation |
| no_rope_layers contains non-binary entries | validation | error | config-validation, mlx, rope, muse-glimmer |
| PEFT lora_alpha must be a positive integer | validation | error | lora, peft, validation, config |
| thinking.budget_tokens must be >= 1024 | validation | error | anthropic, thinking, validation, request-validation, pydantic |
| UlyssesAttention's all-to-all spans the combined sequence… | exception | critical | attention, ring-parallelism, sequence-parallel, distributed, not-implemented |
| Using a slow tokenizer. This might cause a significant… | console | warning | tokenizer, performance, huggingface, startup |
| Download failed for after attempts due to download errors… | panic | critical | network, download, huggingface, retry-exhausted, weights |
| Mamba storage zero-copy requires page_first layout, got | validation | error | sglang, mamba, layout, zero-copy, hierarchical-cache |
| MiniMax-H3 on MPS requires synchronous layerwise offload for | validation | error | minimax-h3, mps, apple-silicon, memory-offload, server-args |
| unsupported input for Sana fused bias-SiLU | exception | error | sana, diffusion, triton, memory-format, channels-last |
| Attention backend override | validation | error | attention, backend-override, configuration, enum-mismatch |
| Cannot determine attention scale for | error_code | error | mlx, attention, model-compatibility, patching |
| denoising_strength must be positive | validation | error | scheduler, flow-matching, validation, diffusion, config |
| Expected CHW image tensor, got shape | exception | error | multimodal, image-processing, tensor-shape, step3-vl |
| load_path is required for STA_inference mode | validation | error | sta, attention, sparse-tuning, kwargs-validation |
| The checkpoint declares quantization, but the model did not… | validation | error | quantization, model-mismatch, linear-layers, text-encoder, silent-failure-guard |
| DSV4 ragged verify does not support context parallel (CP)… | exception | critical | deepseek-v4, ragged-verify, context-parallelism, env-var, speculative-decoding, sglang |
| Encoder produced tokens, but preprocessor metadata expected | http | error | multimodal, token-count, encoder, preprocessor, mismatch |
| language_model does not support get_embed_and_head(). | exception | error | speculative-decoding, kimi, attribute-error, model-loading |
| _load_function expects 'pkg.module.symbol', got | exception | error | import, config-validation, debug-utils, sglang |
| Downloaded model files are still corrupted for | panic | critical | download, corruption, huggingface, validation, weights |
| encode metadata not ready | http | error | timeout, metadata, encoder, disaggregation |
| Expected a 3D packed tensor for | exception | critical | lfm2, moe, weight-loading, tensor-shape |
| expert-pack header coverage is inconsistent | exception | critical | moe, expert-pack, binary-format, header-validation, sglang |
| Including the scheme in --host | console | info | url, deprecation, host-config, networking, tests |
| Layer-sharded direct HiCache backup only supports… | validation | error | hicache, direct-io, layout, sglang |
| --mm-feature-transport=cuda_ipc requires NVIDIA CUDA. | validation | error | sglang, cuda, multimodal, ipc, hardware-requirement |
| --speculative-use-rejection-sampling is incompatible with… | validation | error | speculative-decoding, rejection-sampling, determinism, server-args |
| Subclasses of BaseDiT must define | validation | error | sglang, dit, subclass-contract, class-attribute, import-time |
| The memory capacity is unbalanced. Some GPUs may be… | error_code | error | distributed, tp, gpu-memory, resource-conflict |
| Failed to cancel VMM transport slice(s) | exception | error | cuda, vmm, aggregate-error, dispatch |
| unknown residency policy | validation | error | config, validation, layerwise-offload, residency |
| weight_prefix must be 'w13' or 'w2', got | validation | error | quantization, moe, npu, modelslim, validation |
| attn_sink must be float32 with shape | validation | error | mla, sparse-attention, tensor-validation, dtype |
| Checkpoint provides gate weight | exception | critical | laguna, weight-loading, config-mismatch, gating |
| KV cache dtype mismatch: prefill server has kv_cache_dtype= | exception | critical | disaggregation, pd-disagg, kv-cache-dtype, config-mismatch, quantization |
| NIXL KV transfer has no KV memory segments | validation | error | nixl, disaggregation, hicache, memory-registration, pd-disaggregation |
| NPU detected, but torchair package is not installed. Please… | exception | error | npu, ascend, torchair, torch-compile, import-error |
| queued MiniMax H3 jobs require pre-queue resolved temporal… | validation | critical | minimax-h3, queue-invariant, frame-count, internal |
| Unsupported model architecture | exception | error | registry, alias, architecture, unsupported-architecture |
| Ascend A2/A3 NPU does not support nvfp4… | validation | error | ascend, npu, deepep, nvfp4, quantization, hardware-support |
| canonical request missing | validation | error | minimax-h3, schema-validation, required-field |
| Could not access latents of provided encoder_output | validation | error | vae, encoder-output, attribute-access, diffusers-compat |
| DFLASH requires draft num_hidden_layers in config. Got… | validation | error | speculative-decoding, dflash, missing-config-field |
| Each permutation group must reside on the same gpu | validation | error | marlin, tensor-parallel, tile-alignment, shape-validation |
| Expected CHW image tensor with 1 or 3 channels, got shape | exception | error | multimodal, image-processing, channels, step3-vl |
| Hunyuan3D requires 'image_path' input. | validation | error | hunyuan3d, input-validation, multimodal, missing-argument |
| LPLB fused solver unavailable | exception | critical | lplb, backend-unavailable, jit, cuda |
| No model weights found in | exception | critical | mimo-audio, model-loading, file-not-found, huggingface, weights |
| Weight output_partition_size = | validation | error | fp8, quantization, tensor-parallel, shape-mismatch |
| Z-Image transformer has no `rotary_emb`. It likely loaded… | validation | critical | z-image, model-loading, fallback, rotary-embeddings, diffusers |
| Can not import FA3 in sgl_kernel. Please check your… | exception | critical | sglang, flash-attention, import-error, native-extension, environment |
| ffprobe returned invalid JSON for MiniMax H3 output | exception | error | minimax-h3, ffprobe, json-parse, validation |
| kv-canary: launch_canary_plan_kernels_torch_reference… | validation | error | kv-cache, missing-argument, ragged-tensor, validation |
| Serve backend factory returned ; expected… | exception | critical | sglang, cli, plugin-api-mismatch, type-check |
| Unsupported params_dtype | validation | error | quantization, dtype, npu, modelslim, w8a8 |
| Kimi-K3 MLA K projection must remain GGUF Q4_0 | exception | error | gguf, quantization, weight-loading, kimi, mla |
| move_kv_cache is not yet supported for MiniMaxSparseKVPool… | exception | error | sglang, kv-cache, not-implemented, speculative-decoding, minimax |
| --prefill-only-disable-kv-cache currently requires… | validation | error | sglang, kv-cache, prefill, embedding, server-args |
| seq_len < used rows | validation | error | validation, sequence-length, alignment, minimax-h3 |
| Serialized W4A8 layer | validation | error | quantization, dimension-mismatch, checkpoint |
| Unsupported image type | validation | error | multimodal, input-validation, type-error, image-processing |
| unsupported input for wan_rmsnorm_silu | validation | error | memory-format, triton, vae, wan, validation |
| DFlash layer selection requires num_target_layers >= 4. Got… | validation | error | speculative-decoding, dflash, layer-selection, config-validation |
| --hicache-host-memory-mode buffer_only requires an SWA host… | validation | error | hicache, buffer-only, swa, memory-sizing, validation |
| Hunyuan3D reference attention requires a shared cache. | validation | error | runtime, diffusion, reference-attention, missing-cache |
| are not all equal | validation | error | npu, quant-scale, weight-loading, per-tensor-quant, allclose |
| The 'enable-mixed-chunk' feature is currently unsupported… | exception | error | ascend, npu, mixed-chunk, mla, deepseek, not-implemented, huawei |
| Timesteps must be provided | validation | error | hunyuan3d, pipeline-order, missing-state, scheduler |
| Unsupported Kimi-K3 encoder media item | validation | error | multimodal, kimi-k3, input-validation |
| cached keyframe preparation disagrees with the resolved plan | exception | error | minimax-h3, cache-coherence, pipeline, keyframes |
| Cannot configure checkpoint quantization for | validation | error | quantization, checkpoint-parsing, config-error, text-encoder |
| Cannot transition from terminal state to | validation | warning | disaggregation, request-state, state-machine, terminal-state, race-condition |
| delta payload size mismatch: expected | error_code | critical | rocm, allreduce, tensor-parallel |
| --enable-linear-replayssm is not supported under PD… | validation | error | sglang, replayssm, pd-disaggregation, unsupported-feature |
| Humming does not support DeepEP | validation | error | deepep, humming, moe, dtype, config-validation |
| mask_search_files_path_pos, mask_search_files_path_neg, and… | validation | error | sta, attention, sparse-tuning, kwargs-validation, multimodal |
| Only support per-tensor scaling factor for fp8 KV cache | validation | error | kv-cache, fp8, scale-format, checkpoint-validation |
| Please install diso via `pip install diso`, or set mc_algo… | validation | error | python, import-error, missing-dependency, diso, marching-cubes, hunyuan3d |
| Rank-local TP shard produced for DTensor parameter | exception | error | tensor-parallel, fsdp, dtensor, distributed, weight-loading |
| SANA-WM streaming does not support CFG parallel; run… | exception | error | sana-wm, streaming, cfg-parallel, notimplementederror, server-args |
| unknown norm_type | validation | error | config, validation, diffusion, norm |
| req_to_token table is empty but gather mask is non-empty. | exception | error | sglang, speculative-decoding, dflash, kv-cache, memory-pool, internal-invariant |
| Cosmos3 rollout supports T2V/T2I only; I2V/V2V… | validation | error | cosmos3, rollout, i2v, sde, rl-sampling |
| Currently only attention backends are supported for… | validation | error | sglang, deterministic-inference, mla, attention-backend, deepseek |
| DeepSeek-V4 GGUF mapping collision | panic | error | gguf, deepseek, name-collision, weight-mapping |
| DSA indexer weights_proj LoRA is incompatible with… | exception | error | dsa, lora, piecewise-cuda-graph, prefill, deepseek, sglang |
| Expected Hunyuan3D2PipelineConfig, got | validation | error | hunyuan3d, pipeline-config, type-mismatch, sglang |
| fl2va keyframe preparation requires cached pre-queue probe… | exception | error | minimax-h3, fl2va, pipeline-ordering, missing-metadata |
| Host-pool retraction does not support pure-SWA models. | validation | error | disaggregation, retraction, sliding-window, swa, unified-cache, value-error |
| --mamba-max-states-per-path must be -1 (unlimited) or a… | validation | error | sglang, mamba, server-args, validation, startup |
| MIXED_PRECISION layer group | exception | error | quantization, mixed-precision, quark, mxfp4, config-validation, python |
| NIXL heterogeneous-TP direct-to-host KV transfer is not… | exception | error | nixl, disaggregation, heterogeneous-tp, hicache, pd-disaggregation |