sgl-project/sglang
Documented errors, page 6 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| Weight not found in params_dict | validation | critical | checkpoint-loading, vision-tower, weight-mapping |
| Block-quantized lm_head is not supported; use channel or… | exception | error | quantization, lm-head, block-quantization, weight-loading, tp-sharding |
| Both dtype and regex are set. Only dtype will be used. dtype | console | warning | dtype, regex, structured-output, conflict, sampling-params |
| Either quant_config or online_scheme must be provided | exception | error | quantization, quark, constructor, required-argument, python |
| _block_cnt and _block_idx must be on the same device | validation | error | block-sparse, device-mismatch, attention, cuda |
| page-major layout has no per-layer contiguous regions; KV… | exception | error | kv-cache, page-major, disaggregation, not-implemented |
| RadixKey index out of range | exception | error | radix-cache, index-error, off-by-one, prefix-cache |
| Serialized quantized component weights cannot use a stacked… | validation | error | quantization, state-dict, parameter-mapping, text-encoder |
| action output must have shape [H, D] or [B, H, D], got | validation | error | action-inference, shape-validation, numpy |
| `callback_on_step_end_tensor_inputs` has to be in | validation | error | glm-image, callback, tensor-inputs, input-validation, valueerror |
| --enable-two-batch-overlap is not supported with DSA… | validation | error | tbo, dsa, deepseek, flag-conflict |
| Expected base64-encoded bytes | validation | error | serialization, base64, msgspec |
| Invalid quantization choice | exception | error | modelopt, quantization, invalid-argument, config-validation, sglang |
| MiniMax H3 video generation produced | exception | error | minimax-h3, output-count, validation, video-generation |
| MLX auxiliary-state radix cache does not support… | error_code | error | mlx, mamba, unsupported-feature, server-args |
| {output_batch.error} | exception | critical | scheduler, inference-failure, diffusion, runtime-error |
| SANA-WM streaming decode requires… | validation | error | sana-wm, streaming, vae, component-paths, ltx2 |
| structure_info not used for JSON schema constraints | exception | error | function-call, not-implemented, json-schema, parser |
| The arguments disaggregation-decode-enable-offload-kvcache… | validation | error | sglang, pd-disaggregation, retraction, kv-offload, mutually-exclusive |
| Unsupported activation | exception | error | activation, config-validation, model-loading, kimi |
| video block token counts and timestamps must align | validation | error | minimax-h3, ref2va, video-presentation, validation |
| Failed to load serve backend | exception | critical | sglang, cli, entry-points, plugin-load-failure, import-error |
| gguf package does not provide the DeepSeek name map | panic | error | gguf, deepseek, version-mismatch, dependency |
| HiCache does not support Inkling MTP draft state yet. | exception | error | sglang, speculative-decoding, hicache, mtp, not-implemented |
| NCCL only supports CUDA, ROCm and MUSA backends. | validation | error | nccl, pytorch, cuda, backend-not-supported, distributed, sglang |
| Backward pass is not implemented yet and we do not have… | exception | error | pytorch, autograd, backward, linear-attention, not-implemented |
| Cosmos3 action request produced no action tensor | exception | error | cosmos3, action-generation, runtime, internal-error |
| fused IPM kernel needs | validation | error | shared-memory, capacity, lplb, cuda-kernel |
| indices must be on q's device | validation | error | cuda, device-mismatch, sparse-attention, sglang |
| MiniMax H3 --model-variant and --model-subfolder select… | validation | error | minimax-h3, model-variant, model-subfolder, conflicting-config, sglang |
| Pi05 noise must have shape | validation | error | pi05, vla, noise, shape-validation, flow-matching, torch |
| requires #senders == #recvs but got #senders= | exception | error | distributed, collective, grafter |
| SplitKV partial output (mO) must be Float32 | validation | error | cuda, dtype, flash-attention, cutlass, split-kv |
| Unknown serve backend | validation | error | sglang, cli, entry-points, plugin-registry, invalid-argument |
| Expected hidden_size to be at least | validation | error | layernorm, variance-override, shape-validation, gdn |
| Expected scalar scale for fused-in-checkpoint merged-column… | validation | error | weight-loading, per-tensor-scale, quantization, fused-checkpoint, shard-id |
| exponential_shift enabled but exponential_shift_mu is… | exception | error | scheduler, flow-matching, missing-parameter, exponential-shift |
| Failed to find available port after | exception | error | network, port-allocation, server-startup, configuration |
| GGUF BF16 payload does not have a byte-pair layout | validation | error | gguf, bf16, data-layout, weights |
| --hicache-host-memory-mode buffer_only does not support… | validation | error | hicache, buffer-only, deepseek-v4, sidecar-pool, configuration |
| Integrity check failed | exception | critical | integrity, checksum, model-files |
| Invalid messages at | exception | error | deepseek, tool-calls, conversation-history |
| kv-canary: scatter_req_token_ids bs+1= | validation | error | kv-cache, batch-size-limit, triton, capacity |
| MiniMax H3 AdaLN cache is only compatible with unquantized… | validation | error | adaln, quantization, checkpoint, config-conflict |
| PR # revert is not registered; available | exception | error | env-var, configuration, debug-utils, sglang |
| SP-sharded LTX-2 TI2V expected raw seq_len divisible by… | exception | critical | ltx-2, sequence-parallel, video-generation, shape-validation |
| sparse_mla_q8kv8_prefill_fwd requires h_q padded to a… | validation | error | shape-validation, tensor-parallel, sparse-attention |
| The argument disaggregation-decode-enable-offload-kvcache… | validation | error | sglang, pd-disaggregation, kv-offload, hicache, missing-backend |
| top_p must be in | validation | error | sampling-params, top-p, nucleus-sampling, validation, sglang |
| expected 2D positions, got shape= . | exception | error | sglang, speculative-decoding, dflash, internal-invariant, shape-mismatch |
| Invalid Pi05 precision | validation | error | pi05, dtype, precision, config |
| invalid Rust extension build mode | validation | error | rust-extension, build-mode, env-var, validation, sglang |
| SGLANG_USE_MLX requires stable Torch 2.13.x and MLX >=… | exception | error | mlx, apple-silicon, dependency-missing, mps, sglang |
| `ssm_state_indices` must be 1D for packed decode | exception | error | fla, fused-recurrent, state-cache, shape-validation |
| c2ws_plucker_emb shape must match hidden_states shape, got | exception | error | shape-validation, camera-conditioning, plucker, lingbot |
| Cannot find CUTLASS headers required for JIT compilation… | exception | error | cutlass, jit-kernel, missing-dependency, flashinfer, deep-gemm, build |
| Cannot msgpack encode object of type | exception | error | serialization, msgpack, sglang |
| .start_time_seconds is only allowed for video or… | validation | error | minimax-h3, request-validation, video, conditions |
| Currently, only group size 128 and -1 (channelwise) is… | validation | error | marlin, gptq, awq, group-size, config-validation |
| {detail} | http | error | api, http-400, request-validation, error-wrapper |
| EBNF is not officially supported by OpenAI endpoints… | console | info | ebnf, openai, grammar-constraint, warning, backend |
| Kimi-K3 DCP with decode_attention_backend='cutedsl_mla'… | exception | error | kimi-k3, flashinfer, version-mismatch, dcp, sglang |
| kv-canary: VerifyPlan verify_capacity must be positive, got | validation | error | kv-canary, argument-validation, capacity |
| match text not found in source:\n | exception | error | patching, text-matching, source-patcher, sglang |
| memory_video_len must be a multiple of latent_height *… | validation | error | joy-echo, memory, rope, alignment, video-tokens |
| Missing ModelSlim MoE quantization description for layer | validation | error | modelslim, quantization, moe, missing-keys, config |
| ModelOpt export functionality is not available. Please… | exception | error | modelopt, export, version-mismatch, import-error, sglang |
| output_format must be 'list' or 'numpy' | validation | error | action-endpoint, output-format, enum-validation |
| Ring Attention requires a backend whose kernel exposes the… | error_code | critical | attention, ring-parallelism, backend-support, lse, initialization |
| SGLANG_USE_MLX requires stable Torch 2.13.x and MLX >=… | exception | error | mlx, version-mismatch, torch, mps, sglang |
| Short-conv hybrid models (ZAYA1 CCA, LFM2 / LFM2-MoE) are… | validation | critical | npu, ascend, short-conv, lfm2, zaya1, attention-backend, not-implemented, sglang |
| SP DMD renoise requires packed video… | validation | error | joyecho, sequence-parallel, dmd, latent-shape, validation |
| This use case is not supported. For OpenAI chat models… | exception | error | frontend, openai, chat-model, streaming, program-structure, sglang |
| CUDA VMM feature transport requires each feature field to… | exception | error | cuda, vmm, multimodal, validation, type-error |
| DeepSeekV4 only supports interleave CP strategy, got | validation | error | deepseek, context-parallel, sglang, config-validation |
| Error while loading data | exception | error | multimodal, loading, http, input-validation |
| Failed to get server info. | error_code | error | http, server-info, startup, network |
| GLM-Image AR batch returned an unexpected response: expected | exception | error | glm-image, external-server, batch-mismatch, response-validation, runtimeerror |
| Invalid write range: local= | error_code | critical | kv-cache, index-out-of-range, assertion, chunking |
| Invalid threshold_type | validation | error | validation, value-error, moba, attention, config |
| kv-canary: launch_canary_plan_kernels requires… | validation | error | kv-canary, sliding-window, argument-validation |
| module has no attribute | exception | error | import, lazy-loading, attribute-error, diffusion |
| No frames were recorded | error_code | error | ngram, ffi, buffer-size, cuda-kernel, tensor-shape |
| No kernels registered for op | exception | error | kernels, registry, invalid-argument, sglang |
| Qwen-VL position_ids do not match the attention input | exception | error | qwen-vl, rope, position-ids, shape-mismatch, multimodal |
| The arguments enable-hierarchical-cache and… | validation | error | sglang, hicache, radix-cache, mutually-exclusive, config-validation |
| Unexpected initial_state_source shape | validation | error | gdn, linear-attention, shape-validation, cutedsl |
| Unsupported activation | validation | error | moe, activation, fallback-kernel, validation |
| Detected some but not all shards of | validation | error | modelslim, quantization, fused-layers, mixed-precision, config |
| DFLASH mask_token must be a non-empty string, got | validation | error | sglang, speculative-decoding, dflash, config-validation, tokenizer |
| flashinfer_sparse_mla supports only GLM DSA with FP8 KV… | validation | error | attention-backend, config-validation, sm120, fp8, glm |
| HiSparse supports DSA | validation | error | hisparse, dsa, attention-backend, kv-cache-dtype, sglang |
| Hunyuan3D SD2.1 UNet requires four channel stages. | validation | error | stable-diffusion, unet, config-validation, hunyuan3d |
| Inkling shared-sink LoRA rank dimensions do not match | validation | error | sglang, lora, rank-validation, shape-validation, inkling |
| Layer-sharded MLA HiCache backup with page_first layout… | validation | error | sglang, mla, hicache, jit-kernel, sgl-kernel, build |
| max_seqlen should be prepared for vision flashinfer_cudnn… | exception | error | sglang, vision-transformer, flashinfer-cudnn, kwargs-validation, max-seqlen |
| projection_cls = , not implemented | validation | error | phi4, multimodal, projection, not-implemented, config-validation |
| return_flat_raw_top_logprobs requires rectangular top… | validation | error | sglang, logprobs, data-shape, input-validation |
| Stage-1 latent has frames but sink_size= . | validation | error | sana-wm, refiner, latent-shape, sink-frame, valueerror |
| tool call function name must be a string | validation | error | inkling, tool-calls, type-validation, message-rendering |