ErrLookup › sgl-project/sglang
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models. · Python · 3,697 source files
Analyzed at 7ed29eba80 on 2026-09-03. 3197 documented errors.
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| Kimi-K3 deferred feature length does not match image grids | exception | error | kimi-k3, multimodal, shape-mismatch, vision |
| `d_cache` must have shape [slots, HV, L, V]. | exception | error | helion, kda, replayssm, tensor-shape, cache-layout |
| Unknown image_vae_encoding_position | validation | error | config-validation, pipeline, multimodal, typo |
| Destination MLA KV descriptors do not match prefill pp… | validation | critical | disaggregation, prefill-decode, pipeline-parallel, mla, kv-cache |
| is not officially supported by cache-dit. Supported… | validation | error | cache-dit, dit, block-adapter, unsupported-model, valueerror |
| not found next to the native Metal extension at | exception | error | metal, macos, installation, import-error, sgl-kernel |
| CUDART error | error_code | critical | cuda, gpu, out-of-memory, invalid-device, runtime |
| cuMulticastGetGranularity failed for FlashInfer workspace… | error_code | error | cuda-driver, multicast, flashinfer, workspace-preflight, nvswitch |
| Grid dim ( ) not found in | validation | error | multimodal, preprocessor, grid-metadata, kimi, validation |
| MXFP8 fused prologue requires interleaved K/V scale buffers… | exception | error | mxfp8, scale-buffers, tensor-shape, kv-cache, inkling |
| num_tokens ( ) exceeds num_max_dispatch_tokens_per_rank (… | exception | critical | moe, all-to-all, token-dispatcher, pplx, batch-size, capacity |
| --speculative-use-rejection-sampling is only supported for… | exception | error | speculative-decoding, rejection-sampling, server-args, eagle |
| GLM-Image AR returned too few output_ids: got | exception | critical | glm-image, autoregressive, token-length, image-generation, runtimeerror |
| In ps_version 'v1', the height and width have not been… | console | warning | internvl, vision-encoder, pixel-shuffle, multimodal, deprecation |
| pairs must be a torch.Tensor of shape [N, 2] | validation | error | pytorch, tensor-shape, flow-matching, scheduler, validation |
| Unsupported integer dtype | exception | error | kda, cutlass, dtype, index-tensor, prefill |
| Nemotron 3.5 DFLASH draft requires its checkpoint embedding. | exception | critical | sglang, speculative-decoding, dflash, nemotron, checkpoint, embedding |
| CUDA error ( ) | panic | critical | cuda, cuda-graph, gpu, driver |
| trtllm_mla does not forward the cyclic DCP metadata to its… | exception | error | sglang, attention-backend, context-parallelism, not-implemented, mla, speculative-decoding |
| `k_cache` must have shape [slots, H, L, K]. | exception | error | helion, kda, replayssm, tensor-shape, cache-layout |
| MXFP8 fused prologue requires head_dim-aligned Q/K/V. | exception | error | mxfp8, quantization, head-dim, attention-prologue, inkling |
| MiniMax H3 = must be divisible by TP size . | validation | critical | tensor-parallel, divisibility, config, minimax-h3, launch |
| unknown qk_norm: . Should be one of None, 'layer_norm'… | validation | error | config, validation, diffusion, qk-norm |
| next_token_logits row count mismatch. Expected | validation | critical | sglang, speculative-decoding, dflash, shape-mismatch, tensor-validation |
| Pi05 action fallback must run on the action root | error_code | error | pi05, distributed, action-root, sequence-parallel |
| Expected packed image latents [B, S0, D]. | validation | error | ltx-2, image-encoding, latent-shape, packing, torch |
| [internvl] Cannot process raw images/videos with… | validation | error | multimodal, internvl, pre-tokenized-input, dynamic-tiling |
| The number of initial states is expected to be equal to the… | validation | error | fla, gated-delta-rule, varlen, batch-mismatch, shape-validation |
| Comfy W4A4 layer has unsupported convrot_groupsize= | validation | error | quantization, comfy, w4a4, checkpoint-validation, safetensors |
| MXFP8 KV cache requires K and V scale tensors. | validation | critical | mxfp8, kv-cache, quantization, scale-tensors, sglang |
| BatchedDecodeContext requires full_kv_pool_index_by_layer… | validation | error | mlx, speculative-decode, aot-kernel, kv-cache, config-validation |
| Requantization into is not supported, from the original… | exception | error | quantization, quark, requantization, not-implemented, checkpoint, python |
| HiRadixCache only supports MHA, MLA, DSA, and MSA models | validation | error | hicache, hiradixcache, unsupported-architecture, model-support |
| Missing NIXL destination KV memory kind | exception | error | nixl, disaggregation, internal-invariant, transfer-worker |
| NIXL PD transfer does not support HiSparse combined with… | exception | error | nixl, disaggregation, speculative-decoding, hicache, pd-disaggregation |
| --disaggregation-decode-retraction-backup=host_pool… | validation | error | sglang, pd-disaggregation, priority-scheduling, preemption, retraction |
| `g_cache` must have shape [slots, HV, L, K]. | exception | error | helion, kda, replayssm, tensor-shape, gate-cache |
| --speculative-use-rejection-sampling is incompatible with… | validation | error | speculative-decoding, rejection-sampling, accept-threshold, server-args |
| Short read for | error_code | error | hicache, file-io, cache-corruption, truncated-file |
| GPTQ act_order on XPU requires each group_size block of… | exception | error | xpu, gptq, quantization, tensor-parallel, act-order, not-implemented |
| --speculative-use-rejection-sampling with multi-layer EAGLE… | validation | error | speculative-decoding, rejection-sampling, multi-layer-eagle, server-args |
| MiniMaxH3VisualEncodingStage cannot encode material chains | validation | error | minimax-h3, visual-encoding, unsupported-feature, validation |
| InklingMultimodalProcessor: required config field | validation | critical | multimodal, model-config, inkling, missing-config-field |
| --disagg-server-addr is required for --disagg-role | validation | error | disaggregated-serving, cli-args, missing-argument, validation |
| The hpc_ops attention backend does not support logit cap. | validation | error | attention-backend, hpc-ops, logit-cap, unsupported-feature, sglang |
| NIXL memory registration failed for | exception | critical | nixl, memory-registration, startup, rdma, disaggregation |
| draft sampler set but the draft forward has no… | exception | critical | sglang, speculative-decoding, cuda-graph, draft-sampler, hidden-states, runtime |
| MXFP8 fused prologue requires K/V scale buffers. | exception | error | mxfp8, quantization, scale-buffers, attention-prologue, inkling |
| GGUF tensor declares original shape , which contains… | validation | error | gguf, checkpoint, tensor-shape, validation, diffusion |
| KDA cutedsl: safe_gate (lower_bound) not yet supported | exception | error | kda, cutedsl, safe-gate, not-implemented, linear-attention |
| Pi05 weight load failed | error_code | critical | pi05, weight-loading, checkpoint, state-dict |
| {error_msg} | http | error | encoder, pipeline, error-wrapper, epd, multimodal |
| LONG GARBAGE COLLECTION DETECTED | Generation | console | warning | gc, latency-jitter, performance, scheduler, python |
| Quanto tensor/map prefixes do not match: missing metadata= | validation | error | quantization, quanto, checkpoint, int8, multimodal |
| Pi05 action state broadcast returned None | error_code | error | pi05, distributed, broadcast, sequence-parallel, nccl |
| sgl_kernel.metal is importable, but the native Metal… | error_code | critical | mlx, metal, aot-kernel, apple-silicon, install |
| Pi05 action state is missing on single-rank run | error_code | error | pi05, distributed, sequence-parallel, runtime-state |
| Shared-sink LoRA pool shape changed after initialization… | exception | error | sglang, lora, pool-allocation, shape-validation, runtime-error, inkling |
| must be a mapping or expose to_dict(), got | validation | error | quantization, type-validation, checkpoint, metadata |
| Unsupported fused vision RoPE inputs: q= | validation | error | vision-rope, triton, gpu-capability, tensor-validation |
| tool_choice 'required' or a named tool cannot be combined… | validation | error | openai-api, tool-calling, structured-output, constrained-decoding, sglang |
| cuMemGetAllocationGranularity failed for FlashInfer… | error_code | error | cuda-driver, flashinfer, workspace-preflight, granularity, driver-version |
| material URI base64 payload must be ASCII | validation | error | minimax-h3, base64, ascii, material-io |
| Mooncake's batch register requires a newer version of… | exception | error | mooncake, version-mismatch, batch-register, upgrade-required, sglang |
| nframes should in interval | validation | error | multimodal, video-preprocessing, ernie-4-5-vl, frame-sampling |
| {selection_error}{component_suffix} | validation | critical | attention-backend, config, multimodal, sglang |
| NoOpMHATokenToKVPool.set_kv_buffer was called. This pool is… | exception | critical | kv-cache, attention-backend, embedding, prefill-only |
| candidates and next_token_logits must be on the same… | validation | error | sglang, speculative-decoding, dflash, cuda, device-mismatch |
| `force_flush` must be on the same device as the inputs. | exception | error | helion, kda, replayssm, device-placement, force-flush |
| Given arguments mismatch the SGL function signature | validation | error | sglang, batch, arguments, signature-validation |
| kv-canary: must be positive, got | exception | error | kv-canary, validation, capacities, value-error |
| [internvl][qwen] image_data provided but no images parsed… | validation | error | multimodal, internvl, qwen, placeholder-mismatch |
| Petit is not installed. Please install it with `pip install… | exception | error | quantization, nvfp4, petit, missing-dependency, python |
| A GGUF encoder checkpoint cannot be combined with a second… | validation | error | gguf, quantization, text-encoder, multimodal, checkpoint |
| decoder_stage_channels | validation | error | model-config, channel-dimensions, config-validation, ltx-2 |
| Image aspect ratio must be smaller than 200 | validation | error | dots3, vision, image-preprocessing, aspect-ratio, multimodal |
| The hpc_ops attention backend with an fp8_e4m3 KV cache… | exception | error | hpc-ops, fp8, kv-cache-dtype, hunyuan, attention-backend, sglang |
| action_mode= requires --raw-action-dim. | validation | error | cosmos3, action-conditioning, server-args, configuration |
| FlashInfer allreduce fusion mnnvl backend requires a… | validation | error | flashinfer, allreduce-fusion, gpu-architecture, mnnvl, backend-selection |
| [ ] has entries; expected (per-prompt) or (per-sample). | validation | error | batch-validation, conditioning, multimodal |
| tool_choice references tool | http | error | anthropic, tool-choice, tools, validation, request-conversion |
| DeepGEMM Kernels compilation timeout.\n\nFeel free and… | exception | critical | sglang, deep-gemm, timeout, server-startup, compilation |
| Dual transformers for | validation | error | cache-dit, dual-transformer, model-introspection, attribute-mismatch |
| --linear-attn-verify-backend flashinfer on SM100+ requires… | validation | error | sglang, linear-attention, flashinfer, blackwell, dtype-validation |
| queued MiniMax H3 jobs require pre-queue resolved_v2… | validation | critical | minimax-h3, queue-invariant, geometry, internal |
| Inkling relative attention requires the vendored FA4 CUTE… | exception | critical | sglang, import-error, flashattention, cute, cuda, inkling |
| File does not exist. | validation | error | config, mooncake, ib-devices, file-not-found |
| InklingBatchDenseMLPWithLoRA is ineligible | validation | error | sglang, lora, triton-backend, bf16, eligibility-check, inkling |
| keyframe visual preparation requires one or two ordered… | validation | error | minimax-h3, keyframes, fl2va, payload-validation |
| unsupported time format | validation | error | multimodal, config-validation, video-qa, dots-note-omni |
| Invalid request body | http | error | http-400, request-validation, video-generation, openai-api |
| --speculative-use-rejection-sampling requires… | validation | error | speculative-decoding, rejection-sampling, eagle-topk, server-args |
| Dual-transformer cache-dit is only supported for | validation | error | cache-dit, dual-transformer, model-name-registry, valueerror |
| {e} | http | error | http-400, sampling-params, video-generation, out-of-range |
| Please install mooncake by following the instructions at… | exception | critical | mooncake, dependency, import-error, kv-cache-transfer, sglang |
| When using multiple prompts with multiple input images… | validation | error | multimodal, input-validation, image-input, diffusion |
| CUDA VMM pool has no occupied slice at control offset | exception | error | cuda, vmm, double-free, memory-pool |
| hybrid ref2va layout only supports first/last keyframe… | validation | error | validation, keyframe, video, minimax-h3 |
| kv-canary: must be on 's device , got | validation | error | kv-cache, device-mismatch, torch, validation |
| MXFP8 fused prologue requires contiguous interleaved… | exception | error | mxfp8, scale-buffers, contiguity, kv-cache, inkling |