ErrLookup › sgl-project/sglang

sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models. · Python · 3,697 source files

Analyzed at 7ed29eba80 on 2026-09-03. 3197 documented errors.

Code / MessageTypeSeverityTags
Kimi-K3 deferred feature length does not match image grids
exception error kimi-k3, multimodal, shape-mismatch, vision
`d_cache` must have shape [slots, HV, L, V].
exception error helion, kda, replayssm, tensor-shape, cache-layout
Unknown image_vae_encoding_position
validation error config-validation, pipeline, multimodal, typo
Destination MLA KV descriptors do not match prefill pp…
validation critical disaggregation, prefill-decode, pipeline-parallel, mla, kv-cache
is not officially supported by cache-dit. Supported…
validation error cache-dit, dit, block-adapter, unsupported-model, valueerror
not found next to the native Metal extension at
exception error metal, macos, installation, import-error, sgl-kernel
CUDART error
error_code critical cuda, gpu, out-of-memory, invalid-device, runtime
cuMulticastGetGranularity failed for FlashInfer workspace…
error_code error cuda-driver, multicast, flashinfer, workspace-preflight, nvswitch
Grid dim ( ) not found in
validation error multimodal, preprocessor, grid-metadata, kimi, validation
MXFP8 fused prologue requires interleaved K/V scale buffers…
exception error mxfp8, scale-buffers, tensor-shape, kv-cache, inkling
num_tokens ( ) exceeds num_max_dispatch_tokens_per_rank (…
exception critical moe, all-to-all, token-dispatcher, pplx, batch-size, capacity
--speculative-use-rejection-sampling is only supported for…
exception error speculative-decoding, rejection-sampling, server-args, eagle
GLM-Image AR returned too few output_ids: got
exception critical glm-image, autoregressive, token-length, image-generation, runtimeerror
In ps_version 'v1', the height and width have not been…
console warning internvl, vision-encoder, pixel-shuffle, multimodal, deprecation
pairs must be a torch.Tensor of shape [N, 2]
validation error pytorch, tensor-shape, flow-matching, scheduler, validation
Unsupported integer dtype
exception error kda, cutlass, dtype, index-tensor, prefill
Nemotron 3.5 DFLASH draft requires its checkpoint embedding.
exception critical sglang, speculative-decoding, dflash, nemotron, checkpoint, embedding
CUDA error ( )
panic critical cuda, cuda-graph, gpu, driver
trtllm_mla does not forward the cyclic DCP metadata to its…
exception error sglang, attention-backend, context-parallelism, not-implemented, mla, speculative-decoding
`k_cache` must have shape [slots, H, L, K].
exception error helion, kda, replayssm, tensor-shape, cache-layout
MXFP8 fused prologue requires head_dim-aligned Q/K/V.
exception error mxfp8, quantization, head-dim, attention-prologue, inkling
MiniMax H3 = must be divisible by TP size .
validation critical tensor-parallel, divisibility, config, minimax-h3, launch
unknown qk_norm: . Should be one of None, 'layer_norm'…
validation error config, validation, diffusion, qk-norm
next_token_logits row count mismatch. Expected
validation critical sglang, speculative-decoding, dflash, shape-mismatch, tensor-validation
Pi05 action fallback must run on the action root
error_code error pi05, distributed, action-root, sequence-parallel
Expected packed image latents [B, S0, D].
validation error ltx-2, image-encoding, latent-shape, packing, torch
[internvl] Cannot process raw images/videos with…
validation error multimodal, internvl, pre-tokenized-input, dynamic-tiling
The number of initial states is expected to be equal to the…
validation error fla, gated-delta-rule, varlen, batch-mismatch, shape-validation
Comfy W4A4 layer has unsupported convrot_groupsize=
validation error quantization, comfy, w4a4, checkpoint-validation, safetensors
MXFP8 KV cache requires K and V scale tensors.
validation critical mxfp8, kv-cache, quantization, scale-tensors, sglang
BatchedDecodeContext requires full_kv_pool_index_by_layer…
validation error mlx, speculative-decode, aot-kernel, kv-cache, config-validation
Requantization into is not supported, from the original…
exception error quantization, quark, requantization, not-implemented, checkpoint, python
HiRadixCache only supports MHA, MLA, DSA, and MSA models
validation error hicache, hiradixcache, unsupported-architecture, model-support
Missing NIXL destination KV memory kind
exception error nixl, disaggregation, internal-invariant, transfer-worker
NIXL PD transfer does not support HiSparse combined with…
exception error nixl, disaggregation, speculative-decoding, hicache, pd-disaggregation
--disaggregation-decode-retraction-backup=host_pool…
validation error sglang, pd-disaggregation, priority-scheduling, preemption, retraction
`g_cache` must have shape [slots, HV, L, K].
exception error helion, kda, replayssm, tensor-shape, gate-cache
--speculative-use-rejection-sampling is incompatible with…
validation error speculative-decoding, rejection-sampling, accept-threshold, server-args
Short read for
error_code error hicache, file-io, cache-corruption, truncated-file
GPTQ act_order on XPU requires each group_size block of…
exception error xpu, gptq, quantization, tensor-parallel, act-order, not-implemented
--speculative-use-rejection-sampling with multi-layer EAGLE…
validation error speculative-decoding, rejection-sampling, multi-layer-eagle, server-args
MiniMaxH3VisualEncodingStage cannot encode material chains
validation error minimax-h3, visual-encoding, unsupported-feature, validation
InklingMultimodalProcessor: required config field
validation critical multimodal, model-config, inkling, missing-config-field
--disagg-server-addr is required for --disagg-role
validation error disaggregated-serving, cli-args, missing-argument, validation
The hpc_ops attention backend does not support logit cap.
validation error attention-backend, hpc-ops, logit-cap, unsupported-feature, sglang
NIXL memory registration failed for
exception critical nixl, memory-registration, startup, rdma, disaggregation
draft sampler set but the draft forward has no…
exception critical sglang, speculative-decoding, cuda-graph, draft-sampler, hidden-states, runtime
MXFP8 fused prologue requires K/V scale buffers.
exception error mxfp8, quantization, scale-buffers, attention-prologue, inkling
GGUF tensor declares original shape , which contains…
validation error gguf, checkpoint, tensor-shape, validation, diffusion
KDA cutedsl: safe_gate (lower_bound) not yet supported
exception error kda, cutedsl, safe-gate, not-implemented, linear-attention
Pi05 weight load failed
error_code critical pi05, weight-loading, checkpoint, state-dict
{error_msg}
http error encoder, pipeline, error-wrapper, epd, multimodal
LONG GARBAGE COLLECTION DETECTED | Generation
console warning gc, latency-jitter, performance, scheduler, python
Quanto tensor/map prefixes do not match: missing metadata=
validation error quantization, quanto, checkpoint, int8, multimodal
Pi05 action state broadcast returned None
error_code error pi05, distributed, broadcast, sequence-parallel, nccl
sgl_kernel.metal is importable, but the native Metal…
error_code critical mlx, metal, aot-kernel, apple-silicon, install
Pi05 action state is missing on single-rank run
error_code error pi05, distributed, sequence-parallel, runtime-state
Shared-sink LoRA pool shape changed after initialization…
exception error sglang, lora, pool-allocation, shape-validation, runtime-error, inkling
must be a mapping or expose to_dict(), got
validation error quantization, type-validation, checkpoint, metadata
Unsupported fused vision RoPE inputs: q=
validation error vision-rope, triton, gpu-capability, tensor-validation
tool_choice 'required' or a named tool cannot be combined…
validation error openai-api, tool-calling, structured-output, constrained-decoding, sglang
cuMemGetAllocationGranularity failed for FlashInfer…
error_code error cuda-driver, flashinfer, workspace-preflight, granularity, driver-version
material URI base64 payload must be ASCII
validation error minimax-h3, base64, ascii, material-io
Mooncake's batch register requires a newer version of…
exception error mooncake, version-mismatch, batch-register, upgrade-required, sglang
nframes should in interval
validation error multimodal, video-preprocessing, ernie-4-5-vl, frame-sampling
{selection_error}{component_suffix}
validation critical attention-backend, config, multimodal, sglang
NoOpMHATokenToKVPool.set_kv_buffer was called. This pool is…
exception critical kv-cache, attention-backend, embedding, prefill-only
candidates and next_token_logits must be on the same…
validation error sglang, speculative-decoding, dflash, cuda, device-mismatch
`force_flush` must be on the same device as the inputs.
exception error helion, kda, replayssm, device-placement, force-flush
Given arguments mismatch the SGL function signature
validation error sglang, batch, arguments, signature-validation
kv-canary: must be positive, got
exception error kv-canary, validation, capacities, value-error
[internvl][qwen] image_data provided but no images parsed…
validation error multimodal, internvl, qwen, placeholder-mismatch
Petit is not installed. Please install it with `pip install…
exception error quantization, nvfp4, petit, missing-dependency, python
A GGUF encoder checkpoint cannot be combined with a second…
validation error gguf, quantization, text-encoder, multimodal, checkpoint
decoder_stage_channels
validation error model-config, channel-dimensions, config-validation, ltx-2
Image aspect ratio must be smaller than 200
validation error dots3, vision, image-preprocessing, aspect-ratio, multimodal
The hpc_ops attention backend with an fp8_e4m3 KV cache…
exception error hpc-ops, fp8, kv-cache-dtype, hunyuan, attention-backend, sglang
action_mode= requires --raw-action-dim.
validation error cosmos3, action-conditioning, server-args, configuration
FlashInfer allreduce fusion mnnvl backend requires a…
validation error flashinfer, allreduce-fusion, gpu-architecture, mnnvl, backend-selection
[ ] has entries; expected (per-prompt) or (per-sample).
validation error batch-validation, conditioning, multimodal
tool_choice references tool
http error anthropic, tool-choice, tools, validation, request-conversion
DeepGEMM Kernels compilation timeout.\n\nFeel free and…
exception critical sglang, deep-gemm, timeout, server-startup, compilation
Dual transformers for
validation error cache-dit, dual-transformer, model-introspection, attribute-mismatch
--linear-attn-verify-backend flashinfer on SM100+ requires…
validation error sglang, linear-attention, flashinfer, blackwell, dtype-validation
queued MiniMax H3 jobs require pre-queue resolved_v2…
validation critical minimax-h3, queue-invariant, geometry, internal
Inkling relative attention requires the vendored FA4 CUTE…
exception critical sglang, import-error, flashattention, cute, cuda, inkling
File does not exist.
validation error config, mooncake, ib-devices, file-not-found
InklingBatchDenseMLPWithLoRA is ineligible
validation error sglang, lora, triton-backend, bf16, eligibility-check, inkling
keyframe visual preparation requires one or two ordered…
validation error minimax-h3, keyframes, fl2va, payload-validation
unsupported time format
validation error multimodal, config-validation, video-qa, dots-note-omni
Invalid request body
http error http-400, request-validation, video-generation, openai-api
--speculative-use-rejection-sampling requires…
validation error speculative-decoding, rejection-sampling, eagle-topk, server-args
Dual-transformer cache-dit is only supported for
validation error cache-dit, dual-transformer, model-name-registry, valueerror
{e}
http error http-400, sampling-params, video-generation, out-of-range
Please install mooncake by following the instructions at…
exception critical mooncake, dependency, import-error, kv-cache-transfer, sglang
When using multiple prompts with multiple input images…
validation error multimodal, input-validation, image-input, diffusion
CUDA VMM pool has no occupied slice at control offset
exception error cuda, vmm, double-free, memory-pool
hybrid ref2va layout only supports first/last keyframe…
validation error validation, keyframe, video, minimax-h3
kv-canary: must be on 's device , got
validation error kv-cache, device-mismatch, torch, validation
MXFP8 fused prologue requires contiguous interleaved…
exception error mxfp8, scale-buffers, contiguity, kv-cache, inkling