sgl-project/sglang

Documented errors, page 5 of 32. Back to sgl-project/sglang

Code / MessageTypeSeverityTags
prefetch_threshold must be int, got
validation error hicache, config-validation, prefetch, type-error
ResolvedPlan requires one or two ordered image/keyframe…
validation error minimax-h3, keyframe, frame-index, plan-validation
time must have shape [batch]
validation error pi05, tensor-rank, input-validation
Unsupported dots note omni video_config fields
validation error multimodal, video-config, unknown-fields, valueerror
Unsupported preprocessed video item
validation error multimodal, video, schema-mismatch, version-drift
DeepEP v2 MoE is not validated as a speculative draft…
validation error speculative-decoding, deepep, moe, draft-model, server-args
DSPARK aux hidden capture requires PP=1.
exception error dspark, pipeline-parallel, speculative-decoding, kimi
Failed to render chat template for embedding input
validation error jinja, embeddings, chat-template, type-error
`g_cache` must have dtype torch.float32.
exception error helion, kda, replayssm, dtype, gate-cache
generated MiniMax H3 MP4 size does not match the resolved…
exception error minimax-h3, resolution, geometry-mismatch, validation
Inkling shared-sink LoRA outer factors must have expert…
validation error sglang, lora, shape-validation, moe, shared-experts, inkling
Model architectures failed to be inspected. Please check…
validation error registry, model-resolution, architecture-unsupported
RaggedVerifyLayout requires at least one request
validation error speculative-decoding, validation, batch-empty
refiner cu_seqlens live text length must be in
validation error minimax-h3, dit, shape-validation, text-embeddings
Unknown MLX async mode
validation error mlx, async-scheduler, internal-contract, sglang
Weight input_size_per_partition =
validation error marlin, gptq, tensor-parallel, shape-validation
Cannot resolve the consumer rank before parallel state…
exception critical distributed, parallel-state, cuda-ipc, initialization-order
Comfy INT8 layer needs I8 weights and F32 scales, got and
validation error quantization, int8, dtype-mismatch, checkpoint-validation
Cosmos3 forward_dynamics produces video; use /v1/videos…
validation error cosmos3, action-mode, endpoint-routing, video-generation
FlashAttention combine kernel cannot be implemented with…
error_code critical cuda, attention, split-kv, kernel-config
No frames before start_time
validation error video, timestamps, segment, out-of-range
RPC error on
exception error rpc, debug-utils, remote-error, sglang
The argument disaggregation-decode-enable-offload-kvcache…
validation error sglang, pd-disaggregation, kv-offload, decode-only, config-validation
Unsupported activation
validation error exaone, activation, model-config, unsupported-operation
Cosmos3 batched prompts must tokenize to the same length…
validation error cosmos3, multimodal-gen, batching, tokenization, validation
cuMemGetAddressRange
exception error cuda, vmm, cuda-graph, invalid-pointer
--dcp-replicate-q-proj requires --dcp-size > 1.
validation error sglang, distributed, dcp, argument-validation
--enable-tp-lm-head-all-to-all requires an available PyNCCL…
validation critical distributed, nccl, p2p, tp, lm-head
Hugging Face weight URL has no filename
validation error huggingface, weights, url-parsing
Multimodal data is corrupted or cannot be decoded
exception error multimodal, wrapper, decode, request
QwenImage text conditioning mask has shape
validation error qwen-image, shape-mismatch, mask-validation, text-embedding
[Staging] KV transfer via staging buffer failed
exception error disaggregation, staging-buffer, transfer-failure, wrapper-exception
Action endpoint requires SamplingParams or…
validation error action-endpoint, sampling-params, model-registry, subclass-validation
conflicting safetensors LoRA alpha metadata
validation error lora, safetensors, metadata, conflict
--enable-int8-mamba-checkpoint is not supported together…
validation error sglang, mamba, int8, hierarchical-cache, incompatible-flags, server-args
{failure_msg}: {error_msg}
exception error lora, scheduler, adapter-loading, runtime-error
Humming quantization for MoE only supports…
validation error moe, humming, runner-backend, config-validation
Ideogram4DenoisingStage applies its custom scheduler step
exception error ideogram, scheduler, unsupported-operation, diffusers
invalid compressor checkpoint name
validation error gguf, deepseek, compressor, name-mapping
Invalid config pair (missing '=')
validation error config, validation, debug-utils
Invalid latent H/W computed from batch.height/width
validation error ltx-2, video-generation, resolution-validation, sequence-parallelism
NPU packed attention does not support a sequence that is…
exception error npu, ascend, varlen, ring-attention, not-implemented
SANA-WM refiner text encoder must return per-layer…
exception error sana-wm, text-encoder, hidden-states, runtimeerror, model-output
top_logprobs_num exceeds disaggregation metadata capacity …
validation error sglang, disaggregation, logprobs, metadata-buffer, capacity
Unsupported quantized embedding marker for
validation critical quantization, checkpoint, embedding, model-load
You have passed a list of generators of length
validation error mova, generator, batch-size, diffusers, validation
All memory audio latents must share batch and channel…
validation error joy-echo, audio, latent, memory-bank, concat
All ranges must be within
validation error attention, varlen, mask-metadata, range-validation, multimodal
config.json claims a packaged Muse Glimmer MLX artifact
validation error checkpoint-format, packaging, weight-keys
flash-attn is not installed. Please install it, e.g., `pip…
exception critical flash-attn, import-error, dependency-missing, attention, cuda
Speculative decoding with --enable-unified-memory is only…
validation error speculative-decoding, unified-memory, hybrid-swa, kv-cache
Unknown data: . You may need to set `--data-type` if using…
exception error data-schema, polars, cli, text-comparison
Comfy NVFP4 layer needs a 2D packed weight, got
validation error quantization, nvfp4, rank-mismatch, conv-weights, checkpoint-validation
config not published; cannot read a config leaf
validation error config, runtime-context, initialization-order, sglang
Eagle3 MLA draft post_load_weights only supports float…
exception error eagle3, weight-loading, dtype, quantization, speculative-decoding
--enable-metrics requires smg-grpc-servicer ≥ 0.5.3 (the…
exception error grpc, metrics, version-mismatch, dependency, sglang
`fuse_qkv_projections()` is not supported for models having…
validation error fusion, attention, optimization, value-error
Grouped pipeline returned fewer outputs than requests.
error_code critical runtime, pipeline, inference, internal-error, multimodal
MiniMax-H3 adaln_t_table must have shape [N, D] with N >=…
validation error minimax-h3, safetensors, checkpoint, tensor-shape, validation
MiniMaxH3TextEncodingStage direct Qwen3VL encoder forward…
validation error minimax-h3, legacy-api, not-implemented, request-format
PD disaggregation for MiniMax sparse layers with index…
validation error disaggregation, minimax, sparse-attention, not-implemented
SGLANG_ENABLE_EPLB_BALANCEDNESS_METRIC is no longer…
validation error env-var, eplb, metrics, migration, server-args, deprecated
task does not allow condition role= type=
validation error minimax-h3, condition-rules, task-profile, validation
tokenizer missing required special token
exception critical tokenizer, checkpoint-mismatch, vocab, asr, initialization
unsupported input for Sana fused bias-GLU
exception error sana, diffusion, glu, channels-last, triton
Cannot call `set_default_attn_processor` when attention…
validation error attention, processor, flux2, state-error
Cannot resolve total_kv_heads: kv_args has neither…
exception error disaggregation, staging-buffer, kv-heads, missing-metadata
content hash mismatch for media_data
validation error multimodal, artifact-cache, content-hash, cache-invalidation
DFLASH sliding_attention layers require…
validation error speculative-decoding, dflash, sliding-window, hf-config
--diff-threshold with a single argument must be a float…
validation error cli, parsing, float, threshold
H3 conditioning projection
validation critical minimax-h3, conditioning-projection, shape-mismatch, checkpoint
Missing config.sliding_window for Mellum sliding_attention…
exception critical mellum, sliding-window, config-validation
Missing rope_parameters
exception critical mellum, config-validation, rope, model-loading
q shape must be [num_tokens, num_qo_heads, head_dim], got
validation error shape-validation, rope, gqa, metal, sgl-kernel
quantized tensor maps to a non-weight parameter
validation error gguf, deepseek, quantization, weight-mapping
SGLang does not recognize target_modules=
validation error lora, target-modules, peft-config, sglang
mask/position shape mismatch: vs .
exception error sglang, speculative-decoding, dflash, internal-invariant, mask-shape
DFLASH config.layer_types must be a sequence of attention…
validation error speculative-decoding, dflash, hf-config, type-validation
Language ' ' not recognized. Use full name (e.g.…
exception error whisper, audio, language-code, validation
must use component=value entries
validation error config, cli, validation, layerwise-offload
rmsnorm_hf: unsupported hidden_size=
validation error rmsnorm, shape-validation, cuda-kernel, layernorm
The batch size is expected to be 1 rather than
validation error kda, fla, varlen, batch-shape, attention
action policy returned no output
exception error action-inference, empty-output, scheduler
Cosmos3 AVAE dec_strides product must equal hop_size…
validation error avae, config, cosmos3, init, value-error
--enable-deepseek-v4-fp4-indexer requires SM100 or SM120…
validation error sglang, deepseek, fp4, gpu-architecture, deepgemm
Hf3fsClient.check
validation error hf3fs, validation, alignment, batch-io
Language ' ' is not in this Whisper model's vocabulary. The…
exception error whisper, audio, tokenizer-vocab, model-version
MiniMax H3 loaded checkpoint partition does not match…
validation error minimax-h3, checkpoint-mismatch, model-variant, integrity-check, sglang
MIXED_PRECISION layer group
exception error quantization, mixed-precision, quark, unsupported-algo, mxfp4, python
MXFP8 dense GEMM requested via…
error_code error quantization, mxfp8, gemm-backend, flashinfer, hardware-compatibility
no supported CUDA VMM allocation handle type
exception critical cuda, vmm, handle, driver-capability, environment
Reasoning parser ' ' is always-on and cannot be disabled…
validation error anthropic-api, reasoning, toggle
Unsupported Pi05 dtype
validation error pi05, dtype, checkpoint, config
byte_tensor must be a sufficiently large contiguous uint8…
validation error multimodal, cuda, tensor-validation, memory-pool
config not published for role
exception error runtime-context, initialization-order, roles, sglang
NIXL memory registration failed for aux tensors
exception critical nixl, memory-registration, startup, disaggregation
op has no backend usable on device (registered: )
validation error kernels, device-eligibility, environment, missing-dependency, sglang
PD state transfer failed: unknown state_type=
exception error disaggregation, state-transfer, dispatch, unsupported-feature
The block-wise quantization only supports dynamic…
validation error quantization, fp8, block-quant, activation-scheme
The block-wise quantization only supports fp8-serialized…
validation error quantization, fp8, block-quant, checkpoint-config