sgl-project/sglang

Documented errors, page 7 of 32. Back to sgl-project/sglang

Code / MessageTypeSeverityTags
DSV4 target and draft pools must share SWA ring geometry…
validation error disaggregation, deepseek-v4, swa, geometry-mismatch
Inkling shared-sink gate-up A and down B must use the same…
validation error sglang, lora, shape-validation, consistency-check, moe, inkling
Unknown stop criteria
exception error configuration, simulation, schedule-simulator, sglang
USPAttention masked path supports ring parallelism only for…
exception error attention, ring-parallelism, attention-mask, batch-size, fa-backend, not-implemented
VLA action expert should not share the prefix TP layout…
validation error pi05, vla, parallelism-strategy, tp, config, sglang
[weight_cache: ] quantization method is not verified for…
validation error quantization, weight-cache, ipc, config
--dcp-replicate-q-proj only applies to the a2a/fi_a2a DCP…
validation error sglang, distributed, dcp, argument-validation, incompatible-flags
DeepSeek-V4 FP4 experts require torch.float4_e2m1fn_x2…
exception critical quantization, fp4, pytorch-version, deepseek
Error in stream_executor
console error interpreter, stream-executor, async, frontend, nested-exception
KDA num_heads ( ) must be divisible by shard tp_size ( )
exception error kimi-linear, tensor-parallel, divisibility, startup-validation
routed-expert tensors were not loaded (sample: ). Expected…
exception critical laguna, moe, weight-loading, missing-weights
MiniMax H3 latent preparation requires pre-queue resolved…
validation error minimax-h3, temporal-dimensions, plan-validation
Mismatched ModelSlim quantization for W13 in layer
validation error modelslim, quantization, moe, fused-weights, config-mismatch
override_server_args: unknown ServerArgs field(s)
validation error config, override, typo, sglang
acknowledgements support one consumer or the complete…
validation error cuda-ipc, acknowledgement, protocol-validation, multimodal-transport
Server failed to start within the timeout period.
exception error frontend, timeout, server-startup, spawn, sglang
The current scheduler class
validation error qwen-image, diffusers, scheduler, timesteps, scheduler-unsupported
The original encoder only has
validation error siglip, vision-encoder, layer-override
uniform_samples_for_final_sampling shape mismatch. Expected
validation error sglang, speculative-decoding, dflash, shape-mismatch, rng
Unknown request_id
validation error disaggregation, request-state, unknown-id, stale-request
video rows must be divisible by t*h*w for latent_shape=
validation error validation, tensor-shape, unpatchify, row-count, minimax-h3
Waiting for main node timeout!
exception error sglang, deep-gemm, timeout, multi-node, compilation
Anthropic thinking is not supported for reasoning parser
validation error anthropic-api, reasoning, unsupported-feature
Comfy W4A8 layer has incompatible weight/scale shapes: , …
validation error quantization, shape-mismatch, w4a8, group-size
Cosmos3 action batch size
validation error cosmos3, batching, capacity-limit, server-args
Cosmos3 action endpoint supports action_mode='policy' or…
validation error cosmos3, action-mode, enum-validation
DeepEP v2 MoE has not validated fused shared experts yet…
validation error deepep, shared-experts, fusion, moe, server-args
Failed to load LoRA adapter
exception critical lora, startup, load-failure, sglang
Invalid LoRA merge mode
validation error lora, config-validation, diffusion-pipeline
repetition_penalty must be in (0, 2] (1.0 = no penalty), got
validation error sampling-params, repetition-penalty, validation, sglang
token_ids_logprob must be a flat list of integers.
validation error sglang, validation, logprob, request-validation
attn_res_fused_tma requires SM100+ excluding SM12x; SM
error_code error cuda, tma, sm100, architecture-check, jit
is not bundled or cached, and Rust extension build mode is…
exception error rust-extension, missing-module, build-cache, env-var, sglang
failed to move modules to
error_code critical cuda-oom, device-movement, rollback, memory-offload
hd256 forward varlen expects k rank 3 or 5, got rank
exception error cuda, flash-attention, tensor-shape, sm100, cutlass
Hunyuan3D Paint expects square latents and a matching view…
validation error runtime, shape-mismatch, multiview, diffusion
`initial_state` must be a 4D tensor
validation error kda, helion, shape-validation, tensor-ndim
No call message found for
exception error harmony, tool-calls, conversation-history, validation, sglang
plucker_emb token count
validation error sana-wm, shape-mismatch, camera-embedding, validation
Unsupported input shape
validation error fla, cumsum, shape-validation, head-first
`a`/`b` must have shape [B, HV] with HV=
exception error fla, fused-recurrent, head-mismatch, tensor-parallel
Cannot put argument inside a f-string. This is not…
validation error sglang, f-string, tracer, typeerror
dynamic_batch_seeds must be a list with one seed per prompt
validation error seed, batch-validation, input-validation
expert-pack is not identity triplet layout
exception critical moe, expert-pack, binary-format, flags, layout-mismatch
.position_ids is required
validation error input-validation, position-ids, multimodal, forward
LingBot causal sequence sharding currently requires…
exception error not-implemented, sequence-parallelism, kv-cache, lingbot
MiMoV2 fused qkv_proj checkpoint is TP=
exception critical mimo-v2, tensor-parallel, weight-loading, checkpoint-layout
Only CUDA and MUSA support GGUF quantization currently.
console warning sglang, gguf, rocm, quantization, platform-support
Triton is not supported on current platform, roll back to…
console warning triton, cuda, device-detection, cpu-fallback, fla
Unknown router
exception error cli-arguments, simulation, schedule-simulator, sglang
Unsupported config option
validation error sglang, yaml, config, argparse, unsupported-option
Unsupported model type
validation error comfyui, diffusion, model-loading, unsupported-architecture
Unsupported runner backend
exception error quantization, moe, dispatch, version-skew
A scheme must be defined for each layer
validation error modelslim, quantization, scheme-uninitialized, runtime
Cannot use the fast tokenizer in slow tokenizer mode.
validation error tokenizer, configuration, sglang
Decode out of memory. Try to lower your batch size.\nTry to…
exception critical sglang, memory, decode, paged-kv
f"Adapter weights at
exception error model-loading, state-dict, adapter, weight-mismatch
image_mode=' ' is not supported with multiple images (got…
exception error multimodal, ocr, multi-image, config-validation
Incorrect type of pixel values. Got type
validation error multimodal, vision, type-validation, deepseek-ocr
Inkling shared-sink LoRA expert count does not match
validation error sglang, lora, shape-validation, moe, shared-experts, inkling
kernel dispatch requires at least one tensor argument
validation error kernel-dispatch, api-misuse, assertion, triton
extents exceed BUMPARENA_MAX_EXTENTS ( )
validation error cuda, vmm, memory, capacity-limit
LTX-2 token latents seq_len=
validation error ltx-2, video-generation, sequence-parallelism, tensor-shape-mismatch
Neighborhood attention requires each dim to be at least its…
validation error attention, input-shape, video-generation, ltx-2
No DeepSeek-V4 checkpoint mapping for
panic error gguf, deepseek, unmapped-tensors, weight-mapping
No files found in HF repo
exception error huggingface, checksum, model-files
`out` must have shape
exception error fla, fused-recurrent, output-buffer, shape-validation
PD peers must connect matching DCP ranks, got prefill=
exception critical disaggregation, dcp, parallelism, bootstrap
requires --distilled-lora-path…
validation error ltx2, distilled-lora, missing-path, component, sglang
SGLANG_USE_MLX requires an available PyTorch MPS device
exception error mlx, mps, torch, apple-silicon, device-unavailable
Slice size exceeds destination token capacity for TP slice…
validation critical disaggregation, tensor-parallel, heterogeneous-tp, kv-cache
torch.distributed must be initialised before…
exception critical disaggregation, multi-node, torch-distributed, initialization-order
Unsupported activation type
validation error phi4, activation, glu, config-validation
/v1/models
http error radix-tree, hicache, not-implemented, mem-cache
Cosmos3 policy input requires an observation image
validation error cosmos3, policy-mode, missing-image, input-validation
--enable-linear-replayssm-spec requires a linear draft chain
validation error sglang, replayssm, speculative-decoding, eagle, config-conflict
For Fused MoE layers, only
validation error quantization, moe, mxint4, compressed-tensors, model-config
Invalid spatial patching for packed token latents. Expected…
validation error ltx-2, video-generation, resolution-validation, divisibility
JoyImage conditioning batch mismatch: hidden_states batch=
exception error batch-mismatch, cfg-conditioning, joyimage, multimodal
MiniMax H3 TP-local heads
validation critical ulysses, sequence-parallel, divisibility, tensor-parallel
Ngram speculative decoding only supports CUDA or CPU…
validation error speculative-decoding, ngram, device-support, rocm, server-args
PD decode DCP requires an MLA or hybrid-MLA KV pool.
exception critical disaggregation, dcp, context-parallel, mla, unsupported-feature
rope.inv_freq must stay fp32 after load, got
validation error dtype, fp32, rope, weight-loading
is missing duck-typed methods from SpeculativeAlgorithm: …
validation error speculative-decoding, plugin-api, duck-typing
The pointers must be multiple of 16 bytes.
validation error sglang, cuda-kernel, alignment, silu, shape-validation
Unexpected return type from apply_chat_template
error_code error cosmos3, tokenizer, apply-chat-template, transformers, type-mismatch
attention group range
validation error cuda, vmm, parallelism, world-size-mismatch
--enable-unified-memory with PD disaggregation does not…
validation error unified-memory, pd-disaggregation, hybrid-swa, kv-cache, boot-config
f"recurrent_kda state pool breaks the compiled stride…
validation error sglang, kda, alignment, memory-layout, flashinfer
Failed to load LoRA adapter
validation error lora, duplicate, adapter, sglang
Got but expected positional dim
validation error config, validation, rope, diffusion
invalid Kimi-K3 attention-residual target
validation error kimi-k3, gguf, weight-conversion, validation
Invalid mode: , must be one of 'write', 'read', 'skip
validation error sglang, glm-image, kv-cache, mode-validation, enum-value
KVTransferError(self.bootstrap_room, failure_reason)
exception critical sglang, nixl, kv-transfer, disaggregation, distributed-inference
LoRA adapter ' ' contains target modules that are not…
validation error lora, target-modules, subset-validation, sglang
Mooncake's batch transfer requires mooncake-transfer-engine…
exception error mooncake, version-mismatch, batch-transfer, upgrade-required, sglang
MXFP8 dense GEMM requested via…
error_code error quantization, mxfp8, gemm-backend, hardware-compatibility, flashinfer
predict_num_frames supports a single prediction only, got…
validation error batching, duration-head, shape-validation, ltx-2
Timeout while waiting for event
error_code error timeout, concurrency, async, meta-info
Unexpected Ascend TND softmax LSE shape: expected
exception critical ascend, npu, flash-attention, lse, shape-mismatch