sgl-project/sglang
Documented errors, page 28 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| NPU packed attention requires matching Q/K head dimensions | validation | error | npu, ascend, head-dim, shape-validation |
| raw_latent_shape must be (T, H, W) or (B, C, T, H, W) for… | validation | error | shape-validation, latent-shape, video-generation |
| SGLANG_DISAGG_STAGING_BUFFER with pp_size > 1 is only… | validation | error | sglang, disaggregation, pipeline-parallelism, mooncake, nixl, config-validation |
| Unknown device module | validation | error | device, torch, validation |
| Unsupported compute capability | validation | error | flash-attention, gpu-arch, unsupported-hardware, cuda |
| An input image is required for mesh generation | http | error | http, missing-parameter, mesh-generation, unprocessable-entity |
| audio must be a str, bytes, tuple, torch.Tensor, or… | validation | error | multimodal, audio, type-validation, mimo |
| audio_num_frames must be provided for RoPE coordinate… | exception | error | rope, audio, missing-argument, ltx-2, forward |
| batchMatch received an empty token tail | exception | error | ngram, empty-input, batch-validation |
| chunk_plucker shape mismatch for SANA-WM: expected | validation | error | sglang, sana-wm, chunk-plucker, latent-resolution, shape-mismatch |
| Cosmos3 I2V image list is empty | validation | error | cosmos3, i2v, empty-list, image-input |
| f"flash attention version | validation | error | flash-attention, version-mismatch, dispatch, config |
| FlashInfer KDA kernel (recurrent_kda) is not available… | exception | error | sglang, flashinfer, kda, gpu-compatibility, sm100 |
| has entries, but rank_index= | validation | error | umbp, config-validation, rank-mismatch |
| Inkling reasoning_effort must be finite and in [0.0, 0.99] | validation | error | inkling, reasoning-effort, range-validation |
| LSE partial tensor must have 3 or 4 dimensions… | validation | error | flash-attention, shape-mismatch, validation |
| Model directory does not contain model_index.json. Only… | validation | error | model-format, diffusers, validation |
| model_index.json for | validation | error | diffusers, config-validation |
| (num_frames - 1) must be divisible by num_frame_per_block… | validation | error | causal-denoising, frame-count, divisibility |
| Out-of-tree serve backends cannot replace reserved or… | exception | critical | rocm, allreduce, quick-allreduce, tensor-parallel |
| Passing integer indices (e.g. from `enumerate(timesteps)`)… | validation | error | scheduler, timestep-type, flow-matching |
| quality must be one of | validation | error | minimax-h3, quality-level, request-validation, sampling-params |
| Rank 0 produced no embedding for | http | critical | encoder, disaggregation, mooncake, staging |
| recent_window_tokens must be non-negative or None | validation | error | kv-cache, sliding-window, argument-validation |
| SANA-WM camera_conditions must be sampled at latent frames… | validation | error | sana-wm, camera-conditions, shape-mismatch, latent-frames |
| token_tags must cover the full packed sequence | validation | error | minimax-h3, token-tags, packed-sequence |
| topk_length must be a CUDA tensor | validation | error | cuda, device-mismatch, topk |
| TRTLLM MHA backend for decode is only supported on Hopper… | validation | error | sglang, gpu-architecture, attention-backend, trtllm, decode |
| Unknown recipient | exception | error | harmony, recipient-routing, message-parsing, sglang |
| Unsupported num_bits = | validation | error | quantization, marlin, bit-width, unsupported-format |
| Unsupported weight strategy= | validation | error | quantization, fp8, w8a16, compressed-tensors, strategy |
| audio must be a 2D tensor, but got | validation | error | multimodal, audio, tensor-shape, mimo |
| Harmony does not support reasoning effort | validation | error | gpt-oss, harmony, reasoning, openai-api, sglang |
| kv must have shape (s_kv, h_kv, d_qk), got | validation | error | rank-validation, shape-validation, sparse-mla, kv-cache |
| message content must be a string or a sequence of parts | validation | error | inkling, type-error, content |
| must be contiguous | validation | error | contiguity, output-buffer, sparse-mla, strides |
| Native gRPC does not yet support --tokenizer-worker-num >… | validation | error | grpc, tokenizer-workers, config-validation |
| original_sr must be a positive number, but got | validation | error | multimodal, audio, sample-rate, mimo |
| SANA-WM CFG requires negative prompt embeds. | validation | error | cfg, negative-prompt, sana-wm |
| requires a CUDA tensor | validation | error | multimodal, cuda, tensor-device |
| Serialized W4A4 layer | validation | critical | quantization, dtype, validation, w4a4 |
| text QKV shapes must match | validation | error | shape, attention, hunyuan, rope |
| TGV cute_ext tactic out of range | validation | error | gemm, tgv, tactic, out-of-range |
| The fastokens package is required when… | exception | error | tokenizer, missing-dependency, installation |
| batching config rule requires model or model_contains | validation | error | config, batching, validation, selector |
| cannot parse boolean batching config value | validation | error | config, batching, boolean-parsing, validation |
| Dots note omni video preprocessing requires one request's… | validation | error | multimodal, sampling-params, type-mismatch, valueerror |
| DSpark speculative_num_draft_tokens must equal gamma + 1 | validation | error | speculative-decoding, dspark, conflicting-arguments, num-draft-tokens |
| DSpark with dp attention supports moe_a2a_backend 'none'… | validation | error | speculative-decoding, dspark, moe, a2a-backend, dp-attention |
| evictable_size() is not implemented; use… | exception | error | swa, radix-cache, not-implemented, api-misuse |
| Exactly one of 'prompt' or 'messages' must be provided. | validation | error | tokenization, validation, openai-api, sglang |
| FA4 path does not support rotary embedding. | exception | error | flash-attention, fa4, rotary-embedding, unsupported-operation |
| `final_sigmas_type` must be one of 'zero', or 'sigma_min'… | validation | error | scheduler, diffusion, sigmas, karras, unipc |
| full pack verification requested, but manifest has no full… | validation | error | kimi, moe, expert-pack, sha256, missing-field |
| FusedNormScaleShift cuda not available, using native… | console | warning | cuda, kernel-fallback, layernorm, performance, shape-constraint |
| port duplicates port and --strict-ports is enabled. | exception | critical | network, ports, startup, config |
| pair_postprocess must return a torch.Tensor | validation | error | scheduler, type-mismatch, callback |
| attention_head_dim must be positive. | validation | critical | config, model-architecture, validation, minimax-h3 |
| composite_input event payload requires | validation | error | realtime, event-validation, composite-input, missing-field |
| video block token count must be positive | validation | error | minimax-h3, video-presentation, validation |
| Cosmos3CausalAttention requires num_key_value_heads… | validation | error | sglang, cosmos3, tensor-parallel, gqa, kv-heads |
| cutedsl_mla backend can only be used with MLA models. | validation | error | attention-backend, cutedsl-mla, mla, model-arch-mismatch, sglang |
| flash_attn_varlen_func_op_lse is out+lse op… | validation | error | flash-attention, varlen, api-misuse, lse |
| Inkling reasoning_effort must be in [0.0, 0.99] | validation | error | inkling, reasoning, parameter-out-of-range, sglang |
| Invalid k_out shape for fused KV materialization: got | validation | error | shape-validation, kv-cache, speculative-decoding |
| kv-canary: length must be , got | validation | error | kv-canary, length-mismatch, validation |
| Model directory does not contain a transformer/ directory. | validation | error | diffusers, model-files, directory-layout |
| _block tensors must live on CUDA | validation | error | block-sparse, cuda, cpu-tensor, attention |
| sglang.srt.layers.attention.nsa.transform_index is… | console | warning | deprecation, sglang, nsa, import |
| --sidecar requires SGLang's native gRPC server; it cannot… | validation | error | grpc, sidecar, flag-conflict, config-validation |
| --ssl-keyfile requires --ssl-certfile to be specified as… | validation | error | sglang, ssl, tls, server-config, argument-validation |
| The MLX tensor bridge supports CPU and MPS targets, got | validation | error | sglang, mlx, device, target |
| ThinkingMode: , invalid message without… | exception | error | deepseek, thinking-mode, multi-turn |
| Unknown weight cache transport backend | exception | error | config, transport, invalid-argument |
| unsupported composite_input type | validation | error | realtime, event-validation, composite-input, unsupported-type |
| unsupported Inkling message role | validation | error | inkling, role-validation |
| unsupported intrinsics shape | validation | error | numpy, intrinsics, shape-validation, sana-wm |
| unsupported MiniMax H3 decoder task | validation | error | minimax-h3, task, enum-value-invalid |
| --asr-max-buffer-seconds must be positive | validation | error | sglang, asr, transcription, server-args, validation |
| browser.open requires a cursor or url | validation | error | browser, open, argument-validation, tool-calling |
| Config file path not set. Please set | exception | error | env-var, mooncake, config, missing-configuration |
| Connection closed while reading message header | error_code | error | protocol, socket, eof, weight-cache |
| Cosmos3 action generation requires --domain-id or… | validation | error | cosmos3, action-generation, missing-parameter, domain-id |
| --enable-http2 requires the 'granian' package. Install it… | validation | error | http2, missing-dependency, pip, server-args |
| ffprobe returned invalid stream metadata | exception | error | minimax-h3, ffprobe, stream-metadata, validation |
| flash_attn at sgl-kernel is only supported on sm90 and above | validation | error | flash-attention, fa3, gpu-compatibility, sm90, hardware-unsupported |
| Inkling reasoning_effort must be a number | validation | error | inkling, reasoning-effort, type-error |
| mask_strategy[0] cannot be None for SlidingTileAttention | validation | error | sliding-tile-attention, mask-strategy, validation, forward |
| model_config_parser= | exception | error | gguf, config, server-args |
| contains .gguf files; name the one to serve:\n | exception | error | gguf, huggingface, model-selection |
| Only neox-style RoPE is supported. | validation | error | unsupported-feature, rope, speculative-decoding |
| .aspect_ratio for task must be 'auto' or one of , got | validation | error | minimax-h3, aspect-ratio, enum-validation |
| .duration_seconds is required | validation | error | minimax-h3, duration, required-field |
| Rust workspace for was not found at | exception | error | rust, file-not-found, discovery, cargo |
| sglang.srt.layers.attention.nsa_backend is deprecated; use… | console | warning | deprecation, sglang, attention-backend, import |
| The length of cache_salt should be equal to the batch size. | validation | error | sglang, cache-salt, batching, input-validation |
| Unsupported VAE encode output for SANA-WM first-frame… | validation | error | vae, encode-output, compatibility, sana-wm |
| additional customized generation output is not supported by… | validation | error | rust-egress, output-streaming, feature-incompatibility |
| Calculated target_latency= | validation | error | pipeline-parallel, profiling, data-quality |
| Found unknown quantization= | exception | error | mistral, quantization, config |