sgl-project/sglang
Documented errors, page 14 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| Runai Model Streamer Loader does not support ModelOpt… | exception | error | runai-streamer, modelopt, not-implemented, unsupported-combination, sglang |
| The last dimension ( ) x itemsize ( ) must be a multiple of… | validation | error | sglang, cuda-kernel, alignment, quick-gelu, shape-validation |
| The W8A8Int8 Fused MoE scheme is implemented only for NPU… | exception | error | quantization, moe, int8, npu, hardware-support |
| This browser does not support worker image decoding | error_code | critical | rocm, allreduce, deterministic, tensor-parallel, float32 |
| pool is too small after control metadata: pool= , control= | validation | error | multimodal, memory-pool, capacity-planning |
| Unsupported KV cache type for decode offload | validation | error | disaggregation, kv-offload, hicache, unsupported-pool-type |
| auto_duration was requested but this checkpoint has no… | exception | error | ltx-2, auto-duration, checkpoint-capability, version-mismatch |
| cannot mix lambda extractor with static kwargs | exception | error | decorator, api-misuse, debug-utils |
| . : expected , got | validation | error | type-check, dataclass, runtime-validation, config |
| condition_frame_indexes= | validation | error | cosmos3, video, frame-index, out-of-range, validation |
| Cosmos3 action generation does not support CFG parallel yet | error_code | error | cosmos3, cfg-parallel, action-generation, not-implemented |
| Dit target weight name | validation | error | lora, checkpoint, duplicate-key |
| empty match text | exception | error | patching, validation, source-patcher, sglang |
| --enable-int8-mamba-checkpoint only supports the built-in… | validation | error | sglang, mamba, int8, radix-cache, incompatible-flags, server-args |
| Expected outputs, got from scheduler | exception | critical | scheduler, consistency-check, batch-generation, diffusion |
| Fused QK-Norm + RoPE kernel only supports float16/bfloat16… | validation | error | dtype, rope, fused-kernel, joyimage, multimodal |
| H3 conditioning projection tap | exception | critical | minimax-h3, conditioning-projection, layer-index, out-of-range |
| Hunyuan3D expects image_path as str, got | validation | error | hunyuan3d, type-error, input-validation |
| image_rotary_emb must be cos_sin_cache tensors | validation | error | qwen-image, rope, cos-sin-cache, format-validation |
| Invalid modality | validation | error | multimodal, modality, enum-validation, receiver |
| mega MoE: num_tokens= | exception | critical | moe, deepgemm, kimi-k3, mega-moe, buffer-capacity, env-var |
| MiniMax-H3 ring parallelism requires the FlashAttention… | validation | error | minimax-h3, ring-parallelism, attention-backend, server-args |
| n_q/n_k must be one of | validation | error | hyperparameter, validation, sparse-attention, constructor |
| Not a canonical UMMA_K Layout: Expected profile failure. | validation | error | cutlass, sm100, layout, shared-memory |
| O partial tensor must match dtype_partial | validation | error | cuda, dtype, flash-attention, split-kv, combine |
| Stochastic rounding for the Mamba SSM cache with… | validation | error | sglang, mamba, triton, sm100, gpu-architecture, stochastic-rounding |
| Unsupported down_block_type | validation | error | python, value-error, model-config, hunyuan-vae, unsupported-block |
| Unsupported Hugging Face weight URL | validation | error | huggingface, weights, unsupported-url-shape |
| unsupported MiniMax H3 model variant | validation | error | minimax-h3, model-variant, unsupported-value, sglang |
| Unsupported online_scheme | exception | error | quantization, quark, config-validation, online-requantization, python |
| video must be a tuple of (video_tensor, timestamps), but got | validation | error | video, contract, type-error, decoding |
| assistant reasoning_content must be a string for Inkling… | validation | error | inkling, reasoning, type-error |
| audio_sampling_rate must be set in processor_config or… | validation | error | audio, config, sampling-rate, initialization |
| backend must be a non-empty string | validation | error | sampler, validation, configuration, sglang |
| canonical request has unknown fields | validation | error | minimax-h3, unknown-fields, schema-drift, resolved-plan |
| conditions[ ]: video references are not supported in v1 for… | validation | error | minimax-h3, video-input, unsupported-type, request-validation |
| Cross attention is not supported in the hpc_ops attention… | validation | error | attention-backend, hpc-ops, encoder-decoder, cross-attention, npu, sglang |
| Every extra_key should be a string. | validation | error | sglang, extra-key, cache-key, input-validation, batching |
| HiCache native hash is only supported on little-endian Linux | exception | error | native-extension, platform-support, linux-only, endianess |
| Humming quantization requires `humming-kernels`. Please… | exception | error | humming, quantization, missing-dependency, import-error, installation |
| Invalid quantization method on CPU | validation | error | quantization, cpu, amx, platform-support |
| kv-canary: max_seq_len_per_req must be positive, got | exception | error | kv-canary, validation, seq-len, value-error |
| LoRA with name already exists. Loaded LoRAs | validation | error | lora, registry, duplicate, register, sglang |
| MiniMax H3 attention heads must be divisible by TP size | exception | critical | minimax-h3, tensor-parallel, config-validation, startup |
| MiniMax H3 task must be a non-empty string | validation | error | minimax-h3, task-validation, invalid-argument |
| MLX model has no supported attention layers | error_code | error | mlx, kv-cache, layout, model-compatibility |
| must be non-decreasing | validation | error | npu, ascend, varlen, packed-sequences, validation |
| O partial tensor must have 4 or 5 dimensions: (num_splits… | validation | error | flash-attention, shape-mismatch, rank-error, validation |
| `out` must be contiguous. | exception | error | fla, fused-recurrent, contiguity, output-buffer |
| --prefill-only-disable-kv-cache is incompatible with… | validation | error | sglang, hisparse, sparse-attention, kv-cache, server-args |
| Reference attention was not initialized. | panic | error | runtime, initialization, reference-attention, diffusion |
| rollout_noise_level must be a number, got | validation | error | rl-rollout, sampling-params, type-validation, config |
| SGLANG_DSA_TOPK_BROADCAST requires PyNCCL during CUDA graph… | exception | critical | dsa, pynccl, cuda-graph, broadcast, tp, env-var, sglang |
| target.duration_seconds is required, or exactly one audio… | validation | error | minimax-h3, duration, target, request-validation |
| temp_set_env should not be used for sglang env vars | validation | warning | environment-variables, testing, conventions, sglang |
| Unsupported IO backend | validation | error | sglang, io-backend, kv-cache, hierarchical-cache, configuration |
| Video sampling strategy not specified | validation | error | video, config, sampling, missing-value |
| Could not decode video | validation | error | video, codec, decode, torchcodec, decord |
| Decode context parallel (decode_context_parallel_size > 1)… | exception | error | parallelism, decode-context-parallel, platform-support, cuda, rocm, sglang |
| every request must verify the anchor (verify_len >= 1), got | validation | error | speculative-decoding, validation, off-by-one |
| GQA/MQA requires query heads to be a multiple of KV heads… | validation | error | sage-attention, gqa, head-mismatch, validation |
| Invalid = . | validation | error | speculative-decoding, dflash, config-parsing, type-validation |
| Invalid join mode | validation | error | validation, fork-join, dsl, enum-value |
| k shape must be [num_tokens, num_kv_heads, head_dim], got | validation | error | shape-validation, gqa, rope, metal, sgl-kernel |
| layer_id= is not a sparse attention layer; sparse layers | validation | error | kv-cache, sparse-attention, layer-id, mapping, minimax, sglang |
| LingBot causal sequence sharding requires… | exception | error | sequence-parallelism, missing-attribute, forward-batch, lingbot |
| MiniMax H3 text payload broadcast failed | error_code | critical | minimax-h3, broadcast, tensor-dict, collective, distributed-communication |
| MiniMaxH3TextEncodingStage direct encode requires a… | validation | error | minimax-h3, missing-component, tokenizer |
| MM inputs where only some items are precomputed. | exception | error | sglang, multimodal, precomputed-embeddings, not-implemented |
| must be a floating point tensor | validation | error | diffusion, scheduler, dtype, timestep |
| Namespace ' ' for rank not initialized | exception | error | hf3fs, metadata, initialization-order, local-client |
| Regular expression is not supported in the OpenAI backend. | console | warning | regex, openai, structured-output, constraint-dropped, warning |
| SANA-WM does not support tensor parallelism yet. Use… | validation | error | sana-wm, tensor-parallelism, unsupported-feature, sglang |
| SGLANG_RUST_SERVER does not yet apply… | validation | error | sglang, rust-server, config-conflict, startup |
| store_kv: token_ids has | validation | error | flexkv, kv-cache, validation, off-by-one |
| subsampling_conv_chunking_factor should be -1, 1, or a… | validation | error | subsampling, chunking, audio, config-validation, phi4 |
| The messages should be a list of dict. | validation | error | chat-api, request-validation, messages |
| The output size is not aligned with the quantized weight… | validation | error | gptq, tensor-parallel, shape-mismatch, pack-factor, cpu, amx |
| This browser does not support gzip stream decoding | error_code | error | radix-tree, hicache, not-implemented, io-commit |
| Whisper expects exactly 1 audio input, got | exception | error | whisper, audio, single-input-constraint |
| attn_res: nvb must be in | validation | error | kimi-k3, tuning, argument-validation, range-check |
| Cannot parse schema . The schema must be either a Pydantic… | validation | error | json-schema, structured-output, validation |
| data URI must use ;base64 encoding | validation | error | minimax-h3, data-uri, base64, material-io |
| Environment variable | console | warning | deprecation, environment-variables, migration, startup |
| Expected encoder_outputs to be a list when select_layers is… | validation | error | siglip2, api-misuse, layer-selection |
| expert-pack stats flush interval cannot be negative | exception | error | moe, expert-pack, configuration, interval, argument-validation |
| flattened_bucket 'metadata' must be a list. | validation | error | weights-update, flattened-bucket, metadata, multimodal |
| input contains occurrences of embed_override_token_id= … | validation | error | sglang, embed-overrides, multimodal, count-mismatch |
| MiniMax H3 text payload positive.text_len must match the… | validation | error | minimax-h3, text-encoding, length-mismatch |
| MXFP8 KV cache requires head_dim divisible by | validation | error | kv-cache, mxfp8, quantization, head-dim |
| --prefill-only-disable-kv-cache does not currently support… | validation | error | sglang, kv-cache, mxfp8, quantization, server-args |
| qprep_bf16_fp8_sm90 requires an SM90 (Hopper) GPU | exception | error | cuda, sm90, hopper, jit-kernel, hardware-unsupported |
| rank_consensus() got unexpected keyword argument(s) | validation | error | decorator, distributed, rank-consensus, api-misuse, sglang |
| Received request with | validation | error | lora, multi-lora, limits, configuration |
| The requested FlashAttention forward configuration exceeds… | validation | error | flash-attention, sm120, shared-memory, config-validation |
| Unsupported video input type for EPD encoder | validation | error | video, input-validation, type-error, multimodal |
| attn_sink must be on q's device | validation | error | mla, sparse-attention, multi-gpu, device-mismatch |
| Block sparse tensors | validation | error | block-sparse, block-size, shape-mismatch |
| DSpark speculative_num_draft_tokens must be >= 2 (= gamma +… | validation | error | sglang, dspark, speculative-decoding, config-validation, gamma |
| --enable-strict-thinking requires a grammar backend with… | validation | critical | grammar, strict-thinking, xgrammar, startup |