Lightning-AI/pytorch-lightning
Documented errors, page 2 of 5. Back to Lightning-AI/pytorch-lightning
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| Accessing the device mesh before processes have initialized… | exception | error | pytorch-lightning, model-parallel, device-mesh, initialization-order |
| `pruning_fn` is expected to be a str in | exception | error | pytorch-lightning, pruning, pruning-fn, type-error |
| Device IDs (GPU/TPU) must be an int, a string, a sequence… | validation | error | devices, typeerror, numpy, sequence-validation |
| Found more than one stateful callback of type | validation | error | lightning, callbacks, state-dict, checkpointing, trainer-init |
| Skipping backward by returning `None` from your… | exception | error | xla, training-step, backward, pytorch-lightning, misconfiguration |
| Device IDs (GPU/TPU) must be an int, a string, a sequence… | validation | error | devices, typeerror, null-config, lightning |
| `precision='bf16-mixed'` does not use a scaler, found | exception | error | pytorch-lightning, amp, bf16, mixed-precision, gradscaler, plugin |
| The model contains a key | validation | error | lightning, fabric, checkpoint, state-dict, strict-loading |
| To spawn processes with the | validation | error | lightning, fabric, ddp-spawn, launch, xla |
| You set `strategy= ` but strategies from the DDP family are… | validation | error | pytorch-lightning, mps, apple-silicon, ddp, strategy-mismatch |
| Some provided `parameters_to_prune` don't exist in the… | exception | error | pytorch-lightning, pruning, parameters-to-prune, model-mismatch |
| The launcher can only create subprocesses once. | exception | error | pytorch-lightning, subprocess, launcher, ddp, lifecycle |
| str(_TRANSFORMER_ENGINE_AVAILABLE) | exception | error | transformer-engine, fp8, missing-dependency, gpu, pytorch-lightning |
| Epochs indexing from 1, epoch | validation | error | lightning, gradient-accumulation, unreachable, defensive-code |
| PyTorch `BasePruningMethod` is currently only supported… | exception | error | pytorch-lightning, pruning, base-pruning-method, unsupported-operation |
| ` ` does not support the… | exception | error | pytorch-lightning, combined-loader, multi-dataloader, training |
| Unable to find 'latest' file at | exception | error | pytorch-lightning, deepspeed, checkpoint, conversion |
| `val_dataloader` must be implemented to be used with the… | exception | critical | pytorch-lightning, lightning-module, dataloader, validation, hook-not-implemented |
| Found invalid type for plugin | validation | error | pytorch-lightning, trainer, plugins, config-validation |
| SpikeDetection requires `torchmetrics>=1.0.0` Please… | exception | error | lightning, torchmetrics, version-mismatch |
| Device should be MPS, got | validation | error | pytorch-lightning, trainer, val-check-interval, parse-error |
| In manual optimization, `training_step` must either return… | exception | error | pytorch-lightning, manual-optimization, training-step, return-type |
| `Trainer(strategy= )` is not compatible with an interactive… | validation | error | pytorch-lightning, ddp, jupyter, interactive-environment |
| Expected lengths ( ) to be greater or equal than samples ( ) | validation | error | pytorch-lightning, throughput, validation, argument-mismatch |
| `local_world_size` should be >= 1, got | exception | error | pytorch-lightning, validation, num-workers, world-size |
| Unknown state_dict_type | exception | error | xla, fsdp, config, invalid-value, lightning-fabric |
| You cannot set both `activation_checkpointing` and… | validation | error | fsdp, activation-checkpointing, config, mutually-exclusive, pytorch-lightning |
| Could not find a XLAFSDP model in the provided checkpoint… | exception | error | xla, fsdp, checkpoint, load, model-wrapping, lightning-fabric |
| ReduceLROnPlateau conditioned on metric | exception | error | pytorch-lightning, lr-scheduler, reducelronplateau, monitor-metric, logging |
| The Kubeflow environment can't be detected automatically. | exception | error | pytorch-lightning, kubeflow, cluster-environment, distributed |
| A single `Optimizer` cannot have multiple parameter groups… | validation | error | lightning, lr-monitor, optimizer, duplicate-name |
| Found multiple XLAFSDP modules in the given state. Loading… | exception | error | xla, fsdp, checkpoint, load, multiple-models, lightning-fabric |
| Skipping backward by returning `None` from your… | exception | error | pytorch-lightning, deepspeed, training-step, backward, automatic-optimization |
| The optimizer does not seem to reference any FSDP… | validation | error | fsdp, optimizer, flat-params, setup-order |
| AMP and the LBFGS optimizer are not compatible. | exception | error | pytorch-lightning, amp, lbfgs, optimizer, gradscaler, optimizer-step |
| Combination of parameters every_n_train_steps= | validation | error | pytorch-lightning, modelcheckpoint, mutually-exclusive-args, config-validation |
| The checkpoint path exists and is a directory | exception | error | fsdp, save-checkpoint, is-a-directory, paths |
| You are trying to use `ScheduleWrapper` which require… | exception | error | profiler, kineto, dependency-missing, pytorch |
| DataFetcher is unsupported for | exception | error | pytorch-lightning, internal, running-stage, data-fetcher |
| Do not set `gradient_accumulation_steps` in the DeepSpeed… | exception | error | deepspeed, gradient-accumulation, config-conflict |
| f"You requested to check | validation | error | pytorch-lightning, limit-batches, validation-loop, misconfiguration |
| The `configure_payload` method needs to be overridden. | exception | error | lightning, serving, not-implemented, pytorch |
| The dirpath has changed from | console | warning | model-checkpoint, resume, dirpath, lightning |
| To use the `DeepSpeedStrategy`, you must have DeepSpeed… | exception | critical | deepspeed, strategy, dependencies, installation |
| `. (ckpt_path="hpc")` is set but no HPC checkpoint was… | validation | error | lightning, hpc, resume, checkpoint, slurm |
| Automatic gradient clipping is not supported for manual… | validation | error | pytorch-lightning, manual-optimization, gradient-clipping, configuration |
| . It looks like you passed the path to a subfolder. Try to… | error_code | error | deepspeed, checkpoint, path, load |
| Directory ' ' doesn't exist | exception | error | pytorch-lightning, deepspeed, checkpoint, file-not-found |
| outputs have to be of type torch.Tensor or Mapping, got | exception | error | spike-detection, training-step, typeerror, callback |
| The `ModelParallelStrategy` does not support `Fabric | validation | error | pytorch-lightning, model-parallel, precision, unsupported-combination |
| `Trainer(strategy='deepspeed', precision= )` is not… | exception | error | pytorch-lightning, deepspeed, precision, config-validation |
| Post-localSGD algorithm is used, but model averaging period… | exception | error | pytorch, lightning, ddp, distributed, post-local-sgd |
| Found multiple FSDP models in the given state. Saving… | validation | error | fsdp, checkpoint, save, pytorch-lightning, distributed |
| Support for ` ` has been removed in v2.0.0. ` ` implements… | exception | error | pytorch-lightning, migration, v2-breaking-change, legacy-hooks |
| The selected device indices | error_code | critical | deepspeed, cuda, device-selection, multi-gpu |
| You have configured optimizers but the checkpoint contains… | exception | critical | fsdp, resume, optimizer-states, checkpoint-mismatch |
| GPUs should be a list | exception | error | pytorch-lightning, gpu, device-parser, internal-api |
| Unsupported | exception | error | lightning, checkpoint, not-implemented, lightningmodule |
| `activation_checkpointing_policy` must be a set, found | exception | error | xla, fsdp, activation-checkpointing, type-error, lightning-fabric |
| DeepSpeed and the LBFGS optimizer are not compatible. | exception | error | pytorch-lightning, deepspeed, lbfgs, optimizer, distributed |
| In automatic optimization, `training_step` must return a… | exception | error | pytorch-lightning, training-step, return-type, automatic-optimization |
| The CombinedLoader has | validation | critical | combined-loader, checkpoint, state-mismatch, resume |
| Instantiating your model under the `init_module` context… | validation | error | bitsandbytes, quantization, init-module, pytorch-lightning |
| ModelCheckpoint(save_top_k= , monitor=None) is not a valid… | validation | error | pytorch-lightning, modelcheckpoint, monitor, config-validation |
| PyTorch >= 2.6 requires DeepSpeed >= 0.16.0. Detected… | exception | critical | deepspeed, version-mismatch, pytorch-2-6, compatibility |
| You are trying to `self.log()` but the loop's result… | exception | error | pytorch-lightning, self-log, predict-step, logging, misconfiguration |
| DeepSpeed currently only supports single optimizer, single… | exception | error | deepspeed, lr-scheduler, configure-optimizers |
| f"`Trainer(barebones=True, log_every_n_steps= )` was… | validation | error | trainer, barebones, logging, misconfiguration, pytorch-lightning |
| Schedule should return a `torch.profiler.ProfilerAction`… | exception | error | profiler, schedule, type-error, pytorch |
| Invalid seed specified via PL_GLOBAL_SEED | validation | error | lightning, seeding, environment-variables |
| m.format("on_epoch", on_epoch, fx_name… | validation | error | pytorch-lightning, logging, on-epoch, misconfiguration |
| The given dataset must implement the `__len__` method. | exception | error | pytorch-lightning, distributed, distributed-sampler, iterable-dataset |
| You are trying to `self.log()` but it is not managed by the… | exception | error | pytorch-lightning, self-log, control-flow, logging, misconfiguration |
| You need to set up the model first before you can call… | exception | error | lightning, fabric, no-backward-sync, ddp, setup-order |
| Schedule should be a callable. Found | exception | error | profiler, schedule, type-error, pytorch-lightning |
| Blocking backward sync is only possible if the module… | validation | error | ddp, gradient-accumulation, type-mismatch, pytorch-lightning |
| Found modules that are wrapped with… | exception | error | fsdp, fsdp2, legacy-api, pytorch-version |
| `name` must be a str, found | validation | error | pytorch-lightning, trainer, barebones, checkpointing |
| `predict_dataloader` must be implemented to be used with… | exception | error | pytorch-lightning, lightning-module, dataloader, inference, hook-not-implemented |
| Found multiple distributed models in the given state… | validation | error | lightning, fabric, model-parallel, checkpoint, multiple-models |
| `save_to_log_dir=False` only makes sense when subclassing… | exception | error | lightning-cli, save-config, save-to-log-dir, misuse-guard |
| You provided multiple | exception | error | pytorch-lightning, dataloader-idx, hook-signature, multi-dataloader |
| str(_XLA_AVAILABLE) | exception | error | xla, tpu, missing-dependency, pytorch-lightning |
| The optimizer has references to the model's meta-device… | exception | error | lightning, fabric, fsdp, meta-device, init-module |
| `trainer.predict()` only supports the… | exception | error | pytorch-lightning, predict, combined-loader, multi-dataloader |
| Experiment is not initialized | exception | error | litlogger, lightning, logger, lazy-initialization |
| You passed in a path to a DeepSpeed config but the path… | error_code | error | deepspeed, config, file-not-found, path |
| `. (ckpt_path="best")` is set but `ModelCheckpoint` is not… | validation | error | lightning, checkpoint, best-model, validate, test, predict |
| `setup_optimizers` requires at least one optimizer as input. | validation | error | pytorch-lightning, fabric, optimizer, validation |
| Trainer was configured with `enable_checkpointing=False`… | validation | error | lightning, trainer-init, checkpointing, callbacks, config-conflict |
| You selected an invalid strategy name: `strategy= | validation | error | pytorch-lightning, trainer, strategy, invalid-argument, ddp |
| Your LightningModule code tried to access `self.trainer. | exception | error | fabric, trainer, attribute-error, lightning |
| In automatic_optimization, when `training_step` returns a… | exception | error | pytorch-lightning, training-step, loss, automatic-optimization |
| expected to NOT exist. Aborting to avoid overwriting… | exception | error | lightning-cli, save-config, overwrite, file-exists, idempotent-rerun |
| TensorRT only supports CUDA devices. The current device is | exception | error | tensorrt, cuda, device-mismatch, lightning |
| The DeepSpeed strategy is only supported on CUDA GPUs but | exception | critical | deepspeed, accelerator, cuda, hardware-requirement |
| The loss returned in `training_step` is | exception | critical | pytorch-lightning, nan-loss, numerical-stability, training |
| `ModelCheckpoint(monitor= )` could not find the monitored… | validation | error | lightning, model-checkpoint, metric-not-found, monitor |
| The FSDP strategy can only work with the `FSDPPrecision`… | exception | error | fsdp, precision-plugin, mixed-precision, type-mismatch |
| `Trainer.save_checkpoint(..., storage_options=...)` with… | exception | error | deepspeed, save-checkpoint, storage-options, checkpointio |