What problem does it solve?
This Skill helps you provision and operate Lambda Labs GPU cloud instances without manually juggling console steps, SSH setup, storage attachment, or cluster commands.
Core Features & Use Cases
- Instance lifecycle management: Launch, inspect, and terminate dedicated GPU instances for experiments, training runs, and inference jobs.
- Persistent storage workflows: Attach Lambda filesystems so datasets, checkpoints, models, and outputs survive instance termination.
- Distributed ML operations: Run single-node and multi-node training with SSH tunneling, JupyterLab, TensorBoard, Slurm, and API automation.
- Use Case: A machine learning engineer can spin up an H100 instance, mount shared storage, launch a fine-tuning job, monitor logs remotely, and cleanly terminate the instance after checkpoints are saved.
Quick Start
Ask the skill to launch a Lambda Labs GPU instance for my training job and guide me through connecting, mounting storage, and shutting it down safely.