What problem does it solve?
This Skill simplifies the complex process of interacting with the Stanford Marlowe HPC cluster, ensuring users can submit, monitor, and manage their jobs effectively and compliantly.
Core Features & Use Cases
- Cluster State Verification: Discovers and verifies live cluster facts before proposing commands, preventing guesswork.
- Safe Job Submission: Guides users through
sbatch, salloc, and srun with Marlowe-specific account and partition requirements.
- Monitoring & Diagnosis: Helps track job status, diagnose pending reasons, and review finished jobs using Slurm commands.
- GPU-Hour Tracking: Provides guidance on monitoring GPU-hour consumption for relevant projects.
- Use Case: A researcher needs to submit a GPU-accelerated job on the Marlowe cluster. They can use this Skill to ensure they are using the correct account suffix, partition, and loading the necessary modules, then submit the job safely and monitor its progress.
Quick Start
Use the marlowe-slurm-operator skill to verify the current state of the 'preempt' partition on the Marlowe cluster.