What problem does it solve?
Manually gathering accurate, up-to-date GPU cluster hardware specifications and account-specific job limits for LUMI and Snellius supercomputers is time-consuming and error-prone, often leading to failed scale-up experiment submissions or quota overruns due to stale documentation.
Core Features & Use Cases
- Live Job Limit Queries: Pulls real-time Slurm account-specific running and pending job caps, active job counts, and remaining headroom instead of relying on outdated published defaults.
- Static Hardware Summaries: Provides comparable, side-by-side hardware details for LUMI (ROCm/MI250X) and Snellius (CUDA/A100/H100) clusters including GPU memory per node, node counts, and partition rules.
- Use Case: Before submitting a large-scale ML training run, use this skill to quickly compare LUMI and Snellius capacity, confirm your account has sufficient job headroom, and select the correct partition to avoid submission failures.
Quick Start
Use the sue-cluster-info skill to get a summary of LUMI and Snellius GPU cluster hardware and your account's live job limits for scale-up experiment planning.