What problem does it solve?
This Skill helps diagnose and resolve NVIDIA Network Operator and Kubernetes networking failures, reducing downtime caused by broken OFED drivers, missing virtual functions, failed pods, RDMA issues, and misconfigured network resources.
Core Features & Use Cases
- Cluster Troubleshooting: Inspect Network Operator pods, NicClusterPolicy status, SR-IOV node states, device plugins, logs, events, and network resources.
- Sosreport Analysis: Navigate diagnostic dumps covering cluster metadata, CRDs, operator components, node resources, networking state, and collection errors.
- Failure Remediation: Identify likely causes and corrective actions for OFED crashes, VF creation failures, IP allocation problems, missing NetworkAttachmentDefinitions, stuck pods, and discovery failures.
- Use Case: When a workload is stuck in ContainerCreating and cannot obtain a secondary network, use this Skill to inspect pod events, Multus logs, NetworkAttachmentDefinitions, CNI binaries, and SR-IOV resource availability.
Quick Start
Ask the k8s-launch-kit-troubleshoot skill to analyze the provided sosreport and identify the root cause of the Network Operator failure.