k8s-network-engineer

Configure and troubleshoot NVIDIA networking deployments on Kubernetes.

14|5|Updated Nov 4, 2025
One-click install
npx skills add https://github.com/NVIDIA/k8s-launch-kit --skill k8s-network-engineer
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: k8s-network-engineer
Source: https://github.com/NVIDIA/k8s-launch-kit/tree/main/skills/k8s-network-engineer
Command: npx skills add https://github.com/NVIDIA/k8s-launch-kit --skill k8s-network-engineer

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill helps users design, deploy, validate, and troubleshoot high-performance NVIDIA networking on Kubernetes without navigating complex hardware, fabric, Network Operator, and deployment-profile choices manually.

Core Features & Use Cases

  • Cluster Discovery and Configuration: Identify NIC hardware, topology, capabilities, and profile settings with k8s-launch-kit.
  • Profile Selection and Manifest Generation: Choose and generate configurations for SR-IOV, RDMA Shared, Host Device, InfiniBand, and Spectrum-X deployments.
  • Deployment Safety and Troubleshooting: Preview changes, apply coordinated operator settings, validate releases, collect diagnostics, and analyze failures.
  • Use Case: Configure a multi-rail ConnectX-8 GPU cluster for Spectrum-X, select an appropriate multiplane mode, generate manifests, preview the deployment, and validate the resulting Network Operator resources.

Quick Start

Ask the k8s-network-engineer skill to discover your cluster, recommend a suitable NVIDIA networking profile, generate a dry-run configuration, and explain any deployment prerequisites.

Frequently Asked Questions about k8s-network-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I configure RDMA and SR-IOV networking on Kubernetes?

Configuring RDMA and SR-IOV networking on Kubernetes requires using the NVIDIA Network Operator to select hardware-aware profiles and generate deployment manifests. This Skill guides you through profile selection, manifest generation, and safe dry-run practices for your specific NIC topology.

What is the best way to deploy Spectrum-X on a Kubernetes cluster?

The best way to deploy Spectrum-X on Kubernetes is to discover your ConnectX NIC topology, select a multiplane mode, and generate manifests via the NVIDIA Network Operator. This Skill facilitates fabric and NIC compatibility checks to validate the resulting deployment resources.

How do I troubleshoot InfiniBand network operator failures in Kubernetes?

Troubleshooting InfiniBand Network Operator failures in Kubernetes involves collecting diagnostics and analyzing deployment resources against operator release constraints. This Skill helps you preview coordinated operator settings, validate releases, and analyze failures using safe dry-run practices.

Does the NVIDIA Network Operator support multi-rail ConnectX deployments?

Yes, the NVIDIA Network Operator supports multi-rail ConnectX deployments by allowing you to select appropriate multiplane modes and generate coordinated configurations. This Skill helps identify hardware capabilities, validate multi-rail profile settings, and ensure compatibility across your cluster.

When do I need to use DOCA and BlueField profiles for Kubernetes networking?

You need DOCA and BlueField profiles for Kubernetes networking when deploying high-performance Spectrum-X or RDMA workloads requiring hardware-aware configuration and offloaded data paths. This Skill determines the appropriate profile based on your discovered NIC topology and capabilities.

Why is my SR-IOV network operator deployment not working on Kubernetes?

Your SR-IOV Network Operator deployment on Kubernetes may fail due to mismatched hardware capabilities, incorrect profile selection, or unmet operator release constraints. This Skill helps preview changes, validate prerequisites, and analyze deployment failures using coordinated diagnostics.