hyperpod-node-debugger

Diagnose and remediate per-node issues on HyperPod clusters.

15|9|Updated Aug 16, 2025
One-click install
npx skills add https://github.com/haozhx23/HyperPod-InstantStart --skill hyperpod-node-debugger
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: hyperpod-node-debugger
Source: https://github.com/haozhx23/HyperPod-InstantStart/tree/main/.claude/skills/hyperpod-node-debugger
Command: npx skills add https://github.com/haozhx23/HyperPod-InstantStart --skill hyperpod-node-debugger

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires aws, python3, bash, kubectl, session-manager-plugin, unbuffer, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill enables you to diagnose and remediate per-node issues on a HyperPod cluster, including identifying and resolving problems with specific nodes, such as hardware failures, connectivity issues, and resource constraints.

Core Features & Use Cases

  • Node Health Monitoring: Checks the health status of individual nodes, including GPU and EFA device status, network connectivity, and resource usage.
  • Problem Detection: Identifies issues such as hardware failures, network connectivity problems, and resource exhaustion.
  • Remediation Recommendations: Provides step-by-step instructions for resolving detected issues, including rebooting or replacing nodes, and configuring security groups.
  • Use Case: If a node in your HyperPod cluster is unresponsive or experiencing performance issues, this Skill can help you diagnose the problem and take appropriate action to resolve it.

Quick Start

Use the hyperpod-node-debugger skill to run diagnostics on a specific node within the HyperPod cluster.

Frequently Asked Questions about hyperpod-node-debugger

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I diagnose per-node issues and hardware failures on a HyperPod cluster?

To diagnose per-node issues on a HyperPod cluster, you need to check the health status of individual nodes, including GPU and EFA device status, network connectivity, and resource usage. This process identifies hardware failures and resource exhaustion to pinpoint specific node problems.

What is the best way to remediate unresponsive nodes or network connectivity problems in HyperPod?

The best way to remediate unresponsive nodes in HyperPod is to follow step-by-step remediation recommendations, which include rebooting or replacing problematic nodes and configuring security groups to resolve detected network connectivity problems and hardware failures.

Do I need AWS access and EKS or Slurm familiarity to troubleshoot HyperPod node health?

Yes, troubleshooting HyperPod node health requires access to AWS services like EC2, S3, and CloudWatch. You also need familiarity with HyperPod architecture and EKS or Slurm configurations to effectively identify and resolve resource constraints and connectivity issues.

Can I monitor GPU and EFA device status for resource constraints on individual HyperPod cluster nodes?

Yes, you can monitor GPU and EFA device status for individual nodes within a HyperPod cluster. Node health monitoring checks specific device statuses, network connectivity, and overall resource usage to detect and resolve hardware failures and resource exhaustion.

Why does my HyperPod cluster node experience performance issues and how can I resolve them?

HyperPod cluster nodes experience performance issues due to hardware failures, network connectivity problems, or resource exhaustion. You can resolve these issues by running diagnostics to identify the root cause and executing recommended remediation steps like rebooting or replacing the affected node.

What are the limitations of using kubectl and session-manager-plugin for HyperPod node diagnosis?

Diagnosing HyperPod nodes with kubectl and session-manager-plugin requires unbuffer and python3 dependencies. Limitations include the need for direct AWS service access and familiarity with EKS or Slurm configurations to accurately resolve hardware failures and network connectivity problems.