perf-moe-hardware-configs

Provide hardware-specific MoE training configurations and performance benchmarks.

852|445|Updated May 21, 2025
One-click install
npx skills add https://github.com/NVIDIA-NeMo/Megatron-Bridge --skill perf-moe-hardware-configs
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: perf-moe-hardware-configs
Source: https://github.com/NVIDIA-NeMo/Megatron-Bridge/tree/main/skills/perf-moe-hardware-configs
Command: npx skills add https://github.com/NVIDIA-NeMo/Megatron-Bridge --skill perf-moe-hardware-configs

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) and assets (resource) components.

What problem does it solve?

This Skill provides detailed hardware-specific training playbooks and performance benchmarks for MoE models, helping users optimize training configurations.

Core Features & Use Cases

  • Performance Benchmarking: Offers approximate throughput ranges and MFU metrics for various hardware platforms.
  • Configuration Guidance: Provides recommended parallelism, routing, and recompute strategies based on hardware.
  • Use Case: A researcher tuning MoE models can consult this Skill to select the best configuration for H100 or B200 systems tailored to their model size.

Quick Start

Request the optimal hardware configuration for training a 685B MoE model on an H100 system.

Frequently Asked Questions about perf-moe-hardware-configs

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I optimize MoE training configurations for specific hardware?

To optimize MoE training, select effective parallelism, routing, and recompute strategies based on your specific hardware specs and model sizes. This configuration guidance ensures you achieve optimal performance across diverse systems.

What parallelism and routing strategies work best for large MoE models?

Effective parallelism and routing strategies for large MoE models depend on your hardware platform and model size. Recommended configurations provide specific tuning advice to maximize training throughput and hardware utilization.

Can I get performance benchmarks for MoE training on H100 systems?

Performance benchmarks for MoE training on H100 systems provide approximate throughput ranges and MFU metrics. These metrics guide deployment planning and help validate your hardware-specific training configuration.

How do I choose the right recompute strategy for MoE model training?

Choosing the right recompute strategy for MoE training requires evaluating your hardware specs and model size alongside performance benchmarks. This ensures an effective balance between memory savings and computational throughput.

Does MoE hardware configuration require detailed performance metrics?

MoE hardware configuration requires detailed configuration data and performance metrics to effectively guide deployment planning. This data enables the selection of optimal parallelism and recompute strategies tailored to your system.