slime-rl-training

Guide RL-based post-training of LLMs with Megatron-LM and SGLang.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/helix4u/hermes-agent --skill slime-rl-training-helix4u
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/helix4u/hermes-agent/tree/main/skills/mlops/training/slime
Command: npx skills add https://github.com/helix4u/hermes-agent --skill slime-rl-training-helix4u

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

slime provides guidance for RL-based post-training of LLMs using the slime framework, helping researchers and engineers structure experiments and scale RL training workflows.

Core Features & Use Cases

  • Megatron-LM integration with SGLang-based rollout for scalable RL post-training.
  • Data-generation workflow templates and buffer management for RL tasks.
  • End-to-end guidance for configuring experiments, monitoring progress, and debugging RL pipelines.

Quick Start

Launch slime training with your model checkpoint and prompt data to begin RL post-training.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I configure RL post-training for LLMs using Megatron-LM and SGLang?

RL post-training for LLMs using Megatron-LM and SGLang is configured by launching the slime framework with your model checkpoint and prompt data. This provides end-to-end guidance for structuring scalable training experiments across GPUs.

Do I need a specific environment to run slime RL training workflows?

Yes, running slime RL training workflows requires a Megatron-LM environment, the slime framework, and access to model checkpoints and prompt data. This setup is necessary for researchers and engineers performing scalable post-training.

What is the best way to manage data generation and buffers for RL tasks?

The best way to manage data generation and buffers for RL tasks is using the slime framework's custom data-generation workflow templates. These templates help structure data pipelines and buffer management for reinforcement learning.

How does SGLang-based rollout work with Megatron-LM for scalable training?

SGLang-based rollout integrates with Megatron-LM to enable scalable RL post-training across GPUs. The slime framework provides the necessary templates to configure experiments, monitor progress, and debug these distributed pipelines.

Why use slime for reinforcement learning post-training instead of other tools?

Use slime for reinforcement learning post-training to get end-to-end guidance on configuring experiments, monitoring progress, and debugging pipelines. It specifically integrates Megatron-LM and SGLang for scalable custom data generation workflows.

Can I debug RL pipelines while monitoring training progress?

Yes, you can debug RL pipelines while monitoring training progress using the slime framework. It provides end-to-end guidance for configuring experiments, tracking metrics, and troubleshooting Megatron-LM and SGLang deployments.