slime-rl-training

Implement RL post-training for GLM models with slime, Megatron-LM, and SGLang.

Updated Apr 26, 2026
One-click install
npx skills add https://github.com/dawsonblock/HERMY --skill slime-rl-training-dawsonblock
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/dawsonblock/HERMY/tree/main/hermes-agent-2026.4.23/optional-skills/mlops/slime
Command: npx skills add https://github.com/dawsonblock/HERMY --skill slime-rl-training-dawsonblock

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Provides guided, structured methodologies for reinforcement-learning-based post-training of large language models using slime, enabling researchers to implement end-to-end RL pipelines with Megatron-LM and SGLang.

Core Features & Use Cases

  • Framework integration: Seamlessly connects slime with Megatron-LM and SGLang for RL post-training workflows.
  • Custom data generation and rollout management: Supports data generation, rollout strategies, and evaluation hooks for RL experiments.
  • Use Case: Researchers can prototype RL-based post-training for GLM-family models with custom reward functions and multi-turn interactions.

Quick Start

Install slime and run the included quick-start example to begin RL post-training with Megatron-LM and SGLang.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up RL post-training for LLMs using slime and Megatron-LM?

To set up RL post-training for LLMs using slime and Megatron-LM, install slime and run the included quick-start example to begin configuring distributed training and rollout optimization workflows.

Can I use custom reward functions for on-policy and off-policy training with slime?

Yes, slime supports custom reward functions and multi-turn interactions for both on-policy and off-policy training, allowing researchers to implement tailored rollout-based optimization for GLM-family models.

Does slime integrate with SGLang for rollout data handling?

Slime integrates with SGLang to manage rollout data handling and routing, providing structured data generation workflows and evaluation hooks for RL post-training experiments.

What is the best way to configure Megatron-LM for GLM-family model training?

The best way to configure Megatron-LM for GLM-family model training is using slime's guided methodologies, which provide structured configuration for rollout-based optimization and custom data generation.

Are there limitations when using slime for RL post-training in research environments?

Slime is designed for research environments and specifically targets GLM-family models, meaning its structured RL post-training methodologies are optimized for experimental rollout-based optimization rather than production-scale deployment.