slime-rl-training

Guides RL post-training of large language models using Megatron-LM and SGLang integration.

Updated Apr 10, 2026
One-click install
npx skills add https://github.com/overviewlabs/WHOX --skill slime-rl-training-overviewlabs
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/overviewlabs/WHOX/tree/main/optional-skills/mlops/slime
Command: npx skills add https://github.com/overviewlabs/WHOX --skill slime-rl-training-overviewlabs

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

slime-rl-training provides structured guidance for performing reinforcement learning-based post-training of large language models using the slime framework, enabling efficient adaptation of GLMs with RL signals.

Core Features & Use Cases

  • End-to-end RL post-training guidance for Megatron-LM + SGLang integration on GLM-family models.
  • Supports custom data generation workflows, rollout orchestration, and tool/use-case integration for scalable RL.
  • Provides model configuration presets, data handling patterns, and evaluation hooks for research and production workflows.

Quick Start

Install slime and follow the configuration steps to begin RL post-training with Megatron-LM and SGLang.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up reinforcement learning post-training for large language models using Megatron-LM and SGLang?

Reinforcement learning post-training for large language models using Megatron-LM and SGLang is set up by configuring slime's rollout orchestration, data buffers, and model configuration presets to enable scalable RL training integration.

Can I use slime for RL training on GLM family models and similar architectures?

Yes, slime supports RL training for GLM family models and similar architectures, applying Megatron-LM and SGLang integration to adapt large language models efficiently with RL signals.

What is the best way to orchestrate rollouts and custom data generation workflows for scalable RL training?

The best way to orchestrate rollouts and custom data generation for scalable RL training is using slime's custom generate functions and data handling patterns to manage data buffers and rollout workflows.

Does slime provide evaluation hooks and model script integration for end-to-end RL post-training workflows?

Yes, slime provides evaluation hooks and model script integration to satisfy end-to-end RL post-training workflow requirements, covering model configuration scripts, rollout orchestration, and data generation for research and production environments.

Why use slime's Megatron-LM and SGLang integration for reinforcement learning instead of other post-training approaches?

Slime's Megatron-LM and SGLang integration provides structured guidance for scalable RL training, enabling efficient adaptation of GLMs with custom data generation and tool use integration for academic and production environments.

When do I need slime for reinforcement learning post-training workflows?

You need slime for reinforcement learning post-training workflows when scaling RL training for large language models requires end-to-end orchestration, including model configuration, data buffers, custom generation, and evaluation hooks.