slime-rl-training

Automate end-to-end RL post-training workflows for GLM models with Megatron-LM and SGLang.

Updated Apr 3, 2026
One-click install
npx skills add https://github.com/handsomelong922/my-codex-skills --skill slime-rl-training-handsomelong922
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/handsomelong922/my-codex-skills/tree/main/skills/slime
Command: npx skills add https://github.com/handsomelong922/my-codex-skills --skill slime-rl-training-handsomelong922

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

slime-rl-training provides a complete guide for post-training LLMs with RL, integrating Megatron-LM and SGLang to scale data generation, rollout, and evaluation.

Core Features & Use Cases

  • Megatron-LM based training with SGLang rollout
  • High-throughput data generation and rollout orchestration
  • Broad model support (GLM, Qwen3, Llama3, DeepSeek)
  • End-to-end RL post-training workflows with evaluation and monitoring

Quick Start

Install slime-rl-training, configure your Megatron-LM + SGLang environment, and run the standard GRPO workflow to start RL-based post-training for your LLM.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I scale reinforcement learning post-training for large language models using Megatron-LM and SGLang?

Reinforcement learning post-training for large language models is scaled by integrating Megatron-LM for training and SGLang for high-throughput rollouts. This workflow automates data generation, rollout orchestration, and evaluation for models like GLM, Qwen3, and Llama3.

What prerequisites do I need to run GRPO workflows for LLM post-training?

GRPO workflows for LLM post-training require a configured Megatron-LM and SGLang environment. You also need access to training data and model configuration to orchestrate the training, rollout, and evaluation steps effectively.

How does high-throughput data generation work during LLM rollout orchestration?

High-throughput data generation during LLM rollout orchestration works by coordinating SGLang with Megatron-LM. This integration automates custom data generation and rollouts, enabling efficient end-to-end RL post-training workflows in research and production settings.

Does this RL training workflow support models outside the GLM family?

Yes, this RL training workflow supports models outside the GLM family. It provides broad model support including GLM, Qwen3, Llama3, and DeepSeek, applying Megatron-LM based training and SGLang rollout across these architectures.

What is the best way to automate end-to-end RL post-training workflows with evaluation?

The best way to automate end-to-end RL post-training workflows is using slime to orchestrate training, rollout, and evaluation. This applies custom data generation and Megatron-LM plus SGLang integration for large language models.

Can I use SGLang for rollout orchestration in production RL training settings?

Yes, you can use SGLang for rollout orchestration in production RL training settings. It integrates with Megatron-LM to support high-throughput data generation and end-to-end evaluation for large language model post-training.