slime-rl-training

Orchestrates RL post-training for GLM models with Megatron-LM and SGLang.

78|16|Updated Apr 23, 2026
One-click install
npx skills add https://github.com/sheawinkler/hermes-agent-ultra --skill slime-rl-training-sheawinkler
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/sheawinkler/hermes-agent-ultra/tree/main/optional-skills/mlops/slime
Command: npx skills add https://github.com/sheawinkler/hermes-agent-ultra --skill slime-rl-training-sheawinkler

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

slime-rl-training provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.

Core Features & Use Cases

  • RL-oriented post-training orchestration for LLMs
  • Megatron-LM + SGLang integrated training and rollout
  • Data buffering, reward modeling, and evaluation workflows for production RL

Quick Start

Run the slime-rl-training workflow using your Megatron-LM and SGLang setup to begin RL post-training for GLM models.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I orchestrate RL post-training for large language models?

To orchestrate RL post-training for LLMs, use the slime framework which provides structured guidance for data generation, rollout routing, reward modeling, and policy enforcement. It integrates Megatron-LM and SGLang to handle GLM-scale model training workflows.

Can I use Megatron-LM and SGLang together for LLM rollout and training?

Yes, Megatron-LM and SGLang can be used together for integrated training and rollout. The slime framework specifically combines them to enable data buffering, custom generate functions, and evaluation routines within a production RL training stack.

What is needed to configure Megatron-LM for reinforcement learning workflows?

Configuring Megatron-LM for RL workflows requires setting up data generation, rollout routing, and reward modeling pipelines. The slime framework satisfies these requirements by providing structured tooling for custom generate functions and evaluation routines across Linux and macOS environments.

Does RL post-training with slime support macOS environments?

Yes, RL post-training with slime supports macOS environments. The framework applies data generation, rollout routing, reward modeling, and policy enforcement across both Linux and macOS systems for GLM-scale models.

How do I implement custom data generation workflows for RL training?

To implement custom data generation workflows for RL training, use the slime framework's structured tooling. It provides data buffering, custom generate functions, and evaluation routines that integrate with Megatron-LM and SGLang pipelines for production-grade post-training.