slime-rl-training

Coordinate RL-based post-training for LLMs using slime's Megatron-LM and SGLang integration.

Updated Apr 12, 2026
One-click install
npx skills add https://github.com/thisismynewfmail-ui/Monika-agent --skill slime-rl-training-thisismynewfmail-ui
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/thisismynewfmail-ui/Monika-agent/tree/main/optional-skills/mlops/slime
Command: npx skills add https://github.com/thisismynewfmail-ui/Monika-agent --skill slime-rl-training-thisismynewfmail-ui

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

RL-based post-training for large language models requires coordinating model training, rollout generation, and reward-based optimization. This skill guides users through building scalable RL workflows with slime (Megatron-LM + SGLang) to streamline post-training pipelines.

Core Features & Use Cases

  • End-to-end RL post-training setup for GLM-like models
  • Flexible data generation and rollout orchestration with Megatron-LM and SGLang
  • Tools for monitoring, evaluation, and iterative improvement in production RL tasks

Quick Start

Train a RL-based post-training workflow for your GLM models using slime to coordinate Megatron-LM and SGLang.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I scale RL post-training for LLMs using Megatron-LM?

Scale RL post-training for LLMs by using slime to coordinate Megatron-LM training and SGLang rollout generation, enabling modular configuration and data buffering for production-scale reward-based optimization.

What is the best way to set up end-to-end RL workflows for GLM models?

Set up end-to-end RL workflows for GLM models by configuring slime's Megatron-LM and SGLang integration to orchestrate rollout generation, data buffering, and reward-based optimization.

Can I use slime for custom data generation workflows in reinforcement learning?

Yes, slime supports custom data generation workflows in reinforcement learning by providing flexible rollout orchestration and data buffering to coordinate generation and training pipelines.

Does slime RL post-training work with architectures beyond GLM-family models?

Slime RL post-training applies to GLM-family models and related architectures, coordinating Megatron-LM scale training and SGLang integration to support compatible reinforcement learning workflows.

What components are needed to configure modular RL training with slime?

Configuring modular RL training with slime requires setting up Megatron-LM and SGLang integration, utilizing data buffering and optional reference components to enable end-to-end post-training workflows.