slime-rl-training

Automates scalable RL post-training for LLMs with Megatron-LM and SGLang.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/AVOI-CEO/avoi-agent --skill slime-rl-training-avoi-ceo
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/AVOI-CEO/avoi-agent/tree/main/optional-skills/mlops/slime
Command: npx skills add https://github.com/AVOI-CEO/avoi-agent --skill slime-rl-training-avoi-ceo

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires sglang-router>=0.2.3, ray, torch>=2.0.0, transformers>=4.40.0, and includes references (resource) components.

What problem does it solve?

Slime provides a scalable RL post-training workflow for large language models by integrating Megatron-LM training with SGLang rollout and a modular data buffer to streamline experiment management.

Core Features & Use Cases

  • Megatron-LM integrated training with SGLang rollout for high-throughput RL experiments
  • Multi-turn agentic training with tool use and RL feedback loops
  • Flexible data buffering, configuration management, and reference material for reproducibility
  • Use Case: Researchers can run end-to-end RL post-training for GLMs such as SLIME-based setups

Quick Start

Launch a slime RL training run using your configured data and model settings.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I scale RL post-training for large language models using Megatron-LM?

Scalable RL post-training for large language models is achieved by connecting Megatron-LM training with SGLang-driven rollout and a modular data buffer to streamline high-throughput experiment management.

What is the best way to run multi-turn agentic training with tool use and RL feedback loops?

Multi-turn agentic training with tool use and RL feedback loops is automated through pre-configured model scripts that integrate high-throughput data generation with flexible data buffering.

Does SGLang rollout work with Megatron-LM for high-throughput data generation?

Yes, SGLang rollout integrates directly with Megatron-LM training to provide high-throughput data generation, satisfying requirements for model scripting and data buffering in RL pipelines.

Can I use this RL training pipeline for GLM-style models in a research setting?

Yes, the RL training pipeline is ideal for GLM-style models and related architectures, providing researchers with pre-configured scripts and reference material for reproducible experiments.

Do I need Ray and PyTorch to configure scalable RL post-training workflows?

Yes, scalable RL post-training workflows require Ray and PyTorch along with Transformers and SGLang router to automate data buffering, configuration management, and model scripting.