slime-rl-training

Implement end-to-end RL post-training workflows for LLMs with slime.

Updated May 2, 2026
One-click install
npx skills add https://github.com/AlvaroBiano/hermes-agent --skill slime-rl-training-alvarobiano
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/AlvaroBiano/hermes-agent/tree/main/optional-skills/mlops/slime
Command: npx skills add https://github.com/AlvaroBiano/hermes-agent --skill slime-rl-training-alvarobiano

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires sglang-router>=0.2.3, ray, torch>=2.0.0, transformers>=4.40.0, and includes references (resource) components.

What problem does it solve?

Solves the challenge of implementing and scaling RL-based post-training workflows for large language models using slime.

Core Features & Use Cases

  • Comprehensive guidance on slime integration with Megatron-LM, data generation workflows, and RL scaling for GLMs.
  • Best practices for configuring slime workflows, data buffers, and evaluation in production-style experiments.
  • Use Case: Researchers and engineers building robust RL post-training pipelines for GLM family models.

Quick Start

Run slime to begin RL-based post-training on your GLM model.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I implement reinforcement learning post-training for large language models using slime?

Reinforcement learning post-training for LLMs with slime involves configuring end-to-end workflows, from model setup to data buffering and evaluation. This Skill provides actionable guidance, best practices, and configurations for scaling slime workflows.

Can I integrate slime with Megatron-LM for RL training on GLM-family models?

Yes, slime integrates with Megatron-LM for RL training on GLM-family models. The Skill provides comprehensive guidance on this integration, enabling researchers and engineers to build robust post-training pipelines.

What dependencies are required to run slime workflows for RL post-training?

Running slime workflows requires sglang-router, ray, torch, and transformers. These dependencies support high-throughput data generation and reinforcement learning scaling for large language models.

How does high-throughput data generation work in slime RL training workflows?

High-throughput data generation in slime workflows involves configuring data buffers and generation pipelines. The Skill offers best practices for setting up these production-style experiments to ensure efficient RL post-training.

What are the best practices for configuring data buffers and evaluation in slime workflows?

Best practices for configuring data buffers and evaluation in slime workflows include setting up production-style experiments and applying actionable guidance. This ensures robust RL post-training pipelines for GLM family models.

Is slime suitable for researchers building RL post-training pipelines for GLM models?

Yes, slime is suitable for researchers and engineers building robust RL post-training pipelines for GLM family models. It provides comprehensive guidance on slime integration with Megatron-LM and RL scaling.