grpo-rl-training
Fine-tune language models with GRPO-based RL using TRL and structured reward signals.
npx skills add https://github.com/tangzheng202202/hermes-skills --skill grpo-rl-training-tangzheng202202
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill. Skill: grpo-rl-training Source: https://github.com/tangzheng202202/hermes-skills/tree/main/03-mlops/mlops/training/grpo-rl-training Command: npx skills add https://github.com/tangzheng202202/hermes-skills --skill grpo-rl-training-tangzheng202202