rloo
Reduce gradient variance in reinforcement learning policy optimization with leave-one-out baselines.
npx skills add https://github.com/atrawog/overthink-plugins --skill rloo
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill. Skill: rloo Source: https://github.com/atrawog/overthink-plugins/tree/main/overthink-jupyter/skills/rloo Command: npx skills add https://github.com/atrawog/overthink-plugins --skill rloo