What problem does it solve?
This Skill helps you review reinforcement-learning reward changes before they are approved, reducing the risk of reward hacking, misaligned objectives, and unstable training behavior.
Core Features & Use Cases
- Task Alignment Review: Checks whether a proposed reward change still matches the study objective and expected success criteria.
- Safety and Stability Screening: Flags privileged signals, guardrail conflicts, discontinuities, saturation, and other reward design risks.
- Approval Readiness Assessment: Produces a risk assessment, suggested fixes, and a recommendation for whether the change is ready to validate.
- Use Case: Use it when a teammate proposes a reward parameter tweak or reward-function edit and you need a conservative review before running experiments.
Quick Start
Ask the assistant to review your proposed reward change against the task objective, available training signals, guardrails, and approval criteria, then summarize risks and fixes.