rlxp-audit-results

Audit reinforcement-learning improvement claims for validity across RLXP experiments.

1|Updated May 14, 2026
One-click install
npx skills add https://github.com/junhyekh/rlxp --skill rlxp-audit-results
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: rlxp-audit-results
Source: https://github.com/junhyekh/rlxp/tree/main/plugins/rl-experiment-assistant/skills/rlxp-audit-results
Command: npx skills add https://github.com/junhyekh/rlxp --skill rlxp-audit-results

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps you decide whether an RLXP improvement claim is trustworthy before you promote a candidate or accept an autoloop result.

Core Features & Use Cases

  • Checks task and study identity so results are matched to the correct contract and scope.
  • Verifies metric provenance, held-out evaluation, checkpoint selection, guardrails, and output consistency.
  • Use it when reviewing a new experiment, validating an autoloop proposal, or rejecting a claim that relies only on training reward.

Quick Start

Use the rlxp-audit-results skill to audit the attached candidate, contract, metrics, run summary, and report claim for promotion readiness.

Frequently Asked Questions about rlxp-audit-results

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I validate reinforcement learning improvement claims before promoting a candidate model?

Reinforcement learning audit checks verify that evaluation protocols remain unchanged, held-out splits are properly maintained, checkpoint selection is valid, guardrails are enforced, and training-reward substitution is absent to confirm claim validity.

How do I audit autoloop outputs for task and study identity consistency?

Rejecting reinforcement learning claims based solely on training reward is necessary because training-reward substitution does not reflect true held-out evaluation performance, making the promotion claim invalid without proper metric provenance verification.

What is the best way to verify metric provenance in a reinforcement learning experiment audit?

To audit a reinforcement learning experiment for promotion readiness, attach the candidate, contract, metrics, run summary, and report claim, then verify task identity, metric provenance, held-out splits, checkpoint selection, and guardrails.

Can I use an experiment audit to check if held-out splits and evaluation protocols were changed?

A reinforcement learning experiment audit requires the candidate model, contract, metrics, run summary, and report claim to verify task identity, metric provenance, checkpoint selection, guardrails, and the absence of training-reward substitution.