honest-ship-eval

Execute multi-suite adversarial and structural evaluation gates for SLM ship readiness.

1|Updated Jul 12, 2026
One-click install
npx skills add https://github.com/Tyler-R-Kendrick/slm-training --skill honest-ship-eval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: honest-ship-eval
Source: https://github.com/Tyler-R-Kendrick/slm-training/tree/main/.agents/skills/honest-ship-eval
Command: npx skills add https://github.com/Tyler-R-Kendrick/slm-training --skill honest-ship-eval

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill addresses the critical need for objective, non-negotiable ship readiness validation, preventing the premature deployment of models that rely on soft metrics or hidden data leakage.

Core Features & Use Cases

  • Multi-Suite Validation: Executes comprehensive evaluation across smoke, held-out, adversarial, OOD, and rico_held suites.
  • Honesty-Constrained Gates: Enforces strict parsing and fidelity thresholds to ensure production claims are backed by verifiable evidence.
  • Use Case: Use this tool when preparing a model for production to ensure it meets the rigorous structural and semantic density requirements defined in the project's goal law.

Quick Start

Run the honest-ship-eval skill to execute the full ship-gates evaluation suite against your current model run.

Frequently Asked Questions about honest-ship-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I validate model readiness before shipping to production?

Model readiness validation requires executing multi-suite adversarial evaluation gates across smoke, held-out, and OOD suites. This ensures production claims are grounded in verifiable parse, fidelity, and component recall metrics rather than soft metrics.

What is adversarial testing for model deployment gates?

Adversarial testing for deployment gates involves running rigorous evaluation against out-of-distribution and held-out suites to prevent premature deployment. It enforces strict parsing and fidelity thresholds to ensure claims are backed by verifiable evidence.

How do I set up constrained decoding evaluation for SLM training outcomes?

Constrained decoding evaluation requires strict adherence to decoding invariants and version-stamped evidence contracts during the assessment process. This validates structural and semantic density requirements defined in the project's goal law.

Can I use automated ship gates to prevent premature model deployment?

Automated ship gates enforce non-negotiable readiness validation by applying multi-suite structural and adversarial testing. They prevent models relying on hidden data leakage from passing by requiring strict adherence to fidelity thresholds.

What are the limitations of soft metrics in model quality assurance?

Soft metrics lack the verifiable evidence needed for production readiness, allowing premature deployment of models with hidden data leakage. Rigorous evaluation gates replace them with objective parse, fidelity, and component recall metrics.

Does honest-ship-eval work without external dependencies for quality assurance?

Honest-ship-eval operates as a standalone script component without external dependencies to validate ship readiness. It executes the full evaluation suite against your current model run to generate version-stamped evidence contracts.