the-crypt-keeperthe-crypt-keeperCommunityยท2 Agent Skills Included

reasonscape

Parametric evaluation and forensic analysis of reasoning models

Generates parametric reasoning tests, runs model evaluations at scale, and analyzes results with statistical rigor. Replaces single leaderboard scores with capability shapes showing where models break, at what token cost, and why. Imports new model results and simulates lower context windows to compare performance across configurations.
npx skills add the-crypt-keeper/reasonscape --all -g -y
Available:

Instructs the agent on the Position-Profile-Probe research methodology, statistical verification norms, and the five-stage evaluation pipeline for analyzing reasoning model behavior.

All Skills in This Repository (2)

Pure Emerald Level Indicators

Frequently Asked Questions

FAQPage Schema
How to install ReasonScape?โ–ผ

Run `npx skills add the-crypt-keeper/reasonscape --all -g -y` in your terminal to install all skills in this suite globally.

What does ReasonScape measure?โ–ผ

It measures how reasoning models process information using parametric test generation, tracking accuracy, token cost, truncation, and reasoning quality instead of a single benchmark score.

How to compare models at different context lengths?โ–ผ

Use the context-clip skill to create context-limited copies of an evaluation, such as 8k or 12k versions, then compare performance across them with the analysis commands.

How to add a new model to the r12 dataset?โ–ผ

Use the import skill, which finds your result folders, fetches model metadata, builds the cohort entry, and verifies the import before evaluation.

Can I explore results without running evaluations?โ–ผ

Yes. You can download the public r12 dataset covering 67 models and over 116,000 evaluation points, then explore it with the built-in leaderboard and analysis tools.

Related Repositories in Data & Analytics

View All in Data & Analyticsโ†’