results-analysis

Analyze ML experiment artifacts with descriptive and inferential statistics to generate analysis reports.

Updated Mar 27, 2026
One-click install
npx skills add https://github.com/EmaRimoldi/Claude-scholar-extended --skill results-analysis-emarimoldi
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: results-analysis
Source: https://github.com/EmaRimoldi/Claude-scholar-extended/tree/main/skills/results-analysis
Command: npx skills add https://github.com/EmaRimoldi/Claude-scholar-extended --skill results-analysis-emarimoldi

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill enables researchers to produce a strict, evidence-first analysis bundle for ML/AI experimental results, replacing ad-hoc summaries with verifiable statistics and artifacts.

Core Features & Use Cases

  • Validates experiment artifacts and performs descriptive and inferential statistics.
  • Generates real scientific figures and a complete analysis footprint, including an analysis-report, a stats appendix, and a figure catalog.
  • Produces ready-to-use outputs for results reporting and for handoff to subsequent skills such as results-report.

Quick Start

Analyze your experiment artifacts to produce a complete, publication-ready analysis bundle.

Frequently Asked Questions about results-analysis

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I analyze experimental results with reproducible statistical methods?

Statistical analysis for multi-seed experiments involves validating artifacts and applying descriptive and inferential statistics to compare baselines and ablations. This generates verifiable statistics, real scientific figures, and an analysis footprint with explicit caveats and blockers.

What's the best way to generate publication-quality figures from ML ablation results?

Generating publication-quality figures requires applying descriptive and inferential statistics to validated multi-seed experiment artifacts. This produces real scientific figures, a figure catalog, and a complete analysis footprint with explicit caveats and blockers for rigorous reporting.

How do I validate experiment artifacts before running inferential statistics?

Validating experiment artifacts requires checking multi-seed experiment data, baseline comparisons, and ablation records before applying inferential statistics. This validation step ensures reproducibility and generates a strict, evidence-first analysis bundle with verifiable statistics.

Can I use this approach for multi-seed experiments comparing baselines and ablations?

Yes, multi-seed experiments comparing baselines and ablations are explicitly supported. The analysis applies strict descriptive and inferential statistics to validated artifacts, producing an analysis report, stats appendix, and figure catalog with explicit caveats and blockers.

What outputs do I get from a complete statistical analysis of experimental results?

A complete statistical analysis produces three main artifacts: analysis-report.md, stats-appendix.md, and figure-catalog.md. These outputs provide verifiable descriptive and inferential statistics, real scientific figures, and explicit caveats for publication-ready reporting and downstream handoff.

Why should I replace ad-hoc experiment summaries with verifiable statistics?

Replacing ad-hoc summaries with verifiable statistics ensures reproducibility and evidence-first experimental reporting. Applying strict descriptive and inferential statistics to validated artifacts generates a complete analysis footprint with explicit caveats and blockers, preventing misleading conclusions.