What problem does it solve?
Summarize observational signals about how skills perform across contributor sessions so teams can prioritize transcript review and recommend improvements without mistaking correlation for causation.
Core Features & Use Cases
- Read and parse session metrics to extract per-skill effectiveness signals, sample sizes, and outcome distributions for exploratory monitoring.
- Provide provider- and time-window scoping, per-skill filtering, low-sample confidence warnings, and dominant outcome summaries to surface candidates for manual review.
- Support an improvement mode that delegates to a skill-effectiveness-analyzer, cite session evidence and confounders, and write dashboard snapshots under .claude/skill-metrics for historical tracking.
- Use cases include triaging which skills to inspect in transcript review, corroborating docs-check or lab/eval results, and generating prioritized, evidence-framed recommendations.
Quick Start
Run the skill to generate an observational dashboard from .claude/session-metrics/metrics.jsonl for a chosen time window, provider, or specific skill.