What problem does it solve?
This Skill helps product and ML teams move from vague data readiness to a concrete, actionable data strategy for AI features. It removes uncertainty about what data is needed, how to label and maintain it, and how to operationalize feedback loops and retraining so models remain performant and compliant.
Core Features & Use Cases
- Data inventory & audit: catalog internal sources (logs, events, content, transactions, support), identify owners, access patterns, PII, and gaps.
- Data quality & representativeness: define metrics (completeness, accuracy, timeliness, consistency) and thresholds for production readiness.
- Labeling and pipeline design: recommend human, semi-supervised, weak supervision, synthetic, and LLM-as-labeller approaches plus annotation guidelines and QA.
- Data flywheel & feedback loops: map how user actions generate training signals, estimate velocity, and design explicit/implicit feedback capture.
- Governance, versioning, and retraining: specify PII handling, consent, retention, benchmark creation, dataset/version tracking, and retraining cadence (schedule- or drift-triggered).
- Use case example: produce a prioritized plan to prepare training data, labeling, benchmarks, and retraining for a new recommendation or conversational feature.
Quick Start
Define an AI data strategy for a conversational assistant using product logs, user messages, and support tickets.