senior-data-scientist

Design and run production-grade ML experiments and analytics workflows.

4|Updated Feb 22, 2026
One-click install
npx skills add https://github.com/GeneralReasoning/env-skillsbench --skill senior-data-scientist-generalreasoning
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: senior-data-scientist
Source: https://github.com/GeneralReasoning/env-skillsbench/tree/main/powerlifting-coef-calc/environment/skills/senior-data-scientist
Command: npx skills add https://github.com/GeneralReasoning/env-skillsbench --skill senior-data-scientist-generalreasoning

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

Senior data scientists need to design and run production-grade experiments and analytics workflows that yield trustworthy, reproducible insights for business decisions.

Core Features & Use Cases

  • Experiment design and A/B testing planning across Python/R ecosystems, including hypothesis specification, sample size estimation, and power analysis.
  • Feature engineering, model evaluation, deployment-ready pipelines, and reproducibility with traceability.
  • Causal inference, time-series analysis, and BI-ready reporting for stakeholders.

Quick Start

Provide a dataset and objective, and this skill will generate a complete experimental design and analytics plan ready for production.

Frequently Asked Questions about senior-data-scientist

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design and run production-grade A/B testing for ML systems?

Production-grade A/B testing for ML systems requires structured experimental design, hypothesis specification, and sample size estimation. This skill generates complete testing plans with power analysis to ensure reproducible, scalable results across Python and R environments.

What is the best way to apply causal inference in data science workflows?

Applying causal inference in data science workflows involves isolating treatment effects from confounding variables. This skill integrates causal inference techniques with time-series analysis to produce trustworthy insights for business decisions and stakeholder reporting.

How do I build deployment-ready pipelines for model evaluation?

Building deployment-ready pipelines for model evaluation requires rigorous feature engineering, reproducibility, and traceability. This skill designs scalable evaluation workflows that maintain governance standards while preparing models for production ML environments.

Can I use this skill for time-series analysis and BI reporting in SQL?

Yes, this skill supports time-series analysis and BI-ready reporting across Python, R, and SQL environments. It formats analytical outputs specifically for stakeholder communication, ensuring data-driven initiatives translate into clear business intelligence.

When do I need power analysis and sample size estimation for experimentation?

Power analysis and sample size estimation are needed before running experiments to ensure statistical validity and avoid inconclusive results. This skill calculates these metrics during experimental design to guarantee your A/B tests have sufficient statistical power.

How do I ensure reproducibility and governance in ML analytics workflows?

Ensuring reproducibility and governance in ML analytics workflows requires traceable feature engineering, documented experimental design, and scalable pipelines. This skill structures your workflows to satisfy governance requirements while maintaining analytical reproducibility.