data-scientist

Automates and standardizes end-to-end data science workflows from data gathering to deployment.

1|Updated Apr 21, 2025
One-click install
npx skills add https://github.com/LAI-YEN-CHUN/VSCode-Settings --skill data-scientist-lai-yen-chun
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-scientist
Source: https://github.com/LAI-YEN-CHUN/VSCode-Settings/tree/main/.github/skills/data-scientist
Command: npx skills add https://github.com/LAI-YEN-CHUN/VSCode-Settings --skill data-scientist-lai-yen-chun

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Data science tasks can be complex and time-consuming, requiring a cohesive, end-to-end workflow from data preparation to model deployment; this Skill provides best practices, checklists, and structured guidance to streamline analytics projects and deliver actionable insights.

Core Features & Use Cases

  • Statistical analysis, experimental design, and causal inference for rigorous research
  • Machine learning, model selection, feature engineering, and deployment-ready pipelines
  • Data exploration, visualization, and storytelling to communicate findings
  • Domain-oriented applications across marketing, finance, operations, and product analytics
  • End-to-end workflow guidance from problem framing to production monitoring

Quick Start

Outline a complete data science workflow for a given problem, from data collection to model deployment, including goals and evaluation criteria.

Frequently Asked Questions about data-scientist

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I structure an end-to-end data science workflow from data collection to model deployment?

An end-to-end data science workflow standardizes stages from data gathering and exploratory analysis to feature engineering, model selection, evaluation, and deployment. Structured guidance provides checklists and best practices to streamline these analytics projects into reproducible pipelines.

What's the best way to perform exploratory data analysis and feature engineering for machine learning?

Exploratory data analysis and feature engineering for machine learning are best performed through standardized data exploration, visualization, and storytelling. This process communicates findings and prepares deployment-ready pipelines by applying statistical methods and best practices for reproducible research.

Do I need to know Python and statistical methods to use this data science workflow?

Yes, you need knowledge of Python, machine learning techniques, and statistical methods. The workflow requires understanding best practices for reproducible research to effectively automate data gathering, model selection, evaluation, and deployment across analytics projects.

Can I apply this data exploration and model deployment workflow to domain-specific analytics like marketing or finance?

Yes, this data exploration and model deployment workflow supports domain-oriented applications across marketing, finance, operations, and product analytics. It provides structured guidance for statistical analysis, causal inference, and experimental design tailored to rigorous domain research.

How does statistical analysis and causal inference fit into the machine learning pipeline?

Statistical analysis and causal inference fit into the machine learning pipeline by providing experimental design and rigorous research foundations. They complement model selection and feature engineering to ensure deployment-ready pipelines deliver actionable, trustworthy insights.

Why should I standardize my machine learning model selection and deployment process?

Standardizing machine learning model selection and deployment prevents fragmented analytics workflows and ensures reproducible research. It streamlines complex data science tasks into cohesive pipelines, delivering actionable insights while reducing time and errors across projects.