data-scientist

Automate data analysis and predictive modeling with Pandas, NumPy, Scikit-learn, and XGBoost.

Updated May 4, 2026
One-click install
npx skills add https://github.com/luokai25/luo-ai-skills-market --skill data-scientist-luokai25
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-scientist
Source: https://github.com/luokai25/luo-ai-skills-market/tree/main/09-data-and-ai%20%28by%20Luo%20Kai%29/14-other-ai/data-scientist
Command: npx skills add https://github.com/luokai25/luo-ai-skills-market --skill data-scientist-luokai25

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires pandas, numpy, scikit-learn, xgboost, statsmodels, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates complex data analysis and modeling tasks, helping you uncover insights and make data-driven decisions more efficiently.

Core Features & Use Cases

  • Exploratory Analysis: Profiling, distribution analysis, correlation studies, and outlier detection.
  • Statistical Modeling: Hypothesis testing, regression analysis, time series modeling, and survival analysis.
  • Machine Learning: Problem formulation, feature engineering, algorithm selection, and model training.
  • Feature Engineering: Application of domain knowledge, transformation techniques, and dimensionality reduction.
  • Model Evaluation: Performance metrics, validation strategies, and bias detection.
  • Use Case: Use this Skill to analyze customer purchasing patterns, predict future sales, or identify trends in customer feedback.

Quick Start

Use the data-scientist skill to perform exploratory analysis on the 'customer_data.csv' file.

Frequently Asked Questions about data-scientist

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate exploratory data analysis and statistical modeling in Python?

Automate exploratory data analysis and statistical modeling using this Skill to perform profiling, distribution analysis, and regression. It leverages Pandas and Statsmodels to handle data profiling, hypothesis testing, and correlation studies for business intelligence tasks.

Can I use Scikit-learn and XGBoost for predictive modeling tasks?

Yes, you can use Scikit-learn and XGBoost for predictive modeling. This Skill supports problem formulation, feature engineering, algorithm selection, and model training to predict future sales or analyze customer purchasing patterns.

What's the best way to perform feature engineering and model evaluation?

The best way to perform feature engineering and model evaluation is through transformation techniques, dimensionality reduction, and validation strategies. This Skill applies domain knowledge to engineer features and uses performance metrics to detect model bias.

Does this data analysis approach support time series modeling and survival analysis?

Yes, this data analysis approach supports time series modeling and survival analysis. It provides statistical modeling capabilities for hypothesis testing and regression analysis, enabling effective market analysis and research data evaluation.

How do I generate data visualizations and detect outliers in my dataset?

Generate data visualizations and detect outliers by running exploratory analysis on your dataset. This Skill automates distribution analysis and correlation studies to identify trends in customer feedback and uncover actionable insights.