SYSTEM DOCUMENTATION & REQUIREMENTS
💡 This Skill requires pandas, numpy, scipy, matplotlib, seaborn, sklearn, xgboost, lightgbm, catboost, tensorflow, pytorch, keras, huggingface, spark, dask, ray, vaex, plotly, bokeh, altair, streamlit, mlflow, weights-biases, neptune, dvc, joblib, mlflow.sklearn, scipy.stats, sklearn.ensemble, sklearn.metrics, sklearn.pipeline, sklearn.model_selection, sklearn.preprocessing, plotly.express, plotly.graph_objects, plotly.subplots, and includes scripts (resource) and references (resource) components.
What problem does it solve?
This Skill addresses the challenge of extracting meaningful insights from complex datasets, building predictive models, and implementing machine learning solutions for data-driven decision-making.
Core Features & Use Cases
- Exploratory Data Analysis (EDA): Perform in-depth analysis, identify patterns, and understand data distributions.
- Machine Learning Pipeline Development: Build, train, and evaluate various ML models for classification and regression tasks.
- Data Visualization: Create interactive plots and dashboards for clear communication of findings.
- Production ML Systems: Deploy, monitor, and manage ML models in production environments.
- Use Case: Analyze customer transaction data to build a model that predicts customer churn, enabling proactive retention strategies.
Quick Start
Use the data-scientist skill to perform a comprehensive EDA on the provided dataset.