data-scientist

Document production AI/ML stack usage, LLM calls, pipelines, and cost drivers.

Updated Mar 25, 2026
One-click install
npx skills add https://github.com/ouakar/ubinarys-dental --skill data-scientist-ouakar
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-scientist
Source: https://github.com/ouakar/ubinarys-dental/tree/main/skills/forgewright/skills/data-scientist
Command: npx skills add https://github.com/ouakar/ubinarys-dental --skill data-scientist-ouakar

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Production-grade internal AI/ML systems require cohesive design, rigorous optimization, and measurable ROI across LLM usage, data pipelines, and cost.

Core Features & Use Cases

  • LLM optimization: cost, latency, and quality improvements for production prompts and pipelines.
  • RAG pipeline design: end-to-end retrieval-augmented generation with vector stores and caching.
  • Vector database architecture: scalable storage and retrieval for embeddings and features.
  • AI agent orchestration: multi-agent coordination for automated workflows and decision making.
  • ML pipeline management: end-to-end training, deployment, monitoring, and retraining.
  • Evaluation frameworks & cost modeling: metrics, experiments, and ROI calculation.
  • Use Case: Deploy a production-grade AI system that reduces manual intervention and improves decision speed.

Quick Start

Initialize a baseline production AI stack, run the audit, and implement the top-priority optimization plan.

Frequently Asked Questions about data-scientist

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I optimize LLM costs and latency in production AI pipelines?

End-to-end RAG pipeline design requires integrating vector databases for scalable embedding storage and retrieval, applying robust caching mechanisms, and establishing evaluation frameworks to measure generation quality and ROI.

Can I use vector databases for scalable embedding storage and retrieval in ML pipelines?

Yes, vector databases provide scalable storage and retrieval for embeddings and features within ML pipelines, enabling efficient retrieval-augmented generation and supporting production-grade reliability for AI systems.

What is the best way to orchestrate multi-agent AI workflows for automated decision making?

The best way to orchestrate multi-agent AI workflows is to establish a coordinated agent orchestration framework that automates complex workflows, reduces manual intervention, and improves decision speed across production systems.

How do I establish an evaluation framework for production ML models and calculate ROI?

Establish an evaluation framework for production ML models by defining specific metrics and running experiments to measure performance, which directly enables concrete ROI calculation and cost modeling for your AI stack.