llm-architect

Design and optimize end-to-end LLM systems for production deployment.

8|11|Updated Feb 15, 2026
One-click install
npx skills add https://github.com/belokonm/claude-supercode-skills --skill llm-architect-belokonm
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: llm-architect
Source: https://github.com/belokonm/claude-supercode-skills/tree/main/llm-architect-skill
Command: npx skills add https://github.com/belokonm/claude-supercode-skills --skill llm-architect-belokonm

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires pyyaml, rouge-score, nltk, transformers, peft, datasets, fastapi, uvicorn, torch, chromadb, sentence-transformers, and includes scripts (resource) and references (resource) components.

What problem does it solve?

The LLM Architect helps teams design, evaluate, and deploy production-grade large language model systems with strong guardrails, monitoring, and cost optimization.

Core Features & Use Cases

  • Architecture guidance for model selection, deployment strategies, RAG pipelines, and performance optimization.
  • Safety, compliance, and governance patterns across scalable LLM deployments.
  • Real-world scenarios: building chat assistants, code assistants, and enterprise QA systems with reliable latency and observability.

Quick Start

Outline an end-to-end LLM system architecture for production deployment.

Frequently Asked Questions about llm-architect

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a scalable LLM architecture for production deployment?

To design scalable LLM architecture for production, you must outline model selection, RAG pipelines, deployment strategies, and cost optimization. The system enforces safety guardrails, evaluation, and monitoring for reliable, maintainable large language model deployments.

What is the best way to optimize RAG pipelines for enterprise QA systems?

Optimizing RAG pipelines for enterprise QA systems involves integrating retrieval frameworks with evaluation and safety guardrails. This architecture ensures reliable latency and observability across scalable large language model deployments.

Can I use FastAPI and ChromaDB to build production LLM systems?

Yes, you can use FastAPI and ChromaDB to build production LLM systems. The architecture supports integrating these frameworks to enforce guardrails, manage vector retrieval, and maintain reproducible workflows for scalable large language model deployments.

How do I implement safety guardrails and compliance patterns for LLM deployments?

To implement safety guardrails and compliance patterns for LLM deployments, apply governance rules across scalable architectures. The system enforces guardrails, continuous evaluation, and monitoring to maintain reliable, maintainable large language model operations.

How do I reduce inference costs when deploying large language models?

To reduce inference costs when deploying large language models, apply cost optimization strategies within the system architecture. This involves selecting efficient models and enforcing reproducible workflows to maintain performance while lowering operational expenses.