llm-architect

Design production-grade LLM systems for scalable, safe deployment.

Updated Apr 27, 2026
One-click install
npx skills add https://github.com/Tnemo65/template --skill llm-architect-tnemo65
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: llm-architect
Source: https://github.com/Tnemo65/template/tree/main/.cursor/skills/11-system-design/llm-architect
Command: npx skills add https://github.com/Tnemo65/template --skill llm-architect-tnemo65

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Guides teams in designing, implementing, and operating production-grade LLM systems with strong performance, safety, and scalability.

Core Features & Use Cases

  • Architecture planning for serving patterns, multi-model routing, and monitoring
  • Fine-tuning, RAG integration, and cost-aware serving optimizations
  • Safety, compliance, and auditing through checklists and governance
  • End-to-end workflow from design to validation and deployment

Quick Start

Provide a scalable LLM system blueprint based on your current stack and latency targets.

Frequently Asked Questions about llm-architect

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a scalable LLM serving architecture for enterprise workloads?

Design scalable LLM serving architecture by defining serving patterns, multi-model routing, and monitoring workflows to achieve target latency and throughput across enterprise environments.

What's the best way to integrate RAG into a production LLM system?

Integrate RAG into a production LLM system through structured integration workflows that combine retrieval pipelines with cost-aware serving optimizations to maintain performance and accuracy.

How do I set up monitoring and safety protocols for LLM deployment?

Set up LLM deployment monitoring and safety protocols by applying architecture checklists, governance auditing, and compliance validation to ensure safe and observable production operations.

Can I use multi-model orchestration to route requests across different LLMs?

Multi-model orchestration routes requests across different LLMs by applying architecture planning patterns that balance cost, latency, and performance targets for enterprise serving environments.

What do I need to plan before fine-tuning and deploying an LLM system?

Plan LLM system fine-tuning and deployment by preparing architecture blueprints, defining performance targets, and establishing safety checklists to validate end-to-end production workflows.