quality-performance

Audit AICP runtime performance across profiles and hardware configurations.

Updated Mar 26, 2026
One-click install
npx skills add https://github.com/cyberpunk042/devops-expert-local-ai --skill quality-performance
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: quality-performance
Source: https://github.com/cyberpunk042/devops-expert-local-ai/tree/main/.claude/skills/quality-performance
Command: npx skills add https://github.com/cyberpunk042/devops-expert-local-ai --skill quality-performance

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Audits AICP's runtime performance across multiple models and configurations to surface latency, throughput, and stability issues affecting multi-model orchestration.

Core Features & Use Cases

  • Benchmark infrastructure verification and baseline capture across profiles and hardware configurations.
  • Per-profile latency (cold-start and warm), memory + VRAM footprint, and active backend counts.
  • Support for threshold comparisons, trend analysis, and hot-path planning for investigation.

Quick Start

Run a baseline performance audit on your AICP deployment to capture dry-run metrics.

Frequently Asked Questions about quality-performance

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark runtime performance for multi-model orchestration?

Benchmark runtime performance by auditing latency and throughput gaps across multiple profiles and hardware configurations. The audit captures cold-start, router, and profile-switch costs to identify stability issues affecting multi-model orchestration.

What latency metrics should I profile to identify cold-start bottlenecks?

Profile latency metrics by collecting per-profile timings for both cold-start and warm states. The audit enforces measurement requirements and thresholds, capturing active backend counts and memory footprints to surface cold-start bottlenecks.

Can I measure VRAM and RAM footprints across different hardware configurations?

Yes, you can measure VRAM and RAM footprints across hardware configurations. The audit spans multiple profiles to collect memory footprints and active backend counts, enforcing threshold comparisons to verify infrastructure baselines.

What's the best way to audit profile-switch costs and router latency?

Audit profile-switch costs and router latency by running a baseline performance audit on your deployment. The process enforces measurement requirements and reporting formats, collecting per-profile timings to surface throughput gaps and hot paths.

Does Docker affect runtime performance benchmarking thresholds?

Docker environments are included in runtime performance benchmarking across hardware configurations. The audit captures dry-run metrics and enforces threshold comparisons to identify latency and throughput gaps within your specific Docker deployment setup.

When do I need to run a GPU performance audit for baseline capture?

Run a GPU performance audit for baseline capture when verifying infrastructure across hardware configurations. The audit enforces measurement requirements, collecting VRAM footprints and active backend counts to support trend analysis and hot-path planning.