ai-observability

Analyze AI model performance, GPU utilization, and OpenShift cluster health.

48|31|Updated Feb 2, 2026
One-click install
npx skills add https://github.com/RHEcosystemAppEng/agentic-plugins --skill ai-observability-rhecosystemappeng
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-observability
Source: https://github.com/RHEcosystemAppEng/agentic-plugins/tree/main/rh-ai-engineer/skills/ai-observability
Command: npx skills add https://github.com/RHEcosystemAppEng/agentic-plugins --skill ai-observability-rhecosystemappeng

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires ai-observability, rhoai, openshift, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill helps you analyze AI model performance, GPU utilization, and OpenShift cluster health, enabling you to monitor and optimize your AI infrastructure.

Core Features & Use Cases

  • Model Performance Analysis: Check the performance of AI models, including latency, throughput, and error rates.
  • GPU Utilization: Monitor GPU inventory and utilization across the cluster.
  • Cluster Health: Analyze OpenShift cluster health metrics by category.
  • Tracing: Trace slow inference requests with distributed tracing.
  • Correlation: Correlate signals across logs, metrics, traces, and alerts.
  • Custom PromQL Queries: Run custom PromQL queries against cluster Prometheus.
  • Use Case: If you need to understand the performance of a specific AI model or the health of your OpenShift cluster, this Skill can provide valuable insights.

Quick Start

Use the ai-observability skill to analyze the performance of a specific AI model in your OpenShift cluster.

Frequently Asked Questions about ai-observability

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I monitor AI model performance and GPU utilization on an OpenShift cluster?

You can monitor AI model performance and GPU utilization on an OpenShift cluster by analyzing latency, throughput, error rates, and GPU inventory to optimize your overall AI infrastructure.

What is the best way to trace slow inference requests in OpenShift AI deployments?

Tracing slow inference requests in OpenShift AI deployments is done using distributed tracing to identify bottlenecks in your AI models and correlate them with cluster health metrics.

Can I run custom PromQL queries against Prometheus to analyze OpenShift cluster health?

Yes, you can run custom PromQL queries against cluster Prometheus to analyze OpenShift cluster health metrics by category and correlate signals across logs, metrics, traces, and alerts.

Do I need access to the AI Observability MCP server to correlate signals across logs and metrics?

Yes, you need access to the AI Observability MCP server and related tools to correlate signals across logs, metrics, traces, and alerts for comprehensive infrastructure monitoring.

Why should I use korrel8r to correlate signals for AI observability?

Korrel8r correlates signals across logs, metrics, traces, and alerts to provide valuable insights into AI model performance and OpenShift cluster health, enabling infrastructure optimization.