What problem does it solve?
This Skill helps you monitor the performance, cost, and quality of your Large Language Models (LLMs) and agentic applications by analyzing data ingested into Elastic.
Core Features & Use Cases
- Performance Monitoring: Track latency, throughput, and error rates for LLM operations.
- Cost & Token Tracking: Analyze token usage and estimate costs associated with LLM calls.
- Response Quality: Identify issues related to response quality, content filtering, and errors.
- Workflow Orchestration: Analyze call chaining and agentic workflows to understand execution flow and identify bottlenecks.
- Use Case: You can use this Skill to answer questions like "What is the average token usage for our OpenAI calls yesterday?" or "Are there any LLM operations experiencing high latency or error rates?"
Quick Start
Use the observability-llm-obs skill to find the total input tokens used by the 'gpt-4' model in the last 24 hours.