Toby
Community@bagelhole · NY
Cloud | Software | Security Engineer
Agent Skills by Toby
Showing 152 vetted skills indexed across 1 GitHub repositories.
semantic-versioning
Automate version bumps and changelog generation from Conventional Commits.
git-workflow
Implement Git branching strategies, pull request workflows, and automated release management patterns.
feature-flags
Implement feature flag systems for progressive rollout and A/B testing.
blue-green-deploy
Configure blue-green, canary, and rolling deployments on Kubernetes with Istio and Argo Rollouts.
opentelemetry
Instrument applications and infrastructure with OpenTelemetry for unified traces, metrics, and logs.
loki-logging
Configure Grafana Loki, Promtail, and Grafana for log aggregation in Kubernetes and Docker.
prometheus-grafana
Deploy Prometheus and Grafana with Docker Compose for metrics collection and visualization.
datadog
Configure Datadog agents, integrations, logs, tracing, metrics, dashboards, and alerts.
elk-stack
Deploy and manage the ELK Stack for centralized log aggregation and analysis.
alerting-oncall
Configure Prometheus alert rules and on-call rotations with PagerDuty or Grafana OnCall.
new-relic
Configure New Relic agents, NRQL queries, dashboards, and alerts.
rag-observability-evals
Monitor RAG systems with retrieval and generation quality metrics.
llmops-platform-engineering
Design and operate production LLMOps platforms with CI/CD, evaluation gates, and rollback.
agent-evals
Build automated evaluation suites for AI agents with golden datasets and regression gates.
llm-caching
Implement multi-layer LLM caching with Redis, GPTCache, Qdrant, and provider-side prompt caching.
agent-observability
Instrument AI agents with trace IDs and child spans for LLM calls.
ai-pipeline-orchestration
Orchestrate AI/ML pipelines for ingestion, training, inference, and RAG indexing.
model-registry-governance
Enforce model registry governance with metadata schemas and approval workflows.
llm-cost-optimization
Optimizes LLM API costs via model selection, caching, and batch processing.
ai-sre-incident-response
Define AI incident classes, severity frameworks, and response playbooks for LLM outages.
kubernetes-ops
Deploy, scale, and manage Kubernetes workloads with deployments, services, and configurations.
model-serving-kubernetes
Deploy and manage ML models on Kubernetes with KServe and Triton.
openshift
Manage Red Hat OpenShift clusters and deployments using the oc CLI.
argocd-gitops
Define ArgoCD applications, projects, and sync policies for GitOps deployments.