evaluate-cortex-agent

Evaluate Cortex Agents in Snowflake and review results in Snowsight.

Updated Mar 7, 2026
One-click install
npx skills add https://github.com/randoneering/nix-flake-mirror --skill evaluate-cortex-agent
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluate-cortex-agent
Source: https://github.com/randoneering/nix-flake-mirror/tree/main/home/programs/opencode/skills/snowflake/agent_optimization/evaluate-cortex-agent
Command: npx skills add https://github.com/randoneering/nix-flake-mirror --skill evaluate-cortex-agent

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill provides a structured, repeatable workflow to evaluate Cortex Agents using Snowflake’s native Agent Evaluations, enabling objective benchmarking and comparison of agent performance across configurations.

Core Features & Use Cases

  • Define evaluation datasets for Cortex Agents and track metrics such as correctness, tool_selection_accuracy, tool_execution_accuracy, and logical_consistency.
  • Automate setup of evaluation runs in Snowflake and generate Snowsight reports.
  • Support scenario-based comparisons to measure improvements after prompts, tool changes, or configuration updates.

Quick Start

Configure the target agent, select metrics, build or choose a dataset, run the evaluation, and review results in Snowsight.

Frequently Asked Questions about evaluate-cortex-agent

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark a Cortex Agent in Snowflake?

To benchmark a Cortex Agent in Snowflake, you select metrics like correctness and logical_consistency, supply an evaluation dataset, run the evaluation, and review the performance results in Snowsight.

What metrics are available for Cortex Agent evaluation?

Cortex Agent evaluation metrics include correctness, tool_selection_accuracy, tool_execution_accuracy, and logical_consistency to objectively measure and compare agent performance across configurations.

Can I compare Cortex Agent performance after prompt or tool changes?

Yes, you can compare Cortex Agent performance after prompt or tool changes by running scenario-based evaluations to measure improvements and benchmark configurations against previous results in Snowsight.

Do I need to prepare a dataset to evaluate a Cortex Agent?

Yes, you need to prepare or supply a dataset to evaluate a Cortex Agent, which is necessary to measure metrics like correctness and tool_selection_accuracy during the evaluation run.

What is the best way to track logical consistency in AI agents?

The best way to track logical consistency in AI agents is using Snowflake native Agent Evaluations to benchmark performance, which requires selecting the logical_consistency metric and running an evaluation with a prepared dataset.