evaluating-cosmos-policy

Evaluate NVIDIA Cosmos Policy on LIBERO and RoboCasa simulation environments.

Updated Mar 30, 2026
One-click install
npx skills add https://github.com/KappTech88/AI-RESEARCH-SKILLS-MCP --skill evaluating-cosmos-policy
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluating-cosmos-policy
Source: https://github.com/KappTech88/AI-RESEARCH-SKILLS-MCP/tree/main/skills/cosmos-policy
Command: npx skills add https://github.com/KappTech88/AI-RESEARCH-SKILLS-MCP --skill evaluating-cosmos-policy

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Evaluating NVIDIA Cosmos Policy on LIBERO and RoboCasa simulation environments to benchmark robotics policies and validate performance in headless setups.

Core Features & Use Cases

  • LIBERO & RoboCasa evaluation workflows for end-to-end policy testing.
  • Headless GPU setup with EGL rendering for scalable benchmarking.
  • Reproducible configurations and command surfaces to ensure consistent results across runs.

Quick Start

Begin a smoke evaluation using the LIBERO or RoboCasa workflow to verify setup and baseline performance.

Frequently Asked Questions about evaluating-cosmos-policy

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate robotics policies in LIBERO and RoboCasa simulation environments?

You can benchmark NVIDIA Cosmos Policy in LIBERO and RoboCasa by running the cosmos_policy module with explicit evaluation entry points and configuration files to generate reproducible performance metrics.

Can I run headless GPU evaluations for Cosmos Policy benchmarking?

Yes, headless GPU evaluation is supported using EGL rendering, allowing scalable benchmarking across multiple LIBERO task suites and RoboCasa tasks without requiring a display server.

What dependencies do I need to benchmark Cosmos Policy on LIBERO?

Benchmarking requires the Cosmos Policy repository, LIBERO and RoboCasa simulation environments, and explicit evaluation entry points via the cosmos_policy module with appropriate configuration files.

How do I profile latency for robotics policies in simulation environments?

Latency profiling is handled through evaluation workflows applying reproducible configurations and command surfaces to ensure consistent latency measurements across LIBERO and RoboCasa runs.

Does this Cosmos Policy evaluation workflow support reproducible benchmarks across multiple task suites?

Yes, the workflow supports reproducible benchmarks across multiple LIBERO task suites and RoboCasa tasks by applying reproducible configurations and command surfaces to ensure consistent results across runs.