evaluating-cosmos-policy

Evaluate NVIDIA Cosmos Policy deployments on LIBERO and RoboCasa simulations with latency profiling.

Updated Mar 18, 2026
One-click install
npx skills add https://github.com/tadod12/fraud-detection-research --skill evaluating-cosmos-policy-tadod12
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluating-cosmos-policy
Source: https://github.com/tadod12/fraud-detection-research/tree/main/.agent/skills/18-multimodal/cosmos-policy
Command: npx skills add https://github.com/tadod12/fraud-detection-research --skill evaluating-cosmos-policy-tadod12

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill helps researchers quickly benchmark and validate NVIDIA Cosmos Policy deployments in robotics simulators, enabling reliable comparisons and reproducible results across experiments.

Core Features & Use Cases

  • LIBERO and RoboCasa evaluation workflows to benchmark policy behavior in simulated environments.
  • Headless EGL rendering and latency profiling to quantify performance on GPU clusters or local hardware.
  • Reproducibility guidance and deployment notes for end-to-end Cosmos Policy evaluations in research projects.

Quick Start

Run the LIBERO evaluation workflow with a Cosmos Policy checkpoint to start a headless benchmark.

Frequently Asked Questions about evaluating-cosmos-policy

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate NVIDIA Cosmos Policy in LIBERO and RoboCasa simulations?

To evaluate NVIDIA Cosmos Policy in LIBERO and RoboCasa, you run headless GPU benchmarks using this skill to quantify deployment performance, measure latency, and ensure reproducible results across simulated robotic environments.

What dependencies do I need to run headless Cosmos Policy benchmarks on a GPU cluster?

Running headless Cosmos Policy benchmarks requires PyTorch, MuJoCo, Robosuite, the Robocasa fork, transformers, and the Cosmos Policy repositories to execute end-to-end evaluations and generate latency profiles.

Can I use this skill for latency profiling of robotics policies in simulation?

Yes, this skill supports latency profiling for robotics policies by utilizing headless EGL rendering to quantify performance metrics on GPU clusters or local hardware during LIBERO and RoboCasa evaluation workflows.

Does this skill support reproducible setups for RoboCasa evaluation workflows?

Yes, the skill provides reproducibility guidance and deployment notes for RoboCasa evaluation workflows, allowing researchers to reliably benchmark Cosmos Policy behavior and validate results across multiple simulation experiments.

What is the best way to benchmark a Cosmos Policy checkpoint in a simulated environment?

The best way to benchmark a Cosmos Policy checkpoint is to run the LIBERO evaluation workflow in a headless environment, which quickly validates policy behavior and quantifies performance metrics in simulated robotics tasks.