What problem does it solve?
This Skill solves the problem of evaluating NVIDIA Cosmos Policy on LIBERO and RoboCasa robot-manipulation simulation environments with reproducible results and measurable performance.
Core Features & Use Cases
- Evaluate on LIBERO: Run smoke and full benchmarks using the public evaluation entrypoint for LIBERO.
- Evaluate on RoboCasa: Run single-task and multi-task evaluations for robot kitchens and object interaction scenarios.
- Headless GPU rendering (EGL): Configure MuJoCo/EGL and OpenGL platform variables to run safely on cluster or local headless GPUs.
- Performance profiling support: Establish a repeatable evaluation harness for inference latency/throughput checks tied to the same simulator and configuration.
Use Case: You want to measure Cosmos Policy’s success rate and runtime on LIBERO across multiple task suites (or on RoboCasa across tasks) on a headless GPU node, then compare runs by holding seed, task configuration, and rendering settings constant.
Quick Start
In an EGL-capable runtime, run the official LIBERO evaluation module with the provided flags and a matching Cosmos Policy checkpoint and config.