benchclaw-stage3-simulator-annotation

Process and annotate simulation evidence for BenchClaw Stage 3 benchmarks.

Updated May 7, 2026
One-click install
npx skills add https://github.com/EurecaMoment/BenchClaw --skill benchclaw-stage3-simulator-annotation
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: benchclaw-stage3-simulator-annotation
Source: https://github.com/EurecaMoment/BenchClaw/tree/main/BenchClaw/skills/benchmark-stage3-evidence-compiler/skills/simulator-evidence-compilation/subskills/annotation
Command: npx skills add https://github.com/EurecaMoment/BenchClaw --skill benchclaw-stage3-simulator-annotation

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

BenchClaw Stage 3 Simulator Annotation Skill addresses the need for efficient processing and annotation of simulation evidence, streamlining the benchmark construction workflow for embodied AI and robotics.

Core Features & Use Cases

  • Simulator Evidence Processing: Handles the processing of simulation evidence for the BenchClaw benchmarking system.
  • GT/Annotation Integration: Integrates with Ground Truth (GT) data and annotation records to ensure accurate evidence preparation.
  • Visual Pseudo-Annotation: Optionally supports visual pseudo-annotation for enhanced GT generation when specified.
  • Environment and Simulation Tracking: Records environment versions and simulation details for reproducibility and auditing.

Quick Start

Run the benchclaw-stage3-simulator-annotation Skill with the specified subskill and parameters: /benchclaw-subskill /path/to/SKILL.md -name benchclaw-stage3-simulator-annotation -params ...

Frequently Asked Questions about benchclaw-stage3-simulator-annotation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate simulation evidence annotation for embodied AI benchmark construction?

Simulation evidence annotation for AI benchmarks is automated by processing simulator outputs, integrating Ground Truth (GT) data, and optionally applying visual pseudo-annotation. This Skill streamlines benchmark construction workflows for embodied AI and robotics by ensuring accurate evidence preparation and tracking environment versions for reproducibility.

What is visual pseudo-annotation in robotics simulation and when do I need it?

Visual pseudo-annotation in robotics simulation is an optional feature that enhances Ground Truth (GT) generation from simulation evidence. You need it when your benchmark construction workflow requires supplementary visual data labeling to ensure accurate evidence preparation for embodied AI evaluation.

How do I process Ground Truth data and annotation records for BenchClaw Stage 3?

To process Ground Truth data and annotation records for BenchClaw Stage 3, run the simulator annotation subskill via a subagent execution model with specific input parameters. This integrates GT data with simulation evidence and records environment versions for auditing and reproducibility.

Does simulation evidence annotation require a subagent execution model?

Yes, simulation evidence annotation for BenchClaw Stage 3 requires a subagent execution model. You must invoke the annotation subskill with specific input parameters, such as the path to the SKILL.md file, to properly process Ground Truth data and simulation evidence.

What's the best way to track environment versions for reproducible AI benchmarking?

The best way to track environment versions for reproducible AI benchmarking is to use an annotation workflow that automatically records simulation details and environment versions. This ensures auditing capability and reproducibility during Ground Truth integration for embodied AI and robotics benchmarks.

Simulation evidence processing not working without Ground Truth records, what are the limitations?

A limitation of simulation evidence processing is that it requires Ground Truth data and annotation records to function correctly. Without proper GT integration and a subagent execution model, the annotation workflow cannot ensure accurate evidence preparation or reproducibility for embodied AI benchmarks.