What problem does it solve?
This Skill automates the process of creating benchmark trap tasks for the CentralGauge platform from approved Azure DevOps pull requests, ensuring that the tasks are representative of real-world scenarios and challenging for large language models.
Core Features & Use Cases
- Pull Request Analysis: Extracts and analyzes code comments and decisions from pull requests to identify potential traps for large language models.
- Task Creation: Automatically creates self-contained or base-app-faithful tasks based on identified traps.
- Discrimination Probing: Ensures that the tasks are challenging by requiring both a correct and a naive solution to pass or fail, respectively.
- Use Case: Use this Skill to generate tasks that can be used in the CentralGauge benchmarking platform to evaluate the performance of large language models on Business Central AL code generation.
Quick Start
Use the extract-trap-task skill to create a benchmark task from an Azure DevOps pull request with ID '12345'.