run-deep-swe

Score AI models on the 113-task DeepSWE benchmark via OpenRouter.

61|11|Updated Jun 15, 2026
One-click install
npx skills add https://github.com/Matymatyk-business/david-skills --skill run-deep-swe
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: run-deep-swe
Source: https://github.com/Matymatyk-business/david-skills/tree/main/skills/run-deep-swe
Command: npx skills add https://github.com/Matymatyk-business/david-skills --skill run-deep-swe

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill enables automated, reproducible scoring of AI models on the 113-task DeepSWE coding-agent benchmark via OpenRouter, allowing fair comparisons and verification of vendor claims.

Core Features & Use Cases

  • OpenRouter-driven evaluation with mini-swe-agent for model-agnostic benchmarking.
  • Supports single-task, subset, and full 113-task runs, with optional leaderboard submission.
  • Includes setup, OpenRouter wiring, and run workflows to produce comparable results.

Quick Start

Run a full 113-task DeepSWE evaluation for a selected model via OpenRouter and review the resulting score.

Frequently Asked Questions about run-deep-swe

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark AI models on the DeepSWE coding-agent benchmark?

To benchmark AI models on DeepSWE, you can use this skill to run automated, reproducible scoring via the OpenRouter API. It executes single-task, subset, or full 113-task runs end-to-end using a compatible mini-swe-agent driver.

What do I need to run DeepSWE evaluations through OpenRouter?

Running DeepSWE evaluations requires OpenRouter API access, a compatible mini-swe-agent driver, and a configured environment with Docker, uv, and a model slug. This setup ensures tasks execute correctly end-to-end within the sandbox.

Can I evaluate a subset of DeepSWE tasks instead of the full 113-task corpus?

Yes, you can evaluate a subset of DeepSWE tasks instead of the full 113-task corpus. The skill supports single-task, subset, and full-corpus runs to provide flexible model evaluation options via OpenRouter.

What is DeepSWE benchmark evaluation used for?

DeepSWE benchmark evaluation is used for automated, reproducible scoring of AI models on a 113-task coding-agent benchmark. It allows fair comparisons and verification of vendor claims through model-agnostic testing.

Does this DeepSWE evaluation tool support leaderboard submission?

Yes, this DeepSWE evaluation tool supports optional leaderboard submission. After running your model evaluation through OpenRouter, you can submit the resulting comparable scores directly to the leaderboard.

Why use OpenRouter and Docker for model evaluation on DeepSWE?

Using OpenRouter and Docker for model evaluation on DeepSWE ensures model-agnostic benchmarking and sandboxed task execution. This approach guarantees reproducible scoring by isolating the environment and standardizing API access.