SURUS avatar

SURUS

Official

@surus-lat · Argentina

0Followers
|
19Public Repos
|
23Published Skills

end-to-end AI for latam

Skills Distribution
DomainAI Models & ...Model Benchmarking (40%)Dataset Engineering (30%)Performance Evalua.. (30%)

Agent Skills by SURUS

Showing 23 vetted skills indexed across 1 GitHub repositories.

surus-latsurus-lat
8

workshop-benchmark-together-model

Integrate and benchmark Together AI models in Benchy with structured extraction tasks.

Official
Intermediate
surus-latsurus-lat
8

workshop-define-benchmark

Create and validate custom AI benchmarking tasks in the Benify framework.

Official
Intermediate
surus-latsurus-lat
8

workshop-submit-to-latamboard

Package benchmark run outputs into submission folders and create GitHub pull requests.

Official
Intermediate
surus-latsurus-lat
8

validate

Validate YAML benchmark specification files for completeness and structural integrity.

Official
Intermediate
surus-latsurus-lat
8

rebuild-leaderboard

Automate benchmark re-evaluation and publish results to HuggingFace datasets.

Official
Advanced
surus-latsurus-lat
8

qwen3-asr-howto

Deploy and benchmark Qwen3-ASR models with local, vLLM, or DashScope inference.

Official
Advanced
surus-latsurus-lat
8

define-task

Generate benchmark task YAML configurations for the Benchy engine.

Official
Intermediate
surus-latsurus-lat
8

transcription-benchmark

Benchmark local ASR models on FLEURS datasets with WER and CER tables.

Official
Advanced
surus-latsurus-lat
8

submit-to-latamboard

Package benchmark outputs into JSON summaries and submit via GitHub pull requests.

Official
Intermediate
surus-latsurus-lat
8

setup-data

Configure and validate benchmark datasets by mapping CSV or JSONL files to required schemas.

Official
Intermediate
surus-latsurus-lat
8

evaluate

Executes standardized AI benchmarking workflows with smoke tests and configurable interfaces.

Official
Advanced
surus-latsurus-lat
8

interpret-run

Parse run_outcome.json to diagnose AI benchmark failures and metrics.

Official
Intermediate
surus-latsurus-lat
8

add-task

Generate benchmark task templates and register them in the Benchy evaluation engine.

Official
Intermediate
surus-latsurus-lat
8

define-scoring

Standardize scoring metric definitions for AI benchmarks in Benchy YAML configurations.

Official
Basic
surus-latsurus-lat
8

synthesize-data

Generate synthetic benchmark datasets as JSONL input-output pairs from task specifications.

Official
Intermediate
surus-latsurus-lat
8

best-part-is-no-part

Evaluate software design proposals to minimize component count using a five-lens audit framework.

Official
Advanced
surus-latsurus-lat
8

read-results

Translate benchmark outcome and summary JSON files into human-readable performance reports.

Official
Basic
surus-latsurus-lat
8

add-provider

Integrate AI model providers into the Benchy benchmarking framework via YAML configuration.

Official
Intermediate
surus-latsurus-lat
8

oracle-plan

Convert implementation snapshots to structured oracle-quality design plans and back.

Official
Advanced
surus-latsurus-lat
8

configure-model

Generate YAML schemas for AI model benchmarking targets in Benchy.

Official
Intermediate
surus-latsurus-lat
8

run-benchmark

Execute end-to-end AI benchmarking workflows with spec validation and smoke testing.

Official
Advanced
surus-latsurus-lat
8

push-to-latamboard

Merge local benchmark results into the HuggingFace LatamBoard leaderboard dataset.

Official
Advanced
surus-latsurus-lat
8

whisper-benchmark

Benchmark Whisper-family ASR models on FLEURS datasets with WER and CER metrics.

Official
Advanced

Frequently Asked Questions About SURUS

FAQPage Schema
What specific tasks can I perform using these benchmarking capabilities?

You can define custom evaluation tasks, validate YAML configurations, execute performance smoke tests, and generate human-readable reports from outcome files. The framework supports end-to-end processing from dataset mapping to final leaderboard submission.

Which target personas benefit from these benchmarking resources?

These resources are designed for machine learning engineers, data scientists, and researchers focused on model evaluation, transcription accuracy, and performance auditing within the Latam region.

What are the prerequisites for running these evaluation tasks?

Execution requires a configured environment capable of processing YAML specifications and interacting with the Benchy engine. Users must provide valid input datasets in CSV or JSONL formats and define model targets through the standardized configuration schemas.