SURUS
Official@surus-lat · Argentina
end-to-end AI for latam
Agent Skills by SURUS
Showing 23 vetted skills indexed across 1 GitHub repositories.
workshop-benchmark-together-model
Integrate and benchmark Together AI models in Benchy with structured extraction tasks.
workshop-define-benchmark
Create and validate custom AI benchmarking tasks in the Benify framework.
workshop-submit-to-latamboard
Package benchmark run outputs into submission folders and create GitHub pull requests.
validate
Validate YAML benchmark specification files for completeness and structural integrity.
rebuild-leaderboard
Automate benchmark re-evaluation and publish results to HuggingFace datasets.
qwen3-asr-howto
Deploy and benchmark Qwen3-ASR models with local, vLLM, or DashScope inference.
define-task
Generate benchmark task YAML configurations for the Benchy engine.
transcription-benchmark
Benchmark local ASR models on FLEURS datasets with WER and CER tables.
submit-to-latamboard
Package benchmark outputs into JSON summaries and submit via GitHub pull requests.
setup-data
Configure and validate benchmark datasets by mapping CSV or JSONL files to required schemas.
evaluate
Executes standardized AI benchmarking workflows with smoke tests and configurable interfaces.
interpret-run
Parse run_outcome.json to diagnose AI benchmark failures and metrics.
add-task
Generate benchmark task templates and register them in the Benchy evaluation engine.
define-scoring
Standardize scoring metric definitions for AI benchmarks in Benchy YAML configurations.
synthesize-data
Generate synthetic benchmark datasets as JSONL input-output pairs from task specifications.
best-part-is-no-part
Evaluate software design proposals to minimize component count using a five-lens audit framework.
read-results
Translate benchmark outcome and summary JSON files into human-readable performance reports.
add-provider
Integrate AI model providers into the Benchy benchmarking framework via YAML configuration.
oracle-plan
Convert implementation snapshots to structured oracle-quality design plans and back.
configure-model
Generate YAML schemas for AI model benchmarking targets in Benchy.
run-benchmark
Execute end-to-end AI benchmarking workflows with spec validation and smoke testing.
push-to-latamboard
Merge local benchmark results into the HuggingFace LatamBoard leaderboard dataset.
whisper-benchmark
Benchmark Whisper-family ASR models on FLEURS datasets with WER and CER metrics.
Frequently Asked Questions About SURUS
FAQPage SchemaWhat specific tasks can I perform using these benchmarking capabilities?▼
You can define custom evaluation tasks, validate YAML configurations, execute performance smoke tests, and generate human-readable reports from outcome files. The framework supports end-to-end processing from dataset mapping to final leaderboard submission.
Which target personas benefit from these benchmarking resources?▼
These resources are designed for machine learning engineers, data scientists, and researchers focused on model evaluation, transcription accuracy, and performance auditing within the Latam region.
What are the prerequisites for running these evaluation tasks?▼
Execution requires a configured environment capable of processing YAML specifications and interacting with the Benchy engine. Users must provide valid input datasets in CSV or JSONL formats and define model targets through the standardized configuration schemas.