RedAI Infra avatar

RedAI Infra

Official

@redai-infra · China

0Followers
|
16Public Repos
|
16Published Skills

Building the infrastructure for large model training, inference, optimization, and serving — empowering creators and developers to harness AI at scale.

Skills Distribution
DomainAI Models & ...Distributed Traini.. (40%)Model Architecture.. (35%)Technical Document.. (25%)

Agent Skills by RedAI Infra

Showing 16 vetted skills indexed across 1 GitHub repositories.

redai-infraredai-infra
585

sync-github

Synchronizes code between internal GitLab and external GitHub via gated cherry-pick workflows.

Official
Advanced
redai-infraredai-infra
585

agentic-rollout

Assess, integrate, and validate external agents for Relax resident Agentic rollout.

Official
Advanced
redai-infraredai-infra
585

nemo-gym-recipe-integration

Integrates NVIDIA NeMo Gym environments into Relax as three-step RL training recipes.

Official
Advanced
redai-infraredai-infra
585

perf-doctor

Diagnoses Relax RL training launch scripts for performance and GPU memory misconfigurations.

Official
Advanced
redai-infraredai-infra
585

sglang-upgrade

Rebases Relax's sglang patch onto a new sglang release and updates the training Docker image.

Official
Advanced
redai-infraredai-infra
585

opd-tuning

Diagnoses and tunes Relax on-policy distillation teacher engine configurations and GPU splits.

Official
Advanced
redai-infraredai-infra
585

ssh-ray-cluster

Debugs remote Ray cluster jobs through an SSH submit, log-check, and fix loop.

Official
Intermediate
redai-infraredai-infra
566

model-integration

Integrate custom model architectures into the Relax training system.

Official
Advanced
redai-infraredai-infra
566

git-commit

Create standardized git commit messages with markdown-formatted bodies.

Official
Basic
redai-infraredai-infra
566

creating-skills

Provides step-by-step instructions for developing Claude Code skills with file organization and metadata configuration.

Official
Basic
redai-infraredai-infra
566

redaccel-to-relax

Converts RedAccel RL training components and scripts to Relax framework equivalents.

Official
Advanced
redai-infraredai-infra
566

doc-writer

Verify source code and write bilingual documentation for the Relax framework.

Official
Advanced
redai-infraredai-infra
566

verl-to-relax

Convert verl reinforcement learning recipes and code to Relax.

Official
Advanced
redai-infraredai-infra
566

code-review

Analyze Python codebases for quality, security, and architectural integrity.

Official
Advanced
redai-infraredai-infra
566

debug-hang

Diagnose distributed training hangs in Ray clusters by analyzing call stacks and task states.

Official
Intermediate
redai-infraredai-infra
566

relax-dev-debug

Submit and monitor reinforcement learning training jobs on remote Ray clusters.

Official
Basic

Frequently Asked Questions About RedAI Infra

FAQPage Schema
What specific tasks can engineers perform using RedAI Infra capabilities?

Engineers can migrate reinforcement learning recipes from RedAccel or verl to the Relax framework, debug distributed training hangs in Ray clusters, and generate bilingual documentation for source code. These capabilities streamline the transition of custom model architectures into production-ready training environments.

Which personas benefit most from these infrastructure skills?

Machine learning engineers, infrastructure architects, and research scientists working on large-scale model training benefit from these skills. The registry specifically supports those managing distributed compute clusters and those tasked with standardizing code quality and documentation within the Relax ecosystem.

What are the prerequisites for running these training and debugging tasks?

Users require an active Ray cluster environment and existing access to the Relax framework. The skills assume familiarity with reinforcement learning architectures and the ability to interface with remote compute nodes for job submission and stack trace analysis.