RedAI Infra
Official@redai-infra · China
Building the infrastructure for large model training, inference, optimization, and serving — empowering creators and developers to harness AI at scale.
Agent Skills by RedAI Infra
Showing 16 vetted skills indexed across 1 GitHub repositories.
sync-github
Synchronizes code between internal GitLab and external GitHub via gated cherry-pick workflows.
agentic-rollout
Assess, integrate, and validate external agents for Relax resident Agentic rollout.
nemo-gym-recipe-integration
Integrates NVIDIA NeMo Gym environments into Relax as three-step RL training recipes.
perf-doctor
Diagnoses Relax RL training launch scripts for performance and GPU memory misconfigurations.
sglang-upgrade
Rebases Relax's sglang patch onto a new sglang release and updates the training Docker image.
opd-tuning
Diagnoses and tunes Relax on-policy distillation teacher engine configurations and GPU splits.
ssh-ray-cluster
Debugs remote Ray cluster jobs through an SSH submit, log-check, and fix loop.
model-integration
Integrate custom model architectures into the Relax training system.
git-commit
Create standardized git commit messages with markdown-formatted bodies.
creating-skills
Provides step-by-step instructions for developing Claude Code skills with file organization and metadata configuration.
redaccel-to-relax
Converts RedAccel RL training components and scripts to Relax framework equivalents.
doc-writer
Verify source code and write bilingual documentation for the Relax framework.
verl-to-relax
Convert verl reinforcement learning recipes and code to Relax.
code-review
Analyze Python codebases for quality, security, and architectural integrity.
debug-hang
Diagnose distributed training hangs in Ray clusters by analyzing call stacks and task states.
relax-dev-debug
Submit and monitor reinforcement learning training jobs on remote Ray clusters.
Frequently Asked Questions About RedAI Infra
FAQPage SchemaWhat specific tasks can engineers perform using RedAI Infra capabilities?▼
Engineers can migrate reinforcement learning recipes from RedAccel or verl to the Relax framework, debug distributed training hangs in Ray clusters, and generate bilingual documentation for source code. These capabilities streamline the transition of custom model architectures into production-ready training environments.
Which personas benefit most from these infrastructure skills?▼
Machine learning engineers, infrastructure architects, and research scientists working on large-scale model training benefit from these skills. The registry specifically supports those managing distributed compute clusters and those tasked with standardizing code quality and documentation within the Relax ecosystem.
What are the prerequisites for running these training and debugging tasks?▼
Users require an active Ray cluster environment and existing access to the Relax framework. The skills assume familiarity with reinforcement learning architectures and the ability to interface with remote compute nodes for job submission and stack trace analysis.