Xiaoyu Zhang
Community@bbuf · XiChang
Working at RadixArk and the creator of GiantPandaLLM official account.
Agent Skills by Xiaoyu Zhang
Showing 74 vetted skills indexed across 1 GitHub repositories.
model-pr-history-knowledge
Query PR-driven model optimization history across SGLang, vLLM, TensorRT-LLM, and TokenSpeed.
sglang-sota-humanize-loop
Runs a Humanize-governed RLCR loop that benchmarks, profiles, and patches SGLang until it matches competing serving frameworks.
vllm-sota-humanize-loop
Runs an autonomous Humanize RLCR loop that benchmarks, profiles, and patches vLLM until it matches the best framework.
sglang-humanize-review
Review SGLang pull requests against a corpus of historical human maintainer review discussions.
model-compute-simulation
Simulates operator-level compute flow and estimates FLOPs and MFU for LLM serving configurations.
llm-pipeline-analysis
Inspect LLM torch profiler traces at forward-pass, layer, and kernel level.
llm-serving-capacity-planner
Parse SGLang and vLLM startup logs to decompose GPU memory and estimate request concurrency.
sglang-model-day0-support
Plans and audits evidence-driven SGLang Day-0 support programs for new model architectures.
h100
SSH into h100_sglang to attach to the sglang_bbuf container for GPU-enabled SGLang development.
h100-sglang-diffusion
Automate diffusion smoke tests and kernel validation on the H100 SGLang remote environment.
llm-serving-auto-benchmark
Benchmark LLM serving frameworks across SGLang, vLLM, and TensorRT-LLM.
model-architecture-diagram
Return public original architecture diagram URLs for specified models.
sglang-prod-incident-triage
Convert live SGLang incidents into replayable debugging workflows.
llm-torch-profiler-analysis
Analyze LLM torch-profiler traces to produce kernel, overlap, and fuse tables.
sglang-sota-performance
Benchmark models across SGLang, vLLM, and TensorRT-LLM to identify performance gaps.
model-pr-diff-dossier
Create diff-reviewed PR cards for model optimization PR history documents.
vllm-nemotron-super-optimization
Identify diff-driven optimizations for Nemotron models in vLLM.
vllm-deepseek-v4-optimization
Audits and documents DeepSeek V4 PRs (#40760, #40811, #40806 progress in vLLM.
vllm-glm45-optimization
Audit PR diffs and generate release-ready dossiers for GLM-4.5 series in vLLM.
vllm-glm-vlm-ocr-optimization
Document production-level GLM VLM OCR optimization changes from vLLM PR histories.
vllm-mixtral-quark-int4fp8-moe-optimization
Audit PR diffs for Mixtral Quark INT4-FP8 MoE optimization in vLLM.
vllm-llama4-optimization
Audit PR-backed Llama4 optimizations in vLLM for runtime and quantization.
vllm-qwen36-optimization
Audit Qwen3.6 optimization patches in vLLM using PR-backed evidence.
vllm-kimi-optimization
Analyze Kimi optimization PR diffs in vLLM for runtime compatibility.