Xiaoyu Zhang avatar

Xiaoyu Zhang

Community

@bbuf · XiChang

2,662Followers
|
39Public Repos
|
74Published Skills

Working at RadixArk and the creator of GiantPandaLLM official account.

Agent Skills by Xiaoyu Zhang

Showing 74 vetted skills indexed across 1 GitHub repositories.

BBufBBuf
783

model-pr-history-knowledge

Query PR-driven model optimization history across SGLang, vLLM, TensorRT-LLM, and TokenSpeed.

Community
Intermediate
BBufBBuf
783

sglang-sota-humanize-loop

Runs a Humanize-governed RLCR loop that benchmarks, profiles, and patches SGLang until it matches competing serving frameworks.

Community
Advanced
BBufBBuf
783

vllm-sota-humanize-loop

Runs an autonomous Humanize RLCR loop that benchmarks, profiles, and patches vLLM until it matches the best framework.

Community
Advanced
BBufBBuf
783

sglang-humanize-review

Review SGLang pull requests against a corpus of historical human maintainer review discussions.

Community
Advanced
BBufBBuf
783

model-compute-simulation

Simulates operator-level compute flow and estimates FLOPs and MFU for LLM serving configurations.

Community
Advanced
BBufBBuf
783

llm-pipeline-analysis

Inspect LLM torch profiler traces at forward-pass, layer, and kernel level.

Community
Advanced
BBufBBuf
783

llm-serving-capacity-planner

Parse SGLang and vLLM startup logs to decompose GPU memory and estimate request concurrency.

Community
Advanced
BBufBBuf
783

sglang-model-day0-support

Plans and audits evidence-driven SGLang Day-0 support programs for new model architectures.

Community
Advanced
BBufBBuf
721

h100

SSH into h100_sglang to attach to the sglang_bbuf container for GPU-enabled SGLang development.

Community
Intermediate
BBufBBuf
721

h100-sglang-diffusion

Automate diffusion smoke tests and kernel validation on the H100 SGLang remote environment.

Community
Advanced
BBufBBuf
721

llm-serving-auto-benchmark

Benchmark LLM serving frameworks across SGLang, vLLM, and TensorRT-LLM.

Community
Advanced
BBufBBuf
721

model-architecture-diagram

Return public original architecture diagram URLs for specified models.

Community
Intermediate
BBufBBuf
721

sglang-prod-incident-triage

Convert live SGLang incidents into replayable debugging workflows.

Community
Advanced
BBufBBuf
721

llm-torch-profiler-analysis

Analyze LLM torch-profiler traces to produce kernel, overlap, and fuse tables.

Community
Intermediate
BBufBBuf
721

sglang-sota-performance

Benchmark models across SGLang, vLLM, and TensorRT-LLM to identify performance gaps.

Community
Advanced
BBufBBuf
721

model-pr-diff-dossier

Create diff-reviewed PR cards for model optimization PR history documents.

Community
Advanced
BBufBBuf
721

vllm-nemotron-super-optimization

Identify diff-driven optimizations for Nemotron models in vLLM.

Community
Advanced
BBufBBuf
721

vllm-deepseek-v4-optimization

Audits and documents DeepSeek V4 PRs (#40760, #40811, #40806 progress in vLLM.

Community
Intermediate
BBufBBuf
721

vllm-glm45-optimization

Audit PR diffs and generate release-ready dossiers for GLM-4.5 series in vLLM.

Community
Advanced
BBufBBuf
721

vllm-glm-vlm-ocr-optimization

Document production-level GLM VLM OCR optimization changes from vLLM PR histories.

Community
Advanced
BBufBBuf
721

vllm-mixtral-quark-int4fp8-moe-optimization

Audit PR diffs for Mixtral Quark INT4-FP8 MoE optimization in vLLM.

Community
Advanced
BBufBBuf
721

vllm-llama4-optimization

Audit PR-backed Llama4 optimizations in vLLM for runtime and quantization.

Community
Intermediate
BBufBBuf
721

vllm-qwen36-optimization

Audit Qwen3.6 optimization patches in vLLM using PR-backed evidence.

Community
Advanced
BBufBBuf
721

vllm-kimi-optimization

Analyze Kimi optimization PR diffs in vLLM for runtime compatibility.

Community
Advanced