XiaHan avatar

XiaHan

Community

@jstzwj · SuZhou,China

104Followers
|
161Public Repos
|
23Published Skills

Learning to learn

Agent Skills by XiaHan

Showing 23 vetted skills indexed across 1 GitHub repositories.

jstzwjjstzwj
4

xla

Compile and optimize machine learning models for CPUs, GPUs, and TPUs.

Community
Advanced
jstzwjjstzwj
4

deepspeed

Optimize large-scale deep learning training with ZeRO and parallelism techniques.

Community
Advanced
jstzwjjstzwj
4

nsight

Profile GPU and CPU performance with timeline visualization and hardware counters.

Community
Advanced
jstzwjjstzwj
4

tile-ir

Generate tile-centric GPU kernels for NVIDIA tensor core execution.

Community
Advanced
jstzwjjstzwj
4

nccl

Provide collective operations for multi-GPU data communication in CUDA applications.

Community
Intermediate
jstzwjjstzwj
4

onnxruntime

Run ONNX model inference and training with hardware acceleration.

Community
Intermediate
jstzwjjstzwj
4

cuda

Write CUDA kernels for NVIDIA GPU computing with memory management.

Community
Advanced
jstzwjjstzwj
4

triton

Create and autotune GPU kernels in Python with MLIR-based compilation.

Community
Advanced
jstzwjjstzwj
4

vllm

Deploy large language models across GPU clusters with distributed serving and OpenAI-compatible APIs.

Community
Advanced
jstzwjjstzwj
4

mlir

Create, transform, and optimize MLIR modules for AI frameworks and compiler backends.

Community
Advanced
jstzwjjstzwj
4

tensorflow

Creates ML models with TensorFlow including optimization and deployment workflows.

Community
Intermediate
jstzwjjstzwj
4

sglang

Orchestrate large model deployment with GPU optimization and distributed serving.

Community
Advanced
jstzwjjstzwj
4

tilelang

Generates GPU kernels for CUDA, HIP, Metal, and CPU via Python.

Community
Advanced
jstzwjjstzwj
4

flash-attention

Implement optimized attention kernels for PyTorch and CUDA GPUs.

Community
Advanced
jstzwjjstzwj
4

pytorch

Create, optimize, and deploy PyTorch deep learning models.

Community
Advanced
jstzwjjstzwj
4

jax

Run numerical routines, automatic differentiation, and hardware acceleration with JAX.

Community
Advanced
jstzwjjstzwj
4

Megatron-LM

Train and deploy large transformer models with tensor, pipeline, and expert parallelism.

Community
Advanced
jstzwjjstzwj
4

cuTile

Define immutable tiles for GPU kernels with automatic memory management.

Community
Advanced
jstzwjjstzwj
4

cutlass

Execute architecture-tuned GPU matrix multiplication kernels via NVIDIA CUTLASS.

Community
Advanced
jstzwjjstzwj
4

tvm

Import models from PyTorch, ONNX, and TFLite into Relax IR for optimization.

Community
Intermediate
jstzwjjstzwj
4

ray

Scale AI workflows across distributed systems with Ray.

Community
Intermediate
jstzwjjstzwj
4

xformers

Provide optimized attention and sparse tensor operations for Transformer models.

Community
Advanced
jstzwjjstzwj
4

bitsandbytes

Quantize weights and optimizer states for PyTorch model training.

Community
Advanced