Peijia Qin avatar

Peijia Qin

Community

@t2ance

22Followers
|
11Public Repos
|
112Published Skills

Agent Skills by Peijia Qin

Showing 112 vetted skills indexed across 1 GitHub repositories.

t2ancet2ance

gguf-quantization

Quantize models to GGUF format for CPU inference with llama.cpp.

Community
Advanced
t2ancet2ance

awq-quantization

Automate 4-bit AWQ quantization of large language models for GPU deployment.

Community
Advanced
t2ancet2ance

hqq-quantization

Quantize large language model weights without calibration data.

Community
Advanced
t2ancet2ance

gptq

Quantize large language models to 4-bit precision for memory reduction and faster inference.

Community
Advanced
t2ancet2ance

optimizing-attention-flash

Apply Flash Attention to transformer models for reduced memory usage and increased speed.

Community
Advanced
t2ancet2ance

quantizing-models-bitsandbytes

Quantize LLMs to 8-bit or 4-bit with bitsandbytes for memory reduction.

Community
Intermediate
t2ancet2ance

mlflow

Track ML experiments and manage model lifecycles across training, registry, and deployment.

Community
Intermediate
t2ancet2ance

tensorboard

Visualize training metrics, graphs, and embeddings to analyze ML experiments.

Community
Advanced
t2ancet2ance

weights-and-biases

Automate ML experiment tracking, visualization, and artifact management with Weights & Biases.

Community
Advanced
t2ancet2ance

nemo-curator

Clean and curate multi-modal LLM training data with GPU-accelerated filtering and deduplication.

Community
Intermediate
t2ancet2ance

ray-data

Distribute data transformations across CPU/GPU clusters with Ray Data.

Community
Intermediate
t2ancet2ance

serving-llms-vllm

Deploy and manage high-throughput LLM serving with vLLM-powered OpenAI-compatible endpoints.

Community
Advanced
t2ancet2ance

tensorrt-llm

Deploy high-throughput LLM inference with TensorRT-LLM on NVIDIA GPUs.

Community
Advanced
t2ancet2ance

sglang

Automate prompt prefix caching for LLM serving with SGLang.

Community
Advanced
t2ancet2ance

llama-cpp

Run LLM inference on CPU and non-NVIDIA hardware with GGUF models.

Community
Intermediate
t2ancet2ance

huggingface-tokenizers

Train and load HuggingFace tokenizers with Rust-based Python bindings.

Community
Intermediate
t2ancet2ance

sentencepiece

Creates SentencePiece subword tokenizers from raw corpora for multilingual NLP preprocessing.

Community
Intermediate
t2ancet2ance

llama-factory

Guides users to fine-tune LLMs via LLaMA-Factory WebUI with no-code configuration.

Community
Advanced
t2ancet2ance

peft-fine-tuning

Fine-tune 7B–70B language models using LoRA and QLoRA adapters.

Community
Advanced
t2ancet2ance

axolotl

Design and execute Axolotl-based fine-tuning workflows with YAML configurations.

Community
Advanced
t2ancet2ance

unsloth

Guide memory-efficient LLM fine-tuning with Unsloth using LoRA and QLoRA workflows.

Community
Intermediate
t2ancet2ance

openrlhf-training

Orchestrate distributed RLHF training for large language models with Ray and vLLM.

Community
Advanced
t2ancet2ance

simpo-training

Optimize LLM alignment with reference-free SimPO preference optimization.

Community
Advanced
t2ancet2ance

slime-rl-training

Guide RL post-training of GLM-scale LLMs with Megatron-LM and SGLang workflows.

Community
Advanced