sglang-mixtral-quark-int4fp8-moe-optimization
Quantize Mixtral MoE models to INT4FP8 with per-expert weight packing on AMD ROCm.
npx skills add https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS --skill sglang-mixtral-quark-int4fp8-moe-optimization
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill. Skill: sglang-mixtral-quark-int4fp8-moe-optimization Source: https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/model-optimization/sglang/sglang-mixtral-quark-int4fp8-moe-optimization Command: npx skills add https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS --skill sglang-mixtral-quark-int4fp8-moe-optimization