What problem does it solve?
This Skill removes the complexity of preparing, calibrating, and executing an end-to-end post-training quantization workflow for ONNX models with AMD Quark.
Core Features & Use Cases
- Full ONNX PTQ workflow: Handles model intake, quantization planning, calibration-script generation, manifest creation, and execution confirmation.
- Model optimization scenarios: Supports common ONNX quantization requests such as XINT8, A8W8, BFP16, and MXFP variants for vision models and LLM weight-only compression.
- Use Case: A user can provide an ONNX model and request a complete quantization pipeline that prepares calibration artifacts and runs the final optimized model generation steps.
Quick Start
Ask the skill to run an end-to-end quantization workflow for your ONNX model and specify the target quantization format and any external weights file if present.