What problem does it solve?
This Skill helps optimize offline AI inference on Apple Silicon by reducing latency and improving the performance of local LLM, speech, and embedding workloads.
Core Features & Use Cases
- Local Model Acceleration: Run and optimize LLMs, Whisper speech recognition, and embedding models with Apple MLX for faster on-device execution.
- Latency Optimization: Configure streaming generation, quantization, model loading, and memory usage for responsive offline voice assistant pipelines.
- Use Case: Use this Skill when building ManuAI's factory-floor copilot to choose between MLX and Ollama, tune Qwen inference, accelerate Whisper transcription, and maintain embedding parity for retrieval.
Quick Start
Use the mlx skill to optimize a local Qwen voice assistant pipeline on an Apple Silicon Mac for faster offline inference.