What problem does it solve?
Removes the need to upload full screenshots or manually operate the macOS UI by providing local, low‑token OCR-based perception and scripted input control so the assistant can find text on screen, move the cursor, click buttons, and type reliably.
Core Features & Use Cases
- Local cursor-proximate OCR using the macOS Vision framework to detect visible text and return clickable coordinates.
- Input emulation via cliclick and osascript for mouse movement, clicks, key presses, and keyboard shortcuts with verification.
- Messaging automation and monitoring workflows for apps like WeChat that include read/send scripts and atomic trigger detection.
- Safety rules and fallbacks: staged permission levels, prohibition on passwords/payments/Terminal control, and screenshot + visual-analysis fallback when OCR fails.
- Use case: ask the assistant to locate and click a "保存" button, fill a text field, or send a WeChat message and verify delivery via OCR.
Quick Start
Ask CC to use the desktop skill to locate and click the "保存" button in the current macOS window.