What problem does it solve?
This Skill solves the challenge of manually analyzing and generating images, audio, videos, and documentation by automating these tasks using AI-powered models.
Core Features & Use Cases
- Multimodal Analysis: Analyze images, audio, video, and documentation using state-of-the-art models like Gemini.
- Image Generation: Create high-quality images based on prompts, using models like Imagen 4 or Nano Banana.
- Audio Analysis and Generation: Extract text from audio files, transcribe speeches, and generate audio using different voices and styles.
- Video Analysis and Generation: Summarize, transcribe, and analyze videos, and generate videos with custom parameters.
- Document Conversion: Convert various document formats (PDF, Office, HTML) to Markdown.
- Use Case: If you have a meeting transcript in audio format and you need to extract key points, this Skill can transcribe the audio and generate a summary.
Quick Start
To analyze an image, run the command: analyze <image-file>.