gem

Processes PDFs, images, videos, and URLs to extract insights and generate images via Gemini models.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/hamelsmu/hamel --skill gem-hamelsmu
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: gem
Source: https://github.com/hamelsmu/hamel/tree/main/plugins/hamel-tools/skills/gem
Command: npx skills add https://github.com/hamelsmu/hamel --skill gem-hamelsmu

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Enable multimodal understanding and image generation for files and links that require vision-enabled AI, removing the friction of manually summarizing documents, extracting information from images or videos, and producing visual assets from text prompts.

Core Features & Use Cases

  • Analyze PDFs, images, videos, and YouTube links to produce summaries, key points, and comparisons across documents.
  • Transcribe and generate chapter timestamps for local MP4 files and YouTube videos, and extract contextual insights from long-form media.
  • Generate and edit images using Nano Banana Pro (default) or alternative Gemini image models for creative assets and visual experimentation.

Quick Start

Summarize the attached PDF and extract the five most important takeaways from document.pdf.

Frequently Asked Questions about gem

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract insights and summarize a PDF using multimodal AI?

To extract insights and summarize a PDF using multimodal AI, you process the document file to generate structured summaries, key points, and comparisons across content. This requires a configured GEMINI_API_KEY to access Google's Gemini vision models for analysis.

Can I generate images from text prompts using Gemini models?

You can generate images from text prompts using Gemini's image-generation models like Nano Banana Pro. This allows you to create and edit visual assets directly from text descriptions for content creation and research tasks.

How do I transcribe a YouTube video and generate chapter timestamps?

To transcribe a YouTube video and generate chapter timestamps, you provide the video URL to the multimodal processor. It extracts the audiovisual content to produce a text transcription alongside contextual timestamps for long-form media.

Do I need a GEMINI_API_KEY to analyze images and videos?

Yes, you need a configured GEMINI_API_KEY to analyze images and videos. Access to Google's Gemini multimodal models is required to perform vision analysis, extract structured insights, and execute image-generation tasks.

What is the best way to compare insights across multiple PDF documents?

The best way to compare insights across multiple PDF documents is using multimodal AI to process and analyze the files collectively. This approach automatically extracts structured summaries and key points to highlight differences and similarities.