gemini-video-analysis

Analyze video URLs and local files with Google GenAI API.

8|Updated Jul 26, 2026
One-click install
npx skills add https://github.com/joonlab/joonlab-claudecode-setting-for-share --skill gemini-video-analysis
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: gemini-video-analysis
Source: https://github.com/joonlab/joonlab-claudecode-setting-for-share/tree/main/claude/skills/gemini-video-analysis
Command: npx skills add https://github.com/joonlab/joonlab-claudecode-setting-for-share --skill gemini-video-analysis

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires google-genai, pydantic, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This skill solves the challenge of manually reviewing long or complex videos by automating the extraction of insights, summaries, structured data, and specific highlights using advanced multimodal AI.

Core Features & Use Cases

  • Multimodal Analysis: Processes both visual frames and audio tracks to provide comprehensive summaries, chapter breakdowns, and marketing insights.
  • Structured Extraction: Converts video content into machine-readable JSON schemas for database integration or automated workflows.
  • Use Case: A marketing team can use this to analyze hundreds of social media reels to identify key messaging, target audience sentiment, and potential improvements without watching every second manually.

Quick Start

Use the gemini-video-analysis skill to analyze the video at the provided URL and generate a summary report.

Frequently Asked Questions about gemini-video-analysis

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I analyze video content and extract marketing insights without watching the entire video?

Video analysis automates insight extraction by processing visual frames and audio tracks to provide summaries, chapter breakdowns, and marketing insights using multimodal AI. This eliminates manual review of long or complex videos.

Can I extract structured data from a YouTube URL for automated workflows?

Yes, you can analyze YouTube URLs and local video files to convert content into machine-readable JSON schemas. This structured extraction enables database integration and automated workflows without manual data entry.

How does multimodal video analysis work for chapter extraction and frame-by-frame inspection?

Multimodal video analysis works by processing both visual frames and audio tracks simultaneously to generate summaries and chapter breakdowns. It supports high-precision frame-by-frame inspection using the Google GenAI API.

Do I need a Google GenAI API key to process multimodal video inputs?

Yes, the Google GenAI API is required to process multimodal inputs. The skill utilizes this dependency to analyze video content with configurable temperature and model selection to balance cost and performance.

What is the best way to analyze social media reels for target audience sentiment at scale?

The best way to analyze social media reels at scale is using automated multimodal video analysis. It identifies key messaging and target audience sentiment across hundreds of videos without requiring manual viewing.

Are there limitations when using Gemini 3 for video summarization?

Limitations of Gemini 3 video summarization include balancing cost and performance through configurable temperature settings. Users must manage API usage and model selection to optimize outputs for diverse analysis tasks.