data-scraper-agent

Scrape public sources, enrich with Gemini Flash, and store results.

Updated Mar 31, 2026
One-click install
npx skills add https://github.com/GGEdu/claude-god-mode-template --skill data-scraper-agent-ggedu
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-scraper-agent
Source: https://github.com/GGEdu/claude-god-mode-template/tree/main/skills/data-scraper-agent
Command: npx skills add https://github.com/GGEdu/claude-god-mode-template --skill data-scraper-agent-ggedu

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Automates end-to-end data collection from public sources by scraping, enriching, and storing results.

Core Features & Use Cases

  • Collect data from public sources on a schedule.
  • Enrich results with AI (Gemini Flash) for scoring and summaries.
  • Store results in Notion, Google Sheets, or Supabase and learn from user feedback.

Quick Start

Set up the project, configure storage, and run the GitHub Actions workflow to start collecting data.

Frequently Asked Questions about data-scraper-agent

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate web scraping and store results in Notion or Google Sheets?

Automate end-to-end public data collection by scraping sources on a schedule, enriching results with Gemini Flash, and storing data directly in Notion, Google Sheets, or Supabase.

Can I use GitHub Actions to schedule web scraping tasks for price monitoring?

Yes, GitHub Actions provides the scheduling mechanism to run automated web scraping tasks for monitoring prices, jobs, news, or sports listings at defined intervals.

How does AI enrichment work for scraped data collection?

AI enrichment uses Gemini Flash within free quotas to process scraped data collection, generating automated scoring and summaries to add context to raw public source results.

Does this data scraping approach require a Python framework to run?

Yes, the automated data scraping workflow requires a Python-based scraper framework to extract public data, enrich it, and store results in your configured database.

What is the best way to collect and enrich GitHub repository data automatically?

The best way is an end-to-end automation framework that scrapes public GitHub repos, applies Gemini Flash AI for summary scoring, and stores results in Supabase or Sheets.

Can scraped data collection improve accuracy using user feedback?

Yes, the automated data collection pipeline includes a feedback loop that learns from user input to continuously refine the scraping, enrichment, and storage processes.