crawl-sites

Crawl configured websites and extract HTML content by URL patterns.

39|6|Updated Jan 5, 2026
One-click install
npx skills add https://github.com/the-agency-ai/the-agency --skill crawl-sites
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: crawl-sites
Source: https://github.com/the-agency-ai/the-agency/tree/main/.claude/skills/crawl-sites
Command: npx skills add https://github.com/the-agency-ai/the-agency --skill crawl-sites

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the process of crawling websites and extracting structured content, reducing manual browsing efforts and enabling comprehensive data collection.

Core Features & Use Cases

  • Site Content Extraction: Crawl specified sites and retrieve HTML content based on URL patterns.
  • Customizable Scope: Target specific sites or entire configurations for targeted data harvesting.
  • Use Case: Imagine monitoring multiple documentation sites for updates. Use this Skill to crawl them regularly and collect new articles or changes automatically.

Quick Start

Use the crawl-sites skill to scan your configured websites and gather all relevant content into your output folder.

Frequently Asked Questions about crawl-sites

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract structured content from websites at scale?

Web content extraction at scale is automated by crawling specified sites and retrieving HTML content based on URL patterns. This reduces manual browsing efforts by supporting multiple providers and site configurations for comprehensive web data collection.

How does automated site scraping handle different site architectures?

Automated site scraping ensures compatibility with different site architectures by supporting multiple providers and customizable site configurations. This allows targeted data harvesting across specific sites or entire configuration sets.

Can I target specific URL patterns for web crawling instead of an entire site?

Web crawling can target specific URL patterns rather than entire sites. The customizable scope allows you to specify individual sites or entire configurations for targeted data harvesting based on your exact requirements.

What is the best way to monitor documentation sites for content updates automatically?

Monitoring documentation sites for updates is handled by configuring the crawler to scan those sites regularly. It automatically collects new articles or changes, gathering all relevant content into your output folder.

Do I need any external dependencies to start web scraping with this tool?

No external dependencies are required to start web scraping. The tool operates independently using its built-in scripts and references to automate website crawling and structured content extraction.