archive-crawler

Crawl allow-listed archive paths to surface high-signal personal content for ingestion.

Updated May 11, 2026
One-click install
npx skills add https://github.com/yunusgungor/gbrain --skill archive-crawler-yunusgungor
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: archive-crawler
Source: https://github.com/yunusgungor/gbrain/tree/main/skills/archive-crawler
Command: npx skills add https://github.com/yunusgungor/gbrain --skill archive-crawler-yunusgungor

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

archive-crawler provides a safety-first engine to explore personal content archives, surfacing high-signal material for review and ingestion while excluding noise.

Core Features & Use Cases

  • Safety-first crawling with an explicit allow-list to prevent accidental ingestion of sensitive content.
  • Gold-filtered ingestion to triage and file personal writing, ideas, and related content into brain pages.
  • Manifest-driven workflow to track inventory, reactions, and progress across sources like local drives, cloud dumps, and Gmail takeouts.

Quick Start

Instruct the agent to crawl the allowed archive paths, surface high-signal personal content, and prepare it for ingestion.

Frequently Asked Questions about archive-crawler

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I crawl personal content archives to surface high-signal writing for review?

To crawl personal content archives, this Skill scans designated local drives, Dropbox, Backblaze, and Gmail takeouts to build manifests and triage high-signal personal content for ingestion. It requires an explicit archive path allow-list in gbrain.yml to enforce safety fences.

Can I use archive-crawler to ingest content from a Gmail takeout or Dropbox dump safely?

Yes, you can ingest content from a Gmail takeout or Dropbox dump safely. The Skill enforces safety fences using an explicit path allow-list in gbrain.yml, preventing over-scanning and ensuring sensitive data is not leaked during the triage process.

What is the best way to triage personal writing and ideas from local drives for brain pages?

The best way to triage personal writing and ideas for brain pages is using a manifest-driven workflow. The Skill filters noise, inventories items across local drives, and queues high-signal personal content for review and ingestion.

How does the manifest-driven workflow track inventory and progress across cloud dumps?

The manifest-driven workflow tracks inventory and progress across cloud dumps by building a comprehensive list of archived items. It records reactions, filters out noise, and queues high-signal personal content for ingestion.

Why does crawling personal archives require an explicit path allow-list in gbrain.yml?

Crawling personal archives requires an explicit path allow-list in gbrain.yml to prevent accidental ingestion of sensitive content. This safety-first mechanism ensures the crawler only scans designated directories, avoiding over-scanning and data leakage.

Do I need an archive path allow-list to prevent over-scanning during content triage?

Yes, you need an archive path allow-list to prevent over-scanning during content triage. The Skill mandates this configuration to enforce strict safety fences, ensuring only explicitly allowed local drives and cloud archives are scanned for ingestion.