archive-crawler

Automate discovery and ingestion of personal content from file archives.

Updated Jun 2, 2026
One-click install
npx skills add https://github.com/Ninatuzi/gbrain --skill archive-crawler-ninatuzi
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: archive-crawler
Source: https://github.com/Ninatuzi/gbrain/tree/main/skills/archive-crawler
Command: npx skills add https://github.com/Ninatuzi/gbrain --skill archive-crawler-ninatuzi

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill solves the problem of digital hoarding and information fragmentation by systematically scanning personal archives to identify and surface high-value content like journals, ideas, and relationship-defining correspondence.

Core Features & Use Cases

  • Intelligent Filtering: Automatically distinguishes between noise (system files, binaries) and high-signal content (personal writing, creative work, reflections).
  • Schema-Generic Ingestion: Maps discovered content into your existing brain structure based on your personal filing rules.
  • Use Case: If you have years of scattered emails, local documents, and old project files, this skill will inventory them, allow you to review the most meaningful items, and file them into your knowledge graph with proper back-links.

Quick Start

Ask the agent to crawl my archive and surface the writing worth keeping to begin the discovery process.

Frequently Asked Questions about archive-crawler

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract high-value personal content from scattered file archives?

To extract high-value personal content from scattered file archives, you can automate discovery and ingestion by applying a gold-filter to identify journals, ideas, and relationship material while enforcing strict safety allow-lists for file access.

Can I ingest personal writing from local directories, Dropbox, and email exports into a knowledge graph?

Yes, you can ingest personal writing from local directories, Dropbox, and email exports into a knowledge graph by applying a schema-generic ingestion process that maps discovered content based on your personal filing rules with proper back-links.

How does intelligent filtering separate noise from high-signal content in personal archives?

Intelligent filtering separates noise from high-signal content in personal archives by automatically distinguishing system files and binaries from high-signal content like personal writing, creative work, and reflections.

Do I need a configured scan_paths allow-list to securely traverse sensitive personal data?

Yes, you need a configured gbrain.yml scan_paths allow-list to securely traverse sensitive personal data, ensuring controlled access and preventing unauthorized file discovery during the archiving process.

What is the best way to surface writing worth keeping from years of scattered documents?

The best way to surface writing worth keeping from years of scattered documents is to crawl your archive, inventory the files, review the most meaningful items, and file them into your personal knowledge graph.

Are there limitations when applying a gold-filter to identify relationship material in file archives?

A limitation when applying a gold-filter to identify relationship material is that file access is strictly bound by your configured safety allow-lists, meaning files outside the allowed paths will not be discovered or ingested.