archive-crawler

Identify meaningful writing and ideas from approved personal file archives.

Updated Jun 20, 2026
One-click install
npx skills add https://github.com/Sigmacodeat/subsumio-web --skill archive-crawler-sigmacodeat
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: archive-crawler
Source: https://github.com/Sigmacodeat/subsumio-web/tree/main/server/skills/archive-crawler
Command: npx skills add https://github.com/Sigmacodeat/subsumio-web --skill archive-crawler-sigmacodeat

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps people find meaningful writing, ideas, relationships, and personal knowledge buried inside large archives while avoiding irrelevant or sensitive files.

Core Features & Use Cases

  • Archive Discovery: Inventory and explore approved file archives from local storage, cloud backups, email exports, and similar sources.
  • High-Value Filtering: Identify personal writing, ideas, conversations, creative work, and origin stories while skipping system files, credentials, and low-value content.
  • Interactive Ingestion: Review surfaced items, capture reactions, and organize valuable discoveries into structured knowledge pages.

Quick Start

Ask the archive-crawler skill to crawl my approved archive paths and surface the most valuable items worth keeping.

Frequently Asked Questions about archive-crawler

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I discover valuable writing and ideas buried in personal file archives?

Surfacing valuable content from personal file archives requires identifying meaningful writing and relationship materials while avoiding irrelevant files. This archive discovery process inventories approved paths from local storage, cloud backups, and email exports to selectively ingest high-value personal knowledge.

How do I filter out sensitive files and system data when extracting personal knowledge from a large archive?

Filtering sensitive files during personal archive ingestion requires explicit path allow-lists and content filtering. This approach skips system files and credentials, selectively identifying high-value personal writing and creative work while tracking manifests for structured knowledge-page generation.

Can I crawl email exports and local drives to find old personal conversations and creative work?

Yes, crawling email exports and local drives to find old conversations and creative work is fully supported. The archive discovery process applies to local storage, cloud backups, and email exports, using manifest tracking to surface relationship materials and historical file collections.

What is the best way to organize discovered items from historical file collections into structured knowledge pages?

Organizing discovered items from historical file collections into structured knowledge pages requires an interactive ingestion workflow. This process lets you review surfaced items, capture reactions, and systematically generate structured personal knowledge outputs from valuable archive discoveries.

Do I need explicit archive path allow-lists to perform selective ingestion on my personal data?

Yes, explicit archive path allow-lists are required to perform selective ingestion on personal data. This prerequisite ensures the crawler only accesses approved file archives, preventing unintended discovery of sensitive files while maintaining manifest tracking for structured knowledge generation.