archive-crawler

Crawl file archives to discover and triage high-value personal content.

Updated Jun 10, 2026
One-click install
npx skills add https://github.com/starlink-awaken/omostation-gbrain --skill archive-crawler-starlink-awaken
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: archive-crawler
Source: https://github.com/starlink-awaken/omostation-gbrain/tree/main/skills/archive-crawler
Command: npx skills add https://github.com/starlink-awaken/omostation-gbrain --skill archive-crawler-starlink-awaken

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill solves the problem of digital hoarding and information fragmentation by systematically surfacing high-value personal content—such as journals, ideas, and meaningful correspondence—from vast, unorganized file archives.

Core Features & Use Cases

  • Universal Archivist: Connects to diverse sources including local drives, Dropbox, Gmail takeouts, and legacy formats like mbox or PST.
  • Gold Filtering: Automatically distinguishes between high-signal personal writing and noise like system files or bulk receipts.
  • Interactive Ingestion: Creates a structured manifest and allows you to review and file content into your brain-first knowledge base with exact-quote capture.

Quick Start

Ask the agent to crawl my archive and surface the writing worth keeping to begin the discovery process.

Frequently Asked Questions about archive-crawler

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I find high-value personal content in unorganized file archives?

To surface high-value personal content from unorganized file archives, automate the discovery and triage process to systematically extract journals, ideas, and meaningful correspondence from local drives, cloud storage, and email exports.

Can I extract personal writing from Gmail takeouts and PST files?

Yes, you can extract personal writing from Gmail takeouts and PST files. The archiving process connects to diverse sources including legacy formats like mbox and PST to identify and ingest meaningful correspondence.

What is the best way to filter noise from system files when archiving personal knowledge?

The best way to filter noise from system files when archiving personal knowledge is applying gold filtering, which automatically distinguishes high-signal personal writing from bulk receipts and irrelevant system files.

How do I ensure privacy when mining personal archives for data?

To ensure privacy when mining personal archives for data, the process requires strict adherence to user-defined allow-lists and schema-generic filing rules, guaranteeing organizational consistency and data privacy.

Does the triage process support interactive ingestion into a personal knowledge base?

Yes, the triage process supports interactive ingestion into a personal knowledge base by creating a structured manifest, allowing you to review and file content with exact-quote capture.