paperless-canonical-ingestion

Automates discovery, deduplication, and upsert of documents from cloud storage and email to Paperless with owner-routing tags.

15|7|Updated Aug 6, 2025
One-click install
npx skills add https://github.com/hvkshetry/StewardOS --skill paperless-canonical-ingestion
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: paperless-canonical-ingestion
Source: https://github.com/hvkshetry/StewardOS/tree/main/skills/personas/chief-of-staff/paperless-canonical-ingestion
Command: npx skills add https://github.com/hvkshetry/StewardOS --skill paperless-canonical-ingestion

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill streamlines the process of ingesting family office records into Paperless, ensuring that documents are deduplicated, canonicalized, and properly tagged for easy retrieval and agent routing.

Core Features & Use Cases

  • Multi-Source Discovery: Finds candidate files across Google Drive, Gmail, OneDrive, and SharePoint.
  • Deduplication & Canonicalization: Identifies and selects the single best version of a document, even with minor variations.
  • Automated Tagging: Normalizes titles, document types, and correspondents, and applies owner-routing tags based on a filing matrix.
  • Use Case: When seeding your Paperless system with historical tax documents, legal agreements, and property records, this Skill ensures that only the most authoritative versions are ingested and correctly categorized.

Quick Start

Use the paperless-canonical-ingestion skill to discover and ingest all tax-related documents from your Google Drive into Paperless.

Frequently Asked Questions about paperless-canonical-ingestion

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I deduplicate family office records when ingesting documents from Google Drive and Gmail?

Canonical ingestion discovers candidate files across Google Drive, Gmail, OneDrive, and SharePoint, then selects the single best version of each document. It normalizes metadata and applies owner-routing tags before upserting the deduplicated records into Paperless.

What is the best way to automate tagging and metadata normalization for historical tax documents in Paperless?

Automated tagging uses a filing matrix to normalize titles, document types, and correspondents for historical tax documents. This process applies owner-routing tags to ensure records are correctly categorized and easily retrieved by agents.

Can I ingest historical legal agreements and property records from Microsoft Office APIs into Paperless?

Yes, you can ingest legal agreements and property records by integrating with Microsoft Office APIs for OneDrive and SharePoint. The Skill discovers historical files, deduplicates them, and upserts canonical versions with normalized metadata into Paperless.

How does document deduplication handle minor variations when seeding a Paperless system?

Deduplication handles minor variations by evaluating discovered candidate files and selecting the single most authoritative version of each document. This ensures that when seeding your Paperless system, only the canonical copy is ingested and categorized.

Do I need a filing matrix to route ingested documents to specific owners in Paperless?

Yes, a filing matrix is utilized to apply owner-routing tags and ensure consistent data enrichment. It maps normalized metadata to specific owners, allowing ingested documents to be properly categorized for easy retrieval and agent routing.