Digital Public Library of America avatar

Digital Public Library of America

Official

@dpla · United States of America

0Followers
|
18Public Repos
|
18Published Skills

DPLA brings together the riches of America’s libraries, archives, and museums, and makes them freely available to the world.

Skills Distribution
DomainData Systems...Metadata Harvesting (40%)Data Pipeline Orch.. (30%)Cloud Storage Mana.. (30%)

Agent Skills by Digital Public Library of America

Showing 18 vetted skills indexed across 1 GitHub repositories.

dpladpla
35

dpla-ingest-status

Aggregate ingest status across hubs from status logs and orchestrator state.

Official
Intermediate
dpladpla
35

dpla-hub-info

Display i3.conf hub configuration details from the hub-info script.

Official
Intermediate
dpladpla
35

dpla-monthly-emails

Generate and distribute monthly DPLA hub scheduling emails from i3.conf.

Official
Advanced
dpladpla
35

dpla-run-ingest

Automate hub ingest runs by selecting runbooks and verifying output artifacts.

Official
Intermediate
dpladpla
35

dpla-verify-and-notify

Verify ingest outcomes and notify stakeholders via Slack or email.

Official
Intermediate
dpladpla
35

dpla-monitor-ingest-remap

Monitor IngestRemap pipeline progress across hubs via _SUCCESS markers.

Official
Intermediate
dpladpla
35

send-email

Send ingest summary emails to hub contacts from i3.conf mapping output.

Official
Intermediate
dpladpla
35

dpla-oai-harvest-watch

Parse OAI harvest logs to report set-by-set progress and ETA.

Official
Intermediate
dpladpla
35

dpla-ingest-debug

Diagnose and fix DPLA hub ingestion failures across pipeline stages.

Official
Advanced
dpladpla
35

dpla-s3-ops

Sync hub data to S3 and verify JSONL exports using the dpla profile.

Official
Intermediate
dpladpla
35

dpla-community-webs-ingest

Automate Community Webs data harvesting and ingestion into the DPLA pipeline via EC2 and SSM.

Official
Intermediate
dpladpla
35

dpla-script-workflow

Enforce POSIX bash compatibility and common.sh usage for shell and Python scripts.

Official
Intermediate
dpladpla
35

dpla-staged-report

Identify hubs with newly staged JSONL data in S3 for a given month.

Official
Intermediate
dpladpla
35

dpla-orchestrator

Run the DPLA ingest orchestrator for harvest, mapping, enrichment, JSONL export, anomaly detection, and S3 sync.

Official
Advanced
dpladpla
35

DPLA Hub Ingest

Automate DPLA hub ingest from harvest to S3 sync on EC2.

Official
Advanced
dpladpla
35

dpla-takedown

Remove DPLA items from the live site and search index.

Official
Advanced
dpladpla
35

dpla-s3-and-aws

Coordinate S3 data synchronization and AWS operations for the DPLA ingestion pipeline.

Official
Intermediate
dpladpla
35

s3-latest

Identify and retrieve the most recent S3 exports for a specified hub.

Official
Basic

Frequently Asked Questions About Digital Public Library of America

FAQPage Schema
What specific tasks can be performed using these ingestion capabilities?

These capabilities enable the harvesting of metadata via OAI-PMH, the orchestration of data mapping and enrichment, the synchronization of JSONL exports to S3, and the management of anomaly detection across distributed library hubs.

Which technical personas are intended to utilize these ingestion resources?

These resources are designed for data engineers, library systems administrators, and metadata specialists responsible for maintaining large-scale cultural heritage repositories and ensuring the synchronization of records between regional hubs and central storage.

What are the primary infrastructure dependencies for running these ingestion processes?

The ingestion processes require an AWS environment, specifically EC2 instances for compute, S3 buckets for data storage, and SSM for configuration management, alongside POSIX-compliant shell environments for executing the provided ingest orchestrator.