lamindb

Manages biological datasets with provenance tracking and ontology-based annotation.

6|Updated Dec 30, 2025
One-click install
npx skills add https://github.com/pur3v4d3r/pur3-pkb-codebase --skill lamindb-pur3v4d3r
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: lamindb
Source: https://github.com/pur3v4d3r/pur3-pkb-codebase/tree/main/.claude/skills/__scientific-skills/lamindb
Command: npx skills add https://github.com/pur3v4d3r/pur3-pkb-codebase --skill lamindb-pur3v4d3r

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

LaminDB provides a unified framework to manage, annotate, and reproduce biological datasets by tracking data provenance, schema validation, and ontology-based metadata.

Core Features & Use Cases

  • End-to-end data management: track artifacts, runs, and transforms across analyses.
  • Ontology-driven curation: annotate data with standardized terms from biological ontologies.
  • Reproducible workflows: versioned data, environment capture, and lineage visualization for audits.

Quick Start

Install LaminDB, initialize a local instance, load a dataset, and start tracking to create a lineage-annotated artifact.

Frequently Asked Questions about lamindb

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I track data lineage for scRNA-seq and spatial workflows?

Track scRNA-seq and spatial data lineage by registering artifacts, runs, and transforms in LaminDB, which captures environment metadata and visualizes data provenance for biological workflows.

What is ontology-based annotation for biological data management?

Ontology-based annotation standardizes biological data by mapping metadata to controlled vocabulary terms. LaminDB enforces this via schema validation, ensuring datasets are consistently curated and reproducible.

Can I use LaminDB with Nextflow and Snakemake pipelines?

Yes, LaminDB integrates with workflow tools like Nextflow and Snakemake, allowing you to capture pipeline transforms, track versioned data inputs, and maintain provenance governance across pipeline runs.

Does LaminDB support schema validation for proteomics and clinical data?

Yes, LaminDB supports schema validation across proteomics and clinical data workflows, applying ontology-driven curation to standardize metadata and ensure biological dataset traceability and reproducibility.

What's the best way to make biological datasets reproducible?

The best way to make biological datasets reproducible is using LaminDB for versioned data storage, environment capture, and lineage visualization, ensuring all analysis transforms are traceable.

Do I need ML platforms to manage biological data lineage?

No, ML platforms are not required to manage biological data lineage, but LaminDB supports integration with ML platforms alongside workflow tools like Nextflow and Snakemake to extend provenance tracking.