lamindb

Automate biological dataset management and lineage tracking with LaminDB artifacts, runs, and transforms.

18|1|Updated Dec 27, 2025
One-click install
npx skills add https://github.com/LogauaEngstrom/claude-scientific-skills --skill lamindb-logauaengstrom
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: lamindb
Source: https://github.com/LogauaEngstrom/claude-scientific-skills/tree/main/scientific-skills/lamindb
Command: npx skills add https://github.com/LogauaEngstrom/claude-scientific-skills --skill lamindb-logauaengstrom

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

LaminDB unifies data management for biology by providing a queryable, traceable, and reproducible framework that tracks artifacts, records, runs, and transforms along with ontology annotations.

Core Features & Use Cases

  • Centralized data lakehouse for biological datasets with artifact versioning and lineage tracking.
  • Ontology integration via cell types, genes, tissues; schema validation and curation workflows.
  • Seamless integrations with workflow managers (Nextflow, Snakemake), ML platforms (MLflow, W&B), and visualization tools (Vitessce).

Quick Start

Install LaminDB, initialize a local instance, and start tracking your first analysis to capture data provenance.

Frequently Asked Questions about lamindb

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I track data lineage and provenance for biological datasets?

Biological data lineage is tracked using artifact, run, and transform models to capture versioning and provenance across datasets. This framework records transformations end-to-end, ensuring scRNA-seq and proteomics workflows remain fully traceable and reproducible.

How do I annotate scRNA-seq data with standard cell type ontologies?

scRNA-seq data is annotated with standard cell type ontologies by applying schema validation and ontology-backed curation workflows. This process integrates controlled vocabularies for genes, tissues, and cell types directly into your data lakehouse for queryable biological metadata.

Can I integrate biological data management with Nextflow and MLflow?

Biological data management integrates seamlessly with workflow and ML platforms like Nextflow, MLflow, W&B, and Snakemake. This connectivity captures pipeline runs and tracks machine learning transforms alongside your biological artifacts without changing your existing tools.

What is the best way to make biomedical datasets FAIR and queryable?

Biomedical datasets are made FAIR and queryable by centralizing them in a data lakehouse with artifact versioning and ontology annotations. This approach enforces schema validation and tracks data lineage, ensuring records are findable, accessible, interoperable, and reusable.

Does LaminDB support multi-user and cloud environments for biological data?

LaminDB supports multi-user and cloud environments for biological data by allowing you to initialize instances across local and cloud storage. This setup maintains consistent schema validation and cross-system integrations for distributed teams analyzing large-scale datasets.

How do I visualize tracked proteomics or scRNA-seq artifacts?

Tracked proteomics and scRNA-seq artifacts are visualized through seamless integration with Vitessce. This integration connects your queryable, ontology-annotated biological records directly to interactive visualization tools for exploratory analysis.