lamindb

Manage biological data assets with lineage tracking in Python.

33.0k|3.2k|Updated Oct 19, 2025
One-click install
npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill lamindb-k-dense-ai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: lamindb
Source: https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/scientific-skills/lamindb
Command: npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill lamindb-k-dense-ai

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

LaminDB provides a Python-based framework to manage biological data assets with full data lineage, provenance tracking, and FAIR metadata, enabling researchers to organize, validate, and reproduce analyses across multi-omics workflows.

Core Features & Use Cases

  • Core concepts and data lineage (Artifacts, Records, Runs, Transforms) for traceable research.
  • Data management, querying, validation, and ontology-driven annotation to enable FAIR data across scRNA-seq, spatial transcriptomics, proteomics, and clinical data.
  • Integrations with workflow managers and MLOps platforms (Nextflow, Snakemake, Weights & Biases, MLflow) plus deployment strategies for local and cloud environments.

Quick Start

Install LaminDB, initialize your instance, and start tracking data lineage.

Frequently Asked Questions about lamindb

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I track data lineage and provenance for scRNA-seq workflows in Python?

You can track scRNA-seq data lineage using a Python framework that captures artifacts, records, runs, and transforms. This ensures full provenance and reproducibility across biological workflows.

What is ontology-driven annotation for biological data management?

Ontology-driven annotation applies standardized vocabularies to biological data assets, enabling FAIR metadata. It allows researchers to query, validate, and organize multi-omics datasets consistently.

How do I make biological data queryable and reproducible across multi-omics pipelines?

To make biological data queryable and reproducible, use a framework that supports schema validation, versioning, and end-to-end lineage capture across spatial omics, proteomics, and clinical data.

Does this data management framework work with existing MLOps and workflow tools?

Yes, this biological data framework integrates with workflow managers and MLOps platforms like Nextflow, Snakemake, Weights & Biases, and MLflow for seamless pipeline tracking.

Can I use this lineage tracking system for both local and cloud environments?

Yes, this data lineage tracking system supports deployment strategies for both local and cloud environments, ensuring traceable research and interoperability across different infrastructures.