tooluniverse-expression-data-retrieval

Retrieve gene expression and multi-omics datasets from ArrayExpress and BioStudies.

1.6k|244|Updated Mar 3, 2025
One-click install
npx skills add https://github.com/mims-harvard/ToolUniverse --skill tooluniverse-expression-data-retrieval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: tooluniverse-expression-data-retrieval
Source: https://github.com/mims-harvard/ToolUniverse/tree/main/skills/tooluniverse-expression-data-retrieval
Command: npx skills add https://github.com/mims-harvard/ToolUniverse --skill tooluniverse-expression-data-retrieval

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill addresses the need to locate, disambiguate, and retrieve gene expression experiments and multi-omics datasets from public repositories, and to produce structured datasets profiles with metadata, sample information, and download links.

Core Features & Use Cases

  • Gene/Experiment Disambiguation: resolves official gene identifiers and aliases to enable precise queries.
  • Data Retrieval & Profiling: searches ArrayExpress (E-MTAB, E-GEOD) and BioStudies (S-BSST), gathers metadata and sample information, and compiles a dataset profile with download links.
  • Multi-Omics Integration: surfaces multi-omics studies from BioStudies and guides data integration.
  • Use Case: A researcher needs RNA-seq data for human liver; the skill retrieves relevant experiments, quality indicators, and links to processed and raw data.

Quick Start

Use the skill to search for "RNA-seq breast cancer" in Homo sapiens and generate a dataset profile with top results and links to data files.

Frequently Asked Questions about tooluniverse-expression-data-retrieval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I retrieve gene expression datasets from ArrayExpress and BioStudies?

Retrieve gene expression datasets by searching ArrayExpress (E-MTAB, E-GEOD) and BioStudies (S-BSST) workflows, which compile structured dataset profiles containing metadata, sample information, and direct download links.

How does gene disambiguation work for multi-omics data retrieval?

Gene disambiguation resolves official gene identifiers and aliases before querying public repositories, ensuring precise multi-omics data retrieval and accurate dataset profiling across different species.

Can I use this to fetch RNA-seq data for a specific species like Homo sapiens?

Yes, you can fetch RNA-seq data for Homo sapiens across species. The retrieval process generates dataset profiles with top results, quality indicators, and links to both raw and processed data files.

What is included in a structured dataset profile for gene expression experiments?

A structured dataset profile includes experiment metadata, sample information, quality assessment indicators, and download links, enforcing consistent reporting to enable reliable data reuse from public repositories.

What's the best way to find multi-omics studies in BioStudies?

Use data profiling workflows to search BioStudies (S-BSST) for multi-omics studies, which surfaces relevant datasets, gathers metadata, and guides subsequent multi-omics data integration.

Why do I need quality assessment reports for public repository data?

Quality assessment reports provide quality indicators for retrieved gene expression experiments, enforcing disambiguation and consistent reporting to ensure reliable data reuse from public repositories.