tooluniverse-sequence-retrieval

Retrieve and consolidate biological sequences across public databases with cross-references.

1.6k|244|Updated Mar 3, 2025
One-click install
npx skills add https://github.com/mims-harvard/ToolUniverse --skill tooluniverse-sequence-retrieval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: tooluniverse-sequence-retrieval
Source: https://github.com/mims-harvard/ToolUniverse/tree/main/skills/tooluniverse-sequence-retrieval
Command: npx skills add https://github.com/mims-harvard/ToolUniverse --skill tooluniverse-sequence-retrieval

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Biologists and bioinformaticians waste time and risk errors when manually searching multiple databases to acquire nucleotide and protein sequences for analysis, annotation, and reporting.

Core Features & Use Cases

  • Unified search across NCBI Nucleotide and ENA GenBank entries for DNA, RNA, and protein sequences.
  • Gene/organism disambiguation to resolve exact sequences for a given gene and species.
  • Cross-database reporting & downloads with metadata, cross-references, and downloadable formats (FASTA, GenBank).
  • Use Case: Retrieve the Homo sapiens BRCA1 mRNA sequence and compare it to a bacterial gene for comparative genomics.

Quick Start

Example: Use the tool to search for Escherichia coli K-12 complete genome, fetch top accessions, and download the FASTA sequence for downstream analysis.

Frequently Asked Questions about tooluniverse-sequence-retrieval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I retrieve a FASTA sequence from NCBI and ENA for a specific gene and organism?

To retrieve a FASTA sequence, this tool disambiguates gene and organism identifiers across NCBI Nucleotide and ENA GenBank entries. It consolidates cross-database references and returns downloadable FASTA formats for downstream analysis.

Can I fetch complete genomes and protein sequences using a single sequence retrieval tool?

Yes, you can fetch complete genomes and protein sequences using this sequence retrieval tool. It applies an internal workflow handling accession prefixes and ENA fallback to return standardized biological data.

What is the best way to get cross-database references for RefSeq and GenBank entries?

The best way to get cross-database references is using a unified search across NCBI and ENA. This tool disambiguates identifiers and reports metadata with cross-references for RefSeq and GenBank entries.

How does sequence retrieval handle disambiguation when searching for genes across multiple public databases?

Sequence retrieval handles disambiguation by resolving exact sequences for a given gene and species across multiple public databases. It applies accession prefix logic and ENA fallback to ensure accurate data extraction.

Does the tool support downloading GenBank formats for comparative genomics analysis?

Yes, the tool supports downloading GenBank formats for comparative genomics analysis. It retrieves and consolidates biological sequences with metadata and cross-references for downstream comparative tasks.