exploratory-data-analysis

Analyze scientific data files across 200+ formats and generate markdown reports.

Updated Mar 15, 2026
One-click install
npx skills add https://github.com/sagunkayastha/claude_skills_collection --skill exploratory-data-analysis-sagunkayastha
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: exploratory-data-analysis
Source: https://github.com/sagunkayastha/claude_skills_collection/tree/main/data-analysis-visualization/exploratory-data-analysis
Command: npx skills add https://github.com/sagunkayastha/claude_skills_collection --skill exploratory-data-analysis-sagunkayastha

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires pandas, numpy, biopython, pillow, scikit-image, h5py, nd2reader, czifile, pydicom, tifffile, pymzml, nmrglue, gemmi, rdkit, ase, mdanalysis, pyBigWig, pybedtools, pyfasta, pyfaidx, pysam, cyvcf2, gffutils, pyarrow, openpyxl, json, yaml, toml, configparser, zipfile, tarfile, gzip, bz2, netCDF4, rasterio, geopandas, scipy, matplotlib, seaborn, plotly, networkx, sympy, matlab, simpy, dask, vaex, fluidsim, sec-filings, fredapi, alpha-vantage, modal, dnanexus, latchbio, omero, opentrons, pytorch-lightning, transformers, scikit-learn, shap, pymc, pydicom, histolab, pathml, esm, glycoengineering, adaptyv, iso13485, uniprot, pdb, pubchem, chembl, ensembl, gnomad, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill automates the process of understanding and analyzing diverse scientific data files, saving researchers significant time and effort in data exploration.

Core Features & Use Cases

  • Automated File Type Detection: Identifies over 200 scientific file formats.
  • Format-Specific Analysis: Performs tailored EDA based on file type (chemistry, biology, imaging, etc.).
  • Comprehensive Reporting: Generates detailed markdown reports with findings and recommendations.
  • Use Case: Upload a .fastq file and get a report detailing read counts, quality scores, and GC content, along with recommendations for downstream analysis.

Quick Start

Use the exploratory-data-analysis skill to analyze the file 'my_data.pdb'.

Frequently Asked Questions about exploratory-data-analysis

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I perform exploratory data analysis on scientific data files like PDB or FASTQ?

You can perform exploratory data analysis by feeding scientific data files into the system, which automatically detects the format, extracts metadata, assesses data quality, and outputs a markdown report with downstream recommendations.

Can I analyze microscopy and spectroscopy files automatically without writing custom scripts?

Yes, automated file type detection identifies over 200 scientific formats including microscopy and spectroscopy files, performing tailored EDA and generating detailed reports without requiring custom scripts.

What scientific file formats are supported for automated data quality assessment?

Supported formats span chemistry, bioinformatics, microscopy, spectroscopy, proteomics, metabolomics, and general scientific data, including PDB, FASTQ, netCDF, DICOM, and various genomic and imaging formats.

Does this tool work with bioinformatics formats like VCF and GFF for genomic data analysis?

Yes, it supports bioinformatics formats like VCF and GFF through specialized dependencies, extracting format-specific metadata and providing downstream analysis recommendations directly within the generated markdown report.

How do I get downstream analysis recommendations for raw metabolomics or proteomics data?

Processing raw metabolomics or proteomics data through the automated analysis pipeline generates a detailed markdown report that includes findings, data quality assessments, and specific recommendations for downstream analysis.

What is the best way to extract metadata and assess data quality from diverse chemistry files?

The best way to extract metadata from diverse chemistry files is using an automated EDA pipeline that identifies formats like MZML or SDF, assesses structural data quality, and documents the results in a comprehensive markdown report.