weights-and-biases

Track ML experiments, run hyperparameter sweeps, and manage model artifacts with Weights & Biases.

Updated Jul 3, 2026
One-click install
npx skills add https://github.com/CHENHUI-X/toolbox --skill weights-and-biases-chenhui-x
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: weights-and-biases
Source: https://github.com/CHENHUI-X/toolbox/tree/main/custom-skills/evaluation/weights-and-biases
Command: npx skills add https://github.com/CHENHUI-X/toolbox --skill weights-and-biases-chenhui-x

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires wandb, and includes references (resource) components.

What problem does it solve? Machine learning teams lose track of experiments, hyperparameters, and model versions when training runs are scattered across notebooks and scripts. This Skill provides complete guidance for instrumenting training code with Weights & Biases so every run, metric, and artifact is logged, comparable, and reproducible. ## Core Features & Use Cases - Experiment Tracking: Log metrics, configs, media, and system stats from PyTorch, TensorFlow, Keras, HuggingFace, and PyTorch Lightning training loops. - Hyperparameter Sweeps: Run grid, random, or Bayesian optimization searches with early termination and parallel agents across GPUs. - Artifacts & Model Registry: Version datasets and models with lineage tracking, aliases, and promotion workflows from staging to production. - Use Case: A team fine-tuning a BERT classifier can launch a Bayesian sweep over learning rate and batch size, compare 50 runs in a shared dashboard, and link the best checkpoint to the production model registry. ## Quick Start Set up Weights & Biases tracking for my PyTorch training script so metrics and the final model checkpoint are logged to a project dashboard.

Frequently Asked Questions about weights-and-biases

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I track PyTorch training experiments with Weights & Biases?

Call wandb.init with your project name and config, then call wandb.log with metrics inside your training loop. Use wandb.watch to automatically log gradients and model parameters, and wandb.save or Artifacts to upload checkpoints.

How do I run a hyperparameter sweep with wandb?

Define a sweep config with a search method (grid, random, or bayes), a target metric, and parameter distributions, then create it with wandb.sweep. Launch one or more agents with wandb.agent pointing to your training function to execute trials.

Does Weights & Biases integrate with HuggingFace Transformers?

Yes. Set report_to="wandb" in TrainingArguments and the HuggingFace Trainer automatically logs metrics, evaluation results, and checkpoints to W&B. You can also add custom WandbCallback subclasses for additional logging.

What is the difference between grid, random, and Bayesian sweeps in wandb?

Grid search exhaustively tries all value combinations, random search samples combinations without learning, and Bayesian optimization uses past results to sample promising regions. Bayesian is recommended for expensive training runs with limited compute budgets.

How do I version datasets and models with W&B Artifacts?

Create a wandb.Artifact with a name and type, add files or directories, and log it with run.log_artifact. W&B automatically versions artifacts, tracks lineage between runs, and supports aliases like latest, staging, and production.

Can I use wandb without an internet connection?

Yes. Set WANDB_MODE=offline before initializing your run, and all metrics are stored locally. Later, run wandb sync on the run directory to upload the logged data to the W&B servers.