fleet-inference

Query local and cloud AI models across multiple providers.

Updated Apr 2, 2026
One-click install
npx skills add https://github.com/phyter1/seed --skill fleet-inference
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: fleet-inference
Source: https://github.com/phyter1/seed/tree/main/.claude/skills/fleet-inference
Command: npx skills add https://github.com/phyter1/seed --skill fleet-inference

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill enables direct querying of any model within a fleet, whether local or cloud-based, reducing latency and simplifying model access for AI tasks.

Core Features & Use Cases

  • Model Discovery and Access: Listed models can be retrieved across various providers like MLX, Ollama, Cerebras, Groq, Gemini, and OpenRouter.
  • Flexible Querying: Send prompts to specific models or allow auto-routing to the best available provider for tasks such as inference, analysis, and generation.
  • Use Case: Easily run large language model inferences on the local network or cloud, such as generating content or analyzing data without manual switching between endpoints.

Quick Start

Query models by specifying a prompt or list all available models, enabling quick integration into AI workflows.

Frequently Asked Questions about fleet-inference

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I query multiple local and cloud AI models from a single interface?

Query multiple local and cloud AI models by sending prompts to specific endpoints or using auto-routing to dynamically select the best available provider for your inference tasks. This simplifies large-scale generation and analysis across diverse deployment environments.

Does this inference approach support both local and cloud providers like Ollama and Gemini?

Yes, this approach supports both local and cloud providers including MLX, Ollama, Cerebras, Groq, Gemini, and OpenRouter. It ensures secure, conditioned access to AI endpoints for enterprise or personal use cases across these diverse environments.

What is the best way to run large language model inferences across a diverse deployment fleet?

The best way to run large language model inferences across a fleet is by facilitating direct and flexible access to multiple local and cloud models. This reduces latency and optimizes performance through dynamic model selection.

Can I auto-route prompts to the best available provider for on-demand inference?

Yes, you can auto-route prompts to the best available provider for on-demand inference tasks. This flexible querying mechanism supports dynamic model selection to optimize performance for inference, analysis, and generation.

Are there limitations when switching between local and cloud model endpoints for inference?

Switching between local and cloud model endpoints does not present major limitations here, as the system is designed to simplify model access and reduce latency. It ensures secure, conditioned access to AI endpoints across diverse deployment environments.