austn-tools

Automate TTS and image generation through local AI services.

Updated Feb 2, 2026
One-click install
npx skills add https://github.com/frogr/austnomaton --skill austn-tools
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: austn-tools
Source: https://github.com/frogr/austnomaton/tree/main/skills/austn-tools
Command: npx skills add https://github.com/frogr/austnomaton --skill austn-tools

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Austn-tools consolidates access to Austin's local GPU-powered AI services for content creation, enabling fast generation of audio and visuals without relying on remote cloud providers.

Core Features & Use Cases

  • TTS: generate speech from text with configurable voices and durations.
  • Image generation: create AI visuals via a local GPU pipeline (ComfyUI) for thumbnails and assets.
  • Background removal, vector tracing, and audio stems for post-production and asset preparation.
  • Use case: a creator wants to produce a short narrated clip with a local thumbnail using only on-device resources.

Quick Start

Use the austn-tools skill to generate a voiceover and an image locally, for example:

  • /austn-tools tts "Hello world"
  • /austn-tools image "Robot mascot, friendly, digital art"

Frequently Asked Questions about austn-tools

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate text-to-speech audio locally without using cloud providers?

Local text-to-speech generation is handled by automating Chatterbox TTS through a local GPU pipeline. You can generate speech from text with configurable voices and durations entirely on-device.

Can I generate AI images locally using ComfyUI for digital assets?

Yes, AI image generation is automated via a local ComfyUI pipeline. It leverages local GPU resources to create AI visuals like thumbnails and assets without remote cloud dependencies.

What's the best way to remove backgrounds and trace vectors for post-production?

Background removal and vector tracing are executed as local post-production asset preparation steps. The workflow automates these tasks alongside audio stem extraction using local GPU infrastructure.

Do I need a running local infrastructure to use text-to-speech and image generation features?

Yes, you need Austn's local infrastructure running and reachable. The solution assumes access to local GPU-powered services like Chatterbox TTS and ComfyUI to execute the content creation workflow.

How does local GPU-powered content generation compare to remote cloud AI services?

Local GPU content generation provides fast creation of audio and visuals without relying on remote cloud providers. This approach keeps the entire workflow on-device for thumbnails, voiceovers, and asset preparation.