media-config

Configure DevClaw vision models and audio transcription services for multimodal messaging.

3|Updated Feb 13, 2026
One-click install
npx skills add https://github.com/jholhewres/devclaw-skills --skill media-config
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: media-config
Source: https://github.com/jholhewres/devclaw-skills/tree/main/skills/media-config
Command: npx skills add https://github.com/jholhewres/devclaw-skills --skill media-config

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill streamlines how DevClaw handles visual and audio content from various communication channels, ensuring efficient and accurate processing of images, videos, and audio messages.

Core Features & Use Cases

  • Vision Configuration: Customize image and video understanding models, detail levels, and size limits.
  • Audio Transcription: Set up preferred models, base URLs, and API keys for converting voice messages to text.
  • Use Case: You can configure DevClaw to use a budget-friendly vision model for quick image descriptions and a fast transcription service for voice notes, optimizing performance and cost.

Quick Start

Configure DevClaw to use the OpenAI GPT-4o model for vision and the Groq Whisper-large-v3-turbo model for transcription.

Frequently Asked Questions about media-config

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I configure audio transcription for voice messages across messaging channels?

Audio transcription for messaging channels is configured by setting preferred models, base URLs, and API keys for services like Groq Whisper-large-v3-turbo or OpenAI to convert voice messages into text efficiently.

Can I customize vision models for image and video processing in DevClaw?

Yes, you can customize vision models for image and video understanding in DevClaw by defining specific models such as OpenAI GPT-4o, detail levels, and size limits to process multimodal inputs accurately.

What is the best way to optimize media processing costs for multimodal inputs?

The best way to optimize media processing costs is by configuring DevClaw to use a budget-friendly vision model for quick image descriptions and a fast transcription service for voice notes, balancing performance and cost.

Does media configuration support both OpenAI and Groq for multimodal inputs?

Media configuration supports both OpenAI and Groq integration, allowing you to define specific API endpoints and models like OpenAI GPT-4o for vision and Groq Whisper for audio transcription within your messaging channels.

How do I set API endpoints for audio transcription and video understanding?

API endpoints for audio transcription and video understanding are set by defining specific models and base URLs within the configuration, enabling efficient handling of multimodal inputs like images, videos, and audio messages.

What limitations exist when configuring size limits for vision models?

Limitations for vision model configuration involve the specific size limits and detail levels you define for image and video understanding, which dictate how efficiently DevClaw processes visual content from communication channels.