video-message

Generate lip-syncing VRM avatar video messages from text or audio as MP4 Telegram video notes.

3|Updated Feb 1, 2026
One-click install
npx skills add https://github.com/thewulf7/openclaw-avatarcam --skill video-message
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: video-message
Source: https://github.com/thewulf7/openclaw-avatarcam/tree/main/skill
Command: npx skills add https://github.com/thewulf7/openclaw-avatarcam --skill video-message

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Generate and send lip-syncing VRM avatar video messages from text or audio, enabling expressive replies and video-based TTS delivery.

Core Features & Use Cases

  • Lip-sync avatar video generation from text or audio for messaging
  • Outputs as Telegram video notes (circular format)
  • Configurable avatar and background with optional headless rendering

Quick Start

Run avatarcam with an audio file or TTS input to generate a circular Telegram video note.

Frequently Asked Questions about video-message

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate a lip-sync avatar video message from text?

To generate a lip-sync avatar video message, you provide text or audio input to the Skill, which then renders a VRM avatar and outputs a 384x384 MP4 video note. This process enables expressive text-based video replies for messaging.

Can I create Telegram video notes using a VRM avatar?

Yes, you can create Telegram video notes using a VRM avatar. The Skill specifically outputs 384x384 MP4 files in a circular format, which is the exact specification required for Telegram video notes.

Do I need ffmpeg and avatarcam to render headless video messages?

Yes, you need both ffmpeg and avatarcam to render headless video messages. These dependencies are required by the Skill to process the audio or text input and generate the final VRM avatar lip-sync video output.

How do I use a VRM avatar for TTS video delivery in OpenClaw?

You can use a VRM avatar for TTS video delivery in OpenClaw by routing your text-to-speech audio through this Skill. It generates a lip-syncing avatar video that can be used as a video-based response within OpenClaw workflows.

Can I configure the background when generating a lip-sync avatar video?

Yes, you can configure the background when generating a lip-sync avatar video. The Skill supports configurable avatars and backgrounds, alongside optional headless rendering, allowing customized visual environments for your video messages.

What are the limitations of rendering VRM avatar video notes?

A limitation of rendering VRM avatar video notes is the fixed output resolution of 384x384 MP4. Additionally, the Skill requires ffmpeg and avatarcam to be installed, meaning it cannot function without these specific external dependencies.