heartmula

Generate 48kHz stereo music tracks from lyrics and descriptive tags.

Updated May 4, 2026
One-click install
npx skills add https://github.com/InverterNetwork/hermes-agent --skill heartmula-inverternetwork
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: heartmula
Source: https://github.com/InverterNetwork/hermes-agent/tree/main/optional-skills/creative/heartmula
Command: npx skills add https://github.com/InverterNetwork/hermes-agent --skill heartmula-inverternetwork

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires datasets, transformers, torchtune, torch, and includes assets (resource) components.

What problem does it solve?

This Skill solves the challenge of generating high-quality, original music from text-based lyrics and descriptive tags without requiring professional audio engineering skills or expensive proprietary software.

Core Features & Use Cases

  • Lyrics-to-Song Generation: Transforms structured lyrics and mood tags into full-length, 48kHz stereo audio tracks.
  • Open-Source Foundation: Provides a local, privacy-focused alternative to commercial music generation platforms.
  • Use Case: A content creator can use this to generate custom background music for a video project by providing specific genre tags and lyrical content directly to the agent.

Quick Start

Use the heartmula skill to generate a song from the lyrics in assets/lyrics.txt and the tags in assets/tags.txt and save the output to output.mp3.

Frequently Asked Questions about heartmula

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from lyrics and descriptive tags locally?

You can generate music from lyrics by inputting structured lyrical text and genre tags into a specialized music language model, which outputs high-fidelity 48kHz stereo audio tracks.

What is text-to-music generation using a music language model and codec?

Text-to-music generation with a music language model and codec conditions neural network inference on lyrical text and descriptive tags to synthesize high-fidelity original audio tracks locally.

Do I need a CUDA-enabled GPU to run local audio synthesis for music generation?

Yes, a CUDA-enabled GPU is required for efficient inference and processing the specific model checkpoints needed to run local audio synthesis for music generation.

Can I generate custom background music for video projects without proprietary software?

Yes, you can generate custom background music for video projects without proprietary software by using an open-source, privacy-focused local foundation that processes your specific genre tags and lyrical content.

What are the limitations of generating high-fidelity music from text tags?

Limitations of generating high-fidelity music from text tags include the strict hardware requirement of a CUDA-enabled GPU for efficient inference and the need to acquire specific model checkpoints for the architecture.