AudioEditor

Transcribe audio with word-level timestamps and remove filler words.

1|Updated Mar 11, 2026
One-click install
npx skills add https://github.com/BishopCodes/OpenPAI --skill audioeditor
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: AudioEditor
Source: https://github.com/BishopCodes/OpenPAI/tree/main/skills/Utilities/AudioEditor
Command: npx skills add https://github.com/BishopCodes/OpenPAI --skill audioeditor

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the complex and time-consuming process of editing audio and video files, making professional-quality sound accessible to everyone.

Core Features & Use Cases

  • Intelligent Transcription: Generates word-level transcripts for precise editing.
  • Automated Editing: Detects and removes filler words, stutters, false starts, and dead air with intelligent cut detection.
  • Audio Polish: Optionally applies advanced post-processing for a professional finish.
  • Use Case: Clean up a podcast recording by removing all "ums," "ahs," and long pauses, then normalize the audio levels for consistent playback.

Quick Start

Use the AudioEditor skill to clean the audio file named 'meeting_recording.mp3'.

Frequently Asked Questions about AudioEditor

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I remove filler words and dead air from a podcast recording?

To remove filler words and dead air from a podcast recording, the AudioEditor skill uses intelligent cut detection to identify stutters, false starts, and pauses, then automatically edits the audio with crossfades for seamless playback.

Can I get word-level timestamps for audio transcription?

Yes, you can get word-level timestamps for audio transcription. The AudioEditor skill generates precise transcripts with word-level timing to enable accurate, granular editing of your audio and video files.

Does automated audio cleanup support mouth sound removal and loudness normalization?

Automated audio cleanup supports mouth sound removal and loudness normalization through an optional cloud-based polishing feature using the Cleanvoice API to achieve a professional finish.

What is the best way to automate audio editing for video files?

The best way to automate audio editing for video files is using a skill like AudioEditor that transcribes the content, detects unwanted sounds or pauses, and applies automated cuts with crossfades to clean up the track.

Do I need the Cleanvoice API to edit out long pauses from my audio?

No, you do not need the Cleanvoice API to edit out long pauses. The API is optional for advanced mouth sound removal and loudness normalization, while core dead air detection and removal work without it.