video-learning-assistant

Automate online video analysis by extracting transcripts, summaries, and visual insights from Douyin and Bilibili using yt-dlp and faster-whisper.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/syzl8512/openclaw-experience --skill video-learning-assistant
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: video-learning-assistant
Source: https://github.com/syzl8512/openclaw-experience/tree/main/03_%E6%A0%B8%E5%BF%83Skill/video-learning-skill
Command: npx skills add https://github.com/syzl8512/openclaw-experience --skill video-learning-assistant

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires yt-dlp, faster-whisper, browser-cookie3, mcporter, ffmpeg, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill automates the tedious process of analyzing online videos, extracting valuable information like spoken content and visual details, and transforming it into structured learning materials.

Core Features & Use Cases

  • Automated Video Analysis: Process videos from platforms like Douyin and Bilibili.
  • Transcription & Summarization: Generate accurate transcripts and concise summaries of video content.
  • Visual Analysis: Extract key frames and analyze visual design elements for deeper understanding.
  • Note Generation: Create structured learning notes, visual analysis reports, and even presentation scripts.
  • Use Case: Quickly create study notes from a lecture video on Bilibili, including a summary, key takeaways, and analysis of the presenter's visual aids.

Quick Start

Analyze this video and generate a learning summary: [video link].

Frequently Asked Questions about video-learning-assistant

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract spoken content and generate learning notes from a Bilibili video?

To extract spoken content and generate learning notes from a Bilibili video, you provide the video link to initiate automated downloading, transcription, and summarization. The system then processes the audio to produce structured study materials.

Can I analyze visual design elements and extract key frames from Douyin videos?

Yes, you can analyze visual design elements from Douyin videos. The skill extracts key frames using ffmpeg and applies visual analysis to understand design details, outputting a comprehensive visual analysis report.

Does video transcription work without manually downloading the video first?

Video transcription works without manual downloading by using yt-dlp to automate the retrieval process. Once the video link is provided, the system handles downloading automatically before passing it to faster-whisper for transcription.

What is the best way to convert a video lecture into a structured presentation script?

The best way to convert a video lecture into a structured presentation script is through automated video analysis. The skill transcribes spoken content and analyzes visual aids to generate structured learning notes and scripts.

Do I need to install ffmpeg and faster-whisper to transcribe online videos?

You need ffmpeg and faster-whisper to transcribe online videos because they handle core audio extraction and speech-to-text processing. The skill relies on these dependencies along with yt-dlp to execute the automated analysis workflow.

Are there limitations when using browser-cookie3 for video downloading on Douyin?

When using browser-cookie3 for video downloading, limitations may arise if your browser session expires or lacks proper authentication for Douyin or Bilibili. Valid cookies are required to bypass access restrictions and ensure successful video retrieval.