Extract subtitles/transcripts from YouTube videos. Triggers: "youtube transcript", "extract subtitles", "video captions", "视频字幕", "字幕提取", "YouTube转文字", "提取字幕".
Use when a user needs YouTube subtitles or transcript text from a URL using yt-dlp only, with no video download and no browser automation fallback.
Transcribe YouTube videos and playlists by extracting auto-generated captions directly from the browser — no API key, no external service, completely free.
Transform YouTube videos into podcast-style voice summaries using ElevenLabs TTS
End-to-end YouTube Shorts production pipeline: script → screen record → narration + captions → assembled 9:16 MP4.
Use when user asks about YouTube video content, wants to know what a video says, needs information from a YouTube URL, or when video transcription would answer their ques — from…
Use when user asks about YouTube video content, wants to know what a video says, needs information from a YouTube URL, or when video transcription would answer their ques — from…
YouTube-Transkripte (Untertitel) und Video-Metadaten abrufen und als Markdown, JSON oder Plaintext ausgeben. Bevorzugt manuelle Untertitel, Fallback auf automatisch generierte.
Music songwriting guide for ACE-Step. Provides professional knowledge on writing captions, lyrics, choosing BPM/key/duration, and structuring songs.
The new frontier of audio: AI-generated music with Suno and Udio, AI sound effects with ElevenLabs, AI voice cloning, and AI audio enhancement.
AI大模型专家|参考音频生视频。帮助视频剪辑师、后期团队、广告制作人与需要复用现有素材的创作者直接完成“参考音频生视频”:提供音乐或声音即可围绕节奏、情绪和声音线索生成画面;通过 AI Hive 使用时,生成前自动上传所需素材,提交后自动保存任务、查询进度并下载成片。适用于 AI 广告、TVC、电商视频、产品展示、带货、种草、短剧、漫剧和社媒内容。 Use…
AI大模型专家|Happy Horse 参考音频生视频。帮助视频剪辑师、后期团队、广告制作人与需要复用现有素材的创作者直接完成“Happy Horse 参考音频生视频”:提供音乐或声音即可围绕节奏、情绪和声音线索生成画面;通过 AI Hive 使用时,生成前自动上传所需素材,提交后自动保存任务、查询进度并下载成片。适用于 AI…
AI大模型专家|MiniMax H3 参考音频生视频。帮助视频剪辑师、后期团队、广告制作人与需要复用现有素材的创作者直接完成“MiniMax H3 参考音频生视频”:提供音乐或声音即可围绕节奏、情绪和声音线索生成画面;通过 AI Hive 使用时,生成前自动上传所需素材,提交后自动保存任务、查询进度并下载成片。适用于 AI…
AI大模型专家|Seedance 2.0 参考音频生视频。帮助视频剪辑师、后期团队、广告制作人与需要复用现有素材的创作者直接完成“Seedance 2.0 参考音频生视频”:提供音乐或声音即可围绕节奏、情绪和声音线索生成画面;通过 AI Hive 使用时,生成前自动上传所需素材,提交后自动保存任务、查询进度并下载成片。适用于 AI…
AI大模型专家|Seedance 2.5 参考音频生视频。帮助视频剪辑师、后期团队、广告制作人与需要复用现有素材的创作者直接完成“Seedance 2.5 参考音频生视频”:提供音乐或声音即可围绕节奏、情绪和声音线索生成画面;通过 AI Hive 使用时,生成前自动上传所需素材,提交后自动保存任务、查询进度并下载成片。适用于 AI…
AI大模型专家|Seedance 参考音频生视频。帮助视频剪辑师、后期团队、广告制作人与需要复用现有素材的创作者直接完成“Seedance 参考音频生视频”:提供音乐或声音即可围绕节奏、情绪和声音线索生成画面;通过 AI Hive 使用时,生成前自动上传所需素材,提交后自动保存任务、查询进度并下载成片。适用于 AI…
AI voice creation skill supporting speech recognition (ASR) and text-to-speech (TTS). Uses qwen3-asr-flash-filetrans, qwen-tts and other models.
This skill should be used when the user asks about "audio production", "ElevenLabs", "voice isolator", "audio post-production", "AI narration", "text to speech production",…
This skill performs transcription factor (TF) footprint analysis using TOBIAS on ATAC-seq data. It corrects Tn5 sequence bias, quantifies TF occupancy at motif sites, generates…
Use this skill when adding audio or sound to a Phaser 4 game. Covers loading audio, playing sounds, music, volume, spatial audio, Web Audio API, and SoundManager.
Record audio from the Pi's USB microphone or play WAV files through the Pi's speaker. The audio subsystem runs on the host (audio bridge on port 18796); this skill triggers…
Identify logical topic shifts in audio and produce chaptered timestamps. Use this skill whenever the user mentions chapter generation, topic segmentation, timestamp, audio…
Use this skill whenever writing any audio cue, haptic signal, Web Audio API initialisation, or interaction feedback code in the Myelin simulator.
Identify the primary spoken language in an audio file before transcription runs. Use this skill whenever the user mentions language detection, language identification, spoken…
Use this skill when creators, video editors, advertisers, e-commerce teams, social-commerce teams, and short-form story producers need to generate visuals around music, r — from…
Transcribe and analyze audio content using Google Gemini. Supports local audio files (mp3, wav, m4a, ogg, flac) and YouTube links up to 9.5 hours long.
This skill helps an LLM generate correct audio code with @ax-llm/ax. Use when the user asks about ai.transcribe(), ai.speak(), signature audio inputs or outputs, agent audio…
Build the Chuqlab Intelligence Engine (CIE) - a multi-tenant correctional intelligence platform. Triggers on requests to build CIE, implement CIE phases, work on CrimeMiner…
Tools, patterns, and utilities for creating music with code. Output as a .mp3 file with realistic instrument sounds.
Transcribe recorded audio files to text via Doubao Seed-ASR 2.0 (豆包录音文件识别模型2.0) from ByteDance/Volcengine. Best-in-class Chinese speech recognition with speaker diarization.
Use this skill when the question involves cursor speed, pointer acceleration curves, or interaction with input devices that vary in precision (mouse, trackpad, stylus,…
Use this skill when designing for the edges and corners of the screen — where the cursor cannot overshoot, making them effectively infinite-sized targets along one axis.
Use this skill when designing for touch — phones, tablets, kiosks, in-car displays, smart-TV remotes, anything where the input is a finger or thumb rather than a precise pointer.
ChIP-Atlas epigenome data query using gcell. Use this skill when users ask about: - Finding ChIP-seq, ATAC-seq, or DNase-seq experiments - Searching for transcription factor…
Expert developer skill for implementing real-time voice and video interactions using the Google Gemini Live API.
ALWAYS read this skill before generating or animating any video, or calling video_generate — text-to-video, image-to-video, a start→end transition, or a reference / motion /…
Discover and install third-party skills from external registries when the user needs a capability that no currently-active skill covers.
Use this skill when creators, video editors, advertisers, e-commerce teams, social-commerce teams, and short-form story producers need to generate visuals around music, r — from…
Use this skill when working with MOSS-TTS-Nano. Triggers when user mentions MOSS-TTS-Nano or imports from it.
This skill should be used when the user asks to "implement voice input", "add speech recognition", "use SFSpeechRecognizer", "configure microphone permissions", "音声入力を実装したい",…
Transcribe a specified local video or audio file into cleaned final `.txt`, `.pdf`, or `.docx` transcripts using speech recognition with Apple Silicon GPU acceleration and…
The `marmot` CLI bundles AI generation (text, image, video, speech, transcription), web retrieval (search, scrape, answer, map, crawl, research, findall), and data lookup (enrich,…
Use this skill when creators, video editors, advertisers, e-commerce teams, social-commerce teams, and short-form story producers need to generate visuals around music, r — from…
Moonshine Voice is a fast on-device speech recognition library for interactive voice applications. This skill helps agents install the Python package, load supported language…
Run multi-stage transcription, translation and localized reformatting across languages. Use this skill whenever the user mentions multilingual transcription, audio translation,…
MUST read this skill BEFORE entering generate mode for music tasks. Covers prompt crafting framework, structure syntax, and multi-clip strategy.
Orchestrates end-to-end OCR text extraction from PDF books and documents. Use this skill ALWAYS when the user requests text extraction, OCR, transcription, book processing, or…
Use this skill when building integrations with the OpenWhispr REST API, calling OpenWhispr endpoints, managing notes/folders/transcriptions programmatically, or connecting to the…
Use this skill whenever the user wants to operate on OpenWhispr notes, folders, transcriptions, or audio from a terminal or shell.
Implements Fireflies AI meeting transcription API access for transcripts, users, summaries, and audio upload using generated clients and MCP tools.
Use this skill when creators, video editors, advertisers, e-commerce teams, social-commerce teams, and short-form story producers need to generate visuals around music, r — from…
Use this skill when creators, video editors, advertisers, e-commerce teams, social-commerce teams, and short-form story producers need to generate visuals around music, r — from…
Use this skill when creators, video editors, advertisers, e-commerce teams, social-commerce teams, and short-form story producers need to generate visuals around music, r — from…
Split long-form audio into chunks for parallelized or high-throughput inference. Use this skill whenever the user mentions audio chunking, segment processing, parallel inference,…
Handle audio messages from Telegram and send voice responses. Trigger when the user sends an audio file (e.g., .ogg), requests a voice response, or when the conversation context…
Stream call audio in real-time, fork media to external destinations, and transcribe speech live. Use for real-time analytics and AI integrations.
Stream call audio in real-time, fork media to external destinations, and transcribe speech live. Use for real-time analytics and AI integrations.
Stream call audio in real-time, fork media to external destinations, and transcribe speech live. Use for real-time analytics and AI integrations.
Stream call audio in real-time, fork media to external destinations, and transcribe speech live. Use for real-time analytics and AI integrations.
Stream call audio in real-time, fork media to external destinations, and transcribe speech live. Use for real-time analytics and AI integrations.