Claude Code Skills·Claude Skills·The open SKILL.md registry for Claude
ClaudSkillsContent › Audio Podcast › Page 8

Audio Podcast (Page 8 of 9)

483 Claude Code skills in the Audio Podcast sub-category of Content.

483 skills · updated 2026-07-23 · showing 421–480 of 483 by quality score

For the full experience including quality scoring and one-click install features for each skill — upgrade to Pro.

Streams audio from PulseAudio or ALSA devices into whisper.cpp for real-time speech-to-text with word-level timestamps.
Enhances OpenAI Whisper transcription output with speaker diarization using pyannote.audio pipeline and speechbrain embeddings.
Implement or adjust background transcription jobs for whisper-lolo. Use when wiring Inngest events, handling long-running jobs, chunking before transcription, persisting…
Use when the user wants to transcribe, caption, subtitle, batch process, or convert speech to text from local audio/video files using faster-whisper.
Transcribe audio and video files to text using OpenAI Whisper. Use when: converting podcasts to blog posts; creating video subtitles; extracting quotes from interviews;…
WhisperX extends OpenAI Whisper with batched inference for 70x realtime transcription, phoneme-based word-level timestamp alignment via wav2vec2, voice activity detection, and…
Expert in Windows 3.1 era sound vocabulary for modern web/mobile apps. Creates satisfying retro UI sounds using CC-licensed 8-bit audio, Web Audio API, and haptic coordination.
Thin orchestrator for the end-to-end video localization pipeline. Routes to the four focused sub-skills — /wjs-transcribing-audio, /wjs-translating-subtitles, /wjs-dubbing-video,…
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language.
Trích xuất transcript narration có timestamp cấp câu và cấp từ; script hiện dùng OpenAI Whisper hoặc API transcription tương thích.
Use when user shares a xiaoyuzhoufm.com link, mentions downloading a podcast episode, or wants to transcribe podcast audio to text
Turn a YouTube link into a polished single-file bilingual (Chinese + original) transcript reading page.
Convert YouTube talks/seminars into a markdown transcript with the speaker's slides interleaved at the correct timestamps.
Extract subtitles/transcripts from YouTube videos. Triggers: "youtube transcript", "extract subtitles", "video captions", "视频字幕", "字幕提取", "YouTube转文字", "提取字幕".
Use when a user needs YouTube subtitles or transcript text from a URL using yt-dlp only, with no video download and no browser automation fallback.
Transcribe YouTube videos and playlists by extracting auto-generated captions directly from the browser — no API key, no external service, completely free.
End-to-end YouTube Shorts production pipeline: script → screen record → narration + captions → assembled 9:16 MP4.
Use when user asks about YouTube video content, wants to know what a video says, needs information from a YouTube URL, or when video transcription would answer their question
YouTube-Transkripte (Untertitel) und Video-Metadaten abrufen und als Markdown, JSON oder Plaintext ausgeben. Bevorzugt manuelle Untertitel, Fallback auf automatisch generierte.
The new frontier of audio: AI-generated music with Suno and Udio, AI sound effects with ElevenLabs, AI voice cloning, and AI audio enhancement.
AI voice creation skill supporting speech recognition (ASR) and text-to-speech (TTS). Uses qwen3-asr-flash-filetrans, qwen-tts and other models.
This skill should be used when the user asks about "audio production", "ElevenLabs", "voice isolator", "audio post-production", "AI narration", "text to speech production",…
Use this skill when adding audio or sound to a Phaser 4 game. Covers loading audio, playing sounds, music, volume, spatial audio, Web Audio API, and SoundManager.
Record audio from the Pi's USB microphone or play WAV files through the Pi's speaker. The audio subsystem runs on the host (audio bridge on port 18796); this skill triggers…
Use this skill whenever writing any audio cue, haptic signal, Web Audio API initialisation, or interaction feedback code in the Myelin simulator.
Transcribe and analyze audio content using Google Gemini. Supports local audio files (mp3, wav, m4a, ogg, flac) and YouTube links up to 9.5 hours long.
Build the Chuqlab Intelligence Engine (CIE) - a multi-tenant correctional intelligence platform. Triggers on requests to build CIE, implement CIE phases, work on CrimeMiner…
Tools, patterns, and utilities for creating music with code. Output as a .mp3 file with realistic instrument sounds.
Transcribe recorded audio files to text via Doubao Seed-ASR 2.0 (豆包录音文件识别模型2.0) from ByteDance/Volcengine. Best-in-class Chinese speech recognition with speaker diarization.
ChIP-Atlas epigenome data query using gcell. Use this skill when users ask about: - Finding ChIP-seq, ATAC-seq, or DNase-seq experiments - Searching for transcription factor…
Expert developer skill for implementing real-time voice and video interactions using the Google Gemini Live API.
Discover and install third-party skills from external registries when the user needs a capability that no currently-active skill covers.
This skill should be used when the user asks to "implement voice input", "add speech recognition", "use SFSpeechRecognizer", "configure microphone permissions", "音声入力を実装したい",…
Transcribe a specified local video or audio file into cleaned final `.txt`, `.pdf`, or `.docx` transcripts using speech recognition with Apple Silicon GPU acceleration and…
The `marmot` CLI bundles AI generation (text, image, video, speech, transcription), web retrieval (search, scrape, answer, map, crawl, research, findall), and data lookup (enrich,…
Moonshine Voice is a fast on-device speech recognition library for interactive voice applications. This skill helps agents install the Python package, load supported language…
MUST read this skill BEFORE entering generate mode for music tasks. Covers prompt crafting framework, structure syntax, and multi-clip strategy.
Orchestrates end-to-end OCR text extraction from PDF books and documents. Use this skill ALWAYS when the user requests text extraction, OCR, transcription, book processing, or…
Use this skill when building integrations with the OpenWhispr REST API, calling OpenWhispr endpoints, managing notes/folders/transcriptions programmatically, or connecting to the…
Use this skill whenever the user wants to operate on OpenWhispr notes, folders, transcriptions, or audio from a terminal or shell.
Implements Fireflies AI meeting transcription API access for transcripts, users, summaries, and audio upload using generated clients and MCP tools.
Handle audio messages from Telegram and send voice responses. Trigger when the user sends an audio file (e.g., .ogg), requests a voice response, or when the conversation context…
Transcribes audio and video files to text using pluggable ASR backends. Default backend is local whisper CLI (openai-whisper).
This skill should be used when the user asks to "clean up a transcript", "fix speech artifacts", "edit interview quotes", "polish transcription", "clean up quotes from a…
Corrects speech-to-text transcription errors using dictionary rules and AI-powered analysis. Builds personalized correction databases that learn from each fix.
Use this skill when the user wants to transcribe a Google Meet recording, generate a commercial proposal from a meeting transcript, or push a proposal to Notion.
Interactive text-to-speech audio generation using the gemini-media MCP (Google Gemini TTS). Use this skill whenever the user asks to convert text to speech, generate spoken audio,…
This skill activates when users ask about text-to-speech setup, configuration, troubleshooting, voice selection, or TTS functionality.
Whishper is an open source self-hosted web app for speech-to-text, translation, and subtitle workflows built around Whisper models.
Extract, transcribe, and summarize audio or video files using OpenAI Whisper. Use this skill whenever the user wants to transcribe audio or video, extract what was said in a…
Use this skill when the user is building with `xsai` or any `@xsai/*` package, or is evaluating xsAI for a small OpenAI-compatible workflow with text generation, streaming, tool…
This skill should be used when the user asks to "make a YouTube intro", "create a channel intro", "build a logo sting/bumper", "design an outro / end screen / end card", "add…
Transcribe and organize a YouTube video into a structured Markdown document using the local yt2doc CLI tool.
Control Yulu (语录), the local-first macOS meeting recorder. Use this skill when the user asks to start or stop recording a meeting, check recording status, look up past…
Speech-to-text (audio transcription) via flow_router /v1/audio/transcriptions. OpenAI-compat shape (multipart form), routes to active STT providers.
Add text-to-speech narration to Claude Code on macOS
Transcribe audio verbatim with speaker attribution
When extracting hardcoded data to JSON, field values silently drift:
Agent D2 - Data Collection Specialist - Interviews, Focus Groups & Observation. Covers protocol development, question design, probing strategies, transcription conventions, and…
Transcribe audio via OpenAI Audio Transcriptions API (Whisper).
All Content skills →
More in ContentStorytelling (831) · Video (524) · Translation (519) · Editorial (310) · Writing (271) · Image Design (176)