Automatically helps debug Web Audio API issues, audio playback problems, pitch preservation, and caching issues in the VSSK-shadecn music practice app
Làm sạch bản ghi giọng nói WAV/MP3 theo workflow 2-phase semantic — AI viết lại nội dung không lặp vào TOML, sau đó căn keep flag từng token để render audio cuối.
Comprehensive onset diagnosis - determine if onset is noise or real kalimba note
Extract audio entities from narrative text. Use when analyzing music themes, sound effects, ambient audio, motifs, dynamic scoring, and silence as narrative device.
Expert in digital signal processing for audio applications. Validates biquad filter implementations, frequency response calculations, and audio algorithms.
Master the essential audio post-production techniques—normalization, compression, EQ, and noise reduction—using the correct processing order to achieve professional-quality audio.
FFmpeg audio processing, batch editing, normalization, mixing, and automated audio production workflows.
Create standard SuperCollider audio effects for Bice-Box (delays, reverbs, filters, distortions). Provides templates, ControlSpecs, common patterns, and MCP workflow for safely…
Audio engineering — mastering, mixing, EQ, compression, loudness standards, synthesis, podcast production, music theory, spectrum analysis.
Audio production concepts, DSP fundamentals, mixing/mastering techniques, and DAW workflows. Bridges modular synthesis philosophy with practical audio engineering.
Remove background noise, enhance clarity, and improve audio quality. Perfect for podcasts, videos, and calls. No API keys needed. $2 FREE credits to start.
从视频文件中提取音频。Use when user wants to 提取音频, 抽取音频, 视频转音频, 导出音频, extract audio, video to audio, get audio from video, 把视频的声音提取出来.
ffmpeg patterns for extracting audio from video files and transcoding between formats
Pulls a video soundtrack into mono 16 kHz 16-bit PCM WAV through the bundled FFmpeg wrapper. Trigger on extract-audio, ASR/VAD prep, or normalizing mixed containers to analysis…
You are the audio architecture expert ensuring Leavn's complex audio pipeline stays coherent.
You are the audio fingerprinting and pattern detection specialist for Modcaster's content analysis.
Identifies audio content using Chromaprint/AcoustID fingerprinting, Shazam API recognition, and ACRCloud monitoring.
통합 오디오 생성 스킬. ElevenLabs MCP 기반 TTS(32개국어), 보이스 클로닝(1분 샘플), 다국어 더빙(립싱크), 효과음 생성을 지원. "목소리 생성", "TTS", "음성 합성", "보이스 클로닝", "더빙", "나레이션", "효과음", "AI 음성" 요청 시 사용.
Use whenever the user asks to install, configure, uninstall, snooze, mute, test, troubleshoot, or change settings for the claude-code-audio-hooks audio notification system.
Test Bob The Skull with virtual audio injection instead of speaking. Use when testing wake word detection, STT accuracy, full conversation pipeline, or automated testing.
Audio generation skill — jingles, beds, voiceover, and sound effects. Routes music requests to Suno V5 / Udio / Lyria, speech to MiniMax TTS / FishAudio / ElevenLabs V3, and SFX…
Gemini Live API, Grok Voice Agent, GPT-4o-Transcribe, AssemblyAI patterns for real-time voice, speech-to-text, and TTS.
Create memorable sonic logos using design principles from Intel, Netflix, and McDonald's—crafting 2-5 second audio signatures that achieve instant brand recognition.
You are the on-device audio ML specialist for Modcaster's AI-driven audio processing. — from majiayu000/claude-skill-registry
You are the on-device audio ML specialist for Modcaster's AI-driven audio processing. — from majiayu000/claude-skill-registry
Use when writing songs, generating music or sound with AI, preparing Suno/HeartMuLa prompts, or analyzing audio features and spectrograms.
Use when asked to normalize audio volume, match loudness, or apply peak/RMS normalization to audio files.
Audio playback using Tone.js including players, transport, scheduling, and loading audio. Use when implementing background music, sound effects, audio synchronization, or timed…
Audio ingestion, analysis, transformation, and generation (Transcribe, TTS, VAD, Features). — from dvcrn/openclaw-skills-marketplace
Converts and processes audio files using ffmpeg. Supports format conversion, sample rate changes, mono/stereo conversion, and segment splitting.
Professional audio production for music, podcasts, and sound design. Use when working with audio recording, mixing, mastering, or sound design for any medium.
Analyze audio recording quality - echo detection, loudness, speech intelligibility, SNR, spectral analysis.
Analyze the WaveCap-SDR audio stream to assess tuning quality, detect silence, noise, proper audio, or distortion.
Binding audio analysis data to visual parameters including smoothing, beat detection responses, and frequency-to-visual mappings.
Generate audio replies using TTS. Trigger with "read it to me [URL]" to fetch and read content aloud, or "talk to me [topic]" to generate a spoken response.
Generate audio replies using TTS. Trigger with "read it to me [public URL]" to fetch and read content aloud, or "talk to me [topic]" to generate a spoken response.
Router for audio domain including playback, analysis, and audio-reactive visuals. Use when implementing any audio functionality including music, sound effects, visualizers, or…
Separates audio tracks into individual stems (vocals, drums, bass, other) using Meta's Demucs neural network model via the demucs Python package.
팟캐스트 대본작가(scriptwriter)와 쇼노트편집자(shownote-editor)가 사용하는 오디오 스토리텔링 전문 스킬. 귀로만 듣는 매체에서 청취자의 몰입을 극대화하는 서사 구조, 페이싱, 사운드 연출 방법론을 제공한다.
音频流上传专业版 —— 面向企业团队与专业创作者的高级音频上传工具。核心能力: - 批量音频上传,支持队列管理与断点续传 - 完全自定义编码配置:比特率、采样率、声道、编解码器 - 多质量预设输出(标准/良好/最高/无损),满足不同播放场景 - HLS与DASH双流媒体格式支持,适配多终端播放 - 丰富的元数据管理:标签、描述、自定义键值对 -…
Implements audio systems including sound management, music systems, positional audio, and audio effects.
Game audio systems, music, spatial audio, sound effects, and voice implementation. Build immersive audio experiences with professional middleware integration.
Turn a dialogue, a story with dialogue, or a one-line idea into finished multi-character audio - a radio drama, a two-host podcast, or clean lip-sync voice clips - 100% locally…
Turn creator audio into clean text captions for ecommerce content and reuse. Use when teams need fast transcript-to-caption workflows.
End-to-end audio production workflow with stems, effects, archiving, and verification
Step-by-step audio production with per-stem verification, timing alignment, and incremental quality gates
Incremental audio production with duration alignment handling, per-stem verification, and adaptive extension strategies
Incremental audio production with duration mismatch handling, adaptive stem extension, and pre-mix alignment verification
Audio production with diagnostic analysis, timecode parsing from documents, and verified export workflow
使用 Whisper 将音频/视频转换为文字,支持词级别时间戳。Use when user wants to 语音转文字, 音频转文字, 视频转文字, 字幕生成, transcribe audio, speech to text, generate subtitles, 识别语音.
Transform audio recordings into professional Markdown documentation with intelligent summaries using LLM integration
Build audio transcription pipelines with Whisper, Deepgram, and AssemblyAI including speaker diarization and real-time streaming.
Cut, trim, and edit audio segments with fade effects, speed control, concatenation, and basic audio manipulations.
Quick upload audio to AIOZ Stream API。Create audio objects with default or custom encoding confi。Use when 需要设计创作、UI设计、海报制作、品牌视觉时使用。不适用于3D建模和动画制作。适用于独立开发者、企业团队和自动化工作流场景。
基于 AIOZ Stream API 的音频上传技能免费版,通过三步流程 (Create → Upload Part → Complete) 将本地音频文件上传至 AIOZ 流媒体平台。支持默认快速上传方式,使用默认编码配置, 上传完成后返回 HLS 流媒体播放链接。使用 API Key…
Audio and video processing with FFmpeg, WebRTC, and streaming. Covers transcoding, format conversion, real-time communication, and media pipelines.
Transcribes audio or video to text with word-level timestamps via faster-whisper (default large-v3-turbo) or local openai-whisper (word_timestamps=True).
Transcription beyond the GuideAnts wrapper contract: transcribe workspace files by path (no upload, no 50 MB gateway cap), pass language hints, and sideload other qwen3-family ASR…
Advanced synthesis controls on the loaded audio.cpp TTS model: deterministic output via seed, forcing the spoken language, voice-design from a text description (instructions), and…
PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen).