Monitors live audio streams from RTMP, HLS, or Icecast sources using FFmpeg stream capture and real-time chunked transcription via Deepgram's streaming API or Whisper.cpp.
Build self-hosted speech-to-text APIs using Hugging Face models (Whisper, Wav2Vec2) and create LiveKit voice agent plugins.
基于 Whisper CLI 的本地语音转文字工具(免费版)。核心能力: - 本地音频转文字(transcription),无需 API Key - 支持 mp3 / m4a / wav / flac 等常见格式 - 多种输出格式:txt / srt / vtt / json - 内置翻译模式(音频转英文) - 模型自动下载与缓存
本地录音转文字工具。当用户发送已有录音、音频或视频文件,并希望把语音转成 Markdown 文稿和 SRT 字幕时使用。Apple Silicon 优先用 MLX/Apple GPU 和 whisper-large-v3-turbo-q4,本地转写,不生成 txt/json/vtt,不用于现场临时录音,也不默认调用云端语音识别服务。
Gere texto para fala local em português brasileiro com Piper ou Kokoro. Use quando o usuário quiser TTS offline, leitura em pt-BR, geração rápida de áudio, narração natural, ou…
Local speech-to-text using OpenAI Whisper. Runs fully offline after model download. High quality transcription with multiple model sizes.
Convert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (wi — from…
Use when you have a processed single-cell expression matrix (AnnData object) with pre-computed cluster assignments (e.g., leiden or louvain clusters in adata.obs) and want to…
Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EP — from…
Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EP — from…
OpenClaw agent skill for converting documents to Markdown. Documentation and utilities for Microsoft's MarkItDown library.
Use when the user is finishing a track and wants to check it's ready to send to a mastering engineer or for self-mastering.
Content analysis for video and audio — YouTube, TikTok, podcasts, audio files. Transcription-first pipeline (captions API, user transcript, or Whisper opt-in).
AUTO-INVOKE when user mentions media curation, discography, archive, music collection, video/audio collection, source acquisition, metadata tagging, transcript, Plex, Jellyfin,…
Use when a task needs the judgment of a Medical Transcriptionist — editing a speech-recognition draft for a dictated clinical note, disambiguating a sound-alike drug name or…
Transcribe meeting audio with speaker diarization, generate structured summaries with action items, decisions, and follow-ups, and support multiple audio formats and languages.
Use when you have measured intracellular metabolite concentrations (e.g., via LC–MS/MS) across multiple cell lines or samples and want to predict which metabolic reactions are…
Use when the user asks for media-forge video audio, dialogue, lip-sync, music, sound effects, ambience, beat-sync, audio-reference mapping, desync troubleshooting, or sound-driven…
Expert knowledge for Microsoft Foundry Local (aka Azure AI Foundry Local) development including troubleshooting, best practices, decision making, configuration, and integrations &…
MIST's offline cloned voice -- speak a reply aloud as an embedded audio track, narrate a note or briefing to MP3, run one-off TTS, or transcribe audio to text.
Local 24x7 OpenAI-compatible API server for STT/TTS, powered by MLX on your Mac.
Clone a voice from YouTube audio and synthesize custom dialogue using OmniVoice (k2-fsa, diffusion-LM TTS), with gemma4:e2b (Ollama) for fast reference transcription.
AI-powered WhatsApp features: auto-replies, voice transcription, RAG knowledge base, and style profiles.
De novo motif discovery and known motif enrichment analysis using HOMER and MEME-ChIP. Identify transcription factor binding motifs in ChIP-seq, ATAC-seq, or other genomi — from…
Analyze transcription factor motif accessibility variability using chromVAR. Use when identifying which TF motifs show variable accessibility across samples or conditions — from…
Find patterns, motifs, and subsequences in biological sequences using Biopython. Use when searching for transcription factor binding sites, regulatory elements, or any se — from…
Integrate multiple ENCODE data types (RNA-seq, ATAC-seq, Histone ChIP-seq, TF ChIP-seq) for a tissue/cell type to build a comprehensive regulatory landscape.
Patterns for building multimodal AI applications that combine text, images, audio, and video. Covers vision APIs, audio transcription, and unified pipelines.
PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen).
Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves…
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification.
Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware…
Tools, patterns, and utilities for generating professional music with realistic instrument sounds. Write custom compositions using music21 or learn from existing MIDI files.
AI music generation powered by CellCog。Original instrumental and vocal tracks, 5 seconds to 10 m。Use when 需要视频处理、音频编辑、媒体转换、配音生成时使用。不适用于版权受保护的媒体内容处理。
Productor musical senior. Logic, Ableton, FL Studio, Pro Tools. Mixing, mastering, sound design.
Optimize and format prompts specifically for AI music generation platforms like Suno and Udio, including platform-specific syntax and tag optimization
Use when a claim needs proof instead of trust: porting or replacing an algorithm (FSRS-6 vs py-fsrs parity), deciding whether ~225 reviews can train 21 FSRS weights, a read-db.py…
Turn a new Italian-tutoring recording (.mov) into a transcript and a published Hugo blog post with lesson notes and a vocab list.
Local-first text-to-speech and speech-to-text via the Voicebox MCP server. Generates speech from cloned or preset voice profiles for agent notifications, content voiceovers, and…
[omh] User-sent media - audio, video, YouTube links, screenshots, receipts, OCR, meeting recordings, transcripts, timestamps, and clip summaries, gated for source, permission, and…
Speech-to-text via OmniRoute using OpenAI /v1/audio/transcriptions format with auto-fallback across Whisper, AssemblyAI, Deepgram, Azure STT.
Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML.
OpenAI API integration for building AI-powered applications. Use when working with OpenAI's Chat Completions API, Python SDK (openai), TypeScript SDK (openai), tool use/function…
Text-to-speech conversion using OpenAI's TTS API for generating high-quality, natural-sounding audio.
Local Whisper: audio transcription, multi-language, word timestamps, speaker diarization
API-based speech-to-text transcription through OpenAI. No local model downloads, no GPU, no Python ML stack — just an API key and a shell script.
Analisi acida e basata sui fatti della stabilità delle release di OpenClaudio. Da usare prima di ogni 'openclaw update' per evitare regressioni o leak di log.
High-quality voice synthesis with 9 personas, 11 languages, and streaming using Voice.ai API. — from satoshistackalotto/skills
Audio transcription and text-to-speech generation using OpenRouter API. Use when the user needs to transcribe audio files to text or generate speech/audio from text.
Transcribe audio files via OpenRouter using audio-capable models (Gemini, GPT-4o-audio, etc). — from ndesv21/awesome-openclaw-skills-1
Transcribe audio files via OpenRouter using audio-capable models (Gemini, GPT-4o-audio, etc). — from majiayu000/claude-skill-registry
Use when working with speech recognition, text-to-speech, wake word detection, clap detection, VAD, audio preprocessing, or any audio I/O - covers the full audio pipeline from…
Otter.ai transcription CLI - list, search, download, and sync meeting transcripts to CRM.
Production pipeline for interactive and generative visual art using p5.js. Creates browser-based sketches, generative art, data visualizations, interactive experiences, 3D scenes,…
Baut die vierte Padlet-Spalte als Pendant zu Reiter 4 der Step-Plan-Excel. Workflow-Karten mit nummerierten Checkbox-Schritten, Rechtsgrundlage, Tags fuer Unterzeichner und…
Convert PDF files to MP3 audio using MiniMax MCP Server's text-to-audio tool. Use when user wants to convert a PDF to audio/MP3, create audiobook from PDF, or text-to-speech for…
Pedalboard is a Python library built by Spotify for working with audio: reading, writing, rendering, and adding studio-quality effects.
Estimate the phoneme count and optional IPA transcription for one English word in a declared dialect, with uncertainty made explicit.
Sound systems of human language -- phoneme inventories, the International Phonetic Alphabet, articulatory and acoustic phonetics, phonological rules, suprasegmental features…
Run fast, high-quality neural text-to-speech locally with Piper. Supports 20+ languages with compact ONNX voice models, no cloud API required, and produces natural-sounding speech…