Generating, converting and transcribing media. Every tool below was read from that server's own
tools/list, so the name and arguments are what the server exposes, not what
its description claims.
| Tool | Server |
|---|---|
| oruk_transcribe_audioanswers Transcribe prerecorded English audio to text with time-ordered segments and word timings. Use this when only the words matter. Accepts a public audio URL or base64 bytes (wav/flac/mp3/m4a/ogg/webm, ≤30 MB / ≤60 min). ... arguments: audio_url, audio_base64, filename, model, detail, api_key | oruk Speech ai.oruk |
| extract_video_thumbnailanswers Extract preview thumbnail images from a video at regular intervals and return them as a ZIP file. 동영상에서 일정 구간마다 미리보기 이미지를 추출해 ZIP 파일로 반환합니다. [호출당 10포인트] arguments: width, count, video_url | APICK Vision app.apick |
| transcribe_audio_proanswers Transcribe audio with Brainiall Speech Pro — multilingual transcription.
Supports 99 languages with automatic language detection, word-level
timestamps, per-word confidence scores, and optional speaker diarization
(i... arguments: audio_base64, language, diarize | Brainiall Pronunciation com.brainiall |
| audio-transcribe-longanswers PAID MCP TOOL — $0.209 USDC per successful call via native x402. Discovery is free. Transcribe LONG audio (15-30 minutes) from any publicly accessible URL using OpenAI Whisper. Supports mp3, mp4, m4a, wav, webm, ogg, ... arguments: url, language, source_rights_confirmed | The Stall ai.intuitek.the-stall |
| audio-transcribe-shortanswers PAID MCP TOOL — $0.039 USDC per successful call via native x402. Discovery is free. Transcribe SHORT audio clips (0-5 minutes) from any publicly accessible URL using OpenAI Whisper. Supports mp3, mp4, m4a, wav, webm, ... arguments: url, language, source_rights_confirmed | The Stall ai.intuitek.the-stall |
| audio-transcribeanswers PAID MCP TOOL — $0.129 USDC per successful call via native x402. Discovery is free. Transcribe MEDIUM-length audio (5-15 minutes) from any publicly accessible URL using OpenAI Whisper. Supports mp3, mp4, m4a, wav, web... arguments: url, language, source_rights_confirmed | The Stall ai.intuitek.the-stall |
| transcribe_audioanswers Transcribe spoken audio (from a public URL) to text. arguments: audio_url | 370.ai — AI Gateway: Video (Seedance 2.0, Wan, HappyHorse), Image, Speech + 100+ Chat Models ai.router |
| kuaishou_submit_video_speech_text_by_video_urlanswers 提交快手作品视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: video_url | SocialDataX 快手 Kuaishou MCP com.52choujiang |
| kuaishou_submit_video_speech_text_by_photo_idanswers 根据快手 photo_id 提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: photo_id | SocialDataX 快手 Kuaishou MCP com.52choujiang |
| kuaishou_get_video_speech_text_jobanswers 查询快手视频口播转文字任务状态;用户已提供有效 job_id 时直接使用,否则使用 submit 工具返回值;不重复提交,每次最多等待 240 秒。 arguments: job_id | SocialDataX 快手 Kuaishou MCP com.52choujiang |
| wechat_submit_video_speech_text_by_video_urlanswers 提交微信视频号视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: video_url | SocialDataX 微信视频号 WeChat Channels MCP com.52choujiang |
| wechat_submit_video_speech_text_by_encrypted_object_idanswers 根据微信视频号 encrypted_object_id 提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: encrypted_object_id | SocialDataX 微信视频号 WeChat Channels MCP com.52choujiang |
| wechat_get_video_speech_text_jobanswers 继续查询用户提供的有效 job_id,或 submit 工具返回的 job_id 对应的微信视频号口播转文字任务状态;每次最多等待 240 秒,不创建新任务或触发重处理。 arguments: job_id | SocialDataX 微信视频号 WeChat Channels MCP com.52choujiang |
| instagram_submit_video_speech_text_by_post_urlanswers 根据 Instagram 普通视频帖子或 Reels 链接提交口播转文字任务;图片帖和轮播帖不支持,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: post_url | SocialDataX Instagram MCP com.52choujiang |
| instagram_submit_video_speech_text_by_post_idanswers 根据 Instagram 普通视频帖子或 Reels 的 post_id 提交口播转文字任务;图片帖和轮播帖不支持,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: post_id | SocialDataX Instagram MCP com.52choujiang |
| instagram_get_video_speech_text_jobanswers 根据有效 job_id 查询 Instagram 口播转文字任务;用户已提供时直接使用,否则使用 submit 工具返回的 job_id;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。 arguments: job_id | SocialDataX Instagram MCP com.52choujiang |
| weibo_submit_video_speech_text_by_post_urlanswers 根据微博帖子链接、短链接或分享文案提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: post_url | SocialDataX 微博 Weibo MCP com.52choujiang |
| weibo_submit_video_speech_text_by_post_idanswers 根据微博 post_id 提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: post_id | SocialDataX 微博 Weibo MCP com.52choujiang |
| weibo_get_video_speech_text_jobanswers 继续查询用户提供的有效 job_id,或微博视频口播转文字 submit 工具返回的 job_id;每次最多等待 240 秒,不创建新任务或触发重处理。 arguments: job_id | SocialDataX 微博 Weibo MCP com.52choujiang |
| bilibili_submit_video_speech_text_by_video_urlanswers 根据 Bilibili 视频链接、短链接或分享文案提交口播转文字任务;每次处理一个分 P,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: video_url | SocialDataX B站 Bilibili MCP com.52choujiang |
| bilibili_submit_video_speech_text_by_bvidanswers 根据 Bilibili BV 号提交 P1 口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: bvid | SocialDataX B站 Bilibili MCP com.52choujiang |
| bilibili_get_video_speech_text_jobanswers 根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询 Bilibili 口播转文字任务状态;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。 arguments: job_id | SocialDataX B站 Bilibili MCP com.52choujiang |
| tiktok_submit_video_speech_text_by_urlanswers 提交 TikTok 视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: video_url | SocialDataX TikTok MCP com.52choujiang |
| tiktok_submit_video_speech_text_by_aweme_idanswers 根据 TikTok 视频 aweme_id 提交口播转文字任务;可直接使用视频 post_id,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: aweme_id | SocialDataX TikTok MCP com.52choujiang |
| tiktok_get_video_speech_text_jobanswers 根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询 TikTok 视频口播转文字任务状态;用于继续未完成任务,每次最多等待 240 秒,不触发重处理,也不要重复提交任务。 arguments: job_id | SocialDataX TikTok MCP com.52choujiang |
| youtube_submit_video_speech_text_by_urlanswers 提交 YouTube 视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: video_url | SocialDataX YouTube MCP com.52choujiang |
| youtube_submit_video_speech_text_by_video_idanswers 根据 YouTube video_id 提交口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: video_id | SocialDataX YouTube MCP com.52choujiang |
| youtube_get_video_speech_text_jobanswers 根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询 YouTube 视频口播转文字任务状态;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。 arguments: job_id | SocialDataX YouTube MCP com.52choujiang |
| douyin_submit_video_speech_text_by_video_urlanswers 提交抖音作品视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: video_url | SocialDataX 抖音 Douyin MCP com.52choujiang |
| douyin_submit_video_speech_text_by_aweme_idanswers 根据抖音 aweme_id 提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: aweme_id | SocialDataX 抖音 Douyin MCP com.52choujiang |
| douyin_get_video_speech_text_jobanswers 根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询抖音视频口播转文字任务状态;用于继续未完成任务,每次最多等待 240 秒,不触发重处理,也不要重复提交任务。 arguments: job_id | SocialDataX 抖音 Douyin MCP com.52choujiang |
| zhihu_submit_video_speech_text_by_video_urlanswers 根据知乎独立视频页或带视频回答页链接提交口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: video_url | SocialDataX 知乎 Zhihu MCP com.52choujiang |
| zhihu_submit_video_speech_text_by_zvideo_idanswers 根据知乎独立视频的数字 zvideo_id 提交口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: zvideo_id | SocialDataX 知乎 Zhihu MCP com.52choujiang |
| zhihu_get_video_speech_text_jobanswers 根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询知乎视频口播转文字任务;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。 arguments: job_id | SocialDataX 知乎 Zhihu MCP com.52choujiang |
| xhs_submit_video_speech_text_by_note_urlanswers 根据小红书视频笔记链接、短链接或分享文案提交口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: note_url | SocialDataX 小红书 Xiaohongshu XHS RedNote MCP com.52choujiang |
| xhs_submit_video_speech_text_by_note_idanswers 根据小红书 note_id 提交视频笔记口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: note_id | SocialDataX 小红书 Xiaohongshu XHS RedNote MCP com.52choujiang |
| xhs_get_video_speech_text_jobanswers 根据用户提供的有效 job_id,或提交工具返回的 job_id 查询任务状态;用于继续未完成任务,每次最多等待 240 秒,不触发重处理,也不要重复提交任务。 arguments: job_id | SocialDataX 小红书 Xiaohongshu XHS RedNote MCP com.52choujiang |
| x_submit_video_speech_text_by_post_urlanswers 根据 X / Twitter 视频帖子链接提交口播转文字任务;只处理当前帖子自身首个可用 MP4 视频,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: post_url | SocialDataX X / Twitter MCP com.52choujiang |
| x_submit_video_speech_text_by_post_idanswers 根据 X / Twitter 视频帖子的 post_id 提交口播转文字任务;只处理当前帖子自身首个可用 MP4 视频,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。 arguments: post_id | SocialDataX X / Twitter MCP com.52choujiang |
| x_get_video_speech_text_jobanswers 根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询 X / Twitter 口播转文字任务;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。 arguments: job_id | SocialDataX X / Twitter MCP com.52choujiang |
| transcribe_audioanswers Fetch an audio file from a URL and transcribe it to text with open-source Whisper (100 languages, self-hosted). Good for voice memos, podcast clips and meeting recordings up to ~15 MB. Example — GET https://ainetcafe.... arguments: url, language | com.ainetcafe/netcafe-docs com.ainetcafe |
| transcribe_audioanswers Transcribe audio to text with word-level timestamps.
Converts spoken English audio into text with optional word-level timestamps
and per-word confidence scores.
Args:
audio_base64: Base64-encoded audio data (WAV... arguments: audio_base64, audio_format, include_timestamps | Brainiall Pronunciation com.brainiall |
| transcribe_audioanswers Transcribe Kurdish audio (Sorani or Kurmanji) to text. Requires an STT API key; usage is metered per audio minute against your plan. Max 4MB of decoded audio over MCP — for larger files call POST https://www.kurdishtt... arguments: audio_base64, dialect, filename, mime_type, context | kurdish-tts-stt com.kurdishtts |
| AI-Video-Generator-Image-to-Videoanswers YouCam AI Video Generator transforms text prompts and images into captivating videos with ease. Powered by advanced AI technology, it creates realistic motion effects that bring your ideas and photos to life. With a w... arguments: request, polling | YouCam for Creators com.makeupar |
| generate_videoanswers Render a RAW video clip from your own prompt and return its served mp4 URL. For finished brand ADS prefer render_ad (it runs the Studio quality pipeline — composited text, clean speech, end card, music); use this for ... arguments: prompt, raw, refImage, refVideo, durationSeconds, aspectRatio | Hermoso io.github.hermoso-ai |
| transcribe_audio Transcribe audio to text with word-level timestamps.
Converts spoken English audio into text with optional word-level timestamps
and per-word confidence scores.
Args:
audio_base64: Base64-encoded audio data (WAV... arguments: audio_base64, audio_format, include_timestamps | Speech AI - Pronunciation, STT & TTS io.github.fasuizu-br |
| transcribe_audio_pro Transcribe audio with Whisper Large V3 Turbo — multilingual STT.
Supports 99 languages with automatic language detection, word-level
timestamps, per-word confidence scores, and optional speaker diarization
(identifie... arguments: audio_base64, language, diarize | Speech AI - Pronunciation, STT & TTS io.github.fasuizu-br |
| speech_analytics_retranscribe Retranscribes an uploaded file. arguments: uuid, callback_url, insights | Wavix io.github.wavix |
| kooma_transcribe_from_urlanswers Transcrit en texte un fichier audio ou video accessible par une URL publique en HTTPS. La langue est detectee automatiquement. Le fichier est plafonne a 25 Mo. arguments: url | Kooma — Bambara AI ai.kooma |
| createVideoanswers Generate a short video clip from a source image and a motion text prompt (image-to-video). The job result is the video URL and its actual duration in seconds - there is no separate polling step. Optionally pass `final... arguments: requestBody | Ludo AI Game Assets ai.ludo |
| createVideoFromReferencesanswers Generate a video from 1-5 reference images and a text prompt (references-to-video). Unlike createVideo, which animates a single source image, this composes a new scene that borrows characters, objects, and style from ... arguments: requestBody | Ludo AI Game Assets ai.ludo |
| oruk_analyze_speechanswers Transcribe English audio AND score how it was said in one call: transcript, tagged transcript, calibrated emotion (15 labels) and speaking-style (16 labels) scores, and time-local segments. Use this when the user care... arguments: audio_url, audio_base64, filename, model, detail, api_key | oruk Speech ai.oruk |
| lip_sync_videoanswers Lip-sync audio onto one of your videos. RECOMMENDED: action="create" with engine="best" + video_url + sound_file (base64 data URI) — syncs the whole clip on the highest-quality engine, no face step needed. Kling flow ... arguments: action, engine, video_url, session_id, face_id, sound_file | ai.switchapp/switch ai.switchapp |
| talking_avatar_videoanswers Turn a face photo into a lip-synced talking-head video that speaks your text (or your audio). Provide image_url (a clear face photo) and either script (text to speak, max 2500 characters) or audio_url. Optional voice_... arguments: image_url, script, audio_url, voice_id, language, voice_settings | ai.switchapp/switch ai.switchapp |
| generate_audioanswers Generate spoken audio from text: narration, a voiceover, a read-aloud script, or a multi-voice dialogue. Pass text (up to 2048 chars) — the words to be spoken. To speak in one of YOUR saved voices, pass voice with the... arguments: text, voice, reference_audio_url, reference_audio_urls, image_url, speech_rate | ai.switchapp/switch ai.switchapp |
| analyze_video_reportanswers Run the FULL Switch Vision analysis on a video, the same premium report the Video Analysis page produces: it watches AND listens in three forensic passes and returns a structured report with every category: overview (... arguments: video_url, duration_seconds, question, force | ai.switchapp/switch ai.switchapp |
| transcribeanswers Transcribe audio or video to text, including per-word timestamps for precise editing. Three-call flow: (1) call with `filename` to receive {job_id, payment_challenge}; (2) pay via MPP, then call with `job_id` + `payme... arguments: filename, job_id, payment_credential | Weftly ai.weftly |
| generate_videoanswers Generate a video clip for the company (xAI Imagine Video 1.5, 3 credits): text-to-video, image-to-video, multi-image reference (up to 7), or video edit with native audio. Use when an operator or agent needs social, pr... arguments: companyId, prompt, duration, aspect_ratio, resolution, image_url | com.getfreedomos/freedom-mcp com.getfreedomos |
| ic_transcribe_submitanswers Queue an audio file for offline transcription + speaker diarization by IC's GPU worker. Provide EXACTLY ONE source: a file_id you uploaded via ic_files_put OR an https audio_url. SIZE: ic_files_put caps at ~3.2MB raw ... arguments: file_id, audio_url, language, num_speakers_hint | Immersive Commons com.immersivecommons |
| get_info_thumbnailsanswers **object** A map of thumbnail images associated with the video. For each object in the map, the key is the name of the thumbnail image, and the value is an object that contains other information about the thumbnail. G... arguments: id | com.jojapi/cloud-api-hub-youtube-downloader com.jojapi |
https://neuronto.com/tools?q=..., or connect an agent to
https://neuronto.com/mcp and call find_tool. No key, no signup.