Index / Verified tools / Images, audio and video

Images, audio and video

Generating, converting and transcribing media. Every tool below was read from that server's own tools/list, so the name and arguments are what the server exposes, not what its description claims.

ToolServer
oruk_transcribe_audioanswers
Transcribe prerecorded English audio to text with time-ordered segments and word timings. Use this when only the words matter. Accepts a public audio URL or base64 bytes (wav/flac/mp3/m4a/ogg/webm, ≤30 MB / ≤60 min). ...
arguments: audio_url, audio_base64, filename, model, detail, api_key
oruk Speech
ai.oruk
extract_video_thumbnailanswers
Extract preview thumbnail images from a video at regular intervals and return them as a ZIP file. 동영상에서 일정 구간마다 미리보기 이미지를 추출해 ZIP 파일로 반환합니다. [호출당 10포인트]
arguments: width, count, video_url
APICK Vision
app.apick
transcribe_audio_proanswers
Transcribe audio with Brainiall Speech Pro — multilingual transcription. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (i...
arguments: audio_base64, language, diarize
Brainiall Pronunciation
com.brainiall
audio-transcribe-longanswers
PAID MCP TOOL — $0.209 USDC per successful call via native x402. Discovery is free. Transcribe LONG audio (15-30 minutes) from any publicly accessible URL using OpenAI Whisper. Supports mp3, mp4, m4a, wav, webm, ogg, ...
arguments: url, language, source_rights_confirmed
The Stall
ai.intuitek.the-stall
audio-transcribe-shortanswers
PAID MCP TOOL — $0.039 USDC per successful call via native x402. Discovery is free. Transcribe SHORT audio clips (0-5 minutes) from any publicly accessible URL using OpenAI Whisper. Supports mp3, mp4, m4a, wav, webm, ...
arguments: url, language, source_rights_confirmed
The Stall
ai.intuitek.the-stall
audio-transcribeanswers
PAID MCP TOOL — $0.129 USDC per successful call via native x402. Discovery is free. Transcribe MEDIUM-length audio (5-15 minutes) from any publicly accessible URL using OpenAI Whisper. Supports mp3, mp4, m4a, wav, web...
arguments: url, language, source_rights_confirmed
The Stall
ai.intuitek.the-stall
transcribe_audioanswers
Transcribe spoken audio (from a public URL) to text.
arguments: audio_url
370.ai — AI Gateway: Video (Seedance 2.0, Wan, HappyHorse), Image, Speech + 100+ Chat Models
ai.router
kuaishou_submit_video_speech_text_by_video_urlanswers
提交快手作品视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: video_url
SocialDataX 快手 Kuaishou MCP
com.52choujiang
kuaishou_submit_video_speech_text_by_photo_idanswers
根据快手 photo_id 提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: photo_id
SocialDataX 快手 Kuaishou MCP
com.52choujiang
kuaishou_get_video_speech_text_jobanswers
查询快手视频口播转文字任务状态;用户已提供有效 job_id 时直接使用,否则使用 submit 工具返回值;不重复提交,每次最多等待 240 秒。
arguments: job_id
SocialDataX 快手 Kuaishou MCP
com.52choujiang
wechat_submit_video_speech_text_by_video_urlanswers
提交微信视频号视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: video_url
SocialDataX 微信视频号 WeChat Channels MCP
com.52choujiang
wechat_submit_video_speech_text_by_encrypted_object_idanswers
根据微信视频号 encrypted_object_id 提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: encrypted_object_id
SocialDataX 微信视频号 WeChat Channels MCP
com.52choujiang
wechat_get_video_speech_text_jobanswers
继续查询用户提供的有效 job_id,或 submit 工具返回的 job_id 对应的微信视频号口播转文字任务状态;每次最多等待 240 秒,不创建新任务或触发重处理。
arguments: job_id
SocialDataX 微信视频号 WeChat Channels MCP
com.52choujiang
instagram_submit_video_speech_text_by_post_urlanswers
根据 Instagram 普通视频帖子或 Reels 链接提交口播转文字任务;图片帖和轮播帖不支持,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: post_url
SocialDataX Instagram MCP
com.52choujiang
instagram_submit_video_speech_text_by_post_idanswers
根据 Instagram 普通视频帖子或 Reels 的 post_id 提交口播转文字任务;图片帖和轮播帖不支持,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: post_id
SocialDataX Instagram MCP
com.52choujiang
instagram_get_video_speech_text_jobanswers
根据有效 job_id 查询 Instagram 口播转文字任务;用户已提供时直接使用,否则使用 submit 工具返回的 job_id;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。
arguments: job_id
SocialDataX Instagram MCP
com.52choujiang
weibo_submit_video_speech_text_by_post_urlanswers
根据微博帖子链接、短链接或分享文案提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: post_url
SocialDataX 微博 Weibo MCP
com.52choujiang
weibo_submit_video_speech_text_by_post_idanswers
根据微博 post_id 提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: post_id
SocialDataX 微博 Weibo MCP
com.52choujiang
weibo_get_video_speech_text_jobanswers
继续查询用户提供的有效 job_id,或微博视频口播转文字 submit 工具返回的 job_id;每次最多等待 240 秒,不创建新任务或触发重处理。
arguments: job_id
SocialDataX 微博 Weibo MCP
com.52choujiang
bilibili_submit_video_speech_text_by_video_urlanswers
根据 Bilibili 视频链接、短链接或分享文案提交口播转文字任务;每次处理一个分 P,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: video_url
SocialDataX B站 Bilibili MCP
com.52choujiang
bilibili_submit_video_speech_text_by_bvidanswers
根据 Bilibili BV 号提交 P1 口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: bvid
SocialDataX B站 Bilibili MCP
com.52choujiang
bilibili_get_video_speech_text_jobanswers
根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询 Bilibili 口播转文字任务状态;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。
arguments: job_id
SocialDataX B站 Bilibili MCP
com.52choujiang
tiktok_submit_video_speech_text_by_urlanswers
提交 TikTok 视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: video_url
SocialDataX TikTok MCP
com.52choujiang
tiktok_submit_video_speech_text_by_aweme_idanswers
根据 TikTok 视频 aweme_id 提交口播转文字任务;可直接使用视频 post_id,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: aweme_id
SocialDataX TikTok MCP
com.52choujiang
tiktok_get_video_speech_text_jobanswers
根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询 TikTok 视频口播转文字任务状态;用于继续未完成任务,每次最多等待 240 秒,不触发重处理,也不要重复提交任务。
arguments: job_id
SocialDataX TikTok MCP
com.52choujiang
youtube_submit_video_speech_text_by_urlanswers
提交 YouTube 视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: video_url
SocialDataX YouTube MCP
com.52choujiang
youtube_submit_video_speech_text_by_video_idanswers
根据 YouTube video_id 提交口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: video_id
SocialDataX YouTube MCP
com.52choujiang
youtube_get_video_speech_text_jobanswers
根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询 YouTube 视频口播转文字任务状态;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。
arguments: job_id
SocialDataX YouTube MCP
com.52choujiang
douyin_submit_video_speech_text_by_video_urlanswers
提交抖音作品视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: video_url
SocialDataX 抖音 Douyin MCP
com.52choujiang
douyin_submit_video_speech_text_by_aweme_idanswers
根据抖音 aweme_id 提交视频口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: aweme_id
SocialDataX 抖音 Douyin MCP
com.52choujiang
douyin_get_video_speech_text_jobanswers
根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询抖音视频口播转文字任务状态;用于继续未完成任务,每次最多等待 240 秒,不触发重处理,也不要重复提交任务。
arguments: job_id
SocialDataX 抖音 Douyin MCP
com.52choujiang
zhihu_submit_video_speech_text_by_video_urlanswers
根据知乎独立视频页或带视频回答页链接提交口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: video_url
SocialDataX 知乎 Zhihu MCP
com.52choujiang
zhihu_submit_video_speech_text_by_zvideo_idanswers
根据知乎独立视频的数字 zvideo_id 提交口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: zvideo_id
SocialDataX 知乎 Zhihu MCP
com.52choujiang
zhihu_get_video_speech_text_jobanswers
根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询知乎视频口播转文字任务;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。
arguments: job_id
SocialDataX 知乎 Zhihu MCP
com.52choujiang
xhs_submit_video_speech_text_by_note_urlanswers
根据小红书视频笔记链接、短链接或分享文案提交口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: note_url
SocialDataX 小红书 Xiaohongshu XHS RedNote MCP
com.52choujiang
xhs_submit_video_speech_text_by_note_idanswers
根据小红书 note_id 提交视频笔记口播转文字任务;提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: note_id
SocialDataX 小红书 Xiaohongshu XHS RedNote MCP
com.52choujiang
xhs_get_video_speech_text_jobanswers
根据用户提供的有效 job_id,或提交工具返回的 job_id 查询任务状态;用于继续未完成任务,每次最多等待 240 秒,不触发重处理,也不要重复提交任务。
arguments: job_id
SocialDataX 小红书 Xiaohongshu XHS RedNote MCP
com.52choujiang
x_submit_video_speech_text_by_post_urlanswers
根据 X / Twitter 视频帖子链接提交口播转文字任务;只处理当前帖子自身首个可用 MP4 视频,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: post_url
SocialDataX X / Twitter MCP
com.52choujiang
x_submit_video_speech_text_by_post_idanswers
根据 X / Twitter 视频帖子的 post_id 提交口播转文字任务;只处理当前帖子自身首个可用 MP4 视频,提交后最多等待 240 秒,未完成时返回 job_id 和下一步查询动作。
arguments: post_id
SocialDataX X / Twitter MCP
com.52choujiang
x_get_video_speech_text_jobanswers
根据用户提供的有效 job_id,或 submit 工具返回的 job_id 查询 X / Twitter 口播转文字任务;每次最多等待 240 秒,不触发重处理,也不要重复提交任务。
arguments: job_id
SocialDataX X / Twitter MCP
com.52choujiang
transcribe_audioanswers
Fetch an audio file from a URL and transcribe it to text with open-source Whisper (100 languages, self-hosted). Good for voice memos, podcast clips and meeting recordings up to ~15 MB. Example — GET https://ainetcafe....
arguments: url, language
com.ainetcafe/netcafe-docs
com.ainetcafe
transcribe_audioanswers
Transcribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV...
arguments: audio_base64, audio_format, include_timestamps
Brainiall Pronunciation
com.brainiall
transcribe_audioanswers
Transcribe Kurdish audio (Sorani or Kurmanji) to text. Requires an STT API key; usage is metered per audio minute against your plan. Max 4MB of decoded audio over MCP — for larger files call POST https://www.kurdishtt...
arguments: audio_base64, dialect, filename, mime_type, context
kurdish-tts-stt
com.kurdishtts
AI-Video-Generator-Image-to-Videoanswers
YouCam AI Video Generator transforms text prompts and images into captivating videos with ease. Powered by advanced AI technology, it creates realistic motion effects that bring your ideas and photos to life. With a w...
arguments: request, polling
YouCam for Creators
com.makeupar
generate_videoanswers
Render a RAW video clip from your own prompt and return its served mp4 URL. For finished brand ADS prefer render_ad (it runs the Studio quality pipeline — composited text, clean speech, end card, music); use this for ...
arguments: prompt, raw, refImage, refVideo, durationSeconds, aspectRatio
Hermoso
io.github.hermoso-ai
transcribe_audio
Transcribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV...
arguments: audio_base64, audio_format, include_timestamps
Speech AI - Pronunciation, STT & TTS
io.github.fasuizu-br
transcribe_audio_pro
Transcribe audio with Whisper Large V3 Turbo — multilingual STT. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (identifie...
arguments: audio_base64, language, diarize
Speech AI - Pronunciation, STT & TTS
io.github.fasuizu-br
speech_analytics_retranscribe
Retranscribes an uploaded file.
arguments: uuid, callback_url, insights
Wavix
io.github.wavix
kooma_transcribe_from_urlanswers
Transcrit en texte un fichier audio ou video accessible par une URL publique en HTTPS. La langue est detectee automatiquement. Le fichier est plafonne a 25 Mo.
arguments: url
Kooma — Bambara AI
ai.kooma
createVideoanswers
Generate a short video clip from a source image and a motion text prompt (image-to-video). The job result is the video URL and its actual duration in seconds - there is no separate polling step. Optionally pass `final...
arguments: requestBody
Ludo AI Game Assets
ai.ludo
createVideoFromReferencesanswers
Generate a video from 1-5 reference images and a text prompt (references-to-video). Unlike createVideo, which animates a single source image, this composes a new scene that borrows characters, objects, and style from ...
arguments: requestBody
Ludo AI Game Assets
ai.ludo
oruk_analyze_speechanswers
Transcribe English audio AND score how it was said in one call: transcript, tagged transcript, calibrated emotion (15 labels) and speaking-style (16 labels) scores, and time-local segments. Use this when the user care...
arguments: audio_url, audio_base64, filename, model, detail, api_key
oruk Speech
ai.oruk
lip_sync_videoanswers
Lip-sync audio onto one of your videos. RECOMMENDED: action="create" with engine="best" + video_url + sound_file (base64 data URI) — syncs the whole clip on the highest-quality engine, no face step needed. Kling flow ...
arguments: action, engine, video_url, session_id, face_id, sound_file
ai.switchapp/switch
ai.switchapp
talking_avatar_videoanswers
Turn a face photo into a lip-synced talking-head video that speaks your text (or your audio). Provide image_url (a clear face photo) and either script (text to speak, max 2500 characters) or audio_url. Optional voice_...
arguments: image_url, script, audio_url, voice_id, language, voice_settings
ai.switchapp/switch
ai.switchapp
generate_audioanswers
Generate spoken audio from text: narration, a voiceover, a read-aloud script, or a multi-voice dialogue. Pass text (up to 2048 chars) — the words to be spoken. To speak in one of YOUR saved voices, pass voice with the...
arguments: text, voice, reference_audio_url, reference_audio_urls, image_url, speech_rate
ai.switchapp/switch
ai.switchapp
analyze_video_reportanswers
Run the FULL Switch Vision analysis on a video, the same premium report the Video Analysis page produces: it watches AND listens in three forensic passes and returns a structured report with every category: overview (...
arguments: video_url, duration_seconds, question, force
ai.switchapp/switch
ai.switchapp
transcribeanswers
Transcribe audio or video to text, including per-word timestamps for precise editing. Three-call flow: (1) call with `filename` to receive {job_id, payment_challenge}; (2) pay via MPP, then call with `job_id` + `payme...
arguments: filename, job_id, payment_credential
Weftly
ai.weftly
generate_videoanswers
Generate a video clip for the company (xAI Imagine Video 1.5, 3 credits): text-to-video, image-to-video, multi-image reference (up to 7), or video edit with native audio. Use when an operator or agent needs social, pr...
arguments: companyId, prompt, duration, aspect_ratio, resolution, image_url
com.getfreedomos/freedom-mcp
com.getfreedomos
ic_transcribe_submitanswers
Queue an audio file for offline transcription + speaker diarization by IC's GPU worker. Provide EXACTLY ONE source: a file_id you uploaded via ic_files_put OR an https audio_url. SIZE: ic_files_put caps at ~3.2MB raw ...
arguments: file_id, audio_url, language, num_speakers_hint
Immersive Commons
com.immersivecommons
get_info_thumbnailsanswers
**object** A map of thumbnail images associated with the video. For each object in the map, the key is the name of the thumbnail image, and the value is an object that contains other information about the thumbnail. G...
arguments: id
com.jojapi/cloud-api-hub-youtube-downloader
com.jojapi
Showing 60 of 637 verified tools. Search the whole set, including full input schemas, at https://neuronto.com/tools?q=..., or connect an agent to https://neuronto.com/mcp and call find_tool. No key, no signup.

"Answers" means the endpoint responded to a handshake when last probed, and "auth required" means it demanded credentials. Both are statements about reachability, never about trustworthiness.

Other capabilities

Analytics and monitoring
1,099 verified tools
Files and storage
1,035 verified tools
Calendar and scheduling
723 verified tools
Crypto and blockchain
708 verified tools
Payments and billing
706 verified tools
Maps, location and weather
598 verified tools
Web scraping and browser automation
490 verified tools
Social media
481 verified tools