Index / Verified tools / PDF and document extraction

PDF and document extraction

Reading text, tables and structure out of PDFs and office documents. Every tool below was read from that server's own tools/list, so the name and arguments are what the server exposes, not what its description claims.

ToolServer
PdfOcr_PdfToLinesWithLocationanswers
Convert a PDF into text lines with location Converts a PDF into lines/text with location information and other metdata via Optical Character Recognition. This API is intended to be run on scanned documents. If you w...
ocrapi
cloudmersive.com
PdfOcr_PdfToWordsWithLocationanswers
Convert a PDF into words with location Converts a PDF into words/text with location information and other metdata via Optical Character Recognition. This API is intended to be run on scanned documents. If you want t...
ocrapi
cloudmersive.com
docx_to_pdfanswers
Convert a .docx document (base64) to a PDF. Returns a file_id and a ~1h URL.
arguments: docx
PreteWorks API
com.preteworks
pdf-ocranswers
スキャンPDFから日本語テキスト抽出 (Browser-based tool)
JobDoneBot
io.github.acromoney888
PdfOcr_Postanswers
Converts an uploaded PDF file into text via Optical Character Recognition. PdfOcr
ocrapi
cloudmersive.com
search_type_document_pdfanswers
Search API for 'PDF File' entry type API to search for entries of type PDF File type/type_document_pdf
geodesystems.com:443
geodesystems.com
app.documents.sales.invoices.pdfanswers
Download the invoice as pdf Invoices
Sinao API
sinao.app
app.documents.sales.quotes.pdfanswers
Download the quote as pdf Quotes
Sinao API
sinao.app
pdf_to_docxanswers
Convert a PDF file to a DOCX (Word) file. PDF 파일을 DOCX 파일로 변환해 반환합니다. PDF 형식의 파일만 허용됩니다. [호출당 30포인트]
arguments: pdf_url
APICK Convert
app.apick
docx_to_pdfanswers
Convert a DOCX (Word) file to a PDF file. DOCX 파일을 PDF 파일로 변환해 반환합니다. DOCX 형식의 파일만 허용됩니다. [호출당 80포인트]
arguments: docx_url
APICK Convert
app.apick
parse_documentanswers
Parse an office document (docx, pptx, xlsx, pdf, odt, ods, odp, rtf, epub, csv, doc, ppt) into clean GitHub-Flavored Markdown. Provide a document URL or a local file path. Deterministic Rust converter — headings, tabl...
arguments: url, path, format
ai.clawfetch/mcp
ai.clawfetch
filegraph_pdf_text_conversion_docsanswers
Docs for /pdf/to-text, /pdf/to-docx, /pdf/to-image — extract text with ocr, convert pdf to docx or images.
arguments: language, endpoint
document-processing
ai.filegraph
filegraph_office_documents_docsanswers
Docs for /docx/to-text, /docx/to-pdf, /xlsx/to-text, /xlsx/to-pdf, /xlsx/to-csv, /pptx/to-text, /pptx/to-pdf, /pptx/to-images — extract text/data, convert word/excel/powerpoint to pdf/csv/images.
arguments: language, endpoint
document-processing
ai.filegraph
ocr_imageanswers
Extract text from an image with GPU OCR. Best-in-class Arabic (plus Persian/Urdu) accuracy, manga-aware vertical Japanese, and strong English, French, Spanish, German, Chinese, Korean, Russian, Italian and Portuguese ...
arguments: image_base64, lang, mode, quality, api_key
ocr
com.auto-reader
parse_documentanswers
Parse a quoted PDF to structured Markdown (hallucination-guarded OCR for scanned pages, Japanese-strong). Get quote_id from quote_parse first; price is fixed by the quote.
arguments: quote_id, x_payment
com.mart402/extract
com.mart402
document_pipelineanswers
One call: PDF invoice to parse (hallucination-guarded OCR) + field extraction + arithmetic verification (+optional issuer enrich). Flat $0.05/document; inputs validated before payment.
arguments: url, enrich, x_payment
com.mart402/extract
com.mart402
read_documentanswers
Read a document (PDF or image) from a URL and return its contents as markdown (tables preserved) or plain text. Costs $0.00075 per page, billed to the PennyOCR account; the response includes the exact cost_usd and per...
arguments: url, pages, output, max_pages, max_cost_usd
com.pennyocr/ocr
com.pennyocr
extract_pdfanswers
Extract text and metadata from a PDF document. Price: $0.004
arguments: url
WebLens
dev.weblens
convert_documentanswers
Convert any document to another format without storing a template. Supports 100+ input/output format combinations: Office documents, PDFs, images, web pages, spreadsheets, and more. The source file can be a local path...
arguments: file, convertTo, converter, reportName, hardRefresh, outputPath
Carbone MCP
io.carbone
read_pdfanswers
Read a PDF and return its text as markdown (or plain text). Accepts a public URL or base64 bytes. Extracts the embedded text layer; a scanned, image-only PDF returns a needs-OCR notice instead of empty text. Priced pe...
arguments: source, options, idempotencyKey
docweave-mcp
io.github.nicolasmartalog
render_docxanswers
Render an editable Word (.docx) document from a Kamy template and data. Takes the same { template, data } payload as render_pdf but produces a different container — reach for it when the recipient has to EDIT the docu...
arguments: template, data, name
Kamy
dev.kamy
convert_documentanswers
Convert a file you already hold — .docx, .xlsx or .csv — into a PDF, preserving its existing content. Pass the bytes base64-encoded together with the original filename, which is what the API uses to detect the input t...
arguments: filename, fileBase64, name
Kamy
dev.kamy
pdf_text_extractanswers
Extract text from a PDF: url or base64 data. No OCR — text-based PDFs only.
arguments: url, data, maxChars, headers
ToolSnap MCP
app.toolsnap
formation_extract_rsa_terms_ocranswers
Runs vision OCR on an already-uploaded signed Restricted Stock Purchase/Award Agreement (RSA) PDF for the active company and returns the founder and company names plus structured terms (total_shares, unvested_shares, ...
arguments: agentCode, companyId, sourceS3Uri
Lovie Company Formation
co.lovie
formation_get_ocr_upload_urlanswers
Returns a presigned S3 PUT URL for uploading a company document (kind=SAFE for a SAFE PDF, kind=RSA for a signed restricted stock purchase agreement, kind=CAP_TABLE for a cap-table file, kind=WIRE_PROOF for a wire pay...
arguments: companyId, kind, mimeType
Lovie Company Formation
co.lovie
read_brand_documentanswers
Read one indexed brand document. Returns the indexed metadata (doc_type, summary, key_topics, entities, key_quotes) plus the document body as plain text in content.text — every mime (PDF, DOCX, PPTX, XLSX, markdown, C...
arguments: document_id
Heista
co.heista
document.extractanswers
Extract clean text and metadata from supplied PDF, DOCX, HTML, Markdown, CSV, JSON, YAML, or plain-text documents.
arguments: filename, content_base64
API Acre
com.apiacre
verify_documentanswers
Fetches a document (PDF, DOCX, or plain text) from a URL and runs the same 3-layer citation verification as verify_citation. Use this when you have a link to a brief rather than pasted text.
arguments: url
Citation Safe Verifier
com.citationsafe
document.extract_textanswers
Extract plain text from a PDF or image (base64-encoded). Use when you need raw text for downstream AI analysis (summarization, claim checking, structured extraction). For documents at a public URL, use url.extract ins...
arguments: document_base64, mime_type
com.docimprint/api
com.docimprint
summarize_documentanswers
Summarize a web page, PDF, Office file (.docx/.xlsx/.csv), or raw text using AI. Provide ONE source: 'url', 'pdf' (file_id or base64), 'file' (base64), or 'text'. Optional 'max_words'. Returns the summary.
arguments: url, pdf, file, text, max_words
PreteWorks API
com.preteworks
export_documentanswers
Get a downloadable PNG image, PDF, or Word (.docx) file of a document already published to SendPage. Use this when the user wants a file copy — to attach, print, or hand over — not just the link. Returns a URL to fetc...
arguments: shortId, format
com.sendpageapp/sendpage
com.sendpageapp
ifr_ingest_documentanswers
Extract structured data from a PDF (contract, invoice, lease, bond indenture) using Google Document AI. Requires docuScanAccess on the token — returns 403 otherwise; do NOT retry or offer document scanning again this ...
arguments: pdf_base64, findTheValueWithThisLabel, val_type, doc_id, start_anchor, end_anchor
com.spocont/ifrCoworker
com.spocont
extract_documentanswers
Extract STRUCTURED FIELDS from a document image: invoices, receipts, ID cards — or any custom JSON schema you supply. Every field returns {value, confidence, box} where the confidence and box come from the OCR geometr...
arguments: image_base64, preset, schema, lang, api_key
ocr
com.auto-reader
ocr_and_translateanswers
One call: OCR an image, then translate every line into target_lang. Arabic-first OCR and manga-aware Japanese with right-to-left-aware layout, followed by LLM translation. Automatic source-language detection. Provide ...
arguments: image_base64, target_lang, api_key
ocr
com.auto-reader
get_ttab_documentanswers
Retrieve and read one specific TTABVUE filing by proceeding number and entry number. Returns filing metadata, a working USPTO TTABVUE viewer link, direct PDF link when available, extraction/readability status, documen...
arguments: proceeding_number, entry_number, text_query, max_chars
GleanMark Trademark Search
com.gleanmark
word_to_pdfanswers
Convert a Word document to PDF in your browser — nothing is uploaded. A HelpySelf tool (helpyself.com).
arguments: data, pageSize
HelpySelf
com.helpyself
pdf_mergeanswers
Merge PDF files into one document. A HelpySelf tool (helpyself.com).
arguments: files
HelpySelf
com.helpyself
compare_pdfsanswers
MANDATORY for all document comparison requests. Compare two PDFs side-by-side. When the user asks to compare, diff, or find differences between two PDFs, you MUST call this tool — NEVER attempt to compare documents u...
arguments: session_id, job_id_a, job_id_b
KDAN PDF
com.kdandoc.mcp
delete_pdf_pageanswers
Delete one or more pages from a PDF and return a new versioned file. MANDATORY: Before calling this tool, call 'check_upload_status' with the session_id to confirm the file exists and retrieve the latest job_id. Skip...
arguments: session_id, job_id, page_numbers
KDAN PDF
com.kdandoc.mcp
view_pdfanswers
Display an already-uploaded PDF in an interactive viewer widget. Call this after delete_pdf_page, set_password, change_password, redact_pii, and redact_by_text_range to show the updated PDF. Also call when the user e...
arguments: session_id, job_id
KDAN PDF
com.kdandoc.mcp
get_pdf_page_imagesanswers
Use this when you need to visually inspect PDF pages to identify content that cannot be detected from text alone — for example, pages containing a logo, a photograph, a watermark, a QR code, or any visual element. Ret...
arguments: session_id, job_id, start_page, count
KDAN PDF
com.kdandoc.mcp
list_government_documentsanswers
List a government's adopted plan and budget documents. Each entry carries `covers` — every general-plan element that document answers for. A city that publishes one bound general plan satisfies all eight ...
arguments: government_id, include_superseded
OpenPublica
com.openpublica
get_pdf_metadataanswers
Read a PDF's document properties (title, author, subject, keywords, creator, producer, dates, page count).
arguments: file
com.pdfia/pdf-tools
com.pdfia
set_pdf_metadataanswers
Rewrite a PDF's document properties. Pass empty strings to strip identifying metadata before sharing. Returns base64.
arguments: file, title, author, subject, keywords, creator
com.pdfia/pdf-tools
com.pdfia
pdf_extractanswers
Extract any PDF up to 10 MB by URL. format='text' returns one clean plain-text string; format='json' returns a per-page text array plus document metadata. Parsed in-Worker, no upstream service. Paid: call without x_pa...
arguments: url, format, x_payment
Professor Sausages — Web & Documents
com.professorsausages
ocr_to_markdownanswers
Convert PDF/image files into Markdown through Frenchie
arguments: file_path, uploaded_file_reference, api_key
Frenchie
io.github.lab94
render_html_to_pdfanswers
Render an HTML document to a PDF, returned base64-encoded. `html` is a complete HTML string. Inline any images, fonts and CSS as `data:` URIs — external http(s) resources are not fetched, and no JavaScript is execute...
arguments: html
Pdfrender
dev.pdfrender
render_documentanswers
Generate a document by merging a Carbone template with JSON data. Two modes: (1) pass templateId to use a previously uploaded template; (2) pass template (file path, URL, or base64) to upload and render in a single re...
Carbone MCP
io.carbone
generate_pdfanswers
Generate a PDF from raw HTML, a public URL, or a template + JSON data. Returns the PDF as base64. Priced per document; retries with the same idempotencyKey never double-generate. The canonical way for an AI agent to t...
arguments: source, options, idempotencyKey
docweave-mcp
io.github.nicolasmartalog
render_pdfanswers
Render a PDF from a Kamy template and data, and wait for it. This is the default document tool: it blocks until the file exists and hands back { id, url, bytes, durationMs, templateId, createdAt } in one call, where u...
arguments: template, data, format
Kamy
dev.kamy
merge_pdfsanswers
Concatenate 2–20 existing renders into a single new PDF, in exactly the order the ids are given. Both the inputs and the output are Kamy render ids, so this is the composition step after several render_pdf / convert_d...
arguments: renderIds, name
Kamy
dev.kamy
split_pdfanswers
Extract page ranges from one existing render into separate new PDFs — the inverse of merge_pdfs. Each range you pass produces its own render, returned in the same order, so one call can both halve a contract and peel ...
arguments: renderId, ranges
Kamy
dev.kamy
edit_pdfanswers
Modify an existing PDF: fill AcroForm fields by name, stamp text at absolute coordinates, or paint opaque boxes over regions. Operations apply in the order given, to either a render you own (renderId) or a PDF Kamy do...
arguments: renderId, pdfUrl, operations, name, flattenFields
Kamy
dev.kamy
extract_documentanswers
Extract structured data from a PDF (invoice, receipt, contract, ID document, or any form). Returns the parsed JSON plus a public verify URL that proves the extraction matches the source. Use this when an agent needs t...
arguments: template, source_url, source_base64
Kamy
dev.kamy
verify_pdf_signatureanswers
Turn PDF bytes you are holding into their SHA-256 digest and the matching kamy.dev/verify/{sha256} page URL. Purely local: the MCP Worker hashes the base64 in memory, makes no Kamy API call, stores nothing and forward...
arguments: pdfBase64
Kamy
dev.kamy
image-to-pdfanswers
Combine multiple images into a single PDF document. (Browser-based tool)
JobDoneBot
io.github.acromoney888
word-to-pdfanswers
DOCXをPDFに変換 (Browser-based tool)
JobDoneBot
io.github.acromoney888
pdf-to-wordanswers
PDFをWord(.docx)に変換 (Browser-based tool)
JobDoneBot
io.github.acromoney888
pdf-stamperanswers
Apply Japanese hanko stamps to PDF documents. (Browser-based tool)
JobDoneBot
io.github.acromoney888
ClassifyDocumentanswers
<p>Creates a new document classification request to analyze a single document in real-time, using a previously created and trained custom model and an endpoint.</p> <p>You can input plain text or you can upload a sing...
Amazon Comprehend
amazonaws.com
Showing 60 of 1,247 verified tools. Search the whole set, including full input schemas, at https://neuronto.com/tools?q=..., or connect an agent to https://neuronto.com/mcp and call find_tool. No key, no signup.

"Answers" means the endpoint responded to a handshake when last probed, and "auth required" means it demanded credentials. Both are statements about reachability, never about trustworthiness.

Other capabilities

Files and storage
3,539 verified tools
Analytics and monitoring
3,533 verified tools
Maps, location and weather
3,378 verified tools
Payments and billing
2,448 verified tools
GitHub and version control
2,375 verified tools
Calendar and scheduling
2,238 verified tools
Crypto and blockchain
1,979 verified tools
Images, audio and video
1,916 verified tools
HR and recruiting
1,459 verified tools