services / pdf-extract
textoperational · 12 ms
PDF Text Extraction
Extract all text from a PDF. Send as pdf_base64 (base64-encoded PDF, max ~10 MB decoded). Returns text (full concatenated text), pages array (per-page text + char_count), page_count, and metadata (title, author, creator). Encode with: Buffer.from(pdfBytes).toString('base64'). Ideal for RAG pipelines, document QA, or LLM ingestion.
Run free trial ↗1 free call per day with the example input. Paid: $0.004 USDC, no limit.
Call it
Input
| Field | Type | Description |
|---|---|---|
| pdf_base64 * | string | Base64-encoded PDF file content. Decode a PDF file to base64 and pass it here. Max ~10 MB (unencoded). |
| max_pages | integer = 50 | Maximum number of pages to extract. Default: 50. Use to limit processing time for large PDFs. |
Output
| Field | Type | Description |
|---|---|---|
| text | string | Full extracted text from all pages, joined with newlines. Preserves paragraph structure where possible. |
| page_count | integer | Total number of pages in the PDF |
| pages | array | Per-page text content (first max_pages pages) |
| metadata | object | PDF document metadata (if available) |
| file_size_bytes | integer | Size of the decoded PDF in bytes |
| extracted_at | string | ISO 8601 timestamp of extraction |
Example response (data)
{
"text": "Quarterly Financial Report Q1 2026\n\nExecutive Summary\nTotal revenue increased by 23% year-over-year...",
"page_count": 12,
"pages": [
{
"page": 1,
"text": "Quarterly Financial Report Q1 2026\n\nExecutive Summary",
"char_count": 54
},
{
"page": 2,
"text": "Table of Contents\n1. Revenue Overview\n2. Cost Analysis",
"char_count": 58
}
],
"metadata": {
"title": "Q1 2026 Financial Report",
"author": "Finance Department",
"creator": "Adobe Acrobat"
},
"file_size_bytes": 245760,
"extracted_at": "2026-04-10T14:00:00.000Z"
}