agentsvc.io

services / pdf-extract

textoperational · 12 ms

PDF Text Extraction

Extract all text from a PDF. Send as pdf_base64 (base64-encoded PDF, max ~10 MB decoded). Returns text (full concatenated text), pages array (per-page text + char_count), page_count, and metadata (title, author, creator). Encode with: Buffer.from(pdfBytes).toString('base64'). Ideal for RAG pipelines, document QA, or LLM ingestion.

Run free trial ↗1 free call per day with the example input. Paid: $0.004 USDC, no limit.

Call it

import { wrapFetchWithPayment, x402Client } from "@x402/fetch";
import { registerExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";

const client = new x402Client();
registerExactEvmScheme(client, { signer: privateKeyToAccount(process.env.EVM_PRIVATE_KEY) });
const payFetch = wrapFetchWithPayment(fetch, client);

const res = await payFetch("https://agentsvc.io/api/v1/proxy/pdf-extract", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({"pdf_base64":"<base64 of a PDF>","max_pages":10}),
});
const { data, payment } = await res.json();  // payment.transaction = on-chain receipt

Input

FieldTypeDescription
pdf_base64 *stringBase64-encoded PDF file content. Decode a PDF file to base64 and pass it here. Max ~10 MB (unencoded).
max_pagesinteger = 50Maximum number of pages to extract. Default: 50. Use to limit processing time for large PDFs.

Output

FieldTypeDescription
textstringFull extracted text from all pages, joined with newlines. Preserves paragraph structure where possible.
page_countintegerTotal number of pages in the PDF
pagesarrayPer-page text content (first max_pages pages)
metadataobjectPDF document metadata (if available)
file_size_bytesintegerSize of the decoded PDF in bytes
extracted_atstringISO 8601 timestamp of extraction

Example response (data)

{
  "text": "Quarterly Financial Report Q1 2026\n\nExecutive Summary\nTotal revenue increased by 23% year-over-year...",
  "page_count": 12,
  "pages": [
    {
      "page": 1,
      "text": "Quarterly Financial Report Q1 2026\n\nExecutive Summary",
      "char_count": 54
    },
    {
      "page": 2,
      "text": "Table of Contents\n1. Revenue Overview\n2. Cost Analysis",
      "char_count": 58
    }
  ],
  "metadata": {
    "title": "Q1 2026 Financial Report",
    "author": "Finance Department",
    "creator": "Adobe Acrobat"
  },
  "file_size_bytes": 245760,
  "extracted_at": "2026-04-10T14:00:00.000Z"
}