The Problem
Critical numbers live in PDFs: annual reports, price lists, regulatory filings, research papers. Copying them out by hand does not scale, and generic PDF libraries return a wall of text with no structure you can query.
Use Cases
01Quick Answer
Point process_document (2 credits) at a PDF URL or file to get text, sections, and metadata, then pass that text to extract_with_llm (3 credits) with a JSON schema to get typed fields back. About 5 credits per document, and extraction defaults to a local Ollama model so contents stay on your machine.
02The brief
Critical numbers live in PDFs: annual reports, price lists, regulatory filings, research papers. Copying them out by hand does not scale, and generic PDF libraries return a wall of text with no structure you can query.
CrawlForge process_document extracts text, sections, and metadata from a PDF URL or file, and extract_with_llm turns that text into the exact JSON shape you asked for -- running against local Ollama by default, so document contents never leave your machine.
03In code
1// Pull structured fields out of a PDF report2const doc = await mcp.process_document({3 source: "https://example.com/2026-annual-report.pdf",4 sourceType: "pdf_url",5 options: { extractText: true, extractMetadata: true },6});7 8// Ask for exactly the fields you need, as JSON9const fields = await mcp.extract_with_llm({10 content: doc.text,11 prompt: "Extract fiscal year, revenue, operating margin, and headcount.",12 schema: {13 type: "object",14 properties: {15 fiscalYear: { type: "string" },16 revenue: { type: "string" },17 operatingMargin: { type: "string" },18 headcount: { type: "number" },19 },20 },21});22 23console.log(doc.metadata.pageCount);24console.log(fields.data);04The pipeline
05Questions
04
Point process_document at the PDF URL or file to get text, sections, and metadata, then send that text to extract_with_llm with a JSON schema describing the fields you want. You get typed JSON back instead of a wall of text.
No. process_document accepts a URL or a local file path via sourceType — url, pdf_url, file, or pdf_file — and options such as pageRange and maxPages let you process only the pages you care about.
extract_with_llm defaults to a local Ollama model, so document contents never leave your machine. Pass provider openai or anthropic with the matching API key when you want a cloud model instead.
About 5 credits per document: 2 for process_document and 3 for extract_with_llm. The free 1,000-credit tier covers roughly 200 documents, enough to validate the pipeline against a real archive before upgrading.
06Keep exploring
Turn documentation sites into clean, chunk-ready markdown for retrieval-augmented generation. Crawl, strip the boilerplate, and summarize before you embed.
Build AI agents that search the web, synthesize findings, and deliver up-to-date research. Every claim cites its source so answers stay verifiable.
Start forging
Every new account gets 1,000 free credits. No credit card required.