extract_ with_ llm
AI-powered extraction that runs on a local Ollama model by default — no LLM API key required. Optionally route to OpenAI or Anthropic when you need a hosted model. Give it a prompt (and optionally a JSON Schema) and get back structured data.
Use Cases
Local-First Extraction
Run extractions against Ollama on your own machine — zero LLM API costs and private by default.
Schema-Driven Data Lakes
Combine a prompt with a JSON Schema to populate typed rows for your warehouse or graph store.
Multi-Provider Failover
Start on local Ollama, fall back to OpenAI or Anthropic for higher-stakes pages by toggling one parameter.
Endpoint
/api/v1/tools/extract_with_llmParameters
`provider` defaults to `"auto"`: On a self-run MCP server, "auto" uses your local Ollama — no LLM API key required. On the hosted CrawlForge API there is no local Ollama, so "openai" or "anthropic" work only when the matching key is configured on the execution backend.
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
url | string | Optional | - | URL to fetch and extract from. Either url or content is required. Example: https://example.com/article/42 |
content | string | Optional | - | Raw text or HTML content to extract from. Either url or content is required. Example: "<html>...</html>" |
prompt | string | Required | - | Natural-language instructions guiding the LLM extraction Example: Extract the headline, author, and three key takeaways |
schema | object | Optional | - | Optional JSON Schema describing the data structure to extract Example: {"type":"object","properties":{"title":{"type":"string"}},"required":["title"]} |
provider | string | Optional | auto | LLM provider: "ollama" (local, default), "openai", "anthropic", or "auto" Example: ollama |
model | string | Optional | - | Model identifier. Defaults per provider: llama3.2, gpt-4o-mini, claude-haiku-4-5-20251001 Example: llama3.2 |
maxTokens | number | Optional | 4096 | Maximum tokens for the LLM response Example: 4096 |
Request Examples
cURL — local Ollama (default, no API key)
curl -X POST https://crawlforge.dev/api/v1/tools/extract_with_llm \
-H "X-API-Key: cf_test_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/article/42",
"prompt": "Extract the headline, author, and three key takeaways",
"provider": "ollama"
}'TypeScript — OpenAI with schema
// npm install crawlforge-sdk
import { CrawlForge } from 'crawlforge-sdk';
const client = new CrawlForge({ apiKey: process.env.CRAWLFORGE_API_KEY });
const result = await client.extractWithLlm({
url: 'https://example.com/product/123',
prompt: 'Extract product name, price in USD, and stock status',
provider: 'openai',
model: 'gpt-4o-mini',
schema: {
type: 'object',
properties: {
title: { type: 'string' },
price: { type: 'number' },
in_stock: { type: 'boolean' },
},
required: ['title', 'price'],
},
});
// result.data is untyped in crawlforge-sdk 0.1 — its shape is the Response Example below.
const { extracted } = result.data as {
extracted: { title: string; price: number; in_stock: boolean };
};
console.log(extracted);Python — Anthropic
# pip install crawlforge
from crawlforge import CrawlForge
client = CrawlForge() # reads CRAWLFORGE_API_KEY
result = client.extract_with_llm(
url='https://example.com/article/42',
prompt='Extract headline, author, publish date (ISO 8601), and tags',
provider='anthropic',
model='claude-haiku-4-5-20251001',
schema={
'type': 'object',
'properties': {
'headline': {'type': 'string'},
'author': {'type': 'string'},
'published_at': {'type': 'string'},
'tags': {'type': 'array'},
},
'required': ['headline'],
},
)
# result.data is a plain dict — its shape is the Response Example below.
print(result.data['extracted'])Response Example
{ "success": true, "data": { "extracted": { "headline": "How Local LLMs Are Changing Data Pipelines", "author": "Jane Doe", "takeaways": [ "Lower cost", "Better privacy", "Faster iteration" ] }, "model_used": "llama3.2", "provider": "ollama" }, "credits_used": 3, "credits_remaining": 997, "processing_time": 3500}data.providerResolved provider — "ollama" when provider is "auto"data.model_usedDefault model per provider unless you specify onecredits_usedFlat 3 credits regardless of providerCredit Cost
Tip: Use list_ollama_models first to discover which local models are available before sending an extraction.
Related Tools
Ready to run LLM extractions on your own machine? Sign up for free and get 1,000 credits.