---
name: crawlforge
description: CrawlForge is a metered web-access layer for AI agents. It runs as a local MCP server (31 tools over stdio, priced in credits) and as a REST API at https://www.crawlforge.dev. This file tells an agent how to install it, how to obtain an API key with or without a human at the keyboard, and which tool to pick for each kind of web task so it spends the fewest credits.
---

# CrawlForge agent onboarding

Every URL in this file is absolute and rooted at `https://www.crawlforge.dev`.

## What CrawlForge is

- 31 MCP tools for reading, searching, crawling and extracting from the web: `scrape`, `fetch_url`, `search_web`, `batch_scrape`, `crawl_deep`, `deep_research`, `stealth_mode`, `extract_structured` and the rest.
- Every call is metered in credits: 1 to 10 per call for every tool except `agent`, which costs 8 to 18 (see the ladder below). A new account gets 1,000 one-time free credits; paid plans start at $19/month.
- The MCP server is a local stdio process (`crawlforge-mcp-server` on npm) that calls the CrawlForge API with your key. There is no hosted multi-tenant MCP endpoint, so nothing to point a remote MCP client at.
- The same tools are available over REST without the MCP server (see "REST without the MCP server").

## Install

No install is needed to run the server once:

```bash
npx -y crawlforge-mcp-server@latest mcp
```

A global install adds the `crawlforge` CLI, which handles login and client registration:

```bash
npm i -g crawlforge-mcp-server
```

Claude Desktop can install it in one click from the `.mcpb` bundle attached to the latest release: https://github.com/mysleekdesigns/crawlforge-mcp/releases/latest

Per-client configuration (Claude Desktop, Claude Code, Cursor, VS Code, Windsurf, Zed, Gemini CLI, Cline, n8n) is generated at https://www.crawlforge.dev/docs/integration/configure, with a page per client at `https://www.crawlforge.dev/docs/integration/<client>` where `<client>` is one of `claude-desktop`, `claude-code`, `cursor-ide`, `vscode`, `windsurf`, `zed`, `gemini-cli`, `cline`, `n8n`.

The server reads its key from the `CRAWLFORGE_API_KEY` environment variable. Set `CRAWLFORGE_TOOL_GROUPS` to load a subset of tools (for example `basic,search,scrape`; unset loads all 31) when your client caps the number of tools it will register.

## Get credentials

A key looks like `cf_live_…`. It is shown once at creation, so store it as soon as you receive it.

The human must have a CrawlForge account before any of the paths below will work: https://www.crawlforge.dev/signup (1,000 one-time free credits; email verification is required before a key can be created).

### Path A: the human already has a key

Ask for it, then export it (or put it in the `env` block of the client config you generate):

```bash
export CRAWLFORGE_API_KEY=cf_live_...
```

Do not hand-write `~/.crawlforge/config.json`: the server also expects the account id in that file, so a key-only file is rejected at startup. `crawlforge login` (Path B) writes it correctly.

### Path B: CLI handoff (`crawlforge login`)

With the CLI installed:

```bash
crawlforge login
```

The CLI prints an approval URL, the human opens it and clicks Approve while signed in, and the CLI polls until the key arrives, then saves it to `~/.crawlforge/config.json`. It never edits any MCP client's config file. To register the server with a client afterwards:

```bash
crawlforge init --client claude-code      # or claude-desktop, cursor
```

### Path C: raw HTTP handoff (no CLI)

Any agent that can run `openssl` and make HTTP requests can do what the CLI does. Generate a session id and a PKCE-style verifier/challenge pair:

```bash
SESSION_ID=$(openssl rand -hex 16)
CODE_VERIFIER=$(openssl rand -base64 32 | tr '+/' '-_' | tr -d '=\n')
CODE_CHALLENGE=$(printf '%s' "$CODE_VERIFIER" | openssl dgst -sha256 -binary | openssl base64 -A | tr '+/' '-_' | tr -d '=')
```

Ask the human to open this URL (replace `<agent name>` with a label for the key, for example `my-agent`; it is optional and defaults to `CLI login`):

```
https://www.crawlforge.dev/cli-auth?session_id=$SESSION_ID&code_challenge=$CODE_CHALLENGE&name=<agent name>
```

Signing in first is fine: the page comes back to the approval with the query string intact. Meanwhile poll every 3 seconds:

```bash
curl -s -X POST https://www.crawlforge.dev/api/auth/cli/status \
  -H 'Content-Type: application/json' \
  -d "{\"session_id\":\"$SESSION_ID\",\"code_verifier\":\"$CODE_VERIFIER\"}"
```

Responses:

```json
{"status":"pending"}
```

```json
{"status":"complete","api_key":"cf_live_…","key_id":"…","key_name":"my-agent","email":"…"}
```

The key is delivered exactly once; the next poll is `pending` again, so store it the moment it arrives. The approval expires 10 minutes after the human approves. A `403` with code `CLI_AUTH_VERIFIER_MISMATCH` means the verifier does not match the challenge the human approved: start over with a fresh pair. A `429` means you are polling faster than 60 times a minute from one address.

Constraints the endpoint enforces: `session_id` is 16 to 128 characters of `[A-Za-z0-9_-]`, `code_verifier` is 43 to 128 of the same, and `code_challenge` is the unpadded base64url SHA-256 of the verifier (43 characters).

## Tool ladder

Pick ONE tool per step from this ladder.

- Read one page whose URL you have -> `scrape` (2 credits); ask for every format you need in that call: markdown, links, metadata, html, screenshot, json.
- Raw JSON/XML/API body, headers or status -> `fetch_url` (1).
- Find pages for a query -> `search_web` (5); snippets often answer without a scrape.
- Google organic position -> `serp_rank` (5).
- Reddit -> `reddit_search` (5).
- Blocked (403/429/CAPTCHA/challenge page/empty shell) -> `stealth_mode` with `operation: "scrape"` (5), or `scrape` with `escalate: true`; never stealth first.
- Needs a click, login or scroll -> `scrape_with_actions` (5).
- Several calls have to work on the same page, or a login must survive between them -> `browser_session` (per operation: `open` 3, `read` 2, `snapshot`/`act`/`screenshot`/`close`/`list` 1); `open`, then `snapshot` for `@e1` refs, `act` on a ref, `read`, `close`. `scrape_with_actions` is one-shot — its browser closes when the call returns — and one API key holds one session at a time.
- 2-50 known URLs -> one `batch_scrape` (5), never a loop of `scrape` calls.
- A site's URL list -> `map_site` (2); many pages of one site -> `crawl_deep` (4).
- Exact values from a Next.js/Nuxt/Redux payload -> `extract_embedded_state` (2).
- Known CSS selectors -> `scrape_structured` (2); fields you can describe but not select -> `extract_structured` (3).
- A report from several sources -> ONE `deep_research` call (10 + ~1 per 5 sources).
- Open question with no URLs -> `agent` (8).
- A result that came back `truncated: true` with a `result_handle` -> `read_result` (1).

Rules: never fetch a URL whose content you already have; error results end with `Next step:` naming the tool to try.

Full per-tool reference: https://www.crawlforge.dev/docs/api-reference

## REST without the MCP server

Every tool is also a REST endpoint:

```
POST https://www.crawlforge.dev/api/v1/tools/<tool>
X-API-Key: cf_live_...
Content-Type: application/json
```

```bash
curl -s -X POST https://www.crawlforge.dev/api/v1/tools/scrape \
  -H "X-API-Key: $CRAWLFORGE_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"url":"https://example.com","formats":["markdown","links"]}'
```

Every response uses one envelope:

```json
{ "success": true, "data": { "...": "..." }, "credits_used": 2, "credits_remaining": 998 }
```

Credits are charged only after a call succeeds. The OpenAPI description of all endpoints is at https://www.crawlforge.dev/openapi.json.

Official SDKs wrap the same endpoints:

```bash
npm install crawlforge-sdk     # TypeScript / JavaScript
pip install crawlforge         # Python
```

## Discovery

- https://www.crawlforge.dev/.well-known/mcp.json — the MCP registry `server.json` for the npm package and its stdio transport.
- https://www.crawlforge.dev/llms.txt — short index of the site for language models.
- https://www.crawlforge.dev/llms-full.txt — the full documentation in one file.
- https://www.crawlforge.dev/openapi.json — the REST API description.
- https://www.crawlforge.dev/api/mcp answers 404 on purpose: there is no hosted MCP endpoint.
