An agent has never seen the page it is about to click. scrape_with_actions asks it to name up to 20 actions anyway — type into #user, click .btn-primary, wait 1000ms — and then closes the browser when the call returns. Guess wrong at step 3 and the chain stops there, the result comes back success: false, and the 5 credits are already gone.
CrawlForge MCP v6.6.0 fixes the loop rather than the guess. The new browser_session tool — our 31st — keeps one real browser page alive between tool calls, so an agent can open a page, look at it, act on what it saw, look again, and read the result. The looking is the part that matters, and it has a name: snapshot.
Table of Contents
- What Shipped
- Why a One-Shot Chain Goes Blind
- Snapshot: The Agent Looks Before It Acts
- The Session Loop, End to End
- Refs Fail Loudly, Never Silently
- Snapshot Landed in scrape_
with_ actions Too - What a Session Costs
- Limits, Stated Plainly
- One Session From the CLI
- When to Use Which
- How to Upgrade
What Shipped
v6.6.0 is one tool and one primitive:
browser_session— a single tool with anoperationenum:open,snapshot,act,read,screenshot,close,list. The page, its cookies and its login survive between calls until the session expires.snapshot— an accessibility-style tree of the live page in which every interactive element carries a stable ref (@e1,@e2). Any action'sselectoraccepts a ref in place of a CSS selector.- Tool count goes from 30 to 31, on both surfaces: MCP and the REST API.
- A
crawlforge-browser-sessionsagent skill and acrawlforge browserCLI command ship alongside it.
Nothing was renamed and no existing tool changed shape or price, so it is a drop-in upgrade.
Why a One-Shot Chain Goes Blind
scrape_with_actions is a good tool with one structural limit: the action array is written before anything has loaded. The agent is choosing #login-email or input[name="email"] or .form-field:first-child from memory of how login forms usually look, not from this login form.
When one of those guesses misses, ActionExecutor stops the chain — continueOnError defaults to false — and the tool returns a result carrying success: false and the error rather than throwing. The billing follows the call, not the verdict, so the 5 credits are spent on a chain that got three actions in.
The failure mode is not the cost. It is that the agent's next move is to guess again, with no more information than it had the first time. A one-shot tool cannot tell it what the page contains, because the browser is gone by the time the result arrives.
Snapshot: The Agent Looks Before It Acts
snapshot walks the live DOM and emits one indented line per meaningful node, modelled on the accessibility tree — role, accessible name, and for interactive nodes a ref:
[document] "Sign in"
@e1 [textbox] "Email"
@e2 [textbox] "Password"
@e3 [button] "Sign in"
[link] "Forgot password?"Two properties make this useful rather than decorative. Structural nodes (headings, landmarks, forms) appear as context but get no ref, because they are not targets. And every ref is backed by a data-cf-ref attribute stamped onto the element during the walk, so @e1 resolves to an ordinary [data-cf-ref="e1"] selector and works with every existing action path — including the stealth browser's human-behaviour code, which takes a raw selector string.
We wrote the walk ourselves rather than wrapping Playwright's ariaSnapshot(), which emits YAML with no element refs at all, and page._snapshotForAI(), which is private API we will not depend on.
interactive_only defaults to true and max_nodes defaults to 200, so a 4,000-element application page comes back as a list an agent can actually read.
The Session Loop, End to End
Here is a login and a read, using the crawlforge-sdk package. Five calls, one browser page, one set of cookies.
// npm install crawlforge-sdk
import { CrawlForge } from 'crawlforge-sdk';
const client = new CrawlForge({ apiKey: process.env.CRAWLFORGE_API_KEY });
// 1. Open the session (3 credits). One API key may hold one session at a time.
const opened = await client.browserSession({
operation: 'open',
url: 'https://app.example.com/login',
ttl: 600
});
const { sessionId } = opened.data as { sessionId: string };
try {
// 2. Look before acting (1 credit). Every interactive element comes back
// with a stable ref, so the next call targets what is actually there.
const seen = await client.browserSession({ operation: 'snapshot', session_id: sessionId });
console.log((seen.data as { snapshot: { tree: string } }).snapshot.tree);
// 3. Act on those refs in a separate call (1 credit) — same page, same cookies.
await client.browserSession({
operation: 'act',
session_id: sessionId,
actions: [
{ type: 'type', selector: '@e1', text: 'user@example.com' },
{ type: 'type', selector: '@e2', text: 'secret123' },
{ type: 'click', selector: '@e3' },
{ type: 'wait', duration: 1000 }
]
});
// 4. Read the page the login landed on (2 credits) — the live DOM with the
// session's cookies, not a fresh fetch of the URL.
const page = await client.browserSession({
operation: 'read',
session_id: sessionId,
formats: ['markdown']
});
const { title, content } = page.data as { title: string; content: { markdown: string } };
console.log(title, content.markdown.length);
} finally {
// 5. Close it rather than waiting for the TTL (1 credit).
await client.browserSession({ operation: 'close', session_id: sessionId });
}Step 4 is worth dwelling on. read extracts from the page the session is holding, not from a re-fetch of its URL — a re-fetch would arrive without the cookies and before everything the session has clicked, which is the whole reason to have one. Formats are markdown, html, text and json, and you can ask for several in one call.
Refs Fail Loudly, Never Silently
Refs are invalidated by navigation. That is not a caveat, it is the design: a ref that survived a page change would point at whatever element happened to be third in the new document, and the agent would click it without knowing.
So the ref table lives in a WeakMap keyed by page and is cleared on every main-frame navigation. Act on a stale ref and you get the reason and the fix, not a mystery:
Stale element ref @e3: the page navigated since the last snapshot —
take a new snapshot before acting on refs.An unknown ref is just as explicit: "Unknown element ref @e9: the current snapshot has 4 refs (@e1-@e4) — take a new snapshot." An agent can act on either message on its own. A silent mis-click is the failure we spent the design budget avoiding.
Snapshot Landed in scrape_with_actions Too
You do not need a session to look first. snapshot also shipped as an action type inside scrape_with_actions, so a one-shot chain can observe, then act on the refs its own snapshot produced, inside the same 5-credit call:
{
"url": "https://news.ycombinator.com/login",
"actions": [
{ "type": "snapshot" },
{ "type": "type", "selector": "@e1", "text": "reader" },
{ "type": "click", "selector": "@e3" }
]
}That covers the common case where the agent needs to see the page but the whole flow still fits in one call. Reach for a session when the flow does not: when a later decision depends on what an earlier step returned, or when a login has to hold across several reads.
What a Session Costs
browser_session prices per operation, because one session is many calls and a flat price would charge the ceiling for every cheap one:
| Operation | Credits | What it does |
|---|---|---|
open | 3 | Launches a browser context and navigates |
read | 2 | Extracts the live DOM into your formats |
snapshot | 1 | One injected walk over the open page |
act | 1 | Up to 20 actions against the open page |
screenshot | 1 | PNG or JPEG, full page or one element |
close | 1 | Releases the page and its context |
list | 1 | Your open sessions and their clocks |
The login flow above is 8 credits: 3 + 1 + 1 + 2 + 1. Add the re-snapshot that the post-login navigation calls for and it is 9. A scrape_with_actions call attempting the same thing is 5 — and, on a form it has never seen, far likelier to spend that 5 on nothing.
The tool reference publishes browser_session at a flat 3 credits. That is open's price, and it is a ceiling rather than a rate: the REST route reserves 3 before the request body is readable, then deducts the operation's real cost once the call succeeds. You are never charged more than the published number, and usually less. Credits come from the same pool as every other tool — see plans and credit packs.
Limits, Stated Plainly
A session holds real infrastructure, so the limits are real too:
- It expires.
ttldefaults to 600 seconds fromopen(range 30-3600) andactivity_ttlto 300 seconds since last use (range 10-3600), whichever comes first. Close sessions when you are done rather than letting them age out. - One at a time on the hosted API. A REST API key may hold one session; stdio and self-hosted installs keep the default of three. That is arithmetic, not caution — one customer holding three sessions would occupy the entire hosted session capacity.
- No
executeJavaScripton the hosted surface. Arbitrary script in a browser on our infrastructure is a different act from the same script on your laptop, so it is refused over a remote transport. Useclick,type,selectandpress, or run the MCP server locally over stdio. - Every navigation is re-gated. A long-lived session is a repeatable navigation primitive, so each in-session
navigatere-runs the SSRF guard, the host blocklist and the robots.txt check. Opening a session does not buy an unchecked hop. - Logins do not persist between sessions. A login holds inside one session. There are no saved profiles carried from one session to the next.
One Session From the CLI
The CLI runs one complete session per invocation — open, snapshot, your steps, read, close — because a session lives in the process that opened it and a CLI process ends when the command does:
# steps.json: [{"operation":"act","actions":[{"type":"click","selector":"@e2"}]}]
crawlforge browser https://news.ycombinator.com \
--steps steps.json \
--read --format markdownWhen a session has to outlive the call that opened it, use the MCP tool or the REST API. There is no crawlforge browser open that hands back an id — the page behind that id would die on exit.
When to Use Which
scrape_with_actions | browser_session | |
|---|---|---|
| Browser lifetime | One call | Across calls, until the TTL |
| Selectors | Named up front, unseen | Refs from a snapshot you read |
| Price | 5 flat | 3 to open, then 1-2 each |
| A wrong selector costs | The whole 5-credit call | The 1 credit of that act |
| Login | Repeated every call | Paid once, held by the session |
| Best for | A flow you can write down in advance | A flow you have to see to navigate |
Neither replaces scrape (2 credits), which is still the right answer for a page that renders without being touched. And browser_session is best for teams whose targets are applications rather than documents: dashboards behind a login, multi-step wizards, filtered search results that only exist after four clicks.
How to Upgrade
npm install -g crawlforge-mcp-server@latest
crawlforge --version # 6.6.0If your MCP client launches the server with npx, it picks up v6.6.0 on the next restart. No schema, output-shape or credit change to any existing tool. The full release history is on the changelog.
Want an agent that looks before it clicks? Start free with 1,000 credits — 125 full login-and-read sessions — and read the browser_