Browser Actions Wait Out Crawl-delay, and Local File Paths Stay Local
6.20.0 makes the browser tools as polite as the fetch tools and closes two quiet gaps. A click, key press, select or checkbox in scrape_with_actions and browser_session now waits out the robots.txt Crawl-delay of the page it is on before it runs, as a navigate already did, and requests that reach one host at the same moment queue a full delay apart. Upgrade note: scrape now takes at most one highlights and one question format per call. A second one is refused and charged nothing, where it used to replace the first.
scrape_with_actions and browser_session: a click, press, select or check waits the Crawl-delay of the page's host before it runs, because the browser sends any request it causes before anything can ask. A navigate and the first load were already spaced
Requests that arrive together for a host with a Crawl-delay now queue a full delay apart on both the MCP server and the REST API. Before, every request that came in during one wait went out at the same moment, and map_site and generate_llms_txt now space their sitemap fetches too
process_document reads local files, and its usage report sent their full path to crawlforge.dev. It now reports only the file type, as [local file].pdf, the website masks paths from older servers when they arrive, and the stored rows were cleaned
scrape: a second highlights or question format in one call is a validation error, charged nothing, instead of silently overwriting the first while the +1 credit add-on was still charged
v6.19.2
Fix
Every Defect From the October Live Test Fixed or Decided
6.19.0 finishes the live test of all 31 tools run on 2026-10-03. Parameters that did nothing now work or are gone, text tools no longer weld paragraphs together, an error names the next step for what actually went wrong, and eight more tools take max_inline_chars. Upgrade note: stealth_mode enable and disable and scrape_with_actions captureScreenshots are removed, browser_session read returns the whole page, and a few oversized fields were trimmed from extract_content, search_web, map_site and scrape_with_actions.
stealth_mode on Chromium had stopped Cloudflare's own challenge from loading: three forced headers made every cross-origin request send a CORS preflight that challenges.cloudflare.com refuses. The challenge now loads, and the page is reported as blocked by cloudflare. 6.19.1 drops the same kind of forced header from localized browser contexts
extract_text tables are a uniform grid with absolute image and link URLs on both the MCP server and the REST API, shared through crawlforge-extractors 1.17.0, and REST scrape markdown keeps its tables
reddit_search comment_count counts every comment returned and comments_collapsed the ones left out; serp_rank documents that a target matches by host and that one lookup is one sample, and the REST API returns se_results_count and check_url. 6.19.2 sets its default depth to 30 on both
Errors carry a next step for the actual failure (robots, SSRF, 404, timeout or a bad parameter), and analyze_content has stop words for eleven more languages
v6.18.1
Feature
extract_text and extract_links Reach Walled Pages and Match the REST API
The two cheapest page readers get scrape's fetch ladder. With escalate: true a walled page (a named vendor, 403, 429 or 444) is re-read through the same stealth stage scrape uses, projected at 6 credits and charged 1 when the plain fetch got the page; a 404 or 5xx never escalates. Without it, a failure now says what the target answered and which vendor walled it. A PDF or other binary fails with a pointer to process_document instead of returning garbage, a JSON body comes back with a note, the fetch timeout is 15 s with one retry after a short Retry-After, and an empty client-rendered shell is flagged with rendered: false. Both tools now return what the REST API returns: extract_text reads one line per block element and takes selector and max_length, and extract_links returns one record shape on both surfaces. Upgrade note: extract_links records no longer carry is_external; read type instead. 6.18.1 fixes what a live test of all 31 tools found: crawl_deep no longer crashes on a refused seed and stops when max_pages is reached, map_site reports a seed it could not read as an error, batch_scrape refuses a PDF instead of returning its bytes as text, and scrape names Fastly's client challenge page as a wall. Upgrade note: crawl_deep pages_crawled now counts fetched pages only, and track_changes monitoringOptions.enabled: false no longer starts a monitor.
The same URL through the MCP server and the REST API returns the same links: 327, 196 and 123 links, first 20 identical, on Wikipedia, Hacker News and BBC News captures
Link records are { href, text, type, domain, rel, original_href }: mailto:, tel: and javascript: links are type "other", a <base href> is honoured, and duplicates are matched without fragment or trailing slash
scrape_with_actions' browserOptions.timeout now bounds the page load, and an over-cap scrape_with_actions or browser_session result keeps its wall report inline
6.18.1: crawl_deep with max_pages: 3 on a link-heavy site took 232 s and processed 2,290 queued links; it now stops when the limit is reached. formAutoFill checks radios and checkboxes and selects options, browser_session act returns a screenshot as a resource URI instead of inline base64, and agent runs no web search for a question scoped to its seed URLs
v6.17.1
Feature
Every Wall a Chain Meets Is Named, and Your Own Proxies on Action Chains
scrape_with_actions now checks every navigation it makes, not only the final page: the initial load and each navigate action are listed under navigations with their HTTP status and, on a wall, the vendor. AWS WAF's challenge interstitial, the 202 page Amazon answers a headless browser with, is now named aws-waf on every surface. A stealth chain can go out through the caller's own residential proxies with browserOptions.proxyRotation (Chromium only; CrawlForge supplies no proxies), with credentials removed from results and usage reports. A per-call proxy is refused on Camoufox, which shares one browser across calls and so would have carried other callers' traffic. 6.17.1 also makes the stealth Cloudflare wait recognise Cloudflare's current "Performing security verification" page, and caps that wait at 10 s as intended: a walled doordash.com chain went from 39.5 s to 19.5 s. Nothing here passes a challenge.
navigations: [{ url, finalUrl, httpStatus, blocked? }] on every chain, and httpStatus plus blocked on each navigate result; top-level success and blocked are still decided by the final page
A caller's proxy whose host is loopback, link-local or a cloud metadata address is refused before a browser launches
The aws-waf rule lives in crawlforge-extractors 1.12.0, so the MCP server and the REST API name it the same way
v6.16.0
Feature
extract_embedded_state Decodes Nuxt 3, SvelteKit, YouTube and Shopify, and Redirects Are Gated
extract_embedded_state now decodes what the plain fetch already brought back, with no browser: Nuxt 3's __NUXT_DATA__ as objects instead of an index array, SvelteKit's page data, ytInitialData and ytInitialPlayerResponse (a YouTube watch page returns its videoDetails), Remix, Inertia's data-page, Shopify's ShopifyAnalytics.meta and product JSON, and large JSON data-* attributes. A global assigned a JavaScript literal rather than JSON is read by a static evaluator, so no page code runs. Next.js Flight rows have their references to other rows resolved, a data_rows index points at the rows that carry data rather than markup, and find returns every path whose key matches, ready to pass back as path. The same release gates every redirect: each hop is asked about the blocklist and robots.txt before it is followed, waits out the host's Crawl-delay, and is signed afresh for Web Bot Auth. The price stays 2.
find returns up to 50 matches as { path, preview } with the value's first 200 characters; raw: true also keeps the undecoded Nuxt array. Both work on the REST API too
A redirect into a path robots.txt disallows, or onto a blocked host, is refused before it is requested, and a call it sinks is not charged. Browser tools check every hop and the URL a page lands on, and a page that scrape_with_actions or browser_session keeps is checked again after every action
Audit-log rows no longer keep credentials from a URL: userinfo is dropped and query parameters named like a key or token are redacted
v6.15.0
Feature
extract_embedded_state: Previews for Big Payloads, and a Stealth Re-read
A result over max_inline_chars now comes back as a preview and a result_handle instead of the whole payload: producthunt.com and zappos.com had returned 1-2 MB inline. The inline part keeps the found list, the byte size, the top-level keys of data and a pretty-printed preview cut at a line boundary, never inside a value, and read_result reads any part of the full payload by json_path. keys_only: true returns two levels of keys with each value replaced by its type, for cheap discovery. With escalate: true, a walled page (403, 429, a challenge page or an empty shell) is re-read through the same stealth stage scrape uses and parsed again, and in the browser the framework globals are read off window after JavaScript ran, under window_state. Without escalate, a wall is now an error that names the vendor rather than a success with nothing found. The price stays 2; escalate projects 7 and charges 2 when the plain fetch gets the page.
Verified at max_inline_chars: 3000: producthunt.com 2,979 and zappos.com 2,970 characters inline, the rest one read_result call away
doordash.com (Cloudflare) returned 103-106 Next.js Flight rows through escalate on impit, Camoufox and Chromium
window_state reads __NEXT_DATA__, __NUXT__, __remixContext, ytInitialData, ytInitialPlayerResponse, the Redux and Apollo stores, __TGT_DATA__ and __PWS_DATA__; on Camoufox the read goes through Firefox's isolated world, so no script runs in the page's own
v6.14.0
Feature
Snapshots Inside Shadow DOM and Iframes, and Cookie Banners Answered
Snapshots now use Playwright's native accessibility snapshot, so elements inside open shadow roots and iframes, cross-origin ones included, appear in the tree with ordinary @e refs that act can click. The tree format is unchanged, the old page walk stays as a fallback, and each snapshot says which ran. scrape_with_actions also takes browserOptions.consent: "reject" runs DuckDuckGo autoconsent's opt-out once after the first navigation and after each navigate action, "accept" its opt-in, and the default "off" leaves banners alone so chains that click them themselves keep working. Prices are unchanged.
source: "aria" | "walk" on every snapshot; browser_session snapshots get the same tree, and on-screen aria-hidden controls are now listed
consent: { cmp, action, ms } on the result, capped at 2 s per navigation; a page with no matching banner reports action: "none" and never fails the chain. Verified on the Guardian's Sourcepoint dialog and on Ecosia under Camoufox
The stealth scroll and reading-time helpers take Playwright selectors, so refs inside shadow roots and frames work on the stealth path too
v6.13.0
Improvement
Snapshots That Wait for the Page, and a Queue Instead of a Refusal
Second round of scrape_with_actions fixes. A snapshot now waits for the page to render before reading it: the load event, then a quiet window with no DOM change for 300 ms and no request in flight for 500 ms, capped and never past the action's timeout. The Amazon homepage snapshot went from 0 nodes to 236, search box included. Clicks, key presses and selects wait the same way afterwards, so a single-page app's route change is seen. A burst of calls over the browser limit now queues for up to 60 s instead of failing with "Maximum concurrent sessions", and browser_session shares the same slots only while it opens a page. Stealth clicks on Camoufox no longer stall: two cursor humanisers had been stacked, and the Ecosia consent click went from 12-20 s to 1.4 s. Prices are unchanged.
Snapshot results report waited_ms and settled_by (quiet, cap or closed); browser_session snapshots get the same wait
CRAWLFORGE_MAX_ACTION_SESSIONS (default 3) browser slots and CRAWLFORGE_ACTION_QUEUE_TIMEOUT_MS (default 60 s) of queueing; results carry queued_ms, and executionTime no longer counts the wait
On Camoufox the click simulator makes one move and lets Camoufox draw the path, with its per-move cap at 0.5 s; simulateClick now honours the action's timeout instead of Playwright's 30 s default
v6.12.1
Fix
App-Error Pages Caught, and Cheaper scrape_with_actions Failures
Follow-ups to 6.11.0's impit step. A short page whose text opens with an app error — "Something went wrong", "Oops!", "An error occurred" or Next.js's client-side exception — is now treated as a soft block even under a normal title, so it goes on to the stealth browser instead of coming back as the page. quora.com's "Something went wrong. Wait a moment and try again." had passed on every path, not only through impit. The rule comes from crawlforge-extractors 1.10.0, which the server now requires. A new CRAWLFORGE_IMPIT=off setting turns the impit step off for a whole deployment: measured from the hosted instance's datacenter IP, it cleared none of the benchmark's walls, so the hosted service runs with it off while npm installs keep it on. Prices are unchanged. 6.12.1 is the first round of scrape_with_actions fixes from a live review: maxRetries now defaults to 0, so a failing chain no longer runs twice (a click on a missing selector went from 20.8 s to about 4 s), and every run a caller does ask for is reported under attempts[]. browserOptions.timeout now reaches every action, screenshotOnError: false is honoured, and screenshot bytes no longer leak inline beside their resource URI.
A document under 200 characters whose text opens with an app error phrase is a soft block whatever its title, so an error fallback is retried in the browser instead of returned as content
CRAWLFORGE_IMPIT=off skips the escalation stage's impit step for a deployment whose exit IP it cannot help, and is read on every call. The hosted service sets it off; npm installs keep it on, where it cleared indeed.com from a residential IP
The REST API's scrape escalation stays browser-only by decision, and replaying a Cloudflare clearance cookie through impit was declined: the cookie is bound to the User-Agent it was earned with, so replaying it would mean sending a browser User-Agent instead of CrawlForge's own
scrape_with_actions: maxRetries defaults to 0 with each run reported in attempts[], browserOptions.timeout is the deadline for every action, error recovery is capped at 3 s per action and skipped for a wait or a selector that matches nothing, and a chain with one screenshot returns under 5 KB instead of about 118 KB
v6.11.0
Feature
A Chrome TLS Try Before the Browser
When scrape runs with escalate: true, or the agent retries a walled page, and the engine is "auto" (the default), a Chrome TLS handshake is now tried before any browser launches. It carries the honest CrawlForge User-Agent. If it gets the page, no browser starts: indeed.com came back in about a second from a residential IP where the plain fetch is blocked by Cloudflare. If it does not, the browser runs as before, and the price is 2 + 5 either way. The step uses impit, a new optional dependency with prebuilt binaries; without it, escalation behaves exactly as in 6.10.0. On Chromium, a stealth render still held by a Cloudflare wall with a Turnstile frame now clicks the checkbox once, a 200 page that only embeds a Turnstile widget is no longer called a block, and every escalation writes a row to the compliance audit log.
impit runs only inside the escalation stage: never on the plain fetch, and never when a caller names playwright or camoufox. A page it gets reports stealth.engine: "impit" with a warning naming it, every redirect hop is SSRF-checked with DNS resolution, and a page with under 200 characters of visible text goes on to the browser
Turnstile checkbox click on Chromium through plain Playwright mouse input, only when a wall is still up after the usual wait and the page has a challenges.cloudflare.com frame. Verified against Cloudflare's forced-interactive test sitekey, which proves the mechanism only: whether a real site accepts the click still depends on the IP and fingerprint it scores
nowsecure.nl answers 200 with a Turnstile widget on Cloudflare's test sitekey, and every "Blocked" recorded for it came from our own verdict layer. A 200 document with a title, text and only widget markers now passes; real walls (403) and empty shells stay blocked
A stealth_escalation row goes to logs/compliance-audit.log for scrape with escalate, the agent's stealth retry and stealth_mode, written after the compliance gate and before the browser navigates, with hashed key and owner ids rather than raw values. robots_override rows no longer read apiKeyId: "anonymous" on a live server
v6.10.0
Feature
Clearance Jar: A Solved Challenge Stays Solved
When a stealth render gets past a Cloudflare or DataDome wall, the vendor's clearance cookies — exactly cf_clearance, __cf_bm and datadome — are kept and replayed to the next stealth context with the same engine, User-Agent and proxy. That covers the scrape escalation stage, stealth_mode, the agent's stealth retry, browser_session and scrape_with_actions. On stackoverflow.com from a residential IP, the first stealth call went through Cloudflare's interstitial in 4.1 seconds; the second went straight to 200 with no challenge in 1.6 seconds. No other cookie is ever kept, so one caller's login cannot reach another caller's context.
Only three named cookies are stored, keyed by engine, User-Agent and proxy. A render that meets the wall again drops that host's clearances, expiry is the cookie's own capped at 24 hours, and the jar holds at most 32 identities with 200 cookies each
Persisted at ~/.crawlforge/stealth-clearance.json with mode 0600 and hashed keys, so a second process reuses a clearance from disk: indeed.com loaded with zero challenge-platform requests. CRAWLFORGE_CLEARANCE_JAR=off disables it
A persistent Chromium profile pool was deliberately not built: a shared profile keeps logins and site storage, and on the hosted instance that would leak between customers
v6.9.0
Feature
agent Retries Walled Pages in the Stealth Browser
Before 6.9.0 the agent tool's act stage ran only the plain fetch. A challenged seed page was dropped and the answer was built from search snippets, which in the review's Indeed test produced "3.3 stars, based on 3.3 reviews". Now a page that hits a challenge, a 403 or 429, an empty shell or a timeout is retried through the same stealth stage scrape's escalate: true uses, with the same compliance gate, engine resolver and proxies. With the Chromium engine the same test read the seed page itself and answered 3.3 stars from 58,942 reviews. Pricing follows what got through: 8 credits, plus 5 for each stealth retry that gets the page.
URLs the caller named are retried first, and a discovered URL only when it has no relevant search snippet. Refusals, 404s and 5xx errors are never retried, a run makes at most 2 retries, and none starts with under 20 seconds of wall clock left
A retry that meets the wall again, or throws, is free. The result reports stealth_retries (attempts) and stealth_retries_charged (billed), evidence from a retry is marked via: "stealth", and the published ceiling is 18 credits, 8 when no retry runs. The REST route charges the same 8 + 5 per charged retry
Stated as measured: on the Camoufox engine from the same IP, Cloudflare blocked the retry, and the unchanged scrape escalation fails the same way there
v6.8.0
Feature
A Stealth Benchmark, and the Leaks It Measured Closed
The 2026-09 stealth review's benchmark had been run by hand; 6.8.0 makes it a command, npm run bench:stealth, which drives the stealth browser and the plain fetch against twelve bot walls and five detector pages and heads its matrix with the host OS, exit-IP class, engine and browser versions. Measured against it, the leaks were closed: navigator.webdriver reads false instead of being deleted and nothing is an own property of navigator, userAgentData drops its HeadlessChrome brand, the Chrome major comes from the installed binary, and a Web Worker now answers what the document answers. Detector self-probes went to 19 pass, 0 fail, 1 skip on both engines. Camoufox, the engine that passes the most walls, is now the default through engine "auto", with a fallback to Chromium that is reported rather than silent.
npm run bench:stealth:ci runs ten in-page self-probes that need no third-party site and fails only on a regression not listed in the baseline file, with a --self-check negative control that forces navigator.webdriver to true and must be caught
A page that only embeds a Turnstile widget is no longer called a Cloudflare block (quora.com now passes on both engines), and a custom-titled interstitial gets its wait-out. Stealth pages no longer intercept requests or randomly drop images, fonts and stylesheets, and --disable-web-security is gone from the stealth browser
scrape's escalate_engine and stealth_mode's engine default to "auto": Camoufox first, Chromium with a warning when its binary is absent. browser_session and scrape_with_actions gain an engine parameter, CRAWLFORGE_STEALTH_PROXIES gives every stealth path a server-level proxy list, and a Camoufox User-Agent whose two version tokens disagreed on about half of launches is fixed
v6.7.0
Fix
Proxy Rotation That Routes, and Camoufox That Stays Firefox
A bot-detection bench run against 6.6.2 recommended routing through residential proxies to get past Cloudflare's IP reputation check. That configuration did nothing, in three separate ways: the proxy went onto a Chromium flag that cannot carry the user:pass every residential proxy is issued with, Camoufox's launch path ran unproxied whatever was asked, and the proxy was read once at launch, so rotation could never elapse. All three are fixed, and stealthConfig.proxyRotation now takes ordinary proxy URLs with credentials applied per browser context. Finding it turned up more of the stealth path working against itself: Camoufox, chosen for having nothing to detect, was being handed a Chrome User-Agent and Chromium's client hints.
proxyRotation accepts http, https, socks4 and socks5 URLs, authenticates on both engines, treats a malformed entry as an error rather than an unproxied request, starts at the first entry, and no longer returns the proxy password from get_stats. CrawlForge supplies no proxies; this routes through yours
Camoufox presents its own Firefox identity, with no sec-ch-ua headers and no Chromium-shaped scripts injected, and runs with its own geoip, block_webrtc and humanize features on. Behind a proxy it takes locale, timezone and location from the exit IP
stealth_mode's scrape waits for the page's load event and for the DOM to stop changing before it reads, reporting waited_for_render_ms, and a Firefox page's own uncaught JavaScript error no longer ends the process
v6.6.2
Feature
browser_session — the 31st Tool
A browser page that outlives the call that opened it. browser_session is one tool with an operation: open launches a session on a URL and returns a session id, snapshot returns the page's accessibility tree with a stable ref on every interactive element (@e1 [textbox] "Email", @e3 [button] "Sign in"), act runs 1-20 browser actions whose selector may be one of those refs, read extracts markdown, HTML, text or metadata from the live DOM as it stands after everything the session has clicked, screenshot captures the page or one element, close ends the session and list names the one an API key is holding. The loop is look, then act: the agent reads the refs a snapshot handed it instead of guessing CSS selectors, and the login is paid for once instead of once per call — scrape_with_actions cannot do that, because its browser closes when the call returns. Operations are priced individually from 1 credit: open 3, read 2, and snapshot, act, screenshot, close and list 1 each, with open's 3 the published rate. Sessions are short-lived by design — ttl (default 600s, range 30-3600) is the absolute clock, activity_ttl (default 300s, range 10-3600) the idle one, and whichever fires first closes the page. The same look-first move reaches one-shot chains too: scrape_with_actions takes a snapshot action, so a chain can list what the page offers before it names a selector. 6.6.1 follows with the first live round run against the tool itself: a blocked page is now named as one instead of being returned as content, executeJavaScript hands back its value, and read is capped like every other content tool. 6.6.2 closes a second live round: a full-page screenshot over the stdio message limit had closed the whole MCP connection, and is now refused with its size and what to take instead.
New browser_session (1-3 credits per operation) keeps one browser page alive across API calls: open, snapshot, act, read, screenshot, close and list, with the session id from open carried on every later call. Each operation is billed on its own price — open 3, read 2, the rest 1 each — so a cheap look never pays for an expensive open, and a login-then-read flow (open, snapshot, act, act, read, close) comes to 9 credits
snapshot returns the page as an accessibility tree with a stable ref on every interactive element, and act points selector at @e1 instead of a guessed CSS selector. A ref belongs to the document its snapshot was taken from: any navigation invalidates every ref, and acting on a stale one fails with an error telling you to snapshot again rather than silently clicking the wrong thing. scrape_with_actions gains the same snapshot action, so one-shot chains can look before they act too
Live on the REST API at POST /api/v1/tools/browser_session under the same envelope as every other tool. A REST API key may hold one session at a time, and a second open while one is still live is refused with a named error rather than queued, so close the session when you are done instead of letting a TTL run out
6.6.1 fixes the failure that mattered most: a bot wall or an HTTP error page no longer reads as a successful session. open and read report httpStatus and name the vendor that blocked them — a DataDome 403 previously came back as success with the single word "g2.com" as the page — executeJavaScript returns its value without being asked to, and read is capped at max_inline_chars with a result_handle instead of returning half a million characters inline
6.6.2 budgets screenshot resources below the stdio message limit — an 18.5 MB full-page PNG had closed the whole MCP session, taking every tool on the connection with it — and respect_robots: false now returns its warning from open and act, audited against browser_session rather than scrape_with_actions
v6.5.0
Fix
Cross-Vertical Sweep: DOCX Support and Seven Fixes
Six defects and four gaps from a live sweep of all 30 tools across roughly 600 URLs in thirteen verticals: real estate, healthcare, finance, education, sports, food, entertainment, law and government, industrial supply, telecom and SaaS pricing, non-profits, outdoors and science. process_document now reads Word documents: a .docx fetched as a URL had come back as a megabyte of ZIP bytes reported as page text; the fetched body now decides how it is read, a PDF served under a plain URL reaches the PDF parser, and an archive or image is refused by name. scrape returned irs.gov's tax-bracket page as its header because the hidden-content pass read a print stylesheet as screen CSS, folded inline styles into a universal rule, and split a :has() selector list on every comma; all three are fixed and the page returns its four rate tables. track_changes scored a feed that gained a whole record as no change, extract_structured failed a table extraction outright when the model overran its output budget, map_site ignored a seed's path and its search, and a truncated scrape dropped its highlights and question answers. As published, 6.5.0 also refuses any reddit.com URL before a fetch, since reddit.com blocks every non-browser client, and names the reddit_search call that reads the same posts; reddit_search now tries Arctic Shift first and PullPush second.
process_document reads DOCX with mammoth and routes a PDF served under sourceType url to the PDF parser; a body it cannot read is refused by name instead of being returned as bytes
scrape's hidden-content pass skips print stylesheets, no longer reads inline styles as universal rules, and splits selector lists on top-level commas only, so irs.gov's tax tables come back; a truncated result keeps its highlights, question and json answers inline
track_changes registers a text-only document that grew by a record; extract_structured keeps the complete rows of a cut-off response as partial: true and names the real fallback reason; map_site scopes to the seed's path and ranks by the search terms in each URL; a long robots.txt Crawl-delay is named in the response
reddit.com URLs are refused before any fetch with USE_REDDIT_SEARCH and pointed at the exact reddit_search call for that URL, and reddit_search tries Arctic Shift first and PullPush second
v6.4.0
Fix
Retail, Travel and Aviation Sweep Fixes
Five defects and five gaps from a live sweep of all 30 tools across 415 retail, travel and aviation URLs. As in the previous round, every defect had returned success with something missing rather than an error. map_site read 75 of boeing.com's 1,878 sitemap URLs because the sitemap uses relative loc paths, which the parser threw on; a loc is now resolved against the sitemap's own URL, a bad entry is skipped on its own, and an empty parse is neither cached nor served from the cache. scrape at its default onlyMainContent dropped WestJet's 6-row checked-bag fee table, too small for the data-table size test and inside a component Readability scores out; a dropped table with th header cells is now recovered whatever its size, and a table that opens with an empty corner cell renders as a pipe table instead of flattened text. scrape_with_actions on a JavaScript help centre returned the placeholder Content not available in markdown format as a success; markdown and html now come from the post-action DOM when the extractor has nothing. agent, asked what Southwest charges for a first checked bag, planned only the bare entity query and answered that the fee is not stated; a plan that stops at one query now gets the task's own words as a second, only the entity query votes for the live root, and the answer is $45 cited to southwest.com. crawl_deep previews were the site's mega-menu on every page; pages go through the same main-content pass scrape uses. No tool changes price.
map_site reads sitemaps written with relative loc paths. boeing.com's 1,878-URL sitemap had parsed as empty because normalizeUrl threw on /, so the tool fell back to crawling the home page's links and returned 75. A loc is resolved against the sitemap URL, one bad entry no longer discards the rest, and an empty parse is never written to or served from the sitemap cache, which had kept answering 75 for an hour after the parser was fixed
Tables and JavaScript pages: scrape recovers a dropped table whose author marked th header cells whatever its size, so WestJet's fee table survives the default onlyMainContent, and an empty corner cell in an otherwise all-th first row is promoted so the table renders with its columns. scrape_with_actions serves markdown and html from the post-action DOM, with markdownSource: body, instead of a placeholder when Readability finds no article. crawl_deep content is the page's main content rather than its menu
agent adds the task's own words as a second search query when a current-state plan stops at the bare entity name, keeps the entity query as the only vote for the live root, and may move past the live page when it does not state the answer. get_batch_results takes max_inline_chars and returns a preview plus a result_handle over the limit; extract_links no longer counts a javascript: pseudo-link; reddit_search names the caller's window and how to narrow it when Arctic Shift times out inside it
v6.3.1
Fix
Nested Validation, Version Authority and Linux Login
Four defects from the R19 live sweep, each a confident wrong answer rather than an error, plus a Linux-only crash. extract_structured validated only one level deep, so an array of stray page text passed a schema asking for an array of objects; the validator now lives in one module, every consumer uses it, the CSS-selector fallback included, and success is false when a required field is present but the wrong shape. The same tool fell back to CSS selectors on a large table with no explanation because the output budget counted an array as one field; the budget scales with the schema's array-ness, a truncated response is retried once at twice the budget, and the failure names the token limit — the workable range goes from about 43 rows to about 156. agent answered a version question from a forum post that passed provenance because the string was on a fetched page; for a version question, the version must now appear in a source that is not a discussion page, and what fails is flagged in provenance.unsupported_versions. analyze_content invented entities and a readability score for non-Latin text; both report notApplicable for Cyrillic, Greek, Arabic and Devanagari, and Russian stop words joined the list. 6.3.1: crawlforge login crashed on Linux with ERR_MODULE_NOT_FOUND because one import used the wrong case, which case-insensitive filesystems hid; a new test compares every relative import against git ls-files.
extract_structured validates nested shapes: countries as an array of strings against a schema of objects no longer reports valid. One validator in src/utils/schemaValidate.js serves every consumer, errors name the path (Field countries.0.capital: expected string, got number) and cap at ten plus a count, and success is false when a required field is present but the wrong shape — a visible change for callers that branch on it
extract_structured's output-token budget scales with the schema's array-ness instead of counting an array as one field, a response cut off mid-object is retried once at twice the budget, and the failure says model response was cut off at the N-token output limit rather than reporting a JSON offset; the workable table range goes from roughly 43 rows to roughly 156. agent no longer answers a version question from a discussion page: the version must appear in a non-forum source, one corrective rewrite runs, and anything that survives is flagged in provenance.unsupported_versions; provenance.checked now reports whether the check ran
analyze_content reports notApplicable for entities and readability on non-Latin text instead of inventing them (rain and gusty had come back as an organization, and a Russian weather report scored Very Easy), extending the CJK guard to Cyrillic, Greek, Arabic and Devanagari. 6.3.1 fixes crawlforge login on Linux: one import of AuthManager.js used the wrong case, which macOS and Windows resolve and Linux does not, so the command had died at import for every Linux user since 6.1.0; tests/unit/importCaseExactness.test.js now checks every relative import against git ls-files
v6.2.0
Feature
Hosted Monitors from the MCP Server
Scheduled monitors on the MCP server now come in two kinds. A local monitor, operation create_scheduled_monitor as before, is persisted under ~/.crawlforge/monitors, fires in-process only while the MCP server process is alive, catches up missed runs on restart, and notifies by webhook or Slack from your machine; crawlforge monitor:run-due from system cron remains the guaranteed local path, and a local monitor never sends email. Pass scheduledMonitorOptions with hosted: true and the server instead registers the monitor with the hosted service at /api/v1/monitors using the configured API key and returns the hosted id and a dashboard URL. From then on it is the same hosted monitor the REST API and the dashboard manage: CrawlForge's own scheduler fetches, diffs, records every check and notifies by email and signed webhook with nothing running on your side, the interval is converted to a cron with a 5-minute floor, and the timezone is UTC. No tool changes price: a hosted create or stop costs nothing, a local create still costs 3 credits, and each hosted check bills the track_changes price per compared target.
hosted: true on scheduled monitors. create_scheduled_monitor registers the monitor with the hosted service through /api/v1/monitors using the configured API key and returns the hosted id and a dashboard URL; email goes to notificationOptions.email.recipients and signed webhooks to notificationOptions.webhook. list_scheduled_monitors lists local and hosted monitors together, each carrying a hosted flag, and stop_scheduled_monitor deletes either. A hosted create or stop from the MCP server costs nothing, like the REST monitors API; a local create still costs 3 credits
Local monitors stop pretending to send email. A local monitor with email notification settings now reports an error pointing at hosted: true instead of a fake success; the local email path was always a placeholder. A local monitor notifies by webhook or Slack from your machine, fires only while the MCP server process is running, catches up missed runs on restart, and crawlforge monitor:run-due from system cron is the guaranteed local path
The track_changes input schema is declared once and shared by the server registration and the tool, with no change in behaviour. AlertNotificationSystem.js, 601 lines imported by nothing, is deleted: its live counterpart is the notifier module, and the hosted service sends the email and signed webhooks
v6.1.0
Fix
Confirmation Prompts on Both Protocol Eras
The five tools that ask before an expensive run — crawl_deep over 500 pages, batch_scrape in sync mode over 25 URLs, deep_research over 50 URLs, agent on the pro model, and extract_structured with no LLM configured and more than three required fields — plus the low-credit warning, were built on the 2025-era inline elicitation/create request, and the 2026-07-28 revision has no server-to-client request channel at all: on that era the prompts silently did not fire and the operation proceeded unasked. They are now multi-round-trip input_required returns, which the SDK serves on both eras, so the prompt reaches you either way. A call that comes back asking for confirmation has done no work, so it costs no credits and writes no usage record; the charge is taken once, when the call actually completes, and declining costs nothing. The same release fixes elicitation over HTTP, which had never worked: both HTTP legs serve from a cloned server instance and only a clone is ever connected, so the helper read undefined client capabilities on every HTTP request and every session proceeded unasked — broken since v3.2.0 and failing safe the whole time. The transport now stamps the serving clone on the request context, and stdio was never affected.
Confirmation prompts now reach 2026-07-28-era clients: crawl_deep over 500 pages, batch_scrape in sync mode over 25 URLs, deep_research over 50 URLs, agent on the pro model, and extract_structured with no LLM configured and more than three required fields, plus the low-credit warning, all ask before they run. They were built on the 2025-era inline elicitation/create request, and the 2026-07-28 revision has no server-to-client request channel, so on that era the prompt silently did not fire and the operation proceeded unasked. Each is now a multi-round-trip input_required return, which the SDK serves on both eras — the prompt reaches the user on a 2025-era client and a 2026-era one alike
A confirmation round trip is never billed. A call that returns asking for confirmation has done no work, so it costs zero credits and produces no usage record; the charge is taken once, when the call actually completes. Declining a confirmation costs nothing, and no tool changes price in this release
Elicitation over HTTP is fixed, and it had never worked: both HTTP legs serve from a cloned server instance and only a clone is ever connected, so the helper read undefined client capabilities on every HTTP request and every session proceeded unasked — broken since v3.2.0 and failing safe the whole time, never asking rather than asking wrongly. The transport now stamps the serving clone on the request context. This is the hosted HTTP endpoint only: a stdio connection, which is how npx crawlforge-mcp-server and every desktop client connects, was never affected. Also fixed: the .mcpb bundle smoke test in CI no longer sets a placeholder API key
New `crawlforge login` command: a browser handoff for API keys. The CLI prints an approval URL, the signed-in user approves on the website, and the key is delivered to the CLI once and saved to ~/.crawlforge/config.json. It never edits a client config; a coding agent can run it and relay the URL so the key is never pasted into a terminal
v6.0.0
Feature
The 2026-07-28 Protocol Migration
A major release, because three things a caller can see are removed or changed. The MCP server now speaks the 2026-07-28 protocol revision statelessly on the same /mcp endpoint that already served the 2025 era, with no handshake, no session id and server/discover in place of initialize. Which era a request belongs to is decided by the MCP SDK's own classifier rather than a rule of our own, so the endpoint can never disagree with the SDK about a borderline request, and both eras are served from one tool definition so they cannot drift apart. This is the HTTP endpoint only: a stdio connection, which is how npx crawlforge-mcp-server and every desktop client connect, still negotiates the 2025 era and nothing about it changes. Removed: the task parameter and async task mode on crawl_deep, batch_scrape, deep_research and agent, because the spec deleted the SDK surface they were built on rather than moving it; and the --legacy-http flag, four minor versions after its own warning said it would go. Also fixed: the hosted HTTP endpoint had never applied the protocol-hygiene pass that stdio has always had, so it returned tools in registration order with no icons and no schema dialect.
The 2026-07-28 revision on the HTTP /mcp endpoint, served statelessly alongside the sessionful 2025 era on the same route. A modern request carries the per-request _meta envelope (protocol version, clientInfo, clientCapabilities) plus the Mcp-Method header, and Mcp-Name on a tool call; it gets its own server instance and no session id. Routing, header/body mismatch checks and the 415 on a non-JSON content type are all the SDK's own code rather than reimplementations, and every one of them refuses the request before the tool runs, so a rejected call is never billed. Authentication is unchanged and identical on both eras — Bearer or X-API-Key on every request — and it runs before era routing, so there is no era-shaped hole. GET /health now lists every revision the endpoint serves
Breaking: params.task and async task mode are gone from crawl_deep, batch_scrape, deep_research and agent. They were built on the MCP SDK's experimental tasks API, and SEP-2663 removed that API outright rather than moving it to an extension — on a 2026-era connection an inbound tasks/get answers -32601 even with a handler registered, so there was nothing left to migrate to. The four tools no longer advertise execution.taskSupport and now run synchronously, which is exactly what every caller who never passed task already got. For long work that should not hold a connection open, use batch_scrape's async webhook mode or the result_handle and read_result pattern. Also breaking: CRAWLFORGE_LEGACY_HTTP and --legacy-http are removed along with the v3.1 stateless mode they gated; 2025-era clients are unaffected, because they are served by the sessionful transport that has been the default since v3.2.0
Fixed: the hosted HTTP endpoint now returns the same tools/list a stdio client has always had. The protocol-hygiene wrappers — alphabetical ordering, icons, the JSON Schema 2020-12 dialect, and the cacheable markers on the ten read-only tools — live on the template server's own protocol instance, not in the registration tables each HTTP session's clone copies, so an HTTP session had been served tools in registration order with none of the four while stdio had all of them. This changes what the hosted endpoint returns on the 2025 era: output is reordered and gains icons and a schema dialect on every tool. It is additive metadata plus a reordering — nothing is removed — but a client that pinned the old ordering will see it move. Separately fixed: elicitation prompts reach clients that had silently stopped getting them since the SDK upgrade, because the SDK's convenience method refused a capability shape its own rule counts as valid
02v5.x
Version 5
11
v5.10.0
Feature
Batch Search and PII Redaction
Two parameters, on the MCP server and the hosted REST API alike. search_web now takes queries: between one and ten searches in a single call, each running the full pipeline a single search runs, and the payloads returned as results_by_query — a list in the order they were sent, one entry per query, so a repeated query appears twice and a query that failed carries error in place of its results. The call costs 5 credits for every query that returned results: the batch projection of 5 x the queries sent is a ceiling, and a batch in which none succeeded still returns 200 and costs nothing. A single-query call is byte-identical to what it was before this release: none of queries, count or results_by_query appears in it. redact_pii removes personal data from the text nine tools return, before the result is stored rather than after, so an oversized result read back later with read_result is already redacted. Its default fast mode is regex over text the call has already fetched and adds nothing to the price.
queries (1 to 10 strings) on search_web as an alternative to query: exactly one of the two is required, and sending both or neither is a validation error. Each query runs the full existing pipeline — expansion, cache, dedupe, rank — and the response carries queries, count and results_by_query, whose entries are in input order and each carry their own query plus that search's fields, or query and error when that one search failed. A search_web call costs 5 credits for every query that returned results — a single query costs 5, and a queries batch is projected at 5 x the number of queries sent, which is a ceiling: each query that failed is not charged, so a batch in which none succeeded costs nothing. Both surfaces bill this identically, and a batch in which none succeeded is a 200 carrying the normal envelope — count as sent, one { query, error } entry per query, and credits_used: 0 — not an error. The single-query path is unchanged: one failing search still errors as it always did
redact_pii (boolean, or { entities, replace_style, mode }) on scrape, extract_content, extract_text, batch_scrape, crawl_deep, stealth_mode (operation "scrape"), scrape_with_actions, process_document and search_web. Four regex classes: EMAIL, PHONE, FINANCIAL (card numbers validated with Luhn, IBANs with mod-97) and SECRET (API keys, bearer tokens, labelled passwords), with PERSON and LOCATION model-only. replace_style is tag (<EMAIL>), mask ([REDACTED]) or remove, and SECRET keeps its label and replaces only the value in every style, remove included. Entity names are case-insensitive, and a name the surface cannot redact is rejected with a 400 naming the accepted classes rather than silently dropped — a silent drop would hand back a page still full of what was asked to be removed
The result carries redaction: { entities, count, mode } at result.redaction on the MCP server and inside data on the REST API, deliberately the same access path. It is present whenever redaction was requested, including entities: {} with count: 0 when the page held nothing, so "found nothing" is distinguishable from "the parameter was ignored"; a class with no hits is omitted rather than reported as 0. mode "fast" adds 0 credits. On the MCP server mode "model" adds 3 credits once per call and runs an Ollama NER pass for PERSON and LOCATION only, reporting model_ran and dropping the charge when no model answered; on the REST API it is a 400, unbilled, naming the MCP server — the same treatment the highlights and question formats' model mode already gets. Two known limits, stated rather than hidden: counters derived from the text (content_length, word_count, character_count) describe it before redaction, and address-like fields (url, link, href, canonical_url) are deliberately never redacted
v5.9.0
Feature
Opt-in Escalation on scrape
A page that answers with a bot wall instead of the page no longer costs the caller a second round trip. Pass escalate: true and the plain fetch still runs first; only when the document comes back as a Cloudflare, Amazon, DataDome, PerimeterX, Akamai or Vercel challenge, an empty shell or a short error-titled placeholder does the same call render it once in the stealth browser and derive every requested format from that render. The projection is the ceiling at 7 credits, and a call the plain fetch served pays the base 2. robots.txt is respected on the escalated path too, matched against the same CrawlForge product token, and no bot defence is handled any differently than stealth_mode already handles it. Escalation saves a round trip; it does not promise the page — a wall the stealth render is refused at too comes back as an honest block carrying escalated: true, charged nothing, and the case it handles most reliably is a page that simply needs JavaScript to render.
escalate (boolean, default false) and escalate_engine (playwright or camoufox, default playwright) on scrape, on the MCP server and the hosted REST API alike: the plain fetch always runs first and only a blocked verdict escalates — a named challenge page, an empty shell or a short error-titled placeholder. markdown, html, rawHtml, text, links, metadata and the query-scoped highlights and question formats are all derived from the rendered page, so an escalated response has the shape of a plain one and needs no separate stealth_mode call
What it costs: the base 2 credits, +1 when a query-scoped highlights or question format is present, and +5 only when the stealth stage actually ran. A call with escalate: true projects 7 and is charged 2 when the plain fetch served the page, so the projection is a ceiling and never a floor; a path robots.txt disallows is refused before anything is fetched and charged nothing, and on the hosted REST API the stage returns 503 and charges nothing where the CrawlForge execution backend is not configured; a wall the stealth render is refused at too is returned as a block with escalated: true and charged nothing
data carries escalated whenever escalate: true was passed, plus stealth: { engine, vendor_detected } when the render ran — vendor_detected names the bot-defence vendor the plain fetch hit (cloudflare, amazon, datadome, perimeterx, akamai, vercel) or null for an empty shell or an error placeholder. On the MCP server only, a host that walled a request is remembered for 24 hours, so the next escalate: true call to that host skips the doomed plain fetch and says so in a warning; escalation reuses the existing stealth path, the same compliance gate and the same CrawlForge identity, adding no new evasion
v5.8.0
Feature
Result Handles and read_result
A result that would flood the context now comes back as a preview and a handle instead. Ten large-output tools take max_inline_chars (default 40,000 characters of JSON); over it the response carries preview, result_handle, total_chars, truncated: true and expires_at, and the new read_result tool — the 30th, 1 credit — slices, searches, paginates by line or reads a JSON path from the stored result. Nothing is re-fetched, and on the MCP server nothing leaves your machine: the store is local, with a 1-hour TTL.
max_inline_chars on scrape, fetch_url, extract_content, crawl_deep, batch_scrape, stealth_mode scrape, scrape_with_actions, process_document, deep_research and extract_embedded_state (1,000–10,000,000, default 40,000): over it data keeps its top-level scalar fields plus preview (the first max_inline_chars characters of the text view — the markdown for scrape, the body for fetch_url, the pretty-printed JSON for crawl_deep, batch_scrape and deep_research), result_handle (res_ + UUID), total_chars, view, view_path, truncated: true, expires_at and a warning naming read_result; extract_embedded_state is never truncated but still returns a handle with truncated: false; error responses are never stored
read_result (1 credit, basic): { handle, operation: slice | search | lines | json_path } — slice returns offset, length, text and has_more (default 10,000 characters); search is a case-insensitive literal substring match, never a regex, returning up to max_matches (1–100, default 20) matches with offset, length and 200 characters of context either side, plus total_matches and truncated; lines returns first_line, line_count, total_lines, char_offset, lines and has_more (default 200 lines, cap 5,000); json_path returns path, value and value_chars, or value: null with a preview and a warning to narrow the path when the value exceeds max_inline_chars; an unknown or expired handle is 404 RESULT_NOT_FOUND and is not billed
Where results live: on the MCP server a local store under ~/.crawlforge/results/ with a 1-hour TTL and a 200 MB LRU, nothing uploaded; on the REST API a per-account value with a 1-hour TTL, readable only by the account that created it; batch_scrape's stored results share the same store and get_batch_results is unchanged; the tool count is 30 on both surfaces
v5.7.0
Feature
Query-Scoped Scrapes: highlights and question
scrape gains two formats that carry a query: highlights returns the sentences, table rows and code blocks that match it, and question returns an answer built from them. Both are extractive by default — verbatim page text with an offset into the markdown of the same call, no model in the path — for 1 credit on top of the scrape, unlike summarising fetch tools that paraphrase a page and invent what they cannot find. Model mode is opt-in, labelled and priced separately, with a grounding check on every number and proper noun.
highlights: add { "type": "highlights", "query": "…" } to formats and get up to max_highlights (1–50, default 10) units ranked by BM25 with phrase and heading boosts, each with text, kind (sentence, table_row or code_block), offset, length and score — offset and length index the markdown format of the same call, so markdown.slice(offset, offset + length) === text
question: add { "type": "question", "question": "…" } and get answer.text assembled from the best evidence with grounded: true and up to 5 evidence units; mode: "model" is opt-in and lets a model synthesise the answer (Ollama first, then a server-side OpenAI or Anthropic key, then MCP sampling) or select among the highlight candidates, and a grounding check marks the answer grounded: false with a warning if any number or proper noun in it is missing from the evidence
Pricing and parity: scrape stays 2 credits; a highlights or question format adds 1 credit once per call, mode: "model" adds 3 more, and when no model route exists the server falls back to extractive, warns and bills only the extractive price; the REST API supports both formats in extractive mode and rejects mode: "model" with 400 UNSUPPORTED_FORMAT, charging nothing; a scrape whose markdown exceeds 40,000 characters now suggests highlights
v5.6.11
Improvement
Six Live Regression Rounds, Smarter Tool Selection, Honest Stealth Results
The 5.6 line: everything six full passes over all 29 tools and 18 templates against live sites turned up, a server that steers the model to the right tool once, a stealth browser that reports an error page as an error, and in 5.6.11 a plain scrape that reports a challenge page, an empty shell or an error placeholder as a failure naming the vendor — the same verdict the REST API returns — with usage reports that carry the real client version and typed form text masked before it leaves your machine.
Data tables survive and silent failures are gone: extract_content, process_document and scrape_with_actions re-attach the tables Readability drops; scrape keeps inline links on natively nested CSS; LLM extraction answers null instead of inventing a value; page fetches send no Accept-Language; every fetch path decodes a page by its declared charset, so kakaku.com no longer arrives as mojibake; the agent checks every version, date and count in its answer against the sources and names any it cannot back up; analyze_content segments Hindi, Finnish and Japanese into words
Honest results on every path: stealth_mode reports Cloudflare, DataDome, PerimeterX and Vercel challenges as blocks and an HTTP error page or a titled soft block as a failure carrying its status, waits up to 8 s for a blank document to paint, keeps a Chromium context alive across a camoufox scrape and spoofs hardware inside Web Workers; scrape_with_actions judges the document its chain ended on; and since 5.6.11 plain scrape does too — a Cloudflare, Amazon, DataDome, PerimeterX, Akamai or Vercel wall, whatever its HTTP status, an empty JavaScript shell or a short error page is success: false with blocked.vendor and a Next step naming stealth_mode, on the MCP server and the REST API alike
crawl_deep fetches URLs with their trailing slashes (globalpetrolprices.com answered the stripped form with 404); deep_research searches a 12-word form of a long topic (12 sources where the paragraph found one); localization sets Accept-Language for the configured country; the shopify-product template reads the product page's JSON-LD when the store refuses its JSON endpoint; reddit_search explains an empty Arctic Shift result; usage reports carry the client's real version, and text typed into a form by scrape_with_actions is masked before the usage report leaves your machine
Tool selection: the server's instructions are a decision ladder with credit costs and a never-re-fetch rule, every tool description says when not to use it and what it costs, scrape, search_web and deep_research load at session start, and every error result ends with a Next step naming the tool to try instead of a retry
v5.5.9
Improvement
Extraction Accuracy and Security Hardening
The 5.5 line: a Reddit-wide comment search, localization that runs the search it used to only describe, and a long run of accuracy fixes — including the code examples that had been vanishing silently from documentation pages.
reddit_search searches comments across all of Reddit, and localize_search now runs the search
Every number an LLM returns is verified against the page source, or comes back null with a reason
SSRF guards on every browser and webhook path, and page text fenced before it reaches a model
v5.4.3
Feature
extract_embedded_state — the 29th Tool
Tier 2 extraction: read the JavaScript state a page ships with instead of its rendered HTML. One fetch, exact values, no LLM in the path. The same release fixed nine defects found by running all 29 tools against live sites.
New extract_embedded_state (2 credits) reads __NEXT_DATA__, React Server Component payloads, Nuxt and Apollo state
extract_metadata gains a json_ld_types filter that matches subtypes, not just exact types
Five track_changes operations that had only ever returned an error now work
v5.3.1
Improvement
One Crawler Identity, robots.txt on Every Path
The compliance and identity work: every fetching tool checks robots.txt, a refused request costs nothing, and outbound requests are signed so a site can verify a crawl really came from CrawlForge.
Every fetching tool checks robots.txt, including the two stealth paths that used to walk past it
Outbound requests signed per RFC 9421 (Web Bot Auth), with a published key directory
A compliance refusal is free — it no longer consumes credits
v5.2.9
Feature
Shopify Template and a Shared Extractor Package
A shopify-product template that reads the storefront's own JSON instead of parsing the rendered page, remote Ollama support, and the template extractors split into their own npm package so both CrawlForge surfaces run one implementation.
shopify-product template: exact price, compare-at price and per-variant stock, with no HTML parsing
track_changes scores a real price move instead of how much of the page it occupies
Ollama registered as a real provider, routing to the best model installed
v5.1.0
Feature
reddit_search — the 28th Tool
Real Reddit search, despite reddit.com blocking every direct path CrawlForge can take. The tool never touches reddit.com; it queries the community-run archives instead. It is free to use, and needs no credentials.
reddit_search (2 credits): keyword search over posts and comments, plus full thread trees
Backed by the Arctic Shift and PullPush archives, with a bounded retry when one sheds load
Results normalized to real permalinks and ISO dates, with archive-freshness caveats attached
v5.0.5
Feature
v5.0 — MCP Spec Adoption and a Security Overhaul
The major release: SSRF, OAuth, secrets and billing hardening, 52 correctness fixes including a rewritten crawl_deep, a rebuilt streamable-HTTP transport, and the MCP spec features clients had started to expect.
Structured output (outputSchema) and async tasks on the long-running tools
CRAWLFORGE_TOOLS and CRAWLFORGE_TOOL_GROUPS expose a subset of the tools to cut context bloat
Zero npm-audit vulnerabilities, and the Node floor raised to 20.16
03v4.x
Version 4
03
v4.10.0
Improvement
Clients Steered Toward CrawlForge for Web Work
The server now ships MCP instructions telling any client to prefer CrawlForge tools over its own built-in web search and fetch. It is guidance, not enforcement — an MCP server cannot disable a client's built-in tools.
Server-level MCP instructions reach every client on the next launch, with no reinstall
Four overlapping tool descriptions reinforced where the model reads them, in tools/list
serp_rank returns the full top-10 organic listing alongside the target's positions
v4.9.0
Feature
serp_rank — Real Google Organic Rank
The 27th tool. It reports where a domain actually ranks in Google's organic results for a keyword — the position a search API cannot give — backed by DataForSEO and billed to your own account.
serp_rank (5 credits) returns the best position, the ranking URL, and every position held
Never fabricates a rank: it reports configured:false when DataForSEO is not set up
Billing hardened — a zero-cost call emits no usage event, and an error bills at half
v4.8.1
Feature
Agent Skills, Enforced SSRF, Scheduled Monitoring
Auto-activating Claude Agent Skills, two advertised safety controls that had been silently non-functional made real, and change monitoring that actually schedules.
Seven Claude Agent Skills with trigger-rich descriptions, so they auto-activate
SSRF enforced on the live scraping path, validated on the initial request and every redirect hop
track_changes scheduled monitors persist and rehydrate their baselines on restart