The Problem
A RAG system is only as good as the text you feed it. Raw HTML carries navigation, ads, and markup that inflate token counts and pollute embeddings, and every documentation site is laid out differently, so one parser never fits them all.
Use Cases
01Quick Answer
Use map_site (2 credits) to enumerate a docs site, scrape (2 credits) to return markdown with navigation and ads stripped, and summarize_content (4 credits) to add a short summary to each chunk. That is roughly 6 credits per document, and the markdown goes straight into your chunker and embedding model.
02The brief
A RAG system is only as good as the text you feed it. Raw HTML carries navigation, ads, and markup that inflate token counts and pollute embeddings, and every documentation site is laid out differently, so one parser never fits them all.
CrawlForge scrape returns markdown, metadata, and links from a single call with boilerplate stripped by Readability, and summarize_content condenses long pages into short summaries you can store as chunk metadata for better retrieval.
03In code
1// Turn a documentation site into RAG-ready markdown2const pages = await mcp.map_site({3 url: "https://docs.example.com",4 include_sitemap: true,5 max_urls: 200,6});7 8// One call returns markdown, metadata, and links per page9const doc = await mcp.scrape({10 url: pages.urls[0],11 formats: ["markdown", "metadata", "links"],12 onlyMainContent: true,13});14 15// Condense long pages into chunk-level summaries16const digest = await mcp.summarize_content({17 text: doc.markdown,18 options: { summaryLength: "short", summaryType: "abstractive" },19});20 21console.log(doc.markdown); // chunk + embed this22console.log(digest.summary); // store as chunk metadata04The pipeline
05Questions
04
Use map_site to list every URL, scrape to return each page as markdown with navigation and ads stripped, and summarize_content to attach a short summary to each chunk. The markdown goes straight into your chunker and embedding model.
Raw HTML spends tokens on markup, menus, and ads, and that noise ends up in your embeddings. scrape runs Readability first and returns only the main content, so chunks stay semantically clean and cost less to embed.
Yes. Re-run map_site and scrape on a schedule, and pair them with track_changes to spot the pages that actually changed, so you re-embed only those instead of rebuilding the whole index.
About 6 credits per document — 2 for scrape and 4 for summarize_content — plus 2 for one map_site call per site. A 200-page documentation site is roughly 1,200 credits, and skipping the summaries halves that.
06Keep exploring
Feed your AI agents live web data with structured extraction and multi-source research. Build pipelines with 31 MCP tools, from fetch_url to deep_research.
Collect and structure large-scale web datasets for fine-tuning and training AI models. Crawl entire sites, extract clean text, and export training-ready data.
Start forging
Every new account gets 1,000 free credits. No credit card required.