The Problem
AI answer engines now send real traffic, but they see a different site than Google does. Without a curated map of your important pages and usage terms, crawlers guess -- and summarize the wrong pages, or skip you entirely.
Use Cases
01Quick Answer
Run generate_llms_txt (5 credits) against your domain to produce llms.txt and llms-full.txt -- a curated map of your key pages plus usage guidelines for AI crawlers -- then map_site (2 credits) to confirm what those crawlers can reach. About 7 credits per site, and you re-run it whenever the site changes.
02The brief
AI answer engines now send real traffic, but they see a different site than Google does. Without a curated map of your important pages and usage terms, crawlers guess -- and summarize the wrong pages, or skip you entirely.
CrawlForge generate_llms_txt crawls your site and writes both llms.txt and the longer llms-full.txt, with your organization name, contact, and custom guidelines. map_site then shows what an AI crawler can actually reach before you publish.
03In code
1// Generate llms.txt and llms-full.txt for your site2const llms = await mcp.generate_llms_txt({3 url: "https://example.com",4 format: "both",5 complianceLevel: "standard",6 analysisOptions: { maxPages: 200, maxDepth: 3, respectRobots: true },7 outputOptions: {8 organizationName: "Example Inc",9 contactEmail: "ai@example.com",10 customGuidelines: ["Cite the canonical URL", "Pricing changes monthly"],11 },12});13 14// Confirm what an AI crawler can actually reach15const map = await mcp.map_site({16 url: "https://example.com",17 include_sitemap: true,18 max_urls: 500,19});20 21console.log(llms.llmsTxt);22console.log(`Discoverable pages: ${map.urls.length}`);04The pipeline
05Questions
04
llms.txt is a plain-text file at your site root that points AI crawlers at your important pages and states how the content may be used. It is a proposed standard rather than a ranking guarantee, but it is cheap to publish and replaces guesswork with a curated map.
Run generate_llms_txt against your domain. It crawls up to 500 pages, respects robots directives, and writes llms.txt plus the longer llms-full.txt, including your organization name, contact email, and any custom guidelines you pass.
llms.txt is the short index — sections and key links. llms-full.txt inlines far more detail for models that can take it. Pass format "both" to get the pair, or pick one with "llms-txt" or "llms-full-txt".
Run map_site with include_sitemap to list every reachable URL, then compare that against the pages listed in your llms.txt. Gaps show where a crawler will miss content, or where your sitemap has gone stale.
06Keep exploring
Audit your site and competitors for metadata, broken links, content gaps, and ranking opportunities. Returns a structured report across every crawled page.
Turn documentation sites into clean, chunk-ready markdown for retrieval-augmented generation. Crawl, strip the boilerplate, and summarize before you embed.
Start forging
Every new account gets 1,000 free credits. No credit card required.