The Problem
Training and fine-tuning AI models requires large, clean datasets from diverse web sources. Collecting this data manually is impractical, and raw HTML is too noisy for model training.
Use Cases
01Quick Answer
Use CrawlForge batch_scrape (5 credits) to fetch hundreds of URLs in parallel, then extract_content (2 credits) to return clean, boilerplate-free text or markdown ready for a training pipeline. You collect structured content instead of noisy raw HTML, which improves dataset quality and lowers preprocessing cost -- about 7 credits per document.
02The brief
Training and fine-tuning AI models requires large, clean datasets from diverse web sources. Collecting this data manually is impractical, and raw HTML is too noisy for model training.
CrawlForge batch_scrape processes hundreds of URLs in parallel for scale, while extract_content returns clean, structured text ready for training pipelines. Build datasets from any web source.
03In code
1// Collect training data from documentation sites2const batch = await mcp.batch_scrape({3 urls: [4 "https://docs.example.com/guide/intro",5 "https://docs.example.com/guide/setup",6 "https://docs.example.com/guide/advanced",7 // ... hundreds more URLs8 ],9 format: "markdown",10});11 12// Extract clean content for each page13const dataset = await Promise.all(14 batch.results.map(page =>15 mcp.extract_content({16 url: page.url,17 format: "text",18 remove_navigation: true,19 })20 )21);22 23console.log(`Collected ${dataset.length} documents`);04The pipeline
05Questions
04
Use CrawlForge batch_scrape to fetch hundreds of URLs in parallel, then extract_content to return clean, boilerplate-free text ready for a training pipeline. You get structured content instead of noisy raw HTML.
Raw HTML is full of navigation, ads, and markup that adds noise and wastes tokens. extract_content uses a readability pass to return only the main content as clean text or markdown, which improves dataset quality and lowers preprocessing cost.
Yes. batch_scrape at 5 credits per batch parallelizes fetching across hundreds of URLs, and extract_content at 2 credits cleans each one. Combine with map_site to enumerate a source first, then batch the resulting URLs.
CrawlForge honors robots directives and you control which sources you crawl. You are responsible for the rights to data you collect, so target sites you are permitted to use for training and keep your crawl scope deliberate.
06Keep exploring
Feed your AI agents live web data with structured extraction and multi-source research. Build pipelines with 31 MCP tools, from fetch_url to deep_research.
Extract and restructure content from legacy sites for migration to modern platforms. Keeps formatting, media, and internal links intact as clean Markdown.
Start forging
Every new account gets 1,000 free credits. No credit card required.