7m
Shipping an MCP Server Inside Your Product: Five Decisions
Ten tools out of twenty-nine, off by default, version-pinned and swappable — the architecture behind an AI feature that cites pages it actually opened.
Blog
13 ai engineering
◆7m
Ten tools out of twenty-nine, off by default, version-pinned and swappable — the architecture behind an AI feature that cites pages it actually opened.
◆4m
search_web finds the pages; a local Ollama model reads them. CrawlForge picks that model by measured accuracy rather than parameter count — and a 4B model beats a 20B one.
◆4m
Most "self-hosted" search MCP servers only run the process locally. CrawlForge can point search_web at your own SearXNG instance — and here is the part that still isn't self-hosted.
◆12m
Every AI crawler that matters in 2026: 31 bots, their robots.txt tokens, and what blocking each one actually costs you. Copy-paste file included.
◆6m
Reddit is where the unfiltered opinions live, and it is the most agent-hostile mainstream site on the web. Here is how to give your AI agent working Reddit search — posts, comments, and full threads — through one MCP tool.
◆13m
An agent scraper follows a goal, not a selector. What that means, the three ways to build one, working code, the real failure modes, and the credit math.
◆9m
A July 2026 study found 91.8% of audited MCP servers lack authentication. Here is why web-scraping servers leak cloud credentials -- and how to stop it.
◆11m
The best web scraping tools for AI agents in 2026, ranked by agent-readiness: MCP-native tool discovery, typed schemas, and token-efficient output.
◆9m
No API keys, no cloud, no data leaving your machine. Use extract_with_llm with local Ollama to pull structured data from any site.
◆10m
Learn how the Model Context Protocol works, why it matters for AI agents, and how to build MCP servers and clients with architecture diagrams and code.
◆11m
Build a production RAG pipeline that crawls websites, extracts content, chunks text, generates embeddings, and serves retrieval-augmented answers.
◆14m
Technical deep-dive into anti-bot detection systems and how CrawlForge's stealth mode features help you scrape protected websites ethically and effectively.
◆14m
Learn how to collect, clean, and structure web data for AI training. Best practices for ethical scraping, data quality, and LLM fine-tuning preparation.