On this page
Most keyword research in this industry stops at a volume estimate and a difficulty score out of 100. That tells you almost nothing useful, because it collapses two very different questions into one number: how many people search this, and who currently owns the answer.
So we measured the second question directly. We pulled the real Google organic top ten for the keywords a web scraping company would actually want, using live SERP data rather than a search API's own ordering. What came back was not a difficulty gradient. It was two separate markets that happen to share vocabulary.
Table of contents
- What we measured, and what we did not
- The legacy terms are a commercial fortress
- The MCP terms have no commercial defender
- Six versus zero
- Reddit ranks top three on every keyword
- Firecrawl ranks for its own alternative term
- The intent trap hiding in llm scraper
- The directories that rank are carrying stale data
- What we would do with this
- Where this study is weak
- Frequently asked questions
What we measured, and what we did not
Every position in this article is a real Google organic rank, pulled on 23 August 2026 for a United States desktop search, scanning the top ten results.
That distinction matters more than it sounds. A search API returns results in its order, which is not Google's organic order. CrawlForge's own search_web re-ranks what it retrieves using BM25, semantic similarity, authority and freshness signals — useful for research, useless as a ranking measurement. Reading positions off a search API and calling them ranks is one of the most common ways keyword research quietly becomes fiction.
So the domain lists below come from serp_rank, which reports Google's actual rank_group value. Where we cite how many pages exist for a term, that is Google's own rough index estimate — it is a crowding signal, not search volume. We do not have volume data and have not pretended to.
The legacy terms are a commercial fortress
Here is the verified organic top ten for "web scraping api", a term with roughly 33.1 million indexed pages competing for it:
| # | Domain | Page type |
|---|---|---|
| 1 | scraperapi.com | Homepage |
| 2 | firecrawl.dev | Homepage |
| 3 | reddit.com | Forum thread |
| 4 | scrapingbee.com | Homepage |
| 5 | docs.apify.com | Vendor docs |
| 6 | decodo.com | Product page |
| 7 | zenrows.com | Vendor blog |
| 8 | oxylabs.io | Product page |
| 9 | scrapingdog.com | Homepage |
Eight of the nine are vendor-owned properties. Six are a homepage or a product page — the pages companies fight hardest for, backed by years of link acquisition and, in several cases, nine-figure funding.
There is no clever content angle into this. A new blog post does not displace ScraperAPI's homepage. The honest read is that this term is closed, and any plan whose first step is "rank for web scraping api" is not a plan.
The MCP terms have no commercial defender
Now the verified organic top ten for "mcp scraping", roughly 615,000 indexed pages:
| # | Domain | Page type |
|---|---|---|
| 1 | mcpservers.org | Directory listing |
| 2 | reddit.com | Forum thread |
| 3 | github.com/firecrawl/firecrawl-mcp-server | Code repository |
| 4 | brightdata.com | Third-party blog |
| 5 | skyvern.com | Third-party blog |
| 6 | scrapling.readthedocs.io | Docs page |
| 7 | mcpmarket.com | Directory listing |
| 8 | fastcrw.com | Third-party blog |
"web scraping mcp server" (roughly 390,000 pages) returns nearly the same shape at the top: mcpservers.org, then Reddit, then the Firecrawl MCP repository.
Look at what is missing. Not one vendor homepage. Not one product page. Not one pricing page. The two highest positions belong to a directory and a forum thread — pages that exist because someone catalogued the category or complained about it, not because a company decided to own it.
The closest thing to a commercial presence is a GitHub README and a Read the Docs page. Both rank because they are genuinely the best available documentation, not because anyone optimised them.
Six versus zero
That is the whole study in one line.
On "web scraping api", six of nine top-ten results are a vendor's commercial page. On "mcp scraping", that number is zero.
These are the same buyers. Someone evaluating an MCP scraping server is the same developer who would have searched "web scraping api" two years ago, with the same budget and the same problem. The difference is that one query leads to a wall of companies who have been defending their position since 2016, and the other leads to a directory nobody maintains and a Reddit thread from last July.
The MCP terms are smaller — roughly 615,000 competing pages against 33.1 million. But "smaller and undefended" beats "larger and sealed" every time, and the gap will not stay open. Every quarter that passes, the odds rise that a funded competitor notices the same thing.
Reddit ranks top three on every keyword
We checked four keywords. Reddit is in the top three on all four:
| Keyword | Reddit position |
|---|---|
| firecrawl alternative | 1 |
| mcp scraping | 2 |
| web scraping mcp server | 2 |
| web scraping api | 3 |
This is not a quirk of one SERP. For evaluation and comparison queries in this category, Google is systematically preferring a thread of developers arguing over a vendor's own description of itself.
The practical consequence is uncomfortable but clear: your position in these conversations is part of your search presence whether you participate or not. The r/ClaudeAI thread ranking second for "mcp scraping" is titled "What's the best most reliable MCP to let Claude Code scrape a website?" — a buying question, answered by strangers, ranking above every product page in the category.
Firecrawl ranks for its own alternative term
The organic top ten for "firecrawl alternative" contains firecrawl.dev at position 4.
That is deliberate and it works. Rather than cede the term to competitors writing takedown posts, Firecrawl publishes content that ranks for people actively trying to leave. Apify does the same thing at position 9 with a dedicated alternatives page.
Note who is at position 1, though: a Reddit thread from r/LocalLLaMA titled "What is the best scraper tool right now? Firecrawl is great, but I want to explore more options". The rest of the page is Nimbleway, a GitHub topic page, WebCrawlerAPI, Context, Bright Data, and Thunderbit — six companies competing for the traffic of a seventh company's unhappy customers.
Branded comparison terms are the one place in this category where small vendors reliably outrank large ones, because the term is too specific for the giants to bother defending and too commercial for anyone else to ignore.
The intent trap hiding in llm scraper
"llm scraper" looks like an obvious target. Roughly 329,000 pages, clearly on-topic, clearly AI-adjacent.
It is a trap. A large share of that result set is people trying to block scrapers, not buy one:
- A YunoHost forum thread titled "Prevent LLM scrapers/trawlers?"
- An r/sysadmin thread, "Fighting LLM scrapers is getting harder, and I need some advice"
- Akamai's "The Rise of the LLM AI Scrapers: What It Means for Bot Management"
- A Lobsters discussion of JavaScript proof-of-work anti-scraper systems
Same three words, opposite intent, and a chunk of that audience is actively hostile to the product. Ranking there would produce traffic that converts at approximately zero and a bounce rate that teaches Google the page is a poor answer.
Volume and difficulty scores cannot see this. Only reading the actual results can.
The directories that rank are carrying stale data
Two of the eight results for "mcp scraping" are directories: mcpservers.org at 1 and mcpmarket.com at 7.
We are listed in both. Both listings are wrong. One describes CrawlForge as having 18 tools; the current count is 27. Directory listings are the single highest-leverage fix available here, because the page is already ranking — the work is a correction, not a campaign.
If you sell in this category, go and read your own directory entries right now. The pages beating you may already be describing you, badly.
What we would do with this
Five conclusions we are acting on:
- Stop treating the legacy head terms as targets. They are useful for understanding the market and useless as an acquisition channel. Compete there and you spend years losing to homepages.
- Publish the commercial page nobody has written. The MCP terms have no vendor landing page in the top ten. That is not a content gap, it is an unclaimed position.
- Fix the directory listings first. They rank today, they describe you today, and correcting them costs an afternoon.
- Read the SERP before committing to a keyword. "llm scraper" passes every automated filter and fails on contact with the actual results.
- Treat forums as ranking surfaces. Reddit is top three on everything we measured. Being absent from those threads is a ranking decision, just an unintentional one.
Where this study is weak
Four honest limits.
It is a single snapshot from 23 August 2026. SERPs move, and the MCP results — driven by directories and forum threads — will move faster than the legacy ones.
It is United States desktop only. Mobile results and other locales will differ, sometimes a lot.
It is four keywords with verified positions, plus three more where we looked only at which domains appear. That is enough to establish the pattern, not enough to size the opportunity.
And we have no search volume. Everything here describes who owns the answer, not how many people ask the question. Both matter. We have measured one of them.
Reproduce it yourself
Every number here came from two CrawlForge tools: search_web to find what exists, and serp_rank for the real organic positions. The method is written up step by step in How to Run a Keyword Gap Analysis with an MCP Server, and serp_rank is documented in the API reference.
If you want the wider view of this category first, our ranked roundup of MCP scraping servers and the complete guide to MCP web scraping cover the landscape these keywords describe.
Start free with 1,000 credits — a full four-keyword study like this one costs 20 credits.
Try this yourself — no signup needed
Run any of CrawlForge's 28 scraping and extraction tools in the playground, then start free with 1,000 credits.
1,000 free credits • One-time • No credit card required
Tags
About the Author
Stay updated with the latest insights
Get tutorials, product updates, and web scraping tips delivered to your inbox.
No spam. Unsubscribe anytime.