Skip to content

CrawlForge TeamEngineering Team

5 min read

Market Research API for Consultants: Zero to Sourced

A consultant's market research starts over with every engagement. New category, new vendor list, and a client who will ask on slide fourteen where that number came from. Last quarter's spreadsheet does not transfer. The method does.

A market research api for consultants has to work from zero, because on day one there is no list of vendors to point it at, and it has to hand back facts with a source and a date, because "we checked" is not an answer a partner can give. It also has to cost something you can put against one client rather than into overhead.

Where do you start when there is no vendor list yet?

With two calls that return URLs rather than answers. search_web costs 5 credits a query and takes a queries array, so six phrasings of "who sells X to Y" run in one call for 30 credits. The snippets are usually enough to tell a vendor from a listicle. When you want the passages too, deep_research costs 10 credits a run: it searches, fetches up to ten sources and returns the best-matching passages verbatim, with a sources list carrying each URL, its title and whether it was actually fetched or only seen in a snippet. Nothing is synthesised on the hosted API, which is the point. What comes back is a long-list of domains and the sentences that put them on it.

The next step is finding the right page on each domain, and the wrong way is guessing that pricing lives at /pricing. map_site costs 2 credits a domain and reads the site's sitemap, falling back to a bounded crawl when there is none, so you get the URL list and pick out pricing, product, customers and careers from it. Thirty vendors is 60 credits and you never fetch a 404.

How does a fact get into the deck with its source?

One scrape per page. Ask for markdown, metadata and links together: every format comes from the same fetch and the call costs 2 credits however many you request. The response carries a scraped_at timestamp, which is the date column your deliverable needs.

Bash
curl -X POST https://crawlforge.dev/api/v1/tools/scrape \
  -H "X-API-Key: cf_live_NORTHWIND_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://vendor-a.com/pricing",
    "formats": ["markdown", "metadata", "links"]
  }'

The row you write is vendor, page URL, the claim, the sentence it rests on, and scraped_at. Quote the sentence rather than paraphrasing it. A cell that reads "SSO on all plans" is an opinion until the sentence beside it says so, and the due diligence post covers what happens when a model fills that cell instead.

If the same question has to be answered for every vendor, add a question format for 1 credit more and the call returns the five sentences that best match it, each with an offset into the markdown. The RevOps post is about running that across a category; this one is about getting the category in the first place.

How do you keep one client's spend separate from another's?

Create one API key per client and name it after the engagement. An account holds up to ten, which is more concurrent engagements than most practices run. Every call made with that key writes a row to the request log carrying the tool, the target URL, the credits charged and the key it came from, and the log filters by key and exports as CSV.

That does two things nobody asks for until they need it. The engagement's spend becomes a query rather than a reconstruction, and the client's research is separable from everyone else's when they ask what you hold. At close-out, delete the key. Deletion revokes it and keeps the rows, so the record of what was fetched for whom survives the engagement.

The example above puts the client in the key's name for the same reason: the name is what you will be filtering on.

What does one engagement cost?

A 30-vendor landscape, done the way above:

StepCallsCredits
Long-list: six search phrasings, one deep_research run1 + 140
Find the pages: map_site per vendor3060
Read four pages per vendor120240
One question asked of every pricing page3090
Engagement total430

The free tier's one-time 1,000 credits cover two of those. After that a 1K pack is $3 and a 5K pack is $14, bought when an engagement needs them rather than on a monthly plan, which is a separate argument. At the pack rate the whole landscape above costs about a dollar and a half, small enough to carry inside the fee and auditable if a client asks what the line covers.

What won't it do?

  • No market-size figures. It sizes nothing and forecasts nothing. If the slide needs a TAM, that comes from a report or a dataset.
  • No contact or firmographic data. It reads pages you point it at. It is not a company database and does not overlap with one.
  • No client login. One account, keys per client, and the client sees what you send them.
  • robots.txt is respected by default. A disallowed page is refused before anything is fetched and costs nothing, and turning the check off is recorded against your key.

When the engagement turns into a retainer and the question becomes "tell me when this changes", the answer is a monitor per client, not a scheduled re-run of this.

Start free with 1,000 credits, create a key named after your next engagement, and build one landscape before you build anything around it.

Try this yourself — no signup needed

Explore all 31 CrawlForge scraping and extraction tools in the playground, then start free with 1,000 credits.

1,000 free credits • One-time • No credit card required

Tags

  • consultants
  • market research
  • scrape
  • map_site
  • deep_research
  • b2b

About the Author

CrawlForge Team

Engineering Team

Building the most comprehensive web scraping MCP server. We create tools that help developers extract, analyze, and transform web data for AI applications.

Newsletter

Stay updated with the latest insights

Get tutorials, product updates, and web scraping tips delivered to your inbox.

No spam. Unsubscribe anytime.

FAQ

Frequently asked questions

01Is CrawlForge a market research data provider for consultants?

No. It holds no market-size estimates, no firmographics, no intent signals and no contact records. It reads public pages you point it at and returns what they say, with the URL and a timestamp. For a consulting engagement that makes it the layer that builds and cites the landscape from vendors' own sites, alongside whatever syndicated report or dataset the firm already licenses.

02How do I find the vendors in a category I have never researched?

search_web costs 5 credits a query and accepts a queries array, so several phrasings run in one call and are charged per query that returned results. deep_research costs 10 credits a run and returns verbatim passages plus a sources list with each URL, its title and whether it was fetched. Both return URLs rather than a written answer, which is what you want at the long-list stage. map_site then lists each domain's pages for 2 credits so you pick the pricing or product page instead of guessing its path.

03Can I separate one client's usage from another's?

Yes. Create a named API key per client; an account holds up to ten. Every call logs the key it came from alongside the tool, the URL and the credits charged, and the request log in the dashboard filters by key and exports as CSV. Deleting a key at close-out revokes it and keeps its rows, so the engagement's record survives the key.

04Do I need a subscription to run one engagement?

No. The free tier is a one-time 1,000 credits, enough for two 30-vendor landscapes at the cost worked out above. Beyond that, one-time credit packs start at $3 for 1,000 credits and $14 for 5,000, so a practice that runs research in bursts can buy per engagement rather than pay monthly. Paid plans exist for firms that want a monthly allocation with rollover.

05Why not use batch_scrape for the whole vendor list at once?

On the hosted REST API batch_scrape charges 5 credits per URL attempted and returns plain text, whereas scrape charges 2 per page and returns markdown, metadata and links from one fetch. Batch is one round trip and stores its results for 24 hours, which is convenient, but for a landscape you will read carefully a loop of scrape calls costs less than half as much and gives you more to cite.

Keep reading

Related Articles