A consultant's market research starts over with every engagement. New category, new vendor list, and a client who will ask on slide fourteen where that number came from. Last quarter's spreadsheet does not transfer. The method does.
A market research api for consultants has to work from zero, because on day one there is no list of vendors to point it at, and it has to hand back facts with a source and a date, because "we checked" is not an answer a partner can give. It also has to cost something you can put against one client rather than into overhead.
Where do you start when there is no vendor list yet?
With two calls that return URLs rather than answers. search_web costs 5 credits a query and takes a queries array, so six phrasings of "who sells X to Y" run in one call for 30 credits. The snippets are usually enough to tell a vendor from a listicle. When you want the passages too, deep_research costs 10 credits a run: it searches, fetches up to ten sources and returns the best-matching passages verbatim, with a sources list carrying each URL, its title and whether it was actually fetched or only seen in a snippet. Nothing is synthesised on the hosted API, which is the point. What comes back is a long-list of domains and the sentences that put them on it.
The next step is finding the right page on each domain, and the wrong way is guessing that pricing lives at /pricing. map_site costs 2 credits a domain and reads the site's sitemap, falling back to a bounded crawl when there is none, so you get the URL list and pick out pricing, product, customers and careers from it. Thirty vendors is 60 credits and you never fetch a 404.
How does a fact get into the deck with its source?
One scrape per page. Ask for markdown, metadata and links together: every format comes from the same fetch and the call costs 2 credits however many you request. The response carries a scraped_at timestamp, which is the date column your deliverable needs.
curl -X POST https://crawlforge.dev/api/v1/tools/scrape \
-H "X-API-Key: cf_live_NORTHWIND_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://vendor-a.com/pricing",
"formats": ["markdown", "metadata", "links"]
}'The row you write is vendor, page URL, the claim, the sentence it rests on, and scraped_at. Quote the sentence rather than paraphrasing it. A cell that reads "SSO on all plans" is an opinion until the sentence beside it says so, and the due diligence post covers what happens when a model fills that cell instead.
If the same question has to be answered for every vendor, add a question format for 1 credit more and the call returns the five sentences that best match it, each with an offset into the markdown. The RevOps post is about running that across a category; this one is about getting the category in the first place.
How do you keep one client's spend separate from another's?
Create one API key per client and name it after the engagement. An account holds up to ten, which is more concurrent engagements than most practices run. Every call made with that key writes a row to the request log carrying the tool, the target URL, the credits charged and the key it came from, and the log filters by key and exports as CSV.
That does two things nobody asks for until they need it. The engagement's spend becomes a query rather than a reconstruction, and the client's research is separable from everyone else's when they ask what you hold. At close-out, delete the key. Deletion revokes it and keeps the rows, so the record of what was fetched for whom survives the engagement.
The example above puts the client in the key's name for the same reason: the name is what you will be filtering on.
What does one engagement cost?
A 30-vendor landscape, done the way above:
| Step | Calls | Credits |
|---|---|---|
| Long-list: six search phrasings, one deep_research run | 1 + 1 | 40 |
| Find the pages: map_site per vendor | 30 | 60 |
| Read four pages per vendor | 120 | 240 |
| One question asked of every pricing page | 30 | 90 |
| Engagement total | 430 |
The free tier's one-time 1,000 credits cover two of those. After that a 1K pack is $3 and a 5K pack is $14, bought when an engagement needs them rather than on a monthly plan, which is a separate argument. At the pack rate the whole landscape above costs about a dollar and a half, small enough to carry inside the fee and auditable if a client asks what the line covers.
What won't it do?
- No market-size figures. It sizes nothing and forecasts nothing. If the slide needs a TAM, that comes from a report or a dataset.
- No contact or firmographic data. It reads pages you point it at. It is not a company database and does not overlap with one.
- No client login. One account, keys per client, and the client sees what you send them.
robots.txtis respected by default. A disallowed page is refused before anything is fetched and costs nothing, and turning the check off is recorded against your key.
When the engagement turns into a retainer and the question becomes "tell me when this changes", the answer is a monitor per client, not a scheduled re-run of this.
Start free with 1,000 credits, create a key named after your next engagement, and build one landscape before you build anything around it.