The field guide
66 essential terms covering web scraping, AI agents, the Model Context Protocol, and data extraction.
66terms
4guides
01Guides
Browse by category
01 · Web Scraping
Web Scraping Terms
The 19 core concepts behind fetching, rendering, and parsing web pages — from robots.txt and rate limiting to headless browsers, bot detection, and proxies.
1919 terms
02 · AI / MCP
AI and MCP Terms
The 19 terms that connect large language models to live web data — tokens, embeddings, function calling, MCP tools, search APIs, and the protocol that ties them together.
1919 terms
03 · Data
Data and API Terms
The 14 formats and protocols that move scraped data between systems — REST and JSON, webhooks, SERP APIs, embedded page state, and the markup search engines read.
1414 terms
04 · Industry
Web Data Industry Terms
The 14 business terms behind web data work — managed scraping, credit-based pricing, real-time data, price monitoring, lead enrichment, and governance.
1414 terms
02A–Z
All terms A–Z
A
AI AgentAn AI agent is an autonomous system powered by a large language model that can reason about tasks, make decisions, and take actions by using tools. Agents go beyond simple chatbots by planning and executing multi-step workflows.AI / MCPAI Search APIA web search API for AI agents returns search results as structured data, such as titles, URLs and snippets, so a model can find current pages before reading them.AI / MCPAPI EndpointAn API endpoint is a specific URL where an API receives requests. Each endpoint performs a specific function, like retrieving data, creating records, or triggering actions.Data
B
Batch ScrapingBatch scraping fetches a list of known URLs in one request instead of one call per page, which cuts round trips and makes large collection jobs easier to manage.DataBot DetectionWebsites detect scrapers by scoring signals such as request rate, IP reputation, browser fingerprint and JavaScript behaviour, then block or challenge traffic that looks automated.Web ScrapingBrowser AgentA browser agent is an AI system that operates a real web browser, looking at a page, then clicking, typing and scrolling to complete a task the way a person would.AI / MCP
C
CAPTCHA SolvingCAPTCHA solving refers to automated techniques for bypassing CAPTCHA challenges that websites use to distinguish humans from bots. This includes image recognition, token-based solving, and browser fingerprint emulation.Web ScrapingCompetitive IntelligenceCompetitive intelligence is the systematic collection and analysis of information about competitors, market trends, and industry dynamics. It informs strategic decisions about pricing, positioning, and product development.IndustryContent MigrationContent migration is the process of moving content from one platform or system to another. It involves extracting content from the source, transforming it to match the target format, and loading it into the new system.IndustryContext WindowThe context window is the maximum amount of text (measured in tokens) that a language model can process in a single request. It includes both the input prompt and the generated output.AI / MCPCredit-Based PricingCredit-based pricing charges each API call a set number of credits according to how much work it does, and you buy credits through a plan or a one-off top-up.IndustryCSS SelectorA CSS selector is a pattern used to select and target specific HTML elements on a web page. In web scraping, selectors identify exactly which data to extract from a page's structure.Web Scraping
D
Data GovernanceData governance is the framework of policies, procedures, and standards that ensures data is managed properly throughout its lifecycle. It covers data privacy, compliance, access control, and quality standards.IndustryData PipelineA data pipeline is an automated sequence of steps that collects, processes, transforms, and delivers data from sources to destinations. It enables continuous data flow between systems without manual intervention.IndustryData QualityData quality measures how well a dataset meets the requirements of its intended use. Key dimensions include accuracy, completeness, consistency, timeliness, and validity of the data.IndustryDOM ParsingDOM parsing is the process of converting raw HTML into a structured Document Object Model tree. This tree representation allows programs to navigate and extract specific elements from a web page.Web ScrapingDynamic ContentDynamic content is web content that is loaded or generated by JavaScript after the initial page load. This includes single-page applications, AJAX-loaded data, and client-side rendered content.Web Scraping
E
Embedded Page StateEmbedded page state is the JSON data a website ships inside its own HTML, such as Next.js __NEXT_DATA__ or a Redux store, which the page's JavaScript reads to render itself.DataEmbeddingsEmbeddings are dense numerical vector representations of text, images, or other data. They capture semantic meaning in a format that enables similarity search, clustering, and other machine learning operations.AI / MCPETL (Extract, Transform, Load)ETL is a data integration process that extracts data from sources, transforms it into a suitable format, and loads it into a target system. It is the standard approach for moving data between systems.Industry
F
Fine-TuningFine-tuning is the process of further training a pre-trained language model on a specific dataset to specialize its behavior for a particular task or domain. It adapts general-purpose models to targeted use cases.AI / MCPFunction CallingFunction calling is the ability of language models to invoke external functions or APIs during a conversation. The model decides when to call a function, generates the appropriate arguments, and processes the returned results.AI / MCP
G
Geo-Targeted ScrapingGeo-targeted scraping collects a page or search result as it appears to visitors in a specific country, city or language, since prices, results and content vary by location.Web ScrapingGraphQLGraphQL is a query language for APIs that allows clients to request exactly the data they need. Unlike REST, a single GraphQL endpoint serves all queries, with the client specifying the data shape.Data
H
Headless BrowserA headless browser is a web browser without a graphical user interface that can be controlled programmatically. It executes JavaScript and renders pages exactly like a regular browser, but runs in the background.Web ScrapingHTML ParsingHTML parsing is the process of analyzing HTML markup to extract its structure and content. Parsers convert raw HTML strings into navigable tree structures that programs can query and manipulate.DataHTTP HeadersHTTP headers are key-value pairs sent with HTTP requests and responses that provide metadata about the communication. In scraping, headers like User-Agent, Accept, and Cookie are critical for successful requests.Web Scraping
J
JSONJSON (JavaScript Object Notation) is a lightweight data interchange format that is easy for humans to read and machines to parse. It is the standard format for API responses and structured data exchange.DataJSON-LDJSON-LD (JSON for Linking Data) is a method of encoding structured data using JSON format. It is the preferred format for embedding schema.org markup in web pages for search engine understanding.Data
L
Large Language Model (LLM)A large language model is a neural network trained on vast amounts of text data that can understand and generate human language. LLMs power AI assistants, code generators, and autonomous agents.AI / MCPLead EnrichmentLead enrichment is the process of supplementing basic lead information with additional data points like company size, industry, technology stack, and social profiles. It helps sales teams prioritize and personalize outreach.IndustryLLM-Ready DataLLM-ready data is web content converted into a clean format a language model can use directly, usually markdown or JSON, with navigation, ads and markup removed.AI / MCP
M
Managed Scraping ServiceA managed scraping service runs the browsers, retries and parsers for you, so a team calls an API instead of maintaining its own Playwright or Puppeteer infrastructure.IndustryMarkdownMarkdown is a lightweight markup language that uses plain text formatting syntax. It is widely used for documentation, content creation, and as a clean intermediate format for extracted web content.DataMCP ClientAn MCP client is an application or AI model that connects to MCP servers to discover and invoke tools. It sends tool call requests and processes the structured responses returned by the server.AI / MCPMCP ServerAn MCP server is a service that exposes tools and resources through the Model Context Protocol. It registers available functions, handles incoming tool calls from AI clients, and returns structured results.AI / MCPMCP ToolAn MCP tool is a single named action with a typed input schema that an MCP server exposes for an AI model to call, such as fetching a URL or running a web search.AI / MCPModel Context Protocol (MCP)The Model Context Protocol is an open standard that enables AI models to interact with external tools and data sources through a unified interface. It provides a structured way for LLMs to call functions, access APIs, and retrieve real-time information.AI / MCP
N
P
PaginationPagination is the practice of dividing content across multiple pages. Handling pagination in web scraping means automatically navigating through all pages to collect complete datasets.Web ScrapingPrice MonitoringPrice monitoring is the automated tracking of product and service prices across websites over time. It enables businesses to respond to competitor pricing changes, optimize their own pricing, and identify market trends.IndustryPrompt EngineeringPrompt engineering is the practice of designing and refining instructions given to language models to achieve desired outputs. It involves crafting system prompts, few-shot examples, and structured queries.AI / MCPProxy RotationProxy rotation is the practice of cycling through multiple proxy IP addresses when making web requests. This distributes requests across different IPs to avoid rate limits and IP-based blocking.Web Scraping
R
Rate LimitingRate limiting is a technique used by websites and APIs to control the number of requests a client can make within a given time period. It prevents server overload and defends against abusive scraping.Web ScrapingReal-Time Web DataReal-time web data is information fetched from the live web at the moment it is needed, rather than read from a stored dataset or a model's training data.IndustryResidential ProxyA residential proxy routes requests through an IP address an internet provider assigned to a home connection, so traffic looks like an ordinary visitor, not a data center.Web ScrapingREST APIA REST API (Representational State Transfer) is a web service architecture that uses standard HTTP methods to perform operations on resources. It is the most common API style for web services.DataRetrieval-Augmented Generation (RAG)RAG is an AI architecture that combines information retrieval with text generation. It first retrieves relevant documents from external sources, then uses them as context for the language model to generate accurate, grounded responses.AI / MCPRobots.txtRobots.txt is a standard text file placed at the root of a website that tells web crawlers which pages they are allowed or disallowed from accessing. It is part of the Robots Exclusion Protocol.Web Scraping
S
Schema MarkupSchema markup is a vocabulary of tags (from schema.org) that you add to HTML to improve how search engines read and represent your page. It defines types like Product, Article, Organization, and their properties.DataSchema-Based ExtractionSchema-based data extraction pulls specific fields from a web page into a structure you define in advance, usually a JSON Schema, so every page returns the same shape.DataSEO AuditAn SEO audit is a comprehensive analysis of a website's search engine optimization performance. It evaluates technical SEO, on-page content, metadata, site structure, and identifies opportunities for improvement.IndustrySERP APIA SERP API returns a search engine's results page as structured data, including each result's position, title and URL, so rankings can be tracked and analysed in code.DataSitemapA sitemap is an XML file that lists all the URLs on a website, along with metadata like last modification date and priority. It helps search engines and crawlers discover and index all pages efficiently.Web ScrapingStructured DataStructured data is information organized in a predefined format that makes it easy for machines to parse and understand. On the web, it typically refers to schema.org markup embedded in HTML pages.DataStructured OutputStructured output refers to data returned in a predictable, machine-readable format like JSON, rather than free-form text. It enables reliable downstream processing by AI agents and data pipelines.AI / MCP
T
TokenA token is the basic unit of text that language models process. Text is split into tokens (roughly 4 characters or 0.75 words each) before being processed by the model. Token counts determine costs and context limits.AI / MCPTool UseTool use is the capability of AI models to interact with external tools, APIs, and services to accomplish tasks beyond text generation. It extends model capabilities to include web browsing, code execution, data retrieval, and more.AI / MCP
U
V
W
Web CrawlerA web crawler is a program that systematically browses the web by following links from page to page. Crawlers discover and index content across entire websites or domains.Web ScrapingWeb DataWeb data is any information that is publicly accessible on the internet. It includes website content, social media posts, public APIs, government records, and any other data available through web protocols.IndustryWeb ScrapingWeb scraping is the automated extraction of data from websites. It involves programmatically fetching web pages and parsing their content to collect structured information.Web ScrapingWeb Scraping APIA web scraping API is a hosted service that fetches a web page on request and returns its content as clean data, so you don't run your own browsers or parsers.Web ScrapingWebhookA webhook is an HTTP callback that delivers data to a specified URL when an event occurs. Unlike polling, webhooks push data in real-time, enabling event-driven architectures.Data