Search APIs and scraping APIs are often discussed as if they solve the same problem, but they do not.
A search API helps you find relevant pages, while a scraping API helps you get the content from a page you already know about.
That sounds like a small distinction, but it changes how you should build a workflow.
For example, if you already have the URL to a product page and want its price, description, and availability, you don't need a search API. You already know where the data lives, so you can scrape the page directly.
But if you want to find companies that recently launched an open source observability product and then collect information about each one, scraping alone is not enough. You first need a way to discover the relevant pages.
That is where search comes in.
In practice, many useful web data workflows need one, the other, or both.
Use a search API when you know what you want, but not where it is
Suppose you want to find "open source alternatives to Datadog". You know the topic, but you do not have a list of pages yet.
This is a search problem. You could make a request like:
const result = await spidra.search({
query: "open source alternatives to Datadog",
sources: ["web"],
});The response gives you the relevant results, including their URLs and metadata.
At this point, the search API has answered which pages I should look at. It has not necessarily answered what is actually on those pages.
That is an important difference.
Traditional SERP APIs such as SerpApi are built mainly around returning structured search-engine result data such as titles, links, snippets, rankings, and other SERP metadata.
Newer search products often go further and can return extracted page content with the search results, but the underlying job is still discovery.
Use a scraping API when you already know the URL
Now suppose you already have:
https://spidra.io/pricingand you want the pricing table on that page.
There is no reason to search for the page first.
You can scrape it directly.
A scraping API takes the known URL, loads the page, handles whatever rendering is required, and returns the content in a usable format such as Markdown, HTML, JSON, screenshots, or structured fields.
Search and scraping are often two steps in the same workflow
The interesting part is what happens when you combine them.
Imagine you are building a lead enrichment workflow and want to find companies working on AI observability tools.
You could begin with Search:
const result = await spidra.search({
query: "AI observability startups",
sources: ["web"],
});That gives you the pages.
Then you can scrape the useful results to get the actual information you care about.
For example:
- company name;
- product description;
- founders;
- funding;
- pricing;
- GitHub URL;
- contact information.
The workflow becomes:
This pattern shows up everywhere:
- A research tool might search for recent papers, then extract the full text.
- A sales tool might search for companies that recently raised funding, then scrape their websites for product and contact information.
- A monitoring system might search for newly published pages, then scrape them before deciding whether anything changed.
- An AI agent might search for sources first, then read only the pages it thinks are useful.
Search alone is enough when you only need the results
Not every search needs to turn into scraping. Sometimes the result metadata is already enough.
For example, if you want to build a page that shows the latest articles about a company, you may only need:
- title;
- URL;
- publication date;
- snippet;
- source.
There is no reason to pay the extra cost and latency of fetching every article if your application does not need the full text.
The same applies to images.
If you are searching for:
mountain wallpaper larger:1920x1080and all you need is a list of image results, scraping the pages behind them may add no value.
Scraping alone is enough when discovery has already happened
The reverse is also true. If your application already receives URLs from somewhere else, adding Search would only make the workflow more complicated.
For example, imagine a user pastes five competitor pricing pages into your product and asks you to track them.
You already have the URLs.
The same applies to:
- a list of ecommerce product pages;
- company websites from a CRM;
- URLs from a sitemap;
- documentation pages from a known domain;
- links submitted by users.
Search is useful when you need discovery. If discovery has already happened, skip it.
Crawl is different again
Search and scraping are not the only two options. Sometimes you know the website, but you do not know every page inside it.
For example, you might know that all the useful documentation is somewhere under:
https://docs.example.combut you do not have every documentation URL.
That is a crawl problem.
A crawler starts from one page, follows links across the site, and discovers the pages for you.
So the simplest way to separate the three is:
What you know | Use |
You know the topic or question | Search |
You know the exact URL | Scrape |
You know the site, but not every page | Crawl |
That distinction becomes useful because it stops you from using a more expensive or complicated tool than the job requires.
Search plus scrape is especially useful for AI agents
This distinction becomes even more important when the user is not a person manually clicking links.
An agent often starts with a goal such as:
There is no URL in that instruction.
The agent first needs discovery.
It can search for relevant sources, inspect the results, then fetch the pages worth reading.
A useful workflow might be:
Search gives the agent breadth, and scraping gives it depth.
How this works in Spidra
Spidra now gives you the different pieces separately because not every workflow needs all of them.
You can use:
- Search when you need to discover pages;
- Scrape when you have a known URL;
- Batch when you have several known URLs;
- Crawl when you want to discover pages inside a known site;
- Browser actions when the page needs interaction;
- Structured extraction when you need specific fields instead of raw content.
You can start visually in the dashboard or use the API, SDKs, and MCP when you want to put the same workflow into an application or agent.
Try Search in Spidra or read the Search docs.
