What is web scraping, and when is an API cheaper?
Web scraping means fetching a public web page and extracting structured data from it. The part that decides your cost is not the fetching, it is the parsing: a proxy returns you HTML and leaves the parser your problem, while a scraper API returns records and absorbs the maintenance when the site changes.

The two jobs hiding inside one word
Scraping sounds like one activity. It is two, and they behave completely differently.
Fetching is getting the page. It is a solved problem with a known cost: requests, addresses, and the occasional blocked one. It scales linearly and it rarely surprises you.
Parsing is turning that page into records. It is not solved, because the page is a moving target. A class name changes, a field moves inside a nested object, a counter starts arriving as “3.4K” instead of 3400, and the parser that worked on Friday returns nulls on Monday.
Almost every disappointing scraping project is a project that budgeted for the first job and got billed by the second.
Proxy API or scraper API
The two things sold under the same banner divide exactly along that line.
| Proxy API | Scraper API | |
|---|---|---|
| Returns | Raw HTML | Typed records |
| You write | The parser | The query |
| Breaks when | The markup changes | Nothing, from your side |
| Priced by | Request or bandwidth | Result |
| Best for | Many sites, shallow | Few sites, deep |
Neither is the right answer in general. The question is how many distinct sites you target.
If you need one field from four hundred different sites, no vendor has a parser for all of them and a proxy plus your own generic extraction is the only thing that scales. If you need everything from five sites, forever, then maintaining five parsers is a permanent tax you can simply buy your way out of.
The number that decides it
Per-result pricing makes buying easy to forecast: in this catalogue it runs from $0.30 per 1,000 TikTok posts up to $1.00 per 1,000 YouTube transcripts, and you can multiply that by the volume you need before spending anything.
Building has no equivalent line. The honest comparison is not “API cost versus zero”, it is API cost versus the engineering hours that go into the parser, the proxy pool, the retry logic, the monitoring that tells you the parser broke, and the afternoon someone loses to fixing it. That last item is the one that never appears in the estimate and never stops recurring.
A rough test: if the data matters enough that somebody would notice within a day of it breaking, you are buying an SLA as much as a parser, and that is the case where the per-result price is cheap.
Public does not mean unrestricted
Three constraints apply regardless of whether a page is publicly reachable, and they are worth separating because people tend to collapse them.
The law. Scraping public pages is not inherently illegal in most jurisdictions. That is a much narrower statement than it sounds, and nothing here is legal advice.
Data protection. If the records describe identifiable people, the GDPR and its equivalents apply whether or not the page was public. A username and a follower count is personal data. Being visible on the open web does not change that.
Terms of service. A site can prohibit automated access as a matter of contract, independently of the law. That is a separate question from legality, and it is the one most likely to affect you commercially.
Practically: collect what you need rather than what you can, keep it only as long as it is useful, and take advice for your own situation rather than relying on a blog post.
Where to start
Pick one source and one question. Run a small scrape, look at the actual fields rather than the marketing description of them, and check what is missing before you design anything around it. Most of the expensive surprises in this work are fields that turn out to be optional, or numbers that turn out to be strings.
Questions people ask
Is web scraping legal?
Scraping public pages is not inherently illegal in most jurisdictions, but that is not the whole question and this is not legal advice. What matters in practice is what you collect, what you do with it, and whose terms you agreed to. Personal data brings the GDPR and similar regimes into play regardless of whether the page was public, and a site's terms of service may prohibit automated access independently of the law. Take advice for your own case.
What is the difference between a proxy API and a scraper API?
A proxy API fetches the page and hands you the raw HTML. You write the parser and you maintain it. A scraper API fetches and parses, and hands you typed records. The proxy is cheaper per request and more expensive per month, because the parser is where the ongoing work lives.
Do I need proxies to scrape?
Only if you are fetching the pages yourself. Proxy pools exist because sites rate-limit by address. If you use a scraper API the pool is the vendor's problem and its cost is inside the per-result price.
How much does scraping cost?
Per-result pricing across this catalogue runs from $0.30 per 1,000 TikTok posts to $1.00 per 1,000 YouTube transcripts. Building it yourself is not free either: the cost moves from a per-result line to engineering time, which is harder to forecast and does not stop when the run does.