# What is web scraping, and when is an API cheaper?

Web scraping means fetching a public web page and extracting structured data from it. The part that decides your cost is not the fetching, it is the parsing: a proxy returns you HTML and leaves the parser your problem, while a scraper API returns records and absorbs the maintenance when the site changes.

[![apidojo](/blog/authors/apidojo.png)**Apidojo** *15+ years of experience in data retrieval.*](/blog/authors/apidojo/)

**28 September 2026** Published

**3 min** Read

![Cover card reading "What is web scraping, and when is an API cheaper?"](/blog/covers/what-is-web-scraping.png)

-   Scraping is two jobs, fetching and parsing, and almost all the ongoing cost sits in the second one.
-   A proxy API sells you the fetch. A scraper API sells you the parse. They are priced differently because they are different products.
-   Build your own when you target many sites shallowly. Buy when you target a few sites deeply.
-   Public data is not the same as unrestricted data. Terms of service, personal data rules and rate limits all still apply.

## The two jobs hiding inside one word

Scraping sounds like one activity. It is two, and they behave completely differently.

**Fetching** is getting the page. It is a solved problem with a known cost: requests, addresses, and the occasional blocked one. It scales linearly and it rarely surprises you.

**Parsing** is turning that page into records. It is not solved, because the page is a moving target. A class name changes, a field moves inside a nested object, a counter starts arriving as “3.4K” instead of 3400, and the parser that worked on Friday returns nulls on Monday.

Almost every disappointing scraping project is a project that budgeted for the first job and got billed by the second.

## Proxy API or scraper API

The two things sold under the same banner divide exactly along that line.

|  | Proxy API | Scraper API |
| --- | --- | --- |
| Returns | Raw HTML | Typed records |
| You write | The parser | The query |
| Breaks when | The markup changes | Nothing, from your side |
| Priced by | Request or bandwidth | Result |
| Best for | Many sites, shallow | Few sites, deep |

Neither is the right answer in general. The question is how many distinct sites you target.

If you need one field from four hundred different sites, no vendor has a parser for all of them and a proxy plus your own generic extraction is the only thing that scales. If you need everything from five sites, forever, then maintaining five parsers is a permanent tax you can simply buy your way out of.

## The number that decides it

Per-result pricing makes buying easy to forecast: in this catalogue it runs from $0.30 per 1,000 TikTok posts up to $1.00 per 1,000 YouTube transcripts, and you can multiply that by the volume you need before spending anything.

Building has no equivalent line. The honest comparison is not “API cost versus zero”, it is API cost versus the engineering hours that go into the parser, the proxy pool, the retry logic, the monitoring that tells you the parser broke, and the afternoon someone loses to fixing it. That last item is the one that never appears in the estimate and never stops recurring.

A rough test: if the data matters enough that somebody would notice within a day of it breaking, you are buying an SLA as much as a parser, and that is the case where the per-result price is cheap.

## Public does not mean unrestricted

Three constraints apply regardless of whether a page is publicly reachable, and they are worth separating because people tend to collapse them.

**The law.** Scraping public pages is not inherently illegal in most jurisdictions. That is a much narrower statement than it sounds, and nothing here is legal advice.

**Data protection.** If the records describe identifiable people, the GDPR and its equivalents apply whether or not the page was public. A username and a follower count is personal data. Being visible on the open web does not change that.

**Terms of service.** A site can prohibit automated access as a matter of contract, independently of the law. That is a separate question from legality, and it is the one most likely to affect you commercially.

Practically: collect what you need rather than what you can, keep it only as long as it is useful, and take advice for your own situation rather than relying on a blog post.

## Where to start

Pick one source and one question. Run a small scrape, look at the actual fields rather than the marketing description of them, and check what is missing before you design anything around it. Most of the expensive surprises in this work are fields that turn out to be optional, or numbers that turn out to be strings.

## The scrapers behind this

-   [Tweet Scraper V2 Tweets from searches, profiles, Lists and URLs; 49–64 tweets/sec; minimum 50 per query $0.40 / 1K tweets](/scrapers/twitter-scraper/)
-   [TikTok Scraper Profiles, videos, hashtags, music and locations; 460 posts/sec measured $0.30 / 1K posts](/scrapers/tiktok-scraper/)

## Questions people ask

### Is web scraping legal?

Scraping public pages is not inherently illegal in most jurisdictions, but that is not the whole question and this is not legal advice. What matters in practice is what you collect, what you do with it, and whose terms you agreed to. Personal data brings the GDPR and similar regimes into play regardless of whether the page was public, and a site's terms of service may prohibit automated access independently of the law. Take advice for your own case.

### What is the difference between a proxy API and a scraper API?

A proxy API fetches the page and hands you the raw HTML. You write the parser and you maintain it. A scraper API fetches and parses, and hands you typed records. The proxy is cheaper per request and more expensive per month, because the parser is where the ongoing work lives.

### Do I need proxies to scrape?

Only if you are fetching the pages yourself. Proxy pools exist because sites rate-limit by address. If you use a scraper API the pool is the vendor's problem and its cost is inside the per-result price.

### How much does scraping cost?

Per-result pricing across this catalogue runs from $0.30 per 1,000 TikTok posts to $1.00 per 1,000 YouTube transcripts. Building it yourself is not free either: the cost moves from a per-result line to engineering time, which is harder to forecast and does not stop when the run does.

1.  [API Dojo scraper catalogue and per-result pricing](https://apify.com/apidojo) apify.com

![apidojo](/blog/authors/apidojo.png)

## [Apidojo](/blog/authors/apidojo/)

15+ years of experience in data retrieval.

API Dojo is always ready for you. With 15+ years of experience, data acquisition has never been that easy. Cheapest prices with high availability.

-   🏯 15+ years of experience in data retrieval.
-   💰 Cheapest prices, qualified data.
-   ⚙️ High availability, extreme maintenance.
-   🐉 Hyper communicative, excellent support.

-   [apidojo.io](https://apidojo.io)

## More in this cluster

-   [What web scraping is and when to use it The vocabulary, the legal lines, and the choice between a proxy that returns HTML and an API that returns records. Start here if you are new to this.](/blog/web-scraping/)
