# YouTube Comments Scraper: videos and Shorts

YouTube Comments Scraper reads the comments on any public YouTube video or Short and returns each one with its author, its like and reply counts, and when it was posted.

[Run this scraper free](https://apify.com/apidojo/youtube-comments-scraper)

[See the three jobs](#jobs)

apidojo/youtube-comments-scraper

**Videos**

**Comments returned**

Estimated run cost

$0.00

Free Apify plan: 5 runs a month, 10 items a run. No credit card.

**$0.001** per video queried

**~20** comments included in that

**$0.0005** per comment after

**99.99%** run success, last 30 days

Read on 27 September 2026.

## Five fields, and the two that change the answer

The URL decides what you read and the ceiling decides how much. Between them sit a sort order that changes which comments you get, and a switch that changes how you are billed.

-   `startUrls`

    $0.001per video

    ~20 comments free

    A video or Shorts URL. Both are first-class here, the listing benchmarks them separately at 251 and 267 comments a second.

    `https://youtu.be/Y9Um-8nPnVQ`

-   `sort`

    top | newfree to change

    Defaults to top

    top returns what YouTube considers most relevant; new returns newest first. Same price, different comments.

    `sort: "new"`

-   Bills differently from every other comment scraper here

    `includeReplies`

    +$0.001per thread

    A query, not an item

    Off by default. On, each top-level comment’s reply thread is fetched, and the schema says each thread is charged as its own query event.

    `includeReplies: true`

-   `maxItems`

    ÷ $0.0005past the free 20

    Run-wide

    Comments across the whole run. The listing’s troubleshooting names it first when a run returns less than expected.

    `1000`

Sorting is worth a moment: the listing notes that top and new may surface different comment counts on the same video, so a thin result is sometimes the sort order rather than the ceiling.

**Apify actor:** apidojo/youtube-comments-scraper

**Inputs:** startUrls · sort · includeReplies · maxItems · customMapFunction

**Accepts:** Video and Shorts URLs

**Sorting:** top for most relevant, new for newest first

**Counts:** Strings, not numbers, likeCount arrives as "3.4K"

**Timestamps:** Relative only, "5 months ago", with no absolute date

**Replies:** Charged as a query event per thread, not as items

**Reliability:** 99.99% of 118,137 runs succeeded in the last 30 days

## Most relevant, most recent, or everything under a batch

Three runs. The first two cost the same and answer different questions; the third is where the per-video free allowance starts to matter more than the per-comment rate.

**01 Most relevant $0.001 + $0.0005 past ~20**

**02 Most recent Same price, different order**

**03 A batch of videos One allowance per video**

### How do you scrape comments from a YouTube video?

Paste the URL and leave sort at top. The video query is $0.001 and brings about twenty comments; after that each is $0.0005. A thousand comments from one video costs $0.49.

Input

```
{
  "startUrls": ["https://youtu.be/Y9Um-8nPnVQ"],
  "sort": "top",
  "maxItems": 1000
}
```

$0.001 + 980 × $0.0005 \= **$0.49**

What comes back

The **comment record**, Author, text, counts and a relative time, all seven guaranteed fields.

```
{
  "text": "This one hit you in the feels? Check out the full video…",
  "likeCount": "3.4K",
  "replyCount": "8",
  "publishedTime": "5 months ago",
  "author": {
    "id": "UCa90xqK2odw1KV5wHU9WRhg",
    "name": "@TheOffice"
  }
}
```

likeCount is the string "3.4K", not the number 3400

The listing’s field table gives likeCount and replyCount as strings, and the sample shows the abbreviated form. Sorting or thresholding on them means parsing first, and the abbreviation has already thrown away the exact figure, "3.4K" could be anything from 3,350 to 3,449.

### How do you get the newest YouTube comments?

Set sort to new. It costs exactly the same as top and is the right choice for monitoring, because relevance ranking will bury a comment posted an hour ago beneath one from last year.

Input

```
{
  "startUrls": ["https://youtu.be/Y9Um-8nPnVQ"],
  "sort": "new",
  "maxItems": 500
}
```

$0.001 + 480 × $0.0005 \= **$0.24**

What comes back

The **conditional fields**, Two flags the listing marks as sometimes present, not always.

```
{
  "isPinned": true,
  "isHearted": false
}
```

Relative times cannot be compared across runs

"5 months ago" is measured from the moment of scraping, so the same comment reads differently next week and two runs cannot be merged on it. Record your own run timestamp and resolve the phrase against it, or the dataset loses its chronology.

### How do you scrape comments across many videos?

List the URLs together. Each video carries its own free twenty, so ten videos sharing a thousand comments cost $0.41 where one video taking a thousand alone costs $0.49, breadth is slightly subsidised.

Input

```
{
  "startUrls": [
    "https://youtu.be/Y9Um-8nPnVQ",
    "https://www.youtube.com/shorts/abc123"
  ],
  "sort": "top",
  "maxItems": 1000
}
```

10 × $0.001 + 800 × $0.0005 \= **$0.41**

What comes back

The **author block**, A channel ID on every comment, which is what makes cross-video tracking possible.

```
{
  "author": {
    "id": "UCa90xqK2odw1KV5wHU9WRhg",
    "name": "@TheOffice",
    "thumbnails": [
      { "height": 48, "width": 48, "url": "https://yt3.ggpht.com/…" }
    ]
  }
}
```

Nothing on a row names the video it came from

The guaranteed fields cover the comment and its author, not the source URL. On a multi-video run, either submit one video per run or add the URL yourself through customMapFunction, which exists for exactly this kind of reshaping.

## Seven fields guaranteed, two conditional

The listing publishes an availability table rather than just a sample, which is unusually careful and worth taking literally: seven fields are marked always present, two are marked sometimes, and the types are not what you would assume.

-   4

    Always present

    The comment

    -   text, the full body
    -   likeCount, string
    -   replyCount, string
    -   publishedTime, relative string
-   3

    Always present

    author

    -   id, the channel ID
    -   name, e.g. @TheOffice
    -   thumbnails, array, each with url, width, height
-   2

    Sometimes present

    Marked conditional

    -   isPinned, boolean
    -   isHearted, creator hearted it
-   0

    Not returned

    Plan around these

    -   No absolute timestamp
    -   No numeric engagement
    -   No source video on the row

**text:** The comment body, newlines and links intact

**likeCount:** A string in YouTube’s own abbreviated form, e.g. "3.4K", parse before you compare

**replyCount:** Also a string. "8" is a string eight, not an integer

**publishedTime:** Relative phrase such as "5 months ago", measured from when the run happened

**author.id:** The commenter’s channel ID, the only stable key for tracking someone across videos

**author.name:** Display handle, usually @-prefixed

**author.thumbnails:** Array of avatar sizes, each with url, width and height

**isPinned:** Present sometimes. Pinned comments are usually the creator’s own

**isHearted:** Present sometimes. The creator’s heart is the cheapest endorsement signal on the platform

author.id is the field to build on. Names change and display handles are not unique, but the channel ID lets you follow one commenter across every video you have ever scraped, which is how repeat critics and repeat advocates get found.

## What to settle before you store any of it

Four of these come from the listing’s own documentation of its output and inputs. They matter at schema-design time rather than at run time.

-   01

    Counts are strings and abbreviated

    The availability table types likeCount and replyCount as string, and the sample shows "3.4K". Store the raw string as well as anything you parse out of it, because the parse is lossy and cannot be undone.

-   02

    There is no absolute timestamp

    publishedTime is relative to the moment of scraping. Stamp every run with its own time and resolve the phrase against that, or two runs of the same video cannot be reconciled.

-   03

    Replies bill as queries, not items

    The input schema states that with includeReplies on, each reply thread is charged as a query event. That is a different mechanic from the TikTok comments actor, where a reply is an ordinary dataset item.

-   04

    Sort order changes the count

    The listing’s troubleshooting lists sorting among the reasons a run returns fewer comments than expected, top and new can surface different totals on the same video.

-   05

    isPinned and isHearted are not guaranteed

    Both are marked as sometimes available. Treat their absence as unknown rather than false, or a pinned comment with the field missing will be miscounted.

Free accounts run in demo mode: 5 runs a month, 10 items each.

## Six scenarios from the listing

All six reproduce from the two charge events: $0.001 a video with roughly twenty comments inside it, then $0.0005 each. The pattern to notice is that videos are cheap and comments are not.

| Run | Method | Calculation | Cost |
| --- | --- | --- | --- |
| 20 comments | 1 video, the free page | $0.001 | $0.001 |
| 100 comments | 1 video | $0.001 + 80 × $0.0005 | $0.041 |
| 500 comments | 1 video | $0.001 + 480 × $0.0005 | $0.24 |
| 500 comments | 5 videos | 5 × $0.001 + 400 × $0.0005 | $0.21 |
| 1,000 comments | 1 video | $0.001 + 980 × $0.0005 | $0.49 |
| 1,000 comments | 10 videos | 10 × $0.001 + 800 × $0.0005 | $0.41 |

Three of these carry a third decimal in the listing, $0.241, $0.205 and $0.491, and round to the cent here and in the estimator. Turning includeReplies on adds $0.001 for every thread fetched, which is not part of these figures because the number of threads is a property of the video rather than of your input.

## Who reads a comment section at volume?

The listing names sentiment analysis, brand monitoring, market research and NLP training data. What unites them is wanting text in bulk, which is fortunate, because text is the one field here that arrives clean.

-   NLP and training data

    Top comments, many videos

    Collect comment text in volume, where the string types cost nothing because the text is the payload

    $0.41 per 1,000 across 10 videos

-   Brand monitoring

    Newest first, scheduled

    Watch what arrives under a launch video, using sort: new so recency is not buried by relevance

    $0.24 per 500

-   Creator analytics

    Top comments, one video

    Find what the creator hearted and pinned, which is their own read on what mattered

    $0.49 per 1,000

-   Audience research

    A batch, then group by author.id

    Track individual commenters across a channel’s uploads using the one stable identifier in the record

    $0.041 per 100-comment probe

-   Shorts research

    Top comments, Shorts URLs

    Read Shorts comment sections at the same rate as long-form, the listing benchmarks Shorts slightly faster

    $0.001 per Short sampled

## Parsing what comes back

The actor is apidojo/youtube-comments-scraper. Write the count parser and the timestamp resolver once, at ingest, and keep the originals, everything downstream depends on getting that layer right.

**Python**

**JavaScript**

**cURL**

```
import re
from datetime import datetime, timezone
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

def parse_count(s):
    """'3.4K' -> 3400. Lossy on purpose: keep the original alongside."""
    s = (s or "0").strip()
    mult = {"K": 1_000, "M": 1_000_000, "B": 1_000_000_000}.get(s[-1:].upper(), 1)
    return int(float(re.sub(r"[^\d.]", "", s) or 0) * mult)

run = client.actor("apidojo/youtube-comments-scraper").call(run_input={
    "startUrls": ["https://youtu.be/Y9Um-8nPnVQ"],
    "sort": "top",
    "maxItems": 1000,
})

scraped_at = datetime.now(timezone.utc)          # relative times need an anchor
rows = []
for c in client.dataset(run.default_dataset_id).iterate_items():
    rows.append({
        **c,
        "likes_n": parse_count(c["likeCount"]),   # derived
        "likes_raw": c["likeCount"],              # original, never overwritten
        "scraped_at": scraped_at.isoformat(),
    })

for c in sorted(rows, key=lambda r: r["likes_n"], reverse=True)[:10]:
    print(c["likes_raw"], c["author"]["name"], c["text"][:60])
```

```
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });

const run = await client.actor('apidojo/youtube-comments-scraper').call({
  startUrls: ['https://youtu.be/Y9Um-8nPnVQ'],
  sort: 'new',
  maxItems: 500,
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();

// isPinned and isHearted are documented as sometimes-present: absence is
// unknown, not false.
const hearted = items.filter((c) => c.isHearted === true);
const known = items.filter((c) => c.isHearted !== undefined);
console.log(`${hearted.length} hearted of ${known.length} where the flag was present`);

// author.id is the only stable key for following someone across videos.
const repeat = Object.entries(Object.groupBy(items, (c) => c.author.id))
 .filter(([, cs]) => cs.length > 1);
console.log(repeat.length, 'commenters appeared more than once');
```

```
curl -X POST \
  "https://api.apify.com/v2/acts/apidojo~youtube-comments-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"startUrls":["https://youtu.be/Y9Um-8nPnVQ"],"sort":"top","maxItems":1000}'
```

The listing benchmarks 251 comments a second on videos and 267 on Shorts. Export as JSON or CSV, JSON keeps the nested author block and the thumbnail array intact.

## YouTube Comments Scraper FAQ

### How much does it cost to scrape YouTube comments?

$0.001 for the video query, which includes roughly the first twenty comments, then $0.0005 each. A thousand comments from one video is $0.49; spread across ten videos it is $0.41.

### Are the like counts numbers?

No. The listing types likeCount and replyCount as strings, and they arrive in YouTube’s abbreviated form, "3.4K" rather than 3400. Parse them at ingest and keep the original string.

### Does the output include the date a comment was posted?

Only as a relative phrase such as "5 months ago", measured from when the run happened. There is no absolute timestamp, so record your own run time and resolve against it.

### How are replies charged?

As query events. The input schema states that with includeReplies on, each reply thread is charged as a query event per thread, not as ordinary dataset items.

### Can I scrape comments on YouTube Shorts?

Yes. startUrls takes both video and Shorts URLs, and the listing benchmarks Shorts slightly faster at 267 comments a second against 251.

### What is the difference between top and new?

top returns what YouTube ranks as most relevant, new returns newest first. They cost the same, and the listing notes they can surface different comment counts on the same video.

### Can I tell which video a comment came from?

Not from the row, the guaranteed fields cover the comment and its author only. Run one video per job, or add the source URL yourself with customMapFunction.

### Are isPinned and isHearted always there?

No. The listing marks both as sometimes available, so treat a missing field as unknown rather than as false.

## Where YouTube Comments Scraper fits

[Parent scraper YouTube Scraper All seven YouTube scrapers and what each is for. This is the comment one, the only member of the family whose output is YouTube’s own rendered text rather than typed values, which changes how you store it.](/scrapers/youtube-scraper/)

Sibling scrapers

-   [YouTube Channel Scraper](/scrapers/youtube-channel-information-scraper/)
-   [YouTube Transcript Scraper](/scrapers/youtube-transcript-scraper/)
-   [YouTube Playlist Scraper](/scrapers/youtube-playlist-scraper/)

Guides

-   Parsing YouTube’s abbreviated counts without losing data*Coming soon*
-   Anchoring relative timestamps across scheduled runs*Coming soon*
-   Tracking a commenter across a channel with author.id*Coming soon*

## Read one comment section

A tenth of a cent, twenty comments in. Shorts included.

`startUrls: [one video URL] · sort: "top"`

[Run this scraper free](https://apify.com/apidojo/youtube-comments-scraper)

Apidojo is not affiliated with YouTube or Google LLC.

## Guides for this scraper

-   [YouTube scraping guides Videos, channels, playlists, comments and transcripts: what YouTube's data actually contains, and which fields arrive as text rather than as numbers.](/blog/youtube/)
-   [Parsing YouTube's abbreviated counts YouTube Comments Scraper returns likeCount as the string "3.4K". Two other YouTube actors return integers. Here is how to handle both without lying.](/blog/youtube/youtube-count-strings/)
