YouTube Comments Scraper: videos and Shorts

YouTube Comments Scraper reads the comments on any public YouTube video or Short and returns each one with its author, its like and reply counts, and when it was posted.

apidojo/youtube-comments-scraper

Estimated run cost

$0.00

Free Apify plan: 5 runs a month, 10 items a run. No credit card.

$0.001
per video queried
~20
comments included in that
$0.0005
per comment after
99.99%
run success, last 30 days

Read on 27 September 2026.

Five fields, and the two that change the answer

The URL decides what you read and the ceiling decides how much. Between them sit a sort order that changes which comments you get, and a switch that changes how you are billed.

  • startUrls

    $0.001per video

    ~20 comments free

    A video or Shorts URL. Both are first-class here, the listing benchmarks them separately at 251 and 267 comments a second.

    https://youtu.be/Y9Um-8nPnVQ

  • sort

    top | newfree to change

    Defaults to top

    top returns what YouTube considers most relevant; new returns newest first. Same price, different comments.

    sort: "new"

  • Bills differently from every other comment scraper here

    includeReplies

    +$0.001per thread

    A query, not an item

    Off by default. On, each top-level comment’s reply thread is fetched, and the schema says each thread is charged as its own query event.

    includeReplies: true

  • maxItems

    ÷ $0.0005past the free 20

    Run-wide

    Comments across the whole run. The listing’s troubleshooting names it first when a run returns less than expected.

    1000

Sorting is worth a moment: the listing notes that top and new may surface different comment counts on the same video, so a thin result is sometimes the sort order rather than the ceiling.

YouTube Comments Scraper specificationsapidojo/youtube-comments-scraper
Apify actorapidojo/youtube-comments-scraper
InputsstartUrls · sort · includeReplies · maxItems · customMapFunction
AcceptsVideo and Shorts URLs
Sortingtop for most relevant, new for newest first
CountsStrings, not numbers, likeCount arrives as "3.4K"
TimestampsRelative only, "5 months ago", with no absolute date
RepliesCharged as a query event per thread, not as items
Reliability99.99% of 118,137 runs succeeded in the last 30 days

Most relevant, most recent, or everything under a batch

Three runs. The first two cost the same and answer different questions; the third is where the per-video free allowance starts to matter more than the per-comment rate.

How do you scrape comments from a YouTube video?

Paste the URL and leave sort at top. The video query is $0.001 and brings about twenty comments; after that each is $0.0005. A thousand comments from one video costs $0.49.

Input

{
  "startUrls": ["https://youtu.be/Y9Um-8nPnVQ"],
  "sort": "top",
  "maxItems": 1000
}

$0.001 + 980 × $0.0005=$0.49

What comes back

The comment record, Author, text, counts and a relative time, all seven guaranteed fields.

{
  "text": "This one hit you in the feels? Check out the full video…",
  "likeCount": "3.4K",
  "replyCount": "8",
  "publishedTime": "5 months ago",
  "author": {
    "id": "UCa90xqK2odw1KV5wHU9WRhg",
    "name": "@TheOffice"
  }
}

likeCount is the string "3.4K", not the number 3400

The listing’s field table gives likeCount and replyCount as strings, and the sample shows the abbreviated form. Sorting or thresholding on them means parsing first, and the abbreviation has already thrown away the exact figure, "3.4K" could be anything from 3,350 to 3,449.

How do you get the newest YouTube comments?

Set sort to new. It costs exactly the same as top and is the right choice for monitoring, because relevance ranking will bury a comment posted an hour ago beneath one from last year.

Input

{
  "startUrls": ["https://youtu.be/Y9Um-8nPnVQ"],
  "sort": "new",
  "maxItems": 500
}

$0.001 + 480 × $0.0005=$0.24

What comes back

The conditional fields, Two flags the listing marks as sometimes present, not always.

{
  "isPinned": true,
  "isHearted": false
}

Relative times cannot be compared across runs

"5 months ago" is measured from the moment of scraping, so the same comment reads differently next week and two runs cannot be merged on it. Record your own run timestamp and resolve the phrase against it, or the dataset loses its chronology.

How do you scrape comments across many videos?

List the URLs together. Each video carries its own free twenty, so ten videos sharing a thousand comments cost $0.41 where one video taking a thousand alone costs $0.49, breadth is slightly subsidised.

Input

{
  "startUrls": [
    "https://youtu.be/Y9Um-8nPnVQ",
    "https://www.youtube.com/shorts/abc123"
  ],
  "sort": "top",
  "maxItems": 1000
}

10 × $0.001 + 800 × $0.0005=$0.41

What comes back

The author block, A channel ID on every comment, which is what makes cross-video tracking possible.

{
  "author": {
    "id": "UCa90xqK2odw1KV5wHU9WRhg",
    "name": "@TheOffice",
    "thumbnails": [
      { "height": 48, "width": 48, "url": "https://yt3.ggpht.com/…" }
    ]
  }
}

Nothing on a row names the video it came from

The guaranteed fields cover the comment and its author, not the source URL. On a multi-video run, either submit one video per run or add the URL yourself through customMapFunction, which exists for exactly this kind of reshaping.

Seven fields guaranteed, two conditional

The listing publishes an availability table rather than just a sample, which is unusually careful and worth taking literally: seven fields are marked always present, two are marked sometimes, and the types are not what you would assume.

  • 4

    Always present

    The comment

    • text, the full body
    • likeCount, string
    • replyCount, string
    • publishedTime, relative string
  • 3

    Always present

    author

    • id, the channel ID
    • name, e.g. @TheOffice
    • thumbnails, array, each with url, width, height
  • 2

    Sometimes present

    Marked conditional

    • isPinned, boolean
    • isHearted, creator hearted it
  • 0

    Not returned

    Plan around these

    • No absolute timestamp
    • No numeric engagement
    • No source video on the row
Field referenceDocumented in the actor README
textThe comment body, newlines and links intact
likeCountA string in YouTube’s own abbreviated form, e.g. "3.4K", parse before you compare
replyCountAlso a string. "8" is a string eight, not an integer
publishedTimeRelative phrase such as "5 months ago", measured from when the run happened
author.idThe commenter’s channel ID, the only stable key for tracking someone across videos
author.nameDisplay handle, usually @-prefixed
author.thumbnailsArray of avatar sizes, each with url, width and height
isPinnedPresent sometimes. Pinned comments are usually the creator’s own
isHeartedPresent sometimes. The creator’s heart is the cheapest endorsement signal on the platform

author.id is the field to build on. Names change and display handles are not unique, but the channel ID lets you follow one commenter across every video you have ever scraped, which is how repeat critics and repeat advocates get found.

What to settle before you store any of it

Four of these come from the listing’s own documentation of its output and inputs. They matter at schema-design time rather than at run time.

  • 01

    Counts are strings and abbreviated

    The availability table types likeCount and replyCount as string, and the sample shows "3.4K". Store the raw string as well as anything you parse out of it, because the parse is lossy and cannot be undone.

  • 02

    There is no absolute timestamp

    publishedTime is relative to the moment of scraping. Stamp every run with its own time and resolve the phrase against that, or two runs of the same video cannot be reconciled.

  • 03

    Replies bill as queries, not items

    The input schema states that with includeReplies on, each reply thread is charged as a query event. That is a different mechanic from the TikTok comments actor, where a reply is an ordinary dataset item.

  • 04

    Sort order changes the count

    The listing’s troubleshooting lists sorting among the reasons a run returns fewer comments than expected, top and new can surface different totals on the same video.

  • 05

    isPinned and isHearted are not guaranteed

    Both are marked as sometimes available. Treat their absence as unknown rather than false, or a pinned comment with the field missing will be miscounted.

Free accounts run in demo mode: 5 runs a month, 10 items each.

Six scenarios from the listing

All six reproduce from the two charge events: $0.001 a video with roughly twenty comments inside it, then $0.0005 each. The pattern to notice is that videos are cheap and comments are not.

What a run costsQuery fee + $0.0005 × billable items
RunMethodCalculationCost
20 comments1 video, the free page$0.001$0.001
100 comments1 video$0.001 + 80 × $0.0005$0.041
500 comments1 video$0.001 + 480 × $0.0005$0.24
500 comments5 videos5 × $0.001 + 400 × $0.0005$0.21
1,000 comments1 video$0.001 + 980 × $0.0005$0.49
1,000 comments10 videos10 × $0.001 + 800 × $0.0005$0.41

Three of these carry a third decimal in the listing, $0.241, $0.205 and $0.491, and round to the cent here and in the estimator. Turning includeReplies on adds $0.001 for every thread fetched, which is not part of these figures because the number of threads is a property of the video rather than of your input.

Who reads a comment section at volume?

The listing names sentiment analysis, brand monitoring, market research and NLP training data. What unites them is wanting text in bulk, which is fortunate, because text is the one field here that arrives clean.

  • NLP and training data

    Top comments, many videos

    Collect comment text in volume, where the string types cost nothing because the text is the payload

    $0.41 per 1,000 across 10 videos

  • Brand monitoring

    Newest first, scheduled

    Watch what arrives under a launch video, using sort: new so recency is not buried by relevance

    $0.24 per 500

  • Creator analytics

    Top comments, one video

    Find what the creator hearted and pinned, which is their own read on what mattered

    $0.49 per 1,000

  • Audience research

    A batch, then group by author.id

    Track individual commenters across a channel’s uploads using the one stable identifier in the record

    $0.041 per 100-comment probe

  • Shorts research

    Top comments, Shorts URLs

    Read Shorts comment sections at the same rate as long-form, the listing benchmarks Shorts slightly faster

    $0.001 per Short sampled

Parsing what comes back

The actor is apidojo/youtube-comments-scraper. Write the count parser and the timestamp resolver once, at ingest, and keep the originals, everything downstream depends on getting that layer right.

import re
from datetime import datetime, timezone
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

def parse_count(s):
    """'3.4K' -> 3400. Lossy on purpose: keep the original alongside."""
    s = (s or "0").strip()
    mult = {"K": 1_000, "M": 1_000_000, "B": 1_000_000_000}.get(s[-1:].upper(), 1)
    return int(float(re.sub(r"[^\d.]", "", s) or 0) * mult)

run = client.actor("apidojo/youtube-comments-scraper").call(run_input={
    "startUrls": ["https://youtu.be/Y9Um-8nPnVQ"],
    "sort": "top",
    "maxItems": 1000,
})

scraped_at = datetime.now(timezone.utc)          # relative times need an anchor
rows = []
for c in client.dataset(run.default_dataset_id).iterate_items():
    rows.append({
        **c,
        "likes_n": parse_count(c["likeCount"]),   # derived
        "likes_raw": c["likeCount"],              # original, never overwritten
        "scraped_at": scraped_at.isoformat(),
    })

for c in sorted(rows, key=lambda r: r["likes_n"], reverse=True)[:10]:
    print(c["likes_raw"], c["author"]["name"], c["text"][:60])
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });

const run = await client.actor('apidojo/youtube-comments-scraper').call({
  startUrls: ['https://youtu.be/Y9Um-8nPnVQ'],
  sort: 'new',
  maxItems: 500,
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();

// isPinned and isHearted are documented as sometimes-present: absence is
// unknown, not false.
const hearted = items.filter((c) => c.isHearted === true);
const known = items.filter((c) => c.isHearted !== undefined);
console.log(`${hearted.length} hearted of ${known.length} where the flag was present`);

// author.id is the only stable key for following someone across videos.
const repeat = Object.entries(Object.groupBy(items, (c) => c.author.id))
 .filter(([, cs]) => cs.length > 1);
console.log(repeat.length, 'commenters appeared more than once');
curl -X POST \
  "https://api.apify.com/v2/acts/apidojo~youtube-comments-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"startUrls":["https://youtu.be/Y9Um-8nPnVQ"],"sort":"top","maxItems":1000}'

The listing benchmarks 251 comments a second on videos and 267 on Shorts. Export as JSON or CSV, JSON keeps the nested author block and the thumbnail array intact.

YouTube Comments Scraper FAQ

How much does it cost to scrape YouTube comments?

$0.001 for the video query, which includes roughly the first twenty comments, then $0.0005 each. A thousand comments from one video is $0.49; spread across ten videos it is $0.41.

Are the like counts numbers?

No. The listing types likeCount and replyCount as strings, and they arrive in YouTube’s abbreviated form, "3.4K" rather than 3400. Parse them at ingest and keep the original string.

Does the output include the date a comment was posted?

Only as a relative phrase such as "5 months ago", measured from when the run happened. There is no absolute timestamp, so record your own run time and resolve against it.

How are replies charged?

As query events. The input schema states that with includeReplies on, each reply thread is charged as a query event per thread, not as ordinary dataset items.

Can I scrape comments on YouTube Shorts?

Yes. startUrls takes both video and Shorts URLs, and the listing benchmarks Shorts slightly faster at 267 comments a second against 251.

What is the difference between top and new?

top returns what YouTube ranks as most relevant, new returns newest first. They cost the same, and the listing notes they can surface different comment counts on the same video.

Can I tell which video a comment came from?

Not from the row, the guaranteed fields cover the comment and its author only. Run one video per job, or add the source URL yourself with customMapFunction.

Are isPinned and isHearted always there?

No. The listing marks both as sometimes available, so treat a missing field as unknown rather than as false.

Where YouTube Comments Scraper fits

Read one comment section

A tenth of a cent, twenty comments in. Shorts included.

startUrls: [one video URL] · sort: "top"

Run this scraper free

Apidojo is not affiliated with YouTube or Google LLC.