Parsing YouTube's abbreviated counts

On YouTube Comments Scraper, likeCount and replyCount are typed as strings and arrive in YouTube's own abbreviated form, so "3.4K" could be anything from 3,350 to 3,449. Parse it at ingest, keep the original alongside, and never let the derived number be the only copy.

Cover card reading "Parsing YouTube's abbreviated counts without losing data"

One platform, two conventions

This is the thing to settle before you design a YouTube schema, because getting it wrong produces a pipeline that works on one actor and silently corrupts another.

Actor views / likes duration timestamps
Comments strings, abbreviated n/a relative phrase
Channel Videos integers seconds, integer n/a in the sample
Playlist integers seconds, integer n/a in the sample

The comments actor hands back what YouTube painted on the screen. The other two hand back data. Both are documented, neither is wrong, and a team using both will write one parser and apply it twice unless somebody says this out loud.

Abbreviation is lossy and permanent

“3.4K” is not 3,400. It is somewhere between roughly 3,350 and 3,449, and the information that would let you narrow it was discarded before you received the string.

That has a practical consequence: the parsed number must never be the only copy you keep. If you store 3,400 and throw away “3.4K”, you have converted a known approximation into a false precision, and six months later nobody will remember which columns were exact.

import re

def parse_count(s):
    """'3.4K' -> 3400. Approximate by construction."""
    s = (s or "0").strip()
    mult = {"K": 1_000, "M": 1_000_000, "B": 1_000_000_000}.get(s[-1:].upper(), 1)
    return int(float(re.sub(r"[^\d.]", "", s) or 0) * mult)

row = {
    **comment,
    "likes_n": parse_count(comment["likeCount"]),  # derived, approximate
    "likes_raw": comment["likeCount"],             # original, never overwritten
}

Two columns. One you sort on, one you can defend.

The timestamp problem is worse

publishedTime is a relative phrase: “5 months ago”. It is measured from the moment of the scrape, which means the same comment reads differently every time you run it, and two runs cannot be merged on it at all.

There is no absolute timestamp in the record. So the run’s own time has to come from you:

scraped_at = datetime.now(timezone.utc)
rows = [{**c, "scraped_at": scraped_at.isoformat()} for c in items]

Without that, a comment dataset has no chronology. With it, “5 months ago” resolves to a real date and stays resolved.

Where to be strict

Use the parsed numbers for ranking, filtering and thresholds. That is what they are good for and the error does not matter when you are asking which comments are the most liked.

Do not use them for totals you report to somebody. Summing four hundred abbreviated counts compounds four hundred roundings, and the answer looks authoritative in a way it has not earned.

When a figure has to be exact, the video’s own engagement from the Channel Videos actor is an integer, and that is the number to quote.

Questions people ask

Are YouTube comment like counts numbers?

Not on the comments actor. Its field availability table types likeCount and replyCount as strings, and the sample shows the abbreviated form: "3.4K" for likes and "8" for replies. The eight is a string eight.

Do all the YouTube scrapers do this?

No, and that is the trap. The Channel Videos and Playlist actors type views, likes and duration as numbers and mark them always present. Same platform, same publisher, opposite convention.

How precise is an abbreviated count?

Not very. "3.4K" covers everything from about 3,350 to 3,449, so any figure you derive from it carries roughly a one percent error. That is fine for ranking and wrong for reporting an exact total.

Sources

  1. YouTube Comments Scraper, field availabilityapify.com
  2. YouTube Channel Scraper, output exampleapify.com