Reading account age in X reply data

Each reply arrives with an author block holding the account's creation date, follower and following counts, and lifetime post and media totals. Together those make a useful filter for weighting a conversation, but none of them identifies a bot on its own, and treating any single one as proof produces confident nonsense.

Cover card reading "Reading account age in X reply data"

What arrives for free

A reply scrape returns the responder’s account alongside the reply itself. That matters for cost as much as for analysis: the account data is already paid for, so filtering a conversation by who is in it costs nothing beyond the scrape you were doing anyway.

The author block carries the handle and display name, the account’s creation date, follower and following counts, lifetime post and media counts, the favourites count, and both verification booleans.

Age on its own says nothing

An account created last week is not suspicious. Accounts are created constantly, by people. What makes age informative is pairing it with volume, because the ratio between the two describes a behaviour rather than a property.

from datetime import datetime, timezone

def posts_per_day(author):
    created = datetime.strptime(author["createdAt"], "%a %b %d %H:%M:%S %z %Y")
    days = max((datetime.now(timezone.utc) - created).days, 1)
    return author["statusesCount"] / days

# A two-week-old account with 4,000 posts is a different object from a
# two-week-old account with nine.

Three hundred posts a day sustained over a year is not a person typing. Nine posts over two weeks is somebody who joined and has not decided whether they like it. Age alone cannot separate those; age against volume does it in one line.

The ratio everyone reaches for

Followers divided by following is the most commonly cited signal and the weakest one, because it is the easiest to manipulate and the most affected by the account simply being new. It is worth collecting and worth distrusting.

Signal What it suggests Why it is weak alone
Account age How established the account is New accounts are mostly ordinary people
Posts per day Sustained automation High-volume humans exist
Follower ratio Reciprocity of the account Trivially gamed, and skewed by age
Media count Whether it posts original content Text-only accounts are normal

Read together they describe a shape. Read individually they produce confident nonsense, and the confident nonsense is worse than no filter at all because it survives into a summary that somebody acts on.

Weighting, not classifying

The honest use of these fields is weighting. When a thread’s sentiment is being summarised, a reply from an account that has existed for six years and posts twice a week is evidence about an audience in a way that a reply from a two-day-old account posting hourly is not.

That is a scale, not a boundary. Nothing in a single scrape supports drawing a line and calling everything on one side of it automated. The data is good enough to rank a conversation by how much each voice should count, and it is not good enough to accuse anybody of anything.

Questions people ask

Which account fields come with a reply?

The author block carries userName, name, id, url, the account creation date, follower and following counts, lifetime statuses and media counts, favourites count, and both verification flags. No second request is needed to get them.

Can I detect bots with this data?

Not reliably, and this post does not claim you can. What these fields support is weighting: deciding how much a reply should count in an aggregate. Classification needs behavioural data over time that a single scrape does not contain.

What is the difference between isVerified and isBlueVerified?

They are separate booleans in the output and mean different things. Treat them as two distinct signals rather than collapsing them into one "verified" field.

Sources

  1. Twitter Replies Scraper, output object referenceapify.com
  2. Twitter User Scraper, account fieldsapify.com