When the X replies endpoint under-returns

If a reply scrape returns fewer rows than the tweet claims, check maxItems first, because a low ceiling looks exactly like a quiet thread. If the ceiling was not the problem, switch to the search-based flow, which queries the conversation directly and reaches replies the standard endpoint misses on large or old threads.

Cover card reading "When the X replies endpoint under-returns"

Three reasons a thread comes back thin

A reply scrape that returns forty rows for a tweet showing four hundred replies is not necessarily broken. There are three explanations, and they are worth ruling out in this order because that is the order of how cheap they are to check.

The ceiling. maxItems caps the number of items for the whole run, not per tweet. A batch of ten tweets under a ceiling of 200 shares those 200 rows between them, so the first tweet looks complete and the rest look empty. This is the single most common cause and costs nothing to rule out.

The counter. X counts replies that a scrape cannot return. Deleted replies, replies from suspended or protected accounts, and replies hidden by the author all contribute to replyCount without ever appearing in a result set. A gap of a few percent is normal and permanent.

The endpoint. The standard replies endpoint walks the conversation the way the app does. On very large or very old threads that walk can stop short of what the conversation actually contains.

Measuring before you spend

The root tweet carries its own replyCount, so the check is arithmetic rather than judgement:

rows = replies(tweet_id)                    # default flow
root = next(r for r in rows if not r.get("inReplyToId"))
got, claimed = len(rows), root["replyCount"]

print(f"{got} of {claimed} ({got / claimed:.0%})")

A result in the high nineties is a complete capture. Somewhere in the eighties is normal attrition. Below about half, something is wrong with the run rather than with the thread.

What the second flow actually changes

Setting useSearch to true stops the actor using the replies endpoint. It runs a conversation_id: search instead, which queries the conversation as a corpus rather than walking it as a tree.

The practical difference is coverage on the awkward cases. The listing is explicit that the search route can retrieve results not available through the standard endpoint, and names large and old threads as where that happens.

Default flow Search flow
Cost per tweet $0.014 $0.016
Items included about 40 about 40
Route replies endpoint conversation_id search
Best on ordinary threads large or old threads

Fourteen percent more per tweet sounds like a decision. On a single deep thread it is two tenths of a cent, which is not a decision at all. Across fifty tweets it is a tenth of a dollar, which is the price of not missing a conversation you already paid to collect.

The order that saves money

Run the cheap flow. Compare the count. Raise the ceiling if it was binding and run again. Only then pay the premium, and only for the tweets that actually came back short, not for the whole batch.

Questions people ask

Why did my reply scrape return fewer replies than X shows?

Usually the run ceiling. maxItems applies across the entire run rather than per tweet, so a batch shares it. If the ceiling was not binding, the thread may be large or old enough that the standard replies endpoint under-returns, which is what the search-based flow exists for.

Does the search flow cost more?

Yes, $0.016 a tweet against $0.014 on the default flow, which is about 14% more. Both include roughly the first 40 items per query on a paid plan, so on a deep thread the difference is a rounding error and on a wide batch it is a few cents.

Should I just always use the search flow?

No. It is the second attempt, not the default. Run the cheaper flow, measure what came back against the tweet's replyCount, and switch only when the gap is real.

Sources

  1. Twitter Replies Scraper, flows and pricingapify.com