How to rebuild an X conversation tree from replies

A reply scrape returns a flat array, not a tree. Group the rows by inReplyToId to get each comment's direct children, use conversationId to find the root, and you have the thread back. The two fields are not interchangeable, and that is where most reconstructions go wrong.

Cover card reading "How to rebuild an X conversation tree from replies"

Why the output is flat

Every reply scraper returns rows, not a tree. That is the right choice: a tree is one of many shapes you might want, and a flat array is the only shape that survives being written to CSV, loaded into a warehouse or streamed through a queue.

The threading is not missing, though. It is carried on two fields, and the whole job is knowing which does what.

The two fields that carry the structure

conversationId is the tweet that started the thread. It is identical on every row in that conversation, including the root itself. It answers “which discussion is this part of”.

inReplyToId is the tweet that this particular row replies to. On a top-level reply that is the root. On a reply to a reply it is the other reply. It answers “what is directly above me”.

That distinction is the entire article. Group by conversationId and you get one flat bucket per thread, which is what you already had. Group by inReplyToId and you get an adjacency list, which is a tree.

Field Same for every row in a thread Use it to
conversationId Yes Find the root, split a multi-thread pull
inReplyToId No Build the parent and child edges

Building the tree in one pass

Group the rows once, then walk down from the root. There is no need to sort first and no need for recursion over the whole set.

from collections import defaultdict

def build_tree(rows):
    children = defaultdict(list)
    for row in rows:
        children[row.get("inReplyToId")].append(row)

    root_id = rows[0]["conversationId"]
    known = {row["id"] for row in rows}

    # A reply whose parent never arrived would vanish from the walk. Hang it
    # off the root instead: an orphan in the thread beats a missing row.
    for parent_id, group in list(children.items()):
        if parent_id and parent_id not in known and parent_id != root_id:
            children[root_id].extend(group)
            del children[parent_id]

    def walk(node_id, depth=0):
        for child in children.get(node_id, []):
            yield depth, child
            yield from walk(child["id"], depth + 1)

    return list(walk(root_id))

for depth, reply in build_tree(rows):
    print("  " * depth, reply["author"]["userName"], reply["text"][:60])

The orphan step is the part people skip. A capture is always partial: a ceiling cut the run short, a parent was deleted, an account went private between the reply and the scrape. Dropping those rows silently is how a dataset ends up with fewer replies than the tweet claims to have.

Check the shape before you trust it

The root tweet carries its own replyCount. Compare it against the number of rows you actually received. A small gap is normal, since deleted and hidden replies are counted by X but not returned. A large one means something is wrong with the run rather than with the thread.

On very large or old threads the default collection flow can come back thin. That is what the search-based flow exists for, and it is worth one re-run before you conclude a conversation was quiet.

Questions people ask

What is the difference between conversationId and inReplyToId?

conversationId is the ID of the tweet that started the whole thread and is the same on every row in it. inReplyToId is the ID of the specific tweet that one row answers, which may be the root or may be another reply. Using conversationId to build the tree collapses every reply to depth one.

Why do some replies have a parent that is not in my data?

A thread is only ever a partial capture. If a reply's parent was not returned, whether because of a ceiling, a deletion or a protected account, that row is an orphan. Re-parent orphans to the root so they stay in the dataset rather than disappearing from your counts.

How do I know if I got the whole thread?

Compare the number of rows you received against the root tweet's own replyCount. A large gap usually means the ceiling was too low, or that the thread is big enough to need the search-based collection flow instead of the default one.

Sources

  1. Twitter Replies Scraper, input and output referenceapify.com
  2. Twitter Scraper Unlimited, conversation fieldsapify.com