How to rebuild an X conversation tree from replies
A reply scrape returns a flat array, not a tree. Group the rows by inReplyToId to get each comment's direct children, use conversationId to find the root, and you have the thread back. The two fields are not interchangeable, and that is where most reconstructions go wrong.

Why the output is flat
Every reply scraper returns rows, not a tree. That is the right choice: a tree is one of many shapes you might want, and a flat array is the only shape that survives being written to CSV, loaded into a warehouse or streamed through a queue.
The threading is not missing, though. It is carried on two fields, and the whole job is knowing which does what.
The two fields that carry the structure
conversationId is the tweet that started the thread. It is identical on
every row in that conversation, including the root itself. It answers “which
discussion is this part of”.
inReplyToId is the tweet that this particular row replies to. On a
top-level reply that is the root. On a reply to a reply it is the other
reply. It answers “what is directly above me”.
That distinction is the entire article. Group by conversationId and you get
one flat bucket per thread, which is what you already had. Group by
inReplyToId and you get an adjacency list, which is a tree.
| Field | Same for every row in a thread | Use it to |
|---|---|---|
conversationId |
Yes | Find the root, split a multi-thread pull |
inReplyToId |
No | Build the parent and child edges |
Building the tree in one pass
Group the rows once, then walk down from the root. There is no need to sort first and no need for recursion over the whole set.
from collections import defaultdict
def build_tree(rows):
children = defaultdict(list)
for row in rows:
children[row.get("inReplyToId")].append(row)
root_id = rows[0]["conversationId"]
known = {row["id"] for row in rows}
# A reply whose parent never arrived would vanish from the walk. Hang it
# off the root instead: an orphan in the thread beats a missing row.
for parent_id, group in list(children.items()):
if parent_id and parent_id not in known and parent_id != root_id:
children[root_id].extend(group)
del children[parent_id]
def walk(node_id, depth=0):
for child in children.get(node_id, []):
yield depth, child
yield from walk(child["id"], depth + 1)
return list(walk(root_id))
for depth, reply in build_tree(rows):
print(" " * depth, reply["author"]["userName"], reply["text"][:60])
The orphan step is the part people skip. A capture is always partial: a ceiling cut the run short, a parent was deleted, an account went private between the reply and the scrape. Dropping those rows silently is how a dataset ends up with fewer replies than the tweet claims to have.
Check the shape before you trust it
The root tweet carries its own replyCount. Compare it against the number of
rows you actually received. A small gap is normal, since deleted and hidden
replies are counted by X but not returned. A large one means something is
wrong with the run rather than with the thread.
On very large or old threads the default collection flow can come back thin. That is what the search-based flow exists for, and it is worth one re-run before you conclude a conversation was quiet.
Questions people ask
What is the difference between conversationId and inReplyToId?
conversationId is the ID of the tweet that started the whole thread and is the same on every row in it. inReplyToId is the ID of the specific tweet that one row answers, which may be the root or may be another reply. Using conversationId to build the tree collapses every reply to depth one.
Why do some replies have a parent that is not in my data?
A thread is only ever a partial capture. If a reply's parent was not returned, whether because of a ceiling, a deletion or a protected account, that row is an orphan. Re-parent orphans to the root so they stay in the dataset rather than disappearing from your counts.
How do I know if I got the whole thread?
Compare the number of rows you received against the root tweet's own replyCount. A large gap usually means the ceiling was too low, or that the thread is big enough to need the search-based collection flow instead of the default one.