Comment-level Topic Drift Analysis in the Reddit Corpus
We present a novel application of embedding-based dynamic topic modeling techniques to detect and quantify topic drift at the comment level in a massive corpus. By leveraging pretrained language models to generate contextualized semantic embeddings for short text, we analyzed 12.7 billion Reddit comments spanning 2006 to 2022. Using unsupervised methods on these embeddings, we identify dynamically evolving topic clusters over time. Our primary contribution is a methodology for analysis of semantic drift and discourse evolution in the embedding space itself. We also demonstrate modifications to
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Steven Morse →
“Comment-level Topic Drift Analysis in the Reddit Corpus”
- LinkedLinked via arxiv author · 85%Daniel Runfola →
“Comment-level Topic Drift Analysis in the Reddit Corpus”
- LinkedLinked via arxiv author · 85%Trenton W. Ford →
“Comment-level Topic Drift Analysis in the Reddit Corpus”
- FuzzySimilar title/name (fuzzy) · 87%MaartenGr/BERTopic →
“Fuzzy title match (0.94): “Comment-level Topic Drift Analysis in the Reddit Corpus” ≈ “MaartenGr/BERTopic””
