LeadLag LDA: Estimating Topic Specific Leads and Lags of Information Outlets

Ramesh Nallapati; Xiaolin Shi; Daniel McFarland; Jure Leskovec; Daniel Jurafsky

doi:10.1609/icwsm.v5i1.14147

Authors

Ramesh Nallapati Stanford University
Xiaolin Shi Stanford University
Daniel McFarland Stanford University
Jure Leskovec Stanford University
Daniel Jurafsky Stanford University

DOI:

https://doi.org/10.1609/icwsm.v5i1.14147

Abstract

Identifying which outlet in social media leads the rest in disseminating novel information on specific topics is an interesting challenge for information analysts and social scientists. In this work, we hypothesize that novel ideas are disseminated through the creation and propagation of new or newly emphasized key words, and therefore lead/lag of outlets can be estimated by tracking word usage across these outlets. First, we demonstrate the validaty of our hypothesis by showing that a simple TF-IDF based nearest-neighbors approach can recover generally accepted lead/lag behavior on the outlets pair of ACM journal articles and conference papers. Next, we build a new topic model called LeadLag LDA that estimates the lead/lag of the outlets on specific topics. We validate the topic model using the lead/lag results from the TF-IDF nearest neighbors approach. Finally, we present results from our model on two different outlet pairs of blogs vs. news media and grant proposals vs. research publications that reveal interesting patterns.

LeadLag LDA: Estimating Topic Specific Leads and Lags of Information Outlets

Authors

DOI:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information