Open science
Datasets
This page is a collection of datasets used in my research activity.
When you contact me for data requests, please provide information about your academic status (your home institution, your current position) and explain in brief how you will use the data. Datasets won’t be shared with parties for commercial purposes. Two terms of usage apply:- The appropriate papers are cited in any research product based on these datasets.
- The datasets cannot be redistributed without obtaining my permission.
Eating Disorder Content on TikTok (EDTok)
TikTok video IDs for content related to eating disorders, collected through TikTok’s Research API. Only the video IDs are redistributed, to comply with TikTok’s terms of service. Collected by Charles Bickham.Dataset: https://github.com/cbickham3232/EDTok-A-Dataset-for-Eating-Disorder-Content-on-TikTok
Paper: EDTok: A Dataset for Eating Disorder Content on TikTok
Full dataset page: Eating Disorder Content on TikTok (EDTok)
2024 U.S. Election: Telegram
The largest public Telegram dataset on the 2024 U.S. presidential election, covering political discussion in the lead-up to the vote. Collected by Leonardo Blas.Dataset: https://github.com/leonardo-blas/usc-tg-24-us-election
Distilled version: Hugging Face
Paper: Blas, L., Luceri, L., & Ferrara, E. (2025). Unearthing a Billion Telegram Posts about the 2024 U.S. Presidential Election: Development of a Public Dataset. Companion Proceedings of the ACM Web Conference 2025.
Full dataset page: 2024 U.S. Election: Telegram
2024 U.S. Election: Twitter/X
A large-scale collection of discourse on X (formerly Twitter) about the 2024 U.S. presidential election and the dynamics of political engagement around it. Collected by Ashwin Balasubramanian.Dataset: https://github.com/sinking8/usc-x-24-us-election
Paper: A Public Dataset Tracking Social Media Discourse about the 2024 U.S. Presidential Election on Twitter/X
Full dataset page: 2024 U.S. Election: Twitter/X
2024 U.S. Election: Truth Social
1.5 million Truth Social posts published between February 2022 and October 2024, capturing posts, comments, and user interactions related to the 2024 U.S. presidential election. Collected by Kashish Shah.Dataset: https://github.com/kashish-s/TruthSocial_2024ElectionInitiative
Mirror: Kaggle
Full dataset page: 2024 U.S. Election: Truth Social
2024 U.S. Election: TikTok
Video IDs and Whisper-generated transcripts for TikTok content related to the 2024 U.S. presidential election, collected through TikTok’s Research API. Only the video IDs are redistributed, to comply with TikTok’s terms of service. Collected by Gabriela Pinto.Dataset: https://github.com/gabbypinto/US2024PresElectionTikToks
Archive: Zenodo (doi:10.5281/zenodo.14868880)
Paper: Tracking the 2024 US Presidential Election Chatter on TikTok: A Public Multimodal Dataset
Full dataset page: 2024 U.S. Election: TikTok
2022 Attempted Coup in Peru on TikTok (GET-Tok)
Multimodal TikTok data documenting the 2022 attempted coup in Peru, enriched with generative AI: Whisper transcriptions and GPT-4 annotations. Collected by Gabriela Pinto.Dataset: https://github.com/gabbypinto/GET-Tok-Peru
Paper: GET-Tok: A GenAI-Enriched Multimodal TikTok Dataset Documenting the 2022 Attempted Coup in Peru
Full dataset page: 2022 Attempted Coup in Peru on TikTok (GET-Tok)
Ukraine and Russia Conflict Tweets
An ongoing collection of tweet IDs on the war between Ukraine and Russia, covering the discourse from 17 February 2022 onward. Only tweet IDs are released, for non-commercial research use. Collected by Emily Chen.Dataset: https://github.com/echen102/ukraine-russia
Paper: Tweets in Time of Conflict: A Public Dataset Tracking the Twitter Discourse on the War Between Ukraine and Russia
Full dataset page: Ukraine and Russia Conflict Tweets
COVID-19 Vaccine Misinformation Labels
Large-scale misinformation-labeled datasets of COVID-19 vaccine discourse on Twitter, constructed by refining weak labels over retweet cascades rather than hand-annotating individual posts.Dataset: https://github.com/USC-Melady/Constructing-Misinformation-Datasets-WWW-2022
Paper: Sharma, K., Ferrara, E., & Liu, Y. (2022). Construction of Large-Scale Misinformation Labeled Datasets from Social Media Discourse using Label Refinement. Proceedings of the ACM Web Conference 2022, 3755–3764.
Full dataset page: COVID-19 Vaccine Misinformation Labels
COVID-19 Tweets
Since January 2020, we collected hundreds of millions of tweets related to COVID-19.Dataset: https://github.com/echen102/COVID-19-TweetIDs
Paper: https://publichealth.jmir.org/2020/2/e19273/
Full dataset page: COVID-19 Tweets
2020 U.S. Election Tweets
We collected hundreds of millions of election-related tweets for the 2020 U.S. Presidential election.Dataset: https://github.com/echen102/us-pres-elections-2020
Paper: https://arxiv.org/abs/2010.00600
Full dataset page: 2020 U.S. Election Tweets
Individual Performance in Team-based Online Games
The Dataset used in this study has been deposited in the Harvard Dataverse repository (doi:10.7910/DVN/B0GRWX), and is available at the following URL: Access the League of Legends Dataset The Code used to obtain the results presented in the paper is openly available in four IPython Jupyter Notebooks:- Performance Modeling: RQ1-RQ3
- Performance Prediction: RQ4-Model1 RQ4-Model2 RQ4-Model3
Full dataset page: Individual Performance in Team-based Online Games
Extremist Propaganda Dataset
This dataset is associated with the paper “Contagion dynamics of extremist propaganda in social networks” (Information Sciences). PDFFull dataset page: Extremist Propaganda Dataset
Instagram Dataset
Source: Public media and user information from Instagram.com (through Instagram API). Crawling period: Jan 20 – Feb 17, 2014. Description: The media dataset contains records of the form: the anonymized media ID, the anonymized ID of the user who created the media, the timestamp of media creation, the set of tags assigned to the media, the number of likes and the number of comments it received. The anonymized user network contains asymmetric relations (A follows B); each edge is associated with #likes (by A to media created by B), #comments and the list of comments’ timestamps.Size
Media dataset: 1.7M media associated to 2K users, with 9M tags, 1200M likes, and 41M comments. User network: about 45K vertices and 678K edges.Request data
- Media dataset: about 51MB (200MB uncompressed).
- User network: about 21MB (7MB uncompressed).
Please cite:
Emilio Ferrara, Roberto Interdonato, Andrea Tagarelli. Online Popularity and Topical Interests through the lens of Instagram. In Proc. 25th ACM Conference on Hypertext and Social Media, September 1–4, 2014, Santiago, Chile.Full dataset page: Instagram Dataset