Next Article in Journal
Extremal Matching Energy of Random Polyomino Chains
Next Article in Special Issue
Energy and Entropy Measures of Fuzzy Relations for Data Analysis
Previous Article in Journal
Testing the Beta-Lognormal Model in Amazonian Rainfall Fields Using the Generalized Space q-Entropy
Previous Article in Special Issue
Capturing Causality for Fault Diagnosis Based on Multi-Valued Alarm Series Using Transfer Entropy
Article Menu
Issue 12 (December) cover image

Export Article

Open AccessArticle
Entropy 2017, 19(12), 686;

Do We Really Need to Catch Them All? A New User-Guided Social Media Crawling Method

Department of Computer Science and Engineering, Blekinge Institute of Technology, 371 79 Karlskrona, Sweden
Department of Computational Intelligence, Wrocław University of Science and Technology, 50-370 Wrocław, Poland
Author to whom correspondence should be addressed.
Received: 18 October 2017 / Revised: 28 November 2017 / Accepted: 11 December 2017 / Published: 13 December 2017
(This article belongs to the Special Issue Entropy and Complexity of Data)
PDF [774 KB, uploaded 13 December 2017]


[-15]With the growing use of popular social media services like Facebook and Twitter it is challenging to collect all content from the networks without access to the core infrastructure or paying for it. Thus, if all content cannot be collected one must consider which data are of most importance. In this work we present a novel User-guided Social Media Crawling method (USMC) that is able to collect data from social media, utilizing the wisdom of the crowd to decide the order in which user generated content should be collected to cover as many user interactions as possible. USMC is validated by crawling 160 public Facebook pages, containing content from 368 million users including 1.3 billion interactions, and it is compared with two other crawling methods. The results show that it is possible to cover approximately 75% of the interactions on a Facebook page by sampling just 20% of its posts, and at the same time reduce the crawling time by 53%. In addition, the social network constructed from the 20% sample contains more than 75% of the users and edges compared to the social network created from all posts, and it has similar degree distribution. View Full-Text
Keywords: social media; social networks; sampling; crawling; interestingness social media; social networks; sampling; crawling; interestingness

Figure 1

This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Supplementary material


Share & Cite This Article

MDPI and ACS Style

Erlandsson, F.; Bródka, P.; Boldt, M.; Johnson, H. Do We Really Need to Catch Them All? A New User-Guided Social Media Crawling Method. Entropy 2017, 19, 686.

Show more citation formats Show less citations formats

Note that from the first issue of 2016, MDPI journals use article numbers instead of page numbers. See further details here.

Related Articles

Article Metrics

Article Access Statistics



[Return to top]
Entropy EISSN 1099-4300 Published by MDPI AG, Basel, Switzerland RSS E-Mail Table of Contents Alert
Back to Top