Skip to Content
SymmetrySymmetry
  • Article
  • Open Access

19 November 2018

26 Pages

Online Social Networks (OSN) Evolution Model Based on Homophily and Preferential Attachment

and
School of Electronics and Information Engineering, Korea Aerospace University, Deogyang-gu, Goyang-si, Gyeonggi-do 412-791, Korea
*
Author to whom correspondence should be addressed.

Abstract

In this paper, we propose a new scale-free social networks (SNs) evolution model that is based on homophily combined with preferential attachments. Our model enables the SN researchers to generate SN synthetic data for the evaluation of multi-facet SN models that are dependent on users’ attributes and similarities. Homophily is one of the key factors for interactive relationship formation in SN. The synthetic graph generated by our model is scale-invariant and has symmetric relationships. The model is dynamic and sustainable to changes in input parameters, such as number of nodes and nodes’ attributes, by conserving its structural properties. Simulation and evaluation of models for large-scale SN applications need large datasets. One way to get SN data is to generate synthetic data by using SN evolution models. Various SN evolution models are proposed to approximate the real-life SN graphs in previous research. These models are based on SN structural properties such as preferential attachment. The data generated by these models is suitable to evaluate SN models that are structure dependent but not suitable to evaluate models which depend on the SN users’ attributes and similarities. In our proposed model, users’ attributes and similarities are utilized to synthesize SN graphs. We evaluated the resultant synthetic graph by analyzing its structural properties. In addition, we validated our model by comparing its measures with the publicly available real-life SN datasets and previous SN evolution models. Simulation results show our resultant graph to be a close representation of real-life SN graphs with users’ attributes.

1. Introduction

Online social networks (OSNs) have experienced dramatic growth, due to recent advancements in information and communication technologies. OSNs today have become the most important platforms for social interaction, communication, information processing, social influence and information diffusion, user/social opinion analysis, user behavior and personality investigation, business analytics and other commercial applications and services. Due to their wide range of applications, social network analysis has grabbed the attention of researchers, not only from the computer sciences, but also from social sciences, psychology, mathematics, computational linguistics, artificial intelligence and machine learning.
Understanding of social network (SN) structure, data distribution, and evolution helps in the development of SN models for different applications, such as recommendation systems [1], trust and reputation modeling [2], defense against Sybil attacks [3], community clustering [4,5], trust and interest based modeling [6], disaster detection [7] and other commercial applications. The literature suggests that SN structures along with user attributes, such as user profile information, provide a more in-depth understanding of SN and many useful patterns and trends, extracted from SN data, are used for applications in different domains. Several studies have been conducted to analyze SN. However, considering the wide range of SN applications, there is room for research to provide more comprehensive and promising models for improving the performance of SN applications and services.
The SN are usually composed of two entities, namely nodes (i.e., users or items) and edges, representing the relationship, between nodes. The relationship can be attributed in terms as one of the interaction types, such as friendship, co-authorship, SN activities between nodes and group membership. However, an extended definition of SN can be explored in the area of synthetic SN, i.e., SNs are sets of people or groups of people with common interests and interactions patterns, usually having similar attributes [8]. In SNs, users tend to connect with similar users based on spatial and demographic similarity or mutual interests despite influential users. In most SN evolution studies, the network topological structure is considered to be the basis of changes in the network. Topologically, due to continuous addition of new nodes and their preferential attachments with high degree nodes, the nodes connectivity follows the property of scale-free power law distribution [9,10].
The SN topology based evolution models produce synthetic networks that can structurally represent real-life SN and can be useful to evaluate topology based models and applications, such as modularity based community clustering [11], influence modeling [12], and information diffusion modeling [13]. However, these evaluation models are not suitable for applications where users’ attributes and the relationship strengths are critical, such as recommendation systems [14], trust models [15], and interest-based models [16]. Moreover, such models also have limitations when dealing with new users having low attachment degrees by assigning them low attachment probabilities. However, in real-life SN, new users are recommended to other users with more preference and many similar users are suggested to new users as well, and the probability of establishing relationships among new users is much higher than the existing ones.
Moreover, in SN analysis, access to real-life, complex SN data is critical to evaluating and validating SN models and applications. However, SNs provide APIs for data acquisition, and the process is generally time- and resource-consuming. Moreover, to ensure users’ privacy, social networks, such as Facebook, Instagram, and Twitter, provide a limited amount of information. To train and evaluate SN models, several datasets are available for different SN applications. However, these datasets are more application oriented and cannot be used to evaluate multi-facet models and frameworks.
In this paper, in order to deal with the limitations of topological evaluation models, we propose a novel evolution model based on the principal of homophily and preferential attachment by considering user attribute similarity as the basis for the connection formation in SN. Moreover, using the proposed model, personalized activities are generated, which, in combination with the relationship graph and profile information, results in a synthetic dataset. The proposed model is evaluated, by comparison with real SN datasets and models, to authenticate its feasibility as an approximation of real SN.
The main contributions of the paper can be summarized as:
  • We propose a novel SN evaluation model based on the principles of homophily combined with preferential attachments.
  • We generated a synthetic dataset with the proposed model that can be used as an approximation of data from real-life SN and can be used for the evaluation of the SN models for different application domains.
The rest of the paper is organized as follows. In Section 2, we provide a detailed review of the literature. Section 3 provides a detailed description of the proposed model. In Section 4, we provide detailed description of the conducted experiments, experimental results and detailed analysis of the results. Section 5 draws conclusions of the work.

3. OSN Evalution Model Based on Homophily and Preferential Attachment

In this section, we discuss the challenges associated with synthetic network generation, and also discuss how the proposed model can cope with such challenges.

3.1. Challenges in Synthetic Network Generation

In SN like synthetic network generation, many challenges must be considered while generating graphs and distributing attributes. These challenges are:
  • Attributes distribution: What is the distribution of attribute values? In SN, the node attributes can have high diversity in values. Therefore, it is necessary to determine what the percentage of different values for each target attribute is. These percentages can be obtained from real-life SN datasets, and SN statistics, e.g., the attribute gender has two possible values—male and female, and their percentage on Facebook is 47 % and 53 % , respectively [74].
  • Profile data distribution: What are the trends in the combination of user attributes to form different profiles? This is also an important factor that needs to be considered while generating synthetic SN data and graphs. In SNs, some node attributes can be used to predict the values of other attributes. These attributes are referred as inter-related attributes, e.g., if the age is in the range of 65 + , there is a high probability of having interest in news.
  • Communities structure: What is the community structure? There are many bases to form communities in SN. These bases range from structural parameters to profile similarity. Selecting the basis for community formation is application dependent, such as with information spreading, where the connectivity can be the basis, while, for recommendation systems, the interest similarity is a better choice.
  • Synthetic network topology: What is the topology? Many SN topologies are presented in the literature. SN topology can be obtained from real-life data sets. Previously, it was deduced from many studies that generally SNs have scale-free power-law degree distribution and have small world properties.
  • Activities distribution: What are the activities distribution? From studies in SNA, it has been observed that the social activities obey the power-law distribution. In [9], it was concluded that, in SN, we do not need to follow all users in a group, and, out of all, only a proportion of users generate about 80 % of activities. These trends in activities’ generation can be extracted from SN datasets, previous research, and surveys.
  • Correlation between attributes distribution: What is the correlation between these distributions? These distributions are associated with one another. These correlations can be deduced from SN datasets and surveys.

3.2. Preliminaries

The synthetic SN graph is represented as G ( V , E ) , where V is set of nodes, and E is the set of edges between the nodes. In this work, we have represented SN node attributes into three categories, i.e., spatial information, represented by S p , demographic information, represented by D m , and interest information, represented by I n . S p is a set of location information of each user in the network, where, for the ith user, the location is represented by S p i , and each S p i contains two values, i.e., latitude ( ϕ ) and longitude ( φ ) values. D m , represents a set of demographic information for SN users, with its elements for each ith user as D m i , the D m i contains the age, gender, occupation, religion, language, political view, and initial interests’ information of each user i, which they enter while joining SN. The set of items interests is represented by I n , where I n i represents the interest information of every ith user in the network. The spatial, demographic and initial interests information combines to form user profile ( P r ), and, for every ith node, the profile is represented by P r i , where i ranges from [ 1 , N ] , and N is the number of users in the network. As this is a dynamic/evolving network, the number of nodes can increase by the addition of new nodes over time; therefore, the value of N is incremented by the number of nodes newly added to the network. With the addition of new nodes, the demographic, spatial and initial interest information of new nodes are populated to their corresponding sets. For items’ interests, we used Filmtrust, a publicly available dataset, which contains user-item ratings.
The set S I contains the social interactions generated by users on SN. These social interactions are posts, likes, comments, shares and other information, generated by users on SN. The social interactions are further divided into two categories: i.e., social actions and social responses. The social actions, ( S I i j ) , are the actions performed by a user i to user/items j and the social responses ( S I j i ) are the responses that a user i receives from user j. The social interactions may be in the form of likes, comments, tags, posts, shares, re-shares, tweets, re-tweets, followings, followers, tweets, and re-tweets. The properties, like allowable contents, lengths, and formats, of these interactions may vary due to the business model of different SN. Some interactions may be present in one media but absent in other, depending on the SN applications domain. Therefore, we categorized these interactions into the aforementioned two broad categories. There exist other social media content that the SN users generate, but here we have only considered the social interactions. We have combined users into different groups, and users with similar properties are assigned to the same group, known as communities, based on homophily and demoted by C k ∈ C , where C, is the set of all communities and C k , is the kth community in the network. These notations are listed with brief description in Table 1.
Table 1. List of Notations and Description.

3.3. Proposed SN Evaluation Model

Figure 1 provides a block diagram of the proposed model. The proposed model is composed of seven phases, namely (i) node generation, (ii) data generation, (iii) data combination or profile generation, (iv) data population, (v) similarity based clustering, (vi) network graph generation, and (vii) social activities generation.
Figure 1. Proposed social network operational/dynamic model based on homophily, preferential attachment and personalized activities generation.
As a first step, we generated N number of nodes in nodes generation phase of Figure 1. Example of labelled generated nodes are shown in Figure 2. The node generation process is followed by the data generation phase. The data generation phase is application dependant, and represent SN nodes’ demographic, spatial and interest attributes.
Figure 2. Example of Social Networks nodes generated.
In order to provide an approximation of real-life SN, for the selection of our model’s attributes, we rely on the statistics from public sources, such as US government census data [72], and publicly available SN statistics and insights, such as Facebook [74,75]. Table 2 shows the example attributes and their proportion. There exists some attribute with inter-related values. The user’s initial interests’ information in SN can be associated with users’ demography like age, gender, profession, and location information. In addition, interests in SN can be inferred based on user information, such as social interactions, and user’s neighbors’ interests. In SN user’s interests can be predicted from other users, with higher demographic similarity [76]. Similarly, in [77], the effect of various SN user’s information on interest similarity is investigated by video based user profiling. We need to generate lookup tables to assign such attributes. For example, nodes in age group 18–25, it is more likely to have profession = student and marital status = single, and interests = Sports Teams. Another option, to get these proportions, is to use publicly available SN datasets from sources like kaggle, SNAP, UCInet, dataworld, etc. We generated random geo-locations in the North-America region as shown in Figure 3.
Table 2. Example attributes, attribute-values, their spatial, demographic and interest proportions.
Figure 3. Random geo-location generation in the North-American region; each point on the map consists of latitude and longitude value.
After the nodes and attributes generation, we combined the attributes according to the trends and patterns extracted from real-life datasets and studies to generate profiles. The SNs are global platforms and can have different combinations of these attributes and form a diverse range of profiles. In this work, as an example, we have considered a limited number of nodes, attributes and attribute values. These parameters can be extended based on the application area where this model is used. This model is a generic one and can be adopted according to the application. In Figure 4, the combination of social attributes referred to as users’ profile are assigned to each generated node. In the Figure 4 example, values are assigned to each attribute of the users’ profiles, e.g., three interest, age range, gender and location. The locations in this example are labelled as L1, L2, L3, L4 and L5. These locations contain latitude and longitude values of the location where the user belongs. In profile generation, we also consider the inter-dependency of attributes. For instance, for users with age in the range of 18–25, it is more likely for them to be single, and, if the gender is male, it is more probable to have interest in sports teams. On the other hand, the probability for having interests in brands is higher for females. For the aged user, it is more likely to have an interest in news.
Figure 4. Profiles’ generation and assignment; the profiles are generated by combining the SN users’ attributes.
After the SN nodes generation and profile information population, we calculated similarities between nodes to generate a similarity matrix. The similarity matrix depicts how much a node is similar to all other nodes. The similarity is calculated by using Equation (1) [14]:
U S i m ( u , v ) = α . S p S i m ( u , v ) + β . D m S i m ( u , v ) + γ . I n S i m ( u , v ) + w . O p S i m ( u , v ) ,
where U S i m ( u , v ) is the normalized weighted cumulative similarity between users u and v. This is weighted sum of spatial similarity, demographic similarity, and interest similarities. The α , β , γ , and w are the weights assigned to each similarity based on their importance in SN relation formation.
For computing spatial similarity ( S p S i m ( u , v ) ) between users u, and v, we used cosine similarity [78]. We normalized the spatial similarity by the geo-distance between users, as shown in Equation (2):
S p S i m ( u , v ) = 1 1 + D i s t ( u , v ) 1000 S p u . S p v ∥ S p u ∥ . ∥ S p v ∥ ,
where D i s t ( u , v ) is the geo-distance between the users u and v. Equation (3) finds the geo-distance by using Harvesine formula [79,80]:
a = sin 2 Δ ϕ 2 + cos ( ϕ S p u ) . cos ( ϕ S p v ) . sin 2 Δ φ 2 ; w h e r e   Δ ϕ = ∣ ϕ S p u − ϕ S p v ∣ , Δ φ = φ S p u − φ S p v c = 2 arctan a 1 − a D i s t ( u , v ) = R × c .
In Equation (2), the S p u and S p v are the GPS coordinates of the users u and v and consists of pair of latitude ( ϕ ) and longitude ( φ ) value. ϕ S p u and ϕ S p v are the latitudes of the users u and v and φ S p u and φ S p v are the longitudes of users u and v, respectively.
The demographic similarity ( D m S i m ( u , v ) ) between users u and v is calculated by using cosine similarity [78], as shown in Equation (4):
D m S i m ( u , v ) = ∑ i = 1 D D m i u . D m i v ∑ i = 1 D D m i u 2 . ∑ i = 1 D D m i v 2 ,
where D m i u and D m i v are the ith demographic attribute of users u and v and i ranges from 1 to D, i.e., total number of demographic attributes in D m .
The users’ interest similarity are the users’ interests in different categories such as some users like sports, some users’ have interests in books and movies, some may like to follow celebrities, and some may have interest in politics. Equation (5) computes the interests similarity ( I n S i m ( u , v ) ) between users u and v by using a Jaccard similarity formula [81]. In Equation (5), the interest similarity of two users is computed by dividing the number of common interests with the total number of interests that both users have. When the number of users’ common interests increases, the interest similarity is increased and vice versa. The users’ interest information is binary, e.g., a user is interested in sports or not; therefore, we used Jaccard similarity to measure the interest similarity of users:
I n S i m ( u , v ) = [ I n u ⋂ I n v ] [ I n u ⋃ I n v ] ,
where I n u and I n v are the interest information of users u and v.
The user-items opinion similarity is computed from the Filmtrust dataset by using Equation (6). In SN, the users adopt items based on the rating and opinions of other users with similar interests, and such users are more prone to forming a connection. The user-item opinion similarity is computed from ratings which the SN users give to different items such as movies, food and other products:
O p S i m ( u , v ) = 1 − ∑ i ∈ ( I S e t u ⋂ I S e t v ) . ∣ R u , i − R v , i ∣ R m a x ∣ I S e t u ⋂ I S e t v ∣ ,
where I S e t u and I S e t v are the sets of items rated by users u and v, respectively. R u , i and R v , i are the ratings given by users u and v to item i and R m a x is the maximum rating.
After computation of similarity matrix, K seed nodes are selected, based on the similarity measures, to generate K clusters. The initial nodes are selected in such a way that the seed nodes selected must be similar to more nodes in the network. In addition, all the seed nodes must be dissimilar from each other. An example of seed nodes is shown in Figure 5. The nodes are then clustered based on the similarity measures in the previous phase. Figure 6 shows nodes assignment to their respective kth clusters C k . Each cluster contains a set of similar users. The initial seed nodes allow the nodes to lie in their respective communities based on their similarities. The initial seed nodes and initial links generation are important to initialize the network based on similarity, rather than random initialization. When the number of seeds increased, the network convergence time is slightly decreased.
Figure 5. Seed nodes selection for clustering. The selected seeds are similar to most of the generated nodes and each seed node attributes are different from other selected seed nodes.
Figure 6. Similarity based clustering. Clusters are represented by different colors.
To initialize the graph, links are established between the seed nodes and other member nodes of their corresponding clusters. After the network initial nodes, more links are generated between the nodes other than the seed nodes based on their similarity and preferential attachment. The similarity is the primary parameter for the connection establishment with the degree of the nodes as the secondary parameters, i.e., nodes with higher similarity and degree are more likely to form connection. Figure 7 shows an example of initial links establishment and Figure 8, shows an example of links other than within the community. We limit the number of links in the synthetic graph based on the minimum degree that each node must achieve and average degree of the network, and each node is assigned a random number of links according to the aforementioned criteria, shown in Equation (7). In simulations, the minimum degree and average degree, whichever occurs first, was set as the model convergence criteria. The simulation results are shown in Table 4 in Section 4:
P a t t a c h = w s . U S i m + w d e g . P d e g ( j ) .
Figure 7. Initial links generation within the clusters. Nodes in each cluster are connected with their respective cluster centers.
Figure 8. More links generation, outside and within communities, based on similarity and preferential attachments.
In Equation (7), the P a t t a c h is the probability of a node to attach to other nodes, U S i m is the user similarity, computed by using Equation (1), and the P d e g ( j ) is the degree based attachment probability. The w s and w d e g are the weights of the similarity and degree-based attachments, respectively. The computational formula for degree based attachment [10], is shown in Equation (8):
P d e g ( j ) = d e g ( j ) + 1 ∑ k = 1 N d e g ( k ) ,
where d e g ( j ) is the degree of the target node and the d e g ( k ) is the degree of all other nodes, where k ranges from 1 to N, and N is the total number of nodes in the network.
We considered the average degree to be 100 for a network of 1000 nodes with minimum degree set to 10. The number of links depends on these initial conditions. When we increase the minimum degree, each node must achieve the number of links in the network will increase, and vice versa. Similar is the case of the average degree of the network. We generated random connectivity degree d e g ( j ) for each node and selected the top d e g ( j ) most similar nodes for each jth node. The proposed model is dynamic, and, in our implementation, 10 new nodes are added to the network at each iteration, where each node is connected to the most similar node in their respective community, as their initial connection.
The connectivity of SN users is a projection of profile similarity of the users with other existing users and the newly added users, and the human characteristic, in order to manage/keep limited friends at a time.
After the network generation, we generated user activities. According to [9], user activities also obey the power law distribution, i.e., there are only a small proportion of users who produce a significant proportion of activities over SN. In [10], it was observed that almost 80% of user activities, within an SN group, are produced by 20% of the users in the group. We produced different types of user activities/interactions with a proportion from [9,10,82].
The time complexity of the proposed systems depends on the number of nodes N, number of initial seeds K and number of users’ attributes. The numerical results are shown in Section 4, Table 5.

4. Evaluation and Simulation Results

In this section, we evaluated our model by generating a synthetic SN graph, based on homophily and preferential attachment. From simulations, we observed that the synthetic SN graph obey the SN principles of “Birds of feather flocks together”, and the “Rich get richer”. We validated our results by comparing the closeness of the structural and similarity properties of our synthetic network with that of real-life SN, obtained in different studies. We found that the properties, i.e., degree distribution, clustering coefficients, modularity, spatial similarity, demographic similarity, interest similarity, network diameter, shortest path, of our synthetic network are similar to that of the real-life SN.
We have generated a test network of 1000 nodes with a random distribution of SN nodal attributes as discussed in Section 3. For each node we have generated, random latitudes and longitudes in the US region, demographic attributes, and interest information. The set of demographic attributes considered in the simulation of our model consists of age, gender, religion, language, marital status, profession political orientation and a set of initial interests. These attributes are generally specified by the users at the time of joining an SN. The distribution of demographic information is shown in Table 2. We generated a network of 46,780 edges by connecting the nodes based on similarity and preferential attachment. For graph representation and statistical analysis, we exported our synthetic network to Gephi. Figure 9 shows modularity based clustered synthetic SN graph, generated in Gephi.
Figure 9. Synthetic network graph.
Previously, it has been observed that real-life SN are scale-free and they obey the degree distribution obey power law distribution. From the Figure 10 and Figure 11, it can be observed that the degree distribution of our synthetic network graph follows the power law distribution, which is similar to that of the real-life SN. As most of the synthetic graph nodes are low degree nodes, the average degree of our undirected synthetic graph is 94.56, and this value corresponds to 50% of the CDF, as can be seen in Figure 10. This phenomenon is due to the people psychology to make friendship mostly with similar people only and there is a limited number of friends they make. Such pattern can be observed both in online and offline friendships, as discussed in Section 3. Figure 12 shows the distribution and relationship between age and connectivity. It can be seen in Figure 12 that most of the friends for a user belong to the same age group, i.e., young age users are more likely to connect with other young age users and the elder users are more probable to connect with other old age users. This age-based connectivity is due to their common interests. In [51], a study on the Facebook social graph, it was concluded that users on SN make friends with other users of same age group, but this pattern is more prominent in young individuals, and the age range for neighbors of young users is smaller compared to that of the older ones. We observed similar properties in our synthetic social graph. The connectivity of users with other age group users in the synthetic SN graph is due to attribute similarity other than age i.e., gender, religion, language, profession and initial interests.
Figure 10. Degree distribution; CDF of nodes degree distribution in the synthetic graph.
Figure 11. Degree distribution; nodes degree distribution representation in log–log graphs. The synthetic network obeys the power-law degree distribution.
Figure 12. Age and relationships distribution; similar age users are more likely to connect with each other.
In [51], it is concluded that, in SNs, users establish links with other users that belong to the same locality. Figure 13 shows the geodesic distance between connected users. By comparing the geodesic distances between connected nodes in our synthetic SN graph, it is found to have a similar distribution. In Figure 13, it can be observed that most of the users are connected with other users, which belongs to the nearby geo-location. In [51], it was observed that, on Facebook, a user is more likely to establish a connection with a user from the same country and, from simulation, it was concluded that about 84% of the total edges are within the country. Therefore, the SN users can be divided into clusters of locations for analysis.
Figure 13. Location distribution and relationships; users with less geo-distance are more probable to connect.
Figure 14 shows the relationship of demographic similarity with the links’ establishment between nodes in our synthetic SN. It can be observed that, like real-life SNs, a high proportion of links are established between nodes that are demographically more similar. The links’ establishment in our model depends on multiple attribute and demographic similarity is one of the contributing factors and is not totally dependent on demographic similarity. In addition, the demographic attributes were distributed randomly, and it is more likely for most of the nodes to have a different combination of demographic attributes and only a small fraction of users’ have demographic similarity in the range 0.8–1. Figure 15 shows the interest similarity of connected users in our synthetic SN graph. It is observed that the connected users have high interest similarity. This property, also referred to as “birds of a feather flock together”, is observed in real-life SNs, i.e., Facebook, Twitter, and Instagram. This pattern is the base for many marketing and business models on SN. In the real-life SNs, new friends are recommended to the users based on their common or shared interests. However, the fraction of links for very small interests similarity are comparable because there are some items that both connecting users’ have not rated/experienced. In addition, the connection formation in our model depends on multiple attributes and, as other attributes, the opinion similarity also contributes to the cumulative similarity of all attributes, and there exist users’ having very high demographics, and spatial similarity but low opinion similarity. Such users are more likely to connect due to their demographic and spatial similarities, and vice versa.
Figure 14. Demographic similarity and relationship distribution; users with high demographic similarity are more probable to connect with each other.
Figure 15. Opinion similarity and relationships distribution; users with the same opinion have a high tendency to connect.
Figure 16 shows a CDF of clustering coefficients against a fraction of users, for our graph in comparison with the BA model. From Figure 16, we can see that the clustering coefficient is improved by our model because, in our model, the connections are established between the most similar users that are more likely to lie in the same community. From Figure 16, we can observe that most of the users in our synthetic graph have relatively high clustering coefficients and the range of clustering coefficient for more than 90 % users is from [ 0.4 , 0.7 ] .
Figure 16. Clustering coefficient distribution; comparison of clustering co-efficient of the proposed model with the BA-model (preferential attachment based approach).
In Figure 17, the distribution of social interaction is shown. It is observed that, like real-life SN, the interactions generated by our model obeys the power-law distribution, i.e., and only a small proportion of users are highly active and produce most of the SN interactions. Such users are the most influential users in SN. The distribution of our synthetic social interactions was close to the real-life SN interactions, as observed by [10].
Figure 17. Social interactions distribution; most of the interactions are generated by only a small fraction of users. Such users are referred to as the most active users.
In Table 3, the synthetic SN is compared with the real-life SN datasets to validate their structural closeness based on standard graph measures, i.e., average degree, average path length, average clustering coefficient, modularity, graph density, and graph diameter. We used SN datasets publicly available on Stanford Large Network Dataset Collection [83] and Arizona State University Social Computing Data Repository [84]. The datasets we used are Livejournal [85], Facebook [86], Twitter [84], Friendster [85], Amazon [85], Douban [84], Digg [84], Karate Club [87], Les Mesrible [88], Netsciences [11], Enron email [89,90], CollegeMsg [91], contact network [92], and random graph generated in Gephi. These datasets are accessible and widely used. From Table 3, we can see that, based on the comparison measures, the performance of our model is comparable to that of the real-life SN datasets.
Table 3. Comparison of proposed social network graph with real-life social networks data, and existing SN evolution models based on various metrics and attributes.
As discussed in Section 3, the convergence of our model depends on the minimum degree that each node needs to achieve and the network average degree. Table 4 shows the simulation results of our model with different initial conditions. From Table 4, it can be observed that the number of links in the network is increased as the value of minimum degree and average degree increased. Our model converges when one of the two initial limiting conditions is achieved. For the simulation with no limit on initial conditions, a time-out was set to converge the model.
Table 4. Number of links’ dependency on initial conditions, i.e., minimum node degree and network average degree for network of 1000 nodes.
As discussed in Section 3, the time complexity of our proposed model depends on the number of nodes N, number of similarity parameters, and slightly on the number of initial seeds K. Table 5 shows results of simulation time for our model. We have simulated our model in two types of settings, i.e., pre-computed similarity matrix and with similarities computation. We have found from simulation that the similarity matrix computation is very time-consuming and it hugely increases the time complexity of our model. The model simulation time is also raised as the number of nodes is increased. We have simulated our model for thousands of nodes. The simulation was done in MATLAB 2018 (The MathWorks, Inc., Natick, MA, USA). The system we used for simulation was LG (LG Electronics Nanjing Displays Co. Ltd., Nanjing, China), Core i5, 2.3 GHz processor, 8 GHz memory, and Windows 10 Pro 64-bit (Microsoft, Redmond, WA, USA). For space complexity, on the described system setting, when we increase the number of nodes to more than 20,000, a memory error is shown by MATLAB and also the model needs several days to simulate for 20,000 nodes; therefore, we have shown the results of our model for up to 10,000 nodes. Therefore, for current settings and settings, the memory issue can occur.
Table 5. Time complexity of the proposed model with different numbers of nodes, and number of seeds.
For implementation, we have generated random geo-locations in the North-American region, and user interests’ information using MATLAB built-in functions and commands. We have generated seven demographic attributes with different percentages and distribution, as discussed in Section 3 and shown in Table 2, by using MATLAB. The main program consists of nodes profile generation, similarities computation, similarity based clustering and synthetic graph generation.
As a result of the simulations, a synthetic dataset is produced that can be used by researchers to evaluate their social network models. Our model is scalable and general, and researchers can generate N number of nodes with different attributes depending on their target applications.

5. Conclusions

In this paper, we proposed an SN evolution model based on the property of homophily and preferential attachments. The proposed model is simulated, to generate a synthetic SN graph, evaluated, based on SN structural and similarity properties, and validated by using many graph measures and being compared with real-life SN datasets.
From simulations, it is observed that the network graph, synthesized by our model, bears the real-life OSN properties. Like real-life SN, our synthetic graph has a scale-free nature and obeys the power-law degree distribution. The synthetic SN graph comprises the small-world effect, also referred to as six degrees of separation. The average distance between the nodes was found to be 2.49 and the graph diameter was 6, which is comparable to that of the real-life SN. The resultant graph was found to be fully connected, having short average path length, comparable clustering coefficient, average degree, modularity, and graph diameter. The networks with these properties are categorized to be small world networks [38].
We also observed the relationship between SN structure and user attributes. These attributes are demographic attributes, geodesic distance, and user interest. In real-life SNs, like Facebook, it is observed that the users are more likely to establish a connection with other users of similar demographic attributes and belong to the same locality [51]. From simulation, similar patterns have been observed in our synthetic graphs.
Moreover, as a result of this work, a generic SN synthetic dataset is obtained, which contains SN nodal attributes, such as demographic attributes, spatial information and user interests. The resultant SN graph obeys the SN structural properties, which exhibits its closeness to real-life SN graphs.
The proposed model is scalable, generic, and is based on the phenomenon and psychology behind the real-life SN links formation. This model can be used, by the researcher in the area of SNA, for SN synthetic datasets to evaluate their models and applications in many SN domains. The researchers, while using this model for graph generation, can set attributes, and attribute values distribution and weights according to their target applications.

Author Contributions

The authors contributed equally to this work overall.

Funding

This research was supported by Basic Science Research Programs through the National Research Foundation of Korea (NRF) funded by the Ministry of Education (NRF-2017R1A2B1010817).

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Rose, D.E.; Bornstein, J.J.; Tiene, K.; Ponceleón, D.B. System for Ranking the Relevance of Information Objects Accessed by Computer Users. U.S. Patent Application 10/388,362, 26 October 2010. [Google Scholar]
  2. Mui, L. Computational Models of Trust and Reputation: Agents, Evolutionary Games, and Social Networks. Ph.D. Thesis, Massachusetts Institute of Technology, Cambridge, MA, USA, 2002. [Google Scholar]
  3. Yu, H.; Kaminsky, M.; Gibbons, P.B.; Flaxman, A. Sybilguard: Defending against sybil attacks via social networks. ACM SIGCOMM Comput. Commun. Rev. 2006, 36, 267–278. [Google Scholar] [CrossRef] [Scilit]
  4. Girvan, M.; Newman, M.E. Community structure in social and biological networks. Proc. Natl. Acad. Sci. USA 2002, 99, 7821–7826. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Lancichinetti, A.; Fortunato, S. Community detection algorithms: A comparative analysis. Phys. Rev. E 2009, 80, 056117. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Ullah, F.; Lee, S. Community clustering based on trust modeling weighted by user interests in online social networks. Chaos Solitons Fractals 2017, 103, 194–204. [Google Scholar] [CrossRef] [Scilit]
  7. Ahmad, K.; Pogorelov, K.; Riegler, M.; Conci, N.; Halvorsen, P. Social media and satellites. Multimed. Tools Appl. 2018, 1–39. [Google Scholar] [CrossRef] [Scilit]
  8. Radicchi, F.; Castellano, C.; Cecconi, F.; Loreto, V.; Parisi, D. Defining and identifying communities in networks. Proc. Natl. Acad. Sci. USA 2004, 101, 2658–2663. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Wilson, C.; Sala, A.; Puttaswamy, K.P.; Zhao, B.Y. Beyond social graphs: User interactions in online social networks and their implications. ACM Trans. Web (TWEB) 2012, 6, 17. [Google Scholar] [CrossRef] [Scilit]
  10. Durr, M.; Protschky, V.; Linnhoff-Popien, C. Modeling social network interaction graphs. In Proceedings of the 2012 International Conference on Advances in Social Networks Analysis and Mining (ASONAM 2012), Istanbul, Turkey, 26–29 August 2012; IEEE Computer Society: Washington, DC, USA, 2012; pp. 660–667. [Google Scholar]
  11. Newman, M.E. Modularity and community structure in networks. Proc. Natl. Acad. Sci. USA 2006, 103, 8577–8582. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Kempe, D.; Kleinberg, J.; Tardos, É. Maximizing the spread of influence through a social network. In Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Washington, DC, USA, 24–27 August 2006; ACM: New York, NY, USA, 2003; pp. 137–146. [Google Scholar]
  13. Guille, A.; Hacid, H.; Favre, C.; Zighed, D.A. Information diffusion in online social networks: A survey. ACM Sigmod Rec. 2013, 42, 17–28. [Google Scholar] [CrossRef] [Scilit]
  14. Ullah, F.; Lee, S. Social content recommendation based on spatial-temporal aware diffusion modeling in social networks. Symmetry 2016, 8, 89. [Google Scholar] [CrossRef] [Scilit]
  15. Al Qundus, J.; Paschke, A. Investigating the Effect of Attributes on User Trust in Social Media. In Proceedings of the International Conference on Database and Expert Systems Applications, Regensburg, Germany, 3–6 September 2018; Springer: Berlin, Germany, 2018; pp. 278–288. [Google Scholar]
  16. Bai, Y.; Deng, G.; Zhang, L.; Wang, Y. A Measuring Method for User Similarity based on Interest Topic. Int. J. Perform. Eng. 2018, 14, 691–698. [Google Scholar] [CrossRef] [Scilit]
  17. Sampson, S. A Novitiate in a Period of Change: An Experimental and Case Study of Social Relationships. Ph.D. Thesis, Cornell University, Ithaca, NY, USA, 1968. [Google Scholar]
  18. Burt, M.R. Cultural myths and supports for rape. J. Personal. Soc. Psychol. 1980, 38, 217. [Google Scholar] [CrossRef]
  19. Johnson, J.C. Social networks and innovation adoption: A look at Burt’s use of structural equivalence. Soc. Netw. 1986, 8, 343–364. [Google Scholar] [CrossRef] [Scilit]
  20. Johnsen, E.C. Structure and process: Agreement models for friendship formation. Soc. Netw. 1986, 8, 257–306. [Google Scholar] [CrossRef] [Scilit]
  21. Kumar, R.; Novak, J.; Raghavan, P.; Tomkins, A. Structure and evolution of blogspace. Commun. ACM 2004, 47, 35–39. [Google Scholar] [CrossRef] [Scilit]
  22. Liben-Nowell, D.; Novak, J.; Kumar, R.; Raghavan, P.; Tomkins, A. Geographic routing in social networks. Proc. Natl. Acad. Sci. USA 2005, 102, 11623–11628. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Dodds, P.S.; Muhamad, R.; Watts, D.J. An experimental study of search in global social networks. Science 2003, 301, 827–829. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Adamic, L.; Adar, E. How to search a social network. Soc. Netw. 2005, 27, 187–203. [Google Scholar] [CrossRef] [Scilit]
  25. Wasserman, S.; Faust, K. Social Network Analysis: Methods and Applications; Cambridge University Press: Cambridge, UK, 1994; Volume 8. [Google Scholar]
  26. Strogatz, S.H. Exploring complex networks. Nature 2001, 410, 268–276. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Albert, R.; Barabási, A.L. Statistical mechanics of complex networks. Rev. Mod. Phys. 2002, 74, 47. [Google Scholar] [CrossRef] [Scilit]
  28. Newman, M.E. The structure and function of complex networks. SIAM Rev. 2003, 45, 167–256. [Google Scholar] [CrossRef] [Scilit]
  29. Dorogovtsev, S.N.; Mendes, J.F. Evolution of networks. Adv. Phys. 2002, 51, 1079–1187. [Google Scholar] [CrossRef] [Scilit]
  30. Dorogovtsev, S.N.; Mendes, J.F. Evolution of Networks: From Biological Nets to the Internet and WWW; OUP Oxford: Oxford, UK, 2013. [Google Scholar]
  31. Kleinberg, J. Complex networks and decentralized search algorithms. In Proceedings of the International Congress of Mathematicians (ICM), Madrid, Spain, 22–30 August 2006; Volume 3, pp. 1019–1044. [Google Scholar]
  32. Kumar, R.; Novak, J.; Raghavan, P.; Tomkins, A. On the bursty evolution of blogspace. World Wide Web 2005, 8, 159–178. [Google Scholar] [CrossRef] [Scilit]
  33. Fetterly, D.; Manasse, M.; Najork, M.; Wiener, J.L. A large-scale study of the evolution of Web pages. Softw. Pract. Exp. 2004, 34, 213–237. [Google Scholar] [CrossRef] [Scilit]
  34. Ntoulas, A.; Cho, J.; Olston, C. What’s new on the web? The evolution of the web from a search engine perspective. In Proceedings of the 13th International Conference on World Wide Web, New York, NY, USA, 17–20 May 2004; ACM: New York, NY, USA, 2004; pp. 1–12. [Google Scholar]
  35. Newman, M.E.; Watts, D.J.; Strogatz, S.H. Random graph models of social networks. Proc. Natl. Acad. Sci. USA 2002, 99, 2566–2572. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Erdds, P.; Rényi, A. On random graphs I. Publ. Math. Debr. 1959, 6, 290–297. [Google Scholar]
  37. Bollobás, B.; Fulton, W.; Katok, A.; Kirwan, F.; Sarnak, P. Cambridge Studies in Advanced Mathematics; Random Graphs; Cambridge University Press: New York, NY, USA, 2001; Volume 73. [Google Scholar]
  38. Watts, D.J.; Strogatz, S.H. Collective dynamics of ‘small-world’networks. Nature 1998, 393, 440–442. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Barabási, A.L.; Albert, R. Emergence of scaling in random networks. Science 1999, 286, 509–512. [Google Scholar] [PubMed]
  40. Barabási, A.L. Scale-free networks: A decade and beyond. Science 2009, 325, 412–413. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Chakrabarti, D.; Zhan, Y.; Faloutsos, C. R-MAT: A recursive model for graph mining. In Proceedings of the 2004 SIAM International Conference on Data Mining, Lake Buena Vista, FL, USA, 22–24 April 2004; pp. 442–446. [Google Scholar]
  42. Lancichinetti, A.; Fortunato, S.; Radicchi, F. Benchmark graphs for testing community detection algorithms. Phys. Rev. E 2008, 78, 046110. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Kleinberg, J.M. Authoritative sources in a hyperlinked environment. J. ACM 1999, 46, 604–632. [Google Scholar] [CrossRef] [Scilit]
  44. Chung, F.; Lu, L. Connected components in random graphs with given expected degree sequences. Ann. Comb. 2002, 6, 125–145. [Google Scholar] [CrossRef] [Scilit]
  45. Chung, F.; Lu, L. The average distance in a random graph with given expected degrees. Internet Math. 2004, 1, 91–113. [Google Scholar] [CrossRef] [Scilit]
  46. Waxman, B.M. Routing of multipoint connections. IEEE J. Sel. Areas Commun. 1988, 6, 1617–1622. [Google Scholar] [CrossRef] [Scilit]
  47. Redner, S. How popular is your paper? An empirical study of the citation distribution. Eur. Phys. J. B-Condens. Matter Complex Syst. 1998, 4, 131–134. [Google Scholar] [CrossRef] [Scilit]
  48. Menczer, F. Evolution of document networks. Proc. Natl. Acad. Sci. USA 2004, 101, 5261–5265. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  49. McPherson, M.; Smith-Lovin, L.; Cook, J.M. Birds of a feather: Homophily in social networks. Annu. Rev. Sociol. 2001, 27, 415–444. [Google Scholar] [CrossRef] [Scilit]
  50. Şimşek, Ö.; Jensen, D. Navigating networks by using homophily and degree. Proc. Natl. Acad. Sci. USA 2008. [Google Scholar] [CrossRef] [Scilit]
  51. Ugander, J.; Karrer, B.; Backstrom, L.; Marlow, C. The anatomy of the facebook social graph. arXiv, 2011; arXiv:1111.4503.
  52. Papadopoulos, F.; Kitsak, M.; Serrano, M.Á.; Boguná, M.; Krioukov, D. Popularity versus similarity in growing networks. Nature 2012, 489, 537–540. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Huber, G.A.; Malhotra, N. Political homophily in social relationships: Evidence from online dating behavior. J. Politics 2017, 79, 269–283. [Google Scholar] [CrossRef] [Scilit]
  54. Neyer, F.J.; Lang, F.R. Blood is thicker than water: Kinship orientation across adulthood. J. Personal. Soc. Psychol. 2003, 84, 310. [Google Scholar] [CrossRef]
  55. Doherty, N.A.; Feeney, J.A. The composition of attachment networks throughout the adult years. Pers. Relationsh. 2004, 11, 469–488. [Google Scholar] [CrossRef] [Scilit]
  56. Gerich, J.; Lehner, R. Collection of ego-centered network data with computer-assisted interviews. Methodology 2006, 2, 7–15. [Google Scholar] [CrossRef] [Scilit]
  57. Van Tilburg, T. Losing and gaining in old age: Changes in personal network size and social support in a four-year longitudinal study. J. Gerontol. Ser. B Psychol. Sci. Soc. Sci. 1998, 53, S313–S323. [Google Scholar] [CrossRef] [Scilit]
  58. Said, A.; De Luca, E.W.; Albayrak, S. How social relationships affect user similarities. In Proceedings of the 2010 International Conference on Intelligent User Interfaces Workshop on Social Recommender Systems, Hong Kong, China, 7–10 February 2010. [Google Scholar]
  59. Hanani, U.; Shapira, B.; Shoval, P. Information filtering: Overview of issues, research and systems. User Model. User-Adapt. Interact. 2001, 11, 203–259. [Google Scholar] [CrossRef] [Scilit]
  60. Bobadilla, J.; Ortega, F.; Hernando, A.; Gutiérrez, A. Recommender systems survey. Knowl.-Based Syst. 2013, 46, 109–132. [Google Scholar] [CrossRef] [Scilit]
  61. Xie, J.; Li, X. Make best use of social networks via more valuable friend recommendations. In Proceedings of the 2012 2nd International Conference on Consumer Electronics, Communications and Networks (CECNet), Yichang, China, 21–23 April 2012; IEEE: Piscataway, NJ, USA, 2012; pp. 1112–1115. [Google Scholar]
  62. Chin, A.; Xu, B.; Wang, H. Who should I add as a friend? A study of friend recommendations using proximity and homophily. In Proceedings of the 4th International Workshop on Modeling Social Media, Prague, Czech Republic, 23 September 2013; ACM: New York, NY, USA, 2013; p. 7. [Google Scholar]
  63. Pouli, V.; Kafetzoglou, S.; Tsiropoulou, E.E.; Dimitriou, A.; Papavassiliou, S. Personalized multimedia content retrieval through relevance feedback techniques for enhanced user experience. In Proceedings of the 2015 13th International Conference on Telecommunications (ConTEL), Graz, Austria, 13–15 July 2015; IEEE: Piscataway, NJ, USA, 2015; pp. 1–8. [Google Scholar]
  64. Stai, E.; Kafetzoglou, S.; Tsiropoulou, E.E.; Papavassiliou, S. A holistic approach for personalization, relevance feedback & recommendation in enriched multimedia content. Multimed. Tools Appl. 2018, 77, 283–326. [Google Scholar]
  65. Gauch, S.; Speretta, M.; Chandramouli, A.; Micarelli, A. User profiles for personalized information access. In The Adaptive Web; Springer: Berlin, Germany, 2007; pp. 54–89. [Google Scholar]
  66. Ajrouch, K.J.; Blandon, A.Y.; Antonucci, T.C. Social networks among men and women: The effects of age and socioeconomic status. J. Gerontol. Ser. B Psychol. Sci. Soc. Sci. 2005, 60, S311–S317. [Google Scholar] [CrossRef] [Scilit]
  67. Quinn, D.; Chen, L.; Mulvenna, M. Does age make a difference in the behaviour of online social network users? In Proceedings of the Internet of Things (iThings/CPSCom), 2011 International Conference on and 4th International Conference on Cyber, Physical and Social Computing, Dalian, China, 19–22 October 2011; IEEE: Piscataway, NJ, USA, 2011; pp. 266–272. [Google Scholar]
  68. Cornwell, B.; Laumann, E.O.; Schumm, L.P. The social connectedness of older adults: A national profile. Am. Sociol. Rev. 2008, 73, 185–203. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. Cornwell, B. Age trends in daily social contact patterns. Res. Aging 2011, 33, 598–631. [Google Scholar] [CrossRef] [Scilit]
  70. Marcum, C.S. Age differences in daily social activities. Res. Aging 2013, 35, 612–640. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Pérez-Rosés, H.; Sebé, F. Synthetic generation of social network data with endorsements. J. Simul. 2015, 9, 279–286. [Google Scholar] [CrossRef] [Scilit]
  72. Nettleton, D.F. A synthetic data generator for online social network graphs. Soc. Netw. Anal. Min. 2016, 6, 44. [Google Scholar] [CrossRef] [Scilit]
  73. Sagduyu, Y.E.; Grushin, A.; Shi, Y. Synthetic Social Media Data Generation. IEEE Trans. Comput. Soc. Syst. 2018, 5, 605–620. [Google Scholar] [CrossRef]
  74. Facebook Analytics. Available online: https://www.facebook.com/analytics/1701892993437661/?-section=people_demographics/ (accessed on 31 July 2018).
  75. Fan Page List. 2015. Available online: http://www.fanpagelist.com/category/top_users/ (accessed on 6 August 2018).
  76. Koren, Y.; Bell, R.; Volinsky, C. Matrix factorization techniques for recommender systems. Computer 2009, 42, 30–37. [Google Scholar] [CrossRef] [Scilit]
  77. Han, X.; Wang, L.; Park, S.; Cuevas, A.; Crespi, N. Alike people, alike interests? A large-scale study on interest similarity in social networks. In Proceedings of the 2014 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining, Beijing, China, 17–20 August 2014; IEEE Press: Piscataway, NJ, USA, 2014; pp. 491–496. [Google Scholar]
  78. Dillon, M. Introduction to modern information retrieval: G. Salton and M. McGill; McGraw-Hill: New York, NY, USA, 1983; pp. 491–496. ISBN 0-07-054484-0. [Google Scholar]
  79. Sinnott, R.W. Virtues of the Haversine. Sky Telesc. 1984, 68, 159. [Google Scholar]
  80. Palmer, M.C. Calculation of distance traveled by fishing vessels using GPS positional data: A theoretical evaluation of the sources of error. Fish. Res. 2008, 89, 57–64. [Google Scholar] [CrossRef] [Scilit]
  81. Jaccard, P. Distribution de la flore alpine dans le bassin des Dranses et dans quelques régions voisines. Bull. Soc. Vaud. Sci. Nat. 1901, 37, 241–272. [Google Scholar]
  82. Huberman, B.A.; Romero, D.M.; Wu, F. Social networks that matter: Twitter under the microscope. arXiv, 2008; arXiv:0812.1045.
  83. Leskovec, J.; Krevl, A. {SNAP Datasets}:{Stanford} Large Network Dataset Collection; Stanford University: Stanford, CA, USA, USA, 2014; Available online: http://snap.stanford.edu/data (accessed on 19 November 2018).
  84. Zafarani, R.; Liu, H. Social Computing Data Repository at ASU [“http://socialcomputing. asu. edu/”]. Tempe, AZ: Arizona State University, School of Computing. Inform. Decis. Syst. Eng. 2009.
  85. Yang, J.; Leskovec, J. Defining and evaluating network communities based on ground-truth. Knowl. Inf. Syst. 2015, 42, 181–213. [Google Scholar] [CrossRef] [Scilit]
  86. Leskovec, J.; Mcauley, J.J. Learning to discover social circles in ego networks. In Advances in Neural Information Processing Systems; Neural Information Processing Systems (NIPS) Foundation, Inc.: La Jolla, CA, USA, 2012; pp. 539–547. [Google Scholar]
  87. Zachary, W.W. An information flow model for conflict and fission in small groups. J. Anthropol. Res. 1977, 33, 452–473. [Google Scholar] [CrossRef] [Scilit]
  88. Knuth, D.E. The Stanford GraphBase: A Platform for Combinatorial Computing; ACM Press: New York, NY, USA, 1993. [Google Scholar]
  89. Klimt, B.; Yang, Y. Introducing the Enron Corpus. In Proceedings of the 2004 CEAS, Mountain View, CA, USA, 30–31 July 2004. [Google Scholar]
  90. Leskovec, J.; Lang, K.J.; Dasgupta, A.; Mahoney, M.W. Community structure in large networks: Natural cluster sizes and the absence of large well-defined clusters. Internet Math. 2009, 6, 29–123. [Google Scholar] [CrossRef] [Scilit]
  91. Panzarasa, P.; Opsahl, T.; Carley, K.M. Patterns and dynamics of users’ behavior and interaction: Network analysis of an online community. J. Am. Soc. Inf. Sci. Technol. 2009, 60, 911–932. [Google Scholar] [CrossRef] [Scilit]
  92. Stehlé, J.; Voirin, N.; Barrat, A.; Cattuto, C.; Isella, L.; Pinton, J.F.; Quaggiotto, M.; Van den Broeck, W.; Régis, C.; Lina, B.; et al. High-resolution measurements of face-to-face contact patterns in a primary school. PLoS ONE 2011, 6, e23176. [Google Scholar] [CrossRef] [Scilit] [PubMed]

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.