Spider Community Detection: Seeded Geodesic Expansion with Modularity-Guided Refinement and Greedy Merge Matching
Abstract
1. Introduction
1.1. Background and Motivation
1.2. Challenges in Community Detection
1.3. Contributions of This Work
- We introduce a new seed-selection strategy combining degree, triangle participation, and local clustering into a composite score that favors structurally cohesive starting points.
- We define the spider graph expansion mechanism, which incrementally grows communities using modularity and triangle-aware decision rules, naturally capturing multi-layered network topology.
- We integrate this process with a global refinement pipeline—including Louvain-style local moves, -based merging, and small-fragment cleanup—that jointly improve boundary precision and structural stability.
1.4. Paper Organization
2. Literature Review and Background
2.1. Global Modularity Optimization
2.2. Local Expansion and Seed-Based Methods
2.3. Geodesic Distance Modularity and Local Quality
2.4. Pairwise and Information-Theoretic Evaluation
2.5. Comparison with Hybrid Approaches
2.5.1. Structural Differences from Existing Methods
- OSLOM and significance-based methods.
- Core-periphery and local spectral methods.
- Louvain variants and modularity-driven refinement.
- Flow-based and information-theoretic methods.
2.5.2. Why Spider’s Design Leads to Unique Gains
- Low conductance + high modularity: By accepting nodes only when they satisfy both connectivity ratio and modularity gain thresholds, Spider produces communities with sharp boundaries while maintaining global coherence. Pure modularity optimizers (Leiden) achieve high Q but exhibit higher conductance; local methods (Label Propagation) achieve low conductance locally but fragment the graph.
- Geodesic compactness: The triangle-closure criterion and depth-bounded expansion naturally select nodes that close short cycles, resulting in communities with low internal geodesic distances. This is reflected in Spider’s superior weighted GDM scores, a property not optimized by edge-density or flow-based methods.
- Balanced granularity: Farthest-first seed selection prevents the clustering of seeds in dense regions, avoiding both the over-fragmentation of label propagation (hundreds of micro-communities on sparse graphs, and the over-coarsening of Infomap (collapsing heterogeneous structures into single modules). This balance is achieved without manual tuning through adaptive parameter selection based on graph statistics.
2.6. Dynamic and Temporal Network Community Detection
3. Spider Community Detection Algorithm
3.1. Preliminaries and Notation
3.2. Seed Quality via Triangle–Degree–Clustering Scoring
- The number of incident triangles ;
- The degree ;
- The local clustering coefficient , computed as the ratio of closed triangles to possible triangles among the neighbors of v.
| Algorithm 1 Seed scoring |
| Require: Graph , sparsity threshold Ensure: Seed score for all 1: Compute degree , triangle count , and local clustering coefficient for all 2: Normalize: , , 3: Compute average degree 4: if then▹ Sparse graph 5: for all v ▹ Emphasize degree 6: else ▹ Dense graph 7: for all v ▹ Balanced weights 8: return s |
| Algorithm 2 Diversified center selection |
| Require: Unlabeled nodes U, seed scores s, chosen centers , randomness , pool size K Ensure: Next spider center c 1: Sort U by in descending order and select top-K as pool P 2: if then 3: return highest-scoring node from P (with probability ) or random from P (with probability ) 4: Sample candidate set (deterministically or with exploration) 5: for each do 6: using limited BFS ▹ Geodesic separation 7: return ▹ Maximize distance and quality |
3.3. Local Objective and Spider Expansion
3.4. Depth-Bounded Geodesic Expansion
| Algorithm 3 Spider expansion |
| Require: Graph , seed , max depth R, acceptance parameters Ensure: Community C grown from seed 1: Initialize , track , 2: Collect candidates via BFS up to depth R from 3: for each depth level do 4: Sort candidates at depth ℓ by descending 5: for each candidate v at depth ℓ do 6: Compute , , , 7: Evaluate acceptance criteria based on , , , , and sparsity 8: if then 9: , update and 10: return C |
3.5. Straggler Attachment and Local Refinement
| Algorithm 4 Straggler attachment |
| Require: Graph , labeling , soft tolerance Ensure: Updated labeling with all nodes assigned 1: unlabeled vertices 2: for each do 3: Find neighboring communities 4: 5: if then 6: ▹ Attach to best community 7: else 8: Create singleton community for u 9: return |
| Algorithm 5 Local modularity refinement |
| Require: Graph , labeling , max passes , threshold Ensure: Refined labeling 1: for pass to do 2: False 3: Select candidates (boundary nodes in early passes, all nodes in later passes) 4: for each in random order do 5: 6: Compute removal gain from 7: ▹ Best neighbor community 8: if then 9: , True 10: if not improved then 11: break 12: return |
3.6. Greedy Merge Matching of Neighboring Communities
- The number of edges crossing between and ;
- The volumes and ;
- The internal edge counts and .
| Algorithm 6 Adaptive merge matching |
| Require: Graph , labeling , merge parameters based on sparsity and fragmentation Ensure: Merged labeling 1: Determine (rounds), (size threshold), (edge threshold), (conductance) based on graph structure 2: for round to do 3: Build candidate merge pairs with: 4: • Inter-community edges 5: • At least one community 6: • Positive modularity gain 7: • Merged conductance below 8: if no candidates then 9: break 10: Perform greedy maximum-weight matching on candidates by 11: Apply merges and refine boundaries 12: return |
| Algorithm 7 Tiny community absorption |
| Require: Graph , labeling , size threshold Ensure: Updated labeling 1: for each community c with do 2: Find neighboring communities 3: 4: if exists then 5: Merge c into 6: return |
3.7. Algorithm Outline and Complexity
| Algorithm 8 Spider Community Detection |
| Require: Graph , randomness parameter , depth bound R 1: Compute , , (clustering coefficient) and seed scores for all 2: Initialize labeling for all v 3: , 4: while there exist unlabeled vertices do 5: Select a new center c among top-K unlabeled vertices by diversified seed scoring 6: Grow spider C from c by depth-bounded BFS with adaptive multi-tier modularity- and triangle-based acceptance 7: for all do 8: 9: , 10: Attach unlabeled vertices using modularity-guided straggler rule with soft tolerance 11: Perform local Louvain-style refinement of 12: Determine adaptive merge parameters based on fragmentation ratio and graph sparsity 13: Initialize round counter 14: while there exists a profitable merge and do 15: Compute candidate community merges and their 16: Select a greedy maximum-weight matching of profitable merges 17: Apply merges and perform a short local refinement 18: 19: Merge tiny communities below threshold 20: Perform post-merge refinement 21: Final tiny merge pass with reduced threshold 22: Final refinement 23: return Final labeling |
4. Experimental Evaluation
4.1. Datasets and Experimental Setup
4.2. Evaluation Metrics
- Ground-truth recovery metrics.
- Structural quality metrics.
- Geodesic Distance Modularity aggregation.
4.3. Performance on Real-World Networks
4.3.1. Ground-Truth Recovery: NMI, ARI, and F1
4.3.2. Structural Quality: Modularity and Conductance
4.3.3. Geodesic Cohesion: Weighted Geodesic Distance Modularity
4.3.4. Community Granularity: Number and Average Size of Communities
4.3.5. Qualitative Spider Progression on the Primary School Network
4.4. Performance on LFR Benchmark Networks
Robustness Analysis with Respect to Mixing Parameter
5. Conclusions and Future Work
- Theoretical analysis: Establishing provable guarantees for Spider remains an interesting challenge. For example, one could study conditions under which the local modularity and triangle-based objective recovers planted communities in stochastic block models or geometric random graphs.
- Overlapping and hierarchical communities: The current formulation yields a hard partition of the vertex set. Extending Spider to allow overlapping or hierarchical spider structures could better model networks where nodes naturally participate in multiple modules.
- Scalability and dynamic graphs: While the present implementation already scales well to medium-to-large networks, there is room for further optimization, e.g., through incremental updates of local statistics or streaming variants that handle edge insertions and deletions.
- Integration with sparsification: Combining Spider with metric-backbone or other sparsification techniques prior to community detection is a promising avenue to reduce computational cost and sharpen community boundaries.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Fortunato, S. Community detection in graphs. Phys. Rep. 2010, 486, 75–174. [Google Scholar] [CrossRef] [Scilit]
- Newman, M.E.J. Modularity and community structure in networks. Proc. Natl. Acad. Sci. USA 2006, 103, 8577–8582. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Newman, M.E.J.; Girvan, M. Finding and evaluating community structure in networks. Phys. Rev. E 2004, 69, 026113. [Google Scholar] [CrossRef] [Scilit]
- Blondel, V.D.; Guillaume, J.L.; Lambiotte, R.; Lefebvre, E. Fast unfolding of communities in large networks. J. Stat. Mech. Theory Exp. 2008, 2008, P10008. [Google Scholar] [CrossRef] [Scilit]
- Ng, A.Y.; Jordan, M.I.; Weiss, Y. On spectral clustering: Analysis and an algorithm. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 3–5 December 2002; Volume 14. [Google Scholar]
- Fortunato, S.; Barthélemy, M. Resolution limit in community detection. Proc. Natl. Acad. Sci. USA 2007, 104, 36–41. [Google Scholar] [CrossRef] [Scilit]
- Raghavan, U.N.; Albert, R.; Kumara, S. Near linear time algorithm to detect community structures in large-scale networks. Phys. Rev. E 2007, 76, 036106. [Google Scholar] [CrossRef] [Scilit]
- Rosvall, M.; Bergstrom, C.T. Maps of random walks on complex networks reveal community structure. Proc. Natl. Acad. Sci. USA 2008, 105, 1118–1123. [Google Scholar] [CrossRef] [Scilit]
- Lancichinetti, A.; Radicchi, F.; Ramasco, J.J.; Fortunato, S. Finding statistically significant communities in networks. PLoS ONE 2011, 6, e18961. [Google Scholar] [CrossRef] [Scilit]
- Shen, H.; Cheng, X.; Cai, K.; Hu, M.B. Detect overlapping and hierarchical community structure in networks. Phys. A Stat. Mech. Its Appl. 2009, 388, 1706–1712. [Google Scholar] [CrossRef] [Scilit]
- Fortunato, S.; Hric, D. Community detection in networks: A user guide. Phys. Rep. 2016, 659, 1–44. [Google Scholar] [CrossRef] [Scilit]
- Ahn, Y.Y.; Bagrow, J.P.; Lehmann, S. Link communities reveal multiscale complexity in networks. Nature 2010, 466, 761–764. [Google Scholar] [CrossRef] [Scilit]
- Lancichinetti, A.; Fortunato, S. Community detection algorithms: A comparative analysis. Phys. Rev. E 2009, 80, 056117. [Google Scholar] [CrossRef] [Scilit]
- Benson, A.R.; Gleich, D.F.; Leskovec, J. Higher-order organization of complex networks. Science 2016, 353, 163–166. [Google Scholar] [CrossRef] [Scilit]
- Bakhtar, S.; Harutyunyan, H.A. A new metric to compare local community detection algorithms in social networks using geodesic distance. J. Comb. Optim. 2022, 44, 2809–2831. [Google Scholar] [CrossRef] [Scilit]
- Danon, L.; Díaz-Guilera, A.; Duch, J.; Arenas, A. Comparing community structure identification. J. Stat. Mech. Theory Exp. 2005, 2005, P09008. [Google Scholar] [CrossRef] [Scilit]
- Andersen, R.; Chung, F.; Lang, K. Local graph partitioning using PageRank vectors. In Proceedings of the 47th Annual IEEE Symposium on Foundations of Computer Science (FOCS’06), Berkeley, CA, USA, 21–24 October 2006; pp. 475–486. [Google Scholar]
- Andersen, R.; Lang, K.J. Communities from seed sets. In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Paris, France, 28 June–1 July 2009; pp. 223–232. [Google Scholar]
- Traag, V.A.; Waltman, L.; van Eck, N.J. From Louvain to Leiden: Guaranteeing well-connected communities. Sci. Rep. 2019, 9, 5233. [Google Scholar] [CrossRef] [Scilit]
- Li, D.; Kosugi, S.; Zhang, Y.; Okumura, M.; Xia, F.; Jiang, R. Revisiting Dynamic Graph Clustering via Matrix Factorization. In Proceedings of the ACM on Web Conference 2025, Taipei, Taiwan, 13–17 April 2025; pp. 1342–1352. [Google Scholar] [CrossRef] [Scilit]
- Li, D.; Ma, X.; Gong, M. Joint Learning of Feature Extraction and Clustering for Large-Scale Temporal Networks. IEEE Trans. Cybern. 2023, 53, 1653–1666. [Google Scholar] [CrossRef] [Scilit]
- Zachary, W.W. An information flow model for conflict and fission in small groups. J. Anthropol. Res. 1977, 33, 452–473. [Google Scholar] [CrossRef] [Scilit]
- Mastrandrea, R.; Fournet, J.; Barrat, A. Contact patterns in a high school: A comparison between data collected using wearable sensors, contact diaries and friendship surveys. PLoS ONE 2015, 10, e0136497. [Google Scholar] [CrossRef] [Scilit]
- Stehlé, J.; Voirin, N.; Barrat, A.; Cattuto, C.; Isella, L.; Pinton, J.F.; Quaggiotto, M.; Van den Broeck, W.; Regis, C.; Lina, B. High-resolution measurements of face-to-face contact patterns in a primary school. PLoS ONE 2011, 6, e23176. [Google Scholar] [CrossRef] [Scilit]
- Adamic, L.A.; Glance, N. The political blogosphere and the 2004 US election: Divided they blog. In Proceedings of the 3rd International Workshop on Link Discovery, Chicago, IL, USA, 21 August 2005; pp. 36–43. [Google Scholar]
- Sen, P.; Namata, G.; Bilgic, M.; Getoor, L.; Galligher, B.; Eliassi-Rad, T. Collective classification in network data. AI Mag. 2008, 29, 93. [Google Scholar] [CrossRef] [Scilit]
- Leskovec, J.; Krevl, A. SNAP: Stanford Large Network Dataset Collection. 2014. Available online: https://snap.stanford.edu/data (accessed on 15 September 2025).
- Ley, M. DBLP—Some lessons learned. In Proceedings of the VLDB Endowment, Hong Kong, China, 30 August–2 September 2002; pp. 1493–1500. [Google Scholar]
- Lancichinetti, A.; Fortunato, S.; Kertesz, J. Detecting the overlapping and hierarchical community structure in complex networks. New J. Phys. 2009, 11, 033015. [Google Scholar] [CrossRef] [Scilit]
- Clauset, A. Finding local community structure in networks. Phys. Rev. E 2005, 72, 026132. [Google Scholar] [CrossRef] [Scilit] [PubMed]










| Regime | |||
|---|---|---|---|
| Bootstrap () | 1 | ||
| Growth () | 2 | ||
| Large () | 0 | 3 | |
| Sparse override () | 0 | 2 |
| Dataset | Nodes | Edges | GT Comm. | Avg. Degree |
|---|---|---|---|---|
| Karate Club | 34 | 78 | 2 | 4.59 |
| High School | 327 | 5818 | 9 | 35.58 |
| Primary School | 242 | 8317 | 11 | 68.74 |
| Political Blogs | 1222 | 16,717 | 2 | 27.36 |
| CiteSeer | 2110 | 3668 | 6 | 3.48 |
| Cora | 2485 | 5069 | 7 | 4.08 |
| Wiki Schools | 4403 | 100,382 | 16 | 45.60 |
| DBLP | 3762 | 17,587 | 193 | 9.35 |
| Amazon | 8035 | 183,663 | 195 | 45.72 |
| Dataset | Nodes | Edges | wGDM |
|---|---|---|---|
| Karate Club | 34 | 78 | 0.9745 |
| High School | 327 | 5818 | 0.8838 |
| Primary School | 242 | 8317 | 0.8007 |
| Political Blogs | 1222 | 16,717 | 0.9782 |
| CiteSeer | 2110 | 3668 | 0.4035 |
| Cora | 2485 | 5069 | 0.3747 |
| Wiki Schools | 4403 | 100,382 | 0.8355 |
| DBLP | 3762 | 17,587 | 0.4460 |
| Amazon | 8035 | 183,663 | 0.4915 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Harutyunyan, H.A.; Kamalipour, P. Spider Community Detection: Seeded Geodesic Expansion with Modularity-Guided Refinement and Greedy Merge Matching. Computers 2026, 15, 83. https://doi.org/10.3390/computers15020083
Harutyunyan HA, Kamalipour P. Spider Community Detection: Seeded Geodesic Expansion with Modularity-Guided Refinement and Greedy Merge Matching. Computers. 2026; 15(2):83. https://doi.org/10.3390/computers15020083
Chicago/Turabian StyleHarutyunyan, Hovhannes A., and Parsa Kamalipour. 2026. "Spider Community Detection: Seeded Geodesic Expansion with Modularity-Guided Refinement and Greedy Merge Matching" Computers 15, no. 2: 83. https://doi.org/10.3390/computers15020083
APA StyleHarutyunyan, H. A., & Kamalipour, P. (2026). Spider Community Detection: Seeded Geodesic Expansion with Modularity-Guided Refinement and Greedy Merge Matching. Computers, 15(2), 83. https://doi.org/10.3390/computers15020083

