Community Detection for Contextual-LSBM: Theoretical Limitations of Misclassification Rate and Efficient Algorithms

Jin, Dian; Zhang, Yuqian; Zhang, Qiaosheng

Statistics > Machine Learning

arXiv:2501.11139 (stat)

[Submitted on 19 Jan 2025 (v1), last revised 27 Jan 2025 (this version, v3)]

Title:Community Detection for Contextual-LSBM: Theoretical Limitations of Misclassification Rate and Efficient Algorithms

Authors:Dian Jin, Yuqian Zhang, Qiaosheng Zhang

View PDF HTML (experimental)

Abstract:The integration of network information and node attribute information has recently gained significant attention in the community detection literature. In this work, we consider community detection in the Contextual Labeled Stochastic Block Model (CLSBM), where the network follows an LSBM and node attributes follow a Gaussian Mixture Model (GMM). Our primary focus is the misclassification rate, which measures the expected number of nodes misclassified by community detection algorithms. We first establish a lower bound on the optimal misclassification rate that holds for any algorithm. When we specialize our setting to the LSBM (which preserves only network information) or the GMM (which preserves only node attribute information), our lower bound recovers prior results. Moreover, we present an efficient spectral-based algorithm tailored for the CLSBM and derive an upper bound on its misclassification rate. Although the algorithm does not attain the lower bound, it serves as a reliable starting point for designing more accurate community detection algorithms (as many algorithms use spectral method as an initial step, followed by refinement procedures to enhance accuracy).

Comments:	online version for Isit-25 submission
Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as:	arXiv:2501.11139 [stat.ML]
	(or arXiv:2501.11139v3 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2501.11139

Submission history

From: Dian Jin [view email]
[v1] Sun, 19 Jan 2025 18:51:15 UTC (24 KB)
[v2] Thu, 23 Jan 2025 20:26:54 UTC (25 KB)
[v3] Mon, 27 Jan 2025 17:38:03 UTC (26 KB)

Statistics > Machine Learning

Title:Community Detection for Contextual-LSBM: Theoretical Limitations of Misclassification Rate and Efficient Algorithms

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Community Detection for Contextual-LSBM: Theoretical Limitations of Misclassification Rate and Efficient Algorithms

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators