A high-performance clustering algorithm based on searched experiences.

Clustering is a traditional data mining problem that has attracted researchers from different disciplines because its solution can be applied to many useful problems in our daily life. Since the era of big data is coming, how to "reduce the computing time" of an "effective clustering algorithm" has...

Descripción completa

Detalles Bibliográficos
Publicado en:Computers in Human Behavior Vol. 100; pp. 231 - 242
Autores principales: Tsai, Chun-Wei, Ding, Yong-Chun, Liu, Shi-Jui, Chiang, Ming-Chao, Yang, Chu-Sing
Formato: Artículo
Publicado: Elsevier B.V. Nov2019
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Clustering is a traditional data mining problem that has attracted researchers from different disciplines because its solution can be applied to many useful problems in our daily life. Since the era of big data is coming, how to "reduce the computing time" of an "effective clustering algorithm" has been a promising research issue in recent years. Thus, this paper presents an effective clustering algorithm, by using the so-called searched information to determine later search directions, and then has it implemented on Spark to accelerate its response time for analyzing large-scale datasets. Simulation results show that the proposed algorithm provides a better result than the other clustering algorithms compared in this paper because it is less sensitive to the initial solutions. The simulation results further show that cloud computing platform is capable of enhancing the performance of the proposed algorithm. • This study presents a high-performance clustering algorithm based on searched experiences. • The results show that it is able to find a better result than traditional clustering algorithms. • We apply it to Spark to show its possibility to accelerate the response time on a cloud system.