Title: Parallel processing for stepwise generalisation method on multi-core PC cluster

Authors: Shinpei Yagi; Keiichi Tamura; Hajime Kitakami

Addresses: Graduate School of Information Sciences, Hiroshima City University, 3-4-1, Ozuka-Higashi, Asa-Minami-Ku, Hiroshima, 731-3194, Japan. ' Graduate School of Information Sciences, Hiroshima City University, 3-4-1, Ozuka-Higashi, Asa-Minami-Ku, Hiroshima, 731-3194, Japan. ' Graduate School of Information Sciences, Hiroshima City University, 3-4-1, Ozuka-Higashi, Asa-Minami-Ku, Hiroshima, 731-3194, Japan

Abstract: An approximate query, which is an approximate pattern matching in sequence databases, is one of the most important techniques for many different areas, such as computational biology, text mining, web intelligence and pattern recognition; it returns many similar sub-sequences. In this paper, we refer to a set of such similar sub-sequences as a mismatch cluster. To support users who execute an approximate query on a sequence database to find the regularities of approximate patterns that similar to the query pattern, we have developed the stepwise generalisation method that extracts a reduced expression, called a minimum generalised set, from a mismatch cluster. This paper proposes a novel parallelisation model with a hierarchical task pool for the parallel processing of the stepwise generalisation method on a multi-core PC cluster. To manage tasks efficiently on multi-core CPUs, the proposed model uses the hierarchical task pool and an efficient hierarchical dynamic load balancing technique. We evaluate the proposed method using real protein sequences on an actual multi-core PC cluster. Experimental results confirm that the proposed method performs well on multi-core CPUs and on a multi-core PC cluster.

Keywords: text mining; parallel processing; multi-core PC clusters; dynamic load balancing; approximate query; approximate pattern matching; sequence databases; multi-core CPUs.

DOI: 10.1504/IJKWI.2012.050282

International Journal of Knowledge and Web Intelligence, 2012 Vol.3 No.2, pp.88 - 109

Published online: 04 Sep 2014 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article