Title: Parallel, massive processing in SuperMatrix: a general tool for distributional semantic analysis of corpora

Authors: Bartosz Broda; Maciej Piasecki

Addresses: Institute of Informatics, Wrocław University of Technology, 50-370 Wrocław, Poland ' Institute of Informatics, Wrocław University of Technology, 50-370 Wrocław, Poland

Abstract: This article presents an extended version of the SuperMatrix system - a general tool supporting automatic acquisition of lexical semantic relations from corpora. Extensions focus mainly on parallel processing of massive amounts of data. The construction of the system is discussed. Three distributed parts of the system are presented, i.e., distributed construction of co-incidence matrices from corpora, computation of similarity matrix and parallel solving of synonymy tests. An evaluation of a proposed approach to parallel processing is shown. Parallelisation of similarity matrix computation demonstrates almost linear speedup. The smallest improvements were achieved for construction of matrices, as this process is mostly bound by reading huge amounts of data. Areas of application of the system are described.

Keywords: SuperMatrix; distributional semantics; parallel processing; semantic analysis; lexical semantic relations; corpora; co-incidence matrices; similarity matrix; synonyms.

DOI: 10.1504/IJDMMM.2013.051924

International Journal of Data Mining, Modelling and Management, 2013 Vol.5 No.1, pp.1 - 19

Published online: 29 Jul 2014 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article