Mr. Latte


Compression-based Anomaly Detection using K-means Clustering

This English page is a machine translation of the Korean original, so some wording may read a little off. Read the Korean original

This study presents a new method for storing large log data, and simultaneously, detecting anomaly data. To achieve this, the well-known K-means clustering algorithm is used for the anomaly detection. In K-means algorithm, the dissimilarity between data is calculated on the space transformed by the Logpack compression algorithm. We also performed a feature selection using genetic algorithms to obtain an informative subset of features relevant to anomaly events. Through various tests, it is observed that the proposed method is superior to conventional algorithms.

View details

Looking for a product partner or a lecture? Founders, teams, businesses: from problem framing to launch, plus development lectures shaped for teams and schools.