Title: Clusterdv: a simple density-based clustering method that is robust, general and automatic

Journal Article · · Bioinformatics

Abstract Motivation How to partition a dataset into a set of distinct clusters is a ubiquitous and challenging problem. The fact that data vary widely in features such as cluster shape, cluster number, density distribution, background noise, outliers and degree of overlap, makes it difficult to find a single algorithm that can be broadly applied. One recent method, clusterdp, based on search of density peaks, can be applied successfully to cluster many kinds of data, but it is not fully automatic, and fails on some simple data distributions. Results We propose an alternative approach, clusterdv, which estimates density dips between points, and allows robust determination of cluster number and distribution across a wide range of data, without any manual parameter adjustment. We show that this method is able to solve a range of synthetic and experimental datasets, where the underlying structure is known, and identifies consistent and meaningful clusters in new behavioral data. Availability and implementation The clusterdv is implemented in Matlab. Its source code, together with example datasets are available on: https://github.com/jcbmarques/clusterdv. Supplementary information Supplementary data are available at Bioinformatics online.

Sponsoring Organization:
USDOE Office of Nuclear Energy (NE), Fuel Cycle Technologies (NE-5); USDOE Office of Nuclear Energy (NE), Nuclear Fuel Cycle and Supply Chain
OSTI ID:
1562941
Journal Information:
Bioinformatics, Journal Name: Bioinformatics Journal Issue: 12 Vol. 35; ISSN 1367-4803
Publisher:
Oxford University PressCopyright Statement
Country of Publication:
United Kingdom
Language:
English

References (26)

Structure of the Zebrafish Locomotor Repertoire Revealed with Unsupervised Behavioral Clustering journal January 2018
Iterative shrinking method for clustering problems journal May 2006
Robust path-based spectral clustering journal January 2008
Data clustering: 50 years beyond K-means journal June 2010
:{Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data journal January 2003
Comparing the performance of biomedical clustering methods journal September 2015
Clustering by fast search and merge of local density peaks for gene expression microarray data journal April 2017
Variable Kernel Estimates of Multivariate Densities journal May 1977
An Information Flow Model for Conflict and Fission in Small Groups journal December 1977
On the shortest spanning subtree of a graph and the traveling salesman problem journal January 1956
SLINK: An optimally efficient algorithm for the single-link cluster method journal January 1973
Chameleon: hierarchical clustering using dynamic modeling journal January 1999
Gradient-based learning applied to document recognition journal January 1998
Complex Wavelet Structural Similarity: A New Image Similarity Index journal November 2009
Least squares quantization in PCM journal March 1982
Survey of Clustering Algorithms journal May 2005
A maximum variance cluster algorithm journal September 2002
Estimating the number of clusters in a data set via the gap statistic
  • Tibshirani, Robert; Walther, Guenther; Hastie, Trevor
  • Journal of the Royal Statistical Society: Series B (Statistical Methodology), Vol. 63, Issue 2, p. 411-423 https://doi.org/10.1111/1467-9868.00293
journal May 2001
Clustering by Passing Messages Between Data Points journal February 2007
Clustering by fast search and find of density peaks journal June 2014
Clustering aggregation journal March 2007
OPTICS: ordering points to identify the clustering structure journal June 1999
Lower Bounds for the Partitioning of Graphs journal September 1973
Fast clustering using adaptive density peak detection journal October 2015
FLAME, a novel fuzzy clustering method for the analysis of DNA microarray data journal January 2007
Sensorimotor Gating in Larval Zebrafish journal May 2007