TY - GEN
T1 - An Indicator for the Number of Clusters
T2 - 29th Annual Conference of the German Classification Society - Gesellschaft fur Klassifikation, GfKl 2005
AU - Weber, Marcus
AU - Rungsarityotin, Wasinee
AU - Schliep, Alexander
N1 - Publisher Copyright:
© 2026, Springer Science and Business Media Deutschland GmbH. All rights reserved.
PY - 2006
Y1 - 2006
N2 - The problem of clustering data can be formulated as a graph partitioning problem. In this setting, spectral methods for obtaining optimal solutions have received a lot of attention recently. We describe Perron Cluster Cluster Analysis (PCCA) and establish a connection to spectral graph partitioning. We show that in our approach a clustering can be efficiently computed by mapping the eigenvector data onto a simplex. To deal with the prevalent problem of noisy and possibly overlapping data we introduce the Min-chi indicator which helps in confirming the existence of a partition of the data and in selecting the number of clusters with quite favorable performance. Furthermore, if no hard partition exists in the data, the Min-chi can guide in selecting the number of modes in a mixture model. We close with showing results on simulated data generated by a mixture of Gaussians.
AB - The problem of clustering data can be formulated as a graph partitioning problem. In this setting, spectral methods for obtaining optimal solutions have received a lot of attention recently. We describe Perron Cluster Cluster Analysis (PCCA) and establish a connection to spectral graph partitioning. We show that in our approach a clustering can be efficiently computed by mapping the eigenvector data onto a simplex. To deal with the prevalent problem of noisy and possibly overlapping data we introduce the Min-chi indicator which helps in confirming the existence of a partition of the data and in selecting the number of clusters with quite favorable performance. Furthermore, if no hard partition exists in the data, the Min-chi can guide in selecting the number of modes in a mixture model. We close with showing results on simulated data generated by a mixture of Gaussians.
UR - https://www.scopus.com/pages/publications/105046559775
U2 - 10.1007/3-540-31314-1_11
DO - 10.1007/3-540-31314-1_11
M3 - Conference contribution
AN - SCOPUS:105046559775
SN - 9783540313137
T3 - Studies in Classification, Data Analysis, and Knowledge Organization
SP - 103
EP - 110
BT - From Data and Information Analysis to Knowledge Engineering - Proceedings of the 29th Annual Conference of the Gesellschaft für Klassifikation
A2 - Spiliopoulou, Myra
A2 - Kruse, Rudolf
A2 - Borgelt, Christian
A2 - Nürnberger, Andreas
A2 - Gaul, Wolfgang
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 9 March 2005 through 11 March 2005
ER -