A variable-selection heuristic for K-means clustering

被引:98
|
作者
Brusco, MJ [1 ]
Cradit, JD [1 ]
机构
[1] Florida State Univ, Coll Business, Dept Mkt, Tallahassee, FL 32306 USA
关键词
cluster analysis; K-means partitioning; variable selection; heuristics;
D O I
10.1007/BF02294838
中图分类号
O1 [数学];
学科分类号
0701 ; 070101 ;
摘要
One of the most vexing problems in cluster analysis is the selection and/or weighting of variables in order to include those that truly define cluster structure, while eliminating those that might mask such structure. This paper presents a variable-selection heuristic For nonhierarchical (K-means) cluster analysis based on the adjusted Rand index for measuring cluster recovery. The heuristic was subjected to Monte Carlo testing across more than 2200 datasets with known cluster structure. The results indicate the heuristic is extremely effective at eliminating masking variables. A cluster analysis of real-world financial services data revealed that using the variable-selection heuristic prior to the K-means algorithm resulted in greater cluster stability.
引用
收藏
页码:249 / 270
页数:22
相关论文
共 50 条
  • [1] A variable-selection heuristic for K-means clustering
    Michael J. Brusco
    J. Dennis Cradit
    Psychometrika, 2001, 66 : 249 - 270
  • [2] A Variable Selection Procedure for K-Means Clustering
    Kim, Sung-Soo
    KOREAN JOURNAL OF APPLIED STATISTICS, 2012, 25 (03) : 471 - 483
  • [3] Variable Selection and Outlier Detection for Automated K-means Clustering
    Kim, Sung-Soo
    COMMUNICATIONS FOR STATISTICAL APPLICATIONS AND METHODS, 2015, 22 (01) : 55 - 67
  • [4] Selection of K in K-means clustering
    Pham, DT
    Dimov, SS
    Nguyen, CD
    PROCEEDINGS OF THE INSTITUTION OF MECHANICAL ENGINEERS PART C-JOURNAL OF MECHANICAL ENGINEERING SCIENCE, 2005, 219 (01) : 103 - 119
  • [5] A heuristic K-means clustering algorithm by kernel PCA
    Xu, MT
    Fränti, P
    ICIP: 2004 INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, VOLS 1- 5, 2004, : 3503 - 3506
  • [6] Stability and model selection in k-means clustering
    Ohad Shamir
    Naftali Tishby
    Machine Learning, 2010, 80 : 213 - 243
  • [7] Deterministic Feature Selection for k-Means Clustering
    Boutsidis, Christos
    Magdon-Ismail, Malik
    IEEE TRANSACTIONS ON INFORMATION THEORY, 2013, 59 (09) : 6099 - 6110
  • [8] Stability and model selection in k-means clustering
    Shamir, Ohad
    Tishby, Naftali
    MACHINE LEARNING, 2010, 80 (2-3) : 213 - 243
  • [9] Variable neighborhood search algorithm for k-means clustering
    Orlov, V. I.
    Kazakovtsev, L. A.
    Rozhnov, I. P.
    Popov, N. A.
    Fedosov, V. V.
    IX INTERNATIONAL MULTIDISCIPLINARY SCIENTIFIC AND RESEARCH CONFERENCE MODERN ISSUES IN SCIENCE AND TECHNOLOGY / WORKSHOP ADVANCED TECHNOLOGIES IN AEROSPACE, MECHANICAL AND AUTOMATION ENGINEERING, 2018, 450
  • [10] Automated variable weighting in k-means type clustering
    Huang, JZX
    Ng, MK
    Rong, HQ
    Li, ZC
    IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2005, 27 (05) : 657 - 668