期刊文献+

属性加权的类属型数据非模聚类 被引量:7

Non-Mode Clustering of Categorical Data with Attributes Weighting
在线阅读 下载PDF
导出
摘要 类属型数据广泛分布于生物信息学等许多应用领域,其离散取值的特点使得类属数据聚类成为统计机器学习领域一项困难的任务.当前的主流方法依赖于类属属性的模进行聚类优化和相关属性的权重计算.提出一种非模的类属型数据统计聚类方法.首先,基于新定义的相异度度量,推导了属性加权的类属数据聚类目标函数.该函数以对象与簇之间的平均距离为基础,从而避免了现有方法以模为中心导致的问题.其次,定义了一种类属型数据的软子空间聚类算法.该算法在聚类过程中根据属性取值的总体分布,而不仅限于属性的模,赋予每个属性衡量其与簇类相关程度的权重,实现自动的特征选择.在合成数据和实际应用数据集上的实验结果表明,与现有的基于模的聚类算法和基于蒙特卡罗优化的其他非模算法相比,该算法有效地提高了聚类结果的质量. While categorical data are widely used in many applications such as Bioinformatics, clustering categorical data is a difficult task in the filed of statistical machine learning due to the characteristic of the data which can only take discrete values. Typically, the mainstream methods are dependent on the mode of the categorical attributes in order to optimize the clusters and weight the relevant attributes. A non-mode approach is proposed for statistically clustering of categorical data in this paper. First, based on a newly defined dissimilarity measure, an objective function with attributes weighting is derived for categorical data clustering. The objective function is defined based on the average distance between the objects and the clusters, therefore overcomes the problems in the existing methods based on the mode category. Then, a soft-subspace clustering algorithm is proposed for clustering categorical data. In this algorithm, each attribute is assigned with weights measuring its degree of relevance to the clusters in terms of the overall distribution of categories instead of the mode category, enabling automatic feature selection during the clustering process. Experimental results carried out on some synthetic datasets and real-world datasets demonstrate that the proposed method significantly improves clustering quality.
出处 《软件学报》 EI CSCD 北大核心 2013年第11期2628-2641,共14页 Journal of Software
基金 国家自然科学基金(61175123)
关键词 聚类 类属型数据 属性加权 clustering categorical data mode attribute weighting
  • 相关文献

参考文献3

二级参考文献32

共引文献1163

同被引文献48

  • 1ALPAYDIN E.机器学习导论[M].北京:机械工业出版社,2009:245-251.
  • 2Roiger R J,Geatz M W.数据挖掘教程[M].翁敬农,译.北京:清华大学出版社,2003.
  • 3孙吉贵,刘杰,赵连宇.聚类算法研究[J].软件学报,2008(1):48-61. 被引量:1108
  • 4GETHSIYAL AUGASTA M, KATHIRVALAVAKUMAR T. A new discretization algorithm based on range coefficient of dispersion and skewers for neural networks classifier [ J]. Applied Soft Computing, 2012, 12(2): 619-625.
  • 5WONG T-T. A hybrid discretization method for naive Bayesian c/as- sifiers [ J]. Pattern Recognition, 2012, 45(6) : 2321 - 2325.
  • 6HUANG W, PAN Y, WU J. Supervised discretization with GK-t [ J]. Procedia Computer Science, 2013, 17:114 -120.
  • 7FERREIRA A J, FERREIRA A J. An unsupervised approach to feature discretization and selection [J]. Pattern Recognition, 2012, 45(9): 3048 -3060.
  • 8GUPTA A, MEHROTRAB K G, MOHANB C. A clustering-based discretization for supervised learning [ J]. Statistics and Probability Letters, 2010, 80(9): 816-824.
  • 9CHEN S, TANG L, LIU W, et al. An improved method of disereti- zation of continuous attributes [ J]. Procedia Environmental Sci- ences, 2011, 11(A) : 213 -217.
  • 10MONTALVAO J, CANUTOB J. Clustering ensembles and space dis- cretizatiou--a new regard toward diversity and consensus [ J]. Pat- tern Recognition Letters, 2010, 15(1) : 2415 -2424.

引证文献7

二级引证文献68

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部