Evaluation of clustering algorithms for gene expression data

被引:30
作者
Datta, Susmita [1 ]
Datta, Somnath [1 ]
机构
[1] Univ Louisville, Dept Bioinformat & Biostat, Louisville, KY 40202 USA
关键词
D O I
10.1186/1471-2105-7-S4-S17
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Background: Cluster analysis is an integral part of high dimensional data analysis. In the context of large scale gene expression data, a filtered set of genes are grouped together according to their expression profiles using one of numerous clustering algorithms that exist in the statistics and machine learning literature. A closely related problem is that of selecting a clustering algorithm that is "optimal" in some sense from a rather impressive list of clustering algorithms that currently exist. Results: In this paper, we propose two validation measures each with two parts: one measuring the statistical consistency (stability) of the clusters produced and the other representing their biological functional congruence. Smaller values of these indices indicate better performance for a clustering algorithm. We illustrate this approach using two case studies with publicly available gene expression data sets: one involving a SAGE data of breast cancer patients and the other involving a time course cDNA microarray data on yeast. Six well known clustering algorithms UPGMA, K-Means, Diana, Fanny, Model-Based and SOM were evaluated. Conclusion: No single clustering algorithm may be best suited for clustering genes into functional groups via expression profiles for all data sets. The validation measures introduced in this paper can aid in the selection of an optimal algorithm, for a given data set, from a collection of available clustering algorithms.
引用
收藏
页数:9
相关论文
共 17 条
  • [1] ABBA MC, 2004, BMC BIOINFORMATICS, V6, P5
  • [2] [Anonymous], 1998, MODERN APPL STAT S P
  • [3] MODEL-BASED GAUSSIAN AND NON-GAUSSIAN CLUSTERING
    BANFIELD, JD
    RAFTERY, AE
    [J]. BIOMETRICS, 1993, 49 (03) : 803 - 821
  • [4] The transcriptional program of sporulation in budding yeast
    Chu, S
    DeRisi, J
    Eisen, M
    Mulholland, J
    Botstein, D
    Brown, PO
    Herskowitz, I
    [J]. SCIENCE, 1998, 282 (5389) : 699 - 705
  • [5] Comparisons and validation of statistical clustering techniques for microarray gene expression data
    Datta, S
    Datta, S
    [J]. BIOINFORMATICS, 2003, 19 (04) : 459 - 466
  • [6] DATTA S, 2002, ADV STAT COMBINATORI, P63
  • [7] Methods for evaluating clustering algorithms for gene expression data using a reference set of functional classes
    Datta, Susmita
    Datta, Somnath
    [J]. BMC BIOINFORMATICS, 2006, 7 (1)
  • [8] Dudoit S, 2002, GENOME BIOL, V3
  • [9] Scoring clustering solutions by their biological relevance
    Gat-Viks, I
    Sharan, R
    Shamir, R
    [J]. BIOINFORMATICS, 2003, 19 (18) : 2381 - 2389
  • [10] Hartigan J. A., 1979, Applied Statistics, V28, P100, DOI 10.2307/2346830