Evaluation of clustering algorithms for gene expression data

被引：30

作者：

Datta, Susmita ^{[1
]}

Datta, Somnath ^{[1
]}

机构：

[1] Univ Louisville, Dept Bioinformat & Biostat, Louisville, KY 40202 USA

来源：

BMC BIOINFORMATICS | 2006年 / 7卷 / Suppl 4期

关键词：

D O I：

10.1186/1471-2105-7-S4-S17

中图分类号：

Q5 [生物化学];

学科分类号：

071010 ; 081704 ;

摘要：

Background: Cluster analysis is an integral part of high dimensional data analysis. In the context of large scale gene expression data, a filtered set of genes are grouped together according to their expression profiles using one of numerous clustering algorithms that exist in the statistics and machine learning literature. A closely related problem is that of selecting a clustering algorithm that is "optimal" in some sense from a rather impressive list of clustering algorithms that currently exist. Results: In this paper, we propose two validation measures each with two parts: one measuring the statistical consistency (stability) of the clusters produced and the other representing their biological functional congruence. Smaller values of these indices indicate better performance for a clustering algorithm. We illustrate this approach using two case studies with publicly available gene expression data sets: one involving a SAGE data of breast cancer patients and the other involving a time course cDNA microarray data on yeast. Six well known clustering algorithms UPGMA, K-Means, Diana, Fanny, Model-Based and SOM were evaluated. Conclusion: No single clustering algorithm may be best suited for clustering genes into functional groups via expression profiles for all data sets. The validation measures introduced in this paper can aid in the selection of an optimal algorithm, for a given data set, from a collection of available clustering algorithms.

引用

页数：9

共 17 条

[1] ABBA MC, 2004, BMC BIOINFORMATICS, V6, P5
[2] [Anonymous], 1998, MODERN APPL STAT S P
[3] MODEL-BASED GAUSSIAN AND NON-GAUSSIAN CLUSTERING
BANFIELD, JD
RAFTERY, AE
[J]. BIOMETRICS, 1993, 49 (03) : 803 - 821
[4] The transcriptional program of sporulation in budding yeast
Chu, S
DeRisi, J
Eisen, M
Mulholland, J
Botstein, D
Brown, PO
Herskowitz, I
[J]. SCIENCE, 1998, 282 (5389) : 699 - 705
[5] Comparisons and validation of statistical clustering techniques for microarray gene expression data
Datta, S
Datta, S
[J]. BIOINFORMATICS, 2003, 19 (04) : 459 - 466
[6] DATTA S, 2002, ADV STAT COMBINATORI, P63
[7] Methods for evaluating clustering algorithms for gene expression data using a reference set of functional classes
Datta, Susmita
Datta, Somnath
[J]. BMC BIOINFORMATICS, 2006, 7 (1)
[8] Dudoit S, 2002, GENOME BIOL, V3
[9] Scoring clustering solutions by their biological relevance
Gat-Viks, I
Sharan, R
Shamir, R
[J]. BIOINFORMATICS, 2003, 19 (18) : 2381 - 2389
[10] Hartigan J. A., 1979, Applied Statistics, V28, P100, DOI 10.2307/2346830

← 1 2 →