Problems in gene clustering based on gene expression data

被引:36
作者
Bryan, J
机构
[1] Univ British Columbia, Dept Stat, Vancouver, BC V6M 1L2, Canada
[2] Univ British Columbia, Biotechnol Lab, Vancouver, BC V6M 1L2, Canada
基金
加拿大自然科学与工程研究理事会;
关键词
cluster analysis; microarrays; confidence; bootstrap;
D O I
10.1016/j.jmva.2004.02.011
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
In this work, we assess the suitability of cluster analysis for the gene grouping problem confronted with microarray data. Gene clustering is the exercise of grouping genes based on attributes, which are generally the expression levels over a number of conditions or subpopulations. The hope is that similarity with respect to expression is often indicative of similarity with respect to much more fundamental and elusive qualities, such as function. By formally defining the true gene-specific attributes as parameters, such as expected expression across the conditions, we obtain a well-defined gene clustering parameter of interest, which greatly facilitates the statistical treatment of gene clustering. We point out that genome-wide collections of expression trajectories often lack natural clustering structure, prior to ad hoc gene filtering. The gene filters in common use induce a certain circularity to most gene cluster analyses: genes are points in the attribute space, a filter is applied to depopulate certain areas of the space, and then clusters are sought (and often found!) in the "cleaned" attribute space. As a result, statistical investigations of cluster number and clustering strength are just as much a study of the stringency and nature of the filter as they are of any biological gene clusters. In the absence of natural clusters, gene clustering may still be a worthwhile exercise in data segmentation. In this context, partitions can be fruitfully encoded in adjacency matrices and the sampling distribution of such matrices can be studied with a variety of bootstrapping techniques. (C) 2003 Elsevier Inc. All rights reserved.
引用
收藏
页码:44 / 66
页数:23
相关论文
共 26 条