Sample size planning for developing classifiers using high-dimensional DNA microarray data

被引:85
作者
Dobbin, Kevin K. [1 ]
Simon, Richard M. [1 ]
机构
[1] NCI, Biometr Res Branch, Rockville, MD 20852 USA
关键词
gene expression; microarrays; prediction; predictive inference; sample size;
D O I
10.1093/biostatistics/kxj036
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
Many gene expression studies attempt to develop a predictor of pre-defined diagnostic or prognostic classes. If the classes are similar biologically, then the number of genes that are differentially expressed between the classes is likely to be small compared to the total number of genes measured. This motivates a two-step process for predictor development, a subset of differentially expressed genes is selected for use in the predictor and then the predictor constructed from these. Both these steps will introduce variability into the resulting classifier, so both must be incorporated in sample size estimation. We introduce a methodology for sample size determination for prediction in the context of high-dimensional data that captures variability in both steps of predictor development. The methodology is based on a parametric probability model, but permits sample size computations to be carried out in a practical manner without extensive requirements for preliminary data. We find that many prediction problems do not require a large training set of arrays for classifier development.
引用
收藏
页码:101 / 117
页数:17
相关论文
共 16 条
[1]  
Carlin B. P., 2001, BAYES EMPIRICAL BAYE
[2]   Sample size determination in microarray experiments for class comparison and prognostic classification [J].
Dobbin, K ;
Simon, R .
BIOSTATISTICS, 2005, 6 (01) :27-38
[3]   OPTIMAL PREDICTIVE LINEAR DISCRIMINANTS [J].
ENIS, P ;
GEISSER, S .
ANNALS OF STATISTICS, 1974, 2 (02) :403-410
[4]   How many samples are needed to build a classifier: a general sequential approach [J].
Fu, WJJ ;
Dougherty, ER ;
Mallick, B ;
Carroll, RJ .
BIOINFORMATICS, 2005, 21 (01) :63-70
[5]   Molecular classification of cancer: Class discovery and class prediction by gene expression monitoring [J].
Golub, TR ;
Slonim, DK ;
Tamayo, P ;
Huard, C ;
Gaasenbeek, M ;
Mesirov, JP ;
Coller, H ;
Loh, ML ;
Downing, JR ;
Caligiuri, MA ;
Bloomfield, CD ;
Lander, ES .
SCIENCE, 1999, 286 (5439) :531-537
[6]  
Hocking R.R., 1996, METHODS APPL LINEAR
[7]   Determination of minimum sample size and discriminatory expression patterns in microarray data [J].
Hwang, DH ;
Schmitt, WA ;
Stephanopoulos, G ;
Stephanopoulos, G .
BIOINFORMATICS, 2002, 18 (09) :1184-1193
[9]   A well-conditioned estimator for large-dimensional covariance matrices [J].
Ledoit, O ;
Wolf, M .
JOURNAL OF MULTIVARIATE ANALYSIS, 2004, 88 (02) :365-411
[10]   Prediction error estimation: a comparison of resampling methods [J].
Molinaro, AM ;
Simon, R ;
Pfeiffer, RM .
BIOINFORMATICS, 2005, 21 (15) :3301-3307