Classification using generalized partial least squares

被引:50
作者
Ding, BY
Gentleman, R
机构
[1] Amgen Inc, Med Affairs Biostat, Newbury Pk, CA 91320 USA
[2] Fred Hutchinson Canc Res Ctr, Program Computat Biol, Div Publ Hlth Sci, Seattle, WA 98104 USA
关键词
cross-validation; Firth's procedure; gene expression; iteratively reweighted partial least squares; (quasi) separation; two-stage PLS;
D O I
10.1198/106186005X47697
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
Advances in computational biology have made simultaneous monitoring of thousands of features possible. The high throughput technologies not only bring about a much richer information context in which to study various aspects of gene function, but they also present the challenge of analyzing data with a large number of covariates and few samples. As an integral part of machine learning, classification of samples into two or more categories is almost always of interest to scientists. We address the question of classification in this setting by extending partial least squares (PLS), a popular dimension reduction tool in chemometrics, in the context of generalized linear regression, based on a previous approach, iteratively reweighted partial least squares, that is, IRWPLS. We compare our results with two-stage PLS and with other classifiers. We show that by phrasing the problem in a generalized linear model setting and by applying Firth's procedure to avoid (quasi)separation, we often get lower classification error rates.
引用
收藏
页码:280 / 298
页数:19
相关论文
共 37 条
[1]  
ALBERT A, 1984, BIOMETRIKA, V71, P1
[2]   Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays [J].
Alon, U ;
Barkai, N ;
Notterman, DA ;
Gish, K ;
Ybarra, S ;
Mack, D ;
Levine, AJ .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 1999, 96 (12) :6745-6750
[3]   Selection bias in gene extraction on the basis of microarray gene-expression data [J].
Ambroise, C ;
McLachlan, GJ .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 2002, 99 (10) :6562-6566
[4]   PLS regression methods [J].
Höskuldsson, Agnar .
Journal of Chemometrics, 1988, 2 (03) :211-228
[5]   Random forests [J].
Breiman, L .
MACHINE LEARNING, 2001, 45 (01) :5-32
[6]  
Breiman L., 2002, MANUAL SETTING USING, V1
[7]  
Chen Y, 1997, J Biomed Opt, V2, P364, DOI 10.1117/12.281504
[8]   Comparison of discrimination methods for the classification of tumors using gene expression data [J].
Dudoit, S ;
Fridlyand, J ;
Speed, TP .
JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 2002, 97 (457) :77-87
[9]   Classification of microarray data with penalized logistic regression [J].
Eilers, PHC ;
Boer, JM ;
van Ommen, GJ ;
van Houwelingen, HC .
MICROARRAYS: OPTICAL TECHNOLOGIES AND INFORMATICS, 2001, 4266 :187-198
[10]  
Fahrmeir L., 2001, MULTIVARIATE STAT MO