Distributed learning with bagging-like performance

被引:40
作者
Chawla, NV
Moore, TE
Hall, LO
Bowyer, KW
Kegelmeyer, WP
Springer, C
机构
[1] Univ S Florida, Dept Comp Sci & Engn, Tampa, FL 33620 USA
[2] Univ Notre Dame, Dept Comp Engn & Sci, Notre Dame, IN 46556 USA
[3] Sandia Natl Labs, Biosyst Res Dept, Livermore, CA 94551 USA
关键词
distributed learning; bagging; large data sets; ensembles; multiple classifiers;
D O I
10.1016/S0167-8655(02)00269-6
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Bagging forms a committee of classifiers by bootstrap aggregation of training sets from a pool of training data. A simple alternative to bagging is to partition the data into disjoint subsets. Experiments with decision tree and neural network classifiers on various datasets show that, given the same size partitions and bags, disjoint partitions result in performance equivalent to, or better than, bootstrap aggregates (bags). Many applications (e.g., protein structure prediction) involve use of datasets that are too large to handle in the memory of the typical computer. Hence, bagging with samples the size of the data is impractical. Our results indicate that, in such applications, the simple approach of creating a committee of n classifiers from disjoint partitions each of size 1/n (which will be memory resident during learning) in a distributed way results in a classifier which has a bagging-like performance gain. The use of distributed disjoint partitions in learning is significantly less complex and faster than bagging. (C) 2002 Elsevier Science B.V. All rights reserved.
引用
收藏
页码:455 / 471
页数:17
相关论文
共 29 条
[1]  
[Anonymous], 1999, P 5 ACM SIGKDD INT C
[2]  
[Anonymous], MACHINE LEARNING
[3]   The Protein Data Bank [J].
Berman, HM ;
Westbrook, J ;
Feng, Z ;
Gilliland, G ;
Bhat, TN ;
Weissig, H ;
Shindyalov, IN ;
Bourne, PE .
NUCLEIC ACIDS RESEARCH, 2000, 28 (01) :235-242
[4]  
Blake C.L., 1998, UCI repository of machine learning databases
[5]  
Bowyer KW, 2000, IEEE SYS MAN CYBERN, P1888, DOI 10.1109/ICSMC.2000.886388
[6]   Pasting small votes for classification in large databases and on-line [J].
Breiman, L .
MACHINE LEARNING, 1999, 36 (1-2) :85-103
[7]   Bagging predictors [J].
Breiman, L .
MACHINE LEARNING, 1996, 24 (02) :123-140
[8]  
CHAN P, 1996, 9 FLOR ART INT RES S, P151
[9]  
Chan P. K., 1993, P AAAI WORKSH KNOWL, P227
[10]  
CHAN PK, 1995, P 1 INT C KNOWL DISC, P39