Sample selection bias and presence-only distribution models: implications for background and pseudo-absence data

被引:2219
作者
Phillips, Steven J. [1 ]
Dudik, Miroslav [2 ]
Elith, Jane [3 ]
Graham, Catherine H. [4 ]
Lehmann, Anthony [5 ]
Leathwick, John [6 ]
Ferrier, Simon [7 ]
机构
[1] AT&T Labs Res, Florham Pk, NJ 07932 USA
[2] Princeton Univ, Dept Comp Sci, Princeton, NJ 08544 USA
[3] Univ Melbourne, Sch Bot, Parkville, Vic 3010, Australia
[4] SUNY Stony Brook, Dept Ecol & Evolut, Stony Brook, NY 11794 USA
[5] Univ Geneva, CH-1227 Carogue, Switzerland
[6] NIWA, Hamilton, New Zealand
[7] New S Wales Dept Environm & Climate Change, Armidale, NSW 2350, Australia
关键词
background data; presence-only distribution models; niche modeling; pseudo-absence; sample selection bias; species distribution modeling; target group; SPECIES DISTRIBUTION MODELS; NEW-ZEALAND; REGRESSION; MUSEUM; BIODIVERSITY; PREDICTION; TREES; RISK;
D O I
10.1890/07-2153.1
中图分类号
Q14 [生态学(生物生态学)];
学科分类号
071012 ; 0713 ;
摘要
Most methods for modeling species distributions from occurrence records require additional data representing the range of environmental conditions in the modeled region. These data, called background or pseudo-absence data, are usually drawn at random from the entire region, whereas occurrence collection is often spatially biased toward easily accessed areas. Since the spatial bias generally results in environmental bias, the difference between occurrence collection and background sampling may lead to inaccurate models. To correct the estimation, we propose choosing background data with the same bias as occurrence data. We investigate theoretical and practical implications of this approach. Accurate information about spatial bias is usually lacking, so explicit biased sampling of background sites may not be possible. However, it is likely that an entire target group of species observed by similar methods will share similar bias. We therefore explore the use of all occurrences within a target group as biased background data. We compare model performance using target-group background and randomly sampled background on a comprehensive collection of data for 226 species from diverse regions of the world. We find that target-group background improves average performance for all the modeling methods we consider, with the choice of background data having as large an effect on predictive performance as the choice of modeling method. The performance improvement due to target-group background is greatest when there is strong bias in the target-group presence records. Our approach applies to regression-based modeling methods that have been adapted for use with occurrence data, such as generalized linear or additive models and boosted regression trees, and to Maxent, a probability density estimation method. We argue that increased awareness of the implications of spatial bias in surveys, and possible modeling remedies, will substantially improve predictions of species distributions.
引用
收藏
页码:181 / 197
页数:17
相关论文
共 57 条
[1]   Real vs. artefactual absences in species distributions:: tests for Oryzomys albigularis (Rodentia: Muridae) in Venezuela [J].
Anderson, RP .
JOURNAL OF BIOGEOGRAPHY, 2003, 30 (04) :591-605
[2]   Prediction of potential areas of species distributions based on presence-only data [J].
Argáez, JA ;
Christen, JA ;
Nakamura, M ;
Soberón, J .
ENVIRONMENTAL AND ECOLOGICAL STATISTICS, 2005, 12 (01) :27-44
[3]   Evaluating resource selection functions [J].
Boyce, MS ;
Vernier, PR ;
Nielsen, SE ;
Schmiegelow, FKA .
ECOLOGICAL MODELLING, 2002, 157 (2-3) :281-300
[4]  
Busby J. R., 1991, Plant Protection Quarterly, V6, P8
[5]  
Cadman M.D., 2008, The atlas of the breeding birds of Ontario, 2001-2005
[6]   DOMAIN - A FLEXIBLE MODELING PROCEDURE FOR MAPPING POTENTIAL DISTRIBUTIONS OF PLANTS AND ANIMALS [J].
CARPENTER, G ;
GILLISON, AN ;
WINTER, J .
BIODIVERSITY AND CONSERVATION, 1993, 2 (06) :667-680
[7]   Regional vegetation mapping in Australia: a case study in the practical use of statistical modelling [J].
Cawsey, EM ;
Austin, MP ;
Baker, BL .
BIODIVERSITY AND CONSERVATION, 2002, 11 (12) :2239-2274
[8]  
De'ath G, 2007, ECOLOGY, V88, P243, DOI 10.1890/0012-9658(2007)88[243:BTFEMA]2.0.CO
[9]  
2
[10]   Bias in butterfly distribution maps: the influence of hot spots and recorder's home range [J].
Dennis, R. L. H. ;
Thomas, C. D. .
JOURNAL OF INSECT CONSERVATION, 2000, 4 (02) :73-77