The NBP Negative Binomial Model for Assessing Differential Gene Expression from RNA-Seq

被引:115
作者
Di, Yanming
Schafer, Daniel W.
Cumbie, Jason S.
Chang, Jeff H.
机构
基金
美国食品与农业研究所;
关键词
RNA-Seq; overdispersion; negative binomial; Poisson;
D O I
10.2202/1544-6115.1637
中图分类号
Q5 [生物化学]; Q7 [分子生物学];
学科分类号
071010 ; 081704 ;
摘要
We propose a new statistical test for assessing differential gene expression using RNA sequencing (RNA-Seq) data. Commonly used probability distributions, such as binomial or Poisson, cannot appropriately model the count variability in RNA-Seq data due to overdispersion. The small sample size that is typical in this type of data also prevents the uncritical use of tools derived from large-sample asymptotic theory. The test we propose is based on the NBP parameterization of the negative binomial distribution. It extends an exact test proposed by Robinson and Smyth (2007, 2008). In one version of Robinson and Smyth's test, a constant dispersion parameter is used to model the count variability between biological replicates. We introduce an additional parameter to allow the dispersion parameter to depend on the mean. Our parametric method complements nonparametric regression approaches for modeling the dispersion parameter. We apply the test we propose to an Arabidopsis data set and a range of simulated data sets. The results show that the test is simple, powerful and reasonably robust against departures from model assumptions.
引用
收藏
页数:29
相关论文
共 21 条
[11]   GENERALIZED LINEAR MODELS [J].
NELDER, JA ;
WEDDERBURN, RW .
JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES A-GENERAL, 1972, 135 (03) :370-+
[12]  
Pournelle G. H., 1953, Journal of Mammalogy, V34, P133, DOI 10.1890/0012-9658(2002)083[1421:SDEOLC]2.0.CO
[13]  
2
[14]   The roles of conditioning in inference [J].
Reid, N .
STATISTICAL SCIENCE, 1995, 10 (02) :138-157
[15]   Small-sample estimation of negative binomial dispersion, with applications to SAGE data [J].
Robinson, Mark D. ;
Smyth, Gordon K. .
BIOSTATISTICS, 2008, 9 (02) :321-332
[16]   Moderated statistical tests for assessing differences in tag abundance [J].
Robinson, Mark D. ;
Smyth, Gordon K. .
BIOINFORMATICS, 2007, 23 (21) :2881-2887
[17]   A scaling normalization method for differential expression analysis of RNA-seq data [J].
Robinson, Mark D. ;
Oshlack, Alicia .
GENOME BIOLOGY, 2010, 11 (03)
[18]   edgeR: a Bioconductor package for differential expression analysis of digital gene expression data [J].
Robinson, Mark D. ;
McCarthy, Davis J. ;
Smyth, Gordon K. .
BIOINFORMATICS, 2010, 26 (01) :139-140
[19]  
Smyth GK., 2004, Statistical Applications in Genetics and Molecular Biology, P3, DOI [10.2202/1544-6115.1027, DOI 10.2202/1544-6115.1027]
[20]   Statistical significance for genomewide studies [J].
Storey, JD ;
Tibshirani, R .
PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, 2003, 100 (16) :9440-9445