Simple statistical models predict C-to-U edited sites in plant mitochondrial RNA

被引:36
作者
Cummings, MP [1 ]
Myers, DS [1 ]
机构
[1] Univ Maryland, Ctr Bioinformat & Computat Biol, College Pk, MD 20742 USA
关键词
D O I
10.1186/1471-2105-5-132
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Background: RNA editing is the process whereby an RNA sequence is modified from the sequence of the corresponding DNA template. In the mitochondria of land plants, some cytidines are converted to uridines before translation. Despite substantial study, the molecular biological mechanism by which C-to-U RNA editing proceeds remains relatively obscure, although several experimental studies have implicated a role for cis-recognition. A highly non-random distribution of nucleotides is observed in the immediate vicinity of edited sites (within 20 nucleotides 5' and 3'), but no precise consensus motif has been identified. Results: Data for analysis were derived from the the complete mitochondrial genomes of Arabidopsis thaliana, Brassica napus, and Oryza sativa; additionally, a combined data set of observations across all three genomes was generated. We selected datasets based on the 20 nucleotides 5' and the 20 nucleotides 3' of edited sites and an equivalently sized and appropriately constructed null-set of non-edited sites. We used tree-based statistical methods and random forests to generate models of C-to-U RNA editing based on the nucleotides surrounding the edited/non-edited sites and on the estimated folding energies of those regions. Tree-based statistical methods based on primary sequence data surrounding edited/non-edited sites and estimates of free energy of folding yield models with optimistic re-substitution-based estimates of similar to0.71 accuracy, similar to0.64 sensitivity, and similar to0.88 specificity. Random forest analysis yielded better models and more exact performance estimates with similar to0.74 accuracy, similar to0.72 sensitivity, and similar to0.81 specificity for the combined observations. Conclusions: Simple models do moderately well in predicting which cytidines will be edited to uridines, and provide the first quantitative predictive models for RNA edited sites in plant mitochondria. Our analysis shows that the identity of the nucleotide -1 to the edited C and the estimated free energy of folding for a 41 nt region surrounding the edited C are the most important variables that distinguish most edited from non-edited sites. However, the results suggest that primary sequence data and simple free energy of folding calculations alone are insufficient to make highly accurate predictions.
引用
收藏
页数:7
相关论文
共 37 条
[1]   RNA EDITING IN WHEAT MITOCHONDRIA [J].
ARAYA, A ;
BLANC, V ;
BEGU, D ;
CRABIER, F ;
MOURAS, A ;
LITVAK, S .
BIOCHIMIE, 1995, 77 (1-2) :87-91
[2]   GenBank: update [J].
Benson, DA ;
Karsch-Mizrachi, I ;
Lipman, DJ ;
Ostell, J ;
Wheeler, DL .
NUCLEIC ACIDS RESEARCH, 2004, 32 :D23-D26
[3]   RNA EDITING IN WHEAT MITOCHONDRIA PROCEEDS BY A DEAMINATION MECHANISM [J].
BLANC, V ;
LITVAK, S ;
ARAYA, A .
FEBS LETTERS, 1995, 373 (01) :56-60
[4]   SmcHD1, containing a structural-maintenance-of-chromosomes hinge domain, has a critical role in X inactivation [J].
Blewitt, Marnie E. ;
Gendrel, Anne-Valerie ;
Pang, Zhenyi ;
Sparrow, Duncan B. ;
Whitelaw, Nadia ;
Craig, Jeffrey M. ;
Apedaile, Anwyn ;
Hilton, Douglas J. ;
Dunwoodie, Sally L. ;
Brockdorff, Neil ;
Kay, Graham F. ;
Whitelaw, Emma .
NATURE GENETICS, 2008, 40 (05) :663-669
[5]   Random forests [J].
Breiman, L .
MACHINE LEARNING, 2001, 45 (01) :5-32
[6]   Random forests [J].
Breiman, L .
MACHINE LEARNING, 2001, 45 (01) :5-32
[7]  
BREIMAN L, 2001, 567 U CAL DEP STAT
[8]   RNA editing status of nad7 intron domains in wheat mitochondria [J].
Carrillo, C ;
Bonen, L .
NUCLEIC ACIDS RESEARCH, 1997, 25 (02) :403-409
[9]   Cross-competition in transgenic chloroplasts expressing single editing sites reveals shared cis elements [J].
Chateigner-Boutin, AL ;
Hanson, MR .
MOLECULAR AND CELLULAR BIOLOGY, 2002, 22 (24) :8448-8456
[10]  
Clark L. A., 1993, STAT MODELS S