Interactively optimizing signal-to-noise ratios in expression profiling: project-specific algorithm selection and detection p-value weighting in Affymetrix microarrays

被引:97
作者
Seo, J
Bakay, M
Chen, YW
Hilmer, S
Shneiderman, B
Hoffman, EP [1 ]
机构
[1] Univ Maryland, Childrens Natl Med Ctr, Res Ctr Genet Med, College Pk, MD 20742 USA
[2] Univ Maryland, Human Comp Interact Lab, College Pk, MD 20742 USA
[3] Univ Maryland, Dept Comp Sci, College Pk, MD 20742 USA
关键词
D O I
10.1093/bioinformatics/bth280
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Motivation: The most commonly utilized microarrays for mRNA profiling (Affymetrix) include 'probe sets' of a series of perfect match and mismatch probes (typically 22 oligonucleotides per probe set). There are an increasing number of reported 'probe set algorithms' that differ in their interpretation of a probe set to derive a single normalized 'signal' representative of expression of each mRNA. These algorithms are known to differ in accuracy and sensitivity, and optimization has been done using a small set of standardized control microarray data. We hypothesized that different mRNA profiling projects have varying sources and degrees of confounding noise, and that these should alter the choice of a specific probe set algorithm. Also, we hypothesized that use of the Microarray Suite (MAS) 5.0 probe set detection p-value as a weighting function would improve the performance of all probe set algorithms. Results: We built an interactive visual analysis software tool (HCE2W) to test and define parameters in Affymetrix analyses that optimize the ratio of signal (desired biological variable) versus noise (confounding uncontrolled variables). Five probe set algorithms were studied with and without statistical weighting of probe sets using the MAS 5.0 probe set detection p-values. The signal-to-noise ratio optimization method was tested in two large novel microarray datasets with different levels of confounding noise, a 105 sample U133A human muscle biopsy dataset (11 groups: mutation-defined, extensive noise), and a 40 sample U74A inbred mouse lung dataset (8 groups: little noise). Performance was measured by the ability of the specific probe set algorithm, with and without detection p-value weighting, to cluster samples into the appropriate biological groups (unsupervised agglomerative clustering with F-measure values). Of the total random sampling analyses, 50% showed a highly statistically significant difference between probe set algorithms by ANOVA [F(4,10) > 14, p < 0.0001], with weighting by MAS 5.0 detection p-value showing significance in the mouse data by ANOVA [F(1,10) > 9, p < 0.013] and paired t-test [t(9) = -3.675, p = 0.005]. Probe set detection p-value weighting had the greatest positive effect on performance of dChip difference model, ProbeProfiler and RMA algorithms. Importantly, probe set algorithms did indeed perform differently depending on the specific project, most probably due to the degree of confounding noise. Our data indicate that significantly improved data analysis of mRNA profile projects can be achieved by optimizing the choice of probe set algorithm with the noise levels intrinsic to a project, with dChip difference model with MAS 5.0 detection p-value continuous weighting showing the best overall performance in both projects. Furthermore, both existing and newly developed probe set algorithms should incorporate a detection p-value weighting to improve performance.
引用
收藏
页码:2534 / 2544
页数:11
相关论文
共 40 条
  • [1] *AFF, 2001, DAT AN FUND
  • [2] *AFF, 2001, MICR SUIT US GUID VE
  • [3] *AFF, 2001, STAT ALG REF GUID
  • [4] [Anonymous], 1994, SIGIR
  • [5] A web-accessible complete transcriptome of normal human and DMD muscle
    Bakay, M
    Zhao, P
    Chen, J
    Hoffman, EP
    [J]. NEUROMUSCULAR DISORDERS, 2002, 12 : S125 - S141
  • [6] Sources of variability and effect of experimental approach on expression profiling data interpretation
    Bakay, M
    Chen, YW
    Borup, R
    Zhao, P
    Nagaraju, K
    Hoffman, EP
    [J]. BMC BIOINFORMATICS, 2002, 3 (1)
  • [7] Quantitative analysis of mRNA amplification by in vitro transcription
    Baugh, L. R.
    Hill, A. A.
    Brown, E. L.
    Hunter, Craig P.
    [J]. NUCLEIC ACIDS RESEARCH, 2001, 29 (05)
  • [8] BJORNAR L, 1999, P 5 ACM SIGKDD INT C, P16, DOI DOI 10.1145/312129.312186
  • [9] A comparison of normalization methods for high density oligonucleotide array data based on variance and bias
    Bolstad, BM
    Irizarry, RA
    Åstrand, M
    Speed, TP
    [J]. BIOINFORMATICS, 2003, 19 (02) : 185 - 193
  • [10] The PEPR GeneChip data warehouse, and implementation of a dynamic time series query tool (SGQT) with graphical interface
    Chen, J
    Zhao, P
    Massaro, D
    Clerch, LB
    Almon, RR
    DuBois, DC
    Jusko, WJ
    Hoffman, EP
    [J]. NUCLEIC ACIDS RESEARCH, 2004, 32 : D578 - D581