Human genomes as email attachments

被引:99
作者
Christley, Scott [1 ]
Lu, Yiming [1 ]
Li, Chen [1 ]
Xie, Xiaohui [1 ,2 ]
机构
[1] Univ Calif Irvine, Dept Comp Sci, Irvine, CA 92697 USA
[2] Univ Calif Irvine, Inst Genom & Bioinformat, Irvine, CA 92697 USA
基金
美国国家科学基金会;
关键词
D O I
10.1093/bioinformatics/btn582
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
The amount of genomic sequence data being generated and made available through public databases continues to increase at an ever-expanding rate. Downloading, copying, sharing and manipulating these large datasets are becoming difficult and time consuming for researchers. We need to consider using advanced compression techniques as part of a standard data format for genomic data. The inherent structure of genome data allows for more efficient lossless compression than can be obtained through the use of generic compression programs. We apply a series of techniques to James Watson's genome that in combination reduce it to a mere 4MB, small enough to be sent as an email attachment.
引用
收藏
页码:274 / 275
页数:2
相关论文
共 5 条
  • [1] DNACompress: fast and effective DNA sequence compression
    Chen, X
    Li, M
    Ma, B
    Tromp, J
    [J]. BIOINFORMATICS, 2002, 18 (12) : 1696 - 1698
  • [2] The International HapMap Project
    Gibbs, RA
    Belmont, JW
    Hardenbol, P
    Willis, TD
    Yu, FL
    Yang, HM
    Ch'ang, LY
    Huang, W
    Liu, B
    Shen, Y
    Tam, PKH
    Tsui, LC
    Waye, MMY
    Wong, JTF
    Zeng, CQ
    Zhang, QR
    Chee, MS
    Galver, LM
    Kruglyak, S
    Murray, SS
    Oliphant, AR
    Montpetit, A
    Hudson, TJ
    Chagnon, F
    Ferretti, V
    Leboeuf, M
    Phillips, MS
    Verner, A
    Kwok, PY
    Duan, SH
    Lind, DL
    Miller, RD
    Rice, JP
    Saccone, NL
    Taillon-Miller, P
    Xiao, M
    Nakamura, Y
    Sekine, A
    Sorimachi, K
    Tanaka, T
    Tanaka, Y
    Tsunoda, T
    Yoshino, E
    Bentley, DR
    Deloukas, P
    Hunt, S
    Powell, D
    Altshuler, D
    Gabriel, SB
    Qiu, RZ
    [J]. NATURE, 2003, 426 (6968) : 789 - 796
  • [3] A METHOD FOR THE CONSTRUCTION OF MINIMUM-REDUNDANCY CODES
    HUFFMAN, DA
    [J]. PROCEEDINGS OF THE INSTITUTE OF RADIO ENGINEERS, 1952, 40 (09): : 1098 - 1101
  • [4] The complete genome of an individual by massively parallel DNA sequencing
    Wheeler, David A.
    Srinivasan, Maithreyan
    Egholm, Michael
    Shen, Yufeng
    Chen, Lei
    McGuire, Amy
    He, Wen
    Chen, Yi-Ju
    Makhijani, Vinod
    Roth, G. Thomas
    Gomes, Xavier
    Tartaro, Karrie
    Niazi, Faheem
    Turcotte, Cynthia L.
    Irzyk, Gerard P.
    Lupski, James R.
    Chinault, Craig
    Song, Xing-zhi
    Liu, Yue
    Yuan, Ye
    Nazareth, Lynne
    Qin, Xiang
    Muzny, Donna M.
    Margulies, Marcel
    Weinstock, George M.
    Gibbs, Richard A.
    Rothberg, Jonathan M.
    [J]. NATURE, 2008, 452 (7189) : 872 - U5
  • [5] Compressing DNA sequence databases with coil
    White, W. Timothy J.
    Hendy, Michael D.
    [J]. BMC BIOINFORMATICS, 2008, 9 (1)