Close
Help
Need Help?



SOLiDzipper: A High Speed Encoding Method for the Next-Generation Sequencing Data

Submit a Paper



Publication Date: 10 Mar 2011

Type: Software or database review

Journal: Evolutionary Bioinformatics

Citation: Evolutionary Bioinformatics 2011:7 1-6

doi: 10.4137/EBO.S6618

Abstract

Background: Next-generation sequencing (NGS) methods pose computational challenges of handling large volumes of data. Although cloud computing offers a potential solution to these challenges, transferring a large data set across the internet is the biggest obstacle, which may be overcome by efficient encoding methods. When encoding is used to facilitate data transfer to the cloud, the time factor is equally as important as the encoding efficiency. Moreover, to take advantage of parallel processing in cloud computing, a parallel technique to decode and split compressed data in the cloud is essential. Hence in this review, we present SOLiDzipper, a new encoding method for NGS data.

Methods: The basic strategy of SOLiDzipper is to divide and encode. NGS data files contain both the sequence and non-sequence information whose encoding efficiencies are different. In SOLiDzipper, encoded data are stored in binary data block that does not contain the characteristic information of a specific sequence platform, which means that data can be decoded according to a desired platform even in cases of Illumina, Solexa or Roche 454 data.

Results: The main calculation time using Crossbow was 173 minutes when 40 EC2 nodes were involved. In that case, an analysis preparation time of 464 minutes is required to encode data in the latest DNA compression method like G-SQZ and transmit it on a 183 Mbit/s bandwidth. However, it takes 194 minutes to encode and transmit data with SOLiDzipper under the same bandwidth conditions. These results indicate that the entire processing time can be reduced according to the encoding methods used, under the same network bandwidth conditions. Considering the limited network bandwidth, high-speed, high-efficiency encoding methods such as SOLiDzipper can make a significant contribution to higher productivity in labs seeking to take advantage of the cloud as an alternative to local computing.

Availability: http://szipper.dinfree.com. Academic/non-profit: Binary available for direct download at no cost. For-profit: Submit request for for-profit license from the web-site.


Downloads

PDF  (595.67 KB PDF FORMAT)

RIS citation   (ENDNOTE, REFERENCE MANAGER, PROCITE, REFWORKS)

BibTex citation   (BIBDESK, LATEX)

XML

PMC HTML


Sharing

Our Service Promise

  • Prompt Processing (Less Than 3 Weeks)
  • Fair & Comprehensive Peer Review
  • Professional Author Service
  • Leading Editors in Chief
  • Extensive Indexing
  • High Readership & Impact
  • What Your Colleagues Say

Quick Links

Follow Us We make it easy to find new research papers.
Email AlertsRSS Feeds
FacebookGoogle+Twitter
PinterestTumblrYouTube

SUBJECT HUBS
Our Testimonials
Publishing in Air, Soil and Water and Water Research was the best experience I have had so far in an academic context.  The review process was fair, quick and efficient.  I congratulate the team at Libertas Academica for a very well managed journal.
Magnus Karlsson (IVL Swedish Environmental Research Institute, Stockholm, Sweden) What Your Colleagues Say