Navigation Links
Researchers develop tool to evaluate genome sequencing method
Date:1/2/2013

Advances in bio-technologies and computer software have helped make genome sequencing much more common than in the past. But still in question are both the accuracy of different sequencing methods and the best ways to evaluate these efforts. Now, computer scientists have devised a tool to better measure the validity of genome sequencing.

The method, which is described in the journal PLOS ONE, allows for the evaluation of a wide range of genome sequencing procedures by tracking a small group of key statistical features in the basic structure of the assembled genome. Such sequence-assembly algorithm lays out the individual short reads (strings of DNA's four nucleic acid bases sampled from the target genome) to put together the complete genome sequencemuch like a complex jig-saw puzzle. The method uses techniques from statistical inference and learning theory to select the most significant features. Surprisingly, the method concludes that many features thought by human experts to be the most important were actually highly misleading.

The work was conducted by researchers at New York University's Courant Institute of Mathematical Sciences, NYU School of Medicine, Sweden's KTH Royal Institute of Technology, and Cold Spring Harbor Laboratory.

Current evaluation methods of genome sequencing are typically imprecise. They rely on what amounts to "crowd sourcing," with scientists weighing in on the accuracy of a sequencing method. Other evaluations use apples-to-oranges comparisons in making assessments, thus limiting their value.

In the PLOS ONE work, the researchers expanded upon an earlier system they created, Feature Response Curve (FRCurve), which offers a global picture of how genome-sequencing methods, or assemblers, are able to deal with different regions and different structures in a large complex genome. Specifically, it points out how an assembler might have traded off one kind of quality measure at the expense of another kind. For instance, it shows how aggressively a genome assembler might have tried to pull together a group of genes into a contiguous piece of the genome, while incorrectly rearranging their correct order and copy numbers.

However, FRCurve has a significant limitationit can only gauge the accuracy of certain kinds of assemblers at one time, thereby excluding comparisons among the range of sequencing methods currently being employed. Many of these methods, where the original FRCurve failed, are becoming highly popular, as they are specifically designed to work with the most established next-generation sequencing technologies and are able to perform some error correction and data compression. However, by doing so, they also discard the original signature of key statistical features (e.g., position and orientation of the reads used to generate the candidate sequence) that FRCurve needs for evaluation.

The work reported in PLOS ONE unveils a new method, FRCbam, which has the capability to evaluate a much wider class of assemblers. It does so by reverse engineering the latent structures that were obscured by error-correction and data compression; and it performs this operation rapidly by using efficient and scalable mapping algorithms.

Instead of assumption-ridden simulation or expensive auxiliary methods, FRCbam validates its analysis by examining a large ensemble of assemblers working on a large ensemble of genomes, selected from crowd-sourced competitions like GAGE and Assemblathons. This way, FRCbam can characterize the statistics that are expected and then validate any individual system with respect to it.

FRCbam and FRCurve are expected to be used routinely to rank and evaluate future genome projects. This method is currently employed to evaluate the sequence assembly of the Norway Spruce, one of the largest genomes sequenced so farit is seven times longer than the human genome.


'/>"/>

Contact: James Devitt
james.devitt@nyu.edu
212-998-6808
New York University
Source:Eurekalert

Related biology news :

1. Study by UC Santa Barbara researchers suggests that bacteria communicate by touch
2. UC Santa Barbara researchers discover genetic link between visual pathways of hydras and humans
3. Researchers attempt to solve problems of antibiotic resistance and bee deaths in one
4. UNH researchers find African farmers need better climate change data to improve farming practices
5. Ottawa researchers to lead world-first clinical trial of stem cell therapy for septic shock
6. Researchers uncover molecular pathway through which common yeast becomes fungal pathogen
7. Researchers print live cells with a standard inkjet printer
8. Columbia Engineering and Penn researchers increase speed of single-molecule measurements
9. Researchers reveal how a single gene mutation leads to uncontrolled obesity
10. Researchers discover novel therapy for Crohns disease
11. New paper by Notre Dame researchers describes method for cleaning up nuclear waste
Post Your Comments:
*Name:
*Comment:
*Email:
(Date:5/3/2016)... , May 3, 2016  Neurotechnology, a provider ... MegaMatcher Automated Biometric Identification System (ABIS) , ... multi-biometric projects. MegaMatcher ABIS can process multiple complex ... any combination of fingerprint, face or iris biometrics. ... SDK and MegaMatcher Accelerator , which ...
(Date:4/19/2016)... 2016 The new GEZE SecuLogic ... web-based "all-in-one" system solution for all door components. It ... the door interface with integration authorization management system, and ... The minimal dimensions of the access control and the ... installations offer considerable freedom of design with regard to ...
(Date:3/31/2016)... , March 31, 2016   ... ("LegacyXChange" or the "Company") LegacyXChange is excited ... of its soon to be launched online site for ... https://www.youtube.com/channel/UCyTLBzmZogV1y2D6bDkBX5g ) will also provide potential shareholders a ... DNA technology to an industry that is notorious for ...
Breaking Biology News(10 mins):
(Date:5/25/2016)... ... 2016 , ... The Ankle Plating System 3 and Small ... fractures of the distal tibia and fibula. This system marks Acumed's continued commitment ... is composed of seven plate families that span the lateral, medial, and posterior ...
(Date:5/25/2016)... ... May 25, 2016 , ... The American Medical Informatics Association ... the National Coordinator for Health IT (ONC) outlining a measurement approach to interoperability ... available when and where it was needed. The organization of health informatics professionals ...
(Date:5/25/2016)... ... May 25, 2016 , ... WEDI, the nation’s leading authority on the use ... W. Stellar has been named by the WEDI Board of Directors as WEDI’s president ... executive leader with more than 35 years of experience in healthcare, association management and ...
(Date:5/25/2016)... ... 25, 2016 , ... Biohaven Pharmaceutical Holding Company Ltd. (Biohaven) ... company’s orphan drug designation request covering BHV-4157 for the treatment of Spinocerebellar Ataxia ... , Spinocerebellar ataxia is a rare, debilitating neurodegenerative disorder that is estimated ...
Breaking Biology Technology: