Use of simulated data sets to evaluate the fidelity of Metagenomicprocessing methods

Mavromatis, Konstantinos; Ivanova, Natalia; Barry, Kerri; Shapiro, Harris; Goltsman, Eugene; McHardy, Alice C; Rigoutsos, Isidore; Salamov, Asaf; Korzeniewski, Frank; Land, Miriam; Lapidus, Alla; Grigoriev, Igor; Richardson, Paul; Hugenholtz, Philip; Kyrpides, Nikos C

Title: Use of simulated data sets to evaluate the fidelity of Metagenomicprocessing methods

Journal Article · Fri Dec 01 00:00:00 EST 2006 · Nature Methods

OSTI ID:920356

Mavromatis, Konstantinos; Ivanova, Natalia; Barry, Kerri; Shapiro, Harris; Goltsman, Eugene; McHardy, Alice C; Rigoutsos, Isidore; Salamov, Asaf; Korzeniewski, Frank; Land, Miriam; Lapidus, Alla; Grigoriev, Igor; Richardson, Paul; Hugenholtz, Philip; Kyrpides, Nikos C

Metagenomics is a rapidly emerging field of research for studying microbial communities. To evaluate methods presently used to process metagenomic sequences, we constructed three simulated data sets of varying complexity by combining sequencing reads randomly selected from 113 isolate genomes. These data sets were designed to model real metagenomes in terms of complexity and phylogenetic composition. We assembled sampled reads using three commonly used genome assemblers (Phrap, Arachne and JAZZ), and predicted genes using two popular gene finding pipelines (fgenesb and CRITICA/GLIMMER). The phylogenetic origins of the assembled contigs were predicted using one sequence similarity--based (blast hit distribution) and two sequence composition--based (PhyloPythia, oligonucleotide frequencies) binning methods. We explored the effects of the simulated community structure and method combinations on the fidelity of each processing step by comparison to the corresponding isolate genomes. The simulated data sets are available online to facilitate standardized benchmarking of tools for metagenomic analysis.

View Journal Article

Cite

Export

Save

Research Organization:: Lawrence Berkeley National Lab. (LBNL), Berkeley, CA (United States)

Sponsoring Organization:: USDOE Director. Office of Science. Biological andEnvironmental Research

DOE Contract Number:: DE-AC02-05CH11231

OSTI ID:: 920356

Report Number(s):: LBNL-62596; R&D Project: 626869; BnR: KP1103010; TRN: US200818%%1157

Journal Information:: Nature Methods, Vol. 4; Related Information: Journal Publication Date: 04/29/2007

Country of Publication:: United States

Language:: English

Similar Records

Use of simulated data sets to evaluate the fidelity of metagenomic processing methods

Journal Article · Mon Jan 01 00:00:00 EST 2007 · Nature Methods · OSTI ID:920356

Mavromatis, K; Ivanova, N; Barry, Kerrie; +11 more

Final Report for LDRD Project 02-ERD-069: Discovering the Unknown Mechanism(s) of Virulence in a BW, Class A Select Agent

Technical Report · Thu Feb 06 00:00:00 EST 2003 · OSTI ID:920356

Chain, P; Garcia, E

Binning sequences using very sparse labels within a metagenome

Journal Article · Mon Apr 28 00:00:00 EDT 2008 · BMC Bioinformatics · OSTI ID:920356

Chan, Chon-Kit; Hsu, Arthur L.; Halgamuge, Saman K.; +1 more

Related Subjects

59 BASIC BIOLOGICAL SCIENCES
COMMUNITIES
CONTIGS
DISTRIBUTION
GENES
OLIGONUCLEOTIDES
PIPELINES
PROCESSING

Title: Use of simulated data sets to evaluate the fidelity of Metagenomicprocessing methods

Citation Formats

Similar Records

Related Subjects