skip to main content
OSTI.GOV title logo U.S. Department of Energy
Office of Scientific and Technical Information

Title: The Scientific Data Management Center: Available Technologies and Highlights

Abstract

Managing scientific data has been identified by the scientific community as one of the most important emerging needs because of the sheer volume and increasing complexity of data being collected. Effectively generating, managing, and analyzing this information requires a comprehensive, end-to-end approach to data management that encompasses all of the stages from the initial data acquisition to the final analysis of the data. Based on community input, we have identified three significant requirements. First, more efficient access to storage systems is needed. In particular, parallel file system and I/O system improvements are needed to write and read large volumes of data without slowing a simulation. Second, scientists require technologies to facilitate better understanding of their data, in particular the ability to effectively perform complex data analysis and searches over extremely large data sets. Furthermore, exploratory analysis requires techniques for efficiently selecting subsets of the data. Third, generating the data, collecting and storing the results, keeping track of data provenance, data post-processing, and analysis of results is a tedious, fragmented process. Tools for automation of this process in a robust, tractable, and recoverable fashion are required to enhance scientific exploration.

Authors:
; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; more »; ; ; ; ; ; ; ; ; ; ; ; ; ; « less
Publication Date:
Research Org.:
Pacific Northwest National Lab. (PNNL), Richland, WA (United States)
Sponsoring Org.:
USDOE
OSTI Identifier:
1036433
Report Number(s):
PNNL-SA-80672
KJ0403000; TRN: US201206%%306
DOE Contract Number:  
AC05-76RL01830
Resource Type:
Conference
Resource Relation:
Conference: Proceedings of the Scientific Discovery through Advanced Computing Conference (SciDAC 2011), July 10-14, 2011, Denver, Colorado
Country of Publication:
United States
Language:
English
Subject:
99 GENERAL AND MISCELLANEOUS//MATHEMATICS, COMPUTING, AND INFORMATION SCIENCE; AUTOMATION; DATA ACQUISITION; DATA ANALYSIS; EXPLORATION; MANAGEMENT; SIMULATION; STORAGE

Citation Formats

Shoshani, Arie, Altintas, Ilkay, Chen, Jin, Chin, George, Choudhary, Alok, Crawl, Daniel, Critchlow, Terence J., Gao, K., Grimm, B., Iyer, H., Kamath, Chandrika, Khan, Ayla, Klasky, S., Koehler, Sven, Lang, Rob, Latham, Robert J., Li, J. W., Liao, Wei-keng, Ligon, J., Liu, Q., Ludaescher, Bertram T., Mouallem, Pierre, Nagappan, Mie, Podhorszki, Norbert, Ross, Rob, Rotem, Doron, Samatova, Nagiza F., Silva, C., Sim, A., Tchoua, Roselynne, Thakur, R., Vouk, M., Wu, J., and Yu, Weikuan. The Scientific Data Management Center: Available Technologies and Highlights. United States: N. p., 2011. Web.
Shoshani, Arie, Altintas, Ilkay, Chen, Jin, Chin, George, Choudhary, Alok, Crawl, Daniel, Critchlow, Terence J., Gao, K., Grimm, B., Iyer, H., Kamath, Chandrika, Khan, Ayla, Klasky, S., Koehler, Sven, Lang, Rob, Latham, Robert J., Li, J. W., Liao, Wei-keng, Ligon, J., Liu, Q., Ludaescher, Bertram T., Mouallem, Pierre, Nagappan, Mie, Podhorszki, Norbert, Ross, Rob, Rotem, Doron, Samatova, Nagiza F., Silva, C., Sim, A., Tchoua, Roselynne, Thakur, R., Vouk, M., Wu, J., & Yu, Weikuan. The Scientific Data Management Center: Available Technologies and Highlights. United States.
Shoshani, Arie, Altintas, Ilkay, Chen, Jin, Chin, George, Choudhary, Alok, Crawl, Daniel, Critchlow, Terence J., Gao, K., Grimm, B., Iyer, H., Kamath, Chandrika, Khan, Ayla, Klasky, S., Koehler, Sven, Lang, Rob, Latham, Robert J., Li, J. W., Liao, Wei-keng, Ligon, J., Liu, Q., Ludaescher, Bertram T., Mouallem, Pierre, Nagappan, Mie, Podhorszki, Norbert, Ross, Rob, Rotem, Doron, Samatova, Nagiza F., Silva, C., Sim, A., Tchoua, Roselynne, Thakur, R., Vouk, M., Wu, J., and Yu, Weikuan. Fri . "The Scientific Data Management Center: Available Technologies and Highlights". United States.
@article{osti_1036433,
title = {The Scientific Data Management Center: Available Technologies and Highlights},
author = {Shoshani, Arie and Altintas, Ilkay and Chen, Jin and Chin, George and Choudhary, Alok and Crawl, Daniel and Critchlow, Terence J. and Gao, K. and Grimm, B. and Iyer, H. and Kamath, Chandrika and Khan, Ayla and Klasky, S. and Koehler, Sven and Lang, Rob and Latham, Robert J. and Li, J. W. and Liao, Wei-keng and Ligon, J. and Liu, Q. and Ludaescher, Bertram T. and Mouallem, Pierre and Nagappan, Mie and Podhorszki, Norbert and Ross, Rob and Rotem, Doron and Samatova, Nagiza F. and Silva, C. and Sim, A. and Tchoua, Roselynne and Thakur, R. and Vouk, M. and Wu, J. and Yu, Weikuan},
abstractNote = {Managing scientific data has been identified by the scientific community as one of the most important emerging needs because of the sheer volume and increasing complexity of data being collected. Effectively generating, managing, and analyzing this information requires a comprehensive, end-to-end approach to data management that encompasses all of the stages from the initial data acquisition to the final analysis of the data. Based on community input, we have identified three significant requirements. First, more efficient access to storage systems is needed. In particular, parallel file system and I/O system improvements are needed to write and read large volumes of data without slowing a simulation. Second, scientists require technologies to facilitate better understanding of their data, in particular the ability to effectively perform complex data analysis and searches over extremely large data sets. Furthermore, exploratory analysis requires techniques for efficiently selecting subsets of the data. Third, generating the data, collecting and storing the results, keeping track of data provenance, data post-processing, and analysis of results is a tedious, fragmented process. Tools for automation of this process in a robust, tractable, and recoverable fashion are required to enhance scientific exploration.},
doi = {},
journal = {},
number = ,
volume = ,
place = {United States},
year = {2011},
month = {9}
}

Conference:
Other availability
Please see Document Availability for additional information on obtaining the full-text document. Library patrons may search WorldCat to identify libraries that hold this conference proceeding.

Save / Share: