skip to main content
OSTI.GOV title logo U.S. Department of Energy
Office of Scientific and Technical Information

Title: Imputation for multisource data with comparison and assessment techniques

Journal Article · · Applied Stochastic Models in Business and Industry
DOI:https://doi.org/10.1002/asmb.2299· OSTI ID:1416279

Missing data are prevalent issue in analyses involving data collection. The problem of missing data is exacerbated for multisource analysis, where data from multiple sensors are combined to arrive at a single conclusion. In this scenario, it is more likely to occur and can lead to discarding a large amount of data collected; however, the information from observed sensors can be leveraged to estimate those values not observed. We propose two methods for imputation of multisource data, both of which take advantage of potential correlation between data from different sensors, through ridge regression and a state-space model. These methods, as well as the common median imputation, are applied to data collected from a variety of sensors monitoring an experimental facility. Performance of imputation methods is compared with the mean absolute deviation; however, rather than using this metric to solely rank themethods,we also propose an approach to identify significant differences. Imputation techniqueswill also be assessed by their ability to produce appropriate confidence intervals, through coverage and length, around the imputed values. Finally, performance of imputed datasets is compared with a marginalized dataset through a weighted k-means clustering. In general, we found that imputation through a dynamic linearmodel tended to be the most accurate and to produce the most precise confidence intervals, and that imputing the missing values and down weighting them with respect to observed values in the analysis led to the most accurate performance.

Research Organization:
Los Alamos National Laboratory (LANL), Los Alamos, NM (United States)
Sponsoring Organization:
USDOE
Grant/Contract Number:
AC52-06NA25396
OSTI ID:
1416279
Report Number(s):
LA-UR-17-23333
Journal Information:
Applied Stochastic Models in Business and Industry, Vol. 34, Issue 1; ISSN 1524-1904
Publisher:
WileyCopyright Statement
Country of Publication:
United States
Language:
English
Citation Metrics:
Cited by: 1 work
Citation information provided by
Web of Science

References (17)

The Bayesian elastic net journal March 2010
Foci and Bases of Employee Commitment: Implications for job Performance. journal April 1996
An Entropy Weighting k-Means Algorithm for Subspace Clustering of High-Dimensional Sparse Data journal August 2007
Methods for imputation of missing values in air quality data sets journal June 2004
Missing value estimation methods for DNA microarrays journal June 2001
Ridge Regression: Biased Estimation for Nonorthogonal Problems journal February 1970
A Bayesian Study of the Error-in-Variables Model journal August 1981
An analysis of four missing data treatment methods for supervised learning journal May 2003
Bi-level multi-source learning for heterogeneous block-wise missing data journal November 2014
Regularization Paths for Generalized Linear Models via Coordinate Descent journal January 2010
Sequential Imputations and Bayesian Missing Data Problems journal March 1994
Spectral methods for imputation of missing air quality data journal December 2015
An Introduction to Statistical Learning book January 2013
Review: A gentle introduction to imputation of missing values journal October 2006
The Calculation of Posterior Distributions by Data Augmentation journal June 1987
Sensor fusion potential exploitation-innovative architectures and illustrative applications journal January 1997
An introduction to multisensor data fusion journal January 1997

Figures / Tables (17)