Optimal Bayesian supervised domain adaptation for RNA sequencing data
- Department of Electrical & Computer Engineering, Texas A&M University, College Station, TX 77843, USA; OSTI
- Department of Electrical & Computer Engineering, Texas A&M University, College Station, TX 77843, USA; TEES-AgriLife Center for Bioinformatics & Genomic Systems Engineering, Texas A&M University, College Station, TX 77843, USA
- Department of Electrical & Computer Engineering, Texas A&M University, College Station, TX 77843, USA
When learning to subtype complex disease based on next-generation sequencing data, the amount of available data is often limited. Recent works have tried to leverage data from other domains to design better predictors in the target domain of interest with varying degrees of success. But they are either limited to the cases requiring the outcome label correspondence across domains or cannot leverage the label information at all. Moreover, the existing methods cannot usually benefit from other information available a priori such as gene interaction networks.
ResultsIn this article, we develop a generative optimal Bayesian supervised domain adaptation (OBSDA) model that can integrate RNA sequencing (RNA-Seq) data from different domains along with their labels for improving prediction accuracy in the target domain. Our model can be applied in cases where different domains share the same labels or have different ones. OBSDA is based on a hierarchical Bayesian negative binomial model with parameter factorization, for which the optimal predictor can be derived by marginalization of likelihood over the posterior of the parameters. We first provide an efficient Gibbs sampler for parameter inference in OBSDA. Then, we leverage the gene-gene network prior information and construct an informed and flexible variational family to infer the posterior distributions of model parameters. Comprehensive experiments on real-world RNA-Seq data demonstrate the superior performance of OBSDA, in terms of accuracy in identifying cancer subtypes by utilizing data from different domains. Moreover, we show that by taking advantage of the prior network information we can further improve the performance.
Availability and implementationThe source code for implementations of OBSDA and SI-OBSDA are available at the following link. https://github.com/SHBLK/BSDA.
Supplementary informationSupplementary data are available at Bioinformatics online.
- Research Organization:
- Duke Univ., Durham, NC (United States)
- Sponsoring Organization:
- USDOE Office of Science (SC)
- DOE Contract Number:
- SC0019393;
- OSTI ID:
- 1853183
- Journal Information:
- Bioinformatics, Journal Name: Bioinformatics Journal Issue: 19 Vol. 37; ISSN 1367-4803
- Publisher:
- International Society for Computational Biology - Oxford University Press
- Country of Publication:
- United States
- Language:
- English
Similar Records
ACTINN: automated identification of cell types in single cell RNA sequencing