Skip to main content
U.S. Department of Energy
Office of Scientific and Technical Information

Understanding Generative AI Content with Embedding Models

Dataset ·
DOI:https://doi.org/10.25584/2481996· OSTI ID:2481996

The construction of high-quality numerical features is critical to any quantitative data analysis. Feature engineering has been historically addressed by carefully hand-crafting data representations based on domain expertise. This work views the internal representations of modern deep neural networks (DNNs), called embeddings, as an implicit form of traditional feature engineering. For trained DNNs, we show that these embeddings can reveal interpretable, high-level concepts in unstructured sample data. We use these embeddings in natural language and computer vision tasks to uncover both inherent heterogeneity in the underlying data and human-understandable explanations for it. In particular, we find empirical evidence that there is inherent separability between real data and those generated from AI models.

Research Organization:
Pacific Northwest National Laboratory 2
Sponsoring Organization:
DOE
DOE Contract Number:
AC05-76RL01830
OSTI ID:
2481996
Country of Publication:
United States
Language:
English

Similar Records

Data and Code for Understanding Generative AI Content with Embedding Models
Dataset · Mon Aug 25 00:00:00 EDT 2025 · OSTI ID:2587970

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information
Journal Article · Wed Apr 06 00:00:00 EDT 2022 · IEEE Transactions on Nuclear Science · OSTI ID:1902232

Related Subjects