DOE PAGES title logo U.S. Department of Energy
Office of Scientific and Technical Information

Title: Providing performance portable numerics for Intel GPUs

Journal Article · · Concurrency and Computation. Practice and Experience
DOI: https://doi.org/10.1002/cpe.7400 · OSTI ID:1894739
ORCiD logo [1]; ORCiD logo [1]; ORCiD logo [2]
  1. Steinbuch Centre for Computing Karlsruhe Institute of Technology Karlsruhe Baden‐Württemberg Germany
  2. Steinbuch Centre for Computing Karlsruhe Institute of Technology Karlsruhe Baden‐Württemberg Germany, The Innovative Computing Laboratory University of Tennessee Knoxville Tennessee

Summary With discrete Intel GPUs entering the high‐performance computing landscape, there is an urgent need for production‐ready software stacks for these platforms. In this article, we report how we enable the Ginkgo math library to execute on Intel GPUs by developing a kernel backed based on the DPC++ programming environment. We discuss conceptual differences between the CUDA and DPC++ programming models and describe workflows for simplified code conversion. We evaluate the performance of basic and advanced sparse linear algebra routines available in Ginkgo's DPC++ backend in the hardware‐specific performance bounds and compare against routines providing the same functionality that ship with Intel's oneMKL vendor library.

Research Organization:
University of Tennessee, Knoxville, TN (United States)
Sponsoring Organization:
USDOE; USDOE National Nuclear Security Administration (NNSA); USDOE Office of Science (SC)
OSTI ID:
1894739
Journal Information:
Concurrency and Computation. Practice and Experience, Journal Name: Concurrency and Computation. Practice and Experience Journal Issue: 20 Vol. 35; ISSN 1532-0626
Publisher:
Wiley Blackwell (John Wiley & Sons)Copyright Statement
Country of Publication:
United Kingdom
Language:
English

References (13)

Multiprecision Block-Jacobi for Iterative Triangular Solves book January 2020
Preparing Ginkgo for AMD GPUs – A Testimonial on Porting CUDA Code to HIP book January 2021
Porting Sparse Linear Algebra to Intel GPUs book January 2022
A quantitative roofline model for GPU kernel performance estimation using micro-benchmarks and hardware metric profiling journal September 2017
ParILUT - A Parallel Threshold ILU for GPUs conference May 2019
Batched Generation of Incomplete Sparse Approximate Inverses on GPUs conference November 2016
GMRES: A Generalized Minimal Residual Algorithm for Solving Nonsymmetric Linear Systems journal July 1986
Fine-Grained Parallel Incomplete LU Factorization journal January 2015
ParILUT---A New Parallel Threshold ILU Factorization journal January 2018
Algorithm 915, SuiteSparseQR journal November 2011
Khronos SYCL for OpenCL conference January 2015
Adaptive Precision Block-Jacobi for High Performance Preconditioning in the Ginkgo Linear Algebra Software journal April 2021
Ginkgo: A high performance numerical linear algebra library journal August 2020