Measuring Thread Timing to Assess the Feasibility of Early-Bird Message Delivery Across Systems and Scales

Marts, W. Pepper; Dosanjh, Matthew G. F.; Schonbein, Whit; Levy, Scott; Bridges, Patrick G.

doi:10.1002/cpe.8342

Measuring Thread Timing to Assess the Feasibility of Early-Bird Message Delivery Across Systems and Scales

Journal Article · Wed Dec 11 23:00:00 EST 2024 · Concurrency and Computation. Practice and Experience

DOI:https://doi.org/10.1002/cpe.8342· OSTI ID:2516801

^[1]; ^[2]; Schonbein, Whit ^[2]; ^[2]; Bridges, Patrick G. ^[3]

Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States); Univ. of New Mexico, Albuquerque, NM (United States)
Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States)
Univ. of New Mexico, Albuquerque, NM (United States)

Early-bird communication is a communication/computation overlap technique that leverages fine-grained communication to improve application run-time. Communication is divided such that each individual thread can initiate transmission of its portion of the data upon completion rather than waiting for a dedicated communication phase. The benefit of early-bird communication depends on the completion timing of the individual threads: On the one hand, if all threads are complete at nearly the same time, the overheads of sending multiple messages will accumulate, leading to performance that is worse than if a single message had been sent. On the other hand, if thread completions are spread out in time, those that complete earlier can send data while others continue working, leading to performance that is better than if a single message had been sent. The challenge is that the completion times are currently unknown and can vary based on application, problem size, system software, and underlying hardware. In this paper, we address this lacuna by measuring and evaluating the potential overlap afforded by early-bird communication for a selection of proxy applications. These measurements help us understand whether a given application could benefit from early-bird communication. Here, we present our technique for gathering this data and evaluate data collected from three proxy applications: MiniFE, MiniMD, and MiniQMC. Each application is run on three systems with distinct CPU architectures and strong scales across three run sizes. To characterize the behavior of these workloads, we study the trends of thread timings at both a macro level, across all threads across all runs of an application, and a micro level, that is, within a single process of a single run. We observe that our tested applications exhibit significantly different thread arrival distributions. The machine used had a significant impact, with the window of potential overlap varying by as much as an order of magnitude.

View Accepted Manuscript (DOE)

Research Organization:: Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States)

Sponsoring Organization:: USDOE Office of Science (SC), Advanced Scientific Computing Research (ASCR); USDOE National Nuclear Security Administration (NNSA)

Grant/Contract Number:: NA0003525; NA0003966

OSTI ID:: 2516801

Alternate ID(s):: OSTI ID: 2483373

Report Number(s):: SAND--2025-01664J

Journal Information:: Concurrency and Computation. Practice and Experience, Journal Name: Concurrency and Computation. Practice and Experience Journal Issue: 1 Vol. 37; ISSN 1532-0626

Publisher:: WileyCopyright Statement

Country of Publication:: United States

Language:: English

References (17)

A survey of MPI usage in the US exascale computing project: A survey of MPI usage in the U. S. exascale computing project Bernholdt, David E.; Boehm, Swen; Bosilca, George Concurrency and Computation: Practice and Experience https://doi.org/10.1002/cpe.4851	journal	September 2018
LAMMPS - a flexible simulation tool for particle-based materials modeling at the atomic, meso, and continuum scales Thompson, Aidan P.; Aktulga, H. Metin; Berger, Richard Computer Physics Communications, Vol. 271 https://doi.org/10.1016/j.cpc.2021.108171	journal	February 2022
Test suite for evaluating performance of multithreaded MPI communication Thakur, Rajeev; Gropp, William Parallel Computing, Vol. 35, Issue 12 https://doi.org/10.1016/j.parco.2008.12.013	journal	December 2009
EDF Statistics for Goodness of Fit and Some Comparisons Stephens, M. A. Journal of the American Statistical Association, Vol. 69, Issue 347 https://doi.org/10.1080/01621459.1974.10480196	journal	September 1974
An omnibus test of normality for moderate and large size samples D'Agostino, Ralph B. Biometrika, Vol. 58, Issue 2 https://doi.org/10.1093/biomet/58.2.341	journal	January 1971
CMB: A Configurable Messaging Benchmark to Explore Fine-Grained Communication Marts, W. Pepper; Kruse, Donald A.; Dosanjh, Matthew G. F. 2024 IEEE 24th International Symposium on Cluster, Cloud and Internet Computing (CCGrid) https://doi.org/10.1109/CCGrid59990.2024.00013	conference	May 2024
A Quantitative Analysis of OS Noise Morari, Alessandro; Gioiosa, Roberto; Wisniewski, Robert W. 2011 IEEE International Parallel & Distributed Processing Symposium https://doi.org/10.1109/IPDPS.2011.84	conference	May 2011
Scalable load-balance measurement for SPMD codes Gamblin, Todd; de Supinski, Bronis R.; Schulz, Martin 2008 SC - International Conference for High Performance Computing, Networking, Storage and Analysis https://doi.org/10.1109/SC.2008.5222553	conference	November 2008
Enabling Efficient Multithreaded MPI Communication through a Library-Based Implementation of MPI Endpoints Sridharan, Srinivas; Dinan, James; Kalamkar, Dhiraj D. SC14: International Conference for High Performance Computing, Networking, Storage and Analysis https://doi.org/10.1109/SC.2014.45	conference	November 2014
A new approach for performance analysis of openMP programs Liu, Xu; Mellor-Crummey, John; Fagan, Michael Proceedings of the 27th international ACM conference on International conference on supercomputing https://doi.org/10.1145/2464996.2465433	conference	June 2013
Enabling MPI interoperability through flexible communication endpoints Dinan, James; Balaji, Pavan; Goodell, David Proceedings of the 20th European MPI Users' Group Meeting on - EuroMPI '13 https://doi.org/10.1145/2488551.2488553	conference	January 2013
Improving concurrency and asynchrony in multithreaded MPI applications using software offloading Vaidyanathan, Karthikeyan; Kalamkar, Dhiraj D.; Pamnany, Kiran Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis on - SC '15 https://doi.org/10.1145/2807591.2807602	conference	January 2015
Grain graphs Muddukrishna, Ananya; Jonsson, Peter A.; Podobas, Artur ACM SIGPLAN Notices, Vol. 51, Issue 8 https://doi.org/10.1145/3016078.2851156	journal	February 2016
Modeling and Benchmarking the Potential Benefit of Early-Bird Transmission in Fine-Grained Communication Schonbein, Whit; Levy, Scott; Dosanjh, Matthew G. F. Proceedings of the 52nd International Conference on Parallel Processing https://doi.org/10.1145/3605573.3605618	conference	August 2023
Measuring Thread Timing to Assess the Feasibility of Early-bird Message Delivery Marts, William Pepper; Dosanjh, Matthew G. F.; Schonbein, Whit Proceedings of the 52nd International Conference on Parallel Processing Workshops https://doi.org/10.1145/3605731.3605884	conference	August 2023
Characterizing Task-Based OpenMP Programs Muddukrishna, Ananya; Jonsson, Peter A.; Brorsson, Mats PLOS ONE, Vol. 10, Issue 4 https://doi.org/10.1371/journal.pone.0123545	journal	April 2015
WOMBAT: A Scalable and High-performance Astrophysical Magnetohydrodynamics Code Mendygral, P. J.; Radcliffe, N.; Kandalla, K. The Astrophysical Journal Supplement Series, Vol. 228, Issue 2 https://doi.org/10.3847/1538-4365/aa5b9c	journal	February 2017

Similar Records

Message passing with queues and channels

Patent · Tue Sep 24 00:00:00 EDT 2013 · OSTI ID:1096003

Send-side matching of data communications messages

Patent · Tue Jun 17 00:00:00 EDT 2014 · OSTI ID:1134199

Send-side matching of data communications messages

Patent · Tue Jul 01 00:00:00 EDT 2014 · OSTI ID:1136754

Related Subjects

97 MATHEMATICS AND COMPUTING

Measuring Thread Timing to Assess the Feasibility of Early-Bird Message Delivery Across Systems and Scales

Citation Formats

References (17)

Similar Records

Related Subjects