Abstract
In this article, we summarize the deployment of the Air Force Weather (AFW) HPC11 system at Oak Ridge National Laboratory (ORNL) including the process followed to successfully complete acceptance testing of the system. HPC11 is the first HPE/Cray EX 3000 system that has been successfully released to its user community in a federal facility. HPC11 consists of two identical 800-node supercomputers, Fawbush and Miller, with access to two independent and identical lustre parallel file systems. HPC11 is equipped with Slingshot 10 interconnect technology and relies on the HPE Performance Cluster Manager software for system configuration. ORNL has a clearly defined acceptance testing process used to ensure that every new system deployed can provide the necessary capabilities to support user workloads. We worked closely with HPE and AFW to develop a set of tests that used the United Kingdom's Meteorological Office's Unified Model and 4-dimensional variational data assimilation. We also included benchmarks and applications from the Oak Ridge Leadership Computing Facility portfolio to fully exercise the HPE/Cray programming environment and evaluate the functionality and performance of the system. Acceptance testing of HPC11 required parallel execution of each element on Fawbush and Miller. In addition, careful coordination was needed to ensure successful acceptance of the newly deployed lustre file systems alongside the compute resources. In this work, we present test results from specific system components and provide an overview of the issues identified, challenges encountered, and the lessons learned along the way.
Original language | English |
---|---|
Article number | e7914 |
Journal | Concurrency and Computation: Practice and Experience |
Volume | 36 |
Issue number | 3 |
DOIs | |
State | Published - Feb 1 2024 |
Funding
The authors would like to thank the Cray/HPE team for their invaluable contributions to acceptance testing of HPC11. In particular, Jeff Beckleheimer, Adam Sachitano, Cathy Willis, Pete Johnsen, Eric Dolven, Dave Londo, and Kim Kafka. We would also like to thank Matt Ezell and Don Maxwell from ORNL for lending their expertise to make the transition to HPCM a success. This research used resources of the Oak Ridge Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE‐AC05‐00OR22725.
Funders | Funder number |
---|---|
Office of Science | DE‐AC05‐00OR22725 |
Keywords
- acceptance testing
- system test