Accelerating large-scale plane-wave hybrid functional calculations via low-rank approximations and high-performance computing on the Sugon supercomputer
Description
The rapidly increasing computational demand of large-scale density functional theory (DFT) simulations poses significant performance and scalability challenges for traditional CPU-based supercomputers. Recent advances in GPU-like accelerator architectures, featuring massive on-chip parallelism and high memory bandwidth, are reshaping the computational landscape of large-scale first-principles calculations. In this work, we port the plane-wave DFT software PWDFT to a GPU-like deep computing unit (DCU) heterogeneous platform on the Sugon supercomputer and achieve significant acceleration for hybrid functional calculations through low-rank approximations and high-performance computing optimizations. We integrate the adaptively compressed exchange (ACE) approximation with the projected commutator direct inversion of the iterative subspace (PCDIIS) scheme, thereby reducing the computational time per self-consistent step to only about four times that of a corresponding PBE calculation under the benchmark settings. We further exploit DCU hardware acceleration through optimized linear algebra and FFT libraries, as well as communication-aware kernel optimizations. We demonstrate that plane-wave hybrid DFT calculations for a 2,048-atom silicon system can be performed using only 32 DCUs, and that the framework efficiently scales to employ 4096 DCUs for simulations of a 4,096-atom system. These advances establish an efficient computational framework for large-scale plane-wave hybrid functional calculations at the thousand-atom scale.