About $1.35M to let users choose power-performance tradeoffs when running HPC and AI — including reporting energy consumed for every run (Ohio State University)
Current exascale HPC and AI systems use powerful accelerators and many-core processors and consume significant power. Prior work has focused on achieving maximum performance under power constraints, leaving open whether applications can run with reduced energy while tolerating minimal performance degradation. This project pursues that tradeoff within the MPI library.
Grant overview (primary data)
- Award amount$1,348,501
- RecipientOHIO STATE UNIVERSITY, THE (Ohio)
- ProgramSoftware Institutes
- Period2026-08-01 〜 2029-07-31
- FunderU.S. National Science Foundation (NSF) / NSF
Key points
- It addresses the open question of whether HPC and AI applications can run with reduced energy while tolerating minimal performance degradation.
- Three needs are stated: MPI libraries achieving power-performance tradeoffs, reporting energy consumption per run, and an easy mechanism for specifying tradeoff modes.
- Prior work aimed at maximum performance under power constraints; this project aims to let users choose the exchange rate between speed and energy.
- The work modifies the MVAPICH MPI library and communication runtime, where moving data among nodes consumes power and time.
- Estimated total and obligated amount are both $1,348,501, running from 2026-08-01 to 2029-07-31.
- Three requirements depend on one another in order: create room to trade, report energy per run, and let the user choose the mode.
1Aiming at choice rather than the fastest
Prior work has aimed at extracting maximum performance within a given power envelope. What this project sets out to replace is that premise. If a slight loss of performance is permitted, how far can energy consumption fall? Rather than maximizing speed, the goal is letting users choose the exchange rate between speed and energy.
With the constraint on computing resources becoming power itself, what is chosen as the object of optimization is changing.
2The problem of invisible power
The second of the three stated requirements is reporting energy consumption to the end user for each application run. Put the other way, how much power a computation used is not currently returned to the user. What cannot be seen cannot be reduced. Speed is evident to anyone as elapsed time, whereas power goes unnoticed without machinery to measure and convey it.
That the third requirement is an easy-to-use mechanism for specifying tradeoff modes follows the same judgment: options mean nothing unless they can actually be used.
3The communication library as the place to work
What is modified is the MPI library and communication runtime. In large-scale computation, moving data among many nodes consumes power and time more than the computation itself. Specific design targets named in the record include transport protocols, on-the-fly compression, non-contiguous transfers and load-imbalance awareness.
Rather than rewriting applications, changing the common layer beneath them makes the effect reach many computations at once. The 120 NSF awards this site holds as of 2026-08-31 span 67 programs, and Software Institutes, where this belongs, addresses the foundations of research software.
4Three requirements that run in one line
The three requirements named here are not separate improvements; they depend on one another in order. Create room to trade, show what was consumed, then let the user choose — drop any one and the rest go unused.
- 1Create room to tradeIf a small loss of performance is permitted, how much energy falls? Make that ratio something a design can work with
- 2Show the power that was usedReport energy consumption to the user for each application run. What cannot be seen cannot be reduced
- 3Let the user chooseProvide a usable mechanism for specifying the trade-off mode, so the options exist in a form people can actually reach
That the work targets the MPI library and the communication runtime fits the same design. At scale, moving data between nodes consumes more power and time than the computation itself. Changing the common layer underneath, rather than rewriting applications, reaches many workloads at once.
Why it matters
As the constraint on computing shifts to power itself, the target of optimization moves from speed to an exchange rate. Without power consumed being visible per run, no decision to reduce it can be made. Working at a shared library layer is a referenceable way to reach broadly without rewriting individual applications.
FAQ
Why consider giving up performance?
Why report energy consumption?
Why the MPI library?
Sources (primary)
Source: NSF Award Search (U.S. National Science Foundation, public domain). Amounts are the obligated amount. For privacy, we do not handle principal investigator names.
- NSF Award (original, official)
- NSF Award ID: 2608412