gshc2020.com

Council Post: The New Currency: Rethinking Computing Power In The Age Of AI

tags:
@ 03/08/2026

By Dr. Steven Woo, fellow and distinguished inventor at Rambus.

getty

In the era of gigawatt AI, the systems that succeed will be those that treat every watt as a strategic asset.

​For the past several years, the AI industry has been focused on one thing: faster AI processing. More powerful GPUs and bigger models became standard measures of progress.​

But as AI systems scale, especially for large language model (LLM) training and inference, power is emerging as one of the industry’s biggest constraints. The rapid growth of inference workloads in particular, combined with rising infrastructure demands, is forcing the industry to rethink how AI systems scale. This shift is already reshaping how AI infrastructure is designed, deployed and measured.​

From Performance Scaling To Efficiency Scaling​

Performance and power have always been connected in computing. Historically, higher performance came with higher power consumption, and system designs often followed that tradeoff. In AI, however, that approach is becoming less sustainable.​

At scale, maximizing performance alone is no longer enough. The focus is increasingly on how much power systems consume during computation, and the industry’s focus is on reducing the energy required to process tokens and run inference workloads. AI deployments are now measured as much by power consumption as they are performance, with infrastructure planning increasingly happening at the scale of hundreds of megawatts up to gigawatt-scale campuses. At these levels, power efficiency has become central to the economics of AI.​

Memory And Data Movement Drive Power Consumption​

Surprisingly, one of the key contributors to power consumption in AI systems is data movement. In many architectures, moving data between chips and across systems consumes much more power than the mathematical operations that form the core of AI computation. As models grow in size and context, the volume of data accessed and moved through the system increases significantly, which drives up power consumption in turn. At the same time, compute capabilities have scaled faster than the ability to efficiently move data, creating a growing imbalance between computation and data movement.​

This shift is putting far more pressure on memory systems. Memory can no longer be simply thought of as supporting compute; it is increasingly becoming a primary determinant of overall power efficiency and performance. In many systems, memory and data movement account for a substantial share of total power consumption. As a result, minimizing data movement, reducing the distance data must travel and improving the efficiency of memory accesses have become key focus areas for the industry.​

Infrastructure Is Evolving Around Efficiency​

This pressure to improve power efficiency is driving substantial changes in AI system architectures. For example, specialized hardware tailored to different phases of LLM inference have been created that increase performance and reduce power consumption. The LLM prefill phase is compute-intensive, while the generation phase is more memory bandwidth-intensive. Because these phases have different requirements, using a single type of processor for both can lead to inefficiencies. As a result, the industry is beginning to adopt architectures that optimize compute and memory resources separately for each phase.​

The industry is also revisiting architectures designed to reduce data movement within AI systems, including processing closer to memory and integrating compute and memory more tightly to improve performance while limiting power growth.​

The Edge Makes Power Constraints More Visible​

While data centers highlight the scale of the power challenge, edge and client devices are often more tightly constrained due to the use of batteries and small form factors.​

While large data centers can afford to increase power consumption to achieve higher performance, edge devices such as smartphones must operate within fixed battery power budgets and strict thermal limits due to the lack of fans. As AI continues to move closer to the edge, these constraints are becoming increasingly important in shaping client and edge architectures.​

From More Compute To More Efficient Compute​

Together, these shifts are changing how AI systems are built and evaluated. Power is no longer a secondary design consideration. Rather, it is a first-class design constraint that influences everything from system architecture to hardware specialization to memory system design.​

As AI models grow larger and inference workloads expand, power efficiency has become just as important as raw performance. Rising power densities are also driving changes beyond the chip itself, with the increased adoption of liquid cooling in data centers to manage the heat generated in modern AI platforms.​

AI systems have always been shaped by tradeoffs between compute, memory and architecture. Increasingly, power is determining how those tradeoffs are made, and how AI infrastructure will scale going forward. In the era of gigawatt AI, the systems that succeed will be those that treat every watt as a strategic asset.


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?