Energy efficiency is the measure of useful computational work delivered per unit of energy consumed. In AI infrastructure, accelerated computing is energy-efficient computing—maximizing tokens, tasks, and performance per watt to drive revenue within fixed power budgets.
Energy efficiency is the practice of maximizing computational output while minimizing energy consumption, measured in work completed per unit of energy—expressed as work per kilowatt-hour (kWh).
It’s important to distinguish between power and energy, as the terms are often used interchangeably. Power is how fast energy is used at a given moment, measured in watts. Energy is the total electricity consumption over time, measured in kilowatt-hours or megajoules. Performance per watt measures throughput efficiency—for example, tokens per second per watt—while energy efficiency measures completed work per unit of energy, such as tokens per kilowatt-hour.
Energy and power are related quantities, but they’re different. Energy is power times time, whereas power is energy per unit of time.
Energy efficiency goes beyond power efficiency by emphasizing the highest level of computational work with minimal total energy use. Accelerated computing completes workloads significantly faster than traditional computing, returning hardware to a low-power idle state sooner and consuming less energy per task—even when peak active power is higher.
Hardware Architecture and GPU Domain Size
Successive GPU generations deliver compounding efficiency gains with every release. Domain size—the number of GPUs connected over a high-speed scale-up interconnect—also matters: Larger GPU domains serve today’s mixture of experts (MoE) frontier AI models more efficiently than smaller ones. Every watt in an accelerated AI factory can generate up to 50x more throughput per watt for specific AI workloads.
Inference Software Optimization and Quantization
Techniques such as FP4 quantization, disaggregated serving, expert parallelism, and key-value (KV) cache offloading stack together to increase tokens generated per watt. Reducing model-weight precision from 8-bit to 4-bit reduces memory footprint, bandwidth, and energy required per token. Software improvements compound over time.
Cooling and Power Management
In a typical data center, only about 60% of electricity drawn from the grid becomes useful AI compute; the rest is lost to cooling and power conversion overhead. Liquid-cooled systems raise rack density, cut cooling overhead, can improve water efficiency by over 300%, and increase performance per rack within existing facility constraints. With a 45°C coolant design, operators can further boost energy efficiency within fixed power budgets while relying on a closed-loop coolant system with zero water consumption.
Quick Links
Energy efficiency is a priority for any organization operating compute-intensive workloads—from hyperscale AI factories to governments and regulated industries operating within constrained infrastructure budgets.
Listed below are a few examples of the many costs and solutions data center operators can adopt to optimize and budget for energy-efficient data center strategies. Ultimately, the best approach will depend on the specific needs and constraints of each data center.
Quick Links
Energy efficiency measures how much useful work a system completes per unit of energy consumed. In AI inference, it is measured in tokens per joule or tokens per kilowatt hour.
Energy efficiency emphasizes the electricity required to complete a unit of work.
Performance per watt measures how much throughput a system delivers for a specific workload within a given power draw. In AI inference, it is measured in tokens per second per watt, and it emphasizes operating within a power, thermal, or rack-capacity constraint.
Power usage effectiveness (PUE) measures how much total energy a data center uses for every unit of energy delivered to its IT equipment. PUE should be applied when metrics need to reflect total data-center electricity rather than IT equipment electricity alone:
PUE = facility energy/IT energy
Divide IT-level tokens/kWh and tokens/s/W by PUE to obtain facility-level values.
Multiply IT energy-per-token by PUE to obtain facility energy-per-token.
Energy costs in AI training are shaped by model size and architecture, since larger parameter counts and more complex designs directly increase the compute required per training step.
Dataset size and the number of experimental runs compound this further, as pretraining, ablations, hyperparameter sweeps, and fine-tuning iterations each add their own energy footprint on top of the final production run.
The scale and power density of the AI infrastructure, that is, how many GPUs are deployed and how tightly they’re packed per rack, sets the baseline draw, while communication overhead from multi-GPU and multi-node parallelism (interconnect fabrics, network switching) adds energy that scales with cluster size rather than useful computation.
Energy costs in AI inference are shaped by model choice and size. Serving a smaller, task-appropriate model consumes far less energy per query than routing everything through a large frontier model.
Context length and token volume can be major multipliers as long-reasoning or agentic queries can consume much more energy per query than short standard queries.
Additionally, image or video generation tasks can be orders of magnitude more energy-intensive than simple text queries.
Request volume and usage patterns matter because inference runs continuously at product scale, so aggregate energy grows directly with traffic even as per-query energy improves.
Hardware efficiency, batching, and runtime optimizations (such as continuous batching, quantization, speculative decoding, and KV-cache management) can cut inference energy substantially without degrading user experience.
Across both training and inference, a few factors behave as shared multipliers on total energy cost.
Cooling overhead determines how much non-compute energy is spent removing heat from dense GPU clusters, with liquid cooling typically being more efficient than air cooling.
Electricity price and demand changes affect operating cost directly and create opportunities for hyperscalers to shift flexible workloads to times of abundant, cheaper renewable power.
Power infrastructure, which is the capacity and reliability of grid connections, on-site generation, and distribution to increasingly power-dense racks sets a hard ceiling on how much compute can be deployed at all.
Finally, utilization ties everything together: Running hardware at high, sustained utilization amortizes fixed power draw (cooling, networking, idle GPU baseline) across more useful work.
PUE measures total facility energy divided by IT-equipment energy, so a lower value means less overhead spent on cooling and power delivery relative to the energy actually powering the servers.
The global industry average PUE is approximately 1.5–1.6.
Germany’s Energy Efficiency Act has set a bar of 1.2 or below for new data centers opening after July 2026.
Facilities using containment, liquid cooling, and modern power distribution can realistically achieve a PUE of 1.2 or below.
To capture how effectively AI infrastructure converts electricity into useful AI output, PUE should be combined with workload-level metrics such as tokens per joule.
Running coolant warmer makes a data center more efficient because it shrinks the temperature gap that cooling systems must close, letting the facility rely on passive, low-energy heat rejection instead of energy-intensive refrigeration.
Traditional systems actively refrigerate water below ambient air temperature so it can absorb heat from IT equipment, but once coolant runs at 35 to 45°C, outdoor dry coolers or cooling towers can often reject that heat using just fans and ambient air, without engaging chillers at all.
This is possible because direct-to-chip liquid cooling captures heat right at the processor rather than cooling an entire room of air, so chips stay within safe limits even when coolant enters around 45°C and exits around 55°C.
The savings compound quickly: Raising chiller-plant temperatures by just one degree can cut cooling energy costs by roughly 4%.
Separately, a 50-megawatt facility could save more than $4 million a year in cooling-related energy and water costs by shifting to warmer liquid cooling, though hot or humid climates still require chillers as backup for peak heat days.
Grid timing is the practice of aligning data center energy use with when electricity on the grid is cleanest and cheapest.
By shifting flexible AI workloads to hours when the grid is less congested and renewable generation is abundant, data centers can lower electricity costs and reduce emissions without changing their hardware.
Reducing or curtailing load during peak-demand periods helps avoid expensive grid upgrades and can improve system-wide reliability and rates.
Utilities and regulators increasingly value “power-flexible” AI factories that can dynamically ramp their consumption up or down in response to real-time grid conditions.
NVIDIA offers two power profile configurations to match different operational goals. Max-P prioritizes maximum performance, allowing higher power consumption to achieve the greatest throughput.
Max-Q prioritizes efficiency, balancing performance and power consumption to maximize work per unit of energy consumed. Data center deployments evaluate the trade-off between these profiles based on revenue targets and power constraints.
Explore NVIDIA AI infrastructure and learn how NVIDIA’s full-stack approach maximizes performance per watt from AI factory design to silicon and software.
By advancing energy‑efficient AI and accelerated computing, NVIDIA is reducing product and operational emissions, and partnering globally to use AI to save energy and help society adapt to climate change.