The energy consumption numbers for training frontier AI models have crossed a threshold that makes them difficult to ignore. Training a single large language model now consumes between 50 and 100 gigawatt-hours of electricity, depending on model size and training duration. That is roughly equivalent to the annual electricity consumption of 5,000 to 10,000 households. The carbon emissions from that energy consumption depend on the energy grid where the training occurs, but they are significant in any location.
These numbers are not new in kind. Researchers have been measuring AI training energy consumption for years. What has changed is the scale. As models have grown from billions to trillions of parameters, and as training runs have extended from weeks to months, the energy consumption has grown proportionally. The latest research, published across multiple independent groups this quarter, converges on the conclusion that the current trajectory of model scaling is unsustainable without significant changes to either model efficiency or energy infrastructure.
What the research actually shows
Three papers published this quarter deserve specific attention because they move beyond headline numbers to provide actionable analysis.
The first paper measures the energy consumption of training runs across different model architectures and sizes. The key finding is that energy consumption scales roughly linearly with parameter count and super-linearly with training data volume. This means that doubling the training data for a fixed model size more than doubles the energy cost, because the additional data requires additional training steps, each of which consumes energy proportional to the model size. The practical implication is that the cost of data collection and curation (measured in energy, not just dollars) is a significant and growing component of the total cost of training.
The second paper compares the energy efficiency of different hardware platforms for training. The finding is that newer hardware is substantially more energy-efficient than older hardware for the same training task. Training the same model on NVIDIA’s latest architecture consumes approximately 40 percent less energy than training it on hardware that is two generations old. This is good news for the environmental case, because it means hardware upgrades partially offset the growth in model size. It does not fully offset it, but it narrows the gap.
The third paper examines the carbon footprint of inference, not just training. Training gets most of the attention because it is concentrated and dramatic. A single training run consuming months of compute. But inference, because it runs continuously across millions of users, may consume more total energy over the lifetime of a model than the training run that produced it. The paper estimates that for a widely deployed model, cumulative inference energy consumption exceeds training energy consumption within three to six months of deployment. This finding shifts the focus from training efficiency to inference efficiency as the more important lever for reducing AI’s carbon footprint.
The efficiency imperative
The research creates a clear imperative for data teams: efficiency matters more than raw performance. A model that is 5 percent less accurate but 50 percent more energy-efficient may be the better choice for production deployment, particularly when the accuracy gap is on metrics that do not affect your specific use case.
This is a shift in how many teams evaluate models. The prevailing approach has been to choose the most capable model that fits within the budget, where capability is measured by benchmark accuracy. The efficiency-aware approach adds energy consumption and carbon footprint as evaluation criteria alongside accuracy and cost. For some use cases, the accuracy requirement is non-negotiable. For many production use cases, the accuracy requirement is met by models that are smaller and more efficient than the frontier models that dominate the benchmarks.
Model distillation, training a smaller model to replicate the behaviour of a larger model, is the most direct route to efficiency gains. A distilled model that is one-tenth the size of its teacher model may retain 90 percent of the teacher’s accuracy on domain-specific tasks while consuming a fraction of the energy for both training and inference. The trade-off is that distillation requires a large teacher model to start from, which means the initial energy investment is still large. But the ongoing inference savings can be substantial.
Quantisation, reducing the numerical precision of model weights, offers another efficiency lever. Running a model at 8-bit precision instead of 16-bit precision roughly halves the memory footprint and energy consumption per inference, with minimal accuracy loss for most tasks. For inference at scale, this is a meaningful reduction. The accuracy loss is typically concentrated on tasks that require precise numerical reasoning, which many production workloads do not.
The reporting and compliance angle
Environmental reporting requirements for AI are beginning to appear in corporate sustainability frameworks. The EU’s Corporate Sustainability Reporting Directive requires large companies to disclose energy consumption and carbon emissions across their operations, including digital infrastructure. AI training and inference are part of that digital infrastructure, and they are growing as a share of total energy consumption.
Companies that are subject to sustainability reporting requirements should start measuring AI energy consumption now, even if reporting is not yet mandatory for their specific category. The measurement infrastructure (energy monitoring on GPU clusters, carbon intensity tracking for cloud workloads, lifecycle analysis for hardware) takes time to build. Companies that wait until reporting is mandatory will be scrambling to collect data they should have been collecting all along.
Cloud providers are beginning to offer carbon footprint reporting for AI workloads. These reports are useful but incomplete. They typically cover the energy consumed by the compute instances but not the energy consumed by cooling, networking, or storage infrastructure that supports those instances. For a more complete picture, you need to apply power usage effectiveness ratios to the reported compute energy, which increases the total by 30 to 60 percent depending on the data centre’s efficiency.
The bounded recommendation
Add energy efficiency to your model evaluation criteria. For each new AI workload, estimate the annual inference energy consumption at your expected scale. If a smaller or quantised model meets your accuracy requirements, use it instead of the larger alternative. Start measuring energy consumption for your existing AI workloads, even if the measurement is approximate. You will need these numbers for sustainability reporting, and having them early gives you time to optimise before the reporting deadline.