← All articles

The ESG Illusion Behind AI Training, Inferencing and Power Use

·5 min read

  • AI Agents
  • AI Consulting
  • AI Environmental Cost
  • AI Ethics
  • AI Infrastructure
  • AI Regulation & Governance
  • AI Sustainability
  • AI Training
  • AI and Artificial Intelligence
  • AI and Sustainability
  • Big Data
  • Carbon Footprint
  • Climate Change
  • Climate Impact
  • Data Center
  • Data Center Energy
  • Decarbonization
  • Deepseek
  • ESG
  • Energy Consumption
  • Energy Efficiency
  • Enterprise AI
  • Environmental Impact
  • GPU Cluster
  • GPU and NVIDIA
  • Geopolitics
  • Green Computing
  • Model Distillation
  • Net Zero
  • Sales and Marketing Automation
  • Tech Industry ESG
  • gpu
  • model-scaling
  • nvidia
  • sustainable AI

As AI advances at breakneck speed, the battle for compute power is escalating. Tech giants and nations are amassing GPU (Graphics Processing Unit) clusters, turning AI infrastructure into a geopolitical arms race. With OpenAI, Anthropic, DeepSeek, Meta, and others vying for dominance, power-hungry datacenter clusters now consume as much electricity as small cities.

Ironically, the same companies publish glowing Environmental, Social, and Governance (ESG) reports touting their commitment to a greener earth. Microsoft's emissions surged nearly 25%, compared to 2020 baseline, largely due to indirect emissions from constructing and outfitting new datacenters. AI's hunger for energy is only growing, yet the industry remains silent on its true impact. Is this innovation, or hypocrisy?

The Rise of AI Clusters

Mega datacenter clusters are at the heart of AI breakthroughs, enabling high-performance training and inference. Traditionally, data centers handled distributed computing tasks like search queries, ads, and database access. However, AI inference and training demand vastly different architectures. Making predictions or inference involves real-time processing, where user requests or prompts are sent to a datacenter, processed on a cluster of GPUs, and results are returned. Model pre-training, on the other hand, is a computationally intensive process, often requiring thousands of GPUs working simultaneously to build and refine large-scale AI models.

This raises concerns not just about raw compute usage, but also about the sustainability of AI model training. AI model Training, especially for large language models, is a critical yet energy-intensive process that requires substantial GPU resources.

AlexNet, a deep convolutional neural network that revolutionized image recognition by winning the 2012 ImageNet competition with superior accuracy, was trained on two to four GPUs, a major breakthrough at the time. GPT-3 used 20,000 A100 GPUs, consuming millions of dollars in compute power. GPT-4 scaled even further, with estimates reaching over 100,000 GPUs. NVIDIA's A100 GPU has a power consumption of approximately 400 watts, while its successor, the H100 GPU, requires around 700 watts.

OpenAI and Anthropic have clusters with 60,000–100,000 H100 equivalent GPUs

Meta disclosed buying ~400,000 GPUs but only uses a fraction for training. Llama 3 was trained on 16,000 H100 GPUs.

DeepSeek had 10,000 GPUs in 2021, now estimated at 50,000 GPUs.

Elon Musk's xAI is building the world's largest cluster in Memphis with 200,000 GPUs, estimated to consume over 140 Megawatts. NVIDIA's next-generation GPU architecture called Blackwell, will consume around 1200 watts. OpenAI's Stargate datacenter cluster in Texas is estimated to consume 2.2 Gigawatts.

Meta is building two natural gas plants for their Louisiana data center cluster. Gas plants are still significant CO₂ emitters, and they also emit methane. It is the second most significant greenhouse gas (GHG) after carbon dioxide (CO₂) in terms of its contribution to global warming. Nuclear would be a better option but building a nuclear power plant simply takes too long, hence it is not an option.

Size matters and Bigger is Better!

One of the most striking trends in AI development is the adherence to scaling laws. While early models expanded gradually, the leap from GPT-3 to GPT-4 followed a near-exponential curve, reinforcing the notion that bigger models yield better intelligence, but with caveats. The answer lies in the Chinchilla/Hoffman scaling laws paper, which states that model performance improves as the number of parameters and/or the size of the training dataset increases.

Scaling up models leads to so-called emergent abilities or new capabilities. These tasks often involve complex reasoning, problem-solving, or understanding nuanced language without specific training or finetuning. While scaling up models has led to significant breakthroughs, it's not a universal solution. Even the largest models still struggle with certain tasks. Additionally, the performance gains start to diminish as the model size and compute resources grow excessively, meaning that the cost/performance ratio becomes less favourable. Training models like GPT-4 cost USD +100 millions. DeepSeek apparently was able to train its models at a fraction of GPT-4's training cost.

Monte Carlo Tree Search (MCTS) and Model Distillation

AI companies are optimizing training to reduce costs and increase efficiency. One of the most promising developments is Monte Carlo Tree Search (MCTS), a technique used in AlphaZero and now being explored for large language models. Instead of running one inference path, models now run multiple parallel samples and choose the best output. This technique significantly improves accuracy without increasing cost linearly.

Instead of training massive models, some smaller models are trained by learning from larger ones. DeepSeek reportedly used this method to train competitive models on fewer GPUs. This technique is called model distillation. Some argue that this shortcuts the effort and resources put into training massive models. This can create a perception of unfair advantage, especially when companies leverage open-weight models without bearing the full computational cost of their creation. Moreover, managing distillation is challenging:

  1. Data Bias Transfer: any biases in the original model cascade into the distilled version.
  2. Opaque Learning Process: it is hard to trace and regulate how much knowledge is transferred.
  3. Intellectual Property Questions: who owns the distilled knowledge, especially when built on public models

AI progress seems to be about compute power, memory, infrastructure, and scale. The one controlling the GPUs is likely to control the next era of intelligence! It is for that reason that the US imposes export restrictions on NVIDIA GPUs. Nvidia to launch cheaper Blackwell AI chip for China after US export curbs.

Organizations implementing large-scale AI deployments increasingly require specialized AI Consulting to navigate the complex trade-offs between performance, cost, and environmental impact while ensuring their AI strategies align with both business objectives and sustainability commitments.

The Reality of Enterprise AI Implementation

For organizations considering AI deployment, understanding the environmental implications becomes part of responsible implementation. Companies must balance the benefits of AI with sustainability goals and stakeholder expectations.

The industry is beginning to recognize the need for more sustainable approaches to AI development and deployment. This includes exploring more efficient architectures, optimizing model selection for specific use cases and considering the environmental impact of AI training and inference in strategic planning. Understanding these trade-offs becomes particularly important as organizations implement comprehensive AI strategies that balance innovation with responsibility.

Rainmakers SG helps small and medium businesses design safe and scalable Agentic AI systems that provide immediate ROI!

Want this working in your business?

We help Singapore SMEs and executives turn AI into measurable results.

Book a conversation