digital-humans.org

The honest answer is that ChatGPT is neither an environmental catastrophe nor a free lunch. Training a frontier model consumes an enormous, concentrated burst of energy – GPT-4-class training runs have been estimated in the tens of gigawatt-hours – while each individual query you send costs a small fraction of a watt-hour. The environmental question that actually matters is not any single prompt but the aggregate: billions of inferences a day, running on data centers that draw power and water at industrial scale, in a sector the International Energy Agency projects will roughly double its electricity demand before the decade is out. The nuance lives in that gap between the tiny per-query cost and the vast total.

This piece separates the two cost structures that get carelessly blended in most coverage – the one-time training bill and the recurring inference bill – grounds each in peer-reviewed and vendor-published figures where they exist, and sets AI's footprint against streaming, crypto and aviation so the numbers mean something.

The honest energy answer

Two distinct energy events sit behind every ChatGPT response, and they behave nothing alike.

The first is training: the compute-heavy process of learning model weights from data. It happens once per model version, runs for weeks or months across thousands of accelerators, and consumes a large, fixed quantity of energy that is then amortized over the model's entire operational life.

The second is inference: the energy spent generating a reply each time someone uses the model. Per query this is small – comparable to running a household appliance for a few seconds – but it recurs with every interaction, and ChatGPT serves an enormous volume of interactions.

The practical consequence is that as a model matures, inference comes to dominate its lifetime energy footprint. A single training run is expensive, but a model answering hundreds of millions of prompts a day will, over months, spend far more energy serving users than it ever spent learning. So "is ChatGPT bad for the environment" resolves into a scaling question: the individual cost is trivial, the aggregate is not, and the aggregate grows with adoption.

Training energy costs

Training a large language model is the single most legible piece of AI's carbon story, because it is bounded in time and increasingly documented. Meta's model cards for the Llama 3 family, published on the Meta AI Llama site, report training compute in GPU-hours and give estimated emissions figures – a rare instance of a lab disclosing the number directly rather than leaving analysts to reconstruct it.

OpenAI has not published comparable figures for GPT-4, so all public estimates are external reconstructions built from assumed cluster size, training duration and hardware efficiency. Treat any specific GPT-4 training-energy number you encounter as an educated estimate rather than a disclosed fact, and check the assumptions behind it. As of this writing, the most-cited figures put a frontier training run in the tens of gigawatt-hours, but the error bars are wide.

DeepSeek R1 and its predecessor DeepSeek-V3 complicated the story in early 2025 by reporting training costs far below the assumed frontier norm, claiming strong reasoning performance from a training budget a fraction of what Western labs were believed to spend. The reported figures have been contested, and the accounting – whether it includes prior research runs, failed experiments and data preparation – is not fully transparent. Even so, the episode made a real point: architectural and training-efficiency choices can move the energy cost of a capable model by an order of magnitude, and headline training numbers are only as trustworthy as the methodology beneath them.

Three caveats apply to every training figure. First, energy is not emissions: a run powered by hydroelectricity in Quebec and an identical run on a coal-heavy grid produce wildly different carbon totals, so the local grid mix matters as much as the kilowatt-hours. Second, published numbers rarely include the embodied carbon of manufacturing the chips themselves. Third, training is a moving target – each new model generation tends to be larger, but efficiency gains partly offset the growth, so the trend line is neither flatly rising nor falling.

Inference energy costs

The per-query cost is where public fear and physical reality diverge most sharply. The most rigorous open work here comes from Hugging Face, particularly the research led by Sasha Luccioni and the Hugging Face AI Energy Score project, which benchmarks the energy consumption of models on standardized inference tasks and publishes the methodology so results can be reproduced.

The headline finding across this body of work is that a single text-generation query from a large model costs on the order of a few watt-hours or less – roughly the energy of running an LED bulb for a short spell, or a small fraction of what boiling a kettle consumes. The exact figure depends heavily on model size, output length, hardware and batching efficiency, which is precisely why the Energy Score approach measures under controlled conditions rather than quoting a universal constant. Image generation costs meaningfully more than text; long, reasoning-heavy responses cost more than short ones; and models running in efficient, well-batched production settings cost far less per query than the same model run naively.

The relevant subtlety is task-appropriateness. Luccioni's research has repeatedly shown that using a giant general-purpose model for a task a small specialized model could handle wastes energy by a large multiple. Sending a frontier reasoning model a query that a compact model would answer just as well is the inference equivalent of commuting in a lorry. This is where individual "chatgpt issues" around energy become tractable: the lever is not abstaining from AI but matching model size to task difficulty.

Aggregate inference is the number that should worry anyone. A few watt-hours per query is negligible alone. Multiplied across the hundreds of millions of daily interactions that a service at ChatGPT's scale handles – and the retries and rate-limit backoffs documented in our guide to 429 Too Many Requests errors hint at that volume – it becomes a serious, continuous industrial load.

Data center water use

Electricity is only half the resource story. Data centers running AI workloads reject heat, and much of that heat is carried away by water, either through evaporative cooling towers or, indirectly, through the water consumed generating the electricity in the first place.

The water intensity of a given data center varies enormously by climate, cooling design and local grid. A facility using evaporative cooling in a hot, dry region can consume substantial volumes of fresh water; one in a cool climate using air-side economization or closed-loop cooling may consume very little on-site. The published sustainability reports from major hyperscalers – Microsoft, Google and Amazon Web Services – now disclose water-usage effectiveness figures, though the granularity and comparability of these disclosures remain uneven.

Because water deserves its own treatment, we cover the cooling-water accounting, the on-site-versus-off-site distinction and the regional-stress question in detail in our companion analysis of how much water ChatGPT uses. The short version: water can be the binding environmental constraint in stressed regions even where energy is comparatively clean, and siting decisions matter as much as any efficiency gain.

AI vs other tech

Context turns a scary number into a proportionate one. The IEA's work on data centers and AI, published on the International Energy Agency site, is the most authoritative attempt to size the sector, and it consistently places total data-center electricity demand – AI included – at a low single-digit percentage of global electricity as of the mid-2020s, with a projected steep rise.

Set against everyday digital behaviour, individual AI use is modest. Streaming an hour of high-definition video draws on data-center, network and device energy that comfortably exceeds many dozens of text queries. Against cryptocurrency, the contrast is starker still: proof-of-work Bitcoin mining has for years consumed electricity on the scale of a mid-sized country, with no per-transaction utility gain from the energy spent, whereas AI inference at least produces a variable, useful output per joule. Against aviation, a single long-haul flight emits carbon that would take an individual an implausible number of chatbot queries to match.

None of this excuses AI's footprint. It reframes it. The problem is not that any one query is wasteful – it is that the sector is scaling faster than the grid is decarbonizing, and rapid load growth concentrated in specific regions can strain local infrastructure regardless of the global average. That local strain, and the competition for grid capacity, is the genuine environmental controversy, not the cost of your afternoon's prompts.

The mitigation story

The efficiency trend is real and works in several directions at once. Hardware improves generation over generation, delivering more inference per watt. Model architecture improves: techniques like mixture-of-experts activate only part of a model per query, and quantization shrinks the compute needed to run a given model. Model size itself is being deliberately reduced, with capable small models now doing work that required far larger models a year or two earlier.

Procurement is shifting too. The major hyperscalers have signed large renewable power-purchase agreements and, more recently, nuclear and geothermal deals, aiming to match or genuinely supply their consumption with low-carbon electricity – though "matched annually" is not the same as "clean every hour", and the honest accounting distinguishes the two.

The most structurally interesting shift is toward on-device inference. Running compact models locally on phones and laptops removes the data-center round trip entirely for many tasks, trading centralized grid load for the far smaller draw of the device already in your hand. We explore that architectural move and its trade-offs in our primer on what edge AI is and why it matters. Edge inference will not absorb frontier reasoning workloads, but it can offload a large share of routine queries.

Is ChatGPT bad for the environment?

ChatGPT has a real but often exaggerated environmental footprint. A single ChatGPT query consumes a small amount of energy – on the order of a few watt-hours or less, according to Hugging Face's measured research – while training the underlying model consumes a large one-time quantity of energy. The environmental concern is the aggregate load of billions of daily queries running in data centers, not any individual prompt.

Why is ChatGPT bad for the environment, when people raise concerns?

The concerns about ChatGPT's environmental impact center on scale and siting rather than per-use cost. Data centers running AI workloads draw large, continuous amounts of electricity and cooling water, concentrated in specific regions where they can strain local grids and water supplies, and much of that electricity is not yet from low-carbon sources. The worry is that AI demand is growing faster than grids are decarbonizing.

How much energy does one ChatGPT query use?

One ChatGPT query uses roughly a few watt-hours of energy or less, based on controlled inference measurements from the Hugging Face AI Energy Score project, though the exact figure depends on model size, response length and hardware. Longer, reasoning-heavy responses and image generation cost considerably more than short text replies.

What is likely to change

The trajectory to 2027-2030 is a race between two curves. Demand rises as AI adoption spreads, agentic systems make multiple model calls per task, and reasoning models spend more compute per answer. Efficiency and decarbonization push the other way through better silicon, smaller models, on-device execution and cleaner grids. The IEA's projections point to data-center electricity demand rising sharply over this window, with AI the fastest-growing component, even as per-query costs continue to fall.

Where this lands depends on choices that are governance questions as much as engineering ones: whether labs disclose training and inference energy transparently, whether load growth is met with new clean generation or existing fossil capacity, and whether the industry routes routine work to right-sized models instead of defaulting to the largest available. The environmental verdict on ChatGPT is not fixed – it is being written by procurement contracts, model-architecture decisions and regulatory disclosure rules being negotiated now. That places it squarely in the ethics conversation about how these systems are built and run, alongside questions of labor, safety and the broader economic impact of AI.

For anyone weighing the question in practice, the actionable conclusion is unglamorous: the footprint of your own use is small, the honest levers are collective, and the numbers here are time-sensitive. Verify current figures against the IEA reports, the Hugging Face AI Energy Score and the hyperscalers' own sustainability disclosures before you cite them, because in this sector the ground moves quarter by quarter.