If you are weighing an AI assistant for your practice, the environmental question deserves a real answer, not a slogan. This page shows what one AI response consumes, what a year of drafted notes adds up to, and the design choices that make BastionGPT a low-footprint way to bring AI into healthcare. Every figure is sourced and dated so you can check our work.
Last reviewed August 9, 2026. Updated as AI providers publish new measurements.
One AI text response uses about a quarter of a watt-hour of electricity and about five drops of water. Google measured its median prompt at 0.24 watt-hours, 0.26 milliliters of water, and 0.03 grams of CO2 equivalent, about the same energy as nine seconds of television. OpenAI reports an average of 0.34 watt-hours, about what an efficient lightbulb uses in a couple of minutes. A peer-reviewed measurement published in the journal Joule in April 2026 found a median of 0.31 watt-hours.
If the figures you have seen shared among colleagues are much higher, check their date. Before AI companies published measurements, outside researchers had to estimate, and the widely shared 2023 figure of roughly 3 watt-hours per query is about ten times higher than what the 2025 and 2026 measurements show. The gap keeps widening: Google reports that the energy of a median prompt fell 33-fold in the twelve months to May 2025.
Between 0.24 and 0.34 watt-hours for a typical text response, based on the three primary measurements published to date. Two are the providers' own production figures, and one is independently peer reviewed.
Primary published measurements, read August 2026.
| Measurement | Value | Sourceand date |
|---|---|---|
| Median text prompt, energy | 0.24 Wh | Google, measured on production systems, August 2025 |
| Median query, energy | 0.31 Wh | Microsoft Research, peer reviewed in Joule, April 2026 |
| Average query, energy | 0.34 Wh | OpenAI, June 2025 |
| Median text prompt, water | 0.26 mL | Google, August 2025. About five drops |
| Median text prompt, emissions | 0.03 g CO2e | Google, August 2025 |
Efficiency is improving quickly. Google reports that the energy of a median prompt fell 33-fold, and its carbon footprint 44-fold, between May 2024 and May 2025. Stanford's AI Index records a roughly 280-fold fall in the cost of a fixed level of AI capability between late 2022 and late 2024, and serving cost tracks computing effort closely. The same question, asked today, costs a fraction of what it did two years ago.
About 0.26 milliliters for a median prompt, the volume of five drops, measured by Google across its production fleet in August 2025. That figure counts the water its data centers consume on site, mostly for cooling. Estimates that also include the water behind upstream electricity generation run higher, which is why every comparison on this page applies a tenfold safety margin to the AI figures before drawing a conclusion.
A drafted note takes roughly two to five AI responses: about one watt-hour, less energy than a minute of television, and under a quarter of a teaspoon of water. A heavy year of use, forty responses a day across 230 working days, adds up to about 2.8 kilowatt-hours of electricity and about 2.4 liters of on-site water. That is less electricity than a typical home draws in three hours.
The assumptions behind these figures are listed under How we calculated our figures, and each one is set conservatively.
In two measurable ways, yes.
BastionGPT's purpose is to shorten documentation time. An hour less at a laptop and monitor saves roughly 50 to 90 watt-hours at your desk. The AI responses behind that day's notes consume about 10 to 15 watt-hours in the data center. Under the measured 2025 and 2026 per-response figures, the electricity removed at the desk is three to five times the electricity added in the data center. Estimate your own hours with our time savings calculator.
Lifecycle estimates put one printed A4 sheet at 2 to 13 liters of embedded water and 5 to 10 grams of CO2. One ten-page printed intake packet can embed more water than a full year of AI-assisted documentation, even after the tenfold margin on the AI side. Each workflow that moves from print and fax to a drafted digital document saves more than the computing spent drafting it.
Five design choices, each one verifiable in the product or in our infrastructure providers' public commitments.
Each model generation delivers more quality per watt than the one before it. That is what Google's 33-fold single-year improvement measures. BastionGPT adopts new model generations within weeks of release and retires older ones, so your requests run on the most efficient models available rather than last year's. Serving models from more than one provider also lets us adopt whichever lab currently leads on efficiency at a given quality level.
BastionGPT's Autodetect model selection considers every request, and quality comes first: the request goes to a model that can deliver the best available answer for the task. Among models that clear that bar, the more efficient one answers. This matters most for deep-reasoning modes: the peer-reviewed Joule study estimates that long reasoning queries raise energy use by more than an order of magnitude, and independent benchmarks measure the heaviest reasoning models at 10 to 70 times the energy of efficient standard models. BastionGPT escalates to that depth when a task requires it, and only then. Quality decides which model answers. Efficiency breaks the tie.
70 percent of the tokens BastionGPT processes are served from work the platform has already completed rather than computed again. That reuse cuts the electricity and cooling water of our AI processing by more than 60 percent compared with recomputing every request in full. The arithmetic is under How we calculated our figures.
BastionGPT serves your requests on foundation models whose training is shared across hundreds of millions of users worldwide, so the energy cost of training spreads across all of them and our own footprint is the serving of your requests. The same architecture protects your data: patient information is never used to train anyone's models.
BastionGPT runs on Microsoft Azure infrastructure, with AI inference served by Microsoft and Google platforms, entirely in United States, Canada, and Australia data centers. Microsoft has contracted 34 gigawatts of renewable energy across 24 countries, has committed to becoming carbon negative and water positive by 2030, and every Microsoft data center design since August 2024 consumes zero water for cooling, avoiding more than 125 million liters of water per year per site. Google, whose serving infrastructure produced the per-prompt measurements at the top of this page, signed agreements for more than 12 gigawatts of new clean energy in 2025, its largest annual total, and replenished roughly 78 percent of its freshwater consumption through water stewardship projects that same year.
The measured footprint is small. A year of heavy AI-assisted documentation uses about 2.8 kilowatt-hours of electricity, less than a typical home draws in three hours, and about 2.4 liters of on-site water. It also displaces work that costs far more: printed pages, faxed packets, and hours of screen time. Judge any tool by its numbers, and ask for sourced ones.
We can only publish our own numbers, so this page shows them: measured per-response figures, 70 percent processing reuse, efficiency-first model routing, no training on customer data, and named commitments from the providers that run our infrastructure. The section below turns those into six questions you can put to any vendor. Compare answers, not assurances.
No. BastionGPT does not train or fine-tune AI models on customer or patient data, and requests are served on shared foundation models. Choosing not to train also means no training footprint is attributable to your account.
On Microsoft Azure infrastructure, with AI inference on Microsoft and Google platforms, entirely in data centers in the United States, Canada, and Australia, under the environmental commitments described above and the security controls described on our security page.
Most viral figures trace back to estimates published in 2023, before AI companies released measurements, when researchers had to model older and less efficient systems. The measured 2025 and 2026 disclosures come in roughly ten times lower per response, and per-response energy continues to fall. When you see a figure, check its date.
Any tool you evaluate for your practice should be able to answer these six questions. Ours are answered above.
A vendor who cannot answer them has not measured.
If a figure here is wrong or out of date, tell us and we will correct it, the same standard we apply on our comparison methodology page.
BastionGPT is the HIPAA-compliant AI assistant built for healthcare, serving solo therapy practices through large health systems, on the market since 2023 and trusted by more than 10,000 healthcare professionals. Read how we think about building responsibly in our AI principles, or start a 7-day free trial.