Blog / AI

I Optimized AWS Bills for Years. AI Bills Have the Same Disease.

By Saurav Sharma||5 min read

I have spent years cleaning up AWS bills, including six years at Amazon before going independent. Now I spend my days looking at AI bills, and I keep having the same feeling: I have seen this disease before.

Different meters. Same illness.

Cloud waste regenerates. AI waste will too.

Recently I cut my monthly AWS bill by 22% in one conversation with Claude Code. Under an hour. I never opened the console.

In raw dollars it was not a huge number for my account. But 22% is 22%. Run that percentage against your own bill.

Here is the part that matters: this was not my first cleanup. I did a bigger pass a few months back. This 22% was a mix of waste that crept back in since then, and older stuff that pass never caught. Some of it dated back to 2019.

Cloud waste is almost never one big line item. It is a pile of small ghosts nobody remembers creating:

  • Toll-free numbers stuck in "pending" since 2021, billing every month without sending a single text
  • 90+ GB of screenshots and logs from a synthetic canary I set up in 2023 and forgot
  • An RDS snapshot for a database I deleted years ago
  • ECR images tagged "test" in 2019, still sitting there seven years later
  • S3 buckets full of stale logs
  • A WAF attached to a dead project

Half of it was hiding in linked child accounts I almost never log into.

Your AI bill is accumulating the same ghosts right now. Agents that retry in loops. Workflows someone set up for a demo and never turned off. Premium models answering questions a cheap model could handle. Duplicate subscriptions across teams. Nobody created any of this on purpose. That is exactly why it grows.

The disease is not technical

Yes, the technical levers matter: caching, model routing, batching, prompt design, evals, choosing the right model for the task. I use all of them.

But after years in cloud cost work, I can tell you where the money actually leaks: visibility, ownership, budgets, defaults, and review loops. AI cost optimization depends on the same things. Good intentions are not enough. Teams need mechanisms that make the right behavior easy to repeat.

The same fixes transfer almost one-to-one:

Build cost awareness, not fear. You do not want people scared to use AI. You want them to know where AI spend creates value and where it does not.

Train people on how AI costs actually work. Most waste is not intentional. People just do not know that long context, retries, agent loops, huge outputs, and premium models add up fast. The cloud version of this was engineers who did not know an idle GPU instance billed all weekend.

Make premium AI a conscious choice. Some use cases deserve the best model available. Most do not. Leaders should make that distinction explicit so teams are not guessing.

Make the approved path easy. If the official AI setup is slow or locked down, people will bring their own tools. That is how shadow IT happened in cloud, and shadow AI is worse: duplicate subscriptions, hidden usage, spend nobody can measure.

Keep humans accountable. Every recurring AI workflow should have an owner who knows roughly what it costs, checks whether it is still useful, and has a fallback when it breaks. In cloud we called this tagging and ownership. The names change. The mechanism does not.

What agents change, and what they don't

That 22% AWS session was FinOps and cloud engineering at the same time. Claude found the waste, and it also made the engineering calls. It checked what each resource was wired to before touching it. One bucket looked like an old dump I could safely delete. It turned out to be my live CloudTrail destination, so it left it alone.

The manual grind is collapsing. A week of clicking through accounts is now a prompt and a review. But someone still has to know what is production and what breaks if it quietly disappears. The agent runs the commands. A human still owns the call.

That is the honest shape of this work now, on both the cloud bill and the AI bill.

Why most AI cost programs will fail anyway

Most enterprise AI projects fail. MIT's NANDA report, "The GenAI Divide", found 95% of generative AI pilots delivered no measurable P&L impact, and S&P Global's 2025 survey found 42% of companies abandoned most of their AI initiatives, up from 17% the year before. From what I have seen, it starts the same way every time: leadership says "go do some AI," a team spins up a pilot, no workflow changes, the project dies.

Cost programs fail the same way. A dashboard gets built, nobody owns it, the waste regenerates.

What works is smaller and more boring: take an existing workflow, make it slightly faster or slightly better with AI, give teams room to experiment without demanding ROI in 90 days, and put an owner on every recurring cost. McKinsey's research argues for spending on people at several times the rate you spend on the technology itself. That ratio was true for cloud adoption too.

When did you last go ghost hunting?

If nobody in your company can answer "what did AI cost us last month and was it worth it," you have the disease. It is treatable, but not with a one-time cleanup, because the waste comes back.

This is what I do as a fractional AI advisor: AI cost audits, architecture second opinions, and training across both the technical and human side of AI spend. If your team needs help, book a call at cloudyeti.io/chat. Sorting out your own AI stack and costs as an individual? Grab a 1:1 at cloudyeti.io/meet.

Book a call