Cloud and AI bill reduction
Wasted cloud spend hit 29 percent this year, the first rise in five years, and AI bills compound on top of it. I’m Saša Tomić. I read invoices the way I read systems: line by line, until I can account for every number. Then I make the cuts, and the next invoice is the receipt. Efficiency has been the job since my storage firmware days at IBM.
Who this is for
You spend seven figures a year across cloud providers, observability tools, and AI APIs. Your platform team owns the features; no one owns what they cost to run. The observability bill grows faster than your traffic, and the LLM invoices doubled in six months. Finance has started asking what the company gets for it. If that is your quarter, this work pays for itself out of the waste it finds.
What you get
Three builds, and each one ends with a smaller number.
1. The autopsy, in two weeks
- A line-item read of your cloud, telemetry, and AI invoices against actual usage.
- A ranked kill list: idle resources, zombie storage, overprovisioned instances, orphaned snapshots, wasted log ingestion, redundant model calls.
- Each item priced: what it costs you per month and what it takes to kill.
- I execute quick wins during the audit, so the first savings land before the report does.
2. The execution
- Rightsizing and a commitments ladder, so you stop paying on-demand prices for steady load.
- A telemetry diet: sampling, cardinality cuts, and retention rules that keep your incident debugging sharp while the bill shrinks.
- An LLM diet: model routing, prompt caching, and per-team cost attribution, so each team sees what its AI features cost.
- I roll out each change behind a rollback path, because a savings project that causes an outage has failed.
3. The proof
- A before/after report anchored to your actual invoices.
- A per-team cost dashboard, so the number stays visible after I leave.
- A success-fee option: my fee scales with the savings your invoices show in the first year.
- If the invoices do not drop, you do not pay the success fee.
Why me
At DFINITY I later led Mainnet Reliability for a production network of 1,400 nodes, where idle capacity came out of a real budget; I treated waste the way I treated outages, as something you find and kill. At IBM before that, I built the telemetry pipeline behind Storage Insights, which is where I learned how fast observability data compounds when no one prices it. More on my about page.
How pricing works
The audit is a fixed price, quoted after you share one recent invoice. Execution runs at fixed scope, or on the success fee if you want my pay tied to the delta on the invoice. Either way, you see the kill list and its numbers before you commit to any of it.
Questions I hear
“We already have FinOps dashboards.” Dashboards rank the waste; someone still has to delete the resources, renegotiate the contract, and rewire the pipeline. I do the second half, which is the half that saves money.
“Won’t cutting telemetry blind us during incidents?” The diet keeps the traces your on-call opens at 3am and drops the ones no one has queried in a year. I tune sampling and cardinality rules against your real queries, then verify them against your next incident.
“Is the success fee real?” Yes. We agree on the baseline invoice and the measurement before work starts. The savings come off invoices you can read; my fee is a share of the first-year delta.
“Who sees our data?” An NDA comes first. Invoices and usage data stay in your environment; the report covers cost figures only.
Next step
One LinkedIn message is enough: your rough monthly cloud and AI spend, and which bill annoys you most. You get a written note on where a first audit would dig, within two days, useful even if we never work together.
Sources
- Flexera 2026 State of the Cloud Report: wasted cloud spend rose to 29 percent, the first increase in five years, and 76 percent of large enterprises spend more than $5 million monthly on public cloud.
- Harness, FinOps in Focus 2025: an estimated $44.5 billion of enterprise cloud infrastructure spend was wasted in 2025, about 21 percent of the total.
- Menlo Ventures, 2025 Mid-Year LLM Market Update: enterprise LLM spend more than doubled in six months, from $3.5 billion in November 2024 to $8.4 billion by mid-2025.
- DoiT, Datadog pricing explained: most teams find their actual bill two to three times higher than their initial estimate once logs, APM, and custom metrics compound.