AI Cost Audit

Independent audit of LLM and GPU spend

Your company is overpaying for AI. I'll show you by how much — and prove it can be cheaper without losing quality.

A 3-week audit of your LLM and GPU spend. Typical result: a 20–60% lower bill, verified by quality evals on your own data — not by promises.

Book a 30-minute call

No strings attached. If an audit isn't worth it for you, I'll say so.

From read-only access to the report
3 weeks
Typical reduction in the bill
20–60%
Vendor licences I resell
0

Where the money goes

Three signs you're overpaying

Each one is visible in your invoice and your logs, before anyone touches a line of code.

  • 25–50×

    Price spread between model tiers

    A flagship model for everything

    The most expensive model also handles the trivial work — classification, field extraction, rewriting a sentence — that a cheaper one does equally well. Nobody chose this; it is what happens when the default from the first prototype is never revisited.

  • −90%

    Cost of repeated context, once cached

    Zero caching

    Your agents resend the same system prompt, the same tool definitions and the same retrieved documents hundreds of times a day, at full price every time. Prompt caching cuts that by roughly 90%. One field in your logs tells you whether you are paying it.

  • The premium for an answer nobody waits for

    Night jobs at day rates

    Indexing, enrichment, re-scoring, nightly reports — work with no user on the other end — runs through synchronous APIs and pays roughly twice what the batch path costs. The queue is measured in hours; the deadline is tomorrow morning.

Order of magnitude

What this could be worth to you

Two numbers you already know are enough for a first estimate. The audit replaces the estimate with a figure measured on your own traffic.

Model API invoices, GPU inference and inference hosting.

The more repeated context you send, the more there is to recover through caching.

Estimated savings per month

€12,000€32,400

Per year

€144,000€388,800

That is 20–54% of today's bill.

At €60,000 per month with 40% of traffic from agents, estimated savings are €12,000 to €32,400 per month, or €144,000 to €388,800 per year.

An estimate. The audit replaces this with a measured number.

The floor is the cautious scenario: routing and batching alone. The ceiling rises with your agent share, because repeated context is what caching pays back fastest.

Book a 30-minute call

Method

How I work

Four steps, three weeks, one rule: nothing goes into the report that hasn't been measured.

  1. Read-only access

    Billing, inference logs, GPU metrics. No changes to your code, no production access, no deployment rights. If your policy requires it, an exported dataset works just as well.

  2. A frozen baseline

    Unit cost per feature — per request, per document, per user — plus a quality eval on your own data with metrics agreed up front. This is the reference every later claim points back to.

  3. Hypothesis testing

    A cheaper model here, caching there, batch mode, a different instance size. Each hypothesis is measured against the baseline on both axes: cost and quality. If quality drops, it doesn't make the report.

  4. The report

    Savings in money per month, not in percentages. An implementation plan, the risks, and what I would leave alone. Every number reproducible from your own data.

Deliverables

What you get

Documents your engineers can act on and your finance team can check.

  • A cost map with priorities

    Your bill broken down by feature, model and team, split into quick wins and work that needs real engineering time — with the effort of each one estimated.

  • Eval results

    A before-and-after comparison for every recommendation, run on your dataset against the metrics we agreed before the first change was proposed.

  • An implementation plan for your team

    The order of changes, time estimates, rollback points. Written to be executed by your engineers — the knowledge stays in-house when I leave.

  • Savings potential per month

    One number on top and the full arithmetic underneath, including the part I recommend not touching and why.

Objections

The four questions I always get

About

Who you would be working with

I build AI applications and specialize exclusively in inference economics — tokens, caching, model routing, GPU energy measurement. I don't sell implementations or any vendor's licences; an independence clause is in every contract.

For data centres and companies running their own hardware, I also prepare energy reporting under the EU Energy Efficiency Directive: PUE, WUE, ERF and REF, measured rather than estimated.

Book a 30-minute call