ServicesCost optimisation

Cost optimisation that starts with a measurement.

Inference cost per million tokens, model routing, caching, right-sizing and reservations, applied to AI workloads and to the cloud estate around them. Fixed-fee assessment; savings reported against a baseline you agreed.

Who it is forFinance and engineering leaders whose AI or cloud bill grew faster than the business behind it, and teams that migrated recently and have not yet optimised.

The problem

AI spend is unusual: it scales with usage, it hides inside a few line items, and the levers are technical. Switching a model, routing simple requests to a smaller one, caching repeated prompts or moving inference to different silicon can each move the bill by double-digit percentages. Nobody in finance can pull those levers, and most engineering teams have not been asked to.

We measure first, so that every change is reported against a baseline rather than a feeling. Then we work through the levers in order of return, with your team, inside your accounts.

What we deliver

  • Cost baseline: per workload, per model, per million tokens, with the top ten drivers
  • Model routing and fallback design: the right model per request class
  • Prompt and response caching, batching, context trimming
  • Inference placement: Bedrock, Inferentia, Trainium, GPU, by workload
  • Cloud estate review: right-sizing, reservations and savings plans, idle resources
  • Cost dashboard and monthly governance cadence handed to your team

How we work

Baseline

Two weeks of measurement. Where the money goes, by workload and by driver, in a report your CFO can read.

Levers

Each lever quantified: expected saving, effort, risk. You choose the order.

Apply

Changes made in your tenancy with quality checked against your evaluation suite, so savings never come at the cost of answers.

Govern

Dashboard, budget alerts and a monthly review handed to your team.

What backs it

The sample benchmark report on our home page is the shape of what you receive: cost, throughput, latency and quality side by side. We report savings only against the agreed baseline, and we write down the ones we could not confirm.

Common questions

How do you charge?

The assessment is a fixed fee. Implementation is fixed scope per lever. We do not take a percentage of savings, because that rewards the wrong behaviour.

Will quality suffer?

Not without you knowing. Every change is checked against your evaluation suite before it is kept, and the report shows the quality numbers next to the cost numbers.

Is this only for AI workloads?

No. Because we also migrate general workloads to AWS, the estate review covers compute, storage and data services around the AI systems.

Talk to an engineer about this

Thirty minutes, no slides. Bring the workload and we will tell you what we would do and what it would cost.

Talk to us[email protected]