Note

Published

2026-08-09

Reading

13 min

Data

Synthetic figures

Headline figure

High

Claim

A cost model that ties out exactly is hiding something

A cost model that reconciles to the invoice to the cent has almost certainly hidden its error inside somebody's cost per customer. Five residual sources are mechanical and none of them cancel, so the honest output of a reconciliation is a named remainder, and a harness that fails loudly when the remainder grows is worth more than the model it checks.

13Reading · minutes
2026-08-09Published

Topics: method · AWS · Azure · GCP

Every figure in this note is invented and labelled synthetic. No client data was used.

Argument

A bill is not a mystery

A cloud bill is not a mystery. It is a dataset with a bad schema: complete, correct, arithmetically consistent, and organised around the wrong primary key. Every row is a meter reading. None of them is a customer, a feature, or a team. The work is not investigation, it is a join, and the only way to know whether the join is right is to add it back up and compare it to the invoice.

That comparison is the single most useful test in this discipline, and it is almost always performed dishonestly. Not fraudulently: dishonestly in the mild sense that the person doing it wants the columns to tie out, and there are four easy ways to make them tie out that all destroy the information the test was supposed to give you.

What reconciliation actually means

One sentence, and it has to be falsifiable: the sum of the modelled rows equals the invoice, per cloud, per billing month, at one stated cost basis, within a stated tolerance.

Every clause in that sentence is load-bearing, and dropping any of them is how a model passes a check it should have failed.

Four reconciliations that look like the real one
What was actually checkedWhy it passes anyway
The total only, not per service Two errors in opposite directions cancel in the total and both survive
A basis that was never named Unblended cost ties to the invoice while every unit cost is wrong by your commitment discount
A trailing 30 days, not a billing month Credits, tax and support post on billing-period boundaries and fall outside the window
A total with the remainder plugged in The residual was distributed across the allocation, so by construction it is zero

The first row is the one worth demonstrating, because it is the least obvious and the most common.

SYNTHETIC
A total that reconciles at 0.2% while one service is out by 11%
ServiceInvoiced
USD
Modelled
USD
Residual
Compute41,90041,730−0.4%
Managed database18,40020,424+11.0%
Object storage6,2006,181−0.3%
Data transfer9,7007,760−20.0%
Inference APIs14,30014,214−0.6%
Total90,50090,309−0.2%
Synthetic example. Not a client. No client data was used. The database line was modelled on the on-demand rate while a reservation covered most of it; the data-transfer line was missing the inter-AZ meter entirely. The two errors are unrelated and they nearly cancel.

A total-only check reports −0.2% and everybody goes home. Per service, two material bugs are visible in ten seconds. This is why reconciliation is reported per service and per billing month rather than as one number, and why “we tie out to the invoice” is not an answer to the question “does your model work”.

One basis, named

The second row of that table, a basis that was never named, deserves its own, because four numbers describe the same hour and they are not close. This table was relocated from /methodology, which now states only which basis it uses.

Four bases, one used
BasisForUsed here
ListSizing a discount, not a marginno
BilledCash on the invoiceno
UnblendedNothing. Wrong by exactly your commitment discountno
Amortized effectiveEvery unit-cost numberyes

Unblended is the trap, because it ties to the invoice while every unit cost derived from it is wrong by your commitment discount. Which provider column carries each basis is worked through in what FOCUS does not normalize.

Five residual sources, and none of them cancel

Once the model is genuinely right, the remainder does not go to zero, because five mechanical effects sit between a usage record and an invoice line. They are the same five every time.

Where the remainder comes from, and what to do about each
SourceMechanismWhat it is not
Sub-hourly boundaries Usage straddling an hour or a month edge lands wholly on one side of the bucket you grouped by. Per-second billing does not remove it; it moves it to whichever timestamp your date_trunc chose. Not a rounding error you can shrink by using more decimal places
Credit and refund timing A credit row can post in a later period than the usage it applies to. GCP delivers credits as separate negative rows, and a support refund lands when it is processed. Not a discount you failed to apply
Provider unit and token rounding Each record is rounded at the provider before you see it. Negligible per request, visible across ten million of them, and always in the same direction. Not noise. It has a sign, so it does not average out
Tax and support These attach to an account or a billing profile, not to a usage line, so there is no resource to allocate them to without a rule you supply. Not allocable from the data alone
AI provider usage lag Usage and cost APIs report hours to a day behind. The most recent bucket looks wrong when it is merely incomplete. Not a discrepancy. Re-run the query tomorrow before believing it

The important structural point: these have signs, not just magnitudes. Token rounding is systematically one way. Credit timing moves cost between adjacent periods rather than creating or destroying it. They do not cancel into zero, and a model that reports exactly zero has therefore either been plugged or is not being compared to the invoice at all.

The plug, and what it destroys

Here is the move. The residual is 2.4%. The table is going into a board pack tomorrow. So the 2.4% is spread across teams, pro rata on their existing allocation, which feels neutral, and the columns now sum to the invoice exactly.

What just happened is that the one number capable of telling a reader whether to trust the other numbers was deleted, and the error was moved into every customer’s cost to serve, proportionally, where nobody will ever find it. If the 2.4% was in fact a missing inter-AZ transfer meter belonging almost entirely to one service, the plug has quietly overcharged four teams to undercharge one.

So the rule is blunt: report the residual, never distribute it. It ships as its own line, with its sources named and sized, in whatever the board reads. A named 2.4% remainder is a model with a known error bar. A silent 0.0% is a model with an unknown one.

Why the tolerance is 3%, and what 3% is good for

Three percent, per cloud, per billing month. The number is mine, not a standard, and its justification is the decision it has to support.

SYNTHETIC
The same ±3% band against three different decisions
DecisionFigure±3% bandDecidable?
Is this plan tier losing money?Cost $4,600 vs revenue $1,800±$138Yes, overwhelmingly
Which of two models is cheaper per resolved task?$0.41 vs $0.19±$0.012Yes
Set a list price at a 4% gross margin floorMargin 4.0%±2.9 ptsNo. Do not price from this
Synthetic example. Not a client. The point is the third row: the same model that decisively answers the first two questions must not be used to answer the third.

Good enough to price with. Not good enough to bill from. If you intend to invoice a customer directly from a derived unit cost, you need a metering pipeline with a different design and a much harder audit trail, and the correct thing to say is that this method is not it.

The harness, and what a failing harness catches

A model is a snapshot. A harness is a scheduled job that recomputes the reconciliation every billing month, compares the residual to the tolerance, and fails loudly when it is exceeded: a red build, a message in the channel that owns the number. It lives in your repository from day one. Every figure it produces ships with the query that produced it.

The difference between a harness and a maintained spreadsheet is not accuracy. It is that a harness can fail. A spreadsheet cannot; it can only be wrong, and it will be wrong silently from the first month that something upstream changes. Most hand-maintained cost models have no check of this kind at all, which is why they are usually approximately right for a quarter and then quietly wrong forever.

What the check is actually for is upstream change, and the list is unglamorous and finite:

Ten things that break a working model, and how the residual shows it
Upstream changeHow it shows up
A new service adopted, untaggedResidual grows month over month, concentrated in one service line
A cost allocation tag key deactivatedAllocated share drops; unattributed line jumps; total still ties
Export schema change or column renameA column reads null and a whole cost class silently becomes zero
A new account or project outside the payerInvoice exceeds the model by exactly that account's spend
A credit type not in the filter listAmortized total drifts from the invoice in one direction only
Mid-window price changeModelled cost is right for part of the month and wrong for the rest
A model added at the gateway, not to the rate cardToken counts reconcile, dollar cost does not
A new region or a machine-family migrationRegion-locked commitment credits stop landing where they used to
Timezone drift at the month boundaryA stable residual of roughly one day's spend, every month
A backfill re-run without dedupeModelled cost exceeds the invoice. The only failure that overstates

Note the last row. Almost every failure mode understates cost, which is the dangerous direction: an understated model makes margin look better than it is and nobody complains about good news. That asymmetry is the argument for the harness rather than for reviewing the model when someone remembers to.

-- The whole harness is one assertion, run per cloud per billing month.
-- It is deliberately per service: a total-only check hides paired errors.
with modelled as (
  select billing_period, service, sum(effective_cost) as modelled_cost
  from cost_model
  group by 1, 2
),
invoiced as (
  select billing_period, service, sum(billed_cost) as invoiced_cost
  from invoice_detail                 -- FOCUS 1.4 Invoice Detail, or the
  group by 1, 2                       -- provider's own invoice export
)
select
  i.billing_period,
  i.service,
  i.invoiced_cost,
  m.modelled_cost,
  m.modelled_cost - i.invoiced_cost                        as residual,
  abs(m.modelled_cost - i.invoiced_cost)
    / nullif(i.invoiced_cost, 0)                           as residual_share,
  case
    when abs(m.modelled_cost - i.invoiced_cost)
           / nullif(i.invoiced_cost, 0) > 0.03 then 'FAIL'
    else 'PASS'
  end                                                      as verdict
from invoiced i
full outer join modelled m
  on  i.billing_period = m.billing_period
 and  i.service        = m.service
order by residual_share desc;

Two details in that query are the ones people leave out. It is a full outer join, so a service present in the invoice and absent from the model (the new untagged service, the account outside the payer) appears as a row rather than vanishing. And the verdict is computed in the query rather than eyeballed, because a check that requires a human to read it is a check that stops happening in month four.

FOCUS 1.4 made this materially easier: the Invoice Detail and Billing Period datasets, published June 2026, mean the right-hand side of the join can be a specified dataset rather than each vendor’s own schema.1

What this method cannot tell you

Four limits, stated here rather than discovered in a readout.

Whether a fix is worth doing. A 40,000-a-year saving can cost two engineers a quarter. The model prices the saving; it has no view on the cost of the work, and the ranked waste list is ordered by dollars, not by whether you should care.

Idle capacity from deliberate headroom. In a billing export they are identical. Capacity held for a launch, a failover region, or a seasonal peak looks exactly like waste, and only you know which it is. Every waste line therefore has to be claimed by a human owner before it counts as waste, which is what the ninety minutes with an engineer in the teardown schedule is for.

The cost of an outage. Rightsizing to an observed p99 with no headroom is a reliability decision dressed as a cost decision. This method cannot price the incident it might cause, so it does not get to make that call.

Whether the number will change anything. A merged pull request still needs your approval, and a cost model with no owner on the other side of it produces a very well-reconciled report that nothing happens to.

What would change our mind

If providers shipped a first-class invoice-reconciliation surface (a specified dataset per billing period, per service, at a named basis, with the allocation dimension already in it) then reconciliation becomes a configuration check and the interesting work moves entirely to the join and the split rules. FOCUS 1.4 is a real step in that direction, and the honest reading is that the arithmetic in this note gets easier every year.

What does not get easier: deciding which residual you are willing to publish, and who owns the line that has no owner.

Two residuals, itemized

Synthetic

Relocated from the Spend Teardown page. One estate, 90 days, USD.

LineFigure
Invoice total, three months1,284,600
Model total, amortized effective cost1,262,110
Residual · credits applied outside the window14,300
Residual · sub-hourly boundary effects5,190
Residual · unexplained3,000
Residual as share of invoice1.75%
Synthetic example. Not a client. No client data was used. The residual is reported line by line, never distributed across teams, and 1.75% is inside the stated 3% tolerance.
Synthetic

Relocated from /methodology: the 477.00 USD residual on the one-month, three-tenant derivation published there.

SourceAmount USDOf invoice
Credit and refund application timing165.000.20%
Sub-hourly boundary effects at window edges120.000.15%
Tax and support, no usage line102.000.12%
Provider token rounding, 10.0M requests90.000.11%
Reported residual477.000.58%
Synthetic example. Not a client. No client data was used. Four named sources, one reported line, nothing plugged.

Sources

  1. FOCUS specification, native provider adoption, and the 1.4 Invoice Detail and Billing Period datasets (June 2026): focus.finops.org. High confidence. The field-level consequences are worked through in what FOCUS does not normalize.

Sources are listed at the foot of this note with their confidence stated. Medium and low confidence figures say so in the copy.

Every query here runs against a read-only role you create, scope and revoke. Access policy.

Close

The rung this argument belongs to

Instrumentation Build · Fixed fee per stage

Check the arithmetic yourself

The full derivations, with the queries written out so they run under your own read-only credentials, in your own console.