Note
Published
2026-08-09
Reading
13 min
Data
Synthetic figures
Headline figure
High
Claim
A cost model that ties out exactly is hiding something
A cost model that reconciles to the invoice to the cent has almost certainly hidden its error inside somebody's cost per customer. Five residual sources are mechanical and none of them cancel, so the honest output of a reconciliation is a named remainder, and a harness that fails loudly when the remainder grows is worth more than the model it checks.
Topics: method · AWS · Azure · GCP
Every figure in this note is invented and labelled synthetic. No client data was used.
Argument
A bill is not a mystery
A cloud bill is not a mystery. It is a dataset with a bad schema: complete, correct, arithmetically consistent, and organised around the wrong primary key. Every row is a meter reading. None of them is a customer, a feature, or a team. The work is not investigation, it is a join, and the only way to know whether the join is right is to add it back up and compare it to the invoice.
That comparison is the single most useful test in this discipline, and it is almost always performed dishonestly. Not fraudulently: dishonestly in the mild sense that the person doing it wants the columns to tie out, and there are four easy ways to make them tie out that all destroy the information the test was supposed to give you.
What reconciliation actually means
One sentence, and it has to be falsifiable: the sum of the modelled rows equals the invoice, per cloud, per billing month, at one stated cost basis, within a stated tolerance.
Every clause in that sentence is load-bearing, and dropping any of them is how a model passes a check it should have failed.
| What was actually checked | Why it passes anyway |
|---|---|
| The total only, not per service | Two errors in opposite directions cancel in the total and both survive |
| A basis that was never named | Unblended cost ties to the invoice while every unit cost is wrong by your commitment discount |
| A trailing 30 days, not a billing month | Credits, tax and support post on billing-period boundaries and fall outside the window |
| A total with the remainder plugged in | The residual was distributed across the allocation, so by construction it is zero |
The first row is the one worth demonstrating, because it is the least obvious and the most common.
| Service | Invoiced USD | Modelled USD | Residual |
|---|---|---|---|
| Compute | 41,900 | 41,730 | −0.4% |
| Managed database | 18,400 | 20,424 | +11.0% |
| Object storage | 6,200 | 6,181 | −0.3% |
| Data transfer | 9,700 | 7,760 | −20.0% |
| Inference APIs | 14,300 | 14,214 | −0.6% |
| Total | 90,500 | 90,309 | −0.2% |
A total-only check reports −0.2% and everybody goes home. Per service, two material bugs are visible in ten seconds. This is why reconciliation is reported per service and per billing month rather than as one number, and why “we tie out to the invoice” is not an answer to the question “does your model work”.
One basis, named
The second row of that table, a basis that was never named, deserves its own, because four numbers describe the same hour and they are not close. This table was relocated from /methodology, which now states only which basis it uses.
| Basis | For | Used here |
|---|---|---|
| List | Sizing a discount, not a margin | no |
| Billed | Cash on the invoice | no |
| Unblended | Nothing. Wrong by exactly your commitment discount | no |
| Amortized effective | Every unit-cost number | yes |
Unblended is the trap, because it ties to the invoice while every unit cost derived from it is wrong by your commitment discount. Which provider column carries each basis is worked through in what FOCUS does not normalize.
Five residual sources, and none of them cancel
Once the model is genuinely right, the remainder does not go to zero, because five mechanical effects sit between a usage record and an invoice line. They are the same five every time.
| Source | Mechanism | What it is not |
|---|---|---|
| Sub-hourly boundaries | Usage straddling an hour or a month edge lands wholly on one side of the bucket you grouped by. Per-second billing does not remove it; it moves it to whichever timestamp your date_trunc chose. |
Not a rounding error you can shrink by using more decimal places |
| Credit and refund timing | A credit row can post in a later period than the usage it applies to. GCP delivers credits as separate negative rows, and a support refund lands when it is processed. | Not a discount you failed to apply |
| Provider unit and token rounding | Each record is rounded at the provider before you see it. Negligible per request, visible across ten million of them, and always in the same direction. | Not noise. It has a sign, so it does not average out |
| Tax and support | These attach to an account or a billing profile, not to a usage line, so there is no resource to allocate them to without a rule you supply. | Not allocable from the data alone |
| AI provider usage lag | Usage and cost APIs report hours to a day behind. The most recent bucket looks wrong when it is merely incomplete. | Not a discrepancy. Re-run the query tomorrow before believing it |
The important structural point: these have signs, not just magnitudes. Token rounding is systematically one way. Credit timing moves cost between adjacent periods rather than creating or destroying it. They do not cancel into zero, and a model that reports exactly zero has therefore either been plugged or is not being compared to the invoice at all.
The plug, and what it destroys
Here is the move. The residual is 2.4%. The table is going into a board pack tomorrow. So the 2.4% is spread across teams, pro rata on their existing allocation, which feels neutral, and the columns now sum to the invoice exactly.
What just happened is that the one number capable of telling a reader whether to trust the other numbers was deleted, and the error was moved into every customer’s cost to serve, proportionally, where nobody will ever find it. If the 2.4% was in fact a missing inter-AZ transfer meter belonging almost entirely to one service, the plug has quietly overcharged four teams to undercharge one.
So the rule is blunt: report the residual, never distribute it. It ships as its own line, with its sources named and sized, in whatever the board reads. A named 2.4% remainder is a model with a known error bar. A silent 0.0% is a model with an unknown one.
Why the tolerance is 3%, and what 3% is good for
Three percent, per cloud, per billing month. The number is mine, not a standard, and its justification is the decision it has to support.
| Decision | Figure | ±3% band | Decidable? |
|---|---|---|---|
| Is this plan tier losing money? | Cost $4,600 vs revenue $1,800 | ±$138 | Yes, overwhelmingly |
| Which of two models is cheaper per resolved task? | $0.41 vs $0.19 | ±$0.012 | Yes |
| Set a list price at a 4% gross margin floor | Margin 4.0% | ±2.9 pts | No. Do not price from this |
Good enough to price with. Not good enough to bill from. If you intend to invoice a customer directly from a derived unit cost, you need a metering pipeline with a different design and a much harder audit trail, and the correct thing to say is that this method is not it.
The harness, and what a failing harness catches
A model is a snapshot. A harness is a scheduled job that recomputes the reconciliation every billing month, compares the residual to the tolerance, and fails loudly when it is exceeded: a red build, a message in the channel that owns the number. It lives in your repository from day one. Every figure it produces ships with the query that produced it.
The difference between a harness and a maintained spreadsheet is not accuracy. It is that a harness can fail. A spreadsheet cannot; it can only be wrong, and it will be wrong silently from the first month that something upstream changes. Most hand-maintained cost models have no check of this kind at all, which is why they are usually approximately right for a quarter and then quietly wrong forever.
What the check is actually for is upstream change, and the list is unglamorous and finite:
| Upstream change | How it shows up |
|---|---|
| A new service adopted, untagged | Residual grows month over month, concentrated in one service line |
| A cost allocation tag key deactivated | Allocated share drops; unattributed line jumps; total still ties |
| Export schema change or column rename | A column reads null and a whole cost class silently becomes zero |
| A new account or project outside the payer | Invoice exceeds the model by exactly that account's spend |
| A credit type not in the filter list | Amortized total drifts from the invoice in one direction only |
| Mid-window price change | Modelled cost is right for part of the month and wrong for the rest |
| A model added at the gateway, not to the rate card | Token counts reconcile, dollar cost does not |
| A new region or a machine-family migration | Region-locked commitment credits stop landing where they used to |
| Timezone drift at the month boundary | A stable residual of roughly one day's spend, every month |
| A backfill re-run without dedupe | Modelled cost exceeds the invoice. The only failure that overstates |
Note the last row. Almost every failure mode understates cost, which is the dangerous direction: an understated model makes margin look better than it is and nobody complains about good news. That asymmetry is the argument for the harness rather than for reviewing the model when someone remembers to.
-- The whole harness is one assertion, run per cloud per billing month.
-- It is deliberately per service: a total-only check hides paired errors.
with modelled as (
select billing_period, service, sum(effective_cost) as modelled_cost
from cost_model
group by 1, 2
),
invoiced as (
select billing_period, service, sum(billed_cost) as invoiced_cost
from invoice_detail -- FOCUS 1.4 Invoice Detail, or the
group by 1, 2 -- provider's own invoice export
)
select
i.billing_period,
i.service,
i.invoiced_cost,
m.modelled_cost,
m.modelled_cost - i.invoiced_cost as residual,
abs(m.modelled_cost - i.invoiced_cost)
/ nullif(i.invoiced_cost, 0) as residual_share,
case
when abs(m.modelled_cost - i.invoiced_cost)
/ nullif(i.invoiced_cost, 0) > 0.03 then 'FAIL'
else 'PASS'
end as verdict
from invoiced i
full outer join modelled m
on i.billing_period = m.billing_period
and i.service = m.service
order by residual_share desc;
Two details in that query are the ones people leave out. It is a full outer join, so a service present in the invoice and absent from the model (the new
untagged service, the account outside the payer) appears as a row rather than
vanishing. And the verdict is computed in the query rather than eyeballed,
because a check that requires a human to read it is a check that stops
happening in month four.
FOCUS 1.4 made this materially easier: the Invoice Detail and Billing Period datasets, published June 2026, mean the right-hand side of the join can be a specified dataset rather than each vendor’s own schema.1
What this method cannot tell you
Four limits, stated here rather than discovered in a readout.
Whether a fix is worth doing. A 40,000-a-year saving can cost two engineers a quarter. The model prices the saving; it has no view on the cost of the work, and the ranked waste list is ordered by dollars, not by whether you should care.
Idle capacity from deliberate headroom. In a billing export they are identical. Capacity held for a launch, a failover region, or a seasonal peak looks exactly like waste, and only you know which it is. Every waste line therefore has to be claimed by a human owner before it counts as waste, which is what the ninety minutes with an engineer in the teardown schedule is for.
The cost of an outage. Rightsizing to an observed p99 with no headroom is a reliability decision dressed as a cost decision. This method cannot price the incident it might cause, so it does not get to make that call.
Whether the number will change anything. A merged pull request still needs your approval, and a cost model with no owner on the other side of it produces a very well-reconciled report that nothing happens to.
What would change our mind
If providers shipped a first-class invoice-reconciliation surface (a specified dataset per billing period, per service, at a named basis, with the allocation dimension already in it) then reconciliation becomes a configuration check and the interesting work moves entirely to the join and the split rules. FOCUS 1.4 is a real step in that direction, and the honest reading is that the arithmetic in this note gets easier every year.
What does not get easier: deciding which residual you are willing to publish, and who owns the line that has no owner.
Two residuals, itemized
Relocated from the Spend Teardown page. One estate, 90 days, USD.
| Line | Figure |
|---|---|
| Invoice total, three months | 1,284,600 |
| Model total, amortized effective cost | 1,262,110 |
| Residual · credits applied outside the window | 14,300 |
| Residual · sub-hourly boundary effects | 5,190 |
| Residual · unexplained | 3,000 |
| Residual as share of invoice | 1.75% |
Relocated from /methodology: the 477.00 USD residual on the one-month, three-tenant derivation published there.
| Source | Amount USD | Of invoice |
|---|---|---|
| Credit and refund application timing | 165.00 | 0.20% |
| Sub-hourly boundary effects at window edges | 120.00 | 0.15% |
| Tax and support, no usage line | 102.00 | 0.12% |
| Provider token rounding, 10.0M requests | 90.00 | 0.11% |
| Reported residual | 477.00 | 0.58% |
Sources
- FOCUS specification, native provider adoption, and the 1.4 Invoice Detail and Billing Period datasets (June 2026): focus.finops.org. High confidence. The field-level consequences are worked through in what FOCUS does not normalize.
Sources are listed at the foot of this note with their confidence stated. Medium and low confidence figures say so in the copy.
Every query here runs against a read-only role you create, scope and revoke. Access policy.
Close
More notes
The rung this argument belongs to
Instrumentation Build · Fixed fee per stage
Check the arithmetic yourself
The full derivations, with the queries written out so they run under your own read-only credentials, in your own console.