Note
Published
2026-07-09
Reading
15 min
Data
Synthetic figures
Headline figure
MEDIUM CONFIDENCE
Claim
The node is the billing unit, the namespace is the ownership unit
Kubernetes cost allocation fails for a structural reason, not a tooling one: the cloud bills you for a node and your org chart owns a namespace. Bridging the two means decomposing idle capacity twice, because the two idles have different owners and different fixes.
Topics: Kubernetes · method · AWS · GCP · Azure
Every figure in this note is invented and labelled synthetic. No client data was used.
Argument
The claim
Your cloud invoice contains a line for an m7i.4xlarge that ran for 719 hours.
Your engineering organisation contains a team that owns the checkout namespace.
No console bridges those two sentences, and the reason is not that the console
authors were careless. The billing system meters an instance because that is the
thing it rents you. Kubernetes packs many owners onto that instance and keeps no
opinion about money. The bill is correct, the ownership model is correct, and
they do not share a primary key.
Everything difficult about cluster chargeback follows from that one sentence, including its organisational symptom: the platform lead becomes the cost helpdesk. Every question about which team caused the cluster’s growth arrives at the one person who can read both sides, answering it by hand is most of a day, and the answer has to be produced again next month.
Why the bill hides the problem instead of showing it
Here is the trap that makes this worse than a normal allocation exercise: you paid for the node either way. If a team requests 8 CPUs and uses 1, your invoice does not change. The waste is real, it is expensive, and it is invisible in every cost report that starts from the invoice, because the invoice already netted it out into a flat node-hour.
That is why the reported numbers in this discipline look so bad. Roughly 70% of requested Kubernetes CPU and memory is never used, and 68% of enterprises report that Kubernetes costs are partially or entirely invisible in their existing tooling.1 Chargeback adoption sits around 14%, showback around 13%.1 MEDIUM CONFIDENCE: secondary sources citing primary surveys, not verified against the report text
We quote those figures with the flag attached because they are the most-repeated statistics in this field and the least well sourced. The structural claim does not depend on them: even at 30% unused requests, the accounting problem is identical.
Idle is two different problems wearing one word
This is the part most allocation setups get wrong, and it is the part that decides whether anyone acts on your report.
node capacity ── what you rent, what you pay for
├─ Σ pod requests ── what the scheduler reserved
│ ├─ Σ actual usage ── what ran
│ └─ requests − usage ── IDLE 2: a manifest problem
└─ capacity − Σ requests ── IDLE 1: a provisioning problem
Idle 1: node capacity minus the sum of pod requests. Nobody asked for this capacity. It exists because of instance-type granularity, autoscaler headroom, a node that cannot drain, or a stale node selector holding a machine nobody schedules onto. The owner is the platform team. The fix is a NodePool configuration, a consolidation policy, or a different instance family.
Idle 2: pod requests minus actual usage. A team asked for this and did not use it. The owner is the team that owns the manifest. The fix is a pull request against Helm values.
Report those two as one number and the result is a meeting where the platform team and the product teams each correctly explain that it is not their number. Report them separately, with names attached, and each row has exactly one owner who can merge exactly one change. The most-cited unsolved problem in this whole discipline is getting engineers to act on recommendations2, and a recommendation addressed to nobody in particular is a recommendation nobody acts on.
The same decomposition, in dollars
The shape above is easier to argue with once it has money in it. Here is one month of one node pool, decomposed.
| Layer | vCPU-hours | Cost USD | Owner | The fix |
|---|---|---|---|---|
| Node capacity billed | 184,320 | 31,340 | Nobody. This is the invoice | · |
| Σ pod requests | 129,024 | 21,938 | Allocatable to namespaces | · |
| Σ actual usage | 41,287 | 7,021 | Allocatable to namespaces | · |
| Idle 1 · capacity − requests | 55,296 | 9,402 | Platform team | NodePool shape, consolidation policy |
| Idle 2 · requests − usage | 87,737 | 14,917 | The team that owns the manifest | A pull request against Helm values |
| Reported as waste | 143,033 | 24,319 | Two rows, two owners | Two separate changes |
The single most common error here is subtracting usage from capacity and calling the difference waste. That number (24,319 above) is real, but it is addressed to nobody, and it also overstates the safely recoverable amount, because some of Idle 2 is headroom somebody chose on purpose.
The wiring, per cloud
None of this requires new software. It requires enabling things that are off by default, and the enabling is not retroactive on any of the three clouds.
- AWS / EKS. Split cost allocation data has to be turned on for pod-level cost to appear in the Cost and Usage Report. OpenCost or Kubecost covers namespace allocation. The CUR is the reconciliation baseline, not Cost Explorer, which retains hourly and resource-level detail for 14 days and therefore cannot answer a 90-day question at all.
- GCP / GKE. GKE cost allocation has to be enabled to emit namespace and
workload labels into the BigQuery billing export, and only the
detailed usage export (
gcp_billing_export_resource_v1_<ACCOUNT_ID>) carries resource-level names. The cluster management fee and system workloads need an explicit split rule of their own. - Azure / AKS. The Cost Analysis add-on does in-cluster allocation once enabled and is genuinely good at it. Reconcile at the billing scope, not the subscription. Unused reservation charges land on the billing profile or enrollment, not on the subscription that should have consumed the capacity, so subscription-level showback silently hides reservation waste.
GKE Autopilot inverts the whole story
One important exception, because it changes which fix is even relevant.
Autopilot bills pod resource requests directly, with per-pod minimums and memory-to-CPU ratio floors that round small pods up. On Autopilot, requests are the bill. Idle 1 mostly disappears, because you are not renting the node, and Idle 2 becomes a direct, immediate line-item saving.
The practical consequence: on GKE Standard, a rightsizing pull request is a packing change and saves money only if it lets the autoscaler remove a node. On Autopilot the same pull request is a billing change and saves money the hour it merges. Same diff, different economics, and a report that does not distinguish the two will overstate savings on one cluster and understate them on the other.
The shared-cost rule is a political artifact, so write it down first
Some cluster cost belongs to no namespace: the control plane fee, kube-system,
ingress controllers, cross-AZ traffic generated by service topology, and the
observability agents running on every node. Note that several of those meters
carry no user tags at all. NAT gateway per-GB data processing and cross-AZ
transfer are the classic examples, so they cannot be attributed even in
principle without a rule you supply.
There are four defensible rules, and each one is defensible only for a particular kind of pool. The definitions matter more than the names, because the failure mode is applying a rule outside the case it is honest for.
| Rule | Driver | Honest for | Never used for |
|---|---|---|---|
| Usage-weighted | A measured meter: CPU-seconds, byte-seconds, vectors stored, log bytes ingested | Any pool where a meter exists and scales with the thing being shared | Fixed platform cost, because it punishes the team that grew |
| Request-weighted | Application requests reaching the shared component | Caches, embedding indexes, ingress and gateways | Cost that scales with residency rather than with traffic |
| Headcount | People, or licensed seats | Per-seat tooling that people rather than workloads consume | Infrastructure, ever. Headcount does not cause node-hours |
| Flat | None. Equal shares. | Joint cost small enough that arguing about it costs more than it does | Anything a team can change by shipping a diff |
| No defensible rule | None | Reported unattributed, as its own named line | A plausible-looking weight chosen to make the table complete |
The last row is the one that gets skipped. A pool with no defensible driver (NAT gateway per-GB processing with no user tags, cross-AZ transfer generated by service topology nobody owns) is reported as unattributed. Inventing a weight for it produces a complete-looking table whose completeness is the lie.
And the choice moves real money. The same 9,400 USD of shared cluster cost, under each of these rules:
| Team | Direct | Usage-weighted | Request-weighted | Headcount | Flat |
|---|---|---|---|---|---|
| checkout | $21,400 | $4,700 | $3,100 | $1,880 | $2,350 |
| search | $14,900 | $2,600 | $4,400 | $2,820 | $2,350 |
| ml-serving | $38,200 | $1,600 | $1,300 | $1,880 | $2,350 |
| internal-tools | $2,100 | $500 | $600 | $2,820 | $2,350 |
| Spread, highest ÷ lowest team | · | 9.4× | 7.3× | 1.5× | 1.0× |
| Shared cost allocated | · | $9,400 | $9,400 | $9,400 | $9,400 |
internal-tools pays 500 under one defensible rule and 2,820 under another
defensible rule. Both rules are honest. Neither is discoverable from the data.
Which means the rule is a decision someone has to own, in writing, before the
first report circulates. Publish the alternatives alongside it, so a team that
dislikes its number can check the arithmetic instead of suspecting the motive.
Our default is usage-weighted for infrastructure that scales with traffic and flat for fixed platform overhead like the control plane fee, on the grounds that a fixed cost split by usage punishes the team that grew. But the default is less important than the fact that it is stated.
Node-layer fixes, and the one that quietly fails
Once Idle 1 is separated out, the node layer has an owner and a lever:
- Karpenter with a consolidation policy, plus
do-not-disruptdiscipline on stateful workloads and a stated spot-interruption posture. The failure mode worth knowing: consolidation stops dead on a single unschedulable PodDisruptionBudget, and the symptom is not an error. It is a bill that does not fall after you shipped the fix. - AKS node autoprovisioning with spot pools and an eviction policy you have written down rather than inherited.
- GKE Standard node pools, or Autopilot’s request floors, per the inversion above.
Sequence the manifest changes so nothing gets OOMKilled on a Friday. A memory request cut that pages someone at 2am costs more than it saved, and it ends the programme regardless of the arithmetic.
Acceptance is a number, not an opinion
A cluster allocation that cannot be checked is a cluster allocation nobody has to believe. The test we hold this work to is the same one the rest of the method uses: allocated cluster cost reconciles to the cluster’s billed cost, within 3%, for a full billing month, with the residual reported as its own line rather than smeared across namespaces to make the total tie.
That single sentence does more work than it looks. It fails if EKS split cost
allocation data was never enabled, because pod-level cost never reaches the CUR.
It fails if the control plane fee and kube-system were quietly dropped instead
of being split under a stated rule. It fails if the cluster spans two node pools
and only one was modelled. Each of those is a real bug that a namespace-level
dashboard displays without complaint. The reconciliation is what turns them into
a failing check. Why an exact tie-out is the warning
sign covers the general case.
Where this method is weakest
Two places, both worth stating before someone finds them.
Requests-versus-usage is a lower bound on safe savings, not a target. Usage measured over 30 days does not contain the traffic spike that has not happened yet. Rightsizing to the observed p99 with no headroom is a reliability decision disguised as a cost decision, and this method has no opinion about the value of reliability. It cannot price an outage.
Memory is harder than CPU and the tooling is honest about it less often than it should be. AWS Compute Optimizer, for instance, needs 14 days of CloudWatch metrics and cannot see memory at all without the CloudWatch agent installed, so its over-provisioned verdicts are effectively CPU-only, and following them unmodified will recommend downsizes that OOMKill.
What would change our mind
If the managed Kubernetes services began emitting a first-class ownership dimension into the billing export by default (namespace and workload labels on every row, with the shared-cost split configurable in the console) the allocation half of this work disappears. AKS Cost Analysis and GKE cost allocation are both steps in that direction, and we would rather say so than pretend the consoles are useless.
What would remain is the part that was never a data problem: deciding the split rule, sequencing the manifest changes, and getting them merged.
One shared pool, priced under each of the four rules
Relocated from /methodology, where the four split rules are stated. The same pool, 17,181.00 USD for the month, under each rule. Seats 180 / 90 / 60.
| Rule | Driver | Tenant A | Tenant B | Tenant C | Pool |
|---|---|---|---|---|---|
| Usage-weighted | CPU-sec + byte-sec consumed | 5,602.50 | 7,470.00 | 4,108.50 | 17,181.00 |
| Request-weighted | 9.29M application requests | 7,619.78 | 6,214.37 | 3,346.85 | 17,181.00 |
| Headcount | 330 licensed seats | 9,371.45 | 4,685.73 | 3,123.82 | 17,181.00 |
| Flat | none | 5,727.00 | 5,727.00 | 5,727.00 | 17,181.00 |
| Range, highest minus lowest | how much the choice moves it | 3,768.95 | 2,784.27 | 2,603.18 | · |
Sources
- Kubernetes cost visibility, chargeback and showback adoption, and unused request share: spendark.com, cast.ai. MEDIUM CONFIDENCE: both are secondary sources citing primary surveys; the figures were not verified against the primary report text.
- “Getting engineers to take action on recommendations” as the top-rated practitioner challenge: holori.com, data.finops.org/2025-report. MEDIUM CONFIDENCE.
Sources are listed at the foot of this note with their confidence stated. Medium and low confidence figures say so in the copy.
Every query here runs against a read-only role you create, scope and revoke. Access policy.
Close
More notes
The rung this argument belongs to
Instrumentation Build · Fixed fee per stage
Check the arithmetic yourself
The full derivations, with the queries written out so they run under your own read-only credentials, in your own console.