Why

Methodology

How a cost per customer is built, how shared cost is split, and where the residual comes from.

±3%Target
4Split rules

Coverage

Coverage

2 / 2Deep / scoped
Depth per estate
PageDepthExport of record
AWSdeepCUR 2.0 to S3, FOCUS 1.1 as cross-check
AI · LLM · GPUdeepBedrock invocation logs, Admin API usage, DCGM
AzurescopedCost Management to ADLS Gen2 · amortized AND actual
GCPscopedDetailed usage export to BigQuery, plus pricing

Basis

Cost basis

Amortized effective cost, under every unit-cost number here.

The four bases compared: list, billed, unblended, amortized effective.

Derivation

Cost per customer, derived

  1. 01
    Reconcile to invoiceMiss it and nothing above holds.
  2. 02
    Remove directly attributableDedicated resources leave at cost.
  3. 03
    Decompose idle: two ownersProvisioning, then manifests.
  4. 04
    Split on a measured driverA month's mix is not an hour's.
  5. 05
    Report the remainderNAT and cross-AZ carry no tag.

Pass 03 exists because roughly 70% of requested capacity is unused1MEDIUM CONFIDENCE.

5Passes
1 hResolution

Why SM-active, not GPU-UTIL.

Worked

Synthetic

One month, three tenants

One AWS account, one month. Invoice:$82,617.00. Amortized effective cost.

Derivation, line by line
LineDriverTenant ATenant BTenant CNon-tenantTotal
Dedicated infratenant tag6,400.00···6,400.00
LLM tokensrequest log8,845.647,213.923,886.071,524.3721,470.00
GPU node groupSM-active hours3,394.353,930.301,607.85·8,932.50
Shared clusterCPU-sec + byte-sec5,602.507,470.004,108.501,494.0018,675.00
Shared vector indexstored vectors1,622.40967.20530.40·3,120.00
Log ingestionlog bytes479.601,068.20261.60370.602,180.00
Attributed·26,344.4920,649.6210,394.423,388.9760,777.50
GPU idleheld, no SM···10,917.5010,917.50
Cluster idlecapacity − requests···6,225.006,225.00
NAT + cross-AZno owner exists···1,940.001,940.00
Untagged remaindercontrol plane, S3···2,280.002,280.00
Model total·26,344.4920,649.6210,394.4224,751.4782,140.00
Invoicesame month····82,617.00
Residualinside ±3%····(477.00)
Synthetic. Not a client.Cost per tenant, from these rows.

Queries

The join, in the query

Pass 1 · amortized effective cost
-- Athena over CUR 2.0 (Parquet, Glue partition projection). One payer account,
-- one billing month. Nothing is filtered out: credits, tax and support stay in,
-- because the model has to reconcile to the invoice, not to a flattering subset.
WITH effective AS (
  SELECT
    line_item_usage_start_date                    AS ts,
    line_item_resource_id                         AS resource_id,
    line_item_product_code                        AS service,
    resource_tags['user_tenant']                  AS tenant_tag,
    resource_tags['user_namespace']               AS namespace_tag,
    -- Amortized effective cost. Unblended cost is wrong for unit economics by
    -- exactly your commitment discount, so it is never the basis.
    CASE line_item_line_item_type
      WHEN 'SavingsPlanCoveredUsage'  THEN savings_plan_savings_plan_effective_cost
      WHEN 'DiscountedUsage'          THEN reservation_effective_cost
      WHEN 'SavingsPlanUpfrontFee'    THEN 0        -- amortized into covered usage
      WHEN 'SavingsPlanRecurringFee'  THEN 0
      WHEN 'RIFee'                    THEN 0
      ELSE line_item_unblended_cost
    END                                           AS effective_cost
  FROM cur2.line_items
  WHERE billing_period = '2026-07'
)
Pass 2 · hourly shares from your logs
-- Shared pool: effective cost with no tenant dimension of its own.
, shared_pool AS (
  SELECT date_trunc('hour', ts) AS h, namespace_tag, SUM(effective_cost) AS cost
  FROM effective
  WHERE tenant_tag IS NULL AND namespace_tag IS NOT NULL
  GROUP BY 1, 2
)

-- Weights come from the application's own request logs, hour by hour, because
-- the tenant mix inside a month is not the tenant mix of the month. Allocating
-- on a monthly total silently charges the night shift to the day shift.
, weights AS (
  SELECT date_trunc('hour', request_ts) AS h,
         namespace,
         tenant_id,
         SUM(cpu_core_seconds + ram_byte_seconds / 4.294967296e9) AS w
  FROM app.request_log            -- your logs. Never leaves your account.
  WHERE request_ts >= DATE '2026-07-01' AND request_ts < DATE '2026-08-01'
  GROUP BY 1, 2, 3
)

, shares AS (   -- per hour, per namespace, the shares sum to exactly 1.0
  SELECT h, namespace, tenant_id,
         w / SUM(w) OVER (PARTITION BY h, namespace) AS share
  FROM weights
)

SELECT s.tenant_id,
       SUM(p.cost * s.share) AS allocated_effective_cost
FROM shared_pool p
JOIN shares s ON s.h = p.h AND s.namespace = p.namespace_tag
GROUP BY 1
ORDER BY 2 DESC;
Pass 3 · GPU idle, measured
# Allocated GPU-hours per namespace. A pod holding a GPU is billed for it
# whether or not a kernel ever runs. 1-minute scrape interval assumed; state
# yours, because this divisor is the whole measurement.
sum by (exported_namespace) (
  count_over_time(DCGM_FI_DEV_FB_USED{exported_namespace!=""}[1h])
) / 60

# SM-active-weighted GPU-hours: the fraction of that held time a kernel was
# actually resident on the SMs.
#
# DCGM_FI_DEV_GPU_UTIL is NOT usable for this. It reads near 100% whenever any
# single kernel is resident, so a pod using one SM of 132 reports as saturated.
sum by (exported_namespace) (
  avg_over_time(DCGM_FI_PROF_SM_ACTIVE{exported_namespace!=""}[1h])
)

Shared

Shared cost

The rule chosen is recorded beside the number it produces.One pool, each rule priced.

Split rules
RuleDriverUsed forNever for
Usage-weightedCPU-sec, byte-secAny metered pool·
Request-weightedRequestsCachesResidency-scaled cost
HeadcountLicensed seatsTooling seatsInfrastructure, ever
FlatnoneToo small to argueAnything a team can change
No defensible rulenoneIts own lineA plausible weight
4Rules

Tolerance

Error bars, as policy

The residual is reported as its own line, never spread across the allocation.

The five residual sources, itemized.

±3%Tolerance · per cloud, per month
5Residual sources
NeverResidual spread silently

Good enough to price with, not to bill from.

Limits

What this cannot say

Four limits
Cannot sayWhy
Whether a fix is worth itA five-figure saving can cost two engineers a quarter
Idle or deliberate headroomIdentical in an export
What a customer is worthIn no export
Whether engineers will actA merged PR needs approval

Scoped to 500K to 5M a year2MEDIUM CONFIDENCE. Below it, the residual beats the finding.

4Limits

Recon

This page, reconciled

Figures asserted
2
Sourced
2
Medium confidence
2
Synthetic, labelled
1
Client figures
none used
Tolerance
±3%

Sources

  1. Roughly 70% of requested capacity is never used.MEDIUM CONFIDENCEspendark.com;cast.ai.
  2. Outsourced help framed for the 500K to 5M spend band.MEDIUM CONFIDENCEinfracost.io;finops.org.
NoneUnsourced