Note

Published

2026-06-18

Reading

18 min

Data

Synthetic figures

Headline figure

High

Claim

Per-tenant AI cost is a join, not a report

No provider console can tell you what a customer costs, because the provider bills an API key and only your application knows which customer was behind the request. Cost per tenant is a join between provider usage records and your own request logs, and the join key has to be designed before anyone asks for the number.

18Reading · minutes
2026-06-18Published

Topics: AI and GPU · method · AWS · Azure · GCP

Every figure in this note is invented and labelled synthetic. No client data was used.

Argument

The claim

You sell seats or usage. You pay per token and per GPU-hour. Nothing in either provider’s data model connects the two, and no amount of dashboard configuration will fix that, because the missing field does not exist anywhere in the provider’s records.

A provider knows three things about a request: which credential authenticated it, which model served it, and how many tokens moved. It does not know which of your customers caused it, which feature of your product issued it, whether it was a first attempt or the third retry of a failed tool call, or whether the 40,000-token prefix was a cache write that will never be read again. Those are application-level facts. They live in your request logs and nowhere else.

So cost per customer is not a report that someone has failed to build. It is a join, and like every join it is only as good as the key. The work is deciding where in your architecture the tenant dimension gets stamped onto a provider call, then proving the joined total reconciles against the invoice.

What each provider actually gives you

Start from what is on offer, because the shape of the seam depends on it.

Anthropic and OpenAI both expose Admin API usage and cost endpoints. The grain is per API key and per workspace, per time bucket. That is a real allocation dimension, if your key hygiene is already right. If one production key serves every tenant, the endpoint can tell you your total and nothing else, and no retroactive fix exists. Key structure is an allocation design decision you make months before finance asks the question.

Amazon Bedrock does not log invocations by default. Model invocation logging has to be enabled to CloudWatch Logs or S3 before it records anything, and it is not retroactive. Once on, the useful fields are inputTokenCount, outputTokenCount and cacheReadInputTokenCount. For splitting spend rather than counting tokens, the supported mechanism is an application inference profile carrying cost allocation tags, which is the closest thing any provider currently ships to “bill this call to that dimension.”

Azure OpenAI reports ProcessedPromptTokens and GeneratedTokens per deployment through Azure Monitor, with no per-tenant dimension at all. Worse for the arithmetic: a Provisioned Throughput Unit deployment bills reserved capacity, not tokens. A PTU deployment sized for a peak that arrives twice a month is expensive and completely invisible in any token-based report, because the waste is in the capacity, not the traffic. The practical attribution seam is API Management in front of the deployment, with a token-emitting policy and per-tenant subscription keys.

Vertex AI has request logging, and its own separate waste mode: an online prediction endpoint bills per node-hour whether or not anything calls it. An undeployed-but-not-deleted endpoint is the most common Vertex line item nobody claims.

Four providers, four different grains, and not one of them has a tenant column. The join keys that do exist, side by side:

What each provider gives you to join on, and what you have to add
ProviderUsage recordNative grainYou must supply
Anthropic, OpenAI Admin API usage and cost endpoints API key, workspace, time bucket Key-to-tenant mapping, decided before the keys were issued
Amazon Bedrock Model invocation logs (off by default) Invocation, model, token counts Application inference profile with cost allocation tags, or gateway metadata
Azure OpenAI Azure Monitor metrics Deployment, per metric interval API Management in front, token-emitting policy, per-tenant subscription keys
Vertex AI Request logging to BigQuery Request, endpoint, model Tenant in request metadata, plus endpoint-to-tenant mapping
Self-hosted GPU DCGM metrics plus your scheduler Node, pod, GPU device Namespace-to-tenant mapping and a stated GPU split rule

Capacity metering is a second bill that tokens cannot see

Before the join, one trap that invalidates the whole exercise if you miss it. Two of the four providers meter in two incompatible modes, and only one of them is about tokens.

Azure OpenAI standard and global standard deployments bill tokens. Provisioned Throughput Unit deployments bill reserved capacity, whether or not a token passes through. A PTU deployment sized for a peak that arrives twice a month is pure waste and it is invisible in every token-based cost report you own, because the waste is in the capacity and not in the traffic. Bedrock has the same shape: on-demand and Provisioned Throughput model units are different products with different waste modes. Vertex online prediction endpoints are the extreme case: they bill per node-hour at zero traffic, which is why an undeployed-but-not- deleted endpoint is the most common Vertex line item nobody claims.

The practical consequence for the model: capacity cost has to enter the allocation as a pool split under a stated rule, not as a per-token rate. If you divide PTU cost by tokens served to get a blended rate, a quiet month makes every tenant look expensive and a busy month makes the deployment look efficient, and neither reading is about anything a tenant did.

The three seams that can carry a tenant

There are only three places the dimension can come from, and choosing between them is an architecture decision, not a reporting one.

  1. Per-tenant or per-workspace credentials. Cheapest to reason about, worst to operate at scale, and it breaks the moment one tenant’s traffic is served by a shared background job.
  2. A gateway or middleware that stamps metadata on every provider call: tenant, feature, model, cache-hit class, retry count. This is the only seam that survives adding a fourth provider, because the dimension lives in your code rather than in a vendor’s schema.
  3. Provider-native tagging where it exists: Bedrock application inference profiles with cost allocation tags. Correct and precise, and it does not generalise off AWS.

Most estates end up with the gateway as the source of truth and the provider’s own records as the reconciliation check. That ordering matters: your logs are the allocation dimension, the provider’s invoice is the total, and the gap between them is the thing you have to publish rather than hide.

The arithmetic, on synthetic figures

Below is the derivation we run, on invented numbers. Nothing here came from a client estate.

SYNTHETIC
Cost per tenant, one 30-day window, one model family
LineBasisTenant ATenant B
Input tokensgateway log, tenant-stamped184,300,00041,900,000
Cache read tokensgateway log, cache-hit class96,200,0001,100,000
Cache write tokensgateway log, cache-hit class7,400,0009,800,000
Output tokensgateway log, tenant-stamped22,600,0008,300,000
Direct token costtokens × per-model rate card$1,842$913
Retry multiplierattempts ÷ resolved tasks1.07×1.41×
Direct cost after retriesabove, restated$1,971$1,287
Shared embedding indexquery volume weighted$318$96
GPU serving, own namespaceSM-active-weighted GPU-hours$2,240$0
Observability overheadlog volume weighted$74$31
Cost to servesum of the above$4,603$1,414
Plan revenuebilling system$12,000$1,800
Gross marginrevenue − cost to serve62%21%
Synthetic example. Not a client. No client data was used.

Three things in that table are the whole argument.

The retry multiplier is a per-tenant fact, not a global one. Tenant B costs 41% more than its token count implies because its traffic shape triggers tool-call failures and escalations. A provider dashboard cannot see the difference between an attempt and a resolution, so it reports both tenants as healthy. This is also why model routing has to be priced on cost per resolved task: a cheaper model at 2.3 average attempts plus one human escalation costs more than the expensive model that got it right the first time.

The cache line is where money quietly disappears. Tenant B pays a cache write premium on 9.8M tokens and reads almost none of it back. That prompt shape is pure loss, and it does not raise an error anywhere. A cache-key refactor that includes a build hash will destroy a hit rate silently, which is why the monitor that matters is the ratio of cache-read to cache-creation tokens, watched over deploys rather than checked once.

The shared index needs a stated rule, not a default. Query-volume weighting is a choice. Headcount weighting, request weighting and a flat split each move these two numbers materially, and if the rule is not written down before the numbers are shown, the first team to dislike its allocation will argue that the split was rigged. They will be right to ask.

Prompt caching has a break-even, and it is one equation

The cache line above is worth doing properly, because caching is sold as a straightforward discount and it is actually a bet with a break-even point.

Providers price a cached prefix in two parts: a write premium the first time the prefix is stored, and a read discount every time it is reused inside the time-to-live window. Write the multipliers as fractions of the ordinary input rate (write costs w × base, read costs r × base) and the arithmetic falls out. For a prefix written once and read k times inside the TTL:

uncached cost   = (k + 1) · N          -- every call pays the full input rate
cached cost     = w · N  +  k · r · N  -- one write, k discounted reads

break-even when   w + k·r = k + 1
                       k  = (w − 1) / (1 − r)

That is the whole thing. k is not “how often is this prompt used”. It is reads per TTL window, which is a much harsher denominator. A prefix reused forty times a day across a five-minute TTL may never be read twice inside one window, in which case every call pays a write premium and receives no discount.

SYNTHETIC
Break-even reads per TTL window, at illustrative multipliers
Prompt shapeWrite wRead rBreak-even kMeasured reads / windowVerdict
Shared system prompt, short TTL1.25×0.10×0.2814.2Cache it
Per-tenant document context, short TTL1.25×0.10×0.281.9Cache it
Per-tenant context, long TTL2.00×0.10×1.111.9Marginal. Measure again
Per-conversation history prefix1.25×0.10×0.280.2Paying the write premium for nothing
Synthetic example. Not a client. Multipliers are illustrative. Write premiums, read discounts and TTLs differ by provider and by cache tier, and they change; the break-even is recomputed against your current rate card, not assumed.

The bottom row is the case that costs real money quietly. It raises no error, it shows up in no dashboard, and the only signal is the ratio of cache-read tokens to cache-creation tokens. Watch that ratio over deploys, not once: a cache-key refactor that folds a build hash into the prefix destroys the hit rate on the release that ships it, and the bill moves a fortnight before anyone connects the two.

Routing priced on resolved tasks, not on tokens

Cost per token is the wrong denominator and it recommends the wrong model with great consistency. The right denominator is a resolved task: one unit of work the user actually wanted, including every retry it took and the human who finished it when the model could not.

SYNTHETIC
Four routing options, priced per resolved task
OptionToken cost / attemptMean attemptsUnresolvedEscalation costCost / resolved task
Small model only$0.0112.312.0%$1.080$1.105
Mid model only$0.0431.45.0%$0.450$0.510
Large model only$0.1801.11.5%$0.135$0.333
Small, escalating to large on failure$0.0111.6 + 0.31.5%$0.135$0.212
Synthetic example. Not a client. Escalation cost is 12% (or 5%, 1.5%) of tasks × a $9.00 loaded human support cost, which is a figure you supply. It is the input that decides the answer, and it is not a cloud number.

The small model is roughly sixteen times cheaper per token than the large one and more than three times more expensive per resolved task. Every routing decision priced on the first number is wrong, and the arithmetic that exposes it needs exactly two application-level facts no provider has: attempts per task, and what an unresolved task costs you.

Two conditions on this table before anybody acts on it. Routing changes run against your existing eval suite, not against an opinion. If there is no eval suite, that is a prerequisite you own, and it should be said on the first call rather than in week two. And the escalation cost has to be a real loaded figure from your support organisation; guessing it moves the winner.

GPU-hours, measured honestly

For self-hosted inference and training, the equivalent trap is the utilization metric everyone screenshots.

DCGM_FI_DEV_GPU_UTIL reports whether any kernel was resident on the device during the sampling interval. It reads near 100% for a single small kernel on an H100 with 95% of its streaming multiprocessors idle. It is close to useless as a waste signal and it is the number most often produced to prove a cluster is busy.

The two that carry information are DCGM_FI_PROF_SM_ACTIVE, which reports streaming-multiprocessor occupancy, and DCGM_FI_DEV_FB_USED, which reports framebuffer in use. Together they show the headroom that actually exists.

The dollar figure is then a subtraction, not a percentage:

idle GPU dollars
  = (allocated GPU-hours − SM-active-weighted GPU-hours) × your node rate

where SM-active-weighted GPU-hours
  = Σ over intervals ( interval_hours × DCGM_FI_PROF_SM_ACTIVE )

The output is dollars with a derivation attached, priced at your real node rate rather than at list. It splits into named, separately owned modes:

GPU waste, by mode
ModeSignalOwner
Idle nodes held by a stale selectorAllocated, zero SM-active, zero podsPlatform
Over-requested pods8 GPUs requested, one device with SM-active above noiseThe owning team
MIG or time-slicing headroomSmall models on whole devices, framebuffer far below capacityPlatform
Jobs holding capacity after completionSM-active at zero with the pod still RunningThe owning team
Orphaned checkpoints and block storageVolumes with no pod, from runs that endedThe owning team
Dev keys and dev endpoints on productionCost in the production project with no production trafficWhoever issued the key

The inference waste catalogue, ranked

Across everything above, the recurring modes in rough order of how much money they have been worth. The ordering is our judgement, not a measurement.

Eight ways inference spend leaves without an error
#Waste modeMechanismWhat detects it
1Oversized PTU or provisioned throughputReserved capacity billed at a peak that arrives twice a monthCapacity cost ÷ tokens served, plotted by day
2Retry stormsFailed tool calls retried inside an agent loop, each attempt billedAttempts ÷ resolved tasks, per tenant and per feature
3Unbounded agent fan-outOne user action expanding into dozens of sub-calls with no ceilingCalls per user action, p99 rather than mean
4Full-corpus re-embeddingIndex key includes the build hash, so every deploy re-embeds everythingEmbedding token volume correlated with deploy events
5Cache writes never readPrefix shape below the break-even reuse countCache-read ÷ cache-creation token ratio, over deploys
6Streaming reconnects billed twiceClient reconnects mid-stream; the provider counts a second generationProvider token total exceeding your gateway total
7Undeleted Vertex endpointsOnline prediction endpoints billing per node-hour at zero trafficEndpoint cost with no request-log rows
8Dev keys on the production projectNon-production traffic inside the production cost centreCost by key with no matching production tenant

Six of the eight are only visible in a join between your logs and the provider’s records. That is the practical case for building the seam before anyone asks for cost per customer: the seam is also the detector.

The join, in SQL

The shape is always the same: aggregate your own logs to the grain you want to bill at, aggregate the provider’s records to the same time grain, then price and reconcile. Roughly:

-- 1. your logs: the only place tenant exists
with app as (
  select
    date_trunc('day', request_ts)        as usage_day,
    tenant_id,
    feature,
    model_id,
    cache_class,                          -- read | write | none
    count(*)                              as attempts,
    count(distinct task_id)               as resolved_tasks,
    sum(input_tokens)                     as input_tokens,
    sum(output_tokens)                    as output_tokens
  from request_log
  where request_ts >= current_date - interval '90' day
  group by 1,2,3,4,5
),

-- 2. the provider: the authoritative total, no tenant column
provider as (
  select
    date_trunc('day', invocation_ts)      as usage_day,
    model_id,
    sum(input_token_count)                as input_tokens,
    sum(output_token_count)               as output_tokens,
    sum(cache_read_input_token_count)     as cache_read_tokens
  from bedrock_invocation_log
  group by 1,2
)

-- 3. the reconciliation, printed before any allocation is trusted
select
  a.usage_day,
  a.model_id,
  a.input_tokens                              as app_input,
  p.input_tokens                              as provider_input,
  a.input_tokens - p.input_tokens             as residual_tokens,
  abs(a.input_tokens - p.input_tokens)
    / nullif(p.input_tokens, 0)               as residual_share
from app a
full outer join provider p
  on a.usage_day = p.usage_day
 and a.model_id  = p.model_id
group by 1,2,3,4,5,6
order by residual_share desc;

The reconciliation query runs first and prints loudly. If your logs and the provider’s records disagree by more than a few percent on total tokens, every per-tenant number downstream is decoration. Common causes, in the order we find them: streamed responses counted once in the app and twice at the provider after a client reconnect, background jobs bypassing the gateway, a dev key billing to the production project, and usage-API reporting lag of several hours to a day that makes the last bucket look wrong when it is simply incomplete.

Error bars are part of the deliverable

A cost-per-customer figure quoted to the cent off a lagging API is a fabrication. The residual sources we report by name, every time: provider token rounding, usage-API reporting lag of hours to a day, mid-window model price changes, and the shared-index amortisation choice. Our working tolerance is within 3% of the invoice, and the residual gets reported rather than distributed silently across tenants to make the table add up.

Distributing the residual is the single most common way a unit-cost model becomes untrustworthy, because it hides the one number that tells you whether to believe the rest. The general argument is in a cost model that ties out exactly is hiding something.

Why this is a market gap and not just a hard afternoon

Organisations tracking AI cost went from 31% in 2024 to 63% in 2025 to 98% in 2026, and granular monitoring of AI spend (tokens, LLM requests, GPU utilisation) is the single most-requested capability in the 2026 practitioner survey.1 Inference is now roughly 80% of AI GPU spend industry-wide2 MEDIUM CONFIDENCE: secondary source

Demand is not the constraint. The constraint is that the last mile of this problem is inside the customer’s application, which is exactly the part a vendor cannot ship. A tool can offer you a gateway, a schema and a rate card. It cannot decide what a tenant is in your data model, which of your three background job runners should be billed to whom, or whether a retry storm is the customer’s fault or yours.

What would change our mind

One thing, and we watch for it quarterly: if every major provider accepted an arbitrary caller-supplied metadata dimension on each request and reported cost broken down by it, this work collapses from an engagement into a console checkbox. Bedrock application inference profiles are already most of the way there on one cloud. Anthropic and OpenAI per-workspace cost reporting is genuinely sufficient today if your key hygiene was designed correctly from the start. Saying so is more useful to you than pretending otherwise.

The part that survives is not measurement. It is the decisions measurement enables: routing policy, cache architecture, capacity shape, and pre-launch cost modelling for a feature that does not exist yet. Those are architecture judgments, and no console is close to them.

Where this number is weakest

The GPU serving line in the table above. Allocating GPU-hours to a tenant assumes the tenant had a namespace of its own, which is true in the synthetic example and often false in reality. On a shared inference deployment, GPU cost per tenant is an allocation rule with error bars, not a measurement, and it should be labelled that way in whatever the board eventually reads.

Cost per tenant, from the /methodology derivation

Synthetic

The methodology page derives one month of a multi-tenant AI application line by line against an 82,617.00 USD invoice. This is the number that derivation was for.

TenantAttributed cost USD/moRequests, millionsCost / 1,000 req USDShare of invoice
Tenant A26,344.494.126.3931.9%
Tenant B20,649.623.366.1525.0%
Tenant C10,394.421.815.7412.6%
Named tenants57,388.539.296.1869.5%
Synthetic example. Not a client. No client data was used. 69.5% of the invoice has a named tenant. Raising that share is the deliverable; the remaining 30.5% is idle, untagged and un-ownable cost, reported as its own lines rather than spread across the three tenants to make the columns tie out.

Prompt caching break-even on one prompt shape

Synthetic

Relocated from the Inference Cost Controls page.

LineFigure
Cache write premium over base input+25%
Cache read discount against base input−90%
Reuse count at which the shape turns profitable2
Shapes currently writing and never reading11
Monthly loss on those 11 shapes3,410
Break-even reuse count, this shape2 reads
Synthetic example. Not a client. No client data was used. Every figure is invented; the real premium and discount ratios are read from your provider's own price list at build time.

Sources

  1. State of FinOps, AI cost tracking and top requested capability: data.finops.org, summarised at nops.io/blog/state-of-finops-2026. High confidence.
  2. Inference share of AI GPU spend: spheron.network. MEDIUM CONFIDENCE: secondary source, not a primary survey.

Sources are listed at the foot of this note with their confidence stated. Medium and low confidence figures say so in the copy.

Every query here runs against a read-only role you create, scope and revoke. Access policy.

Close

The rung this argument belongs to

Inference Cost Controls · Fixed fee

Check the arithmetic yourself

The full derivations, with the queries written out so they run under your own read-only credentials, in your own console.