Note
Published
2026-06-18
Reading
18 min
Data
Synthetic figures
Headline figure
High
Claim
Per-tenant AI cost is a join, not a report
No provider console can tell you what a customer costs, because the provider bills an API key and only your application knows which customer was behind the request. Cost per tenant is a join between provider usage records and your own request logs, and the join key has to be designed before anyone asks for the number.
Topics: AI and GPU · method · AWS · Azure · GCP
Every figure in this note is invented and labelled synthetic. No client data was used.
Argument
The claim
You sell seats or usage. You pay per token and per GPU-hour. Nothing in either provider’s data model connects the two, and no amount of dashboard configuration will fix that, because the missing field does not exist anywhere in the provider’s records.
A provider knows three things about a request: which credential authenticated it, which model served it, and how many tokens moved. It does not know which of your customers caused it, which feature of your product issued it, whether it was a first attempt or the third retry of a failed tool call, or whether the 40,000-token prefix was a cache write that will never be read again. Those are application-level facts. They live in your request logs and nowhere else.
So cost per customer is not a report that someone has failed to build. It is a join, and like every join it is only as good as the key. The work is deciding where in your architecture the tenant dimension gets stamped onto a provider call, then proving the joined total reconciles against the invoice.
What each provider actually gives you
Start from what is on offer, because the shape of the seam depends on it.
Anthropic and OpenAI both expose Admin API usage and cost endpoints. The grain is per API key and per workspace, per time bucket. That is a real allocation dimension, if your key hygiene is already right. If one production key serves every tenant, the endpoint can tell you your total and nothing else, and no retroactive fix exists. Key structure is an allocation design decision you make months before finance asks the question.
Amazon Bedrock does not log invocations by default. Model invocation
logging has to be enabled to CloudWatch Logs or S3 before it records anything,
and it is not retroactive. Once on, the useful fields are inputTokenCount,
outputTokenCount and cacheReadInputTokenCount. For splitting spend rather
than counting tokens, the supported mechanism is an application inference
profile carrying cost allocation tags, which is the closest thing any provider
currently ships to “bill this call to that dimension.”
Azure OpenAI reports ProcessedPromptTokens and GeneratedTokens per
deployment through Azure Monitor, with no per-tenant dimension at all. Worse for
the arithmetic: a Provisioned Throughput Unit deployment bills reserved
capacity, not tokens. A PTU deployment sized for a peak that arrives twice a
month is expensive and completely invisible in any token-based report, because
the waste is in the capacity, not the traffic. The practical attribution seam is
API Management in front of the deployment, with a token-emitting policy and
per-tenant subscription keys.
Vertex AI has request logging, and its own separate waste mode: an online prediction endpoint bills per node-hour whether or not anything calls it. An undeployed-but-not-deleted endpoint is the most common Vertex line item nobody claims.
Four providers, four different grains, and not one of them has a tenant column. The join keys that do exist, side by side:
| Provider | Usage record | Native grain | You must supply |
|---|---|---|---|
| Anthropic, OpenAI | Admin API usage and cost endpoints | API key, workspace, time bucket | Key-to-tenant mapping, decided before the keys were issued |
| Amazon Bedrock | Model invocation logs (off by default) | Invocation, model, token counts | Application inference profile with cost allocation tags, or gateway metadata |
| Azure OpenAI | Azure Monitor metrics | Deployment, per metric interval | API Management in front, token-emitting policy, per-tenant subscription keys |
| Vertex AI | Request logging to BigQuery | Request, endpoint, model | Tenant in request metadata, plus endpoint-to-tenant mapping |
| Self-hosted GPU | DCGM metrics plus your scheduler | Node, pod, GPU device | Namespace-to-tenant mapping and a stated GPU split rule |
Capacity metering is a second bill that tokens cannot see
Before the join, one trap that invalidates the whole exercise if you miss it. Two of the four providers meter in two incompatible modes, and only one of them is about tokens.
Azure OpenAI standard and global standard deployments bill tokens. Provisioned Throughput Unit deployments bill reserved capacity, whether or not a token passes through. A PTU deployment sized for a peak that arrives twice a month is pure waste and it is invisible in every token-based cost report you own, because the waste is in the capacity and not in the traffic. Bedrock has the same shape: on-demand and Provisioned Throughput model units are different products with different waste modes. Vertex online prediction endpoints are the extreme case: they bill per node-hour at zero traffic, which is why an undeployed-but-not- deleted endpoint is the most common Vertex line item nobody claims.
The practical consequence for the model: capacity cost has to enter the allocation as a pool split under a stated rule, not as a per-token rate. If you divide PTU cost by tokens served to get a blended rate, a quiet month makes every tenant look expensive and a busy month makes the deployment look efficient, and neither reading is about anything a tenant did.
The three seams that can carry a tenant
There are only three places the dimension can come from, and choosing between them is an architecture decision, not a reporting one.
- Per-tenant or per-workspace credentials. Cheapest to reason about, worst to operate at scale, and it breaks the moment one tenant’s traffic is served by a shared background job.
- A gateway or middleware that stamps metadata on every provider call: tenant, feature, model, cache-hit class, retry count. This is the only seam that survives adding a fourth provider, because the dimension lives in your code rather than in a vendor’s schema.
- Provider-native tagging where it exists: Bedrock application inference profiles with cost allocation tags. Correct and precise, and it does not generalise off AWS.
Most estates end up with the gateway as the source of truth and the provider’s own records as the reconciliation check. That ordering matters: your logs are the allocation dimension, the provider’s invoice is the total, and the gap between them is the thing you have to publish rather than hide.
The arithmetic, on synthetic figures
Below is the derivation we run, on invented numbers. Nothing here came from a client estate.
| Line | Basis | Tenant A | Tenant B |
|---|---|---|---|
| Input tokens | gateway log, tenant-stamped | 184,300,000 | 41,900,000 |
| Cache read tokens | gateway log, cache-hit class | 96,200,000 | 1,100,000 |
| Cache write tokens | gateway log, cache-hit class | 7,400,000 | 9,800,000 |
| Output tokens | gateway log, tenant-stamped | 22,600,000 | 8,300,000 |
| Direct token cost | tokens × per-model rate card | $1,842 | $913 |
| Retry multiplier | attempts ÷ resolved tasks | 1.07× | 1.41× |
| Direct cost after retries | above, restated | $1,971 | $1,287 |
| Shared embedding index | query volume weighted | $318 | $96 |
| GPU serving, own namespace | SM-active-weighted GPU-hours | $2,240 | $0 |
| Observability overhead | log volume weighted | $74 | $31 |
| Cost to serve | sum of the above | $4,603 | $1,414 |
| Plan revenue | billing system | $12,000 | $1,800 |
| Gross margin | revenue − cost to serve | 62% | 21% |
Three things in that table are the whole argument.
The retry multiplier is a per-tenant fact, not a global one. Tenant B costs 41% more than its token count implies because its traffic shape triggers tool-call failures and escalations. A provider dashboard cannot see the difference between an attempt and a resolution, so it reports both tenants as healthy. This is also why model routing has to be priced on cost per resolved task: a cheaper model at 2.3 average attempts plus one human escalation costs more than the expensive model that got it right the first time.
The cache line is where money quietly disappears. Tenant B pays a cache write premium on 9.8M tokens and reads almost none of it back. That prompt shape is pure loss, and it does not raise an error anywhere. A cache-key refactor that includes a build hash will destroy a hit rate silently, which is why the monitor that matters is the ratio of cache-read to cache-creation tokens, watched over deploys rather than checked once.
The shared index needs a stated rule, not a default. Query-volume weighting is a choice. Headcount weighting, request weighting and a flat split each move these two numbers materially, and if the rule is not written down before the numbers are shown, the first team to dislike its allocation will argue that the split was rigged. They will be right to ask.
Prompt caching has a break-even, and it is one equation
The cache line above is worth doing properly, because caching is sold as a straightforward discount and it is actually a bet with a break-even point.
Providers price a cached prefix in two parts: a write premium the first time
the prefix is stored, and a read discount every time it is reused inside the
time-to-live window. Write the multipliers as fractions of the ordinary input
rate (write costs w × base, read costs r × base) and the arithmetic falls
out. For a prefix written once and read k times inside the TTL:
uncached cost = (k + 1) · N -- every call pays the full input rate
cached cost = w · N + k · r · N -- one write, k discounted reads
break-even when w + k·r = k + 1
k = (w − 1) / (1 − r)
That is the whole thing. k is not “how often is this prompt used”. It is
reads per TTL window, which is a much harsher denominator. A prefix reused
forty times a day across a five-minute TTL may never be read twice inside one
window, in which case every call pays a write premium and receives no discount.
| Prompt shape | Write w | Read r | Break-even k | Measured reads / window | Verdict |
|---|---|---|---|---|---|
| Shared system prompt, short TTL | 1.25× | 0.10× | 0.28 | 14.2 | Cache it |
| Per-tenant document context, short TTL | 1.25× | 0.10× | 0.28 | 1.9 | Cache it |
| Per-tenant context, long TTL | 2.00× | 0.10× | 1.11 | 1.9 | Marginal. Measure again |
| Per-conversation history prefix | 1.25× | 0.10× | 0.28 | 0.2 | Paying the write premium for nothing |
The bottom row is the case that costs real money quietly. It raises no error, it shows up in no dashboard, and the only signal is the ratio of cache-read tokens to cache-creation tokens. Watch that ratio over deploys, not once: a cache-key refactor that folds a build hash into the prefix destroys the hit rate on the release that ships it, and the bill moves a fortnight before anyone connects the two.
Routing priced on resolved tasks, not on tokens
Cost per token is the wrong denominator and it recommends the wrong model with great consistency. The right denominator is a resolved task: one unit of work the user actually wanted, including every retry it took and the human who finished it when the model could not.
| Option | Token cost / attempt | Mean attempts | Unresolved | Escalation cost | Cost / resolved task |
|---|---|---|---|---|---|
| Small model only | $0.011 | 2.3 | 12.0% | $1.080 | $1.105 |
| Mid model only | $0.043 | 1.4 | 5.0% | $0.450 | $0.510 |
| Large model only | $0.180 | 1.1 | 1.5% | $0.135 | $0.333 |
| Small, escalating to large on failure | $0.011 | 1.6 + 0.3 | 1.5% | $0.135 | $0.212 |
The small model is roughly sixteen times cheaper per token than the large one and more than three times more expensive per resolved task. Every routing decision priced on the first number is wrong, and the arithmetic that exposes it needs exactly two application-level facts no provider has: attempts per task, and what an unresolved task costs you.
Two conditions on this table before anybody acts on it. Routing changes run against your existing eval suite, not against an opinion. If there is no eval suite, that is a prerequisite you own, and it should be said on the first call rather than in week two. And the escalation cost has to be a real loaded figure from your support organisation; guessing it moves the winner.
GPU-hours, measured honestly
For self-hosted inference and training, the equivalent trap is the utilization metric everyone screenshots.
DCGM_FI_DEV_GPU_UTIL reports whether any kernel was resident on the device
during the sampling interval. It reads near 100% for a single small kernel on an
H100 with 95% of its streaming multiprocessors idle. It is close to useless as a
waste signal and it is the number most often produced to prove a cluster is
busy.
The two that carry information are DCGM_FI_PROF_SM_ACTIVE, which reports
streaming-multiprocessor occupancy, and DCGM_FI_DEV_FB_USED, which reports
framebuffer in use. Together they show the headroom that actually exists.
The dollar figure is then a subtraction, not a percentage:
idle GPU dollars
= (allocated GPU-hours − SM-active-weighted GPU-hours) × your node rate
where SM-active-weighted GPU-hours
= Σ over intervals ( interval_hours × DCGM_FI_PROF_SM_ACTIVE )
The output is dollars with a derivation attached, priced at your real node rate rather than at list. It splits into named, separately owned modes:
| Mode | Signal | Owner |
|---|---|---|
| Idle nodes held by a stale selector | Allocated, zero SM-active, zero pods | Platform |
| Over-requested pods | 8 GPUs requested, one device with SM-active above noise | The owning team |
| MIG or time-slicing headroom | Small models on whole devices, framebuffer far below capacity | Platform |
| Jobs holding capacity after completion | SM-active at zero with the pod still Running | The owning team |
| Orphaned checkpoints and block storage | Volumes with no pod, from runs that ended | The owning team |
| Dev keys and dev endpoints on production | Cost in the production project with no production traffic | Whoever issued the key |
The inference waste catalogue, ranked
Across everything above, the recurring modes in rough order of how much money they have been worth. The ordering is our judgement, not a measurement.
| # | Waste mode | Mechanism | What detects it |
|---|---|---|---|
| 1 | Oversized PTU or provisioned throughput | Reserved capacity billed at a peak that arrives twice a month | Capacity cost ÷ tokens served, plotted by day |
| 2 | Retry storms | Failed tool calls retried inside an agent loop, each attempt billed | Attempts ÷ resolved tasks, per tenant and per feature |
| 3 | Unbounded agent fan-out | One user action expanding into dozens of sub-calls with no ceiling | Calls per user action, p99 rather than mean |
| 4 | Full-corpus re-embedding | Index key includes the build hash, so every deploy re-embeds everything | Embedding token volume correlated with deploy events |
| 5 | Cache writes never read | Prefix shape below the break-even reuse count | Cache-read ÷ cache-creation token ratio, over deploys |
| 6 | Streaming reconnects billed twice | Client reconnects mid-stream; the provider counts a second generation | Provider token total exceeding your gateway total |
| 7 | Undeleted Vertex endpoints | Online prediction endpoints billing per node-hour at zero traffic | Endpoint cost with no request-log rows |
| 8 | Dev keys on the production project | Non-production traffic inside the production cost centre | Cost by key with no matching production tenant |
Six of the eight are only visible in a join between your logs and the provider’s records. That is the practical case for building the seam before anyone asks for cost per customer: the seam is also the detector.
The join, in SQL
The shape is always the same: aggregate your own logs to the grain you want to bill at, aggregate the provider’s records to the same time grain, then price and reconcile. Roughly:
-- 1. your logs: the only place tenant exists
with app as (
select
date_trunc('day', request_ts) as usage_day,
tenant_id,
feature,
model_id,
cache_class, -- read | write | none
count(*) as attempts,
count(distinct task_id) as resolved_tasks,
sum(input_tokens) as input_tokens,
sum(output_tokens) as output_tokens
from request_log
where request_ts >= current_date - interval '90' day
group by 1,2,3,4,5
),
-- 2. the provider: the authoritative total, no tenant column
provider as (
select
date_trunc('day', invocation_ts) as usage_day,
model_id,
sum(input_token_count) as input_tokens,
sum(output_token_count) as output_tokens,
sum(cache_read_input_token_count) as cache_read_tokens
from bedrock_invocation_log
group by 1,2
)
-- 3. the reconciliation, printed before any allocation is trusted
select
a.usage_day,
a.model_id,
a.input_tokens as app_input,
p.input_tokens as provider_input,
a.input_tokens - p.input_tokens as residual_tokens,
abs(a.input_tokens - p.input_tokens)
/ nullif(p.input_tokens, 0) as residual_share
from app a
full outer join provider p
on a.usage_day = p.usage_day
and a.model_id = p.model_id
group by 1,2,3,4,5,6
order by residual_share desc;
The reconciliation query runs first and prints loudly. If your logs and the provider’s records disagree by more than a few percent on total tokens, every per-tenant number downstream is decoration. Common causes, in the order we find them: streamed responses counted once in the app and twice at the provider after a client reconnect, background jobs bypassing the gateway, a dev key billing to the production project, and usage-API reporting lag of several hours to a day that makes the last bucket look wrong when it is simply incomplete.
Error bars are part of the deliverable
A cost-per-customer figure quoted to the cent off a lagging API is a fabrication. The residual sources we report by name, every time: provider token rounding, usage-API reporting lag of hours to a day, mid-window model price changes, and the shared-index amortisation choice. Our working tolerance is within 3% of the invoice, and the residual gets reported rather than distributed silently across tenants to make the table add up.
Distributing the residual is the single most common way a unit-cost model becomes untrustworthy, because it hides the one number that tells you whether to believe the rest. The general argument is in a cost model that ties out exactly is hiding something.
Why this is a market gap and not just a hard afternoon
Organisations tracking AI cost went from 31% in 2024 to 63% in 2025 to 98% in 2026, and granular monitoring of AI spend (tokens, LLM requests, GPU utilisation) is the single most-requested capability in the 2026 practitioner survey.1 Inference is now roughly 80% of AI GPU spend industry-wide2 MEDIUM CONFIDENCE: secondary source
Demand is not the constraint. The constraint is that the last mile of this problem is inside the customer’s application, which is exactly the part a vendor cannot ship. A tool can offer you a gateway, a schema and a rate card. It cannot decide what a tenant is in your data model, which of your three background job runners should be billed to whom, or whether a retry storm is the customer’s fault or yours.
What would change our mind
One thing, and we watch for it quarterly: if every major provider accepted an arbitrary caller-supplied metadata dimension on each request and reported cost broken down by it, this work collapses from an engagement into a console checkbox. Bedrock application inference profiles are already most of the way there on one cloud. Anthropic and OpenAI per-workspace cost reporting is genuinely sufficient today if your key hygiene was designed correctly from the start. Saying so is more useful to you than pretending otherwise.
The part that survives is not measurement. It is the decisions measurement enables: routing policy, cache architecture, capacity shape, and pre-launch cost modelling for a feature that does not exist yet. Those are architecture judgments, and no console is close to them.
Where this number is weakest
The GPU serving line in the table above. Allocating GPU-hours to a tenant assumes the tenant had a namespace of its own, which is true in the synthetic example and often false in reality. On a shared inference deployment, GPU cost per tenant is an allocation rule with error bars, not a measurement, and it should be labelled that way in whatever the board eventually reads.
Cost per tenant, from the /methodology derivation
The methodology page derives one month of a multi-tenant AI application line by line against an 82,617.00 USD invoice. This is the number that derivation was for.
| Tenant | Attributed cost USD/mo | Requests, millions | Cost / 1,000 req USD | Share of invoice |
|---|---|---|---|---|
| Tenant A | 26,344.49 | 4.12 | 6.39 | 31.9% |
| Tenant B | 20,649.62 | 3.36 | 6.15 | 25.0% |
| Tenant C | 10,394.42 | 1.81 | 5.74 | 12.6% |
| Named tenants | 57,388.53 | 9.29 | 6.18 | 69.5% |
Prompt caching break-even on one prompt shape
Relocated from the Inference Cost Controls page.
| Line | Figure |
|---|---|
| Cache write premium over base input | +25% |
| Cache read discount against base input | −90% |
| Reuse count at which the shape turns profitable | 2 |
| Shapes currently writing and never reading | 11 |
| Monthly loss on those 11 shapes | 3,410 |
| Break-even reuse count, this shape | 2 reads |
Sources
- State of FinOps, AI cost tracking and top requested capability: data.finops.org, summarised at nops.io/blog/state-of-finops-2026. High confidence.
- Inference share of AI GPU spend: spheron.network. MEDIUM CONFIDENCE: secondary source, not a primary survey.
Sources are listed at the foot of this note with their confidence stated. Medium and low confidence figures say so in the copy.
Every query here runs against a read-only role you create, scope and revoke. Access policy.
Close
More notes
The rung this argument belongs to
Inference Cost Controls · Fixed fee
Check the arithmetic yourself
The full derivations, with the queries written out so they run under your own read-only credentials, in your own console.