Pricing
Fixed feequoted at the teardown debriefUSD · fixed
Payment
Half at signature, half on acceptance. Net 15.
Prerequisite
Spend Teardown, same estate

Rung

Inference Cost Controls

Routing, caching, capacity shape and per-request cost attribution, shipped behind flags.

Routing, caching and capacity shape are architecture work, so they lose to feature work every sprint. The build puts a seam in the request path, ships routing and caching behind flags, and writes the capacity decision down.

3%Tolerance

LLM providers · GPU

Delivered

What arrives

ArtifactLands in
Per-request cost attribution behind one gateway seamyour repo
Model routing policy in code, priced per resolved taskyour repo
Prompt caching where break-even supports it, plus a monitoryour repo
Capacity-shape decision: on-demand, PTU, or self-hostedwritten memo
CI gate commenting projected cost delta on pull requestsyour repo
Budget and anomaly alerts from measured varianceyour channel

Not included Eval harness authoring (your eval suite) · Deep GPU serving work (separately quoted) · Terraform cost estimates (Infracost) · The migration itself (the memo, not the move)

Your side

BlockYour hoursNeeds
Week 15Repository and CI access, evals runnable.
Week 24Routing review, fallback-ladder decision.
Week 34Flag rollout with on-call, then handover.
Total13

The boundary · the arithmetic

Guarantee

Every change ships behind a flag with a measured before-and-after. Any change that misses its projected saving in the next billing period is reverted or refixed at no charge.

Cap20 hours, for 60 days after handover.

The exact terms.

Write the emailSpend Teardown firstAll four rungs