- Pricing
- Fixed feequoted at the teardown debriefUSD · fixed
- Payment
- Half at signature, half on acceptance. Net 15.
- Prerequisite
- Spend Teardown, same estate
Rung
Inference Cost Controls
Routing, caching, capacity shape and per-request cost attribution, shipped behind flags.
Routing, caching and capacity shape are architecture work, so they lose to feature work every sprint. The build puts a seam in the request path, ships routing and caching behind flags, and writes the capacity decision down.
LLM providers · GPU
Delivered
What arrives
| Artifact | Lands in |
|---|---|
| Per-request cost attribution behind one gateway seam | your repo |
| Model routing policy in code, priced per resolved task | your repo |
| Prompt caching where break-even supports it, plus a monitor | your repo |
| Capacity-shape decision: on-demand, PTU, or self-hosted | written memo |
| CI gate commenting projected cost delta on pull requests | your repo |
| Budget and anomaly alerts from measured variance | your channel |
Not included Eval harness authoring (your eval suite) · Deep GPU serving work (separately quoted) · Terraform cost estimates (Infracost) · The migration itself (the memo, not the move)
Your side
| Block | Your hours | Needs |
|---|---|---|
| Week 1 | 5 | Repository and CI access, evals runnable. |
| Week 2 | 4 | Routing review, fallback-ladder decision. |
| Week 3 | 4 | Flag rollout with on-call, then handover. |
| Total | 13 |
Guarantee
Every change ships behind a flag with a measured before-and-after. Any change that misses its projected saving in the next billing period is reverted or refixed at no charge.
Cap20 hours, for 60 days after handover.