We work with inference and agentic platform teams to access the compute market efficiently — reading supply as it stands, and structuring commitments that map tightly to the technical requirement.
Seven dimensions — each one decides something the market prices:
This is the spec we price against the market. We don't set these numbers — you do. What we add is what they cost: what does each 100ms of p99 budget you refuse to spend cost you per year in foregone batching gain? Most teams can't answer that today — and it's a finance question, not an engineering one.
Two lines of that spec — fits one node, and coupling — decide which market you're in.
What one job needs — not your aggregate volume, not your company's size — sets the smallest unit of compute you can buy.
And that's partly an engineering choice: every leftward move on the map takes a workload out of rack-sized purchases and into node-sized ones.
And the one lever that isn't free: batching buys throughput and spends tail latency — which is why the last chart in this sequence needs your p99 SLO on it.
Batching is the cheapest lever you have, and it spends tail latency to buy throughput. Your p99 SLO binds while real throughput gain is still on the table — the shaded region isn't waste; it's gain your SLO forecloses. That's the number the requirement asked you to price.
FLOPS is a theoretical ceiling, and a fair proxy when a workload is compute-bound. Decode — the phase that dominates most serving bills — is memory-bandwidth-bound: each generated token re-reads the model's weights and that request's KV cache from memory, and the arithmetic per token is small relative to the data moved. Two parts with identical peak FLOPS can produce materially different cost per token.
Illustrative: a 70B-parameter model in FP16 holds ~140GB of weights. At batch size 1 on ~3.35TB/s of memory bandwidth, every token requires re-reading those weights: ~42ms per token, ~24 tokens/sec — with the compute units mostly idle. At batch size 32, the same weight-read serves 32 requests at once, and throughput rises toward ~700–800 tokens/sec. The lever was batching, not FLOPS.
The supply side is dynamic and capacity remains genuinely constrained. Four structural facts decide what your requirement can clear against — and each one shapes the commitment you should sign:
Rack-scale systems deploy as single 120kW+ liquid-cooled units from a short list of qualified hosts; node-scale capacity is sold across the whole market in small steps. Why it matters: your increment regulates how much of the market is open to you at all.
Commitments are denominated in tokens — a guaranteed rate, the provider owns the hardware — or in hardware, where you do. Why it matters: the form decides who carries the cost of idle capacity, before any price comparison starts.
Sub-six-month reserved terms are broadly unavailable; discounts concentrate at 1–2 year commitments, with prepayment norms beyond a year. Why it matters: tenor committed is a price lever — and any discount secured should reflect the lock-in and potential idle-capacity risk.
Hardware generations turn over faster than the contract terms you sign, and residual-value assumptions vary widely across operators. Why it matters: the divergence between the accelerating release of new hardware generations and the assumed residual value of your contracted hardware is reflected on your balance sheet.
The current read on all four — our monthly Market pulse — is published on the compute procurement overview.
A short list of hosts · rack-sized commitments · longer terms
The whole market · small, reversible steps
This is why sizing comes first: your workload profile picks your market — and your market picks the commitments you can sign.
Your requirement sets the target; we secure and optimize the structure that delivers it.
Provisioned throughput — sold in provider-defined units; Azure's "PTUs" (provisioned throughput units) popularized the form, and it has since spread across the market — means committing to a guaranteed token rate while the provider owns the hardware behind it.
| Rung | Typical commitment | You're buying | You're giving up |
|---|---|---|---|
| Serverless / shared API | None | Zero capacity risk, instant elasticity | Highest unit price, no capacity guarantee |
| Provisioned throughput (PTU-style) | Monthly – 1 year | A guaranteed token rate; the provider owns the hardware problem | Pay for the reserved rate whether you consume it or not — a pure utilization break-even |
Lower unit price as you descend either ladder; more of the idle transfers to you.
| Rung | Typical commitment | You're buying | You're giving up |
|---|---|---|---|
| Spot / preemptible | None | The lowest prices available, often by a wide margin | Interruption at the provider's discretion — only viable if the workload survives preemption |
| On-demand instances | None to hourly | Flexibility, fast experimentation | Premium pricing, no guaranteed availability |
| Reserved capacity (GPU-hours) | 6 months – 3 years | Materially lower unit price, guaranteed capacity | A locked demand assumption; over-commitment risk |
| Dedicated bare metal | Multi-year | Isolation, performance predictability, full topology control | Longer lock-in, more operational responsibility |
| Owned hardware in colocation | Capex + facility term | Lowest marginal cost at high sustained utilization | Capital, depreciation risk, obsolescence exposure |
The rungs are the market's public taxonomy. The priced version of this ladder — current tenor curves, provider terms, and what actually moves in negotiation — is client material.
What we structure, at your commitment size:
For a first commitment, the useful question isn't "what's the cheapest rung." It's "what's the smallest commitment that clears my SLOs — and what does it cost me to be wrong?"
The provider invoice is not the cost structure.
Utilization is the multiplier.
We build the numbers diligence underwrites.
Prices four sourcing rungs per successful task, not per token — every contested assumption is a switch on the result card, and the self-host rung shows as a band until your inputs narrow it. If a calculator hides its contested assumptions, distrust it.
Built on the CPST model with full calibration provenance.
Your SLOs, your traffic profile, and — if you hold capacity — realized cost-to-serve and consumed-versus-allocated utilization against what you've contracted.
Commitment structure, tenor laddering, and workload shaping — built against the market read.
Compute needs recur — renewals, expansions, rebalancing. We run procurement as a standing process, on a systematic cadence rather than as a fire drill.
Unit economics your board, lenders, and diligence can rely on, stated defensibly.
Spruce Street can support you at any stage, from inception to enterprise-grade maturity — whether you are:
Signing your first compute commitment: We bring the market intelligence and price the trade-offs on either side of it: committing too little to support growth and burstiness, or locking in excess capacity at rates the market may later undercut.
Scaling through a step-change: When your compute requirement jumps by an order of magnitude, we provide an independent perspective and the tools to help you optimize the structure.
Operating a substantial existing compute fleet: We work alongside your teams — freeing up their bandwidth and adding an edge across the process: the market read, renewal and renegotiation timing, invoice and contract reconciliation, and diligence-ready reporting.