For teams fine-tuning, post-training, running RL, or training simulation and robotics policies — from first dedicated cluster to multi-site fleets. Reliability, time-to-train, and workload fit come first; cost follows from how the structure is built. (Frontier-scale pretraining appears as context; it's not a claim we advise those runs.)
Seven dimensions — each one decides something the market prices:
This is the spec we price against the market. Two of these lines do most of the work: cadence decides your contract structure, and preemption tolerance decides whether the cheapest capacity on the market is open to you at all. Most teams procure as if they were continuous when their actual cadence is campaign-shaped — and pay for the gap.
Two lines of that spec — single-run scale and coupling — decide which market you're in. The test is the run, not the word "training."
Physical AI. Simulation and robotics-policy training is typically campaign-shaped and highly parallel across many individually small jobs — node-sized procurement, with between-run utilization as the dominant economic question rather than topology.
Training buys almost entirely on the hardware side of the market, and three structural facts shape it right now:
Rack increments are long-dated. If pretraining-scale work genuinely forces racks, the commitment is made long before the first run starts — a premium on being certain it's forced.
Burst instruments exist for campaign demand. Pricing a burst window against a standing commitment is the comparison most teams never run.
Spot depth is a training asset. For preemption-tolerant runs, checkpoint discipline turns an interruptible instrument into dependable capacity — and opens the deepest discounts on the market.
Our monthly market pulse — contract forms, tenor economics, generational cadence — is published on the compute procurement overview and applies here in full.
You set the cadence; we structure the contract against it.
Against campaign-shaped demand, we structure three things.
Multi-year commitments priced against runs that arrive in bursts is the default failure mode. The alternative has a market instrument — price the two against each other before either is signed.
The sizing decided which runs qualify; the structure decides how much of the fleet rides the cheap end of the market.
Hardware generations turn over faster than facility terms — the mismatch lands on your balance sheet. Stagger expiries; don't cliff-date them.
MFU divides your bill. The same run at 35% versus 55% realized MFU costs roughly 1.6x as much — that's division, not a benchmark. After MFU:
Unlike inference decode, training is genuinely compute-bound — FLOPS-per-dollar is a legitimate first-cut comparison here.
The grey is recoverable — and our first deliverable is making it a number you can state.
Depreciation, residuals, and financing.
Run cadence, realized MFU band, between-run idle, and commitment coverage.
Commitment tenor against your actual campaign rhythm; spot/reserved blend; backfill structure.
Runs recur — campaigns, renewals, backfill. A standing process on a systematic cadence, not a fire drill.
Cost per run and idle cost stated defensibly for boards, lenders, and diligence.
Optional: send your spec ahead — the seven lines above — and we'll come priced.