Compute Procurement — Training & Post-Training

Securing the compute structure your training runs require

For teams fine-tuning, post-training, running RL, or training simulation and robotics policies — from first dedicated cluster to multi-site fleets. Reliability, time-to-train, and workload fit come first; cost follows from how the structure is built. (Frontier-scale pretraining appears as context; it's not a claim we advise those runs.)

The metric that matters here: cost per completed run, at the time-to-train your research velocity requires.
1

Sizing your requirement

Seven dimensions — each one decides something the market prices:

Single-run scale & couplingone run's footprint; gradient all-reduce, expert routing
decides node or rack — which market you're in
Run cadencecampaign-shaped or continuous
decides contract form — standing commitment or burst instruments
Time-to-train floorthe wall-clock time your research velocity tolerates
decides cluster size and interconnect class
Preemption tolerancewhich runs survive interruption
decides your spot share — the cheapest capacity on the market
Checkpoint cadence & failure rateoverhead at scale
decides your effective cost per run
Realized MFUif you hold capacity: fraction of theoretical peak achieved
decides how much fleet your runs actually need
Placement & timingwhere the cluster must sit; when it must be live
decides site options — against 40+ week rack lead times

This is the spec we price against the market. Two of these lines do most of the work: cadence decides your contract structure, and preemption tolerance decides whether the cheapest capacity on the market is open to you at all. Most teams procure as if they were continuous when their actual cadence is campaign-shaped — and pay for the gap.

Two lines of that spec — single-run scale and coupling — decide which market you're in. The test is the run, not the word "training."

SINGLE-RUN SCALE What actually forces rack-scale — and what doesn't A run crosses only if BOTH hold: it exceeds one node's memory, and its pieces must talk constantly — gradient all-reduce, expert routing. Node-scale buy in small, reversible, node-sized steps LoRA & adapter fine-tuning Post-training & preference tuning Most RL experimentation Simulation & robotics policy training Ablations & evaluation runs many parallel, individually small jobs Rack-scale coherent memory across the rack Large-scale pretraining Frontier MoE training 120kW+ liquid-cooled units, a short list of qualified hosts, longer terms, lumpy increments Placements are categorical, not exhaustive. The test is the run, not the word "training." Everything on the left buys in reversible steps — training doesn't mean racks.
The classification for training work: what crosses into rack territory, and what never needed to.

Physical AI. Simulation and robotics-policy training is typically campaign-shaped and highly parallel across many individually small jobs — node-sized procurement, with between-run utilization as the dominant economic question rather than topology.

2

Reading the market

Training buys almost entirely on the hardware side of the market, and three structural facts shape it right now:

Rack increments are long-dated. If pretraining-scale work genuinely forces racks, the commitment is made long before the first run starts — a premium on being certain it's forced.

Burst instruments exist for campaign demand. Pricing a burst window against a standing commitment is the comparison most teams never run.

Spot depth is a training asset. For preemption-tolerant runs, checkpoint discipline turns an interruptible instrument into dependable capacity — and opens the deepest discounts on the market.

Our monthly market pulse — contract forms, tenor economics, generational cadence — is published on the compute procurement overview and applies here in full.

3

Structuring the commitment

You set the cadence; we structure the contract against it.

  • Denomination first — commitments are in tokens or in hardware; training lives on the hardware ladder: spot, on-demand, reserved, dedicated, owned
  • Every step down the ladder — a lower unit price, traded for ownership of the idle

Against campaign-shaped demand, we structure three things.

Burst windows vs. standing reserves

Multi-year commitments priced against runs that arrive in bursts is the default failure mode. The alternative has a market instrument — price the two against each other before either is signed.

Spot blending

The sizing decided which runs qualify; the structure decides how much of the fleet rides the cheap end of the market.

Tenor laddering

Hardware generations turn over faster than facility terms — the mismatch lands on your balance sheet. Stagger expiries; don't cliff-date them.

4

Optimizing the cost structure

MFU divides your bill. The same run at 35% versus 55% realized MFU costs roughly 1.6x as much — that's division, not a benchmark. After MFU:

  • Sequence length — attention cost grows super-linearly with context
  • Parallelism split efficiency — a mismatched split leaves the fleet waiting on the network
  • Checkpoint & failure overhead — a material share of effective cost per run at scale

Unlike inference decode, training is genuinely compute-bound — FLOPS-per-dollar is a legitimate first-cut comparison here.

RUN CADENCE Committed capacity bills continuously. Runs don't. What you pay for — committed capacity, billed every hour of the term Run 1 Run 2 Run 3 abl. What actually runs idle — paid, unused backfill: inference · spot The gaps are recoverable — internal experiments, preemption-tolerant work, inference serving between runs. Week 0 Week 12 one quarter, illustrative Cadence and run lengths are illustrative — the structure is the point. Most teams procure as if continuous when their cadence is campaign-shaped — the grey is the cost.
Between-run idle is the largest unexamined line in most training budgets — and most teams cannot state theirs.

The grey is recoverable — and our first deliverable is making it a number you can state.

  • Schedule compaction — close the gaps before buying around them
  • Internal experimentation — idle capacity opened to research queues
  • Spot blending — for the preemption-tolerant share of runs
  • Backfill with inference serving — often the single highest-value structural change for teams doing both

Depreciation, residuals, and financing.

  • A live dispersion — assumptions vary widely across operators; not a settled number
  • We make yours defensible — to auditors, lenders, and diligence, across a mixed-generation fleet
Spruce
Street

Where we plug in

1

Evaluate

Run cadence, realized MFU band, between-run idle, and commitment coverage.

2

Configure

Commitment tenor against your actual campaign rhythm; spot/reserved blend; backfill structure.

3

Operate

Runs recur — campaigns, renewals, backfill. A standing process on a systematic cadence, not a fire drill.

4

Report

Cost per run and idle cost stated defensibly for boards, lenders, and diligence.

Find out what your gaps cost.

Book a training compute review

Optional: send your spec ahead — the seven lines above — and we'll come priced.