Pricing

Start with a pilot. Scale when it proves out.

Transparent, fixed-scope pricing — no "contact us for a quote" wall. Run a real model on your data in 72 hours, then take it to production only when the numbers make sense.

Best first step

Pilot

One model, live on your data in 72 hours.

$4,900

fixed price

  • 72-hour deployment
  • One open-weight model, your choice
  • OpenAI-compatible endpoint
  • Reference evaluation on your data
  • Readout + production scoping
Start the pilot

Production Deployment

Your full private inference stack, in production.

from $25,000

one-time · hardware separate

  • Model selection + GPU sizing
  • Optimization: quantization, batching, caching
  • Monitoring, RBAC, audit logging
  • OpenAI-compatible API, inside your network
  • Handover + documentation
Book a scoping call

Confidential Compute

Run on GPUs you don't own — unseen by the owner.

Custom

scoped to your workload

  • NVIDIA CC (H100 / H200 / Blackwell)
  • Encrypted VRAM + CPU TEE
  • Hardware-verified attestation
  • For regulated data & borrowed hardware
  • How it works →
Talk to us

Keep it running — Managed Ops

We run the whole thing: model and driver updates, monitoring, on-call, and SLA tiers — so your team doesn't have to become an MLOps team. Add it to any deployment.

from $1,500

per GPU node / month

Add managed ops

Hardware is yours. Service pricing is separate from GPUs — you own the capex, so there's nothing to rent back from us. We size it and can source it at cost, no markup. As a ballpark, a single H100 (80GB) runs roughly $25–30k and a 4×H100 node roughly $180–240k; smaller models run happily on far less.

Not sure where you'll land?

Estimate your own cloud-vs-on-prem cost and break-even before you talk to anyone.

Open the TCO calculator

Pricing questions

What's included in the $4,900 pilot?+

A 72-hour deployment of one open-weight model on your own or rented GPUs, an OpenAI-compatible endpoint, a reference evaluation on your data, and a readout to scope production. Fixed price, no surprises.

Is hardware included?+

No — hardware is yours, and you own the capex. We advise on sizing and can source it at cost with no markup. Service pricing (pilot, deployment, managed) is separate from hardware.

Can I start with the pilot and scale later?+

Yes — that's the intended path. The pilot proves it on your data, production deployment (from $25,000) takes it live, and optional managed ops (from $1,500 per GPU node per month) keeps it running.

Ready to see it on your data?

Book a scoping call, or kick off the fixed-price pilot today.

Schedule a Discovery Call