AI Budgeting

Don't use a frontier model for a non-frontier problem.

Most AI bills are one mistake repeated a million times: the biggest model answering the smallest question. ZeroH sends every job to the cheapest model that can actually do it, meters the run to the cent, and stops at the cap you set.

“what does our parental leave policy say?”one ordinary question, priced on three lanes
Frontier
€0.42
claude-opus-5
accepted 98%
Hosted small
€0.03
claude-haiku-4.5
accepted 98%
Open weights, your GPU
€0.004
nemotron-3-nano
accepted 98%
same answer · same acceptance · 100× the bill · illustrative figures, your ledger reports its own
Every job, metered

The board your leaders watch. Now read down the model column.

Twelve jobs, six models, one policy. The frontier model turns up twice, on the two jobs that earned it. Everything else runs on a small hosted model, a model tuned on your own corpus, or open weights on a GPU you rent by the hour, under the same governance, row for row.

sandboxedpolicy firstproof
Spend / mo
€982
↓ 54% after routing
Avg cost / job
€0.12
frontier-only: €1.06
Jobs on open weights
61%
one rented GPU
Against the cap
€982 / €1,500
hard cap, not a hope
JobRun byModelAcceptedCostReceipt
every row links to its signed ledger event · cost is the metered actual for that run, never an estimate · the model column is a policy outcome, not a default · numbers here are illustrative · your board reports its own
Auto-routing

The job picks the model. Not the person, and not the vendor.

Nobody in your company should have to know what a token costs. The routing happens before the first one is spent, from what the job already declared about itself.

the job, as it was hired

Draft the LinkedIn week

difficulty: routinedata class: internalbudget left: €188bar: 92% accepted
policy inputs, not a prompt, the router never reads the message
lanes it is allowed to use
Open weights · your GPUnemotron-3-nano€0.01under the brand-voice bar, for now
Hosted smallclaude-haiku-4.5€0.03policy pick
Frontierclaude-opus-5€0.34clears the bar, over the ceiling
↑ escalates one tier, once, only if the eval fails
receipt
  • claude-haiku-4.5
  • 4,180 tokens
  • no escalation
  • €0.03 to Marketing
signed · ev-3103
01

The job says what it is

Not the prompt, the job. Difficulty, data class, deadline, and the budget it has left this month, all of it declared when the job was hired, not guessed at runtime.

02

Policy picks the lane

The router takes the cheapest lane that clears the bar for that job, inside the ceiling its owner set. A masked-governed job never leaves the on-soil lane, whatever it would cost.

03

The answer is checked

Every lane runs against the job’s own eval set. A run that fails the bar escalates exactly one tier and re-runs, and both attempts are on the receipt.

04

The receipt carries the price

Model, tokens, lane, escalation, cost to the cent. Metered per run, attributed to the agent, the job, and the department that owns the budget.

Model policy

The new joiner on open weights. The designer on the good model.

A model policy is a line of governance, not a config file: who, doing which job, on which lane, with what ceiling. Written once by an owner, changed with an approval, and every change is an audit line.

WhoJobLaneWhy this laneCeiling
New joinerDay-to-day assistant: questions, summaries, first draftsopennemotron-3-nanoOn-soil open weightsFast, private, and effectively free per turn. Nobody’s first week needs a frontier model.€15 / mo
Marketing · designerCampaign concepts and art directionfrontierclaude-opus-5Rented, hostedTaste and originality are the deliverable. This is the job that earns the frontier price.€120 / mo
Marketing · content writerBlog and social from an approved briefsmallclaude-haiku-4.5Rented, hostedThe voice comes from the brand pack and the brief, not from model size.€40 / mo
Finance analystAnomaly reads across the ledgeropennemotron-3-superOn-soil open weightsStructured reasoning over numbers that are never allowed to leave the building.€60 / mo
ComplianceContract and control readsfrontierclaude-opus-5Rented, hostedOne missed clause costs more than a year of tokens.€200 / mo
Ask AliShariah answers from the governed corpustunedzeroh-ft-3Tuned, on-soilTuned on your own corpus, cited every time, and cheaper than any hosted model.€90 / mo

the balance is the whole point · pay frontier prices only where a frontier answer changes the outcome · ceilings and lanes here are illustrative

Evaluations

The cheapest model that still passes.

Every job carries its own eval set, built from the work it has already done and from what people accepted, sent back, or rewrote. It re-runs against the lane below and the lane above every time a new model ships, so the routing keeps getting cheaper without anyone re-tuning it by hand.

Draft the LinkedIn week40 finished jobs, replayed on every lane it is allowed to use
Lane belownemotron-3-nano94%€0.01
Runs here todayclaude-haiku-4.594%€0.03
Lane aboveclaude-opus-595%€0.34
The lane below scores the same for a third of the price. That becomes a proposal, not a change, a named owner ratifies it, and the decision is logged.
Upgrade

Move Steve’s contract reads up to the frontier lane.

Steve, your Compliance Agent, had 3 of his last 14 redline reads rewritten by a human. Replayed on the frontier lane, all 14 pass the job’s eval set.

+€0.55 / job · +€19 / mo
Downgrade

Run the LinkedIn week on nemotron-3-nano.

Across the last 40 drafts, nano and haiku scored the same on the brand-voice eval (94% accepted) and no editor could tell them apart in the blind read.

−€38 / mo, same acceptance
New model

nemotron-3-super now clears the finance eval.

A new open-weight model shipped. It was replayed against 120 finished finance jobs before anyone was asked to approve anything.

−€61 / mo at equal acceptance

proposed by the workforce · ratified by a named owner · every ratification is a ledger event · figures illustrative

Simulated orgs, then yours

Your routing starts smart, then stops being ours.

Before one of your jobs runs, the router already has a defensible starting point, because we run a simulation plane: invented companies across industries, sizes and working styles, whose seats do realistic multi-step work against the same governed spine you run. Every simulated action is a real governed turn, so what it teaches is real, and none of it is anyone's data. Then the same loop runs on your side, on your jobs, and takes over.

what the simulation varies
IndustryIslamic bankinsurerfintechfamily officeregulatorretail group
Company shape12 seats300 seats4,000 seatsone branchsix countries
Working stylethe cautious DPOthe impatient AEthe meticulous auditorthe CEO who writes in bullets
Loada normal Tuesdaythe Monday spikequarter-end closean audit week
Data classpublicinternalpersonalregulatedboard-only
every simulated action is a real governed turn · mask → relay → audit → rehydrate
what a new tenant inherits

Lane defaults per job type

Which work a small model handles well, and which work only the frontier gets right, decided from simulated runs before it is ever decided from yours.

Policy presets per role

The starting ceilings and lanes for a new joiner, a designer, a compliance officer. You inherit them on day one and change what you disagree with.

Eval baselines

A first acceptance bar per job type, so the router has something to measure against before your own history exists.

The failure drills

What actually happens at a hard cap, a failed escalation, a refused tool call, rehearsed on simulated orgs so your first one is not the rehearsal.

how much of the routing decision comes from your own evidence
Day one
all inherited

You inherit the defaults

Lanes, ceilings and acceptance bars that came out of simulated orgs shaped like yours. Nothing about your work is known yet, and nothing about it is assumed.

Week two
35% yours

Your own jobs start voting

Every accepted, edited or sent-back result is evidence. Your finished jobs get replayed against every lane you allow, and the first proposals arrive for a named owner to ratify.

Month three
all yours

The routing is yours

Your lanes move on your evidence alone, your writing, your reviewers, your risk appetite. The simulated defaults are gone, and every change that replaced them is a ledger event with a name on it.

no customer data is used, or needed, to build the defaults · what your own jobs teach stays on your soil and moves only your lanes, it is never pooled, and it never becomes another tenant's starting point

Open weights

Rent one open model. Serve hundreds of jobs on it.

NVIDIA's Nemotron 3 family is built for exactly this: a mixture-of-experts design where only about a tenth of the parameters fire per token, so a serious model serves a whole company off one GPU. Open weights also means it runs where your data already lives, nothing leaves, and there is no per-token bill at all.

Nemotron 3 Nano

Dec 2025
3.6B fire per token31.6B loaded
1M context

The day-to-day workhorse: questions, summaries, meeting notes, first drafts.

Nemotron 3 Super

Mar 2026
12.7B fire per token120B loaded
1M context

Analysis and structured reads: screening, anomaly reads, long documents.

Nemotron 3 Ultra

Jun 2026
55B fire per token550B loaded
1M context

The open frontier: coding, migrations, long-horizon reasoning.

Nemotron 3 Nano Omni

Apr 2026
visionaudiolanguage

Scans, forms, screenshots and call audio, without a second vendor.

parameter counts, context windows and ship dates are NVIDIA's published figures for Nemotron 3 · we serve the weights, we do not restate the benchmarks

Where the GPU comes from

Three ways to put a GPU under it.

Rent it by the hour, rent someone else's serving stack, or buy the box outright. The lane above does not change when you switch, only the bill does.

1

Azure GPU in your own subscription

Rented · hourly · your tenant

The model runs in the same subscription and region as the rest of your ZeroH deployment. Same soil, same keys, same ledger. Scale the SKU up for a busy month, down for a quiet one.

Best when you already run Offer 2 on your Azure.
2

NVIDIA NIM or a Hugging Face endpoint

Rented · packaged inference · per hour

Someone else operates the serving stack; you point the open lane at the endpoint. Fastest route from “we should try open weights” to a first token, with no ops on your side.

Best for the pilot, and for bursts you do not want to own.
3

Buy the box

Owned · capex once · fully on soil

A DGX Spark-class desk machine, or a rack if the volume earns it. No per-token bill at all after the purchase, and the data never leaves the room. You carry the operations, we run one ourselves, so we will tell you honestly what that costs.

Best at steady volume, or when the data class allows nothing else.
Capacity

One model, hundreds of jobs, until it isn't.

The same metering that prices a job also watches the box it ran on. When throughput starts slipping, auto-routing protects the SLA that same minute, and the platform tells you what it would take to fix properly, priced from your own numbers, not a vendor's.

GPU utilization · nano lane
83% up from 61% last month
Queue wait · p95
12.4s above the 8s job SLA for 3 days
Jobs served by this one model
2,510 / mo 61% of everything the company ran
what the platform is telling you

The nano lane is out of headroom. While you decide, overflow auto-routes to the hosted small lane so nobody waits, about €0.004 a job, already showing on the meter. To fix it properly: add a second GPU (about €310 a month), or move the three heaviest job types up to the super lane on the same box.

raised by the workforce · ratified by a named owner · both options priced from your own ledger
Flat-rate lane

Some jobs are cheaper on a seat than on a meter.

A multi-hour job, a repo migration, a deep research sweep, a whole-corpus rewrite, can burn more per call than a monthly subscription seat costs. So send it to the seat, inside a governed sandbox.

one job · migrate a 40k-line service · ~6 hours of model work
metered, per token
€68

Priced per call, for one run of one job. Do it four more times this quarter and it is €340.

vs
a flat seat, driven in a sandbox
€22 / mo

The same job, run by a governed sandbox against a team subscription, and the next twenty jobs are free.

the platform prices both lanes before it starts · the cheaper one wins unless the data class forbids it
masked inputs inshort-lived scoped tokens, minted host-sideallowlisted egressthe work leaves as a pull request, never a direct push
Rented · managed

Managed microVMs (E2B)

Seconds to a working sandbox, nothing to operate. Cleared for public-content work: repo audits, SEO sweeps, anything with no governed data in it.

Rented · self-run

Your VM + Firecracker

A warm pool of microVMs on a VM in your region, on your terms. This is the lane governed content is allowed to use.

Owned · on soil

The box you already bought

The same hardware serving the open lane runs sandboxes between jobs. Utilization earns the slot; the driver behind the seam does not change.

one sandbox seam, vendors are drivers: provision → inject masked inputs → run → collect → mask → audit → destroy · secrets never enter a sandbox · read the architecture note

Guardrails

A warning line, and a line nothing crosses.

Budgets are set per agent, per job, per department and per month. At the warning line the owner hears about it. At the hard cap the run is refused before a model is called, and the refusal is a ledger event like everything else.

Company · Julywithin budget
€982 of €1,500on track, 7 days left
Marketingwithin budget
€412 of €600warning line at €480
Financewarning line crossed
€238 of €250owner notified at 80%
Frontier lanehard cap reached
€300 of €300capped, new frontier requests fall back to the small lane
At the warning line

The owner hears it first

80% of the cap, or a spend curve that will cross it before month end. The named owner gets it in Teams with the three jobs driving the burn, while there is still time to do something.

At the hard cap

Refused, not overspent

The turn stops in the governed spine before the relay. No model call, no token, no charge. A cap is a policy decision, so crossing it takes an approval from the person who owns the budget.

Held, not lost

The work waits for a decision

A capped job queues instead of vanishing: raise the budget, move it to a cheaper lane, or let it run next period. Whichever you choose is on the record with your name on it.

Bring your last AI invoice. We will show you the routed version.