Don't use a frontier model for a non-frontier problem.
Most AI bills are one mistake repeated a million times: the biggest model answering the smallest question. ZeroH sends every job to the cheapest model that can actually do it, meters the run to the cent, and stops at the cap you set.
The board your leaders watch. Now read down the model column.
Twelve jobs, six models, one policy. The frontier model turns up twice, on the two jobs that earned it. Everything else runs on a small hosted model, a model tuned on your own corpus, or open weights on a GPU you rent by the hour, under the same governance, row for row.
The job picks the model. Not the person, and not the vendor.
Nobody in your company should have to know what a token costs. The routing happens before the first one is spent, from what the job already declared about itself.
Draft the LinkedIn week
- claude-haiku-4.5
- 4,180 tokens
- no escalation
- €0.03 to Marketing
The job says what it is
Not the prompt, the job. Difficulty, data class, deadline, and the budget it has left this month, all of it declared when the job was hired, not guessed at runtime.
Policy picks the lane
The router takes the cheapest lane that clears the bar for that job, inside the ceiling its owner set. A masked-governed job never leaves the on-soil lane, whatever it would cost.
The answer is checked
Every lane runs against the job’s own eval set. A run that fails the bar escalates exactly one tier and re-runs, and both attempts are on the receipt.
The receipt carries the price
Model, tokens, lane, escalation, cost to the cent. Metered per run, attributed to the agent, the job, and the department that owns the budget.
The new joiner on open weights. The designer on the good model.
A model policy is a line of governance, not a config file: who, doing which job, on which lane, with what ceiling. Written once by an owner, changed with an approval, and every change is an audit line.
the balance is the whole point · pay frontier prices only where a frontier answer changes the outcome · ceilings and lanes here are illustrative
The cheapest model that still passes.
Every job carries its own eval set, built from the work it has already done and from what people accepted, sent back, or rewrote. It re-runs against the lane below and the lane above every time a new model ships, so the routing keeps getting cheaper without anyone re-tuning it by hand.
Move Steve’s contract reads up to the frontier lane.
Steve, your Compliance Agent, had 3 of his last 14 redline reads rewritten by a human. Replayed on the frontier lane, all 14 pass the job’s eval set.
Run the LinkedIn week on nemotron-3-nano.
Across the last 40 drafts, nano and haiku scored the same on the brand-voice eval (94% accepted) and no editor could tell them apart in the blind read.
nemotron-3-super now clears the finance eval.
A new open-weight model shipped. It was replayed against 120 finished finance jobs before anyone was asked to approve anything.
proposed by the workforce · ratified by a named owner · every ratification is a ledger event · figures illustrative
Your routing starts smart, then stops being ours.
Before one of your jobs runs, the router already has a defensible starting point, because we run a simulation plane: invented companies across industries, sizes and working styles, whose seats do realistic multi-step work against the same governed spine you run. Every simulated action is a real governed turn, so what it teaches is real, and none of it is anyone's data. Then the same loop runs on your side, on your jobs, and takes over.
Lane defaults per job type
Which work a small model handles well, and which work only the frontier gets right, decided from simulated runs before it is ever decided from yours.
Policy presets per role
The starting ceilings and lanes for a new joiner, a designer, a compliance officer. You inherit them on day one and change what you disagree with.
Eval baselines
A first acceptance bar per job type, so the router has something to measure against before your own history exists.
The failure drills
What actually happens at a hard cap, a failed escalation, a refused tool call, rehearsed on simulated orgs so your first one is not the rehearsal.
You inherit the defaults
Lanes, ceilings and acceptance bars that came out of simulated orgs shaped like yours. Nothing about your work is known yet, and nothing about it is assumed.
Your own jobs start voting
Every accepted, edited or sent-back result is evidence. Your finished jobs get replayed against every lane you allow, and the first proposals arrive for a named owner to ratify.
The routing is yours
Your lanes move on your evidence alone, your writing, your reviewers, your risk appetite. The simulated defaults are gone, and every change that replaced them is a ledger event with a name on it.
no customer data is used, or needed, to build the defaults · what your own jobs teach stays on your soil and moves only your lanes, it is never pooled, and it never becomes another tenant's starting point
Rent one open model. Serve hundreds of jobs on it.
NVIDIA's Nemotron 3 family is built for exactly this: a mixture-of-experts design where only about a tenth of the parameters fire per token, so a serious model serves a whole company off one GPU. Open weights also means it runs where your data already lives, nothing leaves, and there is no per-token bill at all.
Nemotron 3 Nano
Dec 2025The day-to-day workhorse: questions, summaries, meeting notes, first drafts.
Nemotron 3 Super
Mar 2026Analysis and structured reads: screening, anomaly reads, long documents.
Nemotron 3 Ultra
Jun 2026The open frontier: coding, migrations, long-horizon reasoning.
Nemotron 3 Nano Omni
Apr 2026Scans, forms, screenshots and call audio, without a second vendor.
parameter counts, context windows and ship dates are NVIDIA's published figures for Nemotron 3 · we serve the weights, we do not restate the benchmarks
Three ways to put a GPU under it.
Rent it by the hour, rent someone else's serving stack, or buy the box outright. The lane above does not change when you switch, only the bill does.
Azure GPU in your own subscription
The model runs in the same subscription and region as the rest of your ZeroH deployment. Same soil, same keys, same ledger. Scale the SKU up for a busy month, down for a quiet one.
NVIDIA NIM or a Hugging Face endpoint
Someone else operates the serving stack; you point the open lane at the endpoint. Fastest route from “we should try open weights” to a first token, with no ops on your side.
Buy the box
A DGX Spark-class desk machine, or a rack if the volume earns it. No per-token bill at all after the purchase, and the data never leaves the room. You carry the operations, we run one ourselves, so we will tell you honestly what that costs.
One model, hundreds of jobs, until it isn't.
The same metering that prices a job also watches the box it ran on. When throughput starts slipping, auto-routing protects the SLA that same minute, and the platform tells you what it would take to fix properly, priced from your own numbers, not a vendor's.
The nano lane is out of headroom. While you decide, overflow auto-routes to the hosted small lane so nobody waits, about €0.004 a job, already showing on the meter. To fix it properly: add a second GPU (about €310 a month), or move the three heaviest job types up to the super lane on the same box.
Some jobs are cheaper on a seat than on a meter.
A multi-hour job, a repo migration, a deep research sweep, a whole-corpus rewrite, can burn more per call than a monthly subscription seat costs. So send it to the seat, inside a governed sandbox.
Priced per call, for one run of one job. Do it four more times this quarter and it is €340.
The same job, run by a governed sandbox against a team subscription, and the next twenty jobs are free.
Managed microVMs (E2B)
Seconds to a working sandbox, nothing to operate. Cleared for public-content work: repo audits, SEO sweeps, anything with no governed data in it.
Your VM + Firecracker
A warm pool of microVMs on a VM in your region, on your terms. This is the lane governed content is allowed to use.
The box you already bought
The same hardware serving the open lane runs sandboxes between jobs. Utilization earns the slot; the driver behind the seam does not change.
one sandbox seam, vendors are drivers: provision → inject masked inputs → run → collect → mask → audit → destroy · secrets never enter a sandbox · read the architecture note
A warning line, and a line nothing crosses.
Budgets are set per agent, per job, per department and per month. At the warning line the owner hears about it. At the hard cap the run is refused before a model is called, and the refusal is a ledger event like everything else.
The owner hears it first
80% of the cap, or a spend curve that will cross it before month end. The named owner gets it in Teams with the three jobs driving the burn, while there is still time to do something.
Refused, not overspent
The turn stops in the governed spine before the relay. No model call, no token, no charge. A cap is a policy decision, so crossing it takes an approval from the person who owns the budget.
The work waits for a decision
A capped job queues instead of vanishing: raise the budget, move it to a cheaper lane, or let it run next period. Whichever you choose is on the record with your name on it.