A gateway sees the front door

An API gateway knows /api/chat was called. It can count that, rate-limit it, even reject it before it runs — some gateways do real pre-call checks. What none of them have is what happens after: whether that one request becomes one model call or nine, whether a tool gets retried, whether an agent loops.

The problem isn't that gateways can't check anything in advance. It's that a request count was never the unit that mattered. One request doesn't mean one model call, one cost, or one predictable amount of work.

Why request counts stop working

For a typical API, request counts are a fair proxy for cost: each one does roughly the same work. AI breaks that. Two calls to the same endpoint can differ by 10x in cost, depending on the model, the input and output size, how many tool calls or retries ran, and which provider handled it.

That's how a product ends up with real usage, happy customers, and a line of requests that cost more than the plan they're on — and nothing in a request count would show you which ones.

Block on the worst case, or adapt on the real one

The simplest version of this check does one thing: estimate the worst this call could cost, reserve that much, and reject if the balance can't cover it. That protects you, but it's blunt. A customer with 1,000 credits left can be turned away for a call that reserves 1,400 — even if it would have actually used 200.

Vibe does that ceiling check too; nothing runs without clearing it. But the ceiling isn't the only thing it tracks. Each call settles with what it actually used, not the estimate, and that real number is what the next call sees:

cust_acme: 1,000 credits
  → call 1 (tool lookup):   reserve 300, settle 90   → balance 910
  → call 2 (model answer):  reserve 400, settle 210  → balance 700
  → call 3: checked against 700, not the first estimate

That's the part that matters for agents. A loop that calls a model, calls a tool, and calls the model again isn't one cost decision — it's several. A workflow tier evaluated on call 3 sees what calls 1 and 2 actually spent, so it can let a cheap loop keep running or route a getting-expensive one to a cheaper model, instead of pre-judging the whole loop on one worst-case number up front.

This is measured, not instant. Settlement lands right after each call finishes, not necessarily before your code sees that call's result — so two calls fired back-to-back can briefly see a balance that hasn't caught up yet. It's real consumption, not a live meter with zero lag, and the worst-case ceiling on each individual call is still what keeps a surprise from reaching the provider.

When the ceiling does reject a call, this is what it looks like. Your code gets a VibeRejectionError, and you decide what the customer sees:

try {
  await usage.chat({ identity: 'cust_acme', provider: 'openai', model: 'gpt-5', messages });
} catch (err) {
  if (err instanceof VibeRejectionError) {
    return showUpgradePrompt(); // out of credits: offer a top-up
  }
  throw err;
}

That catch block is the whole product moment. It's where "you're out of credits" turns into "here's how to get more."

Runaway agents, made concrete

Picture a support agent: it calls a model, calls a search tool, and on a low-confidence answer, calls the model again with what it found. That's three calls for one question. Give it a bad retry rule, or a tool that keeps returning low-confidence results, and it can run that loop many times over before anyone notices — all for the one question a customer actually asked.

A workflow tier stops that regardless of how it got there: once the identity crosses the threshold, the next call in the loop is rejected, not the one that finally gets noticed in a report.

The balance also isn't one-size-fits-all. credits() reports both the account's overall plan cap and this customer's own balance, so a single runaway identity can be capped without touching the account-wide ceiling, or the other way around.

A policy can change what runs

It doesn't have to say no. A tier can route the call to a cheaper model instead: your code asks for claude-opus-5, the policy decides claude-sonnet-5 runs instead once the balance is low, and result.model tells you which one actually ran. The call in your code doesn't change — only the policy does.

← PreviousVibe Workflows: what happens to identity's usage Next →Sell more credits with Stripe

FAQ

Does blocking add latency?

Yes — one check before the call, not nothing. That check is also what keeps a denied call from ever reaching the provider, so it's spending time to save a worse cost.

Can I block one customer without affecting others?

Yes. Usage and limits are tracked per identity, so one customer hitting their balance doesn't touch anyone else's.

Is a denied call billed?

Not by the provider — a denial never reaches it. Internally, nothing is reserved or settled against the customer's balance either; the policy check simply doesn't let the call through.