Where the simple model breaks

In a simple API, a request is a fair unit. Reserve 1, do the lookup, settle 1.

Try that on an AI route and it falls apart:

POST /api/agent
  ├── model call
  ├── search tool ── retry
  ├── database lookup
  └── model call

One request from the outside. Four operations inside, two of them expensive, one of them repeated. Next time it may be nine.

So reserve a ceiling, settle the truth

Vibe handles this by splitting the count into two steps. Before the call, it reserves the worst case: your input size plus your output cap. After the call, it settles the real token usage. If the model used less than the ceiling, the customer is only charged for what was used.

Same shape as the simple API. Just a variable amount instead of a fixed 1.

Enter credits

Once cost varies, you need a unit that covers tokens, tool calls and retries without your customer having to understand any of them. That's a credit.

A simple API: 1 lookup = 1. An AI product: 1 model call = whatever it cost. Same engine, same policy, different unit.

A policy can change what runs

An AI policy can say more than no. When credits are low, it can send the call to a cheaper model, cap the output, or turn off an expensive tool. A limit stops the call. A policy decides what the call becomes.

← PreviousEnforce limits where the work happens Next →Turn usage into a customer budget

FAQ

Why not count tokens directly?

You can, but customers don't think in tokens. Credits let you set the price of things, including tools and images.

What happens if the estimate is too high?

Nothing is lost. Settlement replaces the estimate with real usage.

Can a policy change the model?

Yes. A policy can route or downgrade the call, and result.model reports what actually ran.