Where the simple model breaks
In a simple API, a request is a fair unit. Reserve 1, do the lookup, settle 1.
Try that on an AI route and it falls apart:
POST /api/agent
├── model call
├── search tool ── retry
├── database lookup
└── model call
One request from the outside. Four operations inside, two of them expensive, one of them repeated. Next time it may be nine.
So reserve a ceiling, settle the truth
Vibe handles this by splitting the count into two steps. Before the call, it reserves the worst case: your input size plus your output cap. After the call, it settles the real token usage. If the model used less than the ceiling, the customer is only charged for what was used.
Same shape as the simple API. Just a variable amount instead of a fixed 1.
Enter credits
Once cost varies, you need a unit that covers tokens, tool calls and retries without your customer having to understand any of them. That's a credit.
A simple API: 1 lookup = 1. An AI product: 1 model call = whatever it cost. Same engine, same policy, different unit.
A policy can change what runs
An AI policy can say more than no. When credits are low, it can send the call to a cheaper model, cap the output, or turn off an expensive tool. A limit stops the call. A policy decides what the call becomes.
FAQ
Why not count tokens directly?
You can, but customers don't think in tokens. Credits let you set the price of things, including tools and images.
What happens if the estimate is too high?
Nothing is lost. Settlement replaces the estimate with real usage.
Can a policy change the model?
Yes. A policy can route or downgrade the call, and result.model reports what actually ran.