A request isn't a price
For twenty years, "one API call" was a fair unit. It cost about the same as the last one, so you could rate-limit it, count it and charge for it.
AI broke that.
A user clicks "summarize" once. Behind the button, your app calls a model, searches a tool, hits a rate limit, retries, then calls the model again with everything it learned. From the outside it's one request. On that particular call, your provider invoice might show six line items — on a longer one, nineteen.
Counting requests tells you how often a customer used your app. It doesn't tell you how much they consumed, or what that consumption cost you.
Credits are the abstraction
Not "you get 1,000 requests." Nobody knows what that means for an AI feature, including you.
What you want to say is: "You have 10,000 credits. Things cost credits. When they run out, you can buy more." A credit is a unit you define, that maps to what things actually cost you — your product, your pricing, not something Vibe dictates.
Runtime control is the differentiator
A monthly report tells you someone burned through their allowance three weeks ago. That's a story, not control.
A system that only reserves the worst case can reject a request the customer actually had the balance for. Vibe measures what each call actually consumes, then uses the updated balance to guide what happens next. When credits run low, your policy can route the next call to a cheaper model instead of rejecting it outright.
Where Vibe fits
You still own the product and the customers. Vibe is the layer underneath that keeps the count. Your app identifies the customer. Vibe measures AI usage, updates their balance, and enforces your policies as consumption changes.
const result = await usage.chat({
identity: 'cust_acme', // whose credits these are
provider: 'anthropic',
model: 'claude-sonnet-5',
messages: [{ role: 'user', content: 'Summarize this ticket.' }],
});
Your customer never hears the word Vibe. They see "9,680 credits left."
FAQ
What is a credit in an AI app?
A unit you define for AI usage, so one number covers tokens, model calls, tools and retries. Customers get a balance, and each AI call spends from it.
Why not just rate-limit requests?
Because requests vary enormously in cost. One request can trigger many model calls, and a limit on requests won't protect you from the expensive ones.