The order things happen in
When your code calls usage.chat(), three things happen before and after the model does its work:
- Reserve. Vibe works out the worst this call could cost (your input size plus the output cap) and asks: does this customer have that much left? If not, the call stops here and the provider is never contacted.
- Run. The call goes to Anthropic or OpenAI as normal.
- Settle. Vibe replaces the estimate with the real token count. If the model used less than the ceiling, the customer only pays for what was used.
Check, execute, settle. The balance is never a guess after the fact.
Set it up
Vibe ships as a client in three languages. Pick the one your service is already written in — the reserve/run/settle behavior above is identical across them.
TypeScript
npm install @usageflow/vibe
export USAGEFLOW_API_KEY='your-api-key'
import { usage } from '@usageflow/vibe';
usage.init({ apiKey: process.env.USAGEFLOW_API_KEY! });
const result = await usage.chat({
identity: 'cust_acme',
workflow: 'support-agent',
provider: 'anthropic',
model: 'claude-sonnet-5',
messages: [{ role: 'user', content: 'Summarize this ticket.' }],
});
Python
pip install usageflow-vibe
export USAGEFLOW_API_KEY='your-api-key'
from usageflow.vibe import VibeClient, Message
client = VibeClient() # reads USAGEFLOW_API_KEY
result = client.chat(
identity="cust_acme",
workflow="support-agent",
model="claude-sonnet-5",
messages=[Message("user", "Summarize this ticket.")],
)
Go
go get github.com/usageflow/usageflow-go-middleware/v2/pkg/vibe
client, _ := vibe.New(vibe.Options{})
result, err := client.Chat(ctx, vibe.ChatRequest{
Identity: "cust_acme",
Workflow: "support-agent",
Provider: vibe.ProviderAnthropic,
Model: "claude-sonnet-5",
Messages: []vibe.Message{{Role: "user", Content: "Summarize this ticket."}},
})
identity is the customer. It's the thing the credits belong to. workflow is optional and picks which policy applies.
You don't need a policy to start. Make the call with an explicit provider and model, see the traces, and add limits when you're ready.
Show the customer what's left
You'll want a "credits remaining" number in your UI. Reading it costs nothing and reserves nothing:
const c = await usage.credits({ identity: 'cust_acme', workflow: 'support-agent' });
c.workflows[0].remaining; // credits left before this workflow blocks
c.workflows[0].blocked; // true once the limit is reached
This needs a UsageFlow server new enough to support it; against an older one, credits() throws VibeCreditsUnsupportedError instead of a number.
Spending credits outside a chat call
Not every cost is a model call. A nightly batch, an export, a correction. withdraw() and credit() move the balance directly:
await usage.withdraw({
identity: 'cust_acme',
amount: 500,
idempotencyKey: 'job-42-attempt-1',
reason: 'nightly export',
});
idempotencyKey is required and travels with the ledger event for audit, but the server doesn't dedupe on it yet — a blind retry can still double-deduct, so make it unique per attempt and don't rely on it alone for safety.
What you'll see in the Console
Each call shows up as a trace: which customer, which model, how many tokens, what it cost, and whether it was allowed. That's the record you'll open the first time a customer asks "where did my credits go?"
FAQ
Do I need to change my provider setup?
No. Vibe builds the Anthropic and OpenAI clients for you, using your existing provider keys.
What if the call fails?
Settlement uses real usage only. A call that never completed doesn't spend the customer's credits.
Is the balance per workflow or per customer?
Per customer. Every workflow's limit is measured against that customer's one usage number.