The order things happen in

When your code calls usage.chat(), three things happen before and after the model does its work:

  1. Reserve. Vibe works out the worst this call could cost (your input size plus the output cap) and asks: does this customer have that much left? If not, the call stops here and the provider is never contacted.
  2. Run. The call goes to Anthropic or OpenAI as normal.
  3. Settle. Vibe replaces the estimate with the real token count. If the model used less than the ceiling, the customer only pays for what was used.

Check, execute, settle. The balance is never a guess after the fact.

Set it up

Vibe ships as a client in three languages. Pick the one your service is already written in — the reserve/run/settle behavior above is identical across them.

TypeScript

npm install @usageflow/vibe
export USAGEFLOW_API_KEY='your-api-key'
import { usage } from '@usageflow/vibe';

usage.init({ apiKey: process.env.USAGEFLOW_API_KEY! });

const result = await usage.chat({
  identity: 'cust_acme',
  workflow: 'support-agent',
  provider: 'anthropic',
  model: 'claude-sonnet-5',
  messages: [{ role: 'user', content: 'Summarize this ticket.' }],
});

Python

pip install usageflow-vibe
export USAGEFLOW_API_KEY='your-api-key'
from usageflow.vibe import VibeClient, Message

client = VibeClient()  # reads USAGEFLOW_API_KEY

result = client.chat(
    identity="cust_acme",
    workflow="support-agent",
    model="claude-sonnet-5",
    messages=[Message("user", "Summarize this ticket.")],
)

Go

go get github.com/usageflow/usageflow-go-middleware/v2/pkg/vibe
client, _ := vibe.New(vibe.Options{})

result, err := client.Chat(ctx, vibe.ChatRequest{
  Identity: "cust_acme",
  Workflow: "support-agent",
  Provider: vibe.ProviderAnthropic,
  Model:    "claude-sonnet-5",
  Messages: []vibe.Message{{Role: "user", Content: "Summarize this ticket."}},
})

identity is the customer. It's the thing the credits belong to. workflow is optional and picks which policy applies.

You don't need a policy to start. Make the call with an explicit provider and model, see the traces, and add limits when you're ready.

Show the customer what's left

You'll want a "credits remaining" number in your UI. Reading it costs nothing and reserves nothing:

const c = await usage.credits({ identity: 'cust_acme', workflow: 'support-agent' });
c.workflows[0].remaining; // credits left before this workflow blocks
c.workflows[0].blocked;   // true once the limit is reached

This needs a UsageFlow server new enough to support it; against an older one, credits() throws VibeCreditsUnsupportedError instead of a number.

Spending credits outside a chat call

Not every cost is a model call. A nightly batch, an export, a correction. withdraw() and credit() move the balance directly:

await usage.withdraw({
  identity: 'cust_acme',
  amount: 500,
  idempotencyKey: 'job-42-attempt-1',
  reason: 'nightly export',
});

idempotencyKey is required and travels with the ledger event for audit, but the server doesn't dedupe on it yet — a blind retry can still double-deduct, so make it unique per attempt and don't rely on it alone for safety.

What you'll see in the Console

Each call shows up as a trace: which customer, which model, how many tokens, what it cost, and whether it was allowed. That's the record you'll open the first time a customer asks "where did my credits go?"

← PreviousWhy AI apps need credits Next →Stop AI costs before they happen

FAQ

Do I need to change my provider setup?

No. Vibe builds the Anthropic and OpenAI clients for you, using your existing provider keys.

What if the call fails?

Settlement uses real usage only. A call that never completed doesn't spend the customer's credits.

Is the balance per workflow or per customer?

Per customer. Every workflow's limit is measured against that customer's one usage number.