AI agent metering UsageFlow

AI agent metering that stops a run at the cap

Give every agent run a customer identity. UsageFlow counts the tokens, checks the limit before each model call, and bills what actually ran.

Free: 1 workflow with live policy. No credit card.

An agent is a loop, and loops are hard to price

You ask it one question. It plans, searches, reads, hits a tool error, tries again, drafts, reviews and revises. Eight model calls for one answer, and every retry sends the whole conversation again.

A request counter sees eight separate requests. It doesn't know they belong to one customer's run, or that this customer has already spent most of their day's allowance. The first time you see the total is the invoice.

What a request counter sees

Eight calls, no owner, no limit.

POST /v1/messages  200
POST /v1/messages  200
POST /v1/messages  200
POST /v1/messages  200
POST /v1/messages  200
POST /v1/messages  200
POST /v1/messages  200
POST /v1/messages  200

What the ledger sees

One customer, one workflow, one running total.

identity   cust_acme
workflow   research-agent
calls      7 ran, 1 blocked
usage      202,000 tokens
policy     route at 120,000
            block at 200,000
result     stopped at the cap

The same run, with a meter

You change the model call, not the agent. The rules live in the Console.

  1. Name whose budget the run spends

    Replace the provider call with usage.chat and pass the customer's identity and a workflow name. Use the same identity you use in Stripe.

  2. UsageFlow checks before every call

    It reads the customer's running total and applies the policy: allow, route to a cheaper model, degrade, or block. Prompts still go from your server straight to the provider.

  3. Handle the stop like any other result

    A blocked call raises a rejection error and the provider is never contacted. Catch it and end the run with a clear message.

    TypeScript
    import { usage, VibeRejectionError } from '@usageflow/vibe';
    
    async function runAgent(customerId: string, task: string) {
      const messages = [{ role: 'user', content: task }];
    
      for (let step = 0; step < 12; step++) {
        try {
          const turn = await usage.chat({
            identity: customerId,         // whose budget this run spends
            workflow: 'research-agent',   // the policy that sets the cap
            provider: 'anthropic',
            model: 'claude-sonnet-5',
            messages,
          });
          messages.push({ role: 'assistant', content: turn.content });
          if (isDone(turn)) return turn.content;   // your own stop condition
        } catch (err) {
          if (err instanceof VibeRejectionError) {
            return 'This run reached your usage limit.';
          }
          throw err;
        }
      }
    }
  4. Bill what ran

    Add a Stripe meter effect to the policy and each call's usage is reported to your Stripe meter. What you limit and what you invoice come from the same numbers.

Change the cap without a deploy

Thresholds, models and who they apply to are edited in the Console and apply to the next call.

What changes for your team

Questions about agent metering

Straight answers for the category questions.

What is AI agent metering?

AI agent metering counts what each agent run consumes, per customer, and enforces a limit before the next model call. UsageFlow does this inside your app's runtime, so retries and loops are counted too.

How does UsageFlow stop an agent that is looping?

Every model call goes through the UsageFlow SDK with a customer identity and a workflow. Before the call, UsageFlow reads the customer's usage and applies your policy. When the cap is reached the call is blocked, the provider is never contacted, and your agent receives a rejection error it can handle.

How much code changes?

You replace your model call with usage.chat and add an identity and a workflow name. Limits, thresholds and models live in the Console, so you change them without a deploy.

Can I bill agent usage through Stripe?

Yes. Add a Stripe meter effect to the policy and UsageFlow reports each call's usage to your Stripe meter. Use the same identity as your Stripe customer so usage lines up with invoices.

Do my prompts go through UsageFlow?

No. Prompts and responses go from your server straight to the AI provider using your own keys. UsageFlow only sees usage numbers and the identifiers you send.

What does it cost?

The free plan includes one workflow with live policy and needs no credit card.

Stop the next runaway run before it starts

Free: 1 workflow with live policy. No credit card. TypeScript, Python and Go, with OpenAI and Anthropic.