AI usage metering UsageFlow

AI usage metering per customer, with limits before the call

See who used how much AI, across providers, then route, degrade or block before the model call runs.

Free: 1 workflow with live policy. No credit card.

Your provider shows one total. You need one per customer.

The provider's dashboard shows usage for a key. At the end of the month there is one invoice, and no answer to which customer spent what.

That is hard to price, hard to limit, and hard to explain to a customer who asks why they were charged.

What a key-level view shows

One number for everyone.

key        sk-...a41f
this month 3,480,000 tokens
by customer not available

What a per-customer ledger shows

A running total per identity.

cust_acme      412,000
cust_globex     62,000
cust_initech   104,000
cust_umbrella   18,000
cust_hooli       9,000

What gets metered

Every call through the SDK adds to the customer's ledger. This is exactly what is counted.

Call What is added to the ledger
Chat and streaming Input tokens plus output tokens actually used
Embeddings Input tokens
Withdraw and credit The amount you set. You choose the unit; it is not converted.

The ledger resets on the schedule your policy sets: daily, monthly, yearly, or never.

How it works

Four ideas, in the order you meet them.

  1. Identity: whose usage is this

    Pass your customer ID with every call. Use the same ID you use in Stripe. The first call with a new identity creates its ledger automatically.

  2. Ledger: how much they have used

    One running total per identity. Policies compare against this number.

  3. Policy: what happens as it grows

    A short ladder of tiers you edit in the Console. Tiers can route to a cheaper model, degrade the request, or block it, and they apply to the next call without a deploy.

  4. Settlement: what really happened

    UsageFlow reserves a worst-case amount before the call, then records what the provider actually used.

    
                  

What you do with the numbers

Metering only matters if something happens when a customer uses a lot.

Limit Set how much each customer may use per day, month or year.
Route Send heavy users to a cheaper model, even on another provider.
Degrade Keep serving with lower-cost behavior. Your code applies it.
Block Reject the call. The provider is never contacted.
Bill Report the same usage to a Stripe meter. See Stripe billing for AI.

How to choose an AI usage metering tool

Seven questions to ask any tool, and how UsageFlow answers each.

Ask Why it matters UsageFlow
Counts per customer? A per-key total can't price or limit individual customers. One ledger per identity.
Acts before the call? Alerts after the fact don't stop the spend. Policy runs before the provider is contacted. A blocked call costs nothing.
Covers your providers? Teams mix models and vendors. OpenAI and Anthropic through one call shape.
Settles to real usage? Estimates drift from the invoice. Reserves a worst-case amount, then records what ran.
Keeps prompts private? Prompts can be sensitive. Prompts and responses go from your server to the provider, not through UsageFlow.
Connects to billing? Limits and invoices should match. Stripe meter effect on the same usage.
Changes without a deploy? Pricing and limits change often. Edit policies in the Console and they apply to the next call.

Metering your own API routes instead?

UsageFlow also publishes packages that meter the HTTP routes of your own API: Express, Fastify and NestJS for Node.js, FastAPI and Flask for Python, and Gin for Go. They are separate from the UsageFlow SDK.

Different job, different setup

Use the UsageFlow SDK to meter and limit AI model calls per customer. Use a framework package to meter your API's routes. See the docs for both.

Keep going

Questions about AI usage metering

Straight answers for the category questions.

What is AI usage metering?

AI usage metering counts how much AI each customer uses, such as tokens per call, and keeps a running total per customer. UsageFlow then applies your policy to that total before each model call.

How is it different from my provider's dashboard?

A provider dashboard shows usage for a key or workspace. UsageFlow keeps a ledger per customer identity, across providers, and can act on it before the call by routing, degrading or blocking.

What does it count?

Chat and streaming calls count input plus output tokens actually used. Embeddings count input tokens. Withdraw and credit adjustments count the amount you specify.

Do my prompts go through UsageFlow?

No. Prompts and responses go from your server straight to the AI provider using your own keys. UsageFlow only sees usage numbers and the identifiers you send.

Can I meter more than one provider?

Yes. The UsageFlow SDK supports OpenAI and Anthropic through one call shape, and usage lands in one ledger per customer.

What does it cost?

The free plan includes one workflow with live policy and needs no credit card.

Know who used what before the invoice does

Free: 1 workflow with live policy. No credit card. TypeScript, Python and Go, with OpenAI and Anthropic.