# Stop teams eating the budget

> Stop one team's runaway usage from becoming an end-of-month bill nobody can explain.

When every caller shares one provider key, model spend is a single number that nobody can break down. One runaway integration or a single busy team can consume the budget, and you often only find out when the invoice arrives. Without per-caller attribution, there is no way to cap usage in advance or charge it back to the department that incurred it.

## What you can do

- Set monthly USD budgets and token caps per member, team, or surface so that a runaway integration is stopped before it produces a large bill.

- Attribute every token and dollar to a member and their teams so that you can answer who spent what, on which model.

- Export per-team chargeback reports as CSV or JSON so that finance can bill each department without splitting a shared total by guesswork.

- Route cheaper requests to smaller models and reserve frontier models for the hard ones so that you right-size the cost of every call.

- Get alerted on cost spikes and threshold breaches so that a sudden anomaly reaches the people who can act, not just a dashboard nobody is watching.

## How it works

Every caller resolves to a Member, which acts as a virtual key carrying a model allow-list, per-minute request and token limits, and monthly ceilings. Members are grouped into Teams, and once a surface is set to contribute to team quotas, spend and tokens are attributed to both the Member and its Team across every eligible surface. Cost, tokens, and latency are metered per stage, so the LLM call, the Judge, and the Jury each show their own share rather than one combined figure. On each request the Team Gate resolves the caller, checks the model allow-list, then enforces per-minute limits and monthly ceilings before the provider is called. Attributed usage rolls up by Team, Member, and surface into chargeback reports you can export as CSV or JSON.

Usage limits and breach alerts are separate mechanisms: a limit is a hard ceiling, while a breach alert only notifies. A request that would exceed a monthly ceiling or a per-minute limit is rejected with a 429 before it reaches the provider, and an access or model-allow-list failure returns a 403.

## Related

- [Cost and usage governance](/products/affinidi-trust-fabric/agent-stream/concepts/cost-and-usage-governance.md): How per-stage metering, monthly ceilings, and anomaly alerts turn spend into something you can cap.

- [Cost and attribution](/products/affinidi-trust-fabric/agent-stream/concepts/teams-and-attribution.md): How members and teams act as virtual keys and roll spend up for chargeback.

- [Cap spend and rate per member](/products/affinidi-trust-fabric/agent-stream/how-to-guides/teams/cap-spend-and-rate-per-member.md): Step-by-step configuration of per-member budgets, token caps, and rate limits.

- [Alert on budget breaches and cost spikes](/products/affinidi-trust-fabric/agent-stream/how-to-guides/cost-and-usage/alert-on-budget-breaches-and-cost-spikes.md): How to route threshold and anomaly notifications to your existing integrations.
