From 1eacfdd3d00fa514ff39aa30fe4c2b5582229e8a Mon Sep 17 00:00:00 2001 From: Alex Date: Mon, 21 Sep 2026 12:03:02 +0100 Subject: [PATCH] docs: usage quotas How the instance default, team allowances and user overrides resolve (including users in several teams), the quota window, who is charged for agent traffic, how cost budgets price models and what happens to unpriced ones, and the admin and user API. --- .env-template | 6 ++ docs/content/Deploying/Access-Control.mdx | 4 +- docs/content/Deploying/Usage-Quotas.mdx | 109 ++++++++++++++++++++++ docs/content/Deploying/_meta.js | 4 + 4 files changed, 122 insertions(+), 1 deletion(-) create mode 100644 docs/content/Deploying/Usage-Quotas.mdx diff --git a/.env-template b/.env-template index 5e868f5e..7f3c05a6 100644 --- a/.env-template +++ b/.env-template @@ -109,3 +109,9 @@ MICROSOFT_AUTHORITY=https://{tenantId}.ciamlogin.com/{tenantId} # PAT_MAX_LIFETIME_DAYS=365 # PAT_ALLOW_NON_EXPIRING=false # PAT_MAX_PER_USER=25 + +# Usage quotas (set limits in Admin → Quotas). Usage is counted per calendar +# day, week or month in UTC. Models without a declared price are recorded at $0 +# unless a fallback [input, output] USD rate per 1M tokens is given. +# QUOTA_PERIOD=month +# QUOTA_UNPRICED_RATE_PER_MILLION=[0.5, 1.5] diff --git a/docs/content/Deploying/Access-Control.mdx b/docs/content/Deploying/Access-Control.mdx index 1c452004..dc238349 100644 --- a/docs/content/Deploying/Access-Control.mdx +++ b/docs/content/Deploying/Access-Control.mdx @@ -88,6 +88,7 @@ Admins get a dashboard backed by a REST surface under `/api/admin` (every endpoi | `GET` | `/api/admin/audit` | Authentication/admin audit feed. | | `GET` | `/api/admin/devices/audit` | Remote-device audit feed. | | `GET` | `/api/admin/teams` | Instance-wide oversight of all teams. | +| `GET` `PUT` `DELETE` | `/api/admin/quotas/...` | [Usage quotas](/Deploying/Usage-Quotas) for the instance, teams and users. | Deactivating a user via the dashboard works for any auth type, while OIDC deployments can also offboard through [SCIM](/Deploying/OIDC-SSO#scim-user-provisioning). Both revoke live sessions immediately. @@ -141,9 +142,10 @@ Sharing rules: ## Audit log -Access-control actions are appended to the `auth_events` table alongside the [authentication events](/Deploying/OIDC-SSO#login-auditing). This includes admin actions — `admin_user_activated` / `admin_user_deactivated`, `admin_sessions_revoked`, `role_granted` / `role_revoked` (with `metadata.source` = `manual` or `oidc_group`) — and team events (`team.create`, `team.member_add`, `team.member_role`, `team.member_remove`, `team.share`, `team.unshare`, `team.transfer_owner`, `team.delete`). The acting admin is recorded in the event metadata. +Access-control actions are appended to the `auth_events` table alongside the [authentication events](/Deploying/OIDC-SSO#login-auditing). This includes admin actions — `admin_user_activated` / `admin_user_deactivated`, `admin_sessions_revoked`, `role_granted` / `role_revoked` (with `metadata.source` = `manual` or `oidc_group`), `quota_policy_set` / `quota_policy_deleted` — and team events (`team.create`, `team.member_add`, `team.member_role`, `team.member_remove`, `team.share`, `team.unshare`, `team.transfer_owner`, `team.delete`). The acting admin is recorded in the event metadata. ## Related - [SSO with OIDC](/Deploying/OIDC-SSO) — sign-in, group allowlists, and the `auth_events` table. +- [Usage Quotas](/Deploying/Usage-Quotas) — token and cost limits per user and per team. - [App Configuration](/Deploying/DocsGPT-Settings) — the full settings reference. diff --git a/docs/content/Deploying/Usage-Quotas.mdx b/docs/content/Deploying/Usage-Quotas.mdx new file mode 100644 index 00000000..701b9c83 --- /dev/null +++ b/docs/content/Deploying/Usage-Quotas.mdx @@ -0,0 +1,109 @@ +--- +title: Usage Quotas +description: Cap how many tokens or dollars each user may spend per day, week or month, with an instance default, per-team allowances and per-user overrides. +--- + +import { Callout } from 'nextra/components' + +# Usage Quotas + +An instance admin can limit how much each user spends on language models. A quota has two independent budgets: + +- **Tokens** — prompt plus generated tokens. Works for every model, including local ones. +- **Cost (USD)** — tokens priced at the model's catalog rate. Only sees models that declare a price. + +Set either, both or neither. Quotas are managed from **Admin → Quotas**, or through the [API](#api). With no quota set, nothing is limited. + +## Layers + +Limits are set at three layers. For each budget, the first layer that says something wins: + +1. **User override** — one user's own limit. +2. **Team allowance** — what each member of a team gets. +3. **Instance default** — everyone else. + +At each layer a budget is either *not set* (defer to the next layer), a *limit*, or *unlimited*. A limit of `0` blocks the user. The two budgets resolve separately, so a user's token limit can come from their team while their cost limit comes from the instance default. + +### Teams + +A team allowance is **per member**, not a pool the team shares: if the allowance is 2M tokens, each member may use 2M. + +A user in several teams gets the **most generous** allowance among them, and allowances are never added together. Usage is always counted per user, whichever teams they belong to. To hold one person below their team's allowance, give them a user override. + + +Team membership can change without an instance admin — team admins, OIDC group sync and SCIM all add members — so joining a team can only raise a user's allowance to what you granted that team, never lower it. Only instance admins set allowances; team admins cannot. + + +## Windows and enforcement + +Usage is counted over a calendar window in UTC, chosen for the whole instance with [`QUOTA_PERIOD`](/Deploying/Settings-Reference#quotas): `day` (from 00:00), `week` (from Monday) or `month` (from the 1st, the default). Windows are worked out when a request arrives, so there is no reset job to run. + +The quota is checked **before** a request starts. The request that crosses a limit completes; the next one is refused with HTTP `429`: + +```json +{ + "success": false, + "error_code": "quota-exceeded", + "message": "Usage quota reached (1,000,000 of 1,000,000 tokens). It resets at 2026-10-01T00:00:00+00:00.", + "dimension": "tokens", + "unit": "tokens", + "usage": 1000000, + "limit": 1000000, + "bucket": "all", + "source": "instance", + "resets_at": "2026-10-01T00:00:00+00:00" +} +``` + +The response carries a `Retry-After` header. The check covers chat, the agent and OpenAI-compatible APIs, scheduled runs (recorded as `budget_exceeded`) and webhook runs. If the quota check itself fails, the request is allowed. + +Who is charged: + +| Traffic | Charged to | +| --- | --- | +| Chat without an agent | The user | +| A user's own agent, its API key, webhooks and schedules | The agent's owner | +| An agent shared with the user | The user | + +Per-agent token and request limits still apply on top of the owner's quota. + +Users with a quota see their usage and the reset time under **Settings → Analytics**. + +## Pricing + +Cost budgets use the rates in the [model catalog](/Models/cloud-providers), in USD per million tokens: + +```yaml +models: + - id: my-model + input_cost_per_million: 3.0 + output_cost_per_million: 15.0 + cached_input_cost_per_million: 0.3 # optional, prompt-cache reads + cache_write_cost_per_million: 3.75 # optional, prompt-cache writes +``` + +The built-in catalogs ship list prices for hosted models. Override or add rates by dropping a YAML with the same model `id` into `MODELS_CONFIG_DIR`. The cost of each call is stored with its usage row when the call is made, so later price changes do not rewrite history. + + +A model with no declared price is recorded at $0, so a cost budget cannot see it. The Quotas tab lists such models once they have been used. Either limit them with a token budget, declare their rates, or set [`QUOTA_UNPRICED_RATE_PER_MILLION`](/Deploying/Settings-Reference#quotas) to charge a fallback rate. Models a user adds with their own API key are always $0, but their tokens still count. + + +## API + +Every admin endpoint requires the admin role, and every change is written to the [audit log](/Deploying/Access-Control#audit-log) as `quota_policy_set` or `quota_policy_deleted`. + +| Method | Path | Description | +| --- | --- | --- | +| `GET` | `/api/admin/quotas` | All policies by layer, the current window, and used models without a price. | +| `PUT` `DELETE` | `/api/admin/quotas/instance` | The instance default. | +| `GET` `PUT` `DELETE` | `/api/admin/quotas/teams/` | A team's per-member allowance. | +| `GET` `PUT` `DELETE` | `/api/admin/quotas/users/` | A user's override. `GET` also returns the limits the user ends up with, the layer each came from, and their usage. | +| `GET` | `/api/user/quota` | The caller's own limits, usage and reset time. | + +A `PUT` body sets, per budget, a limit or the unlimited flag; leave both out to defer to the next layer: + +```json +{ "token_limit": 2000000, "cost_unlimited": true, "note": "Research team" } +``` + +`bucket` (default `all`) narrows a policy to `direct` traffic (chat without an agent) or `agent` traffic (anything that runs through an agent). A request must fit both its own bucket and `all`. The dashboard edits `all`. diff --git a/docs/content/Deploying/_meta.js b/docs/content/Deploying/_meta.js index 105b9efd..9afb460d 100644 --- a/docs/content/Deploying/_meta.js +++ b/docs/content/Deploying/_meta.js @@ -15,6 +15,10 @@ export default { "title": "👥 Access Control & Teams", "href": "/Deploying/Access-Control" }, + "Usage-Quotas": { + "title": "📊 Usage Quotas", + "href": "/Deploying/Usage-Quotas" + }, "Docker-Deploying": { "title": "🛳️ Docker Setup", "href": "/Deploying/Docker-Deploying"