> For the complete documentation index, see [llms.txt](https://docs.nerovasystems.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.nerovasystems.com/guides/observe-and-account/track-and-limit-usage.md).

# Track and limit usage by your plans

Nerova bills your organization by tokens, always. Nerova does not define plans or tiers for your merchants: you sell your own plans, and you decide what each one includes. This guide shows how to count what each merchant (a Nerova tenant) uses, and how to enforce your plans with the per-tenant controls Nerova already exposes.

> **Experimental.** The `/api/v1/**` operations and webhook events on this page are callable with your Live key (`nrv_live_`), which runs on the free testing allowance until a commercial agreement is signed or a card is saved. See [Lifecycle and availability](https://docs.nerovasystems.com/documentation/api/lifecycle).

## What you'll build

A per-merchant usage loop:

1. **Count** AI conversations and AI bookings for each merchant.
2. **Compare** the counts with the plan the merchant bought from you.
3. **Enforce**: cap tokens, suspend or resume the merchant's AI, or stop AI booking, using Nerova's tenant controls.

Nerova never sees your prices or your plan names. It carries counts, tokens, and switches; your system owns the plan logic.

## Prerequisites

* An activated tenant per merchant; see [Provision and activate tenants](/guides/provision/partner-playbook.md).
* An API key with `usage:read`, `merchant:provision`, `mandate:write`, and `webhook:manage`; see [Authentication](https://docs.nerovasystems.com/documentation/getting-started/authentication).
* A webhook receiver that verifies signatures and deduplicates on the envelope `id`; see [Receive events with webhooks](/guides/build-well/webhooks.md).

## What counts

* **Period.** Every counter below is per **UTC calendar month**. If your plans bill on another cycle, count from the webhook events instead of the monthly summary.
* **AI conversation.** A conversation starts when the AI handles a customer's message and no conversation start was recorded for the same tenant, customer, and channel in the previous 24 hours. A voice call counts as one conversation. Messages inside the 24-hour window do not start a new one.
* **AI booking.** A booking the AI created on the merchant's platform through the Provider API. A cancellation the AI made is counted separately.
* **Tokens.** Input plus output tokens for AI turns, with the voice minutes surcharge included. Tokens are what Nerova bills you for.

## 1. Count with webhooks

Create one endpoint per merchant tenant and subscribe to the three counting events:

```bash
curl --request POST \
  --url "https://api.nerovasystems.com/api/v1/tenants/{tenantId}/webhooks" \
  --header "Accept: application/json" \
  --header "Authorization: ******" \
  --header "Content-Type: application/json" \
  --header "Idempotency-Key: usage-hook-{tenantId}" \
  --data '{
    "url": "https://platform.example.com/nerova/usage",
    "description": "Plan usage counters",
    "eventFilters": ["conversation.started", "booking.created", "booking.cancelled"]
  }'
```

The events carry these `data.object` snapshots. Deduplicate on the envelope `id`, because delivery is at-least-once:

```json
{
  "conversation": "cnv_01J…",
  "channel": "whatsapp",
  "started_at": "2026-08-10T20:14:02Z"
}
```

```json
{
  "booking": "rcpt_01J…",
  "external_appointment_id": "apt_9001",
  "channel": "whatsapp",
  "conversation": "cnv_01J…",
  "created_at": "2026-08-10T20:16:41Z"
}
```

`conversation.started` carries `channel` of `whatsapp` or `voice`. `booking.created` and `booking.cancelled` carry `channel` of `whatsapp`, `voice`, or `flow`, and a `conversation` id that is `null` when the booking did not come from a conversation. `booking.cancelled` has the same shape with `cancelled_at` in place of `created_at`. `external_appointment_id` is the id on your platform, so you can join the event to your own appointment.

## 2. Reconcile with the usage summary

Webhooks tell you when something happened; the summary is the authoritative total for the month. Read it for each merchant at your billing cutoff, and whenever a webhook was missed:

```bash
curl --request GET \
  --url "https://api.nerovasystems.com/api/v1/tenants/{tenantId}/usage/summary?Year=2026&Month=8" \
  --header "Accept: application/json" \
  --header "Authorization: ******"
```

The response is counts only, never Nerova rates or amounts:

* `conversationCount`: AI conversations started in the month.
* `bookingCount`: bookings created in the month, including AI bookings made on your platform through the Provider API.
* `bookingsCancelled`: bookings the AI cancelled in the month.
* `inputTokens`, `outputTokens`, `totalTokens`, and `voiceSeconds`.
* `periodStart`, `periodEnd`, `asOf`, and `meterVersion`, so a reading can be reproduced later.

Omit `Year` and `Month` for the current month; when you send one, send both.

## 3. Attribute bookings in your own system

When Nerova creates, reschedules, or cancels an appointment on your platform, the Provider API request can carry an optional `origin` object:

```json
{
  "origin": {
    "source": "nerova",
    "channel": "whatsapp",
    "conversationId": "cnv_01J…"
  }
}
```

Store it with the appointment. It lets you count AI bookings per merchant inside your own database, flag them in your merchant's calendar, and match them to the `booking.*` events. Treat the field as optional: appointments made by staff or by your own customers carry no `origin`. See [Step 3: Implement Provider API v1](/guides/zero-to-live/step-3-implement-provider-api.md).

## 4. Enforce your plans

Nerova gives you three switches per tenant. Use whichever your plan needs.

### Cap tokens

Set monthly token caps with `PUT /api/v1/tenants/{tenantId}/usage/limits`. The **hard cap** stops the tenant's AI for the rest of the UTC month once reached. The **soft cap** never blocks; it only alerts you. Send `null` to clear a cap. The soft cap must not exceed the hard cap.

```bash
curl --request PUT \
  --url "https://api.nerovasystems.com/api/v1/tenants/{tenantId}/usage/limits" \
  --header "Accept: application/json" \
  --header "Authorization: ******" \
  --header "Content-Type: application/json" \
  --header "Idempotency-Key: limits-{tenantId}-2026-08" \
  --data '{ "hardCapMonthlyTokens": 2000000, "softCapMonthlyTokens": 1600000 }'
```

Subscribe to `usage.threshold_crossed` to hear about it without polling. It fires at most once per tenant, per cap, per UTC month:

```json
{
  "threshold": "soft_cap",
  "unit": "tokens",
  "month": "2026-08",
  "tokens_used": 1604212,
  "cap": 1600000,
  "crossed_at": "2026-08-19T13:27:55Z"
}
```

`threshold` is `soft_cap` when usage reaches or passes the soft cap, and `hard_cap` when usage reaches the hard cap. The event fires on the AI turn that carries the month's usage to the cap. A few details:

* The counter and the once-per-month marker reset at the start of the next UTC month, so a tenant that crosses again next month fires again.
* Changing a cap re-arms it for the current month. Sending the same values again does not.
* If you set a hard cap at or below the tokens already used this month, the tenant's AI stops immediately and no event fires, because no further AI turn runs.

### Suspend and resume

`POST /api/v1/tenants/{tenantId}/lifecycle/Suspended` stops the tenant's AI immediately, and `POST /api/v1/tenants/{tenantId}/lifecycle/Active` resumes it. Each call takes an audited `reason`:

```bash
curl --request POST \
  --url "https://api.nerovasystems.com/api/v1/tenants/{tenantId}/lifecycle/Suspended" \
  --header "Accept: application/json" \
  --header "Authorization: ******" \
  --header "Content-Type: application/json" \
  --header "Idempotency-Key: suspend-{tenantId}-2026-08" \
  --data '{ "reason": "Starter conversation allowance used for August" }'
```

Use it for plans with a hard conversation allowance: when your counter reaches the limit, suspend; at the start of the next month, or when the merchant upgrades, resume. Suspending is all or nothing: it switches the tenant's AI off, not just AI booking.

### Stop AI booking

If a plan includes conversations but not AI booking, set the `bookingsAndReschedules` duty to `Never` with `PUT /api/v1/tenants/{tenantId}/mandate`, or to `AskFirst` to keep a human in the loop. Read the mandate first and send its current `version`; the request shape is in [Set the autonomy mandate](/guides/operate-the-digital-employee/mandate.md).

```bash
curl --request PUT \
  --url "https://api.nerovasystems.com/api/v1/tenants/{tenantId}/mandate" \
  --header "Accept: application/json" \
  --header "Authorization: ******" \
  --header "Content-Type: application/json" \
  --header "Idempotency-Key: mandate-{tenantId}-no-ai-booking" \
  --data '{
    "version": "v1:8c7d6e5f4a3b2c1d0e9f8a7b",
    "changes": [{ "verb": "bookingsAndReschedules", "level": "Never" }],
    "actor": { "actorReference": "system:plan-enforcement", "delegationReference": null }
  }'
```

## Worked example: Starter and Pro

The numbers below are illustrative choices for your own plans, not Nerova defaults.

|                            | Starter   | Pro        |
| -------------------------- | --------- | ---------- |
| AI conversations per month | 150       | 1,000      |
| AI booking                 | Off       | On         |
| Token hard cap             | 2,000,000 | 10,000,000 |
| Token soft cap             | 1,600,000 | 8,000,000  |

When a merchant buys a plan, apply it once:

1. `PUT …/usage/limits` with the plan's caps.
2. `PUT …/mandate` with `bookingsAndReschedules` set to `Never` for Starter, or your chosen level for Pro.
3. Subscribe the tenant's endpoint to `conversation.started`, `booking.*`, and `usage.threshold_crossed`.

Then run the loop:

* On `conversation.started`, increment the merchant's counter for the month. At 80 percent of the allowance, tell the merchant; at the allowance, `POST …/lifecycle/Suspended` on Starter, or offer an upgrade.
* On `booking.created` and `booking.cancelled`, update the merchant's AI booking numbers and show them on the Pro merchant's dashboard.
* On `usage.threshold_crossed` with `soft_cap`, warn the merchant that their AI is close to the token limit. On `hard_cap`, show that the AI has paused until next month or an upgrade, and raise the cap on upgrade.
* On the first day of each UTC month, reset your counters and, if you suspended a tenant for a used-up allowance, `POST …/lifecycle/Active`.
* At your billing cutoff, read `usage/summary` and correct any counter that drifted from Nerova's numbers.

## Next steps

* Meter what Nerova bills you per tenant: [Meter usage for billing](/guides/observe-and-account/usage.md).
* Deliver events reliably: [Receive events with webhooks](/guides/build-well/webhooks.md).
* Prove what the AI did: [Audit work and receipts](/guides/observe-and-account/work-and-receipts.md).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.nerovasystems.com/guides/observe-and-account/track-and-limit-usage.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
