AI spend anomaly detection: baselines and investigation
A spend anomaly is a departure from an expected workload, not a fraud classification. To explain it, separate changes in request count, tokens per request, model mix, and pricing. Then connect the change to jobs and authorization records.
Use comparable baselines
Compare the same principal and workload across equivalent periods. An overnight batch should not be judged against a quiet afternoon. Mark deployments, promotions, scheduled evaluations, and model migrations. For a new account with little history, say that its baseline is immature instead of presenting a precise confidence score.
Keep recorded usage separate from invoice amounts. Cached-token pricing, retries, provider adjustments, and missing cost data can change reconciliation. Fix the accounting unit before interpreting the anomaly.
Decompose the increase
| Change | Check |
|---|---|
| More requests | New customers, retries, repeated jobs, or unfamiliar traffic |
| More input tokens per request | Context accumulation or larger legitimate documents |
| More output tokens per request | Changed output limits or task type |
| A more expensive model mix | Routing changes, fallback, or a new permitted model |
| Different reported cost with similar usage | Pricing configuration, cached usage, currency, and billing adjustments |
A worked calculation
Synthetic example: a job makes 1,000 comparable requests at an average recorded cost of $0.01, totaling $10. The next run makes 2,000 at $0.03, totaling $60. At the earlier average cost, the extra 1,000 requests account for $10. Applying the $0.02 average-cost increase to the new 2,000-request volume accounts for the remaining $40 increase. This decomposition explains the arithmetic, not the cause.
Investigate the model and token mix behind the average before saying prices tripled. Also check whether the apparent request increase is duplicate log delivery. Deduplicating events must not accidentally discard real paid retry attempts.
Turn a baseline into a review policy
Choose thresholds using known workloads, review capacity, and the cost of missed events. Evaluate absolute impact as well as relative change: doubling a tiny test account may matter less than a modest increase on a large production workload. Record why a case crossed the threshold.
Keep the baseline window, minimum data requirement, comparison method, and exclusions with the result. Avoid automatically absorbing a suspected incident into the normal baseline while it is still under investigation.
Report the result honestly
Show total recorded cost, the portion explained by known changes, and the portion still under review. Do not multiply an anomaly score by spend to invent fraud loss. Track confirmed explanations and unresolved cases so later threshold changes can be evaluated against the same evidence.
For an active spike, follow the incident triage sequence. For consumption controls, use the rate-limit and budget guide. For the fields needed, see security logging.