LLM token theft, fraud, and unauthorized resale
In an inference service, “token theft” usually describes unauthorized consumption of model access or allowances. “Token fraud” can refer to several different problems: stolen credentials, repeated promotional claims, or prohibited resale. Define the suspected mechanism before attaching a fraud label to traffic.
Reselling access is not automatically abuse
A gateway may legitimately buy upstream inference and sell a managed service to downstream customers. Different prices, multiple client networks, or a shared provider credential do not establish that its supply is stolen. The relevant questions are whether the upstream access was authorized and whether the downstream use complies with the applicable agreement.
An operator investigating unauthorized resale needs both usage evidence and authorization context. Metadata can show that one credential appears to serve different workloads. It cannot, by itself, establish the contractual relationship between those workloads.
Separate the suspected loss mechanisms
| Hypothesis | Evidence that would strengthen it | Evidence that is insufficient alone |
|---|---|---|
| A stolen credential is consuming paid inference | Confirmed exposure, owner-disputed requests, and corroborating access records | A cost spike |
| An account is serving unapproved downstream customers | Delegation records inconsistent with permitted use and reconciled downstream activity | Many observed IP addresses |
| A group repeatedly claims a limited allowance | Linked entitlement claims and reviewed evidence that the eligibility rule was violated | Accounts sharing a network |
| A legitimate workload has become expensive | Job traces explaining retries, model changes, or larger requests | An anomaly score described as fraud |
Investigate the credential and its route
- Identify the billing boundary. Establish whether the event belongs to an upstream provider key, a gateway virtual key, a team, or an end-user account.
- Reconstruct the time window. Compare request count, model mix, tokens, outcomes, and cost with equivalent earlier periods.
- Separate observable workloads. Use reliable client or workload identifiers where present. Treat network changes as supporting context.
- Ask for reconciliation. Check whether the authorized owner can associate the activity with jobs, customers, or deployments.
- Record the conclusion. Distinguish confirmed unauthorized consumption, explained activity, and unresolved attribution.
For LiteLLM deployments, virtual-key records can help connect spend and access scope to a key or team. A downstream virtual key and an upstream provider key represent different control points; make the distinction explicit in the incident timeline.
Estimate exposure without inventing a loss figure
Synthetic example: a credential records $100 of usage in a review window. The owner reconciles $70 to known workloads and disputes $30. Report $30 of disputed usage pending verification, not $100 of confirmed fraud. If retries or missing cost records affect the calculation, include those limitations.
Gross recorded usage, upstream cost, customer billing, and recoverable loss are different measures. Label the number you actually have. Do not assign every request in a flagged account to an attacker.
Reduce future investigation gaps
Scope credentials to identifiable workloads, keep ownership records current, and retain request correlations across gateway hops. Apply budgets and rate limits appropriate to legitimate use, while recognizing that low-volume unauthorized use can remain below them. For confirmed compromise, contain the credential under the incident process and verify revocation.
Continue with API-key investigation, gateway provenance and delegation, or the inference abuse overview. This guide addresses LLM inference access, not cryptocurrency tokens or payment-fraud classification.