LLM API key abuse detection
To investigate possible misuse of an LLM API key, connect its request history to the workloads authorized to use it. Look for changes in timing, models, token usage, cost, and observed networks, then check those changes against deployments and credential records. An unusual pattern opens a case; it does not establish who used the key or whether they had permission.
If the credential is exposed, contain it promptly
If you have confirmed exposure or ongoing unauthorized use, follow your incident response procedure to revoke or rotate the affected credential and update legitimate workloads. Preserve existing logs alongside containment where practical; do not wait for a complete behavioral analysis while misuse continues. Record the action time and verify that the old credential no longer grants access.
Establish which credential was exposed. A gateway virtual key and its upstream provider key may authorize different routes and workloads. Changing one does not necessarily invalidate the other. Use your own administrative records to identify the scope before assuming the incident is contained.
Collect the evidence needed to reconstruct usage
Work from a restricted, pseudonymized export. Keep the mapping from the pseudonymous credential to its owner inside your organization. Do not put a raw API key in a spreadsheet, ticket, or analysis export.
| Evidence | What it helps answer | Limitation |
|---|---|---|
| Stable credential identifier and UTC timestamps | Which requests belong to the key, and when did activity change? | A shared key may represent multiple authorized workloads. |
| Request and trace identifiers | Can gateway, provider, and billing records be reconciled? | Retries may create multiple attempts for one logical operation. |
| Model, token counts, and recorded cost | Did usage volume or the workload mix change? | Missing usage is unknown, not zero; estimated cost is not necessarily the invoice. |
| Outcome, errors, and duration | Was the increase successful inference, repeated failure, or slow requests? | An error does not by itself establish whether upstream work was billed. |
| Pseudonymous observed network identifier | Did requests arrive through a new network? | NAT, failover, and gateways can hide or change the originating client. |
| Issuance, expiry, revocation, owner, and deployment records | Was this use expected and authorized at the time? | Request behavior cannot substitute for authorization records. |
Start with the last known normal period and the suspected change window. Compare equivalent hours and workload schedules. Where retention permits, include enough history to see weekly batch jobs and normal deployment changes. Document gaps instead of treating an incomplete export as the whole incident.
Compare the key with its own established workload
- Find the first meaningful change. Plot request counts, tokens, and cost over consistent time intervals. Check whether the increase is mostly more requests, larger requests, or a different model mix.
- Separate distinct activity. Compare observed networks and trustworthy workload identifiers. Check whether established traffic continued while unfamiliar traffic appeared. A network alone is not a person or device.
- Check operational explanations. Ask about launches, migrations, new regions, retries, scheduled evaluations, and changed model routing. Align deployment timestamps with the activity.
- Reconcile authorization. Ask the workload owner to identify the relevant jobs or traces. Record which requests can be explained, which remain unresolved, and what evidence is missing.
A worked example: a spike that needs investigation
Synthetic example, not an observed customer incident. A service key normally makes 120 requests per hour from one observed network. At 14:00, it makes 480 requests: 120 from its established network and 360 from a previously unseen network.
This supports the narrow finding that a new traffic population contributed to the increase. It does not prove theft. A scheduled evaluation, a deployment to another region, or an approved integration could produce the same pattern.
Check whether the owner can match the 360 requests to authorized work. Compare their model mix and timing with the declared job, and reconcile request identifiers where available. If the export records $18 for that group, describe it as $18 of recorded usage under investigation. Calling the whole amount fraud loss would require additional evidence.
What changes when you use LiteLLM?
LiteLLM's standard logging payload can include request identifiers, timing, token counts, cost, and virtual-key identity metadata. Identity fields can be absent, and some callbacks can lack the standard payload. Verify your deployed version and configuration before depending on a field. See the LiteLLM logging specification.
Virtual keys provide a way to scope model access and associate spend with keys, users, or teams. Use those records to identify the affected workload; a gateway's upstream provider credential may cover a broader population. See LiteLLM virtual-key documentation.
Inspect the fields your gateway actually emits. If the downstream client disappears behind a shared upstream key, document that attribution gap. No detector can reconstruct an identity that was never recorded.
Close the case with a reproducible explanation
Keep the observation window, affected pseudonymous credential, evidence references, operational explanation, containment actions, and unresolved questions together. Classify the outcome as explained authorized activity, confirmed unauthorized use supported by additional evidence, or unresolved. A quiet period after rotation alone does not establish what happened earlier.
For the next step, use the gateway security logging guide to check evidence coverage, the credential exposure response guide for containment planning, or the credential sharing guide when several workloads share access.