AWS Bedrock Claude excessive use of tokens
Description
Detects identities generating anomalously large model responses relative to their own historical baseline. For each identity, computes the average, maximum, and standard deviation of output token counts across all invocations, then flags any identity whose single largest response exceeds two standard deviations above their own mean. A statistically significant output spike from a single identity may indicate bulk data extraction, successful prompt injection producing verbose output, or a runaway agentic loop hitting context limits.
Query · spl
`aws_bedrock_claude`
| spath output="out_tokens" path="output.outputBodyJson.usage.output_tokens"
| eval user = replace('identity.arn', ".*/", "")
| stats count AS invocations,
avg(out_tokens) AS avg_out,
max(out_tokens) AS max_out,
stdev(out_tokens) AS stdev_out
BY user, identity.arn
| eval stdev_out = coalesce(stdev_out, 0)
| eval threshold = avg_out + (2 * stdev_out)
| where max_out > threshold
| table
user,
identity.arn,
invocations,
avg_out,
max_out,
stdev_out,
threshold
| sort -max_out
| `aws_bedrock_claude_excessive_use_of_tokens_filter`
Implementation guide
You must install and configure the Splunk Add-on for AWS (https://splunkbase.splunk.com/app/1876). Enable Amazon Bedrock model invocation logging in AWS so that Claude request/response payloads are delivered to S3 and/or CloudWatch Logs (see https://docs.aws.amazon.com/bedrock/latest/userguide/model-invocation-logging.html for setup steps), then ingest those logs into Splunk via the AWS TA. Configure the aws_bedrock_claude macro to point to the index and sourcetype (json_no_timestamp) where these logs land.
Known false positives
- This detection may produce false positives for identities with low invocation history, legitimate large document summarization tasks, or automated pipeline sessions with naturally variable output token counts.
Analyst notes
Known false positives: This detection may produce false positives for identities with low invocation history, legitimate large document summarization tasks, or automated pipeline sessions with naturally variable output token counts.