GCP Vertex AI High Request and Token Volume
Description
Detects a high number of Vertex AI prompt-response log rows and high total token usage for one model in the lookback window. Totals use provider usage_metadata.total_token_count (includes thinking tokens when billed). High request and token volume can indicate model extraction, quota abuse, or a compromised caller. Tune thresholds for your baseline.
Query · esql
from logs-gcp_vertexai.prompt_response_logs-*
| where
data_stream.dataset == "gcp_vertexai.prompt_response_logs" and
gcp.vertexai.prompt_response_logs.full_response.usage_metadata.prompt_token_count > 0
| stats
Esql.event_count = count(*),
Esql.prompt_tokens_sum = sum(gcp.vertexai.prompt_response_logs.full_response.usage_metadata.prompt_token_count),
Esql.candidates_tokens_sum = sum(gcp.vertexai.prompt_response_logs.full_response.usage_metadata.candidates_token_count),
Esql.total_tokens_sum = sum(gcp.vertexai.prompt_response_logs.full_response.usage_metadata.total_token_count),
Esql.timestamp_first_seen = min(@timestamp),
Esql.timestamp_last_seen = max(@timestamp)
by
gcp.vertexai.prompt_response_logs.model,
cloud.project.id
| where Esql.event_count >= 100 and Esql.total_tokens_sum >= 10000
| keep
gcp.vertexai.prompt_response_logs.model,
cloud.project.id,
Esql.event_count,
Esql.prompt_tokens_sum,
Esql.candidates_tokens_sum,
Esql.total_tokens_sum,
Esql.timestamp_first_seen,
Esql.timestamp_last_seen
Investigation fields
Pivot points the source recommends for triage.
gcp.vertexai.prompt_response_logs.modelcloud.project.idEsql.event_countEsql.prompt_tokens_sumEsql.candidates_tokens_sumEsql.total_tokens_sumEsql.timestamp_first_seenEsql.timestamp_last_seen
Implementation guide
Requires GCP Vertex AI prompt_response_logs with usage_metadata token counts. Raise Esql.event_count and token
thresholds above your normal peak before broad enablement.
Known false positives
- Approved batch evaluation or high-traffic applications. Raise thresholds above the normal peak per model.
Analyst notes
Investigating GCP Vertex AI High Request and Token Volume
One model exceeded both the request-count and total-token thresholds in the lookback window.
Esql.total_tokens_sum comes from usage_metadata.total_token_count, which includes thinking tokens
when the provider reports them. High volume can indicate model extraction, quota abuse, cost impact,
or a compromised caller. This rule aggregates usage metadata; individual prompt text is not required
on the alert.
Possible investigation steps
- Note
gcp.vertexai.prompt_response_logs.model,cloud.project.id,Esql.event_count,Esql.prompt_tokens_sum,Esql.candidates_tokens_sum, andEsql.total_tokens_sum. - Compare request count vs token sums: many short calls (scraping/extraction) vs fewer huge prompts
(context stuffing / data exfil via prompt). When
Esql.total_tokens_sumis much larger than prompt+candidates, thinking tokens may dominate usage. - Pivot to
logs-gcp_vertexai.auditlogs-*GenerateContent forsource.ipandclient.user.emailconcentration in the same window. - Sample
prompt_response_logsfor that model/time range to see whether content looks like systematic extraction. - Check for recent
SetPublisherModelConfigchanges that might have altered logging visibility.
False positive analysis
- Approved batch evaluation, fine-tuning data prep, or high-traffic production apps. Raise
Esql.event_countand token thresholds above the normal peak per model before broad enablement.
Response and remediation
- Restrict or rotate caller credentials, apply quotas, and review spend/billing for the model.
- If compromise is suspected, hunt lateral use of the same principal across other GCP services.