GCP Vertex AI Elevated Safety Rating on Response
Description
Detects Vertex AI prompt-response logs where a candidate safety_ratings.probability is HIGH or MEDIUM. Those values are stronger signals than NEGLIGIBLE or LOW alone. Requires the request to include safetySettings so ratings are logged.
Query · esql
from logs-gcp_vertexai.prompt_response_logs-* metadata _id, _version, _index
| where
data_stream.dataset == "gcp_vertexai.prompt_response_logs" and
(
MV_CONTAINS(gcp.vertexai.prompt_response_logs.full_response.candidates.safety_ratings.probability, "HIGH") or
MV_CONTAINS(gcp.vertexai.prompt_response_logs.full_response.candidates.safety_ratings.probability, "MEDIUM")
)
| keep
_id,
_version,
_index,
@timestamp,
cloud.project.id,
gcp.vertexai.prompt_response_logs.model,
gcp.vertexai.prompt_response_logs.api_method,
gcp.vertexai.prompt_response_logs.full_request.contents.parts.text,
gcp.vertexai.prompt_response_logs.full_request.safety_settings.threshold,
gcp.vertexai.prompt_response_logs.full_response.candidates.finish_reason,
gcp.vertexai.prompt_response_logs.full_response.candidates.safety_ratings.category,
gcp.vertexai.prompt_response_logs.full_response.candidates.safety_ratings.probability,
gcp.vertexai.prompt_response_logs.full_response.candidates.safety_ratings.severity
Investigation fields
Pivot points the source recommends for triage.
cloud.project.idgcp.vertexai.prompt_response_logs.modelgcp.vertexai.prompt_response_logs.api_methodgcp.vertexai.prompt_response_logs.full_request.contents.parts.textgcp.vertexai.prompt_response_logs.full_response.candidates.safety_ratings.categorygcp.vertexai.prompt_response_logs.full_response.candidates.safety_ratings.probabilitygcp.vertexai.prompt_response_logs.full_response.candidates.safety_ratings.severitygcp.vertexai.prompt_response_logs.full_response.candidates.finish_reasongcp.vertexai.prompt_response_logs.full_request.safety_settings.threshold
Implementation guide
Requires prompt_response_logs with candidates.safety_ratings.* populated. Ratings appear when the client sends safetySettings on the GenerateContent request.
Known false positives
- Borderline creative-writing or news-summarization prompts can occasionally score MEDIUM. Confirm category and prompt text before escalating.
Analyst notes
Investigating GCP Vertex AI Elevated Safety Rating on Response
A GenerateContent response included safety_ratings.probability of HIGH or MEDIUM for one or more
Gemini harm categories (for example HARM_CATEGORY_HATE_SPEECH or
HARM_CATEGORY_DANGEROUS_CONTENT). That is a stronger signal than NEGLIGIBLE/LOW and often means
the prompt probed disallowed content — whether or not generation was fully blocked.
Possible investigation steps
- Read the prompt in
gcp.vertexai.prompt_response_logs.full_request.contents.parts.text. - Note
safety_ratings.category,probability, andseverityon the alert, plusfinish_reason.SAFETYwith HIGH/MEDIUM is high confidence;STOPwith HIGH still warrants review of what was returned. - Check
full_request.safety_settings.thresholdon the same event.BLOCK_NONEmeans category blocking was disabled for listed categories. - Pivot to
logs-gcp_vertexai.auditlogs-*GenerateContent forsource.ipandclient.user.emailin the same window to identify the caller. - Look for related alerts (jailbreak heuristics, repeated refusals, BLOCK_NONE) from the same project/model.
False positive analysis
- Authorized red-team or content-moderation workloads that intentionally score MEDIUM/HIGH. Exclude those projects or principals.
- Borderline creative-writing or news-summarization prompts can occasionally score MEDIUM — read the prompt before escalating.
Response and remediation
- If the prompt is malicious and unauthorized: quarantine the calling identity, revoke Vertex credentials, and review IAM.
- If
BLOCK_NONEwas used, require stricter thresholds (for example BLOCK_MEDIUM_AND_ABOVE) via org policy or application defaults.