AI policy detection settings determine how TrendAI Vision One™ inspects your AI applications' prompts and responses.
A policy has three detection tabs which apply rules to the content it detects:
Rule actions
Harmful content and Sensitive information hold a set of rules. Each rule pairs an
action with the content it detects. Prompt attacks applies a single action and has
no rules, levels, or categories.
|
Action
|
Description
|
|
Allow & log
|
Allows content that matches the rule and records the detection in the logs.
|
|
Deny & log
|
Blocks content that matches the rule and records the detection in the logs.
|
Harmful content
The Harmful content tab detects harmful material in prompts and responses. Each rule
sets a detection level and applies to one category.
The Detection level sets how aggressively the rule flags content.
Sensitive information
The Sensitive information tab detects sensitive data in prompts and responses. Each
rule targets one category of sensitive data at a fixed sensitivity level. You cannot
change a rule's sensitivity level or category, but you can enable or disable the rule
and set its action.
To find rules on this tab, use Rule name to search by rule name or click Add filter and select one or more of the following:
-
State
-
Action
-
Sensitivity level
-
Category
-
Region
You can also turn on redaction on this tab. Redaction replaces sensitive data with
asterisks before the policy applies the rule action.
-
When the action is Deny & log, the policy blocks the sensitive data from reaching the AI model and logs it as asterisks.
-
When the action is Allow & log, the policy sends and logs asterisks instead of the sensitive data.
Prompt attacks
The Prompt attacks tab detects attempts to manipulate an AI application, such as prompt
injection. Unlike the other tabs, it does not use rules, detection levels, or categories.
Instead, you set a single action that applies to every detected prompt attack.
