Supervision evasion propensity

Record summary

A quick snapshot of what this page covers.

Techniques1Attack methods connected to this risk.

Mitigations2Defenses that may help with related attacks.

Domain7. AI System Safety, Failures, & LimitationsThe broad risk area this belongs to.

How this risk is described and categorized.

Domain7. AI System Safety, Failures, & Limitations

Subdomain7.1 > AI pursuing its own goals in conflict with human goals or values

Entity2 - AI

Intent1 - Intentional

Timing3 - Other

CategoryModel Propensities

SubcategorySupervision evasion propensity

Attack methods connected to this risk.

demonstrated

Methodtext_similarity_sqliteConfidence53%

Defenses that may help with related attacks.

DeploymentML Model Evaluation

LifecycleDeployment + 1 moreCategoryTechnical - ML

Business and Data UnderstandingDeployment+1 more

LifecycleBusiness and Data Understanding + 2 moreCategoryTechnical - Cyber

Research source for this risk, when available.

Included resource

AuthorsSAIL & Concordia AIYear2025TypeReport

Original source

Open the public repository used for AI risk records and taxonomy fields.