Adversarial attacks targeting explainable AI techniques

Record summary

A quick snapshot of what this page covers.

Techniques3Attack methods connected to this risk.

Mitigations1Defenses that may help with related attacks.

Domain2. Privacy & SecurityThe broad risk area this belongs to.

Risk profile

How this risk is described and categorized.

"Adversarial attacks can affect not only the model’s output but also its corresponding explanation. Current adversarial optimization techniques can intro- duce imperceptible noise to the input image, so that the model’s output does not change but the corresponding explanation is arbitrarily manipulated [61]. Such manipulations are harder to notice, as they are less commonly known compared to standard adversarial attacks targeting the model’s output."

Domain2. Privacy & Security

Subdomain2.2 > AI system security vulnerabilities and attacks

Entity1 - Human

Intent1 - Intentional

Timing3 - Other

CategoryModel Evaluations (Interpretability/Explainability)

SubcategoryAdversarial attacks targeting explainable AI techniques

Related techniques

Attack methods connected to this risk.

AML.T0010.000 - Hardware

feasible

Methodtaxonomy_keyword_ruleConfidence55%

AML.T0001 - Search Open AI Vulnerability Analysis

demonstrated

Methodtext_similarity_sqliteConfidence54%

AML.T0003 - Search Victim-Owned Websites

demonstrated

Methodtext_similarity_sqliteConfidence54%

Suggested mitigations

Defenses that may help with related attacks.

Limit Public Release of Information

Business and Data Understanding

LifecycleBusiness and Data UnderstandingCategoryPolicy

Source

Research source for this risk, when available.

Included resource

Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems

AuthorsGipiškis et al.Year2024TypeJournal Article

DOIhttps://doi.org/10.48550/arXiv.2410.23472 URLhttps://arxiv.org/abs/2410.23472

Original source

MIT AI Risk Repository

Open the public repository used for AI risk records and taxonomy fields.

Repositoryhttps://airisk.mit.edu/