Uppsats

Robustness of Concept Steering Attacks in Vision-Language Models

Yrkesexamen på avancerad nivå

Uppsala universitet/Avdelningen för beräkningsvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Vision-language models are multi-modal large language models that can answer questions about and make language-prompted analysis on images and videos. This added modality, also introduces novel vulnerabilities. This thesis focuses specifically on concept steering attacks on the vision encoder. In the attack, an image is provided together with a description of the source and target concept. The attack then aims to fool the vision-language model into thinking that the images depict the target concept instead of the source. Optimizable filters, originally designed to defeat watermarking of AI-generated images, have shown some success in defeating diffusion-based purification for attacks against image classifiers. Therefore, this project aims to determine the effect of optimizable filters on attacks against vision-language models. Specifically, for concept steering attacks with a variety of denoising defenses. This study found that the optimizable filters improved the imperceptibility of the images, but severely decreased the adversarial examples' robustness against denoising defenses.

Information

Författare
Carlsson, Jesper
Lärosäte / institution
Uppsala universitet/Avdelningen för beräkningsvetenskap
Publiceringsdatum
2026
Uppsatstyp
Yrkesexamen på avancerad nivå
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.