Uppsats
LLM-Based Adversarial Text Anonymization: Evaluating Attacker Architectures via Explicit and Implicit Signal Reasoning
Master-uppsats
Stockholms universitet/Institutionen för data- och systemvetenskap
Publicerad: 2026
Språk: Engelska
Sammanfattning
Introduction: Large language models can infer personal attributes such as age, occupation, income, and location from everyday online text, even when explicit identifiers are removed. These inferences rely not only on directly stated information but also on implicit signals embedded in writing style, vocabulary, and broader language-use patterns. This creates privacy risks for individuals who share narrative text online and highlights the limitations of existing anonymization methods. Research Question: This thesis investigates to what extent separating an adversarial attacker into signal-specialized components affects post-anonymization privacy protection compared with a unified attacker. Method: This thesis adopts an exploratory experimental methodology based on a controlled comparative evaluation of adversarial anonymization architectures. Four adversarial anonymization pipeline architectures were implemented and evaluated: an explicit-only baseline, a combined single-prompt attacker, a parallel dual-attacker architecture, and a sequential coordinated dual-attacker architecture. All configurations used GPT-4o as attacker and anonymizer and were evaluated over two anonymization rounds on 100 synthetic profiles from the SynthPAI benchmark, using five metrics: Top-1 adversarial accuracy, Top-3 adversarial accuracy, evidence rate, average attacker certainty, and combined text utility. Results: All four configurations converge to a similar post-anonymization Top-3 adversarial accuracy range of approximately 0.24–0.27. Increasing attacker specialization and coordination does not substantially improve anonymization performance relative to the explicit-only baseline. However, the dual-attacker architectures reveal a consistent asymmetry: explicit textual evidence decreases after anonymization, whereas implicit writing-style signals persist more strongly across configurations. McNemar’s test confirmed that no pairwise configuration comparison reached statistical significance. Discussion: The findings suggest that the primary limitation of the evaluated adversarial anonymization framework lies less in attacker awareness than in the anonymizer’s limited ability to suppress distributed stylistic signals through localized rewriting. Future improvements in LLM-based anonymization may therefore depend more on redesigning anonymizers to address broader language patterns, through style-transfer methods or controllable generation, than on increasing the complexity of attackers.
Information
- Författare
- Ebrahimitofighi, Negin
- Lärosäte / institution
- Stockholms universitet/Institutionen för data- och systemvetenskap
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Kandidat-uppsats, Högskolan i Gävle/Avdelningen för datavetenskap och samhällsbyggnad
Vambe, Vimbainaishe
Publicerad: 2026
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Hansson, Martin
Publicerad: 2026
Kandidat-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Bäckman, Caroline, Khodarahmi, Nick
Publicerad: 2026
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Saleh, Abdelrahman
Publicerad: 2026
Master-uppsats, Högskolan i Borås/Akademin för bibliotek, information, pedagogik och IT
Anneling, Marie
Publicerad: 2026
Master-uppsats, Lunds universitet/Institutionen för elektro- och informationsteknik
Du, Mingtong
Publicerad: 2026