Uppsats

LLM-Based Adversarial Text Anonymization: Evaluating Attacker Architectures via Explicit and Implicit Signal Reasoning

Master-uppsats

Stockholms universitet/Institutionen för data- och systemvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Introduction: Large language models can infer personal attributes such as age, occupation, income, and location from everyday online text, even when explicit identifiers are removed. These inferences rely not only on directly stated information but also on implicit signals embedded in writing style, vocabulary, and broader language-use patterns. This creates privacy risks for individuals who share narrative text online and highlights the limitations of existing anonymization methods. Research Question: This thesis investigates to what extent separating an adversarial attacker into signal-specialized components affects post-anonymization privacy protection compared with a unified attacker. Method: This thesis adopts an exploratory experimental methodology based on a controlled comparative evaluation of adversarial anonymization architectures. Four adversarial anonymization pipeline architectures were implemented and evaluated: an explicit-only baseline, a combined single-prompt attacker, a parallel dual-attacker architecture, and a sequential coordinated dual-attacker architecture. All configurations used GPT-4o as attacker and anonymizer and were evaluated over two anonymization rounds on 100 synthetic profiles from the SynthPAI benchmark, using five metrics: Top-1 adversarial accuracy, Top-3 adversarial accuracy, evidence rate, average attacker certainty, and combined text utility. Results: All four configurations converge to a similar post-anonymization Top-3 adversarial accuracy range of approximately 0.24–0.27. Increasing attacker specialization and coordination does not substantially improve anonymization performance relative to the explicit-only baseline. However, the dual-attacker architectures reveal a consistent asymmetry: explicit textual evidence decreases after anonymization, whereas implicit writing-style signals persist more strongly across configurations. McNemar’s test confirmed that no pairwise configuration comparison reached statistical significance. Discussion: The findings suggest that the primary limitation of the evaluated adversarial anonymization framework lies less in attacker awareness than in the anonymizer’s limited ability to suppress distributed stylistic signals through localized rewriting. Future improvements in LLM-based anonymization may therefore depend more on redesigning anonymizers to address broader language patterns, through style-transfer methods or controllable generation, than on increasing the complexity of attackers.

Information

Lärosäte / institution
Stockholms universitet/Institutionen för data- och systemvetenskap
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.