Uppsats

Designing for Privacy in Tabular Federated Learning

Master-uppsats

Jönköping University/JTH, Avdelningen för datavetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

In federated learning (FL), private data leakage through gradient inversion depends strongly on design choices and data properties. We study this risk under an honest-but-curious server threat model, with a focus on clinical tabular data from MIMIC-IV and complementary benchmark tasks for binary classification, multiclass classification, and regression. Unlike image reconstruction, tabular reconstruction cannot be judged well by visual inspection. Evaluation must account for numerical tolerances, categorical exact matches, baseline recoverability, and the difference between feature level recovery and exact row recovery. We evaluate FedSGD gradients and FedAvg model deltas using an exposure aligned training protocol, where attacked models are compared after matched client data exposure rather than matched communication rounds. We compare MLP, ResNet, and FT-Transformer models, and isolate architecture effects further through an MLP grid over width, depth, activation, normalization, and dropout. The results show that local aggregation and client data exposure are among the strongest determinants of tabular gradient leakage. Small client batches and updates representing few distinct records are most vulnerable, while stronger local aggregation and greater client data exposure usually reduce reconstruction in ways that vary across architectures and datasets. FT-Transformer is consistently harder to invert than the MLP and ResNet baselines, suggesting that embedded categorical representations change the geometry of tabular gradient inversion. At the same time, the MLP architecture grid shows that reconstructability also changes substantially within a single model family, so architecture affects privacy beyond the transformer comparison alone. Aggregate reconstruction accuracy can overstate record recovery, especially on MIMIC-IV, where sparse and repeated clinical features produce a strong client marginal prior. Strict exact match results show that complete row recovery is concentrated in the most exposed one-hot-encoded settings, while FT-Transformer can have nonzero feature recovery without exact row recovery. Overall, the findings show that gradient inversion risk in tabular FL is shaped by model architecture, FL training configuration, data exposure, and input data properties.

Information

Lärosäte / institution
Jönköping University/JTH, Avdelningen för datavetenskap
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.