Uppsats

Benchmarking Self-Supervised Log Embeddings Across Log Modalities and Evaluation Targets

H

Chalmers tekniska högskola / Institutionen för industri- och materialvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Log data are increasingly used to support monitoring, troubleshooting, anomalydetection, and operational decision making in software, industrial, and process systems.These records contain heterogeneous information, including structured fields,message text, event sequences, timestamps, and numerical values. Embedding methodscan convert such records into fixed-dimensional vectors for downstream analysis,but the usefulness of an embedding depends on which parts of the original inputunit remain accessible after representation construction.Evaluating log embeddings is challenging because log data do not form a singlehomogeneous modality. Input units can range from individual rows to fixed windowsand larger operational traces, and each scale exposes different event-level, sequencelevel,numerical, and behavioural structure. At the same time, downstream scorescan reflect only one aspect of an embedding space. A classifier score, retrieval result,clustering metric, numerical probe, weak outcome target, or runtime measurementmay therefore lead to different conclusions about the same representation.The aim of this thesis is to benchmark and analyse log embeddings across heterogeneouslog settings, with emphasis on what different representation methods preserve.The study compares strong transparent baselines with self-supervised learned embeddingsbased on reconstruction, contrastive, and hybrid objectives. Four datasetsare used to cover structured row-level logs, operational row and window logs, blocklevelsystem event sequences, and process-oriented subprocess logs. The embeddingsare evaluated using probes, retrieval, clustering, numerical recovery, weak outcometargets, and runtime.The results show that log embedding quality is conditional on the relation betweenthe input unit, the available observable structure, and the evaluation target. Simplebaselines remain strong when the target is close to explicit fields, tokens, counts, ornumerical summaries. Learned embeddings provide clearer benefits when the targetdepends on context or relations across events. The main conclusion is methodological:log embeddings should be evaluated as profiles of preserved information, withmethod selection grounded in the intended analysis use.

Information

Författare
Zhong, Hongliang
Lärosäte / institution
Chalmers tekniska högskola / Institutionen för industri- och materialvetenskap
Publiceringsdatum
2026
Uppsatstyp
H
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.