Uppsats

Forensic Analysis of Computer File Systems Using Vector Search and Large Language Models (LLMs) For Enhanced Cybersecurity and Data Recovery

Magister-uppsats

Högskolan i Skövde/Institutionen för informationsteknologi

Publicerad: 2025

Språk: Engelska

Sammanfattning

This research presents the design, implementation, and evaluation of an Artificial Intelligence (AI)-enhanced forensic tool for detecting sensitive and hidden information across modern computer file systems. It addresses limitations of traditional methods—often reliant on filenames, metadata, or static pattern matching—that struggle against renamed, encrypted, or obfuscated files common in contemporary cybersecurity incidents. The tool integrates semantic vector search and Large Language Models (LLMs) for advanced content discovery. Vector embeddings, generated using the all-MiniLM-L6-v2 Sentence Transformers model, enable semantic similarity matching against investigator queries. Two inference modes were implemented: (1) local/offline processing via Cohere’s command-r-plus model deployed on Oracle’s GenAI Demo-in-a-Box, and (2) cloud-based inference via Cohere’s command-nightly model on Oracle Cloud Infrastructure (OCI) for scalable multilingual classification. Deterministic regular expressions complement LLM analysis by extracting structured entities with contextual interpretation provided by the LLM. To ensure a methodologically sound comparison, the evaluation incorporated both the AI-driven solutions and traditional forensic tools. The Sleuth Kit (TSK) was included for its established capabilities in disk image acquisition and analysis, providing a structural and metadata-based baseline. Recognizing that TSK addresses a broader forensic scope than targeted content search, two additional tools—bulk extractor and ripgrep—were selected for their direct alignment with the AI tool’s core function of detecting content-based artefacts such as email addresses, IP addresses, and PIIin heterogeneous datasets. This dual-selection approach strengthens the validity of the evaluation by enabling both complementary and methodologically equivalent comparisons, ensuring a fair and comprehensive assessment across distinct forensic paradigms. The evaluation used the following dataset: a mixed-format multi-file set (CSV, DOC, DOCX, PDF, TXT JSON) assembled from multiple datasets referenced in the study. Evaluation includes deliberately renamed and partially corrupted files, showed that semantic retrieval with LLM reasoning outperformed traditional tools in identifying contextually relevant threats and revealing obfuscated artifacts, while preserving compatibility with established forensic workflows. Key contributions include: (i) the development of a forensic tool integrating semantic and contextual analysis, (ii) a hybrid detection pipeline combining vectorsearch, regex, and LLM inference, and (iii) an evaluation framework that benchmarks AI methods against traditional forensic tools.

Information

Lärosäte / institution
Högskolan i Skövde/Institutionen för informationsteknologi
Publiceringsdatum
2025
Uppsatstyp
Magister-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.