Uppsats

ORSExplorer: An LLM-Assisted Rule System Extraction Tool – Semi-Automated Extraction of Regulatory Entities and Relationships for Knowledge Graph Creation

Master-uppsats

Stockholms universitet/Institutionen för data- och systemvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Introduction: Legal and regulatory domains contain large amounts of information that are difficult to interpret and structure. Regulatory rules are often connected to other rules, concepts and definitions, which makes them difficult to analyze as isolated statements. Knowledge graphs can represent this type of knowledge as entities and relations but the construction of regulatory knowledge graphs is still largely manual and dependent on expert interpretation. This thesis addresses this problem by designing, implementing and evaluating ORSExplorer, a semi-automated tool that uses Large Language Models to support the extraction of entities and semantic relationships from regulatory documents. Research Question: The primary research question is: “How can a semiautomated Large Language Model-based tool support expert users in extracting and structuring entities and relationships from regulatory documents for knowledge graph creation?” Method: This thesis follows a Design Science Research methodology. Requirements were collected through expert workshops and supervision meetings and were used to guide the iterative development of ORSExplorer. The artifact was implemented as a web-based tool that supports ontology-guided extraction, human-in-the-loop review, graph visualization, manual editing and Wikibase-compatible export. The evaluation was conducted through expert assessment and semi-structured interviews with five expert participants. The interview material was analyzed using thematic analysis and selected requirements were also assessed through technical verification. Results: The results show that ORSExplorer can support expert users by generating ontology-guided candidate entities and relationships from regulatory text. These outputs can be inspected, edited and accepted before they are used for knowledge graph creation. The evaluation showed that participants valued the incremental workflow, the possibility to review reasoning and evidence and the ability to manually control the generated output. The main findings were grouped into four themes: usability, workflow efficiency, trust and control and perceived accuracy and output quality. The results also showed limitations in first-time usability, graph readability, the need for ontology knowledge and limited history tracking between iterations. Discussion: The findings suggest that semi-automated Large Language Model-based tools can support regulatory knowledge graph creation when they are designed as support tools rather than fully automated systems. The artifact helped reduce parts of the manual structuring work but expert validation remained necessary because the generated output could not be treated as final legal interpretation. The main contribution of this thesis is ORSExplorer together with design knowledge on ontology guidance, transparency, incremental construction and human-in-the-loop validation. Future research could evaluate the artifact against a manually created gold standard, improve graph readability, history tracking and test the approach in larger regulatory knowledge graph workflows.

Information

Lärosäte / institution
Stockholms universitet/Institutionen för data- och systemvetenskap
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.