Sammanfattning

Artificial Intelligence (AI), specifically the Large language models (LLM) are powerful tools that are rapidly spreading to many domains in industry, enabling increased productivity and efficiency in the workplace. However, due to pre-training on general data as well as its probabilistic nature , limitations arise due to lack of knowledge on specific domains as well as hallucinations, and false interpretation of input queries. This thesis investigates the performance and application of an LLM in the Swedish legal system, by fine tuning the LLM, Gemma 3-4B and implementing a Retrieval Augmented Generation (RAG) system which is then subjected through qualitative and quantitative evaluations in collaboration with law students , law professors and legal experts. The quantitative results gathered demonstrates the limitation of the model due to its shortcomings in terms of specific section location and understanding of more complex legal queries and legal calculations. F1 scores where this was the case were observed to go to a low of 0,03. It is however performing significantly better when dealing with legal term based query where it scores a high of 0,56. The qualitative results point towards a great improvement and potential in the LLMs query results with RAG implementation. It is also highly praised for its reasoning capabilities and transparency in terms of the steps it took to gather the final answer, although the model still presented limitations, presenting room for further improving the models performance and accuracy.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.