Uppsats

AI driven Metadataextraktion från PDF filer : Analysverktyg för översättning av metadata till bokindustrin

Yrkesexamen på grundnivå

Luleå tekniska universitet/Institutionen för ekonomi, teknik, konst och samhälle

Publicerad: 2026

Språk: Svenska

Sammanfattning

When a book is going to be sold online, it needs to have the right information attached to it – what it is called, who wrote it, what it is about and so on. This information is called metadata, and it is essential for books to be discoverable on the internet. The problem is that this information is often entered manually today, which takes time and is easy to get wrong.Publit Sweden AB is a Swedish company that helps publishers and authors release and distribute books digitally. They wanted to explore whether AI could be used to automate this work – letting a computer read a PDF file and extract the necessary information on its own, instead of a person having to type everything in manually.That is what this thesis is about. The goal was to build a tool that can do exactly that, and that also presents information in the format used by the publishing industry, primarily something called ONIX and Thema. ONIX is a standard format for sharing book information between publishers, distributors and retailers. Thema is a system for categorising books by subject, a bit like a digital bookshelf structure.To reach a good result the work followed a clear process. First, knowledge was gathered through conversations with people at Publit and a UX expert at Luleå University of Technology, by reading research on information design and AI, and by testing five similar tools already available on the market. What became clear was that none of the existing tools were built specifically for the publishing industry. They were either too simple and lacked AI, or too complicated for someone without a technical background.Ideas were then developed, and two different prototypes were built and tested with six users. Based on their feedback the best solution was selected and improved. The final tool was then built as a web application, meaning it works directly in the browser without needing to install anything. Here is how the tool works in practice: the user uploads a PDF file, the AI reads through it and suggests metadata in simple editable fields, the user checks the suggestions and corrects anything that is wrong, and then the information can be downloaded in the right format. No information is stored on any external server, which is important since many book manuscripts are confidential and not yet published.The most important thing this project shows is that something that previously took a long time and required specialist knowledge can be made much faster and simpler with the help of AI. The technology is not the hard part – what was needed was to design the tool in a way that works for the people who are going to use it.Keywords: metadata, PDF analysis, artificial intelligence, ONIX, Thema, information design, publishing industry, web application, user-centred design.

Information

Lärosäte / institution
Luleå tekniska universitet/Institutionen för ekonomi, teknik, konst och samhälle
Publiceringsdatum
2026
Uppsatstyp
Yrkesexamen på grundnivå
Språk
Svenska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.