Uppsats

A Defense Tier Classification of REP Efficacy Against Generative AI Web Crawlers

Kandidat-uppsats

Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)

Publicerad: 2026

Språk: Engelska

Sammanfattning

Generative AI relies on web crawlers to extract training data, leading to significanttension between the intellectual property rights of digital publishers and generativeAI developers. While alternative protocols and algorithms are researched to addressthe new Web crawlers, most websites still rely upon an outdated protocol developedin the 90s, the Robots Exclusion Protocol (REP), as no alternative has been widely adopted in the web landscape. While the protocol was developed to address tradi-tional web crawlers, it has a structural weakness that leads to vulnerability against AI web crawlers. Using Design Science Research, this study develops an Analyzerto assess the effectiveness of REP as a valid means of web crawlers’ access controlbeyond the standard binary syntax used in many implementations. Additionally, thethesis proposes a classification tier of the domain-level defenses into a 5 Tier System.An algorithmic audit of 809 European news publisher web domains, validated using a random sampling of manually-audited audits, reveals a considerable level of non-compliance: 66.1% of the domains analyzed do not have an effective means of pre-venting the use of AI, and 20 of the 26 countries analyzed show a rate of greater than 50% for failed compliance with opt-outs. The results provide empirical evidence thatthe complexities associated with configuration increase the structural degradation ofsecurity. Further evidence demonstrates that REP remains an inadequate protocol toaddress the new AI web crawlers.

Information

Lärosäte / institution
Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.