Uppsats

Escaping the Trap: Text Analysis for Nepenthes Detection

Kandidat-uppsats

Stockholms universitet/Institutionen för data- och systemvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Web crawlers and scrapers are useful tools for data collection, including data for training artificial intelligence models. Data scraping is not appreciated by everyone. In response, website owners have implemented various countermeasures to crawling and scraping, including Nepenthes, a web crawler trap that generates endless pages with machine-generated nonsense text. The problem addressed in this study is the uncertainty of whether Nepenthes can be accurately detected. It investigates whether text analysis techniques can reliably detect machine-generated text produced by Nepenthes. The research strategy is experimental, focusing on text analysis using three statistical measures as a detection method. The work includes collecting and processing data to construct a dataset used to analyze the statistical properties of human- and machine-generated texts. The detector is subsequently created and optimized, followed by a live test to evaluate its ability to detect machine-generated texts by Nepenthes. The results show that there is a considerable variation in the effectiveness of different statistical measures for detection. The live test shows impressive accuracy in detecting machine-generated texts. The conclusion is that text analysis can be an effective method for detecting Nepenthes, and that combining different statistical measures can improve detection accuracy.

Information

Lärosäte / institution
Stockholms universitet/Institutionen för data- och systemvetenskap
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.