Uppsats

Towards Taxonomic Consistency:A reproducible pipeline from raw reads to annotated OTUs using a standardised reference database

Master-uppsats

Uppsala universitet/Institutionen för biologisk grundutbildning

Publicerad: 2026

Språk: Engelska

Sammanfattning

Metabarcoding of environmental DNA (eDNA) has the potential to be an important tool for effective monitoring of fungal biodiversity, without the need of direct observation. Accurate monitoring, however, depends on two critical components: an effective and reproducible bioinformatic preprocessing pipeline as well as a taxonomically consistent and accurate reference database. Fungal refence databases that contain user-sourced sequences such as UNITE are prone to taxonomic inconsistencies, including outdated naming and synonyms. These inconsistencies also affect the species hypothesis (SH), which is UNITE’s internal sequence clustering system. Here, we developed a Conda-packaged pipeline for processing raw PacBio ITS2 reads into annotated OTUs. The pipeline was validated on six eDNA samples sequenced from soil litter. Additionally, we build a Python-based taxonomy standardisation pipeline that updates the taxonomy of UNITE records against the Catalogue of Life taxonomic backbone. The six samples were annotated twice for direct comparison: first using the original UNITE database, and secondly using the taxonomically standardised version. The preprocessing pipeline performed consistently across all samples, retaining the majority of reads through each step. The taxonomy standardisation algorithm modified 4.97% of records in the UNITE database, reducing 865 synonyms and completing missing family-level classifications for over 35 000 records. Reducing the number of synonyms in the reference database did not significantly affect the species richness reported in the samples, but it did alter the annotated families and bootstrap confidence levels. This suggests that, although name standardization may have limited impact on species-level annotations, it can influence studies analysing community composition at higher taxonomic levels, such as family. We expect the effects of taxonomy standardization to become even more pronounced in large comparative studies, where naming inconsistencies may accumulate and introduce biases in biodiversity estimates, especially at higher taxonomic ranks. The proposed approach is computationally efficient and may contribute to more reliable and standardized reference data. In addition, it could be applied more broadly as a general strategy for many types of reference sequence databases, not only fungal databases.

Information

Lärosäte / institution
Uppsala universitet/Institutionen för biologisk grundutbildning
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.