Uppsats
Assessing and mitigating the effect of varying sequencing depth in 16S amplicon data for microbiota studies.
Yrkesexamen på avancerad nivå
Uppsala universitet/Mikrobiologi och immunologi
Publicerad: 2025
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
When performing biodiversity analysis of 16S rRNA gene amplicon sequencing data, onestep to ensure as accurate results as possible is utilising normalisation. With each samplehaving different sequencing depth, or library size, the normalisation step is a way to preventerrors caused by possible biases in the difference in sequencing depth. For example, a samplewith more sequences might appear more diverse than one with a lower sequencing depth. The normalisation step is however a point of contention. There are multiple ways to performit, though one of the more common is to apply rarefaction. Simply, random subsampling ofeach sample in a dataset is performed to a set size. However, this means discarding data. The aim of this project was to apply repeated rarefaction, to keep the benefits of performingnormalisation, while also avoiding discarding crucial data. By keeping more data, a lowerthreshold on this normalisation step would be feasible, and in other words open the possibilityfor inclusion of samples with sequencing depth normally thought of as too low to be used. For this, a script was written to perform the repeated rarefaction and produce graphs andstatistics for analysis. The script can perform the repeated rarefaction on any dataset input thatfits a certain format. Two datasets, one from mosquito tissues and one from mosquitoartificial breeding site water samples, were tested with this script. When running differentamounts of repeats and thresholds for normalisation, a pseudo-F statistic was calculated toinvestigate how the different variables changed the clustering. For both datasets low thresholds (175 sequences and 75 sequences, for the mosquito tissueand mosquito artificial breeding site water respectively), were possible without losingseparation between known clusters in the graph. Also, repeated rarefaction made the graphs more consistent, especially for low threshold runs.Without rarefaction, the pseudo-F statistic index varied greatly, but by the introduction of afew repeats the index became a lot more consistent.
Information
- Författare
- Pölder, Magdalena
- Lärosäte / institution
- Uppsala universitet/Mikrobiologi och immunologi
- Publiceringsdatum
- 2025
- Uppsatstyp
- Yrkesexamen på avancerad nivå
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Yrkesexamen på avancerad nivå, Uppsala universitet/Molekylär evolution
Andersson, Evelina
Publicerad: 2025
Master-uppsats, SLU/Other
Walsh, Benjamin
Publicerad: 2025
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Warnerfjord, Marcus
Publicerad: 2025
Kandidat-uppsats, KTH/Proteinvetenskap
Veljkovic, Anna, Reshef, Lihi, Taresh, Yasmin
Publicerad: 2025
Kandidat-uppsats, KTH/Proteinvetenskap
Cronstrand Elvung, Moa, Käck, Emil, Klangby, Sofia
Publicerad: 2025
Master-uppsats, Uppsala universitet/Institutionen för biologisk grundutbildning
Sarcani, Bianca Ioana
Publicerad: 2026