Uppsats

Assessing and mitigating the effect of varying sequencing depth in 16S amplicon data for microbiota studies.

Yrkesexamen på avancerad nivå

Uppsala universitet/Mikrobiologi och immunologi

Publicerad: 2025

Språk: Engelska

Sammanfattning

When performing biodiversity analysis of 16S rRNA gene amplicon sequencing data, onestep to ensure as accurate results as possible is utilising normalisation. With each samplehaving different sequencing depth, or library size, the normalisation step is a way to preventerrors caused by possible biases in the difference in sequencing depth. For example, a samplewith more sequences might appear more diverse than one with a lower sequencing depth. The normalisation step is however a point of contention. There are multiple ways to performit, though one of the more common is to apply rarefaction. Simply, random subsampling ofeach sample in a dataset is performed to a set size. However, this means discarding data. The aim of this project was to apply repeated rarefaction, to keep the benefits of performingnormalisation, while also avoiding discarding crucial data. By keeping more data, a lowerthreshold on this normalisation step would be feasible, and in other words open the possibilityfor inclusion of samples with sequencing depth normally thought of as too low to be used. For this, a script was written to perform the repeated rarefaction and produce graphs andstatistics for analysis. The script can perform the repeated rarefaction on any dataset input thatfits a certain format. Two datasets, one from mosquito tissues and one from mosquitoartificial breeding site water samples, were tested with this script. When running differentamounts of repeats and thresholds for normalisation, a pseudo-F statistic was calculated toinvestigate how the different variables changed the clustering. For both datasets low thresholds (175 sequences and 75 sequences, for the mosquito tissueand mosquito artificial breeding site water respectively), were possible without losingseparation between known clusters in the graph. Also, repeated rarefaction made the graphs more consistent, especially for low threshold runs.Without rarefaction, the pseudo-F statistic index varied greatly, but by the introduction of afew repeats the index became a lot more consistent.

Information

Lärosäte / institution
Uppsala universitet/Mikrobiologi och immunologi
Publiceringsdatum
2025
Uppsatstyp
Yrkesexamen på avancerad nivå
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.