Uppsats

From Legacy Databases to Modern Standards: A Bioinformatics Pipeline for Migrating Immunodeficiency Variant Data to LOVD 3.0 with GRCh38 Standardisation

Master-uppsats

Lunds universitet/Bioteknik

Publicerad: 2026

Språk: Engelska

Sammanfattning

Primary immunodeficiency diseases (PIDs) are a heterogeneous group of inherited immune system disorders caused by pathogenic variants in over 400 genes, collectively affecting an estimated one in 10,000 individuals worldwide. The IDbases collection, maintained on the legacy MUTbase platform at Lund University, represents one of the most comprehensively curated repositories of such variants, cataloguing disease-causing variants across 136 immunodeficiency-related genes with 6,259 patient records accumulated over more than three decades of expert curation. Despite their scientific value, these databases remain in a format incompatible with current genomic standards: variant coordinates are anchored to obsolete partial ENA clone reference sequences, nomenclature does not conform to Human Genome Variation Society (HGVS) recommendations, and the data cannot be integrated with modern interoperable platforms. This thesis presents a systematic seven-step bioinformatics pipeline for migrating the IDbases data into the Leiden Open Variation Database 3.0 (LOVD 3.0), the leading open-source platform for locus-specific variant databases. The pipeline extracted variants from all 136 gene databases, resolved reference sequence identities through an empirical offset algorithm, lifted over chromosomal coordinates to GRCh38, and normalised all variant descriptions to HGVS nomenclature through three parallel Mutalyzer 3 annotation tracks. Track A (NG_IDRefseq genomic route) achieved a 99.3% normalisation success rate (3,582/3,606 variants). Track B (NM_MANE) achieved 94.1% (4,496/4,775). Track C (NM_IDRefseq) served as a tertiary traceability fallback at 59.5% (2,862/4,811). A central finding of this work is that IDbases IDRefSeq sequences are partial ENA clone fragments contained within full LRG NG_ reference sequences - a previously undocumented structural relationship that causes systematic out-of-bounds coordinate failures for variants added after the original database construction date. The empirical offset algorithm developed to resolve this achieved a match rate of ≥90% for 92 of 94 testable genes (genes where an IDRefSeq FASTA was available) in the genomic track. Following priority-based merging of the three tracks, 7,776 patient–variant records were resolved and, after deduplication, correspond to 2,240 distinct variants across 115 genes, linked to 2,729 patients (5,773 patient–variant pairs). A further 698 distinct variants remained unresolved, of which roughly three-quarters are automatically rescuable, giving an automated resolution rate of 76.2% with an estimated addressable ceiling near 94%. Systematic data quality issues uncovered during the migration - including strand orientation inconsistencies in 43 reverse-strand genes, missing insertion sequences, legacy IVS intronic notation, and transcript version obsolescence - are catalogued as findings to inform future variant database design. The pipeline is implemented as a series of documented, modular Python scripts and is designed for reuse in similar legacy database migration projects. The work makes a large, expertly curated immunodeficiency variant dataset available in a modern, interoperable format aligned with GRCh38 and MANE (Matched Annotation from the NCBI and EMBL-EBI) Select transcript standards.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.