Uppsats
Improving AI-based Synthesizability Scores for Next-Generation Protein Degradation Drugs
Magister-uppsats
Göteborgs universitet/Institutionen för data- och informationsteknik
Publicerad: 2026-06-29
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
PROteolysis TArgeting Chimeras (PROTACs) are an emerging class of therapeuticmodality that enables targeted protein degradation. A PROTAC consists of threecomponents: a target-binding warhead, a chemical linker, and an E3 ligase ligand.By simultaneously binding a Protein Of Interest (POI) and an E3 ubiquitin ligase,the PROTAC induces ubiquitination of the target protein, leading to its recognitionand degradation by the proteasome. Despite their therapeutic potential, the modularstructure of PROTACs often necessitates complex multi-step synthesis. Evaluatingtheir synthesizability typically requires extensive laboratory synthesis and validation,making the process time-consuming, costly, and expertise-intensive.In this thesis, we present a systematic framework for PROTAC synthesizability assessment based on retrosynthetic planning and machine learning. A dataset with 27,099PROTACs was collected, curated, and preprocessed to generate retrosynthesis-basedsynthesizability labels and scores. PROTAC-Splitter was used to decompose eachPROTAC into its constituent warhead, linker, and E3 ligase ligand components, whileAiZynthFinder was employed to evaluate retrosynthetic solvability and generate synthesizability scores for both complete molecules and individual components. Relevantfeatures, including molecular fingerprints, molecular descriptors, and component-levelsynthesizability information, were subsequently used to train machine learning modelsbased on Random Forest, XGBoost, and Multi-Layer Perceptron architectures.Retrosynthetic analysis revealed a 57% agreement between component-level andwhole-PROTAC synthesizability, demonstrating that component-level assessmentcaptures meaningful information about overall synthetic feasibility. Linkers wereidentified as the primary source of synthetic difficulty, highlighting their importancein determining PROTAC synthesizability. Overall, our best ML classification modelachieved a ROC-AUC of 0.958, the best regression model achieved an (R2) score of0.497, increasing to 0.661 after filtering noisy labels. These results demonstrate thatcomponent-level retrosynthetic information can be leveraged to construct efficientsurrogate models for PROTAC synthesizability assessment, providing a scalablealternative to computationally expensive retrosynthetic planning.
Information
- Författare
- Zhu, Jia Xin, Mo, Tingting
- Lärosäte / institution
- Göteborgs universitet/Institutionen för data- och informationsteknik
- Publiceringsdatum
- 2026-06-29
- Uppsatstyp
- Magister-uppsats
- Språk
- Engelska