Uppsats
Beyond Transcripts : Speech-Based Retrieval-Augmented Generation for L2 English Pronunciation Diagnosis and Feedback
Master-uppsats
Stockholms universitet/Institutionen för lingvistik
Publicerad: 2026
Språk: Engelska
Sammanfattning
This thesis investigates whether speech embeddings can be used as retrieval evidence for L2 English pronunciation diagnosis and feedback generation in a RAG-style framework. In this thesis, L2 refers specifically to English as the target second language, while the experiments compare learners from different L1 backgrounds. It extends the existing PER-MDD (Mispronunciation Detection and Diagnosis) pipeline into a three-stage speechbased retrieval system. Stage 1 tests whether utterance-level speech embeddings can retrieve learners’ L1 background. Stage 2 compares a global mixed-L1 pool with L1-specific pools for phoneme-level pronunciation diagnosis. Stage 3 builds a prototype pipeline for feedback generation. The experiments use L2-ARCTIC and L2-ARCTIC-Plus, and compare Whisper, HuBERT, and wav2vec 2.0 encoders. The results show that Whisper embeddings perform well in utterance-level L1 retrieval and clearly outperform an ASR-transcript baseline, which is close to random guessing. This suggests that important acoustic information may be lost when speech is converted into text. At the phoneme level, HuBERT-based retrieval can support pronunciation analysis, but L1-specific pools do not show a clear advantage over the global mixed-L1 pool. Due to time limits, the feedback pipeline is only implemented as an unevaluated prototype. The results show the technical feasibility of RAG-style MDD, but a practical tutor would still need more reliable diagnosis, human evaluation, and testing on spontaneous learner speech.
Information
- Författare
- Sha, Maoxuan
- Lärosäte / institution
- Stockholms universitet/Institutionen för lingvistik
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska