Uppsats

Beyond Transcripts : Speech-Based Retrieval-Augmented Generation for L2 English Pronunciation Diagnosis and Feedback

Master-uppsats

Stockholms universitet/Institutionen för lingvistik

Publicerad: 2026

Språk: Engelska

Sammanfattning

This thesis investigates whether speech embeddings can be used as retrieval evidence for L2 English pronunciation diagnosis and feedback generation in a RAG-style framework. In this thesis, L2 refers specifically to English as the target second language, while the experiments compare learners from different L1 backgrounds. It extends the existing PER-MDD (Mispronunciation Detection and Diagnosis) pipeline into a three-stage speechbased retrieval system. Stage 1 tests whether utterance-level speech embeddings can retrieve learners’ L1 background. Stage 2 compares a global mixed-L1 pool with L1-specific pools for phoneme-level pronunciation diagnosis. Stage 3 builds a prototype pipeline for feedback generation. The experiments use L2-ARCTIC and L2-ARCTIC-Plus, and compare Whisper, HuBERT, and wav2vec 2.0 encoders. The results show that Whisper embeddings perform well in utterance-level L1 retrieval and clearly outperform an ASR-transcript baseline, which is close to random guessing. This suggests that important acoustic information may be lost when speech is converted into text. At the phoneme level, HuBERT-based retrieval can support pronunciation analysis, but L1-specific pools do not show a clear advantage over the global mixed-L1 pool. Due to time limits, the feedback pipeline is only implemented as an unevaluated prototype. The results show the technical feasibility of RAG-style MDD, but a practical tutor would still need more reliable diagnosis, human evaluation, and testing on spontaneous learner speech.

Information

Författare
Sha, Maoxuan
Lärosäte / institution
Stockholms universitet/Institutionen för lingvistik
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska