Sammanfattning

Gaze estimation technology identifies where a subject is looking within a 3D space or a 2D screen interface. Variations in eye appearance due to differences in eye structure and environmental conditions often challenge the accuracy of these models. Data Augmentation (DA), particularly in Deep Learning (DL) models, enhances robustness by artificially increasing the variability of training data. However, conventional DA techniques may alter critical eye details, potentially degrading model performance. This thesis investigates the use of Computer-Generated Imagery (CGI) to create realistic occlusion items---such as hats, eyeglasses, and handheld objects---for integration into gaze estimation model training post-capture. By augmenting the ETH-XGaze dataset with these CGI occlusions, we aimed to improve the generalization performance of gaze estimation models. Our study involved generating various occlusion scenarios within a controlled 3D environment and incorporating these CGI occlusions into training datasets. We assessed the effectiveness of models trained with CGI-augmented data. We then compared their performance to models trained on non-augmented data and proprietary, non-public datasets containing real-world occlusions. The results demonstrated that training with CGI occlusions enhances the ’models ability to handle occluded data compared to baseline models without compromising performance on non-occluded data. These findings suggest that CGI techniques can effectively bridge the diversity gap in existing training datasets, offering a practical solution for developing robust gaze estimation models capable of operating under varied real-world conditions.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.