Uppsats

Navigating Human Pose Estimation : A Comparative Study of Solution Pipelines

Master-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2024

Språk: Engelska

Sammanfattning

Human pose estimation, an important domain within computer vision and artificial intelligence, involves detecting and estimating the positions of individuals in images or videos. This field has witnessed rapid advancement in recent years, marked by the emergence of novel and innovative methodologies, leading to a diverse field where a variety of approaches are being utilized. This diversity within the field extends to both model architectures and solution pipelines. Given the rapid development of the field of human pose estimation, critical analysis of newly emerged solution approaches is sorely needed. This Master’s thesis contributes by conducting a comparative study of two prevalent solution pipeline approaches, the top-down and bottom-up methods. The objective is to provide further insights into the optimal circumstances for the utilization of each pipeline, thereby bridging knowledge gaps in human pose estimation and facilitating continued growth within the field. This project implemented one state-of-the-art human pose estimation model for each pipeline category and evaluated its performance across various challenging image scenarios. Specifically, we examine performance in crowded scenes, occluded environments, and individuals of varying scales. Our findings indicate that no singular pipeline universally excels; rather, the optimal choice depends on the characteristics of the input images. The top-down approach demonstrates superiority in handling highly crowded and occluded images. The results also suggest that top-down models are well- suited for estimating human poses on smaller-scale individuals. Moreover, we observe that the inference time of top-down models is directly proportional to the number of individuals in the input, whereas bottom-up networks display consistent inference times irrespective of crowd density. Consequently, top- down models offer faster performance for images with fewer individuals, whereas bottom-up models excel in more crowded settings.

Information

Författare
Sjöblom, Jonas
Lärosäte / institution
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publiceringsdatum
2024
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.