Uppsats
Topology-Aware Evaluation of Core-to-Core Communication and Application Scaling on Chiplet and Monolithic Hybrid Processors
Master-uppsats
Uppsala universitet/Institutionen för informationsteknologi
Publicerad: 2026
Språk: Engelska
Sammanfattning
Modern multicore processors increasingly expose non-uniform performance behavior within a single socket due to chiplet organization, deep cache hierarchies, and heterogeneous core designs. This thesis studies how such topology-dependent latency effects appear in both microbenchmarks and application-level performance. The evaluation covers two AMD chiplet platforms, the EPYC 9454P and Ryzen Threadripper PRO 9975WX, and one monolithic hybrid platform, the Intel Core i9–14900K. The study first uses cache-latency and core-to-core latency microbenchmarks to characterize local memory-access behavior and communication latency between cores. The AMD chiplet platforms show clear CCD-level communication structure: intra-CCD communication is consistently lower latency than inter-CCD communication across the tested core-to-core variants. The Intel platform does not show CCD-level structure, but still exhibits non-uniformity due to P-core, E-core, and E-core-cluster relationships. On the AMD chiplet platforms, inter-CCD communication is typically several times slower than intra-CCD communication, with grouped inter/intra ratios ranging from about 3.7× to 7.2× across the tested variants. The thesis then evaluates SPLASH-3, SPLASH-4, NPB, and XSBench under topology-aware CPU-set placements on the AMD chiplet platforms. The results show that placement sensitivity is strongly workload dependent. Some workloads may benefit from compact placement, which reduces cross-CCD communication, while others perform better when threads are spread across CCDs, likely due to improved use of distributed cache and memory-system resources. Several workloads, including XSBench under the tested configuration, show little placement sensitivity. Finally, the thesis compares application scaling across the three platforms using a frequency-adjusted runtime comparison over a common thread-count range. These comparisons show that frequency adjustment helps make scaling trends more comparable, but does not eliminate platform-dependent differences. Overall, the results demonstrate that topology-dependent latency is an important factor in modern multicore performance, but it is not by itself a complete predictor of application runtime. Application behavior depends on the interaction between communication patterns, cache capacity, memory-system behavior, thread placement, and platform organization.
Information
- Författare
- Bo, Xu
- Lärosäte / institution
- Uppsala universitet/Institutionen för informationsteknologi
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska