Uppsats

Enhancing Multi-Modal 3D Object Detection with Attention-Based Feature Fusion

Master-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2025

Språk: Engelska

Sammanfattning

Advances in autonomous driving technology have intensified the demand for reliable and robust 3D object detection systems. To enable autonomous vehicles to perceive and understand their surroundings with high accuracy, modern systems integrate multiple sensors. Among these, LiDAR sensors provide precise spatial measurements by generating sparse but highly accurate 3D point clouds, while RGB cameras offer dense semantic information with rich textures and colors. Effectively fusing these complementary modalities is a key research direction in autonomous driving. However, many existing fusion methods rely on simple feature concatenation and lack adaptive crossmodal interactions. This report proposes an improved fusion strategy that leverages attention mechanisms to enhance feature-level interactions between LiDAR point clouds and camera images. Specifically, three attention modules are introduced at different stages: (1) a transformer-based attention mechanism applied to image features; (2) a self-attention mechanism applied to fused features after image and point cloud fusion; and (3) a channel attention mechanism applied to LiDAR features using a Squeeze-and-Excitation (SE) module. Experiments are conducted on two large-scale public datasets, KITTI and Waymo, and evaluated using standard 3D and 2D Average Precision (AP) metrics at multiple Intersection over Union (IoU) thresholds. The results show that applying a transformer-based attention mechanism to image features alone yields limited benefits and can even degrade performance on complex datasets. In contrast, incorporating attention mechanisms into fusion features and LiDAR features consistently improves detection accuracy and robustness. Notably, the best-performing model achieves up to a 0.3% improvement in AP over the baseline on both datasets, without introducing significant computational overhead. This study confirms that attention-based fusion strategies can enhance the expressiveness of multimodal features, leading to more accurate and robust 3D object detection in autonomous driving scenarios.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.