Uppsats

Video Stitching for Remote Operation of Cabinless Autonomous Trucks

Yrkesexamen på avancerad nivå

Uppsala universitet/Avdelningen för beräkningsvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Remote supervision of cabinless autonomous vehicles requires a panoramic view of the vehicle’s surroundings. No single camera covers the full surround, so adjacent views must be fused into one video stream under strict latency. This thesis presents a LiDAR-guided Thin-Plate Spline (TPS) stitching pipeline for the three forward-facing ring cameras (ring_front_left, ring_front_center, ring_front_right) of the Argoverse 2 (AV2) sensor dataset [26]. The pipeline uses LiDAR points as 3-D geometric correspondences. Points visible in both adjacent cameras are projected into each camera’s image plane and onto the shared cylindrical canvas, yielding pixel correspondences that drive the TPS warp. The centre camera (FC) is held at its rotationbaseline canvas position and the two side cameras carry the full TPS displacement budget, so FC’s pixels stay at predictable canvas locations. A dynamic-programming seam search composites the warped images, routing the cut through textureless regions, and per-channel gain compensation reduces illumination mismatches across cameras beforehand. The canvas is sized dynamically from the azimuth and elevation span of the visible LiDAR returns, with a cylindrical projection scaled to the centre camera’s focal length. Per-camera rotation-only remaps are precomputed once, and TPS corrects the residual parallax the pure-rotation model leaves behind. Dynamic canvas sizing removes any fixed field-of-view assumption and adapts automatically to any combination of cameras or log-specific calibration. On Argoverse 2 sensor logs and a laptop-class NVIDIA RTX 3050 Ti the pipeline runs end-to-end at ≈ 28 ms per frame, well below the 50 ms budget of the camera rig’s 20 Hz input cadence. Each log yields over 2000 shared points per frame in both overlaps, an order of magnitude above the deployed M=50 control-point budget. Across ten AV2 logs the rotation baseline leaves adjacent cameras misaligned by 14/24 px on average. The LiDAR-fitted TPS warp drives the cross-camera disagreement at held-out 3-D points to 1.1/1.5 px mean per overlap, per-log values spanning 0.2–2.8 px, while the ORB image-content residual on the same overlaps stays at 5.9/4.7 px median (1.3–17 px across logs). The held-out reprojection is a geometric self-consistency check on the LiDAR-fitted field; the ORB residual measures image-feature alignment and additionally absorbs parallax on near-field structures and the inter-camera shutter stagger. Photometric agreement holds at 24.2/24.2 dB channel-weighted PSNR on the 10-log mean, with log-to-log std ±2.3/±1.5 dB, consistent with the warp extrapolating smoothly beyond the fitted control-point set. Two intrinsic ceilings are analytically bounded: depth parallax on near-field structures inherent to any 2-D warp, and the ≈ 12.5 ms inter-camera shutter stagger without motion compensation. The correctable components, per-channel gain and TPS regularisation, are quantified in the evaluation. The evaluation covers daytime driving on ten AV2 logs spanning four cities (Pittsburgh, Miami, Washington DC, and Detroit) and a wide range of conditions: dry sunny, wet pavement, fog, bridge-underpass HDR, and overcast. Generalisation to night, highway, and other rigs is not demonstrated.

Information

Författare
Näslund, Erik
Lärosäte / institution
Uppsala universitet/Avdelningen för beräkningsvetenskap
Publiceringsdatum
2026
Uppsatstyp
Yrkesexamen på avancerad nivå
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.