ICRA 2026poster0 citations

Multi-Task Visual Perception with Temporal Feature Fusion for Autonomous Driving

Huei-Yung Lin, Shih-Han Wei

Abstract

With the rapid developments of autonomous driving technologies, accurate scene perception has become essential for safe and efficient navigation. The key perception tasks such as lane detection, semantic segmentation of road markings and road area, and object detection directly impact vehicle decision-making and obstacle avoidance. However, most existing methods are trained on a single-task dataset, limiting data diversity and reducing performance in complex scenarios or under occlusion and illumination variation. In this work we propose a multi-task perception network based on image sequence input, integrating lane detection, road marking and road area segmentation, and object detection into a unified framework. The network model employs multi-task learning to share features and improve the computational efficiency, and adopts the cross-dataset training paradigm to enhance generalization across tasks. Furthermore, the temporal information from adjacent frames is leveraged to compensate visual degradation in current frames. Experimental results obtained on multiple datasets demonstrate the proposed technique achieves competitive performance compared to state-of-the-art approaches. Code is available at https://github.com/(removed_for_review).

Autonomous Vehicle NavigationIntelligent Transportation SystemsComputer Vision for Transportation
Multi-Task Visual Perception with Temporal Feature Fusion for Autonomous Driving · ICRA 2026