← Search

Shaohua Dong

6 accepted papers

2026

DMTrack: Spatio-Temporal Multimodal Tracking Via Dual-Adapter

ICRA 2026poster

In this paper, we explore adapter tuning and introduce a novel dual-adapter architecture for spatio-temporal multimodal tracking, dubbed DMTrack. The key of our DMTrack lies in two simple yet effective modules, including a spatio-temporal modality adapter (STMA) and a progressive modality complement…

2025

Efficient and Accurate Low-Resolution Transformer Tracking

IROS 2025

High-performance Transformer trackers have exhibited excellent results, yet they often bear a heavy computational load. Observing that a smaller input can immediately and conveniently reduce computations without changing the model, an easy solution is to adopt a low-resolution input for efficient Tr

Cited by 0SourcecodeScholar
2024

Beyond MOT: Semantic Multi-Object Tracking

ECCV 2024poster

"Current multi-object tracking (MOT) aims to predict trajectories of targets (, “where”) in videos. Yet, knowing merely “where” is insufficient in many crucial applications. In comparison, semantic understanding such as fine-grained behaviors, interactions, and overall summarized captions (, “what”)…

2024

Efficient Multimodal Semantic Segmentation via Dual-Prompt Learning

IROS 2024

Multimodal (e.g., RGB-Depth/RGB-Thermal) fusion has shown great potential for improving semantic segmentation in complex scenes (e.g., indoor/low-light conditions). Existing approaches often fully fine-tune a dual-branch encoder-decoder framework with a complicated feature fusion strategy for achiev

Cited by 45SourcecodeScholar
2024

VastTrack: Vast Category Visual Object Tracking

NeurIPS 2024poster

In this paper, we propose a novel benchmark, named VastTrack, aiming to facilitate the development of general visual tracking via encompassing abundant classes and videos. VastTrack consists of a few attractive properties: (1) Vast Object Category. In particular, it covers targets from 2,115 categor…

2022

Edge-Aware Guidance Fusion Network for RGB–Thermal Scene Parsing

AAAI 2022technical

RGB–thermal scene parsing has recently attracted increasing research interest in the field of computer vision. However, most existing methods fail to perform good boundary extraction for prediction maps and cannot fully use high-level features. In addition, these methods simply fuse the features fro…