← Search

Huansheng Song

4 accepted papers

2025

Beyond Human Perception: Understanding Multi-Object World from Monocular View

CVPR 2025poster

Language and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often availab…

2025

MTTM: Memory-Augmented with Mamba for 3D Medical Images Analysis

ICASSP 2025accepted

The rapid advancement of artificial intelligence has propelled the healthcare industry into a new era of diagnostic precision. A pivotal component of this evolution is the accurate classification of 3D medical images, which necessitates extracting robust feature representations capable of effectivel…

Cited by 0SourceScholar
2025

Mono3DVLT: Monocular-Video-Based 3D Visual Language Tracking

CVPR 2025poster

Visual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal wi…

2020

Simultaneous Detection and Tracking with Motion Modelling for Multiple Object Tracking

ECCV 2020poster

Deep learning based Multiple Object Tracking (MOT) currently relies on off-the-shelf detectors for tracking-by-detection. This results in deep models that are detector biased and evaluations that are detector influenced. To resolve this issue, we introduce Deep Motion Modeling Network (DMM-Net) that…

Cited by 72SourcePDFScholar