← Search

Shaoxiang Chen

10 accepted papers

2024

ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model

NeurIPS 2024poster

Visual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language~(VL) trackers have proposed to utilize additional natural language descriptions to enhance versatility in various applications. However, VL trackers are still infer…

Cited by 6SourcePDFScholar
2024

Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting Learning

AAAI 2024technical

Camera-based bird-eye-view (BEV) perception paradigm has made significant progress in the autonomous driving field. Under such a paradigm, accurate BEV representation construction relies on reliable depth estimation for multi-camera images. However, existing approaches exhaustively predict depths fo…

2024

Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models

NeurIPS 2024poster

Large Multimodal Model (LMM) is a hot research topic in the computer vision area and has also demonstrated remarkable potential across multiple disciplinary fields. A recent trend is to further extend and enhance the perception capabilities of LMMs. The current methods follow the paradigm of adaptin…

2024

Making Large Language Models Better Planners with Reasoning-Decision Alignment

ECCV 2024oral

"Data-driven approaches for autonomous driving (AD) have been widely adopted in the past decade but are confronted with dataset bias and uninterpretability. Inspired by the knowledge-driven nature of human driving, recent approaches explore the potential of large language models (LLMs) to improve un…

Cited by 12SourcePDFScholar
2023

MSMDFusion: Fusing LiDAR and Camera at Multiple Scales With Multi-Depth Seeds for 3D Object Detection

CVPR 2023poster

Fusing LiDAR and camera information is essential for accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features from two drastically different modalities. Recent approaches aim at e…

2022

MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes

ECCV 2022poster

"3D dense captioning is a recently-proposed novel task, where point clouds contain more geometric information than the 2D counterpart. However, it is also more challenging due to the higher complexity and wider variety of inter-object relations contained in point clouds. Existing methods only treat…

2021

Towards Bridging Event Captioner and Sentence Localizer for Weakly Supervised Dense Event Captioning

CVPR 2021poster

Dense Event Captioning (DEC) aims to jointly localize and describe multiple events of interest in untrimmed videos, which is an advancement of the conventional video captioning task (generating a single sentence description for a trimmed video). Weakly Supervised Dense Event Captioning (WS-DEC) goes…

Cited by 85PDFcodeScholar
2020

Hierarchical Visual-Textual Graph for Temporal Activity Localization via Language

ECCV 2020poster

Temporal Activity Localization via Language (TALL) in video is a recently proposed challenging vision task, and tackling it requires fine-grained understanding of the video content, however, this is overlooked by most of the existing works. In this paper, we propose a novel TALL method which builds…

2020

Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in Videos

ECCV 2020poster

Automatically generating sentences to describe events and temporally localizing sentences in a video are two important tasks that bridge language and videos. Recent techniques leverage the multimodal nature of videos by using off-the-shelf features to represent videos, but interactions between modal…

Cited by 123SourcePDFScholar