← Search

Rong Li

14 accepted papers

2026

EventDrive: Event Cameras for Vision-Language Driving Intelligence

CVPR 2026

Event cameras sense the world through asynchronous brightness changes with microsecond latency and high dynamic range, offering motion fidelity far beyond frame-based sensors and capturing temporal structure that conventional exposures often miss. These properties make events a powerful complement t

Cited by 0SourceScholar
2026

Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration

RA-L 2026

Deployable service and delivery robots struggle to navigate multi-floor buildings to reach object goals, as existing systems fail due to single-floor assumptions and requirements for offline, globally consistent maps. Multi-floor environments pose unique challenges including cross-floor transitions

Cited by 3SourcecodeScholar
2025

ConSense: Continually Sensing Human Activity with WiFi via Growing and Picking

AAAI 2025technical

WiFi-based human activity recognition (HAR) holds significant application potential across various fields. To handle dynamic environments where new activities are continuously introduced, WiFi-based HAR systems must adapt by learning new concepts without forgetting previously learned ones. Furthermo…

2025

Global-Aware Monocular Semantic Scene Completion with State Space Models

ICCV 2025poster

Monocular Semantic Scene Completion (MonoSSC) reconstructs and interprets 3D environments from a single image, enabling diverse real-world applications. However, existing methods are often constrained by the local receptive field of Convolutional Neural Networks (CNNs), making it challenging to hand…

Cited by 0SourcePDFScholar
2025

Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change Captioning

AAAI 2025technical

Change captioning aims to describe the differences between two similar images using natural language, significantly aiding in understanding and monitoring changes. This challenging task requires a fine-grained understanding of subtle changes while resisting disturbances like viewpoint shifts and ill…

Cited by 0SourcePDFScholar
2025

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

CVPR 2025poster

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and predefined object categories, limiting scalability and adaptability. To overcome…

2025

Structure Balance and Gradient Matching-Based Signed Graph Condensation

AAAI 2025technical

Training graph neural networks (GNNs) for graph representation has received increasing concerns due to its outstanding performance in the link prediction and node classification tasks, but it incurs much time and storage for tackling large-scale graphs. To alleviate this issue, graph condensation ha…

2025

Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras

NeurIPS 2025spotlight

Event cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challenge. We introduce Talk2Event, the first large-scale benchmark for language-driven…

Cited by 0SourceScholar
2024

An Examination of the Compositionality of Large Generative Vision-Language Models

NAACL 2024long

With the success of Large Language Models (LLMs), many Generative Vision-Language Models (GVLMs) have been constructed via multimodal instruction tuning. However, the performance of GVLMs in multimodal compositional reasoning remains under-explored. In this paper, we examine both the evaluation metr…

2023

Towards Content-based Pixel Retrieval in Revisited Oxford and Paris

ICCV 2023poster

This paper introduces the first two landmark pixel retrieval benchmarks. Like semantic segmentation extends classification to the pixel level, pixel retrieval is an extension of image retrieval and offers information about which pixels are related to the query object. In addition to retrieving image…

Cited by 4PDFcodeScholar
2022

TalkingFlow: Talking Facial Landmark Generation with Multi-Scale Normalizing Flow Network

ICASSP 2022accepted

Deterministic models dominate the field of talking facial land-mark generation by directly mapping speech signals to a certain lip-sync facial landmark sequence, which often suffer from regression to the mean face. In contrast, probability generative models are more beneficial to handle complex data…

Cited by 0SourceScholar
2021

Perception-Aware Multi-Sensor Fusion for 3D LiDAR Semantic Segmentation

ICCV 2021poster

3D LiDAR (light detection and ranging) semantic segmentation is important in scene understanding for many applications, such as auto-driving and robotics. For example, for autonomous cars equipped with RGB cameras and LiDAR, it is crucial to fuse complementary information from different sensors for…

Cited by 223PDFcodeScholar