← Search

Heming Du

19 accepted papers

2026

Déjà Vu: Unlocking Transparent Action Reasoning for Object-Goal Navigation Via Large Language Models

ICRA 2026poster

The remarkable interaction and reasoning capabilities of Large Language Models (LLMs) make them promising in collaborative Embodied AI tasks, particularly for Object-goal Navigation (ObjNav) tasks that require both decision-making and transparent explanation. However, existing work mainly uses LLMs …

Cited by 0Scholar
2026

FreDN: Spectral Disentanglement for Time Series Forecasting via Learnable Frequency Decomposition

AAAI 2026technical

Time series forecasting is essential in a wide range of real world applications. Recently, frequency-domain methods have attracted increasing interest for their ability to capture global dependencies. However, when applied to non-stationary time series, these methods encounter the spectral entanglem

Cited by 0SourcePDFScholar
2026

InclusiveVidPose: Bridging the Pose Estimation Gap for Individuals with Limb Deficiencies in Video-Based Motion

ICLR 2026poster

Approximately 445.2 million individuals worldwide are living with traumatic amputations, and an estimated 31.64 million children aged 0–14 have congenital limb differences, yet they remain largely underrepresented in human pose estimation (HPE) research. Accurate HPE could significantly benefit this…

Cited by 0SourcecodeScholar
2026

Position: Human-Centric Vision Requires Topological Generalization Beyond Fixed Skeletal Topologies

ICML 2026poster

In this position paper, we argue that human-centric vision requires skeletal-topology generalization beyond fixed skeletons. Mainstream pose and body pipelines enforce a fixed skeleton graph with an indexed joint list and fixed adjacency, so the fixed joint inventory does not cover structural absenc…

Cited by 0SourceScholar
2026

ResiHMR: Residual-Limb Aware Single-Image 3D Human Mesh Recovery for Individuals with Limb Loss

CVPR 2026

Single-image human mesh recovery provides a compact 3D, person-centric representation that supports analysis, animation, AR and VR, rehabilitation, and human-computer interaction. However, prevailing systems impose an intact-limb prior and degrade on people with limb loss, because fixed-topology mod

Cited by 0SourceScholar
2026

Robust Test-time Video-Text Retrieval: Benchmarking and Adapting for Query Shifts

ICLR 2026poster

Modern video-text retrieval (VTR) models excel on in-distribution benchmarks but are highly vulnerable to real-world *query shifts*, where the distribution of query data deviates from the training domain, leading to a sharp performance drop. Existing image-focused robustness solutions are inadequate…

Cited by 0SourceScholar
2025

LDPose: Towards Inclusive Human Pose Estimation for Limb-Deficient Individuals in the Wild

ICCV 2025poster

Human pose estimation aims to predict the location of body keypoints and enable various practical applications. However, existing research focuses solely on individuals with full physical bodies and overlooks those with limb deficiencies. As a result, current pose estimation methods cannot be genera…

Cited by 0SourcePDFScholar
2025

M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world Settings

CVPR 2025poster

Human pose estimation is a critical task in computer vision for applications in sports analysis, healthcare monitoring, and human-computer interaction. However, existing human pose datasets are collected either from custom-configured laboratories with complex devices or they only include data on sin…

Cited by 0SourcePDFScholar
2025

Multimodal Retina Image Analysis Survey: Datasets, Tasks and Methods

IJCAI 2025

Retina images provide a noninvasive view of the central nervous system and microvasculature, making it essential for clinical applications. Changes in the retina often indicate both ophthalmic and systemic diseases, aiding in diagnosis and early intervention. While deep learning algorithms have adva

Cited by 0SourcePDFScholar
2025

Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval via Uncertainty Minimization

ICCV 2025poster

Despite recent advances, Text-to-video retrieval (TVR) is still hindered by multiple inherent uncertainties, such as ambiguous textual queries, indistinct text-video mappings, and low-quality video frames. Although interactive systems have emerged to address these challenges by refining user intent…

2025

When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions

NeurIPS 2025poster

Existing Moment retrieval (MR) methods focus on Single-Moment Retrieval (SMR). However, one query can correspond to multiple relevant moments in real-world applications. This makes the existing datasets and methods insufficient for video temporal grounding. By revisiting the gap between current MR…

Cited by 0SourcecodeScholar
2024

MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset

NeurIPS 2024poster

Isolated Sign Language Recognition (ISLR) focuses on identifying individual sign language glosses. Considering the diversity of sign languages across geographical regions, developing region-specific ISLR datasets is crucial for supporting communication and research. Auslan, as a sign language specif…

Cited by 0SourcePDFScholar
2023

Auslan-Daily: Australian Sign Language Translation for Daily Communication and News

NeurIPS 2023poster

Sign language translation (SLT) aims to convert a continuous sign language video clip into a spoken language. Considering different geographic regions generally have their own native sign languages, it is valuable to establish corresponding SLT datasets to support related communication and research.…

Cited by 18SourcePDFScholar
2023

Object-Goal Visual Navigation via Effective Exploration of Relations Among Historical Navigation States

CVPR 2023poster

Object-goal visual navigation aims at steering an agent toward an object via a series of moving steps. Previous works mainly focus on learning informative visual representations for navigation, but overlook the impacts of navigation states on the effectiveness and efficiency of navigation. We observ…

Cited by 27SourcePDFScholar
2023

RVD: A Handheld Device-Based Fundus Video Dataset for Retinal Vessel Segmentation

NeurIPS 2023poster

Retinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina fluctuation, resulting in diminished dataset richness, and the usage of bench-top devices further restricts dataset scal…

Cited by 10SourcePDFScholar
2023

SEFormer: Structure Embedding Transformer for 3D Object Detection

AAAI 2023technical

Effectively preserving and encoding structure features from objects in irregular and sparse LiDAR points is a crucial challenge to 3D object detection on the point cloud. Recently, Transformer has demonstrated promising performance on many 2D and even 3D vision tasks. Compared with the fixed and ri…

2022

Monocular Camera-Based Point-Goal Navigation by Learning Depth Channel and Cross-Modality Pyramid Fusion

AAAI 2022technical

For a monocular camera-based navigation system, if we could effectively explore scene geometric cues from RGB images, the geometry information will significantly facilitate the efficiency of the navigation system. Motivated by this, we propose a highly efficient point-goal navigation framework, dubb…

Cited by 10SourcePDFScholar
2020

Learning Object Relation Graph and Tentative Policy for Visual Navigation

ECCV 2020poster

Target-driven visual navigation aims at navigating an agent towards a given target based on the observation of the agent. In this task, it is critical to learn informative visual representation and robust navigation policy. Aiming to improve these two components, this paper proposes three complement…