← Search

Heng Li

29 accepted papers

2026

Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Model

RSS 2026poster

Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for evaluating Vision-Language-Action (VLA) systems. However, their reliability as evaluation benchmarks is often limited by the representation gap between low-dimensional actions and high-dimensional video synthesis. Thi…

Cited by 0SourceScholar
2026

D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarping

CVPR 2026

Document image dewarping remains a challenging task in the deep learning era. While existing methods have improved by leveraging text line awareness, they typically focus only on a single horizontal dimension. In this paper, we propose a fine-grained deformation perception model that focuses on Dual

Cited by 0SourcecodeScholar
2026

DeLightMono: Enhancing Self-Supervised Monocular Depth Estimation in Endoscopy by Decoupling Uneven Illumination

AAAI 2026technical

Self-supervised monocular depth estimation serves as a key task in the development of endoscopic navigation systems. However, performance degradation persists due to uneven illumination inherent in endoscopic images, particularly in low-intensity regions. Existing low-light enhancement techniques fa

Cited by 0SourcePDFScholar
2026

MMDIR: Multimodal Instruction-Driven Framework for Mixed-Degradation Document Image Restoration

CVPR 2026

Restoring degraded document image is essential for both improving visual quality and optimizing performance in downstream document analysis tasks. Although existing methods have demonstrated substantial improvements in restoration outcomes, they primarily address single-type degradation scenarios. C

Cited by 0SourcecodeScholar
2026

MotionHiFlow: Text-to-Motion via Hierarchical Flow Matching

CVPR 2026

Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce complex and natural movements, they usually operate at only one temporal scale, which

Cited by 2SourcecodeScholar
2026

Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding

ICLR 2026oral

Verification is a key bottleneck in improving inference speed while maintaining distribution fidelity in Speculative Decoding. Recent work has shown that sequence-level verification leads to a higher number of accepted tokens compared to token-wise verification. However, existing solutions often rel…

Cited by 0SourcecodeScholar
2026

VisAssist: A Visually Impaired-Captured Video Question Answering Benchmark for Assistive Systems

AAAI 2026technical

We present VisAssist, the first large-scale video question-answering dataset with 13,413 real-world videos captured by visually impaired users, addressing a critical gap in assistive vision research. Unlike existing benchmarks relying on third-person footage, VisAssist provides authentic first-perso

Cited by 0SourcePDFScholar
2025

AIF-SFDA: Autonomous Information Filter Driven Source-Free Domain Adaptation for Medical Image Segmentation

AAAI 2025technical

Decoupling domain-variant information (DVI) from domain-invariant information (DII) serves as a prominent strategy for mitigating domain shifts in the practical implementation of deep learning algorithms. However, in medical settings, concerns surrounding data collection and privacy often restrict a…

2025

CSMT: Combining Snoring and Metadata-based Text for Sleep Apnea Severity Classification

ICASSP 2025accepted

Sleep apnea is a common sleep disorder that, if untreated, can lead to serious health issues. Snoring is a typical symptom of sleep apnea and can be utilized to develop a noncontact automatic detection method for sleep apnea severity classification (SASC). However, due to patient heterogeneity, the…

Cited by 0SourceScholar
2025

MaxSup: Overcoming Representation Collapse in Label Smoothing

NeurIPS 2025oral

Label Smoothing (LS) is widely adopted to reduce overconfidence in neural network predictions and improve generalization. Despite these benefits, recent studies reveal two critical issues with LS. First, LS induces overconfidence in misclassified samples. Second, it compacts feature representations…

Cited by 0SourcecodeScholar
2025

OG-Gaussian: Occupancy Based Street Gaussians for Autonomous Driving

ICRA 2025

Accurate and realistic 3D scene reconstruction enables the lifelike creation of autonomous driving simulation environments. With advancements in 3D Gaussian Splatting (3DGS), previous studies have applied it to reconstruct complex dynamic driving scenes. These methods typically require expensive LiD

Cited by 5SourceScholar
2025

OccMamba: Semantic Occupancy Prediction with State Space Models

CVPR 2025poster

Training deep learning models for semantic occupancy prediction is challenging due to factors such as a large number of occupancy cells, severe occlusion, limited visual cues, complicated driving scenarios, etc. Recent methods often adopt transformer-based architectures given their strong capability…

2025

PanTS: The Pancreatic Tumor Segmentation Dataset

NeurIPS 2025poster

PanTS is a large-scale, multi-institutional dataset curated to advance research in pancreatic CT analysis. It contains 36,390 CT scans from 145 medical centers, with expert-validated, voxel-wise annotations of over 993,000 anatomical structures, covering pancreatic tumors, pancreas head, body, and t…

Cited by 0SourceScholar
2025

RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection

ACL 2025long

Large language models (LLMs) have demonstrated remarkable capabilities in various domains, including radiology report generation. Previous approaches have attempted to utilize multimodal LLMs for this task, enhancing their performance through the integration of domain-specific knowledge retrieval. H…

2025

STDArm: Transfer Visuomotor Policy From Static Data Training to Dynamic Robot Manipulation

RSS 2025poster

Learning visuomotor policy from human demonstrations serves as an effective method for robots to acquire complex tasks. However, data collection on mobile platforms such as drones is extremely challenging, resulting in most research being conducted with robots in stationary conditions for data colle…

Cited by 0PDFScholar
2025

SeePhys: Does Seeing Help Thinking? – Benchmarking Vision-Based Physics Reasoning

NeurIPS 2025poster

We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In cont…

Cited by 0SourcecodeScholar
2025

Universal Features Guided Zero-Shot Category-Level Object Pose Estimation

AAAI 2025technical

Object pose estimation, crucial in computer vision and robotics applications, faces challenges with the diversity of unseen categories. We propose a zero-shot method to achieve category-level 6-DOF object pose estimation, which exploits both 2D and 3D universal features of input RGB-D image to estab…

Cited by 0SourcePDFScholar
2024

Efficient Backdoor Attacks for Deep Neural Networks in Real-world Scenarios

ICLR 2024poster

Recent deep neural networks (DNNs) have came to rely on vast amounts of training data, providing an opportunity for malicious attackers to exploit and contaminate the data to carry out backdoor attacks. However, existing backdoor attack methods make unrealistic assumptions, assuming that all trainin…

2024

Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions

NeurIPS 2024spotlight

Vision-and-Language Navigation (VLN) aims to develop embodied agents that navigate based on human instructions. However, current VLN frameworks often rely on static environments and optimal expert supervision, limiting their real-world applicability. To address this, we introduce Human-Aware Vision-…

2024

OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous Driving

ICRA 2024poster

Visual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep learning to transform 2D camera images into 3D semantic occupa…

Cited by 8SourcecodeScholar
2024

TDeLTA: A Light-Weight and Robust Table Detection Method Based on Learning Text Arrangement

AAAI 2024technical

The diversity of tables makes table detection a great challenge, leading to existing models becoming more tedious and complex. Despite achieving high performance, they often overfit to the table style in training set, and suffer from significant performance degradation when encountering out-of-distr…

2023

DENSE RGB SLAM WITH NEURAL IMPLICIT MAPS

ICLR 2023poster

There is an emerging trend of using neural implicit functions for map representation in Simultaneous Localization and Mapping (SLAM). Some pioneer works have achieved encouraging results on RGB-D SLAM. In this paper, we present a dense RGB SLAM method with neural implicit map representation. To reac…

2023

Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech Enhancement

ICASSP 2023accepted

In denoising and de-reverberation tasks, the dominant methods are complex spectral masking and complex spectral mapping. To combine advantages and improve speech enhancement performance, we propose a two-stage UNet (TSUNet) to estimate complex spectral masking and complex spectral mapping. We use a…

Cited by 0SourceScholar
2022

FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network

ICASSP 2022accepted

In recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of full-band (48 kHz) speech signals is still very challenging. Becaus…

Cited by 0SourceScholar
2022

Recent Advances in Concept Drift Adaptation Methods for Deep Learning

IJCAI 2022poster

In the ``Big Data'' age, the amount and distribution of data have increased wildly and changed over time in various time-series-based tasks, e.g weather prediction, network intrusion detection. However, deep learning models may become outdated facing variable input data distribution, which is called…

2021

End-to-End Rotation Averaging With Multi-Source Propagation

CVPR 2021poster

This paper presents an end-to-end neural network for multiple rotation averaging in SfM. Due to the manifold constraint of rotations, conventional methods usually take two separate steps involving spanning tree based initialization and iterative nonlinear optimization respectively. These methods can…

Cited by 29PDFcodeScholar