← Search

LI He

29 accepted papers

2026

DeepWriter: A Multi-Agent Collaboration Framework for Information-rich Ultra-long Book Writing

AAAI 2026technical

Long-form books are among the most information-rich and structurally complex forms of written content, often exceeding 100,000 words. While recent methods have enabled basic long-text generation, they remain limited in two key aspects: the inability to generate ultra-long content at book scale, and

Cited by 0SourcePDFScholar
2026

GO-PRE:Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction

ICML 2026poster

Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing methods rely on surrogate signals—such as parameter uncertainty or geometric heuristics—which are often misaligned with the ultimate goal: the fidelity o…

Cited by 0SourceScholar
2026

Hierarchical Visual Relocalization with Nearest View Synthesis from Feature Gaussian Splatting

CVPR 2026

Visual relocalization is a fundamental task in the field of 3D computer vision, estimating a camera's pose when it revisits a previously known scene. While point-based hierarchical relocalization methods have shown strong scalability and efficiency, they are often limited by sparse image observation

Cited by 0SourceScholar
2026

Unifying Stable Optimization and Reference Regularization in RLHF

ICLR 2026poster

Reinforcement Learning from Human Feedback (RLHF) has advanced alignment capabilities significantly but remains hindered by two core challenges: reward hacking and stable optimization. Current solutions independently address these issues through separate regularization strategies, specifically a KL-…

Cited by 0SourcecodeScholar
2025

PF-TEB: Timed Elastic Band-Based Human-Aware Robot Navigation Framework in Crowded Environments

RA-L 2025

To enhance the social navigation performance of mobile robots in crowded environments, we propose a novel framework—Prediction and Fuzzy Timed Elastic Band (PF-TEB) for robot social navigation. Our framework incorporates predicted pedestrian trajectories into pedestrian proxemics modeling as a socia

Cited by 2SourceScholar
2024

Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signals

NeurIPS 2024poster

Invasive brain-computer interfaces with Electrocorticography (ECoG) have shown promise for high-performance speech decoding in medical applications, but less damaging methods like intracranial stereo-electroencephalography (sEEG) remain underexplored. With rapid advances in representation learning,…

2024

GigaTraj: Predicting Long-term Trajectories of Hundreds of Pedestrians in Gigapixel Complex Scenes

CVPR 2024poster

Pedestrian trajectory prediction is a well-established task with significant recent advancements. However existing datasets are unable to fulfill the demand for studying minute-level long-term trajectory prediction mainly due to the lack of high-resolution trajectory observation in the wide field of…

Cited by 3SourcePDFScholar
2024

LiteTrack: Layer Pruning with Asynchronous Feature Extraction for Lightweight and Efficient Visual Tracking

ICRA 2024poster

The recent advancements in transformer-based visual trackers have led to significant progress, attributed to their strong modeling capabilities. However, as performance improves, running latency correspondingly increases, presenting a challenge for real-time robotics applications, especially on edge…

Cited by 20SourcecodeScholar
2024

PISR: Polarimetric Neural Implicit Surface Reconstruction for Textureless and Specular Objects

ECCV 2024poster

"Neural implicit surface reconstruction has achieved remarkable progress recently. Despite resorting to complex radiance modeling, state-of-the-art methods still struggle with textureless and specular surfaces. Different from RGB images, polarization images can provide direct constraints on the azim…

2024

Person Re-Identification for Robot Person Following With Online Continual Learning

RA-L 2024

Robot person following (RPF) is a crucial capability in human-robot interaction (HRI) applications, allowing a robot to persistently follow a designated person. In practical RPF scenarios, the person can often be occluded by other objects or people. Consequently, it is necessary to re-identify the p

Cited by 18SourceScholar
2024

STAGP: Spatio-Temporal Adaptive Graph Pooling Network for Pedestrian Trajectory Prediction

RA-L 2024

Predicting how pedestrians will move in the future is crucial for robot navigation, autonomous driving, and video surveillance. The complex interactions among pedestrians make it difficult to predict their future trajectory. Previous studies have primarily focused on modeling the interaction feature

Cited by 23SourceScholar
2024

SWCF-Net: Similarity-weighted Convolution and Local-global Fusion for Efficient Large-scale Point Cloud Semantic Segmentation

IROS 2024poster

Large-scale point cloud consists of a multitude of individual objects, thereby encompassing rich structural and underlying semantic contextual information, resulting in a challenging problem in efficiently segmenting a point cloud. Most existing researches mainly focus on capturing intricate local f…

Cited by 2SourcecodeScholar
2023

AMR-TST: Abstract Meaning Representation-based Text Style Transfer

ACL 2023findings

Abstract Meaning Representation (AMR) is a semantic representation that can enhance natural language generation (NLG) by providing a logical semantic input. In this paper, we propose the AMR-TST, an AMR-based text style transfer (TST) technique. The AMR-TST converts the source text to an AMR graph a…

2023

Beyond Reward: Offline Preference-guided Policy Optimization

ICML 2023poster

This study focuses on the topic of offline preference-based reinforcement learning (PbRL), a variant of conventional reinforcement learning that dispenses with the need for online interaction or specification of reward functions. Instead, the agent is provided with fixed offline trajectories and hum…

2023

CEIL: Generalized Contextual Imitation Learning

NeurIPS 2023poster

In this paper, we present ContExtual Imitation Learning (CEIL), a general and broadly applicable algorithm for imitation learning (IL). Inspired by the formulation of hindsight information matching, we derive CEIL by explicitly learning a hindsight embedding function together with a contextual polic…

Cited by 23SourcePDFScholar
2023

Combating Bilateral Edge Noise for Robust Link Prediction

NeurIPS 2023poster

Although link prediction on graphs has achieved great success with the development of graph neural networks (GNNs), the potential robustness under the edge noise is still less investigated. To close this gap, we first conduct an empirical study to disclose that the edge noise bilaterally perturbs bo…

2023

Combining Scene Coordinate Regression and Absolute Pose Regression for Visual Relocalization

ICRA 2023poster

Visual relocalization is a fundamental problem in computer vision and robotics. Recently, regression-based methods become popular and they can be categorized into two classes: absolute pose regression and scene coordinate regression. In this work, we present a combined regression network that jointl…

Cited by 4SourceScholar
2023

Exploring Model Dynamics for Accumulative Poisoning Discovery

ICML 2023poster

Adversarial poisoning attacks pose huge threats to various machine learning applications. Especially, the recent accumulative poisoning attacks show that it is possible to achieve irreparable harm on models via a sequence of imperceptible attacks followed by a trigger batch. Due to the limited data-…

2023

Robot Person Following Under Partial Occlusion

ICRA 2023poster

Robot person following (RPF) is a capability that supports many useful human-robot-interaction (HRI) applications. However, existing solutions to person following often as-sume full observation of the tracked person. As a consequence, they cannot track the person reliably under partial occlusion whe…

Cited by 20SourcecodeScholar
2022

NDD: A 3D Point Cloud Descriptor Based on Normal Distribution for Loop Closure Detection

IROS 2022poster

Loop closure detection is a key technology for long-term robot navigation in complex environments. In this paper, we present a global descriptor, named Normal Distribution Descriptor (NDD), for 3D point cloud loop closure detection. The descriptor encodes both the probability density score and entro…

Cited by 14SourcecodeScholar
2022

Perspective Phase Angle Model for Polarimetric 3D Reconstruction

ECCV 2022poster

"Current polarimetric 3D reconstruction methods, including those in the well-established shape from polarization literature, are all developed under the orthographic projection assumption. In the case of a large field of view, however, this assumption does not hold and may result in significant reco…

2021

Robust Improvement in 3D Object Landmark Inference for Semantic Mapping

ICRA 2021poster

Recent works on semantic Simultaneous Localization and Mapping (SLAM) utilizing object landmarks have shown superiority in terms of robustness and accuracy in tracking and localization. 3D object landmarks represented by a cubic or quadric surface are inferred from 2D object bounding boxes which are…

Cited by 4SourceScholar
2020

Keypoint Description by Descriptor Fusion Using Autoencoders

ICRA 2020poster

Keypoint matching is an important operation in computer vision and its applications such as visual simultaneous localization and mapping (SLAM) in robotics. This matching operation heavily depends on the descriptors of the keypoints, and it must be performed reliably when images undergo conditional…

Cited by 5SourceScholar
2019

A Comparison of CNN-Based and Hand-Crafted Keypoint Descriptors

ICRA 2019poster

Keypoint matching is an important operation in computer vision and its applications such as visual simultaneous localization and mapping (SLAM) in robotics. This matching operation heavily depends on the descriptors of the keypoints, and it must be performed reliably when images undergo condition ch…

Cited by 34SourceScholar
2019

Improving Keypoint Matching Using a Landmark-Based Image Representation

ICRA 2019poster

Motivated by the need to improve the performance of visual loop closure verification via multi-view geometry (MVG) under significant illumination and viewpoint changes, we propose a keypoint matching method that uses landmarks as an intermediate image representation in order to leverage the power of…

Cited by 7SourceScholar