← Search

Zhi Li

70 accepted papers

2026

A Hitchhiker's Guide to Poisson Gradient Estimation

ICML 2026poster

Poisson-distributed latent variable models are widely used in computational neuroscience, but differentiating through discrete stochastic samples remains challenging. Two approaches address this: *Exponential Arrival Time* (EAT) simulation and *Gumbel-SoftMax* (GSM) relaxation. We provide the first …

Cited by 0SourceScholar
2026

Active Scene Reconstruction with Topological Reasoning and Semantic-Augmented Reinforcement Learning

ICRA 2026poster

Active scene reconstruction aims to autonomously recover the fine-grained appearance and structural details of a complex unknown scenes. Existing approaches based on 2D topological or voxel-based abstractions often scale poorly to large environments and rely heavily on handcrafted features and heuri…

Cited by 0Scholar
2026

E³SAM2: Entropy-Aware and Edge-Guided Adaptation of SAM2 for Echocardiography Video Segmentation

AAAI 2026technical

Foundation segmentation models, such as SAM and its video-oriented variant SAM2, have achieved remarkable success in natural image and video segmentation. However, their direct application to echocardiography video is challenged by structural uncertainty arising from severe speckle noise and blurry

Cited by 0SourcePDFScholar
2026

Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control

ICML 2026poster

Diffusion models are effective for waypoint prediction in visual navigation, but standard sampling and test time guidance can produce unsafe or inefficient trajectories when updates drift off the training manifold. We propose Fisher Preserving Guidance with Outer Product Span Projection, a training-…

Cited by 0SourceScholar
2026

From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection

CVPR 2026

Zero-shot anomaly detection (ZSAD) aims to identify unseen anomalies without abnormal supervision, which is essential for open-world scenarios. Recent vision-language models such as CLIP enable anomaly reasoning through shared visual-textual embeddings, but existing methods often rely on coarse prom

Cited by 0SourceScholar
2026

From Observations to Events: Event-Aware World Models for Reinforcement Learning

ICLR 2026poster

While model-based reinforcement learning (MBRL) improves sample efficiency by learning world models from raw observations, existing methods struggle to generalize across structurally similar scenes and remain vulnerable to spurious variations such as textures or color shifts. From a cognitive scienc…

Cited by 0SourcecodeScholar
2026

GPTS-Nav: End-to-End Robot Navigation in Dynamic Environments With Graph-Privileged Teacher-Student Reinforcement Learning

RA-L 2026

Reinforcement learning (RL) has renewed interest in end-to-end robot navigation, yet dynamic, crowd-like scenes remain difficult due to partial observability and the fragility of online tracking. We introduce GPTS-Nav, a Graph-Privileged Teacher-Student RL framework for LiDAR-only navigation. A grap

Cited by 0SourceScholar
2026

PROB-EMOE: A Probabilistic Ensemble Mixture-of-Experts Framework for Metro Network Expansion Forecasting

IJCAI 2026

Forecasting Origin-Destination (OD) demand for new metro lines is critical for sustainable infrastructure planning but faces spatiotemporal out-of-distribution challenges. Existing models often struggle to capture heterogeneous interaction patterns in changing topologies and overlook inherent uncert

Cited by 0Scholar
2026

RCPU: Rotation-Constrained Error Compensation for Structured Pruning of a Large Language Model

ICLR 2026poster

In this paper, we propose a rotation-constrained compensation method to address the errors introduced by structured pruning of large language models (LLMs). LLMs are trained on massive datasets and accumulate rich semantic knowledge in their representation space. In contrast, pruning is typically…

Cited by 0SourcecodeScholar
2026

RPE-PAD: Relative Pose Estimation for Pose-agnostic Anomaly Detection

AAAI 2026technical

Pose-agnostic Anomaly Detection (PAD) aims to detect anomalies when the poses of query images are unknown and differ from those in the training set. Therefore, accurately estimating the camera poses for the query images in the test set is critical for this task. Existing query-specific framework met

Cited by 0SourcePDFScholar
2026

WebWorld: A Large-Scale World Model for Web Agent Training

ICML 2026poster

Web agents require massive trajectories to generalize, yet real-world training is constrained by network latency, rate limits, and safety risks. We introduce \textbf{WebWorld} series, the first open-web simulator trained at scale. While existing simulators are restricted to closed environments with …

Cited by 0SourceScholar
2025

ArenaSim: A High-Performance Simulation Platform for Multi-Robot Self-Play Learning

RA-L 2025

In this letter, we introduce <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">ArenaSim</i>, a novel simulation platform designed for realistic and efficient self-play learning in multi-robot cooperative-competitive games. Compared to previous simulati

Cited by 0SourceScholar
2025

Cyclic Vision-Language Manipulator: Towards Reliable and Fine-Grained Image Interpretation for Automated Report Generation

IJCAI 2025

Despite significant advancements in automated report generation, the opaqueness of text interpretability continues to cast doubt on the reliability of the content produced. This paper introduces a novel approach to identify specific image features in X-ray images that influence the outputs of report

Cited by 0SourcePDFScholar
2025

Development of a Stick-Slip Dielectric Elastomer Actuator for Robotic Applications

RA-L 2025

Dielectric elastomer actuators (DEAs) face a performance tradeoff between achieving large displacements and high driving speeds, which limits their use in precision actuation scenarios requiring both rapid response and a wide motion range. To address these limitations, this study introduces a novel

Cited by 1SourceScholar
2025

FFBGNet: Full-Flow Bidirectional Feature Fusion Grasp Detection Network Based on Hybrid Architecture

RA-L 2025

Effectively integrating the complementary information from RGB-D images presents a significant challenge in robotic grasping. In this letter, we propose a full-flow bidirectional feature fusion grasp detection network (FFBGNet) based on a hybrid architecture to generate accurate grasp poses from RGB

Cited by 3SourceScholar
2025

GlobalTomo: A global dataset for physics-ML seismic wavefield modeling and FWI

NeurIPS 2025poster

Global seismic tomography, taking advantage of seismic waves from natural earthquakes, provides essential insights into the earth's internal dynamics. Advanced Full-Waveform Inversion (FWI) techniques, whose aim is to meticulously interpret every detail in seismograms, confront formidable computatio…

Cited by 0SourcecodeScholar
2025

Human-Robot Collaboration for the Remote Control of Mobile Humanoid Robots With Torso-Arm Coordination

ICRA 2025

Recently, many humanoid robots have been increasingly deployed in various facilities, including hospitals and assisted living environments, where they are often remotely controlled by human operators. Their kinematic redundancy enhances reachability and manipulability, enabling them to navigate comp

Cited by 2SourceScholar
2025

LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English Writing

ACL 2025long

The paper explores the performance of LLMs in the context of multi-dimensional analytic writing assessments, i.e. their ability to provide both scores and comments based on multiple assessment criteria. Using a corpus of literature reviews written by L2 graduate students and assessed by human expert…

2025

Learning to Explore Efficiently: Heterogeneous Topological Graphs and Lightweight Global Reasoning for Robotic Exploration

RA-L 2025

Autonomous exploration in large-scale, unknown environments remains a significant challenge in mobile robotics. In this paper, we propose a scalable exploration framework that integrates heterogeneous topological representations, lightweight global-local graph reasoning, and reinforcement learning.

Cited by 1SourceScholar
2025

PNetGPT: Proprietary Protocol Network Traffic Generation with Pre-trained Transformer

ICASSP 2025accepted

Generative pre-trained transformers are exceedingly effective as generative models and classifiers, widely used in natural language processing and computer vision. This work contributes to the exploration of generative pre-trained transformer-based models in the proprietary protocol network traffic.…

Cited by 0SourceScholar
2024

EDM: Synthetic Data from Exemplar Diffusion Model Improves Non-Communicable Diseases Detection

ICASSP 2024accepted

There have been researches revealing obvious associations between facial phenotypes and non-communicable diseases (NCDs), which enables effective health assessment with the integration of model-based learning methods. However, the paucity and poor quality of available datasets hinder the development…

Cited by 0SourceScholar
2024

Expansion-GRR: Efficient Generation of Smooth Global Redundancy Resolution Roadmaps

IROS 2024poster

Global redundancy resolution (GRR) roadmap is a novel concept in robotics that facilitates the mapping from task space paths to configuration space paths in a legible, predictable, and repeatable way. Such roadmaps could find widespread utility in applications such as safe teleoperation, consistent…

Cited by 0SourceScholar
2024

I-AM-G: Interest Augmented Multimodal Generator for Item Personalization

EMNLP 2024main

The emergence of personalized generation has made it possible to create texts or images that meet the unique needs of users. Recent advances mainly focus on style or scene transfer based on given keywords. However, in e-commerce and recommender systems, it is almost an untouched area to explore user…

2024

Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency

COLING 2024main

Despite large language models (LLMs) have demonstrated impressive performance in various tasks, they are still suffering from the factual inconsistency problem called hallucinations. For instance, LLMs occasionally generate content that diverges from source article, and prefer to extract information…

2024

Multi-agent Collaborative Perception via Motion-aware Robust Communication Network

CVPR 2024poster

Collaborative perception allows for information sharing between multiple agents such as vehicles and infrastructure to obtain a comprehensive view of the environment through communication and fusion. Current research on multi-agent collaborative perception systems often assumes ideal communication a…

Cited by 4SourcePDFScholar
2024

Optimizing Instruction Synthesis: Effective Exploration of Evolutionary Space with Tree Search

EMNLP 2024finding

Instruction tuning is a crucial technique for aligning language models with humans’ actual goals in the real world. Extensive research has highlighted the quality of instruction data is essential for the success of this alignment. However, creating high-quality data manually is labor-intensive and t…

2024

ReDiffuser: Reliable Decision-Making Using a Diffuser with Confidence Estimation

ICML 2024poster

The diffusion model has demonstrated impressive performance in offline reinforcement learning. However, non-deterministic sampling in diffusion models can lead to unstable performance. Furthermore, the lack of confidence measurements makes it difficult to evaluate the reliability and trustworthiness…

Cited by 15SourcePDFScholar
2024

Towards Autonomous Tool Utilization in Language Models: A Unified, Efficient and Scalable Framework

COLING 2024main

In recent research, significant advancements have been achieved in tool learning for large language models. Looking towards future advanced studies, the issue of fully autonomous tool utilization is particularly intriguing: given only a query, language models can autonomously decide whether to emplo…

2024

Triple Feature Disentanglement for One-Stage Adaptive Object Detection

AAAI 2024technical

In recent advancements concerning Domain Adaptive Object Detection (DAOD), unsupervised domain adaptation techniques have proven instrumental. These methods enable enhanced detection capabilities within unlabeled target domains by mitigating distribution differences between source and target domains…

Cited by 5SourcePDFScholar
2023

A Shared Autonomous Nursing Robot Assistant with Dynamic Workspace for Versatile Mobile Manipulation

IROS 2023poster

This paper presents a novel integration of a shared autonomous mobile humanoid robot for remote nursing assistance. The proposed nursing robot has a motorized versatile supporting structure to allow flexible integration of the system components, autonomously adjust its mobile manipulation workspace…

Cited by 6SourceScholar
2023

Adaptive Normalization for Non-stationary Time Series Forecasting: A Temporal Slice Perspective

NeurIPS 2023poster

Deep learning models have progressively advanced time series forecasting due to their powerful capacity in capturing sequence dependence. Nevertheless, it is still challenging to make accurate predictions due to the existence of non-stationarity in real-world data, denoting the data distribution rap…

2023

Autonomous Exploration and Mapping for Mobile Robots via Cumulative Curriculum Reinforcement Learning

IROS 2023poster

Deep reinforcement learning (DRL) has been widely applied in autonomous exploration and mapping tasks, but often struggles with the challenges of sampling efficiency, poor adaptability to unknown map sizes, and slow simulation speed. To speed up convergence, we combine curriculum learning (CL) with…

Cited by 10SourcecodeScholar
2023

Human Preferred Augmented Reality Visual Cues for Remote Robot Manipulation Assistance: from Direct to Supervisory Control

IROS 2023poster

When humans control or supervise remote robot manipulation, augmented reality (AR) visual cues overlaid on the remote camera video stream can effectively enhance human's remote perception of task and robot states, and comprehension of the robot autonomy's capability and intent. In this work, we cond…

Cited by 4SourceScholar
2023

MST-Q: Micro Suction Tape Quadruped Robot With High Payload Capacity

RA-L 2023

Payload capacity is a crucial factor for climbing robots, as it directly affects their ability to carry and transport heavy loads during various climbing tasks. However, many dry adhesion-based legged robots prioritize foot design from a bionic perspective to accomplish various climbing tasks while

Cited by 14SourceScholar
2023

Motor Unit Action Potential Based Classification of Hand and Arm Motions

IROS 2023poster

While motion classification architectures have improved in accuracy and robustness in recent years, computationally expensive approaches and sophisticated hardware dependencies limit their real-world applicability. To overcome these challenges, we have designed a lightweight, realtime architecture f…

Cited by 1SourceScholar
2023

PFT-SSR: Parallax Fusion Transformer for Stereo Image Super-Resolution

ICASSP 2023accepted

Stereo image super-resolution aims to boost the performance of image super-resolution by exploiting the supplementary information provided by binocular systems. Although previous methods have achieved promising results, they did not fully utilize the information of cross-view and intra-view. To furt…

Cited by 0SourceScholar
2023

Rehearsal-Free Domain Continual Face Anti-Spoofing: Generalize More and Forget Less

ICCV 2023oral

Face Anti-Spoofing (FAS) is recently studied under the continual learning setting, where the FAS models are expected to evolve after encountering data from new domains. However, existing methods need extra replay buffers to store previous data for rehearsal, which becomes infeasible when previous da…

Cited by 25PDFcodeScholar
2023

The Impacts of Unreliable Autonomy in Human-Robot Collaboration on Shared and Supervisory Control for Remote Manipulation

RA-L 2023

This work compared human-robot shared and supervisory control of remote robots for dexterous manipulation, and examined how the reliability of robot autonomy affects human operator performance, workload, and preference for robot assistance. Specifically, we implemented two human-robot collaboration

Cited by 9SourceScholar
2023

Virtual Try-On with Pose-Garment Keypoints Guided Inpainting

ICCV 2023poster

Virtual try-on is an important technology supporting online apparel shopping, which provides consumers with a virtual experience to fit garments without physically wearing them. Recently, the image-based virtual try-on has received growing research attention. However, the synthetic results of existi…

Cited by 32PDFcodeScholar
2022

Comparison of Haptic and Augmented Reality Visual Cues for Assisting Tele- manipulation

ICRA 2022poster

Robot teleoperation via human motion tracking has been proven to be easy to learn, intuitive to operate, and facilitate faster task execution than existing baselines. However, precise control while performing the dexterous telemanipulation tasks is still a challenge. In this paper, we implement sens…

Cited by 28SourceScholar
2022

Design of a Soft Gripper With Improved Microfluidic Tactile Sensors for Classification of Deformable Objects

RA-L 2022

Tactile object recognition is vital for robotic handling systems; however, existing technologies that concentrate on tactile sensors with high modulus are not suitable for soft grippers to classify deformable objects. In this letter, we integrated an indenter layer into the traditional microfluidic

Cited by 19SourceScholar
2022

Fully Adaptive Framework: Neural Computerized Adaptive Testing for Online Education

AAAI 2022technical

Computerized Adaptive Testing (CAT) refers to an efficient and personalized test mode in online education, aiming to accurately measure student proficiency level on the required subject/domain. The key component of CAT is the "adaptive" question selection algorithm, which automatically selects the b…

2022

HULC: 3D HUman Motion Capture with Pose Manifold SampLing and Dense Contact Guidance

ECCV 2022poster

"Marker-less monocular 3D human motion capture (MoCap) with scene interactions is a challenging research topic relevant for extended reality, robotics and virtual avatar generation. Due to the inherent depth ambiguity of monocular settings, 3D motions captured with existing methods often contain sev…

Cited by 31SourcePDFScholar
2022

Improving Maximum Likelihood Difference Scaling Method To Measure Inter Content Scale

ICASSP 2022accepted

The goal of most subjective studies is to place a set of stimuli on a perceptual scale. This is mostly done directly by rating, e.g. using single or double stimulus methodologies, or indirectly by ranking or pairwise comparison. All these methods estimate the perceptual magnitudes of the stimuli on…

Cited by 0SourceScholar
2022

Sen-Glove: A Lightweight Wearable Glove for Hand Assistance with Soft Joint Sensing

ICRA 2022poster

Perception and portability are critical issues for wearable gloves in hand assistive engineering. However, available wearable gloves either lack flexible sensing or are bulky. In this paper, we present a tendon-driven lightweight wearable glove with soft joint sensing, Sen-Glove. Sen-Glove is equipp…

Cited by 12SourceScholar
2022

Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic Actions

ECCV 2022poster

"Action understanding has evolved into the era of fine granularity, as most human behaviors in real life have only minor differences. To detect these fine-grained actions accurately in a label-efficient way, we tackle the problem of weakly-supervised fine-grained temporal action detection in videos…

2021

Active Telepresence Assistance for Supervisory Control: A User Study with a Multi-Camera Tele-Nursing Robot

ICRA 2021poster

Supervisory control of a humanoid robot in a manipulation task requires coordination of remote perception with robot action, which becomes more demanding with multiple moving cameras available for task supervision. We explore the use of autonomous camera control and selection to reduce operator work…

Cited by 15SourceScholar
2021

Cross-Oilfield Reservoir Classification via Multi-Scale Sensor Knowledge Transfer

AAAI 2021technical

Reservoir classification is an essential step for the exploration and production process in the oil and gas industry. An appropriate automatic reservoir classification will not only reduce the manual workloads of experts, but also help petroleum companies to make optimal decisions efficiently, which…

Cited by 8SourcePDFScholar
2021

UnClE: Explicitly Leveraging Semantic Similarity to Reduce the Parameters of Word Embeddings

EMNLP 2021finding

Natural language processing (NLP) models often require a massive number of parameters for word embeddings, which limits their application on mobile devices. Researchers have employed many approaches, e.g. adaptive inputs, to reduce the parameters of word embeddings. However, existing methods rarely…

2020

Design of a High-level Teleoperation Interface Resilient to the Effects of Unreliable Robot Autonomy

IROS 2020poster

High-level control is generally preferred for the control of complex robot platforms and by users inexperienced with robot teleoperation. However, high-level teleoperation interfaces can be less effective if the robot autonomy is not reliable. To address this problem, it is important to understand h…

Cited by 10SourceScholar
2020

Learning the Compositional Visual Coherence for Complementary Recommendations

IJCAI 2020poster

Complementary recommendations, which aim at providing users product suggestions that are supplementary and compatible with their obtained items, have become a hot topic in both academia and industry in recent years. Existing work mainly focused on modeling the co-purchased relations between two item…

Cited by 0SourcePDFScholar
2020

Shared Autonomous Interface for Reducing Physical Effort in Robot Teleoperation via Human Motion Mapping

ICRA 2020poster

Motion mapping is an intuitive method of teleoperation with a low learning curve. Our previous study investigates the physical fatigue caused by teleoperating a robot to perform general-purpose assistive tasks and this fatigue affects the operator's performance. The results from that study indicate…

Cited by 46SourceScholar
2020

Unseen Face Presentation Attack Detection with Hypersphere Loss

ICASSP 2020accepted

Presentation attack is one of the main threats to face verification systems and attracts great attention of research community. Recent methods achieve great success in intra-database test. However, the problem is more complex in practical scenario as the type of attack could be unseen to system desi…

Cited by 0SourceScholar
2019

Physical Fatigue Analysis of Assistive Robot Teleoperation via Whole-body Motion Mapping

IROS 2019poster

Robot teleoperation via motion mapping has been demonstrated to be an efficient and intuitive approach for controlling and teaching the whole-body motion coordination of humanoid robots. However, the physical fatigue in the usage of such robot teleoperation interfaces may prevent this approach to be…

Cited by 36SourceScholar
2018

Learning Coordinated Vehicle Maneuver Motion Primitives from Human Demonstration

IROS 2018poster

High-fidelity computational human models provide a safe and cost-efficient method for studying driver experience in vehicle maneuvers and for validation of vehicle design. Compared to passive human models, active human models capable of reproducing the decision-making, as well as vehicle maneuver mo…

Cited by 6SourceScholar
2017

Development of a tele-nursing mobile manipulator for remote care-giving in quarantine areas

ICRA 2017poster

During outbreaks of contagious diseases, healthcare workers are at high risk for infection due to routine interaction with patients, handling of contaminated materials, and challenges associated with safely removing protective gear. This poses an opportunity for the use of remote-controlled robots t…

Cited by 118SourceScholar