← Search

He Li

51 accepted papers

2026

ARGH-Mark: Anchor-Synchronized Watermarking with Hamming Correction for Robust and Quality-Preserving LLM Attribution

AAAI 2026technical

The proliferation of large language models has intensified demands for reliable content attribution, yet existing watermarking techniques face a fundamental trilemma: they cannot simultaneously optimize for robustness against attacks, minimal text quality degradation, and detection efficiency. To re

Cited by 0SourcePDFScholar
2026

DeepTracer: Tracing Stolen Model via Deep Coupled Watermarks

AAAI 2026technical

Model watermarking techniques can embed watermark information into the protected model for ownership declaration by constructing specific input-output pairs. However, existing watermarks are easily removed when facing model stealing attacks, and make it difficult for model owners to effectively veri

Cited by 0SourcePDFScholar
2026

FedSDR: Federated Graph Learning with Structural Noise Detection and Reconstruction

CVPR 2026

Federated Graph Learning (FGL) has emerged as a principled framework for decentralized training of Graph Neural Networks (GNNs) while preserving data privacy. In subgraph-FL scenarios, however, structural noise arising from data collection and storage can damage the GNN message-passing scheme of cli

Cited by 0SourcecodeScholar
2026

Generalizing from References using a Multi-Task Reference and Goal-Driven RL Framework

RSS 2026poster

Learning agile humanoid behaviors from human motion offers a powerful route to natural, coordinated control, but existing approaches face a persistent trade-off: reference-tracking policies are often brittle outside the demonstration dataset, while purely task-driven Reinforcement Learning (RL) can …

Cited by 0SourceScholar
2026

LLM-Driven Scenario-Aware Planning for Autonomous Driving

ICASSP 2026poster

Hybrid planner switching framework (HPSF) for autonomous driving needs to reconcile high-speed driving efficiency with safe maneuvering in dense traffic. Existing HPSF methods often fail to make reliable mode transitions or sustain efficient driving in congested environments, owing to heuristic scen…

Cited by 0SourcePDFScholar
2026

MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs

ICML 2026poster

Multimodal large language models (MLLMs) are trained on massive multimodal data, making data unlearning increasingly important as data owners may request the removal of specific content. In practice, these requests often arrive sequentially over time, giving rise to the challenging problem of *MLLM …

Cited by 0SourceScholar
2026

PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space

ICML 2026spotlight

The remarkable success of Chain-of-Thought (CoT), which enhances performance by scaling generation steps at test-time, inspires us to ask: can we leverage a similar scaling of computational steps during pretraining to improve the generation of each individual token? To address this, we propose a nov…

Cited by 0SourceScholar
2026

PonderLM: Pretraining Language Models to Ponder in Continuous Space

ICLR 2026poster

Humans ponder before articulating complex sentence elements, enabling deeper cognitive processing through focused effort. In this work, we introduce this pondering process into language models by repeatedly invoking the forward process within a single token generation step. During pondering, instead…

Cited by 0SourcecodeScholar
2026

Rethinking Federated Prompt Learning for Medical Images: From Textual Tuning to Visual Manifold Anchoring

ICML 2026poster

Federated Prompt Learning (FPL) adapts Vision-Language Models to privacy-sensitive medical imaging, typically via a textual tuning paradigm that assumes the frozen visual encoder provides a discriminative feature geometry. We argue this assumption breaks down in medical settings, leading to two geom…

Cited by 0SourceScholar
2026

ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving

ICML 2026poster

Safety-critical scenarios are central to evaluating autonomous driving systems, yet their rarity in naturalistic logs makes simulation-based stress testing indispensable. Most scenario generation methods treat surrounding agents as adversaries, but they either (i) induce failures without explicitly …

Cited by 0SourceScholar
2026

Towards Realistic Lifelong Re-identification: Identity Recurrence with Changing Clothes

ICML 2026poster

Existing lifelong person re-identification (Re-ID) methods assume that each identity maintains a relatively stable appearance distribution over time. However, in real-world scenarios, identities often reappear asynchronously with substantial clothing changes, which is not modeled in existing lifelon…

Cited by 0SourceScholar
2026

WHU-MARS: A Multispectral Aerial-Ground Benchmark Towards Any-Scenario Person Re-Identification

CVPR 2026

Recent person re-identification (ReID) leverages heterogeneous sensing with multiple modalities and viewpoints to improve robustness across diverse conditions. However, most approaches target predefined scenario pairs (e.g., visible-infrared or aerial-ground) and train separate task-specific models.

Cited by 0SourcecodeScholar
2025

$S^2$FGL: Spatial Spectral Federated Graph Learning

ICML 2025poster

Federated Graph Learning (FGL) combines the privacy-preserving capabilities of federated learning (FL) with the strong graph modeling capability of Graph Neural Networks (GNNs). Current research addresses subgraph-FL only from the structural perspective, neglecting the propagation of graph signals o…

2025

A sEMG-based Muscle-Compensation-Error Elimination Module for the Home-Based Bilateral Rehabilitation System

RA-L 2025

Home-based bilateral training using exoskeletons is considered an effective rehabilitation method to restore the normal motion function of stroke patients' limbs. However, due to the imbalance of limb muscle strength, muscle compensation of the functional upper limb will cause surface electromyograp

Cited by 1SourceScholar
2025

Be Confident: Uncovering Overfitting in MLLM Multi-Task Tuning

ICML 2025poster

Fine-tuning Multimodal Large Language Models (MLLMs) in multi-task learning scenarios has emerged as an effective strategy for achieving cross-domain specialization. However, multi-task fine-tuning frequently induces performance degradation on open-response datasets. We posit that free-form answer g…

Cited by 0SourcePDFScholar
2025

Catch Your Emotion: Sharpening Emotion Perception in Multimodal Large Language Models

ICML 2025spotlight

Multimodal large language models (MLLMs) have achieved impressive progress in tasks such as visual question answering and visual understanding, but they still face significant challenges in emotional reasoning. Current methods to enhance emotional understanding typically rely on fine-tuning or manua…

Cited by 0SourcePDFScholar
2025

Cheb-GR: Rethinking K-nearest Neighbor Search in Re-ranking for Person Re-identification

CVPR 2025poster

Person re-identification (ReID) is the task of matching individuals across different camera views. Existing approaches typically employ neural networks to extract discriminative features, ranking gallery images based on their similarities to probe images. While effective, these methods are often enh…

2025

DKDR: Dynamic Knowledge Distillation for Reliability in Federated Learning

NeurIPS 2025poster

Federated Learning (FL) has demonstrated a promising future in privacy-friendly collaboration but it faces the data heterogeneity problem. Knowledge Distillation (KD) can serve as an effective method to address this issue. However, challenges arise from the unreliability of existing distillation met…

Cited by 0SourcecodeScholar
2025

EAGLES: Towards Effective, Efficient, and Economical Federated Graph Learning via Unified Sparsification

ICML 2025poster

Federated Graph Learning (FGL) has gained significant attention as a privacy-preserving approach to collaborative learning, but the computational demands increase substantially as datasets grow and Graph Neural Network (GNN) layers deepen. To address these challenges, we propose $\textbf{EAGLES}$, a…

Cited by 0SourcePDFScholar
2025

FedSPA: Generalizable Federated Graph Learning under Homophily Heterogeneity

CVPR 2025poster

Federated Graph Learning (FGL) has emerged as a solution to address real-world privacy concerns and data silos in graph learning, which relies on Graph Neural Networks (GNNs). Nevertheless, the homophily level discrepancies within the local graph data of clients, termed homophily heterogeneity, sign…

2025

Horae: A Domain-Agnostic Language for Automated Service Regulation

IJCAI 2025

Artificial intelligence is rapidly encroaching on the field of service regulation. However, existing AI-based regulation techniques are often tailored to specific application domains and thus are difficult to generalize in an automated manner. This paper presents Horae, a unified specification langu

2025

Learn from Downstream and Be Yourself in Multimodal Large Language Models Fine-Tuning

ICML 2025poster

Multimodal Large Language Model (MLLM) has demonstrated strong generalization capabilities across diverse distributions and tasks, largely due to extensive pre-training datasets. Fine-tuning MLLM has become a common practice to improve performance on specific downstream tasks. However, during fine-t…

Cited by 9SourcePDFScholar
2025

MoodAngels: A Retrieval-augmented Multi-agent Framework for Psychiatry Diagnosis

NeurIPS 2025poster

The application of AI in psychiatric diagnosis faces significant challenges, including the subjective nature of mental health assessments, symptom overlap across disorders, and privacy constraints limiting data availability. To address these issues, we present MoodAngels, the first specialized multi…

Cited by 0SourceScholar
2025

NightReID: A Large-Scale Nighttime Person Re-Identification Benchmark

AAAI 2025technical

Person re-identification (Re-ID) is crucial for intelligent surveillance systems, facilitating the identification of individuals across multiple camera views. While significant advancements have been made for daytime scenarios, ensuring reliable Re-ID performance during nighttime remains a significa…

2025

Pixel-wise Divide and Conquer for Federated Vessel Segmentation

IJCAI 2025

Accurate vessel segmentation is essential for diagnosing and managing vascular and ophthalmic diseases. Traditional learning-based vessel segmentation methods heavily rely on high-quality, pixel-level annotated datasets. However, segmentation performance suffers significantly when applied in federat

Cited by 0SourcePDFScholar
2025

Prototype-guided Knowledge Propagation with Adaptive Learning for Lifelong Person Re-identification

IJCAI 2025

Lifelong Person Re-identification (LReID) is essential in dynamic camera networks, which continually adapts to new environments while preserving previously acquired knowledge. Existing LReID techniques often preserve samples from past datasets to maintain old knowledge, potentially leading to privac

2025

TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification

IROS 2025

Visual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information. However, existing methods remain limited in indoor settings due to the highly repetitive structures inherent in such environments.

Cited by 1SourcecodeScholar
2025

Transformer-Based Spatial-Temporal Counterfactual Outcomes Estimation

ICML 2025poster

The real world naturally has dimensions of time and space. Therefore, estimating the counterfactual outcomes with spatial-temporal attributes is a crucial problem. However, previous methods are based on classical statistical models, which still have limitations in performance and generalization. Thi…

2025

Watermarking with Low-Entropy POS-Guided Token Partitioning and Z-Score-Driven Dynamic Bias for Large Language Models

EMNLP 2025

Texts generated by large language models (LLMs) are increasingly widespread online. Due to the lack of effective attribution mechanisms, the enforcement of copyright and the prevention of misuse remain significant challenges in the context of LLM-generated content. LLMs watermark emerges as a crucia

Cited by 0SourcePDFScholar
2024

Autoregressive Image Generation without Vector Quantization

NeurIPS 2024spotlight

Conventional wisdom holds that autoregressive models for image generation are typically accompanied by vector-quantized tokens. We observe that while a discrete-valued space can facilitate representing a categorical distribution, it is not a necessity for autoregressive modeling. In this work, we pr…

2024

CipherDM: Secure Three-Party Inference for Diffusion Model Sampling

ECCV 2024poster

"Diffusion Models (DMs) achieve state-of-the-art synthesis results in image generation and have been applied to various fields. However, DMs sometimes seriously violate user privacy during usage, making the protection of privacy an urgent issue. Using traditional privacy computing schemes like Secur…

2024

Multi-Uncertainty Aware Autonomous Cooperative Planning

IROS 2024poster

Autonomous cooperative planning (ACP) is a promising technique to improve the efficiency and safety of multi-vehicle interactions for future intelligent transportation systems. However, realizing robust ACP is a challenge due to the aggregation of perception, motion, and communication uncertainties.…

Cited by 1SourceScholar
2024

Parameter Disparities Dissection for Backdoor Defense in Heterogeneous Federated Learning

NeurIPS 2024poster

Backdoor attacks pose a serious threat to federated systems, where malicious clients optimize on the triggered distribution to mislead the global model towards a predefined target. Existing backdoor defense methods typically require either homogeneous assumption, validation datasets, or client optim…

Cited by 3SourcePDFScholar
2024

Seamless Virtual Reality With Integrated Synchronizer and Synthesizer for Autonomous Driving

RA-L 2024

Virtual reality (VR) is a promising data engine for autonomous driving (AD). However, data fidelity in this paradigm is often degraded by VR inconsistency, for which the existing VR approaches become ineffective, as they ignore the inter-dependency between low-level VR synchronizer designs (i.e., da

Cited by 8SourceScholar
2024

Self-Driven Entropy Aggregation for Byzantine-Robust Heterogeneous Federated Learning

ICML 2024poster

Federated learning presents massive potential for privacy-friendly collaboration. However, the performance of federated learning is deeply affected by byzantine attacks, where malicious clients deliberately upload crafted vicious updates. While various robust aggregations have been proposed to defen…

Cited by 5SourcePDFScholar
2024

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

NeurIPS 2024poster

Causal reasoning capability is critical in advancing large language models (LLMs) towards artificial general intelligence (AGI). While versatile LLMs appear to have demonstrated capabilities in understanding contextual causality and providing responses that obey the laws of causality, it remains unc…

2023

A Unified Perspective on Multiple Shooting In Differential Dynamic Programming

IROS 2023poster

Differential Dynamic Programming (DDP) is an efficient computational tool for solving nonlinear optimal control problems. It was originally designed as a single shooting method and thus is sensitive to the initial guess supplied. This work considers the extension of DDP to multiple shooting (MS), im…

Cited by 16SourceScholar
2023

Rethinking Federated Learning With Domain Shift: A Prototype View

CVPR 2023poster

Federated learning shows a bright promise as a privacy-preserving collaborative learning technique. However, prevalent solutions mainly focus on all private data sampled from the same domain. An important challenge is that when distributed data are derived from diverse domains. The private model pre…

2023

Subject-Independent Estimation of Continuous Movements Using CNN-LSTM for a Home-Based Upper Limb Rehabilitation System

RA-L 2023

Exoskeleton-assisted home-based rehabilitation plays a vital role in the upper limb rehabilitation of stroke patients in early stage. The surface electromyography (sEMG)-based control can facilitate friendly interactions between individuals and rehabilitation exoskeletons. The exoskeleton can also m

Cited by 30SourceScholar
2023

Versatile Real-Time Motion Synthesis via Kino-Dynamic MPC With Hybrid-Systems DDP

ICRA 2023poster

Specialized motions such as jumping are often achieved on quadruped robots by solving a trajectory optimization problem once and executing the trajectory using a tracking controller. This approach is in parallel with Model Predictive Control (MPC) strategies that commonly control regular gaits via o…

Cited by 18SourcecodeScholar
2022

Mini Cheetah, the Falling Cat: A Case Study in Machine Learning and Trajectory Optimization for Robot Acrobatics

ICRA 2022poster

Seemingly in defiance of basic physics, cats consistently land on their feet after falling. In this paper, we design a controller that lands the Mini Cheetah quadruped robot on its feet as well. Specifically, we explore how trajectory optimization and machine learning can work together to enable hig…

Cited by 45SourceScholar
2022

Zero-Shot Retargeting of Learned Quadruped Locomotion Policies Using Hybrid Kinodynamic Model Predictive Control

IROS 2022poster

Reinforcement Learning (RL) has witnessed great strides for quadruped locomotion, with continued progress in the reliable sim-to-real transfer of policies. However, it remains a challenge to reuse a policy on another robot, which could save time for retraining. In this work, we present a framework f…

Cited by 14SourceScholar
2021

Decentralized, Unlabeled Multi-Agent Navigation in Obstacle-Rich Environments using Graph Neural Networks

IROS 2021poster

We propose a decentralized, learning-based solution to the challenging problem of unlabeled multi-agent navigation among obstacles, where robots need to simultaneously tackle the problems of goal assignment, local collision avoidance, and navigation. Our method has each robot infer their desired act…

Cited by 20SourceScholar
2018

Robust Video Content Alignment and Compensation for Rain Removal in a CNN Framework

CVPR 2018poster

Rain removal is important for improving the robustness of outdoor vision based systems. Current rain removal methods show limitations either for complex dynamic scenes shot from fast moving cameras, or under torrential rain fall with opaque occlusions. We propose a novel derain algorithm, which appl…

Cited by 209SourcePDFScholar
2016

A Rotary-Percussive Ultrasonic Drill for planetary rock sampling

IROS 2016poster

Conventional ultrasonic drills can drill into rocks using high frequency axial vibration, which feature lower power and lower preload force, thus they could become a very attractive solution for future deep-space exploration. However, drill cuttings cannot be removed effectively with the increase of…

Cited by 6SourceScholar