← Search

Huaimin Wang

24 accepted papers

2026

CoPE: A Framework for Optimizing Coordination between Planning and Execution in LLM-based Agents

ICML 2026poster

Fine-tuning Large Language Models (LLMs) as autonomous agents on domain-specific data has emerged as a promising paradigm for tackling interactive, real-world tasks. However, existing studies have overlooked the critical coordination between long-term planning and multi-step execution in optimizing …

Cited by 0SourceScholar
2026

Decouple Your Discovery and Memory in Continual Generalized Category Discovery

CVPR 2026

Continual Generalized Category Discovery (C-GCD) seeks to incrementally discover new categories from unlabeled data and memorize old categories' knowledge, fostering model adaptability in real-world scenarios. Especially, the unlabeled data is from both old and new classes, requiring the model to re

Cited by 0SourceScholar
2026

Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual Learning

CVPR 2026

Continual Visual Question Answering (Continual VQA) poses unique challenges for multimodal continual learning, requiring models to incrementally acquire new knowledge while preserving visual-semantic grounding across tasks. However, existing benchmarks hinder fair and robust evaluation of such capab

Cited by 0SourcecodeScholar
2025

Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

AAAI 2025technical

Agents significantly enhance the capabilities of standalone Large Language Models (LLMs) by perceiving environments, making decisions, and executing actions. However, LLM agents still face challenges in tasks that require multiple decision-making steps. Estimating the value of actions in specific ta…

Cited by 8SourcePDFScholar
2025

Improving the Continuity of Goal-Achievement Ability via Policy Self-Regularization for Goal-Conditioned Reinforcement Learning

ICML 2025poster

This paper addresses the challenge of discontinuity in goal-achievement capabilities observed in Goal-conditioned Reinforcement Learning (GCRL) algorithms. Through a theoretical analysis, we identify that the reuse of successful trajectories or policies during training can aid in achieving adjacent…

Cited by 0SourcePDFScholar
2025

Knowledge Memorization and Rumination for Pre-trained Model-based Class-Incremental Learning

CVPR 2025poster

Class-Incremental Learning (CIL) enables models to continuously learn new classes while mitigating catastrophic forgetting. Recently, Pre-Trained Models (PTMs) have greatly enhanced CIL performance, even when fine-tuning is limited to the first task. This advantage is particularly beneficial for CIL…

2025

Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental Learning

AAAI 2025technical

Logit-based knowledge distillation (KD) is commonly used to mitigate catastrophic forgetting in class-incremental learning (CIL) caused by data distribution shifts. However, the strict match of logit values between student and teacher models conflicts with the cross-entropy (CE) loss objective of le…

2025

V-Pilot: A Velocity Vector Control Agent for Fixed-Wing UAVs from Imperfect Demonstrations

ICRA 2025

This paper addresses the challenge of Velocity Vector Control (VVC) for fixed-wing UAVs using Reinforcement Learning (RL) in the presence of imperfect demonstrations. The multi-objective and long-horizon nature of VVC introduces significant spatial and temporal complexities, complicating RL's explor

Cited by 0SourceScholar
2025

VVC-Gym: A Fixed-Wing UAV Reinforcement Learning Environment for Multi-Goal Long-Horizon Problems

ICLR 2025poster

Multi-goal long-horizon problems are prevalent in real-world applications. The additional goal space introduced by multi-goal problems intensifies the spatial complexity of exploration; meanwhile, the long interaction sequences in long-horizon problems exacerbate the temporal complexity of explorati…

Cited by 0SourcePDFScholar
2024

C3F: Constant Collaboration and Communication Framework for Graph-Representation Dynamic Multi-Robotic Systems

RA-L 2024

Deep reinforcement learning (DRL) methods have been widely applied in distributed multi-robotic systems and successfully realized autonomous learning in many fields. In these fields, robots need to communicate and collaborate with other robots in real time, and reach agreed cognition for task assign

Cited by 0SourceScholar
2024

Iterative Regularized Policy Optimization with Imperfect Demonstrations

ICML 2024poster

Imitation learning heavily relies on the quality of provided demonstrations. In scenarios where demonstrations are imperfect and rare, a prevalent approach for refining policies is through online fine-tuning with reinforcement learning, in which a Kullback–Leibler (KL) regularization is often employ…

2024

Optimistic Model Rollouts for Pessimistic Offline Policy Optimization

AAAI 2024technical

Model-based offline reinforcement learning (RL) has made remarkable progress, offering a promising avenue for improving generalization with synthetic model rollouts. Existing works primarily focus on incorporating pessimism for policy optimization, usually via constructing a Pessimistic Markov Decis…

Cited by 1SourcePDFScholar
2024

Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language Tasks

NeurIPS 2024poster

Continual learning (CL) empowers pre-trained vision-language (VL) models to efficiently adapt to a sequence of downstream tasks. However, these models often encounter challenges in retaining previously acquired skills due to parameter shifts and limited access to historical data. In response, recent…

2023

Complementary Learning System Based Intrinsic Reward in Reinforcement Learning

ICASSP 2023accepted

Deep reinforcement learning has achieved encouraging performance in many realms. However, one of its primary challenges is the sparsity of extrinsic rewards, which is still far from solved. Complementary learning system theory suggests that effective human learning relies on two complementary learni…

Cited by 0SourceScholar
2023

Diversifying Message Aggregation in Multi-Agent Communication Via Normalized Tensor Nuclear Norm Regularization

ICASSP 2023accepted

The use of graph attention networks (GAT) in communication-enhanced multi-agent reinforcement learning (Comm-MARL) has become prevalent. While successful, GAT can lead to homogeneity in the strategies of message aggregation, which can severely limit multi-agent coordination. To address this challeng…

Cited by 0SourceScholar
2022

CRMRL: Collaborative Relationship Meta Reinforcement Learning for Effectively Adapting to Type Changes in Multi-Robotic System

RA-L 2022

Multi-agent reinforcement learning methods have been widely used for multi-robotic systems, and meta-learning methods are also applied to help robots reuse prior experiences to guide new tasks learning. But in some multi-robotic tasks, the robot types cannot be determined in advance or may dynamical

Cited by 7SourceScholar
2022

Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast

IJCAI 2022poster

We present an approach to learn voice-face representations from the talking face videos, without any identity labels. Previous works employ cross-modal instance discrimination tasks to establish the correlation of voice and face. These methods neglect the semantic content of different videos, introd…

2021

Improving Ultrasound Tongue Contour Extraction Using U-Net and Shape Consistency-Based Regularizer

ICASSP 2021accepted

B-mode ultrasound tongue imaging is widely used to visualize the tongue motion, due to its appearing properties. Extracting the tongue surface contour in the B-mode ultrasound image is still a challenge, while it is a prerequisite for further quantitative analysis. Recently, deep learning-based appr…

Cited by 0SourceScholar
2020

Online Meta-Critic Learning for Off-Policy Actor-Critic Methods

NeurIPS 2020poster

Off-Policy Actor-Critic (OffP-AC) methods have proven successful in a variety of continuous control tasks. Normally, the critic's action-value function is updated using temporal-difference, and the critic in turn provides a loss for the actor that trains it to take actions with higher expected retur…

2019

Denoising Convolutional Autoencoder Based B-mode Ultrasound Tongue Image Feature Extraction

ICASSP 2019accepted

B-mode ultrasound tongue imaging is widely used in the speech production field. However, efficient interpretation is in a great need for the tongue image sequences. Inspired by the recent success of unsupervised deep learning approach, we explore unsupervised convolutional network architecture for t…

Cited by 0SourceScholar
2019

Predicting Tongue Motion in Unlabeled Ultrasound Videos Using Convolutional Lstm Neural Networks

ICASSP 2019accepted

A challenge in speech production research is to predict future tongue movements based on a short period of past tongue movements. This study tackles speaker-dependent tongue motion prediction problem in unlabeled ultrasound videos with convolutional long short-term memory (ConvLSTM) networks. The mo…

Cited by 0SourceScholar
2018

How Many Robots are Enough: A Multi-Objective Genetic Algorithm for the Single-Objective Time-Limited Complete Coverage Problem

ICRA 2018poster

Complete coverage, which is the foundation of many robotic applications, aims to cover an area as quickly as possible. This study investigates the time-limited version of multi-robot complete coverage problem, that is, to find the least number of robots and allocate tasks properly to them such that…

Cited by 19SourceScholar