← Search

Donglin Wang

73 accepted papers

2026

Activation-wise Propagation: A One-Timestep Strategy for Spiking Neural Networks

AAAI 2026technical

Spiking neural networks (SNNs) have demonstrated significant potential in real-time multi-sensor perception tasks due to their event-driven and parameter-efficient characteristics. A key challenge is the timestep-wise iterative update of neuronal hidden states (membrane potentials), which complicate

Cited by 0SourcePDFScholar
2026

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

AAAI 2026technical

Vision-Language-Action (VLA) models based on flow matching have shown excellent performance in general-purpose robotic manipulation tasks. However, the action accuracy of these models on complex downstream tasks is unsatisfactory. One important reason is that these models rely solely on the post-tra

Cited by 0SourcePDFScholar
2026

CUBic: Coordinated Unified Bimanual Perception and Control Framework

CVPR 2026

Recent advances in visuomotor policy learning have enabled robots to perform control directly from visual inputs. Yet, extending such end-to-end learning from single-arm to bimanual manipulation remains challenging due to the need for both independent perception and coordinated interaction between a

Cited by 0SourceScholar
2026

Development of a Mixed-Control Ankle Assist Device with Sensor-Fusion-Based Phase Recognition for Walking Exercise Promotion

ICRA 2026poster

"Frail" elderly often experience walking impairments that limit independence and sustained physical activity. Although various assistive devices exist, many rely on single-mode control, limiting adaptability, responsiveness to gait variability, and voluntary motion. To improve, we developed a wearab…

Cited by 0Scholar
2026

Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics

ICML 2026poster

Video action models are a promising foundation for Vision–Language–Action (VLA) because they can learn rich visual dynamics directly from video. However, likelihood-oriented training of diffusion predictors emphasizes globally plausible futures and does not guarantee precision-critical visual dynami…

Cited by 0SourceScholar
2026

Empowering Precise Embodied Agents with Executable Analytic Concepts as Semantic-Physical Blueprints

IJCAI 2026

A core challenge for embodied agents is the ``semantic-to-physical gap"—the difficulty of mapping symbolic reasoning to precise execution. While Vision-Language Models (VLMs) enhance agent task planning, they often fail in problem classes requiring accurate alignment between functional geometry and

Cited by 0Scholar
2026

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation and thus suffering from temporal myopia that degrades long-horizon coherence. In

Cited by 0SourcecodeScholar
2026

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL

ICML 2026poster

Offline goal-conditioned RL (GCRL) learns goal-reaching policies from static datasets, but real-world datasets are often partially observable and history-dependent, exhibiting a mix of Markovian and non-Markovian that violate standard RL assumptions. History-aware sequence models such as Decision Tr…

Cited by 0SourceScholar
2026

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

AAAI 2026technical

Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. Instead, visual attention is always dispe

Cited by 0SourcePDFScholar
2026

Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach

AAAI 2026technical

Compared to single-target adversarial attacks, multi-target attacks have garnered significant attention due to their ability to generate adversarial images for multiple target classes simultaneously. However, existing generative approaches for multi-target attacks primarily encode target labels into

Cited by 0SourcePDFScholar
2026

Rethinking the Practicality of Vision-Language-Action Model: A Comprehensive Benchmark and an Improved Baseline

ICRA 2026poster

Vision-Language-Action (VLA) models have emerged as a generalist robotic agent. However, existing VLAs are hindered by excessive parameter scales, prohibitive pre-training requirements, and limited applicability to diverse embodiments. To improve the practicality of VLAs, we propose a comprehensive …

2026

Robust Online Residual Refinement Via Koopman-Guided Dynamics Modeling

ICRA 2026poster

Imitation learning (IL) enables efficient skill acquisition from demonstrations but often struggles with long-horizon tasks and high-precision control due to compounding errors. Residual policy learning offers a promising, model-agnostic solution by refining a base policy through closed-loop correct…

2026

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model

ICLR 2026poster

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data, which lack accurate spatial awareness and hinder their abili…

Cited by 0SourcecodeScholar
2026

TrajBooster: Boosting Humanoid Whole-Body Manipulation Via Trajectory-Centric Learning

ICRA 2026poster

Recent Vision-Language-Action (VLA) models show potential to generalize across embodiments but struggle to quickly align with a new robot’s action space when high-quality demonstrations are scarce, especially for bipedal humanoids. We present TrajBooster, a cross-embodiment framework that leverages …

2026

Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Diffusion Diffusion Process

ICLR 2026poster

Vision-language-action (VLA) models aim to understand natural language instructions and visual observations and execute corresponding actions as an embodied agent. Recent advancements have integrated future images into the understanding-action loop, enabling foresight-driven policies that reduce abs…

Cited by 0SourcecodeScholar
2026

VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model

AAAI 2026technical

Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how

Cited by 0SourcePDFScholar
2025

Boundary-to-Region Supervision for Offline Safe Reinforcement Learning

NeurIPS 2025poster

Offline safe reinforcement learning aims to learn policies that satisfy predefined safety constraints from static datasets. Existing sequence-model-based methods condition action generation on symmetric input tokens for return-to-go and cost-to-go, neglecting their intrinsic asymmetry: RTG serves as…

Cited by 0SourceScholar
2025

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

ICCV 2025accepted

In robotic visuomotor policy learning, diffusion-based models have achieved significant success in improving the accuracy of action trajectory generation compared to traditional autoregressive models. However, they suffer from inefficiency due to multiple denoising steps and limited flexibility from…

Cited by 0SourcePDFScholar
2025

CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls

AAAI 2025technical

Lyric-to-melody generation is a highly challenging task in the field of AI music generation. Due to the difficulty of learning strict yet weak correlations between lyrics and melodies, previous methods have suffered from weak controllability, low-quality and poorly structured generation. To address…

2025

Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

AAAI 2025technical

In recent years, applying multi-modal large language models (MLLMs) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, MLLMs comprise the well-known Transformer network, which has a less efficient quadratic computation complexity. In this s…

2025

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

ICLR 2025poster

With the rapid development of embodied artificial intelligence, significant progress has been made in vision-language-action (VLA) models for general robot decision-making. However, the majority of existing VLAs fail to account for the inevitable external perturbations encountered during deployment.…

Cited by 2SourcePDFScholar
2025

Integrating Trajectory Optimization and Reinforcement Learning for Quadrupedal Jumping with Terrain-Adaptive Landing

IROS 2025

Jumping constitutes an essential component of quadruped robots’ locomotion capabilities, which includes dynamic take-off and adaptive landing. Existing quadrupedal jumping studies mainly focused on the stance and flight phase by assuming a flat landing ground, which is impractical in many real world

Cited by 1SourceScholar
2025

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation

CoRL 2025poster

Vision-Language-Action (VLA) models have become a cornerstone in robotic policy learning, leveraging large-scale multimodal data for robust and scalable control. However, existing VLA frameworks primarily address short-horizon tasks, and their effectiveness on long-horizon, multi-step robotic manipu…

Cited by 0SourceScholar
2025

MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models

ICRA 2025

Developing versatile quadruped robots that can smoothly perform various actions and tasks in real-world environments remains a significant challenge. This paper introduces a novel vision-language-action (VLA) model, mixture of robotic experts (MoRE), for quadruped robots that aim to introduce reinfo

Cited by 24SourceScholar
2025

Multi-Task Multi-Agent Reinforcement Learning via Skill Graphs

RA-L 2025

Multi-task multi-agent reinforcement learning (M T-MARL) has recently gained attention for its potential to enhance MARL's adaptability across multiple tasks. However, it is challenging for existing multi-task learning methods to handle complex problems, as they are unable to handle unrelated tasks

Cited by 3SourcecodeScholar
2025

PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding

IROS 2025

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in

Cited by 60SourceScholar
2025

Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning

ICRA 2025

This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the l

Cited by 1SourcecodeScholar
2025

ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning

ICML 2025poster

Vision-Language-Action (VLA) models have shown great potential in general robotic decision-making tasks via imitation learning. However, the variable quality of training data often constrains the performance of these models. On the other hand, offline Reinforcement Learning (RL) excels at learning r…

Cited by 0SourcePDFScholar
2025

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation

ICML 2025poster

Behavior Cloning (BC) is a widely adopted visual imitation learning method in robot manipulation. Current BC approaches often enhance generalization by leveraging large datasets and incorporating additional visual and textual modalities to capture more diverse information. However, these methods ove…

Cited by 0SourcePDFScholar
2025

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

NeurIPS 2025poster

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for integrating spatial cues, such as point clouds or depth, either require specialized sensors or fail to effectively exploit d…

Cited by 0SourcecodeScholar
2025

Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport

ICML 2025poster

Diffusion policies have shown promise in learning complex behaviors from demonstrations, particularly for tasks requiring precise control and long-term planning. However, they face challenges in robustness when encountering distribution shifts. This paper explores improving diffusion-based imitation…

2025

Stay Hungry, Keep Learning: Sustainable Plasticity for Deep Reinforcement Learning

ICML 2025poster

The integration of Deep Neural Networks in Reinforcement Learning (RL) systems has led to remarkable progress in solving complex tasks but also introduced challenges like primacy bias and dead neurons. Primacy bias skews learning towards early experiences, while dead neurons diminish the network's c…

Cited by 0SourcePDFScholar
2025

VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot Manipulation

ICLR 2025poster

Vision-language-action models (VLAs) have recently become highly prevalent in robot manipulation due to its end-to-end architecture and impressive performance. However, current VLAs are limited to processing human instructions in textual form, neglecting the more natural speech modality for human in…

2024

Beyond OOD State Actions: Supported Cross-Domain Offline Reinforcement Learning

AAAI 2024technical

Offline reinforcement learning (RL) aims to learn a policy using only pre-collected and fixed data. Although avoiding the time-consuming online interactions in RL, it poses challenges for out-of-distribution (OOD) state actions and often suffers from data inefficiency for training. Despite many effo…

2024

DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation

ICML 2024poster

In this paper, we propose a novel approach called DIffusion-guided DIversity (DIDI) for offline behavioral generation. The goal of DIDI is to learn a diverse set of skills from a mixture of label-free offline data. We achieve this by leveraging diffusion probabilistic models as priors to guide the l…

2024

Expressive Forecasting of 3D Whole-Body Human Motions

AAAI 2024technical

Human motion forecasting, with the goal of estimating future human behavior over a period of time, is a fundamental task in many real-world applications. However, existing works typically concentrate on foretelling the major joints of the human body without considering the delicate movements of the…

2024

GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot

IROS 2024poster

Multi-task robot learning holds significant importance in tackling diverse and complex scenarios. However, current approaches are hindered by performance issues and difficulties in collecting training datasets. In this paper, we propose GeRM (Generalist Robotic Model). We utilize offline reinforceme…

Cited by 13SourcecodeScholar
2024

Graph-Based Environment Representation for Vision-and-Language Navigation in Continuous Environments

ICASSP 2024accepted

The Vision-and-Language Navigation in Continuous Environments (VLN-CE) task requires an agent to follow a language instruction in a realistic environment. Understanding the environment is crucial, yet current methods are relatively simple and direct, without delving into the interplay between langua…

Cited by 0SourceScholar
2024

Improving Cross-Domain Few-Shot Classification with Multilayer Perceptron

ICASSP 2024accepted

Cross-domain few-shot classification (CDFSC) is a challenging and tough task due to the significant distribution discrepancies across different domains. To address this challenge, many approaches aim to learn transferable representations. Multilayer perceptron (MLP) has shown its capability to learn…

Cited by 0SourceScholar
2024

Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation

CVPR 2024poster

This study focuses on a novel task in text-to-image (T2I) generation namely action customization. The objective of this task is to learn the co-existing action from limited data and generalize it to unseen humans or even animals. Experimental results show that existing subject-driven customization m…

2024

Nash CoT: Multi-Path Inference with Preference Equilibrium

EMNLP 2024main

Chain of thought (CoT) is a reasoning framework that can enhance the performance of large language models (LLMs) on complex inference tasks. In particular, among various studies related to CoT, multi-path inference stands out as a simple yet effective improvement. However, there is no optimal settin…

2024

PiTe: Pixel-Temporal Alignment for Large Video-Language Model

ECCV 2024oral

"Fueled by the Large Language Models (LLMs) wave, Large Visual-Language Models (LVLMs) have emerged as a pivotal advancement, bridging the gap between image and text. However, video making it challenging for LVLMs to perform adequately due to the complexity of the relationship between language and s…

2024

Prompt-Based Distribution Alignment for Unsupervised Domain Adaptation

AAAI 2024technical

Recently, despite the unprecedented success of large pre-trained visual-language models (VLMs) on a wide range of downstream tasks, the real-world unsupervised domain adaptation (UDA) problem is still not well explored. Therefore, in this paper, we first experimentally demonstrate that the unsupervi…

2024

RL2AC: Reinforcement Learning-based Rapid Online Adaptive Control for Legged Robot Robust Locomotion

RSS 2024poster

Dynamic fast adaptation is one of the basic capabilities that enables the animals to timely and properly adjust its locomotion reacting to the unpredictable changes. Such capability is also essential for the quadruped robot, when working in the unforseen environment. While reinforcement learning (RL…

Cited by 6SourcePDFScholar
2024

Reinformer: Max-Return Sequence Modeling for Offline RL

ICML 2024poster

As a data-driven paradigm, offline reinforcement learning (RL) has been formulated as sequence modeling that conditions on the hindsight information including returns, goal or future trajectory. Although promising, this supervised paradigm overlooks the core objective of RL that maximizes the return…

2024

Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning

CVPR 2024poster

Recent compositional zero-shot learning (CZSL) methods adapt pre-trained vision-language models (VLMs) by constructing trainable prompts only for composed state-object pairs. Relying on learning the joint representation of seen compositions these methods ignore the explicit modeling of the state and…

2024

VGDIFFZERO: Text-To-Image Diffusion Models Can Be Zero-Shot Visual Grounders

ICASSP 2024accepted

Large-scale text-to-image diffusion models have shown impressive capabilities for generative tasks by leveraging strong vision-language alignment from pre-training. However, most vision-language discriminative tasks require extensive fine-tuning on carefully-labeled datasets to acquire such alignmen…

Cited by 0SourceScholar
2023

A Composite Control Strategy for Quadruped Robot by Integrating Reinforcement Learning and Model-Based Control

IROS 2023poster

Locomotion in the wild requires the quadruped robot to have strong capabilities in adaptation and robustness. The deep reinforcement learning (DRL) exhibits the huge potential in environmental adaptability, while its stability issues remain open. On the other hand, the quadruped robot dynamic model…

Cited by 7SourceScholar
2023

Beyond Reward: Offline Preference-guided Policy Optimization

ICML 2023poster

This study focuses on the topic of offline preference-based reinforcement learning (PbRL), a variant of conventional reinforcement learning that dispenses with the need for online interaction or specification of reward functions. Instead, the agent is provided with fixed offline trajectories and hum…

2023

CEIL: Generalized Contextual Imitation Learning

NeurIPS 2023poster

In this paper, we present ContExtual Imitation Learning (CEIL), a general and broadly applicable algorithm for imitation learning (IL). Inspired by the formulation of hindsight information matching, we derive CEIL by explicitly learning a hindsight embedding function together with a contextual polic…

Cited by 23SourcePDFScholar
2023

Design from Policies: Conservative Test-Time Adaptation for Offline Policy Optimization

NeurIPS 2023poster

In this work, we decouple the iterative bi-level offline RL (value estimation and policy extraction) from the offline training phase, forming a non-iterative bi-level paradigm and avoiding the iterative error propagation over two levels. Specifically, this non-iterative paradigm allows us to conduct…

Cited by 10SourcePDFScholar
2023

VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal Retrieval

CVPR 2023poster

Many recent studies leverage the pre-trained CLIP for text-video cross-modal retrieval by tuning the backbone with additional heavy modules, which not only brings huge computational burdens with much more parameters, but also leads to the knowledge forgetting from upstream models. In this work, we p…

2022

DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement Learning

ICLR 2022poster

Offline reinforcement learning algorithms promise to be applicable in settings where a fixed dataset is available and no new experience can be acquired. However, such formulation is inevitably offline-data-hungry and, in practice, collecting a large offline dataset for one specific task over one spe…

Cited by 48SourcePDFScholar
2022

Domain Generalized Few-Shot Image Classification via Meta Regularization Network

ICASSP 2022accepted

In few-shot image classification scenarios, meta-learning methods aim to learn transferable feature representations extracted from seen domains (base classes) in the meta-training phase and quickly adapt to unseen domains (novel classes) in the meta-testing phase. However, when seen and unseen domai…

Cited by 0SourceScholar
2022

End-to-End Open-Set Semi-Supervised Node Classification with Out-of-Distribution Detection

IJCAI 2022poster

Out-Of-Distribution (OOD) samples are prevalent in real-world applications. The OOD issue becomes even more severe on graph data, as the effect of OOD nodes can be potentially amplified by propagation through the graph topology. Recent works have considered the OOD detection problem, which is critic…

Cited by 17SourcePDFScholar
2022

Learn Goal-Conditioned Policy with Intrinsic Motivation for Deep Reinforcement Learning

AAAI 2022technical

It is of significance for an agent to autonomously explore the environment and learn a widely applicable and general-purpose goal-conditioned policy that can achieve diverse goals including images and text descriptions. Considering such perceptually-specific goals, one natural approach is to reward…

Cited by 23SourcePDFScholar
2022

Tree Structure-Aware Few-Shot Image Classification via Hierarchical Aggregation

ECCV 2022poster

"In this paper, we mainly focus on the problem of how to learn additional feature representations for few-shot image classification through pretext tasks (e.g., rotation or color permutation and so on). This additional knowledge generated by pretext tasks can further improve the performance of few-s…

2021

Attributes-Guided and Pure-Visual Attention Alignment for Few-Shot Recognition

AAAI 2021technical

The purpose of few-shot recognition is to recognize novel categories with a limited number of labeled examples in each class. To encourage learning from a supplementary view, recent approaches have introduced auxiliary semantic modalities into effective metric-learning frameworks that aim to learn a…

2021

Deep Transfer Tensor Decomposition with Orthogonal Constraint for Recommender Systems

AAAI 2021technical

Tensor decomposition is one of the most effective techniques for multi-criteria recommendations. However, it suffers from data sparsity when dealing with three-dimensional (3D) user-item-criterion ratings. To mitigate this issue, we consider effectively incorporating the side information and cross-d…

Cited by 52SourcePDFScholar
2021

Hierarchical Terrain-Aware Control for Quadrupedal Locomotion by Combining Deep Reinforcement Learning and Optimal Control

IROS 2021poster

Quadruped robots possess advantages on different terrains over other types of mobile robots by virtue of their flexible choices of foothold points. It is crucial to integrate terrain perception with motion planning to exploit the potential of quadruped robots. We propose a novel hierarchical terrain…

Cited by 10SourceScholar
2021

Terrain-Aware Risk-Assessment-Network-Aided Deep Reinforcement Learning for Quadrupedal Locomotion in Tough Terrain

IROS 2021poster

When it comes to the control system of quadruped robots, deep reinforcement learning (DRL) is considered to be a promising solution. Despite years of development in this field, difficulties remain in guaranteeing the action stability of DRL-based quadruped robots’ locomotion, especially in tough ter…

Cited by 6SourceScholar
2021

Unsupervised Domain Adaptation with Dynamics-Aware Rewards in Reinforcement Learning

NeurIPS 2021poster

Unsupervised reinforcement learning aims to acquire skills without prior goal representations, where an agent automatically explores an open-ended environment to represent goals and learn the goal-conditioned policy. However, this procedure is often time-consuming, limiting the rollout in some poten…

Cited by 22SourcePDFScholar
2020

Independent Skill Transfer for Deep Reinforcement Learning

IJCAI 2020poster

Recently, diverse primitive skills have been learned by adopting the entropy as intrinsic reward, which further shows that new practical skills can be produced by combining a variety of primitive skills. This is essentially skill transfer, very useful for learning high-level skills but quite challen…