← Search

Wei Liang

34 accepted papers

2026

COLA: Learning Human-Humanoid Coordination for Collaborative Object Carrying

ICRA 2026poster

Human-humanoid collaboration shows significant promise for applications in healthcare, domestic assistance, and manufacturing. While compliant robot-human collaboration has been extensively developed for robotic arms, enabling compliant human-humanoid collaboration remains largely unexplored due to …

Cited by 0Scholar
2026

RadarMP: Motion Perception for 4D mmWave Radar in Autonomous Driving

AAAI 2026technical

Accurate 3D scene motion perception significantly enhances the safety and reliability of an autonomous driving system. Benefiting from its all-weather operational capability and unique perceptual properties, 4D mmWave radar has emerged as an essential component in advanced autonomous driving. Howeve

Cited by 0SourcePDFScholar
2025

CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks

CoRL 2025poster

Humanoid robot teleoperation plays a vital role in demonstrating and collecting data for complex interactions. Current methods suffer from two key limitations: (1) restricted controllability due to decoupled upper- and lower-body control, and (2) severe drift caused by open-loop execution. These iss…

Cited by 0SourceScholar
2025

FloNa: Floor Plan Guided Embodied Visual Navigation

AAAI 2025technical

Humans naturally rely on floor plans to navigate in unfamiliar environments, as they are readily available, reliable, and provide rich geometrical guidance. However, existing visual navigation settings overlook this valuable prior knowledge, leading to limited efficiency and accuracy. To eliminate t…

Cited by 0SourcePDFScholar
2025

Heterogeneous Adversarial Play in Interactive Environments

NeurIPS 2025poster

Self-play constitutes a fundamental paradigm for autonomous skill acquisition, whereby agents iteratively enhance their capabilities through self-directed environmental exploration. Conventional self-play frameworks exploit agent symmetry within zero-sum competitive settings, yet this approach prove…

Cited by 0SourceScholar
2025

IPP-Net: A Generalizable Deep Neural Network Model for Indoor Pathloss Radio Map Prediction

ICASSP 2025accepted

In this paper, we propose a generalizable deep neural network model for indoor pathloss radio map prediction (termed as IPP-Net). IPP-Net is based on a UNet architecture and learned from both large-scale ray tracing simulation data and a modified 3GPP indoor hotspot model. The performance of IPP-Net…

Cited by 0SourceScholar
2025

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

ICCV 2025poster

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus on grounding objects in static observations from 3D reconstru…

Cited by 0SourcePDFScholar
2025

STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

ICASSP 2025accepted

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned V…

Cited by 0SourceScholar
2024

A Lightweight Powered Knee Prosthesis Replicating Early-Stance Knee Flexion During Level Walking

RA-L 2024

Powered knee prostheses promise to improve the mobility of transfemoral amputees by imitating the biomechanics of the missing knee joint. Unfortunately, the heavy weight and short battery life severely limit the application of powered prostheses. Here, we present a lightweight powered knee prosthesi

Cited by 2SourceScholar
2024

Mastering Scene Rearrangement with Expert-Assisted Curriculum Learning and Adaptive Trade-Off Tree-Search

IROS 2024poster

Scene Rearrangement Planning (SRP) has recently emerged as a crucial interior scene task; however, current approaches still face two primary issues. First, prior works define the action space of SRP using handcrafted coarse-grained actions, which are inflexible for scene arrangement transition and i…

Cited by 0SourcecodeScholar
2024

Move as You Say Interact as You Can: Language-guided Human Motion Generation with Scene Affordance

CVPR 2024highlight

Despite significant advancements in text-to-motion synthesis generating language-guided human motion within 3D environments poses substantial challenges. These challenges stem primarily from (i) the absence of powerful generative models capable of jointly modeling natural language 3D scenes and huma…

2024

Prompt-guided Precise Audio Editing with Diffusion Models

ICML 2024poster

Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio t…

Cited by 2SourcePDFScholar
2024

Unsupervised Learning of Facial Optical Flow via Occlusion-Aware Global-Local Matching

ICASSP 2024accepted

Estimating optical flow from facial videos is an essential preprocessing step for many applications. However, it is a challenging task as the facial videos contain rich expressions, large displacements, and complex occlusions. Obtaining the ground truth optical flow for facial videos is very difficu…

Cited by 0SourceScholar
2024

Visual Loop Closure Detection with Thorough Temporal and Spatial Context Exploitation

IROS 2024poster

Despite advancements in visual Simultaneous Localization and Mapping (SLAM), prevailing visual Loop Closure Detection (LCD) methods primarily rely on computationally intensive image similarity comparisons, neglecting temporal-spatial context during long-term exploration. To address this issue, we pr…

Cited by 0SourceScholar
2023

DREAMWALKER: Mental Planning for Continuous Vision-Language Navigation

ICCV 2023poster

VLN-CE is a recently released embodied task, where AI agents need to navigate a freely traversable environment to reach a distant target location, given language instructions. It poses great challenges due to the huge space of possible strategies. Driven by the belief that the ability to anticipate…

Cited by 38PDFcodeScholar
2023

Diffusion-Based Generation, Optimization, and Planning in 3D Scenes

CVPR 2023poster

We introduce SceneDiffuser, a conditional generative model for 3D scene understanding. SceneDiffuser provides a unified model for solving scene-conditioned generation, optimization, and planning. In contrast to prior works, SceneDiffuser is intrinsically scene-aware, physics-based, and goal-oriented…

2023

Discovering the Real Association: Multimodal Causal Reasoning in Video Question Answering

CVPR 2023poster

Video Question Answering (VideoQA) is challenging as it requires capturing accurate correlations between modalities from redundant information. Recent methods focus on the explicit challenges of the task, e.g. multimodal feature extraction, video-text alignment and fusion. Their frameworks reason th…

2023

MEWL: Few-shot multimodal word learning with referential uncertainty

ICML 2023poster

Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning…

2022

Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation

CVPR 2022poster

Since the rise of vision-language navigation (VLN), great progress has been made in instruction following -- building a follower to navigate environments under the guidance of instructions. However, far less attention has been paid to the inverse task: instruction generation -- learning a speaker to…

Cited by 62PDFcodeScholar
2022

Design and Validation of a Polycentric Hybrid Knee Prosthesis With Electromagnet-Controlled Mode Transition

RA-L 2022

A hybrid knee prosthesis is proposed in this letter, which consists of a polycentric structure in passive mode for low-torque activities and a single-axis structure in active mode for high-torque activities. A novel mode transition mechanism controls self-holding electromagnets for switching modes b

Cited by 5SourceScholar
2022

HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes

NeurIPS 2022accept

Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characters of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality and lack semantics. To fill in the gap, we propose a large-scale an…

2021

Structured Scene Memory for Vision-Language Navigation

CVPR 2021poster

Recently, numerous algorithms have been developed to tackle the problem of vision-language navigation (VLN), i.e., entailing an agent to navigate 3D environments through following linguistic instructions. However, current VLN agents simply store their past experiences/observations as latent states i…

Cited by 136PDFcodeScholar
2020

Active Visual Information Gathering for Vision-Language Navigation

ECCV 2020poster

Vision-language navigation (VLN) is the task of entailing an agent to carry out navigational instructions inside photo-realistic environments. One of the key challenges in VLN is how to conduct a robust navigation by mitigating the uncertainty caused by ambiguous instructions and insufficient observ…

2017

Kinematic chain based multi-joint capturing using monocular visual-inertial measurements

IROS 2017poster

Combining light-weight visual and inertial modalities for motion capturing has been popular in robotics researches. There exist scale ambiguity, inaccurate pose estimation with little or no baseline, incremental drifts over time in visual-inertial fusion. Thus, in this paper, we propose a robust mot…

Cited by 2SourceScholar