← Search

Huy Le

9 accepted papers

2026

PAWS: Preference Learning with Advantage-Weighted Segments

ICML 2026poster

Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods typically train utility functions on trajectory or segment-level preferences while relying on per-step utility estimates…

Cited by 0SourceScholar
2026

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

AAAI 2026technical

As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In non-Markovian settings, critical decision

Cited by 0SourcePDFScholar
2026

Trust-Region Diffusion Policies for Massively Parallel On-Policy RL

ICML 2026poster

Reinforcement learning with massively parallel simulations has become an emerging trend; however, most existing approaches still rely on simple Gaussian policy parameterizations. Diffusion models provide a more expressive policy class and have shown strong performance on challenging control problems…

Cited by 0SourceScholar
2025

Enhancing Exploration With Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation

RA-L 2025

Learning diverse policies for non-prehensile manipulation is essential for improving skill transfer and generalization to out-of-distribution scenarios. In this work, we enhance exploration through a two- fold approach within a hybrid framework that tackles both discrete and continuous action spaces

Cited by 3SourcecodeScholar
2025

Geometry-aware RL for Manipulation of Varying Shapes and Deformable Objects

ICLR 2025oral

Manipulating objects with varying geometries and deformable objects is a major challenge in robotics. Tasks such as insertion with different objects or cloth hanging require precise control and effective modelling of complex dynamics. In this work, we frame this problem through the lens of a heterog…

2024

Pseudo Labeling and Contextual Curriculum Learning for Online Grasp Learning in Robotic Bin Picking

ICRA 2024poster

The prevailing grasp prediction methods predominantly rely on offline learning, overlooking the dynamic grasp learning that occurs during real-time adaptation to novel picking scenarios. These scenarios may involve previously unseen objects, variations in camera perspectives, and bin configurations,…

Cited by 1SourceScholar
2024

WAVER: Writing-Style Agnostic Text-Video Retrieval Via Distilling Vision-Language Models Through Open-Vocabulary Knowledge

ICASSP 2024accepted

Text-video retrieval, a prominent sub-field within the domain of multimodal information retrieval, has witnessed remarkable growth in recent years. However, existing methods assume video scenes are consistent with unbiased descriptions. These limitations fail to align with real-world scenarios since…

Cited by 0SourceScholar