← Search

Min Cheng

6 accepted papers

2026

Diffusion Blend: Inference-Time Multi-Preference Alignment for Diffusion Models

ICLR 2026poster

Reinforcement learning (RL) algorithms have been used recently to align diffusion models with downstream objectives such as aesthetic quality and text-image consistency by fine-tuning them to maximize a single reward function under a fixed KL regularization. However, this approach is inherently rest…

Cited by 0SourcecodeScholar
2026

Learning Surgical Robotic Manipulation with 3D Spatial Priors

CVPR 2026

Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or enhance multi-view features by adding wrist-mounted cameras to supplement the d

Cited by 0SourceScholar
2025

A Sequential Approach for Accurate Parameters Identification of Heavy-Duty Hydraulic Manipulators Ensuring Physical Feasibility

RA-L 2025

Accurate identification of dynamic parameters is essential for precise motion control and autonomous operation of heavy-duty hydraulic manipulators. However, due to their low-speed motion property, conventional approaches fail to simultaneously excite all parameters. To overcome this issue, a sequen

Cited by 3SourceScholar
2024

Provable Policy Gradient Methods for Average-Reward Markov Potential Games

AISTATS 2024poster

We study Markov potential games under the infinite horizon average reward criterion. Most previous studies have been for discounted rewards. We prove that both algorithms based on independent policy gradient and independent natural policy gradient converge globally to a Nash equilibrium for the aver…

Cited by 8SourcePDFScholar
2023

Natural Actor-Critic for Robust Reinforcement Learning with Function Approximation

NeurIPS 2023poster

We study robust reinforcement learning (RL) with the goal of determining a well-performing policy that is robust against model mismatch between the training simulator and the testing environment. Previous policy-based robust RL algorithms mainly focus on the tabular setting under uncertainty sets th…