← Search

Chuhan Wang

2 accepted papers

2026

Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization

ICLR 2026poster

Direct Preference Optimization (DPO) has recently been extended from text-only models to vision-language models. However, existing methods rely on oversimplified pairwise comparisons, generating a single negative image via basic perturbations or similarity-based retrieval, which fail to capture the…

Cited by 0SourceScholar
2026

Scheduling Adaptive Imitation Learning for Long-Horizon Dexterous Robot Micromanipulation of Deformable Cell

RA-L 2026

Robots performing collaborative, long-horizon dexterity cell micromanipulation tasks are challenging and practically significant, such as stripping intact cell membranes, which is considered as one of the most technically demanding procedures. Imitation learning approach is expected to address the c

Cited by 2SourceScholar