← Search

Weixuan Liu

2 accepted papers

2026

Learning from Comparison: Constrained Projection Policy Optimization for Pareto-Front Improvement

ICML 2026poster

Constrained multi-objective reinforcement learning aims to discover a diverse set of feasible trade-offs, yet scalarization and signed, normalized group-relative advantages can be brittle under objective-scale drift, near-ties, and feasibility scarcity. We propose constrained projection policy optim…

Cited by 0SourceScholar
2026

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

ICLR 2026poster

Chain-of-Thought (CoT) reasoning has proven effective in enhancing large language models by encouraging step-by-step intermediate reasoning, and recent advances have extended this paradigm to Multimodal Large Language Models (MLLMs). In the medical domain, where diagnostic decisions depend on nuance…

Cited by 0SourceScholar