← Search

Yuexin Bian

4 accepted papers

2026

DiffOP: Reinforcement Learning of Optimization-Based Control Policies via Implicit Policy Gradients

AAAI 2026technical

Real-world control systems require policies that are not only high-performing but also interpretable and robust. A promising direction toward this goal is model-based control, which learns system dynamics and cost functions from historical data and then uses these models to inform decision-making. B

Cited by 0SourcePDFScholar
2026

LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts

ICLR 2026poster

Recent studies have shown that combining parameter-efficient fine-tuning (PEFT) with mixture-of-experts (MoE) is an effective strategy for adapting large language models (LLMs) to the downstream tasks. However, most existing approaches rely on conventional TopK routing, which requires careful hyperp…

Cited by 0SourcecodeScholar
2026

RN-D: Discretized Categorical Actors with Regularized Networks for On-Policy Reinforcement Learning

ICML 2026poster

On-policy deep reinforcement learning remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradients are noisy and policy updates must be conservative. In this paper, w…

Cited by 0SourceScholar
2025

Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

NeurIPS 2025poster

Reinforcement learning (RL) has shown promise in enhancing large language model (LLM) reasoning, yet progress towards broader capabilities is limited by the availability of high-quality, multi-domain datasets. This work introduces \ours, a 92K RL-for-reasoning dataset designed to address this gap, c…

Cited by 0SourceScholar