← Search

Biao Luo

8 accepted papers

2026

Autonomous Partner Selection for Cooperative Multi-Agent Reinforcement Learning

AAAI 2026technical

In cooperative Multi-Agent Reinforcement Learning (MARL), the subgroup-wise learning is employed to assign sub-tasks to agents towards the enhancement of team collaboration. However, the present work is dependent on manually defined allocation criteria, which hinders its capacity to adapt to environ

Cited by 0SourcePDFScholar
2026

Beyond Monotonicity: Revisiting Factorization Principles in Multi-Agent Q-Learning

AAAI 2026technical

Value decomposition is a central approach in multi-agent reinforcement learning (MARL), enabling centralized training with decentralized execution by factorizing the global value function into local values. To ensure individual-global-max (IGM) consistency, existing methods either enforce monotonici

Cited by 0SourcePDFScholar
2026

HPS: Hyperspherical Parameter Sharing for Efficient Multi-Agent Reinforcement Learning

ICML 2026poster

Parameter Sharing (PS) is widely used to improve efficiency in Multi-Agent Reinforcement Learning (MARL), but it can limit behavioral diversity and degrade performance. This limitation stems from gradient conflicts among agents on shared weights, which hinders effective policy learning. To fully cha…

Cited by 0SourceScholar
2026

Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning

ICML 2026poster

Multi-objective reinforcement learning (MORL) seeks policies that effectively balance conflicting objectives. However, presenting many diverse policies without accounting for the decision maker’s (DM’s) preferences can overwhelm the decision-making process. On the other hand, accurately specifying p…

Cited by 0SourceScholar
2026

RIDER: 3D RNA Inverse Design with Reinforcement Learning-Guided Diffusion

ICLR 2026poster

The inverse design of RNA three-dimensional (3D) structures is crucial for engineering functional RNAs in synthetic biology and therapeutics. While recent deep learning approaches have advanced this field, they are typically optimized and evaluated using native sequence recovery, which is a limited…

Cited by 0SourcecodeScholar
2025

Thinking Before Decision: Efficient Interactive Visual Navigation Based on Local Accessibility Prediction

RA-L 2025

Embodied AI has made prominent advances in interactive visual navigation tasks based on deep reinforcement learning. In the pursuit of higher success rates in navigation, previous work has typically focused on training embodied agents to push away interactable objects on the ground. However, such in

Cited by 1SourceScholar
2024

PA2D-MORL: Pareto Ascent Directional Decomposition Based Multi-Objective Reinforcement Learning

AAAI 2024technical

Multi-objective reinforcement learning (MORL) provides an effective solution for decision-making problems involving conflicting objectives. However, achieving high-quality approximations to the Pareto policy set remains challenging, especially in complex tasks with continuous or high-dimensional sta…

Cited by 2SourcePDFScholar