← Search

Tianmeng Hu

5 accepted papers

2026

Beyond Monotonicity: Revisiting Factorization Principles in Multi-Agent Q-Learning

AAAI 2026technical

Value decomposition is a central approach in multi-agent reinforcement learning (MARL), enabling centralized training with decentralized execution by factorizing the global value function into local values. To ensure individual-global-max (IGM) consistency, existing methods either enforce monotonici

Cited by 0SourcePDFScholar
2026

Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning

ICML 2026poster

Multi-objective reinforcement learning (MORL) seeks policies that effectively balance conflicting objectives. However, presenting many diverse policies without accounting for the decision maker’s (DM’s) preferences can overwhelm the decision-making process. On the other hand, accurately specifying p…

Cited by 0SourceScholar
2026

RIDER: 3D RNA Inverse Design with Reinforcement Learning-Guided Diffusion

ICLR 2026poster

The inverse design of RNA three-dimensional (3D) structures is crucial for engineering functional RNAs in synthetic biology and therapeutics. While recent deep learning approaches have advanced this field, they are typically optimized and evaluated using native sequence recovery, which is a limited…

Cited by 0SourcecodeScholar
2024

PA2D-MORL: Pareto Ascent Directional Decomposition Based Multi-Objective Reinforcement Learning

AAAI 2024technical

Multi-objective reinforcement learning (MORL) provides an effective solution for decision-making problems involving conflicting objectives. However, achieving high-quality approximations to the Pareto policy set remains challenging, especially in complex tasks with continuous or high-dimensional sta…

Cited by 2SourcePDFScholar