← Search

Ze Gong

9 accepted papers

2026

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR

ICML 2026poster

Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have empowered large language models (LLMs) to tackle challenging reasoning tasks such as mathematics and programming. Despite its promise, the RLVR paradigm poses significant challenges, as existing methods often suffer from s…

Cited by 0SourceScholar
2026

Learning Ordinal Probabilistic Reward from Preferences

ICLR 2026poster

Reward models are crucial for aligning large language models (LLMs) with human values and intentions. Existing approaches follow either Generative (GRMs) or Discriminative (DRMs) paradigms, yet both suffer from limitations: GRMs typically demand costly point-wise supervision, while DRMs produce unca…

Cited by 0SourceScholar
2026

RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning

ICML 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a prevailing paradigm for enhancing reasoning in Multimodal Large Language Models (MLLMs). However, relying solely on outcome supervision risks reward hacking, where models learn spurious reasoning patterns to satisfy final answer …

Cited by 4SourceScholar
2026

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels

ICML 2026poster

Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. In many settings, supervision is limited to coarse approvals or rejections of whole trajectories (e.g., whether a rollout remained within an unknown safety thresh…

Cited by 0SourceScholar
2025

Offline Safe Reinforcement Learning Using Trajectory Classification

AAAI 2025technical

Offline safe reinforcement learning (RL) has emerged as a promising approach for learning safe behaviors without engaging in risky online interactions with the environment. Most existing methods in offline safe RL rely on cost constraints at each time step (derived from global cost constraints) and…

2022

Explicable Policy Search

NeurIPS 2022accept

Human teammates often form conscious and subconscious expectations of each other during interaction. Teaming success is contingent on whether such expectations can be met. Similarly, for an intelligent agent to operate beside a human, it must consider the human’s expectation of its behavior. Disrega…

Cited by 6SourcePDFScholar
2021

Order Matters: Generating Progressive Explanations for Planning Tasks in Human-Robot Teaming

ICRA 2021poster

Prior work on generating explanations in a planning context has focused on providing the rationale behind an AI agent’s decision-making. While these methods offer the right explanations, they fail to heed the cognitive requirement of understanding an explanation from the explainee or human’s perspec…

Cited by 15SourceScholar
2020

Online Explanation Generation for Planning Tasks in Human-Robot Teaming

IROS 2020poster

As AI becomes an integral part of our lives, the development of explainable AI, embodied in the decision-making process of an AI or robotic agent, becomes imperative. For a robotic teammate, the ability to generate explanations to justify its behavior is one of the key requirements of explainable ag…

Cited by 14SourceScholar