← Search

Junge Zhang

33 accepted papers

2026

Agentic Reinforcement Learning with Implicit Step Rewards

ICLR 2026poster

Large language models (LLMs) are increasingly developed as autonomous agents using reinforcement learning (agentic RL) that reason and act in interactive environments. However, sparse and sometimes unverifiable rewards make it extremely challenging to assign credit when training LLM agents that serv…

Cited by 0SourceScholar
2026

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

CVPR 2026

The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has sparked interest in extending these models to 3D scenes, with the goal of human-like visual-spatial intelligence. However, achieving deep spatial understanding comparable to human capabilities remains challenging fo

Cited by 0SourcecodeScholar
2025

EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning

ACL 2025long

Large Language Models (LLMs) have shown impressive reasoning capabilities in well-defined problems with clear solutions, such as mathematics and coding. However, they still struggle with complex real-world scenarios like business negotiations, which require strategic reasoning—an ability to navigate…

2025

IDEA-Bench: How Far are Generative Models from Professional Designing?

CVPR 2025poster

Recent advancements in image generation models enable the creation of high-quality images and targeted modifications based on textual instructions. Some models even support multimodal complex guidance and demonstrate robust task generalization capabilities. However, they still fall short of meeting…

2025

Sequential Preference Optimization: Multi-Dimensional Preference Alignment with Implicit Reward Modeling

AAAI 2025technical

Human preference alignment is critical in building powerful and reliable large language models (LLMs). However, current methods either ignore the multi-dimensionality of human preferences (e.g. helpfulness and harmlessness) or struggle with the complexity of managing multiple reward models. To addre…

2025

Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes

ICCV 2025poster

Reconstructing 3D scenes from a single image is a fundamentally ill-posed task due to the severely under-constrained nature of the problem. Consequently, when the scene is rendered from novel camera views, particularly in unseen regions far away from the input camera, existing single image to 3D rec…

Cited by 0SourcePDFScholar
2024

ADMN: Agent-Driven Modular Network for Dynamic Parameter Sharing in Cooperative Multi-Agent Reinforcement Learning

IJCAI 2024poster

Parameter sharing is a common strategy in multi-agent reinforcement learning (MARL) to make the training more efficient and scalable. However, applying parameter sharing among agents indiscriminately hinders the emergence of agents diversity and degrades the final cooperative performance. To better…

Cited by 1SourcePDFScholar
2024

BadRL: Sparse Targeted Backdoor Attack against Reinforcement Learning

AAAI 2024technical

Backdoor attacks in reinforcement learning (RL) have previously employed intense attack strategies to ensure attack success. However, these methods suffer from high attack costs and increased detectability. In this work, we propose a novel approach, BadRL, which focuses on conducting highly sparse b…

2024

NeRF-LiDAR: Generating Realistic LiDAR Point Clouds with Neural Radiance Fields

AAAI 2024technical

Labelling LiDAR point clouds for training autonomous driving is extremely expensive and difficult. LiDAR simulation aims at generating realistic LiDAR data with labels for training and verifying self-driving algorithms more efficiently. Recently, Neural Radiance Fields (NeRF) have been proposed for…

2024

Position: Foundation Agents as the Paradigm Shift for Decision Making

ICML 2024poster

Decision making demands intricate interplay between perception, memory, and reasoning to discern optimal policies. Conventional approaches to decision making face challenges related to low sample efficiency and poor generalization. In contrast, foundation models in language and vision have showcased…

2024

ProAgent: Building Proactive Cooperative Agents with Large Language Models

AAAI 2024technical

Building agents with adaptive behavior in cooperative tasks stands as a paramount goal in the realm of multi-agent systems. Current approaches to developing cooperative agents rely primarily on learning-based methods, whose policy generalization depends heavily on the diversity of teammates they int…

2024

TAPE: Leveraging Agent Topology for Cooperative Multi-Agent Policy Gradient

AAAI 2024technical

Multi-Agent Policy Gradient (MAPG) has made significant progress in recent years. However, centralized critics in state-of-the-art MAPG methods still face the centralized-decentralized mismatch (CDM) issue, which means sub-optimal actions by some agents will affect other agent's policy learning. Whi…

2024

Task-Wise Prompt Query Function for Rehearsal-Free Continual Learning

ICASSP 2024accepted

Continual learning (CL) aims to enable a model to retain knowledge of old tasks while learning new ones. One effective approach to CL is based on data rehearsal method. However, this approach increases the cost of storing data and cannot be used when data from old tasks is unavailable for some reaso…

Cited by 0SourceScholar
2023

Subspace-Aware Exploration for Sparse-Reward Multi-Agent Tasks

AAAI 2023technical

Exploration under sparse rewards is a key challenge for multi-agent reinforcement learning problems. One possible solution to this issue is to exploit inherent task structures for an acceleration of exploration. In this paper, we present a novel exploration approach, which encodes a special structur…

Cited by 8SourcePDFScholar
2021

Coordinated Proximal Policy Optimization

NeurIPS 2021poster

We present Coordinated Proximal Policy Optimization (CoPPO), an algorithm that extends the original Proximal Policy Optimization (PPO) to the multi-agent setting. The key idea lies in the coordinated adaptation of step size during the policy update process among multiple agents. We prove the monoton…

2021

Learning to Reweight Imaginary Transitions for Model-Based Reinforcement Learning

AAAI 2021technical

Model-based reinforcement learning (RL) is more sample efficient than model-free RL by using imaginary trajectories generated by the learned dynamics model. When the model is inaccurate or biased, imaginary trajectories may be deleterious for training the action-value and policy functions. To allevi…

2021

SOFT: Softmax-free Transformer with Linear Complexity

NeurIPS 2021spotlight

Vision transformers (ViTs) have pushed the state-of-the-art for various visual recognition tasks by patch-wise image tokenization followed by self-attention. However, the employment of self-attention modules results in a quadratic complexity in both computation and memory usage. Various attempts on…

Cited by 198SourcePDFScholar
2019

Transductive Zero-Shot Learning with Visual Structure Constraint

NeurIPS 2019poster

To recognize objects of the unseen classes, most existing Zero-Shot Learning (ZSL) methods first learn a compatible projection function between the common semantic space and the visual space based on the data of source seen classes, then directly apply it to the target unseen classes. However, in re…

2018

A2-RL: Aesthetics Aware Reinforcement Learning for Image Cropping

CVPR 2018poster

Image cropping aims at improving the aesthetic quality of images by adjusting their composition. Most weakly supervised cropping methods (without bounding box supervision) rely on the sliding window mechanism. The sliding window mechanism requires fixed aspect ratios and limits the cropping region w…

2018

Discriminative Learning of Latent Features for Zero-Shot Recognition

CVPR 2018poster

Zero-shot learning (ZSL) aims to recognize unseen image categories by learning an embedding space between image and semantic representations. For years, among existing works, it has been the center task to learn the proper mapping matrices aligning the visual and semantic space, whilst the importanc…

Cited by 196SourcePDFScholar
2015

Beyond Tree Structure Models: A New Occlusion Aware Graphical Model for Human Pose Estimation

ICCV 2015poster

Occlusion is a main challenge for human pose estimation, which is largely ignored in popular tree structure models. The tree structure model is simple and convenient for exact inference, but short in modeling the occlusion coherence especially in the case of self-occlusion. We propose an occlusion a…

Cited by 28PDFScholar
2015

GRSA: Generalized Range Swap Algorithm for the Efficient Optimization of MRFs

CVPR 2015poster

Markov Random Field (MRF) is an important tool and has been widely used in many vision tasks. Thus, the optimization of MRFs is a problem of fundamental importance. Recently, Veskler and Kumar et. al propose the range move algorithms, which are one of the most successful solvers to this problem. How…

Cited by 8SourcePDFScholar