← Search

Junyu Zhang

22 accepted papers

2026

Counterfactual Planning for Generalizable Agents’ Actions

AAAI 2026technical

Large language models have revolutionized agent planning by serving as the engine of heuristic guidance. However, LLM-based agents often struggle to generalize across complex environments and to adapt to stochastic feedback arising from environment–action interactions. We propose Counterfactual Plan

Cited by 0SourcePDFScholar
2026

Global Policy-Space Response Oracles for Two-Player Zero-Sum Games

ICML 2026poster

The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy set using deep reinforcement learning (DRL). A central challenge is to construct, under limited computational budgets, a small strategy population …

Cited by 0SourceScholar
2025

AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time

EMNLP 2025

This paper presents AlphaOne ( 𝛼1 ), a universal framework for modulating reasoning progress in large reasoning models (LRMs) at test time. 𝛼1 first introduces 𝛼 moment, which represents the scaled thinking phase with a universal parameter 𝛼 .Within this scaled pre- 𝛼 moment phase, it dynamically sc

2025

DiffuseHigh: Training-Free Progressive High-Resolution Image Synthesis Through Structure Guidance

AAAI 2025technical

Large-scale generative models, such as text-to-image diffusion models, have garnered widespread attention across diverse domains due to their creative and high-fidelity image generation. Nonetheless, existing large-scale diffusion models are confined to generating images of up to 1K resolution, whic…

2025

DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

ICLR 2025poster

The rapid advancements in Vision-Language Models (VLMs) have shown great potential in tackling mathematical reasoning tasks that involve visual context. Unlike humans who can reliably apply solution steps to similar problems with minor modifications, we found that state-of-the-art VLMs like GPT-4o c…

Cited by 16SourcePDFScholar
2025

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

ICML 2025oral

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents remain underexplored due to the lack of comprehensive evaluat…

2025

Towards Multi-Table Learning: A Novel Paradigm for Complementarity Quantification and Integration

NeurIPS 2025spotlight

Multi-table data integrate various entities and attributes, with potential interconnections between them. However, existing tabular learning methods often struggle to describe and leverage the underlying complementarity across distinct tables. To address this limitation, we propose the first unified…

Cited by 0SourceScholar
2024

An Improved Finite-time Analysis of Temporal Difference Learning with Deep Neural Networks

ICML 2024poster

Temporal difference (TD) learning algorithms with neural network function parameterization have well-established empirical success in many practical large-scale reinforcement learning tasks. However, theoretical understanding of these algorithms remains challenging due to the nonlinearity of the act…

Cited by 1SourcePDFScholar
2024

Nebnet: Exploiting Node-Edge Bi-Level Network for Gene Expression Prediction

ICASSP 2024accepted

Spatial Transcriptomics (ST) has made great progress in breast cancer due to it captures gene expression with fine-grained spots. It has always been low-throughout owing to reliance on special and pricey technologies. Recently, numerous types of models focus on predicting gene expression in windows…

Cited by 0SourceScholar
2023

Offline Meta Reinforcement Learning with In-Distribution Online Adaptation

ICML 2023poster

Recent offline meta-reinforcement learning (meta-RL) methods typically utilize task-dependent behavior policies (e.g., training RL agents on each individual task) to collect a multi-task dataset. However, these methods always require extra information for fast adaptation, such as offline context for…

2023

Symmetry-Aware Robot Design with Structured Subgroups

ICML 2023poster

Robot design aims at learning to create robots that can be easily controlled and perform tasks efficiently. Previous works on robot design have proven its ability to generate robots for various tasks. However, these works searched the robots directly from the vast design space and ignored common str…

2022

Multi-Agent Reinforcement Learning with General Utilities via Decentralized Shadow Reward Actor-Critic

AAAI 2022technical

We posit a new mechanism for cooperation in multi-agent reinforcement learning (MARL) based upon any nonlinear function of the team's long-term state-action occupancy measure, i.e., a general utility. This subsumes the cumulative return but also allows one to incorporate risk-sensitivity, explorati…

Cited by 12SourcePDFScholar
2021

Generalization Bounds for Stochastic Saddle Point Problems

AISTATS 2021poster

This paper studies the generalization bounds for the empirical saddle point (ESP) solution to stochastic saddle point (SSP) problems. For SSP with Lipschitz continuous and strongly convex-strongly concave objective functions, we establish an $O\left(1/n\right)$ generalization bound by using a probab…

Cited by 43SourcePDFScholar
2021

On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient Method

NeurIPS 2021spotlight

Policy gradient (PG) gives rise to a rich class of reinforcement learning (RL) methods. Recently, there has been an emerging trend to augment the existing PG methods such as REINFORCE by the \emph{variance reduction} techniques. However, all existing variance-reduced PG methods heavily rely on an u…

Cited by 85SourcePDFScholar
2020

Variational Policy Gradient Method for Reinforcement Learning with General Utilities

NeurIPS 2020spotlight

In recent years, reinforcement learning systems with general goals beyond a cumulative sum of rewards have gained traction, such as in constrained problems, exploration, and acting upon prior experiences. In this paper, we consider policy optimization in Markov Decision Problems, where the objective…

Cited by 177SourcePDFScholar