← Search

Pengyu Zhao

9 accepted papers

2026

DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use

ICML 2026poster

Recent work increasingly synthesizes agentic tasks for post-training tool-using LLMs, yet robust generalization under shifts in tasks and toolsets remains an open challenge. We trace this brittleness to insufficient diversity in synthesized training tasks. Scaling diversity is difficult because trai…

Cited by 0SourceScholar
2026

Transferring Policy of Offline Reinforcement Learning From Hybrid Dataset to Real World via Progressive Neural Network

RA-L 2026

Offline reinforcement learning (Offline RL) provides a compelling solution for applying RL in high-risk or resourceconstrained real-world domains such as healthcare, autonomous driving, and robotic manipulation, where online exploration can be unsafe or impractical. However, Offline RL faces critica

Cited by 0SourceScholar
2026

Transferring Policy of Offline Reinforcement Learning from Hybrid Dataset to Real World Via Progressive Neural Network

ICRA 2026poster

Offline reinforcement learning (Offline RL) provides a compelling solution for applying RL in high-risk or resource-constrained real-world domains such as healthcare, autonomous driving, and robotic manipulation. However, Offline RL faces critical challenges arising from limited data coverage and po…

Cited by 0SourceScholar
2025

SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond

NeurIPS 2025poster

Recent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). While open-source replication efforts have primarily focused on mathematical and coding domains, methods and resources for…

Cited by 0SourcecodeScholar
2021

AMEIR: Automatic Behavior Modeling, Interaction Exploration and MLP Investigation in the Recommender System

IJCAI 2021poster

Recently, deep learning models have been widely explored in recommender systems. Though having achieved remarkable success, the design of task-aware recommendation models usually requires manual feature engineering and architecture engineering from domain experts. To relieve those efforts, we explor…

Cited by 17SourcePDFScholar
2020

Adversarial Oracular Seq2seq Learning for Sequential Recommendation

IJCAI 2020poster

Recently, sequential recommendation has become a significant demand for many real-world applications, where the recommended items would be displayed to users one after another and the order of the displays influences the satisfaction of users. An extensive number of models have been developed for se…

Cited by 0SourcePDFScholar
2020

Differentiable Feature Aggregation Search for Knowledge Distillation

ECCV 2020poster

Knowledge distillation has become increasingly important in model compression. It boosts the performance of a miniaturized student network with the supervision of the output distribution and feature maps from a sophisticated teacher network. Some recent works introduce multi-teacher distillation to…

Cited by 54SourcePDFScholar
2020

Preference-Aware Mask for Session-Based Recommendation with Bidirectional Transformer

ICASSP 2020accepted

User profiles are not always visible in E-commerce scenarios, in which case the recommender systems can only summarize users' preferences through sessions of historical records. However, the items in a session might be irrelevant to users' preferences or become the disturbances for modelling the use…

Cited by 0SourceScholar
2019

LadderNet: Knowledge Transfer Based Viewpoint Prediction in 360◦ Video

ICASSP 2019accepted

In the past few years, virtual reality (VR) has become an enabling technique, not only for enriching our visual experience but also for providing new channels for businesses. Untethered mobile devices are the main players for watching 360-degree content, thereby the precision of predicting the futur…

Cited by 0SourceScholar