← Search

Jingyue Gao

6 accepted papers

2025

MARGE: Improving Math Reasoning with Guided Exploration

ICML 2025poster

Large Language Models (LLMs) exhibit strong potential in mathematical reasoning, yet their effectiveness is often limited by a shortage of high-quality queries. This limitation necessitates scaling up computational responses through self-generated data, yet current methods struggle due to spurious c…

Cited by 0SourcePDFScholar
2023

Decentralized Motor Skill Learning for Complex Robotic Systems

RA-L 2023

Reinforcement learning (RL) has achieved remarkable success in complex robotic systems (eg. quadruped locomotion). In previous works, the RL-based controller was typically implemented as a single neural network with concatenated observation input. However, the corresponding learned policy is highly

Cited by 9SourceScholar
2022

Reinforcement learning with Demonstrations from Mismatched Task under Sparse Reward

CoRL 2022poster

Reinforcement learning often suffer from the sparse reward issue in real-world robotics problems. Learning from demonstration (LfD) is an effective way to eliminate this problem, which leverages collected expert data to aid online learning. Prior works often assume that the learning agent and the ex…

Cited by 6SourceScholar
2021

Learning Groupwise Explanations for Black-Box Models

IJCAI 2021poster

We study two user demands that are important during the exploitation of explanations in practice: 1) understanding the overall model behavior faithfully with limited cognitive load and 2) predicting the model behavior accurately on unseen instances. We illustrate that the two user demands correspond…

2021

U-BERT: Pre-training User Representations for Improved Recommendation

AAAI 2021technical

Learning user representation is a critical task for recommendation systems as it can encode user preference for personalized services. User representation is generally learned from behavior data, such as clicking interactions and review comments. However, for less popular domains, the behavior data…

Cited by 152SourcePDFScholar