← Search

Deqiang Ouyang

5 accepted papers

2026

Improving Generalization in Offline Meta-Reinforcement Learning via Cross-task Contexts

AAAI 2026technical

Context-based offline meta-reinforcement learning (meta-RL) is a paradigm that integrates meta-learning with offline reinforcement learning. It learns a strategy to extract task-specific contexts from trajectories of meta-training tasks and leverages this strategy for adapting to unseen target tasks

Cited by 0SourcePDFScholar
2026

MetaGameBO: Hierarchical Game-Theoretic Driven Robust Meta-Learning for Bayesian Optimization

AAAI 2026technical

Meta-learning for Bayesian optimization accelerates optimization by leveraging knowledge from previous tasks, but existing methods optimize for average performance and fail on challenging outlier tasks critical in practice. These limitations become particularly severe when target tasks exhibit distr

Cited by 0SourcePDFScholar
2025

Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation

NeurIPS 2025poster

The growing demand for efficient deep learning has positioned dataset distillation as a pivotal technique for compressing training dataset while preserving model performance. However, existing inner-loop optimization methods for dataset distillation typically rely on random truncation strategies, wh…

Cited by 0SourceScholar
2025

Learning Robust Neural Processes with Risk-Averse Stochastic Optimization

ICML 2025poster

Neural processes (NPs) are a promising paradigm to enable skill transfer learning across tasks with the aid of the distribution of functions. The previous NPs employ the empirical risk minimization principle in optimization. However, the fast adaption ability to different tasks can vary widely, and…

Cited by 0SourcePDFScholar
2020

Exploring Parameter Space with Structured Noise for Meta-Reinforcement Learning

IJCAI 2020poster

Efficient exploration is a major challenge in Reinforcement Learning (RL) and has been studied extensively. However, for a new task existing methods explore either by taking actions that maximize task agnostic objectives (such as information gain) or applying a simple dithering strategy (such as noi…

Cited by 0SourcePDFScholar