← Search

Qitong Gao

9 accepted papers

2025

Variational Adversarial Training Towards Policies with Improved Robustness

AISTATS 2025poster

Reinforcement learning (RL), while being the benchmark for policy formulation, often struggles to deliver robust solutions across varying scenarios, leading to marked performance drops under environmental perturbations.~Traditional adversarial training, based on a two-player max-min game, is known t…

Cited by 0SourceScholar
2024

Off-Policy Selection for Initiating Human-Centric Experimental Design

NeurIPS 2024poster

In human-centric applications like healthcare and education, the \textit{heterogeneity} among patients and students necessitates personalized treatments and instructional interventions. While reinforcement learning (RL) has been utilized in those tasks, off-policy selection (OPS) is pivotal to close…

Cited by 0SourcePDFScholar
2024

On Trajectory Augmentations for Off-Policy Evaluation

ICLR 2024poster

In the realm of reinforcement learning (RL), off-policy evaluation (OPE) holds a pivotal position, especially in high-stake human-involved scenarios such as e-learning and healthcare. Applying OPE to these domains is often challenging with scarce and underrepresentative offline training trajectories…

Cited by 4SourcePDFScholar
2024

Steering Decision Transformers via Temporal Difference Learning

IROS 2024poster

Decision Transformers (DTs) have been highly effective for offline reinforcement learning (RL) tasks, successfully modeling the sequences of actions in a given set of demonstrations. However, DTs may perform poorly in stochastic environments, which are prevalent in robotics scenarios. In this paper,…

Cited by 0SourceScholar
2023

Off-Policy Evaluation for Human Feedback

NeurIPS 2023poster

Off-policy evaluation (OPE) is important for closing the gap between offline training and evaluation of reinforcement learning (RL), by estimating performance and/or rank of target (evaluation) policies using offline trajectories only. It can improve the safety and efficiency of data collection and…

Cited by 8SourcePDFScholar
2022

A Reinforcement Learning-Informed Pattern Mining Framework for Multivariate Time Series Classification

IJCAI 2022poster

Multivariate time series (MTS) classification is a challenging and important task in various domains and real-world applications. Much of prior work on MTS can be roughly divided into neural network (NN)- and pattern-based methods. The former can lead to robust classification performance, but many o…

2022

Gradient Importance Learning for Incomplete Observations

ICLR 2022poster

Though recent works have developed methods that can generate estimates (or imputations) of the missing entries in a dataset to facilitate downstream analysis, most depend on assumptions that may not align with real-world applications and could suffer from poor performance in subsequent tasks such as…

2020

Deep Imitative Reinforcement Learning for Temporal Logic Robot Motion Planning with Noisy Semantic Observations

ICRA 2020poster

In this paper, we propose a Deep Imitative Q-learning (DIQL) method to synthesize control policies for mobile robots that need to satisfy Linear Temporal Logic (LTL) specifications using noisy semantic observations of their surroundings. The robot sensing error is modeled using probabilistic labels…

Cited by 9SourceScholar