← Search

Justin Fu

19 accepted papers

2026

MAGNIFIED: RL Fine-Tuning of Multimodal Large Language Models for Motion Planning

ICRA 2026poster

Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in semantic understanding and common sense reasoning, making them promising candidates for solving planning problems in autonomous driving. However, the next-token text prediction objectives traditionally used in pre…

2024

Improving Agent Behaviors with RL Fine-tuning for Autonomous Driving

ECCV 2024poster

"A major challenge in autonomous vehicle research is modeling agent behaviors, which has critical applications including constructing realistic and reliable simulations for off-board evaluation and forecasting traffic agents motion for onboard planning. While supervised learning has shown success in…

2023

Imitation Is Not Enough: Robustifying Imitation with Reinforcement Learning for Challenging Driving Scenarios

IROS 2023poster

Imitation learning (IL) is a simple and powerful way to use high-quality human driving data, which can be collected at scale, to produce human-like behavior. However, policies based on imitation learning alone often fail to sufficiently account for safety and reliability concerns. In this paper, we…

Cited by 106SourceScholar
2023

Waymax: An Accelerated, Data-Driven Simulator for Large-Scale Autonomous Driving Research

NeurIPS 2023poster

Simulation is an essential tool to develop and benchmark autonomous vehicle planning software in a safe and cost-effective manner. However, realistic simulation requires accurate modeling of multi-agent interactive behaviors to be trustworthy, behaviors which can be highly nuanced and complex. To ad…

Cited by 116SourcePDFScholar
2022

CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning

NAACL 2022long

Conventionally, generation of natural language for dialogue agents may be viewed as a statistical learning problem: determine the patterns in human-provided data and generate appropriate responses with similar statistical properties. However, dialogue can also be regarded as a goal directed process,…

2022

Context-Aware Language Modeling for Goal-Oriented Dialogue Systems

NAACL 2022findings

Goal-oriented dialogue systems face a trade-off between fluent language generation and task-specific control. While supervised learning with large language models is capable of producing realistic text, how to steer such responses towards completing a specific task without sacrificing language quali…

Cited by 27SourcePDFScholar
2022

Hierarchical Model-Based Imitation Learning for Planning in Autonomous Driving

IROS 2022poster

We demonstrate the first large-scale application of model-based generative adversarial imitation learning (MGAIL) to the task of dense urban self-driving. We augment standard MGAIL using a hierarchical model to enable generalization to arbitrary goal routes, and measure performance using a closed-lo…

Cited by 60SourceScholar
2021

Benchmarks for Deep Off-Policy Evaluation

ICLR 2021poster

Off-policy evaluation (OPE) holds the promise of being able to leverage large, offline datasets for both evaluating and selecting complex policies for decision making. The ability to learn offline is particularly important in many real-world domains, such as in healthcare, recommender systems, or ro…

2021

Learning to Reach Goals via Iterated Supervised Learning

ICLR 2021oral

Current reinforcement learning (RL) algorithms can be brittle and difficult to use, especially when learning goal-reaching behaviors from sparse rewards. Although supervised imitation learning provides a simple and stable alternative, it requires access to demonstrations from a human supervisor. In…

2019

From Language to Goals: Inverse Reinforcement Learning for Vision-Based Instruction Following

ICLR 2019poster

Reinforcement learning is a promising framework for solving control problems, but its use in practical situations is hampered by the fact that reward functions are often difficult to engineer. Specifying goals and tasks for autonomous machines, such as robots, is a significant challenge: conventiona…

Cited by 153SourcePDFScholar
2019

Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction

NeurIPS 2019poster

Off-policy reinforcement learning aims to leverage experience collected from prior policies for sample-efficient learning. However, in practice, commonly used off-policy approximate dynamic programming methods based on Q-learning and actor-critic methods are highly sensitive to the data distribution…

Cited by 1291SourcePDFScholar
2019

When to Trust Your Model: Model-Based Policy Optimization

NeurIPS 2019poster

Designing effective model-based reinforcement learning algorithms is difficult because the ease of data generation must be weighed against the bias of model-generated data. In this paper, we study the role of model usage in policy optimization both theoretically and empirically. We first formulate a…

2018

Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition

NeurIPS 2018poster

The design of a reward function often poses a major practical challenge to real-world applications of reinforcement learning. Approaches such as inverse reinforcement learning attempt to overcome this challenge, but require expert demonstrations, which can be difficult or expensive to obtain in prac…

Cited by 155SourcePDFScholar
2017

EX2: Exploration with Exemplar Models for Deep Reinforcement Learning

NeurIPS 2017spotlight

Deep reinforcement learning algorithms have been shown to learn complex tasks using highly general policy classes. However, sparse reward problems remain a significant challenge. Exploration methods based on novelty detection have been particularly successful in such settings but typically require g…

Cited by 197SourcePDFScholar
2017

Generalizing Skills with Semi-Supervised Reinforcement Learning

ICLR 2017poster

Deep reinforcement learning (RL) can acquire complex behaviors from low-level inputs, such as images. However, real-world applications of such methods require generalizing to the vast variability of the real world. Deep networks are known to achieve remarkable generalization when provided with massi…

Cited by 93SourceScholar
2016

One-shot learning of manipulation skills with online dynamics adaptation and neural network priors

IROS 2016poster

One of the key challenges in applying reinforcement learning to complex robotic control tasks is the need to gather large amounts of experience in order to find an effective policy for the task at hand. Model-based reinforcement learning can achieve good sample efficiency, but requires the ability t…

Cited by 181SourceScholar