← Search

Kai Yan

11 accepted papers

2025

MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?

NeurIPS 2025poster

The ability to recognize patterns from examples and apply them to new ones is a primal ability for general intelligence, and is widely studied by psychology and AI researchers. Many benchmarks have been proposed to measure such ability for Large Language Models (LLMs); however, they focus on few-sho…

Cited by 0SourceScholar
2025

Non-Autoregressive Multimodal Machine Translation

ICASSP 2025accepted

Performing better text translation by integrating auxiliary inputs from visual information has gained widespread attention in recent years. While existing methods outperform the text-only translation models, the step-by-step generative style reduces the inference speed, which limits their applicabil…

Cited by 0SourceScholar
2024

Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models

ICML 2024poster

While language models (LMs) have shown potential across a range of decision-making tasks, their reliance on simple acting processes limits their broad deployment as autonomous agents. In this paper, we introduce Language Agent Tree Search (LATS) -- the first general framework that synergizes the cap…

2024

Offline Imitation from Observation via Primal Wasserstein State Occupancy Matching

ICML 2024poster

In real-world scenarios, arbitrary interactions with the environment can often be costly, and actions of expert demonstrations are not always available. To reduce the need for both, offline Learning from Observations (LfO) is extensively studied: the agent learns to solve a task given only expert st…

2024

Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision Transformers

NeurIPS 2024spotlight

Decision Transformers have recently emerged as a new and compelling paradigm for offline Reinforcement Learning (RL), completing a trajectory in an autoregressive way. While improvements have been made to overcome initial shortcomings, online finetuning of decision transformers has been surprisingl…

2023

A Simple Solution for Offline Imitation from Observations and Examples with Possibly Incomplete Trajectories

NeurIPS 2023poster

Offline imitation from observations aims to solve MDPs where only task-specific expert states and task-agnostic non-expert state-action pairs are available. Offline imitation is useful in real-world scenarios where arbitrary interactions are costly and expert actions are unavailable. The state-of-th…

2023

Global Posture Stabilization for the Kinematic Model of a Rear-Axle Driven Car-Like Mobile Robot Considering Obstacle Avoidance

RA-L 2023

The car-like mobile robot (CLMR) is a kind of wheeled robot that has many applications in industry, transportation, security, etc. The CLMR considered in this study has rear-wheel driving. The robot is subject to nonholonomic constraints, and the robotic system is classified as a nonlinear and under

Cited by 8SourceScholar
2023

Neural-PBIR Reconstruction of Shape, Material, and Illumination

ICCV 2023poster

Reconstructing the shape and spatially varying surface appearances of a physical-world object as well as its surrounding illumination based on 2D images (e.g., photographs) of the object has been a long-standing problem in computer vision and graphics. In this paper, we introduce an accurate and hig…

Cited by 30PDFcodeScholar
2022

CEIP: Combining Explicit and Implicit Priors for Reinforcement Learning with Demonstrations

NeurIPS 2022accept

Although reinforcement learning has found widespread use in dense reward settings, training autonomous agents with sparse rewards remains challenging. To address this difficulty, prior work has shown promising results when using not only task-specific demonstrations but also task-agnostic albeit som…

2021

A Surrogate Objective Framework for Prediction+Programming with Soft Constraints

NeurIPS 2021poster

Prediction+optimization is a common real-world paradigm where we have to predict problem parameters before solving the optimization problem. However, the criteria by which the prediction model is trained are often inconsistent with the goal of the downstream optimization problem. Recently, decision…

Cited by 7SourcePDFScholar