← Search

Xuyang Chen

11 accepted papers

2026

Finite-Time Analysis of Actor-Critic Methods with Deep Neural Network Approximation

ICLR 2026poster

Actor–critic (AC) algorithms underpin many of today’s most successful reinforcement learning (RL) applications, yet their finite-time convergence in realistic settings remains largely underexplored. Existing analyses often rely on oversimplified formulations and are largely confined to linear functi…

Cited by 0SourceScholar
2026

VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning

ICML 2026poster

Offline reinforcement learning (RL) learns effective policies from pre-collected datasets, offering a practical solution for applications where online interactions are risky or costly. Model-based approaches are particularly advantageous for offline RL, owing to their data efficiency and generalizab…

Cited by 0SourceScholar
2025

BRIGHT-VO: Brightness-Guided Hybrid Transformer for Visual Odometry with Multi-modality Refinement Module

IJCAI 2025

Visual odometry (VO) plays a crucial role in autonomous driving, robotic navigation, and other related tasks by estimating the position and orientation of a camera based on visual input. Significant progress has been made in data-driven VO methods, particularly those leveraging deep learning techniq

2024

Box2Poly: Memory-Efficient Polygon Prediction of Arbitrarily Shaped and Rotated Text

AAAI 2024technical

Recently, Transformer-based text detection techniques have sought to predict polygons by encoding the coordinates of individual boundary vertices using distinct query features. However, this approach incurs a significant memory overhead and struggles to effectively capture the intricate relationship…

2024

Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator

IJCAI 2024poster

Actor-critic methods have achieved state-of-the-art performance in various challenging tasks. However, theoretical understandings of their performance remain elusive and challenging. Existing studies mostly focus on practically uncommon variants such as double-loop or two-timescale stepsize actor-cr…

Cited by 1SourcePDFScholar
2023

Efficient Q-Learning over Visit Frequency Maps for Multi-Agent Exploration of Unknown Environments

IROS 2023poster

The robot exploration task has been widely studied with applications spanning from novel environment mapping to item delivery. For some time-critical tasks, such as rescue catastrophes, the agent is required to explore as efficiently as possible. Recently, Visit Frequency-based map representation ac…

Cited by 5SourceScholar
2023

Global Convergence of Two-Timescale Actor-Critic for Solving Linear Quadratic Regulator

AAAI 2023technical

The actor-critic (AC) reinforcement learning algorithms have been the powerhouse behind many challenging applications. Nevertheless, its convergence is fragile in general. To study its instability, existing works mostly consider the uncommon double-loop variant or basic models with finite state and…

Cited by 12SourcePDFScholar