← Search

Guang Tan

17 accepted papers

2026

COVR: Collaborative Optimization of VLMs and RL Agent for Visual-Based Control

AAAI 2026technical

Visual reinforcement learning (RL) suffers from poor sample efficiency due to high-dimensional observations in complex tasks. While existing works have shown that vision-language models (VLMs) can assist RL, they often focus on knowledge distillation from the VLM to RL, overlooking the potential of

Cited by 0SourcePDFScholar
2026

DynamicTree: Interactive Real Tree Animation via Sparse Voxel Spectrum

CVPR 2026

Generating dynamic and interactive 3D trees has wide applications in virtual reality, games, and world simulation. However, existing methods still face various challenges in generating structurally consistent and realistic 4D motion for complex real trees. In this paper, we propose DynamicTree, the

Cited by 0SourcecodeScholar
2026

GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature Space

ICLR 2026poster

In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data. A key challenge lies in handling heterogeneous features from agents equipped with different sensing modalities or model architectures, which complicates data fusion.…

Cited by 0SourcecodeScholar
2026

HouseTune: Two-Stage Floorplan Generation with LLM Assistance

AAAI 2026technical

This paper proposes a two-stage text-to-floorplan generation framework that combines the reasoning capability of Large Language Models (LLMs) with the generative power of diffusion models. In the first stage, we leverage a Chain-of-Thought (CoT) prompting strategy to guide an LLM in generating an in

Cited by 0SourcePDFScholar
2026

Resolving the Stability-Plasticity Dilemma in Reinforcement Learning via Complementary Continual Critics

CVPR 2026

This paper proposes the Continual Dual-Critic with Cross-Attention (CD-CCA) framework for visual reinforcement learning to address the plasticity-stability conflict. Our method introduces continual learning techniques into the visual RL architecture, constructing two complementary critics using Cont

Cited by 0SourcecodeScholar
2025

BEVSync: Asynchronous Data Alignment for Camera-based Vehicle-Infrastructure Cooperative Perception Under Uncertain Delays

AAAI 2025technical

Vehicle-to-infrastructure (V2I) cooperative perception systems can enhance the sensing abilities of autonomous vehicles. Existing V2I solutions often consider LiDARs devices instead of cameras, the most prevalent sensors with low cost and wide installation. In addition, a major challenge that has be…

Cited by 0SourcePDFScholar
2025

Exploiting Continuous Motion Clues for Vision-Based Occupancy Prediction

AAAI 2025technical

Occupancy networks aim to reconstruct the surroundings with occupied semantic voxels. However, frequent object occlusions often occur in dynamic real-world scenarios, which cannot be captured by independent frames. Most existing occupancy networks generate results without explicitly considering past…

2025

VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement Learning

CVPR 2025poster

Vision-based Reinforcement Learning (VRL) attempts to establish associations between visual inputs and optimal actions through interactions with the environment. Given the high-dimensional and complex nature of visual data, it becomes essential to learn policy upon high-quality state representation.…

Cited by 0SourcePDFScholar
2024

DMR: Decomposed Multi-Modality Representations for Frames and Events Fusion in Visual Reinforcement Learning

CVPR 2024poster

We explore visual reinforcement learning (RL) using two complementary visual modalities: frame-based RGB camera and event-based Dynamic Vision Sensor (DVS). Existing multi-modality visual RL methods often encounter challenges in effectively extracting task-relevant information from multiple modaliti…

2024

Density-Adaptive Model Based on Motif Matrix for Multi-Agent Trajectory Prediction

CVPR 2024poster

Multi-agent trajectory prediction is essential in autonomous driving risk avoidance and traffic flow control. However the heterogeneous traffic density on interactions which caused by physical laws social norms and so on is often overlooked in existing methods. When the density varies the number of…

Cited by 1SourcePDFScholar
2024

InterCoop: Spatio-Temporal Interaction Aware Cooperative Perception for Networked Vehicles

ICRA 2024poster

In autonomous driving, cooperative perception through vehicle-to-vehicle (V2V) communication is considered crucial for enhancing traffic safety and efficiency. However, existing methods often simplify the handling of perception data from multiple vehicles. In these approaches, the egovehicle aggrega…

Cited by 2SourceScholar
2024

Partitioning Message Passing for Graph Fraud Detection

ICLR 2024poster

Label imbalance and homophily-heterophily mixture are the fundamental problems encountered when applying Graph Neural Networks (GNNs) to Graph Fraud Detection (GFD) tasks. Existing GNN-based GFD models are designed to augment graph structure to accommodate the inductive bias of GNNs towards homophil…

Cited by 30SourcePDFScholar
2024

Resolving Loop Closure Confusion in Repetitive Environments for Visual SLAM through AI Foundation Models Assistance

ICRA 2024poster

In visual SLAM (VSLAM) systems, loop closure plays a crucial role in reducing accumulated errors. However, VSLAM systems relying on low-level visual features often suffer from the problem of perceptual confusion in repetitive environments, where scenes in different locations are incorrectly identifi…

Cited by 5SourceScholar