← Search

Long Yang

21 accepted papers

2026

DriveAgent: Multi-Agent Structured Reasoning with LLM and Multimodal Sensor Fusion for Autonomous Driving

ICRA 2026poster

We introduce DriveAgent, a modular multi-agent autonomous driving framework that leverages large language model (LLM) reasoning combined with multimodal sensor fusion for autonomous driving. DriveAgent orchestrates specialized agents operating on camera, Light Detection and Ranging (LiDAR), Inertial…

2026

GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation

CVPR 2026

Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require precise 3D reasoning. We propose GeoPredict, a geometry-aware VLA framework that augments a continuous-action policy with

Cited by 0SourcecodeScholar
2025

DriveAgent: Multi-Agent Structured Reasoning With LLM and Multimodal Sensor Fusion for Autonomous Driving

RA-L 2025

We introduce DriveAgent, a modular multi-agent autonomous driving framework that leverages large language model (LLM) reasoning combined with multimodal sensor fusion for autonomous driving. DriveAgent orchestrates specialized agents operating on camera, Light Detection and Ranging (LiDAR), Inertial

Cited by 16SourceScholar
2025

Low-Dimension-to-High-Dimension Generalization and Its Implications for Length Generalization

ICML 2025poster

Low-Dimension-to-High-Dimension (LDHD) generalization, a subset of Out-of-Distribution (OOD) generalization, involves training on a low-dimensional subspace and testing in a high-dimensional space. Assuming instances are generated from latent variables reflecting problem scale, LDHD generalization c…

Cited by 1SourcePDFScholar
2025

MSE-based Sampling of Bandlimited Product Graph Signals via Joint Low-pass Impulse Responses

ICASSP 2025accepted

Matrix graph signals, which are associated with two factor graphs, are ubiquitous in daily life, such as time-varying physical signals in sensor networks and rating matrices in recommendation systems. In practice, due to the row-wise and column-wise smoothness, they are modeled as bandlimited (BL) g…

Cited by 0SourceScholar
2025

SimLauncher: Launching Sample-Efficient Real-World Robotic Reinforcement Learning via Simulation Pre-Training

IROS 2025

Autonomous learning of dexterous, long-horizon robotic skills has been a longstanding pursuit of embodied AI. Recent advances in robotic reinforcement learning (RL) have demonstrated remarkable performance and robustness in real-world visuomotor control tasks. However, applying RL in the real world

Cited by 3SourceScholar
2025

UniTac2Pose: A Unified Approach Learned in Simulation for Category-level Visuotactile In-hand Pose Estimation

CoRL 2025poster

Accurate estimation of the in-hand pose of an object based on its CAD model is crucial in both industrial applications and everyday tasks—ranging from positioning workpieces and assembling components to seamlessly inserting devices like USB connectors. While existing methods often rely on regression…

Cited by 0SourceScholar
2024

FlagVNE: A Flexible and Generalizable Reinforcement Learning Framework for Network Resource Allocation

IJCAI 2024poster

Virtual network embedding (VNE) is an essential resource allocation task in network virtualization, aiming to map virtual network requests (VNRs) onto physical infrastructure. Reinforcement learning (RL) has recently emerged as a promising solution to this problem. However, existing RL-based VNE met…

2024

LVDiffusor: Distilling Functional Rearrangement Priors From Large Models Into Diffusor

RA-L 2024

Object rearrangement, a fundamental challenge in robotics, demands versatile strategies to handle diverse objects, configurations, and functional needs. To achieve this, the AI robot needs to learn functional rearrangement priors to specify precise goals that meet the functional requirements. Previo

Cited by 12SourceScholar
2024

Langevin Policy for Safe Reinforcement Learning

ICML 2024poster

Optimization and sampling based algorithms are two branches of methods in machine learning, while existing safe reinforcement learning (RL) algorithms are mainly based on optimization, it is still unclear whether sampling based methods can lead to desirable performance with safe policy. This paper f…

Cited by 1SourcePDFScholar
2024

Optimizing over Multiple Distributions under Generalized Quasar-Convexity Condition

NeurIPS 2024poster

We study a typical optimization model where the optimization variable is composed of multiple probability distributions. Though the model appears frequently in practice, such as for policy problems, it lacks specific analysis in the general setting. For this optimization problem, we propose a new s…

Cited by 0SourcePDFScholar
2023

A Semi-Automatic Oriental Ink Painting Framework for Robotic Drawing From 3D Models

RA-L 2023

Creating visually pleasing stylized ink paintings from 3D models is a challenge in robotic manipulation. We propose a semi-automatic framework that can extract expressive strokes from 3D models and draw them in oriental ink painting styles by using a robotic arm. The framework consists of a simulati

Cited by 2SourceScholar
2023

Augmented Proximal Policy Optimization for Safe Reinforcement Learning

AAAI 2023technical

Safe reinforcement learning considers practical scenarios that maximize the return while satisfying safety constraints. Current algorithms, which suffer from training oscillations or approximation errors, still struggle to update the policy efficiently with precise constraint satisfaction. In this a…

Cited by 21SourcePDFScholar
2023

VOCE: Variational Optimization with Conservative Estimation for Offline Safe Reinforcement Learning

NeurIPS 2023poster

Offline safe reinforcement learning (RL) algorithms promise to learn policies that satisfy safety constraints directly in offline datasets without interacting with the environment. This arrangement is particularly important in scenarios with high sampling costs and potential dangers, such as autonom…

2022

Constrained Update Projection Approach to Safe Policy Optimization

NeurIPS 2022accept

Safe reinforcement learning (RL) studies problems where an intelligent agent has to not only maximize reward but also avoid exploring unsafe areas. In this study, we propose CUP, a novel policy optimization method based on Constrained Update Projection framework that enjoys rigorous safety guarantee…

2022

Penalized Proximal Policy Optimization for Safe Reinforcement Learning

IJCAI 2022poster

Safe reinforcement learning aims to learn the optimal policy while satisfying safety constraints, which is essential in real-world applications. However, current algorithms still struggle for efficient policy updates with hard constraint satisfaction. In this paper, we propose Penalized Proximal Pol…

2022

Policy Optimization with Stochastic Mirror Descent

AAAI 2022technical

Improving sample efficiency has been a longstanding goal in reinforcement learning. This paper proposes VRMPO algorithm: a sample efficient policy gradient method with stochastic mirror descent. In VRMPO, a novel variance-reduced policy gradient estimator is presented to improve sample efficiency. W…

Cited by 40SourcePDFScholar
2021

On Convergence of Gradient Expected Sarsa(λ)

AAAI 2021technical

We study the convergence of Expected Sarsa(λ) with function approximation. We show that with off-line es- timate (multi-step bootstrapping) to ExpectedSarsa(λ) is unstable for off-policy learning. Furthermore, based on convex-concave saddle-point framework, we propose a con- vergent Gradient Expecte…

Cited by 4SourcePDFScholar
2018

Texture Mapping for 3D Reconstruction With RGB-D Sensor

CVPR 2018poster

Acquiring realistic texture details for 3D models is important in 3D reconstruction. However, the existence of geometric errors, caused by noisy RGB-D sensor data, always makes the color images cannot be accurately aligned onto reconstructed 3D models. In this paper, we propose a global-to-local cor…

Cited by 102SourcePDFScholar
2017

Distinguishing the Indistinguishable: Exploring Structural Ambiguities via Geodesic Context

CVPR 2017spotlight

A perennial problem in structure from motion (SfM) is visual ambiguity posed by repetitive structures. Recent disambiguating algorithms infer ambiguities mainly via explicit background context, thus face limitations in highly ambiguous scenes which are visually indistinguishable. Instead of analyzin…

Cited by 38PDFcodeScholar