← Search

Jun Luo

47 accepted papers

2026

Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models

AAAI 2026technical

Vision-Language Continual Learning (VLCL) has attracted significant research attention for its robust capabilities, and the adoption of Parameter-Efficient Fine-Tuning (PEFT) strategies is enabling these models to achieve competitive performance with substantially reduced resource consumption. Howev

Cited by 0SourcePDFScholar
2026

Cheating Stereo Matching in Full-Scale: Physical Adversarial Attack Against Binocular Depth Estimation in Autonomous Driving

AAAI 2026technical

Though deep neural models adopted to realize the perception of autonomous driving have proven vulnerable to adversarial examples, known attacks often leverage 2D patches and target mostly monocular perception. Therefore, the effectiveness of Physical Adversarial Examples (PAEs) on stereo-based binoc

Cited by 0SourcePDFScholar
2026

Decoupled Heuristic Multi-Vehicle Emergency Trajectory Planning for Sudden Obstacles

RA-L 2026

The emergence of sudden obstacles can significantly reduce the feasible space and may induce locally non-convex or fragmented space, especially in densely clustered scenarios, making vehicle trajectory planning remarkably challenging. Current methods face computational bottlenecks when generating em

Cited by 0SourceScholar
2026

PMSPO: Progressive Matching and Semantic-Aware Policy Optimization for Camouflaged Object Detection

ICML 2026poster

Reinforcement learning-based Multimodal Large Language Models (MLLMs) provide new perspectives for visual grounding, yet face significant challenges in Camouflaged Object Detection (COD) where objects blend seamlessly with backgrounds. This stems primarily from: difficulties in multi-object matching…

Cited by 0SourceScholar
2025

An Evaluation Framework for Product Images Background Inpainting Based on Human Feedback and Product Consistency

AAAI 2025technical

In product advertising applications, the automated inpainting of backgrounds utilizing AI techniques in product images has emerged as a significant task. However, the techniques still suffer from issues such as inappropriate background and inconsistent product in generated product images, and existi…

2025

Bio-inspired Shape Self-Assembly in Large-Scale Swarm Robots Under Information Asymmetry *

IROS 2025

This study investigates the problem of large-scale swarm robots shape self-assembly problem under conditions of information asymmetry. Existing methods assume complete sharing of global information; however, this assumption has significant limitations in terms of resource consumption and swarm emerg

Cited by 0SourceScholar
2025

Co-Reinforcement Learning for Unified Multimodal Understanding and Generation

NeurIPS 2025spotlight

This paper presents a pioneering exploration of reinforcement learning (RL) via group relative policy optimization for unified multimodal large language models (ULMs), aimed at simultaneously reinforcing generation and understanding capabilities. Through systematic pilot studies, we uncover the sign…

Cited by 0SourcecodeScholar
2025

Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning

ICCV 2025poster

Recent advancements in multimodal large language models (MLLMs) have demonstrated exceptional performance in multimodal perception and understanding. However, leading open-source MLLMs exhibit significant limitations in complex and structured reasoning, particularly in tasks requiring deep reasoning…

2025

Do Not DeepFake Me: Privacy-Preserving Neural 3D Head Reconstruction Without Sensitive Images

AAAI 2025technical

While 3D head reconstruction is widely used for modeling, existing neural reconstruction approaches rely on high-resolution multi-view images, posing notable privacy issues. Individuals are particularly sensitive to facial features, and facial image leakage can lead to many malicious activities, suc…

Cited by 0SourcePDFScholar
2025

GeoPro-Net: Learning Interpretable Spatiotemporal Prediction Models Through Statistically-Guided Geo-Prototyping

AAAI 2025technical

The problem of forecasting spatiotemporal events such as crimes and accidents is crucial to public safety and city management. Besides accuracy, interpretability is also a key requirement for spatiotemporal forecasting models to justify the decisions. Merely presenting predicted scores fails to conv…

2025

LOPT: Learning Optimal Pigovian Tax in Sequential Social Dilemmas

NeurIPS 2025poster

Multi-agent reinforcement learning (MARL) has emerged as a powerful framework for modeling autonomous agents that independently optimize their individual objectives. However, in mixed-motive MARL environments, rational self-interested behaviors often lead to collectively suboptimal outcomes situatio…

Cited by 0SourceScholar
2025

MaSS13K: A Matting-level Semantic Segmentation Benchmark

CVPR 2025poster

High-resolution semantic segmentation is essential for applications such as image editing, bokeh imaging, AR/VR, etc. Unfortunately, existing datasets often have limited resolution and lack precise mask details and boundaries. In this work, we build a large-scale, matting-level semantic segmentation…

2025

Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models

ICLR 2025poster

Federated prompt learning benefits federated learning with CLIP-like Vision-Language Model's (VLM's) robust representation learning ability through prompt learning. However, current federated prompt learning methods are habitually restricted to the traditional FL paradigm, where the participating cl…

2025

On the Performance Analysis of Momentum Method: A Frequency Domain Perspective

ICLR 2025poster

Momentum-based optimizers are widely adopted for training neural networks. However, the optimal selection of momentum coefficients remains elusive. This uncertainty impedes a clear understanding of the role of momentum in stochastic gradient methods. In this paper, we present a frequency domain anal…

Cited by 0SourcePDFScholar
2025

Reinforcement Learning-Based Energy-Efficient and Obstacle-Free Path Planning for Magnetic Microrobots in Dynamic Environments

IROS 2025

Online path planning for magnetic microrobots actuated by electromagnetic system in dynamic flow field presents significant challenges due to time-varying fluid dynamics, energy constraints, and collision risks. Traditional path planning approaches, which often rely on static flow assumptions or sim

Cited by 0SourceScholar
2025

Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents

NeurIPS 2025poster

Large Language Models (LLMs) are increasingly deployed as autonomous agents in multi-agent systems, and promising coordination has been demonstrated in handling complex tasks under predefined roles and scripted workflows. However, significant challenges remain in open-ended environments, where agent…

Cited by 0SourceScholar
2025

Zero Shot Domain Adaptive Semantic Segmentation by Synthetic Data Generation and Progressive Adaptation

IROS 2025

Deep learning-based semantic segmentation models achieve impressive results yet remain limited in handling distribution shifts between training and test data. In this paper, we present SDGPA (Synthetic Data Generation and Progressive Adaptation), a novel method that tackles zero-shot domain adaptive

Cited by 3SourcecodeScholar
2024

A Framework for Real-time Generation of Multi-directional Traversability Maps in Unstructured Environments

ICRA 2024poster

In complex unstructured environments, accurate terrain traversability analysis is a fundamental requirement for the successful execution of any movements of ground robots, especially given that terrain traversability often exhibits anisotropy. However, the difficulty in obtaining multi-directional t…

Cited by 1SourceScholar
2024

A Rigid-Flexible-Soft Coupled Dexterous Hand With Sliding Tactile Perception and Feedback

RA-L 2024

The human hand is capable of executing a wide range of complex movements due to its biomechanical structure and skin sensing system. Designing an anthropomorphic hand that mimics the biomechanical structure of the human and incorporates skin sensing, presents a long-term challenge in the field of ro

Cited by 8SourceScholar
2024

BE-SLAM: BEV-Enhanced Dynamic Semantic SLAM with Static Object Reconstruction

IROS 2024poster

The quality of a robot’s environmental perception determines whether it can achieve more intelligent applications, such as semantic interaction with humans. SLAM, on the other hand, is one of the crucial capabilities for a robot to perceive its environment. However, when only a monocular image is pr…

Cited by 0SourceScholar
2023

Ask Language Model to Clean Your Noisy Translation Data

EMNLP 2023long findings

TTransformer models have demonstrated remarkable performance in neural machine translation (NMT). However, their vulnerability to noisy input poses a significant challenge in practical implementation, where generating clean output from noisy input is crucial. The MTNT dataset is widely used as a ben…

Cited by 0SourceScholar
2023

Dynamic Decision Frequency with Continuous Options

IROS 2023poster

In classic reinforcement learning algorithms, agents make decisions at discrete and fixed time intervals. The duration between decisions becomes a crucial hyperparameter, as setting it too short may increase the problem's difficulty by requiring the agent to make numerous decisions to achieve its go…

Cited by 9SourcecodeScholar
2023

FedPerfix: Towards Partial Model Personalization of Vision Transformers in Federated Learning

ICCV 2023poster

Personalized Federated Learning (PFL) represents a promising solution for decentralized learning in heterogeneous data environments. Partial model personalization has been proposed to improve the efficiency of PFL by selectively updating local model parameters instead of aggregating all of them. How…

Cited by 22PDFcodeScholar
2023

OCHID-Fi: Occlusion-Robust Hand Pose Estimation in 3D via RF-Vision

ICCV 2023poster

Hand Pose Estimation (HPE) is crucial to many applications, but conventional cameras-based CM-HPE methods are completely subject to Line-of-Sight (LoS), as cameras cannot capture occluded objects. In this paper, we propose to exploit Radio-Frequency-Vision (RF-vision) capable of bypassing obstacles…

Cited by 7PDFcodeScholar
2023

PGFed: Personalize Each Client's Global Objective for Federated Learning

ICCV 2023oral

Personalized federated learning has received an upsurge of attention due to the mediocre performance of conventional federated learning (FL) over heterogeneous data. Unlike conventional FL which trains a single global consensus model, personalized FL allows different models for different clients. Ho…

Cited by 13PDFcodeScholar
2022

A Simple Decentralized Cross-Entropy Method

NeurIPS 2022accept

Cross-Entropy Method (CEM) is commonly used for planning in model-based reinforcement learning (MBRL) where a centralized approach is typically utilized to update the sampling distribution based on only the top-$k$ operation's results on samples. In this paper, we show that such a centralized approa…

2022

Multiagent Q-learning with Sub-Team Coordination

NeurIPS 2022accept

In many real-world cooperative multiagent reinforcement learning (MARL) tasks, teams of agents can rehearse together before deployment, but then communication constraints may force individual agents to execute independently when deployed. Centralized training and decentralized execution (CTDE) is in…

Cited by 10SourcePDFScholar
2022

Offline Learning of Counterfactual Predictions for Real-World Robotic Reinforcement Learning

ICRA 2022poster

We consider real-world reinforcement learning (RL) of robotic manipulation tasks that involve both visuomotor skills and contact-rich skills. We aim to train a policy that maps multimodal sensory observations (vision and force) to a manipulator's joint velocities under practical considerations. We p…

Cited by 7SourceScholar
2022

Understanding and mitigating the limitations of prioritized experience replay

UAI 2022poster

Prioritized Experience Replay (ER) has been empirically shown to improve sample efficiency across many domains and attracted great attention; however, there is little theoretical understanding of why such prioritized sampling helps and its limitations. In this work, we take a deep look at the priori…

Cited by 25SourcePDFScholar
2021

Design and Experimental Evaluation of a Multi-Mode Mobile Robot Based on Eccentric Paddle Mechanism

RA-L 2021

In order to improve climbing performance of the robot in complex environments, the wheel-legged mechanism is gradually being widely used. In this letter, we simplify the previous eccentric paddle mechanism (ePaddle), and propose a 2-DOF ePaddle-based robot which integrates stability and maneuverabil

Cited by 4SourceScholar
2021

Graph-SIM: A Graph-based Spatiotemporal Interaction Modelling for Pedestrian Action Prediction

ICRA 2021poster

One of the most crucial yet challenging tasks for autonomous vehicles in urban environments is predicting the future behaviour of nearby pedestrians, especially at points of crossing. Predicting behaviour depends on many social and environmental factors, particularly interactions between road users.…

Cited by 27SourcecodeScholar
2021

Learning robust driving policies without online exploration

ICRA 2021poster

We propose a multi-time-scale predictive representation learning method to efficiently learn robust driving policies in an offline manner that generalize well to novel road geometries, and damaged and distracting lane conditions which are not covered in the offline training data. We show that our pr…

Cited by 2SourceScholar
2021

Prediction by Anticipation: An Action-Conditional Prediction Method Based on Interaction Learning

ICCV 2021poster

In autonomous driving (AD), accurately predicting changes in the environment can effectively improve safety and comfort. Due to complex interactions among traffic participants, however, it is very hard to achieve accurate prediction for a long horizon. To address this challenge, we propose predictio…

Cited by 4PDFcodeScholar
2021

Self-Supervised Simultaneous Multi-Step Prediction of Road Dynamics and Cost Map

CVPR 2021poster

In this paper we propose a system consisting of a modular network and a trajectory planner. The network simultaneously predicts Occupancy Grid Maps (OGMs) and estimates space-time cost maps (CMs) corresponding to the areas around the vehicle. The trajectory planner computes the cost of a set of pred…

Cited by 4PDFcodeScholar
2020

Adaptive Hierarchical Down-Sampling for Point Cloud Classification

CVPR 2020poster

Deterministic down-sampling of an unordered point cloud in a deep neural network has not been rigorously studied so far. Existing methods down-sample the points regardless of their importance for the network output and often address down-sampling the raw point cloud before processing. As a result, s…

Cited by 172PDFScholar
2020

SMARTS: An Open-Source Scalable Multi-Agent RL Training School for Autonomous Driving

CoRL 2020

Interaction is fundamental in autonomous driving (AD). Despite more than a decade of intensive R&D in AD, how to dynamically interact with diverse road users in various contexts still remains unsolved. Multi-agent learning has recently seen big breakthroughs and has much to offer towards solving rea

2016

An automated system for investigating sperm orientation in fluid flow

ICRA 2016

Mammalian sperms reorient against fluid flow in the female reproductive tract, known as rheotaxis. Compared to chemotaxis that provides short-distance guidance, rheotaxis provides long-distance guidance for a sperm to find the egg cell. However, only a low number of sperms are capable of rheotaxis a

Cited by 6SourceScholar
2016

Studying of rectilinear locomotion for a two-segment system with anisotropic dry friction model

IROS 2016poster

This paper contributes to the understanding of the fundamental properties of rectilinearly locomotion of an one-dimensional system travelling on the horizontal plane, where dry Coulomb friction acting between it and surface. We discuss an approximate steady-state motion on a simplified two-segment s…

Cited by 0SourceScholar
2015

Distortion invariant joint-feature for visual tracking in catadioptric omnidirectional vision

ICRA 2015poster

Central catadioptric omnidirectional images exhibit serious nonlinear distortions due to quadratic mirrors involved. Conventional visual features developed based on the perspective model are hard to achieve a satisfactory performance when directly applied to the distorted omnidirectional image. This…

Cited by 7SourceScholar