← Search

Xudong Zhang

23 accepted papers

2026

Aligning Cross-View Visual Geometries in LVLMs Through Human-Like Reasoning Learning

AAAI 2026technical

Spatial understanding is a critical capability for LVLMs (Large Vision-Language Models) to advance embodied AI applications. Existing works primarily focus on enhancing spatial understanding within a single frame, i.e., injecting 3D spatial concepts into LVLMs under single coordinate system. However

Cited by 0SourcePDFScholar
2026

Efficient Reinforcement Learning for Zero-Shot Coordination in Evolving Games

AAAI 2026technical

Zero-shot coordination(ZSC), a key challenge in multi-agent game theory, has become a hot topic in reinforcement learning (RL) research recently, especially in complex evolving games. It focuses on the generalization ability of agents, requiring them to coordinate well with collaborators from a dive

Cited by 0SourcePDFScholar
2026

Global Policy-Space Response Oracles for Two-Player Zero-Sum Games

ICML 2026poster

The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy set using deep reinforcement learning (DRL). A central challenge is to construct, under limited computational budgets, a small strategy population …

Cited by 0SourceScholar
2026

JoPPO: Hierarchical Photography Assessment via Contrastive Joint Conditional Probabilistic Reinforcement Learning

CVPR 2026

With the advancement of Vision-Language Models (VLMs), employing VLM-as-a-Judge for visual evaluation has become a widely adopted metric in vision research. However, existing VLM-as-a-Judge approaches suffer from biased scoring outcomes with low discrimination and lack the capacity for unified multi

Cited by 0SourcecodeScholar
2026

The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation

ICRA 2026poster

While scaling laws for imitation learning have primarily focused on generalization in open-world settings, the relationship between data and precision in closed-world tasks like robotic assembly remains largely unexplored. This paper systematically investigates this relationship and introduces a nov…

Cited by 0Scholar
2026

The Sword, Shield, and Achilles’ Heel: Characterizing the Linguistic Inductive Bias of Large Language Models for Spatial Reasoning in Navigation Planning

IJCAI 2026

Large Language Model (LLM)-based navigation systems have commonly constructed expli cit spatial representations (e.g., topological graphs, semantic raster maps) and translated them into textual descriptions as LLMs’ inputs. However, the linguistic structures of such text-based spatial representation

Cited by 0Scholar
2025

MsRAG: Knowledge Augumented Image Captioning with Object-level Multi-source RAG

IJCAI 2025

Language-Visual Large Models (LVLMs) have made significant strides in enhancing visual understanding capabilities. However, these models often struggle with knowledge-based visual tasks due to constrains in their pre-training data scope and timeliness. Existing Retrieval-Augmented Generation (RAG) m

Cited by 0SourcePDFScholar
2025

RCP-Bench: Benchmarking Robustness for Collaborative Perception Under Diverse Corruptions

CVPR 2025poster

Collaborative perception enhances single-vehicle perception by integrating sensory data from multiple connected vehicles. However, existing studies often assume ideal conditions, overlooking resilience to real-world challenges such as adverse weather and sensor malfunctions, which is critical for sa…

2025

Training Deep Neural Networks with Virtual Smoothing Classes

AAAI 2025technical

Learning with softmax cross-entropy on one-hot labels often leads to overconfidence on the correct class. While label smoothing regulates this overconfidence by redistributing some confidence from the correct class to other incorrect classes, it compromises the representation in the logits about the…

2024

Extending Test-Time Augmentation with Metamorphic Relations for Combinatorial Problems

ICML 2024spotlight

The application of machine learning methods to solve combinatorial problems has garnered considerable research interest. In this paper, we propose MAgg (**M**etamorphic **Agg**regation), a method to augment machine learning models for combinatorial problems at inference time using metamorphic relati…

Cited by 0SourcePDFScholar
2024

Monocular Localization with Semantics Map for Autonomous Vehicles

ICRA 2024poster

Accurate and robust localization remains a significant challenge for autonomous vehicles. The cost of sensors and limitations in local computational efficiency make it difficult to scale to large commercial applications. Traditional vision-based approaches focus on texture features that are suscepti…

Cited by 0SourceScholar
2024

Robust Communicative Multi-Agent Reinforcement Learning with Active Defense

AAAI 2024technical

Communication in multi-agent reinforcement learning (MARL) has been proven to effectively promote cooperation among agents recently. Since communication in real-world scenarios is vulnerable to noises and adversarial attacks, it is crucial to develop robust communicative MARL technique. However, exi…

Cited by 5SourcePDFScholar
2023

Promoting Cooperation in Multi-Agent Reinforcement Learning via Mutual Help

ICASSP 2023accepted

Multi-agent reinforcement learning (MARL) has achieved great progress in cooperative tasks in recent years. However, in the local reward scheme, where only local rewards for each agent are given without global rewards shared by all the agents, traditional MARL algorithms lack sufficient consideratio…

Cited by 0SourceScholar
2023

SELVO: A Semantic-Enhanced Lidar-Visual Odometry

IROS 2023poster

In the face of complex external environment, single sensor information can no longer meet the accuracy requirements of low-drift SLAM. In this paper, we focus on the fusion scheme of cameras and lidar, and explore the gain of semantic information to SLAM system. A Semantic-Enhanced Lidar-Visual Odom…

Cited by 2SourceScholar
2022

Point Cloud Change Detection With Stereo V-SLAM: Dataset, Metrics and Baseline

RA-L 2022

Localization and navigation are basic robotic tasks requiring an accurate and up-to-date map to finish these tasks, with crowdsourced data to detect map changes posing an appealing solution. Collecting and processing crowdsourced data requires low-cost sensors and algorithms, but existing methods re

Cited by 4SourcecodeScholar
2022

Pose Refinement with Joint Optimization of Visual Points and Lines

IROS 2022poster

High-precision camera re-localization technology in a pre-established 3D environment map is the basis for many tasks, such as Augmented Reality, Robotics and Autonomous Driving. The point-based visual re-localization approaches are well-developed in recent decades, but are insufficient in some featu…

Cited by 22SourceScholar
2021

Retrieval and Localization with Observation Constraints

ICRA 2021poster

Accurate visual re-localization is very critical to many artificial intelligence applications, such as augmented reality, virtual reality, robotics and autonomous driving. To accomplish this task, we propose an integrated visual re-localization method called RLOCS by combining image retrieval, seman…

Cited by 10SourceScholar
2020

Few-Shot Semantic Segmentation with Democratic Attention Networks

ECCV 2020poster

Few-shot segmentation has recently generated great popularity, addressing a challenging yet important problem of segmenting objects from unseen categories with scarce annotated support images. The crux of few-shot segmentation is to extract object information from the support image and then propagat…

Cited by 234SourcePDFScholar
2020

Stabilizing Multi-Agent Deep Reinforcement Learning by Implicitly Estimating Other Agents' Behaviors

ICASSP 2020accepted

Deep reinforcement learning (DRL) is able to learn control policies for many complicated tasks, but it's power has not been unleashed to handle multi-agent circumstances. Independent learning, where each agent treats others as part of the environment and learns its own policy without considering oth…

Cited by 0SourceScholar
2020

Vision Global Localization with Semantic Segmentation and Interest Feature Points

IROS 2020poster

In this work, we present a vision-only global localization architecture for autonomous vehicle applications, and achieves centimeter-level accuracy and high robustness in various scenarios. We first apply pixel-wise segmentation to the front-view mono camera and extract the semantic features, e.g. p…

Cited by 5SourceScholar
2019

Efficient Multi-agent Cooperative Navigation in Unknown Environments with Interlaced Deep Reinforcement Learning

ICASSP 2019accepted

This work addresses a multi-agent cooperative navigation problem that multiple agents work together in an unknown environment in order to reach different targets without collision and minimize the maximum navigation time they spend. Typical reinforcement learning-based solutions directly model the c…

Cited by 0SourceScholar
2017

Automatic radar waveform recognition based on time-frequency analysis and convolutional neural network

ICASSP 2017accepted

In this paper, we apply the idea of deep learning to radar waveform recognition. Since the frequency variation with time is the most essential distinction among radar signals with different modulation types, we transform one-dimensional radar signals into time-frequency images (TFIs) using time-freq…

Cited by 0SourceScholar