← Search

Xiaolin Zhang

26 accepted papers

2026

AEFGL: Reverse Auction and Value Evaluation-Based Federated Graph Learning Incentive Mechanism (Student Abstract)

AAAI 2026technical

Federated Graph Learning enables multiple clients to collaboratively train graph models while protecting local private data. However, most studies have assumed that all clients contribute data voluntarily and actively. Without reasonable incentives, clients are often reluctant to contribute personal

Cited by 0SourcePDFScholar
2025

Distributed Bundle Adjustment Based on Penalty Function Method

RA-L 2025

Bundle Adjustment (BA) aims to estimate the camera poses and build maps utilizing the nonlinear optimization algorithm. The update step of the optimization is obtained by solving a linear system, which is the bottleneck of the BA efficiency. Many works perform bundle adjustment in a distributed mann

Cited by 1SourceScholar
2025

Martin: Mobility-Aware Reputation Mechanism for Federated Learning

ICASSP 2025accepted

The rapid development of the Internet of Things (IoT) has resulted in an increasing volume of data, much of which contains sensitive private information. Federated Learning (FL) allows clients to train models without sharing raw data, showing significant potential for privacy protection. However, de…

Cited by 0SourceScholar
2025

Optimized Design and Calibration of a Human-Eye-Sized Active Binocular Vision System Based on Spherical Parallel Mechanism

RA-L 2025

The Active Binocular Vision System (ABVS), resembling the human eye, demonstrates potential for improving visual perception in robotic systems, especially in dynamic and complex environments. In this letter, we present an optimized design of a three degree-of-freedom (DoF) Active Monocular Vision Sy

Cited by 2SourceScholar
2024

BEE-Net: Bridging Semantic and Instance with Gated Encoding and Edge Constraint for Efficient Panoptic Segmentation

ICRA 2024poster

Panoptic segmentation is a challenging perception task, which can help robots to comprehensively perceive the surrounding environment. In the task, we notice that semantic, instance, and panoptic have rich relations, however, which are rarely explored. In this work, we propose a novel panoptic, inst…

Cited by 0SourceScholar
2024

CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers

ICRA 2024poster

With the increasing demands for perception accuracy in autonomous driving, there is a growing focus on fine-grained 3D semantic occupancy prediction. Effectively representing detailed three-dimensional scenes has become a significant challenge in the development of this task. In this paper, we prese…

Cited by 0SourceScholar
2024

Efficient Solution to PnP Problem Based on Vision Geometry

RA-L 2024

Perspective-n-Point (PnP) problem aims to estimate pose from known 3D map points and their projections. Efficient PnP (EPnP), one of the classical PnP solvers, represents camera pose with control points, which are easier to estimate utilizing the least square (LS) formulation. However, the geometry

Cited by 9SourceScholar
2024

Pixel-Level Precision Saccade Control of Arbitrary 3D Spatial Points in Active Binocular Vision Systems

RA-L 2024

We propose a vision-based method to achieve Pixel-Level precision saccadic movements in an active binocular vision system (ABVS). Traditional methods for precise saccadic motion rely heavily on extensive training or accurate kinematic models. However, extensive pre-training reduces the flexibility o

Cited by 0SourceScholar
2024

Self-supervised Scale Recovery for Decoupled Visual-inertial Odometry

RA-L 2024

Accurate localization for intelligent robots remains a significant challenge, and self-supervised visual-inertial odometry (VIO) has emerged as a promising solution. However, existing self-supervised VIO works consider inertial information as the ordinary data input, losing its ability to recover ab

Cited by 3SourceScholar
2023

CM-CS: Cross-Modal Common-Specific Feature Learning For Audio-Visual Video Parsing

ICASSP 2023accepted

The weakly-supervised audio-visual video parsing (AVVP) task aims to parse duration and categories of each snippet when only the video-level event labels are provided. Most methods either leverage attention mechanisms to explore cross-modal and cross-video event semantics or alleviate label noise to…

Cited by 0SourceScholar
2023

Fast Extrinsic Calibration for Multiple Inertial Measurement Units in Visual-Inertial System

ICRA 2023poster

In this paper, we propose a fast extrinsic calibration method for fusing multiple inertial measurement units (MIMU) to improve visual-inertial odometry (VIO) localization accuracy. Currently, data fusion algorithms for MIMU highly depend on the number of inertial sensors. Based on the assumption tha…

Cited by 4SourceScholar
2023

FeatDANet: Feature-level Domain Adaptation Network for Semantic Segmentation

IROS 2023poster

Unsupervised domain adaptation (UDA) is proposed to better adapt the network trained on labeled synthetic data to unlabeled real-world data for addressing the annotation cost. However, most of these methods pay more attention to domain distributions in input and output stages while ignoring the impo…

Cited by 3SourceScholar
2023

G2Pxy: Generative Open-Set Node Classification on Graphs with Proxy Unknowns

IJCAI 2023poster

Node classification is the task of predicting the labels of unlabeled nodes in a graph. State-of-the-art methods based on graph neural networks achieve excellent performance when all labels are available during training. But in real-life, models are of ten applied on data with new classes, which…

2022

J-RR: Joint Monocular Depth Estimation and Semantic Edge Detection Exploiting Reciprocal Relations

IROS 2022poster

Depth estimation and semantic edge detection are two key tasks in computer vision, which have made great progress. To date, how to associatively predict the depth and the semantic edge is rarely explored. In this work, we first propose a flexible two-branch framework that can make the two tasks take…

Cited by 3SourceScholar
2022

Simultaneous Calibration of Multiple Revolute Joints for Articulated Vision Systems via SE(3) Kinematic Bundle Adjustment

RA-L 2022

We propose a vision-based approach to calibrate kinematic structure of low degree-of-freedom (DoF) articulated systems. Standard hand-eye calibration yields excellent eye-to-hand relations by explicitly estimating end-effector mounted camera poses from the Perspective-n-Point (PnP) problem of a sing

Cited by 7SourceScholar
2022

Spatiotemporally Enhanced Photometric Loss for Self-Supervised Monocular Depth Estimation

IROS 2022poster

Recovering depth information from a single image is a long-standing challenge, and self-supervised depth estimation methods have gradually attracted attention due to not relying on high-cost ground truth. Constructing an accurate photometric loss based on photometric consistency is crucial for these…

Cited by 7SourceScholar
2021

Camera Parameters Aware Motion Segmentation Network with Compensated Optical Flow

IROS 2021poster

Learning to distinguish independent moving objects from the observed optical flow with a moving camera remains challenging. In this work, we first present a novel camera pose compensation (CPC) scheme. With the help of ingenious geometric analysis, it breaks the observed optical flow into patterns t…

Cited by 2SourceScholar
2020

3DCFS: Fast and Robust Joint 3D Semantic-Instance Segmentation via Coupled Feature Selection

ICRA 2020poster

We propose a novel fast and robust 3D point clouds segmentation framework via coupled feature selection, named 3DCFS, that jointly performs semantic and instance segmentation. Inspired by the human scene perception process, we design a novel coupled feature selection module, named CFSM, that adaptiv…

Cited by 16SourcecodeScholar
2020

RegionNet: Region-feature-enhanced 3D Scene Understanding Network with Dual Spatial-aware Discriminative Loss

IROS 2020poster

Neural networks have recently achieved impressive success in semantic and instance segmentation on 2D images. However, their capabilities have not been fully explored to address semantic instance segmentation on unstructured 3D point cloud data. Digging into the regional feature representation to bo…

Cited by 3SourceScholar
2020

Richer Aggregated Features for Optical Flow Estimation with Edge-aware Refinement

IROS 2020poster

Recent CNN-based optical flow approaches have a separated structure of feature extraction and flow estimation. The core task of optical flow is finding the corresponding points while rich representation is just the key part of such matching problems. However, the prior work usually pays more attenti…

Cited by 1SourceScholar
2019

SSF-DAN: Separated Semantic Feature Based Domain Adaptation Network for Semantic Segmentation

ICCV 2019poster

Despite the great success achieved by supervised fully convolutional models in semantic segmentation, training the models requires a large amount of labor-intensive work to generate pixel-level annotations. Recent works exploit synthetic data to train the model for semantic segmentation, but the dom…

Cited by 208PDFScholar
2018

3D Recurrent Neural Networks with Context Fusion for Point Cloud Semantic Segmentation

ECCV 2018poster

Semantic segmentation of 3D unstructured point clouds remains an open research problem. Recent works predict semantic labels of 3D points by virtue of neural networks but take limited context knowledge into consideration. In this paper, a novel end-to-end approach for unstructured point cloud semant…

Cited by 362SourcePDFScholar
2018

Adversarial Complementary Learning for Weakly Supervised Object Localization

CVPR 2018poster

In this work, we propose Adversarial Complementary Learning (ACoL) to automatically localize integral objects of semantic interest with weak supervision. We first mathematically prove that class localization maps can be obtained by directly selecting the class-specific feature maps of the last convo…

Cited by 728SourcePDFScholar
2018

Self-produced Guidance for Weakly-supervised Object Localization

ECCV 2018poster

Weakly supervised methods usually generate localization results based on attention maps produced by classification networks. However, the attention maps exhibit the most discriminative parts of the object which are small and sparse. We propose to generate Self-produced Guidance (SPG) masks which sep…