← Search

Jiamao Li

21 accepted papers

2026

Energy-guided Dual Domain-invariant Prompting Framework with Fourier Regularization for Generalized Few-Shot Medical Segmentation

AAAI 2026technical

Precise segmentation of organ and tissue lesions is essential for clinical diagnosis and treatment. Despite the progress of deep learning and foundation segmentation models, their domain generalization capability remains limited particularly when dealing with cross-domain scenarios or unseen data, l

Cited by 0SourcePDFScholar
2025

$\mathbf{F}{2} \mathbf{R}{2}$: Frequency Filtering-Based Rectification Robustness Method for Stereo Matching

ICRA 2025

Most stereo matching networks assume that the stereo images are perfectly rectified, ignoring the perturbation of extrinsic parameters due to collisions, mechanical vibrations, and thermal expansion. This leads to poor rectification robustness in real-world stereo systems. That is, even minor rectif

Cited by 0SourceScholar
2025

Distributed Bundle Adjustment Based on Penalty Function Method

RA-L 2025

Bundle Adjustment (BA) aims to estimate the camera poses and build maps utilizing the nonlinear optimization algorithm. The update step of the optimization is obtained by solving a linear system, which is the bottleneck of the BA efficiency. Many works perform bundle adjustment in a distributed mann

Cited by 1SourceScholar
2025

PCGE: Boosting 3D Visual Grounding via Progressive Comprehension and Geometric-topology Perception Enhancement

IROS 2025

The 3D visual grounding task aims to establish correspondences between the 3D physical world and textual descriptions. Despite significant progress having been made, it still suffers from some challenges that need to be solved. a) Scene-agnostic text reasoning causes misaligned target region concent

Cited by 0SourceScholar
2024

BEE-Net: Bridging Semantic and Instance with Gated Encoding and Edge Constraint for Efficient Panoptic Segmentation

ICRA 2024poster

Panoptic segmentation is a challenging perception task, which can help robots to comprehensively perceive the surrounding environment. In the task, we notice that semantic, instance, and panoptic have rich relations, however, which are rarely explored. In this work, we propose a novel panoptic, inst…

Cited by 0SourceScholar
2024

CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers

ICRA 2024poster

With the increasing demands for perception accuracy in autonomous driving, there is a growing focus on fine-grained 3D semantic occupancy prediction. Effectively representing detailed three-dimensional scenes has become a significant challenge in the development of this task. In this paper, we prese…

Cited by 0SourceScholar
2024

Efficient Solution to PnP Problem Based on Vision Geometry

RA-L 2024

Perspective-n-Point (PnP) problem aims to estimate pose from known 3D map points and their projections. Efficient PnP (EPnP), one of the classical PnP solvers, represents camera pose with control points, which are easier to estimate utilizing the least square (LS) formulation. However, the geometry

Cited by 9SourceScholar
2024

Pixel-Level Precision Saccade Control of Arbitrary 3D Spatial Points in Active Binocular Vision Systems

RA-L 2024

We propose a vision-based method to achieve Pixel-Level precision saccadic movements in an active binocular vision system (ABVS). Traditional methods for precise saccadic motion rely heavily on extensive training or accurate kinematic models. However, extensive pre-training reduces the flexibility o

Cited by 0SourceScholar
2024

Self-supervised Scale Recovery for Decoupled Visual-inertial Odometry

RA-L 2024

Accurate localization for intelligent robots remains a significant challenge, and self-supervised visual-inertial odometry (VIO) has emerged as a promising solution. However, existing self-supervised VIO works consider inertial information as the ordinary data input, losing its ability to recover ab

Cited by 3SourceScholar
2023

CM-CS: Cross-Modal Common-Specific Feature Learning For Audio-Visual Video Parsing

ICASSP 2023accepted

The weakly-supervised audio-visual video parsing (AVVP) task aims to parse duration and categories of each snippet when only the video-level event labels are provided. Most methods either leverage attention mechanisms to explore cross-modal and cross-video event semantics or alleviate label noise to…

Cited by 0SourceScholar
2023

Fast Extrinsic Calibration for Multiple Inertial Measurement Units in Visual-Inertial System

ICRA 2023poster

In this paper, we propose a fast extrinsic calibration method for fusing multiple inertial measurement units (MIMU) to improve visual-inertial odometry (VIO) localization accuracy. Currently, data fusion algorithms for MIMU highly depend on the number of inertial sensors. Based on the assumption tha…

Cited by 4SourceScholar
2023

FeatDANet: Feature-level Domain Adaptation Network for Semantic Segmentation

IROS 2023poster

Unsupervised domain adaptation (UDA) is proposed to better adapt the network trained on labeled synthetic data to unlabeled real-world data for addressing the annotation cost. However, most of these methods pay more attention to domain distributions in input and output stages while ignoring the impo…

Cited by 3SourceScholar
2022

J-RR: Joint Monocular Depth Estimation and Semantic Edge Detection Exploiting Reciprocal Relations

IROS 2022poster

Depth estimation and semantic edge detection are two key tasks in computer vision, which have made great progress. To date, how to associatively predict the depth and the semantic edge is rarely explored. In this work, we first propose a flexible two-branch framework that can make the two tasks take…

Cited by 3SourceScholar
2022

Simultaneous Calibration of Multiple Revolute Joints for Articulated Vision Systems via SE(3) Kinematic Bundle Adjustment

RA-L 2022

We propose a vision-based approach to calibrate kinematic structure of low degree-of-freedom (DoF) articulated systems. Standard hand-eye calibration yields excellent eye-to-hand relations by explicitly estimating end-effector mounted camera poses from the Perspective-n-Point (PnP) problem of a sing

Cited by 7SourceScholar
2022

Spatiotemporally Enhanced Photometric Loss for Self-Supervised Monocular Depth Estimation

IROS 2022poster

Recovering depth information from a single image is a long-standing challenge, and self-supervised depth estimation methods have gradually attracted attention due to not relying on high-cost ground truth. Constructing an accurate photometric loss based on photometric consistency is crucial for these…

Cited by 7SourceScholar
2021

Camera Parameters Aware Motion Segmentation Network with Compensated Optical Flow

IROS 2021poster

Learning to distinguish independent moving objects from the observed optical flow with a moving camera remains challenging. In this work, we first present a novel camera pose compensation (CPC) scheme. With the help of ingenious geometric analysis, it breaks the observed optical flow into patterns t…

Cited by 2SourceScholar
2020

3DCFS: Fast and Robust Joint 3D Semantic-Instance Segmentation via Coupled Feature Selection

ICRA 2020poster

We propose a novel fast and robust 3D point clouds segmentation framework via coupled feature selection, named 3DCFS, that jointly performs semantic and instance segmentation. Inspired by the human scene perception process, we design a novel coupled feature selection module, named CFSM, that adaptiv…

Cited by 16SourcecodeScholar
2020

RegionNet: Region-feature-enhanced 3D Scene Understanding Network with Dual Spatial-aware Discriminative Loss

IROS 2020poster

Neural networks have recently achieved impressive success in semantic and instance segmentation on 2D images. However, their capabilities have not been fully explored to address semantic instance segmentation on unstructured 3D point cloud data. Digging into the regional feature representation to bo…

Cited by 3SourceScholar
2020

Richer Aggregated Features for Optical Flow Estimation with Edge-aware Refinement

IROS 2020poster

Recent CNN-based optical flow approaches have a separated structure of feature extraction and flow estimation. The core task of optical flow is finding the corresponding points while rich representation is just the key part of such matching problems. However, the prior work usually pays more attenti…

Cited by 1SourceScholar
2018

3D Recurrent Neural Networks with Context Fusion for Point Cloud Semantic Segmentation

ECCV 2018poster

Semantic segmentation of 3D unstructured point clouds remains an open research problem. Recent works predict semantic labels of 3D points by virtue of neural networks but take limited context knowledge into consideration. In this paper, a novel end-to-end approach for unstructured point cloud semant…

Cited by 362SourcePDFScholar