← Search

Guanghui Zhang

14 accepted papers

2025

$\mathbf{F}{2} \mathbf{R}{2}$: Frequency Filtering-Based Rectification Robustness Method for Stereo Matching

ICRA 2025

Most stereo matching networks assume that the stereo images are perfectly rectified, ignoring the perturbation of extrinsic parameters due to collisions, mechanical vibrations, and thermal expansion. This leads to poor rectification robustness in real-world stereo systems. That is, even minor rectif

Cited by 0SourceScholar
2025

CalibMutiL: Online Calibration Of LiDAR-Camera Based On Multi-level Visual Feature Fusion

IROS 2025

Multi-sensor fusion is a key technology in the field of autonomous driving and robotics. Traditional offline multi-sensor fusion calibration methods rely on manual operations and fail to meet real-time requirements, while recent online calibration technologies have limited generalization capabilitie

Cited by 0SourcecodeScholar
2025

Distributed Bundle Adjustment Based on Penalty Function Method

RA-L 2025

Bundle Adjustment (BA) aims to estimate the camera poses and build maps utilizing the nonlinear optimization algorithm. The update step of the optimization is obtained by solving a linear system, which is the bottleneck of the BA efficiency. Many works perform bundle adjustment in a distributed mann

Cited by 1SourceScholar
2024

A Deep Reinforcement Learning Approach to Balance Viewport Prediction and Video Transmission in 360° Video Streaming

IJCAI 2024poster

360° video streaming has seen tremendous growth in past years. However, our measurement reveals a dilemma that severely limits QoE. On the one hand, viewport prediction requires the shortest possible prediction distance for high predicting accuracy; On the other hand, video transmission requires mor…

2024

BEE-Net: Bridging Semantic and Instance with Gated Encoding and Edge Constraint for Efficient Panoptic Segmentation

ICRA 2024poster

Panoptic segmentation is a challenging perception task, which can help robots to comprehensively perceive the surrounding environment. In the task, we notice that semantic, instance, and panoptic have rich relations, however, which are rarely explored. In this work, we propose a novel panoptic, inst…

Cited by 0SourceScholar
2024

CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers

ICRA 2024poster

With the increasing demands for perception accuracy in autonomous driving, there is a growing focus on fine-grained 3D semantic occupancy prediction. Effectively representing detailed three-dimensional scenes has become a significant challenge in the development of this task. In this paper, we prese…

Cited by 0SourceScholar
2024

Efficient Solution to PnP Problem Based on Vision Geometry

RA-L 2024

Perspective-n-Point (PnP) problem aims to estimate pose from known 3D map points and their projections. Efficient PnP (EPnP), one of the classical PnP solvers, represents camera pose with control points, which are easier to estimate utilizing the least square (LS) formulation. However, the geometry

Cited by 9SourceScholar
2024

Self-supervised Scale Recovery for Decoupled Visual-inertial Odometry

RA-L 2024

Accurate localization for intelligent robots remains a significant challenge, and self-supervised visual-inertial odometry (VIO) has emerged as a promising solution. However, existing self-supervised VIO works consider inertial information as the ordinary data input, losing its ability to recover ab

Cited by 3SourceScholar
2024

patchDPCC: A Patchwise Deep Compression Framework for Dynamic Point Clouds

AAAI 2024technical

When compressing point clouds, point-based deep learning models operate points in a continuous space, which has a chance to minimize the geometric fidelity loss introduced by voxelization in preprocessing. But these methods could hardly scale to inputs with arbitrary points. Furthermore, the point c…

2023

CM-CS: Cross-Modal Common-Specific Feature Learning For Audio-Visual Video Parsing

ICASSP 2023accepted

The weakly-supervised audio-visual video parsing (AVVP) task aims to parse duration and categories of each snippet when only the video-level event labels are provided. Most methods either leverage attention mechanisms to explore cross-modal and cross-video event semantics or alleviate label noise to…

Cited by 0SourceScholar
2023

FeatDANet: Feature-level Domain Adaptation Network for Semantic Segmentation

IROS 2023poster

Unsupervised domain adaptation (UDA) is proposed to better adapt the network trained on labeled synthetic data to unlabeled real-world data for addressing the annotation cost. However, most of these methods pay more attention to domain distributions in input and output stages while ignoring the impo…

Cited by 3SourceScholar
2022

J-RR: Joint Monocular Depth Estimation and Semantic Edge Detection Exploiting Reciprocal Relations

IROS 2022poster

Depth estimation and semantic edge detection are two key tasks in computer vision, which have made great progress. To date, how to associatively predict the depth and the semantic edge is rarely explored. In this work, we first propose a flexible two-branch framework that can make the two tasks take…

Cited by 3SourceScholar
2022

Spatiotemporally Enhanced Photometric Loss for Self-Supervised Monocular Depth Estimation

IROS 2022poster

Recovering depth information from a single image is a long-standing challenge, and self-supervised depth estimation methods have gradually attracted attention due to not relying on high-cost ground truth. Constructing an accurate photometric loss based on photometric consistency is crucial for these…

Cited by 7SourceScholar
2020

RegionNet: Region-feature-enhanced 3D Scene Understanding Network with Dual Spatial-aware Discriminative Loss

IROS 2020poster

Neural networks have recently achieved impressive success in semantic and instance segmentation on 2D images. However, their capabilities have not been fully explored to address semantic instance segmentation on unstructured 3D point cloud data. Digging into the regional feature representation to bo…

Cited by 3SourceScholar