← Search

Zhenbo Xu

15 accepted papers

2025

FruitMMBench: A Multi-modal Benchmark for Fruit Quality Assessment

ICASSP 2025accepted

The rapid advancement of Large Vision-Language Models (LVLMs) has brought notable improvements in tasks like visual recognition and multi-modal understanding, demonstrating significant potential in real-world applications. However, their performances on issues related to daily life such as fruit qua…

Cited by 0SourceScholar
2025

Improving Food Recognition with Retrieval-Augmented and Domain-Adaptive LVLMs

ICASSP 2025accepted

Food recognition is pivotal in enhancing intelligent food recommendation systems and nutritional management, contributing to balanced diets and overall health. Although Large Vision-Language Models (LVLMs) have demonstrated impressive performances across various domains, their performance on the foo…

Cited by 0SourceScholar
2023

A Solution to Co-occurence Bias: Attributes Disentanglement via Mutual Information Minimization for Pedestrian Attribute Recognition

IJCAI 2023poster

Recent studies on pedestrian attribute recognition progress with either explicit or implicit modeling of the co-occurence among attributes. Considering that this known a prior is highly variable and unforeseeable regarding the specific scenarios, we show that current methods can actually suffer in g…

Cited by 9SourcePDFScholar
2023

One-Shot Neural Band Selection for Spectral Recovery

ICASSP 2023accepted

Band selection has a great impact on the spectral recovery quality. To solve this ill-posed inverse problem, most band selection methods adopt hand-crafted priors or exploit clustering or sparse regularization constraints to find most prominent bands. These methods are either very slow due to the co…

Cited by 0SourceScholar
2022

Shape Prior Guided Attack: Sparser Perturbations on 3D Point Clouds

AAAI 2022technical

Deep neural networks are extremely vulnerable to malicious input data. As 3D data is increasingly used in vision tasks such as robots, autonomous driving and drones, the internal robustness of the classification models for 3D point cloud has received widespread attention. In this paper, we propose a…

Cited by 22SourcePDFScholar
2021

Adversarial Attacks on Object Detectors with Limited Perturbations

ICASSP 2021accepted

Deep convolutional neural networks are widely witnessed vulnerable to adversarial attacks. Recently, great progress has been achieved in attacking object detectors. However, current attacks neglect the practical utility and rely on global perturbations on the target image with a large number of patc…

Cited by 0SourceScholar
2021

Continuous Copy-Paste for One-Stage Multi-Object Tracking and Segmentation

ICCV 2021poster

Current one-step multi-object tracking and segmentation (MOTS) methods lag behind recent two-step methods. By separating the instance segmentation stage from the tracking stage, two-step methods can exploit non-video datasets as extra data for training instance segmentation. Moreover, instances belo…

Cited by 29PDFcodeScholar
2021

MDANet: Multi-Modal Deep Aggregation Network for Depth Completion

ICRA 2021poster

Depth completion aims to recover the dense depth map from sparse depth data and RGB image respectively. However, due to the huge difference between the multi-modal signal input, vanilla convolutional neural network and simple fusion strategy cannot extract features from sparse data and aggregate mul…

Cited by 17SourcecodeScholar
2021

Mask4D: 4D Convolution Network for Light Field Occlusion Removal

ICASSP 2021accepted

Current light field (LF) occlusion removal approaches usually select only a part of sub-aperture images (SAIs) or simply stack all SAIs to reconstruct the center view, which destroys the spatial layout of SAIs. In this paper, we present a simple yet effective LF occlusion removal method name Mask4D,…

Cited by 0SourceScholar
2021

Pointer Networks for Arbitrary-Shaped Text Spotting

ICASSP 2021accepted

Current text spotting methods perform text detection and text recognition separately. However, in complex scenes where bounding boxes of texts with various shapes are often overlapped, text detection becomes error-prone. By contrast, character detection is more non-ambiguous and easier to learn. In…

Cited by 0SourceScholar
2021

Revealing the Reciprocal Relations Between Self-Supervised Stereo and Monocular Depth Estimation

ICCV 2021poster

Current self-supervised depth estimation algorithms mainly focus on either stereo or monocular only, neglecting the reciprocal relations between them. In this paper, we propose a simple yet effective framework to improve both stereo and monocular depth estimation by leveraging the underlying complem…

Cited by 34PDFScholar
2021

VK-Net: Category-Level Point Cloud Registration with Unsupervised Rotation Invariant Keypoints

ICASSP 2021accepted

In this paper, we propose VK-Net, a neural network that learns to discover a set of category-specific keypoints from a single point cloud in an unsupervised manner. VK-Net is able to generate semantically consistent and rotation invariant keypoints across objects of the same category and different v…

Cited by 0SourceScholar
2020

Associate-3Ddet: Perceptual-to-Conceptual Association for 3D Point Cloud Object Detection

CVPR 2020poster

Object detection from 3D point clouds remains a challenging task, though recent studies pushed the envelope with the deep learning techniques. Owing to the severe spatial occlusion and inherent variance of point density with the distance to sensors, appearance of a same object varies a lot in point…

Cited by 115PDFScholar
2020

Segment as Points for Efficient Online Multi-Object Tracking and Segmentation

ECCV 2020poster

Current multi-object tracking and segmentation (MOTS) methods follow the tracking-by-detection paradigm and adopt convolutions for feature extraction. However, as affected by the inherent receptive field, convolution based feature extraction inevitably mixes up the foreground features and the backgr…

2018

Towards End-to-End License Plate Detection and Recognition: A Large Dataset and Baseline

ECCV 2018poster

Most current license plate (LP) detection and recognition approaches are evaluated on a small and usually unrepresentative dataset since there are no publicly available large diverse datasets. In this paper, we introduce CCPD, a large and comprehensive LP dataset. All images are taken manually by wo…